calcolatori-bench: The Benchmark That Makes LLMs Read Italian, Boot a Kernel, and Prove They Weren’t Cheating

A deep dive into the University of Pisa-style systems exam harness that turns agent evaluation into a full-stack stress test: translation, compilation, emulation, and exact-output verification.

8 min read • View on GitHub • More from galatolofederico

A wide editorial illustration of a technical gauntlet. An Italian exam PDF spills from a file cabinet on the left, a sealed Docker and QEMU chamber sits in the center, and a terminal with assembly symbols and a PASS or FAIL stamp stands behind locked glass on the right. The scene explains that the benchmark is an end-to-end exam pipeline, not a simple coding test.
The benchmark is really a controlled exam circuit. The model has to cross language, tooling, and sandbox boundaries before it can claim success.
Key Takeaways

The smartest thing about calcolatori-bench is that it refuses to behave like a normal benchmark. It does not ask an LLM to solve a tidy prompt and move on. It drops the model into a constrained systems lab, feeds it Italian exam material, and then checks whether the result survives real compilation and verification.

That makes the repository feel closer to a controlled experiment than a test suite. The benchmark is designed to expose where agentic systems break: translation, task decomposition, tool use, environment handling, and exactness under pressure.

The benchmark that forces an LLM to pass a systems exam

Federico Galatolo describes it plainly in the repository itself: “A simple benchmarking tool for the 'Calcolatori Elettronici' course at the University of Pisa.” That undersells the design. In practice, the project turns a course exam into a reproducible agent-evaluation harness.

A simple benchmarking tool for the 'Calcolatori Elettronici' course at the University of Pisa.

Federico Galatolo, Author/Maintainer · calcolatori-bench GitHub Repository Description

The core idea is simple to describe and hard to execute: the model must read a problem statement, usually from a PDF, understand the course conventions, generate code for a low-level environment, and then prove the output is correct. That is a very different game from writing a function in a notebook and calling it done.

Why this is harder than a normal coding benchmark

The real trapdoor is the double-bind. The instructions are in Italian, but the implementation work spans C++, assembly, custom kernel tooling, and emulated hardware. The model has to translate meaning before it can optimize behavior.

That matters because most benchmarks isolate one skill. This one couples several together. If the model misreads the prompt, the code is wrong. If it writes the right idea in the wrong environment, the build fails. If it gets both of those right but the output is slightly off, the grader still rejects it.

A close-up editorial illustration of a narrow bridge between an Italian PDF page and a compiler terminal. Small code tokens are threaded across the bridge, and an inspection lens checks the exact output on the far side. The image explains the benchmark's core challenge, which is translating one formal system into another without losing meaning.
The benchmark’s hardest move is not coding or translation alone. It is converting a constrained prompt into a constrained program without drifting in either direction.

How the evaluation loop actually works

The repository’s credibility comes from the loop, not from any single script. The flow is roughly: extract text from the exam PDF, inject it into the agent’s context, let the agent work inside a Docker sandbox, compile and run the result in the emulated environment, then verify the output again on the host side. That last step matters. It makes cheating much harder and makes failures more informative.

This is the project’s real architecture. The benchmark is the pipeline, and each stage can fail for a different reason.

That architecture does more than test whether a model can generate code. It tests whether the model can operate in a staged environment with tooling, isolation, and verification. The benchmark measures the whole chain of action, not just the final answer.

The sandbox is the benchmark

The Docker setup is not a neutral wrapper. It is part of the problem definition. The repository’s container layer pulls in course-specific tooling such as libce and a patched qemu-ce, and it uses AUTOCORR=1 to route virtual machine output into something the harness can capture and score.

# Conceptual shape of the sandbox
export AUTOCORR=1
./evaluate.sh --harness opencode --exam exams/2025-01-27_08
# build, run, normalize, compare

That setup is why the benchmark feels more like a lab than a library. The environment is custom, the execution path is constrained, and the output path is intentionally rigid. If any link in the chain bends, the score tells you exactly where.

Why exact-output grading matters

This project does not rely on generous partial credit. The output either matches the expected result or it does not. That sounds harsh, but for systems work it is often the most honest choice.

ApproachWhat it rewardsWhat it missesFit for agentic evaluation
Exact-output gradingCorrect behavior under constraintsSmall near-misses that may still be usefulVery high
Unit tests with broad toleranceFunctional correctnessFragility in boundary casesMedium
Timing-only benchmarkingPerformance under loadWhether the program is actually rightLow
Human reviewContext and judgmentScale and reproducibilityMedium

For a benchmark about kernel-level and emulator-backed tasks, exactness is not a bug. It is the point. If the model cannot preserve semantics all the way through the pipeline, then the benchmark has done its job.

What the leaderboard really measures

The results pipeline does more than sort submissions into pass and fail buckets. It also tracks efficiency through turn counts, which turns the benchmark into a measure of search strategy as much as correctness. A model that eventually solves the task after thrashing around is still revealing something important.

That detail changes the interpretation of the leaderboard. Success is not just whether the model found an answer. It is how many moves it needed to get there, how much environment friction it absorbed, and how well it navigated a narrow academic toolchain.

Where calcolatori-bench fits in the benchmark landscape

The nearest familiar tools are not really competitors, because they answer different questions. Criterion tests C and C++ code. Google Benchmark measures timing. Generic Python and Bash grading scripts can automate workflows. None of them are built to turn an academic systems exam into a controlled agent environment.

ToolPrimary purposeWhat it measuresWhat it missesDistance from calcolatori-bench
CriterionUnit testing for C and C++Program behaviorAgent workflow, emulation, exam contextFar
Google BenchmarkMicrobenchmarkingRuntime performanceCorrectness under constrained executionFar
Custom Python or Bash scriptsAd hoc gradingWhatever the script author encodesConsistency, reproducibility, agent behaviorNear in shape, far in ambition
calcolatori-benchAgent evaluation for a systems courseTranslation, compilation, sandbox use, exact outputGeneral-purpose portabilityThe reference point

That is why the project is interesting. It is not trying to become a universal benchmark. It is making a specific kind of task legible: a multilingual, low-level, tightly verified systems exam for agents.

The bigger idea: benchmarks are becoming environments

The deeper lesson is architectural. The repo treats evaluation as a world with rules, tools, and consequences, not as a static collection of prompts. That is a useful direction for agent assessment, because real work rarely happens in a vacuum.

Seen that way, calcolatori-bench is a small but sharp proof of concept. It shows how evaluation can combine documents, sandboxes, compilers, emulators, and grading logic into one controlled loop. The more that loop resembles reality, the more honest the benchmark becomes.