calcolatori-bench: The Benchmark That Makes LLMs Read Italian, Boot a Kernel, and Prove They Weren’t Cheating
A deep dive into the University of Pisa-style systems exam harness that turns agent evaluation into a full-stack stress test: translation, compilation, emulation, and exact-output verification.
- calcolatori-bench is less a coding benchmark than a closed exam environment that pressures an LLM across reading, planning, compiling, and verification.
- Its real innovation is the double-bind: the agent must parse Italian instructions while operating in a low-level C++ and assembly stack.
- The sandbox is not a wrapper around the task, it is the task, because Docker, QEMU, and host-side checks enforce honesty at every step.
- Exact-output grading turns brittle reasoning into visible failure, which makes the benchmark useful for studying agent reliability instead of just answer quality.
The smartest thing about calcolatori-bench is that it refuses to behave like a normal benchmark. It does not ask an LLM to solve a tidy prompt and move on. It drops the model into a constrained systems lab, feeds it Italian exam material, and then checks whether the result survives real compilation and verification.
That makes the repository feel closer to a controlled experiment than a test suite. The benchmark is designed to expose where agentic systems break: translation, task decomposition, tool use, environment handling, and exactness under pressure.
The benchmark that forces an LLM to pass a systems exam
Federico Galatolo describes it plainly in the repository itself: “A simple benchmarking tool for the 'Calcolatori Elettronici' course at the University of Pisa.” That undersells the design. In practice, the project turns a course exam into a reproducible agent-evaluation harness.
A simple benchmarking tool for the 'Calcolatori Elettronici' course at the University of Pisa.
The core idea is simple to describe and hard to execute: the model must read a problem statement, usually from a PDF, understand the course conventions, generate code for a low-level environment, and then prove the output is correct. That is a very different game from writing a function in a notebook and calling it done.
Why this is harder than a normal coding benchmark
The real trapdoor is the double-bind. The instructions are in Italian, but the implementation work spans C++, assembly, custom kernel tooling, and emulated hardware. The model has to translate meaning before it can optimize behavior.
That matters because most benchmarks isolate one skill. This one couples several together. If the model misreads the prompt, the code is wrong. If it writes the right idea in the wrong environment, the build fails. If it gets both of those right but the output is slightly off, the grader still rejects it.
How the evaluation loop actually works
The repository’s credibility comes from the loop, not from any single script. The flow is roughly: extract text from the exam PDF, inject it into the agent’s context, let the agent work inside a Docker sandbox, compile and run the result in the emulated environment, then verify the output again on the host side. That last step matters. It makes cheating much harder and makes failures more informative.
That architecture does more than test whether a model can generate code. It tests whether the model can operate in a staged environment with tooling, isolation, and verification. The benchmark measures the whole chain of action, not just the final answer.
The sandbox is the benchmark
The Docker setup is not a neutral wrapper. It is part of the problem definition. The repository’s container layer pulls in course-specific tooling such as libce and a patched qemu-ce, and it uses AUTOCORR=1 to route virtual machine output into something the harness can capture and score.
# Conceptual shape of the sandbox
export AUTOCORR=1
./evaluate.sh --harness opencode --exam exams/2025-01-27_08
# build, run, normalize, compare
That setup is why the benchmark feels more like a lab than a library. The environment is custom, the execution path is constrained, and the output path is intentionally rigid. If any link in the chain bends, the score tells you exactly where.
Why exact-output grading matters
This project does not rely on generous partial credit. The output either matches the expected result or it does not. That sounds harsh, but for systems work it is often the most honest choice.
| Approach | What it rewards | What it misses | Fit for agentic evaluation |
|---|---|---|---|
| Exact-output grading | Correct behavior under constraints | Small near-misses that may still be useful | Very high |
| Unit tests with broad tolerance | Functional correctness | Fragility in boundary cases | Medium |
| Timing-only benchmarking | Performance under load | Whether the program is actually right | Low |
| Human review | Context and judgment | Scale and reproducibility | Medium |
For a benchmark about kernel-level and emulator-backed tasks, exactness is not a bug. It is the point. If the model cannot preserve semantics all the way through the pipeline, then the benchmark has done its job.
What the leaderboard really measures
The results pipeline does more than sort submissions into pass and fail buckets. It also tracks efficiency through turn counts, which turns the benchmark into a measure of search strategy as much as correctness. A model that eventually solves the task after thrashing around is still revealing something important.
That detail changes the interpretation of the leaderboard. Success is not just whether the model found an answer. It is how many moves it needed to get there, how much environment friction it absorbed, and how well it navigated a narrow academic toolchain.
Where calcolatori-bench fits in the benchmark landscape
The nearest familiar tools are not really competitors, because they answer different questions. Criterion tests C and C++ code. Google Benchmark measures timing. Generic Python and Bash grading scripts can automate workflows. None of them are built to turn an academic systems exam into a controlled agent environment.
| Tool | Primary purpose | What it measures | What it misses | Distance from calcolatori-bench |
|---|---|---|---|---|
| Criterion | Unit testing for C and C++ | Program behavior | Agent workflow, emulation, exam context | Far |
| Google Benchmark | Microbenchmarking | Runtime performance | Correctness under constrained execution | Far |
| Custom Python or Bash scripts | Ad hoc grading | Whatever the script author encodes | Consistency, reproducibility, agent behavior | Near in shape, far in ambition |
| calcolatori-bench | Agent evaluation for a systems course | Translation, compilation, sandbox use, exact output | General-purpose portability | The reference point |
That is why the project is interesting. It is not trying to become a universal benchmark. It is making a specific kind of task legible: a multilingual, low-level, tightly verified systems exam for agents.
The bigger idea: benchmarks are becoming environments
The deeper lesson is architectural. The repo treats evaluation as a world with rules, tools, and consequences, not as a static collection of prompts. That is a useful direction for agent assessment, because real work rarely happens in a vacuum.
Seen that way, calcolatori-bench is a small but sharp proof of concept. It shows how evaluation can combine documents, sandboxes, compilers, emulators, and grading logic into one controlled loop. The more that loop resembles reality, the more honest the benchmark becomes.