napkin-math: The Repo That Turns Guesswork Into Hardware Sense

A Rust-powered collection of benchmarks, scripts, and case studies that teaches engineers how to estimate systems in units that actually matter.

8 to 10 min read • View on GitHub • More from sirupsen

An engineer stands at a workbench surrounded by measuring tools, CPU cores, cache lines, and lab instruments. One large gauge dominates the scene, marked as <span class=time to process 1 GiB, while smaller readouts fade into the background. The image explains the repo’s core move: translating low-level benchmark noise into a practical system-sizing unit." data-prompt="Create an editorial illustration rendered entirely in black ink on a pure white background. The image depicts a wide laboratory workbench where an engineer calibrates a large precision gauge labeled by shape only, not text, so it clearly reads as the idea of time to process 1 GiB. Around the bench are small benchmark readouts, cache-line blocks, CPU cores arranged like mechanical parts, and a few technical instruments suggesting measurement discipline. The scene should feel like an old scientific engraving repurposed for modern systems engineering. The artist uses tightly packed crosshatching lines layered at different angles to build up shadow and form, with clean open areas of white for highlights. The line work has the quality of a classic metal engraving, precise and deliberate, with varied line weights where bold contour lines define shapes and finer interior lines create tonal depth. The overall style evokes vintage newspaper editorial illustrations from The Economist or Wall Street Journal. No color, no gradients, no grey fills. Only black lines on white. The background MUST be pure white #FFFFFF. No paper texture, no cream, no off-white, no noise, no grain. Perfectly clean flat white background." loading="lazy">
The repo’s real trick is not faster code. It is a better unit of thought.
Key Takeaways

The best thing about sirupsen/napkin-math is that it refuses the usual benchmark theater. It does not celebrate a raw number and call it insight. It asks a more useful question: what does this machine actually buy me in the units my team can plan around?

The smallest useful unit is not nanoseconds

The repo’s signature move is to collapse performance into something senior engineers can reason about fast: time to process 1 GiB. That is bigger than a CPU cycle and more honest than a per-op microbenchmark. It is close enough to the work people actually ship, from logs to scans to storage pipelines.

A useful benchmark is not just a run. It is a controlled environment that makes the number meaningful.

That framing matters because it changes the question from “is this fast?” to “is this fast enough for the shape of the job?” A system design conversation gets much better when you can say, with a straight face, that a workload will cost 12 seconds per GiB instead of waving at latency distributions and hoping everyone nods along.

Napkin math is a technique to quickly estimate the feasibility and performance of a system using a few key numbers and simple calculations. It's about being approximately right, rather than precisely wrong.

Simon Eskildsen, Author · Napkin Math

Why napkin math needs a laboratory

The repo earns its credibility by being suspicious of its own measurements. It does not trust a benchmark that was allowed to drift through page cache warmth, scheduler noise, or cross-core interference. The scripts and benchmark harnesses are there to remove excuses before they remove uncertainty.

use std::hint::black_box;

fn bench(input: &[u8]) -> usize {
    let mut sum = 0usize;
    for &b in input {
        sum = sum.wrapping_add(black_box(b as usize));
    }
    black_box(sum)
}

// In the repo, this kind of guardrail matters because it prevents
// the compiler from erasing the work you think you are measuring.

That is the right instinct. Benchmarks are only trustworthy when the surrounding machine is pinned down. Cache flushing, core affinity, and careful use of black_box are not performance details. They are measurement hygiene.

Rust measures the machine, Go measures the tax

LensWhat it tells youBest use
Napkin mathHow much work a system can plausibly absorbEarly sizing and architecture decisions
Formal performance modelingHow a queue or service behaves under assumptionsWhen you need rigor and can afford the setup
Load testing and benchmark toolsWhat happened in a real runVerification, regression testing, and tuning

The Rust side of the repo is about the machine itself. The Go experiments are about the tax you pay for convenience: pointer chasing, GC scanning, and abstraction overhead. Put together, they make a strong case that performance is never just software or hardware. It is the shape of the boundary between them.

A split-scene editorial illustration contrasts two benchmark environments. On the left, fans spin, doors are open, and results scatter into noise. On the right, the machine is sealed, cores are pinned, cache is flushed, and one stable number emerges from the experiment. The image explains why environmental control is the difference between a guess and a usable measurement.
The same code can yield two very different truths depending on how carefully the environment is controlled.

This is where the repo stops being a cheat sheet and starts acting like a lab notebook. The message is not that microbenchmarks are useless. It is that they are easy to misunderstand unless you reduce the number of moving parts around them.

Estimation is a critical skill for senior engineers. It allows you to quickly evaluate design trade-offs and identify bottleneck before they become problems.

Simon Eskildsen, Author · Estimation

The case studies are the product

The newsletter archive matters because it shows the method in motion. MySQL pagination, cold-start behavior, and other system questions are not decorative examples. They are the place where the repo’s measurement habits become reusable judgment.

That makes the project feel less like documentation and more like institutional memory. The code tells you how to measure. The case studies tell you what to do with the result when the question is messy, real, and worth arguing about.

What this repo is really teaching

The deeper lesson is that senior engineers do not need more raw facts. They need calibrated intuition. napkin-math sits between formal models and casual guesses, giving you enough structure to make a decision without pretending the world is cleaner than it is.

ApproachStrengthBlind spot
Napkin mathFast, decision-oriented, grounded in hardware realityIt is approximate by design
Formal modelingHigh rigor and repeatabilityCan be expensive to build and hard to maintain
Load testingShows real behavior under pressureArrives late in the process

That placement is the point. The repo is not trying to replace deeper analysis. It is trying to keep teams from starting in the wrong place. If you can estimate the shape of a problem well enough on paper, you will ask better questions when the expensive tooling finally arrives.