napkin-math: The Repo That Turns Guesswork Into Hardware Sense
A Rust-powered collection of benchmarks, scripts, and case studies that teaches engineers how to estimate systems in units that actually matter.
time to process 1 GiB, while smaller readouts fade into the background. The image explains the repo’s core move: translating low-level benchmark noise into a practical system-sizing unit." data-prompt="Create an editorial illustration rendered entirely in black ink on a pure white background.
The image depicts a wide laboratory workbench where an engineer calibrates a large precision gauge labeled by shape only, not text, so it clearly reads as the idea of time to process 1 GiB. Around the bench are small benchmark readouts, cache-line blocks, CPU cores arranged like mechanical parts, and a few technical instruments suggesting measurement discipline. The scene should feel like an old scientific engraving repurposed for modern systems engineering. The artist uses tightly packed crosshatching lines layered at different angles to build up shadow and form, with clean open areas of white for highlights. The line work has the quality of a classic metal engraving, precise and deliberate, with varied line weights where bold contour lines define shapes and finer interior lines create tonal depth. The overall style evokes vintage newspaper editorial illustrations from The Economist or Wall Street Journal. No color, no gradients, no grey fills. Only black lines on white. The background MUST be pure white #FFFFFF. No paper texture, no cream, no off-white, no noise, no grain. Perfectly clean flat white background." loading="lazy">
- Napkin-math turns benchmark results into a sizing language that engineers can use before a system exists.
- Its real contribution is methodological: it treats benchmarking like a controlled experiment, not a vibe check.
- The repo bridges hardware limits, runtime overhead, and real-world case studies into one practical estimation workflow.
- It sits earlier in the decision tree than formal modeling or load testing, not in competition with them.
The best thing about sirupsen/napkin-math is that it refuses the usual benchmark theater. It does not celebrate a raw number and call it insight. It asks a more useful question: what does this machine actually buy me in the units my team can plan around?
The smallest useful unit is not nanoseconds
The repo’s signature move is to collapse performance into something senior engineers can reason about fast: time to process 1 GiB. That is bigger than a CPU cycle and more honest than a per-op microbenchmark. It is close enough to the work people actually ship, from logs to scans to storage pipelines.
That framing matters because it changes the question from “is this fast?” to “is this fast enough for the shape of the job?” A system design conversation gets much better when you can say, with a straight face, that a workload will cost 12 seconds per GiB instead of waving at latency distributions and hoping everyone nods along.
Napkin math is a technique to quickly estimate the feasibility and performance of a system using a few key numbers and simple calculations. It's about being approximately right, rather than precisely wrong.
Why napkin math needs a laboratory
The repo earns its credibility by being suspicious of its own measurements. It does not trust a benchmark that was allowed to drift through page cache warmth, scheduler noise, or cross-core interference. The scripts and benchmark harnesses are there to remove excuses before they remove uncertainty.
use std::hint::black_box;
fn bench(input: &[u8]) -> usize {
let mut sum = 0usize;
for &b in input {
sum = sum.wrapping_add(black_box(b as usize));
}
black_box(sum)
}
// In the repo, this kind of guardrail matters because it prevents
// the compiler from erasing the work you think you are measuring.
That is the right instinct. Benchmarks are only trustworthy when the surrounding machine is pinned down. Cache flushing, core affinity, and careful use of black_box are not performance details. They are measurement hygiene.
Rust measures the machine, Go measures the tax
| Lens | What it tells you | Best use |
|---|---|---|
| Napkin math | How much work a system can plausibly absorb | Early sizing and architecture decisions |
| Formal performance modeling | How a queue or service behaves under assumptions | When you need rigor and can afford the setup |
| Load testing and benchmark tools | What happened in a real run | Verification, regression testing, and tuning |
The Rust side of the repo is about the machine itself. The Go experiments are about the tax you pay for convenience: pointer chasing, GC scanning, and abstraction overhead. Put together, they make a strong case that performance is never just software or hardware. It is the shape of the boundary between them.
This is where the repo stops being a cheat sheet and starts acting like a lab notebook. The message is not that microbenchmarks are useless. It is that they are easy to misunderstand unless you reduce the number of moving parts around them.
Estimation is a critical skill for senior engineers. It allows you to quickly evaluate design trade-offs and identify bottleneck before they become problems.
The case studies are the product
The newsletter archive matters because it shows the method in motion. MySQL pagination, cold-start behavior, and other system questions are not decorative examples. They are the place where the repo’s measurement habits become reusable judgment.
That makes the project feel less like documentation and more like institutional memory. The code tells you how to measure. The case studies tell you what to do with the result when the question is messy, real, and worth arguing about.
What this repo is really teaching
The deeper lesson is that senior engineers do not need more raw facts. They need calibrated intuition. napkin-math sits between formal models and casual guesses, giving you enough structure to make a decision without pretending the world is cleaner than it is.
| Approach | Strength | Blind spot |
|---|---|---|
| Napkin math | Fast, decision-oriented, grounded in hardware reality | It is approximate by design |
| Formal modeling | High rigor and repeatability | Can be expensive to build and hard to maintain |
| Load testing | Shows real behavior under pressure | Arrives late in the process |
That placement is the point. The repo is not trying to replace deeper analysis. It is trying to keep teams from starting in the wrong place. If you can estimate the shape of a problem well enough on paper, you will ask better questions when the expensive tooling finally arrives.