dmtrKovalenko/oxcaml-test: OCaml, Stripped Down to the Metal

A tiny OxCaml scratchpad where the same image-diff workload is rewritten from boxed records to unboxed fields to SIMD lanes, then checked against Zig as a performance foil.

7 min read • View on GitHub • More from dmtrKovalenko

A wide workbench with the same pixel-difference pipeline repeated three times. The left side is cluttered with boxed values and heap cells, the center is reduced to bare numeric fields, and the right side streams through SIMD-like lanes in neat parallel rows. It explains that the repository is really about removing runtime overhead layer by layer.
The repo is a representation ladder disguised as a benchmark.
Key Takeaways

The same algorithm, three bodies

oxcaml-test reads like a lab notebook for one question: how much runtime machinery can you peel away before OCaml stops feeling like OCaml? The same pixel-diff workload shows up in several forms, from ordinary boxed records to unboxed fields to a batched version that leans on SIMD-shaped arithmetic. The point is not that one file wins. The point is that each file removes a different layer of friction.

ImplementationRepresentationRuntime shapeWhat it proves
odiff.mlBoxed OCaml recordsReadable, allocation-heavy baselineThe workload works before any tricks are added
odiff_optimized.mlUnboxed floats and recordsLess heap traffic, fewer tags, tighter dataOxCaml can keep hot values closer to registers
odiff_fast.mlUnboxed data plus batch processingManually shaped for wide arithmeticThe compiler can be given a much easier vectorization target
odiff_fast.zigPacked, systems-language baselineLow-friction control versionThe OCaml path can be judged against a lean reference

That is why the repository feels narrower than a benchmark suite and sharper than a demo. Each file is a note about how far a functional language can move toward systems-level performance when the data layout stops fighting the workload.

What OxCaml changes under the hood

OxCaml's value here is practical, not mystical. Unboxed floats and unboxed records let the compiler keep hot values out of heap-allocated wrappers, which cuts tagging, allocation, and garbage-collector pressure in the middle of a tight loop. In the SIMD demo inside main.ml, the same idea appears in miniature: process chunks manually, then send the leftovers down a scalar tail path. The extra code is not the story. The missing overhead is.

The useful mental model is a ladder, not a switch. Each step removes one kind of runtime friction while keeping the same core algorithm recognizable.

A close-up of a record card labeled with four numeric fields being peeled out of a cardboard box and set directly into a shallow register tray. It explains how unboxed records reduce allocation and let the compiler keep data in a tighter form.
Unboxed values keep the hot path light.

Inside the hot loop

The real pressure point is calculateBatchColorDeltas. That is where the repo stops talking about compiler theory and starts arranging arithmetic so the optimizer has a straight shot at the work: unroll the loop, keep the batch dense, and push the leftover pixels into a separate tail path. The code also carries a quiet detail that matters more than it sounds like it should. Kahan summation keeps the final result numerically honest while the loop gets faster.

type pixel = { r : float#; g : float#; b : float#; a : float# }

let lane0 = Ocaml_simd_sse.Float64x2.extract ~idx:0 vc

That shape tells the whole story. The repo is not chasing a magic flag or a clever syntax stunt. It is arranging data so the compiler has fewer excuses to box, branch, or spill.

A split illustration showing two workbenches side by side. The left side is cluttered with boxed records, scalar cleanup, and allocator scribbles, while the right side is arranged as a cleaner production line with the same algorithm flowing through a low-friction path. It explains why Zig is present as a control reference, not as the headline.
Zig is the control line, not the punchline.

Why Zig is in the same folder

Zig is not here to win a culture contest. It is the control line that tells you what the optimized OCaml path looks like when the language stops getting in its own way. odiff_fast.zig uses the same workload to make the comparison concrete. If the OCaml version gets close, that is interesting. If it does not, the gap is still useful because it shows exactly where the cost remains.

QuestionOxCaml answerZig answer
How is the pixel data represented?Unboxed fields and batched valuesPacked structs and direct layout
What happens to the hot loop?It gets shaped around what the compiler can keep in registersIt starts from a low-level, low-friction baseline
Where does the extra work go?Into explicit tail handling and loop structureInto fewer runtime abstractions overall
What is the point of the comparison?To measure how far OCaml can be pushedTo make that distance visible

The useful comparison is not ideological. It is mechanical. The repo asks whether a functional language can approach systems-language efficiency when the hottest path is written with the machine in mind.

A scratchpad with a thesis

The most interesting thing about dmtrKovalenko/oxcaml-test is how unapologetically narrow it is. It is not trying to be a polished library or a universal framework. It is trying to answer one hard question with several versions of the same algorithm: how much of the runtime can disappear before the code stops being a useful abstraction? That is a real thesis, and it is visible in the files.

The answer here is not that OCaml should become Zig. The answer is that a functional language can move much closer to metal when boxing, allocation, extraction, and branchy tails stop dominating the hot path. In this repo, the performance story is the representation story.