dmtrKovalenko/oxcaml-test: OCaml, Stripped Down to the Metal
A tiny OxCaml scratchpad where the same image-diff workload is rewritten from boxed records to unboxed fields to SIMD lanes, then checked against Zig as a performance foil.
- This repo treats performance as a representation problem, not a language war.
- OxCaml matters here because it removes boxing and heap traffic from a hot pixel loop.
- The interesting part is not one fast file, but the ladder from ordinary OCaml to SIMD-shaped batches.
- Zig is the control line that keeps the OCaml experiment honest.
The same algorithm, three bodies
oxcaml-test reads like a lab notebook for one question: how much runtime machinery can you peel away before OCaml stops feeling like OCaml? The same pixel-diff workload shows up in several forms, from ordinary boxed records to unboxed fields to a batched version that leans on SIMD-shaped arithmetic. The point is not that one file wins. The point is that each file removes a different layer of friction.
| Implementation | Representation | Runtime shape | What it proves |
|---|---|---|---|
| odiff.ml | Boxed OCaml records | Readable, allocation-heavy baseline | The workload works before any tricks are added |
| odiff_optimized.ml | Unboxed floats and records | Less heap traffic, fewer tags, tighter data | OxCaml can keep hot values closer to registers |
| odiff_fast.ml | Unboxed data plus batch processing | Manually shaped for wide arithmetic | The compiler can be given a much easier vectorization target |
| odiff_fast.zig | Packed, systems-language baseline | Low-friction control version | The OCaml path can be judged against a lean reference |
That is why the repository feels narrower than a benchmark suite and sharper than a demo. Each file is a note about how far a functional language can move toward systems-level performance when the data layout stops fighting the workload.
What OxCaml changes under the hood
OxCaml's value here is practical, not mystical. Unboxed floats and unboxed records let the compiler keep hot values out of heap-allocated wrappers, which cuts tagging, allocation, and garbage-collector pressure in the middle of a tight loop. In the SIMD demo inside main.ml, the same idea appears in miniature: process chunks manually, then send the leftovers down a scalar tail path. The extra code is not the story. The missing overhead is.
Inside the hot loop
The real pressure point is calculateBatchColorDeltas. That is where the repo stops talking about compiler theory and starts arranging arithmetic so the optimizer has a straight shot at the work: unroll the loop, keep the batch dense, and push the leftover pixels into a separate tail path. The code also carries a quiet detail that matters more than it sounds like it should. Kahan summation keeps the final result numerically honest while the loop gets faster.
type pixel = { r : float#; g : float#; b : float#; a : float# }
let lane0 = Ocaml_simd_sse.Float64x2.extract ~idx:0 vc
That shape tells the whole story. The repo is not chasing a magic flag or a clever syntax stunt. It is arranging data so the compiler has fewer excuses to box, branch, or spill.
Why Zig is in the same folder
Zig is not here to win a culture contest. It is the control line that tells you what the optimized OCaml path looks like when the language stops getting in its own way. odiff_fast.zig uses the same workload to make the comparison concrete. If the OCaml version gets close, that is interesting. If it does not, the gap is still useful because it shows exactly where the cost remains.
| Question | OxCaml answer | Zig answer |
|---|---|---|
| How is the pixel data represented? | Unboxed fields and batched values | Packed structs and direct layout |
| What happens to the hot loop? | It gets shaped around what the compiler can keep in registers | It starts from a low-level, low-friction baseline |
| Where does the extra work go? | Into explicit tail handling and loop structure | Into fewer runtime abstractions overall |
| What is the point of the comparison? | To measure how far OCaml can be pushed | To make that distance visible |
The useful comparison is not ideological. It is mechanical. The repo asks whether a functional language can approach systems-language efficiency when the hottest path is written with the machine in mind.
A scratchpad with a thesis
The most interesting thing about dmtrKovalenko/oxcaml-test is how unapologetically narrow it is. It is not trying to be a polished library or a universal framework. It is trying to answer one hard question with several versions of the same algorithm: how much of the runtime can disappear before the code stops being a useful abstraction? That is a real thesis, and it is visible in the files.
The answer here is not that OCaml should become Zig. The answer is that a functional language can move much closer to metal when boxing, allocation, extraction, and branchy tails stop dominating the hot path. In this repo, the performance story is the representation story.