Strata: The Inference Engine That Runs a 125B MoE Like a Memory Cache

How a C++ and CUDA engine turns expert routing, graph capture, and aggressive quantization into a practical way to run server-class models on consumer GPUs.

8 to 10 min read • View on GitHub • More from Niko1221

A giant mechanical library built around a rotating selector arm, with hot experts packed tightly inside a VRAM vault near the center and cooler experts shelved farther away in RAM and SSD drawers. The image explains Strata’s core idea: MoE inference works when experts behave like cache lines, not a monolithic blob.
Strata treats a 125B MoE model like a memory hierarchy, not a single block of weights.

Strata does something different. It stops thinking in layers and starts thinking in experts. ... Strata keeps the most-used experts on the GPU, in a cache. And here is the clever part. It does not guess which experts those are. It learns.

Andrew Zhu, Developer/Tech Blogger · 100 Tokens/Second on One RTX 3090 Ti
Key Takeaways

The 125B Model That Behaves Like It Fits

Strata starts from a premise that sounds impossible until you stop thinking in parameters. A 125B Mixture-of-Experts model is not one giant weight blob to cram into VRAM. It is a population of specialists, and the trick is to keep the right specialists close while everything else stays out of the hot path.

That is the article’s real inversion. Strata is not a “make the model smaller” project first. It is a project about making the model behave like a selectively resident working set, where memory residency matters as much as arithmetic.

A WSJ-style hedcut portrait of Niko1221, the maintainer of Strata, rendered in black ink on white. The portrait grounds the project in a recognizable builder and introduces the human behind the systems work.

Why Strata’s Real Unit Is the Expert, Not the Layer

Most inference engines think in layers. Strata thinks in experts. That sounds like a small vocabulary change, but it changes the whole optimization problem.

If you offload by layer, you move a lot of irrelevant weight along with the useful part. If you cache by expert, you can rank residency by actual usage, keep hot experts in VRAM, and let cold ones drift into cheaper memory tiers until they are needed again.

An interactive residency map makes the key idea visible: Strata manages experts like a cache, not a monolith.

DimensionLayer-based offloadStrata’s expert cache
Offloading unitWhole layersIndividual experts
Residency policyMostly staticRanked by usage and locality
Fast-path memoryGeneric VRAM allocationHot experts pinned close to the GPU
Cold-path handlingSwap whole blocksMove only the experts that are not needed
Main bottleneckData movement and coordinationCache misses and residency decisions

Inside the Engine Room: Arena, Poisoning, and Graph Capture

Under the hood, Strata’s core files read like a defense against waste. The memory arena pre-plans allocations instead of improvising them. That matters because a 125B-class MoE model leaves almost no room for guessing.

The poison path is the opposite of silent failure. Fill the arena with bad values, surface uninitialized access immediately, and stop pretending a broken memory plan is a valid inference path. In a constrained system, correctness is a performance feature.

Then there is graph capture. CUDA graphs let Strata record a launch sequence once and replay it, which trims CPU orchestration overhead when the token flow is repetitive. The machine spends more time computing and less time negotiating with the host.

A small but telling implementation pattern

// Conceptual shape of the engine, not exact source
CapturedGraph graph = registry.get(layer_type, n_tokens);
if (!graph.valid()) {
  graph.capture(plan, kernels);
}
graph.replay(session_state);
A close-up of a planning desk with an allocation sheet, a poisoned memory vial, and a GPU graph replay spool. The scene explains how Strata combines pre-planned memory, fail-fast debugging, and replayed execution to keep inference stable under tight constraints.
Strata’s mechanics are about preventing surprises before they become latency.

The Hidden Trick: Caching Experts Like Hot Data

The expert cache is where Strata stops looking like a model runner and starts looking like a storage system. The engine watches usage, ranks what deserves proximity, and turns memory residency into a live policy instead of a fixed layout.

That makes the whole project feel like a memory hierarchy scheduler wearing an LLM mask. The model is not simply fit into hardware. It is continuously negotiated into place.

What changesConventional local inferenceStrata
Residency signalMostly manual tuningUsage-driven expert ranking
Memory tiersVRAM plus generic spilloverVRAM, RAM, and deeper storage with expert awareness
Behavior under pressurePerformance falls off a cliffHot experts stay protected longer
User mental modelLoad model, run promptManage expert residency and reuse
Optimization targetGeneral throughputMoE-specific locality

How It Avoids the Usual Local-LLM Tax

Strata is opinionated about usability because serious systems tools usually pay a tax in setup friction. Here, the one-click install and hardware calibration do real editorial work: they hide the plumbing without hiding the mechanism.

That matters because the project is not aimed only at systems engineers. It is trying to make a difficult class of inference feel reachable on ordinary gaming hardware, with a narrow enough installation path that the user can focus on the model, not the scaffolding.

The ready-made engine now includes code for RTX 20 cards (sm_75), and setup accepts them. ... Answers in Chinese, Japanese or Korean are written faster, RTX 20 cards are supported, and five fixes.

Where Strata Differs from llama.cpp, vLLM, and the Usual Suspects

This is not a universal winner story. It is a niche story with unusually sharp edges. llama.cpp is the baseline for local inference. vLLM is the throughput-first server reference. Strata lives in a narrower, stranger lane.

ProjectPrimary strengthMemory strategyBest fit
llama.cppBroad local inference supportLayer-based offload and general-purpose flexibilityGeneral local usage and portability
vLLMServer throughput and batchingPagedAttention on data-center style hardwareHigh-concurrency serving
StrataMoE specialization on consumer GPUsExpert-granular caching and residency managementLarge MoE models on limited VRAM

The point is not that Strata replaces the others. The point is that it solves a different problem. It assumes the model is MoE-heavy, the hardware is constrained, and the real victory comes from being obsessive about what stays resident.

What This Project Suggests About the Next Wave of Local AI

Strata’s deeper lesson is that local AI progress is shifting downward into the system stack. Better kernels matter. Better quantization matters. But memory orchestration, residency policy, and CPU-GPU coordination now matter just as much.

That is why Strata is interesting beyond the benchmark screenshots. It is a proof that the next breakthrough may not be a bigger model or a cleverer prompt wrapper. It may be a better scheduler for the model you already have.