Strata: The Inference Engine That Runs a 125B MoE Like a Memory Cache
How a C++ and CUDA engine turns expert routing, graph capture, and aggressive quantization into a practical way to run server-class models on consumer GPUs.

Strata does something different. It stops thinking in layers and starts thinking in experts. ... Strata keeps the most-used experts on the GPU, in a cache. And here is the clever part. It does not guess which experts those are. It learns.
- Strata’s breakthrough is not raw model compression. It is expert-level residency management that keeps the hottest MoE specialists close to the GPU and pushes the rest out of the way.
- The engine’s real performance story is orchestration. Memory planning, poison-mode failure checks, and CUDA graph replay cut the CPU out of the fast path.
- Strata solves a narrower problem than general local inference stacks, but it solves that problem more aggressively by treating MoE routing like a cache hierarchy.
- The project suggests that the next gains in local AI will come from scheduling and memory policy as much as from kernels and quantization.
The 125B Model That Behaves Like It Fits
Strata starts from a premise that sounds impossible until you stop thinking in parameters. A 125B Mixture-of-Experts model is not one giant weight blob to cram into VRAM. It is a population of specialists, and the trick is to keep the right specialists close while everything else stays out of the hot path.
That is the article’s real inversion. Strata is not a “make the model smaller” project first. It is a project about making the model behave like a selectively resident working set, where memory residency matters as much as arithmetic.
Why Strata’s Real Unit Is the Expert, Not the Layer
Most inference engines think in layers. Strata thinks in experts. That sounds like a small vocabulary change, but it changes the whole optimization problem.
If you offload by layer, you move a lot of irrelevant weight along with the useful part. If you cache by expert, you can rank residency by actual usage, keep hot experts in VRAM, and let cold ones drift into cheaper memory tiers until they are needed again.
| Dimension | Layer-based offload | Strata’s expert cache |
|---|---|---|
| Offloading unit | Whole layers | Individual experts |
| Residency policy | Mostly static | Ranked by usage and locality |
| Fast-path memory | Generic VRAM allocation | Hot experts pinned close to the GPU |
| Cold-path handling | Swap whole blocks | Move only the experts that are not needed |
| Main bottleneck | Data movement and coordination | Cache misses and residency decisions |
Inside the Engine Room: Arena, Poisoning, and Graph Capture
Under the hood, Strata’s core files read like a defense against waste. The memory arena pre-plans allocations instead of improvising them. That matters because a 125B-class MoE model leaves almost no room for guessing.
The poison path is the opposite of silent failure. Fill the arena with bad values, surface uninitialized access immediately, and stop pretending a broken memory plan is a valid inference path. In a constrained system, correctness is a performance feature.
Then there is graph capture. CUDA graphs let Strata record a launch sequence once and replay it, which trims CPU orchestration overhead when the token flow is repetitive. The machine spends more time computing and less time negotiating with the host.
A small but telling implementation pattern
// Conceptual shape of the engine, not exact source
CapturedGraph graph = registry.get(layer_type, n_tokens);
if (!graph.valid()) {
graph.capture(plan, kernels);
}
graph.replay(session_state);
The Hidden Trick: Caching Experts Like Hot Data
The expert cache is where Strata stops looking like a model runner and starts looking like a storage system. The engine watches usage, ranks what deserves proximity, and turns memory residency into a live policy instead of a fixed layout.
That makes the whole project feel like a memory hierarchy scheduler wearing an LLM mask. The model is not simply fit into hardware. It is continuously negotiated into place.
| What changes | Conventional local inference | Strata |
|---|---|---|
| Residency signal | Mostly manual tuning | Usage-driven expert ranking |
| Memory tiers | VRAM plus generic spillover | VRAM, RAM, and deeper storage with expert awareness |
| Behavior under pressure | Performance falls off a cliff | Hot experts stay protected longer |
| User mental model | Load model, run prompt | Manage expert residency and reuse |
| Optimization target | General throughput | MoE-specific locality |
How It Avoids the Usual Local-LLM Tax
Strata is opinionated about usability because serious systems tools usually pay a tax in setup friction. Here, the one-click install and hardware calibration do real editorial work: they hide the plumbing without hiding the mechanism.
That matters because the project is not aimed only at systems engineers. It is trying to make a difficult class of inference feel reachable on ordinary gaming hardware, with a narrow enough installation path that the user can focus on the model, not the scaffolding.

The ready-made engine now includes code for RTX 20 cards (sm_75), and setup accepts them. ... Answers in Chinese, Japanese or Korean are written faster, RTX 20 cards are supported, and five fixes.
Where Strata Differs from llama.cpp, vLLM, and the Usual Suspects
This is not a universal winner story. It is a niche story with unusually sharp edges. llama.cpp is the baseline for local inference. vLLM is the throughput-first server reference. Strata lives in a narrower, stranger lane.
| Project | Primary strength | Memory strategy | Best fit |
|---|---|---|---|
| llama.cpp | Broad local inference support | Layer-based offload and general-purpose flexibility | General local usage and portability |
| vLLM | Server throughput and batching | PagedAttention on data-center style hardware | High-concurrency serving |
| Strata | MoE specialization on consumer GPUs | Expert-granular caching and residency management | Large MoE models on limited VRAM |
The point is not that Strata replaces the others. The point is that it solves a different problem. It assumes the model is MoE-heavy, the hardware is constrained, and the real victory comes from being obsessive about what stays resident.
What This Project Suggests About the Next Wave of Local AI
Strata’s deeper lesson is that local AI progress is shifting downward into the system stack. Better kernels matter. Better quantization matters. But memory orchestration, residency policy, and CPU-GPU coordination now matter just as much.
That is why Strata is interesting beyond the benchmark screenshots. It is a proof that the next breakthrough may not be a bigger model or a cleverer prompt wrapper. It may be a better scheduler for the model you already have.