Flash-MoE: The Apple Silicon Engine That Makes a 397B Model Behave Like It Has More Memory

A pure C and Metal inference stack for Apple Silicon that streams MoE experts from SSD, leans on macOS page cache instead of custom caching, and turns hardware constraints into the architecture.

9 min read • View on GitHub • More from danveloper

A wide editorial illustration of a laptop-like machine split into layers. Tokens flow into a compact inference pipeline on top, CPU and GPU cores work in parallel in the middle, and an SSD warehouse of expert modules sits underneath with only a few modules lifted upward at a time. The image explains that Flash-MoE treats storage as part of the runtime, not as a passive disk.
Flash-MoE turns the machine itself into a memory hierarchy for model inference.
Key Takeaways

Flash-MoE is not trying to win the generic inference framework race. It is trying to make one absurdly large MoE model behave on one very specific machine, and that choice changes everything.

The MacBook Is Not the Bottleneck. Memory Is.

The headline feat is simple to state and hard to believe: a 397B parameter model can be used on consumer Apple Silicon by refusing to keep the whole thing resident in RAM. Flash-MoE loads only the active experts for each token, then lets the rest stay on SSD until they are needed.

That reframe matters. Once you stop treating the model as something that must fit in memory, the problem becomes a systems question about overlap, latency hiding, and where each part of the machine is best used.

Expert weights (209GB at 4-bit) are read from NVMe SSD on demand via parallel pread() with GCD dispatch groups. Only the K=4 active experts per layer are loaded (~6.75MB each). The OS page cache manages caching — no custom cache needed ("Trust the OS" principle).

danveloper, Author/Maintainer · GitHub - danveloper/flash-moe

Why Flash-MoE Bets on macOS Instead of Fighting It

Most systems would put a bespoke cache between the runtime and the disk. Flash-MoE does the opposite. It trusts macOS page cache, uses ordinary pread(), and lets Darwin handle memory pressure.

That is the contrarian move at the center of the project. The repo treats the operating system as an optimization layer, not an obstacle to route around.

Flash-MoE overlaps routing, loading, and execution so the GPU does not wait on the disk.

int fd = open(model_path, O_RDONLY);
ssize_t n = pread(fd, expert_buf, expert_bytes, expert_offset);
if (n != expert_bytes) {
    perror("pread");
    return -1;
}

// The point is not custom caching.
// The point is to let the OS decide what stays hot.

The Pipeline: Projections, Routing, Experts

Inside infer.m, the model is organized around a triple-buffered command pipeline. One stage handles projections, another handles routing, and a third handles expert execution. The point is to keep those stages overlapped across layers instead of serializing them into a single wait loop.

That matters because Flash-MoE is not just moving data. It is scheduling data movement around compute so the GPU, CPU, and SSD each stay busy doing the work they are best at.

A close-up editorial illustration of a GPU kernel workbench. On one side, a nibble-packed 4-bit matrix is being unpacked. On the other, a small algebraic rearrangement turns multiply-and-add steps into a fused lane. Expert weights sit in a threadgroup buffer while reduction lines converge into a single output vector. The image explains how tiny arithmetic and memory-layout choices make the hot path fast.
The performance story lives in the kernel details, not just the architecture diagram.

Inside the Hot Path: The Metal Kernel That Does the Heavy Lifting

The sharpest code in the repo lives in shaders.metal, especially dequant_matvec_4bit_v3. Its job is straightforward to say and hard to do well: unpack 4-bit weights, apply scaling, and push the work through Metal without wasting bandwidth.

The trick is a rearranged fused multiply-add path. Instead of treating dequantization and multiply as separate steps, the kernel folds them together so the GPU does less bookkeeping and more useful arithmetic.

It also leans on SIMD-group reductions and threadgroup memory. That keeps intermediate values close to the compute units and reduces the cost of talking to global memory every time a partial sum appears.

Why This Is Not a General-Purpose Framework

Flash-MoE is deliberately narrow. It is built around one model family, one GPU backend, one storage strategy, and one hardware target. That is not a weakness. It is the reason the repo can be so aggressive.

ProjectPlatform focusModel scopeStorage strategySpecialization
Flash-MoEApple SiliconOne giant MoE model familySSD streaming with OS page cacheVery high
llama.cppBroad, cross-platformMany model familiesMostly resident memory, with experiments around offloadHigh portability
ExLlamaV2NVIDIA GPUsGPU-friendly LLMsGPU-first loading and kernelsPerformance tuned

Specialization buys simplicity in the runtime and speed in the hot path. The tradeoff is obvious: this is not the engine you pick if your first priority is portability.

How It Stacks Up Against llama.cpp and ExLlamaV2

The comparison is not about who is better in the abstract. It is about what each project is optimized to prove. llama.cpp proves breadth. ExLlamaV2 proves GPU performance on NVIDIA hardware. Flash-MoE proves that a very large MoE model can be made usable on Apple Silicon if the runtime is willing to live inside the machine's memory hierarchy.

That is a different kind of claim. It is narrower, but it is also more radical.

The Origin Story: A Human, an AI, and a 24-Hour Build

The repo's backstory fits the project's tone: fast, technical, and a little absurd in the best way. Dan Woods describes Flash-MoE as a pure C and Metal inference engine that runs Qwen3.5-397B-A17B on a MacBook Pro with 48GB of RAM at more than 4 tokens per second.

The broader lesson is not that a laptop is secretly a supercomputer. It is that a focused runtime, built around a specific model and a specific memory strategy, can punch far above what the raw specs suggest.

Pure C/Metal inference engine that runs Qwen3.5-397B-A17B (a 397 billion parameter Mixture-of-Experts model) on a MacBook Pro with 48GB RAM at 4.4+ tokens/second with production-quality output including tool calling.

danveloper, Author/Maintainer · GitHub - danveloper/flash-moe

What Flash-MoE Suggests About the Next Generation of Local AI

Flash-MoE is a clue about where local AI may be headed. The winning stack may not be a giant general-purpose framework with every backend under the sun. It may be a set of highly opinionated runtimes that understand a specific model, a specific chip, and a specific storage path.

In that world, the operating system is not just there to support the model. It is part of the model runtime.