oMLX: The SSD-Backed Memory Tier for Apple Silicon

How a native macOS inference server uses tiered KV caching and continuous batching to eliminate the "Cold Start" problem for local agents.

8 min read • View on GitHub • More from solatticus

A vintage library where a librarian is sliding a massive, glowing drawer out of a stone wall and onto a polished reading desk. The image represents the architectural concept of tiered memory, moving saved context from a cold SSD to hot RAM.
oMLX treats the Mac's fast NVMe SSD as a cold storage tier for LLM context, allowing massive sessions to be restored instantly without re-computation.
Key Takeaways

The Context Eviction Problem

The "Cold Start" is the silent killer of local AI. Apple Silicon provides massive pools of Unified Memory. However, loading three large models with a 100,000-token context window will still exhaust a 64GB machine.

When memory pressure peaks, standard local runners like Ollama are forced to evict inactive models. The KV cache (the mathematical "memory" of your ongoing conversation) is completely destroyed in the process. When you switch back to that model, the engine must re-evaluate the entire prompt from scratch. This burns compute cycles and frustrates users.

oMLX solves this by rethinking the storage hierarchy. It introduces a two-tier memory system optimized specifically for macOS. Active context lives in hot RAM. Inactive context is written to the cold SSD.

Parking the Brain: Persistent Sessions

The standout feature of oMLX is "Persistent Sessions." By leveraging the ultra-fast speeds of Mac NVMe drives, the server can park an active session to disk using the lightweight safetensors format.

This operates exactly like a save state in a video game. You can chat with a coding agent for an hour, switch to a vision model to analyze an image, and then restore the coding agent. The engine pulls the KV blocks back from the SSD into RAM. Zero tokens are re-computed. The session resumes instantly.

A two-tier architecture diagram showing the Memory Tier Toggle. On the left

The Protocol Translator

Building a fast inference engine is only half the battle. The other half is speaking the fragmented languages of modern AI clients. The omlx/api/adapters directory contains a sophisticated translation layer that unifies these dialects.

Whether a client expects Anthropic XML formats or OpenAI JSON schemas, the server normalizes the request into an InternalRequest object. This decoupling allows oMLX to act as a drop-in backend for agentic tools like Claude Code or OpenClaw without requiring custom wrappers.

The adapter pattern extends to reasoning models. DeepSeek and Qwen models often inject their internal logic into the output stream using specialized thinking tags. The oMLX ThinkingParser intercepts these streams in real-time. It buffers partial tags across network packets, cleanly separating the model's internal monologue from its actual response.

A complex mechanical funnel sorting jagged shapes into uniform spheres. This illustrates the adapter layer translating fragmented LLM API schemas into a unified internal protocol.
The API adapter layer buffers and filters incoming streams, ensuring reasoning tags and tool calls are parsed flawlessly before reaching the client.

Continuous Batching for One

Consumer inference runners typically process one request at a time. If a large generation is running, subsequent requests must wait in line. oMLX brings enterprise-grade scheduling to the desktop.

The core engine utilizes a BatchGenerator. This continuous batching system allows multiple concurrent requests to share compute cycles simultaneously. You can serve an LLM handling a long writing task while a lightweight reranker sorts search results in the background.

By maximizing the utilization of the Apple Silicon GPU, oMLX behaves less like a personal chatbot and more like a production-ready microservice.

A request pipeline diagram illustrating the stateful Thinking Stream. On the left

The "Zero-Config" Guardrails

Running local AI models aggressively can easily lock up a computer. Mac Unified Memory is shared between the CPU and GPU. If an inference server requests too much RAM, the operating system will freeze.

oMLX includes dynamic hardware guardrails in its settings.py file. Instead of relying on static limits, the system calculates an adaptive reserve. It automatically caps model memory usage at 90% of available RAM, ensuring the OS always has enough breathing room to function smoothly.

A close-up of a mechanical pressure gauge with the needle held back by a sturdy iron bolt. This represents the dynamic memory guardrails preventing system crashes.
Dynamic memory calculations ensure the server never requests more RAM than the operating system can safely provide.

The Local Runner Spectrum

The local LLM ecosystem offers distinct paths depending on the user's needs. oMLX targets the intersection of native performance and developer convenience, filling the gap between CLI tools and consumer apps.

Runner Framework Caching Strategy Concurrency
Ollama llama.cpp (GGUF) Stateless (RAM only) Serial
MLX-LM Apple MLX Stateless (RAM only) Serial
LM Studio llama.cpp (GGUF) Stateless (RAM only) Serial
oMLX Apple MLX Tiered (RAM + SSD) Continuous Batching

As autonomous coding agents become standard desktop tools, the demands on local hardware will shift. Stateless inference engines are sufficient for simple chat queries. For persistent, long-running agent workflows, treating the SSD as an extension of the model's brain is a necessary evolution.


Sources: solatticus/omlx repository and architecture documentation.