oMLX: The SSD-Backed Memory Tier for Apple Silicon
How a native macOS inference server uses tiered KV caching and continuous batching to eliminate the "Cold Start" problem for local agents.
- oMLX uses the Mac NVMe SSD as a secondary memory tier to store inactive KV caches and eliminate re-computation during context switching.
- The server implements a translation layer that normalizes varied API schemas from OpenAI and Anthropic into a unified internal protocol.
- Continuous batching allows the engine to process multiple concurrent requests on Apple Silicon GPUs rather than queuing them serially.
- Dynamic hardware guardrails protect system stability by automatically capping model memory usage at 90% of available RAM.
The Context Eviction Problem
The "Cold Start" is the silent killer of local AI. Apple Silicon provides massive pools of Unified Memory. However, loading three large models with a 100,000-token context window will still exhaust a 64GB machine.
When memory pressure peaks, standard local runners like Ollama are forced to evict inactive models. The KV cache (the mathematical "memory" of your ongoing conversation) is completely destroyed in the process. When you switch back to that model, the engine must re-evaluate the entire prompt from scratch. This burns compute cycles and frustrates users.
oMLX solves this by rethinking the storage hierarchy. It introduces a two-tier memory system optimized specifically for macOS. Active context lives in hot RAM. Inactive context is written to the cold SSD.
Parking the Brain: Persistent Sessions
The standout feature of oMLX is "Persistent Sessions." By leveraging the ultra-fast speeds of Mac NVMe drives, the server can park an active session to disk using the lightweight safetensors format.
This operates exactly like a save state in a video game. You can chat with a coding agent for an hour, switch to a vision model to analyze an image, and then restore the coding agent. The engine pulls the KV blocks back from the SSD into RAM. Zero tokens are re-computed. The session resumes instantly.
The Protocol Translator
Building a fast inference engine is only half the battle. The other half is speaking the fragmented languages of modern AI clients. The omlx/api/adapters directory contains a sophisticated translation layer that unifies these dialects.
Whether a client expects Anthropic XML formats or OpenAI JSON schemas, the server normalizes the request into an InternalRequest object. This decoupling allows oMLX to act as a drop-in backend for agentic tools like Claude Code or OpenClaw without requiring custom wrappers.
The adapter pattern extends to reasoning models. DeepSeek and Qwen models often inject their internal logic into the output stream using specialized thinking tags. The oMLX ThinkingParser intercepts these streams in real-time. It buffers partial tags across network packets, cleanly separating the model's internal monologue from its actual response.
Continuous Batching for One
Consumer inference runners typically process one request at a time. If a large generation is running, subsequent requests must wait in line. oMLX brings enterprise-grade scheduling to the desktop.
The core engine utilizes a BatchGenerator. This continuous batching system allows multiple concurrent requests to share compute cycles simultaneously. You can serve an LLM handling a long writing task while a lightweight reranker sorts search results in the background.
By maximizing the utilization of the Apple Silicon GPU, oMLX behaves less like a personal chatbot and more like a production-ready microservice.
The "Zero-Config" Guardrails
Running local AI models aggressively can easily lock up a computer. Mac Unified Memory is shared between the CPU and GPU. If an inference server requests too much RAM, the operating system will freeze.
oMLX includes dynamic hardware guardrails in its settings.py file. Instead of relying on static limits, the system calculates an adaptive reserve. It automatically caps model memory usage at 90% of available RAM, ensuring the OS always has enough breathing room to function smoothly.
The Local Runner Spectrum
The local LLM ecosystem offers distinct paths depending on the user's needs. oMLX targets the intersection of native performance and developer convenience, filling the gap between CLI tools and consumer apps.
| Runner | Framework | Caching Strategy | Concurrency |
|---|---|---|---|
| Ollama | llama.cpp (GGUF) | Stateless (RAM only) | Serial |
| MLX-LM | Apple MLX | Stateless (RAM only) | Serial |
| LM Studio | llama.cpp (GGUF) | Stateless (RAM only) | Serial |
| oMLX | Apple MLX | Tiered (RAM + SSD) | Continuous Batching |
As autonomous coding agents become standard desktop tools, the demands on local hardware will shift. Stateless inference engines are sufficient for simple chat queries. For persistent, long-running agent workflows, treating the SSD as an extension of the model's brain is a necessary evolution.
Sources: solatticus/omlx repository and architecture documentation.