Skulk: The Local AI Cluster That Plans for Failure
A hard-forked Mac inference stack that pools memory, survives master loss, and keeps token flow coherent across a messy local network.
- Skulk’s real product is not a model runner but a failure-tolerant control plane for local Macs that can lose a node and keep going.
- Its core edge is operational discipline: memory-aware placement, master failover, and event-sourced state turn consumer hardware into something cluster-like.
- Cross-node speculative decoding is the flashy feature, but the harder problem is keeping draft and verify steps coherent across uneven machines and jittery networks.
- The project sits in the awkward middle between single-machine local tools and datacenter serving stacks, which is exactly why its niche matters.
Why Skulk Exists
Most local inference tools solve one of two problems. They make a single machine easier to use, or they make a datacenter cluster faster. Skulk is aimed at the gap in between: a pile of consumer Macs that can fit a large model only if the software is strict about state, placement, and recovery.
That is why the repo reads less like a model wrapper and more like an operations system. It is built to pool unified memory across heterogeneous Apple Silicon devices, expose an OpenAI-compatible API, and keep the cluster alive when one node or one coordinator falls over.
If you have multiple Macs, Exo lets you pool them into a single inference cluster. The constraint is total unified memory across devices — and the network connecting them.
The Real Trick: It Survives the Boring Failures
Skulk’s differentiator is not that it can shard a model. Plenty of projects can do something like that. The difference is that Skulk treats boring failures as first-class events: master crashes, memory pressure, abandoned requests, and network jitter are all things the system expects and shapes around.
The architecture leans on event-sourced cluster state, which means a new master can seed itself from replicated history instead of starting blind. That matters because a silent outage in a local cluster is usually worse than a loud one. Skulk prefers the loud one.
How a Request Moves Through the Cluster
The request path is split into two jobs. The control plane finds peers, decides where the model should live, and coordinates lifecycle events. The data plane moves token work fast enough, and consistently enough, that a distributed model still feels like one model.
client -> master: request
master -> workers: placement
workers -> cluster: draft tokens
cluster -> verifier: accept or reject
cluster -> log: commit state
if node fails:
elect new master from replicated state
resume placement or reroute request
That split is why the networking stack matters. The repo’s Rust layer handles peer discovery and transport concerns, while Python stays close to orchestration and model lifecycle. In other words, the system separates coordination from inference, then gives each part the tools it needs.
Speculative Decoding, But Across Nodes
Speculative decoding is simple to describe and hard to make reliable. A smaller draft model predicts a few tokens ahead. A larger target model verifies them in fewer passes than generating each token one by one.
Skulk makes that harder by doing it across nodes. Now the bottleneck is not just model quality. It is straggler control, token ordering, and ensuring that a fast machine does not run ahead of a slow one and poison the sequence.
| Layer | Single-node speculative decoding | Skulk’s cluster version |
|---|---|---|
| Drafting | One machine proposes ahead locally | A smaller engine proposes while work is coordinated across nodes |
| Verification | The target model accepts or rejects in place | Verification has to stay aligned across the cluster ring |
| Failure mode | A bad request slows one box | A slow or dropped node can wedge the whole flow unless coordination is strict |
| Payoff | Lower latency on one machine | A distributed cluster that stays coherent under uneven hardware |
That is the point of the design. Skulk is not trying to look clever by distributing tokens. It is trying to keep a distributed token stream from drifting apart.
Memory Is the Boss on Apple Silicon
On Apple Silicon, memory management is not a footnote. Unified memory makes large models possible, but it also makes admission control unforgiving. Wired pages, anonymous memory, and the compressor can turn a machine that looks free into a machine that is about to fall over.
Skulk responds with conservative placement. The master checks capacity, the worker checks again before spawn, and the cluster treats a failed fit as a normal outcome instead of a surprise. That kind of caution sounds dull until you compare it with the alternative: a Metal OOM and a reboot.
Why Rust Sits Next to Python
The repo is polyglot for a reason. Python is close to the orchestration surface, the API, and the model lifecycle. Rust is where the project goes when it needs lower-level networking, transport, and stronger control over the data plane.
That split makes the architecture legible. It also hints at the project’s priorities. Skulk is not chasing a fancy frontend first. It is trying to keep the cluster reliable enough that the frontend can be simple.
| Area | Python | Rust |
|---|---|---|
| Job | API, orchestration, lifecycle | Networking, transport, peer coordination |
| Strength | Fast iteration and control flow | Low-level reliability and predictable performance |
| Risk if overused | Turns into a scripting tangle | Turns into overengineered systems code |
| Why it fits Skulk | Keeps the cluster usable | Keeps the cluster coherent |
What Skulk Is Compared With
Skulk makes sense only once you compare it to the obvious alternatives. Single-device local tools optimize convenience. Datacenter serving stacks optimize throughput. Skulk sits in the awkward middle: a local cluster of uneven Macs that needs to behave more like infrastructure than a toy.
| System | Best at | Hardware model | Failure handling | Where Skulk differs |
|---|---|---|---|---|
| Ollama or llama.cpp style local runners | Simple single-machine setup | One box | Usually local process recovery | Skulk spans multiple Macs and treats the cluster as the unit |
| vLLM-style serving | High-throughput production inference | Datacenter GPUs | Production-grade, but for homogeneous fleets | Skulk targets ad hoc consumer hardware instead |
| Skulk | Pooling local Macs into one resilient cluster | Heterogeneous Apple Silicon | Master failover, conservative placement, replicated state | It is built around survivability before speed |
The closest sibling in spirit is exo, because Skulk comes out of that local clustering world. But Skulk’s emphasis is different. It reads like the team decided that reliability, observability, and recovery should be the headline features, not the supporting cast.
The Bet
Skulk is making a strong claim about where local AI is headed. If people are going to scale models on consumer hardware, the software has to behave like a system operator, not a demo harness. It has to understand failure, memory pressure, and the messy physics of a real network.
That is a narrower ambition than “run models locally,” but it is also a more durable one. Skulk’s real innovation is not that it makes Macs do something impressive once. It is that it makes them keep doing it after something goes wrong.