The Math of Broken Servers: Unpacking Netflix/chaosmonkey
How an interface-driven Go binary uses statistical probability, Spinnaker, and standard Linux cron to enforce architectural resilience at global scale.

Before we embarked on the upgrade, the Simian Army project, which houses the previous version of Chaos Monkey, had gotten to a state where it had become difficult for us to make changes to it. The internal version of the Simian Army imported the open source version and included some Netflix-specific behavior in a way that made it difficult for us to reason about the impact of changes we wanted to make.
- Chaos Monkey v2 abandons the custom daemon model, relying entirely on standard Linux cron for distributed execution reliability.
- The system uses a biased coin-flip algorithm based on Mean Time Between Kills to turn unpredictable outages into statistical certainties.
- By offloading cloud provider integrations to Spinnaker, the Go binary remains completely agnostic to underlying infrastructure like AWS or Kubernetes.
- Built-in safety valves, such as the Outage interface and Leashed mode, ensure the tool never compounds an active production incident.
The Probability of Disaster
Netflix didn't just invent chaos engineering; they institutionalized it through math. The core of Chaos Monkey isn't a wildcard script that wreaks havoc at random. It is a highly calculated, scheduled auditor.
In schedule/schedule.go, the tool uses a biased coin-flip algorithm called shouldKillInstance. This function relies on a configured "Mean Time Between Kills" to determine the probability of an instance dying on any given day. By mathematically distributing failures, unpredictable outages become a statistical certainty that engineering teams can plan around.
At our scale it is guaranteed that servers on our cloud platform will sometimes suddenly fail or disappear without warning. If we don’t have proper redundancy and automation, these disappearing servers could cause service problems.
Cron as a Distributed Scheduler
The most surprising architectural decision in Chaos Monkey v2 is its execution engine. Instead of running a complex, stateful background daemon, it delegates scheduling entirely to the operating system.
The Go binary calculates a daily schedule and simply writes it to the local Linux crontab via the registerWithCron function. If the Chaos Monkey process crashes, the OS-level cron daemon ensures the scheduled terminations still execute flawlessly.
Delegating the Cloud to Spinnaker
Chaos Monkey contains zero AWS, GCP, or Kubernetes specific code. It achieves multi-cloud execution by heavily coupling itself to Spinnaker, Netflix's continuous delivery platform.
The spinnaker/spinnaker.go integration acts as the bridge. Chaos Monkey queries Spinnaker's API to understand the infrastructure topology. When a cron job fires, the termination command is sent back to Spinnaker, which translates it into the appropriate cloud-specific API call.
The Safety Switch
Chaos engineering requires psychological safety. Chaos Monkey implements this through the Outage interface and a configuration parameter called Leashed.
When Leashed is set to true, the tool selects victims and logs the intended termination without actually executing the kill command, allowing teams to dry-run their chaos. Furthermore, if the Outage interface detects an active production incident, the tool immediately halts operations to avoid compounding the crisis.
From Simian Army to Surgical Strike
The original Simian Army was a heavy Java monolith. Chaos Monkey v2 was rewritten in Go, reflecting an industry shift toward specialized, decoupled micro-tools.
| Feature | Chaos Monkey v1 (Simian Army) | Chaos Monkey v2 |
|---|---|---|
| Language | Java Monolith | Go Binary |
| Cloud Awareness | Direct AWS API calls | Cloud-agnostic via Spinnaker |
| Scheduling | Internal daemon loop | Standard Linux Cron |
| Scope | Included Janitor/Conformity monkeys | Single-purpose instance termination |