The Math of Broken Servers: Unpacking Netflix/chaosmonkey

How an interface-driven Go binary uses statistical probability, Spinnaker, and standard Linux cron to enforce architectural resilience at global scale.

7 min read • View on GitHub • More from Netflix

A massive clockwork engine with a single severed gear tooth contained by a firewall line, while the rest of the mechanism turns smoothly. This illustrates the concept of controlled, localized failure within a functioning system.
Chaos Monkey enforces resilience by injecting controlled, mathematically probable failures into production systems.

Before we embarked on the upgrade, the Simian Army project, which houses the previous version of Chaos Monkey, had gotten to a state where it had become difficult for us to make changes to it. The internal version of the Simian Army imported the open source version and included some Netflix-specific behavior in a way that made it difficult for us to reason about the impact of changes we wanted to make.

Lorin Hochstein, Engineer at Netflix · Interview with InfoQ
Key Takeaways

The Probability of Disaster

Netflix didn't just invent chaos engineering; they institutionalized it through math. The core of Chaos Monkey isn't a wildcard script that wreaks havoc at random. It is a highly calculated, scheduled auditor.

In schedule/schedule.go, the tool uses a biased coin-flip algorithm called shouldKillInstance. This function relies on a configured "Mean Time Between Kills" to determine the probability of an instance dying on any given day. By mathematically distributing failures, unpredictable outages become a statistical certainty that engineering teams can plan around.

At our scale it is guaranteed that servers on our cloud platform will sometimes suddenly fail or disappear without warning. If we don’t have proper redundancy and automation, these disappearing servers could cause service problems.

Netflix Technology Blog · Netflix Chaos Monkey Upgraded

Cron as a Distributed Scheduler

The most surprising architectural decision in Chaos Monkey v2 is its execution engine. Instead of running a complex, stateful background daemon, it delegates scheduling entirely to the operating system.

The Go binary calculates a daily schedule and simply writes it to the local Linux crontab via the registerWithCron function. If the Chaos Monkey process crashes, the OS-level cron daemon ensures the scheduled terminations still execute flawlessly.

Delegating the Cloud to Spinnaker

Chaos Monkey contains zero AWS, GCP, or Kubernetes specific code. It achieves multi-cloud execution by heavily coupling itself to Spinnaker, Netflix's continuous delivery platform.

Chaos Monkey relies entirely on Spinnaker to map infrastructure and execute API calls to the underlying cloud providers.

The spinnaker/spinnaker.go integration acts as the bridge. Chaos Monkey queries Spinnaker's API to understand the infrastructure topology. When a cron job fires, the termination command is sent back to Spinnaker, which translates it into the appropriate cloud-specific API call.

The Safety Switch

Chaos engineering requires psychological safety. Chaos Monkey implements this through the Outage interface and a configuration parameter called Leashed.

A close-up of a heavy industrial circuit breaker switch labeled OUTAGE, locked by a thick mechanical safety pin with a dangling paper tag reading LEASHED.
The 'Leashed' mode allows teams to dry-run chaos experiments without executing actual terminations.

When Leashed is set to true, the tool selects victims and logs the intended termination without actually executing the kill command, allowing teams to dry-run their chaos. Furthermore, if the Outage interface detects an active production incident, the tool immediately halts operations to avoid compounding the crisis.

From Simian Army to Surgical Strike

The original Simian Army was a heavy Java monolith. Chaos Monkey v2 was rewritten in Go, reflecting an industry shift toward specialized, decoupled micro-tools.

FeatureChaos Monkey v1 (Simian Army)Chaos Monkey v2
LanguageJava MonolithGo Binary
Cloud AwarenessDirect AWS API callsCloud-agnostic via Spinnaker
SchedulingInternal daemon loopStandard Linux Cron
ScopeIncluded Janitor/Conformity monkeysSingle-purpose instance termination
WSJ hedcut-style portrait of Lorin Hochstein