storage_ballast_helper: Predictive Self-Healing for the AI Agent Era
How a Rust-based daemon uses industrial control theory and sacrificial files to prevent disk exhaustion during high-velocity coding loops.

The problem: agent swarms and build systems can fill disks faster than humans can react, causing failed builds, stuck daemons, and crashed workflows.
- Sacrificial ballast files provide instantaneous disk relief to prevent system crashes during high-velocity AI agent workflows.
- A PID controller uses industrial control theory to predict disk exhaustion based on the acceleration of storage consumption.
- The daemon uses a deterministic scoring system and evidence ledger to ensure automated file deletion remains safe and auditable.
- A synchronous Rust architecture avoids async overhead to guarantee the monitor remains responsive during extreme I/O pressure.
The 100% Disk Death Spiral
A full disk is a terminal event for modern operating systems. When storage hits 100%, databases corrupt, logs fail to write, and recovery becomes a manual nightmare of booting from rescue media. In the past, this was a rare occurrence caused by a runaway server log. Today, the rise of AI coding agents has changed the math. An autonomous coding loop can generate ten gigabytes of temporary git clones and build artifacts in the time it takes a human developer to get a cup of coffee.
The Nautical Maneuver: Throwing Ballast Overboard
Traditional disk cleanup tools wait until a problem exists, then scan the filesystem to find things to delete. This takes time. When a disk is at 99.9% capacity, time is the one resource you do not have. The storage_ballast_helper (sbh) daemon borrows a concept from naval engineering and high-availability systems: ballast. By pre-allocating massive dummy files on the disk, the system creates an artificial buffer. When critical pressure is reached, deleting a 5GB ballast file is instantaneous, providing immediate breathing room for the system to survive while the slower, deterministic cleanup process begins.
Math as a Shield: PID Controllers and EWMA
The intelligence of sbh lies in its predictive engine. It does not simply look at total disk usage. Instead, it measures the velocity and acceleration of consumption using Exponentially Weighted Moving Averages (EWMA). This math filters out the noise of a single large file download to find the true underlying trend. A Proportional-Integral-Derivative (PID) controller then takes this velocity and calculates a pressure score. If a rogue process is rapidly allocating space, the PID controller will escalate the alert level to enforce mode long before the disk actually fills up.
The Evidence Ledger: Why Did You Delete That?
Automated deletion is inherently dangerous. To build trust, sbh relies on a deterministic scoring system and an Evidence Ledger. Every file considered for deletion is scored from 0.0 to 1.0 based on age, size, and artifact-likeness (such as living in a target directory). Crucially, the system implements hard vetoes. A directory containing a git repository or an active file handle will never be deleted, regardless of the pressure score. Every action is logged to an SQLite database, allowing developers to query exactly why a specific path was cleared.
| Maturity Level | Tool Type | Approach | Reaction Time |
|---|---|---|---|
| Level 1 (Manual) | ncdu, du | Human intervention required to find and delete large files. | Minutes to Hours |
| Level 2 (Reactive) | tmpreaper, cron | Blindly deletes files based on age or location on a fixed schedule. | Hours to Days |
| Level 3 (Predictive) | sbh | Forecasts exhaustion using control theory and uses sacrificial ballast. | Milliseconds |
Architecture: Rust Without the Async
For a system-level utility designed to run when resources are completely exhausted, runtime overhead is a liability. The author chose to build sbh in 100% Safe Rust, but deliberately avoided the async ecosystem. Instead of a heavy tokio runtime, the daemon uses raw OS threads and crossbeam channels. This synchronous, message-passing architecture guarantees predictable latency and ensures the monitoring thread is never starved by an I/O-bound scanning thread.