franken_whisper: The Bayesian Hypervisor for the Agentic Ear
How a zero-unsafe Rust orchestrator unified the fragmented Whisper ecosystem into a high-reliability streaming engine for AI agents.

Agent workflows need structured, streaming, machine-readable output, not human-oriented terminal decorations that break when piped.
- The orchestrator replaces human-readable terminal output with a ten-step NDJSON streaming pipeline for autonomous agents.
- An adaptive router uses a Bayesian decision matrix to select the optimal backend based on hardware and accuracy requirements.
- The zero-unsafe Rust core prioritizes native audio decoding to minimize reliance on external system dependencies.
- Integrated SQLite persistence transforms every transcription into a durable and debuggable database transaction.
Listening for Machines
The standard approach to speech recognition treats transcription as a human-facing task. A user runs a command line tool, waits, and reads the text output on their screen. But as autonomous AI agents become the primary consumers of audio data, printing raw text to a terminal is no longer sufficient.
franken_whisper was built to solve this exact impedance mismatch. It abandons human-readable terminal decorations in favor of a strict, event-driven architecture. Every stage of its ten-step pipeline emits Newline Delimited JSON (NDJSON), providing a live telemetry stream that allows an agent to react to partial transcripts or hardware state changes in real time.
The Bayesian Decision Matrix
The most significant technical departure from standard Whisper wrappers is how franken_whisper routes audio. Instead of forcing a developer to hardcode a specific backend, the engine uses an Adaptive Backend Router.
This router treats backend selection as a probabilistic optimization problem. It evaluates the host hardware, the requested speed, and the required accuracy to dynamically choose between Python-based JAX, C++ implementations, or specialized diarization engines.

Adaptive backend routing: Bayesian decision contract selects the best engine per-request with explicit loss matrix, posterior calibration, and deterministic fallback
107,000 Lines of Certainty
Machine learning orchestration is notoriously brittle, often relying on unstable C bindings and fragile Python environments. franken_whisper counters this with a massive Rust codebase governed by a strict zero-unsafe policy.
This defensive posture extends to audio processing. Rather than defaulting to external FFmpeg binaries, the orchestrator attempts a native Rust decode using the symphonia crate. It only falls back to FFmpeg when absolutely necessary, and can even auto-provision a static FFmpeg binary if the host system lacks one.
The Memory of the Machine
Most transcription tools are ephemeral. They process audio, print the result, and exit. franken_whisper treats speech recognition as a database transaction.
By integrating a specialized SQLite implementation, every transcription run becomes a durable record. The orchestrator logs the full event trace, system parameters, and intermediate states, allowing agent developers to deterministically debug and replay audio processing failures.
| Feature | Standard CLI Tools | franken_whisper |
|---|---|---|
| Output Format | Stdout Text / GUI | Streaming NDJSON |
| Persistence | Ephemeral | SQLite Run History |
| Backend Logic | Hardcoded Engine | Bayesian Routing |
| Audio Decoding | Requires System FFmpeg | Native Rust (Symphonia fallback) |