circuitbreakerlabs/cli: The Automated Interrogation Room for LLMs

How a Rust-based man-in-the-middle proxy uses WebSockets and Rhai scripting to run adaptive, multi-turn adversarial attacks in CI/CD.

6 min read • View on GitHub • More from circuitbreakerlabs

A mechanical arm plays chess against a black box, connected by a single wire. This represents the automated attacker probing the target model.
The CLI acts as the battleground where an automated attacker probes a black-box model.
Key Takeaways

The Multi-Turn Battleground

Static benchmarks fail to catch jailbreaks that unfold over a conversation. The Circuit Breaker Labs CLI is not a static tester. It functions as an active proxy. It facilitates a live, multi-turn interaction where an attacker model adapts its prompts based on the target model's previous defenses.

A Man-in-the-Middle Built in Rust

The core engine uses a trait-based strategy pattern to manage complex async loops. It simultaneously listens for incoming WebSocket messages from the remote orchestrator and spawns tokio tasks to fetch responses from the defending LLM.

The Adaptive Attack Loop: The CLI sits in the middle, translating WebSocket streams into HTTP requests.

The Universal Translator: Rhai Scripting

The most surprising technical choice is how the CLI handles custom model integrations. Instead of Python, it uses Rhai, a lightweight Rust scripting language. Rhai scripts map internal protocol types to any proprietary JSON schema on the fly, keeping the main binary dependency-free.

A mechanical magnifying glass over a machine-code Rosetta stone reveals meshing gears. This illustrates Rhai scripting acting as a universal translation layer.
Rhai scripts allow developers to map custom APIs without touching the core Rust binary.

Safety-as-Code for the Pipeline

Operational realities require robust tooling. The CLI uses ratatui for real-time terminal progress tracking and strict error handling to prevent network stutters from crashing large evaluation runs.

Circuit Breaker Labs CLI (`cbl`) is a command-line tool for running comprehensive AI safety evaluations on language models. It helps you test your AI systems against unsafe prompts and adversarial scenarios to ensure they respond safely and appropriately.

Circuit Breaker Labs CLI Documentation, Official Project Documentation · Introduction - Circuit Breaker Labs CLI

Managed Adversaries vs. Static Scanners

The system acts as a client for a managed safety service rather than a standalone scanner. This shifts the burden of generating robust adversarial test cases away from the user.

Featurecircuitbreakerlabs/cliGarakPromptfoo
Attack GenerationAdaptive APILocal StaticUser-defined
ExtensibilityRhai scriptsPython modulesYAML/JS
Multi-turn SupportNativeLimitedManual
DeploymentSingle BinaryPip InstallNPM