actions: Breaking the Build on Jailbreaks

How a suite of GitHub Actions brings traditional CI/CD discipline to non-deterministic LLM safety.

7 min read • View on GitHub • More from circuitbreakerlabs

An automated factory assembly line where a mechanical arm inspects gears and rejects a fractured one. This represents the automated CI/CD pipeline catching vulnerable LLM prompts before deployment.
By shifting AI red-teaming left, vulnerabilities are caught during the build process rather than in production.
Key Takeaways

Red Teaming on Every Commit

Software engineering relies on automated testing to catch regressions before they reach production. Artificial intelligence development, however, often relies on manual red-teaming phases that happen right before deployment. This creates a dangerous bottleneck. Security audits slow down feature velocity, or worse, get skipped entirely.

The circuitbreakerlabs/actions repository forces a paradigm shift. It brings standard software engineering discipline to prompt engineering and model fine-tuning. By treating LLM jailbreaks and safety violations exactly like syntax errors or broken unit tests, it ensures that vulnerable prompts simply cannot be merged.

The Anatomy of a Hard Fail

At its core, the repository provides a bridge between GitHub CI/CD workflows and an external safety evaluation API. It transforms abstract safety policies into a concrete fail-action-threshold. If a prompt's vulnerability score exceeds this threshold, the action triggers a hard system exit.

The execution pipeline of a Circuit Breaker composite action, culminating in a threshold evaluation.

This is not just an observability tool that logs warnings. It is a gatekeeper. By mapping YAML inputs to a robust Python CLI environment, the composite actions handle complex API negotiations while keeping the developer-facing configuration trivial.

Simulating the Long Game

Single-turn evaluations are straightforward, but modern jailbreaks rarely rely on a single malicious prompt. Attackers use sequence and context to slowly corner a model. Catching these sophisticated attacks requires simulating extended conversations and managing complex state.

Multi-turn evaluations detect sophisticated attacks that unfold over multiple conversation turns

Circuit Breaker Labs, Organization · Changelog - Circuit Breaker Labs GitHub Actions
A close-up of a chessboard where a robotic hand counters a human's knight move. Dotted lines trace the knight's previous moves, illustrating a multi-turn attack strategy.
Multi-turn evaluations simulate the sophisticated, multi-step strategies attackers use to bypass AI guardrails.

The multiturn-evaluate actions enforce strict architectural rules, such as requiring an even number of turns to maintain a valid user-assistant exchange pattern. This ensures the simulated attacks accurately reflect real-world adversarial behavior.

Built for the Build Pipeline

A CI step that takes ten minutes to install dependencies is useless to a fast-moving engineering team. The repository addresses this by leveraging modern Python tooling to guarantee speed and reproducibility.

The actions utilize Astral's uv, a lightning-fast Python package installer and resolver. By executing logic via uv run, the overhead of setting up the evaluation environment is reduced to milliseconds. Furthermore, the inclusion of Nix flakes ensures absolute deterministic parity across developer machines and the CI runner.

The Evaluation Ecosystem

The landscape of AI evaluation tooling is crowded, but highly specialized. While other tools focus on general quality metrics or runtime validation, Circuit Breaker Labs owns the specific niche of automated attack simulation at build time.

Featurecircuitbreakerlabs/actionsPromptfooGuardrails AI
Primary FocusAutomated attack simulationComprehensive prompt metricsRuntime output validation
Execution PhaseCI/CD Build TimeLocal / CI TestingApplication Runtime
CI IntegrationNative GitHub ActionsConfigured CLI scriptsCustom implementation
Key MechanismAPI-driven red-teamingLocal LLM judgingSchema validation

By integrating directly into the pull request workflow, these actions prevent vulnerable AI configurations from ever reaching production environments. They represent a critical maturation step in how the industry secures large language models.