camel-prompt-injection: CaMeL draws a hard line between data and commands
Google Research's prompt injection defense uses a quarantined LLM, capability tags, and a custom interpreter to keep untrusted inputs from steering privileged actions.
- CaMeL does not try to make the model spot bad instructions. It changes the execution path so tainted data cannot become authority.
- The repository's custom interpreter turns provenance into a security primitive, which lets the system block unsafe tool calls at runtime.
- The quarantined LLM and redacted error loop close two easy backchannels in agent systems: hostile content and helpful error messages.
- The project is a research artifact, but it is a clear blueprint for moving agent security from prompt hygiene to enforced information flow control.
Prompt injection looks like a text problem until you watch an agent act on the wrong sentence. The failure is not that the model reads hostile content. The failure is that hostile content can borrow the model's authority. CaMeL's answer is blunt: separate what the system can trust from what it can merely inspect.
In the two and a half years that we’ve been talking about prompt injection attacks I’ve seen alarmingly little progress towards a robust solution. The new paper Defeating Prompt Injections by Design from Google DeepMind finally bucks that trend. This one is worth paying attention to.
The architecture splits trust before it splits tasks
CaMeL, short for Capability-based Model Language, comes from Google Research, Google DeepMind, and ETH Zurich. The paper treats prompt injection like a systems bug, then reaches for older security ideas: capability-based access and information flow control. The user's trusted request goes to a privileged LLM. Untrusted content, like emails or web pages, goes to a quarantined LLM that can extract facts but cannot steer tools.
The repo's src/camel/interpreter/ package is the real choke point. It parses a Python-like subset, tracks provenance on values, and intercepts operations through protocols that can deny unsafe flows. In practice that means a value is not only a string or number, it is also a trust label with readers attached. If the label is wrong for the target tool, the call stops.
def base_security_policy(args, side_effect: bool) -> bool:
if not side_effect:
return True
return all(is_public(arg) for arg in args)
Upstream, the system prompt generator reflects over tool signatures and Pydantic models, then asks the model to write code against those interfaces. That turns the LLM from a free-form talker into a constrained programmer. The benefit is not elegance. It is enforcement. The model can plan, but the interpreter gets the final say.
Why filters lose
| Approach | How it decides | What it trusts | Weak spot |
|---|---|---|---|
| Input filtering | Scans prompts for suspicious text | That bad intent can be detected from words alone | Wrapped instructions still slip through |
| Dual LLM pattern | Separates trusted and untrusted models | The boundary is mostly conceptual | No hard execution fence |
| CaMeL | Tracks provenance in a custom interpreter | Data carries readers and capabilities | Research artifact, not production ready |
That contrast is the point. Early evaluations against AgentDojo suggest the design can block attacks while keeping useful behavior alive, which is a better outcome than simply making the model more suspicious of everything. CaMeL is not trying to win a text classification contest. It is trying to stop untrusted data from becoming a command.
The redaction loop is the subtle move that makes the design feel serious. If an exception contains untrusted text, CaMeL does not automatically hand that text back to the model as helpful context. It redacts first, then decides what can safely be said. That closes a backchannel that many agent systems never notice until it is already exploited.
This is a research artifact released to reproduce the results in our paper. The interpreter implementation likely contains bugs (e.g., it might throw uncaught exceptions and crash) and the implementation might not be fully secure. This is **not** a Google product, and we are not planning to provide support for and/or maintain this codebase.
That warning is the right frame. This is a research artifact, not a production drop. But as a blueprint, it is unusually clear: if agents are going to read hostile text and still act safely, the answer is not to decorate the prompt. It is to build a boundary the data cannot cross.