camel-prompt-injection: CaMeL draws a hard line between data and commands

Google Research's prompt injection defense uses a quarantined LLM, capability tags, and a custom interpreter to keep untrusted inputs from steering privileged actions.

12 min read • View on GitHub • More from google-research

A wide editorial scene shows a trusted document entering a guarded airlock while a messy inbox and web page are diverted into a separate quarantine chamber. The image explains CaMeL's central move: trust is separated before data reaches the part of the system that can act on it.
CaMeL treats untrusted input like contaminated cargo, not like a sentence to be politely ignored.
Key Takeaways

Prompt injection looks like a text problem until you watch an agent act on the wrong sentence. The failure is not that the model reads hostile content. The failure is that hostile content can borrow the model's authority. CaMeL's answer is blunt: separate what the system can trust from what it can merely inspect.

In the two and a half years that we’ve been talking about prompt injection attacks I’ve seen alarmingly little progress towards a robust solution. The new paper Defeating Prompt Injections by Design from Google DeepMind finally bucks that trend. This one is worth paying attention to.

Simon Willison, Developer and Blogger · Simon Willison blog

The architecture splits trust before it splits tasks

CaMeL, short for Capability-based Model Language, comes from Google Research, Google DeepMind, and ETH Zurich. The paper treats prompt injection like a systems bug, then reaches for older security ideas: capability-based access and information flow control. The user's trusted request goes to a privileged LLM. Untrusted content, like emails or web pages, goes to a quarantined LLM that can extract facts but cannot steer tools.

CaMeL's trust boundary lives in the interpreter, not in the prompt.

The repo's src/camel/interpreter/ package is the real choke point. It parses a Python-like subset, tracks provenance on values, and intercepts operations through protocols that can deny unsafe flows. In practice that means a value is not only a string or number, it is also a trust label with readers attached. If the label is wrong for the target tool, the call stops.

def base_security_policy(args, side_effect: bool) -> bool:
    if not side_effect:
        return True
    return all(is_public(arg) for arg in args)

Upstream, the system prompt generator reflects over tool signatures and Pydantic models, then asks the model to write code against those interfaces. That turns the LLM from a free-form talker into a constrained programmer. The benefit is not elegance. It is enforcement. The model can plan, but the interpreter gets the final say.

Why filters lose

ApproachHow it decidesWhat it trustsWeak spot
Input filteringScans prompts for suspicious textThat bad intent can be detected from words aloneWrapped instructions still slip through
Dual LLM patternSeparates trusted and untrusted modelsThe boundary is mostly conceptualNo hard execution fence
CaMeLTracks provenance in a custom interpreterData carries readers and capabilitiesResearch artifact, not production ready

That contrast is the point. Early evaluations against AgentDojo suggest the design can block attacks while keeping useful behavior alive, which is a better outcome than simply making the model more suspicious of everything. CaMeL is not trying to win a text classification contest. It is trying to stop untrusted data from becoming a command.

A close-up editorial illustration shows a printed exception message being cut by a metal redaction blade before it reaches a terminal and speaker cone. It explains how CaMeL treats error messages as a possible injection path, not just a debugging aid.
Even error messages can be hostile if they echo tainted content back into the loop.

The redaction loop is the subtle move that makes the design feel serious. If an exception contains untrusted text, CaMeL does not automatically hand that text back to the model as helpful context. It redacts first, then decides what can safely be said. That closes a backchannel that many agent systems never notice until it is already exploited.

This is a research artifact released to reproduce the results in our paper. The interpreter implementation likely contains bugs (e.g., it might throw uncaught exceptions and crash) and the implementation might not be fully secure. This is **not** a Google product, and we are not planning to provide support for and/or maintain this codebase.

GitHub Repository Warning, Official Project Documentation · google-research/camel-prompt-injection README

That warning is the right frame. This is a research artifact, not a production drop. But as a blueprint, it is unusually clear: if agents are going to read hostile text and still act safely, the answer is not to decorate the prompt. It is to build a boundary the data cannot cross.