Auditing the Autonomous Spender: Inside tomerhakak/agentprobe
How a local-first testing framework uses record-and-replay to tame infinite tool loops, catch prompt injections, and bring software engineering rigor to LLM agents.

Cost is the silent killer. Most developers have no idea what their agents cost per run. I've seen agents that cost $5 per query because of redundant LLM calls.
- AgentProbe treats AI agents as financial actors, introducing strict cost assertions to prevent runaway API billing during infinite tool loops.
- The framework's record-and-replay architecture captures non-deterministic agent runs into deterministic .aprobe traces for local testing.
- A specialized testing DSL replaces strict string equality with fuzzy schema matching and behavioral validators suited for LLM outputs.
- Built-in fuzzing executes 55 parallel prompt injection attacks locally to red-team agent defenses before production.
The Infinite Loop Liability
Agents are not just code. They are autonomous spenders. A standard deterministic function either works or throws an error. An LLM agent experiencing a hallucination might subtly enter an infinite loop, repeatedly calling a search API and burning through credits until a hard timeout occurs. Testing these systems requires a paradigm shift from checking correctness to auditing behavior.
AgentProbe addresses this with an integrated X-Ray feature and a dedicated CostCalculator. It tracks input and output tokens per step, allowing developers to write direct assertions against the financial footprint of a specific workflow.
from agentprobe import record, assert_trace
@record
def my_expensive_agent(query: str):
# Agent logic running LangChain or CrewAI
return agent.execute(query)
# Assert the agent didn't burn the budget
assert_trace(my_expensive_agent).total_cost().is_less_than(0.05)
The Flight Data Recorder
To test an agent reliably, you must first freeze its state. AgentProbe introduces a record-and-replay architecture via the agentprobe/core/recorder.py module. By wrapping an agent function with the @record decorator, the framework intercepts all LLM prompts, tool executions, and final outputs.
These interactions are serialized into a standardized .aprobe file format. This creates a deterministic artifact from a fundamentally non-deterministic process. You can then run assertions against this trace offline, saving time and API costs.
Asserting Against Chaos
Traditional unit tests rely on strict string equality. This fails instantly when an LLM decides to add a polite greeting to its JSON response. AgentProbe provides a specialized Assertion DSL designed for probabililty.
The agentprobe/core/asserter.py module uses functional matchers like _Contains and _LessThan. More importantly, it features fuzzy schema matching. It validates the structural integrity and semantic meaning of the output rather than expecting a byte-for-byte match.
Automated Paranoia
Testing for correct behavior is only half the battle. Testing for malicious behavior is where most agent deployments fail. The agentprobe/fuzz module acts as an automated red team for your local development environment.

Security is an afterthought. Most agents are vulnerable to basic prompt injection. A simple "ignore previous instructions" breaks 60% of agents I've tested.
The framework comes pre-loaded with 55 distinct prompt injection attacks across five categories. Developers can run these fuzzing routines against their agent's defenses locally, ensuring that an instruction to drop a database table is safely ignored before the code ever reaches production.
The Paradigm Shift: Trace Testing
Moving from unit testing to trace testing requires rethinking what a test actually measures. AgentProbe shifts the focus from the code itself to the execution path and the financial cost of that path.
| Feature | Traditional Unit Testing | Agent Trace Testing (AgentProbe) |
|---|---|---|
| Target | Deterministic functions | Non-deterministic probabilistic workflows |
| Mocking | Manual hardcoded fixtures | Auto-captured .aprobe playback files |
| Assertions | Strict string/value equality | Fuzzy schema matching & behavioral checks |
| Metrics Tracked | Execution time & coverage | Token usage, total cost, & tool call sequences |