Auditing the Autonomous Spender: Inside tomerhakak/agentprobe

How a local-first testing framework uses record-and-replay to tame infinite tool loops, catch prompt injections, and bring software engineering rigor to LLM agents.

7 min read • View on GitHub • More from tomerhakak

A vintage mechanical cash register with tangled ticker tape exploding from it, while a robotic hand tries to stamp APPROVED on the paper. This represents the financial chaos of runaway AI agents and the framework's attempt to audit them.
AgentProbe aims to tame the unpredictable financial cost of autonomous tool loops.

Cost is the silent killer. Most developers have no idea what their agents cost per run. I've seen agents that cost $5 per query because of redundant LLM calls.

Tomer Hakak, Creator of AgentProbe · DEV Community Blog
Key Takeaways

The Infinite Loop Liability

Agents are not just code. They are autonomous spenders. A standard deterministic function either works or throws an error. An LLM agent experiencing a hallucination might subtly enter an infinite loop, repeatedly calling a search API and burning through credits until a hard timeout occurs. Testing these systems requires a paradigm shift from checking correctness to auditing behavior.

Hedcut portrait of Tomer Hakak, creator of AgentProbe.

AgentProbe addresses this with an integrated X-Ray feature and a dedicated CostCalculator. It tracks input and output tokens per step, allowing developers to write direct assertions against the financial footprint of a specific workflow.

from agentprobe import record, assert_trace

@record
def my_expensive_agent(query: str):
    # Agent logic running LangChain or CrewAI
    return agent.execute(query)

# Assert the agent didn't burn the budget
assert_trace(my_expensive_agent).total_cost().is_less_than(0.05)

The Flight Data Recorder

To test an agent reliably, you must first freeze its state. AgentProbe introduces a record-and-replay architecture via the agentprobe/core/recorder.py module. By wrapping an agent function with the @record decorator, the framework intercepts all LLM prompts, tool executions, and final outputs.

These interactions are serialized into a standardized .aprobe file format. This creates a deterministic artifact from a fundamentally non-deterministic process. You can then run assertions against this trace offline, saving time and API costs.

The record-and-replay architecture intercepts API calls, storing them locally for free, deterministic test execution.

Asserting Against Chaos

Traditional unit tests rely on strict string equality. This fails instantly when an LLM decides to add a polite greeting to its JSON response. AgentProbe provides a specialized Assertion DSL designed for probabililty.

The agentprobe/core/asserter.py module uses functional matchers like _Contains and _LessThan. More importantly, it features fuzzy schema matching. It validates the structural integrity and semantic meaning of the output rather than expecting a byte-for-byte match.

A close-up of a magnifying glass held over a pool of spilled, chaotic ink. Through the lens, the ink resolves into perfectly structured geometric shapes and architectural blueprints.
The Assertion DSL brings strict schema validation to fuzzy, unpredictable LLM outputs.

Automated Paranoia

Testing for correct behavior is only half the battle. Testing for malicious behavior is where most agent deployments fail. The agentprobe/fuzz module acts as an automated red team for your local development environment.

Security is an afterthought. Most agents are vulnerable to basic prompt injection. A simple "ignore previous instructions" breaks 60% of agents I've tested.

Tomer Hakak, Creator of AgentProbe · DEV Community Blog

The framework comes pre-loaded with 55 distinct prompt injection attacks across five categories. Developers can run these fuzzing routines against their agent's defenses locally, ensuring that an instruction to drop a database table is safely ignored before the code ever reaches production.

A heavy steel bank vault door being tested by dozens of mechanical skeleton keys simultaneously, inserted into various cracks and seams.
The built-in fuzzing engine executes parallel prompt injection attacks to red-team agent defenses.

The Paradigm Shift: Trace Testing

Moving from unit testing to trace testing requires rethinking what a test actually measures. AgentProbe shifts the focus from the code itself to the execution path and the financial cost of that path.

FeatureTraditional Unit TestingAgent Trace Testing (AgentProbe)
TargetDeterministic functionsNon-deterministic probabilistic workflows
MockingManual hardcoded fixturesAuto-captured .aprobe playback files
AssertionsStrict string/value equalityFuzzy schema matching & behavioral checks
Metrics TrackedExecution time & coverageToken usage, total cost, & tool call sequences