One trace, three evaluators, one failure loop
Interactive Agent Trace Evaluation Cycle
Pause
Reset
Node Info
Description text goes here.
01 / Agent Execution
02 / Trace Capture
03 / Evaluators
04 / Outcomes
05 / Revision
trace
retriever
tool hops
failure feedback loop
Agent Run
Recipe Chatbot
USER PROMPT
TOOL CALLS
FINAL ANSWER
"Add 2 cups flour..."
Trace Capture
Nested Child Spans
span: retriever_query
span: tool_hop (search->recipe)
span: answer_synthesis
llm_call & output parsing
Human Annotation
End-to-End Quality Rating
LLM-as-a-Judge
Prompt & Tool Adherence
Retrieval Metrics
Context Precision / Recall
Transition Analysis
Tool-to-Tool State Graph
PASS / FAIL
PARTIAL MISMATCH
FAIL (RETRIEVAL)
BRITTLE TRANSITION
Revision Node
System Feedback
Rubric Changes
Prompt Tuning
Tool-chain Fixes