One trace, three evaluators, one failure loop

Interactive Agent Trace Evaluation Cycle

Node Info
Description text goes here.
01 / Agent Execution 02 / Trace Capture 03 / Evaluators 04 / Outcomes 05 / Revision trace retriever tool hops failure feedback loop Agent Run Recipe Chatbot USER PROMPT TOOL CALLS FINAL ANSWER "Add 2 cups flour..." Trace Capture Nested Child Spans span: retriever_query span: tool_hop (search->recipe) span: answer_synthesis llm_call & output parsing Human Annotation End-to-End Quality Rating LLM-as-a-Judge Prompt & Tool Adherence Retrieval Metrics Context Precision / Recall Transition Analysis Tool-to-Tool State Graph PASS / FAIL PARTIAL MISMATCH FAIL (RETRIEVAL) BRITTLE TRANSITION Revision Node System Feedback Rubric Changes Prompt Tuning Tool-chain Fixes