langchain-ai/ai-evals-hw-langsmith: The Homework Repo That Teaches You to Debug AI Like a System

A five-part LangSmith curriculum that starts with a recipe bot and ends with trace replay, judge alignment, retrieval metrics, and transition-matrix failure analysis.

8 to 10 min read • View on GitHub • More from langchain-ai

A wide editorial scene of a control room where paper trace logs feed into a central workbench. A recipe bot sits on one side as a mechanical kitchen instrument, while a judge clipboard, annotation tags, and feedback loops connect the system together. The image explains that the repo treats AI evaluation as an engineering loop built from traces, judgments, and iteration.
The repo’s trick is simple: it turns a blank eval exercise into a living trace archive you can inspect, score, and improve.
Key Takeaways

A Fake Production Dashboard That Teaches the Real Thing

The most interesting thing in this repo is not the homework structure. It is the illusion of production. By replaying traces, feedback, and failures into LangSmith, it gives you a living system to inspect before you ever see real traffic.

LangSmith Evaluation lets you evaluate agent performance, grounded in real production trace data and aligned to human judgment.

LangChain Documentation, Official Resource · LangSmith Evaluation Overview

That matters because most eval tutorials start with abstract metrics. This repo starts with evidence. You are not told to trust a score. You are shown the trace behind it, the annotation that challenged it, and the workflow change that followed.

Why This Is More Than a Prompting Course

The course is really about lifecycle thinking. Prompting is only the first move. The bigger discipline is observe, annotate, judge, compare, and revise until the agent becomes a system you can reason about.

That shift changes the mental model. Instead of asking whether a prompt is clever, you ask whether the agent is measurable. Instead of asking whether a response sounds good, you ask which failure mode produced it and which test would catch it next time.

The Five Homeworks Form a Full AI Debugging Pipeline

One trace can support several evaluators, and each one catches a different kind of failure.

The homeworks are sequenced like a lab, not a lecture. hw1 builds the agent and the trace. hw2 introduces annotation and error analysis. hw3 aligns the judge. hw4 measures retrieval. hw5 inspects transition patterns across tools.

HomeworkMain questionPrimary lens
hw1Can the agent run and trace cleanly?Basic observability
hw2What kinds of mistakes are humans seeing?Manual annotation
hw3Can an LLM judge match human judgment?Judge alignment
hw4Is retrieval finding the right evidence?Recall@K and MRR
hw5Where do multi-tool chains become brittle?Transition matrices

The Judge Has to Be Evaluated Too

The sharpest idea in the repo is that the evaluator is itself a model artifact. If you let an LLM grade output without calibration, you are just moving uncertainty into a different box.

Every company when building a mission-critical agent will build benchmarks for that agent.

Harrison Chase, CEO of LangChain · Harrison Chase Interview - Agent Benchmarks

That is why hw3 matters. It does not just ask you to score responses. It asks you to compare scores against human review in the Evaluator Playground, then tighten the rubric until the judge becomes useful instead of merely confident.

How the Repo Moves Data Through the Loop

A close-up editorial illustration showing one trace splitting into several branches. One branch becomes annotation cards, another becomes a judge score sheet, and a third becomes a lattice of tool hops like a small transition matrix. The image explains how a single recorded run can generate multiple evaluation lenses.
The same trace powers different questions: what happened, who judged it, and where the workflow bent.

The mechanics are straightforward and useful. Traces are uploaded, feedback is attached, and synthetic datasets stand in for real traffic. That makes the course feel like production without requiring production.

# Conceptual flow, based on the repo's scripts
agent_run = run_agent(user_prompt)
trace_id = upload_traces(agent_run)
attach_feedback(trace_id, human_review)
judge_score = llm_judge(trace_id)
retrieval_metrics = compute_recall_at_k(trace_id)
transition_profile = build_transition_matrix(trace_id)

# Then use the failures to improve the next run
revise_rubric()
revise_prompt()
revise_tool_chain()

That loop is the whole point. Observability becomes input, not output.

RAG Metrics and Transition Matrices Make Failure Visible

hw4 and hw5 widen the frame. hw4 asks whether the retriever finds the right recipe evidence, using Recall@K and MRR. hw5 asks a different question: once the agent starts moving between tools, where does the sequence become fragile?

That distinction matters. Retrieval metrics tell you whether the right context was available. Transition matrices tell you whether the agent used that context well enough to move through the task without getting stuck or derailed.

LensWhat it measuresWhat it misses
Retrieval metricsWhether the right evidence was surfacedWhether the agent used that evidence well
Transition analysisHow the agent moves between toolsWhether the final answer is fluent
Final answer gradingSurface qualityProcess quality and hidden brittleness

Where It Fits in the Eval Landscape

This repo does not try to out-platform Langfuse, Phoenix, DeepEval, or Braintrust. It does something narrower and rarer: it teaches a method inside the LangChain ecosystem.

ProjectPrimary strengthWhat makes it different here
LangSmith homework repoTrace-first pedagogyTeaches the workflow, not just the tool
LangfuseOpen observability platformBroader platform, less instructional sequencing
PhoenixRAG and drift analysisStrong analysis, less curriculum structure
DeepEvalOffline testing styleGood for unit-test thinking, not trace replay
BraintrustHigh-throughput eval executionEnterprise eval muscle, different teaching goal

The distinction is clarity of method. The repo is not trying to be the biggest suite. It is trying to make the mental model obvious.

Why This Repo Matters

The broader lesson is simple. AI debugging is becoming trace-driven, not prompt-driven. The useful unit of work is no longer a clever string. It is a recorded behavior you can replay, judge, annotate, and improve.

That is why this repo lands. It gives learners a realistic failure archive, a scaffold for evaluating it, and a way to turn the result back into better tests. That is not just coursework. It is the shape of the job.