skillgrade: Unit Testing the English Language
As AI agents move from chat interfaces to autonomous coworkers, Minko Gechev's new framework replaces "prompt vibes" with repeatable, dockerized evidence.
The problem? There’s no way to know if they actually work. You write a text file, hand it to an agent, and hope for the best. When you tweak the instructions, you have no signal telling you whether that change made things better or worse. You’re flying blind.
- The framework runs evaluations across multiple trials to provide a mathematical success rate for non-deterministic AI agents.
- Isolated Docker containers prevent state pollution by providing a fresh environment for every individual test execution.
- A hybrid grading system combines deterministic shell scripts with LLM-based rubrics to measure both binary outcomes and qualitative reasoning.
- The tool specifically validates an agent's ability to discover and execute modular instructions stored in local skill files.
The Fluke Factor
Building AI agents feels like alchemy. You tweak an English prompt, run it once, and if it works, you ship it. But one successful run in a non-deterministic system is statistically meaningless. A set of instructions that works today might fail completely tomorrow.
This framework treats an agent's instructions as a compiled binary. It runs evaluations across dozens of trials to calculate a reliable success rate. If you change a sentence in your prompt, skillgrade provides mathematical proof of whether you caused a regression.
A Sandbox for Every Trial
A fair test requires a clean environment. If an agent creates a file during its first attempt, the second attempt cannot inherit that state. State pollution ruins benchmarks. skillgrade solves this with a strict isolation model using its Provider architecture.
The framework spins up a fresh Docker container for every single execution. It builds a base image, injects the targeted skill instructions into specific hidden directories, and spawns isolated containers for parallel runs. The agent can execute arbitrary shell commands to solve the task without risking the host machine.
Determinism Meets Nuance
Grading an AI agent is notoriously difficult. Traditional unit tests check for binary outcomes. Did the code compile or did it crash? This deterministic approach is fast but entirely blind to the quality of the process.
skillgrade introduces a hybrid grading system. Developers write bash scripts to verify hard requirements, ensuring the agent actually completed the requested task. Then, they layer an LLM-based rubric on top to evaluate the agent's reasoning. The final score is a weighted combination of absolute facts and qualitative analysis.
| Grading Strategy | Mechanism | Best Used For |
|---|---|---|
| Deterministic | Executes a shell script (e.g., npm test) and checks exit codes or JSON output. |
Verifying compilation, file existence, or strict syntax requirements. |
| LLM Rubric | Passes the full terminal transcript to a judge model (like Claude 3.5 Sonnet) with a grading prompt. | Evaluating efficiency, politeness, or adherence to complex procedural rules. |
The SKILL.md Standard
The framework aligns with a critical shift in how we build for AI. Instead of stuffing every possible rule into a massive, fragile system prompt, developers are moving toward progressive disclosure. They write focused SKILL.md files that act as modular documentation.
skillgrade specifically tests tool discovery. It places these skill files in the container environment and asks the agent to perform a high-level task. The evaluation only passes if the agent successfully searches the filesystem, reads the correct skill file, and executes its exact steps. It is an integration test for autonomous behavior.
# eval.yaml
task: "Refactor the authentication service to use the new token schema."
provider:
type: docker
image: node:20
skills:
- path: .agents/skills/auth-schema/SKILL.md
graders:
- type: deterministic
command: npm run test:auth
weight: 0.8
- type: llm_rubric
rubric: "Did the agent avoid modifying the legacy user table?"
weight: 0.2
Sources:
- mgechev/skillgrade on GitHub
- Unit Tests for AI Agent Skills by Minko Gechev