skillgrade: Unit Testing the English Language

As AI agents move from chat interfaces to autonomous coworkers, Minko Gechev's new framework replaces "prompt vibes" with repeatable, dockerized evidence.

7 min read • View on GitHub • More from mgechev

A single cracked gear tooth on an otherwise pristine clockwork mechanism, contained by a thin firewall line.
skillgrade brings deterministic testing to the probabilistic world of AI agents.
Portrait of Minko Gechev

The problem? There’s no way to know if they actually work. You write a text file, hand it to an agent, and hope for the best. When you tweak the instructions, you have no signal telling you whether that change made things better or worse. You’re flying blind.

Minko Gechev, Creator (Unit Tests for AI Agent Skills)
Key Takeaways

The Fluke Factor

Building AI agents feels like alchemy. You tweak an English prompt, run it once, and if it works, you ship it. But one successful run in a non-deterministic system is statistically meaningless. A set of instructions that works today might fail completely tomorrow.

This framework treats an agent's instructions as a compiled binary. It runs evaluations across dozens of trials to calculate a reliable success rate. If you change a sentence in your prompt, skillgrade provides mathematical proof of whether you caused a regression.

A scientist watching a single gold coin land perfectly on its edge while thousands of other coins land randomly.
A single successful run is a statistical anomaly. True confidence requires repeated, isolated trials.

A Sandbox for Every Trial

A fair test requires a clean environment. If an agent creates a file during its first attempt, the second attempt cannot inherit that state. State pollution ruins benchmarks. skillgrade solves this with a strict isolation model using its Provider architecture.

The framework spins up a fresh Docker container for every single execution. It builds a base image, injects the targeted skill instructions into specific hidden directories, and spawns isolated containers for parallel runs. The agent can execute arbitrary shell commands to solve the task without risking the host machine.

A horizontal pipeline showing the lifecycle of an agent evaluation trial. Stage 1: "Provision" shows a pristine box (Docker container) being created. Stage 2: "Inject" shows a document labeled SKILL.md and a folder of fixtures being inserted into the box. Stage 3: "Execute" shows an AI core (Gemini/Claude) interacting with the box

Determinism Meets Nuance

Grading an AI agent is notoriously difficult. Traditional unit tests check for binary outcomes. Did the code compile or did it crash? This deterministic approach is fast but entirely blind to the quality of the process.

skillgrade introduces a hybrid grading system. Developers write bash scripts to verify hard requirements, ensuring the agent actually completed the requested task. Then, they layer an LLM-based rubric on top to evaluate the agent's reasoning. The final score is a weighted combination of absolute facts and qualitative analysis.

Grading Strategy Mechanism Best Used For
Deterministic Executes a shell script (e.g., npm test) and checks exit codes or JSON output. Verifying compilation, file existence, or strict syntax requirements.
LLM Rubric Passes the full terminal transcript to a judge model (like Claude 3.5 Sonnet) with a grading prompt. Evaluating efficiency, politeness, or adherence to complex procedural rules.
A magnifying glass split down the middle. One side shows green/red binary pixels; the other shows handwritten notes on a scroll.
Hybrid grading combines the binary certainty of a shell script with the contextual understanding of an LLM judge.

The SKILL.md Standard

The framework aligns with a critical shift in how we build for AI. Instead of stuffing every possible rule into a massive, fragile system prompt, developers are moving toward progressive disclosure. They write focused SKILL.md files that act as modular documentation.

skillgrade specifically tests tool discovery. It places these skill files in the container environment and asks the agent to perform a high-level task. The evaluation only passes if the agent successfully searches the filesystem, reads the correct skill file, and executes its exact steps. It is an integration test for autonomous behavior.

# eval.yaml
task: "Refactor the authentication service to use the new token schema."
provider:
  type: docker
  image: node:20
skills:
  - path: .agents/skills/auth-schema/SKILL.md
graders:
  - type: deterministic
    command: npm run test:auth
    weight: 0.8
  - type: llm_rubric
    rubric: "Did the agent avoid modifying the legacy user table?"
    weight: 0.2

Sources: