mattbeane/research-quals: The Competency Gate for the AI Era

How a Python CLI uses LLM evaluators to withhold advanced automation until researchers prove they can do the work manually.

8 min read • View on GitHub • More from mattbeane

A massive bank vault door slightly ajar, revealing automated gears inside, while a human hand with a magnifying glass inspects a manuscript outside. This illustrates the concept of human manual effort unlocking automated research tools.
In research-quals, manual human effort is the cryptographic key required to unlock the automation engine.
Key Takeaways

The Automation Embargo

While the rest of the software world is building AI agents that do the work for you, research-quals uses AI to stop you from automating your work. It is an educational lockbox. The core of this system is an abstraction called the UNLOCK_MAP, located inside cli/competency_gate.py. This file acts as a bouncer for your terminal.

Attempting to run an advanced research command without the prerequisite student level results in a hard block. The software actively refuses to execute commands. It demands that you prove your foundational tacit knowledge first. You must do the work manually, submit it, and let the machine grade your human effort.

Grading the Vibe

Qualitative research is notoriously subjective. To solve this, cli/evaluator.py uses Few-Shot Prompting and Anthropic to compare a student's Markdown submission against an expert_baseline.md. The programmatic rubrics are unforgiving. For instance, the system has a hard-coded requirement demanding that students find at least 50 percent of disconfirming evidence to pass.

The system does not just look for keywords. It compares the distribution of findings against what a senior scholar would identify. This is a technical implementation of adversarial evidence, a core tenet of rigorous qualitative research. If you fail, the system enforces a cooldown period of 7 to 30 days to prevent brute-forcing the LLM evaluator.

The Competency Skill Tree visualizes the UNLOCK_MAP dependency graph and the evaluation rubrics required to progress.

Cryptography for Pedagogy

Once the LLM evaluator approves the work, the state must be maintained. The lib/credentials architecture handles this by minting UUID-backed badges. The issuer.py script writes these badges to a local record.json file. This creates a portable proof of work for academic competency.

{
  "student_id": "usr_9a8b7c6d",
  "badges": [
    {
      "type": "DOMAIN",
      "domain_id": "domain-1",
      "uuid": "8f92a1b3-4c5d-6e7f-8a9b-0c1d2e3f4a5b",
      "issued_at": "2024-01-15T10:30:00Z",
      "score": 85
    }
  ],
  "current_level": 1
}

This JSON file is the ultimate source of truth for the local environment. It is a simple, stateless approach to managing user progression without requiring a complex database backend.

The Theory-Forge Ecosystem

The ultimate payoff is access to the theory-forge ecosystem. The research-quals project is not a standalone application but a licensing engine. Contrast this with fully autonomous pipelines that take a prompt and output a formatted paper. Those systems risk hallucination and weak theoretical foundations.

A vintage jeweler's balance scale. The left tray holds a single solid iron weight labeled 'Disconfirming Evidence'. The right tray holds a massive pile of feathers labeled 'Cherry-Picked Quotes'. The scale tips toward the heavy iron weight. This illustrates the strict rubric demanding quality over volume.
The evaluator heavily weights disconfirming evidence over sheer volume of findings, penalizing cherry-picked data.

Gated AI represents a sustainable path forward for rigorous disciplines. By forcing humans to remain in the loop during the foundational phases, the system ensures that when the advanced AI agents are finally unlocked, they are guided by a competent human orchestrator.