autoresearch: The 630-Line Replacement for the Machine Learning Engineer

Inside karpathy/autoresearch, where the human writes the Markdown protocol and the AI writes the PyTorch.

6 min read • karpathy/autoresearch

A human hand holding a pencil connected by a thread to a mechanical hand typing equations.
The role of the engineer shifts from writing the code to directing the agent that writes the code.
Portrait of Andrej Karpathy

The training code here is a simplified single-GPU implementation of nanochat. The core idea is that you're not touching any of the Python files like you normally would as a researcher. Instead, you are programming the `program.md` Markdown files that provide context to the AI agents and set up your autonomous research org.

— Andrej Karpathy, karpathy/autoresearch
Key Takeaways

The Markdown Manager

The most fascinating aspect of Andrej Karpathy's autonomous research framework is not that an AI is writing code. It is the intentional, extreme constraint placed on the human operator. In this system, the engineer no longer writes PyTorch. Their job is entirely compressed into editing a single text file.

That file is program.md. It acts as the cognitive architecture for an autonomous research organization. The human sets the high-level policy, dictates the areas of exploration, and establishes the rules of engagement. The AI agent, acting as the tireless researcher, reads this protocol and mutates the actual training code overnight.

This approach flips the traditional development loop. The human is no longer debugging tensor shapes or managing memory allocations. The human is the director, and the agent is the execution layer.

The Five-Minute Physics Engine

If the human provides the direction, the environment must provide the constraints. The genius of the framework lies in its localized evolutionary simulator, built on two immutable rules: a strict five-minute wall-clock limit and a single evaluation metric.

Instead of training for a fixed number of epochs or steps, every experiment runs for exactly 300 seconds. This forces the agent to optimize for hardware efficiency. If an architectural change makes the model smarter but slower, it might process fewer tokens in five minutes, resulting in a worse final score. The agent naturally discovers optimizations specific to the user's exact GPU.

The scoring metric itself is val_bpb, or validation bits per byte. Unlike standard cross-entropy loss, bits per byte is information-theoretically grounded and independent of vocabulary size. The agent can experiment with different tokenization strategies without cheating the scoreboard.

A flowchart showing the "5-Minute Hill-Climbing Loop". Three main nodes: "The Brain (program.md)"

The Three Pillars of Autonomous Research

The repository achieves this loop through extreme minimalism. The entire system is effectively 630 lines of code divided into three functional pillars. This prevents the LLM from getting lost in a complex directory tree.

First is the immutable environment, defined in prepare.py. This handles data ingestion and tokenization. The agent is strictly forbidden from modifying this file. It represents the physics of the universe, ensuring all evaluations remain fair and comparable.

Second is the mutable research surface, train.py. This single file contains the complete GPT architecture, the optimizer, and the training loop. By centralizing the logic, the agent has a clear, bounded action space to apply its mutations.

Finally, the system uses Git as the agent's short-term memory and undo button. When the agent completes a five-minute run, it checks the metric. If the score improves, it commits the changes. If the experiment fails, it simply executes a git reset, snapping the code back to its previous pristine state.

Code Evolution over PDF Generation

This minimalist, code-first approach stands in stark contrast to other autonomous AI research frameworks. The current landscape is dominated by heavyweight pipelines aiming to simulate the entire academic publishing process.

Feature karpathy/autoresearch Heavyweight Frameworks (e.g., The AI Scientist)
Primary Output Optimized PyTorch code Conference-ready LaTeX PDF
Complexity Single-file mutation (~630 lines) Multi-stage pipeline (20+ steps)
State Management Git commits and branch resets Vector databases and persistent memory
Core Goal Hardware-aware micro-optimization Academic peer-review simulation

Projects like Sakana AI's The AI Scientist attempt to handle literature reviews, LaTeX compilation, and peer-review simulation to produce a final PDF. Karpathy's framework ignores the paper entirely. It focuses solely on the inner loop of machine learning research: mutating code and measuring the loss curve.

A massive printing press compared to a small, glowing forge.
Heavyweight frameworks focus on generating academic papers, while minimalist agents focus purely on iterating the code.

By stripping away the overhead of formatting and prose generation, the agent can run hundreds of rapid, focused experiments overnight. It is a pragmatic reduction of the research process down to its most fundamental components: one GPU, one file, and one metric.


Sources: