autoresearch: The 630-Line Replacement for the Machine Learning Engineer
Inside karpathy/autoresearch, where the human writes the Markdown protocol and the AI writes the PyTorch.
The training code here is a simplified single-GPU implementation of nanochat. The core idea is that you're not touching any of the Python files like you normally would as a researcher. Instead, you are programming the `program.md` Markdown files that provide context to the AI agents and set up your autonomous research org.
- The framework shifts the engineer's role from writing PyTorch code to managing a high-level Markdown protocol.
- A strict five-minute wall-clock limit forces the agent to optimize for hardware efficiency and real-world speed.
- The system uses Git commits and resets as a minimalist memory mechanism to preserve successful code mutations.
- The project prioritizes rapid code evolution over the generation of academic papers and LaTeX documentation.
The Markdown Manager
The most fascinating aspect of Andrej Karpathy's autonomous research framework is not that an AI is writing code. It is the intentional, extreme constraint placed on the human operator. In this system, the engineer no longer writes PyTorch. Their job is entirely compressed into editing a single text file.
That file is program.md. It acts as the cognitive architecture for an autonomous research organization. The human sets the high-level policy, dictates the areas of exploration, and establishes the rules of engagement. The AI agent, acting as the tireless researcher, reads this protocol and mutates the actual training code overnight.
This approach flips the traditional development loop. The human is no longer debugging tensor shapes or managing memory allocations. The human is the director, and the agent is the execution layer.
The Five-Minute Physics Engine
If the human provides the direction, the environment must provide the constraints. The genius of the framework lies in its localized evolutionary simulator, built on two immutable rules: a strict five-minute wall-clock limit and a single evaluation metric.
Instead of training for a fixed number of epochs or steps, every experiment runs for exactly 300 seconds. This forces the agent to optimize for hardware efficiency. If an architectural change makes the model smarter but slower, it might process fewer tokens in five minutes, resulting in a worse final score. The agent naturally discovers optimizations specific to the user's exact GPU.
The scoring metric itself is val_bpb, or validation bits per byte. Unlike standard cross-entropy loss, bits per byte is information-theoretically grounded and independent of vocabulary size. The agent can experiment with different tokenization strategies without cheating the scoreboard.
The Three Pillars of Autonomous Research
The repository achieves this loop through extreme minimalism. The entire system is effectively 630 lines of code divided into three functional pillars. This prevents the LLM from getting lost in a complex directory tree.
First is the immutable environment, defined in prepare.py. This handles data ingestion and tokenization. The agent is strictly forbidden from modifying this file. It represents the physics of the universe, ensuring all evaluations remain fair and comparable.
Second is the mutable research surface, train.py. This single file contains the complete GPT architecture, the optimizer, and the training loop. By centralizing the logic, the agent has a clear, bounded action space to apply its mutations.
Finally, the system uses Git as the agent's short-term memory and undo button. When the agent completes a five-minute run, it checks the metric. If the score improves, it commits the changes. If the experiment fails, it simply executes a git reset, snapping the code back to its previous pristine state.
Code Evolution over PDF Generation
This minimalist, code-first approach stands in stark contrast to other autonomous AI research frameworks. The current landscape is dominated by heavyweight pipelines aiming to simulate the entire academic publishing process.
| Feature | karpathy/autoresearch | Heavyweight Frameworks (e.g., The AI Scientist) |
|---|---|---|
| Primary Output | Optimized PyTorch code | Conference-ready LaTeX PDF |
| Complexity | Single-file mutation (~630 lines) | Multi-stage pipeline (20+ steps) |
| State Management | Git commits and branch resets | Vector databases and persistent memory |
| Core Goal | Hardware-aware micro-optimization | Academic peer-review simulation |
Projects like Sakana AI's The AI Scientist attempt to handle literature reviews, LaTeX compilation, and peer-review simulation to produce a final PDF. Karpathy's framework ignores the paper entirely. It focuses solely on the inner loop of machine learning research: mutating code and measuring the loss curve.
By stripping away the overhead of formatting and prose generation, the agent can run hundreds of rapid, focused experiments overnight. It is a pragmatic reduction of the research process down to its most fundamental components: one GPU, one file, and one metric.
Sources:
- Repository and documentation: karpathy/autoresearch
- The AI Scientist paper generation framework: sakanaai/ai-scientist