karpathy/nanochat: The $100 Full-Stack LLM
How Andrej Karpathy collapsed the entire generative AI pipeline from raw tokens to a conversational web UI into a single hackable codebase.
Excited to release new repo: nanochat! (it’s among the most unhinged I’ve written). Unlike my earlier similar repo nanoGPT which only covered pretraining, nanochat is a minimal, from scratch, full-stack training/inference pipeline
- The repository collapses the entire AI development lifecycle into a single pipeline that trains a functional chat model for roughly $100.
- A single depth parameter automatically scales all other hyperparameters using Chinchilla-optimal laws to simplify model architecture.
- The system replaces the traditional AdamW optimizer with Muon to achieve faster mathematical convergence through orthogonal weight updates.
- The dataloader prioritizes sample quality by discarding a third of the raw data to ensure every training sequence is perfectly aligned.
The $100 Speedrun
Training a capable chat model used to require a dedicated infrastructure team and a massive computational budget. That reality is shifting rapidly. Algorithmic efficiency and better datasets have turned Large Language Model training from a corporate endeavor into a competitive sport for solo developers.
Released in October 2025, nanochat is a minimal experimental harness designed for training LLMs from scratch on a single GPU node. It is not just another inference engine or a pretraining script. It is the entire vertical slice of AI development.
The project collapses tokenization, pretraining, supervised fine-tuning, reinforcement learning, and a web UI into one pipeline. The goal is simple but radical. A user with a single cloud GPU can train and run a functional ChatGPT-style interface in about four hours for approximately $100.
The "One Dial" Philosophy
Most machine learning repositories expose a bewildering array of hyperparameters. Users must manually tune learning rates, attention heads, and hidden dimensions. Karpathy replaced this dark art with a single abstraction.
The core of the nanochat user experience is the --depth parameter. Instead of guessing the right architectural balance, the user simply sets the model depth. The system uses Chinchilla-optimal scaling laws to calculate the rest automatically.
If you want a larger model, you dial up the depth. The width, heads, learning rate, and training horizons scale in mathematically proven proportion. This is the breakthrough that makes the repository accessible to developers outside the AI research bubble.
The Engine Room: Muon and FP8
Reaching GPT-2 performance on a $100 budget requires ruthless optimization. The repository leans heavily on bleeding-edge techniques that have only recently become standard. The architecture incorporates Flash Attention 3 for high-performance memory access on Hopper GPUs and a minimalist 150-line implementation of FP8 precision training.
The most significant divergence from legacy models is the optimizer. For years, AdamW was the default choice for training neural networks. nanochat replaces it with Muon for internal matrix parameters.
Muon (Orthogonalized Gradient Descent) treats weight updates as orthogonalizations rather than simple scaled steps. This accelerates convergence significantly. It is the secret engine that makes the aggressive training timelines possible.
The 100% Utilization Dataloader
Data processing in AI training usually involves greedy packing. Documents are jammed into fixed-width arrays, often resulting in sentences being sliced in half across rows. This creates confusing sequences where a document starts without a Beginning-Of-Sequence (BOS) token.
The nanochat dataloader takes a different approach. It implements a BOS-aligned Best-Fit packing algorithm. It maintains a buffer and searches for the exact document that perfectly fits the remaining space in a sequence.
To ensure perfect batch packing where every row starts cleanly, the algorithm deliberately discards roughly 35% of the raw dataset. It is a calculated trade-off. The system prioritizes perfect context alignment and sample quality over raw data volume.
The Landscape: Full-Stack vs. Piecemeal
To understand the value of nanochat, you have to look at the surrounding ecosystem. The open-source AI landscape is highly fragmented.
| Project | Primary Focus | Key Optimization | Target User |
|---|---|---|---|
| nanochat | Full lifecycle (Train to UI) | Chinchilla scaling dial | Researchers, Solo Devs |
| modded-nanogpt | Pretraining speedruns | Extreme kernel fusion | Optimization Engineers |
| llama2.c | Portability | Pure C implementation | Embedded Systems Devs |
| FastChat | Production serving | UI and API routing | MLOps Teams |
Projects like modded-nanogpt push the absolute limits of pretraining speed but stop there. FastChat provides excellent serving infrastructure but expects you to bring a pre-trained model. nanochat is unique because it bridges these worlds. It provides the only minimal, transparent tool covering the entire vertical slice of modern AI.
Sources: