karpathy/nanochat: The $100 Full-Stack LLM

How Andrej Karpathy collapsed the entire generative AI pipeline from raw tokens to a conversational web UI into a single hackable codebase.

By Repo Explainer | 12 min read | karpathy/nanochat

A sprawling industrial factory being squeezed by a mechanical vise into a small pocket watch, representing the compression of LLM training complexity.
The entire lifecycle of a generative AI product is now compressed into a single, highly readable codebase.
Portrait of Andrej Karpathy

Excited to release new repo: nanochat! (it’s among the most unhinged I’ve written). Unlike my earlier similar repo nanoGPT which only covered pretraining, nanochat is a minimal, from scratch, full-stack training/inference pipeline

Andrej Karpathy, Creator. Source
Key Takeaways

The $100 Speedrun

Training a capable chat model used to require a dedicated infrastructure team and a massive computational budget. That reality is shifting rapidly. Algorithmic efficiency and better datasets have turned Large Language Model training from a corporate endeavor into a competitive sport for solo developers.

Released in October 2025, nanochat is a minimal experimental harness designed for training LLMs from scratch on a single GPU node. It is not just another inference engine or a pretraining script. It is the entire vertical slice of AI development.

The project collapses tokenization, pretraining, supervised fine-tuning, reinforcement learning, and a web UI into one pipeline. The goal is simple but radical. A user with a single cloud GPU can train and run a functional ChatGPT-style interface in about four hours for approximately $100.

The "One Dial" Philosophy

Most machine learning repositories expose a bewildering array of hyperparameters. Users must manually tune learning rates, attention heads, and hidden dimensions. Karpathy replaced this dark art with a single abstraction.

The core of the nanochat user experience is the --depth parameter. Instead of guessing the right architectural balance, the user simply sets the model depth. The system uses Chinchilla-optimal scaling laws to calculate the rest automatically.

A slider labeled "Model Depth" (e.g.

If you want a larger model, you dial up the depth. The width, heads, learning rate, and training horizons scale in mathematically proven proportion. This is the breakthrough that makes the repository accessible to developers outside the AI research bubble.

The Engine Room: Muon and FP8

Reaching GPT-2 performance on a $100 budget requires ruthless optimization. The repository leans heavily on bleeding-edge techniques that have only recently become standard. The architecture incorporates Flash Attention 3 for high-performance memory access on Hopper GPUs and a minimalist 150-line implementation of FP8 precision training.

The most significant divergence from legacy models is the optimizer. For years, AdamW was the default choice for training neural networks. nanochat replaces it with Muon for internal matrix parameters.

A close-up of two mechanical gears. One gear is perfectly machined with sharp, orthogonal teeth. The other gear has slightly rounded, imprecise teeth.
Muon treats weight updates as orthogonalizations, providing sharper, more precise mathematical convergence compared to the traditional rounding of AdamW.

Muon (Orthogonalized Gradient Descent) treats weight updates as orthogonalizations rather than simple scaled steps. This accelerates convergence significantly. It is the secret engine that makes the aggressive training timelines possible.

The 100% Utilization Dataloader

Data processing in AI training usually involves greedy packing. Documents are jammed into fixed-width arrays, often resulting in sentences being sliced in half across rows. This creates confusing sequences where a document starts without a Beginning-Of-Sequence (BOS) token.

The nanochat dataloader takes a different approach. It implements a BOS-aligned Best-Fit packing algorithm. It maintains a buffer and searches for the exact document that perfectly fits the remaining space in a sequence.

A stream of text blocks (documents of varying lengths) arrives. In "Standard Packing"

To ensure perfect batch packing where every row starts cleanly, the algorithm deliberately discards roughly 35% of the raw dataset. It is a calculated trade-off. The system prioritizes perfect context alignment and sample quality over raw data volume.

The Landscape: Full-Stack vs. Piecemeal

To understand the value of nanochat, you have to look at the surrounding ecosystem. The open-source AI landscape is highly fragmented.

Project Primary Focus Key Optimization Target User
nanochat Full lifecycle (Train to UI) Chinchilla scaling dial Researchers, Solo Devs
modded-nanogpt Pretraining speedruns Extreme kernel fusion Optimization Engineers
llama2.c Portability Pure C implementation Embedded Systems Devs
FastChat Production serving UI and API routing MLOps Teams

Projects like modded-nanogpt push the absolute limits of pretraining speed but stop there. FastChat provides excellent serving infrastructure but expects you to bring a pre-trained model. nanochat is unique because it bridges these worlds. It provides the only minimal, transparent tool covering the entire vertical slice of modern AI.


Sources: