GuppyLM: The Fish That Makes LLMs Legible

A tiny transformer, synthetic conversations, and browser deployment turn one open-source repo into a clean tour of how language models are actually built.

8 min read • View on GitHub • More from arman-bd

A fish tank on a workbench feeds a compact machine and a laptop. The scene explains that GuppyLM turns synthetic data into a tiny transformer and then into a browser demo, making the whole path easy to follow.
GuppyLM is less a model than a fully visible pipeline, with persona, architecture, and deployment lined up in one small repo.
Key Takeaways

GuppyLM is not interesting because it is small. It is interesting because it makes the entire LLM pipeline legible. The fish persona is the clue. Once you see that the model learns from 60,000 synthetic conversations, the repo stops looking like a magic trick and starts looking like a set of design choices.

Built a ~9M param LLM from scratch to understand how they actually work. Vanilla transformer, 60K synthetic conversations, ~130 lines of PyTorch. Trains in 5 min on a free Colab T4. The fish thinks the meaning of life is food. Fork it and swap the personality for your own character.

armanified, Project Creator · Show HN thread

Persona is the product

Most chat demos hide their training data behind a generic assistant voice. GuppyLM does the opposite. The model is trained to be a specific fish, with lowercase replies, tank life, food, water temperature, and short practical thoughts.

That narrowness is the point. Because the persona is encoded in the data, the behavior becomes inspectable. You can trace a weird answer back to the template engine in generate_data.py, not to some vague emergent quality in a giant corpus.

Code is not always self-documenting, and can often tell you how it was written, but not why.

arcanemachiner, Hacker News Commenter · Show HN thread

A plain transformer under the glass

Under the fish is a deliberately ordinary Transformer. model.py uses attention blocks, a feed-forward network, learned positional embeddings, and ReLU. There is no RoPE, no SwiGLU, and no grouped query attention.

self.lm_head = nn.Linear(config.d_model, config.vocab_size, bias=False)
self.lm_head.weight = self.tok_emb.weight

That restraint matters. Weight tying between the token embedding and the output head keeps the footprint small. Combined with the roughly 8.7 million parameter scale, it makes the whole model easy to train, export, and inspect.

The training loop keeps the same discipline. train.py uses AdamW, cosine decay, mixed precision, and a held-out validation set that picks the best checkpoint. Nothing here is exotic. The pedagogy is in the combination, not the novelty.

From notebook to browser

The repo's best lesson is that the path stays visible from start to finish. Synthetic data becomes a tiny model, the model becomes exported weights, and the weights become a browser demo through ONNX and WebAssembly. Nothing disappears into a managed service.

The model stays understandable because the path from synthetic conversations to browser runtime never gets hidden.

This projects shares similarities with Minix. Minix is still used at universities as an educational tool for teaching operating system design. Minix is the operating system that taught Linus Torvalds how to design (monolithic) operating systems. Similarly having students adding capabilities to GuppyLM is a good way to learn LLM design.

thomasfl, Hacker News Commenter · Show HN thread

Why the comparison matters

ProjectMain focusWhat you learn
GuppyLMOne fish persona, a tiny vanilla transformer, and browser deployment.How training data, architecture, and runtime line up into one understandable system.
nanoGPTA minimal GPT training stack for flexible experimentation.The general mechanics of GPT training without the persona layer.
Typical tutorial repoA notebook that gets you to first output fast.How to run a demo, but not how the model's behavior is shaped.

Compared with nanoGPT, GuppyLM trades breadth for a single, memorable persona. Compared with generic tutorial repos, it trades convenience for an end-to-end story. The fish works because it gives the learner one shape to track all the way through the stack.