Noodle: The Rust Transformer That Turns LLMs Into a Readable System

A minimal GPT-style model, built from scratch, that shows how training, fine-tuning, sampling, and local chat can all live in one typed Rust codebase.

11 min read • View on GitHub • More from penberg

A workbench scene with a small transformer opened like a precision watch, surrounded by a Rust toolchain stamp, typed model schematics, a laptop, and a cloud GPU crate linked by a thin cable. It explains that Noodle treats an LLM as a legible system, not a mysterious black box.
Noodle’s trick is not that it hides complexity. It arranges the whole stack so the complexity becomes readable, then gives you a local path and a cloud path when the workload demands it.

Noodle is a minimal language model implementation written in Rust. It's designed to be simple, educational, and easy to understand — a from-scratch implementation of a transformer-based language model that you can train on your own data.

Pekka Enberg, Author/Maintainer · penberg/noodle: README.md
Key Takeaways

Noodle’s most interesting claim is not that it is a transformer in Rust. It is that the entire lifecycle, from data prep to sampling, has been pulled back into view. If you have ever felt that modern LLM tooling hides the mechanics inside framework glue, this repo takes the opposite bet.

That matters because Noodle is aimed at developers who want to understand the machine, not just call it. The repository describes itself as a minimal language model implementation written in Rust, and that word minimal does a lot of work. The codebase stays small enough to read, but complete enough to train, fine-tune, and chat.

Rust Is the Point, Not Just the Language

Rust changes how this project feels. The model is generic over a backend, which means the same core logic can run against different hardware targets without rewriting the architecture itself. That is a small detail with a big effect: the model stops being a pile of tensor calls and starts looking like an actual library.

WSJ hedcut-style portrait of Pekka Enberg based on his GitHub avatar. It identifies the maintainer behind Noodle and supports the article’s origin context with a verified reference image.

The architecture is simple on the surface and unusually disciplined underneath: one model core, multiple backends, several lifecycle commands.

pub struct Model<B: Backend> {
    token_emb: Embedding<B>,
    pos_emb: Embedding<B>,
    blocks: Vec<TransformerBlock<B>>,
    ln_f: LayerNorm<B>,
    head: Linear<B>,
}

impl<B: Backend> Model<B> {
    pub fn forward(&self, tokens: Tensor<B, 2>) -> Tensor<B, 3> {
        // embeddings -> blocks -> final projection
        todo!()
    }
}

The important line is the generic parameter, Model<B: Backend>. That is the whole thesis of the stack in one signature. The model is not welded to one accelerator story, so the rest of the repository can stay focused on the actual language model instead of backend translation.

The Model Is a GPT-2 Style Decoder, Kept Plain on Purpose

Noodle follows the classic decoder-only transformer pattern: token embeddings, positional embeddings, a stack of transformer blocks, final normalization, and an output projection. The design is intentionally unflashy. That is the point. It wants to show how little machinery you need before the system starts feeling like a language model.

token ids -> token embeddings + position embeddings -> transformer blocks -> layer norm -> logits -> next token

The design document makes the scale explicit: 4 transformer blocks, 256 hidden dimension, 4 attention heads, 256 tokens of context, and about 29 million parameters. That is small enough to read about, but large enough to behave like a real model. It is a teaching model that still has to earn its outputs.

Fine-Tuning Is Handled Like a Text Parsing Problem

This is where Noodle gets especially clear-eyed. Instead of treating instruction tuning as some opaque training mode, the code parses examples into instruction and response sections, then masks the instruction tokens so the loss only applies to the answer. That makes supervised fine-tuning feel like string handling plus loss accounting, which is exactly what it is.

A close-up of an instruction-response training card on a white desk. The instruction half is muted under a masking stencil, the response half is sharply inked, and a thin loss line touches only the response tokens. A small sampler valve sits above it, indicating the next stage of token selection. It explains how Noodle turns instruction tuning into explicit parsing and masking rather than hidden training magic.temperature, top-k, top-p, and repetition penalty without any text. The artist uses tightly packed crosshatching lines layered at different angles to build up shadow and form, with clean open areas of white for highlights. The line work has the quality of a classic metal engraving, precise and deliberate, with varied line weights where bold contour lines define shapes and finer interior lines create tonal depth. The overall style evokes vintage newspaper editorial illustrations from The Economist or Wall Street Journal. No color, no gradients, no grey fills. Only black lines on white. The background MUST be pure white #FFFFFF. No paper texture, no cream, no off-white, no noise, no grain. Perfectly clean flat white background." loading="lazy">
Fine-tuning becomes legible when the prompt is masked and the loss is aimed only at the response. The code is doing bookkeeping, not mysticism.
fn prepare_example(sample: &str) -> PreparedExample {
    let instruction = extract_section(sample, "Instruction");
    let response = extract_section(sample, "Response");

    PreparedExample {
        input_ids: tokenize(&instruction),
        labels: mask_prompt_tokens(&instruction, &response),
    }
}

That explicitness pays off. You can see where the prompt ends, where the target begins, and why the model is not punished for failing to predict the question it was given. For a developer coming from a different domain, this is the kind of implementation detail that turns a training pipeline into something you can trust.

Sampling Is Where Small Models Become Pleasant to Use

A small model can sound clever for a few tokens and then fall into a loop. Noodle addresses that with a layered sampling stack: temperature, top-k, top-p, and repetition penalty. The repetition penalty is the unsung part. It is what keeps generation from circling the same phrase until the chat feels broken.

let next = sample_next_token(
    logits,
    Temperature::new(0.8),
    TopK::new(40),
    TopP::new(0.9),
    RepetitionPenalty::new(1.1),
);

In practice, that means inference is not just “run the model.” It is a pipeline that shapes uncertainty before the next token is chosen. The result is a more usable small model, which is exactly what an educational repo should care about.

The Modal Bridge Makes the Project Honest

Noodle does not pretend a laptop is the right place for every job. The repository includes a bridge to Modal for the heavier training path, while keeping the local Rust experience intact for model code, inference, and chat. That is a better story than purity for its own sake. It says: use the local machine for understanding, and use the cloud when the training run needs it.

ProjectPrimary languageMain goalTraining supportInference focusWhat it is best at
NoodleRustEducational, full lifecycle model stackYesYesReadable Rust code that spans training, fine-tuning, and chat
nanoGPTPythonMinimal GPT learning referenceYesSomeThe clearest Python path into GPT-style training
llama.cppC/C++Fast local inferenceNoYesOptimized on-device inference for many model variants

That comparison is the key to understanding Noodle’s niche. nanoGPT is the canonical teaching repo in Python. llama.cpp is the local inference heavyweight. Noodle sits between them: training-capable, Rust-native, and intentionally small enough to read without losing the whole plot.

What Noodle Is Really Teaching

Noodle is not trying to win benchmark theater. It is trying to make the mechanics of an LLM feel inspectable again. That matters because the project collapses the gap between concept and implementation. You can trace the model, the training loop, the sampler, and the CLI as one coherent artifact, which is rare and valuable.

The bigger lesson is architectural, not just educational. AI stacks do not have to be opaque, and they do not have to be scattered across a tangle of scripts. Noodle argues that a modern language model can be typed, portable, and understandable without giving up the ability to do real work.