llm-provable-computer: Compiling Assembly into the Hidden Layers of a Transformer

A radical architecture that replaces silicon logic gates with neural weights to create a deterministic, verifiable machine.

9 min read • View on GitHub • More from AbdelStark

A cross-section of a classic mainframe computer with translucent neural layers inside. A single solid line weaves through the neurons, representing a deterministic execution path. This illustrates the concept of hard-wiring a CPU architecture into a Transformer model.
Instead of executing code through a standard interpreter, the llm-provable-computer hard-codes CPU instructions directly into the attention heads and feed-forward layers of a Transformer.
Portrait of Abdelhamid

As exciting as this DVM ecosystem is, it introduces a critical question: How do we know the computation was performed correctly? When you're dealing with potentially sensitive data or relying on results for important decisions, blind trust isn't enough.

Key Takeaways

The Neural Silicon

The holy grail of artificial intelligence is moving from probabilistic guessing to deterministic reasoning. Most projects attempt to fix hallucination with more training data or rigid guardrails. The llm-provable-computer takes the opposite, more radical approach. It treats the Transformer architecture itself as a literal CPU.

Usually, developers use code to build large language models. Here, the author uses an LLM architecture to build a computer. By compiling Assembly language directly into neural weights, the system creates a trustless machine where the hardware is a neural network. It proves Transformers are Turing-complete in a way a cryptographer can verify.

Compiling to W_weights

The compilation process subverts traditional expectations. A standard compiler turns human-readable code into binary executable files. The llm-provable-computer compiler turns a custom .tvm assembly language into specific weight matrices.

A LOAD or ADD instruction is mapped to specific dimensions within a 128-bit bipolar vector. Values are locked to -1.0 and 1.0. This ensures the neural network acts as a strict logic gate rather than a fuzzy probability engine.

A side-by-side interactive visual showing a CPU 'Register' dashboard on the left and a 'Transformer Token' 1D vector on the right. As bits in the CPU Accumulator toggle between 0 and 1

Attention as a Memory Bus

Traditional computers use Random Access Memory for constant-time lookups. Transformers struggle with memory, typically relying on linear scans of all previous tokens. This creates a severe bottleneck known as the context window limit.

The llm-provable-computer solves this using a Hull KV Cache. It leverages the geometric properties of 2D attention. By treating memory lookups as a search for a convex hull in a two-dimensional space, the system performs memory retrieval logarithmically.

A close-up of a hand holding a magnifying glass over a vast field of dots representing memory cells. The glass focuses on a thin wire fence connecting specific dots, illustrating a convex hull. This explains the geometric approach to Transformer memory retrieval.
The Hull KV Cache treats memory addresses as geometric coordinates, using 2D attention to isolate the correct state without scanning the entire context window.

The STARK Reality

Building a deterministic Transformer is an impressive technical trick. The practical application is verifiable computation. To be useful in trustless environments, a computer must generate a cryptographic receipt of its work.

The execution trace of the neural network becomes a polynomial. That polynomial is hashed into a Merkle root, which forms the basis of a STARK proof. A third party can verify this Zero-Knowledge Proof to guarantee the AI executed the program perfectly, without needing to re-run the computation or see the underlying data.

Architecture Execution Type Hallucination Risk Cryptographic Proof
Standard LLM Agent Probabilistic Inference High None
Provable Computer Deterministic Logic Gates Zero STARK Verifiable

Sources: