SakanaAI/sparser-faster-llms: How TwELL Makes Sparse Transformers Run Like They Were Built for Hopper

A deep dive into the repo that couples activation sparsity with H100-tuned CUDA kernels, so the model’s zeros become real speed, lower memory use, and lower energy.

8 min read • View on GitHub • More from SakanaAI

A massive GPU drawn like a cathedral of tiled chambers, with dense blocks entering one side and neatly packed sparse tiles exiting the other in aligned lanes. The image explains the repo’s core idea: sparsity only matters when the hardware can consume it without breaking the execution path.
TwELL is not just a sparse format. It is a way to make sparsity travel cleanly through Hopper-class kernels.
Key Takeaways

The Zeroes Are Not the Point

The interesting thing about SakanaAI/sparser-faster-llms is not that it makes models sparse. It is that it tries to make sparsity survive contact with a GPU. That sounds obvious until you remember how many sparse ideas die the moment wall-clock time enters the conversation.

The repo is built around a simple but stubborn problem: fewer active weights do not automatically mean less latency. Dense GPU kernels are brilliant at regularity. Sparse models often destroy that regularity, then pay for it in divergence, awkward memory access, and kernel overhead.

A close-up split scene shows irregular sparse blocks on one side and neatly stacked TwELL tiles on the other entering a straight kernel lane. The contrast explains why generic sparsity often fails to produce speed, while structured packing preserves the GPU’s preferred rhythm.
Generic sparsity can look smaller on paper and still run badly. TwELL is built to keep the machine in its comfort zone.

Why Dense GPUs Make Sparse Models Slow

The mismatch is structural. GPUs want tiled, predictable, high-throughput work. Generic sparse formats tend to introduce exactly the opposite: uneven work distribution, extra indirection, and wasted movement through memory.

The core trick is not pruning. It is converting sparsity into a layout the kernel can execute without stalling.

ApproachKernel regularityMemory trafficSpeedup potentialMaps cleanly to GPU tilingComplexity lives in
Dense baselineHighHighNone from sparsityYesLarge matmuls
Generic sparse representationLowOften still highUnreliableUsually notFormat conversion and branching
TwELL on HopperHighLowerReal wall-clock gainYesPacking format and custom CUDA kernels

TwELL: The Packing Format That Fits the Machine

TwELL, short for Tile-wise ELLPACK, is the repo’s answer to the sparse-performance gap. The point is not just to store fewer values. The point is to store them in a way that keeps the compute path tiled, aligned, and friendly to Hopper-class execution.

That is why the repository leans so hard on custom CUDA. The packing format and the kernels are not separate stories. TwELL exists because the kernel needs a predictable shape, and the kernel works because the data has been repackaged to match that shape.

In practice, the repository’s kernel layer handles dense-to-TwELL and TwELL-to-dense paths, plus fused MLP execution. That matters because every extra pass over memory is another chance to lose the gains sparsity promised in the first place.

The Repo’s Two Jobs: Train Sparsity, Then Cash It In

The repo is not just an inference demo. It includes sparse model definitions, training code, benchmark scripts, and optional energy measurement. That makes it feel like a complete engineering artifact, not a one-off kernel drop.

# Typical repo-level flow, simplified
python train.py --config-name sparse_llama
python benchmark_inference.py --model checkpoints/sparse
python benchmark_inference.py --measure-energy true

That separation is the right one. Training has to discover and preserve useful sparsity. Inference then has to cash that sparsity in through custom execution paths that do not waste the savings on bookkeeping.

How do we make LLMs faster and lighter? Don’t force the GPU to adapt to sparsity. Reshape the sparsity to fit the GPU! ⚡️ Excited to share our new #ICML2026 paper in collaboration with @NVIDIA: "Sparser, Faster, Lighter Transformer Language Models". This work introduces new http

Sakana AI, SakanaAILabs, 70,467 followers · @SakanaAILabs on X

What the Paper Claims, What the Repo Ships

The paper makes the case for high activation sparsity with limited quality loss. The repo makes that claim testable by shipping code paths for checkpoints, benchmarking, and kernel execution rather than leaving the idea trapped in a chart.

That distinction matters. A lot of efficiency papers are strongest in presentation and weakest in reproducibility. Here, the implementation layer is the argument. The reader can inspect the packing, the kernels, and the measurement machinery that ties speed to actual GPU behavior.

What you getPaperRepo
Sparsity targetClaims strong sparsity with limited downstream impactImplements sparse model paths and checkpoints
Speed storyReports throughput and efficiency gainsProvides benchmark scripts and kernel code
Energy storyFrames lower energy as part of the valueIncludes energy measurement tooling
Engineering proofMethodology and resultsExecutable implementation

Where This Fits in the Efficiency Stack

TwELL does not replace quantization, pruning, or other efficiency methods. It sits in a narrower lane: making activation sparsity executable on Hopper without giving back the wall-clock gains to kernel overhead.

That is a useful correction to the usual efficiency conversation. Quantization reduces precision. TwELL changes the shape of the work itself. The payoff is not abstract elegance. It is that the GPU spends more of its time doing useful tiled math and less time untangling a sparse mess.

Bottom Line

SakanaAI/sparser-faster-llms is interesting because it treats sparsity as a systems problem. The repo’s real contribution is not the idea that zeros are good. It is the machinery that lets zeros become speed on actual hardware.

That is why TwELL matters. It is the missing middle layer between a sparse model and a fast one, and the repo ships that middle layer instead of assuming it will appear on its own.