The PC Part Picker for Local AI: Unpacking llmfit
How a Rust-based diagnostic engine uses bandwidth math and hardware probes to solve the 'will it run' problem before you download a single weight.
- llmfit eliminates local AI trial and error by mathematically proving model compatibility before any weights are downloaded.
- The Rust engine dynamically calculates bits-per-parameter and handles MoE active parameters to prevent overestimating VRAM requirements.
- Multi-vendor hardware probes correctly distinguish between discrete VRAM and Apple Silicon unified memory.
- A predictive 'Plan' mode estimates tokens-per-second to guide future hardware upgrades based on specific model targets.
Curing VRAM Anxiety
The local AI ecosystem is plagued by trial and error. Developers download massive Hugging Face weights only to find their GPU lacks the memory to load the context window. This out-of-memory crash is a visceral shared experience for anyone running local models. llmfit attacks this inefficiency by moving compatibility checks to the very beginning of the workflow.
Instead of reactive testing, the tool acts as a preemptive diagnostic engine. It cross-references your specific hardware against a database of hundreds of models. The result is a mathematical proof of compatibility that saves hours of wasted bandwidth.
The Sieve of Quantization
The core intelligence of the project lives in its quantization logic. The engine treats quantization as a dynamic variable rather than a static file. It uses heuristics like bits per parameter to calculate exactly how much memory a specific target requires.
This logic is particularly crucial for complex architectures. When evaluating Mixture-of-Experts models, simple tools often overestimate requirements by counting total parameters. The llmfit engine calculates based on active parameters, ensuring users do not artificially limit their model choices.
One detail I appreciate is how it handles Mixture-of-Experts (MoE) models. For example, Mixtral 8x7B has around 46.7B total parameters, but only about 12.9B are active per token during inference. Many tools still treat it like a full 46B parameter model when estimating requirements. This tool accounts for the actual active parameters.
Interrogating the Silicon
Accurate math requires accurate inputs. The hardware probe cascades through checks for NVIDIA, AMD, and Apple Silicon. It queries system interfaces to map the exact memory topology of the host machine.
A critical distinction is its handling of unified memory. On Apple Silicon and modern APUs, shared RAM performs differently than discrete VRAM. The engine understands this architecture and routes memory calculations accordingly, categorizing hardware into backend enums that dictate the inference runtime.
Predicting the Upgrade Delta
Diagnostic tools look at the present. The Plan mode looks at the future. By using memory bandwidth lookups to estimate tokens-per-second, it transforms from a diagnostic utility into a hardware advisor.
Users can simulate hardware upgrades to see exactly what GPU they need to buy. It calculates the theoretical maximum speed based on memory bandwidth and model size, giving concrete purchasing guidance for specific AI workloads.
Preemptive vs. Reactive Tooling
The standard approach to local AI relies on generic runners that require manual testing. Benchmarking tools require executing the model to gauge performance. llmfit stands out by providing a zero-download mathematical proof.
| Feature | llmfit | llm-checker | Trial & Error (Ollama/HF) |
|---|---|---|---|
| Verification Method | Mathematical Proof | Runtime Benchmark | Raw Execution |
| Bandwidth Cost | Zero bytes | Full Model Download | Full Model Download |
| Hardware Awareness | Deep Multi-Vendor Probe | Generic Runner | OS Default |
| MoE Parameter Logic | Active Parameters Only | Total Parameters | Total Parameters |