A streaming Mixture-of-Experts (MoE) inference engine written in pure C, that runs models larger than your RAM on ordinary consumer hardware - MiniMax-M2 (230B) - GPT-OSS (20B and 120B) and Qwen3-MoE
10 stars
C
Your first custom repo explanation is free. Reading existing public explainers always stays free.
This will take 10-20 minutes. You can close the tab and come back later.