Tom Turney

Working on LLM inference systems, KV cache compression, and kernel-level optimizations (TurboQuant).

GitHub 46 repos 602 followers

Explained projects

Yet to be explained

vllm-swift
vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon
Python274 stars
Explain
offlabel
HTML21 stars
Explain
longctx
Open long-context inference stack: retrieval + open weights, no closed parts. pip install longctx.
Python7 stars
Explain
macos-ci-burst-runner
Opt-in menu-bar controller for safely using any Mac as a GitHub Actions burst runner
Shell4 stars
Explain
pastors-pocket-spurgeon
Offline Spurgeon study companion — pastoral counsel, sermon prep, and Sword & Trowel sermon grading. Gemma-4-12B fine-tune served on TurboQuant llama.cpp.
Python3 stars
Explain
ffai
FFAI — F*cking Fast AI. Dual-engine inference: Swift (Apple/iPhone, native Metal) + Rust (cross-platform: CUDA/Vulkan/ROCm). Shared kernels via metaltile.
Swift2 stars
Explain
zero
Rust2 stars
Explain
ds4-turboquant_plus
Private backup of TQ+ Metal port work on antirez/ds4. Branch tom/turbo3-kv-cuda (turbo2/3/4 + Wave M3 inline-dequant + half-tile + CUDA stubs + h8 Flash WIP).
C1 stars
Explain
tqkit
Unified toolkit for benchmarking and integrating TurboQuant+ KV-cache compression across llama.cpp, vLLM, MLX, and vllm-swift.
Python1 stars
Explain
melee
SSBM GALE01 decomp progress fork (doldecomp/melee). Provide your own main.dol.
C0 stars
Explain
homebrew-tap
Homebrew tap for TheTom projects (vllm-swift)
Ruby0 stars
Explain