Benchmarks & Eval

LLM benchmarks, eval harnesses, model comparison tools, and agent performance testing

97 explainers
Pixieology: Steering the Fae in the Machine
Benchmarks & Eval
Pixieology: Steering the Fae in the Machine
How MoralityLabAI uses mechanistic interpretability to surgically toggle between cold logic and lyrical whimsy.
7 min read