SkateBench: The Benchmark That Knows a Kickflip from an Inward Heel
A tiny TypeScript evaluation lab for seeing whether models understand skateboarding terminology, and what that precision costs.

This project was created using `bun init` in bun v1.2.0. Bun is a fast all-in-one JavaScript runtime.
- SkateBench turns skateboarding vocabulary into a stress test for model precision.
- Its evaluator is deliberately simple because exact terminology can be judged without a second model.
- The repo is a full benchmark lab, with a fast Bun runner and a dashboard that prices accuracy against throughput and cost.
- Its real value is specificity, which makes it a sharp complement to general-purpose evals.
At first glance, SkateBench looks like a joke. It asks LLMs to recognize skateboarding tricks, then ranks them on whether they know the difference between a kickflip and an inward heel. The real move is more serious. It uses a tiny vocabulary with unforgiving definitions to expose when a model is precise, not just fluent.
The repository comes from t3dotgg, and the README states the project's bias toward speed and iteration plainly. That matters, because SkateBench is not a polished research artifact. It is a practical loop for running models, checking answers, and learning which systems are accurate enough to trust.
The quote captures the whole posture of the repo. Bun gives the benchmark runner a fast start, and the rest of the stack turns that speed into a repeatable workflow. SkateBench is built for the kind of evaluation you can rerun, compare, and extend without turning the project into a lab of its own.
The repo is a benchmark lab, not a single script
The codebase is split cleanly. /bench handles execution, model orchestration, and result writing. /visualizer turns the output into a dashboard with rankings, charts, and cost-aware views of model performance. That split is the real clue. SkateBench is meant to be run over and over, not admired once.
bench/
cli.tsx
index.ts
constants.ts
tests/*.json
visualizer/
app/page.tsx
data/benchmark-results.json
The scoring trick is brutally simple
The evaluator is not trying to be clever. A response passes if it includes one of the accepted answers strings and none of the negative_answers strings. That is enough because skateboarding terms are precise. If the prompt asks for one trick and the model names another, the mismatch is the point.
const isCorrect = (response: string, test: TestCase) =>
test.answers.some(answer => response.includes(answer)) &&
!test.negative_answers.some(answer => response.includes(answer))
The strength of that rule is what it excludes. SkateBench does not need a judge model to infer intent or reward vague proximity. It only needs the domain to be clean enough that a wrong trick name is plainly wrong. That makes the benchmark less theatrical and more trustworthy.
The dashboard turns accuracy into a decision
The visualizer changes the question from “who got it right” to “who got it right at a sane price.” It tracks successRate, tokensPerSecond, and averageCostPerTest, then puts those numbers in the same frame. That matters because a model that is a little better but much slower or more expensive is not automatically the better choice.
This is where SkateBench becomes operational. The benchmark is not just about factual recall. It is about deciding which model gives you the best mix of precision, latency, and cost for a narrow task where mistakes are easy to spot and easy to explain.
Where SkateBench fits among modern evals
| Project | What it tests | Evaluation unit | Scoring style | Why it is different from SkateBench |
|---|---|---|---|---|
| SkateBench | Skateboarding trick terminology | Individual trick definitions | Exact inclusion plus forbidden-term filtering | A narrow domain where terminology is the signal |
| MMLU | Broad general knowledge | Multiple-choice questions | Accuracy over many subjects | Wide coverage, but less domain pressure |
| SWE-bench | Software engineering problem solving | Real GitHub issues and patches | Task success through tests and fixes | Agentic coding, not vocabulary recall |
| SkillEval | Prompt skill quality | Custom skill files and prompts | A/B evaluation and judge scoring | Optimizes instructions, not inherent knowledge |
| TetrisBench | Real-time game play and planning | Game states over episodes | Performance over a sequence | Dynamic decision-making instead of static recall |
SkateBench is not trying to replace broad benchmarks. It is trying to ask a narrower question with far more confidence. That is the useful pattern here. As models improve, the most interesting evals may be the ones that stop pretending one score can tell the whole story.