SkateBench: The Benchmark That Knows a Kickflip from an Inward Heel

A tiny TypeScript evaluation lab for seeing whether models understand skateboarding terminology, and what that precision costs.

11 min read • View on GitHub • More from T3-Content

A wide editorial scene of a halfpipe built from stacked benchmark cards, with tiny model tokens riding the ramp. Some land cleanly on exact trick labels, while others slip away when the vocabulary becomes too precise. The image explains how SkateBench turns a niche domain into a stress test for model accuracy.
SkateBench uses skateboarding jargon as a precision test, not a novelty prompt.

This project was created using `bun init` in bun v1.2.0. Bun is a fast all-in-one JavaScript runtime.

t3dotgg, Primary Contributor and Maintainer · skatebench README
Key Takeaways

At first glance, SkateBench looks like a joke. It asks LLMs to recognize skateboarding tricks, then ranks them on whether they know the difference between a kickflip and an inward heel. The real move is more serious. It uses a tiny vocabulary with unforgiving definitions to expose when a model is precise, not just fluent.

The repository comes from t3dotgg, and the README states the project's bias toward speed and iteration plainly. That matters, because SkateBench is not a polished research artifact. It is a practical loop for running models, checking answers, and learning which systems are accurate enough to trust.

WSJ hedcut portrait of t3dotgg based on the GitHub avatar. It anchors the article in the maintainer's voice and signals that SkateBench is a builder's tool, not a detached academic benchmark.

The quote captures the whole posture of the repo. Bun gives the benchmark runner a fast start, and the rest of the stack turns that speed into a repeatable workflow. SkateBench is built for the kind of evaluation you can rerun, compare, and extend without turning the project into a lab of its own.

The repo is a benchmark lab, not a single script

The codebase is split cleanly. /bench handles execution, model orchestration, and result writing. /visualizer turns the output into a dashboard with rankings, charts, and cost-aware views of model performance. That split is the real clue. SkateBench is meant to be run over and over, not admired once.

bench/
  cli.tsx
  index.ts
  constants.ts
  tests/*.json
visualizer/
  app/page.tsx
  data/benchmark-results.json

The scoring trick is brutally simple

The evaluator is not trying to be clever. A response passes if it includes one of the accepted answers strings and none of the negative_answers strings. That is enough because skateboarding terms are precise. If the prompt asks for one trick and the model names another, the mismatch is the point.

A close tabletop view of a benchmark card under a magnifying lens, with one answer phrase framed as correct and a nearby forbidden phrase crossed out. The scene explains why SkateBench can use exact string matching when the domain vocabulary is narrow and unambiguous.
When the domain language is exact, the score can stay exact too.

A simple flow from test case to model response to pass or fail explains why the evaluator can stay lightweight.

const isCorrect = (response: string, test: TestCase) =>
  test.answers.some(answer => response.includes(answer)) &&
  !test.negative_answers.some(answer => response.includes(answer))

The strength of that rule is what it excludes. SkateBench does not need a judge model to infer intent or reward vague proximity. It only needs the domain to be clean enough that a wrong trick name is plainly wrong. That makes the benchmark less theatrical and more trustworthy.

The dashboard turns accuracy into a decision

The visualizer changes the question from “who got it right” to “who got it right at a sane price.” It tracks successRate, tokensPerSecond, and averageCostPerTest, then puts those numbers in the same frame. That matters because a model that is a little better but much slower or more expensive is not automatically the better choice.

This is where SkateBench becomes operational. The benchmark is not just about factual recall. It is about deciding which model gives you the best mix of precision, latency, and cost for a narrow task where mistakes are easy to spot and easy to explain.

Where SkateBench fits among modern evals

ProjectWhat it testsEvaluation unitScoring styleWhy it is different from SkateBench
SkateBenchSkateboarding trick terminologyIndividual trick definitionsExact inclusion plus forbidden-term filteringA narrow domain where terminology is the signal
MMLUBroad general knowledgeMultiple-choice questionsAccuracy over many subjectsWide coverage, but less domain pressure
SWE-benchSoftware engineering problem solvingReal GitHub issues and patchesTask success through tests and fixesAgentic coding, not vocabulary recall
SkillEvalPrompt skill qualityCustom skill files and promptsA/B evaluation and judge scoringOptimizes instructions, not inherent knowledge
TetrisBenchReal-time game play and planningGame states over episodesPerformance over a sequenceDynamic decision-making instead of static recall

SkateBench is not trying to replace broad benchmarks. It is trying to ask a narrower question with far more confidence. That is the useful pattern here. As models improve, the most interesting evals may be the ones that stop pretending one score can tell the whole story.