The DIY Chatbot Arena: Inside instructkr/elo-leaderboard-archive

How a localized LLM leaderboard uses manual stream parsing and a mathematical "underdog bias" to solve the Elo cold-start problem.

7 min read · instructkr/elo-leaderboard-archive

A classic boxing ring viewed from slightly above, with two identical, featureless mechanical automatons facing each other. A judge stands outside holding a clipboard.
The blind test arena relies on strict anonymity and a human-in-the-loop judge to determine Elo rankings.
Key Takeaways

The Localized Arena

While global platforms like LMSYS dominate the AI ranking conversation, they are inherently biased toward English. Evaluating models on linguistic nuances, such as Korean honorifics and specific cultural context, requires a localized arena.

The repository reveals the "glue code" required to build a blind LLM test from scratch. It serves as a sovereign evaluation tool, ensuring that models are judged fairly on the specific cultural axes that matter to local developers.

Solving the Cold Start

When a new model is introduced to an Elo ranking system, it needs a baseline number of battles to establish an accurate rating. Without intervention, a random selection process would leave new models languishing with high variance.

The system solves this with a clever matchmaking algorithm. It applies a manual bias using a maxChatCount * 1.08 formula. By subtracting a model's current chat count from this artificially inflated ceiling, the system generates a selection weight that heavily favors newcomers. This mathematically forces untested models into the arena until their sample size matches the veterans.

The Underdog Bias Generator in action, forcing new models to converge on an Elo score faster.

The Raw Reality of Stream Parsing

High-level AI SDKs are elegant, but they abstract away the messy reality of streaming text. In the core engine, the developer ditches standard OpenAI or LangChain libraries to maintain absolute control over the proxying process.

The code intercepts Server-Sent Events, buffers the fragments using a custom synthesizer, and manually strips the data prefixes. This low-level approach ensures the Next.js UI receives smooth, predictable text updates regardless of the underlying provider's quirks.

synthesizer.toString('utf8').replace("data: ", "").split("\n\ndata: ").forEach(chunk => {
  // Manual chunk parsing and UI buffer updates
});

Polling Over WebSockets

Real-time bidirectional communication is typically the domain of WebSockets. However, maintaining persistent connections introduces significant architectural complexity and scaling challenges.

The frontend opts for a brutal but highly effective alternative: a 200ms timeout polling loop. It repeatedly strikes the completion endpoint to fetch the latest text deltas for Model A and Model B, trading network efficiency for stateless simplicity.

A split composition. Left: a thick taut cable connecting two telegraph machines. Right: a mechanical arm dropping a bucket down a well on a fast timer.
WebSockets maintain a continuous connection (left), while the arena's polling mechanism repeatedly dips into the server for updates (right).

The PocketBase Pragmatism

Building an AI evaluation tool requires reliable state management for models, chat logs, and fluctuating Elo scores. Instead of spinning up a heavy relational database cluster, the project utilizes PocketBase.

This Go-based SQLite wrapper gives the Next.js application an instant backend and an out-of-the-box admin UI. It is the ultimate pragmatic stack for an AI researcher: maximum velocity, zero DevOps overhead, and immediate visibility into the underlying battle data.