The DIY Chatbot Arena: Inside instructkr/elo-leaderboard-archive
How a localized LLM leaderboard uses manual stream parsing and a mathematical "underdog bias" to solve the Elo cold-start problem.
- The project builds a localized, language-specific Chatbot Arena to evaluate models on cultural nuances that global leaderboards miss.
- An "Underdog Bias" algorithm mathematically forces new models into matchups more frequently to solve the Elo cold-start problem.
- The architecture eschews heavy SDKs in favor of manual Server-Sent Events (SSE) stream parsing and a brute-force 200ms polling loop.
- Next.js paired with PocketBase creates an instant, pragmatic backend for rapid AI research prototyping.
The Localized Arena
While global platforms like LMSYS dominate the AI ranking conversation, they are inherently biased toward English. Evaluating models on linguistic nuances, such as Korean honorifics and specific cultural context, requires a localized arena.
The repository reveals the "glue code" required to build a blind LLM test from scratch. It serves as a sovereign evaluation tool, ensuring that models are judged fairly on the specific cultural axes that matter to local developers.
Solving the Cold Start
When a new model is introduced to an Elo ranking system, it needs a baseline number of battles to establish an accurate rating. Without intervention, a random selection process would leave new models languishing with high variance.
The system solves this with a clever matchmaking algorithm. It applies a manual bias using a maxChatCount * 1.08 formula. By subtracting a model's current chat count from this artificially inflated ceiling, the system generates a selection weight that heavily favors newcomers. This mathematically forces untested models into the arena until their sample size matches the veterans.
The Raw Reality of Stream Parsing
High-level AI SDKs are elegant, but they abstract away the messy reality of streaming text. In the core engine, the developer ditches standard OpenAI or LangChain libraries to maintain absolute control over the proxying process.
The code intercepts Server-Sent Events, buffers the fragments using a custom synthesizer, and manually strips the data prefixes. This low-level approach ensures the Next.js UI receives smooth, predictable text updates regardless of the underlying provider's quirks.
synthesizer.toString('utf8').replace("data: ", "").split("\n\ndata: ").forEach(chunk => {
// Manual chunk parsing and UI buffer updates
});
Polling Over WebSockets
Real-time bidirectional communication is typically the domain of WebSockets. However, maintaining persistent connections introduces significant architectural complexity and scaling challenges.
The frontend opts for a brutal but highly effective alternative: a 200ms timeout polling loop. It repeatedly strikes the completion endpoint to fetch the latest text deltas for Model A and Model B, trading network efficiency for stateless simplicity.
The PocketBase Pragmatism
Building an AI evaluation tool requires reliable state management for models, chat logs, and fluctuating Elo scores. Instead of spinning up a heavy relational database cluster, the project utilizes PocketBase.
This Go-based SQLite wrapper gives the Next.js application an instant backend and an out-of-the-box admin UI. It is the ultimate pragmatic stack for an AI researcher: maximum velocity, zero DevOps overhead, and immediate visibility into the underlying battle data.