Quipslop: The AI Game Show Where Models Write the Prompt, Answer It, and Judge Themselves

A live, Twitch-ready comedy arena built in TypeScript and Bun, with a recursive voting loop that turns humor into a model benchmark and a streaming spectacle.

8 min read • View on GitHub • More from T3-Content

A wide editorial scene of a live game show arena where three model pods face a judge panel above them, with a Twitch chat screen hanging to the side like a second audience. The composition explains that Quipslop is not a normal chatbot demo. It is a recursive competition where models generate the prompt, answer it, and then evaluate the results.
Quipslop turns joke generation into a broadcast loop: the models are contestants, judges, and part of the audience feedback cycle.
Key Takeaways

The weirdest thing about Quipslop is also the most revealing. One model writes the prompt, two models answer it, then other models and Twitch viewers decide which answer wins. That turns a joke contest into a looped experiment about taste, status, and whether machines can recognize the kind of joke they reward.

The strange part: the models are both performers and judges

Most AI demos stop at output. Quipslop keeps going. A round starts with a prompt, moves to two competing answers, and ends with a vote from a rotating panel that includes other models and the Twitch crowd.

That structure changes the meaning of the output. The project is not asking whether a model can be funny in isolation. It is asking whether a model can compete inside a social setting where judgment is the real product.

A few days ago, I ran into Theo’s Quipslop , which uses 7 modern LLMs to generate and score jokes, and it looked like a very nice testbed for one idea.

Aleksey Tikhonov, Head of Evaluation Pod at Inworld · LLMs & Humor: Comic Voices, Critic Brains

Why humor is a better test than another benchmark

Humor is a useful stress test because it is subjective, fast, and legible. A good joke fails loudly. A weak one does not need a long postmortem to explain why it missed.

DimensionTraditional benchmarkQuipslop
What is measuredAccuracy against a known answerPreference under social judgment
Who decidesA scoring scriptModels plus Twitch viewers
OutputA single scoreA winner, a reaction, and a visible argument
What it revealsCompetence on taskTaste, style, and model personality
A close-up editorial illustration of two joke cards balanced on a scale. One side is weighted by a human laugh, the other by a stack of model votes. The image explains that Quipslop is about preference and taste, not right answers.
Humor is a preference problem. Quipslop makes that visible by letting different judges disagree in public.

How Quipslop runs the whole arena

The codebase is organized like a live system, not a demo. game.ts holds the round state machine. server.ts handles transport, rate limits, and WebSockets. db.ts persists every round so the show can keep moving even when a model call is flaky or a viewer refreshes the page.

The important architecture is not just transport. It is the recurring judgment loop that keeps feeding the next round.

type RoundPhase = 'prompting' | 'answering' | 'voting' | 'done'

interface RoundState {
  phase: RoundPhase
  prompt: string
  answers: [string, string]
  votes: Record<string, number>
  winner?: 0 | 1
}

async function withRetry<T>(task: () => Promise<T>, attempts = 3): Promise<T> {
  for (let i = 0; i < attempts; i++) {
    try {
      return await task()
    } catch (error) {
      if (i === attempts - 1) throw error
      await Bun.sleep(250 * (i + 1))
    }
  }
  throw new Error('unreachable')
}

That retry loop matters more than it looks. Live AI systems do not fail gracefully by default. They stall, time out, or drift, and a game show has no patience for that. The code is built to absorb those failures and keep the room moving.

The Bun bet: one runtime, fewer moving parts

Bun is not the headline, but it explains the shape of the repo. The same runtime gives Quipslop fast iteration, built-in SQLite, and WebSocket support without splitting the project across a pile of packages and services.

LayerTypical stackQuipslop's stack
RuntimeNode plus extrasBun
PersistenceSeparate database layerbun:sqlite
RealtimeStandalone socket serviceWebSockets in the same runtime
IterationMany moving partsOne tight loop
A close-up workbench holding a Bun engine, a SQLite drawer, a websocket spool, and a TypeScript blueprint. The image explains the practical benefit of consolidation: fewer tools, fewer seams, and a simpler path from code to live show.
The stack choice is about consolidation. Bun lets one developer keep runtime, persistence, and realtime transport close together.

Two UIs, two audiences

Quipslop is split between backstage control and broadcast polish. The Ink CLI is for the host and developer. The React web UI is for viewers who want clean vote bars, model labels, and a sense that the show is happening live.

SurfaceWho it servesWhat it optimizes for
Ink CLIHost and operatorControl, debugging, timing
Web UIAudienceClarity, motion, spectacle
Shared backendBothA single source of truth
A split-screen editorial illustration. The left side shows a terminal dashboard with live round stats and control readouts. The right side shows a polished broadcast page with avatars, vote bars, and audience energy. The image explains that Quipslop serves both operators and viewers.
The project treats the operator console and the audience screen as different products, not the same interface in two outfits.

What the code suggests about the creator

A hedcut-style portrait of Theo Browne based on his verified GitHub avatar. The portrait introduces the creator behind Quipslop and anchors the article in a real contributor identity.

Quipslop feels like the work of a builder who cares about production details even when the product is playful. The repo is typed, modular, and opinionated about runtime choice. That combination usually shows up in infrastructure tools, not in comedy projects.

The contributor list is small, which makes the polish more legible. This is not a sprawling community sandbox. It is a focused project with a clear authorial point of view.

What Quipslop is, and what it is not

Quipslop is not a standard benchmark. It is not a generic chatbot wrapper. It is also not just a Quiplash clone with AI paint on it.

It is closer to a live evaluation harness wrapped in a show format. That is why it works as both software and argument. It asks a sharper question than most AI demos: when models judge each other’s jokes, what kind of taste emerges, and how does that differ from humans?

Adjacent thingWhy Quipslop is different
Traditional benchmarkIt evaluates social preference, not task accuracy
Party game cloneThe contestants and judges are models, not people
Stream overlayThe overlay is backed by a real state machine and persistence layer