drq: When the benchmark fights back

Digital Red Queen uses an LLM to evolve Core War warriors in an adversarial loop, turning a retro programming game into a lab for open-ended adaptation.

10 min read · SakanaAI/drq

A circular arena of memory with tiny code warriors fighting across a ring. It shows how Digital Red Queen turns Core War into a moving battlefield where survival, not static accuracy, decides who stays in the game.
DRQ treats the evaluation loop as the product. Each generation creates the next opponent, so the test never sits still.
Key Takeaways

Most AI benchmarks are static. They freeze the world, run a score, and declare a winner. Digital Red Queen, or DRQ, does the opposite: it keeps the opponent changing, so the only durable skill is adaptation.

Why Sakana AI chose Core War

The repo comes from Sakana AI and MIT collaborators, but the more important detail is the arena they picked. Core War gives them a Turing-complete battleground where "warriors" are assembly-like programs fighting for control of circular memory. That makes the project feel less like a toy benchmark and more like a controlled laboratory for hostile adaptation.

In the game Core War, assembly-like programs called “warriors” fight for control of a virtual computer. Warriors may employ sophisticated strategies including targeted self-replication, data bombing, and massive multithreading, in order to crash other programs, and dominate the machine.

Sakana AI, Organization · Sakana AI blog

What makes the setup interesting is not nostalgia. It is the shape of the objective. A warrior is not rewarded for a single clean answer. It is rewarded for surviving a population that is actively trying to kill it.

How the loop works

The code splits cleanly. `src/` orchestrates the LLM loop. `corewar/` provides the simulator, Redcode parser, and memory model. `gpt_warriors/` keeps the generated programs as a record of what survived each round, while multiprocessing and pygame handle scale and visualization.

DRQ is a feedback loop, not a one-shot generator. The archive shapes the prompt, the prompt shapes the candidates, and the candidates are judged by a simulator that keeps the pressure moving.

for round_idx in range(num_rounds):
    candidates = llm_mutate(champions, prompt_bank)
    battles = pool.map(run_single_round, candidates)
    scores = fitness_from_survival(battles)
    champions = select_top_k(candidates, scores)
    archive.extend(champions)

Inside the simulator, the rules are intentionally old school. `corewar/mars.py` steps warriors in round-robin order, `corewar/core.py` wraps memory with modulo arithmetic, and a warrior dies when its task queue goes empty. That combination makes the environment hard to game in a simple, linear way.

A close view of a Redcode program being cut, copied, and recombined on a drafting table. It explains how the LLM acts less like code autocomplete and more like a targeted mutator that preserves strategy while changing implementation.
DRQ is strongest when it treats the model as a strategist, not just a writer. The mutation step matters because the search is about preserving behavior under pressure.

Why the LLM is more than autocomplete

This is where the LLM matters. DRQ is not asking the model to autocomplete a Redcode file from scratch and stop there. It feeds the winner back into the system as a semantic mutator, so the model can preserve a strategy while rewriting the implementation.

We find that this dynamic adversarial process leads to the emergence of increasingly general strategies and reveals an intriguing form of convergent evolution, where different code implementations settle into similar high-performing behaviors.

Sakana AI, Organization · Sakana AI blog

That is a more interesting search problem than code generation alone. It is closer to evolutionary pressure: different code paths can converge on the same behavior because the environment rewards function, not style.

DRQ versus a frozen benchmark

ModelSelection pressureWhat changesWhat you learn
Static benchmarkOne fixed task setNothing between runsA score on a frozen test
Prompted code generationA human prompt and revisionMostly the promptWhether a model can draft useful code
Digital Red QueenA growing archive of prior warriorsThe opponent and the objectiveWhich strategies still work when the world adapts

That is why the repo feels bigger than the game it uses. Core War supplies the physics, but DRQ supplies the thesis: the most useful evaluation loop may be the one that refuses to sit still. For security, robotics, or any adversarial domain, that is the right stress test.