drq: When the benchmark fights back
Digital Red Queen uses an LLM to evolve Core War warriors in an adversarial loop, turning a retro programming game into a lab for open-ended adaptation.
- Digital Red Queen turns Core War into a moving target by judging each generation against the warriors that came before it.
- The LLM is acting less like autocomplete and more like a semantic mutator that can preserve strategy while rewriting implementation.
- The repo's real machinery is a circular-memory virtual machine, round-robin scheduling, and parallel battles that make adversarial pressure concrete.
- The project matters because it turns open-ended adaptation into a controlled sandbox with obvious relevance to security and other arms races.
Most AI benchmarks are static. They freeze the world, run a score, and declare a winner. Digital Red Queen, or DRQ, does the opposite: it keeps the opponent changing, so the only durable skill is adaptation.
Why Sakana AI chose Core War
The repo comes from Sakana AI and MIT collaborators, but the more important detail is the arena they picked. Core War gives them a Turing-complete battleground where "warriors" are assembly-like programs fighting for control of circular memory. That makes the project feel less like a toy benchmark and more like a controlled laboratory for hostile adaptation.
In the game Core War, assembly-like programs called “warriors” fight for control of a virtual computer. Warriors may employ sophisticated strategies including targeted self-replication, data bombing, and massive multithreading, in order to crash other programs, and dominate the machine.
What makes the setup interesting is not nostalgia. It is the shape of the objective. A warrior is not rewarded for a single clean answer. It is rewarded for surviving a population that is actively trying to kill it.
How the loop works
The code splits cleanly. `src/` orchestrates the LLM loop. `corewar/` provides the simulator, Redcode parser, and memory model. `gpt_warriors/` keeps the generated programs as a record of what survived each round, while multiprocessing and pygame handle scale and visualization.
for round_idx in range(num_rounds):
candidates = llm_mutate(champions, prompt_bank)
battles = pool.map(run_single_round, candidates)
scores = fitness_from_survival(battles)
champions = select_top_k(candidates, scores)
archive.extend(champions)
Inside the simulator, the rules are intentionally old school. `corewar/mars.py` steps warriors in round-robin order, `corewar/core.py` wraps memory with modulo arithmetic, and a warrior dies when its task queue goes empty. That combination makes the environment hard to game in a simple, linear way.
Why the LLM is more than autocomplete
This is where the LLM matters. DRQ is not asking the model to autocomplete a Redcode file from scratch and stop there. It feeds the winner back into the system as a semantic mutator, so the model can preserve a strategy while rewriting the implementation.
We find that this dynamic adversarial process leads to the emergence of increasingly general strategies and reveals an intriguing form of convergent evolution, where different code implementations settle into similar high-performing behaviors.
That is a more interesting search problem than code generation alone. It is closer to evolutionary pressure: different code paths can converge on the same behavior because the environment rewards function, not style.
DRQ versus a frozen benchmark
| Model | Selection pressure | What changes | What you learn |
|---|---|---|---|
| Static benchmark | One fixed task set | Nothing between runs | A score on a frozen test |
| Prompted code generation | A human prompt and revision | Mostly the prompt | Whether a model can draft useful code |
| Digital Red Queen | A growing archive of prior warriors | The opponent and the objective | Which strategies still work when the world adapts |
That is why the repo feels bigger than the game it uses. Core War supplies the physics, but DRQ supplies the thesis: the most useful evaluation loop may be the one that refuses to sit still. For security, robotics, or any adversarial domain, that is the right stress test.