image-editing-arena: Image Editing Arena: Making Image Models Agree Long Enough to Compare Them
A Cloudflare Worker, a schema shim, and a cost-aware UI turn incompatible image-editing models into a controlled side-by-side test.
- Image Editing Arena is a benchmark disguised as a product, because it makes cost, latency, and visual quality compete in the same frame.
- The hardest problem in the repo is not generating an edit, but translating one user request into several incompatible model schemas.
- A Cloudflare Worker keeps the browser thin and the API key hidden, which makes the whole thing feel edge-native instead of backend-heavy.
- The arena format matters because image-editing models win on different tasks, so a single leaderboard leaves too much information out.
Most image editors ask a simple question: which model should I run? This repo asks a harder one: which model should I trust for this specific edit, at this price, with this latency? That shift turns the interface into a benchmark. The product is the comparison.
A benchmark that behaves like a product
That is the trick. Image Editing Arena does not just render outputs from multiple models, it stages a fair contest. One prompt, one source image, several model backends, then a shared comparison layer that shows what each model actually cost and how long it took.
With so many options, it can be hard to figure out which one works best for your needs. In this post, we’re putting them head to head and evaluating each across a range of image editing tasks. By the end, you should have a clear sense of which one fits your workflow.
The real problem is schema mismatch
The neat front end hides a messy truth. Replicate image models do not agree on input shape, so the repo has to normalize one user action into several dialects. Some models want image, others want image_input, and some expect arrays like [image]. The adapter layer is what makes the arena feel uniform.
// Conceptual shape of the adapter layer
function createPrediction(model, input) {
const payload = { ...input };
if (model.owner === 'black-forest-labs') {
payload.input_images = [input.image];
} else if (model.name.includes('edit')) {
payload.image = input.image;
} else {
payload.image_input = [input.image];
}
return proxyToWorker(payload);
}
async function pollPrediction(id) {
while (true) {
const res = await fetch(`/api/replicate/predictions/${id}`);
if (res.status === 200 && res.done) return res.output;
await wait(nextBackoff());
}
}
That adapter is the repo’s most important piece of engineering. It hides model-specific quirks behind a common request path, then hands the request off to a polling loop that keeps checking until the result is ready. In other words, it makes disagreement legible.
Why the browser never talks to Replicate directly
The browser stays clean because the secret stays out of it. A Cloudflare Worker sits in front of Replicate, injects authorization, and forwards the request at the edge. That keeps the app deployable as a lightweight single-page experience without turning the front end into a key vault.
This matters for more than security. It makes the demo feel instant, reduces backend ceremony, and keeps the architecture easy to reason about. The worker is not an incidental proxy. It is what lets the product be a product.
The UI makes trade-offs visible
The interface does something many benchmark pages avoid. It puts pricing and performance next to the image results, so the user can judge output quality in context. The StatsPanel, model selector, and cost fields turn a visual task into a decision problem.
| Model | Arena signal | Best observed strength | Main trade-off |
|---|---|---|---|
| GPT-image-1 | Lowest price in the comparison | Cheap experimentation | Slowest turnaround |
| FLUX.1 Kontext [dev] | Fastest generation | Speed | Quality can slip under heavy optimization |
| SeedEdit 3.0 | Winner on object removal | Background reconstruction | Not the cheapest or fastest |
| Qwen Image Edit | Winner on object removal | Strong edit fidelity | Less of a speed story |
| FLUX.1 Kontext [pro] | Struggled on the bridge test | Still useful for simple edits | Left obvious artifacts in the benchmark task |
The cheapest is GPT-image-1 from OpenAI which starts $0.01 per image, but it has the longest generation time (around 40 seconds). FLUX.1 Kontext [dev] (optimized by Pruna AI) is the fastest at 1.9s per generation and is also one of the cheaper options, but there is, of course, a trade off with image editing quality for hyper-optimized models.
Why this matters beyond one repo
The bigger pattern is arena-fication. Text model leaderboards taught people to compare systems side by side, but image editing has its own constraints, its own schemas, and its own failure modes. A single global winner is too blunt an instrument for that reality.
This repo sits between a benchmark and a demo. That is why it is useful. It does not just say which model is good. It shows how quickly a supposedly simple image edit becomes a question of payload shape, edge proxying, polling logic, and cost awareness.
| Pattern | What it optimizes for | What it hides | Why Image Editing Arena is different |
|---|---|---|---|
| Single-model editor | Fast output | Cross-model trade-offs | It answers one model at a time |
| Generic benchmark page | Scoring and rankings | Product feel | It often ignores the workflow around the test |
| Arena-style editor | Comparison under one prompt | Nothing important, by design | It makes model choice itself the product |