VIGA: The AI That Reconstructs Images by Writing the Code Behind Them
A dual-agent system turns visual matching into a loop of rendering, verification, and revision, making image reconstruction editable instead of opaque.
- VIGA reframes vision as a debugging loop, where code generation, rendering, and verification replace one-shot image guessing.
- Its real output is an editable scene program, which makes the result easier to inspect, adjust, and trust than a flat raster image.
- The verifier is not a passive critic, because it can change viewpoints and send targeted corrections back into the next iteration.
- MCP gives the agents real tool access, so the system behaves less like a demo and more like a general reconstruction platform.
VIGA is easiest to understand if you stop thinking about image generation and start thinking about debugging. It takes a target image, writes code for a scene, renders the result, inspects the mismatch, and tries again. That is not a single prediction. It is a visual build loop.
VIGA Does Not Draw. It Compiles.
The core move is simple and unusual. VIGA treats a picture as something you can reconstruct by writing executable graphics code, not something you must hallucinate in pixels. In practice, that means Blender scenes for 3D tasks and PowerPoint manipulation for document layout tasks, all produced as editable artifacts.
VIGA is an analysis-by-synthesis code agent for programmatic visual reconstruction. It approaches vision-as-inverse-graphics through an iterative loop of generating, rendering, and verifying scenes against target images.
That editability matters. A raster image is a dead end. A scene file can be opened, moved, relit, re-rendered, or reused. VIGA aims at the kind of output engineers and designers actually want when they care about structure, not just appearance.
Why Editable Output Changes the Game
| Mode | Passes | Tool access | Memory | Output | Failure mode |
|---|---|---|---|---|---|
| One-shot VLM code generation | 1 | None or thin wrappers | None | Usually a single script or guess | Spatial mistakes stay baked in |
| Memory-less baseline like BlenderAlchemy | Multiple | Limited | No evolving context | Procedural reconstruction | Forgets prior corrections |
| VIGA | Multiple iterative rounds | MCP tool servers | Yes, via prompt memory | Editable scene programs | Can revise against evidence |
The practical difference is not subtle. If the camera is wrong, the lighting is off, or the object is shifted a few pixels, VIGA can inspect that error and target the correction. That makes the output feel closer to software development than to image synthesis.
The Generator and the Verifier Work Like a Build Loop
The generator writes code. The verifier does not just grade the result. It inspects the render, looks for specific discrepancies, and feeds correction back into the next pass. That is the loop that makes VIGA feel disciplined rather than speculative.
Under the hood, the agents are connected through Model Context Protocol tool servers. That gives the system actual hands. Blender, PowerPoint, and segmentation tools can be discovered and invoked without hardwiring every task into a bespoke interface.
MCP Gives the Agents Hands
This is one of the cleanest parts of the design. The repository separates reasoning from execution. The agent thinks in prompts and code. The tool layer handles the messy reality of rendering engines, slide decks, and scene inspection. That separation is what makes the system feel reusable across domains.
# Conceptual shape of the loop
code = generator.write(target_image, memory=history)
rendered = blender.execute(code)
delta = verifier.inspect(rendered, target_image)
code = generator.revise(code, feedback=delta, memory=history)
Prompt Memory Keeps the Search From Forgetting Itself
VIGA is not stateless. The prompt builder keeps the target image, initial code, initial render, and previous attempts in a sliding memory window. That matters because reconstruction is a search problem, and search gets weaker when it forgets where it has already been.
This is also why the verifier matters more than a simple scoring model. It can compare successive candidates against the same target, preserve the trail of mistakes, and keep the generator from wandering back into old errors. The system is trying to converge, not merely react.
The Benchmark Story: More Than Static Scenes
| Task family | Tooling | What VIGA changes | Why it matters |
|---|---|---|---|
| BlenderBench | Blender + render loop | Reconstructs editable 3D scenes | Tests spatial reasoning and scene structure |
| SlideBench | PowerPoint + layout tools | Rebuilds document layouts as code | Shows the loop works for 2D composition too |
| Dynamic scenes | Physics-aware scene execution | Handles interaction and motion | Pushes from inverse graphics toward inverse physics |
The broader point is that the architecture is not narrowly tied to one benchmark. The same write-run-compare-revise pattern can move from static scenes to slide layouts and even dynamic interactions. That is a stronger claim than image generation usually makes.
What VIGA Is Really Selling: Inverse Graphics as Software
VIGA’s thesis is bigger than reconstruction quality. It suggests that visual understanding can be operationalized as software: inspectable, editable, and improved through iteration. That is a different mental model from prompt in, pixels out.
The result is a system that behaves more like an engineer than a dream machine. It writes, tests, checks the evidence, and revises the implementation. For anyone trying to build trustworthy multimodal tools, that is the part worth copying.