VIGA: The AI That Reconstructs Images by Writing the Code Behind Them

A dual-agent system turns visual matching into a loop of rendering, verification, and revision, making image reconstruction editable instead of opaque.

9 to 11 min read • View on GitHub • More from Fugtemypt123

A wide editorial illustration shows a reconstruction loop in a visual debugging lab. A target image sits on the left, a generator writes scene code in the middle, and a verifier inspects a rendered scene on the right with a magnifying lens.
VIGA does not stop at a first guess. It keeps rendering, checking, and revising until the scene matches the target more closely.
Key Takeaways

VIGA is easiest to understand if you stop thinking about image generation and start thinking about debugging. It takes a target image, writes code for a scene, renders the result, inspects the mismatch, and tries again. That is not a single prediction. It is a visual build loop.

VIGA Does Not Draw. It Compiles.

The core move is simple and unusual. VIGA treats a picture as something you can reconstruct by writing executable graphics code, not something you must hallucinate in pixels. In practice, that means Blender scenes for 3D tasks and PowerPoint manipulation for document layout tasks, all produced as editable artifacts.

VIGA is an analysis-by-synthesis code agent for programmatic visual reconstruction. It approaches vision-as-inverse-graphics through an iterative loop of generating, rendering, and verifying scenes against target images.

That editability matters. A raster image is a dead end. A scene file can be opened, moved, relit, re-rendered, or reused. VIGA aims at the kind of output engineers and designers actually want when they care about structure, not just appearance.

Why Editable Output Changes the Game

ModePassesTool accessMemoryOutputFailure mode
One-shot VLM code generation1None or thin wrappersNoneUsually a single script or guessSpatial mistakes stay baked in
Memory-less baseline like BlenderAlchemyMultipleLimitedNo evolving contextProcedural reconstructionForgets prior corrections
VIGAMultiple iterative roundsMCP tool serversYes, via prompt memoryEditable scene programsCan revise against evidence

The practical difference is not subtle. If the camera is wrong, the lighting is off, or the object is shifted a few pixels, VIGA can inspect that error and target the correction. That makes the output feel closer to software development than to image synthesis.

A side-by-side engraving-style comparison shows a sealed raster image on one side and an editable scene file on the other. The editable side includes visible camera, lighting, and object controls, emphasizing that the result can be changed after generation.
VIGA’s output is not just something to view. It is something to edit, which changes the economics of reconstruction.

The Generator and the Verifier Work Like a Build Loop

The interesting part is not that VIGA uses tools. It is that the tools sit inside a closed visual reasoning loop.

The generator writes code. The verifier does not just grade the result. It inspects the render, looks for specific discrepancies, and feeds correction back into the next pass. That is the loop that makes VIGA feel disciplined rather than speculative.

Under the hood, the agents are connected through Model Context Protocol tool servers. That gives the system actual hands. Blender, PowerPoint, and segmentation tools can be discovered and invoked without hardwiring every task into a bespoke interface.

MCP Gives the Agents Hands

This is one of the cleanest parts of the design. The repository separates reasoning from execution. The agent thinks in prompts and code. The tool layer handles the messy reality of rendering engines, slide decks, and scene inspection. That separation is what makes the system feel reusable across domains.

# Conceptual shape of the loop
code = generator.write(target_image, memory=history)
rendered = blender.execute(code)
delta = verifier.inspect(rendered, target_image)
code = generator.revise(code, feedback=delta, memory=history)

Prompt Memory Keeps the Search From Forgetting Itself

VIGA is not stateless. The prompt builder keeps the target image, initial code, initial render, and previous attempts in a sliding memory window. That matters because reconstruction is a search problem, and search gets weaker when it forgets where it has already been.

This is also why the verifier matters more than a simple scoring model. It can compare successive candidates against the same target, preserve the trail of mistakes, and keep the generator from wandering back into old errors. The system is trying to converge, not merely react.

The Benchmark Story: More Than Static Scenes

Task familyToolingWhat VIGA changesWhy it matters
BlenderBenchBlender + render loopReconstructs editable 3D scenesTests spatial reasoning and scene structure
SlideBenchPowerPoint + layout toolsRebuilds document layouts as codeShows the loop works for 2D composition too
Dynamic scenesPhysics-aware scene executionHandles interaction and motionPushes from inverse graphics toward inverse physics

The broader point is that the architecture is not narrowly tied to one benchmark. The same write-run-compare-revise pattern can move from static scenes to slide layouts and even dynamic interactions. That is a stronger claim than image generation usually makes.

What VIGA Is Really Selling: Inverse Graphics as Software

VIGA’s thesis is bigger than reconstruction quality. It suggests that visual understanding can be operationalized as software: inspectable, editable, and improved through iteration. That is a different mental model from prompt in, pixels out.

The result is a system that behaves more like an engineer than a dream machine. It writes, tests, checks the evidence, and revises the implementation. For anyone trying to build trustworthy multimodal tools, that is the part worth copying.