MapTrace: Teaching models to walk a map without breaking it

MapTrace is a synthetic data machine for spatial reasoning. It teaches multimodal models to respect walls, corridors, and connectivity instead of drawing straight lines through whatever looks shortest.

13 min read • View on GitHub • More from google-research

A folded indoor map with a thin ink route that tries to cross a wall, then bends into a corridor and reaches a red endpoint. The scene shows that the real task is not spotting objects on a map, but following valid connectivity through space.
MapTrace trains models on the difference between a visible route and a valid route.
Key Takeaways

Most multimodal models can name the objects on a map. They fail when the question becomes physical: can you get from here to there without walking through a wall? MapTrace exists because that gap is still wide, and it is more useful to treat it as a data problem than as a model mysticism problem.

You can show a modern multimodal model a zoo map and it’ll tell you all about elephants and giraffes. But ask it to trace a walking route from the entrance to the reptile house? There’s a decent chance it’ll draw a line straight through an animal enclosure or a gift shop.

Ghazi Khan, Author · mgks.dev blog

Why MapTrace exists

The paper behind MapTrace is blunt about the failure mode. It says multimodal large language models are strong at visual and textual reasoning, but still limited at fine-grained spatial understanding, especially route tracing on maps. That is the hole the project tries to fill, and it does it with synthetic supervision instead of a hand-built annotation factory.

While Multimodal Large Language Models have achieved human-like performance on many visual and textual reasoning tasks, their proficiency in fine-grained spatial understanding, such as route tracing on maps remains limited.

Artemis Panagopoulou, et al., Researchers · arXiv paper

What is in the repo

The repository reads like a focused research pipeline, not a product shell. The engine room sits in src/, where postprocess_data.py cleans and compresses route data, finetune_gemma27b.py handles tuning, and vllm_inference.py handles evaluation and serving. Around that are the usual open-source edges, plus a distributed training config, and the stack is pure Python: PyTorch, Hugging Face transformers, trl, peft, pandas, pyarrow, shapely, Pillow, and vLLM.

The pipeline is the story: MapTrace turns synthetic maps into validated route traces, then compresses them into training targets.

The diagram matters because the project is not just about making maps. It is about building a closed loop where synthetic scenes, graph-based routes, and critic models combine to produce better labels than a human team could feasibly draw at scale. That is the real differentiator here, not any single model checkpoint.

The data pipeline is the product

The most interesting code lives in src/postprocess_data.py. It draws start and end markers directly onto the image, then simplifies the path with a Ramer-Douglas-Peucker style pass through shapely.geometry.LineString.simplify, which keeps the route shape while cutting it down to something a language model can actually learn. The coordinates are normalized to a 0.0 to 1.0 range and rounded to four decimals, which saves tokens without throwing away the geometry.

A close-up of a jagged ink route being compressed into a cleaner chain of anchor points. It shows how MapTrace turns pixel-level traces into compact training targets that fit a model context window without losing the shape of the route.
Route simplification is how the project turns geometry into language-friendly supervision.

How it trains and serves

On the tuning side, finetune_gemma27b.py uses LoRA to hit the useful projection layers, then leans on Flash Attention 2 and FSDP to make a 27B model feasible without absurd hardware. The use of streaming dataset loading is a practical tell. This repo is built for a 210GB-scale dataset, not for a demo-sized notebook. On the inference side, vllm_inference.py uses multimodal processor settings and LoRARequest so the adapter can be applied cleanly at serving time.

ApproachWhat it does wellWhat it fails to do
Prompt-only multimodal modelDescribes map objects and reads landmarksOften ignores walls, shortcuts through forbidden space, and loses topology
MapTrace-trained multimodal modelLearns semantically valid routes from synthetic supervisionStill depends on training coverage and image quality
Classic pathfinding graphFinds optimal routes on an explicit graphDoes not understand raw images or natural language prompts
Manual annotationCan capture nuance and intentDoes not scale to millions of routes

That comparison is the point. MapTrace does not replace pathfinding, and it does not pretend that language models should magically infer geometry on their own. It teaches a multimodal model to inherit some of pathfinding’s discipline while still reading messy images and answering in language.