troyvis: Troy-VIS: Turning Any Text Prompt Into a Real-Time Video Track

Google Research’s open-vocabulary video segmentation stack pairs language grounding, memory, and deployment-minded engineering so the hard part is not just accuracy, but speed.

10 min read • View on GitHub • More from google-research

A wide editorial scene shows a sequence of video frames laid out like panels on a workbench, while a typed prompt acts like a beam that locks onto the same object across each frame. The image explains Troy-VIS as a system that uses language to select an object and then preserve that object's identity through time.
Troy-VIS is interesting because the prompt is not the end of the story. It is the start of a tracking process that has to survive motion, occlusion, and frame changes.
Key Takeaways

A model that tracks what you name

Most video segmentation systems still feel like closed libraries. They know their classes, they know their limits, and they fail politely when the object is outside the script. Troy-VIS takes a broader bet: give it text, and it should find and track the thing you meant, even if that category was never the point of the original dataset.

Troy-VIS is the first efficient foundation model family for open-vocabulary object perception. It can detect and segment objects of any class in images and track objects of any class in videos.

google-research/troyvis GitHub Repository, Project Documentation · troyvis README

That framing is the real hook. The repo is not promising a nicer demo on a narrow benchmark. It is trying to make open-vocabulary video instance segmentation feel like a practical interface: name the object, then keep that object stable as the video moves.

The trick is not just language, it is time

The central idea is temporal continuity. Troy-VIS uses language to find the object, then uses memory and frame-to-frame propagation to keep finding the same object.

Open-vocabulary is only half the problem. The other half is identity. Once the model has grounded a prompt, it has to preserve that match across motion, blur, and occlusion without re-solving the entire scene from scratch on every frame.

That is why the paper’s speed story matters. The repo points to three ingredients that make the system feel deployable: a decoupled attention feature enhancer, Flash Embedding Memory, and kernel interpolation. Together, they reduce the cost of moving information through the model and exploit the fact that video is not a random sequence of images.

In this paper, we address the challenge of performing open-vocabulary video instance segmentation (OV-VIS) in real-time. We analyze the computational bottlenecks of state-of-the-art foundation models that performs OV-VIS, and propose a new method, TROY-VIS, that significantly improves processing speed while maintaining high accuracy.

Bin Yan, Martin Sundermeyer, David Joseph Tan, Huchuan Lu, Federico Tombari, Authors · TROY-VIS paper

How Troy-VIS is assembled

Under the hood, the repo is organized like a system, not a demo. The core EVAP architecture coordinates a vision backbone, text encoding, feature fusion, and set-based prediction. The codebase supports multiple backbones, including Swin, ResNet, InternImage, and EVA variants, which tells you the team is treating the visual encoder as a tunable component rather than a single fixed bet.

The split between training and deployment is equally important. The deployment variants, including the DPL files, strip away training-only machinery such as the Hungarian matcher and criterion logic so inference stays lean. That is a quiet but meaningful signal: this repo is designed to be run, not just studied.

The text side is not an afterthought either. The model can project CLIP-style embeddings into visual feature space, then fuse them early enough that the prompt influences how the image is interpreted, not just how the final class label is chosen.

Why the data pipeline is part of the product

A close editorial illustration shows several dataset streams pouring from labeled containers into a single training chamber, with some streams entering as still images and others as video clips. One branch is transformed through a small mechanical bridge into pseudo-video before merging, which explains how Troy-VIS learns from heterogeneous sources without losing category consistency.
The model only looks simple once the data has already been normalized. In practice, the repo spends a lot of effort making heterogeneous datasets behave like one training universe.