asset-harvester: NVIDIA Asset Harvester: The Pipeline That Turns AV Logs Into 3D Assets

A geometry-aware system for recovering camera pose, synthesizing consistent novel views, and lifting them into Gaussian splats for simulation.

10 to 12 min read • View on GitHub • More from NVIDIA

A roadside driving scene is transformed into a clean 3D asset pipeline. On one side, a sparse autonomous vehicle log shows an occluded car, a pedestrian, and scattered roadside clutter. On the other side, the same scene appears as a compact cluster of Gaussian splats ready for simulation. The image explains that the system is about converting messy capture into reusable simulation assets, not just generating pretty views.
Asset Harvester’s core move is operational, not artistic: turn imperfect driving logs into simulation-ready 3D assets.
Key Takeaways

The interesting thing about NVIDIA Asset Harvester is not that it can make 3D from images. Plenty of systems can do some version of that. The interesting thing is that it treats autonomous driving logs as raw material for a production pipeline, then solves the geometry problem early so the rest of the stack can stay fast and useful.

The real product is not a model. It is a conversion machine.

The repo is organized like an assembly line. NCore logs are parsed, camera pose is estimated, sparse views are generated, and the result is lifted into 3D Gaussian splats for simulation. That shape matters. It says the target is not a benchmark demo or a research artifact. It is a reusable asset factory for AV workflows.

Asset Harvester is a staged system. Each module solves a different representation problem before handing off to the next.

Why geometry comes first

Asset Harvester does not ask diffusion to improvise in 3D and hope the result lines up later. It conditions generation on camera pose and Plücker ray embeddings, which keeps the synthesized views consistent enough to be lifted into a stable 3D representation. That choice is the whole story.

# Packing variable numbers of views into a fixed batch
# so attention can run without letting padding dominate.
views = _unaggregate(packed_views, mask)
views = transformer(views, camera_pose, plucker_rays)
packed = _aggregate(views, mask)

# The important part is not the helper names.
# It is that the model can handle sparse, irregular view counts
# while still preserving geometric alignment.

The repository’s multiview core, SparseViewDiT, is doing two jobs at once. It generates plausible images and enforces spatial agreement between them. That is why the model can start from one or two sparse observations and still produce a set of views that behave like they belong to the same scene.

A close-up of a transformer block receiving image tokens and geometric rays at the same time. Sparse input views are packed through a masked batch on one side, while Plücker ray lines thread through the block and emerge as 16 aligned camera views on the other side. The image explains that geometry is used as a constraint, not decoration.
SparseViewDiT binds generation to geometry, which is what makes the downstream 3D lifting practical.

The pipeline, stage by stage

The codebase splits the hard problem into smaller ones. First, NCore parsing turns raw logs into structured samples. Then the camera estimator recovers pose and intrinsics. SparseViewDiT synthesizes consistent multiview imagery. TokenGS lifts those views into 3D Gaussian splats. Benchmark scripts close the loop by checking whether the assets are actually usable.

WorkflowPose requiredHuman laborSpeedFit for sparse AV logs
Manual asset creationNo, but heavy authoring is neededHighSlowPoor
Traditional reconstructionUsually yes or strongly helpfulMedium to highSlow to moderateMixed
Asset HarvesterEstimated inside the pipelineLowFast enough for an operational loopStrong

TokenGS is the part that makes this practical

The reason this stack feels operational instead of academic is the lifting stage. Traditional 3D reconstruction often leans on slow optimization. TokenGS is feed-forward, which means the generated views can become splats in seconds rather than after a long refinement loop. That changes where the bottleneck lives.

In other words, the system does not stop at believable synthesis. It tries to get to a usable 3D object that can be dropped into simulation tooling without a human cleaning up every edge and hole.

Why this stack beats the usual suspects for AV logs

ApproachStrengthWeaknessWhy Asset Harvester differs
Manual meshesHigh controlToo much labor for log-scale useAutomates extraction from real driving data
NeRF-style pipelinesStrong visual qualityOften slow and optimization-heavyUses faster lifting into splats
Plain multiview generationGood-looking outputsCan drift geometricallyConditions generation on pose and rays
Asset HarvesterEnd-to-end log-to-asset flowDepends on estimated geometry qualityCombines parsing, geometry, generation, and lifting

What the benchmark suggests

The presence of a benchmark suite matters. It signals that NVIDIA is not treating this as a one-off demo for a splashy paper figure. It is building a platform with enough internal structure to evaluate whether the extracted assets really work in the NuRec-AV-Object setting.

That is the strategic read: the future of AV simulation may be less about hand-built worlds and more about harvesting reusable assets from real driving logs. Asset Harvester is a bet on that shift, and the codebase is already shaped like a production-minded answer to it.