asset-harvester: NVIDIA Asset Harvester: The Pipeline That Turns AV Logs Into 3D Assets
A geometry-aware system for recovering camera pose, synthesizing consistent novel views, and lifting them into Gaussian splats for simulation.
- Asset Harvester is a conversion machine for messy AV logs, not a single image-to-3D model.
- Its central trick is geometry-first generation, which keeps sparse-view synthesis usable for downstream lifting.
- The pipeline matters because each stage solves a different representation problem, from parsing logs to producing Gaussian splats.
- NVIDIA’s bet is that simulation assets should be extracted from driving data, not handcrafted one by one.
The interesting thing about NVIDIA Asset Harvester is not that it can make 3D from images. Plenty of systems can do some version of that. The interesting thing is that it treats autonomous driving logs as raw material for a production pipeline, then solves the geometry problem early so the rest of the stack can stay fast and useful.
The real product is not a model. It is a conversion machine.
The repo is organized like an assembly line. NCore logs are parsed, camera pose is estimated, sparse views are generated, and the result is lifted into 3D Gaussian splats for simulation. That shape matters. It says the target is not a benchmark demo or a research artifact. It is a reusable asset factory for AV workflows.
Why geometry comes first
Asset Harvester does not ask diffusion to improvise in 3D and hope the result lines up later. It conditions generation on camera pose and Plücker ray embeddings, which keeps the synthesized views consistent enough to be lifted into a stable 3D representation. That choice is the whole story.
# Packing variable numbers of views into a fixed batch
# so attention can run without letting padding dominate.
views = _unaggregate(packed_views, mask)
views = transformer(views, camera_pose, plucker_rays)
packed = _aggregate(views, mask)
# The important part is not the helper names.
# It is that the model can handle sparse, irregular view counts
# while still preserving geometric alignment.
The repository’s multiview core, SparseViewDiT, is doing two jobs at once. It generates plausible images and enforces spatial agreement between them. That is why the model can start from one or two sparse observations and still produce a set of views that behave like they belong to the same scene.
The pipeline, stage by stage
The codebase splits the hard problem into smaller ones. First, NCore parsing turns raw logs into structured samples. Then the camera estimator recovers pose and intrinsics. SparseViewDiT synthesizes consistent multiview imagery. TokenGS lifts those views into 3D Gaussian splats. Benchmark scripts close the loop by checking whether the assets are actually usable.
| Workflow | Pose required | Human labor | Speed | Fit for sparse AV logs |
|---|---|---|---|---|
| Manual asset creation | No, but heavy authoring is needed | High | Slow | Poor |
| Traditional reconstruction | Usually yes or strongly helpful | Medium to high | Slow to moderate | Mixed |
| Asset Harvester | Estimated inside the pipeline | Low | Fast enough for an operational loop | Strong |
TokenGS is the part that makes this practical
The reason this stack feels operational instead of academic is the lifting stage. Traditional 3D reconstruction often leans on slow optimization. TokenGS is feed-forward, which means the generated views can become splats in seconds rather than after a long refinement loop. That changes where the bottleneck lives.
In other words, the system does not stop at believable synthesis. It tries to get to a usable 3D object that can be dropped into simulation tooling without a human cleaning up every edge and hole.
Why this stack beats the usual suspects for AV logs
| Approach | Strength | Weakness | Why Asset Harvester differs |
|---|---|---|---|
| Manual meshes | High control | Too much labor for log-scale use | Automates extraction from real driving data |
| NeRF-style pipelines | Strong visual quality | Often slow and optimization-heavy | Uses faster lifting into splats |
| Plain multiview generation | Good-looking outputs | Can drift geometrically | Conditions generation on pose and rays |
| Asset Harvester | End-to-end log-to-asset flow | Depends on estimated geometry quality | Combines parsing, geometry, generation, and lifting |
What the benchmark suggests
The presence of a benchmark suite matters. It signals that NVIDIA is not treating this as a one-off demo for a splashy paper figure. It is building a platform with enough internal structure to evaluate whether the extracted assets really work in the NuRec-AV-Object setting.
That is the strategic read: the future of AV simulation may be less about hand-built worlds and more about harvesting reusable assets from real driving logs. Asset Harvester is a bet on that shift, and the codebase is already shaped like a production-minded answer to it.