inference-mirror: The GPU Document Factory Behind marker
Datalab's wrapper turns PDF-to-Markdown into a parallel pipeline, splitting big documents across workers, sharing GPU capacity efficiently, and reassembling structured output at the end.
- inference-mirror treats a PDF as schedulable work, which is why it scales better than a straight synchronous parser.
- The repo's real advantage is the orchestration around marker, not a new model, because queueing, shared storage, and reassembly keep the pipeline moving.
- NVIDIA MPS turns one GPU into a shared engine, so multiple workers can stay busy without wasting accelerator capacity.
- The project sits between a raw library and a full cluster, removing the glue code teams would otherwise have to build and maintain.
A PDF is not one job
Most PDF pipelines still behave like a single long request: ingest, parse, wait, return. inference-mirror changes the unit of work. A file becomes chunks, chunks become queued jobs, and jobs are farmed out to workers so the GPU stays occupied instead of idle.
That shift matters when documents get large. One 400-page PDF can tie up a naive service for too long, especially if the model has to handle layout detection, OCR, and cleanup in sequence. Here, the boring parts of ingestion become the thing the system is optimized around.
How inference-mirror fans work out and back in
The shape is familiar if you have built any queue-backed service. The server accepts an upload, validates the file, persists it to shared disk, and splits the PDF into page groups or logical chunks. Each chunk becomes a message in RabbitMQ, workers pull those jobs independently, and the results land back on a shared output volume for reassembly in page order.
That shared-volume detail matters. The PDF does not get copied back and forth as a giant blob each time a worker needs it. The server and workers see the same storage, which keeps the network out of the hottest path and makes the whole system feel much more like infrastructure than a wrapper.
NVIDIA MPS is the leverage point. Instead of giving each worker a fully isolated CUDA context and letting them fight over expensive accelerator time, the service can share one GPU more efficiently across multiple processes. In practice, that means more useful work per card, less context-switching overhead, and a better fit for long, messy documents.
From script to service
The repo feels production-minded because the deployment itself is generated, not hand-edited. generate_compose.py builds the Docker Compose topology around the number of GPUs and workers_per_gpu, while the container setup mounts shared volumes so the server and workers can see the same files without extra transport.
The rest of the scaffolding reinforces the same idea. start_mps.sh brings up NVIDIA MPS for shared GPU execution, integration_test.py behaves like a benchmark harness by watching resource usage and pages per second, and .env.example exposes tuning knobs such as recognition and detection batch sizes. This is not just packaging. It is repeatable capacity planning.
Who built it, and why that matters
The project sits inside Datalab's document machine-readability effort, which is a useful clue. This is not a hobby wrapper around marker. It looks like software built by people who needed document ingestion to work reliably at production scale, for the kind of PDFs that matter in retrieval pipelines and model feeding systems.
That context also explains the shape of the code. The repo favors orchestration, shared storage, and throughput over novelty. It is the kind of open source project that solves the slow, unglamorous part of an AI system, then gets out of the way.
What this replaces
| Approach | Setup | Parallelism | GPU efficiency | Large PDFs | Best fit |
|---|---|---|---|---|---|
| Plain marker invocation | One process or one request path | Low | Good for single jobs, weak for saturation | Works, but can block for a long time | Local experiments and quick tests |
| inference-mirror | Dockerized API with RabbitMQ, shared storage, and worker pools | High | Uses MPS to keep one GPU busy across workers | Built to auto-chunk and reassemble | Production ingestion and document throughput |
| Hand-rolled queue plus workers | You build the server, queue, volumes, and retry logic yourself | Variable | Depends on how much tuning you do | Possible, but easy to get wrong | Teams that want maximum control and minimum dependency |
The real win is not that inference-mirror beats marker on algorithmic cleverness. It is that it wraps a strong parser in the runtime shape production wants. That is what most teams actually need: a document pipeline that behaves like a service, not a script.