inference-mirror: The GPU Document Factory Behind marker

Datalab's wrapper turns PDF-to-Markdown into a parallel pipeline, splitting big documents across workers, sharing GPU capacity efficiently, and reassembling structured output at the end.

7 min read • View on GitHub • More from datalab-to

A single oversized PDF enters a document factory as a heavy slab, then gets cut into smaller stacks and routed to several worker stations before becoming a clean Markdown scroll on the far side. The scene explains the article's central idea: one document can be broken into many jobs so the system can work in parallel instead of waiting on one long parse.
The core move is simple to say and hard to engineer: one document becomes many scheduled tasks, then returns as one ordered result.
Key Takeaways

A PDF is not one job

Most PDF pipelines still behave like a single long request: ingest, parse, wait, return. inference-mirror changes the unit of work. A file becomes chunks, chunks become queued jobs, and jobs are farmed out to workers so the GPU stays occupied instead of idle.

This diagram shows the repo's real insight: the expensive part is not just parsing a PDF, it is scheduling the work so the hardware stays full.

That shift matters when documents get large. One 400-page PDF can tie up a naive service for too long, especially if the model has to handle layout detection, OCR, and cleanup in sequence. Here, the boring parts of ingestion become the thing the system is optimized around.

How inference-mirror fans work out and back in

The shape is familiar if you have built any queue-backed service. The server accepts an upload, validates the file, persists it to shared disk, and splits the PDF into page groups or logical chunks. Each chunk becomes a message in RabbitMQ, workers pull those jobs independently, and the results land back on a shared output volume for reassembly in page order.

That shared-volume detail matters. The PDF does not get copied back and forth as a giant blob each time a worker needs it. The server and workers see the same storage, which keeps the network out of the hottest path and makes the whole system feel much more like infrastructure than a wrapper.

A close view of one GPU drawn like a shared engine, with several worker pipes feeding into it through a manifold and pressure gauges showing load balance. It explains why MPS matters here, because several processes can share one accelerator more efficiently than a single process hogging the device.
The hardware trick is not glamourous, but it is the difference between a GPU that looks busy and a GPU that actually is busy.

NVIDIA MPS is the leverage point. Instead of giving each worker a fully isolated CUDA context and letting them fight over expensive accelerator time, the service can share one GPU more efficiently across multiple processes. In practice, that means more useful work per card, less context-switching overhead, and a better fit for long, messy documents.

From script to service

The repo feels production-minded because the deployment itself is generated, not hand-edited. generate_compose.py builds the Docker Compose topology around the number of GPUs and workers_per_gpu, while the container setup mounts shared volumes so the server and workers can see the same files without extra transport.

The rest of the scaffolding reinforces the same idea. start_mps.sh brings up NVIDIA MPS for shared GPU execution, integration_test.py behaves like a benchmark harness by watching resource usage and pages per second, and .env.example exposes tuning knobs such as recognition and detection batch sizes. This is not just packaging. It is repeatable capacity planning.

Who built it, and why that matters

The project sits inside Datalab's document machine-readability effort, which is a useful clue. This is not a hobby wrapper around marker. It looks like software built by people who needed document ingestion to work reliably at production scale, for the kind of PDFs that matter in retrieval pipelines and model feeding systems.

That context also explains the shape of the code. The repo favors orchestration, shared storage, and throughput over novelty. It is the kind of open source project that solves the slow, unglamorous part of an AI system, then gets out of the way.

What this replaces

ApproachSetupParallelismGPU efficiencyLarge PDFsBest fit
Plain marker invocationOne process or one request pathLowGood for single jobs, weak for saturationWorks, but can block for a long timeLocal experiments and quick tests
inference-mirrorDockerized API with RabbitMQ, shared storage, and worker poolsHighUses MPS to keep one GPU busy across workersBuilt to auto-chunk and reassembleProduction ingestion and document throughput
Hand-rolled queue plus workersYou build the server, queue, volumes, and retry logic yourselfVariableDepends on how much tuning you doPossible, but easy to get wrongTeams that want maximum control and minimum dependency

The real win is not that inference-mirror beats marker on algorithmic cleverness. It is that it wraps a strong parser in the runtime shape production wants. That is what most teams actually need: a document pipeline that behaves like a service, not a script.