datalab-to/oss_container: The One-Container Cluster That Turns PDFs Into Markdown at GPU Speed

A FastAPI front end, RabbitMQ queue, and GPU workers all live inside one image, with chunking and merge logic that make large document conversion feel parallel, predictable, and self-hosted.

8 min read • View on GitHub • More from datalab-to

A shipping container opened like a cutaway diorama, with three internal zones: an API desk in front, a conveyor of queued jobs in the middle, and GPU worker bays at the back. PDFs enter on one side and a clean ribbon of Markdown exits on the other, showing how the repository packages distributed processing into a single deployable unit.
The container is the product. API, queue, and workers are collapsed into one image so the system behaves like a cluster without forcing the user to operate a cluster.
Key Takeaways

The PDF Bottleneck Nobody Wants to Talk About

PDF ingestion is where a lot of AI systems quietly slow down. Clean Markdown, JSON, or HTML is rarely the hard part on paper. The hard part is turning a large, ugly document into something structured, fast enough for RAG, evaluation, and batch workflows without turning the infrastructure into a science project.

That is the useful trick in datalab-to/oss_container. It does not just wrap Marker. It packages the boring parts of distributed inference, then hides them behind a single Docker image that feels simple to run.

Datalab Inference Service. Containerized inference service for marker. This is not production-ready, and is only for evaluation purposes.

Datalab Maintainers, Project Maintainers · datalab-to/oss_container README

One Container, Three Jobs

The repository’s shape is the first clue. The API lives in inference/server, the consumer in inference/worker, and the orchestration glue sits in run.sh and supervisord. That is a familiar microservice split, except it ships as one image.

The repo compresses an ordinary distributed stack into one deployable object. The user sees one container, but inside it behave the API, queue, and workers like separate services.

Project roleDeployment modelSelf-hosted?Output focusBest fit
oss_containerSingle Docker image with API, queue, and workersYesMarkdown, JSON, HTMLGPU teams that want a self-contained ingestion appliance
UnstructuredLibrary and platform toolsYesConnectors and preprocessing across many file typesBroader document pipelines that need many formats
NougatResearch modelYesAcademic PDF text extractionSpecialized papers and math-heavy PDFs
Mistral OCRCloud APINoOCR and document understandingTeams that want a strong hosted black box
olmOCROpen-source pipelineYesLarge-scale PDF to MarkdownHigh-volume markdown conversion at dataset scale

That design choice matters because it changes who owns the complexity. The operator does not assemble a queue, a server, and a worker pool by hand. The repo makes that topology feel like a product feature instead of an ops burden.

How the Container Sizes Itself to the GPU

This is the sharpest engineering detail in the repo. run.sh checks available VRAM, computes how many workers it can afford, and then supervises the process tree accordingly. The container does not guess at throughput. It reads the machine and adapts.

# Pseudocode shape of the orchestration
vram_gb=$(get_gpu_vram)
workers=$(( vram_gb / DATALAB_VRAM_PER_WORKER ))
start_mps
exec supervisord -c /etc/supervisor/conf.d/supervisord.conf
A close-up mechanical gauge marked with GPU memory levels, feeding a set of worker slots below it. As the gauge moves across different VRAM amounts, different numbers of worker figures light up, and a switch labeled MPS routes one GPU into multiple parallel lanes. The image explains hardware-aware worker sizing rather than fixed process counts.
Worker count is computed from the GPU, not hardcoded into the image. That makes the container portable across very different machines.

That is not a cosmetic detail. It is what lets the same container behave reasonably on a modest card or a larger enterprise GPU. The repo also uses NVIDIA MPS, which matters because multiple workers on one device only help if the GPU can share attention without thrashing.

Why Chunking Changes the Throughput Story

The next step is even more important: the server rewrites the work graph. In maybe_chunk_pdf, long PDFs are split into smaller jobs before they hit the queue. That means one large input can be processed by several workers at once instead of waiting behind a single serial conversion path.

This is the part that makes the repository feel more like infrastructure than a parser. It changes latency from a single long task into a parallel workflow. That is a very different product promise, especially when the input is a 300-page PDF with awkward tables and mixed layout.

Reassembling the Document Without Losing Structure

Chunking only works if the merge step is equally careful. The repository’s merge.py handles Markdown, JSON, and HTML differently, so the final output is not a naive concatenation job. It stitches content back together in a format-aware way.

# Conceptual shape of the merge layer
if output_format == "markdown":
    merged = "\n".join(chunk_texts)
elif output_format == "json":
    merged = merge_children_nodes(chunk_json_blobs)
elif output_format == "html":
    merged = stitch_body_tags_with_beautifulsoup(chunk_html_docs)
FormatMerge behaviorWhy it matters
MarkdownConcatenates ordered chunk textKeeps the final artifact readable and LLM-friendly
JSONMerges nested document nodesPreserves structure for downstream tools
HTMLStitches body content with DOM-aware parsingAvoids broken markup after parallel processing

That is the hidden constraint in the whole design. Parallelism is only useful if reconstruction is trustworthy. The repo earns its speed by being careful at both ends of the pipeline.

How It Compares to the Rest of Document AI

Against the wider document-AI field, this project is easiest to understand as a packaging win. Unstructured is broader. Nougat is specialized. Mistral OCR is strong but hosted. olmOCR is relevant at scale. oss_container is the one that turns a GPU-backed document pipeline into something a small team can run and reason about locally.

ProjectDeployment modelHardware controlOutput styleWhat it wins on
oss_containerSingle self-hosted containerHighMarkdown, JSON, HTMLPackaging, deployability, parallel chunking
UnstructuredLibrary plus platformMediumMany file types and connectorsBreadth and ecosystem
NougatResearch modelLow to mediumAcademic text extractionMath-heavy PDFs
Mistral OCRHosted APILowOCR and understandingConvenience and model quality
olmOCROpen pipelineMediumMarkdown at scaleDataset-scale conversion

That is why the project feels distinct. It is not trying to beat every alternative on model novelty. It is making the operational shape of document AI much easier to own.

Why This Matters for the Next Wave of RAG Infrastructure

The larger lesson is simple. Document AI is moving from a model-only conversation to an operations conversation. Teams care about self-hosting, throughput, deterministic output, and how much of the pipeline they actually control.

That is where oss_container is interesting. It makes Marker feel like a scalable product while still shipping as one container. The real innovation is not that it can parse a PDF. It is that it can behave like a distributed system without making you run one.