datalab-to/datalab-on-prem: The LLM-ification of Document Parsing

How a handful of Bash scripts orchestrate multi-GPU Vision-Language Models to bring air-gapped, state-of-the-art OCR to the enterprise.

6 min read • View on GitHub • More from datalab-to

A massive industrial printing press being operated by a sophisticated robotic arm, all contained within a thick, pristine glass vault. This illustrates the concept of heavy-duty modern AI models operating within a secure, air-gapped environment.
The modern OCR workload requires massive infrastructure, securely contained.
Key Takeaways

The Tensor Parallelism Surprise

Optical Character Recognition (OCR) used to be simple. You pointed a tool like Tesseract at a scanned image, and the CPU churned out text. But the datalab-on-prem repository reveals a fundamental shift in the industry: state-of-the-art document parsing has become a heavy-duty Large Language Model (LLM) workload.

The surprise lies within the core orchestration script, run-datalab-inference-container.sh. Among the standard Docker flags, you find variables like TENSOR_PARALLEL_SIZE and MAX_NUM_SEQS. These are not parameters for a lightweight library; these are the tuning dials of a vLLM inference engine.

# Excerpt from run-datalab-inference-container.sh
docker run -d \
  --name "$CONTAINER_NAME" \
  --gpus all \
  --ipc=host \
  -e DATALAB_LICENSE_KEY="$DATALAB_LICENSE_KEY" \
  -e TENSOR_PARALLEL_SIZE="${TENSOR_PARALLEL_SIZE:-1}" \
  -e MAX_NUM_SEQS="${MAX_NUM_SEQS:-256}" \
  -v "$SERVICE_ACCOUNT_KEY_FILE:/app/sa.json:ro" \
  "$IMAGE_NAME"

This configuration indicates that Datalab’s flagship model, Chandra, is a massive Vision-Language Model (VLM). To achieve acceptable latency or even fit the model into VRAM, it must be physically split across multiple GPUs using tensor parallelism. This fundamentally redefines the operational footprint of "running OCR" on-premise.

How environment variables dictate the architectural footprint of the container.

The License-Gated Delivery Pattern

The repository itself is intentionally lean—a bootstrap kit consisting of a few Bash scripts. However, it operates as a highly secure delivery mechanism for Datalab's proprietary technology. The code here is open, but its utility is zero without a license key and a Google Cloud Service Account.

The script abstracts the complexity of authenticating with the Google Artifact Registry and managing the local Docker daemon. It provides a "SaaS-like" deployment experience on the customer's own hardware, ensuring that the proprietary model weights remain protected while the customer retains full control over their infrastructure.

A Tale of Three Models

The internal routing logic of the scripts reveals a tiered approach to document parsing. Based on the DATALAB_MODEL environment variable, the scripts dynamically switch their target repository, pulling one of three distinct model architectures.

TierTarget WorkloadHardware ProfileUnderlying Tech
markerStandard text extraction, high throughput.CPU or single GPU. Low VRAM.Legacy pipeline (Surya-based).
chandra-smallBalanced performance for typical enterprise PDFs.Modern NVIDIA GPU (e.g., T4/A10G).Optimized VLM.
chandraComplex layouts, handwriting, math, maximum accuracy.Multi-GPU cluster (A100/H100), Tensor Parallelism.State-of-the-Art VLM.

This separation allows organizations to choose the right tool for the job. For simple text extraction, the legacy pipeline suffices. But for complex documents, Datalab realized that the pipeline approach—segmenting the page, inferencing blocks separately, and reassembling them—was insufficient.

We decided that we needed to move to full page decoding.

Vik Paruchuri, Creator · Introducing our newest model: Chandra

The Air-Gapped Imperative

In a world of cheap cloud APIs from OpenAI and Anthropic, the existence of this repository highlights a critical business reality: medical, legal, and financial institutions cannot legally send unredacted, sensitive PDFs over the public internet.

An extreme close-up of a cleanly severed fiber-optic cable. The right side is dark, while the left side has rooted into mechanical soil, sprouting a self-contained clockwork ecosystem. This represents an air-gapped system operating independently of the internet.
Air-gapped systems must thrive entirely disconnected from the broader internet.

Competitors like JetBrains' Datalore or Project Nomad address similar needs for "offline-first" AI. However, datalab-on-prem is laser-focused on the specific problem of high-fidelity document parsing. It acts as the bridge that allows these institutions to own the infrastructure while renting the intelligence.

The Commercial Open Source Playbook

The repository is the culmination of a broader strategic playbook. Datalab built immense goodwill and a massive user base (over 50k GitHub stars) with open-source tools like Surya and Marker. These tools solved the "good enough" problem for many developers.

But when the pipeline approach hit its limits, Datalab developed Chandra. By open-sourcing the orchestration layer (this repository) while keeping the massive VLM proprietary, they offer enterprise customers the ultimate combination: the security of an air-gapped, on-premise deployment with the performance of a state-of-the-art model.