datalab-to/datalab-on-prem: The LLM-ification of Document Parsing
How a handful of Bash scripts orchestrate multi-GPU Vision-Language Models to bring air-gapped, state-of-the-art OCR to the enterprise.
- Modern state-of-the-art OCR is no longer a lightweight CPU task; it requires multi-GPU orchestration and tensor parallelism.
- Datalab-on-prem acts as a highly secure delivery mechanism, using open-source wrapper scripts to deploy proprietary, license-gated models.
- The repository abstracts complex hardware allocation and vLLM parameters into simple environment variables for enterprise deployment.
- Datalab’s strategy relies on open-sourcing pipeline components while commercializing the heavy-duty inference engines for air-gapped environments.
The Tensor Parallelism Surprise
Optical Character Recognition (OCR) used to be simple. You pointed a tool like Tesseract at a scanned image, and the CPU churned out text. But the datalab-on-prem repository reveals a fundamental shift in the industry: state-of-the-art document parsing has become a heavy-duty Large Language Model (LLM) workload.
The surprise lies within the core orchestration script, run-datalab-inference-container.sh. Among the standard Docker flags, you find variables like TENSOR_PARALLEL_SIZE and MAX_NUM_SEQS. These are not parameters for a lightweight library; these are the tuning dials of a vLLM inference engine.
# Excerpt from run-datalab-inference-container.sh
docker run -d \
--name "$CONTAINER_NAME" \
--gpus all \
--ipc=host \
-e DATALAB_LICENSE_KEY="$DATALAB_LICENSE_KEY" \
-e TENSOR_PARALLEL_SIZE="${TENSOR_PARALLEL_SIZE:-1}" \
-e MAX_NUM_SEQS="${MAX_NUM_SEQS:-256}" \
-v "$SERVICE_ACCOUNT_KEY_FILE:/app/sa.json:ro" \
"$IMAGE_NAME"
This configuration indicates that Datalab’s flagship model, Chandra, is a massive Vision-Language Model (VLM). To achieve acceptable latency or even fit the model into VRAM, it must be physically split across multiple GPUs using tensor parallelism. This fundamentally redefines the operational footprint of "running OCR" on-premise.
The License-Gated Delivery Pattern
The repository itself is intentionally lean—a bootstrap kit consisting of a few Bash scripts. However, it operates as a highly secure delivery mechanism for Datalab's proprietary technology. The code here is open, but its utility is zero without a license key and a Google Cloud Service Account.
The script abstracts the complexity of authenticating with the Google Artifact Registry and managing the local Docker daemon. It provides a "SaaS-like" deployment experience on the customer's own hardware, ensuring that the proprietary model weights remain protected while the customer retains full control over their infrastructure.
A Tale of Three Models
The internal routing logic of the scripts reveals a tiered approach to document parsing. Based on the DATALAB_MODEL environment variable, the scripts dynamically switch their target repository, pulling one of three distinct model architectures.
| Tier | Target Workload | Hardware Profile | Underlying Tech |
|---|---|---|---|
| marker | Standard text extraction, high throughput. | CPU or single GPU. Low VRAM. | Legacy pipeline (Surya-based). |
| chandra-small | Balanced performance for typical enterprise PDFs. | Modern NVIDIA GPU (e.g., T4/A10G). | Optimized VLM. |
| chandra | Complex layouts, handwriting, math, maximum accuracy. | Multi-GPU cluster (A100/H100), Tensor Parallelism. | State-of-the-Art VLM. |
This separation allows organizations to choose the right tool for the job. For simple text extraction, the legacy pipeline suffices. But for complex documents, Datalab realized that the pipeline approach—segmenting the page, inferencing blocks separately, and reassembling them—was insufficient.
We decided that we needed to move to full page decoding.
The Air-Gapped Imperative
In a world of cheap cloud APIs from OpenAI and Anthropic, the existence of this repository highlights a critical business reality: medical, legal, and financial institutions cannot legally send unredacted, sensitive PDFs over the public internet.
Competitors like JetBrains' Datalore or Project Nomad address similar needs for "offline-first" AI. However, datalab-on-prem is laser-focused on the specific problem of high-fidelity document parsing. It acts as the bridge that allows these institutions to own the infrastructure while renting the intelligence.
The Commercial Open Source Playbook
The repository is the culmination of a broader strategic playbook. Datalab built immense goodwill and a massive user base (over 50k GitHub stars) with open-source tools like Surya and Marker. These tools solved the "good enough" problem for many developers.
But when the pipeline approach hit its limits, Datalab developed Chandra. By open-sourcing the orchestration layer (this repository) while keeping the massive VLM proprietary, they offer enterprise customers the ultimate combination: the security of an air-gapped, on-premise deployment with the performance of a state-of-the-art model.