Datalab

Datalab

Developing state of the art document intelligence models.

GitHub 11 repos 936 followers

Explained projects

datalab-to/oss_container: The One-Container Cluster That Turns PDFs Into Markdown at GPU Speed
datalab-to/oss_container: The One-Container Cluster That Turns PDFs Into Markdown at GPU Speed
A FastAPI front end, RabbitMQ queue, and GPU workers all live inside one image, with chunking and merge logic that make large document conversion feel parallel, predictable, and self-hosted.
8 min read · Aug 5, 2026
inference-mirror: The GPU Document Factory Behind marker
inference-mirror: The GPU Document Factory Behind marker
Datalab's wrapper turns PDF-to-Markdown into a parallel pipeline, splitting big documents across workers, sharing GPU capacity efficiently, and reassembling structured output at the end.
7 min read · Apr 6, 2026
The Stateful Document: Unpacking the Datalab SDK
The Stateful Document: Unpacking the Datalab SDK
How a resilient Python wrapper turns long-running PDF extraction into a predictable, asynchronous pipeline.
8 min read · Mar 31, 2026
The Geometry of Reading: Unpacking pdftext
The Geometry of Reading: Unpacking pdftext
How an Apache-2.0 mandate forced developers to reconstruct PDF layouts from raw C++ character dumps.
7 min read · Mar 31, 2026
datalab-to/datalab-on-prem: The LLM-ification of Document Parsing
datalab-to/datalab-on-prem: The LLM-ification of Document Parsing
How a handful of Bash scripts orchestrate multi-GPU Vision-Language Models to bring air-gapped, state-of-the-art OCR to the enterprise.
6 min read · Mar 30, 2026
datalab-to/marker: The Vision-First End to the PDF Bottleneck
datalab-to/marker: The Vision-First End to the PDF Bottleneck
How a custom OCR pipeline and polygon geometry turned thirty years of printer-optimized documents into structured data for AI agents.
8 min read · Mar 30, 2026
Surya and the Vision-First Reconstruction of the Document
Surya and the Vision-First Reconstruction of the Document
How Vik Paruchuri's high-performance toolkit replaced legacy OCR with a multimodal understanding of layout, logic, and reading flow.
· Mar 28, 2026
Chandra and the End of the OCR Assembly Line
Chandra and the End of the OCR Assembly Line
How full-page multimodal decoding is replacing twenty years of fragile document pipelines to finally solve the messy reality of handwriting and math.
· Mar 26, 2026

Yet to be explained

lift
Extract structured data from documents quickly and accurately.
Python1.2k stars
Explain
omni_extract_bench
A consistent, fair, benchmark for structured extraction.
Python26 stars
Explain
pykatex
Python4 stars
Explain