The Stateful Document: Unpacking the Datalab SDK

How a resilient Python wrapper turns long-running PDF extraction into a predictable, asynchronous pipeline.

8 min read • View on GitHub • More from datalab-to

A massive mechanical printing press feeds a single glowing thread into a glass vault, which outputs neat stacks of paper. This represents the Datalab SDK's checkpoint architecture, converting heavy PDF processing into a reusable cached state.
By caching the intermediate state of a parsed document, the SDK eliminates the need to re-run expensive OCR models for every extraction query.

We decided that we needed to move to full page decoding. This led us to look around at the existing open source models. Unfortunately, nothing quite hit the set of features and accuracy that we needed.

Vik Paruchuri, Co-Founder of Datalab · Introducing our newest model: Chandra
Key Takeaways

The Millisecond Mismatch

Modern AI engineering suffers from a fundamental disconnect. Developers expect API wrappers to return text instantly. Turning complex, unstructured PDFs into clean Markdown requires heavy GPU inference that shatters standard HTTP timeouts. The datalab-to/sdk repository abandons the illusion of instantaneous document parsing. Instead, it treats conversion as a long-running, asynchronous workload that demands a robust polling lifecycle.

The SDK's underlying polling architecture utilizes exponential backoff to monitor long-running inference tasks without overwhelming the API.

Polling for Reality

The heart of the SDK lives in datalab_sdk/client.py. It features a dual-client architecture built around an AsyncDatalabClient that uses aiohttp for non-blocking requests. A synchronous wrapper allows simpler script execution without async overhead. Because document intelligence takes minutes rather than milliseconds, the client relies heavily on the tenacity library. This handles 429 Rate Limit errors and introduces exponential backoff jitter. This jitter prevents the catastrophic thundering herd problem when processing massive document batches.

@retry(
    retry=retry_if_exception_type(RateLimitError),
    wait=wait_exponential_jitter(initial=1, max=60),
    stop=stop_after_attempt(10)
)
async def _poll_result(self, endpoint: str, max_polls: int = 120):

The Checkpoint Paradigm

The most powerful feature of the SDK data models is the checkpoint_id. Traditional API wrappers treat document parsing as a stateless action: you upload a file, get text back, and the server discards the context. If you want to change your extraction prompt, you must re-upload and re-parse the entire 100-page PDF.

Datalab shifts this to a stateful paradigm. The initial upload converts the document and caches its intermediate state. Developers can then run multiple lightweight extraction or segmentation queries against that saved state. This drastically reduces GPU inference time and API costs.

The checkpoint architecture allows multiple distinct processing steps to query a single, cached document state.

Escaping the Pipeline

To understand the SDK architecture, you have to look at the evolution of the underlying models. Datalab co-founder Vik Paruchuri previously built the widely adopted open-source tools Marker and Surya. These tools utilized a pipeline-based approach that segmented pages into blocks before running inference. That approach struggled with highly complex documents containing math, varied layouts, and handwriting.

Hedcut portrait of Datalab Co-Founder Vik Paruchuri.

The result was Chandra, a new OCR model capable of full-page decoding and layout awareness. The sheer computational power required for full-page decoding is exactly why the Python SDK requires such a heavy-duty, state-aware polling mechanism to function reliably at scale.

A side-by-side visual metaphor. Left: tweezers painstakingly dropping individual metal letterpress blocks into a grid. Right: a flat photographic plate capturing an intricate blueprint in a single flash. This illustrates the shift from pipeline-based OCR to full-page decoding.
The shift from block-by-block pipeline processing (left) to layout-aware full-page decoding (right) required fundamentally different API architecture to handle increased inference duration.

JSON as the Orchestrator

The SDK provides a glimpse into its future through the recipes/workflows directory. By acting as a Directed Acyclic Graph executor, the SDK allows developers to chain Conversion, Segmentation, and Extraction steps via a declarative JSON configuration. This moves the repository away from being a simple wrapper and positions it as a complete orchestration engine for document intelligence pipelines.

{
  "name": "Invoice Processing Pipeline",
  "steps": [
    {
      "id": "convert_to_markdown",
      "action": "convert"
    },
    {
      "id": "extract_totals",
      "action": "extract",
      "depends_on": ["convert_to_markdown"]
    }
  ]
}