Chandra and the End of the OCR Assembly Line

How full-page multimodal decoding is replacing twenty years of fragile document pipelines to finally solve the messy reality of handwriting and math.

datalab-to/chandra

An eye scanning a massive blueprint, outputting a clean scroll of text, representing a Vision Language Model reading a full document at once.
Chandra treats the entire page as a single visual canvas, bypassing the fragile segmentation steps of traditional OCR.

We decided that we needed to move to full page decoding. This led us to look around at the existing open source models. Unfortunately, nothing quite hit the set of features and accuracy that we needed.

Vik Paruchuri, Author · Introducing our newest model: Chandra

Key Takeaways

The Fragility of the Segmented Page

For twenty years, turning a document into data required a fragile assembly line. The standard Optical Character Recognition (OCR) pipeline works by chopping a page into discrete blocks, reading the text inside those blocks, and then attempting to stitch everything back together. It is a Frankenstein approach to document processing.

This method works perfectly for a clean, single-column novel. It fails catastrophically when confronted with the messy reality of human documents. A handwritten medical form overlapping a printed border, a complex financial table with nested columns, or a dense mathematical proof will shatter a traditional segmentation pipeline. When the pipeline loses the spatial context of the page, it produces a meaningless bag of words. The reading order is destroyed.

A conveyor belt where a document is chopped into cubes by robot arms, and a worker fails to glue them back together.
Traditional OCR pipelines destroy spatial context by segmenting the page before reading it.

Decoding the Whole Canvas

Chandra represents a fundamental paradigm shift. Instead of a pipeline of specialized micro-models, Chandra uses a single Vision Language Model (VLM) to perform full-page decoding. It looks at the document exactly as a human does.

Portrait of Vik Paruchuri.

By processing the entire image in one pass, Chandra inherently preserves layout. It understands that a caption belongs to the image above it. It recognizes that a handwritten note in the margin relates to the paragraph beside it. The core abstraction in the repository is a strict HTML schema defined in prompts.py. This forces the VLM to output deterministic, structured data complete with bounding box coordinates for every element.

Full-page decoding bypasses the error-prone segmentation step entirely.

Scaling the Vision

Feeding high-resolution documents into a VLM presents a massive computational challenge. VLMs have strict sequence length limits. Standard OCR tools simply resize the image to fit, which blurs fine details like subscripts in a mathematical formula or the loops of cursive handwriting.

Chandra solves this with a sophisticated grid-block scaling algorithm found in util.py. Instead of arbitrary resizing, it calculates dimensions in 28-pixel grid blocks. If an image exceeds the model's token limits, a refinement loop analyzes the aspect ratio distortion. It mathematically decides whether trimming width or height blocks will best preserve the legibility of the specific document's strokes.

From Pixels to Production

The architecture of Chandra is designed for scale. The InferenceManager acts as a unified gateway, abstracting the execution engine. Developers can run it locally using HuggingFace Transformers for debugging, or deploy it to a high-throughput vLLM server for production batch processing.

The most brilliant piece of defensive engineering lies in the vLLM integration. VLMs are prone to hallucination loops where they repeat the same phrase endlessly. Chandra includes a token monitor that detects these repetitions. When triggered, it automatically increments the model's temperature to shake the VLM out of its loop, ensuring reliable extraction without human intervention.

Chandra detects hallucination loops and dynamically adjusts temperature to recover.

The New Gold Standard

This architectural rigor translates directly to performance. Chandra has established a new baseline for what open-source document intelligence can achieve, competing directly with proprietary giants while offering the control required for enterprise deployments.

FeatureChandraTesseractolmOCR
ApproachFull Page VLMSegmented PipelineFull Page VLM
Output FormatMarkdown/HTML/JSONPlain Text/hOCRMarkdown
HandwritingExcellentPoorGood
Math (LaTeX)Native SupportNoneHigh
Layout AwarenessFull Bounding BoxesBasic BlocksText-only

By treating document analysis as a vision problem rather than a text extraction problem, Chandra finally bridges the gap between how a machine reads a page and how a human understands it.