Chandra and the End of the OCR Assembly Line
How full-page multimodal decoding is replacing twenty years of fragile document pipelines to finally solve the messy reality of handwriting and math.

We decided that we needed to move to full page decoding. This led us to look around at the existing open source models. Unfortunately, nothing quite hit the set of features and accuracy that we needed.
- Chandra replaces fragmented OCR pipelines with a single Vision Language Model that decodes entire pages at once to preserve spatial context.
- A specialized grid-block scaling algorithm preserves the legibility of fine details like mathematical subscripts and handwriting by avoiding destructive image resizing.
- The system enforces structured data output by using a strict HTML schema and providing bounding box coordinates for every extracted element.
- An automated token monitor prevents hallucination loops by dynamically adjusting the model temperature when repetitive text is detected.
The Fragility of the Segmented Page
For twenty years, turning a document into data required a fragile assembly line. The standard Optical Character Recognition (OCR) pipeline works by chopping a page into discrete blocks, reading the text inside those blocks, and then attempting to stitch everything back together. It is a Frankenstein approach to document processing.
This method works perfectly for a clean, single-column novel. It fails catastrophically when confronted with the messy reality of human documents. A handwritten medical form overlapping a printed border, a complex financial table with nested columns, or a dense mathematical proof will shatter a traditional segmentation pipeline. When the pipeline loses the spatial context of the page, it produces a meaningless bag of words. The reading order is destroyed.
Decoding the Whole Canvas
Chandra represents a fundamental paradigm shift. Instead of a pipeline of specialized micro-models, Chandra uses a single Vision Language Model (VLM) to perform full-page decoding. It looks at the document exactly as a human does.
By processing the entire image in one pass, Chandra inherently preserves layout. It understands that a caption belongs to the image above it. It recognizes that a handwritten note in the margin relates to the paragraph beside it. The core abstraction in the repository is a strict HTML schema defined in prompts.py. This forces the VLM to output deterministic, structured data complete with bounding box coordinates for every element.
Scaling the Vision
Feeding high-resolution documents into a VLM presents a massive computational challenge. VLMs have strict sequence length limits. Standard OCR tools simply resize the image to fit, which blurs fine details like subscripts in a mathematical formula or the loops of cursive handwriting.
Chandra solves this with a sophisticated grid-block scaling algorithm found in util.py. Instead of arbitrary resizing, it calculates dimensions in 28-pixel grid blocks. If an image exceeds the model's token limits, a refinement loop analyzes the aspect ratio distortion. It mathematically decides whether trimming width or height blocks will best preserve the legibility of the specific document's strokes.
From Pixels to Production
The architecture of Chandra is designed for scale. The InferenceManager acts as a unified gateway, abstracting the execution engine. Developers can run it locally using HuggingFace Transformers for debugging, or deploy it to a high-throughput vLLM server for production batch processing.
The most brilliant piece of defensive engineering lies in the vLLM integration. VLMs are prone to hallucination loops where they repeat the same phrase endlessly. Chandra includes a token monitor that detects these repetitions. When triggered, it automatically increments the model's temperature to shake the VLM out of its loop, ensuring reliable extraction without human intervention.
The New Gold Standard
This architectural rigor translates directly to performance. Chandra has established a new baseline for what open-source document intelligence can achieve, competing directly with proprietary giants while offering the control required for enterprise deployments.
| Feature | Chandra | Tesseract | olmOCR |
|---|---|---|---|
| Approach | Full Page VLM | Segmented Pipeline | Full Page VLM |
| Output Format | Markdown/HTML/JSON | Plain Text/hOCR | Markdown |
| Handwriting | Excellent | Poor | Good |
| Math (LaTeX) | Native Support | None | High |
| Layout Awareness | Full Bounding Boxes | Basic Blocks | Text-only |
By treating document analysis as a vision problem rather than a text extraction problem, Chandra finally bridges the gap between how a machine reads a page and how a human understands it.