Surya and the Vision-First Reconstruction of the Document

How Vik Paruchuri's high-performance toolkit replaced legacy OCR with a multimodal understanding of layout, logic, and reading flow.

datalab-to/surya

A giant mechanical eye hovers over a chaotic stack of translucent floating documents, projecting a grid that snaps the pages into a perfectly ordered architectural blueprint.
Surya treats document understanding as a holistic vision task, rather than just extracting a bag of words.

Surya is a document OCR toolkit that does: OCR in 90+ languages that benchmarks favorably vs cloud services. Line-level text detection in any language. Layout analysis (table, image, header, etc detection). Reading order detection. Table recognition (detecting rows/columns). LaTeX OCR.

Vik Paruchuri, Author · Repository: datalab-to/surya

Key Takeaways

The Column-Inch Catastrophe

For decades, optical character recognition was a solved problem with a fatal flaw. Legacy engines like Tesseract were brilliant at identifying individual letters but terrible at understanding how those letters fit together on a page. If fed a modern academic paper with two columns, a sidebar, and a chart, traditional OCR simply reads left to right across the entire page. The result is a flattened, garbled mess. In the era of Retrieval-Augmented Generation (RAG), feeding this "bag of words" to a Large Language Model guarantees hallucination.

This is the "reading order" problem. To an AI agent, text is useless without its spatial context. Surya emerged to solve this by treating optical character recognition not as a text extraction task, but as a vision-first document reconstruction problem.

A split scene comparing traditional OCR to Surya. On the left, a person reads a scroll through a tiny vertical slit. On the right, a person looks down at a massive illuminated map of the entire scroll.
Traditional OCR engines lack the spatial awareness required to parse complex multi-column layouts.

How Surya determines the sequence of text blocks to preserve logical reading order.

Beyond the Bounding Box

Surya abandons the rigid rectangular bounding boxes standard in older OCR libraries. Instead, it utilizes an EfficientViT vision transformer to generate high-resolution semantic heatmaps of the document. These heatmaps identify text density and structural elements before a single character is read.

The post-processing layer then converts these glowing probability clouds into precise four-point polygons. By employing dynamic thresholding that calculates the average intensity of the top ten percent of pixels, the system adapts to varying image contrasts seamlessly. This polygon-based approach allows Surya to wrap text that is skewed, warped, or densely packed in scientific notation without clipping neighboring columns.

The transformation from raw model output (heatmap) to precise geometric data (polygons).

A Multilingual Vision Backbone

Built on PyTorch and deeply integrated with the Hugging Face transformer ecosystem, Surya is designed for high-performance extraction. The architecture relies on a hybrid model configuration, utilizing a specialized EfficientViT for rapid detection and a Qwen2-based decoder for language comprehension. This foundation gives it an inherent advantage in multilingual contexts.

An ink portrait of Vik Paruchuri in a WSJ hedcut style.
FeatureTesseractAWS TextractSurya
ArchitectureLSTM / LegacyProprietary CloudVision Transformer (EfficientViT)
Layout AwarenessPoorExcellentExcellent
Reading OrderLinear OnlySupportedNative / Core Focus
DeploymentLocal / OfflineAPI OnlyLocal / Open Weights

The Hallucination Filter

Generative OCR models have a well-known failure mode. Because they are autoregressive, they can occasionally get stuck in a feedback loop, endlessly repeating the same phrase or hallucinating text that does not exist on the page. Surya tackles this directly in its recognition post-processing layer.

A typewriter ribbon tangled into an infinite loop, about to be cut precisely by a pair of sharp ink scissors.
Surya's post-processing logic acts as a strict cutoff for the autoregressive loops common in transformer-based OCR.

The toolkit includes aggressive truncation logic to identify and sever these repetitions. It also features tag balancing to ensure that structured outputs, like LaTeX equations or HTML-style layout tags, are syntactically closed. This cleanup step is what transforms raw transformer output into production-ready data.

The latest version of Surya OCR has a new architecture, and is trained on significantly more data than before.

By open-sourcing the weights under a Rail-M license and providing a robust Python API, the project has democratized access to document intelligence. Developers no longer have to choose between the high latency of cloud APIs and the structural blindness of legacy local tools. Surya proves that vision models are the only reliable way to read a document.