Surya and the Vision-First Reconstruction of the Document
How Vik Paruchuri's high-performance toolkit replaced legacy OCR with a multimodal understanding of layout, logic, and reading flow.

Surya is a document OCR toolkit that does: OCR in 90+ languages that benchmarks favorably vs cloud services. Line-level text detection in any language. Layout analysis (table, image, header, etc detection). Reading order detection. Table recognition (detecting rows/columns). LaTeX OCR.
- Surya uses vision transformers to reconstruct document layouts rather than just extracting raw text.
- The system replaces rigid bounding boxes with semantic heatmaps and polygons to handle skewed or complex text.
- Post-processing logic prevents the common generative AI failure of autoregressive hallucination loops.
- The toolkit provides local, open-weight alternatives to proprietary cloud services for multilingual OCR and reading order detection.
The Column-Inch Catastrophe
For decades, optical character recognition was a solved problem with a fatal flaw. Legacy engines like Tesseract were brilliant at identifying individual letters but terrible at understanding how those letters fit together on a page. If fed a modern academic paper with two columns, a sidebar, and a chart, traditional OCR simply reads left to right across the entire page. The result is a flattened, garbled mess. In the era of Retrieval-Augmented Generation (RAG), feeding this "bag of words" to a Large Language Model guarantees hallucination.
This is the "reading order" problem. To an AI agent, text is useless without its spatial context. Surya emerged to solve this by treating optical character recognition not as a text extraction task, but as a vision-first document reconstruction problem.
Beyond the Bounding Box
Surya abandons the rigid rectangular bounding boxes standard in older OCR libraries. Instead, it utilizes an EfficientViT vision transformer to generate high-resolution semantic heatmaps of the document. These heatmaps identify text density and structural elements before a single character is read.
The post-processing layer then converts these glowing probability clouds into precise four-point polygons. By employing dynamic thresholding that calculates the average intensity of the top ten percent of pixels, the system adapts to varying image contrasts seamlessly. This polygon-based approach allows Surya to wrap text that is skewed, warped, or densely packed in scientific notation without clipping neighboring columns.
A Multilingual Vision Backbone
Built on PyTorch and deeply integrated with the Hugging Face transformer ecosystem, Surya is designed for high-performance extraction. The architecture relies on a hybrid model configuration, utilizing a specialized EfficientViT for rapid detection and a Qwen2-based decoder for language comprehension. This foundation gives it an inherent advantage in multilingual contexts.
| Feature | Tesseract | AWS Textract | Surya |
|---|---|---|---|
| Architecture | LSTM / Legacy | Proprietary Cloud | Vision Transformer (EfficientViT) |
| Layout Awareness | Poor | Excellent | Excellent |
| Reading Order | Linear Only | Supported | Native / Core Focus |
| Deployment | Local / Offline | API Only | Local / Open Weights |
The Hallucination Filter
Generative OCR models have a well-known failure mode. Because they are autoregressive, they can occasionally get stuck in a feedback loop, endlessly repeating the same phrase or hallucinating text that does not exist on the page. Surya tackles this directly in its recognition post-processing layer.
The toolkit includes aggressive truncation logic to identify and sever these repetitions. It also features tag balancing to ensure that structured outputs, like LaTeX equations or HTML-style layout tags, are syntactically closed. This cleanup step is what transforms raw transformer output into production-ready data.

The latest version of Surya OCR has a new architecture, and is trained on significantly more data than before.
By open-sourcing the weights under a Rail-M license and providing a robust Python API, the project has democratized access to document intelligence. Developers no longer have to choose between the high latency of cloud APIs and the structural blindness of legacy local tools. Surya proves that vision models are the only reliable way to read a document.