datalab-to/marker: The Vision-First End to the PDF Bottleneck

How a custom OCR pipeline and polygon geometry turned thirty years of printer-optimized documents into structured data for AI agents.

8 min read • View on GitHub • More from datalab-to

A chaotic mountain of paper documents being fed into a glowing prism, emerging as perfectly aligned geometric blocks of code. This illustrates Marker turning unstructured legacy PDFs into structured machine data.
Marker solves the data ingestion bottleneck by treating legacy documents as visual puzzles rather than text streams.

Marker converts documents to markdown, JSON, chunks, and HTML quickly and accurately.

Vik Paruchuri, Primary Maintainer, Datalab · datalab-to/marker
Key Takeaways

The Illusion of Text

For thirty years, the Portable Document Format has served a single master: the printer. PDFs were never meant to be read by machines. They lack semantic structure. Instead of containing paragraphs or tables, a PDF is essentially a list of absolute coordinates telling a print head exactly where to drop ink on a page.

This absolute positioning creates a massive bottleneck for modern artificial intelligence. When developers build Retrieval-Augmented Generation (RAG) applications, they usually attempt to extract the underlying text stream from these files. The result is often garbage. Multi-column layouts get scrambled, tables turn into unreadable text dumps, and mathematical equations break entirely. The biggest hurdle in AI today is not the intelligence of the model, but the quality of the data fed into it.

Hedcut portrait of Vik Paruchuri, creator of Marker.

The Vision-First Reversal

Marker solves the text extraction problem by abandoning text extraction entirely. Instead of parsing the PDF text stream, the engine rasterizes the document. It treats every page as an image.

Using the custom-built Surya vision model, Marker looks at the page exactly as a human reader would. It identifies the reading order, multi-column layouts, and tables by analyzing the visual geometry. Once the structure is visually mapped, it performs optical character recognition to reconstruct the document into perfect Markdown.

A split screen comparing a tangled ribbon of ticker tape to mechanical eyes scanning an architectural grid. This shows the difference between traditional text-stream extraction and Marker's top-down structural approach.
Traditional parsers read sequentially and get lost. Marker maps the visual hierarchy first.

Polygons and Processors

The architecture relies on a strict pipeline. The DocumentBuilder orchestrates three distinct vision passes. First, the LayoutBuilder identifies where images, tables, and text blocks sit. Next, the LineBuilder sweeps across these blocks to identify individual lines of text. Finally, the OcrBuilder extracts the characters.

Standard parsers use simple rectangular bounding boxes to isolate text. Marker uses PolygonBox geometry. This allows the engine to draw tight, irregular borders around complex elements, easily handling skewed text and overlapping images without capturing unwanted noise.

The three-stage Builder pipeline isolates layout elements before attempting character recognition.

A close-up of a magnifying glass over a mathematical equation, with an irregular polygon drawn tightly around it to avoid an overlapping photograph. This illustrates why polygon geometry outperforms rectangular bounding boxes.
Irregular polygon mapping prevents overlapping elements from corrupting the extracted text.

The Hybrid Escape Hatch

Marker operates primarily on fast, local models to keep compute costs low and throughput high. It can process roughly 25 pages per second on an H100 GPU. However, certain document tasks defy local heuristics. Merging a single financial table that splits across two separate PDF pages is notoriously difficult.

To solve this, Marker includes an explicit flag to route complex edge cases to an external Large Language Model.

For the highest accuracy, pass the `--use_llm` flag to use an LLM alongside marker.

Vik Paruchuri, Primary Maintainer, Datalab · datalab-to/marker

The hybrid routing system allows external LLMs to stitch together fragmented data blocks that local heuristics miss.

The Open Source Reality Check

The document intelligence space is crowded. IBM Research backs Docling, which excels at native LangChain integration and preserving semantic hierarchy. Proprietary cloud APIs like LlamaParse offer zero-setup convenience for a subscription fee. Marker occupies the high-performance middle ground.

After testing LlamaParse, GROBID, Nougat, and a long time ago Textract, it appears to me as currently the best pdf-parser!

relsas, User/Contributor · V2 is a huge improvement · Issue #121

Marker is built for developers who need raw speed, massive format versatility, and the ability to run the entire pipeline locally without sending sensitive documents to a third-party server.

FeatureMarkerDoclingLlamaParse
Primary ArchitectureVision-first (Surya OCR)Semantic Hierarchy ModelsProprietary Cloud API
Compute ModelLocal Hybrid (Optional LLM)Heavy Local CPU/GPUCloud Subscription
Core StrengthRaw speed and format versatilityNative LangChain/LlamaIndex integrationZero-setup ease of use