datalab-to/marker: The Vision-First End to the PDF Bottleneck
How a custom OCR pipeline and polygon geometry turned thirty years of printer-optimized documents into structured data for AI agents.

Marker converts documents to markdown, JSON, chunks, and HTML quickly and accurately.
- Marker solves the retrieval-augmented generation bottleneck by treating PDFs as images rather than text streams, avoiding standard parsing errors.
- The pipeline relies on custom Surya vision models and polygon geometry to accurately map complex layouts like multi-column text and equations.
- A hybrid architecture uses fast local models for the majority of the work while routing edge cases like multi-page tables to external LLMs.
- By prioritizing raw speed and format versatility, Marker provides a high-performance alternative to enterprise tools like Docling or cloud APIs like LlamaParse.
The Illusion of Text
For thirty years, the Portable Document Format has served a single master: the printer. PDFs were never meant to be read by machines. They lack semantic structure. Instead of containing paragraphs or tables, a PDF is essentially a list of absolute coordinates telling a print head exactly where to drop ink on a page.
This absolute positioning creates a massive bottleneck for modern artificial intelligence. When developers build Retrieval-Augmented Generation (RAG) applications, they usually attempt to extract the underlying text stream from these files. The result is often garbage. Multi-column layouts get scrambled, tables turn into unreadable text dumps, and mathematical equations break entirely. The biggest hurdle in AI today is not the intelligence of the model, but the quality of the data fed into it.
The Vision-First Reversal
Marker solves the text extraction problem by abandoning text extraction entirely. Instead of parsing the PDF text stream, the engine rasterizes the document. It treats every page as an image.
Using the custom-built Surya vision model, Marker looks at the page exactly as a human reader would. It identifies the reading order, multi-column layouts, and tables by analyzing the visual geometry. Once the structure is visually mapped, it performs optical character recognition to reconstruct the document into perfect Markdown.
Polygons and Processors
The architecture relies on a strict pipeline. The DocumentBuilder orchestrates three distinct vision passes. First, the LayoutBuilder identifies where images, tables, and text blocks sit. Next, the LineBuilder sweeps across these blocks to identify individual lines of text. Finally, the OcrBuilder extracts the characters.
Standard parsers use simple rectangular bounding boxes to isolate text. Marker uses PolygonBox geometry. This allows the engine to draw tight, irregular borders around complex elements, easily handling skewed text and overlapping images without capturing unwanted noise.
The Hybrid Escape Hatch
Marker operates primarily on fast, local models to keep compute costs low and throughput high. It can process roughly 25 pages per second on an H100 GPU. However, certain document tasks defy local heuristics. Merging a single financial table that splits across two separate PDF pages is notoriously difficult.
To solve this, Marker includes an explicit flag to route complex edge cases to an external Large Language Model.

For the highest accuracy, pass the `--use_llm` flag to use an LLM alongside marker.
The Open Source Reality Check
The document intelligence space is crowded. IBM Research backs Docling, which excels at native LangChain integration and preserving semantic hierarchy. Proprietary cloud APIs like LlamaParse offer zero-setup convenience for a subscription fee. Marker occupies the high-performance middle ground.
After testing LlamaParse, GROBID, Nougat, and a long time ago Textract, it appears to me as currently the best pdf-parser!
Marker is built for developers who need raw speed, massive format versatility, and the ability to run the entire pipeline locally without sending sensitive documents to a third-party server.
| Feature | Marker | Docling | LlamaParse |
|---|---|---|---|
| Primary Architecture | Vision-first (Surya OCR) | Semantic Hierarchy Models | Proprietary Cloud API |
| Compute Model | Local Hybrid (Optional LLM) | Heavy Local CPU/GPU | Cloud Subscription |
| Core Strength | Raw speed and format versatility | Native LangChain/LlamaIndex integration | Zero-setup ease of use |