The Geometry of Reading: Unpacking pdftext

How an Apache-2.0 mandate forced developers to reconstruct PDF layouts from raw C++ character dumps.

7 min read • View on GitHub • More from datalab-to

A large printing press locked behind an iron gate, while a smaller clockwork mechanism outside rebuilds metal letters on a conveyor belt. This represents rebuilding structured text extraction outside of restrictive AGPL licenses.
Rebuilding the printing press outside the walled garden.
Key Takeaways

The AGPL Trap

Extracting text from a PDF is notoriously difficult. Extracting text while preserving the actual reading order, paragraph blocks, and font styles is a monumental task. For years, developers have relied on PyMuPDF for this capability. It works exceptionally well, but it carries a significant catch. The open-source version of PyMuPDF is licensed under the GNU Affero General Public License (AGPL).

For enterprise developers and commercial SaaS products, AGPL is often considered radioactive. It requires any software interacting with it over a network to also release its source code. This licensing trap forced developers to choose between paying hefty commercial licensing fees, writing their own complex layout parsers, or settling for basic text dumpers that destroy document formatting.

Text extraction like PyMuPDF, but without the AGPL license. PDFText extracts plain text or structured blocks and lines. It's built on pypdfium2, so it's fast, accurate, and Apache licensed.

datalab-to/pdftext README, Project Documentation · Repository: datalab-to/pdftext

Reassembling the Bag of Characters

To escape the AGPL trap, the team at datalab-to turned to Google's PDFium engine via the Python binding pypdfium2. PDFium is lightning fast and safely licensed under Apache-2.0. However, PDFium does not understand paragraphs. It does not even understand words. It merely reports that a specific glyph exists at a floating-point X and Y coordinate.

This leaves the developer with a disorganized bag of characters. The core achievement of pdftext is its spatial reconstruction pipeline. It acts as a geometric detective, grouping these disconnected coordinates back into a logical human reading order.

The Spatial Reconstruction Pipeline groups raw character coordinates into semantic blocks.

The Invisible Math of Page Layouts

The transition from raw coordinates to semantic text happens primarily via heuristics. In late 2024, the project completely replaced its old decision tree with a purely heuristic-based segmentation model. This required writing custom spatial geometry logic in Python.

A close-up of an antique brass caliper measuring the microscopic gap between two metal typography blocks, specifically an 'f' and an 'i'. The caliper dial shows precise geometric coordinates.
Extracting reading order requires precise geometric measurement of character proximity.

One of the most fascinating details is the loosebox heuristic. PDFium can provide a tight bounding box around a character or a loose one that includes ascenders and descenders. pdftext dynamically toggles this loosebox calculation based on rotation and specific characters like apostrophes. Without this geometry, a single tall quote mark could trick the engine into thinking a line of text spans multiple vertical rows.

We now remove the `span, line, block` decision tree from the pipeline and use heuristics instead. Also adds a Windows test to CI.

C++ Muscle, Python Orchestration

Python is not known for its blazing speed in mathematical loops. To keep performance competitive with native C libraries, pdftext relies on a clever concurrency model. It uses a ProcessPoolExecutor to map extraction tasks across multiple CPU cores.

The true optimization lies in the worker initialization hook. Instead of repeatedly opening and closing the heavy PDF file for every single page, pdftext pins the document into memory globally for each worker process. The C++ engine handles the raw glyph iteration, and Python orchestrates the final layout assembly.

Where It Fits in the Stack

The document extraction landscape is crowded. Heavyweight vision models like Marker use Large Language Models to achieve near-perfect semantic accuracy, but they require GPUs and significant processing time. Basic libraries like pypdf are fast but completely blind to layout.

pdftext carves out the middle ground. It provides the structured, layout-aware output of premium tools without the licensing headaches or hardware requirements.

LibraryLicenseCore EngineLayout AwarenessSpeed
pdftextApache-2.0pypdfium2 (C++)High (Heuristics)Fast
PyMuPDFAGPL/CommercialMuPDF (C)High (Native)Fastest
MarkerGPLVision LLMsPerfect (Semantic)Slow (GPU)