The Geometry of Reading: Unpacking pdftext
How an Apache-2.0 mandate forced developers to reconstruct PDF layouts from raw C++ character dumps.
- pdftext bypasses the restrictive AGPL license of PyMuPDF by building a custom spatial geometry engine on top of the permissively licensed pypdfium2.
- The library reconstructs reading order and layout structure from a raw, disorganized soup of geometric character coordinates.
- Custom heuristics like the loosebox toggle allow the engine to correctly map complex ligatures and punctuation without breaking line detection.
- Performance remains competitive by delegating page-by-page character extraction to C++ while orchestrating the process via a Python memory pool.
The AGPL Trap
Extracting text from a PDF is notoriously difficult. Extracting text while preserving the actual reading order, paragraph blocks, and font styles is a monumental task. For years, developers have relied on PyMuPDF for this capability. It works exceptionally well, but it carries a significant catch. The open-source version of PyMuPDF is licensed under the GNU Affero General Public License (AGPL).
For enterprise developers and commercial SaaS products, AGPL is often considered radioactive. It requires any software interacting with it over a network to also release its source code. This licensing trap forced developers to choose between paying hefty commercial licensing fees, writing their own complex layout parsers, or settling for basic text dumpers that destroy document formatting.
Text extraction like PyMuPDF, but without the AGPL license. PDFText extracts plain text or structured blocks and lines. It's built on pypdfium2, so it's fast, accurate, and Apache licensed.
Reassembling the Bag of Characters
To escape the AGPL trap, the team at datalab-to turned to Google's PDFium engine via the Python binding pypdfium2. PDFium is lightning fast and safely licensed under Apache-2.0. However, PDFium does not understand paragraphs. It does not even understand words. It merely reports that a specific glyph exists at a floating-point X and Y coordinate.
This leaves the developer with a disorganized bag of characters. The core achievement of pdftext is its spatial reconstruction pipeline. It acts as a geometric detective, grouping these disconnected coordinates back into a logical human reading order.
The Invisible Math of Page Layouts
The transition from raw coordinates to semantic text happens primarily via heuristics. In late 2024, the project completely replaced its old decision tree with a purely heuristic-based segmentation model. This required writing custom spatial geometry logic in Python.
One of the most fascinating details is the loosebox heuristic. PDFium can provide a tight bounding box around a character or a loose one that includes ascenders and descenders. pdftext dynamically toggles this loosebox calculation based on rotation and specific characters like apostrophes. Without this geometry, a single tall quote mark could trick the engine into thinking a line of text spans multiple vertical rows.
We now remove the `span, line, block` decision tree from the pipeline and use heuristics instead. Also adds a Windows test to CI.
C++ Muscle, Python Orchestration
Python is not known for its blazing speed in mathematical loops. To keep performance competitive with native C libraries, pdftext relies on a clever concurrency model. It uses a ProcessPoolExecutor to map extraction tasks across multiple CPU cores.
The true optimization lies in the worker initialization hook. Instead of repeatedly opening and closing the heavy PDF file for every single page, pdftext pins the document into memory globally for each worker process. The C++ engine handles the raw glyph iteration, and Python orchestrates the final layout assembly.
Where It Fits in the Stack
The document extraction landscape is crowded. Heavyweight vision models like Marker use Large Language Models to achieve near-perfect semantic accuracy, but they require GPUs and significant processing time. Basic libraries like pypdf are fast but completely blind to layout.
pdftext carves out the middle ground. It provides the structured, layout-aware output of premium tools without the licensing headaches or hardware requirements.
| Library | License | Core Engine | Layout Awareness | Speed |
|---|---|---|---|---|
| pdftext | Apache-2.0 | pypdfium2 (C++) | High (Heuristics) | Fast |
| PyMuPDF | AGPL/Commercial | MuPDF (C) | High (Native) | Fastest |
| Marker | GPL | Vision LLMs | Perfect (Semantic) | Slow (GPU) |