Beyond the Pixel: Reconstructing the Spine of a Book with pdf-craft

How a local-first pipeline uses DeepSeek's vision intelligence to transform scanned chaos into structured Markdown without the LLM "correction tax."

8 min read • View on GitHub • More from oomol-lab

A vintage printing press stamping glowing geometric wireframes onto a clean digital tablet.
pdf-craft reconstructs the digital structure of a physical book by analyzing geometric page layouts.
Key Takeaways

The Death of the Correction Loop

The standard playbook for processing complex PDF documents has always been a clunky two-step dance. First, you run a basic Optical Character Recognition engine to extract raw text coordinates. Second, you pass that garbled, fragmented text to a Large Language Model and beg it to fix the formatting.

This approach is slow, expensive, and leaks private data to server-side APIs. The pdf-craft project outright eliminates this cycle. By integrating DeepSeek's Vision-Language Models directly into a local Python pipeline, it treats document conversion not as text recognition, but as semantic reconstruction.

"Version 1.0.0 represents a fundamental shift in pdf-craft's text recognition strategy. Prior to this release, the system used a two-stage approach: basic OCR followed by LLM-based text correction. The new architecture eliminates the correction stage entirely, relying on DeepSeek OCR's native accuracy to produce final text output."

Split screen showing a messy Rube Goldberg machine for OCR correction versus a single sleek mechanical eye outputting a perfect book.
The architectural pivot removed the expensive LLM correction tax in favor of native vision intelligence.

Seeing the Gutter: Geometric Document Analysis

A flat image of a page does not natively understand columns or reading order. If you naively read a two-column academic paper from left to right, the text becomes meaningless. The library solves this using a module called reading_serials.py.

The code uses X-axis projections to build a weighted histogram of the page. It looks for the "valleys" in the histogram to identify the gutters of empty white space between columns. This geometric approach allows pdf-craft to slice the page into logical serials before attempting to read the words.

A two-column academic paper is shown as a flat image. Below it

The Jointer: Stitching the Narrative Thread

Books have page breaks, images that interrupt paragraphs, and floating tables. A standard OCR engine treats every page as an isolated island. The jointer.py module acts as the bridge.

It maintains a stateful "tail" pattern. When a paragraph ends abruptly at the bottom of a page, the jointer holds that incomplete sentence in memory. As the next page is processed, the jointer snaps the held tail directly onto the head of the new sequence, resulting in a continuous, unbroken Markdown paragraph.

Two separate blocks of text represent the bottom of page one and the top of page two. An incomplete sentence at the bottom of the first block glows softly. A visual connection line pulls the top of the second block upward

Math as a First-Class Citizen

Most open-source extraction tools fail spectacularly when they encounter LaTeX formulas. They mangle equations into random alphanumeric strings. pdf-craft treats mathematics as an atomic unit.

The expression.py parser distinguishes between standard text and LaTeX delimiters. Once it identifies a mathematical expression, it locks the content within those boundaries. When the surrounding text is reflowed or joined, the mathematical integrity is perfectly preserved.

A complex LaTeX equation being held safely by digital tweezers while surrounding text flows like water.
Mathematical expressions are parsed and protected as atomic units during structural reflowing.

Scaling the "Gundam"

Running high-end vision models locally requires serious resource management. The page_ref.py module calculates the maximum allowable DPI based on a target file size to prevent memory crashes on heavy architectural drawings.

The project offers multiple tiers of the DeepSeek model, allowing developers to trade speed for accuracy depending on their available hardware. The "Gundam" tier provides the highest semantic accuracy, but the smaller models offer rapid iteration for CPU-bound environments.

Tool Architecture Formula Support Privacy
pdf-craft Local Vision-Language Model High (Native LaTeX) 100% Offline
Stirling-PDF Server-side (Java/Docker) Basic Structure Self-Hosted
OpenDataLoader Extraction Pipeline Moderate Local/Cloud Hybrid
Tesseract CLI Legacy OCR Poor 100% Offline

Sources: