Beyond the Pixel: Reconstructing the Spine of a Book with pdf-craft
How a local-first pipeline uses DeepSeek's vision intelligence to transform scanned chaos into structured Markdown without the LLM "correction tax."
- The library replaces the traditional OCR-plus-LLM correction loop with a single-stage vision-language model to eliminate processing fees and data leakage.
- Geometric histogram analysis identifies white space "valleys" to accurately reconstruct multi-column layouts before text extraction.
- A stateful joining module stitches fragmented sentences across page breaks to maintain a continuous narrative flow in the final Markdown.
- Mathematical formulas are treated as protected atomic units to prevent LaTeX equations from being mangled during document reflowing.
The Death of the Correction Loop
The standard playbook for processing complex PDF documents has always been a clunky two-step dance. First, you run a basic Optical Character Recognition engine to extract raw text coordinates. Second, you pass that garbled, fragmented text to a Large Language Model and beg it to fix the formatting.
This approach is slow, expensive, and leaks private data to server-side APIs. The pdf-craft project outright eliminates this cycle. By integrating DeepSeek's Vision-Language Models directly into a local Python pipeline, it treats document conversion not as text recognition, but as semantic reconstruction.
"Version 1.0.0 represents a fundamental shift in pdf-craft's text recognition strategy. Prior to this release, the system used a two-stage approach: basic OCR followed by LLM-based text correction. The new architecture eliminates the correction stage entirely, relying on DeepSeek OCR's native accuracy to produce final text output."
Seeing the Gutter: Geometric Document Analysis
A flat image of a page does not natively understand columns or reading order. If you naively read a two-column academic paper from left to right, the text becomes meaningless. The library solves this using a module called reading_serials.py.
The code uses X-axis projections to build a weighted histogram of the page. It looks for the "valleys" in the histogram to identify the gutters of empty white space between columns. This geometric approach allows pdf-craft to slice the page into logical serials before attempting to read the words.
The Jointer: Stitching the Narrative Thread
Books have page breaks, images that interrupt paragraphs, and floating tables. A standard OCR engine treats every page as an isolated island. The jointer.py module acts as the bridge.
It maintains a stateful "tail" pattern. When a paragraph ends abruptly at the bottom of a page, the jointer holds that incomplete sentence in memory. As the next page is processed, the jointer snaps the held tail directly onto the head of the new sequence, resulting in a continuous, unbroken Markdown paragraph.
Math as a First-Class Citizen
Most open-source extraction tools fail spectacularly when they encounter LaTeX formulas. They mangle equations into random alphanumeric strings. pdf-craft treats mathematics as an atomic unit.
The expression.py parser distinguishes between standard text and LaTeX delimiters. Once it identifies a mathematical expression, it locks the content within those boundaries. When the surrounding text is reflowed or joined, the mathematical integrity is perfectly preserved.
Scaling the "Gundam"
Running high-end vision models locally requires serious resource management. The page_ref.py module calculates the maximum allowable DPI based on a target file size to prevent memory crashes on heavy architectural drawings.
The project offers multiple tiers of the DeepSeek model, allowing developers to trade speed for accuracy depending on their available hardware. The "Gundam" tier provides the highest semantic accuracy, but the smaller models offer rapid iteration for CPU-bound environments.
| Tool | Architecture | Formula Support | Privacy |
|---|---|---|---|
| pdf-craft | Local Vision-Language Model | High (Native LaTeX) | 100% Offline |
| Stirling-PDF | Server-side (Java/Docker) | Basic Structure | Self-Hosted |
| OpenDataLoader | Extraction Pipeline | Moderate | Local/Cloud Hybrid |
| Tesseract CLI | Legacy OCR | Poor | 100% Offline |
Sources: