PATH A: PIPELINE OCR PATH B: CHANDRA (VLM) crop recognize heuristics full image generate Document Image Segment Block 1 Segment Block 2 OCR Text 1 OCR Text 2 Reconstruct Layout (Chaotic Heuristics) Vision Language Model HTML/MD + Bounding Boxes