Document-to-Answer Pipeline
Endee Vector Architecture & RAG Retrieval Path
STAGE 1/5 DOCUMENT INGESTION & PARSING: Raw files (PDF, DOCX, TXT, HTML) are continuously streamed into the unified document processor.
01. Ingestion & Preprocessing 02. Vector Indexing (Endee) 03. Contextual Query & Retrieval Path chunks attrs tokens batch knn index vector top-k context BATCH INSERT GROUP [4] #C1 #C2 #C3 #C4 PDF FILE DOCX FILE TXT FILE HTML FILE PROCESSOR CHUNKING META ENRICH EMBEDDING ENDEE DB USER QUERY VECTOR SEARCH TOP CHUNKS CITED ANSWER
NODE INSPECTOR: PROCESSOR Click nodes or batch chunks to view properties
Component Unified Doc Parser
Status / Operation Extracting layout & clean text
Output Target Chunker & Metadata Engine