Advanced RAG

96 explainers
MarkItDown: Microsoft’s Document Converter That Treats Markdown as an AI Ingestion Layer
Advanced RAG
MarkItDown: Microsoft’s Document Converter That Treats Markdown as an AI Ingestion Layer
A stream-based, plugin-driven pipeline that turns messy Office files, PDFs, images, and more into structure LLMs can actually use.
9 min read
Rescuing Data from the PDF Trap: Inside Docling
Advanced RAG
Rescuing Data from the PDF Trap: Inside Docling
How a hybrid architecture of deterministic parsing and small vision models is fixing the ingestion bottleneck for RAG pipelines.
8 min read
PageIndex: Teaching RAG to Read the Table of Contents Before It Reads the Page
Advanced RAG
PageIndex: Teaching RAG to Read the Table of Contents Before It Reads the Page
A vectorless retrieval framework that replaces chunk-and-embed guesswork with hierarchical document maps, reasoning-based navigation, and verification loops for complex PDFs and Markdown.
10 min read
datalab-to/marker: The Vision-First End to the PDF Bottleneck
Advanced RAG
datalab-to/marker: The Vision-First End to the PDF Bottleneck
How a custom OCR pipeline and polygon geometry turned thirty years of printer-optimized documents into structured data for AI agents.
8 min read
Grounding the Black Box: How google/langextract Forces LLMs to Show Their Work
Advanced RAG
Grounding the Black Box: How google/langextract Forces LLMs to Show Their Work
Extracting structured data from unstructured text is easy. Proving exactly where that data came from down to the character offset is the hard part. Here is how Google built an audit trail for generative extraction.
8 min read
RAG_Techniques: The Repo That Teaches You Retrieval Is a Decision System
Advanced RAG
RAG_Techniques: The Repo That Teaches You Retrieval Is a Decision System
From HyDE to agentic loops and evaluation, this collection shows why advanced RAG is less about prompting a model and more about shaping what the model sees, when it sees it, and how you know it worked.
12 min read
Unlimited-OCR: The OCR Project That Tries to Read the Whole Document at Once
Advanced RAG
Unlimited-OCR: The OCR Project That Tries to Read the Whole Document at Once
Baidu’s long-horizon parser replaces the usual detect-recognize-stitch pipeline with a VLM-style stream that can keep context across pages, resist repetition, and emit structured output from raw document images.
8 min read
RAG-Anything: The Framework That Refuses to Pretend Documents Are Text
Advanced RAG
RAG-Anything: The Framework That Refuses to Pretend Documents Are Text
A tour of the ingestion pipeline, multimodal processors, and resilience machinery that turn messy files into graph-ready knowledge.
9 min read
Surya and the Vision-First Reconstruction of the Document
Advanced RAG
Surya and the Vision-First Reconstruction of the Document
How Vik Paruchuri's high-performance toolkit replaced legacy OCR with a multimodal understanding of layout, logic, and reading flow.
last30days-skill: The High-Stakes Architecture of Real-Time Research
Advanced RAG
last30days-skill: The High-Stakes Architecture of Real-Time Research
How a modular signal engine triangulates truth across prediction markets, social sentiment, and the open web to kill the LLM knowledge cutoff.
The Great PDF Triage: How OpenDataLoader PDF Ends the RAG Accuracy Tax
Advanced RAG
The Great PDF Triage: How OpenDataLoader PDF Ends the RAG Accuracy Tax
By blending deterministic Java heuristics with surgical AI strikes, Hancom’s hybrid engine solves the multi-column "scrambled text" problem without a GPU-heavy budget.
8 min read
Azure-Samples/azure-search-openai-demo: The RAG Template That Shows Its Work
Advanced RAG
Azure-Samples/azure-search-openai-demo: The RAG Template That Shows Its Work
A deep look at the Azure sample that turns enterprise search, semantic reranking, secure citations, and multimodal ingestion into one production-minded reference app.
10 min read
Beyond the Pixel: Reconstructing the Spine of a Book with pdf-craft
Advanced RAG
Beyond the Pixel: Reconstructing the Spine of a Book with pdf-craft
How a local-first pipeline uses DeepSeek's vision intelligence to transform scanned chaos into structured Markdown without the LLM "correction tax."
8 min read
Memgraph: The Graph Database That Treats Memory Like a Runtime, Not a Cache
Advanced RAG
Memgraph: The Graph Database That Treats Memory Like a Runtime, Not a Cache
A C++20 graph engine with polyglot extensions, strict memory control, and a design that makes in-memory databases feel less fragile and more programmable.
11 min read
The End of the DIY RAG Script: Inside langflow-ai/openrag
Advanced RAG
The End of the DIY RAG Script: Inside langflow-ai/openrag
How a unified stack of Docling, OpenSearch, and Langflow is turning semantic search into a plug-and-play agentic tool.
6 min read
The Living Database: Unpacking RuVector
Advanced RAG
The Living Database: Unpacking RuVector
How a Rust-based vector engine uses Graph Neural Networks and reinforcement learning to rewrite its own indexes in real time.
8 min read
Decoding the Noise: Inside llm_aided_ocr
Advanced RAG
Decoding the Noise: Inside llm_aided_ocr
How a hybrid pipeline pairs classic extraction with language models to turn garbage scans into perfect Markdown.
6 min read
geo-optimizer-skill: GEO Optimizer: The Open-Source Lighthouse for AI Citability
Advanced RAG
geo-optimizer-skill: GEO Optimizer: The Open-Source Lighthouse for AI Citability
It audits robots.txt, llms.txt, schema, and brand signals, then turns research on generative engine optimization into a tool an AI agent can actually use.
8 min read
mini-rag: The RAG Tutorial That Teaches the Parts Everyone Skips
Advanced RAG
mini-rag: The RAG Tutorial That Teaches the Parts Everyone Skips
A production-minded walkthrough of the workers, factories, vector stores, and Arabic-first prompting that turn a basic retrieval demo into a real application.
8 min read
The Geometry of Reading: Unpacking pdftext
Advanced RAG
The Geometry of Reading: Unpacking pdftext
How an Apache-2.0 mandate forced developers to reconstruct PDF layouts from raw C++ character dumps.
7 min read
chokepoint-atlas: Chokepoint Atlas Turns Supply-Chain Research Into a Testable System
Advanced RAG
chokepoint-atlas: Chokepoint Atlas Turns Supply-Chain Research Into a Testable System
A Python pipeline that grades evidence, scores bottlenecks, and turns AI infrastructure research into structured artifacts instead of hand-wavy thesis decks.
8 min read
The LLM as Database Admin: Unpacking chroma-mcp
Advanced RAG
The LLM as Database Admin: Unpacking chroma-mcp
How the official Model Context Protocol server for Chroma replaces hardcoded RAG pipelines by giving AI agents autonomous control over their own vector memory.
6 min read
The Agentic Search Factory: How context-1-data-gen Solves Context Rot
Advanced RAG
The Agentic Search Factory: How context-1-data-gen Solves Context Rot
Inside Chroma's specialized pipeline for generating high-fidelity, multi-hop retrieval tasks that teach models to ignore the noise.
RAG-Tutorials Shows Why Good Retrieval Is a Systems Problem
Advanced RAG
RAG-Tutorials Shows Why Good Retrieval Is a Systems Problem
A modular tour through ingestion, FAISS, agentic retrieval, and vectorless document navigation that shows how RAG matures beyond naive chunk search.
8 min read
santifer/cv-santiago: The CV That Behaves Like a Production AI System
Advanced RAG
santifer/cv-santiago: The CV That Behaves Like a Production AI System
An interactive portfolio that turns a personal resume into a live demo of agentic RAG, voice latency control, security defenses, and eval-driven LLMOps.
8 min read
lambda-RLM: λ-RLM: The Y-Combinator for LLMs
Advanced RAG
lambda-RLM: λ-RLM: The Y-Combinator for LLMs
A typed functional runtime that replaces agentic chaos with deterministic recursion, socket-isolated inference, and mathematically chosen partitioning for long-context reasoning.
10 min read
Paperclip: The Biomedical Corpus That Behaves Like a Unix Filesystem
Advanced RAG
Paperclip: The Biomedical Corpus That Behaves Like a Unix Filesystem
A CLI and MCP server that lets humans and AI agents grep, read, and synthesize millions of papers as if they were files on disk.
9 min read
bb25: When BM25 Stops Being a Score and Starts Being a Probability
Advanced RAG
bb25: When BM25 Stops Being a Score and Starts Being a Probability
A Rust implementation of Bayesian BM25 that calibrates lexical relevance into probabilities, then fuses sparse and dense signals with log-odds instead of guesswork.
8 min read
The Anatomy of a RAG App: Inside gpt4-pdf-chatbot-langchain-chroma
Advanced RAG
The Anatomy of a RAG App: Inside gpt4-pdf-chatbot-langchain-chroma
How a Next.js boilerplate moved AI development past the single-prompt wrapper by bringing vector math and local-first document retrieval to the TypeScript ecosystem.
8 min read
The Database That Writes Its Own Firmware: Inside chroma-core/chroma-cookbooks
Advanced RAG
The Database That Writes Its Own Firmware: Inside chroma-core/chroma-cookbooks
Why the creators of Chroma built a type-safe, self-correcting orchestration kernel hiding in plain sight as a repository of tutorials.
8 min read
The Illusion of Zero Latency: Unpacking frankensearch
Advanced RAG
The Illusion of Zero Latency: Unpacking frankensearch
How a Rust-based hybrid engine uses structured concurrency and a two-tier progressive pipeline to eliminate retrieval bottlenecks for local agents.
7 min read
bayesian-bm25: Bayesian BM25: Turning Search Scores Into Probabilities You Can Trust
Advanced RAG
bayesian-bm25: Bayesian BM25: Turning Search Scores Into Probabilities You Can Trust
How Cognica’s library converts raw BM25 into calibrated relevance signals, then uses log-odds fusion and Bayesian priors to make hybrid search less ad hoc.
10 min read
The End of the Vector API Key: Inside chroma-swift
Advanced RAG
The End of the Vector API Key: Inside chroma-swift
How a polyglot stack of Swift, Rust, and MLX is moving retrieval-augmented generation from the cloud directly onto Apple Silicon.
8 min read
aisqlgen: The AI SQL Notebook That Edits, Not Just Generates
Advanced RAG
aisqlgen: The AI SQL Notebook That Edits, Not Just Generates
A local-first Next.js tool that narrows schemas before prompting, remembers prior query edits, and turns raw database answers into charts.
8 min read
The Remote Brain for Coding Agents: Inside chroma-core/package-search
Advanced RAG
The Remote Brain for Coding Agents: Inside chroma-core/package-search
How a Python pipeline and a GitHub repository orchestrate a massive, continuously updated vector database of the world's open-source packages.
6 min read
repo-explainer: Repo Explainer: The Anti-RAG Tool That Reads the Whole Codebase
Advanced RAG
repo-explainer: Repo Explainer: The Anti-RAG Tool That Reads the Whole Codebase
A tiny Streamlit app uses Gemini 1.5 Pro’s massive context window to turn an entire GitHub repo into something you can interrogate like a senior engineer.
7 min read
Beyond the Grid: How laravel-excel-to-x Turns Spreadsheets into AI Context
Advanced RAG
Beyond the Grid: How laravel-excel-to-x Turns Spreadsheets into AI Context
Moving past the "Import to DB" bottleneck with a high-speed transformation engine for Laravel.
6 min read
PID_Symbol_Detection: Why P&ID Symbol Extraction Works Better as Detect, Then Compare
Advanced RAG
PID_Symbol_Detection: Why P&ID Symbol Extraction Works Better as Detect, Then Compare
A two-stage pipeline uses sliced inference to find tiny symbols on huge drawings, then few-shot learning to label them without needing massive class-specific datasets.
9 min read
`google-research/ecology-georeferencing`: the repo that teaches an LLM to find the smallest map that matters
Advanced RAG
`google-research/ecology-georeferencing`: the repo that teaches an LLM to find the smallest map that matters
A cost-aware Gemini pipeline filters out non-maps, reads ecology figures with surrounding context, and turns PDF pages into georeferenced data instead of dead images.
8 min read
LLMOps_series: MultiDocChat: The Session-Scoped RAG Template That Treats Ingestion, Retrieval, and Evaluation as One System
Advanced RAG
LLMOps_series: MultiDocChat: The Session-Scoped RAG Template That Treats Ingestion, Retrieval, and Evaluation as One System
A close look at how `yashprogrammer/LLMOps_series` turns multi-document chat into a portable, testable FastAPI stack with per-session FAISS indexes, LCEL retrieval, and LangSmith evaluations.
8 min read
The Stateful Document: Unpacking the Datalab SDK
Advanced RAG
The Stateful Document: Unpacking the Datalab SDK
How a resilient Python wrapper turns long-running PDF extraction into a predictable, asynchronous pipeline.
8 min read
chroma-core/chroma_datasets: The Missing Standard Library for Vector Data
Advanced RAG
chroma-core/chroma_datasets: The Missing Standard Library for Vector Data
How a lightweight Python wrapper turns Hugging Face into a portable save state for Chroma collections, cutting the RAG cold-start from hours to seconds.
6 min read
UQA: The Database Repo That Treats SQL, Graphs, and Vector Search as the Same Algebra
Advanced RAG
UQA: The Database Repo That Treats SQL, Graphs, and Vector Search as the Same Algebra
A theory-heavy Python engine that uses posting lists, functors, and Bayesian calibration to turn multi-model database design into one unified execution model.
11 min read
The Database That Thinks in Sentences: Unpacking musetronstar/tagd
Advanced RAG
The Database That Thinks in Sentences: Unpacking musetronstar/tagd
How an obscure C++ semantic engine bypasses SQL entirely, using a custom parser and an axiomatic ontology to turn SQLite into a knowledge graph.
7 min read
bigindexer: BGI: The Code Indexer That Reads Behavior, Not Imports
Advanced RAG
bigindexer: BGI: The Code Indexer That Reads Behavior, Not Imports
A look at a static analysis tool that builds architectural maps from intent, turns cluster overflow into a refactoring signal, and exposes the result to AI agents through MCP.
9 min read
wittgenstein-reader: Wittgenstein Reader: An AI That Understands the Space Between the Propositions
Advanced RAG
wittgenstein-reader: Wittgenstein Reader: An AI That Understands the Space Between the Propositions
A close reading of a hybrid Python and TypeScript pipeline that turns OCR noise into context-aware philosophical commentary, with a reader UI built to keep the argument moving.
10 min read
Appwrite `sdk-for-md`: The Docs SDK That Turns Documentation Into a Dependency
Advanced RAG
Appwrite `sdk-for-md`: The Docs SDK That Turns Documentation Into a Dependency
A Markdown-first package for AI agents, built to make API knowledge searchable, versioned, and safe to load on demand.
6 min read
tsukineko-grimoire: Tsukineko Grimoire: The RAG System That Turns arXiv Into a Bilingual Research Shelf
Advanced RAG
tsukineko-grimoire: Tsukineko Grimoire: The RAG System That Turns arXiv Into a Bilingual Research Shelf
A smart archive for papers, figures, and cited answers, built to bridge Japanese researchers into an English-dominated literature stream.
9 min read