InsightLedger: The AI Equity Analyst That Knows When Not to Trust Itself
A hybrid research pipeline uses Gemini to extract financial facts from annual reports, then hands the numbers to deterministic code for scoring, ranking, and peer comparison.
- InsightLedger is interesting because it uses AI to read annual reports and deterministic code to decide what those readings mean.
- The repo’s real strength is its split between messy document extraction and strict numeric ranking.
- Progress tracking and schema-constrained prompts make the pipeline feel like a research tool, not a demo.
- Its edge is focus: it targets NIFTY-style equity workflows instead of pretending to be a universal stock oracle.
Why this is not just another AI stock picker
Most AI finance demos make the same mistake: they ask the model to do everything. InsightLedger takes the opposite approach. It treats Gemini like a document parser, then hands the extracted fields to NumPy and rule-based code for the part that should never be improvisational: the math.
That split matters because equity research has two different jobs. One is reading messy annual reports, management commentary, and disclosure tables. The other is calculating ratios, normalizing peers, and ranking them with discipline. InsightLedger separates those jobs cleanly.
The pipeline is built like a research factory
The orchestration lives in pipeline.py, and it reads like someone has actually run this workflow on real data. There is a demo mode for smaller runs, a batch mode for broader coverage, and progress tracking that lets the job resume when a report fails or a network call breaks.
That resumability is not a nicety. It is the difference between a usable research tool and an expensive rerun machine. If you are processing many PDFs and making repeated LLM calls, a saved progress state is a cost-control feature as much as a developer convenience.
# pipeline.py concept
progress = load_progress()
if mode == "demo":
symbols = demo_universe()
else:
symbols = nifty_100_universe()
for symbol in symbols:
if progress.is_done(symbol):
continue
report = fetch_report(symbol)
extracted = extract_with_gemini(report, schema="schema_individual.json")
scores = compute_ratios(extracted)
save_progress(symbol, scores)
Gemini is used as a parser, not a pundit
The nicest technical idea in the repo is the prompt pipeline in scraper/ai_extractor.py. Prompts live in Markdown, output is constrained by schema_individual.json, and the LLM is pushed toward a single job: turn unstructured report text into structured fields that the rest of the system can trust.
That is a better use of an LLM than open-ended financial advice. It turns fuzzy narrative into repeatable JSON. In other words, the model is being asked to read, not to speculate.
The math layer is the real trust engine
Once the fields are extracted, the system switches to deterministic computation. quantitative_fetcher.py calculates figures like free cash flow and ROIC. scraper/peer_evaluator.py uses NumPy to compute Z-scores across peers, which turns qualitative-looking output into a ranked comparison with explicit math behind it.
This is the anti-hallucination layer, and it is meaningful because it does not merely add a disclaimer. It structurally limits where uncertainty can live. The model can misread a sentence, but it does not get to invent the ranking.
| Workflow | What it is good at | Where it fails | Trust model | Best use case |
|---|---|---|---|---|
| Traditional analyst | Reading nuance, asking follow-up questions, spotting context | Slow, costly, inconsistent at scale | Human judgment plus review | Deep work on a small number of companies |
| Generic LLM finance app | Fast summaries and conversational access | Can blur facts, math, and speculation | User trust, often too loose | Broad exploration and idea generation |
| InsightLedger | Structured extraction and repeatable ranking | Still depends on report quality and schema design | LLM for reading, code for judgment | Repeatable screening across a defined universe |
What makes InsightLedger different
The project is not trying to be the smartest narrator in the room. It is trying to be the most disciplined workflow. That gives it a clearer product shape than a generic chatbot wrapped around market data, and it fits the realities of equity research better than an all-purpose AI wrapper ever could.
The market focus helps too. Targeting Indian filings and NIFTY-style workflows makes the repo feel grounded in a real operational context, not a synthetic demo. That specificity is a feature, not a limitation.
Where the project feels strong, and where it still feels early
The architecture is modular and thoughtful. The separation between ingestion, extraction, and evaluation suggests somebody cared about failure modes, not just screenshots. At the same time, the project still reads like a functional prototype rather than a fully hardened product.
That is not a criticism so much as a read on maturity. The bones are good. The next step would be turning this disciplined pipeline into something that can survive messy real-world usage without hand-tuning.