InsightLedger: The AI Equity Analyst That Knows When Not to Trust Itself

A hybrid research pipeline uses Gemini to extract financial facts from annual reports, then hands the numbers to deterministic code for scoring, ranking, and peer comparison.

8 min read • View on GitHub • More from navyabalaji

A wide split desk scene. On one side, an analyst’s workspace is buried under annual reports, highlighted pages, ratio sheets, and sticky notes. On the other side, a clean machine pipeline turns a PDF into structured financial cards, with a thin tether linking the two halves. It explains the article’s core idea: AI reads the documents, but math makes the judgment.
InsightLedger is built around a split-brain workflow: interpretation first, computation second.
Key Takeaways

Why this is not just another AI stock picker

Most AI finance demos make the same mistake: they ask the model to do everything. InsightLedger takes the opposite approach. It treats Gemini like a document parser, then hands the extracted fields to NumPy and rule-based code for the part that should never be improvisational: the math.

That split matters because equity research has two different jobs. One is reading messy annual reports, management commentary, and disclosure tables. The other is calculating ratios, normalizing peers, and ranking them with discipline. InsightLedger separates those jobs cleanly.

A close-up editorial scene of a report being stamped into structured fields, then pushed through a second mechanical device that compares peers and produces a ranked ladder. The first stage looks interpretive and text-driven. The second stage looks colder and more deterministic, showing how extracted facts become ranked judgments.
The project’s key move is simple: extract with AI, rank with math.

The pipeline is built like a research factory

The orchestration lives in pipeline.py, and it reads like someone has actually run this workflow on real data. There is a demo mode for smaller runs, a batch mode for broader coverage, and progress tracking that lets the job resume when a report fails or a network call breaks.

That resumability is not a nicety. It is the difference between a usable research tool and an expensive rerun machine. If you are processing many PDFs and making repeated LLM calls, a saved progress state is a cost-control feature as much as a developer convenience.

A three-stage loop makes the design philosophy explicit: ingest, extract, judge, then resume if anything breaks.

# pipeline.py concept
progress = load_progress()
if mode == "demo":
    symbols = demo_universe()
else:
    symbols = nifty_100_universe()

for symbol in symbols:
    if progress.is_done(symbol):
        continue
    report = fetch_report(symbol)
    extracted = extract_with_gemini(report, schema="schema_individual.json")
    scores = compute_ratios(extracted)
    save_progress(symbol, scores)

Gemini is used as a parser, not a pundit

The nicest technical idea in the repo is the prompt pipeline in scraper/ai_extractor.py. Prompts live in Markdown, output is constrained by schema_individual.json, and the LLM is pushed toward a single job: turn unstructured report text into structured fields that the rest of the system can trust.

That is a better use of an LLM than open-ended financial advice. It turns fuzzy narrative into repeatable JSON. In other words, the model is being asked to read, not to speculate.

The math layer is the real trust engine

Once the fields are extracted, the system switches to deterministic computation. quantitative_fetcher.py calculates figures like free cash flow and ROIC. scraper/peer_evaluator.py uses NumPy to compute Z-scores across peers, which turns qualitative-looking output into a ranked comparison with explicit math behind it.

This is the anti-hallucination layer, and it is meaningful because it does not merely add a disclaimer. It structurally limits where uncertainty can live. The model can misread a sentence, but it does not get to invent the ranking.

A hedcut-style portrait of the contributor based on a verified GitHub avatar. The portrait provides attribution context for the repo’s visible maintainer while keeping the article’s visual language consistent with editorial line art.
WorkflowWhat it is good atWhere it failsTrust modelBest use case
Traditional analystReading nuance, asking follow-up questions, spotting contextSlow, costly, inconsistent at scaleHuman judgment plus reviewDeep work on a small number of companies
Generic LLM finance appFast summaries and conversational accessCan blur facts, math, and speculationUser trust, often too looseBroad exploration and idea generation
InsightLedgerStructured extraction and repeatable rankingStill depends on report quality and schema designLLM for reading, code for judgmentRepeatable screening across a defined universe

What makes InsightLedger different

The project is not trying to be the smartest narrator in the room. It is trying to be the most disciplined workflow. That gives it a clearer product shape than a generic chatbot wrapped around market data, and it fits the realities of equity research better than an all-purpose AI wrapper ever could.

The market focus helps too. Targeting Indian filings and NIFTY-style workflows makes the repo feel grounded in a real operational context, not a synthetic demo. That specificity is a feature, not a limitation.

Where the project feels strong, and where it still feels early

The architecture is modular and thoughtful. The separation between ingestion, extraction, and evaluation suggests somebody cared about failure modes, not just screenshots. At the same time, the project still reads like a functional prototype rather than a fully hardened product.

That is not a criticism so much as a read on maturity. The bones are good. The next step would be turning this disciplined pipeline into something that can survive messy real-world usage without hand-tuning.