TL-DR: How a Local LLM Turns Chaotic WhatsApp Chats Into Reliable Events
A privacy-first Python pipeline uses windowing, fuzzy matching, and anchor-based date resolution to extract deadlines from messy Hinglish group chats without sending data to the cloud.
- TL-DR’s real trick is not better prompting, but refusing to trust the model until deterministic code has cleaned up its output.
- Sliding chat windows with overlap let the system preserve context across boundaries without forcing the model to process an entire export at once.
- Fuzzy matching and numeric mismatch checks stop small-model hallucinations from turning similar assignments into the same event.
- Anchor-based date resolution makes words like kal and parso historically correct by tying them to the message timestamp, not the machine clock.
WhatsApp group chats are terrible storage for deadlines. They mix memes, side quests, half-finished plans, and language switches into one endless scroll. TL-DR exists to pull the useful parts back out and turn them into something you can actually put on a calendar.
That matters because the problem is not summarization. It is extraction under noise. The repo treats a local LLM as a noisy parser, then surrounds it with Python code that repairs IDs, carries context forward, and resolves dates against the message that spawned them.
Why this is more than a chat summarizer
The obvious version of this project would be a chatbot that says, "Here is what your group discussed." TL-DR aims higher. It identifies deadlines, events, updates, and cancellations, then emits a calendar-friendly result that can survive real-world chat chaos.
The key constraint is privacy. Instead of sending sensitive group messages to a cloud API, the repo runs locally with Ollama and gemma4:e4b. That choice shapes everything else in the pipeline, because a smaller model needs help to stay useful.
The core idea: trust the model less, the pipeline more
TL-DR is built around a simple wager: a small local model can still be valuable if you do not let it be the final authority. The model extracts candidates. The pipeline decides what survives.
That shows up immediately in the repo structure. Windowing keeps prompts manageable. Known items carry state forward. Matching repairs bad IDs. Date resolution anchors relative language to the message that produced it. Each stage compensates for a different failure mode.
Small models invent ids (DBMS_Assignment_2), put a title where an id belongs, or copy one item's title onto another's messages. I solve this by generating a unique hash for each item based on its first-seen content.
How TL-DR keeps context across messy chat windows
The windowing strategy is the first guardrail. Chats are split when they hit 40 messages or a 3-hour gap, with a 5-message overlap to keep handoffs from snapping context in half. That overlap matters when a deadline gets changed just as a new window begins.
The extractor then receives a list of known items from earlier windows. That gives the LLM a chance to issue updates or cancellations instead of inventing duplicates. It is a practical way to give a stateless model a memory without pretending it actually has one.
# Conceptual flow from the repo
windows = make_windows(chat, max_messages=40, max_gap_hours=3, overlap=5)
known_items = []
for window in windows:
result = extract_window(window, known_items=known_items)
known_items = reconcile(known_items, result.items)
# Later windows can update or cancel earlier items instead of duplicating them
This is why the project feels more like evidence processing than chat generation. It is not asking the model to narrate the conversation. It is asking the model to help build a durable record of what changed, when it changed, and which earlier item it belongs to.
Why fuzzy matching matters when the model invents IDs
This is the defensive-programming section of the repo. The matcher does not blindly accept whatever the model says. It compares titles, checks overlap, and refuses obvious numeric collisions. That is how it avoids merging Assignment 1 with Assignment 2 just because the model got lazy with IDs.
| Approach | Input style | Handles messy human language? | Preserves privacy? | Handles updates and cancellations? | Works with Hinglish? | Calendar-ready output |
|---|---|---|---|---|---|---|
| Regex or command bots | Strict commands | Poor | Usually yes | Weak | Poor | Limited |
| Cloud LLM summarizers | Free-form chat | Good | No | Inconsistent | Mixed | Usually summary-first |
| TL-DR | Free-form WhatsApp exports | Good | Yes | Strong | Strong | Structured events plus .ics |
The point of the matcher is not elegance for its own sake. It is to keep the pipeline honest when the model confuses names, titles, and IDs. In a small-model system, that kind of plumbing is the product.
Anchors beat the system clock
Relative time is where chat tools usually break. A word like kal can mean tomorrow, but only if you know which message said it and when that message was sent. TL-DR resolves those phrases against the message timestamp, not the machine’s current date.
That sounds like a small detail until you think about old exports. If you run the tool next week, a naive parser can misread the whole timeline. Anchoring date resolution to the chat itself makes the output historically correct instead of merely current.
The same logic applies to Hinglish time words like aaj, kal, and parso. The repo also handles time-of-day cues such as subah, shaam, and raat, which is exactly the kind of local nuance that global AI demos often flatten.
Why Hinglish is the real benchmark
If this repo only worked on polished English, it would be much less interesting. The real test is whether it can interpret casual Indian student chat, where deadlines are wrapped in slang, shorthand, and mixed-language phrasing.
That is where the project becomes culturally specific in a useful way. TL-DR is not trying to be a universal assistant. It is trying to be a reliable utility for a very common local workflow: students trying not to miss an assignment buried in a flood of messages.
This is also why the local-first design matters. The privacy story is good, but the deeper story is fit. The tool is shaped around the language and habits of the people who actually need it.
Where TL-DR beats the obvious alternatives
| Approach | Best at | Weakness | Why TL-DR wins |
|---|---|---|---|
| Regex-based WhatsApp bots | Rigid commands | Breaks on real conversation | TL-DR can read messy natural chat |
| Cloud summarizers | Generic summaries | Not privacy-first and not calendar-native | TL-DR keeps data local and outputs structured events |
| Generic local assistants | Broad convenience | Usually not tuned for reconciliation | TL-DR carries memory across windows and repairs IDs |
The comparison is not about raw intelligence. It is about workflow fit. TL-DR beats simpler tools because it understands that extraction, reconciliation, and date normalization are separate problems, and each one needs a different safeguard.
What this repo says about building with small models
TL-DR is a clean example of how to build with model weakness instead of around it. The system does not assume the model will be precise, consistent, or memoryful. It assumes the opposite, then compensates with structure.
That makes the project bigger than a WhatsApp parser. It is a design pattern for local AI: keep the model small, keep the data private, and make the pipeline strict enough to earn trust.
Group chats bury the one real deadline under hundreds of messages, memes and "bhai sab aa gaye?". tl;dr reads a WhatsApp export, extracts only the deadlines, events and changes of plan, and gives you a clean summary plus a calendar file you can import.