DataChat: The Text-to-SQL Assistant That Refuses to Trust the Model

A supply-chain analytics tool that turns questions into PostgreSQL queries, then checks every column against the live schema before it answers.

8 min read • View on GitHub • More from Umrazahmed26

A wide editorial scene shows a supply chain manager asking a question at a desk while a multi-stage machine processes the request. The question passes through planner, writer, validator, executor, and synthesiser stages before a final answer emerges. It explains that DataChat is a controlled pipeline rather than a direct chat interface.
DataChat’s core idea is not just text-to-SQL. It is text-to-SQL with checkpoints, so the model has to pass through a sequence of guarded stages before a user sees an answer.
Key Takeaways

Most text-to-SQL tools sell the same promise: ask in English, get data back. DataChat is interesting because it does not let the model drive straight into PostgreSQL. It inserts a defensive middle layer that plans, writes, validates, executes, and only then explains. That changes the product from a chatbot into a constrained compiler for business questions.

Why DataChat feels more like a compiler than a chatbot

The strongest signal in the repo is not the UI. It is the refusal to trust a single model pass. The backend breaks the task into stages so each one has a narrow job, and each one can fail safely. That is a much more serious answer to text-to-SQL than a prompt template with a database connection.

The goal of DataChat is to make data accessible to everyone, regardless of their technical background.

Umraz Ahmed, Project Creator · DataChat: Converse with Your Data

How the pipeline works

The key design choice is separation. DataChat does not ask one model to do everything. It splits the work so each step can be inspected, constrained, or swapped out.

The pipeline is simple to describe and powerful in practice. First, planner logic extracts intent and entities from the user’s question. Then the SQL writer turns that into PostgreSQL-specific query text. Next, the validator checks whether the query actually matches the live schema. Only after that does execution happen, followed by a synthesis step that turns rows into a short business answer.

StageJobFailure it prevents
PlannerDetect intent and entities before SQL generationAmbiguous prompts becoming vague queries
SQL writerDraft PostgreSQL-specific SQLA model skipping relational structure
ValidatorCheck tables and columns against live schemaHallucinated column names or fake joins
ExecutorRun only approved SQLUnsafe or broken queries reaching the database
SynthesiserConvert rows into a business summaryRaw data dumps that users cannot read

The part that matters most: schema-aware validation

If DataChat has a technical thesis, it lives in the validator. The repo does not stop at generic prompt hygiene. It inspects the database schema and checks that tables and columns actually exist before a query can move forward. That is the difference between a demo and a tool a team could plausibly rely on.

A close-up shows a handwritten SQL query being pushed through a mesh filter inside a machine. On one side, tokens and identifiers are compared against a visible schema board with tables and columns. Invalid names snag on the filter, while approved names pass through cleanly. It explains how schema-aware validation protects the database from hallucinated SQL.
The validator is the real gatekeeper. It does not just look for suspicious words. It checks the query against the live schema, which is what makes the pipeline feel defensible.
# Conceptual shape of the validator
# Strip literals, inspect identifiers, compare against live schema

clean_sql = strip_string_literals(sql)
tokens = extract_identifiers(clean_sql)
valid_tables = inspect(engine).get_table_names()
valid_columns = get_live_columns(engine)

for token in tokens:
    if token in RESERVED_WORDS:
        continue
    if token not in valid_tables and token not in valid_columns:
        raise ValidationError(f"Unknown identifier: {token}")

This is why the ThoughtLog matters. It is not theater. It is a visible audit trail for the staging process, which helps users understand why the system arrived at a query and helps developers find where it failed. In an LLM product, that is not a nice-to-have. It is the core of trust.

This is a game-changer for our team. We spend too much time writing ad-hoc queries.

A User on LinkedIn, Data Analyst · DataChat: Converse with Your Data

Why this is tuned for supply chain, not just SQL

DataChat is not trying to be a generic wrapper around a database. The repository shows domain-specific mapping between business language and schema terms, which is what makes the tool feel aimed at supply chain workflows. That matters because most of the pain in enterprise analytics is translation, not computation.

Business languageMapped schema ideaWhy it helps
buyercustomerMatches how managers ask questions
skuproductReduces vocabulary mismatch
shipmentorder or delivery recordConnects ops terms to transactional data
regionlocation fieldMakes filtering feel natural

That specialization is the hidden advantage. Generic text-to-SQL systems often fail because they leave users to guess the schema in their own heads. DataChat narrows that gap by teaching the system the language of the domain, not just the syntax of SQL.

How DataChat compares to the field

ApproachStrengthWeaknessTrust modelBest for
Generic LLM wrapper for SQLFast to prototypeEasy to hallucinate queriesMostly implicitDemos and experiments
Framework-based build with LangChain or LlamaIndexHighly flexibleMore assembly requiredDepends on the builderTeams with custom requirements
Dashboard-first BIFamiliar and stableStatic and slow to adaptHuman-curatedRecurring reporting
DataChatOpinionated and inspectableNarrower scopeSchema-aware pipelineBusiness users asking live questions

That middle position is the point. DataChat is more turnkey than a framework build, but more transparent than a black-box chatbot. It is not trying to replace BI, and it is not pretending the model is reliable on its own. It turns trust into a sequence of checks.

What DataChat gets right, and where it still feels like a prototype

The engineering reads as thoughtful. The separation of concerns is clean, the fallback logic is pragmatic, and the observability mindset is real. At the same time, the repository feels like a high-quality prototype or beta system more than a widely hardened product, which is exactly why the architecture deserves attention. It shows what serious text-to-SQL design looks like before scale forces compromise.

If you are building conversational analytics, the lesson is blunt: do not ask the model to be correct in one leap. Break the task into smaller jobs, validate the dangerous parts against reality, and expose enough intermediate state that humans can understand the result. DataChat is a neat example of that discipline.