Swiggy-End-To-End-Data-Analytics: Swiggy End-To-End Data Analytics: When a Churn Model Becomes a Retention Playbook

A portfolio project that moves from SQL and Power BI to machine learning and a Flask app, then turns customer risk into concrete business action.

8 min read • View on GitHub • More from parlhad

A control room made of ink and white space, with a central switchboard routing customer signals into three action bins. The scene explains the article’s core idea: analytics is only valuable here when it ends in a decision, not a score.
The interesting move is not prediction. It is the translation from churn probability into a retention action.
Key Takeaways

Most portfolio analytics repos stop when the model produces a probability. This one keeps going. In parlhad/Swiggy-End-To-End-Data-Analytics, the useful output is not just churn risk. It is a suggested response: discount, monitor, or push a premium offer. That is the difference between predictive analytics and prescriptive analytics.

Not a churn score. A retention decision.

The strongest detail in the repo is in `app.py`. The app does not just surface a score and leave the user to interpret it. It turns that score into an action band. That small design choice changes the entire project from a dashboard demo into a decision system.

The primary objective of this project is to create an interactive Power BI dashboard for analyzing food delivery data from Swiggy.

Pralhad K., Project Creator · Swiggy-End-To-End-Data-Analytics README.md

How the repo turns raw orders into a business system

The pipeline is cleanly layered. Raw CSVs live in `Data/`. SQL defines the schema and business queries. Power BI handles diagnostic reporting. The notebook trains a churn model. Flask and Docker wrap the result so a non-technical user can interact with it.

The repo is not a straight analytics pipeline. It is a decision pipeline, with thresholds that convert prediction into retention policy.

Project typeWhat it producesWhat decision it supportsWhere it stops shortWhy this repo goes further
Dashboard-only portfolioCharts and KPIsDescribes the businessStops at diagnosisAdds prediction and action
Notebook-only churn modelA probability scoreRanks riskStops at modelingMaps risk to a response
Generic end-to-end demoA stitched stackProves toolingStops at integrationCloses the loop from data to retention playbook

The database is the first model

`SQL/schema.sql` is doing more than storage work. It defines the source of truth, cleans inconsistent IDs, and precomputes business questions that matter to a revenue team. That includes customer value, restaurant performance, geography, and rankings. In other words, the database is already expressing a point of view about the business.

-- Schema and cleaning logic shape the analytical backbone
CREATE TABLE users (
  user_id INT PRIMARY KEY,
  name VARCHAR(100),
  city VARCHAR(100),
  gender VARCHAR(20)
);

-- Example of the cleanup pattern used in the project
NULLIF(REGEXP_REPLACE(restaurant_id, '[^0-9]', '', 'g'), '')::INT AS restaurant_id_clean;

-- A business query pattern seen throughout the file
SELECT r.city, r.restaurant_name,
       RANK() OVER (PARTITION BY r.city ORDER BY COUNT(o.order_id) DESC) AS city_rank
FROM orders o
JOIN restaurants r ON o.restaurant_id = r.restaurant_id
GROUP BY r.city, r.restaurant_name;

That matters because analytics projects often treat SQL as a necessary prelude. Here it is part of the product logic. It defines what counts, how it is joined, and which comparisons the business can trust.

The dashboard asks the right questions

Power BI is not the main event, but it is not window dressing either. It gives the project a diagnostic layer before the model tries to make a prediction. That is useful. A team should understand where the problem lives before it decides how to intervene.

The dashboard’s value is organizational, not cosmetic. It separates users, cities, restaurants, and overall insights into views that are easy to scan. For a PM or analyst, that means the conversation starts with the business shape of the problem instead of the ML output.

A close-up ink drawing of a single customer signal moving from left to right through a compact model box, then splitting into three action cards. The image explains how a probability becomes a business response instead of stopping as an abstract score.
The key innovation is the handoff from model output to policy. That handoff is where prescriptive analytics begins.

The notebook turns history into risk

The notebook is where the project becomes predictive. It uses feature engineering, label encoding, and a tree-based classifier to estimate churn. That is standard enough. The interesting part is the context: the README signals a high churn rate, which makes class imbalance central, not incidental.

That explains the modeling choices. When churn is skewed, accuracy is a trap. ROC-AUC is the right kind of evaluation because it asks whether the model can rank risky users above safer ones. That is more useful for a retention workflow than a single yes-or-no label.

from sklearn.preprocessing import LabelEncoder
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import roc_auc_score

# Encode categorical features
le_city = LabelEncoder()
df['City_enc'] = le_city.fit_transform(df['City'])

# Train churn model
model = RandomForestClassifier(n_estimators=200, random_state=42)
model.fit(X_train, y_train)

# Evaluate ranking quality
probs = model.predict_proba(X_test)[:, 1]
auc = roc_auc_score(y_test, probs)

Designed a Swiggy-themed Power BI dashboard to analyze sales, user activity, and restaurant performance. Used Power Query and DAX to clean data, create KPIs, and deliver interactive visual insights.

Pralhad K., Project Creator · parlhad/Swiggy-End-To-End-Data-Analytics

Why this portfolio project feels complete

Compared with a dashboard-only repo, this one has a real downstream use. Compared with a notebook-only churn model, it has a user-facing interface. Compared with many end-to-end demos, it does not just prove that the stack works. It shows how the stack supports a business decision.

SignalWhat it suggests
SQL schema and cleaningThe project starts with data integrity
Power BI viewsThe project understands the business problem
RandomForest modelThe project can rank churn risk
Flask app with action bandsThe project is meant to trigger action

There are limits, and they matter. This is still a solo portfolio repo, not a team-hardened product. The notebook environment appears fragile, and the public footprint is small. But those limits do not weaken the central insight. They make it easier to see the idea in its pure form.

The idea is simple and strong: analytics is only valuable when it changes a decision. This repo gets there by letting SQL define the facts, Power BI frame the business, machine learning rank the risk, and Flask tell the operator what to do next.