Smart-Factory-Insights: The 7-Day Warning System Hidden Inside an Industrial IoT Pipeline

From Spark preprocessing to XGBoost predictions and operator-facing dashboards, this repo turns sensor noise into a maintenance decision the factory can use.

8-10 min read • View on GitHub • More from hardiksd7

A factory control desk where raw sensor sheets and telemetry streams enter from the left, pass through a central processing engine, and emerge on the right as a circled maintenance date and a machine status indicator. The image explains how noisy industrial data becomes a concrete maintenance deadline.
The repo’s real product is not a model score. It is a maintenance deadline that gives operators time to act.
Key Takeaways

The real product is a deadline

Most predictive maintenance demos stop at a score. This one tries to do something more useful: tell a maintenance team when to act. The repo’s defining move is a 7-day warning window, which turns model output into a scheduling problem instead of a curiosity.

That is why the project feels different from a generic industrial dashboard. It is not asking, “Will this machine fail?” in the abstract. It is asking whether the answer arrives early enough to order parts, plan downtime, and avoid a surprise stop.

56% less value than predicted. That is what the average large technology project delivers. The promise is rarely the problem. The delivery is. What great delivery means, the economics behind it, and 11 disciplines to get it right, including AI at scale. https://t.co/qpgNnni51f

Will Conaway, WillConaway1 · @WillConaway1 on X

Why 7 days is the sweet spot

A warning horizon is always a tradeoff. Too early, and the signal becomes expensive noise. Too late, and the plant has no time to intervene.

Generic predictive maintenanceSmart-Factory-Insights
Predicts failure somedayPredicts failure within 7 days
Single model outputBinary classification plus RUL regression
Raw score for analystsActionable warning for operators
Generic dashboard patternCustom UI plus Tableau
Notebook workflow tied to one environmentSpark preprocessing in Colab for easier reproduction

That seven-day band is the editorial center of the repo. It is just specific enough to be operational, but not so narrow that it turns into guesswork.

From raw telemetry to training data

The heavy lift lives in Spark_Preprocessing.ipynb. The notebook installs its own Spark runtime in Google Colab, pulls in raw CSV sensor data, defines schemas, and prepares a feature table for modeling. That choice matters: it makes a big-data style workflow runnable without a dedicated cluster.

The same feature table feeds two different questions: will it fail soon, and how soon is soon?

Spark is not here because the repo is pretending to be a Hadoop platform. It is here because sensor data gets messy fast, and the preprocessing step needs to feel like industrial ETL instead of a notebook toy.

# Representative setup pattern from the preprocessing notebook
!apt-get install -y openjdk-8-jdk-headless
!wget -q https://downloads.apache.org/spark/spark-3.2.1/spark-3.2.1-bin-hadoop3.2.tgz
!tar -xzf spark-3.2.1-bin-hadoop3.2.tgz

from pyspark.sql import SparkSession
spark = SparkSession.builder.master('local[*]').appName('SmartFactoryInsights').getOrCreate()

Two models, one decision

The modeling design is more interesting than the algorithm choice. The repo uses one branch to classify whether failure is imminent and another to estimate remaining useful life in days. XGBoost is a pragmatic fit here because tabular sensor features usually reward strong tree-based methods more than exotic architectures.

A close-up of two hands hovering over the same maintenance dashboard, one pointing to a failure warning card and the other to a remaining-useful-life gauge. A thin bridge connects both outputs, showing that the repository treats classification and regression as two parts of one decision.
The project does not choose between a warning and a forecast. It uses both to sharpen the maintenance call.
Binary failure modelRUL regression model
Answers yes or noAnswers how many days remain
Optimized for urgencyOptimized for planning
Useful for immediate interventionUseful for scheduling maintenance
Communicates threshold riskCommunicates trajectory
Best read with a deadlineBest read with a calendar

That pairing is the heart of the repo’s trust story. A binary alert says something is wrong. A remaining-life estimate says how much room the team has to respond.

The handoff from notebook to operator

The last mile matters here. The repository pairs a custom UI for instant predictions with a Tableau layer for fleet-level visibility, which keeps the operator workflow separate from the reporting workflow.

That split is smart. It avoids forcing one interface to do everything, and it respects the difference between a technician checking one machine and a manager scanning a plant.

What this repo gets right, and what it leaves open

StrengthTradeoff
Portable Colab-based Spark setupNot a true distributed deployment
Clear maintenance horizonHard-coded business assumption
Dual outputs for decision supportMore moving parts to validate
UI plus Tableau splitMore surface area to maintain
Clean, readable notebook pipelineEarly-stage project maturity

The project is strongest as a prototype of decision architecture. It compresses uncertainty into a short maintenance deadline, then gives that deadline two different machine-learning views and two different user-facing surfaces. What it does not yet prove is production scale, governance, or long-term model monitoring.

That limitation is not a flaw in the article’s core idea. It is the point. Smart-Factory-Insights shows how far a well-shaped pipeline can get when the goal is not just prediction, but action.