Incident-Governance-Platform: The Alerting System That Decides Whether You Should Panic

A Django-backed incident engine that turns raw telemetry into governed severity, explainable escalation, and context-aware triage.

8 min read • View on GitHub • More from Dnyaneshwari40

A wide control-room scene shows a single telemetry pulse entering a sequence of mechanical gates. Each gate refines the signal into baseline, stress, severity, and explanation before it exits as a stamped incident verdict. The image explains that the repo is not just detecting anomalies, it is governing decisions.
Raw telemetry is not the end of the story. In this design, it becomes an incident only after it passes through a sequence of policy gates.
Key Takeaways

Why alerts become incidents only after interpretation

Most alerting systems are good at one thing: shouting. They can tell you that CPU crossed a line or latency spiked, but they still leave the hardest decision to a human. This repo is interesting because it treats that decision as software.

That is a meaningful shift. Instead of encoding a fixed threshold like CPU > 90%, the platform asks whether the signal is unusual relative to its own recent history, whether there is enough data to trust the judgment, and what response should follow. That is not monitoring. That is governance.

A split editorial illustration compares a rigid alarm switch on the left with a calibrated decision bench on the right. The left side fires whenever one dial crosses a hard red line. The right side weighs baseline, deviation, and sample size before producing a severity verdict. The image explains the difference between static thresholds and adaptive governance.
Static thresholds are simple. Adaptive governance is harder to fool.

The orchestrator: where raw telemetry turns into a verdict

The conceptual center of the codebase is the IncidentOrchestrator. It is the place where raw telemetry stops being a number and starts becoming a decision. Baseline computation, stress scoring, severity inference, and explanation all meet there.

One incident flows through context, scoring, and a final governed verdict.

The value of the orchestrator is not just that it connects modules. It makes the decision path legible. You can trace how a recent CPU spike becomes a stress score, how that score influences severity, and how the system explains itself in the final record.

class IncidentOrchestrator:
    def process(self, metric, context):
        baseline = self.baseline_engine.compute(context)
        stress = self.stress_engine.score(metric, baseline)
        severity, reason = self.severity_engine.decide(
            metric=metric,
            baseline=baseline,
            stress_score=stress,
            window_size=len(context),
        )
        return {
            "severity": severity,
            "stress_score": stress,
            "reason": reason,
        }

Bootstrapping is the smartest thing in the code

The best detail in the repository is the quietest one: MIN_WINDOW_SIZE = 30. Before the system has enough history, it refuses to pretend that it knows what normal looks like. That is the opposite of a brittle alerting rule.

This matters because early data is noisy by definition. A small sample can make a harmless spike look like a crisis. The fallback severity logic prevents the platform from overreacting before its own baseline has matured.

Generating illustration...

The system’s caution at startup is a feature, not a delay.

How the service layer turns one event into a governed record

The service layer is where the architecture becomes operational. In incident_service.py, the platform saves the raw metric, fetches recent context, runs orchestration, then writes the decision back into the database. That sequence is simple, but it is the difference between a demo and a traceable system.

This is the right split. Django handles persistence and delivery. The core engines stay stateless and focused on judgment. That separation makes the code easier to test, easier to reason about, and easier to swap out later.

def ingest_incident(payload):
    incident = Incident.objects.create(**payload)
    recent = Incident.objects.filter(
        created_at__gte=timezone.now() - ROLLING_WINDOW
    )
    verdict = orchestrator.process(payload, recent)
    incident.severity = verdict["severity"]
    incident.stress_score = verdict["stress_score"]
    incident.reason = verdict["reason"]
    incident.save()
    return incident

That pattern also creates an audit trail. Every incident can show both the raw values and the derived judgment that followed from them.

Explainability is not an afterthought here

The reason field is small, but it changes the product. The platform does not just label an incident LOW or CRITICAL. It records the logic that produced the label, which makes review possible.

That matters in incident response, where trust is earned by explanation. If the system says a metric is serious, operators need to know whether that is because the score was high, the baseline was low, the sample window was stable, or some combination of all three.

Escalation is a different problem from severity

Severity is about one incident. Escalation is about the system around it. The escalation engine looks for clusters of critical incidents inside a window, which is much closer to an SRE escalation policy than a simple alert rule.

DimensionStatic thresholdingAdaptive governance
Decision ruleFixed line in the sandBaseline, deviation, and context
Context sensitivityLowHigh
Early data handlingUsually noneBootstrapped fallback
ExplainabilityThreshold onlyReason string plus scores
Escalation behaviorOften separate and manualCluster-aware and policy driven
Operational fitNoisy environments struggleBetter for triage and review

That distinction is the whole product thesis. A single spike may deserve a warning. A cluster of critical spikes may deserve an escalation. The repo treats those as different questions.

What this repo gets right, and what still marks it as a prototype

The architecture is the strongest argument for the project. The split between core/, services/, and Django persistence is clean. The code reads like someone cared about single responsibility, which is rare enough to note.

The prototype signals are also obvious. Debug defaults, hardcoded credentials, and in-memory statistics point to an early-stage system. That does not weaken the idea. It just tells you where the work still needs to happen before the platform can carry real operational load.

What worksWhat still needs work
Clear separation between decision logic and persistenceReplace local statistical scans with streaming aggregates
Explainable severity recordsHarden configuration and secret handling
Bootstrapped fallback for cold startAdd scaling primitives for high throughput
Readable service-layer flowMove beyond prototype deployment defaults