Incident-Governance-Platform: The Alerting System That Decides Whether You Should Panic
A Django-backed incident engine that turns raw telemetry into governed severity, explainable escalation, and context-aware triage.
- Incident-Governance-Platform turns monitoring into a decision layer, not a prettier dashboard.
- Its core novelty is adaptive severity, which treats urgency as a function of context and sample size.
- The orchestrator and service layer separate raw metric ingestion from governed incident creation, which makes the logic easy to audit.
- The repo is strongest as a prototype of explainable triage, but its production gap is still visible in the usual places: hardcoded config, debug defaults, and in-memory statistics.
Why alerts become incidents only after interpretation
Most alerting systems are good at one thing: shouting. They can tell you that CPU crossed a line or latency spiked, but they still leave the hardest decision to a human. This repo is interesting because it treats that decision as software.
That is a meaningful shift. Instead of encoding a fixed threshold like CPU > 90%, the platform asks whether the signal is unusual relative to its own recent history, whether there is enough data to trust the judgment, and what response should follow. That is not monitoring. That is governance.
The orchestrator: where raw telemetry turns into a verdict
The conceptual center of the codebase is the IncidentOrchestrator. It is the place where raw telemetry stops being a number and starts becoming a decision. Baseline computation, stress scoring, severity inference, and explanation all meet there.
The value of the orchestrator is not just that it connects modules. It makes the decision path legible. You can trace how a recent CPU spike becomes a stress score, how that score influences severity, and how the system explains itself in the final record.
class IncidentOrchestrator:
def process(self, metric, context):
baseline = self.baseline_engine.compute(context)
stress = self.stress_engine.score(metric, baseline)
severity, reason = self.severity_engine.decide(
metric=metric,
baseline=baseline,
stress_score=stress,
window_size=len(context),
)
return {
"severity": severity,
"stress_score": stress,
"reason": reason,
}
Bootstrapping is the smartest thing in the code
The best detail in the repository is the quietest one: MIN_WINDOW_SIZE = 30. Before the system has enough history, it refuses to pretend that it knows what normal looks like. That is the opposite of a brittle alerting rule.
This matters because early data is noisy by definition. A small sample can make a harmless spike look like a crisis. The fallback severity logic prevents the platform from overreacting before its own baseline has matured.
Generating illustration...
How the service layer turns one event into a governed record
The service layer is where the architecture becomes operational. In incident_service.py, the platform saves the raw metric, fetches recent context, runs orchestration, then writes the decision back into the database. That sequence is simple, but it is the difference between a demo and a traceable system.
This is the right split. Django handles persistence and delivery. The core engines stay stateless and focused on judgment. That separation makes the code easier to test, easier to reason about, and easier to swap out later.
def ingest_incident(payload):
incident = Incident.objects.create(**payload)
recent = Incident.objects.filter(
created_at__gte=timezone.now() - ROLLING_WINDOW
)
verdict = orchestrator.process(payload, recent)
incident.severity = verdict["severity"]
incident.stress_score = verdict["stress_score"]
incident.reason = verdict["reason"]
incident.save()
return incident
That pattern also creates an audit trail. Every incident can show both the raw values and the derived judgment that followed from them.
Explainability is not an afterthought here
The reason field is small, but it changes the product. The platform does not just label an incident LOW or CRITICAL. It records the logic that produced the label, which makes review possible.
That matters in incident response, where trust is earned by explanation. If the system says a metric is serious, operators need to know whether that is because the score was high, the baseline was low, the sample window was stable, or some combination of all three.
Escalation is a different problem from severity
Severity is about one incident. Escalation is about the system around it. The escalation engine looks for clusters of critical incidents inside a window, which is much closer to an SRE escalation policy than a simple alert rule.
| Dimension | Static thresholding | Adaptive governance |
|---|---|---|
| Decision rule | Fixed line in the sand | Baseline, deviation, and context |
| Context sensitivity | Low | High |
| Early data handling | Usually none | Bootstrapped fallback |
| Explainability | Threshold only | Reason string plus scores |
| Escalation behavior | Often separate and manual | Cluster-aware and policy driven |
| Operational fit | Noisy environments struggle | Better for triage and review |
That distinction is the whole product thesis. A single spike may deserve a warning. A cluster of critical spikes may deserve an escalation. The repo treats those as different questions.
What this repo gets right, and what still marks it as a prototype
The architecture is the strongest argument for the project. The split between core/, services/, and Django persistence is clean. The code reads like someone cared about single responsibility, which is rare enough to note.
The prototype signals are also obvious. Debug defaults, hardcoded credentials, and in-memory statistics point to an early-stage system. That does not weaken the idea. It just tells you where the work still needs to happen before the platform can carry real operational load.
| What works | What still needs work |
|---|---|
| Clear separation between decision logic and persistence | Replace local statistical scans with streaming aggregates |
| Explainable severity records | Harden configuration and secret handling |
| Bootstrapped fallback for cold start | Add scaling primitives for high throughput |
| Readable service-layer flow | Move beyond prototype deployment defaults |