Incident Severity Classification
The shared scale that turns a vague sense of urgency into a defined response — who is woken, who is told, and how fast, decided before the incident rather than during it.
Definition
Incident severity classification is a small, published set of levels — commonly SEV1 to SEV4 — each defined by customer impact and each attached to a specific response: who is paged, who leads, who communicates, how often updates are issued, and whether a formal review follows. The scale exists so that the first person to see a problem does not have to negotiate the response while the problem is still growing.
Why It Matters
The expensive minutes in most incidents are the first ten, and they are usually spent deciding whether this is serious. An engineer sees elevated errors, wonders whether waking the database lead is justified, checks another dashboard, messages a colleague who is asleep, and twelve minutes disappear. A severity scale removes that deliberation by pre-answering it. It also protects people in the other direction: a defined SEV3 means nobody is paged at 3 am for something that can wait for business hours, which is what keeps an on-call rotation sustainable.
A Workable Scale
- SEV1 — critical. Core function unavailable or data at risk for many customers. Immediate page, incident commander assigned, customer communication within 30 minutes, updates every 30 minutes, blameless review mandatory.
- SEV2 — major. Significant degradation or a critical function broken for a subset. Page during and outside hours, named lead, hourly updates, review required.
- SEV3 — minor. Limited impact with a workaround. Business-hours response, tracked to resolution, no page.
- SEV4 — low. Cosmetic or single-user issues. Normal backlog handling.
Two rules make the scale work. Anyone may declare any severity, and over-declaring is explicitly not punished. And severity can be downgraded once impact is understood — declaring SEV1 and reclassifying at minute eight is a success, not an error.
Real-World Example
A logistics SaaS platform had a four-level scale defined in a wiki page nobody had read since onboarding. During a checkout outage, the first responder judged it "probably a SEV2" because only one region was affected, so no incident commander was appointed and no customer notice went out. The outage was in fact global for anyone using a particular payment method — around 34% of revenue — and ran for 71 minutes before someone senior noticed the revenue graph and escalated to SEV1.
The blameless review found the definitions were written in infrastructure language ("multiple availability zones affected") rather than customer language, so the responder had no way to map what they saw to a level. The rewrite defined each level by customer-visible outcome — "customers cannot complete a purchase" — added a revenue-impact hint, and put the scale on a single card in the incident channel topic. Median time to declare across the next quarter fell from roughly 14 minutes to under 4.
Practical Lessons Learned
- Define severity by customer impact, not infrastructure state. Responders can see symptoms, not architecture diagrams.
- Make over-declaration safe and cheap. If declaring SEV1 triggers an executive inquest, people will hesitate exactly when hesitation costs most.
- Four levels is plenty. Six-level scales produce arguments about the boundary between them.
- Put the scale where the incident happens. A wiki page is not a control; a pinned card in the incident channel is.
- Track time-to-declare as a metric. It is often a bigger lever than time-to-fix.
Expert Tips
- Add a plain-language example to each level drawn from a real past incident. Examples classify faster than criteria.
- Automate what the declaration does: creating the channel, paging the rota, opening the timeline document. Manual choreography loses minutes.
- Review a sample of last quarter's incidents against the scale. Systematic under-classification is common and easy to see in hindsight.
- Tie communication obligations to severity explicitly, so customer updates are not a separate judgement call under pressure.
- Let the incident commander, not the engineer fixing it, own severity changes during the incident.
Common Mistakes
- Writing severity definitions in internal system terms that responders cannot map to symptoms.
- Treating a SEV1 declaration as an accusation, which trains people to under-declare.
- Having no defined communication or review obligations per level, so the scale changes nothing.
- Allowing severity to drift upward silently without appointing a commander.
- Maintaining the scale in documentation only, never in the tooling that runs the response.
Key Takeaways
- Severity exists to pre-decide the response, not to describe the fault.
- Define levels by customer impact in plain language, with real examples.
- Anyone can declare; over-declaring must be safe.
- Attach paging, leadership, communication and review obligations to each level.
- Measure time-to-declare — it is often where the minutes go.
Related Concepts
Works with Incident Commander Role, Incident Management, Blameless Postmortem, and Alert Fatigue Management.
Frequently Asked Questions
How many severity levels should a team have?
Four is the practical sweet spot: one critical, one major, one minor and one low. Fewer forces genuinely different situations into the same response; more creates boundary debates during the exact minutes when nobody should be debating anything.Who is allowed to declare an incident?
Anyone who sees a problem, at any severity. Gating declaration behind seniority adds delay at the worst possible moment. The incident commander can adjust severity once impact is understood, which is the correct place for that judgement.Is downgrading a severity a sign of a mistake?
No — it is the system working. Declaring high and reclassifying within minutes costs a little disruption; declaring low and discovering an hour later that revenue was down costs a great deal more. Teams should say so publicly to reinforce the behaviour.Should severity be based on customer impact or system impact?
Customer impact, always. Responders can observe symptoms — checkout failing, dashboards blank, uploads rejected — far more reliably than they can assess architectural blast radius. Infrastructure detail belongs in the diagnosis, not the classification.How does severity relate to priority?
Severity describes current impact; priority describes the order of work. They usually align during an incident but diverge afterwards: a resolved SEV1 may leave a low-priority cleanup task, while a persistent SEV3 with no workaround can warrant high priority in the backlog.What should each severity level trigger automatically?
At minimum: the paging path, an incident channel, a timeline document, and the communication cadence. Automating those removes several minutes of choreography per incident and makes the scale a working control rather than a document.Which calculators on PMMilestone.org apply to Incident Severity Classification?
What is a common misconception about Incident Severity Classification?
That the topic is well-defined across all references. In practice, definitions vary between PMBOK, PRINCE2, AACE and ISO 21500 — this entry uses the definition most aligned with field practice on capital projects, and flags where the standards diverge.Which related encyclopedia entries should I read alongside Incident Severity Classification?
Read Earned Value Management, Critical Path Method and the DCMA 14-point assessment next. The full A–Z is available in the PMMilestone Encyclopedia, and quick one-line definitions live in the PM Glossary on the flagship platform.How does Dr. Hassan Eliwa's research treat Incident Severity Classification?
Dr. Hassan Eliwa's research focuses on owner-side project controls, schedule integrity and forensic delay analysis on capital construction and power programmes. Incident Severity Classification is treated through that lens — what a planning or controls engineer is expected to do with it on a live project, not its textbook definition alone. See the full research library at PMMilestone Research Articles.How is Incident Severity Classification defined on PMMilestone Research & Insights?
The shared scale that turns a vague sense of urgency into a defined response — who is woken, who is told, and how fast, decided before the incident rather than during it. For the full treatment, see the definition, principles, applications and related entries above — every encyclopedia entry follows the same research-grade structure.
People also ask
Follow-up questions practitioners search for next — each one points to the calculator, template or reference entry that answers it.
Where is this in the glossary?
Quick-lookup definitions across 1,200+ PM terms. PM Glossary on PMMilestone.org ↗
Which learning track covers this end-to-end?
Structured tracks from beginner planner to programme controls director. Project Controls Academy ↗
Which book goes deeper than this entry?
Practitioner field handbooks with worked numerical examples. Books & Publications ↗
Which calculator on PMMilestone.org applies here?
The integrated EVM workbook covers most cost-schedule diagnostics. EVM Calculator ↗
Related Entries
More in Engineering Practice
- Letter AAPI Versioning Strategy
The agreed way an interface evolves without breaking the clients already depending on it — a product decision that determines how much of the team's future is spent on backward compatibility.
- Letter DDeveloper Onboarding Ramp
The deliberate path from a new engineer's first day to independent, confident contribution — measured in days to first merge and weeks to first on-call shift.
Further reading on PMMilestone.org
Curated companion resources hosted on the flagship platform, PMMilestone.org.
- For practitioners who want to go deeper, the Project Controls Academy.
- Engineers researching this topic typically continue with the Learning Tracks.
- A practical companion to this entry is the Books & Publications.
- Closely related on the flagship platform is the EVM Calculator.
- Useful alongside this article is the Schedule Health Checker.
- Many readers follow this up with the PMMilestone.org knowledge hub.