DevOps / SRE · Letter I

Incident Commander Role

The single named coordinator who runs a major incident — directing responders, owning communication and making decisions — so the best engineers can fix the problem instead of chairing a forty-person call.

By Dr. Hassan Eliwa, PhD · Founder of PMMilestone.org and PMMilestone.com · Updated 2026-08-21

Definition

The incident commander (IC) is the one person formally accountable for the response to a major incident — not for fixing the technical fault, but for running the effort that fixes it. The IC declares severity, assigns roles (tech lead, scribe, communications), maintains the incident timeline, decides when to escalate or roll back, and owns stakeholder communication until stand-down. The role borrows directly from emergency services command structures, and its first principle is counter-intuitive: the IC keeps their hands off the keyboard.

Why It Matters

Uncommanded incidents fail in predictable ways. Forty people join the call; the three engineers who could fix it spend the hour answering status questions; two responders independently restart the same service; nobody tells support what to say to customers; and an executive, starved of updates, starts directing technical work from the bridge. Mean time to mitigate stretches not because the fault is hard but because the response is unmanaged. A trained IC converts that crowd into an organisation: one decision-maker, one timeline, one voice outward, and specialists left alone to specialise.

How the Role Works in Practice

  1. Declare and take command explicitly — "I'm IC, tech lead is Ana, scribe is Ben, comms goes out every 15 minutes from me."
  2. Set the severity and the structure — one bridge or channel, roles named, bystanders thanked and released.
  3. Keep the rhythm — regular update cadence to stakeholders, even when the update is "no change, next update at :30."
  4. Assign, don't absorb — every hypothesis gets a named owner and a deadline; the IC tracks, the owner investigates.
  5. Make the hard calls — roll back versus fix forward, fail over versus ride it out — with the information available, at the time it is needed.
  6. Hand over or stand down cleanly — long incidents get structured IC handovers; every incident ends with a declared close and a review scheduled.

Real-World Example

A payments platform suffered a database failover that left checkout erroring at 09:12 on a Tuesday. Within minutes forty people had joined the incident call, including two VPs. The three engineers who understood the replication topology spent forty-five minutes narrating their screens to management while two well-meaning responders applied contradictory fixes. Mitigation came at 74 minutes. The post-incident review introduced a trained IC rotation. Three months later a similar failure was declared at 14:03, roles assigned by 14:05, customer comms out by 14:20, rollback decision made at 14:16 — and checkout was healthy at 14:21. Same systems, same engineers, same fault class. The eighteen-minute outcome was not better technology; it was one person whose entire job was to run the response.

Practical Lessons Learned

  • The IC must not debug. The moment the commander starts reading logs, nobody is commanding — hand the hypothesis to an owner and keep the picture.
  • Announce the cadence and keep it. Stakeholders who know the next update lands at :15 and :30 stop interrupting; stakeholders in the dark interrupt constantly.
  • A scribe doubles the IC's effectiveness. The timeline, decisions and actions captured live are the post-incident review half-written.
  • Small is fast. Release everyone not actively working a task; a twelve-person bridge is a meeting, not a response.
  • Severity drives structure. Not every page needs an IC — a Sev-3 with one responder just needs a ticket. Declare command proportionally.

Expert Tips

  • Run game days quarterly where the IC role rotates through engineers and engineering-adjacent staff — command is a practised skill, and the quiet hour of a drill is where people learn it.
  • Give the IC explicit authority in writing: they may pull anyone from any meeting, and their rollback decision stands unless a named executive overrules on the record.
  • Script the first five minutes — declaration phrases, role assignments, comms template. Under stress, nobody improvises good process.
  • Hand over command formally for incidents crossing hours: current state, open actions, next decision point, next comms slot. Ten minutes of structure saves an hour of re-discovery.
  • Close every incident with the same three artifacts: timeline, customer impact statement, review date. Future-you, writing the board report, will be grateful.

Common Mistakes

  • The most senior engineer appoints themselves IC and then fixes the fault personally — the response has no commander precisely when it needs one.
  • Executives freelancing in the channel, asking for updates off-cadence and redirecting responders mid-investigation.
  • No declared severity, so a company-wide structure gets applied to a minor fault — or worse, no structure to a major one.
  • The IC making decisions silently, so responders work at cross purposes because the plan lives in one head.
  • Skipping the review because "it was only twenty minutes" — short incidents are the cheapest training data you will ever get.

Key Takeaways

  • Major incidents need one named commander whose product is decisions and communication, not code.
  • The IC assigns and tracks; specialists investigate and fix. Mixing the roles halves both.
  • A visible update cadence is what keeps stakeholders out of the responders' way.
  • Authority, scripts and drills make command real — titles alone do not.
  • Every incident closes with a timeline, an impact statement and a scheduled review.

Related Concepts

Pairs with Incident Management, Post-Incident Review, On-Call Rotation, and Blameless Postmortem.

Frequently Asked Questions

  • Should the incident commander be the most senior engineer?
    Usually no. Command and deep diagnosis are different skills, and your strongest diagnostician is more valuable on the tools. The IC needs decisiveness, communication and calm — train a broad rotation, and deliberately assign senior engineers as tech leads rather than commanders during major events.
  • Does every incident need an incident commander?
    No — proportionality matters. A single-responder Sev-3 needs a ticket, not a command structure. ICs earn their overhead at Sev-1 and Sev-2, where multiple responders, customer impact and stakeholder pressure make coordination the binding constraint. Declare command when the response, not the fault, becomes complex.
  • Isn't appointing an IC just adding bureaucracy to an emergency?
    That is the common objection, and the data runs the other way: teams that adopt incident command consistently report large drops in time-to-mitigate. The apparent overhead — declaration, roles, cadence — replaces a far larger hidden cost: forty uncoordinated people interrupting the three who can fix it.
  • What if the IC makes the wrong call mid-incident?
    A reasonable decision made on time beats a perfect decision made late. The IC decides with the information available and records the reasoning via the scribe; the post-incident review is where calls get examined, blamelessly, with better data. Punishing a wrong-but-reasonable call guarantees no one accepts command again.
  • How do you hand over incident command on a long incident?
    Formally: the outgoing IC briefs current state, open actions with owners, the next decision point and the comms cadence, then announces the transfer on the bridge and in the stakeholder channel. Ten structured minutes beats an hour of the new commander re-discovering context while responders wait.
  • Who writes customer communications during an incident?
    The IC owns outbound communication but usually delegates drafting to a comms lead — support or product — using pre-approved templates. Engineers should never be writing customer copy mid-incident; their attention is the scarcest resource in the room, and comms is a specialist task with its own skills.
  • Which calculators on PMMilestone.org apply to Incident Commander Role?
    For Incident Commander Role, the most relevant tools on the flagship platform are the EVM, SPI and CPI calculators on PMMilestone.org. They reproduce the formulas referenced in this entry against your own project data.
  • What is a common misconception about Incident Commander Role?
    That the topic is well-defined across all references. In practice, definitions vary between PMBOK, PRINCE2, AACE and ISO 21500 — this entry uses the definition most aligned with field practice on capital projects, and flags where the standards diverge.
  • Which related encyclopedia entries should I read alongside Incident Commander Role?
    Read Earned Value Management, Critical Path Method and the DCMA 14-point assessment next. The full A–Z is available in the PMMilestone Encyclopedia, and quick one-line definitions live in the PM Glossary on the flagship platform.
  • How does Dr. Hassan Eliwa's research treat Incident Commander Role?
    Dr. Hassan Eliwa's research focuses on owner-side project controls, schedule integrity and forensic delay analysis on capital construction and power programmes. Incident Commander Role is treated through that lens — what a planning or controls engineer is expected to do with it on a live project, not its textbook definition alone. See the full research library at PMMilestone Research Articles.
  • How is Incident Commander Role defined on PMMilestone Research & Insights?
    The single named coordinator who runs a major incident — directing responders, owning communication and making decisions — so the best engineers can fix the problem instead of chairing a forty-person call. For the full treatment, see the definition, principles, applications and related entries above — every encyclopedia entry follows the same research-grade structure.

People also ask

Follow-up questions practitioners search for next — each one points to the calculator, template or reference entry that answers it.

Related Entries

Browse more in this category

More in DevOps / SRE

View all DevOps / SRE entries →

Further reading on PMMilestone.org

Curated companion resources hosted on the flagship platform, PMMilestone.org.

Related Encyclopedia Entries
Research Articles
Career Guides
Tools on PMMilestone.org
Buy me a coffee