The Decision Pre Mortem: Find the Missing Signal Before You Commit

A decision pre mortem for support operations helps you catch missing signals before a workflow change ships. You’ll get a reusable one-page artifact, branch-level metric validation, and guardrails that prevent “green dashboard, red reality” surprises.

Lucía Ferrer
Lucía Ferrer
18 min read·

Run the “meeting-before-the-meeting”: a 30-minute pre mortem for any support ops change

Most support operations changes don’t fail loudly. They fail politely.

The dashboard stays green. Leadership moves on. Frontline folks build a workaround, stop mentioning it, and quietly absorb the tax. Then a month later you realize you traded one visible pain for three invisible ones: higher recontact, more escalations, quieter churn, and a team that now tiptoes around the new workflow like it’s a pothole nobody is allowed to acknowledge.

That’s the job of a decision pre mortem for support operations: a short, structured conversation you run before you commit, where the team assumes the change has already failed and asks: “What happened, what did we miss, and what signal would have warned us early?”

This isn’t pessimism. It’s operational honesty. Support is a business of fragments: one angry email, one confusing bot exchange, one region melting down while the global average looks fantastic.

What “missing signal” looks like in support (clean dashboards, messy reality)

A missing signal is anything that would change the decision if you could see it clearly, but you can’t—because it’s unmeasured, aggregated away, biased by tagging, delayed until customers are already upset, or trapped in agent notes nobody reads.

Concrete example: you enable auto close after 72 hours to reduce backlog. Aggregate metrics show “backlog down 18%” and “time to first response steady.” Everyone high-fives.

The missing signal is that reopen rate and repeat contact jumps for one queue—often complex billing—because customers reply on day four with the detail you needed. You didn’t reduce work. You deferred it, duplicated it, and made it angrier.

Another: you tighten a macro to be more direct (“per policy, we can’t…”) and you see fewer follow-ups. AHT drops. But escalation language spikes because the message reads like a parking ticket taped to someone’s windshield. The missing signal wasn’t efficiency. It was tone impact.

A useful (slightly annoying) question whenever an aggregate win gets celebrated: “Which branch could be failing while this number still looks good?” If nobody can answer, you just found your next pre mortem.

The trigger list: changes that deserve a pre mortem (automation, staffing, routing, SLAs, policy)

Use a pre mortem when the change is hard to reverse socially, even if it’s technically reversible.

It earns its keep for:

  • Automation: bots, auto close, deflection, macros, reply suggestions.
  • Staffing/coverage: weekends, on-call, tier splits, new schedules.
  • Routing: new queues, skills, regions, priority rules.
  • SLAs/KPIs: new targets, breach logic, escalation timing.
  • Policy: refunds, identity verification, what “resolved” means.

Decision framing that works: if the change affects (a) what customers experience first, or (b) what agents do most often, treat it as pre mortem-worthy. Those are the levers that create fast, silent failure.

The pre mortem question set you can copy into the calendar invite

Keep it small and slightly inconvenient. Thirty minutes is usually enough to expose weak confidence.

Use three prompts:

  • “It’s 30 days after launch and this change is considered a mistake. What’s the story of how it went wrong?”
  • “What metric is going to tell us everything is fine while reality is getting worse?”
  • “What is the one missing signal we need before we approve this?”

Decision rule for support ops: if the group can’t name at least one plausible failure story and one measurable early warning signal, you’re not ready to ship. You’re ready to hope.

One move that changes the tone fast: ask the Decider (or whoever is most excited) to say this out loud before discussion starts: “If we see X, I will pause.” Vague optimism hates deadlines.

Ship the one-page pre mortem artifact: roles, timebox, and the minimum inputs

The most common way teams ruin a support operations pre mortem is turning it into either a leadership vibe check or an analyst report-out.

Both fail for the same reason: they concentrate authority in one place and leave blind spots untouched.

The fix is a one-page artifact produced inside a timebox, with roles that create the right friction. Your goal is speed plus accountability—not paperwork.

Roles that reduce bias (Decider, Facilitator, Data Buddy, Frontline Rep, Skeptic)

Five roles. In a small team, one person can wear two hats, but don’t drop the Skeptic or the Frontline Rep. Those are your smoke alarms.

Decider owns the call and explicitly states what would change their mind. If they can’t say that, the meeting is theater.

Facilitator protects the timebox, pulls quiet voices in, and blocks solutioneering. This is the person who says, “We’re not fixing it yet. We’re naming how it fails.”

Data Buddy brings the minimum decision-grade numbers and can explain what’s noisy, what’s lagging, and what’s not tracked. Not “the person who built the dashboard.” The person who can tell you what the dashboard can’t.

Frontline Rep brings reality from the queue. Pick someone who still touches tickets or reviews quality weekly—not someone who “used to.”

Skeptic plays hostile investor for the hour. If you don’t assign this, the room will politely avoid conflict and then complain later in private, which is the most expensive form of feedback.

This is where teams get burned: running the pre mortem only with leadership because it feels “strategic.” You get alignment and confidence—and almost no contact with edge cases. Better mix: one leader, one operator, one frontline rep, one skeptic voice.

A 45–60 minute agenda with hard stops (and what to do if you run out of time)

Start with a clear promise: you’re here to decide whether the change is ready to ship—not whether it’s a good idea in the abstract.

A flow that holds up in real ops:

  • Set the decision statement and scope (what’s changing; which queues/channels/customers).
  • Name assumptions (each person offers one “this must be true”).
  • Write failure narratives silently, then share (silent writing prevents the loudest voice from setting the plot).
  • Translate narratives into missing signals (earliest observable warning for each story).
  • Add guardrails and owners (pause, rollback, escalation).
  • Make the call (approve, approve with conditions, or stop pending signal integrity).

If time gets tight, don’t “extend” and call it collaboration. Pick one: schedule a second session, or the Decider makes a provisional call with explicit unknowns recorded. Anything else becomes a slow-motion yes.

The one-page template: decision, assumptions, failure narratives, signals, guardrails, owners

The artifact should be short enough that someone can read it in two minutes and understand why you did what you did.

Use these sections:

  • Decision and scope
  • What success looks like (plain language)
  • Assumptions we are betting on
  • Failure narratives (3–5 “how it goes wrong” stories)
  • Missing signals and how we will detect them
  • Guardrails and kill criteria (pause/rollback triggers)
  • Owners and check cadence
  • Open unknowns (labeled explicitly)

Two filled examples:

Assumption: “Auto close at 72 hours will reduce backlog without hurting customer outcomes.”

Signal: “Reopen rate and repeat contact within 7 days by queue, with billing tracked separately.”

Owner: “Support ops lead reviews daily week one and posts a one-paragraph summary.”

Assumption: “New routing that sends ‘login issues’ to Tier 1 will reduce time to first response.”

Signal: “Escalation rate from Tier 1 to Tier 2 by region/language, plus top misroute reasons from sampling.”

Owner: “Routing owner + frontline rep review 20 conversations daily across branches for three days.”

Pre-work rules: what data is allowed, what gets labeled as “unknown,” and why

Pre-work should be strict. Otherwise you show up with a 40-slide deck that nobody reads and everyone pretends they did.

Rules that keep it decision-grade:

  • Bring only data that can influence a decision in the next two weeks. If it won’t change scope, guardrails, or monitoring, it’s background noise.
  • Any metric that isn’t trustworthy at the branch level gets labeled unknown.
  • Separate lagging indicators from leading ones. If proof of harm arrives after churn, you need a different signal.

Common failure: letting “unknown” become “we’ll look later.” Later is where accountability goes to nap. Instead, connect unknowns to a condition: “We ship only to one queue until tagging health is proven,” or “We pause until we can break out the metric by region.”

If you want broader context on the original pre mortem framing (and why it works), this is a useful skim: [1]

Truth-test branch signals before you trust the aggregate (queues, regions, teams)

Control Where it lives What to set What breaks if it’s wrong
Set: Guardrail: Pause decision if signal integrity issues cannot be resolved within a defined timeframe Decision-making process, risk management plan A maximum time limit — e.g., 48 hours to address signal issues before halting the decision process Rushing into poorly informed decisions, compounding errors, loss of credibility
Set: Define clear branch-level pass/fail tests — e.g., per queue, region, team Pre-mortem checklist, decision brief At least 5 specific, measurable tests for each branch, with clear thresholds Misleading aggregate metrics, incorrect resource allocation, biased decision-making
Set: Identify potential misleading aggregates and their underlying branches Pre-mortem analysis, data review sessions Specific examples of aggregates that hide branch-level issues — e.g., 'overall CSAT is high, but Region X is failing' Failure to address critical issues hidden by averages, wasted effort on non-problems
Set: Establish a 'kill criteria' for signal integrity failures Decision framework, pre-mortem protocol A clear rule: 'If X% of branch signals are invalid / missing, decision is paused / reverted' Committing to decisions based on unreliable data, loss of trust in metrics
Validate data collection methods at the branch level Data governance documentation, audit logs Regular checks on data input processes, sample sizes, and reporting consistency for each branch Inaccurate data, biased samples, inability to compare branches fairly
Set: Review for missing coverage in critical branches — e.g., new markets, niche products Metric dashboards, reporting requirements Mandatory reporting for all critical branches, even if data is sparse. flag 'no data' as a risk Blind spots in decision-making, neglecting high-risk or high-potential areas

Aggregates are comforting. They also lie in very specific, repeatable ways.

When someone says, “Support is fine, our SLA is fine,” they often mean, “The biggest queue looks okay.” Meanwhile a smaller queue is melting down, the language mix shifted, or one region is getting systematically misrouted.

This is where branch level metric validation stops being a nerdy analytics habit and becomes a risk control.

Branch-level truth tests: denominator sanity, tagging health, mix-shift detection

These truth tests are decision-grade because they have real pass/fail meaning:

  • Denominator sanity: do the counts match reality? Sudden drops in contact volume by branch without an operational explanation often mean tracking broke.
  • Tagging health: are tags/dispositions used consistently? If one team uses “customer education” as a bucket for everything, your slices are fiction.
  • Mix shift detection: did the customer mix change? A workflow update can look like a win if the queue got easier that week.
  • Branch parity: do key ratios look roughly similar across branches unless you have a known reason? Stable overall escalation rate can hide a single-region spike.
  • Latency symmetry: did one channel improve while another degraded? Email improving while chat worsens is a classic side effect.
  • Reopen vs recontact split: don’t trust one “resolution rate.” Separate “resolved and stayed resolved” from “resolved and came back.”

Don’t wait for perfect dashboards. A scrappy branch check beats an elegant global average—as long as you’re explicit about what’s noisy and what’s missing.

Coverage checks: what you are not measuring (hours, languages, channels, severity)

Missing signals often come from missing coverage. You measure what’s convenient, not what’s risky.

Two slices that uncover surprises fast:

Region and language. A routing change that helps English tickets can quietly harm non-English tickets because skill rules get rigid, translation workflows differ, and coverage isn’t symmetrical.

Channel and severity. A deflection change can reduce chat volume while pushing high-severity users into email, where they wait longer and escalate harder.

Rule that saves teams: if you can’t break out the metric by the branch that will carry the risk, you don’t have a metric. You have a mood.

And yes—this is where teams get burned—assuming “no data” means “no problem.” In support ops, “no data” usually means “the branch isn’t instrumented,” which is a risk by itself.

Conversation sampling that doesn’t lie (how to avoid survivor bias and “resolved-only” reviews)

Sampling is where teams accidentally gaslight themselves.

If you only review “resolved” conversations, you select for success. If you only review escalations, you select for drama. You need the boring middle.

A lightweight cadence most teams can sustain:

  • For each affected branch, review 10 conversations per day for the first three days, then 10 per branch twice in the next week.
  • Make sure at least two per branch are reopened, escalated, or transferred.
  • Include at least one auto-closed or bot-deflected conversation if automation is involved.

If the team can’t sustain that cadence, that’s not a moral failing. It’s a capacity signal. Shrink rollout scope.

The “green dashboard, red reality” patterns: when aggregates hide concentrated failures

Classic scenario: you split a queue into “Account access” and “Billing” to improve routing. Aggregate first response improves by 12%. Great.

But the billing sub-queue now has a backlog slope that climbs every day because routing rules undercount severity, and the few billing specialists can’t keep up. The missing signal was branch backlog slope by queue—not overall first response time.

Stop rules matter here. If you can’t trust the branch numbers, you pause.

A practical stop rule: if two or more core branch truth tests fail and can’t be resolved within one working day, don’t approve broad rollout. Narrow scope or delay.

This logic maps cleanly to kill criteria as a discipline: [2]

Decide what to automate vs keep human: where bots and macros amplify blind spots

Automation in support is like adding a faster conveyor belt to a factory. If the sorting is wrong, you just deliver mistakes at a more impressive pace.

A decision pre mortem is where you decide not only “can we automate?” but “what kind of failure does automation create, and will we notice it in time?”

A decision rule: automate when the failure is cheap and observable; keep human when it’s expensive and silent

Here’s the rule operators can actually use:

Automate when the failure is cheap and observable.

Keep a human when the failure is expensive and silent.

Cheap and observable looks like a wrong macro suggestion an agent can ignore, with a clear feedback loop.

Expensive and silent looks like a bot confidently applying the wrong eligibility rule, “resolving” the case, and you only find out later when refunds spike or cancellations rise.

Tradeoff to say out loud: automation often improves speed and consistency, but it can also hide error behind a clean resolution status. Humans surface nuance faster, but they introduce variance and “hero workflows.” You’re choosing which failure mode you prefer—and which you can detect.

Automation failure modes in support: exception collapse, tone drift, misrouting, and false resolution

Automation misses aren’t random. They cluster.

Exception collapse: edge cases get forced into the closest category, and the system stops admitting uncertainty. Early warnings: rising “other” tags, rising transfers, longer internal notes.

Tone drift: macros and bots become technically correct and emotionally wrong. Early warnings: sentiment and complaint language rising even if handle time drops.

Misrouting: routing rules optimize for speed but ignore skill nuance. Early warnings: escalation/transfer rate up in one region, team, or language.

False resolution: auto close or deflection reduces visible backlog while increasing reopens and repeat contacts. Early warnings: “resolved” rising while recontact rises.

This is why “AHT down” isn’t a victory song. It’s a clue. Sometimes a good clue, sometimes a trap.

Human judgment failure modes: inconsistency, over-escalation, and “local hero” workflows

Humans aren’t automatically safer. They fail differently.

Inconsistency shows up as policy drift: two customers get two answers. Early warning: higher variance in outcomes across agents and teams.

Over-escalation happens when agents are punished for being wrong, so they escalate to protect themselves. Early warning: escalations rise while quality doesn’t.

Local hero workflows are the sneakiest because they look like excellence. One experienced agent invents a workaround that saves the day, and suddenly your system depends on a person, not a process. Early warning: a subset of cases only resolves when one or two people touch them.

In your pre mortem, always ask: “Where does this require hero behavior to look good?” If the answer is “nowhere,” you might be rounding up.

How to run an exceptions-first review (and what exceptions tell you about missing signals)

Instead of reviewing average cases, review the exceptions first. Exceptions are where missing signals live.

Two examples that show up constantly:

A macro change adds firmer language to reduce back-and-forth. AHT improves because agents send fewer follow-ups. But escalations and complaint language rise because the macro sounds like a parking ticket. The missing signal wasn’t AHT. It was escalation reason codes—or a simple qualitative flag from daily sampling: “customer reacted negatively to tone.”

An auto routing change sends “login issues” to Tier 1. First response improves. But Tier 1 transfer rate spikes for one product line because “login issue” is often a permissions issue. The missing signal was branch transfer rate by product area, plus a quick scan of the top misroute reasons.

You don’t need new tools to capture this. You need an exceptions log for week one. Keep it simple: branch, what failed, customer impact, how you detected it, owner, what you changed.

Some teams use AI as an adversarial partner because it has no career incentive to agree with the boss. Whether you use AI or not, the adversarial posture is the point: [3]

Pre-commit checks: counterfactuals, guardrails, and what to monitor in the first week

The difference between a brave decision and a reckless one is rarely intent. It’s whether you planned the exit.

Support ops changes feel reversible because they’re “just workflows.” In practice, reversals are painful: agents retrain, customers adapt, reporting breaks, and leadership loses confidence. Pre-commit checks make reversibility real, not theoretical.

Counterfactuals: ‘If we do nothing, what happens?’ and ‘If we reverse it, what breaks?’

Counterfactuals expose missing signal risk fast.

Ask before approval:

  • If we do nothing for 30 days, what happens to backlog, SLA risk, and customer pain?
  • If we reverse this change after one week, what breaks—training, routing rules, customer expectations, reporting continuity?

If you can’t answer the reversal question without hand-waving, you’re about to ship a one-way door disguised as a toggle.

Guardrails that matter: kill switches, thresholds, and staged rollout criteria

Guardrails aren’t a list of metrics. Guardrails are action triggers.

A pre-commit set that actually holds up in week one:

  • Kill switch plan: who can pause, where it’s communicated, what “paused” means for agents today.
  • Scope control: start with one queue or one region until signal integrity is proven.
  • Threshold guardrails: tied to action, not debate.
  • Escalation path: who to pull in if a branch starts failing.
  • Customer comms stance: what you tell customers if the change causes confusion.

Example threshold (tune numbers to your baseline, keep the structure):

If reopen rate increases by a team-defined threshold in the affected queue for two consecutive days, you pause rollout, review 20 sampled conversations from that branch, and decide within one business day whether to rollback or adjust.

Guardrails for support automation should be measurable, owned, and timebound. Otherwise they’re just comforting words you say while shipping anyway.

First 24 hours vs first week: leading indicators (quality, recontact, backlog slope, escalations)

Don’t treat week one as a single blob. The first 24 hours is about safety. The first week is about learning.

In the first 24 hours, watch fast-moving indicators:

  • Backlog slope by branch
  • Escalations and transfers by branch
  • Top exception types from the exceptions log
  • Repeating agent feedback (patterns, not one-offs)

In the first week, add indicators that reveal customer impact:

  • Recontact within 7 days for the same issue category
  • Reopen rate split by queue and severity
  • Quality review outcomes from sampled conversations
  • Complaint language frequency in notes/tags (even if informal)

Notice what isn’t your north star: average handle time.

AHT is useful, but it’s a classic “we got faster by getting worse” metric unless you pair it with recontact and quality.

Practical decision rule: pick two “must not degrade” signals and treat them as your stop line. Teams that try to monitor ten things usually monitor none.

How to assign owners and cadence so monitoring actually happens

Monitoring fails for one boring reason: nobody owns the calendar.

Assign owners the way you assign incident roles.

A cadence that’s lightweight but real:

  • Daily for the first three days: Data Buddy posts a short branch snapshot; Frontline Rep posts two patterns from sampling; Decider confirms continue or pause.
  • Twice in week one: Facilitator updates the one-page artifact with what was learned, including new unknowns.
  • End of week one: 20-minute review of guardrails, exceptions, and whether to expand scope.

Document in the same one-page artifact with a dated addendum. If it lives somewhere else, it will be forgotten. If it’s forgotten, it might as well not exist.

Light humor, because we all need one: rolling out a support workflow change without guardrails is like “testing in production,” except your customers are the test suite and they don’t come with helpful error messages.

After you commit: turn every pre mortem into a decision library (so you stop relearning the same lesson)

A pre mortem compounds when you can retrieve it. Otherwise it’s just a one-time ritual that makes everyone feel mature.

The goal is a small decision library that makes the next approval faster because you can say, “We’ve seen this movie before, and we know which scene goes wrong.” Over time, pre mortems also work as calibration tools because you can compare what you predicted with what actually happened: [4]

The 10-minute post-decision addendum: what was true, what was missing, what changed

Keep a standing 10-minute addendum one week after launch.

Do three things:

  • Mark which assumptions were true, false, or still unknown.
  • Write the missing signal you discovered, especially if it surprised you.
  • Update guardrails for next time based on what actually moved.

This is where “we should remember that” turns into “we will not forget that.”

How to store and reuse pre mortems (tags: queue, change type, risk pattern, signals)

Don’t overthink tooling. A shared folder and consistent naming usually beat an ambitious system nobody maintains.

Use a few tags at the top of the artifact: queues affected; change type (automation, routing, staffing, SLA, policy); risk pattern (misrouting, tone drift, false resolution, mix shift); missing signals discovered; guardrails used; outcome after one week and one month.

Common mistake: letting pre mortems become write-only documents. If nobody can find the last three, you’re not building a library—you’re building a junk drawer.

Coaching loop: how to onboard new leads/operators using past misses

Here is a copyable example library entry:

Decision: “Enable auto close after 72 hours for low priority email queue.”

Branches affected: “Email, low priority, billing adjacent.”

Missing signal discovered: “Reopen rate by billing adjacent tag was the early warning.”

Guardrail: “Pause if reopen rate rises above baseline threshold for two days.”

Lesson: “Backlog went down, repeat contact went up, so we narrowed scope and changed the close message.”

This prevents the next bad commitment when someone proposes auto close again and claims, “We did this before and it was fine.” You can respond, calmly and with receipts, “It was fine overall. It was not fine in that branch. Here is what we watch this time.”

To make this real on Monday, don’t build a program. Do one thing.

Schedule the next “meeting before the meeting” and paste the one-page pre mortem headings into the invite. Then hold yourself to three priorities: (1) name one misleading aggregate you are currently trusting, (2) define two branch slices you’ll validate before shipping, and (3) set one guardrail with a clear pause action.

Your production bar is modest: one-page artifact, one owner per signal, and a first-week monitoring note that is written down where the team can find it.

Sources

  1. expectedvalue.co.uk — expectedvalue.co.uk
  2. howtothink.ai — howtothink.ai
  3. pre-mortem.ai — pre-mortem.ai
  4. howtothink.ai — howtothink.ai