The Pre Mortem for Decisions: Questions That Expose Bad Assumptions and Missing Signals

A practical decision pre mortem for support operations leaders. Use decision pre mortem questions for support operations to surface risky assumptions, spot missing signals, set decision gates, and avoid quiet failures in staffing, routing, macros/automation, escalations, and SLAs.

Lucía Ferrer
Lucía Ferrer
18 min read·

When a support-ops decision fails quietly, it’s usually missing signals—not bad intent

Everyone has seen the “successful” support ops change that somehow makes life worse.

You tweak routing to improve first response time. The dashboard obliges. FRT drops. Leadership relaxes. Then, two weeks later, the floor feels… different. Agents are doing more transfers. Reopens creep up. The escalation channel has more traffic than usual. One queue starts aging, but not loudly enough to trigger an incident. Nothing is “broken,” so it quietly becomes the new normal.

That’s the pattern: support operations decisions fail quietly when you’re missing signals, not because the team is reckless. Intent is usually good. The plan is often reasonable. Visibility is what’s broken.

A decision pre mortem is a short, structured exercise where you assume the decision failed in the future, then work backward to name the most likely reasons. In support ops terms, it forces the room to say out loud where the work will go, which customers get the worst version, and what you’ll see first.

A fast pre‑mortem beats a slow rollback

In support operations, “rollback” is rarely a clean reversal. Customers already experienced the change. Agents already built coping strategies. Stakeholders already got used to the new numbers. Undoing it becomes slow, political, and expensive.

A 30‑minute pre mortem is the closest thing you get to rehearsing failure without paying full price.

If you’ve ever read a postmortem and thought, “How did we not see that coming,” this is the moment you try to see it coming.

The four decision types that deserve a pre‑mortem (staffing, routing, macros/automation, escalations/SLAs)

Not every tweak needs ceremony. But these changes usually do, because they reshape the work and the incentives:

Staffing and capacity (schedules, shrinkage assumptions, adding chat), routing (queue logic, priority rules, language splits), macros/automation (templates, deflection, auto‑closes), and escalations/SLAs (handoff thresholds, breach policy, promise language).

You don’t run a pre mortem because you’re anxious. You run it because these changes can “work” on paper while quietly shifting pain somewhere else.

How ‘polished noise’ shows up in dashboards and conversation samples

Support teams drown in polished noise: numbers and anecdotes that look crisp but aren’t decision‑grade.

An FRT win can hide a transfer spiral. Flat CSAT can hide a segment meltdown if your sample is skewed. Conversation reviews can become polished noise too if you only look at clean tickets from happy paths.

A pre mortem gives the room a useful interruption when someone drops a tidy metric:

Is that decision‑grade, or is it polished noise?

Keep one concrete quiet failure in mind as you read. A routing change improves FRT by sending more tickets to generalists. Misroutes increase. Misroutes create transfers. Transfers stretch time to resolution. Longer resolution drives reopens. Reopens trigger escalations. The top‑line metric looks better while the customer experience gets worse.

That’s exactly what decision pre mortem questions for support operations are meant to catch.

A practical habit that pays: when a metric “improves,” ask what had to get worse for that improvement to happen. If the room can’t answer, you probably haven’t found the cost yet.

Run the 30‑minute pre‑mortem: scope the decision, timebox, and assign roles so it stays sharp

Pre mortems get a bad reputation for one reason: the meeting drifts. People brainstorm vague risks, debate hypotheticals, and leave with nothing you can operate.

A support ops pre mortem has one job: produce a one‑page output with assumptions, missing signals, gates, monitoring, and rollback. If it doesn’t do that, it’s just a discussion with a timer.

Write the decision as a reversible bet (what changes, for whom, by when)

Start by writing the decision in one sentence with constraints. The goal is a reversible bet, not a motivational poster:

Decision statement: “We will change X for Y segment through Z channel, starting date, to improve primary outcome, while keeping guardrail metrics within bounds, and we will review at time window for ship, adjust, or rollback.”

In support ops, constraints are where the plan quietly breaks, so name them up front: hours of coverage, languages, regulated queues, VIP tiers, peak season, tooling limits, and any “we can’t do that in this system” reality.

If you can’t name the segment and time window, you’re not making a decision. You’re making a wish.

Pick the smallest ‘blast radius’ you can test

Teams sometimes ask, “How do we test routing without building an experiment platform?” Most of the time you don’t need fancy tooling. You need a smaller blast radius.

Constrain one dimension (channel, region/language, issue category, or tier) so monitoring and rollback are real. This is where teams get burned: they launch broadly to “keep it simple,” then discover their dashboards can’t explain what changed. Small blast radius isn’t cautious for cautious’ sake. It’s what makes learning possible.

Roles: facilitator, skeptic, data scribe, frontline reality check

A 30‑minute pre mortem stays sharp when roles are explicit. Not because you love process—because it prevents the two classic failure modes: endless debate and wishful thinking.

The facilitator keeps time and stops solutioning too early. The skeptic is allowed to be annoying on purpose, but must propose a signal for every critique (“If that’s true, what would we see by Wednesday?”). The data scribe writes the one‑page output in real time, including owners and cadence. The frontline reality check (agent, lead, QA partner) is the person who can say, “That’s not how tickets arrive,” or “That macro will create three follow‑ups.”

Four to six people is the sweet spot: one ops owner, one frontline rep, one QA/enablement partner, one affected stakeholder (product/engineering/sales). More than that, you start getting performance instead of truth.

A common mistake: inviting only managers. You’ll get polished‑noise risks like “customer confusion” instead of concrete breakpoints like “transfers spike because the tags overlap.” Bring someone who still touches real conversations weekly.

The one-page pre‑mortem output: assumptions, missing signals, gates, rollback

You don’t need a complicated ritual. You need a rhythm that forces specificity before you ship.

Start with the decision statement and blast radius. Then do a short silent write: “It failed. What happened first?” People write failure modes and the first signal they’d expect. Share, cluster, and for each cluster name (1) the missing signal and (2) the action you’ll take. Finish by pre‑committing gates and rollback criteria with owners.

The rule that keeps this honest: if you can’t name the signal, don’t ship the change. That doesn’t mean perfect data. It means you don’t launch on vibes alone.

Two setup examples, because abstraction is cheap and tickets are not.

Staffing/capacity example: you want to reduce weekend coverage to save cost. A pre mortem surfaces hidden assumptions like “weekend volume is predictable” and “backlog clears on Monday without impacting SLA.” Missing signals usually include backlog aging by hour‑of‑week, not just total backlog.

Routing + macro rollout example: you’re routing billing issues to a new specialist queue and rolling out a macro bundle for those specialists. A pre mortem surfaces assumptions like “tags are accurate enough to route” and “specialists have capacity for the inflow.” Missing signals tend to be misroute rate, transfer rate, and the ugly end of handle‑time distribution for billing issues. If you don’t track those, you’re setting up a slow‑motion incident.

Don’t leave with “we’ll monitor it.” Leave with who’s looking, how often, and what they’ll do if it moves.

A solid baseline reference for the premortem format is here: [1]

The question framework: 6 buckets that expose bad assumptions and missing signals before you ship

Assignment strategy Best for Advantages Risks Recommended when
Assumption Challenge: What is our riskiest assumption, and how can we validate it? Decisions built on untested hypotheses (e.g., 'agents adapt quickly') Exposes weak points. encourages data-driven validation. reduces speculative risk Endless debate. validation costly/time-consuming Before implementing new tools, policies, or training programs
Rollback Plan: What is the exact process to revert this decision if it fails? Any decision with potential negative impact or uncertainty Reduces failure fear. safety net. clarifies recovery steps Rollback complex/impossible. false security if untested Any change to critical support workflows (routing, escalation, SLA)
Silent Failure Indicators: How would we know if this decision is failing quietly? Detecting gradual degradation or non-obvious negative outcomes Reveals hidden problems. early intervention. protects customer experience Requires deep system understanding. hard to define Any change affecting agent efficiency, customer satisfaction, or system performance
Decision Gate: What specific criteria must be met before this decision can proceed? High-impact changes — e.g., new routing, staffing model, SLA adjustment Clear success metrics. prevents premature deployment. identifies blockers early Slows agile teams. bureaucracy if overused. subjective criteria Decision has significant cost, customer impact, or irreversible consequences
Sampling Bias: Is our current data representative of all customer segments or scenarios? Decisions based on A/B tests, surveys, or limited pilots Ensures fairness. prevents skewed results. identifies underrepresented groups Over-complication of data collection. difficulty achieving true randomness Assessing impact of macro changes, new self-service options, or staffing models
Missing Signal: What data points are we NOT collecting that would indicate failure? Identifying blind spots in monitoring and reporting Uncovers silent failures. improves data quality. proactive risk detection Analysis paralysis. collecting irrelevant data Evaluating new metrics, dashboards, or post-launch monitoring plans

Most teams ask the wrong pre mortem questions because they ask them like strategists, not operators.

An operator asks: where will the work go, which segment gets the worst version, what breaks first, and what would make us stop.

Use these six buckets every time. Run them in order when the change is big. Cherry‑pick when it’s smaller. The outcome stays the same: fewer surprises, faster detection, cleaner reversals.

Bucket 1: What must be true for this to work?

These are your assumption grenades. If one is wrong, the decision fails in practice.

What must be true about tagging quality for this routing change to work? About forecast accuracy and shrinkage for this staffing plan to hold? About agent proficiency for this macro rollout to reduce resolution time rather than increase reopens? About engineering response time for a new escalation rule to mean anything?

Decision rule: if the riskiest assumption is untestable (or only testable after full rollout), treat the decision as higher risk and shrink the blast radius.

Bucket 2: What breaks first (and where does the work move)?

Work doesn’t disappear in support. It migrates.

If the new specialist queue overloads, does work spill back to generalists, into escalations, or into backlog aging? If you shorten an SLA, does the work move into quick replies, extra follow‑ups, and more reopens? If automation deflects easy tickets, does what remains become higher complexity—and therefore slower—even if volume drops?

This bucket prevents the classic outcome: you “improved” metric X by creating metric Y that you didn’t plan to staff for.

Bucket 3: What are we optimizing, and what will we accidentally worsen?

This is where you stop pretending you can improve everything at once.

If you improve FRT, what quality metric gets worse first? If you reduce cost per contact, which customer segment pays? If you increase specialization, what happens to flexibility during spikes?

A practical rule: pick one explicit guardrail that protects customers, and one that protects the team. Teams get burned when they treat burnout as a “people issue” instead of a predictable outcome of moving work around.

Bucket 4: Which segment will experience the worst version?

Support changes rarely fail evenly. They fail on the edges.

Which customers have the highest‑effort path through the new routing? Which language, region, tier, or issue type will be last to benefit? Which channel gets the most templated responses and the lowest empathy after a macro rollout?

If you can name the “worst version” customer, you can usually name the first signals too.

Bucket 5: What are we not measuring (or measuring wrong)?

This is the missing signal bucket. It’s usually where the pre mortem pays for itself.

Are you measuring FRT without measuring transfers, misroutes, and reopens? Are tags drifting because agents are incentivized to pick the fastest category? Are you staring at CSAT averages without checking the distribution and the “very dissatisfied” tail? Are you sampling only solved tickets and calling it representative?

Bucket 6: What would make us stop or roll back?

This is the part most teams skip, then regret.

What is the decision gate to proceed past pilot? What is the rollback trigger that’s specific enough to use under pressure? Who owns the call, and what’s the review cadence?

Two examples that show how one question exposes a missing signal.

Routing change example: you ask, “What must be true about categorization accuracy?” The team says, “Tags are fine.” The frontline reality check says, “We pick the closest tag because the right one is buried.” Now you realize you have no baseline misroute rate, and you’re about to scale routing built on optimistic tagging. The fix isn’t philosophical: pilot where tagging is audited, and watch transfers and reopens daily.

SLA change example: you ask, “If we tighten response time, what behavior will agents adopt?” Someone says, “They’ll send more quick acknowledgements.” Great—maybe that helps perceived responsiveness. Or maybe it increases customer effort because they get more back‑and‑forth without progress. The decision is speed vs quality vs backlog. Pre‑commit guardrails: don’t accept a faster SLA if reopens and time to resolution worsen beyond a set band for two weeks.

A broader cultural take on premortems (still worth reading, even if you keep your own prompts operator‑tight) is here: [2]

What to trust vs what to measure: separating decision-grade signals from ‘polished noise’

Support ops is full of data. The problem isn’t volume. It’s trust.

Decision‑grade signals are reliable enough that you’d bet a launch on them. Polished noise is what looks neat in a meeting and collapses under pressure.

If you want your support ops pre mortem to actually prevent problems, you need a filter for what to trust, what to validate, and what to stop pretending is meaningful.

Dashboards lie by omission: metric definitions, tagging drift, and survivorship bias

Most “dashboard lies” are lies by omission.

FRT improves while transfers and reopens worsen (you answered quickly, then bounced the customer). Average handle time improves while the tail gets worse (specialists are overloaded). Backlog looks stable while backlog aging explodes (customers feel aging, not totals). CSAT stays flat while detractors concentrate in one segment (averages hide meltdowns).

A simple practice that saves teams: every metric you use as a gate gets a plain‑language definition next to it. If you can’t define it, you can’t gate on it.

This is where teams get burned in a sneakier way: they trust a metric that changed definition mid‑stream—new tagging rules, a tooling migration, a different CSAT trigger—then treat before/after like clean truth. If the definition isn’t stable, it’s not a gate. It’s trivia.

Conversation sampling that doesn’t fool you (how many, which segments, what to code)

Conversation reviews are powerful—and dangerously easy to bias.

Polished noise example: someone pulls ten “representative” tickets, but they’re all solved, all daytime, all from the same product area, and none include angry customers or multi‑touch cases. The macro rollout looks “fine.” Meanwhile, night shift and non‑English queues absorb the pain.

You don’t need a research lab. You need intentional variety: include at least one high‑value tier, one long‑tail tier, one under‑represented region/language, and one messy issue type; review both good and bad outcomes (reopens, escalations, transfers); and code with a small repeatable lens: customer effort, clarity of next step, empathy/tone, and correctness.

Before you conclude anything, ask: what didn’t we look at that might be the worst version? That one line prevents the feel‑good sample set.

Leading indicators for trouble: where backlog and quality regressions show up first

Support leaders often wait for lagging indicators like CSAT, churn, or quarterly retention. By the time those move, you already paid the price.

Leading indicators that tend to show trouble early include reopen rate (especially within 72 hours), transfer rate (including multi‑hop chains), escalation rate (and reason codes), backlog aging (not just size), QA defect rate (correctness and policy adherence), and handle‑time tails (not just averages). In conversation samples, watch for “I already told you” and “still waiting.” Those phrases show up before the quarterly deck does.

For any high‑impact change, set a short early warning window—first 3–5 business days—where you look daily. Quiet failures don’t wait politely for your weekly metrics meeting.

A simple ‘signal ladder’ to decide if the data is good enough to ship

Here’s a compact way to decide if a signal is decision‑grade.

Level one: it’s defined and stable. Level two: you can segment it by channel, tier, and issue type. Level three: you can trace it to inspectable artifacts (ticket history, conversation samples). Level four: it’s hard to game.

If key gate metrics are stuck at level one, be cautious. If they’re at level three or four, you can ship with more confidence.

A ship/no‑ship rule that holds up: if the primary outcome looks good but guardrail signals are missing or contaminated, pause. If you can’t reliably measure transfers after a routing change because events aren’t logged consistently, you’re effectively blind. Fix the signal or shrink the blast radius until you can observe.

For a crisp reminder that premortems should be structured, not vibes, this framing is useful: [3]

Failure modes and tradeoffs: what breaks first after staffing, routing, macro, escalation, or SLA changes

If you lead support ops long enough, you start to see the same movie with different actors.

Act one: optimism. Act two: a small crack. Act three: you realize the crack was load‑bearing.

The decision pre mortem exists to keep you from watching act three on repeat.

Work doesn’t disappear—it moves (to escalations, reopens, backlog aging, or agent burnout)

Support work is like laundry: you can move it from the chair to the bed, but it’s still laundry.

Routing changes usually break first in misroutes and transfers; specialist queues either starve or flood, creating backlog aging in one pocket. Macro/automation changes tend to break first in tone and correctness; deflection can rise while sentiment drops, then reopens and escalations climb because customers feel brushed off. Staffing changes often break in off‑hours aging and burnout because the same headcount now handles higher‑complexity work. SLA and escalation rule changes often break in gaming behavior (fast acknowledgements, slow progress) and in upstream overload when specialist or engineering teams get flooded.

The point isn’t to catalog every failure. It’s to predict the first breakpoints so you can place guardrails where failure actually starts.

Second-order effects by change type: staffing vs routing vs macros vs escalation rules vs SLAs

Different changes have different second‑order effects, and that’s where “we did everything right” turns into “why are customers mad?”

Staffing changes often show up in the wait‑time distribution first: the oldest tickets get older, and more tickets cross uncomfortable thresholds. Routing changes often show up in bounce patterns first: a generalist becomes a triage layer, FRT improves, transfers spike, specialists get flooded with borderline cases, then inconsistent answers appear and reopens rise. Macro and automation changes often show up as a tradeoff between speed and human trust: handle time drops, but customers feel unheard, so they escalate or reopen. Escalation rule changes often break trust if thresholds are too low (escalations become a second support queue) or too high (frontline loses confidence and customers suffer). SLA changes often create a world where every ticket gets a fast reply and a slow resolution—metric‑centric, not customer‑centric.

Decision gates that prevent ‘irreversible drift’

Irreversible drift is when you ship a change, then slowly adapt processes around it until nobody remembers what “before” felt like. At that point, rollback isn’t really possible because the organization already contorted itself.

Gates prevent drift because they force a decision while the change is still a pilot, not a new identity.

A pre‑commit guardrail set should include three parts: a threshold, a time window, and an owner. Example: “Transfer rate must not increase more than X percent for two consecutive business days in the pilot segment. Owner is support ops. Reviewed daily for the first week.”

One practical rule: make one guardrail about quality and one about workload movement. If you only gate on speed, you’ll optimize for fast wrong answers.

Rollback criteria that are specific enough to be used under pressure

Rollback criteria fail when they’re vague. “If CSAT drops” is vague, and it’s late.

A better pattern uses a leading indicator, a lagging indicator, and a time window. For an SLA change: “Roll back the new first response SLA for chat if reopen rate increases by more than X percent for three consecutive days (leading), and if CSAT for chat in the pilot segment drops below Y for two consecutive weeks (lagging). Review daily for the first week, then twice weekly.”

This is where teams get burned socially: moving the goalposts after launch. The change ships, a guardrail trips, and someone says, “That metric was always noisy,” or “Let’s give it another week,” without acknowledging the pre‑commitment. The fix isn’t technical. It’s governance: write gates in advance, name the owner, and put the review on the calendar before you ship.

Make it repeatable: a one-page pre‑mortem output you can reuse for every support-ops change

Pre mortems compound when you treat them like a reusable ritual, not a one‑off event.

The goal isn’t to predict everything. It’s to build reflexes: name assumptions, demand decision‑grade signals, and pre‑commit to gates and rollback.

The pre‑mortem deliverable: assumptions, missing signals, gates, monitoring, rollback, owner

Keep the deliverable brutally short: decision statement and constraints; blast radius; top assumptions; missing signals; decision gate to expand; guardrails with threshold/time window/owner; monitoring cadence (daily in week one, then weekly); rollback criteria using leading + lagging indicators; rollback owner; and the predicted first breakpoints (what breaks first, where work moves).

One filled‑field example, because this is where people stay vague.

Rollback criterion for an SLA change: “If backlog aging over 72 hours in the affected queue increases by more than X percent for five consecutive days, and QA correctness defects rise above Y for two weeks, revert to the prior SLA promise and re‑evaluate staffing assumptions.”

That’s specific enough to use when the pressure hits.

How to store and revisit decisions (so you learn across changes)

Store pre mortems where your team actually looks, not where documents go to retire.

Keep a lightweight log with three lines per decision: what you changed, what you predicted would break first, and what actually happened. That comparison is where the learning lives. Not a full postmortem—just a pattern check.

Over time, you build your own library of support ops failure modes. That becomes a strategic asset because future changes get faster and safer.

A lightweight cadence: before major changes, and a mini-version for weekly tweaks

Not every tweak deserves 30 minutes. For faster changes, use a mini version that still forces commitment: what’s changing (for whom, starting when), what breaks first and what signal you’ll check tomorrow, one guardrail that would make you pause, and who owns the check‑in in 48 hours.

If you can’t answer those, don’t ship. Or ship only to a tiny blast radius.

End with a Monday plan, because rituals only stick when they hit the calendar.

On Monday, pick one upcoming change—a routing tweak, a macro rollout, or an SLA adjustment—and run a 30‑minute pre mortem with four people.

Write a reversible decision statement with a time window. Name at least two leading indicators that would reveal silent failure. Pre‑commit one decision gate and one rollback trigger with an owner.

If you leave with a one‑page doc that includes at least one missing signal you discovered—and you either add that signal or shrink the blast radius until you can observe—you’re doing what grown‑up support operations looks like.

Sources

  1. frameworklist.com — frameworklist.com
  2. medium.com — medium.com
  3. strategicdecisionsolutions.com — strategicdecisionsolutions.com