[{"data":1,"prerenderedAt":47},["ShallowReactive",2],{"/en/blog/the-decision-pre-mortem-find-the-missing-signal-before-you-commit":3,"/en/blog/the-decision-pre-mortem-find-the-missing-signal-before-you-commit-surround":38},{"id":4,"locale":5,"translationGroupId":6,"availableLocales":7,"alternates":8,"_path":9,"path":9,"title":10,"description":11,"date":12,"modified":12,"meta":13,"seo":23,"topicSlug":28,"tags":29,"body":31,"_raw":36},"9e2078da-51ef-4ea8-ae5a-0637a76f80e3","en","f3b6802e-9057-423c-9ccf-cf510acbfd93",[5],{"en":9},"/en/blog/the-decision-pre-mortem-find-the-missing-signal-before-you-commit","The Decision Pre Mortem: Find the Missing Signal Before You Commit","A decision pre mortem for support operations helps you catch missing signals before a workflow change ships. You’ll get a reusable one-page artifact, branch-level metric validation, and guardrails that prevent “green dashboard, red reality” surprises.","2026-07-26T09:18:03.148Z",{"date":12,"badge":14,"authors":17},{"label":15,"color":16},"New","primary",[18],{"name":19,"description":20,"avatar":21},"Lucía Ferrer","Calypso AI · Clear, expert-led guides for operators and buyers",{"src":22},"https://api.dicebear.com/9.x/personas/svg?seed=calypso_expert_guide_v1&backgroundColor=b6e3f4,c0aede,d1d4f9,ffd5dc,ffdfbf",{"title":24,"description":25,"ogDescription":25,"twitterDescription":25,"canonicalPath":9,"robots":26,"schemaType":27},"The Decision Pre Mortem: Find the Missing Signal Before You","A decision pre mortem for support operations helps you catch missing signals before a workflow change ships. You’ll get a reusable one page artifact, branch","index,follow","BlogPosting","decision_systems_researcher",[30],"the-decision-pre-mortem-find-the-missing-signal-before-you-commit",{"toc":32,"children":34,"html":35},{"links":33},[],[],"\u003Ch2>Run the “meeting-before-the-meeting”: a 30-minute pre mortem for any support ops change\u003C/h2>\n\u003Cp>Most support operations changes don’t fail loudly. They fail politely.\u003C/p>\n\u003Cp>The dashboard stays green. Leadership moves on. Frontline folks build a workaround, stop mentioning it, and quietly absorb the tax. Then a month later you realize you traded one visible pain for three invisible ones: higher recontact, more escalations, quieter churn, and a team that now tiptoes around the new workflow like it’s a pothole nobody is allowed to acknowledge.\u003C/p>\n\u003Cp>That’s the job of a \u003Cstrong>decision pre mortem for support operations\u003C/strong>: a short, structured conversation you run \u003Cem>before\u003C/em> you commit, where the team assumes the change has already failed and asks: “What happened, what did we miss, and what signal would have warned us early?”\u003C/p>\n\u003Cp>This isn’t pessimism. It’s operational honesty. Support is a business of fragments: one angry email, one confusing bot exchange, one region melting down while the global average looks fantastic.\u003C/p>\n\u003Ch3>What “missing signal” looks like in support (clean dashboards, messy reality)\u003C/h3>\n\u003Cp>A missing signal is anything that would change the decision if you could see it clearly, but you can’t—because it’s unmeasured, aggregated away, biased by tagging, delayed until customers are already upset, or trapped in agent notes nobody reads.\u003C/p>\n\u003Cp>Concrete example: you enable \u003Cstrong>auto close after 72 hours\u003C/strong> to reduce backlog. Aggregate metrics show “backlog down 18%” and “time to first response steady.” Everyone high-fives.\u003C/p>\n\u003Cp>The missing signal is that \u003Cstrong>reopen rate and repeat contact\u003C/strong> jumps for one queue—often complex billing—because customers reply on day four with the detail you needed. You didn’t reduce work. You deferred it, duplicated it, and made it angrier.\u003C/p>\n\u003Cp>Another: you tighten a macro to be more direct (“per policy, we can’t…”) and you see fewer follow-ups. AHT drops. But escalation language spikes because the message reads like a parking ticket taped to someone’s windshield. The missing signal wasn’t efficiency. It was tone impact.\u003C/p>\n\u003Cp>A useful (slightly annoying) question whenever an aggregate win gets celebrated: “Which branch could be failing while this number still looks good?” If nobody can answer, you just found your next pre mortem.\u003C/p>\n\u003Ch3>The trigger list: changes that deserve a pre mortem (automation, staffing, routing, SLAs, policy)\u003C/h3>\n\u003Cp>Use a pre mortem when the change is hard to reverse socially, even if it’s technically reversible.\u003C/p>\n\u003Cp>It earns its keep for:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Automation:\u003C/strong> bots, auto close, deflection, macros, reply suggestions.\u003C/li>\n\u003Cli>\u003Cstrong>Staffing/coverage:\u003C/strong> weekends, on-call, tier splits, new schedules.\u003C/li>\n\u003Cli>\u003Cstrong>Routing:\u003C/strong> new queues, skills, regions, priority rules.\u003C/li>\n\u003Cli>\u003Cstrong>SLAs/KPIs:\u003C/strong> new targets, breach logic, escalation timing.\u003C/li>\n\u003Cli>\u003Cstrong>Policy:\u003C/strong> refunds, identity verification, what “resolved” means.\u003C/li>\n\u003C/ul>\n\u003Cp>Decision framing that works: if the change affects (a) what customers experience first, or (b) what agents do most often, treat it as pre mortem-worthy. Those are the levers that create fast, silent failure.\u003C/p>\n\u003Ch3>The pre mortem question set you can copy into the calendar invite\u003C/h3>\n\u003Cp>Keep it small and slightly inconvenient. Thirty minutes is usually enough to expose weak confidence.\u003C/p>\n\u003Cp>Use three prompts:\u003C/p>\n\u003Cul>\n\u003Cli>“It’s 30 days after launch and this change is considered a mistake. What’s the story of how it went wrong?”\u003C/li>\n\u003Cli>“What metric is going to tell us everything is fine while reality is getting worse?”\u003C/li>\n\u003Cli>“What is the one missing signal we need before we approve this?”\u003C/li>\n\u003C/ul>\n\u003Cp>Decision rule for support ops: if the group can’t name at least one plausible failure story \u003Cem>and\u003C/em> one measurable early warning signal, you’re not ready to ship. You’re ready to hope.\u003C/p>\n\u003Cp>One move that changes the tone fast: ask the Decider (or whoever is most excited) to say this out loud before discussion starts: “If we see X, I will pause.” Vague optimism hates deadlines.\u003C/p>\n\u003Ch2>Ship the one-page pre mortem artifact: roles, timebox, and the minimum inputs\u003C/h2>\n\u003Cp>The most common way teams ruin a support operations pre mortem is turning it into either a leadership vibe check or an analyst report-out.\u003C/p>\n\u003Cp>Both fail for the same reason: they concentrate authority in one place and leave blind spots untouched.\u003C/p>\n\u003Cp>The fix is a one-page artifact produced inside a timebox, with roles that create the right friction. Your goal is speed plus accountability—not paperwork.\u003C/p>\n\u003Ch3>Roles that reduce bias (Decider, Facilitator, Data Buddy, Frontline Rep, Skeptic)\u003C/h3>\n\u003Cp>Five roles. In a small team, one person can wear two hats, but don’t drop the Skeptic or the Frontline Rep. Those are your smoke alarms.\u003C/p>\n\u003Cp>\u003Cstrong>Decider\u003C/strong> owns the call and explicitly states what would change their mind. If they can’t say that, the meeting is theater.\u003C/p>\n\u003Cp>\u003Cstrong>Facilitator\u003C/strong> protects the timebox, pulls quiet voices in, and blocks solutioneering. This is the person who says, “We’re not fixing it yet. We’re naming how it fails.”\u003C/p>\n\u003Cp>\u003Cstrong>Data Buddy\u003C/strong> brings the minimum decision-grade numbers and can explain what’s noisy, what’s lagging, and what’s not tracked. Not “the person who built the dashboard.” The person who can tell you what the dashboard \u003Cem>can’t\u003C/em>.\u003C/p>\n\u003Cp>\u003Cstrong>Frontline Rep\u003C/strong> brings reality from the queue. Pick someone who still touches tickets or reviews quality weekly—not someone who “used to.”\u003C/p>\n\u003Cp>\u003Cstrong>Skeptic\u003C/strong> plays hostile investor for the hour. If you don’t assign this, the room will politely avoid conflict and then complain later in private, which is the most expensive form of feedback.\u003C/p>\n\u003Cp>This is where teams get burned: running the pre mortem only with leadership because it feels “strategic.” You get alignment and confidence—and almost no contact with edge cases. Better mix: one leader, one operator, one frontline rep, one skeptic voice.\u003C/p>\n\u003Ch3>A 45–60 minute agenda with hard stops (and what to do if you run out of time)\u003C/h3>\n\u003Cp>Start with a clear promise: you’re here to decide whether the change is ready to ship—not whether it’s a good idea in the abstract.\u003C/p>\n\u003Cp>A flow that holds up in real ops:\u003C/p>\n\u003Cul>\n\u003Cli>Set the decision statement and scope (what’s changing; which queues/channels/customers).\u003C/li>\n\u003Cli>Name assumptions (each person offers one “this must be true”).\u003C/li>\n\u003Cli>Write failure narratives silently, then share (silent writing prevents the loudest voice from setting the plot).\u003C/li>\n\u003Cli>Translate narratives into missing signals (earliest observable warning for each story).\u003C/li>\n\u003Cli>Add guardrails and owners (pause, rollback, escalation).\u003C/li>\n\u003Cli>Make the call (approve, approve with conditions, or stop pending signal integrity).\u003C/li>\n\u003C/ul>\n\u003Cp>If time gets tight, don’t “extend” and call it collaboration. Pick one: schedule a second session, or the Decider makes a provisional call with explicit unknowns recorded. Anything else becomes a slow-motion yes.\u003C/p>\n\u003Ch3>The one-page template: decision, assumptions, failure narratives, signals, guardrails, owners\u003C/h3>\n\u003Cp>The artifact should be short enough that someone can read it in two minutes and understand why you did what you did.\u003C/p>\n\u003Cp>Use these sections:\u003C/p>\n\u003Cul>\n\u003Cli>Decision and scope\u003C/li>\n\u003Cli>What success looks like (plain language)\u003C/li>\n\u003Cli>Assumptions we are betting on\u003C/li>\n\u003Cli>Failure narratives (3–5 “how it goes wrong” stories)\u003C/li>\n\u003Cli>Missing signals and how we will detect them\u003C/li>\n\u003Cli>Guardrails and kill criteria (pause/rollback triggers)\u003C/li>\n\u003Cli>Owners and check cadence\u003C/li>\n\u003Cli>Open unknowns (labeled explicitly)\u003C/li>\n\u003C/ul>\n\u003Cp>Two filled examples:\u003C/p>\n\u003Cp>Assumption: “Auto close at 72 hours will reduce backlog without hurting customer outcomes.”\u003C/p>\n\u003Cp>Signal: “Reopen rate and repeat contact within 7 days by queue, with billing tracked separately.”\u003C/p>\n\u003Cp>Owner: “Support ops lead reviews daily week one and posts a one-paragraph summary.”\u003C/p>\n\u003Cp>Assumption: “New routing that sends ‘login issues’ to Tier 1 will reduce time to first response.”\u003C/p>\n\u003Cp>Signal: “Escalation rate from Tier 1 to Tier 2 by region/language, plus top misroute reasons from sampling.”\u003C/p>\n\u003Cp>Owner: “Routing owner + frontline rep review 20 conversations daily across branches for three days.”\u003C/p>\n\u003Ch3>Pre-work rules: what data is allowed, what gets labeled as “unknown,” and why\u003C/h3>\n\u003Cp>Pre-work should be strict. Otherwise you show up with a 40-slide deck that nobody reads and everyone pretends they did.\u003C/p>\n\u003Cp>Rules that keep it decision-grade:\u003C/p>\n\u003Cul>\n\u003Cli>Bring only data that can influence a decision in the next two weeks. If it won’t change scope, guardrails, or monitoring, it’s background noise.\u003C/li>\n\u003Cli>Any metric that isn’t trustworthy at the branch level gets labeled \u003Cstrong>unknown\u003C/strong>.\u003C/li>\n\u003Cli>Separate lagging indicators from leading ones. If proof of harm arrives after churn, you need a different signal.\u003C/li>\n\u003C/ul>\n\u003Cp>Common failure: letting “unknown” become “we’ll look later.” Later is where accountability goes to nap. Instead, connect unknowns to a condition: “We ship only to one queue until tagging health is proven,” or “We pause until we can break out the metric by region.”\u003C/p>\n\u003Cp>If you want broader context on the original pre mortem framing (and why it works), this is a useful skim: \u003Ca href=\"#ref-1\" title=\"expectedvalue.co.uk — expectedvalue.co.uk\">[1]\u003C/a>\u003C/p>\n\u003Ch2>Truth-test branch signals before you trust the aggregate (queues, regions, teams)\u003C/h2>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Control\u003C/th>\n\u003Cth>Where it lives\u003C/th>\n\u003Cth>What to set\u003C/th>\n\u003Cth>What breaks if it’s wrong\u003C/th>\n\u003C/tr>\n\u003C/thead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>Set: Guardrail: Pause decision if signal integrity issues cannot be resolved within a defined timeframe\u003C/td>\n\u003Ctd>Decision-making process, risk management plan\u003C/td>\n\u003Ctd>A maximum time limit — e.g., 48 hours to address signal issues before halting the decision process\u003C/td>\n\u003Ctd>Rushing into poorly informed decisions, compounding errors, loss of credibility\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Set: Define clear branch-level pass/fail tests — e.g., per queue, region, team\u003C/td>\n\u003Ctd>Pre-mortem checklist, decision brief\u003C/td>\n\u003Ctd>At least 5 specific, measurable tests for each branch, with clear thresholds\u003C/td>\n\u003Ctd>Misleading aggregate metrics, incorrect resource allocation, biased decision-making\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Set: Identify potential misleading aggregates and their underlying branches\u003C/td>\n\u003Ctd>Pre-mortem analysis, data review sessions\u003C/td>\n\u003Ctd>Specific examples of aggregates that hide branch-level issues — e.g., &#39;overall CSAT is high, but Region X is failing&#39;\u003C/td>\n\u003Ctd>Failure to address critical issues hidden by averages, wasted effort on non-problems\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Set: Establish a &#39;kill criteria&#39; for signal integrity failures\u003C/td>\n\u003Ctd>Decision framework, pre-mortem protocol\u003C/td>\n\u003Ctd>A clear rule: &#39;If X% of branch signals are invalid / missing, decision is paused / reverted&#39;\u003C/td>\n\u003Ctd>Committing to decisions based on unreliable data, loss of trust in metrics\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Validate data collection methods at the branch level\u003C/td>\n\u003Ctd>Data governance documentation, audit logs\u003C/td>\n\u003Ctd>Regular checks on data input processes, sample sizes, and reporting consistency for each branch\u003C/td>\n\u003Ctd>Inaccurate data, biased samples, inability to compare branches fairly\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Set: Review for missing coverage in critical branches — e.g., new markets, niche products\u003C/td>\n\u003Ctd>Metric dashboards, reporting requirements\u003C/td>\n\u003Ctd>Mandatory reporting for all critical branches, even if data is sparse. flag &#39;no data&#39; as a risk\u003C/td>\n\u003Ctd>Blind spots in decision-making, neglecting high-risk or high-potential areas\u003C/td>\n\u003C/tr>\n\u003C/tbody>\u003C/table>\n\u003Cp>Aggregates are comforting. They also lie in very specific, repeatable ways.\u003C/p>\n\u003Cp>When someone says, “Support is fine, our SLA is fine,” they often mean, “The biggest queue looks okay.” Meanwhile a smaller queue is melting down, the language mix shifted, or one region is getting systematically misrouted.\u003C/p>\n\u003Cp>This is where \u003Cstrong>branch level metric validation\u003C/strong> stops being a nerdy analytics habit and becomes a risk control.\u003C/p>\n\u003Ch3>Branch-level truth tests: denominator sanity, tagging health, mix-shift detection\u003C/h3>\n\u003Cp>These truth tests are decision-grade because they have real pass/fail meaning:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Denominator sanity:\u003C/strong> do the counts match reality? Sudden drops in contact volume by branch without an operational explanation often mean tracking broke.\u003C/li>\n\u003Cli>\u003Cstrong>Tagging health:\u003C/strong> are tags/dispositions used consistently? If one team uses “customer education” as a bucket for everything, your slices are fiction.\u003C/li>\n\u003Cli>\u003Cstrong>Mix shift detection:\u003C/strong> did the customer mix change? A workflow update can look like a win if the queue got easier that week.\u003C/li>\n\u003Cli>\u003Cstrong>Branch parity:\u003C/strong> do key ratios look roughly similar across branches unless you have a known reason? Stable overall escalation rate can hide a single-region spike.\u003C/li>\n\u003Cli>\u003Cstrong>Latency symmetry:\u003C/strong> did one channel improve while another degraded? Email improving while chat worsens is a classic side effect.\u003C/li>\n\u003Cli>\u003Cstrong>Reopen vs recontact split:\u003C/strong> don’t trust one “resolution rate.” Separate “resolved and stayed resolved” from “resolved and came back.”\u003C/li>\n\u003C/ul>\n\u003Cp>Don’t wait for perfect dashboards. A scrappy branch check beats an elegant global average—as long as you’re explicit about what’s noisy and what’s missing.\u003C/p>\n\u003Ch3>Coverage checks: what you are not measuring (hours, languages, channels, severity)\u003C/h3>\n\u003Cp>Missing signals often come from missing coverage. You measure what’s convenient, not what’s risky.\u003C/p>\n\u003Cp>Two slices that uncover surprises fast:\u003C/p>\n\u003Cp>\u003Cstrong>Region and language.\u003C/strong> A routing change that helps English tickets can quietly harm non-English tickets because skill rules get rigid, translation workflows differ, and coverage isn’t symmetrical.\u003C/p>\n\u003Cp>\u003Cstrong>Channel and severity.\u003C/strong> A deflection change can reduce chat volume while pushing high-severity users into email, where they wait longer and escalate harder.\u003C/p>\n\u003Cp>Rule that saves teams: if you can’t break out the metric by the branch that will carry the risk, you don’t have a metric. You have a mood.\u003C/p>\n\u003Cp>And yes—this is where teams get burned—assuming “no data” means “no problem.” In support ops, “no data” usually means “the branch isn’t instrumented,” which is a risk by itself.\u003C/p>\n\u003Ch3>Conversation sampling that doesn’t lie (how to avoid survivor bias and “resolved-only” reviews)\u003C/h3>\n\u003Cp>Sampling is where teams accidentally gaslight themselves.\u003C/p>\n\u003Cp>If you only review “resolved” conversations, you select for success. If you only review escalations, you select for drama. You need the boring middle.\u003C/p>\n\u003Cp>A lightweight cadence most teams can sustain:\u003C/p>\n\u003Cul>\n\u003Cli>For each affected branch, review \u003Cstrong>10 conversations per day\u003C/strong> for the first three days, then \u003Cstrong>10 per branch\u003C/strong> twice in the next week.\u003C/li>\n\u003Cli>Make sure at least \u003Cstrong>two\u003C/strong> per branch are reopened, escalated, or transferred.\u003C/li>\n\u003Cli>Include at least \u003Cstrong>one\u003C/strong> auto-closed or bot-deflected conversation if automation is involved.\u003C/li>\n\u003C/ul>\n\u003Cp>If the team can’t sustain that cadence, that’s not a moral failing. It’s a capacity signal. Shrink rollout scope.\u003C/p>\n\u003Ch3>The “green dashboard, red reality” patterns: when aggregates hide concentrated failures\u003C/h3>\n\u003Cp>Classic scenario: you split a queue into “Account access” and “Billing” to improve routing. Aggregate first response improves by 12%. Great.\u003C/p>\n\u003Cp>But the billing sub-queue now has a backlog slope that climbs every day because routing rules undercount severity, and the few billing specialists can’t keep up. The missing signal was branch backlog slope by queue—not overall first response time.\u003C/p>\n\u003Cp>Stop rules matter here. If you can’t trust the branch numbers, you pause.\u003C/p>\n\u003Cp>A practical stop rule: if \u003Cstrong>two or more\u003C/strong> core branch truth tests fail and can’t be resolved within \u003Cstrong>one working day\u003C/strong>, don’t approve broad rollout. Narrow scope or delay.\u003C/p>\n\u003Cp>This logic maps cleanly to kill criteria as a discipline: \u003Ca href=\"#ref-2\" title=\"howtothink.ai — howtothink.ai\">[2]\u003C/a>\u003C/p>\n\u003Ch2>Decide what to automate vs keep human: where bots and macros amplify blind spots\u003C/h2>\n\u003Cp>Automation in support is like adding a faster conveyor belt to a factory. If the sorting is wrong, you just deliver mistakes at a more impressive pace.\u003C/p>\n\u003Cp>A decision pre mortem is where you decide not only “can we automate?” but “what kind of failure does automation create, and will we notice it in time?”\u003C/p>\n\u003Ch3>A decision rule: automate when the failure is cheap and observable; keep human when it’s expensive and silent\u003C/h3>\n\u003Cp>Here’s the rule operators can actually use:\u003C/p>\n\u003Cp>Automate when the failure is \u003Cstrong>cheap and observable\u003C/strong>.\u003C/p>\n\u003Cp>Keep a human when the failure is \u003Cstrong>expensive and silent\u003C/strong>.\u003C/p>\n\u003Cp>Cheap and observable looks like a wrong macro suggestion an agent can ignore, with a clear feedback loop.\u003C/p>\n\u003Cp>Expensive and silent looks like a bot confidently applying the wrong eligibility rule, “resolving” the case, and you only find out later when refunds spike or cancellations rise.\u003C/p>\n\u003Cp>Tradeoff to say out loud: automation often improves speed and consistency, but it can also hide error behind a clean resolution status. Humans surface nuance faster, but they introduce variance and “hero workflows.” You’re choosing which failure mode you prefer—and which you can detect.\u003C/p>\n\u003Ch3>Automation failure modes in support: exception collapse, tone drift, misrouting, and false resolution\u003C/h3>\n\u003Cp>Automation misses aren’t random. They cluster.\u003C/p>\n\u003Cp>\u003Cstrong>Exception collapse:\u003C/strong> edge cases get forced into the closest category, and the system stops admitting uncertainty. Early warnings: rising “other” tags, rising transfers, longer internal notes.\u003C/p>\n\u003Cp>\u003Cstrong>Tone drift:\u003C/strong> macros and bots become technically correct and emotionally wrong. Early warnings: sentiment and complaint language rising even if handle time drops.\u003C/p>\n\u003Cp>\u003Cstrong>Misrouting:\u003C/strong> routing rules optimize for speed but ignore skill nuance. Early warnings: escalation/transfer rate up in one region, team, or language.\u003C/p>\n\u003Cp>\u003Cstrong>False resolution:\u003C/strong> auto close or deflection reduces visible backlog while increasing reopens and repeat contacts. Early warnings: “resolved” rising while recontact rises.\u003C/p>\n\u003Cp>This is why “AHT down” isn’t a victory song. It’s a clue. Sometimes a good clue, sometimes a trap.\u003C/p>\n\u003Ch3>Human judgment failure modes: inconsistency, over-escalation, and “local hero” workflows\u003C/h3>\n\u003Cp>Humans aren’t automatically safer. They fail differently.\u003C/p>\n\u003Cp>\u003Cstrong>Inconsistency\u003C/strong> shows up as policy drift: two customers get two answers. Early warning: higher variance in outcomes across agents and teams.\u003C/p>\n\u003Cp>\u003Cstrong>Over-escalation\u003C/strong> happens when agents are punished for being wrong, so they escalate to protect themselves. Early warning: escalations rise while quality doesn’t.\u003C/p>\n\u003Cp>\u003Cstrong>Local hero workflows\u003C/strong> are the sneakiest because they look like excellence. One experienced agent invents a workaround that saves the day, and suddenly your system depends on a person, not a process. Early warning: a subset of cases only resolves when one or two people touch them.\u003C/p>\n\u003Cp>In your pre mortem, always ask: “Where does this require hero behavior to look good?” If the answer is “nowhere,” you might be rounding up.\u003C/p>\n\u003Ch3>How to run an exceptions-first review (and what exceptions tell you about missing signals)\u003C/h3>\n\u003Cp>Instead of reviewing average cases, review the exceptions first. Exceptions are where missing signals live.\u003C/p>\n\u003Cp>Two examples that show up constantly:\u003C/p>\n\u003Cp>A \u003Cstrong>macro change\u003C/strong> adds firmer language to reduce back-and-forth. AHT improves because agents send fewer follow-ups. But escalations and complaint language rise because the macro sounds like a parking ticket. The missing signal wasn’t AHT. It was escalation reason codes—or a simple qualitative flag from daily sampling: “customer reacted negatively to tone.”\u003C/p>\n\u003Cp>An \u003Cstrong>auto routing change\u003C/strong> sends “login issues” to Tier 1. First response improves. But Tier 1 transfer rate spikes for one product line because “login issue” is often a permissions issue. The missing signal was branch transfer rate by product area, plus a quick scan of the top misroute reasons.\u003C/p>\n\u003Cp>You don’t need new tools to capture this. You need an exceptions log for week one. Keep it simple: branch, what failed, customer impact, how you detected it, owner, what you changed.\u003C/p>\n\u003Cp>Some teams use AI as an adversarial partner because it has no career incentive to agree with the boss. Whether you use AI or not, the adversarial posture is the point: \u003Ca href=\"#ref-3\" title=\"pre-mortem.ai — pre-mortem.ai\">[3]\u003C/a>\u003C/p>\n\u003Ch2>Pre-commit checks: counterfactuals, guardrails, and what to monitor in the first week\u003C/h2>\n\u003Cp>The difference between a brave decision and a reckless one is rarely intent. It’s whether you planned the exit.\u003C/p>\n\u003Cp>Support ops changes feel reversible because they’re “just workflows.” In practice, reversals are painful: agents retrain, customers adapt, reporting breaks, and leadership loses confidence. Pre-commit checks make reversibility real, not theoretical.\u003C/p>\n\u003Ch3>Counterfactuals: ‘If we do nothing, what happens?’ and ‘If we reverse it, what breaks?’\u003C/h3>\n\u003Cp>Counterfactuals expose missing signal risk fast.\u003C/p>\n\u003Cp>Ask before approval:\u003C/p>\n\u003Cul>\n\u003Cli>If we do nothing for 30 days, what happens to backlog, SLA risk, and customer pain?\u003C/li>\n\u003Cli>If we reverse this change after one week, what breaks—training, routing rules, customer expectations, reporting continuity?\u003C/li>\n\u003C/ul>\n\u003Cp>If you can’t answer the reversal question without hand-waving, you’re about to ship a one-way door disguised as a toggle.\u003C/p>\n\u003Ch3>Guardrails that matter: kill switches, thresholds, and staged rollout criteria\u003C/h3>\n\u003Cp>Guardrails aren’t a list of metrics. Guardrails are action triggers.\u003C/p>\n\u003Cp>A pre-commit set that actually holds up in week one:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Kill switch plan:\u003C/strong> who can pause, where it’s communicated, what “paused” means for agents today.\u003C/li>\n\u003Cli>\u003Cstrong>Scope control:\u003C/strong> start with one queue or one region until signal integrity is proven.\u003C/li>\n\u003Cli>\u003Cstrong>Threshold guardrails:\u003C/strong> tied to action, not debate.\u003C/li>\n\u003Cli>\u003Cstrong>Escalation path:\u003C/strong> who to pull in if a branch starts failing.\u003C/li>\n\u003Cli>\u003Cstrong>Customer comms stance:\u003C/strong> what you tell customers if the change causes confusion.\u003C/li>\n\u003C/ul>\n\u003Cp>Example threshold (tune numbers to your baseline, keep the structure):\u003C/p>\n\u003Cp>If \u003Cstrong>reopen rate\u003C/strong> increases by a team-defined threshold in the affected queue for \u003Cstrong>two consecutive days\u003C/strong>, you \u003Cstrong>pause rollout\u003C/strong>, review \u003Cstrong>20 sampled conversations\u003C/strong> from that branch, and decide within \u003Cstrong>one business day\u003C/strong> whether to rollback or adjust.\u003C/p>\n\u003Cp>Guardrails for support automation should be measurable, owned, and timebound. Otherwise they’re just comforting words you say while shipping anyway.\u003C/p>\n\u003Ch3>First 24 hours vs first week: leading indicators (quality, recontact, backlog slope, escalations)\u003C/h3>\n\u003Cp>Don’t treat week one as a single blob. The first 24 hours is about safety. The first week is about learning.\u003C/p>\n\u003Cp>In the \u003Cstrong>first 24 hours\u003C/strong>, watch fast-moving indicators:\u003C/p>\n\u003Cul>\n\u003Cli>Backlog slope by branch\u003C/li>\n\u003Cli>Escalations and transfers by branch\u003C/li>\n\u003Cli>Top exception types from the exceptions log\u003C/li>\n\u003Cli>Repeating agent feedback (patterns, not one-offs)\u003C/li>\n\u003C/ul>\n\u003Cp>In the \u003Cstrong>first week\u003C/strong>, add indicators that reveal customer impact:\u003C/p>\n\u003Cul>\n\u003Cli>Recontact within 7 days for the same issue category\u003C/li>\n\u003Cli>Reopen rate split by queue and severity\u003C/li>\n\u003Cli>Quality review outcomes from sampled conversations\u003C/li>\n\u003Cli>Complaint language frequency in notes/tags (even if informal)\u003C/li>\n\u003C/ul>\n\u003Cp>Notice what isn’t your north star: average handle time.\u003C/p>\n\u003Cp>AHT is useful, but it’s a classic “we got faster by getting worse” metric unless you pair it with recontact and quality.\u003C/p>\n\u003Cp>Practical decision rule: pick two “must not degrade” signals and treat them as your stop line. Teams that try to monitor ten things usually monitor none.\u003C/p>\n\u003Ch3>How to assign owners and cadence so monitoring actually happens\u003C/h3>\n\u003Cp>Monitoring fails for one boring reason: nobody owns the calendar.\u003C/p>\n\u003Cp>Assign owners the way you assign incident roles.\u003C/p>\n\u003Cp>A cadence that’s lightweight but real:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Daily for the first three days:\u003C/strong> Data Buddy posts a short branch snapshot; Frontline Rep posts two patterns from sampling; Decider confirms continue or pause.\u003C/li>\n\u003Cli>\u003Cstrong>Twice in week one:\u003C/strong> Facilitator updates the one-page artifact with what was learned, including new unknowns.\u003C/li>\n\u003Cli>\u003Cstrong>End of week one:\u003C/strong> 20-minute review of guardrails, exceptions, and whether to expand scope.\u003C/li>\n\u003C/ul>\n\u003Cp>Document in the same one-page artifact with a dated addendum. If it lives somewhere else, it will be forgotten. If it’s forgotten, it might as well not exist.\u003C/p>\n\u003Cp>Light humor, because we all need one: rolling out a support workflow change without guardrails is like “testing in production,” except your customers are the test suite and they don’t come with helpful error messages.\u003C/p>\n\u003Ch2>After you commit: turn every pre mortem into a decision library (so you stop relearning the same lesson)\u003C/h2>\n\u003Cp>A pre mortem compounds when you can retrieve it. Otherwise it’s just a one-time ritual that makes everyone feel mature.\u003C/p>\n\u003Cp>The goal is a small decision library that makes the next approval faster because you can say, “We’ve seen this movie before, and we know which scene goes wrong.” Over time, pre mortems also work as calibration tools because you can compare what you predicted with what actually happened: \u003Ca href=\"#ref-4\" title=\"howtothink.ai — howtothink.ai\">[4]\u003C/a>\u003C/p>\n\u003Ch3>The 10-minute post-decision addendum: what was true, what was missing, what changed\u003C/h3>\n\u003Cp>Keep a standing 10-minute addendum one week after launch.\u003C/p>\n\u003Cp>Do three things:\u003C/p>\n\u003Cul>\n\u003Cli>Mark which assumptions were true, false, or still unknown.\u003C/li>\n\u003Cli>Write the missing signal you discovered, especially if it surprised you.\u003C/li>\n\u003Cli>Update guardrails for next time based on what actually moved.\u003C/li>\n\u003C/ul>\n\u003Cp>This is where “we should remember that” turns into “we will not forget that.”\u003C/p>\n\u003Ch3>How to store and reuse pre mortems (tags: queue, change type, risk pattern, signals)\u003C/h3>\n\u003Cp>Don’t overthink tooling. A shared folder and consistent naming usually beat an ambitious system nobody maintains.\u003C/p>\n\u003Cp>Use a few tags at the top of the artifact: queues affected; change type (automation, routing, staffing, SLA, policy); risk pattern (misrouting, tone drift, false resolution, mix shift); missing signals discovered; guardrails used; outcome after one week and one month.\u003C/p>\n\u003Cp>Common mistake: letting pre mortems become write-only documents. If nobody can find the last three, you’re not building a library—you’re building a junk drawer.\u003C/p>\n\u003Ch3>Coaching loop: how to onboard new leads/operators using past misses\u003C/h3>\n\u003Cp>Here is a copyable example library entry:\u003C/p>\n\u003Cp>Decision: “Enable auto close after 72 hours for low priority email queue.”\u003C/p>\n\u003Cp>Branches affected: “Email, low priority, billing adjacent.”\u003C/p>\n\u003Cp>Missing signal discovered: “Reopen rate by billing adjacent tag was the early warning.”\u003C/p>\n\u003Cp>Guardrail: “Pause if reopen rate rises above baseline threshold for two days.”\u003C/p>\n\u003Cp>Lesson: “Backlog went down, repeat contact went up, so we narrowed scope and changed the close message.”\u003C/p>\n\u003Cp>This prevents the next bad commitment when someone proposes auto close again and claims, “We did this before and it was fine.” You can respond, calmly and with receipts, “It was fine overall. It was not fine in that branch. Here is what we watch this time.”\u003C/p>\n\u003Cp>To make this real on Monday, don’t build a program. Do one thing.\u003C/p>\n\u003Cp>Schedule the next “meeting before the meeting” and paste the one-page pre mortem headings into the invite. Then hold yourself to three priorities: (1) name one misleading aggregate you are currently trusting, (2) define two branch slices you’ll validate before shipping, and (3) set one guardrail with a clear pause action.\u003C/p>\n\u003Cp>Your production bar is modest: one-page artifact, one owner per signal, and a first-week monitoring note that is written down where the team can find it.\u003C/p>\n\u003Ch2>Sources\u003C/h2>\n\u003Col>\n\u003Cli>\u003Ca href=\"https://expectedvalue.co.uk/blog/pre-mortem-decision-making\">expectedvalue.co.uk\u003C/a> — expectedvalue.co.uk\u003C/li>\n\u003Cli>\u003Ca href=\"https://www.howtothink.ai/learn/kill-criteria\">howtothink.ai\u003C/a> — howtothink.ai\u003C/li>\n\u003Cli>\u003Ca href=\"https://pre-mortem.ai/ai-pre-mortem\">pre-mortem.ai\u003C/a> — pre-mortem.ai\u003C/li>\n\u003Cli>\u003Ca href=\"https://www.howtothink.ai/learn/pre-mortem-as-a-calibration-tool\">howtothink.ai\u003C/a> — howtothink.ai\u003C/li>\n\u003C/ol>\n",{"body":37},"## Run the “meeting-before-the-meeting”: a 30-minute pre mortem for any support ops change\n\nMost support operations changes don’t fail loudly. They fail politely.\n\nThe dashboard stays green. Leadership moves on. Frontline folks build a workaround, stop mentioning it, and quietly absorb the tax. Then a month later you realize you traded one visible pain for three invisible ones: higher recontact, more escalations, quieter churn, and a team that now tiptoes around the new workflow like it’s a pothole nobody is allowed to acknowledge.\n\nThat’s the job of a **decision pre mortem for support operations**: a short, structured conversation you run *before* you commit, where the team assumes the change has already failed and asks: “What happened, what did we miss, and what signal would have warned us early?”\n\nThis isn’t pessimism. It’s operational honesty. Support is a business of fragments: one angry email, one confusing bot exchange, one region melting down while the global average looks fantastic.\n\n### What “missing signal” looks like in support (clean dashboards, messy reality)\n\nA missing signal is anything that would change the decision if you could see it clearly, but you can’t—because it’s unmeasured, aggregated away, biased by tagging, delayed until customers are already upset, or trapped in agent notes nobody reads.\n\nConcrete example: you enable **auto close after 72 hours** to reduce backlog. Aggregate metrics show “backlog down 18%” and “time to first response steady.” Everyone high-fives.\n\nThe missing signal is that **reopen rate and repeat contact** jumps for one queue—often complex billing—because customers reply on day four with the detail you needed. You didn’t reduce work. You deferred it, duplicated it, and made it angrier.\n\nAnother: you tighten a macro to be more direct (“per policy, we can’t…”) and you see fewer follow-ups. AHT drops. But escalation language spikes because the message reads like a parking ticket taped to someone’s windshield. The missing signal wasn’t efficiency. It was tone impact.\n\nA useful (slightly annoying) question whenever an aggregate win gets celebrated: “Which branch could be failing while this number still looks good?” If nobody can answer, you just found your next pre mortem.\n\n### The trigger list: changes that deserve a pre mortem (automation, staffing, routing, SLAs, policy)\n\nUse a pre mortem when the change is hard to reverse socially, even if it’s technically reversible.\n\nIt earns its keep for:\n\n- **Automation:** bots, auto close, deflection, macros, reply suggestions.\n- **Staffing/coverage:** weekends, on-call, tier splits, new schedules.\n- **Routing:** new queues, skills, regions, priority rules.\n- **SLAs/KPIs:** new targets, breach logic, escalation timing.\n- **Policy:** refunds, identity verification, what “resolved” means.\n\nDecision framing that works: if the change affects (a) what customers experience first, or (b) what agents do most often, treat it as pre mortem-worthy. Those are the levers that create fast, silent failure.\n\n### The pre mortem question set you can copy into the calendar invite\n\nKeep it small and slightly inconvenient. Thirty minutes is usually enough to expose weak confidence.\n\nUse three prompts:\n\n- “It’s 30 days after launch and this change is considered a mistake. What’s the story of how it went wrong?”\n- “What metric is going to tell us everything is fine while reality is getting worse?”\n- “What is the one missing signal we need before we approve this?”\n\nDecision rule for support ops: if the group can’t name at least one plausible failure story *and* one measurable early warning signal, you’re not ready to ship. You’re ready to hope.\n\nOne move that changes the tone fast: ask the Decider (or whoever is most excited) to say this out loud before discussion starts: “If we see X, I will pause.” Vague optimism hates deadlines.\n\n## Ship the one-page pre mortem artifact: roles, timebox, and the minimum inputs\n\nThe most common way teams ruin a support operations pre mortem is turning it into either a leadership vibe check or an analyst report-out.\n\nBoth fail for the same reason: they concentrate authority in one place and leave blind spots untouched.\n\nThe fix is a one-page artifact produced inside a timebox, with roles that create the right friction. Your goal is speed plus accountability—not paperwork.\n\n### Roles that reduce bias (Decider, Facilitator, Data Buddy, Frontline Rep, Skeptic)\n\nFive roles. In a small team, one person can wear two hats, but don’t drop the Skeptic or the Frontline Rep. Those are your smoke alarms.\n\n**Decider** owns the call and explicitly states what would change their mind. If they can’t say that, the meeting is theater.\n\n**Facilitator** protects the timebox, pulls quiet voices in, and blocks solutioneering. This is the person who says, “We’re not fixing it yet. We’re naming how it fails.”\n\n**Data Buddy** brings the minimum decision-grade numbers and can explain what’s noisy, what’s lagging, and what’s not tracked. Not “the person who built the dashboard.” The person who can tell you what the dashboard *can’t*.\n\n**Frontline Rep** brings reality from the queue. Pick someone who still touches tickets or reviews quality weekly—not someone who “used to.”\n\n**Skeptic** plays hostile investor for the hour. If you don’t assign this, the room will politely avoid conflict and then complain later in private, which is the most expensive form of feedback.\n\nThis is where teams get burned: running the pre mortem only with leadership because it feels “strategic.” You get alignment and confidence—and almost no contact with edge cases. Better mix: one leader, one operator, one frontline rep, one skeptic voice.\n\n### A 45–60 minute agenda with hard stops (and what to do if you run out of time)\n\nStart with a clear promise: you’re here to decide whether the change is ready to ship—not whether it’s a good idea in the abstract.\n\nA flow that holds up in real ops:\n\n- Set the decision statement and scope (what’s changing; which queues/channels/customers).\n- Name assumptions (each person offers one “this must be true”).\n- Write failure narratives silently, then share (silent writing prevents the loudest voice from setting the plot).\n- Translate narratives into missing signals (earliest observable warning for each story).\n- Add guardrails and owners (pause, rollback, escalation).\n- Make the call (approve, approve with conditions, or stop pending signal integrity).\n\nIf time gets tight, don’t “extend” and call it collaboration. Pick one: schedule a second session, or the Decider makes a provisional call with explicit unknowns recorded. Anything else becomes a slow-motion yes.\n\n### The one-page template: decision, assumptions, failure narratives, signals, guardrails, owners\n\nThe artifact should be short enough that someone can read it in two minutes and understand why you did what you did.\n\nUse these sections:\n\n- Decision and scope\n- What success looks like (plain language)\n- Assumptions we are betting on\n- Failure narratives (3–5 “how it goes wrong” stories)\n- Missing signals and how we will detect them\n- Guardrails and kill criteria (pause/rollback triggers)\n- Owners and check cadence\n- Open unknowns (labeled explicitly)\n\nTwo filled examples:\n\nAssumption: “Auto close at 72 hours will reduce backlog without hurting customer outcomes.”\n\nSignal: “Reopen rate and repeat contact within 7 days by queue, with billing tracked separately.”\n\nOwner: “Support ops lead reviews daily week one and posts a one-paragraph summary.”\n\nAssumption: “New routing that sends ‘login issues’ to Tier 1 will reduce time to first response.”\n\nSignal: “Escalation rate from Tier 1 to Tier 2 by region/language, plus top misroute reasons from sampling.”\n\nOwner: “Routing owner + frontline rep review 20 conversations daily across branches for three days.”\n\n### Pre-work rules: what data is allowed, what gets labeled as “unknown,” and why\n\nPre-work should be strict. Otherwise you show up with a 40-slide deck that nobody reads and everyone pretends they did.\n\nRules that keep it decision-grade:\n\n- Bring only data that can influence a decision in the next two weeks. If it won’t change scope, guardrails, or monitoring, it’s background noise.\n- Any metric that isn’t trustworthy at the branch level gets labeled **unknown**.\n- Separate lagging indicators from leading ones. If proof of harm arrives after churn, you need a different signal.\n\nCommon failure: letting “unknown” become “we’ll look later.” Later is where accountability goes to nap. Instead, connect unknowns to a condition: “We ship only to one queue until tagging health is proven,” or “We pause until we can break out the metric by region.”\n\nIf you want broader context on the original pre mortem framing (and why it works), this is a useful skim: [[1]](#ref-1 \"expectedvalue.co.uk — expectedvalue.co.uk\")\n\n## Truth-test branch signals before you trust the aggregate (queues, regions, teams)\n\n| Control | Where it lives | What to set | What breaks if it’s wrong |\n| --- | --- | --- | --- |\n| Set: Guardrail: Pause decision if signal integrity issues cannot be resolved within a defined timeframe | Decision-making process, risk management plan | A maximum time limit — e.g., 48 hours to address signal issues before halting the decision process | Rushing into poorly informed decisions, compounding errors, loss of credibility |\n| Set: Define clear branch-level pass/fail tests — e.g., per queue, region, team | Pre-mortem checklist, decision brief | At least 5 specific, measurable tests for each branch, with clear thresholds | Misleading aggregate metrics, incorrect resource allocation, biased decision-making |\n| Set: Identify potential misleading aggregates and their underlying branches | Pre-mortem analysis, data review sessions | Specific examples of aggregates that hide branch-level issues — e.g., 'overall CSAT is high, but Region X is failing' | Failure to address critical issues hidden by averages, wasted effort on non-problems |\n| Set: Establish a 'kill criteria' for signal integrity failures | Decision framework, pre-mortem protocol | A clear rule: 'If X% of branch signals are invalid / missing, decision is paused / reverted' | Committing to decisions based on unreliable data, loss of trust in metrics |\n| Validate data collection methods at the branch level | Data governance documentation, audit logs | Regular checks on data input processes, sample sizes, and reporting consistency for each branch | Inaccurate data, biased samples, inability to compare branches fairly |\n| Set: Review for missing coverage in critical branches — e.g., new markets, niche products | Metric dashboards, reporting requirements | Mandatory reporting for all critical branches, even if data is sparse. flag 'no data' as a risk | Blind spots in decision-making, neglecting high-risk or high-potential areas |\n\nAggregates are comforting. They also lie in very specific, repeatable ways.\n\nWhen someone says, “Support is fine, our SLA is fine,” they often mean, “The biggest queue looks okay.” Meanwhile a smaller queue is melting down, the language mix shifted, or one region is getting systematically misrouted.\n\nThis is where **branch level metric validation** stops being a nerdy analytics habit and becomes a risk control.\n\n### Branch-level truth tests: denominator sanity, tagging health, mix-shift detection\n\nThese truth tests are decision-grade because they have real pass/fail meaning:\n\n- **Denominator sanity:** do the counts match reality? Sudden drops in contact volume by branch without an operational explanation often mean tracking broke.\n- **Tagging health:** are tags/dispositions used consistently? If one team uses “customer education” as a bucket for everything, your slices are fiction.\n- **Mix shift detection:** did the customer mix change? A workflow update can look like a win if the queue got easier that week.\n- **Branch parity:** do key ratios look roughly similar across branches unless you have a known reason? Stable overall escalation rate can hide a single-region spike.\n- **Latency symmetry:** did one channel improve while another degraded? Email improving while chat worsens is a classic side effect.\n- **Reopen vs recontact split:** don’t trust one “resolution rate.” Separate “resolved and stayed resolved” from “resolved and came back.”\n\nDon’t wait for perfect dashboards. A scrappy branch check beats an elegant global average—as long as you’re explicit about what’s noisy and what’s missing.\n\n### Coverage checks: what you are not measuring (hours, languages, channels, severity)\n\nMissing signals often come from missing coverage. You measure what’s convenient, not what’s risky.\n\nTwo slices that uncover surprises fast:\n\n**Region and language.** A routing change that helps English tickets can quietly harm non-English tickets because skill rules get rigid, translation workflows differ, and coverage isn’t symmetrical.\n\n**Channel and severity.** A deflection change can reduce chat volume while pushing high-severity users into email, where they wait longer and escalate harder.\n\nRule that saves teams: if you can’t break out the metric by the branch that will carry the risk, you don’t have a metric. You have a mood.\n\nAnd yes—this is where teams get burned—assuming “no data” means “no problem.” In support ops, “no data” usually means “the branch isn’t instrumented,” which is a risk by itself.\n\n### Conversation sampling that doesn’t lie (how to avoid survivor bias and “resolved-only” reviews)\n\nSampling is where teams accidentally gaslight themselves.\n\nIf you only review “resolved” conversations, you select for success. If you only review escalations, you select for drama. You need the boring middle.\n\nA lightweight cadence most teams can sustain:\n\n- For each affected branch, review **10 conversations per day** for the first three days, then **10 per branch** twice in the next week.\n- Make sure at least **two** per branch are reopened, escalated, or transferred.\n- Include at least **one** auto-closed or bot-deflected conversation if automation is involved.\n\nIf the team can’t sustain that cadence, that’s not a moral failing. It’s a capacity signal. Shrink rollout scope.\n\n### The “green dashboard, red reality” patterns: when aggregates hide concentrated failures\n\nClassic scenario: you split a queue into “Account access” and “Billing” to improve routing. Aggregate first response improves by 12%. Great.\n\nBut the billing sub-queue now has a backlog slope that climbs every day because routing rules undercount severity, and the few billing specialists can’t keep up. The missing signal was branch backlog slope by queue—not overall first response time.\n\nStop rules matter here. If you can’t trust the branch numbers, you pause.\n\nA practical stop rule: if **two or more** core branch truth tests fail and can’t be resolved within **one working day**, don’t approve broad rollout. Narrow scope or delay.\n\nThis logic maps cleanly to kill criteria as a discipline: [[2]](#ref-2 \"howtothink.ai — howtothink.ai\")\n\n## Decide what to automate vs keep human: where bots and macros amplify blind spots\n\nAutomation in support is like adding a faster conveyor belt to a factory. If the sorting is wrong, you just deliver mistakes at a more impressive pace.\n\nA decision pre mortem is where you decide not only “can we automate?” but “what kind of failure does automation create, and will we notice it in time?”\n\n### A decision rule: automate when the failure is cheap and observable; keep human when it’s expensive and silent\n\nHere’s the rule operators can actually use:\n\nAutomate when the failure is **cheap and observable**.\n\nKeep a human when the failure is **expensive and silent**.\n\nCheap and observable looks like a wrong macro suggestion an agent can ignore, with a clear feedback loop.\n\nExpensive and silent looks like a bot confidently applying the wrong eligibility rule, “resolving” the case, and you only find out later when refunds spike or cancellations rise.\n\nTradeoff to say out loud: automation often improves speed and consistency, but it can also hide error behind a clean resolution status. Humans surface nuance faster, but they introduce variance and “hero workflows.” You’re choosing which failure mode you prefer—and which you can detect.\n\n### Automation failure modes in support: exception collapse, tone drift, misrouting, and false resolution\n\nAutomation misses aren’t random. They cluster.\n\n**Exception collapse:** edge cases get forced into the closest category, and the system stops admitting uncertainty. Early warnings: rising “other” tags, rising transfers, longer internal notes.\n\n**Tone drift:** macros and bots become technically correct and emotionally wrong. Early warnings: sentiment and complaint language rising even if handle time drops.\n\n**Misrouting:** routing rules optimize for speed but ignore skill nuance. Early warnings: escalation/transfer rate up in one region, team, or language.\n\n**False resolution:** auto close or deflection reduces visible backlog while increasing reopens and repeat contacts. Early warnings: “resolved” rising while recontact rises.\n\nThis is why “AHT down” isn’t a victory song. It’s a clue. Sometimes a good clue, sometimes a trap.\n\n### Human judgment failure modes: inconsistency, over-escalation, and “local hero” workflows\n\nHumans aren’t automatically safer. They fail differently.\n\n**Inconsistency** shows up as policy drift: two customers get two answers. Early warning: higher variance in outcomes across agents and teams.\n\n**Over-escalation** happens when agents are punished for being wrong, so they escalate to protect themselves. Early warning: escalations rise while quality doesn’t.\n\n**Local hero workflows** are the sneakiest because they look like excellence. One experienced agent invents a workaround that saves the day, and suddenly your system depends on a person, not a process. Early warning: a subset of cases only resolves when one or two people touch them.\n\nIn your pre mortem, always ask: “Where does this require hero behavior to look good?” If the answer is “nowhere,” you might be rounding up.\n\n### How to run an exceptions-first review (and what exceptions tell you about missing signals)\n\nInstead of reviewing average cases, review the exceptions first. Exceptions are where missing signals live.\n\nTwo examples that show up constantly:\n\nA **macro change** adds firmer language to reduce back-and-forth. AHT improves because agents send fewer follow-ups. But escalations and complaint language rise because the macro sounds like a parking ticket. The missing signal wasn’t AHT. It was escalation reason codes—or a simple qualitative flag from daily sampling: “customer reacted negatively to tone.”\n\nAn **auto routing change** sends “login issues” to Tier 1. First response improves. But Tier 1 transfer rate spikes for one product line because “login issue” is often a permissions issue. The missing signal was branch transfer rate by product area, plus a quick scan of the top misroute reasons.\n\nYou don’t need new tools to capture this. You need an exceptions log for week one. Keep it simple: branch, what failed, customer impact, how you detected it, owner, what you changed.\n\nSome teams use AI as an adversarial partner because it has no career incentive to agree with the boss. Whether you use AI or not, the adversarial posture is the point: [[3]](#ref-3 \"pre-mortem.ai — pre-mortem.ai\")\n\n## Pre-commit checks: counterfactuals, guardrails, and what to monitor in the first week\n\nThe difference between a brave decision and a reckless one is rarely intent. It’s whether you planned the exit.\n\nSupport ops changes feel reversible because they’re “just workflows.” In practice, reversals are painful: agents retrain, customers adapt, reporting breaks, and leadership loses confidence. Pre-commit checks make reversibility real, not theoretical.\n\n### Counterfactuals: ‘If we do nothing, what happens?’ and ‘If we reverse it, what breaks?’\n\nCounterfactuals expose missing signal risk fast.\n\nAsk before approval:\n\n- If we do nothing for 30 days, what happens to backlog, SLA risk, and customer pain?\n- If we reverse this change after one week, what breaks—training, routing rules, customer expectations, reporting continuity?\n\nIf you can’t answer the reversal question without hand-waving, you’re about to ship a one-way door disguised as a toggle.\n\n### Guardrails that matter: kill switches, thresholds, and staged rollout criteria\n\nGuardrails aren’t a list of metrics. Guardrails are action triggers.\n\nA pre-commit set that actually holds up in week one:\n\n- **Kill switch plan:** who can pause, where it’s communicated, what “paused” means for agents today.\n- **Scope control:** start with one queue or one region until signal integrity is proven.\n- **Threshold guardrails:** tied to action, not debate.\n- **Escalation path:** who to pull in if a branch starts failing.\n- **Customer comms stance:** what you tell customers if the change causes confusion.\n\nExample threshold (tune numbers to your baseline, keep the structure):\n\nIf **reopen rate** increases by a team-defined threshold in the affected queue for **two consecutive days**, you **pause rollout**, review **20 sampled conversations** from that branch, and decide within **one business day** whether to rollback or adjust.\n\nGuardrails for support automation should be measurable, owned, and timebound. Otherwise they’re just comforting words you say while shipping anyway.\n\n### First 24 hours vs first week: leading indicators (quality, recontact, backlog slope, escalations)\n\nDon’t treat week one as a single blob. The first 24 hours is about safety. The first week is about learning.\n\nIn the **first 24 hours**, watch fast-moving indicators:\n\n- Backlog slope by branch\n- Escalations and transfers by branch\n- Top exception types from the exceptions log\n- Repeating agent feedback (patterns, not one-offs)\n\nIn the **first week**, add indicators that reveal customer impact:\n\n- Recontact within 7 days for the same issue category\n- Reopen rate split by queue and severity\n- Quality review outcomes from sampled conversations\n- Complaint language frequency in notes/tags (even if informal)\n\nNotice what isn’t your north star: average handle time.\n\nAHT is useful, but it’s a classic “we got faster by getting worse” metric unless you pair it with recontact and quality.\n\nPractical decision rule: pick two “must not degrade” signals and treat them as your stop line. Teams that try to monitor ten things usually monitor none.\n\n### How to assign owners and cadence so monitoring actually happens\n\nMonitoring fails for one boring reason: nobody owns the calendar.\n\nAssign owners the way you assign incident roles.\n\nA cadence that’s lightweight but real:\n\n- **Daily for the first three days:** Data Buddy posts a short branch snapshot; Frontline Rep posts two patterns from sampling; Decider confirms continue or pause.\n- **Twice in week one:** Facilitator updates the one-page artifact with what was learned, including new unknowns.\n- **End of week one:** 20-minute review of guardrails, exceptions, and whether to expand scope.\n\nDocument in the same one-page artifact with a dated addendum. If it lives somewhere else, it will be forgotten. If it’s forgotten, it might as well not exist.\n\nLight humor, because we all need one: rolling out a support workflow change without guardrails is like “testing in production,” except your customers are the test suite and they don’t come with helpful error messages.\n\n## After you commit: turn every pre mortem into a decision library (so you stop relearning the same lesson)\n\nA pre mortem compounds when you can retrieve it. Otherwise it’s just a one-time ritual that makes everyone feel mature.\n\nThe goal is a small decision library that makes the next approval faster because you can say, “We’ve seen this movie before, and we know which scene goes wrong.” Over time, pre mortems also work as calibration tools because you can compare what you predicted with what actually happened: [[4]](#ref-4 \"howtothink.ai — howtothink.ai\")\n\n### The 10-minute post-decision addendum: what was true, what was missing, what changed\n\nKeep a standing 10-minute addendum one week after launch.\n\nDo three things:\n\n- Mark which assumptions were true, false, or still unknown.\n- Write the missing signal you discovered, especially if it surprised you.\n- Update guardrails for next time based on what actually moved.\n\nThis is where “we should remember that” turns into “we will not forget that.”\n\n### How to store and reuse pre mortems (tags: queue, change type, risk pattern, signals)\n\nDon’t overthink tooling. A shared folder and consistent naming usually beat an ambitious system nobody maintains.\n\nUse a few tags at the top of the artifact: queues affected; change type (automation, routing, staffing, SLA, policy); risk pattern (misrouting, tone drift, false resolution, mix shift); missing signals discovered; guardrails used; outcome after one week and one month.\n\nCommon mistake: letting pre mortems become write-only documents. If nobody can find the last three, you’re not building a library—you’re building a junk drawer.\n\n### Coaching loop: how to onboard new leads/operators using past misses\n\nHere is a copyable example library entry:\n\nDecision: “Enable auto close after 72 hours for low priority email queue.”\n\nBranches affected: “Email, low priority, billing adjacent.”\n\nMissing signal discovered: “Reopen rate by billing adjacent tag was the early warning.”\n\nGuardrail: “Pause if reopen rate rises above baseline threshold for two days.”\n\nLesson: “Backlog went down, repeat contact went up, so we narrowed scope and changed the close message.”\n\nThis prevents the next bad commitment when someone proposes auto close again and claims, “We did this before and it was fine.” You can respond, calmly and with receipts, “It was fine overall. It was not fine in that branch. Here is what we watch this time.”\n\nTo make this real on Monday, don’t build a program. Do one thing.\n\nSchedule the next “meeting before the meeting” and paste the one-page pre mortem headings into the invite. Then hold yourself to three priorities: (1) name one misleading aggregate you are currently trusting, (2) define two branch slices you’ll validate before shipping, and (3) set one guardrail with a clear pause action.\n\nYour production bar is modest: one-page artifact, one owner per signal, and a first-week monitoring note that is written down where the team can find it.\n\n## Sources\n\n1. [expectedvalue.co.uk](https://expectedvalue.co.uk/blog/pre-mortem-decision-making) — expectedvalue.co.uk\n2. [howtothink.ai](https://www.howtothink.ai/learn/kill-criteria) — howtothink.ai\n3. [pre-mortem.ai](https://pre-mortem.ai/ai-pre-mortem) — pre-mortem.ai\n4. [howtothink.ai](https://www.howtothink.ai/learn/pre-mortem-as-a-calibration-tool) — howtothink.ai\n",[39,43],{"_path":40,"path":40,"title":41,"description":42},"/en/blog/the-three-question-test-for-any-metric-trust-it-fix-it-or-ignore-it","The Three Question Test for Any Metric: Trust It, Fix It, or Ignore It","A practical support metrics trust test you can run in a live ops review. Use three questions to decide whether to trust, fix, or ignore a KPI, spot definition drift, routing and mix illusions, survey/",{"_path":44,"path":44,"title":45,"description":46},"/en/blog/why-your-best-researchers-still-get-it-wrong-the-hidden-failure-modes-in-decisio","Why Your Best Researchers Still Get It Wrong: The Hidden Failure Modes in Decision Systems","Support leaders keep making confident decisions off clean dashboards and smart analysis, then paying for it in churn risk and rework. This article breaks down hidden failure modes in support decision",1785947701932]