The Pre-Mortem for Metrics: A Workflow to Predict How Your Signals Will Mislead You

A practical metrics pre-mortem workflow for support leaders to validate CSAT, first response time, deflection, backlog, and QA signals before acting on trends. Learn guardrails, tripwires, and an exec

Lucía Ferrer
Lucía Ferrer
18 min read·

The moment to distrust a “better” dashboard (and run the pre-mortem)

You know the moment. The dashboard turns greener, the line slopes the “right” way, and someone says, “Great, we fixed it.” Meanwhile your team lead is Slacking you that agents are drowning, customers are angrier, and the backlog is getting weird.

Here is a concrete version I have seen more than once: first response time drops by 40 percent in a week. Leadership celebrates. Then you find out a new auto reply policy went live, or routing started firing an instant bot touch that counts as a response. The metric improved without the customer experience improving. It is not fraud. It is just a number doing exactly what you accidentally told it to do.

That is the moment to run a pre mortem for metrics.

A metrics pre mortem is a short, structured exercise where you assume your metric led you to a wrong decision, and you list the most likely ways the metric could have misled you before you act on it.

This is not a one off forensic investigation after the damage. It is a repeatable workflow you can run before you present a trend, change an SLA, cut headcount, reroute chat to a bot, or declare victory.

The point is to treat metrics like any other operational system with failure modes. Dashboards are not truth. They are instruments. And like any instrument, they drift, they get bumped, and sometimes they are wired to the wrong thing.

A familiar story: the metric improves, the floor reality doesn’t

Most “improvement without improvement” comes from ordinary changes: policy tweaks, channel mix shifts, tagging drift, or subtle definition edits. A small change in what counts as an eligible ticket, or when the clock starts, can move your trendline more than a month of coaching.

What a pre-mortem is: assume failure, then predict the causes

Pre mortems are a classic risk tool popularized in project settings, and the same logic maps cleanly to measurement. If you want the mindset in its original form, see the general pre mortem framing here: [1].

Scope: which support signals this workflow covers (CSAT, FRT, deflection, backlog, QA)

This workflow is meant for support signals that routinely drive big decisions: CSAT trend validation, first response time, deflection, backlog health, and QA scores. It also helps with adjacent signals like reopen rate, escalations, and transfers, but we will keep the focus on the five most abused charts.

Set the pre-mortem up: define the decision, the metric’s job, and what would count as “misleading”

Control Where it lives What to set What breaks if it’s wrong
Set: Metric's Job Decision Memo, Metric Catalog Why metric exists, decision it informs Irrelevant metrics, 'vanity' dashboards
Set: Exception: Contextual Override Decision Memo, Incident Report Conditions to ignore metric (e.g., holidays, outages) Panic over normal fluctuations, wasted investigation
Set: Decision Memo (one-paragraph) Shared doc (Confluence, Notion) Who decides, what changes, what risk Misaligned actions, wasted effort
Set: Misleading Criteria Decision Memo, Pre-Mortem template Segmentation sensitivity, instrumentation drift, response bias False positives/negatives, incorrect conclusions
Set: Pre-Mortem Timebox Meeting invite 60-90 min session Incomplete analysis, missed insights
Set: Pre-Mortem Participants Meeting invite Ops lead, team lead, QA, analytics partner Missed perspectives, no actionable outcomes
Set: Decision Rule: Action Trigger SOP, Playbook Metric change that triggers specific action Inconsistent responses, reactive decisions
Set: Guardrail: Metric Thresholds Dashboard, Alerting system Upper/lower bounds for 'healthy' behavior Delayed incident response, unnoticed degradation

The fastest way to turn metrics into theater is to start with the dashboard. The disciplined way is to start with the decision. As Ashwinikumar Patil argues in “Your metric has a design flaw,” the flaws show up when the metric meets real incentives and real humans: [2].

Write the decision memo first: what are you about to change if the metric is ‘up’ or ‘down’?

Before you debate whether the trend is “real,” write a one paragraph decision memo that makes the stakes explicit:

Decision memo template: “If

[metric] moves

[direction] for

[time window], then

[decider] will

[action] to

[what part of support], because we believe it means

[interpretation]. The main risk if we are wrong is

[customer or cost risk]. We will not act if

[stop the line condition].”

Concrete anchor number one: “If FRT stays under 30 minutes for two consecutive weeks, the support director will reduce weekend coverage.” That is a high risk decision. Your validation bar must be higher than “the line went down.”

Concrete anchor number two: “If deflection goes up after we add a help center prompt, the product VP will expand the bot to billing.” That is also high risk, because deflection can be ticket suppression, not successful self serve.

Define the metric’s job: detect demand, speed, effort, quality, or trust?

Metrics fail when they get promoted to jobs they cannot do.

First response time is mostly a speed signal. It is not a quality signal.

CSAT is mostly a trust and sentiment signal. It is not a demand signal.

Deflection is an interaction design signal. It is not automatically a cost savings signal.

Backlog is a capacity and flow signal. It is not automatically “customer pain,” unless you connect it to aging and tier.

QA is a consistency signal. It is not automatically an outcome signal.

When you name the job, you also name what the metric is allowed to ignore. That is how you stop yourself from using one chart to answer five different questions.

Freeze definitions and boundaries: numerator, denominator, time window, and eligible tickets

Most teams try to validate support metrics after the debate starts. Do the opposite. Freeze definitions up front.

Decide what is in and out. Are social tickets eligible? Are bot only conversations excluded? Are reopened tickets counted twice? Does “pending” stop the clock? What time zone defines a day?

Two stop the line conditions you should treat as non negotiable:

First, if the definition changed mid period. Even a “small” change like counting bot touches as first response can invalidate your comparison.

Second, if eligibility changed. If you moved VIP customers to a separate queue, your CSAT and FRT are now describing a different population.

A practical tip: keep a simple definitions page in your support metrics glossary and update it like you would update a policy. If the glossary is missing, your dashboard is effectively running on tribal knowledge.

The 7-step metrics pre-mortem workflow (table)

The table below is the metrics pre mortem workflow you can reuse across CSAT, FRT, deflection, backlog, and QA.

Set: Metric's Job. One sentence that prevents you from asking the metric to do five jobs.

Set: Decision Memo (one-paragraph). If you cannot name the action, you are not ready to argue about the trend.

Set: Misleading Criteria. The list that gives you permission to say “we are not acting yet.”

Set: Pre-Mortem Timebox. Thirty minutes is enough to catch the big lies without turning it into a committee.

Trace the metric supply chain: where the number gets made (and where it gets bent)

A metric is not a number. It is a supply chain. It starts as messy customer demand and ends as a clean looking chart. The bending happens in the middle.

Map the lifecycle: ticket created → categorized → queued → responded → resolved → surveyed → scored

If you want to know whether you can trust a support metric, map the lifecycle in plain language. You are not diagramming for fun. You are asking where the number is born, and where it can be distorted.

A vendor neutral “metric supply chain” checklist looks like this:

  1. Entry: where did the customer come from, and what counts as a ticket versus a message?

  2. Identity: how do you deduplicate, merge, or split conversations?

  3. Categorization: what tags, forms, or contact reasons exist, and who chooses them?

  4. Queuing: what routing rules and priority policies decide who sees it first?

  5. Touches: what counts as first response and next response, including automation?

  6. Resolution: what counts as solved, and what statuses pause work without pausing time?

  7. Reopen: what triggers reopening, and how is it counted?

  8. Survey: who receives CSAT, when, and at what send rate?

  9. Scoring: what is excluded from reporting, and what gets backfilled later?

The audit artifacts to pull before you argue about the trend are boring, and that is why they work: policy change log, routing rules summary, taxonomy change notes, survey send rules, and the current QA rubric version.

Common mistake number one: trusting memory instead of artifacts. “I do not think we changed anything” is not evidence. People forget. Tools change. Defaults update. Pull the notes.

Channel mix and eligibility shifts (chat vs email vs social) that break comparability

Channel mix is the quiet assassin of comparability.

Concrete example: you launch chat and aggressively promote it in app. Suddenly 30 percent of your volume moves from email to chat. First response time looks amazing because chat is staffed differently and customers abandon sooner. CSAT may rise because chat issues are simpler, or it may fall because the handoff feels rushed. Your dashboard is now describing a different world.

This is where segmentation saves you. Always check trend direction by channel, then by tier, then by issue type. If the overall line improves while the email segment worsens, you did not “improve support.” You shifted demand.

A practical tip: track channel share as a first class tripwire. If channel share moves, treat most other comparisons as provisional until you explain the mix shift.

Tagging drift and reclassification: when the taxonomy changes without permission

Tagging drift is what happens when your taxonomy changes in practice even if nobody “approved” a change.

Concrete example: a team lead tells agents to start using “login issue” instead of “auth issue” because it reads better. Another lead trains a new cohort that “billing question” is the safest tag when unsure. Within a month, your top contact reasons change. Not because customers changed, but because humans did.

You can detect tagging drift without fancy tooling. Look for sudden jumps in tag distribution, a rise in “other,” or a shift in the mix of contact reasons inside the same product area. If you already have a tagging taxonomy audit guide, use it here and treat the output as part of your pre mortem packet.

Common mistake number two: cleaning the dashboard instead of fixing the taxonomy behavior. If you hide messy tags, you lose the very signal that tells you your measurement is degrading.

Timestamp traps: first response, next response, reopen, ‘pending’, and bot touches

Time based metrics are the easiest to “improve” accidentally.

If you batch triage twice a day, first response time can look better if you send quick acknowledgements, but time to resolution can worsen because real work starts later. One operational change can move multiple metrics in opposite directions.

Also watch for clock stopping states. “Pending” can mean “waiting for customer,” but it can also mean “we parked it to hit SLA.” If your backlog policy encourages parking, your backlog can look stable while your aging distribution quietly rots.

A useful heuristic: any time a metric depends on timestamps, ask, “Which events are human, which are automated, and which are policy driven?” That is where distortion lives.

Decide what to trust: tradeoffs, guardrails, and decision rules that stop KPI theater

Support leaders are not unethical because they optimize a metric. They are doing what the organization implicitly asked them to do. The fix is not shaming. The fix is making the tradeoffs explicit and putting guardrails around the optimization.

If you like the broader framing of predicting metric failure before the rollout, “Canary metrics lie more than you think” makes the same point in a different domain: [3].

Pick the ‘primary’ metric per decision—and name the sacrificed dimension

Every decision has a primary metric. Pick one, and write down what you are willing to sacrifice.

If the decision is staffing, your primary might be backlog aging or time to first human response. The sacrificed dimension might be “handwritten empathy” or “non urgent follow ups.” Name it, so you can manage it.

If the decision is automation, your primary might be deflection, but the sacrificed dimension might be “white glove handling for edge cases.” If you pretend you can have both, you will create a bot that pleases a chart and annoys a customer.

Light humor, because we all need it: a dashboard is a speedometer, not a physician. It will tell you you are going 80. It will not tell you why the engine is smoking.

Add guardrails: what must not get worse while you optimize (quality, repeat contacts, escalations)

Guardrails are the metrics you refuse to sacrifice.

Here are concrete guardrail pairings that work in real operations:

First response time with repeat contact rate. If FRT improves but repeat contacts rise, you are acknowledging faster, not solving better. This is a common first response time misleading pattern.

Deflection with escalation or complaint rate. If deflection rises and escalations rise, you may be suppressing tickets or forcing customers into higher friction paths. This is the classic deflection metric pitfalls scenario.

QA score with customer outcomes like reopen rate or post resolution CSAT. If QA rises but reopens rise, your rubric may be rewarding politeness over effectiveness.

Backlog size with backlog aging distribution. If the count is stable but the tail is growing, you are building a graveyard of old tickets.

A practical tip: keep guardrails to two or three for any one decision. Too many guardrails turns into “we can never change anything.”

Decision rules: what evidence threshold turns a trend into action

You need a trust rubric that turns “interesting” into “decision grade.” Without it, pressure will force you to overclaim.

A simple trust rubric with threshold examples:

Definition stability: no metric definition changes during the comparison window.

Eligibility stability: channel and tier eligibility unchanged, or changes explicitly segmented.

Survey stability for CSAT: send rate and response rate within a narrow band, not swinging because of throttling or timing changes.

Staffing mix stability: major changes in new hire share or outsourcing flagged, because they change both speed and sentiment.

Volume context: major demand spikes or product incidents noted, because they can swamp process effects.

Action gating example that should be policy: do not ship a staffing reduction based on improved FRT unless repeat contacts and backlog aging are stable, and the channel mix did not shift toward easier work.

Counter-metrics for the big five (CSAT, FRT, deflection, backlog, QA)

For each primary metric, choose counter metrics that tell you whether you are “winning on paper.”

CSAT counter metrics: response rate, survey send rate, and ticket mix by difficulty. If response rate collapses, your CSAT trend is less trustworthy.

FRT counter metrics: time to resolution, transfers, and reopens. If you respond quickly but bounce customers between teams, you are paying the cost elsewhere.

Deflection counter metrics: escalation rate, contact us clicks after self serve, and complaint volume. If self serve is “successful,” frustration should not spike.

Backlog counter metrics: aging distribution and VIP share. A backlog with the same size but older tickets is not healthy.

QA counter metrics: reviewer agreement and outcome proxies like reopens. QA without consistency is performance art.

Segmentation pitfall to avoid: slicing until you find a win. If you look at 20 segments, one will look better by luck. The rule that prevents this is simple: pre register your segments in the pre mortem, and limit them to the ones that match the decision. Channel, tier, and top contact reasons are usually enough.

Failure modes: how support metrics ‘improve’ while support gets worse (and the tripwires to catch it)

A good pre mortem does not just say “metrics can lie.” It predicts how they will lie in your environment, then puts tripwires in place so you catch distortion early.

The operational pre mortem idea has been written about in several contexts, including observability and risk prevention. If you want extra perspective, see [4] and [5].

CSAT failure modes: response bias, survey throttling, and ‘easy ticket’ skew

Failure mode one: response bias. If only very happy or very angry customers answer, your CSAT becomes a mood ring for extremes.

Failure mode two: survey throttling or timing changes. If you start sending fewer surveys, or you send them only after certain macros, your CSAT trend can “improve” because you selected the respondents.

Failure mode three: easy ticket skew. If you shift simple password resets to chat and those are overrepresented in survey sends, CSAT rises while complex issues remain painful.

Tripwire ideas that catch this: survey send rate, response rate, and the share of surveys by contact reason. If those move, do CSAT trend validation before you claim improvement.

FRT failure modes: autoresponders, bot touches, triage batching, and clock-stopping states

Failure mode one: autoresponders and bot touches counted as first response. It makes FRT look heroic while customers still wait for a human.

Failure mode two: triage batching. You can send quick acknowledgements to “stop the clock,” then batch real work later. FRT improves while time to resolution worsens.

Failure mode three: clock stopping statuses. If “pending” or “waiting” is used as a parking lot, you can meet SLA and still provide slow service.

Tripwires that catch this: share of tickets with an automated first touch, time to first human response, and the distribution of time to resolution, not just the average.

Deflection failure modes: ticket suppression vs true self-serve success

Failure mode one: suppression dressed up as deflection. You hide contact options, add friction, or push customers into dead ends. Ticket volume drops. Customer effort rises.

Failure mode two: misattribution. Customers read an article and still open a ticket, but your analytics credits the article view as success.

Here is the trust harming scenario to take seriously: aggressive deflection increases frustration, which increases escalations and public complaints. You “saved” tickets and spent trust. That is a terrible trade.

Tripwires that catch this: escalation rate, complaint rate, and contact us clicks after self serve. If you have a deflection measurement guide that distinguishes suppression versus success, this is where it belongs in the workflow.

Backlog & QA failure modes: hiding work in statuses, cherry-picking reviews, rubric drift

Backlog failure mode one: hiding work in statuses. Tickets move into “pending,” “on hold,” or internal queues that are not counted in the headline backlog.

Backlog failure mode two: closing and reopening cycles. If you close aggressively to keep backlog down, reopens rise and customers lose confidence.

QA failure mode one: cherry picking reviews. If reviewers sample only easy interactions, QA rises while real quality does not.

QA failure mode two: rubric drift. The rubric changes, or reviewers interpret it differently over time, making trend comparisons shaky.

Tripwires that catch this: aging distribution (especially the long tail), reopen rate, the share of tickets in each status, and reviewer agreement checks.

Tripwire monitoring: the small set of ‘always-on’ signals that flag distortion early

Tripwires are your always on support KPI guardrails. You do not need many. You need the right ones, owned by someone, checked on a cadence.

Here are seven that cover most distortion:

  1. Survey send rate and 2) survey response rate, owned by the CSAT owner, checked weekly.

  2. Channel share, owned by support ops, checked weekly.

  3. Tag distribution shift, owned by the taxonomy owner or ops, checked monthly, and weekly during big launches. If you want a simple approach, watch for sudden changes in top tags and in “other.”

  4. Reopen rate, owned by the team lead, checked weekly.

  5. Transfer and escalation rate, owned by the team lead plus QA, checked weekly.

  6. Backlog aging distribution, owned by ops, checked weekly.

When a tripwire triggers, do not panic and do not rationalize. Run a small investigation play that matches the distortion:

  1. Confirm whether a policy, routing, taxonomy, survey rule, or rubric changed in the last two weeks.

  2. Segment the metric by channel, tier, and contact reason to see where the movement actually lives.

  3. Pull a small sample of real conversations from the moving segment and sanity check whether the customer experience improved or just the instrument.

That is your support dashboard audit checklist in practice. It is not glamorous. It is how you avoid confident wrong decisions.

Make it routine: a 30-minute cadence, a one-page artifact, and how to brief leadership without overclaiming

Most teams fail at metric validation for one reason: they treat it as extra work. The trick is to make it a short cadence that happens at the moments when you are most likely to overclaim.

The cadence: when to run the pre-mortem (before QBRs, headcount changes, policy shifts)

Run the metrics pre mortem workflow before QBRs, before headcount changes, before SLA policy shifts, and before major routing or automation rollouts. If you are about to make a decision that is hard to reverse, you want decision grade metrics, not vibes.

A lightweight agenda for a 30 minute pre mortem meeting looks like this: spend 5 minutes on the decision memo, 10 minutes on definition and eligibility freeze, 10 minutes on guardrails and tripwires, and 5 minutes agreeing on what would block action.

The one-page artifact: assumptions, definition freeze, tripwire status, and decision recommendation

Keep a one page artifact that lives next to the dashboard. It should have:

Assumptions and decision memo.

Metric job and definition freeze, including eligibility boundaries.

Misleading criteria and stop the line conditions.

Segmentation plan.

Guardrails and their current status.

Tripwires with owners and last check date.

Decision recommendation: proceed, proceed with caution, or blocked.

This one page is also where your internal links belong. If your org has a support metrics glossary, a tagging taxonomy audit guide, a backlog triage framework, a deflection measurement guide, and a QA scorecard design article, reference them from the artifact so people stop reinventing definitions.

Language that preserves credibility: how to present ‘directional’ vs ‘decision-grade’ metrics

When you brief leadership, use a script that makes uncertainty concrete instead of evasive.

Pattern: “Claim:

[what moved]. Confidence level: directional or decision grade. Risks:

[the top two ways it could be misleading]. Next validation step:

[what you will check and when].”

You can also say, plainly, “This is blocked by a definition change mid period,” or “This is directional because channel eligibility shifted.” That is professional rigor, not excuse making.

To make this real on Monday: first action, copy the 7 step metrics pre mortem workflow table into a doc template and run it on one metric, ideally first response time since it is famously easy to game. Your three priorities are to freeze the definition, add two guardrails (repeat contacts and backlog aging are good defaults), and add two tripwires (channel share and survey response rate catch a lot). Your production bar is simple: in 30 minutes, you should produce a one page artifact that lets you either act with confidence or explicitly say “not yet” with reasons.

Primary CTA: Download or copy the 7 step metrics pre mortem workflow as a doc template and run it before your next metrics review.

Secondary CTA: Start with one metric this week, add two guardrails plus two tripwires, then expand to the full set next month.

Sources

  1. vikasmalpani.com — vikasmalpani.com
  2. tightmargins.substack.com — tightmargins.substack.com
  3. medium.com — medium.com
  4. linkedin.com — linkedin.com
  5. activecollab.com — activecollab.com