The meeting is tomorrow: run a 20-minute metric pre-mortem before you trust the dashboard
You know the meeting. A dashboard is on the screen, the numbers look clean, and someone says, “So we can cut two heads, right?” because deflection is up and first response time is down. In support ops, that is how decisions get made: fast, confident, and sometimes painfully wrong.
Here is the uncomfortable truth: metrics are fallible witnesses. They do not lie on purpose, but they can be coached, confused, or quietly changed when workflows change. If you have ever celebrated a CSAT bump and then discovered it came from a survey trigger change, you already get it.
A concrete example I have seen more than once: deflection “improves” after a help center refresh and leadership approves a staffing cut. Two weeks later inbound contacts are unchanged, escalations tick up, and agents report more angry customers. What happened? The deflection metric was counting “article viewed” as “issue avoided,” while customers still came back through chat. The dashboard looked right. The decision did not.
That is why pre mortems for metrics work. A metric pre mortem is a short, structured exercise where you assume the metric will mislead you next month, then you list the most likely reasons why. You are not debating the number. You are stress testing its trustworthiness before it drives a high stakes call.
In 20 minutes, the ritual should produce three outputs you can walk into the meeting with.
First, the assumptions the metric depends on, like what counts as “responded” or who is “in scope.” Second, the failure modes that would make the metric look healthier than reality, like channel leakage, reclassification, or survey gaming. Third, decision rules for leadership, such as “we can proceed if these two cross checks agree” or “we pause if we see a step change aligned to a workflow release.”
Why numbers that look right create wrong confidence
Support metrics tend to be “right looking” because they are precise, not because they are true. The dashboard will happily compute first response time down to the minute, even if your start timestamp quietly shifted last week.
The pre mortem promise: identify what would make the metric lie
Pre mortems are widely used as a cheap risk audit before launches and commitments, because they surface the few risks that matter instead of the twenty generic worries everyone can generate on demand. The same idea applies to metrics: the cheapest time to find out your KPI is compromised is before you act on it, not after you reorganize your team.
If you want a broader primer on the pre mortem technique outside metrics, this skill write up is a solid reference point: [1]
What you walk into the meeting with (not just a number)
Walk in with the number, yes. Also walk in with a one paragraph caveat, the two fastest validation checks you ran, and the one thing you will not let leadership do until you resolve it. That is what separates “reporting” from decision support.
Practical tip: if you only have time for one move, print the dashboard (or paste a screenshot into the deck) and write one sentence under each headline metric that starts with “This could be artificially up or down if…” Then pick the likeliest cause and validate it.
Write the 'metric contract' first: what the number actually includes, excludes, and rewards
Most teams try to do a pre mortem with a metric that is not fully defined. That is like arguing about a movie when you are not sure everyone watched the same cut.
The fix is a “metric contract.” It is a short, plain language agreement that says what the number includes, excludes, and rewards. It turns fuzzy dashboard labels into something you can actually validate.
If you take one idea from experimentation culture, steal this one: decide what counts as a win before you stare at the results. Locking the definition early reduces motivated reasoning later, even when everyone involved is well intentioned. (This is one reason pre registration exists in experiments.) A readable piece on that mindset is here: [2]
Definitions that drift: numerator, denominator, windows, timestamps
Support metrics drift because someone “makes a small improvement” to workflow that changes what gets counted. Drift is not always malicious. It is often a side effect of reasonable operational change.
Here is a metric contract checklist you can paste into your doc. Keep it short enough that someone will actually read it.
Metric name and decision it supports. Example: “First response time for staffing and coverage decisions.”
Exact definition of numerator and denominator. What is being counted, and out of what universe.
Time window and reporting cadence. Example: “Rolling 7 days, reviewed weekly in ops, monthly in QBR.”
Start and stop timestamps. Example: “Start when a customer message creates a new conversation” versus “start when ticket enters an agent queue.”
Inclusions and exclusions. Example: which ticket types, which statuses, which spam filters, which merged conversations.
Segments that must be shown alongside the overall number. Example: chat versus email, enterprise versus SMB, region and language.
Ownership. One metric owner who updates the contract when workflows change.
Known failure modes and the top two cross checks.
“Do not use this metric for…” guardrail. Example: “Do not use deflection alone to justify staffing cuts.”
Common mistake number one: teams treat the dashboard label as the definition. “Resolution time” sounds universal, so people assume it is universal. It never is. Do the contract first, then analyze.
Segmentation that changes the story (channel, tier, issue type, region, language)
Segmentation is not decoration. It is how you catch channel leakage and gaming.
Imagine your first response time looks great overall. Then you break it out by language and discover English chat improved while Spanish email doubled. That is not “noise.” That is a decision changing signal.
Two concrete anchors that matter in real support orgs:
First, channel mix shifts. If chat volume rises because you nudged customers away from email, your response time may improve simply because chat is staffed differently.
Second, tiering changes. If enterprise tickets get reclassified into a higher priority queue, overall resolution time may worsen while your most valuable customers are happier. Or the reverse: you optimize overall and quietly harm your highest revenue segment.
Practical tip: pick no more than four default segments you always show. For most teams, channel, tier, issue type, and region or language is enough to surface the big problems without turning the review into a data festival.
Incentive check: what behavior this metric silently pushes
Every KPI is a tiny policy. It tells people what “good” looks like.
This is where trustworthiness of CSAT and response time often breaks. If you reward speed without a quality counterweight, people will close fast and clean up later. If you reward CSAT alone, people will cherry pick “safe” tickets and avoid hard conversations.
Here are two contract clauses you can use, with tradeoffs spelled out.
Clause example for first response time:
“First response time excludes automated acknowledgements and bot messages. It measures time to first human agent reply in the customer’s channel. Reason: we are making staffing decisions, and auto replies do not reduce customer anxiety. Tradeoff: if automation genuinely resolves simple issues, this clause may understate the customer experience improvement, so we will pair it with containment rate for bot resolved conversations.”
Clause example for resolution time:
“Resolution time pauses during customer waiting periods after an agent asks a question. It resumes when the customer replies. Reason: we want to measure internal throughput and remove customer delay from the operational metric. Tradeoff: this can hide customer frustration with back and forth, so we will also monitor number of touches and reopen rate as a customer proxy.”
Now the part everyone forgets: reclassification and status changes alter who is in scope.
If agents re tag “billing bug” to “billing question” to route it faster, your bug backlog drops and your question volume rises. If you introduce a “pending customer” status, you may “improve” resolution time without actually solving anything faster.
Practical tip: add one line in the contract that says, “If we change taxonomy, routing rules, or automation that touches this metric, we annotate the dashboard that week.” It is boring. It saves you.
Run the pre-mortem prompts: if this metric misled us next month, what went wrong?
A good pre mortem is not a brainstorming party. It is a time boxed ritual with a facilitator who keeps it crisp and a decision owner who actually uses the output.
If you have never run one, borrow the structure that shows up in strong product and policy pre mortems: assume failure, list reasons individually first, then converge on the few that matter and turn them into mitigations and tests. This framing is widely used because it avoids groupthink and surfaces cross functional risks early.
Facilitator script: roles, timebox, and ground rules
Start by stating the decision the metric will drive. “We are using deflection and backlog to decide whether to expand automation and reduce weekend coverage.” Make the stakes explicit.
Then set two ground rules.
First, no defending the metric during the writing phase. You are generating failure modes, not litigating.
Second, every concern must imply an observable signal. If someone says “CSAT is biased,” ask “What would we expect to see if that is true?”
You need four roles.
The facilitator runs time and keeps the group honest.
The metric owner answers definition questions and owns follow up.
A skeptical reviewer, usually someone from QA, analytics, or an experienced team lead, plays the “assume it is broken” role.
The decision owner commits to a decision rule, not just “interesting, thanks.”
Prompt sets by metric family (CSAT, speed, deflection, backlog) and the diagnostic signals they imply
Use prompts that match common support KPI failure modes. Here are prompts that reliably catch the real issues, with diagnostic signals called out on several so you can move fast.
- Definition drift: “What part of the workflow changed that could shift the start or stop time without changing actual performance?”
Diagnostic signal: a step change on the exact day a routing, SLA, or status policy changed.
- Ticket reclassification: “Did we change tags, forms, or categories in a way that moves work between buckets?”
Diagnostic signal: sudden volume drop in one issue type and matching rise in another, with total contacts steady.
- Channel leakage: “Did customers move to another channel that is not counted, or counted differently?”
Diagnostic signal: deflection up while inbound contacts across all channels do not fall, or while phone complaints rise.
- Sampling bias in CSAT: “Who is no longer getting surveyed, and why?”
Diagnostic signal: response rate drops, but CSAT rises, especially in one channel or segment.
- Survey gaming: “Did agents learn how to avoid surveys or nudge ratings?”
Diagnostic signal: CSAT climbs but negative comments remain intense, or the share of 5 star ratings spikes while 3 and 4 collapse.
- Automation side effects: “Did we add bots, macros, or auto closures that change the pace of the queue?”
Diagnostic signal: variance drops toward zero, or first response time improves while reopen rate climbs.
- Backlog optics: “Did backlog shrink because we closed more, or because we redefined what counts as open?”
Diagnostic signal: backlog down while the aging tail worsens, meaning fewer old tickets get resolved.
- SLA pressure and Goodhart: “If we tightened SLAs, what did people do to hit them that might hurt quality?”
Diagnostic signal: resolution time down and escalations or reopens up.
Merge and duplicate behavior: “Are we merging or splitting conversations differently?”
Segment imbalance: “Which segment could be driving the overall movement, and would that change the decision?”
You do not need all ten every time. Pick six to eight based on the metric that will drive the decision.
Turn hypotheses into quick tests you can run before the meeting
The difference between a useful pre mortem and a depressing one is whether you convert “could be wrong because…” into a quick test with an owner.
Use this simple format for each top hypothesis: “If X is true, we will see Y. We will test by comparing A versus B. Owner is C. Due by D.”
Two quick test examples that work in the real world:
First, to validate automation side effects on response time, compare human only conversations versus automation assisted conversations over the same week. Confirming result: the “improvement” exists mostly in automation assisted flows, while human only stays flat. Disconfirming result: both improve similarly.
Second, to catch reclassification effects, audit a sample of tickets from two weeks before and two weeks after a policy or form change, and compare the distribution of categories and statuses. Confirming result: the same customer problems are now landing in different buckets, or spending more time in “pending customer.” Disconfirming result: category mix is stable.
Practical tip: keep the tests pre meeting sized. If it takes two weeks and a data project, it is not a pre mortem test, it is a roadmap item.
When you document this for leadership, do not bury it. Put a one slide “metric caveats and checks” page directly after the dashboard screenshot. It forces the right conversation.
Diagnostic signals to assume first: the patterns that mean your support metric is lying (and what to test fastest)
| Assignment strategy | Best for | Advantages | Risks | Recommended when |
|---|---|---|---|---|
| Definition Drift (CSAT) | Identifying changes in what customers consider 'satisfactory' | Catches evolving customer expectations or product changes | Misinterpreting genuine sentiment shifts as noise | CSAT drops without obvious service changes. after product updates |
| Channel Leakage (Deflection) | Understanding if self-service pushes customers to other channels | Reveals hidden costs and poor customer experience from deflection | Difficult to track cross-channel journeys accurately | Deflection rate increases but overall contact volume doesn't decrease |
| Sampling Bias (CSAT) | Ensuring survey respondents represent the true customer base | Prevents skewed results from unrepresentative samples | Excluding valid feedback by over-correcting for bias | CSAT scores are unusually high/low compared to qualitative feedback |
| Survey Gaming (CSAT) | Identifying attempts to artificially inflate or deflate scores | Protects metric integrity from internal or external manipulation | False positives can lead to distrust in legitimate high scores | Unusual patterns in survey completion times or identical responses |
| Automation Side-Effects (Deflection) | Assessing if automation creates new, harder-to-solve problems | Highlights unintended negative consequences of new tech | Blaming automation for unrelated issues | Deflection rises but complex ticket volume also increases |
| Ticket Reclassification (Speed) | Detecting agents manipulating ticket categories for better metrics | Uncovers gaming behavior that inflates speed metrics | Over-monitoring can reduce agent autonomy and trust | Average handling time (AHT) drops sharply for specific categories |
| Mixed Causes (All Metrics) | Recognizing that single explanations are often insufficient | Encourages holistic investigation, avoids premature conclusions | Can lead to analysis paralysis if not managed | Initial tests for single failure modes are inconclusive |
If you have ever watched a leadership team fixate on a single line trending up, you know why this section matters. Your job is not to be cynical about metrics. Your job is to notice when the metric is acting like a witness who changed their story.
The fastest way is to memorize a few red flag shapes and divergences. Metrics rarely fail in subtle ways. They fail in patterns.
Red-flag shapes in the trend (step changes, cliff effects, flatlines, and sudden variance drops)
A step change is the classic “something changed in measurement, not reality.” Treat any step change aligned to a workflow release as a caveat trigger.
A cliff effect is when performance looks fine until it suddenly does not. This often shows up when a backlog passes a threshold and a queue policy kicks in.
A flatline is underrated. If CSAT never moves while comments swing wildly, your survey may be sampling the wrong set.
A sudden variance drop is a hallmark of automation or process gating. Humans are messy. If the metric becomes too smooth, ask why.
Decision threshold you can adopt: any step change that lines up with a known workflow or automation change requires a caveat, even if the step is “good news.”
Cross-metric divergences that expose hidden tradeoffs (resolution time down while reopens up)
Support KPI failure modes often reveal themselves when two metrics stop agreeing.
If resolution time improves while reopen rate rises, you are probably closing faster, not solving better.
If CSAT rises while escalation volume rises, you may be surveying only the easy cases.
If deflection rises while inbound contacts do not fall, you are likely counting activity, not avoided work.
Here is the Goodhart moment, specific to support: when you tighten an SLA without a quality counterweight, people will hit the clock by pushing work into statuses that stop the timer or by closing borderline tickets. The dashboard celebrates. Customers reopen. This is not because your team is bad. It is because your metric became the target.
Decision threshold you can adopt: if reopen rate increases meaningfully right after an SLA change, pause any staffing reduction until you validate.
Segment breakouts that reveal leakage or gaming (channel, tier, issue type, language and region)
Segment gaps tell the truth your overall average hides.
Decision threshold you can adopt: if any segment gap doubles week over week, treat the overall number as caveated until you explain the shift.
Now put it all together in a framework you can reuse.
Mixed causes are common. A bot rollout can create channel leakage and sampling bias at the same time. Single factor stories are comforting and often wrong.
Two mini audit recipes you can run quickly without turning this into a data science project:
First audit recipe for survey bias: pull a simple comparison of surveyed versus not surveyed tickets over the same period, focusing on issue type, segment, and severity proxy (like escalation or refund involvement). Confirming result: surveyed set is systematically easier or different. Disconfirming result: surveyed set looks similar.
Second audit recipe for deflection validity: compare the change in self service activity to the change in contacts for the same topics. Confirming result: article views rise but topic specific contacts do not fall, suggesting “view” is not “deflect.” Disconfirming result: topic specific contacts drop in line with increased self service.
After you apply the table, call out a few controls explicitly so people remember them.
Definition Drift (CSAT): treat any CSAT jump aligned to survey logic changes as caveated.
Channel Leakage (Deflection): treat “deflection up with contacts flat” as a validation requirement, not a win.
Sampling Bias (CSAT): never interpret CSAT without response rate and mix.
Survey Gaming (CSAT): watch for unnatural rating distribution shifts.
Automation Side-Effects (Deflection): monitor complaint themes and failed contact paths, not just containment.
Decision rules: when to trust the metric, when to caveat it, and when to stop-the-line
The goal is not to become the person who distrusts every number. That person is exhausting and eventually ignored. The goal is to be precise about confidence.
A simple confidence tiering works because leadership can act on it.
A simple confidence tiering (trust / caveat / pause) driven by signals + tests
Tier 1, Trust: the metric definition is stable, no red flag shapes, and your top two cross checks agree.
Tradeoff: you may move slower than a team that blindly acts, but you will waste less time undoing bad calls.
Tier 2, Caveat: you see one or more diagnostic signals, but you have enough validation to interpret directionally. You proceed, but you constrain the decision.
Tradeoff: you keep momentum, but you avoid irreversible actions like cutting headcount or locking an SLA promise.
Tier 3, Pause: you hit a stop the line trigger. You do not use the metric for the decision until a specific validation is done.
Tradeoff: it feels bureaucratic in the moment. It is cheaper than being confidently wrong.
Here are explicit stop the line triggers that should force caveats or a pause, each tied to what you would see on the dashboard.
Trigger one: definition change in the metric contract. Dashboard signal: step change or discontinuity in the trend. Action: annotate, reset baselines, and present pre and post separately.
Trigger two: workflow or routing change that affects timestamps or queue entry. Dashboard signal: sudden improvement in FRT without a matching change in staffing or volume. Action: run the human only comparison check.
Trigger three: automation rollout without a comparison group or holdout thinking. Dashboard signal: variance drops and speed improves while reopens or complaints rise. Action: separate automation assisted flows in reporting until validated.
Common mistake number two: teams add caveats in the footnotes after the deck is approved. The correct move is to change the decision rule, not just add a disclaimer.
Triangulation: what to pair with CSAT, speed, deflection, and backlog
Triangulation is how you avoid being fooled by a single compromised metric. Pair one “efficiency” metric with one “outcome” metric.
For speed metrics like first response time and resolution time, pair with reopen rate or escalation rate. If FRT improves but escalations rise, do not cut staffing. That divergence says you are answering faster but solving worse, or pushing complexity upward.
For deflection, pair with inbound contacts and complaint volume, ideally topic specific. This is the heart of how to validate deflection metric claims. If deflection rises but inbound is unchanged, treat deflection as an engagement metric, not avoided work, and do not use it to justify coverage reductions.
For backlog, pair total backlog with aging percentiles. A smaller backlog that has a worse long tail is not a win. It is a risk.
Two concrete triangulation examples tied to real decisions:
Staffing decision: you want to reduce weekend coverage because FRT improved.
Pairing: FRT plus escalation rate and reopen rate.
Divergence that changes the decision: FRT improves but escalations or reopens rise. Decision: keep staffing steady and investigate closure quality.
Automation expansion decision: you want to expand bot containment because deflection improved.
Pairing: deflection plus inbound contacts for the same topics and complaint themes.
Divergence that changes the decision: deflection up but topic contacts flat or complaints up. Decision: slow expansion and fix contact paths and bot handoff.
Ongoing monitoring: alerts for drift, gaming, and automation side-effects
Pre mortems are not a one time cleanse. Metrics decay because the business changes.
Set a light cadence.
Weekly ops review: metric owner calls out any workflow changes and watches for red flag shapes.
Monthly decision review: skeptical reviewer brings one “here is how this could be wrong” item, even when things look good.
Quarterly contract review: decision owner signs off on whether the metric is still fit for the decisions it is being used to justify.
Practical tip: assign ownership for the metric contract the same way you assign ownership for an on call rotation. If nobody owns it, drift wins by default.
Make it a ritual: a lightweight agenda, roles, and artifacts you can reuse every month
Rituals fail for one of two reasons. They are too heavy, so people skip them when things get busy. Or they are too performative, so people do them after the decisions are already made.
You can avoid both with a short agenda, clear roles, and reusable artifacts.
The recurring agenda (15–30 minutes) and who owns each part
Open with the decision context, not the charts.
A reusable agenda template that fits in 20 to 30 minutes:
Two minutes: facilitator states the decision and the metrics that will drive it.
Three minutes: metric owner reads the metric contract changes since last review.
Five minutes: silent individual write. “Assume this metric misleads us next month. Why?”
Seven minutes: group cluster and name the top failure modes.
Five minutes: pick the top two tests and assign owners and due dates.
Three minutes: decision owner states the decision rule for the upcoming meeting. Trust, caveat, or pause.
Light humor helps here. I tell teams: “We are not here to roast the dashboard. We are here to stop it from roasting us in front of the CFO.”
Artifacts: the metric contract, pre-mortem log, and caveat slide
Keep it to one page each.
First, the metric contract, updated when definitions or workflows change.
Second, a pre mortem log: date, decision, top risks, tests run, and what you learned. This becomes institutional memory.
Third, a caveat slide you can drop into leadership decks.
Here is a concrete example of how to present caveats in a leadership meeting, without sounding defensive.
Deflection is up 12% month over month. Confidence: caveated.
Why: help center tracking changed on July 3 and bot handoff was updated on July 5.
Cross checks: topic level contacts are flat and complaints about “cannot reach a person” increased.
Decision rule: do not reduce staffing based on deflection this month. We will re evaluate after the two validation checks complete next week.
Practical tip: store the diagnostic signal list inside the metric contract doc, pinned near the top, or appended to your monthly review template. If it lives in someone’s notes, it dies there.
How to keep the ritual from becoming performative
The most common anti pattern is doing the pre mortem after the deck is final. At that point it is theater, not risk control.
The fix is simple: schedule the pre mortem 24 hours before the decision meeting, and require that the deck includes either “no caveats” or “caveats and tests” as a standard slide. If it is missing, the deck is not done.
Now a concrete Monday plan you can actually execute.
First action: pick one high stakes metric you will use this month, and write a one page metric contract for it.
Then focus on three priorities.
Identify the top two failure modes that would change the decision, not every possible flaw.
Run two fast validation checks, one segmentation breakout and one cross metric divergence check.
Agree on a decision rule with your decision owner: what would make you trust, caveat, or pause.
Set a realistic production bar: you are done when leadership can repeat your caveat in one sentence and the metric owner has two named tests with owners and due dates. Anything beyond that is optional polish, not the core work.
If you want one extra push, run the pre mortem before the next QBR and bring one caveated insight to the meeting. It is the easiest way to build trust as the person who keeps the org from making expensive decisions on “right looking” numbers.
Sources
- github.com — github.com
- atticusli.com — atticusli.com

