How to Run a Pre Mortem on Your Metrics Before They Run Your Strategy

A metrics pre mortem is a short, structured meeting where you assume a KPI target went wrong and work backward to find definition drift, automation inflation, and gaming risks before OKRs lock it in.

Lucía Ferrer
Lucía Ferrer
17 min read·

Spot the moment your KPI stops measuring and starts steering

You rarely notice when a KPI graduates from “measurement” to “management.” It happens quietly, like a thermostat that starts negotiating with you about what “comfortable” means.

Support teams see this faster than most. One quarter you add Average Handle Time (AHT) to a scorecard “just to understand where time goes.” Next quarter someone asks why AHT is up. Staffing shifts. Queue policies tighten. Suddenly the team is optimizing for speed.

Then the predictable horror story lands: AHT improves, backlog looks healthier, and CSAT still slips because reopens and escalations rise. You didn’t set out to build a factory line. The metric did it for you.

That’s why a metrics pre mortem is worth your calendar. It’s a short, structured meeting where you pretend it’s a few months from now and the KPI target caused real damage. Then you ask, calmly and specifically, “How did we get here?”

This isn’t “be negative.” It’s basic risk management for incentives. Once a number becomes a target, people will find the most literal path to the target. That’s not a character flaw. It’s just how systems work under pressure.

In support, “innocent” KPIs like CSAT, AHT, deflection, FCR, backlog, and escalation rate can quickly become staffing policy, channel strategy, and QA standards—without anyone ever calling it strategy.

The hidden hand: when scorecards become staffing and channel strategy

The steering starts when the scorecard becomes the language of tradeoffs:

“Deflection up” becomes “push customers to the help center first.”

“Backlog down” becomes “close anything without a reply in 24 hours.”

“Escalations down” becomes “stop escalating, just tell them no.”

Those are strategic choices wearing a measurement costume.

A useful framing: treat any KPI you plan to put into OKRs like a product launch. If it can change behavior, it deserves a risk review.

A fast self-check: signals your metrics are already running the show

A few tells show up early:

People ask “what does this do to the metric?” more often than “what does this do to the customer?”

Teams build local workarounds: “Just mark it solved and they can reopen,” or “Route that to email so it won’t count against chat.” When you hear those, you’re not dealing with a morals problem. You’re dealing with measurement design.

Definitions drift. The number stays stable, but what it represents shifts as routing, tagging, and automation evolve.

What a pre mortem produces: decisions, not dashboards

A metrics pre mortem should end with three usable artifacts:

  1. A decision: OKR-ready, pilot, or diagnostic-only.

  2. A definition sheet that locks what the KPI means right now (including the annoying edge cases).

  3. A short risk-and-guardrails plan so you prevent KPI gaming in support instead of discovering it in a painful postmortem.

If you want the basic pre-mortem structure, these are solid primers: Atlassian pre mortem play and Asana premortem. The move is applying that structure to metrics and incentives—not projects.

Run the pre mortem meeting: timebox, roles, and the artifacts you need on the table

The most common failure is trying to do this solo in a spreadsheet. The risk lives in the gaps between teams: ops, QA, frontline reality, channel tooling, and whoever quietly owns routing rules and bots.

A support-focused metrics pre mortem meeting should feel like a structured argument with receipts. Healthy conflict is the feature. If everyone agrees immediately, you probably missed the failure mode that will burn you later.

The mindset Kathryn Fulton describes—let the plan meet resistance before reality forces it to—is exactly what you want here: The secret to smarter strategy: pretend it already failed.

Invite list by failure mode

Invite based on how the KPI can fail, not based on org-chart politeness.

Bring:

Support ops (reporting, queue health, tooling quirks).

QA or a team lead (what “good resolution” looks like in real conversations).

Two frontline agents (the people who can describe the Friday-at-4:55 workarounds without blinking).

Channel owners (chat vs email vs phone is never apples-to-apples).

Whoever owns automation rules and bot behavior (automation is the fastest way to make a metric look better than reality).

This is where teams get burned: only inviting managers. Managers talk policy. Agents talk actual behavior. You need both.

Prework pack: light, real, and shared

Keep prework small but concrete. You’re not aiming for analytics perfection. You’re aiming for a shared reality.

Bring 30 days of the KPI trend plus 2–3 “likely movers.”

If the KPI is AHT: bring reopens, escalations, CSAT.

If it’s deflection: bring contact rate, unresolved follow-ups, CSAT.

Bring a small audit set: 20 real interactions across channels (e.g., 6 email, 6 chat, 4 phone, 4 social/community) across easy and hard cases. Not to judge agents—this is to test whether your metric logic survives real work.

Bring the “rules of the game”: tagging taxonomy, macros that auto-apply tags, auto-closure timers, bot routing rules, and channel mapping notes (like “chat converted to email becomes a new ticket”). That’s where definition drift hides.

Quick test: if you can’t find the current metric definition in under five minutes, your first pre mortem finding is “we don’t have a definition.”

Agenda: “Assume the KPI ruined our quarter. How did it happen?”

Two timeboxes that actually work:

60–90 minutes for one KPI you already track and want to promote into OKRs.

2 hours for cross-channel or automation-heavy KPIs (deflection and FCR love complexity).

Run it in four beats:

  1. Intent: “What decision will this KPI drive if we hit or miss it?” If nobody can answer, keep it diagnostic.

  2. Assume failure: “It’s next quarter. We hit the number and customers are worse off. What did we do to make that happen?”

  3. Assume confusion: “The number moved, but we can’t tell if we improved.” (This is where instrumentation drift and attribution bugs show up.)

  4. Guardrails: “If we chase this number, what must also stay healthy?”

Outputs: one pager, risk log, OKR readiness

Leave with:

A metric one-pager: purpose, exact definition, data sources, exclusions, owner.

A risk log: plausible failure modes with an early detection signal.

A decision: pass, pilot, or park.

Decision rule: if you can’t write the definition and top three failure modes in plain language, the KPI isn’t ready to steer incentives.

For extra “assume the metric is guilty until proven innocent” energy, this is a strong companion: Your metric has a design flaw.

Lock definitions before you lock targets: prevent definition and instrument drift

Support leaders are often surprised by how fast a metric stops meaning what they think it means. That’s not incompetence. That’s entropy.

Two drifts matter most:

Definition drift: the words change meaning. “Resolution” becomes “agent clicked solved.” “Escalation” becomes “tagged as escalation,” which is not the same as “moved to tier two.”

Instrument drift: the measuring mechanism changes. Routing updates, channel migrations, bot flows, auto-closure rules, taxonomy refactors—your trend line becomes a story about tooling, not performance.

Never lock a target until you can defend the nouns.

Define the nouns (so edge cases don’t define them for you)

Write definitions that survive real life. These templates work because they force decisions.

Ticket: a customer-initiated request that enters a staffed queue and requires agent/specialist action. Includes new contacts created by channel conversion if the customer has to restate the issue. Excludes spam, duplicate system alerts, internal tasks unless they block resolution.

Escalation: a deliberate transfer of ownership from frontline to a specialist queue with different service expectations. Counted at ownership transfer—not when a tag is applied. Courtesy consults don’t count. Reassignments within the same tier don’t count.

Resolution: the customer’s stated issue is addressed with a clear outcome and next step, and either confirmed by customer response or passes a defined cooling period without return contact on the same issue.

FCR: resolved within the first staffed interaction for that issue, with no follow-up within a defined window (e.g., 7 days) across all channels.

This is where teams get burned: defining metrics in tool-language instead of customer-language. “Solved” is a field. “Resolved” is an outcome. Define the outcome first, then map fields to it.

Channel and branch attribution: one journey, multiple metric events

A customer journey can touch chat, email, and phone in the same afternoon. If you don’t decide how those transitions count, your KPI will quietly punish the wrong teams.

Classic failure: a customer starts in chat, agent requests logs, customer emails logs, the system creates a new ticket. One issue just became two “first contacts.” FCR drops, volume rises, and everyone argues about performance when the real problem is accounting.

Branch/region scorecards create another trap. If one branch routes complex cases to a shared queue, their AHT looks worse and their escalation rate looks better (or vice versa) depending on how you count. You end up rewarding routing politics, not support quality.

Write the mapping in plain language. You don’t need perfection. You need consistency.

Edge cases that break comparability

Decide on the repeat offenders and put them in an appendix on the one-pager:

Merges: does the parent inherit time and first response? If you don’t decide, AHT will jump when agents clean duplicates.

Splits: one ticket becomes two issues. If you count one resolution, you inflate performance. If you count two without time allocation rules, you punish thorough work.

Transfers: who owns time and outcome when an issue bounces between queues? If it’s “everyone,” it becomes “no one.”

Duplicate contacts: anxious customers create volume. That’s why you’ll need paired metrics and audits—otherwise “volume down” becomes “make it harder to reach us.”

Decision rules: OKR-grade vs diagnostic-only

A metric is OKR-grade only if it passes four tests:

Stability: definition/instrumentation is unlikely to change this quarter, or you have a change log and re-baselining plan.

Interpretability: a leader can understand what a move means without three caveats.

Controllability: the team can influence it through legitimate improvement, not just workarounds.

Abuse resistance: you can name at least three ways it can be gamed—and you have guardrails to detect it early.

Fail any test? Keep it visible, keep it useful, but don’t attach incentives.

Assume you ‘won’ the metric: find the automation and incentive paths that make performance look better than reality

If you want to prevent KPI gaming in support, stop relying on “good judgment” as your control plan. People use judgment until the system puts them in a corner. Then they do what the system rewards.

A metrics pre mortem works because it gives the room permission to be creatively pessimistic on purpose: “If we had to make this number look great by Friday, what would we do?” The loopholes appear fast.

Automation vs human judgment: where the numbers inflate

Automation is a gift and a liar. It speeds work up, and it can quietly rewrite what your KPI counts.

Auto-closure reduces backlog but increases reopens.

Add a rule: close tickets after 48 hours of no customer reply. Backlog drops. Leadership celebrates. Two weeks later reopens spike because customers reply late or didn’t understand the last message. The “improvement” was just moving work into the future—like sweeping dust under the rug and calling it minimalist design.

Auto-tagging inflates FCR.

If a bot or macro applies “resolved” when a knowledge-base link is sent, you may count “attempted deflection” as “resolved outcome.” FCR goes up. CSAT goes down. Your dashboard looks heroic while customers feel dismissed.

Bot routing changes the case mix.

A bot pushes easy issues to self-serve. Deflection improves. The customers who reach agents are now the hardest cases. AHT rises, CSAT gets more volatile. If leaders only see deflection, they’ll cut staffing at exactly the wrong time.

That’s why “who owns automation rules” is never optional.

May Mor’s framing is useful here: treat KPIs like something you ship and then monitor, because they change behavior like product changes do: Ship KPIs like features.

How deflection gets gamed (and how it accidentally lies)

Deflection is vulnerable because it’s partly counterfactual: you’re measuring something that didn’t happen.

It gets gamed when “handled” means “a bot sent a message,” or when “resolved” means “the customer didn’t contact support again” without checking whether they churned, complained elsewhere, or just gave up.

The most common “deflection win” that backfires: adding friction to contact options. The help center becomes a maze. Contact rate drops. Deflection looks fantastic. Meanwhile social complaints rise and your account team starts fielding angry emails. The metric didn’t lie. It measured the wrong success condition.

If you use deflection as a KPI target, require one qualitative signal alongside it—small weekly samples of “did this self-serve path actually solve the issue?” Otherwise you’ll optimize for silence.

AHT and backlog: speed traps that create reopens, escalations, and quiet churn

AHT is useful, but it’s a magnet for bad behavior once it becomes a target.

The easy way to improve AHT is to shorten conversations. The hard way is to reduce complexity through tooling, training, and product fixes. If you only reward the number, the easy way wins.

This creates the “fast but wrong” failure mode: agents rush, customers reopen, escalations rise because frontline didn’t take two extra minutes to gather context, and CSAT declines. Worse, some customers don’t reopen. They churn quietly. Your dashboard looks clean. Revenue doesn’t.

Backlog has a similar trap. When backlog is a target, people close or defer work to protect the number. That’s not laziness. That’s rational survival.

Pre mortem prompts: the Support KPI Dirty Dozen

Use these questions in the meeting. They surface most failure modes without turning your org into a measurement bureaucracy.

  1. If we hit this KPI by changing labels/statuses, would the customer experience improve?

  2. What’s the simplest workaround an agent could use to make this number look good today?

  3. What automation rule could inflate this metric without anyone noticing?

  4. What would a smart, stressed person do if their bonus depended on this KPI?

  5. If we were gaming it, what other metric would worsen?

  6. Where can channel switching create double-counting or missed counting?

  7. What happens to this KPI when volume spikes and we go into survival mode?

  8. What happens when routing/queue ownership/bot flows change mid-quarter?

  9. Does this KPI punish teams handling complex cases (enterprise, regulated, high-severity)?

  10. Can a new manager understand this metric in two sentences without being misled?

  11. If the number moves, what’s the first “customer story” we’ll tell ourselves that might be wrong?

  12. What would make us retire this as an OKR target?

The meeting only pays off if you connect the answers to guardrails. Otherwise it’s just a spooky campfire story.

Install guardrails and cadence: paired metrics, red-flag thresholds, and audit samples that keep KPIs honest

Assignment strategy Best for Advantages Risks Recommended when
Red-Flag Thresholds (e.g., Churn Rate > 5% for 3 consecutive days) Early warning of critical performance degradation or anomalies. Clear, actionable triggers for investigation. prevents minor issues from escalating. False positives if thresholds are too sensitive. alert fatigue. KPIs are critical to business health and have predictable ranges.
Instrumentation Change Log (e.g., tracking changes to event tags) Understanding impact of data collection changes on historical trends. Explains unexpected metric shifts. aids in debugging and data validation. Requires diligent documentation. can be overlooked in fast-paced environments. Any KPI relying on tracked user behavior or system events.
Paired Metrics (e.g., Conversion Rate + Avg Order Value) Detecting gaming or unintended consequences of a single KPI. Provides a more holistic view. harder to game both metrics simultaneously. Can add complexity. correlation doesn't always imply causation. KPIs have direct, measurable counter-metrics or balancing metrics.
Definition Lock & Versioning — e.g., 'Active User' definition v2.1 in data dictionary Preventing metric drift and ensuring consistent interpretation across teams. Clear communication. reduces debate over what a metric means. Can be rigid. slow to adapt to evolving product or business models. Any shared KPI used by multiple teams or for strategic decisions.
Manual Audit Samples (e.g., 5% of new user sign-ups reviewed weekly) Catching data quality issues, definition drift, or gaming not visible in aggregates. Uncovers qualitative insights. ensures data integrity at the source. Resource-intensive. small sample size might miss issues. High-impact KPIs, new data sources, or suspected data manipulation.
Operator Reality Check (e.g., weekly check-in with front-line staff) Identifying discrepancies between reported metrics and ground truth. Uncovers gaming or unintended behaviors not captured by data alone. Qualitative and anecdotal. can be dismissed without quantitative backing. KPIs directly reflect user experience or operational processes.

A pre mortem is only valuable if it changes how you operate after the KPI goes live.

Guardrails do two jobs:

They stop you from celebrating the wrong win.

They give frontline teams psychological safety. When people know you’ll look at paired metrics and audits, they feel less pressure to “perform for the dashboard.”

Paired metric guardrails (keep one KPI from being “won” the wrong way)

Pairs aren’t for metric hoarding. They’re constraints.

CSAT: pair with reopens or escalation rate. If CSAT rises while escalations rise, you may be surveying only easy cases.

AHT: pair with reopens plus a small QA pass-rate sample. If AHT drops and reopens climb, you got faster at being incomplete.

FCR: pair with “repeat contacts per issue” across channels. FCR alone gets fooled by splitting and channel conversion.

Deflection: pair with contact rate by issue type plus a small sample of self-serve outcomes. If deflection improves while unresolved follow-ups rise, you’re building dead ends.

Backlog: pair count with aging (median and tail). “500 at 1 day” is a different world than “200 with a 14-day tail.”

Escalations: pair with severity mix and segment. Escalations can drop because frontline improved—or because they stopped escalating hard cases.

Red flags you pre-commit to investigate

Red flags should be boring, explicit, and agreed in advance. Starting points many support orgs can adapt:

If AHT improves more than ~15% week-over-week, validate with QA sample and reopens before you celebrate.

If backlog drops more than ~20% in a week without clear staffing/volume changes, check for auto-closure or status hygiene shifts.

If deflection improves while CSAT drops ~0.3 in the same period, review self-serve paths for dead ends.

If FCR jumps right after a tagging/macro change, assume instrument drift until proven otherwise.

The thresholds don’t need to be perfect. They need to exist—because “too good to be true” is a recurring production bug.

Audit sampling: small, consistent, and focused on the guardrails

You don’t need a dedicated analytics squad. You need a sampling habit you’ll keep.

A feasible default: 10 tickets per week per major channel, plus 5 from whichever queue moved the most. Split review between a QA lead and a rotating team lead so it doesn’t become one person’s unpaid second job.

If that’s heavy, start with 5 per channel per week. Consistency beats ambition.

Review what matches your risk:

AHT: completeness and actual resolution, not just speed.

Deflection: whether customers returned through another channel.

Backlog: whether closures were legitimate or “dashboard cleaning.”

Important warning: don’t turn audits into a witch hunt. Score the system, not the agent. If audits create blame, people will hide edge cases—and your data quality will get worse.

Change log discipline (the unsexy thing that saves executive meetings)

Keep a lightweight change log for anything that can shift measurement without real performance change: routing changes, bot releases, auto-closure timers, macro/tagging updates, taxonomy revisions, channel migrations.

Rule of thumb: if a change could move the KPI more than ~5% by itself, you either re-baseline or annotate the trend and temporarily suspend target comparisons. Otherwise you’ll spend exec time arguing about whether you “improved” when you really just shipped a new bot flow.

Close the loop: a two-week rollout to pilot the KPI, document the risks, and only then bake it into OKRs

Most KPI disasters come from skipping the boring middle. Leaders see a metric, set a target, and assume the organization will “figure it out.” The organization does figure it out—usually in the most literal way possible.

A simple two-week rollout is enough to catch definition drift, instrument drift, and incentive loopholes before OKRs turn the KPI into a religion.

Pilot plan: one branch or one channel, then expand

Week one: pilot in one branch/queue/channel where you can learn fast. For deflection, choose one issue category. For AHT, pick a team with stable volume.

Week two: expand to a second segment and validate whether your guardrails still hold.

Don’t pilot only in your best-behaved team. Pilot where the edge cases live. That’s where definitions break.

What to document (so the learning survives the meeting)

By end of week one: metric one-pager + risk register with the top failure modes you found.

By end of week two: guardrail triggers that were actually used at least once in review, plus two audit samples completed.

If you want a KPI-specific pre mortem companion, keep this nearby: Pre Mortem KPIs: understanding what could go wrong.

How to communicate (so people don’t optimize in silence)

Your message matters more than your dashboard.

A script that works in real teams:

“We’re adding AHT as a target because we believe we can remove friction for customers and agents. We’re not rewarding speed at the expense of quality. That’s why we will watch reopens, escalations, and a weekly QA sample alongside AHT. If you see pressure to close early or avoid hard cases, raise it. That’s a system problem, not a performance problem.”

That last line is the difference between honest reporting and quiet gaming.

Retirement clause: when a KPI stops being OKR-worthy

Give every KPI an exit ramp.

OKR readiness gate (after two weeks): OKR-ready only if (1) the definition sheet is locked with channel mapping and edge cases, (2) no red-flag thresholds triggered without an explainable cause, and (3) audit samples match the dashboard story at least ~80% of the time.

Retirement clause: if exec reviews require repeated caveats, or teams keep finding new gaming paths faster than you can add guardrails, retire it as an OKR target and move it back to diagnostic-only.

Your Monday plan

First action: put a 75-minute metrics pre mortem on the calendar for one KPI you’re tempted to target—usually AHT or deflection.

Three priorities for that meeting: (1) bring a 20-ticket audit set across channels, (2) lock definitions for at least ticket, resolution, and escalation, and (3) choose two paired metrics plus one red-flag threshold you will investigate every time.

Realistic production bar: by Friday, you should have a one-page definition sheet, a short risk log, and a weekly audit cadence that someone actually owns. If you cannot get those three, the KPI is not ready to run your strategy, and that is a perfectly acceptable answer.