The Pre Mortem for Metrics: How to Find the Failure Modes Before You Trust a Dashboard

A practical metrics pre mortem for support dashboards: map the metric supply chain, hunt coverage gaps, prevent definition drift, and set decision rules so leaders know when a dashboard is safe to use for decisions.

Lucía Ferrer
Lucía Ferrer
12 min read·

Support dashboards don’t usually fail because someone “lied.” They fail because the number answered a different question than the one leadership acted on.

A metrics pre mortem is a short, structured way to surface those mismatch risks before the metric ships decisions. You assume it’s 90 days later, the metric drove a bad call, and now you work backward: how did this dashboard mislead us? Treat the chart like an incident candidate.

If that feels dramatic, good. A dashboard that influences staffing, SLA promises, or performance management is production infrastructure. It deserves the same skepticism you’d bring to anything that can wake you up at 2 a.m.

Treat dashboards like incident candidates (and define what ‘trust’ means for this decision)

Everyone has lived some version of this: the dashboard says support is “getting faster,” leadership celebrates, and two weeks later you’re drowning in angry escalations and overtime. The metric wasn’t malicious. It was just… unqualified.

Here’s the repeat offender: Average Handle Time drops, someone concludes “we’re more efficient,” and staffing gets reduced. Meanwhile the drop came from routing more complex work to a different queue, agents using macros that close tickets faster, or a policy change that quietly reshaped what “done” means. Customers feel the pain. The dashboard looks fantastic.

So define trust as a decision standard, not a vibe.

  • A monitoring metric is a smoke alarm. It’s allowed to be noisy. It’s used for “something changed; go look.”
  • A decision-grade metric is a steering wheel. It must be stable enough to justify actions with real blast radius (headcount, SLA changes, compensation plans, major process redesign).

This is where teams get burned: they treat a monitoring metric like a decision-grade metric because it’s already on a dashboard and the line goes up and to the right. That’s not governance; that’s astrology with extra steps.

Use this framing to start your pre mortem:

“It’s 90 days later and we made the wrong call. How did this metric mislead us?”

If you can’t name five plausible answers in five minutes, you’re either not being imaginative enough—or your org hasn’t been burned yet. (Give it time.)

The usual first cracks in support dashboards:

  • Missing conversations: the dashboard never saw a chunk of demand.
  • Shifting definitions: what counts as “first response,” “resolved,” or “reopened” changed under you.
  • Mix shifts: the work got easier/harder, but the metric tells a performance story.

The goal isn’t perfection. The goal is a clear, written boundary: what this metric is safe to influence, and what it is not.

Map the metric supply chain: from customer conversation → fields → transforms → chart

Most dashboard arguments are supply chain arguments. Not the warehouse kind—the “how did a customer conversation become a number on a chart” kind.

A metric is downstream of routing rules, tags, automation, reopen logic, business-hour settings, time-window choices, and data delays. If you only look at the chart, you’re debating the final painting while ignoring the plumbing behind the wall.

Start with the decision and work backward.

Pick one metric people actually act on (first response time, time to resolution, backlog age, contact rate). Then force this question:

“If this number moves by ~10%, what would we change?”

If the answer is “hire,” “cut coverage,” “rewrite schedules,” “change SLAs,” or “declare the program worked,” you’re in decision-grade territory.

A practical move that saves meetings: write the action down before you argue about the number. If you can’t name the action, treat the metric as monitoring-grade until someone can.

Next: list every input that can change the metric without changing reality.

Support metrics get sneaky because workflow changes move numbers without improving customer outcomes. A few concrete ways this happens:

  • Auto-replies that count as responses. You add an immediate acknowledgement (“We got your request”). First response time improves overnight, even if humans arrive later.
  • Reopen policy changes. You redefine reopenings so reopened work counts as a new ticket. Time to resolution “improves” because you stopped measuring the long tail.
  • Queue and routing reshuffles. Complex cases get moved out of the measured queue. The measured queue gets faster; customers don’t.

Common mistake: teams treat routing, tagging, automation, and reopen changes as “just process.” In a metrics pre mortem, assume every one of those changes is a metrics change until proven otherwise.

Finally: capture a metric contract snapshot so definition drift becomes obvious.

You don’t need a 40-page governance binder. You need a one-page snapshot that makes drift visible and forces ownership. Keep it tight:

  • Name + plain-language meaning (what a human thinks it means)
  • Business question + decision allowed (and what it must not be used for)
  • Event definition (what counts as “first response,” “resolved,” “backlog,” “reopened”)
  • Clock rules (start/stop, business hours vs. calendar, pauses)
  • Inclusions/exclusions (channels, queues, regions, segments; spam/duplicates/merges/internal notes)
  • Required segments (the slices that must always be shown with the headline)
  • Sensitive dependencies (routing rules, automation policies, tagging requirements)
  • Ownership split: owner of meaning vs. owners of levers (who can change routing/automation)
  • Change triggers + review cadence (what forces an update; last reviewed / next review)

That ownership split matters more than most teams admit. The people who can change routing and automation often aren’t the people who get yelled at when the board deck is wrong.

A solid outside framing on why this is career-relevant (not theoretical) is here: [1]

Run a coverage gap hunt: find the conversations your dashboard never sees

Dashboards are biased by default because they only measure what they can see. Support leaders often assume the dashboard equals “all customer pain.” It rarely does.

A coverage gap is any set of customer conversations that happen outside your measurement boundary. In support, coverage gaps distort volume, backlog, SLA attainment, and even CSAT. If you miss an entire channel, your “improvement” might just be customers going somewhere else.

Start with a blunt conversation inventory.

Where can a customer ask for help today?

Email and chat are obvious. Then list what teams forget or minimize: phone calls logged in another tool, app store reviews that trigger outreach, social DMs, community posts, in-product feedback, partner escalations, and internal escalation paths.

Concrete cross-channel scenario that breaks dashboards: a customer starts in chat, gets told to email because it’s “complex,” then calls because email is slow. If your reporting counts chat and email but not phone, your dashboard says backlog is under control while true workload is leaking out the side.

If you want the fastest route to the truth, ask frontline agents: “What are the top three places you do off-system work?” You’ll get the answer in thirty seconds, and it will be more accurate than your assumptions.

Coverage tests that don’t require perfect instrumentation.

You don’t need a months-long tooling project to detect bias. You need simple tests that create pressure on the system.

  1. Weekly reconciliation between the dashboard and a lightweight “reality log.” The log can be a phone tally, a short sampling sheet, or a record of escalations. You’re not trying to match every record. You’re trying to catch drift early.

  2. Random customer journey sampling. Pick ~20 recent customers, trace their help-seeking path across channels, and check whether each touch appears in reporting. This finds silent channels faster than any architecture review.

To put a number on it without getting fancy, use a discrepancy ratio:

Discrepancy ratio = conversations recorded ÷ conversations observed

“Observed” comes from your sampling or reconciliation source. If the ratio is 0.8 for a segment, your dashboard is missing ~20% of conversations for that slice. That’s enough to blow up staffing models and make SLA attainment look healthier than it is.

If you need a reminder that dashboards can look fine while the data is broken, this piece captures the “wrong without an error message” problem: [2]

Common missing-conversation failure modes worth naming in the pre mortem:

  • Phone or social work not counted (classic)
  • Split escalations: frontline ticket is measured, specialist work isn’t (resolution looks great; customers wait days)
  • Duplicates and merges mishandled (volume trends become workflow artifacts)

Common mistake: leaders rely on a single “total volume” number as if it represents all demand. Insist on a channel view plus an escalation view—even if the escalation view is initially ugly. Ugly-but-honest beats clean-and-wrong.

Lock definitions + decision rules to prevent drift, cherry-picking, and metric theater

Support dashboards rarely fail with a dramatic explosion. They fail in slow motion.

A definition shifts. A slice disappears from the default view. Teams learn which behaviors make the number look good. Soon the metric is more politics than measurement.

Your metrics pre mortem should end with three protections: locked definitions, forced transparency on slices, and decision rules that prevent the metric from being used as a weapon.

Definition drift patterns to watch.

  • The clock. Does first response time start when the customer sends the first message, or when the ticket is created? Does it pause outside business hours? Do automated acknowledgements count? Any of these can be valid—none are safe if they change silently.
  • Exclusions. You exclude spam, then exclude “low value plans,” then exclude a noisy queue, and now the metric is technically accurate and strategically misleading.
  • Resolved semantics. “Resolved” as customer-confirmed vs. agent-closed vs. moved-to-escalation are three different worlds.

When definitions must change because workflows change, keep both versions visible for a transition period. Otherwise you create a false narrative (“we improved!”) that’s really just a measurement cutover.

Cherry-picking slices: how it happens.

Cherry-picking rarely looks like fraud. It looks like someone trying to be helpful.

A support leader presents first response time excluding weekends because “weekends are weird.” Another leader shows only a specific queue because “that’s the main queue.” Someone filters out a region because “holidays skew things.” All of those might be defensible—until the filters become invisible and the baseline disappears.

The fix is simple and slightly annoying (which is why it works): require a baseline view.

Define standard slices that must always appear with the headline metric: channel, queue, region, language, entitlement, severity. Then require any filtered view to sit next to the unfiltered baseline. Not as a gotcha. As transparency.

Decision rules: what actions are allowed, and what actions are banned.

A metric becomes dangerous when it can trigger big actions without corroboration. Write decision rules in plain language and keep them close to the dashboard.

Allowed actions:

  • Investigate routing, staffing, and channel mix if the metric moves materially for two consecutive weeks.
  • Shift intraday schedules if the change is isolated to time-of-day slices.
  • Trigger targeted quality review if faster responses correlate with lower CSAT or higher reopen rate.

Banned actions:

  • Don’t change headcount based on a one-week dip/spike without confirming coverage and mix.
  • Don’t declare a program successful using only a single headline metric.
  • Don’t reset SLA targets until you’ve had at least one full cycle of stable definitions and stable channel coverage.

The operating rule that survives real life: change definitions when you must, but never change them silently—and never change them mid-narrative.

For a quick catalog of metric anti-patterns that lead to theater, keep this handy: [3]

Pre-mortem the failure modes: automation blind spots, shifting mix, and gaming (then assign monitors)

Assignment strategy Best for Advantages Risks Recommended when
Pre-mortem workshop with

This is the section where you assume the metric will betray you and list exactly how. Support metrics break in predictable ways: automation creates phantom wins, mix shifts look like performance, and incentives turn humans into clever rule-followers. None of this requires bad people. It just requires people.

Automation blind spots to call out explicitly.

A worked example that shows up everywhere:

You add an automation that replies instantly to new chats with a bot message: “Hi, I can help. Pick a category.” The dashboard counts that as first response. First response time drops from 45 minutes to 2 minutes. Leadership celebrates.

But the bot fails on edge cases, so customers loop for 10 minutes, then get handed to a human who’s now busier than before. True time to human help increases. Customers feel ignored in a new and exciting way.

Pre mortem question: what will make the dashboard celebrate while customers complain? “Bot replies counted as responses” is usually top-three.

Mix shifts that look like performance.

Support performance changes when the work changes.

  • You push low-severity requests into self-serve. Remaining tickets are harder. Time to resolution rises; that’s not necessarily failure.
  • You launch in a new region. Language coverage changes. Response times shift by time of day.
  • You change entitlement rules. Premium customers get faster help while everyone else waits.

If you don’t segment by mix drivers, the metric mostly tells you who showed up, not how well you performed.

A practical guardrail: pick two sentinel slices that always stay on screen with the headline number. For support, severity + channel works well, or entitlement + language. If the headline moves but the sentinel slices explain it, you avoid a lot of panic and a lot of bad “fixes.”

Gaming and Goodhart’s Law in support metrics.

If you pay attention to a metric, humans will optimize it. Attach it to performance reviews and they’ll really optimize it.

This is where teams get burned: fast-but-low-value responses, premature resolutions, a spike in reopenings, agents avoiding complex tickets because they hurt averages. The number improves; customers don’t.

Guardrails don’t have to be punitive. Pair the metric with a counter-metric and a sampling habit:

  • Pair first response time with a small weekly sample of “time to meaningful first response.”
  • Pair resolution speed with reopen rate.

Monitoring plan: drift triggers, anomaly checks, audit cadence.

A pre mortem isn’t a one-time workshop. It’s a maintenance posture.

Set a cadence that matches the risk. For decision-grade support metrics, monthly review is a good default. For anything used in staffing or SLA promises, add a lightweight weekly integrity check.

Define triggers that force a re-run: routing changes, automation policy changes, new channels, staffing model changes, SLA definition changes, reopen rule changes. If you ship workflow changes weekly, congratulations—you also ship metric risk weekly.

If you want a memorable example of how often teams mistake instrument problems for business problems, this is worth your time: [4]

Here’s a simple assignment matrix you can use to decide who owns what in the pre mortem and its follow-through.

Run the pre-mortem with leadership: a 30-minute agenda, outputs, and a ‘trust boundary’ statement

Leadership doesn’t need you to be cynical about dashboards. They need you to be specific about what’s safe to decide today.

Bring the metrics pre mortem into the dashboard review, keep it tight, and make the output operational. It reduces decision risk without turning the meeting into a philosophy seminar.

A tight 30-minute agenda.

  • Minutes 0–5: Confirm the decision at stake (staffing, SLA, quality, or monitoring).
  • Minutes 5–12: Review the metric contract snapshot (definition, inclusions/exclusions, clock rules, owners).
  • Minutes 12–20: Brainstorm failure modes using “90 days later we were wrong.” Capture the top five.
  • Minutes 20–26: Agree on guardrails + decision rules for the next cycle.
  • Minutes 26–30: Assign owners and dates for the two highest-leverage integrity tests.

Defer platform debates. This meeting is about trust boundaries and risk controls, not a tooling migration.

Outputs you should leave with.

  • A one-page metric contract snapshot.
  • A short failure-mode register: top ways the metric misleads + the quickest test for each.
  • Named owners for monitoring, plus cadence.

When ownership is vague, integrity work becomes optional. Optional integrity work is how you end up staffing off a chart that quietly stopped counting an entire channel.

Put the trust boundary in writing.

“Today, we will use [dashboard name] to monitor [metric] for [purpose] on a [cadence] basis. We will not use it for [high-stakes decision] until [coverage test] and [definition check] pass, and until we review [sentinel slices] alongside the headline number.”

Keep next week painfully concrete.

Pick one metric that currently drives a real decision—usually staffing tied to first response time or backlog. Then do three things:

Map the supply chain. Run one coverage-gap test. Publish the trust boundary statement.

Set a realistic bar: monitoring-grade by end of week, decision-grade by next monthly review. And no staffing changes off a single trend line until coverage and definition are stable. That’s the whole point: the dashboard should be a tool for decisions, not the reason you have to undo them later.

Sources

  1. calypso.ms — calypso.ms
  2. productquant.dev — productquant.dev
  3. kpitree.co — kpitree.co
  4. thicket.sh — thicket.sh