The pre-meeting test: name one decision each metric can change (or cut it)
A familiar failure: the dashboard “wins” while the customer experience loses
You have seen this movie. The weekly support dashboard looks healthier than last week. Average handle time is down. SLA compliance is up. Someone says, “Nice, the process changes worked.” Then, twenty minutes later, an angry product leader is in your inbox because escalations spiked, backlog is older than it has been in months, and your most experienced agents are spending their afternoons untangling reopened tickets.
One of my favorite versions of this failure is when AHT drops and everyone celebrates, but reopen rate quietly climbs and the backlog ages. What really happened is rarely “the team got faster.” More often, the definition moved, the channel mix changed, or a macro encouraged premature closures. The dashboard still improved. Reality did not.
The decision change rule (and what counts as a real decision)
Here is the sanity check that fixes more support reporting than any new dashboard ever will: If a metric did not change the decision, you measured the wrong thing. Or, more precisely, you measured something that belongs in a diagnostic appendix, not in the decision seat.
A “real decision” is not “we should keep an eye on it.” A real decision has two options, an owner, a time horizon, and a consequence. Examples: do we staff weekend coverage next month or not, do we freeze deflection experiments this week or keep pushing them, do we reroute billing tickets to a specialist queue or spread them across pods.
What “decision grade” means for weekly ops vs QBR
To keep this practical, label every metric in your support dashboard sanity check as one of three types.
“Decision grade” means it can credibly flip an action in the next weekly ops cycle or in the next QBR planning window.
“Diagnostic only” means it helps explain what happened, but you would not steer with it without additional context or validation.
“Vanity” means it looks impressive, moves a lot, and changes nothing. It can be fun at parties, but it should not run your team.
The promise: you can do a 10 to 20 minute pre meeting workflow that produces a one page output for leaders. It prevents definition debates during the meeting and stops “support KPI sanity check” from becoming a monthly ritual of confused nodding.
Start from the decision, not the KPI: write the two options and the threshold that flips them
Turn KPIs into choices: hire vs reallocate, backlog burn vs quality push, deflection expand vs rollback
Most teams build dashboards like a scrapbook: everything gets pasted in, because it might be useful someday. Experienced support operators do the opposite. They start with the decisions they expect to make, then pull in only the metrics that can legitimately change those decisions.
Pick two or three decisions that actually matter this month. Staffing and capacity is usually one. Workflow and routing changes is another. Backlog strategy is a third. Now force each into a binary choice.
For staffing: do we hire contractors for six weeks, or do we reallocate senior agents away from project work?
For backlog: do we run a backlog burn week, or do we keep focusing on quality and coaching even if backlog grows?
For deflection: do we expand self service and bot coverage, or do we roll back because it is displacing complexity and increasing escalations?
A useful support metrics sanity check workflow feels almost annoyingly specific here. That is the point.
The “flip threshold”: what number would actually change your plan this week?
Now the part most teams skip, and the reason their support dashboard sanity check devolves into commentary. You need a flip threshold.
A flip threshold is the number that, if you saw it on Monday, would change what you do by Friday.
Example tied to action: “If backlog age over 7 days exceeds 15 percent of open tickets, we freeze deflection experiments for one week and staff backlog burn using two senior agents plus one rotating specialist.”
Another: “If escalation rate rises above 2.5 percent of solved tickets for two consecutive weeks, we stop optimizing AHT and run a quality push with targeted coaching and reopen analysis.”
Common mistake number one: setting thresholds that are emotionally comforting instead of operationally useful. If your threshold is “we will act when things are really bad,” congratulations, you just built a lagging indicator with a nice suit.
When a metric is only diagnostic (and how to label it without deleting it)
You do not have to delete diagnostic metrics. You just have to stop pretending they are steering wheels.
CSAT is a good example. It is an outcome metric, it lags, and it is vulnerable to sampling bias. It still belongs in your review, but often as “diagnostic only” unless you have stable survey coverage and a clear action you are willing to take.
A healthier pattern is to pair an outcome metric with a steerable process metric. For example, keep CSAT, but steer on reopen rate, escalation rate, and sampled quality checks. That gives you an earlier signal that your “wins” are not quietly hurting customers.
If you like the framing, Calypso’s “If you cannot explain the decision, do not ship the metric” idea is worth keeping in your team’s vocabulary: [1]
A fast score: decision impact × controllability × time to effect
When you have twenty candidate KPIs, you need a quick triage that does not become its own meeting.
Use a simple 1 to 3 score on three dimensions.
Impact: 1 means it rarely changes customer experience or cost, 3 means leaders will feel it within a month.
Controllability: 1 means it is mostly driven by product changes or seasonality, 3 means support can meaningfully move it.
Time to effect: 1 means it takes a quarter to respond, 3 means you can change it within one to two weeks.
Metrics that score high across all three are your decision grade core. Metrics that score low are either diagnostic only or vanity.
Here is a decision template you can paste into a doc. It is the heart of a decision driven support metrics practice, and it also makes QBR metrics sanity checks less theatrical.
- Decision
- Option A
- Option B
- Metric
- Flip threshold
- Owner
- Next action
Concrete support decision examples:
For routing: Decision is whether to reroute “billing refunds” tickets to a specialist queue. Option A is keep distributed assignment. Option B is specialist queue. Metric is escalation rate and backlog age in that category. Flip threshold is “if backlog age in refunds exceeds 5 days or escalations exceed 3 percent, we switch routing for two weeks.”
For staffing: Decision is whether to add weekend coverage. Option A is no weekend coverage. Option B is two agents on Saturday. Metric is Monday backlog carryover and first response time on weekend created tickets. Flip threshold is “if Monday carryover exceeds 20 percent of weekly volume for three weeks, we pilot weekend coverage.”
Tradeoff to keep you honest: outcome metrics lag and are harder to game, but they move slowly. Process metrics are steerable and faster, but people can accidentally optimize them in ways that make customers miserable. Your workflow should explicitly hold both truths.
Trust the signal before you trust the trend: the 10–15 minute data sanity workflow (coverage, definitions, mix)
| Assignment strategy | Best for | Advantages | Risks | Recommended when |
|---|---|---|---|---|
| 1. Coverage Check | Validating data completeness and source integrity | Quickly identifies missing data or broken pipelines. high confidence in metric source | Can miss subtle definition shifts. requires clear data source mapping | If coverage < 95% then metric is 'diagnostic-only' this week. Annotate and quarantine. |
| 3. Mix Check (e.g., Channel Mix Shift) | Detecting underlying shifts in composition that impact aggregate metrics | Explains sudden metric changes not due to performance. highlights external factors | Requires understanding of contributing segments. can be mistaken for true performance change | If channel mix shift > 5% — e.g., new intake channel, annotate metric as 'mix-adjusted' or 'segment-specific'. |
| 6. Decision Relevance Check | Confirming the metric still informs a specific business decision | Ensures metrics remain actionable. prevents 'metric bloat' | Can lead to premature metric retirement. requires clear decision frameworks | If metric does not change a decision, label as 'informational' or 'deprecate'. |
| 2. Definition Check | Ensuring metric meaning aligns with current business context | Catches changes in how data is collected or interpreted. prevents misinterpretation | Requires up-to-date documentation. can be time-consuming for complex metrics | If definition changed, re-run with corrected definition or label as 'historical' for comparison. |
| 4. Outlier/Anomaly Check | Identifying unusual data points that skew overall trends | Prevents overreaction to noise. focuses attention on true signals | Can dismiss real but unexpected changes. requires clear outlier thresholds | If an anomaly is detected, investigate root cause. If data error, quarantine. If real, understand impact. |
| 5. Comparison Check (e.g., Missing Intake Channel) | Validating metric against related or expected data points | Reveals inconsistencies or gaps in data collection. builds confidence in metric accuracy | Requires reliable comparison data. can be misleading if comparison data is also flawed | If a critical intake channel is missing or misrouted, flag metric as 'incomplete' and do not use for decision-making. |
Coverage checks: missing tickets, hidden channels, and silent failure queues
Before you argue about whether a metric is up or down, confirm it is measuring the same world as last week. Most “detect bad support data signals” work comes down to a few brutally boring checks that people skip because the chart looks pretty.
Coverage is first. Are you actually counting all tickets that customers created? The fastest way to get fooled is to lose an intake source without noticing. A common real world example: a new in app form routes to a different system, or a partner channel starts emailing a shared inbox that nobody ingests into your reporting. Your backlog “improves” because those tickets are now invisible, while escalations rise because the only people noticing are execs.
Stop rule: if coverage is below 95 percent for a core channel or queue, downgrade the affected metric to diagnostic only for the week. Do not steer staffing decisions with a blindfold.
Containment: annotate the dashboard, quarantine the trend line, and re run after intake is reconciled.
Definition drift checks: tags, categories, macro mapping, and SLA clock rules
Definition drift is the quieter cousin of coverage failure. Nothing “breaks,” but everything changes.
Tags get renamed. Categories get merged. A macro starts applying a new reason code. Someone adjusts what counts as “first response,” or changes which statuses pause the SLA clock. The metric moves and everyone thinks performance changed.
Stop rule: if definitions changed in the last reporting window and you cannot restate the before and after in one sentence, the metric is diagnostic only.
Containment: add a note right on the metric for this week and keep a change log entry so you do not forget why the line bent.
Channel mix and complexity shifts: why AHT, CSAT, SLA move even if performance didn’t
Here is the channel mix trap that blows up support KPI sanity checks.
Let’s say chat share rises from 25 percent to 45 percent because you launched chat more prominently in the product. Chat is faster, so AHT falls. Meanwhile, the hardest issues now land in email because customers who can wait tend to be the ones with complex problems. Email gets slower, backlog ages, and escalations rise. Your blended AHT chart looks like a win while your system is actually under more strain.
Stop rule: if the channel mix shifts materially, any blended metric becomes diagnostic only until you view it by channel and by complexity band.
Containment: split the metric in the meeting view, and do not compare blended AHT week over week like it is the same thing.
Deflection attribution reality check: what got removed from the denominator?
Deflection metrics are especially sensitive to denominator tricks, sometimes accidental, sometimes… optimistic.
A classic example: email survey delivery breaks, so CSAT is now mostly collected on chat, where customers are in the moment and response rates are higher. CSAT jumps. Leaders cheer. In reality, you stopped hearing from the angriest segment because they were in the channel that lost surveys.
Another: you improve self service, but it mostly deflects easy tickets. The remaining intake is harder, so SLA compliance drops and AHT rises. If you do not adjust for complexity, you will misdiagnose the team as “getting worse” right after you removed the easy work.
Stop rule: if the denominator for any metric changed, you must name what changed before you interpret the trend.
Containment: annotate and bring a like for like view, such as “within chat only,” “within billing only,” or “excluding deflected categories.”
Below is a support metrics sanity check workflow table you can run before weekly ops, QBRs, or any time someone says “the numbers look off.”
Coverage Check: If you cannot account for where tickets entered, do not trust any downstream trend.
Definition Check: When the clock rules or tags move, your “improvement” might be paperwork.
Mix Check (e.g., Channel Mix Shift): Blended metrics lie when customers shift channels.
Outlier/Anomaly Check: A surprise spike is a question, not a victory lap.
Automation vs human judgment: where workflows, macros, and AI summaries are safe—and where they quietly rewrite the metric
Automation safe metrics: counts, timeliness, and routing throughput (with guardrails)
Automation is great at doing the same thing the same way every time. That makes it excellent for operational throughput, and dangerous for meaning.
Counts, timestamps, and routing throughput are usually safe to automate and track. Volume by channel. First response time. Ticket touch counts. Queue transfers. These are “what happened” metrics, and the system can usually record them without interpretation.
The guardrail is simple: when automation changes the workflow, it can change the measured world. So you need to treat automation rollouts like product launches, not like background noise.
Judgment required metrics: sentiment, resolution quality, and ‘reason’ codes
The moment a metric depends on human judgment, you need to assume drift.
Reason codes are the classic trap. Agents pick whatever tag gets them through the form fastest. Then leadership makes product decisions from that taxonomy as if it came from a careful analyst. If you want “top contact drivers” to be decision grade, you need periodic sampling and calibration. Otherwise, it is diagnostic only.
Sentiment and quality are similar. AI summaries can help reviewers move faster, but they can also nudge what reviewers notice. If your QA score suddenly rises right after you shipped a new rubric or an AI assistant, it might be the tool rewriting the grade, not the team improving.
Change log discipline: the simplest way to stop ‘process updates’ from masquerading as performance
If you do only one thing from this article, do this: keep a lightweight change log and review it alongside KPI movement every week.
It does not need to be fancy. It needs to exist.
A simple template that works:
Date. Change. Scope. Owner. Expected metric impact. Metrics at risk. When to re evaluate.
Now the concrete anchor you have probably lived through. You deploy a macro that inserts a polished troubleshooting checklist. Agents love it because it reduces typing. AHT drops. Two weeks later, reopen rate rises because the macro pushes customers through steps that do not fit edge cases, and agents close tickets with “reply if you still need help.” The macro did exactly what it was designed to do. Your metric story was wrong.
That is why “support dashboard sanity check” needs a process lens. Automation can change customer behavior, agent behavior, and categorization behavior, all at once.
Tradeoffs: speed vs quality, consistency vs nuance, and how to pick the right leading indicator
Automation always trades nuance for consistency. That can be a great trade when the alternative is chaos. It is a bad trade when the nuance is the customer’s actual problem.
A practical tip that saves pain: when you roll out automation that should reduce time, pair it with a quality leading indicator you trust. If you want AHT to fall, commit to watching reopen rate, escalations, and a small QA sample for the affected categories. Otherwise, you will optimize for speed the way a teenager “cleans” their room by shoving everything into the closet.
Decision rule for the meeting: if a major process update shipped this week and it touches intake, routing, macros, AI summaries, or surveys, freeze any metric driven operational change that depends on those measures until you have at least one clean comparison window. Use the week for diagnosis, not celebration.
Five common instrumentation movers to put in your change log, because each can bias a KPI:
First, macro updates can lower handle time while increasing reopen rate.
Second, routing changes can improve SLA for one queue by dumping complexity into another.
Third, survey changes can create instant CSAT lifts by changing who gets asked and when.
Fourth, AI summarization rollout can change QA outcomes by changing what reviewers focus on.
Fifth, launching a new channel can distort blended AHT, SLA, and CSAT by shifting mix.
If you want a broader evaluation mindset for automation decisions, this is a helpful read: [2]
Failure modes: the ways ‘good’ support metrics get better while your system gets worse
Sudden CSAT lifts: sampling bias, timing changes, and who stopped receiving surveys
A CSAT spike is not inherently good news. It is a prompt to ask, “Who did we just stop hearing from?”
Concrete anchor: your CSAT jumps from 4.2 to 4.6 in a week. Escalations also rise. One likely explanation is not “everyone got nicer.” It is that survey delivery changed. Maybe email surveys failed, maybe you switched from 24 hours later to immediately after closure, maybe the trigger now excludes certain ticket types.
Diagnostic sequence for a CSAT spike (3 to 5 checks):
First, check survey coverage by channel and category. Who received surveys this week versus last week?
Second, check response rate. A higher score with a lower response rate is often a different audience.
Third, check timing. Asking immediately after a quick chat is a different emotional moment than asking after a multi day email thread.
Fourth, check whether any high severity or refund categories were excluded.
Fifth, triangulate with escalations and reopen rate. If those rose, treat CSAT as diagnostic only until explained.
AHT drops: deflection, macro shortcuts, premature closes, or complexity displacement
AHT dropping is another “could be great, could be fake” moment.
Concrete anchor: AHT falls 18 percent after you launched new self service flows. At the same time, backlog age worsens in email and reopen rate ticks up. The likely story is complexity displacement. Easy issues got deflected. The remaining work got harder. Your blended AHT looks better because chat grew, not because resolution improved.
Diagnostic sequence for an AHT drop (3 to 5 checks):
First, split AHT by channel and by category. If only chat improved, it might be mix.
Second, check deflection and intake mix. Did volume drop in easy categories?
Third, check reopen and escalation rates. Shorter tickets with more reopens is not “efficiency.”
Fourth, check closure reasons and macro usage. Did you introduce a “closing script” that ends conversations faster?
Fifth, sample ten tickets from the fastest cohort and read them. The truth is usually there.
SLA and backlog optics: splitting tickets, pausing clocks, queue hiding, or backlog aging
SLA and backlog are fertile ground for accidental gaming. Sometimes nobody is trying to cheat. They are just trying to survive.
Failure modes you should assume exist somewhere in your system:
SLA improves because you changed clock pause rules, not because response got faster.
Backlog shrinks because tickets are being split into smaller tickets, inflating “solved” while the customer still has one unresolved issue.
First response time improves because agents send a placeholder reply, then the real work starts later.
Backlog looks fine because one queue is hidden or misrouted, so aging happens off dashboard.
AHT drops because of premature closes and “reply if you still need help” patterns.
Deflection looks great because hard issues were excluded from the tracked denominator.
CSAT rises because fewer negative segments receive surveys or because timing moved.
Quality score rises because the rubric changed or reviewers changed, not because behavior changed.
Escalations rise while headline KPIs improve, because complexity is being displaced to senior staff.
What to do when you spot a failure mode: contain, annotate, and pick a safer proxy
When you catch a failure mode, do not start with blame. Start with containment.
Use a simple mapping in the meeting so you do not spiral into anecdotes.
Red flag: CSAT up sharply. Likely cause: sampling bias or timing change. Next check: survey coverage and response rate by channel. Meeting decision: downgrade CSAT to diagnostic only and steer on escalations plus QA sample this week.
Red flag: AHT down sharply. Likely cause: channel mix shift or premature closures. Next check: reopen rate, escalation rate, and AHT by channel. Meeting decision: pause AHT targets, run a quality push in the affected categories.
Red flag: SLA compliance up while backlog age worsens. Likely cause: clock rules or focus on first response only. Next check: time to resolution, backlog age bands, and oldest ticket review. Meeting decision: prioritize backlog burn and measure aged backlog reduction, not just first response.
Safer proxies to keep handy when a headline metric is contaminated:
Reopen rate is a great lie detector for “speed improvements.”
Escalation rate is a great lie detector for “everything is fine.”
Backlog age bands beat total backlog because they expose slow rot.
A small first contact resolution sample beats a single blended percentage.
If you want a philosophical nudge to remove metrics that never pay rent, this is a fun read: [3]
Bring a one-page “decision delta”: what changed, what it means, what you’ll do next week
The minimum pre-read: decisions, flip thresholds, trusted metrics, and annotations
The best support metrics meetings are not longer. They are pre wired.
Bring a one page “decision delta” instead of a twenty slide tour of charts. The job of the pre read is to answer: what changed, can we trust it, what decision does it affect, and what are we doing next.
Here is a copy friendly one page outline you can paste into a doc.
Decision delta, week of [date]
- Decisions to make this week
- Decision
- Options
- Flip threshold
- Owner
- Decision grade metrics only
- Metric, current value, prior value
- One sentence interpretation
- Confidence label: decision grade or diagnostic only
- Diagnostics and annotations
- What failed the support data quality checks
- What changed in definitions, coverage, mix
- What we are doing to fix or re run
- Change log highlights
- Process, routing, macro, survey, AI summary changes this period
- Metrics at risk
- Next actions and follow ups
- Action
- Owner
- Due date
- What we will check next week
How to present uncertainty without losing credibility
Leaders do not lose trust because you say “we are not sure.” They lose trust because you pretend certainty and get caught.
Use crisp language: “CSAT is diagnostic only this week due to survey coverage drop in email. We are steering on escalation rate and QA sample until coverage is restored.” That reads as competence, not weakness.
Concrete anchor example of downgrading: “First response time improved 22 percent, but Coverage Check found in app tickets were not ingested for two days. Metric is diagnostic only. Decision impact is paused for staffing reallocation until corrected counts are confirmed.”
A tight meeting close: owners, next checks, and what to remove from the dashboard
Close the meeting by removing something. Seriously. Every week you keep every metric, your dashboard becomes a junk drawer.
Your cadence can be simple.
Run the support metrics sanity check workflow before the meeting.
Update the change log.
Make the decisions, with flip thresholds.
Follow up next week on whether the decision was correct, and whether the metric was actually decision grade.
Monday plan that works in the real world:
First action: schedule a 25 minute working session with support ops and one team lead to label your current dashboard metrics as decision grade, diagnostic only, or remove.
Three priorities for the week: (1) pick two decisions and write flip thresholds, (2) run the coverage, definition, and mix checks on your core metrics, (3) start the change log and review it next to every big KPI movement.
Realistic production bar: if you can ship a one page decision delta for next week’s ops review and downgrade at least one shaky metric to diagnostic only with a clear annotation, you are already operating above the industry average. Copy the template, run it once, and iterate. That is how this becomes a system, not a pep talk.
Sources
- calypso.ms — calypso.ms
- topickz.com — topickz.com
- counting-stuff.com — counting-stuff.com

