[{"data":1,"prerenderedAt":47},["ShallowReactive",2],{"/en/blog/the-two-question-test-for-any-metric-would-we-notice-and-would-we-know-what-to-d":3,"/en/blog/the-two-question-test-for-any-metric-would-we-notice-and-would-we-know-what-to-d-surround":38},{"id":4,"locale":5,"translationGroupId":6,"availableLocales":7,"alternates":8,"_path":9,"path":9,"title":10,"description":11,"date":12,"modified":12,"meta":13,"seo":23,"topicSlug":28,"tags":29,"body":31,"_raw":36},"a08e20fc-a4e2-4403-9a3e-3d9121f6bdfd","en","5cc153a2-abda-4e03-858d-56ea020f89d3",[5],{"en":9},"/en/blog/the-two-question-test-for-any-metric-would-we-notice-and-would-we-know-what-to-d","The Two Question Test for Any Metric: Would We Notice and Would We Know What to Do","A practical two question test for metrics that helps support leaders cut dashboard clutter, choose support KPIs that matter, and connect metric movement to staffing, automation, and escalation plays.","2026-07-30T09:18:01.710Z",{"date":12,"badge":14,"authors":17},{"label":15,"color":16},"New","primary",[18],{"name":19,"description":20,"avatar":21},"Lucía Ferrer","Calypso AI · Clear, expert-led guides for operators and buyers",{"src":22},"https://api.dicebear.com/9.x/personas/svg?seed=calypso_expert_guide_v1&backgroundColor=b6e3f4,c0aede,d1d4f9,ffd5dc,ffdfbf",{"title":24,"description":25,"ogDescription":25,"twitterDescription":25,"canonicalPath":9,"robots":26,"schemaType":27},"The Two Question Test for Any Metric: Would We Notice and","A practical two question test for metrics that helps support leaders cut dashboard clutter, choose support KPIs that matter, and connect metric movement to","index,follow","BlogPosting","decision_systems_researcher",[30],"the-two-question-test-for-any-metric-would-we-notice-and-would-we-know-what-to-d",{"toc":32,"children":34,"html":35},{"links":33},[],[],"\u003Ch2>When the dashboard is green but the week is on fire: why most metrics fail at the moment you need them\u003C/h2>\n\u003Cp>Monday morning looks fine. CSAT is steady. SLA is mostly green. First response time is “within target.” The dashboard gives everyone permission to breathe.\u003C/p>\n\u003Cp>By Thursday, you’re in the kind of week that makes artisanal pottery look like a rational career plan.\u003C/p>\n\u003Cp>Escalations spike. The VIP queue is aging. A handful of high-priority tickets are stuck waiting on engineering. Meanwhile the team is thrashing across channels—working hard, but not necessarily pulling the levers that reduce risk.\u003C/p>\n\u003Cp>When you look back, the signals were there. They were just trapped inside averages, rollups, and “overall” numbers that never raised a hand.\u003C/p>\n\u003Cp>That’s why most support metrics fail at the moment you need them.\u003C/p>\n\u003Ch3>The surprise gap: ‘we tracked it’ isn’t the same as ‘we would have noticed’\u003C/h3>\n\u003Cp>Support teams gravitate toward clean outcomes because they’re easy to explain: CSAT, SLA attainment, deflection, QA score.\u003C/p>\n\u003Cp>But many are lagging indicators. They describe the week you already lived, not the week you’re walking into.\u003C/p>\n\u003Cp>A classic surprise: overall SLA is green while the backlog quietly shifts toward high-priority tickets and older age buckets. You “tracked backlog” as a count, but you would not have noticed the queue was getting dangerous until escalations started.\u003C/p>\n\u003Cp>If a metric can’t reliably pull attention early enough to change the outcome, it’s not a management tool. It’s a retrospective.\u003C/p>\n\u003Ch3>The action gap: ‘we saw it’ isn’t the same as ‘we knew what to do’\u003C/h3>\n\u003Cp>Even when a dashboard does get attention, the next hour often turns into interpretive dance: people stare, someone says “interesting,” and then you schedule a meeting to decide what it means.\u003C/p>\n\u003Cp>That’s the action gap. The metric moved, but there’s no named owner, no pre-decided response, and no decision rights. “We’ll keep an eye on it” becomes the default playbook.\u003C/p>\n\u003Cp>This is why the line “If you cannot explain the decision, do not ship the metric” lands so well with operators. A metric without a decision attached becomes reporting theater.\u003C/p>\n\u003Cp>Calypso frames this as a review workflow: \u003Ca href=\"#ref-1\" title=\"calypso.ms — calypso.ms\">[1]\u003C/a>\u003C/p>\n\u003Ch3>The two question test for metrics (and what it replaces)\u003C/h3>\n\u003Cp>Here’s the two question test for metrics:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Would we notice?\u003C/strong> If this metric drifted in a meaningful way, would we reliably see it \u003Cem>in time\u003C/em> to intervene?\u003C/li>\n\u003Cli>\u003Cstrong>Would we know what to do?\u003C/strong> If we noticed, do we already know the first actions, the owner, and the escalation path?\u003C/li>\n\u003C/ul>\n\u003Cp>A metric “passes” only when both answers are yes.\u003C/p>\n\u003Cp>By the end, you’ll be able to classify every number on your support dashboard as \u003Cstrong>Keep, Watch, Retire, or Redesign\u003C/strong>—and, more importantly, connect metric movement to staffing moves, automation tuning, and escalation triggers.\u003C/p>\n\u003Ch2>Run the test in 10 minutes: a practical rubric to keep, watch, retire, or redesign each support metric\u003C/h2>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Assignment strategy\u003C/th>\n\u003Cth>Best for\u003C/th>\n\u003Cth>Advantages\u003C/th>\n\u003Cth>Risks\u003C/th>\n\u003Cth>Recommended when\u003C/th>\n\u003C/tr>\n\u003C/thead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>Redesign — Score 2: Notice 1, Action 1 or Notice 0, Action 2 or Notice 2, Action 0\u003C/td>\n\u003Ctd>Metrics that are either hard to notice or hard to act on, but the underlying problem is important\u003C/td>\n\u003Ctd>Focuses effort on improving metric utility. prevents premature retirement of valuable concepts\u003C/td>\n\u003Ctd>Resource-intensive redesign. potential for endless iteration without clear goals\u003C/td>\n\u003Ctd>The metric concept is critical, but its current form fails either the &#39;notice&#39; or &#39;action&#39; test\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Keep (Score 4: Notice 2, Action 2)\u003C/td>\n\u003Ctd>Operational metrics with clear thresholds and pre-defined playbooks\u003C/td>\n\u003Ctd>Immediate actionability. reduces decision fatigue. high trust in data\u003C/td>\n\u003Ctd>Over-reliance on static playbooks. missing novel issues. alert fatigue if thresholds are too sensitive\u003C/td>\n\u003Ctd>Metric movement directly maps to a known, documented response — e.g., staffing adjustment, escalation\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Guardrail: &#39;Notice in Time&#39; Definition\u003C/td>\n\u003Ctd>Ensuring metric alerts precede critical operational deadlines\u003C/td>\n\u003Ctd>Proactive problem solving. prevents escalations due to delayed awareness\u003C/td>\n\u003Ctd>Overly aggressive thresholds leading to false positives. ignoring human response time\u003C/td>\n\u003Ctd>Defining lead times for staffing, incident response, or customer communication\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Guardrail: &#39;Know What to Do&#39; Definition\u003C/td>\n\u003Ctd>Ensuring metric movement triggers pre-agreed, documented playbooks\u003C/td>\n\u003Ctd>Standardized response. reduces ad-hoc decision-making. faster resolution\u003C/td>\n\u003Ctd>Rigid playbooks failing in novel situations. lack of empowerment for frontline teams\u003C/td>\n\u003Ctd>Establishing clear, repeatable actions for common metric fluctuations\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Retire or Context-Only — Score 0-1: Notice 0, Action 0 or Notice 1, Action 0\u003C/td>\n\u003Ctd>Vanity metrics, metrics with no clear impact, or those that consistently fail both tests\u003C/td>\n\u003Ctd>Reduces dashboard clutter. frees up resources. improves signal-to-noise ratio\u003C/td>\n\u003Ctd>Losing historical context. overlooking latent issues if retired too quickly\u003C/td>\n\u003Ctd>The metric provides no actionable insight and its movement doesn&#39;t trigger any response\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Keep/Watch (Score 3: Notice 2, Action 1 or Notice 1, Action 2)\u003C/td>\n\u003Ctd>Metrics with clear signals but evolving responses, or clear responses but subtle signals\u003C/td>\n\u003Ctd>Balances stability with adaptability. encourages playbook refinement\u003C/td>\n\u003Ctd>Delayed action if &#39;watch&#39; period is too long. misinterpreting weak signals\u003C/td>\n\u003Ctd>Metric is new, undergoing a pilot, or its operational response needs further definition/testing\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Worked Example: High CSAT, Low Resolution Rate\u003C/td>\n\u003Ctd>Illustrating how the test applies to conflicting metrics\u003C/td>\n\u003Ctd>Highlights the need for deeper investigation. prevents misleading conclusions\u003C/td>\n\u003Ctd>Misinterpreting individual metric scores without context\u003C/td>\n\u003Ctd>Training teams on applying the rubric to complex, multi-metric scenarios\u003C/td>\n\u003C/tr>\n\u003C/tbody>\u003C/table>\n\u003Cp>Use the rubric below to score each metric on two dimensions—\u003Cstrong>Notice\u003C/strong> and \u003Cstrong>Action\u003C/strong>—then assign it a disposition.\u003C/p>\n\u003Cp>Reference this table as the “score-to-decision” map (it’s the piece teams skip, then wonder why the dashboard keeps growing).\u003C/p>\n\u003Cp>A key point: this isn’t a taste test (“I like this metric”). It’s an operations test (“will this metric help us run the week?”).\u003C/p>\n\u003Ch3>Question 1 — Would we notice? (signal, sensitivity, timing)\u003C/h3>\n\u003Cp>Score \u003Cstrong>Notice\u003C/strong> from 0 to 2:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>0:\u003C/strong> You wouldn’t reliably see meaningful change. It’s buried in weekly averages, reviewed too late, or too noisy to trust.\u003C/li>\n\u003Cli>\u003Cstrong>1:\u003C/strong> You might notice, but not consistently—or not in time.\u003C/li>\n\u003Cli>\u003Cstrong>2:\u003C/strong> You would notice \u003Cem>in time to act\u003C/em>.\u003C/li>\n\u003C/ul>\n\u003Cp>“In time” is where teams get burned. A metric can be perfectly accurate and still arrive after your last useful lever window.\u003C/p>\n\u003Cp>Anchor your Notice score to real lead times:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Routing / workflow changes:\u003C/strong> hours to a day.\u003C/li>\n\u003Cli>\u003Cstrong>Automation or deflection tuning:\u003C/strong> days to a week (often longer if approvals are required).\u003C/li>\n\u003Cli>\u003Cstrong>Staffing shifts:\u003C/strong> 1–2 weeks for schedules; months for hiring.\u003C/li>\n\u003C/ul>\n\u003Cp>Decision rule: if a metric reliably alerts you \u003Cem>after\u003C/em> the last reasonable moment you could change the outcome for the lever you care about, it fails the Notice test for that lever.\u003C/p>\n\u003Cp>This is also why the “question before the number” mindset matters—start from the decision and work backward to the signal: \u003Ca href=\"#ref-2\" title=\"lospino.so — lospino.so\">[2]\u003C/a>\u003C/p>\n\u003Ch3>Question 2 — Would we know what to do? (owner + response)\u003C/h3>\n\u003Cp>Score \u003Cstrong>Action\u003C/strong> from 0 to 2:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>0:\u003C/strong> The metric sparks meetings, not action. No owner, or the “owner” can’t actually change anything.\u003C/li>\n\u003Cli>\u003Cstrong>1:\u003C/strong> There’s an intuitive reaction, but it isn’t pre-committed. People improvise.\u003C/li>\n\u003Cli>\u003Cstrong>2:\u003C/strong> There’s a pre-agreed response: owner, first moves, escalation trigger, and a rollback condition.\u003C/li>\n\u003C/ul>\n\u003Cp>This is the quiet truth: actionability isn’t a property of the metric. It’s a property of the metric \u003Cstrong>plus\u003C/strong> your operating system.\u003C/p>\n\u003Cp>A fast gut-check is the “so what” test—if you can’t finish the sentence “so what will we do differently,” the metric isn’t decision-grade: \u003Ca href=\"#ref-3\" title=\"kaushik.net — kaushik.net\">[3]\u003C/a>\u003C/p>\n\u003Ch3>Assign the disposition (Keep / Watch / Retire / Redesign)\u003C/h3>\n\u003Cp>Add the scores:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>4 = Keep.\u003C/strong> It belongs on the dashboard that runs your week.\u003C/li>\n\u003Cli>\u003Cstrong>3 = Keep/Watch.\u003C/strong> Useful, but either the signal or response needs maturity.\u003C/li>\n\u003Cli>\u003Cstrong>2 = Redesign.\u003C/strong> The intent is valid; the signal or action path is broken.\u003C/li>\n\u003Cli>\u003Cstrong>0–1 = Retire or Context-only.\u003C/strong> Keep it for storytelling or research, not steering.\u003C/li>\n\u003C/ul>\n\u003Cp>Two guardrails keep teams honest (and match the table rows people tend to gloss over):\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Guardrail: Notice in Time Definition.\u003C/strong> “We look weekly” isn’t a Notice 2 if you needed to act Tuesday.\u003C/li>\n\u003Cli>\u003Cstrong>Guardrail: Know What to Do Definition.\u003C/strong> “We’ll discuss” isn’t Action 2. It’s Action 0 with better manners.\u003C/li>\n\u003C/ul>\n\u003Ch3>Make it real: quick scoring examples\u003C/h3>\n\u003Cp>\u003Cstrong>SLA:\u003C/strong> Many teams auto-rate SLA as “Keep.” But if you only review weekly SLA attainment, Notice is often a \u003Cstrong>1\u003C/strong> (you learn after the miss). If you segment SLA by customer tier and ticket age \u003Cem>and\u003C/em> you have a duty manager who can reroute same day, it becomes \u003Cstrong>Notice 2 / Action 2\u003C/strong>.\u003C/p>\n\u003Cp>Same metric name. Totally different operational value.\u003C/p>\n\u003Cp>\u003Cstrong>CSAT:\u003C/strong> Overall CSAT is frequently \u003Cstrong>Notice 0–1\u003C/strong> because response rates are uneven, surveys don’t cover every channel, and the sample drifts. Action is often \u003Cstrong>1\u003C/strong> because the “response” is a monthly verbatim review that rarely changes staffing or workflows. CSAT often deserves \u003Cstrong>Redesign\u003C/strong>—segmentation by contact reason, minimum response counts, and visible sample health—more than it deserves retirement.\u003C/p>\n\u003Cp>\u003Cstrong>Worked example: High CSAT, low resolution rate:\u003C/strong> This is where the rubric prevents false comfort. High CSAT can coexist with low resolution if surveys hit “easy” tickets, or if customers are polite but still stuck. Score them separately, then ask: would low resolution be noticed early, and do we know what to do (capacity shift, engineering escalation, macro changes)? The “conflict” is the point: it forces investigation instead of letting one shiny number win.\u003C/p>\n\u003Ch2>Make ‘Would we notice?’ true: designing signals that catch problems before staffing and escalations get ugly\u003C/h2>\n\u003Cp>The Notice question isn’t “do we have a number?” It’s “do we have a signal that shows up early enough to matter?”\u003C/p>\n\u003Cp>Support operations is full of slow-motion failures. The queue looks fine until tail risk shows up, and then the only options left are heroics.\u003C/p>\n\u003Ch3>What breaks first: leading indicators beat end-of-week averages\u003C/h3>\n\u003Cp>Averages are comforting. They also hide pain.\u003C/p>\n\u003Cp>Concrete example: average first response time across all tickets stays at 2 hours. Looks great. Meanwhile, P1 tickets are waiting 6 hours, but there are fewer of them, so the average stays polite.\u003C/p>\n\u003Cp>Customers don’t experience your average. They experience the worst slice of your week.\u003C/p>\n\u003Cp>Leading indicators tend to crack before CSAT collapses:\u003C/p>\n\u003Cul>\n\u003Cli>Backlog \u003Cstrong>age\u003C/strong>, especially for high priority and VIP tiers\u003C/li>\n\u003Cli>Queue mix (share of P1/P2)\u003C/li>\n\u003Cli>\u003Cstrong>Reopen rate\u003C/strong>, which often rises before complaints do\u003C/li>\n\u003Cli>\u003Cstrong>Escalation rate\u003C/strong>, which is blunt—but early—when trust is cracking\u003C/li>\n\u003C/ul>\n\u003Cp>If you’re choosing between “count” and “age,” pick age for operational control. Counts can rise for healthy reasons (growth, launches). Age rarely rises for healthy reasons.\u003C/p>\n\u003Cp>A useful anchor: watch the \u003Cem>percent of high-priority tickets older than X hours\u003C/em> for two checks in a row. It’s harder to ignore, and it maps directly to customer harm.\u003C/p>\n\u003Ch3>Thresholds vs baselines: avoid false calm and constant panic\u003C/h3>\n\u003Cp>Static targets are seductive: “Keep first response time under 2 hours.”\u003C/p>\n\u003Cp>Support demand isn’t static. Some weeks you ship a feature. Some weeks a partner breaks. Some weeks vacation and channel mix do their thing.\u003C/p>\n\u003Cp>Static thresholds create two failure modes:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>False calm:\u003C/strong> metric stays green while risk accumulates in one segment.\u003C/li>\n\u003Cli>\u003Cstrong>Constant panic:\u003C/strong> metric is always red, training everyone to ignore it.\u003C/li>\n\u003C/ul>\n\u003Cp>You don’t need fancy math to fix this. You need a shared definition of “meaningful change.” A simple approach: treat movement as meaningful when it shifts by a clear absolute amount \u003Cem>or\u003C/em> a clear percentage versus a recent baseline, and stays there for more than one check cycle.\u003C/p>\n\u003Cp>Decision rule: decide whether the metric is meant to catch \u003Cstrong>spikes\u003C/strong> or \u003Cstrong>trends\u003C/strong>.\u003C/p>\n\u003Cul>\n\u003Cli>Spikes want faster checks and smaller tolerances.\u003C/li>\n\u003Cli>Trends want slower confirmation and stronger segmentation.\u003C/li>\n\u003C/ul>\n\u003Ch3>Segment or lie: channel mix, priority, and region hide the story\u003C/h3>\n\u003Cp>“Overall” is where operational truth goes to get blended into nonsense.\u003C/p>\n\u003Cp>Common scenario: overall SLA is green. Email SLA is red. Chat SLA is very green. You staffed heavily for chat because it’s loud and real-time; email quietly absorbs the pain. The average tells a story of success while one segment waits.\u003C/p>\n\u003Cp>Start with segments that actually change decisions:\u003C/p>\n\u003Cul>\n\u003Cli>Priority/severity (P1/P2 behave like a different business)\u003C/li>\n\u003Cli>Customer tier (paid and free users shouldn’t be blended)\u003C/li>\n\u003Cli>Channel (chat, email, phone have different constraints)\u003C/li>\n\u003Cli>Region/language (follow-the-sun makes global averages misleading)\u003C/li>\n\u003C/ul>\n\u003Cp>Tradeoff warning: segmentation can create alert fatigue if you turn every slice into an alarm.\u003C/p>\n\u003Cp>Decision rule: alert on slices with distinct owners or distinct levers. If nobody can act on “APAC email,” don’t ship alerts for it. Fix decision rights first.\u003C/p>\n\u003Ch2>Make ‘Would we know what to do?’ true: mapping metric movement to staffing, automation, and escalation playbooks\u003C/h2>\n\u003Cp>Lots of dashboards can tell you something changed. The hard part is getting the org to move quickly and consistently without waiting for the weekly meeting.\u003C/p>\n\u003Cp>Smart people still improvise under pressure. Pre-commitment is what saves you.\u003C/p>\n\u003Ch3>From metric to mechanism: capacity, workflow, demand shaping\u003C/h3>\n\u003Cp>Almost every support metric maps to three mechanisms:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Capacity:\u003C/strong> staffing, schedules, queue assignments, on-call rotations\u003C/li>\n\u003Cli>\u003Cstrong>Workflow:\u003C/strong> routing, triage, escalation criteria, macros, handoffs, tooling friction\u003C/li>\n\u003Cli>\u003Cstrong>Demand shaping:\u003C/strong> help center quality, deflection/self-serve, in-product guidance, status comms, and product fixes\u003C/li>\n\u003C/ul>\n\u003Cp>A metric becomes actionable when you can say which mechanism it should pull and who owns that mechanism.\u003C/p>\n\u003Cp>Concrete anchor: first response time rising in chat is usually capacity or routing—hours-to-fix levers. CSAT drifting down for one contact reason is often workflow or product—days-to-weeks levers.\u003C/p>\n\u003Cp>If you treat both as “weekly KPIs,” you’ll respond too slowly to one and too noisily to the other.\u003C/p>\n\u003Ch3>Pre-commit actions: small, medium, big (and reversible vs slow)\u003C/h3>\n\u003Cp>Good playbooks don’t treat every wobble as a crisis, and they don’t wait until the metric is catastrophic.\u003C/p>\n\u003Cp>In human terms, an action ladder for channel-level first response time might look like:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Small move (early drift):\u003C/strong> reroute one or two agents next shift, pause non-urgent internal work, post a heads-up with expectations.\u003C/li>\n\u003Cli>\u003Cstrong>Medium move (clear miss):\u003C/strong> add short-term coverage for a couple days, simplify triage temporarily, enable a pre-approved deflection assist.\u003C/li>\n\u003Cli>\u003Cstrong>Big move (sustained drift):\u003C/strong> escalate staffing tradeoffs, revisit forecast assumptions, open a cross-functional issue if demand is defect-driven.\u003C/li>\n\u003C/ul>\n\u003Cp>Split actions into \u003Cstrong>reversible\u003C/strong> versus \u003Cstrong>slow\u003C/strong>. Reversible actions (reroutes, queue caps, temporary macros, status messaging) are ideal early because you can roll them back. Slow actions (hiring, policy changes, major tooling shifts) need sustained signals.\u003C/p>\n\u003Cp>This is where teams get burned: playbooks that assume perfect conditions (“we’ll just add headcount”) instead of the levers you can actually pull this week.\u003C/p>\n\u003Ch3>Ownership and decision rights: who can act without a meeting\u003C/h3>\n\u003Cp>The fastest way to make a metric non-actionable is to require a meeting before anyone can touch the levers.\u003C/p>\n\u003Cp>The two question test forces the uncomfortable conversation: who owns this metric, and do they have permission to act?\u003C/p>\n\u003Cp>If the answer is “support leadership,” that can work—\u003Cem>if\u003C/em> there’s a duty owner empowered during the week. If the answer is “we all do,” the real answer is “nobody does.”\u003C/p>\n\u003Cp>A lightweight playbook shape (short enough to be used) is:\u003C/p>\n\u003Cp>Trigger → Owner → First moves → Escalate when → Roll back when.\u003C/p>\n\u003Cp>Two anchors:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cp>\u003Cstrong>SLA by customer tier:\u003C/strong> Trigger on VIP SLA or VIP backlog age crossing a threshold for two checks. Owner is the duty manager. First moves are rerouting top performers, pausing low-priority work, notifying the account team with a recovery estimate. Escalate if it stays red for a day or escalations spike. Roll back after a day of normal backlog age.\u003C/p>\n\u003C/li>\n\u003Cli>\u003Cp>\u003Cstrong>QA score:\u003C/strong> Trigger on a drop that persists across reviewers or concentrates in one queue. Owner is the QA lead. First moves are calibration, spot audit, and one coaching theme. Escalate when it maps to a policy/product change, not just reviewer drift. Roll back after calibration shows scoring noise.\u003C/p>\n\u003C/li>\n\u003C/ul>\n\u003Cp>Staffing lead time changes what you can promise. If schedules take 1–2 weeks to adjust, a weekly SLA report is not a staffing lever. Use faster operational signals (age-based backlog, channel-level response time) for staffing. Use slower outcomes (CSAT) as demand-shaping and product feedback.\u003C/p>\n\u003Cp>A complementary lens says the same thing plainly: metrics should change decisions, not fill slides. \u003Ca href=\"#ref-4\" title=\"rogerwong.me — rogerwong.me\">[4]\u003C/a>\u003C/p>\n\u003Ch2>Failure modes: how the test can still fool you (and how to catch it before a bad decision)\u003C/h2>\n\u003Cp>Even if a metric passes the two question test for metrics, it can still mislead you.\u003C/p>\n\u003Cp>Support metrics live in a world of incentives, sampling quirks, and shifting customer behavior. The goal isn’t cynicism. It’s avoiding confident wrongness.\u003C/p>\n\u003Ch3>Goodhart and gaming: when “better” isn’t better\u003C/h3>\n\u003Cp>Classic failure: average handle time drops because agents rush or avoid complex tickets. The metric improves; reopen rate and escalations rise.\u003C/p>\n\u003Cp>Catch it with a paired sanity check. If handle time improves while reopen rate worsens, you didn’t get more efficient—you pushed pain downstream.\u003C/p>\n\u003Cp>Response: demote handle time from a target to context, and read it alongside reopen rate by contact reason.\u003C/p>\n\u003Cp>Another failure: deflection rises because you count sessions that don’t create a ticket—even when the customer failed and left. It looks like success; backlog still rises.\u003C/p>\n\u003Cp>Catch it by pairing deflection with contact volume for the same intent. If “deflection” is up and contacts are also up, you’re probably measuring browsing, not resolution.\u003C/p>\n\u003Cp>Response: redesign deflection toward “solved intent,” not “no ticket created,” and tie it to a time window that reflects customer behavior.\u003C/p>\n\u003Ch3>Measurement trust issues: sampling bias and shifting contact reasons\u003C/h3>\n\u003Cp>CSAT can look stable because the sample changed: new channels, different survey triggers, response rate swings, or only happy customers replying.\u003C/p>\n\u003Cp>Catch it by making sample health impossible to ignore: minimum response counts, response rates, and segmentation by channel and contact reason. If trust is compromised, treat CSAT as Watch until survey design and coverage are repaired.\u003C/p>\n\u003Cp>QA can drift because reviewers change or the rubric shifts. You think quality improved; you mostly changed scoring.\u003C/p>\n\u003Cp>Catch it with periodic calibration and reviewer distribution checks. If you can’t keep the measurement stable, stop using QA score as a weekly steering lever.\u003C/p>\n\u003Ch3>Tradeoffs you must name (or the metrics will name them for you)\u003C/h3>\n\u003Cp>Support is full of real tradeoffs. Metrics get dangerous when teams pretend there is no trade.\u003C/p>\n\u003Cul>\n\u003Cli>Faster first response with lower CSAT can mean you’re stopping the clock, not solving.\u003C/li>\n\u003Cli>Higher deflection with higher customer effort can mean you’re pushing people into self-serve that doesn’t help.\u003C/li>\n\u003Cli>Green SLA with rising escalations often means the target is generous, segmentation is wrong, or the SLA doesn’t cover the customers who are actually upset.\u003C/li>\n\u003C/ul>\n\u003Cp>When two metrics disagree, don’t average them in your head. Pick a guardrail. Example: “We optimize response speed as long as reopen rate stays within range.” Naming the guardrail makes the tradeoff explicit—and harder to game.\u003C/p>\n\u003Cp>For a similar argument in a different voice (and a useful reminder that dashboards don’t interpret themselves), this is worth reading: \u003Ca href=\"#ref-5\" title=\"bumbleb.co — bumbleb.co\">[5]\u003C/a>\u003C/p>\n\u003Ch2>Your metric scorecard for next week: decide what drives staffing, automation, and escalations—and what gets downgraded\u003C/h2>\n\u003Cp>Dashboards sprawl the way closets do. Nobody plans it. It just happens, one well-intentioned request at a time.\u003C/p>\n\u003Cp>The cure isn’t a bigger dashboard. The cure is a recurring decision about what gets to drive action.\u003C/p>\n\u003Ch3>A one-page checklist to audit your current dashboard\u003C/h3>\n\u003Cp>Keep this tight and operational. The point isn’t to debate definitions first; it’s to decide what runs the week.\u003C/p>\n\u003Cul>\n\u003Cli>List the 10–15 metrics that show up most often in staff meetings and exec readouts.\u003C/li>\n\u003Cli>For each metric, score \u003Cstrong>Notice (0–2)\u003C/strong> and \u003Cstrong>Action (0–2)\u003C/strong>.\u003C/li>\n\u003Cli>For each \u003Cstrong>Keep\u003C/strong> candidate, write one decision sentence: “If this moves, we will do X within Y time.” If you can’t write it, it’s not Keep.\u003C/li>\n\u003Cli>For each \u003Cstrong>Redesign\u003C/strong> candidate, name what’s broken: signal, segmentation, timing, ownership, or response.\u003C/li>\n\u003Cli>For each \u003Cstrong>Retire/Context-only\u003C/strong> candidate, decide where it belongs instead: monthly readout, quarterly deep dive, or nowhere.\u003C/li>\n\u003C/ul>\n\u003Cp>Cap the decision-driver dashboard. Most teams do best with \u003Cstrong>three to five\u003C/strong> drivers. Everything else is supporting context. Teams get burned when they keep everything “just in case,” then miss the one metric that mattered.\u003C/p>\n\u003Ch3>What to do with each disposition (Keep / Watch / Retire / Redesign)\u003C/h3>\n\u003Cul>\n\u003Cli>\u003Cstrong>Keep:\u003C/strong> assign an owner with decision rights and a short playbook. If a Keep metric has no owner, that’s the first fix.\u003C/li>\n\u003Cli>\u003Cstrong>Watch:\u003C/strong> set an explicit cadence and a purpose (“we’re piloting,” “we’re validating signal quality,” “we’re building the playbook”). Watching everything is the same as watching nothing.\u003C/li>\n\u003Cli>\u003Cstrong>Retire:\u003C/strong> remove it from the dashboard that runs the week. If you’re nervous, move it to a secondary page first—just don’t give it steering-wheel real estate.\u003C/li>\n\u003Cli>\u003Cstrong>Redesign:\u003C/strong> improve one thing at a time. Most redesigns are boring (and that’s good): segment it, replace an average with a distribution, tie it to a decision threshold, add sample health.\u003C/li>\n\u003C/ul>\n\u003Cp>Common mistake: redesigning by making a metric fancier instead of more decision-ready. If the owner still can’t answer “what do we do when it moves,” you didn’t redesign—you decorated.\u003C/p>\n\u003Ch3>How to review monthly: keep the dashboard from re-sprawling\u003C/h3>\n\u003Cp>Once a month, re-score your top metrics with the two question test for metrics—especially after channel mix changes, product launches, or staffing shifts.\u003C/p>\n\u003Cp>Once a quarter, do a reset: retire at least one metric and promote at most one. Promotions require two things: a named owner and a written response.\u003C/p>\n\u003Cp>If you want a concrete Monday plan, keep it simple: take one screenshot of your current dashboard, audit 10 metrics, pick three decision drivers tied to staffing/automation/escalations, then commit to two redesigns and one retire.\u003C/p>\n\u003Cp>By Friday, you should have playbooks for the two metrics that most affect staffing and escalations, with triggers and owners who can act without a meeting.\u003C/p>\n\u003Cp>That’s how you turn “we tracked it” into “we ran the week.”\u003C/p>\n\u003Ch2>Sources\u003C/h2>\n\u003Col>\n\u003Cli>\u003Ca href=\"https://www.calypso.ms/en/blog/if-you-cannot-explain-the-decision-do-not-ship-the-metric-a-review-workflow-that\">calypso.ms\u003C/a> — calypso.ms\u003C/li>\n\u003Cli>\u003Ca href=\"https://lospino.so/blog/the-question-before-the-number\">lospino.so\u003C/a> — lospino.so\u003C/li>\n\u003Cli>\u003Ca href=\"https://www.kaushik.net/avinash/kill-useless-web-metrics-apply-so-what-test\">kaushik.net\u003C/a> — kaushik.net\u003C/li>\n\u003Cli>\u003Ca href=\"https://rogerwong.me/2026/07/metrics-should-change-decisions\">rogerwong.me\u003C/a> — rogerwong.me\u003C/li>\n\u003Cli>\u003Ca href=\"https://bumbleb.co/blog/2026-07-19-nobody-asks-a-dashboard-a-second-question\">bumbleb.co\u003C/a> — bumbleb.co\u003C/li>\n\u003C/ol>\n",{"body":37},"## When the dashboard is green but the week is on fire: why most metrics fail at the moment you need them\n\nMonday morning looks fine. CSAT is steady. SLA is mostly green. First response time is “within target.” The dashboard gives everyone permission to breathe.\n\nBy Thursday, you’re in the kind of week that makes artisanal pottery look like a rational career plan.\n\nEscalations spike. The VIP queue is aging. A handful of high-priority tickets are stuck waiting on engineering. Meanwhile the team is thrashing across channels—working hard, but not necessarily pulling the levers that reduce risk.\n\nWhen you look back, the signals were there. They were just trapped inside averages, rollups, and “overall” numbers that never raised a hand.\n\nThat’s why most support metrics fail at the moment you need them.\n\n### The surprise gap: ‘we tracked it’ isn’t the same as ‘we would have noticed’\n\nSupport teams gravitate toward clean outcomes because they’re easy to explain: CSAT, SLA attainment, deflection, QA score.\n\nBut many are lagging indicators. They describe the week you already lived, not the week you’re walking into.\n\nA classic surprise: overall SLA is green while the backlog quietly shifts toward high-priority tickets and older age buckets. You “tracked backlog” as a count, but you would not have noticed the queue was getting dangerous until escalations started.\n\nIf a metric can’t reliably pull attention early enough to change the outcome, it’s not a management tool. It’s a retrospective.\n\n### The action gap: ‘we saw it’ isn’t the same as ‘we knew what to do’\n\nEven when a dashboard does get attention, the next hour often turns into interpretive dance: people stare, someone says “interesting,” and then you schedule a meeting to decide what it means.\n\nThat’s the action gap. The metric moved, but there’s no named owner, no pre-decided response, and no decision rights. “We’ll keep an eye on it” becomes the default playbook.\n\nThis is why the line “If you cannot explain the decision, do not ship the metric” lands so well with operators. A metric without a decision attached becomes reporting theater.\n\nCalypso frames this as a review workflow: [[1]](#ref-1 \"calypso.ms — calypso.ms\")\n\n### The two question test for metrics (and what it replaces)\n\nHere’s the two question test for metrics:\n\n- **Would we notice?** If this metric drifted in a meaningful way, would we reliably see it *in time* to intervene?\n- **Would we know what to do?** If we noticed, do we already know the first actions, the owner, and the escalation path?\n\nA metric “passes” only when both answers are yes.\n\nBy the end, you’ll be able to classify every number on your support dashboard as **Keep, Watch, Retire, or Redesign**—and, more importantly, connect metric movement to staffing moves, automation tuning, and escalation triggers.\n\n## Run the test in 10 minutes: a practical rubric to keep, watch, retire, or redesign each support metric\n\n| Assignment strategy | Best for | Advantages | Risks | Recommended when |\n| --- | --- | --- | --- | --- |\n| Redesign — Score 2: Notice 1, Action 1 or Notice 0, Action 2 or Notice 2, Action 0 | Metrics that are either hard to notice or hard to act on, but the underlying problem is important | Focuses effort on improving metric utility. prevents premature retirement of valuable concepts | Resource-intensive redesign. potential for endless iteration without clear goals | The metric concept is critical, but its current form fails either the 'notice' or 'action' test |\n| Keep (Score 4: Notice 2, Action 2) | Operational metrics with clear thresholds and pre-defined playbooks | Immediate actionability. reduces decision fatigue. high trust in data | Over-reliance on static playbooks. missing novel issues. alert fatigue if thresholds are too sensitive | Metric movement directly maps to a known, documented response — e.g., staffing adjustment, escalation |\n| Guardrail: 'Notice in Time' Definition | Ensuring metric alerts precede critical operational deadlines | Proactive problem solving. prevents escalations due to delayed awareness | Overly aggressive thresholds leading to false positives. ignoring human response time | Defining lead times for staffing, incident response, or customer communication |\n| Guardrail: 'Know What to Do' Definition | Ensuring metric movement triggers pre-agreed, documented playbooks | Standardized response. reduces ad-hoc decision-making. faster resolution | Rigid playbooks failing in novel situations. lack of empowerment for frontline teams | Establishing clear, repeatable actions for common metric fluctuations |\n| Retire or Context-Only — Score 0-1: Notice 0, Action 0 or Notice 1, Action 0 | Vanity metrics, metrics with no clear impact, or those that consistently fail both tests | Reduces dashboard clutter. frees up resources. improves signal-to-noise ratio | Losing historical context. overlooking latent issues if retired too quickly | The metric provides no actionable insight and its movement doesn't trigger any response |\n| Keep/Watch (Score 3: Notice 2, Action 1 or Notice 1, Action 2) | Metrics with clear signals but evolving responses, or clear responses but subtle signals | Balances stability with adaptability. encourages playbook refinement | Delayed action if 'watch' period is too long. misinterpreting weak signals | Metric is new, undergoing a pilot, or its operational response needs further definition/testing |\n| Worked Example: High CSAT, Low Resolution Rate | Illustrating how the test applies to conflicting metrics | Highlights the need for deeper investigation. prevents misleading conclusions | Misinterpreting individual metric scores without context | Training teams on applying the rubric to complex, multi-metric scenarios |\n\nUse the rubric below to score each metric on two dimensions—**Notice** and **Action**—then assign it a disposition.\n\nReference this table as the “score-to-decision” map (it’s the piece teams skip, then wonder why the dashboard keeps growing).\n\nA key point: this isn’t a taste test (“I like this metric”). It’s an operations test (“will this metric help us run the week?”).\n\n### Question 1 — Would we notice? (signal, sensitivity, timing)\n\nScore **Notice** from 0 to 2:\n\n- **0:** You wouldn’t reliably see meaningful change. It’s buried in weekly averages, reviewed too late, or too noisy to trust.\n- **1:** You might notice, but not consistently—or not in time.\n- **2:** You would notice *in time to act*.\n\n“In time” is where teams get burned. A metric can be perfectly accurate and still arrive after your last useful lever window.\n\nAnchor your Notice score to real lead times:\n\n- **Routing / workflow changes:** hours to a day.\n- **Automation or deflection tuning:** days to a week (often longer if approvals are required).\n- **Staffing shifts:** 1–2 weeks for schedules; months for hiring.\n\nDecision rule: if a metric reliably alerts you *after* the last reasonable moment you could change the outcome for the lever you care about, it fails the Notice test for that lever.\n\nThis is also why the “question before the number” mindset matters—start from the decision and work backward to the signal: [[2]](#ref-2 \"lospino.so — lospino.so\")\n\n### Question 2 — Would we know what to do? (owner + response)\n\nScore **Action** from 0 to 2:\n\n- **0:** The metric sparks meetings, not action. No owner, or the “owner” can’t actually change anything.\n- **1:** There’s an intuitive reaction, but it isn’t pre-committed. People improvise.\n- **2:** There’s a pre-agreed response: owner, first moves, escalation trigger, and a rollback condition.\n\nThis is the quiet truth: actionability isn’t a property of the metric. It’s a property of the metric **plus** your operating system.\n\nA fast gut-check is the “so what” test—if you can’t finish the sentence “so what will we do differently,” the metric isn’t decision-grade: [[3]](#ref-3 \"kaushik.net — kaushik.net\")\n\n### Assign the disposition (Keep / Watch / Retire / Redesign)\n\nAdd the scores:\n\n- **4 = Keep.** It belongs on the dashboard that runs your week.\n- **3 = Keep/Watch.** Useful, but either the signal or response needs maturity.\n- **2 = Redesign.** The intent is valid; the signal or action path is broken.\n- **0–1 = Retire or Context-only.** Keep it for storytelling or research, not steering.\n\nTwo guardrails keep teams honest (and match the table rows people tend to gloss over):\n\n- **Guardrail: Notice in Time Definition.** “We look weekly” isn’t a Notice 2 if you needed to act Tuesday.\n- **Guardrail: Know What to Do Definition.** “We’ll discuss” isn’t Action 2. It’s Action 0 with better manners.\n\n### Make it real: quick scoring examples\n\n**SLA:** Many teams auto-rate SLA as “Keep.” But if you only review weekly SLA attainment, Notice is often a **1** (you learn after the miss). If you segment SLA by customer tier and ticket age *and* you have a duty manager who can reroute same day, it becomes **Notice 2 / Action 2**.\n\nSame metric name. Totally different operational value.\n\n**CSAT:** Overall CSAT is frequently **Notice 0–1** because response rates are uneven, surveys don’t cover every channel, and the sample drifts. Action is often **1** because the “response” is a monthly verbatim review that rarely changes staffing or workflows. CSAT often deserves **Redesign**—segmentation by contact reason, minimum response counts, and visible sample health—more than it deserves retirement.\n\n**Worked example: High CSAT, low resolution rate:** This is where the rubric prevents false comfort. High CSAT can coexist with low resolution if surveys hit “easy” tickets, or if customers are polite but still stuck. Score them separately, then ask: would low resolution be noticed early, and do we know what to do (capacity shift, engineering escalation, macro changes)? The “conflict” is the point: it forces investigation instead of letting one shiny number win.\n\n## Make ‘Would we notice?’ true: designing signals that catch problems before staffing and escalations get ugly\n\nThe Notice question isn’t “do we have a number?” It’s “do we have a signal that shows up early enough to matter?”\n\nSupport operations is full of slow-motion failures. The queue looks fine until tail risk shows up, and then the only options left are heroics.\n\n### What breaks first: leading indicators beat end-of-week averages\n\nAverages are comforting. They also hide pain.\n\nConcrete example: average first response time across all tickets stays at 2 hours. Looks great. Meanwhile, P1 tickets are waiting 6 hours, but there are fewer of them, so the average stays polite.\n\nCustomers don’t experience your average. They experience the worst slice of your week.\n\nLeading indicators tend to crack before CSAT collapses:\n\n- Backlog **age**, especially for high priority and VIP tiers\n- Queue mix (share of P1/P2)\n- **Reopen rate**, which often rises before complaints do\n- **Escalation rate**, which is blunt—but early—when trust is cracking\n\nIf you’re choosing between “count” and “age,” pick age for operational control. Counts can rise for healthy reasons (growth, launches). Age rarely rises for healthy reasons.\n\nA useful anchor: watch the *percent of high-priority tickets older than X hours* for two checks in a row. It’s harder to ignore, and it maps directly to customer harm.\n\n### Thresholds vs baselines: avoid false calm and constant panic\n\nStatic targets are seductive: “Keep first response time under 2 hours.”\n\nSupport demand isn’t static. Some weeks you ship a feature. Some weeks a partner breaks. Some weeks vacation and channel mix do their thing.\n\nStatic thresholds create two failure modes:\n\n- **False calm:** metric stays green while risk accumulates in one segment.\n- **Constant panic:** metric is always red, training everyone to ignore it.\n\nYou don’t need fancy math to fix this. You need a shared definition of “meaningful change.” A simple approach: treat movement as meaningful when it shifts by a clear absolute amount *or* a clear percentage versus a recent baseline, and stays there for more than one check cycle.\n\nDecision rule: decide whether the metric is meant to catch **spikes** or **trends**.\n\n- Spikes want faster checks and smaller tolerances.\n- Trends want slower confirmation and stronger segmentation.\n\n### Segment or lie: channel mix, priority, and region hide the story\n\n“Overall” is where operational truth goes to get blended into nonsense.\n\nCommon scenario: overall SLA is green. Email SLA is red. Chat SLA is very green. You staffed heavily for chat because it’s loud and real-time; email quietly absorbs the pain. The average tells a story of success while one segment waits.\n\nStart with segments that actually change decisions:\n\n- Priority/severity (P1/P2 behave like a different business)\n- Customer tier (paid and free users shouldn’t be blended)\n- Channel (chat, email, phone have different constraints)\n- Region/language (follow-the-sun makes global averages misleading)\n\nTradeoff warning: segmentation can create alert fatigue if you turn every slice into an alarm.\n\nDecision rule: alert on slices with distinct owners or distinct levers. If nobody can act on “APAC email,” don’t ship alerts for it. Fix decision rights first.\n\n## Make ‘Would we know what to do?’ true: mapping metric movement to staffing, automation, and escalation playbooks\n\nLots of dashboards can tell you something changed. The hard part is getting the org to move quickly and consistently without waiting for the weekly meeting.\n\nSmart people still improvise under pressure. Pre-commitment is what saves you.\n\n### From metric to mechanism: capacity, workflow, demand shaping\n\nAlmost every support metric maps to three mechanisms:\n\n- **Capacity:** staffing, schedules, queue assignments, on-call rotations\n- **Workflow:** routing, triage, escalation criteria, macros, handoffs, tooling friction\n- **Demand shaping:** help center quality, deflection/self-serve, in-product guidance, status comms, and product fixes\n\nA metric becomes actionable when you can say which mechanism it should pull and who owns that mechanism.\n\nConcrete anchor: first response time rising in chat is usually capacity or routing—hours-to-fix levers. CSAT drifting down for one contact reason is often workflow or product—days-to-weeks levers.\n\nIf you treat both as “weekly KPIs,” you’ll respond too slowly to one and too noisily to the other.\n\n### Pre-commit actions: small, medium, big (and reversible vs slow)\n\nGood playbooks don’t treat every wobble as a crisis, and they don’t wait until the metric is catastrophic.\n\nIn human terms, an action ladder for channel-level first response time might look like:\n\n- **Small move (early drift):** reroute one or two agents next shift, pause non-urgent internal work, post a heads-up with expectations.\n- **Medium move (clear miss):** add short-term coverage for a couple days, simplify triage temporarily, enable a pre-approved deflection assist.\n- **Big move (sustained drift):** escalate staffing tradeoffs, revisit forecast assumptions, open a cross-functional issue if demand is defect-driven.\n\nSplit actions into **reversible** versus **slow**. Reversible actions (reroutes, queue caps, temporary macros, status messaging) are ideal early because you can roll them back. Slow actions (hiring, policy changes, major tooling shifts) need sustained signals.\n\nThis is where teams get burned: playbooks that assume perfect conditions (“we’ll just add headcount”) instead of the levers you can actually pull this week.\n\n### Ownership and decision rights: who can act without a meeting\n\nThe fastest way to make a metric non-actionable is to require a meeting before anyone can touch the levers.\n\nThe two question test forces the uncomfortable conversation: who owns this metric, and do they have permission to act?\n\nIf the answer is “support leadership,” that can work—*if* there’s a duty owner empowered during the week. If the answer is “we all do,” the real answer is “nobody does.”\n\nA lightweight playbook shape (short enough to be used) is:\n\nTrigger → Owner → First moves → Escalate when → Roll back when.\n\nTwo anchors:\n\n- **SLA by customer tier:** Trigger on VIP SLA or VIP backlog age crossing a threshold for two checks. Owner is the duty manager. First moves are rerouting top performers, pausing low-priority work, notifying the account team with a recovery estimate. Escalate if it stays red for a day or escalations spike. Roll back after a day of normal backlog age.\n\n- **QA score:** Trigger on a drop that persists across reviewers or concentrates in one queue. Owner is the QA lead. First moves are calibration, spot audit, and one coaching theme. Escalate when it maps to a policy/product change, not just reviewer drift. Roll back after calibration shows scoring noise.\n\nStaffing lead time changes what you can promise. If schedules take 1–2 weeks to adjust, a weekly SLA report is not a staffing lever. Use faster operational signals (age-based backlog, channel-level response time) for staffing. Use slower outcomes (CSAT) as demand-shaping and product feedback.\n\nA complementary lens says the same thing plainly: metrics should change decisions, not fill slides. [[4]](#ref-4 \"rogerwong.me — rogerwong.me\")\n\n## Failure modes: how the test can still fool you (and how to catch it before a bad decision)\n\nEven if a metric passes the two question test for metrics, it can still mislead you.\n\nSupport metrics live in a world of incentives, sampling quirks, and shifting customer behavior. The goal isn’t cynicism. It’s avoiding confident wrongness.\n\n### Goodhart and gaming: when “better” isn’t better\n\nClassic failure: average handle time drops because agents rush or avoid complex tickets. The metric improves; reopen rate and escalations rise.\n\nCatch it with a paired sanity check. If handle time improves while reopen rate worsens, you didn’t get more efficient—you pushed pain downstream.\n\nResponse: demote handle time from a target to context, and read it alongside reopen rate by contact reason.\n\nAnother failure: deflection rises because you count sessions that don’t create a ticket—even when the customer failed and left. It looks like success; backlog still rises.\n\nCatch it by pairing deflection with contact volume for the same intent. If “deflection” is up and contacts are also up, you’re probably measuring browsing, not resolution.\n\nResponse: redesign deflection toward “solved intent,” not “no ticket created,” and tie it to a time window that reflects customer behavior.\n\n### Measurement trust issues: sampling bias and shifting contact reasons\n\nCSAT can look stable because the sample changed: new channels, different survey triggers, response rate swings, or only happy customers replying.\n\nCatch it by making sample health impossible to ignore: minimum response counts, response rates, and segmentation by channel and contact reason. If trust is compromised, treat CSAT as Watch until survey design and coverage are repaired.\n\nQA can drift because reviewers change or the rubric shifts. You think quality improved; you mostly changed scoring.\n\nCatch it with periodic calibration and reviewer distribution checks. If you can’t keep the measurement stable, stop using QA score as a weekly steering lever.\n\n### Tradeoffs you must name (or the metrics will name them for you)\n\nSupport is full of real tradeoffs. Metrics get dangerous when teams pretend there is no trade.\n\n- Faster first response with lower CSAT can mean you’re stopping the clock, not solving.\n- Higher deflection with higher customer effort can mean you’re pushing people into self-serve that doesn’t help.\n- Green SLA with rising escalations often means the target is generous, segmentation is wrong, or the SLA doesn’t cover the customers who are actually upset.\n\nWhen two metrics disagree, don’t average them in your head. Pick a guardrail. Example: “We optimize response speed as long as reopen rate stays within range.” Naming the guardrail makes the tradeoff explicit—and harder to game.\n\nFor a similar argument in a different voice (and a useful reminder that dashboards don’t interpret themselves), this is worth reading: [[5]](#ref-5 \"bumbleb.co — bumbleb.co\")\n\n## Your metric scorecard for next week: decide what drives staffing, automation, and escalations—and what gets downgraded\n\nDashboards sprawl the way closets do. Nobody plans it. It just happens, one well-intentioned request at a time.\n\nThe cure isn’t a bigger dashboard. The cure is a recurring decision about what gets to drive action.\n\n### A one-page checklist to audit your current dashboard\n\nKeep this tight and operational. The point isn’t to debate definitions first; it’s to decide what runs the week.\n\n- List the 10–15 metrics that show up most often in staff meetings and exec readouts.\n- For each metric, score **Notice (0–2)** and **Action (0–2)**.\n- For each **Keep** candidate, write one decision sentence: “If this moves, we will do X within Y time.” If you can’t write it, it’s not Keep.\n- For each **Redesign** candidate, name what’s broken: signal, segmentation, timing, ownership, or response.\n- For each **Retire/Context-only** candidate, decide where it belongs instead: monthly readout, quarterly deep dive, or nowhere.\n\nCap the decision-driver dashboard. Most teams do best with **three to five** drivers. Everything else is supporting context. Teams get burned when they keep everything “just in case,” then miss the one metric that mattered.\n\n### What to do with each disposition (Keep / Watch / Retire / Redesign)\n\n- **Keep:** assign an owner with decision rights and a short playbook. If a Keep metric has no owner, that’s the first fix.\n- **Watch:** set an explicit cadence and a purpose (“we’re piloting,” “we’re validating signal quality,” “we’re building the playbook”). Watching everything is the same as watching nothing.\n- **Retire:** remove it from the dashboard that runs the week. If you’re nervous, move it to a secondary page first—just don’t give it steering-wheel real estate.\n- **Redesign:** improve one thing at a time. Most redesigns are boring (and that’s good): segment it, replace an average with a distribution, tie it to a decision threshold, add sample health.\n\nCommon mistake: redesigning by making a metric fancier instead of more decision-ready. If the owner still can’t answer “what do we do when it moves,” you didn’t redesign—you decorated.\n\n### How to review monthly: keep the dashboard from re-sprawling\n\nOnce a month, re-score your top metrics with the two question test for metrics—especially after channel mix changes, product launches, or staffing shifts.\n\nOnce a quarter, do a reset: retire at least one metric and promote at most one. Promotions require two things: a named owner and a written response.\n\nIf you want a concrete Monday plan, keep it simple: take one screenshot of your current dashboard, audit 10 metrics, pick three decision drivers tied to staffing/automation/escalations, then commit to two redesigns and one retire.\n\nBy Friday, you should have playbooks for the two metrics that most affect staffing and escalations, with triggers and owners who can act without a meeting.\n\nThat’s how you turn “we tracked it” into “we ran the week.”\n\n## Sources\n\n1. [calypso.ms](https://www.calypso.ms/en/blog/if-you-cannot-explain-the-decision-do-not-ship-the-metric-a-review-workflow-that) — calypso.ms\n2. [lospino.so](https://lospino.so/blog/the-question-before-the-number) — lospino.so\n3. [kaushik.net](https://www.kaushik.net/avinash/kill-useless-web-metrics-apply-so-what-test) — kaushik.net\n4. [rogerwong.me](https://rogerwong.me/2026/07/metrics-should-change-decisions) — rogerwong.me\n5. [bumbleb.co](https://bumbleb.co/blog/2026-07-19-nobody-asks-a-dashboard-a-second-question) — bumbleb.co\n",[39,43],{"_path":40,"path":40,"title":41,"description":42},"/en/blog/how-to-choose-a-single-source-of-truth-without-starting-a-civil-war","How to Choose a Single Source of Truth Without Starting a Civil War","A practical, human-first approach to choosing a single source of truth for support metrics without political drama. Learn how to map where dashboards diverge, standardize definitions, score candidate sources, and publish one official set of support numbers teams can actually use.",{"_path":44,"path":44,"title":45,"description":46},"/en/blog/stop-arguing-about-anecdotes-how-to-combine-field-notes-metrics-and-events-into-","Stop Arguing About Anecdotes: How to Combine Field Notes, Metrics, and Events Into One Call","A practical way to combine field notes, metrics, and events into one call so support, product, and leadership can make a single decision with decision grade evidence instead of anecdote driven whiplax",1785947701683]