When “good-looking” support metrics are actually a trap (and how the Decision Audit changes the game)
You know the week I mean. The dashboard looks clean, leadership is happy, and you can finally breathe. CSAT ticked up, AHT dropped, and deflection is “working.” Then Friday arrives with a spike in escalations, your backlog swells in the one queue nobody watches, and a senior leader asks why customers are suddenly furious in reviews.
That is the polished noise problem: improvement without truth. Metrics are not lying to you, but they can absolutely be telling a story that only covers the parts of the support system you are currently measuring. A classic support ops scenario looks like this: AHT is down 18 percent and deflection is up, so the week gets labeled a win. Meanwhile, escalations climb because more complex cases are getting misrouted or prematurely closed, and the backlog grows in the high value segment because those customers bypass the bot and head straight to human support.
A Decision Audit is a weekly decision audit workflow that forces the team to name the assumptions sitting underneath “we are doing better.” In this context, an assumption is a belief you are treating as true in order to make or defend a decision, even though you have not recently verified it. It is not the same thing as a guess. It is the thing your plan quietly depends on.
This is also not a status meeting with nicer formatting. A status meeting lists work. A Decision Audit interrogates claims and decides what to do about uncertainty.
The promise is simple and practical. Every week, you leave with three types of outputs: Decide, Park, or Instrument. Decide means you commit to an action because the assumptions are strong enough. Park means you stop burning time on it until a specific trigger happens. Instrument means you assign the smallest measurement or check that would make next week’s decision easier. Each Instrument outcome gets tripwires, which are the specific signals that will force you to reopen the question.
Weekly is the right cadence for support because the system moves weekly. Channel mix shifts weekly. Policy changes show up weekly. Tag drift happens weekly. Waiting a month turns learning into archaeology.
For more background on why a lightweight weekly assumption ritual works, the framing in this piece is compatible with what you will find in [1].
How to capture assumptions fast: turn claims into an Assumption Ledger you can actually review weekly
Most teams do not struggle with decisions. They struggle with unspoken premises. Someone says “deflection is up, so we should invest more in the bot,” and everyone nods because it sounds reasonable. The assumptions do not get spoken, so they do not get tested, so they come back later as surprise escalations.
The move that changes everything is a simple translation pattern:
Claim → Assumptions → Expected signals → Disconfirming signals
You are not trying to predict the future perfectly. You are trying to make your assumptions reviewable. A practical rule for writing assumptions in falsifiable language is:
If X is true, we should observe Y by Z.
That last part, “by Z,” is what keeps the meeting from becoming philosophy club. Time bounds are kindness.
Here are three claim to assumption translations that show up constantly in weekly support ops decision review.
First translation: “Deflection improved after the new help center launch.”
The assumptions might be: customers are actually resolving their issue without contacting support; deflected customers are not later reopening via email; and deflection is not simply shifting volume into social or app store reviews. Expected signals: a sustained drop in new ticket arrival rate for the deflected topics, stable or improving CSAT for those topics, and no corresponding rise in “where is my order” escalations. Disconfirming signals: backlog growth in the email queue, increased repeat contacts within 7 days, or higher escalation rate for the same issue.
Second translation: “Backlog is down, so the queue is healthy.”
Assumptions: the reduction is not driven by bulk closes, policy changes, or redefining what counts as backlog. Expected signals: stable reopen rate, stable first contact resolution, and no increase in customer follow ups. Disconfirming signals: a spike in reopened tickets, higher “agent did not help” comments, or an increase in supervisor escalations.
Third translation: “CSAT is up, customers are happier.”
Assumptions: survey delivery did not change; response mix is not shifting toward easier channels; and higher CSAT is not being driven by faster but lower quality resolutions. Expected signals: stable quality audits, stable escalation rate, and stable sentiment in verbatims. Disconfirming signals: AHT down plus escalations up, or CSAT up while negative verbatims cluster around “got bounced” and “had to repeat myself.”
One “what people get wrong” moment: teams treat a win metric as proof, instead of as a prompt.
Example: AHT drops sharply and everyone celebrates. The unspoken assumption is “we got more efficient at solving problems.” But in support, AHT can drop because agents are deflecting more, closing faster, or avoiding complex cases. If escalations rise at the same time, the more honest story is “we got faster at ending conversations, not faster at resolving issues.” The Decision Audit exists to force that distinction.
Now, the minimal Assumption Ledger template. Keep it so light that people will actually use it.
You want fields that support weekly review and decision log and tripwires, not a novel:
- Claim being made. One sentence.
- Decision at stake. What will we do differently if we believe this.
- Assumptions. Two to five statements written as “If X, we should observe Y by Z.”
- Expected signals. The two or three metrics or observations that should move if the claim is true.
- Disconfirming signals. The two or three things that would make us stop and rethink.
- Confidence this week. High, medium, low.
- Owner. Who is responsible for updating the signals.
- Tripwire and next review date. What would force a reopen and when we will check again.
What to skip: long background, screenshots, and defensive rationales. The ledger is not a court deposition.
To collect assumptions async, use a simple pre meeting habit: every metric owner posts one claim and its assumptions the day before. Limit it to the top two claims that might influence leadership decisions that week. If you do nothing else, enforce that limit. Scarcity keeps the ritual sharp.
When the claim is vague, do not let it slide. “We fixed it” and “volume is down” are not claims, they are vibes.
A practical salvage pattern is to ask for the noun and the boundary. What exactly changed, and where does it apply.
If someone says “we fixed it,” ask: which customer problem, which channel, and which segment. If someone says “volume is down,” ask: new tickets or total touches, and in which category. If someone says “customers are happier,” ask: which evidence, CSAT verbatims or churn risk signals.
If you want a richer version of assumption first reviews, the spirit is similar to what [2] argues: compare what was assumed to what occurred, using the same structure each time, so learning compounds instead of resetting.
Trust checks before you believe the dashboard: coverage gaps, definition drift, and proxy traps
Support metrics are slippery because the system is alive. Channels change, policies change, customers change, and your measurement pipeline changes even when nobody thinks it did. If you do not run trust checks, you end up doing performance art with numbers.
A support metric trust checklist does not need to be heavy. It needs to be consistent. Here is a short trust check sequence you can run weekly in 10 minutes before you decide anything:
- Coverage check: what is included and excluded this week.
- Mix check: did channel, language, or segment mix shift.
- Definition check: did the metric definition change, even subtly.
- Process check: did an operational policy change move the metric.
- Proxy check: if this metric moves, do we expect customer pain to move.
- Sanity check: do two independent signals agree, even directionally.
Coverage gaps are the first quiet killer. Many “global” dashboards only include email and chat, while escalations travel via a different pathway, and your largest enterprise customers use a dedicated channel that is not normalized. If deflection is up but only in English web traffic, you may be blind to what happened in app users, non English locales, or paid tiers.
Concrete example: backlog looks flat overall, but the VIP queue is growing. Leadership hears “backlog stable,” then gets surprised by churn risk escalations. The metric was not wrong. It was incomplete.
Definition drift is the second killer. CSAT, AHT, backlog, and escalations often mean something different than last month.
Three common trust breakers that cause drift:
- Channel mix shift. You moved volume from email to chat, which changes AHT and CSAT baselines.
- Tag taxonomy changes. You renamed categories, so trend lines are not comparable.
- Survey delivery changes. You changed when or to whom CSAT surveys are sent, which can inflate scores.
Add a fourth because it bites teams constantly: reopen policy changes. If “reopen” becomes a new ticket, your resolution rate looks better and your customers feel worse.
Proxy traps are where experienced operators get humble. A proxy is a metric you use because you cannot measure the real thing easily. Proxies are fine until you forget they are proxies.
Deflection is a proxy for “customers solved it themselves.” AHT is a proxy for “efficiency.” Backlog is a proxy for “queue health.” None of these are the customer experience itself.
Here is the uncomfortable example. You launch a bot flow. Deflection climbs. AHT drops. Everyone cheers. But the underlying customer pain did not move. Customers just learned to mash “agent” faster, or they abandoned and came back angrier. Your proxy improved. Reality did not.
This is why “how to validate CSAT AHT deflection signals” matters as a weekly habit, not as a one time analytics project. You validate by checking whether the expected second order signals agree.
If deflection is real, you should see fewer tickets in the deflected topics, lower repeat contact, and stable or improving CSAT comments for those topics. If you only see deflection up and nothing else, treat it as untrusted.
Now the tradeoff: speed versus certainty.
In support ops, waiting for perfect certainty is also a decision, and it is often the wrong one. The rule I use is: if the decision is reversible and low blast radius, proceed with lower certainty, but set tight tripwires. If the decision is hard to reverse, like changing staffing model, changing a customer promise, or rolling out deflection to high value customers, you pause or instrument until the metrics are decision grade.
Decision rules that work in practice:
First, proceed when two independent signals agree and trust checks pass. Example: CSAT up and negative verbatims down, while survey delivery and channel mix are stable.
Second, pause when a trust breaker is present. Example: AHT dropped the same week you changed routing or introduced macros that close faster.
Third, instrument when the decision is blocked by a missing signal. Example: you cannot tell if deflection reduced repeat contacts, so you add a simple weekly sample or a lightweight check on repeat contact rate for deflected topics.
One more “what people get wrong” moment: teams argue about the number instead of arguing about whether the number is comparable.
Do not debate whether CSAT is 4.4 or 4.6 if you changed survey timing. Start with comparability, then interpretation.
Run the Decision Audit meeting: roles, timeboxes, and how to keep narratives from hijacking decisions
| Assignment strategy | Best for | Advantages | Risks | Recommended when |
|---|---|---|---|---|
| Pre-circulate a brief with key assumptions and data points | Informed discussion and efficient use of meeting time | Attendees come prepared. reduces time spent on background context | Reliance on attendees reading beforehand. can lead to pre-conceived notions | Decisions with significant impact or complex underlying data |
| Assign a dedicated facilitator (not the decision-maker) | Neutral moderation and process adherence | Ensures fairness. keeps the meeting on track without bias from the decision-maker | Requires a skilled facilitator. can be seen as an extra overhead | All Decision Audit meetings to ensure objective process |
| A rule for “one claim at a time” and how to park rabbit holes | Maintaining clarity and preventing tangential discussions | Reduces cognitive load. ensures each point is fully explored before moving on | Can interrupt natural conversation flow. requires strong facilitation | Any discussion where multiple ideas or issues are being raised simultaneously |
| A timeboxed agenda with explicit prompts/questions | Ensuring focus and covering all critical areas | Prevents meeting sprawl. keeps discussions on track. ensures all topics are addressed | Can feel rigid. might stifle emergent important discussions if not managed well | Every Decision Audit meeting to maintain efficiency and output |
| Examples of Decide / Park / Instrument outputs relevant to support | Clarifying expected outcomes and actions from the meeting | Provides concrete guidance. helps team understand the 'so what' of discussions | Examples might not cover all scenarios. can limit creative solutions if too prescriptive | Onboarding new team members to the Decision Audit. when outputs are unclear |
| Document all 'Parked' items in a visible backlog | Ensuring no good ideas are lost and follow-up occurs | Builds trust that ideas will be revisited. prevents re-discussion of parked items | Backlog can grow unmanageably large. requires regular review and prioritization | Any meeting where tangents are common or time is limited |
The meeting fails for one reason: stories feel satisfying. Stories also let us avoid admitting what we do not know. The Decision Audit meeting is built to keep narratives in their place, which is as hypotheses, not conclusions.
You need four roles. One person can wear two hats in a small team, but name them anyway.
Facilitator: keeps the workflow moving and enforces the rules. This should not be the decider. If the boss facilitates, everyone performs.
Metric owner: brings the claim, the data, and the current confidence.
Skeptic: the friendly breaker of assumptions. Their job is to ask “what would have to be true” and “how might this be wrong.”
Decider: the person who can commit resources or policy. Sometimes this is support leadership, sometimes it is shared with product or engineering.
A 30 to 45 minute decision audit cadence for support is enough if you do not let it become story time. The secret rule is one claim at a time. Everything else gets parked.
Here is a timeboxed agenda that works, followed by the reusable workflow table.
Start with a two minute reminder of the point: we are not here to review work, we are here to decide what we believe and what we will do.
Then take the top one or two claims from the Assumption Ledger. For each, do trust checks first, then decide, park, or instrument.
When narrative hijack starts, the facilitator needs a phrase that is polite and sharp.
Two examples you can steal:
First: “That’s plausible. What would have to be true for that story to be the correct one this week?”
Second: “Let’s pause the explanation and name the assumption. What evidence would change our mind by next Friday?”
Now the workflow table.
Four controls to call out by name because they keep the meeting from rotting.
Pre circulate a brief with key assumptions and data points.
Assign a dedicated facilitator (not the decision maker).
A rule for “one claim at a time” and how to park rabbit holes.
A timeboxed agenda with explicit prompts/questions.
How do you phrase decisions so they survive leadership review. Make them falsifiable and bounded. “We will expand deflection to billing questions for English web chat for two weeks because repeat contact is stable and CSAT verbatims improved; we will revisit if escalations rise above X or CSAT drops below Y.”
That last part is not pessimism. It is professional.
If you want to go deeper on why capturing a decision trail changes behavior, this is aligned with the thinking in [3].
Also, a small humor line you can keep in your back pocket: dashboards are like bathroom scales. They are useful, but if you stand on them eight times a day, you are not getting healthier, you are just getting anxious.
Follow-through that prevents re-learning: assign owners, set tripwires, and revisit assumptions (plus the failure modes that kill the ritual)
A Decision Audit only compounds if next week is meaningfully easier than this week. That does not happen by magic. It happens because every assumption has an owner, a due by, and a decision boundary.
Owner means one name, not a group. If it is shared, it is owned by nobody.
Due by means a date that exists before the next audit, so the result can show up in the room.
Decision boundary means you specify what the result will change. Otherwise you will collect data forever.
Tripwires are the mechanism that turns weekly reviews into a living system. A tripwire is a small set of signals that force a reopen, because they indicate your assumptions might be wrong.
You only need three to five tripwires per major decision. More than that, people ignore them.
Tripwire examples tied to support metrics, covering multiple dimensions:
Volume and arrival rate: “Reopen this decision if new ticket volume for the deflected topics rebounds to within 5 percent of baseline for two consecutive weeks.”
Quality and CSAT: “Reopen if CSAT for the affected queue drops by 0.2 or more, or if negative verbatims mentioning ‘had to contact twice’ increase week over week.”
Efficiency and AHT: “Reopen if AHT drops while first contact resolution also drops, or if transfer rate rises, suggesting we are getting faster by bouncing.”
Risk, escalations, backlog: “Reopen if escalation rate rises by 15 percent, or if backlog age over a defined threshold increases, even if total backlog count looks fine.”
A simple weekly to weekly loop keeps this from becoming paperwork. Next week, you only recheck two things:
First, any Instrument outcomes from last week. Did the measurement happen, and did it reduce uncertainty.
Second, any tripwire that fired. If none fired, you do not reopen decisions out of boredom.
Now, the six failure modes that kill Decision Audits, plus what to do instead.
Failure mode one: metric worship. The highest number wins, even if it is a proxy. Recovery move: require one expected signal and one disconfirming signal for every claim.
Failure mode two: endless instrumentation. The team keeps asking for “more data” as a way to avoid making a call. Recovery move: add a decision boundary to every instrument request. If the data will not change the decision, do not collect it.
Failure mode three: blame seeking. Audits turn into “who caused CSAT to drop.” Recovery move: move language from fault to assumptions. Ask, “What did we assume that did not hold?”
Failure mode four: all green complacency. When things look good, the team stops questioning definitions and coverage. Recovery move: run trust checks even when metrics are improving. Especially then.
Failure mode five: meeting overload. The audit becomes another recurring meeting that everyone resents. Recovery move: cap it at two claims per week and force async prep. If the prep does not happen, Park the claim.
Failure mode six: silent data changes. Tag taxonomy changes, routing changes, survey changes, but nobody notes it. Recovery move: maintain a tiny “changes log” that the facilitator reads in 30 seconds at the start. If you cannot explain a trend, assume a measurement change until proven otherwise.
Two more that are worth watching because they are sneaky.
Failure mode seven: the loudest narrative wins. Someone tells a compelling story and the room moves on. Recovery move: use the facilitator phrases. “That’s plausible. What would have to be true?” and “Name the assumption.”
Failure mode eight: the ritual becomes performative for leadership. People optimize for looking certain, not being correct. Recovery move: reward explicit uncertainty paired with tripwires. The goal is not to be right in the meeting. The goal is to be less wrong next week.
If you want a useful mental model here, I like the way [4] argues that good decisions disappear quickly unless you capture them while context is fresh. The Decision Audit is that idea applied to support metrics and assumptions, weekly.
Your next weekly audit: a 30-minute starter version you can run before the next leadership meeting
You do not need a full program to start. You need one clean rep.
The minimum viable agenda is 30 minutes and one claim. Pick the claim most likely to show up in your next leadership meeting, usually something like “deflection is working,” “AHT is improving,” or “backlog is under control.”
What to prepare asynchronously, 10 minutes each:
- The metric owner writes the claim and two to three assumptions in the “If X, we should observe Y by Z” format.
- The skeptic notes any trust breakers from the past week, like channel mix shifts, tag taxonomy changes, survey delivery changes, or reopen policy changes.
- The facilitator confirms the decision at stake and writes the three possible outcomes on the doc: Decide, Park, Instrument.
A concrete first week output example that is realistic:
Instrument: “We believe deflection reduced tickets for password reset, but we cannot validate repeat contact. Owner will bring next week a simple weekly check of repeat contacts within 7 days for that topic. Tripwire is repeat contact rising week over week.”
Decide: “We will keep the bot flow for password reset active for another two weeks because ticket arrival for that topic is down and escalations are stable.”
Park: “We are parking the ‘CSAT improved’ claim because survey timing changed and results are not comparable this week. We revisit after one clean week of stable delivery.”
What success looks like after 4 weeks should be measurable, not inspirational.
Week 1: you stop arguing about opinions and start writing assumptions. You also leave with at least one Instrument task that has an owner and due by.
Week 2: you have fewer metric definition disputes because trust breakers get surfaced early.
Week 3: you see fewer surprise escalations because tripwires force earlier rechecks.
Week 4: decisions get faster and calmer. A practical success signal is that leadership questions shift from “why did this happen” to “what did we assume and what are we watching.” Another measurable signal is fewer last minute escalations that contradict the dashboard narrative.
Monday plan, realistic version.
First action: pick one claim that is currently steering your support ops narrative and write it at the top of a doc titled “Weekly Decision Audit.”
Three priorities: (1) capture the assumptions in falsifiable language, (2) run the short support metric trust checklist before interpreting any chart, (3) end with Decide, Park, or Instrument plus one to three tripwires.
Production bar: in 30 minutes, you should produce one logged decision or one instrument task that will make next week easier. If you do not, the meeting was theater. Park it, tighten the scope, and try again with one claim and stricter facilitation.
Sources
- futureponder.com — futureponder.com
- sufficientcertainty.com — sufficientcertainty.com
- standin.co — standin.co
- kevintpayne.substack.com — kevintpayne.substack.com

