Research, signal design, and decision systems

How do we redesign our dashboards and incentives so teams can’t game good-looking activity metrics (busyness) and leadership only sees the signal?

Lucía Ferrer
Lucía Ferrer
13 min read·

Answer

Stop using raw activity counts as success metrics and remove them from compensation. Redesign dashboards around a North Star outcome, a small set of leading indicators, and a few guardrails that prevent trading quality for volume. Then add metric governance so every KPI has an owner, a definition, and a regular audit for Goodhart’s Law behavior.

Diagnose where busyness is leaking into decision making

Most organizations do not wake up and decide to reward noise. It happens because activity is easy to count, easy to explain in a meeting, and fast to improve, even when it is disconnected from customer value. This is the reporting gap between activity and actual results, and it is why “everything is green” while outcomes feel stuck (see [1]).

Here are 9 diagnostic signals that busyness metrics are driving decisions instead of outcomes.

  1. Activity rises while outcomes are flat or declining. For example, more calls but win rate and pipeline conversion do not move.

  2. You see compliance spikes right before reviews, board meetings, or bonus deadlines.

  3. Proxies diverge. The KPI says performance improved, but customer retention, defect rates, or revenue quality says otherwise.

  4. More “work” creates more downstream rework. Tickets closed goes up, reopen rate goes up too.

  5. The metric improves fastest where it is easiest to manipulate. Think of logging more touches, splitting work into smaller items, or inflating story points.

  6. Distribution looks suspicious. Median performance is stable but there is a sudden growth in perfect scores or threshold hugging.

  7. Teams optimize locally and harm adjacent teams. One group hits SLA by pushing work back upstream or downstream.

  8. The same person can produce the metric and validate it. That is an invitation to self grading.

  9. Leaders ask “how many did we do” more often than “what changed for the customer or business.”

A quick audit checklist you can run this week on your dashboards and incentive plan is below.

  1. For every KPI, write the decision it is supposed to support. If there is no decision, it is probably vanity.

  2. Identify whether the metric is an outcome, a driver of outcomes, or a mere log of activity.

  3. Check controllability. Can the team move the number without delivering value. If yes, it needs guardrails or demotion.

  4. Look for targets with cliffs. If people get paid for crossing a single line, you should expect threshold gaming.

  5. Compare KPI trends to a lagging truth metric. For sales this might be retained revenue, not just bookings.

  6. Ask who benefits if the metric is wrong. If the answer is “the metric owner,” add independence and audits.

  7. Verify the denominator. Counts without normalization, like per rep, per customer, per week, are usually noise.

  8. Review the last three “green” months. List what decisions were made because of the dashboard, and whether those decisions worked.

Common mistake moment: teams often respond by adding more KPIs to “cover the gaps.” That creates dashboard clutter and gives people more surfaces to game. Do the opposite. Reduce KPIs, make each one sharper, and add a few carefully chosen guardrails.

Redefine success: outcomes, value, and constraints (the metric design brief)

If you want signal, you need a metric design brief before you touch the dashboard. This is the forcing function that stops “we track what we can” and moves you to “we measure what matters,” with explicit constraints and tradeoffs. Calypso’s guidance on auditing and redesigning KPIs for Goodhart’s Law behavior is a useful reference point [2].

Use this template for every metric that leadership will look at.

Objective. What business problem are we trying to solve.

User or customer value. What improves for the customer when we do well.

Time horizon. How quickly should this metric respond, and what lagging metric will validate it.

Controllability boundary. What the team can directly influence versus what they can only contribute to.

Primary metric. The one number we are trying to improve.

Counter metrics or guardrails. What must not get worse while we improve the primary metric.

Data source and definition. Exact query logic, inclusion rules, and refresh cadence.

Ownership. One accountable owner for definition, quality, and interpretation.

Tip: write the “how this gets gamed” paragraph inside the brief. If you cannot describe how someone could cheat, you have not thought about it hard enough.

Build a balanced measurement system (primary + guardrails + diagnostics)

A practical model that holds up under pressure has three layers.

Layer 1 is the North Star outcome. This is the thing leadership ultimately cares about, such as retained revenue, renewal rate, time to value, or availability. It is harder to game, but slower to move.

Layer 2 is a small set of leading indicators, typically two to five. These are operational drivers that predict the outcome early enough to act. They should be directional, not a substitute for outcomes.

Layer 3 is guardrails and diagnostics. Guardrails prevent perverse optimization, like improving speed by reducing quality. Diagnostics explain what is happening, like segmentation, cohorts, and pipeline stage flow.

Keep KPI count low. If you need a composite score, make it transparent and decomposable, or you will create a scoreboard that nobody trusts. Adam Analytics makes the point plainly: dashboards fail when the KPI becomes the target and stops being a measure [3].

Use ranges instead of single point targets when the system is noisy. A target band reduces the incentive to manipulate tiny movements just to “hit the number.”

Leading Indicators (Inputs): use them for steering, not for declaring victory.

Focus on Outcome Metrics (North Star): keep leadership anchored on value, not motion.

Paired Metrics (Quantity + Quality): stop volume games by making quality visible.

Guardrail Metrics (Counter-metrics): protect what you refuse to trade away.

Dashboards that surface signal: from activity logs to decision dashboards

If your dashboard reads like an activity diary, leaders will manage what they see. The goal is a decision dashboard: it answers “are we winning, why, and what do we do next.” Pulse RevOps emphasizes building sales ops dashboards that connect rep activity to outcomes and forecasting, rather than rewarding surface level motion [4].

A clean layout that works for execs has four zones.

Executive summary panel. North Star outcome, trendline versus baseline, forecast, and a short callout of what changed since last review.

Leading indicators. Two to five drivers with context, not raw counts. Show conversion rates, time to event, or cost per outcome.

Guardrails. A small set of “must not worsen” metrics. Make them visually loud when they break.

Narrative annotations and drilldowns. A sentence or two on causes, plus the ability to segment by region, product, channel, or customer cohort.

Visual design rules that reduce noise.

Use trendlines over single numbers. A single number is easy to cherry pick.

Always show a baseline and seasonality context when relevant.

Segment aggressively. Averages hide the truth. Cohorts often reveal it.

Prefer rates and normalized measures over totals.

Make definitions visible. If people debate what the metric means, the meeting is already lost.

Example of a bad widget: “Emails sent this week.” It encourages spammy behavior and says nothing about impact.

Example of a good widget: “Incremental meetings booked per 100 targeted accounts, with show rate and downstream opportunity creation,” plus a guardrail for unsubscribe or complaint rate.

Tip: add a “decision log” panel. If a KPI moved and no decision followed, ask whether the KPI belongs on the exec view.

Hardening metrics against gaming (Goodhart proofing techniques)

You do not need perfect measurement. You need measurement that is expensive to fake and easy to validate. The Calypso article on picking signals people cannot easily fake is a solid framing for this mindset [5].

Use these techniques, with tradeoffs in mind.

Paired metrics, quantity plus quality. For support, pair tickets resolved with reopen rate and CSAT. For engineering, pair throughput with defect escape and availability.

Lagging validation. If you use leading indicators, explicitly validate them against outcomes monthly or quarterly. If correlation breaks, downgrade the indicator.

Cohorts and time to event. Measure whether improvements sustain over time, not just in the current week. “Time to first value” is harder to game than “number of onboarding calls.”

Cost per outcome. Add an efficiency lens, like cost per retained customer, cost per qualified pipeline dollar, or hours per shipped feature that sticks.

Normalized denominators. Move from “tickets closed” to “tickets closed per active customer,” or “per agent hour,” so volume changes do not masquerade as performance.

Sampling and spot checks. Randomly audit a small sample for quality and data integrity. Scrapes.us discusses how to instrument without harm and avoid incentives that push people into performative behaviors [6].

Anomaly detection as a governance tool. You do not need fancy math to start. Flag sudden spikes, threshold hugging, and distribution shifts, then ask for explanation.

Tradeoff to name out loud: adding guardrails increases complexity. The art is selecting guardrails that protect real risk, not building a museum of metrics.

One light joke, because it is true: counting keystrokes is like judging a restaurant by how many plates it washes. Busy, yes. Better, not necessarily.

Incentive redesign: rewarding outcomes without creating perverse incentives

If you tie compensation to a metric, assume it will be optimized, sometimes creatively. The safest rule is simple: do not pay people on raw activity logs. Use activity for coaching and capacity planning, not for bonuses.

Patterns that tend to work.

Shared outcome pools. Pay a meaningful portion based on team outcomes, like retained revenue, customer health, or shipped value. This reduces internal gaming and encourages collaboration.

Threshold plus slope, no cliffs. Avoid “hit 100 and get paid, hit 99 and get nothing.” Use a continuous curve, so there is less reason to manipulate timing.

Balanced scorecards with caps. If you must use multiple metrics, cap the influence of any single one and include a quality gate.

Team based plus individual mix. Individual incentives can drive ownership, but keep them tied to outcome metrics the individual can influence. Use guardrails at the team level to prevent local optimization.

Deferred components. For roles where quality shows up later, defer part of payout until lagging indicators validate the outcome. This is particularly useful in sales and customer success.

Practical tip: run incentives in parallel for one cycle. Shadow calculate payouts using the new plan while paying the old one, then review who would have been overpaid for activity and who would have been underpaid for value.

Operating cadence and governance: keeping metrics honest over time

Metrics degrade. People learn the system, the market changes, and what used to be a signal becomes a target. So treat metrics as a product with ongoing stewardship.

A workable cadence.

Monthly metric council. Review North Star outcomes, leading indicators, and guardrails. Ask what is becoming gameable and what is drifting.

Quarterly KPI rotation review. Retire metrics that no longer predict outcomes or that cause perverse behaviors. Add new diagnostics sparingly.

Pre mortems for new metrics. Before rollout, ask “how will this be gamed” and “what might it break.”

Post mortems after misses. When outcomes miss, audit whether the leading indicators gave early warning or provided false comfort.

Simple RACI for metric change control.

Responsible: analytics or ops team maintains definitions and dashboards.

Accountable: functional leader owns the metric brief and behavior impact.

Consulted: finance, HR, and adjacent functions affected by incentives.

Informed: executives and managers who consume the dashboard.

Function by function examples (swap busyness metrics for outcome systems)

Below are examples across six common functions. The pattern is consistent: demote activity, promote outcomes, and add guardrails.

Sales.

Gamed busyness metric: calls made, emails sent.

Outcome: qualified pipeline created that converts to retained revenue.

Leading indicators: connect rate, meeting to opportunity conversion, stage progression velocity.

Guardrails: win rate, discount rate, churn of sold accounts, forecast accuracy.

Marketing.

Gamed busyness metric: content pieces published, impressions.

Outcome: incremental revenue or qualified pipeline influenced.

Leading indicators: conversion rate by channel, cost per qualified lead, cohort retention of acquired customers.

Guardrails: brand complaint rate, unsubscribe rate, lead to close rate quality.

Customer support.

Gamed busyness metric: tickets closed.

Outcome: time to resolution with customer satisfaction and low repeat contact.

Leading indicators: first response time, backlog age distribution.

Guardrails: reopen rate, escalation rate, CSAT, defect creation downstream.

Engineering.

Gamed busyness metric: story points, commits, lines of code.

Outcome: reliable delivery of valuable changes.

Leading indicators: cycle time, deployment frequency where relevant.

Guardrails: defect escape rate, change failure rate, availability, incident frequency.

HR and recruiting.

Gamed busyness metric: interviews scheduled.

Outcome: quality hires who stay and perform.

Leading indicators: time to shortlist, candidate experience.

Guardrails: first year attrition, hiring manager satisfaction, diversity and fairness checks.

Finance and procurement.

Gamed busyness metric: number of vendor negotiations.

Outcome: total cost of ownership reduction with service levels maintained.

Leading indicators: contract cycle time, compliance rate.

Guardrails: supplier performance, risk exposure, downtime from cost cutting.

Change management: adoption, data trust, and cultural reset

Redesigning metrics is political because it changes status, rewards, and narratives. If you do it like a surprise audit, people will resist. If you do it like a shared upgrade to decision quality, adoption follows.

A rollout plan that typically works.

Pilot with one or two teams. Pick an area with clear outcomes and reasonable data quality.

Parallel run. For four to eight weeks, keep the old dashboard but review the new one in leadership meetings. Track where they disagree and why.

Calibration period for targets. Use historical ranges and adjust slowly. Do not set aggressive targets until definitions are stable.

Manager training. Teach how to coach from leading indicators without paying on them. This is metric literacy, not math.

Single source of truth. Lock definitions, document them, and resist spreadsheet side quests.

Address fear directly. People often worry transparency means punishment. Make it explicit that activity metrics are for operational improvement, while outcomes and guardrails are for performance.

Practical tip: publish a one page “metric contract” for each KPI. It states purpose, owner, definition, and what actions leaders should and should not take based on it.

Executive questions that reliably separate signal from noise

When leadership asks better questions, gaming becomes harder. Here are 12 that work across functions.

  1. If this metric goes up, what specific customer or business outcome should improve, and by when.

  2. What is the lagging metric that validates this leading indicator.

  3. What guardrail tells us we are not buying speed with quality.

  4. Where can a team move this number without creating value.

  5. What changed in behavior after we set the target.

  6. Show me the distribution, not the average. Who is improving and who is not.

  7. How does this look by cohort. Do improvements persist for customers acquired or onboarded in the same period.

  8. Are we seeing threshold hugging or end of period spikes.

  9. What did we stop doing to make this number go up, and what did that cost.

  10. What is the cost per outcome, and is it improving.

  11. If we removed this metric from compensation, would we still want to track it.

  12. What decision will we make differently if this number is red next week.

If you do only one thing first, do this: pick one North Star outcome per function, add two to four leading indicators, add two guardrails, then remove all raw activity counts from compensation. That single change shifts the organization from looking busy to getting better, which is the entire point.

Option Best for What you gain What you risk Choose if
Leading Indicators (Inputs) Operational teams, short-term adjustments, forecasting Early warning signals, ability to course-correct quickly Can be gamed if not tied to lagging outcomes, may not reflect true impact You need to understand drivers of future performance and empower teams.
Cohort Analysis Understanding user behavior over time, impact of changes Reveals true retention/engagement, isolates effects of interventions More complex data setup, requires consistent tracking You need to see how different groups perform or react over their lifecycle.
Focus on Outcome Metrics (North Star) Strategic alignment, long-term vision Clear direction, reduced gaming, true impact measurement Slow feedback, difficulty attributing short-term actions You need to measure ultimate business value, not just activity.
Paired Metrics (Quantity + Quality) Balancing output with standards (e.g., sales calls + conversion rate) Discourages gaming one metric at the expense of another Can be complex to track, requires clear definitions for both You suspect teams are prioritizing volume over value or vice-versa.
Guardrail Metrics (Counter-metrics) Preventing unintended negative consequences Protects against optimizing one area at the cost of another — e.g., speed vs. quality Can add dashboard clutter, requires careful selection You have critical non-negotiables or potential for perverse incentives.
Random Audits & Spot Checks Validating data integrity, deterring metric manipulation Increased data trustworthiness, reinforces ethical behavior Resource intensive, can feel punitive if not framed correctly You observe suspicious metric spikes or need to ensure compliance.

Sources


Last updated: 2026-07-02 | Calypso

Sources

  1. webresults.io — webresults.io
  2. calypso.ms — calypso.ms
  3. adam-analytics.com — adam-analytics.com
  4. pulserevops.com — pulserevops.com
  5. calypso.ms — calypso.ms
  6. scrapes.us — scrapes.us

Tags

signal-vs-noise-why-organizations-misread-data