Agility9 min read

Measuring for Decisions, Not Judgment

Most measurement conversations quietly turn into performance conversations. That's the moment a metric stops helping — people start optimizing the number instead of the work it was supposed to describe.

Two different reasons to look at a number

Measure for a decision

  • Where is the problem, and how big is it?
  • Did the change produce the effect we expected?
  • Which work is aging or getting stuck?
  • What's the safest next experiment?

Measure for judgment (avoid this)

  • Ranking individuals
  • Comparing teams with one number
  • Treating the metric as the goal itself
  • Optimizing local output over system health

Whenever a metric becomes a reward or punishment target, people stop optimizing the system and start optimizing the number. The test for any metric you're about to introduce: name the exact decision it's meant to change.

Here's how the slide from decision to judgment usually happens in practice. A support manager starts sharing each agent's average first-response time in a weekly table — the intent is good, faster responses for customers. Within two months the average first-response time really does drop. But customer satisfaction scores drop too. What happened: agents learned to close out complex requests with a fast, empty reply like “looking into it, will follow up soon,” then push the real resolution later. The number the manager was watching looked great. The thing it was supposed to represent — customers actually getting help quickly — got worse. The metric wasn't wrong; using it for judgment turned it into a target, and the target got gamed.

Four lenses, not one number

Optimizing any one of these alone tends to quietly break another. Speed without quality produces rework. Satisfaction without flow visibility hides a growing backlog. A balanced view watches all four together, and makes the trade-offs between them visible instead of accidental.

Value

User outcome, business or risk impact, adoption.

Flow

Cycle time, throughput, work in progress.

Quality

Escaped defects, rework, technical health.

Team health

Focus, trust, sustainable pace.

A concrete version of this: near the end of a sprint, a team relaxes its code review standard to close more items — flow speeds up, throughput climbs, the burndown looks great. Three weeks later, production has three incidents back to back, and the time spent firefighting them nearly stalls new feature work. They optimized flow without watching quality; once quality broke, flow broke with it. Had they been watching both lenses together, the same week's rise in throughput next to a rise in escaped defects would have been visible immediately — early enough to tighten the WIP limit or hold the review bar and simply reset the speed expectation, instead of finding out three weeks later.

This builds on the output–outcome–impact chain

If you haven't already, it's worth reading Output vs Outcome vs Impact first — the "value" lens above is exactly that chain applied to a metric. An output number (shipped, closed, deployed) is necessary but never sufficient on its own; it has to be read alongside an outcome or impact signal before it means anything about value.

Frequently asked questions

Isn't velocity a metric too — why isn't it a decision metric?

Velocity can be a decision metric inside one team's own capacity planning. It stops being one the moment it's used to rank teams or set a target — at that point people start optimizing the number instead of the work.

Tracking four lenses at once sounds complicated — is it?

It doesn't mean four dashboards. It means before you read any single number, you ask what it might be trading off against. In practice this is usually one or two signals per lens, reviewed together, not in isolation.

Where do I even start choosing metrics?

Start from the decision, not the data you already have. Ask 'what decision am I actually trying to make better?' first, then pick the smallest signal that would change that decision.

Cookie policy

See privacy policy