Which Test Metrics Actually Matter and Which Ones Are Vanity?

Particle41 Team
October 6, 2026

Your last engineering review probably included a code coverage number. Someone reported it, a chart trended up and to the right, and everyone nodded. Coverage is the most-reported testing metric in the industry. It’s also one of the least useful, and the gap between those two facts is costing you.

Code coverage measures one thing: what percentage of your code executed while your tests ran. That’s it. It says nothing about whether those tests would actually catch a bug. You can write a test that calls a function, asserts nothing, and you’ve “covered” that code. Teams do this constantly—sometimes deliberately to hit a coverage gate, sometimes because AI generated tests optimized for coverage rather than correctness. The number goes up. The quality does not.

A metric you can game without improving quality is a vanity metric. And if your team is steering by vanity metrics, you’re optimizing for a dashboard instead of for outcomes. Let’s separate the numbers that flatter you from the numbers that tell you the truth.

Why Coverage Fools Smart Teams

Coverage is seductive because it’s easy to compute, easy to chart, and feels like rigor. But it measures execution, not verification—and that distinction is everything.

Consider what high coverage doesn’t tell you. It doesn’t tell you whether your assertions are meaningful or whether they’d notice a regression. It doesn’t tell you whether you tested the right code—you can have 95% coverage and zero coverage on the one payment path that actually matters. And in 2026 it’s actively misleading, because AI can generate tests that maximize coverage while verifying almost nothing. We’ve reviewed suites at 90%+ coverage that would have caught maybe a third of realistic bugs.

This doesn’t mean coverage is worthless. Low coverage on critical code is a genuine red flag—it tells you something is untested. But high coverage is not a green light. It’s the absence of one specific problem, not the presence of quality. Treat it as a floor to investigate, never a goal to celebrate.

The Metrics That Actually Predict Quality

The metrics worth putting on your dashboard share one trait: they measure outcomes the business feels, and you can’t game them without genuinely improving. Four matter most.

Escaped defects. The count of bugs that reached production—the ones your testing should have caught and didn’t. This is the most honest measure of whether your quality process works, because it measures failures your customers experienced. Track the trend and the severity, not just the raw count. A rising escaped-defect rate means your testing is losing ground no matter what your coverage says.

Change failure rate. The percentage of deployments that cause a failure requiring remediation—a rollback, a hotfix, an incident. This is one of the core DORA metrics for a reason: it directly measures how often shipping breaks things. In the 2025 DORA benchmarks, the top performers keep change failure rate under 5%. If yours is much higher, your quality gates are letting bad changes through, and that’s a far more actionable signal than any coverage figure.

Mean time to recovery (MTTR). When something does break, how long until you’ve restored service—what DORA now frames as failed deployment recovery time, with top teams restoring service in under an hour. Because no testing catches everything, your ability to recover fast is a quality metric in its own right. A team that recovers in minutes is in a fundamentally stronger position than one that takes hours, even at the same defect rate. MTTR measures resilience, which matters precisely because perfection is impossible.

Flake rate. The percentage of your tests that pass and fail non-deterministically. This is a meta-metric—it measures whether you can trust your other metrics. A high flake rate corrupts everything, because it trains your team to ignore red and lets real failures hide among false ones. Keep it under 1%. A trustworthy suite is the foundation that makes every other number meaningful.

How These Connect to the Business

What makes these four metrics matter isn’t that they’re technical. It’s that they map directly to outcomes a CEO understands.

Escaped defects are customer trust and support cost. Change failure rate is how confidently you can ship—and therefore how fast you can deliver value. MTTR is how much a failure actually costs you in downtime and reputation. Flake rate is how much engineering time you’re burning on false alarms and how much risk is hiding in plain sight.

Coverage maps to none of these. You can move coverage from 80% to 90% and have zero effect on a single business outcome. You can drive change failure rate from 20% to 10% and your whole organization ships more confidently. That’s the test of a real metric: when it improves, does the business get better? For the four above, yes. For coverage, not reliably.

Don’t Trade One Vanity Metric for Another

Once teams hear “coverage is vanity,” they sometimes overcorrect into new vanity metrics. A few to watch for.

Total test count. “We have 12,000 tests” tells you nothing about quality and often signals a bloated, slow, hard-to-maintain suite. More tests is not better; better tests is better.

Pass rate in isolation. A suite that’s always 100% green might be excellent or might be testing nothing of consequence. Pass rate only means something alongside escaped defects—green tests plus production bugs means your tests aren’t testing the right things.

Test execution time as a brag. Fast is good, but a suite that’s fast because it’s shallow is fast for the wrong reason. Speed matters in service of trustworthy coverage, not as an end in itself.

The pattern: any metric you can improve without improving real outcomes is a candidate for vanity. Always ask what business result moves when the number moves. If the answer is “nothing,” stop reporting it as a goal.

Build a Dashboard Worth Steering By

If you’re an engineering manager rebuilding your quality dashboard, here’s the shape we recommend.

Lead with outcome metrics: escaped defects (trend and severity), change failure rate, and MTTR. These three answer the question leadership actually cares about—is our software reliable and getting more so?

Support with health metrics: flake rate and inner-loop test speed. These tell you whether your testing process itself is healthy enough to trust.

Keep coverage as a diagnostic, not a target. Watch it for gaps—critical code with low coverage is worth investigating. But never set a coverage target, because the moment you do, your team optimizes the number instead of the quality. Goodhart’s law is undefeated: when a measure becomes a target, it ceases to be a good measure.

A useful sanity check: for any metric on your dashboard, ask whether a clever engineer could improve it without making the software better. If yes, it’s a diagnostic at best and demote it from your goals.

What Good Looks Like

The teams we work with that have the highest real quality tend to share a profile. Change failure rate in the single digits to low teens. MTTR measured in minutes because they’ve invested in observability and fast rollback. A flake rate under 1% so they trust their own suite completely. And a steadily falling escaped-defect rate because every production bug becomes a permanent test.

Notably, these teams often don’t have remarkable coverage numbers. They have coverage where it counts and don’t obsess over the global percentage. They optimized for outcomes and let the vanity metric fall where it may.

At Particle41, when we stand up a delivery team’s quality practice, we instrument the outcome metrics first and treat coverage as a diagnostic from day one. Pairing senior engineers with AI agents lets us generate test volume quickly—but we measure success by escaped defects and change failure rate, never by a coverage figure that AI could inflate on its own. The goal is software that holds up in production, and only outcome metrics tell you whether you’re getting it.

The Bottom Line

Coverage tells you your tests ran. Escaped defects, change failure rate, MTTR, and flake rate tell you your tests work. One of those is what you want to know.

Stop celebrating the number that goes up without making anything better. Build a dashboard around outcomes the business actually feels, demote coverage to a gap-finding diagnostic, and steer by the metrics you can’t game. Your software—and the people who depend on it—will be measurably better for it.

Sources

  1. Are you an Elite DevOps performer? Find out with the Four Keys Project, Google Cloud (2020)
  2. DORA Metrics: A Full Guide to Elite Performance Engineering, Multitudes (2025) — citing the 2025 DORA Report
  3. Goodhart’s law, Wikipedia