← All posts

QA metrics that matter: a scorecard for CTOs and engineering leaders

· 4 min read

The QA metrics that matter to engineering leaders are the ones that change a decision: requirement coverage, pass rate on the release candidate, open blockers per release, escaped defects, flaky test rate, time from code-complete to release, mean time to diagnose a failure, and the share of QA time spent on maintenance. Test counts and code coverage on their own look impressive and decide nothing.

Vanity metrics vs decision metrics

Vanity metric Why it misleads Better question
Number of test cases More tests can mean more duplication Is every requirement covered?
Code coverage alone Lines executed is not behaviour verified Do tests check the expected result?
Tests executed per week Measures activity, not safety Did the release build pass?
Bugs found by QA Rewards finding over preventing How many bugs escaped to users?

The scorecard: 8 metrics

1. Requirement coverage

Definition: requirements with at least one test ÷ requirements in scope. Why it matters: an untested requirement is an unknown, and unknowns are what break releases. Watch for: coverage with only happy-path tests. Track negative and boundary coverage too.

2. Pass rate on the release candidate

Definition: passing tests ÷ executed tests, on the exact build you plan to ship. Why it matters: it is the closest single number to "is this release safe?" Watch for: a pass rate from an older build or a different environment.

3. Open blockers per release

Definition: failing or untested high-priority requirements at the go/no-go point. Why it matters: this is the number the release decision actually depends on. See the release readiness checklist.

4. Escaped defects

Definition: bugs found in production that tests should have caught, per release. Why it matters: it measures the outcome users feel. Watch for: always ask which requirement the escaped bug belonged to, and whether it had a test.

5. Flaky test rate

Definition: tests with inconsistent results on unchanged code ÷ all tests. Why it matters: flakiness destroys trust in every other metric. See flaky tests.

6. Time from code-complete to release

Definition: time between "feature done" and "feature live". Why it matters: this is where QA shows up as the critical path, or doesn't.

7. Mean time to diagnose a failure

Definition: time from a failed test to knowing whether it is a product bug, a test problem or an environment problem. Why it matters: a failure nobody can explain quickly gets ignored.

8. Maintenance share

Definition: share of QA effort spent fixing existing tests rather than covering new work. Why it matters: when maintenance dominates, coverage stops growing no matter how hard the team works.

What good looks like

There are no universal targets, and any benchmark you read should come with its source and context. Aim for trends:

  • Requirement coverage going up, especially for high-priority requirements.
  • Open blockers at go/no-go going down.
  • Escaped defects going down, and each one traced to a requirement.
  • Time from code-complete to release going down without escaped defects rising.

How to report QA to the board or leadership

Keep it to three lines per release:

  1. Scope and coverage: "42 requirements in scope, 40 tested, 2 accepted as untested (low priority)."
  2. Result: "All high-priority requirements passing on the release build; 1 medium issue accepted."
  3. Trend: escaped defects and time-to-release versus the last few releases.

(The numbers above are an illustration of the format, not a benchmark.)

Where the data comes from

Most of these metrics need one thing: every test linked to a requirement, and every result linked to a test. With those links, coverage, blockers and escaped-defect tracing fall out automatically. Without them, someone rebuilds the picture by hand in a traceability matrix before every release.

How testdart gives leaders this view

testdart links every test to its requirement from the moment it's generated, and every run writes its results back. The result is one live view of every requirement, test and result, with pass rate and a reason for every failure, so the answer to "are we ready?" is already on screen and nobody chases a QA status update. It is the reporting half of AI-native QA. See the view for engineering leaders.

FAQ

Is code coverage a useless metric? No. It is useful for finding untested code. It is misleading as a measure of whether requirements are verified.

How often should QA metrics be reviewed? Blockers and pass rate at every release; trends monthly or quarterly.

Who should own QA metrics? Engineering leadership owns the targets; QA owns the data and the explanations.

See it on your own flow.

testdart reads your requirements, writes the test cases, runs them in a real browser and shows you what broke.