1. Definition and why it matters
Metrics make engineering legible and improvable, and the wrong metrics actively harm. v1.0 framed the competency as setting and using meaningful measures; v3.0 adds interpretation and the measurement of AI impact, because the pressure to show a number is now strongest exactly where the numbers are least settled. The competency is choosing a small balanced set of measures that describe the delivery system rather than the people in it, reading them as trends and conversation starters rather than verdicts, refusing the uses that corrupt them — individual productivity scores, cross-team leaderboards, a single number for the board — and, at executive scope, declaring in advance how the organisation will know whether its strategy is working. It matters because every metric becomes a target the moment it is attached to a reward or a comparison, and because unmeasured growth is indistinguishable from bloat. The examinations test it through the productivity score per engineer, the league table of teams, the metric that improved while nothing else did, and the AI investment that must show a return the honest number does not yet support.
2. Core principles
- Measure the system, not the people. The validated delivery measures describe the team's delivery system — throughput and stability — and are never individual productivity scores.
- A measure that becomes a target stops measuring. Activity metrics — lines of code, commit counts, story points per engineer — are the easiest to game and the most destructive to weaponise. Anticipate the gaming before attaching a target.
- Small, balanced, and paired with context. A few measures across throughput and stability so that improving one does not silently wreck another; numbers alongside the qualitative signals that explain them.
- Trend a team against itself; never rank teams against each other. Contexts differ, story points are not a currency, and a ranking teaches every team to optimise the number.
- Instrument the system between teams. Above team scope, what a manager uniquely owns is the queues, handoffs and dependencies between teams; measure those wholesale rather than each team's internals retail.
- Declare the success measure when the strategy is set. A strategy with no falsifiable measure cannot be shown to be failing, and its "success" will be judged by effort expended. Report AI's effect the same way: honestly, with uncertainty, including where the effect is negative.
3. Models and evidence
Measurement is the area of this standard with the most research and the most misuse of it, and the unit is careful to grade both.
DORA metrics research
The four delivery measures validated by the research programme in Nicole Forsgren, Jez Humble and Gene Kim, Accelerate: The Science of Lean Software and DevOps (2018) — deployment frequency and lead time for changes (throughput), change failure rate and time to restore service (stability) — and the finding that they move together in high-performing organisations rather than trading off. They are deliberately team-level measures of the delivery system, meant for trending a team against itself. Used as individual measures or as a cross-team leaderboard, they become targets and the research no longer applies.
The SPACE framework practice
The framework in Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Thomas Zimmermann, Brian Houck and Jenna Butler, The SPACE of Developer Productivity: There's More to It Than You Think (2021), from several of the same researchers, arguing that developer productivity cannot be captured by a single metric and proposing five dimensions — satisfaction and wellbeing, performance, activity, communication and collaboration, efficiency and flow — to be measured in balance. Its practical use is as a checklist against one-dimensional measurement: any proposed productivity metric that sits in only one dimension, and especially in activity alone, is incomplete by construction. It is a framework paper rather than an empirical finding and is graded as such.
Goodhart's law practice
When a measure becomes a target, it ceases to be a good measure, in the form given by Marilyn Strathern, 'Improving Ratings': Audit in the British University System (1997). The governing hazard of everything in this unit, at every scope: the engineer who games a personal metric, the team that optimises a leaderboard, the organisation whose activity dashboard is healthy while outcomes are flat. It is not an argument against measuring; it is an argument for balanced sets, outcome-anchored measures, and metrics treated as conversation starters rather than verdicts.
Measuring AI impact contested
The evidence introduced in DE-5 Designing how work flows: a controlled experiment, Sida Peng, Eirini Kalliamvakou, Peter Cihon and Mert Demirer, The Impact of AI on Developer Productivity: Evidence from GitHub Copilot (2023), found a defined task completed substantially faster with an AI assistant; a randomised trial, Joel Becker, Nate Rush, Elizabeth Barnes and David Rein (METR), Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (2025), found experienced developers slower in their own repositories while believing they were faster. For measurement the lesson is specific: self-reported speed-up is not a measure, adoption and licence counts are not measures, and pull-request counts are activity. The effect is measured in the organisation's own conditions, against outcomes, with the uncertainty stated, and the honest report to a board may be "not yet, and here is why".
The falsifiable scorecard practice
Not a named model but the executive practice this unit treats as the answer to "is the strategy working": a small, stable scorecard declared when the strategy is set — the business outcomes the investment was meant to move, leading indicators chosen in advance, delivery-system trends at organisation level, and organisational health — reviewed on a cadence, revised when the strategy is revised, and capable of proving the strategy wrong. Its failure is the dashboard that swells until nobody reads it, and the board meeting at which "working" turns out never to have been defined.
4. Practice
The team's four numbers
Deployment frequency, lead time, change failure rate and time to restore, trended monthly for the team against its own history, and read in the retrospective alongside what the team says happened. The numbers find problems; the conversation explains them. Nobody's name appears next to any of them.
The gaming pre-mortem
Before any metric is attached to a target, a reward or a comparison, the manager asks: how would a reasonable person hit this number without doing the thing it stands for? If the answer is easy, the target is withdrawn or the measure is paired with one that would expose the gaming.
The between-teams instruments
For a group of teams, measures of the system between them: queue times at shared dependencies, handoff latency, how long cross-team requests wait. These are the manager of managers' own instruments, and they are where capacity drains and platform decay show up before delivery does.
Measuring AI in the team's conditions
An AI-assistance rollout is measured against outcomes — cycle time, change failure rate, rework — in the organisation's own codebases over a period long enough to see effects on understanding as well as speed, with self-reports collected but not treated as the measure. The report states the uncertainty.
The scorecard, declared in advance
When a strategy or a major investment is set, the executive writes down how the organisation will know whether it is working: the outcomes, the leading indicators, the delivery and organisational health trends, and the date of the first review. The scorecard is one page and it is the page the board sees.
5. Scaling note
At team scope the object is one team's metrics and what they hide, and the manager measures the delivery system, pairs throughput with stability, and refuses individual scores. At organisational scope the temptation the manager never faced before is comparing teams; the object becomes measures that aggregate or compare without inviting the gaming that comparison creates — teams trended against themselves, the system between teams instrumented — and measured evidence carried into every capacity argument. At executive scope the question is whether the strategy is working: the measures the business uses to judge engineering, chosen so that the organisation is not made worse by hitting them, a falsifiable scorecard declared in advance, and AI's effect reported honestly to the board. The pattern is in How Judgment Scales; the flow the measures describe is DE-5 Designing how work flows.
6. Judgment
- Measures team-level delivery outcomes and pairs throughput with stability. team
- Uses metrics to find problems, not to rank people. team
- Anticipates gaming before attaching a target, and stays alert to it afterward.
- Declines to use a single productivity number for individuals, and can say why.
- Failure mode — individual output as "productivity". team
- Failure mode — chasing one metric until another breaks. team
- Failure mode — metrics as surveillance. team
- Failure mode — numbers without context. team
- Trends each team against itself — cycle time and its direction, predictability, change failure rate — and never leaderboards them. org
- Instruments the system between teams: queues, handoffs, dependencies. org
- Carries measured evidence into every capacity, drain and platform-decay argument. org
- Measures the effect of AI assistance honestly across teams, in their own conditions, including where it is not helping. org
- Failure mode — cross-team velocity leaderboards. org
- Failure mode — normalising story points into a fake currency. org
- Failure mode — measuring team internals while the queues between them go dark. org
- Failure mode — an AI rollout whose only moved number is pull-request count. org
- Declares the strategy's success criteria when the strategy is set, and keeps the scorecard small and stable. exec
- Chooses the measures by which the business judges engineering, and refuses measures that would make the organisation worse to hit. exec
- Reports AI's effect to the board with uncertainty, including a negative or not-yet result. exec
- Reads growth without measured outcomes as the bloat it is indistinguishable from, and fixes the measurement rather than the narrative. exec
- Failure mode — declaring victory by effort expended. exec
- Failure mode — a strategy with no falsifiable success measure. exec
- Failure mode — dashboard sprawl. exec
- Failure mode — a single productivity number for the board. exec
7. Tensions
Legibility versus distortion. Every measure that makes engineering legible to the business also invites the organisation to optimise it. The resolution is a balanced set and the gaming pre-mortem, not the absence of measurement.
Comparison versus context. Leadership wants to compare teams, and comparison is what turns a measure into a target. Trending each team against itself gives leadership what it needs — is this team improving — without the leaderboard.
Evidence versus timing. The honest measurement of an AI investment takes longer than the quarter in which the board wants a return. The judgment is in reporting the uncertainty rather than the convenient number, and in setting the review date when the investment is made.
Simplicity versus completeness. A single number is what executives ask for and what corrupts fastest; a full dashboard is complete and unread. The one-page scorecard, declared in advance, is the compromise that survives both.
Measuring versus trusting. A manager who instruments everything signals distrust; one who instruments nothing manages by anecdote. The line is the system versus the people: measure the delivery system thoroughly, and never the individual.
8. Worked scenario
A senior manager who leads six teams is asked by their vice-president to produce a league table of the teams by velocity, to be reviewed monthly, so that leadership can see which teams are performing. Two of the teams work on a mature product with heavy compliance requirements; two on a new product with none; two are platform teams whose output is consumed by the other four. The vice-president's intent is reasonable — they cannot tell which teams are in trouble — and the request is specific.
Producing the table is the easy answer and the one that damages all six teams. Story points are not a currency across teams; the contexts differ so much that the ranking would measure the product, not the team; and every team would learn within a month to optimise the number, which is the one thing leadership does not want them doing. Refusing the request is the other wrong answer: it leaves the vice-president unable to see what they need to see and reads as engineering resisting accountability.
The senior manager answers the need rather than the request. They propose, and build within a month, a view that gives leadership what it actually wants: each team trended against itself on cycle time, predictability (committed against delivered) and change failure rate, so that a team whose cycle time is rising or whose predictability is falling is visible without being ranked against a team in a different context. Alongside it, the measures of the system between the teams — how long requests to the platform teams wait, how often cross-team dependencies slip — because that is where the trouble the vice-president suspects usually lives, and it is the part of the system the senior manager uniquely owns.
They also show the vice-president what the league table would have done, using two months of history: the "worst" team by velocity was the compliance team, whose cycle time was flat and whose change failure rate was the lowest of the six, and the "best" was a new-product team whose predictability had collapsed. The table would have rewarded the team in trouble and punished the one performing well.
The vice-president accepts the trended view. What the senior manager does not do is produce a ranking to be seen to comply, or refuse and leave the need unmet.
9. Related competencies
- DE-5 Designing how work flows — the flow the measures describe; Goodhart's law in the operating cadence.
- DE-3 Reliability, incident response, and operational ownership — the stability measures and the research on throughput and stability together.
- SV-1 Engineering strategy — the strategy whose success measure is declared when it is set.
- SV-2 Communicating and partnering across functions — carrying measured evidence into the conversation with the board.
- TJ-5 AI-assisted engineering — the AI adoption whose effect is being measured.
10. Self-check
- Why are the DORA metrics team-level rather than individual?
Answer
Because they describe the delivery system — throughput and stability — not a person's output, and because the research validating them applies to that use. As individual measures they become targets and stop measuring. - What question is asked before any metric is attached to a target?
Answer
How would a reasonable person hit this number without doing the thing it stands for? If the answer is easy, withdraw the target or pair the measure with one that would expose the gaming. - Why is trending a team against itself sound where ranking teams against each other is not? org
Answer
A trend answers "is this team improving" in its own context; a ranking compares contexts that differ, treats story points as a currency, and teaches every team to optimise the number. - What does a manager of managers uniquely own, and therefore measure? org
Answer
The system between teams — queues at shared dependencies, handoff latency, cross-team requests waiting — which is where capacity drains and platform decay show before delivery does. - An AI-tooling rollout must show a return and the only number that moved is pull-request count. What is the honest position? org
Answer
Pull-request count is activity, not outcome. Measure against cycle time, change failure rate and rework in the organisation's own conditions, over a period long enough to see effects on understanding, with self-reports collected but not treated as the measure — and report the uncertainty. - What does it mean for a strategy to have a falsifiable success measure, and what happens at the board when it does not? exec
Answer
Outcomes and leading indicators declared when the strategy is set, capable of showing the strategy is failing. Without them, "working" is judged by effort expended, and the board discovers that success was never defined — at which point growth is indistinguishable from bloat. - The board wants a single number for engineering productivity. What is the answer? exec
Answer
A one-page scorecard declared in advance — strategic outcomes, delivery-system trends, organisational health — instead of a single number, with the reason given: any single number would make the organisation worse to hit.
Sources
- Nicole Forsgren, Jez Humble and Gene Kim, Accelerate: The Science of Lean Software and DevOps (2018) — the DORA delivery metrics and the research behind them.
- Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Thomas Zimmermann, Brian Houck and Jenna Butler, The SPACE of Developer Productivity: There's More to It Than You Think (2021) — the SPACE framework: productivity is not one number.
- Marilyn Strathern, 'Improving Ratings': Audit in the British University System (1997) — Goodhart's law in its standard form.
- Sida Peng, Eirini Kalliamvakou, Peter Cihon and Mert Demirer, The Impact of AI on Developer Productivity: Evidence from GitHub Copilot (2023) — the controlled experiment on AI-assisted task completion.
- Joel Becker, Nate Rush, Elizabeth Barnes and David Rein (METR), Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (2025) — the randomised trial on experienced developers with AI tools.