İçeriğe geç
wedevit

September 28, 2026 · 10 min read · software

İlhan Buğra Aslan

How do you measure a software team's performance? DORA metrics and what not to count


You cannot tell how well a software team is doing from lines of code, commit counts or closed tickets. Those count effort, not result. The defensible version measures the delivery flow rather than the people in it, and that is what DORA's five metrics do: change lead time, deployment frequency and failed deployment recovery time show throughput, while change fail rate and deployment rework rate show instability. Read together, they answer one question well, which is how fast and how safely the team ships. Read one at a time, or broken down per developer, they quietly corrupt the thing you set out to improve.

Start with what not to measure

Lines of code, commits, pull requests opened, sprint velocity. All four make the same mistake: they count the volume of work produced without asking whether the work was any good. Make velocity a target and the points inflate within a sprint. Measure lines of code and nobody deletes anything again. This is Goodhart's law in its software form, and DORA's own guidance puts the warning right next to the metrics: turning a metric into a goal pushes teams to move the number rather than improve the system.

The reference point for this argument is the summer of 2023. McKinsey published a piece titled "Yes, you can measure software developer productivity," and two weeks later Kent Beck and Gergely Orosz published a two-part response. Their central objection was that the proposed framework measures effort and output but not outcome or impact. Knowing how much work a developer produced in a week tells you nothing about whether that work turned into revenue or customer value.

In practice, "how is the team performing" usually hides three separate questions. Are we fast enough? Is quality slipping? Are we getting anything for the money? DORA metrics answer the first two well. The third needs entirely different data, and we will get there at the end.

What the five metrics actually count

There were four metrics in 2014 and there are five now. The names and definitions changed along the way, so write down which definition you are using before you start collecting anything.

Change lead time. The time from a change being committed to version control until it is running in production. Note where the clock starts: at the commit, not at the idea. "We asked for it, it went live six weeks later" is a different measurement, often a more useful one for the product side, but it is not this one. Teams that blur the two mistake queueing time in analysis for engineering slowness.

Deployment frequency. How many deployments reach production in a given period, or the time between them.

Failed deployment recovery time. How long it takes to recover from a deployment that failed and needed immediate intervention. This one split off from the older mean time to recover in 2023, because MTTR swept in events that had nothing to do with your change, like a data centre outage. The current definition counts only failures your own deployment caused.

Change fail rate. The share of deployments that require immediate intervention after they go out: a rollback, an emergency patch, a hands-on fix to keep the system up.

Deployment rework rate. The fifth metric, added in 2024. The share of deployments that were unplanned and happened because of an incident in production. The difference from change fail rate matters: change fail rate catches deployments that blow up on the spot, rework rate catches the unplanned release you cut the next morning to close a bug a user found. A low change fail rate next to a high rework rate means the team is not having emergencies, it is just permanently cleaning up.

The first three describe throughput, the last two describe instability. Keeping them apart matters, because any single throughput number is easy to improve on its own. Turn off the tests and deployment frequency goes up.

The data is already in your systems

Before buying a tool, look at what you have. Deployment frequency falls out of your pipeline's production job history. Change lead time falls out of the timestamps on the commits included in each release. The remaining three need incident and rollback records.

That is usually where teams get stuck, and the missing piece is a definition rather than a tool. If you have not written down in one sentence what makes a deployment "failed," two engineers will classify the same event differently and the ratio you get at month end means nothing. The simplest workable rule: if a deployment needs unplanned intervention afterwards (rollback, hotfix, killing a flag), it counts as failed. Tag unplanned releases the same way. One extra field in the pipeline saves you from reconstructing the rework rate by hand six months later. Where metric definitions should live and who owns them is the subject of the reporting and single source of truth post.

On scope, DORA is explicit: apply the metrics at the application or service level, and do not aggregate across teams or applications. Averaging an accounting integration that ships weekly with a web app that ships five times a day produces a number that misrepresents both teams.

What counts as a good number?

Benchmark tables are tempting and mostly the wrong thing to chase in year one, for two reasons. Context is the first: embedded software, a regulated system and a SaaS product cannot share a release rhythm. The second is that handing down an external target as an internal mandate ("everyone deploys several times a day") is a failure mode DORA warns about by name.

The healthier route is to establish your own baseline and watch the trend. DORA's free Quick Check is enough to set a starting point. After that, what you are reading is direction rather than absolute value: did lead time shorten over the quarter, is the fail rate climbing with recent releases, is recovery time stretching on one particular service?

Here is a useful test. If the number does not put anyone in a defensive position in a meeting, you are using it correctly. "Why is our lead time 11 days?" is not an accusation, it is a search for the queue. The answer is rarely typing speed. It is pull requests waiting for review, manual steps in the release, or an approval board that meets once a week.

Speed and stability are not a trade-off

The common assumption is that a team shipping less often is being careful. The oldest and most consistent finding in DORA's research says the opposite: teams that deliver quickly also tend to run more stably, because the same practices produce both.

Batch size does most of the work. DORA's guidance on working in small batches gives a concrete threshold: any batch of code that takes longer than a week to complete and check is too big, and developers should be checking changes into trunk at least daily. A large batch means a release that is hard to roll back, a failure that is hard to diagnose and a recovery that takes longer.

Approval process does the rest. DORA's research found that external change advisory boards have a negative impact on delivery performance, and found no evidence that a more formal external review process lowers change fail rates. What works instead is peer review during development, automated testing and decent monitoring. The board's useful job is coordination and process design, not reading diffs.

Everything beyond that is familiar engineering work: building a CI/CD pipeline that a small team can maintain, getting test automation into a sensible shape, preparing zero-downtime releases and a rollback path, rolling risky changes out behind feature flags. Metrics do not replace any of that. They tell you which of it paid off.

Which way AI is pushing these numbers

The 2025 DORA report draws on close to 5,000 respondents and more than 100 hours of interviews, and the picture it paints is specific. Ninety percent of respondents use AI at work and more than 80% say it has made them more productive, while 30% report little or no trust in the code it produces. The report finds AI adoption associated with higher delivery throughput and lower delivery stability at the same time.

Its main argument is that AI works as an amplifier: it magnifies what an organisation already does well, and magnifies the mess where there is mess. DORA's May 2026 report on the return from AI investment adds two ideas worth borrowing. One is the verification tax, the cost of reviewing generated code eating into the gains early on. The other is a J-curve of value realisation, where adoption dips before it pays.

The practical translation: if you handed your team a coding assistant, do not celebrate the deployment frequency line alone. If change fail rate and rework rate climb in the same quarter, the speed you gained is being spent on fixes later. The acceptance criteria in the reviewing AI-generated code post are what keep that from happening.

What the five metrics cannot see

DORA metrics measure whether you are building the thing right. They say nothing about whether it is the right thing. Shipping a feature nobody uses three times a day looks excellent on this dashboard.

Two other views close the gap. Reliability is one: if releases are frequent but users complain the app is slow, these metrics stay quiet and your SLOs and error budget do not. Developer experience is the other. The SPACE framework, published in 2021, argues that productivity cannot be reduced to a single number and proposes five dimensions: satisfaction and well-being, performance, activity, collaboration and communication, and efficiency and flow. Its working rule is to measure at team level and pick metrics from at least three of those dimensions.

The third view is the product itself: usage, conversion, support volume, renewals. If delivery metrics improve while those stay flat, the bottleneck was never engineering speed. It was what you chose to build, which is a scoping problem (what belongs in the first release).

Three easy ways to ruin the numbers

The first is a per-developer dashboard. Hand out commits, PRs or lead time by individual and behaviour changes on day one: nobody takes on a large refactor, nobody spends an afternoon reviewing someone else's branch, pull requests get split artificially. Keep the measurement at team level.

The second is ranking teams against each other. Different products, architectures and risk profiles do not line up side by side. If you want a comparison, compare a team to its own history.

The third is attaching the numbers to bonuses, performance reviews or supplier penalties. Put a target on deployment frequency and you get empty deployments. Put a target on change fail rate and incidents stop being logged. That is another reason to read both instability metrics together: rework rate still surfaces some of what never made it into the incident log.

If someone else writes your software

These metrics work in a supplier relationship too, but they belong in the contract as a visibility clause, not a threshold. What you want is not a promise of three releases a week. It is access to deployment records, incident records and the pipeline, five metrics reported monthly against the same definitions, and a short written review after an outage.

That access doubles as handover insurance: when the records and the pipeline sit on your side, a change of team does not erase the measurement or its history. What else belongs in writing is covered in the maintenance and support contract post, and how risk splits between the parties in fixed price versus time and materials.

One more caveat. If your supplier's lead time is long, a good share of it is probably yours: decisions waiting for approval, a test environment delivered late, a steering meeting that happens weekly. The metric is not a report card on one party, it shows a queue both parties are standing in.

Three steps that fit in a week

One: put the last 90 days of production releases into a table with the date, the service, the timestamp of the earliest commit included, and whether anything was rolled back. Three of the five metrics come straight out of that table, and in most teams the real surprise is how the releases cluster on the calendar.

Two: write one sentence each defining a failed deployment and an unplanned deployment, then tag both in the pipeline. Data collected without a definition has to be reinterpreted three months later.

Three: pick a single service and just watch it for a quarter without setting a target. A goal set before a baseline exists produces number management, not improvement.

At Wedevit we reconstruct a team's delivery flow from records that already exist, write the definitions down, add the measurement points to the pipeline, and turn the result into a map of bottlenecks rather than a scorecard on people. A concrete place to start: when was the oldest commit in your last release written, and how long does it take you to find that out?


Need help with this topic?

get in touch →← all posts