Definition of Done
Definition of Done
Definition: The Definition of Done (DoD) is a single, explicit checklist of quality criteria that every backlog item must satisfy before the team may call it complete. It applies uniformly to all work the team produces, not to any one feature, and it encodes the minimum bar required for an increment to be genuinely releasable. When the DoD is met, no hidden remainder of work is left behind; when it is not, the team has quietly created an obligation it has not recorded anywhere.
The DoD is a shared commitment, not a preference. It converts the word “done” from a subjective status report into a verifiable, binary state that anyone on the team, or outside it, can check.
How It Works
What Belongs on the List
A DoD contains criteria that are universal, verifiable, and within the team’s control. Universal means it applies to every item: a bug fix, a feature, a spike deliverable. Verifiable means an outsider could confirm it without asking the author’s opinion. Within the team’s control means the team can actually satisfy it during the sprint rather than waiting on a queue it does not own.
A typical mature list reads something like this:
| Criterion | Verification | Why it is on the list |
|---|---|---|
| Code merged to trunk | Git history | Unmerged code is inventory, not value |
| Unit and integration tests written and passing | CI pipeline green | Untested behavior is unproven behavior |
| Peer reviewed and approved | Pull request record | Catches design drift and knowledge silos |
| Static analysis and linting clean | CI gate | Prevents slow erosion of consistency |
| No new critical or high severity defects | Scanner plus manual triage | Prevents shipping known harm |
| Deployed to a staging environment | Deployment log | Proves the artifact actually builds and runs |
| Observability in place: logs, metrics, alerts | Dashboard exists | Unmonitored features fail silently |
| User-facing documentation updated | Docs commit | Undocumented features generate support load |
| Accessibility checks passed | Automated plus spot check | Legal and ethical baseline |
| Product Owner has seen it working | Verbal or recorded confirmation | Closes the interpretation gap |
What does not belong: anything specific to one feature (that is acceptance criteria), anything the team cannot influence, and anything unfalsifiable. “Code is clean” is not a DoD criterion. “Passes the agreed lint configuration with zero errors” is.
Who Owns It
The development team owns the DoD. This is the single most misunderstood ownership question in agile practice. Managers do not hand it down; the Product Owner does not set it; the team writes it, because the team is the group that must live inside it every day.
The exception is organizational floor conditions. If the enterprise mandates that all code touching payment data undergo a security review before release, that constraint enters the DoD as a non-negotiable item. The team may not remove it. The team may still add to it, and usually should. The rule of thumb: the organization sets a floor, the team sets its own ceiling above that floor, and the ceiling only ever rises.
Where multiple teams contribute to one product, an organizational DoD acts as a baseline all teams must meet, with each team free to hold a stricter local version. This matters enormously for Scaled Agile (SAFe and LeSS) settings, where an integrated increment is only as done as the weakest contributing team’s definition.
How It Is Enforced
A DoD enforced by memory and goodwill decays within two sprints. Enforcement must be mechanical wherever possible.
The strongest pattern: every criterion that can be a pipeline gate is a pipeline gate. Tests, coverage thresholds, static analysis, dependency vulnerability scans, build reproducibility, deployment to staging. These are checks the machine performs on every commit, and a red pipeline means the item cannot move. This is why DoD maturity and CI-CD maturity are effectively the same curve viewed from two angles.
Criteria that resist automation, such as “the Product Owner has seen it working” or a manual exploratory test pass, need a human ritual. The most reliable is a per-item checklist rendered directly on the ticket, with the item physically unable to enter the Done column until every box is ticked. In Kanban this maps neatly onto explicit column exit policies. In Scrum the sprint review becomes the moment where an unmet DoD is exposed publicly, which is uncomfortable by design.
The failure mode to guard against is the “done except” conversation. “It is done except the tests.” “Done except deploying it.” An item is done or it is not. There is no partial credit, because partial credit is precisely the mechanism by which unfinished work escapes the board.
How It Evolves
The DoD is a living artifact and should tighten over time. The natural forum is the Sprint Retrospective, where the team examines what leaked into production or into the next sprint and asks whether a criterion would have caught it.
Tightening follows capability. A team without automated deployment cannot put “deployed to staging” on the list honestly; putting it there before the capability exists produces either a criterion everyone ignores, or a sprint spent doing manual deploys by hand. So the sequence runs: build the capability, then encode it as a criterion. As Test Pyramid and TDD discipline improves, “unit tests written” can graduate to “unit tests written, integration tests cover the primary path, coverage on changed lines meets threshold.” As CI-CD matures, “merged to trunk” can graduate to “deployed behind a flag to production.” (See Feature Flags for how the latter becomes safe.)
Loosening the DoD is almost always a mistake, and it is usually proposed under deadline pressure. The honest alternative is to reduce scope, not to reduce the bar. A team that ships fewer items that are genuinely done is in a far better position than one that ships more items that are not.
Every arrow back to “Stays In Progress” is the point of the artifact. The gate has no side door.
Undone Work and the “Done-Done” Anti-Pattern
Watch for teams that have invented a vocabulary of gradations: done, done-done, really-done, dev-done, QA-done. The proliferation of qualifiers is diagnostic. It means the word “done” has stopped carrying information, so people bolt on modifiers to recover the meaning it used to have. The cure is not a better adjective; it is a real DoD that makes one word sufficient.
The residue these teams accumulate has a name: undone work. It is the gap between what the team calls done and what would actually be required to put the increment in front of users. If the DoD stops at “merged and reviewed” but shipping additionally requires a regression pass, a performance test, a security review, and a release note, then every completed item carries an invisible tail of four activities.
Undone work has three properties that make it uniquely corrosive:
| Property | Consequence |
|---|---|
| It is proportional to output | The more the team delivers, the larger the tail grows; productivity and risk rise together |
| It is not tracked anywhere | It appears in no backlog, no estimate, and no forecast, so no one can plan for it |
| It is discovered late and all at once | It surfaces at release time as a “hardening sprint” or “stabilization phase” of unpredictable length |
The hardening sprint is the classic symptom. A team runs eight clean sprints, then spends four unplanned weeks making the result actually shippable. Those four weeks were not a new cost; they were the accumulated tail of the previous eight sprints being paid in one lump, at the worst possible moment, with the least information about which item caused which problem.
Two responses are legitimate. Shrink the tail by pulling activities into the DoD as fast as capability allows. Where an activity genuinely cannot move yet, make the remaining tail explicit: write it down, estimate it per item, and reserve capacity for it in the plan. What is never legitimate is leaving it undocumented and hoping.
Why It Matters
- It makes “increment” a meaningful word. Without a DoD, the sum of the sprint’s items is not a releasable thing; it is a pile of partially finished work whose remaining cost is unknown. The DoD is what converts a collection of tickets into something a business can actually decide to ship.
- It makes velocity honest. Velocity and Burndown Charts are only meaningful if the unit being counted is consistent. Points claimed for work that is 80 percent complete inflate the number and destroy its forecasting value. A strict DoD is the calibration standard behind the metric.
- It removes the largest single source of end-of-sprint surprise. The classic sprint collapse, where five items are “almost done” on the last day and none can ship, is a DoD failure, not an estimation failure.
- It prevents invisible technical debt. Work marked done but not really done is debt taken out without anyone signing for it. There is no ticket, no register entry, no owner, and no interest rate anyone can see, but the obligation is real and it compounds. See Code Refactoring and Technical Debt.
- It distributes quality responsibility across the team. When testing, review, and documentation are inside the definition of done rather than downstream activities owned by other people, the developer who writes the code owns its quality.
- It shortens the feedback loop with the Product Owner. “Product Owner has seen it working” as a criterion forces the interpretation conversation to happen inside the sprint rather than at review, when it is expensive.
- It gives new team members an operational description of the culture. A DoD tells a joiner more about how a team actually works than any onboarding document.
- It makes trade-off conversations explicit and adult. When someone asks to skip the tests to hit a date, a written DoD turns that from a quiet individual decision into a visible team decision with a named owner.
- It is the precondition for continuous delivery. You cannot deploy on every merge if “merged” and “releasable” are different states. Tightening the DoD until they converge is the actual work of getting to continuous deployment.
Definition of Done vs Acceptance Criteria
These two are confused more often than any other pair in agile practice, and the confusion is expensive because it lets teams believe they have a DoD when they have only ever written acceptance criteria.
The distinction is scope. The Definition of Done is one list that applies to every item the team produces. Acceptance criteria are a different list for every item, describing what that particular item must do. The DoD is about the quality and completeness of the work; acceptance criteria are about the behavior of the feature.
Put another way: acceptance criteria answer “did we build the right thing?” The DoD answers “did we build the thing right, and completely?” An item can satisfy every acceptance criterion and still be nowhere near done, because it has no tests, no documentation, and lives on an unmerged branch.
Side by Side
| Dimension | Definition of Done | Acceptance Criteria |
|---|---|---|
| Scope | One list, applies to all items | One list per item |
| Author | The development team | Product Owner, refined with the team |
| Changes | Rarely, deliberately, in retrospectives | Constantly, per item, during refinement |
| Question answered | Is the work complete and releasable? | Does the feature do what was asked? |
| Subject matter | Process and quality: tests, review, docs, deploy | Behavior: inputs, outputs, rules, edge cases |
| Typical form | Checklist of engineering and quality gates | Given/When/Then scenarios or a bullet list |
| Verified by | Pipeline gates plus a team checklist | Product Owner acceptance, tests written from them |
| Failure symptom when missing | Hidden unfinished work, unstable velocity | Wrong feature built, rework after review |
| Relationship to a User Story | External to it, applies to all stories | Written on the story itself |
A Worked Example: Password Reset
Take a single story: As a returning user, I want to reset my password by email so that I can regain access to my account.
Acceptance criteria (specific to this story, written by the Product Owner during Product Backlog and Refinement):
| # | Criterion |
|---|---|
| 1 | Submitting a registered email address sends a reset link within 60 seconds |
| 2 | Submitting an unregistered address shows the same confirmation message, revealing nothing about account existence |
| 3 | The reset link expires 30 minutes after issue |
| 4 | The reset link is single use; a second visit shows an expired-link page |
| 5 | New passwords must satisfy the existing password policy, with inline validation |
| 6 | A successful reset invalidates all existing sessions for that account |
| 7 | Rate limit: at most 3 reset requests per address per hour |
Definition of Done (the same list that applied to last sprint’s dashboard widget and next sprint’s billing fix):
| # | Criterion |
|---|---|
| 1 | All acceptance criteria demonstrated and accepted by the Product Owner |
| 2 | Unit tests cover new logic; integration test covers the end-to-end reset flow |
| 3 | Pull request approved by at least one other engineer |
| 4 | CI green: build, tests, lint, dependency vulnerability scan |
| 5 | No new critical or high severity issues in the security scanner |
| 6 | Merged to trunk and deployed to staging |
| 7 | Structured logs and a metric for reset-request volume and failure rate |
| 8 | Help centre article updated; support team notified of behavior change |
| 9 | Accessibility check passed on both new pages |
| 10 | Feature flag defined with an owner and a removal date |
Notice that criterion 1 of the DoD references acceptance criteria. That is the correct relationship: the DoD contains a slot into which each item’s specific criteria are plugged. They are not competitors; they compose. And notice that criteria 2 through 10 would be word-for-word identical on a story about export to CSV, which is exactly why they live in one shared document rather than being retyped on every ticket.
The pathological pattern to watch for: a team that writes excellent acceptance criteria and has no DoD ships features that behave correctly and are simultaneously untested, undocumented, unmonitored, and unmergeable. Every individual demo goes well. The system degrades anyway.
Where Bugs, Spikes, and Chores Fit
Because the DoD applies to everything the team produces, it must survive contact with items that are not user stories. Teams frequently trip here and quietly exempt whole categories of work, which reopens the same leak.
| Item type | Acceptance criteria | Definition of Done |
|---|---|---|
| User story | Behavioral scenarios for the feature | Applies in full |
| Bug fix | The reported defect no longer reproduces, under the stated conditions | Applies in full, plus a regression test proving the specific defect stays fixed |
| Spike or research | A question to answer and a decision to enable | A tailored subset: findings written down and shared, recommendation recorded, throwaway code discarded rather than merged |
| Chore or upgrade | The dependency, version, or config reaches the target state | Applies in full; upgrades are exactly the work most likely to break something silently |
Spikes are the one honest exception, because a spike produces knowledge rather than a shippable increment and gating it on “deployed to staging” is meaningless. Handle that by writing a short, explicit spike DoD rather than by waving spikes through with no bar at all. Bugs and chores get no exception whatsoever: an untested bug fix is the purest form of undone work, since the team has now touched fragile code and left behind no evidence that the fix holds.
Definition of Ready, and the Argument Against It
The mirror-image concept is the Definition of Ready (DoR): a checklist an item must satisfy before the team will pull it into a Sprint. Typical contents: the item has a clear description, acceptance criteria exist, dependencies are identified, it is estimated, it is small enough to finish inside a sprint, any required design assets are attached.
The case for it is real. Teams that repeatedly start work on half-specified items burn days mid-sprint chasing clarification, and their sprint failure rate reflects it. A DoR gives the team a legitimate, non-personal way to say “not yet” to an item that would otherwise stall.
The case against it is also real, and worth taking seriously rather than dismissing. The objections cluster into four:
| Objection | The argument |
|---|---|
| It recreates a stage gate | A hard readiness gate reintroduces a phase boundary between analysis and development, which is precisely the handoff that iterative methods were designed to remove |
| It encourages upfront specification | Teams start writing detail into items that may never be built, which is inventory in exactly the sense Lean Software Development warns about |
| It substitutes for conversation | A checklist on a ticket is a weak replacement for two people talking. “Ready” documents can be complete and still leave everyone unclear |
| It creates a queue of blame | “Not ready” becomes a way to return work to the Product Owner rather than a prompt to go and resolve the ambiguity together |
The workable middle ground, and the one most experienced teams land on, is to treat the DoR as a heuristic held by the team rather than a policy enforced against the backlog. It is a conversation prompt during refinement, not a gate with a bouncer. Where the DoD must be binary and strict, the DoR should be soft and negotiable. That asymmetry is deliberate: being wrong about readiness costs a conversation, while being wrong about doneness costs a release.
A useful diagnostic: if your DoR contains “all questions answered,” you have built a waterfall gate. If it contains “we understand this well enough to start and we know who to ask,” you have built a heuristic.
The Maturity Ladder
A DoD is a snapshot of engineering capability, so it should read differently for a team six months into automation than for one running continuous deployment. The ladder below is not a maturity model to be scored against; it is a description of what each rung becomes honest to write down once the underlying capability exists.
| Rung | Delivery criterion | Testing criterion | Prerequisite capability |
|---|---|---|---|
| 1 | Merged to a shared branch | Manually verified by the author | Version control and a code review habit |
| 2 | Builds reproducibly in CI | Unit tests exist for new logic | A working CI server |
| 3 | Deployed automatically to staging | Integration test covers the primary path | Scripted, repeatable deployment |
| 4 | Deployed to production behind a flag | Coverage threshold on changed lines enforced | Feature Flags with owners and expiry |
| 5 | Released to a percentage of real traffic | Contract tests plus an automated rollback trigger | Telemetry good enough to detect harm in minutes |
Two rules govern movement on this ladder. First, build the capability, then encode the criterion. A criterion the team cannot meet is worse than no criterion, because it teaches everyone that the list is aspirational. Second, move one rung at a time and let it settle for several sprints. Adding three criteria at once makes it impossible to tell which one caused the throughput change you are now arguing about.
The escape-driven loop matters more than the ladder itself. The best source of new DoD criteria is not a template but your own production incidents: for each escaped defect, ask what single verifiable check would have stopped it, and whether that check is cheap enough to run on every item forever. Most are not, which is why the list stays short. The ones that are become permanent.
The Debt Mechanism: How a Weak DoD Compounds
Consider what actually happens when an item is marked done with tests skipped.
The loop is self-reinforcing, which is what makes it dangerous. Each turn of it increases the pressure that caused the shortcut in the first place. And because the missing work was never recorded anywhere, the team’s own metrics actively lie about the situation: velocity rises exactly when quality falls, and falls later for reasons that appear unrelated.
This is the sense in which the DoD is a financial control. Recorded technical debt is a decision: someone chose a shortcut, wrote it down, and accepted the interest. Unrecorded debt from a violated DoD is not a decision at all. It is a liability created without a signature, which no one can price, prioritize, or pay down deliberately. The register in Code Refactoring and Technical Debt only works if things actually reach it, and a strict DoD is what forces them there: if the tests are not written, the item is not done, so either the tests get written or the shortfall becomes a visible, ticketed piece of work.
The practical rule that follows: if the team decides to violate the DoD, that is permitted only if the shortfall becomes a backlog item in the same conversation. Not “we will get to it.” A ticket, with a description of what is missing, created before the item moves.
Comparison
| Concept | Scope | Owner | Changes how often | Purpose |
|---|---|---|---|---|
| Definition of Done | Every item, uniformly | Development team | Rarely; tightens over time | Guarantees the increment is genuinely releasable |
| Acceptance Criteria | One specific item | Product Owner with team | Every item, during refinement | Specifies what that feature must do |
| Definition of Ready | Every item, before pull | Team, softly | Occasionally | Reduces mid-sprint clarification stalls |
| Sprint Goal | The whole sprint | Product Owner with team | Every sprint | Gives the sprint a coherent purpose and a basis for trade-offs |
| Release Criteria | A release candidate | Product plus operations | Per release train | Gates a deploy to real users; often a superset of the DoD |
The relationship worth internalizing: acceptance criteria fold into the DoD (criterion one), the DoD applies to every item contributing to the Sprint Goal, and if the DoD is strong enough, release criteria become nearly empty because the DoD already covers them. A team whose release checklist is long and whose DoD is short has put its quality bar in the wrong place, at the point where fixing anything is most expensive.
Real-World Use Cases
- A payments team adds “PCI scope review completed for any change touching cardholder data” to its DoD. It is a genuine organizational constraint with an external dependency, so the team also adds a refinement rule: any item likely to hit PCI scope is flagged during Product Backlog and Refinement so the review is requested on day one rather than day nine.
- A mobile team cannot deploy on merge because of app store review latency, so its DoD says “merged to trunk, included in a signed internal build, verified on the two lowest-supported OS versions.” The criterion respects the real constraint instead of pretending it away.
- A platform team whose customers are other engineers puts “the change is documented in the changelog with a migration note, and the version is bumped per Semantic Versioning” on the list, because for a library, undocumented behavior change is the primary failure mode.
- A team adopting trunk-based development graduates its DoD from “merged to a release branch” to “merged to trunk behind a flag with an owner and expiry date,” directly coupling the DoD change to the new Feature Flags capability.
- A startup post Product-Market Fit discovers its support load is dominated by undocumented features and adds a docs criterion. Cycle time rises by roughly a day per item; support tickets fall by a third. The trade was made deliberately and measured.
- A regulated healthcare team includes “clinical safety case updated and signed by the clinical safety officer.” It is slow, external, and non-negotiable, so the team keeps a standing weekly slot with that officer rather than treating each request as a surprise.
- A team running Kanban encodes its DoD as an explicit exit policy printed above the Done column, with a work-in-progress limit on the column immediately before it, so items physically pile up and become visible when the gate is not being met.
- A team recovering from a bad quarter deliberately shrinks scope rather than the bar, halving planned items for two sprints while adding a coverage-on-changed-lines gate. Velocity dips, then recovers above its prior level once unplanned defect work stops consuming capacity.
- A Scaled Agile (SAFe and LeSS) program publishes a baseline DoD all eight teams must meet for integration, then leaves each team free to hold a stricter local list. Integration failures drop because the weakest link now has a floor.
- A team practicing strong Test Pyramid and TDD removes “unit tests written” from its DoD entirely, because tests-first is now how code is written and the criterion had become a tautology. The list should contain live constraints, not fossils.
Common Pitfalls
- Confusing it with acceptance criteria. The most common failure of all. Teams write detailed per-story criteria, call that their Definition of Done, and are then baffled when features that pass every demo turn out to be untested, unmonitored, and undocumented. One list for everything, one list per item; they are different documents doing different jobs.
- Writing criteria the team cannot verify. “Code is maintainable,” “performance is acceptable,” “the design is good.” These feel important and are useless as gates because two reasonable people will disagree. Replace each with something a machine or an unambiguous human check can answer: a lint rule, a latency budget with a number, a review approval.
- Including things outside the team’s control. Putting “approved by the architecture board” on the list when that board meets monthly guarantees that nothing is ever done inside a sprint. Either get the constraint changed, or restructure how the team engages with it, or accept that your sprint boundary is fictional.
- Treating it as negotiable under pressure. The DoD exists precisely for the moments when it is inconvenient. A definition that is suspended whenever a deadline approaches is not a definition; it is an aspiration. If the bar genuinely must be lowered, change the written DoD explicitly, in the open, with the team, and change it back on a named date.
- Letting it go stale. A DoD written eighteen months ago that still says “attach a screenshot to the ticket” while the team now has automated visual regression testing has become theatre. Every criterion in a stale list teaches the team that the list is decorative, which corrodes compliance with the criteria that still matter.
- Making it so long that nobody reads it. Thirty criteria is not a higher bar than twelve; it is a lower one, because thirty criteria get skimmed. Automate everything automatable so the human-checked residue stays short enough to actually check.
- Enforcing it only at the sprint review. Discovering at review that four items fail the DoD is far too late. The gate belongs at the moment the item moves, on the board, with the pipeline doing most of the work continuously.
- Allowing “done except.” The phrase should be treated as a signal, not a status. An item is done or in progress. Partial credit is the exact mechanism by which unfinished work escapes into the past, where it is nobody’s job to find it.
- Copying another team’s list wholesale. A DoD reflects a specific team’s capability, stack, risk profile, and constraints. Adopting one from a book or a neighboring team produces criteria the team cannot meet and does not believe in. Start from what you already do reliably, write that down, then add one criterion at a time.
- Assuming a strict DoD alone produces quality. It sets a floor and prevents a specific failure mode. It does not substitute for good design, real user feedback, or the judgment covered in Code Review and Static Analysis. A team can meet every criterion and still ship something nobody wanted, which is what acceptance criteria and the Sprint Goal are for.
Related Terms
- Scrum
- Sprint
- Product Backlog and Refinement
- Sprint Retrospective
- Velocity and Burndown Charts
- Test Pyramid and TDD
- Code Refactoring and Technical Debt
- CI-CD
Example
A fintech team of seven maintained a reconciliation service. Their board looked healthy: 40 to 45 points per sprint, burndown lines that reached zero, demos that went well. Their production incident rate had roughly tripled over nine months, and nobody could explain the discrepancy. Their written DoD had three items: code merged, PR approved, story demoed.
The retrospective that changed things started from a single number. An engineer counted how many of the previous quarter’s 61 completed items had test coverage on the code they changed. The answer was 22. Nothing in the DoD required tests, so during any busy sprint, tests were the first thing cut, and because the omission created no ticket, it was invisible in every metric the team reported. Their velocity of 43 was, on inspection, roughly 43 points of feature work plus an unrecorded and steadily growing liability. Work marked done but not really done is debt taken out without anyone signing for it, and this team had been signing nothing for nine months while the balance compounded.
They did two things. First, they wrote the missing work down: a triage pass produced 19 backlog items for untested code paths in the highest-risk modules, which for the first time made the shortfall a number leadership could see and prioritize. Second, they rewrote the DoD to ten items, with seven enforced by the pipeline: coverage on changed lines above 70 percent, integration test on any code path touching ledger balances, dependency scan clean, deployed to staging automatically, structured logging on new branches, runbook entry for any new alert, and Product Owner confirmation on staging rather than on a developer laptop. They did not add a criterion for anything they could not yet automate; a “performance regression check” idea was deferred a quarter until the benchmark harness existed.
Velocity fell to 26 for two sprints and the team held its nerve, mostly because they had agreed in advance that the drop was the measurement correcting rather than the team slowing. By sprint five it was back to 38, and by sprint eight it sat at 44 with a materially different composition: unplanned defect work had fallen from roughly a quarter of capacity to under 8 percent. The number on the chart was almost identical to where they started. The difference was that it now meant something, because every point behind it referred to work that was actually finished.
Referenced by
- Agile Manifesto
- Extreme Programming (XP)
- Iterative and Incremental Development
- Kanban
- Product Backlog and Refinement
- Requirements Engineering
- Scaled Agile (SAFe and LeSS)
- Scrum
- Software Development Methodology Terms MOC
- Sprint
- Sprint Retrospective
- Story Points and Estimation
- V-Model
- Velocity and Burndown Charts
- Waterfall Model