Extreme Programming (XP)
Extreme Programming (XP)
Definition: Extreme Programming (XP) is an agile software development framework, created by Kent Beck in the late 1990s, that prescribes a specific set of technical engineering practices — test-driven development, pair programming, continuous integration, relentless refactoring — alongside a lightweight planning process. Unlike frameworks that govern how work is organized and tracked, XP governs how code is actually written. Its central bet is that the cost of changing software can be kept flat rather than rising exponentially over time, provided the team maintains extremely tight feedback loops and a permanently clean codebase. The name comes from taking known-good practices and turning the dial to the extreme: if code review is good, review continuously through pairing; if testing is good, test before writing every line; if integration is good, integrate many times a day.
XP is the only mainstream agile framework that will tell you what to do at 2 p.m. on a Tuesday with your hands on a keyboard.
How It Works
XP is built as three concentric layers: values (why), principles (how to reason), and practices (what to do). The practices are the visible surface, but they are not a menu — they are deliberately interlocking, and removing one degrades several others.
The Five Values
The values are the framework’s constitution. Every practice traces back to at least one of them.
| Value | Meaning in practice | What it rules out |
|---|---|---|
| Communication | Prefer face-to-face and code-as-conversation over documents and tickets | Handoff-by-spec, silent disagreement, tribal knowledge silos |
| Simplicity | Build the simplest thing that could possibly work, today | Speculative generality, framework-first architecture, “we’ll need it later” |
| Feedback | Shorten every loop until problems surface within minutes | Long integration branches, quarterly releases, deferred QA |
| Courage | Delete code, tell the truth about estimates, refactor the scary module | Defensive coding, padded estimates, avoiding the hard conversation |
| Respect | Every member’s contribution matters; nobody ships breakage onto teammates | Hero culture, blame retros, broken builds left overnight |
Courage is the one people underestimate. XP asks you to throw away a day’s work when a simpler design emerges, to say “that will take three weeks, not one,” and to rewrite the module everyone is afraid of. Without the safety net of tests and pairing, courage is recklessness; with it, courage is just competence.
The Principles Layer
Values are too abstract to settle an argument; practices are too concrete to cover a novel situation. The principles are the translation layer — the reasoning tools you use when the practices do not obviously apply.
- Humanity. Software is written by people with needs for safety, accomplishment, belonging, and growth. A process that ignores those needs will be routed around.
- Economics. Software must produce value; the time value of money and the option value of deferred decisions are real inputs to design choices, not afterthoughts.
- Mutual benefit. Every practice must benefit the person doing it now, not only some future maintainer. This is why XP favors tests and refactoring over documentation: documentation costs you today and pays someone else later, while a test pays you within minutes.
- Self-similarity. Structures that work at one scale usually work at another. The red-green-refactor loop of TDD is the same shape as the plan-build-review loop of an iteration, and the same shape as a release cycle.
- Improvement. There is no perfect process, only a better one than yesterday’s. Ship, then improve — do not wait for excellence to begin.
- Diversity. Teams need differing perspectives to see problems whole. Conflict of opinion is the mechanism, not the failure mode; the practices exist to make that conflict productive.
- Reflection. Good teams examine how they work as well as what they produce, which is the same instinct later formalized as the Sprint Retrospective.
- Flow. Deliver a steady stream of value rather than batching into large phases. The lineage from this principle to Kanban and to DevOps Culture is direct.
- Opportunity. Treat problems as chances to change, not merely as things to survive.
- Redundancy. Critical problems get several defenses at once. Defects are caught by pairing, by unit tests, by acceptance tests, and by CI — overlapping nets, deliberately.
- Failure. If you do not know which of three designs is right, try one. Failure that produces knowledge is not waste.
- Quality. Quality is not a dial you trade against speed. Lowering quality does not make projects faster; past a shallow threshold it makes them slower, because rework and fear consume the time supposedly saved.
- Baby steps. Small changes deployed frequently beat large changes deployed rarely — the risk of a step is roughly proportional to its size.
- Accepted responsibility. Work is signed up for, not assigned. The person who accepts a story owns its estimate, its tests, and its design.
The Twelve Core Practices, Grouped
Beck’s original Extreme Programming Explained (1999) laid out twelve practices. They divide cleanly into four thematic clusters.
Cluster 1 — Fine-scale feedback (the code loop)
- Test-Driven Development (TDD). Write a failing test, write the minimum code to pass it, refactor. The test is a specification written in executable form. TDD is not primarily about defect count — it is a design technique. Code that is hard to test is code with bad coupling, and TDD makes you feel that pain before the design ossifies. See Test Pyramid and TDD.
- Pair Programming. Two engineers, one keyboard. The driver types; the navigator thinks one level up — about edge cases, naming, whether this is the right approach at all. Roles swap frequently. Pairing is continuous code review, continuous knowledge transfer, and continuous design discussion, collapsed into the moment the code is written rather than deferred to a pull request days later.
- Whole Team / On-Site Customer. A real customer representative with actual authority sits with the team, available to answer questions in seconds rather than days. This practice is what makes XP’s minimal requirements documentation viable: you do not need a 40-page spec if you can turn around and ask.
- The Planning Game. Business writes stories and sets priority; engineering estimates and sets scope per iteration. Neither side overrides the other. Business cannot dictate estimates; engineering cannot dictate priority. This clean separation of authority is one of XP’s most underrated contributions.
Cluster 2 — Continuous process
- Continuous Integration. Every pair integrates to the mainline multiple times per day, and the full test suite runs on every integration. Integration ceases to be a phase and becomes a non-event. This practice is the direct ancestor of everything now filed under CI-CD and CI-CD Best Practices.
- Refactoring (Design Improvement). Improve structure without changing behavior, continuously, as part of every task — not as a scheduled “cleanup sprint.” The premise is that design decays under feature pressure unless actively maintained, and that a comprehensive test suite makes structural change cheap and safe. See Code Refactoring and Technical Debt.
- Small Releases. Ship to real users in the smallest useful increments, on a cadence measured in weeks or less. Small releases shrink the blast radius of any single mistake and generate real market feedback rather than imagined feedback.
Cluster 3 — Shared understanding
- Simple Design. The codebase should pass four criteria, in priority order: passes all tests; reveals intention; contains no duplication; uses the fewest elements possible. Anything beyond that is speculative and will probably be wrong. This is the practical form of YAGNI — “You Aren’t Gonna Need It.”
- Collective Code Ownership. Anyone may change any file. There are no gatekeepers and no “that’s Dave’s module.” Combined with pairing and a strong test suite, this eliminates the bus-factor problem and removes the queueing delay of waiting for a specific person.
- Coding Standards. A shared, agreed style so that collective ownership does not produce a patchwork. Today this is largely automated by formatters and linters — see Code Review and Static Analysis — but the principle stands: the code should look like one person wrote it.
- System Metaphor. A shared, plain-language story of how the system works (“the codebase is a shipping warehouse: orders arrive, get picked, get packed, get dispatched”) that guides naming and structure. This is the least-adopted XP practice and the one Beck himself later de-emphasized, but the underlying need — a coherent shared mental model — is real.
Cluster 4 — Programmer welfare
- Sustainable Pace (originally “40-Hour Week”). No sustained overtime. XP treats crunch not as a moral issue but as an engineering one: tired programmers write code with more defects, make worse design decisions, and skip the practices that keep the system healthy. Overtime borrows velocity from next month at a punitive interest rate.
Primary and Corollary Practices
The 2004 second edition of Extreme Programming Explained reorganized the twelve into a larger set split by adoption order — a response to the observation that teams were failing by attempting everything simultaneously.
| Tier | Practices | Adoption guidance |
|---|---|---|
| Primary | Sit together, whole team, informative workspace, energized work, pair programming, stories, weekly cycle, quarterly cycle, slack, ten-minute build, continuous integration, test-first programming, incremental design | Safe to adopt individually and in any order; each yields value on its own |
| Corollary | Real customer involvement, incremental deployment, team continuity, shrinking teams, root-cause analysis, shared code, code and tests as the primary artifacts, single code base, daily deployment, negotiated scope contract, pay-per-use | Attempt only after the primary practices are solid; adopting these early usually backfires |
Three additions in that edition deserve attention because they are less known than the original twelve:
- Slack. Deliberately include low-priority, droppable work in every cycle so commitments can be met without heroics. Slack is what makes sustainable pace survive contact with a deadline — it is the buffer that absorbs estimation error instead of transferring it onto people’s evenings.
- The ten-minute build. A hard, specific number: the entire system must build and run all tests in under ten minutes. This is stated as a practice rather than an aspiration because everything else — continuous integration, refactoring confidence, fast feedback — degrades the moment the build gets slower than a coffee break.
- Informative workspace. The room itself should report project status at a glance: story cards on a wall, build status visible, defect counts posted. Information radiators beat information refrigerators, an idea that survives today in dashboards and build-status lights.
The primary/corollary split is XP’s own admission that big-bang adoption fails. It is also the most practically useful part of the framework for a team starting out: begin with the ten-minute build, continuous integration, and test-first programming, because those three make every other practice cheaper.
The Nested Feedback Loops
The structural heart of XP is a set of concentric feedback cycles, each an order of magnitude longer than the one inside it. Nothing waits long to be checked.
Read it as a signal-latency chart. A typo caught by your navigator costs seconds. A design flaw caught by a unit test costs minutes. A misread requirement caught by an acceptance test costs days. The same flaw discovered after a six-month waterfall handoff costs months. XP’s entire economic argument is that pushing detection leftward on this scale is worth the overhead the practices impose.
The Cost-of-Change Curve
Classical software economics — the argument behind the Waterfall Model and the V-Model — holds that the cost of changing a requirement rises exponentially with how late it is discovered, therefore you must specify everything up front. Beck’s counter-claim was that the exponential curve is not a law of nature but an artifact of the practices in use.
This is the intellectual crux. If the curve really is flat, deferring decisions is not sloppiness — it is optimal, because a decision made later is made with strictly more information. Every XP practice exists to hold that curve down.
The Planning Game and the Development Cycle
XP’s planning process is deliberately thin, and its one strong opinion is about authority: business decides what and in which order, engineering decides how long and how. Neither may overrule the other, which removes the single most common dysfunction in software planning — the negotiated-down estimate.
Two details carry most of the weight. First, velocity is measured, never assigned — the team commits to what the last cycle actually produced, which converts planning from negotiation into arithmetic (see Velocity and Burndown Charts). Second, the inner red-green-refactor loop and the outer plan-build-demo loop are the same shape at different scales, which is the self-similarity principle made concrete: a broken test and a misunderstood story are handled identically, just at different latencies.
Why It Matters
- It is the engineering half of agile. Scrum and Kanban tell you how to organize and visualize work; neither says a word about how code should be written. XP fills exactly that gap, which is why teams that adopt process-only agile so often end up delivering the same low-quality software on a faster cadence.
- It originated or popularized practices now considered baseline. TDD, continuous integration, refactoring-as-discipline, collective ownership, user stories, and the planning-poker lineage of Story Points and Estimation all trace to the XP community.
- It gives quality a mechanism, not an aspiration. “We value quality” is a poster. “No code reaches mainline without a test that failed first, written by a pair, integrated within four hours” is a system.
- It makes emergent architecture defensible. Without a test harness and refactoring discipline, “we’ll evolve the design” is wishful thinking. With them, it is a repeatable practice.
- It attacks the bus factor structurally. Pairing plus collective ownership means knowledge is distributed as a side effect of normal work, not through documentation drives that never happen.
- It reframes technical debt as an operational hazard. XP does not schedule debt paydown; it refuses to accumulate the debt in the first place, on the grounds that a codebase resistant to change destroys the team’s ability to respond to the market.
- It hardwires humane pace into the method. Sustainable pace is a named practice, not a wellness initiative — burnout is treated as a defect source, which makes it an engineering concern that survives budget scrutiny.
- Its ideas underpin modern delivery. The DORA research program’s high-performing-team metrics (deployment frequency, lead time, change-failure rate, restore time) measure almost exactly what XP practices optimize. See DevOps Culture.
- It is falsifiable at the team level. Unlike vaguer methodologies, you can walk into a room and objectively determine whether a team does XP: is there a red-green-refactor cycle, are people pairing, does mainline build in under ten minutes?
Deep Dive: Why XP Lost the Name and Won the Substance
This is the most interesting thing about XP, and it is widely misunderstood.
By around 2010, XP had effectively vanished as a named framework in enterprise adoption. Certification bodies, consultancies, and job titles coalesced around Scrum. Surveys of agile adoption consistently show Scrum and Scrum-derived hybrids dominating, with pure XP in the low single digits. On the scoreboard of framework adoption, XP lost decisively.
And yet look at what a competent 2020s engineering team actually does:
| XP practice (1999) | Its 2020s descendant | Ubiquity today |
|---|---|---|
| Continuous integration | CI-CD pipelines, CI-CD Best Practices | Near-universal |
| Test-driven development | Test Pyramid and TDD, test-first culture, coverage gates | Very common |
| Refactoring as routine work | Code Refactoring and Technical Debt as a first-class concept | Universal vocabulary |
| Small releases | Continuous deployment, Feature Flags, trunk-based development | Very common |
| Collective code ownership | Pull-request culture, shared repos, no per-file owners | Near-universal |
| Coding standards | Auto-formatters, linters, Code Review and Static Analysis | Near-universal |
| Pair programming | PR review, mob/ensemble sessions, AI pairing tools | Partially adopted |
| User stories from the planning game | User Story, backlog items | Universal vocabulary |
| Simple design / YAGNI | YAGNI as an idiom every senior engineer uses | Universal vocabulary |
| On-site customer | Embedded PM, product trio | Weakly adopted |
Nearly every practice won. The framework’s label is what died.
Several mechanisms explain the split:
Scrum was easier to sell and easier to adopt. Scrum asks managers to change meeting structures. XP asks engineers to change how they type, and asks the business to seat a decision-maker with the team full-time. The first is a training purchase; the second is an organizational renegotiation. Adoption follows the path of lower resistance, and Scrum’s certification economy gave it a distribution engine XP never had.
XP’s practices are separable; XP as a system is not. Because each practice has standalone value, organizations cherry-picked. CI is valuable on its own, so CI spread. TDD is valuable on its own, so TDD spread (partially). Pair programming has its value concentrated in second-order effects — knowledge diffusion, defect prevention, design quality — which are hard to attribute in a spreadsheet, and its cost is glaringly visible as two salaries on one task. So pairing did not spread. XP was disassembled for parts, and the parts were good enough to survive independently.
The tooling absorbed the discipline. Practices that could be automated became infrastructure and stopped needing a methodology to enforce them. Nobody says “we practice continuous integration” anymore; they say “CI is green.” Coding standards became a pre-commit hook. When a practice becomes a tool, it loses its association with the philosophy that produced it — which is precisely why most engineers who run a red-green-refactor loop every day have never read a word of Beck.
The synthesis is what actually shipped. The dominant real-world configuration is Scrum ceremonies wrapped around XP engineering practices: sprints, standups, and retros for coordination; TDD, CI, refactoring, and code review for the code. Most teams running this configuration have no idea half of it is XP. That is the sense in which XP won: its ideas are so thoroughly naturalized that they read as “just how software is built” rather than as one framework’s contested proposal.
The practical lesson for a team today: if your agile adoption feels like it produced faster meetings but not better software, the missing half is almost certainly the XP half. Adding engineering practices to a process framework is usually a larger improvement than switching process frameworks.
Deep Dive: Why Full XP Adoption Is Genuinely Rare
Honesty matters here, because XP advocates often frame low adoption as pure organizational cowardice. Some of it is. But two practices are genuinely expensive and organizationally hard, and pretending otherwise damages the argument.
Full-time pair programming. The evidence is broadly favorable — studies typically find pairs take somewhat more total effort per task (commonly in the 15-60% range depending on task novelty and pair experience) while producing fewer defects and better designs. But that trade is hard to defend under budget pressure, because the cost is immediate and legible while the benefit is deferred and diffuse. Pairing is also genuinely draining: sustained pairing is more cognitively intense than solo work, and it suits some people far better than others. Introverted engineers, neurodivergent engineers, and engineers in different time zones can all find full-time pairing corrosive. Most teams land on a pragmatic middle — pair on hard, novel, or high-risk work; solo on routine work; mob on architectural decisions — which is defensible even though it is not XP as written.
The on-site customer. XP assumes a customer representative with real decision authority, embedded with the team, available continuously. In most organizations no such person exists or can be spared. Product managers own multiple teams. Real customers are external, numerous, and contradictory. In regulated or enterprise-sales contexts, requirements come from committees and contracts, not from a person you can ask. Teams substitute a part-time PM, and the practice degrades into “we have a PM in Slack” — which is not the same thing at all, because the whole point was sub-minute latency on requirement questions.
Secondary frictions compound these:
- XP presumes a small, co-located, generalist team — roughly two to twelve people in one room. Distributed teams, deep specialization, and organizations of hundreds all strain the model, and XP has no native scaling story of its own (see Scaled Agile (SAFe and LeSS)).
- The practices are mutually load-bearing. Adopt collective code ownership without a strong test suite and you get chaos. Adopt refactoring without CI and you get long-lived broken branches. Partial adoption can be worse than none, which makes incremental rollout risky.
- There is no certification economy. No credential to sell means no consultants evangelizing it into procurement budgets, and no manager-legible proof of adoption.
- The name aged badly. “Extreme” signalled edginess in 1999 and reads as unserious in an enterprise procurement document in 2026. Naming genuinely affected diffusion.
The reasonable conclusion is not “XP failed” but “XP’s separable practices were adopted at rates proportional to how cheap they were to adopt, and its expensive practices were not.” That is a perfectly rational market outcome, and it leaves a real question open: teams that skip pairing and the on-site customer are running XP with its two highest-bandwidth communication channels removed, and they should expect to pay for that somewhere else — usually in review latency and requirement churn.
Comparison
| Dimension | Extreme Programming (XP) | Scrum | Kanban | Lean Software Development |
|---|---|---|---|---|
| Primary layer | Engineering practice | Process management | Workflow / flow management | Philosophy and principles |
| Prescribes how to write code | Yes, extensively | No | No | No |
| Prescribes roles | Coach, customer, programmer (light) | Product Owner, Scrum Master, Developers | None required | None |
| Cadence | 1-2 week iterations, small releases | Fixed-length Sprint | Continuous flow, no iterations | Continuous flow |
| Change mid-cycle | Allowed if a story of equal size is swapped out | Discouraged within a sprint | Fully allowed, pull-based | Allowed |
| Core control mechanism | Feedback loops and test suites | Timeboxes and ceremonies | WIP limits and cycle time | Eliminating waste, amplifying learning |
| Quality mechanism | Built in via TDD, pairing, CI | Delegated to Definition of Done | Not specified | Build integrity in |
| Typical team size | 2-12, co-located | 3-9 | Any | Any |
| Adoption difficulty | High — changes daily engineer behavior | Moderate — changes meetings | Low — overlays existing process | Low as philosophy, hard in practice |
| Certification market | Effectively none | Very large | Moderate | Small |
The key point the table makes: XP and Scrum are not competitors. They operate at different layers and compose cleanly. Scrum has a Definition of Done but never tells you how to reach it; XP is a concrete answer to that question. A team can run Scrum ceremonies — Sprint planning, standups, Sprint Retrospective, Velocity and Burndown Charts — while using XP practices for all actual development. This hybrid, sometimes called “Scrum-XP” or just left unnamed, is arguably the most common effective configuration in the industry.
Where XP does genuinely differ from Scrum on process: XP’s planning game allows swapping an unstarted story for another of equal estimate mid-iteration, whereas classical Scrum protects sprint scope; and XP explicitly assigns estimation authority to engineers and prioritization authority to the customer, a separation Scrum implies but states less forcefully.
Real-World Use Cases
- Chrysler C3 payroll (1996-2000) — the project where Kent Beck, Ron Jeffries, and Martin Fowler assembled XP. It shipped payroll for roughly 10,000 employees and became the canonical XP case study, though the project was ultimately cancelled — a fact XP critics cite and advocates attribute to organizational rather than technical causes.
- Early-stage startups pre-Product-Market Fit — small co-located teams, an available founder acting as the on-site customer, and requirements that change weekly. This is the environment XP was designed for and where it fits most naturally.
- Financial trading and pricing systems — several London investment banks became well-known XP shops in the 2000s. High correctness requirements plus fast-changing business rules make TDD and continuous integration directly profitable, and the salary structure makes pairing’s cost less salient.
- Payment and billing platforms — domains where a defect is a financial loss and a compliance event. TDD-derived executable specifications double as audit evidence for how a rule is implemented.
- Legacy system rescue — characterization tests plus disciplined refactoring is the standard playbook for taming an untested legacy codebase, a technique that grew directly out of the XP community (Michael Feathers’ work in particular).
- Rewriting a high-churn core service — collective code ownership plus pairing lets several engineers work concurrently on interlocking modules without the merge queue and review latency that a per-owner model imposes.
- Onboarding at scale — teams facing rapid headcount growth use pairing and rotating pairs as the primary onboarding mechanism, cutting time-to-first-meaningful-commit from weeks to days.
- Regulated and safety-adjacent software — medical device software, clinical systems, and similar domains use XP’s test-first discipline to produce traceable, executable requirement coverage, layering formal documentation on top rather than replacing it.
- Platform and internal tooling teams — where the “customer” genuinely does sit nearby (they are other engineers in the same building), the on-site-customer practice is actually achievable, making full XP more viable than in product teams.
- Ensemble/mob programming adoptions — a modern descendant where the whole team works on one story at one screen; teams typically adopt it for the hardest 20% of work while pairing or soloing the rest.
Common Pitfalls
- Cherry-picking the cheap practices. Adopting CI and stand-ups while skipping TDD, pairing, and refactoring gives you fast delivery of code nobody has verified. The practices reinforce each other; the ones teams skip are usually the ones doing the load-bearing work.
- Treating “simple design” as permission to skip design. YAGNI means do not build for imagined futures — not “do not think.” Simple design still requires the four rules (tests pass, intent revealed, no duplication, minimal elements), and the last three demand real design effort.
- Emergent architecture without refactoring discipline. If the team defers design decisions but never actually restructures, you get accidental architecture: the shape of the system becomes whatever the first six stories happened to require. Emergent design is a claim about continuous restructuring, not about deferral alone.
- Pairing as surveillance or as a tutorial. A senior watching a junior type is not pairing; neither is one person driving for four hours while the other reads Slack. Real pairing requires role rotation, active navigation, and genuine peer engagement. Bad pairing is worse than solo work because it costs double and produces less.
- The proxy customer. Substituting a busy product manager, a business analyst without authority, or a Slack channel for the on-site customer preserves the ceremony while destroying the practice. The value was sub-minute latency on requirement questions; a two-day turnaround provides none of it.
- Slow or unreliable test suites. A suite that takes 40 minutes will not be run before every commit; a suite with flaky tests will be ignored when red. Either failure silently converts XP back into ad-hoc development while the team still believes it is doing XP. Test suite speed and reliability are load-bearing infrastructure, not hygiene.
- Confusing test coverage with TDD. Writing tests after the fact to hit a coverage number gets you the metric without the design benefit. TDD’s value is that the test constrains the design before the design exists; tests written afterwards mostly ratify whatever you already built.
- Sustainable pace abandoned at the first deadline. The practice exists precisely for deadline pressure — that is the only time it costs anything. A team that drops it under pressure has not adopted it; it has merely enjoyed it during calm periods.
- Collective ownership without standards or tests. Anyone-can-change-anything plus no shared style plus weak tests produces a codebase that degrades from every direction at once. Collective ownership is safe only inside the safety net the other practices provide.
- Scaling XP by adding people. XP’s communication practices assume everyone can talk to everyone. Past roughly a dozen people the model breaks and needs an explicit structural answer — see Team Topologies and Conway’s Law and Scaled Agile (SAFe and LeSS) — rather than simply more pairs in more rooms.
- Declaring XP because the team pairs occasionally. Partial adoption is fine and often correct, but call it what it is. Teams that believe they are doing XP stop looking for the gaps that are actually hurting them.
Related Terms
- Agile Manifesto
- Scrum
- Test Pyramid and TDD
- Code Refactoring and Technical Debt
- CI-CD
- Code Review and Static Analysis
- Iterative and Incremental Development
- User Story
Example
A seven-person team at a mid-size insurance company owns the quoting engine — the service that turns an applicant’s details into a price. It is nine years old, has 4% test coverage, and takes three weeks to change a rating rule because nobody can predict what a change will break. Underwriting wants rules changed weekly. The team has been running Scrum for two years: they have sprints, a groomed Product Backlog and Refinement process, burndown charts, and a stable velocity. They also have a 22% change-failure rate and a quoting engine everyone is afraid to touch. The process is healthy; the code is not. This is the exact shape of the problem XP addresses, and no amount of better sprint planning will touch it.
They keep Scrum and add the XP layer underneath it. Month one: characterization tests. Working in pairs, they wrap the twelve highest-traffic rating rules in tests that assert current behavior — not correct behavior, current behavior, because the first goal is a net, not a verdict. They wire the suite into CI with a hard rule that a red mainline stops all other work until it is green. The suite takes 90 seconds. Month two: TDD becomes mandatory for all new rules, and pairs refactor opportunistically inside whatever code their story touches — no separate cleanup tickets, no “refactoring sprint.” They adopt collective ownership and drop the informal convention that Priya owns the rating tables. They negotiate two half-days a week of an actual underwriter sitting with the team, which is not a true on-site customer but cuts the median requirement question from four days to fifteen minutes. They do not adopt full-time pairing; they pair on rating logic and anything touching money, and solo on UI and reporting. They also skip the system metaphor entirely.
Six months in, coverage on the rating path is 71%, the median rule change ships in two days instead of three weeks, and change-failure rate is 6%. Velocity, measured in Story Points and Estimation, initially dropped 30% during the test-writing months and then rose to roughly 1.4x the pre-XP baseline — the classic J-curve. Nothing about their Scrum practice changed: same sprints, same standups, same Sprint Retrospective. What changed was that the Definition of Done finally had a mechanism behind it. Asked what framework they use, the team says “Scrum.” They are running XP, and like most teams in that position, they do not know it.
Referenced by