Team Topologies and Conway's Law
Team Topologies and Conway’s Law
Definition: Conway’s Law states that any organization designing a system will produce a design whose structure mirrors the organization’s own communication structure. Team Topologies is a practical framework built on that observation, defining four fundamental team types and three interaction modes, and using team cognitive load as the constraint that sets viable team and service boundaries. Together they reframe architecture as an organizational problem: if you want a different system shape, change the team shape first. The technique for doing so deliberately is called the Inverse Conway Manoeuvre.
How It Works
The law itself, stated accurately
Melvin Conway published “How Do Committees Invent?” in 1968 (the paper was written in 1967 and rejected by Harvard Business Review before Datamation ran it). The claim is narrow and precise:
Any organization that designs a system will produce a design whose structure is a copy of the organization’s communication structure.
Three words in that sentence carry the weight and are routinely misread.
“Communication structure,” not org chart. The relevant graph is who actually talks to whom — who sits in the same standup, who shares a Slack channel, who can get a question answered in ten minutes versus ten days. A formal reporting line with no communication traffic does not shape the architecture. An undocumented back-channel between two engineers absolutely does. Remote and hybrid work changed the topology of these graphs far more than it changed the org chart.
“A copy,” not “an influence.” Conway’s argument is not soft. It is a homomorphism claim: every module boundary in the system corresponds to a communication boundary in the organization, because a design decision that spans two modules requires the people responsible for those modules to negotiate, and negotiation is expensive. Where communication is cheap, interfaces stay fluid and modules merge. Where communication is expensive, the interface calcifies into a hard boundary — and a hard boundary is exactly what a module is.
“Will produce,” not “should produce.” The law is descriptive, not normative. It makes no claim that the resulting architecture is good. It says only that this is what you will get, whether or not anyone intended it.
Why the coupling is causal, not coincidental
The mechanism is mundane. Consider two engineers who must decide whether a piece of validation logic lives in the API layer or the client. If they sit together, the conversation costs five minutes and they will happily move the code back and forth as they learn more. If they are on different teams in different time zones with different backlogs and different managers, the same conversation costs a week, a meeting, and possibly a ticket in someone else’s sprint. The rational response is to freeze an interface and stop asking. That frozen interface is now a permanent architectural seam — created not by design reasoning, but by coordination cost.
Repeat this across thousands of decisions over three years and the system’s decomposition is a fossil record of the organization’s communication costs at the moment each decision was made. This is why architecture diagrams from mature companies so often look like nobody designed them: nobody did. The org chart designed them.
Evidence in the wild
Conway’s own illustration was a compiler. Assign a compiler to four teams and you will get a four-pass compiler — not because four passes is optimal for that language, but because four passes is what four teams can build without renegotiating boundaries. Assign the same compiler to two teams and you get two passes.
The modern version is depressingly reliable. A company with a frontend team, a backend team, and a DBA team ships a three-tier architecture. Every time. Whether or not three tiers is right for the problem, whether or not a vertical slice per feature would deliver faster, whether or not the “backend” is actually two unrelated products wearing a trench coat. The architecture was decided by the recruiting plan.
| Organizational shape | Architecture you reliably get | Was it chosen? |
|---|---|---|
| Frontend / backend / DBA teams | Three-tier, with the seams at the team lines | No |
| One team per business capability | Services aligned to business capabilities | Usually yes |
| Offshore team owning “maintenance” | A legacy monolith nobody refactors | No |
| A separate QA department | Testing as a late gate, poor test automation in code | No |
| A separate ops department | Deploy as a handoff ritual, weak observability | No |
| One team owning a shared internal library | A library with a stable, well-documented API | Often yes |
| Four teams building a compiler | A four-pass compiler | No |
Notice the pattern in the third column. When teams are aligned to value (a business capability, a product, a user journey), the law works for you. When teams are aligned to technology layers or job functions, the law works against you and produces an architecture with seams in exactly the places where change is most frequent.
The Inverse Conway Manoeuvre
The actionable insight is that fighting the law is a losing game. Teams have tried it: architects publish a target diagram, the org stays the same, and eighteen months later the code has drifted back into the shape of the org chart. The law is not a tendency you can overcome with discipline. It is the aggregate of thousands of individually rational decisions about coordination cost.
The Inverse Conway Manoeuvre (coined by Thoughtworks, popularized by Team Topologies) inverts the causality deliberately:
- Decide what architecture you want — the module and service boundaries that match how the business actually changes.
- Restructure the teams so that their communication structure is a copy of that target architecture.
- Let the law do the work.
If you want six independently deployable services owned end to end, create six teams with clear, non-overlapping ownership and make cross-team communication deliberately more expensive than intra-team communication (via APIs and documentation rather than meetings). The architecture will converge on the shape you set up. You are not overriding Conway’s Law; you are aiming it.
Read that diagram in both directions. Left to right it is Conway’s Law: three teams produce three services with hard interfaces where the team boundaries are. Right to left it is the Inverse Manoeuvre: you wanted three services, so you built three teams.
The caveat that makes this honest: reorganizing people is disruptive, slow, and emotionally costly in a way that refactoring code is not. The manoeuvre is a heavy instrument. Use it when the architecture you have is genuinely blocking the business, not to chase a fashionable diagram.
Fracture planes: choosing where the seams go
If team boundaries become architectural boundaries, then choosing team boundaries is the architectural decision, and it needs the same rigor. Team Topologies calls the candidate split lines fracture planes — natural seams along which a system can be divided with minimal ongoing coupling. The list is worth memorizing, because most organizations only ever consider the first two and end up with layer-based teams.
| Fracture plane | Split along | Use when |
|---|---|---|
| Business domain | Bounded contexts, user journeys, capabilities | Almost always the first choice; changes cluster within a domain |
| Regulatory compliance | Systems under PCI, HIPAA, SOX scope | Compliance overhead would otherwise contaminate every team |
| Change cadence | Parts that change daily versus quarterly | Coupling a fast-moving surface to a slow core taxes both |
| Risk profile | Experimental versus safety-critical | Different testing, review, and release rigor per side |
| Performance isolation | Components with extreme latency or throughput needs | The specialist tuning work would swamp a generalist team |
| Technology | Genuinely different stacks (embedded, ML training) | Last resort — this is the plane that produces layer teams |
| User persona | Distinct user types with distinct workflows | The product truly serves separable audiences |
The diagnostic for a badly chosen plane is simple: count how many teams a typical feature touches. If the answer is routinely more than one, the seam is cutting across the direction of change rather than along it. Splitting by technology is the classic offender — every user-visible change crosses every layer, so every change crosses every team.
The good seams share a property: they follow how the business changes, not how the software is built. See Requirements Engineering for how domain boundaries surface in the first place, and Product Roadmap for the change-cadence signal.
Why It Matters
- It explains architecture failures that look technical but are not. A team that cannot ship independently usually has an organizational dependency, not a technical one. Rewriting the code will not fix it; the new code will grow the same seams.
- It makes microservice migrations predictable. Splitting a monolith without splitting the teams produces a distributed monolith: the same coupling, now over the network, with added latency and failure modes. Most failed microservice programs are this exact mistake.
- It converts a vague culture argument into a design lever. “We should communicate better” is unactionable. “Team A owns this service end to end and Team B consumes it through a versioned API” is a decision you can make on a Tuesday.
- It sets a ceiling on delivery speed. Flow through a system is limited by the number of hand-offs and wait states. Each cross-team dependency in the critical path adds queueing time measured in days, not hours, regardless of how fast anyone codes.
- It predicts where quality problems concentrate. Defects cluster at module boundaries that correspond to organizational boundaries, because that is where shared understanding is thinnest and where nobody feels full ownership.
- It exposes the true cost of shared ownership. A service owned by three teams has no owner. Every change requires negotiation, so nobody refactors, so entropy accumulates faster than in a service with one clear owner.
- It gives a principled basis for team size. Cognitive load, not headcount budgets, should determine how much system a team owns — which turns “how big should this team be?” into an answerable question.
- It reframes platform investment. A platform team exists to reduce the cognitive load of stream-aligned teams. That is a measurable purpose, unlike “own the infrastructure.”
- It gives architects a job that scales. Instead of drawing diagrams and enforcing them, architects design team boundaries and interaction modes — an intervention that keeps working after they leave the room.
Cognitive Load: The Real Constraint on Team Boundaries
Team Topologies borrows John Sweller’s cognitive load theory and applies it at team scale. A team can only own as much system as it can hold in its collective head. Exceed that and the team stops improving the system and starts merely surviving it.
Sweller’s three categories map cleanly onto software work:
| Load type | In software terms | What to do about it |
|---|---|---|
| Intrinsic | The inherent difficulty of the domain — distributed consensus, tax law, video codecs | Reduce with training, hiring, and better abstractions. It never goes to zero. |
| Extraneous | Accidental difficulty — flaky deploy scripts, tribal knowledge, five ways to provision a database | Eliminate ruthlessly. This is what a platform team exists to absorb. |
| Germane | The valuable work of building a mental model of the business domain | Protect and maximize. This is the load you want your team carrying. |
The design goal is simple to state and hard to do: drive extraneous load toward zero, keep intrinsic load within the team’s capability, and give the freed capacity to germane load.
Why exceeding it is the real reason large services rot
The usual story about an unmaintainable service is technical: bad abstractions, missing tests, accumulated shortcuts. The organizational story underneath is that the system grew past the point where any one team could model it. Once that happens, the failure mode is mechanical:
- No individual understands the whole, so every change is made locally with incomplete information.
- Local changes accumulate contradictions the team cannot see, because seeing them would require the mental model nobody has.
- Refactoring becomes too risky to attempt, because nobody can predict blast radius.
- The team optimizes for not breaking things rather than for improving them.
- Delivery slows, so more people are added — which increases communication paths and makes the shared model harder, not easier, to maintain.
Step 5 is Brooks’s Law arriving as a direct consequence of exceeded cognitive load. See Code Refactoring and Technical Debt for the code-level view of the same spiral.
Measuring it well enough to act
You cannot measure cognitive load precisely, but you can measure it usefully. The Team Topologies authors suggest simply asking the team, on a scale of one to five: “Is the domain(s) you work on too large for the team to handle comfortably?” Aggregated over time, that self-report is a better signal than any proxy metric. Supporting indicators:
| Signal | What high load looks like |
|---|---|
| Onboarding time | New engineer takes more than 6-8 weeks to ship independently |
| Bus factor | Only one person can safely change large parts of the system |
| Context switching | Team handles more than 2-3 distinct domains in one sprint |
| Interrupt rate | Support and questions consume more than 20-30% of capacity |
| Change failure rate | Rising, with post-mortems that say “we didn’t know that was connected” |
| Refactoring frequency | Near zero, with the team describing areas as “scary” |
The practical rule: limit the number of domains per team. A team can typically hold one complicated domain, or two to three simple ones — not both. If a team owns a complicated domain, that is the whole job.
Why the two-pizza team heuristic works
Amazon’s rule — a team small enough to be fed by two pizzas, roughly 5-9 people — is usually explained as being about meeting efficiency. That is a symptom, not the cause.
The real reason is that communication paths grow as n(n-1)/2:
| Team size | Communication paths | Practical effect |
|---|---|---|
| 3 | 3 | Everyone knows everything, effortlessly |
| 5 | 10 | Shared mental model holds naturally |
| 8 | 28 | Holds with intentional practices (pairing, rotation, docs) |
| 12 | 66 | Sub-groups form; the shared model starts to fracture |
| 20 | 190 | Two teams pretending to be one; the split is already happening informally |
| 50 | 1225 | Formal process replaces shared understanding entirely |
Around 8-9 people the cost of maintaining a shared mental model starts to exceed the benefit of the extra hands. Dunbar’s numbers make the same point at larger scales — roughly 15 for close trust, 50 for a group with a shared identity, 150 for a group where people know each other by name. Team Topologies uses those thresholds to size not just teams but the groupings above them.
Crucially, the two-pizza rule constrains team size and therefore system size, because a team should own no more system than it can model. The heuristic is really a cap on service scope wearing a catering metaphor. Amazon paired it with a second, less-quoted rule that made it work: every team owns its services in production, end to end. Small team plus full ownership plus bounded scope is the actual formula. See DevOps Culture for why the ownership half is non-negotiable.
The Four Fundamental Team Types
Team Topologies (Skelton and Pais, 2019) argues that the sprawl of team names in most organizations — component teams, feature teams, infrastructure, tooling, architecture, SRE, CoE — collapses into exactly four types. Anything else is one of these four with a confusing name, or a team with no clear purpose.
| Team type | Purpose | Owns | Duration | Success looks like |
|---|---|---|---|---|
| Stream-aligned | Deliver a continuous flow of value for one business domain or user journey | A slice of the product, end to end, in production | Long-lived | Ships independently without waiting on other teams |
| Enabling | Help stream-aligned teams acquire missing capabilities | Nothing in production | Long-lived team, short engagements (weeks) | It makes itself unnecessary and leaves |
| Complicated-subsystem | Own a part requiring deep specialist knowledge | One hard subsystem behind a clean interface | Long-lived, but justified continuously | Stream teams consume it without understanding it |
| Platform | Reduce the cognitive load of stream-aligned teams | Internal products with self-service APIs | Long-lived | Teams choose to use it because it is genuinely easier |
Stream-aligned is the default and should be the overwhelming majority — typically 70-90% of your engineering teams. Every other type exists to make stream-aligned teams more effective. If you have more platform and enabling teams than stream-aligned ones, you have built an organization that services itself.
Enabling teams are the most commonly missing type. They are coaches, not doers: a testing-practices enabling team spends six weeks with a stream team, raises their capability with Test Pyramid and TDD, and leaves. The failure mode is an enabling team that starts doing the work instead of teaching it, at which point it becomes a permanent dependency and a bottleneck.
Complicated-subsystem teams should be rare and justified. The bar is that the subsystem needs genuine specialist knowledge — a pricing engine with actuarial math, a physics kernel, a video transcoding pipeline, a hardware driver. “This code is messy” is not sufficient justification; that is a refactoring problem. Each such team is a permanent dependency, so create one only when the specialist knowledge is real and irreducible.
Platform teams must be judged as product teams. The measure is adoption without mandate. If stream teams have to be forced onto the platform, the platform is adding cognitive load rather than removing it. The Team Topologies phrase is “thinnest viable platform” — build the smallest thing that removes real friction, which is often a curated wiki page and three scripts before it is a control plane. Tooling such as CI-CD pipelines, Docker base images, and deployment templates is typical platform surface area.
The Three Interaction Modes
Team Topologies is equally insistent that team interactions be explicit and few. Every pair of teams should be in exactly one of three modes at a time, and the mode should be a deliberate, stated choice with an expected duration.
| Mode | What it means | Communication cost | Right duration | Best for |
|---|---|---|---|---|
| Collaboration | Two teams work closely together, shared responsibility, blurred boundary | High | Weeks to a few months, time-boxed | Discovering an unknown interface; rapid innovation at a new boundary |
| X-as-a-Service | One team consumes something the other provides, via a clear versioned interface | Low | Indefinite — this is the steady state | Predictable delivery at scale; anything commoditized |
| Facilitating | One team helps another improve, without doing the work | Medium | Weeks, explicitly time-boxed | Capability gaps, adopting a new practice, unblocking a struggling team |
The design intent is to move interactions toward X-as-a-Service over time. Collaboration is valuable precisely when you do not yet know where the boundary should be — two teams work in the messy middle, discover the right interface, then stop collaborating and formalize it. Collaboration that never ends is not collaboration; it is two teams that should be one team, or a boundary drawn in the wrong place.
Three concrete failure signatures:
- Permanent collaboration. Two teams in a standing weekly sync for eighteen months. Either merge them or find the interface they have been avoiding.
- X-as-a-Service without service quality. A platform team declares itself as-a-Service but has no docs, no SLOs, and answers questions in a ticket queue with a four-day latency. Consumers route around it, and the promised load reduction never materializes.
- Facilitating that turned into doing. The enabling team took the work “just this once” and is now on the critical path for every release.
A useful discipline: write the interaction mode down, next to each dependency, with an expiry date. Modes that have no expiry and are not X-as-a-Service are organizational debt.
The Team API
X-as-a-Service only reduces cognitive load if consuming the service is genuinely cheaper than understanding it. Team Topologies makes this concrete with the idea of a Team API — the full published surface a team exposes to everyone else, of which code is only one part:
| Element | What it answers | Typical failure |
|---|---|---|
| Code interfaces | How do I call this? | Undocumented endpoints; breaking changes without Semantic Versioning |
| Documentation | How do I use it correctly? | A wiki last edited two years ago |
| Ways of working | How do I get a change made? | “File a ticket” with no visible queue or SLA |
| Roadmap | What is coming, and should I wait? | No public plan, so consumers build workarounds |
| Support model | Who do I ask, and how fast will they reply? | A private channel and one overloaded person |
| Environments | Can I try it without a meeting? | A sandbox that requires manual provisioning |
The bar is that a consuming team should be able to go from “I need this” to “it works in my staging environment” without talking to a human. Every step that requires synchronous conversation is a step that reintroduces coordination cost — which, by Conway’s Law, will eventually reintroduce coupling. Techniques like Feature Flags and versioned contracts exist largely to let a provider change without scheduling a conversation with every consumer.
A useful exercise: ask each team to write its Team API on one page, then ask its three biggest consumers to grade it. The gap between the two documents is a precise measure of how much load that team is actually offloading versus how much it merely claims to.
Comparison
| Concept | Core claim | Unit of change | Primary constraint | Relationship |
|---|---|---|---|---|
| Team Topologies / Conway’s Law | Architecture mirrors communication structure; change teams to change architecture | Team boundaries and interaction modes | Cognitive load per team | The frame this note describes |
| Scaled Agile (SAFe and LeSS) | Large-scale delivery needs explicit coordination ceremonies across many teams | Cadence, ceremonies, and roles layered above teams | Alignment across a large program | Adds coordination on top of existing teams |
| DevOps Culture | The dev/ops handoff is the bottleneck; the team that builds it runs it | Ownership and incentives | Feedback loop latency | Supplies the incentive argument for stream-aligned ownership |
| Microservices | Independently deployable services enable independent delivery | Service boundaries in code | Distributed-system complexity | The architecture Conway’s Law will only give you if the org matches |
| Domain-Driven Design | Model boundaries should follow business domains (bounded contexts) | Bounded contexts in the model | Domain complexity | Supplies the content of the boundaries; Team Topologies supplies the people |
The sharpest contrast: Team Topologies versus SAFe
Both frameworks answer the same question — how do you deliver software with 200 engineers instead of 20? — and they answer it in opposite directions.
Scaled Agile (SAFe and LeSS) treats the existing team structure as given and adds machinery to coordinate it: Agile Release Trains, PI Planning, Release Train Engineers, a Program Board, System Demos, an Inspect and Adapt cadence. If forty teams have dependencies on each other, SAFe gives you a two-day quarterly event to surface and sequence those dependencies, a visible board to track them, and a role whose job is to chase them.
Team Topologies makes the opposite move: if you need a two-day event to untangle dependencies, the dependencies are the problem, not the scheduling of them. Redraw the team boundaries so most of those dependencies disappear, and the coordination machinery becomes unnecessary.
| Dimension | SAFe / LeSS | Team Topologies |
|---|---|---|
| Treats dependencies as | A fact to be managed | A defect to be designed out |
| Primary intervention | Add coordination ceremonies and roles | Change team boundaries and ownership |
| Optimizes for | Alignment and predictability across many teams | Fast flow within an autonomous team |
| Signature artifact | PI Planning, the Program Board | Team API, interaction-mode map |
| Cost when misapplied | Coordination overhead becomes the work; teams plan more than they build | Reorgs that disrupt people without fixing the underlying boundary |
| Underlying view of Conway’s Law | Implicit; the org shape is a given | Explicit; the org shape is the lever |
The fair reading is that they operate at different layers and are not strictly exclusive — LeSS in particular shares Team Topologies’ preference for feature teams over component teams, and plenty of organizations run PI Planning while quietly redrawing team boundaries underneath it. But when they conflict, the diagnostic question is sharp: are we scheduling this dependency, or eliminating it? Every quarter spent scheduling a dependency you could have eliminated is a quarter of compounding coordination cost. Coordination overhead scales superlinearly with team count; autonomy scales flat.
Real-World Use Cases
- Amazon’s 2002 services mandate. Bezos required every team to expose data and functionality only through service interfaces, with no direct database access and no back doors, and to design those interfaces as if they would be externalized. This is a pure Inverse Conway Manoeuvre: it changed how teams could communicate, and the service-oriented architecture followed. AWS emerged from the platform capability it forced into existence.
- Netflix’s full-cycle developers. Netflix moved from a central operations team to teams owning build, deploy, run, and support for their services, backed by a strong internal platform. The architecture (small, independently deployable services with strong resilience patterns) is a direct consequence of the ownership model.
- Spotify’s squad model. Widely copied and widely misunderstood — even Spotify has said the model was a snapshot, not a prescription. The genuinely transferable ideas were vertical squads owning a user-facing slice, and chapters/guilds as low-bandwidth knowledge networks that did not become delivery dependencies. Copying the vocabulary without the ownership produces nothing.
- The distributed monolith. A company splits a monolith into thirty services but keeps the same three layer-based teams. Every feature still touches all three teams, releases are still coordinated, and now there are network calls in the middle. The org did not change, so the architecture did not really change either. This is the single most common microservices failure.
- Platform team as a rebranded ops team. The ops team gets renamed “Platform,” keeps its ticket queue, and now product teams file tickets to get a database. Cognitive load did not move; only the sign on the door did. The test is self-service: can a stream team provision what it needs, at 2am, without a human in the loop?
- The pricing engine as a complicated-subsystem team. An insurance company keeps four actuarial specialists as a dedicated team owning the rating engine behind a clean API. Product teams call it without understanding the math. This is the correct use of the pattern: the specialist knowledge is genuine and irreducible.
- Enabling team for test automation. A six-week engagement embeds two coaches with a struggling team, establishes patterns from Test Pyramid and TDD and Code Review and Static Analysis, then leaves. Six months later the team’s change failure rate is down and the coaches are not in the room.
- Merging two teams that could not stop collaborating. A checkout team and a payments team spend a year in permanent sync. The fix was not better meetings; it was merging them into one stream-aligned team owning the whole purchase journey, and later splitting on the boundary they discovered rather than the one they inherited.
- Post-acquisition integration. The acquired company’s product remains architecturally separate for years because the teams remain organizationally separate. Integration timelines that ignore this reliably slip; the ones that succeed start by merging teams around shared domains rather than by writing adapters.
- Conway’s Law in open source. A project with a single maintainer has a tightly coupled core. Once maintainership splits across subsystems, the plugin boundaries harden into a stable API — often within one release cycle, and usually without anyone writing an architecture document.
Common Pitfalls
- Reorganizing without a target architecture. A reorg with no explicit architectural goal is just churn. Decide what boundaries you want first; the org change is the mechanism, not the objective. Reorgs are expensive in trust — spend that budget once, deliberately.
- Splitting the code without splitting the teams. The most expensive microservices mistake. If three teams still touch every service, you have paid for distribution and received none of the autonomy. Conway’s Law will pull the coupling straight back.
- Calling everything a platform team. Renaming the ops team does not reduce anyone’s cognitive load. A platform is a product with users who could choose otherwise, self-service interfaces, documentation, and an SLO. Without those, it is a ticket queue with better branding.
- Letting enabling teams become permanent dependencies. An enabling team that does the work rather than teaching it turns into a bottleneck, and worse, into a reason the stream team never learns. Time-box every engagement and define its exit condition on day one.
- Ignoring cognitive load when assigning ownership. “Team A can take this too” is how a functioning team becomes a firefighting team. Every new domain added to a team’s plate should displace something, or the team’s germane capacity — the part that actually improves the system — goes to zero first.
- Copying Spotify’s model as a template. Squads, tribes, chapters, and guilds were one company’s snapshot at one moment. Copying the names without copying the ownership, autonomy, and platform investment produces the same architecture you had, with new vocabulary.
- Treating the Inverse Conway Manoeuvre as free. Moving people has real human cost: broken relationships, lost context, damaged trust, six months of reduced output. Use it when the architecture is genuinely blocking the business, not to chase a diagram.
- Assuming the org chart is the communication structure. The real graph is who talks to whom. Two teams that share a lunch table are effectively one team. A “shared ownership” arrangement with a monthly sync is effectively no ownership at all. Map the actual traffic before redrawing anything.
- Shared ownership of a critical service. Three teams owning one service means no team owns it. Nobody refactors, nobody sets standards, and everyone assumes someone else is watching the error rate. One owner, published interface, consumers as clients.
- Forgetting that the law applies to the fix. The architecture team that redesigns your system will produce a design mirroring their communication structure, not yours. Design decisions made by a group disconnected from delivery reliably produce boundaries the delivery teams cannot honor.
Related Terms
- DevOps Culture — the incentive argument for end-to-end ownership; stream-aligned teams are unworkable without it
- Scaled Agile (SAFe and LeSS) — the contrasting approach: coordinate existing teams rather than redraw them
- Scrum — the delivery cadence a stream-aligned team most often runs internally
- Kanban — makes cross-team dependency wait time visible, which is the evidence a boundary is wrong
- Lean Software Development — the flow-efficiency thinking that Team Topologies applies at the organizational level
- Code Refactoring and Technical Debt — the code-level symptom of a team that has exceeded its cognitive load
- CI-CD Best Practices — independent deployability is what makes team autonomy real rather than nominal
- Sprint Retrospective — where teams surface the dependency pain that signals a boundary problem
Example
A logistics company had 90 engineers organized by technology: a mobile team, a web team, an API team, a data team, and a small “core services” group. Every feature — say, adding proof-of-delivery photos — required a ticket in four backlogs, four sprint plannings, and a release coordinated across four teams. Lead time from idea to production averaged eleven weeks, of which fewer than three were spent writing code; the rest was queueing. Predictably, the architecture was a four-tier stack with a shared database, and the seams were exactly where the team boundaries were. Nobody had designed it that way. The recruiting plan had.
The first attempt at a fix was architectural: a target diagram of twelve domain services, an ADR process, and a migration plan. Eighteen months later, four services existed, all four called the same shared database, and the API team still gated every change. The diagram had lost to the org chart, as it always does.
The second attempt inverted the causality. Leadership picked six value streams from the customer’s perspective — driver experience, shipper booking, tracking and notifications, billing, warehouse operations, and partner integrations — and rebuilt teams around them, each with mobile, backend, and data skills inside the team. Two supporting teams were created: a platform team responsible for deployment pipelines, observability, and environment provisioning as genuine self-service, and a complicated-subsystem team of four owning the route optimization engine, which required operations-research expertise nobody else had. Interaction modes were written down explicitly: platform to stream teams as X-as-a-Service, route optimization as X-as-a-Service behind a versioned API, and one deliberately time-boxed eight-week collaboration between the billing and partner-integrations teams, who did not yet know where their shared boundary belonged.
The first four months were worse, not better. Velocity dropped as people learned unfamiliar parts of the system, and two engineers left. By month nine, five of the six streams could deploy without coordinating with another team, and lead time had fallen to nine days. The database was still shared in three places — the last coupling to go — but the architecture had begun to look like the six-service diagram the first attempt had failed to produce, and nobody had needed to enforce it. The billing and partner-integration teams ended their collaboration at week ten with an agreed interface and stopped meeting. The most telling metric was not deployment frequency but onboarding: a new engineer’s time to first production change dropped from seven weeks to six days, because there was finally an amount of system small enough for one person to hold in their head.
Referenced by