Velocity and Burndown Charts

Velocity and Burndown Charts

Definition: Velocity is the amount of estimated work — usually measured in story points — that a specific team completes in a single iteration, used as an empirical input for forecasting how much future work will fit into future iterations. A burndown chart plots remaining work against time within a sprint or release, showing whether the team is trending toward finishing. Together they form the core empirical feedback loop of iterative planning: velocity says how fast this team historically moves, and the burndown says whether the current iteration matches that history. Neither is a measure of productivity, quality, or effort, and treating them as such is the single most common way teams destroy their own planning data.

How It Works

Computing Velocity

Velocity is deliberately crude. At the end of a sprint, sum the estimates of every backlog item that met the Definition of Done. That sum is the velocity for that sprint. Nothing else enters the calculation.

Three rules make the number meaningful:

  1. Binary credit. An item is either done or it is not. A story estimated at 8 points that is 90% complete contributes zero. There is no partial credit, ever. Partial credit is the fastest way to make a burndown lie, because “90% done” is the most unreliable self-report in software.
  2. Original estimate, not revised. If the team estimated a story at 5 points and it turned out to be a 13-point slog, it still counts as 5. The estimate is what was used for planning; recording the actual would make velocity a measure of effort rather than a calibration of estimates.
  3. Only this team, only this cadence. Velocity is unitless and team-relative. A velocity of 34 has no meaning outside the team that produced it, and it silently changes meaning if the team’s composition, sprint length, or estimation baseline changes.

Formally, for sprint ii with completed items DiD_i:

vi=∑s∈Dipoints(s)v_i = \sum_{s \in D_i} \text{points}(s)

The rolling band that matters for forecasting is not the mean but the empirical range across the last kk sprints (typically k=6k = 6 to 1212):

vmin⁡=min⁡(vi−k+1,…,vi)vmax⁡=max⁡(vi−k+1,…,vi)v_{\min} = \min(v_{i-k+1}, \ldots, v_i) \qquad v_{\max} = \max(v_{i-k+1}, \ldots, v_i)

Some teams trim the extremes — drop the single best and single worst sprint before taking the range — to avoid letting one catastrophic outage sprint or one unusually clean sprint dominate the forecast. That is a reasonable adjustment as long as it is applied consistently and not re-tuned every time the answer is inconvenient.

Reading a Burndown

A sprint burndown has time on the x-axis (usually days of the sprint) and remaining work on the y-axis. Remaining work is the sum of estimates for items not yet done. Two lines appear:

  • The ideal line, a straight descent from the sprint commitment to zero on the last day. This is a reference, not a target. No real sprint follows it.
  • The actual line, which steps down each time an item reaches done. Because credit is binary, the actual line is a staircase, not a smooth curve. Larger stories make bigger, less frequent steps.

The information in a burndown is entirely in the shape of the actual line relative to the ideal, not in whether it ends at zero. A sprint that ends at 3 points remaining but descended steadily is healthier than one that ends at zero after a vertical cliff on the final day.

The first line is the actual, the second is the ideal. Note the plateau at days 5 through 7: nothing finished for three days. That plateau is the most useful part of the chart, because it is a question the team can answer at standup. Maybe a dependency was blocked. Maybe three people were all mid-story simultaneously. Maybe one story was badly estimated. The chart does not tell you which — it tells you where to look.

What Goes on the Y-Axis

The y-axis choice changes what the chart can tell you, and teams often inherit one without deciding.

Story points. The default, and the only option consistent with velocity. Because credit is binary and stories vary in size, the line is a coarse staircase. This is a feature: it makes work-in-progress problems visible as long flat stretches. The drawback is that a sprint with one 13-point story and one 2-pointer produces a chart with almost no resolution.

Remaining task hours. Some teams decompose stories into tasks with hour estimates and burn down hours. The line is much smoother and updates daily, which feels more informative. It usually is not. Hour estimates are re-estimated daily by the person doing the work, so the chart tracks confidence rather than completion, and the classic failure is a task stuck at “2 hours remaining” for four days. Hours also tempt teams into tracking individual utilization, which is a different metric wearing the same clothes.

Item count. Ignore estimates entirely and burn down the number of unfinished stories. Crude, but honest, and surprisingly accurate when the team slices work to consistent small sizes. This is the burndown equivalent of the Kanban throughput approach and requires no estimation at all.

The practical guidance: burn down points at the story level, and if you want daily resolution, get it from smaller stories rather than from hour-level task decomposition. A sprint containing eight 2-to-5-point stories produces a legible chart with no extra bookkeeping; a sprint containing three 13-pointers needs hour-tracking to look like anything, which is a signal to slice differently rather than to instrument harder.

Burnup vs Burndown

A burnup chart plots the same sprint or release with two separate lines: completed work rising from zero, and total scope as its own line across the top. The gap between them is the remaining work.

This decomposition is why burnup is strictly more informative. In a burndown, a single line encodes two independent quantities — work finished and work added — and they cancel. If the team completes 5 points and someone adds 5 points of scope on the same day, the burndown is perfectly flat and looks like a stalled team. In a burnup, the completed line rises by 5 and the scope line rises by 5, and both facts are visible.

For release-level tracking across many sprints, burnup is close to mandatory. Releases almost always change scope, and a release burndown that mixes discovery-driven scope growth with delivery progress produces a chart nobody trusts and everybody ignores. For a two-week sprint where scope is supposed to be stable, burndown is simpler and adequate — and a rising burndown is itself a useful alarm that the sprint boundary was violated.

BurndownBurnup
Lines plottedRemaining work (one line)Completed work and total scope (two lines)
Scope changeHidden — absorbed into the single lineExplicit — the scope line moves
Best forSingle sprint with fixed scopeReleases, epics, multi-sprint work
Reads “done” asLine reaches zeroCompleted line meets scope line
Failure modeFlat line is ambiguousSlightly busier to read

Forecasting With a Range, Not an Average

The point of collecting velocity is to answer “when will this be done?” The wrong way to answer is with a single number derived from a mean. A mean produces a single date, a single date reads as a commitment, and a commitment derived from noisy historical data is a promise the team did not make.

The right method: take the last several sprints, use the low and high of that band, and produce two dates. Report the range.

Given remaining scope RR points and a velocity band [vmin⁡,vmax⁡][v_{\min}, v_{\max}], the sprint count is:

npessimistic=⌈Rvmin⁡⌉noptimistic=⌈Rvmax⁡⌉n_{\text{pessimistic}} = \left\lceil \frac{R}{v_{\min}} \right\rceil \qquad n_{\text{optimistic}} = \left\lceil \frac{R}{v_{\max}} \right\rceil

Worked example. A team’s last eight sprints, most recent last:

Sprint1415161718192021
Velocity2934223136273330

The mean is 30.25. The temptation is to divide and report a date. Instead, look at the band: the low is 22, the high is 36. Trimming the single best and worst gives a working band of 27 to 34.

Remaining scope on the release backlog is 248 points. Sprints are two weeks. Today is the start of sprint 22.

  • Pessimistic: ⌈248/27⌉=⌈9.19⌉=10\lceil 248 / 27 \rceil = \lceil 9.19 \rceil = 10 sprints → 20 weeks
  • Optimistic: ⌈248/34⌉=⌈7.29⌉=8\lceil 248 / 34 \rceil = \lceil 7.29 \rceil = 8 sprints → 16 weeks

The honest statement is: “16 to 20 weeks, assuming scope does not change.” That last clause is not boilerplate. It is the dominant source of error, and it is why the same team should also be tracking scope growth. If the release backlog has grown an average of 6 points per sprint for the last five sprints, the effective velocity against a moving target is not 27–34 but 21–28, and the forecast stretches to 9–12 sprints. Modeling scope growth explicitly turns a forecast that is quietly wrong into one that is roughly right.

The unestimated tail. Real release backlogs are never fully estimated — the near items are refined and pointed, the far items are titles. Dividing only the estimated portion by velocity produces a forecast that is confidently short, because it silently assumes the unrefined half is free. Two defensible fixes: assign the unestimated items the median point value of comparable already-estimated items and flag them as low confidence, or forecast only to the boundary of what is estimated and state plainly that everything beyond it is unforecast. Both are honest. Quietly ignoring the tail is not, and it is the most common reason a forecast that looked solid in month one is off by 40% in month four.

A more rigorous version replaces the band with a Monte Carlo simulation: sample historical velocities with replacement a few thousand times, accumulate until scope is exhausted, and read percentiles off the resulting distribution. The output is “85% confidence we finish within 11 sprints,” which is a genuinely better artifact than a range. But the simple band captures most of the value and requires no tooling, so start there.

The bar chart makes the variability legible in a way the mean never does. Sprint 16 at 22 points and sprint 18 at 36 points differ by more than 60%. Any forecast that pretends this team reliably delivers 30 points is fiction dressed as arithmetic.

Stabilization and When Velocity Becomes Usable

Velocity is not usable data until it stabilizes, which typically takes three to six sprints with a stable team. Before that, the estimation baseline is still moving — the team has not yet converged on what a “3” means — so early velocities measure calibration drift as much as delivery.

Velocity also resets, silently, whenever any of the following change: team membership, sprint length, the Definition of Done, or the estimation scale. A team that adds two engineers does not get a proportionally higher velocity next sprint; it usually gets a lower one for two or three sprints while onboarding costs are paid, then a higher one. Reading the dip as a productivity decline is a misread of an entirely predictable ramp.

Why It Matters

  • It replaces negotiation with observation. Before empirical velocity, delivery dates came from whoever argued hardest in the room. Velocity makes the conversation about what the team has actually done for the last eight sprints, which is much harder to argue with than an opinion.
  • It makes forecast uncertainty explicit rather than hidden. A single-date commitment hides its own error bars. A range from a velocity band puts the uncertainty on the table where stakeholders can plan around it.
  • It surfaces problems mid-sprint instead of on the last day. A burndown that plateaus on day 4 is a signal available six days before the sprint review. Without the chart, the same problem surfaces as a surprise at demo.
  • It quantifies the cost of scope changes. When someone adds work mid-release, a burnup shows the scope line stepping up and the projected completion date moving with it. That converts “can we just squeeze this in?” from a social question into an arithmetic one.
  • It exposes work-in-progress problems. A staircase with rare, huge steps means the team runs many items in parallel and finishes them all at the end. That is a flow problem, and the burndown shape is often the first place it becomes visible.
  • It anchors Sprint Retrospective discussions in data. “Sprint 16 dropped to 22 — what happened?” is a far more productive retro opener than “how did everyone feel about the sprint?”
  • It makes estimation calibration checkable. If sprint after sprint the team commits to 40 and delivers 28, the problem is not effort — it is a systematically optimistic estimation baseline, and velocity is the evidence.
  • It supports capacity-aware planning. With a known band, a team can plan a sprint with two people on vacation by scaling expected capacity, rather than committing to the usual load and failing.
  • It gives Product Roadmap conversations a reality check. A roadmap with four quarters of features and a team velocity implying eighteen months of work is a roadmap that will break — and the arithmetic proves it before anyone commits publicly.

Diagnostic Patterns: Reading the Shape

The chart’s value is in its shape. These are the recurring shapes and what each one actually means.

The Cliff — Flat Early, Vertical at the End

The line barely moves for the first 70% of the sprint, then drops almost to zero on the last day or two.

What it means: work was integrated late. Everything was in progress simultaneously and nothing crossed the done line until the final push. This is the most common unhealthy pattern and it is frequently misread as a strong finish.

Why it is dangerous: the “done” on the last day is usually not real done. Testing was compressed, code review was rubber-stamped, and the last item merged an hour before demo. The defects surface next sprint, where they consume capacity that nobody planned for. The cliff sprint effectively borrows against the next one.

What to do: cap work in progress. Reduce the number of items a team can have open at once — if three people can only have two stories in flight, the third person pairs or reviews instead of starting a fourth. Split large stories so they can finish mid-sprint. Verify the Definition of Done was genuinely met for the last-day items, and if it was not, carry them and take the velocity hit honestly.

The Rise — The Line Goes Up

Remaining work increases at some point mid-sprint.

What it means: scope was added after the sprint started, or an item was re-estimated upward, or a story was split into pieces summing to more than the original.

Why it matters: a sprint’s central promise is a stable scope boundary for a short, fixed window. A rising burndown is direct evidence that the boundary was violated. It is not automatically bad — an urgent production incident is a legitimate reason — but it should never be invisible.

What to do: find out who added the work and through which path. If it was an emergency, name it and consider whether the team needs an explicit reserve capacity for interrupts. If it was a stakeholder inserting a “small” request directly to an engineer, that is a process failure to fix in the retro, not an estimation failure. And if the sprint is now clearly overloaded, remove something explicitly rather than letting the team quietly fail on everything.

The Perfect Line

The actual line tracks the ideal almost exactly, hitting zero precisely on the last day, sprint after sprint.

What it means: almost certainly not excellence. Real work is lumpy — stories are different sizes, dependencies block, people get sick, and code review adds latency. A staircase that perfectly mimics a straight line is a statistical improbability repeated.

Why it is suspicious: the realistic explanations are all forms of number management. Items are being marked done before they meet the Definition of Done to keep the line on track. Estimates are being padded so the sprint is comfortably under-committed and the line lands neatly. Someone is editing remaining-work numbers to match the ideal. Or the team is sandbagging: consistently committing to less than they can do, which makes the chart pretty and the forecast useless.

What to do: stop treating a smooth chart as a success signal, and say so out loud — the shape is often smooth precisely because someone believes it is being graded. Check whether committed scope has crept downward over time. Compare estimates against actual cycle times for a few stories. If items are being closed prematurely, tighten the done criteria and expect the chart to get uglier and more honest.

The Plateau

The line drops normally, then goes flat for several consecutive days, then resumes.

What it means: something blocked. A dependency on another team, a broken build, an environment outage, a decision waiting on someone unavailable, or a single story that was drastically underestimated and is consuming everyone.

What to do: the plateau’s start date is the diagnostic. Look at what changed that day. Blocked-on-external-dependency plateaus are a Team Topologies and Conway’s Law problem and will recur until the dependency is restructured. A plateau caused by one giant story is an estimation and slicing problem. Both are fixable; neither is fixed by asking people to work harder.

The Slow Bleed

The actual line descends but consistently sits above the ideal, ending the sprint with a meaningful chunk of work unfinished — every sprint.

What it means: systematic over-commitment. The team consistently pulls more than its velocity band supports.

What to do: this is the easiest pattern to fix and the most commonly ignored. Plan to the low end of the velocity band for two or three sprints. Finishing early is not a failure; it means the team can pull one more item. Chronically finishing late erodes trust with stakeholders far more than a smaller, reliable commitment does.

The Early Zero

The line reaches zero on day 7 of a 10-day sprint.

What it means: under-commitment, or a sprint where several items turned out much simpler than estimated.

What to do: once is noise; three times is a pattern worth correcting. Pull the next item from the Product Backlog and Refinement rather than idling, and adjust the next sprint’s commitment upward. Watch for the alternative explanation: the team is deliberately sandbagging because velocity is being used to evaluate them — which is the failure mode in the next section.

Velocity Is a Planning Tool, Not a Performance Metric

This is the single largest failure mode, and it is worth being blunt: the moment velocity is used to evaluate a team, it stops measuring anything.

The mechanism is Goodhart’s Law — when a measure becomes a target, it ceases to be a good measure. Velocity is unusually vulnerable because the team controls both sides of the ratio. They produce the estimates and they deliver against them. There is no external calibration. If a manager asks for velocity to increase 20% next quarter, the team can satisfy that request perfectly without changing a single thing about how they work: estimate what used to be a 3 as a 5. The number rises, the manager is satisfied, and precisely nothing has improved. Worse, the team’s forecasting ability is now degraded, because the historical data no longer shares a baseline with new estimates.

Three specific anti-patterns follow from targeting velocity:

Cross-team comparison. Velocity is unitless and team-relative. Team A’s 40 and Team B’s 25 are measured in different, incomparable currencies. Ranking teams by velocity, putting velocity in a dashboard next to team names, or setting a org-wide velocity target rewards whichever team inflates fastest. Within six months every team’s estimation baseline has drifted independently and the numbers are pure noise. Any Scaled Agile (SAFe and LeSS) rollout that aggregates velocity across teams into a single portfolio number has built a metric with no denominator.

Velocity as an improvement goal. “Increase velocity” is not an improvement goal, it is an instruction to inflate estimates. Real improvement goals are things like reducing cycle time, reducing defect escape rate, reducing the number of items carried between sprints, or reducing the size of the largest story. Those are externally observable and much harder to game. If velocity rises as a side effect of removing a bottleneck, that is genuine — but it is the bottleneck removal you were measuring, not the velocity.

Velocity in individual performance reviews. Velocity is a team-level property emerging from collaboration, pairing, review, and mutual help. Attributing it to individuals penalizes exactly the behaviors that make a team effective: the engineer who spends two days unblocking three colleagues has a terrible individual velocity and enormous individual value.

The healthy framing: velocity is the team’s own instrument, owned by the team, used to answer their own planning questions. It should be visible to stakeholders as a forecasting input — “based on the last eight sprints, this is 16 to 20 weeks” — and never as a scorecard. When leadership asks about it, translate the question into a forecast rather than reporting the raw number.

The Flow Metrics Alternative

Kanban takes a different position: skip estimation entirely and measure the work items themselves as they move through the system.

  • Cycle time — elapsed time from when work starts on an item to when it is done. Measured in days, directly, with no estimation step.
  • Throughput — number of items completed per unit time. Just a count.
  • Work in progress — how many items are in flight right now.
  • Age of in-flight work — how long each currently-open item has been open, which is the most actionable real-time signal in the set.

These relate through Little’s Law:

Cycle Time=Work In ProgressThroughput\text{Cycle Time} = \frac{\text{Work In Progress}}{\text{Throughput}}

which yields the most reliable lever in the whole discipline: to reduce cycle time, reduce WIP. This holds regardless of estimation practice.

Flow metrics have three advantages over velocity. First, no estimation step, so nothing to inflate — cycle time is measured by timestamps, not judgment, which makes it substantially harder to game. Second, they produce distributions rather than averages: instead of “typically 4 days,” you can say “85% of items finish within 9 days,” which is directly usable as a service-level expectation. Third, they work with continuous flow and do not require a fixed iteration boundary.

The cost is that flow forecasting assumes items are roughly comparable in size. If a team’s work ranges from one-hour tweaks to three-week rewrites, item counts forecast badly. In practice, teams that slice work consistently small find item-count forecasting is as accurate as points-based forecasting — which is a strong argument for slicing small in the first place.

Many mature teams end up using both: Story Points and Estimation for coarse release-level sizing where the work is still vague, and cycle time plus throughput for operational decisions where the work is concrete. The two are not in conflict.

Adjacent Charts Worth Knowing

Two charts sit next to the burndown and answer questions it cannot.

Cumulative flow diagram (CFD). A stacked area chart where each band is a workflow state — backlog, in progress, in review, done — plotted over time. The vertical thickness of a band at any point is the number of items sitting in that state; the horizontal distance between the top and bottom of the stack is approximate cycle time. A CFD makes bottlenecks impossible to miss: if the “in review” band steadily thickens while “done” grows slowly, code review is the constraint, and no amount of pulling more work into “in progress” will help. A burndown would show only a disappointing line and no cause. CFDs pair naturally with Kanban and with any team practicing serious Code Review and Static Analysis discipline, since review queues are the most common invisible bottleneck.

Monte Carlo forecast distribution. Rather than dividing remaining scope by a velocity band, simulate: draw a velocity at random from the team’s history, subtract it from remaining scope, repeat until scope is exhausted, record the sprint count, and run the whole thing ten thousand times. The output is a histogram of completion dates, from which percentiles read directly — “50% chance by sprint 9, 85% by sprint 11, 95% by sprint 12.”

This is better than a range for two reasons. It handles a skewed velocity history correctly, where a min-max band silently overweights a single bad sprint. And it produces a confidence level, which is the right currency for a date conversation: a stakeholder who hears “85% confident by week 22” understands they are being told about risk, whereas “week 22” is heard as a promise. The simulation is a few dozen lines of code or a spreadsheet, and it consumes exactly the same input data the team already has.

Both charts share a property worth noticing: they add information without adding process. Neither requires the team to estimate anything new, attend a new meeting, or update a new field. That is the bar any additional metric should clear.

Comparison

ConceptWhat it measuresTime frameGameable?Primary use
VelocityEstimated points completed per iterationOne sprint, tracked across manyHighly — team owns the estimatesForecasting how much fits in future sprints
Burndown chartRemaining work over timeWithin a single sprintYes, via premature “done”Mid-sprint course correction
Burnup chartCompleted work and total scope as separate linesSprint or full releaseHarder — scope changes are visibleRelease tracking with changing scope
Cycle timeElapsed days from start to done, per itemContinuous, per itemHard — measured by timestampsPredicting individual item delivery
ThroughputItems completed per unit timeContinuousSomewhat, via item splittingForecasting without estimation
CapacityAvailable person-hours in the iterationOne sprint, forward-lookingNot reallyAdjusting commitment for leave and holidays

Real-World Use Cases

  • Release date forecasting for a contractual commitment. A team with 248 points remaining and a 27–34 velocity band reports “16 to 20 weeks” to a customer rather than a single date, and the contract is negotiated against the pessimistic end.
  • Mid-sprint intervention. A Scrum team notices the burndown flat from day 4 to day 6, discovers at standup that all four engineers are blocked on one unmerged refactor, and pairs down to unblock it — recovering three days that would otherwise have been lost to a cliff finish.
  • Making an interrupt cost visible. A support escalation adds 13 points mid-sprint. The burnup’s scope line steps up, and the product owner can see the two stories that will now not ship, so the trade is made deliberately rather than discovered at review.
  • Detecting an onboarding ramp. Velocity drops from 32 to 21 for two sprints after two new hires join. The chart makes the expected ramp legible so it is not misdiagnosed as a delivery problem, and it recovers to 38 by sprint four.
  • Retrospective evidence. A Sprint Retrospective opens with three sprints of burndowns side by side, all showing the same day-8 cliff — turning a vague sense that “sprints feel rushed” into a specific, fixable integration-timing problem.
  • Calibrating an over-optimistic estimation baseline. A team commits to 45 points for six consecutive sprints and delivers 28, 31, 26, 30, 29, 27. The data ends the argument: the commitment is wrong, not the team.
  • Sizing a quarterly roadmap. A Product Roadmap proposal totals roughly 400 points against a band of 27–34, implying 12–15 sprints against a 6-sprint quarter. The arithmetic forces prioritization before the quarter starts rather than triage in week ten.
  • Justifying technical debt work. A team shows velocity declining from 38 to 24 over eight sprints while headcount is flat, correlated with rising defect counts, and uses it to argue for a Code Refactoring and Technical Debt investment against a specific module.
  • Capacity-adjusted sprint planning. With three of eight engineers out for a holiday week, the team scales its commitment from the usual 30 to 19 and hits it — instead of committing to 30 and carrying 11 points into the next sprint.
  • Detecting a slicing problem. A burndown with only three enormous steps reveals that the sprint held three 13-point stories rather than eight smaller ones, prompting a change in how Product Backlog and Refinement slices work.

Common Pitfalls

  • Treating velocity as a productivity score. The instant it appears in a performance review or a cross-team dashboard, teams inflate estimates and the number decouples from reality. It is a planning input owned by the team, not a scorecard.
  • Comparing velocity between teams. The unit is arbitrary and team-local. Team A’s 40 and Team B’s 25 are not comparable in any direction, and building a portfolio metric from them produces confident nonsense.
  • Giving partial credit. Counting a 13-point story as “8 points done” makes the burndown descend smoothly and mean nothing. The story either met the Definition of Done or it did not.
  • Using a single average to forecast. A mean converts noisy history into one date, and a date reads as a promise. Report the band. If sprints ranged 22 to 36, the forecast has a range too.
  • Ignoring scope change in a burndown. A flat sprint burndown can mean nobody finished anything or that completion exactly matched additions. Use a burnup for anything longer than one sprint, and mark scope changes on sprint burndowns explicitly.
  • Chasing the ideal line. The straight line is a reference for comparison, not a target to hit. Teams that manage the chart to match it start closing items early, and the chart becomes a fiction maintained for an audience.
  • Re-estimating completed work. Recording actual effort instead of the original estimate turns velocity into an effort log and destroys its use as an estimation calibration signal. The original number stays.
  • Forgetting that velocity resets. Team changes, sprint-length changes, a stricter Definition of Done, or a rebaselined estimation scale all invalidate historical velocity. Continuing to forecast off pre-change data produces confidently wrong dates.
  • Forecasting from an unstable baseline. The first three to six sprints of a new team measure calibration drift, not throughput. Do not build a release commitment on them.
  • Counting unfinished carryover twice. If a 5-point story slips to the next sprint, it counts in the sprint where it finishes and nowhere else. Splitting the credit across both sprints inflates total velocity and corrupts every forecast built on it.
  • Letting the chart substitute for conversation. A burndown identifies where to look; it never explains why. A team that reads the chart at standup instead of talking to each other has automated the symptom and kept the disease.

Example

A payments team at a mid-size fintech had eleven sprints of history and a release commitment they were quietly worried about. Their board showed the last eight velocities: 29, 34, 22, 31, 36, 27, 33, 30. Leadership had been told “about 30 points a sprint,” and the release backlog stood at 248 points, so the number circulating in planning meetings was eight sprints — sixteen weeks, comfortably inside the quarter-plus-one they had promised a partner bank.

The team’s tech lead rebuilt the forecast as a band. Trimming the 22 and the 36, the working range was 27 to 34, giving 8 to 10 sprints, or 16 to 20 weeks. That alone changed the conversation, because the promised date sat exactly on the optimistic edge. Then she plotted the release backlog’s total size sprint by sprint as a burnup and found the second, worse problem: the scope line had risen steadily, averaging about six points per sprint of newly discovered work — integration edge cases from the partner’s API that nobody could have specified up front. Against a moving target, effective velocity was closer to 21–28, pushing the forecast to 9 to 12 sprints. The optimistic case now barely met the commitment and the pessimistic case missed it by six weeks.

The sprint burndowns explained the rest. Five of the last six showed the same shape: flat through day 6, then a near-vertical drop on days 9 and 10. Everything was integrating at the end. And the carryover pattern confirmed the cost — each sprint opened with defects from the previous sprint’s last-day merges, consuming roughly four points of unplanned capacity that never appeared in any commitment. The team made three changes. They capped work in progress at four items so stories finished mid-sprint instead of piling up at the boundary. They planned to 27, the low end of the band, rather than 30. And they took the burnup to the partner with an explicit 9-to-12-sprint range and a named list of the scope that had been added since the original estimate. The partner renegotiated the milestone, splitting it in two. Three sprints later the burndowns had lost their cliff — the shape was a rough staircase descending across the whole sprint — and velocity had settled at 28, 31, 29. Lower than the 30 that had been quoted for months, and for the first time, a number anyone could plan against.

Dig deeper