How to Customize Your Team Flow for the AI Age
A companion to the Kanban Guide for Scrum Teams for AI-native product teams. What changes in your Definition of Workflow, WIP limits, flow metrics, and Scrum events once agents do the building, and what happens to scaling when the old dependencies start dissolving.
Click image to open full size A companion to the Kanban Guide for Scrum Teams, written for teams and leaders running an AI-native, spec-driven lifecycle. Its sibling covers the same question for Scrum itself.
Does flow discipline still matter when the machine never gets tired?
There is a version of the AI story in which flow stops being interesting. If an agent can generate a feature in an afternoon, why limit work in progress, and why hold a Definition of Workflow when the workflow is a prompt and a pull request? Flow discipline matters more now, and the flow metrics are the fastest honest test of whether an AI investment is producing value or AI Theater. AI relaxed one constraint without touching the rest of the system, so the constraint moved somewhere else. In nearly every AI-native team I work with it moves to where humans understand, review, decide, and validate. That place has finite capacity, now fed several times faster than before. A Kanban system is the instrument that makes it visible. Your throughput will go up. Watch your work item age.
Kanban is a strategy for optimizing the flow of value, and the Kanban Guide for Scrum Teams keeps a minimal set of practices that complement what a professional Scrum Team already does. An AI development lifecycle changes what those existing practices make visible, and it raises the stakes on all of them. The Definition of Workflow becomes the document that decides whether your metrics are trustworthy. A WIP limit becomes load-bearing. The same logic runs past the team boundary, where an AI lifecycle quietly dissolves some of the dependencies scaling frameworks exist to manage while making the rest measurably worse.
What AI actually does to a team’s workflow
Start with the physics. Before agents, the expensive state was building, and almost every capacity conversation with a leader was some version of hiring more developers. Agentic coding inverts that: building collapses toward zero, and the states that used to be rounding errors become the whole system. Those states are specification, review, verification, integration, release, and adoption. The new bottlenecks are bottlenecks of attention, and attention is far less visible than hands on keyboards. A developer waiting on a build is obviously waiting. A developer with eleven open agent-generated pull requests looks extremely busy, and genuinely is, while nothing in that picture is finishing. I have written elsewhere about why the bottleneck moved rather than disappeared. The flow consequence is straightforward. A team that can see its work in progress and its aging items will locate the new constraint within a Sprint or two. Without that visibility, a year of throughput celebrations can pass with nothing reaching a customer sooner.
Redraw the Definition of Workflow before you touch anything else
The Definition of Workflow is the Scrum Team members’ explicit shared understanding of what their policies are for following the Kanban practices. That framing matters, because the document captures the team’s own answer in the team’s own terms. It names what a work item is, where it starts and finishes, and the states in between. It also covers how work in progress is controlled, the policies governing each state, and a Service Level Expectation. Every element deserves a second look once agents are in the system, because every metric a leader later sees inherits its accuracy from this document.
The definition of a work item changes first. Product Backlog items move up a level in an AI lifecycle, from fine-grained stories toward features and capabilities, with specs carrying the detail stories used to carry. If you keep tracking stories while the team is pulling specs, your board is measuring a unit that no longer corresponds to anything a human decides about. Pick the altitude at which a person makes a real decision, which is usually the spec.
Where work starts is the decision that will bite you. The temptation is to start the clock when an agent begins generating, and that is exactly how teams manufacture flattering cycle times. Start it when a human commits attention to the item and finish it when the work is done by your Definition of Done and ideally in the hands of users. If your started-to-finished span excludes specification at the front and review, integration, and release at the back, you have defined a workflow that measures the one part AI already solved. That workflow will survive dashboard reviews easily.
Make review and verification first-class states. Most pre-AI boards had a single “In Review” column nobody looked at hard, because review was rarely where work sat. Now it is exactly where work sits, so split that column into awaiting human review, in review, awaiting integration, and awaiting release. You cannot manage a queue you have never made visible. In an agentic system these queues form fast and quietly.
Explicit policies now have to cover humans and agents alike: pull criteria, entry and exit conditions, what “ready for review” means for agent-generated code, and when an agent may proceed autonomously. The agreements you leave implicit are now the ones violated silently and at machine speed, which is why they need to live somewhere an agent can read them, the subject of the operating system companion.
The four practices, under an AI DLC
The Kanban Guide for Scrum Teams keeps a deliberately small set of practices, and its own names for them are precise, so I use them here. They are Visualization of the Workflow, Limiting Work in Progress, Active management of work items in progress, and Inspecting and adapting the team’s Definition of Workflow. Their definitions hold under an AI lifecycle, and their emphasis shifts hard. Two of them, visualization and limiting work in progress, get materially harder to do well.
Visualization gets harder because it now has to include work you did not personally do. Agents running overnight and pull requests open that nobody has looked at are both work in progress. The common failure mode is a board showing the team’s intentions while the real system state lives in a Git provider and scattered agent sessions. Give the queue of things waiting on a person the most prominent place on the board, because that queue is the constraint.
Limiting work in progress gets simultaneously harder and more valuable. It is harder because starting a fourth thing used to mean doing a fourth thing, and now it means typing a sentence. It is more valuable because a WIP limit is the main thing standing between a team and unbounded review debt. Each independent pod carries its own WIP limit. A pod is the smallest unit that can carry a spec from shaping through review without waiting on anyone outside it. Start at one active spec per pod, with a second slot the pod has to earn on evidence. That looks absurdly low until you notice which capacity binds. A pod can start three specs in an afternoon and cannot realistically review three in a week, so sizing WIP against generation capacity guarantees a review queue that grows every day. Open the second slot only after the pod has run at one for a few weeks with review latency and work item age both flat. The piece on calculating WIP limits for the AI age walks through deriving that number from your own data. Somebody will object that the agents can handle more, which is precisely the sentence that tells you the limit is doing its job.
Actively managing work items in progress shifts from unblocking builds to unblocking people. In an AI lifecycle the most common impediment is an item sitting in a queue waiting for a person to read something. This is also the only practice positioned to catch failure modes specific to this way of working, such as an agent that loops without converging or a spec that drifts from what the team meant. Generated work also arrives technically complete and directionally wrong.
Inspecting and adapting the Definition of Workflow moves from occasional hygiene to a standing agenda item, because the underlying capability changes every few months. A yearly revisit leaves the team running a model of its own process that expired two model releases ago.
The flow metrics are the honest test
The four flow metrics are where AI Theater becomes measurable. Defined carefully, they are much harder to game than anything else a leader is likely to be shown, and the longer treatment for the agentic lifecycle is here. Work in progress moves first. If WIP has climbed since the team adopted agents and cycle time has not improved proportionally, what the team bought was parallelism. Parallelism on its own puts nothing in front of a user any sooner.
Work item age matters most. If a team adopts only one metric, this is the one. Cycle time is a lagging indicator available only after something finishes, while age is a leading indicator for everything in flight. In an agentic system items age for human reasons: a review, a decision, a stakeholder, an integration slot. A board where items are young and steadily finishing is the clearest picture available of a healthy AI-native team, and it is the picture worth defending. When age climbs while throughput looks excellent, the team is manufacturing inventory.
Cycle time tells you whether acceleration reached the customer. It is measured between the start and end points your own Definition of Workflow declares, so two teams’ cycle times are not comparable unless their definitions agree. The useful pair is the moment a human commits to an item and the moment a user can use it. First prompt to last commit measures something narrower. If generation got ten times faster and cycle time improved by fifteen percent, that is a precise measurement telling you the constraint sits downstream in review capacity or the release process. Throughput is the metric most likely to be weaponized, because it counts items finished. If items got smaller or the definition of finished got looser, the number rises without anything improving.
These metrics still measure the human system. Dashboards that count agent actions or tokens consumed are output measurement wearing a flow costume, and they will tell you the story you were hoping to hear. Moving the conversation from activity to impact is the whole reason to instrument flow, and the fuller argument is here.
What is your Service Level Expectation worth now?
The Service Level Expectation is a forecast derived from your own cycle time history. It takes the form “eighty-five percent of items finish within nine days,” and it holds up under an AI lifecycle. It is still the most credible answer when a leader asks what the investment bought. A team that went after its review capacity and cleared it can say something like “our eighty-fifth percentile went from twenty-four days to nine.” I am inventing those two numbers to show what such a claim looks like. The shape matters because the claim is falsifiable. Anyone who asks can check the cycle time history behind it, over a stated window and against a stated Definition of Workflow. A claim like “our developers report a thirty percent productivity gain” is a survey result about how the work felt.
What changes is the shape of the distribution behind the number. Coding time used to be the bulk of cycle time, and it was reasonably well behaved. Now human review latency dominates, and it is lumpy. It depends on who is available and how many items sit ahead of it. The tail fattens, so the median tightens beautifully while the eighty-fifth percentile moves far less. That is what I expect to see in teams adopting agents without touching review capacity, and the expectation is falsifiable. A team that shows a tightening median and an equally tightening tail, without having changed anything about how review works, is evidence against it. Pulling the tail down takes deliberate work on review capacity, which is exactly why the eighty-fifth percentile is the honest number to quote.
You can test this on your own team. Measure four things over a window of at least two months before and two months after. Those are the eighty-fifth percentile cycle time, the age distribution of items currently in progress, throughput per Sprint, and the WIP the team actually held, which is usually some distance from the WIP it agreed to. Anything shorter than that window is noise, and any single one of those four can be made to look good while the other three rot. Re-baseline on a rolling window measured in weeks, and quote the tail alongside the median. Treat a breach as a trigger for a conversation about what is blocking the item.
The Scrum accountabilities get sharper, not smaller
The Product Owner is accountable for maximizing value. Flow data keeps that claim checkable when generation is cheap. Ordering the Product Backlog is now mostly about protecting the team’s capacity to specify well and to validate outcomes. A Product Owner who orders by what agents produce quickly fills the review queue with work nobody asked for. One who watches work item age keeps discovering that the constraint is their own decision latency.
The Definition of Workflow belongs to the whole Scrum Team, and it gets misattributed constantly. The Scrum Team creates the Definition of Done, the quality bar the Increment has to clear, and the Developers are required to conform to it. Where the organization already sets one as a standard, every Scrum Team follows it as a minimum. The Definition of Workflow is a different animal, a shared agreement about how the team pulls and finishes work, with the Product Owner and Scrum Master inside it rather than adjacent to it. That distinction earns its keep under an AI lifecycle. The workflow’s newest policies decide when an agent proceeds alone and when a person has to be in the loop, which sets how much review load lands on the team and how much risk reaches the customer.
Developers still carry the harder half of the new bottleneck, because reading generated code carefully is its own skill. The plausible-looking wrong answer shows up far more often than it does in a colleague’s pull request, and the volume arriving per hour exceeds what any review culture was built to absorb. Treat review capacity as a limited, visible resource and you can plan around it. Left as slack time between real work, it comes back as a queue. The Scrum Master’s share is holding that WIP limit when starting is free and keeping the Definition of Workflow current as the tooling shifts underneath it. It also means making the review queue visible to people who would rather look at the throughput chart.
Running the Scrum events on flow data
The Kanban Guide for Scrum Teams leaves the Scrum events where they are and changes what they look at. I covered the four flow metrics across the Scrum events in the pre-AI version of that argument. Size Sprint Planning against your review and validation capacity. Teams that plan against generation capacity instead fill the Sprint with items that finish generating and never finish finishing. The Daily Scrum changes most. Agents produce results several times a day, so a team that looks at its board once every twenty-four hours is sampling a fast system at a slow rate. The two questions worth asking are which items are aging and which ones are waiting on a person. The event itself stays where the Guide puts it, fifteen minutes at the same time and place every working day. The Guide has already answered the sampling problem. It says plainly that the Daily Scrum “is not the only time Developers are allowed to adjust their plan” and that they often meet throughout the day to re-plan the Sprint’s work. So keep the event daily and let the board push continuous signals into that space.
The Sprint Review puts flow data in front of the people who can act on it. That includes the uncomfortable version, the one that arrives when the team got much faster at building while the organization got no faster at releasing. That is acceleration whiplash, and it is rarely visible anywhere else. The Sprint Retrospective works the inside of the same problem. The richest material there is the human-agent operating model: where WIP limits held and where they were overridden, and which specs aged and why. Then ask which parts of the Definition of Workflow no longer describe reality.
Where the flow work and the spec work meet
Inside the Sprint the team runs a continuous inner loop: shape a capability into specs, hand them to agents, verify what comes back. Three heuristics keep that loop healthy, and all three are Kanban ideas. Finish before you start, meaning open specs and unreviewed pull requests get completed before new ones are opened, which is the hardest discipline to hold when starting costs nothing. Keep specs small, because batch size discipline now has a second justification on top of the flow one. Monolithic specs degrade agent performance, inflate context costs, and make human review slower and less reliable. Small batches are now better for the machine as well as for flow. Cap active work per pod, and find the real number empirically.
The objection that spec-driven development is big up-front design in new packaging comes up in every room, and the resemblance is genuine: a spec and a requirements document both describe intended behavior before it exists. I have argued this in why spec-driven development is not waterfall unless you use it that way. The flow version is simpler. Batch size and feedback latency separate them, and those are the two things a Kanban system exists to manage. A requirements document is a large batch with a feedback loop measured in months. A spec that covers one capability, gets executed the same day, and gets validated against real usage inside a week is a small batch with a short loop. Short loops are the mechanism empiricism runs on. Spec-driven development turns into waterfall at the moment you let the batch grow and the loop stretch, and your Kanban system shows it as rising work item age on things everyone still describes as in progress.
What changes when it is more than one team
Everything above is a team-level configuration, and most teams get real value from stopping there. The interesting question arrives a quarter later, when several teams are running this way and somebody asks whether the coordination apparatus still earns its keep. Scaling has always been dependency management, and it helps to be specific about the kinds. There are knowledge dependencies, where somebody knows something you do not, and skill dependencies, where somebody can do something you cannot. There are also work dependencies, where somebody has to change code you do not own, and decision dependencies, where somebody has to approve, fund, prioritize, or sign off. Most scaling frameworks treat these as one category and build a single apparatus to handle all of them, which was defensible when they all moved at similar speeds.
An AI development lifecycle destroys that symmetry. Knowledge and skill dependencies dissolve fast, because an agent with the right context closes much of the gap that used to require a specialist’s calendar. Work dependencies partially dissolve, though code ownership and deployment coupling are stubbornly real. Decision dependencies get worse, because more decisions now arrive per unit of time at the same set of humans. The machinery built for knowledge and skill dependencies therefore becomes overhead, while the machinery for decision dependencies becomes the binding constraint exactly as its load increases.
Your multi-team board is now a review and decision queue. Organization-level WIP is the number of specs in flight across all the teams, which is almost certainly higher than anyone has counted. Most scaled dashboards measure the predictability of plans and never look at the age of work, which is a fine way to look organized while items rot in a queue. Adopt aggregate work item age, which tells you whether the coordination layer is helping items finish or watching them wait.
Localize the collaboration instead of coordinating it
Here the AI lifecycle offers something genuinely new. The classic reason for a handoff between teams was some version of “we do not know that codebase” or “we do not have that skill,” and both are precisely what agents are good at closing. The move now is to make it cheaper for a team to do the work safely inside someone else’s area than to queue a request to that team.
The way to make it safe is encoded context. The team that owns an area publishes its architectural constraints, its test contract, its review policy, and its Definition of Done in a form agents can actually read. That might be an AGENTS.md, a spec library, or a shared skill. That converts a collaboration dependency into a context dependency, and context dependencies are asynchronous and need nobody’s calendar. Localize the generation and keep the accountability. Team A generates against Team B’s encoded rules, and Team B still reviews. That review is smaller and better bounded, and the round trip is a single queue, where a handoff used to cost a whole planning cycle. Invest in contract tests to match. This is the AI-era version of loosely coupled and tightly aligned, with alignment enforced by machine-readable policy rather than a recurring meeting.
Two failure modes deserve naming. The first is everyone changing everyone’s code with agents and nobody owning anything. The owning team’s review policy has to gate that at the merge, because an optional Team B review leaves you with entropy and good intentions. The second is subtler and costs more. Some coordination meetings were never really about work sequencing; they were the only forum where people from different teams shared context, spotted architectural problems early, or built relationships that make hard conversations possible later. Cut those and you will not notice the loss for six months, and then you will notice it all at once. Keep the forum, change its agenda.
How scaling patterns evolve
These are stages an organization moves through, and none of them is a target state. It starts as classic dependency management, with a coordination event that surfaces dependencies and a planning cadence that exists mostly to sequence work. Then teams get smaller and broader, because the same scope needs fewer people once each person’s reach expands, and that is the descaling everyone talks about. Next the agenda turns to where the review constraint sits and whether the generated work is architecturally coherent. Most organizations keep running the old agenda for a year after the questions changed. Eventually encoded context becomes the primary coordination method and whatever scheduled coordination survives turns its attention to outcomes.
Then you hit the real ceiling. Descaling works on procedural complexity, which is coordination overhead created by how we organized ourselves. It stalls on structural complexity: regulatory gates, shared platforms, legacy systems with real coupling, live customers with contractual expectations. None of that responds to better tooling. It sits in contracts and system coupling that took years to accumulate and will take years to unwind, so point whatever coordination apparatus you keep at exactly those places. Treat the stages as a diagnosis. If you remove a mechanism and the decision it used to produce stops getting made, put it back. Teams that get this wrong usually ran descaling as a cost-takeout program and counted only the roles they removed.
Your first move this week
Book forty-five minutes and rewrite your Definition of Workflow against the elements the Kanban Guide for Scrum Teams says your visualization should include. The “AI DLC version” column is where the argument lives, and the “Who decides” column is there because ownership confusion is the most common reason these agreements never hold.
| DoW element | Pre-AI default | AI DLC version | Who decides |
|---|---|---|---|
| What “started” means | A developer opens a branch on a story | A human commits attention to a spec and it leaves the pool anyone else could pull from | Scrum Team |
| What “finished” means | Merged to main and meeting the DoD | Meets the DoD, is released, and has been observed in real use at least once | Scrum Team, which owns both the DoD and where flow ends |
| Workflow items and states | Story: To Do, In Progress, In Review, Done | Spec: Shaping, Generating, Awaiting Human Review, In Review, Awaiting Integration, Awaiting Release | Scrum Team |
| WIP limits per state | Three or four per team, mostly aspirational | One active spec per pod, second slot earned by evidence, plus a hard cap on Awaiting Human Review | Scrum Team |
| How items are pulled | Whoever is free takes the top item | Nothing enters Generating while anything sits in Awaiting Human Review; review beats generation | Scrum Team, published where agents can read it |
| Service Level Expectation | 85% within N days, recomputed quarterly | 85th percentile quoted with its tail, re-baselined on a rolling six-week window | Scrum Team, reported by the PO at Sprint Review |
Run this as a working session with the whole team and the current board on screen. A document drafted alone and circulated for comment will never surface the arguments you need to have. Fill the third column out loud and argue about the cells you disagree on. Write down a first version even where it is obviously wrong, because a visible wrong policy gets corrected within a Sprint and an implicit one never does. Then put the result where an agent can read it, and set a date six weeks out to do it again. The payoff is that you can finally see where your work waits, which is where a speed problem gets solved.
Generation got cheap. Comprehension did not, and flow is how you find out which of the two your organization is actually short of.
Frequently asked questions
Should we still limit work in progress if agents are doing most of the building?
Yes, and it matters more than before. A WIP limit was always about protecting the system from starting more than it can finish, and AI made starting nearly free while leaving the cost of reviewing and validating where it was. Set it on each independent pod, and start at one active spec. The second slot gets earned once review latency and work item age have stayed flat, because review capacity is the binding constraint and it is much smaller than generation capacity.
What is the single most useful flow metric in an AI development lifecycle?
Work item age. Cycle time only tells you about work that already finished, and throughput will look good regardless. Age tells you right now which in-flight items are stuck, and in an agentic system they are almost always stuck for a human reason. It is the closest proxy available for human decision latency, the real constraint in most AI-native teams.
Do we need Kanban if we are already doing Scrum?
You need the flow practices, and the Kanban Guide for Scrum Teams is a deliberately minimal way to add them without adding a second framework. Scrum gives you the empirical loop and the accountabilities; the flow layer tells you whether that loop is producing value faster or just producing more. Without flow visibility, an AI lifecycle means running the Scrum events on opinion rather than data.
Does an AI development lifecycle mean we can stop doing scaled planning events?
Not automatically, and not all at once. Some of what those events coordinate is dissolving, particularly knowledge and skill dependencies, and some is getting worse, notably decision dependencies and architectural coherence. Change the agenda before cutting the event, and measure aggregate work item age across teams. Keep whatever is load-bearing for structural complexity.
Practical thinking on turning AI pilots, adoption, and portfolio work into business impact - by finding the constraint, changing the work, and proving value as you go.
Yuval Yeret helps product and tech leaders move from agile theater to evidence-informed delivery. Work with Yuval →
- 01 How to Customize Your Scrum for the AI Age 15 min
- 02 How to Customize Your SAFe for the AI Age 16 min
- 03 An Operating System for AI-Native Teams 14 min