How to Customize Your SAFe for the AI Age
Most of your scaling apparatus is now optional. Your first principles are not. A companion for leaders running SAFe while adopting AI-native, spec-driven development, and trying to tell the difference.
Click image to open full size Scaled Agile’s answer to the AI era adds surfaces to the framework; this is a companion for leaders working out which surfaces they can now take away.
Is a scaling framework still the right answer when AI is descaling the work?
If you lead product or engineering in a SAFe organization, you are probably living inside a contradiction. AI is making your teams faster than your operating model knows how to absorb, and you suspect a meaningful share of your coordination machinery is now ceremony. But you came up through agile at scale. You have seen what happens when an organization trades discipline for “move fast,” and you will not abandon flow and alignment to chase a demo. Evidence-based steering is not up for trade either.
My answer is that most of your scaling apparatus is now optional, your first principles are not, and telling the two apart is the entire job. The mechanics of a scaling framework exist because organizations were complex. AI dissolves one kind of that complexity and leaves the other untouched. Good SAFe implementations always said to use the simplest version that solves a real coordination problem and keep simplifying; what has changed is how small “simplest” can be, and how fast that window is moving.
Scaled Agile’s own answer arrived on 23 June 2026 as AI-Native SAFe. Scaled Agile describes it as a companion to the Core SAFe Operating Model in one breath and as a reinvention in the next, and the honest reading is that nobody knows yet which it will turn out to be. It identifies a dedicated “AI value architect” role, though the release is Early Access and framework guidance for that role has not yet been published. It also raises three surfaces to framework level: Responsible AI; AI-Native Intent, Specifications, and Context; and Curated Data. It adds apparatus where AI itself creates new structural complexity, which is the right move, and it leaves the subtraction question, the one this post is about, for you to answer inside your own implementation.
Diagnose before you dismantle
The descaling pattern repeats at every altitude in the framework. A person backed by agents takes on what used to require a team, a team takes on what used to require a train, and a train takes on what used to require a Solution Train. Teams get smaller and broader because one person now spans more disciplines, and spec-driven development turns story-and-task overhead into a conversation between a person and their agents. That overhead relocates upward to the spec and the feature, lower in volume and higher in value.
So the interesting question is what happened to the complexity that made the apparatus necessary. What descaling dissolves is procedural: the coordination, handoffs, and decomposition an organization invented to move work through itself. What it leaves untouched is structural: legacy debt, acquisition sprawl, regulatory load, security and data-governance gates, live customers who cannot be disrupted. Conway’s law was not repealed by a model release.
Step one is therefore a diagnosis: for each piece of apparatus, name the coordination problem it solves. If you cannot name it, or if the problem was procedural and AI dissolved it, dismantle it. Structural problems keep their apparatus, which should concentrate where the irreducible coordination lives rather than spreading out of habit.
Before you touch any single event, work the coarsest dial you have. Essential is the base, and the two layers above it, Large Solution and Portfolio, are independent add-ons rather than rungs on a ladder. That matters for descaling, because it means they come off separately. Drop the Large Solution layer once one descaled ART can deliver what previously needed a Solution Train. Drop the portfolio layer once the remaining value streams can be funded and steered without a separate portfolio apparatus. Take both off and you are at Essential; take one off and you are at Portfolio or Large Solution depending on which one you kept. Start here, because removing a layer retires whole clusters of roles and events at once and saves you from litigating each of them separately.
Teams and the train
The Agile Team gets smaller and broader, a shift the Scrum companion works through at team level. A team of four backed by agents can plausibly own what a team of nine owned two years ago, across a wider slice of the stack. That concentrates key-person risk and review capacity in fewer heads, and it leaves less slack to absorb a wrong direction.
The Agile Release Train is where the pressure lands hardest. Its cadence-and-synchronization machinery is largely a response to cross-team dependencies, even though alignment and flow are purposes in their own right; when the dependencies thin, the machinery thins with them while the alignment work survives. If AI has genuinely made your teams more independent, some of your ARTs are now larger than the coordination problem they solve. The right move is then fewer and smaller trains, or teams that no longer belong on a train at all. Apply a hard test here, because a feeling of autonomy proves nothing: can a team take a change from idea to production, in front of real customers, without waiting on anyone outside it?
A payments organization I worked with ran an eight-team train and still spent two full days of PI Planning resolving dependencies. Three of those teams had spent two quarters absorbing the front-end and data work they used to hand off. They split the train, kept five teams on one ART around the shared ledger, let the other three run without one, and cut planning to a half-day outcomes session. Feature lead time on the independent three dropped by about a third in a quarter. The smaller ledger ART got faster too, because its planning conversation was finally about one coherent problem.
The roles
Product Management is the clearest case of freed capacity. Much of the classic legwork, writing feature descriptions and keeping the backlog clean, is now substantially assisted by AI. Be precise about the altitude: Features are Product Management’s unit, while Capabilities are a Large Solution construct owned by Solution Management, and both roles get the same relief from the same mechanism. The relief buys them the work they rarely had time for: the customer problem and whether the last bet paid off.
The Product Owner faces the sharpest question in the framework. Where teams are smaller and specs replace stories, the traditional workload of story writing and acceptance largely evaporates. The role then either moves up toward genuine outcome ownership for a slice of the product, or it becomes hard to justify. I would put that question on the table now, with the people currently in the role, while there is still room to design the answer with them.
The Release Train Engineer runs less coordination machinery, because there is less of it to run. The job that grows in its place is testing whether the descaling is real: which dependencies actually dissolved, and which merely went undocumented. Most RTEs were hired for facilitation skill, and this work asks for analysis and organizational design, so expect to invest in the transition.
The System Architect gets substantially more important, and this is the role most organizations will under-invest in. When agents generate code at volume, architectural intent is what keeps the output coherent, so someone has to define the boundaries, patterns, and guardrails agents work within, and encode them where agents can read them. Architectural review then moves out of the milestone calendar and into the generation loop itself, which changes the operating rhythm as much as the role scope.
The events
PI Planning is worth exactly as much as the dependency and alignment problem it resolves. If AI has thinned your dependency web, a two-day event whose main output was an ART Planning Board strung with red yarn is solving a problem you no longer have. The half that survives is the one most organizations undervalued: people in a room arguing about strategy and leaving aligned on outcomes. That half is the alignment work, and it belongs on a shorter horizon. A quarter is a long time to hold a plan steady when the tooling under it turns over two or three times inside that window.
Iterations still bound risk, and the scarce resource to plan around inside one is validation capacity, since generation capacity has stopped being the constraint. The System Demo earns more of its slot than before, because agent errors are subtler and more confidently wrong than the errors junior developers make. An integrated demo in front of people who can judge it is where those errors get caught.
Inspect and Adapt becomes the primary vehicle for the descaling work itself, as a standing examination of which coordination problems remain and where the constraint has moved. Put the operating-model change on a real backlog with owners and dates, and run a few structural experiments per PI. ART Sync deserves precision, because SAFe treats it as two halves that are typically combined in practice. Coach Sync works delivery and impediments across teams, and PO Sync exists to gain visibility into the ART’s progress toward its PI Objectives and make the adjustments that visibility calls for. Both halves exist largely to manage procedural complexity, so where that complexity has dissolved, drop the half that no longer produces a decision you can name and keep the half that does.
Backlogs, prioritization, and quality
Backlog altitude rises, because stories stop being the unit anyone above the team tracks. Decomposition is now a private matter between a person and their agents, so Features become the working currency at ART level and Capabilities at Large Solution.
WSJF becomes more consequential, because prioritization used to be partly a rationing exercise. When delivery capacity is abundant, the quality of that decision becomes the primary determinant of outcomes. The denominator is job duration, with job size standing in as the usual estimation proxy, and that proxy loses meaning as build cost collapses. Cost of delay gains weight in the same move.
On flow I will be brief, because the team flow companion covers the mechanics. Flow velocity will look excellent, and it is the number most likely to end up on a steering committee slide with nothing behind it. Bring flow efficiency and the wait-state data instead. Flow efficiency is value-added time as a percentage of total elapsed time, and paired with where the waiting actually happens it is the hardest reading on the board to fake. It is also the clearest evidence of whether AI sped up the work or only sped up the queuing. Flow predictability answers a question flow efficiency cannot, which is whether the organization can still commit to something and hit it while the tooling underneath keeps turning over. It is reported at team, ART, and portfolio altitudes, so keep the metric distinct from the instrument you use to produce it. At ART level that instrument is the ART Predictability Measure, which averages each team’s achievement score for the PI. A team’s score compares the business value Business Owners assigned to its committed PI Objectives at PI Planning against the value those Business Owners award at Inspect and Adapt. Uncommitted objectives sit outside that score by design, as a deliberate buffer, and SAFe caps them at two or three for exactly that reason. The failure mode is a train that quietly widens that buffer. The plan and the score come out of the same room, and a train that sandbags its commitments posts a healthy number while delivering less than it could have.
Built-In Quality gets stronger out of necessity, because human vigilance at review time cannot carry the volume agents generate. The move is toward tighter automated gates and a Definition of Done the pipeline enforces.
The portfolio layer
Lean Portfolio Management is the layer most likely to be under-disrupted, and that is a problem. The mapping between value streams and structure is now stale, because capacity that required a Solution Train may fit inside a single ART. More importantly, the portfolio is where the cost-takeout failure mode either gets caught or gets institutionalized. Descaling should mean the same people taking on ambition that was previously out of reach; capacity that gets quietly banked as headcount reduction is just the feature factory with a smaller payroll. If cost takeout is the only move you make, your best people will notice before your board does.
The portfolio question for each quarter has moved past how to fund the existing trains. Given what our teams can now do, what were we not attempting because it was too expensive, and is it still too expensive? The portfolio backlog and the portfolio Kanban system around it also become the most interesting flow surface you have. Once story-level work collapses into a single person’s working session with their agents, the flow questions that need managing live at the Epic level.
Where to start
Treat this as a diagnosis followed by a few experiments, and keep the word transformation out of it. Pick one team and let it operate as though its dependencies have dissolved, and measure what actually blocks it. Cut one sync and see which decision fails to get made. Hold the steering discipline while you strip the machinery. AI makes it easy to go fast in a straight line, and a scaling framework with its double-loop learning removed will happily accelerate an organization toward outcomes nobody checked.
Your first move this week
| SAFe apparatus | What it was compensating for | Still needed if… | Retire or shrink if… |
|---|---|---|---|
| PI Planning | Hundreds of people unable to agree on a quarter | Strategy is contested and cross-team bets collide | Its main output was a dependency map |
| ART Planning Board | Story-level handoffs between specialist teams | Dependencies run through shared platforms or regulated gates | Most red yarn is handoffs agents now absorb |
| ART Sync (Coach Sync plus PO Sync) | Weekly drift between delivery and content decisions | Each half makes a decision nothing else makes | One half is pure status; drop it, keep the other |
| System Demo | Integration surprises found too late | People who can judge agent output attend and can reject it | Your pipeline already shows those same people integrated behavior continuously |
| Solution Intent | Design and compliance knowledge scattered across teams and suppliers | Agents generate against it and someone maintains it as an executable constraint | It is a milestone document nobody reads and no pipeline enforces |
| RTE role | Full-time facilitation of coordination machinery | Someone must own flow measurement and operating-model change | It is calendar administration; move the person into redesigning team boundaries |
| Solution Train | One solution no single ART could assemble | Integration spans several ARTs or suppliers | One descaled ART can now deliver it end to end |
| LPM portfolio Kanban | Invisible, unfunded, uncapped Epic flow | Epics compete for the same scarce validation and architecture capacity | One value stream funds everything and Epics are already visible in the ART backlog |
Score your own train honestly on this table before you change a single event. Most leadership teams find two or three rows where they cannot name the coordination problem, and one or two rows they had written off as ceremony that turn out to be doing real work.
Frequently asked questions
Should we abandon SAFe if AI is descaling our work?
No. That framing hides the real decision, which is which coordination problems you still have. Some organizations carry heavy structural complexity: regulated releases, shared platforms, acquisition sprawl. They keep most of the apparatus and concentrate it where coordination is irreducible. Everyone else ends up with a much lighter SAFe.
How does this square with AI-Native SAFe?
Scaled Agile released AI-Native SAFe on 23 June 2026, in Early Access, with guidance landing every two weeks and the full release due at the SAFe Summit in September. It identifies an “AI value architect” role, raises three surfaces to framework level (Responsible AI; AI-Native Intent, Specifications, and Context; and Curated Data), and adds a Building an AI Organization competency inside the Leadership and Culture discipline. Scaled Agile calls it a companion to the Core SAFe Operating Model in some materials and a reinvention in others, so treat the positioning as unsettled rather than quoting whichever version suits your argument.
That release adds apparatus while this post argues for taking apparatus away, and both hold at once. AI generates real new structural complexity: a retrieval corpus that nobody owns, an evaluation suite that nobody runs before a model swap, a generated migration script that no named person signed. Structural complexity has always earned apparatus, so the additions are well aimed. The coordination machinery in your current implementation, on the other hand, was built for procedural complexity, and procedural complexity is the part AI dissolves. Read the June release as permission to move governance weight toward where the AI itself is risky, and read this post as the other half of that trade. Do only the adding and you will end up carrying both operating models at full weight.
What happens to the RTE role?
The role changes shape, and the new shape is more analytical. Facilitation load drops as the machinery thins, and stewardship of the operating model replaces it: tracking where flow is actually blocked, and pressing whether the train boundary still fits the problem.
Scaled Agile says the same thing about the Scrum Master and Team Coach, naming both roles alongside the RTE as coaching roles that must evolve. At team level the evolution runs in the same direction. There is less facilitation of events that have thinned and more attention to where human review is the constraint. The role gains a willingness to say out loud that a coordination habit the team kept is no longer buying anything.
Do we still need PI Planning?
Usually yes, in a much shorter form. Split the event into its two jobs and price them separately: dependency resolution is the half AI erodes, so if your board is mostly procedural handoffs, cut it close to zero. Shared understanding of strategy and bets gets more valuable as teams get faster.
How do we know if our dependencies are procedural or structural?
Run the idea-to-production test on one team for a few weeks and log every wait. Procedural dependencies are waits on a person to write or review something an agent can now produce, or to hand it onward, and they vanish once you stop scheduling them. Structural dependencies are waits on a shared platform, a data contract, or a compliance gate.
Ask, every quarter, whether you are organized for the complexity you have or the complexity you used to have.
Practical thinking on turning AI pilots, adoption, and portfolio work into business impact - by finding the constraint, changing the work, and proving value as you go.
Yuval Yeret helps product and tech leaders move from agile theater to evidence-informed delivery. Work with Yuval →
- 01 How to Customize Your Team Flow for the AI Age 23 min
- 02 How to Customize Your Scrum for the AI Age 15 min
- 03 An Operating System for AI-Native Teams 14 min