Help me think through how this episode applies to my situation. Start by asking what I am trying to change. Separate the episode's claims from your suggestions, and say when the notes do not support a claim. Use the transcript to find passages, then check the audio before quoting.
Transcript: https://yuvalyeret.com/podcast/episodes/tiny-teams-same-dependencies-can-ai-flatten-your-organization/transcript.md
## Published episode notes
Consulting firms are telling engineering leaders that coding agents mean you can flatten the org, cut coordination layers, and move to autonomous tiny teams. But in the enterprise trenches, most teams aren't decoupled feature teams, they are system teams with heavy cross-team dependencies. 10x-ing code generation without fixing organizational coupling doesn't eliminate scaling overhead; it just floods downstream queues with unintegrated PRs and dependency waits.
In this solo episode, Yuval Yeret breaks down what scaling frameworks (SAFe, LeSS) actually do in the age of AI, why procedural coordination collapses while structural complexity remains, and how to use coding agents at your system bottlenecks to continuously descale without breaking delivery.
Notable Quotes:
"If your organization requires 4 portfolios, 8 value chains, and 16 pods to align before a single customer epic can ship, 10x-ing code generation just produces more unintegrated PRs and longer dependency queues. That isn't throughput, it's activity theater."
"Scaling frameworks exist for one reason: to manage the coordination overhead forced on you by your current dependencies. As long as those dependencies exist, you need to manage them somewhere."
"Don't just create smaller teams, call them tiny, and assume they will 10x without fixing the architecture and the system around them."
Links & Resources:
- Read the companion article: https://yuvalyeret.com/blog/most-of-your-scaling-apparatus-is-now-optional
- Organizing teams around outcomes: https://yuvalyeret.com/blog/when-and-why-do-we-need-a-product-operating-model
- Kevin Fox's Blue Light (Theory of Constraints): https://theoryofconstraints.blogspot.com/2007/06/toc-stories-2-blue-light-creating.html
Scaling AI: From Activity to Impact , For leaders trying to turn AI activity into real business impact.
Insights (and help/advice) on scaling AI from activity to impact , yuvalyeret.com/insights
Yuval's Linkedin – https://www.linkedin.com/in/yuvalyeret/
## Transcript
Riverside's published transcript for this episode. Speaker labels are machine-generated; check the audio before quoting.
Episode: https://yuvalyeret.com/scaling-ai-podcast/tiny-teams-same-dependencies-can-ai-flatten-your-organization/
Source RSS GUID: b8fdb16e-d611-4e51-a819-9ebb5effa35f
Source: Riverside podcast:transcript URL
## Transcript
Yuval Yeret: Does this sound familiar? Your engineers or teams who are adopting an Agentique software development lifecycle see work being compressed. Tasks that took hours now take AI agents minutes or even seconds sometimes. They aren't even worth managing. Stories that took days now at most take a few hours, but often also flow in minutes and are created and managed by AI. These engineers They complain that the current ways of working, whether they are a form of Scrum, Safe, Team Level Kanban, whatever approach you use to manage the work, on Jira, Azure DevOps, Linear, all of these feel wasteful and irrelevant. It sometimes feels like they were just looking for the perfect excuse to get rid of the process theater overhead, doesn't it? Your research into best practices in organizing an AI-native product engineering organization shows similar advice from all the usual sources, McKinsey, Deloitte, Gartner, all of them boiling down to the notion of organizing into tiny teams with minimal or non-existent coordination layers, which mean flatter, leaner organization, which always sounds attractive. With a lot of the human costs, especially in the middle layers, and a lot of the process overhead replaced by token spend or made redundant by token spend. And even putting aside the conversation around are tokens still the free capacity that we thought of them as, there's a pertinent question here that you're asking yourself. Does all of this mean that there's no need to manage the work? The way we've done it before? All of the ways we've scaled our ways of working, the edge-all approaches, sprints, reviews, planning, story points, PI planning, release planning, are all of these redundant? Can we just staff these tiny teams, give them tokens, decide on a spec ribbon harness, and that's it? Can it be this simple? Hi, I'm Yuval Garrett, and this is Scaling AI from Activity to Impact. So let's take a step back. What's the right product team size? Let's go back to the history, what's the guidance, why did we land on the team size that we've had until AI and what changes with AI? So product team size is always an optimization or trade-off. On one hand, we know that the smaller the team, the easier it is to collaborate within the team and to get to a high-performing state. On the other hand, a small team that's always blocked is useless. We want the team to be empowered, to be able to close the loop, to deliver values, sense and respond with minimal or even no dependencies ideally outside the team. And even if we never get to that ideal North Star. The point is to minimize the number of dependencies because any dependency has the potential of locking or slowing down the team or adding a lot of coordination work that is wasteful. So it might be something that we have to do, it might be a necessary evil, but we are trying to optimize around it. We often refer to this as being cross-functional, the team having access to all of the functions or disciplines that are needed. So identify, define, deliver, test and realize value. When we when we look out there to most of the organizations, if we put aside AI native startups, okay, most product teams, at least I see out there in the trenches, are yes, up to around 10 people, but they're not cross-functional enough. Therefore, they often depend on other teams to deliver something valuable. And and that makes sense. The typical organization has so much complexity and scope in their products, in their portfolio of products, that it's almost unrealistic, you might say, to staff all of the expertise needed to do everything across this portfolio into one team. So you end up having to optimize for something. And the reality is that in some of the cases, the organization didn't even work very hard to optimize. They found a shortcut in the system to product team transformation. They just called their systems products and voila, they have product teams. But these teams fail the empower test. They cannot really deliver a real feature end-to-end. They might say that they can deliver a feature, but those features are just a bunch of stories. We used to ask, I used to ask a team, an organization, show me what you can deliver on your own without the dependencies. Can you deliver features? And that's a useful sniff test for whether it's an effective and powered team. But then the shortcut gang found a way around it. They define features as a bunch of work that the team can do, not as something valuable and marketable. And then the value of the term, the pattern feature team, went down the drain, and it requires a lot of validation that it's a real feature team, not just a feature theater team or a theater feature team, whatever you want to call it. So, if we assume that a lot of our teams are somewhere on this continuum, but not really fully empowered product teams, there will be dependencies if we're trying to deliver meaningful, useful value. And that's where the need for scaling patterns comes in. If the teams can't fully deliver value on their own, We need some mechanism to coordinate and manage the collaboration across these teams to deliver value. And that's the purpose of scaling frameworks. Whether it's safe, less, Nexus, even the Spotify model that isn't a framework, they all at their core try to answer the question: what's the minimal amount of management or coordination scaffolding or overhead that we can get away with to ensure predictable, effective delivery of value across a set of teams? To create essentially alignment and coordination across a set of teams. Each of these frameworks makes a different assumption on the state of the organization. Less annexes, for example, assume product teams that are much closer to the ideal feature team, so that they can get away with less coordination and scaling overhead. Safe also prefers such an organization with teams that are organized around value. But if we're If we're looking at which organizations choose to work with SAFE, SAFE is a solution for main market organizations, mainstream organizations that struggle to make this transition from systems and project teams all the way to empowered product teams. Which means that SAFE does need to include more elements to manage the coordination overhead than the lighter weight frameworks. Okay. So we established why we had scaling frameworks in the first place. Let's talk about what happens when product teams start to use AI to build software. AI accelerates the software development cycle. Engineers that use coding agents effectively can be much more productive. They tackle stories that used to take several days in a few hours or even minutes, but beyond increasing the story output, the real effect. Depends on engineering team versatility and the code-based scale and health. Imagine this. If your organization requires four portfolios, eight value streams, 16 teams or pods to align and build something and then integrate before a single customer initiative epic feature can ship, 10x in code generation, even across all of these teams. Just produces more unintegrated PRs and longer dependency queues. 10Xing a few of these teams might not necessarily have any effect. It really depends which of these teams and are they the Balmic. You do not get faster time to market an amplified throughput if you just accelerate development inside each one of these teams. You just get a new version of Activity Theater. So let's go back to the concept, the silver bullet concept of tiny teams. What needs to be true for a tiny team topology to work? Or in other terms, what does it look like to be ready for AI amplification and acceleration, real acceleration in software development? What does AI readiness in software development look like? So tiny teams work well when Engineers can spend wide across their code base. They aren't afraid to go anywhere. They're allowed and even encouraged to go anywhere. And they're capable of providing judgment, perspective to AI agents that are working anywhere in the code base that they need to touch to deliver value. And again, a lot of the time, the value, especially when you have a portfolio of products, is in the connection between these different products. It's to leverage the fact that you brought these products together into a portfolio that you get some advantage compared to a best of read product in a certain vertical. You have to do work that cuts across in order to be competitive. You cannot just say, no, no, I'll do all of the work in a certain The the next thing that enables tiny teams to work well is when the coding agents can see and work across the code base, not just in a specific repo. I'm working with several coaching clients that were where that's not a trivial conclusion where they, as developers or even leaders of teams, they're focused on their own repo. They don't have access and for sure they don't have a strong context. For all of the code repositories of the teams that they depend on. Some of the time it's just early days in figuring this out, some of the times it's security, some time it's trusting between groups in the organization. Some of the times it's just the technical challenge of how do we give how do we get coding agents to be able to ingest the huge context from A portfolio of repositories. But when the coding agent can see across the code base of the product portfolio across multiple products, when they have access to all of these repositories, to the relevant test environments, and to the knowledge about the context the product and its adjacent products is living in, then they can do more meaningful work. The the third criteria is when humans and their agents can own the entire value stream, from intent or what's the problem or opportunity, all the way to value realization, not just building, not just even integrating or even testing or deploying. It's about stewarding the product, the software, through adoption, retention, value creation. So if that's the if that's what we want to see for tiny teams to work, then you can already imagine, I guess, that when we try to implement a tiny team in the typical organization that now raises the let's go AI native flag, the the reality is that they don't have Teams that can do that. Even their 10-person teams cannot do that right now. Their coding agents are often more of a single player or single-team coding agent that looks at a small set of repositories, it cannot look across all of them. The the typical organization didn't make enough headway into a real agile product operating model, so their teams and architecture don't provide the conditions right now for tiny team success. In these environments, you might be able to reduce team size, but until you deal with the same limitations of the pre-AI product teams you had, which aren't really empowered product teams, you would still need the same level of coordination across these teams. So don't throw your scaling patterns yet. I I still believe, and and I've been telling teams and practitioners and leaders for More than a decade now. Descale when you can, scale when you must. And and that pattern applies in this context as well. Do you still need scaling patterns? In this case, you probably do need quite a bit of your current scaling patterns, unfortunately. As long as your teams need to collaborate to deliver value, you need some mechanism to coordinate and manage priorities. Execution. Delivery all across these teams. If, for example, a true feature still needs the involvement of multiple teams, you still need some mechanism for planning and coordinating the work, the the work breakdown of that feature. And you might even still need to manage user stories, believe it or not. Not because the agents cannot work through these stories, but if those stories are the way teams Interact with each other and tell each other this is something I need from you, then you might still need to manage these interactions and manage these dependencies. Now, maybe at some point my coding agent will interact around stories with your coding agent, and that's fine, but we still need that level of. What your scaling apparatus looks like at this point really depends on what you need. Which is good advice I give to all my clients and students. Again, descale when you can, scale when you must. Try to avoid solution trains, agile release trains. Try to even avoid teams. If one person can do everything, great, they can work as an individual. That's typically not realistic. But regardless of how you structure right now, what you want is to think towards the future, to establish an AI readiness runway. To start to think about, okay, we have to work this way right now. But what needs to be true for us to be able to descale? Or at least enable a lighter scaling apparatus? The same that was true before AI. We need to work on team independence and decoupling. The new interesting thing with AI is that AI is actually helpful for some of these moves. So let's think through this. First of all, collective ownership. The notion that everyone can code anywhere. The typical barriers to collective code ownership or versatility are the lack of safety nets, the amount of technical debt we have. The level of complexity. There's the notion of here be dragons, areas where engineers are afraid to get anywhere close to. Joshua Karievsky talks about technical safety. The fact that people don't feel safe to go into a certain area of the code. And when they don't, they won't. They won't go there. And the owners of that code don't feel safe to let. Others touch the code that they're responsible for because they know they're gonna be on the hook for fixing it. Coding agents can be leveraged to refactor, reduce technical debt, and create safety through test automation and code review workflows. That essentially go into some of these dragon caves and clear them up. Coding agents with their ability to get up to speed on new code faster than any human and support, be a copilot for a human that has to go into an uncharted territory from their perspective can be a great tool to help them work in these foreign areas more productively, more safely. The next challenge or the next approach that we might have to create to create empowered product teams that are smaller is to create an architecture where a lot of the dependencies are replaced with consumable self-service platforms. So that much more of the functionality that I might need to deliver to develop my feature. I can self-serve myself using APIs or other techniques that don't require new development in the platform. So I won't have something that I need to ask another team to do for me before I can continue. I can do more of the work myself, or the APIs already support me in what I need. And this isn't anything new, to be honest. Many of the organizations I work with, most of them, have a clear vision for what's a decoupled architecture that will set them free for dependencies. But they simply never had the capacity to invest properly in the architectural runway. Even when they're doing a cloud transformation, multi-year project, the pressure is often so high to finish the transformation that they go into a lift and shift mode. Rather than a proper re-architecture to microservices or something along those lines. The speed and throughput of coding agents can be leveraged to pursue this architectural runway. This sort of work is actually great for coding agents because they can work in a closed loop. The customer is an internal system. It's easy to observe and to define the boundaries. It's much easier to do this sort of work. then establish something that is desirable from the perspective of real world customers. Coding agents can even take on gauntlet goals where they re-architect entire systems assuming they truly have access to the right context and environment. Alright, so there are a couple of different things you could do to pursue AI readiness, but the reality is Rome wasn't built in a day. An AI ambitious organization won't become AI native or even AI ready in a day either, even in a quarter. Even with coding agents, you have limited capacity and limited throughput you can dedicate to working on this runway. So, surprise, surprise, you might still need to prioritize. Let's talk a bit about how do you prioritize your AI readiness from. One of the techniques I like to use is to focus on the constraint. Instead of throwing away your scaling apparatus, leverage it to find the constraints that you have in the system. The constraints in the form of dependencies, for example. Here are some of the ways I like to do that. So, whether it is for a dependencies board retrospective or mining the information in your Jira or AVO, you want to review your dependencies. And here again is one area where I can really help. In analysis of a lot of data and finding correlations and dependencies in Look at why the work is getting blocked. Dependencies are often the reason, whether explicitly or implicitly. And again, AI can analyze your data. Even if you don't have a formal block reason for all of these situations, it can dive deeper into comments and find some of these patterns. You can also look at flow metrics. Items with the highest cycle or flow time or lowest flow efficiency. Typically hide some dependency on a scarce and unpredictable resource. This will help you figure out the constraint, the bottleneck. It's often the team or system that everyone depends on, that is often unpredictable, whose capacity a lot of time is spent in rework, maintenance, serving demand, where they can barely get time to be proactive to go on the offense. And to be honest, more often than not, people already know to point to this constraint without the data. The next step is to exploit the constraint. What does that mean? Once you identify your constraint, it's time to move 20% of your engineering capacity to that team. Well, I'm just kidding. A more practical and economical solution these days is to focus your coding agent adoption in this area. This is where even local acceleration can actually shift the performance of the whole system. Because if you're accelerating the bottleneck, you might be affecting the, you are probably affecting the whole system. Those coding agents can deal with the ongoing simple demand, refactoring, defect fixing, automation. So it can free some human capacity to think and drive ways to improve capacity further. It's in work the architectural one way that this team really has to make progress on. Items that are focused on dependency elimination, whether it's The creation of platforms, self-service APIs, enabling other teams to self-serve through knowledge sharing, and the creation of context aware agents, digital twins, whatever you want to call them, that are specialized for this system. So that they can either advise or join other teams that are doing work that that integrates with this area, or even serve them with much smoother service levels. The the in general, what we want to do is to optimize for these bottlenecks, these specialized, scarce knowledge, scarce capacity people and teams. Kevin Fox tells a classic story in the world of system thinking and theory of constraints that's called the the blue light. The essence of the story is that if you realized that the welder is the constraint of your Value stream of your production line, you want to see the blue light of a weld in action coming out of the welding booth as much of the time as possible. And of course, there are a few assumptions about utilization, efficiency, effectiveness here, but if we assume that the welder is really working only on the prioritized things, if there's no light, the welder isn't welding and the system is blocked. If we apply it to our weld, You want to make sure that these specialized teams can be as effective as possible. Every time they waste effort because someone asks an incomplete question, brings in work with missing context, or they get stuck on their own dependencies, it's a big deal. It's a bigger deal than teams that aren't the ball neck that are doing some rework. So you want to focus on that. And at this point, you want to think about what you can do to keep the blue light on. A specific question in the age of AI is how can I help how can AI help us here? Even for the cases where you want a specialized knowledge the humans are bringing in the loop, AI can help facilitate efficient intake. It can help as a copilot for these humans. Hopefully, there's some things that the AI can do on its own and reduce, you know, route around these people altogether, reducing some of the offloading some of the capacity. So let's go back to the scaling apparatus. What are the scaling patterns that make sense? Those that help you manage and optimize the flow through your system, the flow through your constraints. This isn't that different from the perspective I started the article with. The scaling mechanism should be the minimum viable coordination overhead that is needed and sufficient. To tackle the collaboration and coordination across people and teams required because of your current dependencies. Should you use an agile release frame? PI planning? Are you ready to shift from full story-based PI planning to planning around features and outcomes, maybe? It's something along the lines of what the AI Native Safe is introducing with outcome planning. The question shouldn't be. Which scaling framework was cooler, newer, or has a stronger ecosystem around it at this point? It should be which has the patterns that are a good fit for where you are right now, and has the vision and runway to both nudge you as well as support you on the way to working the way you want to in the age of AI. Arts are great, portfolios are great, solution trains can be a great solution for a certain context. But then again, try to avoid them if at all possible. Because collaboration is much easier the smaller the team or the group or the organization you have. Tiny teams are a great idea. And if you're ready for them and for reimagining how you do the work and the full on the scaling, I say go for it. If you're not currently ready, but believe you can be and need a push, one thing you could do is say, okay, let's go for tiny teams, burn the ships. That's the burn the ships approach to change. And let's deal with all of the dependencies, all of the problems, and only bring back the scaling patterns that we really need. That can work as well. So there are Different things that you can do here just going into this process, thinking about what's the minimal scaling patterns that you need. Either you know, bring them up from zero or take, you know, take out the patterns that you're currently using and go to lighter weight that you currently have. Both of these approaches can work, and maybe you can try to do a thinking exercise of coming at it from two different perspectives and see where you land. The the way I approach it with my clients is often looks something like this. And this is something you could take and do on your own. So take your last couple of customer visible outcomes that are meaningful to the business, to product, to the organization, and count how many teams did each one have to cross. Not how many teams touched the code, how many teams had to agree. on something wait for each other before it shipped. AI can help you mine Jira or ADO for what's really going on here. So it can save some time these days. If the answer is one or two, then a lot of your apparatus probably is optional. You're already in a good spot. If the answer is five, then agents could make each one of these five teams faster, but they will likely still depend on each other to deliver anything useful unless you make a different change. And what you can do then is run a thinking or simulation exercise to consider some what-if scenarios. What if we implemented tiny teams along the existing structure or topology? Would we still see dependencies? What if we invested in creating some platforms? What if we created digital twins? What if we did this? What if we did that? And then you can see which of these potential interventions have the biggest impact. On simplifying your structure and what's the cost of and the runway towards achieving them. AI can also help you with running these scenarios, of course, but only if you give them enough context about what's really going on. What are the different products? How does the historical work look like? What do you have in your backlog? What are the typical dependencies? All of that context. Can help AI help you figure out what are your options and play these white if scenarios. At the end of this process, you have a couple of different alternatives, a couple of different potential experiments to run or things you're sure you want to do, and then you can create an improvement roadmap and start to execute on it on the path to really high readiness. Hopefully, AI can be used to help you achieve AI readiness as well. The only gotcha here is don't just create smaller teams, call them tiny, and assume they will 10x without taking care of how they're organized, the architecture they're working on, and the system around them. That is just tiny team theater. I hope this was useful. I'm Yval. Join me again for more insights on scaling AI from activity to imp.
## Source boundary
These are the published show notes from the podcast feed. They are a starting point for discussion, not a verbatim record of the conversation. The transcript is machine-generated and may contain errors or unlabeled speakers. Check the audio before quoting anyone.