Don't Redesign Your Process Yet. Change Who Writes the Artifacts.
Next Insurance broke its whole product development lifecycle into agent skills and deliberately kept sprints, roadmaps, Jira and the role split. Why authorship is the first thing to change in an agentic rollout, and how you know when the process change is finally due.
Click image to open full size The first thing to change is who writes the artifacts, not how you work
“We didn’t change the process. We still have a quarterly planning and we have sprint planning and we have roadmaps and a list of features that we want to deliver in a quarter.” That is Shay Mandel, describing an organization of about two hundred people that has spent three or four months putting an agentic development lifecycle into production. On purpose, they didn’t say let’s jump to the end and let’s change everything, all the processes and everything. What they changed instead is who writes the artifacts: in order to create good requirements there are a set of steps you need to do, and they created a skill for each one of them. The people in engineering, and in product, are not supposed to write anything on their own anymore. They just need to fix the skills or fix the context.
It’s clear that the model is not going to be perfect. But you’re trying not just to fix the model’s outputs, or even to avoid fixing the model’s outputs. You’re fixing the system instead. The inputs into the model, whether that’s the context or the skills. So essentially everybody is developing this system that is the agentic development lifecycle, which is the right mindset. That is the change I would make first, and it’s a much smaller change than the one most leadership teams reach for after a demo. The process change is real and it comes second. You can tell when it comes due: it’s when your board stops telling you where the work is.
What they changed, and what they left alone on purpose
Quarterly planning is unchanged. So is sprint planning, the roadmap, and the list of features they want to deliver in a quarter. So are Jira and Confluence, and the split between engineering, product management and insurance product.
“I think what we did is we didn’t change the process. We still have a quarterly planning and we have sprint planning and we have roadmaps and a list of features that we want to deliver in a quarter.” : Shay Mandel
They didn’t change the processes much. What they did was break them down. In Shay’s words: in order to create good requirements, there are a set of steps you need to do, and they created a skill for each one of them. Describing the problem. Then analyzing the problem, getting some data, maybe doing some user research, creating wireframes. All of these are things that you usually do. And then eventually you articulate everything in one PRD where you define the solution. There is an uber skill that helps you route. You say you want to do this PRD. It tells you which stage you’re in, and then it directs you. It’s interactive. It asks you how you measure it, what you expect the impact to be, and it interviews you until it can write the section.
Then the same treatment on the far side of the handoff, where the cost usually is:
“We know that usually we hand off to engineering and then there are a lot of reviews and questions and there is kind of a lot of feedback. We created a skill for that, so we call it PRD review by engineering.” : Shay Mandel
It goes and looks at everything and makes sure that for every domain, the requirements actually answer the questions or the details that they need. Then it gives the product manager the feedback, and suggestions on how to fix it. It preempts the engineering review, in the hope that when engineering looks at it there’s less rework. It tries to do this as automatically as possible. Each domain mapped and created their own context. That makes it a little bit more accurate, and also more efficient in how to do the review, the tech design, the implementation and the testing. But eventually, after this stage, it’s ready for engineering review for real, by a human. They assume that there are humans in the loop and they want them still to look at it until they get the confidence.
None of that required a new process. It’s the same lifecycle, with different hands on the pen.
Everyone’s job is now fixing the thing that writes the artifact
They’re running this in a few squads so far, and the others are adopting it pretty fast. They understand the value, they understand the efficiency, so almost all the engineers are contributing to it and the PMs are contributing too. What they are contributing is not documents.
“They are not supposed to write anything on their own anymore. They just need to fix the skills or fix the context. So everything that is written will be written in a better quality.” : Shay Mandel
They fix the system, and then see that the model is able to create the right solution. Shay put it to his own people more bluntly:
“We need a new state of mind. We are now developers, not of the code, we are developers of the system.” : Shay Mandel
At the moment they’re still reviewing the PRs, the code changes. But they are not the initial reviewer, they are the last reviewer, after however many automatic iterations it took. So hopefully it’s zero. You have zero feedback, that’s the goal. And if you do have feedback, the engineer should go in and change the guidelines, the context, whatever. Then next time the reviewer, or even the designer, goes in and does the work better the first time.
Those skill files are the executable specification of how you work, which is the argument I make at more length about what spec-driven harnesses actually are.
Decide which decisions stay human, and say which ones out loud
Once agents write the artifacts, “keep a human in the loop” stops being a policy and becomes a design question, because the loop has a lot of places to sit. It’s an interesting dilemma. At the extreme, I could say that how to code something is a decision. And if I want humans to make those decisions, then I’m back all the way to humans at least reviewing all the code. There could also be a decision of what’s important. Just choosing what’s the feature, what’s the architecture, what’s the outcome that we want to see, the leading indicators, what’s the system behaviour that we want to see. The question is where the right altitude is for the decisions moving forward.
When you execute there are a lot of decisions to make, and they’re trying to let the machine do those. But there are important decision junctions where they still think there’s a place for human judgment. Deciding on the priorities and selecting the features is one. Most of the work on creating the PRD might be automatic, but they expect a human to read it and approve it. On the engineering side they let the agent create the technical design and the plan, and they want a human review before it starts coding. Then the agent runs the PR review, and once it says the review is okay, they still want a human to look at it before it goes to production.
He’s talking about important changes, and in his domain the reason is concrete. Making a mistake there can be very harmful, and it might take a long time to see the output. In insurance you introduce a new product and you start getting claims only after a year. Then you see if it’s too many claims or too little, and whether you should adjust something. Deciding whether a button should be on the top or the bottom is not that. There the agent can go all the way: create the two variants, put the test in production, and come back with the result saying this one is winning. Today they still want to review it and switch manually. Maybe in the future they can push that automatically: if it’s very clear that one solution is better than the other, then they let it go. They’re not there yet. They need to build their own confidence in the system first.
Expectations are the other half of that job, and Shay’s framing is the one I’d hand to anyone whose management has just watched a demo. From idea to POC, like always, it’s very quick. From idea to production level, writing it in a way that will pass all the reviews and so on, takes time.
“We talk about it kind of as grades, so we say the initial version will probably be sixty percent accurate or seventy. And we’ll probably get to ninety percent or ninety-five. And the extra five percent is why we need a human in the loop.” : Shay Mandel
So the first thing to do is to set the expectations: give the demo, but also say that it will take time until you get to this level. That’s leadership work, and it’s cheaper than the alternative.
Your harness is a product, and its hardest users are not engineers
One of the things I’ve seen in other organizations that I work with and talk to is this pattern: the harness is created by engineering. A spec-driven harness that is very engineering oriented, you might say, both in what it focuses on and what it doesn’t focus on, but also in how it’s structured. And when you give it to anybody that’s not an engineer, or even to anybody who isn’t the engineer that wrote it, people don’t really know what to do with it. They don’t even understand what the benefit of it is.
Shay hit the concrete version of that immediately, because his rollout crosses functions by design:
“For product managers and insurance experts, they live in a different place. When you tell them open a terminal and just write these commands, you lost them.” : Shay Mandel
Engineers usually can do some admin or root access on their machines and the others don’t, and it makes sense. So they just need to make sure that the skills they write work for everyone, and that the expectation is set in the right way. And they need very clear getting-started guides. In some cases those are different for the different audiences, or have different levels. You need to think from a product perspective, even about your AI harness and your AI capabilities. You need to make sure it’s something that people find easy to activate and easy to retain, especially when you start to move beyond engineering. It’s the same argument I make about treating your AI context as a product.
The failure mode Shay is watching for is the one that looks like enthusiasm. He talked with other companies, and what worried him is that some of them have a lot of skills because everyone builds their own. They don’t use the generic skill. And then when you look at the statistics, every skill is used by maybe one or two users. A lot of what you build, all these brains, is not really utilized. Next’s answer is that everyone should improve the brain that we have in the company, and that it all sits on top of the same brain. Shay leads a squad for what they call AI enablement. It builds the infrastructure, the guidelines and the methodologies. And they put one of their people inside the business squads, to make sure they’re building the right infrastructure for them, that those squads actually use it, and that they can contribute back. A forward deployed engineer.
And then the detail that pays for all of it:
“Part of their definition of work for this squad is not to deliver the feature, it’s to build a system that can deliver more features in the future.” : Shay Mandel
They shifted some of that squad’s OKRs, with approval from management. It’s okay that the first quarter will be slower, because it will accelerate the quarters after it. This quarter the focus is building the machine. And the machine has to be built so other teams can extend it and other squads can use it. If every squad and every engineer has their own way of working, it won’t be productive and it’ll be very hard to maintain. If you want people building the machine, you have to stop measuring them on the features they didn’t ship while they were building it. That’s a funding decision, and it’s the one place leadership can’t delegate.
When the board stops telling you where the work is, change the process
Next isn’t there yet, and they know it. They think they need to build this kind of dashboard. Right now they look at the epic levels and they see the completion of the epics, and they’re starting to think about how to see everything that’s going on, assuming much more will be going on, and what the right way is to visualize which projects are in flight and which stage they’re in.
One of the things I’ve seen for many years, and I think it applies well here too, is an epic-level Kanban board. First of all, epic-level statuses that reflect all of the touch points between agents and humans. If you think about the PRD: the agent finished writing it, and it’s now waiting for human review. That’s a distinct state change. So all of these state changes, especially when it’s a handoff between humans and agents, show them separately. Then use a Kanban board to reflect where things are, and where things are waiting for the humans. You also get the benefit of flow metrics that show you telemetry of how long it takes in each one of these steps, so you can start to calculate flow efficiency and how much time things are really waiting.
The signal that the process itself needs to change turned up in something Shay described almost as an aside. The extra capacity doesn’t only show up as doing more of the same, faster. It also shows up in what they’re willing to scope. Instead of saying okay, that’s the MVP, that’s phase one, let’s make it as simple as possible, in some cases they can dream bigger. Phase one can be a bit bigger. A bigger change, something more bold that they can move forward. The moment your capacity starts changing what you’re willing to scope, your quarterly planning is working from the wrong assumptions and your board is describing a flow that doesn’t exist any more. Fix the board first. It’s the cheapest instrument you have, and most boards stop too early even before agents get involved.
”Isn’t that just moving slowly?”
It’s the fair objection, and no. Change two things at once and you can’t tell which one is failing, which is where most AI programs are right now. Beyond that, any improvement away from the constraint is meaningless. Deleting a ceremony is almost always away from the constraint. Changing who writes the PRD, and what a person does when it comes back wrong, is right at it.
The fact that settles it for me is that Next’s product and engineering organization isn’t the cautious corner of that company. They implemented AI in support and claims and the customer-facing side about a year ago and have kept improving it since, with product managers deployed there and a team supporting those teams. So when it comes to going agentic, product and engineering is actually behind some of the rest of the organization. That’s worth sitting with if you assumed, as most technology leaders do, that engineering leads this and the rest of the company follows.
The bigger opportunity runs the other way. We look at product and engineering like a factory, like a pipeline, and we understand that there’s a value stream, there is a flow, there’s a set of stages. We create skills for each one of the stages, and we work on optimizing this system with the intent that it becomes more and more agentic. The question is what other factory lines are material to the business. Claims processing, closing the books each month, whatever it is. Then you apply the same concept there. The mindset you need, that engineering mindset of constantly tuning the process, is somewhat more difficult to find in other business processes, in my experience. So applying AI outside of engineering requires bringing an engineering mindset to the rest of the organization, whether that’s by applying engineering capacity to it or by building those skills and that mindset elsewhere.
Listen to the full conversation
Shay Mandel leads product and AI enablement at Next Insurance. He’s on LinkedIn at linkedin.com/in/shaymandel, and he runs the Product Leaders AI meetup, which is Bay Area local and also on Zoom.
The full episode goes further on the auto-improvement loops they point at a KPI, the ask-engineering and ask-product agents they built for a team split across time zones, and where the actuaries, finance and HR come into this: Developers of the System, Not the Code: Inside Next Insurance’s Agentic DLC, also on Spotify.
If you’re standing this up right now: change authorship first, hold the process still enough that you can read the results, and treat the harness as a product with users who aren’t you. The operating model will need to change. It’ll be a much easier argument to make once your own board is showing you where the work is stuck.
Practical thinking on turning AI pilots, adoption, and portfolio work into business impact - by finding the constraint, changing the work, and proving value as you go.
Yuval Yeret helps product and tech leaders move from agile theater to evidence-informed delivery. Work with Yuval →