Clemens Adolphs on Getting AI Investments Right
AI initiatives die in the proof-of-concept stage when run as fixed-scope projects. Clemens Adolphs and I unpack why AI work needs real product discovery.
Click image to open full size Where do AI initiatives actually go to die?
AI initiatives have a well-known graveyard: the proof-of-concept stage. Something impressive gets built, everyone is excited for a week, and then it gathers dust.
I dug into why with Clemens Adolphs, co-founder of AIce Labs, on the Scaling with Agility podcast. He describes the job as helping enterprises get AI initiatives done right end to end — “and not just have everything die in the proof of concept stage and gather dust there.” A lot of the pitfalls he sees overlap with what agility is supposed to help with, which is why the conversation was worth having.
The risk nobody tests
The failure usually gets blamed on feasibility. Clemens puts the weight somewhere else:
With everything you ever build there is that market risk broadly — will anybody actually want it once you build it? And market risk doesn’t necessarily need to mean, oh, you’re building a consumer-facing product and they have to buy it. It could also just be an internal tool that you roll out and you want your people to use, but they don’t use it. There’s too much friction, and they just rather copy paste from ChatGPT.
That is the part organizations skip. I’m tempted to say an AI-native CRM that doesn’t require the salespeople to enter information manually is obviously great, everybody’s going to love that — but I’ve seen so often that even these obvious assumptions eventually fall flat. How often do we think things are really desirable, a field of dreams that people would really come to, and then we jump straight to technology feasibility?
Interestingly, Clemens’s counter-example is about where market risk is genuinely low:
Let’s say the last big project — it already had some internal validation, because they had a basic tool that did the thing and that had adoption. If you’re improving on something that already exists along the dimension that tool was being used, then that typically has very low market risk. It’s like, oh, we’re paying for this tool because it saves us five minutes.
So the question isn’t “is there risk,” it’s “which risk is the live one here.” Improving an adopted tool is a different bet from launching something nobody has asked for.
Prototyping is the new wizard behind the screen
There’s an interesting opportunity in how the AI ecosystem is evolving. If you look at vibe coding, or even things that aren’t vibe coding, there are so many opportunities to prototype, to create preliminary versions of things that work. They’re not proving feasibility from a technology perspective and they’re not necessarily going to be production code ever — but they’re the new version of the Wizard of Oz test.
The old version was a human behind the screen making things work. The new version is ChatGPT or Claude with some gum and glue and a few prompts, and things work, and we can prove some things very cheaply. If those things work, we can make it a more robust solution.
Take the internal market seriously
Here’s the realization I keep coming back to: you have an internal market, and you need to treat it like a market, where people have options. Even when we are developing systems of record, how people use them matters. Your CRM is a system of record, but are people really going to use this feature or that feature, or are they going to do the minimum they have to in order to get away with it? We need to apply the product management and product discovery techniques to the internal market as well. People have a choice to use the products we’re building for them internally — and they also have a choice not to.
You can mandate that they use it, but you can’t mandate that they embrace it. Then they’ll find their workarounds and back channels, and that should tell you something. As Clemens put it, if they’re not using your shiny awesome tool, maybe it’s worth digging into that and treating it as “we’re getting a bad NPS, let’s do some user interviews.”
Or we’re not seeing usage — so let’s look at the pirate metrics internally. Are people aware of our tool? Are they activating? Once they’re activating, are they actually using it? Do they continue using it? Are they paying for it?
That last one is the interesting one, because internally people don’t typically pay for anything. So how do you know people are actually willing to invest in that tool? Maybe it’s enough that they invest their time as a replacement for money. Maybe at some point we’ll see people paying for AI tokens for usage — are they really willing to pay for the AI consumption of their department’s tooling?
Clemens pointed at where that experiment is already running:
It’s interesting in the developer space, where famously developers are cheap when it comes to tooling. But with Claude Code and Cursor and this and that, everybody is shelling out money even if it’s not reimbursed by the employer, because they just see so much value in what the tools do for them. I still think the business should pay for the tool, but it would be an interesting thought experiment — if we took this tool away from you, how much would you pay to get it back?
That relates to the product-market-fit survey question. How dissatisfied would you be if we took this tool away from you? If we took away the AI capabilities on your CRM, how unhappy would you be — or would you actually be very glad?
I know that from my own space. Few people running agility transformations are willing to ask the teams they work with: are we a keeper? Would you fight to keep using the process, the operating system? Would you fight to keep having access to us as an internal consulting service?
If you’re an internal consultant, whether that’s an AI consultant or an agile consultant, it would be very interesting to know what the people you work with actually think. Are they willing to tolerate you? Are they going to fight for your time? Are they going to use the first opportunity to throw you under the bus because you’re not useful? It might be a tough mirror to face, but it’s better to know early than to be surprised when the budget for the enabling function is cut — which happens very often.
The AI-waterfall anti-pattern
I asked what separates the projects that work from the ones that go sideways.
What works is if the objectives and values are clear — what we want to achieve — but there’s freedom and leverage and flexibility in the how.
And the anti-pattern is the one where you have a Gantt chart without calling it a Gantt chart.
And it’s exactly that, and the whole idea of estimates. What does it matter that we estimate something? You told us this needs to happen then. So it needs to happen then. Or velocity charts — what does it matter that we track velocity? You told us how many stories and which ones you need done.
Let me poke at that, because I’m with him but I want to play devil’s advocate. If I’m the business owner or the sponsor, I want some sense of where this is going. I’m going to pay you, or fund my people, and they’re going to go away and do their thing.
What is it that we know? We know we need to achieve a certain outcome. What does that outcome really look like? What are the hard points and the soft points? Of the hard points, the use cases we want to show — how many of those are working already at each point in time, and how do we think those use cases will spread over the two or six months we’re going to work on this?
That framing is useful, and I’ve seen it be useful against scope creep. When you don’t have anything like it, it’s tempting for people to add more and more use cases: this is cool, while we’re at it let’s add this capability and that capability, because we don’t know when this is going to finish anyway, so let’s just keep delivering value. But there are diminishing returns. Is it more valuable to deliver the base outcome after two months, or to deliver with these additional use cases after three? If we’re not even having those conversations, we’re just letting the team figure it out — which might be fine if they have the right context, but only then.
Clemens had the better version of what a status update should sound like:
Last time we showed you it couldn’t do this, now it can do that. That’s way more interesting than just giving them a report that says last week we completed 37 story points.
Velocity is output. Traction is the needle moving.
I try to talk about the difference between velocity and traction. Velocity is output — not activity. Activity is that we spent six hours on this. Velocity is output. Traction is that this is moving the needle towards where we want to be. We’re not just showing you working stuff; it’s working stuff that does what you want it to do, at least a small part of it.
That requires the team to be empowered enough to ask whether the use case you added, or the user story you threw in, is actually going to lead to the outcome you want. The best version of the contract between sponsor and team is: tell us what you need, tell us what you really really want, and we will figure out how. You don’t have to write the user stories — the team that’s going to do the work will slice it into the work they need to do. Tell us the big story of what you’re trying to achieve. Tell us the outcome.
Clemens noted the industry is having its own argument about this:
Recently there was this pushback against even using user stories. Linear, the issue tracker, have their whole philosophy written up online, and their point is write issues, don’t write stories. Maybe that’s the backlash against how they have just been used.
Another recent read was Robert Martin’s Clean Agile, where he says user stories are supposed to be very simple, just a placeholder for a conversation that needs to take place just in time when you actually work on the thing. But they have become this thing where they’re in Jira and they’re three pages long.
And ready months in advance — which creates a lot of waste, because by the time you get there half the acceptance criteria might not be applicable anymore. It’s all spec’d out, the designer already made the design for the fully functional thing in Figma, and now you have this super complicated thing with lots of tabs and dashboards and popups. Good luck doing nice modular vertical slicing on that. Everything is meshed together into a giant story, and because of that you have a feature branch working on that story that lives for two weeks.
There is a lot to fix in this ecosystem.
Principles travel; practices do not
Trying to converge, I put my read back to Clemens: these projects are the classic place to be agile in how you develop them, to think from a product perspective, to think about an internal market, to try things at the right level of agility. He got there first with the better formulation — that these projects are the place to be agile, but not the place to do agile. I don’t think there’s any place to do agile, to be honest.
In this case in particular, the principles very strongly still apply. But if you take those principles and you apply them to a given context, and use that to derive practices, those practices might make sense in that context — and then you port those to a different context and it doesn’t work anymore, and you have to go back to the principles.
That is exactly what we believe in here, and it earned him his badge as a friend of the no-BS, nuance-first version of this podcast.
Watch the conversation
This article is based on my Scaling with Agility conversation with Clemens Adolphs, co-founder of AIce Labs. The full episode goes deeper on engagement models for AI delivery, the Swiss-cheese model of AI capability, and where user stories went wrong.
Prefer audio? Listen on the episode page or directly on Spotify. Find Clemens on LinkedIn.
You can mandate that they use it, but you can’t mandate that they embrace it.
Practical thinking on turning AI pilots, adoption, and portfolio work into business impact - by finding the constraint, changing the work, and proving value as you go.
Yuval Yeret helps product and tech leaders move from agile theater to evidence-informed delivery. Work with Yuval →