How to Measure AI Impact Beyond Token Caps
Every enterprise is arguing about the right per-employee AI spend cap. Tomer Elias and I compared notes on why that number is unanswerable until you can attribute the spend to something.
The token cap is the least interesting number in your AI budget
Somewhere in your company right now, someone is trying to decide whether an employee should get $100, $300, or $1,000 a month of AI spend. It is a real question with a real invoice behind it, and it is also the wrong place to start, because there is no cap number that is defensible on its own. Set it low and you throttle the experiments that would have told you something. Set it high and you have no way to distinguish a team burning tokens on work that shipped from a team burning tokens on sessions nobody ever looked at again. I got into this with Tomer Elias, a product executive who has spent fifteen years around AI and data, including the first AI lab in Israel, a cybersecurity unicorn, and a committee defining agentic identity standards alongside OpenAI, AWS, and Cloudflare. He has been going company by company asking enterprises how they actually decide.
His answer, and mine, converged fast: the cap is downstream of attribution. Until you can trace a block of AI spend to a piece of work, and that work to a stated objective, the cap is a guess dressed up as governance. What follows is the method underneath that claim. It is a ladder from token spend to business impact, and every rung of it is an old measurement problem that AI made expensive enough to finally address. The uncomfortable part, which Tomer names directly, is that the ladder mostly measures your organization, not your AI.
Why can nobody agree on the right cap?
Start with the mechanics of the decision, because they explain the deadlock. Tomer put it in the terms the finance side of the house recognizes:
“AI puts the credit card at the hand of the employees. But without the oversight and training, costs can spiral. But high usage doesn’t really mean a bad thing. It’s not good or bad. The question is what’s the impact that you get out of that.” — Tomer Elias
That last sentence is the whole problem. High usage is not a signal in either direction. A startup burning enormous amounts of tokens is not obviously in trouble, and an enterprise with disciplined, modest per-seat usage is not obviously healthy. It might just be an organization where nobody has found anything worth doing yet. Without a way to evaluate whether the spend was wasteful or impactful, as Tomer said, it does not really matter what the cap is.
I pushed back on one part of this in the conversation, and I want to keep the pushback in, because it changes what you should do about it. This is not a new problem that AI introduced. I have been helping organizations define outcomes, measure them, and close the loop on them for the better part of a decade before any of this. The tooling was different and the bill was smaller, but the failure was identical: teams commit to activity because activity is safe to commit to, and outcomes are not.
There is a diagnostic I use for this. I have a tool I vibe-coded that connects to an organization’s Jira, Linear, ADO, or GitHub and classifies everything in flight into four buckets. Activity: let us do this thing. Output: let us build this feature, this screen, this integration. Outcome: let us enable this kind of user, internal or external, to do something they could not do before. Impact: this moves a business number. When I asked Tomer to guess what share of a typical organization’s managed work sits at outcome level or above, he said twenty percent. Twenty percent is a good organization. Most are below it, and most of what is below it is output masquerading as outcome in the ticket title.
So when a leader says they cannot measure AI impact, the honest reading is usually that they could not measure impact before AI either. What changed is that the incapacity now has a monthly invoice attached to it.
What does AI actually expose about your organization?
Tomer’s framing of this is the line from the conversation I keep coming back to:
“Once you implement AI in your organization, it bubbles up or self-surfaces your DNA and the organizational culture that didn’t change for a while and now needs to change if you really want to push impact with AI.” — Tomer Elias
This is why I think the cap debate is a distraction rather than merely a hard problem. The cap is a spend-control question. What the spend is exposing is a design question about how your organization defines value, who is allowed to decide, and whether anyone closes the loop. Those were true before you bought seats. AI is the contrast dye.
Tomer comes at organizations as an industrial engineer, which produces a useful lens: every company is a factory whose output happens to be a service, a product, or software. Factories went from humans on the line, to humans and machines, to machines, and at each step there were definitions of what good looked like, what waste looked like, what you spent per line and what you got back. The same definitions should exist when you start inviting digital helpers into your workflows. His caveat is the important one, though: an LLM is not a machine, because a machine’s output is the same every time and an agent’s is not. That non-determinism is exactly why you need observability and guardrails rather than a fixed cost-per-unit assumption.
I would add the second half of the factory lens, which is where I think most AI programs lose their money. Goldratt’s rule holds here with no modification: any improvement away from the constraint is meaningless, and any improvement directly at the constraint is a multiplier. The pattern I keep seeing is an engineering organization that gets genuinely faster at writing code while the code review queue, the QA pass, the release approval, and the downstream decision-making are untouched, usually because those steps belong to people who are less comfortable with the tools. You have moved the bottleneck, paid for the privilege, and shipped nothing faster. If the people deploying AI in your organization showed up as continuous improvement engineers rather than as tool champions, they would find the use cases that actually pay.
The attribution ladder, one rung at a time
Here is the practical core. Instead of asking “what is our AI ROI,” which nobody can answer, ask which rung of this ladder you can currently stand on. Tomer’s advice was to start deliberately small, and I think he is right that the sequencing matters more than the sophistication:
“It’s really hard to say, hey, your ARR tripled just because you implemented AI. It could be that it was tripled because of economical changes. And maybe you have now a great sales team. It’s not just because of AI, but the question is whether you can attribute first.” — Tomer Elias
Rung one: can you tie spend to a piece of work? Not to a person, not to a department, but to an initiative. Five people, ten sessions, a thousand dollars of tokens, all pointed at this OKR or this goal. If you cannot produce that sentence today, nothing above this rung is available to you, and this is the rung to build first.
Rung two: did anything leave the session? Tomer’s version of the first honest measurement is unglamorous: did the agentic work produce something that was eventually embedded in your product, or was it a wasteful session, not an experiment but plain waste? A merged PR, a shipped prototype, a document that someone else used. This is the rung that separates token-maxing from work.
Rung three: is anyone using the thing? Take the example we worked through in the episode, because it is the one I see everywhere: an organization decides to aggregate all its meeting recordings and make them available as context. First you can observe that you collected them, which is genuinely better than nothing. Then you can observe telemetry that people are accessing them. Both are cheap. Both are more than most organizations have, which should tell you something about how early this all still is.
Rung four: is it changing decisions? This is where the value hypothesis has to have been written down in advance. Somebody funded the transcript pile expecting fewer meetings, or faster decisions, or better ones. If nobody wrote down which, you cannot close the loop on it, and the honest answer is that you funded a capability without a claim attached.
Rung five: does a business objective move? Tomer’s example here was an approval process that used to take a month, compressed by agents that arrive carrying the full context a human approver needs. Even when a person still signs off, they are no longer going back and forth gathering information. That is the full kit in Lean terms. You want to see the white light of the welder welding, not the welder walking back to the parts bin.
Most organizations I talk to are trying to argue about rung five while standing on nothing. Start at rung one. It is a data plumbing problem, not a philosophy problem.
Fund AI experiments the way a VC funds companies
The other half of the cap conversation is what happens when someone runs out of budget. Today, in most places, they ask and they get more, automatically. That is not a terrible default. It beats throttling curiosity, and it does not turn a finance policy into a referendum on whether someone is working hard enough, but it wastes the moment. The request is the only point in the whole loop where someone is naturally motivated to explain what they are doing.
So change what the request costs. Not more money, more evidence. When you ask for the next tranche, you say what you are working on, what you have to show, and why continuing is a better bet than stopping. This is staged funding: pre-seed, then a checkpoint, then a real conversation. My preference is that the pre-seed round has no token budget at all, so people use their existing subscription. If they come back with a prototype, a proof of concept, an MVP, anything that constitutes evidence rather than enthusiasm, there is a lightweight business case to update. Sometimes the right move is to continue because the option value is worth holding even though you are not sure it will work, which is real options thinking and innovation accounting doing their job. Sometimes the right move is to kill it.
Which means the criteria have to be agreed up front, when nobody is invested. Tomer asked the sharpest version of the question in the conversation: how long would you let the experiment run? The answer being “it depends” is exactly why kill criteria belong in the conversation on day one rather than in a budget meeting three months later.
I would not have a human being sitting in that loop, though. The instinct is to appoint someone to approve token requests, and you have just created a single point of failure whose calendar becomes the constraint. The version I want is an agent that grills you: here is what I want the tokens for, here is what I have so far. It pushes back, it points out that the thing already exists as a subscription, it tells you to go ahead without escalating. The decision rights stay with leadership. The bottleneck does not get a desk.
The build-versus-buy question nobody asks first
Tomer took the conversation somewhere I did not expect, and it is the most immediately actionable thing in the episode. Before any of the measurement machinery matters, most organizations are skipping a question:
“That build mode drove everyone to say, I’ll just use Claude and build it myself. Without really thinking about the experimental stage, you know, how much time will it take you to ramp up to do something that is not part of your core business?” — Tomer Elias
His point is about total cost of ownership, and it is not the usual TCO lecture. It is that your manpower is not built to maintain this, your DNA is not built to maintain this, and the ramp-up plus the iterative support plus making it work at scale is the part that never makes it into the decision. The question to answer before you build is whether the outcome is better than a product you could have bought.
I would sequence it slightly differently. Desirability comes first: would anyone actually use this? Until that is settled, build versus buy is premature optimization. Once it is settled, feasibility and viability come in, and that is exactly where total cost of ownership lives.
But here is what struck me. I have looked at every spec-driven framework and harness I can find, and none of them ask this. They were built by engineers who want to engineer. Nowhere in the flow, not before the spec and certainly not after it, does anything stop and ask: is there an open source project for this? Is there a subscription service? Why are we building this ourselves? That is a gap in the tooling, and right now it has to be filled by a person or a prompt, because it will not be filled by the harness.
Tomer sees the person filling it as an emerging role. Some tech-forward companies he has spoken with have already named someone per business unit to make that call before anyone starts building. He calls the general version a business engineer, the successor to the go-to-market engineer who wires Zapier, Salesforce, and Gong together to make a revenue process run. Their projects are framed as what we want to automate, and their judgment call is build versus buy, weighed against what is actually painful and what will generalize across the company rather than becoming another point solution to maintain.
Give the business engineer a digital twin of your stack
If you are going to make that role work at any scale, it cannot depend on one person’s memory of your tech estate. I coached a go-to-market engineer who was frustrated that AI had told them to build something, and asked me why it had not suggested doing it in Slack, which is already connected to Salesforce. My answer was the boring one: does it know you have Slack connected to Salesforce? It does not have the context.
So we built a skill on the fly that mapped the ecosystem: go look at Slack, map everything you can see, and if you cannot see it, grill me until you do. It found Salesforce and a great deal more. The output is not a diagram for humans. It is a queryable map of what already exists and what considerations a business engineer would apply, so that anyone working on an AI use case, and eventually agents hunting for use cases on their own, can ask whether this makes sense, whether we already have a way to do it, and if not, what is off the shelf.
Tomer’s caution here is real and it is the reason this is harder in a large enterprise than the paragraph above suggests. The blocker is not the mapping, it is the knowledge base underneath it: updated, curated, tagged, classified. He spent years looking at how organizations manage structured and unstructured data across cloud and on-prem, and that is what convinced him enterprises would take longer to adopt AI properly than the hype cycle assumes. Smaller companies can close that gap. Enterprises are carrying a data debt that predates every AI decision they are about to make. I asked him whether anyone is solving this well at enterprise scale, and neither of us could name one. That is either a gap in our knowledge or a genuinely open market, and I suspect the CMDB and asset-management incumbents should be more worried about it than they appear to be.
Measure whether your ways of working are still moving
There is one more loop, and it is the one almost nobody instruments. When I asked Tomer what a manager should actually set as an AI-related goal for their team, his answer was grounded and correct: for a product team, look at how much code was written by AI, how much of it passed review and made it into the product, and whether regressions surfaced that were caused by the way the agents were used. Feed that back to whoever initiated the tooling so the next round is better informed. That is a solid baseline.
I want to add a different altitude, from a Siemens engagement in Israel years ago. We did not call them OKRs at the time, but one of the goals the VP gave the organization was: I want to see experiments in how we work. I want to see us changing the way we work. And specifically, I want to see failed experiments, because a quarter with no failed experiments means nobody left their comfort zone. Today you would call that compound engineering, and it is measurable.
The measurement I would put in front of every engineering leader sits on the files that describe how the agents should work here. Pull requests landing on your CLAUDE.md and AGENTS.md, skills being curated rather than written once, anyone actually using them. Those are the meta-loops, and they are cheap to instrument compared to everything else in this article.
The culture question underneath the cap
I want to close where the cap conversation actually ends up, because it is not a finance conversation.
The Goldilocks problem in front of leaders is real: encourage people to experiment without turning it into “burn tokens at all costs.” What decides which way that lands is not the policy. It is whether people feel safe. In a healthy organization, people experiment, report what did not work, and improve their own throughput without fear. In a toxic one, you get toxic token usage, activity theater with a credit card, and worse, people who quietly avoid improving their own throughput because they think they are cutting the branch they are sitting on. Tomer named the same fear from the enterprise side: performance reviews are already asking employees whether they used AI to hit their KPIs, and the honest answer to that question depends entirely on whether the honest answer is safe.
Which brings the whole thing back around. The token cap looks like a budgeting decision. It is actually a question about whether your organization can state what it wants, trace what it spent, and tell the truth about what came back. Tomer’s advice for leaders starting out was education first, then defining the use cases and goals you want people to attempt, then reviewing the process monthly rather than waiting for the quarter. Mine is narrower: pick one AI investment you are currently funding and walk it up the ladder as far as you can get. Wherever you fall off is the rung to build next, and it will tell you more about your operating model than any cap ever will.
Listen to the full conversation
This article came out of my conversation with Tomer Elias on Scaling AI: From Activity to Impact. The episode runs about 54 minutes and goes further on agentic identity and security guardrails, the data infrastructure blocker, and who ends up building AI capability outside of product and engineering.
Listen on the episode page or directly on Spotify. Find Tomer on LinkedIn.
Watch the full conversation:
You cannot set a defensible cap on spend you cannot attribute. Pick one AI investment, walk it up the ladder, and let the rung you fall off tell you what to fix.
Practical thinking on turning AI pilots, adoption, and portfolio work into business impact - by finding the constraint, changing the work, and proving value as you go.
Yuval Yeret helps product and tech leaders move from agile theater to evidence-informed delivery. Work with Yuval →
- 01 Zoetis CTO on AI Operating-Model Change 18 min