Help me think through how this episode applies to my situation. Start by asking what I am trying to change. Separate the episode's claims from your suggestions, and say when the notes do not support a claim. Use the transcript to find passages, then check the audio before quoting.
Transcript: https://yuvalyeret.com/podcast/episodes/beyond-token-caps-w-tomer-elias-how-enterprises-actually-measure-ai-impact/transcript.md
## Published episode notes
Most enterprise AI conversations get stuck on the wrong number: how many tokens an employee should be allowed to burn. Tomer Elias argues the cap is the least interesting question in the room. The hard part is attribution, knowing whether any of that spend moved a business metric at all. We compare notes on what enterprises are actually doing right now, why AI keeps exposing organizational problems that predate it, and what has to be in place before "AI impact" means anything.
About the Guest
Tomer Elias is a product executive with more than 15 years leading startups from zero to one and one to 100, through unicorn and IPO stages. His focus has been AI and data throughout. He was part of the first AI lab in Israel, helped build a cybersecurity unicorn, and served on a committee defining agentic identity standards alongside OpenAI, AWS, and Cloudflare. He is currently mapping how enterprises adopt AI and where they sit on the maturity curve.
Chapters
00:00 Why this conversation
00:54 Tomer's background: AI labs, a cyber unicorn, agentic identity standards
02:26 What actually changed with gen AI: ambiguity
02:50 Human operating systems and agent operating systems
06:17 Setting a token cap, and why the number isn't the point
07:19 Measuring outcomes was broken long before AI
08:43 Activity, output, outcome, impact: what the work really looks like
10:55 AI surfaces the organizational DNA you never fixed
11:19 The factory lens: an industrial engineer's view of the enterprise
14:31 Goldratt and the constraint: where AI actually pays
15:56 Two levels of observability: spec conformance vs. value
19:00 A worked example: meeting transcripts as context
20:33 Kill criteria and staged funding for AI experiments
22:41 Build vs. buy, and the question nobody asks first
25:09 The rise of the business engineer
25:50 The digital twin that knows your stack
28:06 Data infrastructure: the real enterprise blocker
30:05 Closing the loop: telemetry, usage, decisions
34:12 Psychological safety and toxic token usage
39:32 Security guardrails: MCP safety, DLP, sandboxes
42:54 Who builds beyond product and engineering?
46:54 Teaching product skills instead of staffing PMs
Notable Quotes
"AI puts the credit card at the hand of the employees. But without the oversight and training, costs can spiral. But high usage doesn't really mean a bad thing. It's not good or bad. The question is what's the impact that you get out of that." , Tomer Elias
"Once you implement AI in your organization, it surfaces your DNA and the organizational culture that didn't change for a while and now needs to change if you really want to push impact with AI." , Tomer Elias
"Any improvement away from the bottleneck or from the constraint is meaningless, and any improvement directly at the constraint is a real multiplier." , Yuval Yeret
"If it's a healthy DNA, people would feel safe to experiment. And if it's a toxic culture, you would get toxic token usage and activity theater." , Yuval Yeret
Links and Resources
Tomer Elias on LinkedIn: https://www.linkedin.com/in/tomer-elias1
If you're working on turning AI activity into business impact, I write about this every week in the Scaling w/ Agility newsletter.
Scaling AI: From Activity to Impact , yuvalyeret.com
The Scaling w/ Agility Newsletter – yuvalyeret.com/insights
Yuval's Linkedin – https://www.linkedin.com/in/yuvalyeret/
How to Measure AI Impact Beyond Token Caps – https://yuvalyeret.com/blog/how-to-measure-ai-impact-beyond-token-caps
Is Your Agent Harness Still Learning? Check Its Last Updated Time. – https://yuvalyeret.com/blog/your-agent-config-is-standard-work
## Transcript
Automatic transcript of the published podcast audio. Recognition errors are possible. Speakers are unlabeled; do not attribute a passage to Yuval or a guest without checking the audio.
Episode: https://yuvalyeret.com/scaling-ai-podcast/beyond-token-caps-w-tomer-elias-how-enterprises-actually-measure-ai-impact/
Source RSS GUID: 056c8877-b280-4cad-ac85-9b1b29504332
Source: published Riverside RSS audio enclosure
Transcription: faster-whisper base.en, English
## Transcript
[00:00:00] Welcome to the Scaling AI from Activity to Impact podcast.
[00:00:05] I'm Yvall Yerit, and today with me is Tomer Elias.
[00:00:09] I met Tomer at a private product leadership community,
[00:00:14] and it was pretty obvious very quickly
[00:00:18] that we share an interest area in exploring
[00:00:22] what are people doing out there to evolve beyond token maxing
[00:00:28] and AI activity and starting to get some real outcomes
[00:00:33] and some real impact.
[00:00:34] So I thought let's invite Tomer here,
[00:00:37] and let's compare notes.
[00:00:40] What I'm seeing, what you are seeing Tomer.
[00:00:43] So welcome to the podcast for people
[00:00:45] that aren't familiar with you.
[00:00:47] Would you share a bit about your background
[00:00:50] and what brings you to be interested in this challenge?
[00:00:54] Sure, so it's great to be here.
[00:00:56] My background is more than 15 years of product executive.
[00:01:00] I've led multiple startups from 0 to 1,
[00:01:03] from 1 to 100 in unicorns and IPO.
[00:01:05] My focus was always around AI and data.
[00:01:10] That's something that I found myself the most passionate about.
[00:01:17] I was part of the first AI lab in Israel,
[00:01:20] built a cybersecurity unicorn at Big AD,
[00:01:24] and I also was part of a comedy that helped define
[00:01:28] what would be an energetic identity, right?
[00:01:31] The best practice, the protocol together with OpenAI, AWS,
[00:01:36] CloudFlare, and recently I started to see
[00:01:39] how everyone are talking about tokens and cost of tokens,
[00:01:42] and I tried to figure out how enterprises today adopt AI
[00:01:47] and what's their maturity level, started to talk about it
[00:01:50] within the community that we are at,
[00:01:52] but also reached out to friends of mine and people on LinkedIn
[00:01:56] to get them to share a little bit more on the stage that they are at.
[00:02:01] And cool, that's very cool.
[00:02:02] So I mean, I guess the first question, not the main,
[00:02:06] but maybe the first question that I would ask you
[00:02:08] since you've been around AI,
[00:02:12] I'm curious to hear how different is this age of gen AI
[00:02:19] and what's going on around it from, let's call it,
[00:02:23] classic traditional machine learning style.
[00:02:26] I would start with ambiguity, you know, machine learning
[00:02:29] or other algorithms are mostly definitive.
[00:02:33] When it comes to a lot of labs or the use of authentic tools,
[00:02:36] the outcomes are never the same.
[00:02:38] You need to put up a lot of effort to make it produce the same outcome
[00:02:43] or same level of quality of outcome across workflows
[00:02:48] and use cases.
[00:02:49] Sounds like people.
[00:02:51] Like, yeah, I mean, if you look at my background,
[00:02:55] like a lot of my work for people that aren't familiar is in managing,
[00:03:00] it's called a human operating system.
[00:03:02] So my background initially was in computer operating systems,
[00:03:06] Linux, kernel stuff, networking, routing.
[00:03:11] But then as from becoming an engineering leader of organizations
[00:03:16] building software operating systems,
[00:03:19] I eventually started to focus more on the human operating system
[00:03:23] and culture hacking.
[00:03:26] And, you know, I realized humans are pretty unpredictable as well.
[00:03:31] And there's a lot of ambiguity when leading human systems
[00:03:35] and you need to create certain guardrails to still get value
[00:03:39] out of human systems, but you also need to provide agency.
[00:03:43] And there's an interesting, ICN, actually an interesting corollary
[00:03:51] similarity between what works, what I've seen work
[00:03:56] to unleash the potential of people versus what works to unleash
[00:04:01] the potential of agents and what guardrails do you need?
[00:04:05] Like, you could give a developer four months to work on something.
[00:04:09] You would waste a lot of money to give them a developer
[00:04:13] team, you know, a product team, you know, multiple months,
[00:04:17] you're wasting a lot of dollars and you're not sure
[00:04:20] that you're getting something.
[00:04:21] So there's a lot of similarities between that and it's quite similar.
[00:04:25] But in terms of the waste or the cost itself, with humans,
[00:04:30] you have something that is definitive, you have time and money
[00:04:33] and manpower with AI, it's not definitive.
[00:04:36] It could be that because of an error or a loop or something
[00:04:40] that was not really configured correctly.
[00:04:42] Or because people don't have the site, AI can cost you way more
[00:04:47] in very short period of time if you don't have the right guardrails
[00:04:51] and observability.
[00:04:52] Yeah.
[00:04:53] Yeah.
[00:04:54] So part of the conversation, and I was chatting to somebody else
[00:04:59] in the space a couple of weeks ago, and we were talking about
[00:05:05] this rogue AI, you know, agents that are cost endless tokens.
[00:05:12] And I think part of that conversation is, you know,
[00:05:16] what's the model that you even use?
[00:05:19] I mean, it's a choice to give your agents unlimited budget
[00:05:24] or to give your humans unlimited tokens in a time period.
[00:05:29] Like, I don't work at an enterprise.
[00:05:31] I'm an individual.
[00:05:33] And I know that my, the way I work is I work through the subscriptions,
[00:05:37] right?
[00:05:39] So if I start, if I give cloud code or codecs or tragic,
[00:05:45] if you work as it's called as of today, a goal that is complex,
[00:05:50] what's the worst case scenario that it would hit its window
[00:05:54] and say, okay, I'm going to rest for a bit.
[00:05:57] So I think it's a choice that organizations have or the people
[00:06:02] that get, you know, an unbounded budget.
[00:06:06] Is that how you're thinking about it or how, how, what are people doing about this?
[00:06:11] Yeah, so I'll share a little bit from the conversation that I had with some enterprises.
[00:06:16] First of all, the challenge is to define those, the cap itself, right?
[00:06:21] Would you say 100 a month for an employee or 300 or 1000?
[00:06:27] How do you make sure you do not really block productivity
[00:06:31] or block real product development by putting the cap to low?
[00:06:36] Also, Priceline just say that AI puts the credit card at the hand of the employees,
[00:06:42] but without the oversight and training cost can spiral.
[00:06:45] But high usage doesn't really mean a bad thing.
[00:06:49] It's not good or bad.
[00:06:51] The question is what's the impact that you get out of that?
[00:06:54] And that's something that is really difficult to quantify today.
[00:06:57] And it really resonates with what you say, because you can put the guardrail,
[00:07:01] but if it's 100, 300 or 1000, when you don't have a way to really evaluate what the outcome,
[00:07:10] whether it's wasteful or impactful to organization,
[00:07:14] it doesn't really matter what's a cap.
[00:07:17] So I'm going to push, I mean, yes, but I'm going to push back on this for a second.
[00:07:21] Again, bringing it back to people.
[00:07:25] Yes, but I can tell you that I've been working to help organizations,
[00:07:30] figure out what are the outcomes of what we're doing and measuring outcomes
[00:07:35] and closing the loop on outcomes for almost a decade before AI,
[00:07:40] because you have a lot of KPIs or KPIs, but you have a lot of teams
[00:07:46] that are focused on vanity metrics.
[00:07:49] Like if you look at most engineering organizations,
[00:07:52] most even product managers, they shy away from measuring outcomes and impact,
[00:07:58] because it's harder to commit to outcome and impact.
[00:08:01] So, I mean, it's a, I think it's a big opportunity that AI is shining a light on,
[00:08:11] but it's not necessarily a really new challenge.
[00:08:14] So, and the fact that it's not new doesn't mean that we have a great solution,
[00:08:19] but we've learned a lot of things along the way about what's the bottleneck
[00:08:24] to really focusing on outcomes.
[00:08:27] And in some of my conversations and the work that I'm doing with organizations,
[00:08:32] I see how hard it is for leaders to actually focus on outcomes.
[00:08:39] Like I have this tool that I use, I've quoted with AI that I connect to an organization's
[00:08:48] GRA instance, okay, or whatever they're using, get up linear ADO there.
[00:08:55] And what it does is it analyzes the work that people are focused on,
[00:09:03] whether it's the goals, OKRs, features, GRA, FX, stories, whatever it is.
[00:09:10] And it tells an organization, it gives them visibility to how much of the work is managed
[00:09:17] as let's do this thing as an activity, how much of the work is managed as output,
[00:09:22] let's build this feature, let's do this screen, let's do this button,
[00:09:26] let's add this integration, whatever, how much work is managed as outcomes?
[00:09:31] Like let's enable this type of customer to do a certain thing,
[00:09:35] whether it's an internal customer and internal persona or a real end user or customer.
[00:09:42] And how much is impact?
[00:09:44] Like things that really move the needle from this perspective,
[00:09:46] which is not an ideal way to manage the work that's a side conversation.
[00:09:50] You want to guess, when I look at the typical organization,
[00:09:56] what's the percentage of work that is managed in each one of these categories?
[00:10:03] I would say about 20% outcome.
[00:10:08] Yeah, that's a good organization.
[00:10:11] Yeah, yeah, and at this point, I'm not really surprised.
[00:10:17] It's both hard to think about your customers, but there's also,
[00:10:23] like the fact that you're willing to put it on paper or on record,
[00:10:28] that this is the outcome you're focused on.
[00:10:32] I'll be able to measure it at the end.
[00:10:34] I'll be able to see whether you're actually moving the needle
[00:10:38] and that requires a maturity of the organization that a lot of organizations are struggling with that.
[00:10:43] So you see, I do, so I think that both of us agree,
[00:10:53] once you implement AI in your organization,
[00:10:55] it bubbles up or self-surface your DNA and the organizational culture
[00:11:01] that didn't change for a while and now needs to change if you really want to push impact with AI.
[00:11:07] It just made people be more structured when it comes to the implementation process
[00:11:14] and the evaluation process of those investments.
[00:11:17] My perspective coming from an industrial engineer,
[00:11:20] so I look at enterprises and companies as factories at the end of the day.
[00:11:25] Either the outcome, it can be a service, it can be a product, it can be a software.
[00:11:30] But when factories started, humans were on the line, right?
[00:11:35] They had to do things in order to deliver the product.
[00:11:38] And then it was man and machine, right?
[00:11:41] It was humans and machines and later on, it became only machines.
[00:11:45] Along those ways, you had to add definitions of what's impact, right?
[00:11:50] What's good looks like? What's wasteful?
[00:11:52] How much money you put on that line and what's the outcome that you get and whether it's valuable?
[00:11:59] The same should be when we invite more digital twins or digital helpers
[00:12:06] to be part of our agentic repositories or organizational workflows.
[00:12:11] But the situation coming back to LMs is that LMs are not like a machine
[00:12:18] because the outcome cannot really be trusted.
[00:12:22] It's not like the same every time that the LMs would do one workflow,
[00:12:28] the output would not be the same.
[00:12:31] So in that perspective, you need better observability and better control
[00:12:35] on the actions and the activity of the agent.
[00:12:39] And you need also better monitoring on what's out of your guardrails
[00:12:44] and in terms of costing, in terms of abuse, in terms of waste, and in terms of impact.
[00:12:49] And without having that definition, and that's by the way,
[00:12:52] something that I do see also already happening in enterprises that are not tech forward.
[00:12:57] Some enterprise that I chat with, I won't say their names,
[00:13:02] companies that are not, doesn't really develop software, already started to define OKRs
[00:13:08] and KPIs for each business unit.
[00:13:11] And today, they ask during performance review, they ask the employees whether they utilized AI
[00:13:17] in order to achieve their KPIs and OKRs.
[00:13:19] And later on, when you'll see more agents over there,
[00:13:22] they will try to tie those back to agent outcome or agent impact.
[00:13:29] So you see how those building blocks are stacking along for that future.
[00:13:34] So I 100% agree with the factory metaphor.
[00:13:39] And I think applying some of the techniques from the factory continuous improvement movement,
[00:13:48] Goldrat, you know, lean is what we need to do here.
[00:13:54] I, and I think it applies at multiple levels.
[00:13:58] So one level is that I'm seeing as useful is if people that are deploying AI in the organization,
[00:14:08] those people with those OKRs and KPIs are looking at their business processes
[00:14:15] using the factory lens and they show up as industrial engineers
[00:14:20] or continuous improvement engineers or professionals.
[00:14:23] That's the right mindset to find the AI use case that would really move the needle
[00:14:28] because like Goldrat said, this, you know, any improvement away from the bottleneck
[00:14:33] or from the constraint is meaningless and any improvement directly at the constraint
[00:14:37] is, you know, a real multiplier.
[00:14:40] So that's the thing you want to focus on and that's the way to drive AI usage effectively
[00:14:47] from activity to real impact.
[00:14:50] But if you go back to the, let's call it the design factor,
[00:14:54] where it's the organization that is building these capabilities that we're trying to.
[00:14:58] I mean, I see a couple of levels of what it will need a bucket for experimental, right?
[00:15:04] If you think about it, you have a bucket for production, like what we do with AI
[00:15:09] to move things to production to product.
[00:15:12] But you need to have a bucket for experiments, you know, to build things, new things, right?
[00:15:16] Yeah, but I think even that is a design factory.
[00:15:21] The design mindset still applies.
[00:15:24] Bank driven still applies.
[00:15:25] You still want to think about intent.
[00:15:27] So you give some agency to figure things out.
[00:15:29] And then you want to tighten the design.
[00:15:32] Correct.
[00:15:32] You'll explore, converge, build, all of that still applies even when you're experimenting
[00:15:38] but I guess there are a couple of different guardrails that are important.
[00:15:45] You mentioned observability and I'd love to go into that.
[00:15:48] What I see in a lot of organizations is the first thing that they want,
[00:15:53] that they tackle tightening is they are building what we ask it to build.
[00:16:00] And that is the thing that we, is the thing that it built operating to the specification
[00:16:08] that we aligned on.
[00:16:11] And that's one level of the challenge.
[00:16:13] I would say building something that meets the specification using, you know, reasonable
[00:16:21] budget, you know, time is not that much of an issue.
[00:16:25] And I would say observability at that level is not that.
[00:16:29] I mean, it's not that big of a challenge.
[00:16:32] I mean, it is a challenge to make sure that you do regression testing and all the things that, you know,
[00:16:37] a good solid software factory should be doing.
[00:16:40] And if they're not, AI is amplifying the problem.
[00:16:44] Yeah, it also can help them deal with that.
[00:16:49] But you still see organizations that the developers are focused on AI coding
[00:16:56] and nobody is looking at the downstream bottleneck and improving that because it's people that
[00:17:02] aren't as comfortable with, you know, with leveraging AI.
[00:17:07] But it takes time.
[00:17:09] Yeah.
[00:17:10] Organizations solve that and scale the AI from activity to good output.
[00:17:14] Now it's the real challenge of observability.
[00:17:17] Like, how do you observe whether they actually move the needle the way you wanted it to?
[00:17:24] And the first challenge is that it's often not even defined.
[00:17:29] What did you want to achieve?
[00:17:32] And that's correct.
[00:17:33] It's not necessarily defined in a measurable way.
[00:17:36] And you cannot always attribute everything to AI, right?
[00:17:40] You start with, let's say, small steps towards it.
[00:17:45] It's really hard to say, hey, your AR are tripled just because you implemented AI, right?
[00:17:51] And maybe that it was tripled because of economical changes and maybe you have now great sales
[00:17:58] team.
[00:17:59] It's not just because of AI.
[00:18:01] But the question is whether you can attribute first.
[00:18:04] Like, do you have a way to attribute every kind of change in your organization and tied back
[00:18:10] to an agentic workflow or AI process?
[00:18:14] You start with that, right?
[00:18:17] Maybe to factories, right?
[00:18:19] You need to trace back to the cause to understand whether it was part of or impactful at all.
[00:18:27] So I would say start with that.
[00:18:30] Start with realizing whether your agentic code output something that was at least embedded
[00:18:38] into your product eventually or was it just a wasteful session or multiple sessions that were
[00:18:44] wasteful, even not experimental.
[00:18:46] I think that taking those baby steps would help eventually with measuring impact.
[00:18:51] And that's my kind of advice to everyone while trying to implement AI.
[00:18:56] So let's maybe take a use case and work through it.
[00:19:00] So a common one that I see all over is we're having a lot of meetings in the organization
[00:19:07] and we believe that if we aggregate the recordings from all of these meetings and add that to
[00:19:15] context, it will be useful.
[00:19:19] So if I take the approach that you talked about, the first level would be to measure how many
[00:19:29] people in the organization or maybe it's one person or one team that took on this challenge.
[00:19:34] Okay?
[00:19:35] And you can look, you need to be able to say these sessions that used these amount of tokens
[00:19:44] were used to work on this challenge, on this OKR goal, whatever structure you used to manage
[00:19:52] that.
[00:19:53] And then to be able to say at each point in time, okay, so far we've had five people work on this
[00:20:00] with 10 sessions on it, we've wasted $1,000 in tokens on it and so far there's no output.
[00:20:10] For example, so far it's only games.
[00:20:13] Nobody checked in and EPR.
[00:20:15] Nothing is running as a capability in the organization.
[00:20:19] How long would you allow the experiment to happen?
[00:20:22] That's a good question.
[00:20:24] I mean, I think of course the answer is it depends.
[00:20:28] I'm a consultant.
[00:20:29] The answer is it depends.
[00:20:31] I think the important conversation is up from talking about key criteria.
[00:20:35] So when you manage your portfolio of experiments, you'd say, okay,
[00:20:39] this is one of our top priority things.
[00:20:41] We want to run on it.
[00:20:43] But like we do in VC funding, you get pre-seed money.
[00:20:47] You get a couple of weeks with a certain token budget.
[00:20:51] Maybe initially, by the way, I would say you don't get a token budget.
[00:20:55] You just use your subscription.
[00:20:57] Yeah.
[00:20:58] Try to run with it.
[00:20:59] Try to run with it.
[00:21:00] If you come back and show promising results, hopefully, you know, I wouldn't consider the fact
[00:21:08] that there's nothing to show for us promising results.
[00:21:11] So probably there needs to be some sort of output.
[00:21:16] It can be an MVP.
[00:21:17] It can be a proof of concept, whatever.
[00:21:19] But you come back with it and you show this as promise.
[00:21:22] There's an updated mini lightweight business case.
[00:21:26] Maybe even AI looks at that business case and says, okay.
[00:21:29] Now, even if we're not sure that this will eventually work,
[00:21:34] it's still worthwhile to explore that option.
[00:21:36] That's a real options approach and innovation accounting.
[00:21:40] Now, it makes sense to continue to work on this, maybe.
[00:21:44] Maybe not.
[00:21:45] Maybe we kill it at that point.
[00:21:46] And there's probably a good conversation to have in each organization around
[00:21:51] build care that that's a decision.
[00:21:53] Take place.
[00:21:54] Yeah.
[00:21:55] What do we empower individuals to decide?
[00:21:59] Maybe it's a process of like today, a lot of organizations have a situation
[00:22:04] where you run out of tokens.
[00:22:05] You automatically can ask and get the tokens, which is not a bad approach.
[00:22:10] But I can see a reality where when you ask for tokens, you say,
[00:22:15] this is why I want these tokens.
[00:22:17] This is what I'm working on.
[00:22:18] This is the evidence that I have so far.
[00:22:20] And I can automatically say, you know, and cut away.
[00:22:23] You know, this is interesting.
[00:22:24] Go ahead.
[00:22:25] We don't need to go and talk to anybody.
[00:22:27] You know, I'm just helping you think through this.
[00:22:30] And if you think it's worthwhile, go ahead and do it.
[00:22:35] And go ahead.
[00:22:36] Maybe I should take you even back.
[00:22:39] Lots of conversation starts with whether I should buy or build it, right?
[00:22:44] That build mode drove everyone to say, I'll just use cloud and build it myself.
[00:22:49] Without really thinking about the experimental stage, you know,
[00:22:53] how much time will it take you to ramp up to do something that is not part of your core business, right?
[00:22:59] It's not what you're supposed to do as a business to build your own tools
[00:23:03] because your manpower are not built to do that.
[00:23:06] Your DNA is not built to do that as well.
[00:23:09] And whether you have the right capacity to also not just build,
[00:23:13] but iteratively support and maintain and make it work in scale.
[00:23:18] And you know, those things are, let's say, I wouldn't say ideas,
[00:23:22] but best practices that most wouldn't think about in advance.
[00:23:26] But once you decide to go that route,
[00:23:29] you need to understand how much time it will take to ramp up,
[00:23:32] how much time it will take to trial and make it happen,
[00:23:35] and whether the outcome is better than a product that you could have bought it for.
[00:23:40] So you need to look at desirability.
[00:23:42] Would people even use it?
[00:23:44] And until you know, don't worry too much about build versus buy.
[00:23:49] Once you align on desirability, there is the question of visibility and viability.
[00:23:55] And there you do need to think about total cost of ownership.
[00:23:58] I 100% agree.
[00:23:59] And I would also say that if you look at all of this spec-driven frameworks,
[00:24:04] harnesses that I looked at, they don't ask those questions.
[00:24:08] You try to engineer the tool.
[00:24:10] That created harnesses because they want to engineer.
[00:24:14] They're not asking questions around before you do anything.
[00:24:19] Maybe even before the spec, for sure after the spec,
[00:24:24] is there an open source for this?
[00:24:26] Is there a subscription service for this?
[00:24:29] Why are we doing this ourselves?
[00:24:32] So what I do see in tech-forward companies, those that I spoke with,
[00:24:36] some of them already defined a role for each BU,
[00:24:41] a person that's supposed to evaluate those things before they start rebuilding.
[00:24:46] Think about a person, like today we have a go-to-market engineer.
[00:24:50] They build agent tools in order to drive go-to-market business using automation.
[00:24:56] It doesn't have to be agents or tokens, et cetera.
[00:24:59] It can be also taking things with ZAPR and tying them back altogether
[00:25:04] to make things run smoothly or faster or more efficiently.
[00:25:08] I believe that in the future, you'll see more business engineers.
[00:25:11] Within business units, you'll have business engineers that their projects
[00:25:17] and problems will be tied to what we want to automate
[00:25:22] and make things run smoothly.
[00:25:24] They will decide whether to build or buy based on their overall perspective,
[00:25:28] what's painful, what's really will move them middle,
[00:25:31] and what will allow us to amplify it also across the company
[00:25:35] and not just build a point solution for now that maybe later will cost us more
[00:25:40] because we need to maintain it.
[00:25:42] So let's take it to the next level and assume we're really AI-native.
[00:25:46] What I see happening is that each one of these business engineers
[00:25:50] has a digital twin agent that is aware of what's the current tech stack,
[00:25:59] whether it's through configuration management, assets, mapping.
[00:26:03] It has access to a map of what you currently have.
[00:26:06] Maybe you even have the way to do this internally,
[00:26:10] that the person that is starting to look at it with their AI
[00:26:13] is not thinking about.
[00:26:15] I've coached a go-to-market engineer where AI told them to do something
[00:26:20] and they said, why didn't it suggest to go to do it in Slack
[00:26:24] that is connected to Salesforce?
[00:26:26] Does it know that you have Slack that is connected to Salesforce?
[00:26:30] The context is missing.
[00:26:31] It doesn't have the context.
[00:26:32] So we built on the fly a skill of mapping the ecosystem.
[00:26:37] We told go look at Slack, map everything that you know.
[00:26:40] It found out that there are Salesforce and there's a lot of stuff.
[00:26:43] It's not to have it grill you for a second,
[00:26:46] but understand the ecosystem and try to dump the considerations,
[00:26:51] to map the considerations that the business engineer will take into account,
[00:26:56] so that eventually what you want when it comes to agency
[00:27:00] is everybody that is working on a iOS case.
[00:27:03] And at some point, agents that are trying to find effective use cases
[00:27:09] they can go to this agent and ask it.
[00:27:12] It doesn't make sense to do this.
[00:27:14] Do we have a way to do it right now?
[00:27:16] If we don't, you know, is there something off the shelf that would do it well?
[00:27:20] Like what are the options?
[00:27:23] I like what I don't want to see is bottlenecks
[00:27:29] that are because we're relying on single point of failure people.
[00:27:35] I don't want to see people that become bottlenecks
[00:27:38] and need to be in each one of these loops.
[00:27:41] And now in order to really make it to you build,
[00:27:45] get yourself outside of this loop while still managing the decision-making.
[00:27:50] That for me is the question of scaling.
[00:27:53] So actually now in order to really make it work,
[00:27:56] getting back to the problem that AI surface all the, let's say,
[00:28:00] structural, cultural challenge of an organization,
[00:28:03] making sure that you have the right knowledge base
[00:28:06] that is updated, curated, tagged, classified, et cetera.
[00:28:11] That's a huge problem.
[00:28:13] Enterprises, it will be very hard for them to build something that you mentioned,
[00:28:18] like them mapping the stack and putting the context in the stack
[00:28:21] based on departments and different be used.
[00:28:23] Smaller companies will be able to close the gap and maybe make some order
[00:28:27] and make sure they have the right updated knowledge base
[00:28:31] and their AI will get the context that it will be fresh and correct.
[00:28:35] But the challenge, it was always data.
[00:28:38] Looking back like five years ago when I was part of Big AD,
[00:28:42] looking back at how organizations manage data in structured, unstructured, cloud,
[00:28:49] all of that allowed me to understand that enterprises will,
[00:28:53] they will take more time to really adopt AI properly.
[00:28:58] It's not just the guardrail and governance, it's also to me to make sure
[00:29:02] your infrastructure, your data infrastructure is ready for that.
[00:29:06] And did you see anyone trying to solve that better?
[00:29:10] I haven't talked to anybody that's solving this at the enterprise scale.
[00:29:15] I'm really curious, like, I don't think any of the people that worked at HPSCMDB
[00:29:21] when I coached them on how to do agile 15 years ago is still there.
[00:29:26] But I'd be curious to see here what the CMDB and asset management players are doing about this.
[00:29:33] I can imagine them and service now and these different players having a perspective on this.
[00:29:39] Maybe the other area, all the way downstream in the process that I'm curious what you're seeing is
[00:29:48] observability.
[00:29:50] So observability, not whether the solution works, but whether it's, or not whether the solution can work,
[00:29:58] not regression testing and, you know, and evolves.
[00:30:04] But are people actually using it?
[00:30:07] Is it useful?
[00:30:10] Is it in production creating results that are within boundaries?
[00:30:16] All of that closing the loop on, it's not necessarily impact.
[00:30:21] I agree with you.
[00:30:22] We don't need to be at the point where we can attribute, you know, collecting all of the Zoom
[00:30:28] or Google need transcripts and putting them somewhere.
[00:30:31] It would be hard to attribute that directly to revenue growth or profitability.
[00:30:37] But if we made an assumption that, I don't know, will make better decisions
[00:30:44] if we have access to this or there would be.
[00:30:48] It can take your meetings.
[00:30:50] It can save time.
[00:30:51] I don't know.
[00:30:52] There needs to be an hypothesis about what's the value?
[00:30:56] Why are we doing that thing?
[00:30:58] How do we, how do we close the loop on, are we really getting that value?
[00:31:05] So for example, if we collected all of those transcripts for everybody, we can observe that
[00:31:12] we've collected them.
[00:31:13] That's a good start.
[00:31:14] That's better.
[00:31:15] That's better than nothing.
[00:31:16] And we can see that somebody is using them.
[00:31:19] Exactly.
[00:31:20] So we need to be able to see the limitary that people are accessing this.
[00:31:25] And that's a good start.
[00:31:26] Is it driving decisions?
[00:31:28] That's a next level.
[00:31:30] That's exactly the next level.
[00:31:31] The next step would be to tie it to a business objective in your organization and see whether
[00:31:37] it helped somewhat.
[00:31:40] Either by making it faster, shortening the time, less people need to be involved, faster
[00:31:46] maybe decision making.
[00:31:47] I saw some examples of agents that help with an approval process that took, let's say month
[00:31:54] within an organization and now with agents that come with all the needed context for every,
[00:32:00] even if there's person at the end of the day that needs to approve, they have all the information
[00:32:05] they need and approve the process.
[00:32:07] Exactly.
[00:32:08] It's a full key.
[00:32:09] They don't need to go back and forth.
[00:32:10] It's like the...
[00:32:11] Yeah.
[00:32:12] The world there in Goldrat's example, you always want to see the white light from the world they're
[00:32:19] welding rather than going back and forth.
[00:32:21] That's what you want to measure, right?
[00:32:23] Exactly.
[00:32:24] So we are in an early maturity stage.
[00:32:27] That's what you see at enterprises.
[00:32:30] Smaller companies are adopting it faster.
[00:32:32] And that's why you hear different startups that are saying that they consume a lot of tokens
[00:32:37] and they are burning tokens.
[00:32:39] It's not a bad thing.
[00:32:40] You just need to make sure that you burn them efficiently.
[00:32:43] You know, your people are doing, they're not abusing it, of course, but they are using the
[00:32:48] harness correctly in a way that wouldn't cost you more money.
[00:32:52] But eventually, high usage can drive faster product development, can drive faster decision-making,
[00:33:00] can streamline your business objectives, could be more customers, bigger pipeline or building
[00:33:07] a new pipeline.
[00:33:08] It really depends on your kind of initiatives and use cases and what you evaluate that would
[00:33:13] be more impactful for you.
[00:33:15] Yeah.
[00:33:16] So, all right.
[00:33:17] So maybe let's pause for a second.
[00:33:22] And that's one good advice that you're offering for people out there, which is, it's okay
[00:33:29] to experiment, it's okay to see usage.
[00:33:34] I would say, while that's going on, have some default guardrails.
[00:33:39] It's okay to extend if there's outcome, but it's safer to start with some guardrails.
[00:33:49] And I would say be careful with how you navigate the conversation around what people do with
[00:33:57] their subscription.
[00:33:59] So it's the tricky, goal deluxe of encouraging people to try and experiment, but it's not
[00:34:07] about use the tokens at all costs or whatever.
[00:34:12] It's all about psychological safety, I would say, eventually it depends on the DNA of the
[00:34:16] organization.
[00:34:17] Exactly.
[00:34:18] If it's a healthy DNA, people would feel safe to experiment.
[00:34:21] Even if it's a toxic culture, you'd get toxic token usage and activity theater and people
[00:34:29] that are afraid to improve their throughput because they're thinking that they're cutting
[00:34:36] the branch that they're sitting on.
[00:34:38] But let's assume you're an organization and you're at that point.
[00:34:45] You're seeing people using, they're starting to use more and more, what would be the first
[00:34:51] next step that you would suggest for these organic data streams?
[00:34:56] So it starts with education, making sure people are educated enough to use that.
[00:35:03] And instead of telling them, hey, just use it, you need to help them as a manager to define
[00:35:08] use cases or OKRs that you would like them to try and achieve using those tools.
[00:35:14] And similarly to how you do quarterly reviews, you would like to do maybe a monthly review
[00:35:22] on the process itself and give them helping hand in that kind of process.
[00:35:28] So let's go.
[00:35:29] I agree with you.
[00:35:30] Let's do maybe an example of what such an OKR that a manager gives their team might look
[00:35:38] like.
[00:35:39] At what altitude do you think that would be?
[00:35:42] So an easy thing would be product, right?
[00:35:46] When you see a scrum team and you ask scrum team to build that capability into the product,
[00:35:52] you would like to see them using AI in order to do that, right?
[00:35:56] And evaluating how much pieces of code was written by AI, right?
[00:36:01] Evaluating how much of those pieces of code passed the review, right?
[00:36:05] Went from PR and QA and really is part of the product.
[00:36:10] And eventually you would like to also monitor whether there are regressions testing that found
[00:36:15] things that were not supposed to happen because of the use of AI.
[00:36:19] Having those loop and having sure that there is a feedback and sharing that feedback with
[00:36:25] whoever initiative the usage of those coding agents will learn and be educated about what
[00:36:32] should happen next in order to improve it, that's the gold I think standard here for a
[00:36:37] start.
[00:36:40] A later one, a lot of one for the product in this.
[00:36:44] Yeah.
[00:36:45] I worked with Siemens in Israel many years ago and one of the OKRs, we didn't call them
[00:36:51] OKRs then, but one of the OKRs that the VP gave the organization was, I want to see experiments.
[00:37:01] I want to see compound engineering, you would call it today.
[00:37:05] I want to see us changing the way we work all the time.
[00:37:11] So I don't want to just see experiments, I want to see failed experiments.
[00:37:16] I want to see that we're challenging our comfort zone and, you know, we're trying different things
[00:37:21] in how we work with AI all the time.
[00:37:25] That would be another interesting, interesting example.
[00:37:28] And you can start with a Slack channel or a town hall call when everyone shares, but when
[00:37:34] you scale and have lots of people, you need a kind of a system or a software that will
[00:37:39] help you monitor that in real time.
[00:37:41] Understand where the organization moved from experiment or what's the bucket on experiment
[00:37:48] and what's the bucket on real production value.
[00:37:51] And you would like to kind of make sure it has kind of breaks and, you know, so let's be
[00:37:58] clear what I'm talking about is experiments in how we work.
[00:38:02] Yeah.
[00:38:03] And if you think about your Cloud MD file or your cadence and the file, there's the story
[00:38:10] about the Lean Sensei that came to a manager and saw the workstation and said, you're cheating
[00:38:18] the company out of money, you're, you know, you're not doing your job.
[00:38:25] Because what they saw is that the standard work to comment, which describes the standard
[00:38:31] operating procedure for that whole process, it was printed on Fox paper for this reason.
[00:38:38] And they saw that it was yellow.
[00:38:40] You could barely see.
[00:38:42] And that indicates sort of like I would look at your Cloud MD file and see that it's from
[00:38:47] a year ago.
[00:38:49] No chance that the Cloud MD file from a year ago reflects, you know, compound engineering
[00:38:55] and continuous improvement.
[00:38:57] There's no chance that an agent that is configured with guidance from a year ago is the most effective
[00:39:05] way to do things today.
[00:39:07] Yeah.
[00:39:08] So that's something else that you can measure.
[00:39:10] And you know, that you can do with how many PRs do you have on your preferences files?
[00:39:16] Are your skills evolving?
[00:39:18] Are you creating new skills?
[00:39:19] Are people using these things?
[00:39:20] Yeah.
[00:39:22] Metal level, metal loops that I think organizations.
[00:39:27] We need to start to look at.
[00:39:30] In order to really do that, we didn't mention that there is a need also for some security
[00:39:34] tools, right?
[00:39:35] Not just pure observability like data dog, but really having the ability to define what
[00:39:40] MCPs are safe, what skills are safe, whether they're integrations that are unsafe for
[00:39:45] organization, DLP, you know, making sure the data is not leaked from highly sensitive
[00:39:50] location into your agents, your coding, et cetera.
[00:39:54] Those are things that are still being built.
[00:39:56] I see startups that are still in stealth that are building towards that.
[00:40:01] And that's great.
[00:40:02] It will allow more organizations to adopt AI.
[00:40:05] I do see that even without those guardrails and definitions are enterprise that I talk with
[00:40:10] that told me that they are starting with sandbox environments, you know, in order to allow
[00:40:15] people to experiment and have kind of safe zone without corrupted kind of outcomes.
[00:40:21] So again, still early maturity phase, we are talking about real future here.
[00:40:27] The reason that we don't see all those challenges yet is because there are not yet those security
[00:40:33] guardrails as well implemented and working seamlessly in the organization.
[00:40:38] Yeah.
[00:40:39] I mean, there are a lot of missing pieces, both on the technical stack side, as well as
[00:40:45] from a human fluency perspective, like most people in most organizations are barely making
[00:40:51] their first steps out of chat mode.
[00:40:55] Like I think it would be interesting to see now that both flawed and chat GPT integrate
[00:41:02] co-work or work mode closer to chat, I think I'm guessing.
[00:41:08] They just are doing is they're trying to drive adoption of a gentic mode by making it
[00:41:16] closer to what people are doing, and it will be very interesting to see will that move the
[00:41:20] needle and drive more people, especially out of the coding engineering organization to explore
[00:41:28] a grant interfaces more.
[00:41:31] Two things that really happen this week is that cloud desktop to co-work and match that
[00:41:36] into the chat, and also allow the people to continue conversation also on the mobile,
[00:41:43] even when your computer is off.
[00:41:47] You will not be able to have the ability to control files in your desktop when your computer
[00:41:52] is off, but it's still allowing you to do much more than before.
[00:41:57] I do see companies that are intentionally trying to shift from driving cloud within just the
[00:42:05] development side to business, and that's the most interesting thing.
[00:42:09] You don't see a lot of waste, you don't see a lot of tokens that are being burned because
[00:42:14] it's so early in maturity, but you do see some exciting experiment and some outcomes that
[00:42:21] are not being shared online.
[00:42:24] There's a lot of slope outside.
[00:42:27] You see some kind of nuggets of success in organizations that are trying to drive business
[00:42:32] people to use AI encoding, even to develop their own workflows and automations.
[00:42:39] Yeah, and I go back to your comment about build versus buy, and I think there's another
[00:42:46] variant of that, which is when you look at the people outside of product and engineering,
[00:42:54] it's not clear to them who's going to build.
[00:42:58] If you look at the typical organization, I mean, I talked to a cheese product technology
[00:43:05] officer for a very successful Israeli company a couple of months ago, and he said, my organization,
[00:43:14] like, we're on top of this shit.
[00:43:18] We do spec driven, it's a gen, we're 10x, all good, and I asked him, okay, so what's going
[00:43:24] on beyond your organization?
[00:43:28] And he said, you know what, I don't know.
[00:43:30] I don't know.
[00:43:31] And I talked to some other people elsewhere in the organization and, you know, they're still
[00:43:37] thinking about do we use glean, do we use this, do we do that, it's not clear who's going
[00:43:44] to work on making closing the books faster, integrating young and the other tools and Sales
[00:43:52] course beyond what the tools provide out of the box, which is great, which is probably
[00:43:56] a great starting point.
[00:43:58] What do you do beyond that?
[00:43:59] Who is going to tackle that?
[00:44:00] So is it go to market engineers?
[00:44:03] What does that look like?
[00:44:05] What about organizations that don't have, you know, the tech talent inside their go to market
[00:44:11] organization to do this?
[00:44:12] Don't need to hire.
[00:44:13] It is coming from the vendors, is that the right approach, is it just consultants?
[00:44:20] It's a mix.
[00:44:21] Oh, the one certainly on how that's going to evolve.
[00:44:25] Eventually it's a mix that puts lots of opportunity outside for people in advisory as well as other
[00:44:31] people that would like to be more product managers, because you see how people try, how
[00:44:37] organization try to nurture and make people ascend to become product managers in those
[00:44:43] business units.
[00:44:44] So together with a go to market engineer or business engineer, you'll see products.
[00:44:49] Are you seeing that dynamic?
[00:44:52] You do see some.
[00:44:53] It's not yet there.
[00:44:54] You do see some that are trying.
[00:44:57] Not everyone are real product managers, eventually you can say, Hey, I'll give you the title and
[00:45:02] you have maybe some knowledge about my organization, but eventually you need to have the right kind
[00:45:07] of a set of skills to do that.
[00:45:09] But you see more and more the need of product managers and engineers in other BUNIs units
[00:45:15] in order to drive that adoption, because it's not just the build mode, you don't just
[00:45:20] build those agents or build those tools for the organization, whether it's an application
[00:45:25] or you need a way to also maintain it and put it in production, you know, in business production
[00:45:30] to allow others to use.
[00:45:32] And in order to do that, you want to ask, you know, your best seller to do that and build
[00:45:38] that.
[00:45:39] You may need your advice.
[00:45:40] I would say even if it's buy mode.
[00:45:43] If it's buy mode, you want somebody with the product hat, as well as an engineering hat
[00:45:49] to think through what is it that we're buying?
[00:45:53] Does it make sense?
[00:45:54] How it fits our ecosystem?
[00:45:55] How it fits our flows?
[00:45:56] What are the outcomes that we're focusing on, closing the measurement loop that this is
[00:46:01] actually driving value rather than just we spend some money on a vendor?
[00:46:06] I mean, five.
[00:46:09] As long as it's not seat based and it's a usage based, you'll see more and more organization
[00:46:15] taking that stuff, concept of whether I should pay now or just use tokens for that.
[00:46:22] Yeah.
[00:46:23] So, I mean, I think it's an interesting dilemma.
[00:46:26] You talked about growing product managers or identifying and staffing product managers
[00:46:33] and engineers outside of the product and engineering organization, outside of IT.
[00:46:39] One of the things I'm trying to figure out is, okay, that's one option.
[00:46:43] Another option is not necessarily staffing or defining a product management role, but
[00:46:54] teaching product skills and engineering, meaning, let's call it workable, serviceable product
[00:47:01] and engineering skills to people in the business.
[00:47:04] I think there's an interesting opportunity that AI can play in that.
[00:47:09] So, could you teach now that people are going out of jet into work mode?
[00:47:17] Could you teach work mode to apply some of these techniques, to teach people to think
[00:47:22] like product managers, to think about engineering aspects, could an AI agent help people find
[00:47:30] where the value is, help you manage your consideration and investment in a way that maximizes
[00:47:38] value realization from AI.
[00:47:40] That's like getting back, getting back to context, that AI agent will need a good context and
[00:47:45] you'll need to build a rug that represents where it should look for and make sure it's
[00:47:50] updated, again, not a year old skills needed, product skills, this is fascinating.
[00:47:58] Let's keep talking about this as we each look at the market from different perspectives,
[00:48:03] but I think there's a lot of alignment on this.
[00:48:06] So thank you for staying with us, dear listener.
[00:48:11] Thank you, Tomer, for coming and see you next time on Skilling AI from Activity to Impact.
## Source boundary
These are the published show notes from the podcast feed. They are a starting point for discussion, not a verbatim record of the conversation. The transcript is machine-generated and may contain errors or unlabeled speakers. Check the audio before quoting anyone.