EPISODE 1955 [INTRODUCTION] [0:00:00] KB: It is widely reported that a gap has emerged between enterprise spending on AI and the durable value captured from that spend. Individual employees have enthusiastically adopted coding assistance and chatbots. Yet those gains do not seem to be transforming businesses at an organizational level. One of the most important questions in the tech industry today is understanding why AI is not yet delivering returns that match the investment and what separates the small number of enterprises succeeding from the many that are not. Scale AI is known for supplying the human-labeled data behind many frontier models. It now also builds AI applications and agents for large enterprises. That combination of working alongside frontier labs and inside enterprise deployments gives the company a rearview of why enterprise AI may be stalling. Emily Xue is the Head of Enterprise AI at Scale AI and previously spent over a decade at Google. In this episode, Emily joins Kevin Ball to discuss the three layers where enterprise AI breaks down, why frontier model benchmarks miss what enterprises actually need, the data foundation problem, how the most successful companies combine internal domain expertise with outside AI specialists, and more. Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc. [INTERVIEW] [0:01:58] KB: Emily, welcome to the show. [0:02:00] EX: Thank you so much for inviting me, Kevin. It's my pleasure to be here. [0:02:03] KB: Yeah, I'm really excited to dig in with you on this. Let's start with a bit of your background. So, can you kind of give us a little bit of your career to date and how you ended up at Scale? [0:02:13] EX: Yeah, I also love to share that. A lot of people were saying like my career is kind of a reverse of many people. So, currently, I'm the Head of Enterprise AI and Scale AI. It's kind of a company, a startup company, but it's not like as small as a startup anymore that people kind of know from the perspective of supplying the data to many of the frontier labs and being an active participation to the GenAI large launch model to where we are. So, it's very interesting work from both understanding the customer requirements on enterprise side as well as driving the innovations research for pushing enterprise AI to land with revenues. Before that, I actually worked for a really large company, Google, one of the hyperscalers and also the frontier labs, two-in-one. I worked there for 11 years, being very hands-on engineers. I'm also early member of Google Brain team. I was a researcher, hands-on model training publishing papers. And I was engineering lead for several of the cloud AI product, Vertex product, tuning, GenAI tuning, evaluation agent engine with my team members who really start with the concept, the idea and see the market needs for it and then put our hands down there, build up the product and bring this to the customer. It's a really fruitful experience working with the team, put by team to bring this to the market to see what the enterprise adoption for this platform products are. That's my experience at Google for about 11 years. But actually before I joined Google, I was a professor. I get my tenure from Vanderbilt University. So my research is kind of combination of applying a lot of the theoretical analysis from optimization statistical learning in understanding certain networking system, network and large scale distribute system and with the application to several of domain separate physical system, healthcare system. After I got tenure from Vanderbilt, I spent my year, sabbatical years at Google and decide to stay with the industry. That's kind of like me. Yeah. [0:04:17] KB: That's awesome. Well, I'd love to dig into a few different pieces of that, but let's start also with a quick overview of scale because you mentioned scale. When I was doing the research for this call, I initially was like, "Oh, I've heard of scale. They do data labeling." Right? But y'all do a lot more than data labeling now, don't you? What is the span of what Scale is tackling? [0:04:36] EX: Absolutely. This is a great question and segue. So people really know about Scale, and this is also like how Scale actually incorporated as a company is really to bring data to machine learning sort of labs and the machine learning companies as a data provider, starting from providing data for image, vision perception models. And later on, the market is actually expanding from the vision models and then getting to language models. This is where we're actually getting into partnership with OpenAI in very early days providing human annotations, human feedback data to make GPT where they are now. That's the core business. But meanwhile, at Scale and also I think the broader market also realized that beyond the foundation model capabilities, the lending of the value of all these models really depends on building reliable applications and agents that powers the applications. So that's actually another big chunk of business for scale. And within that sort of the application business, we have two business units depending on who are the customers we are speaking to. I'm currently leading the enterprise AI side where I'm in the enterprise business unit, where our customers are private sectors, 500 sort of enterprises. And we also have another business unit where it's public sector facing, US and global governments. But the goal is kind of like really building applications and solutions on top of leveraging the intellectual capabilities for models to serve the business needs. [0:06:12] KB: Yeah. Well, and that makes a ton of sense. I think we're all trying to figure out where are the value layers in this new AI world. And it definitely looks like it keeps moving up the stack. Well, I'd love to really dive into the enterprise side. I think there's a lot of interesting things there. And I work at a relatively small startup. The adoption challenges in a startup world are very different than they are in the enterprise world. I think you've actually written a bunch about the different failure modes of enterprise AI and things like that. So let's maybe start with that. What is different about adopting AI in a large enterprise and what are the classic failure modes? [0:06:49] EX: Yeah, absolutely. So this is very insightful question, and there's actually many kind of directions people are thinking about it, right? Maybe what we can do is I would think about to start with is it's a problem of two sides. On one side is on the machine learning AI capability side. People usually say think about from sort of AI perspective. On the other side is actually from the platform system perspective. So in order to have system solution application that truly serves the enterprise customer needs, both side has to meet with each other and be successful. And then on top of the technology layer, that is also adoption layer, which is the management side. It also touches the people side of the enterprise usage. It involves education, leadership, routing and all the stuff. So I just want to sort of maybe anchor this discussion on this people versus technology. And within technology, there's the AI part of it, there's also the system part of it. So maybe we can start with on the AI side of it a little bit. If you think about the AI side, the market the community have a lot of exposure to what foundation model frontier labs are doing, right? I come from Google. Actually, I'm part of the Gemini team. I've been going through like Gemini 1.0 all the way to 2.5. If I look from the insider length when we develop Gemini, right? If you look at it, the development team actually have very limited exposure to what enterprise needs are. And this is exactly what I feel even I'm from cloud enterprise part of business of Google. Fundamentally, every model iteration, which we call data flywheel, understanding where the model gaps are and how to contribute data, train it. The whole engineering practice is pretty much owned by the engineering team who actually do not have the business exposure. And that reflected as - right? People are very benchmark driven. They have very focused goal. Say, "Hey, we actually have different capability pillars; math, reasoning, code, conversational multi-turn safety." And then if you look at that, the enterprise part is becoming a very narrow pillar in all the frontier labs leaderboard progressively into 2025, which means the model knows a lot of the human intellectual capability and knowledge. Many of them is reflected by our flagship benchmark humanity last exam. But in terms of concretely what is the professional requirement? What is the operational workflow? What are the policies that is actually super critical in the enterprise domain? There's no evaluation criteria. There's no intuition knowledge within the people who are training the model. Then from AI perspective, the fundamental capability has a gap from what - [0:09:58] KB: Well, it sounds like not only a capability gap but even a goal or target gap. [0:10:03] EX: Exactly. [0:10:04] KB: It's not clear where it needs to evolve. [0:10:07] EX: Absolutely. Spot on. So we need to set up the North Star. And the enterprise, it's a single word, but it's not as simple as math and code. Math and code are the true kind of - we have a lot of breakthrough in terms of intellectual capability advancement in the past. But if you think about math and code are the things that majority of the people who hands on the code truly understand, right? [0:10:36] KB: There are also domains in which it's, if not trivial, much easier to define what correct looks like. So you can build an optimization function and have an iteration loop and things like that. [0:10:47] EX: Absolutely right. This is where the reinforced learning with verifiable reward is actually bringing a lot of capability advance in I think last year. Let's make it - the time flies. Yeah. [0:11:00] KB: I know, right? A year these days is like five years. [0:11:04] EX: Exactly. I think we're starting from the early of 2025 where we're thinking about reasoning models. So now I'm thinking about this like one year and a half. Yeah. And you're absolutely right, having a precise definition on what is correct, what is right actually bring a lot of benefit in the post-training process are. But then in the enterprise domain and back to what we're discussing about, if you think about enterprise, it's a single word. It's even hard to define what it is. There are so many verticals, right? And then let's say healthcare. And within healthcare, think about there's primary providers, medical centers. There's insurance companies. There's pharmaceutical life science companies. Even with health provider, there's big academic medical centers. There's community healthcare service. They have very diverse needs regarding what they going to use AI for and what they regard as acceptable AI service are. That's kind of like the complexity of the enterprise world. And the complexity is not very well captured and presented to the labs developers. That's one of the root cause for where the capability gaps are. [0:12:20] KB: Yeah, that makes a lot of sense. So tech layer, there's some gaps. Absolutely. But let's not leave it there. Let's look at the other layers too. [0:12:28] EX: The other layers. Yeah. So that's kind of like starting from like the foundation model. But then we also talk about the system side, right? So not just the models. When you develop application, it doesn't matter big or small solutions or reusable applications agents. You need to have it to be integrated with the enterprise context environment. You need to understand where datas are, how to get yourself authenticated. When we say yourself, it's the agent themselves, to get authenticated to the right data. And also, to understand for all the enterprise services, there are APIs, there are data schemas behind it. What do they mean, right? With different table names, with different value names, how do we actually understand the context? A lot of cases people talks about context engineering. How we actually learn those tribe knowledge within the enterprise and bring that knowledge to be part of the context into the agent? And that is actually a non-trivial problem, right? Because it's not like all the information is right there presented to you. [0:13:39] KB: No. Absolutely. I'm dealing with exactly this in a startup environment where we have a handfuls of systems and a few different databases. You look at at an enterprise that's been around for decades and has hundreds of thousands of people, this is a mammoth problem. [0:13:54] EX: Totally. And then you hit another point in terms of the pain point which we is - when we think about enterprise and then we talk about this is a new era of AI waves. But then let's look at in the past two decades, when people are talking about digitalization. Everything from paper to digital form, digital asset. There's lot of technical depth actually in that wave. Because during that wave, the entity who actually are the information receivers are human. Humans are extremely flex accessible and they are able to tolerant. Give me a PDF, give me a handwritten notes. It's not like perfectly digitized in a way that is to the high standard. But I can still get it. This is actually the problem we get, for example, from the engagement with Mayo Clinic, right? We think right now all the electronic medical record system are digitized, right? It conforms a uniform standard called FHIR. But the reality is there are always like patient who are not actually within a network. They go through another system. They come to you with very complicated medical history. They're not going to give you the standardized, normalized electronic medical records. They will give you hundreds of pages of their historical medical records printed in PDF. People can read it. But then in order to actually ingest the data from different modality with the technical depth from the last two decades fragmented received processed by human being in the past. Now you have to plug in an agent that is powered by foundation model, large knowledge models. Then the gap is actually emerge from there. [0:15:45] KB: Yeah. Yeah. Absolutely. Every shortcoming that we all have that a human has just said, "Oh, you know, that's not perfect, but I'll paper over it." Now, we have to go back and revisit and say, "Well, can an agent paper over it? Maybe not. What do we do with that?" [0:15:59] EX: Absolutely. That's kind of like one of the things we experience, how we actually integrate an application, an agent into enterprise environment, really address all these gaps from fragmented information, tribe knowledge. All the things become the barrier, actually prevent accelerated deployment of AI solutions in enterprise. [0:16:22] KB: Well, and this kind of then ties into this last piece, which is the human layer. What needs to adapt in terms of how we're engaging in work? [0:16:31] EX: Absolutely. This is like at the lower layer of the technical issues, but the solution to the technical problems really lies in the human's hands. And this is where, from organization perspective, the leadership commitment, the leadership, because all these motions, all these AI motions, it's not going to be successful if there's not a clear incentive on what is the benefit, what is the ROI. If the leadership does not recognize the value of the ROI, and then when we actually get into all these barriers, then people could easily sort of - will not have a clear actually success criteria. People could easily pause at impressive demo and not pushing forward because there's no incentive to show the fundamental return and the impact, economical impact to the solution, right? It's just drop it in the middle of nowhere and then not get into productionization. From this perspective, the most important thing is the alignment at a high level on a leadership level. We're aligned on a North Star successful metric. What is the expected output of it? And then driven by that, we would have a clear technical solutions, technical roadmap to say now in order to achieve this high level, North Star success metric, what are the things we need to do? What is the path? What's the low-hanging fruit we need to do in the first step to get into the system foundation to be developed? And what's the next step to get to the user engagement? How we actually collect the feedback continuously from the usage to improve the system from the agents? Those are the things that actually defining the driving force to be successful. [0:18:32] KB: Yeah. Having leadership agreeing on where to go definitely seems important. Another thing that I'm pondering about, one of the things that I've seen and I've heard a lot of people struggling with is - least in the coding space, which is the one that I and I think much of our audience are most familiar with, these tools are really focused on individual productivity. [0:18:51] EX: Yes. [0:18:52] KB: But if you just drop a bunch of individual productivity improvements into an organization without rethinking about your organization, you often end up just bottlenecked on the next layer that didn't speed up the same amount. And you don't see nearly the level of organizational improvement that you expect based on the individual. How is that showing up in these large enterprises in terms of change management, or reorganization, or rethinking of how we organize ourselves? [0:19:18] EX: Yeah, this is such a great point. It just kind of really resonates what we see in the field as well. If you go to a lot of the enterprises and you ask them the AI adoptions, the answer could really go two directions depending on how people think about AI adoptions. If we say the people who are using Claude code, GPT, and you're using all these individual AI tools, your adoption is amazingly high. Everybody is using it, right? But they're not using it in the context of like exactly what you captured in contributing in a coordinated enterprise workflow. They just use it to improve individual productivities. And then to do that, there's actually two major problems, kind of very related. The first one is there could be a lot of IP leakage. Because when you interact with an individual tool, all the queries you have, all the questions you have, there are a lot of enterprise confidential information. And even if it's not proprietary data, there's also the way you ask questions and then ask follow-up questions, review a lot of the judgment and the implicit knowledge within people's mind. And those information is actually leaked into the frontier labs. So that's one side of the problem. But on the other side of it is actually these are very valuable enterprise digital assets, the IPs. And when people are actually getting used to just interacting with the individual productivity tools, all this digital asset, all this implicit knowledge are actually not captured and shared across the organization. This is exactly the point you're talking about, right? In many cases, when you work on a project and then when you have experience in the context, the skills, all this expert knowledge, from an enterprise perspective, this information needs to be consolidated across multiple team members and then into a coherent knowledge base. So that the next generation of the people who are working on the project will be inheriting the knowledge, the experience being accumulated from the previous conversation. Those are the things that are actually missing by building this. [0:21:46] KB: So that kind of brings us to the flip side. We've talked about where this breaks down, where the gaps are in different ways. What does success look like? I know you all recently published this report and you said, I think, it was 6% of large enterprises are really right now succeeding at integrating AI and seeing the benefits. So what are the patterns? What are the ways that they get around these gaps? How does that end up looking? [0:22:11] EX: Yeah. So when we have the 6% sort of the reports, the first thing we think about is how we define success, right? In this domain, defining success is actually not easy as well. Because as we said, if you ask them about AI adoptions, they could say some shallow adoptions, and that's not success. The first thing what we did is we really think about what the success criteria are. And then for that, we basically have a pretty rigorous way to say that. It's kind of like how people among all the projects people have been engaged in, how many of them actually go from pilot into production, and then meaningfully integrated with the workflow. And by meaningfully integrated with the current enterprise workflow, that is measured by very concrete usage metrics. So those are kind of like the things we're looking at how we think about the success in a very concise way. We actually look at the characteristics to see what are the characteristics the 6% winners actually share. We are not able to attribute to say, "Hey, these are the factors that are causal factors. If you did this, you'll be successful." We're kind of like doing machine learning. We try to capture what are the shared attributes about all this 6% - [0:23:32] KB: I mean, have you ever been able to make a true causal argument in a human system? That's very, very hard. But yeah, I think that is useful, right? The pattern matching is what we're all looking for. It's like, okay, what are the patterns that I can try this? And maybe that will help move my organization closer. [0:23:48] EX: Yes, that's exactly the sentiment in what we're doing. So there are a lot of findings in the reports, but the three things kind of really strongly reflected from all the statistics and the study we have. The first one actually is not surprising, is the data foundation. And this is actually we've talked about how data is fragmented, multimodal data, connecting data actually to understand the data access policy. What are the governance? How we actually collect the feedbacks from the field? All this data infrastructure foundation is actually a prior, kind of like very strongly related with the success of AI search deployment success. No surprise, right? [0:24:32] KB: It's not surprising, but I feel like this is a thing that comes up. And there's memes that have gone around of like, "Oh, we have this huge data mess. We could solve that. Or we could just sprinkle AI on top and hopefully it will work." And it all comes back to, "No, you actually have to solve your data mess." [0:24:46] EX: It's a chicken-egg problem, right? If you don't have the data, then AI will not do anything. But then when AI is actually getting to the enterprise in the right way, then AI could be helping you with, for example, one of the key technologies we're actually providing to the market is entity resolution. So you have different tables. The keys, the way you name corporate entities are different in different tables. In the past, this is a very traditional machine learning problem that has been there for a long time. But now with the large language model, with a lot of priorly trained word knowledge, they can do this job really better. So now if you bring AI to solve the problem, you have a much better normalized, standardized, linked data asset. Then the next layer of AI agent is able to build on top of it and then discover and unleash more capabilities of solutions on top of it. So it's an iterative process, but you need to have the foundation together first. That's one of the first things we noticed. The second one, as we also chatted about, is the human part of it. Usually, if there's early stage investment in terms of senior leadership sponsorship, and then they are committed in change of management, they provide resources in employee engagement training. Because personal usage versus enterprise workflow AI, a lot of cases have different steps, different UIs. If they actually would get into the co-developed product, co-developed motion, and then envision when AI is co-pilot, work with them in a workflow environment, then the success of adoption is actually much increased versus initially we don't actually imagine what a change of management look like when the AI agency is being deployed. If the enterprise front load or the organization work, that is actually something we noticed among the 6% winners. [0:26:51] KB: No, I think that's really important because I talk to a lot of people in a lot of different environments here and it feels like one of the most common failure modes I hear about is essentially treating this as a technology and sort of magic adoption thing. Everybody must use AI. It's a magic technology. It will make everything better. Just do it. You figure it out. Versus the successful mode of like, "Hey, this is something new and tricky, and we're going to work together and have change management and co-develop and really refine how we're working with this." [0:27:24] EX: Absolutely. And this is actually a great segue to the third point we actually discovered, which is like on the market right now, if you talk to enterprise customers, there's always these questions about buy versus build. And many cases, if they ask you, "Do you have a product that we can buy immediately?" And on the other side, if you get deeper with them, they're thinking, "Oh, we have a pretty strong engineering team. Can we build it together?" Build internally. But the answer to the question is actually not that simple, as you pointed out. It's really we need to bring the expertise from both sides. If you just get a product, even with a customization consulting team to get to the enterprise side, we still face the problem how to integrate the product with the workflow, with the enterprise context. And then how the product is actually going to be used? What's the CUJ look like? What are the touch points for customization? That's kind of like one. On the other side, if they think, "Hey, we have the internal team, let's just build it. It's really going to fit very closely to our internal usage pattern, the workflow pattern." Then what is missing is the internal team. Of course, the URL is very strong in your team with many years of seasoned experience. They usually don't have the broader understanding about AI technologies. That actually, one of the benefits when we work very closely with the frontier labs, we actually have a best practice in terms of how we do evaluation of the foundation models. What are the best patterns in developing agents around on top of these models? And what are the typical failure modes as well? When we actually see the 6% winners, they strategically combine the internal expertise, which is that domain knowledge in terms of when this is useful, how to use it, and what outcome is actually going to change the workflow and decisions with the specialized partners who actually bring in the expertise regarding AI models, AI application development, and system development. When the expertise from both sides are combined together, this actually creates the motion and the foundation to be successful. [0:29:52] KB: Yeah. So one of the things you mentioned there that I'd love to dig into, because I think you have done a fair amount of work in this space, is around evaluation, right? And I think one of the things that I definitely see hiring engineers and looking for folks to work on AI products and things around this is one of the biggest gaps is people who can grapple with non-determinism and those who struggle with it. And the fact that in a non-deterministic environment, you can't do unit tests. You have to have some other way of validating quality and understanding what's within your sort of trained scope, or the training corpus, or the nice fit of the model, and what isn't and all that sort of thing. How do you all approach evaluation? And particularly, I think there's a lot of people thinking about evals at the prompt level. How do you think about evals at an agentic or workflow level? [0:30:38] EX: This is my favorite topic, and then it's a list of questions that worth the time to dive in a bit. So in a very simplistic term, if I think about evaluation, there are two things about evaluation. The first one is what needs to be evaluated, right? And then as you point out, in the foundation model development, what we evaluate is the model. Whether the model is going to work, right? And then that's the subject to build. And then in order to evaluate, there are two things surrounding it we need to do. One is how the model behaves and the expected input. This is actually super important. When we do model development, you test it under a very broad spectrum of queries. And the queries, of course, will be sliced into different pillars, some are coding, some are conversation, and some are adversarial for risk assessment. And then the last part of it is when the output, doesn't matter from the model agent, is produced, how are we actually going to tell is it good or is it actually less desired? And in the indeterministic case, as people recognize, there's another aspect of the robustness. Can you actually reliably, robustly give the consistent output when either you repeat your task or you have small perturbation in the input. These are kind of the three ingredients of it in the context of agent evaluation. The first one is actually the agent self, also the environment you're actually operating in, right? So it's not just what you develop the agent intelligence self and build on top of the model. It is also what is the backend environment? What are the tools that this agency is connecting to? What are the service APIs that the agent has access to? And more importantly, what is the backend state? How many databases you are actually connected to? What are the file systems? What are the knowledge base that enterprise? This actually defines the environment during evaluation, right? Because if you change that, the evaluation result could change, and that's the robustness question we need to get into. People need to understand what are the environment? What is the entity we are evaluating? And then the second part of it, as I said, it's kind of software testing that you need to define your test cases to cover the happy path, the typical, the edge cases. So we need to strategically define what is the set of tasks that cover different dimensions of the capability and the functionality of the expected behavior of your agent, right? Just to give you an example, let's say we are doing patient safety triage agent, right? And then the goal of the agency is going to look at a report to say, "Is this an event I should report?" And then what is the policy term that I need to ground my decision to? So this is the agent task. And now when you get into the evaluation part, the high-level mindset is actually very similar to what software unit testing is doing. Then you need to think about how do you design your unit test to cover different aspects of it, right? Then you need to think about there's the positive case. The event is true. But then you also need to cover different policies, like underwater scenario. Because if you look at the event safety guidance, there could be 1000. Well, that's a little bit too high. 100 different criterias. Then you need to make sure your query actually hit the different categories of the events so that you don't have some missing edge cases where things you don't have the confidence to speak about. And of course, you also need to think about a negative case and a hard-negative case. You need to think about what if there's ambiguity where you're missing evidence. What if there are terms that is not very well specified in the guidance? So you need to think about all these things, cases, and those are the evaluation data set curation. [0:35:05] KB: I totally agree. I want to dig in a little bit more about the environmental piece, because I think looking at building out these types of eval frameworks in different environments, one of the biggest challenges that I have encountered is how do you get a realistic, but also reproducible data environment, right? You don't necessarily want to run this against your prod environment or maybe you do, right? But what does that whole setup look like? [0:35:30] EX: I think this is still open question, to be very honest. And this is actually a research direction we're working very hard. And I'll actually have some result there to be shared with everybody coming up. And hopefully, we'll talk to you again when it's out. But you're actually asking a real hard question. The problem right now is if you look at across the industry at this moment, There's very limited in a conservative, but it's not known, evaluation environment that is backed up by real data. Because the real data from enterprise, it involves a lot of the privacy, legal usage concerns. But if you don't use real data, then all the dependencies, the very complex dependencies within the data will not be truthfully reflected in the environment you are developing, right? [0:36:22] KB: Yes, exactly. This is exactly the problem I feel like we're all grappling with right now. [0:36:26] EX: We're all grappling with. Yeah. The solution I can share a little bit of like how people approach it, you start data vendors. And we're one of them, right? So we're actually looking at data not actively being used but similar in nature. Similar semantic domain, usage domain. But from the past, expired data. And we do DID. We do sort of anonymization over the data so that the sensitive information will not be leaked, but the semantic relationship will still be there. There are lots of research in that domain. People are looking for proxies, basically, surrogate and proxies to build a similar environment. But I agree with you, the sim-to-real is always a gap. Simulation to real environment is always a gap in the machine learning research deployment challenges. Yeah. [0:37:14] KB: I am curious actually. You mentioned this is an area that you all are providing. What are the services that Scale is providing here in the enterprise space? Because we've talked about some of the needs and the gaps, and you've sort of alluded to, we do co-development, we're working with people, you provide some of these simulated data environments. What actually does, if I was reading it correctly, the Gen AI platform or the enterprise platform Scale provide for folks? [0:37:41] EX: Yeah, absolutely. So at very, very high level, we do three things, right? We provide platform, which we call skill generative platform. That's the foundation allowing all the system integration work, all the trustworthy agent execution, and then allowing the agent execution. All the traces to be logged and then for evaluation for auditing purposes that the platform we're providing, which provide the system infrastructure foundation to build a reliable AI solutions. That's one pillar. Another pillar is actually our full deployment motion. This is like we're not just giving you a platform. What we do is we actually have a team. It's a very well-constructed team which have strategists. Purely people coming from management consulting like McKinsey, BCG, they would work with the customer to understand their business north metric and then look at the use cases that allows us to produce the impact for outcome to meet that North Star metrics. And meanwhile, we have the technical people who works very closely in the initial scoping process to say, "What is technically feasible within the time frame we have?" And this we co-define, we discover the use case, co-define the use case to be worked on. Then we get into the delivery motion where for deployment people, the people with applied and as well as system background work together to really wire all the fibers, code all the AI solutions, perform the evaluations, understand the worst gaps are, and hill climb on the quality. So that's kind of the full deployment service motion we provide to enterprises. [0:39:24] KB: Yeah, it's really interesting because I feel like this is a pattern that is starting to - if you look at our industry as a whole, we're evolving from essentially deploying software that was relatively repetitive, but in the enterprises were integrating then, trying to find ways to integrate into their processes. And we're moving to a world where we're deploying a set of capabilities, which has software, but really then is the details are implemented either by the company that it's sold to or by forward deployed engineers. That's probably the most rapid growing software field right now is forward-deployed engineers. [0:40:01] EX: Exactly. And then actually just on top of what you're saying, here's another dimension. When AI capability is being deployed, it's fundamental, it's a learnable system, right? So it's not just like a one-shot thing. What's good for deployment engineering team, like Scale AI, we're not just saying, "Here's the solution." We actually think more long-term as we previously alluded to, how we actually help you to collect all the digital asset that reflected from the usage pattern through the traces, how it would actually help you to identify the quality gaps when we talk about the evaluations, where things fall, and then provide you the research innovation, which was the last piece I was going to talk about, to improve it progressively over time so that the agent becomes smarter over the time of people using it. Because it would capture human knowledge, human skills and then use it in the future sort of. It would incorporate into the memory of the agents and then help its usage moving forward. [0:41:11] KB: That's interesting can we talk more about that? Because I think coming back to our co-development, there's a co-evolution that needs to happen there as well, right? Because the agents become more capable. But then what people are trying to do with them evolves and kind of what that looks like. How do you think about setting that up? [0:41:26] EX: Yes, I think that's actually a very essential piece when we get agent into the enterprise environment, because it's a living thing. And don't think like in the previous era, when we say service, a source kind of era, you connect it, but what you get are just data into it. But then once the agent, agents actually serve as the copilot that works with you, then agent would actually learn from like we first talked about, it will learn your institutional best practices as your institutional knowledge base. In a collaborative environment, let's say we have a group of people who are working on a project in three months. Then when we actually work together on the same agentic system, then the agent is able to actually accumulate our usage. The questions we had, working solutions that failed, would accumulate this project-oriented information. And then over time, it would make our project execution a lot more productive. And meanwhile, on an individual basis, you could also remember what you did yesterday, what is your usage pattern, and so on and so forth. So the agent would have components built-in to learn from the interactions and become smarter and interact with people on the project's execution, workflow execution in a smoother way is actually one of the differentiations that it would bring to the customer enterprise. [0:42:55] KB: Fascinating. So how do you think about the evolution of human in the loop here? And what can get fully absorbed and automated versus where you have people checking in versus where you have this very co-pilot-esque approach? [0:43:12] EX: Absolutely. I think there are two dimensions of things. The first one is when AI gets into the enterprise, it plays a very active role from any aspect of it. But one thing is, who's going to be liable for the output, for the decision, right? Someone needs to be hold. responsible for it. You can't say this is the AI data. But in my humble opinion, I think the human is going to play a long-lasting role in being responsible for the AI decisions and what the work from the AI is. Technically, how that's going to happen, I'll get to that in a minute. But in the long run, human stays. Human will be liable for it. What AI does would be helping AI in making decisions in various ways. But towards the end of the day, we need to have the people who are responsible for the AI decisions in the enterprise environment. That's kind of like a very long-term, sort of like the foundation of human, stay human in the loop. But with that being said, there's also a tactical direction. You know what human does, right? So I believe that part is actually what we see is going to have a path towards initially when the trust between - when we talk about evaluation, right? There's offline evaluation, there's also online evaluation. Initially, before an agent is being used in the production environment for a long time to gain sufficient trust, there's always going to be a human in the loop component where a human will be triage, ambiguous cases, human will be validating the high-stake operations, where the system state got permitted. So right now, a lot of the writing, almost in all the applications, if you see something's been written, and then there's always asked you to confirm, right? So all this risky high-stake operations, human needs to confirm that. And then even for the insight derived by the AI agents, human needs to review it to make sure it really aligns with the business policy and the business standard. Right. But over the time, when the trust is built and then the products is matured in a way that we are going to significantly reduce the cognitive load of human being, to a point, some of the very confident answers, confident operations from the agents should be just automated. But that is a tactical roadmap to initially we have a higher human oversight to pay to get reliable, trustworthy deployment. But this cost is going to reduce gradually when the mutual trust is established. [0:46:10] KB: Yeah. Well, and that fits with what we were talking about in terms of our lack of a truly high fidelity eval environment is you got to put it into the actual production environment and see how it behaves. And while you're doing that sort of live action evaluation, you don't want it changing key data or making big decisions. I think there's something interesting on the sort of long-term vision there. So you highlighted liability and responsibility, which I think is true. That's an important piece, right? I think it's really important that whatever. And we're already seeing this in the coding world, right? People will be like, "Oh, the agent made this decision." Well, who said it was okay? Who committed it? You still made a decision there, even if your decision was to delegate your thinking. But I feel like in that framing, it feels very burdensome, right? It's like the agent's doing the work, but you have the liability and the responsibility. What's the positive side of that? What we're going to be doing as we automate more and more in terms of like this direction, decision-making, exploration. I think there's something on that side of the world as well. [0:47:14] EX: Yeah, absolutely. To your point, I feel like people are still trying to figure out what is the cognitive interface between AI and human, right? Fundamentally, AI is actually expanding human's intelligence and capability, which leads to higher productivity, as well as higher quality decisions, right? But then to your point is, where? What is the interface? Where it actually helps? Instead of asking me all these trivial questions right now, like everybody's in Claude Code, I'm like, yes, confirm, yes, confirm. After a while, I got fatigue, right? And then not going to be making the right decisions anymore. This is actually also from a research perspective, actually people are actively looking into in a sense that when the AI knows what is the point where we truly need human input, right? When he's sure about what he's doing, when the AI is not sure about what he's doing. And that kind of like calibration capability is something that people start to look very much deeper into and it's still kind of like open research at this moment. [0:48:32] KB: Yeah. There's something interesting around that in the - I remember digging in at some point. Large language models, your default interface doesn't actually expose the confidence at all. [0:48:42] EX: Confidence. Exactly. [0:48:44] KB: But you can, especially if you're doing an open model, right? You can look at what was the distribution of the logits on there and see which one - this was very sure versus this was actually a pretty even spread could have ended up in a bunch of places. So there's something fascinating at exposing that, you know? [0:49:01] EX: Exactly. There are research about like how the logics, the logics at token level. And then now you think about like a long sequence, right? [0:49:09] KB: Right. Yeah. What does this look like at an agent level, or a task level, or something? Yeah, yeah, yeah. [0:49:13] EX: Exactly. And is it calibrated, right? If the agent's at 60% confident and that actually empirically calibrate with the statistics on that. This is open research at this moment. And then people are still thinking of like talking about sometimes the agents are overconfident on many things. Yeah, that's actually a really fascinating area. People should really spend more time looking to one, especially when we're getting deeper into productionization and the human agent collaboration. [0:49:43] KB: Yeah, yeah. I've been thinking about this a lot with regards to some of these entity extraction, and context extraction and things like that. Can we add a confidence layer, right? Okay, we are 80% certain this is all one entity, but need some more evidence. How could we validate?" Things like that. [0:50:00] EX: Yes, yes, completely. [0:50:03] KB: Awesome. Well, so we're getting close to the end of our time together. Is there anything we have not talked about yet that you think would be useful or important for us to discuss? [0:50:12] EX: Yeah, this is a really fruitful and productive conversation. Really enjoyed it. Yeah, I think we covered everything. Yeah. What do you think? [0:50:21] KB: Yeah, I mean, I think we covered a lot of ground. [0:50:23] EX: A lot of many, many ground. Yeah. [0:50:25] KB: Awesome. Well, I think let's maybe do one thing back. If you were to look, you just did all this research looking at the different failure modes inside of the enterprise and what these 6% companies are doing really well. What would you say is if someone was working inside of one of those enterprises, but was not maybe in a leadership role, or maybe they are in a leadership role, but it's a software engineer listening to this, or a line manager, or something like that, what's the leverage point you would push them towards to increase their chance of success in terms of enterprise adoption? [0:50:57] EX: Yeah. If there's a single point, it actually, I think, is boiled into get the right pilot use case. And there is a lot of things within the right, right? The right thing needs to be truly valuable. People have the incentive to make it work. It's not a toy use case, right? And then with that, then you could actually organizing. Even like a software engineer, they could convince their leadership because it aligns with their business metric, right? So there's a lot of things within the right part. The second part is the right use case needs to be ready in a sense that you actually have the right data, you have the right resources and asset to make it correct. And the third one on the right part is actually you can think about whether the technology and expertise within the team outside actually would enable this to successfully execute it and land in the room. So it is basically find the sweet spot to start with and the success would compound from the first successful pilot to production launch. [END]