He looked at me in the eye and said, I'm not going to implement AI because I'm afraid of the outcome. We are in a regulated market. If something happens, we cannot control race.
SPEAKER_00So who is the person in the organization on the client side or the internal stakeholder that is most likely the person that will understand the concept, understand the requirements of it, and be able to fulfill on the client side?
SPEAKER_01The group that develops and implements AI solution. Don't feel shy about the output of AI because it's good for everyone.
SPEAKER_00Hernan Lardies is a seasoned tech leader with 25 years in global sales and operations, a COO of Ragmetrics. He helps companies reduce risk and boost trust in generative AI systems. Welcome to Using AI at work. I'm your host, Chris Dagle. Each week we'll be learning how today's business owners, entrepreneurs, and ambitious professionals are getting more done with smart use of tomorrow's tech. Let's get started. Right now, every business leader is asking the same question. What are we going to do about AI? If this is you, ChiefAIOfficer.com has the answer. We give you a simple path forward where we provide executive and team training so your people know exactly how to safely use generative AI in their day-to-day. We also manage the deployment and implementation to make sure tools actually get adopted and deliver results. And we'll also guide company-wide transformation so AI becomes part of your operating system, not just another shiny object. The companies that act now will increase productivity, cut costs, and grow faster than their competitors. Those that wait will get left behind. So if you want to make AI work in your business, visit chiefaiofficer.com and see how we're helping companies of all sizes finally get results from AI. All right, greetings, esteemed listeners. This is Chris. I'm the host of Using AI at Work, where we explore generative AI's application and participation in the work day through a non-technical lens. Now, the topic that we'll be covering today we could get very technical in the weeds, but I'm going to make sure that the conversation sticks to how can I, as an executive, um first off, understand the application of and get benefit from uh model evaluations, the testing of the results that we're getting from AI that we're putting into production. And our guest today is a uh fantastic uh thought leader on that topic. Hernan is a uh he's been working with AI since 1991, was his first exposure as a software development engineer, has worked with a lot of the known brands, multinational brands that you'd know of. But we were recently introduced through a post that he and his partner had put in uh Rachel Wood's group called the AI Exchange. If you're not familiar with that, uh definitely worth checking out. And I know uh Rachel's a friend, but she comes from kind of a geekier background when it comes to AI. And in all of the training that she does, she always stresses evals, evals, evals. Well, as somebody who didn't come from a technical background or a development background where that was important, that was kind of a um a step maybe that I I skipped or didn't pay a lot of attention to until recently. Now we find ourselves at Chief AI officer doing a lot of work with uh clients now, where it's not if there's a mistake, it's not internal to our company. If there's a mistake, it has it has impact, it has uh gravitas on a client company, which we definitely don't want to do. So when I had the opportunity to connect with Hernan and his uh business partner on this topic so that I could understand it, I was excited to share that information with you. And this interview will be a continuation of a basic conversation that Hernan and I had maybe a couple of weeks ago. So with that, Hernan, um, welcome to the show. And uh anything that I left out maybe that you want uh you think people should understand?
SPEAKER_01No, first of all, uh thank you for having me, Chris. A pleasure talking to you today. You define a little bit what what I what I've been doing. So my first interaction with AI was in 1991. At the time was purely academic because there was no processing power. The main difference between now and then is the processing power that we have today. So for sure, in the middle of my career, I've been doing other things, but in the last few years came back to the AI world.
SPEAKER_00Nice. And it's a new habit that I've started at the beginning of the shows. I'm borrowing from Greg Eisenberg and his show. But um I think it's it's a good question to ask. By the end of this episode, considering our audience and the preamble that I gave, what would you want the listeners to walk away with?
SPEAKER_01Why having a measure to control or to look at the output of AI, it's important. And we for sure will get into the details of that. But why why do we need to look at what we are getting from AI, analyze it, and have some kind of statistics behind it?
SPEAKER_00I think it's a fantastic target. Okay, so let's jump into it. Um, Hernandez, we had a brief conversation where I was uh what started our initial like dialogue was that I was interested in exploring your services to bring that into what we're doing for clients so that our chief AI officers had some tools and frameworks to be able to um evaluate, hey, is the solution that we're building for this client tenable? Is it is it stable? Is it going to give them a high degree of accuracy when it comes to the output of that intelligence? And that really um I think that's a great place for us to start. So can you kind of explain to the audience what what does evalu mean in the context of using AI?
SPEAKER_01AI is based on a pro a probabilistic feature. So in order to have an outcome, there's some probabilistic stochastic process uh that is embedded. So that drives to probabilistic outcomes. So as I always say, if you play the law ready enough times, at some point in time you'll win it.
SPEAKER_00I think this is good. Before we go on, explain the difference in your um uh position on uh deterministic versus probabilistic probabilistic, yes.
SPEAKER_01Yeah, and in the old days of in the old days, actually, up to a few years ago, that's the old days of computers, computers programs were deterministic. So there are a series of uh uh algorithms that produce a result. So if you put uh one plus two, you will always get uh you will always get three. In in AI, it's it's probabilistic, so the outcome is not necessarily always the same one, and that's basically the fun of AI. So that also drives that sometimes the outcome is not correct, or it's what we call it's hallucinated. And the nature of AI is that hallucination cannot be avoided, errors cannot be avoided, and that's basically the core of what we want to resolve with what we do at Rugmetrics. Try to understand what is that output of AI and how to measure and evaluate it.
SPEAKER_00Okay, so um as a listener, um I'm sure that you have used you got maybe a few prompts that you use regularly. You pick them up at maybe a meeting or you develop them yourselves, and you've noticed that I use the same prompt, but the output is different, is maybe not wildly different every time, but it's not the exact same every time. And that is a key concept as we continue this conversation. If I have an activity or a workflow or a process in my business that I want to accelerate the result by using AI, I need to be prepared for the fact that if it's a very precise activity, if I don't have uh frameworks in place, I can't always count on the result being meeting that that uh tolerance of precision required for that workflow output or whatever. Hope this is making sense. But it's very important because outside of if you just want to use AI for thought leadership and ideation and working on strategy, that's okay. But if you want to operationalize AI and make sure that it's something that becomes part of your workflows and processes that you can count on and you can say, hey, you know what? We know we can now scale this activity in our business by leveraging AI and we're gonna get those consistent results, then this effort on evaluation is something that's extremely important. So I just I want you guys to know that this is not just some theoretical topic. This is a key piece of being able to operationalize generative AI at scale. Okay. So, Hernan, thank you for that definition. What is the experience that you're having with clients? Who are the type of clients that are looking for, that are even savvy enough to know that they need this as part of their process and that are looking for help solutions like what you guys provide?
SPEAKER_01Yeah, the market is continuously evolving. So we are on the phase where more and more companies and enterprises are starting to implement AI from large enterprises to small uh uh SMBs starting to implement some kind of AI. So anyone that has any risk with customers, or customers could be external or internal, any risk with customers should at least using AI should have an option to control to control that. From simple chatbots uh implementations that give you uh um directions or for for a product, how to set it up, to fintech companies that provide information about financials to healthcare companies, to anything in regulated markets. So everyone that has an AI tool that requires, as you defined it before, a very clear clear precision requires control on that.
SPEAKER_00Yep. Okay. So for those of you listening, some of the activities in your business as you're introducing AI, you may be fine just saying, hey, for this role, I'm gonna give you a license to chat GPT, some basic governance guardrails, a little bit of training, and that's all you may need in this particular role. For other roles, you know that boy, if we could have the bot, the agent, the app that could support this role, that individual in that role could do 10x or 100x productivity. But when you are outsourcing or relying on the models to have that much of an impact on that process or workflow or job in your business, the reliance on consistency is paramount. There's there's no two ways around it, or else you will scale it, it'll break, and then you'll go back to the drawing board and you'll say, hey, this all this AI stuff, it's too hard. Let's just keep with prompting. So Hernan, for um businesses who are like this is a new concept to them, how do they how would you suggest they start to identify those activities where, no question, we need to have an evaluation process in place before we put this into production?
SPEAKER_01Yeah, so let's let's let's recap a little bit. So the evaluation process goes into two areas, if you want. One is to ensure that the the the output is is correct in terms of of what it general correctness means, is not hallucinating, the answer is more or less correct. The other one is what is what you really want to measure? So different businesses have different needs, and those needs rely on the quality specific of what they are trying to implement. So, as an example, a retail entity, a retail chatbot, might be focusing or not on discounts, and users might uh might want to measure the quality of those discounts. So it's it's not only just providing a general answer with that is correct in form and and in context, but also looking at the details behind it. The process to generate a good AI product and to monitor your a good AI product starts on the development phase. The development phase should include testing and optimization. So without technique getting too technical, if you have a RAC system, there are different ways to get information out of the RAC system, different policies. So the optimization of those policies can help to reduce cost, can help to get better throughput, and can help to get better accuracy on the answers. So it is important when the AI solution is being developed to have a good uh um evaluation process or tooling to evaluate during there. Then the the other the other uh uh portion comes is when the the AI solution is it's alive, it's up and running. So we heard in the in the uh in the market uh concepts about hallucination, concepts about drift. The objective is how to detect those. How to detect if something uh is drifting, an AI solution is drifting. So uh one way to do it is to continuously monitor the output of that AI tool in time. And if you're measuring accuracy and you start to see that the accuracy slowly starts to degrade, then you can see that you're starting starting to have a problem. So accuracy, let's say in a scalp for in a scale from one to five might be uh I don't know, 4.9, good, 4.9. Next week 4.9, and then start to come 4.8, 4.75, 4.73. Then you start to see that it's uh uh uh degradation. So having the tools in place is important for that. To start, it's uh basically having the commitment to understanding that there's something that needs to be done. The tools in themselves that are uh out there like ours are easy to implement, are easy to monitor, are easy to handle. So that's not the issue. I think it's more important to have the concept, the idea that monitoring is required and evaluations are required in order to get a better better output of AI.
SPEAKER_00And uh, you know, uh spoiler alert for the listeners, when I I was interested in this for our you know, our own business. And when I inquired about the pricing, I was uh overwhelmed with how generous their pricing is. So just so you guys know, this is something that you need to be doing in your business for any act. Obviously, some activities are more important than others. It's doubtful that many of you have uh, and because listen, only the reason I'm saying this is I'm inside of a lot of businesses as an advisor, as a you know, a leader of bringing AI into their business. I doubt that many of you have any type of evaluation, whether manual or you know, computer assisted, for the AI activities that are occurring in your business. And the main thing I want you to be thinking about right now is where are we at risk? Because the business owners that we talk to who are interested in, hey, we've got to do something about AI, we're ready to go, but we don't understand the risks enough to go all in. If that's you, right, then I've just exposed you to a whole nother vector of risk that you probably hadn't even thought about yet. But the good news is that now that you understand this and now that you're becoming aware of solutions like Ragmetrics, it's not going to be a big deal for you to confidently operationalize AI into a lot more areas, even areas that are uh, you know, there's compliance involved. There's uh, you know, high uh uh high penalties for any type of mistakes and things like that. So the good news is the risk is being mitigated on this call. So, Hernan, um I what we've been doing is manually like human map mapping this on spreadsheet, and it's better than nothing, right? Version one is better than version none. However, it takes a lot of time from both the chief AI officer and the pilot teams. It is open to human error and all the things that you know are the promise of why you'd want AI in the first place. How does a company transition from not doing this at all and being like, oh wow, we didn't even think about that, to maybe those that that are randomly checking outputs and grading them on a on a whatever their their scale is and tracking it in a spreadsheet to going into an environment where is it automated? I mean, what what is the what does that progression look like with the final result being your solution?
SPEAKER_01Yeah, so so yeah, for sure, the the objective is to be to be automated. So um so there are there are three types of of projects or companies working on AI and how they address this issue. There are three different the ones that do nothing and basically say no AI is good enough, let's continue with that. And then the problems come and there are several stories in the market about that. The ones that test manually, uh that I think they from the ones that are testing the few that that test something they test manually, and then uh the automated process like ours. So the process is very is very easy in in concept because um there are two ways, as we said it before, before the product is in production and after the product is in production. Before the product is in production, the process is very easy. It's getting the knowledge base, getting the information that should be used to train the AI tool to understand how the AI tool should behave, getting that information and creating test samples, we call them data sets, test samples, and inject those test samples through the pipeline, the AI pipeline, and see what comes on the other side, and basically create that. And that process is basically just uh, without getting too technical, just injecting information to an API and getting the response to an A from the API, and then we look and we evaluate. And as I said before, there are hundreds of metrics or criteria that we can use to evaluate. So the processing itself, from the technical point of view, is extremely easy, similar to what we call live AI evaluations or or monitoring. That is, we get the information through an API of the context, the answer and the question, or the input and output that that AI agent, AI component had, and we look at that through the same metrics or different metrics or the metrics that are required or needed in that specific point. So the process on the technical point of view is very simple. Um as I as I said before, and I'm I'm thinking I'm becoming repetitive, uh it's it's it that comes with age. And becoming repetitive. The importance here is to define what needs to be tested and how it's going to be tested. And uh, Chris, there are there are hundreds of examples, and probably we should have uh mentioned those at the beginning from Air Canada being sued because a chatbot didn't provide uh right information about a flight and tickets to a fast food company uh um getting an order wrong and putting hundreds of thousands of French fries in in the same ticket, to New York city government providing information to someone to break the law. So there are hundreds and thousands of examples that I can help any day. The most important thing is are we ready to start testing? Are we ready to start the evaluation? And it's uh it's in the leaders of the organization to drive that forward.
SPEAKER_00So, okay, I'm gonna run through this so that I can extract information from you and apply it to the client experience that we're giving for our people, right? Okay. So who is the person in the organization that I need to on the client side or the internal stakeholder that is most likely the person that will understand the concept, understand the requirements of it, and be able to fulfill on the on the client side, on whatever those obligations are for this?
SPEAKER_01Yes. So there are two two audiences. So one of the reasons why I started with with this project is I had a conversation with a CIO of a bank in the Midwest. Customer for a long time in other businesses. So I went to see him to see if he wanted to implement some chatbots and voice bots to handle customer care customer service. He looked at me in the eye and said, Hernan, I'm not going to implement it. I'm not going to implement AI because I'm afraid of the outcome. We are in a regulated market. If something happens, we cannot control risk.
SPEAKER_00Sure.
SPEAKER_01We cannot manage that. So that's how I started here. So that that is one audience. One audience is the business owner, and I mean business owner in the broad sense, could be division manager.
SPEAKER_00Yeah.
SPEAKER_01Yeah. The business owner that wants to have to reduce risk and the risk profile of the services that he's given, either internal or external. So that's that's that's one group. The second group is on the IT side or on the development side. That is, people that are every day working with AI tools and need to develop and need to accelerate the process. And in order to bring a product with quality, they need to test and evaluate. So those are the two audiences. So the business owner and the group that develops and implements AI solutions, AI agents in general.
SPEAKER_00So now I've got my internal representatives for this effort on any AI deployment. Next, my question would be who's going to develop the tests?
SPEAKER_01So the tests are very easy to develop, and we can help our customers, or we help our customers to develop that. The test is very easy to develop, and I sort of gave an indication before. Assume that you want to test a chatbot for retail that deals with retail policies for refunds, exchange, or whatever. So there's a policy that is written, it's probably in a PDF document. So and probably that same policy is the one that has been loaded into the chatbot to provide the answers from. So we get that policy and we create the test scenarios based on that policy. It's as simple as uploading the policy into our tool, our system, and asking the system to generate hundreds of questions based on that policy. So that is phase one. Phase two is to define the criteria that you're going to use. What do you want to measure? Accuracy? Do you want to measure succinctness? Do you want to look at how many tokens it's used? Do you want to measure the quality of the output, the tone? There's several. So that's the second phase. What is what you want, which the criteria, what is what you want to measure? And then the first, the third phase is running the um tests and looking at the results. And as I said before, in general, the tests are run uh through an API, or can the information can be uploaded to our system and our system can run it offline. For that, we also need the prompts, we also need which are the models that you're going to be using, and and and and yeah, and the basic things. Um, but yeah, that's that's basically how how the the it runs. Once you have the results of the test, you can look and define, refine your prompts, or change the model if you want to try another model to see what it's better and things like that. So you have all the tools there to optimize your your platform.
SPEAKER_00So as a listener, I want you to think about this. Could you do all those steps manually? Could you go to the models and say, hey, here's our situation? Can you create these test data? Sure. Could you, as a human, go in there and copy and paste and hit enter, and again, and again, and again, and again, you sure could. However, um, before you got exhausted, maybe you've uh run through X number of scenarios. How but that's not probably gonna touch all of the edge cases that'll come up in the wild with humans interacting with your AI system. But when you have a technology like this, and I would assume, Ernon, that that the one of the main benefits of the API integration is that I can do uh I can do high-speed testing very quickly, a lot of scenarios.
SPEAKER_01That's the objective. It's uh you for sure you can always do it manually. Uh um yeah. But basically is instead of taking a minute per question, two minutes per questions, we can do hundreds of questions almost simultaneously. Uh and you put them there and you look at the answers and you get the statistics on the back end. So you can look first overall statistics. The overall of this is, as we said, 4.9, so it's good enough. The other, the other set that we tested is 4.92, so let's look why this one is better. So you work on the real things that is optimization and getting the things up and running faster and better, not just spending time in creating. The other thing is whatever the model is, you're ensuring that you are having consistency in how it's evaluated. So if you continue evaluations on if you have people, people might look at them differently. If I analyze all the questions, I might have the same rationale. But if I give some to you, you might say that something is a four when I thought it's a three or a five, and there might be differences. So here there's consistency throughout the process. So if you have an update on the product and in two weeks you come and change a few things on the product, then you can test with the same uh uh data sets, the same data sets, and look at how that changed in in your output. It got worse, so the update is not what we needed. It got better, there we go. So that's that's the idea of handed.
SPEAKER_00Yeah, that's very helpful. And I mean, I guess to give it an analogy, I've got to be in San Diego on Saturday. I could walk to San Diego, which would be the equivalent of me manually doing these evals, or of course I can jump on a delta flight, and which would be the equivalent of leveraging technology to get the result faster. Okay. Um so okay, great. I've got this. We we've we've defined it with the the stakeholders internally. We're starting the testing. So I would imagine that the the evaluation of that though, you still want some human in the loop to maybe do a sampling of comparison to what the system says is accurate output, and the human saying, hey, well, wait a minute. I do this every single day. I respond to these tickets, I answer these questions, whatever that is. I don't think that's the same. Where is human in the loop on further evaluating what the what the evaluation system evaluated?
SPEAKER_01So we see human in the loop in a very minimal aspect, not because we want to minimize humans. That's not the idea. But in a very minimal aspect. And it's to retrofit information back into the system in the moments where it cannot be done automatically. And then we can talk a little bit more how to correct a few things automatically. But in a model that cannot be automatically. So you look at numbers, you decide that there's something that needs to be improved, a prompt or things like that. In some cases, needs a human to do those things. Our system, uh, we tested, we did some testings, and and our system is 95%, uh, it has 95% human agreement. That means that 95% of the time a person of the system will rate in the same way. So uh we believe it's it's pretty pretty good and pretty consistent in that in that front. So we try to minimize the use of people uh only for those cases where it's it's it's it's extremely needed. Okay. The AI tool, the AI agent, and work on that, on that front.
SPEAKER_00How are you guys capturing that? Is that sitting with the process owner and documenting, you know, how they do a certain process and what they look for? Or how are you capturing that information?
SPEAKER_01Basically, it's uh we don't develop the agent. So that's part of the process that the agent development has to do. We look at what we call the knowledge base. So that information should be in the knowledge base. Yeah. Yeah. Well, a knowledge base in the broad sense, if you want. But that information should be in the knowledge base. So if if a knowledge base, as an example, could be uh uh an FAQ, a frequent to the ask question. So that means that if a customer says, I want to return this, the first thing that you the the the person asks is when did you buy it? Or uh and things like that. So that is part of the knowledge base, and that is uploaded into the evaluation process.
SPEAKER_00Okay. And again, through the lens of our our practice of doing this for companies, I'm asking these questions, but I think they're applicable to a business owner so that they can understand when they need to plug this type of thing in. What scope of a of a project would you say would be too small for really, oh, you don't really need evals because the human's going to be involved anyway, as compared to something where it's obvious, as we mentioned, there's no way that my team can test nearly as many scenarios as the tech.
SPEAKER_01So there are two as as as we spoke, there are two phases, one pre-production and one post-production. I would say that the post-production area is any. So there's no minimum. Okay. As an example, we we um develop uh uh a node for NA10 and non-code tool for AI agents. That that's basically for we'll say for any any segment, but starting with it with a low, low tier. As soon as an A approves, it's going to be available. But basically, it's we are targeting any segment. And that's the objective of that is what we call the monitoring piece of the live AI evaluations. So that has no no minimum, or okay, nothing's too small from there. From the rest is what you define, what you as a developer or implement uh someone that implements an AI agent, what do you define that your need is? So I can I can give you a very simple example. I develop internally a chatbot uh for for a demo internal, and before using it, I run it uh through some evaluations to measure the accuracy and also to measure the consumption because we are a startup, we don't have uh uh too many funds, so yeah, we measured the consumption. So we realized that some of the things we were doing were consuming more than we should, and I can give you an exact example and explain how how it went. But uh so and basically this was just for a demo internally. Then I ran a 20-minute evaluation and I decided to change the way that the rag was retrieving information, and it was and that basically the new the new way used half of the tokens that the previous way. So it's not only accuracy, it's also to optimize on the cost side. So yeah, it's it's whatever you want. And basically, this was with 25 questions, it's not hundreds, it was with 25 questions, just to have a sample of the consumption. So there's there's no real limiting what you want to do.
SPEAKER_00Okay. Um the evaluation process, if I'm already in in action with some pilot projects, and and at least with chief AI officer, we kind of have a three-phase. We we model the process, then we kind of do like a build and assessment, and then once we feel it's hit an accuracy level, then we roll it out into production, but monitored rollout over 30, 60 days. So in this environment, that would fall into step two for us, where we're actually we've we've built the test thing and we're we're evaluating. Let's say we run it through there, and you mentioned a 95% accuracy target. That's an internal target that we also uh use when we're working working with clients, but again, we're doing it manually. How how long, provided that it it hits that 95% accuracy in the RSS phase, how long is that that evaluation process going to take? Minutes, a few days?
SPEAKER_01Minutes or okay short hours, depending how many how many samples you want, but a couple of hours. Uh so it's in the day range. So it's a few hours during the day, and that's that's basically it. Not it shouldn't take it shouldn't take it shouldn't take more more than that. In my mind, uh, Chris, what it's what it's more important is not only doing an evaluation once, but finding a way to continuously do evaluations, either by monitoring in real time, or every couple of days inject a set of to uh of a data set to evaluate and see the output and see how that is uh changing in time or not. Because again, we go back to drift, we go back to hallucinations, we go back to errors, and that is what we want to prevent. So ensure that there's something always there that is being uh uh it's driving evaluations.
SPEAKER_00So for me, I'm seeing uh a constraint in our model because here's our option now if we're doing it manually. Put it in a checklist and give it to the client and hope fingers crossed that they are they're like, oh yeah, it's working fine. I don't need to test it, right? Let's hope that they don't do that. Um, or a chief AI officer, highly compensated individual with you know specialized talent, is going in there and doing some basic, right? Is this accurate, right? So it's not a high leverage activity for that stakeholder in the company or an external chief AI officer. Could this all be set up to where there was an automatic trigger to run a slight eval every 72 hours or whatever and then report back to like that's all automatable. It's okay.
SPEAKER_01Yeah, basically it's it's yeah, it's you have the data sets, you have the the the experiment that has the the parameters of what you want to evaluate, yeah. Inject inject the data set through an API, get the results on the other side, and and look at the numbers. So at night, every other day, you can you can run it, you can test it. Yeah. And I'm saying night, assuming that that is there's less than sure. Yeah. Yeah. Or the other way, it's having the the real-time evaluation, evaluation process uh uh continuously running. So those those are the two ways to to do it.
SPEAKER_00Depending on the I don't know, the process itself. Some might be uh once a week, once every three days, as compared to some are doing it multiple times per day because there's such a, you know, like we have to make sure this is right every time. Okay, this is very helpful. And actually, it it's uh this is one of the things I love about generative AI. One, there's so many layers to it, right? So we've got listeners to this show that they're they're like, you know what, I'm getting good at prompting. I'm starting, right? And they think that they have like, and again, guys, if you're listening and this is where you are, there's no slight to this. This continues to happen to me, somebody who's immersed in this for, you know, past several years every single day. But you you you do something and then you're like, oh, wait a minute. That's just like the basics. There's this next level. And then you kind of get in there, and then, oh my gosh, there's another level. And that's kind of what I'm experiencing right now with this understanding of, and now that I'm seeing it, it's like, you know, oh, of course. But until you get exposed to an idea, you don't, it's it's not like, oh yeah, of course that makes sense. You just you don't even know that that exists. And that's what's happening here. And I hope for the listeners that you're seeing the wow, oh what this eval concept hadn't even thought about that before. And I mean, I'm a smart guy, I work with a lot of businesses. It's not something that I until I started talking to Heranan and his business partner a couple weeks ago, I didn't realize the impact of this. We were doing it manually. We were trying to hit that 95% accuracy before we we were using the clients, pilot teams to evaluate it and all those sorts of things. And now I'm realizing another layer to this thing. This is so much smarter, better, more accurate, cheaper, faster, everything. So this is good.
SPEAKER_01The example that you gave on the guy that is good at prompting and he knows, and so my question to those guys in general is okay, you change five words on that prompt. Right. How do you know it's better? How do you know it's better than before? Because you run 10 tests and you say, yeah, it's better because this tens, because that is what why you change it for, because you wanted that outcome. But what happens with the other 400,000 uh scenarios that you haven't tested? So the idea of having uh uh automatic testing is to resolve that in 10 minutes, you know, yeah, this is better because this, that, and the other, and and you have confidence on on that. So yeah.
SPEAKER_00So, and just again, through the lens of of the practice that we have, what does the what does spinning this up in an environment look like? Like what's the onboarding for me as a client who would be interested in testing the tools?
SPEAKER_01As I said, the if it is pre-production testing, the onboarding is extremely simple because it's basically using an API and understanding how how it works. So in general, we we we work with customers on that front. It's very it's very simple, it's no no no big deal. The live hallucination has two components, one that we haven't mentioned yet, but I'm going to mention it now because that adds a little bit more uh tricky things, let's put it that way. One component is just getting the monitoring piece. So basically every interaction that goes or every output that goes from an AI agent is being evaluated, and you get an answer. The other component on that is how to retrofit that evaluation. Then what I mean evaluation is a score and a reason why the score is what it is, or why that evaluation failed, or what whatever. So, how to retrofit that into the AI agent to force the AI agent to correct itself. That has a little bit more components because that requires a modification of the prompt and and and other things, but still it's extremely uh uh easy and friendly. It's all a couple of APIs that go from one side to the other. So if someone wants to try the tool now and and has a little bit of knowledge on on the technology side, I'm pretty sure that in an hour he he can be using it without without any problem.
SPEAKER_00So my takeaway from that answer is that it's way more important that we introduce this early in the process as compared to oh, we'll figure it out once it goes live.
SPEAKER_01Like again, seems obvious when I say it, but um Let me let me give you let me give you a real example on on that I was mentioning about that that chatbot. So we were we're we're working in in a specific demo for for a retail customer. So got the retail policy uploaded into a rack, a pine cone rack, nothing too fancy, or something very very simple. So I said, okay, let's get uh five chunks. So the top the top K is uh five chunks, and let's start to work from from there. Nothing too crazy. The chatbot in Python, streamlit, nothing, nothing too too fancy, something very simple, very, very normal. And I said, this is work, it's fine. What happens if it between if the the instead of five chunks, we get two chunks? So top K for the ones that don't know, it's it's two instead of five. That means that basically we are getting less data from the rack to evaluate. So the first thing is if we look only at the output, probably the accuracy it's it's better in the two because it's more precise. As you know, the more data that you get from the rack, the the more you know the more uh accuracy you lose because basically you get what is more accurate first and then you come down the line. Fine, good. But it was interesting when we when we tested the generation piece, also was more accurate in in in when we talked two chunks instead of five. And that's because simple rag, simple questions, simple chatbot, the uh in two chunks it has enough data to provide the correct answer. So the other three chunks were adding noise. So uh it was more accurate in uh with Topk2, so basically two chunks, and additionally, when you were using only when we are using two chunks, we have less data that is transferred to the LLN to generate the response, so less tokens, so it ended to be around half the consumption of tokens. So those are the things that if you are way down the line, you don't necessarily are coming back to test. So the earlier that you can start testing and evaluating, the more that you can drive accuracy to your system.
SPEAKER_00Yeah. That makes perfect sense. Like I get it. This is helpful. This is now like I'm my mind is spinning because I'm thinking about all this training I need to go, not necessarily rework, but enhance with the concept of this new um this new approach, which would be leveraging technology to do high velocity, uh high volume evals as compared to what I've been doing, which is horse and buggy, right? Like the old way of got a spreadsheet, we're doing some tests. This is fantastic. So it's a huge win for not just my my business, but uh the efficiency that will, you know, any client will be able to uh experience from this eval. And uh a confidence on our side, or and if like folks. If you're listening to this and you're the person leading this internally or you're working with an external or an internal person who's leading this, it's important for you to know this. And I would actually take this back to them and say, hey, what's what's our our protocol for evals on our processes? Right. Now that you have a good understanding of what it is and why it's important, um I would be curious, are your people doing this? Right. And if not, and Hernan, now now what I'd like to do is kind of shift to where if people want to find out more where they can go. But let me ask you first, you guys don't run the evals, you have the technology that supports the running of the evals. Yes.
SPEAKER_01Okay, for sure. If someone needs help, we will provide it, there's no no doubt. But we don't run the evals. The idea is that uh the the people implementing, the people monitoring, have their own uh implementation of the evaluation process and they run it and they run it whenever they want, however they want. They can change metrics, they can look if there's some issue. I don't know, they detect that discounts are coming too often. Uh so they can create an evaluation process or monitoring process to look at the discounts in in in in that chatbody, in that agent. So, yes, there uh we allow people to do it.
SPEAKER_00So, okay. One of the things that I would be thinking about if I was listening to this was that's great, but my people don't know how to do that. Where can people, where can teams like pilot teams or implementation teams learn how to set up the evaluation?
SPEAKER_01We we have a very yeah, we have we believe we have very very good information in our website, Ragmetrics.ai. There's a resources area and there's the documentation area. The resources area more tells the story on how to run evals, what is important, different types of evils. So I think I think there's uh a lot of information there. For sure, for sure, they can contact us uh through the website, and we are more than willing to help people. So we believe that this is new in the market, it's a new category that is going to be defined. We are more than willing to help anyone that has a question, anyone that has a comment, doesn't know how to do it, more than willing to jump on the phone and give examples, even provide code of things that we've done internally in order to help them.
SPEAKER_00Okay. Now, I'm preparing to roll out AI or I'm talking to my board next week, and we've gonna come together with thoughts. I've just listened to this podcast episode. I know they're gonna ask about budget. Let's say middle market company, we're gonna start with maybe um one or two pilots across finance, HR, operations, sales, um, and marketing. What's my budget need to be?
SPEAKER_01Go ahead. So so basically, we we uh um the unit of measure is the evaluations. So our our package that we are launching now, it's around $250, starting from there. That that provides around 3,000 evaluations, so it's a good number to start in the process. From there it goes up, and for sure, if someone has a specific need, we are more than willing to address it. But it's not that you need to uh sell the house in order to get an evaluation process. That's that's not the idea. And the more that this runs, the cost for evaluation will will go down, and yeah, and it's going to make it easier for for for everyone.
SPEAKER_00Yeah, I'll tell you, when we when we first had our conversation, and once I realized A, the impact, but also the the technical nature of it, I was expecting a pretty big, I was expecting that this to be inaccessible price-wise, except for companies that were like working with Deloitte or whatever, right? I didn't realize how affordable that was. And when you work out the math, 3,000 evaluations, let's say um, we said, you know what, we're gonna save the 250. We're gonna have our people do it. For your people to run 3,000 evaluations, the payroll on that would be way beyond the 250. Pony up the 250, do it right, remove the human error from it, and let the tech handle it for you in a day as compared to a couple of weeks. So that's my position.
SPEAKER_01And I and again, uh more than willing to so for sure we have models that if you're a large company and want to have something in-house due to privacy, we can host it in-house, on-prem. So we have different models, but but yeah, the idea here is to help the AI developers uh understand what they need to do. And also to your question about the board, I thought you were going to go to the phase, okay. And which is the risk of implementing AI? Why do we want to implement AI? And I think the big part of the risk of implementing AI is that AI hallucinates, AI brings error, that jeopardizes cost, that jeopardize customer retention, loyalty. So by having uh a tool that evaluates and monitors in real time, you are mitigating all those risks. Yeah. I think that in my mind that's that's the most important, important thing.
SPEAKER_00Yeah, and it's uh a low-cost insurance to make sure that all this energy you're putting into introducing AI into your business, that it's working and that it's something you can count on. So like um this has been fantastic for me because now I feel like I've got a whole different level of uh comprehension so that I can now explain this to other businesses and obviously the stuff that we're doing internally at Chief AI officer on why it's important and why that like you you don't want to not do this, right? Like you don't want to put AI into production and just hope that you know your employees are you know spot checking it on gut and getting it right. No, you got a business at stake here. So well Hernan, any any uh like I guess closing thoughts or considerations for everybody um as they go and chew on this?
SPEAKER_01No, just don't feel shy about AI, don't feel shy about the output of AI because it's good for everyone, it helps in a lot of ways, but ensure that you have the correct control points and the correct uh boundaries to understand what's going on in with the AI output. That's basically that's basically it.
SPEAKER_00I think that's fantastic advice. So we're gonna have uh links to ragmetrics.ai, we're gonna have links to your social um connections and all that sort of thing. And I would encourage you, uh listener, dear listener, regardless of where you are in the phase, if you're leading, if you're investigating, if you're somebody who is at the top of the organization, if you're somebody who's kind of peripheral to what's happening, if this isn't being discussed, I would say, hey, like send them the episode or whatever. But but I would I would say this is certainly a consideration that we need to do. And I think I've got a resource that we can start with, right? So awesome. Well, I look forward to uh staying in touch about this subject as you guys continue to make breakthroughs in the startup environment. Um, and I look forward to actually engaging um with Ragmetrics as a client for our clients to make sure that we're, you know, we're we're doing evals in a way that I feel very comfortable and confident with.
SPEAKER_01Yeah, Chris, thank you very much for the opportunity sharing this uh with you and your audience. And yeah, looking forward to continue conversations in all fronts.
SPEAKER_00Thank you very much. Thanks everybody. We'll catch you on the next episode in the meantime. Uh just keep using AI. Thanks for tuning in to using AI at work. Don't forget to subscribe for more conversations about how to use AI at work. And a special thank you to our sponsor, Chief AI Officer, for empowering businesses with AI education and training. Visit their website for free AI readiness assessment and AI strategy guide to help you get started using AI at work. That's www.chiefaiofficer.com. Follow us on Twitter at the handle usingAI at work. And visit www.usingai at work.com for free resources to help you harness AI in your role.