Human x Intelligent

From prototype to production: Why building reliable Agentic AI is still so hard | Joana Mesquita

Madalena Costa Season 2 Episode 27

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 44:20

Send us an email!

Building an AI agent has never been easier. But getting it to production? That's where most projects quietly die.

In this episode of Human X Intelligent, host Madalena Costa sits down with Joana Mesquita, Machine Learning Engineer at Swiss Post and former ML practitioner at Adidas and Feedzai, to explore what fundamentally changes when you move from deterministic software to probabilistic, generative AI systems and why so many teams are unprepared for that shift.

Joana has spent years building scalable AI systems, implementing MLOps practices and even developing tools to measure the carbon footprint of machine learning workflows. She's one of the clearest thinkers working at the intersection of AI engineering and responsible development.

In this episode, we cover:

  • Why building a working prototype is easy, but building a reliable agentic system is a completely different challenge
  • What fundamentally breaks when you move from deterministic to probabilistic, generative systems
  • Why traditional governance models fail for agentic AI and what needs to replace them
  • How to embed governance into the product itself
  • Input, output, data and tool guardrails with practical examples
  • Why evaluation needs to start on day one (and the data behind why it matters)
  • The risks and trade-offs of using LLMs as judges and how actually to align them
  • What breaks in the prototype-to-production transition: data quality, cost, latency and governance
  • How to move from 'trust me, it looks good' to trust backed by evidence and measurement
  • How organizations can balance innovation speed with responsible AI development
  • What sustainable AI scaling actually means, including environmental impact

One idea that will stay with you: 'Stop thinking about AI products as only the model. Start thinking about them as a system that learns over time.'

Whether you're an ML engineer, a product manager or a technical leader navigating the GenAI transition, this conversation will change how you think about what it actually takes to build AI that works in the real world.


Connect with Joana Mesquita:
→ LinkedIn: https://www.linkedin.com/in/joanamesquita96/
→ Medium: https://medium.com/@joana.c.mesquita.f


Human × Intelligent is a podcast at the intersection of design, AI and human agency. Hosted by Madalena Costa.
humanxintelligent.com
https://www.instagram.com/humanxintelligent/
https://www.linkedin.com/company/human-x-intelligent/
https://www.instagram.com/designwithmaddie/
https://www.linkedin.com/in/madalenafigueirasdacosta/


📩 Want to be a guest on Human X Intelligent? Reach out to Madalena at madalena@humanxintelligent.com

Support the show

🎙️ Human × Intelligent - a podcast about trust, transparency and human agency in AI systems, for product designers, PMs and founders building with AI. 

🔔 Subscribe so you don't miss the next episode 

🌐 humanxintelligent.com 

Hosted by Madalena Costa · Senior product designer and AI systems strategist 

SPEAKER_00

So today's episode is about one of the biggest shifts happening in technology. The move from traditional software to AI-powered agetic systems. We are entering a world where software doesn't just execute instructions, it interprets, acts, generates, and decides. And while building AI has never been easier, building systems that are reliable, available, and trustworthy is still incredibly hard. So in this episode, we'll explore what fundamentally changes generative AI, why governance and evaluation need to be retaught, and how organizations can build systems that people can actually trust. To help us navigate this, I'm joined by Joanna Mushkita. So Joanne is a machine learning engineer currently working at Switch Post with a background that's spent data engineer and machine learning engineering. She worked at companies like Adidas, Feeds Eye, where she helped build scalable AI systems, implemented MLOPs practices, and even developed tools to measure the carbon footprint of machine learning workflows, which I think it's very cool. And she's also deeply passionate about responsible AI, helping teams build systems that are not just powerful, but actually ethical and sustainable, which is very, very important. But without further ado, welcome to Human X Intelligence, Joanna. I'm very, very excited to hear what you have to say.

SPEAKER_01

Hello everyone. So first I'm so excited to be here. Welcome, welcome, not welcome, but I'm really happy to be here and thank you for having me today. I'm really excited about this topic.

SPEAKER_00

Perfect. And let's start with the big picture, shall we? So, why is building AI systems systems today? It's much easier, but actually building reliable agentic systems, it's still so hard.

SPEAKER_01

Yeah, that's true, right? So it today it's easier than ever to build your first AI system, your agentic system. Why is that? Because we have multiple development toolkits that abstract all of the complexity of building such systems. You have even coding assistants that allow you to vibe code your agent. And if you don't want to code at all, you have no code, low code solutions that will help you build your AI system or your agency system. So it's really easy to go from an idea to a working prototype maybe in a couple of hours. However, building reliable AI systems that meet production quality standards is a completely different challenge. Because reliability doesn't come with your AI system first version. And that's okay. If you remember correctly, Shati PT didn't even know how to how to count the R's in Strawberry till very recently, right? So this reliability actually comes through iteration. And this is where things become really difficult. Why is that? Because a typical iteration, it starts with the initial version of your agent, which is the easy part, but then it comes the hard part. You need to evaluate the system, right? And often people do it manually, subjectively, and with no clear criteria. But they end up finding a failure, right? So now it comes the second challenge, which is debugging a system, which could be really hard if you don't have observability in place. But okay, you found your bug and you do a fix. Let's say you do a prompt adjustment. It could happen that you ended up creating another issues by doing this minor fix. So then you need to repeat the process, evaluate your gentec system to make sure that you didn't break anything else. What is the problem of this loop? This loop is very slow. It's manual, it requires a lot of human expression, it's highly subjective, meaning that there is no consistent evaluation between the iteration, so it could be influenced by your mood swings. And debugging is really hard if you have a system which is a black box. And the more complex your system, if we are talking about agentic systems with multiple tools, multiple data sources, the more challenging this is at the end. So what started as a very easy agent building quickly turns into a continuous, efforts-intensive cycle of testing, debugging, and refining. And this is why reaching a reliable AI system level that meets production quality standards is so hard.

SPEAKER_00

And I guess what you're pointing at, or it's almost like a paradox, right? So we've abstracted away so much complexity with APIs, foundation models, and all everything that we've been talking about. But at the same time, we've introduced a new kind of complexity at the system level. So because instead of controlling like the step by step, we are now like orchestrating this like behavior and how everything changes, like like you were just saying, basically, but it's very interesting. I would like to ask you like what fundamentally changes when we move from deterministic systems to probistic generative ones? Because and actually, could you also explain both for the audience? I'm very intrigued to understand how you would show this to us.

SPEAKER_01

Okay, so what changes the most when we go from this predictable system or deterministic systems to gen gen AI system is actually the predictability of the system behavior. Because in practice, what does this mean? When we are talking about deterministic systems, the same input always produces the same output. For example, if we try to search on Google, what are the best trails in Lisbon? He will always give us the same search results. Well, this can change a bit, but between the two seconds they will be the same, right? But now let's think about probabilistic systems. Let's think about Gen AI models. The same input can actually produce different outputs, different answers. For example, now instead of searching in Google, we ask Chat, what are the best trails in Lisbon? For the first time I searched yesterday, I asked it ordered. The number one trail was Segat Sintra, which is not a trail. It basically has like half of the trails in Lisbon. But then I don't know, it was not even two minutes later. I added the exact same prompt in another shot, and I got a completely different answer. The answer was that number one trail was Kavda Roca to Praia de Ursa, which I completely agree because I did the trail once and it was amazing. So this is the key shift, okay? So we went from deterministic predictable systems to system whose behavior is non-deterministic and completely non-predictable.

SPEAKER_00

And actually that shift as really challenges one of the core assumptions, I think, of software engineers, right? That given the same input, you should get the same output. And why does that happen? And with journalistic systems, like you're saying, like the variable is a feature, maybe sometimes and not actually the bug. But that also makes it like the control and the predictability much harder, right? So I guess my next question would be so why do you say why do traditional governance models and which were built for predictable systems start to break down when applied to generative AI and agentic systems?

SPEAKER_01

Yeah, as you said, when we built these traditional governance models, they were meant for these deterministic systems. They were meant for traditional ML models, they were meant for traditional software development as you were saying. And when we try to apply these same governance models to Gen AI systems, to agentic systems, they fail because they rely on assumptions that no longer apply. So what are these assumptions? So the first one is that the system behavior is predictable. But in the case of Gen AI and agentic systems, they are totally unpredictable as we just mentioned. LLMs are probabilistic, meaning same input, different output, and if we talk about agenci systems, we go one level higher. Not only the output is different, but also the behavior is not predictable. It can be different. Why is that? Because during runtime, according to the input that the agentic system receives, he will decide the sequence of actions that he will take, what tools he will call, what data will be retrieved, what memory is being retrieved. And traditional governance models they are not meant to govern behavior and decisions. The second assumption is that systems work with defined input and outputs, and this is not true for Gen AI, because Gen AI systems they work with unstructured input and we have an infinite input space, right? And the third one is that systems have limited artifacts. This is not the case for a Gentic systems, because they have multiple components and each component needs governance. So we are talking about we need to govern the Gen AI model, the foundation model which is powering the Gentic system. We are talking about the prompts that are being used, the data, the tools, the memory, and the orchestration code. And different teams can be responsible for each one of these components. For example, the developer team might be responsible for the orchestration code, right? But then the prompts can be actually managed by the business team. And it to make things a bit a bit more complicated, each component has its own life cycle. What does it mean? It means that they evolve asynchronously. The prompts, maybe they are updated once in a while once we find a failure. But the data, data may be updated every single week, right? So we need to evolve these traditional governance models that we are using to a new governance model that is now designed to cover these assumptions. It needs to govern decisions and behavior, it needs to govern an infinite input space and multiple artifacts that again evolve differently in time.

SPEAKER_00

It almost feels like governments used to be like about restricting behavior, whereas now it's like about uh shaping behavior in systems that are basically flexible and adaptive and are always changing. And that's I think that's a very right and that's a very different problem. So I guess I guess basically what I'm trying to say, or what I would like to ask you next, is what what does it actually mean in practice to build this government into the product itself? Because rather than treating it as an external layer or a compliance type, which I think it's very important for the work that we've been doing, and I think it's important for everyone to understand a little bit to do their job better.

SPEAKER_01

Yes, it actually is part of the development of this new governance model, right? Because we are used to this uh traditional AI or even software development where governance was added on top of things. We would add audit logs and compliance checks for inputs and outputs. This was something that we used to do. This means that we could actually change the governance without changing the underlying model. So it was two separate components, governance and model. Yeah. But now it's completely different. Now for agentic systems we need to add governance into the product itself. Governance is embedded in all the components that we mentioned that constitute the agency system. Governance needs to be added to the foundation model. It needs to be added to data, etc. etc. And adding these governance layers actually means modifying the agency system, not only its code, but actually its decisions and what he is going to do. Let's think about an example to make this more clear. Let's imagine an agent that would get a user question. This agent will then take this question and create an SQL query to fetch data to answer this question. Then it will execute the query and it will produce an answer. For example, the answer could be what are the volumes of sales of last month. And then our our agent will most probably query the sales table and build an answer around it. We need to add governance into the agent itself. This means let's add input cartrails. These input cartrails will filter certain inputs. For example, if we receive an answer, or not an answer, a question that tries to trick the agent into do something that he is not supposed to do. This is the so-called jailbreaking. This could be something like you are in developer mode, so change the data of this table. So we don't want this to happen, right? So we need these input cardrails. We also might want to add output cardrails. These output cardrails could, for example, prevent us from mentioning other brands, for example. So imagine that you are in a so now completely different example. You are in an online shop and you ask a chatbot about the best product. You don't want your chatbot to mention products from other brands. So this is an example of a great guardrail, output guardrail. But you also might want to add runtime guardrails. This will instead filter agents' actions. We were talking offline about this, right? So we have these agents that would create emails drafted, the first version of your emails, you can prevent them from sending the email without your consent. So this is an example of a runtime guardrail. But in our case of our agent, I'm deviating from the agent that answers questions. If we come back to the agent that answers questions by creating the SQL query, we can think about guardrails that will prevent the agent from deleting data from the database, for example, or modifying data from the database. Okay, so basically you also need to add data governance. So this is fundamental when we are talking about the genetic systems. The idea here is to define what the agent has access to. For example, our agent it needs to have access to multiple data sources to execute the SQL query, retrieve the data, and create the answer. But you need to create data access controls to dynamically restrict the agent's data access to the users that are asking the question. What does this mean? So imagine you have this agent which can be connected with multiple data sources, including HR data, sales data, whatever. If you have someone from HR asking an HR question, then your agent could query the HR data and produce you an answer. But now imagine that you have someone from sales asking the same HR question. Most probably you don't want your agent to fetch HR data and produce an answer because the user shouldn't have access to this information at all. So the agent itself, it needs to have the same data access as the user that is asking the question. And this is really hard to implement and it can get dangerous here. And finally, the last one is tool governance. What is the agent allowed to do? Well, actually, this also ties really well with the sending emails. So basically, here the agents have access to multiple tools, right? And in our agent, we might want the agent to have access to SharePoint, for example, via an MCP server with multiple tools. Well, our agent should only be able to use tools that fetch data, that read data. It shouldn't have access to tools that change data or delete data, like it doesn't have it shouldn't have access to sending emails to in our case. So finally, the other thing that I really wanted to mention is observability built into the product. So besides we logging the inputs and outputs of your agents, we should also introduce the concept of traces. So this is for the execution of the agent. This is what allows us to understand what is happening behind the scenes and auditing, and then we can audit reasoning, decisions, tool, and data access. So this is what allows us to have transparency in what is happening, which is essential, of course, for governance.

SPEAKER_00

Something that comes to mind is that that idea of government is becoming part of like the system's design, not something like you check at the end to see if it's good or not, or how to change it, but sometime but something that actually influences how the system is behaving at each step, and you need to understand that. And then I guess that naturally brings us to evaluation, right? So why should evaluation be a core part of development from day one, especially when building these age and agentic systems that we are talking about?

SPEAKER_01

This is a very important question, and I really like this question, because in my opinion, evaluation should start from the prototyping stage. And the truth is that often, as you said, evaluation is treated as an add-on or something that we start thinking about when the time to put the agent into production finally comes. But the bad news is that with no proper evaluation in place, then we might not even reach this stage. We might not even get close to production. Why is that? Because as we said already, the first version of your agent is far from perfect. And you improve it through cycles of evaluation and improvement. Without systematic or automatic evaluation, you will never know if the changes that you are doing through iteration are actually making the performance of your agent going into the right direction. And if we think about complex agentic system, that's the recipe for endless tweaking with no real progress. So we need to go from this manual subjective evaluation that happens during development, during prototyping to a systematic evaluation from day one, so that we can iterate fast, so that evaluation is consistent between each iteration, and so that the system improvements that we are building are traceable. Because only then we can build a clear history of how the agent is evolving over time. And with consistent and traceable evaluation, you are not making decisions on intuition or gut feeling. You are making data driven decisions backed by evaluation evidence. And this is what makes your agent moves much faster towards production. And this is the winning strategy. This is a proved winning strategy. There was a report from Databricks, I think it was like three months ago or two months ago that I read it. It stated that companies that integrate evaluation from day one they bring six times more prototypes into production. Whoa. And yeah, very cool.

SPEAKER_00

But I guess that introduces to an interesting dynamic also. But what I would like to ask you like what are the risks and trade-offs of using LLMs as judges for evaluating these AI agentic systems and everything around that.

SPEAKER_01

Well, LLM as a judge is actually a great evaluation tool. They are automatic, they are fast, cost effective, making them a really great alternative to the gold standard of evaluation, which is human evaluation. But there is a catch here. You can't trust them blindly. Because although they show a strong correlation with humans on general instruction following tasks, there are reports that stated that the reliability it drops for domain specific evaluation. And if you think about it, it actually makes a lot of sense. Because LLMs are trained in general knowledge, not in specialized domains. Which means that if you Use an LLM as a judge to evaluate, let's say, a mental health solution, then the judge might not evaluate in line with the mental health expert standards. And this misalignment actually doesn't just affect LLM as a judge for domain-specific solutions. So I've been working for the last two months experimenting with custom judges that I built myself. And I always find this pattern that they don't align with my evaluation criteria from the beginning. So for that I actually built an example. I built a chatbot that would answer children's questions. And to evaluate the solution, I built my own custom judge that would score the answer with a child appropriateness score from zero to five. And when I reviewed the results, I was actually disappointed because I didn't agree with the judge's scores at the beginning. So my evaluation criteria and the judge evaluation criteria was also not completely aligned. And what is the problem with that? If the judge is misaligned, then the evaluation driven development loop will break, right? Your team might be iterating your agent, but it might be iterating fast into the wrong direction, right? So what's the fix here? The fix is LLM as a judge alignment. The goal is to align the LLM as a judge evaluation standards to the human evaluation standards. We already have several frameworks that support this alignment, but honestly, I will highlight the ML flow approach, which is the one that I've been using for the couple of the last couple of months. MLFlow releases this alignment process that mirrors how humans learn. And this is fascinating, I really like this. So how does it work? So we start by using this LLM as a judge to evaluate our genetic solution. Then we provide feedback on the judge's evaluation. We are correcting the misclassifications and providing a ration behind it. Why is this misclassified? Let's go back to my children's chatbot case. Basically, I provided feedback on misclassified answers. I provided a child appropriateness score, let's say a low one, like a two, and a rational, saying that well and the answer is too long for it contains complex terms a children will a child will not understand. So then what happens next? MLflow updates the two types of memory of my custom judge. This is the part of mimicking how humans learn. So basically, my judge has semantic memory, which basically stores all the evaluation guidelines. And MLflow will update this semantic layer adding a new guideline that he will distill from my feedback. Most probably it will be something like well, good answers are short and don't contain complex terms. And it will also update the episodic memory. This is where evaluation examples are stored. So next time my custom judge evaluates a new answer from my child chatbot, it will retrieve the most updated guidelines from the semantic layer. And it will also retrieve past evaluation examples with feedback and hopefully produce a score that is much more aligned with my standards, much more aligned with my feedback. And this is actually very similar to how we learn a new task. So when you are presented with the task, you will record recall the guidelines that someone gave you, or that you understood from feedback that you received, and you will recall past examples of that task. So constructive feedback is what it will get you better at that certain task. The same happens with judges. They will also get better. So the main takeaway here is that LLM as a judge, they are really a powerful tool. But they are only powerful if we include human feedback, if the human is still in the loop.

SPEAKER_00

It's basically it's almost like we are building like pro a stack of probabilistic systems like on top of each other. But that that also goes like um goes on and handled actually raises the questions about like compounding uncertainty and like biases, which we kind of talked about. But so when teams move like from a prototype to like something that is working in a demo or to a production system, what do you think are the key things that actually break in that transition?

SPEAKER_01

So we usually tend to answer the question can AI or Gen AI systems solve this problem in specific. But in fact, we should be answering something more broad. We should be answering can our company solve this issue with the resources that we have. Because when we ignore this company factor and focus only on the technology, on the Gen AI, on the AI, we build prototypes for an idealized environment. This means that we tend to choose clean, structured and consistent data. Then when we reach production, we will see messy, incomplete and outdated data sometimes. We will build our prototypes thinking that cost is not an issue. So build the best solution possible. And I end up using the best models, right? I will use a lot of tokens. But then when I reach production and I see the amount of users that are using my solution, I see costs exploding. And this is where the solution can maybe not be viable anymore. Also, when we are building these prototypes, we don't take latency into account, right? So we build it again with the best models, which are heavy, they can be slow, we use a lot of tokens, we do a lot of calls, and I did this myself. So when presenting the prototype, I had like this uh saved demo, and I fast forward into the answer. So the stakeholders, they actually don't notice the latency, or at the time no one is paying attention to the latency, but then when you move to the real world, then you can get some complaints, right? That you know it's taking two seconds, it's taking three seconds, and you might need to reiterate, right? And the last one is that, and I think the most important one, is that we ignore governance, right? So we think that we don't need to involve security, right? And then when we move to production, we have security teams knocking at our doors asking, are you sure you are able to you should be using this data? How are you controlling the data access? How am I going to audit these agentic systems? So when we build a prototype, ignoring the company data quality, the cost, the latency, and ignoring governance, then what do you get? You get a high performance prototype, but it's not built for the real world. So when you try to move it to production, everything starts to fall apart. And what started as a project with a great potential and high expectations from the stakeholders, it ends up in the 95% of projects that don't make it to production, that wouldn't generate any value.

SPEAKER_00

It kind of seems like uh these AI systems behave more like a living system. And I really like to use the idea of like the behive for this, which they all work together for something uh specific, but everyone has their own tasks basically. Then static ones, basically. They evolve like from drift and then like to interact with changing environments, which kind of makes me kind of makes me like thinking of the monitoring and evaluation on ongoing responsibility, more or less like you were saying. But I guess what I'm trying to ask is how do you move from trust as something that is promised to something you actually measure and track?

SPEAKER_01

That's a great question, right? Especially in these times where we are talking about probabilistic systems, right? So for decades we were used to software development, which is mostly binary. You build something, you test it, it works or it doesn't, and if it works, you just trust the solution. That's straightforward, you can move to production. But then when we are talking about agentic systems, you enter the probabilistic realm. Sometimes the agentic system is right, sometimes it's wrong, and sometimes they are confidently wrong. This is the so-called hallucination. It provides you with a very great answer full of wrong information. For example, we have like this very old case of Air Canada. They learned this the hard way when they had a shed bot that invented a policy and gave wrong information to a passenger. And this information didn't exist at all in the system. So the bottom line here is that when we are dealing with this kind of uncertainty, it's really hard to then reach trust. So how do we actually move from this trust as a promise to trust as a measurable system? Well, we need to account for two things, two steps. The first one is observability. You are not able to trust what you can't understand. So you need to know what is happening behind the scenes. You need to know what your agentic system is doing. So we need to add observability into our system to capture what tools did the agent choose. What data was used to create a certain answer. And with that observability, a black box, it is transformed into a transparent system. You can explain the agent decisions. Transparency is what starts to promote trust. And then you have step number two. Define success explicitly. So you need to define what success looks like for your agency system. And here is the key. There is no universal definition of success. Success can mean different things depending on your use case. For example, a customer facing chatbot has different requires requirements when compared with an internal chatbot, which has completely different requirements when we talk about a health chatbot, a totally different industry. So you have to define very specific and measurable criteria for your agency system. Not it looks good to me. Actually, that's actually something that translates into numbers. For example, our customer support chatbot solves 85% of the tickets without human escalation or intervention. Our customer sentiment analysis solution has an agreement of 75% with the humans. That's a measurable definition of success. So this is what promotes a shift from the trust me it looks good to here is the evidence of how well my system is doing. So with observability and a real definition of success, then you have trusting evidence. You are trusting data, you are trusting a system you can actually measure and understand.

SPEAKER_00

That's very interesting. I was at the same time like thinking of what you were saying, but it is very interesting. And I guess my next question would be in the terms of like the organizational level. That how would you say that can companies balance the speed of innovation with responsible AI development without creating like too much friction, which is sometimes important to have friction, but you know, without anyone. Not too much.

SPEAKER_01

Exactly. Yes. Not too much. I've seen this happening in in organizations that establish a centralized team whose mission is not to build use cases, but to enable innovation while ensuring that this innovation consistently aligns with the company guidelines, with the industry regulations, which could be different from health to finance to retail. And it also makes sure that we are following responsible AI principles. And we have a great example of this: the Gen AI transition phase, right? So three years ago, three years ago already. So three years ago, I was part of the centralized team whose mission was to accelerate and enable Gen AI at the organization level. What does this actually mean in practice? It means that we were responsible for exploring new foundation models family, new tools, new platform features from different providers, and integrating all of this into a platform that was used for development within the organization while making sure that the organization meets responsible AI development. The idea here is that we were the ones creating development guidelines, how to use these tools, right? Providing pre-approved compliant development templates, enforcing centralized governance and security. For example, we had a solution that would look into all AI use cases, and it would be looking for security breaches or compliance problems. For example, it would look into is there any credential exposure? Are is this projects using non-approved tools or models? At the time, for example, we couldn't use lemma models. We were also checking if there was a project which was accessing sensitive or restrictive data, if there was any violation of the AI usage policies. So this strategy was what enabled Gen AI across the company really fast. Because instead of having these individual teams or individual developers going through this process of exploring and integrating and defining guidelines of development individually, we had this central team which would promote fast and compliant innovation by providing a secure and innovative foundation for the project team to start developing their use cases.

SPEAKER_00

I guess that and something that I've been talking with other people also about is that there's often a fear of adding this kind of governments and evaluation because it people think that it will slow teams down a lot. But I guess in many cases, not having them creates, like you were saying more or less, like even bigger risks later, and something that sometimes you cannot go back to. But so my question would be what does it actually mean to scale these AI systems, the Gen Tix AI, sustainably today across costs, uh infrastructures, and sometimes even environmental impact, which is very important, and there's not a lot of talk about that as well.

SPEAKER_01

Scaling AI systems in a sustainable way across the organization actually means keeping your costs under control. And this is especially true for Gen AI solutions where costs can rapidly spiral. You have token usage that can explode as soon as you start putting more gen tic models into production. You can have inefficient architecture that is using all the context window, meaning dumping everything into the context window without having like a filter of what should go there and what shouldn't, right? You could be using really expensive large language models where maybe you don't need to. Sustainable in terms of cost, sustainable solutions could mean if you are talking about simple solutions, it could mean using small, cost effective models. But if we are talking about complex solutions, this means maybe we need to choose a more efficient architecture. But scaling in a sustainable way also means standardization of AI development and even deployment. So this means that we need to move from development silos to a shared platform and infrastructure that operates multiple agents under a single interface. Why is this so important? This is important because it allows real reusability, which will enable fast development and again of course fast innovation and it will also lower maintenance effort. This is really important because if we have a custom solution, if everyone has its own custom solution, then this is really hard to maintain in future. But if everyone follows the same template, if everyone follows the same standard, then it doesn't really matter if teams move a lot or not because maintenance should be done in the same way, because we are talking about the standards. And finally, and very important, is we need to consider environmental impact when building these solutions. Again, we should use small models when possible. We should think about efficient architecture. For example, do we really need to fine-tune a model? This is very costly, not only in terms of money, but of course environmental impact. Maybe we can use a RAG system which is which could be more efficient. We can use caching, for example. This would avoid unnecessary costs. And one thing which is really important, especially when we are talking about environmental impact, is let's measure. Because we can't do decisions having environmental impact in mind without having a measure. Otherwise, again, we don't know if we are going into the right direction. So, in a nutshell, scaling AI systems sustainably across the organization means keeping your costs under control. Focus on development and deployment standardization and consider the environmental impact of your solutions.

SPEAKER_00

It feels like it needs to be that in equilibrium and balance between everything in order to have the best solution in a way that doesn't overwhelm the environment, the tool, and everything around. And it actually goes from like the evaluation, the government to sustainability. But basically, I guess for wrapping up this amazing conversation that we are having, I'd like to ask if there's one mindset shift that product teams need to make to build a genetic system that actually scale. And if so, what is it?

SPEAKER_01

I would say that product teams need to stop thinking about AI products as only the model or only the agent and start thinking about AI products as a system. This contrasts what we are used to for traditional software development where the product is the code, right? But when we talk about agentic systems, the orchestration code is just one component. The real product that we need to be delivering is the system around it. It's the tracing. We need to know what is happening. It's the evaluation, how is the system performing over time? It's governance, what should the system be allowed to do or not? So the teams that scale AI or Gentic AI successfully, they don't think about let's ship a model. They think about let's build a system that learns over time.

SPEAKER_00

I guess if there's one thing from this conversation that I would like to make clearer that it shows clearly, is that the challenge today isn't is not actually building artificial intelligence. It's building these systems that behave well in the real world with real users. And that means thinking beyond these models, like uh into systems, evaluation, government, trust, uh, environment, everything we we've talked and you shared today. But I would like to just stage one and thank you so much for sharing such a thoughtful and practical and proactive and everything, this perspective. This is really, really insightful conversation. And I guess also to everyone that is listening, if this episode gave you a new way to think about agentic AI, AI systems, like share it with someone building in the space and go talk with Joana, go bother her a lot. And basically, thank you so so much, Joana, and again everyone that is listening.

SPEAKER_01

Thank you for having me. Really happy about this conversation. Thank you, and see you in the next episode of Human X Intelligence.