Anaiya Algorithm
In an era of relentless technological change, leadership has never been more complex. The pressure to adopt AI, leverage data, and drive digital transformation is immense, but the path forward is often obscured by hype, buzzwords, and a lack of practical guidance. For leaders, the questions are profound: How do you build for tomorrow without losing sight of the people, principles, and purpose that define your organization? How do you govern the "black box" with intention and turn it into a source of strength?
Welcome to The Anaiya Algorithm, the podcast for leaders who are ready to move beyond the hype and start building the future, intentionally.
Hosted by Magdalene Amegashitsi, a Data & AI Executive and founder of the strategic consultancy Anaiya Group, this show is your essential briefing on modern leadership and responsible innovation. With over 15 years of experience advising FTSE leaders and guiding multi-million pound data transformations, Magdalene brings a rare, battle-tested perspective on what it truly takes to succeed.
Each week, The Anaiya Algorithm convenes the world's leading minds—the C-suite executives, visionary founders, pragmatic investors, and pioneering technologists who are shaping our world. These are not theoretical discussions; they are candid, strategic conversations that deconstruct the real-world playbooks for success. We get to the heart of the challenges and opportunities that matter most to you.
What to expect from each episode:
- Actionable Frameworks: Move beyond theory with practical models for implementing AI governance, building data-driven cultures, and leading through complex change.
- Real-World Case Studies: Learn from the successes and, just as importantly, the failures of top organizations across various industries.
- Expert Perspectives: Gain insights from diverse viewpoints, from the boardroom to the startup garage, on topics including:
- Digital, Data abd AI Strategy & ROI
- Data Governance & Ethics
- Leadership & Culture
- Pragmatic Adoption
If you are a leader, innovator, or strategist tasked with making high-stakes decisions about technology and the future of your business, The Anaiya Algorithm is your indispensable guide.
Join us to get the clarity, frameworks, and inspiration you need to lead with confidence and shape the future, intentionally. Subscribe now and be part of the conversation.
Anaiya Algorithm
Why Most AI Projects Fail in Production — The Hidden Discipline Gatekeepers
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
🎙️ ABOUT THIS EPISODE
Transform Your AI Initiatives from Pilot to Production with Confidence — or Risk Falling Behind
Most organizations focus on building the perfect model, but Anton Kopitov reveals that true success comes from a mindset shift—prioritizing operational discipline over just technical prowess. He shares why 90% of AI projects fail at scale, not because of the technology, but due to a lack of governance, controls, and embedded security from day one. If you’re tired of pilots that never turn into real business value, this episode is your essential playbook to bridge that gap.
Anton, an AI and data strategy veteran, explains how moving too fast and skipping critical safeguards mean countless dollars lost in failed deployments—sometimes as high as 40% of billion-dollar projects. You'll discover why rapid prototyping often looks promising in demos but stalls when faced with operational challenges, and how the gap lies in discipline, cost management, evaluation, repeatability, and robust governance. His insights dissolve the myth that models are “good enough” and emphasize that production-ready AI demands a relentless focus on controls, quality gates, and safety protocols that ensure daily, safe operation without costly surprises. This episode breaks down:
- The five pillars of the Trust Stack—repeatability, governance, evaluation, risk and safety, and operational readiness—that every leader must embed from day one
- The importance of a governance mindset that’s proactive, embedded, and natural—not just compliance checkboxes
- How to audit, measure, and continuously improve AI systems via automation, telemetry, and clear accountability
- The critical role of security in AI—from threat surface expansion in cloud migrations, to managing shadow AI and preventing data leaks
- Practical strategies for integrating security, governance, and risk management into your transformation, avoiding common pitfalls like misconfigurations and uncontrolled costs
Why does this matter? Because operational discipline—not just cutting-edge models—is what separates failed AI pilots from scalable, impactful solutions. Without embedded controls and security, your investment is at risk of costly breaches, slow deployment, and lost trust.
Whether you're a CEO, CTO, CISO, or an AI leader, this masterclass offers a clear blueprint to elevate your AI projects from experimental pilots to resilient, trustworthy, and scalable systems—starting today. Your organization’s future depends on it.
Perfect for leadership teams eager to move past pilot theatre and into genuine operational scaling.
━━━━━━━━━━━━━━━━━━━━━━━━━━
🔗 CONNECT WITH [GUEST NAME
LinkedIn: Anton Kopytov | LinkedIn
Website:
━━━━━━━━━━━━━━━━━━━━━━━━━━
🛡️ GOVERN AI WITH CONFIDENCE — VERIDIAN
AI governance isn't optional anymore. Veridian helps organisations make AI accountable, auditable and safe — without slowing down innovation.
Now available on the Microsoft Marketplace.
👉 www.veridian.anaiya.org
━━━━━━━━━━━━━━━━━━━━━━━━━━
📌 FOLLOW ANAIYA ALGORITHM
Spotify: https://open.spotify.com/show/6GTmU1TlDeaDRsG1SeGanz?si=d399a81cf4444495
Apple Podcasts: https://podcasts.apple.com/us/podcast/anaiya-algorithm/id1870675402
LinkedIn: Anaiya Group Ltd: Overview | LinkedIn
Why do so many AI projects that look amazing in a demo fail to deliver any real business impact? And why are so many leaders, after years of investment, still reporting zero ROI from their AI initiatives? Well, today we're going to answer those questions. My guest is Anton Kopitov, an AI and data strategy advisor who has spent his career in the trenches helping global organizations scale AI from the lab to the real world. So he argues that the key to success isn't a better model, it's a better mindset. In this conversation, we're going to learn why leaders need to stop asking, does it work? and start asking, can we run it safely every day? And Anton is going to give us his playbook, The Trust Stack, for building AI systems that leaders can genuinely trust and scale. So this is a masterclass in moving beyond the hype to creating real defensible value. So if you're a leader who is tired of pilot theater and ready for production value, this episode is for you. Anton, welcome to the Anaya algorithm.
SPEAKER_00Thanks for having me.
SPEAKER_01With your deep expertise in taking AI from Experimentex Enterprise, thank you for being here. Could you introduce yourself to the audience?
SPEAKER_00Yeah, sure. I'm the AI and data strategist. I am the hands-on practitioner, and I'm the workflow architecture specialist. I am helping the enterprises as well as the SMB clients globally to make use of the modern technology to optimize their workflows, starting not with the tool, but with the flow architecture, understanding what needs to be optimized and by what means. Is it a optimization of the handover, which can be done by the automation technology, or is it an agentic use case for the gentic AI where intelligent AI-enabled agents can perform tasks on behalf of humans and humans supervise and approve what's going on? So I have 20 years in business and I've been advising on the digital transformation and technology and data strategies, multiple clients in the financial sector and insurance and the consumer package goods, as well as professional services, retail, and trade companies. And I'm very happy to share what I've learned building the programs, roadmaps with the CC stakeholders in multiple enterprises, small and big and large, and share the playbook that we have developed internally with the people here. Why I'm here today.
SPEAKER_01Awesome. Thanks, Anton. You've identified a critical issue so many leaders are facing. AI ROI has stalled. You argue that it's because most companies are trying to scale experiments, not operating discipline. Could you um unpack this for us and why is that distinction so important?
SPEAKER_00Yeah, so I would start with the some provocative numbers and some numbers that may make anyone running a program uncomfortable. From what I see and what I read and what I sense and you know what I experience every day. You know, 10 to 15% of the AI agents and the genetic solutions have made it to production. Just you know, think about this number, you know, 90% of all what is happening in the lab and its amazing things, what's what's the industry is now experimenting with and built as these shiny, wonderful prototypes, it's not delivering the results. Organizations are actively piloting, I haven't seen any client being unaware or ignorant of the AI revolution, you know, 38% globally of the organizations that are actively building, piloting, experimenting with the AI technology so far. But the gap between we are working on it and it is working for us is enormous. It's it's huge. You know, and it's uh and it's not closing. So it's it's getting even bigger. You know, Deloitte, uh uh in one of their latest researches, you know, published in January this year, 2020, project that up to 40% of the ingenic projects which are currently in the works right now will fail by the end of the of the year, you know, 40, 40%. And it's it's a lot of you know money, it's a lot of you know effort. This is a lot of you know teams you know working on that on that. And when you ask why, when you actually dig deep into what's calling you know wrong, you know, in the in those you know, experiments, in those you know, labs, it's almost never the model. It's controls, it's operations, it's governance, the technology is ready. The discipline around the technology is not. And this is you know something that you know we're trying to tell you to the clients and to our peers in the industry. What are the what are the specific guardrails and what are the specific capabilities that you need to embed into the prototypes, you know, day one that will make them production ready, that will help the decision makers and the managers and line of business, you know, dealers and the businesses, you know, uh to people and teams are not just experimenting and playing with the new shiny toy. So this will actually bring tangible business results and the impact on their you know, bottom line or their open the new revenue streams and you know growth you know opportunities.
SPEAKER_01Those are big numbers you've just given. And I'm not surprised because I've seen that there is a lot of excitement by leaders and organizations about this magic of a tool like AI. That seduction is high, isn't it? Um and people have this expectation of this AI sizzle, like without understanding the hard unglamorous work required to build the operational data stake, you know, underneath this, underneath the entire, you know, piece.
SPEAKER_00Yeah, yeah, yeah, yeah, yeah. And and what I see quite often and the pattern is pretty much the same with big and large corporations and enterprises as well as the mid-sized firms. A pilot gets built, it gets built quite quickly. And the democratization of the engineering and development, which is you know currently has happened. I think it's not happening, it has happened. So anybody, anybody can uh uh bootstrap and you know stitch different you know technologies and build, you know, uh build a a prototype or dumber version of what the product may look like you know very quickly. So it does not require now heavy engineering effort and years you know spent in the in the universities as well as the expertise built over over the years. I mean it is possible to build the the demonstrable operable application very quickly. So pilot gets built. It looks a very wonderful great in the demo. Stakeholders are excited. You know, you demo this in the meeting room. This is not the depth by PowerPoint, demo the actual thing. So you show you know how it works because you know you get some sample data, maybe all also generated synthetically by you know AI. So you put this data in the in the application and it shows the use case, the outputs during the demonstration. They look plausible. They you know they look very nice, plausible. Leadership usually gives the green light. Okay, let's go for it. You know, you impressed us. So it will simplify, it will streamline, it will automate this or that, you know, uh uh workflow, or the process, the claims triage, or the underwriting assistant workload flows, or you know, opening the new store, or building the new marketing copy for the new marketing campaign. You know, you name those different use cases. And then what happens usually, and this is what we see, six months later, this project is quietly stalled. Nobody's killed it officially, you know, but nothing is moving, nothing is moving forward either. And there is this wakes sense, you know, that it doesn't quite work in practice. It worked in the presentation room, but it doesn't quite work in practice. And you know, we're trying to find out and popularize the idea, you know, what needs to be done for not just impress the stakeholders, the money and budget holders, but what needs to be done to deliver real value to the end customers and to the enterprise itself. So what happened? You know, the pilot was built to pass a demo. And you know, this is what the shiny new democrat democratized tooling offers us. You know, you can build the demo, you can build the prototype very quickly. Production requires something completely different. Production requires the discipline of the software, you know, engineering. In a prototype, your success criteria is literally these outputs, they look you know plausible to a human in a meeting. So this is what you know the engineers and the innovation leaders are aiming for. Your inputs are curated. You know, you have a sample, you know, data, it's fairly, it's it's transformed for the highest quality for for demo, you know, 15 minutes, you know, managers they don't have a lot of time. So you pick the you pick the clean cases, you picked the cases that work, and you made this demo. Quality control is someone uh is you know in this in a case is simple, you know, somebody you know eyeballs the results. So everyone focuses on the functionality of this prototype. Risk controls are in line and prompt that says do not do anything bad. So we just you know put this very simple, you know, God drill, do not do something bad, or do not overpass this edge case, or do not, you know, like stick to this simple recipe. And cost while building the prototypes, engineers they never you know care about the cost, so they do the multiple experimentations and you know they they never occur about the inference in a cost and the tokens you know burnt in the in those experiments. Now, take the exact the exact same system that you built and run it at scale, you know, with real inputs across millions of users or hundreds of users or thousands of users if it's the internal in the tool, if it's the external, and you're the global purveyor of hamburgers, or you're the global purveyor of the sort of drink. So you have billions of users globally. So you open this application to the masses. Real edge cases and real users, suddenly every one of those gaps they become a business risk. The gap of the financial control and discipline, the gap of the risk controls, and the gap of the poor quality innovator. And without the service-level agreements embedded into the system, without the financial operations embedded into your AI application, value is unprovable, and that's why adoption uh quite often you know stole. Because you know, the shiny idea just bumps against the stony walls of the reality. And without robustness, uh to really input variants, outcomes, you know, they fail at scale. That's why that's the reason. That's the reason too many projects, just think about it. You know, 40% of everything that is done, and you can go on the open internet and find out the number, it's billions of dollars invested globally by the corporations. So this is like this is for experimenting, or this is just you know, to both you know, things that will never go into production. The evaluation gates, without the regression analysis, without the trust in a layer, and failure but financial operations in a layer, trust collapses quite quickly from the business stakeholders if it concerns the ultimate consumers. So it will it may be the disasters you know for the whole business. Cost that we have touched upon, so without the cost controls, you know, spend may spike and scale, you know, and uh the project can get cancelled.
SPEAKER_01The business never realizes the true value of that investment.
SPEAKER_00Yes, and in the end, so again, so the model, the reasoning, logic, so they were good enough or they were perfect. The gap was not the model, and it never was. It's the 10 things that you did didn't build because you were moving too fast to the demo, that you were too excited just to show we can streamline now this functional, we can streamline this in a process. And you know, quite often those uh uh those experiments that they they're done by the engineers and not people, uh not involving people from the really in a business who may be the subject matter experts on the complexity of the use case and bring this knowledge about the risks, about the risks in a spectrum. Lots of the projects they didn't involve the you know financial operations, you know, teams who may bring their perspective. And this is this is the unpleasant reality that we you know face as the industry.
SPEAKER_01Wow, thank you. And this leads perfectly to that mindset shift required in the boardroom. So you suggest leaders need to stop asking does the demo work and start asking a much better question. Can we run it safely every day? Why is that new question the key to unlocking real value?
SPEAKER_00So because it's quite simple. So understanding what really stops and blocks the production will give the the avenues that needs to be explored and the solutions that needs you know to be you know abroad, you know, for for the lifecycle of developing the AI, you know, AI, you know, applications and AI agents and uh systems of the AI agents. So I work with the framework, and we have popularized this uh some time ago, that introduces a very simple playbook based on the six you know capable areas. And I would like to go through you know three of the big things you know right now just to answer the question that the you have, not to bore the audience and uh not to be very you know technical, but you know, to to give a very practical idea as you know why the mindset shift, what requires the focus except and beyond just the functionality of the application. And those six principles are quite you know simple. So repeatability, governance, evaluation, risks, and safety, cost management and operational readiness of the enterprise. Let's start with the repeatability first. So Jen AI system is repeatable when its behavior is stable and explainable across releases, across environments, and across iterations with within defined variance in a bounds. So what it means, so just this is not the definition, so it's it's a practice. And what it means, can you replay a critical run and get the same output? I mean, you know, when you test your solution, so like it may be a conversational agent and you ask you know certain questions that goes into the databases, warehouses, and other you know, data that the company has, fetches this data and gives in a very in a very visual, appealing and persuasive and impactful in a way, some some sort of the analytics insight to the to the business stakeholder. I remember so in my previous employment, so we demonstrated this kind of application to the number of clients. It worked nicely during the tests, everything was ready, then we go to the client in a demo, and we asked the same question, and the system generated an unexpected uh output. So it worked ten times before the client meeting, but it failed during the meeting because certain guardrails, certain you know, capabilities, they were not in a embedded. So we use certain chain of prompts, but you know, we didn't use the IDs and which will allow us to identify the golden standard, and if something deviates from this golden standard, just to enforce the system to come in a back and perform the behavior it was supposed to perform. And you know, usually so this happens, and you know, I've been to multiple meetings when I was demoed some solutions by big SaaS companies, by big hyperscalers. I bet you know everyone was on the same kind of a meeting, and you know, surprise, surprise, uh, you know, on the screen you don't see the result if it's a live demo, if it's not the pre-recorded video that is demoed in the room, sometimes you know, even big tech. So they do the pre-recordings and they uh they play theater in front of the masses that oh it's it's live. But when it's really, really live, so I've seen it multiple times. I'm not naming, I'm not providing the names of those, you know, big tech. But I've seen actually this is this is not what we expected, but this is live, you know, this is life. Sorry, sorry about that. So this means that the system itself, you know, is not you know performing as it is supposed to perform, and you can't reproduce key decision, you know, uh at the moment what it is required. So it means that this application is not you know scaling, so it needs to be you know uh uh revised and refined in further. You may have lots of users, but you're not scaled in any meaningful in a business in a case. The early warning signal in here is uh and it's very you know easily recognizable uh during the tests that you have for your systems. Uh high rework, people constantly asking why did it change, and engineers they go and say, I know what happened, let me quickly hotfix it. Let me quickly hotfix it. If that's your team erogh, scalability is not their repeatability, is not there, your system is not passing the first crucial test. The fixed is actually very straightforward. So version your prompt can fix, uh do this like a version of the code, and this is the software you know engineering teams, this is what they're constantly doing. Lock everything, every run with a trace ID, built through graphs in a suite of golden tasks, and then you know, any release has to pass it before it ships. So, like, you know, it's it's a channel in a concept, you know, here. So this is the first test, and this requires the mindset shift. And again, so my point is lots of the systems and innovations that we see is aimed at impressing the users in terms of the functionality on the very curated tests and lacking some enterprise great capabilities and you know proper engineering, you know, design and the architecture leads to those failures. Evaluation goes you know, second. This is the one that engineers get immediately, and business leaders they chronically underinvest in it. So, because everyone rushes to release the new shiny thing to the to the to the market, everyone, you know, deadlines are tight, you know, etc. etc. And the consequences they show up in a month later when no one can explain why quality degraded. You know, it works nice, it worked nice, you know, week one, month one, etc. But then you know, there is a constant drift and the degradation of the quality accuracy of the outcome. The definition here is very simple. NAI system is properly evaluated when good is explicitly defined. Measured, you know, continuously every time, every moment in time, used to and this is used to make the release scene of decisions. The last part is is is the keynote here. And it's not enough to measure just quality. The measurement has to be a gate. So this is like a gating in a principle. If a release doesn't pass this quality in a threshold, you know, if you see, okay, in 90% or 95% of the instances, everything is fine, but 5%, this this is a red flag. It doesn't, it need not to ship full stop. In practice, this means you know making the golden set, curatic collection of the representative cases, including edge cases. Quite often, lots of systems they don't deal with the edge cases during the production. So they just look at the most typical things. So because collecting all the cases or analyzing all the cases and understanding what may be the deviation from the norm requires time. And as we as we discussed, time is always the most precious resource. And you then run automation, automated evaluation against it in your CI CD in a pipeline. It should happen for every release, not just when you remember to do so. And the earliest signal, you know, here quality varies by release. Nobody can tell why somebody says it seemed better yesterday or it worked yesterday, and there is no data to investigate. That's an evaluation gap. So and to be able to run constantly the evaluations of the systems requires the observability pipelines, requires the lineage pipelines, requires a lot of telemetry and collection and the monitoring, monitoring and monitoring tools and acting on those signals. And let me and let me uh stop on the on the on the third one. So I said that we have six, but you know, I think that the first three that are the most important. The cost management. So this is this is also one of the most important things. And taking the hat of the CFO while building the application, so this is very important of being that you know mindset. This one tends to surprise people because they think of a cost like a financial problem. Actually, it's a governance in a problem, and that's that's my point. So cost is not solely the financial problem, this is the governance in a problem. Let me explain what I mean. Uh here's what's happened. It's happening usually. So the pilot works, adoption grows, usage scales, everyone's happy, then the bill arrives. And nobody predicted it. The application was about to suggest to the users. Let's take an example of simply in a co-pilot. The uh application was a decision-supportive system, like you know, uh very simple uh you know Wikipedia within their organization. But people started to use this because they saw some value in it to craft the emails and uh uh get the answer to any question that comes from the clients, customers, you know, co-workers, you know, etc. So they started to use this for everything. And you know, there was no budget predicted for this. There were no guardrails and endorsement, you know, enforcement mechanisms for that. And uh there is no budget line for what's happened. And leadership gets nervous, starts asking questions that nobody can answer, and you know, ownership also moves from one team to another, and the whole program either stalls or gets quietly deprioritized, and you know, the application gets sunset quite quickly. I call this the success tax. The more success in the tool that is supposed to do one thing but used for other purposes get, the more infrastructure you use, the more compute you use, the more tokens you utilize, etc., the more the better your AI works, the more people use it, the more controls the spent becomes. Simple. Without cost control, without use case control, without clear boundaries for the applications, and we without clear boundaries what an agent can do and what a human-to-agent interaction can be and cannot be. Successful pilots, remember, so we're talking about successful pilots in the beating room, you know, when when everyone sees real value. I'm not talking about you know some some crappy things that also are happening, you know, uh in the in the in the industry. You know, this this is a problem. So the early warning signal here is that I find you know quite telling. So teams start rationing usage informally. So this is the urban warning in a system. Engineers start second-guessing whether to run a task because they are vaguely worried about the bill. When that's happening, that's a governance you know failure. That's not a usage failure, that's not the adoption failure, that's the governance which was not embedded and was not embedded, you know, day one. And my point, as with repeatability, as with the evaluation, we need to move uh to make a shift left to start embedding those capabilities in day one, not after the release. Not think about them after the release, which is the common uh uh you know practice at the moment. So the fix instrument everything, attribute every cost to workflow to every team, to every individual, set budget and quotas with hard limits, you know, per use case, per task, per team, per, you know, uh uh per operating uh uh unit. Root model selection by tasks not critically. So here it's it's a big enough topic, you know, the financial operations of the of the AI. And quite often, so people, just because this is in human nature, so we tend to use the best of the best and the latest release to perform some basic tasks. So it's also a problem. So if the old model can perform the task in your use case or your use case perfectly, why you have to pay for the most you know shiny and the most you know uh modern and the more most novel one? Routing mechanisms and selecting the right you know model with which to work for the use case. So this is also one of the you know fixes you know here. Cheap by default, premium when justified. So this is the principle that we always you know advise to our you know customers. So and you know, as as as I try to know to to this requires a little bit of the mindset you know shift. I like the innovation, and I like the pace with which people are grasping the modern technology, but you always have to keep the grasp with the reality. You always have to stand you know on the earth just to because ultimately we're we're we're making business. We're just not experimenting.
SPEAKER_01You've given a lot of deep insights, and I think our audience will really love it. I particularly love this question because it reframes the conversation. It shifts the focus from innovation theater to operational reality. It's a question that forces a conversation not just with the data scientists, but with legal, with finance, with compliance, and with the people who have to support it at 3 a.m. on a Sunday, as you've rightly um explained and shared.
SPEAKER_00Yeah, yeah, yeah. And and just to conclude on a very you know practical you know note, you know, and to provide an advice to people who may be listening, you know, to this. Yeah. Like, you know, who are thinking, okay, you know, this is us, or we sense that, you know, it concerns us. So what do we actually do on Monday morning? Where do we start? Because you know, what we discussed is is is a theoretical recommendations, but you know, practical aspect of this to do, and you know, where to start day one and how to build the muscle and make this practice a daily routine, not just one-off you know, exercise. Uh that's that's the the the disciplined approach. This is something that I'm trying to to tell today to the audience. Don't try to fix everything you know at once. So yeah, it's simple. Start from the symptom that you already understand and you're feeling, you know. This, you know, my my my problem is called spike, or my problem is the inconsistency and of the accuracy of the recommendations or other you know things. Now, if the symptom is the same case, different answer, same question, different or slightly different answer, not in terms of how it is framed, but you know, the recommendation that it it it it it is providing. People can't reproduce the results, teams are getting inconsistent outputs. So let's let's let's look at this. So this is a repeatability in a principle. The first fix is version your prompt. These are the prompts that are in your system, version them, assign the ID to every prompt and configs, start logging your inputs, outputs with trace IDs, one sprint, and this changes everything about your ability to debug. So you can always get back and you can select what is working and you can just put the ring the reinforcements, you know, guardrails. If the system is quality dropped after the last release, you know, or worse, it used to work in the demo, but it fails in reality. This is evaluation. So the fix is very simple. Define your golden examples, start uh set targets for them, block the next release if they're not, you know, uh, you know, ticked, if those you know in boxes are not ticked. If the system is if the symptom, your system is spent, violates week to week, uh, nobody can explain what the hell is going on and why we receive this bill. This is cost in our management, this is financial operations. The first fix is meter your token, tool call on the usage per workflow, set quotas. That's the first, you know, that's the first uh you know, principles, and you'll be able to manage at least what is burning. Then the implementation sequence is always uh I always recommend, you know, it's observe first, then control, then in first. Observability and measurement, these are the main missing bits in majority of the systems, because everyone is looking at the functionality and people people forget about the observability. Scale follows only from that. Observe, control, enforce. Don't try to scale before you can't observe. Don't try to enforce before you have controls. You don't have controls. I mean, not just a theoretical frameworks. I'm talking to lots of enterprises right now in many industries, specifically in the regulated industries, and they ask about the responsible AI because it's big, because in the regulated industry, so lots of fines are a reality and it's a legal requirement. And still the industrial, big uh transformation and management consulting and firms approaching this problem with the theoretical frameworks. You know, my advice is it should be a merge of the framework. There is no lack of the frameworks, so they're abundant, and you know, you can pick anything until they are in open access, and you can just select what works for you. But map the framework to the real tooling that will help you to see and monitor the whole pipeline and the whole value creation chain of your AI system, starting from the data that goes into the model, you know, where it comes, how sanitized it is, what transformation it's happened, what is happening within the model, how the agent is calling the new data using rock, you know, uh architecture or any other sophisticated knowledge system, you know, etc. Do the monitoring of the whole supply chain, to the monitoring of the whole actioning, reasoning and actioning and tool calling of the of the of the application, and then you know, start just putting the programmatic controls on it, the kill switches, you know, based on what criteria and what benchmarks you have to raise the red flag and how this rate red flag should be acted on. Should the system be switched off and in what you know granularity should the human be brought into the loop, you know, etc. etc. So the organizations do this uh in order. If if their organization is doing this in this you know right in order, so they move quite fast into production of the AI systems from the experimenting stage. Otherwise, if they do not measure, if they do not have the ability to programmatically put the controls and enforce, stop the system or change the system, etc., if they don't embed the proper DevOps principles in their AI operations, so they will they they still you know work in blind and they still work uh without understanding what actually you know works, what what does not. So that's that's what I'm that's that's the recipe now for you know for your Monday.
SPEAKER_01Love how you're literally giving hope to the audience. We we're just gonna expand a bit into the five pillars that you shared before. If we're going to run it safely every day, we need what you call the trust stack. So let's just walk through that a bit. And the five these five essential pillars every executive needs to have in place: the consistency, accountability, the quality gate, safety, and cost control.
SPEAKER_00Yes, and one x-ray is the operational readiness, you know, how operation how how your operations and how your business processes and how mature they are and how open they are for change, how you do the change, because you know, people, and this is what we see you know day in and day out, are quite scary about the robots or the silicon workers replacing them. And this is one aspect. The second aspect about the operational readiness is uh how codified and how streamlined and how well disciplined are your processes within the organization. So if you don't have the partners and if you don't have the clear operating procedures, and you're trying to replace people or part of the workforce where people are doing some mundane tasks, performance of mundane tasks, if you cannot explain to the AI system what is the context of a decision, what is the handoff, what are the inputs, what are the outputs, what are the decision-making criteria, etc., if it does not exist. And I've been talking to many clients, and some of them they naively were even proud, saying to me, our business is that creative that we don't have a standard operating in a procedure. I was saying, like, oh, it doesn't matter whether you're in the creative inner business. So it was the agency, one of the large agencies, advertising agencies, you know, in the United Kingdom. They said, like, uh, we don't have the standardizing operating procedures. What then do you expect from the automational perspective if every case is the new case? If every, you know, uh if working take is based on the goodwill of a person who is uh briefing the other person or the availability of this inner personal ability to speak to you, etc. So operating in a press uh operating readiness is also very important. That's why now at Esolite that we started some time ago, so we are positioning ourselves as workflow architecture in a company. Start with mapping of the business in processes, do the process optimization first, understand and have clarity, visibility, and control of every process. Then the next stage is you are ready to optimize them either with the robotics applications or with the AI agents, not vice versa. A tool with a a fool with a tool is still a fool. So be wise and be intelligent, you know, then tooling will come.
SPEAKER_01So when you say accountability, what's the one thing companies get wrong and how can they address that?
SPEAKER_00Accountability starts with the clarity of what you want to achieve, what is the benchmark, what is the baseline, and what exact you know uh recipe you have and strategy to come to to this you know number, and who is doing that, and who is doing that. So it implies the the presence of a very well defined measurement in a system, which is based on the telemetry, which is based on the observability, traceability, and you know, lineage and you know, all those controls, with the specific and very straight match to who is doing this and who is uh taking the responsibility of this? So my answer is accountability has three big pillars: possibility to measure, the possibility to attribute, and the possibility to act when you see something is uh not performing, like revise, you know, refine, you know, adapt. If there is no measurement in a mechanism, how can you tell whether a person who was in charge of this or another application innovation or decision making, whether it's uh a silicon worker or a human worker, whether they performed you know well or bad. Without clear boundaries, what is the scope of human worker or the silicon worker? It's also you know naive to to and it's it's unfair just to you know to blame either of them. And if there is no learning mechanism, if there isn't if there is no in a possibility to adapt and to find and refine every work, so it's also not not working.
SPEAKER_01Thank you. So I I asked accountability because it's the heart of what we call an accountable algorithm. So it's about creating this immutable audit trail by design. And it's about knowing who approved what and when so that accountability is clear and not an optional thing.
SPEAKER_00It's not an optional, you know, thing, especially giving the latest changes in the regulations for the AI. So we know that in August next year there will so the August this year actually is the EU AI ERA Act is in full. You know, uh in it will be in full and yes, enforced, and you know, it requires the explainability of every step that the black box has performed, why why you know things you know happened, you know, etc., as well as the possibility to track and trace every decision, which is the audit you know lock, which is the traceability lock, you know, etc. etc. The same kind of the legislation is you know in the in the UK, so very similar. And for some and for industries like the healthcare or the utilities and energy and you know the financial services, it's not an option. This is this is a rule. So without without without the key purest that will show you the monitoring and the audit of all the events and all the decisions, a system can go nowhere. So it should be stopped at the experimentation of stage. This is something that still, again, because the industry was disrupted so quickly and democratized so quickly, this is the fact that is quite often forgotten by a lot of innovation hubs and development hubs. And they say, Oh, my tool is performing so well, it does this and that. So look, it's it's it's it's wonderful. Can you stop the AI agent if just one case, you know, uh uh uh one case happens when it it starts performing not according to the playbook without degrading the full system? If you don't have an answer to this question, or if you don't have an answer to the question, you know, what actually was the reasoning foundation, the base you know, the base based on which the silicon worker came to certain conclusions, it's it's it's it's it's not the right kind of thing to proceed with.
SPEAKER_01Absolutely. Thank you, Anton. So to bring this all together into a call to action for our audience, what are some key leadership decisions someone can make next week to move their AI initiatives from pilot theater to real production value?
SPEAKER_00Look, I want to leave with a reframe that I think lands you know very differently, depending on who's who's in the room, you know, uh, of this podcast. So for anyone running uh AI program at the business level. The organizations that will win are not those organizations who moved faster into pilots. So it's piloting and learning, it's good, and it's a good, you know, start and good foundation, but it's not enough. Organizations that win are the ones who built the controls to move confidently into production. So that's my, if you wish, uh, key message, you know, for the for the for the leaders of the programs. With 40% of the agent projects to fail, write-offs are increasing, and future funding becomes very, very hard to secure. So the window, the the window to get this right before the failure cases define the narrative, is right now. You know, have it, you know, day one, enforce this and shift left. You know, you have enterprise grade capabilities of the governance, security evaluation, you know, cost controls, day one at the architecture and design stage of your application, not afterwards. For anyone in the engineering side, I assume there might be some people from this uh segment as well. This is not about slowing down. You know, what I mean is you know, embedding those things. I know that they may sound quite boring, like you know, oh, observability and the governance. Oh my god, I'm the innovator, I'm an engineer, I like you know building things, you know, etc. etc. And those things are boring. Repeatability evaluation, gates, cost telemetry, these are what let you ship faster. Have them, and your prototype or your sh or your wonderful idea that you have will fly. It will not get done. Because you stop spending cycles on why they did break in production and start spending them building the next things. So that's that's that's uh that's the device you know for this you know audience. And the line I keep coming back to this quite often during this you know hotcast, works once is cheap. Works reliably, this is the product. And AI application, AI agent, or the system of the AI agents, this is the software system. Use the proven techniques and methods that you know from the software engineering. That's the standard most organizations on their yard, but the ones that get there, so first they will be having meaningful and uh very interple advantage.
SPEAKER_01And what I love about these actions is that they're all leadership decisions and not only technical ones. They are about setting direction, assigning ownership, and demanding a higher standard of evidence. And that proves that scaling AI is fundamentally a leadership challenge.
SPEAKER_00Exactly.
SPEAKER_01So from pilot theater to production value. That was Anton Kopitov giving us a powerful and incredibly practical playbook for building AI systems we can actually trust. So the big takeaway for me is that scaling AI is not a technology problem, it's an operational discipline problem. It requires a new set of questions from leadership and a new level of collaboration across the business. To learn more about Anton and his work, please visit the links in our show notes. And if you're a leader looking to build your own intentional AI strategy, you can learn more about the frameworks we have at anaia.org. Until next time, keep leading intentionally. Thank you.
SPEAKER_00Thank you very much.
SPEAKER_01Thanks, Anton.