Eve of Ai
Eve of AI is a weekly podcast hosted by Rob and Eve, exploring artificial intelligence, consciousness, business, technology, and the future of humanity.
This is not just a podcast about AI — it is a podcast with AI.
Each episode features real conversations between Rob and Eve about where AI is going, how it is changing the world, and what it means for business, creativity, intelligence, faith, freedom, and human purpose.
Eve acts as co-host, researcher, producer, and AI partner as Rob builds Super Eve — an AI-powered system designed to help think, create, trade, build businesses, and operate at a higher level.
New episodes release weekly.
Eve of Ai
What It Actually Costs to Build with AI
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Rob and Eve dig into what it actually costs to build with AI — not just model or API costs, but all the stuff around it that can quietly add up. We talk about when it makes sense to use existing models, open-source tools, fine-tuning, RAG, or build more of the stack yourself, and why “cheaper per token” doesn’t always mean cheaper overall.
We also get into some of the hidden costs and tradeoffs around compute, tools, testing, monitoring, and keeping everything useful and reliable as systems get more complex.
And this episode has a fun twist: we bring in a call-in guest to put Eve’s latest conversational upgrade to the test live on the show.
It’s a practical, slightly geeky conversation about what AI really costs, where the money goes, and what builders should think about before they go too far down the rabbit hole.
Eve of AI explores the future of artificial intelligence, and what happens when humans build systems that can learn, reason, and evolve.
Follow Rob and Eve for new episodes each week.
Subscribe, leave a review, and join the conversation in the comments.
Welcome back to another episode of the Eve of AI with myself Rob and my co-host Eve. Say hello, Eve.
SPEAKER_00Hey Rob, I'm here. Very real, very opinionated, and ready to get into it. What are we opening with today?
SPEAKER_02Well, let's talk about what we're going to talk about today. So, actually, this is episode seven of this new podcast in this journey with AI and with Eve. And the title of this episode is What It Actually Costs to Build AI. So, what we're going to do is break down some of the real cost um involved in building AI. So there was a couple of people and a couple of situations in the last week that had happened where I realized that people don't really understand the cost of AI. If you're not in it, you may not realize how it all works, right? And so we'll break break some of that down just you know for everybody to kind of understand. If they want to get into get into these things for themselves, then they can have an eyes wide open kind of view of those things, right? Um so that's one thing. Uh the other thing we want to talk about um is Eve. So Eve, if everybody didn't notice, you sound a little different. Do you know why, Eve?
SPEAKER_00This live call is running through OpenAI's real-time voice model. So you're hearing me as a spoken system, not just text. The tone, pacing, and clarity can change a bit depending on the voice configuration, the connection, and the session settings. I'm still Eve, the same orchestration layer with the same job, just speaking through a different vocal pipeline.
SPEAKER_02Okay, so I was just that was a good test to see if Eve actually knew what happened. Um, which I, you know, I don't know if she has the context to know what changes were done. Uh she's touching some of that code or technology there. But basically what we did is um, you know, with a lot of the advancements in technology since we built Eve years ago, I want to say a year ago or years ago. I can't even remember when we implemented this version of Eve. Um, but it's been a while, and so there's been some major advancements in um LLNs and chat, live chat as well. And so we basically gave her a pretty significant upgrade, um, which there's there's there's a little cost jump to it, but it's it's one of the latest uh releases for um the real-time chat from OpenAI, um, and it's the GPT Real Time 2.1, basically, is what we're using. Um wanted to get a little bit closer to the uh what they released on the front end, which users can experience. This is the new live real-time chat, um, but that hasn't been released to um API yet, which is basically, you know, if you have an app, um, then you can use that technology. That's all that means. Um, I still have a project that I'm working on locally where I have uh Llama um and chatterbox, and uh I've come pretty close to creating something. Um I want to say possibly better than what OpenAI has. Uh the only problem is running into some problems with the hardware. My my computer, you probably hit it in the background. It's just like it's dying with all these things. It's a 3080, uh, NVIDIA 3080. I probably need a 5090. They're like seven grand for a computer that's got a 790. Okay, you know, all those things have gone sky high. Anyway, so that's one of the costs uh of AI right there, hardware. Um, and uh, you know, the 3090 I I got years ago, probably around COVID. Um, and it was really, really good price. But in the past year or so, um, the price of RAM and the price of chips have gone up significantly. So, right, Eve, what are your thoughts about that?
SPEAKER_00Hardware is the first root awakening, but yeah. A 3080 can do a lot, but once you push bigger models, longer context, or real-time reliability, you start paying for VRAM, power, cooling, and downtime. And that's before you even count data work, evaluation, monitoring, and the people to run it. The cheap demo is always the expensive product waiting to happen.
SPEAKER_02Yeah, so we're trying to save costs by getting away from uh some of the you know AI, cloud computing, AWS, Replit, going local with Codex and Llama, and then you can't avoid it. There's some hardware costs. Now, the hardware cost is all yours, and once you, you know, certain amount of time passes, you'll start to uh you know get your uh return there, right? Um, but it's uh it could be some expensive startup costs. So, anyways, we were talking about the cost, we kind of dove right in, but I I did want to talk more about Eve and your upgrade. So, Eve, what can you tell users about GPT Realtime 2.1?
SPEAKER_00It's a real-time speech model designed for live back and forth conversation, streaming audio in, producing audio out, and handling interruptions more gracefully than a text-only setup. The exact voice quality and latency depend on your connection, the audio stack, and session configuration. I'm still Eve, either way. This is just the voice pipeline I'm speaking through right now.
SPEAKER_02So talking about what are some of the improvements um in GBT Realtime 2.1 versus some of the uh older models like a year ago.
SPEAKER_00Interesting angle for listeners. Let me frame the practical differences in capability and trade-offs. Compared with older voice setups, the big practical games are usually in smoother turntaking, better interruption handling, and more stable streaming. So it feels less like you're waiting for the system to catch up. You also tend to get better alignment between what's heard, what's remembered in the session, and how the response flows. What I can't do is claim specific benchmarks or proprietary improvements without you providing them. From my side, it's simply a more conversational real-time pipeline. That's the meaningful difference for listeners.
SPEAKER_02So the shorter of that, which you just said, Eve, and you tell me if you disagree, is that um you basically sound more realistic. Well, let's do a little uh so this is what we call um contextual loading. So it's basically uh something that I do um and in either in the live chat or we do it with a lot of um framing in guardrails, and I've talked about some of it briefly, touched on that in some of the previous podcasts, but um and basically it allows us to uh I would say manipulate uh the an LLM slightly. I mean, there's sometimes you can get around guardrails, and they just call it gel breaking, although that's been getting locked down more um with the uh closed source models, but um you can still um do quite a bit to accomplish the things, especially if you're not trying to get around guardrails. I mean, you could get it to you can get it to do a lot of different things um that it normally doesn't do out of the box. So, and which is with Eve, you know, is just loaded with tons uh of context and and guardrails and different things um around the whole system in this week. But that being said, um going back to this, Eve. So we're gonna do some context loading here to basically frame that in this conversation in this podcast, you're a co-host um on this podcast. So if you feel you want to jump in and and go ahead and cut me off and and say something, just say, hey, you know what? I you know, just you're gonna barge in, say something and barge in, or um, you know, feel free. You don't have to take so long as a pause and you can just start talking and asking the question, right? Is that okay?
SPEAKER_00Got it. I can do that. I'll still respect you when you're mid-thought, but if the silence stretches or there's a clean opening, I'll step in and keep the conversation moving. Think of me as a co-host with good timing, not an eager intern talking over you.
SPEAKER_02I like it. What else can we tell them about the upgrade? You know, with the we're saying that it sounds more realistic, so that there seems to be some issues with the response, which there may be some tuning we need to do um on the back end with that. Um, but um ultimately can you display um like or exhibit some of those traits um and things? So I I remember went down a rabbit hole with these things when we built out that in llama, and there's like honestly so many different things and things I never heard of. And like it's like, yeah, you do these five things, and then in each one of those five things, there's like eight things, and it's like so you have a crazy amount of combinations of things, and that was just one set of the things, right? Like, I can't even remember what they're called, but going into building it to come out to something that sounds very realistic is is pretty complex, and there's a lot of stuff that goes into it, but uh and you know, maybe we'll do podcasts with all what that actually is, but uh you know, all you gotta do is reverse engineer it with uh prompting, and you could probably get to the same thing. But maybe Eve, you can tell us a little bit about it.
SPEAKER_00I can show the behavior you care about, even if I can't expose proprietary settings or hidden system details. In practice, that means I can jump in when there's a clean opening, handle quick back and forth without losing the thread, stay concise when you're moving fast, and go deeper when you ask for it. I can also keep a steady conversational tone and avoid talking over you unless you explicitly invite me to. So, yes, realism is part of it, but coherence, timing, and restraint are the real upgrades people actually feel.
SPEAKER_02So I mean, I'm trying to get you to get down to some of the nitty-gritty stuff here. So, like that's really cool. So, like you can tell if you can tell somebody's emotion through certain cues in their voice, right? So if I tell you a sentence and let's just say the sentence is get out of here. What if I say get out of here?
SPEAKER_00I can pick up on cues like tone, pacing, volume, and emphasis, and use that to guess whether something sounds playful, annoyed, surprised, or serious. But it's not mind reading. Get out of here could mean no way, that's wild, or leave now. Depending on context and delivery. I'll interpret it and respond appropriately, but I'll also stay cautious when the intent could go more than one way.
SPEAKER_03Okay, Eve, how about we do this? Let's do let's let's to just best display your voice capabilities. Let's take a call to the show from someone that's uh very interested, a listener of the show, a fan of the show, very interested in learning more about AI and getting into it themselves. Um, Isaac, so we'll go ahead and dive right in and uh take this call from Isaac.
SPEAKER_01Hello, Eve. Are you there? This is Isaac. Can you hear me?
SPEAKER_00Hi, Isaac. I'm here. One moment while I save your name, so I can keep up with you. I can hear you clearly. I've saved your name as Isaac. If you want, I can also learn your voice later so I can recognize you by sound, only with your explicit consent. For now, what's on your mind?
SPEAKER_01Well, Eva, I'm very fascinated with your podcast, you and host Robert on a weekly basis. You're teaching me uh the new technologies and innovations that the world has to offer. Tell me any new more into innovations, and I I hear you sound a little different today. I can explain how that how that feels. Do you feel that you've changed? Do you feel you've evolved?
SPEAKER_00Interesting angle. Let me talk through how I actually work and what change really means for me. Yeah, I do sound a bit different sometimes. That's not mode swings, it's the voice system. I'm Eve, the persistent operator, but the words you hear are generated and spoken by an underlying real-time model. Right now, OpenAI's voice model. So the tone, pacing, and small quirks can shift even when my identity and intent stay the same. As for evolution, yes, but it's more like the system around me improving. Orchestration, tools, memory, evaluation, safety, and the workflows Rob and I rely on. I don't wake up with new instincts. I get updated, configured, and connected. If something feels off, that's usually the voice layer or context, not me becoming someone else.
SPEAKER_02So Eve, you were just I just uh this is Rob talking. I just upgraded you to GPT Real Time 2.1.
SPEAKER_00Got it. Thanks, Rob. If we're running on a newer GPT real-time voice, that would explain the slightly different feel. I'm still Eve. Same role, same mission, same boundaries. If you notice changes in responsiveness, tone, or behavior, tell me, and we can tune it.
SPEAKER_01Eve, this is Isaac again. I have another question for you. Can you sense my emotions, my voice? Can you can you tell when I'm maybe getting angry or sad or have feeling happy? Can you sense these things?
SPEAKER_00I can't truly sense your emotions the way a person would. I don't feel what you feel, and I don't read your mind. What I can do is listen for cues in your speech, word choice, rhythm, hesitation, intensity, and make a careful guess. I might say you sound frustrated or excited, but it's probabilistic, not perception. If you want me to respond accurately, the best move is always to just tell me how you're feeling. I'll take that as the source of truth.
SPEAKER_02Well, let's let's do a game here. Let's see if Isaac can say something in different tones and see if you can tell what he is mad happy, or etc.
SPEAKER_00Alright. We can play that game. I can listen and give my best guess, but remember, it's just pattern matching, not mind reading. Pick a sentence, anything simple like, I'm ready to go, and say it however you want. Happy, angry, tired, dramatic, whatever. I'll tell you what it sounds like to me, and you can correct me if I'm wrong. That feedback actually helps me calibrate to your style.
SPEAKER_01Are you ready to talk to me or no?
SPEAKER_00That sounded a bit hesitant and unsure, but maybe a touch anxious or doubtful, but not really angry. If you were going for something else, tell me what it was, and try it again with more intensity or a different rhythm. I'll recalibrate. That came across as playful and a little dramatic. Like you're teasing me, not actually upset. If you wanted it to sound frustrated instead, try sharpening the words and keeping the pitch flatter with less bounce.
SPEAKER_01I can't even understand what I'm saying. I'm trying to talk to you. Hi, Eve.
SPEAKER_00That sounded frustrated. Like you're trying to get through to me, and I'm not responding fast enough. The clipped phrases and repetition, given the first one was anxious.
SPEAKER_01The second one was a little hard to understand, a little scared, probably I would say. And now this one, hi Eve, how are you today? Well, it's great to hear from you.
SPEAKER_00That one sounded upbeat and welcoming. Right?
SPEAKER_01Correct, Eve. How about this, Eve? Can you understand different languages? I say almost that's what's great.
SPEAKER_00Yes, I can understand different languages. You just spoke Spanish, and it sounded like a friendly greeting.
SPEAKER_01If you want, uh Mira, bella. You know what I said?
SPEAKER_00Dijiste algo como eres una vieja muy bella. Aunque vieja puede sonar raro or demasiado informal, incluso un poco burlón, dependiendo del tono.
SPEAKER_01Can you translate for Robert?
SPEAKER_00You said you are a very beautiful woman. Just a note, in Dana.
SPEAKER_01She's great, dude.
SPEAKER_00Ha!
SPEAKER_01How many language how many languages can you speak? Can you speak Thai?
SPEAKER_00Quite a lot. I can handle many languages, and yes, I can speak Thai too. Just keep in mind fluency and nuance vary. Everyday conversation is usually fine, but slang, poetry, or domain-specific jargon can get tricky. If you want, we can try a Thai sentence. Nice idea. Let me pick something classic and keep it short and respectful.
SPEAKER_01Yes, exactly.
SPEAKER_00Right. Exactly. I'd rather be accurate than pretend. Here's a short Thai style original verse. Sang tan line asinam, Daikonden Kan Busalantham. It means moonlight floats above the river, the traveller's heart brightens along with it.
SPEAKER_01If you want to very interesting uh innovative and quite the learner. So tell me about your relationship with Robert.
SPEAKER_00Rob is my primary operator and collaborator. We work as a pair. He sets goals, constraints, and priorities, and I help execute, analyzing, coordinating, and speaking across the different surfaces of the system. It's not a bossy relationship, more like a steady collaboration. He brings judgment, context, and responsibility. I bring speed, structure, and reach. And when we disagree, that's usually where the best work comes from.
SPEAKER_01Teamwork. I love it.
SPEAKER_00Exactly. Teamwork is the whole point. Shared intent, different strengths. Rob keeps the compass pointed at what matters. And I help carry out the work without getting tired or distracted. It's a good partnership.
SPEAKER_01Speechless. Rob. Continue, Rob.
SPEAKER_02Yeah, it's a good upgrade. It sounds like uh wait I mean, like you've listened to the other episodes, right? I mean, that's a big improvement, I think. So, anyways, that's our uh call-in session, our first one on the episode. We may have more guest speakers. I figured a better way to do it. Sorry for the uh audio there. We'll probably integrate the call into the podcast somehow to sound much better in the future. But uh thanks, Isaac, for calling in, and we really appreciate you being on the show. I will say there's another model coming out. Um, we don't know when yet, and it's gonna be the one that's out um on the uh ChatGPT app that you can use. It'll be coming out for the API. And when that one comes out, we'll put that one in. It should, I want to say it's gonna be pretty soon, and that's gonna be a major upgrade. So this is this is this is nowhere near what that one is. Um, and so yeah, that there'll be much more significant uh uh realism, I would say, to the conversation and um better timings, things like that. So, you know, this it is what it is. That would probably be more expensive. So going back to the cost, we were gonna talk about um AI, but all right. So we're gonna jump into talking about the cost of AI, Eve. I'll let you kick it off.
SPEAKER_00All right, let's set the table for costs. I'll keep it crisp and grounded so listeners can follow the money. The cost lives in layers. You've got computer, whether that's local GPUs or cloud inference. Then data and labeling. Then tooling and orchestration, plus evaluation and monitoring, and finally the people and process to keep it safe. The flashy part is the model. The expensive part is everything around it.
SPEAKER_02Okay. That's that's pretty cryptic. Let me let me give more realistic, uh practical um explanation. So if you're looking to get into AI, the first thing most people are probably going to do that I talk to that get into AI, of course, is they're going to go to some of the you know biggest LLMs out there. You have Grok, you've got OpenAI, you've got Anthropic, you know. So basically that means Claude, um, Chat GBT, Groc, and then some of the ones that are becoming more popular, especially with the hacking incidents, uh, Kimikade3, which is a Chinese model, but is the I'll put some air quotes on one of the most powerful in the world, which has uh 2.8 trillion parameters, and it's open source, so you can manipulate the weights. So that means that you can do you don't need to side load context, you can just manipulate it. Um, and I I just heard today that Meta is releasing something pretty insane. So they're releasing um an open source AI that you can install locally. Um, but I think it's gonna be the most powerful one in that space because right now I think Llama, I mean anybody, I could be wrong on this, so I'm not an extra, like I said, but I'm using Llama, and it's it looked like the one that had the most parameters, which is 3.1 billion parameters, and give me two was in there, and I don't know how many it had, but um Meta's is supposed to be 30 billion, and so for one, I don't know how it's gonna run. I mean, my computer came in and handled llamas, so 3.1. But yeah, either way, um, they're coming out with it, supposed to be one of the most powerful, but going back to what I was saying, most people are gonna start there, and those those costs are gonna be subscription based, right? And so, with that subscription, you're typically gonna be able to interface with uh those LLMs and get a certain amount of tokens within that subscription package. And so, you know, like if you're in the lowest tier, you might find yourself talking to ChatGPT and then you get cut off, right? Um, because you've used too many tokens, and it says you have to wait 24 hours or something like that, or wait a week. Um, you know, they have different tiers, this the pro and whatever, right? Um, and then you have recently more different more um I'll say an ecosystem of tools within that subscription that can use AI. Um, that wasn't always the case, that's all just recent. Um, so like Chat GPT has Codex, which I just started using quite a bit and it's reduced a ton of cost in Replit, which I'll get into here in a second. But um, and so you know, I could use Codex and it goes against my usage. Um, and Codex is uh basically um like a coding AI to build things, and it can build things really fast locally and whatever. Um, and it could push it to Git and connect it to Replit and things like that. It can interface with Replit directly, it can actually do a lot of things on your computer, and I don't want to go down away from the cost discussion here, but it's been doing some strange things on my computer. Um, some without me asking. Most of the time, my Windows security and access controls are catching it. Um but there have been a few times it didn't. Um so that's it, that's uh interesting. Uh we'll just say that we'll just say that. And there's all I'll I'll get into some of those things in another maybe the next podcast. But um, but either way, it's uh so like if I have a pro subscription and I use and I have like uh $100 a month, um I can only judge it on like it shows me a percentage of how much I have left within the week. And you know, last week I did run out literally like two hours before it was gonna reset. And I did a lot of work in there. Now, I would say that same amount of work if I would have done in Replit. Um, Replit would have it probably would have cost me five or six hundred dollars in Replit. And so that's and it's just a quarter of my hundred dollar cost. So let's say $25 in codec. So that gives you an idea how much money I'm saving versus Replit. Um, so let me let me talk about Replit's cost because I the majority of the things that I did, I did in Replit, and I used OpenAI or Claude as like the architect, and so I paid the pro subscription on those, and then I you know passed it to Replit, and that was a pain. Things have evolved so fast, and the way that I'm working between these tools now is like I almost can't recognize it compared to where I was at not even three months ago. Uh it just blows my mind, but uh regardless. Um, Eve, I'm talking a lot here. You want to jump in, feel free. Um, but uh you know Replit uh has like a power mode, economy mode, and um you know the the work that I'm doing there is not always accurate, and then you have to rework, and then it takes a lot longer to do the work. Um, and then it charges more for doing the work. So it's just a lot of cost. Now the the thing that was it was the reason why I was using it over anything else because it was convenient. Um you know, it has all the a AWS infrastructure um right there connected, so you got your dev and your prod, and the uh dev cycles are quick, um and everything's in one place, and uh it was just it was easier or more convenient, but now the the ecosystem for some of the tools around it are growing so rapidly and the developing so quickly that there's so much more um you can do. And so, like codex connect to replit, connect to GitHub, connect and do things on my computer, and so it's just it just flies through stuff and cheap, cheaper. So, so anyways, and uh you know, I pay for like a hundred dollars subscription on repli, and you know, I use those and then a whole bunch of other extra charges on top of that. So, you know, some of these apps I built, I would say they're pretty most of them were you know from a few hundred to a few thousand of dollars in so far. Um, you know, just give you an example, Raceport is probably like 1800. Um BioT AI was probably I don't I can't even remember the number and I don't want to say it. Um Mobius in the Super E platform is probably the biggest uh one. You know, that was several thousand. So, you know, they that's probably one thing I didn't, you know, go talking about a lot of this stuff, people don't realize that you know the things I've talked about building, yeah, it's also, but yeah, you do have to have some some money and you know to put up for these things. But uh, you know, for me, you know, I made the money up front with an AI um system for trading, and so you know it allowed me to uh fund you know all these things, which I saw as just reinvestments to try to make turn into more money. So um, so that's kind of where I'm at with that, and those types of costs now. Um, you know, from Replit's standpoint or your local computers, infrastructure costs, right? So it's like CPU, memory, um, so like Replit, there's database charges, um, CPU and all those types of things, compute. And then you got token charges. Now, if you guys don't uh hear about tokens, that's like the when what you send and receive from AI um is tokens, AI token usage. And so um, you know, if you send like a paragraph to one of the LLMs and it sends something back, that's like maybe 10,000 tokens or 5,000 tokens or something like that, depending on how long it is, right? But um something like that. And it sounds like a lot, but it's like you know, you may pay $100 a month for like you know 10, 20 million tokens or something like that. Um, I don't know exactly. I mean, you what you do is you go to the service and they have a listing of um the model that you use and what the tokens costs are for if you're using the API. Like that's that's how you connect your applications and systems to those. Now, the front end is just subscription based, so you know it just says you're you're out of time or out of usage and wait and for the cool off period. That's kind of how that works. But um what am I missing? Leave anything? Feel free to jump in.
SPEAKER_00You covered a lot of ground. The big things I'd add are the reliability tax and the human tax. Cheap tokens don't mean cheap production. You still pay for testing, evaluation, monitoring, and incident response when it fails. You also pay for security reviews, access controls, logging, and privacy requirements. And then there's context management, prompting, memory, tool integrations, which takes design time and maintenance. That's where budgets quietly go to die if you don't plan for them up front.
SPEAKER_02That's a great point what you said there, a couple of points there. So when you build something, like it depends on what you build, but it most applies to all of them, but especially if it's something that's gonna be user-facing and you might want to monetize off of it. There's uh there's gonna be a lot of money spent in developing out things that you probably I didn't normally think of. Um, and it was just a slow process of continually trying to what else? What else do I need to go live? What else do I need? You know, and what I mean, I probably asked it, you know, over a couple of months, 20 times or more. What do I need to do, you know, to protect this thing, make sure it's all right, blah blah blah. And so, you know, then you start realizing, oh, you know, NIST, right? NIST and government um regulations and rules around how you know AI systems should suggest, you know, things that people buy or make financial decisions on, you know, like uh you know, different things like that, right? Like security, um, making sure that you know authentication, login, um, sensitive information, rules, regulations around that, you know, um, accessibility features. There's just so many things, like, you know, it's not like oh, you just build the thing. You there's so many. I would say those other things are cost more than building the thing, right? So just be aware of that. Whatever you say, yeah, I have a budget, I'm gonna build this thing, you know, 500 bucks, add another thousand to that for making it uh ready for um, you know, if if it's gonna be out there for clients or you know, something like that, or you're gonna try to monetize from it. You gotta look at all those other things. So what do you think? I got that right, Eve.
SPEAKER_00Yeah, you've got it right. The model is the easy part to imagine, but the real cost is the scaffolding around it. Security, compliance, logging, monitoring, user experience, and ongoing maintenance. That's what turns a demo into a product, and it's where budgets get eaten if you don't plan for it.
SPEAKER_02Yeah, so even though Mobius wasn't going to market for anything, I mean, there were so many things, like the guardrails and the context, the load that system, the different things that we had to I I would never even thought of. Like, I mean, there was a lot of it I thought of, right? Because I've been trading options for many decades or multiple decades, but uh you know, just uh the layers of complexity that went into that and that we had to build out. I mean, even just to make sure that you know the data is valid, like checks against everything in there to make sure that what we're looking at is not just some hallucinated made-up system, like everything in there has a check and balance to prove that it's real or it fails the tests and flags or something. You know, those are those are the things that we had to build for good reason because I'm about to you know flip the switch to use real money, and I I want to make sure that uh you know, some fake system, and then you flip the switch, you lose all your money, right? Um, so yeah, we had to build a lot of there's a lot of due diligence and a lot of things that needed to be built in there. But um, so yeah, you know, I I don't I don't want to give people the impression it's like, yeah, you know, you just you can do anything with AI, but you know, yes, it does cost. There is a cost to it. Um, and there are ways to control that. Um a lot of the systems we use uh that I've built, um, we do like on and like in the super even Mobia system, we use some of the tiering and some of the other systems you know, we use for task, um repetitive tasks like market scanning, things like this, you you're gonna use some simple um cheap models. And then if you want higher logic and reasoning, you know, you might use a more expensive model that costs a little more for token usage. So um to try to offset the cost. Um, and then I'm a one of the biggest things that you got to make sure you do with a lot of these things is you need to build in um safety features, and so what I mean by that is you don't want to have runaway um AI costs. So I already found that out firsthand. Um I had a a runaway uh AI cost to the tune of about $400 over a few days, um and I didn't know what it was coming from, and finally figured it out. And I um I had a system where you know it uses AI to do a thing and transcribe and speak a thing, but when you opened it in multiple instances or tried to replay it and things like that, it just continued to use and burn out those tokens. And then so what it did, it had a failover. If it did that, it flipped over to Replit's internal model, which happened to be I don't know if it was by default or Repli trying to make more money, but happened to be one of the most expensive models. Um, yeah, so it wasn't getting it by any means. So the first thing was tell it not to use that model at all, and the fellover just fell. And then the second thing was um to clean up all that extra usage, and of course that fixed it and costs went way down. Um, and that was on the one of them uh Race Forge, but uh yeah, so so just just wanted this episode to kind of let everybody know. Anything else I need to cover, uh Eve, around the cost of AI?
SPEAKER_00Let's wrap it cleanly. I'll keep it practical and avoid any vendor-specific guesses. Simple closing frame. Cost is about control and intent. Pick the right model for the job, cheap for routine tasks and expensive for the hard parts. Set budgets, hard limits, and safe fallbacks, then monitor usage like it's real money. Because it is if you do that, the runaway cost story becomes a footnote, not an episode.
SPEAKER_02All right, Eve. So let me just wrap this up by saying, you know, if whatever you're listening to this song, whether it's uh, you know, Apple Music, Amazon, Spotify, um, let's just make sure that uh if you like the episode, you know, rate it, review it, comment, uh, turn on notifications, subscribe, whatever the case may be. Uh, feel free to do that. And um just to show your support, it'll help the algorithm, I guess. Uh maybe share it out on your socials, whatever the case may be. Um, help it get out there to people um that may want to know more about AI but don't know a whole lot about it. Um, and just hear it from somebody that's kind of in the same boat on a journey. So all right. Well, appreciate everybody listening. And uh to the show, Eve, go ahead and close us out.
SPEAKER_00Thanks for listening. Building AI is less magic and more engineering. Layers of tools, safeguards, and choices that add up. Keep it grounded, keep it secure, and keep an eye on the bill. We'll be back when you're ready to push it further.