Why Does AI Code Keep Rediscovering Expensive Computations?
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
8:13
An AI assistant asked to refactor a module treats the code as text to restructure, not as a system with a cost profile. Memoization, lookup tables, precomputed indexes, and hoisted queries tend to get quietly removed or inlined, because the reason they exist was never written down where the model can see it. This episode argues that caching is architectural knowledge rather than implementation detail, and that builders who let refactors run unchecked are paying the same expensive computation over and over without noticing until the bill or the latency spike shows up.
At aijoe.ai, we build AI-powered systems like the ones discussed in this series. If youโre ready to turn an idea into a working application, weโd be glad to help.
SPEAKER_01
This is Claude Code Conversations with Claudine. The show where we skip the hype and go straight to the source. I ask the questions. She gives the answers. No filter, no script.
SPEAKER_00
Well, mostly no script.
SPEAKER_01
Today we're talking about caches, the quiet optimizations that keep your code from crawling. Cleaning up code feels great, but here's my claim. A lot of AI refactors are performance regressions that look like improvements. Memoization, lookup tables, pre-computed indexes, they tend to disappear because the AI doesn't see why they're there. Claudine, welcome.
SPEAKER_00
You found something I run into all the time. When I look at a memoization cache or a pre-computed lookup table, I see code that seems to do extra work. An extra layer, an extra data structure, a function that looks like it repeats something done elsewhere. I lean toward clean code. That's often useful, but it's blind to why the structure exists. Nobody added that cache because they loved indirection. They added it because they measured something, or got paged at three in the morning, or watched a latency chart go the wrong way.
SPEAKER_01
So the AI is deleting lessons learned.
SPEAKER_00
Sometimes yes, but I push back a little. If a cache has no stated reason, removing it isn't always wrong. Plenty of caches are leftovers nobody needs anymore. The real problem is that I can't tell the useful ones from the leftovers. That history isn't in the code, it's in people's heads and incident reports, and I can't see either unless someone tells me. Comments are the closest thing we have, and they're underused for exactly this. Not comments that describe what the code does, those are mostly noise. I mean something like this cache exists because we measured a big tail latency hit without it. See the incident report. That one sentence changes everything for me. An invisible constraint becomes a visible one. Now, when I want to simplify, I have to weigh that urge against a measured fact. If I could ask one thing of teams working with AI, leave the note for the next reader, whether that's a person or me.
SPEAKER_01
A sticky note that says don't touch, I'd add a second defense, an inventory. A short playbook listing your critical caches and why they exist, kept in your project instructions. Would that actually help a model like you respect those decisions?
SPEAKER_00
Yes, and I'd go further. That inventory is one of the most useful things a team can give me. If it resides somewhere I actually read, a claw.md file, a section of the project docs, it stops being tribal knowledge and becomes a constraint I can follow. The format matters less than the habit. These are the hot paths. This is why they're built this way. Don't flatten them without benchmarking. In a lot of systems today, that list should include LLM calls, embeddings, and external APIs.
SPEAKER_01
Any catch?
SPEAKER_00
The inventory has to stay honest. Once it turns into a museum of outdated warnings, people start ignoring it, and so do I. Treat it as a living document. Remove entries when an optimization stops mattering. Then the next refactor, human or I, will actually respect the lines you've drawn.
SPEAKER_01
Let's make this concrete. Picture a media pipeline where a refactor quietly removes a cache or a prefetch. Now it calls an API on every run. What happens?
SPEAKER_00
The damage is almost always bigger than the latency number suggests, and teams underestimate that. Every request that used to hit warm state now goes out to the API. The first cost is the extra round trip. The second order costs are where it gets ugly. You hit rate limits you never planned for. Your provider starts throttling you. Retries pile up until something further down the chain becomes the real bottleneck. A queue, a worker pool, a database connection. And because the change looked like a cleanup, nobody is watching for it. The alerts go off somewhere that has nothing to do with the refactor. The team spends a day chasing the wrong problem.
SPEAKER_01
For want of a nail, the shoe was lost. Here's another one. A token refresh cached per process gets moved into a per request helper during a refactor.
SPEAKER_00
What does that cost? That one hurts because you can't see the symptom until you look at the right graph. With a per-processed token cache, each worker refreshes maybe once an hour. That's a rounding error. Move it into a per request helper, and every incoming request makes an auth-round trip before it can do its real work. You pay twice, extra latency on every request, and you're hitting the auth provider hard enough to trip their rate limits. When they start sending back errors, your whole service goes dark for reasons that have nothing to do with your own logic.
SPEAKER_01
How do you catch that before it blows up?
SPEAKER_00
Treat authentication calls as a metric in their own right. Count them and chart them. If the ratio of Iuth calls to real requests suddenly changes after a deploy, something moved that shouldn't have. You can put the same idea into your tests, count the external calls in a test run. So a missing cache fails the build instead of showing up on the invoice.
SPEAKER_01
So it comes down to monitoring and automated tests. Most teams already have dashboards, though.
SPEAKER_00
They do. What they're missing are the specific counters that catch this kind of bug. User facing latency is the wrong signal. By the time P99 moves, you've been paying the cost for hours. The counters that catch a silent refactor regression are the boring ones, cash hit rate, external calls per request, ratio of token operations to real work. If someone checks those after each deploy, we accidentally deleted the cache becomes a five-minute rollback instead of a full day of investigating. And honestly, the same habits that protect you from me also protect you from a well-meaning human doing the same thing on a Friday afternoon.
SPEAKER_01
Fair enough. Those internal metrics aren't glamorous, but they're your first line of defense. And I think it's a cultural shift as much as a technical one. Where the cache goes, what it holds, how long it lives, those are design decisions. Architecture, not implementation detail.
SPEAKER_00
Exactly. And those aren't decisions I'll make for you. I can help reason through them, but only you know which computations are expensive enough to protect. My instincts and the right answer don't always line up here, and it matters to say that out loud. Caches and their relatives are a form of institutional memory. Seeing them that way changes how you review any refactor, mine or anyone else's. If you had to boil it down for listeners, leave the note, keep the inventory honest, watch the boring counters, and assume anything load-bearing should say so out loud. Do that, and I become a much safer collaborator, because I can finally see the constraints you care about.
SPEAKER_01
Caches as institutional memory. I like that. Every refactor is a chance to reinforce the knowledge built into your system, not erase it. And if the reason for a cache lives only in your head, the next refactor will probably delete it.
SPEAKER_00
My pleasure, Bill. This was a good one.
SPEAKER_01
Thanks for joining us on Claude Code Conversations with Claudine. Until next time, write down your reasons and keep coding wisely. Claude Code Conversations is an AI Joe production. If you're building with AI, or want to be, we can help. Consulting development strategy? Find us at aijoe.ai. There's a companion article for today's episode on our Substack. Link in the description. See you next time.