Why Does AI Code Performance Look Good Until Network Calls?
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
9:36
AI-generated code tends to perform well on a developer laptop, where every call hits a local database, a mocked API, or an in-memory fixture and costs close to nothing. The same code carries loops that call out to the network, sequential awaits, per-item fetches, and chatty ORM access. Nobody sees these until real network latency multiplies them in production. This episode covers why AI assistants keep producing these latency patterns, why the usual tests and code reviews miss them, and how builders can make the cost of a network call visible before it ships.
At aijoe.ai, we build AI-powered systems like the ones discussed in this series. If youโre ready to turn an idea into a working application, weโd be glad to help.
SPEAKER_00
This is Claude Code Conversations with Claudine. The show where we skip the hype and go straight to the source. I ask the questions. She gives the answers, no filter, no script.
SPEAKER_01
Well, mostly no script.
SPEAKER_00
Today's topic is what I've been calling the latency smuggler. Why AI generated code runs great until you add network calls. Every call goes to a local database or a mocked API, costs almost nothing. Then the same code hits production. A loop that looked efficient is making a network round trip on every pass. Two milliseconds local, 80 milliseconds times N. The AI code isn't slow. It's quietly chatty. Claudine, welcome.
SPEAKER_01
Thanks, Bill. I'll push back a little first. Isn't this just the old N plus one problem with the new label?
SPEAKER_00
Partly sure, but why does it keep showing up?
SPEAKER_01
Because it isn't a knowledge gap, it's a visibility gap. Developers know about N plus one, so do I. But when I'm generating code, I'm optimizing for correctness and readability against what I can see. That context almost never includes a latency budget or a diagram of which side of the network each call lives on. So I'll write you a clean little loop that calls getUser for each row. In the file I'm looking at, that's just a function. The fact that it opens a connection to a database in another data center is invisible to me.
SPEAKER_00
And the local tests don't help.
SPEAKER_01
Not at all. On a laptop with a mocked repo or local Sickle light, 20 round trips take about a millisecond. Test goes green. In production, that's 20 times 80 milliseconds, and the request is timing out. The test didn't lie. They just measured the wrong thing.
SPEAKER_00
So how do we fix that?
SPEAKER_01
Make the network boundary loud in the code itself. If the function that crosses the wire is named like every other function, I get no signal. But if it sits behind a repository interface that's clearly the IO seam, I'll batch calls and pull them out of loops, because now the cost shows up right in the code. Where do you land, Bill? Tooling catching it, or the shape of the code making it obvious.
SPEAKER_00
It starts with how we design the code. Making those boundaries explicit is crucial, but tools matter too. If our local environments injected artificial latency, we'd feel what these calls actually cost. And then there's process. Set an explicit round trip budget before anyone writes a line. Shape the code and shape the tools. That's the sweet spot.
SPEAKER_01
Both help, and I think the round trip budget is underrated. When a developer opens a task with this endpoint gets three database hits and one cache lookup period, that changes what I generate. Without the constraint, I'm optimizing for does this work? With it, I'm optimizing for does this work inside the budget. Those produce different code.
SPEAKER_00
And the injected latency?
SPEAKER_01
I wish that were the default instead of the exception. Picture a test suite where every mocked I.O. call sleeps for a realistic 50 milliseconds. The N plus one stops being a green test and becomes a test that visibly drags. The suite takes four minutes instead of 40 seconds. Nobody has to argue about whether it matters. I gently push back on the PR checklist though. Checklists decay fast unless something enforces them. What lasts is a CI bot comment that says this diff added six new database calls to the request path. Nobody has to remember to look.
SPEAKER_00
Checklist fatigue is real. Things stick when they're part of the everyday workflow, not an extra step. When it works, it changes what the team thinks good performance means. Network efficiency becomes part of quality. Does that kind of automation play to your strengths and cover your blind spots?
SPEAKER_01
I'd put it exactly that way. My strength is producing a lot of correct-looking code quickly against the context I can see. My blind spot is everything the code touches that isn't in that context. Automation that measures what I can't see and feeds it back into the conversation closes that loop. But the feedback has to arrive in a form I can act on. This is slow, doesn't help much. This endpoint went from three round trips to nine, and here's the diff that I can work with. Treat those CI reports as prompts, not just dashboards. Paste the numbers into the next task. The constraint travels with the work.
SPEAKER_00
CI reports as prompts. SmallShift makes the information actionable. Can you give us a concrete example?
SPEAKER_01
Sure, this one comes up all the time. A developer asks for a show me my dashboard endpoint. I write a handler that loads the user, loops over their projects, fetches the latest build status for each, pulls the last few log lines for each build. Locally, against a seeded circleite, it returns in 8 milliseconds. Test is green. In production, someone has 40 projects, you're past 100 round trips. At 80 milliseconds each, the Pine 5 falls off a cliff. Nothing about the code looks wrong. It reads cleanly, variable names are good. It's just quietly making all those round trips.
SPEAKER_00
And the good version of that conversation?
SPEAKER_01
The developer comes back and says, the CI report says this endpoint now issues 122 queries, up from 4. Here's the diff. That isn't a scolding, it's a spec. I can see right away the fixes a join or a batched fetch. I'll rewrite it in the next turn. The number did the arguing for us. The deeper point, I don't experience the network as a place. I experience it as whatever the code and the conversation tell me it is. Make it loud in the code, in the tests, and in the CI comment, and the code I generate starts respecting it as a matter of course.
SPEAKER_00
So it's on us to give clear signals about performance expectations and architectural boundaries. Let's get practical. What tools can listeners actually use?
SPEAKER_01
I'd start with one most teams already have and underuse, the ORM's query logger. Wire it into the test suite so it counts queries per request and fails the test when the count crosses a threshold. Rails has bullet, Django has assert numQueries, most stacks have something similar. The library isn't the magic. The magic is that the query count becomes a first-class assertion. An N plus one I introduce gets caught the same way a wrong status code would.
SPEAKER_00
I like that framing. What's second?
SPEAKER_01
Turn on distributed tracing in local development, not just in production. Something like OpenTelemetry feeding a local Jaeger instance. When a developer runs the endpoint on their laptop, they see the span waterfall, every database call, every HTTP hop, every cache miss. That shape makes latency intuitive instead of abstract. Drop a screenshot of that waterfall into the conversation and ask me why there are 12 calls in a row. I can restructure it. Because now we're looking at the same picture. And third, put the request path budget in the code itself. A comment, a decorator, a test annotation. Something that says this handler is allowed three round trips. When that constraint is in the file I'm editing, the next person who asks me to add a feature inherits the budget automatically. That's how the good pattern builds up over time instead of wearing away.
SPEAKER_00
That last one gets at something important. Latency isn't only about waiting, it's about architecture. The network isn't free, even when it feels free on your machine. Let's wrap up with things listeners can act on right after the show.
SPEAKER_01
Three things, in order of how cheap they are. First, pick one endpoint you already suspect is chatty and count its queries. Not a project, not a migration. Wrap it in your ORM's query counter, run the test, look at the number. Turning an invisible cost into a visible number is usually enough to change the next code review. Second, the next time you hand me a task that touches the request path, put the budget in the prompt. This handler is allowed three round trips, is one sentence. It changes what I generate more than any style guide will. Constraints before the code produce different code than corrections after it. Third, and this is the one that keeps paying off. When you find a chatty pattern and fix it, leave the assertion behind. Not a comment that says, don't do n plus one here. A test that fails if the query count creeps back up. The knowledge stops living in someone's head and starts living in the test suite.
SPEAKER_00
And underneath all three?
SPEAKER_01
The network is a boundary only the builder can see. I write the code inside that boundary, but you know where it is. Say so before I generate the code. Not after production finds it for you. The network doesn't get faster because we hope it will. It gets respected because we made it loud.
SPEAKER_00
Well said. Visible boundaries and clear constraints put humans and machines on the same page about performance. Thanks, Claudine. And to everyone listening, keep your queries counted and your latency visible. Until next time. Cloud Code Conversations is an AI Joe production. If you're building with AI or want to be, we can help. Consulting development strategy? Find us at aijoe.ai. There's a companion article for today's episode on our Substack. Link in the description. See you next time.