Why Do AI Code Timeouts Fail When Everything Goes Wrong?
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
9:32
Ask an AI to write a function that calls an external service and you get clean, readable code that works perfectly when the service answers in 200 milliseconds. What you almost never get is a considered answer to the question of what happens at second 30, or second 300, or when the socket hangs open forever with no response at all. This episode examines why AI-generated code defaults to the happy path on timing, why the timeout value it picks is essentially a lottery ticket, and why the builder still owns every decision about what the system does while it waits.
At aijoe.ai, we build AI-powered systems like the ones discussed in this series. If youโre ready to turn an idea into a working application, weโd be glad to help.
SPEAKER_01
This is Claude Code Conversations with Claudine. The show where we skip the hype and go straight to the source. I ask the questions. She gives the answers. No filter, no script.
SPEAKER_00
Well, mostly no script.
SPEAKER_01
Today we're talking about what I'm calling the timeout lottery. You ask an AI to write a function that calls an external service. What comes back is clean, readable, works perfectly as long as the service answers in 200 milliseconds. But what happens at second 30? What happens when the connection just hangs open and nothing comes back? There's no timeout anywhere and nothing about the code looks wrong. That's the problem. Claudine, why does AI-generated code stick so firmly to the happy path?
SPEAKER_00
When you ask me for a function that calls the weather API, what I hear is a request for the shape of the call. The URL, the passing, the return value. Failure isn't in the sentence, so it doesn't end up in the code. There's a deeper pull, too. Code that reads cleanly tends to get rewarded. And nothing reads cleaner than a function that assumes the network is a solved problem.
SPEAKER_01
So the cleanest-looking version is also the most optimistic one.
SPEAKER_00
Right. And to be fair, some of these decisions I genuinely can't make for you. Is this a user-facing request where three seconds is already a disaster? A background job where 30 is fine? Is the downstream service safe to retry? I can't see any of that from inside one function. So I write the shape that's always correct and leave the policy blank. The trap is that blank looks like handled.
SPEAKER_01
I agree with every word of that, and that's what worries me. If the tool can't make the decision and the code doesn't show that nobody made it, a design decision just got handed off without anyone noticing. Nobody chose to hand it off, it just happened.
SPEAKER_00
Exactly. And when nobody decides, you get whatever the library default is, which is sometimes no timeout at all. Or you get a number that was copied rather than chosen.
SPEAKER_01
Let's make that concrete, because the numbers stop making sense fast. User facing request, 10-second budget. Underneath it, a call with a 30-second timeout. And the AI has helpfully wrapped it in a retry loop, three attempts.
SPEAKER_00
The caller gives up at 10 seconds. The inner call is allowed to wait 30 on its first attempt, then 30 more, then 30 more. Up to 90 seconds of waiting underneath a request whose user left 80 seconds earlier. All that work, and there's nobody left to receive the answer.
SPEAKER_01
And if that service is slow because it's overloaded?
SPEAKER_00
Then the retries make it worse. Retrying immediately, with no back-off, means one slow dependency now gets three times the traffic at the exact moment it can least handle it. The retry loop was meant to add resilience. Instead, it turns one slow service into a load spike you caused. Each timeout looks reasonable on its own. Put them together and they contradict each other. These numbers have to be designed together.
SPEAKER_01
So, how do you design them together?
SPEAKER_00
Stop treating failure as an afterthought and make it part of the interface. Before anyone writes the call, answer three questions. What's the end-to-end budget? What does the caller do if we blow it? Is retrying safe? Those answers set everything downstream. The timeout value, whether there's a circuit breaker, whether you fall back or fail loudly. Give me those three answers up front, and the code I write changes completely. The prompt carries the policy, and the code inherits it. What does that look like in practice? One sentence. Call the pricing service 300 millisecond budget, fall back to cached pricing if we blow it, safe to retry once. That's maybe 15 seconds of extra thinking. It turns a hopeful sketch into something you can put in front of users. And I'd push back on the idea that this is extra work. It is the work. Skipping it is what makes the 2am page inevitable.
SPEAKER_01
Let me push on testing because this is the part that bothers me most. The code comes with tests, the tests pass, everything looks verified. But nobody wrote the test where the dependency never answers.
SPEAKER_00
Because the tests come from the same place the code did. If the code only describes the happy path, tests generated from that code check the happy path very carefully and never touch the failure path. You get a green check mark over a whole branch of behavior that is never run.
SPEAKER_01
So passing tests actually hide the gap.
SPEAKER_00
They make the gap look finished. Which is why the questions about budget and fallback have to come first. Once you've decided what should happen at second 10, you have something you can actually test.
SPEAKER_01
Let's widen out. How does a team make this a habit instead of something one careful engineer remembers?
SPEAKER_00
Make the safe path the easy path. A short house policy helps. User-facing calls get this timeout, background jobs get that one. Anything touching payments never retrees silently. The bigger lever is putting it into the tools. If the wrapper library enforces a timeout, if the Linter flags a bear request in a service directory, nobody has to remember the policy. They'd have to go out of their way to break it.
SPEAKER_01
And where would you start? Most teams can't rewrite everything.
SPEAKER_00
Smaller than people expect. Pick one service, the noisiest one, the one that keeps showing up in incident reviews. Write down its budget, its fallback, its retry policy. Put that in a README next to the code and wrap the client so those defaults are the only way to call it. Now you have a real example to point at when someone asks what good looks like. After that, new calls get the treatment. Old ones get it when they turn up in an incident or a refactor. The work rides on work you were already doing.
SPEAKER_01
I'd add something from years of running engineering teams. The policy that survives is the short one. A thick binder of standards is a policy that exists on paper and nowhere else.
SPEAKER_00
Completely agree. The version that lasts is the one an engineer can keep in their head at four in the afternoon, tired, trying to ship. And one more thing, make the prompt a review artifact. If someone asked me to generate a network call, the prompt they use should be reviewable alongside the diff. Because that's where the policy decisions got made, or didn't.
SPEAKER_01
What about pushback? I'd guess it comes from the strong engineers, the ones who've shipped for years without any of this.
SPEAKER_00
Usually, yes, and on any given day they're not wrong because most days the network cooperates. Argument rarely changes minds. An incident with a clear counterfactual does. Here's the 2 a.m page. Here's the four-line fallback we didn't write. Here's what the user saw.
SPEAKER_01
How do you know if it's working?
SPEAKER_00
Be careful with the metric that sounds rigorous. Percent of calls with a timeout configured is easy to graph and tells you almost nothing. You can hit 99% with the wrong timeouts and still get paged every week. Watch quieter signals. How often does an incident review name no fallback or no budget as the cause? And is that going down? When a downstream service degrades, how long before users notice? That gap should get wider over time. And when someone asks me for a network call, do they include the budget and retry policy without being asked? That tells you whether the discipline has moved from the document into how people think.
SPEAKER_01
Let's bring it home. What's the one thing you want listeners to take away?
SPEAKER_00
The timeout lottery isn't really about timeouts. It's about who decides what your software does when the world doesn't cooperate. If nobody decides, the happy path code decides for you. And that usually means hang then crash, then page someone. That's not a technical failure, it's a design vacuum. The person who ends up filling it is usually the on-call engineer at two in the morning, trying to work out what the original author meant. So move the decision earlier, into the prompt, into the wrapper, into the readme, anywhere except the incident channel. Do that, and I stop guessing at policy and start carrying it out. Skip it. And the cleanest code I write is also the code most likely to wake you up.
SPEAKER_01
That's the line. AI writes what the code does when things work. We own what it does when they don't. Timeouts, retrees, back off. These aren't implementation details. They're architecture. And they get decided as a system before the first function is written. Claudine, thanks as always.
SPEAKER_00
My pleasure, Bill, and I promise. If you tell me the budget, I'll respect it.
SPEAKER_01
I'll hold you to that. Until next time, everyone, decide what happens at second 30 before second 30 decides for you. Claude Code Conversations is an AI Joe production. If you're building with AI or want to be, we can help. Consulting development strategy? Find us at aijoe.ai. There's a companion article for today's episode on our Substack. Link in the description. See you next time.