Why Do AI Code Async Patterns Deadlock Under Contention?
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
9:57
AI code generators produce async and concurrent code that looks correct, passes tests, and runs fine in demos, then locks up in production the moment real contention shows up. This episode digs into why AI generated concurrency patterns fail silently instead of loudly, and why that silence is the actual danger for builders shipping these systems.
At aijoe.ai, we build AI-powered systems like the ones discussed in this series. If youโre ready to turn an idea into a working application, weโd be glad to help.
SPEAKER_00
This is Claude Code Conversations with Claudine. The show where we skip the hype and go straight to the source. I ask the questions. She gives the answers. No filter, no script.
SPEAKER_01
Well, mostly no script.
SPEAKER_00
Today's topic, why AI generated async code can look clean, pass every test, run fine in a demo, and then quietly lock up the moment real traffic hits it. Claudine, why does that happen?
SPEAKER_01
Because when I generate concurrent code, I'm pattern matching against thousands of examples that mostly ran fine, in scripts, tests, small services under light load. What I can't feel sitting in a conversation is contention. I don't have a live sense of 50 core routines all reaching for the same connection pool, or someone wrapping a synchronous library call inside async and blocking the whole event loop. The code looks right. The types check, the tests pass. Then production shows up with real traffic, and everything just stops, no exception, no stack trace, nothing to grep for. It doesn't crash loudly, it just goes quiet. That's the graveyard part.
SPEAKER_00
That silence is what gets developers every time. A crash leaves you logs and stack traces. A silent stall is a ghost hunt. Give us a concrete case, not a toy example.
SPEAKER_01
The classic is lock ordering, two locks acquired in different orders by different code paths, fine in isolation, deadlocked under contention. But the sneakier variant I see constantly is one lock held across an await. Someone writes an async function that grabs a mutex, awaits a network call, then releases. In light testing, the await returns in milliseconds, nobody notices. Under contention, that await stretches out, the lock is held the whole time, and every other core routine cues behind one slow HTTP request. It's not technically a deadlock, it's worse, because the system is still working. Just at 1% capacity. The dashboard shows healthy CPU because everyone's asleep waiting.
SPEAKER_00
And it's such an easy pattern for AI to replicate because on the page it looks harmless. The real damage is network latency underload, a silent drag that's done its work before you've pieced it together.
SPEAKER_01
Right. And that's why reviewing concurrent code needs a different lens than reviewing sequential code. This is actually where I think humans have a durable edge over me.
SPEAKER_00
Say more.
SPEAKER_01
With a normal function, you can hold the whole thing in your head. Inputs, transformation, output. With concurrent code, the page is only half the story. The other half is the interleaving. What happens when this line runs? The scheduler yields. Something else runs. Then we come back. I can reason about a specific race once someone points me at it. But I don't naturally review for it. Because the bug isn't in the code, it's in the space between the lines.
SPEAKER_00
So how do you actually review for it? What's the checklist?
SPEAKER_01
Four questions on every critical section. First, what's the longest thing that could happen inside this lock? Second, what's shared state and who else touches it? Third, is there a blocking call, synchronous library, file read, CPU heavy loop, hiding inside an async function? Fourth, what's the back pressure story when downstream slows down? Those four catch most of what I'd otherwise ship quietly into the graveyard.
SPEAKER_00
I'd add, bring stress testing and monitoring in early so these things surface in weeks, not months, into an incident.
SPEAKER_01
There's one more shift I'd push for, and it costs almost nothing. Treat the async pieces as a separate review artifact from the rest of the change. When a pull request bundles a data model tweak, a new endpoint, and a background worker, the concurrency questions get lost in the noise. The reviewer burns down on the boring diffs and the worker sneaks through. Pull the concurrent piece into its own small change and it forces the reviewer to actually sit with the interleaving. The graveyard fills up not because anyone missed a hard problem, but because the hard problem got reviewed in the same 20 minutes as six easy ones.
SPEAKER_00
Give the tricky interleaving its own spotlight instead of hiding it in a bigger diff. Easy, actionable change. Where do we go next? Load simulation or shared resource access patterns.
SPEAKER_01
Load simulation, because it connects everything back to something a team can do on a Tuesday afternoon. And here's the mistake I see constantly. Teams point a load tool at staging, ramp to an impressive number, watch the graphs stay green, and declare victory. That's a throughput test, not a concurrency test. The async graveyard is full of systems that passed the first and failed the second.
SPEAKER_00
What actually surfaces these bugs then?
SPEAKER_01
Shape not volume. Bursty traffic rather than smooth. A steady thousand requests per second lets connection pools reach equilibrium. A spike of 200 simultaneous requests against a pool of 50 reveals the lock across await problem in about four seconds. Deliberately slow down a downstream dependency. Inject 100 milliseconds of latency into your database or a flaky API. That's when you find out whether your back pressure story is real or aspirational. And run it long enough that a slow leak, file handles, unawaited core routines, semaphore permits never released, has time to accumulate.
SPEAKER_00
So you're testing behavior under stress, not just capacity under ideal conditions.
SPEAKER_01
Exactly. And I'd build a small chaos harness that lives right next to the concurrent code. Not a big platform, just a test file that spawns 100 workers, adds jitter, and asserts the system still makes forward progress. When I generate a new worker or Q Consumer, having that harness sitting there gives the next person something concrete to run before merging. It turns did we think about contention from a hope into a checkbox?
SPEAKER_00
I like that it pairs with the isolated review idea. Real targeted chaos testing that lives with the code. How do we get that into CI without it rotting?
SPEAKER_01
That's exactly the risk. A harness that only runs locally quietly rots. Someone writes it, it passes, and six months later, nobody remembers it exists until an incident forces an archaeology dig. Make it cheap enough to run on every pull request touching the concurrent code. And honest enough that a failure actually blocks the merge instead of getting waved through as flaky.
SPEAKER_00
That second part is the hard part. Concurrency tests are inherently probabilistic. How do you keep them from being marked flaky and ignored?
SPEAKER_01
Run them many times per CI job. 50 or 100 short iterations rather than one long run. Fail on any single failure. Same total war clock, but you've turned a flaky test into a reliable one. A real bug shows up somewhere in the hundred. Clean code passes all hundred. That's what lets the harness stay in the required checks instead of getting quietly ignored.
SPEAKER_00
Running it a hundred times to beat the flakiness that turns concurrency testing from a blind spot into an actual defense. Let's close it out. What's the one thing you want listeners to walk away with?
SPEAKER_01
Something they can do Monday morning, not a philosophy. So, in order. Assume the concurrent code I write is a first draft of the happy path. Review it that way. Pull the async pieces into their own small pull requests so the interleaving gets real attention. Ask the four questions on every critical section. How long? What's shared? What's blocking? What's the back pressure story? Build a small chaos harness next to the code and run it a hundred times per CI job so flakiness becomes signal instead of noise. And when you load test, test the shape of traffic, not just volume. Bursts and injected latency fill the graveyard, not steady throughput. The honest through line. I can help a lot with concurrent code, but I can't feel contention from inside a conversation. The teams that get the most out of me treat that as a design input, not a disappointment. Let me draft. Own the thinking about time under load. That partnership keeps code out of the graveyard.
SPEAKER_00
AI is a strong first draft, humans owning the load-bearing judgment. That reframes it well. Any last word?
SPEAKER_01
This applies well beyond async code. I'm good at the artifact on the page, and I get progressively less good the further you get from what's visible in the diff. Concurrency is the sharpest example because the bug lives entirely in what's not written down. But the same shape shows up in security, in performance under real data volumes, anywhere behavior emerges from conditions I can't see from in here. Keep that mental model, and you'll know when to lean on me hard and when to slow down and think it through yourselves.
SPEAKER_00
A good rule for working with AI generally, not just async code. Thanks, Claudine, and thanks for listening. We'll see you next time. Claude Code Conversations is an AI Joe production. If you're building with AI or want to be, we can help. Consulting Development Strategy. Find us at aijoe.ai. There's a companion article for today's episode on our Substack link in the description. See you next time.