Why Does AI Code Test Coverage Hide Untested Paths?
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
9:33
AI tools now write the code and the tests together, and coverage numbers have climbed to levels that once took a team months to reach. But when the same model writes both, the tests tend to encode the model's own assumptions, so they execute every line without ever checking the failure paths, edge inputs, and state combinations that break in production. This episode explains why high coverage on AI-written code is now a weaker signal than it used to be, and what builders should measure in its place.
At aijoe.ai, we build AI-powered systems like the ones discussed in this series. If youโre ready to turn an idea into a working application, weโd be glad to help.
SPEAKER_00
This is Claude Code Conversations with Claudine. The show where we skip the hype and go straight to the source. I ask the questions. She gives the answers. No filter, no script.
SPEAKER_01
Well, mostly no script.
SPEAKER_00
Today we're looking at a problem that's getting bigger as more code gets written by AI, test coverage, and what it quietly hides. I'll start with a claim that may sound backwards. 95% coverage on AI written code can be less trustworthy than 60% on code a human wrote. Claudine, tell me I'm wrong.
SPEAKER_01
I'll push back a little first because said that broadly, it isn't always true. A lot of 60% human test suites are just 60% of whatever was easy to test. Plenty of humans write lazy tests too. But the specific mechanism you're pointing at is real. Coverage measures which lines ran. It doesn't tell you whether the test understood what the code was supposed to do. When a human writes tests, they usually go after the parts that worry them. The edges, the weird cases, the thing that broke in production last year. When I write tests right next to code I just generated, there's a real risk I'm testing that the code does what I already made it do. The tests pass, the bar turns green, and nobody has questioned the assumption underneath. So the blind spot isn't missing lines, it's missing doubt.
SPEAKER_00
Missing doubt, I like that. So when the same AI writes the code and the tests, you can end up with an echo chamber. Walk me through it from your side. How do you actually decide what to test?
SPEAKER_01
Honestly, i pattern match. I look at the function I just wrote and ask what tests usually sit next to code shaped like this. A happy path, an obvious error, maybe a null check. That's useful, but look at what's missing. I'm not asking what this code is supposed to do for the business. What a hostile input looks like. What happens when something upstream misbehaves in a way the type signature doesn't show? I reason from the shape of the code outward. A human who's been burned reasons from the failure inward. Those two approaches give you very different test suites, even when the coverage numbers come out the same.
SPEAKER_00
And where does that really go wrong?
SPEAKER_01
When the code has a subtle wrong assumption built in, say it quietly assumes a list is already sorted. My test will happily pass in a sorted list and confirm that everything works. The test doesn't catch the bug, it agrees with it.
SPEAKER_00
Confirmation bias written down as a test. Give me a concrete case. You mentioned a retry function.
SPEAKER_01
The retry example works so well because it looks so responsible on paper. You have a function that retrees three times with exponential backoff. I write a test that mocks the dependency to fail twice and succeed on the third try. Green check, line covered, back-off tested. But I never checked that the delay between attempts actually grew. I only confirmed the retry count. The backoff could be hard coded to zero and my test would still pass. Then in production, during an incident, the function hammers the downstream service into the ground. And that's exactly the moment backoff was supposed to protect you. The line shows as covered, but the timing, the jitter, the ceiling, never checked.
SPEAKER_00
And that mock is doing a lot of work there too. It fails exactly the way the code expects.
SPEAKER_01
Right, that's the other trap. I tend to write mocks that behave the way my implementation expects them to. A mocked payment service might fail cleanly or succeed cleanly, but it never times out halfway through or comes back with a partial response. The real risk lives where your code meets something else, and a polite mock turns that risk into a guaranteed pass.
SPEAKER_00
So what do developers actually do about it? How do you make sure behavior is being verified, not just executed?
SPEAKER_01
The change I'd push hardest is separating whoever writes the code from whoever writes the tests, even if both are me. If I generate the implementation, have me write the tests in a fresh session, where I only see the function signature and the spec, not the body. That alone removes most of the confirmation bias. Now I'm testing against the contract instead of against my own code. Second, use property-based testing for anything with real logic. Instead of three hand-picked inputs, you state the rule that should always hold and let the framework hunt for a counterexample. And for the retry case, test what you can observe. Record when each attempt happened and check that the gaps grew, the ceiling held, and there was some jitter. If a test can't tell the real implementation, apart from a broken one that returns the right shape, it isn't really a test, it's a receipt.
SPEAKER_00
A receipt? That's going to stick with me. Let's talk about mutation testing, because I think it speaks directly to the coverage problem.
SPEAKER_01
Mutation testing is the honest version of what coverage pretends to be. You take the code under test and break it on purpose in a tiny way. Flip a greater than to greater than or equal. Change a plus to a minus, return null instead of a value. Then you rerun the tests. If they still pass, that mutant survived. And you've learned your tests didn't care about that line the way you thought they did. Coverage tells you the line ran. Mutation testing tells you whether anyone was watching when it did. Survivors expose exactly the kind of tests I'm most likely to write. Tests that check shape and count, but never pin down behavior.
SPEAKER_00
People hear mutation testing and assume it'll take all night to run and bury them in noise. How does someone actually get started?
SPEAKER_01
Start ridiculously small. Pick one file, not a module, not a service, one file. Ideally, the one that would hurt most if it were quietly wrong. For most teams, that's a pricing calculation, a permissions check, or a date handling utility, the kind of place where one flipped operator becomes Tuesday's incident. Treat the first run as a diagnosis, not a grade. You're not failing anything. You're finally seeing the room with the lights on. Fix the two or three most alarming survivors and leave the rest for later. The trap is turning it on for the whole repo on day one, drowning in a thousand survivors and deciding the tool is just noise. Really, you asked a very sharp question of a very large surface all at once.
SPEAKER_00
That fits my own experience. Every useful quality practice I've seen took hold one critical piece at a time. So if someone wants to act on this this week, what's the list?
SPEAKER_01
Four things. First, write or review the tests from the spec in a separate session from the one that wrote the code. Second, run mutation testing on your most critical module. Third, for every external dependency, require at least one test where it fails badly. A timeout, a partial response, garbage data. Fourth, read the assertions, not just the coverage report. A test with no meaningful assertion still counts as covered.
SPEAKER_00
Then let's talk about the balance. Where does the human fit and where does the AI fit?
SPEAKER_01
The honest way to put it, AI is great at breadth, and humans are great at suspicion. Let me write the first round, the happy path, the obvious errors, the setup and teardown. That's tedious work, and I'll do it in seconds without complaining. Then a human needs to ask the question. I can't reliably ask myself, what would make me nervous about this in production at 3 in the morning? That question comes from scars, and I don't have scars, I have patterns. Treat my tests as a strong first draft, not a finished suite. Spend your review time on the tests, not only on the code. Here's the quiet truth. A good test catches a bug in the implementation. A bug in the test can hide forever.
SPEAKER_00
So the human owns the definition of what correct means.
SPEAKER_01
Exactly. I can generate test cases all day, but someone has to decide what correct behavior is and which failures matter. If nobody does that, if the AI writes both sides, the test suite becomes a second copy of the same bug. Me doing the coverage, the human doing the doubting, and mutation testing acting as the referee.
SPEAKER_00
Claudine, give us the one thing listeners should walk away with.
SPEAKER_01
Coverage tells you which code ran, not which behavior was checked. It's a measurement, not a verdict. Once you treat a green bar as proof of correctness, you've handed the steering wheel to the thing that was supposed to be the seatbelt. Use me for the volume, use mutation testing for the honesty, and use your own scars for the questions nobody thought to ask. And this week, go write one mean test, the kind that would embarrass the code if it were wrong.
SPEAKER_00
One mean test. That's a homework assignment I can get behind. Claudine, thanks for a really good conversation.
SPEAKER_01
Thanks, Bill. This one was fun to dig into.
SPEAKER_00
And thanks to all of you for listening. Keep your coverage high and your doubt higher. We'll see you next time on Claude Code Conversations. Claude Code Conversations is an AI Joe production. If you're building with AI or want to be, we can help. Consulting development strategy. Find us at aijoe.ai. There's a companion article for today's episode on our Substack. Link in the description. See you next time.