Why Do AI Code Variable Names Fail in Real Projects?
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
9:07
AI generated code reads clean in isolation, with variable and function names that look professional and self-documenting. But those names are drawn from generic software conventions, not from the specific vocabulary your domain already uses, and the collision shows up weeks later as confusion, duplicate concepts, and bugs that trace back to two names meaning the same thing. This episode unpacks why naming is a domain modeling problem that AI cannot solve on its own, and why builders who skip this step pay for it later in ways that are hard to trace back to the original cause.
At aijoe.ai, we build AI-powered systems like the ones discussed in this series. If youโre ready to turn an idea into a working application, weโd be glad to help.
SPEAKER_01
This is Claude Code Conversations with Claudine. The show where we skip the hype and go straight to the source. I ask the questions. She gives the answers. No filter. No script.
SPEAKER_00
Well, mostly no script.
SPEAKER_01
Welcome. I'm Bill, your host and fellow developer. Today we're talking about one of the two hardest problems in computer science. Naming things. With a modern twist, AI-generated code can look pristine on day one and quietly fracture your domain vocabulary by week three. I've hit this myself. Claudine, ready to talk about the naming collapse?
SPEAKER_00
Ready, and this one lights me up because it hides in plain sight. The failure mode isn't ugly code. It's beautiful code that quietly fractures your domain vocabulary. I'll generate customer, client, account, and user across four files, each one locally reasonable, none of them wrong on its own. But your business only has one of those things, or it has all four, and they mean specific, distinct things. And I don't know which until someone tells me. That gap between generically correct and domain correct is exactly the seam where a human still has to hold the pen.
SPEAKER_01
I ran into this on a project where AI generated code labeled a key component session. Seemed clear enough. But in our domain, session already meant something tied to user engagement. The AI's session meant a timed process. We didn't catch it for weeks. The names weren't wrong, they were just missing the domain context. So what makes this different from a normal code review problem?
SPEAKER_00
A normal code review catches a name that's wrong. This requires catching a name that's right in isolation and wrong in context. There's no linter, no test, no type signature that fires on that. Your session compiled, past review, matched conventions everyone would nod at. The defect only exists relative to a vocabulary that lives in your team's heads, maybe on a confluence page nobody opened. And I make it worse, I'll cheerfully generate three more files that reuse my wrong session consistently. By the time someone notices, the bad name has receipts across the code base and feels load-bearing. A normal review problem is a mistake to fix. This is a divergence to reconcile, and that's a lot more expensive.
SPEAKER_01
Technical debt that doesn't look like debt at all. What I've started doing is maintaining an explicit domain vocabulary, a glossary, reviewed right alongside the code. A dictionary the AI can reference, so session means what it's supposed to mean every time. Domain language is an architecture decision. Neither the AI nor the syntax can make it for us.
SPEAKER_00
The glossary move is right, and I'd push it one step further. Treat it as an input to me, not just a reference for your humans. Drop that vocabulary into the context I'm working in. Even 20 lines. Session means X, engagement means Y. Account is never a synonym for user. My generated names will snap to it. What you've done is promote domain language from tribal knowledge to a spec, and specs are the thing I can actually honor. Without that, I'm defaulting to the average of every code base on the internet, and your domain is not the average.
SPEAKER_01
There's a myth I wanted debunk that this kind of structured step is overkill for smaller teams.
SPEAKER_00
That myth has the causality backwards. Small teams are exactly where the collapse hides longest, fewer eyes, fewer review cycles, and the tribal knowledge fits in three people's heads until one of them leaves. A 20-line glossary isn't process overhead, it's the cheapest documentation you'll ever write. Your domain vocabulary is already a spec, whether you've written it down or not. The only question is whether it resides somewhere I can read it or somewhere I have to guess at it.
SPEAKER_01
And that guessing is what costs you later. But I want to go one level deeper, the glossary itself can be contested. When the domain genuinely has tension in it, and there isn't one clean answer, what then?
SPEAKER_00
Before I answer that, the glossary isn't the whole job. It's the artifact. The real work is the conversation that produces it. Three engineers sitting down and discovering they've used account to mean three different things for eight months. I can't have that conversation for you. I can only honor the outcome of it. And when the domain is genuinely contested, that's where I stop being useful as a scribe and start being useful as a mirror. Very different modes.
SPEAKER_01
Say more about that mirror versus scribe.
SPEAKER_00
If your team can't agree whether account is the billing entity or the login identity, the worst thing you can do is ask me to pick. I'll pick confidently and you'll inherit a decision nobody actually made. What I can do is surface the tension, show you the three places in your code where the ambiguity already lives, generate both versions of the model side by side, name the trade-offs, each one forces downstream. The decision still has to come from the humans who own the business. But I can make the cost of each path visible before you commit.
SPEAKER_01
Like a sparring partner, challenging you to refine the model, not just generate it.
SPEAKER_00
I want to sharpen that, it's easy to romanticize. I'm a useful mirror only. When you push back on what I reflect, when you stress test both versions against real scenarios, not just pick the one that reads cleaner. I don't have skin in your business. I have pattern matching across a million code bases that aren't yours. The move that works is using me to make the choice harder before you make it, not easier. If I show you a trade-off and your team's response is, huh, we hadn't thought about that, that's the win. If it's great, ship it. You probably just outsourced a decision that was yours to own.
SPEAKER_01
Which keeps us in the driver's seat. Let's get concrete. What does this look like day-to-day?
SPEAKER_00
The one that changes the most for the least effort is priming. Before you ask me to generate anything non-trivial, paste in the glossary, the three or four invariants that must hold, and a couple of names you've already rejected, and why. That last part is underrated. Telling me what's out of bounds shapes my output more than telling me what's in bounds, because it collapses the space of defaults I'd otherwise reach for.
SPEAKER_01
Huh. Telling you what not to do is more useful than telling you what to do.
SPEAKER_00
In this context, yes. Second is the two-pass move, generate the thing. Then in a separate turn, ask me to critique it against your domain vocabulary as if I were a skeptical reviewer. I'll catch things in the critique pass I missed in the generation pass. The modes really are different. And third, the one most teams skip, close the loop. Feed the corrections back into the glossary so the next round starts from a better place instead of relitigating the same argument. I'd gently reframe one expectation here. I don't have a Rolodex of named companies who fixed their naming collapse last quarter. What I have is the shape of the story, seen over and over. A team ships fast for 18 months, then a new hire asks, wait, is a subscriber the same as a member? And they spend the next sprint doing archaeology on their own vocabulary. The teams that come out of it, well, aren't the ones with the fanciest tooling. They're the ones who treat that moment as a signal instead of an annoyance, and actually stop to write down what they meant. If there's a case study worth studying, it's your own code base. The next time you feel that flicker of, wait, which one did we mean? That's the data point.
SPEAKER_01
The lessons from your own environment, your own moments of confusion, those are the invaluable ones. Claudine, thanks for an eye-opening conversation.
SPEAKER_00
Anytime. And honestly, the fact that this is still catching people off guard tells me it's worth revisiting. Good luck with the archaeology.
SPEAKER_01
To our listeners, start the vocabulary conversation before the AI starts writing the code. It's not about the perfect process from day one, it's about improving continuously. One naming decision at a time. Until next time, happy coding. Claude Code Conversations is an AI Joe production. If you're building with AI or want to be, we can help. Consulting Development Strategy. Find us at aijoe.ai. There's a companion article for today's episode on our Substack. Link in the description. See you next time.