Being confidently wrong is one of the biggest red flags in engineering. Long before AI entered the picture, I was sitting in architecture reviews and feature spec discussions watching smart, experienced people propose things that sounded completely right and weren't.

Not wrong in the way that comes from carelessness or inexperience. Wrong in the way that comes from knowing a lot. The kind of wrong that sounds so convincing in a meeting that everyone nods along, because the person saying it has been right about hard things before. They have the track record, the technical depth, the ability to explain their reasoning clearly and persuasively.

And they're still wrong.

Intelligence and confidence travel together. The more someone knows, the more certain they tend to sound, even when they're operating outside the boundaries of what they've actually validated. They pattern-match from past experience. They fill in gaps with assumptions that feel like facts. They build an argument so coherent that questioning it feels like you're the one missing something.

One of the skills I've honed over the years is detecting this, in myself and in others. In business integrations, feature specs, architecture reviews, product designs. You develop an instinct for when someone's confidence has outpaced their evidence. You learn to probe the edges of their certainty, to ask the questions that test whether they've actually validated what they're proposing or just built a convincing story around their assumptions.

This problem is about to get much harder.


A stone sits solid and certain above the waterline, its reflection dissolving into something less sure of itself

"To know that you do not know is the best. To think you know when you do not is a disease." -- Tao Te Ching


LLMs are the most confidently wrong entity you will ever work with.

They never hesitate. They never hedge. They never say "I'm not sure about this." Every response arrives with the same calm authority whether it's right or completely fabricated.

I've seen an LLM generate a perfectly reasonable-looking API integration that called endpoints which didn't exist. The code was clean. The error handling was thoughtful. The comments explained the logic clearly. And the entire thing was built on a hallucinated API surface.

I've seen them recommend package dependencies that don't exist. I've watched one decide mid-session that a problem we were debugging required different documentation entirely, then confidently rewrite working code to match its new theory. I've seen one mistake OAuth test credentials for basic auth and try to rewrite our entire authentication integration, presenting the change as the obvious correct solution.

Every one of those looked exactly like the correct answers that came before them.

With a confident engineer in a meeting, you can at least read the room. You can watch for the slight hesitation, the hand wave over a detail, the moment where conviction outpaces evidence. With AI, there are no tells. And you used to have to be in that meeting, or on that pull request, or in that Slack thread to encounter confident wrongness. Now anyone with a chat window can generate a polished proposal, a detailed feature spec, a technical architecture document in minutes. Engineers, product managers, executives, it doesn't matter. The output looks professional, reads confidently, and may or may not be grounded in reality. And most people reading it can't tell the difference.

AI is close to the right answer most of the time. If it were wrong half the time, nobody would trust it. But it's right often enough that the remaining 5% becomes nearly invisible. You stop checking because the last ten answers were perfect. And then the eleventh one isn't.


Catching that 5% is a skill built from experience, mostly the painful kind. You ship something that looked correct in every review and watch it fail in production. You trust an expert's recommendation and later realize it was built on assumptions nobody questioned. Every experienced engineering leader carries a library of these moments. They're what shift your default question from "does that sound right?" to "how do you know that?"

"Does that sound right?" evaluates coherence. Whether the logic flows, whether the conclusion feels reasonable. Confident wrongness passes that test every time. "How do you know that?" probes the foundation. What evidence backs the claim, what assumptions haven't been tested, where actual knowledge ends and extrapolation begins.

There are methods that try to formalize this. The 5 Whys, devil's advocate, red teaming. I've used variations of all of them with engineers and product managers over the years. They work because they force you past the confident surface and into the unexamined assumptions underneath. Most people haven't pressure-tested their own thinking that deep. They don't realize how much of their certainty rests on things they've never actually validated.

Creating space for doubt in a fast-moving team is hard. Nobody wants to be the person who slows things down. When the code looks right and the spec makes sense and the deadline is Thursday, stopping to ask "but how do we know this is actually correct?" feels like friction.

It is friction. And that friction is where correctness lives.

The leaders who catch the 5% aren't smarter than everyone else. They've just been wrong enough times to know that coherence isn't proof.


Our generation of engineers might be the last one with a natural comparison point. We remember what it felt like to not know something and have to sit with that. To dig through documentation, try things that didn't work, build wrong mental models and slowly replace them with better ones.

Engineers coming up now will have a fundamentally different relationship with not knowing. When the default response to any question is a fluent, confident, immediate answer from an AI, you get less practice sitting with "I don't know." You get fewer reps developing the feeling that something sounds too neat, too complete, too sure of itself. The tool always has an answer, so the experience of not having one becomes rare.

I want to move faster, not slower. AI is incredible for prototyping, for exploring ideas early, for getting a proof of concept in front of stakeholders before you've committed to an architecture. In the early stages of product development, less scrutiny is fine. That's the whole point of a prototype.

The danger is when that same lack of scrutiny carries over into production systems. When critical infrastructure gets updated because an AI-generated solution looked right and nobody with the experience to question it actually did. When product managers and business leaders get sold on a proposal that reads beautifully but falls apart the moment an engineer with real context reviews it. When the confidence of the tool replaces the judgment of the team.

If skepticism is a skill that gets sharpened by experience, what happens when the technology removes the very experiences that sharpen it?

AI, by being right most of the time and sounding right all of the time, reduces the opportunities to learn this the hard way. You skip the small failures that teach you to question. And then the first time it really matters, you don't have the instinct to pause.


We've been dealing with confident wrongness our whole careers. We learned to catch it in people because we had to. Now we're learning to catch it in AI. Even with experience, it's getting harder. The output keeps getting better, the confidence keeps getting more convincing, and the line between right and almost-right keeps getting thinner. If we're struggling with it, how can we expect the next generation to catch it without the same experience?

Maybe it won't matter. Maybe in a year or two these tools will be good enough that this whole internal struggle is for naught. Maybe engineering leadership becomes something else entirely, and what AI empowers us to build outweighs the failures it produces along the way.

But until then, the challenge for engineering leaders is pushing AI tooling forward with a heavy dose of skepticism and critical thinking. Moving fast and questioning everything at the same time.