AI Code Quality Depends on Who Is Checking the Output
An answer comes back from the model. It is formatted, it compiles, and it has a short explanation attached. It looks exactly as confident as the last answer, which was right.
Whether this one is right is the question every team building with AI now faces dozens of times a day, and the evidence says reading it carefully is not how you find out.
What does the Dunning-Kruger effect have to do with AI code?
The Dunning-Kruger idea, applied to AI, is usually about the person: someone who does not know enough to know they are wrong ships code they cannot judge. That risk is real. It is also only one axis, because the model has a confidence of its own, and it carries no information about whether this answer is right.
The psychology behind Dunning-Kruger is debated, and nothing here depends on it being true. It is useful as a shape everyone recognizes: confidence and competence drifting apart. Add the model as a second, independent source of confidence and the picture becomes a grid.

Read it one cell at a time:
- Error caught. Someone who knows what to check meets a wrong answer and has a real chance of stopping it.
- Leverage. The same person meets a right answer and moves faster than they could alone.
- Fragile luck. Someone new to the work gets a right answer. It works, for now, and nobody can say why.
- Blind leading blind. A wrong answer meets a reader who cannot tell. Nobody in the loop knows anything is wrong.
The model decides the column, fresh on every answer, and it never tells you which one you are in. The row is decided before the work starts, by who does it. Most of what we wrote about why vibe coding hits a wall is the bottom row, described from the inside.
Can you tell when AI code is wrong by reading it?
Not reliably, even when you program for a living. In a July 2026 experiment, 86 Python programmers judged AI-generated checks on code. They were right 74 percent of the time when the check was correct and 49 percent of the time when it was flawed, which is a coin flip, with about the same confidence both times.
The study, by Zhanna Kaufman, Yuriy Brun, Adithya Murali and Madeline Endres, used assertions: short lines that state what the code is supposed to do.[1] Participants had to decide whether each one was correct and complete.
The gap between the two accuracy numbers is not noise. The odds of judging a correct check accurately were nearly three times the odds for a flawed one.
The detail that matters most is what did not move.
Confidence did not fall with accuracy. The readers had no internal signal that they had moved from the right column to the wrong one.
The feeling of having checked is not a check.
This complicates the grid in a useful way. These were programmers, the top row by any hiring standard, and they still could not detect flawed output by reading it. Expertise buys a better starting guess and the habit of checking. It does not buy a detector.
Does an explanation make AI output safer?
No. In the same study, plain-language explanations attached to the AI's output gave no overall improvement in accuracy. Low-quality explanations made readers less accurate while making them more confident, from 3.99 to 4.25 on a five-point scale. The fluent reason next to the answer pushed people toward the wrong call.
Most AI tools now explain themselves by default, and the explanation is usually the most reassuring thing on the screen. It reads like a colleague walking you through the change. The paper's authors put their conclusion plainly: "contrary to common assumptions, AI assistance may not improve the reliability of code comprehension and review."[1]
That is how a confidently wrong answer passes for a confidently right one. The explanation is fluent in both columns, so it cannot tell you which column you are in, and a weak one leaves you surer of the wrong call.
How good is AI-generated code, really?
Good often enough to earn trust, and uneven in ways a reader cannot see. The quality is not uniformly mediocre. It is strong on some kinds of task and weak on others, so a reviewer who watched it succeed twice will reasonably relax right before the third answer fails.
We covered one measured version of that unevenness in the security pass rate that has not moved in two years: some kinds of weakness are handled well almost every time and others almost never, and an average hides the difference.
The grid is the reader's side of the same problem. Uneven output meets a reviewer whose confidence does not track it.
What actually changes the outcome?
Moving the check out of the reader's head. If reading the output and feeling sure does not catch flawed work, the catch has to come from something that does not care how fluent the answer sounds: tests that fail, a checklist that forces specific questions, and a reviewer who has not read the model's explanation.
The people building these tools describe the same answer. Anthropic's Claude Code lead listed what his team uses to hold AI-written code to a higher standard.[2]
Production code written by Claude should have a higher bar than if it was written by a human.
His list was lint rules, tests, end-to-end tests, fuzzers that run daily, and automated code and security reviews. Every item is a check that works whatever the reviewer happens to feel. In practice that comes down to four habits:
- Tests that can fail. Write or keep the test that would break if the answer were wrong, and run it before trusting the change.
- A written checklist. Review against specific questions, not a general impression of whether it looks fine.
- A second reader without the explanation. Hide the model's rationale from whoever reviews the change, since it moves judgment the wrong way.
- Small changes. Keep each change small enough that a wrong answer shows up quickly and cheaply.
We made the architectural version of this argument in why the human in the loop has to be designed, not assumed. The study supplies the number that post was missing.
Who you put on the work still matters
The hiring decision is the row, and it still counts. Someone who knows what to check has a better starting guess about where AI output fails and builds those checks in without being asked. That is why skill still decides outcomes even as the tools improve. What the study rules out is treating their judgment alone as the safety net.
The grid is not an argument against building with AI. It is the reason judgment did not get cheaper when generating code did, and why the best teams now put that judgment into tests and process instead of trusting how the output feels.
References
- ^1.Zhanna Kaufman, Yuriy Brun, Adithya Murali and Madeline Endres (arXiv preprint), “Programmers Are Poor and Overconfident Judges of LLM-Generated Assertions” (July 9, 2026)
- ^
Frequently asked
How good is AI-generated code?›Good enough often enough that the failures are hard to see, which is the real problem.
Can you tell when AI-generated code is wrong just by reading it?›Measurably not, even for programmers. In a controlled experiment with 86 Python programmers, participants judged correct generated checks accurately 74 percent of the time but flawed ones only 49 percent of the time, which is chance.
Does an explanation make AI output more trustworthy?›It makes output feel more trustworthy without making the reader more accurate.
Does the Dunning-Kruger effect apply to AI coding?›Partly, and only along one axis. The familiar version is about the person who does not know enough to know they are wrong.
How should you review AI-generated code?›Move the check somewhere other than your own reading of it. Tests that fail when the code is wrong, a review against a written checklist, a second reader who has not seen the model's explanation, and changes small enough that a wrong answer shows up quickly.
Let’s build it together.
We turn clever prototypes into production systems people can rely on. If you’re building with agents and want a hand making it real, leave your email and we’ll be in touch.