A defense of human code review I wish I saw more often, especially in light of the concerns people have about cognitive/comprehension debt: comprehension redundancy. At the end, if taken seriously, at least two people understand how the feature works (even if that number is, on average, trending closer to between one and zero). Ideally at least one of the two also comes away with a better understanding of the wider system and how the feature fits into or stands out from that landscape.
If only one person knows the code, then PR time isn't going to save you.
I'm tech lead and I basically don't review PRs, and I tell people this, with a caveat - if you can tell me what you specifically want me to review, for what specific purpose, I'm happy to!
So "can you review this bit for race conditions" is great, love it. This forces people to actually think about what in their code they should be suspicious of, if anything.
"Can you review this" [link to PR] is getting a rubber stamp because humans have never been good enough at "just spotting bugs" to make this worth it and now any LLM is better than a human.
And for the purpose of understanding - PR time is too late. I have not reviewed a PR ("for real") in a long time and yet I could tell you how every system my people have built works down to a very fine level of detail. And it's because _we talk to each other!_ We don't just chill in the same slack channel and code independently, we all value each others brains and want each others inputs because we know it will improve our product and we value what perspectives others will bring.
Trying to learn via PR is a sad substitute for real collaboration and teamwork.
I consider PRs to be primarily a defense of the architecture, and to a lesser degree a general sanity check. It’s also a useful opportunity to enforce automations are being run
There's been a lot of talk about the purpose of code review recently. It makes sense in the face of AI. Heres a link that was submitted a little while ago: https://mathstodon.xyz/@mjd/115096720350507897
And in response I wrote a non-exhaustive checklist of things that a code review can look for:
- Does it functionally achieve what it sets out to (as per tacker issue or PR description)?
- Does it have extraneous code? Leftover debug prints, private API keys etc...
- Does it have any obvious defects? Memory leaks, un-handled edge cases, security flaws, obsolete API calls, etc...
- Could it be more understandable? Add/remove abstractions, better variable/method names, more/less functional etc...
- Is the style consistent with the codebase and/or style guidelines?
- Are there obvious performance improvements? Hashset instead of list, lazy evaluations, etc...
- Is it sufficiently well tested?
I think LLMs are okay at most of these, and worst at the first.
“Net new” is one it seems to be particularly bad at.
I don’t think I have ever even once seen an LLM solve a problem related to overengineering by simply removing the overengineering. They always choose to add more epicycles and further compound the complexity.
The thing is, leadership relies on the people actually building the software to provide concrete, accurate feedback about cost. Without that they have no chance to do a decent cost/benefit analysis.
But AI has engendered a collapse in developers’ ability to actually do that. Those of us who are stuck on the vibecoding bandwagon have lost the comprehensive understanding of the systems under our care that we need to understand and explain the quality and maintenance implications of a change.
Worse, if you happen to lose your mind and suggest the initial development cost is anything more than ~zero, your friendly neighborhood Claude keener will publicly shame you for not having sufficient faith in the Glorious Agentic Future. Product leadership will then have no choice but to side with them, not necessarily because they agree, but because they, too, are aware that we’re still in the phase of the hype cycle where openly questioning said hype is a career-limiting move.
I am contemplating code review within my own organization, and the question I return to is:
> Does this organization prioritize human learning?
That has been my primary motivator for code reviews. I want to teach and learn from others, especially given the decreasing levels of collaboration due to increased AI usage.
The sad truth is that all of my feedback just goes straight to agents. Maybe 10% is reacted to by a human, so I’m left wondering if there’s any value to a real review aside from poorly training robots to do my job, and further atrophying the abilities of my team members.
What makes you think so?
It's not my experience, so I'd like to hear what kind of software you work with/teams and what is leading you to understand that code doesn't need review.
In my experience, automated code review is more pointless than ever.
We have all the linters, tests, and AI writing code for us. I don’t need the left hand to tell the right hand it did a good job. I’m very certain my code runs when I push the PR.
What I need now is architectural, long-horizon and business perspective.
For those of us who do utilize the pre-existing tools and the coding agents effectively, yeah automated code reviews eventually become redundant. But they're still around because there are a lot of devs who don't even glance at what the agents are writing for them. Maybe partly because they don't care to streamline their workflow, or maybe because they've never been particularly good at writing code, and don't actually understand what the agent is giving them.
Lately I've seen some pretty glaringly obvious issues caught during the preliminary automated code review, and the issues seem to be coming from individuals who don't actually understand what the code is doing.
For the rest of us who are actually using the whole stack effectively, the automated code review is essentially a CI gate to protect the repo from the devs who don't know what they're doing.
code reviews are just a gateway that can be whatever you want it to be, and is kind of legacy human coder thing now. At its basics it was a point to catch problems that humans were likely to make / would more likely make if they knew there wasn't a review. Now you can target it for AI mistakes. You can build your code review skills (AI skill) to be incredibly thorough. The points made in the article don't really seem like things you need to do at the "legacy" gateway of code review. Things are different now. Code is cheap. Validation, Product Coherence, Governance need to be done early and throughout.
I disagree. Most of my prompts are “implement <issue-id>”. The issue has a user story and acceptance criteria that the LLM uses to write a plan and ultimately produce code. None of that matters since the code artifact is what actually gets pushed in the repo, thus the code is what should be reviewed.
I believe code review is important (the article articulates some of the reasons) but I have to admit, having different models review a PR before passing to a human has proven valuable.
Call for a style guide meeting every time you see feedback like this, and write down what everybody agrees on. You’ll never have to do it again after 3-4 of those. Problem solved.
Still not solved? Guess it was really about the commas and not the value delivered anyway, so do whatever you feel like.
I actually think this might be a false positive. I think it got tripped up by the higher than average use of jargon (which I don't mind here because the article itself flowed well and raised good points).
Gpt-zero scores "human", and I've always found it to be a better judge
I'm tech lead and I basically don't review PRs, and I tell people this, with a caveat - if you can tell me what you specifically want me to review, for what specific purpose, I'm happy to!
So "can you review this bit for race conditions" is great, love it. This forces people to actually think about what in their code they should be suspicious of, if anything.
"Can you review this" [link to PR] is getting a rubber stamp because humans have never been good enough at "just spotting bugs" to make this worth it and now any LLM is better than a human.
And for the purpose of understanding - PR time is too late. I have not reviewed a PR ("for real") in a long time and yet I could tell you how every system my people have built works down to a very fine level of detail. And it's because _we talk to each other!_ We don't just chill in the same slack channel and code independently, we all value each others brains and want each others inputs because we know it will improve our product and we value what perspectives others will bring.
Trying to learn via PR is a sad substitute for real collaboration and teamwork.
And in response I wrote a non-exhaustive checklist of things that a code review can look for:
- Does it functionally achieve what it sets out to (as per tacker issue or PR description)?
- Does it have extraneous code? Leftover debug prints, private API keys etc...
- Does it have any obvious defects? Memory leaks, un-handled edge cases, security flaws, obsolete API calls, etc...
- Could it be more understandable? Add/remove abstractions, better variable/method names, more/less functional etc...
- Is the style consistent with the codebase and/or style guidelines?
- Are there obvious performance improvements? Hashset instead of list, lazy evaluations, etc...
- Is it sufficiently well tested?
I think LLMs are okay at most of these, and worst at the first.
Is there already a pattern or code on in in the existing codebase that handles this functionality,
Do we really need net new code to achieve this functionality?
Can existing code be extended or abstracted to more cleanly implement this feature or functionality.
I don’t think I have ever even once seen an LLM solve a problem related to overengineering by simply removing the overengineering. They always choose to add more epicycles and further compound the complexity.
- Is the change architecturally right?
Particularly the latter LLMs seem still pretty useless at.
Although I do think that LLMs have made it much easier to justify writing low-value code which can make this more common now.
But AI has engendered a collapse in developers’ ability to actually do that. Those of us who are stuck on the vibecoding bandwagon have lost the comprehensive understanding of the systems under our care that we need to understand and explain the quality and maintenance implications of a change.
Worse, if you happen to lose your mind and suggest the initial development cost is anything more than ~zero, your friendly neighborhood Claude keener will publicly shame you for not having sufficient faith in the Glorious Agentic Future. Product leadership will then have no choice but to side with them, not necessarily because they agree, but because they, too, are aware that we’re still in the phase of the hype cycle where openly questioning said hype is a career-limiting move.
> Does this organization prioritize human learning?
That has been my primary motivator for code reviews. I want to teach and learn from others, especially given the decreasing levels of collaboration due to increased AI usage.
The sad truth is that all of my feedback just goes straight to agents. Maybe 10% is reacted to by a human, so I’m left wondering if there’s any value to a real review aside from poorly training robots to do my job, and further atrophying the abilities of my team members.
We have all the linters, tests, and AI writing code for us. I don’t need the left hand to tell the right hand it did a good job. I’m very certain my code runs when I push the PR.
What I need now is architectural, long-horizon and business perspective.
Lately I've seen some pretty glaringly obvious issues caught during the preliminary automated code review, and the issues seem to be coming from individuals who don't actually understand what the code is doing.
For the rest of us who are actually using the whole stack effectively, the automated code review is essentially a CI gate to protect the repo from the devs who don't know what they're doing.
> What I need now is architectural, long-horizon and business perspective.
That's exactly what these tools are now good at. They have a huge gap when fixing these issues properly but they can spot these issues no problem
„The indent is wrong here“
„Comments should end with a period“
Because this kind of feedback is and was always easy.
Still not solved? Guess it was really about the commas and not the value delivered anyway, so do whatever you feel like.
Gpt-zero scores "human", and I've always found it to be a better judge
Something is missing in the new ai bot review paradigm we’ve all sleepwalked into.
I’ve been building Archme.io for this reason. PR reviews for the age of AI