How to Review AI-Generated Code Without Slowing Down

A practical workflow for reviewing AI-generated code: smaller diffs, explicit assumptions, real evidence, and focused human judgment.

2026-10-03 16:38:38 - Mohamad Abuzaid

How to Review AI-Generated Code Without Slowing Down

AI can produce code quickly. It can also produce a large, plausible-looking diff quickly.

Those are not the same achievement.

The uncomfortable part of AI-assisted development is rarely getting a first solution. It is deciding whether that solution deserves to enter a codebase. If the agent gives you 700 changed lines, three renamed abstractions, and a confident summary, careful review can feel slower than writing the change yourself.

That is a workflow problem, not a signal that reviewers should read every generated line with supernatural concentration.

The useful shift is to make the agent deliver a reviewable proposal: a small change, an explanation of the behavior it changes, evidence that it works, and a clear account of what it did not change. The human reviewer can then spend attention where it is uniquely valuable: intent, risk, assumptions, and fit with the existing system.

That is close to the boundary I described in Use AI to Become a Better Software Developer, Not Just a Faster One: the model can propose; engineering judgment decides what the proposal means and whether its evidence is good enough.

Why generated code feels expensive to review

AI-generated code often has a dangerous quality: it looks finished before it is understood.

The syntax is tidy. Names are reasonable. The diff may even come with tests. But a correct-looking implementation can be built on a wrong assumption: a nullable field that is actually required, an authorization check applied after data is fetched, a retry that duplicates a side effect, or a lifecycle detail that only breaks on a real device.

The larger the diff, the easier it is to miss that assumption. A reviewer starts checking names, formatting, and local details because they are visible, while the real question—“Does this change preserve the system’s contract?”—gets less attention.

This is not unique to AI. It is ordinary code-review physics, amplified by an author that can generate volume with almost no friction. Google’s public guidance makes the same trade-off explicit: a change should improve code health, but review should not demand perfection before progress can happen. Small, self-contained changes make that judgment easier.

Ask for a smaller change before asking for a review

Do not start with “implement the feature.” Start with a boundary.

For example:

Add validation for an expired invite token in the existing acceptance endpoint. Do not refactor the invitation service, change the response schema, or rename unrelated code. Show the files you expect to touch before editing.

That instruction does two things. It makes the intended behavior concrete, and it protects the review from opportunistic cleanup that happens to look sensible in a generated diff.

A small diff is not automatically safe. A three-line authorization bug can be much riskier than a 200-line UI refactor. But a small, single-purpose change gives the reviewer a chance to understand the relevant context instead of reconstructing an entire design decision from a broad rewrite.

If the task genuinely needs several changes, ask the agent to propose the slices first. Each slice should have one observable behavior and one reviewable reason to exist. GitHub’s pull-request guidance also supports breaking large changes into smaller, independently reviewable pull requests where that makes sense.

Make the agent explain its own diff

An AI-generated diff without a short explanation is an incomplete handoff.

Before you review it, require an answer to these questions:

1. What behavior changes for a user or caller?
2. Which files changed, and why does each need to change?
3. Which assumptions did you make about existing contracts or data?
4. What did you deliberately not change?
5. Which tests or checks did you run, and what do they prove?
6. What remains unverified or risky?

The point is not to treat the response as proof. Models can explain an incorrect decision fluently. The point is to turn hidden assumptions into things a reviewer can challenge.

If an agent says, “I assumed expired tokens are invalid regardless of the invite state,” you can immediately check the product rule and existing service behavior. Without that statement, the same assumption may sit inside a conditional that looks perfectly normal.

This is also a cheap way to catch scope drift. “I updated the shared token utility for consistency” is a useful warning when the requested work was only one endpoint.

Review behavior before implementation details

Start the review above the diff.

First, restate the contract in your own words: what should happen on the happy path, at the boundaries, and when something fails? Then inspect the code for evidence that it does exactly that.

For a change that accepts an invitation, my first pass would be about questions such as:

Only after that pass do I spend time on implementation choices: whether a helper belongs in this module, whether the naming is clear, or whether the control flow can be simpler.

This ordering matters. It prevents a review from becoming a style conversation while an incorrect product rule slips through. Google’s reviewer guide similarly asks reviewers to consider design, functionality, complexity, tests, naming, comments, style, documentation, and the surrounding system—not merely the changed lines.

For security-sensitive work, add a deliberate threat-model pass. AI is particularly good at producing code that handles the expected path while forgetting authorization, data exposure, injection boundaries, or unsafe defaults. The OWASP Secure Code Review Cheat Sheet is a useful prompt for that pass; it should complement your application’s own security rules, not replace them.

Require evidence, not confidence

“I ran the tests” is not enough review information. Ask what ran, what it covers, and what it does not cover.

For the invite-token example, useful evidence might be:

For a UI change, that may be a before-and-after screenshot plus the focused test. For a migration, it may be a migration plan, rollback note, and a query against representative data. For an API integration, it may be a mocked error response and a trace of the failure path.

Evidence has limits. Passing tests only show what the tests exercised. A green build does not prove that the requirements are right, that a user flow makes sense, or that an unsafe assumption was never introduced. But evidence dramatically narrows the part of the problem a human must simulate mentally.

Automated checks belong here too. CI can run builds, tests, linting, static analysis, and code scanning consistently. GitHub describes status checks as a way to surface results from these systems and, when required for a protected branch, to prevent merging until they pass. That is valuable mechanical pressure—but it is not a substitute for a reviewer deciding whether the change is the right one.

Use pushback as a second review round

The first implementation is a proposal, not a conclusion.

After the agent responds to your initial review, ask a second set of questions:

Is this the simplest correct solution?
Which assumption is most likely to be wrong?
What failure mode is still not covered by the tests?
What smaller alternative would avoid this new abstraction?
What would make this unsafe to deploy today?

This is valuable even when you do not accept the agent’s answer directly. The questions make it compare options, name uncertainty, and inspect its own solution from a different angle. They also create a compact record of why the eventual change was accepted.

Do not turn this into an infinite debate loop. Use it when the change crosses a risk boundary: auth, payments, data migration, concurrency, destructive operations, public APIs, or a core workflow. For a small internal UI fix with strong tests, one well-scoped review may be enough.

Let automation remove mechanical noise

Formatting arguments, import order, obvious lint failures, and routine static checks are poor uses of scarce review attention. They should be handled automatically before a reviewer opens the diff.

That does not mean “automate everything and trust the green check.” It means reserving human review for questions that tools cannot settle reliably:

Teams need to tune that boundary. An early-stage product may accept a narrow, well-tested implementation that is not perfectly generalized. A regulated or high-impact system may require stronger review, independent validation, or manual test evidence. The workflow should adapt to the risk, not apply the same ritual to every line of code.

A review request you can reuse

Here is a prompt I would use before reviewing an agent’s change:

Implement only the requested behavior. Keep the diff focused and avoid unrelated refactoring.

Before presenting the change, provide:
- a one-paragraph behavior summary;
- the files changed and why;
- assumptions about existing contracts;
- tests/checks run, their exact results, and gaps;
- known risks or unverified cases;
- the smallest alternative you considered and why you did not choose it.

If the change touches auth, data, money, concurrency, or a public API, call that out explicitly.

You will still need to read the code. You will still find mistakes. The difference is that you are no longer treating the diff as a mystery novel whose plot must be reconstructed line by line.

AI writes a proposal. The reviewer validates the reasoning and the evidence.

That is how review stays fast without becoming careless.

What evidence would make you comfortable approving an AI-generated change in your codebase—and what would still make you stop the merge?

Sources

More Posts