A patch can be easy to read and still be wrong. Good names, familiar patterns, and a confident explanation make code feel more trustworthy, but they do not establish that it satisfies the project’s actual constraints.

Reviewing AI-generated code benefits from the same discipline as reviewing any unfamiliar contribution. The challenge is keeping that discipline when producing a large patch takes much less time than understanding it.

Establish the intended behavior first

Before reading every changed line, write down the observable result the change is meant to produce.

For a cache change, the requirement might be: “Concurrent requests for the same file should share one read operation.” That is more useful than “Improve caching,” because you can reason about whether a particular implementation satisfies it.

Include the boundaries: what should happen after an error, after cancellation, or when the underlying file changes? If these choices are not clear, record the uncertainty instead of approving an implementation by assumption.

Follow the changed path into unchanged code

Read the changed function’s callers, shared state, and error handling. Many regressions arise at the boundary between a new implementation and an existing expectation.

Consider this simplified cache:

async function load(path: string) {
  if (cache.has(path)) return cache.get(path);

  const document = await readDocument(path);
  cache.set(path, document);
  return document;
}

This can reuse a completed read, but two concurrent calls can both miss the cache before either finishes. If the requirement is to share concurrent work, this implementation does not meet it.

Caching an in-flight promise is a possible design, but it introduces further questions. Is a rejected promise removed? Can one caller’s cancellation affect another? What happens when invalidation occurs during the read?

The review should follow the actual requirement through these states. Recognizing a familiar “promise cache” pattern is only the beginning.

Check claims against the project

Generated patches sometimes rely on an API or convention that exists elsewhere but not in this repository’s version of a library. Verify imported symbols, configuration names, supported platforms, and error contracts against the installed dependencies and current code.

Also inspect the patch’s scope. A focused behavior change should have an understandable reason for each affected area. Unexpected formatting, dependency upgrades, or unrelated refactors make the meaningful change harder to assess.

Separate newly introduced problems from existing ones. If a risky behavior was already present, it may still deserve attention, but it should not be described as a regression caused by this patch.

Ask whether the test could catch the mistake

A test that merely repeats the implementation’s assumptions can pass alongside the bug.

For the concurrent cache example, a useful test would control when the underlying read resolves, start two calls before that point, and observe whether one read or two occurred. A test that calls load twice in sequence verifies a different property.

Look for an assertion about observable behavior, and check that the setup actually reaches the state under discussion. Distinguish a test you inspected from a test you ran; only the latter provides execution evidence for your current environment.

Write findings that someone can verify

A useful review finding connects four things:

  1. The changed location.
  2. The input or state that triggers the problem.
  3. The resulting behavior and its impact.
  4. The evidence or reproduction that supports the conclusion.

For example: “Two requests started before the first read resolves both pass the cache check. Each starts a separate read, so the new implementation does not enforce the requested one-read behavior.”

That is actionable. “The concurrency could be better” leaves the author to guess which assumption failed.

Keep the final judgment yours

An AI reviewer can search widely, explain unfamiliar code, and suggest cases to investigate. Its findings still need verification. Follow the references and check that the stated trigger is possible in the real calling context.

In Yenpo, a fixed Git comparison keeps the review attached to specific revisions and remains read-only. This lets you investigate the proposed behavior before starting a separate editing task.

The aim is a decision you can explain: what the change does, why the evidence supports it, and what remains unverified.