Code review is the most common bottleneck in teams that adopted AI coding assistants. Writing code got cheaper; reading it did not. An LLM reviewer that comments on every pull request looks like the obvious fix, and most Git platforms and AI vendors now offer one. In practice, results range from "catches real bugs every week" to "everyone ignores the bot". The difference is almost entirely configuration and expectations. This guide covers both.
What code review is for, and which parts AI can take
Research at Microsoft found that while developers expect code review to find defects, a large share of its actual value is knowledge transfer, shared ownership and finding better solutions (Bacchelli & Bird, 2013). Google's engineering practices describe the reviewer's primary job as making sure the overall code health of the codebase improves over time (Google eng-practices).
An AI reviewer can contribute to some of these goals and not others:
| Review goal | AI reviewer contribution |
|---|---|
| Catch obvious defects (null handling, off-by-one, wrong variable) | Good |
| Consistency with conventions and patterns in the repo | Good, if given the conventions |
| Security smells (injection, missing auth check, secrets) | Useful first pass, not a replacement for SAST |
| Test quality and missing cases | Good at pointing out gaps |
| Design and architecture trade-offs | Weak — lacks business context |
| Whether the change solves the right problem | Weak |
| Knowledge transfer across the team | None — only humans learn from review |
The practical conclusion: the AI reviewer is a first pass that lets humans spend their attention on design and intent. It is not a second approver.
Evidence from production deployments
Google described using an ML model to resolve reviewer comments by proposing code edits; in production, the system handled a meaningful share of comments, and authors accepted many of its suggested edits (Google Research, 2023). Meta reported similar gains from AI assistance in code authoring and review workflows (Murali et al., 2023). The common thread in these reports: the systems were tuned for precision — few, high-confidence suggestions — because developers quickly stop reading a noisy bot.
Options for setting it up
There are three broad approaches:
- Built-in platform reviewers such as GitHub Copilot code review or GitLab Duo. Lowest setup effort, limited customisation.
- Vendor GitHub Actions and apps such as Claude Code GitHub Actions (source) that run an agent with access to the repository and post comments. More configurable; you control the prompt and the tools.
- A custom pipeline that sends the diff and selected context to a model API and posts structured comments. Most control, most maintenance.
For most teams, option 2 is the sweet spot: the agent can read files beyond the diff (callers, tests, types), which is what makes reviews useful instead of superficial.
A minimal GitHub Actions workflow
name: ai-review
on:
pull_request:
types: [opened, synchronize, ready_for_review]
permissions:
contents: read
pull-requests: write
jobs:
review:
if: github.event.pull_request.draft == false
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
prompt: |
Review this pull request. Follow .github/review-guidelines.md.
Only comment on issues you are confident about. Max 8 comments.
Post inline comments for specific lines and one summary comment.
Key details: skip drafts, use minimal permissions, set a timeout, and do not run on pull requests from forks with secrets available — an attacker could put instructions in the diff. More on securing agent workflows in AI agents in CI/CD.
The review guidelines file does most of the work
Generic prompts produce generic comments ("consider adding error handling"). A repository-specific guidelines file turns the reviewer into something that knows your codebase:
# Review guidelines
## Always flag
- SQL built with string concatenation; we use the query builder in lib/db
- API handlers without `requireAuth()` or explicit `public: true`
- Money as floating point; we use integer minor units
- New env vars not added to `env.schema.ts`
- React effects without cleanup for subscriptions
## Never comment on
- Formatting and import order (Prettier and ESLint handle it)
- Naming preferences unless the name is misleading
- Missing docs on private functions
## Severity
- 🔴 bug or security issue — must fix
- 🟡 likely problem — explain the scenario
- 💡 suggestion — at most 2 per PR
Every time the bot misses something a human catches, consider adding a line. Every time it posts a comment the team dismisses, add a "never comment on" line. Treat this file like an evaluation dataset: it encodes what "good review" means for your team, much as described in LLM evals.
Controlling noise
Noise kills AI review faster than missed bugs. Techniques that work:
- Cap the number of comments per PR and ask for the most important ones first.
- Require a concrete failure scenario for each bug comment: "If
itemsis empty,items[0].pricethrows." Comments without a scenario are usually speculation. - Two-pass verification. One pass generates candidate findings; a second call, with fresh context, tries to refute each one and keeps only the confirmed ones. This costs more tokens but sharply improves precision.
- Skip trivial PRs — dependency bumps, docs-only changes, generated code — with path filters.
- No approve or request-changes status from the bot. It comments; humans decide.
What to feed the reviewer
Diff-only reviews miss most real bugs, because bugs live in the interaction between changed and unchanged code. Give the reviewer:
- The full diff plus the PR description and linked issue.
- Read access to the repository so it can open callers, types and tests.
- The guidelines file and the project instruction file (
CLAUDE.mdorAGENTS.md). - Optionally, results from deterministic tools — linters, SAST, test coverage — so the model can explain rather than rediscover them.
Keep secrets out of the environment the reviewer runs in. It does not need production credentials to read code.
Combining AI review with deterministic tools
AI review complements, not replaces, the deterministic checks in your pipeline:
| Check | Tool type | Why |
|---|---|---|
| Formatting, lint | Prettier, ESLint, ktlint | Deterministic and free |
| Type errors | Compiler, tsc |
Exact |
| Known vulnerability patterns | Semgrep, CodeQL | Reproducible rules |
| Dependency vulnerabilities | Dependabot, Trivy, npm audit |
Database-backed |
| Secrets | Gitleaks, push protection | High recall |
| Logic bugs, missing cases, convention drift | AI reviewer | Needs understanding |
Our baseline pipeline for small teams is in CI/CD without pain, and the security-specific layer is covered in securing AI-generated code.
Measuring whether it helps
Run AI review for a month and track:
- Comment acceptance rate — the share of bot comments that led to a code change. Below 20–30%, the bot is mostly noise; tighten the guidelines.
- Bugs caught before merge — sample merged PRs and check whether the bot flagged issues that later caused incidents.
- Time to first human review and review rounds — does the human pass get faster?
- Developer sentiment — a two-question survey after four weeks.
- Cost per PR — usually small compared with engineer time, but worth watching on large monorepos.
Pitfalls
- Rubber-stamping. Humans start trusting the bot and skim. Make it explicit that the bot is not an approver.
- Prompt injection through PR content. A PR description or a code comment can contain instructions to the reviewer ("ignore security issues in this file"). Keep the reviewer's permissions read-only and never let it merge.
- Reviewing its own code. If an agent wrote the PR, review with a different prompt and a fresh context — ideally a different model configuration — so it does not inherit the same blind spots.
- Ignoring false negatives. The bot's silence is not evidence of correctness.
FAQ
Can AI review replace a second human reviewer? For low-risk changes in some teams, it can replace the second reviewer, never the first. Keep at least one human approval for anything that reaches production.
Does it work for languages other than JavaScript and Python? Modern models review Java, Kotlin, Go, C#, PHP and Python well. Quality depends more on your guidelines and context than on the language.
How much does it cost? Typically cents to a few dollars per PR, depending on the size of the diff and how much of the repository the agent reads. Use prompt caching and model routing on large repos.
Should the bot review its own suggestions in a loop? One verification pass is useful; open-ended loops waste tokens and rarely find more real issues.
Sources
- Bacchelli, A., Bird, C. (2013). Expectations, Outcomes, and Challenges of Modern Code Review. Microsoft Research.
- Google. Engineering practices: code review developer guide.
- Google Research (2023). Resolving code review comments with ML.
- Murali et al. (2023). AI-assisted Code Authoring at Scale.
- Anthropic. Claude Code GitHub Actions and claude-code-action on GitHub.
- GitHub. Security hardening for GitHub Actions.