
How to Review AI-Generated Code for Bugs and Security Risks
Two tests pass. The new invoice export returns a CSV. The pull request looks ready until a reviewer asks one question the existing tests never asked: what happens when a signed-in customer requests an invoice belonging to another tenant?
In a controlled example we ran for this guide, that request returned HTTP 200 and the other tenant’s data. The code compiled, the happy path worked, and the login check was present. None of those facts established that the caller could see the particular invoice.
This article shows how to make that kind of review decision. It applies whether a developer pasted one AI suggestion or an agent prepared a complete branch; the security and behavior standard does not change with authorship.
The generated code may be polished, and the accompanying tests may be green. Your job is to trace what the change allows, compare it with the intended contract, and find evidence for the paths that would matter if the assumption were wrong.
The example here is an author-constructed Python fixture, not code produced by a named AI product or taken from a real customer system. We executed it and retained the source and output. That distinction matters: the observed failure demonstrates a review method and a specific flaw in a small program, not the rate at which coding assistants produce vulnerabilities.
In this guide
- What does reviewing AI-generated code actually mean?
- Start with the contract before opening the diff
- Follow the caller, resource and data beyond changed lines
- The review record: a passing PR that crosses a tenant boundary
- Review correctness before reaching for a vulnerability label
- Map security risks to the path that creates them
- Treat tests and scanners as evidence with a scope
- Decide what to ask for, escalate or approve
- Use W.O.R.T.H. to judge AI assistance and review load
- Make the review repeatable without turning it into theater
- How We Researched and Tested This Guide
- Frequently Asked Questions
- Final Thoughts
What does reviewing AI-generated code actually mean?
Reviewing AI-generated code means deciding whether a change implements the requested behavior safely within the real application’s contracts. That requires more than reading added lines.
A useful review combines the task or issue, the diff, relevant unchanged code, the paths through which a user can reach the new behavior, the checks that actually ran, and the consequences of a wrong result. GitHub’s guide to reviewing AI-generated code recommends functional checks, context and intent review, dependency scrutiny, human collaboration and automated analysis; none should be mistaken for a substitute for the others.
Consider a new endpoint that calls an old helper. The helper may be perfectly correct for the internal billing job it served yesterday. Exposing its output to a customer-facing route today changes the security question without changing the helper’s source file.
A reviewer who limits attention to highlighted lines can miss the new path from an untrusted caller to a sensitive result. Our overview of AI coding assistants explains the broader development workflow; this guide begins after a change has been proposed and asks whether it should be accepted.
The goal is not to prove the absence of every possible bug. A pull-request review has limited time and incomplete knowledge. It should identify the behavior that is meant to be true, prioritize the consequences if it is false, and request enough specific evidence to support an honest merge decision.
A passing suite proves that the tests which ran produced their expected outputs under those conditions. It cannot prove a missing case, and an AI-written test can share the implementation’s mistaken assumption.
There is a useful discipline behind traditional peer review: another person reconstructs the change without relying on the author’s confidence. AI can make code generation faster, but it does not make that independent reconstruction obsolete.
A human author who cannot explain an agent’s changes has not transferred enough understanding to the team. Our guide to how AI coding agents work covers the agent’s tool loop and handoff; here we examine the review of the delivered code and its evidence.
Start with the contract before opening the diff
Write the intended result in a sentence that can be falsified. “Add invoice export” is a feature label, not a reviewable contract.
“An authenticated member of tenant alpha may download alpha’s invoice as CSV; a member of beta may not retrieve it” tells the reviewer which success and denial paths matter. If the requirement says only “users can download invoices,” the missing ownership rule is a product-policy question, not something the model or reviewer should silently invent.
The issue, acceptance criteria, existing API contract and relevant policy may disagree. Resolve the disagreement before approving implementation details. A developer might assume that knowing an invoice ID grants access; the product may instead require tenant membership, an invoice-specific role or a narrower account relationship.
Neither a test generated from the code nor a passing demo can choose the right rule. When the decision affects customer data or obligations, ask the service or policy owner to specify it and record the answer in the PR.
Next, establish the scope of the submission. Read the PR description and file list, note new public entry points, data-model changes, migrations, dependency and lockfile changes, permissions, configuration, workflows and tests.
Ask for the exact commands run and their outputs, not a sentence saying “all checks passed.” If a generator rewrote a large file to add one endpoint, request a smaller diff or an explanation of the mechanical changes; reviewing unrelated churn at the same time raises the chance of missing the actual behavior change.
Find the closest existing implementation in the repository. A neighboring export route may use a scoped query, a service-level permission check, column allowlist and audit event that the new route omitted. The comparison is not a demand for identical style.
It reveals hidden contracts: where the system enforces access, how it returns errors, which sensitive fields it withholds, and what the tests normally assert. An invented parallel implementation is a review finding when it bypasses a control, even if its formatting looks cleaner than the older code.
Then verify any fact that lives outside the PR author’s narrative. A method name must exist in the version in the lockfile, not merely in the latest online documentation. A configuration option must have the assumed default in the deployed environment. A database API must provide the transaction or isolation behavior that the code relies on.
If AI introduced a dependency, confirm its exact package identity, source, license compatibility, release state and necessity before treating a successful install as approval. Our comparison of coding assistants addresses product choice; the review question here is whether the submitted code’s external assumptions are true in this repository.
A strong PR description makes these checks cheaper: it states the intended behavior, important non-goals, affected boundaries, verification commands and known gaps. It should also tell the reviewer which files an agent inspected when that explains a design choice. The description is a map, not evidence by itself. If it claims an integration test passed, the check run or retained output should show the command, environment and result.
Turn the file list into a review map
Read the diff in the order a request travels through the system, not necessarily the alphabetical order of filenames. For a web feature, start with the route or UI action, follow the service call into data access, then inspect the serializer, error path and tests.
A file list may show a controller and a test while hiding the unchanged permission helper and database model that determine the result. Write down those dependencies before commenting on individual lines.
Separate three kinds of change. A behavior change alters who can do what or what result they receive. A mechanical change moves code, renames a symbol or reformats files without intending to change behavior.
An environment change modifies a dependency, schema, feature flag, permission or deployment setting. They require different evidence. A rename can be checked for call-site completeness; a new authorization rule needs policy cases; a schema migration needs data compatibility and rollout analysis.
Watch for broad changes that hide a small security decision. If an agent regenerated an entire API client, ask for the generated source, version and reason, then review the intentional surface separately from generated churn.
If a PR combines an export feature with unrelated refactoring, a reviewer may spend attention on moved lines and miss the new public path. Splitting the patch is useful when it makes each behavior falsifiable and revertible; it is not merely a preference for small line counts.
The comparison with a neighboring route should be precise. Trace the old route’s middleware, scoped query, output fields and denial response; then identify every difference in the new one.
A missing audit event may be a policy requirement or merely an optional observability choice, while a missing object permission can expose data. Classify the difference by consequence before asking for change. This prevents the review from becoming a style debate when the actual issue is a trust boundary.
Name the evidence you expect before more tools run. If the change touches tenant access, request a wrong-tenant route test; if it touches a migration, ask how old records are transformed and how rollout is reversed; if it touches an SDK, verify the lockfile version and the API call against that version. A request for “more tests” without a missing behavior claim usually produces more happy-path tests and little new confidence.

Follow the caller, resource and data beyond changed lines
For a consequential operation, trace one complete path. Identify the caller and selected object, then locate the permission decision, data load, output and failure behavior. A missing link in that chain is a concrete review question.
OWASP’s Secure Code Review Cheat Sheet treats architecture, data flow, trust boundaries, business logic and modified security controls as part of code review. The method works for an AI-assisted diff because the risk lives in the behavior, not in the way the characters were typed.
Begin at the first boundary where the application accepts an external value: route parameter, request body, uploaded file, webhook, queued job, model output or configuration. Follow that value through validation and transformation to database reads, filesystem paths, external calls, logs and responses.
Inspect the actual call sites and definitions of helpers that appear to “handle it.” A helper called getInvoice may retrieve by ID; its name does not prove it filters by tenant. A service called authorizeExport may check only a role; that does not prove the selected record belongs to the caller’s tenant.
Authentication answers “who is the caller?” Authorization answers “may this caller perform this action on this particular resource?” OWASP’s Authorization Cheat Sheet calls for explicit permission decisions, deny-by-default behavior and validation on every request.
In a multi-tenant system, a login guard followed by an unscoped ID lookup is therefore a suspicious boundary: an authenticated person may still ask for another tenant’s object. The identifier’s unpredictability is not a substitute for permission, and hiding a link in the interface does not secure the API.
Check where authorization happens relative to data access and side effects. A route that builds a report before rejecting the caller may have already read sensitive records or triggered an expensive job. A background worker that reuses an internal helper may not inherit the middleware that protected the user-facing request.
A cache keyed only by invoice ID may serve a previously authorized response in the wrong tenant context. These are different mechanisms; investigate the ones the actual change introduces rather than attaching every possible vulnerability label to every PR.
The destination matters as much as the initial read. An endpoint may check ownership before returning a record yet include internal notes in its serialized object. A log statement may expose a token or personal data even when the HTTP response is safe. A CSV export can disclose more columns than the screen displays; CSV content also raises a separate formula-handling question when data is opened in spreadsheet software.
State the expected output fields and inspect the serializer or template that produces them. In our small fixture, the output includes only an ID and amount; the demonstrated defect is unauthorized access to that amount, not a claim that its CSV contains other fields or formulas.
This is why review depth should track the trust boundary crossed. A local formatting change may need focused behavioral and accessibility checks. A new customer-facing download, billing action, file upload or permission rule deserves a trace through access and output paths.
Security review is part of the broader engineering decision; our AI cybersecurity primer covers detection and response at a different layer. Finding the boundary in a PR prevents a problem before the operational security team has to investigate it.
The review record: a passing PR that crosses a tenant boundary
We built a small, author-constructed program to make the evidence problem visible. Its data store has two records: inv-a belongs to alpha and has amount_cents 12500; inv-b belongs to beta and has amount_cents 28700.
An existing lookup_invoice(invoice_id) helper returns a record by ID without knowing the caller. That can be a legitimate internal lookup contract. A new export route authenticates the caller, reuses the helper and emits CSV, but does not check whether the caller’s tenant owns the record.
The unchanged helper performs the equivalent of INVOICES.get(invoice_id): it looks up an invoice by ID and has no actor parameter. The new export function first rejects a missing actor with 401. It then calls the helper, returns 404 for an unknown ID, and otherwise returns 200 with _csv_for(invoice_id, record). These are the exact branches relevant to the review; none compares the record’s tenant with the caller’s tenant.
That missing comparison is easy to miss if the reviewer sees the login guard and stops tracing. The helper’s behavior was not necessarily wrong for its old internal caller. The new route turns an ID-only lookup into a customer-visible response, changing who can benefit from the helper without changing the helper itself. The reviewer has to follow that new use and ask where the resource-level permission is enforced.
We ran the saved fixture with Python 3.12.14 and the standard-library unittest runner on 28 September 2026. The existing checks asserted that alpha could export inv-a and that an unauthenticated caller was rejected. Both passed.
The targeted expectation then asserted that alpha asking for inv-b should receive 403; it failed because the vulnerable route returned 200. A direct call returned the CSV body ‘invoice_id,amount_cents’ followed by ‘inv-b,28700’. The reviewed version adds an ownership comparison before serialization; its own-tenant, unauthenticated and wrong-tenant cases all passed.
| Check actually run | Outcome in the fixture | What the result establishes |
|---|---|---|
| ExistingChecks: same-tenant export and unauthenticated denial | 2 tests passed; exit 0 | The own-record path returns CSV and a missing actor receives 401. It says nothing about another authenticated tenant. |
| BoundaryExpectation: alpha requests beta’s inv-b | Test failed; exit 1; AssertionError: 200 != 403 | The vulnerable route allowed this constructed cross-tenant read. The deliberate failure exposes the missing property. |
| ReviewedFix: own, unauthenticated and cross-tenant cases | 3 tests passed; exit 0 | The added ownership check rejects the demonstrated cross-tenant case under the fixture’s policy. It is not a production security certification. |
The decisive test calls export_invoice_vulnerable(“alpha”, “inv-b”) and expects status 403. It failed as ‘AssertionError: 200 != 403’. In the reviewed route, the condition record[“tenant”] != actor_tenant returns (403, “forbidden”) before _csv_for runs. The equivalent wrong-tenant check then passed, alongside the own-tenant and unauthenticated checks.
The test is deliberately at function level, so it does not exercise a web framework’s session or middleware. A real PR needs a route-level or integration check if the authorization guarantee depends on those layers. That is why the review record separates the observed result from the unverified deployment path instead of calling the three passing tests a complete security proof.
This is the AI Hustle World Boundary-to-Evidence Review Record: a compact way to connect a PR’s claim, the call path, a required boundary, observed checks, remaining uncertainty and a reviewer action. It is an evidence record, not a new security standard or numerical score. The point is that a teammate can see what was reviewed and challenge a missing assumption without reconstructing the whole conversation that produced the code.
| Record field | Filled result for the invoice change | Reviewer action |
|---|---|---|
| Intended claim | Alpha may export inv-a; alpha must not receive beta’s inv-b. | Confirm the tenant policy with the product owner if the issue is ambiguous. |
| Entry point and unchanged dependency | The new export function calls lookup_invoice by ID; the helper does not know actor_tenant. | Read the helper and its other callers before assuming it authorizes. |
| Resource and destination | inv-b belongs to beta, amount_cents 28700; the new route serializes the returned record into a CSV response. | Treat unauthorized disclosure as a merge blocker in this scenario. |
| Evidence supplied | Happy-path and no-login checks pass, but neither passes an authenticated wrong-tenant ID. | Add a denial case at the boundary that actually serves the response. |
| Challenged behavior | Alpha requesting inv-b receives 200 where the fixture expects 403. | Request the ownership fix; do not approve the vulnerable version. |
| Evidence after change | The reviewed function rejects wrong-tenant access and three focused tests pass. | Inspect integration, real policy, logging and response fields before approving a real service. |
The corrected route checks record[“tenant”] against actor_tenant before calling the CSV serializer. An actual application might use a scoped database query rather than a post-fetch comparison, and its policy might intentionally return 404 instead of 403 to avoid revealing that an invoice exists.
Neither status code is universally correct for every deployment; the required property here is that the wrong tenant receives no invoice data. If authorization is enforced centrally, the review must verify that the new route actually participates in that enforcement.
The fixture contains no HTTP server, real session, database, deployment configuration or external AI tool. We did not run a static scanner on it.
Its strength is that the expected and observed results are explicit; its limit is that a three-test local model cannot tell us whether a production system has another path, race, policy exception or logging disclosure. The source file and captured console output are retained with the editorial research record; the code path and outputs above make the core finding reconstructable without implying a real-world benchmark.
What changes when the same rule reaches production?
In a database-backed service, authorization may be applied in the query itself: select the invoice where its ID matches the requested ID and its tenant matches the trusted caller context. That prevents the application from handling another tenant’s record in the first place.
A post-fetch comparison, as in our small fixture, can also deny the HTTP response, but it creates a period in which the wrong record exists in process memory. The reviewer should ask whether any logging, caching, metrics or exception handler touches that record before the check.
Some applications enforce row-level security in the database. If that is the control, the review needs to see how the connection receives the tenant context, how it is cleared between pooled requests, and what role background jobs use.
A local unit test of the route cannot demonstrate that a database policy is configured and active in the deployment. The practical evidence is a database-backed test using distinct tenant identities or an inspected policy and integration path.
Authorization may be centralized in middleware or a service layer instead. Inspect the exact route registration and order of handlers. A test of a similar endpoint is not enough if the new endpoint is mounted under a different router or bypasses the authorization decorator.
If the permission helper checks only a global role, ask where it verifies the relationship to this invoice. “Admin” can mean administrator of one tenant, not administrator of every tenant in the system.
The export may also become asynchronous. A caller starts a job while authorized, then loses access before the worker reads the invoice or a signed download link is used. The product must define whether permission is checked at request, execution, download, or more than one point.
The reviewer cannot infer that answer from a green synchronous route test. Identify the job’s actor context, the lifetime of the generated file, and whether its storage path and download token are scoped to the same policy.
The final response is only one disclosure channel. A cache keyed by invoice ID without tenant context can serve a previous result to the next requester; a trace may record the CSV body even when the route later returns 403.
Review the order of operations and the keys and sinks the change introduces. This is not a demand to inspect every cache in the company. It is a consequence of this particular path taking tenant-owned data from an unscoped lookup to an externally reachable output.
The local example suggests a disciplined production test matrix. Keep the own-tenant success and unauthenticated denial, then exercise an authenticated wrong tenant through the actual route.
Add a tenant admin without invoice permission if roles make that possible, and test the download stage separately if a job creates a file. Each added case corresponds to a distinct policy relationship. More tests with the same actor and same resource would add volume without probing a new boundary.
Review correctness before reaching for a vulnerability label
Many important bugs are failures of the requested business behavior rather than recognizable injection or access-control patterns. A report might include every invoice for the account when the request specified only unpaid invoices; a refund might calculate from the current price instead of the amount originally charged.
Both can compile and return plausible values. State the independent rule, choose a case that distinguishes the intended behavior from the likely wrong one, and trace where the implementation applies the rule.
Boundaries in ordinary data deserve the same attention. If a feature filters through the end of a day, ask whether the endpoint is inclusive and which timezone defines “day.” If pagination accepts an empty cursor, ask whether it repeats the first page or terminates.
If a calculation uses money, check representation, rounding point and whether the result can go negative. These questions should come from the feature’s contract; a reviewer does not need to enumerate every odd input in the language. A single counterexample that would expose a mistaken interpretation is often more useful than ten happy-path snapshots.
Concurrency matters when two requests can observe and modify the same state. An AI assistant may write a clean “read balance, subtract amount, save balance” sequence that passes every single-threaded test yet allows concurrent requests to overspend.
The reviewer needs to inspect the database transaction, conditional update or lock semantics used by this project; simply adding a second unit test with sequential calls would not exercise the race. The relevant proof may be a database-backed integration test, a design argument tied to the storage engine, or a specialist review. Do not claim a unit test proves isolation it never exercised.
Performance and reliability also have specific failure shapes. A loop that queries the database once for each of 500 exported rows can be acceptable in a toy fixture and expensive in a production report. An endpoint without a row limit can consume memory even if its result is correct for five records.
Use expected volumes and service constraints to choose a representative check. Query count, memory use and response time under that load provide better evidence than “looks efficient.”
The quality of the fix matters. A reviewer should be wary of a patch that suppresses a failing test, broadens a catch block, changes a status code without enforcing access, or deletes a validation branch so the suite turns green.
Ask whether the correction enforces the contract at a stable boundary and whether it leaves another route to the same data untouched. A new local condition can repair the shown case while preserving duplicated permission logic that will drift next month; that maintenance cost belongs in the review decision.
Look for failures between the happy path and the rollback
For a multi-step write, draw the point at which each side effect becomes durable. An order handler may reserve inventory, capture payment and enqueue a receipt. If payment succeeds but queuing fails, retrying the entire handler can capture payment twice unless the operation has a stable idempotency key or recorded state transition.
Ask which step can be replayed safely, what the caller sees after an ambiguous timeout, and where the system stores the answer. A generated retry loop without those answers may amplify the failure it was meant to handle.
A database transaction can protect changes within one database, but it cannot automatically roll back an external payment or email. The reviewer should see the boundary between transactional and external work and the compensation or reconciliation path for partial success.
That might mean an outbox, a durable job with deduplication, or an explicit manual recovery process. The specific pattern depends on the system; the requirement is that a failed second step does not silently leave an unaccounted first step.
Schema changes introduce a different timing problem. A migration can pass in a clean test database yet fail against old nulls, duplicate values or rows written by the previous application version. During a rolling deployment, old and new code may run together.
Review whether the new column is nullable or backfilled, whether reads tolerate both shapes, and when a constraint becomes safe to enforce. Ask what happens if deployment stops halfway. “Migration completed locally” describes one environment, not the production data transition.
Error handling can erase evidence. Catching every exception and returning an empty export may look resilient, but the caller cannot distinguish “no invoices” from “database unavailable.” Swallowing a permission exception can be worse if the handler falls through to a default response that includes data.
A reviewer should locate the error category, user-visible response and log event for each consequential failure. If sensitive data appears in the error log, the fix has moved the disclosure rather than removed it.
These cases need selective depth. A document-only change does not require a payment-state analysis; a one-step read does not need an outbox. When a PR introduces a durable write, external call, migration or retry, however, the failure boundary is part of the requested behavior. Require evidence that follows the actual operation rather than adding generic “edge case” tests whose setup never reaches the risky step.

Map security risks to the path that creates them
Begin with what the changed code accepts, reads, executes and emits. If untrusted input reaches a query or command, inspect how the program separates data from instructions. OWASP’s SQL Injection Prevention guidance explains parameterized queries as a primary defense for SQL construction.
A generated query containing a variable is not automatically vulnerable; inspect the actual query API, parameter binding and any dynamically selected identifiers. Conversely, a string that appears sanitized may still be unsafe when inserted into a different interpreter such as a shell, HTML, template or URL.
For response data, decide whether the caller should see each class of field and how it is encoded for the destination. A web view requires context-appropriate output handling; a CSV download has different risks; a webhook should send only the integration’s required fields.
A generated serializer that returns a whole database model can quietly expose internal notes, tokens or metadata even if the main displayed value is correct. Trace the final emitted shape, not just the query’s filter. In the invoice fixture, the amount is sensitive because the requester is the wrong tenant; no claim is made that integer cents or the two-column CSV independently cause injection.
Secrets and credentials require their own route through the change. Look for literal keys in source and tests, new environment variables with permissive fallbacks, tokens written to logs, or prompts that include restricted data. If a secret is committed, removing it from a subsequent diff may not remove it from history or invalidate it; rotate or revoke it under the organization’s incident process.
A scanner can help find known patterns, but a green scan does not establish that a proprietary credential or sensitive prompt was handled correctly. The reviewer should name the exposure and the concrete remediation rather than accepting “scanner passed.”
Dependencies carry execution and supply-chain risk. Confirm that an added package is real, needed, compatible with the project version and approved license policy. Inspect install scripts or transitive changes when the dependency is privileged or the lockfile churn is disproportionate to the feature.
If the code invokes a cloud SDK, verify the exact method and permission scope against the installed release. AI can suggest an API that sounds credible, but the artifact to trust is the package in your build and its primary documentation, not the wording of the generated explanation.
Configuration changes can undermine code-level checks. A new route may be protected locally while an alternate deployment path omits the same middleware. A permissive CORS setting does not itself replace server authorization, but it can broaden browser access in ways the team did not intend.
A debug flag, public bucket setting or additional workflow permission may create an exposure outside the function being reviewed. Follow the change into deployment manifests and CI permissions when those files move, and involve the owner of that environment when the consequences cannot be judged in the repository alone.
Some AI-assisted security risks sit earlier in the workflow. Repository text, retrieved documentation or issue comments can contain instructions aimed at the coding agent; generated changes may include an unnecessary dependency or leak sensitive context into a tool call. OWASP’s Secure Coding with AI guidance covers such AI-specific concerns.
This review still judges the resulting PR’s behavior and artifacts; the agent’s permissions and tool execution history are separate evidence when they affect the outcome. The human-in-the-loop decision guide is useful when a team must decide which consequential actions require an explicit human gate.
An incident-response program has a different job. It can detect suspicious activity after deployment and help investigate exposed systems; it cannot make an unreviewed permission boundary acceptable. Our guide to AI analysis of vulnerabilities and security alerts deals with that operational layer. A good PR review and good monitoring reinforce each other, but neither should be used to excuse a known defect at merge time.
Treat tests and scanners as evidence with a scope
A check result is useful only when you know what it exercised. Record the exact command, environment, commit and outcome for the checks relevant to the change. A build can expose syntax and type errors; focused tests can exercise a specified path; integration tests can show how middleware, storage and serialization work together; a browser test can show a user-facing flow.
Each has blind spots. If the repository’s CI ran only lint and unit tests, “CI green” does not mean a migration was tried or a production permission policy was verified.
Read the tests as claims about behavior, not as decorative coverage. For a new export route, ask which assertion fails if the implementation ignores tenant ownership, returns an extra confidential column, or exports an incorrect filter result.
A test that calls the route with only alpha and inv-a can validate the normal response while leaving every wrong-tenant request unexplored. A test that mocks the authorization helper to always approve may be useful for formatting logic, but it cannot establish the real permission decision. Do not reject mocking indiscriminately; decide which boundary the test is supposed to prove.
The relevant negative case is often smaller than a massive test expansion. In our fixture, the two old checks were insufficient until the reviewer added one actor-resource combination: alpha requesting beta’s inv-b. That case failed before the fix and passed after it.
For another feature the discriminating case might be a revoked role, a duplicate message, a malformed cursor or a downstream timeout. The principle is to choose a plausible wrong implementation and make sure the supplied evidence would expose it. The process of building and maintaining such tests is a separate subject; here the reviewer needs to know what the existing tests do and fail to do.
Static analysis can reveal paths and patterns humans miss, but its finding depends on rules, configuration, language support and data-flow modeling. CodeQL’s documentation on data-flow analysis distinguishes local and global flows and taint tracking.
That does not license a blanket claim that scanning misses authorization: a project can write custom rules, and a scanner may find a concrete access-control pattern. It means the reviewer should ask which query ran and whether it modeled the application’s tenant rule. We did not scan our controlled fixture, so its failure cannot be attributed to a scanner’s blind spot.
Pay attention to which alerts appear in the PR interface. GitHub’s documentation on triaging code-scanning alerts in pull requests describes alerts on lines in a diff. That is a useful signal, but a changed caller can alter the use of an unchanged helper, and an alert view is not a complete map of that behavior.
Investigate findings, false positives and coverage gaps separately. Likewise, a dependency report identifies known package concerns under its data and configuration; it does not decide whether the package is necessary or whether an invented API call works.
A second AI reviewer can generate hypotheses and search for omitted cases. Give it the requirement, relevant code paths and a specific question such as “show how this route establishes that the selected invoice belongs to this actor.” Ask for file and line references and a testable counterexample, then verify the answer yourself.
A model may confidently assert that middleware enforces a rule it has not actually inspected, or share the original generator’s assumption. Its comment is a lead, not a sign-off. Our AI cybersecurity versus traditional security tools analysis examines that broader division of labor; this article’s decision remains tied to evidence for the submitted change.
If a check could not run, write that plainly. “The integration suite requires a database service unavailable in this review environment; the targeted unit cases passed” is a meaningful statement. “All tests pass” would be false.
Decide whether the missing check is essential before merge or can be run in a controlled later stage, and assign an owner. A security-sensitive data export should not gain approval through ambiguity about a skipped environment.

Decide what to ask for, escalate or approve
A useful review comment identifies a claim, a path, an observed result and the change needed to close the gap. For the constructed invoice PR, the comment would say that export_invoice_vulnerable calls an ID-only lookup after checking merely that actor_tenant is present; alpha can request inv-b and receives 200 with beta’s amount.
Request a resource-level ownership check before serialization, a wrong-tenant case at the route boundary, and confirmation of the real service’s denial policy. That is an actionable change request. “Please check security” does not tell an author what failed or how to prove the fix.
Approval is appropriate when the reviewer understands the change, the intended behavior is settled, the consequential paths have relevant evidence, and remaining uncertainty fits the team’s risk tolerance. A request for changes is appropriate for a concrete defect or missing proof that the author can supply.
Escalate when a decision depends on product policy, legal obligations, system architecture or a threat model outside the reviewer’s authority. Pause a large or opaque change when nobody can explain its boundaries well enough to review it. These are practical dispositions, not a promise that an approved PR is risk-free.
The same severity label should not be applied to every omission. A missing format test on an internal admin-only CSV may justify a targeted addition; a cross-tenant disclosure of customer invoices blocks the illustrated change.
A payment path with an ambiguous idempotency guarantee may require a service-owner or payments specialist, even if the PR contains no obvious bug. State the possible harm and the uncertainty; that is more informative than marking every AI-generated patch “high risk.”
Review time is finite, so make risk-based triage explicit. If you approve customer-facing exports, payments or permission changes, use the full actor-resource-data trace and record the evidence before merging.
For an isolated presentation change, focus on behavior and accessibility without filling out a security table that has no relevant boundary. Give deeper attention to internet-facing entry points, privileged data, destructive operations, dependencies with execution privileges and changes that are hard to roll back; that is how the team keeps the review meaningful as AI-assisted change volume rises.
The team’s governance rules also matter. If the organization requires a named owner for access-control policy, record that owner in the PR; if it requires security approval for a new trust boundary, route the change there. Our AI governance framework covers those organizational controls. They should make accountability visible, not become a form that substitutes for reading the implementation.
Use W.O.R.T.H. to judge AI assistance and review load
AI Hustle World’s established W.O.R.T.H. lens means Workflow fit, Output capability, Risk and rights, Total cost, and Human effort. Here it is a decision aid for whether AI assistance improves a team’s coding and review workflow. It is not a new security score, and the five headings do not certify a PR. A tool can produce a complete-looking patch quickly while shifting more work onto the reviewer than it saved for the author.
Workflow fit asks whether generated changes work with the repository’s architecture, checks and ownership rules. If an assistant repeatedly invents parallel permission helpers or a new test framework, the review queue has to reconcile those choices on every PR.
Output capability asks whether the code and its evidence solve the actual behavior. Does the export respect the tenant boundary, and do the tests distinguish the wrong-tenant case? Fluent comments and a long change summary add little if that answer is unknown.
Risk and rights covers who may see the source, prompts, logs and customer data used during generation or review. It also asks who may approve a policy decision. A model can suggest a permission rule, but the owner must define the allowed actor-resource relationship.
If code or customer data is sent to an assistant, check the approved data-handling path before using the tool for generation or review. Our guide to AI data retention addresses the separate question of what happens to prompts and uploaded files; the PR reviewer still needs to inspect what this workflow actually sent.
Total cost includes model usage, CI, security scans, rework, incidents and maintenance of duplicated or weak code. A free suggestion can become expensive if it enlarges the review burden. Human effort counts time to specify, inspect, test, explain and later repair the change; measure the full path from request to accepted behavior.
W.O.R.T.H. becomes useful when a team chooses which tasks to delegate. AI help with a stable, well-specified parser may fit the workflow and remain cheap to validate. A cross-tenant billing route with uncertain ownership semantics may be better handled by a human defining the policy and test cases first, then using AI for bounded implementation and review suggestions.
This is a qualitative choice, not a universal rule against using AI on security-sensitive work. If your organization considers a specialized AI reviewer, evaluate it on the actual repository and failure cases. A product’s marketing claim cannot replace evidence for the submitted change.

Make the review repeatable without turning it into theater
Keep a short PR record that a future maintainer can read: the behavior claim, changed entry points, relevant unchanged dependencies, affected trust boundaries, checks actually run, unresolved checks, and disposition with owner. For a small change, this may be a few sentences in the description.
For a new public data path, a table like the Boundary-to-Evidence Review Record can prevent reviewers from losing the one assumption that matters. The format is secondary; the linkage between claim and observed evidence is the value.
Teams can improve the process by looking at outcomes rather than the number of boxes ticked. Track how often reviewers find a missing actor-resource case, how often a PR reopens because its evidence was incomplete, how many post-merge defects originated in misunderstood contracts, and the time spent reviewing versus repairing generated code.
These measures need consistent definitions and a baseline; an increase in comments alone could mean better detection or merely more noise. Use a sample of real changes to learn which checks catch consequential problems.
Keep feedback close to the work. If reviewers repeatedly find unscoped lookups in new routes, add a repository example or a policy test for that class of route. If generated code keeps using a nonexistent SDK method, provide versioned documentation and lockfile context before generation.
If PRs arrive too large to review, split behavior from mechanical cleanup and make the author explain both. Those changes improve the inputs and the review surface, reducing the amount of detective work needed after every agent run.
Some teams may decide that certain changes need two human reviewers or an AppSec review; others can meet the same risk with strong local controls and a service owner. What matters is the actual control path. A second approval from someone who never inspected the permission rule adds little. A focused reviewer who can trace the request, reproduce the failure and require the correct check contributes more than a ceremonial sign-off.
If nobody changes the process while code-generation volume grows, the likely failure is an expanding queue of plausible PRs with thinner attention per change. The answer is not to ban generated code or to demand a week-long audit for every line.
It is to put review effort where a wrong assumption crosses a meaningful boundary and make the evidence for that boundary easy to inspect. That is also why the old peer-review practice exists: it provides an independent challenge to a story the implementer already believes.
How We Researched and Tested This Guide
We reviewed primary guidance from GitHub and OWASP alongside official data-flow and PR scanning documentation. We also compared competing AI-code-review guides for their coverage of intent, security boundaries, tests and merge decisions. AI assistance supported research and drafting; factual technical claims were checked against the linked primary pages.
The Boundary-to-Evidence Review Record is our synthesis of those review questions and the controlled demonstration. Secure code review, authorization checks and negative testing are established practices. Our contribution is the explicit link from a particular code path to observed evidence and a reviewer decision.
For the first-hand example, we wrote a tiny Python 3.12.14 program with two fake tenant-owned invoices, an ID-only lookup, a vulnerable export route and a reviewed route. On 28 September 2026 we ran ExistingChecks, BoundaryExpectation and ReviewedFix as separate unittest commands.
The outputs were two passes, the deliberate 200-versus-403 failure, and three passes respectively; a direct call also recorded the unauthorized CSV response. The source and console output were saved as ai_code_review_boundary_demo.py and ai_code_review_boundary_demo_results.txt. No coding product, production application, scanner, live customer data or representative repository sample was tested.
The example is deliberately narrow. It does not establish defect prevalence, compare AI tools, or prove that the reviewed route would be safe in a deployed web service.
The article uses published security guidance for general principles and the local run only for the exact behavior it observed. If the service’s real access policy differs, the expected test and response must follow that policy; the wrong tenant must still not receive data it is forbidden to see.
Final Thoughts
The most revealing line in a review may be one that did not change. In our small case, an old lookup was adequate as an ID lookup; the new route made its unscoped result available to a different audience. Two green tests said the happy path and no-login path worked. One targeted wrong-tenant question exposed the missing permission decision.
Use that habit on real changes: state the behavior, trace the caller through the relevant code and data, challenge the highest-consequence assumption, and write down what the evidence proves. Let tools accelerate finding and checking, while keeping the merge decision attached to a person who can explain the path. A review that makes one hidden boundary visible is more valuable than a long checklist that never reaches the actual risk.
What Did the Coding Agent Do Before This PR?
Trace the agent’s tool calls, tests and human handoff so you can judge its submitted code against the right evidence.
See the Coding Agent Workflow →Frequently Asked Questions
Is AI-generated code inherently less secure than human-written code?
Its provenance does not decide whether a specific change is safe. Review the behavior, dependencies and security boundaries under the same engineering standard, while watching for plausible invented APIs, omitted constraints and tests that share the generator’s assumptions. Do not turn one constructed example or a vendor statistic into a universal defect rate.
What should I read before the changed lines?
Read the issue or requirement, acceptance criteria, PR description and file list. Identify the intended user, resource and failure behavior, then find a similar existing implementation and the helpers the new code calls. That context tells you whether the diff is solving the right problem and where the actual controls live.
Why is a passing test suite insufficient?
Passing tests establish only the cases that ran under their configured conditions. In our fixture, the normal export and no-login tests passed while a signed-in wrong-tenant request still returned the invoice. Ask which plausible wrong behavior would make a relevant test fail, and inspect skipped or unavailable checks separately.
How do I review code when an agent changed many files?
First map the requested behavior to entry points, data paths, tests, configuration and dependencies. Separate mechanical churn from behavior changes, ask the author to explain unexplained files, and split the PR if the consequential path cannot be reconstructed. A large diff is not automatically bad, but its scope must remain reviewable.
Should every AI-generated PR receive a full security audit?
No. Match depth to exposure and consequence: a new tenant-facing export warrants resource-level authorization and output review, while a contained presentation change has different risks. Escalate when policy or architecture is unclear; routine, low-risk changes still need ordinary functional and maintenance review.
Can an AI code reviewer approve another AI’s code?
It can suggest overlooked paths, but its answer is advisory. Require references to actual repository code and a testable counterexample, then verify them with deterministic checks and a responsible reviewer. Two models can repeat the same mistaken requirement or infer a control from a reassuring function name.
Does a green static-analysis scan prove the change is safe?
No. A scan reports what its configured rules and models detected on the code it analyzed. Investigate its findings, then separately check application-specific policy, data destination and business behavior. A custom rule may cover a tenant pattern; a default configuration should not be presumed to understand it.
What if the AI-generated test and implementation agree?
Compare both with an independent requirement or authorized policy decision. If the model misunderstood “own invoice” as “any invoice whose ID is known,” code and tests can agree while violating the product contract. Add a case with a different actor-resource relationship before accepting the matching pair as evidence.
What should a reviewer say when a test could not run?
Name the missing command or environment, the checks that did run, and the behavior that remains unverified. Decide whether that missing evidence is required before merge and assign a person to obtain it. An unavailable database integration test cannot be transformed into “all checks passed” by a confident summary.
When is a PR ready to merge after a security fix?
The owner should state the policy, show enforcement on the real path, supply a case that failed before and passed after the fix, and account for relevant integration gaps. Evidence depth depends on the consequence. Approval records a justified decision under stated limits, not a guarantee of zero future defects.
Written by
Muntasir Ahmad Chowdhury
Founder-AI Hustle World
Muntasir Ahmad Chowdhury is the Founder of AI Hustle World, an independent publication dedicated to making Artificial Intelligence practical, trustworthy, and easy to understand. He researches AI tools, automation, customer service, productivity, and real-world business applications, helping readers make smarter technology decisions through research-driven, experience-backed content.
Expertise:
AI Tools • AI Automation • AI Customer Service • AI Productivity • Generative AI • AI Workflows
Get Smarter With AI
Enjoyed this guide? Get practical AI tools, tutorials, and honest reviews delivered to your inbox.