Verification field guide
How to verify AI-generated code changes beyond code review
Static AI code review cannot prove software works. Learn how to bind task intent, run executable checks, and capture evidence for AI code changes.
Moving Beyond Static Opinions in AI Code Verification
Large language models and autonomous coding agents can generate hundreds of lines of code in seconds. However, engineering velocity breaks when developers are asked to approve pull requests they did not author and cannot easily reason about. The default response across many software organizations has been to insert another AI model as a static code reviewer.
Static AI code review is valuable for discovering obvious risk patterns, flagging style inconsistencies, or explaining complex diffs. But static review cannot prove that software actually works. An LLM reviewing a diff reads source code as text; it does not execute the program, initialize a database, manage browser state, or process HTTP headers.
A pull request generated by an agent can look clean, include docstrings, pass linter checks, and still fail silently at runtime. Relying solely on static code review to verify autonomous code changes introduces a dangerous gap between plausible-looking source text and verified operational reality. To establish confidence in AI-generated code, software teams must move beyond code review to execution-backed verification.
The Structural Limits of Code Review for Agent Outputs
To understand why code review alone is insufficient for AI-generated code, consider how autonomous agents construct changes. When an agent attempts a task—such as updating an authentication flow or modifying a database query—it operates against statistical likelihoods and context windows.
This creates specific failure modes that static code review routinely misses:
- Plausible Lookalike Code: Agents frequently generate code patterns that mirror surrounding repository style but use invalid assumptions regarding library APIs, internal helper contracts, or async execution timing.
- Unseen Edge-Case Behavior: A diff modifying an authorization check may look mathematically sound in isolation while failing under specific session states, cookie configurations, or concurrent requests.
- Decoy and Scope Drift: Agents tasked with fixing a bug can inadvertently edit adjacent, working functions or rewrite working tests to fit a broken implementation.
- Environment-Dependent Mutations: Changes involving DOM interactions, local storage, network retries, or database transactions cannot be verified by inspecting syntax trees.
When a human or an LLM reviews these diffs statically, they evaluate plausibility. But software correctness is not a matter of plausibility; it is a matter of executable behavior under defined constraints.
The 5-Part Executable Verification Loop
Execution-backed verification bridges the gap between agent intent and operational truth. Rather than asking a model whether a diff looks correct, executable verification subjects the change to deterministic checks and captures reproducible evidence.
The following is a recommended workflow, not a claim that every CodeVetter receipt path implements every phase. In particular, experimental project-receipt ingestion consumes existing runner evidence; it does not execute checks, discover commands, or persist results into the desktop database.
A verification loop consists of five distinct phases:
[Task Intent] ──> [Agent Change] ──> [Executable Checks] ──> [Captured Evidence] ──> [Measurable Verdict]
1. Freeze Task Intent and Scope
Verification begins by binding the exact requested behavior and its acceptance boundaries. If the original task description is ambiguous, the verifier records that ambiguity explicitly rather than allowing the agent or evaluator to invent a friendly interpretation.
2. Bind the Repository Revision and Change
The exact base Git commit SHA, the workspace state, and the agent's patch must be pinned together. Verification must run against a clean, pinned revision so that external workspace drift cannot alter the outcome.
3. Execute Bounded Behavioral Checks
The verifier runs authoritative, deterministic checks against the modified codebase. These checks prioritize repository-owned unit and integration tests, supplemented by focused browser journeys or API calls when existing test suites do not exercise the changed behavior.
4. Capture Raw Execution Evidence
Record the evidence the runner actually captures: command identity, exit codes, bounded and redacted output, resource measurements, network observations, and failure signatures where available. Mark missing measurements and collection limits explicitly. A command-level memory measurement is not automatically a process-tree measurement, and missing network telemetry does not mean zero egress.
5. Render a Measurable Verdict
Use the chosen verifier's documented status vocabulary. CodeVetter's experimental project-receipt analyzer reports passed, failed, or no_confidence separately for correctness, performance, safety, inventory, and overall status. Missing evidence remains explicit; a passing dimension does not establish that unchecked requirements are safe. Unsupported receipt formats are rejected before analysis.
Establishing Behavioral Boundaries: Browser State and APIs
Code review evaluates syntax; verification evaluates behavioral boundaries. When verifying agent changes, execution must target the exact interface boundaries exposed to users or system callers.
Browser Journeys and User Interface State
For web applications, subtle frontend regressions—such as lost input focus, incorrect state hydration, broken event handlers, or race conditions during async fetches—frequently pass static diff inspection. Executable verification uses headless browser automation (such as Playwright) to run end-to-end flows against local development servers, capturing DOM states, console errors, and visual evidence.
API Contracts and Authorization
When an agent updates API endpoints or middleware, verification requires executing actual HTTP requests against isolated local endpoints. Checks must explicitly test both valid payload paths and boundary conditions, such as unauthorized access attempts and malformed inputs.
Failure Taxonomy: Separating Regressions from Environment Noise
A common flaw in automated verification systems is treating every failed command as an agent bug. If a test fails because a local database port was busy, or because an external API timed out, blaming the coding agent distorts quality metrics.
A useful verification workflow records failure classifications when the evidence supports them. The labels below are design guidance, not proof of causal attribution:
- Agent Defect / Failure: The executed check failed explicitly due to broken logic or unhandled exceptions introduced by the agent's change.
- Possible Regression: A previously passing test failed after the change. Check comparable baseline conditions, repeatability, and environment differences before attributing the failure to the patch.
- Pre-Existing Failure: The test was already failing on the base commit prior to the agent's run.
- Operational / Environment Failure: Execution was aborted due to infrastructure conditions, such as missing binaries, port conflicts, or memory limits.
- Unverified Requirement: No authoritative test or runner existed to exercise the requested acceptance criteria.
Keep uncertain attribution explicit. Before/after observations alone cannot rule out environment changes or flaky tests.
Fix Closure and Re-Check Linkage
When verification identifies a failure in an agent's patch, the failure evidence becomes structured feedback for a corrective agent iteration.
To document a bounded fix and re-check trail:
- Preserve the Initial Failure: The failing run, command, exit code, and failure signature are retained in a persisted evidence record.
- Apply the Corrective Patch: The agent applies a targeted fix in an isolated, clean workspace.
- Re-Run the Failing Check: The verifier re-executes the exact test that previously failed to confirm resolution.
- Link Before-and-After Evidence: Link the original failure to the passing re-check, including revision and environment identities. This establishes the recorded check's outcome, not the absence of every regression. Preserve broader checks and unresolved limitations separately.
Next Action for Technical Teams
Do not rely on second-model LLM review alone to approve agent pull requests. Evaluate your workflow against the five-part loop: freeze task intent, isolate the revision, run behavioral checks, capture evidence, and inspect failure-to-pass re-checks alongside remaining coverage gaps.
To examine how local verification functions in practice, download CodeVetter or inspect our public benchmark methodology.