# Why reproducible evidence matters when an AI agent says a task is done

> Canonical page: https://codevetter.com/articles/why-reproducible-evidence-matters-when-an-ai-agent-says-a-task-is-done

The adoption of AI coding agents has shifted the bottleneck of software engineering. We are no longer constrained by the speed at which we can draft new features. Instead, the primary constraint has become verification. When an autonomous coding agent completes a task and reports "done," how do we actually know the change is correct?

In many workflows, the answer is manual inspection. A developer must read the generated diff, pull the branch, run the test suite, and click through the application to ensure nothing broke. This process is time-consuming and negates the velocity gains the agent was supposed to provide.

To bridge this trust gap, the industry must move beyond trusting an agent's self-reported completion and beyond asking a second model for its opinion. We need reproducible execution evidence.

## The Limits of LLM-as-Judge and Static Review

As the volume of agent-generated code increases, a common response has been to automate the review process using the same technology. This approach involves feeding the diff to another language model and asking it to identify defects. Hosted pull request bots analyze the structure of the change and flag potential issues.

While static AI review can catch syntax errors or stylistic inconsistencies, it suffers from a fundamental limitation: it is not executable proof. An LLM reading a diff is providing an opinion based on pattern matching. It does not know if the code will compile. It does not know if a new API endpoint correctly parses the payload.

Furthermore, static review tools lack the broader execution context of the repository. They view the diff in isolation, perhaps with some surrounding file context, but they do not run the application. When an agent hallucinates a variable name, a static reviewer might miss the error if the hallucination looks plausible.

Executable outcomes outrank model opinion. When the goal is to determine whether a software task was completed correctly, a deterministic test failure is infinitely more valuable than a probabilistic model's approval.

## Defining Reproducible Execution Evidence

If static review is insufficient, what is the alternative? The answer lies in reproducible execution evidence.

Execution evidence is generated by running the code in a deterministic environment and observing its behavior. It moves the verification process from the realm of "looks correct" to "provably works." Reproducible evidence means that any developer, or any other agent, can take the same code, run the same checks, and achieve the exact same results without relying on a hidden cloud state.

In practical terms, reproducible execution evidence includes:

- **Build Logs:** Does the code compile without errors?
- **Test Results:** Do the unit and integration tests pass? Did the agent's change break an unrelated test?
- **Runtime Behavior:** For web applications, does the server start? Can a script interact with the newly created API?

By capturing this evidence deterministically, we create a durable, machine-readable record of the software's state. This record serves as the source of truth for whether the agent succeeded.

## The Verification Loop

To effectively utilize execution evidence, it must be integrated into a tight loop. At CodeVetter, we structure this around a specific process: **task → agent change → executable verification → evidence → measurable verdict**.

1. **Task:** The workflow begins with a clear intent, such as fixing a bug.
2. **Agent Change:** The coding agent executes the task, modifying the local repository.
3. **Executable Verification:** Instead of immediately asking a human to review the diff, the system runs deterministic checks. This might involve running a Playwright script.
4. **Evidence:** The verification step produces concrete artifacts—JSON test reports, build logs, and traces.
5. **Measurable Verdict:** The evidence is synthesized into a clear outcome. Did the tests pass? Was the bug fixed?

If the verdict is a failure, the evidence is fed back to the agent as a concrete diagnostic. The agent is given the exact compiler error or the precise test assertion that failed. This deterministic feedback loop allows the agent to iteratively correct its work.

## A Concrete Example: The API Boundary

Consider a web development task: a developer asks an agent to update a backend to support a new query parameter on a search endpoint.

The agent modifies the controller, adds the logic, and reports the task as complete. A static AI reviewer might look at the diff, confirm that the syntax is valid, and approve the pull request.

However, an execution-backed verification system takes a different approach. It runs a local test script that starts the server and sends an HTTP request with the new parameter.

Imagine the agent made a subtle error: it parsed the parameter as a string, but the database function expects an integer.

The static reviewer missed this because it didn't trace the type through the application hierarchy. But the executable verification catches it instantly. The request returns a `500 Internal Server Error`, and the test framework captures the stack trace indicating a type mismatch.

The system captures this execution evidence and provides a measurable verdict: Failure. The developer (or the agent itself) can now see exactly what went wrong, supported by a reproducible runtime trace. There is no ambiguity.

*

## Local, Bounded, and Reversible Workflows

A critical aspect of reproducible evidence is where and how it is generated. For verification to be fast, secure, and iterative, it should be local to the repository.

Many AI coding tools rely on hosted infrastructure—pushing code to a cloud service where a PR bot analyzes the diff. This introduces latency, requires granting third-party access to proprietary source code, and creates a disconnect between the developer's local environment and the verification environment.

A local, bounded workflow means the execution evidence is gathered exactly where the agent made the changes: on the developer's machine or within an isolated, deterministic local sandbox.

- **Privacy and Security:** Code never leaves the local machine for verification. There is no need to upload proprietary logic to a third-party server just to run a test suite.
- **Speed:** Local execution eliminates network latency. A developer can ask an agent for a change, and the verification loop can run in seconds, right in the same workspace.
- **Reversibility:** By keeping the workflow local and bound to the repository's version control system, changes are easily reversible. If the executable verification shows a catastrophic failure, the local Git worktree can be reset instantly.

At CodeVetter, our architecture reflects this principle. The verification engine is a native desktop binary and Rust core that operates over local SQLite. It does not require a hosted server to evaluate whether a local application functions correctly.

*

## Moving from Trust to Verification

The goal of AI coding agents is not to write code that looks good; it is to write code that works. As long as we rely on manual inspection or non-executable model opinions to bridge the trust gap, we will be bottlenecked by the limits of static analysis.

By demanding reproducible execution evidence, engineering teams can change their relationship with coding agents. Instead of acting as full-time code reviewers, developers can become orchestrators—setting the tasks, defining the boundaries, and letting deterministic verification handle the proof.

If you are incorporating autonomous agents into your workflow, the most important question to ask is not "What model is the agent using?" but rather, "How will I execute and verify the result?"

### Practical Next Action

To start shifting toward execution-backed verification, pick one critical component of your application—such as an authentication flow or a core data transformation function. Write a deterministic, local test script (using a framework like Vitest or Playwright) that exercises this component from end to end.

The next time you use a coding agent to modify that component, do not review the diff first. Run the script. Let the execution evidence tell you whether the agent succeeded.

## Public product links

- [CodeVetter](https://codevetter.com/)
- [Download](https://codevetter.com/download)
- [Documentation](https://codevetter.com/docs/)
- [Source](https://github.com/Codevetter/codevetter)
