← CodeVetter

Verification field guide

CodeVetter vs open-source agent-verification CLIs

Compare CodeVetter with open-source agent-verification CLIs like Verdict, PatchDrill, ProofGate, evigate, didrun, and verity — quoted from their own READMEs.

An open-source wave of agent-verification CLIs now sits between code review and CI: Verdict, PatchDrill, ProofGate, evigate, didrun, and verity each check one narrow, deterministic claim about agent-written code. CodeVetter is a local verification workbench that reviews an agent change, exercises the changed behavior, and packages the evidence into a portable bundle. This page compares the shapes, quoting each project's own README, checked 2026-09-18.

The short comparison

QuestionCodeVetterThe verification CLIs
Primary jobVerify whether an agent task worked and retain evidenceOne deterministic gate per tool: intent-vs-diff, proof requirements, or claims-vs-events
Main surfaceLocal desktop viewer plus CLI/MCP and portable evidence bundlesCLI or GitHub Action, usually wired into CI or a hook
Verdict basisExecuted checks against changed behavior, findings, and uncertaintyStructural or recorded facts — never a model's judgement of the diff
PlatformApple-silicon macOSPlatform-independent
Where code goesLocal machine; a chosen model provider only for review findingsLocal or CI; most make no model call at all
Current CodeVetter availability, checked August 2026Apple-silicon macOS build through GitHub Releases; no CodeVetter subscription is currently declared, and model-provider costs are separateOpen source

What each CLI actually checks

From each repository's own README, checked 2026-09-18:

  • Verdict (github.com/verdict-ci/verdict) — "the merge-time fidelity gate for AI-written code." It reads a pull request's declared intent — the issue, PR description, or linked spec — and compares it against the actual diff, catching intent-mismatch, silent-scope, unbacked-claim, doc-drift, and missing-impl. It runs as a GitHub Action on every PR and as a CLI.
  • PatchDrill (github.com/seungdori/patchdrill) — "the deterministic proof layer between code review and CI." It answers "what proof should exist for THIS diff before merge — and what's missing?" with no model call and no network. Its detectors cover leaked secrets, prompt injection planted in files an agent will read, workflow escalation, missing proof, and dependency drift across roughly 25 ecosystems.
  • ProofGate (github.com/aevryone/proofgate) — a "trust-but-verify merge gate" that blocks only on structural or executed facts: committed secrets, empty or always-true tests, security functions stubbed to a constant, imports that resolve nowhere. Claims inferred from PR prose are warnings, never blockers.
  • evigate (github.com/shiki-yusuke/evigate) — a local CLI that parses a session transcript into tool-observed events and agent-declared claims, then returns proven, contradicted, or unknown verdicts. A claim is proven only when the recorded evidence supports it; ambiguous evidence is reported as unknown rather than guessed.
  • didrun (github.com/nelsonwerd/didrun) — "flight records for agent-written code." It records what an agent actually executed — the command, the exit code, the output, and the exact git tree state — binds claims to a commit, and exports a self-contained HTML report. Its README is precise about scope: "didrun is a recording, not a proof." There is no LLM in the trust path.
  • verity (github.com/pietro-falco/verity) — checks declared claims in a manifest against the filesystem, git HEAD, and command exit codes, stamping each true or false deterministically and offline. It explicitly "checks truth, not quality."

Where they overlap with CodeVetter

All of these projects, CodeVetter included, reject the same premise: that a second model reading the diff counts as verification. They all favor deterministic, inspectable evidence and local or CI execution over a hosted reviewer.

Where they differ

The CLIs are deliberately narrow. Each answers one question well — did the diff do what it claimed (Verdict), does the required proof exist (PatchDrill), are there structural merge blockers (ProofGate), did the session's claims match its events (evigate), what did the agent actually run (didrun), is this declared claim true (verity). None of them tries to answer the broader task question.

CodeVetter aims at that broader question: did the agent complete the requested software task correctly? Its verification loop binds the task and exact change to executable checks — repository tests plus focused browser or API checks — and separates pass, fail, and unverified rather than emitting a single gate verdict. The output is an evidence bundle (JSON, Markdown, or offline HTML) that a reviewer can inspect or replay locally.

When a verification CLI is the better choice

Plainly: often, today.

  • You want a merge gate in CI. These tools were designed to block merges; CodeVetter is a local workbench, not a branch-protection check.
  • You need one narrow guarantee. If the question is exactly "did the agent's report match what it executed," evigate or didrun answers it with less machinery.
  • You are not on a Mac. CodeVetter publishes an Apple-silicon macOS build only; every tool on this page runs anywhere.
  • You want zero dependencies and auditable code. Several of these CLIs are small enough to read in an afternoon.

When CodeVetter is the better choice

  • The question is task completion, not claim fidelity. A diff can faithfully implement its PR description and still break the requested behavior; a claims-checker will not see that.
  • You need evidence for a human reviewer. CodeVetter emits a portable bundle with checks, artifacts, and explicit uncertainty rather than a single exit code.
  • You want the whole review loop local. Findings, re-checks, and history stay on your machine; your model key is used only for review findings.

How this comparison was built

Sources. Each tool's public README at github.com/verdict-ci/verdict, github.com/seungdori/patchdrill, github.com/aevryone/proofgate, github.com/shiki-yusuke/evigate, github.com/nelsonwerd/didrun, and github.com/pietro-falco/verity — all checked 2026-09-18. CodeVetter claims come from this repository and /privacy.

What we did not test. We did not run any of these tools against a shared case set. No catch rate, latency, or cost figure appears on this page because none was measured. These projects are young and change quickly; re-check their READMEs before relying on a specific capability.

What a real head-to-head needs. The same frozen agent tasks, disclosed configuration, repeated runs, and one scorer applied to every tool — the same standard CodeVetter applies to itself at /benchmark.

Where to go next