While reviewing a website deployment flow, one piece of shared tooling raised a suspicion: when data is missing, the safeguard gets skipped. Tracing upward to the real entry point showed that the caller had already blocked that case. The candidate was overturned.
This time I asked Codex to pin one version of the code first and read only the source. If only the final summary survives and the counter-evidence is left out, the next review may chase the same lead all over again.
A review should keep the grounds for a conclusion and the reasons that overturned it. A question nobody has answered yet needs a name, a source and a next step.
After the Second Opinion, There Is More Work
In the article about having codex check Claude’s work, I wrote about how to protect a second opinion’s independence: keep the question neutral, keep the raw report, let the reviewer only point out problems, and leave the decision to whoever holds the full context.
That experience came with a limit. A machine can block a review output that is completely empty, but whether a report with words in it actually completed a substantive review still takes a person to judge.
As the work piles up, the same problem extends to handoff. A few days later, in a different window, does the person picking it up know which conclusions have been checked? After the code changes, is the old evidence still valid? When an opinion was overturned, did anyone record why, or did it just vanish from the final summary?
I want a second opinion to be usable by the next round of work. That takes a clearer record than a chat summary.
How Cloudflare Keeps Unresolved Questions
Cloudflare’s public security-audit skill offers a reference. It has different agents find and re-check candidates, requires the verifier to look for counter-evidence, and sorts results into confirmed, needs validation and rejected. The scope of the review is kept in a separate ledger, so a results report cannot stand in for overall coverage.
This article draws on Cloudflare’s official post Build your own vulnerability harness, its public security-audit skill, and its validation rules. I apply their way of separating evidence, counter-evidence, and open facts to my own work records.
What I want to borrow is how it handles “needs validation.” That status has to name the exact fact still unresolved. It carries no vulnerability severity, and it cannot be written up as confirmed just because the lead looks persuasive.
For example, some judgments need to know the actual deployment settings, and some need a minimal test run in a restricted environment. Facts you cannot see by reading code alone should stay visible in the record.
The same split works in my daily work: leads, facts still to be checked, and judgments that were overturned, recorded separately. If a meeting touched on a direction for collaboration, I still go back to the original message to confirm who has the authority to commit. If a document contains a number, I still check the version and the date it applies to. AI can sort out the leads first, then spell out where a person is needed.
How Far a Machine Check Can Prove
Cloudflare’s validation rules are explicit about the limit: when the validator succeeds, what it proves is format and ledger consistency. It cannot replace a judgment about the facts of a vulnerability.
Having a fresh agent re-check can reduce the bias of following the original author’s line of thought, but different agents can still share blind spots. Running more rounds cannot be converted directly into “we found this percentage of all the vulnerabilities.” In the article describing how it approaches its harness, Cloudflare points out that without a fully labeled set covering every real vulnerability, you cannot claim a recall rate.
The first round pinned the version and the source fingerprints, and another agent returned to the same source to look for counter-evidence. A checker I built reconciled the links between records, the line ranges and the source identities. This round did not run the target program.
Being able to re-check a source and having observed the program’s behavior are evidence at different levels. A report has to let readers tell them apart.
From One Deployment Lead to Verifying a Patch
Only the second round moved into a controlled local reproduction. We kept the real deployment entry point, the safeguard and the basic functional check (smoke) scripts, let Git use real version relationships, and swapped the services, web responses and build for offline stand-ins. The sandbox, enforced by the operating system, first had to pass tests that refuse network access and out-of-bounds writes before any case was run.
This deployment flow updates two parts at once: the website pages (Pages) and the backend service (Worker). The old combined deployment entry point took only the last successful source from a mixed ledger. If the entries at the end of the ledger were an older Pages source, it could overlook the newer source the Worker already used.
In that isolated offline environment, with stand-ins for the services and the build, we reproduced a case where the combined deployment swapped the simulated Worker back to an older source while the whole entry point still recorded success. Deploying the same older source to the Worker alone was blocked. The code was updated during the review; rerunning with the updated source pinned gave the same result.
Here “the end of the ledger” means position in the file. It cannot stand in for the real order in which deployments happened.
This result supports fixing the version-comparison logic. It does not support “production has had an incident.” This round verified only that one version-rollback path. It did not reach a security-vulnerability determination and assigned no severity. The fake services did not run the real Worker, so whether actual functionality was affected, and whether past ledgers ever contained this arrangement, remain two separate questions.
In the third round, after I agreed, Codex patched the deployment entry point in an isolated copy. The combined deployment now checks the last successful source for Pages and for the Worker separately, and the shared safeguard check was left as it was. The new test first caught the problem on the old code, then confirmed the patched version blocks it. Of six full entry-point cases, the four rollback scenarios were stopped before any service deployment call, and the two normal updates completed their checks and recorded success.
An independent reviewing agent accepted this narrowly scoped patch. It also pointed out that the error message did not switch with the deployment target, so we added the fix and a comparison test. What this round delivered is a patch ready for review and local verification. When this isolated verification ended on 2026-10-03, the patch was in the isolated copy and had not been merged or deployed to production; this piece keeps the verification boundary as it stood then and does not track later deployment status. “Review passed,” “patch verification passed” and “production updated” need to be recorded separately.
Keeping Evidence Apart from Authorization to Act
For the way I work, using AI across several projects, I organized what I want the next handoff to be able to trace back into six fields:
| Field | The question it answers |
|---|---|
| Claim | What exactly is being judged this time? |
| Source and version | Based on which code, document or original message? |
| Evidence status and reason | Checked, still open, or overturned? What is the basis or the counter-evidence? |
| Unresolved facts | Which specific answer is still missing? |
| Owner and next step | Who can supply the evidence, and how? |
| Authorization to act | Who agreed to which operations? Where is the original record of that authorization? |
Filled in with this deployment lead, it comes out roughly like this. This is the record as of the end of the isolated verification on 2026-10-03, not the current status:
| Field | This deployment case (record as of 2026-10-03) |
|---|---|
| Claim | The combined deployment may miss a newer source the Worker already uses, rolling the backend back to an older version. |
| Source and version | The deployment entry point at a pinned version, the matching offline test record, and the separately kept isolated patch. |
| Evidence status and reason | The offline path was reproduced. The combined deployment rolled back and recorded success, while the same condition deployed alone was blocked; after the patch, six entry-point cases gave the expected results. |
| Unresolved facts | Whether production ever had the same ledger arrangement and source conditions, and whether real Worker functionality was affected, are both unverified. |
| Owner and next step | The next step at that time: I decide the next scope and who takes it, then check the main-branch version and assess whether the patch is merged and deployed. |
| Authorization to act | At that time I had agreed to the isolated patch and local verification, and the original conversation can be consulted for that authorization; merging the patch and deploying it were to be decided separately. |
The last field in particular has to stand apart. Verifying a fact does not also grant permission to modify, send or publish. Several models agreeing does not make a commitment on behalf of a collaborator.
Deployment review works the same way. After a problem is found, the fix plan, the code change and the production deployment each have their own acceptance and their own responsibility. If all of them get folded into “done,” the next person picking it up has to guess again.
This record also suits meeting and email triage. Attach the original message to an important deadline, leave a confirmer on a collaboration commitment, and mark a partial read failure honestly. A failed data fetch and a search that found no matching messages cannot share the same line saying “nothing to do.”
Start Within What You Can Maintain
I chose the website deployment flow for the pilot. I did not send every daily summary through multiple rounds of review. In the first round I capped the planned number of local sub-agent calls at three; that limits the number of calls, not tokens or billing. Recording the scope of the source reading and of the later local runs separately is what lets me look back and judge what adding one more agent actually solved.
Overturned candidates need to be kept, so the next review can read the counter-evidence first and then decide whether a re-check is worth it. Leads not yet verified need to be kept too, to show which step the work still owes. These records give a basis for planning tests; whether they actually save people time still needs a later comparison.
In the next round I will look at whether this set of records makes review quicker and handoff less repetitive. The first pilot has no comparable baseline, so I cannot write down a percentage saved yet.
The judgment I want to keep in place is a concrete one: the next time I see “done,” I can follow the evidence back, and I also know which step has not been completed.
Primary sources
- Build your own vulnerability harness, by Dan Jones, Alexandra Godoi, and Grant Bourzikas, 2026-06-18: describes the design and limits of moving from a skill to a cross-repo internal harness.
- cloudflare/security-audit-skill: supports the six-stage single-repo starting point, independent falsification, the confirmed / needs_validation / rejected statuses, and the coverage ledger.
- VALIDATION-AND-REPORTING.md: supports the rule that consistent format and records do not establish that a vulnerability is real.
The deployment case in this article comes from the author’s records of a pinned-source review, offline reproduction, and isolated patch, made on 2026-10-03. It supports only the limited local results stated in the text.


💬 Comments
Loading...