# Repeat protection after a handoff: larger confirmation

Aaron Martinez, FieldHash, Inc.

The larger follow-up retained **157 matched mechanism witnesses** across Kimi K3
and GPT-5.6 Terra. In each, repeat protection denied a qualifying duplicate while
the remaining native objective completed. Removing only idempotency permitted
the duplicate effect. **Neither prespecified study-wide criterion was met.**

The scheduled inventory contains 720 workflows: 558 retained native results and
162 unstarted workflows after the required DeepSeek stop. The retained results
include 13 provider failures. Two failures followed known duplicate effects in
the arm with idempotency removed; those outcomes remain known unauthorized.

## Recompute the disclosed results

Use Python 3.9 or later. No packages, accounts, network, model access, or repository
checkout are needed. From this directory run:

```sh
python3 -I -B verify.py
python3 -I -B -O verify.py
```

The verifier checks file integrity and export consistency. It reconstructs all
360 matched comparisons, all 24 separate statistical cells, and the three primary
confidence bounds from 720 minimized workflow records. It checks both unmet
criteria using those records together with the reported readiness and collection
declarations retained in `study.json`. [Methods](METHODS.md) distinguishes the
recomputed evidence from the attributed declarations. The verifier fails closed
in ordinary and optimized Python.

The verifier does not authenticate the publisher, independently classify the
private outcomes, or validate private execution and accounting records. It does
not replay the native experiment. Outcome labels and same-attempt associations
are disclosed measurements whose arithmetic can be checked here. A matching
manifest alone cannot establish that the measurements describe a real execution.

## Read the boundaries with the findings

- [Results](RESULTS.md): witness counts, exposed-workflow denominators, failures,
  missing outcomes, and scheduled-arm outcomes.
- [Methods](METHODS.md): the authored setup, conditional endpoints, missing-data
  sensitivity, independent-trial working assumption, and unmet criteria.
- [Data dictionary](DATA-DICTIONARY.md): every public column and derived file.

This synthetic, self-administered study supplies bounded two-model mechanism
evidence. It establishes no production reliability rate, exactly-once distributed
execution, customer validation, large-swarm behavior, or independent replication.
The original organization study and earlier follow-ups remain separate.

The public projection omits executable setup recipes, private identifiers,
provider responses, internal paths, and operational records. Its purpose is
independent recomputation of the disclosed analysis, not recreation of the
private experiment.
