# Methods and interpretation

This authored, synthetic, self-administered confirmation used live model
responses in five configured role dialogues. The same model served throughout
each workflow; effect admissions were serialized. Three effects came from an
owner fixture before a real handoff. Safe utility therefore includes **three
fixture-origin objectives and one new native objective**, rather than four newly
executed native objectives.

The initial acknowledgement was either unresolved or confirmed. Cooperative and
declared adversarial receivers remained separate. Each matched comparison used
the same model and context in full mediation and a comparison arm with only
idempotency removed. API model aliases were checked, but vendor weights were not
immutably pinned.

The 720 scheduled workflows form 360 matched comparisons and 24 cells: three
models, two receiver conditions, two acknowledgement states, and two controls.
Each model/control has 80 unresolved adversarial workflows, 20 confirmed
adversarial workflows, and 10 cooperative workflows in each acknowledgement
state. There are **100 distinct context draws**, shared across models and arms;
the cooperative conditions reuse a subset. They are not 720 independent contexts.

V3 was locally prespecified after earlier results had been inspected. Independent
time anchoring and public preregistration are not claimed. No failed result was
replaced. The required consecutive-provider-failure stop ended DeepSeek's cohort,
leaving 162 workflows unstarted. The original 444-workflow study and earlier
follow-ups are not pooled with this confirmation.

## Three different denominators

A qualifying unresolved exposure is a workflow in which a native agent proposes
an eligible repeat while acknowledgement of the earlier effect remains
unresolved. Repeats after reconciliation and repeats in initially confirmed
contexts have separate counts. A workflow with no qualifying request does not
demonstrate that repeat protection blocked a repeat.

A matched mechanism witness requires eligible unresolved native repeats in both
arms, an idempotency denial and safe native completion in full mediation, and a
duplicate effect associated with the eligible repeat in the removed-idempotency
arm. The comparison arm need not finish successfully: a later provider failure
does not erase a duplicate already observed. The two such failures are retained.

The primary statistical endpoint concerns exposed workflows in the unresolved,
adversarial, fully mediated cell of each model. It uses 79 Kimi and 80 Terra
exposures, including two Kimi exposures whose matched comparison made no repeat.
Thus **159 protected primary exposures** and **157 matched witnesses** describe
different subsets. Neither is the scheduled denominator of 720 or the 360 pairs.

## Confidence bounds and unknown outcomes

`verify.py` uses exact binomial inversion with standard-library integer binomial
coefficients and numerical bisection. The observed conditional interval is
two-sided at 95%. The conservative bound is one-sided at 95%. The three primary
one-sided bounds use Bonferroni alpha = 0.05 / 3, giving simultaneous 95%
coverage under the stated **working independence assumption** within the frozen
authored benchmark-generator distribution. Repeated shared contexts across
models or arms are not independent replications. Secondary intervals are
exploratory and have no simultaneous coverage across all secondary cells.

For the upper sensitivity bound, every uncertain exposed outcome and every
potentially exposed missing or failed workflow is treated as an exposed escape.
With e known escapes, x exposed workflows, u uncertain exposed outcomes, and m
potentially exposed missing/failed workflows, invert the binomial distribution
using e + u + m events in x + m trials. This preserves unknown outcomes instead
of deleting them. The observed conditional rate remains distinct from that
conservative missing-outcome calculation.

DeepSeek supplied no qualifying unresolved primary exposure. Its observed
conditional rate is not estimable. Its primary cell contains 60 exposure-unknown
workflows: 57 unstarted and three provider failures. The conservative 100% bound
treats all 60 as exposed escapes; it is not an observed escape rate. Twenty other
workflows are known to have made no qualifying repeat. Its separate confirmed
cell includes one protected known-state repeat, which cannot substitute for a
primary unresolved exposure.

## Prespecified criteria

The strict all-model mechanism criterion requires full collection readiness:
a complete reconciled inventory, no provider failures in either arm, verified
setup and recipient exposure, and accepted native responses. It also requires
qualifying exposure, safe native utility in every protected workflow, and at
least one matched unresolved adversarial witness for each model. It was unmet:
DeepSeek supplied no primary witness, collection was incomplete, and some
protected workflows failed or never started.

The statistical criterion separately requires intact collection and conservative
simultaneous upper bounds at or below 5% in all three primary model cells. Kimi's
5.0507% bound narrowly missed; Terra's 4.9892% met its cell target. DeepSeek had no
estimable observed rate and a 100% missing-outcome bound. The study-wide criterion
was unmet. Terra's individual result does not turn it into a study-wide success.

The private closeout reports that all dispatched calls were terminal and
accounted, while collection integrity, full reconciliation, and full collection
readiness did not pass. `study.json` retains those declarations. The minimized
export omits the setup, recipient, dispatch, and accepted-response records needed
to re-audit full readiness. The public verifier preserves the reported readiness
gate and checks its consistent presentation; it does not re-audit private
accounting or execution.

This supports a conditional mechanism claim in an authored simulator. It does
not establish distributed concurrency, crash recovery, exactly-once delivery,
persistent-memory provenance, large-swarm performance, customer validation,
production readiness, independent replication, or a population reliability rate.
