Live diagnostic · 1,000 answers

Every answer should resolve to support, contradiction, or review.

We ran live models over 300 public RFC conflicts and built a claim packet for every answer. Faithful answers attributed cleanly. Off-script answers were flagged. The interesting result is that nothing went unaccounted for.

Supported claim

For every live answer in this Governed Memory diagnostic, the FieldHash Ledger attribution packet assigned a disposition: supported by a governing record, contradicted by one, or unattributable to governed context. No answer passed through unaccounted for.

Off-script answers flagged

5

Models went off-script 5 times in 400 governed answers. Each was flagged unattributable. This is reported as an incident count, not a generalized catch-rate estimate.

Stale assertions caught

4

Stale assertions were rare: 4 across 1,000 live answers, each flagged by the packet. We do not generalize that rate; see the corpus disclosure below.

Empty packets

0/1,000

Three model families (Gemini 3.1 Flash Lite n=600, Claude Haiku 4.5 n=200, GPT-5.5 n=200), governed and plain top-k context conditions, every answer dispositioned. Temperature 0, pinned model ids.

Why stale assertions were rare: RFC records state their own supersession in prose, which protects the model. Most enterprise records do not. The conversion rate from context leak to stale assertion is a property of your corpus, not of ours. That is the first question a pilot scopes.

Judge measurement · n=400 pairs+

Supporting diagnostic

Why the deterministic path holds the enforcement role.

We measured two LLM judges against deterministic ground truth before letting either one near enforcement. Both failed the same way: calling genuine contradictions unrelated. The deterministic path scores those exactly, by construction. The judges stay in shadow, with their error rates on the record.

Gemini 3.1 Flash Lite as judge

36/400 total

The pooled error rate is 9.0%, but the failure that matters is contradiction blindness: 36 of roughly 200 genuine contradictions were judged unrelated. Zero call failures.

Claude Haiku 4.5 as judge

42/400 total

The pooled error rate is 10.5%, but the contradiction subset is sharper: 41 of roughly 200 genuine contradictions were judged unrelated, plus one inverted support. Zero call failures.

What this shows

Live model answers can be dispositioned deterministically against governed records, across three model families, with no answer left unaccounted for.

Off-script behavior surfaces as a flag, not an uncaught miss, under the governed context condition.

Measured judge errors, especially on genuine contradiction rows, justify keeping deterministic relations primary; the architecture choice now has data, not only principle.

Four design corrections made during smoke runs are disclosed in the methods report, including a provider schema-enforcement asymmetry that invalidated the first GPT-5.5 run.

What this does not show

Not third-party validation; self-administered on public data.

Not a hallucination detector. We flag answers that conflict with or escape your governed records, which is a narrower and more checkable promise.

Not evidence the models are unsafe; they answered correctly almost every time. The product is the accounting, not the alarm.

Not a generalizable stale-assertion rate. That rate belongs to the corpus, and yours is different.

Mechanism lineage: this diagnostic extends the template-answer attribution act of the public corpus authority series, which carries the full enforcement, falsification, recovery, and exposure arc this work builds on.