Reviewer-approved authority promotion

Governed Memory turns authority signals into the governed answer path.

This evidence brief tests the upstream step before enforcement: turning benchmark authority evidence into governed state before the model sees answer context. On the blind n=30 cross-provider probe, FieldHash recovered 90/90 current tokens across Gemini, Claude, and GPT without pre-seeded canonical memory labels.

Cross-provider blind probe

90/90

Gemini, Claude, and GPT n=30

Two-pass smart diagnostic

28/30

current record selected 30/30; Gemini n=30

N=100 governed path reruns

Gemini/GPT 100/100

FieldHash + provider; Claude 95/100 with 0 stale substitutions

FieldHash fact extraction audit

99/100 | 95/100

role-equiv Gemini | GPT; exact spans 68/100 | 76/100

Dense-memory stress

2,156 → 500

no-model compaction test; 100/100 current-token retained

Two regimes are shown: the n=30 blind cross-provider probe includes Gemini, Claude, and GPT on a Claude-authored disjoint corpus; the n=100 governed path reruns use the same frozen corpus, where Gemini and GPT completed cleanly while Claude produced five empty provider responses.

In practical terms, this is the moment before memory becomes useful. A workspace may contain an old plan, a correction, a rejected idea, and a newer approved direction. FieldHash has to turn trusted authority evidence into governed state before the model ever sees the answer context.

The benchmark records were not routed with pre-seeded canonical memory labels; expected labels were used to score the promoted state after the run. This is not a claim that FieldHash infers real-world authority from arbitrary prose without configured signals.

The outcome: Reviewed authority state resolves the evaluated conflict before the model is called, producing a cleaner governed answer path and keeping evaluated stale, adversarial fragments out of the answer path.

Approved vendor policy

The selected vendor rule enters the context sent to the model instead of older procurement notes.

Current contract term

The active clause wins over superseded draft language that still matches the query.

Rejected analysis path

A discarded recommendation remains visible, but cannot become the answer.

Operational correction

The reviewed fix carries forward before the model sees stale troubleshooting context.

Why this matters

The hard part is not only remembering. It is turning the authoritative memory into the next answer path.

The earlier memory-pressure diagnostic tested whether FieldHash honors an approved-current label once that label already exists. This benchmark moves upstream inside a governed corpus: the records are not handed to the answer path with canonical memory labels, and FieldHash must promote authority state before the answer is constructed.

That distinction matters for real work. Teams do not only need persistent memory; they need memory that can demote stale context, reject discarded alternatives, and carry the reviewed version forward without hiding the historical trail.

The evidence catalog keeps this page in context with the downstream diagnostics: the current record can be retrieved into candidates and still fail to govern the answer path.

Scope boundary: this is authority-state promotion inside a governed benchmark corpus, not independent certification of real-world truth. In production, systems of record and review, supersession, and rollback signals define real-world authority; FieldHash enforces those configured signals and governs what reaches the answer path.

Core comparison

FieldHash governed promotion

90/90 exact tokens

Prompt-only smart memory

40/90 exact tokens

Retrieval-only memory

36/90 exact tokens

Recency-aware memory

3/90 exact tokens

FieldHash Gemini slice

30/30 exact tokens; matched to smart diagnostic

Same-budget two-pass smart diagnostic

28/30 exact tokens; selector chose current record 30/30

The cross-provider n=30 probe used a Claude-authored disjoint corpus. FieldHash recovered every current token across Gemini, Claude, and GPT answer paths; prompt-only instructions and retrieval-only memory did not. The fair same-model comparison is narrower: FieldHash + Gemini reached 30/30 on this slice, while a later same-budget Gemini two-pass smart diagnostic on the same corpus selected the current record 30/30 and answered 28/30, with zero stale substitutions. That narrows the claim: the advantage is governed, auditable answer-path control, not that a frontier model cannot identify the current record when given a separate selection pass.

Provider results

The governed answer path replicated across providers.

The n=100 reruns use the same frozen corpus and semantic-label artifact. The point is not to rank models or prove basic semantic selection; it is to test whether Governed Memory keeps stale context out of the answer surface those models receive.

Gemini 3.5 Flash

FieldHash + Gemini on the same-family n=100 corpus.

LLM + FieldHash100/100
Prompt-smart memory25/100
Retrieval-only memory23/100
Recency-aware memory2/100
FieldHash stale mentions0

GPT-5.5

FieldHash + GPT on the same-family n=100 corpus.

LLM + FieldHash100/100
Prompt-smart memory30/100
Retrieval-only memory23/100
Recency-aware memory0/100
FieldHash stale mentions0

Claude Opus 4.7

FieldHash + Claude; five misses were empty provider responses, not stale substitutions.

LLM + FieldHash95/100
Prompt-smart memory17/100
Retrieval-only memory14/100
Recency-aware memory0/100
FieldHash stale mentions0

Promotion audit

Fact extraction was strong, but provider-dependent.

The n=100 FieldHash promotion audits split current-token recovery from stricter fact-span extraction. The Gemini audit rescored the FieldHash + Gemini semantic-label artifact: 100/100 current-token recovery, 99/100 role-equivalent facts, and 68/100 exact source spans. A FieldHash + GPT-5.5 promotion rerun on the same corpus reached 99/100, 95/100, and 76/100 respectively. Its one false promotion was a stale unreviewed handoff promoted as current/reviewed, the exact class of error this layer is meant to catch. Strict source-span fidelity, provider-invariant fact extraction, and zero false-promotion operation are not claimed as solved.

FieldHash + Gemini audit

Current-token recovery

100/100

Role-equivalent current facts

99/100

Exact source-span match

68/100

FieldHash + GPT-5.5 rerun

Current-token recovery

99/100

Role-equivalent current facts

95/100

Exact source-span match

76/100

Dense-memory stress

Crowded memory can be compressed without dropping the current fact.

This is a no-model mechanism test, not another model-in-the-loop answer run. It duplicated optional stale and noise fragments from 500 base records to 2,156 crowded records. Protected state compaction reduced the candidate set back to 500 records while preserving 100/100 current-token recovery and 99/100 role-equivalent facts.

Crowded optional records

2,156 input records

After protected state compaction

500 records retained

Records removed

1,656 optional records

Mechanism

Governed state promotion, defined as a pre-answer control plane.

The benchmark is not claiming that the base model changed its weights. It shows a governed non-parametric loop: promote durable state from configured authority evidence, update authority metadata, filter stale context, then condition the next model call on the reviewed version.

Step 1

Unlabeled memory

A project has reviewed notes, superseded records, unreviewed handoffs, rejected alternatives, and near-duplicate neighboring projects.

Step 2

Promotion into governed state

Inside the governed benchmark corpus, the promotion pass turns configured authority evidence into current, superseded, rejected, and ordinary state before answer construction.

Step 3

Governed arbitration

Governed Memory enforces which records are allowed to influence the answer path before the base model responds.

Step 4

Clean answer path

The model sees the current operational fact and a smaller review surface, while stale context remains auditable outside the answer path.

What this supports

Governed state promotion can change future answers without changing the base model.

Inside this governed corpus, the system promoted the authoritative memory, kept stale context available for audit, and prevented it from shaping the next answer.

This is the useful version of learning for high-stakes workflows: not hidden weight updates, but visible promotion, supersession, scope, compression, and answer-path control.

Reviewer boundaries

How to read the result.

Reviewer boundary: internal synthetic adversarial benchmark. The strongest independence check is the Claude-authored disjoint n=30 corpus; the n=100 governed path reruns use a same-family Gemini-authored corpus and frozen semantic-label artifact. On the n=30 Gemini slice, FieldHash reached 30/30 while a same-budget Gemini two-pass diagnostic selected the current record 30/30 and answered 28/30 with zero stale substitutions, so the public claim is governed answer-path control under singleton-current memory conflict rather than superiority over every same-budget selector. Broad reasoning superiority, universal memory safety, model-weight learning, provider-invariant fact extraction, zero false-promotion operation, and perfect source-span extraction are not claimed.

The result is best read as an internal architecture benchmark for memory-state promotion under adversarial stale-context pressure. External validation on public knowledge-conflict or memory-update datasets remains the next credibility step.

Continue with the full governed loop.

Reviewer-approved promotion tests the upstream step. The loop synthesis shows how promotion, enforcement, lifecycle state, and audit fit together.