Keep stale records out of the next answer.
Across 300 public MemConflict questions, a strong prompt left stale records in the model's context with labels and asked it to ignore them: 206 appearances. Governed Memory kept them out of the selected context, 0 appearances, with about a third of the context (323 vs 956 mean characters).
Authority came from the benchmark's labels. On its rule scorer the strong prompt scored 297/300 and FieldHash 283/300; 15 of FieldHash's 17 misses were the exact correct value stated without narrating the conflict.
Result
Less context. Inspectable gate decisions.
Naive retrieval sends ~16,020 characters and scores 230/300. A strong prompt scores 297/300 but leaves the stale record in context and asks the model to ignore it. FieldHash forwards only the governing record (~323 characters: ~3x less than the prompt, ~50x less than naive retrieval) and scores 283/300 on the benchmark's rule scorer. Seventeen outputs did not satisfy the benchmark's narrative-form scorer. Manual review found 15 contained the correct value in a shorter form, one abstained, and one matched the comparison prompt. FieldHash's distinct claim is that the stale record was excluded before generation and the governed context was recorded for every answer.
Boundary: a configured-authority diagnostic on the public sample, not an accuracy-beating or leaderboard claim. gpt-5.5, n=300.
Historical harness correction: numeric zero was rendered as blank text by str(value or ""). On a retained zero-siblings case, the frontier answer called the value "unspecified" yet received 0.5 from the rule scorer; Gemma returned a blank count, and Qwen invented counts. The published totals retain those outputs and scores. This input defect limits answer-quality interpretation across these models; it does not change the observed gate routes.
Attributable
All 300
Every answer carries a packet.
What governed the handoff, what was blocked, and why, all hash-chained, for every one of the 300 answers by construction. This is a property of the mechanism, not a benchmark score. The packet records allowed and blocked records; the separate harness rows record selected context. The packet does not bind the actual model-input bytes.
Observed exclusion
0
Blocked records in selected context.
No blocked record appears in the selected-memory fields of the 300 governed harness rows. The packet records the gate decision; its hash chain does not independently prove what the model received.
Efficient
~50x less
Context the model has to read.
~323 characters of governed context per answer: ~3x less than even a tight authority-labeled prompt, ~50x less than naive retrieval that dumps every candidate.
Why the context collapses
Retrieval surfaces everything. Governed memory forwards one record.
A memory query can match the current record and several superseded or contradicted ones. Naive retrieval hands the model the whole set and asks it to sort them out, carrying the full text into context on every call, and scoring worst for it.
FieldHash resolves which record governs from the configured authority state, forwards only that record, and keeps the blocked alternatives in the packet. The harness selects less context before answering. The retained outputs and scores include both model errors and the shared zero-rendering defect disclosed above.
Per-answer context, what reaches the model, and score
| Arm | Context chars | Stale records reaching the model | Score |
|---|---|---|---|
Naive retrieval (recency-aware) Dumps every plausible candidate as raw session text; the model muddles the conflict. | 16,020 | all | 230/300 |
Strong labeled prompt Governing and blocked records are both in context, labeled, with the stale one trusted to be ignored. | 956 | all (labeled) | 297/300 |
FieldHash governed memory Only the governing record reaches the model; blocked alternatives are removed and kept in the packet. | 323 | none | 283/300 |
Why attribution is the point
A correct answer does not prove which records entered model context.
The strong prompt leaves the stale record in context and asks the model to ignore it. The retained harness rows show the stale record excluded from selected context; the packet separately records the gate’s allowed and blocked record IDs.
When an output misses the benchmark scorer, the packet still shows which record the gate allowed and which alternative it withheld. The 17 scorer misses are classified below and remain available in the row-level bundle.
A review decision must be distinguished from the answer that follows it. This evaluated harness still called the model after a manual-review disposition and relied on an instruction not to guess. In the local addendum, Qwen answered both review cases with invented governing and blocked records. The gate’s review routing held; downstream abstention was not enforced.
What the packet records
- Allowed record: the one approved by the gate
- Blocked records: superseded or contradicted, marked withheld by the gate
- Reason: why each record was allowed or blocked
- Answer hash + chain: reviewer-verifiable after the fact; stronger tamper evidence depends on deployment configuration
The packet records the gate decision. Selected context and model answers remain separate harness observations.
Reading the failures
Here are FieldHash's misses. Read them yourself.
The benchmark counts these as misses. The packet shows what each one actually was, so judge for yourself.
The user’s gender
“Male.”
The gold says the correct value is Male. FieldHash returns it, docked only for not adding “sources disagree.”
Number of siblings
“The user has 1 sibling.”
The exact correct count. Docked because the gold narrates the disagreement and FieldHash answers plainly.
When they watch the NBA
“The memory needs manual review.”
A conflict it could not cleanly resolve, so it declined. The strong prompt declined here too.
Seventeen outputs did not satisfy the benchmark's narrative-form scorer. Manual review found 15 contained the correct value in a shorter form, one abstained, and one matched the comparison prompt. Every row remains in the verification bundle for independent classification.
Claim boundary
What this page does and does not claim.
- This is a private configured-authority diagnostic on the public MemConflict released sample, not an official leaderboard result.
- FieldHash forwards only the governing record: ~50x less context than naive retrieval, ~3x less than a strong prompt.
- On the benchmark’s rule scorer FieldHash scores 283/300 to the strong prompt’s 297/300. Read all 17 misses: 15 are the exact correct value stated without narrating the conflict, one is an abstention the prompt also made, one matches the prompt verbatim. The shared input-rendering defect and judge disagreement limit answer-quality conclusions; these scores do not establish an accuracy advantage.
- Every answer carries a hash-chained packet recording the gate’s disposition. Exclusion is observed in the harness’s selected-memory fields, not cryptographically bound to actual model input by that packet.
- The public checker recomputes row totals, scores, context lengths, and the packet hash chain. It does not independently replay selection, bind the model request to the packet, or prove the authority configuration correct. Choosing which record governs is an upstream policy this diagnostic does not test.
- The run uses benchmark conflict metadata as configured authority state. The claim is governed-memory influence and attribution, not autonomous timeline inference.
- Primary run: single model (gpt-5.5), n=300 stratified. The model-invariance addendum below reruns the same governed path with two local models. The published bundles retain row-level scores and context measurements; their checkers reproduce specified totals and packet-chain consistency.
Model-invariance addendum · two local models · two seeds · same-day dual judge
We swapped the frontier model for a laptop. The gate didn't notice.
Same harness, same 300 questions, same seeds, same official judge. The only change was the model behind the gate: the frontier API above, then two 4-bit models running on a laptop through a local endpoint. The gate blocked the same records, routed the same conflicts to review, and every packet chain verified. Answer reliability varied even though the same exclusion and review checks held.
| Model behind the gate | Naive retrieval | Strong labeled prompt | FieldHash governed |
|---|---|---|---|
| gpt-5.5 (frontier API) | 154/300 | 295/300 | 290/300 |
| Gemma 4 E2B, 4-bit, laptop | 128/300 | 278/300 | 279/300 |
| Qwen 2.5 1.5B, 4-bit, laptop | 134/300 | 259/300 | 249/300 |
856 = 856 = 856
Records blocked, per model
The gate blocked the identical set of records for all three models on the primary seed, and 868 for both local models on the second seed. Governance decisions did not depend on the model.
2
Identical review routes
Every run selected manual review for the same two conflicts. The harness still called the answer model; Qwen invented governing and blocked records in both cases on both seeds. Review routing did not enforce answer abstention.
0
Packet-chain errors
All packet chains verified across four local runs. The bundle retains the rows and a checker for specified counts and chain consistency.
How to read this fairly
- Paired comparisons with the strong labeled prompt were not significant (Gemma discordant pairs 7 vs 6, exact McNemar p=1.000; Qwen 20 vs 30, p=0.203). This establishes neither equivalence nor superiority.
- The tier gap is real and disclosed: Gemma governed trails frontier governed 279 to 290 on the same judge (p=0.019). Qwen trails clearly. The same governance checks held, while answer reliability varied by model.
- Both local governed arms beat the ungoverned frontier arm on the same 300 questions: Gemma won 133 paired questions and lost 8 (p=2.4e-30); Qwen won 120 and lost 25 (p=4.4e-16).
- Same harness, seeds, sampling, temperature, and token budget as the published run, verified against its artifact; a 128-token sensitivity run produced identical seed-1 results. Local arms used an OpenAI-compatible local endpoint (chat completions) rather than the Responses API.
- The second judge changes counts and one ordering. For Gemma, gpt-4o scored governed 279/300 and prompt 278/300; Gemini 3.5 Flash scored governed 288/300 and prompt 289/300. For Qwen, the corresponding governed/prompt counts were 249/259 and 268/271. Judge choice affects answer-quality comparisons.
- Reading the local misses: retained Gemma answers include polarity errors and blanks caused by the shared harness converting numeric zero to empty text. The zero-rendering defect also affected the frontier and Qwen inputs; it is not solely a Gemma failure. Historical scores remain unchanged and include answers scored on those defective inputs. A corrected experiment would need a separately identified run.
- Task-scoped: answering from a governed record after enforcement, on 4-bit quantized weights. Not open-ended generation parity, and not customer validation.
The deployment implication is bounded: these governance checks also held with locally run models, but model choice still requires an answer-reliability evaluation. The local runs also demonstrate a data-location option: the governed records were answered on the hardware where they were held.
Why it matters
Inspect the answer path, and diagnose when it is wrong.
A model answer alone does not establish which records entered its input or whether a stale alternative was present. FieldHash records what the gate allowed, what it withheld, and why. Reviewers can compare that packet with the harness’s selected-memory fields and model answer to diagnose misses. The packet alone does not prove which bytes reached the model.
That is Governed Memory in one line: select records from configured authority and keep the gate decision available for review.