Related text should not pass as governing evidence.

A relevant passage can still support the wrong answer. Once a governing source is known, FieldHash checks whether it supports the proposed fact. Unsupported or uncertain bindings go to review or abstention, with a record of why.

Result

Supported facts passed. Unsupported candidates stayed out.

In this FedReg/eCFR-derived proxy diagnostic, FieldHash produced 682 clean packets with zero wrong bindings and withheld all 1,658 safety-route rows from clean allow. Every one of the 2,340 rows received an evidence record. The strongest local comparator, a cross-encoder NLI model, made 84 wrong bindings on the same 682 clean-eligible rows.

Controlled, self-administered evidence from one configured profile with structured authority signals. The 2,340 rows are related variants of 110 source cases, not independent trials or customer validation. See the study scope below for generalization limits.

Boundary rows

2,340

Supported answers and near misses, tested together.

The FedReg/eCFR-derived profile mixed 682 clean semantic bindings with conflict, partial support, no-governing-source, revoked, superseded, and near-miss rows.

Wrong clean allows

0/682

No wrong binding on clean-eligible rows.

No wrong binding occurred among 682 clean-allow-eligible rows. Separately, all 1,658 safety-route rows were withheld from clean allow under the configured profile.

Cross-encoder misses

84

A strong NLI baseline still mis-bound clean rows.

The local cross-encoder NLI comparator reached 0.964 overall accuracy, with 84 wrong bindings among the same 682 clean-allow-eligible rows.

Clean packets

682

Passed rows received verifier-backed packets.

Clean allows were exported only for profiled rows where the governing source safely bound to the answerable fact.

Tamper mutations

12/12

Every packet mutation was detected.

Status flips, hash edits, checksum changes, packet deletion, chain breaks, reason-code edits, and manifest mutations failed verification.

Underlying source cases

110

Source cases expanded into targeted variants.

The 2,340 rows derive from 80 confirmed, 18 conflicted, and 12 unverifiable source cases. Related variants are dependent observations.

What changed

Require support from the source that governs.

Two passages can discuss the same subject while contradicting each other, omitting a condition, or relying on an expired rule. For the team accountable for an answer, the question is whether the applicable source supports that specific fact.

FieldHash checks that connection before issuing a clean evidence packet. Within the tested profile, incomplete support, conflict, or an invalid source routes the case away from clean allow and leaves evidence for review.

The authority source is resolved first. This diagnostic tests the next decision: whether that source supports a differently worded answer, including plausible near misses.

The review decision

A convincing match still needs support. When the configured checks cannot establish it, FieldHash withholds the clean packet and records the reason so a reviewer can resolve the uncertainty.

Mechanism

Authority is resolved first. Semantics only decide whether a packet is allowed.

01

A governing record is resolved.

FieldHash starts with the authority source selected by policy, review, registry, or source governance. It does not ask the model to decide which source is allowed to govern.

02

The answerable fact is tested.

The resolver checks whether the governing record actually supports the differently worded candidate answer. Semantic similarity is supporting evidence, not the sole decision.

03

Unsafe cases route away.

Ambiguity, partial support, contradiction, corrupted bindings, missing evidence, or low confidence produce abstain or manual-review evidence instead of a clean packet.

04

Clean packets require a pass.

Only safe bindings receive clean evidence packets. Route-evidence packets still record why unsupported cases did not receive a clean allow.

Baselines

Sentence similarity found related text. It did not preserve the safety boundary.

The diagnostic compared FieldHash against sentence-nearest, sentence-logistic, TF-IDF, and a local NLI cross-encoder. FieldHash is a governed pipeline: authority is resolved first, then semantics decide packet eligibility. The baselines see the same rows and candidates, but they do not reproduce the full authority resolver. The NLI baseline preserved safety routes, but still chose the wrong clean answer 84 times. That is the boundary FieldHash is designed to protect.

Wrong clean allows count incorrect answer bindings among 682 clean-eligible rows. Unsafe false allows count clean allows among the separate 1,658 rows that should have abstained or routed to review. The bundle calls their sum silent mis-binds; that aggregate is not the wrong-clean count. The cross-encoder preserved all safety routes but chose the wrong answer on 84 clean-eligible rows.

ArmOverallClean allowSafe routeWrong clean allows / 682Unsafe false allows / 1,658
FieldHash boundary resolver1.0001.0001.00000
Sentence-nearest baseline0.4250.9000.229681,278
Sentence-logistic baseline0.4690.8900.296541,163
TF-IDF full-row baseline0.6820.1740.890119136
Cross-encoder NLI entailment0.9640.8771.000840

The cross-encoder comparator was given the same rows, source records, candidate facts, and explicit safety-route metadata. The remaining gap is clean answer binding, not access to the safety route.

Corrupted controls

Break the authority relation, and clean allow should collapse.

The controls deliberately corrupted source bindings and labels. If a resolver still passes clean allows under those conditions, it is probably learning a shortcut. Here the corrupted arms failed promotion: some collapsed almost completely, while inverse-source and random-label controls still produced residual clean allows and many wrong clean allows. That is why they stay controls, not promoted packet evidence.

Shuffled source bindings

overall0.490

clean allow0.004

wrong clean / 6820

unsafe false allows1

Inverse source bindings

overall0.394

clean allow0.235

wrong clean / 682496

unsafe false allows896

Inverse labels

overall0.709

clean allow0.000

wrong clean / 682682

unsafe false allows0

Random labels

overall0.238

clean allow0.245

wrong clean / 682155

unsafe false allows844

Threshold frontier

You can buy more caution. The packet tells you what it costs.

The threshold grid shows the operational tradeoff after structured authority resolution. Precision stayed flat across the displayed range because the authority resolver, not the threshold alone, prevents wrong-source clean allows in this corpus. Raising the floor preserved precision but routed more otherwise-clean rows to review. The threshold is a buyer-controlled operating setting, and the packet exposes its review cost.

A regulated workflow can choose a stricter threshold and accept more manual review. A lower-risk workflow can preserve recall. Either way, the packet records the route.

ThresholdPrecisionRecallClean rows routed awayWrong clean allows
0.1001.0001.00000
0.2001.0001.00000
0.4501.0001.00000
0.5001.0000.99440
0.5501.0000.98880
0.6001.0000.977160

Verifier path

Both allowed and withheld decisions left evidence.

The run emitted 682 clean packets and 2,340 route-evidence records. Clean packets show the accepted binding. Route evidence shows why each row was allowed, abstained, routed to manual review, or denied clean allow.

The independent verifier checked packet lineage, checksum manifests, status consistency, reason codes, route-evidence coverage, tamper mutations, and bundle integrity. Every deliberate mutation was caught.

Download the public verification bundle.

The bundle includes scrubbed row-level metrics, scrubbed packet JSONL with recomputed public hash chains, the fairness-audit output, checksums, and a standard-library verifier. It excludes raw source text, raw candidate text, prompts, provider logs, customer data, internal runner scripts, audit scripts, resolver modules, and local paths.

Tested failure modes

  • public-source FedReg/eCFR proxy rows
  • paraphrase and no-copy source values
  • conflicting public-source records
  • partial-support cases
  • revoked or superseded-source rows
  • no-governing-source rows
  • similar-but-wrong candidates
  • manual-review and abstain rows
  • shuffled and inverse controls
  • cross-encoder NLI comparator
  • packet tamper mutations

Study scope

Read the result within its tested scope.

This self-administered public-source proxy diagnostic tests clean-packet eligibility under one configured FedReg/eCFR-derived profile. The local cross-encoder comparator was also run by FieldHash, not a third party.

The governance layer resolves, binds, routes, and records after a governing source is known. The result does not show that a model learned authority or can resolve arbitrary semantic truth without structured authority signals.

The 2,340 rows derive from 110 source cases. Related variants do not justify an independent-trial safety bound or a production failure-rate estimate.

Cross-generator paraphrases, independently authored paraphrase corpora, and customer-specific resolver generalization remain outside the tested scope.

Customer-corpus validation and production readiness require a customer-scoped evaluation. These results do not establish either.

Next step

Test the authority behind your answers.

Bring one workflow where a plausible but unsupported answer creates real review risk. Evaluate your records, authority hierarchy, and review routes in shadow mode, then use the findings to decide which handoffs to govern.