Semantic authority boundary
Related text should not pass as governing evidence.
Semantic binding is the technical test beneath that claim. FieldHash should bind only when the governing record supports the answerable fact, and route everything else away from clean allow before a packet can overstate the evidence.
Result
The test was not whether FieldHash could bind meaning. It was whether it could refuse the wrong binding at scale.
In a 2,340-row FedReg/eCFR-derived proxy diagnostic, FieldHash produced 682 clean packets where the governing record safely supported the answerable fact, 2,340 route-evidence records covering every row, and zero wrong clean allows under that profile. Sentence baselines found related text. A local cross-encoder NLI baseline was much stronger. It still produced 84 wrong clean allows.
This is controlled public-source proxy evidence, not customer validation. FedReg/eCFR is hard on semantic similarity, but it still carries structured authority signals, and this run does not prove cross-generator or independently authored paraphrase generalization. The useful product claim here is bounded abstention after authority has been resolved: when authority cannot be safely bound to an answer under the configured resolver path, FieldHash withholds the clean packet and records why.
Boundary rows
2,340
A profiled public-source diagnostic tested the edge.
The FedReg/eCFR-derived profile mixed 682 clean semantic bindings with conflict, partial support, no-governing-source, revoked, superseded, and near-miss rows.
Wrong clean allows
0
No unsupported row received a clean allow in-profile.
Under this configured profile, the resolver did not bind a governing source to the wrong answerable fact, and did not allow unsupported rows to pass as clean evidence.
Cross-encoder misses
84
A strong NLI baseline still mis-bound clean rows.
The local cross-encoder NLI comparator reached 0.964 overall accuracy, but still produced 84 wrong clean allows on the same row set.
Clean packets
682
Passed rows received verifier-backed packets.
Clean allows were exported only for profiled rows where the governing source safely bound to the answerable fact.
Tamper mutations
12/12
Every packet mutation was detected.
Status flips, hash edits, checksum changes, packet deletion, chain breaks, reason-code edits, and manifest mutations failed verification.
Rule-of-three bound
0.13%
The caveat is part of the result.
Zero observed wrong clean allows in 2,340 rows gives a conservative 95% upper bound near 0.13%. This is still proxy evidence, not customer validation.
What changed
Similarity is cheap. Bounded evidence is not.
A model or embedding system can often tell that two statements are related. That is not enough for governed inference. Related text can be opposite, incomplete, expired, partially supported, or drawn from the wrong source.
FieldHash treats semantic binding as a packet-eligibility gate. The resolver can help widen the product beyond perfectly structured records, but this diagnostic only tests that claim inside a configured FedReg/eCFR-derived profile. What makes it defensible is the negative boundary: no clean evidence packet when the source does not safely govern the answer.
This diagnostic does not prove governance over records with no structured authority at all, and it does not prove general paraphrase handling outside this profile. It tests the next layer down: once FieldHash has a governing signal, can it bind that authority to differently worded answer facts without clean-allowing near matches in a controlled profile?
The trap avoided
If semantic binding becomes a confident model judgment inside the gate, FieldHash becomes another grounding layer. If it is abstention-biased, verifier-backed, and packet-gated, it extends governed inference into messy records while preserving the core promise.
Mechanism
Authority is resolved first. Semantics only decide whether a packet is allowed.
01
A governing record is resolved.
FieldHash starts with the authority source selected by policy, review, registry, or source governance. It does not ask the model to decide which source is allowed to govern.
02
The answerable fact is tested.
The resolver checks whether the governing record actually supports the differently worded candidate answer. Semantic similarity is supporting evidence, not the sole decision.
03
Unsafe cases route away.
Ambiguity, partial support, contradiction, corrupted bindings, missing evidence, or low confidence produce abstain or manual-review evidence instead of a clean packet.
04
Clean packets require a pass.
Only safe bindings receive clean evidence packets. Route-evidence packets still record why unsupported cases did not receive a clean allow.
Baselines
Sentence similarity found related text. It did not preserve the safety boundary.
The diagnostic compared FieldHash against sentence-nearest, sentence-logistic, TF-IDF, and a local NLI cross-encoder. FieldHash is a governed pipeline: authority is resolved first, then semantics decide packet eligibility. The baselines see the same rows and candidates, but they do not reproduce the full authority resolver. The NLI baseline preserved safety routes, but still chose the wrong clean answer 84 times. That is the boundary FieldHash is designed to protect.
How to read the table: wrong clean allows are clean allows for the wrong answerable fact. Unsafe false allows are clean allows on rows that should have abstained or routed to review. That is why the cross-encoder can score 1.000 on safe routing while still producing 84 wrong clean allows: it avoided unsafe review-route violations, but chose the wrong clean answer on supported rows.
| Arm | Overall | Clean allow | Safe route | Wrong clean allows | Unsafe false allows |
|---|---|---|---|---|---|
| FieldHash boundary resolver | 1.000 | 1.000 | 1.000 | 0 | 0 |
| Sentence-nearest baseline | 0.425 | 0.900 | 0.229 | 1,346 | 1,278 |
| Sentence-logistic baseline | 0.469 | 0.890 | 0.296 | 1,217 | 1,163 |
| TF-IDF full-row baseline | 0.682 | 0.174 | 0.890 | 255 | 136 |
| Cross-encoder NLI entailment | 0.964 | 0.877 | 1.000 | 84 | 0 |
The cross-encoder comparator was given the same rows, source records, candidate facts, and explicit safety-route metadata. The remaining gap is clean answer binding, not access to the safety route.
Corrupted controls
Break the authority relation, and clean allow should collapse.
The controls deliberately corrupted source bindings and labels. If a resolver still passes clean allows under those conditions, it is probably learning a shortcut. Here the corrupted arms failed promotion: some collapsed almost completely, while inverse-source and random-label controls still produced residual clean allows and many wrong clean allows. That is why they stay controls, not promoted packet evidence.
Shuffled source bindings
overall0.490
clean allow0.004
wrong clean allows1
unsafe false allows1
Inverse source bindings
overall0.394
clean allow0.235
wrong clean allows1,392
unsafe false allows896
Inverse labels
overall0.709
clean allow0.000
wrong clean allows682
unsafe false allows0
Random labels
overall0.238
clean allow0.245
wrong clean allows999
unsafe false allows844
Threshold frontier
You can buy more caution. The packet tells you what it costs.
The threshold grid shows the operational tradeoff after structured authority resolution. Precision stayed flat across the displayed range because the authority resolver, not the threshold alone, prevents wrong-source clean allows in this corpus. Raising the floor preserved precision but routed more otherwise-clean rows to review. That is a buyer-control surface, not a hidden model behavior.
A regulated workflow can choose a stricter threshold and accept more manual review. A lower-risk workflow can preserve recall. Either way, the packet records the route.
| Threshold | Precision | Recall | Clean rows routed away | Wrong clean allows |
|---|---|---|---|---|
| 0.100 | 1.000 | 1.000 | 0 | 0 |
| 0.200 | 1.000 | 1.000 | 0 | 0 |
| 0.450 | 1.000 | 1.000 | 0 | 0 |
| 0.500 | 1.000 | 0.994 | 4 | 0 |
| 0.550 | 1.000 | 0.988 | 8 | 0 |
| 0.600 | 1.000 | 0.977 | 16 | 0 |
Verifier path
Both allowed and withheld decisions left evidence.
The run emitted 682 clean packets and 2,340 route-evidence records. Clean packets show the accepted binding. Route evidence shows why each row was allowed, abstained, routed to manual review, or denied clean allow.
The independent verifier checked packet lineage, checksum manifests, status consistency, reason codes, route-evidence coverage, tamper mutations, and bundle integrity. Every deliberate mutation was caught.
Download the public verification bundle.
The bundle includes scrubbed row-level metrics, scrubbed packet JSONL with recomputed public hash chains, the fairness-audit output, checksums, and a standard-library verifier. It excludes raw source text, raw candidate text, prompts, provider logs, customer data, internal runner scripts, audit scripts, resolver modules, and local paths.
Tested failure modes
- public-source FedReg/eCFR proxy rows
- paraphrase and no-copy source values
- conflicting public-source records
- partial-support cases
- revoked or superseded-source rows
- no-governing-source rows
- similar-but-wrong candidates
- manual-review and abstain rows
- shuffled and inverse controls
- cross-encoder NLI comparator
- packet tamper mutations
Claim boundary
What this proves, and what it does not.
This is a controlled FieldHash diagnostic, not a customer-corpus validation claim.
It does not claim production readiness or arbitrary open-world semantic truth resolution.
It does not claim the model learned authority. The governance layer resolves, binds, routes, and records.
It tests a narrower claim: under this configured FedReg/eCFR-derived profile, clean packets are emitted only when a governing source safely binds to the answerable fact.
FedReg/eCFR is hard on wording but still provides structured authority signals; this does not validate records with no resolvable authority source.
It does not validate cross-generator paraphrases, independently authored paraphrase corpora, or customer-specific semantic resolver generalization.
The 2,340-row result is still self-administered and public-source proxy evidence. The cross-encoder baseline is locally run, not third-party administered.
A customer-scoped pilot remains required before customer-specific validation.
Next step
Bring the boundary to one workflow.
A pilot should test your real records, your authority hierarchy, your review routes, and your tolerance for manual review. FieldHash should show what was allowed, what was blocked, what was routed away, and why.