Evidence and boundary tests
Claims ship with their boundaries.
Every page below shows the exact test, the result, and the limit of the claim. These are diagnostics, not universal product guarantees.
Governed Memory covers which records may shape an answer. Governed Actions covers which agent actions may run. Governed Precedent covers which reviewed decisions may carry forward.
The strongest public bundles include row-level artifacts and a standard-library verifier that recomputes the figures and fails on drift. Negative results and ties are published, not buried.
Start here
The Governed Learning Loop
See how FieldHash turns authorized review into bounded Authority State across Memory, Actions, and Precedent, and how drift returns prior decisions to review.
Flagship proof
Model behavior + enforcement
Governed Actions
Part I tests revoked tool versions on real PyPI yank data. Part II tests unapproved consequential actions and exact-payload approval. Together they show what changes when the gate acts before selection and dispatch.
Context exclusion + answer path
Governed Memory
On the public MemConflict benchmark, FieldHash forwards the governing record, withholds the superseded alternative, and records the context supplied for the answer. Context reduction is secondary to proving exclusion.
Bounded reviewer-decision reuse
Governed Precedent
A reviewed decision can guide the next similar case, but only while its scope and evidence still hold. Drift, expiry, conflict, missing evidence, or tampering sends the case back to review.
Boundary research
Semantic-support boundary
Related Text Is Not Governing Evidence
A 2,340-row FedReg/eCFR-derived hard-negative profile tests whether related language is withheld when it does not support the answerable fact. This is in-profile proxy evidence, not open-world generalization.
Authority-inference boundary
Where Simpler Controls Are Enough
A pre-registered study publishes a selection tie and the conditions where a filter or prompt is sufficient, then separates selection accuracy from lifecycle, review, and accounting.
Public-source ablations
Public Corpus Authority Series
Public RFC conflicts show enforcement when authority is configured, then expose what collapses when the authority signal is hidden.
Answer accounting
Live Answer Attribution
Across 1,000 live model answers, each output resolves to supported, contradicted, or unattributable rather than passing through unaccounted for.
How to read the metrics+
Business failure count
A workflow outcome such as a revoked tool, stale record, or unapproved action reaching the governed path.
Model behavior count
What a model selected or attempted under the stated prompt, context, and model configuration.
Packet-contract verification
Whether generated evidence artifacts satisfy the published packet schema and accounting rules.
Artifact-tamper verification
Whether a deliberate mutation causes the verifier to reject the changed artifact or bundle.
Context reduction
How much candidate material was withheld from the model input under the tested policy.
Internal synthetic diagnostic
A FieldHash-authored mechanism test. Useful for engineering evidence, not customer validation.
Research archive
Supporting diagnostics remain public, without crowding the buyer path.
Internal stress tests, lifecycle diagnostics, context-compression results, ties, and negative controls remain available for technical scrutiny.
FieldHash is architecturally independent of the systems it governs. Its current public diagnostics are self-administered; third-party validation is a separate milestone. Evidence sources, caveats, methods, and verification links live on the individual study pages.