Evidence and boundary tests

Claims ship with their boundaries.

Every page below shows the exact test, the result, and the limit of the claim. These are diagnostics, not universal product guarantees.

Governed Memory covers which records may shape an answer. Governed Actions covers which agent actions may run. Governed Precedent covers which reviewed decisions may carry forward.

The strongest public bundles include row-level artifacts and a standard-library verifier that recomputes the figures and fails on drift. Negative results and ties are published, not buried.

Start here

The Governed Learning Loop

See how FieldHash turns authorized review into bounded Authority State across Memory, Actions, and Precedent, and how drift returns prior decisions to review.

Follow the loop

Flagship proof

Model behavior + enforcement

Governed Actions

Part I tests revoked tool versions on real PyPI yank data. Part II tests unapproved consequential actions and exact-payload approval. Together they show what changes when the gate acts before selection and dispatch.

Context exclusion + answer path

Governed Memory

On the public MemConflict benchmark, FieldHash forwards the governing record, withholds the superseded alternative, and records the context supplied for the answer. Context reduction is secondary to proving exclusion.

Bounded reviewer-decision reuse

Governed Precedent

A reviewed decision can guide the next similar case, but only while its scope and evidence still hold. Drift, expiry, conflict, missing evidence, or tampering sends the case back to review.

Boundary research

Semantic-support boundary

Related Text Is Not Governing Evidence

A 2,340-row FedReg/eCFR-derived hard-negative profile tests whether related language is withheld when it does not support the answerable fact. This is in-profile proxy evidence, not open-world generalization.

Authority-inference boundary

Where Simpler Controls Are Enough

A pre-registered study publishes a selection tie and the conditions where a filter or prompt is sufficient, then separates selection accuracy from lifecycle, review, and accounting.

Public-source ablations

Public Corpus Authority Series

Public RFC conflicts show enforcement when authority is configured, then expose what collapses when the authority signal is hidden.

Answer accounting

Live Answer Attribution

Across 1,000 live model answers, each output resolves to supported, contradicted, or unattributable rather than passing through unaccounted for.

How to read the metrics+

Business failure count

A workflow outcome such as a revoked tool, stale record, or unapproved action reaching the governed path.

Model behavior count

What a model selected or attempted under the stated prompt, context, and model configuration.

Packet-contract verification

Whether generated evidence artifacts satisfy the published packet schema and accounting rules.

Artifact-tamper verification

Whether a deliberate mutation causes the verifier to reject the changed artifact or bundle.

Context reduction

How much candidate material was withheld from the model input under the tested policy.

Internal synthetic diagnostic

A FieldHash-authored mechanism test. Useful for engineering evidence, not customer validation.

Research archive

Supporting diagnostics remain public, without crowding the buyer path.

Internal stress tests, lifecycle diagnostics, context-compression results, ties, and negative controls remain available for technical scrutiny.

Browse research archive

FieldHash is architecturally independent of the systems it governs. Its current public diagnostics are self-administered; third-party validation is a separate milestone. Evidence sources, caveats, methods, and verification links live on the individual study pages.