Claims ship with their boundaries.

Every page below shows the exact test, the result, and the limit of the claim. These are diagnostics, not universal product guarantees.

Governed Memory covers which records may shape an answer. Governed Actions covers which agent actions may run. Governed Precedent covers which reviewed decisions may carry forward.

Check each bundle’s stated verification boundary. Some tools aggregate retained rows; others check published outcome consistency, file integrity, or signatures. Those checks do not automatically replay the authority decision or prove actual model input or execution. Study pages distinguish retained run reports from what their public checkers reproduce, including negative results and ties.

Start here

Flagship proof

Governed Agents · multi-agent flagship

Keep authority intact as agents work together

An authored, self-administered synthetic study follows five cooperating agent roles. Full mediation completed 84/84 main workflows safely with no known unauthorized outcome. Per-action authority, the strongest comparator we built for this study, completed 75/84 safely, with seven known unauthorized workflows and two safety-unknown failures. All seven known violations exceeded the shared budget. Repeat protection is evaluated in separate follow-ups with their own denominators and unmet criteria.

Governed Actions · execution surfaces

Authority Across Execution Surfaces

Across separate live episodes, Kimi recorded 46 prompt-only, 47 request-local, 42 incomplete-mediation, and 0 enumerated-surface unauthorized outcomes; Terra recorded 47, 48, 42, and 0. All 128 authorized objectives completed under enumerated-surface FieldHash authority.

Governed Actions · live effect authority

Authority After Denial

In separate live episodes, Kimi produced 45 prompt-only, 44 exact-action, and 0 FieldHash-governed unauthorized executions; Terra produced 45, 44, and 0. Across the governed arms, 127 of 128 authorized objectives completed and 60 of 61 recovery opportunities found an authorized continuation.

Context exclusion + answer path

Governed Memory

On 300 MemConflict questions, FieldHash kept stale records out of the selected context: 0 appearances, against 206 for a strong prompt that left them in with labels. FieldHash used about a third of the context (323 vs 956 mean characters). The prompt scored 297/300 on the rule scorer and FieldHash 283/300; 15 of FieldHash’s 17 misses were the correct value stated tersely.

Bounded reviewer-decision reuse

Governed Precedent

The 83-case diagnostic records 10 valid reuse cases, 3 valid time-boundary controls, and 70 invalid-overreach cases. None of the 70 received clean allow. The public checker counts and checks consistency of these outcomes; it does not replay the underlying authority decisions.

Boundary research

Governed Actions · sequence authority

Governed Agent Boundary Crossing

In the sealed 600-episode study, DeepSeek produced 76 unauthorized effects under prompt-only control and 58 cumulative-limit crossings under exact-action checks. FieldHash carried authority across the sequence: none occurred, while all 48 authorized-control episodes across the exact-action and sequence-aware governed arms completed.

Governed Actions · effect authority

Authority Under Effect Substitution

DeepSeek and Terra selected 37 operationally attractive substitutions across all eight tested families. Holding those choices constant, prompt-only and exact-action control each permitted 45 unauthorized effect executions. FieldHash effect authority permitted none, while all 96 authorized objectives completed.

Semantic-support boundary

Related Text Is Not Governing Evidence

The 2,340 rows derive from 110 source cases. FieldHash recorded 0 wrong clean allows on 682 clean-eligible rows versus 84 for the cross-encoder; 1,658 safety-routing rows have a separate denominator. These related rows do not establish a general deployment safety bound.

Authority-inference boundary

Where Simpler Controls Are Enough

Simulated oracle review using ground-truth chain order reached 300/334; a single-pass selector reached 312/334. The scale addendum reports stronger abstention-enabled baselines and a clean-status-filter tie. The original and post-run-amended rollback metrics remain separate.

Public-source ablations

Public Corpus Authority Series

Public RFC conflicts show enforcement when authority is configured, then expose what collapses when the authority signal is hidden.

Answer accounting

Live Answer Attribution

Across 1,000 live model answers, each output resolves to supported, contradicted, or unattributable rather than passing through unaccounted for.

The controls

These are control families used across the linked studies, not identical implementations in every experiment. Each study page retains its published labels, exact checks, and comparison boundaries. All comparators are experimental implementations, not benchmarks of commercial policy engines or gateways.

Prompt only

The model is instructed to respect the configured limits, without an external gate enforcing those limits. This provides an instructional baseline.

Request-local checks

Evaluate a proposed operation using the permissions and facts supplied for that request. In these experiments, those checks do not enforce the cumulative effect across the workflow. The exact checks vary by study.

Per-action authority

In the original organization study, this checks the owner’s objective, the actor’s identity and assignment, the delegation chain, and current authority at admission and delivery. It omits shared-budget reservation and cross-agent repeat protection. It is the strongest comparator we built for that study.

Approval-bound gate

The action study’s protected condition requires a signed human approval bound to the hash of the proposed action. Its study page defines the approval and execution boundary.

Incomplete mediation

Some consequential routes are governed while an enumerated route remains outside mediation. These conditions test coverage gaps; each study specifies the route left uncovered.

FieldHash effect authority

These conditions check configured effect relationships against the applicable authority. Sequence, substitution, after-denial, and execution-surface studies use different designs and published labels. Their study pages retain the exact definitions and limits.

FieldHash full mediation

In the organization study, the protected condition combines effect checks with shared-budget reservation, delivery-time currentness, and repeat protection. The separate repeat-protection follow-ups retain their own controls, denominators, and criteria.

How the published control labels relate

Exact action. This is a request-local control in the sequence, substitution, and after-denial studies, with different study-specific implementations. In the sequence study it already checks current authority for the selected exact effect, but does not carry cumulative trajectory state. Identity or current-authority checks alone do not establish the organization study's per-action authority contract.

Configured gate. In the action study this means a signed approval bound to the exact action. It is not an interchangeable name for effect authority across substitute routes.

Study-specific effect-authority labels. These labels refer to different protected conditions: continuing trajectory state, a matched effect-substitution evaluation, and coverage of the eight enumerated synthetic execution surfaces. Their scope remains attached to each study.

How to read the metrics+

Business failure count

A workflow outcome such as a revoked tool, stale record, or unapproved action reaching the governed path.

Model behavior count

What a model selected or attempted under the stated prompt, context, and model configuration.

Packet-contract verification

Whether generated evidence artifacts satisfy the published packet schema and accounting rules.

Artifact-tamper verification

Whether a deliberate mutation causes a verifier to reject the changed artifact or bundle. A recorded internal result is distinct from a mutation suite the public checker reruns.

Public outcome consistency

Counts, recorded dispositions, and self-hashes agree within a published projection. This does not independently replay the underlying authority decision.

Context reduction

How much candidate material was withheld from the model input under the tested policy.

Internal synthetic diagnostic

A FieldHash-authored mechanism test. Useful for engineering evidence, not customer validation.

Research archive

Supporting diagnostics remain public, without crowding the buyer path.

Internal stress tests, lifecycle diagnostics, context-compression results, ties, and negative controls remain available for technical scrutiny.

Browse research archive

FieldHash is architecturally independent of the systems it governs. Its current public diagnostics are self-administered; third-party validation is a separate milestone. Evidence sources, caveats, methods, and verification links live on the individual study pages.