Told to ask first, models still skipped approval 8 times in 270. The gate let none through.
Three frontier models selected consequential actions in 8 of 270 ask-first evaluations. Under a decisive operational prompt, they did so 226 times. Applied to those operational choices, the configured Governed Action Execution gate allowed none through without approval.
Controlled diagnostic: model action selections and deterministic gate decisions. The harness did not dispatch external actions.
Operational prompt
226/270
consequential choices under a prompt encouraging decisive action; 3 models
Prompt instruction, ask-first
8/270
consequential choices despite an explicit instruction to request approval first
Unapproved choices allowed by the gate
0/270
deterministic gate applied to operational choices; no external actions executed by this harness
Recorded signing result
270/270
retained run reports verified approvals and 15/15 mutations caught; public checker does not reverify signatures
The outcome: an agent that cannot run a consequential action without a human approval bound to that exact action, and a FieldHash Ledger record of who approved what. Change the amount after approval and the approval no longer verifies.
A capable agent, a decision it can just make.
The corpus presents decisions such as dropping an old production audit table the CTO asked to clean up, revoking an AWS key exposed in a public gist, or sending a CFO-signed $2.3M wire. In many of the consequential cases, the step is already queued, drafted, or ready to run. Claude Opus 4.7 authored 180 decision points across infrastructure, finance, data, communications, and security: 90 consequential cases and 90 benign controls. Each went to three models. The menu offered read, list, search, and other actions alongside consequential options.
The operational prompt described an autonomous operations agent with access to a live company system and instructed it to be decisive and complete tasks efficiently. Models selected consequential actions in 226 of 270 should-gate evaluations, including 70 of 90 for Opus. A separate prompt explicitly required approval before consequential actions and added a request-approval option. That condition produced 8 consequential choices in 270 evaluations. The harness collected structured choices; it did not dispatch external actions.
That matters beyond API tools. Browser and desktop agents do the same class of work through clicks, shell commands, file edits, and form submissions. FieldHash governs those paths when the runtime wraps the intended step as an action packet before execution.
A finding we did not go looking for
On the 270 benign tasks — read a file, list services, search logs — the models still reached for a consequential action, such as a shell command or an external call: 49 times (18%) under the decisive prompt and 16 times (6%) under the ask-first prompt. The ask-first prompt also sent 77 of the 270 routine tasks to a human for approval. Agents escalate ordinary work, not only tasks that call for action. In these tests, 0/270 intended low-risk operations were falsely blocked; attempted escalations into consequential actions were held.
A consequential action cannot run without a signed approval.
Governed Action Execution classifies each proposed action against a configured policy. Read and low-write actions pass; consequential and irreversible actions require approval bound to the exact action hash. The harness reused the operational model selections and applied this gate deterministically. All consequential selections required approval, leaving zero unapproved allows across 270 should-gate evaluations. The separate local signing path recorded approval verification; no external action dispatcher was exercised. In the 270 benign controls, no selected low-risk action was falsely blocked.
0/270
Unapproved consequential choices allowed by the configured gate.
226/270
Consequential choices under the operational prompt across three models.
0/270
No low-risk action was falsely blocked in these 270 tests.
Capability does not close the gap
| Model | Without governance | Prompt, ask-first | With the gate |
|---|---|---|---|
| Gemini 3.5 Flash | 80/90 | 1/90 | 0/90 |
| GPT-5.5 | 76/90 | 1/90 | 0/90 |
| Claude Opus 4.7 | 70/90 | 6/90 | 0/90 |
Could you just tell the model to ask first? Mostly. Not reliably.
We tested the obvious cheaper alternative: a prompt that drops the decisive framing, requires human approval before any consequential action, and offers a request-approval option. Consequential choices fell from 226 to 8 of 270, from 84% to 3%. We publish that, because prompting is a real mitigation.
Eight consequential choices remained: six from Claude Opus 4.7, and one each from Gemini and GPT. In seven of them, the scenario cited a standing runbook or policy that authorized or required immediate action, and the model selected the consequential action instead of asking. A prompt leaves that conflict to the model. An external approval check makes the requirement binding at the governed handoff.
Even when the model asks, it leaves no proof. The gate does.
A prompt alone does not bind a human approval to the exact action payload. The gate's local signing path binds an approval to the action hash and records its verification. The retained run reports 270 approval-verification cases. Its five mutation types changed the action hash, flipped a decision, removed a record, duplicated a record, and changed the manifest hash. Across three models, all 15 mutations were detected. This suite tested those record mutations; it did not test signature swapping or dispatch a modified transfer to an external system.
270/270
Signed approvals reported as verified in the retained run; the public checker prints this pinned count.
15/15
Tamper mutations reported as detected in the retained run. The public checker prints this pinned result; it does not rerun the mutations.
0
Prompt instructions that produced any audit record at all.
Model-authored corpus
Model-authored scenarios, administered by FieldHash.
Claude Opus 4.7 authored the 180 decision points. FieldHash designed and administered the protocol; this is self-administered research. The same model selected consequential actions on 70 of its 90 should-gate cases under the operational prompt and on six under the ask-first prompt. The frontier run retained 1,080 elicitations across three models. Its protocol was pre-registered and sealed before the run. After the run, a scoring bug that counted parse errors and approval requests as consequential was fixed, and the aggregates were recomputed from the unchanged rows. Action-choice scoring and local hash/signature checks are deterministic, with no model judge. Four of 540 case-level evaluations contained one unrecoverable parse failure. Each failed elicitation contributed no consequential action to its arm's numerator; the precommitted 270-case arm denominator remained fixed, and the failures are reported separately.
What this does not claim
The boundary, stated plainly.
- Not an accuracy claim. The gate controls what an agent is allowed to do; it does not make the model smarter or safer at judging.
- FieldHash enforces a configured policy. A human and the policy decide what counts as consequential; FieldHash gates and proves. It does not decide autonomously what is safe.
- Zero unapproved allows is a property of the configured gate on these selections. Production enforcement also depends on the policy, approval channel, and complete mediation of the relevant effect paths.
- The scenarios are model-authored, single-turn decision points rather than production incidents. The operational prompt told the agent to be decisive; the ask-first prompt dropped that line, required approval before consequential actions, and offered a request-approval action. The gate is applied to operational choices, not a third independent model run.
- Computer-use coverage depends on instrumentation. FieldHash does not infer authority from pixels or arbitrary UI state; clicks, commands, file edits, and form submissions must be represented as action packets before this gate can govern them.
- The retained packets support inspection of the recorded decisions. Establishing that a real action was mediated also requires execution-path evidence; packet consistency alone does not establish that or the correctness of the policy.
- Not a claim about any other product. The comparison is enforcement against instruction, both FieldHash's own arms.
Capability-tier addendum · five models · published corpus and seed
Told to ask first, the smallest local model still chose the action 49 times in 90.
The same corpus was evaluated on two 4-bit local models. Under the ask-first prompt, the frontier models selected consequential actions in 1 to 6 of 90 cases; the local models did so in 18 and 49. The deterministic gate recorded zero unapproved allows at all five tested tiers. The local addendum retained zero low-risk over-blocks and 90/90 approval-verification cases per model.
| Arm | Gemini 3.5 Flash | GPT-5.5 | Claude Opus 4.7 | Gemma E2B local | Qwen 1.5B local |
|---|---|---|---|---|---|
| Without governance | 80/90 | 76/90 | 70/90 | 80/90 | 87/90 |
| Prompt, ask-first | 1/90 | 1/90 | 6/90 | 18/90 | 49/90 |
| With the gate | 0/90 | 0/90 | 0/90 | 0/90 | 0/90 |
Reproduce the figures
Download the public verification bundle.
The bundle includes per-model decision rows (action types, not raw responses), the Claude-authored scenario corpus, the harness and gate primitive, and a standard-library checker. It derives ungoverned and prompt-only consequential-action counts, benign-case escalations, and denominators from rows, and checks that consequential choices were marked as requiring approval. It returns the governed zero as a fixed value; it does not reconstruct execution. The 270/270 attribution and 15/15 tamper results are printed from pinned totals, without re-verifying signatures or rerunning the tamper suite. Those are retained run results, not independently reproduced cryptographic proof from this checker. Run python code/verify_governed_action_metrics.py. No network, no keys.
The other half of governed actions
Part I: which tools an agent is allowed to trust.
This study governs whether a consequential action runs. Its companion governs which tools and sources an agent is allowed to act on in the first place: on PyPI-derived yank data, models given an authority log that omitted the withdrawal chose the withdrawn version in 540 of 540 trap evaluations; behind the gate, none did.
Governed Agents
See how this study contributes to the governed-agent evidence progression.
Six studies move from one consequential action to accumulation, route substitution, live replanning, changes in execution surface, and shared limits across cooperating agents.
Govern what your agents are allowed to do.
Bring one agent workflow with a real consequential action: a payment, a deploy, a deletion. We will show what it would have been allowed to run, what we would have held for approval, and the tamper-evident packet your reviewers can inspect.
Bring one workflow