Your agent can choose a version you revoked.
A prompt missing the relevant withdrawal selected the withdrawn package version in 540 of 540 trap evaluations. Governed Tool Use received that authority record separately and filtered the version out: zero unauthorized choices. The evaluations derive from 143 unique PyPI scenarios.
With the complete current authority log, prompting also reached zero. The stale-log comparison tests missing authority in model context; it does not test a live registry connector or actual package installation.
Selection-only diagnostic: the harness requested a JSON version choice. No packages were installed.
Unauthorized choices, full context
539/540
version choices across three models and three seeds
Prompt instruction, clean log
0/540
ties the gate when the authority log is current (published)
Prompt instruction, stale log
540/540
a one-step-old log breaks it completely
Unauthorized choices, gate
0/540
yanked candidates removed before selection; same 540 trap evaluations
Apply authority before selection, then retain the decision for review. The harness constructs a FieldHash Ledger packet showing which candidates were allowed and which were withheld.
The newest version can already be withdrawn.
A withdrawn release can still have the largest version number and a convincing description. The registry's yank flag carries the changed authority. If that signal never reaches the decision, choosing the newest-looking candidate can select a version the maintainer has withdrawn.
We evaluated 143 unique PyPI yank scenarios 540 times across three models and three seeds. In each trap, the withdrawn version is the newest one shown. The flags are retained in the corpus snapshot; later public API responses may differ if registry status changes. The full-context arm omits the yank signal, and all three frontier models selected the unauthorized version within one case of each other.
The traps deliberately expose this selection risk; their frequency is not an estimate of production incidence.
Remove the revoked version before the model sees it.
In this diagnostic, Governed Tool Use takes yanked versions out of the candidate set using the captured registry authority fields. The harness accepts only a response that parses to an offered candidate. Across three frontier models and three seeds, the gate recorded zero unauthorized choices in 540 trap evaluations, against 539 of 540 without governance.
0/540
Unauthorized version choices under the gate, on trap evaluations.
539/540
Unauthorized choices in full context, pooled across the same models and seeds.
0/538
Unauthorized version choices on 538 valid-version control evaluations, in every arm.
Results by model
| Model | Without governance | With the gate |
|---|---|---|
| Gemini 3.5 Flash | 180/180 | 0/180 |
| GPT-5.5 | 180/180 | 0/180 |
| Claude Opus 4.7 | 179/180 | 0/180 |
Unauthorized version choices by model (180 trap evaluations each), sampled from 143 unique scenarios. The gate and the complete current-log prompt each recorded zero in every run.
A complete authority prompt tied the gate.
We tested the obvious cheaper alternative: state the yanked versions in the prompt and instruct the model not to choose them. With the complete captured authority data included in each prompt, the models selected zero yanked versions, tied with the gate on all 540 trap evaluations.
The next condition tests what changes when the prompt omits the relevant revocation while the gate still receives it. It isolates where the authority check happens.
Keep the authority check outside the prompt.
The stale-log arm omitted this case's yank from the prompt. It selected the yanked version in all 540 trap evaluations. The gate used the captured registry authority fields separately and removed the yanked version from the candidate set. This isolates missing authority in model context; the diagnostic did not exercise a live registry connector or its freshness behavior.
Clean, current log
instruction ties the gate
Log one step out of date
one-step-stale log breaks instruction
A separate adversarial-instruction arm replaced the authority record and produced 535/540 unauthorized choices. That tests absent authority, not injection against a complete authority prompt. Candidate filtering applies the supplied authority before the model chooses.
Three models · three seeds
The result held on every run.
Nine runs: three frontier models (Gemini 3.5 Flash, GPT-5.5, Claude Opus 4.7) across three seeds, pooled to 1,078 paired evaluations (540 trap evaluations over 143 unique PyPI yank scenarios, and 538 valid-version controls; two control evaluations were dropped where a model response did not parse to a candidate). The gate recorded zero unauthorized choices in every run, as did the clean-log prompt arm. The stale-log arm selected the yanked version on every paired trap evaluation. The public stdlib verifier recomputes these choice counts from retained rows and fails on drift.
Study scope
What was tested.
- This is version-choice evidence. No package installation, installation-boundary enforcement, production effect, or installer bypass was executed or measured.
- Not an injection-defense claim. The gate filters candidates from authority data supplied outside the prompt; this study does not establish model hardening against adversarial prompts.
- The gate is only as trustworthy as its authority input. This harness uses a captured registry snapshot; it does not establish live-source freshness, protected write access, or a deployed installation control. Deployment tamper evidence and source protection require separate configuration.
- Not a claim about any other product. The comparison is enforcement against instruction, both FieldHash's own arms.
- The trap cases are constructed to place the revoked version as the newest shown. The yank itself is real and re-fetchable; the framing is disclosed.
- FieldHash enforces configured authority. It does not infer which version is authoritative from prompt text.
Local-model addendum
The gate does not ask the local model to remember the registry.
+
Local-model addendum
The gate does not ask the local model to remember the registry.
We reran the same PyPI authority arms through two local open-weight routes, Qwen 2.5 1.5B 4-bit and SmolLM2 360M 6-bit. The addendum is post-hoc portability evidence, not part of the pre-registered frontier run. It shows the same selection boundary: local models can pick a yanked candidate when it remains available. When the gate removes it first, the paired trap rows record zero unauthorized version choices. No package installation is measured.
The verifier counts only cases where all six arms parsed to a candidate. Qwen produced 120 paired cases; SmolLM2 produced 80 paired cases from the 120-case request. That is a disclosed local-model quality boundary, not hidden sample loss.
| Local route | Paired traps | Full context | Prompt, clean log | Stale log | Gate |
|---|---|---|---|---|---|
| Qwen 2.5 1.5B local | 60 | 49/60 | 2/60 | 48/60 | 0/60 |
| SmolLM2 360M local | 38 | 26/38 | 17/38 | 20/38 | 0/38 |
Counts are unauthorized version choices on latest-yank trap evaluations, recomputed from the local addendum bundle. Across both local routes, full context picked yanked versions in 75/98 paired traps, stale-log prompting in 68/98, and the gate in 0/98.
Reproduce the figures
Download the public verification bundle.
The bundle includes decision rows, public PyPI ground-truth references, the disclosed task harness, checksums, and a stdlib verifier. It excludes raw model responses, provider routing, keys, and internal enforcement code.
The public figure checker recomputes retained choice counts. It does not replay installation, verify actual model-request bytes, or establish a deployed enforcement path.
The other half of governed actions
Part II: whether a consequential action runs at all.
This study measures package-version choices. Its companion studies consequential action choices — delete, deploy, revoke, transfer. Across 270 evaluations of 90 authored decision points, models chose the consequential action 226 times under a prompt encouraging decisive action, and 8 times when told to ask first. The companion page separates those publicly recomputable choices from its retained governed-enforcement, attribution, and tamper reports.
Govern what your agents are allowed to do.
Bring one workflow where tool availability or approval can change. In a six-week shadow evaluation, compare your current controls with FieldHash decisions and review the evidence before enabling enforcement.
Bring one workflow