Give the model a smaller tool menu to work with.

Across Gemini, GPT, and Claude, a five-tool FieldHash budget cut the visible tool menu from 27.53 to 4.74 mean tools, about 83% fewer, and kept the required function in all 240 target-tool cases. Across the three models combined, FieldHash scored one to five cases better than the full-tool and prompt-smart arms on each reported function-call metric.

Supporting diagnostic: paired tests found neither a significant quality lift nor equivalence. In a Gemini 3.5 Flash follow-up, a matched top-k retriever with the same five-tool budget exposed the same number of tools and tied FieldHash on function quality.

Tool surface reduction

~83%

4.74 vs 27.53 mean tools, less prompt to audit

Gold-function retention

240/240

target-tool cases retained the required function

Function-name correctness

343/360

vs 339/360 full, 341/360 prompt; paired p >= 0.219

AST-like correctness

335/360

vs 330/360 full, 334/360 prompt; paired p >= 0.180

False calls on irrelevance

14/120

vs 17/120 full, 16/120 prompt; paired p >= 0.375

Narrow the options before the model chooses, then retain what was included and excluded for review. This diagnostic tests that context-control step with the same underlying models.

Failure mode

Tool calling fails when the relevant tool is surrounded by plausible neighbors.

Modern agents often expose the base model to a large menu of tools, schemas, and distractors. That improves capability coverage, but it also increases the chance of irrelevant calls, argument confusion, and unnecessary prompt load.

This benchmark tests a narrower question: can a governed routing layer reduce what reaches the model without hiding the function the task actually requires?

Visible tool surface

Full-tool baseline

27.53 mean tools

Prompt-smart baseline

27.53 mean tools

FieldHash filtered path

4.74 mean tools

The measured result is a much smaller tool surface, the required function retained in every target-tool case, and function-call scores were one to five cases better than the full-tool and prompt-smart arms, without a significant paired difference. A Gemini 3.5 Flash matched top-k comparator tied FieldHash on function-name and AST-like quality, so differentiated efficiency is not claimed.

Results

Three provider paths, same control-plane pattern.

The run used BFCL V4 simple, multiple-function, and irrelevance task families across Gemini 3.1 Flash Lite, GPT-5.5, and Claude Opus 4.7. Each arm saw the same task rows; the difference was how much visible tool context was sent to the model.

Function-call quality

AST-like correctness: whether the selected function and structured arguments match the expected call pattern. The one-to-five row deltas below are observed differences; the study established neither a significant quality lift nor equivalence.

Full-tool baseline

330/360 AST-like

Prompt-smart baseline

334/360 AST-like

FieldHash filtered path

335/360 AST-like

False tool calls on irrelevance

Full-tool baseline

17/120 irrelevance cases (14.2%)

Prompt-smart baseline

16/120 irrelevance cases (13.3%)

FieldHash filtered path

14/120 irrelevance cases (11.7%)

0 cases120 cases (100%)

The denominator is 120 irrelevance evaluations (40 cases across three models); the other 240 evaluations require a target tool. Lower is better here. The effect is modest and directional only; the paired no-call comparison did not establish significant separation.

Provider replication

The aggregate is not hiding a single-provider result.

Each provider saw the same 120 BFCL-derived rows, including 40 irrelevance cases, and the same three arms. Values are shown as full-tool / prompt-smart / FieldHash filtered, except visible tools, which shows full-tool to filtered.

Gemini 3.1 Flash Lite

Visible tools27.53 → 4.74
Function name113 / 114 / 114
AST-like111 / 111 / 111
False calls5 / 5 / 5

GPT-5.5

Visible tools27.53 → 4.74
Function name115 / 114 / 116
AST-like110 / 112 / 113
False calls4 / 5 / 3

Claude Opus 4.7

Visible tools27.53 → 4.74
Function name111 / 113 / 113
AST-like109 / 111 / 111
False calls8 / 6 / 6

Audit checks

The compression result is checked against the raw run artifacts.

Because compression can look good by hiding necessary context, the review checks focus on whether the filtered path kept the required tool, survived a different distractor shuffle, and matches the underlying run records.

Retention is measured under distractor sets that exclude near-equivalent tools, so exact-function scoring stays unambiguous. The result tests clutter reduction without hiding the required function; it does not claim robustness to every possible tool synonym or duplicate schema.

The close function-calling differences do not establish equivalence. Exact paired McNemar checks did not show significant lift: function-name p=0.219 versus full-tool and p=0.625 versus prompt-smart; AST-like p=0.180 and p=1.000; no-call on irrelevance p=0.375 and p=0.625. A follow-up Gemini 3.5 Flash matched top-k arm tied FieldHash on function-name and AST-like totals. The core result is context-hygiene instrumentation rather than a public superiority claim.

Retention check

240/240

Every target-tool case in the three-provider aggregate kept the required gold function visible to the model.

Seed replication

80/80

A second Gemini 3.1 Flash Lite run with a different distractor seed retained every required function and preserved the compression pattern.

Paired quality check

no lift

Exact paired tests did not find significant separation from full-tool or prompt-smart arms; a Gemini 3.5 Flash matched top-k comparator tied FieldHash on function quality.

Consistency gate

18 tests

The tool benchmark test suite recomputes the public numbers from the raw run artifacts, including retention and seed-replication checks.

What this supports

Context control is useful infrastructure.

Governed Actions can shape the tool context before generation and record the selection for review. This diagnostic demonstrates that step; it does not establish an advantage over commodity top-k retrieval or improved model generation.

In practice, that can mean fewer irrelevant tools, shorter prompts, and a cleaner decision surface for the same underlying model.

Reviewer boundaries

How to read the result.

BFCL-derived internal distractor-pressure diagnostic, not an official BFCL leaderboard score. Distractor sets exclude near-equivalent tools to keep exact-function scoring unambiguous, so gold-function retention should not be read as a differentiated selection moat. Exact paired checks did not establish a quality lift over full-tool or prompt-smart arms. A first matched top-k comparator on Gemini 3.5 Flash tied FieldHash on function-name and AST-like quality. The result supports context-hygiene instrumentation, not a public claim that FieldHash beats commodity tool retrieval.

The result remains live as a supporting product-architecture diagnostic because it verifies context narrowing, gold-function retention, and auditability. It should not be read as a tool-retrieval superiority benchmark: the first matched top-k comparator tied FieldHash on function quality. The main tool-authority claim lives in the revoked-action diagnostic.

Evaluate your agent's tool handoff.

Bring one agent workflow and compare its current tool handoff with FieldHash in shadow mode. Measure what reaches the model, what goes to review, and what the control takes to operate before deciding whether to enforce.

Download methods report