Start with one workflow worth governing.
Test whether an earlier approval, governing record, or reviewed exception still applies. FieldHash runs beside one consequential workflow for six weeks in shadow mode. Production answers and actions stay unchanged.
Build the case for governing one consequential workflow. Record what FieldHash would allow, withhold, or send to review. Production enforcement remains off. Measure where valid review can carry forward and what the control takes to operate. Your reviewers receive the evidence to decide whether to enforce, adjust the authority rules, remain in shadow mode, or stop.
At closeout, you will have the observed false-allow rate across the agreed sample, review volume, median review time, added latency, deployment effort, and the evidence needed to decide whether the control is worth enforcing.
First-workflow fit
Start where the wrong handoff has a clear consequence.
Start with one consequential action or decision, the policy, approval, or reviewed decision that authorizes it, and what can change before that authority is used. The strongest first pilot is bounded, replayable, and owned. It does not need to be your largest workflow or begin with perfect metadata. Evaluate changing approvals, reviewed-decision reuse, instrumented action paths, or per-workflow cumulative limits where configured.
Current pilot scope: The shipped kernel supports per-workflow cumulative budgets. The shared-budget runtime evaluated across cooperating agents is a qualified research candidate, not part of the current pilot deployment. Cross-agent shared-budget requirements need separate scoping.
A strong first workflow
The workflow is already running, in late-stage testing, or has replayable traces.
A wrong record, tool, action, or reused review decision creates a concrete review or operating consequence.
A system or reviewer can establish what is current, approved, revoked, superseded, or out of scope.
One accountable owner can scope the workflow and agree on the pilot acceptance criteria.
The first pilot team
Workflow owner: the AI platform, product, or operations lead who knows the current path.
Authority owner: the person who owns the policy, record system, tool registry, or approval process.
Reviewer: security, model risk, legal, compliance, or another accountable control function.
If the authority signals are incomplete, bring the workflow and the reviewer who can establish them. The first scoping step maps what already exists and what still requires review.
Example first workflow
Keep policy answers current. Let approved exceptions expire.
Connect one policy source and its review process. FieldHash tests whether the current policy governs each answer, whether an approved exception still applies, and whether expired, rejected, or superseded authority stays out. Reviewers receive evidence showing which authority governed and why.
Source-connected pilot path
Test what changes after the first handoff.
Where the workflow can expose authority through a read-only SQL projection or signed event feed, the pilot can test more than a prepared snapshot.
See which workflows and reviewed decisions depend on a changed source. In shadow mode, FieldHash records where that change would stop reuse or return a case to review while production answers and actions remain unchanged.
Managed SaaS connectors remain scoped with the design partner.
MCP tool path
Test the tool boundary without replacing the MCP server.
For an MCP workflow, FieldHash runs between the host and existing server over a versioned stdio or stateless HTTP profile. The pilot records which tools, resources, prompts, and server instructions would remain available, which exact calls would stop or route to review, and what evidence each decision leaves.
The pilot can also test multi-round input, subscriptions, deferred tasks, and cache boundaries when those MCP features are in scope. Transport authentication, upstream permissions, and host containment remain part of the existing stack.
Computer-use path
Test structured browser actions before dispatch.
For a Playwright-compatible workflow, the pilot can govern navigation, clicks, form input, key presses, selection, checkbox changes, and file upload as structured action packets. Consequential steps can require an approval bound to the exact action and a verified Ledger record before execution.
FieldHash does not read pixels or replace browser authentication, session custody, target allowlists, or the host runtime.
Not a fit when+
There is no accountable workflow owner or reviewer who can establish what should govern.
The workflow is neither live nor reproducible from traces, so the current handoff cannot be compared.
A wrong handoff has no meaningful operating, review, or compliance consequence.
The goal is a universal compliance certification rather than a scoped control test.
Commercial path
Pilot first. Enforce where risk is real.
Step 1
Scoping
Map the workflow, authority sources, and review boundary.
Step 2
Baseline capture
Capture the context, prompts, and actions already in use.
Step 3
Path comparison
Compare current outcomes with FieldHash shadow decisions.
Step 4
Verification report
Review allows, blocks, review routes, and packet integrity.
Step 5
Control handover
Recommend the next scope; enable enforcement only after separate acceptance.
Evaluate before enforcement
Shadow-mode pilot
Compare one current workflow against FieldHash decisions while production stays unchanged.
Enforce where risk is real
Production enforcement
Move accepted policies into enforcement. Approved context and actions proceed; rejected alternatives stay out.
Stronger signing, customer-owned logs, external anchoring, and SIEM or GRC export remain deployment-scoped rather than part of the default pilot path.
What the pilot includes
One workflow, one evidence trail.
The single-tenant reference deployment provides an authenticated gate, private operator console, durable review queue, signed ledger, and backup-and-restore tooling. The SOW records which components the pilot uses, their access controls, and the recovery checks required before acceptance.
One scoped AI agent, RAG workflow, enterprise memory source, tool workflow, MCP stdio or stateless HTTP server, or instrumented computer-use path
Baseline capture of the current retrieval, memory, or action handoff
Shadow-mode comparison showing what would proceed, stop, or go to review
Controlled enforcement simulation showing excluded records or actions removed from the packet
Governed Precedent review where prior reviewer decisions may safely carry forward
Governed Learning lineage from the initial packet through review, Authority State, reuse, suspension, and return to review
Source-health and freshness report for connected authority sources where configured
Change-impact report showing affected workflows, precedents, and review routes
Reviewer-ready evidence packet, scorecard, closeout findings, and verification commands where configured
A production-scope recommendation: enforce, revise the authority inputs, continue in shadow mode, or stop
Deployment topology, authentication, persistence, key custody, retention, regional hosting, monitoring, and incident response are reviewed separately and recorded in the SOW.
Review security and deploymentReference operator console
Your reviewers see the governed path before anything changes.
When included in the pilot SOW, the authenticated console shows current Authority State, source health, unresolved bindings, review work, active and suspended precedents, evidence hashes, and the workflows affected by a source change. Reviewers can inspect what would proceed, what would stay out, and what still needs a decision before they recommend enforcement.
The console exposes client-safe authority state, not raw customer content. Production access requires an authenticated admin operator.
Open the sample console

Illustrative client-safe sample. Live pilot data remains subject to the agreed access and retention controls.
Pilot economics
Measure the value and the effort together.
Put a number on invalid decisions caught, valid work preserved, review effort, and operating cost in your workflow. Agree the measures before testing, then use the results to decide whether to enforce. Counts are normalized per 1,000 handoffs where volume permits.
A shadow disagreement is a case for reviewers to investigate, not a prevented incident. Controlled enforcement testing must also examine what happens after a denial, including whether the agent changes its plan. Production enforcement requires a separate acceptance decision.
01
Unsafe influence
False clean allows and stale, revoked, or out-of-scope influence that would have reached the governed path.
02
Review load
Cases routed to review, separated by missing authority, ambiguity, drift, and infrastructure failure.
03
Reviewer effort
Median review time and the share of routed cases resolved without source-system cleanup.
04
Review-to-state conversion
Reviewed cases that produced a valid Authority State update, resolved binding, revocation, narrowed scope, or bounded precedent.
05
Precedent eligibility
Review outcomes eligible for reuse under the workflow's authority policy versus one-off decisions, plus later cases resolved through each approved precedent.
06
Review reduction
Repeat reviews avoided through valid reuse, with stale, expired, revoked, conflicting, or unsupported reuse counted separately.
07
Optional second-workflow extension
If separately scoped, measure authority objects, source bindings, and precedents reused by a second workflow, plus setup and reviewer effort.
08
Over-blocking
Approved records or actions withheld, with reviewer disposition and the recorded reason.
09
Operating cost
Added latency, deployment effort, and whether the customer considers the control worth enforcing.
Review the pilot acceptance categories+
The workflow owner and reviewers set the thresholds before testing. These categories keep the closeout decision tied to the risk and operating burden of the selected workflow.
| Measure | Question | Acceptance rule |
|---|---|---|
| Governing-source accuracy | Did the packet use the reviewer-approved source of authority for each evaluated case? | Set in the SOW. Critical workflows usually require no unresolved governing-source errors. |
| Blocked influence | Were stale, rejected, revoked, rolled-back, or out-of-scope items kept out of the live path? | Set before testing, with a separate ceiling for critical misses. |
| Failure behavior | When authority was missing or conflicting, did the case stop or go to review? | No clean allow for missing authority in configured high-risk scopes. |
| Packet completeness | Could reviewers see what proceeded, what stopped, the reason, and the authority source? | Required packet fields present or documented as out of scope before testing. |
| Precedent reuse | Did a reused decision remain inside its approved scope, evidence, policy, expiry, and dependencies? | No clean reuse after expiry, drift, conflict, revocation, or missing evidence. |
| Review-to-state conversion | Which reviewed outcomes became valid external Authority State, and which correctly remained one-off? | Every state update traces to an authorized review outcome; no model output or score grants authority by itself. |
| Optional second-workflow extension | If separately scoped, did a second workflow reuse existing Authority State or precedent without inheriting scope that did not belong to it? | Shared authority is reused where valid, with no workflow-local scope leakage in the agreed test set. |
| Dependency precision | Did a source change suspend every precedent with a recorded dependency without disrupting unrelated workflows? | No missed recorded dependencies or unrelated suspensions in the agreed test set. |
| MCP enforcement | Were revoked tools and unsupported influence removed, interactive lineage preserved, and unapproved exact calls stopped before upstream dispatch? | No configured blocked call or content reaches the upstream or model-facing path in the agreed test set. |
| Computer-use enforcement | Did every evaluated browser step match its policy, exact approval, page boundary, and evidence record before dispatch? | No configured blocked or approval-gated action executes in the agreed structured-action test set. |
| Operational fit | Did the handoff meet the agreed latency, retention, export, and reviewer-time needs? | Targets set by workflow and deployment path before the pilot starts. |
What we need from you
Bring one workflow and the systems that already define authority.
One workflow where stale, superseded, rejected, or revoked influence creates review risk
Candidate source systems: RAG store, policy system, CLM, GRC, registry, tool catalog, or document platform
Authority signals available today: approval state, version history, effective dates, revocation records, supersession links, or reviewer ownership
Change signals available today: connector updates, schema changes, owner changes, policy hierarchy updates, approval workflow changes, or source-state heartbeats
A small set of example questions, actions, MCP tool calls, computer-use traces, or workflow steps that matter to the evaluating team
The reviewer group that will inspect packet completeness, false allows, false blocks, and production-scope fit
Bring the workflow. We will map the authority signals.
If the signals already live in CLM, policy, GRC, document, registry, or workflow systems, FieldHash can test how they govern the answer, tool, or computer-use action path. If they don't exist yet, we start with propose-then-approve metadata before enforcement.
Scope, fees, deployment posture, support commitments, retention terms, and conversion terms are defined in the applicable order form or statement of work.