Six-week pilot

Start with one workflow worth governing.

Choose one live or near-live workflow where the wrong record, tool, action, or reviewed decision would create rework, risk, or reviewer uncertainty. FieldHash runs beside it in shadow mode for six weeks. Production answers and actions stay unchanged.

Your reviewers see what would have proceeded, what would have stopped, and what would have gone to review. At closeout, they decide whether to enforce, adjust the authority rules, remain in shadow mode, or stop. For computer-use agents, clicks, commands, file edits, and form submissions can be evaluated when represented as action packets.

At closeout, you will have a measured false-clean-allow count and observed rate across the agreed evaluation sample, review-route volume, median review time, added latency, deployment effort, and the evidence needed to decide whether the control is worth enforcing.

First-workflow fit

Start where the wrong handoff has a clear consequence.

The strongest first pilot is bounded, replayable, and owned. It does not need to be your largest workflow or begin with perfect metadata.

A strong first workflow

The workflow is already running, in late-stage testing, or has replayable traces.

A wrong record, tool, action, or reused review decision creates a concrete review or operating consequence.

A system or reviewer can establish what is current, approved, revoked, superseded, or out of scope.

One accountable owner can scope the workflow and agree on the pilot acceptance criteria.

The first pilot team

Workflow owner: the AI platform, product, or operations lead who knows the current path.

Authority owner: the person who owns the policy, record system, tool registry, or approval process.

Reviewer: security, model risk, legal, compliance, or another accountable control function.

If the authority signals are incomplete, bring the workflow and the reviewer who can establish them. The first scoping step maps what already exists and what still requires review.

Example first workflow

Keep policy answers current. Let approved exceptions expire.

Connect one policy source and its review process. FieldHash tests whether the current policy governs each answer, whether an approved exception still applies, and whether expired, rejected, or superseded authority stays out. Reviewers receive evidence showing which authority governed and why.

Source-connected pilot path

Test what changes after the first handoff.

Where the workflow can expose authority through a read-only SQL projection or signed event feed, the pilot can test more than a prepared snapshot.

FieldHash can materialize current authority state, observe source changes in shadow mode, show which workflows and precedents those changes affect, and send invalid reuse back to review without changing production behavior during the evaluation.

Managed SaaS connectors remain scoped with the design partner.

MCP tool path

Test the tool boundary without replacing the MCP server.

For an MCP workflow, FieldHash runs between the host and existing server over a versioned stdio or stateless HTTP profile. The pilot records which tools, resources, prompts, and server instructions would remain available, which exact calls would stop or route to review, and what evidence each decision leaves.

The 2026-07-28 profile can also test governed multi-round input, scope-bound subscriptions, cache boundaries, and a bounded, opt-in Tasks profile. Subscription lineage is durable, but disconnected clients must re-establish the listen request. Transport authentication, upstream OAuth, server permissions, and MCP Apps host containment remain in the existing stack.

Computer-use path

Test structured browser actions before dispatch.

For a Playwright-compatible workflow, the pilot can govern navigation, clicks, form input, key presses, selection, checkbox changes, and file upload as structured action packets. Consequential steps can require an approval bound to the exact action and a verified Ledger record before execution.

FieldHash does not read pixels or replace browser authentication, session custody, target allowlists, or the host runtime.

Not a fit when+

There is no accountable workflow owner or reviewer who can establish what should govern.

The workflow is neither live nor reproducible from traces, so the current handoff cannot be compared.

A wrong handoff has no meaningful operating, review, or compliance consequence.

The goal is a universal compliance certification rather than a scoped control test.

Commercial path

Pilot first. Enforce where risk is real.

Step 1

Scoping

Map the workflow, authority sources, and review boundary.

Step 2

Baseline capture

Capture the context, prompts, and actions already in use.

Step 3

Path comparison

Compare current outcomes with FieldHash shadow decisions.

Step 4

Verification report

Review allows, blocks, review routes, and packet integrity.

Step 5

Control handover

Move accepted policies into live enforcement where warranted.

Evaluate before enforcement

Shadow-mode pilot

Compare one current workflow against FieldHash decisions while production stays unchanged.

Enforce where risk is real

Production enforcement

Move accepted policies into enforcement. Approved context and actions proceed; rejected alternatives stay out.

Stronger signing, customer-owned logs, external anchoring, and SIEM or GRC export remain deployment-scoped rather than part of the default pilot path.

What the pilot includes

One workflow, one evidence trail.

One scoped AI agent, RAG workflow, enterprise memory source, tool workflow, MCP stdio or stateless HTTP server, or instrumented computer-use path

Baseline capture of the current retrieval, memory, or action handoff

Shadow-mode comparison showing what FieldHash would allow, block, caveat, or route to review

Controlled enforcement simulation showing excluded records or actions removed from the packet

Governed Precedent review where prior reviewer decisions may safely carry forward

Governed Learning lineage from the initial packet through review, Authority State, reuse, suspension, and return to review

Source-health and freshness report for connected authority sources where configured

Change-impact report showing affected workflows, precedents, and review routes

Reviewer-ready evidence packet, scorecard, closeout findings, and verification commands where configured

A production-scope recommendation: enforce, revise the authority inputs, continue in shadow mode, or stop

Deployment topology, authentication, persistence, key custody, retention, regional hosting, monitoring, and incident response are reviewed separately and recorded in the SOW.

Review security and deployment

Operator console included

Your reviewers see the governed path before anything changes.

During shadow mode, the authenticated console shows current Authority State, source health, unresolved bindings, review work, active and suspended precedents, evidence hashes, and the workflows affected by a source change. Reviewers can inspect what would proceed, what would stay out, and what still needs a decision before they recommend enforcement.

The console exposes client-safe authority state, not raw customer content. Production access requires an authenticated admin operator.

Open the sample console
Mobile FieldHash operator console showing a case withheld for review

Illustrative client-safe sample. Live pilot data remains subject to the agreed access and retention controls.

Pilot economics

Measure the queue before enforcing the gate.

Fail-closed routing is useful only when the review burden is tolerable. The pilot measures safety and reviewer load together. Counts are normalized per 1,000 handoffs where volume permits; no public diagnostic substitutes for the customer result.

01

Unsafe influence

False clean allows and stale, revoked, or out-of-scope influence that would have reached the governed path.

02

Review load

Cases routed to review, separated by missing authority, ambiguity, drift, and infrastructure failure.

03

Reviewer effort

Median review time and the share of routed cases resolved without source-system cleanup.

04

Review-to-state conversion

Reviewed cases that produced a valid Authority State update, resolved binding, revocation, narrowed scope, or bounded precedent.

05

Precedent eligibility

Review outcomes eligible for reuse under the workflow's authority policy versus one-off decisions, plus later cases resolved through each approved precedent.

06

Review reduction

Repeat reviews avoided through valid reuse, with stale, expired, revoked, conflicting, or unsupported reuse counted separately.

07

Second-workflow reuse

Authority objects, source bindings, and precedents reused by a second workflow, plus the change in setup and reviewer effort.

08

Over-blocking

Approved records or actions withheld, with reviewer disposition and the recorded reason.

09

Operating cost

Added latency, deployment effort, and whether the customer considers the control worth enforcing.

Review the pilot acceptance categories+

The workflow owner and reviewers set the thresholds before testing. These categories keep the closeout decision tied to the risk and operating burden of the selected workflow.

MeasureQuestionAcceptance rule
Governing-source accuracyDid the packet use the reviewer-approved source of authority for each evaluated case?Set in the SOW. Critical workflows usually require no unresolved governing-source errors.
Blocked influenceWere stale, rejected, revoked, rolled-back, or out-of-scope items kept out of the live path?Set before testing, with a separate ceiling for critical misses.
Failure behaviorWhen authority was missing or conflicting, did the case stop or go to review?No clean allow for missing authority in configured high-risk scopes.
Packet completenessCould reviewers see what proceeded, what stopped, the reason, and the authority source?Required packet fields present or documented as out of scope before testing.
Precedent reuseDid a reused decision remain inside its approved scope, evidence, policy, expiry, and dependencies?No clean reuse after expiry, drift, conflict, revocation, or missing evidence.
Review-to-state conversionWhich reviewed outcomes became valid external Authority State, and which correctly remained one-off?Every state update traces to an authorized review outcome; no model output or score grants authority by itself.
Workflow compoundingDid a second workflow reuse existing Authority State or precedent without inheriting scope that did not belong to it?Shared authority is reused where valid, with no workflow-local scope leakage in the agreed test set.
Dependency precisionDid a source change suspend every precedent with a recorded dependency without disrupting unrelated workflows?No missed recorded dependencies or unrelated suspensions in the agreed test set.
MCP enforcementWere revoked tools and unsupported influence removed, interactive lineage preserved, and unapproved exact calls stopped before upstream dispatch?No configured blocked call or content reaches the upstream or model-facing path in the agreed test set.
Computer-use enforcementDid every evaluated browser step match its policy, exact approval, page boundary, and evidence record before dispatch?No configured blocked or approval-gated action executes in the agreed structured-action test set.
Operational fitDid the handoff meet the agreed latency, retention, export, and reviewer-time needs?Targets set by workflow and deployment path before the pilot starts.

What we need from you

Bring one workflow and the systems that already define authority.

One workflow where stale, superseded, rejected, or revoked influence creates review risk

Candidate source systems: RAG store, policy system, CLM, GRC, registry, tool catalog, or document platform

Authority signals available today: approval state, version history, effective dates, revocation records, supersession links, or reviewer ownership

Change signals available today: connector updates, schema changes, owner changes, policy hierarchy updates, approval workflow changes, or source-state heartbeats

A small set of example questions, actions, MCP tool calls, computer-use traces, or workflow steps that matter to the evaluating team

The reviewer group that will inspect packet completeness, false allows, false blocks, and production-scope fit

Bring the workflow. We will map the authority signals.

If the signals already live in CLM, policy, GRC, document, registry, or workflow systems, FieldHash can test how they govern the answer, tool, or computer-use action path. If they don't exist yet, we start with propose-then-approve metadata before enforcement.

Scope, fees, deployment posture, support commitments, retention terms, and conversion terms are defined in the applicable order form or statement of work.