Pursuit ReplayPricingLearnWorkbenchAboutThe SprintResultsFAQBook a Pursuit Replay →
‹ Learn · Measurement

Your AI pilot is cheating.

Not maliciously. Structurally. Most AI pilots are built in a way that makes success almost unavoidable — and tells you almost nothing about production.

Three ways a pilot cheats

It can see the answer.The demo runs on a proposal you already finished. The model — or the person prompting it — has access to the final submission, the clarified requirements, the SME answers that took three weeks to extract. Your team at kickoff had none of that. A test with the answer key in the room measures nothing.

It runs on cleaned-up knowledge.Someone curated twenty good documents into the tool before the demo. Your real library has nine hundred documents, a third of them stale, two of them contradicting each other about your largest contract. The pilot tested a knowledge base you don’t have.

It skips the pressure.No amendment landing on day nine. No page limit. No SME who answers in fragments at 11pm. No reviewer who has to decide whether an unsupported claim goes out under the company’s name. The conditions that make pursuit work hard are exactly the conditions the pilot removed.

This is why the maturity assessments and capability scores keep disappointing. They measure opinions about readiness and benchmarks of model skill. Neither tests whether a redesigned workflow survives contact with your organization.

What an honest test looks like

Time-lock the evidence.Replay a finished pursuit, but at every step the AI sees only what existed at that historical moment. Kickoff tasks get kickoff artifacts. The final submission stays hidden — it’s the evaluation material, never an input.
Separate the conditions.Baseline (what actually happened) vs. model-only on your messy knowledge vs. the governed system vs. your own team on the redesigned workflow. The gaps between them tell you what the model contributed, what fixing the knowledge contributed, and what adoption contributed.
Use held-out data.Don’t develop and test on the same sections. NIST’s evaluation guidance says the quiet part: document your test sets and methods or your results mean nothing outside the room.
Test adversarially.Contradictory amendments. Stale sources. Requirements buried in tables. Instructions hidden inside documents the AI should refuse. Missing information. If the pilot never saw a hard case, the pilot never happened.
Blind the review.Reviewers score paired artifacts without knowing which is historical and which is AI-assisted. Compliance, accuracy, evidence strength, revision required. Deterministic checks where possible; an LLM judge never grades another LLM’s work alone.
Measure reliance, not applause.Adoption isn’t “people used it.” It’s whether people accepted correct suggestions and rejected wrong ones. Automation bias and algorithm disuse are both failure modes, and both are measurable.

End with a verdict, not a vibe

An honest pilot ends with a decision per workflow: retire it, rework it, run it in shadow, promote it to copilot, or delegatenarrow authority to it — against thresholds you agreed before the test ran. Zero missed mandatory requirements. Citation coverage above 95 percent. A tested human fallback. A named owner. Write it down as a certificate that states the exact conditions under which the workflow passed. Anything less is a demo with a scorecard stapled on.

Now go do

Take the last AI demo you saw and ask three questions: What information did it have that our team wouldn’t have had at that point? Whose knowledge base was it running on — ours, or a curated one? What happened when it hit a hard case? If nobody can answer, you watched a performance.

This test exists. We run it.

Pursuit Replay reconstructs a completed pursuit and replays it with AI under time-locked evidence — before you risk a live one. You can inspect a full sample readout right now.

See how a replay runs