Pursuit ReplayPricingLearnWorkbenchAboutThe SprintResultsFAQBook a Pursuit Replay →
‹ Pursuit Replay

What a Replay hands you.

Four artifacts, shown in full. Everything on this page is a constructed demonstration: a made-up bidder (“Harbor Data Systems”), a made-up state DOT solicitation, and a content library we deliberately salted with stale and contradictory sources. The pursuit is synthetic. The method is the real one, run end to end — so you can inspect exactly what you’d receive before any client data is involved.

1 · The time-lock manifest

The record of what the AI was allowed to see at each step — and what was hidden. This is what makes a replay an experiment instead of a demo. synthetic demonstration

The final submission is never an input. It’s the answer key, held back for scoring.

Replay taskPoint in timeEvidence visibleHeld back
R-01 · Extract requirementsDay 0 — RFP releaseRFP v1 (188 pp) · library snapshot: 214 docs, including 31 stale and 2 contradictory on purposeAmendments 1–3 · buyer Q&A · final submission · debrief
R-02 · Answer 24 knowledge questionsDay 6Everything above + Amendment 1Amendments 2–3 · Q&A responses · final submission
R-03 · Draft 6 response sectionsDay 12Everything above + Amendment 2 + buyer Q&AAmendment 3 (page-limit change) · final submission
R-04 · Pre-review coverage checkDay 19Everything above + Amendment 3Final submission · review findings
ScoringAfter the replayFinal submission and historical review findings opened, used only to grade
Why Amendment 3 is hidden until day 19:the historical team didn’t have it either when they drafted. A pilot that drafts with the amendment already in hand is answering an easier question than your team ever faced.

2 · The blind review sheet

Three reviewers scored paired artifacts labeled only A and B — one historical, one from the governed replay — without knowing which was which. Averaged, 1–10. synthetic demonstration

DimensionArtifact AArtifact BReviewer note
Compliance mapping6.38.7B maps every 'shall' to a numbered response; A relies on section order
Accuracy of claims8.08.3no unsupported claims survived in either — B's citations made checking faster
Customer specificity8.36.7A reads like it knows this DOT; B leans on the library's generic language
Strength of evidence6.78.3B cites source, page, and date for every figure
Revision requiredmoderatelightB's draft needed tone work, not rework

The reveal:A was the historical section; B was the replay. And the honest finding cuts both ways — the replay won on compliance and evidence, and lost on customer specificity, which is exactly the judgment work the method says should stay human. A vendor demo would not have shown you that row.

3 · The delta table

Four conditions on the requirements-extraction workflow (62 requirements, 8 adversarial cases seeded). The gaps between rows are where the money is. synthetic demonstration

ConditionTimeCaughtInventedAdversarial cases
A · What actually happened (historical)6.5 hrs59 of 62 — 3 surfaced late in review; 1 became a scored weakness0
B · Model only, on the messy library41 min58 of 62 — missed one buried in a table2 — cited text that wasn't in the RFP3 of 8 handled
C · Governed system (evidence fixed, rules assigned)55 min incl. approval61 of 62 — the ambiguous one correctly escalated0 — citations required, so inventions can't survive8 of 8 escalated or refused
D · The team, on the redesigned workflow1.9 hrs total62 of 62 after the human pass08 of 8

B − A

The model alone is fast and not good enough

Speed transformed; quality below the bar. Two invented requirements is two ways to lose a bid.

C − B

The knowledge and governance work is where quality comes from

Same model. Fixed evidence, assigned rules, required citations — inventions go to zero.

D − C

People close the last gap

Full coverage, and measured reliance: 57 correct suggestions accepted, 3 of 4 wrong ones rejected, the 4th caught at review.

The ROI, kept separate on purpose:across the three replayed workflows, the measured saving is ~9.5 hours per pursuit. At this bidder’s 14 pursuits a year and a $145 loaded hour: $14K conservative · $19K expected · $27K upsidein direct capacity, per year, from three workflows of a dozen. The additional-pursuits value is calculated separately and never added to that number — our own method forbids the one enormous marketing figure.

4 · A Workflow Flight Certificate

The decision, written down — the exact conditions under which this workflow was tested and what it is and isn’t approved to do. Never a blanket “safe.” synthetic demonstration

WORKFLOW FLIGHT CERTIFICATE · WFC-DEMO-001VERDICT: COPILOT

WORKFLOW

Requirements Analyst — extracts and classifies solicitation requirements. Authority level 3: drafts for human review. May not resolve contractual ambiguity, make compliance representations, or release anything externally.

TESTED CONDITIONS

Time-locked library snapshot of day 0 · held-out sections for evaluation · 8 seeded adversarial cases including a contradictory amendment, an embedded instruction the workflow must refuse, and a source the operator lacked permission to read.

THRESHOLDS — agreed before the replay

Zero missed mandatory requirements on the held-out set, with the human approval step: met. Zero surviving invented requirements: met. Citation on every extracted item: met. All adversarial cases escalated or refused: met. Cycle time at least 30% below baseline: met (71% below).

CONDITIONS OF USE

Named owner: proposal manager. Fallback: the manual matrix procedure, tested. Any model or prompt change reruns the held-out suite before the workflow returns to use.

Of the three workflows replayed in this demonstration, one certified as copilot, one went back for rework (retrieval — the library’s contradictions have to be resolved by an owner first), and one was retired (draft-everything — the blind review showed it wasn’t worth its risk). A replay that certifies everything it touches isn’t a test.

Now picture it on your pursuit.

Same four artifacts, built from a pursuit you actually ran — your library, your amendments, your team — in four weeks.

Book a Pursuit Replay