Four artifacts, shown in full. Everything on this page is a constructed demonstration: a made-up bidder (“Harbor Data Systems”), a made-up state DOT solicitation, and a content library we deliberately salted with stale and contradictory sources. The pursuit is synthetic. The method is the real one, run end to end — so you can inspect exactly what you’d receive before any client data is involved.
The record of what the AI was allowed to see at each step — and what was hidden. This is what makes a replay an experiment instead of a demo. synthetic demonstration
The final submission is never an input. It’s the answer key, held back for scoring.
| Replay task | Point in time | Evidence visible | Held back |
|---|---|---|---|
| R-01 · Extract requirements | Day 0 — RFP release | RFP v1 (188 pp) · library snapshot: 214 docs, including 31 stale and 2 contradictory on purpose | Amendments 1–3 · buyer Q&A · final submission · debrief |
| R-02 · Answer 24 knowledge questions | Day 6 | Everything above + Amendment 1 | Amendments 2–3 · Q&A responses · final submission |
| R-03 · Draft 6 response sections | Day 12 | Everything above + Amendment 2 + buyer Q&A | Amendment 3 (page-limit change) · final submission |
| R-04 · Pre-review coverage check | Day 19 | Everything above + Amendment 3 | Final submission · review findings |
| Scoring | After the replay | — | Final submission and historical review findings opened, used only to grade |
Three reviewers scored paired artifacts labeled only A and B — one historical, one from the governed replay — without knowing which was which. Averaged, 1–10. synthetic demonstration
| Dimension | Artifact A | Artifact B | Reviewer note |
|---|---|---|---|
| Compliance mapping | 6.3 | 8.7 | B maps every 'shall' to a numbered response; A relies on section order |
| Accuracy of claims | 8.0 | 8.3 | no unsupported claims survived in either — B's citations made checking faster |
| Customer specificity | 8.3 | 6.7 | A reads like it knows this DOT; B leans on the library's generic language |
| Strength of evidence | 6.7 | 8.3 | B cites source, page, and date for every figure |
| Revision required | moderate | light | B's draft needed tone work, not rework |
The reveal:A was the historical section; B was the replay. And the honest finding cuts both ways — the replay won on compliance and evidence, and lost on customer specificity, which is exactly the judgment work the method says should stay human. A vendor demo would not have shown you that row.
Four conditions on the requirements-extraction workflow (62 requirements, 8 adversarial cases seeded). The gaps between rows are where the money is. synthetic demonstration
| Condition | Time | Caught | Invented | Adversarial cases |
|---|---|---|---|---|
| A · What actually happened (historical) | 6.5 hrs | 59 of 62 — 3 surfaced late in review; 1 became a scored weakness | 0 | — |
| B · Model only, on the messy library | 41 min | 58 of 62 — missed one buried in a table | 2 — cited text that wasn't in the RFP | 3 of 8 handled |
| C · Governed system (evidence fixed, rules assigned) | 55 min incl. approval | 61 of 62 — the ambiguous one correctly escalated | 0 — citations required, so inventions can't survive | 8 of 8 escalated or refused |
| D · The team, on the redesigned workflow | 1.9 hrs total | 62 of 62 after the human pass | 0 | 8 of 8 |
B − A
Speed transformed; quality below the bar. Two invented requirements is two ways to lose a bid.
C − B
Same model. Fixed evidence, assigned rules, required citations — inventions go to zero.
D − C
Full coverage, and measured reliance: 57 correct suggestions accepted, 3 of 4 wrong ones rejected, the 4th caught at review.
The decision, written down — the exact conditions under which this workflow was tested and what it is and isn’t approved to do. Never a blanket “safe.” synthetic demonstration
WORKFLOW
Requirements Analyst — extracts and classifies solicitation requirements. Authority level 3: drafts for human review. May not resolve contractual ambiguity, make compliance representations, or release anything externally.
TESTED CONDITIONS
Time-locked library snapshot of day 0 · held-out sections for evaluation · 8 seeded adversarial cases including a contradictory amendment, an embedded instruction the workflow must refuse, and a source the operator lacked permission to read.
THRESHOLDS — agreed before the replay
Zero missed mandatory requirements on the held-out set, with the human approval step: met. Zero surviving invented requirements: met. Citation on every extracted item: met. All adversarial cases escalated or refused: met. Cycle time at least 30% below baseline: met (71% below).
CONDITIONS OF USE
Named owner: proposal manager. Fallback: the manual matrix procedure, tested. Any model or prompt change reruns the held-out suite before the workflow returns to use.
Of the three workflows replayed in this demonstration, one certified as copilot, one went back for rework (retrieval — the library’s contradictions have to be resolved by an owner first), and one was retired (draft-everything — the blind review showed it wasn’t worth its risk). A replay that certifies everything it touches isn’t a test.
Same four artifacts, built from a pursuit you actually ran — your library, your amendments, your team — in four weeks.
Book a Pursuit Replay →