The Evidence Sprint

You are being asked to commit real budget to AI on the strength of opinions. The sprint replaces them with measurement: ten working days on the workflow where AI promises you the most, tested with and without it against success criteria locked in advance. You end with a go, change or stop decision your board and your finance director can interrogate, and the evidence pack that survives the interrogation.

Ten days to a decision you can defend

Plenty of people can tell you AI is promising. The sprint exists for the harder question: what is it worth, in your operation, on your cases, this quarter? Only measurement can answer that, so measurement is what you get.

You choose the workflow where the promise matters most: quotes that take days to assemble, investigations that close late, reports that senior people rebuild by hand every month. Over ten working days we measure what that workflow actually costs you today, run a controlled trial of AI assistance on your own historical cases, and grade the results blind with your own experts. The sprint ends in a recommendation with only three possible words in it: go, change or stop, judged against thresholds you agreed before any trial ran. Most AI assessments end in a score. Yours ends in a decision, with the evidence attached.

The focus is deliberate. Ten days spent properly measuring the workflow that matters will move a real budget decision; a questionnaire scored across every department will not. And the first measured workflow gives you more than its own answer: you keep the instruments, the baseline discipline and the confidence to point at the next one.

The point is not to make AI look good. The point is to make the decision safe to take. All three verdicts are priced in, and an evidenced stop that protects a budget is as useful a result as an evidenced go.

What you keep

Everything the sprint produces is yours, built to be used without us:

  • The decision report: the workflow as measured, the methodology in plain English, the trial results, the failure modes and what they mean for deployment design, the value model, and the recommendation, with a two-page decision memo your board can read in five minutes.
  • The evidence pack: the locked evaluation plan with its hash, the measurement outputs, the trial run logs and re-run configuration, the grading records, and the analysis itself. It is yours, it depends on no tool of ours, and any competent third party could re-run the analysis from it.
  • A value model built to be attacked: benefits split into cash, capacity and quality, never blended into one flattering total; every assumption stated and editable, so your finance director can substitute their own numbers and the model still stands.
  • The ninety-day plan: what to do with the decision, who owns each benefit, and how the result will be re-measured in production with the same instruments, so the sprint's numbers stay honest after we leave.

If that is the evidence you have been missing, booking takes two minutes. The rest of this page explains exactly how the ten days earn it.

Ten days, four gates

The sprint keeps its ten-day promise because it is gated, not heroic. Four hard gates decide whether the work proceeds, and none of them can be talked around.

Before day one: the start gate

The sprint does not start until the ground is ready: the workflow named, the data requested and arrived or its gaps put in writing, the right people booked into the calendar, and a named budget owner who will sign the baseline. A sprint started on promises instead of data swells into a fifteen-day sprint. Ours starts when the gate says so, which is why ten days means ten days.

Days one to four: measure the ground

We establish what the workflow costs you today, from the strongest evidence your systems can support: system extracts where they exist, case-by-case reconstruction from documents where they do not, and structured observation of the real work alongside the people who do it. We also capture something on day one that pays off at the end: each stakeholder privately predicts the numbers before measurement, and the predictions stay sealed until the readout. On day one, every data gap is raised in writing by the end of the day. By day four the baseline is written down in operational definitions your team has challenged, and your budget owner signs it as a fair basis for judging change. No signature, no trial phase.

In parallel we build a register of how the workflow actually fails today, from your own historical cases, named in your language and graded by severity with your senior practitioner. We also run a practical screen of the workflow's regulatory position, scored and clearly labelled: it is not legal advice, and it says exactly where its authority ends.

Day five: the lock

Left to habit, an AI pilot decides what success means after it has seen the results. Yours will not: we lock the success criteria, the trial cases, the grading rubric, the analysis plan and the go, change and stop thresholds before anything confirmatory runs, the same discipline pre-registered science uses. The locked plan is committed with a cryptographic hash and the hash is shared with you, so whatever the numbers say, you will know they were not chosen to flatter the outcome. Any later amendment carries a new hash and a stated reason, shown in the report.

Days six to eight: trial and analysis

Every trial case runs in both conditions: your current method and the AI-assisted configuration, frozen at the lock. Your own subject-matter experts grade the outputs blind: paired, in randomised order, with conditions masked, using binary checks built from the failure register rather than gut scores. Two graders where feasible, with their agreement measured and reported. Analysis follows the locked plan: paired differences with confidence intervals, failure modes counted by severity, and, most consequentially, how often a failure slipped past the human reviewer, because that number prices the checking a real deployment would need.

Days nine and ten: the gate and the decision

Before anything reaches you, the report passes a quality gate run by a founder who did not write it: every headline number traced to its source, the statistics checked, the locked thresholds applied verbatim. If the gate fails, the readout moves and the fee does not. Then the readout: your sealed day-one forecasts are revealed beside the measured results, the findings are walked through, and the recommendation is argued from the thresholds you agreed, not around them. You leave with the decision, a ninety-day plan, and everything we produced.

The honesty rules

The sprint is only worth buying if you can trust an answer you did not want. These rules are how we earn that trust, and each one is built into the sprint's structure rather than left to good intentions.

  • Every finding is graded by how we know it: measured by us, verified against your artefacts, demonstrated to us, asserted to us, or estimated with stated assumptions. Headlines may only rest on the first two grades.
  • "What we could not verify" is a mandatory section of every report. Ten days cannot verify everything, so the report says exactly what it could not, and what it would take to close each gap.
  • The trial is sized for effects worth deploying. We state the smallest effect the trial can detect before it runs. If the result is not distinguishable from no effect, that has a specified meaning: any effect present is too small to fund, which is a decision, not a failed product.
  • Diffuse gains are named as diffuse. Time savings that cannot be consolidated into a budget line are reported as real but not bankable, above the fold of the report.
  • Descope is stated, never silent. If promised data never arrives, we drop to the next strongest measurement method and state the precision cost in the report. If your usable historical cases run out, the sprint converts to a baseline and test-design engagement rather than manufacturing a result.

What it asks of you

Real measurement needs access, so the sprint asks for specific things and books them before day one: a kick-off with the sponsor, process owner and senior expert; three short role interviews; two two-hour observation sessions in which your operators simply do their normal work while we watch; a baseline sign-off session with the budget owner; two to three hours of blind grading from your experts, splittable across sittings; and a ninety-minute readout. The observation sessions are passive time; the active items are what add up to the eight to ten hours described below. The workflow should run at least weekly, and around fifteen usable historical cases need to exist. If your situation does not fit that shape, say so when you book: the honest answer may be a different starting point, and we will tell you.

The investment

A fixed fee, agreed before you book
One figure covers the full ten-day sprint: baseline measurement, controlled trial, blind grading, error analysis, value model, decision report and evidence pack. The figure you book at is the figure you pay, with nothing metered and nothing added.
About eight to ten hours
Your team's active time across the ten days. Real evidence costs real involvement, so we state that cost openly and honour it: if grading volume threatens the budget, we shrink the trial and say so in the report rather than overdrawing your hours.

After the sprint

The sprint is complete in itself; every route from it is a choice, priced at the readout. When the decision is go, the natural next step is building the production capability, with what operators said they would need to trust the system carried into the build as requirements. When you want the numbers watched after deployment, a measurement retainer keeps the evidence honest: a monthly metric refresh on the workflow, or multi-workflow assurance with quarterly re-measurement and a board-ready pack, each a fixed monthly fee agreed before it starts. And when the sprint's answer raises a bigger question about direction, that conversation belongs to direction and diagnosis.

Book a sprint

Your sprint is led personally by Ballista's founders, the same people whose names go on the report; you can meet them here before you book. Tell us about the workflow: what it is, roughly how often it runs, and what makes it painful. We will reply with a straight answer on whether it can carry a sprint, and the pre-sprint checklist if it can.

Required. Treated as confidential from the first word.

Prefer to talk it through first?

If you are not ready to name a workflow, describe the operational pressure instead. We will help you find the workflow worth measuring, or tell you plainly if a sprint is not the right first step.

Your message goes to the people who would do the work.