Field notes. John Arndt, Soxoa. Published .
An AI Build Sprint can end with nothing installed. That is not the failure mode; it is one of the designed outcomes, and a sprint that cannot end that way was never an evaluation.
This is the part of the offer people ask about most sceptically, usually phrased as: so I might pay you and get nothing? No. You might pay and get a written recommendation not to build, with the evidence attached. That is a different thing, and on the occasions it is the right answer it is worth more than the build would have been.
Agree the bar before anybody looks at a result
Clears the bar has to mean something specific, and the specifics have to be written down before the evaluation runs, while nobody is invested in the answer.
A usable bar names four things. The task, precisely enough that two people would grade an output the same way. The accuracy or quality threshold, on the real example set, stated as a number or a rule rather than as an impression. What happens to the cases that fall short, which is usually a named reviewer with a workable queue. And the operating cost and effort ceiling, because a system that is accurate and unaffordable to run has not cleared anything.
Then the stop condition, which is the same sentence read backwards: if the threshold is not met on the agreed examples, or the residual cases are too many for the reviewer to absorb, or the running cost exceeds the ceiling, we do not install it. Writing that down in week zero takes twenty minutes. Trying to agree it in week six, with a half-built system on the table and people’s enthusiasm attached to it, is not a negotiation anybody wins.
Evaluate on real examples, before installation
The evaluation runs on the client’s own approved examples, including the awkward ones, and it runs before anything is integrated into a working system. That ordering is the whole point. Integration is where most of the cost and nearly all of the political commitment lives, so a decision made after integration is not really a decision.
Two failure patterns show up at this stage, and both are cheap to find and expensive to miss. The first is a system that performs well on the ordinary cases and badly on exactly the cases that made the workflow expensive. If the difficulty was the exceptions and the exceptions are what fails, the workflow has not been improved regardless of the headline number.
The second is a review load that quietly cancels the benefit. A system with a residual rate that sounds small can still generate more checking work than the manual process took, particularly where reviewing costs about as much as doing. That arithmetic should be done during the evaluation, in the reviewer’s actual minutes, not estimated afterwards.
There are also ordinary, non-technical reasons a sprint stops, and they surface at this stage too: the volume turns out to be too low for the maintenance the system would need, the data is scattered across systems nobody has authority to change, or the process is about to be replaced for reasons that have nothing to do with AI.
What a written stop recommendation contains
If the answer is do not build, what you get is a document, and it should be good enough to take to somebody else and have them argue with it.
The bar, as it was agreed
Quoted from the scope document with its date, so the standard being applied is visibly the one set at the start rather than one arrived at once the results were in.
What was tested, and on what
The approach or approaches evaluated, the example set with its size and where it came from, and how outputs were graded. Enough detail that the evaluation could be repeated by someone else in six months when the tools have moved.
Where it fell short, and by how much
Results against the threshold, broken out by case type rather than averaged into a single figure, because the average is usually the least informative number available. Named examples of the failures, since a failure you can look at is a failure a reader can evaluate for themselves.
Why it fell short
A mechanism, not a verdict. The information needed to make the decision is not in the documents. The exceptions are the actual work and they are each different. The reviewer load exceeds the time saved. Different causes point at different remedies, and a stop report that does not distinguish them is not useful.
What would have to change
The conditions under which this becomes worth revisiting: a data source that would have to exist, a volume threshold, an upstream process fix, a capability that is not there yet. This is the part that gets read a year later.
What to do instead, if anything
Often there is a smaller intervention next to the one that failed, and sometimes the honest recommendation is to spend the money on something that is not AI at all. Occasionally it is to do nothing, which is also a recommendation and should be stated as one rather than implied.
Why that is worth paying for
Compare it against the alternative, which is not a free lunch but a build that goes in anyway. That path costs the installation, the integration work, the training, the year of maintenance, and the slow discovery that people are working around it. It also spends the thing I care most about protecting on a client’s behalf: the organization’s willingness to try a second time. A team that watched one system fail in production is a much harder audience for the next proposal, including the proposal that would have worked.
There is a second argument, about incentives. An evaluation that can only come back positive is not an evaluation, and the party running it knows that. If my commercial model required every sprint to end in an installation, you should read my recommendations accordingly. Being able to stop is what makes the recommendations mean anything, which is the same reason I insist on writing down what I will not automate.
I would rather deliver an unwelcome document than an unused system. The document at least leaves you better informed, and it leaves the next decision open.
Where this fits
The stop decision belongs to the AI Build Sprint, where the bar and the stop condition are written into the scope before work begins. If you are earlier than that and want to know whether you are ready to evaluate anything, the AI Readiness Assessment is free and runs in your browser. Otherwise schedule a strategy call and bring the workflow you suspect will not clear the bar; that is usually the most productive half hour available.