Build it in front of them.

Most AI disappointment inside an organization is an ownership and honesty problem, not a model problem.

TitleBuild it in front of them
Issued
StatusOpinion

Opinion. John Arndt, Soxoa. Published .

Most of the organizations I talk to have already tried AI. Someone bought licenses, a handful of people used them seriously for about a month, and the effort settled into a quiet hum of drafted emails and summarized meetings. Leadership looks at the invoice, decides the technology was oversold, and moves on.

That conclusion is usually wrong, and it is expensive, because it writes off the parts that would have worked. The capability sitting in front of these teams is already ahead of what they were asking of it. What was missing was not intelligence. It was that nobody owned the result, and nobody was willing to say out loud which parts of the work should not be touched.

Those are organizational problems. A better model does not fix them, and the next release will not either. What follows is a set of positions rather than a methodology, and I hold them because the alternative keeps producing the same quiet failure.

Build on the real work, in front of the people who do it

A pilot that runs on clean sample data proves that someone can prepare clean sample data. It tells you almost nothing about whether the system will survive Tuesday.

Real work is where both the value and the failure live. The value is in the volume of ordinary cases. The failure is in the ones that are not ordinary: the page that came through upside down, the customer who put three separate requests in one message, the exception somebody in accounting has been handling from memory for six years. A system evaluated on tidy inputs will clear its acceptance criteria and then fail on exactly the cases that made the workflow expensive in the first place.

There is a second reason, and it is political rather than technical. The people who do the work know where the difficulty actually sits, and they will say so in a working session in a way they will not in a requirements interview. They also watch the thing get made. A team that saw the decisions being taken will operate the result; a team handed a finished box will route around it.

In practice: choose one workflow, and have the person who performs it pick thirty to fifty approved real examples, with an explicit instruction to include the ones that ruin their afternoon. Build in a shared room, on a screen everyone can see, and let people argue with the output while it is still cheap to change.

Marking what stays manual is design, not retreat

Every honest drawing has parts marked leave alone. Producing that list early is part of the work, not an admission that the work failed.

Automation economics are not uniform across a workflow. The last stretch of cases often carries most of the risk and a small share of the volume. Those cases cost more to build for, more to monitor, and far more when they go wrong. A workflow where three steps run automatically and one deliberately does not is a finished design, and it will usually beat the version that tried to take all four.

There is a credibility argument too. A consultant who says everything here can be automated has told you something about their commercial model, not about your process. The list of things I will not automate is the part of my advice you can actually check.

In practice: mark the workflow map into three categories, automate, assist, and leave manual, and write the reason beside every manual mark. Circumstances change; a reason lets a future team reopen the decision on purpose instead of by accident.

Every consequential output carries a name

Wherever output reaches a customer, a permanent record, a regulator, or a decision with money attached, a named person reviews it, and that person is chosen before the build starts.

In risk and payments software the operating assumption is that any decision may be examined later by someone who was not in the room: an auditor, a regulator, a customer’s lawyer, a board. That assumption produces better systems, and not mainly because of the examination. It produces better systems because it forces the question of who is answerable to be settled while the design is still on paper.

Designing the review path first changes what gets built. It determines what evidence is logged, what is surfaced to the reviewer, and what they need on screen to decide quickly. Bolted on afterwards, review becomes a queue nobody has time for, and a queue nobody has time for becomes a rubber stamp within a month.

In practice: name the reviewer in the scope document, before the tooling conversation, and give them a deadline, a way to reject, and enough context to decide in under a minute. If nobody will accept the name, you have found something more important than a tooling problem.

Sometimes the drawing should say do not build

A real evaluation has to be able to come back negative, or it was theatre.

The reasons are ordinary: the volume is too low for the maintenance it would need, the exceptions are the actual work, the data is scattered across systems nobody has authority to change, or the process is about to be replaced for reasons unrelated to AI. Building anyway produces a system nobody uses, and spends something scarcer than the budget, which is the organization’s willingness to try a second time.

In practice: write the acceptance bar and the stop condition into the scope before the evaluation runs, while nobody is invested in the answer. Then hold to it. A stop decision reached against a written bar is a result you can act on. A stop decision reached by vibe three months in is a write-off.

A named lead beats a committee

For most organizations the right structure for AI decisions is one accountable person with real authority, not a cross-functional working group.

Committees are good at surfacing concerns and bad at absorbing risk. Every member holds an effective veto and none of them owns the outcome, so the safe move is always to ask for more review. AI decisions are unusually exposed to this, because the risks are easy to articulate and the benefits are specific to a workflow most of the room does not perform. The predictable result is a policy document, a pilot that never ends, and a set of tools the staff use anyway without telling anyone.

One person whose name is on the portfolio behaves differently. They can decline things, which a committee cannot really do, and they can be held to a decision later, which is the part that makes declining credible.

That is the honest argument for a fractional lead, and it comes with a limit worth stating plainly. An outside lead is an advisor, not an officer of your company. Legal, security, and compliance sign-off stays with your own responsible people. What an outside person supplies is someone who does this work continually, who has no vendor relationship to protect, and who can be brought in or let go without restructuring a department.

What you should be holding at the end

The test of an engagement is what still works after the outside party leaves: operating notes written for the person who will run the thing, access that does not depend on a consultant’s account, a named internal owner, and enough understanding in the team to change the system when the work changes, which it will.

If the system only runs while I am in the room, I have not finished. That is the standard I would want applied to me, so it is the one I work to.

If your last attempt at AI produced a subscription and a shrug, the thing to change is not the model. It is who is accountable for the result, and how honest the drawing is about the parts that should stay in human hands.

John ArndtFounder, Soxoa · September 21, 2026

Disagree with something here?

Bring the workflow you think proves me wrong. Thirty minutes, with me, and an honest answer about whether there is anything worth building.