Helium – AI automation agency logo
Helium – AI automation agency logo
Helium – AI automation agency logo
Helium – AI automation agency logo

Where AI Actually Fails in a Small Business

Knowing what will not work is worth more than knowing what will. Four situations where AI reliably disappoints, and what to do in each instead.

Most writing about AI in business describes what it can do. That is the easy half, and it leaves owners to discover the other half at their own expense.

Being precise about the failures is more useful, partly because it saves money and partly because a supplier who tells you where something will not work is easier to believe about where it will.

1. Where the rules were never agreed

The most common failure has nothing to do with capability. It is that the business cannot say what a correct outcome looks like.

Ask two people to qualify the same enquiry, categorise the same expense, or decide whether a job is complete, and you frequently get two defensible answers. That is not a technology problem, it is an unresolved decision, and automating it does not settle it. It encodes one person’s version and hides the disagreement.

What to do instead: write the rule down and get it agreed. An hour in a room. If nobody can produce a rule two people would apply the same way, you have found a management problem wearing a technology costume, and the automation should wait.

2. Where being wrong is expensive and hard to detect

Systems are excellent at doing the same thing every time. They are poor at noticing that the same thing has quietly become the wrong thing.

If an error surfaces immediately, that is manageable. If it surfaces three months later inside a client relationship, or in a compliance context, the cost is disproportionate to whatever was saved.

What to do instead: keep a person in the loop, but only at the point of consequence. Not reviewing everything, which defeats the purpose, but reviewing the cases above a threshold you set deliberately. The boundary should be drawn by what a mistake costs, not by volume.

3. Where the judgement is the product

In professional services particularly, clients are paying for a specific person’s assessment of their situation. Automating that is not efficiency, it is a slow route to becoming interchangeable.

The distinction is between the judgement and the work around it. Reading the material, assembling the background, drafting the routine sections and formatting the output are all work. Deciding what the client should do is the product.

What to do instead: automate up to the decision and stop. Done properly this makes the judgement more valuable, because the person applying it has more time and better prepared material in front of them.

4. Where it runs once a quarter

Frequency decides economics more than complexity does.

A process running four times a year will never repay a build, and by the fourth run the process will have changed anyway. Worse, nobody develops a habit around something that infrequent, so when it does run somebody has forgotten how it works and does it by hand.

What to do instead: keep quarterly work manual and document it properly. The documentation is the actual win there, and it costs a fraction of a build.

The failure that is not AI at all

Worth separating out, because it accounts for a great deal of disappointment.

A system that produces good output which somebody then has to move somewhere by hand has not removed work. It has relocated it. This is not a limitation of the technology, it is a decision made during implementation, and it is the single most common reason a capable system produces no measurable result.

The test is whether the output lands in the place the work actually happens, without a human courier. If it does not, expect the tool to be abandoned regardless of how well it performs.

What the research says about where it does work

Stanford’s AI Index, published in April 2024, reviewed the evidence on assisted work and found that it both speeds up completion and improves output quality, with the largest gains among people who were previously less skilled at the task.

That finding maps neatly onto the failures above. The gains are largest in the work around the judgement, where a less experienced person can now produce something close to what an experienced one would. They are smallest in the judgement itself, which is exactly where businesses should not be automating.

How to choose the first project

Run any candidate against four questions.

  • Can two people state the rule the same way? If not, fix that first.

  • How many times a week does it happen? Under five, be sceptical.

  • What does a mistake cost, and would you notice? This sets where the human review sits.

  • Does the output land where the work happens? If it needs carrying, the value leaks out.

A candidate that survives all four will almost certainly work. One that fails any of them should be fixed rather than built, and the fixing is usually cheaper than the build was going to be.

The failure mode nobody warns you about

There is a fifth situation, and it is the one that catches careful businesses.

A system works well for six months and then stops being right, without breaking. The business changed underneath it. A new service line appeared that does not fit the categories. A supplier altered a format. A stage got renamed. The system continues producing output with complete confidence, and the output is now subtly wrong.

Nothing alerts anybody, because nothing failed. It surfaces when somebody notices a number that looks odd, by which point the wrong output has been flowing for months.

What to do instead: put a name against every system, with fifteen minutes a month to ask whether the business has changed shape. That is the whole governance requirement at this size, and it is skipped almost universally because it feels like nothing.

How to talk to a supplier about this

The questions above also work as a filter on whoever is selling to you.

Ask where they think this will not work. A supplier who cannot name a limitation either does not understand the problem or is not being straight with you, and both cost the same in the end.

Ask what happens when it is wrong, and listen for whether they have thought about detection rather than just accuracy. Ask who owns it in eighteen months. The answers to those three tell you more than any demonstration.

Sources

AI Optimize will tell you when a process should not be automated, which is more often than most suppliers admit. That work sits under Workflow Automation.

Related reading

WHAT WE BUILD

This is the part we solve