Helium – AI automation agency logo
Helium – AI automation agency logo
Helium – AI automation agency logo
Helium – AI automation agency logo

Keeping a Record of What Your AI Did

A client disputes a figure your system produced eight months ago. You need to show what it saw, what it decided, and who checked it. Most businesses cannot.

A client challenges something. An invoice amount, a rejected application, a quote, a categorisation. It was produced by a system in September and it is now May.

The question is not whether the system is generally accurate. It is what happened in that specific case, and whether you can show it.

Why this is different from ordinary software

Conventional software applies a rule. If you know the rule and the input, you can reconstruct the output at any point in the future, because the same input always produces the same result.

A system with an AI step does not work that way. The output depends on the model version, the instructions in place at the time, and what context it was given. Change any of those and the same input may produce something different.

So reconstruction after the fact is not reliable. If it was not recorded when it happened, it is gone, and the honest answer to the client is that you do not know.

What the regulators are asking for

The European Union’s AI Act entered into force on 1 August 2024, with obligations phasing in from 2 February 2025. For systems in its high risk categories, which include employment decisions, creditworthiness and access to essential services, it requires record keeping and human oversight rather than leaving them optional.

It reaches you only where the output of your system is used inside the Union, so for most Canadian and US businesses it does not apply directly. It is still the clearest published description of what a competent standard looks like, and the obligations it sets out are close to what a serious client will ask you for in procurement.

Quebec businesses have a separate reason to care. Law 25 obliges you to be able to say what personal information you hold and what was done with it, and a system that processes client data without leaving a record makes that question unanswerable.

The six things to record
  • The input, as received. The actual document, message or record, not a summary of it.

  • The output, and how confident it was. Confidence is what tells you later whether this was a routine call or a marginal one.

  • The version of the instructions in force. The single most commonly missing item, and the one that makes a case reconstructable rather than guesswork.

  • Whether a person reviewed it, and who. A name and a timestamp. This is the answer to most questions anybody will ask.

  • What the person changed. Overrides are the most valuable data in the whole system, and almost nobody captures them.

  • What happened next. Where the output went and what acted on it.

The overrides are worth more than the rest

Worth its own section because businesses treat overrides as exceptions to be tolerated rather than as information.

Every time somebody corrects the system, they are telling you where it is wrong, in a specific and actionable way. A pattern of corrections in one category is a rule that needs changing. A rising override rate is the earliest possible warning that the business has moved and the system has not.

Businesses that capture overrides improve their systems continuously and cheaply. Businesses that do not rebuild every two years, having learned nothing from the first version.

What this is not

Not monitoring your staff, and it matters that the distinction is made explicitly when the system is introduced.

The record exists so that a decision can be explained, not so that individuals can be assessed on how often they intervene. If people believe their override rate is a performance measure, they will stop overriding, which destroys the most valuable data in the system and removes the human check at the same time.

Say that out loud. Overriding is the job working correctly, not a fault.

The commercial reason, which is larger

Compliance is the obvious argument and it is not the one that pays.

Larger clients increasingly ask, in procurement, what happens when your AI is wrong and how you would demonstrate it. A business with a straight answer wins work against a business that has to go away and find out. That is happening now in professional services and finance and it is moving outward.

The same record is also what makes the system debuggable, which is why we build it in from the start rather than adding it later. When something goes wrong, the difference between an hour and a week is entirely whether the trail exists.

The two questions to ask your supplier

If somebody has already built you an AI system, these settle in five minutes whether you have a record or not.

Show me a specific decision from six months ago. Not a log that it ran. What the input was, what it decided, and what instructions were in force at the time. If that cannot be produced, the trail does not exist, whatever the documentation says.

What happens to the record when the model is updated. A system that overwrites its configuration on each change has destroyed the ability to explain anything that happened before, and this is a common default rather than a rare oversight.

Build it in, do not add it later

Retrofitting a record is expensive and always incomplete, because the period before you added it stays permanently unexplainable. That gap tends to cover exactly the early months when the system was least reliable and most likely to have produced something somebody will later query.

It costs very little at build time. It is one of the things worth insisting on in a first conversation, and how a supplier reacts to the request tells you a good deal about how they work.

How long to keep it

Long enough to cover the period in which somebody could reasonably dispute the decision, which is usually your contractual or limitation period rather than anything to do with technology.

Two practical notes. Keep the record and the personal data separately where you can, so retention obligations on client information do not force you to destroy your operational history. And decide the period deliberately, because keeping everything forever is its own liability and it is what happens by default when nobody decides.

Sources

AI Optimize records what ran, on what input, what it decided, and where a person intervened, on every system we build. That work sits under Custom AI Integrations.

Related reading

WHAT WE BUILD

This is the part we solve