
What to Do With Twenty Years of Paper
There is a room, or a storage unit, holding two decades of files. Scanning all of it is a project nobody will fund. Here is the version that is worth doing.

Every established business has one. A room, a basement, a rented unit, holding boxes of files going back to before anybody currently employed there started.
Somebody proposes digitising it roughly every three years. It gets costed, everybody agrees it is sensible, and it does not happen.
Why the full project never gets approved
Because the business case is genuinely weak when the scope is everything.
Scanning two hundred boxes costs real money and produces a large collection of images that nobody can search. You have converted a room of paper into a folder of PDFs, which is tidier and very little more useful. The value was supposed to come from being able to find things, and scanning alone does not deliver that.
So the proposal is correctly rejected, and the room stays.
Most of it genuinely does not matter
The first thing to establish, because it changes the size of the problem by an order of magnitude.
A large share of any archive is beyond any retention period, relates to clients long gone, or duplicates something already held electronically. It is not an asset. It is a cost you are paying rent on, and in the case of personal information it is a liability rather than a neutral one.
Deciding what to destroy is a smaller and more valuable exercise than deciding what to scan, and it is the one almost nobody starts with.
Three categories, three different answers
Things you must keep and will never read. Statutory records past their useful life. Do not scan these. Store them cheaply, index the boxes so you know what is where, and leave them alone.
Things somebody actually asks for. Usually a small, identifiable set: contracts still in force, records for current clients, anything relating to a live matter or an ongoing obligation. This is the only category worth digitising properly, and it is normally under a fifth of the volume.
Things nobody should be keeping. Past retention, no obligation, no client. Destroy them, with a record of what was destroyed and when.
What AI changed about this
The reason the full project never worked is that scanning gave you images and images are not searchable. Making them useful meant somebody reading and indexing every document, which is what made the cost unjustifiable.
That step is what changed. A scan can now be read rather than merely stored.
AI sorts as it reads. What kind of document this is, which client it belongs to, what date it covers, whether it is still live. That is the indexing that used to be the whole cost, and it is what turns a folder of PDFs into something answerable.
AI reads what OCR could not. Handwritten annotations, forms filled in by hand, faded thermal paper, documents photographed at an angle. Traditional character recognition failed on exactly the material an old archive is full of.
AI identifies what should not be kept. Flagging documents past retention, or containing personal information you have no basis to hold, so the destruction decision is made from evidence rather than by guessing at a box level.
AI makes the archive answerable. What did we agree with this client about liability in 2014, answered with the document cited, rather than somebody spending a morning in a storage unit.
The economics of the middle category changed completely, and the outer two categories still should not be scanned.
Do it when there is a reason
The projects that succeed are attached to something that was going to happen anyway.
A move to smaller premises. A system migration. A client asking for their file. An acquisition where somebody is conducting diligence. A regulator or insurer asking a question. Each of those creates a deadline and a budget that a general tidying exercise never will.
If none of those is happening, do the destruction work anyway. It costs almost nothing, it reduces what you are storing, and it makes the eventual project a third of the size.
The box you should open first
Not the oldest one. The most recent boxes that are still being added to.
If paper is still accumulating, digitising the back catalogue solves a problem that keeps recreating itself. Stop the inflow first: whatever is still arriving on paper and going into a box should be captured at the point it arrives instead. That is a smaller change, it takes effect immediately, and it means the archive stops growing while you decide what to do about the rest.
Businesses that skip this step scan two hundred boxes and have two hundred and four a year later.
What people actually search for
Worth knowing before spending anything, because it is rarely what the policy assumes.
In practice, requests cluster around a small set: proving what was agreed, proving something was done on a date, and finding a document a client has lost. Almost nobody goes to an archive to browse.
That means the indexing that matters is client, date and document type. Getting those three right on the middle category delivers most of the value, and anything more elaborate is usually built for a use case nobody has.
Start with one shelf
Take one box and sort it into the three categories above. It takes an hour and it tells you the ratio for the whole room.
Work out what storage actually costs you, including the rent, per year. Most businesses have never calculated it and it is usually more than they assume.
Ask what people have needed from the archive in the last two years. That list defines the middle category better than any policy.
Destroy the third category, properly, with a record.
One hour and one box will tell you whether this is a project or a morning. In most businesses the genuinely valuable portion turns out to be small enough that the decision stops being difficult.
AI Optimize reads the scans rather than merely storing them, so an archive becomes something you can ask a question of. That work sits under Document Intake & Validation.
Related reading

What to Do With Ten Years of Data Nobody Uses
Every established business is holding a decade of records it has never examined. Most of it is worthless and a small part answers questions the business has been guessing at for years.

What to Do About the Shared Drive Nobody Owns
Every business has one. Fifteen years of files, three folder structures layered on top of each other, and nobody willing to delete anything. It is a liability and a search problem at the same time.
WHAT WE BUILD
