top of page

AI for records retention decisions about what to keep and delete

  • 5 days ago
  • 3 min read

Updated: 4 days ago

Introduction


Most businesses keep everything, because deleting feels risky and storage is cheap. The result is a decade of accumulated documents, emails, spreadsheets and personal data with no idea what is in it. This is simultaneously a liability, because you hold personal data you no longer need and cannot account for, and an operational drag, because nobody can find the current version of anything.

The obstacle is not the deletion; it is the classification. Knowing which of four hundred thousand files are accounting records, employment records, contracts, personal data or drafts is the work, and it is the kind of high-volume sorting where automation is genuinely useful. What it cannot do is decide how long each category should be kept, because that is set by law.


1. AI for records retention decisions classifies, it does not decide the period


An important division.

Sorting documents into categories is mechanical. How long each category must be retained is determined by legal, tax, employment and sector requirements, and those differ by jurisdiction.


2. Write the retention schedule first


The governing document.

Category, retention period, the reason for it, and what happens at the end. This is a short document and it is the thing that turns deletion from a risk into a policy.


3. Classify what you actually hold before deciding anything


The discovery step.

Volume, location, category and age. Businesses are routinely surprised: personal data in old spreadsheets, customer records in a former employee's mailbox, copies of identity documents nobody remembers taking.


4. Treat personal data as the priority


Where the obligation is sharpest.

Holding personal information longer than necessary is a breach in many jurisdictions regardless of whether anything goes wrong, and it increases the harm of any breach that does. This category should be addressed first.


5. Do not delete anything subject to a dispute or investigation


The critical exception.

Live or reasonably anticipated litigation, a regulatory enquiry, an insurance claim or an employment dispute all suspend deletion. Automated deletion running through a dispute is a serious problem, and a hold mechanism has to exist.


6. Keep the deletion decision reviewable


Auditability matters.

A record of what was deleted, when, under which policy rule, and by whose authority. Deletion without a record is indistinguishable from loss, and the distinction matters if it is ever questioned.


7. Watch the copies


Where policies fail in practice.

Backups, archives, personal drives, email attachments, exports and former employees' devices. A retention policy applied only to the main system leaves the data in half a dozen other places.


8. Delete in stages and start with the clear cases


Practical sequencing.

Old marketing lists, expired applications, superseded drafts and duplicate copies are unambiguous. Working through these builds the process before you reach anything contentious.


9. Make retention part of creation


The permanent fix.

New records classified and dated as they are created remove the need for a future clean-up. Every business that has done this exercise once concludes that the real answer was not to accumulate in the first place.

Retention periods for accounting, tax, employment, health and safety, and sector-specific records are set by law and vary considerably by jurisdiction, as do data protection obligations and the rules on responding to requests from individuals. Confirm the periods that apply to you rather than adopting a general schedule.


Conclusion


Automate the classification and take the retention periods from the rules that apply to you.

Write the retention schedule before deleting anything, discover what you actually hold and where, address personal data first because holding it too long is itself a breach, suspend deletion for anything subject to a dispute or investigation, keep an auditable record of what was deleted and under which rule, apply the policy to backups, archives and personal drives as well as the main system, start with the unambiguous categories, and classify new records at the point of creation.


Related reading


 
 
 

Comments


bottom of page