Quite often, organizations archive data without engaging fully with the complexities of how to arrange it, what to preserve, and most importantly why it is being preserved. Organizing archives is an important step in ensuring both ease of future use but also compliance, security and minimized data loads.

The question arises time and again, usually when older data unexpectedly resurfaces. The current event bringing it to the forefront in some circles is a leak of about twelve terabytes worth of varied data from Valve’s Steam. Much of the data is unimportant, of course, but major questions remain and some of them have wider application.
The most immediately obvious question is one of simple security – how an archive of older data ends up becoming publicly accessible. This is something that we have spoken about at length on this blog, and it bears repeating. Keeping your files archived is a necessity, but so is screening the business-critical information from exposure.
A more interesting question is delving into what has ended up in the data, and what it says about retained information. In this case, it includes older forms of projects not previously released to the public (including third-party data from developers who publish through Valve’s storefront). This reality leads us to the most important considerations: what do I retain, how do I decide, and how do I organize it?
“What do I retain?”
The simplest answer would be “whatever is legally required,” as that gives us a good baseline. This is the minimum data load that is required by law to undergo long-term retention. As always, the longer answer is more complex.
One should consider the types of records that are essential to your organization’s history, operations, and decision-making processes:
- What records are needed for business continuity?
- What records fall under legal hold requirements?
- What data demonstrates clear milestones or achievements?
- Is this data important in full or in aggregate (can I keep only a quarterly summary of financial records, or do I keep them in full?)
- Am I likely to need these records to rebuild the project?
How should I organize my archive?
A clear and consistent organization system is essential for several reasons. First, it enables quick and accurate retrieval of documents, saving time and resources. Second, it helps prevent misfiling, loss, or damage of critical records.
Common methods of organizing include simply leaving existing structures intact. For example, associating emails with their corresponding mailbox and generating user is a first step. The key beyond that is to structure things such that it makes sense for the organization’s specific needs, taking into account the type, volume, and sensitivity of the records. For instance, a company might organize its financial records by year, quarter, or project, while a research institution might categorize its documents by researcher, project, or publication. The important part is ensuring it is easy to identify documents with minimal time and effort.
Why does this matter?
Clear organization prevents data from slipping through the cracks and makes it easier to keep accountability. This both makes it easier to retrieve documents when needed, but it also makes it simpler to keep a grasp of what is stored.
Data volume increases and correspondingly rising costs are a continuous problem for organizations worldwide; having a clear picture of your data structure, what you have retained, where and why is the first step to being able to reduce your data volume and get spiraling costs back under control, without compromising on your continuity, security or compliance.
Your Data in Your hands – With TECH-ARROW
by Matúš Koronthály
Image generated by Canva