A spreadsheet with energy consumption, travel data or procurement data is rarely the endpoint. Before a figure appears in a sustainability report, it has been retyped, added up, converted, filtered and sometimes manually corrected. Every step is a place where something can shift relative to the source, and every step that is not documented is a step that no one can retrace afterwards.
The common route from source to report consists of a number of recognisable operations. Data is copied from a source document, often with a manual copying step. Next it is normalised: units are aligned, notations adjusted, missing values filled in or estimated. Then aggregation follows, in which figures from different locations, periods or departments are combined into a single number. In between, corrections take place: a faulty row is adjusted, an outlier is removed, an assumption is applied to an empty cell. At the end of that chain stands the figure that appears in the report.
The problem is not that these operations take place. Normalising and aggregating are necessary to make spreadsheet data usable. The problem is that these steps usually reside in the head of a single employee, or at best in an email exchange that no one can find anymore.
If a figure in a report is called into question, it must be possible to trace where it came from and what was done to it. Without that traceability, every question about a number becomes a search: who adjusted this row, on what basis, and was the same operation applied the year before. With normalisation, for example, the questions are: which conversion factor was used, and has that factor changed since. With aggregation, the questions are: which sources were added together, and was a location accidentally counted twice or skipped altogether.
A spreadsheet without documented operations can produce a different figure the following year from the same source data, simply because someone else performs the normalisation or interprets a correction differently. That is not fraud, it is the absence of a documented process. The result is the same: the figure is not reproducible.
Some organisations think that a change log in the spreadsheet is sufficient. A log records that something has changed, but not why, by whom in what role, and on the basis of which rule. The difference between a log and an audit trail lies in that context: an audit trail makes an operation retraceable, a log only records that something happened.
The absence of a tool is no excuse to skip lineage. Even without specialised software, it is possible to record for each data point which source it comes from, which operations were applied to it and who carried out those operations. This can be done with a fixed structure alongside the spreadsheet itself: a register in which source, operation and responsible party are kept together. How this looks without any tool involved is described on the page about setting up lineage without a tool.
Every operation step needs an owner. Not only of the final data point, but of the operation itself: who decided that this normalisation rule is applied, and who is allowed to change that rule. That ownership is often not established. Two questions can be distinguished here: who owns the definition of a data point, meaning what the number precisely means, and who owns the underlying process, meaning who is responsible for the path along which it comes about. Without having established these two, an operation remains an individual habit rather than an organisational process.
The Data Readiness Scan maps out these operations: which steps take place between source and report, who carries them out and which rules underlie them. This is not a reporting tool or a questionnaire, but a record of the process that precedes the figure.
Once it is clear which operations take place between spreadsheet and report, it also becomes visible which of those steps are repetitive and rule-driven, and therefore eligible to be organised differently. The work scan from FTE TO AI calculates per task which part of the work can be taken over by AI, and thereby connects directly to precisely the operations described here: normalising, aggregating and correcting are steps that, once documented, can be assessed against that question.
Vraag maar waar een datapunt vandaan komt. Dat is meestal de hele vraag.
Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.