csrdready Put me on the waitlist

Kennisbank

What happens between the source and the report

A figure in a sustainability report has almost always been on a journey. It starts as a line on an energy invoice, a meter reading in a production system, an input field in an HR tool. By the time the figure appears in the report, it has been counted, converted, aggregated and sometimes corrected. Those intermediate steps are rarely visible in the end result. Anyone who only looks at the report sees a number. Anyone who looks at the path in between sees a series of choices.

The steps that almost always occur

Between source and report there is usually a fixed sequence of transformations, even if no one has ever written it down.

First, data is collected: exported from a system, typed in from an invoice, copied from a spreadsheet. It is then normalized: litres become cubic metres, kilowatt-hours become gigajoules, local currency becomes a fixed unit of account. Next it is assigned to a category or scope, which is a choice and not an automatic step. Then comes aggregation: figures from locations, months or departments are combined into an annual total. Along the way corrections occur, for double counting, for missing months, for an incorrect unit that someone noticed a year ago and corrected by hand.

Every step is a place where an assumption is made. An emission factor is chosen. An estimate replaces a missing measurement. A rounding is applied. None of that is a problem in itself. The problem arises when no one knows anymore which assumption was made, by whom, and why.

Why this needs to be documented

A reported figure that cannot be traced back to its source is an asserted figure. As soon as a controller, accountant or supervisory authority asks how a number was built up, the answer must be more than "it's in the system". The answer must be able to show the route: this source, this conversion, this aggregation, this correction.

Documenting that route has three direct consequences. First, finding errors becomes a matter of minutes instead of days, because it is clear where a conversion has been applied and where not. Second, handover becomes possible: if the person who manages the spreadsheet leaves, the knowledge of the transformations does not leave with her. Third, a basis for review is created, because an external party can follow the steps without first having to reconstruct them.

Without that documentation, every reporting cycle is a repetition of detective work. Someone calls the previous manager, searches through old emails, guesses at the reason behind a rounding. That work is invisible in the report itself, but it does determine how much trust that report deserves.

Different transformations, different documentation

Not every transformation calls for the same approach. Aggregation, in which figures from multiple sources are combined into a single total, calls for different documentation than normalization, in which units and definitions are aligned. Anyone who wants to know exactly how to document aggregation will find that in an explanation of documenting aggregation steps, and anyone wondering how to handle normalization can read that in a description of documenting normalization. Both are part of the same chain, but the questions they raise are different: aggregation raises questions about completeness, normalization about consistency.

The situation also changes when the source itself is not a system but a spreadsheet. Then there is no automatic export, no system log, no fixed structure, and the documentation has to be built up differently. Anyone dealing with that situation will find pointers in an explanation of source-to-report mapping when the source is a spreadsheet and in an overview of the transformations between source and report specifically for spreadsheet sources. For those who have no budget for a tool and need to build lineage with the resources already at hand, there is an approach to building lineage without specialized software.

More than a logbook

Documentation of transformations is often confused with a logbook: a list of who changed what and when. That is part of the story, but not the whole of it. An audit trail that only records changes does not explain why a choice was made or which rule was applied. The difference between a logbook and a genuine accountability structure is worked out in an explanation of why an audit trail must be more than a logbook.

The role of the Data Readiness Scan

The Data Readiness Scan maps out this route: which source feeds which data point, which transformations lie in between, who owns each step and which quality rule applies to it. This is not a report and not a questionnaire, but the underlying structure that makes both of those reliable in the first place.

Once the route has been documented

Once the transformations between source and report have been described, it becomes visible which steps are fixed manual work: retyping an invoice, applying a fixed conversion factor, combining monthly figures according to a fixed rule. That is exactly the kind of work for which it can be calculated what part can be taken over by AI, without the documentation from source to report losing its function. FTE TO AI calculates per task which part of it is transferable, based on the work scan that makes that difference visible per task.

The waiting list

The Data Readiness Scan is in development. Anyone who wants the transformations between source and report documented as soon as the scan becomes available can sign up for the waiting list.

Marvinde assistent van de Data Readiness Scan

Vraag maar waar een datapunt vandaan komt. Dat is meestal de hele vraag.

Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.