An energy bill states kilowatt-hours, a fuel receipt states liters, a waybill states ton-kilometers. Before these figures can appear together in a sustainability report, someone has to convert them to a common unit and often to a common emission factor. That conversion is called normalization. It often happens in a spreadsheet, with a formula that no one can find again once the question arises of where a figure came from.
Normalization is not a single action but a series of choices. Which source provides the raw value. Which unit serves as the standard. Which emission factor or conversion factor is applied, and from which source or which year that factor comes. Is a correction made for calorific value, for temperature, for a different financial year. Every choice is an assumption, and every assumption determines the result. Two organizations that normalize the same raw data with a different factor or a different base year arrive at a different figure, without either of them doing anything wrong.
The problem does not arise at the first calculation. It arises a year later, when someone asks why the figure for 2023 differs slightly from the figure for 2024, or when a controller wants to know which emission factor was used for a specific energy flow. Without recording, the answer is: we no longer know, or: we'll find out. With recording, the answer is a reference to the rule that describes the conversion. The difference between those two situations is exactly what why an audit trail is more than a logbook is about: an audit trail is not a report made afterwards, it is the recording of the choice at the moment it is made.
Normalization is one of the operations that take place between source and report, alongside adding up, filtering, and redistributing. Which operations these are exactly, and in what order they affect the data, is described in which operations sit between source and report when the source is a spreadsheet. For every operation, the same question applies: can it be established what happened, by whom, based on which rule. With normalization, an extra layer is added, because the rule itself may have an external source — an emission factor database that is updated annually. If that database changes and no one has recorded which version was used for which reporting year, the comparability of figures across years can no longer be substantiated.
For every data point that undergoes a normalization step, a number of details are needed to make that step traceable. The raw value and its source. The target unit and the conversion or emission factor used, including the origin and version of that factor. The date or period to which the factor applies. The person or system that carried out the conversion. And the place in the report where the normalized figure ends up. This is, in effect, a specific form of source-to-report mapping, in which not only the origin of a figure is recorded but also the calculation rule that changed it along the way. What that mapping looks like when the source is a spreadsheet is worked out in what is source-to-report mapping when the source is a spreadsheet.
In practice, knowledge of normalization factors often rests with a single employee who knows which tab uses which factor and why. As soon as that person goes on holiday, becomes ill, or changes roles, that knowledge can no longer be retrieved, only reconstructed — with the risk that the reconstruction produces a different outcome than the original. Recording independent of the person means that the rule is traceable in a register, not in someone's head. That register does not need to be a complicated system; it can be a fixed format that is maintained alongside each spreadsheet. What that looks like without purchasing a tool for it is described in how do you create lineage without a tool, and the approach specifically for normalization in spreadsheets in how do you record normalization when the source is a spreadsheet.
There are tools that automate normalization and display the factors used in an overview. That is useful once the rules are known and the assumptions are fixed. A tool that normalizes based on a factor that no one has checked, or that uses an outdated version of an emission factor database without anyone noticing, produces a neat outcome from an incorrect calculation. The order is: first record which rule is applied and why, only then set up a system that executes that rule consistently.
Recording normalization rules is itself a task that takes time: looking up factors, noting versions, documenting the links between raw value and converted figure. Part of that work is repeatable and rule-bound, which makes it a candidate for support by AI. FTE TO AI calculates in the work scan, per task, which part of the work can be taken over in that way, so that it becomes clear where people remain necessary to make choices and where a system can take over execution.
Vraag maar waar een datapunt vandaan komt. Dat is meestal de hele vraag.
Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.