csrdready Put me on the waitlist

Kennisbank

A data point that gets counted twice

The problem is not visible in the report

A duplicate data point almost never stands out in the end result. It stands out when two business units independently submit the same energy consumption, the same fleet data, or the same waste stream, and no one has noticed that it concerns the same source. The total then checks out visually, but the sum is wrong. This happens especially at organizations that work per business unit with their own spreadsheets, their own input forms and their own people who supply the figures. Every unit neatly supplies what is asked. No one compares the source.

What a data point register records about this

A data point register does not solve this by checking the total, but by recording per data point where it comes from. Not just which figure was entered, but from which system, which file or which measurement it originates, and who supplied it. If two business units point to the same source for the same data point — the same energy contract, the same fleet management system, the same purchasing report from a shared supplier — that becomes visible as soon as the sources are placed side by side. This is exactly what source-to-report lineage per data point is for: not as an after-the-fact check, but as a fixed part of how the register is built up.

How it arises in practice

Duplication usually arises in three ways. The first is a shared source that is read out separately by two units, such as a central energy contract that is recorded by both the site and the head office. The second is a shared activity that is assigned twice, such as a transport service that is registered as its own emissions by both the sending and the receiving unit. The third is a merger or reorganization in which two registers were combined without anyone checking whether the same source was included under two names. None of these situations can be recognized from the figure itself. They can only be recognized from the source behind it.

What needs to be recorded to be able to see this

To be able to recognize a duplication, the register must record at least the following per data point: the exact source, the business unit that supplies the data point, the period it relates to, and the owner responsible for its accuracy. Without that combination, comparing between units is guesswork. With that combination, it is a matter of sorting: placing all data points with the same source next to each other and assessing whether they may genuinely be counted separately or whether there is overlap. Exactly which fields are needed and in what order they are filled in is described at how to set up a data point register that starts with this information.

When you can trust this

A register that shows duplications is not automatically a register in which duplications no longer occur. It is a register in which they can be traced because the source is recorded for each data point. Whether that is sufficient depends on how many of the data points actually have a source filled in — a register in which half the fields are empty cannot possibly show which units overlap. How many source fields are filled in, and what that says about the reliability of the whole, is addressed at how many of your data points have a source. Only once that proportion is large enough is a comparison between business units more than a snapshot.

Not every data point is worth the effort

The temptation is to check everything for duplication, including data points that carry hardly any weight in the final report. That is not where the time is best spent. Which data points are actually needed for reporting and which are superfluous detail is described at which data points you actually need. A smaller register with sources that are correct is more useful than a complete register in which no one has checked whether the sources overlap.

When this is finished across multiple business units

For a single unit, a register is finished as soon as every data point has a source and an owner. For multiple business units, an additional step is added: checking whether the same source is not listed under two names in the register. What that means concretely for finalizing a register that spans units is described at when a register is finished with multiple business units. That moment is not when all fields have been filled in, but when the entries per unit have been checked against each other.

Once the data point register is in place, with sources and ownership per data point, a second question arises: who does the work of periodically checking and updating those sources. Part of that work — retrieving source data, comparing entries, flagging deviations — is repetitive enough to examine what portion of it can be handed over to AI. FTE TO AI's work scan calculates per task which part of the work qualifies for this, not as a replacement for the register but as a follow-up step once the register exists.

Marvinde assistent van de Data Readiness Scan

Vraag maar waar een datapunt vandaan komt. Dat is meestal de hele vraag.

Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.