A duplicate data point rarely stands out by its name. Two rows in the register are usually named differently — 'office energy consumption' and 'electricity consumption NL location' — while they refer to the same meter, the same invoice or the same source. The duplication is not in the text, but in the origin. Anyone searching by name alone will not find the duplicates.
A data point register is not a list of reporting rows, but a register of the underlying data: for each data point, the definition, the source, the owner and the path from source to report. Each data point gets a fixed place with four fixed fields: what it is, where it comes from, who is responsible for it and which processing steps have been applied to it before it ends up in a report. Without those four fields, a register is a glossary, not a control instrument.
The question what is source-to-report mapping describes that third step in detail: the route a figure travels from source system to report row. That same route is the instrument that makes duplications visible. Two data points with a different name but an identical source-to-report line are the same data point, entered twice.
Duplicate data points usually do not arise from carelessness, but from organizational structure. A facilities team records energy consumption for building management, a sustainability team records the same consumption for CSRD reporting. Both teams work from their own spreadsheet, with their own naming, without visibility into each other's records. With multiple locations or business units, this pattern multiplies: each unit records similar data under its own label.
The page where does a data point already exist at multiple business units addresses precisely this mechanism: before you add a data point to the register, the first question is whether it already exists somewhere in the organization, only under a different name or at a different department. Asking that question before registration prevents a large part of the duplications that would otherwise have to be traced afterward.
The reliable way to recognize a duplication is not comparing descriptions, but comparing sources. Two data points that come from the same system, the same table or the same export file, with the same reference date and the same unit, are candidates for merging — even if the naming gives no indication of that. Conversely, two data points with a similar name but a different source may not be a duplication, but two separate measurements that happen to resemble each other.
This distinction can only be made if the register records the source per data point, and not only the final value. A register that merely collects figures without provenance cannot detect duplications — at best, it can suspect them.
Besides the source, ownership is a second indicator. When two data points share the same source but are registered under different owners, that is not in itself a problem — it may indicate a data point that is used by multiple departments. It does become a signal when both owners independently carry out the same processing steps, without knowing of each other's work. In that case, it is not the data point that is duplicated, but the process around it.
The question which processing steps sit between source and report is relevant here: two data points with the same source but different processing steps may produce different figures, while appearing in the report under the same heading. That is a riskier situation than a simple duplication, because the figures can diverge without anyone noticing.
Tracing duplicate data points is not a one-time exercise that ends at a fixed moment. As an organization grows, systems change or reporting boundaries shift, new candidates for duplication arise. The page when is a register complete describes that a register is not complete on a fixed date, but at the moment when every data point has a source, an owner and a verifiable route to the report — including the check on whether that data point already existed elsewhere.
For organizations with multiple business units, the question which data points do you actually need with multiple business units is a good starting point, because that question determines the scope of the register before the search for duplications begins. A smaller, more sharply defined register naturally contains less room for overlap.
Manually searching through sources, owners and processing steps to find duplications is exactly the type of work that can be broken down into repeatable steps: laying data side by side, flagging similarities, drafting a proposal for merging for review. For organizations that want to know which part of that comparison work is transferable to an AI application, and which part continues to require human oversight, the work scan from FTE TO AI calculates this at the task level. The scan gives no verdict on your register, but a breakdown of the tasks it contains — a starting point before you decide how to organize the tracing of duplications.
Vraag maar waar een datapunt vandaan komt. Dat is meestal de hele vraag.
Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.