A data point only gains meaning once it is established what may and may not be entered. The energy consumption of a location cannot be negative. A percentage does not go above one hundred. An emission factor should fall within a certain range, depending on the source used. Without these boundaries, any value entered can flow through to the report, even if it is the result of an incorrect unit, a misplaced decimal, or a swapped column.
A valid value is therefore not a quality judgment about the content of the figure. It is a technical boundary: a range, a type, a mandatory unit, a relationship with another data point. You establish that boundary per data point, not per report. That is the difference between a check that someone performs manually once a year, and a rule that is fixed alongside the data point itself, so that every entry passes through it.
Not every deviation is an error. A consumption figure that deviates sharply from last year's may be an error, but it could also be the result of a renovation, an acquisition, or a different measurement period. It is therefore useful to distinguish between two types of rules.
The first type is a hard boundary: a value that simply cannot occur. Negative where that is not possible, a percentage above one hundred, a date in the future. These values should be blocked before they proceed further.
The second type is a signal: a value that is possible, but deviates from what you would expect. A sharp increase compared to last year, a figure that lies well outside the range of comparable locations, an entry that arrives right at the deadline without explanation. A signal blocks nothing. It calls for a look, and possibly for an explanation before the figure proceeds further. How you set up and distinguish between these two types of rules is described on the page about setting up a signal deviation per data point.
A rule that triggers but is seen by no one is not a rule. Each data point should therefore have not only a boundary value, but also a name: who owns this data point, and who sees the signal when the value deviates. With a data point that has no owner, a signal disappears into a system that no one monitors, or into a spreadsheet that is only opened again at year-end close.
This ownership need not be complicated. It concerns a name next to the data point, not a process. But without that name, a quality rule is a rule on paper only. Which checks logically belong to which type of data point, and why not every data point needs the same check, can be read on the page about which checks belong to a data point.
The boundaries of a data point do not come out of nowhere. They follow from the definition of the data point: what it measures, in what unit, over what period, for which part of the organization. If that definition differs per unit — one location calculates in square meters of floor area, another in gross floor area — then what counts as a valid value also differs, and a boundary that is correct for one location becomes unusable for another. How to handle this is described on the page about what to do with definitions that differ per unit.
There is also another pitfall: a boundary value that suggests a figure is more precise than the source allows. An emission factor with four decimal places behind an estimate that is only accurate to two digits creates a precision that does not exist. Quality rules should catch this false precision rather than confirm it, something explored further on the page about preventing false precision.
It is tempting to acquire a tool that automatically checks for errors. But a tool placed on top of an unorganized process checks rules that have not yet been established, against data points that no one owns, based on definitions that differ per unit. The result is a tidier report about figures that are still not under control.
The Data Readiness Scan therefore first establishes what exists: the data point register, the origin of each figure, the ownership, and the rules that determine what counts as a valid value and what counts as a signal. This is work that is now often scattered across spreadsheets, email threads, and the heads of a handful of people. The tool that supports this is under construction; anyone who wants to get started with this now can sign up for the waiting list.
Once it is clear which rules and checks belong to which data point, a clear picture also emerges of which part of that control work is repeatable and which part requires judgment. That distinction is central to the work scan of FTE TO AI, which calculates per task which part of the work can be taken over by AI and which part remains with a human.
Vraag maar waar een datapunt vandaan komt. Dat is meestal de hele vraag.
Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.