csrdready Put me on the waitlist

Kennisbank

Which check belongs to this data point, and who sees it when the value deviates

A data point without a check is a number that no one contradicts. It sits in the report, it comes out of a system or a spreadsheet, and it is accepted because there is no reason to question it. Only when an external party asks a question, or when a figure suddenly shifts by a factor of ten compared to last year, does it become clear that no one has ever looked at it. By then it is too late to reconstruct where it went wrong.

The question that prevents this is simple to ask and hard to answer without structure: which check belongs to this data point, and who notices when it goes wrong. These are two questions, and both deserve an answer before a figure enters a report.

What a valid value is, and what it is not

Every data point has boundaries. Energy consumption cannot be negative. An fte count is not a decimal. A percentage does not exceed a hundred. This seems obvious, but in practice these kinds of boundaries are recorded nowhere — they live in the head of whoever has been processing the data for years, and disappear the moment that person changes roles.

Recording a valid value does not just mean noting a lower and an upper limit. It also means recording in which unit a data point is supplied, which date formats are allowed, and whether an empty field is a valid outcome or a sign that something is missing. Without that agreement, every exception is judged again, by whoever happens to be looking at that moment. What exactly counts as valid, and when a boundary value is more of an assumption than a rule, is worked out on the page about what a valid value is.

The difference between an error and a signal

Not every deviation is an error. A data point can lie completely within the valid boundaries and still be a signal — a consumption that suddenly drops, a count that is three times as high as last quarter, a supplier that fails to supply data for the first time in two years. These are not invalid values. They are values that call for a look from someone who knows the context.

A signal rule is therefore something different from a validation rule. Validation determines whether a value can exist. A signal determines whether a value, even though it is valid, is still a reason to look. How that threshold is set — a fixed percentage deviation, comparison with a historical series, or a combination — depends on the data point and on how stable the underlying activity normally is. The design of such a rule, including the trade-off between too many and too few signals, is described on the page about setting a signal deviation.

Who notices it is not a technical question

A rule that no one sees is not a rule. If a validation or signal check triggers somewhere in a system, it must be clear who receives the notification and what that person does with it. Is it the person entering the data, the owner of the data point, or someone who oversees the whole before the report is compiled? Without a designated recipient, a signal disappears into a log that no one consults.

This touches on ownership, and ownership touches on a problem that remains unresolved in many organizations: different parts of the organization use different definitions for what is, on paper, the same data point. One location counts fte's including contracted staff, another does not. When a signal triggers that a value deviates, the first question is often not whether the value is correct, but whether everyone shares the same definition. How you deal with that is described on the page about definitions that differ per unit.

Why this does not start with a tool

It is tempting to look for these checks in a software package: something that automatically warns, validates, and reports. But a tool that performs checks on data for which no one has recorded what a valid value is, is not performing a check — it merely produces a neat-looking outcome for a process that has still not been thought through. The rules must exist first, independent of whichever system eventually applies them. This order, and why reversing it usually leads to disappointment, is explained on the page about buying a tool first or setting up the process first.

What this delivers before anything is built

When it is established for every data point what a valid value is, when a signal triggers, and who receives that signal, a register emerges that not only documents but is also usable as a set of requirements. That register is exactly what is needed to determine what a system — built internally or purchased — should be able to do. Without that set of requirements, every purchase becomes a gamble. How you translate that into functional requirements that come from your own process, rather than from a supplier, is described on the page about functional requirements from your own process.

Checks on data points also expose how much of the work around them is repetition: making the same comparison, forwarding the same notification, judging the same exception again. Which part of that repetition is suitable to be taken over by AI, and which part instead requires judgment that cannot be automated, is a question that can be answered per task. The work scan from FTE TO AI calculates that per task.

Marvinde assistent van de Data Readiness Scan

Vraag maar waar een datapunt vandaan komt. Dat is meestal de hele vraag.

Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.