csrdready Put me on the waitlist

Kennisbank

The problem isn't in the spreadsheet, but in what happens before it

There's a fixed moment in almost every conversation about sustainability data: someone points at the spreadsheet. Too many tabs, too much manual work, too much room for error. The conclusion seems obvious: replace the spreadsheet with a system and the problem is solved. That conclusion is usually drawn too soon.

What a spreadsheet does and does not do

A spreadsheet is a surface. It shows figures, adds them up, links them together. What it does not do is explain where a figure comes from, who is responsible for it, or whether it still matches the definition that was set two years ago. Those questions were never put to the spreadsheet — they were never recorded anywhere. The spreadsheet gets the blame for something that already went wrong earlier: during collection, during retyping, in the assumption that a colleague knew which figure was meant.

Replace the spreadsheet with a software package and those questions remain unanswered. The system then shows a tidier overview of the same uncertainty. The reporting looks more professional; the underlying data has not become more reliable. That is the pitfall: buying a tool before it is clear what that tool is supposed to organize.

What is actually going on

Usually it comes down to three things that grew apart from each other. There is no up-to-date overview of which data points an organization needs — that overview was once made for an old reporting standard and never updated. There is no recorded line from source to reported figure, so no one can say with certainty whether a number comes from one system or another, or from an estimate someone once filled in because the real data wasn't available. And there is no owner per data point — the person who supplies the figure is not automatically the person who can explain where it comes from or what its quality is.

These three things have nothing to do with spreadsheets. They would cause the same problem in any system. A spreadsheet simply makes them more visible, because there is no layer on top hiding the mess.

What we can and cannot show

A data point register with source-to-report lineage records where a data point comes from, who is responsible for it, and which quality rules apply to it. That is useful, and it is also limited. The register does not show a substantive judgment on whether a figure is correct — it shows whether the path to that figure is traceable. Two organizations with the same register can still score differently on data quality, because one organization has a source that is itself inaccurate and the other does not. The register makes that difference visible; it does not resolve it.

Also important: not every data point needs the same amount of lineage. For some figures a simple, well-documented source is sufficient; for others more detail is needed because there are more steps between source and report. Which data points an organization actually needs depends on the reporting obligation and the sector, and that is different from assuming that everything deserves the same amount of attention. How many of the existing data points already have a source varies greatly by organization — for one it is recorded in an ERP link, for another it exists only in the memory of a single employee.

Why this goes slower than buying a tool

Setting up a register and lineage is not a matter of switching on a system. It means checking, topic by topic, where a figure originates, who looks at it before it reaches the report, and what happens if that person is no longer there. That work differs per topic: for one topic the source is already available, for another it still needs to be found or reconstructed. Anyone wondering how much time this takes per topic will find a more realistic answer in the estimate of turnaround time per topic than in a tool demo promising that it all happens automatically.

This approach does not produce a report — that is a different piece of tooling, built on this data as its foundation. What it does produce is a structure that remains in place even when the spreadsheet is replaced, even when the employee who knew everything leaves. What that means in practice for whoever looks at the register later is described in who consults the data point register once the person who built it is gone, and what changes once the lineage has been established is covered in the consequences of lineage once it is in place.

The question that remains

Spreadsheets are not the problem, but they are the first visible symptom. Anyone who replaces the spreadsheet without first knowing which data points actually matter, which source belongs to them, and who is responsible for them, simply moves the problem to a more expensive system. What remains then is the question that precedes this work: which data points are actually needed for the reporting obligation, and whether some of them may already exist somewhere in the organization without anyone knowing, as described in where a data point may already exist.

This is work that people currently do largely by hand: tracking down sources, comparing definitions, checking on ownership. Part of that investigative work can be structured and accelerated with AI, part cannot — that distinction is exactly what FTE TO AI's work scan looks at. The work scan calculates, per task, which part of the work can be taken over by AI, giving a more realistic picture than the assumption that a tool solves the entire problem.

Marvinde assistent van de Data Readiness Scan

Vraag maar waar een datapunt vandaan komt. Dat is meestal de hele vraag.

Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.