The hidden challenges behind collecting, validating, calculating, and proving ESG data and a practical framework for turning fragmented information into reliable, disclosure-ready data.
1. Introduction: The ESG Reporting Problem Starts Before the Report
Imagine it is the final month before your company’s ESG disclosure is due. The sustainability team is looking at a number that will eventually appear in a public report: total electricity consumption across the business.
It looks simple. There is a number in the spreadsheet. Someone has checked the arithmetic. The reporting template is waiting.
Then someone asks a seemingly harmless question: “Where did this number come from?”
Suddenly, the spreadsheet stops being the answer. It becomes the beginning of an investigation.
One facility sent a utility bill. Another sent a monthly export from a building-management system. A third reported electricity in a different unit. Finance has a record of the amount paid, but not the physical consumption. Facilities owns the meter readings. Procurement owns several supplier records. The sustainability team owns the disclosure. And somewhere in the middle, an assumption was made six months ago that nobody can quite remember.
This is where ESG reporting really begins.
The report is only the visible end of the process. Behind every disclosed metric sits a chain of sources, people, systems, calculations, assumptions, approvals and evidence.
That is why the real problem is not simply collecting more ESG data. It is collecting data that is accurate, consistent, attributable, traceable, verifiable and defensible.
2. What Is ESG Data Collection and Why Does It Matter?
At its simplest, ESG data collection is the process of gathering and organizing the information an organization needs to understand and disclose its environmental, social and governance performance. In practice, however, the word “collection” makes the process sound much easier than it is.
A modern company may need to bring together sustainability data from finance, operations, HR, facilities, procurement, IT, travel providers and suppliers. The information may arrive through APIs, ERP exports, utility bills, questionnaires, spreadsheets, invoices, meters or manually entered records.
And the metrics themselves span very different parts of the business.
Environmental data
- Energy consumption
- GHG emissions
- Water consumption
- Waste
- Transportation and fuel
- Biodiversity and resource use
Social data
- Workforce and headcount
- Diversity and inclusion
- Health and safety
- Employee turnover
- Training and development
- Human-rights indicators
Governance data
- Board composition
- Ethics and compliance
- Risk management
- Anti-corruption
- Data privacy and controls
The difficult part is what happens between those sources and the final disclosure. A single metric can pass through several departments, systems and transformations before anyone outside the organization sees it.
That is why a serious ESG data collection process is closer to a controlled data supply chain than a spreadsheet exercise.
3. Where ESG Data Actually Comes From
One of the fastest ways to understand ESG data collection is to stop thinking about “the ESG dataset” as a single thing.
There usually isn’t one dataset. There is a network of datasets.
| Data | Potential Source | Typical Owner |
|---|---|---|
| Electricity consumption | Utility bills / meters | Facilities |
| Fuel consumption | Fleet records | Operations |
| Employee turnover | HRIS | HR |
| Business travel | Travel platform | Finance / HR |
| Purchased goods | ERP / procurement | Procurement |
| Supplier emissions | Supplier questionnaires | Procurement / Sustainability |
| Waste | Waste contractor | Facilities / EHS |
| Water consumption | Utility records | Facilities |
That distribution creates a deceptively hard problem: the person responsible for the disclosure may not control the source.
This distinction matters enormously. Reporting ownership and data ownership are not the same thing.
4. The 7 Hidden Failure Points in ESG Data Collection
Most articles stop at “ESG data is difficult.” That is true, but it is not particularly useful.
A better question is: where, exactly, does the process fail?
Once you look at an ESG data point as a journey, seven recurring failure points become visible.
Nobody Clearly Owns the Data
The sustainability team may be accountable for the final disclosure, but it may not own the underlying electricity, workforce, procurement or supplier information.
Consider the simple example of electricity. Sustainability needs the metric. Facilities owns the meter. Finance may own the utility invoice. An external utility company may own the underlying bill. An ESG reviewer eventually signs off on the reported figure.
If nobody has been explicitly assigned responsibility for the metric, the process becomes a chain of assumptions.
Good ESG data collection best practices therefore start with ownership. Every important metric should have someone who knows what it means, where it comes from, when it is due, what evidence supports it and what to do when something looks wrong.
The Data Exists but Not in the Right Format
This is one of the most overlooked problems in sustainability data collection.
Sometimes the data is not missing at all. It is simply incompatible.
One facility reports 125,000 kWh. Another reports 125 MWh. A third sends an electricity bill showing $18,500 of spend.
All three records tell us something. But they are not immediately comparable.
Normalization is the quiet machinery of an ESG data platform. Units, currencies, reporting periods, organizational boundaries and definitions need to be brought into a common structure before calculations can be trusted.
Data availability does not guarantee data usability.
The Data Doesn’t Exist
Then there is the harder case: the number genuinely isn’t available.
A supplier does not provide emissions data. A facility has no historical meter records. An older acquisition comes with incomplete documentation. A small supplier simply does not maintain a sustainability dataset.
This is where estimation enters the process.
Estimation is not inherently a failure. Undocumented estimation is.
A mature process distinguishes between actual data, estimated data and proxy data, and records the assumptions, emission factors, methodology, source, reporting period, confidence or quality assessment and responsible reviewer.
That distinction gives the organization something valuable: visibility into where its data is strong and where better primary data would materially improve the result.
Different Teams Use Different Methodologies
Two teams can work with the same underlying activity and still produce different answers.
One may use a different boundary. Another may apply a different emission factor. A third may use a different reporting period.
Without standard definitions and calculation methodologies, the organization can end up with several numbers that each look reasonable in isolation.
This is why robust ESG data collection software needs more than a place to upload files. The process also needs a methodology layer: consistent definitions, calculation rules, factor libraries and boundaries.
Supplier Data Is One of the Hardest Pieces
Supplier data exposes the central weakness of many ESG programs: the company responsible for the disclosure often does not control the organization that owns the underlying data.
The workflow is rarely a one-time questionnaire. It is closer to a cycle:
This becomes particularly important for Scope 3, where value-chain data can extend beyond direct suppliers. The practical objective is not to demand perfect information from every supplier on day one. It is to prioritize the suppliers and categories where improved primary data can make the greatest difference.
You Have the Number but Can You Prove It?
This is where ESG data collection turns into something more consequential: evidence.
Suppose the final disclosure says a company consumed 125,430 kWh of electricity.
The number may be perfectly correct. But an auditor, assurance provider, customer or internal reviewer could still ask:
- Where did the number come from?
- Who provided it?
- When was it collected?
- Which document supports it?
- Which methodology was used?
- Which emission factor was applied?
- Who reviewed it?
- Was the number changed after approval?
The answers form the metric’s data lineage.
A useful way to think about this is that every material ESG number should carry its own evidence chain: disclosure → calculation → raw data → source document → data owner → period → methodology → review.
ESG Data Changes After You Collect It
Here is the failure point that many ESG processes are not designed for: the number changes.
A supplier revises its submission. An emission factor is updated. A facility is acquired. A facility leaves the reporting boundary. Someone discovers a calculation error. Historical information is restated.
Which version of the number is the official one?
This is why ESG data version control matters.
For every material change, an organization should be able to answer: what changed, why it changed, who changed it, when it changed, and what the previous value was.
Without that history, the organization may have a current number but no reliable story about how it got there.
Case Study: When ESG Data Controls Break Down: The Goldman Sachs Example
ESG data problems rarely announce themselves as a dramatic spreadsheet error. Sometimes the problem is quieter: the company has a methodology, a process and a set of controls on paper but the people collecting or using the information do something different in practice.
A useful U.S. example comes from Goldman Sachs Asset Management (GSAM). In November 2022, the U.S. Securities and Exchange Commission (SEC) charged GSAM over failures in policies and procedures involving ESG research used for two mutual funds and a separately managed account strategy marketed as ESG investments. GSAM agreed to pay a $4 million penalty.
According to the SEC, the issue was not simply that an ESG number was “wrong.” The regulator found that, from April 2017 to February 2020, GSAM had several failures in how its ESG research policies and procedures were established and followed. For one product, the SEC said there were no written ESG research policies and procedures from April 2017 to June 2018. After policies were established, they were not consistently followed.
The SEC also found that questionnaires required for companies being considered for portfolio inclusion were often completed after securities had already been selected. In some cases, personnel relied on earlier ESG research that had been conducted differently from the process required by the stated policies.
That distinction matters for any company building an audit-ready ESG data process. The lesson is not simply “check your ESG numbers.” It is to make the entire journey of the data—and the decisions based on it—traceable.
A reviewer should be able to answer questions such as:
- What data or evidence was required?
- Where did it come from?
- Which methodology or calculation rule was supposed to be used?
- Who was responsible for collecting and reviewing it?
- Was the required information available before the relevant decision was made?
- Which version of the methodology or source was used?
- Can the company demonstrate that the documented process was actually followed?
The most dangerous ESG data problem is not always a wrong number. Sometimes, it is a number—or a decision based on ESG information—whose journey cannot be proven.
This case maps directly to several of the failure points discussed above: unclear ownership, inconsistent methodologies, weak evidence trails and gaps between the documented process and the process actually followed. It also shows why ESG data collection software and governance controls need to do more than store a final value: they need to preserve context, methodology, ownership and evidence around that value.
Primary reference: U.S. Securities and Exchange Commission, “SEC Charges Goldman Sachs Asset Management for Failing to Follow its Policies and Procedures Involving ESG Investments,” November 22, 2022. View the SEC release.
Source note: The SEC described these findings as policies-and-procedures failures and did not characterize the case as proof that GSAM fabricated a particular ESG metric. This article uses the case specifically to illustrate the importance of controlled, traceable ESG data processes.
5. The ESG Data Lifecycle: A Practical Framework
Once the failure points are visible, the solution becomes much clearer. Instead of treating ESG data collection as a single activity, treat it as a lifecycle.
6. What an Audit-Ready ESG Data Process Looks Like
There is a useful difference between having an ESG reporting process and having an ESG data process.
In a fragmented environment, the workflow may look like this:
Fragmented Process
Controlled Process
The organization can still produce a report. It may even produce it on time. But every reporting cycle starts the same way: chasing information, reconciling conflicting spreadsheets, asking people to explain old calculations and trying to reconstruct evidence that should have been captured at the beginning.
A controlled ESG data process changes the sequence:
This is where automated ESG data collection becomes useful—not because automation magically makes ESG data accurate, but because it can make a good process repeatable.
Credibl positions its ESG platform around automated collection, validation, data management and reporting, including structured and unstructured data workflows. See: https://www.crediblesg.com/
7. How US Companies Can Improve ESG Data Collection
- Establish clear data ownership — every material metric should have an accountable owner.
- Create standardized definitions — define exactly what each metric means and what is inside or outside its boundary.
- Centralize ESG data — reduce scattered spreadsheets and disconnected repositories.
- Automate recurring collection — particularly where the same information is requested repeatedly.
- Build validation rules — flag missing values, outliers, duplicate records, incorrect units and unexpected changes.
- Track data lineage — maintain the relationship between source, data, calculation and disclosure.
- Separate actual and estimated data — do not hide estimates inside aggregate numbers.
- Maintain an audit trail — preserve changes, approvals and supporting evidence.
The important word here is “process.” An ESG data platform can support the process, but it cannot replace ownership, definitions or judgment.
8. ESG Data Collection Checklist
Before collecting
- Have we defined the metric?
- Have we identified the source?
- Have we assigned an owner?
During collection
- Is the reporting period clear?
- Are units standardized?
- Is the data complete?
- Is supporting evidence available?
During validation
- Are there anomalies?
- Are calculations consistent?
- Are estimates clearly identified?
- Has the methodology been documented?
Before reporting
- Can we trace the number back to its source?
- Has it been reviewed?
- Can we reproduce the calculation?
- Is there a record of changes?
9. Conclusion — Better ESG Reporting Starts With Better Data
The most important shift is conceptual.
ESG reporting is not where the data story begins. It is where the story becomes visible.
The real work happens earlier: when someone reads a utility bill, when a supplier answers a questionnaire, when Finance exports a transaction file, when HR updates a workforce record, when an emission factor is selected, and when a reviewer decides whether a number is ready to approve.
A company cannot build reliable disclosures from data that is unowned, inconsistent, unverified, unsupported or untraceable.
The goal is more ambitious: build an ESG data process in which every important metric can be traced from disclosure back to its underlying evidence—and forward again into the calculation and reporting outputs that depend on it.
That is what turns sustainability data collection from a recurring administrative burden into an organizational capability.
And it is also where the value of an ESG data platform becomes clearer. The right system does not simply give teams another place to enter numbers. It helps connect sources, owners, evidence, validation, calculation, approval and reporting into one controlled workflow.
Credibl helps organizations centralize ESG data, establish structured collection workflows, improve data quality and maintain traceability across the ESG reporting process. Explore Credibl:
Sources & Further Reading
- Deloitte — ESG data and Scope 3 measurement challenges: https://www.deloitte.com/us/en/services/audit/articles/esg-survey.html
- Credibl — ESG Data Collection & Validation: https://www.crediblesg.com/platforms/act/esg-data-collection-validation/
- Credibl — Supplier Assessment: https://www.crediblesg.com/platforms/adopt/supplier-assessment/
- Credibl — Sustainability Intelligence Platform: https://www.crediblesg.com/