// The Systems Lens

Your Utility Has Been Doing Data Governance Wrong

The industry has been talking about data governance for two decades. The results are almost nowhere. Here is why the gap is about to become very visible, in the most uncomfortable way possible.

Your Utility Has Been Doing Data Governance Wrong

I spent time recently on a call with one of the largest utilities in the country. Smart people. Serious investment. A governance committee that had been meeting for over a year. When I asked a simple question, "If I need to find every dataset related to residential load forecasting, where do I go?" the answer was immediate and honest: "You'd have to ask Sarah. She knows where everything lives."

Sarah is not a data catalog. She is a person who has worked there for eleven years and carries institutional knowledge that exists nowhere else. She is also a single point of failure for a multi-billion dollar operation.

When institutional knowledge lives in one person rather than your data systems, you have already lost the governance battle.

When the most critical governance tool in your organization is a person named Sarah, the documented strategy and the operational reality have never met. When the most critical governance tool in your organization is a person named Sarah, the documented strategy and the operational reality have never met.

What I heard in that answer is not unusual. It is the rule. The utility industry has been talking about data governance for two decades and most organizations are still relying on their version of Sarah. The conversation has been real. The progress has not kept up. And now, with AI moving from boardroom aspiration to regulatory expectation, the gap between where most utilities are and where they need to be is about to become very visible, in the most uncomfortable way possible.


The Honest State of Utility Data

A typical utility data estate looks something like this.

There is a SCADA historian, probably OSIsoft PI, holding years of operational time-series data, accessible only through a proprietary client that three people in the OT group know how to use. There is a GIS system with asset geography, maintained by a team that treats it as their system rather than a shared resource. There is an AMI platform producing enormous volumes of meter data that the analytics team would love to use but struggles to access in any consistent, documented way. And there is an ERP, SAP or Oracle or Maximo, holding asset records and work orders and financial data, with a schema so complex that even the vendor's consultants approach it carefully.

None of these systems were designed to share information with each other. They were designed to do their specific jobs, and they do those jobs well. That is not the problem. The problem is that nobody built the layer above them. The layer that says: here is what we have, here is what it means, here is how it connects, here is who owns it, and here is the regulatory obligation attached to it.

That layer is what data governance is supposed to be. In most utilities, it exists as a set of policies, a steering committee, and a large spreadsheet that was last updated nine months ago.

The numbers behind two decades of governance investment reveal a gap between what utilities documented and what they actually built.

Two decades of governance investment and the gap between what utilities documented and what they actually built has only grown wider. Two decades of governance investment and the gap between what utilities documented and what they actually built has only grown wider.

Governance that lives in a spreadsheet is not governance. It is a photograph of governance. A point-in-time record that begins aging the moment it is saved.

Hardeep Anand

Why Every Attempt So Far Has Fallen Short

Utilities have spent real money on this. Collibra. Alation. Informatica. These are not bad tools. They are horizontal tools, built for every industry, and therefore built perfectly for none of them. A utility that buys a generic data catalog and tries to make it speak the language of substations, feeders, and meters is in for a long, expensive project that may or may not get finished before the person leading it moves on.

There is a deeper problem underneath the tool selection, though. Governance has been treated as a documentation exercise rather than an operational capability. The goal has been to describe the data estate, not to automate its governance. Documentation is always behind reality. The moment a new AMI system is deployed or a dataset is copied into a cloud analytics platform, the documentation is stale. Nobody updates the spreadsheet in real time because nobody can.

Horizontal governance tools built for every industry end up purpose-built for none, and utilities pay the price in misfit implementations.

A horizontal tool configured for every industry serves none of them well, and utilities running on proprietary operational systems feel that mismatch most acutely. A horizontal tool configured for every industry serves none of them well, and utilities running on proprietary operational systems feel that mismatch most acutely.

And frankly, there has not been enough urgency. Data governance has been important but not urgent.

That is changing, and it is changing faster than most utility leadership teams realize.


What Is Bearing Down on the Industry

Regulatory pressure is tightening in ways that make ungoverned data a financial liability.

NERC CIP requirements for BES Cyber System Information classification have existed for years, but audits are becoming more rigorous and the question has shifted. It used to be whether you had a classification. Now it is whether that classification is current, defensible, and attached to your actual systems rather than a spreadsheet that may not match reality. The exposure is up to one million dollars per violation per day. That number focuses the mind.

FERC Order 881 created a new class of real-time operational data that must be governed and traceable. Order 2222 opened wholesale markets to distributed energy resource aggregators and introduced data provenance requirements that most utilities have not fully worked through yet. State public utility commissions, driving grid modernization agendas, are asking utilities to justify their data and modeling choices in integrated resource plan filings in ways they were not doing five years ago.

Federal funding is becoming conditional on data quality. The Department of Energy's grid modernization programs are large and real, and they are increasingly tied to FAIR-compliant data, meaning data that is Findable, Accessible, Interoperable, and Reusable. That framework is not a vendor invention. It is the DOE's own standard, adopted from the scientific community, and it is quietly becoming a prerequisite for collaborative research and grant eligibility. Utilities without a governed data foundation are being passed over for programs their peers are accessing.

And then there is AI. State commissions are starting to ask utilities about their AI strategies. The National Association of Regulatory Utility Commissioners published AI principles built around transparency, accountability, fairness, and security. When a commission asks a utility to explain an AI-driven demand response decision or a predictive maintenance recommendation, the answer requires data lineage, a traceable connection from the decision back to the data and the model that produced it. Without a governed data foundation, that answer does not exist. Most utilities are twelve to eighteen months away from that conversation becoming uncomfortable.

Regulatory compliance, federal funding access, and credible AI deployment all require the same underlying capability: a governed, discoverable, traceable data estate. A utility that builds it once satisfies all three. The foundation is the same. The benefits are compounding.

Regulatory compliance, federal funding, and credible AI deployment are three separate pressures converging on exactly the same underlying requirement. Regulatory compliance, federal funding, and credible AI deployment are three separate pressures converging on exactly the same underlying requirement.


What Good Actually Looks Like

The standards for what a governed data estate should look like are well established. They are not frameworks invented by consultants. They are international standards developed by W3C, endorsed by the DOE, adopted by the European Commission, and validated across large data programs in energy, science, and government. Utilities do not need to invent a new approach. They need to apply the right standards to their specific context.

W3C DCAT, the Data Catalog Vocabulary, is the international standard for describing data assets in a machine-readable way. When a utility's datasets are described using DCAT, any tool, any system, any AI model can discover and understand them without human translation. This is the foundation layer. Everything else sits on top of it.

FAIR is the readiness test. Originally developed for scientific data and now a DOE standard, FAIR asks four questions of any dataset: can a machine find it, access it, join it with other data, and reuse it without human intervention? Most utility data fails on at least three of the four. That failure is not an embarrassment. It is a starting point.

NERC CIP classification, FERC traceability, and state PUC requirements form the compliance floor. The governance problem is not that utilities do not know these standards. They do. The problem is that compliance has been treated as documentation rather than automation. Classification needs to be live, not periodic. It needs to be attached to actual systems, not maintained separately in a spreadsheet that drifts from reality.

NARUC's AI principles are the forward mandate. Transparency, accountability, fairness, security. Each one is an evaluation criterion that a commission can apply to an AI deployment. And each one requires data governance as its evidence base. You cannot demonstrate transparency without lineage. You cannot demonstrate accountability without knowing who owns the training data and when it was last validated.

A governed utility data estate in practice looks like this. Every dataset has a machine-readable description: what it contains, how fresh it is, who owns it, what systems it depends on, what regulatory classification applies. That description lives in a catalog that updates automatically as the estate evolves, not in a spreadsheet that someone updates quarterly when they remember. Lineage is automated. When something changes, the people who need to know are alerted. A NERC auditor's request is answered with a query that runs in seconds, not a project that takes three weeks.

This is achievable. It is not a five-year transformation program. It is a sequenced build that delivers value at each step.


Getting the Sequence Right

The most important thing to understand about data governance and AI is that they are not separate initiatives. Getting AI-ready and getting data-ready are the same work, done in the right order.

A utility that deploys AI without a governed data foundation is building on sand. The models get trained on data nobody can fully describe. The decisions they produce are untraceable. When the commission asks the question, and they will, the answer is not there.

Start with an honest diagnostic. Not a vendor demo and not a policy review, but a real assessment of where your data actually lives, what is governed, and what regulatory risk is hiding in the gaps. Most utilities are surprised by what they find. Then build the foundation: a catalog, automated lineage for the critical pipelines, regulatory classification that is live rather than periodic. Then extend that foundation into AI programs so that every model deployed inherits the governance controls of the data it consumed.

Each step in a sequenced governance build delivers measurable value before the next one begins, which is how you sustain momentum.

A sequenced governance build works because each step delivers measurable value on its own, which is how momentum survives the gap between initiative launch and organizational patience. A sequenced governance build works because each step delivers measurable value on its own, which is how momentum survives the gap between initiative launch and organizational patience.

Each step pays for itself. The catalog reduces audit preparation time. Automated lineage compresses rate case data assembly from weeks to days. Live NERC classification removes the exposure that sits in every ungoverned analytics copy of BCSI data. And the AI governance layer produces the documentation a commission will accept as evidence that the utility's AI strategy is real, not aspirational.

The utilities that will lead the next decade of grid modernization are not the ones with the biggest AI ambitions. They are the ones that build the foundation that makes those ambitions defensible.

Hardeep Anand

The utilities that will lead grid modernization are not the ones with the biggest AI ambitions but the ones that built the foundation first.


A Useful Test

Think of the most important dataset in your operation. The one that feeds the most decisions, the most reports, the most regulatory filings.

Can a new member of your team find it without being told it exists? Can they access it without proprietary software? Can they join it with data from a different system without writing custom integration code? Can they understand what it means, who owns it, and what compliance obligations attach to it, without asking Sarah?

If any of those answers is no, you have a governance gap. And in the current environment, governance gaps are not just technical debt. They are regulatory exposure, funding eligibility risk, and a liability in the AI conversation your commission is about to start having with you.

Build the foundation now. Before the audit. Before the commission filing. Before the AI deployment.

And before Sarah decides she has earned her retirement.


Hardeep Anand works with utilities on data and AI governance strategy, helping them move from ungoverned data estates to FAIR-compliant, AI-ready infrastructure using international standards. He can be reached at hardeep@apas.ai.

This article may be shared and republished with attribution.

💬0

Responses

    // Collateral for this article

    Every asset, brand-matched and ready to publish. Upload the featured image to Substack, the carousel to LinkedIn, the cards to X.

    Brand Cards (Chitra) · 10
    banner-1200x300.png
    banner-1200x300.png
    carousel-01-cover.png
    carousel-01-cover.png
    carousel-02.png
    carousel-02.png
    carousel-03.png
    carousel-03.png
    carousel-04.png
    carousel-04.png
    carousel-05.png
    carousel-05.png
    carousel-06.png
    carousel-06.png
    carousel-07.png
    carousel-07.png
    carousel-08-cta.png
    carousel-08-cta.png
    featured-1200x630.png
    featured-1200x630.png
    Concept Figures (Chitra + Roop) · 5
    figure-01-quote.png
    figure-01-quote.png
    figure-02-comparison.png
    figure-02-comparison.png
    figure-03-sequence.png
    figure-03-sequence.png
    figure-04-stat.png
    figure-04-stat.png
    figure-05-quote.png
    figure-05-quote.png
    X Cards (Mudra) · 4
    card-01-1080.png
    card-01-1080.png
    card-02-1080.png
    card-02-1080.png
    card-03-1080.png
    card-03-1080.png
    hook-1600x900.png
    hook-1600x900.png
    HA
    Hardeep Anand

    Thirty years inside water infrastructure, now building the intelligence layer it was missing. The Systems Lens is where the thinking lives, in plain language, every claim sourced.

    Subscribe to The Systems Lens