Utility operations and the AI transition
Back to Stories

The AI Transition for Utilities: A Strategic Guide for 2026 and Beyond

Regulatory landscape, the agentic turn, connected AI ecosystems, and a practical framework for getting started

NewGen Strategies & Solutions | April 2026 · Fully Updated August 2026

This report accompanies a detailed blog exploration. Read the companion blog post →

Foundation

Executive Summary

Utilities stand at an inflection point. Adoption is no longer the question: Itron's October 2025 survey of 500 North American electric utility executives found 81% already using AI, with 41% describing it as "fully integrated" into operations. The question now is depth. Most utility AI use remains at the individual-productivity layer: drafting, summarizing, research. The gap between using AI and using it well is where competitive separation is happening. MIT Project NANDA's widely cited 2025 "GenAI Divide" study (a preliminary analysis based on 52 interviews, 153 survey responses, and 300+ public initiatives) estimated that despite $30-40B in enterprise GenAI investment, only 5% of integrated pilots produced measurable P&L impact, not because the technology fails but because tools are not integrated into actual workflows. The utilities that move decisively now (implementing structured governance, investing in data foundations, and building connected AI ecosystems) will shape industry standards for the next decade.

The technology itself has fundamentally shifted. The chatbot era (a text box that answers questions from training memory) is over. The defining development of 2025-2026 is agentic AI: systems that use tools, query live databases, write and execute code, and carry out multi-step work with checkpoints. Underpinning this shift is the Model Context Protocol (MCP), launched by Anthropic in late 2024, adopted by OpenAI, Microsoft, and Google in 2025, and donated to the Linux Foundation in December 2025 as an open industry standard. MCP gives organizations a standard way to connect AI to their own systems (customer information, GIS, work management, document repositories), creating an internal AI ecosystem where employees get governed, audited, conversational access to the company's data and tools. In parallel, small language models (SLMs) running on inexpensive on-premise hardware have made private, secure, customized AI practical for organizations of any size. This report covers both developments in depth, because together they change what "AI strategy" means for a utility.

The regulatory environment remains unformed. As of August 2026, no state Public Utilities Commission has issued formal guidance on utility operational AI, and Arizona's Corporation Commission inquiry (opened in March 2026 and still the only state docket of its kind) remains in the information-gathering stage, with a stakeholder workshop expected before year-end. Federal policy has accelerated dramatically in the opposite direction of restriction: the July 2025 "Winning the Race: America's AI Action Plan" made AI infrastructure a national priority, and the Genesis Mission (Executive Order 14363, November 2025, expanded with $5B+ in July 2026) explicitly targets AI-accelerated grid planning and interconnection. The most concrete operational guidance to date is the December 2025 joint publication from CISA, NSA, FBI, and partner agencies from six allied nations, "Principles for the Secure Integration of Artificial Intelligence in Operational Technology," which sets the de facto standard for AI touching SCADA and control systems. The regulatory vacuum on the operational side creates an unusual advantage for early movers: utilities that proactively engage their regulators, document governance practices, and demonstrate measurable benefits establish precedent rather than follow it.

Industry adoption shows clear patterns. Across North America, utilities have successfully deployed AI in seven core use cases: leak detection, demand forecasting, vegetation management, customer service automation, predictive maintenance, water quality monitoring, and grid optimization. Duke Energy's self-healing grid technology avoided nearly 2.4 million customer outages and more than 11 million outage-hours across its six-state territory in 2024. PG&E operates more than 700 HD wildfire cameras as part of a layered mitigation program and reports 2025 as its third consecutive year without a major fire from its equipment. EPRI's Open Power AI Consortium (launched in March 2025 with Duke, Exelon, Southern Company, Constellation, NVIDIA, Microsoft, and AWS among its founders) is building the first power-sector-specific AI models. Yet the pilot-to-production gap persists: most utility AI pilots still fail to scale, primarily due to fragmented data, poor workflow integration, and insufficient change management. Gartner predicts over 40% of agentic AI projects industry-wide will be canceled by end of 2027 for exactly these reasons. The opportunity lies not in selecting AI technology but in solving the structural problems that prevent successful scaling.

Financial deployment is growing faster than forecast. Market estimates vary by scope, but the direction is consistent: the AI-in-energy market is estimated at roughly $23B in 2025, growing to $28B in 2026 (~22% CAGR), with the utility-operations slice alone at $6.2B in 2025 and projected to reach $7.6B in 2026. National Grid Partners' late-2025 survey found 42% of utility leaders planning to deploy AI within two years, and its official summary reports utilities increasingly judging AI investments on financial return and cost savings, a sign the sector has moved past experimentation budgets. Meanwhile, the cost side of the equation has collapsed: Stanford's 2025 AI Index found the cost of GPT-3.5-level performance fell more than 280-fold in under two years, making both cloud AI at scale and always-on local AI economically routine.

Data readiness is the binding constraint, and the integration layer has changed. Practitioners consistently report that the large majority of machine-learning effort goes to data preparation rather than model development or algorithm selection. Utilities historically accumulated operational data in siloed systems: work management platforms, GIS databases, SCADA networks, customer information systems, and Advanced Metering Infrastructure (AMI) deployments, often incompatible and unlinked. Utilities that have succeeded with AI (Duke Energy, National Grid, PG&E) invested first in data integration, establishing unified data foundations before selecting specific use cases. What is new in 2026 is that the integration target has a standard: rather than building bespoke point-to-point pipelines, organizations increasingly expose each system through an MCP server behind a governed gateway, letting one AI layer reach all of them with per-user permissions and full audit logging. The data work is still most of the effort, but it now compounds into an ecosystem rather than a collection of one-off integrations.

Strategic recommendations are fourfold: start immediately despite regulatory uncertainty, invest in data architecture before algorithms, plan a model portfolio rather than a model choice, and learn from other industries that have navigated similar transformations. The model-portfolio point is new to this edition and reflects how the market matured in 2025-2026: frontier cloud models (Claude, GPT, Gemini) for complex reasoning and long-horizon agentic work, and small, fine-tuned, on-premise models for routine, high-volume, and sensitive tasks, including work adjacent to operational technology where cloud connections are structurally excluded. Healthcare regulation (HIPAA) created privacy frameworks that enable AI deployment at scale. Banking introduced model risk management principles in 2011 that are now applied to AI. Manufacturing established digital twin standards that utilities can adapt. The wheel does not require reinvention. This report provides utilities with the regulatory context, adoption data, technology architecture, cross-industry playbooks, and step-by-step implementation framework needed to transition from PoC paralysis to strategic AI deployment.

Governance

The Regulatory Landscape

The federal framework is fragmented but developing. The Department of Energy's April 2024 "AI for Energy" report established the foundational federal position, identifying AI as central to the nation's energy transition and resilience strategy. The report explicitly calls for AI applications in grid operations, demand response, cybersecurity, and renewable integration. DOE's Artificial Intelligence for Interconnection (AI4IX) program, announced in November 2024 with up to $30M available, focuses on using AI to accelerate the generator interconnection process and represents the federal government's most direct utility AI investment. The National Institute of Standards and Technology published the AI Risk Management Framework in January 2023, providing voluntary guidelines for organizations developing, deploying, or using AI systems. NIST's framework emphasizes measurement, accountability, and transparency, principles that utilities increasingly adopt in governance structures.

The Federal Energy Regulatory Commission is engaging AI through the load side, not the operations side. FERC has not issued guidance specific to utility operational AI. Its AI-related activity in 2025-2026 has centered on the demand AI creates: a February 2025 show-cause proceeding on data-center co-location at PJM generators, a December 18, 2025 order finding PJM's tariff unjust and unreasonable as applied to co-located loads (with compliance filings due early 2026), approval of SPP's High Impact Large Load framework, and what the Commission itself describes as aggressive targeted action on large-load integration in 2026. On operational AI, FERC's perspective remains implicit: the Commission prioritizes wholesale market reliability, transparency, and non-discriminatory access, and utilities implementing AI for grid operations must ensure AI systems do not create unfair market advantages or reduce transparency. The EPA similarly lacks operational AI guidance; its focus remains on cybersecurity frameworks and data protection rather than algorithmic decision-making.

Federal policy has pivoted decisively toward acceleration. President Biden's Executive Order 14110 (October 2023) established a structured framework for federal AI governance; President Trump's Executive Order 14179 (January 2025) replaced it with a deregulatory posture. The defining document is now "Winning the Race: America's AI Action Plan" (July 23, 2025), a 90+ action program whose second pillar is AI infrastructure (grid capacity, permitting reform, and data center buildout), accompanied by same-day executive orders including accelerated federal permitting for data center infrastructure. The Genesis Mission (Executive Order 14363, November 24, 2025) directs the Department of Energy and its seventeen national laboratories to apply AI to scientific and infrastructure challenges; the White House added over $5B in July 2026, and DOE's 26 "Genesis Mission AI Grand Challenges" explicitly include AI-accelerated grid planning and interconnection at 20-100x current decision speed. DOE's Speed to Power initiative (September 2025) and roughly $1.9B in SPARK transmission funding (March 2026) round out the picture. A December 2025 White House framework also signaled intent to preempt fragmented state AI regulation. The net effect: the federal government is treating AI as energy infrastructure policy, while operational AI governance remains with states and market participants.

State Public Utilities Commissions remain focused on the demand side. As of August 2026, no state PUC has issued formal guidance on utility operational AI deployment, and Arizona's Corporation Commission remains the only state to have opened a formal examination. Docket AU-00000A-26-0060, opened by Commissioner Lea Márquez Peterson on March 24, 2026, asks regulated electric, gas, and Class A/B water utilities to file on their use of AI in planning, forecasting, storm response, and procurement, with a stakeholder workshop expected before the end of 2026. No decision or guidance has yet issued. Every other state commission's AI attention is consumed by the demand side: roughly 27 states had large-load or data-center legislation pending in early 2026, Idaho's HB 911 (effective July 2026) requires commission-approved contracts for new loads adding 50 MW or more, and Oklahoma enacted HB 2992, the Data Center Customer Protection Act of 2026, in May 2026. The asymmetry is striking: commissions are building sophisticated frameworks for AI's electricity demand while utility use of AI in operations proceeds without guidance. That gap will close; utilities that engage Arizona's workshop and file thoughtful comments will influence the template other states copy.

Industry bodies are filling the void with principles, education, and shared infrastructure. The Water Environment Federation, with Amazon, The Water Center at the University of Pennsylvania, and Leading Utilities of the World, launched the Water-AI Nexus Center of Excellence in September 2025. Its first output, notably, addressed sustainable water use by data centers rather than utility operations. The Edison Electric Institute made AI a central theme of its 2025 and 2026 annual conferences and convened its first Advancing AI for the Electric Sector Summit in September 2025. EPRI's Open Power AI Consortium (March 2025) goes further than principles: it is building shared, power-sector-specific AI models with utility and hyperscaler members. NARUC's activity remains educational ("AI for Utility Regulators" trainings, AI panels at its 2025 Annual Meeting and 2026 Winter Policy Summit), without an adopted resolution on operational AI. The American Water Works Association's 2026 State of the Water Industry survey (2,171 respondents) shows cautious AI exploration dominated by cybersecurity concern, and the Water Research Foundation has active projects on an AI adoption framework for water and wastewater utilities. None of these carry regulatory force, but they establish baseline expectations and industry norms that utilities increasingly adopt.

International regulators have moved earlier than U.S. states. The United Kingdom's Ofgem published guidance on the ethical use of AI in the energy sector (2025), establishing outcomes-based expectations that cover energy companies' own development and operational use of AI. Ofgem's framework identifies four core outcomes: safety (AI must not reduce grid safety), security (AI must not introduce new cybersecurity vectors), fairness (AI must not create unjust customer outcomes), and sustainability (AI must support decarbonization). Ofgem's approach is outcomes-based rather than prescriptive: companies can achieve outcomes through multiple pathways. Ofwat, the water regulator, runs a £600M Water Innovation Fund supporting a broad portfolio that includes AI projects for water quality monitoring, incident response, and leakage reduction. Australia's regulatory bodies have so far published more limited guidance. European Union regulations focus on algorithmic transparency and consumer protection rather than utility operations.

The rate base question remains unresolved. No identified PUC decision has explicitly ruled on whether utility AI investments are recoverable through the rate base. Most utilities are embedding AI investments in broader IT capital expenditure budgets, where they are recovered as part of general information technology capital. As AI spending grows, utilities will face explicit rate case challenges: can utilities recover AI system development costs, training investments, and model maintenance through rates? Arizona's formal inquiry will likely generate the first explicit guidance. Utilities should anticipate that regulators will apply existing IT capital recovery frameworks initially, but as AI becomes operational (not just research), regulators may require separate accounting, performance metrics, and benefit verification.

Regulatory Entity Jurisdiction Current Guidance Sector Focus Key Actions
Department of Energy Federal AI for Energy Report (Apr 2024) All energy systems AI4IX $30M program; interconnection
NIST Federal AI Risk Management Framework All sectors Voluntary; emphasizes accountability
FERC Federal (wholesale) None on operational AI; Dec 2025 PJM co-location order Wholesale markets; large AI loads Data-center co-location; SPP large-load framework
White House / DOE Federal AI Action Plan (Jul 2025); Genesis Mission EO 14363 (Nov 2025) AI infrastructure; grid capacity $5B+ Genesis funding; grid-planning AI Grand Challenges; Speed to Power
CISA + NSA/FBI + international partners Federal (advisory) Principles for Secure Integration of AI in OT (Dec 2025) Operational technology; SCADA Assess AI use in OT; secure integration; governance
EPA Federal Cybersecurity focus; no AI guidance Water/wastewater; environmental Data protection requirements
Arizona Corporation Commission State Formal Inquiry (Docket AU-00000A-26-0060); workshop expected late 2026 Regulated electric, gas, Class A/B water AI in planning, forecasting, storm response, procurement
Ofgem (UK) International Ethical AI guidance (2025) Energy distribution Safety, security, fairness, sustainability
NARUC State association SaaS/Cloud/AI Primer (Dec 2024) Educational; all sectors Commissioner education; governance principles

NewGen Insight: The Regulatory Opportunity

The regulatory vacuum is temporary and tactical. Arizona's inquiry is moving deliberately (a workshop in late 2026, likely guidance in 2027), and when it produces the first state-level framework, other states will adopt modified versions. The five months since this report's first edition confirm the pattern: no second state has opened an operational-AI docket, but commission attention to AI (via data-center load) has intensified everywhere, and the jump from "AI as load" to "AI in operations" is one rate case away. Utilities that document governance, measure outcomes, and engage regulators proactively will shape these frameworks rather than adapt to them. Conversely, utilities that deploy AI without governance transparency will face retroactive compliance requirements and potential rate adjustments. The winning strategy is engagement and transparency now.

Staged Regulatory Engagement: Move Now and Engage Continuously. Utilities face a false choice between "wait for guidance" and "deploy without regulator visibility." The optimal strategy is staged engagement: Deploy Tier 1 (team plans) and Tier 2 (API integration for IT systems) immediately under existing operational authority; these do not require new regulatory approval. Simultaneously, file proactive inquiries with state PUCs requesting clarification on AI cost recovery, governance frameworks, and data handling requirements. Participate in Arizona's inquiry and similar state-level examinations. Large utilities should petition NARUC collectively for explicit AI guidance by end of 2026, presenting pilot data and governance frameworks as evidence of responsible adoption. This approach reduces regulatory risk by engaging early while avoiding the competitive liability of inaction. OT-domain AI integration should be deferred until regulatory precedent is established (likely 2027-2028), but IT-domain AI should proceed immediately with regulator communication, not in silent isolation.

Market Data

The State of Utility AI Adoption

Adoption has crossed the threshold; depth has not. When this report's first edition was drafted, the defining statistic was the "96-to-26 paradox": 96% of utility executives viewed AI as strategically important while only 26% had advanced beyond proof-of-concept. The 2025-2026 survey wave shows the front half of that paradox resolving fast. Itron's "Grid Edge Intelligence" report (October 2025, survey of 500 North American electric utility executives) found 81% already using AI and 41% describing it as fully integrated, a pace few respondents saw coming, given that a year earlier only 27% had expected AI to reach full integration within five years; 57% cite AI for grid balancing and DER management and 53% for safety applications. IBM's Institute for Business Value reports 94% of utility executives expect AI to contribute significantly to revenue growth within three years. But depth remains the differentiator: National Grid Partners' global survey of 166 utility innovation leaders found 66% citing a shortage of skilled personnel as the top obstacle, and production-scale operational deployments remain concentrated among the largest utilities. The gap between using AI and integrating it reflects structural barriers: fragmented data, immature workflows, organizational silos, and competing capital priorities.

Market sizing reflects strong growth on every definition. Estimates vary with scope and come from commercial market-research firms rather than audited data, but the trendlines agree. The broad AI-in-energy market is estimated at roughly $22.8B in 2025, reaching $27.9B in 2026 (~22% CAGR) and projected toward $60B by 2030. The narrower AI-in-utility-operations segment is estimated at $6.2B in 2025, growing to $7.6B in 2026. These growth rates substantially exceed typical utility IT spending growth, indicating AI's strategic centrality. The capital context matters too: utility capex plans have swollen to historic levels (Duke Energy at $103B for 2026-2030, Southern Company at more than $80B over five years, Xcel at $60B over five years), driven substantially by data-center-era load growth, and AI-enabled planning, forecasting, and asset management are increasingly how utilities intend to deliver those programs efficiently.

Use case maturity is distinctly stratified. Seven core use cases have demonstrated production maturity: leak detection (water/wastewater), demand forecasting (all sectors), vegetation management (electric), customer service automation (all sectors), predictive maintenance (all sectors), water quality monitoring (water/wastewater), and grid optimization (electric). In NewGen's assessment of announced deployments, these Tier 1 use cases represent roughly 60% of current utility AI activity. Tier 2 use cases, characterized as "growing" (water treatment optimization, advanced grid control, outage prediction), account for roughly a quarter and are likely to mature into Tier 1 within 12-24 months. Tier 3 use cases (wildfire detection, supply chain optimization, rate design AI) remain experimental and account for the remainder. This stratification reflects a fundamental pattern: utilities first automate observable, repeatable problems with clear metrics, then expand to integration with operational workflows, and finally tackle strategic decision-making.

The pilot-to-production failure rate is unacceptably high, and now rigorously documented. Industry research has long indicated that the large majority of machine-learning pilots fail to reach production; MIT Project NANDA's preliminary 2025 "GenAI Divide" study put a sharper point on it, estimating that only 5% of integrated enterprise GenAI pilots produced measurable P&L impact despite $30-40B invested, diagnosing a "learning gap" (tools that do not adapt to actual workflows) rather than model inadequacy. Gartner predicts more than 40% of agentic AI projects will be canceled by end of 2027 on escalating costs, unclear business value, and inadequate risk controls. MIT also described a "shadow AI economy": workers at over 90% of surveyed companies reported regular personal AI tool use, while only about 40% of companies had purchased official subscriptions, meaning ungoverned adoption is outrunning governed adoption nearly everywhere. Root cause analysis reveals three primary failure modes:

  • Data fragmentation. The bulk of machine-learning effort goes to data preparation, extraction, and validation, and utilities with siloed data systems find production data quality often requires engineering effort several times the initial pilot investment.
  • Workflow integration. AI systems designed in isolation from operational workflows require significant process redesign to integrate with human decision-making.
  • Organizational inertia. Pilots managed by IT or innovation teams without embedded operations leadership produce solutions that do not address operational priorities.

Named utility deployments demonstrate achievable results. Duke Energy's self-healing grid technology avoided nearly 2.4 million customer outages and more than 11 million outage-hours across its six-state territory in 2024; Duke's AWS collaboration aims to cut grid-planning power-flow simulations from weeks to minutes. PG&E operates more than 700 HD wildfire detection cameras (600+ with AI smoke detection covering over 90% of its high-fire-threat areas), embeds AI throughout its 2026-2028 Wildfire Mitigation Plan (weather modeling, drone and aerial inspection analysis across hundreds of thousands of poles), runs grid and asset analytics on a multi-year Palantir Foundry deployment, and reports 2025 as its third consecutive year with zero major fires from its equipment, a record PG&E attributes to the layered program as a whole, not to any single AI tool. National Grid announced an operational-AI collaboration with Air Space Intelligence in May 2026 covering grid planning, resilience, outage detection, and storage siting. Exelon runs autonomous AI drone inspection (cutting inspection-planning time from an hour to 30 seconds, per Exelon) and launched POSEIDON, an AI-driven platform for real-time storm-restoration updates. In water, adoption is broadening (Bluefield Research reports growing utility-led generative AI activity), though independently verified outcome data remains scarcer than in the electric sector. Perhaps most significant structurally: EPRI's Open Power AI Consortium (Duke, Exelon, Southern Company, Constellation, Ameren, NYPA, and MISO alongside NVIDIA, Microsoft, AWS, and Oracle) is developing and benchmarking the first domain-specific generative AI models for the power sector, aiming to spare utilities from starting with generic models.

Cultural factors are as important as technical ones. Experience across industries consistently favors "operational champion" models over top-down mandates: a respected operations professional (not an IT leader) who advocates for AI adoption, directs pilot design, and guides production implementation. Organizations lacking embedded AI champions show lower adoption and higher failure rates. Additionally, in Raftelis interviews with leaders from 18 public utilities, every respondent said they personally used AI tools, yet only four of the 18 utilities had a formal written AI policy. The broader workforce mirrors this: Gallup found 52% of U.S. employees using AI at work at least a few times a year by mid-2026 (the first time above half, and nearly double the share of two years earlier), while MIT's shadow-AI finding (personal AI use reported at over 90% of surveyed companies vs. 40% with official subscriptions) shows bottom-up adoption outrunning top-down governance almost everywhere. This gap creates compliance risk and, equally, opportunity for structured governance to capture value that is already being generated informally.

Utility Primary AI Application Reported Outcome Sector
Duke Energy Self-healing grid; AWS grid-planning collaboration ~2.4M outages avoided, 11M+ outage-hours (six states, 2024); AWS collaboration targets power-flow simulations in minutes Electric
National Grid Predictive maintenance; grid-planning AI (Air Space Intelligence, 2026) Operational-AI partnership covering DER siting, resilience planning, outage mapping, and storage deployment Electric/Gas
PG&E Wildfire detection; grid data platform 700+ HD cameras (600+ AI-equipped); Palantir Foundry deployment; 3 straight years zero major equipment fires (layered program) Electric
Exelon Autonomous drone inspection; storm-restoration AI Edge-AI drone defect detection (inspection planning cut from 1 hour to 30 seconds); POSEIDON restoration-updates platform Electric
Veolia Leak detection; optimization Reports 15% leakage reduction in Prague's network via its Hubgrade AI program Water/Wastewater

NewGen Insight: The Pilot-to-Production Gap as Consulting Opportunity

The high failure rate in production scaling is not an indictment of utility AI or a sign of unsuitability. It reflects the massive engineering and organizational work required to move from controlled pilot environments to operational integration. Consulting firms that specialize in this transition (data architecture design, workflow integration, change management, and governance implementation) are now in high demand. Utilities recognize they cannot solve this alone and are investing in external expertise. The opportunity lies not in selecting AI models but in solving structural transformation barriers.

Technology Shift

The Agentic Turn: From Answering Questions to Doing Work

The most important AI development since this report's first edition is not a bigger model; it is a different kind of system. The chatbot era, in which an AI answers questions from what it absorbed in training, is giving way to the agentic era, in which an AI uses tools: it queries live databases, opens and reads actual files, writes and executes code, operates a browser, and carries out multi-step assignments with checkpoints. The commercial signals are unambiguous. Anthropic's Claude Code (an agent that works directly in files and systems rather than a chat window) became the fastest-growing product in the company's history, reaching a $1B annualized revenue run-rate within roughly six months of launch, and its general-knowledge-work sibling Cowork followed as a research preview in January 2026. Microsoft reported over 30 million paid Microsoft 365 Copilot seats in July 2026, with active agents in the Microsoft 365 ecosystem up 15x year over year. OpenAI's ChatGPT Agent and Google's Gemini Enterprise platform round out a market in which every major vendor's flagship enterprise offering is now an agent platform, not a chat window.

For utilities, the agentic turn matters because it addresses the failure mode that made executives distrust AI in the first place: the confident wrong answer. A language model is a pattern engine, not a calculator. Asked "what was our average residential bill last quarter?", a chatbot can respond two ways: compute the answer (write a query or code that sums and divides actual numbers) or recall one, producing a figure that looks like an average residential bill based on the millions of rate studies and financial documents in its training data. In a plain chat window, nothing forces the first path, and the result is the most dangerous kind of error: a wrong number with the right magnitude and the right number of decimal places. Research quantifies the difference structurally. Program-aided approaches (having the model write and run code rather than reason to an answer in text) raised math-benchmark accuracy from the mid-50s to over 70% in the foundational studies, and grounding responses in retrieved source material measurably reduces unsupported answers in well-designed systems. But grounding narrows fabrication without eliminating it: in Stanford researchers' 2024 preregistered test of 202 legal queries, purpose-built RAG-based legal AI tools still hallucinated on roughly 17-18% of queries despite vendor claims. The conclusion for any organization whose work product is numbers (rate models, cost-of-service studies, regulatory filings) is that instructions ("do not invent numbers") reduce the failure mode, but only mechanisms (forced computation, source-traceable outputs, independent verification passes) control it.

The operating standard: an auditable path for every number. The organizations getting reliable value from AI in analytical work have converged on a common discipline. Every figure in a deliverable must trace to a source cell, a live formula, or code output that can be re-run; if it cannot, it does not ship. Work is staged rather than run end-to-end in one pass: the AI first extracts required values into a checkable artifact with source references, a human verifies that artifact, and only then is the deliverable built from the verified values. And verification is performed adversarially in a fresh session (a second AI pass that locates every number in the finished product back in the source system and flags mismatches), because the session that produced a number will confidently defend it. NewGen has adopted this as firm-wide standard practice for AI-assisted analytics, and utilities should require the same of any consultant or vendor delivering AI-assisted analysis: ask not "did you use AI?" but "show me the path from every number to its source."

The honest scorecard on agents is mixed, and instructive. The gap between announcement and production is wide: while some surveys report near-universal executive enthusiasm, S&P Global's 2026 survey of AI-investing organizations found agentic AI fully integrated at just 18% and adopted in specific departments or projects at another 27% (with 27% still in trial or proof of concept), and Deloitte reports just 11% of surveyed organizations with agentic systems deployed in production. Klarna's arc is the canonical lesson: its AI assistant successfully took on roughly two-thirds of customer-service chats and cut average resolution times from 11 minutes to under 2, generating tens of millions in annual savings. The company then publicly rehired human agents for complex and emotionally sensitive cases after over-cutting staff. The pattern across documented successes is consistent: agents deliver in constrained, well-governed domains (service triage, document processing, IT operations, reconciliation, scheduling), with approval gates on consequential actions, not in open-ended autonomy. For a utility, that maps to a clear near-term agenda: agents for customer-contact triage, regulatory document preparation, work order drafting, and data retrieval across systems, with humans owning every decision that touches a customer bill, a filed document, or a switch.

Dimension Chatbot Era (2023-2024) Agentic Era (2025-2026)
How it answers Recalls plausible text from training patterns Queries systems, runs code, cites sources
Characteristic failure Confident fabrication (hallucination) Overreach: acting beyond intended scope
Primary safeguard Human review of every output Scoped permissions, audit logs, approval gates, verification passes
Data access Whatever is pasted into the chat Governed connections to live systems (MCP)
Utility example Summarize this filed testimony Pull Q2 consumption by class from the CIS, compute revenue variance in code, and draft the variance memo with cell-level sources

NewGen Insight: Structure Beats Instructions

The single most useful mental model for utility leaders evaluating AI workflows: telling a model "do not invent numbers" is an instruction; making it write code that runs against the actual billing database is a mechanism. Instructions shape behavior; mechanisms change what is physically possible. Every high-stakes AI workflow in a utility (rate analysis, regulatory filings, financial reporting) should be designed so that fabrication is structurally difficult, not merely discouraged: forced computation, extraction artifacts a human can verify, and a fresh-session audit that checks every number in the final product against source. Applied this way, AI does not merely match traditional QC; it exceeds it, because a code-driven audit checks every number in minutes where a human under deadline spot-checks a handful.

Architecture

The Connected Enterprise: MCP and the Internal AI Ecosystem

The Model Context Protocol has become the standard that turns "AI adoption" into "AI infrastructure." Launched by Anthropic in November 2024 as an open standard for connecting AI systems to tools and data (think of it as a universal port, so every system needs one connector rather than a bespoke integration per AI product), MCP was adopted by OpenAI in March 2025, then by Google and Microsoft within months. In December 2025 Anthropic donated MCP to the Linux Foundation's newly formed Agentic AI Foundation, whose platinum members include AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft, and OpenAI. That vendor-neutral governance removes the "betting on one vendor's standard" objection. The scale of adoption is remarkable: official SDKs across the major languages, more than 10,000 published servers per the Linux Foundation's donation announcement, native MCP support for Windows in public preview since November 2025, and managed MCP offerings from AWS (Bedrock AgentCore Gateway), Salesforce, Snowflake, and Databricks. A major July 2026 specification release added long-running task support and a stateless core designed for enterprise-scale deployment.

What this makes possible is an internal AI ecosystem: one governed layer through which employees get direct, conversational access to the company's data and tools. Concretely, a utility exposes each core system through an MCP server: the customer information system, GIS, the work management platform, AMI/meter data management, the document repository, and financial systems. An authorized employee can then ask a question no single system can answer ("Which districts had the most water-quality complaints in July, what is the pipe material and age profile there, and do we have crews available to accelerate the planned main replacement?"), and the AI queries the complaint database, GIS, and the work management system in sequence, synthesizes the answer, and cites what it pulled from where. The analyst who once filed a data request and waited a week now gets governed self-service in minutes. This is the same "unified data foundation" this report has always argued for. The difference is that there is now an industry-standard way to build it, and every data integration investment compounds: each new MCP server becomes available to every AI application, rather than serving a single point-to-point pipeline.

The enterprise pattern is a governed gateway, not a free-for-all. Mature deployments do not point AI tools directly at operational systems. They route everything through an MCP gateway: a central endpoint that presents each user a permission-scoped catalog of tools (the AI acting for a billing analyst simply cannot see SCADA data), injects credentials so secrets never reach the model, enforces policy, and logs every call for audit. The reference case is Block (the financial services company behind Square and Cash App), whose internal MCP-native agent "goose" connects GitHub, Jira, Snowflake, Slack, and internal systems; Block reported more than 6,500 of its 10,000+ employees using it weekly, with company-reported savings of eight to ten hours per engineer per week, in a company handling regulated financial data. Atlassian's pattern is equally instructive for utilities: its hosted MCP server executes every AI request under the requesting user's existing permissions, so the AI can never see more than the person asking. These patterns (gateway, scoped catalogs, identity-based permissions, audit) map directly onto the access-control discipline utilities already run for financial and customer systems.

The security record demands respect, not avoidance. An AI agent with access to internal data, exposure to untrusted content, and a channel to communicate externally combines what security researchers call the "lethal trifecta": the ingredients for data exfiltration via prompt injection. These risks are not theoretical: public reporting in 2025 described a cross-tenant exposure flaw in Asana's MCP implementation, and researchers demonstrated private-repository exfiltration through a GitHub MCP integration; a 2025 analysis of 1,899 open-source MCP servers flagged 5.5% for "tool poisoning" patterns, where malicious instructions hide in tool descriptions. The countermeasures are established and specific: never combine all three trifecta elements in one agent session; least-privilege, short-lived, scoped credentials; an internal allowlist of approved servers rather than open registry access; output filtering before tool results reach the model; and gateway-level audit of every call. The protocol itself has hardened substantially: mid-2025 revisions standardized OAuth 2.1-based authorization for servers that implement authentication. Note that the gateway, allowlisting, logging, and injection defenses above are implementation responsibilities, not properties the protocol provides on its own. For utilities the posture should mirror how the sector already treats operational networks: assume compromise attempts, segment aggressively, log everything. These controls are precisely what a state commission will ask about first.

For utilities, the connected-enterprise path runs through the IT domain and stops at the OT boundary. Everything described above applies to IT-domain systems: customer data, asset records, documents, work management, financials. Operational technology is different. Consistent with the December 2025 CISA-led international principles for AI in OT (discussed in the Security section), live control systems should not be MCP-exposed to general-purpose AI. The bridge pattern is the read-only historian replica: SCADA and sensor data flow one-way into an IT-domain data platform, and the MCP server exposes that, giving analysts and AI full visibility into operational history with no path back to control. Utilities can capture most of the ecosystem's value (planning, analytics, maintenance prioritization, customer operations) without touching the control plane.

System Example AI Capability via MCP Access Model Key Controls
Customer Information System (CIS) Consumption analysis, billing variance, arrears patterns Read-mostly; per-user permissions Field-level scoping; PII redaction; audit log
GIS Asset age/material queries, spatial correlation with complaints or outages Read-only CEII classification on critical layers
Work Management / EAM Backlog analysis, crew availability, draft work orders Read + human-approved writes Approval gate on any created record
AMI / MDMS Load profiling, anomaly and theft detection, forecasting inputs Read-only, aggregated by default Customer-level access restricted by role
Document repository Filing research, precedent retrieval, institutional knowledge search Read-only; user's existing permissions Inherited document ACLs
SCADA historian (replica) Operational trend analysis, event reconstruction, maintenance signals Read-only replica in IT domain One-way data flow; no path to control systems

NewGen Insight: The Ecosystem Is the New Data Integration

For a decade, "data integration" meant warehouse projects that consumed budgets and produced dashboards. The MCP ecosystem reframes the same investment: every system you connect becomes usable, in natural language, by every authorized employee and every AI application you deploy afterward. The practical starting point is deliberately modest: stand up two or three read-only MCP servers over your least sensitive, highest-value data (the document repository and GIS are common first choices), put them behind a gateway with per-user permissions and full logging, and let a pilot group work with them for a quarter. The lessons from that quarter (about permissions, data quality, and what employees actually ask) are worth more than any architecture document, and the investment carries forward into every subsequent phase.

Technology Strategy

The Model Portfolio: Frontier Cloud, On-Premise, and Small Language Models

"Which AI should we use?" is now a portfolio question, not a product question. At the frontier, the flagship cloud models (Anthropic's Claude, OpenAI's GPT-5 family, Google's Gemini) compete on a new axis: not how well they answer a question, but how long they can work. The 2025-2026 generation plans multi-step work, drives browsers and terminals, and completes tasks measured in hours rather than seconds, which is what makes the agentic capabilities described earlier real. Meanwhile the price of intelligence has collapsed: Stanford's 2025 AI Index found the cost of GPT-3.5-level performance fell more than 280-fold between late 2022 and late 2024, with per-token prices for a given capability level continuing to drop rapidly. Two consequences follow. First, frontier-model access at enterprise scale is no longer a meaningful budget line for a utility; the constraint is governance and integration, not API cost. Second, the same collapse has made an entirely different deployment model economical: running capable models on your own hardware.

Small language models have made "good enough, private, and cheap" a real strategy. An SLM is a model small enough to run on affordable local hardware, typically 1-30 billion parameters against the frontier's much larger scale. The current generation is remarkably capable: Microsoft's Phi-4 (14B parameters) posts reasoning-benchmark results competitive with far larger models; Google's Gemma 3, Alibaba's Qwen (notable for permissive Apache licensing), IBM's enterprise-focused Granite, Meta's Llama, and NVIDIA's Nemotron line give organizations a deep menu of openly licensed options. NVIDIA Research's influential June 2025 position paper argued that SLMs are "sufficiently powerful, inherently more suitable, and necessarily more economical" for many repetitive agent subtasks (the routine classification, extraction, formatting, and routing work that makes up much of what enterprise AI systems actually do), illustrating roughly 10-30x lower serving cost in one model-pair comparison. Analyst consensus has followed: Gartner predicts that by 2027 organizations' usage of small, task-specific models will be at least three times that of general-purpose LLMs. Critically, SLMs can be customized: fine-tuned on a utility's vocabulary, rate structures, asset naming conventions, and historical documents (often by distilling a frontier model's capability into a small model), then paired with retrieval so facts come from the utility's actual data rather than the model's memory.

The on-premise economics have been transformed. The first edition of this report priced self-hosted AI at $50K-500K with a dedicated ML engineering team. That was accurate in early 2026 for datacenter-class GPU deployments, and already obsolete at the entry level. Appliance-class hardware changed the arithmetic: NVIDIA's DGX Spark, a desktop unit at roughly $4,700, runs models up to ~200B parameters; AMD-based mini-workstations with 128GB of unified memory run capable open models for $1,500-2,500; and a 70B-parameter model (more capable than anything that existed three years ago) runs quantized on hardware costing less than a single crew truck's annual fuel. Serious production deployments still warrant server-class hardware and engineering discipline, but the floor has dropped from "IOU innovation budget" to "any utility's IT refresh cycle." A municipal utility can now run a private, fine-tuned model on premises for less than the cost of one seat-year of enterprise software; the constraint, as everywhere in AI, is skills and governance rather than hardware.

Why on-premise matters specifically for utilities: some data should never leave, and some environments cannot connect. Four drivers make the private tier structurally important in this sector. First, data sovereignty: critical energy infrastructure information (CEII), customer usage data, and security-sensitive asset records carry obligations that make "the data never leaves the building" the cleanest possible compliance answer. Second, operational technology: cloud APIs are structurally excluded from many OT environments: an air-gapped model doing anomaly detection or predictive maintenance on hardware that has never touched the internet aligns naturally with NERC CIP electronic security perimeters and the segmentation practices of NIST SP 800-82 (alignment, not automatic compliance; that depends on the complete architecture, controls, and evidence), and it is the architecture the December 2025 CISA-led OT principles point toward. Third, customization: a small model fine-tuned on utility-specific language and data outperforms a generic giant on narrow recurring tasks. Fourth, cost and latency predictability for high-volume workloads (meter data classification, document routing, transcript summarization), where per-call cloud pricing compounds and a local model runs at effectively fixed cost.

The architecture that wins is hybrid: route by task, escalate by need. The dominant 2026 enterprise pattern is tiered routing: a lightweight local model classifies each request and handles the routine majority of traffic, while escalating genuinely hard problems (long-horizon reasoning, novel analysis, complex drafting) to frontier cloud models. Published routing research reports task-specific cost reductions in the 40-85% range, with small but measurable quality tradeoffs that depend on the workload. For a utility, the mapping is intuitive: the frontier tier drafts rate case testimony and reasons through a complex interconnection analysis; the local tier classifies thousands of customer contacts, extracts fields from scanned records, summarizes shift logs, and screens meter anomalies: privately, at fixed cost, around the clock. The two tiers are complements, not competitors, and the portfolio should be an explicit governance artifact: which classes of task and data run where, and why.

Workload Right Tier Why
Rate case strategy, complex analysis, long-horizon agentic work Frontier cloud (enterprise agreement) Reasoning depth and task-horizon capability unmatched by small models
Document drafting, research, cross-system Q&A via MCP Frontier cloud Benefits from strongest general capability; IT-domain data under enterprise data protections
High-volume classification, extraction, routing, summarization SLM (local or cloud) 10-30x cheaper; fine-tunable to utility vocabulary; quality parity on narrow tasks
CEII, security-sensitive asset data, unredactable customer data On-premise SLM Data never leaves utility infrastructure; cleanest compliance posture
OT-adjacent analytics (anomaly detection, predictive maintenance in segmented environments) Air-gapped SLM appliance Satisfies NERC CIP / NIST SP 800-82 segmentation; aligns with CISA OT principles

NewGen Insight: The Portfolio Is Now Open to Everyone

The most underappreciated shift of 2025-2026 is who can play. When private AI meant half a million dollars and an ML team, it was an IOU-only conversation. A $5,000 appliance running an openly licensed, fine-tuned model, behind the utility's own firewall and trained on its own documents, puts genuinely private AI within reach of a municipal utility or rural cooperative. The failure mode to avoid is symmetrical: small utilities should not assume private AI is beyond them, and large utilities should not build on-premise empires for workloads that enterprise cloud agreements already handle securely and more capably. Match the tier to the task and the data classification, write the mapping down, and revisit it annually; model economics are still moving fast enough that this year's right answer is next year's starting point.

Strategy

Lessons from Other Industries

Healthcare established a regulatory and operational template for privacy-driven AI. The Health Insurance Portability and Accountability Act (HIPAA) created mandatory privacy and security frameworks that, counterintuitively, enabled scalable AI deployment. HIPAA established three critical capabilities that utilities lack. First, Business Associate Agreements (BAAs) create contractual frameworks for data sharing with external vendors while maintaining liability and accountability. Second, de-identification protocols (HIPAA Safe Harbor method) enable organizations to use real data for model training without exposing personally identifiable information. Third, federal cloud-security frameworks (FedRAMP, established 2011 and modernized in 2024) gave regulated organizations a template for vetting cloud AI platforms, though FedRAMP itself governs federal agency procurement and is distinct from HIPAA compliance. These mechanisms were controversial when introduced but are now recognized as enablers of rapid innovation. Healthcare organizations from Mayo Clinic (which has built a multi-petabyte de-identified clinical data platform) to Cleveland Clinic (deploying AI scribes at scale to cut clinician documentation time) scaled AI because they solved privacy and compliance architecturally, not organizationally. FDA's running (non-comprehensive) list of AI-enabled medical devices shows 333 entries with 2025 final decisions versus 235 for 2024 (roughly 42% growth), reflecting mature regulatory pathways for AI authorization. Utilities face an opportunity to adopt healthcare's playbook: establish contractual and architectural privacy frameworks before scaling deployments.

Banking demonstrates decade-long transformation and robust model governance. JPMorgan Chase's COIN (Contract Intelligence) platform reviews commercial loan agreements that previously consumed roughly 360,000 lawyer-hours annually, work now done in seconds; the bank had attributed 80% of its loan-servicing errors to contract-interpretation mistakes, which is exactly the failure mode COIN targets. Bank of America's Erica AI assistant has served roughly 25 million clients and surpassed 3 billion cumulative interactions since its 2018 launch. Capital One began its technology transformation in 2012, exited its last data center in 2020, and describes that modern stack as the foundation for its subsequent ML and AI capabilities. Banking's AI success reflects institutional adoption of Model Risk Management frameworks. The Federal Reserve and Office of the Comptroller of the Currency published SR 11-7, "Guidance on Model Risk Management," in 2011; in April 2026 the federal banking agencies superseded it with revised guidance (SR 26-2) that carries the same principles forward. The framework, originally applied to mortgage valuation, interest rate, and credit risk models, now covers AI tools whenever they meet its model definition. Model Risk Management requires:

  • Model inventory and documentation
  • Independent validation of models before production deployment
  • Ongoing monitoring of model performance
  • Rapid decommissioning when performance degradation is detected
  • Governance committee oversight

Banks using material models are expected to maintain model-risk governance proportionate to their size and complexity. Community-bank industry surveys consistently identify regulatory scrutiny as a top barrier to AI adoption, mirroring the current utility concern. Yet nationally, banks are scaling AI because regulatory frameworks are mature. Utilities can adopt banking's governance model directly: establish Model Risk Management functions, implement independent validation processes, and create governance oversight.

Government agencies demonstrate production-scale AI with compliance frameworks. The U.S. Treasury reported preventing and recovering over $4B in fraud and improper payments in fiscal 2024 across multiple processes, with machine-learning check-fraud detection accounting for roughly $1B of that total. The Social Security Administration has documented AI-supported workflows and workload processing, including planned use of AI to identify determination-ready disability claims. The General Services Administration launched the "FedRAMP 20x" pathway in March 2025 to accelerate cloud authorization; its first pilot authorizations were issued within four months, and GSA later reported review times of roughly five weeks versus over a year previously. The Office of Management and Budget published M-25-22 in 2025, "Advancing Governance, Innovation, and Risk Management for Agency Use of Artificial Intelligence," establishing acquisition guidelines and governance requirements for federal agencies. OMB M-25-22 is essentially a government procurement template that private utilities can adapt: it specifies data governance, security requirements, model documentation, audit trails, and human-in-the-loop decision requirements. Additionally, Gallup found 43% of public-sector employees using AI at least a few times a year as of Q4 2025 (21% daily or several times weekly), suggesting government agencies are scaling AI adoption despite regulatory constraints. The City of Phoenix's myPHX311 provides a web and mobile portal for city-service requests, including water and waste, with a virtual assistant handling routine inquiries. These government examples are particularly relevant to utilities because government agencies operate under comparable privacy, security, and accountability constraints.

Manufacturing established digital twin standards and O&M savings benchmarks. GE Vernova reports more than $1.6B in customer production and mechanical losses avoided through its SmartSignal predictive analytics, whose managed-services team monitors more than 7,000 critical assets. Siemens says its Senseye predictive maintenance platform can reduce unplanned machine downtime by up to 50%. MarketsandMarkets projects the digital twin market to grow from $21.14B (2025) to $149.81B (2030), reflecting industrial adoption momentum. Manufacturing's lessons for utilities are twofold. First, digital twins (virtual replicas of physical systems that receive real-time data and enable predictive modeling) are applicable to all utility infrastructure. Second, manufacturing established clear ROI metrics for digital twins: unplanned downtime, maintenance cost, asset life, and scheduling efficiency. Effect sizes depend on the asset and deployment, but utilities can adopt the measurement framework directly. Additionally, in December 2025 the U.S. Cybersecurity and Infrastructure Security Agency, with the NSA's AI Security Center, the FBI, and cybersecurity agencies from Australia, Canada, Germany, the Netherlands, New Zealand, and the United Kingdom, jointly published "Principles for the Secure Integration of Artificial Intelligence in Operational Technology." The document, covering machine learning, large language models, and AI agents in OT contexts, warns that AI models and their data can be manipulated to bypass safety guardrails, and organizes its guidance around four principles: understand AI, consider AI use in the OT domain, establish AI governance and assurance frameworks, and embed safety and security practices into AI-enabled OT systems. It points squarely toward network segmentation between IT and OT, rigorous pre-deployment testing, human oversight of consequential actions, and hard fail-safes for any autonomous function. This joint guidance is required reading for utilities planning grid or water treatment automation.

The BCG 10/20/70 rule travels well across industries. Across healthcare, banking, government, and manufacturing, Boston Consulting Group's analysis of AI programs finds that roughly 10% of the challenge is algorithms, 20% is technology and data, and 70% is people and process. It is a change-management heuristic rather than a precise budget formula, but the implication for utilities is direct: selecting models is the smallest part of the work. Weighting resources heavily toward change management, training, workflow redesign, and organizational alignment (with most of the remainder going to data architecture, integration, and quality) matches how successful adopters actually spend. Most organizations invert this in practice, putting the bulk of their budget into technology.

Industry Key Regulatory Framework Governance Pattern Scale Achievement Relevance to Utilities
Healthcare HIPAA + FedRAMP Data privacy architecturally solved 295 FDA-authorized AI devices (2025) Privacy framework template
Banking SR 11-7 / SR 26-2 Model Risk Management Independent validation required 360K lawyer-hours automated (JPMorgan COIN) Governance model template
Government OMB M-25-22 + FedRAMP 20x Procurement-driven governance ~$1B ML-attributed fraud prevention (Treasury) Acquisition framework template
Manufacturing CISA joint Principles for AI in OT (Dec 2025) Safety-first; segmentation; human oversight $1.6B avoided losses attributed over 15+ yrs (GE) OT/SCADA integration template

Critical Note on OT/SCADA AI: The December 2025 joint CISA/NSA/international "Principles for the Secure Integration of AI in Operational Technology" is the most authoritative guidance yet for utility AI touching control systems. It is advisory rather than regulatory, but utilities planning any AI application affecting SCADA, grid operations, or treatment systems should treat its core practices as non-negotiable: isolated development and testing, strict IT/OT network segmentation, hard fail-safes on any autonomous function, and human review of consequential actions. Expect these practices to be incorporated into state PUC guidance and insurer expectations as frameworks mature.

Implementation

The AI Deployment Spectrum

AI deployment exists on a spectrum from simple to complex, with proportional cost, risk, and capability trade-offs. Utilities typically follow a progression: starting with team plans, advancing to API integration, then expanding to MCP servers and tool integration, and finally implementing self-hosted models or air-gapped deployments for mission-critical systems. Understanding this spectrum prevents both premature over-engineering and under-investment. (Cost and timeline figures in this section are illustrative planning ranges, not vendor quotes or industry benchmarks; verify current pricing and scope-specific estimates before budgeting.)

Tier 1: Team Plans ($20-30 per user per month). Claude Team, Microsoft 365 Copilot, ChatGPT Team, and similar offerings provide enterprise-grade AI capabilities without infrastructure investment. Distinguishing features typically include enterprise data protection (no training on user data), single sign-on and role-based access control, audit logging, and current SOC 2 Type II reports, though entitlements vary by vendor and plan, so verify each capability for the specific product. Deployment is immediate: utilities activate licenses and distribute to pilot groups. No IT infrastructure, no model training, no data pipeline configuration. Ideal use cases include document drafting, research synthesis, meeting summaries, code assistance, regulatory analysis, and customer inquiry categorization. Payoff timeline is 2-4 weeks. Most utilities should begin here, testing AI workflows with minimal risk. Tier 1 identifies use cases, builds user familiarity, and creates organizational appetite for more sophisticated deployments. Organizations often underestimate Tier 1 impact: self-reported time savings at scale are commonly substantial, though realized economic ROI depends on whether recovered time is actually redeployed; measure both.

Tier 2: API Integration (usage-based, priced per million tokens). This tier connects AI models (via API calls) to existing utility systems: CIS (Customer Information Systems), GIS, work management platforms, billing systems, and read-only SCADA historian replicas on the IT side of the OT boundary. Implementation requires developer resources: API endpoint design, prompt engineering, data pipeline configuration, and integration testing. Typical projects span 2-4 months. Tier 2 unlocks use cases requiring data contextualization: automated regulatory filing preparation (API pulls utility data, AI generates filing, humans review), customer service routing (API characterizes inbound inquiry, AI classifies urgency/type, routes to appropriate team), data analysis and insight generation (API aggregates operational data, AI identifies patterns, generates reports). Tier 2 requires data governance: what data is transmitted to the AI model? Most mature utilities implement data anonymization, extracting only necessary fields. Costs are proportional to token volume, priced per million tokens: 100K calls averaging 2,000 tokens each is 200M tokens a month, which runs from a few hundred to a few thousand dollars depending on the model tier and input/output mix. The practical point: API costs scale with usage and are rarely the binding constraint. Risk is moderate: API breaches expose operational data, mitigated through network segmentation, encryption, and role-based data access controls.

Tier 3: MCP Servers and Tool Integration. The Model Context Protocol (now a Linux Foundation-governed open standard supported by Anthropic, OpenAI, Microsoft, Google, and AWS, covered in depth in The Connected Enterprise section above) allows AI models to directly query operational databases and tools. Instead of utilities manually preparing data and passing it via API, MCP servers let AI request specific data dynamically under per-user permissions. A water utility implementing Tier 3 would expose: (1) treatment plant historian data (turbidity, flow, chemical levels from a read-only replica), (2) the customer complaint database, (3) GIS (pipe location, age, material), (4) the work management system (pending maintenance, crew availability). The AI can then synthesize across systems: "Given the elevated turbidity reading at Plant A, the 12 customer complaints about water clarity in the Southwest district, and the pipe age data showing 40-year-old cast iron in that area, I recommend prioritizing the planned main replacement in the Southwest zone next quarter." What was an emerging standard in early 2026 is now mainstream: managed MCP offerings from AWS, Snowflake, Databricks, and Salesforce mean much of the server-side work can be bought rather than built, and implementation timelines are compressing from 3-6 months toward weeks for common systems. Risk is moderate to high and concentrated in governance: authentication, authorization, allowlisting, and audit logging, as detailed earlier.

Tier 4: Self-Hosted Models ($5K-500K depending on scale). Organizations run openly licensed models (Llama, Qwen, Gemma, Phi, Granite, Mistral) on internal infrastructure, eliminating external data transmission. This tier has been transformed by the SLM and hardware developments covered in The Model Portfolio section: appliance-class hardware from roughly $2K-5K runs highly capable small models, while serious multi-user production deployments on server-class GPUs still run $50K-500K. The realistic pattern for most utilities is no longer "build a frontier-model replacement" but "run fine-tuned small models for specific sensitive or high-volume workloads," which requires competent IT engineering and model-lifecycle discipline, but no longer a dedicated ML research team. Payoff timelines have compressed accordingly: a scoped single-workload deployment (document classification over sensitive records, for example) can prove out in 3-6 months. This tier is now appropriate for any utility with defined data sovereignty needs, not just the federally mandated few.

Tier 5: Air-Gapped Deployments ($200K-2M+). Physical isolation from the internet places models on self-contained networks with no external communication. This tier is appropriate for nuclear utilities, critical grid control centers, and organizations under specific federal security mandates. Implementation requires physically segregated network infrastructure, pre-downloaded models and data, manual security updates, and restricted personnel access. Risk is high operationally: any system failure requires on-site remediation, and models become stale as training knowledge ages, though the modern SLM ecosystem eases this, since updated open models can be periodically transferred in through controlled media. The December 2025 CISA-led principles for AI in OT point to exactly this posture for AI functions in control environments: segmentation, isolated testing, hard fail-safes, and human review of consequential actions.

Tier Monthly Cost Range Setup Time IT Complexity Security Level Best For
1: Team Plans $20-30/user 1-2 weeks Minimal High (SOC 2, ISO 27001) Document work, research, analysis
2: API Integration Usage-based (per-token) 2-4 months Moderate High (network segmentation required) Automated data analysis, customer routing
3: MCP Servers $5K-15K setup + $2K/month 3-6 months High High (auth + audit logging) Real-time operations, integrated analytics
4: Self-Hosted Models (SLMs) $5K-500K initial (scale-dependent) 3-12 months Moderate-High Very High (data never leaves) Fine-tuned SLMs; sensitive/high-volume workloads
5: Air-Gapped $200K-2M+ initial 12-24 months Very High Maximum (isolated) Critical SCADA, nuclear, federal mandates

NewGen Insight: The Typical Progression

Most utilities should plan a progression: Year 1 focuses on Tier 1 (team plans) to build organizational AI capability and identify high-impact use cases. Year 2 implements Tier 2 (API integration) for production applications with defined workflows and stands up the first read-only MCP servers, a timeline that has compressed since this report's first edition, as managed MCP offerings mature and the standard's Linux Foundation governance removes vendor-risk objections. Year 2-3 expands the MCP ecosystem toward integrated operations, and Tier 4 (self-hosted SLMs) now enters the conversation earlier for utilities with defined sovereignty needs, at appliance-scale cost. Tier 5 remains reserved for control-environment and federally mandated contexts. The false choice is between "do nothing until we have perfect governance" and "immediately deploy to Tier 5 for maximum security." The winning strategy is progressive adoption with appropriate governance at each tier.

Strategic Differentiation by Utility Size: The optimal starting tier depends on existing IT maturity. Large IOUs with dedicated data teams and integrated SCADA/GIS/CIS infrastructure can move directly to Tier 2 (API integration) for specific high-ROI use cases like demand forecasting or outage prediction, compressing Year 1 and Year 2 into parallel tracks. Mid-sized utilities with partial integration should follow the canonical path: 6-12 months of Tier 1 (non-critical domains like billing analysis) while assembling an operational integration roadmap in parallel. Small utilities with minimal data integration should remain at Tier 1 for 12-18 months, building organizational capability first; rushing to APIs creates technical debt and pilot failures. This tiered approach reduces risk while enabling appropriate deployment velocity for each utility's constraints.

Risk Management

Security: What Utility Leaders Need to Know

Enterprise AI services solve privacy architecturally; consumer tools do not. The fundamental distinction between consumer AI (ChatGPT free tier, Google Gemini free) and enterprise AI (Claude Team, ChatGPT Team, Azure OpenAI) is data usage policy. Consumer tools may use inputs for model training depending on product and settings. Enterprise services contractually guarantee that user data is not used for model training, with retention limited to defined contractual windows (zero-data-retention configurations exist for some API services). Training exclusion and retention are separate terms; verify both. This is not a minor distinction; it is architecturally meaningful. Healthcare organizations using HIPAA-covered data, financial institutions using customer account numbers, and government agencies using classified information all require enterprise-grade guarantees. Most utilities using team plans for routine analysis do not transmit sensitive data, making consumer tools acceptable for some use cases. However, utilities planning production deployment should exclusively use enterprise services with contractual guarantees.

SOC 2, FedRAMP, and ISO 27001 are auditable compliance frameworks. A SOC 2 Type II report (issued by independent auditors) examines the design and operating effectiveness of a service provider's security controls over the report's stated review period; obtain and inspect the actual report and its scope. FedRAMP (Federal Risk and Authorization Management Program) is a government-led authorization process for cloud services used by federal agencies, applying NIST SP 800-53 controls with continuous monitoring at defined impact levels. GovRAMP (formerly StateRAMP, rebranded in 2025) provides a parallel NIST-based framework for state, local, tribal, and education organizations. ISO 27001 is an international standard for information security management; it requires annual third-party audits. When evaluating AI vendors or cloud platforms, utilities should require: (1) SOC 2 Type II certification (at minimum), (2) FedRAMP certification if the vendor serves federal agencies, (3) ISO 27001 if international operations are planned. Request auditor reports and review the specific controls each framework validates. Enterprise agreement terms should specify that the vendor will maintain these certifications throughout the contract.

Data handling tiers reduce risk through classification. Healthcare's approach to data classification is instructive: public data (published research, general patient education) can be shared broadly; internal data (employee communications, general operational data) requires employee confidentiality agreements; confidential data (customer financial information, insurance data) requires contracts and technical safeguards; and restricted data (medical records, genetic information, psychotherapy notes) requires HIPAA compliance and specific authorization. Utilities should implement similar tiering:

  • Public: utility service territories, rate schedules, outage maps
  • Internal: employee communications, general budget information
  • Confidential: customer names and addresses, billing data, some operational metrics
  • Restricted: customer usage patterns, critical infrastructure vulnerabilities, some SCADA sensor data

This classification then determines what data can be transmitted to Tier 1 systems (public and internal only), Tier 2 systems (public, internal, confidential with anonymization), or higher tiers (all data with encryption and segmentation). Implementing data classification is a modest, high-leverage effort that reduces the most common data-exposure scenarios.

The paranoia is real but solvable. Utility leaders frequently express concern that external AI systems could expose proprietary or sensitive operational data. This concern is not unfounded: data breaches at cloud service providers do occur, and in rare cases, human employees at AI companies have accessed customer data. However, other industries have solved this problem. Healthcare processes billions of patient records annually through cloud-based AI systems with documented HIPAA compliance. Banking routes customer financial data through AI systems daily. Government agencies process sensitive federal information through FedRAMP-authorized AI services. The solution is architectural and contractual:

  • Implement data classification (described above)
  • Select vendors with appropriate certifications
  • Implement data anonymization where possible
  • Segment networks to prevent lateral access
  • Maintain audit logs
  • Conduct annual security assessments

None of these steps is novel; all are standard IT security practice. Utilities should expect their Chief Information Security Officer or external cybersecurity consultants to validate vendor selection and implementation approaches.

OT/SCADA systems require distinct security approaches. Operational Technology (OT) systems controlling power flow, water treatment, or gas pressure are physically dangerous if compromised. The authoritative reference is the December 2025 "Principles for the Secure Integration of Artificial Intelligence in Operational Technology," published jointly by CISA, the NSA's AI Security Center, the FBI, and cybersecurity agencies from seven allied nations. Its four principles (understand AI, consider AI use in the OT domain, establish AI governance and assurance frameworks, and embed safety and security practices) translate into practices utilities should treat as mandatory even though the document is advisory: isolated development and testing environments, strict network segmentation between IT and OT, hard fail-safes on any autonomous function, and human review of consequential actions. The guidance is notable for warning explicitly that AI models and their training data can be manipulated to bypass safety guardrails, a threat class traditional OT security programs do not cover. These practices will likely be incorporated into state regulatory guidance and insurer expectations. Utilities planning any AI application affecting SCADA should consult CISA's published OT guidance and their regional CISA Cybersecurity Advisor, and should hire external OT security consultants to validate architectures.

The IT/OT Boundary Is Decisive for Cloud AI Deployment. Cloud AI platforms (Claude Team, Azure OpenAI, AWS Bedrock) operating under SOC 2 Type II and FedRAMP certifications are appropriate for IT systems processing billing data, customer service, analytics, and regulatory analysis. These platforms have strong, independently audited security programs appropriate for sensitive IT-domain data when proper enterprise agreements and configurations are in place. However, this approval does not extend to OT systems. Cloud AI services connected directly to SCADA systems, demand response, or grid dispatch create unacceptable risk: cloud service dependencies could cascade to operational failures, model drift in safety-critical contexts is not remediable, and no state PUC has established regulatory precedent for cloud-based OT AI. The solution is architectural clarity: designate systems as either IT-domain (appropriate for cloud AI) or OT-domain (appropriate for air-gapped or dedicated on-premise deployment with explicit regulatory approval). Many utilities will find that their highest-value AI applications (demand forecasting, customer service, regulatory filing support) are IT-domain applications that benefit immediately from cloud AI, while OT integration follows on a longer timeline after regulatory precedent is established.

NewGen Insight: Security as Enabler

Utility leaders sometimes view security requirements as constraints that slow AI adoption. The inverse is true. Robust security governance enables accelerated adoption: utilities that establish clear security frameworks, implement vendor requirements, and maintain audit compliance move faster to production deployment because they eliminate retroactive security concerns. Utilities that defer security planning find themselves defending deployments and backtracking to add controls. In NewGen's experience, front-loaded security investment (weeks of governance work, not months) pays for itself many times over in faster, smoother deployment afterward.

Execution

The AI Adoption Roadmap for Utilities

Successful AI adoption follows a phased roadmap emphasizing governance, data, and organizational capability before algorithmic complexity. This roadmap is based on patterns observed at utilities that successfully scaled AI (Duke Energy, National Grid, PG&E) and adapted from healthcare and banking transformation models. The roadmap spans approximately 24 months from initiation to production scaling. Cost, savings, and timeline figures throughout this section are illustrative planning ranges drawn from project experience and industry reporting; calibrate them to your utility's size, procurement rules, and starting point.

Phase 1: Foundation (Months 1-3). The first phase establishes governance, assesses data readiness, and initiates low-risk deployments. Activities include:

  • Establish an AI Governance Committee with cross-functional representation (CIO, Chief Operations Officer, Chief Regulatory Officer, Chief Legal Officer, Finance, Security). The committee meets monthly and has decision authority on AI policy, vendor selection, and escalation issues.
  • Develop an AI Policy and Acceptable Use Guidelines addressing data classification and handling, appropriate use of AI tools, confidentiality and privacy requirements, prohibition on autonomous decisions in critical contexts, and audit requirements. Most utilities complete this step in 3-4 weeks using templates from similar organizations or external counsel.
  • Conduct a data readiness assessment across core systems (Customer Information System, SCADA, GIS, work management, AMI, accounting) to identify what data is available, where it resides, data quality status, and integration barriers. This typically costs $25K-50K and requires 4-6 weeks.
  • Deploy team plan licenses (Claude Team, Microsoft Copilot) to a pilot group of 30-50 employees in research, analysis, and documentation roles, with two hours of training covering appropriate use, confidentiality requirements, and example workflows.
  • Identify and train 2-5 "AI Champions" per major department (Operations, Customer Service, Planning, Finance, Regulatory). Champions are respected operational leaders (not IT staff) who drive adoption within their departments, with 8 hours of training plus monthly community-of-practice meetings.
  • Establish baseline metrics for the outcomes you plan to improve: cost per customer contact, mean time to respond to service requests, annual outage hours, treatment plant efficiency, demand forecasting error, maintenance planning time, and regulatory filing preparation time.

Phase 2: Quick Wins (Months 4-8). This phase identifies specific use cases and measures early ROI. Activities include:

  • Document processing and regulatory filing assistance. AI drafts initial versions of annual reports, cost-of-service analyses, and regulatory filings; humans review and finalize. Typical outcome: 30-40% reduction in preparation time.
  • Research and analysis acceleration. AI synthesizes regulatory changes, competitive benchmarks, and industry trends; analysts review for accuracy and relevance. Typical outcome: 40-50% acceleration in research scope definition.
  • Customer service triage. AI pre-processes inbound contacts (phone, email, chat), categorizes urgency and type, and routes to the appropriate team. Typical outcome: 25-35% reduction in first-contact routing time, improved accuracy.
  • Meeting documentation and work order assistance. AI generates meeting notes, creates work order drafts, and summarizes key decisions. Typical outcome: 15-20% reduction in administrative overhead.
  • Measure and communicate early results to the executive team, documenting time saved, cost reduction, and quality metrics. Early wins build organizational confidence and support for continued investment.

Phase 3: Integration (Months 9-18). This phase expands AI into operational systems through API integration and data pipelines. Activities include:

  • Implement API connections to CIS for automated regulatory reporting and customer analysis
  • Connect AI to work management and GIS systems to enable predictive maintenance and vegetation management pilots
  • Implement MCP servers behind a governed gateway connecting AI to GIS, document repositories, and a read-only SCADA historian replica for operational analytics, keeping the OT control plane isolated per the CISA-led principles
  • Deploy custom tools for rate analysis, demand forecasting, and cost-of-service modeling
  • Launch predictive maintenance pilots on highest-value asset categories (critical pump stations, major transmission lines, large treatment processes)
  • Expand training to all employees in pilot departments; create tiered certification (basic AI literacy, practitioner certification, advanced certifications by domain)
  • Conduct a mid-program review: measure adoption rates, cost savings, risk incidents, and ROI against baseline metrics, and adjust the roadmap based on results

Phase 4: Transformation (Months 18+). This phase scales successful deployments and pursues strategic applications. Activities include:

  • Roll out predictive analytics across all major asset categories and operational areas
  • Implement digital twin technology for complex treatment processes or transmission networks
  • Deploy autonomous systems for bounded processes (predictive maintenance scheduling, demand response optimization, treatment chemical dosing) with human-in-the-loop and kill-switch requirements
  • Develop AI-enhanced rate design and cost-of-service analysis capabilities
  • Engage in industry leadership through AWWA, WEF, and EEI participation; utilities with mature AI capabilities are increasingly asked to lead peer learning and standard-setting
  • Plan for the rate case: quantify AI-driven operational and financial improvements, document investments and benefits, and work with regulatory advisors to establish precedent for AI cost recovery in rate base
  • Continuously assess emerging AI capabilities, and plan for Tier 3 (MCP servers) or Tier 4 (self-hosted models) if appropriate for strategic data sovereignty or specialized domain requirements

NewGen Insight: The Importance of Phase Sequencing

Utilities often want to compress this timeline or skip early phases. The most common pressure is to move directly to high-value use cases (predictive maintenance on $10M assets) without establishing governance or data foundations. This consistently fails. The utilities that succeeded (Duke, National Grid) completed Phases 1-2 in 4-6 months but did not rush. They built governance, identified quick wins, built organizational confidence, and then scaled. Skipping to Tier 3 or 4 deployment before completing Phase 2 almost always results in 12-18 month delays, cost overruns, and governance failures. The roadmap should be viewed as a floor, not a ceiling: utilities with strong data governance and IT maturity may compress timelines, but skipping steps is not recommended.

Next Steps

Recommendations for Utility Leaders

Based on the regulatory landscape, adoption patterns, cross-industry lessons, and implementation frameworks described above, NewGen offers the following recommendations:

  1. Start immediately despite regulatory uncertainty. No state PUC has issued guidance; first movers will shape the conversation rather than adapting to it. Utilities implementing governance now establish competitive advantage and industry leadership. Waiting for regulatory guidance is a strategic error.
  2. Establish an AI Governance Committee before deploying any tools. Cross-functional governance (CIO, COO, Chief Regulatory Officer, Legal, Finance, Security) provides accountability, risk management, and decision authority. This committee should meet monthly and oversee all AI initiatives. Without governance, AI adoption becomes chaotic and increases compliance risk.
  3. Invest in data readiness before selecting algorithms. Conduct a comprehensive data assessment identifying what data exists, where it resides, data quality status, and integration barriers. The bulk of machine-learning effort is data preparation; utilities without integrated data foundations will struggle to scale AI beyond pilots. Data architecture is the binding constraint.
  4. Begin with team plan licenses at Tier 1 deployment. Claude Team, Microsoft Copilot, or equivalent offerings provide enterprise-grade capability with minimal risk and immediate deployment. Use these tools for 60-90 days to identify high-impact use cases, build employee familiarity, and establish baselines. This is the lowest-risk pathway to organizational learning.
  5. Plan a model portfolio, not a model choice. Map task classes and data classifications to deployment tiers: frontier cloud models for complex reasoning and agentic work over IT-domain data, small on-premise models for routine high-volume tasks and sovereignty-bound data, air-gapped deployments for OT-adjacent functions. Write the mapping down as a governance artifact and revisit it annually; model economics are moving fast enough that the right answer changes.
  6. Require an auditable path for every AI-produced number. Adopt the compute-don't-recall standard for analytical work: figures in filings, studies, and board materials must trace to a source cell, live formula, or re-runnable code output, with staged human checkpoints and a fresh-session verification pass before delivery. Hold consultants and vendors to the same standard: ask to see the path from number to source, not just the deliverable.
  7. Identify AI Champions in every major department. Clinical champions (respected operational leaders, not IT staff) drive adoption within departments at 3-4x higher rates than top-down mandates. Invest in champion training and create regular community-of-practice forums. Champions are force multipliers for organizational transformation.
  8. Engage your regulator proactively on governance. Share your AI governance framework, data handling policies, and risk management approaches with your state PUC before deploying production AI. Regulators value transparency, and utilities that demonstrate responsible governance enter future proceedings from a far stronger position. Conversely, utilities that deploy AI without regulator communication face retroactive oversight risk.
  9. Budget by BCG's rule of thumb: roughly 70% people and process, 20% technology and data, 10% algorithms. Most utilities invert this allocation. Rebalance resources toward change management, training, workflow redesign, and organizational restructuring. Algorithm selection is the smallest component of AI success; organizational transformation is the largest.
  10. Measure everything before and after AI deployment. Establish baseline metrics (cost per customer contact, outage hours, planning time, forecasting error, etc.). For each AI implementation, measure the specific outcome expected. Only deploy AI that demonstrably improves defined metrics. Avoid deployment for deployment's sake.
  11. Plan for the rate case now. Document all AI investments, allocate costs to appropriate departments, and measure benefits. When rate cases are filed 2-3 years from now, utilities need documented precedent for AI cost recovery. Regulatory commissions will require evidence that AI investments benefited customers. Begin collecting this evidence immediately. As a general pattern, larger utilities with mature data infrastructure are best positioned to propose AI investments as rate base capital backed by documented success metrics and performance gates; mid-sized utilities will typically want a measured pilot with specific ROI evidence before filing; smaller utilities will often fund AI through operating budgets while the evidence base builds. Rate treatment is jurisdiction- and fact-specific; engage regulatory counsel and commission staff before committing capital on an assumption of recovery. The burden of proof is higher where precedent does not exist, and utilities that build documented evidence during pilots will face less rate case resistance than utilities relying solely on peer examples.
  12. Adopt a hybrid build-partner-buy model for AI capability. Utilities cannot realistically compete with tech-sector talent compensation; recruiting and retaining internal data science teams is structurally difficult for regulated utilities. Instead, employ a differentiated strategy: (1) Buy proven solutions for routine functions (billing analysis, customer service automation, meter reading optimization) from established vendors. (2) Partner with consultants and AI services firms for domain-specific projects where the utility provides operational expertise and the partner provides AI methodology. (3) Build internal capability selectively only for applications that are (a) truly unique to the utility, (b) material in recurring value, and (c) defensibly strategic. For most utilities, the "build" category is nearly empty. This approach enables faster deployment, reduces talent retention risk, and aligns with evidence (including MIT's 2025 interview-based research) suggesting internally built enterprise AI succeeds materially less often than vendor or hybrid approaches. Learn from other industries without reinventing the wheel: healthcare solved privacy governance through HIPAA and FedRAMP, banking through Model Risk Management frameworks, manufacturing through digital twin standards. These frameworks are mature and directly applicable to utilities.
Reference

Appendix A: Utility AI Use Case Matrix

ROI figures below are indicative ranges compiled from vendor and industry reporting, not audited benchmarks; validate against your own baseline before using them in a business case.

Use Case Sector Maturity Documented ROI Representative Deployments
Leak Detection Water/Wastewater Tier 1 (Mature) $2-5M annual savings per utility (indicative) Veolia (reports 15% leakage reduction in Prague); satellite programs (Asterra) across many utilities
Demand Forecasting All sectors Tier 1 (Mature) 5-15% forecasting error reduction PG&E, National Grid, major municipal utilities
Vegetation Management Electric/Gas Tier 1 (Mature) Up to 50% reduction in manual patrol effort (FirstEnergy LiDAR program) FirstEnergy, Duke Energy, AiDash/Overstory deployments
Customer Service Automation All sectors Tier 1 (Mature) 25-35% first-contact resolution improvement (indicative) Duke Energy; Bank of America Erica (3B+ cumulative interactions)
Predictive Maintenance All sectors Tier 1 (Mature) $5-15M annual savings per large utility National Grid, Duke Energy, manufacturing sector
Water Quality Monitoring Water/Wastewater Tier 2 (Growing) Not yet well quantified Emerging deployments (UK, Australia, U.S.)
Grid Optimization Electric Tier 2 (Growing) 10-20% efficiency gains (estimated) Duke Energy (self-healing grid), emerging deployments
Wildfire Detection Electric Tier 3 (Emerging) Not yet quantified PG&E (700+ HD cameras; 3 straight years zero major equipment fires, layered program)
Treatment Optimization Water/Wastewater Tier 2 (Growing) $1-3M annual savings (estimated) Veolia and UK utilities (vendor-reported)
Document Processing All sectors Tier 1 (Mature) 30-40% reduction in processing time JPMorgan COIN (360K lawyer-hours), government agencies
Reference

Appendix B: Deployment Tier Comparison

Dimension Tier 1 Tier 2 Tier 3 Tier 4 Tier 5
Monthly Cost $20-30/user Usage-based (per-token) $5K-15K setup + $2K/mo $5K-500K initial (appliance to server-class) $200K-2M+ initial
Setup Time 1-2 weeks 2-4 months Weeks-6 months 3-12 months 12-24 months
Data Exposure Moderate (vendor-hosted) Moderate (API-based) Moderate (dynamic queries) Low (internal) Very Low (isolated)
IT Complexity Minimal Moderate High Very High Very High
Autonomy Level None (interactive only) Automated workflows Integrated tools Specialized models Full autonomy possible
Best For Research, analysis, drafting Data analysis, customer routing Operations integration, real-time analytics Fine-tuned SLMs; sovereignty-bound and high-volume workloads Critical SCADA, nuclear
Reference

Appendix C: Security Certification Quick Guide

Certification Issuer Scope Key Requirements Validation Frequency
SOC 2 Type II Independent auditor Service provider security controls Security, availability, processing integrity, confidentiality, privacy Annual audit (12-month observation period)
FedRAMP US Government Cloud services for federal agencies NIST SP 800-53 controls; continuous monitoring Annual assessment; continuous monitoring
GovRAMP (formerly StateRAMP) State governments Cloud services for state/local agencies Modified FedRAMP framework; varies by state Varies; typically annual
ISO 27001 Independent auditor Information security management systems Information security governance, risk management, process controls Annual audit; 3-year recertification
HIPAA Compliance Self-assessment + audits Healthcare data protection Privacy Rule, Security Rule, Breach Notification Rule Continuous; periodic audits recommended
Reference

Appendix D: AI Vendor Landscape for Utilities

Hyperscale Cloud Providers: Amazon Web Services (SageMaker, Bedrock, AgentCore Gateway for MCP), Microsoft Azure (OpenAI services, Copilot, Copilot Studio agents), Google Cloud (Vertex AI, Gemini Enterprise), IBM Cloud. These providers offer comprehensive infrastructure, pre-built models, managed services, agent platforms, and enterprise support. Advantages: mature compliance frameworks, extensive integrations, large partner ecosystems. Disadvantages: vendor lock-in risk, complex pricing, broad scope may exceed utility needs.

Power-Sector AI Infrastructure (new in 2025-2026): EPRI's Open Power AI Consortium is developing the first power-sector-specific generative AI models with utility members and NVIDIA, Microsoft, AWS, and Oracle, a membership-based alternative to building domain models alone. Palantir (Foundry) has entered utility grid-data and asset management at scale via its PG&E deployment. Schneider Electric's announced $3.1B agreement to acquire Cognite (June 2026, pending closing) would merge agentic industrial AI with the AVEVA platform, consolidating the industrial data layer many utilities already run. IBM launched a Power Autonomous Operations AI agent in July 2026. On the hardware side, appliance-class inference systems (NVIDIA DGX Spark, AMD-based workstations) enable on-premise small-model deployment at entry costs below $5K.

Specialized Utility AI Platforms: Asterra (formerly Utilis; satellite-based water loss detection), Xylem (Vue platform; water system intelligence), AiDash (satellite-based vegetation management), Overstory (vegetation intelligence for electric utilities), C3 AI (enterprise AI applications for utilities), Innowatts (AMI-driven energy analytics). These vendors provide utility-specific pre-built models and workflows. Advantages: domain expertise, faster implementation, lower customization needed. Disadvantages: less flexibility, limited to specific use cases, smaller vendor ecosystem.

Water/Wastewater Specific: Veolia (integrated solutions), Suez (water analytics), Xylem Digital Solutions, Fracta (machine-learning pipe condition assessment). These vendors focus on water industry applications including leak detection, treatment optimization, water quality prediction. Evaluation criteria: integration with existing SCADA and GIS systems, data validation protocols, regulatory reporting capabilities.

Electric Specific: Uplight (customer engagement), Eaton (grid analytics), Schneider Electric (energy management), GE Vernova (digital twin, O&M optimization), Siemens Energy. These vendors focus on grid operations, demand response, renewable integration, and predictive maintenance. Evaluation criteria: FERC compliance, integration with SCADA and market systems, real-time analytics capabilities.

Evaluation Framework: For any vendor, utilities should verify:

  • SOC 2 Type II report or FedRAMP authorization status
  • Reference accounts of similar size and sector
  • Data protection and privacy provisions in the contract
  • Integration roadmap with the utility's existing systems
  • Support model and SLA terms
  • Pricing transparency and scalability
  • Roadmap alignment with utility strategic priorities

Most utilities should request 60-90 day pilot arrangements before committing to large deployments.

Sources

References

  1. U.S. Department of Energy. (2024). "AI for Energy: Opportunities for a Modern Grid and Clean Energy Economy." April 2024.
  2. National Institute of Standards and Technology. (2023). "Artificial Intelligence Risk Management Framework (AI RMF 1.0)." NIST AI 100-1, January 26, 2023.
  3. Federal Energy Regulatory Commission. (2020). "Order No. 2222: Participation of Distributed Energy Resource Aggregations in Markets Operated by Regional Transmission Organizations and Independent System Operators." September 17, 2020; clarified in Orders 2222-A and 2222-B (2021).
  4. Executive Office of the President. (2025). "Executive Order 14179: Removing Barriers to American Leadership in Artificial Intelligence." January 2025.
  5. Arizona Corporation Commission. (2026). "Formal Inquiry into the Potential Use of Artificial Intelligence in Utilities' Operations." Docket AU-00000A-26-0060, opened March 24, 2026.
  6. Ofgem (UK Office of Gas and Electricity Markets). (2025). "Guidance on the Ethical Use of AI in the Energy Sector."
  7. Water Environment Federation, Amazon, The Water Center at the University of Pennsylvania, and Leading Utilities of the World. (2025). "Water-AI Nexus Center of Excellence." Launched September 25, 2025.
  8. National Association of Regulatory Utility Commissioners and U.S. DOE Natural Gas Partnership. (2020). "Artificial Intelligence for Natural Gas Utilities: A Primer." October 2020.
  9. National Association of Regulatory Utility Commissioners. (2025). "Regulators' Financial Toolbox: Leveraging Software as a Service, Cloud Computing, and Artificial Intelligence in Electric Utilities." Center for Partnerships & Innovation.
  10. The Business Research Company. (2026). "AI in Energy Global Market Report 2026" (July 2026) and "AI in Utilities Operations Market Report 2026" (March 2026, distributed by Research and Markets).
  11. Itron. (2025). "Grid Edge Intelligence: Five Ways Energy Utilities Can Deploy AI Now." Survey of 500 North American electric utility executives, October 2025.
  12. National Grid Partners. (2025). "2025 Utility Innovators Survey." Global online survey of 166 utility innovation leaders, June–July 2025; results released October 9, 2025.
  13. IBM Institute for Business Value. (2025). "Utilities in the AI Era: Powering Ahead to a Smarter Future." November 2025.
  14. U.S. Cybersecurity and Infrastructure Security Agency, NSA AI Security Center, FBI, and international partners. (2025). "Principles for the Secure Integration of Artificial Intelligence in Operational Technology." Joint guidance, December 3, 2025.
  15. Executive Office of the President. (2025). "Winning the Race: America's AI Action Plan." White House, July 23, 2025.
  16. Executive Office of the President. (2025). "Executive Order 14363: The Genesis Mission." November 24, 2025.
  17. Office of Management and Budget. (2025). "M-25-22: Driving Efficient Acquisition of Artificial Intelligence in Government." April 3, 2025.
  18. Board of Governors, Federal Reserve System, OCC, and FDIC. (2026). "SR 26-2: Revised Guidance on Model Risk Management." April 17, 2026; supersedes SR 11-7 (2011).
  19. Boston Consulting Group. (2025). "How to Accelerate the GenAI Revolution in Sales." January 2025; source of the 10/20/70 rule of thumb.
  20. JPMorgan Chase. (2017). "Annual Report 2016 — COIN Contract Intelligence." Investor letter describing 360,000 lawyer-hours of document review automated.
  21. Air Space Intelligence. (2026). "Air Space Intelligence and National Grid Collaborate to Deploy Operational AI to the Electric Grid." May 6, 2026.
  22. Duke Energy. (2025). "How Duke Energy Is Preparing for Another Active Hurricane Season." June 2, 2025; self-healing technology avoided nearly 2.4M customer outages and 11M+ outage-hours across six states in 2024.
  23. PG&E. (2026). "2026-2028 Wildfire Mitigation Plan." AI-enabled camera network, aerial inspection analytics, and grid data platform.
  24. EPRI. (2025). "Open Power AI Consortium." Launched March 2025; domain-specific power-sector AI models.
  25. FirstEnergy. (2025). "Advanced LiDAR Technology Being Used to Enhance Vegetation Management Across FirstEnergy's Footprint." News release.
  26. GE Vernova. "Digital Twin Technology." Company publication; ~$1.6B in avoided losses attributed across 15+ years.
  27. Anthropic. (2024). "Introducing the Model Context Protocol." November 2024; specification revisions through July 2026.
  28. Linux Foundation. (2025). "Linux Foundation Announces the Formation of the Agentic AI Foundation (AAIF), Anchored by New Project Contributions Including Model Context Protocol (MCP), goose and AGENTS.md." December 9, 2025.
  29. MIT Project NANDA. (2025). "The GenAI Divide: State of AI in Business 2025 — Preliminary Findings." July 2025; 52 interviews, 153 survey responses, 300+ public initiatives.
  30. Gartner. (2025). "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027." Press release, June 25, 2025.
  31. Belcak, P. and Heinrich, G., NVIDIA Research. (2025). "Small Language Models are the Future of Agentic AI." arXiv:2506.02153, June 2025.
  32. Schneider Electric. (2026). "Agreement to Acquire Cognite ($3.1B)." June 30, 2026.
  33. Gallup. (2026). "Organizational AI Adoption Jumps Six Points." July 20, 2026; 52% of U.S. employees using AI at work at least a few times a year.
  34. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C.D., and Ho, D.E. (2024). "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools." Stanford/Yale preprint, May 2024.
  35. American Water Works Association. (2026). "2026 State of the Water Industry Report." 2,171 respondents.
  36. U.S. Treasury Department. (2024). "Treasury Announces Over $4 Billion in Fraud Prevention and Recovery in Fiscal 2024." Press release; ~$1B attributed to machine-learning check-fraud detection.

At NewGen Strategies & Solutions, helping utilities adopt new technology responsibly is what we do every day. Whether you need an AI readiness assessment, governance and policy support, or a trusted advisor to evaluate vendor proposals, we'd love to hear from you.