The full power of TDengine, now free forever for up to 5,000 tags.

Explore More

How Utilities Use Grid Data to Detect Faults Earlier

Jim Fan

September 2, 2026 /

As urban grid digitalization advances, the distribution network is no longer just an operation and maintenance scenario with “many devices and lots of data.” It is a tightly coupled operations site where supply reliability, power quality, repair efficiency, and customer satisfaction are all bound together.

For the power utility, the real challenge is not a lack of data. When the main transformer, switchgear, distribution transformer, reactive power compensation devices, and various environmental sensors all keep producing massive volumes of data at once, the operations center still struggles to answer the most critical questions immediately: which segment of line is the anomaly in? Will it keep propagating along the power flow? Is what they see a short-lived fluctuation, or a real fault that is already affecting power quality and the user experience?

This is the increasingly clear dividing line as distribution network O&M moves into its next stage: utilities no longer need only connected data and more SCADA screens. They need data that supports judgment, explains anomalies, and helps O&M crews act faster.

Urban medium-voltage distribution networks: a new judgment challenge across many points and a wide area

Before discussing the data challenges of the distribution network, one thing is worth clarifying: the distribution network is not physically structured like most industrial monitoring scenarios. In a centralized factory, a chemical unit, a longwall mining face, or a car assembly line, all the key equipment usually sits within a few hundred meters. The process flow is chained together, and the anomaly propagation path is plainly visible. The distribution network, by contrast, is spread out like a map: 3 substations at 110kV carry the main grid power flow, 42 feeders at 10kV extend into every corner of the city like blood vessels, and 210 distribution transformers are scattered across industrial parks, commercial centers, office buildings, residential areas, and public facilities, serving more than 500,000 end users in total.

This spread-out form means the O&M semantics of the distribution network differ fundamentally from traditional industrial monitoring.

In spatial terms, equipment is no longer gathered in front of a single control desk but spread across the city’s power corridors. A fault may occur at a distribution transformer in the eastern development zone, while its cause may trace all the way back to mechanical contact aging in the main transformer tap changer at a western substation.

In electrical coupling terms, the main transformer tap changer position determines the 10kV bus voltage, and the bus voltage in turn determines the customer voltage on the low-voltage side of downstream distribution transformers. A rise in a distribution transformer’s three-phase imbalance degree travels through the neutral current and shows up as local overheating in the transformer. Once a reactive power compensation device drops out of service, the power factor and voltage deviation of the entire feeder deteriorate in sync immediately.

In customer terms, the load curves of industrial and commercial customers are completely different. Industrial areas see load double instantly during the day shift, while commercial areas get a second peak in the evening, with a peak-to-valley ratio up to 3:1. The “normal” boundary a single distribution transformer bears differs by time of day and by season.

Because of these differences, the data challenges of the distribution network are not solved by simply applying the “more sensors, more screens” approach. At a scale of more than 8 billion kWh of annual power supply serving 500,000 households directly, what the operations center really needs to answer is not “which device is alarming” but “what does this alarm mean, where will it go, and which segment must be dealt with first.” None of these three questions can be answered from a single-point reading alone.

Alarms are plentiful, but judgment is still slow: what holds operations back is not a lack of data

The distribution network does not lack alarms. SCADA throws up no fewer than a thousand red and yellow flashes every day, and the duty officer’s phone notifications are almost never quiet. But if you watch where frontline O&M crews actually get stuck, the problem is usually not that they did not see something. They saw it and made the wrong first reaction. The following three misjudgments happen almost every day on the distribution operations floor.

Misjudgment one: treating a combined cause as a single cause

An industrial distribution transformer’s oil temperature reading spikes to 88°C. The duty officer’s first reaction is usually “high summer load, normal heat accumulation, keep watching.” But over the same period, the transformer’s three-phase imbalance degree had already climbed from 5% to 21%, and the neutral current exceeded twice the rated value. The real cause is single-phase overloading from an unreported capacity expansion of an electroforming plant’s rectifier units, compounded by dust accumulation on the cooling fins that cut heat dissipation efficiency to 70%. Either cause alone would not be immediately fatal, but when they combine, the oil temperature rises far faster than historical curves would predict. Looking at oil temperature, load rate, or imbalance degree in isolation cannot reach this conclusion. Only when the three are placed on the same time axis does “combined cause” take on operational meaning.

Misjudgment two: mistaking a precursor signal for noise and filtering it out

Starting in the early morning, phase C current of main transformer A-T1 dropped abnormally, from 350A to 172A, and the three-phase imbalance degree rose from 2% to 4.5% with fluctuation. The value never touched any routine alarm threshold. Judged by threshold alone, it was a “minor fluctuation,” possibly even flagged as noise by the rules engine. But it was actually a precursor signal that the phase C contact resistance of the tap changer had already increased by 40%. The real problem is that hours later, when the reactive power compensation device at the commercial center dropped out due to a capacitor fault, this “underlying imbalance” stacked instantly with the reactive power deficit. Bus phase C voltage fell to 9.2kV, downstream mall air conditioners cycled on and off frequently, and electronic equipment in office buildings was damaged. By the time the duty officer reacted, the most economical window for handling it had already passed. What makes many risks dangerous is not that they happen suddenly. It is that they have been evolving for a long time, and no one connected the precursor to the consequence.

Misjudgment three: splitting a chain incident into isolated events and handling them separately

In the previous scenario, the main transformer anomaly, the SVG fault, the bus undervoltage, and the customer complaints are almost inevitably picked up by different crews: the substation maintenance crew handles the main transformer, the reactive power compensation team handles the SVG, dispatch handles the voltage curve, and customer service handles the complaint tickets. Each person sees a small slice of the truth, but no one holds the full causal chain. When the timeline is pieced together in the post-incident review, it turns out that from the earliest phase C current anomaly to the final customer undervoltage, it was the same event completing four causal layers step by step. It was just that, at the time, it had been split into four tickets that did not recognize each other.

These three misjudgments share one trait: the problem is not insufficient monitoring but data not being organized into a form that supports judgment. No matter how many SCADA screens there are or how dense the alarm entries are, if every one stops at the level of “some value on some device exceeded a threshold,” the O&M crew can only keep chasing alarms reactively. What actually drives up business losses is never the number of alarms but the quality of judgment. From an industrial distribution transformer overheating to a 4-hour production stoppage with about US$250,000 in indirect economic loss, and from a main transformer phase C precursor to customer-side undervoltage with about US$15,700 in operating-dispute compensation, the same pattern sits behind every one of these numbers: the data arrived, but the judgment did not.

From passive monitoring to proactive judgment: reorganizing the data chain

Object modeling: scattered measurement points return to a unified business view

The judgment problem of the distribution network does not come from the number of devices or monitoring screens. It comes from how the data is organized. TDengine maps sensors, equipment, transformer districts, regions, and other objects into a clear data catalog through a tree hierarchy, and each node can carry attributes, analyses, panels, events, and related documents. After object modeling, substations, main transformers, switchgear, distribution transformers, and reactive power compensation devices are no longer scattered SCADA measurement points. They are understandable objects on the same power supply chain.

Figure 1: A tree hierarchy organizes plant assets and measurement points into a unified business view

The significance of this change for power utilities is direct. In the past, many anomalies were hard to judge not because historical data was missing but because the same data represented different risks under different regions, seasons, and load characteristics. An oil temperature of 75°C on an industrial park transformer is normal operation, while 75°C on a residential transformer may already mean severe overload. Only after the data structure, asset relationships (substation to main transformer to bus to feeder to distribution transformer to district), and business semantics are sorted out does downstream analysis and judgment have a common foundation. Once a unified entry point is built around objects, the O&M crew no longer sees scattered measurement points but business objects that can be understood together with feeder, district, and upstream and downstream status.

Real-time analyses and event linkage: alarms return to a complete process context

Finding the anomaly alone is not enough to improve O&M response. What the distribution network needs is that, the moment an anomaly appears, the system presents the key context around it at the same time: which feeder the voltage fluctuation occurred on, what load characteristics that time period has, whether upstream and downstream districts are already affected, and whether it is changing together with other quality indicators. On top of data modeling, TDengine adds real-time analysis, event management, and alarm linkage. It continuously monitors the data stream, generates KPIs, detects anomalies, triggers events, and organizes events together with related assets, duration, severity, and contextual trends, instead of just throwing out an isolated alarm.

Take a distribution transformer oil temperature rise as an example. The field needs more than “the oil temperature exceeded the limit.” It needs to see the linked changes in load rate, three-phase imbalance degree, neutral current, and ambient temperature at the same time, to tell whether this is normal heat accumulation during a hot summer period or a sustained single-phase overload stacked with a heat dissipation anomaly. When the main transformer three-phase imbalance alarms, it cannot just stare at the imbalance degree itself. It must further combine tap changer position, downstream bus voltage, reactive power compensation device status, and customer-side voltage changes to judge whether it is evolving into a voltage quality incident. The real value is not how many new alarms were added. It is that anomalies finally have process context that can explain them.

Figure 2: General information settings for real-time analysis

Figure 3: Trigger conditions for real-time analysis

Figure 4: Actions after the real-time analysis triggers

TDengine offers Chat BI, which takes natural-language descriptions of real-time analysis needs, understands them with AI, and creates the real-time analysis tasks, greatly lowering the difficulty and threshold of manual configuration. TDengine also offers Zero Query Intelligence, which senses the scenario and recommends the real-time analysis tasks that should be created for it, further reducing dependence on power industry knowledge and lowering the difficulty of data analysis.

Process analysis and AI-assisted insight: judgment shifts from experience-driven to evidence-driven

Once data is organized into object relationships and anomalies can be explained along the power flow, insight no longer depends only on a few highly experienced veterans. Dispatchers, equipment administrators, repair crews, and electricity inspectors can all share a basis for judgment around the same time axis, the same business objects, and the same set of key indicators. The collaboration model shifts from “everyone looks at their own system” to “everyone forms a consistent judgment around the same facts,” and fault location and handling become faster and more reliable.

TDengine provides natural-language Q&A over process analysis, correlation analysis, regression, batch comparison, anomaly discovery, and panel interpretation, helping users move from “what happened” to “why it happened.” Troubleshooting that used to require repeated confirmation across multiple systems and multiple crews can now be done largely within the same set of objects, events, and analysis chains. The O&M site no longer just faces “knowing that a distribution transformer is abnormal.” It can form judgments backed by evidence faster and turn those judgments into action.

Figure 5: AI interpretation and data mining on the analysis panel

Based on the anomaly events that occur, TDengine supports AI root-cause analysis. It retrieves relevant historical data, forms hypotheses about the cause, verifies them, and generates a structured analysis report, with less manual back-and-forth throughout, greatly reducing dependence on IT skills and on industry knowledge and experience.

Figure 6: AI root-cause analysis of an event

The analysis loop in typical anomaly scenarios

Industrial distribution transformer overheating: the judgment path from oil temperature alarm to production loss

In the industrial distribution transformer overheating scenario, the first signal the system captures is not the final high oil temperature but an abnormal change in load characteristics. Phase A current of distribution transformer IND-TF01 in the industrial park climbed from a 380A baseline to 530A and kept rising to 610A. The three-phase imbalance degree rose from 5% to 21%, crossed the 15% alarm threshold, and stayed there for more than 10 minutes, triggering the Warning alarm in the standing rules. Right behind it, the transformer oil temperature rose from 60°C to 74°C, and within two hours climbed further to 88°C, crossing the 85°C alarm line and triggering a Major alarm. The two standing rules echo each other. The dual-factor confirmation avoids false alarms and gives the operations center a clear entry point right away: this is no longer an isolated load fluctuation or oil temperature rise but a transformer anomaly that needs to be traced immediately.

After the alarm triggers, the field adds the event to the analysis workspace and runs a three-step trace around IND-TF01 and its feeder to judge the nature of the anomaly, its degree of evolution, and its business impact.

Step 1: Confirm the load distribution characteristics

On the IND-TF01 object, add the three-phase current and imbalance degree attributes and observe whether this is even overloading or single-phase overloading. In the chart, phase A current is 685A while phases B and C are only around 380A, the three-phase imbalance degree reaches 31%, and the neutral current is 258A, already 2.3 times the rated value. This shows it is not overall load growth but an unreported expansion on single-phase equipment, most likely the large capacity increase of the electroforming plant’s rectifier units. The nature of the anomaly is now clear: this is a typical customer-side overload in violation of rules.

Step 2: Confirm heat dissipation degradation

Continue to compare oil temperature and load rate on the same time axis. In the chart, at a load rate of only 68%, the oil temperature has already spiked to 102°C. This oil-temperature-to-load ratio clearly deviates from the historical baseline range. This indicates the heat dissipation system itself has a problem. Dust accumulation on the cooling fins has cut dissipation efficiency to about 70%, and single-phase overload combined with insufficient dissipation makes heat accumulate far faster than expected. Stopping the overload alone is no longer enough to pull the temperature back to a safe range.

Step 3: Quantify the downstream impact

The trace extends to the customer side to assess the production loss at the downstream electroforming plant if the fault evolves into a trip. In the chart, the winding hot-spot temperature has risen to 118°C, past the 105°C safety threshold, and overload protection is about to trip. At this point, both the nature of the anomaly and its business impact are confirmed: a single-phase overload in violation of rules is in progress, heat dissipation capacity has failed, and downstream production loss is imminent. The operations center dispatched the load-switch and fin-cleaning work orders more than 3 hours in advance, avoiding about US$250,000 in production-stoppage losses.

Figure 7: Correlation analysis of distribution transformer overheating and insulation aging

Around this analysis loop, several key conclusions become clear. The three-phase imbalance degree and the oil temperature show a typical leading-and-following propagation relationship, directly confirming the classic risk path of single-phase overload to heat accumulation to insulation aging. The deviation of the oil-temperature-to-load ratio lets heat dissipation degradation be inferred from the data side without relying on O&M experience. The leading relationship between winding hot-spot temperature and oil temperature also turns “how much longer can this transformer hold” from an experiential guess into a conclusion backed by data.

Main transformer three-phase imbalance: the judgment path from tap changer anomaly to customer undervoltage

In the main transformer three-phase imbalance scenario, the earliest alarm again comes from a standing rule. Phase C current of A-T1 at substation A had already dropped abnormally in the early morning, from a normal 350 A to 172 A, and the three-phase imbalance degree rose from 2% to 4.5% with fluctuation. The value itself had not crossed the 10% alarm threshold, but it had already entered the analysis system’s “suspected anomaly” list. Hours later, reactive power compensation device COM-SVG01 at the commercial center dropped out of service due to an internal capacitor fault, and the power factor fell from 0.98 to 0.87, triggering a Warning alarm. Right after that, the commercial peak started. A-T1’s three-phase imbalance degree climbed rapidly to 18.3%, and the 10kV bus phase C voltage fell to 9.2kV, triggering Major alarms in succession. One end of the three alarms sits at the equipment layer and the other at the voltage quality line, gathering what could have been split into a “main transformer event,” an “SVG event,” and a “bus voltage event” into a single anomaly that must be traced immediately.

Once the event is added to the analysis workspace, the trace runs backward along the power flow, moving through the three objects A-T1, COM-SVG01, and COM-TF01 in turn.

Step 1: Locate the root cause inside the main transformer

Continue to observe three-phase current and tap changer operating time on A-T1. In the chart, phase C current is persistently lower than the other two phases, and its fluctuation pattern closely overlaps the timing of tap changer switching actions, pointing to poor contact in the phase C tap changer, with contact resistance about 40% higher than normal. This “phase C independently low” curve is not a typical load imbalance response. It is a signal of mechanical contact aging inside the main transformer, meaning A-T1 already meets the conditions for scheduling maintenance immediately.

Step 2: Trace the reactive power system and assess voltage support capacity

Add the reactive output and operating status attributes on COM-SVG01. In the chart, COM-SVG01’s reactive output drops abruptly from +260kvar to 0, and the status bit switches to “out of service.” At the same time, the feeder power factor falls from 0.98 to 0.87, a clear reactive deficit across the feeder. Combined, the two indicators show that reactive power compensation support has disappeared and the 10kV bus voltage has lost its dynamic regulation. Once the commercial peak load starts, the voltage drop is almost unavoidable.

Step 3: Assess the customer-side voltage quality risk

The trace extends to the commercial center’s customer side. On COM-TF01, observe the low-voltage side three-phase voltage. In the chart, phase C voltage falls from 220V to 198V, crossing the 200V undervoltage alarm line. Over the same period, the mall’s central air conditioning control system starts recording frequent on/off cycles, and the office building’s elevator variable frequency drives report voltage anomalies. The conclusion at this step: the voltage quality incident has already propagated from the main transformer layer to the customer layer. If the operators do not immediately switch to the main transformer backup winding and bypass the SVG, complaints and equipment damage compensation will keep growing.

Figure 8: Correlation analysis of main transformer three-phase unbalance and low feeder bus voltage

Following these three steps, a clear root-cause chain is fully reconstructed: the customer-side undervoltage comes from the bus voltage drop; the bus voltage drop comes from the reactive deficit; the reactive deficit comes from the SVG faulting out of service; and the timing of the SVG dropout, which made its impact so severe, is linked to the “underlying imbalance” from poor contact in the phase C tap changer of main transformer A-T1. The phase C current precursor signal that stayed “low for a long time without alarming” strings the two-stage fault mode onto a single chart. The roughly 2.5-hour lead between the SVG dropout and the voltage drop gives the operations center a very valuable prevention window. And this four-layer causal chain, traced back from customer undervoltage to the main transformer tap changer, lets equipment anomalies, reactive power anomalies, and voltage quality anomalies, which are easy to handle separately, finally be explained as a single evolution process within the same analysis path.

Distribution network O&M intelligence: what is needed is not more screens but better connected judgment capability

As power utilities keep pushing distribution network intelligence forward, what really sets the ceiling on the value of digitalization is no longer how much data is connected or how many visualization screens are built. It is whether that data can form actionable judgment on the O&M site. For a scenario like the urban medium-voltage distribution network, many points, wide coverage, and tight coupling, the hard part has never been collecting data. It is that, once enough data is available, how to understand and explain anomalies faster and turn judgment into action.

The value of this shift also shows in the comparison data: transformer fault early warning moved up from “after the fault was found” to more than 72 hours ahead, three-phase imbalance alarm response shortened from 24 hours to within 2 minutes, load-rate exceedance detection timeliness rose from under 60% to over 99%, equipment operating data traceability extended from 3 months to 36 months, and annual unplanned outages dropped from 10-15 per 100 km to within 5. Improvements across these indicators together support the shift from “repair after the fact” to “predictive O&M.”

TDengine organizes scattered data into understandable business objects, turns alarms into explainable process judgments, and further settles data capability into business capability that supports supply reliability, power quality, and repair efficiency, bringing more efficient, precise, and sustainable business results to the distribution network intelligence efforts of power utilities.

TDengine comes with a high-performance, distributed time-series database, Industrial Ontology modeling and an Industrial Agent Runtime, providing a full-stack solution for industrial data streams from collection and storage to real-time analytics, visualization, event management, and root-cause analysis. To learn more about TDengine, visit www.tdengine.com and download it free.

Try it yourself

Install and deploy TDengine Visit the TDengine Download Center, select TDengine All-in-One, choose the deployment platform and architecture that matches your environment, and follow the guided steps to complete the installation.

Load the sample data

On first activation, on the sample data loading screen, select Distribution Network Equipment Status Monitoring and wait for loading to complete.

If you have already activated the product, click your avatar in the top-right corner, select the Management Console, choose Sample Data on the left, then select Distribution Network Equipment Status Monitoring to load it. Wait a few minutes for loading to finish, and you are ready to explore.

Distribution Network Equipment Status Monitoring