The full power of TDengine, now free forever for up to 5,000 tags.

Explore More

Solar PV Operations: Detecting Power Loss Before Alarms Sound

Juno Qiu

September 4, 2026 /

As renewable energy connects to the grid at scale, a solar PV plant is no longer just a large collection of modules and installed capacity. It is a high-precision operating site where generation efficiency, equipment degradation, weather volatility, and economic returns are tightly bound together.

For renewable energy operators, the real challenge is not a lack of data. When string current, combiner box output, inverter MPPT, solar irradiance, and grid connection point parameters keep producing massive volumes of data at the same time, the control room still struggles to answer a few critical questions in a timely way: Why did the PR not reach the level it should have today? Which string is the generation loss actually happening on? Is the power decline in front of you normal fluctuation from cloud shading, or is the equipment quietly degrading?

This is also the dividing line that has become increasingly clear as PV plant operations move into fine-grained asset management: what operators need is no longer just “bringing data in and building a SCADA screen,” but making the data support judgment, explain anomalies, and help the operations crew act faster.

The fundamental difference between PV plants and traditional plants: an asset driven by the weather

Before discussing the data challenges of PV plants, it is worth clarifying one thing first: PV plants are physically different from most traditional generation sites. In thermal and hydro plants, output power is mainly decided by human dispatch. Operators decide how much load the unit carries, and an anomaly shows up as “the load it should have carried was not carried.” The output of a PV plant is not decided by people. It is driven in real time by several natural variables: solar irradiance, module temperature, cloud shading, and wind speed. What operators can do is not “decide how much electricity to generate” but “convert as much of the energy the weather provides into electricity as possible.”

This “driven by the weather” attribute means the operations semantics of a PV plant are completely different from those of a traditional plant.

From the efficiency dimension, the core KPI of PV is not absolute generation but the PR value (Performance Ratio, the ratio of actual generation to theoretical generation). The theoretical peak PR is about 0.83, and measurements on a golden day can reach 0.856. Every 1% of PR lost means about 1.4 million kWh of lost annual generation for a 100 MWp plant, which at a feed-in tariff of about US$0.06/kWh translates to an annual economic loss of about US$81,700. From the fault evolution dimension, most PV problems do not “suddenly break” but slowly get worse. String hot-spot decay can evolve invisibly over weeks. An inverter cooling fan speed can drift slowly from 2,200 rpm down to 1,450 rpm before finally triggering an over-temperature shutdown. This kind of degradation has no obvious fault moment, only an efficiency curve being pulled down little by little. From the spatial dimension, a 100 MWp plant often contains tens of thousands of strings. The 1 MW subset used in this demo already involves 4 inverters, 128 strings, and 2,560 modules, and the problem could be hiding on any module of any string.

Because of these differences, the data challenge of a PV plant cannot be solved simply by applying the “more sensors, more alarms” approach. At the scale of 140 million kWh of annual generation, the operations center does not really need to answer “which device stopped” but “why the PR did not reach the level it should have today, where did the missing generation go, and was it the weather or the equipment.” None of these three questions can be answered by a single point reading alone.

Three paradoxes at the PV operations site: plenty of data, but answers often arrive late

PV plants are not short of data. String current, DC voltage, MPPT voltage, inverter temperature, irradiance, module temperature… a new set of readings enters the control room database every 60 seconds. But if you watch the moments when field operations crews actually get stuck, the problem is usually not that “the data did not arrive” but that the data arrived and you still cannot see where the real problem is. The following three paradoxes happen almost every day at existing PV plants at the same time.

Paradox one: data is abundant, but string degradation can still only be found during the annual inspection. A string inverter connects to 32 strings. What the traditional operations platform shows is “inverter-level average power,” which means after averaging 32 strings, one string dropping 15% gets diluted into an overall drop of under 0.5%, which cannot trigger any alarm at all. Real string hot-spot decay often only gets discovered during the annual infrared thermal imaging inspection, and by then the generation loss has already happened. The data was often there. The problem is that the granularity at which the data is presented does not match the granularity at which the real problem occurs: the problem happens at the individual string level, while monitoring stays at the inverter level.

Paradox two: the weather is clearly fine, but the PR keeps drifting down. The scenario PV operations most easily “explain away” is a low PR. A bit more cloud, a bit higher module temperature, seasonal degradation, module aging… any one of these reasons can explain away an actual efficiency loss. But if you put irradiance and power on the same scatter plot and run a regression, the real problem is usually very clear: at the same 800 W/m² POA irradiance, DC power in a healthy period should reach 240 kW, while in the abnormal period it is only 187 kW, a deviation of 22%, far beyond what any weather fluctuation can explain. Traditional monitoring cannot catch this because it only looks at “whether the absolute value exceeds a threshold,” not “how far this value is from where it should be.” A real efficiency loss never shows up as a value crossing a limit; it shows up as a deviation from the level it should have reached.

Paradox three: no alarm has sounded, but generation loss has already quietly happened. An inverter cooling fan speed slowly slides from 2,200 rpm to 1,820 rpm, and the internal temperature slowly climbs from 42°C to 62°C. Taken individually, none of these values crosses the traditional red threshold line, so the rules engine does not alarm and the duty operator does not get involved. But from that moment on, the inverter conversion efficiency has already started to decline, until one day the temperature breaks 75°C and triggers an over-temperature shutdown, and only then does the whole station operations team realize “something is wrong.” By the time the protection action finally triggers, hundreds of kWh of generation have quietly leaked away in the meantime. What makes many risks truly dangerous is not that they happen suddenly, but that they have been evolving for a long time while no one set the alarm threshold sensitively enough during the gradual change.

These three paradoxes point to the same conclusion: PV operations are never stuck because of “not seeing data,” but because of “not seeing deviation,” “not seeing correlation,” and “not seeing early trends.” No matter how many SCADA panels there are, if every value is only evaluated independently, the operations crew can only chase losses after the fact. And what really drives up economic loss is never the number of alarms: a string hot-spot decay that loses 1,354 kWh cumulatively over 4 days translates to about US$7,900 for a single station; an inverter overheating shutdown that loses 920 kWh over 4 days translates to about US$5,400. Behind all of these numbers is the same pattern: the data arrived but the judgment did not; and by the time the judgment arrived, the action was half a step late.

From passive monitoring to proactive judgment: reorganizing the data pipeline

Object-based modeling: scattered measurement points return to a unified business view

The difficulty of judging what is happening at a PV plant does not come from the number of sensors or monitoring screens. It comes from the way the data is organized. TDengine uses a tree hierarchy to map objects such as weather stations, combiner boxes, inverters, and booster stations into a clear data directory, and each node can carry attributes, analyses, panels, events, and related documents. After object-oriented modeling, the weather monitoring area, the A/B/C/D generation units, and the grid boosting and connection system are no longer scattered time-series measurement points. They become understandable objects sitting on the same “irradiance → string → combiner → conversion → grid connection” chain.

Figure 1: A tree hierarchy organizes plant assets and measurement points into a unified business view

This change means something very direct for renewable energy operators. In the past, many efficiency losses were hard to judge not because there was no historical data, but because the same piece of data represented different health states at different times, under different weather, and in different areas. A module temperature of 60°C is normal at noon under full irradiance, but in the evening as irradiance fades it may already signal a heat dissipation problem. Only when the data structure, the asset relationships (weather station → combiner box → string → inverter → booster station), and the business semantics are straightened out first does later analysis and judgment have a unified foundation. With a unified entry point built around objects, the operations crew no longer sees scattered measurement points, but business objects that can be understood together with irradiance conditions, string numbers, and area PR.

Real-time analysis and event linkage: alarms return to their full process context

Simply detecting an anomaly is not enough to improve PV operations response. What a PV plant needs even more is that, the moment an anomaly appears, the system can present the key context related to it at the same time: which area this power decline happened in, what irradiance conditions it corresponds to, whether it has already affected downstream PR, and whether it is changing together with other efficiency indicators. On top of data modeling, TDengine provides real-time analysis, event management, and alarm linkage. It continuously monitors the data stream, generates KPIs, detects anomalies, and triggers events, and organizes each event together with the related assets, duration, severity, and context trends, instead of just throwing out an isolated alarm.

Take a rising string current dispersion as an example. The site does not only need to know “the current on several strings is low.” It also needs to see at the same time the linked changes in DC bus voltage, MPPT operating voltage, combiner box output power, and real-time PR, to judge whether this is a short disturbance from local cloud shading or ongoing string decay. When an inverter internal temperature rise alarm triggers, you cannot only stare at the temperature value; you need to combine the cooling fan speed, inverter conversion efficiency, DC input power, and the over-temperature alarm state to judge whether it is evolving toward an over-temperature shutdown. The real value is not in how many more alarms were added, but in the fact that an anomaly finally has process context that can be explained.

Figure 2: General information settings for real-time analysis

Figure 3: Trigger conditions for real-time analysis

Figure 4: Actions after the real-time analysis is triggered

TDengine provides a Chat BI capability that lets users describe a real-time analysis need in natural language, then creates the real-time analysis task from that description. TDengine also offers a proactive recommendation capability that senses the scenario and recommends the real-time analysis tasks that should be created for it, reducing reliance on PV industry knowledge and making data analysis easier.

Process analysis and AI-assisted insight: judgment shifts from experience-driven to evidence-driven

Once the data is organized into object relationships and anomalies can be explained along the efficiency chain, insight no longer depends only on a few experienced plant managers. Operations crews, station operators, equipment administrators, and asset managers can all share the same basis for judgment around the same time axis, the same business objects, and the same set of key indicators. Collaboration shifts from “everyone looking at their own screen” to “forming a consistent judgment around the same facts,” and locating and recovering from efficiency losses becomes faster and more stable.

TDengine provides natural language Q&A for process analysis, correlation analysis, regression, batch comparison, anomaly discovery, and panel interpretation, helping users keep going from “what happened” to “why it happened.” A troubleshooting process that used to require repeated confirmation across multiple systems and multiple specialties can now be completed inside the same set of objects, events, and analysis links. The operations site is then no longer facing “we know an inverter has a low PR,” but can form an evidence-backed judgment faster and turn that judgment into action.

Figure 5: AI interpretation and data mining on the analysis panel

Based on the abnormal events that occur, TDengine supports AI root-cause analysis: it retrieves relevant historical data, forms hypotheses about the cause, verifies those hypotheses, and generates a structured analysis report, reducing manual back-and-forth and reliance on IT skills and industry experience.

Figure 6: AI root-cause analysis of an event

The analysis loop in typical abnormal scenarios

String hot-spot decay scenario: the judgment path from a current-dispersion alarm to full-day generation loss

In the string hot-spot decay scenario, the first signal the system catches is not a drop in inverter output power, but an anomaly in a fine-grained, string-level indicator. The string current dispersion on combiner box DCB-A1 in Zone A climbs slowly from its baseline of 0.15 A, breaks through the 0.8 A alarm threshold, and stays there for more than 30 minutes, triggering a Warning alarm in the standing rules. Right after that, the minimum string current drops sharply from 10.6 A to 0.2 A, and the number of abnormal strings rises from 0 to 2, triggering a Major alarm. The two standing rules echo each other and gather a string fault that would otherwise have been diluted by “inverter-level average power” into a single abnormal event that needs to be traced immediately. It is no longer an isolated string fluctuation, but a confirmed event that has already pulled down the PR of the whole inverter.

After the alarm triggers, the site adds the event to the analysis workbench and walks through a three-step trace around DCB-A1 and INV-A to judge the nature of the anomaly, how far it has evolved, and its business impact.

Step 1: locate the abnormal string.

On the DCB-A1 object, the current curves of all 16 strings are shown at the same time. The chart clearly shows the current on string A-S09 dropping sharply from 10.6 A to 0.2 A, while the other 15 strings stay in the normal 10.5-11.0 A range, and the string current dispersion jumps from 0.15 A to 2.31 A. This shows it is not the whole combiner box that has a problem, but that one string has a serious internal fault, most likely an open-circuit failure of the module bypass diode caused by long-term hot-spot stress. The nature of the anomaly is now clear: this is a typical permanent, string-level decay, not an overall power decline caused by weather fluctuation.

Step 2: confirm that MPPT has settled into a sub-optimal point.

The trace continues to the inverter layer. The chart shows INV-A’s MPPT operating voltage falling from a baseline of 785 V to 742 V, with the DC input voltage dropping in parallel. This is the MPPT algorithm actively re-searching for the maximum power point after detecting the string anomaly, but because one string has badly diverged from the curve, it finally locks onto a local sub-optimal point off the global optimum. At this point the inverter conversion efficiency is still above 98%, so on the surface there is “no obvious fault,” but the inverter’s DC input power has already dropped from 257 kW to 175 kW, and its AC output power from 247 kW to 168 kW. Looking only at “whether the inverter has an alarm” cannot find this problem at all.

Step 3: quantify the impact on the full-day PR.

Put the real-time PR value and the weather station’s irradiance curve on the same time axis and compare them. The chart shows INV-A’s real-time PR falling from 0.852 to 0.668, while the POA irradiance over the same period stays above 900 W/m² and the module temperature stays in the normal range. The weather has nothing abnormal; the PR drop comes entirely from the equipment side. At this point both the nature of the anomaly and its business impact are confirmed: bypass diode failure on string A-S09, MPPT stuck at a sub-optimal point, and full-day generation loss still expanding. The operations center issues a work order accordingly. On Day 4, the maintenance crew replaces the diode, the MPPT immediately re-locks onto the optimum point, and the real-time PR recovers to 0.848. The whole event loses 1,354 kWh of generation over 4 days, causing about US$790 of economic loss in the demo subset, which scales to about US$7,900 for the full 100 MWp station.

Figure 7: String hot-spot decay scenario, the judgment path from a current-dispersion alarm to full-day generation loss

Around this analysis loop, several key conclusions become clear. String current dispersion is a leading indicator far more sensitive than inverter average power: it can give early warning while the power of the whole inverter has dropped by less than 5%. A shift in MPPT operating voltage reflects a string anomaly earlier than a drop in conversion efficiency. And the deviation between real-time PR and irradiance lets you infer from the data side, without relying on operations experience, whether the problem is “up in the sky” or “down on the ground.”

Inverter overheating shutdown scenario: the judgment path from fan speed drop to Zone B PR drop

In the inverter overheating shutdown scenario, the earliest alarm again comes from the standing rules, but the entry point is not temperature; it is heat dissipation capacity. The cooling fan speed on inverter INV-B in Zone B slides slowly from a baseline of 2,200 rpm to 1,820 rpm, breaks through the 1,800 rpm alarm threshold, and stays there for more than 15 minutes, triggering a Warning alarm. A few hours later, the inverter internal temperature climbs from 42°C to 65°C, triggering a Major alarm. Another day later, the temperature breaks 75°C, the over-temperature alarm bit flips to true, the inverter switches to over-temperature shutdown (status=3), and the AC output power drops to zero instantly. The three alarms bite together and gather a process that would otherwise have been split into a “fan event,” a “temperature event,” and a “shutdown event” into a single evolution chain that needs to be traced immediately.

After the event is added to the analysis workbench, the trace goes backward along the heat chain, working through three groups of indicators inside INV-B in turn: heat dissipation, temperature, and efficiency.

Step 1: locate the root cause of the falling heat dissipation capacity.

On INV-B, put the cooling fan speed and the inverter internal temperature on the same chart. The chart shows the fan speed following an obvious monotonic downward trend, sliding slowly from 2,200 rpm to 1,450 rpm, a drop of more than 34%. Over the same period the internal temperature follows an obvious monotonic upward trend, rising from 42°C to 75.3°C. The two curves show a typical mirrored negative correlation, directly confirming the classic thermal path of “fan speed drop → heat removal capacity decline → internal temperature rise.” Root cause judgment: the fan blades are clogged with dust accumulated over a long time, mechanical resistance increases, and the speed gradually falls.

Step 2: confirm the actual impact of overheating on efficiency.

Continue by adding conversion efficiency and DC input power attributes to INV-B. The chart shows conversion efficiency sliding slowly from 99.0% in the normal period to 97.3%. On the surface this is a deviation of less than 2 percentage points, but given a power base of 250 kW at full load, this single “efficiency decline” item alone means about 4.25 kWh of extra loss per hour. More critically, when the internal temperature breaks 75°C, the over-temperature protection acts, the inverter goes straight into shutdown, and the AC output power drops instantly from 246.8 kW to 0 kW. The heat buildup has evolved from an “efficiency loss” into a “shutdown loss,” and the two loss modes stack up one after another on the same device.

Step 3: quantify the impact on the area PR.

The trace extends to the overall PR performance of Zone B. The chart shows Zone B’s real-time PR falling from a normal 0.852 to 0.631, forming a clear contrast with Zones A/C/D over the same period. Under the same weather conditions, the PR of Zones A/C/D all stays above 0.84, and only Zone B shows an abnormal decline. This comparison also confirms from another angle that the loss was entirely caused by the single equipment failure of INV-B, unrelated to weather or other zones. At this point the nature of the anomaly, its evolution process, and its business impact all close the loop: clogged fan with accumulated dust, internal temperature rising sharply, conversion efficiency declining, over-temperature shutdown triggered, Zone B PR dragged down. On Day 6, after the fan is replaced, INV-B conversion efficiency recovers to 99.0%, Zone B PR returns to 0.852, and the whole event loses 920 kWh of generation over 4 days, causing about US$54 of economic loss in the demo subset, which scales to about US$5,400 for the full station.

Figure 8: Inverter overheating shutdown scenario, the judgment path from fan speed drop to Zone B PR drop

Walking through these three steps, a clear root-cause chain is fully reconstructed. The Zone B PR drop comes from the lower AC output power of INV-B. The lower AC output of INV-B comes from the falling conversion efficiency and the over-temperature shutdown. The falling conversion efficiency comes from the rising internal temperature. And the rising internal temperature comes from the continuously declining cooling fan speed. The cooling fan speed, a “seemingly unimportant auxiliary parameter,” leads the over-temperature shutdown by about 3 days in time, giving the operations center an extremely valuable preventive maintenance window. And this four-layer causal chain running from the PR drop all the way back to the clogged fan lets efficiency anomalies, temperature anomalies, and shutdown anomalies that would otherwise be handled separately finally be explained as one evolution process in the same analysis path.

Fine-grained PV plant operations: what operators need is not more screens but better connected judgment

As renewable energy operators keep pushing forward with the digitalization of PV plants, what really sets the upper limit of digitalization value is no longer how much data has been connected or how many visualization screens have been built, but whether that data can form executable judgments at the operations site. For a hundred-megawatt-class PV plant, a scenario that is weather-driven, evolves invisibly, and runs at a very fine granularity, the hard part is never collecting data. It is, once there is enough data, how to understand deviations and explain decay faster, and turn judgments into action.

The contrast in the numbers also shows the value of this shift: string-level PR monitoring granularity goes from the inverter level (an average of 32 strings) to independent monitoring at the individual string level; string decay detection goes from annual manual measurement to 24×7 real-time early warning, with alarms triggered by a ±3% deviation; fault location time goes from 2-4 hours of manual troubleshooting to automatic location of the specific string within 5 minutes; and inverter thermal efficiency tracking goes from quarterly inspection to minute-level over-temperature warning. Together, these improvements support the shift from “measuring after the fact” to “keeping the process under control.” Every 1% of PR gain corresponds to 1.4 million kWh of annual generation, about US$81,700 of economic gain, and precise monitoring can recover about US$333,000 to US$667,000 in generation losses each year.

TDengine organizes scattered data into understandable business objects, turns alarms into explainable process judgments, and turns data capabilities into business capabilities that support generation efficiency, equipment health, and asset returns, bringing more efficient, precise, and sustainable results to renewable operators’ PV plants.

TDengine comes with a high-performance, distributed time-series database, Industrial Ontology modeling and an Industrial Agent Runtime, providing a full-stack solution for industrial data streams from collection and storage to real-time analytics, visualization, event management, and root-cause analysis. To learn more about TDengine, visit www.tdengine.com and download it for free.

Try it yourself

Install and deploy TDengine Visit the TDengine Download Center, select TDengine All-in-One, choose the deployment platform and architecture that matches your environment, and follow the guided steps to complete the installation.

Load the sample data

On first activation, on the sample data loading screen, select PV Power Plant Efficiency Analysis and wait for loading to complete.

If you have already activated the product, click your avatar in the top-right corner, select the Management Console, choose Sample Data on the left, then select PV Power Plant Efficiency Analysis to load it. Wait a few minutes for loading to finish, and you are ready to explore.

PV Power Plant Efficiency Analysis