AI Solar Panel
§4 Section 4 of 6 2,959 words · 13 min

Monitoring and Anomaly Detection: Proving Why Your Solar Panels Are Not Producing Enough

Most people discover a fault in their PV system months late, and usually by accident: a quarterly bill looks odd, or they happen to glance at the app on a bright day and see 1.8 kW where 4.2 kW should be. By then you have lost 200 kWh you will never get back, worth roughly £50 at import prices and about £30 at a 15p/kWh export rate.

The fix is not more dashboards. It is a baseline, a small set of ratios that are dimensionless and therefore comparable across days, and an alert rule that only fires when something has genuinely gone wrong. This page covers how to build that, with the data sources that actually exist in the UK and the numbers you should expect to see.

What “not producing enough” has to mean before you can detect it

Raw kWh is useless as a signal. A 6.4 kWp array in Bristol might make 38 kWh on 12 June and 1.9 kWh on 12 December, and both figures can be perfectly healthy. Any alert built on absolute output will either scream every November or stay silent through a failed string in July.

What you need is a quantity that stays roughly constant when the system is healthy. Three candidates, in increasing order of usefulness:

Specific yield (kWh per kWp per day). Removes system size, keeps all the weather. Good for comparing your system against a neighbour’s, useless as a daily alarm.

Performance ratio (PR). Actual energy divided by the energy the array would have made at 100% efficiency given the irradiance that actually landed on the plane of the modules. A UK rooftop system in decent condition sits at 0.78 to 0.86 in spring and autumn, dipping to 0.72 to 0.80 in high summer because of module temperature, and rising above 0.85 in cold bright weather. PR is the workhorse.

Clear-sky index of output (actual output divided by modelled clear-sky output). Needs no irradiance sensor at all, because the clear-sky model is deterministic. On overcast days it sits at 0.15 to 0.35 and tells you nothing. On the handful of genuinely clear days each month it is razor sharp.

Here is the arithmetic on a real day. System: 16 × 400 W modules, 6.4 kWp, 30° tilt, azimuth 170° (10° west of south), Bristol. Plane-of-array irradiation on 12 June was 6.8 kWh/m². Reference yield is therefore 6.8 kWh/kWp, and the array’s theoretical output is 6.4 × 6.8 = 43.5 kWh. The inverter actually reported 28.9 kWh.

PR = 28.9 / 43.5 = 0.664
Expected PR for a clear June day at this site = 0.79
Shortfall = (0.79 - 0.664) / 0.79 = 15.9%

A 16% shortfall is far outside normal scatter. That is a real finding, and it is the kind of thing you can only see once you have divided out the weather.

Where the irradiance number comes from when you have no sensor

Almost nobody has a pyranometer on their roof. Four workable substitutes, roughly in order of accuracy:

PVGIS hourly time series (the European Commission JRC tool). Free, no key, and the seriescalc endpoint returns hourly global in-plane irradiance for your exact lat/lon, tilt and azimuth back to 2005. Set components=1 and pvcalculation=0 and you get G(i), Gb(i), Gd(i) in W/m². The SARAH3 database covers the UK properly. This is historical only, which is fine for retrospective analysis and for building your expected-PR curve.

Open-Meteo, which exposes global_tilted_irradiance with tilt and azimuth parameters on the free forecast and archive endpoints, no API key. Check the azimuth convention before you trust it (it is measured from south, so 170° true becomes +10). Hourly resolution, updated continuously, and the archive reanalysis goes back decades.

Solcast, whose hobbyist tier gives free rooftop-site estimates and actuals with a modest daily call budget. The quality is better than reanalysis because it is satellite-derived at 5 km, and you can request estimated_actuals for the last week. Budget your calls: pull once an hour, not once a minute.

pvlib’s clear-sky model, which needs nothing but your coordinates and a timestamp. This is the cheapest option and the one I would start with:

import pvlib, pandas as pd

site = pvlib.location.Location(51.4545, -2.5879, 'Europe/London', 50)
times = pd.date_range('2026-06-12 04:00', '2026-06-12 22:00',
                      freq='5min', tz='Europe/London')
cs = site.get_clearsky(times)                      # ghi, dni, dhi
solpos = site.get_solarposition(times)
poa = pvlib.irradiance.get_total_irradiance(
        surface_tilt=30, surface_azimuth=170,
        dni=cs.dni, ghi=cs.ghi, dhi=cs.dhi,
        solar_zenith=solpos.apparent_zenith,
        solar_azimuth=solpos.azimuth)

clear_sky_kw = 6.4 * poa.poa_global / 1000 * 0.79  # kWp x kW/m2 x PR

That last line gives you a per-5-minute expected power curve for a cloudless day. Divide your measured power by it and you have a number between 0 and about 1.05 that behaves the same way in December as in June.

Getting the data out of your kit

Cloud APIs are convenient and rate-limited. Local Modbus is unlimited and fiddly. Use both.

SolarEdge: the monitoring API is capped at 300 requests per day per site (and per account token), which sounds generous until you realise energyDetails at QUARTER_OF_AN_HOUR granularity is restricted to a one-month window per call. Poll currentPowerFlow every 5 minutes and you burn 288 calls before lunch. Better: run the SolarEdge Modbus TCP interface (enable it on port 1502 in the inverter’s local web UI), read holding registers directly, and keep the cloud API for weekly backfill. The solaredge_modbus Python library and the matching Home Assistant custom integration both expose per-optimiser data if you have optimisers.

Enphase: Enlighten API v4 uses OAuth and meters your calls monthly on the free plan, so cache aggressively. The per-microinverter data is the killer feature here, because a single failing microinverter is a 1/16th loss that will never show up as a system-level alarm.

GivEnergy: the v1 REST API returns 5-minute data-points for a given day, plus battery SOC, and there is a local Modbus TCP option on port 8899 that most people end up using because it has no rate limit at all. givenergy-modbus and the givtcp container are the standard routes.

Solis, Growatt, Sunsynk, Fox ESS, Huawei: all of these have either a documented cloud API (SolisCloud’s signed-request API, Huawei FusionSolar’s northbound API) or a serviceable RS485/Modbus path via a cheap USB adapter or a Waveshare RS485-to-Ethernet box. Growatt’s ShineServer API is unofficial but well reverse-engineered.

Your smart meter, which is the ground truth for import and export. Two free routes: n3rgy gives you half-hourly consumption and export from the DCC once you consent with your MPAN, and a Hildebrand Glow CAD gives you near-real-time data over MQTT at roughly 10-second intervals for SMETS2 meters. Export data matters because it is what you actually get paid for, and because a mismatch between inverter-reported generation and metered export plus measured house load will expose a CT clamp installed backwards or on the wrong conductor.

Store it somewhere you can query. InfluxDB plus Grafana is the common setup and takes an evening. A SQLite file written by a cron job and read with pandas is entirely sufficient for a single house and much easier to back up.

Reading the shape of the day

Before any statistics, learn the signatures. A 5-minute power curve tells you more than a daily total ever will.

What you seeMost likely causeConfirming check
Flat top at exactly 3,680 WG98 export limitation working as designedDoes the flat top disappear when the immersion or EV charger is on?
Flat top at inverter nameplate (e.g. 5,000 W)Inverter clipping, array oversizedHappens only on clear days near solar noon, April to August
Smooth curve, uniformly 15% low, every daySoiling, module degradation, or a wrong PR assumptionCompare against last year’s same-month PR
Sharp notch 08:00 to 09:30, same time daily, drifting with seasonShading (chimney, aerial, neighbour’s tree)Notch start time should shift by minutes per week
Ragged sawtooth, output collapsing to zero for 60 to 90 seconds then recoveringGrid over-voltage tripLog grid voltage; look for the 10-minute mean approaching 253 V
Whole array drops to zero at 12:40 on a sunny day, returns at 12:55Inverter fault or a frequency eventCheck the event log against the inverter error codes reference
One string at 87% of the other, constant all dayString fault: failed module, bad MC4, blown fuseRatio test, below
Midday output ramping down smoothly on a sunny breezy dayLFSM-O responding to grid frequency above 50.4 HzRare, but real, and it is not a fault

The over-voltage case deserves special attention in the UK because it is common and frequently misdiagnosed as an inverter fault. Statutory supply voltage is 230 V +10% / −6%, so 216.2 V to 253.0 V. G98 and G99 protection settings trip the inverter when the 10-minute mean exceeds 253.0 V, with a fast trip around 262 V. On a long rural spur, or at the far end of a street where half the houses now have PV, midday voltage rise of 6 to 8 V is normal and pushes you over the limit. If your logs show voltage sitting at 249 to 252 V between 11:00 and 15:00 with intermittent dropouts, that is a DNO problem, and the DNO is obliged to investigate. Log it for a fortnight before you call them.

The three ratios worth alerting on

Performance ratio, temperature corrected. Raw PR drops every summer, which will trip a naive threshold in July. Correct it:

PR_corrected = PR / (1 + gamma * (T_cell - 25))
T_cell ≈ T_ambient + (NOCT - 20) / 800 * POA

With γ = −0.0034 /°C (a typical modern mono PERC coefficient) and a cell temperature of 48°C, the correction is 7.8%. A raw PR of 0.74 on a hot day becomes a corrected 0.80, which is right in line with the March figure. Now a single threshold works all year.

String ratio. If you have two strings, or two MPPT inputs, the ratio of their outputs is the single most sensitive diagnostic you have, because both strings see identical weather. Take the median ratio over the middle four hours of each day:

2026-09-14  MPPT1/MPPT2 = 1.004
2026-09-15  MPPT1/MPPT2 = 0.997
2026-09-16  MPPT1/MPPT2 = 1.011
2026-09-17  MPPT1/MPPT2 = 0.872   <-- 
2026-09-18  MPPT1/MPPT2 = 0.869

A step change of that size, persisting, on an 8-module string, is one module out of action (1/8 = 12.5%, and 0.872 is close enough once you account for the remaining modules running at a slightly different operating point). No weather model required, no irradiance data, and it works on a cloudy day.

Battery round-trip efficiency. Sum AC energy in and AC energy out over a full charge-discharge cycle. A hybrid system with a decent LFP pack should return 84% to 90% AC-to-AC. Track it weekly:

week of 2026-08-24   charged 52.1 kWh   discharged 45.0 kWh   RTE 86.4%
week of 2026-08-31   charged 49.8 kWh   discharged 43.1 kWh   RTE 86.5%
week of 2026-09-07   charged 51.3 kWh   discharged 41.2 kWh   RTE 80.3%

That third week is worth investigating: either a cell is drifting and the BMS is cutting the discharge short, or the inverter’s standby draw has gone up, or the system started doing partial cycles that never reach the efficient part of the curve. Separately, measure standby consumption directly, overnight, with the battery idle. A hybrid inverter typically draws 30 to 60 W continuously, which is 260 to 525 kWh a year. If yours reads 140 W, something is running that should not be.

An anomaly rule that does not cry wolf

Start simple. Compute the daily corrected PR, take a rolling 21-day median and a rolling median absolute deviation, and flag a day when it sits more than three robust standard deviations below the median. Median and MAD, not mean and standard deviation, because one storm-damaged day should not widen your bands for three weeks.

import numpy as np

df['pr_med'] = df.pr_corr.rolling(21, min_periods=10).median()
mad = df.pr_corr.rolling(21, min_periods=10).apply(
        lambda x: np.median(np.abs(x - np.median(x))))
df['z'] = (df.pr_corr - df.pr_med) / (1.4826 * mad)
df['flag'] = df.z < -3

Then add persistence, which is the part everyone skips. Do not alert on flag. Alert on three flagged days inside a rolling window of five. Single bad days are almost always a data gap, a reboot, or a curtailment event. Three in five days is a fault.

Filter the input too. Drop days where the daily clear-sky index is below 0.25, because PR is numerically unstable when the denominator is tiny, and drop any day where you recorded a clipping event, because clipped days artificially depress PR by design.

For sub-daily detection, use the 5-minute residual between measured power and clear-sky expected power, but only evaluate it in a window from two hours after sunrise to two hours before sunset, and only on days whose midday clear-sky index exceeds 0.8. On those clear days, a residual that stays below −15% for more than 30 consecutive minutes while the sun is high is worth a notification. That rule catches a failed optimiser within a day or two in summer.

Isolation Forest from scikit-learn earns its place once you have a year of data and more than one variable. Feed it a feature vector per day: corrected PR, string ratio, peak power divided by nameplate, the hour of peak output, inverter internal temperature, and grid voltage 95th percentile. Fit on a known-good 12-month period, score subsequent days, and investigate anything with a score in the worst 1%. The value is not that it beats the threshold rule on PR alone; it is that it catches combinations, such as a normal PR achieved with an abnormally late peak and a high grid voltage, which is what partial curtailment looks like.

Do not use Prophet or an LSTM for this. Seasonal forecasting models fit the fault along with the signal and quietly absorb a slow 8% degradation into their trend term.

Turning a flag into a diagnosis

An alert tells you something changed. The next 20 minutes decide whether it is your problem or someone else’s.

Check whether it was just the weather, using an independent source. Sheffield Solar’s PV_Live API publishes half-hourly estimated PV generation for Great Britain and for each regional group, free and without a key. If national output per installed kW dropped the same amount your system did, the sky did it. PVOutput is the other route: join a team or just compare against nearby systems in the postcode search, and if four systems within ten miles all show the same dip, stand down.

Pull the inverter event log next. Almost every fault that matters leaves a code, and the codes are frequently more informative than the app’s friendly text. A SolarEdge “Error 3x9B” or a Solis “ARC-FAULT”, “OV-G-V01”, or “NO-Grid” tells you exactly which subsystem gave up. Our companion page on inverter error codes and what they mean for your output maps the common Solis, Growatt, SolarEdge, Fox ESS and Sunsynk codes onto the production impact you should expect, which matters because some codes cost you 100% of the day and some cost you 2%.

Look at the per-module data if you have it. Enphase gives you per-microinverter output natively; SolarEdge gives you per-optimiser figures in the layout view and through the Modbus interface; Tigo optimisers report per-module through the CCA. One module reading 40 W while its 15 neighbours read 310 W is a bypass diode or a connector, not a mystery.

Rule out the boring causes before the interesting ones. A CT clamp knocked loose during a boiler service. A firmware update that reset the export limit. An EV charger with solar-diversion mode that quietly consumed the generation you thought had vanished (your myenergi Zappi or Eddi has its own API, and its logs will show it). Scaffolding. A satellite dish installed in March. Lichen along the bottom 3 cm of every module, which in a UK coastal or wooded location can cost 4 to 6% and takes a pole brush and a bucket of deionised water to fix.

A monitoring stack that takes one weekend

Home Assistant handles ingestion and alerting with the least friction, because the integrations already exist for SolarEdge, Enphase, GivEnergy, Solis, Fox ESS, Sunsynk, Growatt, Victron VRM and the Glow CAD, and because the Solcast and Forecast.Solar integrations give you the expected-output series without writing an HTTP client. Push the long-term series into InfluxDB and build the PR and string-ratio panels in Grafana, since the HA recorder is not designed for three years of 5-minute data.

A template sensor for daily corrected PR, an automation that evaluates it at 23:30, and a notification via ntfy or a Telegram bot is about 40 lines of YAML. Give the notification a body that contains the numbers, not just “solar anomaly detected”: today’s PR, the 21-day median, the string ratio, and the peak power. Half the time you will diagnose it from the notification alone.

Run a weekly job that recomputes the full history rather than only appending. Inverter cloud APIs backfill missing intervals hours or days later, and a pipeline that only ever appends will leave permanent holes that your MAD calculation reads as faults.

Set one calendar reminder for early April and one for early October. Those are the months with high irradiance and low module temperature, when PR is at its annual peak and a 5% loss is easiest to see, and they bracket the season when a fault costs you the most.

In this section

The supporting pages under this subject.