The First Month of Monitoring: Building Your Expected-vs-Actual Baseline
Your system went live last week. The inverter app shows a nice green tick, the numbers look plausible, and the installer has driven away. This is the most valuable 30 days your solar array will ever have, and almost everyone wastes it.
Here is why. In year four, when your generation looks a bit low, you will want to know whether it is genuinely low or just cloudy. Without a reference period captured while the system was provably healthy, you have nothing to compare against except modelled estimates and your own memory of what last April felt like. Both are useless. The commissioning month is your chance to bottle a known-good state: real panels, real shading, real inverter, real meter, measured rather than simulated.
What follows is a protocol. Log these things, compute this ratio, store these artefacts. It takes roughly 20 minutes of setup and about five minutes a week after that.
What you’re actually building
The output of month one is not a pretty chart. It is a small set of files you will still be opening in 2032:
- A raw 5-minute or 15-minute generation series, exported and stored outside the manufacturer’s cloud.
- A daily table of expected vs actual, with irradiance attached.
- A computed solar panel performance ratio for the month, plus its daily distribution.
- A written note of every known anomaly (scaffolding, a day the battery was in a weird mode, the afternoon a neighbour’s tree surgeon was in) so future-you doesn’t chase ghosts.
That last one matters more than it sounds. Baselines get poisoned by unlogged one-offs.
Day zero: capture the constants
Before any data arrives, write down the things that will silently change and that you will later misremember. A plain system.yml next to your data does the job:
commissioned: 2026-09-12
dc_capacity_kwp: 5.94 # 18 x 330 W
inverter: SolaX X1-Hybrid-5.0
inverter_ac_limit_kw: 5.0
battery: Pylontech US5000 x2 (9.6 kWh usable ~ 8.6)
arrays:
- name: rear
kwp: 4.29 # 13 panels
azimuth_deg: 172 # near-south
tilt_deg: 38
- name: side
kwp: 1.65 # 5 panels
azimuth_deg: 262 # west
tilt_deg: 38
shading_notes: "Chimney clips rear string until ~09:10 GMT in winter. Oak at 250 deg, ~14 m, clears side array after leaf-fall."
export_limit_kw: 3.68
mpan: "…"
meter: "Landis+Gyr E470, Octopus, 30-min via API"
Then get a modelled expectation. PVGIS is the one to use: the EU JRC tool at re.jrc.ec.europa.eu, free, no login, and its SARAH3 database covers the UK properly. Run each array separately (different azimuths behave very differently) and download the hourly series as CSV, not just the monthly summary. Set system loss to 14% if you have no better information; that is the tool’s default and it roughly covers cabling, soiling, mismatch and inverter conversion.
For a 4.29 kWp south-facing array at 38° in the Midlands, PVGIS will give you something near 3,780 kWh/year, with October around 265 kWh and December near 95 kWh. Keep the hourly file. You will use it for the ratio.
The four data streams to log
Inverter-side generation. Your inverter cloud holds this at 5-minute resolution, typically for 12 months, sometimes less. SolaX, GivEnergy, Fox ESS and Solis all have APIs; GivEnergy’s is the friendliest (api.givenergy.cloud/v1/inverter/{serial}/data-points/{date}, bearer token from the portal). SolarEdge gives you site/{id}/energy with timeUnit=QUARTER_OF_AN_HOUR but rate-limits at 300 calls a day. Pull daily, append to Parquet or CSV, never rely on the vendor keeping it.
Meter-side import and export. If you are on Octopus, the REST endpoint for half-hourly consumption is the reference dataset that settles your bills, and it is the only stream a dispute can be built on. Log both registers.
Irradiance. This is the one people skip, and skipping it makes the whole exercise guesswork. Options, in order of preference: a cheap reference cell or pyranometer on the roof plane (Apogee SP-110 is about £180 and good enough); Open-Meteo’s free archive API, which serves shortwave_radiation and global_tilted_irradiance hourly for your exact lat/lon with no key; or the nearest Met Office site. Open-Meteo is what I would actually use. The global_tilted_irradiance parameter accepts tilt and azimuth, which means it does the plane-of-array transposition for you.
Module temperature, if you have it. Optimiser systems (SolarEdge, Tigo) expose per-module data. Otherwise take ambient from the same Open-Meteo call and apply a NOCT approximation.
Computing the performance ratio
Performance ratio is the ratio of what your system actually produced to what it should have produced given the sunlight that actually fell on it. It normalises away weather, which is exactly what you need if you want to compare October 2026 with October 2031.
PR = E_actual / (G_poa × A_kwp / G_stc)
where G_stc = 1 kW/m²
Concretely: over a day, plane-of-array irradiation of 3.42 kWh/m² on a 4.29 kWp array gives a reference yield of 3.42 × 4.29 = 14.67 kWh. If the inverter reported 11.9 kWh, PR = 11.9 / 14.67 = 0.811.
A healthy UK domestic rooftop sits between 0.78 and 0.86 in temperate months. Anything above 0.90 usually means your irradiance source is wrong, not that your panels are miraculous. Below 0.70 in clear conditions points at shading, soiling, a string fault or clipping.
In pandas, with a daily-joined frame:
df["ref_yield"] = df["poa_kwh_m2"] * KWP
df["pr"] = df["gen_kwh"] / df["ref_yield"]
clean = df[(df["poa_kwh_m2"] > 1.5) & (~df["flagged"])]
print(clean["pr"].describe())
count 24.000000
mean 0.823
std 0.031
min 0.751
25% 0.807
50% 0.826
75% 0.844
max 0.871
That standard deviation of 0.031 is the number to keep. It is your noise floor. Any future month whose mean PR sits more than two of those below 0.823 is worth investigating; anything within it is weather.
The poa_kwh_m2 > 1.5 filter matters. On very dull days the ratio becomes unstable because inverter self-consumption and the MPPT start-up threshold are a large fraction of a small number. You will see PRs of 0.4 on December drizzle days and they mean nothing. IEC 61724 handles this by excluding low-irradiance intervals, and so should you.
Temperature correction, and when to bother
Raw PR drops in summer because panels get hot. A typical mono module loses 0.35%/°C above 25°C, so a cell at 55°C is down about 10.5%. That will show up as a summer PR of 0.76 against a spring 0.85, and it is not a fault.
Temperature-corrected PR fixes this:
GAMMA = -0.0035
df["t_cell"] = df["t_ambient"] + (df["poa_w_m2"] / 800.0) * (44 - 20)
df["pr_temp_corr"] = df["pr"] / (1 + GAMMA * (df["t_cell"] - 25))
If your commissioning month is March or October, plain PR is fine and you can add correction later. Commission in June or July and you must do this, or your baseline is artificially pessimistic and every subsequent month will look great by comparison.
Week-by-week protocol
Week 1. Get the pipeline running and verify the plumbing before you trust any number. Cross-check inverter generation against the meter on a day you were out of the house: if nothing was consumed, export should approximately equal generation minus standby and battery charging. A 3-5% gap is normal. A 12% gap means your CT clamp is on the wrong cable or reversed, which is the single most common commissioning fault and trivially fixable in week one, impossible to untangle in year three.
Week 2. Build the daily expected-vs-actual table and start eyeballing shape, not just totals. Totals hide everything. Plot one clear day’s 5-minute generation as a text sparkline if you like, but what you are looking for is a smooth bell with a notch. Notches at consistent clock times are shading. A flat top at 5.0 kW is inverter clipping (fine, expected, but subtract it from your reference yield or your PR will look low every June).
Week 3. Identify your clear-sky days. Open-Meteo gives you cloud_cover; days where the daily mean is under 15% are your gold standard. You want at least three. On these, record the peak instantaneous power as a percentage of DC capacity. A 5.94 kWp system hitting 4.9 kW at solar noon is healthy (82%); hitting 3.6 kW is not, and September is when you can still see it.
Week 4. Freeze the baseline. Write it out, and treat the file as read-only:
{
"baseline_period": "2026-09-13/2026-10-12",
"pr_mean": 0.823,
"pr_sd": 0.031,
"pr_n_days": 24,
"clearsky_peak_pct_dc": 0.824,
"clearsky_days": ["2026-09-18", "2026-09-27", "2026-10-04"],
"excluded_days": {
"2026-09-21": "scaffolding on rear roof",
"2026-09-30": "inverter firmware update, 3h offline"
},
"irradiance_source": "open-meteo GTI, tilt 38, az 172",
"notes": "Battery in force-charge overnight 00:30-04:30 throughout."
}
What this buys you
Two winters from now your app will tell you October produced 240 kWh against 265 kWh expected. Without a baseline, that is a shrug. With one, you compute October’s PR as 0.79 against a baseline of 0.823 with SD 0.031, see it sits one SD low, check that irradiance was down 8% year-on-year, and conclude the system is fine. Or you compute 0.71, four SDs out, pull the clear-sky peak and find it at 64% of DC capacity instead of 82%, and now you have a specific, evidenced conversation with your installer about a string that has dropped out. That is the difference between suspecting a problem and demonstrating one.
The detection rules that sit on top of this (rolling z-scores, per-string divergence, soiling ramp detection) are covered in the monitoring and anomaly detection pillar. All of them need a baseline underneath, and none of them can retroactively manufacture one.
One practical warning about the calendar. If your system went live in November or December, your first 30 days will produce a thin, noisy baseline: too few days clear the irradiance filter, and the SD will be wide. Capture it anyway, label it provisional, and schedule a proper 30-day re-baseline for the following March. March and September are the best UK months for this: decent irradiance, moderate temperatures, minimal clipping, low correction error.
Set a calendar reminder for the same fortnight next year to re-run the identical script against the new data. Degradation in crystalline silicon runs about 0.5%/year, so year two should land near 0.819 and year five near 0.807. If year five comes back at 0.74, you have found something, and you will have four years of identically-computed numbers to prove it did not happen overnight.