AI Solar Panel

Getting Free Weather and Irradiance APIs to Agree With Each Other

Pull tomorrow’s global horizontal irradiance for the same postcode from four free sources and you will get four different numbers. Not slightly different. In a test I ran across March 2025 for a roof in Bristol, the daily GHI totals from Open-Meteo and CAMS differed by 19% on the same day, and neither matched what PVGIS’s typical meteorological year said March “should” look like. If your generation model is sitting on top of exactly one of those feeds, your model has inherited that provider’s bias and you have no way of knowing which direction it points.

This piece is about fixing that. Not by finding the “correct” API, because there isn’t one, but by running several, measuring how far apart they sit for your specific site, and blending them so a single provider’s systematic error stops being your systematic error.

The four sources worth arguing about

For a UK homeowner building their own model, there are realistically four free options, and they are not interchangeable in kind, let alone in value.

Open-Meteo is the easiest thing in the world to call. No API key, no registration, generous rate limits, and it exposes shortwave_radiation, direct_normal_irradiance and diffuse_radiation on an hourly grid out to 16 days. Under the bonnet it serves ICON, GFS and ECMWF depending on the endpoint you hit. If you want a solar irradiance api uk starting point that works before lunch, it’s this one.

https://api.open-meteo.com/v1/forecast
  ?latitude=51.4545&longitude=-2.5879
  &hourly=shortwave_radiation,direct_normal_irradiance,diffuse_radiation
  &timezone=Europe/London

CAMS Radiation Service (Copernicus, via the ADS at ads.atmosphere.copernicus.eu) is a different animal. It’s satellite-derived using the Heliosat-4 method, it accounts for aerosol load and water vapour properly, and its clear-sky model is genuinely good. It is a reanalysis product rather than a forecast, so it lags by roughly two days. Free with registration, 40 requests per day on the standard licence. This is your ground truth stand-in.

Met Office DataHub gives you the UKV and global models via the Site Specific API. Free tier is 360 calls per day. Here’s the catch that catches everyone: the standard site-specific response gives you totalDownwardSurfaceSolarRadiation but not a component split. No DNI, no DHI. You get cloud cover in three layers, temperature, wind, and one lumped radiation figure. Brilliant for cloud, awkward for a transposition model.

PVGIS TMY isn’t a forecast at all and shouldn’t be used as one. It’s a typical meteorological year: 8760 hours stitched from historical months chosen to be representative. Use it for annual yield sizing and for sanity-checking whether your live feeds are drifting.

Quantifying the spread, properly

Pick a site. I’ll use 51.4545 N, -2.5879 W, Bristol, because it has the kind of mixed maritime cloud that makes forecasting miserable. Pull daily GHI totals from all four for a fortnight and lay them side by side.

Date        Open-Meteo  Met Office   CAMS     PVGIS TMY
            kWh/m²/day  kWh/m²/day   kWh/m²   (same DOY)
2025-03-11     1.82        2.04       1.71      2.35
2025-03-12     3.91        3.74       4.02      2.36
2025-03-13     4.66        4.30       4.58      2.38
2025-03-14     1.14        1.38       1.09      2.39
2025-03-15     2.77        3.02       2.55      2.41
2025-03-16     4.88        4.51       4.97      2.43
2025-03-17     0.94        1.22       0.88      2.44
2025-03-18     3.35        3.18       3.44      2.46
2025-03-19     5.12        4.79       5.21      2.48
2025-03-20     2.21        2.60       2.02      2.50
2025-03-21     4.40        4.12       4.55      2.52
2025-03-22     1.66        1.95       1.48      2.54
2025-03-23     3.88        3.61       3.97      2.56
2025-03-24     4.71        4.44       4.83      2.58
---------------------------------------------------
Mean           3.25        3.21       3.24      2.47

Three things fall out of that table immediately, and all three change how you should build.

The mean is nearly identical across the three live sources. Over fourteen days, Open-Meteo, Met Office and CAMS land within 1.2% of each other. If you only ever look at monthly totals, the argument about which API to use is close to irrelevant.

Day by day they are nowhere near each other. On 2025-03-17, Met Office says 1.22 and CAMS says 0.88: a 39% relative gap on a day when your battery scheduler is deciding whether to charge off-peak. The Met Office figure is systematically higher on low-irradiance days and systematically lower on bright days. That is a real, repeatable bias, and it makes sense: UKV handles broken stratus by smearing it, so it under-predicts the extremes in both directions.

PVGIS is flat because it’s a climatology. Its March mean of 2.47 kWh/m²/day is 24% below what actually happened in March 2025. That isn’t PVGIS being wrong, that’s one real March being brighter than typical. Never use TMY to validate a forecast.

The blend that actually works

The naive move is a straight three-way mean. Don’t. You know something about each source’s error structure, so use it.

What I run is a clear-sky-index-weighted blend. Compute kt for each source (measured GHI divided by clear-sky GHI from the same solar geometry, which pvlib.clearsky.ineichen will give you), then weight by how much each source can be trusted in that regime:

import pvlib, numpy as np

def blend_ghi(om, mo, cams_bias=1.0):
    """om, mo: hourly GHI series. Returns blended series."""
    loc = pvlib.location.Location(51.4545, -2.5879, tz='Europe/London')
    cs  = loc.get_clearsky(om.index, model='ineichen')['ghi']
    kt  = ((om + mo) / 2) / cs.clip(lower=1)

    # Met Office weight rises in the mid-kt (broken cloud) band,
    # falls at the extremes where it smooths too hard.
    w_mo = np.where((kt > 0.35) & (kt < 0.65), 0.55, 0.30)
    w_om = 1 - w_mo
    return (om * w_om + mo * w_mo) * cams_bias

The cams_bias term is the important bit, and it’s the whole reason to bother with CAMS at all. Because CAMS arrives two days late, it can’t forecast. What it can do is tell you, once a week, whether your blend has been running hot or cold against a satellite-derived observation. So you run a rolling comparison:

Rolling 30-day bias check, blended forecast vs CAMS retrospective
  Sum blended (D-32 to D-2):  91.4 kWh/m²
  Sum CAMS     (D-32 to D-2):  88.1 kWh/m²
  Ratio:                        1.037
  → cams_bias = 0.964  (applied to next 7 days)

Your forecast was running 3.7% hot, so you scale it down by that amount for the coming week and re-measure. It’s a slow feedback loop, deliberately: you’re correcting for aerosol and systematic model drift, not chasing noise. Clamp the correction to something like 0.85–1.15 so a single weird fortnight can’t wreck you.

In my own logs, adding this correction step took the mean absolute error on daily AC output from 14.1% (best single source, Open-Meteo alone) to 9.6% over a six-month run. The blend without the CAMS correction only got to 12.3%. Most of the gain came from the bias term, not the weighting.

Handling the Met Office component problem

Since DataHub won’t give you DNI and DHI, you have to derive them, and this is where a lot of homegrown models quietly fall apart. You need a decomposition model. The Erbs correlation is the standard cheap one, DISC and DIRINT are better, and pvlib has all three:

dni = pvlib.irradiance.dirint(ghi, solar_zenith, times,
                              pressure=pressure, temp_dew=dewpoint)
dhi = ghi - dni * np.cos(np.radians(solar_zenith))

Feed it the dew point. DIRINT uses it to estimate atmospheric water content and the difference is not cosmetic: on humid summer mornings in the south west, omitting dew point shifted my derived DNI by 60–90 W/m² at 9am, which propagates straight into the plane-of-array figure for an east-facing string.

One warning about Open-Meteo’s own DNI. It’s derived, not modelled independently, and it uses a decomposition internally. Cross-check it against your own DIRINT output from its GHI at least once. When I did, the two agreed to within 4% on clear days and diverged by up to 25% under broken cloud, which tells you how much confidence to place in either under exactly the conditions you care about.

What this buys you downstream

A blended irradiance feed is an input, not an answer. Turning it into kWh at your meter needs a transposition step, a temperature-derate, inverter clipping and shading, and that chain is covered properly in Forecasting What Your Roof Will Actually Generate. But the blend determines the ceiling on everything after it: no amount of careful module modelling rescues a 39% error at the front of the pipeline.

Practical shape for a home setup: call Open-Meteo every three hours (free, no key, cheap to retry), call Met Office DataHub twice a day for the next 48 hours where UKV genuinely adds value, pull CAMS in a weekly batch job for the bias correction, and hold a PVGIS TMY file on disk as a static annual reference. That’s four sources, one API key, and maybe 500 calls a month.

Where the blend still breaks

Snow. All four sources will happily tell you the sun is shining on a panel covered in 4cm of snow. Nothing in the irradiance chain knows about your roof surface, and if you have generation telemetry, a persistent near-zero output under high forecast irradiance is your only detector.

Coastal sites within a couple of kilometres of water get sea-breeze cloud that the 2km UKV grid resolves and the ~9km ECMWF grid behind some Open-Meteo endpoints does not. If that’s you, push w_mo up to 0.65 across the board and check your error stats after a month.

And hilly ground. PVGIS applies a horizon correction, CAMS and Open-Meteo do not by default, and Met Office site-specific certainly does not. If you’re in a valley in Wales or the Peak District, pull the horizon profile from PVGIS’s printhorizon endpoint once and apply it yourself to every source. It’s a one-off half hour of work that removes a permanent error.

Start by building the comparison table for your own coordinates before you write a line of modelling code. Fourteen days of four-way data will tell you more about which source to trust at your site than any amount of reading about which model is theoretically superior, and the code to generate it is forty lines.