Repository navigation
Load model systematically underestimates on sites with large PV #739
Description
Activity
Correction: the mechanism above is wrong
I built a probe before implementing anything, and it contradicts my own analysis. Recording it rather than quietly changing the issue.
The censoring argument has the sign backwards. Discarding samples where
loadW < 0truncates the lower tail. The mean of what survives is biased high, not low. Whatever is driving an underestimate, it is not that guard — and "fixing" it would have made predictions worse.The real defect is the outlier filter locking the model out. A rejected sample updates neither
MAEnorSamples, so the bandmax(MAE × 10, 200)can never grow in response to being persistently wrong. Combined withm.Samples > 50arming the filter after 51 minutes at the 60 s sample cadence, a model calibrated on one quiet hour rejects the real house permanently.Measured on
NewModel(10000):after one overnight hour at 400 W: samples=60 MAE=57.0 band=570.5 then four hours at 5 kW: accepted=0 rejected=240 prediction for that hour: 1794 W (truth 5000 W) after a full week at 5 kW: 1794 W — MAE still 57.0100% rejection, permanently.
MAEcannot move because only accepted samples update it, and nothing is accepted. The model is confidently wrong: a small reported error alongside an arbitrarily large real one — which is exactly the signature the help report flags, and exactly why it looks like a forecast problem from the outside.Whether this is what happened on the reported site I still cannot prove without their model state —
/api/loadmodelwould settle it. But it is a real defect independent of that report, it reproduces from a clean model in seconds, and its symptom matches.What is being fixed
Three changes, in #741:
- A hard bound that is always on. A sample above
3 × PeakWis a measurement fault and is rejected from the first sample, before any soft filter arms. Derived from configured hardware, never from what the model has learned, so it cannot be talked down by a model that has mislearned. This is what keeps the first day safe once the soft filter arms later. - The soft filter arms after a day, not an hour. 51 minutes calibrates a whole house's plausible-residual band on one arbitrary hour — usually a quiet one, because that is when restarts happen. A day spans at least one night-and-day cycle.
- A run of same-direction rejections widens the band. Ten consecutive rejections the same way is the house saying the level moved, not noise. The band widens by exactly enough to admit the residual (
MAE = max(MAE, |err|/10)) and the ordinary EMA takes over from there. Accepting the one sample and resetting was not enough — the band stayed narrow, the next nine were rejected too, and a day at a new level only reached 1650 of 3000 W.
After the fix, on the same probe: the 3-minute 6 kW spike is still rejected in full, a sustained shift to 3 kW is picked up within ten minutes, and every trained bucket reaches 3000 W.
The four directions in the original post are superseded. Only "surface the discard rate" survives, and it is smaller than the rest — the negative-residual discard is not the bug.
- A hard bound that is always on. A sample above
- added a commit that references this issue
on Aug 3, 2026
Symptom
A site running
active arbitrageplanned every afternoon slot against a house load of 383 W. The house was drawing 7.9 kW. The plan was internally consistent —383 W − 11 720 W + 1 060 W = −10 277 W, matching the plan's ownGRID −10.29 kW— so the optimizer solved the problem it was given correctly. It was given the wrong house.Consecutive slots from that plan:
The user's own energy export puts real consumption in that window at ~3.0 kW (
consumer_use 756.6 Whper 15 min bucket), against a forecast of 383 W. The next day's plan predicted 523 / 771 / 1230 W for the same hours.Downstream, the forecast error drives the visible complaint: real load misses the forecast by more than
TwinDriftLoadW(200 W,mpc/service.go:255) on essentially every tick, so reactive replans fire constantly and the live setpoint keeps diverging from whatever plan the dashboard is showing.Mechanism
Service.sampleAtcomputes the training sample and drops it when it comes out negative —loadmodel/service.go:344:Model.Updateapplies the same guard again (loadmodel/model.go:240).The guard is right about the individual sample: negative load is not physical. But the discarded samples are not a random subset. On a site with 11.7 kW of PV,
loadWis a small difference between two large, independently-sampled quantities. Meter and inverter readings are not synchronised and each carries its own noise, so the residual scatters around the true load with an error comparable to the load itself. Every draw that lands below zero is deleted; every draw that lands low-but-positive is kept and trained on.That is censored sampling, and the resulting estimator is biased low by construction. The bias scales with PV size relative to house load — which is why this site, with the largest array, shows it worst, and why a small-PV site never notices.
Two things then lock it in:
model.go:277) rejects residuals beyondmax(MAE × 10, 200)onceSamples > 50. After the mean has converged low, MAE is small, so genuine high-load samples start looking like outliers and are rejected too. The model becomes confidently wrong: small reported error, large real error.gridWis still used (service.go:334). With one of several strings down,pvWis too small,loadWtoo small, and the sample is still trained on if it stays positive.repairPoisonedBuckets(model.go:185) already exists to undo a related poisoning from the heating-subtraction bug, so the failure mode is not new to this file — just not covered for this cause.Why it matters
Every planner strategy sizes its slots against this forecast. An 8× underestimate does not degrade the plan gracefully; it produces a plan for a different building. The optimizer, the dispatcher and the safety layer then all behave correctly and the result still looks broken to the user — which is exactly what makes it expensive to diagnose from the outside.
Directions
Not prescribing a fix, but the options that look sound:
|pvW|greatly exceeds the residual, the sample carries little information about load; down-weight rather than accept it at full strength.slog.Debug("loadmodel: skip (neg load)")is invisible in production. A counter on/api/loadmodelwould have made this diagnosable without reading the source.Detection
The help report added in #TBD flags this at the top of its Findings when the model's prediction and live load diverge by more than 60%, with both numbers quoted. That makes the symptom visible; it does not fix the model.
Repro
Any site where PV materially exceeds house load for a large part of the day. Compare
/api/loadmodel(samples,mae_w,quality) withload_wversusload_w_predictedfrom/api/statusacross a sunny afternoon. A smallmae_walongside a large live gap is the signature.