Ensemble parameter calibration for process-based ecosystem models.
The package estimates a model's parameters from field observations with Ensemble Kalman Inversion (EKI): it runs an ensemble of parameter sets through the model and moves them toward the values that fit the data. It is not tied to any one model, site, or variable, and it is set up so other calibration methods can be added alongside EKI.
Install from the source directory:
install.packages("calibration", repos = NULL, type = "source")or, during development:
pkgload::load_all("calibration")The SIPNET forward model and the workflow scripts also need PEcAn.
A calibration is one call, calibrate(obs, prior, forward, control). Here the
model is the identity, so the answer is known and we can check the calibration
finds it:
library(calibration)
prior <- prior_from_specs(list(
theta1 = list(distn = "norm", parama = 0, paramb = 5),
theta2 = list(distn = "norm", parama = 0, paramb = 5)
))
y <- c(obs1 = 2, obs2 = -1)
Sigma <- diag(c(0.05, 0.05)); dimnames(Sigma) <- list(names(y), names(y))
obs <- list(y = y, Sigma = Sigma)
forward <- function(U, iteration) { G <- U; colnames(G) <- names(y); G }
result <- calibrate(obs, prior, forward,
calibration_control(n_particles = 300, n_iterations = 3, seed = 1))
colMeans(result$U) # near c(2, -1)You provide four things to calibrate():
obs: the observations,list(y, Sigma, meta), frombuild_obs().prior: the prior over the parameters, from theprior_from_*()constructors.forward: a function that runs the model at a parameter matrix and returns its predictions at each observation, frommake_forward_sipnet()or your own.control: the method and its settings, fromcalibration_control().
The numbered scripts run a calibration end to end from a config:
Rscript scripts/010_prepare_observations.R --config <config.yml>
Rscript scripts/020_build_priors.R --config <config.yml>
Rscript scripts/030_calibrate.R --config <config.yml>
Rscript scripts/040_plot.R --config <config.yml>
examples/1_salinas_soc/ is a worked example (soil carbon at the Salinas organic
cropping systems), and vignettes/calibration_demo.qmd walks through the whole
thing step by step.
The repository carries configuration only. Model inputs, run trees and result
caches live outside it, under the directory named by CALIBRATION_DATA_ROOT:
export CALIBRATION_DATA_ROOT=/path/to/artifactsEvery path in an example config is built from that root
For the joint calibration example, the inputs and the calibration result are published so the run can be reproduced or audited without cluster access:
s3://carb/calibration/cal_val_joint_example/
That package holds the block table as run, the prepared management events, the
meteorological drivers, the initial conditions and the result cache, with
checksums. It does not include the run tree, which is per iteration model output
regenerated by scripts/030_calibrate.R. Unpack it under your
CALIBRATION_DATA_ROOT and see examples/2_joint_soc_n2o/README.md.
The statewide treatment-effect matrix is published the same way:
s3://carb/calibration/statewide_effects_example/
That package holds the treatment event package, the PFT inputs, the expanded
settings, the executable, and the scored results, with checksums. It does not
include the model output tree or the meteorological drivers, which are 13 GB
and 3.5 GB; annual_outputs.csv is the annual reduction the scoring reads. The
run report renders from the package alone. Unpack it under your
CALIBRATION_DATA_ROOT and see examples/3_statewide_effects/README.md.
R/the package: the estimator, the transport maps, the priors, the observations, the SIPNET forward model, scores, and plots.scripts/the numbered workflow.examples/a worked example config and its notes.tests/unit tests.vignettes/the demo.runs/per-site SIPNET run configs for the cal/val sites.tools/event_prep/the scripts that generate theevents.jsonfiles.
One directory per site, plus a combined site_info.csv and a shared template.xml.
| site | crop | PFT | window |
|---|---|---|---|
modesto |
almond | temperate.deciduous |
2018-2019 |
russell_ranch |
corn / tomato / wheat | annual_crop |
1992-2014 |
salinas_socs |
lettuce / broccoli | annual_crop |
2003-2011 |
us_bi1 |
alfalfa | annual_crop |
2017-2023 |
us_bi2 |
corn | annual_crop |
2016-2023 |
us_twt |
rice | annual_crop |
2010-2023 |
modesto is the only temperate.deciduous site, so it keeps its own template.xml;
the other five share runs/template.xml. The two files differ only in <pfts>; every
other setting, including the <model><options> block, is identical so that run settings
stay constant across sites.
runs/site_info.csv carries all six sites; there are no per-site copies. Every
user_config.yaml points at it with site_info_file: "../site_info.csv".
01_ERA5_nc_to_clim.R and 03_xml_build.R both process every row of whatever
site_info they are given, and the run window comes from --start_date / --end_date
rather than from the file. There is no --site option. So invoking one site's config
builds met and settings for all six, using that config's dates.
No two of the six sites share a window, so a single multi-site run is not currently
correct for all of them. Running one site in isolation needs either a site filter in
the CLI or the window moving into site_info.csv as per-row columns. Until then, point
that site's external_paths.site_info_file at a one-row copy. There is no
--site_info_file flag on magic-ensemble: the workflow path comes from
workflow_manifest.yaml and cannot be overridden on the command line, so the only way in
is the file the config stages into the run directory.
Committed are inputs only: *_user_config.yaml, the shared runs/site_info.csv,
template.xml and events.json. settings.xml is not committed - 03_xml_build.R generates it into the
run directory and run-ensembles reads it from there, so a copy here would never be read.
Met, initial conditions and model output live on the cluster and in
s3://carb/calval_sa_inputs/, not in git.
external_paths are relative to the config's own directory, and the CLI resolves them from
INVOCATION_CWD, which defaults to wherever you invoke from. Point it at the site directory:
SITE=$PWD/runs/us_twt
cd /path/to/workflows
INVOCATION_CWD=$SITE ./magic-ensemble prepare --config $SITE/us_twt_user_config.yaml
INVOCATION_CWD=$SITE ./magic-ensemble run-ensembles --config $SITE/us_twt_user_config.yamlRunning from the workflows checkout without setting it fails at the first step with
external_paths.template_file: source file not found. That is the CLI resolving
template.xml against the workflows checkout rather than the site directory.
magic-ensemblereplaces the whole<model>element fromworkflow_manifest.yamlat prepare time, so<revision>,<binary>and<options>set in a template here do not survive. That includesNITROGEN_CYCLEandANAEROBIC. The templates carry them so the intended settings are recorded and consistent, but SIPNET does not receive them until the manifest block gains an<options>element.<host>is replaced the same way, from the config'specan_dispatch.- The two templates cannot yet be collapsed into one.
03_xml_build.Rbuilds a MultiSettings withPEcAn.settings::createMultiSiteSettings(), which varies only<run>per site and keeps<pfts>global, and it never reads thesite.pftcolumn. A single template listing both PFTs would hand every site both. Collapsing needs either a per-site PFT in03_xml_build.Ror the site filter above, so that each run covers one PFT. events.inis derived fromevents.jsonbut nothing checks the two agree; they have drifted before.- Sites with multiple treatments do not yet have one
events.jsonper treatment.