Repository navigation
Enumerate validation tests and null checks #233
Description
Activity
- added a parent issue
on Jul 7, 2026 A first pass at the enumeration, surveyed from the UNIONS 2D release papers (I–IV), KiDS-Legacy (Wright & Stölzner 2025), DES Y3/Y6, and HSC Y3.
This is a catalogue of the option space: pass/fail criteria and final selection are deferred. As tests become active work, they get promoted to their own issues, with this list linking them.
Existing 2D tests reworked for multi-bin — the bulk of the work:
- PSF systematics (ρ/τ statistics, leakage, ξ_sys) per bin / bin-pair
- B-mode suite (pure-EB, COSEBIs, C_ℓ^BB) tomographic
- catalog null tests re-run per catalog version
- per-bin multiplicative & additive bias calibration
- tomographic covariance validation; config↔harmonic cross-checks on mocks
Genuinely new for tomography:
- per-bin n(z) calibration: SOMPZ, Δz priors from mocks, clustering-z
- internal-consistency splits: drop-bin, auto-vs-cross, red–blue, hemisphere (cf. KiDS-Legacy's tiered framework)
- PSF diagnostics not run in the 2D round: mean-shear-vs-PSF slope, brighter-fatter, chromatic residuals
Requires the lens sample (for the planned 3×2pt extension): shear-ratio test, pairwise probe-consistency.
Estimator space is open — ξ±, pseudo-Cℓ, COSEBIs, or several in parallel; many tests come in per-space variants, and running multiple spaces makes cross-statistic consistency a test in itself.
Blinding (→ #232): under a data-vector shift, catalog-level tests are untouched (upstream of the data vector) and the B-mode suite is transparent to the shift (a theory-difference shift is pure E-mode); only parameter-level checks need blinded-chain protocol. Reasoning in the dropdown.
Full test catalogue (9 categories, ~60 tests)
This is a catalogue of candidate tests; selection and pass/fail criteria are deferred. It surveys the validation and null tests used across UNIONS Papers I/II/III/IV (2D release), KiDS-Legacy (Wright 2025, Stölzner 2025), DES Y3 & Y6, and HSC Y3, so that the tomographic round can choose from a known menu. Where a survey stated a threshold it is recorded below, but adopting it here is a later decision.
The tomographic analysis' estimator space is not yet decided — configuration space (ξ±), harmonic space (pseudo-Cℓ), COSEBIs, or several in parallel. Many tests below come in per-space variants (B-modes, goodness-of-fit, scale-cut stability), and running more than one space makes cross-statistic and cross-space consistency tests in their own right (§6).
Scope tags:
[needs lens sample]— applies once the planned 3×2pt extension's lens sample is in hand; everything untagged applies to the cosmic-shear analysis. Continuity tags:✅done in UNIONS 2D ·🆕new-for-tomography ·✅→🆕a 2D analogue exists but the tomographic bin structure makes it new work (typically per-bin / per-bin-pair).
1. PSF systematics
- ρ-statistics (ρ₀–ρ₅) — PSF-model & residual ellipticity/size autocorrelations; input to the (α,β,η) leakage model.
✅→🆕 per-bin
All surveys; UNIONS/KiDS/DES fit α,β,η; DES Y3 χ²/dof≈95/120. Amplitude/SNR requirement à la Mandelbaum+18. - τ-statistics (τ₀,τ₂,τ₅) — galaxy–PSF cross-correlations; solve (α,β,η) via τ = R·Ω + Σ.
✅→🆕 per-bin
UNIONS; DES Y6 (χ²/dof=218/248, p=0.91); HSC. - PSF-leakage params α,β,η — least-squares/emcee fit; the α–η degeneracy needs a joint 2D fit.
✅→🆕
UNIONS Paper I; DES; KiDS α₂≈0.01. Consistency of corrected vs uncorrected checked. - ξ_sys/ξ_+ contamination fraction — additive PSF term / cosmological signal; sets scale cuts.
✅→🆕 per-bin-pair
UNIONS; KiDS (Bacon ξ_sys & Paulin-Henriksson); DES Y6 (~0.01%); HSC. UNIONS/KiDS require <10% of statistical σ. - Mean shear vs PSF ellipticity (leakage slope) — linear fit ⟨e_gal⟩ vs e_PSF (+4th-moment); slope=leakage, intercept=additive.
🆕
DES-style; not in UNIONS 2D. DES Y3 slopes ~−0.001±0.002; DES Y6 ∂e₁/∂e₁,PSF=−0.0057±0.0018 (~2σ). Slopes ≈ 0. - PSF residuals at validation stars (ellipticity, size, 4th moments) — held-out (20%) star residuals centred on zero.
✅
UNIONS Paper I. Qualitative, centred at zero. - Brighter-fatter / size residual vs magnitude — fractional star size/shape residual vs mag.
🆕
DES Y3 (<0.5% except m<16.5). - PSF residuals vs star colour (chromatic) — residuals vs r−z near median galaxy colour.
🆕
DES Y3 (dT/T<0.002, δe<1e-4 at r−z=0.75). - Position–shape λ-statistics — galaxy-position × PSF-residual correlations (Zhang+24); bounds density-shape additive bias.
[needs lens sample]✅
UNIONS Paper I. Not a null test (nonzero expected); bounds GGL contamination. - PSF contamination vs TATT B-mode prediction — PSF-induced B-mode ≪ IA-sourced B-mode.
🆕 (optional)
DES Y3 harmonic (Doux+22). ~1 order of magnitude below TATT A=1.
2. B-modes
- Pure-mode ξ_±^{E/B} (config space) — E/B decomposition of the 2PCF; PTE vs zero.
✅→🆕 tomographic
UNIONS Paper II; DES (2PCF p=0.14/0.24). PTE>0.05 (KiDS/DES use p>0.01). - COSEBIs B_n — real-space log E/B statistic, first 6 (and 20) modes.
✅→🆕
UNIONS II; KiDS (p>0.01; p=0.04/0.09); DES. PTE>0.05. - Harmonic-space C_ℓ^BB (pseudo-Cℓ) — B-mode power spectrum, NaMaster/iNKA cov.
✅→🆕
UNIONS IV (BB PTE=0.13); DES Y6 (χ²/dof=24/30, p=0.76); HSC (BB p=0.50). PTE>0.05. - C_ℓ^EB cross-power — check B-modes don't correlate with E-signal; symmetric scatter about zero.
✅→🆕
UNIONS IV (EB PTE=0.95); HSC (EB p=0.95). Qualitative. - PTE-map / scale-cut stability — grid of angular/ℓ cut combinations; adopted cuts must sit inside a broad passing region.
✅→🆕
UNIONS II/IV; KiDS; HSC. Stable to shifting by several bins. - B-mode-driven data cut investigation — B-mode failure → root-cause (blind) → cut.
✅ (expect to recur per-bin)
UNIONS II (size cut, n_eff 6.48→4.96); KiDS (astrometric masking, 4%, 46.6 deg²); HSC (drop GAMA09H field, ℓ<300 excess). Control: randomized-mask must not restore pass. - Look-elsewhere / fitted-cut check — treat scale-cut boundaries as fitted params (reduce dof), re-check min PTE.
✅
UNIONS II (0.18→0.09). Still >0.05.
3. Catalog-level null tests
- Tangential/cross shear around field/tile/CCD/cell centres — γ_t, γ_x vs radius from processing-unit centres; radial additive systematics.
✅→🆕 re-run per catalog version
UNIONS I; DES Y3/Y6. Reduced χ²≈1 (random-subtracted). - Tangential/cross shear around stars (Gaia / bright-faint / validation) — coherent shape systematics near stars, split by magnitude.
✅ ⚠ open anomaly
UNIONS I (Gaia OK; own validation stars anomalous χ²=2.53 ⚠); DES Y3/Y6. χ²≈1. - Mean shear / ⟨g₁,g₂⟩ vs focal-plane & tile position — χ² null vs zero, binned by CCD/tile/cell coords.
✅→🆕
UNIONS I (tile OK; CCD χ²/dof≈1.13 mild); DES Y3/Y6; KiDS 2D c-term map (|c/σ|≤1%). - Mean shear vs survey properties — ⟨shear⟩ vs depth/seeing/sky/airmass/SNR/size/mag/photo-z maps.
✅ (partial) →🆕 broaden
DES Y3/Y6; UNIONS I sky-maps (n_eff, shape noise, FWHM). No spatial pattern in e1/e2. - Ellipticity–property regression bias b_i — linear regression of e_obs on survey/galaxy properties (Gatti+21); auto-contamination term.
✅
UNIONS I. ⟨bS·bS⟩/⟨ee⟩ < 0.02 on analysis scales. - Star–galaxy cross-correlation — shear × star position/ellipticity; stellar contamination bound.
🆕 (DES/HSC-style)
DES Y3 harmonic; HSC. Negligible vs statistical error. - Binary-star / stellar-locus contamination — high-|e| tail removal via colour/size-mag cut; residual m budget.
✅ (selection)
DES Y3 (~20% of flagged removed); UNIONS star selection. - Masking / selection sanity by magnitude — cumulative-cut mag histograms behave as expected.
✅
UNIONS I. Qualitative.
4. Calibration & shear bias
- Multiplicative bias m from image sims — grid + N-body/realistic placement; recover m at known input shear.
✅→🆕 per-bin m + correlations
UNIONS I/V (m=−0.057, σ tripled→0.014); DES Y3 (m=−2.08%); DES Y6 (m=(3.4±6.1)e-3); HSC. Prior in inference. - Additive bias c₁,c₂ — weighted-mean ellipticity subtracted pre-2PCF, per bin/hemisphere.
✅→🆕 per-bin
UNIONS I (jackknife, ~1e-4); KiDS (per bin/hemisphere); DES Y3. Calibration, not pass/fail. - Metacalibration shear + selection response — R_γ and R_sel per galaxy via artificial shearing; incl. shear-dependent cuts.
✅
UNIONS I; DES Y3. Applied as correction. - Selection-bias calibration (m_sel, a_sel) — image-sim estimate of shear-dependent selection effects.
✅
HSC; DES Y3; UNIONS FLAGS study. - FLAGS / size-cut selection B-mode check — deblended-source inclusion (FLAGS≤2) induced B-modes → reverted.
✅
UNIONS I. Drove FLAGS=0 + size-cut decisions. - Half-light-radius ↔ DES-T size cross-check — verify size definition maps onto DES-Y3 T on overlap.
✅
UNIONS I. Qualitative.
5. Redshift distributions n(z) (largely 🆕 — the heart of tomography)
- SOM n(z) calibration — SOM on spec sample × photometry, bootstrap for uncertainty.
✅→🆕 tomographic SOMPZ
UNIONS III (DEEP2+VVDS+VIPERS × CFHTLenS); KiDS; DES (SOMPZ). Validated on mocks. 2D was single-bin. - Δz bias & uncertainty from mocks — push mock (MICE2/GLASS) through SOM, compare to truth.
✅→🆕 per-bin Δz
UNIONS III (MICE2, Δz=−0.003±0.018); DES. Gaussian Δz prior per bin. - Shear-ratio (SR) test — GGL ratios, same lens / different source bins → geometric z-calibration check.
[needs lens sample]🆕
DES Y6 (lowest lens × highest source bins). Consistency with SOMPZ+WZ. - WZ / clustering-z cross-calibration — cross-correlation redshifts vs photo-z.
[needs lens sample]🆕
DES (SOMPZ+WZ); HSC (CAMIRA-LRG); HSC photo-z secondary-peak rejection.
6. Data-vector internal consistency
- Scale-cut sensitivity scan (θ_min / k_max / ℓ_max) — track S8 shift & χ² as small scales removed.
✅→🆕
UNIONS III/IV; DES; KiDS; HSC. Bin-removal ΔS8 < 0.2σ; k>k_max <10% of signal. - Drop-bin test — remove each tomographic bin (+ its cross-corrs), re-infer.
🆕 (needs bins)
KiDS (max 0.24σ COSEBIs); HSC. Descriptive shift. - Auto- vs cross-correlation split — all autos vs all crosses (probes II vs GI IA).
🆕
KiDS-Legacy Stölzner. p>0.01 tiered metrics. - Small- vs large-scale / ℓ split — split data vector at scale midpoint, independent inference.
✅→🆕
UNIONS IV; KiDS; HSC. Consistency (~0.1σ). - Config- vs harmonic-space S8 consistency — same catalog through both pipelines, correlated-mock significance.
✅ ⚠ known ~2σ offset
UNIONS III/IV (0.06 in S8, 2.18σ ⚠ caveat); KiDS. Mock-based p-value. - IA-model robustness (NLA/TATT/z-dep/A_IA=0) — re-infer under IA variants.
✅→🆕 (IA more active tomographically)
UNIONS III (largest −0.71σ); KiDS (NLA-M fiducial); DES. Justify baseline. - Nonlinear P(k)/baryon robustness — HMCode±feedback vs Halofit; k_max signal contribution.
✅
UNIONS III/IV; KiDS; DES; HSC. Shifts ≲0.4σ. - Tiered consistency framework (evidence / param-shift / TPD) — 3-tier split-model metrics (suspiciousness, KDE shift, translated PPD).
🆕 (KiDS-Legacy machinery)
KiDS-Legacy Stölzner. p>0.01 / N_σ<2.36. - Hemisphere / N–S catalog split — two patches, joint cov, split-model cosmology.
🆕
KiDS (N_σ,S=0.71). p>0.01. - Red–blue colour split — split by galaxy type, separate IA/z params.
🆕
KiDS (2.6–2.8σ, attributed to IA). p>0.01. - Cross-statistic consistency (COSEBIs/bandpowers/2PCF) — joint data vectors, independent param sets.
🆕
KiDS. p>0.01 (baryon-driven tier-1 exceedances).
7. Covariance validation
- Analytic vs mock covariance — iNKA/CosmoCov/OneCovariance vs GLASS/SALMO lognormal & Gaussian sims.
✅→🆕 tomographic cov
All. UNIONS (350 GLASS + 10⁴ Gaussian, 5–10%); KiDS (4224 mocks, ~10%); HSC (1404 mocks, SNR>21); DES. - Masking effect on covariance — masked vs unmasked (Gaussian/NG/SSC) ratio.
✅
UNIONS III (vs Troxel+18). Qualitative. - Density-inhomogeneity impact — homogeneous-density mock isolates variable-depth effect on cov.
✅
UNIONS IV (~20% of iNKA/OneCov discrepancy). - Gaussian-only conservativeness — dropping NG/SSC underestimates variance → PTEs conservative.
✅
UNIONS II; KiDS. Argument + rerun. - Hartlap / debiasing correction — correct inverse-cov for finite realizations.
✅
UNIONS II (≤2%); HSC (0.96). Applied. - Covariance stability under the blind — cov must not depend on the hidden cosmology strongly enough to flip conclusions (under a data-vector shift, cov is computed at the reference cosmology and the catalog is untouched, so this is much weaker than under the 2D n(z)-shift scheme, where cov was recomputed per blind).
✅→🆕
UNIONS III (n(z)-shift era: <10% relative; most-conservative PTE adopted). - Iterative (cosmology-dependent) covariance — iterate cov to best-fit cosmology, re-infer.
🆕
KiDS. - Covariance-prescription robustness — iNKA vs OneCovariance (Gaussian+NG) in inference.
✅
UNIONS IV (0.02σ despite 20–25% error-bar diff).
8. Inference / likelihood validation
- Config↔harmonic pipeline cross-validation on mocks — both pipelines on shared GLASS mocks recover fiducial cosmology.
✅→🆕
UNIONS III/IV (341 GLASS). ΔS8 means <0.5σ — explicit unblinding gate. - PSF nuisance null-recovery on (PSF-free) mocks — α,β recover zero on mocks with no PSF systematics.
✅
UNIONS III. Within 1σ. - Goodness-of-fit via mock χ² distribution / PPD — empirical eff-dof from mocks → data-vector PTE.
✅→🆕
UNIONS III/IV (PTE=0.015/0.76); DES (ΔPPD>0.01); KiDS; HSC. p>0.01. - Pairwise probe-combination consistency (ΔPPD) — e.g. shear vs 2×2pt, w vs γ_t compatible before combining.
[needs lens sample]🆕
DES Y3/Y6 (7 checks, ΔPPD>0.01, one waived for prior-volume). - Nuisance-posterior vs prior pushing — tight-prior params must not push edges (flagged DES lens-bin-2).
🆕
DES Y6. Qualitative, blind. - Two-pipeline agreement — independent codebases on identical synthetic vector.
✅ (analogue)
DES (CosmoSIS vs CosmoLike, Δχ²<0.2); UNIONS config vs harmonic teams. - Sampler cross-check — baseline vs alternate sampler.
✅
UNIONS III (Polychord vs Nautilus); KiDS (CAMB vs CosmoPower). Consistency. - Model-variation Δχ² sweep on synthetic data — each modeling choice's bias; importance-sample check.
✅ (analogue)
DES/HSC. 2D bias <0.3σ (DES) / <0.5σ (HSC); Δχ²<0.2 insignificant. - N_Θ / eff-parameters + metric sensitivity on mocks — calibrate suspiciousness; validate tiered metrics on injected-tension mocks.
🆕
KiDS-Legacy Stölzner. Recover injected tension. - Point-estimate (2D-KDE MAP) validation — MAP recovery on mocks vs alternatives (projection effects).
✅
UNIONS IV. Better input-cosmology recovery.
9. External comparison (post-unblinding, not gates)
Combination with Planck / DESI BAO; comparison to DES/KiDS/HSC contours; tension metrics (tensiometer, suspiciousness). Reported as consistency levels, never pass/fail. All surveys; UNIONS III/IV did this.
✅
Blinding interactions
The live candidate for this round is a data-vector shift blind (Muir et al. 2020 / DES Smokescreen style): ξ± is shifted by Δξ = ξ_model(Θ_blind) − ξ_model(Θ_ref), a difference of two theory predictions at a hidden and a reference cosmology. The consistent principle across DES, KiDS, HSC and UNIONS still holds — anything that is a model-independent statement about the data can and should run fully blinded; only comparison to absolute cosmological parameters is gated — and the data-vector shift maps cleanly onto it.
Catalog-level tests are completely unaffected. All PSF diagnostics (ρ/τ, leakage α,β,η, ξ_sys, mean-shear-vs-PSF, brighter-fatter, colour), all catalog null tests (tangential-shear-around-X, mean-shear-vs-position/property, star–galaxy), and all calibration work (m, c, metacal response, selection bias) live upstream of the data vector — they operate on the shear catalogue itself, which the shift never touches. They run directly on real data.
B-mode tests survive the shift essentially exactly. The shift Δξ is a difference of two theory predictions, and cosmic-shear theory predicts pure E-mode, so Δξ carries no B-mode power. COSEBIs B_n and pure-mode ξ_B are constructed to annihilate any pure-E field; therefore a B-mode statistic computed from the shifted ξ± equals the true one — exactly for the continuous statistic, and up to small broad-bin discretization leakage in practice. If that leakage ever matters at the required precision, the fallback is to compute the B-mode statistics pre-shift and blind afterwards. Either way, B-mode-driven catalog and scale-cut decisions are made upstream of the data vector anyway, so the whole B-mode suite — where the hardest blind systematics work happened in every survey (KiDS astrometric masking, HSC field drop, UNIONS size cut) — runs freely under this scheme. Covariance validation, scale-cut determination, and IA-model selection on synthetic data are likewise blind-safe.
Parameter-level checks run on the shifted vector with the parameter axes obscured — the standard DES practice. Goodness-of-fit PTE / ΔPPD, nuisance-prior pushing, and pairwise probe-combination consistency are all computed on chains run against the shifted data vector: the p-values and consistency metrics are meaningful, while the absolute parameter values they'd reveal sit at Θ_blind, not Θ_ref, and stay hidden until the reveal.
KiDS-Legacy lesson. KiDS-Legacy's triple-catalog shape blind actively hampered diagnosing their B-mode systematic, because blinding the catalog forbade direct comparison against previous unblinded catalogs — a powerful localized-systematic diagnostic. A data-vector shift sidesteps this entirely: the catalog is never blinded, so all catalog-level comparisons (including to earlier unblinded UNIONS catalogs) stay available for exactly this kind of forensics.
(History: the previous UNIONS 2D release used an n(z)-shift blind — three random Δz shifts applied to n(z), with cosmological inference run in triplicate across the three blinds. The data-vector shift replaces that scheme for this round.)
Transparency convention. As in both DES surveys, once parameter-level unblinding occurs, no pipeline changes are permitted except pre-registered contingencies (cf. DES lens-bin-2 exclusion) and post-hoc bug fixes that provably touch nothing else (cf. UNIONS Δz −0.030→−0.003 rerun). This convention carries forward unchanged.
— Claude (Fable) on behalf of Cail.
@sachaguer actually thinking about B modes <--> blinding more, I can't think of why B-mode tests can't be run on blinded data vectors in principle since blinding should only affect the E mode. Unless it's a question of numerical issues.. do you remember what motivated the B-mode tests to be ran on unblinded data vectors in Euclid?
Anyway, having skimmed through this terrifying list of checks I don't see any strong reasons not to adopt Muir-style blinding. Planning to write up a PRD for adopting Smokescreen in #234 next.
I am confused. I think of blinding as acting on the E-mode alone. I am not sure whether the B-mode estimators are products of the SGS or not in Euclid but I think we never asked ourselves the question and just made our B-mode tests on the unblinded catalogue anyway. In a nutshell, I am not sure what motivates the question.
i was thinking about blinding the xi + and - that are used to generate COSEBIs and pure EB xi. you could blind either the raw xi products, or the derived E/B products. the former seems simpler, since you only have two products to blind (xi and Cl).
I see. We should test on mocks if it does not make a difference for the B-mode estimation. If that passes, it's fine by me.
OK sounds good. If we shift the raw xi and Cl we can do inference on blinded COSEBIs even if pyccl doesn't implement COSEBIs :)
From the 2026-07-07 tomography call (notes): write down, in one place, the tests we want to run on the tomographic analysis — null tests, validation checks, data-quality diagnostics. A broad bucket to take a first stab at.
One downstream implication (not the goal of this issue): the list will inform the blinding design, since each check either survives a blinded data vector or needs a workaround.