Skip to content

Migrate cosmo_val + Snakemake to the SACC writers #247

Description

@cailmdaley

Make SACC the sole native output format of the cosmology-validation pipeline. Every two-point product (ξ± coarse and fine, pseudo-Cℓ, COSEBIs, pure-E/B, ρ/τ, n(z)) is written through the sacc_io builders into single-statistic SACC part files, and the Snakemake workflow assembles those parts into one terminal SACC file per catalogue version, {version}.sacc. No bespoke writer remains on the data-vector path.

Desired end state

  • Each cosmo_val statistic's write step is replaced by a call to the matching sacc_io builder, producing a part file that holds exactly one statistic plus the self-describing n(z) tracers it references. A producer writes a part with a single save(s, path) call; it never assembles a terminal file and never emits a bespoke format alongside.
  • A dedicated Snakemake assembly rule reads the parts, concatenates them in canonical order, attaches the BlockDiagonal covariance (analysis blocks plus the dense per-pair fine block), and writes the single terminal {version}.sacc. The terminal file is a pure gather; anything that wants only one statistic targets the corresponding part rule, not a second terminal file.
  • Part filenames are DAG internals: consumers bind parts through the workflow's path helpers, never by hardcoding names. {version}.sacc is the only stable name.
  • Blinding rides the DAG here. Blinding is per part, at birth (PRD: SACC data format + Smokescreen blinding #241 §4; CLIs and the assembly-time assertion are delivered by Smokescreen blinding wiring (fork protocol, three theory backends, custody) #252): a blind-init rule fires once per catalogue version, each blindable part rule (coarse ξ±, fine ξ±, pseudo-Cℓ) is followed by blind-part so only blinded parts persist on disk, COSEBIs/pure-E/B are derived from the blinded fine ξ± (born blinded), and the terminal assembly asserts blind_commitment is identical across parts. This issue authors the rules; Smokescreen blinding wiring (fork protocol, three theory backends, custody) #252 supplies the machinery.
  • The bespoke .txt/custom-FITS/.npz writers on the data-vector path are deleted; nothing downstream reads them.

sacc_io is a fixed dependency and the single source of truth for every builder signature. Metadata keys required on every file: catalogue_version, sp_validation_version, created, npatch.

Acceptance

The workflow produces {version}.sacc per version from parts, with only blinded blindable parts persisted and the commitment assertion exercised; the deleted bespoke writers have no remaining readers; the assembled file matches the sacc_io layout contract.

— Fable on behalf of Cail.

Activity

  1. changed the title [-]`cosmo_val` + Snakemake migration to the SACC writers[/-] [+]Migrate cosmo_val + Snakemake to the SACC writers[/+] on Jul 10, 2026
  2. cailmdaley commented on Jul 21, 2026

    @cailmdaley
    CollaboratorAuthor

    Notes on terminal-file contents and size, from the #245 review discussion — relevant when this PR decides what the final gather contains.

    Size budget (5 bins, 15 pairs, production grids)

    The data vectors are trivial; the covariance sets the cost.

    Component Points Covariance on disk
    Analysis vector (ξ± 20θ, pseudo-Cℓ EE/BB/EB 32ℓ, COSEBIs n≤20, pure-E/B, ρ/τ) ~5,300 ~220 MB dense; ~tens of MB block-diagonal
    Integration-grid ξ± (1000 θ × 2 × 15 pairs) 30,000 ~0.5 GB block-diagonal; ~10 GB if dense

    Two consequences:

    1. assemble_covariance now stores covariances block-diagonally (one FITS table per block, Σ block² on disk instead of dense N²; sacc's concatenate_data_sets preserves the blocks through merge). This removes most of the storage cost of keeping products in one file. In particular, Cℓ BB/EB and ρ/τ are cheap (~17 MB even with full cross-polarisation covariance) — no storage reason to exclude them.
    2. The integration-grid ξ± is the only expensive component. It is an intermediate: consumed by the COSEBIs/pure-E/B derivation before the gather exists. Recommendation: drop it from the terminal {version}.sacc and let it persist as its per-part file — Snakemake provenance covers traceability. That keeps the terminal file at tens-of-MB scale with everything scientifically load-bearing inside.

    Other pieces relevant here

    • Covariance variants (OneCov / NaMaster / …): a SACC file holds one covariance, so variants are separate merged files reusing the same per-part data vectors — cheap under block-diagonal storage. Declaration/naming of variants belongs to this PR.
    • Flexible containers: the sacc_io writers are being relaxed so optional components (e.g. Cℓ EB) need not be present; consumers select what they need and fail loud on absence. The gather logic here should assume partial files are legal.
    • Integration grid size is a parameter, default 1000 θ bins (1k-vs-10k comparison in the B-modes paper found no substantial difference; we may push higher for final production).

    — Claude (Fable) on behalf of Cail.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions