Skip to content

Exposure-level HealSparse maps: exposure count and flagged pixels - #887

Open
cailmdaley wants to merge 150 commits into
developfrom
feat/instrument-defect-map
Open

cailmdaley wants to merge 150 commits into
developfrom
feat/instrument-defect-map

Conversation

@cailmdaley

@cailmdaley cailmdaley commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Builds on #886 (merged). Closes #878.

Two HealSparse maps per campaign, at the UNIONS mask ladder's resolution (nside 131072, 1.6" pixels):

  • nexp_<run>.hsp (uint8): per sky pixel, the exposures whose CCD has a valid PSF model and covers it.
  • nflagged_<run>.hsp (uint16): per sky pixel, the flagged CCD pixels (bad columns, saturated pixels, bleed trails, cosmics) of those exposures. One sky pixel holds ~74 MegaCam pixels.

Neither map applies a cut.

What these maps are for

For Cail and Martin to weigh:

  • Per object, the catalogue already has what shape measurement needs. ngmix zero-weights flagged pixels, and the veto and epoch drops are applied there. The catalogue carries N_EPOCH, NGMIX_N_EPOCH_FAILED and the interpolation count. The maps serve the footprint and the randoms.
  • nexp is the necessary product. Randoms have to follow the same N_EPOCH ≥ k cut as the galaxies.
  • nflagged is a diagnostic of where the detector is bad. It costs nothing extra because it comes from the same fragment. Don't cut on it as is: randoms that mimic the galaxy selection would need "exposures still usable after the defect veto", not raw flagged-pixel counts. Open question: keep it as a diagnostic, replace it later with a usable-exposure map, or drop it.
  • Next step, a cheap validation: compare N_EPOCH with nexp at object positions. They should agree except near defects and where epochs were cut.

How it works

  • exp_maps runs per exposure, after exp_persist, and writes one fragment beside the PSF tar.
    • The fragment is 0 outside the exposure, and 1 + the number of flagged pixels inside each CCD with a PSF model.
    • Those CCDs are the ones with a validation_psf file in the tar.
    • The geometry is each split image's WCS over its DATASEC: the 2048×4612 imaging area, without the overscan border, which the flag image marks with value 3.
  • exposure_maps sums a campaign's fragments into the two maps.
  • Reclaimed exposures are handled as for the star catalogue: the fragment is depended on directly, so the exposure chain is never rebuilt from VOS (durable_edge, now shared by both).
  • clean_exposure waits for the fragment.
  • Fake-PSF sims build neither map.

Relation to #797 and #890

On real data

These figures use smk-g14-fz, tile 189.305, and all 7 of its exposures. The flag splits were re-made on Nibi.

one exposure

tile

  • nexp. Each dither's chip gaps show as strips one exposure lower. Over the tile, nexp ≥ 3 holds for 81% of the area.
  • Flags. Inside DATASEC, 0.28% of pixels are flagged, mostly one-pixel columns and cosmics.
    • nflagged > 0 covers 6.9% of the tile, because a single flagged pixel marks a whole 1.6" sky pixel.
    • Weighted by count, the flagged area is 0.33% of the exposure area.
    • The flags are almost all static: Jaccard 0.994 between exposures 2603236 and 2603237, measured on Nibi. Bad columns recur at each dither offset.
  • Cost.
    • exp_maps: ~8 s and 0.15 GB per exposure. A fully flagged CCD takes 3.8 s and 1.5 GB; the rule requests 3 GB.
    • Merge: ~0.9 s per exposure. At DR6 scale it needs ~70 GB in one job.
  • Against v1.6. Martin's v1.6 coverage.hsp (nside 2048) has nexp ≥ 3 over 60% of this tile, against 81% here. The exposure sets probably differ; this is not checked.

Tests

  • Suite green in the dev image: 967 passed, 1 skipped with -m "not slow"; tests/workflow: 19 passed.
  • The params pin gains exp_maps and exposure_maps; no existing rule's params digest moves.

Claude Opus 5.5 on behalf of Cail

🤖 Generated with Claude Code

https://claude.ai/code/session_012yfkBvkzPv33iJxXZ6jAxd

Cail Daley and others added 29 commits September 9, 2026 08:57
persist_exp.py copies one exposure's named PSF products off /scratch onto the
persistent root and records what went, with sizes. The threat it answers is the
60-day purge, not clean_exposure: run_dir is scratch and products_dir is
/project, so the only way a per-exposure product outlives its campaign is to
leave the filesystem. Exempting files from reclamation would not have done it.

The search is recursive beneath the PSF chain's four module output dirs, because
setools writes into mask/, rand_split/, new_cat/, plot/ and stat/ rather than
flat -- so the config's patterns stay plain file names and the layout stays ours.
A pattern that matches nothing is a recorded warning (setools rejects sparse
CCDs); nothing matching at all is a failure, since a green manifest over an
empty copy is what would let reclamation delete an unsaved exposure.

config.yaml's persist_exp: defaults to validation_psf-*.fits -- the psfex_interp
VALIDATION catalogue, the rho/tau statistics input, the minimum. The opt-in
candidates are documented there with what each buys; sizes are still to be
measured. PSFEx residuals and XML are not candidates as the chain stands: the
committed default.psfex sets CHECKIMAGE_TYPE NONE and WRITE_XML N.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
One rule per exposure, output = ONE manifest on the persistent root at
<products_dir>/exp/<shard>/<exp>/manifests/exp_persist.json. Not a directory()
output: what we want written down is which files were copied and how big each
was, and a directory attests only that a directory exists. Byte-stable, so a
no-op rerun does not move an mtime clean_exposure reads.

Its own rule rather than a cp on the end of exp_psf, and that is the whole
point: the keep list rides on params, so adding a pattern reruns seconds of
copying instead of four hours of PSF fitting per exposure.

A localrule, by the arithmetic that made exp_star_cat one -- a few MB of cp,
~20k of them at DR6 scale, each shorter than the scheduling latency that would
submit it. The mid-chain grouping constraint does not bite: its neighbours are
exp_psf (too heavy to fuse) and clean_exposure (local already).

clean_exposure gains the manifest as an input, so a store is never reclaimed
before its keepers have left scratch -- conditional only on there being a keep
list, since "keep nothing" must not become a dependency on a rule that would
fail for having nothing to copy.

rule all requests the persist manifests DIRECTLY, not only through
clean_exposure: the purge takes the store whether or not clean: is on, so
hanging the copy off reclamation alone would lose everything in a clean:false
campaign. Cleaned exposures are excluded -- their exp_psf manifest is gone, so
asking would rebuild the chain from VOS, and a tombstone already means the copy
happened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… omits it

run_report disk-scans the scratch run_dir, and exp_persist's manifest is the one
exposure manifest that lives on products_dir instead -- the placement that makes
it survive clean_exposure. Listed in EXP_STAGES it would read as "not run" for
every exposure in the campaign, so it is deliberately absent, with the reason on
the line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e copies

Inodes, not bytes, bind on /project (~1 M-file group quota): smk-m2 measured
~200 loose files per exposure with all candidates on — 25k for 64 tiles, ~2 M
at DR6 scale, for 7 GB. persist_exp.py now writes <products_dir>/exp/<shard>/
<exp>/psf/<exp>.tar (uncompressed, flat members, deterministic: ownership
zeroed, sorted, tmp-cmp-mv so a no-op rerun keeps the mtime) and the manifest
lists every member. Manifest path, rule wiring and params are unchanged.

config.yaml's candidate table carries the smk-m2 per-exposure sizes;
psfex_cat and star_stat are marked unmeasured (no live store held them).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015qtLUV3bVLPV5p6un7aTFR
An orphaned tar.tmp on /project is an inode nothing revisits — the leak the
tar design exists to avoid. try/finally around both tmp writes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015qtLUV3bVLPV5p6un7aTFR
develop no longer has exp_star_cat / star_catalogue (PR #847); the three
comments that cited exp_star_cat as the localrule precedent now cite
clean_exposure, which makes the same argument.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Every rule in this workflow was per unit, and the two products downstream
analysis actually opens are per campaign. So a run ended one file short on
each side and both merges were a manual pass afterwards. These are the last
links of chains the workflow already had.

star_cat_merge stacks every exposure's every CCD's validation_psf into one
<products_dir>/full_starcat-0000000.fits — the rho/tau statistics input, at
the path sp_validation hardcodes. It reads the members straight out of the
per-exposure tars exp_persist wrote (tarfile + BytesIO); unpacking ~800k files
to merge them would defeat the tar's whole purpose. The stacking is
MergeStarCatPSFEX, the class the old merge_starcat_runner called, so the
column list keeps exactly one definition. That class gains one thing: an entry
may be [fileobj, name] rather than [path], fits.open taking the first element
and the CCD_NB regex the last — the same string for the runner's one-element
entries.

final_cat_merge collects every ready tile's final_cat into
<products_dir>/final_cat_<campaign>.hdf5: one dataset per tile under a group
named for the campaign, the final_cat.param columns, an n_tiles attribute.
That schema is sp_validation's reader's, so it is fixed. The COLUMN
EXTRACTION reuses scripts/python/create_final_cat.py (read_param_file,
read_data, copy_data) so the column grammar keeps one definition; the file is
written here, because that script's own discovery walks a directory layout
this workflow does not have and groups by a unit ShapePipe v2 has dropped.
Two places where the reference implementation is not reproducible are pinned
down at the call site rather than copied: copy_data leaves every non-requested
column as uninitialised memory, and read_param_file's column order varies with
the process hash seed. bin/sp now snapshots the repo's scripts/ so the
campaign pins that file like everything else it runs.

Neither merge puts its input paths in its shell: ~20k of them is an order of
magnitude over Linux's 128 KiB MAX_ARG_STRLEN for one argv entry. Each job is
handed the two small files the Snakefile itself started from — the tile list
and the run index — and DERIVES the same set from them, through readers that
now live in build_index.py beside the schema. The rule's params carries that
set's fingerprint, which is the rerun trigger, and the equality of the two
sides is what makes the trigger mean anything: a glob over products_dir would
merge tiles or exposures from an earlier, larger tile list sharing the root,
rows no trigger could see.

star_cat_merge depends on a live exposure through its exp_persist manifest and
on a RECLAIMED one through its tar, which no rule declares and which therefore
requires nothing to be built. Requesting a reclaimed exposure's manifest
instead rebuilds its whole chain from VOS, and ancient() does not prevent that:
measured on smk-g6 with one reclaimed exposure given a manifest by hand, the
dry run grew exp_get_images, exp_split, exp_psf and exp_persist jobs.
Reclaimed exposures belong in the star catalogue — carrying their PSF products
off scratch is what exp_persist is for.

Both rebuild rather than append, so the output is a function of its input set:
byte-stable on a no-op rerun (tmp-then-cmp-then-mv), rebuilt when a unit is
appended. Neither is a localrule — one job over ~20k units is real work.
star_cat_merge produces no job, and a parse-time warning rather than a runtime
failure, when persist_exp keeps no validation_psf-shaped file.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
PERSISTENCE KEYED ON THE WRONG FILE, and the consequence was the avalanche it
exists to prevent. persist_targets() skipped an exposure only when its SCRATCH
tombstone was there, but clean_exposure is one of two ways a store disappears
and the other leaves nothing behind: /scratch is purged on a 60-day window
whether or not the workflow reclaimed anything. After a purge — or on any
campaign with clean: false — every exposure looked live, its persist manifest
was requested, its exp_psf manifest was gone, and snakemake rebuilt the whole
exposure chain from VOS.

The test is now exp_store_reclaimed(), used by persist_manifests() and by
star_cat_merge's live/reclaimed split so the two cannot disagree, and it takes
BOTH pieces of evidence that an exposure once had a store: a tombstone, or a
persisted manifest with no exp_psf manifest beside it. Both are needed. The
manifest clause alone regressed the tombstone case — measured on smk-g6, whose
126 exposures were cleaned before exp_persist existed and so have no manifest:
the dry run grew 126 exp_get_images, exp_split, exp_psf and exp_persist jobs,
a whole campaign rebuilt from VOS. The tombstone clause alone is the purge bug
above. Neither file can be dropped, because the parse cannot otherwise tell a
store that is GONE from one not BUILT yet, and a fresh campaign must still be
asked to persist.

OVERLAPPING KEEP PATTERNS FAILED EVERY EXPOSURE. persist_exp treated a file
matched by two patterns as a flat-member name collision, so
'validation_psf-*.fits' alongside '*.fits' — an ordinary way to write a keep
list — aborted the pack. Two DIFFERENT paths on one member name is still
fatal; the same path twice is now one file, recorded under the first pattern
that matched it.

merge_star_cat MATERIALISED THE WHOLE CAMPAIGN before merging a row: every
member's bytes, ~2 MB per exposure, ~40 GB at DR6's ~20k exposures against a
rule asking for 16 GB. TarMembers hands the merge class the same entries one
tar at a time, so peak memory is one member plus the class's own accumulators,
which are the unavoidable term. It keeps __len__ off the manifests so the
count is still logged before a tar is opened.

THE STAR MERGE RERAN ON BOOKKEEPING. Its fingerprint was over input PATHS, and
an exposure's edge flips from its manifest to its tar the moment its store is
reclaimed — so every reclamation pass reran the merge over identical content.
It is over the exposure IDS now, which move only when the set does, and which
are what the job derives on its own side. final_cat_merge's is over tile ids
for the same reason.

merge_final_cat's MISSING-COLUMN CHECK WAS UNREACHABLE. create_final_cat's
read_data wraps its column selection in a bare `except:` that prints and falls
through, so a missing column left its return values unbound and the caller got
UnboundLocalError from the return statement, naming nothing. The columns are
checked against the catalogue's own header before read_data is called, and the
message now names every missing one.

A TILE LISTED TWICE killed the merge on the second create_dataset. The tile
list is appended to by hand, so duplicates happen; campaign_tiles() dedupes it
order-preserving, and the Snakefile's TILES does the same so the fingerprint
and the job's derived set still name the same set.

merge_class OFFERED MCCD AND SETOOLS while only MergeStarCatPSFEX had learned
the [fileobj, name] entry shape. Both now take the entry's name from its last
element like PSFEX does — unchanged for the module runner, whose entries are
[path]. Setools needs one thing more before it can read a tar (it hands
file_io input_file_list[0][0] as a template path), and merge_class says so
rather than implying otherwise.

Also: README no longer lists scripts/sp_rule.py, which does not exist.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Three defects in scripts/python/create_final_cat.py, fixed where they live
rather than worked around in the workflow rule that now calls it. A hand-run
of the tool deserves them as much as the rule does, and files it has already
written carry the first one.

copy_data allocated np.empty with the SOURCE catalogue's full dtype and then
filled only the requested columns, so every column NOT in the parameter file
reached the hdf5 file as uninitialised memory: meaningless values, and
different bytes on every run over the same inputs. It allocates the requested
columns alone now, in the source catalogue's order. The parameter file says
what the merged catalogue is for; those are the columns it gets.

read_param_file returned list(set(...)), whose order varies with the process's
string hash seed. Column order is part of a structured dtype and therefore
part of the file, so two runs over the same inputs disagreed. Ordered dedup
via dict.fromkeys. (The duplicate-count message also only printed for more
than one duplicate, and said {n} literally.)

read_data wrapped its column selection in a bare `except:` that printed and
fell through, leaving its return values unbound — so a missing column surfaced
to the caller as UnboundLocalError from the return statement, naming neither
the file nor the column. It raises a KeyError naming the file and every
missing column, in parameter-file order.

process()'s own create_dataset follows the array copy_data returns rather than
the source dtype, which are no longer the same thing.

merge_final_cat.py drops the equivalents it had been carrying at the call site
and relies on the fixed functions. The fixture hdf5 is byte-identical either
way (md5 6de2d261…): the workaround and the fix produce the same file, which
is the point.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Both merges had a constant mem_mb, which is wrong by however much a campaign
differs from the one it was tuned on — and these are the only two rules whose
single job grows with the whole campaign. Both are now measured slopes,
evaluated against the campaign's own bytes at DAG build, still * attempt.

STAR SIDE, measured on this login node in the campaign container over
synthetic tars, 20 and 80 exposures of 40 CCDs x 400 stars:

    input members      peak RSS (getrusage RUSAGE_CHILDREN)
    32.3 MB            383 MB
    129.0 MB          1313 MB

a slope of 10.1x input bytes over a ~73 MB interpreter floor. Tenfold because
MergeStarCatPSFEX accumulates every column into python LISTS of python floats
before building the output arrays. THE CONSEQUENCE IS A CEILING and the
Snakefile says so: at ~2 MB of members per exposure a 16 GB job merges roughly
800 exposures, and DR6's ~20k would want ~400 GB. A full-survey full_starcat
needs that accumulation changed to preallocated arrays or a two-pass count —
a change to MergeStarCatPSFEX, not to this rule, and not in this PR. The
formula is honest about the slope so the job asks for what it will use and
fails at submission rather than most of the way through.

TILE SIDE, measured against smk-g6's real catalogues, 2 tiles (73.9 MB in,
largest 39.6 MB) and 6 tiles (235.5 MB in, largest 47.7 MB): peak RSS 129 MB
and 139 MB. FLAT in the tile count, because the merge holds one catalogue at a
time — so it is sized on the LARGEST tile at ~3x, not on the total. Runtime is
the total, since every tile is read end to end. On smk-g6's 64 tiles the rule
resolves to mem_mb=1002, runtime=51, against the flat 8000/120 it had.

Sizes come from stat() on the tar or the catalogue, falling back to the
measured per-unit default when a fresh campaign has not produced it yet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Closes the readability half of #844, and makes the
2026-09-08 call's request — keep the PSF model — something you can write down
as `psf_model` rather than `*.psf`.

`persist_exp:` entries are now names from a catalogue in persist_exp.py, which
is the single source of truth for what each one means, what it costs per
exposure and what keeping it buys:

    psf_validation  validation_psf-*.fits        2.0 MB
    psf_model       *.psf                        2.8 MB
    psfex_cat       psfex_cat-*.cat            unmeasured
    star_selection  star_selection-*.fits       24.5 MB
    star_train      star_split_ratio_80-*.fits  19.9 MB
    star_test       star_split_ratio_20-*.fits   7.1 MB
    star_stats      star_stat-*.txt            unmeasured

`persist_exp.py --list-products` renders it, and config.yaml's block IS that
rendering rather than a second copy of it — the old block was a long comment
listing globs and their sizes, maintained by hand beside the code that
actually knew them.

THE DEFAULT BECOMES psf_validation + psf_model, ~4.8 MB per exposure. The
model is the single most capability-adding thing an exposure can keep: with
it the PSF can be re-interpolated at any position later without rebuilding the
chain from VOS, and without it that capability dies with the /scratch purge.

A raw glob is still accepted as an escape hatch for a file the catalogue does
not name yet. The test is syntactic and cheap — a glob metacharacter or a dot
means glob, a bare identifier means name — so `*.psf` and `psf_model` cannot
be confused. An unknown NAME is a parse-time WorkflowError listing the valid
ones, not a silently empty keep or a per-exposure failure an hour in.

The manifest records both: `products` as written, `patterns` resolved, and
each member's own `product`. star_cat_merge's gate and its member glob resolve
through the same catalogue, so adding a product cannot leave the two
disagreeing, and the "add this to persist_exp" hint now names the product.

Tile-side retention is explicitly out of scope and noted as such in
config.yaml: final_cat is the only tile product that persists today, and it is
written straight to products_dir by tile_make_cat.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
All three merge classes built every output column by extending a python list
with one value per star: `x += list(data["X"])`. Four bytes of float32 payload
became a 32-byte python object plus an 8-byte pointer in an overallocating
list, measured end to end at ~10x the input bytes — which put a full-survey
full_starcat (~20k exposures x 40 CCDs) at ~400 GB of RAM and out of reach of
any node.

One array per input catalogue per column, concatenated once at the end. Same
values, same order, same dtypes — np.array() over a list of numpy scalars and
np.concatenate() over the arrays they came from agree on both. The stacking
helper empties the list it is handed, which is half the saving: concatenate
holds the chunks and the result at once, so releasing column by column peaks
at one campaign plus one column rather than two campaigns.

MEASURED on the same two fixture points, 20 and 80 exposures of 40 CCDs x 400
stars:

    input members    peak RSS, before    after
    32.3 MB          383 MB              238 MB
    129.0 MB        1313 MB              740 MB

10.1x -> 5.5x, over a ~62 MB interpreter floor. The rule's mem_mb factor
follows. A 16 GB job now merges ~2200 exposures rather than ~800.

THE REMAINING 5.5x IS THE OUTPUT SIDE: file_io writes every float column as
FITS 1D, so float32 inputs become a float64 table astropy then buffers. That
is a change to the output FORMAT, which is what sp_validation reads, and a
different decision from this one.

BYTE-IDENTICAL OUTPUT, both ways in. The workflow's tar path and the module
runner's plain [path] path produce the same file as before the change, md5
f7caa1cf… on the fixture — the runner path checked by calling
MergeStarCatPSFEX directly with [[path]] entries as merge_starcat_runner
builds them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
`persist_exp:` was doing two jobs. It decided what a campaign keeps for later
— a retention question, and the user's — and it also decided whether the
campaign's star catalogue could be built at all, because star_cat_merge
existed only when the keep list happened to name something
validation_psf-shaped. That made the survey's PSF diagnostics an opt-in, and a
typo away from silently absent.

exp_persist now packs psf_validation for every exposure whatever the config
says. It is the merged catalogue's PROVENANCE — a full_starcat with no
per-exposure inputs beside it cannot be audited, re-cut or recomputed after a
purge — and it is what keeps APPENDING TILES CHEAP, since a tile added next
month brings exposures whose catalogues must join the existing stack and the
alternative is rebuilding their chains from VOS. ~2 MB per exposure: ~40 GB
and ~40k inodes at DR6 scale against a ~1 M-inode group quota, which is the
price of being able to say where the number came from.

`persist_exp:` is therefore purely additive retention, defaulting to
psf_model, and an EMPTY list is now a coherent instruction rather than a
switch that turns persistence off: the tar holds the merge's inputs and
nothing else. The keep-list gate on star_cat_merge and its parse-time warning
are gone with it, as is clean_exposure's conditional edge on exp_persist —
there is no configuration left under which that rule has nothing to wait for.

No transient/cleanup knob: these files are kept, not staged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…t all

final_cat_merge on smk-g6's real catalogues failed on three of the 67 columns
the parameter file asks for. Two of them the file should not have been asking
for.

IMAFLAGS_ISO is DROPPED. The tile-side SExtractor runs with FLAG_IMAGE = False
and DOT_PARAM_FILE = default_noimaflags.param (config_tile_Sx.ini), so the
column is never written into a tile catalogue and asking for it could only
fail. Instrument flags reach the pipeline on the EXPOSURE side, where
exp_split delivers the flag image and SExtractor reads it.

NGMIX_MOM_FAIL is RENAMED to NGMIX_MCAL_TYPES_FAIL, which is what f0fca23
called it in June and what the catalogues carry.

NGMIX_NEIGHBOUR_FLAG STAYS. It was added to make_cat in fa6e001 on
2026-07-12 and it is the blend flag the systematics tests need. smk-g6's
catalogues do not have it — checked on the files — so final_cat_merge still
fails there, and that failure is correct: the campaign's catalogues are
missing a column the analysis wants, which is a fact about the data and not
about this file. Its launch snapshot is gone (only .snakemake survives under
smk-g6-state), so the run's HEAD cannot be read back; what remains is that its
catalogues carry NGMIX_MCAL_TYPES_FAIL (June) but not NGMIX_NEIGHBOUR_FLAG
(July), consistent with a snapshot taken between the two.

NO MASK COLUMN REPLACES IMAFLAGS_ISO, AND THE FILE NOW SAYS WHY. The intended
replacement is make_cat's per-band MASK_<band>, queried from the healsparse
maps named by MASK_EXT_PATHS — and the workflow sets none: config_tile_Mc.ini
has no such entry, save_mask_ext_data is never called, no MASK_<band> column
exists in any catalogue this workflow has produced, and smk-g6's carry none.
Naming one here would fail every merge on every campaign. The merged catalogue
therefore carries no mask information today; that is a CONFIG gap, and closing
it is setting MASK_EXT_PATHS first and adding the column names second. No
healsparse map is staged under /project/def-mjhudson yet.

With NGMIX_NEIGHBOUR_FLAG set aside, the merge runs clean over all 64 of
smk-g6's real catalogues: 2.50 GB read in 20 s at 151 MB peak RSS, producing a
0.94 GB hdf5 of 65 columns. That also confirms the tile-side sizing — the rule
asks for 1002 MB and 51 minutes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
The array accumulation landed a commit ago took the star merge from ~10x the
input bytes to ~5.5x. What was left was the accumulation itself: one array per
input catalogue, then a concatenate that has to hold its inputs and its result
at the same time.

MergeStarCatPSFEX now makes two passes. The first reads only the FITS HEADER
of every input — NAXIS2, the row count — and touches no data block; the second
allocates each output column once, at its exact final length, and fills it
slice by slice. There are no chunks and no concatenate, so peak memory is one
output plus one input catalogue.

The workflow's tar reader hands over the archive's own file object rather than
a BytesIO of the whole member, so the counting pass costs a header rather than
a member. Both it and a plain list of paths are iterable twice, which the two
passes require; a one-shot iterable would fill nothing on the second pass, so
the merge checks that the passes agree on the row count rather than writing a
catalogue padded with uninitialised memory.

MEASURED on the same two fixture points, 20 and 80 exposures of 40 CCDs x 400
stars:

    input members   python lists   arrays+concat   two passes
    32.3 MB          383 MB          238 MB          221 MB
    129.0 MB        1313 MB          740 MB          661 MB
    slope             10.1x           5.5x            4.8x

The rule's mem_mb factor follows. A 16 GB job now merges ~1300 exposures.

THE REMAINING 4.8x IS THE OUTPUT SIDE: file_io writes every float column as
FITS 1D, so float32 inputs become a float64 table astropy then buffers — 141 MB
of table for 78 MB of payload at the 80-exposure point. What stands between
here and a full-survey full_starcat is that format, not the merge.

MergeStarCatMCCD and MergeStarCatSetools keep the array accumulation. Their
process() computes campaign-wide statistics over the same columns, so a
two-pass rewrite there is a larger change with no consumer today — psfex is
what every campaign runs.

BYTE-IDENTICAL OUTPUT, both ways in: the workflow's tar path and the module
runner's plain [path] path both give md5 f7caa1cf… on the fixture, unchanged
through both rewrites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
The rule read every tile's catalogue on every run, because a DAG output must
be a function of its input set and rebuilding is the simple way to guarantee
that. At DR6 scale it is also ~800 GB of IO to add one 35 MB tile.

It now brings the file INTO AGREEMENT with the campaign: a tile with no
dataset is added, a dataset whose tile has left the campaign is deleted, a
dataset whose source catalogue CHANGED is re-read, and one that agrees with
its source is left alone, unread. Each dataset records its source's size and
mtime as attributes, and a mismatch is what changed means — which is also what
keeps the file from drifting from its inputs the way an append-only tool does.
create_final_cat.py's own process() implements only the append-only half of
this, skipping any tile already present whatever the file on disk now says.

WHAT IS AND IS NOT A FUNCTION OF THE INPUT SET, since this is the guarantee
being traded. The file's CONTENT is: the same tiles with the same catalogues
give the same datasets, the same columns and the same n_tiles, whether they
arrived at once or one campaign at a time. Its BYTE LAYOUT is not, because
hdf5 lays a group out in the order things were added. That is the price of not
re-reading the campaign.

UNTOUCHED ON A NO-OP, which is stronger than the byte comparison it replaces
and cheaper to establish: reconciling is PLANNED against a read-only open, and
an empty plan never opens the file for writing, so its mtime cannot move. A
non-empty plan is carried out on a copy which is then moved into place, so a
crash mid-merge leaves the old catalogue intact.

VERIFIED on a three-tile fixture: build (3 added), no-op (unchanged, mtime
identical to the nanosecond), append one tile WHILE AN EXISTING TILE'S
CATALOGUE IS UNREADABLE — chmod 000, which succeeds and reports 1 added, so
the existing tiles were demonstrably not read — rewrite of one catalogue
(1 refreshed), and dropping two tiles from the list (2 removed, datasets gone,
n_tiles 1). A from-scratch build of the same set is byte-stable across reruns.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
OPTIONAL COLUMNS WERE DECIDED ONCE FOR THE WHOLE MERGE. MergeStarCatPSFEX read
MAG/SNR/ACCEPTED out of the first catalogue's dtype and applied that verdict to
every file behind it, so a merge over a mix of ordinary and pix2wcs-converted
catalogues was wrong in both directions: ordinary first raised KeyError on the
first converted file, and converted first SILENTLY ZEROED the real values of
every ordinary file. The dtype now comes from any file that carries the column
and pass 2 asks each file for its own schema, so only the files that actually
lack a column are zero-filled. Both orderings verified on fixtures.

A FAILED psfex_interp COULD GET A GREEN MANIFEST, and clean_exposure takes that
manifest as its go-ahead to delete the store. With the default retention list,
an exposure whose interpolation failed but whose PSFEx model landed had a
non-empty match set, so the pack succeeded and the stars went with the store —
unrecoverable short of rebuilding the chain from VOS. psf_validation is not
optional: nothing matching it now fails the job, while the store is still on
disk. Retention products that match nothing stay warnings.

SHRINKING THE KEEP LIST DELETED PRODUCTS FROM /project. The list rides on
`params`, so editing it reruns the pack — which rewrote the tar without what
had been dropped, on the backed-up filesystem, with the scratch store it came
from usually already reclaimed. RETENTION IS NOW ADDITIVE: an existing tar is a
FLOOR, its members carried into the new one whatever the current list says, so
a config change can only ever add. Removing a product is a deliberate act on
products_dir, not a config edit. Verified: pack with [psf_model], rerun with an
empty list, the .psf is still there under its own product name and the tar is
byte-identical.

THE COLUMN SET REACHED NO RERUN TRIGGER. final_cat_merge's reconcile keyed
staleness on each source catalogue's size and mtime, and the column set is not
a source catalogue: final_cat.param arrives through `params`, and the hash
covered workflow/scripts/ only, not scripts/python/create_final_cat.py. So this
PR's own edit to final_cat.param would have left every dataset in an existing
hdf5 written to the old schema with nothing to notice. The file now carries a
digest of the resolved column list on its root and refreshes every tile when it
moves, and MERGE_FINAL_HASH covers all three files the rule's behaviour comes
from. Verified: build, edit the parameter file, rerun -> 3 refreshed.

STAR_CAT_MERGE'S MEMORY WAS SIZED ON THE TAR, which holds whatever the campaign
retains, while the merge reads the psf_validation members alone. Measured on a
fixture with a 3 MB PSF model kept: the tar is 92x the members it will read,
and the default retention is 2.4x. It also jumped discontinuously as exposures
were packed. The manifests record the product each member came from — exactly
so this is answerable without opening a tar — so the sizing sums those members.

NO REQUEST WAS CAPPED. A mem_mb above the partition maximum is a job SLURM
never schedules and snakemake never diagnoses: it sits PENDING while the
campaign looks alive. Both merge formulas grow with the campaign, so at some
size they cross it. `max_mem_mb:` (default 750000, for Nibi's 766 GB standard
node) caps both, with a parse-time warning naming the rule that was capped.

copy_data ORDERED ITS OUTPUT BY THE SOURCE CATALOGUE, which made the merged
dtype a property of the catalogue rather than of the parameter file: the
ordered dedup added to read_param_file had no effect, and two tiles written by
different ShapePipe versions landed in one group with two different structured
dtypes, which np.concatenate refuses. It orders by param_list now. Verified on
two catalogues with reversed column orders and an extra column: one dtype,
concatenate works. The fixture hdf5 md5 moves with the column order,
d2882294… -> 43ff946d….

RECONCILE LEAKED SPACE. It copied the file and deleted datasets in place, and
HDF5 never reclaims that, so every refresh of a tile grew the file by that
tile. A plan that removes or refreshes anything now builds the tmp fresh,
moving the datasets it keeps across with h5py's own group copy — a
dataset-level copy that never reads a row into numpy — so the result is
compact; pure-append plans still copy and append. Verified: five successive
full refreshes leave the file the same size, and dropping a tile shrinks it.

Also: comments referring to the deleted parse-time keep-list gate are gone;
`-s add` is documented as what it is (accepted by create_final_cat.py's
validator, then falling through to the ordinary walk, so not a way to add one
tile by hand); and three latent issues are noted where they live rather than
fixed — hdu.columns.dtype ignoring TSCAL/TZERO (no validation_psf column is
scaled), MergeStarCatSetools rebinding its ellipticity accumulators so only the
last file's reach the output (pre-existing, setools is not wired to any
workflow path), and the .tmp a SIGKILL can orphan next to the catalogue.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
… the tile one

The campaign's two products behaved differently for no reason anyone chose.
The shear catalogue reconciled — an append read the appended tiles — while the
star catalogue was one flat FITS table that had to be restacked from every
exposure the campaign had ever seen to add one: ~40 GB of members at DR6 scale
to add ~2 MB, held in memory while it happened. Now they are the same thing.

<products_dir>/full_starcat_<campaign>.hdf5, one dataset per exposure at
exposures/<exp>, each holding that exposure's every CCD's rows with a CCD_NB
column, an n_exposures root attribute and the same column digest the tile side
carries. Named for the campaign exactly as the shear catalogue beside it is.

CCD_NB IS AN INT: it is parsed out of the member name, where it is always
digits, so a string buys nothing and costs 8 bytes a row against 4. DTYPES ARE
NATIVE: float32 stays float32, where the FITS writer widened every float column
to 1D, doubling both the file and the peak memory of the job that wrote it for
no information.

MEMORY IS NOW FLAT IN THE CAMPAIGN — one exposure at a time — so the rule is
sized on the largest exposure's members rather than the campaign's, and the
~240 GB a DR6-scale flat table would have wanted is simply not a number any
more. The Snakefile's sizing block keeps the measurements that got us here,
because they are the argument for the format.

THE RECONCILE MACHINERY IS NOW ONE MODULE, workflow/scripts/hdf5_reconcile.py,
used by both merges rather than duplicated: plan against a read-only open,
add/refresh/remove, refresh everything when the column digest moves, compact
rewrite when anything is removed or refreshed, untouched on a no-op. Writing it
twice would have been two chances to disagree about what an output owes its
inputs.

THE WORKFLOW NO LONGER CALLS MergeStarCat* AT ALL, so merge_star_cat.py drops
the shapepipe import and the psf-model switch, and the [fileobj, name] entry
shape those classes learned for it is REVERTED — with no caller it was upstream
surface with nothing behind it. What stays upstream is what fixes the module
runner's own problems: the two-pass allocation, and asking each file for its
own optional columns instead of deciding once for the merge. The runner path is
byte-identical to before all of it, md5 f7caa1cf… on the fixture.

VERIFIED on the fixtures: build (2 added); no-op (unchanged, mtime identical to
the nanosecond); append one exposure with the others' tars at chmod 000, which
succeeds and reports 1 added, so they were demonstrably not read; remove one
(1 removed, dataset gone, n_exposures 2). Every one of the 16 columns equals
the FITS version's values. The tile side's compaction sequence was re-run
against a fix this work exposed — the keep-what-changed path called
Dataset.copy, which does not exist, and only bites when a plan both rewrites
and keeps something.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
path_hash KILLED EVERY INVOCATION over a file one rule needs. It runs at
module level, so a snapshot without scripts/python/ raised FileNotFoundError
during the parse — taking `sp --unlock`, `sp report` and every dry run with
it, and taking them with a bare traceback rather than the diagnosis
merge_final_cat.py already carries for exactly this case, which the job could
never reach. The hash degrades to a sentinel and warns once; the parse
survives, the rule still exists, and the job prints the message written for
it. Verified both halves with the file moved aside.

THE STAR MERGE'S OPTIONAL COLUMNS HAD NO CANONICAL DTYPE. MAG/SNR/ACCEPTED
took the dtype of whichever file carried them, falling back to X's float32
when none did — so an exposure whose files are all pix2wcs-converted got a
float32 ACCEPTED while its neighbours got int32, and datasets under exposures/
differed in dtype. np.concatenate refuses that, and no digest can repair it
because nothing about the schema CHANGED. The three are pinned (int32,
float32, float32) and cast. Verified on two exposures, one carrying them and
one not: identical dtypes, concatenate works.

A MISLABELED MANIFEST ABORTED THE CAMPAIGN. is_member() accepts a member by
product label OR by file name — right for "did this exposure keep the
product" — but read_exposure() selects by name alone, so an entry labelled
psf_validation whose name did not match put the exposure in the merge and then
killed the whole job when the tar held nothing selectable. Membership is now
the name on both sides, with one line saying an entry was labelled and skipped.

RENAMING `campaign:` WOULD HAVE HALF-UPDATED THE FILE. The tile hdf5 carries
the campaign in its GROUP, so a rename pointed the rule at a new group inside
the same file: a second group beside the first, the first frozen and stale,
and n_tiles describing one of them. One file is one campaign — apply refuses
and names what is already there.

APPEND IS CHEAP IN READS, NOT IN WRITES, and the docstrings said otherwise.
The existing file is copied so the result can be moved into place atomically:
one pass over it and, briefly, twice its size on disk. Corrected, and apply
now refuses when the filesystem cannot hold it rather than filling /project
and leaving a truncated tmp beside a catalogue people trust.

A CORRUPT TAR RAISED A RAW ReadError, on both sides. persist_exp now refuses
to write a new tar and says the old one is untouched and may hold products
nothing else has; merge_star_cat names the tar and says not to delete it.

THE 16-COLUMN SCHEMA IS DEFINED TWICE and nothing held the two together.
MergeStarCatPSFEX writes the flat FITS table the module runner emits;
merge_star_cat.py writes the hdf5. Separate implementations are right — only
one of them reads tars, keeps native dtypes and reconciles — but a column
added to one writer would simply be missing from the other's product, found by
whoever next computed rho statistics from the wrong one.
tests/unit/test_star_cat_columns.py asserts the names and their order agree;
verified passing, and verified failing when one list is changed.

Also noted where it lives: adding a retention product re-packs the tar and
moves its mtime, so the star merge refreshes those exposures although their
validation members are byte-for-byte unchanged — seconds per exposure against
per-member bookkeeping on every exposure, which is not a trade worth making.

Stale docs updated to the hdf5 product: the Snakefile's merges header, the
README's star_cat_merge paragraph and scripts list (the hdf5 paragraph written
last round never landed — its edit script aborted before writing), and
config.yaml's psf_validation block. The README now says there are two writers
and names the test that keeps their schema together.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
PR #847 removed ShapePipe's mask generation and left the sky-fixed UNIONS
masks to be queried per object by make_cat via MASK_EXT_PATHS. The committed
config never set it, so no campaign wrote a mask column and the merged shear
catalogue carried no mask information at all.

It does now. The DR6 Aug-2026 ugriz bit ladder is staged group-readably at
/project/def-mjhudson/unions-wl/masks/dr6-2026-08 (11 boolean healsparse maps,
nside 131072, True = masked), reached as a third input root SP_INPUT_MASKS
beside tiles and exposures. config_tile_Mc.ini names all 11 in MASK_EXT_PATHS
and carries the bit table; final_cat.param carries the matching MASK_n1 ..
MASK_n2048, replacing the IMAFLAGS_ISO note.

Two caveats live with the config, because a cut written without them is wrong:
the August regeneration swapped the content of bits 0 and 1, so n1|n2 is only
safe combined; and n2048 is 1 where there is no Pan-STARRS z2 data, so an OR
over every column masks the whole survey. Nothing cuts on these here.

PSF-star selection is untouched: MASK_PATHS in config_exp_psfex.ini stays
commented out.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Two rules carry state across invocations and are correct only over
SEQUENCES: hdf5_reconcile brings a catalogue into agreement with a
campaign that changes under it, and persist_exp packs a tar whose
existing members are a floor. Neither is a claim the example tests can
finish making, so each gets a hypothesis state machine that walks a
random sequence of campaign edits and asserts the model after every one.

hdf5_reconcile: units added, refreshed and removed, and the column set
flipped, against a real file. After every step the datasets are the
campaign's units with their sources' content, the count and digest
attributes agree, every dataset shares one dtype, and a no-op leaves the
mtime alone. Source mtimes are set explicitly, so a same-size rewrite
inside one filesystem tick cannot masquerade as a refresh.

Compaction is asserted where the module actually claims it — the rebuild
path — and stated as "does not grow with history": a rebuild costs ~1.4 kB
more than a from-scratch build (h5py's group copy writes more metadata
than create_dataset does) and that overhead is constant, which
test_repeated_refresh_does_not_grow_the_file pins directly. Plus a crash
injected at the rename, which must leave the previous file byte-identical.

persist_exp: random keep lists of product names, raw globs and
overlapping mixtures over a store that gains and loses products. Members
are additive across packs, never duplicated, and the manifest agrees with
the tar down to the product labels; a missing psf_validation fails
without writing a manifest, an unknown product name is refused before any
work, a corrupt tar is left exactly as it is, and two different sources
with one member name are still fatal.

Both files were checked against five mutants (never rebuild, never
refresh, non-additive retention, optional psf_validation, tolerated
collision); each is caught.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…878)

The exposures' instrument flag images — bad columns, saturated pixels, bleed
trails — are the one masking input this campaign has that never leaves the pixel
domain: exp_split splits them per CCD, SExtractor reads them as IMAFLAGS_ISO,
and that is the end of it. The survey footprint is built from the CCD corner WCS
in the headers, so it cannot subtract them and would silently include defective
pixels. The lost area is percent-level, but it carries exactly the thin
small-scale geometry an accurate window function needs, and only ShapePipe ever
opens these files.

Two rules, both in exposure.smk.

exp_defect_map rasterizes ONE exposure's flag splits into a boolean healsparse
fragment on the persistent root — nside_sparse 131072 over nside_coverage 128,
bit-packed, True = masked, the same form as every map in the UNIONS ladder, so
it drops into that ladder without a resolution change. It hangs off exp_split
rather than exp_psf, so re-rasterizing at a different fidelity never touches the
four-hour PSF chain, and clean_exposure takes its manifest as an input, so
reclamation cannot overtake the copy. The WCS comes from the image split beside
each flag file (headers only, never pixels): the flag mosaic's HDUs carry no WCS
at all — checked on a real exposure — and headers-<num>.npy is a pickled array
of astropy WCS instances a second consumer should not inherit.

Rasterization samples each flagged CCD pixel across its full extent, corners
included, and is CONSERVATIVE: a healpix pixel is masked if any part of the
flagged region touches it, because the centre-based alternative erases a
one-pixel bad column, which is the geometry #878 exists to keep. `oversample`
rides on params with the measured convergence table behind its default of 3.

defect_map_merge unions the campaign's fragments into
<products_dir>/defect_map_<campaign>.hsp. It RECONCILES like final_cat_merge —
a new exposure is OR-ed in on the spot, an exposure that left the campaign or a
fragment that changed forces a rebuild (a union cannot be un-OR-ed), a no-op
leaves the file untouched — against a sidecar that records which exposures are
already in it, declared as a second output because the two are only meaningful
together. Memory is flat in the exposure count: each fragment is read, reduced
to its pixel ids and dropped, so the job holds one accumulator (the campaign's
footprint) and one 2 MB fragment whether the campaign is 127 exposures or 20k.

The map is a campaign PRODUCT, not an input: nothing here reads it back.

Measured on 2079612p inside the container, 40 CCDs, 4.4% of pixels flagged:
37 s and 0.62 GB peak RSS at oversample 3, a 2.0 MB fragment of 359858 healpix
pixels over 13 coverage pixels (95 s / 1.29 GB / 360840 pixels at oversample 5,
so the default is within 0.3% of it at a third of the cost).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…he ladder

config_tile_Mc.ini does NOT gain a MASK_EXT_PATHS entry, and the header now says
why: every path in that ladder exists before the run starts, and the defect map
exists only after it, so wiring it would name a file the first run of a fresh
campaign cannot have and every tile would fail on a missing map. What the header
gains instead is the recipe — where the map lands, copying or symlinking it into
the inputs.masks root, the one MASK_EXT_PATHS entry and the matching
final_cat.param line — plus the two things to know before doing it: the map is
the producing campaign's own footprint, so a catalogue built from a different
tile list reads False (unknown, not clean) outside it, and the rasterization is
conservative, widening a one-pixel bad column to the 1.61" healpix resolution.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Three defects in the two scripts, found reviewing them.

A CCD WHOSE FLAG SPLIT IS MISSING was silently dropped. ccd_files globbed
flag-*.fits and rasterized whatever came back, so an incompletely materialised
split dir — an age-based scratch purge deleting files one at a time, a store
copied or restored half way, a truncated rsync — produced a fragment covering
half the exposure, written with "status": "complete", and nothing downstream
able to notice. Reproduced on a fixture: remove one flag split of two and the
script exited 0 with n_ccds 1. The guard it had was on the opposite, unreachable
direction (a flag without its image), which split_exp cannot produce. The loop
is now driven by the EXPECTED CCD count, which comes in on params from
config_exp_Sp.ini's own N_HDU, and a short split dir is a hard error.

A FULLY FLAGGED CCD blew the rule's memory budget. A MegaCam exposure routinely
carries a dead or saturated chip: 2048 x 4612 = 9.4M flagged pixels, 85M samples
at oversample 3, held in one shot several GB of coordinate arrays and astropy
temporaries against a 2 GB request — so the job OOMed on all three attempts, the
exposure never got a fragment, and because clean_exposure waits on that manifest
it could never be reclaimed either. The 0.62 GB measured on a 4.1%-flagged
exposure said nothing about that case. The flagged-pixel loop is now batched at
CHUNK source pixels and np.unique-accumulated, so peak RSS is set by the batch
and not by how bad the CCD is. Measured on exactly that case, a fully flagged
chip against 2079612p's real WCS: 0.74 GB and 18.3 s at 500k, 1.13 GB and 18.9 s
at 1M — time is flat in the batch size and memory is linear in it, so 500k is
free. The real exposure is unchanged bit for bit (359858 healpix pixels, 2.0 MB)
at 34 s.

THE SIDECAR WENT STALE ON A NO-OP. main() returned on an empty plan without
touching it, but two of its fields describe the CAMPAIGN and not the map: add
tiles whose exposures all lack fragments and campaign_exposures and
exposures_without_fragment are wrong on disk while the union is untouched and
correct. That is exactly the record the docstring promises is "recorded rather
than merely printed". The record is now built apart from writing the map and
rewritten alone when it differs; write_stable keeps the map's mtime where it is.

Also: the stamp is IMPORTED from hdf5_reconcile rather than re-derived, so the
two halves of a campaign's reconciliation answer "did this source change" the
same way; and the docstring now says plainly, in hdf5_reconcile's own terms,
that the map's CONTENT is a function of the input set and its BYTES are not
(measured: 1,586,880 B rebuilt vs 1,589,760 B appended, identical valid_pixels).

tests/unit/test_defect_map_reconcile.py pins the four reconcile branches, the
sidecar refresh and the nside mismatch — fragments made with make_empty, since
the merge half never opens a FITS image or a WCS.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…or 7 GB

Four defects in the DAG wiring and the sizing.

A RERUN TRIGGER RE-AVALANCHED THE CAMPAIGN. defect_map_inputs() named the
manifest of any exposure that had one, reclaimed or not, but exp_defect_map's
input is exp_split's manifest and that went with the scratch store. So any
trigger on the rule — and params is live, carrying the script hash, the nside
and the oversampling this file advertises as a minute's work to change — would
schedule exp_get_images and exp_split from VOS for every reclaimed exposure: a
four-hour chain each, campaign-wide, on a one-character edit. It now mirrors
star_cat_inputs(): a live exposure through its manifest, a reclaimed one through
its FRAGMENT, which is not a declared output of any rule and is therefore a true
leaf. That also gives prod_exp_fragment() its only caller, so the fragment's
location is spelled once here and once in merge_defect_map.fragment_path()
rather than twice with nothing joining them. Verified on a two-exposure compute
fixture: with 2079612 reclaimed and holding a fragment, requesting the map
builds 4 jobs (one exp chain, the live exposure's) where both-live builds 7.

THE FIRST MERGE OF A LARGE CAMPAIGN ASKED FOR ~52 GB. Before there is a sidecar
the coverage-pixel count was estimated as 13 per exposure capped at the FULL
SKY, and 12 * 128^2 = 196608 is reached at ~15k exposures — so a DR6-scale
campaign estimated the whole sky, 25.8 GB of accumulator and ~52 GB of request
(~104 GB on the retry) for a job this file's own comment measures at ~3 GB.
Exposures overlap almost completely; the prior has to be a FOOTPRINT, so it is
now capped at the ~23k coverage pixels the UNIONS ugriz maps measure over the
same sky. A small campaign is still sized on its own exposures, where the
per-exposure figure is the honest one: the two-exposure fixture asks for 806 MB.

THE MERGE'S RUNTIME REACHED 11.4 H WITH NO RETRIES AND NO PARTIAL STATE, so its
mem_mb/runtime attempt scaling was dead code and one walltime overrun eleven
hours in lost the whole thing. retries: 1 is declared, the formula is corrected
to the ~1 s per exposure actually measured, and it is capped below the walltime
Alliance policy lets a job run without checkpointing — with the comment saying
that a campaign needing longer needs a resumable accumulator, not a bigger
number.

SILENCE WAS THE ANSWER when no exposure can contribute. A campaign whose
exposures were all reclaimed by a workflow predating this rule builds no defect
job at all, and the operator saw nothing rather than "no map is possible here".
It now says so once at parse time, like the missing-index note.

Also: a comment pointed at DEFECT_COV_BYTES, which does not exist.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…n one

The README and config.yaml carried the pre-batching numbers and an inode
estimate a third of the real cost.

The DR6 inode figure was ~20k against a ~1M group quota; it is ~60k, because
each exposure gets a defect/ DIRECTORY, the fragment in it and a manifest under
manifests/ — three inodes, not one. The fragment was also described as sitting
"beside its PSF tar", which is not the layout; the path is spelled out.

The timing is the re-measured 34 s (2079612p, 40 CCDs, 4.1% of pixels flagged,
oversample 3, inside the campaign container), and both files now carry the
worst-case memory bound the batching exists for — a fully flagged chip at
0.74 GB — rather than only the average exposure's 0.62 GB.

The README also records what the reconcile does and does not promise (content,
not bytes) and points at the new unit test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
… run

clean_exposure's exp_defect_map edge was unconditional, which reopened the
avalanche defect_map_inputs() is written to avoid. An exposure whose scratch
store went to the 60-day /scratch purge, or to a `clean: false` run, has NO
tombstone — exp_store_reclaimed()'s docstring names that case — so clean_targets()
still asks for its tombstone, the missing exp_defect_map manifest is scheduled,
and its own input (exp_split's manifest) went with the store: exp_get_images and
exp_split from VOS, four hours per exposure, campaign-wide, on the first run of
this branch.

Reproduced on a fixture campaign (tile 000.000, exposures 2079612 live and
2079613 purged-without-tombstone, tile_vignets present): a dry run of 2079613's
tombstone scheduled 13 jobs including exp_get_images and exp_split; with the
edge made conditional it schedules one, clean_exposure. The live exposure still
gets exp_defect_map ordered before its clean, and a reclaimed exposure that
already has a fragment depends on the fragment, a leaf that builds nothing.

ccd_files' short-split error now says what to do about the store it pins: a
damaged split dir is still a hard error rather than a half-footprint fragment
marked complete, but the operator is told that deleting the scratch exp_split
manifest is what rebuilds the split and unblocks the clean.

defect_map_merge's mem_mb also goes through capped_mem(), as star_cat_merge and
final_cat_merge already do: defect_map_cov_bytes() scales as nside^2 and
defect_map.nside is an advertised knob, so one ladder change turns the request
into a job SLURM never schedules and snakemake never diagnoses. Verified on the
fixture with max_mem_mb=700: "defect_map_merge: sized at 806 MB, capped to
max_mem_mb=700" at parse time, mem_mb=700 on the job.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
A rebuild set pixels into a fresh accumulator fragment by fragment, so every
fragment that touched an unseen coverage pixel made healsparse grow — and copy —
the sparse array. union_coverage() now reads each fragment's COVERAGE TABLE
first (HealSparseCoverage.read, kilobytes, never the sparse array) and seeds
make_empty(cov_pixels=...) with the union, so the final array is allocated once.

MEASURED rather than assumed, on synthetic fragments at the campaign's own
resolution (nside 131072 / coverage 128), 400 fragments over 5131 coverage
pixels, a 656 MB accumulator: accumulation 8.4 s unseeded, 7.1 s seeded, with a
0.7 s coverage pre-pass. The reallocations are ~16% of the accumulation, not the
dominant term a naive reading predicts — healsparse grows the array in blocks —
so the review's tens-of-TB estimate does not hold, but the pre-pass is linear
where the copying is not, and the win grows with the accumulator.

Only the rebuild path seeds: an append starts from the map on disk, whose
coverage is already most of the footprint. A fragment at the wrong coverage
resolution abandons the seed so accumulate()'s named error still stands; the
existing nside-mismatch test pins that. tests/unit/test_defect_map_reconcile.py
passes (7/7, run in the container), and an end-to-end fixture merge reproduces
rebuild == OR of fragments, no-op untouched, append, and removal-rebuild.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
persist_exp.py carries its own copy of the exposure PSF stage's run-dir name,
and it still held the pre-98bc0857 spelling. exp_persist would have searched
run_sp_exp_SxSePsfPi, found nothing, and tarred an empty product set -- a
silent loss rather than a failure, since an exposure with no keepable products
is a legitimate state.

This half of the rename lives here rather than in the develop hotfix because
persist_exp.py does not exist on develop; it arrives with this branch.
tests/unit/test_workflow_run_names.py (in the hotfix) checks this file when it
is present and skips the check when it is not, so the guard travels with
whichever branch has something to guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
(cherry picked from commit 5c1d41fe31537c6d22628670de4be8083ded5beb)
cailmdaley and others added 3 commits September 29, 2026 01:18
Both exposure-derived HealSparse maps now share one resolution block,
config.yaml's exposure_maps: (nside 131072 / nside_coverage 128, the mask
ladder's), with defect: and nexp: sub-blocks.

- exp_footprint (localrule, after exp_persist): the sky corners of every CCD
  with a valid PSF model, read off exp_persist.json's validation_psf members
  and headers-<exp>.npy, to <products_dir>/exp/<shard>/<exp>/manifests/
  exp_footprint.json. Requested by rule all whatever the map switch says;
  clean_exposure waits on it for a live store.
- nexp_map (campaign, off unless exposure_maps.nexp.enabled): stamps every
  footprint record on the products root into
  <products_dir>/nexp_map/nexp_map_<run>.hsp, a uint16 count of exposures
  with a valid PSF per pixel, beside nexp_map_<run>.json. Memory is sized on
  the campaign footprint like defect_map_merge (map_cov_pixels()).
- astra.yaml: decision masking.nexp_map_valid_psf_ccds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BeQNTBcjTqu6TPktxysovQ
nexp_map follows the defect map's rule: built whenever the campaign can
support it (MAPS_NEXP = a fitted PSF), skipped without error for
psf_model: fake. exposure_maps.nexp.enabled defaults to true and stays as
an opt-out, since the map is rebuilt whole. The pin gains nexp_map's job.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BeQNTBcjTqu6TPktxysovQ
nexp_map_targets() mirrors defect_map_targets(): an in-scope exposure with
no footprint record (reclaimed before exp_footprint existed) is counted in a
parse-time warning, since the map undercounts there, and with no record at
all the map is not requested, instead of a job that fails every invocation.
docs: an Exposure-level maps page with both chains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BeQNTBcjTqu6TPktxysovQ
@cailmdaley cailmdaley changed the title Emit the instrument flags as a healsparse defect map (exp_defect_map, defect_map_merge) Exposure-level HealSparse maps: instrument defects and exposure count Sep 29, 2026
cailmdaley and others added 11 commits September 29, 2026 15:38
A run config still carrying the retired blocks parsed silently, so
coverage.enabled: false no longer turned the count map off. The error
names the replacement keys under exposure_maps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#897 renames run_config.apply_machine_defaults to apply_defaults; the
ordering assertion now holds on both sides of that merge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The eleven external-mask columns carry the mask type beside the flag
value: MASK_1_Faint_star_halos ... MASK_2048_z2, with the names
sp_validation uses for the same bits. The labels live in MASK_EXT_PATHS
(config_tile_Mc.ini); make_cat still writes MASK_<label> verbatim, and
final_cat.param lists the same eleven names.

A new check ties each label's flag value to the one in its map's file
name, so a label cannot sit on another bit's map.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UDwoWdGTgqnXMfN6sCHUGt
…elop merged in) into feat/instrument-defect-map

Conflicts: tests/workflow/test_dag.py and workflow/README.md take both
sides (the defect and exposure-count rules beside tile_get_catalogue);
tests/workflow/params_pin.json is regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UDwoWdGTgqnXMfN6sCHUGt
cailmdaley and others added 7 commits October 5, 2026 14:37
config.yaml: keep the nibi inputs.masks entry and take develop's
catalogue comment (no segmentation maps since #933). params_pin.json
regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012yfkBvkzPv33iJxXZ6jAxd
…fect-map

test_dag.py: keep this branch's PSF_RULES and DEFECT_RULES alongside
develop's TILE_SHAPE_STORE_RULES. README: this branch's footprint and
nexp-map wording with develop's MCCD refusal; exposure.smk line from here,
tile.smk line from develop (SExtractor joined to the UNIONS catalogue).
params_pin.json regenerated; versus feat/wire-external-masks only this
branch's four new rules change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012yfkBvkzPv33iJxXZ6jAxd
The default cut (masking.mask_default_cut) becomes the OR of MASK_4_Stars,
MASK_8_Manual, MASK_64_r and MASK_1024_Maximask. The halo columns stay in
the catalogue; the cluster test now checks that default plus halos
reproduces mask_r and that the halos add area beyond the default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012yfkBvkzPv33iJxXZ6jAxd
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012yfkBvkzPv33iJxXZ6jAxd

# Conflicts:
#	universes/committed.yaml
make_cat writes all eleven MASK_<flag value>_<name> columns and cuts on
none; which bits to cut on is decided in sp_validation. Remove the
masking.mask_default_cut decision, its universe entry, the cluster test
that tied the default to mask_r, and the default-cut prose in make_cat,
config_tile_Mc.ini and the workflow README. The mask-ext-ladder-columns
contract now hangs off masking.sky_mask_application.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012yfkBvkzPv33iJxXZ6jAxd
…and nflagged

Rebuilds #887 around a single per-exposure product. exp_maps (after
exp_persist) writes one uint8 HealSparse fragment per exposure at the mask
ladder's resolution: 0 outside the exposure's valid-PSF CCDs, 1 + n inside,
n being the flagged CCD pixels whose centres fall in that sky pixel.
exposure_maps sums a campaign's fragments into nexp_<run>.hsp and
nflagged_<run>.hsp. The fragment rides on exp_persist's custody exactly
(durable_edge, now shared with the star catalogue and clean_exposure).

What real exposures (smk-g14-fz, tile 189.305) showed and this fixes:
- 3.6% of every split is the overscan border, flagged 3. Coverage polygons
  from NAXIS included it (+3.7% area, covering the inter-CCD gaps) and the
  defect map then masked it. Coverage and flags now use DATASEC.
- Inside DATASEC 0.28% of pixels are flagged, mostly one-pixel columns and
  cosmics; any-touch rasterization at 1.6" masks 2.2% of the sky for them.
  The fragment counts flagged pixels instead (0.25% of area), so the
  consumer chooses the threshold; nflagged > 0 is the any-touch mask.
- The old defect map covered CCDs without a PSF model; both maps now share
  the valid-PSF CCD set.

Removed: the defect-map reconcile sidecar and digests, the separate
exp_footprint/nexp_map chain, the exposure_maps config block and its
retired-key check, the N_HDU parse, the data-only gate, the oversampling
grid and batching, and the deletion of #797's CANFAR coverage tools.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012yfkBvkzPv33iJxXZ6jAxd
Takes #886's removal of the mask_default_cut decision.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012yfkBvkzPv33iJxXZ6jAxd
@cailmdaley cailmdaley changed the title Exposure-level HealSparse maps: instrument defects and exposure count Exposure-level HealSparse maps: exposure count and flagged pixels Oct 5, 2026
Base automatically changed from feat/wire-external-masks to develop October 6, 2026 14:39
#886 landed as a squash whose tree equals feat/wire-external-masks@baa6430c,
already merged here; the three conflicts (params_pin.json, universes/committed.yaml,
workflow/README.md) take this branch's side, which extends #886's. The merge
adds exactly #940's ngmix change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Nfmxobnf8yQJoqdk6kUAx

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Emit the instrument flags as a healsparse defect map for the survey footprint

1 participant