Collocation and SPARE-ICE updates - #431
Open
olemke wants to merge 28 commits into
Open
Conversation
Improvements and fixes made by Lorena during her Masters' thesis. Collocator (typhon/collocations/collocator.py): - Extend `collocate_dataset` to handle multiple secondary files per primary: collect all secondaries matched to a primary, concatenate them along the time dimension, renumber the time index, and join their file paths with "::" in the `__file` attribute. The first secondary's file attributes are taken as representative for the combined group. - Return early from `collocate_dataset` when no matches are found, and propagate `skip_file_errors` through `FileSet.match`. - Sort the primary and secondary datasets by time at the start of `collocate` so unsorted input files cannot break xarray's stack/sel. - Switch the common-period computation to `pd.to_datetime` so NaT values introduced by concatenating secondaries no longer break the min/max calculation. - Cast the `interval` DataArray to `timedelta64[ns]` and fix the "tempoerally" typo in a comment. Collocations common (typhon/collocations/common.py): - Use `np.int64` explicitly for the row counters in `_rows_for_secondaries` so results stay consistent across platforms. FileSet (typhon/files/fileset.py): - Treat the `find` search interval as closed (no longer subtract a microsecond from `end`) so end-time files are included consistently with `collocate`'s common-period computation. - Add `skip_file_errors` to `match` and search the secondary fileset over an `max_interval`-extended window while keeping the primary search on the original [start, end]. - Append the original file name (plus a temp suffix) to the decompress target path so decompressed HDF4/CloudSat files keep a recognisable name in the temp directory. CloudSat handler (typhon/files/handlers/cloudsat.py): - Rewrite `get_info` to work for both R04 and R05 granules: parse year and DOY from the filename and combine `UTC_start` with the last `Profile_time` value instead of relying on the `start_time`/`end_time` global attributes. - Promote `scnline` to a coordinate (numbered from 1) after reading. - Update the 2C-ICE documentation link to the current CloudSat DPC page. NetCDF4 handler (typhon/files/handlers/common.py): - Add `read_from_list` helper that reads a list of FileInfo objects and concatenates the resulting datasets along a caller-supplied common dimension. AVHRR GAC handler (typhon/files/handlers/tovs.py): - Pass `file_info` into `_interpolate_packed_pixels` so the "too many NaNs" error message can name the affected file, and let the granule be read as None when the interpolation raises. RetrievalProduct (typhon/retrieval/common.py): - Exclude attributes starting with `_repr_` (e.g. `_repr_html_`) from being deep-copied into the product dict. SPAREICE (typhon/retrieval/spareice/common.py): - Mask negative elevations via a `.loc` assignment to avoid chained-assignment warnings. - Convert the input collocations to an xarray Dataset before pulling lat/lon/time/scnpos, and make the retrieval `index` 1-based to match `scnline`. - Add `no_files_error` to `retrieve_from_collocations` and pass it through to `Collocations.map`. Fix `Timer` instantiation (`Timer().start()` rather than the classmethod form) and simplify the elapsed-time log line. - Rework the SPARE-ICE report plots: enable the top/right axes spines for a framed look, save figures as PDF (instead of PNG) with `bbox_inches='tight'`, drop the per-experiment subdirectory, draw a zero line and set xlim on the bias/mfe plots, and stop hard-coding a y-range for the bias plot. Rewrite `_report_ice_cloud` to build the confusion matrix with seaborn `heatmap` (labels=[1,0], row-normalised with NaN-safe denominators), keeping the previous implementation as `_report_ice_cloud_old`.
Collocation end time is now inclusive instead of exclusive.
The secondary dataset sort guard was checking `isinstance(primary, tuple)` in both branches due to a copy-paste bug, so the secondary dataset was never sorted when the primary was not a tuple. Use `isinstance(secondary, tuple)` for the secondary check. Also drop a stray `print(collocated)` debug statement left inside the `collocate` docstring example.
Clean up several leftover print() debug statements added during development: - collocations/common.py: drop "default read_mode is collapse" print issued whenever Collocations.read() is called with the default read_mode. - files/fileset.py: drop print(file_info) on every read in _call_map_function. - retrieval/spareice/common.py: drop print of ice_cloud feature_importances_ in _report_ice_cloud.
FileSet's decompress target names previously relied on the private `tempfile._get_candidate_names()` generator (an underscore-prefixed stdlib API that may change across Python versions). Use the public `secrets.token_hex(8)` instead, which is stable and produces a random filename suffix of the same shape. Applied at both decompress call sites in fileset.py (_get_info_via_handler and FileSet.read).
The interpolation fallback in AVHRR_GAC_HDF.read caught every exception (including KeyboardInterrupt and unrelated bugs) and silently returned None, masking real failures. Narrow it to `except ValueError` (the only exception raised by _interpolate_packed_pixels) and log the skipped file via a module-level logger.warning.
olemke
force-pushed
the
apply-lorena-collocation-updates
branch
from
August 11, 2026 10:16
b3fa04f to
c5ffad5
Compare
_tree_from_dict referenced the deprecated `n_features_` coefficient key, which was renamed to `n_features_in_` in scikit-learn 1.0. Update the deserialization to match the new attribute name so RetrievalProduct trees reconstruct correctly with modern sklearn.
Replace typhon.utils.to_array with direct numpy handling when rebuilding decision trees from stored coefficients. n_features_in_ and n_outputs_ are passed through unchanged, while n_classes_ is coerced with np.atleast_1d/np.asarray to satisfy sklearn's expectation of an integer array.
Trees trained with scikit-learn <1.3 lack the missing_go_to_left field in their node ndarray dtype. Newer sklearn versions reject this in Tree.__setstate__ with an incompatible-dtype error, making trained SPAREICE products unloadable. Rebuild the nodes array with the expected dtype (missing_go_to_left defaulting to 0, matching pre-1.3 behaviour) before calling __setstate__.
Replace `dataset.dims.keys()` with `dataset.sizes.keys()` in common._xarray_rename_fields and AAPP_HDF._test_coords to avoid the FutureWarning about the upcoming change in `Dataset.dims` return type.
Remove the MultiIndex coordinates before assigning a plain integer coordinate, since a MultiIndex cannot be stored to a file.
netCDF4 returns unmasked variables as MaskedArrays. Passing these directly to xarray promotes int types to float (NaN fill), which turned int64 coordinates (e.g. scnline, scnpos, channel) of SPARE-ICE collocation files into float64 on newer xarray/numpy stacks. Unwrap fully unmasked arrays so the original dtype is kept. Add a regression test verifying int64 coordinate dtypes survive a write/read roundtrip.
Subgroups only inherit dimensions when the parent group declares them first, so older readers that look at locally-defined dimensions (group.dimensions) fail to map variables in the subgroups. Sort groups by nesting depth so deeper subgroups are written first and declare their dimensions locally. Add a regression test.
The helper that samples the elevation map computed cell indices from the wrong origin and flipped the latitude direction. Use lat + 90 (instead of 90 - lat) for south→north ordering at the (-90, -180) origin.
fsspec's LocalFileSystem reports symlinks to directories as type other rather than directory, so the trailing-slash glob pattern in _get_matching_dirs skipped them entirely. Drop the trailing slash and filter with isdir(), which follows symlinks. Add a regression test for MHS files read through a symlinked placeholder directory.
Broken MHS AAPP files that lack the scnline dimension crashed the reader. Return None and log a warning instead.
Also compute the mean of the squared data in each grid cell and return it alongside the mean and the number of points.
Only rename variables, coordinates and dimensions whose names still contain the group prefix, and use rename_dims for dimensions.
The bundled standard.json was trained with scikit-learn <1.0 and failed to deserialize with modern versions. Remap the pre-1.0 module paths (sklearn.preprocessing.data, ...) in _import_class, drop constructor params removed from the estimator signatures, and fall back to the old n_features_ key when n_features_in_ is missing or None. Also parse structured dtypes stored in their dictionary representation, which numpy >=2 emits for aligned dtypes such as decision tree nodes.
The inplace Series.replace silently does nothing under pandas 3.0 Copy-on-Write, leaving log10(0) = -inf in the iwp column. Clearsky pixels (iwp=0) were then never NaN, so the training dataset preparation found an empty clearsky subset and crashed. Assign the replaced Series back to the column, which works across all pandas versions.
olemke
force-pushed
the
apply-lorena-collocation-updates
branch
from
August 14, 2026 12:48
f52e2d1 to
0a259dc
Compare
When a file cannot be read (e.g. a freshly written preprocessed file turns out to be corrupt) and skip_file_errors is True, align now yields a None placeholder for the affected match instead of desynchronizing the secondary file generator and raising AlignError for all subsequent matches. _collocate_matches handles the placeholder by skipping the match, logging the affected file, instead of crashing on None.copy() and killing the collocation process. Also fixes an UnboundLocalError in match() when called without max_interval. The filesystem in question is, of course, Lustre on Levante, which occasionally forgets which files it is supposed to have - or their metadata, or both. Truly a wonder of modern storage engineering. Fixes the "[Einstein] ERROR: I got a problem and terminate!" failures.
Add regression tests for the graceful handling of files that cannot be read during alignment and collocation: - align yields None only for the corrupt match (no AlignError from a desynchronised secondary generator) and logs the failed file - match() no longer raises UnboundLocalError without max_interval - _collocate_matches skips the affected match instead of crashing on None.copy() and killing the collocation process
olemke
force-pushed
the
apply-lorena-collocation-updates
branch
from
August 20, 2026 07:29
a23ed9d to
0a41a76
Compare
When skip_file_errors is True and a file cannot be read, the affected match is now skipped gracefully (instead of crashing) and the paths of all files in that match are collected in collocator.skipped_files via a queue shared between the worker processes. This lets the caller detect that data was silently lost and, e.g., regenerate the input files and retry, instead of assuming the search completed cleanly. Add tests covering both the direct queue population and the end-to-end multiprocessing search reporting skipped_files.
The pip-built netCDF4 library on Ubuntu segfaults at the C level when asked to open a file with an unknown format instead of raising a catchable exception. This crashes the interpreter when align reads a corrupt file with skip_file_errors=True in a worker thread. Validate the NetCDF magic signature before handing the file to netCDF4.Dataset and raise OSError for unrecognised formats, keeping such errors catchable by the skip_file_errors machinery. Add a regression test that reading a non-NetCDF file raises OSError.
FileSet.align reads files from a pool of threads. pip-built netCDF4 wheels are often linked against a non-thread-safe HDF5 library, so concurrent netCDF4.Dataset calls can crash the interpreter even for valid files. The flaky CI segfaults on Ubuntu and macOS were this race, triggered by the new corrupt-file align test. Serialize all netCDF4 access in the NetCDF4 handler with a module-level lock so parallel reads stay safe. Keep the magic-byte validation so corrupt files still raise a catchable OSError for skip_file_errors. Add a test that reads files concurrently from multiple threads.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR consolidates a series of updates to the collocation machinery and its satellite data pipeline, many of which were contributed by Lorena Kowalczyk as part of her Master's thesis on SPARE-ICE. The collocator now combines all matching secondary files into a single dataset per primary file (instead of processing only the first match), tolerates
NaTtime values, and sorts inputs by time before collocating. The PR also bundles fixes to theFileSet/NetCDF4 layer that these workflows depend on: symlinked directories are now discovered and read correctly, integer dtypes survive NetCDF4 round-trips, subgroups are written in an order that older readers can parse, decompression uses collision-free temporary file names, and netCDF4 access is hardened against segfaults. SPARE-ICE retrieval benefits from corrected elevation-map indexing, compatibility with models saved by older scikit-learn versions, and updated validation plots. Along the way, the collocation and SPARE-ICE code was modernized to be compatible with Python 3.14.Changes
typhon/collocations/collocator.py: Combine all overlapping secondary files viaxr.concatbefore collocating, mark the combined files in the__filevariable, return early when no file matches are found, and sort input datasets by time to avoid xarray stacking problemstyphon/collocations/collocator.py: Replacepd.Timestamp(values.min().item())withpd.to_datetime(...).min()so datasets containingNaTare handled, cast interval output totimedelta64[ns], and drop the MultiIndex explicitly before storing collocation outputtyphon/collocations/collocator.py+typhon/files/fileset.py: Whenskip_file_errorsis enabled,align()now yields a placeholder instead of silently dropping a match when a file cannot be read (keeping item counts in sync with the collocator), and the collocator collects the paths of the affected files intoself.skipped_filesvia a queue from the worker processestyphon/files/fileset.py: Switchfind()from a semi-open to a closed interval (endinclusive) to matchcollocate(), allowmatch()to skip file-not-found errors viaskip_file_errors, and glob directories without a trailing slash plus anisdir()filter so symlinked directories are found; use unique (secrets.token_hex) decompression targets to avoid name clashes between processes and replace the privatetempfile._get_candidate_nameswithsecretstyphon/files/handlers/common.py: Write NetCDF4 subgroups before their parent groups so dimensions are declared locally, and unwrap fully-unmaskedMaskedArrays before handing them to xarray to preserve integer dtypes (e.g.scnline,channelcoordinates); skip identity renames when writing groups so variables already named without a group prefix are left untouched; add aread_from_list()helper for concatenating multiple files; useDataset.sizesinstead of the deprecatedDataset.dimstyphon/files/handlers/common.py: Harden netCDF4 access against interpreter segfaults — validate a file's NetCDF magic signature before passing it tonetCDF4.Dataset(unknown formats can crash instead of raising, so errors stay catchable viaskip_file_errors), and serialize all netCDF4 access through a module-level threading lock since HDF5 builds are often not thread-safe andFileSet.alignreads from a thread pooltyphon/files/handlers/cloudsat.py: Derive file start/end times from the filename (year/day-of-year) and theUTC_start/Profile_timefields, makingget_infowork for both R04 and R05 data, and exposescnlineas a 1-based coordinatetyphon/files/handlers/tovs.py: Degrade gracefully when AVHRR packed-pixel interpolation fails (log a warning and returnNoneinstead of aborting), report the affected file in the error message, and skip broken MHS files that lack thescnlinedimensiontyphon/retrieval/common.py: Restore SPARE-ICE models saved with old scikit-learn versions — remap un-prefixed sklearn module names (e.g.sklearn.tree.tree→sklearn.tree._classes), fall back ton_features_whenn_features_in_isNone, backfill themissing_go_to_leftnode field, decode structured dtypes in their dictionary representation (numpy ≥ 2), and filter constructor params against the class signature; exclude_repr_*attributes when serializing modelstyphon/retrieval/spareice/common.py: Correct elevation-grid cell indexing (offset to-90/-180origin), zero theelevationcolumn via.loc, replace-infIWP values without in-place mutation, convert collocations withto_xarray()before retrieval, forward ano_files_errorflag so missing files can abort or be skipped, and rework the validation plots (PDF output, confusion-matrix heatmap with feature-importance panel, framing/spine and axis-limit fixes)typhon/geographical.py: Guard against division by zero ingridded_meanso empty cells return 0 instead ofnan, and additionally return the mean of the squared values (useful for variance estimates)typhon/tests/...: Add tests for symlinked directory discovery/reading, locally-declared subgroup dimensions, int-coordinate dtype round-trips, non-NetCDF file rejection, and parallel netCDF4 reads; update test time ranges for the closed-intervalfind()semanticsBreaking Changes
FileSet.find()now treatsendas inclusive (closed interval) instead of semi-open. Code that relied on the previous exclusive-end behavior will see one additional file at the interval boundary.gridded_mean()now returns three values (mean, mean of squares, count) instead of two.