diff --git a/docs/MIM_REVIEWED_RELEASE.md b/docs/MIM_REVIEWED_RELEASE.md index db445f30..d86bcdeb 100644 --- a/docs/MIM_REVIEWED_RELEASE.md +++ b/docs/MIM_REVIEWED_RELEASE.md @@ -12,10 +12,11 @@ old September 21 release. Its `candidate_only` mode controls candidate generatio not permission to publish or silently replace production mappings. Paired production promotion remains the separate reviewed step below. -The [2026-09-27 invalid-CAS mapping review](reviews/mim-invalid-cas-20260927/PROMOTION.md) -records the latest proposed paired staging change and its remaining gates. +The [2026-09-29 Potato scope record](reviews/mediadive-potato-scope-20260929/PROMOTION.md) +records the latest isolated paired staging change and its remaining gates. Paired-change tests, CI, production integration and fresh KG validation remain -separate required gates. The [earlier native-category record](reviews/mim-native-category-20260927/PROMOTION.md) +separate required gates. The [2026-09-27 invalid-CAS mapping review](reviews/mim-invalid-cas-20260927/PROMOTION.md), +[earlier native-category record](reviews/mim-native-category-20260927/PROMOTION.md) and [2026-09-25 acceptance record](reviews/mim-admission-20260925/PROMOTION.md) are preserved as historical evidence, not reassigned to the new candidate. The first hydration candidate exposed three further legacy scope defects (#1169) @@ -194,7 +195,12 @@ The legacy full consolidator now refuses to run while the release pin is present including with `--dry-run` or `--allow-stale-vendored`: its additive seed and raw companion paths would undo the migration. Its narrowly scoped `--identity-policy-only --output ...` operation remains available; it does not -constitute a MIM refresh. +constitute a MIM refresh. That bounded operation names its actual writer, +`scripts/consolidate_chemical_mappings.py`, and its code SHA-256 in the tool/version +metadata. It preserves the description's meaning and adds one current identity +policy fingerprint. Original reconstruction provenance is retained through the +reviewed baseline hash and prior record, not a mismatched old tool/new version. +Unrelated metadata and retained mapping rows remain unchanged. The earlier input-schema repair added explicit empty `verified_date` cells to two canonical chemical rows. That formatting-only repair did not install the @@ -205,9 +211,9 @@ receipts to hide input or code drift. ## Promotion, transforms, and merge Candidate generation does not publish files or run production transforms. Once -the final delta and consumer behavior are accepted, promote **both** the reviewed -unified candidate and the byte-exact supported MIM table to their canonical -mapping paths in one reviewed change: +the final delta and consumer behavior are accepted, verify and promote the +reviewed unified candidate and the byte-exact supported MIM table **together** at +their canonical mapping paths in one reviewed change: - `mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz` - `mappings/ingredient_mappings.sssom.tsv` @@ -215,7 +221,9 @@ mapping paths in one reviewed change: The supported file must match the committed pin's supported-product hash; the unified file must match the final accepted candidate hash. Do not install the upstream source SSSOM, the withheld product, a floating sibling checkout, or -only one member of the pair. `mode=candidate_only` remains the CLI's publication +an unverified counterpart. If the supported table already has the pinned bytes, +retain it unchanged; paired review does not require a meaningless rewrite or a +new MIM release pin. `mode=candidate_only` remains the CLI's publication boundary: paired installation is an explicit reviewed repository operation, not permission for future automatic refreshes. diff --git a/docs/reviews/mediadive-potato-scope-1236.md b/docs/reviews/mediadive-potato-scope-1236.md new file mode 100644 index 00000000..236de863 --- /dev/null +++ b/docs/reviews/mediadive-potato-scope-1236.md @@ -0,0 +1,60 @@ +# Bare Potato is not an evidenced extract identity (#1236) + +The [official EINECS inventory](https://www.mhsr.sk/uploads/files/Zwx10C5G.pdf) +(PDF page 21, checked 2026-09-29) pairs CAS 93348-51-7 and EC 297-194-4 with +`zemiak, Solanum tuberosum aegrotans, extrakt`. Its described scope is extractives +and physically modified derivatives. Preserve the `aegrotans` qualifier; this +is not a declaration that every potato extract or starting tuber is identical. + +The historical nine MediaDive ingredient-1606 observations are retained in +`tests/resources/mediadive/potato_scope.json`, including their exact original +recipe JSON and immutable historical witness/archive hashes. Five say peeled +and cut, one says fresh/washed/peeled/sliced, and three say only Potato. The six +explicit starting-material descriptions do not establish the extract identity; +the three unqualified observations lack the specificity to choose it. Neither +group proves that every eventual prepared extract is chemically different. +All nine observations, including amounts and optional concentration fields, +remain valuable source data and must survive withholding this grounding. + +The source recipe JSON contains no CAS property. The imported unified lexical +claim and both legacy MicroMediaParam rows are separate evidence layers. The +fixture preserves the complete unified row and identifies the legacy rows; +the producer audit must preserve full original mapping rows/input identities, +not invent a CAS assertion inside the unchanged raw recipe record. MediaDive's +[ingredient page](https://bacmedia.dsmz.de/ingredients/1606) also cautions that +identifiers are largely automatically matched. Its +[recipe 3653](https://bacmedia.dsmz.de/solutions/3653) boils and strains the starting +material, which does not establish ingredient-to-extract identity by itself. + +The finite policy therefore withholds **bare Potato → cas:93348-51-7** through +the existing shared name/target guard. It does not ban the CAS identifier, +reject arbitrary strings containing potato, select a replacement identity, or +approve other extract names. Explicit authority labels and direct registry +queries remain eligible for separately supplied evidence. Potato flour, +starch, and other named preparations remain distinct; none is an inferred +replacement for these records. + +The same guard covers the unified reader, strict/hydrate fallback parsing, +embedded CAS-RN aliases, and nested solution-name lookup. With no independent +exact target, MediaDive retains its existing source-local ingredient identity; +no pure compound, botanical taxon, or flour identity is invented. Rebuilding +from a stale synonym or legacy file must not resurrect the held claim. +Candidate quarantine and producer-bound audit retain imported claims separately +from active graph identity. Audit reasons distinguish +`unsupported_material_form_identity` from `insufficient_material_specificity`. + +The policy change alone does not replace mapping artifacts. The separate +[2026-09-29 staging record](mediadive-potato-scope-20260929/PROMOTION.md) +records the reviewed identity-only unified candidate staged in an isolated +worktree; the pinned supported MIM product and release pin remain byte-identical. +Production integration, full replay and release acceptance are separate gates. +Candidate generation and reviewed paired promotion follow +`docs/MIM_REVIEWED_RELEASE.md`; changing a shared policy can stale consumers and +requires fresh producer/merge admission. The historical nine-record fixture is +not an expected count or release acceptance for a future rebuilt graph. + +The primary inventory was read directly, but the fixture is a small cited +transcription, not a retained complete PDF or CAS Registry response. Failed +Common Chemistry/PubChem retrievals supplied no evidence of registry absence. +See [#1236](https://github.com/Knowledge-Graph-Hub/kg-microbe/issues/1236), linked +to the broader [#286](https://github.com/Knowledge-Graph-Hub/kg-microbe/issues/286). diff --git a/docs/reviews/mediadive-potato-scope-20260929/PROMOTION.md b/docs/reviews/mediadive-potato-scope-20260929/PROMOTION.md new file mode 100644 index 00000000..c16b7bd1 --- /dev/null +++ b/docs/reviews/mediadive-potato-scope-20260929/PROMOTION.md @@ -0,0 +1,103 @@ +# Bare-Potato identity-only candidate: isolated staging, 2026-09-29 + +Status: the independently reviewed candidate is staged in the isolated +`fix/mediadive-potato-scope-1236` worktree only. It has **not** replaced the +production checkout's unified artifact. Full MediaDive replay, final whole-suite +gates, CI, production integration, fresh transforms/merge, and KG release review +remain pending at this checkpoint. This record does not certify a release. + +The [scientific scope review](../mediadive-potato-scope-1236.md) and +[#1236](https://github.com/Knowledge-Graph-Hub/kg-microbe/issues/1236) hold only +bare Potato → `cas:93348-51-7`, preserving source observations and imported +grounding claims separately. No replacement chemical, registry identity or +botanical identity is inferred. The +[metadata correction #1237](https://github.com/Knowledge-Graph-Hub/kg-microbe/issues/1237) +makes the identity-only refresh identify its actual writer and preserve the +semantic description without corrupting YAML scalar values. + +## Exact artifacts and origin + +| Role | SHA-256 | +| --- | --- | +| Original unified baseline | `f1545663016b3871d2176ccbdc76caffdefa472838b716956e8d722aaa2caff9` | +| Accepted isolated unified candidate, 13,374,364 bytes | `09c44642ab13b234e89ce113f7510aa9efb56969e4b539cb7843b43dcb425ba7` | +| Unchanged supported MIM table | `6b52b30e018b369aa322d41dfd4e81fcfae0e895e34d7fe48900abf5835815fb` | +| Unchanged reviewed-release pin | `f082c05656a0910c85176eec7b41deeb77c27967aecb80393fddc262819b6d97` | +| Current identity-only writer, `scripts/consolidate_chemical_mappings.py` | `257fe4d8bf16f92eb78d5e375065e030ea56d7c1174d859fecd2baa80ac18c27` | +| Identity-policy fingerprint | `61786a46effa48de9901f77713b172db16a7d6797f35711d5c15c01b0e0ea926` | +| Complete finite candidate review result | `447aff5399cea0dc06c947d67e25f9354ac1c9b98fbb1add204dfa3756c07232` | +| Retained removed full-row TSV, 321 bytes | `07503b97da62d9ef0d5024e167bc55a79a6bea1a6bcb96202b2be658632391b0` | + +The original reconstruction writer was +`scripts/mim_conservative_refresh.py`, SHA-256 +`58ec62b9b01a28f4f3b6b47ff319a11396617d4a07ada409acc2742bcb77ae6a`. +That historical provenance remains bound to the original baseline and the +[previous review](../mim-invalid-cas-20260927/PROMOTION.md); it is not presented +as the writer of the new identity-only bytes. + +The upstream origin remains immutable MIM commit +`1848b0fe521bc2462f165912fcf92d09ad9a8cec`, with reviewed manifest +`9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af`. +This is not a new upstream MIM export or tag. Both mapping products were checked +as a pair; the already correct supported table and pin were not rewritten. + +## Verified finite delta + +The complete streaming comparison found exactly one removed 13-column lexical +row: `kgm.name:potato` / `Potato` / `skos:closeMatch` / +`cas:93348-51-7`, source `mediadive_compounds`, comment `synonym`, mapping date +`2026-09-04`. Its exact original fields remain in +`tests/resources/mediadive/potato_scope.json` (fixture SHA-256 +`b4755925efe8cde9871569b047e28c185ac56e2cba6295fca20fef2901062723`) +and the removed-row evidence above. All retained mapping row bytes and order +were identical. No new mappings or replacement targets were introduced. + +The candidate has 591,949 rows, 120,183 distinct targets, 336,050 exact matches +and 255,899 close matches. Only three metadata fields changed: +`mapping_tool`, `mapping_tool_version`, and `mapping_set_description`. The +description retains the original conservative-MIM manifest statement and adds +exactly one identity-policy fingerprint. All unrelated metadata bytes/order, +assertion dates, historical unified version and legacy direction contract are +unchanged. Supported MIM remains 1,747 exact mappings covering 1,696 names. + +A second identity-only refresh removed/relabelled zero rows and produced the +same compressed candidate SHA-256, not merely equivalent parsed rows. Fresh +runtime lookups returned no target for `Potato`, `potato`, `(Potato)` and +`Pot.ato`; they preserved `CrKSO42 x 12 H2O` → `cas:7788-99-0`, `KH2PO3` → +`cas:13977-65-6`, `TAPSO` → `cas:68399-81-5`, `TitaniumIII chloride` → +`cas:7705-07-9`, and `Potato flour` → `FOODON:03302378`. Explicit extract and +direct-CAS query eligibility does not invent a mapping absent from these inputs. + +The immutable local review directory is +`data/issue1224-quarantine-20260929.xjuS9H/identity-candidate-1236-fixed.rxU4PT/review-complete-01/`. +It contains the hash-bound `result.json`, before/after input guards, +`removed-claim.tsv` / `removed-claim.json`, and exact `second-cycle.sssom.tsv.gz`. +The earlier interrupted candidate remains diagnostic only and is not this +accepted artifact. These candidate checks do not replace producer execution +receipts or a full baseline/replay comparison. + +## Remaining gates + +The targeted native mapping-pair and legacy predicate-semantics modules passed +**19 tests** against the staged candidate: + +```bash +python -m pytest -q tests/test_sssom_asymmetric_direction.py tests/test_sssom_predicate_semantics.py +``` + +The changed test module also passed Ruff, and `git diff --check` passed. These +are targeted checks, not full-suite or CI completion. The existing invalid-CAS, +native-category, supported-pair and historical-fixture controls remain intact; +the same full stream also checks absence of the exact held Potato row. + +The producer audit +must retain actual complete strict/hydrate/embedded claims, source recipes and +their hashes; the historical nine-record fixture is not an assumed current +cohort. Whole-source replay must conserve all fields and occurrence multiplicity +except the reviewed grounding delta, with explicit node review. Shared policy +changes require current freshness assessment, potentially all 15 canonical +producers, followed by fresh merge admission and KG model/path/release reviews. + +Green candidate checks or future green code CI alone do not complete these +production and scientific gates. No freshness receipt may be restamped to make +older outputs appear current. diff --git a/kg_microbe/transform_utils/mediadive/material_scope_audit.py b/kg_microbe/transform_utils/mediadive/material_scope_audit.py new file mode 100644 index 00000000..0c439b0f --- /dev/null +++ b/kg_microbe/transform_utils/mediadive/material_scope_audit.py @@ -0,0 +1,263 @@ +"""Preserve finite Potato/extract grounding candidates separately from identity (#1236).""" + +import csv +import gzip +import hashlib +import io +import json +import re +from pathlib import Path + +from kg_microbe.merge_utils.source_admission import SourceAdmission +from kg_microbe.transform_utils.constants import ( + CAS_RN_KEY, + CAS_RN_PREFIX, + COMPOUND_ID_KEY, + COMPOUND_KEY, + DATA_KEY, + ID_COLUMN, + RECIPE_KEY, + SOLUTION_KEY, + SOLUTIONS_KEY, + SOURCE_ASSERTION_ID_COLUMN, + SOURCE_RECORD_COLUMN, +) +from kg_microbe.utils.ingredient_identity import ( + IDENTITY_POLICY, + _policy_target_key, + _scope_key, + ingredient_mapping_allowed, +) +from kg_microbe.utils.source_finalization import SourceFinalizationRequired +from kg_microbe.utils.sssom_identity_policy import classify_mapping_row +from kg_microbe.utils.transform_fingerprint import _repo_root +from kg_microbe.utils.tsv_io import tsv_writer + +AUDIT_FILENAME = "mediadive_material_scope_quarantine.tsv" +TARGET = "cas:93348-51-7" +AUTHORITY_LABEL = "zemiak, Solanum tuberosum aegrotans, extrakt" +AUTHORITY_URI = "https://www.mhsr.sk/uploads/files/Zwx10C5G.pdf#page=21" +UNIFIED_ROLE = "material_scope_unified" +POLICY_ROLE = "material_scope_policy" +AUDIT_HEADER = ( + SOURCE_ASSERTION_ID_COLUMN, + SOURCE_RECORD_COLUMN, + "source_record_sha256", + "source_input_path", + "source_input_sha256", + "retained_target", + "disposition", + "reason", + "qualifier_evidence", + "candidate_route", + "candidate_input_path", + "candidate_input_sha256", + "candidate_record_locator", + "candidate_target", + "candidate_record", + "policy_input_path", + "policy_input_sha256", + "authority_label", + "authority_uri", +) + + +def _json(value): + """Encode evidence without changing null, false, zero, missing fields or literal pipes.""" + return json.dumps(value, ensure_ascii=False, allow_nan=False, sort_keys=True, separators=(",", ":")) + + +def _potato(value): + """Match only the reviewed bare source spelling, using the policy's normalization.""" + return isinstance(value, str) and _scope_key(re.sub(r"[^\w\s-]", "", value)) == "potato" + + +def _reason(raw): + """Recognize only reviewed explicit forms; unknown qualifiers do not imply fresh tuber.""" + qualifiers = {key: raw[key] for key in ("condition", "attribute") if key in raw} + explicit = (isinstance(raw.get("condition"), str) and raw["condition"].strip().casefold() == "peeled and cut") or ( + isinstance(raw.get("attribute"), str) + and raw["attribute"].strip().casefold() == "fresh, washed, peeled and sliced" + ) + return ( + "unsupported_material_form_identity" if explicit else "insufficient_material_specificity", + _json(qualifiers), + ) + + +def _table_rows(stream, required): + """Read complete rows with physical line locators and reject ambiguous evidence.""" + skipped = 0 + for first in stream: + if not first.startswith("#"): + break + skipped += 1 + else: + raise SourceFinalizationRequired("Missing material-scope mapping header") + from itertools import chain + + reader = csv.DictReader(chain((first,), stream), delimiter="\t") + fields = reader.fieldnames or () + if len(fields) != len(set(fields)) or not required.issubset(fields): + raise SourceFinalizationRequired("Invalid material-scope mapping header") + for ordinal, row in enumerate(reader, 1): + if None in row or any(value is None for value in row.values()): + raise SourceFinalizationRequired("Malformed material-scope mapping row") + yield f"record={ordinal};line_end={skipped + reader.line_num}", row + + +class MaterialScopeAudit: + """Collect only the reviewed grounding candidates; never create or choose an identity.""" + + def __init__(self, transform, media_list): + """Bind selected current input bytes before graph output, only scanning a relevant cohort.""" + self.transform = transform + self.guard = SourceAdmission() + self.claims = [] + self.rows = {} + solutions = { + str(solution[ID_COLUMN]) + for medium in media_list[DATA_KEY] + for solution in transform.media_detailed[str(medium[ID_COLUMN])].get(SOLUTIONS_KEY, []) + } + relevant = False + for identifier in solutions: + recipe = transform.solutions_data[identifier].get(RECIPE_KEY) + # Match the existing producer's absent/non-list recipe behavior. + if isinstance(recipe, list): + relevant |= any( + _potato(item.get(COMPOUND_KEY, item.get(SOLUTION_KEY))) for item in recipe if isinstance(item, dict) + ) + self.active = relevant + if not relevant: + return + # Bind the exact current policy even if the mapping reader has already + # excluded the row. Such rows are candidates, not attempted lookups. + with self._consume(POLICY_ROLE, IDENTITY_POLICY) as stream: + policies = list(_table_rows(stream, {"target_id", "authority_label", "kind", "value", "reason"})) + if not any( + _policy_target_key(row["target_id"]) == TARGET + and row["authority_label"] == AUTHORITY_LABEL + and row["kind"] == "name_pattern" + and row["value"] == "(?i)^potato$" + for _, row in policies + ) or ingredient_mapping_allowed("Potato", TARGET): + raise SourceFinalizationRequired("Reviewed Potato material-scope identity hold is missing") + unified = _repo_root() / transform.DATA_INPUTS[0] + with self._consume(UNIFIED_ROLE, unified) as snapshot: + with gzip.GzipFile(fileobj=snapshot.buffer) as compressed: + with io.TextIOWrapper(compressed, encoding="utf-8", newline="") as stream: + for locator, row in _table_rows(stream, {"subject_id", "subject_label", "object_id"}): + if ( + row["subject_id"].startswith("kgm.name:") + and _potato(row["subject_label"]) + and _policy_target_key(row["object_id"]) == TARGET + and classify_mapping_row(row)[0] in {"canonical_name", "synonym"} + ): + self._claim(UNIFIED_ROLE, locator, row["object_id"], row) + for role in ("micromediaparam_hydrate", "micromediaparam_strict"): + # The already selected optional-input contract binds absence too. + with transform.consume_optional_input(role) as stream: + if stream is None: + continue + for locator, row in _table_rows(stream, {"original", "mapped"}): + if _potato(row["original"]) and _policy_target_key(row["mapped"]) == TARGET: + self._claim(role, locator, row["mapped"], row) + self.guard.verify() + + def _consume(self, role, path): + """Guard the original lexical locator in addition to the immutable parser snapshot.""" + self.guard.bind_path(path) + self.guard.capture(path) + return self.transform.consume_input(role, path) + + def _claim(self, role, locator, target, record): + """Retain distinct row ordinals even when the complete mapping payload is duplicated.""" + snapshot = self.transform.consumed_input_snapshots[role] + self.claims.append((role, snapshot, locator, target, _json(record))) + + def observe(self, occurrence, raw): + """Attach available candidates to an actual emitted occurrence and its selected target.""" + if not self.active or not _potato(raw.get(COMPOUND_KEY, raw.get(SOLUTION_KEY))): + return + if _policy_target_key(occurrence[ID_COLUMN]) == TARGET: + raise SourceFinalizationRequired("Held Potato extract identity escaped source resolution") + candidates = list(self.claims) + identifier = raw.get(COMPOUND_ID_KEY) + if identifier is not None: + embedded = self.transform.compounds_data.get(str(identifier), {}) + value = embedded.get(CAS_RN_KEY) + if value is not None and _policy_target_key(CAS_RN_PREFIX + str(value)) == TARGET: + candidates.append( + ( + "mediadive_compounds", + self.transform.consumed_input_snapshots["mediadive_compounds"], + f"key={_json(str(identifier))};field={CAS_RN_KEY}", + CAS_RN_PREFIX + str(value), + _json(embedded), + ) + ) + source = self.transform.consumed_input_snapshots["mediadive_solutions"] + policy = self.transform.consumed_input_snapshots[POLICY_ROLE] + reason, qualifiers = _reason(raw) + payload = occurrence[SOURCE_RECORD_COLUMN] + for route, snapshot, locator, target, candidate in candidates: + row = ( + occurrence[SOURCE_ASSERTION_ID_COLUMN], + payload, + hashlib.sha256(payload.encode("utf-8")).hexdigest(), + source["path"], + source["sha256"], + occurrence[ID_COLUMN], + "quarantined_grounding_candidate", + reason, + qualifiers, + route, + snapshot["path"], + snapshot["sha256"], + locator, + target, + candidate, + policy["path"], + policy["sha256"], + AUTHORITY_LABEL, + AUTHORITY_URI, + ) + key = (occurrence[SOURCE_ASSERTION_ID_COLUMN], route, locator) + if self.rows.setdefault(key, row) != row: + raise SourceFinalizationRequired("Material-scope occurrence changed across repeated visits") + + def write(self): + """Write complete producer-time evidence, including a valid empty-cohort header.""" + self.guard.verify() + self.transform.verify_consumed_inputs() + path = self.transform.output_dir / AUDIT_FILENAME + if path.is_symlink(): + raise SourceFinalizationRequired("Material-scope audit output must not be a symlink") + with path.open("w", encoding="utf-8", newline="") as stream: + writer = tsv_writer(stream, quoting=csv.QUOTE_NONE, quotechar=None) + writer.writerow(AUDIT_HEADER) + for key in sorted(self.rows): + writer.writerow(self.rows[key]) + self.guard.verify() + self.transform.record_producer_audit(AUDIT_FILENAME) + + +def verify_recorded_material_inputs(report, report_path): + """Require actual canonical evidence origins for the producer-bound candidate sidecar.""" + snapshots = report.get("consumed_inputs", {}) + path = Path(report_path).parent / AUDIT_FILENAME + with path.open(encoding="utf-8", newline="") as stream: + reader = csv.DictReader(stream, delimiter="\t", quoting=csv.QUOTE_NONE) + if tuple(reader.fieldnames or ()) != AUDIT_HEADER: + raise SourceFinalizationRequired("Invalid material-scope producer audit header") + has_rows = next(reader, None) is not None + if not has_rows and not any(role in snapshots for role in (UNIFIED_ROLE, POLICY_ROLE)): + return + expected = { + UNIFIED_ROLE: _repo_root() / "mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz", + POLICY_ROLE: IDENTITY_POLICY, + } + for role, origin in expected.items(): + if snapshots.get(role, {}).get("path") != str(origin.resolve()): + raise SourceFinalizationRequired(f"Missing or wrong material-scope input origin: {role}") diff --git a/kg_microbe/transform_utils/mediadive/mediadive.py b/kg_microbe/transform_utils/mediadive/mediadive.py index 53c79e81..07061a5a 100644 --- a/kg_microbe/transform_utils/mediadive/mediadive.py +++ b/kg_microbe/transform_utils/mediadive/mediadive.py @@ -135,6 +135,11 @@ verify_bulk_inputs, verify_recorded_bulk_inputs, ) +from kg_microbe.transform_utils.mediadive.material_scope_audit import ( + AUDIT_FILENAME, + MaterialScopeAudit, + verify_recorded_material_inputs, +) from kg_microbe.transform_utils.transform import Transform from kg_microbe.utils.chemical_mapping_utils import ChemicalMappingLoader from kg_microbe.utils.dummy_tqdm import DummyTqdm @@ -158,6 +163,7 @@ class MediaDiveTransform(Transform): #: via constants.py (#1035), plus BacDive's intermediate strain-taxid TSV (#1091). TRANSFORM_INPUTS = ("ontologies", BACDIVE) DEFAULT_INPUT_DIR = RAW_DATA_DIR + REQUIRED_AUDIT_FILES = (AUDIT_FILENAME,) REQUIRED_CONSUMED_INPUTS = ("bacdive_taxon_lookup", *REQUIRED_BULK_INPUTS) DATA_INPUTS = ("mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz",) @@ -242,11 +248,16 @@ def producer_native_inputs(self, recorded): def verify_native_inputs(self, *, byte_verified=False, output_dir=None): """Keep the constructor and selected list bound through finalization.""" verify_bulk_inputs(self, byte_verified=byte_verified, output_dir=output_dir) + material_audit = getattr(self, "_material_scope_audit", None) + if material_audit is not None: + material_audit.guard.verify(metadata_only=byte_verified) @classmethod def verify_recorded_native_inputs(cls, report, report_path, *, admission=None): """Admit the original raw JSON identities with the public source validator.""" - return verify_recorded_bulk_inputs(report, report_path, admission=admission) + guard = verify_recorded_bulk_inputs(report, report_path, admission=admission) + verify_recorded_material_inputs(report, report_path) + return guard def _create_node_row( self, @@ -796,6 +807,9 @@ def get_solution_recipe_occurrences(self, id: str): ), } ) + material_audit = getattr(self, "_material_scope_audit", None) + if material_audit is not None: + material_audit.observe(occurrences[-1], item) return occurrences def _ingredient_identity_allowed(self, name: str, target: str) -> bool: @@ -1083,6 +1097,7 @@ def _run(self, data_file: Union[Optional[Path], Optional[str]] = None, show_stat bacdive_df = pd.read_csv(bacdive_file, sep="\t", usecols=[BACDIVE_ID_COLUMN, NCBITAXON_ID_COLUMN]) self.verify_consumed_inputs() self._preflight_bulk_records(input_json) + self._material_scope_audit = MaterialScopeAudit(self, input_json) # Create dictionary lookup for O(1) access instead of O(n) DataFrame filtering bacdive_strain_to_ncbi = dict(zip(bacdive_df[BACDIVE_ID_COLUMN], bacdive_df[NCBITAXON_ID_COLUMN], strict=True)) @@ -1451,6 +1466,7 @@ def _run(self, data_file: Union[Optional[Path], Optional[str]] = None, show_stat ) drop_duplicates(self.output_edge_file) self.verify_consumed_inputs() + self._material_scope_audit.write() # Print data source and API call statistics print("\n" + "=" * 80) diff --git a/mappings/ingredient_identity_exclusions.tsv b/mappings/ingredient_identity_exclusions.tsv index c81a0a63..cb6b2c38 100644 --- a/mappings/ingredient_identity_exclusions.tsv +++ b/mappings/ingredient_identity_exclusions.tsv @@ -1,4 +1,5 @@ target_id authority_label kind value reason +cas:93348-51-7 zemiak, Solanum tuberosum aegrotans, extrakt name_pattern (?i)^potato$ The reviewed bare MediaDive Potato ingredient does not establish the EINECS EC 297-194-4 extract identity; six saved occurrences specify fresh/peeled/cut material and three are unqualified. Preserve observations and imported mapping claims separately, without guessing a replacement or banning the registry identifier. See docs/reviews/mediadive-potato-scope-1236.md (#1236). CHEBI:44356 N-tris(hydroxymethyl)methyl-2-aminoethanesulfonic acid name_pattern (?i)^tris(?:\(hydroxymethyl\)|hydroxymethyl)methylamine$ The observed Tris source is not the native TES sulfonic-acid derivative; preserve source identity without guessing a replacement from a legacy substring match. Source 331 quantities corroborate Tris molecular mass (#1169). CHEBI:63037 triammonium citrate name_pattern (?i)^(?:\(nh4\)|nh4)[ _-]*citrate$ The observed unqualified ammonium citrate source lacks counterion stoichiometry, external IDs and molecular-amount evidence; do not select the native triammonium salt. Explicit diammonium/triammonium names and formulas remain valid (#1169). CHEBI:62946 ammonium sulfate name_pattern (?i)^(?:\(nh4\)|nh4)2s4$ The observed oxygen-free source formula (NH4)2S4 cannot silently become sulfate (NH4)2SO4 without independent source evidence; preserve the unresolved source ingredient and native sulfate identities (#1169). diff --git a/mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz b/mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz index 0818ee06..f145effc 100644 Binary files a/mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz and b/mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz differ diff --git a/scripts/consolidate_chemical_mappings.py b/scripts/consolidate_chemical_mappings.py index 0a58c9f3..e5a0be03 100644 --- a/scripts/consolidate_chemical_mappings.py +++ b/scripts/consolidate_chemical_mappings.py @@ -77,6 +77,7 @@ from typing import Dict, List, Optional, Set, Tuple import pandas as pd +import yaml from kg_microbe.utils.cas import invalid_cas_identifier from kg_microbe.utils.chemical_mapping_utils import ( @@ -3038,6 +3039,92 @@ def ingredient_policy_fingerprint() -> str: return digest.hexdigest() +def _refresh_identity_metadata(lines: List[str], policy_hash: str) -> List[str]: + """Edit three YAML fields semantically while retaining unrelated header bytes.""" + text = "".join(line[1:].removeprefix(" ") for line in lines) + # These are logical YAML newlines but not physical TSV line boundaries. + # Reject unsupported raw forms rather than slice the wrong metadata rows. + if any(separator in text for separator in ("\x85", "\u2028", "\u2029")): + raise ValueError("Unsupported raw Unicode line separator in SSSOM metadata") + try: + # An alias could bind an unrelated field to a rewritten scalar. Do not + # silently change that field or leave a dangling anchor reference. + events = list(yaml.parse(text)) + if any(isinstance(event, yaml.AliasEvent) for event in events): + raise ValueError("SSSOM metadata aliases are not supported by identity refresh") + document_end = next( + (event.start_mark.line for event in events if isinstance(event, yaml.DocumentEndEvent) and event.explicit), + None, + ) + root = yaml.compose(text) + + def validate(node): + """Reject duplicate/non-string mapping keys before constructing metadata.""" + if isinstance(node, yaml.MappingNode): + seen = set() + for key, value in node.value: + if not isinstance(key, yaml.ScalarNode) or key.tag != "tag:yaml.org,2002:str": + raise ValueError("SSSOM metadata keys must be strings") + if key.value in seen: + raise ValueError(f"Duplicate SSSOM metadata key: {key.value}") + seen.add(key.value) + validate(value) + elif isinstance(node, yaml.SequenceNode): + for value in node.value: + validate(value) + + if root is not None: + if not isinstance(root, yaml.MappingNode) or root.flow_style: + raise ValueError("SSSOM metadata must be a top-level block mapping") + validate(root) + metadata = yaml.safe_load(text) + if metadata is None: + metadata = {} + if not isinstance(metadata, dict): + raise ValueError("SSSOM metadata must be a mapping") + except yaml.YAMLError as exc: + raise ValueError("Malformed SSSOM metadata YAML") from exc + + fields = { + "mapping_tool": "kg-microbe/scripts/consolidate_chemical_mappings.py", + "mapping_tool_version": script_fingerprint(), + "mapping_set_description": "", + } + for key in fields: + if key in metadata and not isinstance(metadata[key], str): + raise ValueError(f"SSSOM metadata {key} must be a string") + description = metadata.get("mapping_set_description", "") + description = re.sub(r"(?: Ingredient identity policy sha256:[0-9a-f]{64}\.)+$", "", description) + fields["mapping_set_description"] = description + f" Ingredient identity policy sha256:{policy_hash}." + + replacements = {} + if root is not None: + for key, value in root.value: + if key.start_mark.column != 0: + raise ValueError("SSSOM metadata fields must begin on separate top-level physical lines") + if key.value not in fields: + continue + end = value.end_mark.line + bool(value.end_mark.column) + replacements[key.start_mark.line] = (end, key.value) + # JSON quoting is valid YAML and preserves quotes, Unicode and exact + # newlines for every original YAML scalar style. No assertion is redated. + rendered = {key: f"# {key}: {json.dumps(value)}\n" for key, value in fields.items()} + result, position = [], 0 + while position < len(lines): + if position == document_end: + result.extend(rendered.values()) + rendered.clear() + if position in replacements: + end, key = replacements[position] + result.append(rendered.pop(key)) + position = end + else: + result.append(lines[position]) + position += 1 + result.extend(rendered.values()) + return result + + def refresh_identity_policy(source: Path, output: Path) -> dict: """ Apply structural CAS admission and reviewed identity exclusions without enrichment. @@ -3056,21 +3143,18 @@ def refresh_identity_policy(source: Path, output: Path) -> dict: candidate = Path(scratch) / output.name output_open = open_deterministic_gzip if output.suffix == ".gz" else lambda p: p.open("w", encoding="utf-8", newline="") with open_source(source, "rt", encoding="utf-8", newline="") as incoming, output_open(candidate) as outgoing: + metadata_lines = [] for line in incoming: if not line.startswith("#"): fields = next(csv.reader([line], delimiter="\t")) break - if line.startswith("# mapping_tool_version:"): - line = f'# mapping_tool_version: "{script_fingerprint()}"\n' - elif line.startswith("# mapping_set_description:"): - line = re.sub(r" Ingredient identity policy sha256:[0-9a-f]{64}\.", "", line) - line = line.rstrip("\r\n").removesuffix('"') + f' Ingredient identity policy sha256:{policy_hash}."\n' - outgoing.write(line) + metadata_lines.append(line) else: raise ValueError("SSSOM input has no column header") required = {"subject_id", "subject_label", "predicate_id", "object_id", "object_label", "comment"} if not required.issubset(fields): raise ValueError(f"Missing SSSOM identity columns: {required - set(fields)}") + outgoing.writelines(_refresh_identity_metadata(metadata_lines, policy_hash)) outgoing.write(line) # This artifact's exporter emits literal, sanitized TSV, not CSV # quoting. Preserve every unrelated serialized row byte-for-byte. diff --git a/tests/resources/mediadive/potato_scope.json b/tests/resources/mediadive/potato_scope.json new file mode 100644 index 00000000..4c509f2b --- /dev/null +++ b/tests/resources/mediadive/potato_scope.json @@ -0,0 +1,60 @@ +{ + "version": 1, + "issue": "https://github.com/Knowledge-Graph-Hub/kg-microbe/issues/1236", + "authority": { + "cas": "cas:93348-51-7", + "ec_identifier": "297-194-4", + "label": "zemiak, Solanum tuberosum aegrotans, extrakt", + "label_language": "sk", + "url": "https://www.mhsr.sk/uploads/files/Zwx10C5G.pdf", + "pdf_page_one_based": 21, + "accessed": "2026-09-29", + "scope_paraphrase": "Extractives and physically modified derivatives from Solanum tuberosum aegrotans; not proof of identity for a bare Potato starting ingredient.", + "retrieval_limit": "Primary inventory text was retrieved; this fixture preserves its exact row label and identifiers, not a complete hash-pinned PDF or a direct CAS Registry payload." + }, + "historical": { + "archive_sha256": "b97b026e11dc1394d61d58c0b926c8a65d665ec8fb8e04df5e640abd2bceac9d", + "witness_file": "data/mim-ship-20260925.QDVfoK/cas-inventory-a75-preparation.HJMtnx/result/cas-incident-edges.jsonl", + "witness_sha256": "d73d7c65f997cd7c8e06bfb65c52bdbf9b9f6a0b9387d5c3092164b494c59384", + "scope": "Nine selected historical raw recipe observations, not current output expectations. Raw records themselves contain no CAS field." + }, + "mapping_claims": { + "unified": { + "sha256": "f1545663016b3871d2176ccbdc76caffdefa472838b716956e8d722aaa2caff9", + "physical_line": 591193, + "row": { + "subject_id": "kgm.name:potato", + "subject_label": "Potato", + "predicate_id": "skos:closeMatch", + "object_id": "cas:93348-51-7", + "object_label": "", + "object_source": "obo:cas.owl", + "mapping_justification": "semapv:LexicalMatching", + "source": "mediadive_compounds", + "mapping_date": "2026-09-04", + "confidence": "", + "comment": "synonym", + "object_formula": "", + "object_category": "biolink:ChemicalEntity" + } + }, + "legacy": { + "strict_sha256": "7e100314f1a62f675154aef7537be5b72ed2db9fbb8e8b1785198103aff31c71", + "hydrate_sha256": "d2bc349ae31990cf54ef308a14826c756be0e0ee874d35a0a6068d75ef04b3a3", + "physical_line_each": 13770, + "selected_fields_only": {"medium_id": "jcm_333_composition", "original": "Potato", "mapped": "CAS-RN:93348-51-7"}, + "scope": "Selected identifying fields, not a replacement for preserving complete original legacy rows in a producer audit." + } + }, + "occurrences": [ + {"source_assertion_id": "mediadive.solution:227#recipe/2", "raw": {"amount": 200, "compound": "Potato", "compound_id": 1606, "g_l": 200, "optional": 0, "recipe_order": 2, "unit": "g"}, "review_disposition": "insufficient_material_specificity"}, + {"source_assertion_id": "mediadive.solution:3653#recipe/1", "raw": {"amount": 200, "compound": "Potato", "compound_id": 1606, "condition": " peeled and cut", "g_l": 200, "optional": 0, "recipe_order": 1, "unit": "g"}, "review_disposition": "unsupported_material_form_identity"}, + {"source_assertion_id": "mediadive.solution:3678#recipe/1", "raw": {"amount": 300, "compound": "Potato", "compound_id": 1606, "condition": " peeled and cut", "g_l": 300, "optional": 0, "recipe_order": 1, "unit": "g"}, "review_disposition": "unsupported_material_form_identity"}, + {"source_assertion_id": "mediadive.solution:3679#recipe/1", "raw": {"amount": 30, "compound": "Potato", "compound_id": 1606, "condition": " peeled and cut", "g_l": 30, "optional": 0, "recipe_order": 1, "unit": "g"}, "review_disposition": "unsupported_material_form_identity"}, + {"source_assertion_id": "mediadive.solution:3754#recipe/1", "raw": {"amount": 200, "compound": "Potato", "compound_id": 1606, "condition": " peeled and cut", "g_l": 200, "optional": 0, "recipe_order": 1, "unit": "g"}, "review_disposition": "unsupported_material_form_identity"}, + {"source_assertion_id": "mediadive.solution:4051#recipe/1", "raw": {"amount": 200, "compound": "Potato", "compound_id": 1606, "g_l": 200, "optional": 0, "recipe_order": 1, "unit": "g"}, "review_disposition": "insufficient_material_specificity"}, + {"source_assertion_id": "mediadive.solution:4879#recipe/1", "raw": {"amount": 200, "compound": "Potato", "compound_id": 1606, "condition": " peeled and cut", "g_l": 200, "optional": 0, "recipe_order": 1, "unit": "g"}, "review_disposition": "unsupported_material_form_identity"}, + {"source_assertion_id": "mediadive.solution:5502#recipe/1", "raw": {"amount": 150, "compound": "Potato", "compound_id": 1606, "g_l": 150, "optional": 0, "recipe_order": 1, "unit": "g"}, "review_disposition": "insufficient_material_specificity"}, + {"source_assertion_id": "mediadive.solution:840#recipe/1", "raw": {"amount": 200, "attribute": "fresh, washed, peeled and sliced", "compound": "Potato", "compound_id": 1606, "condition": "", "g_l": 200, "optional": 0, "recipe_order": 1, "unit": "g"}, "review_disposition": "unsupported_material_form_identity"} + ] +} diff --git a/tests/test_identity_refresh_metadata.py b/tests/test_identity_refresh_metadata.py new file mode 100644 index 00000000..68ca7cce --- /dev/null +++ b/tests/test_identity_refresh_metadata.py @@ -0,0 +1,186 @@ +"""Identity-only refresh preserves YAML meaning and truthful writer provenance (#1237).""" + +import gzip +import hashlib +from pathlib import Path + +import pytest +import yaml + +from scripts import consolidate_chemical_mappings as refresh +from tests.test_mim_conservative_refresh import FIELDS, _metadata, _row + +TOOL = "kg-microbe/scripts/consolidate_chemical_mappings.py" +PRIOR_PAIR = "mapping_tool: prior/writer.py\nmapping_tool_version: 'sha256:old'\n" +TAIL = "unrelated: # preserve this comment\n nested: ['two', 'one']\n literal: 'colon: # quoted'\n" + + +def _source(tmp_path, fragment, newline="\n"): + """Write one valid, immutable test assertion and literal metadata controls.""" + path = tmp_path / "baseline.tsv" + prefix = yaml.safe_dump(_metadata(), sort_keys=False) + header = "".join("# " + line + newline for line in (prefix + fragment + TAIL).splitlines()) + row = _row("kgm.name:water", "CHEBI:15377", "water", name="water", comment="canonical_name") + body = ("\t".join(FIELDS) + newline + "\t".join(row[field] for field in FIELDS) + newline).encode() + path.write_bytes(header.encode() + body) + return path, body, "".join("# " + line + newline for line in prefix.splitlines()).encode() + + +def _read(path): + """Retain literal header/data bytes and separately parse the header semantics.""" + payload = gzip.decompress(path.read_bytes()) if path.suffix == ".gz" else path.read_bytes() + lines = payload.splitlines(keepends=True) + stop = next(index for index, line in enumerate(lines) if not line.startswith(b"#")) + header = b"".join(lines[:stop]) + metadata = yaml.safe_load("".join(line.decode()[1:].removeprefix(" ") for line in lines[:stop])) + return metadata, header, b"".join(lines[stop:]) + + +@pytest.mark.parametrize( + "description", + [ + "mapping_set_description: plain text\n", + "mapping_set_description: 'single ''quoted'' text'\n", + 'mapping_set_description: "double \\"quoted\\" text"\n', + "mapping_set_description: >\n folded first\n folded second\n", + "mapping_set_description: >-\n folded first\n folded second\n", + "mapping_set_description: |\n literal first\n literal second\n", + "mapping_set_description: |+\n literal first\n\n", + "mapping_set_description: |-\n literal first\n literal second\n", + "mapping_set_description: 'line one\n line two'\n", + 'mapping_set_description: "unicode \\u03b1 and \\u2028 exact"\n', + ], +) +@pytest.mark.parametrize("newline", ["\n", "\r\n"]) +def test_scalar_meaning_tool_pair_unrelated_bytes_and_fixed_point(tmp_path, description, newline): + """Actual SSSOM validation accepts each style without corrupting its scalar value.""" + source, body, prefix = _source(tmp_path, PRIOR_PAIR + description, newline) + original = source.read_bytes() + original_metadata, _, _ = _read(source) + first, second = tmp_path / "first.tsv.gz", tmp_path / "second.tsv.gz" + expected_stats = {"rows_read": 1, "rows_removed": 0, "rows_relabelled": 0} + assert refresh.refresh_identity_policy(source, first) == expected_stats + assert refresh.refresh_identity_policy(first, second) == expected_stats + metadata, header, actual_body = _read(first) + marker = f" Ingredient identity policy sha256:{refresh.ingredient_policy_fingerprint()}." + assert metadata["mapping_set_description"] == original_metadata["mapping_set_description"] + marker + assert metadata["mapping_set_description"].count(marker) == 1 + assert metadata["mapping_tool"] == TOOL + writer_hash = hashlib.sha256(Path(refresh.__file__).read_bytes()).hexdigest() + assert metadata["mapping_tool_version"] == "sha256:" + writer_hash + affected = {"mapping_tool", "mapping_tool_version", "mapping_set_description"} + assert {key: value for key, value in metadata.items() if key not in affected} == { + key: value for key, value in original_metadata.items() if key not in affected + } + assert header.startswith(prefix) + assert header.endswith("".join("# " + line + newline for line in TAIL.splitlines()).encode()) + assert actual_body == body + assert first.read_bytes() == second.read_bytes() + assert source.read_bytes() == original + + +@pytest.mark.parametrize( + "fragment", + [ + "", + "mapping_tool: old.py\n", + "mapping_tool_version: old-version\n", + "mapping_set_description: existing text\n", + PRIOR_PAIR, + ], +) +def test_missing_fields_gain_actual_writer_pair(tmp_path, fragment): + """Absent or partial provenance must never remain a false or incomplete pair.""" + source, body, _ = _source(tmp_path, fragment) + output = tmp_path / "result.tsv.gz" + refresh.refresh_identity_policy(source, output) + metadata, _, actual_body = _read(output) + assert metadata["mapping_tool"] == TOOL + assert metadata["mapping_tool_version"] == refresh.script_fingerprint() + prior_description = (yaml.safe_load(fragment) or {}).get("mapping_set_description", "") + assert metadata["mapping_set_description"] == prior_description + ( + f" Ingredient identity policy sha256:{refresh.ingredient_policy_fingerprint()}." + ) + assert actual_body == body + + +def test_prior_policy_suffixes_replace_without_losing_original_newlines(tmp_path): + """Only previous trailing policy annotations are replaced, never original content.""" + old = " Ingredient identity policy sha256:" + "a" * 64 + "." + original = 'Original "quote"\nretained newline\n' + fragment = yaml.safe_dump({"mapping_set_description": original + old + old}) + source, _, _ = _source(tmp_path, PRIOR_PAIR + fragment) + output = tmp_path / "result.tsv.gz" + refresh.refresh_identity_policy(source, output) + assert _read(output)[0]["mapping_set_description"] == original + ( + f" Ingredient identity policy sha256:{refresh.ingredient_policy_fingerprint()}." + ) + + +@pytest.mark.parametrize( + "fragment,reason", + [ + (PRIOR_PAIR + "mapping_tool: second.py\n", "Duplicate"), + (PRIOR_PAIR + "mapping_tool_version: other\n", "Duplicate"), + ("mapping_set_description: first\nmapping_set_description: second\n", "Duplicate"), + ("extra:\n nested: one\n nested: two\n", "Duplicate"), + ('mapping_set_description: "unterminated\n', "Malformed"), + ("mapping_set_description: null\n", "must be a string"), + ("mapping_set_description: false\n", "must be a string"), + ("mapping_set_description: 123\n", "must be a string"), + ("mapping_set_description: [a, b]\n", "must be a string"), + ("mapping_set_description: {a: b}\n", "must be a string"), + ("mapping_tool: null\n", "must be a string"), + ("mapping_tool_version: 123\n", "must be a string"), + ("mapping_set_description: &shared valid\nextra: *shared\n", "aliases"), + ("extra: !!python/object:unsafe {}\n", "Malformed"), + ("? [complex, key]\n: invalid\n", "keys must be strings"), + ("---\nnew_document: unsupported\n", "Malformed"), + ], +) +def test_invalid_metadata_never_replaces_existing_candidate(tmp_path, fragment, reason): + """Reject ambiguity and bad scalar types before atomic output publication.""" + source, _, _ = _source(tmp_path, fragment) + original = source.read_bytes() + output = tmp_path / "candidate.tsv.gz" + output.write_bytes(b"previous validated candidate") + with pytest.raises(ValueError, match=reason): + refresh.refresh_identity_policy(source, output) + assert source.read_bytes() == original + assert output.read_bytes() == b"previous validated candidate" + + +@pytest.mark.parametrize("terminator", ["...", "... # end comment"]) +def test_explicit_document_terminator_keeps_new_fields_inside_metadata(terminator): + """Absent fields are inserted before a valid YAML document end marker.""" + original = ["# ---\n", "# unrelated: value\n", f"# {terminator}\n"] + output = refresh._refresh_identity_metadata(original, "a" * 64) + parsed = yaml.safe_load("".join(line[2:] for line in output)) + assert parsed["unrelated"] == "value" + assert parsed["mapping_tool"] == TOOL + assert output[0] == original[0] and output[-1] == original[-1] + + +def test_indented_ellipsis_is_unchanged_scalar_content_not_document_end(): + """An unrelated block literal must not receive metadata insertion in its body.""" + original = ["# unrelated: |\n", "# Title\n", "# ...\n", "# end\n"] + output = refresh._refresh_identity_metadata(original, "a" * 64) + parsed = yaml.safe_load("".join(line[2:] for line in output)) + assert parsed["unrelated"] == "Title\n...\nend\n" + assert output[: len(original)] == original + assert parsed["mapping_tool"] == TOOL + + +@pytest.mark.parametrize("separator", ["\x85", "\u2028", "\u2029"]) +def test_raw_unicode_separator_does_not_shift_physical_header_spans(separator): + """Unsupported raw separators fail closed rather than erase unrelated fields.""" + original = [f'# mapping_set_description: "first{separator}second"\n', "# unrelated: preserve\n"] + with pytest.raises(ValueError, match="Unsupported raw Unicode line separator"): + refresh._refresh_identity_metadata(original, "a" * 64) + + +def test_multiple_keys_hidden_on_one_physical_header_line_fail_closed(): + """A semantic span must not overwrite another field on the same physical row.""" + lines = ["# mapping_set_description: first\u2028unrelated: keep\n"] + with pytest.raises(ValueError, match="Unsupported raw Unicode line separator"): + refresh._refresh_identity_metadata(lines, "a" * 64) diff --git a/tests/test_ingredient_identity_contract.py b/tests/test_ingredient_identity_contract.py index 3bc5789c..f8f23596 100644 --- a/tests/test_ingredient_identity_contract.py +++ b/tests/test_ingredient_identity_contract.py @@ -2,12 +2,15 @@ import csv import gzip +import hashlib import io +import json from pathlib import Path from types import SimpleNamespace import pytest +from kg_microbe.transform_utils import constants as c from kg_microbe.transform_utils.mediadive.mediadive import MediaDiveTransform from kg_microbe.transform_utils.metatraits.metatraits import MetaTraitsTransform from kg_microbe.transform_utils.metatraits_gtdb.metatraits_gtdb import MetaTraitsGTDBTransform @@ -15,6 +18,145 @@ from kg_microbe.utils import chemical_mapping_utils as mapping from kg_microbe.utils.ingredient_identity import ingredient_mapping_allowed from tests.test_consolidate_chemical_mappings import _load_module +from tests.test_mim_conservative_refresh import FIELDS, _metadata, _row, _table + + +def _potato_scope(): + """Load a hash-bound historical source excerpt, never current transformed data.""" + path = Path(__file__).parent / "resources/mediadive/potato_scope.json" + payload = path.read_bytes() + assert hashlib.sha256(payload).hexdigest() == "b4755925efe8cde9871569b047e28c185ac56e2cba6295fca20fef2901062723" + return json.loads(payload) + + +@pytest.mark.parametrize("name", ["Potato", "potato", " POTATO ", "Pot.ato", "(Potato)"]) +@pytest.mark.parametrize("target", ["cas:93348-51-7", "CAS:93348-51-7", "CAS-RN:93348-51-7", "cas-rn:93348-51-7"]) +def test_potato_hold_is_target_scoped_and_covers_existing_normalization(name, target): + """Case/punctuation or legacy prefixes cannot reinstate the one reviewed narrowing.""" + assert not ingredient_mapping_allowed(name, target) + assert ingredient_mapping_allowed(name, c.MEDIADIVE_INGREDIENT_PREFIX + "1606") + + +def test_potato_hold_does_not_ban_registry_or_other_material_names(): + """Eligibility is not new identity evidence for an extract, flour, or starch.""" + scope = _potato_scope() + target = scope["authority"]["cas"] + for name in (scope["authority"]["label"], target, "93348-51-7", "Potato flour", "Potato starch", "Potato extract"): + assert ingredient_mapping_allowed(name, target) + for name, supported in ( + ("KH2PO3", "cas:13977-65-6"), + ("TAPSO", "cas:68399-81-5"), + ("TitaniumIII chloride", "cas:7705-07-9"), + ("CrKSO42 x 12 H2O", "cas:7788-99-0"), + ): + assert ingredient_mapping_allowed(name, supported) + assert ingredient_mapping_allowed("Potato flour", "FOODON:03302378") + assert {row["review_disposition"] for row in scope["occurrences"]} == { + "unsupported_material_form_identity", + "insufficient_material_specificity", + } + assert sum(row["review_disposition"] == "unsupported_material_form_identity" for row in scope["occurrences"]) == 6 + assert len(scope["occurrences"]) == 9 + assert all(c.CAS_RN_KEY not in row["raw"] for row in scope["occurrences"]) + + +@pytest.mark.parametrize("stale_object_label", ["", "Potato"]) +def test_potato_unified_reader_rejects_stale_alias_and_canonical_label(tmp_path, monkeypatch, stale_object_label): + """The historical lexical row cannot confer an extract identity or its node name.""" + scope = _potato_scope() + target, authority = scope["authority"]["cas"], scope["authority"]["label"] + stale = dict(scope["mapping_claims"]["unified"]["row"], object_label=stale_object_label) + rows = [ + stale, + _row("kgm.name:native_extract", target, authority, name=authority, comment="canonical_name"), + _row("kgm.name:registry_literal", target, authority, name=target, comment="synonym"), + _row("MIM:Potato_Flour", "FOODON:03302378", "Potato flour"), + _row("MIM:Tapso", "cas:68399-81-5", "TAPSO"), + ] + path = tmp_path / "potato.tsv" + _table(path, FIELDS, rows, _metadata()) + monkeypatch.setattr(mapping, "_LOADED", False) + monkeypatch.setattr(mapping, "_CACHED_PATH", None) + mapping.load_unified_mappings(path) + for options in ({}, {"synonyms": False}, {"fuzzy_hydrate": True}, {"fuzzy_stereochemistry": True}): + assert mapping.find_chebi_by_name("Potato", **options) is None + assert "Potato" not in mapping.get_synonyms(target) + assert mapping.get_canonical_name(target) == authority + assert mapping.find_chebi_by_name(authority) == target + assert mapping.find_chebi_by_name(target) == target + assert mapping.find_chebi_by_name("Potato flour") == "FOODON:03302378" + assert mapping.find_chebi_by_name("TAPSO") == "cas:68399-81-5" + + +@pytest.mark.parametrize("route", ["unified", "strict", "hydrate", "embedded", "all"]) +@pytest.mark.parametrize("target", ["cas:93348-51-7", "CAS-RN:93348-51-7"]) +def test_potato_mediadive_all_fallbacks_remain_source_local(tmp_path, route, target): + """Both legacy loaders and later embedded evidence cannot bypass the shared hold.""" + transform = MediaDiveTransform.__new__(MediaDiveTransform) + transform.chemical_loader = SimpleNamespace( + find_chebi_by_name=lambda *_: target if route in {"unified", "all"} else None + ) + transform.compound_mappings = {} + for selected in ("hydrate", "strict"): + if route not in {selected, "all"}: + continue + path = tmp_path / f"compound_mappings_{selected}.tsv" + path.write_text(f"original\tmapped\nPotato\t{target}\n", encoding="utf-8") + loaded = transform._load_mapping_file(path, selected) + assert loaded == {} + transform.compound_mappings.update(loaded) + # Simulate a stale caller cache too; read-time validation alone is insufficient. + if route == "all": + transform.compound_mappings["potato"] = target + transform.compounds_data = ( + {"1606": {c.COMPOUND_KEY: "Potato", c.CAS_RN_KEY: "93348-51-7"}} if route in {"embedded", "all"} else {} + ) + transform.using_bulk_data = True + transform.api_calls_avoided = 0 + assert transform.standardize_compound_id("1606", "Potato") == c.MEDIADIVE_INGREDIENT_PREFIX + "1606" + if route == "embedded": + assert transform.standardize_compound_id("1606") == c.MEDIADIVE_INGREDIENT_PREFIX + "1606" + + +@pytest.mark.parametrize("route", ["unified", "legacy", "all"]) +def test_potato_nested_solution_name_does_not_reinstate_registry(route): + """A synthetic nested-solution probe exercises the adjacent non-compound branch.""" + transform = MediaDiveTransform.__new__(MediaDiveTransform) + target = _potato_scope()["authority"]["cas"] + transform.using_bulk_data = True + transform.api_calls_avoided = 0 + transform.translation_table = {} + raw = {"solution_id": 2, "solution": "Potato", "amount": 5, "unit": "ml"} + transform.solutions_data = {"1": {"recipe": [raw]}} + transform.chemical_loader = SimpleNamespace( + find_chebi_by_name=lambda *_: target if route in {"unified", "all"} else None + ) + transform.compound_mappings = {"potato": target} if route in {"legacy", "all"} else {} + occurrence = transform.get_solution_recipe_occurrences("1")[0] + assert occurrence[c.ID_COLUMN] == c.MEDIADIVE_SOLUTION_PREFIX + "2" + assert json.loads(occurrence[c.SOURCE_RECORD_COLUMN]) == raw + + +@pytest.mark.parametrize("filename", ["compound_mappings_strict.tsv", "compound_mappings_strict_hydrate.tsv"]) +def test_potato_legacy_consolidator_cannot_regenerate_held_synonym(tmp_path, monkeypatch, filename): + """Tiny in-memory regeneration tests no production CLI, pin, or export replacement.""" + module = _load_module() + monkeypatch.setattr(module, "_build_mangle_blacklist", lambda *_: set()) + scope = _potato_scope() + path = tmp_path / filename + _table( + path, + ("original", "mapped", "chebi_label"), + [{"original": "Potato", "mapped": "CAS-RN:93348-51-7", "chebi_label": scope["authority"]["label"]}], + ) + original = path.read_bytes() + consolidator = module.ChemicalMappingConsolidator() + consolidator.load_compound_mappings(path) + entity = consolidator.chemicals[scope["authority"]["cas"]] + assert entity["canonical_name"] == scope["authority"]["label"] + assert "Potato" not in entity["synonyms"] + assert "potato" not in consolidator.name_index + assert path.read_bytes() == original @pytest.fixture @@ -173,7 +315,7 @@ def test_bounded_refresh_preserves_nonidentity_rows_and_is_a_fixed_point(identit module.refresh_identity_policy(first, second) assert first.read_bytes() == second.read_bytes() with gzip.open(first, "rt") as handle: - rows = list(csv.DictReader(handle, delimiter="\t")) + rows = list(csv.DictReader((line for line in handle if not line.startswith("#")), delimiter="\t")) assert rows[-2]["predicate_id"] == "skos:broadMatch" assert rows[-1]["comment"] == "recipe_equivalent_hydrate" assert {row["object_id"] for row in rows} >= {"CHEBI:78018", "FOODON:03302071"} diff --git a/tests/test_mediadive_material_scope_audit.py b/tests/test_mediadive_material_scope_audit.py new file mode 100644 index 00000000..c5c92f62 --- /dev/null +++ b/tests/test_mediadive_material_scope_audit.py @@ -0,0 +1,336 @@ +"""Preserve finite imported potato claims without inventing raw assertions (#1236).""" + +import copy +import csv +import gzip +import hashlib +import json +from collections import Counter +from pathlib import Path +from types import SimpleNamespace + +import pytest + +from kg_microbe.transform_utils.mediadive import material_scope_audit as audit +from kg_microbe.transform_utils.mediadive import mediadive as mod +from kg_microbe.utils.producer_audits import verify_producer_audits +from kg_microbe.utils.source_finalization import SourceFinalizationRequired + +RESOURCE = Path(__file__).parent / "resources" +FIXTURE = RESOURCE / "mediadive/potato_scope.json" + + +def _recipes(): + """Reconstruct exact raw positions without inventing a missing historical ingredient.""" + saved = json.loads(FIXTURE.read_text()) + recipes = {} + for occurrence in saved["occurrences"]: + identifier, position = occurrence["source_assertion_id"].split(":")[1].split("#recipe/") + recipes[identifier] = { + "recipe": [{"instruction": "synthetic noningredient position placeholder"}] * (int(position) - 1) + + [occurrence["raw"]] + } + return recipes + + +def _write_mappings(root, *, unified=1, legacy=True): + """Create tiny genuine parser inputs, including duplicate rows and extension columns.""" + mappings = root / "mappings" + mappings.mkdir() + row = json.loads(FIXTURE.read_text())["mapping_claims"]["unified"]["row"] + path = mappings / "kgmicrobe_unified_entity_mappings.sssom.tsv.gz" + with gzip.open(path, "wt", encoding="utf-8", newline="") as stream: + stream.write("# synthetic test snapshot of the saved original row\n") + writer = csv.DictWriter(stream, fieldnames=list(row), delimiter="\t") + writer.writeheader() + writer.writerows([row] * unified) + if legacy: + for filename in (mod.MICROMEDIAPARAM_COMPOUND_MAPPINGS_FILE, mod.MICROMEDIAPARAM_HYDRATE_MAPPINGS_FILE): + (root / "raw" / filename).write_text( + 'original\tmapped\tsource\textra\nPotato\tCAS-RN:93348-51-7\ttest\t"literal|pipe"\n' + ) + return path + + +def _producer(tmp_path, monkeypatch, *, recipes=None, unified=1, legacy=True, embedded=None, selected=None): + """Run the actual bulk readers and producer; replace only unrelated native/lookup services.""" + raw = tmp_path / "raw" + raw.mkdir() + recipes = _recipes() if recipes is None else recipes + mapping_path = _write_mappings(tmp_path, unified=unified, legacy=legacy) + monkeypatch.setattr(audit, "_repo_root", lambda: tmp_path) + medium = {"id": 1, "name": "Scope fixture", "complex_medium": False} + for field in ( + mod.MEDIADIVE_SOURCE_COLUMN, + mod.MEDIADIVE_LINK_COLUMN, + mod.MEDIADIVE_MIN_PH_COLUMN, + mod.MEDIADIVE_MAX_PH_COLUMN, + mod.MEDIADIVE_REF_COLUMN, + mod.MEDIADIVE_DESC_COLUMN, + ): + medium[field] = "" + (raw / "mediadive.json").write_text(json.dumps({"data": [medium]})) + bulk = raw / "mediadive" + bulk.mkdir() + for filename, payload in ( + ( + "media_detailed.json", + {"1": {"solutions": [{"id": int(key), "name": "Fixture"} for key in recipes]}}, + ), + ("media_strains.json", {"1": {}}), + ("solutions.json", recipes), + ("compounds.json", embedded or {}), + ): + (bulk / filename).write_text(json.dumps(payload)) + monkeypatch.setattr(mod, "BACDIVE_TMP_DIR", RESOURCE / "provenance_serialization") + monkeypatch.setattr(mod, "MEDIADIVE_TMP_DIR", tmp_path) + monkeypatch.setattr(mod.MediaDiveTransform, "_load_chebi_roles", lambda self: None) + monkeypatch.setattr(mod.MediaDiveTransform, "_load_chebi_categories", lambda self: None) + loader = SimpleNamespace( + find_chebi_by_name=lambda name: selected, + get_canonical_name=lambda identifier: "", + get_node_enrichment=lambda identifier: {"xref": "", "synonym": ""}, + get_parents=lambda identifier: [], + get_category=lambda identifier: "", + ) + monkeypatch.setattr(mod, "ChemicalMappingLoader", lambda: loader) + producer = mod.MediaDiveTransform(raw, tmp_path / "transformed") + return producer, mapping_path + + +def _rows(producer): + """Read the literal producer-owned audit schema.""" + with (producer.output_dir / audit.AUDIT_FILENAME).open() as stream: + reader = csv.DictReader(stream, delimiter="\t", quoting=csv.QUOTE_NONE) + assert tuple(reader.fieldnames) == audit.AUDIT_HEADER + return list(reader) + + +def test_nine_original_occurrences_and_full_mapping_candidates_survive(tmp_path, monkeypatch): + """Six explicit forms and three unqualified records stay distinct from imported CAS claims.""" + producer, mapping = _producer(tmp_path, monkeypatch, unified=2) + producer.run(show_status=False) + rows = _rows(producer) + assert len(rows) == 36 # two distinct unified rows + strict + hydrate, per occurrence + assert Counter(row["reason"] for row in rows) == { + "unsupported_material_form_identity": 24, + "insufficient_material_specificity": 12, + } + expected = {row["source_assertion_id"]: row["raw"] for row in json.loads(FIXTURE.read_text())["occurrences"]} + for row in rows: + raw = json.loads(row["source_record"]) + assert raw == expected[row["source_assertion_id"]] + assert "CAS-RN" not in raw + assert row["retained_target"] == "mediadive.ingredient:1606" + assert row["source_record_sha256"] == hashlib.sha256(row["source_record"].encode()).hexdigest() + assert row["disposition"] == "quarantined_grounding_candidate" + candidate = json.loads(row["candidate_record"]) + if row["candidate_route"] == audit.UNIFIED_ROLE: + assert candidate == json.loads(FIXTURE.read_text())["mapping_claims"]["unified"]["row"] + assert row["candidate_input_sha256"] == hashlib.sha256(mapping.read_bytes()).hexdigest() + else: + assert candidate["extra"] == "literal|pipe" + before = (producer.output_dir / audit.AUDIT_FILENAME).read_bytes() + producer.get_solution_recipe_occurrences("3653") + producer._material_scope_audit.write() + assert (producer.output_dir / audit.AUDIT_FILENAME).read_bytes() == before + verify_producer_audits(producer) + with producer.output_edge_file.open() as stream: + edges = [row for row in csv.DictReader(stream, delimiter="\t") if row.get("source_assertion_id")] + assert len(edges) == 9 + assert {row["source_assertion_id"]: json.loads(row["source_record"]) for row in edges} == expected + + +@pytest.mark.parametrize("kind", ["no_cohort", "no_candidate"]) +def test_empty_or_removed_candidate_is_valid_header_only(tmp_path, monkeypatch, kind): + """Absence of a historical held row must not make valid local observations fail.""" + kwargs = ( + {"recipes": {"1": {"recipe": [{"compound": "Unrelated", "compound_id": 7}]}}} if kind == "no_cohort" else {} + ) + producer, _ = _producer(tmp_path, monkeypatch, unified=0, legacy=False, **kwargs) + producer.run(show_status=False) + assert _rows(producer) == [] + verify_producer_audits(producer) + + +@pytest.mark.parametrize("recipe", [None, "", {}, 0]) +def test_metadata_only_solution_recipe_keeps_existing_empty_behavior(tmp_path, monkeypatch, recipe): + """The audit does not turn a non-list recipe into a new producer schema requirement.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": recipe}}) + producer.run(show_status=False) + assert _rows(producer) == [] + + +@pytest.mark.parametrize("predicate", ["skos:broadMatch", "skos:narrowMatch", "rdfs:seeAlso"]) +def test_weak_or_annotation_rows_are_not_grounding_candidates(tmp_path, monkeypatch, predicate): + """Preserve nonidentity semantics instead of relabeling every Potato relation as identity.""" + producer, path = _producer(tmp_path, monkeypatch, legacy=False) + row = json.loads(FIXTURE.read_text())["mapping_claims"]["unified"]["row"] + row["predicate_id"] = predicate + with gzip.open(path, "wt", newline="") as stream: + writer = csv.DictWriter(stream, fieldnames=list(row), delimiter="\t") + writer.writeheader() + writer.writerow(row) + producer.run(show_status=False) + assert _rows(producer) == [] + + +def test_actual_embedded_claim_is_preserved_without_inserting_it_in_recipe(tmp_path, monkeypatch): + """An embedded compound record is its own claim origin, not a recipe field.""" + embedded = {"compound": "Potato", "CAS-RN": "93348-51-7", "note": None, "flag": False} + producer, _ = _producer(tmp_path, monkeypatch, unified=0, legacy=False, embedded={"1606": embedded}) + producer.run(show_status=False) + rows = _rows(producer) + assert len(rows) == 9 + assert {row["candidate_route"] for row in rows} == {"mediadive_compounds"} + assert all(json.loads(row["candidate_record"]) == embedded for row in rows) + assert all("CAS-RN" not in json.loads(row["source_record"]) for row in rows) + + +def test_retained_target_is_actual_independent_selection_not_assumed_local(tmp_path, monkeypatch): + """The audit does not choose or invent a replacement for a separately selected target.""" + producer, _ = _producer(tmp_path, monkeypatch, selected="mediadive.ingredient:synthetic-reviewed-control") + producer.run(show_status=False) + assert {row["retained_target"] for row in _rows(producer)} == {"mediadive.ingredient:synthetic-reviewed-control"} + + +@pytest.mark.parametrize("damage", ["missing", "malformed", "duplicate_header"]) +def test_selected_mapping_evidence_fails_before_graph_output(tmp_path, monkeypatch, damage): + """Missing or ambiguous current claim evidence cannot be replaced by historical rows.""" + producer, path = _producer(tmp_path, monkeypatch) + if damage == "missing": + path.unlink() + else: + with gzip.open(path, "wt") as stream: + stream.write( + "subject_id\tsubject_label\tobject_id\n" + if damage == "malformed" + else "subject_id\tsubject_label\tobject_id\tobject_id\n" + ) + stream.write("kgm.name:potato\tPotato\n") + with pytest.raises((SourceFinalizationRequired, FileNotFoundError, ValueError)): + producer.run(show_status=False) + assert not producer.output_edge_file.exists() + assert producer.producer_audit_snapshots == {} + + +@pytest.mark.parametrize("damage", ["change", "delete", "symlink"]) +def test_audit_bytes_are_bound_at_producer_time(tmp_path, monkeypatch, damage): + """Generic mandatory-audit checks must reject mutation rather than restamp it.""" + producer, _ = _producer(tmp_path, monkeypatch) + producer.run(show_status=False) + path = producer.output_dir / audit.AUDIT_FILENAME + original = path.read_bytes() + if damage == "change": + path.write_bytes(original + b"changed\n") + else: + path.unlink() + if damage == "symlink": + replacement = tmp_path / "unowned.tsv" + replacement.write_bytes(original) + path.symlink_to(replacement) + with pytest.raises(SourceFinalizationRequired): + verify_producer_audits(producer) + + +def test_input_drift_or_repeated_occurrence_drift_never_gets_certified(tmp_path, monkeypatch): + """Original inputs and occurrence evidence remain bound for the instance lifetime.""" + producer, path = _producer(tmp_path, monkeypatch) + producer.run(show_status=False) + original = copy.deepcopy(producer.solutions_data["3653"]) + producer.solutions_data["3653"]["recipe"][0]["amount"] = 999 + with pytest.raises(SourceFinalizationRequired, match="repeated visits"): + producer.get_solution_recipe_occurrences("3653") + producer.solutions_data["3653"] = original + path.write_bytes(path.read_bytes() + b"changed") + with pytest.raises((SourceFinalizationRequired, ValueError)): + producer._material_scope_audit.write() + + +def test_unknown_qualifier_does_not_infer_fresh_material(): + """Scope reasons are finite evidence classifications, not a fuzzy phenotype parser.""" + reason, evidence = audit._reason({"condition": "fresh extract, exact scope unreviewed", "attribute": None}) + assert reason == "insufficient_material_specificity" + assert json.loads(evidence)["attribute"] is None + + +@pytest.mark.parametrize("name", ["Pot.ato", "(Potato)", " POTATO "]) +def test_policy_equivalent_punctuation_retains_candidate_audit(tmp_path, monkeypatch, name): + """The collector must not skip labels whose grounding the shared guard rejects.""" + recipes = {"1": {"recipe": [{"compound": name, "compound_id": 1606}]}} + producer, _ = _producer(tmp_path, monkeypatch, recipes=recipes) + producer.run(show_status=False) + rows = _rows(producer) + assert len(rows) == 3 + assert all(json.loads(row["source_record"])["compound"] == name for row in rows) + + +@pytest.mark.parametrize("name", ["Potato flour", "Potato extract", "Sweet potato", "potato starch"]) +def test_nonbare_material_names_do_not_enter_the_finite_hold(name): + """No substring-wide potato policy is introduced by auditing.""" + assert not audit._potato(name) + + +@pytest.mark.parametrize("role", [audit.UNIFIED_ROLE, audit.POLICY_ROLE]) +@pytest.mark.parametrize("damage", ["missing", "wrong_origin"]) +def test_recorded_claims_require_both_original_evidence_origins(tmp_path, monkeypatch, role, damage): + """A public source report cannot substitute a same-content arbitrary claim file.""" + producer, _ = _producer(tmp_path, monkeypatch) + producer.run(show_status=False) + report = {"consumed_inputs": producer.consumed_input_snapshots} + report_path = producer.output_dir / "source_finalization.json" + audit.verify_recorded_material_inputs(report, report_path) + if damage == "missing": + del report["consumed_inputs"][role] + else: + report["consumed_inputs"][role]["path"] = str(tmp_path / "unreviewed-alias.tsv") + with pytest.raises(SourceFinalizationRequired, match="input origin"): + audit.verify_recorded_material_inputs(report, report_path) + + +def test_missing_policy_row_is_not_hidden_by_reader_cache(tmp_path, monkeypatch): + """Actual policy bytes must retain the reviewed finite rule even with an old in-memory cache.""" + producer, _ = _producer(tmp_path, monkeypatch) + policy = tmp_path / "empty-policy.tsv" + policy.write_text("target_id\tauthority_label\tkind\tvalue\treason\n") + monkeypatch.setattr(audit, "IDENTITY_POLICY", policy) + monkeypatch.setattr(audit, "ingredient_mapping_allowed", lambda name, target: False) + with pytest.raises(SourceFinalizationRequired, match="hold is missing"): + producer.run(show_status=False) + assert not producer.output_edge_file.exists() + + +def test_same_byte_unified_symlink_retarget_is_rejected(tmp_path, monkeypatch): + """Original lexical origins remain guarded independently of equal file payloads.""" + producer, path = _producer(tmp_path, monkeypatch) + first = path.with_name("original.gz") + second = path.with_name("substitute.gz") + first.write_bytes(path.read_bytes()) + second.write_bytes(path.read_bytes()) + path.unlink() + path.symlink_to(first) + producer.run(show_status=False) + path.unlink() + path.symlink_to(second) + with pytest.raises(SourceFinalizationRequired, match="locator changed"): + producer._material_scope_audit.write() + + +@pytest.mark.usefixtures("local_source_schema") +def test_finalization_repeat_and_public_admission_preserve_exact_audit(tmp_path, monkeypatch): + """The existing finalization lifecycle carries producer bytes without restamping.""" + from kg_microbe.utils.source_finalization import verify_finalized_source_files + from tests.test_mediadive_recipe_occurrences import _run_transform + + _write_mappings(tmp_path, legacy=False) + monkeypatch.setattr(audit, "_repo_root", lambda: tmp_path) + producer = _run_transform(tmp_path, monkeypatch, _recipes()) + path = producer.output_dir / audit.AUDIT_FILENAME + identity = producer.producer_audit_snapshots[audit.AUDIT_FILENAME] + before = path.read_bytes() + report = producer.finalize(fresh_run=True) + assert report["producer_audit_members"][audit.AUDIT_FILENAME] == identity + assert report["audit_members"][audit.AUDIT_FILENAME] == identity + assert path.read_bytes() == before + assert producer.finalize() == report + verify_finalized_source_files([producer.output_node_file, producer.output_edge_file]) diff --git a/tests/test_mediadive_raw_inputs.py b/tests/test_mediadive_raw_inputs.py index 92c14ae8..7907c2fe 100644 --- a/tests/test_mediadive_raw_inputs.py +++ b/tests/test_mediadive_raw_inputs.py @@ -16,7 +16,13 @@ from kg_microbe.utils.optional_consumed_inputs import optional_input_paths, verify_recorded_optional_inputs from kg_microbe.utils.source_finalization import SourceFinalizationRequired from tests.test_mediadive_bulk_inputs import write_bulk_inputs -from tests.test_merge_source_freshness import FIXTURES, merge_config, prepare_source, record_source +from tests.test_merge_source_freshness import ( + FIXTURES, + merge_config, + prepare_source, + record_source, + write_empty_mediadive_audit, +) ROOT = Path(__file__).resolve().parents[1] RESOURCES = Path(__file__).parent / "resources" @@ -199,6 +205,7 @@ def admitted(tmp_path, monkeypatch, local_source_schema, build, request): read_lookup(value) for kind in ("nodes", "edges"): shutil.copyfile(FIXTURES / f"{kind}.tsv", value.output_dir / f"{kind}.tsv") + write_empty_mediadive_audit(value) value.finalize(fresh_run=True) record_source(value) config = merge_config(tmp_path, [value]) diff --git a/tests/test_merge_source_freshness.py b/tests/test_merge_source_freshness.py index 0261a30b..b57557cc 100644 --- a/tests/test_merge_source_freshness.py +++ b/tests/test_merge_source_freshness.py @@ -12,8 +12,10 @@ from kg_microbe.merge_utils import merge_kg from kg_microbe.run import main from kg_microbe.transform import DATA_SOURCES +from kg_microbe.transform_utils.constants import DATA_KEY +from kg_microbe.transform_utils.mediadive.material_scope_audit import AUDIT_FILENAME, AUDIT_HEADER, MaterialScopeAudit from kg_microbe.transform_utils.transform import Transform -from kg_microbe.utils.source_finalization import SourceFinalizationRequired +from kg_microbe.utils.source_finalization import SourceFinalizationRequired, verify_finalized_source_files from kg_microbe.utils.transform_fingerprint import upstream_fingerprint, write_fingerprint ROOT = Path(__file__).resolve().parents[1] @@ -45,6 +47,17 @@ def read_mediadive_bulk_inputs(transform, *, create=False): return tuple(paths) +def write_empty_mediadive_audit(transform): + """Run the actual audit writer only for the consumed immutable zero-cohort fixture.""" + payloads = json.loads((FIXTURES.parent / "mediadive_bulk_inputs.json").read_text()) + with transform.consume_bulk_input("mediadive_media_list", transform.input_base_dir / "mediadive.json") as reader: + media_list = json.load(reader) + assert media_list == payloads["mediadive_media_list"] + assert media_list[DATA_KEY] == [], "Nonempty media require the actual producer's occurrence processing" + transform._material_scope_audit = MaterialScopeAudit(transform, media_list) + transform._material_scope_audit.write() + + def record_source(transform): """Write the actual registered producer's metadata over the isolated fixture graph.""" cls = type(transform) @@ -81,12 +94,40 @@ def prepare_source(tmp_path, source, *, prefix="", marker=True, output_name=None with transform.consume_optional_input(role) as reader: if reader is not None: reader.read() + if source == "mediadive": + write_empty_mediadive_audit(transform) transform.finalize(file_prefix=prefix, fresh_run=True) if marker: record_source(transform) return transform +def test_empty_mediadive_fixture_has_real_producer_audit(tmp_path): + """An empty cohort still retains the writer's exact bytes through public source admission.""" + transform = prepare_source(tmp_path, "mediadive") + path = transform.output_dir / AUDIT_FILENAME + assert path.read_bytes() == ("\t".join(AUDIT_HEADER) + "\n").encode() + identity = transform.producer_audit_snapshots[AUDIT_FILENAME] + report = json.loads((transform.output_dir / "source_finalization.json").read_text()) + assert report["producer_audit_members"][AUDIT_FILENAME] == identity + assert report["audit_members"][AUDIT_FILENAME] == identity + verify_finalized_source_files([transform.output_node_file, transform.output_edge_file]) + + +@pytest.mark.parametrize("damage", ["missing", "unrecorded"]) +def test_empty_mediadive_fixture_cannot_omit_producer_evidence(tmp_path, damage): + """Neither finalized metadata nor a header-only file replaces producer-time evidence.""" + transform = prepare_source(tmp_path, "mediadive") + if damage == "missing": + (transform.output_dir / AUDIT_FILENAME).unlink() + else: + transform._producer_audit_snapshots.clear() + before = {path.name: path.read_bytes() for path in transform.output_dir.iterdir()} + with pytest.raises(SourceFinalizationRequired, match="producer audit"): + transform.finalize(fresh_run=True) + assert {path.name: path.read_bytes() for path in transform.output_dir.iterdir()} == before + + def merge_config(tmp_path, transforms, *, diagnostic=False): """Select explicit fixture paths, using unrelated config labels to prevent path-name inference.""" config = tmp_path / "merge.yaml" diff --git a/tests/test_merge_upstream_review.py b/tests/test_merge_upstream_review.py index 41e09d6d..e69a113a 100644 --- a/tests/test_merge_upstream_review.py +++ b/tests/test_merge_upstream_review.py @@ -13,7 +13,7 @@ from kg_microbe.transform_utils.transform import Transform from kg_microbe.utils.source_finalization import SourceFinalizationRequired from kg_microbe.utils.transform_fingerprint import upstream_fingerprint, write_fingerprint -from tests.test_merge_source_freshness import read_mediadive_bulk_inputs +from tests.test_merge_source_freshness import read_mediadive_bulk_inputs, write_empty_mediadive_audit ROOT = Path(__file__).resolve().parents[1] FIXTURES = Path(__file__).parent / "resources/merge_source_freshness" @@ -55,6 +55,8 @@ def _prepare(tmp_path, name): for role, _ in getattr(cls, "OPTIONAL_RAW_CONSUMED_INPUTS", ()): with transform.consume_optional_input(role) as reader: assert reader is None + if name == "mediadive": + write_empty_mediadive_audit(transform) transform.finalize(fresh_run=True) _record(transform) return transform diff --git a/tests/test_mim_conservative_refresh.py b/tests/test_mim_conservative_refresh.py index cc1b0a75..34697bd3 100644 --- a/tests/test_mim_conservative_refresh.py +++ b/tests/test_mim_conservative_refresh.py @@ -3,7 +3,9 @@ import csv import gzip import io +import json from collections import Counter +from pathlib import Path import pytest import yaml @@ -189,6 +191,33 @@ def _read(path): return list(csv.DictReader((line for line in handle if not line.startswith("#")), delimiter="\t")) +def test_potato_historical_mapping_is_losslessly_quarantined_and_not_reasserted(inputs, monkeypatch): + """The supported-only candidate retains old full rows, not a generic-to-extract identity.""" + fixture = json.loads((Path(__file__).parent / "resources/mediadive/potato_scope.json").read_text()) + held = fixture["mapping_claims"]["unified"]["row"] + target, label = fixture["authority"]["cas"], fixture["authority"]["label"] + positive = _row("kgm.name:explicit_extract", target, label, name=label, comment="canonical_name") + _table(inputs["baseline"], FIELDS, [*refresh._rows(inputs["baseline"]), held, dict(held), positive], _metadata()) + original = inputs["baseline"].read_bytes() + first = refresh.build_conservative_candidate(**inputs) + quarantined = [row for row in _read(first.quarantine_path) if row["subject_id"] == held["subject_id"]] + assert len(quarantined) == 2 + assert all(row.pop("quarantine_reason") == "reviewed_identity_policy_name" for row in quarantined) + assert quarantined == [held, held] + assert positive in _read(first.candidate_path) + assert not any(row["subject_id"] == held["subject_id"] for row in _read(first.candidate_path)) + assert inputs["baseline"].read_bytes() == original + second = refresh.build_conservative_candidate( + **dict(inputs, baseline=first.candidate_path, output_directory=inputs["output_directory"].with_name("second")) + ) + assert first.candidate_path.read_bytes() == second.candidate_path.read_bytes() + monkeypatch.setattr(runtime, "_LOADED", False) + monkeypatch.setattr(runtime, "_CACHED_PATH", None) + runtime.load_unified_mappings(second.candidate_path) + assert runtime.find_chebi_by_name("Potato") is None + assert runtime.find_chebi_by_name(label) == target + + def test_conservative_refresh_rebuilds_connected_claims_and_preserves_independent_relations(inputs): """Reset mixed/propagated claims while retaining separately supported identity.""" result = refresh.build_conservative_candidate(**inputs) diff --git a/tests/test_sssom_asymmetric_direction.py b/tests/test_sssom_asymmetric_direction.py index e507afc5..be93a92e 100644 --- a/tests/test_sssom_asymmetric_direction.py +++ b/tests/test_sssom_asymmetric_direction.py @@ -10,6 +10,7 @@ import gzip import hashlib +import json from collections import Counter from pathlib import Path from unittest import TestCase @@ -25,6 +26,7 @@ UNIFIED = REPO_ROOT / "mappings" / "kgmicrobe_unified_entity_mappings.sssom.tsv.gz" RELEASE_PIN = REPO_ROOT / "mappings" / "mim_reviewed_release.json" PRIOR_CLAIMS = REPO_ROOT / "tests" / "resources" / "cas_mapping_promotion" / "prior_claims.tsv" +POTATO_SCOPE = REPO_ROOT / "tests" / "resources" / "mediadive" / "potato_scope.json" # Byte identities accepted together in the immutable-export promotion review. # Updating a release requires reviewing both products, not regenerating these @@ -32,10 +34,12 @@ SOURCE_COMMIT = "1848b0fe521bc2462f165912fcf92d09ad9a8cec" MANIFEST_SHA256 = "9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af" SUPPORTED_SHA256 = "6b52b30e018b369aa322d41dfd4e81fcfae0e895e34d7fe48900abf5835815fb" -UNIFIED_SHA256 = "f1545663016b3871d2176ccbdc76caffdefa472838b716956e8d722aaa2caff9" -BUILDER_SHA256 = "58ec62b9b01a28f4f3b6b47ff319a11396617d4a07ada409acc2742bcb77ae6a" +UNIFIED_SHA256 = "09c44642ab13b234e89ce113f7510aa9efb56969e4b539cb7843b43dcb425ba7" +IDENTITY_REFRESH_WRITER_SHA256 = "257fe4d8bf16f92eb78d5e375065e030ea56d7c1174d859fecd2baa80ac18c27" +IDENTITY_POLICY_SHA256 = "61786a46effa48de9901f77713b172db16a7d6797f35711d5c15c01b0e0ea926" RELEASE_PIN_SHA256 = "f082c05656a0910c85176eec7b41deeb77c27967aecb80393fddc262819b6d97" PRIOR_CLAIMS_SHA256 = "629198d090f7e45f7f17ce97a124eb9cad879b0bb066930b11ea8cf80ce15b2c" +POTATO_SCOPE_SHA256 = "b4755925efe8cde9871569b047e28c185ac56e2cba6295fca20fef2901062723" def _sha256(path): @@ -141,13 +145,20 @@ def test_prior_claim_fixture_is_fixed_before_candidate_reconstruction(self): ) self.assertTrue(all(row["mapping_date"] == "2026-09-04" for row in rows)) - def test_unified_header_identifies_the_pinned_review_and_builder(self): - """Bind the accepted reconstruction to its manifest, not a floating input.""" + def test_unified_header_identifies_the_pinned_review_and_identity_refresh_writer(self): + """Keep the original reviewed origin and name the actual bounded-update writer.""" metadata = _metadata(UNIFIED) pin = load_release_pin(REPO_ROOT) - self.assertEqual(metadata["mapping_tool"], "kg-microbe/scripts/mim_conservative_refresh.py") - self.assertEqual(metadata["mapping_tool_version"], "sha256:" + BUILDER_SHA256) - self.assertIn("reviewed manifest sha256:" + pin["manifest_sha256"], metadata["mapping_set_description"]) + self.assertEqual(metadata["mapping_tool"], "kg-microbe/scripts/consolidate_chemical_mappings.py") + self.assertEqual(metadata["mapping_tool_version"], "sha256:" + IDENTITY_REFRESH_WRITER_SHA256) + self.assertEqual( + metadata["mapping_set_description"], + "Conservative MIM candidate; reviewed manifest sha256:" + + pin["manifest_sha256"] + + ". Ingredient identity policy sha256:" + + IDENTITY_POLICY_SHA256 + + ".", + ) # The conservative builder preserves historical baseline metadata # (#1170). Absence retains the reader's legacy direction contract; the # accepted unified set has no asymmetric rows to interpret. Do not @@ -163,6 +174,11 @@ def test_unified_counts_and_native_categories_match_the_reviewed_candidate(self) entities = set() predicates = Counter() native_rows = Counter() + self.assertEqual(_sha256(POTATO_SCOPE), POTATO_SCOPE_SHA256) + potato_claim = json.loads(POTATO_SCOPE.read_text(encoding="utf-8"))["mapping_claims"]["unified"]["row"] + self.assertEqual(len(potato_claim), 13) + held_potato_row = _row_key(potato_claim) + retained_potato_claims = 0 prior_rows = list(_iter_sssom_rows(PRIOR_CLAIMS)) prior_invalid = { _row_key(row) @@ -197,6 +213,8 @@ def test_unified_counts_and_native_categories_match_the_reviewed_candidate(self) rows += 1 entities.add(row["object_id"]) predicates[row["predicate_id"]] += 1 + if row["object_id"] == potato_claim["object_id"] and _row_key(row) == held_potato_row: + retained_potato_claims += 1 for key in ("subject_id", "object_id"): if invalid_cas_identifier(row[key]): invalid_endpoints[(key, row[key])] += 1 @@ -211,13 +229,14 @@ def test_unified_counts_and_native_categories_match_the_reviewed_candidate(self) added_provenance[full_row] += 1 if row["object_id"] in {"NCIT:C16883", "NCIT:C71939"}: native_rows[(row["subject_id"], row["predicate_id"], row["object_id"], row["object_category"])] += 1 - self.assertEqual(rows, 591893 - 33 + 90) - self.assertEqual(len(entities), 120184) - self.assertEqual(predicates, {"skos:exactMatch": 336050, "skos:closeMatch": 255900}) + self.assertEqual(rows, 591893 - 33 + 90 - 1) + self.assertEqual(len(entities), 120183) + self.assertEqual(predicates, {"skos:exactMatch": 336050, "skos:closeMatch": 255899}) self.assertEqual(predicates["skos:broadMatch"], 0) self.assertEqual(predicates["skos:narrowMatch"], 0) self.assertEqual(native_rows, expected_native_rows) self.assertFalse(invalid_endpoints) self.assertFalse(retained_invalid) + self.assertEqual(retained_potato_claims, 0) self.assertEqual(retained_lexical, Counter({key: 1 for key in prior_lexical})) self.assertEqual(added_provenance, expected_provenance) diff --git a/tests/test_transform_category_alignment.py b/tests/test_transform_category_alignment.py index fc4c9219..aa1154b3 100644 --- a/tests/test_transform_category_alignment.py +++ b/tests/test_transform_category_alignment.py @@ -1,171 +1,128 @@ -"""Integration tests for transform category alignment.""" +"""Hermetic category alignment checks using real isolated source loaders (#1244).""" -import csv +import builtins +from collections import Counter from pathlib import Path import pytest +import requests -from kg_microbe.transform_utils.bacdive.bacdive import BacDiveTransform -from kg_microbe.transform_utils.constants import CHEBI_NODES_FILE +from kg_microbe.transform_utils.bacdive import bacdive as bacdive_module +from kg_microbe.transform_utils.constants import INGREDIENT_CATEGORY, METABOLITE_CATEGORY from kg_microbe.transform_utils.mediadive import mediadive as mediadive_module -from kg_microbe.transform_utils.mediadive.mediadive import MediaDiveTransform - -@pytest.fixture -def mediadive_categories(monkeypatch): - """Exercise the category loader without unrelated API/ontology orchestration.""" - fixture = Path(__file__).parent / "resources/transform_category_alignment/chebi_nodes.tsv" - monkeypatch.setattr(mediadive_module, "CHEBI_NODES_FILE", fixture) - transform = MediaDiveTransform.__new__(MediaDiveTransform) - transform.chebi_categories = {} - transform._load_chebi_categories() - assert len(transform.chebi_categories) == 3 - return transform - - -class TestTransformCategoryAlignment: - """Test that transforms correctly align categories with ontologies transform.""" - - @pytest.fixture - def chebi_categories_from_ontologies(self): - """Load CHEBI categories from ontologies transform for comparison.""" - categories = {} - chebi_nodes_file = CHEBI_NODES_FILE - - if not chebi_nodes_file.exists(): - pytest.skip("CHEBI nodes file not found - ontologies transform not run") - - with open(chebi_nodes_file) as f: - reader = csv.DictReader(f, delimiter="\t") - for row in reader: - if row["id"].startswith("CHEBI:"): - categories[row["id"]] = row["category"] - - return categories - - def test_bacdive_loads_chebi_categories(self): - """Test that BacDive transform loads CHEBI categories from ontologies.""" - if not CHEBI_NODES_FILE.exists(): - pytest.skip("CHEBI nodes file not found - ontologies transform not run") - bacdive = BacDiveTransform() - - # Check that categories were loaded - assert len(bacdive.chebi_categories) > 0, "BacDive should load CHEBI categories from ontologies transform" - print(f"\n✓ BacDive loaded {len(bacdive.chebi_categories):,} CHEBI categories") - - def test_mediadive_loads_chebi_categories(self): - """Test that MediaDive transform loads CHEBI categories from ontologies.""" - if not CHEBI_NODES_FILE.exists(): - pytest.skip("CHEBI nodes file not found - ontologies transform not run") - mediadive = MediaDiveTransform() - - # Check that categories were loaded - assert len(mediadive.chebi_categories) > 0, "MediaDive should load CHEBI categories from ontologies transform" - print(f"\n✓ MediaDive loaded {len(mediadive.chebi_categories):,} CHEBI categories") - - def test_bacdive_get_chebi_category_specific_examples(self): - """Test that BacDive returns correct categories for specific CHEBI IDs.""" - bacdive = BacDiveTransform() - - if "CHEBI:16828" in bacdive.chebi_categories: - category = bacdive._get_chebi_category("CHEBI:16828") - print(f"\n✓ CHEBI:16828 category: {category}") - assert category == bacdive.chebi_categories["CHEBI:16828"] - - def test_mediadive_get_chebi_category_specific_examples(self, mediadive_categories): - """Test that MediaDive returns correct categories for specific CHEBI IDs.""" - assert mediadive_categories._get_chebi_category("CHEBI:16828") == "biolink:ChemicalEntity" - - def test_bacdive_category_matches_ontologies(self, chebi_categories_from_ontologies): - """Test that BacDive categories match ontologies transform categories.""" - bacdive = BacDiveTransform() - - # Compare samples (not all 224k, just a sample) - sample_ids = list(bacdive.chebi_categories.keys())[:100] - mismatches = [] - - for chebi_id in sample_ids: - bacdive_cat = bacdive.chebi_categories[chebi_id] - ontologies_cat = chebi_categories_from_ontologies.get(chebi_id) - - if ontologies_cat and bacdive_cat != ontologies_cat: - mismatches.append((chebi_id, bacdive_cat, ontologies_cat)) - - assert len(mismatches) == 0, f"Found {len(mismatches)} category mismatches between BacDive and ontologies" - print(f"\n✓ Verified {len(sample_ids)} CHEBI categories match ontologies transform") - - def test_mediadive_category_matches_ontologies(self, chebi_categories_from_ontologies): - """Test that MediaDive categories match ontologies transform categories.""" - mediadive = MediaDiveTransform() - - # Compare samples (not all 224k, just a sample) - sample_ids = list(mediadive.chebi_categories.keys())[:100] - mismatches = [] - - for chebi_id in sample_ids: - mediadive_cat = mediadive.chebi_categories[chebi_id] - ontologies_cat = chebi_categories_from_ontologies.get(chebi_id) - - if ontologies_cat and mediadive_cat != ontologies_cat: - mismatches.append((chebi_id, mediadive_cat, ontologies_cat)) - - assert len(mismatches) == 0, f"Found {len(mismatches)} category mismatches between MediaDive and ontologies" - print(f"\n✓ Verified {len(sample_ids)} CHEBI categories match ontologies transform") - - def test_bacdive_fallback_to_metabolite_category(self): - """Test that BacDive falls back to METABOLITE_CATEGORY for unknown CHEBI IDs.""" - from kg_microbe.transform_utils.constants import METABOLITE_CATEGORY - - bacdive = BacDiveTransform() - - # Test with a CHEBI ID that definitely doesn't exist - fake_chebi = "CHEBI:99999999" - category = bacdive._get_chebi_category(fake_chebi) - - assert category == METABOLITE_CATEGORY - print(f"\n✓ BacDive correctly falls back to {METABOLITE_CATEGORY} for unknown CHEBI IDs") - - def test_mediadive_fallback_to_ingredient_category(self, mediadive_categories): - """Test that MediaDive falls back to INGREDIENT_CATEGORY for unknown CHEBI IDs.""" - from kg_microbe.transform_utils.constants import INGREDIENT_CATEGORY - - mediadive = mediadive_categories - - # Test with a CHEBI ID that definitely doesn't exist - fake_chebi = "CHEBI:99999999" - category = mediadive._get_chebi_category(fake_chebi) - - assert category == INGREDIENT_CATEGORY - print(f"\n✓ MediaDive correctly falls back to {INGREDIENT_CATEGORY} for unknown CHEBI IDs") - - def test_category_distribution(self, chebi_categories_from_ontologies): - """Test distribution of CHEBI categories in ontologies transform.""" - from collections import Counter - - category_counts = Counter(chebi_categories_from_ontologies.values()) - - print("\n✓ CHEBI category distribution in ontologies transform:") - for category, count in category_counts.most_common(10): - percentage = (count / len(chebi_categories_from_ontologies)) * 100 - print(f" {category}: {count:,} ({percentage:.1f}%)") - - assert "biolink:ChemicalEntity" in category_counts - assert category_counts["biolink:ChemicalEntity"] > 100000, "Most CHEBI compounds should be ChemicalEntity" - - def test_mediadive_classify_ingredient_uses_chebi_category(self, mediadive_categories): - """Test that MediaDive _classify_ingredient_category uses CHEBI categories for CHEBI IDs.""" - mediadive = mediadive_categories - - # Test with a known CHEBI ID from ontologies - test_chebi_ids = [cid for cid in list(mediadive.chebi_categories.keys())[:5] if cid.startswith("CHEBI:")] - assert len(test_chebi_ids) == 3 - - for chebi_id in test_chebi_ids: - expected_category = mediadive.chebi_categories[chebi_id] - actual_category = mediadive._classify_ingredient_category(chebi_id, "test compound") - - assert actual_category == expected_category, ( - f"MediaDive _classify_ingredient_category should use CHEBI category for {chebi_id}" - ) - - print("\n✓ MediaDive _classify_ingredient_category correctly uses CHEBI categories") +RESOURCE = Path(__file__).parent / "resources/transform_category_alignment/chebi_nodes.tsv" +EXPECTED = { + "CHEBI:16828": "biolink:ChemicalEntity", + "CHEBI:50906": "biolink:ChemicalRole", + "CHEBI:60004": "biolink:ChemicalEntity", +} +PRODUCERS = { + "bacdive": (bacdive_module, bacdive_module.BacDiveTransform, METABOLITE_CATEGORY), + "mediadive": (mediadive_module, mediadive_module.MediaDiveTransform, INGREDIENT_CATEGORY), +} + + +@pytest.fixture(params=["absent", "populated"]) +def category_loader(tmp_path, monkeypatch, request): + """Exercise real loaders while rejecting full initialization and unrelated I/O.""" + immutable = RESOURCE.read_bytes() + selected = tmp_path / "selected-chebi-nodes.tsv" + selected.write_bytes(immutable) + defaults = tmp_path / "unrelated-defaults" + defaults.mkdir() + if request.param == "populated": + for relative in ( + "data/raw/mediadive/media_detailed.json", + "data/raw/bacdive/record.yaml", + "data/transformed/ontologies/chebi_nodes.tsv", + ): + path = defaults / relative + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text("Poison default input: unit tests must never read this.\n") + monkeypatch.chdir(defaults) + opened = [] + forbidden = [] + real_open = builtins.open + + def reject_initialization(*args, **kwargs): + """Reject full constructors, mapping/ontology setup and network requests.""" + forbidden.append("initialization or network") + raise AssertionError("Category unit test invoked unrelated initialization or network") + + def category_only_open(path, mode="r", *args, **kwargs): + """Admit only the selected fixture in read-only mode; never production files.""" + if Path(path).resolve() != selected.resolve() or mode not in ("r", "rt"): + forbidden.append(str(path)) + raise AssertionError(f"Category unit test attempted unrelated file I/O: {path}") + opened.append(str(Path(path).resolve())) + return real_open(path, mode, *args, **kwargs) + + for module, producer, _ in PRODUCERS.values(): + monkeypatch.setattr(producer, "__init__", reject_initialization) + monkeypatch.setattr(module, "CHEBI_NODES_FILE", selected) + monkeypatch.setattr(module, "open", category_only_open, raising=False) + monkeypatch.setattr(module, "ChemicalMappingLoader", reject_initialization) + monkeypatch.setattr(requests.sessions.Session, "request", reject_initialization) + + def load(source, *, missing=False): + """Call the unmodified loader without constructor side effects or stubbed results.""" + module, producer, _ = PRODUCERS[source] + if missing: + monkeypatch.setattr(module, "CHEBI_NODES_FILE", tmp_path / "missing-selected-category.tsv") + value = producer.__new__(producer) + value.chebi_categories = {} + before = len(opened) + value._load_chebi_categories() + assert len(opened) == before + (0 if missing else 1) + assert not forbidden, "A real loader swallowed an unrelated-I/O rejection" + return value + + yield load + assert forbidden == [] + assert selected.read_bytes() == immutable + assert RESOURCE.read_bytes() == immutable + + +@pytest.mark.parametrize("source", PRODUCERS) +def test_loads_exact_native_category_fixture(category_loader, source): + """Both real loaders must retain every fixture category, not only the default type.""" + assert category_loader(source).chebi_categories == EXPECTED + + +@pytest.mark.parametrize("source", PRODUCERS) +@pytest.mark.parametrize("identifier", EXPECTED) +def test_each_known_category_matches_native_fixture(category_loader, source, identifier): + """Check all selected IDs so a loader returning only fallback ChemicalEntity fails.""" + assert category_loader(source)._get_chebi_category(identifier) == EXPECTED[identifier] + + +@pytest.mark.parametrize("source", PRODUCERS) +def test_exact_category_distribution(category_loader, source): + """Use a finite fixture distribution, never a graph-scale count or conditional skip.""" + assert Counter(category_loader(source).chebi_categories.values()) == { + "biolink:ChemicalEntity": 2, + "biolink:ChemicalRole": 1, + } + + +@pytest.mark.parametrize("source", PRODUCERS) +def test_unknown_identifier_uses_source_fallback(category_loader, source): + """Unknown IDs retain each source-specific declared fallback contract.""" + assert category_loader(source)._get_chebi_category("CHEBI:99999999") == PRODUCERS[source][2] + + +@pytest.mark.parametrize("source", PRODUCERS) +def test_missing_selected_fixture_uses_fallback_without_default_input(category_loader, source): + """A missing explicit category file never authorizes unrelated default data access.""" + value = category_loader(source, missing=True) + assert value.chebi_categories == {} + assert value._get_chebi_category("CHEBI:16828") == PRODUCERS[source][2] + + +@pytest.mark.parametrize("identifier", EXPECTED) +def test_mediadive_ingredient_classification_uses_native_category(category_loader, identifier): + """The real ingredient classifier must preserve the selected ChEBI category.""" + assert ( + category_loader("mediadive")._classify_ingredient_category(identifier, "test compound") == EXPECTED[identifier] + )