diff --git a/docs/reviews/bacdive-ec-substrate-corrections-1249-1250.md b/docs/reviews/bacdive-ec-substrate-corrections-1249-1250.md new file mode 100644 index 00000000..c2d4516f --- /dev/null +++ b/docs/reviews/bacdive-ec-substrate-corrections-1249-1250.md @@ -0,0 +1,70 @@ +# Four finite BacDive EC/substrate corrections (#1249, #1250) + +These corrections address four complete historical mapping rows, not a global +CHEBI:54684 replacement, a CAS identity export, or the complete #286 backlog. +Raw `bacdive_mappings.tsv` remains unchanged. The canonical policy retains all +eight original cells, including the trailing space in the glucoside name. +Actual matching uses complete cells, not historical physical line numbers. + +| Original source line (reviewed snapshot) | Exact source assay | Target | Scope | +|---|---|---|---| +|119|API_ID32E_alpha GLU|CHEBI:91122|alpha-D-glucopyranoside / EC:3.2.1.20| +|28|API_rID32A_alpha GAL|CHEBI:546840|alpha-D-galactoside / EC:3.2.1.22| +|120|API_ID32E_alpha GAL|CHEBI:546840|same molecule; contradictory KEGG cell is audit-only| +|271|API_rID32STR_alpha GAL|CHEBI:546840|alpha-D-galactoside / EC:3.2.1.22| + +The immutable test fixture retains complete original rows, complete native +ChEBI records and incident edges, selected native TSV rows, primary structural +facts/URLs, and scientific input hashes. Native CHEBI:91122's exact IUPAC +synonym and full InChI match [manufacturer 487506](https://www.sigmaaldrich.com/US/en/product/mm/487506), +which identifies alpha-glucosidase substrate use. Native CHEBI:546840's distinct +stereostructure matches [manufacturer N0877](https://www.sigmaaldrich.com/US/en/product/sigma/n0877). +The two stereospecific InChIKeys differ; equal formulas and CAS annotations alone +would not establish this identity. Official EC nomenclature corroborates +[alpha-glucosidase](https://iubmb.qmul.ac.uk/enzyme/EC3/2/1/20.html) and +[alpha-galactosidase](https://iubmb.qmul.ac.uk/enzyme/EC3/2/1/22.html). + +Evidence caveats remain explicit: the glucoside ChEBI prose mentions beta-D- +glucopyranose despite its exact IUPAC name, parent, stereostructure and supplier +evidence agreeing on alpha glucoside. N0877's marketing subtitle names +alpha-glucosidase, while its formal product/application sections and EC authority +support alpha-galactosidase. The original conflicting text is documented, not +used as corroboration. Supplier records do not prove which commercial batch an +API kit used or substrate specificity of every enzyme in an EC class. + +Line120's original [KEGG:C01083](https://www.kegg.jp/entry/C01083) denotes +trehalose, not nitrophenyl galactoside. `withheld_source_fields=KEGG_ID` clears +only that cell in the corrected in-memory row; its complete original claim +remains in the raw file, curated rule, immutable fixture and audit. No KEGG or +CAS identity edge, assay-to-reagent assertion, or taxon phenotype is introduced. + +The selected native export does not declare CHEBI:54684 or an official alias/ +replacement. Native absence is not proof of historical obsolescence, and digit +similarity to CHEBI:546840 is not the correction's basis. The raw ontology SHA +is `a5d40380ab78bde0e8b5a704dbee3cba2bcaa7608be46eebc2f4a92932516c9b`; +the reviewed legacy source SHA is +`f39e9753fe2879877b2cba51c9896f44515f3a5785e790deb0d0bd5ca02c9379`. +These are review witnesses, not runtime pins forcing future source snapshots. + +## Producer boundary + +The real run consumes the finite canonical policy and legacy mapping through +immutable snapshots before opening graph outputs. Malformed/duplicate policy +fields fail; changes within a recognized reviewed scope require new review. +Exact already-corrected rows are idempotent, and an absent cohort is normal. +Unrelated old-ID uses do not acquire either reviewed target. Source row +multiplicity is retained through correction, before ordinary graph deduplication. + +`ec_substrate_corrections.tsv` is mandatory producer audit output. It retains +each original complete source row, actual physical ordinal, original and emitted +seven-field edge claims, reason, primary evidence URLs, source/policy locations +and original read hashes. The producer's final input guards verify the same +admission, including symlink targets; finalization and merge enforce the original +audit bytes through existing producer-audit hooks. Even an empty cohort writes +a header-only audit and binds both required consumed roles. + +The existing EC edge predicate, relation and provenance tier are unchanged. +Ontology dependencies and ordinary finalization still govern native endpoint +closure. No unified artifact, immutable supported MIM table, release pin, raw +record, or shared finalizer is edited. A fresh real source rebuild and graph +reviews remain necessary; focused fixture tests are not production acceptance. diff --git a/docs/reviews/combined-material-scope-20260929/PROMOTION.md b/docs/reviews/combined-material-scope-20260929/PROMOTION.md new file mode 100644 index 00000000..4a9d21cd --- /dev/null +++ b/docs/reviews/combined-material-scope-20260929/PROMOTION.md @@ -0,0 +1,104 @@ +# Combined material-scope candidate: isolated staging, 2026-09-29 + +Status: the accepted derived unified artifact and reviewed native-pair test are +staged in the isolated combined-materials worktree, after source integration at +`620c9c9ea9c5d3e5ab777cb635ab85f9438cd550`. The coordinator preserved the +production checkout and its mapping artifact, archive and user statistics. +The 19 native paired-artifact/predicate tests and 584 broader source-integration +tests passed after staging. Actual combined source replay/comparison, final +repository gates and CI remain pending at this checkpoint. This is not release +acceptance or evidence of fresh production transforms. + +## Exact paired artifacts + +| Role | SHA-256 | +| --- | --- | +| Previous Potato-only unified baseline | `09c44642ab13b234e89ce113f7510aa9efb56969e4b539cb7843b43dcb425ba7` | +| Staged unified candidate, 13,374,300 bytes | `67c48e1bf6bed1f1fef03a0da1d7d1b56af9fc72374dddd703c222fb36df3cd4` | +| Unchanged supported MIM table | `6b52b30e018b369aa322d41dfd4e81fcfae0e895e34d7fe48900abf5835815fb` | +| Unchanged reviewed MIM release pin | `f082c05656a0910c85176eec7b41deeb77c27967aecb80393fddc262819b6d97` | +| Unchanged identity-only writer | `257fe4d8bf16f92eb78d5e375065e030ea56d7c1174d859fecd2baa80ac18c27` | +| Identity-policy fingerprint | `4670cbb9255bdac7e9654fc415c9bfd35e55955fea5f27865a74e55849810b4c` | +| Applied native paired-artifact test | `e77a1c342549097e94542dc8fcc4e424bb475eb94ddff0c47f6a781ff47bc7ad` | + +Upstream remains immutable MIM commit +`1848b0fe521bc2462f165912fcf92d09ad9a8cec`, manifest +`9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af`: +1,747 supported exact mappings and 1,696 names. This is not a new upstream +release, repin or supported-MIM rewrite. The policy fingerprint binds the +curated exclusions plus the shared identity and CAS implementations. + +## Completed candidate checks + +The complete streaming delta removed exactly three generic CHEBI:86658 claims +(two exact, one close) and changed only `object_label` on four retained native +structural-synonym rows. All other row bytes/order and unrelated metadata stayed +identical. Only the description's policy fingerprint changed; the writer, +assertion dates, historical unified version and legacy predicate semantics did +not. The seven original complete rows and native synonym witness remain in the +[immutable P3556 fixture](../../../tests/resources/mediadive/p3556_scope.json), +SHA-256 `7c06f4e2179356ddf1d917a6938cba5e3bdf6669f42a29019b93963e4a70646c`. + +The candidate contains **591,946 rows**, **120,183 distinct targets**, +**336,048 exact** and **255,898 close** matches, with no broad/narrow rows. +The coordinator independently confirmed these counts by a full gzip/CSV stream +under unchanged before/after candidate hashes. + +Three serial native consumer phases completed with exit zero and settled +process groups: a second identity-only export and fresh reader processes with +synonyms enabled and disabled. The second export removed/relabelled zero rows +and reproduced the exact compressed candidate bytes. Both reader modes passed +their finite native controls. Generic lexical phosphatidylcholine can still +reach a class in synonym mode; that is not approval to assign the whole P3556 +product to that class. The source-qualified local guard and real source replay +remain necessary. + +The completed checks ran at frozen pre-BacDive-integration head +`b8f18b838f5797c9dc4f4a64451c27298713e7ef`; they are not represented as new +execution at the subsequent integrated/staged head. Local evidence resides in +the primary checkout under +`data/issue1224-quarantine-20260929.xjuS9H/combined-identity-candidate.9eHuMu/`: + +- [Finite delta receipt](../../../data/issue1224-quarantine-20260929.xjuS9H/combined-identity-candidate.9eHuMu/review-delta-01/result.json): + `a4ec8a0c7a3d30d46208a16ac034301bdb1dae8209cbb951681856b596f0af5d`. +- [Complete original/changed-row ledger](../../../data/issue1224-quarantine-20260929.xjuS9H/combined-identity-candidate.9eHuMu/review-delta-01/complete-row-delta.json): + `13fbefd64155d220a75731dfecfd9e5ab64df3f29e42e561c705d266aad6e50c`. +- [Completed consumer receipt](../../../data/issue1224-quarantine-20260929.xjuS9H/combined-identity-candidate.9eHuMu/consumer-review-01/result.json): + `fa785548ca557e87057acf2d19b4e543437d3e39e28e7b160641617dd90f1100`, + status `PASS_FINITE_CANDIDATE_CONSUMER_CHECKS_ONLY`. + +These project-relative evidence links refer to retained local data, not files +duplicated into the isolated worktree or a published release. + +## Separate source changes and remaining acceptance + +The seven-row mapping delta must not be conflated with recipe/assay changes: + +- [P3556](../mediadive-p3556-scope-1241.md) retains the qualified whole product + on its existing local ID and preserves full imported candidate evidence. +- [Qualified mixed Sugar](../mediadive-sugar-context-1245.md) uses its exact + source-context decision. Generic supported MIM Sugar remains unchanged. +- [Two finite peptone holds](../mediadive-peptone-scope-1248.md) reject only the + reviewed name/CID combinations; independent alternative targets are not + globally banned. Their [26 original occurrences](../../../tests/resources/mediadive/peptone_scope.json) + and all original matching legacy claims remain preserved. +- [Tetramethyl ammonium chloride](../tetramethylammonium-chloride-20260929.md) + gains its finite existing-name route; unqualified tetramethylammonium is not + the salt. +- [BacDive EC/substrate corrections](../bacdive-ec-substrate-corrections-1249-1250.md) + match four complete original rows and preserve their full mandatory audit, + including the withheld contradictory KEGG cell. They do not change this + unified artifact or establish a global identifier replacement. + +The [historical Potato promotion](../mediadive-potato-scope-20260929/PROMOTION.md) +and earlier MIM promotion records are unchanged. No historical receipt is +restamped as current acceptance. The updated paired-artifact test preserves +their CAS, Potato, native-provenance and generic Sugar/Mucin controls. + +Next require approved real source replay/comparison with full +raw/quantity/provenance preservation and manual +node-delta review. Replay may precede or accompany frozen-head repository gates. +All four full gates, green exact-head CI and independent review are required +before PR merge. A fresh all-15-source transform batch, candidate-only merge, +source-to-merged evidence retention and final KG reviews remain separate release +requirements. This checkpoint does not close the broader #286 backlog. diff --git a/docs/reviews/mediadive-p3556-scope-1241.md b/docs/reviews/mediadive-p3556-scope-1241.md new file mode 100644 index 00000000..82d47f9b --- /dev/null +++ b/docs/reviews/mediadive-p3556-scope-1241.md @@ -0,0 +1,104 @@ +# Source-qualified Sigma P3556 material scope (#1241) + +## Evidence and scientific limit + +The saved MediaDive observation `mediadive.solution:629#recipe/1` names +`L-α-Phosphatidylcholine`, compound 2082, and explicitly supplies +`attribute: SIGMA P3556`. Its amount is 5 mg and its source `g_l` value is 5. +Those original values are retained, not recalculated or silently corrected. +The source does **not** assert a CAS identifier. + +The [supplier product sheet](https://www.sigmaaldrich.com/deepweb/assets/sigmaaldrich/product/documents/152/475/p3556pis.pdf) +describes an egg-yolk material with several fatty-acid constituents, not one +pure acyl species. The supplier's [drug-delivery brochure](https://b2b.sigmaaldrich.com/deepweb/assets/sigmaaldrich/marketing/global/documents/709/352/polymeric-drug-delivery-techniques-web.pdf) +also depicts variable fatty-acid residues for P3556 (PDF page index 18, +printed page 17). The [product page](https://www.sigmaaldrich.com/US/en/product/sigma/p3556) +additionally supplies a fixed structure. That catalog structure is not proof +of whole-product molecular purity; its stereolayer differs from the native +ChEBI witness. These primary documents were read online. Local PDF downloads +failed, so no local PDF checksum is claimed. + +Native `CHEBI:86658` has a generic preferred label but explicit synonyms, +SMILES and InChI identifying a particular 16:0/(9E,12E)-18:2 structure. +The immutable fixture `tests/resources/mediadive/p3556_scope.json` preserves +that structured witness, all eight selected full original mapping rows, +the exact raw occurrence, and hashes of the original raw/mapping inputs. +The error is assigning the **whole product quantity** to that molecule. +This is not a claim that the molecule cannot occur in the material. + +## Transform and mapping behavior + +- Two finite existing-policy exclusions reject generic `L-alpha`/`L-α` + phosphatidylcholine names and the exact `MIM:L-alpha-Phosphatidylcholine` + pair for `CHEBI:86658`. The policy authority label is an actual native + structured synonym, not the misleadingly generic preferred label. +- The explicit source qualifier `SIGMA P3556` (case and whitespace only) + keeps a recipe occurrence on its existing MediaDive ingredient/solution + ID before unified, legacy or embedded identity selection. The current + observation therefore remains `mediadive.ingredient:2082` with its raw + source record, assertion ID, quantities and observation/manual provenance. + Neither ingredient number, recipe number, a generic name nor a fuzzy + catalog-code match triggers this product decision. +- A supplied occurrence, including `{}`, never inherits another recipe's + embedded product qualifier. Direct calls without an occurrence can use an + explicit qualifier in the embedded source compound record. +- Local category remains the existing broad `biolink:ChemicalEntity` for + the current source name. This change does not assert new native mixture + typing, a replacement ChEBI charge/class, CAS or other molecular identity. +- Explicitly structured native molecular routes and unrelated generic + phosphatidylcholine classification claims remain eligible. There is no + global ban on `CHEBI:86658`, `CHEBI:16110`, `CHEBI:49183` or CAS 8002-43-5. + Eligibility alone is not new identity approval. + +## Reversible candidate evidence + +The existing required producer sidecar +`mediadive_material_scope_quarantine.tsv` retains available candidate rows +separately from graph identity. Its schema, producer-time byte snapshot, +source-finalization copying and public-admission checks are unchanged. +Potato/extract and P3556 candidates are separate finite profiles. + +For a P3556 observation, the audit preserves applicable current unified and +legacy rows (matching the actual occurrence spelling), the original exact +generic MIM claim, and actual embedded identifiers when present. Duplicate +input rows have separate locators. Full original row columns, raw record, +source/candidate path and hash, retained local target, qualifier, authority +URI and reason are retained. A contradictory structure-specific display +name cannot override the product qualifier; its applicable mapping claims +are still auditable. These rows are **available candidates**, not assertions +that every lookup was attempted, and are never fabricated raw CAS claims. + +Review #1243 corrected canonical object-label coverage: eligible identity, +attribute, canonical-name and synonym rows can supply the reader's object +label; synonym rows additionally supply the subject label. A row matching +both fields is retained once per original locator, while duplicate physical +rows remain distinct. Broader, hydrate and annotation rows do not become +grounding candidates. The old fixture's four structured-synonym rows also +carry the generic object label, so they remain auditable until a reviewed +derived export relabels that metadata. Audit row counts are not historical +constants. + +Conditional consumed-input role `material_scope_supported` binds the +canonical unchanged `mappings/ingredient_mappings.sssom.tsv` only when the +selected source contains P3556. It is audit evidence, not an active grounding +source. It preserves the complete original imported assertion after a future +derived unified export removes the held pair. The supported MIM table and +release pin remain immutable. Empty source cohorts or removed candidates +are valid; a genuinely empty audit retains its header. + +## Verification and deployment boundary + +Hermetic tests cover native reader/direct-pair denial, explicit structure +controls, legacy and bounded/conservative regeneration, all occurrence +fallbacks, raw quantities, duplicate full-row audit evidence, profile +isolation, input-origin/drift protection, and real source finalization/reuse/ +public admission. The tiny identity-only refresh removes three generic rows +and replaces the misleading object label on four retained structural rows; +its second cycle is byte-identical. Conservative regeneration preserves full +quarantined originals and multiplicity. + +This source change alone is not a rebuilt graph or release acceptance. +A separate reviewed derived-mapping candidate, source replay, transform +freshness checks and merged-graph reviews remain required. No immutable +supported MIM release, production raw data or production archive is modified +by this implementation stage. diff --git a/docs/reviews/mediadive-peptone-scope-1248.md b/docs/reviews/mediadive-peptone-scope-1248.md new file mode 100644 index 00000000..03ea64ca --- /dev/null +++ b/docs/reviews/mediadive-peptone-scope-1248.md @@ -0,0 +1,76 @@ +# Two finite peptone material identity holds (#1248) + +## Scientific scope + +`Soy peptone` and `Vitamin-free casamino acids` do not establish exact identity +with the particular structural record PubChem:167312541. Keep the original +material observations and imported lookup claims separately. No replacement +CAS, ChEBI, PubChem or inferred constituent identity is approved here. + +The [PubChem record](https://pubchem.ncbi.nlm.nih.gov/compound/167312541) really +does carry soybean-peptone annotations and casein-digest synonyms; preferred +name disagreement alone is not the exclusion criterion. It also describes a +specific multicomponent structure, formula C290H252N8O72. Its annotations do not +demonstrate that either variable digest is exactly that structure. +[PubChem distinguishes](https://pubchem.ncbi.nlm.nih.gov/docs/glossary) unique +compound structures from deposited substances. + +The manufacturer's [vitamin-assay casamino acids](https://www.thermofisher.com/order/catalog/product/228820) +is an acid hydrolysate of casein; [Bacto Soytone](https://www.thermofisher.com/order/catalog/product/kr/en/243620) +is an enzymatic soybean digest. Two saved soy occurrences name Bacto/BD Soytone; +these qualifiers remain specific to their own records, not inferred for the +other 23. The evidence supports distinct source materials, not an exact +supplier or composition for every occurrence, and not a global ban on the CID +or its registry annotations. + +## Implementation boundary + +Two anchored `name_pattern` rules in `ingredient_identity_exclusions.tsv` +reject only the reviewed name/target combinations. Existing case, punctuation +and separator normalization is retained, including the `PubChem` and +`pubchem.compound` prefix spellings. Direct CID references, unrelated material +names and independently supported other targets remain eligible. There is no +new raw compound-ID rule or unconditional supplier/product override. + +The existing MediaDive ingredient checks enforce these rules on unified names, +legacy strict/hydrate lookups, embedded identifiers and adjacent solution-name +resolution. A missing identity retains the existing source-local ID. The two +observed compound IDs 75 and 654 describe this saved cohort; they are not +hardcoded policy keys. + +The existing producer-owned `mediadive_material_scope_quarantine.tsv` gains a +finite peptone profile. It records only eligible claims to the reviewed CID, +not unrelated alternative targets or weak/nonidentity rows. It preserves +complete original mapping columns, duplicate row ordinals, source records, +locators, original qualifiers, selected retained targets and consumed-input +hashes. The reason is `digest_material_not_demonstrated_structural_identity`. +Repeated visits do not duplicate evidence; distinct recipe positions remain +distinct. No historical row is invented when the current input lacks it. + +The canonical shared policy and unified input remain required audit origins. +Producer-time audit binding, source finalization and public merge admission +continue to enforce the existing mandatory audit contract. No audit schema, +merge rewriting or additional finalizer behavior is introduced. + +## Immutable fixture and regression scope + +`tests/resources/mediadive/peptone_scope.json`, SHA-256 +`41a556f70cc0ae696d01895371214af7f25bb62c9d7a7b9a73c0fe589fed1fa3`, +retains all 26 saved observations (25 soy, one vitamin-free), their complete +old edge rows, and all ten original matching strict/hydrate claims (four soy +and one vitamin-free in each table). It preserves source-file hashes and exact +original physical rows; unrelated multi-line source cells were parsed with +the actual CSV dialect during extraction. This is a bounded fixture, not a +complete source snapshot or a newly accepted rebuild. + +Tests exercise actual cold lookup, both target-prefix spellings, unified and +legacy fallback, embedded IDs, nested solution names, unaffected direct CID +and other-name routes, original quantities, duplicate observations, current +policy origins, immutable audit bytes and real finalization/admission. A tiny +identity-only exporter regression verifies stable second-cycle behavior; it +does not regenerate the real candidate. + +The unified archive, supported MIM TSV and immutable MIM release pin are not +changed by this source patch. Derived candidate review, actual MediaDive replay, +full tests/CI, all-source freshness/rebuild and merged-KG review remain separate +gates. This does not close umbrella #286 or certify a release. diff --git a/docs/reviews/mediadive-sugar-context-1245.md b/docs/reviews/mediadive-sugar-context-1245.md new file mode 100644 index 00000000..78675f4e --- /dev/null +++ b/docs/reviews/mediadive-sugar-context-1245.md @@ -0,0 +1,59 @@ +# MediaDive's qualified mixed-sugar stock (#1245) + +The original JCM 537 recipe describes a stock solution containing xylose, +maltose and cellobiose, each at 250 mM. The saved MediaDive record at +`mediadive.solution:4338#recipe/16` names `Sugar`, compound 1795, and preserves +that exact qualifier with an addition of `0.1 ml`. It supplies no CAS field. +Primary source: https://www.jcm.riken.jp/cgi-bin/jcm/jcm_grmd?GRMD=537; +MediaDive display: https://mediadive.dsmz.de/medium/J537. + +The immutable generic MIM Sugar assertion targets NCIT:C71939. Its native +Food semantics do not demonstrate that this specific three-sugar solution +is equivalent to that identity. This is a finite withheld grounding, not +proof that NCIT excludes every nonsucrose substance, a regulatory food-grade +claim, or a global Sugar/NCIT prohibition. + +## Source decision and evidence + +`mappings/canonical/mediadive_material_context_dispositions.tsv` has one exact +source-name/attribute decision, with case/whitespace-only recognition. +The withheld target is an example explaining the review, not a global ban. +The source occurrence remains on its existing local MediaDive ingredient +(or nested-solution) ID before unified, legacy or embedded fallback. An +explicit empty occurrence never borrows the cached compound's qualifier. +Unqualified or differently qualified Sugar observations retain normal lookup. +No component assertions, replacement molecular/CAS identity, or inferred +utilization are added. The existing broad ChemicalEntity category is retained; +this patch does not claim newly implemented mixture-specific categorization. + +The selected decision is consumed as `material_scope_context_policy` before +graph output and guarded with the existing SourceAdmission/snapshot contract. +Missing, duplicate or changed decision content fails closed. Finalization's +existing consumed-input ledger and finalization-input fingerprint include the +exact selected file. Sugar audit rows name this actual policy and hash; +Potato and P3556 retain their original identity-exclusion policy bindings. +Conditional supported-MIM evidence applies to both P3556 and Sugar, including +selected-context/header-only cases after all grounding candidates disappear. + +The existing material audit retains complete eligible unified, supported, +legacy and real embedded candidate rows, their physical ordinals, source raw +JSON, quantity, provenance and assertion position. Duplicate physical input +rows remain distinct. These are available original claims, not claims that +the held source invoked each lookup. The supported MIM table, release pin, +unified artifact and raw inputs are unchanged. + +## Verification boundary + +The immutable `tests/resources/mediadive/sugar_context.json` records the saved +raw observation and all three original mapping rows with origin digests. +Tests use only that fixture and temporary inputs. They exercise ordinary +generic Sugar lookups, every current fallback namespace, explicit empty and +other contexts, physical claim/occurrence multiplicity, selected file drift, +actual new policy origin, coexistence with Potato/P3556, and a tiny real +source graph through finalization/repeat/public admission. + +Full code gates, independent review, actual source replay, source-to-merged +preservation and final KG release checks remain separate. This document is +not a claim of completed production rebuild or release acceptance. The local +chemical-mapping skill and reviewed MIM runbook required retaining immutable +upstream inputs rather than refreshing or rewriting their generic evidence. diff --git a/docs/reviews/tetramethylammonium-chloride-20260929.md b/docs/reviews/tetramethylammonium-chloride-20260929.md new file mode 100644 index 00000000..1794b1f3 --- /dev/null +++ b/docs/reviews/tetramethylammonium-chloride-20260929.md @@ -0,0 +1,38 @@ +# Finite tetramethyl ammonium chloride alias (#1242) + +This reviewed name route addresses a current coverage gap found while tracing +#286's recovered historical CAS cohort. It does not validate every historical +CAS assignment or authorize a graph release. + +Current saved MediaDive compound 720 and solution 1577, recipe position 4, +explicitly name `Tetramethyl ammonium chloride`, with amount 1 g and g_l 1. +The source has no CAS field. Its finalized pre-change target is local +`mediadive.ingredient:720`. Native active CHEBI:7070 declares the complete +chloride salt, with preferred name `N,N,N-Trimethylmethanaminium chloride`, +synonym `Tetramethylammonium chloride`, and CAS annotation `cas:75-57-0`. + +[MediaDive's ingredient groups](https://www.bacmedia.dsmz.de/ingredient-groups?asc=0&limit=100&order=name) +independently associates the exact spaced spelling with CAS 75-57-0 and formula +C4H12N.Cl. It separately lists unqualified `Tetramethyl ammonium`, which is not +the chloride salt. [PubChem CID 6379](https://pubchem.ncbi.nlm.nih.gov/compound/6379) +corroborates the chloride identity and distinguishes parent CID 6380. Public +records were checked on 2026-09-29; no immutable web-content checksum is claimed. +Independent agent review corroborated the source, native target, and finite alias. + +Only one existing `ingredient_name_scopes.tsv` entry is added. It uses the +existing bounded name-scope normalization, not a new global whitespace-removal +algorithm. The target must already be declared by the selected mapping input. +The unqualified raw compound 1649 / solution 3797 recipe position 1 stays +`kgmicrobe.compound:tetramethyl_ammonium`; other counterions and hydrates receive +no new mapping. No public CAS value is inserted into either raw source record. + +`tests/resources/tetramethylammonium_chloride.json` retains the two exact selected +raw records, their positions, prior targets, a labeled native-node projection, +and SHA-256 identities of the complete original source/native files. Tests use +small declared-target inputs; they never consult a live service. The fixture's +native xref is an annotation, not a newly emitted equivalence edge. + +Immutable supported MIM and pin bytes are unchanged. Full producer replay must +still reconcile every changed target and node while preserving quantities, +source assertions and raw evidence. Final all-source freshness and merged-KG +review remain separate gates; focused tests are not build acceptance. diff --git a/kg_microbe/transform_utils/bacdive/bacdive.py b/kg_microbe/transform_utils/bacdive/bacdive.py index 1937afc6..81267ba5 100644 --- a/kg_microbe/transform_utils/bacdive/bacdive.py +++ b/kg_microbe/transform_utils/bacdive/bacdive.py @@ -25,6 +25,21 @@ import yaml from tqdm import tqdm +from kg_microbe.transform_utils.bacdive.ec_substrate_corrections import ( + AUDIT_FIELDS as EC_SUBSTRATE_AUDIT_FIELDS, +) +from kg_microbe.transform_utils.bacdive.ec_substrate_corrections import ( + AUDIT_FILE as EC_SUBSTRATE_AUDIT_FILE, +) +from kg_microbe.transform_utils.bacdive.ec_substrate_corrections import ( + POLICY_RELATIVE as EC_SUBSTRATE_POLICY, +) +from kg_microbe.transform_utils.bacdive.ec_substrate_corrections import ( + REQUIRED_INPUTS as EC_SUBSTRATE_INPUTS, +) +from kg_microbe.transform_utils.bacdive.ec_substrate_corrections import ( + prepare_corrections, +) from kg_microbe.transform_utils.bacdive.emission import ( DEPOSIT_CONFLICT_HEADER, RESOLUTION_COLLAPSED, @@ -267,7 +282,10 @@ class BacDiveTransform(Transform): DATA_INPUTS = ( "mappings/isolation_source_to_ontology.tsv", "mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz", + EC_SUBSTRATE_POLICY, ) + REQUIRED_CONSUMED_INPUTS = EC_SUBSTRATE_INPUTS + REQUIRED_AUDIT_FILES = (EC_SUBSTRATE_AUDIT_FILE,) def __init__( self, @@ -2016,6 +2034,7 @@ def _emit_metabolite_utilization( def run(self, data_file: Union[Optional[Path], Optional[str]] = None, show_status: bool = True): """Run the transformation.""" + self.begin_consumed_inputs() # Resolve the ontology adapter before any output file is opened. The # adapters are lazy, so without this the first lookup happens deep inside # the write loop — and a fatal ontology error there leaves the previous, @@ -2023,6 +2042,9 @@ def run(self, data_file: Union[Optional[Path], Optional[str]] = None, show_statu # before the truncation costs a run; failing after costs the outputs. resolve_adapter(self.ncbi_impl) self._prepare_assay_outputs() + ec_substrate_rows, ec_substrate_audit, ec_substrate_admission = prepare_corrections( + self, BACDIVE_TMP_DIR / BACDIVE_MAPPING_FILE + ) # replace with downloaded data filename for this source input_file = os.path.join(self.input_base_dir, "bacdive_strains.json") # must exist already # Read the JSON file into the variable input_json @@ -2123,7 +2145,6 @@ def run(self, data_file: Union[Optional[Path], Optional[str]] = None, show_statu with ( open(str(BACDIVE_TMP_DIR / "bacdive.tsv"), "w") as tsvfile_1, open(str(BACDIVE_TMP_DIR / "bacdive_physiology_metabolism.tsv"), "w") as tsvfile_2, - open(str(BACDIVE_TMP_DIR / BACDIVE_MAPPING_FILE), "r") as tsvfile_3, open(str(BACDIVE_TMP_DIR / "bacdive_name_tax_classification.tsv"), "w") as tsvfile_4, # Atomic: a run that dies here used to leave a truncated pair on # disk with nothing marking it partial -- a SIGTERM during the @@ -2133,6 +2154,7 @@ def run(self, data_file: Union[Optional[Path], Optional[str]] = None, show_statu # before it. Same rule as lpsn (#820) and lpsn_api (#985). atomic_write(self.output_node_file, newline="") as node, atomic_write(self.output_edge_file, newline="") as edge, + atomic_write(Path(self.output_dir) / EC_SUBSTRATE_AUDIT_FILE, newline="") as ec_audit, # This source-local diagnostic is not covered by the ordinary graph # receipt: postbuild acceptance must hash/admit it explicitly (#1185). atomic_write(Path(self.output_dir) / CHEMICAL_IDENTITY_CONFLICTS_FILE) as chemical_conflicts, @@ -2145,6 +2167,11 @@ def run(self, data_file: Union[Optional[Path], Optional[str]] = None, show_statu writer_2.writerow(PHYS_AND_META_COL_NAMES) writer_3 = tsv_writer(tsvfile_4) writer_3.writerow(NAME_TAX_CLASSIFICATION_COL_NAMES) + ec_audit_writer = tsv_writer(ec_audit, quoting=csv.QUOTE_NONE, quotechar=None) + ec_audit_writer.writerow(EC_SUBSTRATE_AUDIT_FIELDS) + ec_audit_writer.writerows( + [entry[column] for column in EC_SUBSTRATE_AUDIT_FIELDS] for entry in ec_substrate_audit + ) node_writer = tsv_writer(node) node_writer.writerow(self.node_header) @@ -2192,18 +2219,14 @@ def run(self, data_file: Union[Optional[Path], Optional[str]] = None, show_statu node_writer.writerows(self.assay_target_nodes_generated) custom_curie_data = yaml.safe_load(cc_file) - bacdive_mappings_list_of_dicts = list(csv.DictReader(tsvfile_3, delimiter="\t")) # Generate EC→substrate edges from bacdive_mappings.tsv # These represent enzymatic reactions where an enzyme (EC number) acts on a substrate (ChEBI) # Note: EC and ChEBI nodes are normally created by the ontologies transform. - # Obsolete CHEBI IDs (e.g. CHEBI:54684 4-Nitrophenyl-alpha-D-galactoside) are - # missing from chebi_nodes.tsv, so we emit a labelled stub here using the - # `substrate` column from bacdive_mappings.tsv to prevent KGX from creating - # a biolink:NamedThing stub at merge time. - ec_substrate_edges, ec_substrate_stub_nodes = self._generate_ec_substrate_rows( - bacdive_mappings_list_of_dicts - ) + # A missing native ID is not proof of obsolescence or a replacement. + # Exact reviewed row corrections run before graph output (#1249); + # other legacy claims retain their original endpoint and label. + ec_substrate_edges, ec_substrate_stub_nodes = self._generate_ec_substrate_rows(ec_substrate_rows) # Write EC→substrate edges if ec_substrate_edges: @@ -3717,6 +3740,11 @@ def run(self, data_file: Union[Optional[Path], Optional[str]] = None, show_statu with open(METABOLITE_MAPPING_FILE, "w") as f: json.dump(METABOLITE_MAP, f, indent=4) + ec_substrate_admission.verify() + self.verify_consumed_inputs() + + self.record_producer_audit(EC_SUBSTRATE_AUDIT_FILE) + # Write non-matching media links to a file media_links_file = os.path.join(self.output_dir, "bacdive_media_links.txt") with atomic_write(media_links_file) as f: @@ -3761,6 +3789,8 @@ def run(self, data_file: Union[Optional[Path], Optional[str]] = None, show_statu dedup_on_sort_column=True, ) drop_duplicates(self.output_edge_file) + ec_substrate_admission.verify() + self.verify_consumed_inputs() logger.info( "[bacdive] LPSN cross-refs: %s matched, %s unmatched, %s ambiguous", diff --git a/kg_microbe/transform_utils/bacdive/ec_substrate_corrections.py b/kg_microbe/transform_utils/bacdive/ec_substrate_corrections.py new file mode 100644 index 00000000..7ae8c017 --- /dev/null +++ b/kg_microbe/transform_utils/bacdive/ec_substrate_corrections.py @@ -0,0 +1,197 @@ +"""Apply finite reviewed legacy EC/substrate corrections without raw or global aliases.""" + +import csv +import json +import re +from pathlib import Path + +from kg_microbe.merge_utils.source_admission import SourceAdmission +from kg_microbe.transform_utils.constants import ( + AGENT_TYPE_COLUMN, + ENZYME_TO_SUBSTRATE_EDGE, + HAS_INPUT_RELATION, + KNOWLEDGE_ASSERTION, + KNOWLEDGE_LEVEL_COLUMN, + MANUAL_AGENT, + OBJECT_COLUMN, + PREDICATE_COLUMN, + PRIMARY_KNOWLEDGE_SOURCE_COLUMN, + RELATION_COLUMN, + SUBJECT_COLUMN, +) + +POLICY_RELATIVE = "mappings/canonical/bacdive_ec_substrate_corrections.tsv" +POLICY_PATH = Path(__file__).resolve().parents[3] / POLICY_RELATIVE +AUDIT_FILE = "ec_substrate_corrections.tsv" +REQUIRED_INPUTS = ("ec_substrate_correction_policy", "ec_substrate_legacy_mapping") +SOURCE_FIELDS = ("CHEBI_ID", "substrate", "KEGG_ID", "CAS_RN_ID", "EC_ID", "enzyme", "pseudo_CURIE", "reaction_name") +POLICY_FIELDS = SOURCE_FIELDS + ( + "correction_id", + "corrected_chebi_id", + "withheld_source_fields", + "reason", + "evidence_uri", +) +AUDIT_FIELDS = ( + "correction_id", + "status", + "source_file", + "source_sha256", + "source_line", + "source_row_json", + "original_edge_json", + "emitted_edge_json", + "withheld_source_fields", + "reason", + "evidence_uri", + "policy_file", + "policy_sha256", +) + + +def _rows(stream, *, policy=False): + """Reject ambiguous TSV shape while retaining literal source values and physical ordinals.""" + reader = csv.reader(stream, delimiter="\t") + header = next(reader, None) + required = set(POLICY_FIELDS if policy else ("CHEBI_ID", "EC_ID", "substrate")) + if ( + not header + or len(set(header)) != len(header) + or any(not value for value in header) + or not required <= set(header) + or (policy and set(header) != required) + ): + raise ValueError("Invalid EC substrate correction policy/source header") + previous_line = reader.line_num + for cells in reader: + first_line = previous_line + 1 + previous_line = reader.line_num + if len(cells) != len(header): + raise ValueError(f"Malformed EC substrate row at line {first_line}") + yield first_line, dict(zip(header, cells, strict=True)) + + +def load_policy(stream): + """Read explicit complete-row curation, never partial names or global old-ID replacements.""" + policies = [] + seen_ids, seen_rows, seen_assays = set(), set(), set() + for _, rule in _rows(stream, policy=True): + original = {key: rule[key] for key in SOURCE_FIELDS} + identity = tuple(original.values()) + if ( + not re.fullmatch(r"[a-z0-9][a-z0-9_-]+", rule["correction_id"]) + or not re.fullmatch(r"CHEBI:[1-9][0-9]*", rule["CHEBI_ID"]) + or not re.fullmatch(r"CHEBI:[1-9][0-9]*", rule["corrected_chebi_id"]) + or rule["corrected_chebi_id"] == rule["CHEBI_ID"] + or rule["withheld_source_fields"] not in ("", "KEGG_ID") + or (rule["withheld_source_fields"] and not rule["KEGG_ID"]) + or not re.fullmatch(r"EC:[0-9]+(?:\.[0-9-]+){3}", rule["EC_ID"]) + or any(not rule[key].strip() for key in ("substrate", "enzyme", "pseudo_CURIE", "reaction_name", "reason")) + or not rule["pseudo_CURIE"].startswith("kgmicrobe.assay:") + or not rule["evidence_uri"] + or any(character in rule["reason"] for character in "\t\r\n") + or any( + not uri.startswith("https://") or any(c.isspace() for c in uri) + for uri in rule["evidence_uri"].split("|") + ) + or rule["correction_id"] in seen_ids + or identity in seen_rows + or rule["pseudo_CURIE"] in seen_assays + ): + raise ValueError("Invalid or duplicate EC substrate correction rule") + seen_ids.add(rule["correction_id"]) + seen_rows.add(identity) + seen_assays.add(rule["pseudo_CURIE"]) + policies.append(rule) + if not policies: + raise ValueError("Empty EC substrate correction policy") + return policies + + +def _edge(row, knowledge_source): + """Describe the exact existing emitter's scientific/provenance fields for a source claim.""" + return { + SUBJECT_COLUMN: row["EC_ID"].strip(), + PREDICATE_COLUMN: ENZYME_TO_SUBSTRATE_EDGE, + OBJECT_COLUMN: row["CHEBI_ID"].strip(), + RELATION_COLUMN: HAS_INPUT_RELATION, + PRIMARY_KNOWLEDGE_SOURCE_COLUMN: knowledge_source, + KNOWLEDGE_LEVEL_COLUMN: KNOWLEDGE_ASSERTION, + AGENT_TYPE_COLUMN: MANUAL_AGENT, + } + + +def _json(value): + """Encode complete original cells reversibly inside a literal TSV cell.""" + return json.dumps(value, sort_keys=True, separators=(",", ":"), ensure_ascii=True) + + +def apply_corrections(source_rows, policies, *, knowledge_source): + """Correct only complete reviewed rows; known-scope drift requires a new review.""" + corrected, audit = [], [] + for line_number, row in source_rows: + selected, known_scope = [], False + for rule in policies: + # These anchors identify a potential reviewed claim, never authorize + # its correction. Blanking one identity/record cell cannot evade the + # complete-row comparison. Unrelated uses of the old ChEBI ID stay put. + candidate = row.get("pseudo_CURIE") == rule["pseudo_CURIE"] or ( + row.get("CHEBI_ID") == rule["CHEBI_ID"] + and (row.get("EC_ID") == rule["EC_ID"] or row.get("CAS_RN_ID") == rule["CAS_RN_ID"]) + ) + known_scope = known_scope or candidate + original = {key: rule[key] for key in SOURCE_FIELDS} + target_corrected = {**original, "CHEBI_ID": rule["corrected_chebi_id"]} + expected = dict(target_corrected) + if rule["withheld_source_fields"]: + expected[rule["withheld_source_fields"]] = "" + if row in (original, target_corrected, expected): + selected.append((rule, expected)) + if not selected: + if known_scope: + raise ValueError(f"Reviewed EC substrate scope changed at line {line_number}; review required") + corrected.append(dict(row)) + continue + if len(selected) != 1: + raise ValueError(f"Ambiguous EC substrate correction at line {line_number}") + rule, expected = selected[0] + status = "already_corrected" if row == expected else "corrected" + corrected.append(expected) + audit.append( + { + "correction_id": rule["correction_id"], + "status": status, + "source_line": str(line_number), + "source_row_json": _json(row), + "original_edge_json": _json(_edge(row, knowledge_source)), + "emitted_edge_json": _json(_edge(expected, knowledge_source)), + "withheld_source_fields": rule["withheld_source_fields"], + "reason": rule["reason"], + "evidence_uri": rule["evidence_uri"], + } + ) + return corrected, audit + + +def prepare_corrections(transform, legacy_path, *, policy_path=POLICY_PATH): + """Bind actual policy/source reads before opening graph outputs, retaining the original admission.""" + admission = SourceAdmission() + admission.capture(policy_path) + admission.capture(legacy_path) + with transform.consume_input(REQUIRED_INPUTS[0], policy_path) as policy_stream: + policies = load_policy(policy_stream) + with transform.consume_input(REQUIRED_INPUTS[1], legacy_path) as source_stream: + rows, audit = apply_corrections(_rows(source_stream), policies, knowledge_source=transform.knowledge_source) + snapshots = transform.consumed_input_snapshots + for entry in audit: + entry.update( + { + "source_file": snapshots[REQUIRED_INPUTS[1]]["path"], + "source_sha256": snapshots[REQUIRED_INPUTS[1]]["sha256"], + "policy_file": snapshots[REQUIRED_INPUTS[0]]["path"], + "policy_sha256": snapshots[REQUIRED_INPUTS[0]]["sha256"], + } + ) + admission.verify() + transform.verify_consumed_inputs() + return rows, audit, admission diff --git a/kg_microbe/transform_utils/mediadive/material_scope_audit.py b/kg_microbe/transform_utils/mediadive/material_scope_audit.py index 0c439b0f..f2a630dd 100644 --- a/kg_microbe/transform_utils/mediadive/material_scope_audit.py +++ b/kg_microbe/transform_utils/mediadive/material_scope_audit.py @@ -1,4 +1,4 @@ -"""Preserve finite Potato/extract grounding candidates separately from identity (#1236).""" +"""Preserve reviewed material grounding candidates separately from identity.""" import csv import gzip @@ -12,21 +12,32 @@ from kg_microbe.transform_utils.constants import ( CAS_RN_KEY, CAS_RN_PREFIX, + CHEBI_KEY, + CHEBI_PREFIX, COMPOUND_ID_KEY, COMPOUND_KEY, DATA_KEY, ID_COLUMN, + KEGG_KEY, + KEGG_PREFIX, + MEDIADIVE_INGREDIENT_PREFIX, + MEDIADIVE_SOLUTION_PREFIX, + PUBCHEM_KEY, + PUBCHEM_PREFIX, RECIPE_KEY, + SOLUTION_ID_KEY, SOLUTION_KEY, SOLUTIONS_KEY, SOURCE_ASSERTION_ID_COLUMN, SOURCE_RECORD_COLUMN, ) +from kg_microbe.utils.chemical_mapping_utils import normalize_name from kg_microbe.utils.ingredient_identity import ( IDENTITY_POLICY, _policy_target_key, _scope_key, ingredient_mapping_allowed, + ingredient_xref_allowed, ) from kg_microbe.utils.source_finalization import SourceFinalizationRequired from kg_microbe.utils.sssom_identity_policy import classify_mapping_row @@ -39,6 +50,33 @@ AUTHORITY_URI = "https://www.mhsr.sk/uploads/files/Zwx10C5G.pdf#page=21" UNIFIED_ROLE = "material_scope_unified" POLICY_ROLE = "material_scope_policy" +SUPPORTED_ROLE = "material_scope_supported" +CONTEXT_ROLE = "material_scope_context_policy" +CONTEXT_POLICY = "mappings/canonical/mediadive_material_context_dispositions.tsv" +SUGAR_ATTRIBUTE = "250 mM each of xylose, maltose and cellobiose" +SUGAR_REASON = "qualified_sugar_solution_not_demonstrated_food_identity" +SUGAR_URI = "https://www.jcm.riken.jp/cgi-bin/jcm/jcm_grmd?GRMD=537" +CONTEXT_FIELDS = { + "source_name", + "source_attribute", + "disposition", + "reason", + "evidence_uri", + "authority_label", + "withheld_target_example", +} +P3556_MIM = "MIM:L-alpha-Phosphatidylcholine" +P3556_TARGET = "CHEBI:86658" +P3556_AUTHORITY_LABEL = "Sigma P3556 egg-yolk phosphatidylcholine material (variable fatty-acid composition)" +P3556_AUTHORITY_URI = "https://www.sigmaaldrich.com/deepweb/assets/sigmaaldrich/product/documents/152/475/p3556pis.pdf" +PEPTONE_TARGET = "pubchem.compound:167312541" +PEPTONE_AUTHORITY_LABEL = "Glycerides, C8-10 mono-and di-" +PEPTONE_AUTHORITY_URI = "https://pubchem.ncbi.nlm.nih.gov/compound/167312541" +PEPTONE_REASON = "digest_material_not_demonstrated_structural_identity" +PEPTONE_PATTERNS = { + "Soy peptone": "(?i)^soy[ _-]+peptone$", + "Vitamin-free casamino acids": "(?i)^vitamin[ _-]+free[ _-]+casamino[ _-]+acids$", +} AUDIT_HEADER = ( SOURCE_ASSERTION_ID_COLUMN, SOURCE_RECORD_COLUMN, @@ -72,6 +110,45 @@ def _potato(value): return isinstance(value, str) and _scope_key(re.sub(r"[^\w\s-]", "", value)) == "potato" +def p3556_material(raw): + """Recognize only an explicit supplier/product qualifier, never a name or source ID.""" + value = raw.get("attribute") if isinstance(raw, dict) else None + return isinstance(value, str) and " ".join(value.split()).casefold() == "sigma p3556" + + +def _peptone(value): + """Match only the two reviewed digest names, including shared policy normalization.""" + return isinstance(value, str) and _scope_key(re.sub(r"[^\w\s-]", "", value)) in { + "soy peptone", + "vitamin free casamino acids", + } + + +def sugar_material(raw, name=None): + """Match only the reviewed source name and stock qualifier, with case/whitespace normalization.""" + if not isinstance(raw, dict): + return False + name = raw.get(COMPOUND_KEY, raw.get(SOLUTION_KEY, name)) + attribute = raw.get("attribute") + return ( + isinstance(name, str) + and " ".join(name.split()).casefold() == "sugar" + and isinstance(attribute, str) + and " ".join(attribute.split()).casefold() == SUGAR_ATTRIBUTE.casefold() + ) + + +def _profile(raw): + """Give explicit product evidence priority over a possibly contradictory display name.""" + if p3556_material(raw): + return "p3556" + if sugar_material(raw): + return "sugar" + if _peptone(raw.get(COMPOUND_KEY, raw.get(SOLUTION_KEY))): + return "peptone" + return "potato" if _potato(raw.get(COMPOUND_KEY, raw.get(SOLUTION_KEY))) else None + + def _reason(raw): """Recognize only reviewed explicit forms; unknown qualifiers do not imply fresh tuber.""" qualifiers = {key: raw[key] for key in ("condition", "attribute") if key in raw} @@ -120,86 +197,226 @@ def __init__(self, transform, media_list): for medium in media_list[DATA_KEY] for solution in transform.media_detailed[str(medium[ID_COLUMN])].get(SOLUTIONS_KEY, []) } - relevant = False + self.names = {"potato": set(), "p3556": set(), "sugar": set(), "peptone": set()} for identifier in solutions: recipe = transform.solutions_data[identifier].get(RECIPE_KEY) # Match the existing producer's absent/non-list recipe behavior. if isinstance(recipe, list): - relevant |= any( - _potato(item.get(COMPOUND_KEY, item.get(SOLUTION_KEY))) for item in recipe if isinstance(item, dict) - ) - self.active = relevant - if not relevant: + for item in recipe: + if isinstance(item, dict) and (profile := _profile(item)): + name = item.get(COMPOUND_KEY, item.get(SOLUTION_KEY)) + self.names[profile].add(name if isinstance(name, str) else "") + self.active = any(self.names.values()) + if not self.active: return # Bind the exact current policy even if the mapping reader has already # excluded the row. Such rows are candidates, not attempted lookups. - with self._consume(POLICY_ROLE, IDENTITY_POLICY) as stream: - policies = list(_table_rows(stream, {"target_id", "authority_label", "kind", "value", "reason"})) - if not any( - _policy_target_key(row["target_id"]) == TARGET - and row["authority_label"] == AUTHORITY_LABEL - and row["kind"] == "name_pattern" - and row["value"] == "(?i)^potato$" - for _, row in policies - ) or ingredient_mapping_allowed("Potato", TARGET): + policies = [] + if self.names["potato"] or self.names["p3556"] or self.names["peptone"]: + with self._consume(POLICY_ROLE, IDENTITY_POLICY) as stream: + policies = list(_table_rows(stream, {"target_id", "authority_label", "kind", "value", "reason"})) + if self.names["sugar"]: + with self._consume(CONTEXT_ROLE, _repo_root() / CONTEXT_POLICY) as stream: + decisions = list(_table_rows(stream, CONTEXT_FIELDS)) + if len(decisions) != 1: + raise SourceFinalizationRequired("Reviewed Sugar context requires one exact decision") + self.sugar_decision = decisions[0][1] + if set(self.sugar_decision) != CONTEXT_FIELDS or self.sugar_decision != { + "source_name": "Sugar", + "source_attribute": SUGAR_ATTRIBUTE, + "disposition": "retain_existing_local", + "reason": SUGAR_REASON, + "evidence_uri": SUGAR_URI, + "authority_label": "JCM 537 mixed xylose, maltose and cellobiose stock solution", + "withheld_target_example": "NCIT:C71939", + }: + raise SourceFinalizationRequired("Reviewed Sugar context decision changed or is missing") + if self.names["potato"] and ( + not any( + _policy_target_key(row["target_id"]) == TARGET + and row["authority_label"] == AUTHORITY_LABEL + and row["kind"] == "name_pattern" + and row["value"] == "(?i)^potato$" + for _, row in policies + ) + or ingredient_mapping_allowed("Potato", TARGET) + ): raise SourceFinalizationRequired("Reviewed Potato material-scope identity hold is missing") + if self.names["p3556"] and ( + ingredient_mapping_allowed("L-alpha-Phosphatidylcholine", P3556_TARGET) + or ingredient_xref_allowed(P3556_MIM, P3556_TARGET) + or not all( + any( + row["target_id"] == P3556_TARGET and row["kind"] == kind and row["value"] == value + for _, row in policies + ) + for kind, value in ( + ("name_pattern", "(?i)^l[ _-]*(?:alpha|α)[ _-]*phosphatidylcholine$"), + ("xref", P3556_MIM), + ) + ) + ): + raise SourceFinalizationRequired("Reviewed P3556 material-scope identity hold is missing") + if self.names["peptone"] and not all( + not ingredient_mapping_allowed(name, PEPTONE_TARGET) + and any( + _policy_target_key(row["target_id"]) == PEPTONE_TARGET + and row["authority_label"] == PEPTONE_AUTHORITY_LABEL + and row["kind"] == "name_pattern" + and row["value"] == pattern + for _, row in policies + ) + for name, pattern in PEPTONE_PATTERNS.items() + ): + raise SourceFinalizationRequired("Reviewed peptone material-scope identity holds are missing") unified = _repo_root() / transform.DATA_INPUTS[0] with self._consume(UNIFIED_ROLE, unified) as snapshot: with gzip.GzipFile(fileobj=snapshot.buffer) as compressed: with io.TextIOWrapper(compressed, encoding="utf-8", newline="") as stream: for locator, row in _table_rows(stream, {"subject_id", "subject_label", "object_id"}): - if ( - row["subject_id"].startswith("kgm.name:") - and _potato(row["subject_label"]) - and _policy_target_key(row["object_id"]) == TARGET - and classify_mapping_row(row)[0] in {"canonical_name", "synonym"} - ): - self._claim(UNIFIED_ROLE, locator, row["object_id"], row) + self._mapping_claims(UNIFIED_ROLE, locator, row) + if self.names["p3556"] or self.names["sugar"]: + # Original imported evidence remains available after the derived + # unified row is removed. This never feeds an identity lookup. + with self._consume(SUPPORTED_ROLE, _repo_root() / "mappings/ingredient_mappings.sssom.tsv") as stream: + for locator, row in _table_rows(stream, {"subject_id", "subject_label", "object_id"}): + self._mapping_claims(SUPPORTED_ROLE, locator, row) for role in ("micromediaparam_hydrate", "micromediaparam_strict"): # The already selected optional-input contract binds absence too. with transform.consume_optional_input(role) as stream: if stream is None: continue for locator, row in _table_rows(stream, {"original", "mapped"}): - if _potato(row["original"]) and _policy_target_key(row["mapped"]) == TARGET: - self._claim(role, locator, row["mapped"], row) + if ( + self.names["potato"] + and _potato(row["original"]) + and _policy_target_key(row["mapped"]) == TARGET + ): + self._claim("potato", None, role, locator, row["mapped"], row) + for profile in ("p3556", "sugar"): + for name in self.names[profile]: + if name and name.lower().strip() == row["original"].lower().strip(): + self._claim(profile, name, role, locator, row["mapped"], row) + if _policy_target_key(row["mapped"]) == PEPTONE_TARGET: + for name in self.names["peptone"]: + if name.lower().strip() == row["original"].lower().strip(): + self._claim("peptone", name, role, locator, row["mapped"], row) self.guard.verify() + def _mapping_claims(self, role, locator, row): + """Select only reader-eligible lexical or exact imported claims for the actual spelling.""" + route = classify_mapping_row(row)[0] + if route in {"canonical_name", "synonym"}: + if ( + self.names["potato"] + and _potato(row["subject_label"]) + and _policy_target_key(row["object_id"]) == TARGET + ): + self._claim("potato", None, role, locator, row["object_id"], row) + # All four recognized primary routes can supply canonical object + # metadata; only synonym rows additionally index subject_label (#1243). + # Preserve eligible original claims even when the identity policy now + # rejects that label. These are available candidates, not lookup calls. + if route not in {"identity", "attribute", "canonical_name", "synonym"}: + return + if _policy_target_key(row["object_id"]) == PEPTONE_TARGET: + for name in self.names["peptone"]: + if normalize_name(name) == normalize_name(row.get("object_label", "")) or ( + route == "synonym" and normalize_name(name) == normalize_name(row["subject_label"]) + ): + self._claim("peptone", name, role, locator, row["object_id"], row) + for profile, subject, target in ( + ("p3556", P3556_MIM, P3556_TARGET), + ("sugar", "MIM:Sugar", "NCIT:C71939"), + ): + if ( + self.names[profile] + and route == "identity" + and row["subject_id"] == subject + and row["object_id"] == target + ): + self._claim(profile, None, role, locator, row["object_id"], row) + continue + for name in self.names[profile]: + if name and ( + normalize_name(name) == normalize_name(row.get("object_label", "")) + or (route == "synonym" and normalize_name(name) == normalize_name(row["subject_label"])) + ): + self._claim(profile, name, role, locator, row["object_id"], row) + def _consume(self, role, path): """Guard the original lexical locator in addition to the immutable parser snapshot.""" self.guard.bind_path(path) self.guard.capture(path) return self.transform.consume_input(role, path) - def _claim(self, role, locator, target, record): + def _claim(self, profile, name, role, locator, target, record): """Retain distinct row ordinals even when the complete mapping payload is duplicated.""" snapshot = self.transform.consumed_input_snapshots[role] - self.claims.append((role, snapshot, locator, target, _json(record))) + self.claims.append((profile, name, (role, snapshot, locator, target, _json(record)))) def observe(self, occurrence, raw): """Attach available candidates to an actual emitted occurrence and its selected target.""" - if not self.active or not _potato(raw.get(COMPOUND_KEY, raw.get(SOLUTION_KEY))): + profile = _profile(raw) + if not self.active or not profile: return - if _policy_target_key(occurrence[ID_COLUMN]) == TARGET: + if profile == "potato" and _policy_target_key(occurrence[ID_COLUMN]) == TARGET: raise SourceFinalizationRequired("Held Potato extract identity escaped source resolution") - candidates = list(self.claims) + if profile == "peptone" and _policy_target_key(occurrence[ID_COLUMN]) == PEPTONE_TARGET: + raise SourceFinalizationRequired("Held peptone structural identity escaped source resolution") + if profile in {"p3556", "sugar"}: + local = ( + MEDIADIVE_INGREDIENT_PREFIX + str(raw[COMPOUND_ID_KEY]) + if raw.get(COMPOUND_ID_KEY) is not None + else MEDIADIVE_SOLUTION_PREFIX + str(raw[SOLUTION_ID_KEY]) + ) + if occurrence[ID_COLUMN] != local: + raise SourceFinalizationRequired("Held source material identity escaped source resolution") + name = raw.get(COMPOUND_KEY, raw.get(SOLUTION_KEY)) + candidates = [claim for group, spelling, claim in self.claims if group == profile and spelling in (None, name)] identifier = raw.get(COMPOUND_ID_KEY) if identifier is not None: embedded = self.transform.compounds_data.get(str(identifier), {}) - value = embedded.get(CAS_RN_KEY) - if value is not None and _policy_target_key(CAS_RN_PREFIX + str(value)) == TARGET: + for field, prefix in ( + (CHEBI_KEY, CHEBI_PREFIX), + (KEGG_KEY, KEGG_PREFIX), + (PUBCHEM_KEY, PUBCHEM_PREFIX), + (CAS_RN_KEY, CAS_RN_PREFIX), + ): + value = embedded.get(field) + target = prefix + str(value) + if value is None or (profile == "potato" and _policy_target_key(target) != TARGET): + continue + if profile == "peptone" and _policy_target_key(target) != PEPTONE_TARGET: + continue candidates.append( ( "mediadive_compounds", self.transform.consumed_input_snapshots["mediadive_compounds"], - f"key={_json(str(identifier))};field={CAS_RN_KEY}", - CAS_RN_PREFIX + str(value), + f"key={_json(str(identifier))};field={field}", + target, _json(embedded), ) ) source = self.transform.consumed_input_snapshots["mediadive_solutions"] - policy = self.transform.consumed_input_snapshots[POLICY_ROLE] - reason, qualifiers = _reason(raw) + policy_role = CONTEXT_ROLE if profile == "sugar" else POLICY_ROLE + policy = self.transform.consumed_input_snapshots[policy_role] + reason, qualifiers = ( + ("whole_product_not_molecular_identity", _json({"attribute": raw["attribute"]})) + if profile == "p3556" + else _reason(raw) + ) + if profile == "sugar": + reason, qualifiers = self.sugar_decision["reason"], _json({"attribute": raw["attribute"]}) + if profile == "peptone": + reason = PEPTONE_REASON + authority_label, authority_uri = AUTHORITY_LABEL, AUTHORITY_URI + if profile == "p3556": + authority_label, authority_uri = P3556_AUTHORITY_LABEL, P3556_AUTHORITY_URI + elif profile == "peptone": + authority_label, authority_uri = PEPTONE_AUTHORITY_LABEL, PEPTONE_AUTHORITY_URI + elif profile == "sugar": + authority_label, authority_uri = self.sugar_decision["authority_label"], self.sugar_decision["evidence_uri"] payload = occurrence[SOURCE_RECORD_COLUMN] for route, snapshot, locator, target, candidate in candidates: row = ( @@ -220,8 +437,8 @@ def observe(self, occurrence, raw): candidate, policy["path"], policy["sha256"], - AUTHORITY_LABEL, - AUTHORITY_URI, + authority_label, + authority_uri, ) key = (occurrence[SOURCE_ASSERTION_ID_COLUMN], route, locator) if self.rows.setdefault(key, row) != row: @@ -246,18 +463,38 @@ def write(self): def verify_recorded_material_inputs(report, report_path): """Require actual canonical evidence origins for the producer-bound candidate sidecar.""" snapshots = report.get("consumed_inputs", {}) + context_path = str((_repo_root() / CONTEXT_POLICY).resolve()) + context_recorded = any(item.get("path") == context_path for item in report.get("inputs", ())) path = Path(report_path).parent / AUDIT_FILENAME with path.open(encoding="utf-8", newline="") as stream: reader = csv.DictReader(stream, delimiter="\t", quoting=csv.QUOTE_NONE) if tuple(reader.fieldnames or ()) != AUDIT_HEADER: raise SourceFinalizationRequired("Invalid material-scope producer audit header") - has_rows = next(reader, None) is not None - if not has_rows and not any(role in snapshots for role in (UNIFIED_ROLE, POLICY_ROLE)): + has_rows = False + needs_supported = False + needs_context = False + needs_identity = False + for row in reader: + has_rows = True + sugar = row.get("reason") == SUGAR_REASON + needs_context |= sugar + needs_identity |= not sugar + needs_supported |= sugar or row.get("reason") == "whole_product_not_molecular_identity" + if ( + not has_rows + and not context_recorded + and not any(role in snapshots for role in (UNIFIED_ROLE, POLICY_ROLE, SUPPORTED_ROLE, CONTEXT_ROLE)) + ): return expected = { UNIFIED_ROLE: _repo_root() / "mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz", - POLICY_ROLE: IDENTITY_POLICY, } + if needs_identity or POLICY_ROLE in snapshots: + expected[POLICY_ROLE] = IDENTITY_POLICY + if needs_context or context_recorded or CONTEXT_ROLE in snapshots: + expected[CONTEXT_ROLE] = _repo_root() / CONTEXT_POLICY + if needs_supported or needs_context or context_recorded or CONTEXT_ROLE in snapshots or SUPPORTED_ROLE in snapshots: + expected[SUPPORTED_ROLE] = _repo_root() / "mappings/ingredient_mappings.sssom.tsv" for role, origin in expected.items(): if snapshots.get(role, {}).get("path") != str(origin.resolve()): raise SourceFinalizationRequired(f"Missing or wrong material-scope input origin: {role}") diff --git a/kg_microbe/transform_utils/mediadive/mediadive.py b/kg_microbe/transform_utils/mediadive/mediadive.py index 07061a5a..b0699f0a 100644 --- a/kg_microbe/transform_utils/mediadive/mediadive.py +++ b/kg_microbe/transform_utils/mediadive/mediadive.py @@ -138,6 +138,8 @@ from kg_microbe.transform_utils.mediadive.material_scope_audit import ( AUDIT_FILENAME, MaterialScopeAudit, + p3556_material, + sugar_material, verify_recorded_material_inputs, ) from kg_microbe.transform_utils.transform import Transform @@ -762,7 +764,9 @@ def get_solution_recipe_occurrences(self, id: str): if isinstance(item[COMPOUND_KEY], str) else item[COMPOUND_KEY] ) - ingredient_id = self.standardize_compound_id(str(item[COMPOUND_ID_KEY]), source_name) + ingredient_id = self.standardize_compound_id( + str(item[COMPOUND_ID_KEY]), source_name, source_record=item + ) elif SOLUTION_ID_KEY in item and item[SOLUTION_ID_KEY] is not None: # Resolve the original scope; normalize only the display key. if isinstance(item[SOLUTION_KEY], str): @@ -776,10 +780,14 @@ def get_solution_recipe_occurrences(self, id: str): solution_name_normalized = source_name.lower().strip() # Check if solution name can be mapped to ontology via unified or legacy mappings - candidates = [ - self.chemical_loader.find_chebi_by_name(source_name), - self.compound_mappings.get(solution_name_normalized), - ] + candidates = ( + [] + if p3556_material(item) or sugar_material(item) + else [ + self.chemical_loader.find_chebi_by_name(source_name), + self.compound_mappings.get(solution_name_normalized), + ] + ) solution_id = next( ( candidate @@ -823,7 +831,7 @@ def _ingredient_identity_allowed(self, name: str, target: str) -> bool: label = getter(target) if getter else "" return ingredient_hydration_compatible(name, label, getattr(self, "chebi_hydrate_names", {}).get(target, ())) - def standardize_compound_id(self, id: str, compound_name: str = None): + def standardize_compound_id(self, id: str, compound_name: str = None, *, source_record: dict = None): """ Get standardized IDs via unified chemical mappings, bulk data, or legacy mappings. @@ -835,8 +843,14 @@ def standardize_compound_id(self, id: str, compound_name: str = None): :param id: MediaDive compound ID. :param compound_name: Compound name for mapping lookup. + :param source_record: Actual occurrence evidence; even an empty dictionary overrides embedded qualifiers. :return: Standardized ID. """ + # A whole supplier product is not a pure molecular species. The actual + # occurrence wins over another recipe's cached embedded record (#1241). + evidence = source_record if source_record is not None else getattr(self, "compounds_data", {}).get(id, {}) + if p3556_material(evidence) or sugar_material(evidence, compound_name): + return MEDIADIVE_INGREDIENT_PREFIX + id if compound_name: # Check unified chemical mappings by compound name. The unified # mapping is not restricted to CHEBI — it also holds FOODON diff --git a/mappings/canonical/bacdive_ec_substrate_corrections.tsv b/mappings/canonical/bacdive_ec_substrate_corrections.tsv new file mode 100644 index 00000000..e9f114a2 --- /dev/null +++ b/mappings/canonical/bacdive_ec_substrate_corrections.tsv @@ -0,0 +1,5 @@ +CHEBI_ID substrate KEGG_ID CAS_RN_ID EC_ID enzyme pseudo_CURIE reaction_name correction_id corrected_chebi_id withheld_source_fields reason evidence_uri +CHEBI:54684 4-Nitrophenyl alpha-D-glucopyranoside CAS-RN:3767-28-0 EC:3.2.1.20 alpha-glucosidase kgmicrobe.assay:API_ID32E_alpha GLU alpha-Glucosidase bacdive-glucoside-1249 CHEBI:91122 Exact native IUPAC synonym and supplier stereostructure support glucoside, not galactoside. https://www.ebi.ac.uk/chebi/searchId.do?chebiId=CHEBI:91122|https://www.sigmaaldrich.com/US/en/product/mm/487506|https://iubmb.qmul.ac.uk/enzyme/EC3/2/1/20.html +CHEBI:54684 4-Nitrophenyl-alpha-D-galactoside CAS-RN:7493-95-0 EC:3.2.1.22 alpha-galactosidase kgmicrobe.assay:API_rID32A_alpha GAL alpha-Galactosidase bacdive-galactoside-rid32a-1250 CHEBI:546840 Native exact name and supplier stereostructure support galactoside; no old-ID replacement is inferred. https://www.sigmaaldrich.com/US/en/product/sigma/n0877|https://iubmb.qmul.ac.uk/enzyme/EC3/2/1/22.html +CHEBI:54684 4-Nitrophenyl-alpha-D-galactoside KEGG:C01083 CAS-RN:7493-95-0 EC:3.2.1.22 alpha-galactosidase kgmicrobe.assay:API_ID32E_alpha GAL alpha-Galactosidase bacdive-galactoside-id32e-1250 CHEBI:546840 KEGG_ID Native and supplier structure support galactoside; original KEGG:C01083 denotes trehalose and is audit-only. https://www.sigmaaldrich.com/US/en/product/sigma/n0877|https://iubmb.qmul.ac.uk/enzyme/EC3/2/1/22.html|https://www.kegg.jp/entry/C01083 +CHEBI:54684 4-Nitrophenyl-alpha-D-galactoside CAS-RN:7493-95-0 EC:3.2.1.22 alpha-galactosidase kgmicrobe.assay:API_rID32STR_alpha GAL alpha-Galactosidase bacdive-galactoside-rid32str-1250 CHEBI:546840 Native exact name and supplier stereostructure support galactoside; no old-ID replacement is inferred. https://www.sigmaaldrich.com/US/en/product/sigma/n0877|https://iubmb.qmul.ac.uk/enzyme/EC3/2/1/22.html diff --git a/mappings/canonical/mediadive_material_context_dispositions.tsv b/mappings/canonical/mediadive_material_context_dispositions.tsv new file mode 100644 index 00000000..ab7ef992 --- /dev/null +++ b/mappings/canonical/mediadive_material_context_dispositions.tsv @@ -0,0 +1,2 @@ +source_name source_attribute disposition reason evidence_uri authority_label withheld_target_example +Sugar 250 mM each of xylose, maltose and cellobiose retain_existing_local qualified_sugar_solution_not_demonstrated_food_identity https://www.jcm.riken.jp/cgi-bin/jcm/jcm_grmd?GRMD=537 JCM 537 mixed xylose, maltose and cellobiose stock solution NCIT:C71939 diff --git a/mappings/ingredient_identity_exclusions.tsv b/mappings/ingredient_identity_exclusions.tsv index cb6b2c38..51d6edb4 100644 --- a/mappings/ingredient_identity_exclusions.tsv +++ b/mappings/ingredient_identity_exclusions.tsv @@ -1,4 +1,8 @@ target_id authority_label kind value reason +pubchem.compound:167312541 Glycerides, C8-10 mono-and di- name_pattern (?i)^soy[ _-]+peptone$ The reviewed soy digest does not establish the specific multicomponent PubChem structure, despite genuine soybean-peptone annotations on that record. Preserve source-local material observations and imported mapping claims; do not ban direct CID references or infer a replacement registry identity. See docs/reviews/mediadive-peptone-scope-1248.md (#1248). +pubchem.compound:167312541 Glycerides, C8-10 mono-and di- name_pattern (?i)^vitamin[ _-]+free[ _-]+casamino[ _-]+acids$ The reviewed vitamin-free acid casein digest does not establish the specific multicomponent PubChem structure or identity with the separate soy digest. Preserve source-local material observations and complete imported claims without inferring a replacement CAS. See docs/reviews/mediadive-peptone-scope-1248.md (#1248). +CHEBI:86658 [(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxy-propyl] 2-(trimethylammonio)ethyl phosphate name_pattern (?i)^l[ _-]*(?:alpha|α)[ _-]*phosphatidylcholine$ The generic source name does not establish the native specific 16:0/(9E,12E)-18:2 molecular structure. Preserve explicit structured names and source-qualified Sigma P3556 whole-product observations separately; do not infer another charge, CAS or ChEBI identity. See docs/reviews/mediadive-p3556-scope-1241.md (#1241). +CHEBI:86658 [(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxy-propyl] 2-(trimethylammonio)ethyl phosphate xref MIM:L-alpha-Phosphatidylcholine The generic reviewed MIM ingredient identifier does not establish the specific native acyl/stereochemical structure. Retain the immutable imported claim as evidence without admitting it as an exact identity. See docs/reviews/mediadive-p3556-scope-1241.md (#1241). cas:93348-51-7 zemiak, Solanum tuberosum aegrotans, extrakt name_pattern (?i)^potato$ The reviewed bare MediaDive Potato ingredient does not establish the EINECS EC 297-194-4 extract identity; six saved occurrences specify fresh/peeled/cut material and three are unqualified. Preserve observations and imported mapping claims separately, without guessing a replacement or banning the registry identifier. See docs/reviews/mediadive-potato-scope-1236.md (#1236). CHEBI:44356 N-tris(hydroxymethyl)methyl-2-aminoethanesulfonic acid name_pattern (?i)^tris(?:\(hydroxymethyl\)|hydroxymethyl)methylamine$ The observed Tris source is not the native TES sulfonic-acid derivative; preserve source identity without guessing a replacement from a legacy substring match. Source 331 quantities corroborate Tris molecular mass (#1169). CHEBI:63037 triammonium citrate name_pattern (?i)^(?:\(nh4\)|nh4)[ _-]*citrate$ The observed unqualified ammonium citrate source lacks counterion stoichiometry, external IDs and molecular-amount evidence; do not select the native triammonium salt. Explicit diammonium/triammonium names and formulas remain valid (#1169). diff --git a/mappings/ingredient_name_scopes.tsv b/mappings/ingredient_name_scopes.tsv index c8c3171c..6fcb2b50 100644 --- a/mappings/ingredient_name_scopes.tsv +++ b/mappings/ingredient_name_scopes.tsv @@ -8,6 +8,7 @@ name Rifamycin SV CHEBI:29673 Explicit SV retains the specific chemical identity name Xanthine CHEBI:15318 Bare name has generic scope; an explicit MIM subject, 9H label or reviewed CAS selects CHEBI:17712. name 9H-xanthine CHEBI:17712 Explicit 9H representation; preserve its native subclass relationship to CHEBI:15318. name Sorbitan Monooleate kgmicrobe.ingredient:sorbitan_monooleate Unspecified recipe material; exact molecular identity and CAS remain withheld pending product evidence. +name Tetramethyl ammonium chloride CHEBI:7070 Finite reviewed chloride-salt spelling; native ChEBI and MediaDive CAS 75-57-0 agree. Unqualified parent and other counterions remain separate. See reviews/tetramethylammonium-chloride-20260929.md (#1242). cas cas:1404-26-8 NCIT:C61894 Verified NCIT P210 annotation for the B1/B2 mixture; no equivalence to B1 or sulfate. cas cas:4135-11-9 CHEBI:8309 Verified ChEBI CAS annotation for B1 only. cas cas:6998-60-3 CHEBI:29673 Verified ChEBI CAS annotation for SV only; the rifamycin family has no single CAS. diff --git a/mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz b/mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz index f145effc..bdef18f9 100644 Binary files a/mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz and b/mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz differ diff --git a/tests/resources/bacdive_ec_substrate_corrections/authority.json b/tests/resources/bacdive_ec_substrate_corrections/authority.json new file mode 100644 index 00000000..565d10b2 --- /dev/null +++ b/tests/resources/bacdive_ec_substrate_corrections/authority.json @@ -0,0 +1,577 @@ +{ + "review_date": "2026-09-29", + "native_raw_sha256": "a5d40380ab78bde0e8b5a704dbee3cba2bcaa7608be46eebc2f4a92932516c9b", + "native_tsv_sha256": "c161f8aea7ffeb1a2557f24100b3b3b67bfd62196d484b33c675921154c8b152", + "legacy_mapping_sha256": "f39e9753fe2879877b2cba51c9896f44515f3a5785e790deb0d0bd5ca02c9379", + "complete_native_scan": { + "nodes": 218301, + "edges": 417074 + }, + "raw_selected_nodes_or_exact_identifier_references": [ + { + "ordinal": 160861, + "node": { + "id": "http://purl.obolibrary.org/obo/CHEBI_546840", + "lbl": "4-nitrophenyl alpha-D-galactoside", + "type": "CLASS", + "meta": { + "definition": { + "val": "An α-D-galactopyranoside having a 4-nitrophenyl substituent at the anomeric position." + }, + "subsets": [ + "http://purl.obolibrary.org/obo/chebi/3_STAR" + ], + "synonyms": [ + { + "synonymType": "http://purl.obolibrary.org/obo/chebi/IUPAC_NAME", + "pred": "hasExactSynonym", + "val": "4-nitrophenyl-alpha-D-galactopyranoside", + "xrefs": [ + "IUPAC" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "1-O-(4-nitrophenyl)-alpha-D-galactopyranose", + "xrefs": [ + "chebi" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "1-O-(4-nitrophenyl)-alpha-D-galactose", + "xrefs": [ + "chebi" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "1-O-(p-nitrophenyl)-alpha-D-galactopyranose", + "xrefs": [ + "chebi" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "1-O-(p-nitrophenyl)-alpha-D-galactose", + "xrefs": [ + "chebi" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "4-nitrophenyl-alpha-D-galactopyranoside", + "xrefs": [ + "chebi" + ] + }, + { + "synonymType": "http://purl.obolibrary.org/obo/chebi/IUPAC_NAME", + "pred": "hasRelatedSynonym", + "val": "4-nitrophenyl-alpha-D-galactopyranoside", + "xrefs": [ + "iupac" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "4-nitrophenyl-alpha-D-galactoside", + "xrefs": [ + "chebi" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "p-Nitrophenyl alpha-D-galactopyranoside", + "xrefs": [ + "chemidplus" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "p-nitrophenyl alpha-D-galactoside", + "xrefs": [ + "chebi" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "p-nitrophenyl-alpha-D-galactopyranoside", + "xrefs": [ + "chebi" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "p-nitrophenyl-alpha-D-galactoside", + "xrefs": [ + "chebi" + ] + } + ], + "xrefs": [ + { + "val": "cas:7493-95-0" + }, + { + "val": "pubmed:19913595" + }, + { + "val": "reaxys:92214" + } + ], + "basicPropertyValues": [ + { + "pred": "http://www.geneontology.org/formats/oboInOwl#hasOBONamespace", + "val": "chebi_ontology" + }, + { + "pred": "https://w3id.org/chemrof/charge", + "val": "0" + }, + { + "pred": "https://w3id.org/chemrof/generalized_empirical_formula", + "val": "C12H15NO8" + }, + { + "pred": "https://w3id.org/chemrof/inchi_key_string", + "val": "IFBHRQDFSNCLOZ-IIRVCBMXSA-N" + }, + { + "pred": "https://w3id.org/chemrof/inchi_string", + "val": "InChI=1S/C12H15NO8/c14-5-8-9(15)10(16)11(17)12(21-8)20-7-3-1-6(2-4-7)13(18)19/h1-4,8-12,14-17H,5H2/t8-,9+,10+,11-,12+/m1/s1" + }, + { + "pred": "https://w3id.org/chemrof/mass", + "val": "301.251" + }, + { + "pred": "https://w3id.org/chemrof/monoisotopic_mass", + "val": "301.07977" + }, + { + "pred": "https://w3id.org/chemrof/smiles_string", + "val": "O=[N+]([O-])c1ccc(O[C@H]2O[C@H](CO)[C@H](O)[C@H](O)[C@H]2O)cc1" + } + ] + } + } + }, + { + "ordinal": 208951, + "node": { + "id": "http://purl.obolibrary.org/obo/CHEBI_91122", + "lbl": "4-nitrophenyl alpha-D-glucoside", + "type": "CLASS", + "meta": { + "definition": { + "val": "An α-D-glucoside that is β-D-glucopyranose in which the anomeric hydroxy hydrogen is replaced by a 4-nitrophenyl group." + }, + "subsets": [ + "http://purl.obolibrary.org/obo/chebi/3_STAR" + ], + "synonyms": [ + { + "synonymType": "http://purl.obolibrary.org/obo/chebi/IUPAC_NAME", + "pred": "hasExactSynonym", + "val": "4-nitrophenyl alpha-D-glucopyranoside", + "xrefs": [ + "IUPAC" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "4-Nitrophenyl alpha-glucoside", + "xrefs": [ + "chemidplus" + ] + }, + { + "synonymType": "http://purl.obolibrary.org/obo/chebi/IUPAC_NAME", + "pred": "hasRelatedSynonym", + "val": "4-nitrophenyl alpha-D-glucopyranoside", + "xrefs": [ + "iupac" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "p-nitrophenyl alpha-D-glucoside", + "xrefs": [ + "chebi" + ] + } + ], + "xrefs": [ + { + "val": "cas:3767-28-0" + }, + { + "val": "pubmed:16495668" + }, + { + "val": "reaxys:92212" + } + ], + "basicPropertyValues": [ + { + "pred": "http://www.geneontology.org/formats/oboInOwl#hasOBONamespace", + "val": "chebi_ontology" + }, + { + "pred": "https://w3id.org/chemrof/charge", + "val": "0" + }, + { + "pred": "https://w3id.org/chemrof/generalized_empirical_formula", + "val": "C12H15NO8" + }, + { + "pred": "https://w3id.org/chemrof/inchi_key_string", + "val": "IFBHRQDFSNCLOZ-ZIQFBCGOSA-N" + }, + { + "pred": "https://w3id.org/chemrof/inchi_string", + "val": "InChI=1S/C12H15NO8/c14-5-8-9(15)10(16)11(17)12(21-8)20-7-3-1-6(2-4-7)13(18)19/h1-4,8-12,14-17H,5H2/t8-,9-,10+,11-,12+/m1/s1" + }, + { + "pred": "https://w3id.org/chemrof/mass", + "val": "301.251" + }, + { + "pred": "https://w3id.org/chemrof/monoisotopic_mass", + "val": "301.07977" + }, + { + "pred": "https://w3id.org/chemrof/smiles_string", + "val": "O=[N+]([O-])c1ccc(O[C@H]2O[C@H](CO)[C@@H](O)[C@H](O)[C@H]2O)cc1" + } + ] + } + } + } + ], + "raw_incident_edges": [ + { + "ordinal": 274660, + "edge": { + "sub": "http://purl.obolibrary.org/obo/CHEBI_546840", + "pred": "is_a", + "obj": "http://purl.obolibrary.org/obo/CHEBI_46953" + } + }, + { + "ordinal": 274661, + "edge": { + "sub": "http://purl.obolibrary.org/obo/CHEBI_546840", + "pred": "is_a", + "obj": "http://purl.obolibrary.org/obo/CHEBI_63367" + } + }, + { + "ordinal": 401927, + "edge": { + "sub": "http://purl.obolibrary.org/obo/CHEBI_91122", + "pred": "is_a", + "obj": "http://purl.obolibrary.org/obo/CHEBI_22390" + } + }, + { + "ordinal": 401928, + "edge": { + "sub": "http://purl.obolibrary.org/obo/CHEBI_91122", + "pred": "is_a", + "obj": "http://purl.obolibrary.org/obo/CHEBI_35716" + } + }, + { + "ordinal": 401929, + "edge": { + "sub": "http://purl.obolibrary.org/obo/CHEBI_91122", + "pred": "is_a", + "obj": "http://purl.obolibrary.org/obo/CHEBI_63367" + } + }, + { + "ordinal": 401930, + "edge": { + "sub": "http://purl.obolibrary.org/obo/CHEBI_91122", + "pred": "http://purl.obolibrary.org/obo/RO_0000087", + "obj": "http://purl.obolibrary.org/obo/CHEBI_75050" + } + }, + { + "ordinal": 401931, + "edge": { + "sub": "http://purl.obolibrary.org/obo/CHEBI_91122", + "pred": "http://purl.obolibrary.org/obo/RO_0018038", + "obj": "http://purl.obolibrary.org/obo/CHEBI_16836" + } + } + ], + "native_chebi_rows": { + "header": [ + "id", + "category", + "name", + "description", + "xref", + "provided_by", + "synonym", + "deprecated", + "same_as" + ], + "selected_rows": [ + { + "line_number": 160862, + "original_line": "CHEBI:546840\tbiolink:ChemicalEntity\t4-nitrophenyl alpha-D-galactoside\tAn α-D-galactopyranoside having a 4-nitrophenyl substituent at the anomeric position.\tcas:7493-95-0|pubmed:19913595|reaxys:92214\tchebi.json\t1-O-(4-nitrophenyl)-alpha-D-galactopyranose|1-O-(4-nitrophenyl)-alpha-D-galactose|1-O-(p-nitrophenyl)-alpha-D-galactopyranose|1-O-(p-nitrophenyl)-alpha-D-galactose|4-nitrophenyl-alpha-D-galactopyranoside|4-nitrophenyl-alpha-D-galactoside|p-Nitrophenyl alpha-D-galactopyranoside|p-nitrophenyl alpha-D-galactoside|p-nitrophenyl-alpha-D-galactopyranoside|p-nitrophenyl-alpha-D-galactoside\t\t\n", + "row": { + "id": "CHEBI:546840", + "category": "biolink:ChemicalEntity", + "name": "4-nitrophenyl alpha-D-galactoside", + "description": "An α-D-galactopyranoside having a 4-nitrophenyl substituent at the anomeric position.", + "xref": "cas:7493-95-0|pubmed:19913595|reaxys:92214", + "provided_by": "chebi.json", + "synonym": "1-O-(4-nitrophenyl)-alpha-D-galactopyranose|1-O-(4-nitrophenyl)-alpha-D-galactose|1-O-(p-nitrophenyl)-alpha-D-galactopyranose|1-O-(p-nitrophenyl)-alpha-D-galactose|4-nitrophenyl-alpha-D-galactopyranoside|4-nitrophenyl-alpha-D-galactoside|p-Nitrophenyl alpha-D-galactopyranoside|p-nitrophenyl alpha-D-galactoside|p-nitrophenyl-alpha-D-galactopyranoside|p-nitrophenyl-alpha-D-galactoside", + "deprecated": "", + "same_as": "" + } + }, + { + "line_number": 208952, + "original_line": "CHEBI:91122\tbiolink:ChemicalEntity\t4-nitrophenyl alpha-D-glucoside\tAn α-D-glucoside that is β-D-glucopyranose in which the anomeric hydroxy hydrogen is replaced by a 4-nitrophenyl group.\tcas:3767-28-0|pubmed:16495668|reaxys:92212\tchebi.json\t4-Nitrophenyl alpha-glucoside|4-nitrophenyl alpha-D-glucopyranoside|p-nitrophenyl alpha-D-glucoside\t\t\n", + "row": { + "id": "CHEBI:91122", + "category": "biolink:ChemicalEntity", + "name": "4-nitrophenyl alpha-D-glucoside", + "description": "An α-D-glucoside that is β-D-glucopyranose in which the anomeric hydroxy hydrogen is replaced by a 4-nitrophenyl group.", + "xref": "cas:3767-28-0|pubmed:16495668|reaxys:92212", + "provided_by": "chebi.json", + "synonym": "4-Nitrophenyl alpha-glucoside|4-nitrophenyl alpha-D-glucopyranoside|p-nitrophenyl alpha-D-glucoside", + "deprecated": "", + "same_as": "" + } + } + ] + }, + "native_ec_rows": { + "header": [ + "id", + "category", + "name", + "description", + "xref", + "provided_by", + "synonym", + "deprecated", + "same_as" + ], + "selected_rows": [ + { + "line_number": 9376, + "original_line": "EC:3.2.1.20\tbiolink:MolecularActivity|biolink:Protein\talpha-glucosidase\t\t\tec.json\tacid maltase|glucoinvertase|glucosidosucrase|lysosomal alpha-glucosidase|maltase|maltase-glucoamylase\t\t\n", + "row": { + "id": "EC:3.2.1.20", + "category": "biolink:MolecularActivity|biolink:Protein", + "name": "alpha-glucosidase", + "description": "", + "xref": "", + "provided_by": "ec.json", + "synonym": "acid maltase|glucoinvertase|glucosidosucrase|lysosomal alpha-glucosidase|maltase|maltase-glucoamylase", + "deprecated": "", + "same_as": "" + } + } + ] + }, + "legacy_rows": { + "header": [ + "CHEBI_ID", + "substrate", + "KEGG_ID", + "CAS_RN_ID", + "EC_ID", + "enzyme", + "pseudo_CURIE", + "reaction_name" + ], + "selected_rows": [ + { + "line_number": 28, + "original_line": "CHEBI:54684\t4-Nitrophenyl-alpha-D-galactoside\t\tCAS-RN:7493-95-0\tEC:3.2.1.22\talpha-galactosidase\tkgmicrobe.assay:API_rID32A_alpha GAL\talpha-Galactosidase\n", + "row": { + "CHEBI_ID": "CHEBI:54684", + "substrate": "4-Nitrophenyl-alpha-D-galactoside", + "KEGG_ID": "", + "CAS_RN_ID": "CAS-RN:7493-95-0", + "EC_ID": "EC:3.2.1.22", + "enzyme": "alpha-galactosidase", + "pseudo_CURIE": "kgmicrobe.assay:API_rID32A_alpha GAL", + "reaction_name": "alpha-Galactosidase" + } + }, + { + "line_number": 119, + "original_line": "CHEBI:54684\t4-Nitrophenyl alpha-D-glucopyranoside \t\tCAS-RN:3767-28-0\tEC:3.2.1.20\talpha-glucosidase\tkgmicrobe.assay:API_ID32E_alpha GLU\talpha-Glucosidase\n", + "row": { + "CHEBI_ID": "CHEBI:54684", + "substrate": "4-Nitrophenyl alpha-D-glucopyranoside ", + "KEGG_ID": "", + "CAS_RN_ID": "CAS-RN:3767-28-0", + "EC_ID": "EC:3.2.1.20", + "enzyme": "alpha-glucosidase", + "pseudo_CURIE": "kgmicrobe.assay:API_ID32E_alpha GLU", + "reaction_name": "alpha-Glucosidase" + } + }, + { + "line_number": 120, + "original_line": "CHEBI:54684\t4-Nitrophenyl-alpha-D-galactoside\tKEGG:C01083\tCAS-RN:7493-95-0\tEC:3.2.1.22\talpha-galactosidase\tkgmicrobe.assay:API_ID32E_alpha GAL\talpha-Galactosidase\n", + "row": { + "CHEBI_ID": "CHEBI:54684", + "substrate": "4-Nitrophenyl-alpha-D-galactoside", + "KEGG_ID": "KEGG:C01083", + "CAS_RN_ID": "CAS-RN:7493-95-0", + "EC_ID": "EC:3.2.1.22", + "enzyme": "alpha-galactosidase", + "pseudo_CURIE": "kgmicrobe.assay:API_ID32E_alpha GAL", + "reaction_name": "alpha-Galactosidase" + } + }, + { + "line_number": 271, + "original_line": "CHEBI:54684\t4-Nitrophenyl-alpha-D-galactoside\t\tCAS-RN:7493-95-0\tEC:3.2.1.22\talpha-galactosidase\tkgmicrobe.assay:API_rID32STR_alpha GAL\talpha-Galactosidase\n", + "row": { + "CHEBI_ID": "CHEBI:54684", + "substrate": "4-Nitrophenyl-alpha-D-galactoside", + "KEGG_ID": "", + "CAS_RN_ID": "CAS-RN:7493-95-0", + "EC_ID": "EC:3.2.1.22", + "enzyme": "alpha-galactosidase", + "pseudo_CURIE": "kgmicrobe.assay:API_rID32STR_alpha GAL", + "reaction_name": "alpha-Galactosidase" + } + } + ] + }, + "native_assay_wells": [ + { + "kit_index": 8, + "well_index": 23, + "kit_name": "API ID32E", + "well": { + "name": "alpha GLU", + "label": [ + "α-glucosidase" + ], + "type": [ + "enzyme" + ], + "description": [ + "Tests for α-glucosidase activity" + ], + "chebi_id": [], + "chebi_name": [], + "pubchem_cid": [], + "pubchem_name": [], + "inchi": [], + "smiles": [], + "enzyme_name": [ + "α-glucosidase" + ], + "ec_number": [ + "3.2.1.20" + ], + "ec_name": [], + "go_terms": [], + "go_names": [], + "kegg_ko": [], + "kegg_reaction": [], + "rhea_ids": [], + "metacyc_reaction": [], + "metacyc_pathway": [] + } + }, + { + "kit_index": 8, + "well_index": 30, + "kit_name": "API ID32E", + "well": { + "name": "alphaMAL", + "label": [ + "α-maltosidase" + ], + "type": [ + "enzyme" + ], + "description": [ + "Tests for α-maltosidase activity" + ], + "chebi_id": [], + "chebi_name": [], + "pubchem_cid": [], + "pubchem_name": [], + "inchi": [], + "smiles": [], + "enzyme_name": [ + "α-maltosidase" + ], + "ec_number": [ + "3.2.1.20" + ], + "ec_name": [], + "go_terms": [], + "go_names": [], + "kegg_ko": [], + "kegg_reaction": [], + "rhea_ids": [], + "metacyc_reaction": [], + "metacyc_pathway": [] + } + } + ], + "primary_structure_witnesses": [ + { + "url": "https://www.sigmaaldrich.com/US/en/product/mm/487506", + "product": "487506", + "cas": "3767-28-0", + "name": "p-Nitrophenyl-alpha-D-glucopyranoside", + "inchi": "InChI=1S/C12H15NO8/c14-5-8-9(15)10(16)11(17)12(21-8)20-7-3-1-6(2-4-7)13(18)19/h1-4,8-12,14-17H,5H2/t8-,9-,10+,11-,12+/m1/s1", + "inchi_key": "IFBHRQDFSNCLOZ-ZIQFBCGOSA-N", + "application_summary": "Manufacturer identifies use as alpha-glucosidase substrate." + }, + { + "url": "https://www.sigmaaldrich.com/US/en/product/sigma/n0877", + "product": "N0877", + "cas": "7493-95-0", + "name": "4-Nitrophenyl-alpha-D-galactopyranoside", + "inchi": "InChI=1S/C12H15NO8/c14-5-8-9(15)10(16)11(17)12(21-8)20-7-3-1-6(2-4-7)13(18)19/h1-4,8-12,14-17H,5H2/t8-,9+,10+,11-,12+/m1/s1", + "inchi_key": "IFBHRQDFSNCLOZ-IIRVCBMXSA-N", + "application_summary": "Formal product/application sections identify alpha-galactosidase use; marketing subtitle conflicts by calling it alpha-glucosidase substrate." + } + ], + "authority_caveats": [ + "CHEBI:54684 absent from the selected complete native node/edge scan; historical obsolescence or official replacement is not established.", + "CHEBI:91122 prose definition mentions beta-D-glucopyranose, while exact IUPAC synonym, parent, stereostructure and independent supplier identify alpha glucoside.", + "Supplier products are independent molecular evidence, not proof of the physical API kit batch.", + "Source physical line120 KEGG:C01083 is contradicted by the primary KEGG trehalose record and is withheld from corrected in-memory data." + ], + "primary_ec_witnesses": [ + { + "id": "EC:3.2.1.20", + "name": "alpha-glucosidase", + "url": "https://iubmb.qmul.ac.uk/enzyme/EC3/2/1/20.html" + }, + { + "id": "EC:3.2.1.22", + "name": "alpha-galactosidase", + "url": "https://iubmb.qmul.ac.uk/enzyme/EC3/2/1/22.html" + } + ], + "withheld_kegg_witness": { + "id": "KEGG:C01083", + "name": "alpha,alpha-Trehalose", + "formula": "C12H22O11", + "url": "https://www.kegg.jp/entry/C01083" + } +} diff --git a/tests/resources/mediadive/p3556_scope.json b/tests/resources/mediadive/p3556_scope.json new file mode 100644 index 00000000..37f3958d --- /dev/null +++ b/tests/resources/mediadive/p3556_scope.json @@ -0,0 +1,277 @@ +{ + "version": 1, + "source_assertion_id": "mediadive.solution:629#recipe/1", + "raw": { + "recipe_order": 1, + "compound": "L-α-Phosphatidylcholine", + "compound_id": 2082, + "attribute": "SIGMA P3556", + "amount": 5, + "unit": "mg", + "g_l": 5, + "optional": 0 + }, + "input_sha256": { + "compounds": "ccc1130355354a55729d31618b152d1c340d75379dfea49352e9075f3419cea1", + "solutions": "9ab0fbf72c2e5ac25ecab3ada268dfe524db971daa86d67532dea6a5412a5c44", + "unified": "09c44642ab13b234e89ce113f7510aa9efb56969e4b539cb7843b43dcb425ba7", + "supported": "6b52b30e018b369aa322d41dfd4e81fcfae0e895e34d7fe48900abf5835815fb", + "chebi_json": "a5d40380ab78bde0e8b5a704dbee3cba2bcaa7608be46eebc2f4a92932516c9b" + }, + "mapping_claims": [ + { + "record_key": "unified:record=582794:line=582872", + "table": "unified", + "row": { + "subject_id": "MIM:L-alpha-Phosphatidylcholine", + "subject_label": "", + "predicate_id": "skos:exactMatch", + "object_id": "CHEBI:86658", + "object_label": "L-alpha-Phosphatidylcholine", + "object_source": "obo:chebi.owl", + "mapping_justification": "semapv:ManualMappingCuration", + "source": "mediaingredientmech_reviewed[manifest=9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af]", + "mapping_date": "2026-09-24", + "confidence": "", + "comment": "", + "object_formula": "", + "object_category": "biolink:ChemicalEntity" + }, + "original_physical_line_utf8": "MIM:L-alpha-Phosphatidylcholine\t\tskos:exactMatch\tCHEBI:86658\tL-alpha-Phosphatidylcholine\tobo:chebi.owl\tsemapv:ManualMappingCuration\tmediaingredientmech_reviewed[manifest=9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af]\t2026-09-24\t\t\t\tbiolink:ChemicalEntity\n" + }, + { + "record_key": "unified:record=582795:line=582873", + "table": "unified", + "row": { + "subject_id": "kgm.name:2r-3-1-oxohexadecoxy-2-9e12e-1-oxooctadeca-912-dienoxypropyl_2-trimethylammonioethyl_phosphate", + "subject_label": "[(2R)-3-(1-oxohexadecoxy)-2-[(9E,12E)-1-oxooctadeca-9,12-dienoxy]propyl] 2-(trimethylammonio)ethyl phosphate", + "predicate_id": "skos:closeMatch", + "object_id": "CHEBI:86658", + "object_label": "L-alpha-Phosphatidylcholine", + "object_source": "obo:chebi.owl", + "mapping_justification": "semapv:LexicalMatching", + "source": "native_ontology:chebi", + "mapping_date": "2026-09-24", + "confidence": "", + "comment": "synonym", + "object_formula": "", + "object_category": "biolink:ChemicalEntity" + }, + "original_physical_line_utf8": "kgm.name:2r-3-1-oxohexadecoxy-2-9e12e-1-oxooctadeca-912-dienoxypropyl_2-trimethylammonioethyl_phosphate\t[(2R)-3-(1-oxohexadecoxy)-2-[(9E,12E)-1-oxooctadeca-9,12-dienoxy]propyl] 2-(trimethylammonio)ethyl phosphate\tskos:closeMatch\tCHEBI:86658\tL-alpha-Phosphatidylcholine\tobo:chebi.owl\tsemapv:LexicalMatching\tnative_ontology:chebi\t2026-09-24\t\tsynonym\t\tbiolink:ChemicalEntity\n" + }, + { + "record_key": "unified:record=582796:line=582874", + "table": "unified", + "row": { + "subject_id": "kgm.name:2r-3-hexadecanoyloxy-2-9e12e-octadeca-912-dienoyloxy-propyl_2-trimethylammonioethyl_phosphate", + "subject_label": "[(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxy-propyl] 2-(trimethylammonio)ethyl phosphate", + "predicate_id": "skos:closeMatch", + "object_id": "CHEBI:86658", + "object_label": "L-alpha-Phosphatidylcholine", + "object_source": "obo:chebi.owl", + "mapping_justification": "semapv:LexicalMatching", + "source": "native_ontology:chebi", + "mapping_date": "2026-09-24", + "confidence": "", + "comment": "synonym", + "object_formula": "", + "object_category": "biolink:ChemicalEntity" + }, + "original_physical_line_utf8": "kgm.name:2r-3-hexadecanoyloxy-2-9e12e-octadeca-912-dienoyloxy-propyl_2-trimethylammonioethyl_phosphate\t[(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxy-propyl] 2-(trimethylammonio)ethyl phosphate\tskos:closeMatch\tCHEBI:86658\tL-alpha-Phosphatidylcholine\tobo:chebi.owl\tsemapv:LexicalMatching\tnative_ontology:chebi\t2026-09-24\t\tsynonym\t\tbiolink:ChemicalEntity\n" + }, + { + "record_key": "unified:record=582797:line=582875", + "table": "unified", + "row": { + "subject_id": "kgm.name:2r-3-hexadecanoyloxy-2-9e12e-octadeca-912-dienoyloxy-propyl_2-trimethylazaniumylethyl_phosphate", + "subject_label": "[(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxy-propyl] 2-(trimethylazaniumyl)ethyl phosphate", + "predicate_id": "skos:closeMatch", + "object_id": "CHEBI:86658", + "object_label": "L-alpha-Phosphatidylcholine", + "object_source": "obo:chebi.owl", + "mapping_justification": "semapv:LexicalMatching", + "source": "native_ontology:chebi", + "mapping_date": "2026-09-24", + "confidence": "", + "comment": "synonym", + "object_formula": "", + "object_category": "biolink:ChemicalEntity" + }, + "original_physical_line_utf8": "kgm.name:2r-3-hexadecanoyloxy-2-9e12e-octadeca-912-dienoyloxy-propyl_2-trimethylazaniumylethyl_phosphate\t[(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxy-propyl] 2-(trimethylazaniumyl)ethyl phosphate\tskos:closeMatch\tCHEBI:86658\tL-alpha-Phosphatidylcholine\tobo:chebi.owl\tsemapv:LexicalMatching\tnative_ontology:chebi\t2026-09-24\t\tsynonym\t\tbiolink:ChemicalEntity\n" + }, + { + "record_key": "unified:record=582798:line=582876", + "table": "unified", + "row": { + "subject_id": "kgm.name:2r-3-hexadecanoyloxy-2-9e12e-octadeca-912-dienoyloxypropyl_2-trimethylazaniumylethyl_phosphate", + "subject_label": "[(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxypropyl] 2-(trimethylazaniumyl)ethyl phosphate", + "predicate_id": "skos:closeMatch", + "object_id": "CHEBI:86658", + "object_label": "L-alpha-Phosphatidylcholine", + "object_source": "obo:chebi.owl", + "mapping_justification": "semapv:LexicalMatching", + "source": "native_ontology:chebi", + "mapping_date": "2026-09-24", + "confidence": "", + "comment": "synonym", + "object_formula": "", + "object_category": "biolink:ChemicalEntity" + }, + "original_physical_line_utf8": "kgm.name:2r-3-hexadecanoyloxy-2-9e12e-octadeca-912-dienoyloxypropyl_2-trimethylazaniumylethyl_phosphate\t[(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxypropyl] 2-(trimethylazaniumyl)ethyl phosphate\tskos:closeMatch\tCHEBI:86658\tL-alpha-Phosphatidylcholine\tobo:chebi.owl\tsemapv:LexicalMatching\tnative_ontology:chebi\t2026-09-24\t\tsynonym\t\tbiolink:ChemicalEntity\n" + }, + { + "record_key": "unified:record=582799:line=582877", + "table": "unified", + "row": { + "subject_id": "kgm.name:l--phosphatidylcholine", + "subject_label": "L-α-Phosphatidylcholine", + "predicate_id": "skos:closeMatch", + "object_id": "CHEBI:86658", + "object_label": "L-alpha-Phosphatidylcholine", + "object_source": "obo:chebi.owl", + "mapping_justification": "semapv:LexicalMatching", + "source": "mediaingredientmech_reviewed[manifest=9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af]", + "mapping_date": "2026-09-24", + "confidence": "", + "comment": "synonym", + "object_formula": "", + "object_category": "biolink:ChemicalEntity" + }, + "original_physical_line_utf8": "kgm.name:l--phosphatidylcholine\tL-α-Phosphatidylcholine\tskos:closeMatch\tCHEBI:86658\tL-alpha-Phosphatidylcholine\tobo:chebi.owl\tsemapv:LexicalMatching\tmediaingredientmech_reviewed[manifest=9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af]\t2026-09-24\t\tsynonym\t\tbiolink:ChemicalEntity\n" + }, + { + "record_key": "unified:record=582800:line=582878", + "table": "unified", + "row": { + "subject_id": "kgm.name:l-alpha-phosphatidylcholine", + "subject_label": "L-alpha-Phosphatidylcholine", + "predicate_id": "skos:exactMatch", + "object_id": "CHEBI:86658", + "object_label": "L-alpha-Phosphatidylcholine", + "object_source": "obo:chebi.owl", + "mapping_justification": "semapv:LexicalMatching", + "source": "mediaingredientmech_reviewed[manifest=9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af]|native_ontology:chebi", + "mapping_date": "2026-09-24", + "confidence": "", + "comment": "canonical_name", + "object_formula": "", + "object_category": "biolink:ChemicalEntity" + }, + "original_physical_line_utf8": "kgm.name:l-alpha-phosphatidylcholine\tL-alpha-Phosphatidylcholine\tskos:exactMatch\tCHEBI:86658\tL-alpha-Phosphatidylcholine\tobo:chebi.owl\tsemapv:LexicalMatching\tmediaingredientmech_reviewed[manifest=9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af]|native_ontology:chebi\t2026-09-24\t\tcanonical_name\t\tbiolink:ChemicalEntity\n" + }, + { + "record_key": "supported:record=950:line=987", + "table": "supported", + "row": { + "subject_id": "MIM:L-alpha-Phosphatidylcholine", + "subject_label": "L-alpha-Phosphatidylcholine", + "predicate_id": "skos:exactMatch", + "object_id": "CHEBI:86658", + "object_label": "L-alpha-Phosphatidylcholine", + "object_source": "obo:chebi.owl", + "mapping_justification": "semapv:LexicalMatching", + "source": "MIM:EBI OLS exact search|MIM:curator=refresh_occurrence_statistics|kgm:culturebotai_reviewed|kgm:mediaingredientmech_reviewed|kgm:mediaingredientmech_reviewed[curator=unmapped_ols_exact_audit]", + "mapping_date": "2026-08-27", + "confidence": "0.99", + "comment": "", + "other": "L-α-Phosphatidylcholine", + "validation_method": "OAK+OLS:chebi|SYNONYM_ENRICH|2026-07-07" + }, + "original_physical_line_utf8": "MIM:L-alpha-Phosphatidylcholine\tL-alpha-Phosphatidylcholine\tskos:exactMatch\tCHEBI:86658\tL-alpha-Phosphatidylcholine\tobo:chebi.owl\tsemapv:LexicalMatching\tMIM:EBI OLS exact search|MIM:curator=refresh_occurrence_statistics|kgm:culturebotai_reviewed|kgm:mediaingredientmech_reviewed|kgm:mediaingredientmech_reviewed[curator=unmapped_ols_exact_audit]\t2026-08-27\t0.99\t\tL-α-Phosphatidylcholine\tOAK+OLS:chebi|SYNONYM_ENRICH|2026-07-07\n" + } + ], + "native_structure": [ + { + "node_ordinal": 204443, + "node": { + "id": "http://purl.obolibrary.org/obo/CHEBI_86658", + "lbl": "L-alpha-Phosphatidylcholine", + "type": "CLASS", + "meta": { + "subsets": [ + "http://purl.obolibrary.org/obo/chebi/1_STAR" + ], + "synonyms": [ + { + "pred": "hasRelatedSynonym", + "val": "[(2R)-3-(1-oxohexadecoxy)-2-[(9E,12E)-1-oxooctadeca-9,12-dienoxy]propyl] 2-(trimethylammonio)ethyl phosphate", + "xrefs": [ + "chebi" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "[(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxy-propyl] 2-(trimethylammonio)ethyl phosphate", + "xrefs": [ + "chebi" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "[(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxy-propyl] 2-(trimethylazaniumyl)ethyl phosphate", + "xrefs": [ + "chebi" + ] + }, + { + "pred": "hasRelatedSynonym", + "val": "[(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxypropyl] 2-(trimethylazaniumyl)ethyl phosphate", + "xrefs": [ + "chebi" + ] + } + ], + "basicPropertyValues": [ + { + "pred": "http://www.geneontology.org/formats/oboInOwl#hasOBONamespace", + "val": "chebi_ontology" + }, + { + "pred": "https://w3id.org/chemrof/charge", + "val": "0" + }, + { + "pred": "https://w3id.org/chemrof/generalized_empirical_formula", + "val": "C42H80NO8P" + }, + { + "pred": "https://w3id.org/chemrof/inchi_key_string", + "val": "JLPULHDHAOZNQI-JLOPVYAASA-N" + }, + { + "pred": "https://w3id.org/chemrof/inchi_string", + "val": "InChI=1S/C42H80NO8P/c1-6-8-10-12-14-16-18-20-21-23-25-27-29-31-33-35-42(45)51-40(39-50-52(46,47)49-37-36-43(3,4)5)38-48-41(44)34-32-30-28-26-24-22-19-17-15-13-11-9-7-2/h14,16,20-21,40H,6-13,15,17-19,22-39H2,1-5H3/b16-14+,21-20+/t40-/m1/s1" + }, + { + "pred": "https://w3id.org/chemrof/mass", + "val": "758.075" + }, + { + "pred": "https://w3id.org/chemrof/monoisotopic_mass", + "val": "757.56216" + }, + { + "pred": "https://w3id.org/chemrof/smiles_string", + "val": "CCCCC/C=C/C/C=C/CCCCCCCC(=O)O[C@H](COC(=O)CCCCCCCCCCCCCCC)COP(=O)([O-])OCC[N+](C)(C)C" + } + ] + } + } + } + ], + "primary_evidence": [ + { + "uri": "https://www.sigmaaldrich.com/deepweb/assets/sigmaaldrich/product/documents/152/475/p3556pis.pdf", + "scope": "Supplier P3556 product sheet: egg-yolk fatty-acid composition has 16:0, 18:0, 18:1, 18:2 and minor other residues. This does not establish one pure acyl species." + }, + { + "uri": "https://b2b.sigmaaldrich.com/deepweb/assets/sigmaaldrich/marketing/global/documents/709/352/polymeric-drug-delivery-techniques-web.pdf", + "scope": "PDF page index 18 (printed 17) depicts variable fatty-acid residues R and R-prime for P3556." + }, + { + "uri": "https://www.sigmaaldrich.com/US/en/product/sigma/p3556", + "scope": "Catalog includes a fixed structure alongside egg-yolk product scope. It is not whole-product purity evidence; the catalog stereolayer also differs from native CHEBI:86658." + } + ], + "limitations": "Primary documents were read online, not locally byte-pinned. No claim that the specific molecule is absent from the mixture; no new CAS, charge or chemical identity is approved." +} diff --git a/tests/resources/mediadive/peptone_scope.json b/tests/resources/mediadive/peptone_scope.json new file mode 100644 index 00000000..5c720c3c --- /dev/null +++ b/tests/resources/mediadive/peptone_scope.json @@ -0,0 +1,1403 @@ +{ + "authority_uris": [ + "https://pubchem.ncbi.nlm.nih.gov/compound/167312541", + "https://www.thermofisher.com/order/catalog/product/228820", + "https://www.thermofisher.com/order/catalog/product/kr/en/243620" + ], + "inputs": { + "data/issue1224-quarantine-20260929.xjuS9H/issue286-display-join-mo25m5zm/groups.json": "56d0c510b84106de7a124f927f4428bab95a3879593bdca970424093e7092ec4", + "data/raw/compound_mappings_strict.tsv": "7e100314f1a62f675154aef7537be5b72ed2db9fbb8e8b1785198103aff31c71", + "data/raw/compound_mappings_strict_hydrate.tsv": "d2bc349ae31990cf54ef308a14826c756be0e0ee874d35a0a6068d75ef04b3a3" + }, + "mapping_claims": [ + { + "header_original_line_utf8": "medium_id\toriginal\tmapped\tchebi_label\tchebi_formula\tvalue\tconcentration\tunit\tmmol_l\toptional\tsource\tnormalized_compound\thydration_number\tchebi_match\tchebi_id\tchebi_original_name\tsimilarity_score\tmatch_confidence\tmatching_method\tmapping_status\tmapping_quality\tbase_compound\tbase_formula\twater_molecules\thydrate_formula\tbase_chebi_id\tbase_chebi_label\tbase_chebi_formula\thydration_state\thydration_parsing_method\thydration_confidence\tbase_compound_for_mapping\tbase_molecular_weight\twater_molecular_weight\thydrated_molecular_weight\tcorrected_mmol_l\n", + "input": "data/raw/compound_mappings_strict.tsv", + "input_sha256": "7e100314f1a62f675154aef7537be5b72ed2db9fbb8e8b1785198103aff31c71", + "line_number": 12, + "original_line_utf8": "3\tsoy peptone\tPubChem:167312541\t\t\t3.0\t\tg/L\t\t\tjson\t\t0\t\t\t\t\t\t\toriginal_only\tgood\tsoy peptone\tsoypeptone\t\tsoypeptone\t\t\t\t\tno_hydration\thigh\t\t100.0\t0.0\t100.0\t30.0\n", + "record": 11, + "row": { + "base_chebi_formula": "", + "base_chebi_id": "", + "base_chebi_label": "", + "base_compound": "soy peptone", + "base_compound_for_mapping": "", + "base_formula": "soypeptone", + "base_molecular_weight": "100.0", + "chebi_formula": "", + "chebi_id": "", + "chebi_label": "", + "chebi_match": "", + "chebi_original_name": "", + "concentration": "", + "corrected_mmol_l": "30.0", + "hydrate_formula": "soypeptone", + "hydrated_molecular_weight": "100.0", + "hydration_confidence": "high", + "hydration_number": "0", + "hydration_parsing_method": "no_hydration", + "hydration_state": "", + "mapped": "PubChem:167312541", + "mapping_quality": "good", + "mapping_status": "original_only", + "match_confidence": "", + "matching_method": "", + "medium_id": "3", + "mmol_l": "", + "normalized_compound": "", + "optional": "", + "original": "soy peptone", + "similarity_score": "", + "source": "json", + "unit": "g/L", + "value": "3.0", + "water_molecular_weight": "0.0", + "water_molecules": "" + } + }, + { + "header_original_line_utf8": "medium_id\toriginal\tmapped\tchebi_label\tchebi_formula\tvalue\tconcentration\tunit\tmmol_l\toptional\tsource\tnormalized_compound\thydration_number\tchebi_match\tchebi_id\tchebi_original_name\tsimilarity_score\tmatch_confidence\tmatching_method\tmapping_status\tmapping_quality\tbase_compound\tbase_formula\twater_molecules\thydrate_formula\tbase_chebi_id\tbase_chebi_label\tbase_chebi_formula\thydration_state\thydration_parsing_method\thydration_confidence\tbase_compound_for_mapping\tbase_molecular_weight\twater_molecular_weight\thydrated_molecular_weight\tcorrected_mmol_l\n", + "input": "data/raw/compound_mappings_strict.tsv", + "input_sha256": "7e100314f1a62f675154aef7537be5b72ed2db9fbb8e8b1785198103aff31c71", + "line_number": 4974, + "original_line_utf8": "dsmz_1658_composition\tSoy peptone\tPubChem:167312541\t\t\t5.0\t\tg\t\t\tjson\t\t0\t\t\t\t\t\t\toriginal_only\tgood\tSoy peptone\tSoypeptone\t\tSoypeptone\t\t\t\t\tno_hydration\thigh\t\t100.0\t0.0\t100.0\t50.0\n", + "record": 4971, + "row": { + "base_chebi_formula": "", + "base_chebi_id": "", + "base_chebi_label": "", + "base_compound": "Soy peptone", + "base_compound_for_mapping": "", + "base_formula": "Soypeptone", + "base_molecular_weight": "100.0", + "chebi_formula": "", + "chebi_id": "", + "chebi_label": "", + "chebi_match": "", + "chebi_original_name": "", + "concentration": "", + "corrected_mmol_l": "50.0", + "hydrate_formula": "Soypeptone", + "hydrated_molecular_weight": "100.0", + "hydration_confidence": "high", + "hydration_number": "0", + "hydration_parsing_method": "no_hydration", + "hydration_state": "", + "mapped": "PubChem:167312541", + "mapping_quality": "good", + "mapping_status": "original_only", + "match_confidence": "", + "matching_method": "", + "medium_id": "dsmz_1658_composition", + "mmol_l": "", + "normalized_compound": "", + "optional": "", + "original": "Soy peptone", + "similarity_score": "", + "source": "json", + "unit": "g", + "value": "5.0", + "water_molecular_weight": "0.0", + "water_molecules": "" + } + }, + { + "header_original_line_utf8": "medium_id\toriginal\tmapped\tchebi_label\tchebi_formula\tvalue\tconcentration\tunit\tmmol_l\toptional\tsource\tnormalized_compound\thydration_number\tchebi_match\tchebi_id\tchebi_original_name\tsimilarity_score\tmatch_confidence\tmatching_method\tmapping_status\tmapping_quality\tbase_compound\tbase_formula\twater_molecules\thydrate_formula\tbase_chebi_id\tbase_chebi_label\tbase_chebi_formula\thydration_state\thydration_parsing_method\thydration_confidence\tbase_compound_for_mapping\tbase_molecular_weight\twater_molecular_weight\thydrated_molecular_weight\tcorrected_mmol_l\n", + "input": "data/raw/compound_mappings_strict.tsv", + "input_sha256": "7e100314f1a62f675154aef7537be5b72ed2db9fbb8e8b1785198103aff31c71", + "line_number": 5038, + "original_line_utf8": "dsmz_1669_composition\tSoy peptone\tPubChem:167312541\t\t\t1.0\t\tg\t\t\tjson\t\t0\t\t\t\t\t\t\toriginal_only\tgood\tSoy peptone\tSoypeptone\t\tSoypeptone\t\t\t\t\tno_hydration\thigh\t\t100.0\t0.0\t100.0\t10.0\n", + "record": 5035, + "row": { + "base_chebi_formula": "", + "base_chebi_id": "", + "base_chebi_label": "", + "base_compound": "Soy peptone", + "base_compound_for_mapping": "", + "base_formula": "Soypeptone", + "base_molecular_weight": "100.0", + "chebi_formula": "", + "chebi_id": "", + "chebi_label": "", + "chebi_match": "", + "chebi_original_name": "", + "concentration": "", + "corrected_mmol_l": "10.0", + "hydrate_formula": "Soypeptone", + "hydrated_molecular_weight": "100.0", + "hydration_confidence": "high", + "hydration_number": "0", + "hydration_parsing_method": "no_hydration", + "hydration_state": "", + "mapped": "PubChem:167312541", + "mapping_quality": "good", + "mapping_status": "original_only", + "match_confidence": "", + "matching_method": "", + "medium_id": "dsmz_1669_composition", + "mmol_l": "", + "normalized_compound": "", + "optional": "", + "original": "Soy peptone", + "similarity_score": "", + "source": "json", + "unit": "g", + "value": "1.0", + "water_molecular_weight": "0.0", + "water_molecules": "" + } + }, + { + "header_original_line_utf8": "medium_id\toriginal\tmapped\tchebi_label\tchebi_formula\tvalue\tconcentration\tunit\tmmol_l\toptional\tsource\tnormalized_compound\thydration_number\tchebi_match\tchebi_id\tchebi_original_name\tsimilarity_score\tmatch_confidence\tmatching_method\tmapping_status\tmapping_quality\tbase_compound\tbase_formula\twater_molecules\thydrate_formula\tbase_chebi_id\tbase_chebi_label\tbase_chebi_formula\thydration_state\thydration_parsing_method\thydration_confidence\tbase_compound_for_mapping\tbase_molecular_weight\twater_molecular_weight\thydrated_molecular_weight\tcorrected_mmol_l\n", + "input": "data/raw/compound_mappings_strict.tsv", + "input_sha256": "7e100314f1a62f675154aef7537be5b72ed2db9fbb8e8b1785198103aff31c71", + "line_number": 7400, + "original_line_utf8": "dsmz_545_composition\tSoy peptone\tPubChem:167312541\t\t\t2.5\t\tg\t\t\tjson\t\t0\t\t\t\t\t\t\toriginal_only\tgood\tSoy peptone\tSoypeptone\t\tSoypeptone\t\t\t\t\tno_hydration\thigh\t\t100.0\t0.0\t100.0\t25.0\n", + "record": 7395, + "row": { + "base_chebi_formula": "", + "base_chebi_id": "", + "base_chebi_label": "", + "base_compound": "Soy peptone", + "base_compound_for_mapping": "", + "base_formula": "Soypeptone", + "base_molecular_weight": "100.0", + "chebi_formula": "", + "chebi_id": "", + "chebi_label": "", + "chebi_match": "", + "chebi_original_name": "", + "concentration": "", + "corrected_mmol_l": "25.0", + "hydrate_formula": "Soypeptone", + "hydrated_molecular_weight": "100.0", + "hydration_confidence": "high", + "hydration_number": "0", + "hydration_parsing_method": "no_hydration", + "hydration_state": "", + "mapped": "PubChem:167312541", + "mapping_quality": "good", + "mapping_status": "original_only", + "match_confidence": "", + "matching_method": "", + "medium_id": "dsmz_545_composition", + "mmol_l": "", + "normalized_compound": "", + "optional": "", + "original": "Soy peptone", + "similarity_score": "", + "source": "json", + "unit": "g", + "value": "2.5", + "water_molecular_weight": "0.0", + "water_molecules": "" + } + }, + { + "header_original_line_utf8": "medium_id\toriginal\tmapped\tchebi_label\tchebi_formula\tvalue\tconcentration\tunit\tmmol_l\toptional\tsource\tnormalized_compound\thydration_number\tchebi_match\tchebi_id\tchebi_original_name\tsimilarity_score\tmatch_confidence\tmatching_method\tmapping_status\tmapping_quality\tbase_compound\tbase_formula\twater_molecules\thydrate_formula\tbase_chebi_id\tbase_chebi_label\tbase_chebi_formula\thydration_state\thydration_parsing_method\thydration_confidence\tbase_compound_for_mapping\tbase_molecular_weight\twater_molecular_weight\thydrated_molecular_weight\tcorrected_mmol_l\n", + "input": "data/raw/compound_mappings_strict.tsv", + "input_sha256": "7e100314f1a62f675154aef7537be5b72ed2db9fbb8e8b1785198103aff31c71", + "line_number": 8099, + "original_line_utf8": "dsmz_668_composition\tVitamin-free casamino acids\tPubChem:167312541\t\t\t1.0\t\tg\t\t\tjson\t\t0\t\t\t\t\t\t\toriginal_only\tgood\tVitamin-free casamino acids\tVitamin-freecasaminoacids\t\tVitamin-freecasaminoacids\t\t\t\t\tno_hydration\thigh\t\t100.0\t0.0\t100.0\t10.0\n", + "record": 8094, + "row": { + "base_chebi_formula": "", + "base_chebi_id": "", + "base_chebi_label": "", + "base_compound": "Vitamin-free casamino acids", + "base_compound_for_mapping": "", + "base_formula": "Vitamin-freecasaminoacids", + "base_molecular_weight": "100.0", + "chebi_formula": "", + "chebi_id": "", + "chebi_label": "", + "chebi_match": "", + "chebi_original_name": "", + "concentration": "", + "corrected_mmol_l": "10.0", + "hydrate_formula": "Vitamin-freecasaminoacids", + "hydrated_molecular_weight": "100.0", + "hydration_confidence": "high", + "hydration_number": "0", + "hydration_parsing_method": "no_hydration", + "hydration_state": "", + "mapped": "PubChem:167312541", + "mapping_quality": "good", + "mapping_status": "original_only", + "match_confidence": "", + "matching_method": "", + "medium_id": "dsmz_668_composition", + "mmol_l": "", + "normalized_compound": "", + "optional": "", + "original": "Vitamin-free casamino acids", + "similarity_score": "", + "source": "json", + "unit": "g", + "value": "1.0", + "water_molecular_weight": "0.0", + "water_molecules": "" + } + }, + { + "header_original_line_utf8": "medium_id\toriginal\tmapped\tchebi_label\tchebi_formula\tvalue\tconcentration\tunit\tmmol_l\toptional\tsource\tnormalized_compound\thydration_number\tchebi_match\tchebi_id\tchebi_original_name\tsimilarity_score\tmatch_confidence\tmatching_method\tmapping_status\tmapping_quality\tbase_compound\tbase_formula\twater_molecules\thydrate_formula\tbase_chebi_id\tbase_chebi_label\tbase_chebi_formula\thydration_state\thydration_parsing_method\thydration_confidence\tbase_compound_for_mapping\tbase_molecular_weight\twater_molecular_weight\thydrated_molecular_weight\tcorrected_mmol_l\thydrated_chebi_id\thydrated_chebi_label\thydrate_mapping_source\n", + "input": "data/raw/compound_mappings_strict_hydrate.tsv", + "input_sha256": "d2bc349ae31990cf54ef308a14826c756be0e0ee874d35a0a6068d75ef04b3a3", + "line_number": 12, + "original_line_utf8": "3\tsoy peptone\tPubChem:167312541\t\t\t3.0\t\tg/L\t\t\tjson\t\t0\t\t\t\t\t\t\toriginal_only\tgood\tsoy peptone\tsoypeptone\t\tsoypeptone\t\t\t\t\tno_hydration\thigh\t\t100.0\t0.0\t100.0\t30.0\t\t\t\n", + "record": 11, + "row": { + "base_chebi_formula": "", + "base_chebi_id": "", + "base_chebi_label": "", + "base_compound": "soy peptone", + "base_compound_for_mapping": "", + "base_formula": "soypeptone", + "base_molecular_weight": "100.0", + "chebi_formula": "", + "chebi_id": "", + "chebi_label": "", + "chebi_match": "", + "chebi_original_name": "", + "concentration": "", + "corrected_mmol_l": "30.0", + "hydrate_formula": "soypeptone", + "hydrate_mapping_source": "", + "hydrated_chebi_id": "", + "hydrated_chebi_label": "", + "hydrated_molecular_weight": "100.0", + "hydration_confidence": "high", + "hydration_number": "0", + "hydration_parsing_method": "no_hydration", + "hydration_state": "", + "mapped": "PubChem:167312541", + "mapping_quality": "good", + "mapping_status": "original_only", + "match_confidence": "", + "matching_method": "", + "medium_id": "3", + "mmol_l": "", + "normalized_compound": "", + "optional": "", + "original": "soy peptone", + "similarity_score": "", + "source": "json", + "unit": "g/L", + "value": "3.0", + "water_molecular_weight": "0.0", + "water_molecules": "" + } + }, + { + "header_original_line_utf8": "medium_id\toriginal\tmapped\tchebi_label\tchebi_formula\tvalue\tconcentration\tunit\tmmol_l\toptional\tsource\tnormalized_compound\thydration_number\tchebi_match\tchebi_id\tchebi_original_name\tsimilarity_score\tmatch_confidence\tmatching_method\tmapping_status\tmapping_quality\tbase_compound\tbase_formula\twater_molecules\thydrate_formula\tbase_chebi_id\tbase_chebi_label\tbase_chebi_formula\thydration_state\thydration_parsing_method\thydration_confidence\tbase_compound_for_mapping\tbase_molecular_weight\twater_molecular_weight\thydrated_molecular_weight\tcorrected_mmol_l\thydrated_chebi_id\thydrated_chebi_label\thydrate_mapping_source\n", + "input": "data/raw/compound_mappings_strict_hydrate.tsv", + "input_sha256": "d2bc349ae31990cf54ef308a14826c756be0e0ee874d35a0a6068d75ef04b3a3", + "line_number": 4974, + "original_line_utf8": "dsmz_1658_composition\tSoy peptone\tPubChem:167312541\t\t\t5.0\t\tg\t\t\tjson\t\t0\t\t\t\t\t\t\toriginal_only\tgood\tSoy peptone\tSoypeptone\t\tSoypeptone\t\t\t\t\tno_hydration\thigh\t\t100.0\t0.0\t100.0\t50.0\t\t\t\n", + "record": 4971, + "row": { + "base_chebi_formula": "", + "base_chebi_id": "", + "base_chebi_label": "", + "base_compound": "Soy peptone", + "base_compound_for_mapping": "", + "base_formula": "Soypeptone", + "base_molecular_weight": "100.0", + "chebi_formula": "", + "chebi_id": "", + "chebi_label": "", + "chebi_match": "", + "chebi_original_name": "", + "concentration": "", + "corrected_mmol_l": "50.0", + "hydrate_formula": "Soypeptone", + "hydrate_mapping_source": "", + "hydrated_chebi_id": "", + "hydrated_chebi_label": "", + "hydrated_molecular_weight": "100.0", + "hydration_confidence": "high", + "hydration_number": "0", + "hydration_parsing_method": "no_hydration", + "hydration_state": "", + "mapped": "PubChem:167312541", + "mapping_quality": "good", + "mapping_status": "original_only", + "match_confidence": "", + "matching_method": "", + "medium_id": "dsmz_1658_composition", + "mmol_l": "", + "normalized_compound": "", + "optional": "", + "original": "Soy peptone", + "similarity_score": "", + "source": "json", + "unit": "g", + "value": "5.0", + "water_molecular_weight": "0.0", + "water_molecules": "" + } + }, + { + "header_original_line_utf8": "medium_id\toriginal\tmapped\tchebi_label\tchebi_formula\tvalue\tconcentration\tunit\tmmol_l\toptional\tsource\tnormalized_compound\thydration_number\tchebi_match\tchebi_id\tchebi_original_name\tsimilarity_score\tmatch_confidence\tmatching_method\tmapping_status\tmapping_quality\tbase_compound\tbase_formula\twater_molecules\thydrate_formula\tbase_chebi_id\tbase_chebi_label\tbase_chebi_formula\thydration_state\thydration_parsing_method\thydration_confidence\tbase_compound_for_mapping\tbase_molecular_weight\twater_molecular_weight\thydrated_molecular_weight\tcorrected_mmol_l\thydrated_chebi_id\thydrated_chebi_label\thydrate_mapping_source\n", + "input": "data/raw/compound_mappings_strict_hydrate.tsv", + "input_sha256": "d2bc349ae31990cf54ef308a14826c756be0e0ee874d35a0a6068d75ef04b3a3", + "line_number": 5038, + "original_line_utf8": "dsmz_1669_composition\tSoy peptone\tPubChem:167312541\t\t\t1.0\t\tg\t\t\tjson\t\t0\t\t\t\t\t\t\toriginal_only\tgood\tSoy peptone\tSoypeptone\t\tSoypeptone\t\t\t\t\tno_hydration\thigh\t\t100.0\t0.0\t100.0\t10.0\t\t\t\n", + "record": 5035, + "row": { + "base_chebi_formula": "", + "base_chebi_id": "", + "base_chebi_label": "", + "base_compound": "Soy peptone", + "base_compound_for_mapping": "", + "base_formula": "Soypeptone", + "base_molecular_weight": "100.0", + "chebi_formula": "", + "chebi_id": "", + "chebi_label": "", + "chebi_match": "", + "chebi_original_name": "", + "concentration": "", + "corrected_mmol_l": "10.0", + "hydrate_formula": "Soypeptone", + "hydrate_mapping_source": "", + "hydrated_chebi_id": "", + "hydrated_chebi_label": "", + "hydrated_molecular_weight": "100.0", + "hydration_confidence": "high", + "hydration_number": "0", + "hydration_parsing_method": "no_hydration", + "hydration_state": "", + "mapped": "PubChem:167312541", + "mapping_quality": "good", + "mapping_status": "original_only", + "match_confidence": "", + "matching_method": "", + "medium_id": "dsmz_1669_composition", + "mmol_l": "", + "normalized_compound": "", + "optional": "", + "original": "Soy peptone", + "similarity_score": "", + "source": "json", + "unit": "g", + "value": "1.0", + "water_molecular_weight": "0.0", + "water_molecules": "" + } + }, + { + "header_original_line_utf8": "medium_id\toriginal\tmapped\tchebi_label\tchebi_formula\tvalue\tconcentration\tunit\tmmol_l\toptional\tsource\tnormalized_compound\thydration_number\tchebi_match\tchebi_id\tchebi_original_name\tsimilarity_score\tmatch_confidence\tmatching_method\tmapping_status\tmapping_quality\tbase_compound\tbase_formula\twater_molecules\thydrate_formula\tbase_chebi_id\tbase_chebi_label\tbase_chebi_formula\thydration_state\thydration_parsing_method\thydration_confidence\tbase_compound_for_mapping\tbase_molecular_weight\twater_molecular_weight\thydrated_molecular_weight\tcorrected_mmol_l\thydrated_chebi_id\thydrated_chebi_label\thydrate_mapping_source\n", + "input": "data/raw/compound_mappings_strict_hydrate.tsv", + "input_sha256": "d2bc349ae31990cf54ef308a14826c756be0e0ee874d35a0a6068d75ef04b3a3", + "line_number": 7400, + "original_line_utf8": "dsmz_545_composition\tSoy peptone\tPubChem:167312541\t\t\t2.5\t\tg\t\t\tjson\t\t0\t\t\t\t\t\t\toriginal_only\tgood\tSoy peptone\tSoypeptone\t\tSoypeptone\t\t\t\t\tno_hydration\thigh\t\t100.0\t0.0\t100.0\t25.0\t\t\t\n", + "record": 7395, + "row": { + "base_chebi_formula": "", + "base_chebi_id": "", + "base_chebi_label": "", + "base_compound": "Soy peptone", + "base_compound_for_mapping": "", + "base_formula": "Soypeptone", + "base_molecular_weight": "100.0", + "chebi_formula": "", + "chebi_id": "", + "chebi_label": "", + "chebi_match": "", + "chebi_original_name": "", + "concentration": "", + "corrected_mmol_l": "25.0", + "hydrate_formula": "Soypeptone", + "hydrate_mapping_source": "", + "hydrated_chebi_id": "", + "hydrated_chebi_label": "", + "hydrated_molecular_weight": "100.0", + "hydration_confidence": "high", + "hydration_number": "0", + "hydration_parsing_method": "no_hydration", + "hydration_state": "", + "mapped": "PubChem:167312541", + "mapping_quality": "good", + "mapping_status": "original_only", + "match_confidence": "", + "matching_method": "", + "medium_id": "dsmz_545_composition", + "mmol_l": "", + "normalized_compound": "", + "optional": "", + "original": "Soy peptone", + "similarity_score": "", + "source": "json", + "unit": "g", + "value": "2.5", + "water_molecular_weight": "0.0", + "water_molecules": "" + } + }, + { + "header_original_line_utf8": "medium_id\toriginal\tmapped\tchebi_label\tchebi_formula\tvalue\tconcentration\tunit\tmmol_l\toptional\tsource\tnormalized_compound\thydration_number\tchebi_match\tchebi_id\tchebi_original_name\tsimilarity_score\tmatch_confidence\tmatching_method\tmapping_status\tmapping_quality\tbase_compound\tbase_formula\twater_molecules\thydrate_formula\tbase_chebi_id\tbase_chebi_label\tbase_chebi_formula\thydration_state\thydration_parsing_method\thydration_confidence\tbase_compound_for_mapping\tbase_molecular_weight\twater_molecular_weight\thydrated_molecular_weight\tcorrected_mmol_l\thydrated_chebi_id\thydrated_chebi_label\thydrate_mapping_source\n", + "input": "data/raw/compound_mappings_strict_hydrate.tsv", + "input_sha256": "d2bc349ae31990cf54ef308a14826c756be0e0ee874d35a0a6068d75ef04b3a3", + "line_number": 8099, + "original_line_utf8": "dsmz_668_composition\tVitamin-free casamino acids\tPubChem:167312541\t\t\t1.0\t\tg\t\t\tjson\t\t0\t\t\t\t\t\t\toriginal_only\tgood\tVitamin-free casamino acids\tVitamin-freecasaminoacids\t\tVitamin-freecasaminoacids\t\t\t\t\tno_hydration\thigh\t\t100.0\t0.0\t100.0\t10.0\t\t\t\n", + "record": 8094, + "row": { + "base_chebi_formula": "", + "base_chebi_id": "", + "base_chebi_label": "", + "base_compound": "Vitamin-free casamino acids", + "base_compound_for_mapping": "", + "base_formula": "Vitamin-freecasaminoacids", + "base_molecular_weight": "100.0", + "chebi_formula": "", + "chebi_id": "", + "chebi_label": "", + "chebi_match": "", + "chebi_original_name": "", + "concentration": "", + "corrected_mmol_l": "10.0", + "hydrate_formula": "Vitamin-freecasaminoacids", + "hydrate_mapping_source": "", + "hydrated_chebi_id": "", + "hydrated_chebi_label": "", + "hydrated_molecular_weight": "100.0", + "hydration_confidence": "high", + "hydration_number": "0", + "hydration_parsing_method": "no_hydration", + "hydration_state": "", + "mapped": "PubChem:167312541", + "mapping_quality": "good", + "mapping_status": "original_only", + "match_confidence": "", + "matching_method": "", + "medium_id": "dsmz_668_composition", + "mmol_l": "", + "normalized_compound": "", + "optional": "", + "original": "Vitamin-free casamino acids", + "similarity_score": "", + "source": "json", + "unit": "g", + "value": "1.0", + "water_molecular_weight": "0.0", + "water_molecules": "" + } + } + ], + "occurrences": [ + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 69305, + "original_line_utf8": "mediadive.solution:1139\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t1.0\t\t\tmediadive.solution:1139#recipe/3\t{\"amount\":1,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":1,\"optional\":0,\"recipe_order\":3,\"unit\":\"g\"}\tg\t1.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "1.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:1139#recipe/3", + "source_record": "{\"amount\":1,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":1,\"optional\":0,\"recipe_order\":3,\"unit\":\"g\"}", + "subject": "mediadive.solution:1139", + "unit": "g", + "value": "1.0" + } + } + ], + "raw": { + "amount": 1, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 1, + "optional": 0, + "recipe_order": 3, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:1139#recipe/3" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 69737, + "original_line_utf8": "mediadive.solution:1185\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t3.0\t\t\tmediadive.solution:1185#recipe/2\t{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t3.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "3.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:1185#recipe/2", + "source_record": "{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:1185", + "unit": "g", + "value": "3.0" + } + } + ], + "raw": { + "amount": 3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 3, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:1185#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 69892, + "original_line_utf8": "mediadive.solution:1203\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t3.0\t\t\tmediadive.solution:1203#recipe/2\t{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t3.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "3.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:1203#recipe/2", + "source_record": "{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:1203", + "unit": "g", + "value": "3.0" + } + } + ], + "raw": { + "amount": 3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 3, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:1203#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 71429, + "original_line_utf8": "mediadive.solution:1390\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t5.0\t\t\tmediadive.solution:1390#recipe/2\t{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t5.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "5.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:1390#recipe/2", + "source_record": "{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:1390", + "unit": "g", + "value": "5.0" + } + } + ], + "raw": { + "amount": 5, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 5, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:1390#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 72789, + "original_line_utf8": "mediadive.solution:1537\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t8.0\t\t\tmediadive.solution:1537#recipe/2\t{\"amount\":8,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":8,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t8.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "8.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:1537#recipe/2", + "source_record": "{\"amount\":8,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":8,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:1537", + "unit": "g", + "value": "8.0" + } + } + ], + "raw": { + "amount": 8, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 8, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:1537#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 73213, + "original_line_utf8": "mediadive.solution:159\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t3.0\t\t\tmediadive.solution:159#recipe/2\t{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t3.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "3.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:159#recipe/2", + "source_record": "{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:159", + "unit": "g", + "value": "3.0" + } + } + ], + "raw": { + "amount": 3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 3, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:159#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 80004, + "original_line_utf8": "mediadive.solution:2339\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t1.0\t\t\tmediadive.solution:2339#recipe/2\t{\"amount\":1,\"attribute\":\"e.g. Bacto Soytone\",\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":1,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t1.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "1.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:2339#recipe/2", + "source_record": "{\"amount\":1,\"attribute\":\"e.g. Bacto Soytone\",\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":1,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:2339", + "unit": "g", + "value": "1.0" + } + } + ], + "raw": { + "amount": 1, + "attribute": "e.g. Bacto Soytone", + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 1, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:2339#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 82746, + "original_line_utf8": "mediadive.solution:2668\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t5.0\t\t\tmediadive.solution:2668#recipe/2\t{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t5.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "5.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:2668#recipe/2", + "source_record": "{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:2668", + "unit": "g", + "value": "5.0" + } + } + ], + "raw": { + "amount": 5, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 5, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:2668#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 83396, + "original_line_utf8": "mediadive.solution:2745\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t3.0\t\t\tmediadive.solution:2745#recipe/2\t{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t3.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "3.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:2745#recipe/2", + "source_record": "{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:2745", + "unit": "g", + "value": "3.0" + } + } + ], + "raw": { + "amount": 3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 3, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:2745#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 83887, + "original_line_utf8": "mediadive.solution:2815\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t3.3\t\t\tmediadive.solution:2815#recipe/7\t{\"amount\":3.3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3.3,\"optional\":0,\"recipe_order\":7,\"unit\":\"g\"}\tg\t3.3\n", + "row": { + "agent_type": "manual_agent", + "g_l": "3.3", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:2815#recipe/7", + "source_record": "{\"amount\":3.3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3.3,\"optional\":0,\"recipe_order\":7,\"unit\":\"g\"}", + "subject": "mediadive.solution:2815", + "unit": "g", + "value": "3.3" + } + } + ], + "raw": { + "amount": 3.3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 3.3, + "optional": 0, + "recipe_order": 7, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:2815#recipe/7" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 86205, + "original_line_utf8": "mediadive.solution:3119\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t1.0\t\t\tmediadive.solution:3119#recipe/3\t{\"amount\":1,\"attribute\":\"BD Soytone\",\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":1,\"optional\":0,\"recipe_order\":3,\"unit\":\"g\"}\tg\t1.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "1.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:3119#recipe/3", + "source_record": "{\"amount\":1,\"attribute\":\"BD Soytone\",\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":1,\"optional\":0,\"recipe_order\":3,\"unit\":\"g\"}", + "subject": "mediadive.solution:3119", + "unit": "g", + "value": "1.0" + } + } + ], + "raw": { + "amount": 1, + "attribute": "BD Soytone", + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 1, + "optional": 0, + "recipe_order": 3, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:3119#recipe/3" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 86246, + "original_line_utf8": "mediadive.solution:313\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t5.0\t\t\tmediadive.solution:313#recipe/3\t{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":3,\"unit\":\"g\"}\tg\t5.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "5.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:313#recipe/3", + "source_record": "{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":3,\"unit\":\"g\"}", + "subject": "mediadive.solution:313", + "unit": "g", + "value": "5.0" + } + } + ], + "raw": { + "amount": 5, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 5, + "optional": 0, + "recipe_order": 3, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:313#recipe/3" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 86337, + "original_line_utf8": "mediadive.solution:3145\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t3.0\t\t\tmediadive.solution:3145#recipe/2\t{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t3.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "3.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:3145#recipe/2", + "source_record": "{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:3145", + "unit": "g", + "value": "3.0" + } + } + ], + "raw": { + "amount": 3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 3, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:3145#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 86624, + "original_line_utf8": "mediadive.solution:3170\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t3.0\t\t\tmediadive.solution:3170#recipe/2\t{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t3.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "3.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:3170#recipe/2", + "source_record": "{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:3170", + "unit": "g", + "value": "3.0" + } + } + ], + "raw": { + "amount": 3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 3, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:3170#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 88350, + "original_line_utf8": "mediadive.solution:3439\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t5.0\t\t\tmediadive.solution:3439#recipe/2\t{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t5.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "5.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:3439#recipe/2", + "source_record": "{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:3439", + "unit": "g", + "value": "5.0" + } + } + ], + "raw": { + "amount": 5, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 5, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:3439#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 88581, + "original_line_utf8": "mediadive.solution:3464\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t1.0\t\t\tmediadive.solution:3464#recipe/2\t{\"amount\":1,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":1,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t1.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "1.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:3464#recipe/2", + "source_record": "{\"amount\":1,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":1,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:3464", + "unit": "g", + "value": "1.0" + } + } + ], + "raw": { + "amount": 1, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 1, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:3464#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 88790, + "original_line_utf8": "mediadive.solution:3500\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t3.0\t\t\tmediadive.solution:3500#recipe/2\t{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t3.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "3.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:3500#recipe/2", + "source_record": "{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:3500", + "unit": "g", + "value": "3.0" + } + } + ], + "raw": { + "amount": 3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 3, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:3500#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 89197, + "original_line_utf8": "mediadive.solution:3553\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t0.3\t\t\tmediadive.solution:3553#recipe/2\t{\"amount\":0.3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":0.3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t0.3\n", + "row": { + "agent_type": "manual_agent", + "g_l": "0.3", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:3553#recipe/2", + "source_record": "{\"amount\":0.3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":0.3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:3553", + "unit": "g", + "value": "0.3" + } + } + ], + "raw": { + "amount": 0.3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 0.3, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:3553#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 98993, + "original_line_utf8": "mediadive.solution:477\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t\t\t\tmediadive.solution:477#recipe/2\t{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t5.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:477#recipe/2", + "source_record": "{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:477", + "unit": "g", + "value": "5.0" + } + } + ], + "raw": { + "amount": 5, + "compound": "Soy peptone", + "compound_id": 75, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:477#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 99199, + "original_line_utf8": "mediadive.solution:479\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t5.0\t\t\tmediadive.solution:479#recipe/2\t{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t5.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "5.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:479#recipe/2", + "source_record": "{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:479", + "unit": "g", + "value": "5.0" + } + } + ], + "raw": { + "amount": 5, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 5, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:479#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 99287, + "original_line_utf8": "mediadive.solution:48\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t2.0\t\t\tmediadive.solution:48#recipe/3\t{\"amount\":2,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":2,\"optional\":0,\"recipe_order\":3,\"unit\":\"g\"}\tg\t2.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "2.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:48#recipe/3", + "source_record": "{\"amount\":2,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":2,\"optional\":0,\"recipe_order\":3,\"unit\":\"g\"}", + "subject": "mediadive.solution:48", + "unit": "g", + "value": "2.0" + } + } + ], + "raw": { + "amount": 2, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 2, + "optional": 0, + "recipe_order": 3, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:48#recipe/3" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 111177, + "original_line_utf8": "mediadive.solution:6398\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t3.0\t\t\tmediadive.solution:6398#recipe/2\t{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t3.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "3.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:6398#recipe/2", + "source_record": "{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:6398", + "unit": "g", + "value": "3.0" + } + } + ], + "raw": { + "amount": 3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 3, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:6398#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 111399, + "original_line_utf8": "mediadive.solution:6469\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t10.0\t\t\tmediadive.solution:6469#recipe/1\t{\"amount\":10,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":10,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t10.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "10.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:6469#recipe/1", + "source_record": "{\"amount\":10,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":10,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:6469", + "unit": "g", + "value": "10.0" + } + } + ], + "raw": { + "amount": 10, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 10, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:6469#recipe/1" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 112217, + "original_line_utf8": "mediadive.solution:72\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t3.0\t\t\tmediadive.solution:72#recipe/2\t{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t3.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "3.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:72#recipe/2", + "source_record": "{\"amount\":3,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":3,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:72", + "unit": "g", + "value": "3.0" + } + } + ], + "raw": { + "amount": 3, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 3, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:72#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 113259, + "original_line_utf8": "mediadive.solution:900\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t5.0\t\t\tmediadive.solution:900#recipe/2\t{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t5.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "5.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:900#recipe/2", + "source_record": "{\"amount\":5,\"compound\":\"Soy peptone\",\"compound_id\":75,\"g_l\":5,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:900", + "unit": "g", + "value": "5.0" + } + } + ], + "raw": { + "amount": 5, + "compound": "Soy peptone", + "compound_id": 75, + "g_l": 5, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:900#recipe/2" + }, + { + "observed_edge_rows": [ + { + "header_original_line_utf8": "subject\tpredicate\tobject\trelation\tprimary_knowledge_source\tknowledge_level\tagent_type\tg_l\tmmol_l\tpublications\tsource_assertion_id\tsource_record\tunit\tvalue\n", + "line_number": 71635, + "original_line_utf8": "mediadive.solution:1410\tbiolink:has_part\tPubChem:167312541\tBFO:0000051\tinfores:mediadive\tobservation\tmanual_agent\t1.0\t\t\tmediadive.solution:1410#recipe/2\t{\"amount\":1,\"compound\":\"Vitamin-free casamino acids\",\"compound_id\":654,\"g_l\":1,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}\tg\t1.0\n", + "row": { + "agent_type": "manual_agent", + "g_l": "1.0", + "knowledge_level": "observation", + "mmol_l": "", + "object": "PubChem:167312541", + "predicate": "biolink:has_part", + "primary_knowledge_source": "infores:mediadive", + "publications": "", + "relation": "BFO:0000051", + "source_assertion_id": "mediadive.solution:1410#recipe/2", + "source_record": "{\"amount\":1,\"compound\":\"Vitamin-free casamino acids\",\"compound_id\":654,\"g_l\":1,\"optional\":0,\"recipe_order\":2,\"unit\":\"g\"}", + "subject": "mediadive.solution:1410", + "unit": "g", + "value": "1.0" + } + } + ], + "raw": { + "amount": 1, + "compound": "Vitamin-free casamino acids", + "compound_id": 654, + "g_l": 1, + "optional": 0, + "recipe_order": 2, + "unit": "g" + }, + "source_assertion_id": "mediadive.solution:1410#recipe/2" + } + ], + "scope": "26 saved raw observations plus all ten exact name/target legacy claim rows; no new identity" +} diff --git a/tests/resources/mediadive/sugar_context.json b/tests/resources/mediadive/sugar_context.json new file mode 100644 index 00000000..d02fbd82 --- /dev/null +++ b/tests/resources/mediadive/sugar_context.json @@ -0,0 +1,15 @@ +{ + "raw": {"amount": 0.1, "attribute": "250 mM each of xylose, maltose and cellobiose", "compound": "Sugar", "compound_id": 1795, "optional": 0, "recipe_order": 16, "unit": "ml"}, + "source_assertion_id": "mediadive.solution:4338#recipe/16", + "source_solutions_sha256": "9ab0fbf72c2e5ac25ecab3ada268dfe524db971daa86d67532dea6a5412a5c44", + "source_compounds_sha256": "ccc1130355354a55729d31618b152d1c340d75379dfea49352e9075f3419cea1", + "saved_mapping_evidence_sha256": "801e802d8cecead0af2b11486a35d79ecc9dcf02786329887e739b7d8ad25617", + "primary_recipe_uri": "https://www.jcm.riken.jp/cgi-bin/jcm/jcm_grmd?GRMD=537", + "native_definition_fixture": "tests/resources/ncit_category_projection/native.json", + "native_definition_fixture_sha256": "92e77fc97f4b0e5c1253251ed0868d62b6b163f1232dbba816e06388cefd54f4", + "mapping_claims": [ + {"table": "unified", "original_record": 590651, "original_line_end": 590729, "row": {"subject_id": "MIM:Sugar", "subject_label": "", "predicate_id": "skos:exactMatch", "object_id": "NCIT:C71939", "object_label": "Sugar", "object_source": "obo:ncit.owl", "mapping_justification": "semapv:ManualMappingCuration", "source": "mediaingredientmech_reviewed[manifest=9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af]", "mapping_date": "2026-09-24", "confidence": "", "comment": "", "object_formula": "", "object_category": "biolink:Food"}}, + {"table": "unified", "original_record": 590652, "original_line_end": 590730, "row": {"subject_id": "kgm.name:sugar", "subject_label": "Sugar", "predicate_id": "skos:exactMatch", "object_id": "NCIT:C71939", "object_label": "Sugar", "object_source": "obo:ncit.owl", "mapping_justification": "semapv:LexicalMatching", "source": "mediaingredientmech_reviewed[manifest=9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af]|native_ontology:ncit", "mapping_date": "2026-09-24", "confidence": "", "comment": "canonical_name", "object_formula": "", "object_category": "biolink:Food"}}, + {"table": "supported", "original_record": 1559, "original_line_end": 1596, "row": {"subject_id": "MIM:Sugar", "subject_label": "Sugar", "predicate_id": "skos:exactMatch", "object_id": "NCIT:C71939", "object_label": "Sugar", "object_source": "obo:ncit.owl", "mapping_justification": "semapv:LexicalMatching", "source": "MIM:mim-queue|MIM:curator=refresh_occurrence_statistics", "mapping_date": "2026-08-27", "confidence": "0.99", "comment": "", "other": "", "validation_method": "none|UNKNOWN_TERM|2026-07-07"}} + ] +} diff --git a/tests/resources/merge_source_freshness/bacdive_ec_substrates.tsv b/tests/resources/merge_source_freshness/bacdive_ec_substrates.tsv new file mode 100644 index 00000000..d8f6b999 --- /dev/null +++ b/tests/resources/merge_source_freshness/bacdive_ec_substrates.tsv @@ -0,0 +1 @@ +CHEBI_ID substrate KEGG_ID CAS_RN_ID EC_ID enzyme pseudo_CURIE reaction_name diff --git a/tests/resources/merge_source_freshness/ec_substrate_corrections.tsv b/tests/resources/merge_source_freshness/ec_substrate_corrections.tsv new file mode 100644 index 00000000..99cc37f5 --- /dev/null +++ b/tests/resources/merge_source_freshness/ec_substrate_corrections.tsv @@ -0,0 +1 @@ +correction_id status source_file source_sha256 source_line source_row_json original_edge_json emitted_edge_json withheld_source_fields reason evidence_uri policy_file policy_sha256 diff --git a/tests/resources/tetramethylammonium_chloride.json b/tests/resources/tetramethylammonium_chloride.json new file mode 100644 index 00000000..2fd45c41 --- /dev/null +++ b/tests/resources/tetramethylammonium_chloride.json @@ -0,0 +1,31 @@ +{ + "origin": { + "compounds_sha256": "ccc1130355354a55729d31618b152d1c340d75379dfea49352e9075f3419cea1", + "solutions_sha256": "9ab0fbf72c2e5ac25ecab3ada268dfe524db971daa86d67532dea6a5412a5c44", + "native_nodes_sha256": "c161f8aea7ffeb1a2557f24100b3b3b67bfd62196d484b33c675921154c8b152", + "source_edges_sha256": "05ecb2df49afa66607393fbc0d6f9e15980315ba7bd088d38793d5c329fbf787" + }, + "native_projection": { + "id": "CHEBI:7070", + "name": "N,N,N-Trimethylmethanaminium chloride", + "synonym": "Tetramethylammonium chloride", + "category": "biolink:ChemicalEntity", + "xref": "cas:75-57-0|kegg.compound:C11335" + }, + "occurrences": [ + { + "solution_id": "1577", + "array_index": 3, + "previous_target": "mediadive.ingredient:720", + "expected_target": "CHEBI:7070", + "raw": {"recipe_order":4,"compound":"Tetramethyl ammonium chloride","compound_id":720,"amount":1,"unit":"g","g_l":1,"optional":0} + }, + { + "solution_id": "3797", + "array_index": 0, + "previous_target": "kgmicrobe.compound:tetramethyl_ammonium", + "expected_target": "kgmicrobe.compound:tetramethyl_ammonium", + "raw": {"recipe_order":1,"compound":"Tetramethyl ammonium","compound_id":1649,"amount":1,"unit":"g","g_l":1,"optional":0} + } + ] +} diff --git a/tests/test_bacdive_ec_substrate_corrections.py b/tests/test_bacdive_ec_substrate_corrections.py new file mode 100644 index 00000000..a127b2bc --- /dev/null +++ b/tests/test_bacdive_ec_substrate_corrections.py @@ -0,0 +1,339 @@ +"""Finite source-row corrections preserve original evidence, independent assay outputs and admission.""" + +import copy +import csv +import hashlib +import io +import json +from collections import Counter +from pathlib import Path + +import pytest + +from kg_microbe.transform_utils.bacdive import bacdive as producer +from kg_microbe.transform_utils.bacdive import ec_substrate_corrections as correction +from kg_microbe.utils.producer_audits import verify_producer_audits +from kg_microbe.utils.source_finalization import SourceFinalizationRequired, graph_rows, verify_finalized_source_files +from tests.test_bacdive_run_references import FIXTURE as RECORDS +from tests.test_bacdive_run_references import _run_fixture + +FIXTURE = Path(__file__).parent / "resources/bacdive_ec_substrate_corrections/authority.json" + + +def _evidence(): + """Read complete immutable original/native records without production data access.""" + raw = FIXTURE.read_bytes() + assert hashlib.sha256(raw).hexdigest() == "9dac3cd68874e887e116426e26f120b3ff6e8e04ae7d24210ae0f049a054abb6" + return json.loads(raw) + + +def _originals(): + """Return the complete reviewed legacy cells, including the trailing space and contradictory KEGG.""" + return [(row["line_number"], row["row"]) for row in _evidence()["legacy_rows"]["selected_rows"]] + + +def _policy(): + """Read the actual finite canonical policy through its production parser.""" + with correction.POLICY_PATH.open() as stream: + return correction.load_policy(stream) + + +def _apply(rows, policies=None): + """Exercise production correction, never a copied mapping loop.""" + return correction.apply_corrections( + rows, _policy() if policies is None else policies, knowledge_source="infores:bacdive" + ) + + +def _write_source(path, rows): + """Serialize a tiny selected source cohort with real schema and original cell values.""" + with path.open("w", newline="") as stream: + writer = csv.DictWriter(stream, fieldnames=correction.SOURCE_FIELDS, delimiter="\t") + writer.writeheader() + writer.writerows(row for _, row in rows) + + +def _run(tmp_path, monkeypatch, rows, records=None, during_prepare=None): + """Run the existing tiny real producer fixture, injecting only its finite legacy source table.""" + real_prepare = correction.prepare_corrections + captured = [] + + def prepare(transform, legacy_path): + """Use actual snapshots and correction after writing a synthetic source before its read.""" + _write_source(legacy_path, rows) + captured.append(transform) + result = real_prepare(transform, legacy_path) + if during_prepare: + during_prepare(transform, legacy_path, result) + return result + + monkeypatch.setattr(producer, "prepare_corrections", prepare) + edges, nodes = _run_fixture(tmp_path, monkeypatch, [] if records is None else records) + return captured[0], edges, nodes + + +def test_policy_exact_original_cells_and_independent_stereostructures(): + """Four original rows match four finite rules; the two molecular stereoisomers never collapse.""" + evidence = _evidence() + assert {tuple(rule[key] for key in correction.SOURCE_FIELDS) for rule in _policy()} == { + tuple(row[key] for key in correction.SOURCE_FIELDS) for _, row in _originals() + } + native = { + item["node"]["id"].rsplit("_", 1)[-1]: item["node"] + for item in evidence["raw_selected_nodes_or_exact_identifier_references"] + } + for source, identifier in zip(evidence["primary_structure_witnesses"], ("91122", "546840"), strict=True): + properties = {p["pred"].rsplit("/", 1)[-1]: p["val"] for p in native[identifier]["meta"]["basicPropertyValues"]} + assert properties["inchi_string"] == source["inchi"] + assert properties["inchi_key_string"] == source["inchi_key"] + assert native["91122"]["meta"]["definition"]["val"].startswith("An α-") + assert "β-" in native["91122"]["meta"]["definition"]["val"] # Original prose caveat retained. + assert evidence["withheld_kegg_witness"]["name"] == "alpha,alpha-Trehalose" + + +def test_four_complete_rows_correct_only_their_reviewed_targets_and_withhold_wrong_kegg(): + """Line120's ancillary contradiction is audit-only; no raw cells are mutated.""" + originals = _originals() + before = copy.deepcopy(originals) + rows, audit = _apply(originals) + assert originals == before + assert [row["CHEBI_ID"] for row in rows] == ["CHEBI:546840", "CHEBI:91122", "CHEBI:546840", "CHEBI:546840"] + assert [entry["source_line"] for entry in audit] == ["28", "119", "120", "271"] + assert rows[2]["KEGG_ID"] == "" + assert json.loads(audit[2]["source_row_json"])["KEGG_ID"] == "KEGG:C01083" + assert audit[2]["withheld_source_fields"] == "KEGG_ID" + assert all(entry["status"] == "corrected" for entry in audit) + for row, entry in zip(rows, audit, strict=True): + edge = json.loads(entry["emitted_edge_json"]) + assert edge == { + "subject": row["EC_ID"], + "predicate": "biolink:has_input", + "object": row["CHEBI_ID"], + "relation": "RO:0002233", + "primary_knowledge_source": "infores:bacdive", + "knowledge_level": "knowledge_assertion", + "agent_type": "manual_agent", + } + + +@pytest.mark.parametrize("field", correction.SOURCE_FIELDS) +@pytest.mark.parametrize("index", range(4)) +def test_one_changed_original_cell_cannot_inherit_reviewed_correction(field, index): + """A known row cannot evade exact review by changing or blanking any source field.""" + rows = _originals() + rows[index][1][field] = "" if rows[index][1][field] else "changed" + with pytest.raises(ValueError, match="scope changed"): + _apply(rows) + + +def test_unrelated_old_id_is_not_a_global_alias_and_does_not_gain_new_identity(): + """An unrelated EC/source record keeps every cell; policy does not rewrite an ontology ID globally.""" + row = {**_originals()[0][1], "EC_ID": "EC:1.1.1.1", "CAS_RN_ID": "", "pseudo_CURIE": "kgmicrobe.assay:other"} + rows, audit = _apply([(7, row)]) + assert rows == [row] and audit == [] + + +def test_duplicates_and_already_corrected_rows_keep_full_multiplicity(): + """Reordered physical rows stay scoped by complete cells, not a historical ordinal allowlist.""" + original = _originals()[1][1] + first, audit = _apply([(11, original), (9, original)]) + assert len(first) == len(audit) == 2 + assert [entry["source_line"] for entry in audit] == ["11", "9"] + second, second_audit = _apply(list(enumerate(first, 40))) + assert second == first + assert all(entry["status"] == "already_corrected" for entry in second_audit) + + +@pytest.mark.parametrize( + "field,value", + [ + ("corrected_chebi_id", "CHEBI:54684"), + ("corrected_chebi_id", "CHEBI:x"), + ("withheld_source_fields", "EC_ID"), + ("evidence_uri", ""), + ("reason", ""), + ("pseudo_CURIE", ""), + ("EC_ID", "not-an-EC"), + ], +) +def test_malformed_policy_rejected(field, value): + """Policy metadata cannot turn partial/ambiguous context into a scientific correction.""" + rule = dict(_policy()[0], **{field: value}) + text = io.StringIO() + writer = csv.DictWriter(text, fieldnames=correction.POLICY_FIELDS, delimiter="\t") + writer.writeheader() + writer.writerow(rule) + text.seek(0) + with pytest.raises(ValueError, match="Invalid"): + correction.load_policy(text) + + +@pytest.mark.parametrize("damage", ["duplicate-rule", "duplicate-header", "short-row", "empty-policy", "extra-header"]) +def test_ambiguous_policy_shape_rejected(damage): + """Dictionary parsing cannot silently discard duplicate or malformed authoritative fields.""" + text = correction.POLICY_PATH.read_text() + header, *lines = text.splitlines() + if damage == "duplicate-rule": + text += lines[0] + "\n" + elif damage == "duplicate-header": + text = header + "\tCHEBI_ID\n" + lines[0] + "\tCHEBI:54684\n" + elif damage == "short-row": + text = header + "\nshort\n" + elif damage == "empty-policy": + text = header + "\n" + else: + text = header + "\tunused\n" + lines[0] + "\tvalue\n" + with pytest.raises(ValueError): + correction.load_policy(io.StringIO(text)) + + +@pytest.mark.parametrize("damage", ["duplicate-header", "short-row", "extra-cell"]) +def test_malformed_source_shape_rejected(damage): + """Complete-row matching never proceeds after a malformed source header or record.""" + if damage == "duplicate-header": + text = "EC_ID\tCHEBI_ID\tsubstrate\tCHEBI_ID\n" + elif damage == "short-row": + text = "EC_ID\tCHEBI_ID\tsubstrate\nEC:1\tCHEBI:1\n" + else: + text = "EC_ID\tCHEBI_ID\tsubstrate\nEC:1\tCHEBI:1\twater\textra\n" + with pytest.raises(ValueError): + list(correction._rows(io.StringIO(text))) + + +def test_actual_run_retains_assay_observations_and_full_audit(tmp_path, monkeypatch): + """Only the finite EC/substrate assertions change; independent strain observations remain identical.""" + records = json.loads(RECORDS.read_text()) + baseline = tmp_path / "baseline" + baseline.mkdir() + candidate = tmp_path / "candidate" + candidate.mkdir() + _, old_edges, _ = _run(baseline, monkeypatch, [], records) + transform, new_edges, nodes = _run(candidate, monkeypatch, _originals(), records) + additions = [row for row in new_edges if row["predicate"] == "biolink:has_input"] + assert {(row["subject"], row["object"]) for row in additions} == { + ("EC:3.2.1.20", "CHEBI:91122"), + ("EC:3.2.1.22", "CHEBI:546840"), + } + assert Counter( + tuple(sorted(row.items())) for row in new_edges if row["predicate"] != "biolink:has_input" + ) == Counter(tuple(sorted(row.items())) for row in old_edges) + assert not any(row["object"].startswith(("CAS", "cas:", "KEGG:")) for row in additions) + assert not any(node["id"] == "CHEBI:54684" for node in nodes) + audit = list(graph_rows(transform.output_dir / correction.AUDIT_FILE)) + assert len(audit) == 4 + assert [json.loads(row["source_row_json"]) for row in audit] == [row for _, row in _originals()] + for row in audit: + assert row["source_sha256"] == hashlib.sha256(Path(row["source_file"]).read_bytes()).hexdigest() + assert row["policy_sha256"] == hashlib.sha256(Path(row["policy_file"]).read_bytes()).hexdigest() + assert set(transform.consumed_input_snapshots) == set(correction.REQUIRED_INPUTS) + verify_producer_audits(transform) + + +def test_absent_cohort_still_binds_policy_source_and_header_only_audit(tmp_path, monkeypatch): + """Removing a reviewed source claim is normal, not a forced historical count or hidden fixture bypass.""" + transform, edges, nodes = _run(tmp_path, monkeypatch, []) + assert not edges and not nodes + assert list(graph_rows(transform.output_dir / correction.AUDIT_FILE)) == [] + assert set(transform.consumed_input_snapshots) == set(correction.REQUIRED_INPUTS) + verify_producer_audits(transform) + + +@pytest.mark.parametrize("damage", ["policy", "source", "symlink-retarget"]) +def test_changed_prepared_inputs_abort_without_replacing_prior_graph(tmp_path, monkeypatch, damage): + """The original admission and immutable snapshots remain guards through atomic output publication.""" + output = tmp_path / "transformed/bacdive" + output.mkdir(parents=True) + prior = {} + for name in ("nodes.tsv", "edges.tsv", correction.AUDIT_FILE): + path = output / name + path.write_text("previous complete artifact\n") + prior[path] = path.read_bytes() + selected_policy = tmp_path / "policy.tsv" + selected_policy.write_bytes(correction.POLICY_PATH.read_bytes()) + real_prepare = correction.prepare_corrections + + def prepare(transform, legacy): + """Mutate only temporary fixture inputs after their actual read.""" + _write_source(legacy, _originals()) + if damage == "symlink-retarget": + original = legacy.with_suffix(".original") + original.write_bytes(legacy.read_bytes()) + legacy.unlink() + legacy.symlink_to(original) + result = real_prepare(transform, legacy, policy_path=selected_policy) + if damage == "symlink-retarget": + replacement = legacy.with_suffix(".replacement") + replacement.write_bytes(legacy.read_bytes()) + legacy.unlink() + legacy.symlink_to(replacement) + else: + target = selected_policy if damage == "policy" else legacy + target.write_text(target.read_text() + "changed\n") + return result + + monkeypatch.setattr(producer, "prepare_corrections", prepare) + with pytest.raises((SourceFinalizationRequired, ValueError, RuntimeError)): + _run_fixture(tmp_path, monkeypatch, []) + assert all(path.read_bytes() == content for path, content in prior.items()) + + +@pytest.mark.usefixtures("local_source_schema") +def test_actual_finalization_repeat_and_public_admission_require_original_audit(tmp_path, monkeypatch): + """A normal real producer with no reviewed cohort still publishes a mandatory consumed/audit contract.""" + transform, _, _ = _run(tmp_path, monkeypatch, []) + audit = transform.output_dir / correction.AUDIT_FILE + before = audit.read_bytes() + report = transform.finalize(fresh_run=True) + assert ( + report["producer_audit_members"][correction.AUDIT_FILE] + == transform.producer_audit_snapshots[correction.AUDIT_FILE] + ) + assert set(report["consumed_inputs"]) == set(correction.REQUIRED_INPUTS) + assert transform.finalize() == report + verify_finalized_source_files([transform.output_node_file, transform.output_edge_file]) + assert audit.read_bytes() == before + audit.write_bytes(before + b"tampered\n") + with pytest.raises(SourceFinalizationRequired, match="audit"): + verify_finalized_source_files([transform.output_node_file, transform.output_edge_file]) + + +@pytest.mark.usefixtures("local_source_schema") +def test_actual_corrected_graph_finalizes_with_native_targets_and_unchanged_audit(tmp_path, monkeypatch): + """The real finite graph and original audit survive native closure and strict source admission.""" + transform, _, _ = _run(tmp_path, monkeypatch, _originals()) + native = transform.output_base_dir / "ontologies" + native.mkdir() + evidence = _evidence() + chebi_rows = evidence["native_chebi_rows"] + (native / "chebi_nodes.tsv").write_text( + "\t".join(chebi_rows["header"]) + + "\n" + + "".join(entry["original_line"] for entry in chebi_rows["selected_rows"]) + ) + (native / "ec_nodes.tsv").write_text( + "id\tcategory\tname\tprovided_by\n" + "EC:3.2.1.20\tbiolink:MolecularActivity|biolink:Protein\talpha-glucosidase\tinfores:ec\n" + "EC:3.2.1.22\tbiolink:MolecularActivity|biolink:Protein\talpha-galactosidase\tinfores:ec\n" + ) + (transform.input_base_dir / "chebi.json").write_text( + json.dumps( + { + "graphs": [ + { + "nodes": [ + entry["node"] for entry in evidence["raw_selected_nodes_or_exact_identifier_references"] + ] + } + ] + } + ) + ) + before = (transform.output_dir / correction.AUDIT_FILE).read_bytes() + report = transform.finalize(fresh_run=True) + assert (transform.output_dir / correction.AUDIT_FILE).read_bytes() == before + assert len(list(graph_rows(transform.output_edge_file))) == 2 + assert ( + report["producer_audit_members"][correction.AUDIT_FILE] + == transform.producer_audit_snapshots[correction.AUDIT_FILE] + ) + verify_finalized_source_files([transform.output_node_file, transform.output_edge_file]) diff --git a/tests/test_mediadive_p3556_scope.py b/tests/test_mediadive_p3556_scope.py new file mode 100644 index 00000000..9981c742 --- /dev/null +++ b/tests/test_mediadive_p3556_scope.py @@ -0,0 +1,453 @@ +"""Keep a source-qualified supplier material distinct from a pure acyl species (#1241).""" + +import copy +import csv +import gzip +import hashlib +import json +from collections import Counter +from pathlib import Path + +import pytest + +from kg_microbe.transform_utils.mediadive import material_scope_audit as audit +from kg_microbe.utils import chemical_mapping_utils as mapping +from kg_microbe.utils.ingredient_identity import ( + ingredient_authority_label, + ingredient_mapping_allowed, + ingredient_xref_allowed, +) +from kg_microbe.utils.producer_audits import verify_producer_audits +from kg_microbe.utils.source_finalization import SourceFinalizationRequired +from tests import test_mediadive_material_scope_audit as audit_tests +from tests import test_mim_conservative_refresh as refresh_tests +from tests.test_consolidate_chemical_mappings import _load_module +from tests.test_mediadive_material_scope_audit import _producer, _rows +from tests.test_mediadive_recipe_occurrences import _resolver +from tests.test_mim_conservative_refresh import FIELDS, _metadata, _table + +FIXTURE = Path(__file__).parent / "resources/mediadive/p3556_scope.json" +TARGET = "CHEBI:86658" +inputs = refresh_tests.inputs +bundle = refresh_tests.bundle + + +def _saved(): + """Bind the immutable original source, native structure and complete imported rows.""" + data = FIXTURE.read_bytes() + assert hashlib.sha256(data).hexdigest() == "7c06f4e2179356ddf1d917a6938cba5e3bdf6669f42a29019b93963e4a70646c" + return json.loads(data) + + +def _claims(table): + """Retain full columns from the selected immutable mapping excerpt.""" + return [item["row"] for item in _saved()["mapping_claims"] if item["table"] == table] + + +def _mapping_files(root, *, unified=None, supported=None): + """Write tiny synthetic snapshots, never a repository mapping export.""" + unified = _claims("unified") if unified is None else unified + supported = _claims("supported") if supported is None else supported + directory = root / "mappings" + directory.mkdir(exist_ok=True) + with gzip.open(directory / "kgmicrobe_unified_entity_mappings.sssom.tsv.gz", "wt", newline="") as stream: + writer = csv.DictWriter(stream, fieldnames=FIELDS, delimiter="\t", lineterminator="\n") + writer.writeheader() + writer.writerows(unified) + path = directory / "ingredient_mappings.sssom.tsv" + _table(path, tuple(_claims("supported")[0]), supported) + return path + + +@pytest.mark.parametrize( + "name", ["L-alpha-Phosphatidylcholine", "L-α-Phosphatidylcholine", " l ALPHA phosphatidylcholine "] +) +def test_generic_alias_is_not_native_specific_structure(name): + """The authority's generic display label must not defeat its explicit molecular structure.""" + assert not ingredient_mapping_allowed(name, TARGET) + assert not ingredient_xref_allowed(audit.P3556_MIM, TARGET) + assert not ingredient_xref_allowed(TARGET, audit.P3556_MIM) + for row in _claims("unified")[1:5]: + assert ingredient_mapping_allowed(row["subject_label"], TARGET) + assert ingredient_mapping_allowed(name, "CHEBI:16110") + assert ingredient_mapping_allowed(name, "CHEBI:49183") + assert ingredient_mapping_allowed(name, "cas:8002-43-5") + assert ingredient_xref_allowed("cas:8002-43-5", "CHEBI:16110") + + +def test_native_reader_retains_structure_but_not_generic_or_direct_identity(tmp_path, monkeypatch): + """Actual reader policies close both lexical and MIM paths without banning the molecule.""" + path = tmp_path / "fixture.tsv" + _table(path, FIELDS, _claims("unified"), _metadata()) + monkeypatch.setattr(mapping, "_LOADED", False) + monkeypatch.setattr(mapping, "_CACHED_PATH", None) + mapping.load_unified_mappings(path) + for name in ("L-alpha-Phosphatidylcholine", "L-α-Phosphatidylcholine"): + assert mapping.find_chebi_by_name(name) is None + assert mapping.find_chebi_by_xref(audit.P3556_MIM) is None + assert mapping.get_canonical_name(TARGET) == ingredient_authority_label(TARGET) + for row in _claims("unified")[1:5]: + assert mapping.find_chebi_by_name(row["subject_label"]) == TARGET + + +@pytest.mark.parametrize("target", [TARGET, "CHEBI:16110", "CHEBI:49183", "CAS-RN:8002-43-5", "PubChem:123"]) +@pytest.mark.parametrize("kind", ["compound", "solution"]) +def test_product_qualifier_precedes_every_name_route(target, kind): + """The whole product stays local even when the displayed name is structurally specific.""" + raw = dict(_saved()["raw"], compound=_claims("unified")[2]["subject_label"]) + if kind == "solution": + raw["solution_id"] = raw.pop("compound_id") + raw["solution"] = raw.pop("compound") + transform = _resolver({"629": {"recipe": [raw]}}) + calls = [] + transform.chemical_loader.find_chebi_by_name = lambda name: calls.append(name) or target + transform.compound_mappings[raw.get("compound", raw.get("solution")).lower()] = target + transform.compounds_data = { + "2082": {"ChEBI": "86658", "KEGG-Compound": "C1", "PubChem": 123, "CAS-RN": "8002-43-5"} + } + occurrence = transform.get_solution_recipe_occurrences("629")[0] + assert calls == [] + assert occurrence["id"] == f"mediadive.{'ingredient' if kind == 'compound' else 'solution'}:2082" + assert json.loads(occurrence["source_record"]) == raw + assert {key: occurrence[key] for key in ("amount", "unit", "g_l", "mmol_l")} == { + "amount": 5, + "unit": "mg", + "g_l": 5, + "mmol_l": None, + } + + +@pytest.mark.parametrize("qualifier", ["SIGMA P3556", " sigma p3556 "]) +def test_direct_call_uses_explicit_embedded_evidence_but_empty_occurrence_does_not(qualifier): + """A source ID never borrows another occurrence's material qualifier when one was supplied.""" + transform = _resolver({}) + transform.compounds_data = {"2082": {"attribute": qualifier, "ChEBI": "16110"}} + transform.chemical_loader.find_chebi_by_name = lambda _: "CHEBI:16110" + assert transform.standardize_compound_id("2082", "Generic") == "mediadive.ingredient:2082" + assert transform.standardize_compound_id("2082", "Generic", source_record={}) == "CHEBI:16110" + assert transform.standardize_compound_id("2082", "Generic", source_record={"attribute": "Other"}) == "CHEBI:16110" + + +@pytest.mark.parametrize("qualifier", [None, "", "SIGMA P35560", "P3556", "SIGMA P3556 alternative", "SIGMA P3555"]) +def test_no_product_or_id_only_hold(qualifier): + """Neither ingredient number nor a fuzzy catalog code imposes this finite scientific decision.""" + raw = dict(_saved()["raw"], attribute=qualifier) + transform = _resolver({"629": {"recipe": [raw]}}) + transform.chemical_loader.find_chebi_by_name = lambda _: "CHEBI:16110" + assert transform.get_solution_recipe_occurrences("629")[0]["id"] == "CHEBI:16110" + + +def test_actual_writer_preserves_original_observation_and_all_available_claims(tmp_path, monkeypatch): + """Audit originals are candidates, not invented raw CAS assertions or attempted lookups.""" + raw = _saved()["raw"] + producer, _ = _producer( + tmp_path, monkeypatch, recipes={"629": {"recipe": [raw]}}, legacy=False, selected="CHEBI:16110" + ) + supported = _mapping_files(tmp_path) + before = supported.read_bytes() + producer.run(show_status=False) + rows = _rows(producer) + # All seven unified rows carry the old generic canonical object label; + # preserve them as available claims even when the new policy rejects it. + assert len(rows) == 8 + assert Counter(row["candidate_route"] for row in rows) == {audit.UNIFIED_ROLE: 7, audit.SUPPORTED_ROLE: 1} + originals = {audit._json(row) for row in _claims("unified") + _claims("supported")} + for row in rows: + assert row["candidate_record"] in originals + assert json.loads(row["source_record"]) == raw + assert row["retained_target"] == "mediadive.ingredient:2082" + assert row["source_assertion_id"] == _saved()["source_assertion_id"] + assert row["reason"] == "whole_product_not_molecular_identity" + assert row["authority_uri"] == audit.P3556_AUTHORITY_URI + assert row["source_record_sha256"] == hashlib.sha256(row["source_record"].encode()).hexdigest() + assert "CAS-RN" not in json.loads(row["source_record"]) + with producer.output_edge_file.open() as stream: + # The producer's pandas deduplication still writes CSV-quoted fields; + # source finalization supplies the canonical literal TSV boundary. + edges = [row for row in csv.DictReader(stream, delimiter="\t") if row.get("source_assertion_id")] + assert len(edges) == 1 + edge = edges[0] + assert edge["object"] == "mediadive.ingredient:2082" + assert json.loads(edge["source_record"]) == raw + assert (edge["primary_knowledge_source"], edge["knowledge_level"], edge["agent_type"]) == ( + "infores:mediadive", + "observation", + "manual_agent", + ) + assert (edge["value"], edge["unit"], edge["g_l"], edge["mmol_l"]) == ("5.0", "mg", "5.0", "") + with producer.output_node_file.open() as stream: + node = next( + row + for row in csv.DictReader(stream, delimiter="\t", quoting=csv.QUOTE_NONE) + if row["id"] == "mediadive.ingredient:2082" + ) + assert node["category"] == "biolink:ChemicalEntity" + assert node["provided_by"] == "infores:mediadive" + assert supported.read_bytes() == before + verify_producer_audits(producer) + + +def test_actual_spelling_and_duplicate_candidate_rows_are_preserved(tmp_path, monkeypatch): + """A contradictory structural display still holds the product and retains its own original claims.""" + specific = _claims("unified")[2] + raw = dict(_saved()["raw"], compound=specific["subject_label"], custom=None, optional=False) + embedded = {"2082": dict(raw, ChEBI="86658", PubChem=0, **{"KEGG-Compound": "C1", "CAS-RN": "8002-43-5"})} + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [raw]}}, legacy=False, embedded=embedded) + _mapping_files(tmp_path, unified=[specific, specific, *_claims("unified")[-2:]]) + # Legacy optional selection happened at constructor-time: create it first in + # a separate test below; this test exercises selected unified and embedded inputs. + producer.run(show_status=False) + rows = _rows(producer) + assert Counter(row["candidate_route"] for row in rows) == { + audit.UNIFIED_ROLE: 2, + audit.SUPPORTED_ROLE: 1, + "mediadive_compounds": 4, + } + assert { + json.loads(row["candidate_record"])["subject_label"] + for row in rows + if row["candidate_route"] == audit.UNIFIED_ROLE + } == {specific["subject_label"]} + assert len({row["candidate_record_locator"] for row in rows if row["candidate_route"] == audit.UNIFIED_ROLE}) == 2 + assert all(json.loads(row["source_record"]) == raw for row in rows) + + +@pytest.mark.parametrize("removed", [False, True]) +def test_removed_unified_claim_keeps_original_supported_evidence_or_valid_empty_audit(tmp_path, monkeypatch, removed): + """Removing an identity claim is a normal state, not a reason to invent a candidate.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + _mapping_files(tmp_path, unified=[], supported=[] if removed else None) + producer.run(show_status=False) + assert len(_rows(producer)) == (0 if removed else 1) + assert audit.SUPPORTED_ROLE in producer.consumed_input_snapshots + verify_producer_audits(producer) + report = {"consumed_inputs": producer.consumed_input_snapshots} + audit.verify_recorded_material_inputs(report, producer.output_dir / "source_finalization.json") + del report["consumed_inputs"][audit.SUPPORTED_ROLE] + if not removed: + with pytest.raises(SourceFinalizationRequired, match="origin"): + audit.verify_recorded_material_inputs(report, producer.output_dir / "source_finalization.json") + + +def test_supported_origin_drift_and_repeated_visits_are_guarded(tmp_path, monkeypatch): + """The supported table is audit-only evidence with the same immutable-consumption guarantees.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + supported = _mapping_files(tmp_path) + producer.run(show_status=False) + path = producer.output_dir / audit.AUDIT_FILENAME + original = path.read_bytes() + producer.get_solution_recipe_occurrences("1") + producer._material_scope_audit.write() + assert path.read_bytes() == original + report = {"consumed_inputs": copy.deepcopy(producer.consumed_input_snapshots)} + report["consumed_inputs"][audit.SUPPORTED_ROLE]["path"] = str(tmp_path / "other.tsv") + with pytest.raises(SourceFinalizationRequired, match="origin"): + audit.verify_recorded_material_inputs(report, producer.output_dir / "source_finalization.json") + supported.write_text(supported.read_text() + "changed") + with pytest.raises(Exception, match="changed|drift|identity"): + producer._material_scope_audit.write() + + +def test_identity_refresh_keeps_explicit_structures_and_is_idempotent(tmp_path): + """The ordinary bounded refresher cannot reinstate generic aliases or the MIM identity pair.""" + module = _load_module() + source, first, second = (tmp_path / name for name in ("source.tsv", "first.tsv.gz", "second.tsv.gz")) + _table(source, FIELDS, _claims("unified"), _metadata()) + before = source.read_bytes() + result = module.refresh_identity_policy(source, first) + assert result["rows_removed"] == 3 + assert result["rows_relabelled"] == 4 + module.refresh_identity_policy(first, second) + assert first.read_bytes() == second.read_bytes() + with gzip.open(first, "rt") as stream: + rows = list(csv.DictReader((line for line in stream if not line.startswith("#")), delimiter="\t")) + expected = [dict(row, object_label=ingredient_authority_label(TARGET)) for row in _claims("unified")[1:5]] + assert rows == expected + assert source.read_bytes() == before + + +@pytest.mark.parametrize("name", ["L-alpha-Phosphatidylcholine", "L-α-Phosphatidylcholine"]) +def test_legacy_regeneration_does_not_restore_generic_alias(tmp_path, monkeypatch, name): + """Legacy import retains the native molecule but must not propagate its rejected generic label.""" + module = _load_module() + monkeypatch.setattr(module, "_build_mangle_blacklist", lambda *_: set()) + path = tmp_path / "compound_mappings_strict.tsv" + _table( + path, + ("original", "mapped", "chebi_label"), + [{"original": name, "mapped": TARGET, "chebi_label": "L-alpha-Phosphatidylcholine"}], + ) + consolidator = module.ChemicalMappingConsolidator() + consolidator.load_compound_mappings(path) + assert consolidator.chemicals[TARGET]["canonical_name"] == ingredient_authority_label(TARGET) + assert name not in consolidator.chemicals[TARGET]["synonyms"] + + +def test_conservative_refresh_losslessly_quarantines_original_generic_rows(inputs): + """Ordinary regeneration preserves full rejected claims and row multiplicity across cycles.""" + refresh = refresh_tests.refresh + held = [_claims("unified")[index] for index in (0, 5, 6)] + _table(inputs["baseline"], FIELDS, [*refresh._rows(inputs["baseline"]), *held, *held], _metadata()) + before = inputs["baseline"].read_bytes() + first = refresh.build_conservative_candidate(**inputs) + rows = [row for row in refresh_tests._read(first.quarantine_path) if row["object_id"] == TARGET] + assert len(rows) == 6 + assert {row.pop("quarantine_reason") for row in rows} == { + "reviewed_identity_policy_name", + "reviewed_identity_policy_xref", + } + assert Counter(audit._json(row) for row in rows) == Counter(audit._json(row) for row in held * 2) + assert all(row["object_id"] != TARGET for row in refresh_tests._read(first.candidate_path)) + second = refresh.build_conservative_candidate( + **dict(inputs, baseline=first.candidate_path, output_directory=inputs["output_directory"].with_name("second")) + ) + assert first.candidate_path.read_bytes() == second.candidate_path.read_bytes() + assert inputs["baseline"].read_bytes() == before + + +def test_mixed_profiles_legacy_candidates_and_full_payloads_do_not_cross(tmp_path, monkeypatch): + """Available legacy row multiplicity is retained only on its matching material occurrence.""" + original = audit_tests._write_mappings + + def write_with_legacy(root, **kwargs): + """Populate selected optional files before the producer chooses its immutable inputs.""" + result = original(root, **kwargs) + for filename in ( + audit_tests.mod.MICROMEDIAPARAM_COMPOUND_MAPPINGS_FILE, + audit_tests.mod.MICROMEDIAPARAM_HYDRATE_MAPPINGS_FILE, + ): + with (root / "raw" / filename).open("a") as stream: + for _ in range(2): + stream.write('L-α-Phosphatidylcholine\tCHEBI:16110\tsaved alternative\t"literal|pipe"\n') + return result + + monkeypatch.setattr(audit_tests, "_write_mappings", write_with_legacy) + raw = _saved()["raw"] + recipes = {"1": {"recipe": [raw, {"compound": "Potato", "compound_id": 1606}]}} + producer, _ = _producer(tmp_path, monkeypatch, recipes=recipes) + potato = json.loads(audit_tests.FIXTURE.read_text())["mapping_claims"]["unified"]["row"] + _mapping_files(tmp_path, unified=[*_claims("unified"), potato]) + producer.run(show_status=False) + rows = _rows(producer) + product = [row for row in rows if row["reason"] == "whole_product_not_molecular_identity"] + tuber = [row for row in rows if row["reason"] != "whole_product_not_molecular_identity"] + assert len(product) == 12 and len(tuber) == 3 + assert {row["candidate_target"] for row in tuber} == {"CAS-RN:93348-51-7", "cas:93348-51-7"} + assert {row["candidate_target"] for row in product} == {TARGET, "CHEBI:16110"} + assert Counter(row["candidate_route"] for row in product) == { + audit.UNIFIED_ROLE: 7, + audit.SUPPORTED_ROLE: 1, + "micromediaparam_hydrate": 2, + "micromediaparam_strict": 2, + } + assert all( + json.loads(row["candidate_record"])["extra"] == "literal|pipe" + for row in product + if row["candidate_route"].startswith("micromediaparam") + ) + + +@pytest.mark.parametrize("predicate", ["skos:broadMatch", "skos:narrowMatch", "rdfs:seeAlso"]) +def test_nonidentity_candidates_are_not_misreported_as_product_groundings(tmp_path, monkeypatch, predicate): + """A product hold is not permission to misclassify unrelated parent or annotation evidence.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + rows = [dict(row, predicate_id=predicate) for row in _claims("unified")] + _mapping_files(tmp_path, unified=rows, supported=[]) + producer.run(show_status=False) + assert _rows(producer) == [] + + +def test_supported_symlink_retarget_fails_even_for_same_bytes(tmp_path, monkeypatch): + """The selected canonical evidence locator cannot silently change its physical origin.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + path = _mapping_files(tmp_path) + one, two = tmp_path / "one.tsv", tmp_path / "two.tsv" + one.write_bytes(path.read_bytes()) + two.write_bytes(path.read_bytes()) + path.unlink() + path.symlink_to(one) + producer.run(show_status=False) + path.unlink() + path.symlink_to(two) + with pytest.raises(Exception, match="changed|binding|resolved|retarget"): + producer._material_scope_audit.write() + + +@pytest.mark.usefixtures("local_source_schema") +def test_p3556_finalization_repeat_and_public_admission_keep_original_audit(tmp_path, monkeypatch): + """Real native closure/finalization retains product-local rows and immutable candidate evidence.""" + from kg_microbe.utils.source_finalization import graph_rows, verify_finalized_source_files + from tests.test_mediadive_recipe_occurrences import _run_transform + + _mapping_files(tmp_path) + monkeypatch.setattr(audit, "_repo_root", lambda: tmp_path) + producer = _run_transform(tmp_path, monkeypatch, {"629": {"recipe": [_saved()["raw"]]}}) + path = producer.output_dir / audit.AUDIT_FILENAME + before = path.read_bytes() + identity = producer.producer_audit_snapshots[audit.AUDIT_FILENAME] + report = producer.finalize(fresh_run=True) + assert report["producer_audit_members"][audit.AUDIT_FILENAME] == identity + assert report["audit_members"][audit.AUDIT_FILENAME] == identity + assert audit.SUPPORTED_ROLE in report["consumed_inputs"] + assert path.read_bytes() == before + assert producer.finalize() == report + verify_finalized_source_files([producer.output_node_file, producer.output_edge_file]) + rows = [row for row in graph_rows(producer.output_edge_file) if row.get("source_assertion_id")] + assert len(rows) == 1 and rows[0]["object"] == "mediadive.ingredient:2082" + assert json.loads(rows[0]["source_record"]) == _saved()["raw"] + + +@pytest.mark.parametrize("route", ["identity", "attribute", "canonical_name", "synonym"]) +def test_canonical_object_label_candidate_matches_the_actual_reader(tmp_path, monkeypatch, route): + """All recognized primary routes can index object metadata, without a subject-label match (#1243).""" + raw = _saved()["raw"] + row = dict( + _claims("unified")[0], + subject_id="MIM:independent-material", + subject_label="Independent source spelling", + object_id="CHEBI:16110", + object_label=raw["compound"], + comment="", + ) + if route == "attribute": + row["subject_id"] = row["object_id"] + elif route in {"canonical_name", "synonym"}: + row.update(subject_id="kgm.name:independent-spelling", comment=route) + producer, path = _producer(tmp_path, monkeypatch, recipes={"629": {"recipe": [raw]}}, legacy=False) + _mapping_files(tmp_path, unified=[row], supported=[]) + monkeypatch.setattr(mapping, "_LOADED", False) + monkeypatch.setattr(mapping, "_CACHED_PATH", None) + mapping.load_unified_mappings(path) + assert mapping.find_chebi_by_name(raw["compound"]) == "CHEBI:16110" + producer.run(show_status=False) + rows = _rows(producer) + assert len(rows) == 1 + assert rows[0]["candidate_record"] == audit._json(row) + assert rows[0]["retained_target"] == "mediadive.ingredient:2082" + + +def test_subject_and_object_match_is_one_claim_per_original_ordinal(tmp_path, monkeypatch): + """One row with two applicable labels is not two claims; duplicate source rows still are.""" + raw = _saved()["raw"] + row = dict(_claims("unified")[-2], subject_label=raw["compound"], object_label=raw["compound"]) + producer, _ = _producer(tmp_path, monkeypatch, recipes={"629": {"recipe": [raw]}}, legacy=False) + _mapping_files(tmp_path, unified=[row, row], supported=[]) + producer.run(show_status=False) + rows = _rows(producer) + assert len(rows) == 2 + assert {item["candidate_record"] for item in rows} == {audit._json(row)} + assert len({item["candidate_record_locator"] for item in rows}) == 2 + + +def test_canonical_row_subject_only_is_not_a_reader_name_route(tmp_path, monkeypatch): + """Canonical-name rows use object_label; only synonym rows contribute subject_label.""" + raw = _saved()["raw"] + row = dict(_claims("unified")[-1], object_id="CHEBI:16110", object_label="Different material") + producer, path = _producer(tmp_path, monkeypatch, recipes={"629": {"recipe": [raw]}}, legacy=False) + _mapping_files(tmp_path, unified=[row], supported=[]) + monkeypatch.setattr(mapping, "_LOADED", False) + monkeypatch.setattr(mapping, "_CACHED_PATH", None) + mapping.load_unified_mappings(path) + assert mapping.find_chebi_by_name(raw["compound"]) is None + producer.run(show_status=False) + assert _rows(producer) == [] diff --git a/tests/test_mediadive_peptone_scope.py b/tests/test_mediadive_peptone_scope.py new file mode 100644 index 00000000..b1c14102 --- /dev/null +++ b/tests/test_mediadive_peptone_scope.py @@ -0,0 +1,367 @@ +"""Hold only two unsupported digest/CID identities while retaining complete observations (#1248).""" + +import copy +import csv +import gzip +import hashlib +import json +from collections import Counter +from pathlib import Path + +import pytest + +from kg_microbe.transform_utils.mediadive import material_scope_audit as audit +from kg_microbe.utils import chemical_mapping_utils as mapping +from kg_microbe.utils.ingredient_identity import ingredient_mapping_allowed, ingredient_xref_allowed +from kg_microbe.utils.producer_audits import verify_producer_audits +from kg_microbe.utils.source_finalization import SourceFinalizationRequired, graph_rows, verify_finalized_source_files +from tests import test_mediadive_material_scope_audit as audit_tests +from tests.test_consolidate_chemical_mappings import _load_module +from tests.test_mediadive_recipe_occurrences import _resolver, _run_transform +from tests.test_mim_conservative_refresh import FIELDS, _metadata, _row, _table + +FIXTURE = Path(__file__).parent / "resources/mediadive/peptone_scope.json" +TARGETS = ("pubchem.compound:167312541", "PubChem:167312541") +NAMES = ("Soy peptone", "Vitamin-free casamino acids") + + +def _saved(): + """Pin the complete original 26-occurrence and ten-claim excerpt.""" + payload = FIXTURE.read_bytes() + assert hashlib.sha256(payload).hexdigest() == "41a556f70cc0ae696d01895371214af7f25bb62c9d7a7b9a73c0fe589fed1fa3" + return json.loads(payload) + + +def _recipes(): + """Reconstruct exact original array positions with explicitly synthetic noningredient gaps.""" + result = {} + for occurrence in _saved()["occurrences"]: + solution, position = occurrence["source_assertion_id"].split(":", 1)[1].split("#recipe/") + result[solution] = { + "recipe": [{"instruction": "synthetic position placeholder"}] * (int(position) - 1) + [occurrence["raw"]] + } + return result + + +def _unified_rows(target=TARGETS[0]): + """Declare one synthetic target and two stale aliases without endorsing their identity.""" + return [ + _row( + "kgm.name:authority", + target, + audit.PEPTONE_AUTHORITY_LABEL, + name=audit.PEPTONE_AUTHORITY_LABEL, + comment="canonical_name", + ), + *[ + _row( + "kgm.name:" + str(index), + target, + audit.PEPTONE_AUTHORITY_LABEL, + name=name, + comment="synonym", + predicate="skos:closeMatch", + ) + for index, name in enumerate(NAMES) + ], + ] + + +def _mapping_files(root, rows): + """Write only tiny test-owned unified claims, not the real mapping artifact.""" + directory = root / "mappings" + directory.mkdir(exist_ok=True) + path = directory / "kgmicrobe_unified_entity_mappings.sssom.tsv.gz" + with gzip.open(path, "wt", encoding="utf-8", newline="") as stream: + writer = csv.DictWriter(stream, fieldnames=FIELDS, delimiter="\t", lineterminator="\n") + writer.writeheader() + writer.writerows(rows) + return path + + +def _producer(tmp_path, monkeypatch, *, recipes=None, unified=(), legacy=True, **kwargs): + """Install full immutable legacy excerpts before the real transform consumes its inputs.""" + original_writer = audit_tests._write_mappings + + def write(root, **ignored): + """Keep complete original columns/physical rows, including duplicate-name claims.""" + original_writer(root, unified=0, legacy=False) + path = _mapping_files(root, unified) + if legacy: + claims = _saved()["mapping_claims"] + for name in ("compound_mappings_strict.tsv", "compound_mappings_strict_hydrate.tsv"): + selected = [item for item in claims if item["input"].endswith("/" + name)] + (root / "raw" / name).write_bytes( + ( + selected[0]["header_original_line_utf8"] + + "".join(item["original_line_utf8"] for item in selected) + ).encode() + ) + return path + + monkeypatch.setattr(audit_tests, "_write_mappings", write) + return audit_tests._producer( + tmp_path, monkeypatch, recipes=_recipes() if recipes is None else recipes, legacy=False, **kwargs + ) + + +@pytest.mark.parametrize("target", TARGETS) +@pytest.mark.parametrize( + "name", [*NAMES, " SOY_PEPTONE ", "Soy (peptone)", "Vitamin_free_casamino_acids", " vitamin free casamino acids "] +) +def test_two_name_scopes_hold_all_existing_policy_normalizations(name, target): + """Legacy namespace and normalization variants cannot resurrect the exact reviewed claims.""" + assert not ingredient_mapping_allowed(name, target) + assert audit._peptone(name) + assert ingredient_mapping_allowed(name, "mediadive.ingredient:independent-control") + + +@pytest.mark.parametrize( + "name", + [ + "Peptone", + "Soytone", + "Soy protein", + "Casein", + "Casamino acids", + "Soy peptone extract", + "Vitamin-free casamino acids solution", + *TARGETS, + audit.PEPTONE_AUTHORITY_LABEL, + ], +) +def test_unreviewed_names_and_direct_identifiers_are_not_globally_banned(name): + """Neither a CID-wide blacklist nor a substring-wide material policy is introduced.""" + assert ingredient_mapping_allowed(name, TARGETS[0]) + assert not audit._peptone(name) + assert ingredient_xref_allowed(TARGETS[1], TARGETS[0]) + + +@pytest.mark.parametrize("target", TARGETS) +@pytest.mark.parametrize("synonyms", [True, False]) +def test_actual_cold_reader_rejects_stale_names_and_retains_target(tmp_path, monkeypatch, target, synonyms): + """Both indexes enforce finite exclusions without deleting the declared structure.""" + path = tmp_path / "tiny.tsv" + _table(path, FIELDS, _unified_rows(target), _metadata()) + monkeypatch.setattr(mapping, "_LOADED", False) + monkeypatch.setattr(mapping, "_CACHED_PATH", None) + mapping.load_unified_mappings(path) + assert all(mapping.find_chebi_by_name(name, synonyms=synonyms) is None for name in NAMES) + assert mapping.find_chebi_by_name(audit.PEPTONE_AUTHORITY_LABEL, synonyms=synonyms) == target + assert mapping.get_canonical_name(target) == audit.PEPTONE_AUTHORITY_LABEL + + +@pytest.mark.parametrize("name", NAMES) +@pytest.mark.parametrize("target", TARGETS) +@pytest.mark.parametrize("route", ["unified", "legacy", "embedded"]) +def test_every_compound_fallback_and_nested_solution_keeps_local(name, target, route): + """All old fallback entry points consult the policy, without a new hardcoded material guard.""" + raw = {"compound": name, "compound_id": 75, "amount": 0, "unit": "g", "g_l": 0, "mmol_l": None} + nested = {"solution": name, "solution_id": 17, "amount": 2, "unit": "ml", "note": False} + resolver = _resolver({"1": {"recipe": [raw, nested]}}) + if route == "unified": + resolver.chemical_loader.find_chebi_by_name = lambda query: target + elif route == "legacy": + resolver.compound_mappings = {name.lower(): target} + else: + # The embedded schema's key is PubChem; its value is the unprefixed CID. + resolver.compounds_data = {"75": {"compound": name, "PubChem": 167312541}} + actual = resolver.get_solution_recipe_occurrences("1") + assert [item["id"] for item in actual] == ["mediadive.ingredient:75", "mediadive.solution:17"] + assert [json.loads(item["source_record"]) for item in actual] == [raw, nested] + assert actual[0]["amount"] == actual[0]["g_l"] == 0 and actual[0]["mmol_l"] is None + assert resolver.standardize_compound_id("75", name, source_record=raw) == "mediadive.ingredient:75" + + +@pytest.mark.parametrize("name", [*TARGETS, "Unrelated material"]) +def test_explicit_cid_and_other_name_fallbacks_remain_unchanged(name): + """An available unrelated route is not suppressed by the two-name source audit profile.""" + resolver = _resolver({}) + resolver.compound_mappings = {name.lower(): TARGETS[1]} + assert resolver.standardize_compound_id("7", name) == TARGETS[1] + + +def test_all26_occurrences_and_all10_legacy_claims_survive(tmp_path, monkeypatch): + """Eight original soy claims per use and two vitamin-free claims remain losslessly auditable.""" + producer, _ = _producer(tmp_path, monkeypatch) + producer.run(show_status=False) + rows = audit_tests._rows(producer) + assert len(rows) == 202 + expected = {item["source_assertion_id"]: item["raw"] for item in _saved()["occurrences"]} + claim_rows = [item["row"] for item in _saved()["mapping_claims"]] + assert Counter(row["retained_target"] for row in rows) == { + "mediadive.ingredient:75": 200, + "mediadive.ingredient:654": 2, + } + assert {row["reason"] for row in rows} == {audit.PEPTONE_REASON} + for row in rows: + assert json.loads(row["source_record"]) == expected[row["source_assertion_id"]] + assert json.loads(row["candidate_record"]) in claim_rows + assert row["candidate_target"] == TARGETS[1] + assert row["source_record_sha256"] == hashlib.sha256(row["source_record"].encode()).hexdigest() + assert json.loads(row["qualifier_evidence"]) == { + key: expected[row["source_assertion_id"]][key] + for key in ("condition", "attribute") + if key in expected[row["source_assertion_id"]] + } + with producer.output_edge_file.open() as stream: + edges = [row for row in csv.DictReader(stream, delimiter="\t") if row.get("source_assertion_id")] + assert len(edges) == 26 + for edge in edges: + raw = expected[edge["source_assertion_id"]] + assert json.loads(edge["source_record"]) == raw + assert edge["object"] == "mediadive.ingredient:" + str(raw["compound_id"]) + original = next( + item for item in _saved()["occurrences"] if item["source_assertion_id"] == edge["source_assertion_id"] + ) + assert edge == dict(original["observed_edge_rows"][0]["row"], object=edge["object"]) + before = (producer.output_dir / audit.AUDIT_FILENAME).read_bytes() + for identifier in _recipes(): + producer.get_solution_recipe_occurrences(identifier) + producer._material_scope_audit.write() + assert (producer.output_dir / audit.AUDIT_FILENAME).read_bytes() == before + verify_producer_audits(producer) + + +@pytest.mark.parametrize("route", ["identity", "attribute", "canonical_name", "synonym"]) +@pytest.mark.parametrize("target", TARGETS) +def test_audit_collects_only_reader_eligible_held_target_claims(tmp_path, monkeypatch, route, target): + """Match actual primary-label indexing, not all imported annotations or other targets.""" + row = _row("MIM:independent", target, NAMES[0], name="Other subject spelling") + if route == "attribute": + row["subject_id"] = target + elif route in {"canonical_name", "synonym"}: + row.update(subject_id="kgm.name:alias", comment=route) + if route == "synonym": + row.update(subject_label=NAMES[0], object_label=audit.PEPTONE_AUTHORITY_LABEL) + unrelated = dict(row, object_id="CHEBI:16110") + weak = dict(row, predicate_id="skos:broadMatch") + producer, _ = _producer( + tmp_path, + monkeypatch, + recipes={"1": {"recipe": [{"compound": NAMES[0], "compound_id": 75}]}}, + unified=[row, row, unrelated, weak], + legacy=False, + ) + producer.run(show_status=False) + rows = audit_tests._rows(producer) + assert len(rows) == 2 and len({item["candidate_record_locator"] for item in rows}) == 2 + assert all(json.loads(item["candidate_record"]) == row for item in rows) + + +def test_embedded_target_only_claim_and_independent_replacement_are_preserved(tmp_path, monkeypatch): + """Do not quarantine an unrelated embedded claim or insist on a local target unnecessarily.""" + embedded = {"compound": NAMES[0], "PubChem": 167312541, "ChEBI": 16110, "note": None, "flag": False} + producer, _ = _producer( + tmp_path, + monkeypatch, + recipes={"1": {"recipe": [{"compound": NAMES[0], "compound_id": 75}]}}, + legacy=False, + embedded={"75": embedded}, + selected="mediadive.ingredient:reviewed-control", + ) + producer.run(show_status=False) + rows = audit_tests._rows(producer) + assert len(rows) == 1 and rows[0]["candidate_target"] == TARGETS[1] + assert json.loads(rows[0]["candidate_record"]) == embedded + assert rows[0]["retained_target"] == "mediadive.ingredient:reviewed-control" + + +def test_repeated_raw_occurrences_are_not_collapsed(tmp_path, monkeypatch): + """Equal payloads and quantities remain different original source positions.""" + raw = {"compound": NAMES[0], "compound_id": 75, "amount": 0, "unit": "g", "g_l": 0, "mmol_l": None} + producer, _ = _producer( + tmp_path, + monkeypatch, + recipes={"1": {"recipe": [raw, copy.deepcopy(raw)]}}, + unified=_unified_rows(), + legacy=False, + ) + producer.run(show_status=False) + rows = audit_tests._rows(producer) + assert len(rows) == 2 and {row["source_assertion_id"] for row in rows} == { + "mediadive.solution:1#recipe/1", + "mediadive.solution:1#recipe/2", + } + with producer.output_edge_file.open() as stream: + edges = [row for row in csv.DictReader(stream, delimiter="\t") if row.get("source_assertion_id")] + assert len(edges) == 2 and all(json.loads(row["source_record"]) == raw for row in edges) + assert all(float(row["value"]) == float(row["g_l"]) == 0 and row["mmol_l"] == "" for row in edges) + + +def test_no_current_held_claim_does_not_invent_an_audit_row(tmp_path, monkeypatch): + """A reviewed name is not itself evidence that a quarantined mapping still exists.""" + producer, _ = _producer( + tmp_path, + monkeypatch, + recipes={"1": {"recipe": [{"compound": NAMES[0], "compound_id": 75}]}}, + unified=[], + legacy=False, + ) + producer.run(show_status=False) + assert audit_tests._rows(producer) == [] + verify_producer_audits(producer) + + +@pytest.mark.parametrize("damage", ["missing_rule", "missing_origin", "wrong_origin", "tampered_audit"]) +def test_required_policy_and_audit_evidence_never_get_restamped(tmp_path, monkeypatch, damage): + """The same producer/finalizer authority rules apply to the new finite profile.""" + producer, _ = _producer(tmp_path, monkeypatch) + if damage == "missing_rule": + policy = tmp_path / "policy.tsv" + lines = audit.IDENTITY_POLICY.read_text().splitlines(keepends=True) + policy.write_text("".join(line for line in lines if "^soy[ _-]+peptone$" not in line)) + monkeypatch.setattr(audit, "IDENTITY_POLICY", policy) + with pytest.raises(SourceFinalizationRequired, match="holds are missing"): + producer.run(show_status=False) + assert not producer.output_edge_file.exists() + return + producer.run(show_status=False) + report = {"consumed_inputs": copy.deepcopy(producer.consumed_input_snapshots)} + report_path = producer.output_dir / "source_finalization.json" + audit.verify_recorded_material_inputs(report, report_path) + if damage == "tampered_audit": + (producer.output_dir / audit.AUDIT_FILENAME).write_text("replaced") + with pytest.raises(SourceFinalizationRequired): + verify_producer_audits(producer) + return + if damage == "missing_origin": + del report["consumed_inputs"][audit.POLICY_ROLE] + else: + report["consumed_inputs"][audit.POLICY_ROLE]["path"] = str(tmp_path / "wrong-policy.tsv") + with pytest.raises(SourceFinalizationRequired, match="input origin"): + audit.verify_recorded_material_inputs(report, report_path) + + +@pytest.mark.usefixtures("local_source_schema") +def test_real_finalization_keeps_raw_rows_quantities_and_bound_audit(tmp_path, monkeypatch): + """Finalization and public admission retain distinct observations and original audit bytes.""" + _mapping_files(tmp_path, _unified_rows()) + monkeypatch.setattr(audit, "_repo_root", lambda: tmp_path) + producer = _run_transform(tmp_path, monkeypatch, _recipes()) + path = producer.output_dir / audit.AUDIT_FILENAME + before, identity = path.read_bytes(), producer.producer_audit_snapshots[audit.AUDIT_FILENAME] + report = producer.finalize(fresh_run=True) + assert report["producer_audit_members"][audit.AUDIT_FILENAME] == identity + assert report["audit_members"][audit.AUDIT_FILENAME] == identity + assert path.read_bytes() == before and producer.finalize() == report + verify_finalized_source_files([producer.output_node_file, producer.output_edge_file]) + assert {row["object"] for row in graph_rows(producer.output_edge_file) if row.get("source_assertion_id")} == { + "mediadive.ingredient:75", + "mediadive.ingredient:654", + } + + +def test_tiny_identity_export_is_stable_and_keeps_unrelated_direct_rows(tmp_path, monkeypatch): + """The existing bounded writer enforces new names without rewriting supported MIM or a real export.""" + module = _load_module() + metadata = _metadata() + metadata["curie_map"]["pubchem.compound"] = "https://pubchem.ncbi.nlm.nih.gov/compound/" + rows = [*_unified_rows(), _row("MIM:unrelated", TARGETS[0], audit.PEPTONE_AUTHORITY_LABEL)] + source, first, second = (tmp_path / name for name in ("source.tsv", "first.tsv.gz", "second.tsv.gz")) + _table(source, FIELDS, rows, metadata) + before = source.read_bytes() + result = module.refresh_identity_policy(source, first) + assert result == {"rows_read": 4, "rows_removed": 2, "rows_relabelled": 0} + assert module.refresh_identity_policy(first, second) == {"rows_read": 2, "rows_removed": 0, "rows_relabelled": 0} + assert first.read_bytes() == second.read_bytes() and source.read_bytes() == before diff --git a/tests/test_mediadive_sugar_context.py b/tests/test_mediadive_sugar_context.py new file mode 100644 index 00000000..3adb4341 --- /dev/null +++ b/tests/test_mediadive_sugar_context.py @@ -0,0 +1,428 @@ +"""Preserve the explicitly qualified mixed-sugar stock without changing generic Sugar (#1245).""" + +import copy +import csv +import gzip +import hashlib +import json +from collections import Counter +from pathlib import Path + +import pytest + +from kg_microbe.transform_utils.mediadive import material_scope_audit as audit +from kg_microbe.utils import chemical_mapping_utils as mapping +from kg_microbe.utils.producer_audits import verify_producer_audits +from kg_microbe.utils.source_finalization import SourceFinalizationRequired +from tests import test_mediadive_material_scope_audit as audit_tests +from tests.test_mediadive_material_scope_audit import _producer, _rows +from tests.test_mediadive_recipe_occurrences import _resolver + +FIXTURE = Path(__file__).parent / "resources/mediadive/sugar_context.json" +POLICY = Path(__file__).resolve().parents[1] / audit.CONTEXT_POLICY +TARGET = "NCIT:C71939" + + +def _saved(): + """Read exact saved raw and full imported evidence, not a live endpoint or mapping table.""" + payload = FIXTURE.read_bytes() + assert hashlib.sha256(payload).hexdigest() == "09784226fecea0f2b8fc5b3d80748c5a58031c0aa3398e60c292c53a2ff9202f" + return json.loads(payload) + + +def _claims(table): + """Select full immutable native mapping rows without interpreting identity.""" + return [item["row"] for item in _saved()["mapping_claims"] if item["table"] == table] + + +def _files(root, *, unified=None, supported=None): + """Create ordinary tiny selected inputs at canonical locations in the test repository.""" + directory = root / "mappings" + directory.mkdir(exist_ok=True) + context = root / audit.CONTEXT_POLICY + context.parent.mkdir(exist_ok=True) + context.write_bytes(POLICY.read_bytes()) + for table, filename, data in ( + ( + "unified", + "kgmicrobe_unified_entity_mappings.sssom.tsv.gz", + _claims("unified") if unified is None else unified, + ), + ("supported", "ingredient_mappings.sssom.tsv", _claims("supported") if supported is None else supported), + ): + path = directory / filename + opener = gzip.open if path.suffix == ".gz" else open + with opener(path, "wt", encoding="utf-8", newline="") as stream: + writer = csv.DictWriter(stream, tuple(_claims(table)[0]), delimiter="\t", lineterminator="\n") + writer.writeheader() + writer.writerows(data) + return context + + +@pytest.mark.parametrize("kind", ["compound", "solution"]) +@pytest.mark.parametrize("target", [TARGET, "CHEBI:17992", "CAS-RN:57-50-1", "KEGG:C00089", "PubChem:5988"]) +def test_reviewed_context_wins_all_candidate_targets_without_inventing_components(kind, target): + """Source context wins before all candidates, retaining original quantity and no inferred components.""" + raw = _saved()["raw"] + if kind == "solution": + raw["solution"] = raw.pop("compound") + raw["solution_id"] = raw.pop("compound_id") + transform = _resolver({"4338": {"recipe": [raw]}}) + calls = [] + transform.chemical_loader.find_chebi_by_name = lambda name: calls.append(name) or target + transform.compound_mappings = {"sugar": target} + transform.compounds_data = { + "1795": {"ChEBI": "17992", "KEGG-Compound": "C00089", "PubChem": 5988, "CAS-RN": "57-50-1"} + } + actual = transform.get_solution_recipe_occurrences("4338") + assert len(actual) == 1 and calls == [] + assert actual[0]["id"] == f"mediadive.{'ingredient' if kind == 'compound' else 'solution'}:1795" + assert (actual[0]["amount"], actual[0]["unit"], actual[0]["g_l"], actual[0]["mmol_l"]) == (0.1, "ml", None, None) + assert json.loads(actual[0]["source_record"]) == raw + + +@pytest.mark.parametrize( + "name,attribute,held", + [ + ("Sugar", audit.SUGAR_ATTRIBUTE, True), + (" SUGAR ", " 250 mM each of xylose, maltose AND cellobiose ", True), + ("Sugar", None, False), + ("Sugar", "", False), + ("Sugar", "250 mM xylose", False), + ("Sugar", audit.SUGAR_ATTRIBUTE + ".", False), + ("S.u.g.a.r", audit.SUGAR_ATTRIBUTE, False), + ("Sucrose", audit.SUGAR_ATTRIBUTE, False), + ("Food Sugar", audit.SUGAR_ATTRIBUTE, False), + ], +) +def test_exact_case_whitespace_scope_not_ingredient_id_or_generic_name(name, attribute, held): + """Other qualifiers and spellings remain ordinary lookup routes, including unrelated food context.""" + raw = dict(_saved()["raw"], compound=name, attribute=attribute) + transform = _resolver({"1": {"recipe": [raw]}}) + transform.chemical_loader.find_chebi_by_name = lambda _: TARGET + expected = "mediadive.ingredient:1795" if held else TARGET + assert transform.get_solution_recipe_occurrences("1")[0]["id"] == expected + + +def test_direct_call_does_not_borrow_embedded_qualifier_when_occurrence_supplied(): + """Explicit empty and differently qualified occurrence evidence outranks cached compound context.""" + transform = _resolver({}) + transform.compounds_data = {"1795": dict(_saved()["raw"], ChEBI="17992")} + transform.chemical_loader.find_chebi_by_name = lambda _: TARGET + assert transform.standardize_compound_id("1795", "Sugar") == "mediadive.ingredient:1795" + for raw in ({}, {"compound": "Sugar"}, {"compound": "Sugar", "attribute": "food sweetener"}): + assert transform.standardize_compound_id("1795", "Sugar", source_record=raw) == TARGET + + +def test_actual_writer_binds_new_policy_and_full_original_claims(tmp_path, monkeypatch): + """Preserve the reviewed raw-list ordinal, full candidates, source amount and seven-field profile.""" + raw = _saved()["raw"] + recipe = [{"instruction": "synthetic ordinal placeholder"}] * 15 + [raw] + producer, _ = _producer(tmp_path, monkeypatch, recipes={"4338": {"recipe": recipe}}, legacy=False, selected=TARGET) + policy = _files(tmp_path) + supported = tmp_path / "mappings/ingredient_mappings.sssom.tsv" + supported_before = supported.read_bytes() + producer.run(show_status=False) + actual = _rows(producer) + assert len(actual) == 3 + originals = Counter(audit._json(row) for row in _claims("unified") + _claims("supported")) + assert Counter(row["candidate_record"] for row in actual) == originals + assert audit.POLICY_ROLE not in producer.consumed_input_snapshots + for row in actual: + assert row["policy_input_path"] == str(policy.resolve()) + assert row["policy_input_sha256"] == hashlib.sha256(policy.read_bytes()).hexdigest() + assert row["source_assertion_id"] == _saved()["source_assertion_id"] + assert json.loads(row["source_record"]) == raw and "CAS-RN" not in raw + assert row["retained_target"] == "mediadive.ingredient:1795" + assert row["reason"] == audit.SUGAR_REASON and row["authority_uri"] == audit.SUGAR_URI + with producer.output_edge_file.open() as stream: + edges = [row for row in csv.DictReader(stream, delimiter="\t") if row.get("source_assertion_id")] + assert len(edges) == 1 + edge = edges[0] + assert ( + edge["subject"], + edge["predicate"], + edge["object"], + edge["relation"], + edge["primary_knowledge_source"], + edge["knowledge_level"], + edge["agent_type"], + ) == ( + "mediadive.solution:4338", + "biolink:has_part", + "mediadive.ingredient:1795", + "BFO:0000051", + "infores:mediadive", + "observation", + "manual_agent", + ) + assert (edge["value"], edge["unit"], edge["g_l"], edge["mmol_l"]) == ("0.1", "ml", "", "") + with producer.output_node_file.open() as stream: + nodes = list(csv.DictReader(stream, delimiter="\t", quoting=csv.QUOTE_NONE)) + node = next(row for row in nodes if row["id"] == "mediadive.ingredient:1795") + assert (node["name"], node["category"], node["provided_by"]) == ( + "Sugar", + "biolink:ChemicalEntity", + "infores:mediadive", + ) + assert not node["xref"] and not node["synonym"] + assert supported.read_bytes() == supported_before + verify_producer_audits(producer) + + +def test_actual_reader_generic_sugar_identity_still_available(tmp_path, monkeypatch): + """The source-context fix is not a global NCIT/Sugar or immutable MIM exclusion.""" + _files(tmp_path) + monkeypatch.setattr(mapping, "_LOADED", False) + monkeypatch.setattr(mapping, "_CACHED_PATH", None) + mapping.load_unified_mappings(tmp_path / "mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz") + assert mapping.find_chebi_by_name("Sugar") == TARGET + assert mapping.find_chebi_by_xref("MIM:Sugar") == TARGET + + +@pytest.mark.parametrize("removed", [False, True]) +def test_removed_candidates_keep_original_supported_or_header_only(tmp_path, monkeypatch, removed): + """Absent candidates do not erase selected-context policy or invent grounded identities.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + _files(tmp_path, unified=[], supported=[] if removed else None) + producer.run(show_status=False) + assert len(_rows(producer)) == (0 if removed else 1) + assert audit.CONTEXT_ROLE in producer.consumed_input_snapshots + assert audit.SUPPORTED_ROLE in producer.consumed_input_snapshots + verify_producer_audits(producer) + + +@pytest.mark.parametrize( + "damage", + [ + "missing", + "duplicate", + "source_name", + "source_attribute", + "disposition", + "reason", + "evidence_uri", + "withheld_target_example", + ], +) +def test_missing_or_changed_selected_context_policy_fails_before_output(tmp_path, monkeypatch, damage): + """A missing reviewed decision is never silently adopted as a new source rule.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + policy = _files(tmp_path) + if damage == "missing": + policy.unlink() + else: + with policy.open() as stream: + reader = csv.DictReader(stream, delimiter="\t") + fields, rows = reader.fieldnames, list(reader) + if damage == "duplicate": + rows.append(dict(rows[0])) + else: + rows[0][damage] = "wrong" + with policy.open("w", newline="") as stream: + writer = csv.DictWriter(stream, fields, delimiter="\t", lineterminator="\n") + writer.writeheader() + writer.writerows(rows) + with pytest.raises((FileNotFoundError, SourceFinalizationRequired, ValueError), match="Sugar|Missing|missing"): + producer.run(show_status=False) + assert not producer.output_edge_file.exists() + + +def test_context_drift_and_same_byte_symlink_retarget_cannot_publish_audit(tmp_path, monkeypatch): + """Selected origin and bytes remain guarded after the exact record has been resolved.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + policy = _files(tmp_path) + first, second = tmp_path / "first.tsv", tmp_path / "second.tsv" + first.write_bytes(policy.read_bytes()) + second.write_bytes(policy.read_bytes()) + policy.unlink() + policy.symlink_to(first) + producer.run(show_status=False) + policy.unlink() + policy.symlink_to(second) + with pytest.raises(Exception, match="changed|binding|resolved|retarget"): + producer._material_scope_audit.write() + + +@pytest.mark.parametrize("role", [audit.CONTEXT_ROLE, audit.SUPPORTED_ROLE]) +def test_public_origin_verifier_requires_actual_policy_and_supported_roles(tmp_path, monkeypatch, role): + """A generic ingredient policy or wrong MIM locator cannot masquerade as the context decision.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + _files(tmp_path) + producer.run(show_status=False) + report = {"consumed_inputs": copy.deepcopy(producer.consumed_input_snapshots)} + report["consumed_inputs"][role]["path"] = "/wrong/context-or-supported-file" + with pytest.raises(SourceFinalizationRequired, match="origin"): + audit.verify_recorded_material_inputs(report, producer.output_dir / "source_finalization.json") + + +@pytest.mark.usefixtures("local_source_schema") +def test_real_finalization_repeat_and_public_admission_preserve_context_audit(tmp_path, monkeypatch): + """Validate a tiny actual graph and finalized audit through normal admission, without live data.""" + from kg_microbe.utils.source_finalization import graph_rows, verify_finalized_source_files + from tests.test_mediadive_recipe_occurrences import _run_transform + + policy = _files(tmp_path) + monkeypatch.setattr(audit, "_repo_root", lambda: tmp_path) + producer = _run_transform(tmp_path, monkeypatch, {"4338": {"recipe": [_saved()["raw"]]}}) + before = (producer.output_dir / audit.AUDIT_FILENAME).read_bytes() + report = producer.finalize(fresh_run=True) + assert report["consumed_inputs"][audit.CONTEXT_ROLE]["path"] == str(policy.resolve()) + assert producer.finalize() == report + verify_finalized_source_files([producer.output_node_file, producer.output_edge_file]) + assert (producer.output_dir / audit.AUDIT_FILENAME).read_bytes() == before + edges = [row for row in graph_rows(producer.output_edge_file) if row.get("source_assertion_id")] + assert len(edges) == 1 and edges[0]["object"] == "mediadive.ingredient:1795" + assert json.loads(edges[0]["source_record"]) == _saved()["raw"] + + +@pytest.mark.parametrize("route", ["unified", "legacy", "ChEBI", "KEGG-Compound", "PubChem", "CAS-RN"]) +@pytest.mark.parametrize("held", [True, False]) +def test_each_original_fallback_is_protected_only_for_qualified_occurrence(route, held): + """An alternative namespace cannot evade context, while the original ordinary fallback stays available.""" + transform = _resolver({}) + transform.compounds_data = {"1795": dict(_saved()["raw"])} + if route == "unified": + transform.chemical_loader.find_chebi_by_name = lambda _: TARGET + expected = TARGET + elif route == "legacy": + transform.compound_mappings["sugar"] = TARGET + expected = TARGET + else: + literal, expected = { + "ChEBI": ("17992", "CHEBI:17992"), + "KEGG-Compound": ("C00089", "KEGG:C00089"), + "PubChem": (5988, "PubChem:5988"), + "CAS-RN": ("57-50-1", "CAS-RN:57-50-1"), + }[route] + transform.compounds_data["1795"][route] = literal + actual = transform.standardize_compound_id("1795", "Sugar", source_record=_saved()["raw"] if held else {}) + assert actual == ("mediadive.ingredient:1795" if held else expected) + + +def test_mixed_occurrences_and_duplicate_original_candidate_rows_stay_separate(tmp_path, monkeypatch): + """Same compound ID and repeated recipe order cannot transfer qualifiers or merge source records.""" + original_write = audit_tests._write_mappings + + def selected_legacy(root, **kwargs): + """Install duplicate original legacy rows before ordinary optional-input selection.""" + path = original_write(root, **kwargs) + for filename in ( + audit_tests.mod.MICROMEDIAPARAM_COMPOUND_MAPPINGS_FILE, + audit_tests.mod.MICROMEDIAPARAM_HYDRATE_MAPPINGS_FILE, + ): + with (root / "raw" / filename).open("a") as stream: + stream.write('Sugar\tNCIT:C71939\toriginal legacy\t"full|claim"\n' * 2) + return path + + monkeypatch.setattr(audit_tests, "_write_mappings", selected_legacy) + qualified = _saved()["raw"] + unqualified = dict(qualified) + unqualified.pop("attribute") + another = dict(qualified, amount=0, optional=False, nullable=None) + producer, _ = _producer( + tmp_path, + monkeypatch, + recipes={"1": {"recipe": [qualified, unqualified, another]}}, + embedded={"1795": dict(qualified, **{"CAS-RN": "57-50-1"})}, + selected=TARGET, + ) + _files(tmp_path, unified=[*_claims("unified"), _claims("unified")[0]]) + producer.run(show_status=False) + rows = _rows(producer) + assert len(rows) == 18 # 3 unified + 1 supported + 4 legacy + 1 real embedded, on two occurrences. + assert Counter(row["source_assertion_id"] for row in rows) == { + "mediadive.solution:1#recipe/1": 9, + "mediadive.solution:1#recipe/3": 9, + } + assert all(row["retained_target"] == "mediadive.ingredient:1795" for row in rows) + assert ( + len({(row["source_assertion_id"], row["candidate_route"], row["candidate_record_locator"]) for row in rows}) + == 18 + ) + assert all("CAS-RN" not in json.loads(row["source_record"]) for row in rows) + assert { + json.loads(row["candidate_record"])["extra"] for row in rows if row["candidate_route"].startswith("micromedia") + } == {"full|claim"} + occurrences = producer.get_solution_recipe_occurrences("1") + assert [row["id"] for row in occurrences] == ["mediadive.ingredient:1795", TARGET, "mediadive.ingredient:1795"] + assert json.loads(occurrences[-1]["source_record"]) == another + + +@pytest.mark.parametrize("role", [audit.CONTEXT_ROLE, audit.SUPPORTED_ROLE]) +def test_header_only_selected_context_cannot_drop_required_origin(tmp_path, monkeypatch, role): + """No remaining imported claim still requires the producer's actual selected decision input.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + _files(tmp_path, unified=[], supported=[]) + producer.run(show_status=False) + assert _rows(producer) == [] + snapshots = copy.deepcopy(producer.consumed_input_snapshots) + report = {"consumed_inputs": snapshots, "inputs": list(copy.deepcopy(snapshots).values())} + snapshots.pop(role) + with pytest.raises(SourceFinalizationRequired, match="origin"): + audit.verify_recorded_material_inputs(report, producer.output_dir / "source_finalization.json") + + +def test_unrelated_context_does_not_require_sugar_decision_or_supported_input(tmp_path, monkeypatch): + """Ordinary unqualified Sugar does not acquire a source-specific hold or new audit dependency.""" + raw = {"compound": "Sugar", "compound_id": 1795, "amount": 1, "unit": "g"} + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [raw]}}, legacy=False, selected=TARGET) + producer.run(show_status=False) + assert _rows(producer) == [] + assert audit.CONTEXT_ROLE not in producer.consumed_input_snapshots + assert audit.SUPPORTED_ROLE not in producer.consumed_input_snapshots + + +@pytest.mark.parametrize("route", ["identity", "attribute", "canonical_name", "synonym"]) +def test_sugar_candidate_object_metadata_and_duplicate_locators(tmp_path, monkeypatch, route): + """All four eligible object-label routes remain auditable without claiming lookups were attempted.""" + row = dict( + _claims("unified")[0], subject_id="MIM:other", subject_label="Different spelling", object_id="CHEBI:17992" + ) + if route == "attribute": + row["subject_id"] = row["object_id"] + if route in {"canonical_name", "synonym"}: + row.update(subject_id="kgm.name:other", comment=route) + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + _files(tmp_path, unified=[row, row], supported=[]) + producer.run(show_status=False) + actual = _rows(producer) + assert len(actual) == 2 + assert len({item["candidate_record_locator"] for item in actual}) == 2 + assert all(json.loads(item["candidate_record"]) == row for item in actual) + + +@pytest.mark.parametrize("predicate", ["skos:broadMatch", "skos:narrowMatch", "rdfs:seeAlso"]) +def test_sugar_does_not_turn_weak_or_annotation_rows_into_identity_candidates(tmp_path, monkeypatch, predicate): + """Matching source display does not change the scientific relation of an original row.""" + producer, _ = _producer(tmp_path, monkeypatch, recipes={"1": {"recipe": [_saved()["raw"]]}}, legacy=False) + _files(tmp_path, unified=[dict(row, predicate_id=predicate) for row in _claims("unified")], supported=[]) + producer.run(show_status=False) + assert _rows(producer) == [] + + +def test_three_material_profiles_keep_their_actual_policy_origins(tmp_path, monkeypatch): + """Sugar never labels its decision as the preexisting P3556/Potato identity exclusions.""" + pc = json.loads(FIXTURE.with_name("p3556_scope.json").read_text()) + potato = json.loads(FIXTURE.with_name("potato_scope.json").read_text()) + recipes = {"1": {"recipe": [_saved()["raw"], pc["raw"], {"compound": "Potato", "compound_id": 1606}]}} + producer, _ = _producer(tmp_path, monkeypatch, recipes=recipes, legacy=False) + _files( + tmp_path, + unified=[ + *_claims("unified"), + *(item["row"] for item in pc["mapping_claims"] if item["table"] == "unified"), + potato["mapping_claims"]["unified"]["row"], + ], + supported=[ + *_claims("supported"), + *(item["row"] for item in pc["mapping_claims"] if item["table"] == "supported"), + ], + ) + producer.run(show_status=False) + actual = _rows(producer) + assert len(actual) == 12 # three Sugar, eight P3556, one Potato original. + for row in actual: + selected = tmp_path / audit.CONTEXT_POLICY if row["reason"] == audit.SUGAR_REASON else audit.IDENTITY_POLICY + assert row["policy_input_path"] == str(selected.resolve()) + assert row["policy_input_sha256"] == hashlib.sha256(selected.read_bytes()).hexdigest() diff --git a/tests/test_merge_source_freshness.py b/tests/test_merge_source_freshness.py index b57557cc..a70e7154 100644 --- a/tests/test_merge_source_freshness.py +++ b/tests/test_merge_source_freshness.py @@ -12,6 +12,7 @@ from kg_microbe.merge_utils import merge_kg from kg_microbe.run import main from kg_microbe.transform import DATA_SOURCES +from kg_microbe.transform_utils.bacdive import ec_substrate_corrections from kg_microbe.transform_utils.constants import DATA_KEY from kg_microbe.transform_utils.mediadive.material_scope_audit import AUDIT_FILENAME, AUDIT_HEADER, MaterialScopeAudit from kg_microbe.transform_utils.transform import Transform @@ -58,6 +59,25 @@ def write_empty_mediadive_audit(transform): transform._material_scope_audit.write() +def prepare_bacdive_ec_correction_inputs(transform): + """Bind a zero-cohort fixture through real policy/read/audit contracts, without claiming a producer run.""" + legacy = transform.input_base_dir / "bacdive_ec_substrates.tsv" + original = FIXTURES / "bacdive_ec_substrates.tsv" + if not legacy.exists(): + shutil.copyfile(original, legacy) + assert legacy.read_bytes() == original.read_bytes() + transform.knowledge_source = "infores:bacdive" + rows, audit, admission = ec_substrate_corrections.prepare_corrections(transform, legacy) + assert rows == [] and audit == [], "Nonempty correction cohorts require the actual producer's graph/audit emission" + header = (FIXTURES / "ec_substrate_corrections.tsv").read_bytes() + assert header == ("\t".join(ec_substrate_corrections.AUDIT_FIELDS) + "\n").encode() + (transform.output_dir / ec_substrate_corrections.AUDIT_FILE).write_bytes(header) + transform.record_producer_audit(ec_substrate_corrections.AUDIT_FILE) + admission.verify() + transform.verify_consumed_inputs() + return ec_substrate_corrections.REQUIRED_INPUTS + + def record_source(transform): """Write the actual registered producer's metadata over the isolated fixture graph.""" cls = type(transform) @@ -82,8 +102,9 @@ def prepare_source(tmp_path, source, *, prefix="", marker=True, output_name=None for kind in ("nodes", "edges"): shutil.copyfile(FIXTURES / f"{kind}.tsv", transform.output_dir / f"{prefix}{kind}.tsv") bulk_roles = read_mediadive_bulk_inputs(transform, create=True) if source == "mediadive" else () + correction_roles = prepare_bacdive_ec_correction_inputs(transform) if source == "bacdive" else () for role in getattr(cls, "REQUIRED_CONSUMED_INPUTS", ()): - if role in bulk_roles: + if role in (*bulk_roles, *correction_roles): continue assert role == "bacdive_taxon_lookup", f"Fixture needs an explicit immutable input for {role}" lookup = tmp_path / f"{source}-bacdive-lookup.tsv" @@ -102,6 +123,37 @@ def prepare_source(tmp_path, source, *, prefix="", marker=True, output_name=None return transform +def test_empty_bacdive_fixture_has_real_correction_input_and_audit_contract(tmp_path): + """An inert graph fixture still exercises actual policy preflight and immutable audit registration.""" + transform = prepare_source(tmp_path, "bacdive", marker=False) + filename = ec_substrate_corrections.AUDIT_FILE + path = transform.output_dir / filename + assert path.read_bytes() == (FIXTURES / filename).read_bytes() + identity = transform.producer_audit_snapshots[filename] + report = json.loads((transform.output_dir / "source_finalization.json").read_text()) + assert set(report["consumed_inputs"]) == set(ec_substrate_corrections.REQUIRED_INPUTS) + assert report["producer_audit_members"][filename] == identity + assert report["audit_members"][filename] == identity + verify_finalized_source_files([transform.output_node_file, transform.output_edge_file]) + + +@pytest.mark.parametrize("damage", ["missing", "unrecorded", "changed"]) +def test_empty_bacdive_fixture_cannot_omit_or_restamp_correction_audit(tmp_path, damage): + """The merge fixture cannot bypass the same mandatory evidence checks as the actual producer.""" + transform = prepare_source(tmp_path, "bacdive", marker=False) + path = transform.output_dir / ec_substrate_corrections.AUDIT_FILE + if damage == "missing": + path.unlink() + elif damage == "changed": + path.write_bytes(path.read_bytes() + b"changed\n") + else: + transform._producer_audit_snapshots.clear() + before = {path.name: path.read_bytes() for path in transform.output_dir.iterdir()} + with pytest.raises(SourceFinalizationRequired, match="audit"): + transform.finalize(fresh_run=True) + assert {path.name: path.read_bytes() for path in transform.output_dir.iterdir()} == before + + def test_empty_mediadive_fixture_has_real_producer_audit(tmp_path): """An empty cohort still retains the writer's exact bytes through public source admission.""" transform = prepare_source(tmp_path, "mediadive") diff --git a/tests/test_merge_upstream_review.py b/tests/test_merge_upstream_review.py index e69a113a..39753455 100644 --- a/tests/test_merge_upstream_review.py +++ b/tests/test_merge_upstream_review.py @@ -13,7 +13,11 @@ from kg_microbe.transform_utils.transform import Transform from kg_microbe.utils.source_finalization import SourceFinalizationRequired from kg_microbe.utils.transform_fingerprint import upstream_fingerprint, write_fingerprint -from tests.test_merge_source_freshness import read_mediadive_bulk_inputs, write_empty_mediadive_audit +from tests.test_merge_source_freshness import ( + prepare_bacdive_ec_correction_inputs, + read_mediadive_bulk_inputs, + write_empty_mediadive_audit, +) ROOT = Path(__file__).resolve().parents[1] FIXTURES = Path(__file__).parent / "resources/merge_source_freshness" @@ -44,8 +48,9 @@ def _prepare(tmp_path, name): for kind in ("nodes", "edges"): shutil.copyfile(FIXTURES / f"{kind}.tsv", transform.output_dir / f"{kind}.tsv") bulk_roles = read_mediadive_bulk_inputs(transform, create=True) if name == "mediadive" else () + correction_roles = prepare_bacdive_ec_correction_inputs(transform) if name == "bacdive" else () for role in getattr(cls, "REQUIRED_CONSUMED_INPUTS", ()): - if role in bulk_roles: + if role in (*bulk_roles, *correction_roles): continue assert role == "bacdive_taxon_lookup", f"An explicit immutable fixture is required for {role}" lookup = tmp_path / f"{name}-bacdive-lookup.tsv" diff --git a/tests/test_sssom_asymmetric_direction.py b/tests/test_sssom_asymmetric_direction.py index be93a92e..475eb98c 100644 --- a/tests/test_sssom_asymmetric_direction.py +++ b/tests/test_sssom_asymmetric_direction.py @@ -27,6 +27,7 @@ RELEASE_PIN = REPO_ROOT / "mappings" / "mim_reviewed_release.json" PRIOR_CLAIMS = REPO_ROOT / "tests" / "resources" / "cas_mapping_promotion" / "prior_claims.tsv" POTATO_SCOPE = REPO_ROOT / "tests" / "resources" / "mediadive" / "potato_scope.json" +P3556_SCOPE = REPO_ROOT / "tests" / "resources" / "mediadive" / "p3556_scope.json" # Byte identities accepted together in the immutable-export promotion review. # Updating a release requires reviewing both products, not regenerating these @@ -34,12 +35,16 @@ SOURCE_COMMIT = "1848b0fe521bc2462f165912fcf92d09ad9a8cec" MANIFEST_SHA256 = "9bb29d5605d93dea351be9624d99c5d8ada57d4b9831b22764957b784bd685af" SUPPORTED_SHA256 = "6b52b30e018b369aa322d41dfd4e81fcfae0e895e34d7fe48900abf5835815fb" -UNIFIED_SHA256 = "09c44642ab13b234e89ce113f7510aa9efb56969e4b539cb7843b43dcb425ba7" +UNIFIED_SHA256 = "67c48e1bf6bed1f1fef03a0da1d7d1b56af9fc72374dddd703c222fb36df3cd4" IDENTITY_REFRESH_WRITER_SHA256 = "257fe4d8bf16f92eb78d5e375065e030ea56d7c1174d859fecd2baa80ac18c27" -IDENTITY_POLICY_SHA256 = "61786a46effa48de9901f77713b172db16a7d6797f35711d5c15c01b0e0ea926" +IDENTITY_POLICY_SHA256 = "4670cbb9255bdac7e9654fc415c9bfd35e55955fea5f27865a74e55849810b4c" RELEASE_PIN_SHA256 = "f082c05656a0910c85176eec7b41deeb77c27967aecb80393fddc262819b6d97" PRIOR_CLAIMS_SHA256 = "629198d090f7e45f7f17ce97a124eb9cad879b0bb066930b11ea8cf80ce15b2c" POTATO_SCOPE_SHA256 = "b4755925efe8cde9871569b047e28c185ac56e2cba6295fca20fef2901062723" +P3556_SCOPE_SHA256 = "7c06f4e2179356ddf1d917a6938cba5e3bdf6669f42a29019b93963e4a70646c" +P3556_REVIEWED_LABEL = ( + "[(2R)-3-hexadecanoyloxy-2-[(9E,12E)-octadeca-9,12-dienoyl]oxy-propyl] 2-(trimethylammonio)ethyl phosphate" +) def _sha256(path): @@ -179,6 +184,25 @@ def test_unified_counts_and_native_categories_match_the_reviewed_candidate(self) self.assertEqual(len(potato_claim), 13) held_potato_row = _row_key(potato_claim) retained_potato_claims = 0 + # Full original claims predate this candidate; only the reviewed native + # object-label replacement is licensed on the four retained synonyms. + self.assertEqual(_sha256(P3556_SCOPE), P3556_SCOPE_SHA256) + p3556_fixture = json.loads(P3556_SCOPE.read_text(encoding="utf-8")) + p3556_originals = [claim["row"] for claim in p3556_fixture["mapping_claims"] if claim["table"] == "unified"] + self.assertEqual(len(p3556_originals), 7) + self.assertTrue(all(len(row) == 13 and row["object_id"] == "CHEBI:86658" for row in p3556_originals)) + p3556_original_keys = {_row_key(row) for row in p3556_originals} + self.assertEqual(len(p3556_original_keys), 7) + p3556_structured = [row for row in p3556_originals if row["source"] == "native_ontology:chebi"] + self.assertEqual(len(p3556_structured), 4) + self.assertTrue(all(row["comment"] == "synonym" for row in p3556_structured)) + self.assertIn( + P3556_REVIEWED_LABEL, + {entry["val"] for entry in p3556_fixture["native_structure"][0]["node"]["meta"]["synonyms"]}, + ) + expected_p3556 = Counter(_row_key({**row, "object_label": P3556_REVIEWED_LABEL}) for row in p3556_structured) + retained_p3556_originals = Counter() + relabelled_p3556 = Counter() prior_rows = list(_iter_sssom_rows(PRIOR_CLAIMS)) prior_invalid = { _row_key(row) @@ -215,6 +239,12 @@ def test_unified_counts_and_native_categories_match_the_reviewed_candidate(self) predicates[row["predicate_id"]] += 1 if row["object_id"] == potato_claim["object_id"] and _row_key(row) == held_potato_row: retained_potato_claims += 1 + if row["object_id"] == "CHEBI:86658": + full_row = _row_key(row) + if full_row in p3556_original_keys: + retained_p3556_originals[full_row] += 1 + if full_row in expected_p3556: + relabelled_p3556[full_row] += 1 for key in ("subject_id", "object_id"): if invalid_cas_identifier(row[key]): invalid_endpoints[(key, row[key])] += 1 @@ -229,14 +259,16 @@ def test_unified_counts_and_native_categories_match_the_reviewed_candidate(self) added_provenance[full_row] += 1 if row["object_id"] in {"NCIT:C16883", "NCIT:C71939"}: native_rows[(row["subject_id"], row["predicate_id"], row["object_id"], row["object_category"])] += 1 - self.assertEqual(rows, 591893 - 33 + 90 - 1) + self.assertEqual(rows, 591946) self.assertEqual(len(entities), 120183) - self.assertEqual(predicates, {"skos:exactMatch": 336050, "skos:closeMatch": 255899}) + self.assertEqual(predicates, {"skos:exactMatch": 336048, "skos:closeMatch": 255898}) self.assertEqual(predicates["skos:broadMatch"], 0) self.assertEqual(predicates["skos:narrowMatch"], 0) self.assertEqual(native_rows, expected_native_rows) self.assertFalse(invalid_endpoints) self.assertFalse(retained_invalid) self.assertEqual(retained_potato_claims, 0) + self.assertFalse(retained_p3556_originals) + self.assertEqual(relabelled_p3556, expected_p3556) self.assertEqual(retained_lexical, Counter({key: 1 for key in prior_lexical})) self.assertEqual(added_provenance, expected_provenance) diff --git a/tests/test_tetramethylammonium_chloride.py b/tests/test_tetramethylammonium_chloride.py new file mode 100644 index 00000000..ec006266 --- /dev/null +++ b/tests/test_tetramethylammonium_chloride.py @@ -0,0 +1,149 @@ +"""Resolve a reviewed salt spelling without guessing its parent or other salts (#1242).""" + +import json +from pathlib import Path + +import pytest + +from kg_microbe.utils import chemical_mapping_utils as runtime +from kg_microbe.utils.ingredient_identity import ingredient_cas_annotations, ingredient_name_target +from tests import test_mim_conservative_refresh as refresh_tests +from tests.test_mediadive_recipe_occurrences import _resolver +from tests.test_mim_conservative_refresh import FIELDS, _metadata, _row, _table + +FIXTURE = Path(__file__).parent / "resources/tetramethylammonium_chloride.json" +SALT_NAME = "Tetramethyl ammonium chloride" +PARENT_NAME = "Tetramethyl ammonium" +PARENT_ID = "kgmicrobe.compound:tetramethyl_ammonium" +inputs = refresh_tests.inputs +bundle = refresh_tests.bundle + + +def _load_selected(tmp_path, monkeypatch, *, salt=True, reverse=False, wrong_alias=False): + """Use only two explicitly declared targets and the native salt synonym.""" + evidence = json.loads(FIXTURE.read_text()) + native = evidence["native_projection"] + rows = [_row("kgm.name:parent", PARENT_ID, PARENT_NAME, name=PARENT_NAME, comment="canonical_name")] + if wrong_alias: + rows.append(_row("kgm.name:wrong_old_alias", PARENT_ID, PARENT_NAME, name=SALT_NAME, comment="synonym")) + if salt: + rows.extend( + _row("kgm.name:" + key, native["id"], native["name"], name=native[key], comment=comment) + for key, comment in (("name", "canonical_name"), ("synonym", "synonym")) + ) + if reverse: + rows.reverse() + path = tmp_path / "selected-declarations.tsv" + _table(path, FIELDS, rows, _metadata()) + monkeypatch.setattr(runtime, "_LOADED", False) + monkeypatch.setattr(runtime, "_CACHED_PATH", None) + return runtime.ChemicalMappingLoader(path) + + +@pytest.mark.parametrize("synonyms", [True, False]) +@pytest.mark.parametrize("reverse", [True, False]) +@pytest.mark.parametrize("salt_first", [True, False]) +def test_finite_alias_and_parent_are_order_independent(tmp_path, monkeypatch, synonyms, reverse, salt_first): + """Both lookup modes honor the finite spelling without changing its parent.""" + loader = _load_selected(tmp_path, monkeypatch, reverse=reverse) + queries = [(SALT_NAME, "CHEBI:7070"), (PARENT_NAME, PARENT_ID)] + if not salt_first: + queries.reverse() + for name, expected in queries: + assert loader.find_chebi_by_name(name, synonyms=synonyms) == expected + assert loader.find_chebi_by_name("Tetramethylammonium chloride") == "CHEBI:7070" + assert ingredient_name_target(PARENT_NAME) is None + + +@pytest.mark.parametrize("synonyms", [True, False]) +def test_missing_salt_declaration_stays_unresolved(tmp_path, monkeypatch, synonyms): + """A reviewed scope is a query route, not permission to manufacture a target.""" + loader = _load_selected(tmp_path, monkeypatch, salt=False) + assert loader.find_chebi_by_name(SALT_NAME, synonyms=synonyms) is None + assert loader.find_chebi_by_name(PARENT_NAME, synonyms=synonyms) == PARENT_ID + + +@pytest.mark.parametrize("salt", [True, False]) +def test_wrong_historical_alias_cannot_replace_approved_scope(tmp_path, monkeypatch, salt): + """A stale parent alias cannot win or invent a registry relationship.""" + loader = _load_selected(tmp_path, monkeypatch, salt=salt, wrong_alias=True) + for synonyms in (True, False): + assert loader.find_chebi_by_name(SALT_NAME, synonyms=synonyms) == ("CHEBI:7070" if salt else None) + assert loader.find_chebi_by_name(PARENT_NAME, synonyms=synonyms) == PARENT_ID + assert loader.find_chebi_by_xref("cas:75-57-0") is None + assert loader.find_chebi_by_name("75-57-0") is None + assert ingredient_cas_annotations("CHEBI:7070") == [] + + +def test_candidate_cycles_preserve_the_same_finite_runtime_route(inputs, monkeypatch): + """The real tiny conservative refresh retains scope semantics and is byte-stable.""" + native = json.loads(FIXTURE.read_text())["native_projection"] + metadata = _metadata() + metadata["curie_map"]["kgmicrobe.compound"] = "https://w3id.org/kg-microbe/compound/" + rows = [ + *refresh_tests.refresh._rows(inputs["baseline"]), + _row("kgm.name:parent", PARENT_ID, PARENT_NAME, name=PARENT_NAME, comment="canonical_name"), + _row("kgm.name:wrong_old_alias", PARENT_ID, PARENT_NAME, name=SALT_NAME, comment="synonym"), + _row("kgm.name:salt", native["id"], native["name"], name=native["name"], comment="canonical_name"), + ] + _table(inputs["baseline"], FIELDS, rows, metadata) + before = inputs["baseline"].read_bytes() + first = refresh_tests.refresh.build_conservative_candidate(**inputs) + second = refresh_tests.refresh.build_conservative_candidate( + **dict(inputs, baseline=first.candidate_path, output_directory=inputs["output_directory"].with_name("second")) + ) + assert inputs["baseline"].read_bytes() == before + assert first.candidate_path.read_bytes() == second.candidate_path.read_bytes() + monkeypatch.setattr(runtime, "_LOADED", False) + monkeypatch.setattr(runtime, "_CACHED_PATH", None) + loader = runtime.ChemicalMappingLoader(second.candidate_path) + assert loader.find_chebi_by_name(SALT_NAME) == native["id"] + assert loader.find_chebi_by_name(PARENT_NAME) == PARENT_ID + assert loader.find_chebi_by_name("Tetramethyl ammonium bromide") is None + assert loader.find_chebi_by_name(SALT_NAME + " monohydrate", fuzzy_hydrate=True) is None + assert loader.find_chebi_by_xref("cas:75-57-0") is None + + +@pytest.mark.parametrize( + "name", + [ + "Tetramethyl ammonium bromide", + "Tetramethyl ammonium hydroxide", + "Tetramethyl ammonium chloride monohydrate", + "Trimethyl ammonium chloride", + "Tetra methyl ammonium chloride", + ], +) +def test_scope_does_not_guess_other_counterions_hydrates_or_spellings(tmp_path, monkeypatch, name): + """No substring match, generic space deletion, or parent-to-salt inference is added.""" + loader = _load_selected(tmp_path, monkeypatch) + assert ingredient_name_target(name) is None + assert loader.find_chebi_by_name(name, fuzzy_hydrate=True, fuzzy_stereochemistry=True) is None + + +def test_actual_occurrences_preserve_raw_record_position_and_quantities(tmp_path, monkeypatch): + """Change the supported salt target only; do not insert public CAS metadata.""" + evidence = json.loads(FIXTURE.read_text()) + loader = _load_selected(tmp_path, monkeypatch) + recipes = {} + for occurrence in evidence["occurrences"]: + # Only selected raw records are copied from the source. Padding keeps + # their positions and is explicitly synthetic, not invented ingredients. + padding = [{"instruction": "Synthetic position padding"}] * occurrence["array_index"] + recipes[occurrence["solution_id"]] = {"recipe": padding + [occurrence["raw"]]} + transform = _resolver(recipes) + transform.chemical_loader = loader + for occurrence in evidence["occurrences"]: + records = transform.get_solution_recipe_occurrences(occurrence["solution_id"]) + assert len(records) == 1 + actual, raw = records[0], occurrence["raw"] + assert actual["id"] == occurrence["expected_target"] + assert actual["source_assertion_id"] == ( + f"mediadive.solution:{occurrence['solution_id']}#recipe/{occurrence['array_index'] + 1}" + ) + assert json.loads(actual["source_record"]) == raw + assert actual["name"] == raw["compound"] + for column in ("amount", "unit", "g_l", "mmol_l"): + assert actual[column] == raw.get(column) + assert not any("cas" in field.lower() for field in json.loads(actual["source_record"])) + assert transform.solutions_data == recipes