RO-Crate static-map arm: retarget the interface-table rows that no longer resolve, and test every row against the schema (#2915) - #4382
Open
realmarcin wants to merge 6 commits into
Conversation
…nger resolve, and test every row against the schema (#2915) 54 of the 136 rows in data/ro-crate_mapping/d4d_rocrate_interface_mapping.tsv named a slot or class the schema does not have, so `d4d rocrate map` dropped every value they read and reported a bare "unplaceable". Twelve rows had an obvious successor and now name it: - Dataset.vulnerable_populations -> Dataset.at_risk_populations (renamed in #129), with the crate side corrected from rai:atRiskPopulations, which no crate carries, to d4d:atRiskPopulations, the slot's own slot_uri, which CM4AI carries and the FAIRSCAPE converter already reads - Dataset.external_resource -> Dataset.external_resources - Dataset.machine_annotation_analyses -> Dataset.machine_annotation_tools - MachineAnnotation.tool_name -> MachineAnnotationTools.tools - DatasetCollection.data_governance_committee -> DataGovernance.committee_name - Maintenance.frequency -> UpdatePlan.frequency - Subset.* -> DataSubset.*, Variable.* -> VariableMetadata.* - SamplingStrategy.strategy_type/details -> strategies/description The other 42 declare why they place nowhere in two new trailing columns, Unplaced (out_of_scope: 11; owner_question: 31) and Unplaced_Reason, which the provenance report now shows beside the schema's reason and counts in its Outcome legend. Rows 67 and 87 keep their mapping columns as they are (#4043); row 87 only gains its declaration. TestTableAgainstSchema judges every row's D4D path by the placement map_crate applies and fails on a row that places nowhere without a declaration, or that places and keeps one. A corpus-lane test ties the committed reports to the current table's declarations. Regenerated data/ro-crate_packages/{CHORUS,CM4AI,VOICE}/processed/ (filled rows CHORUS 32->33, CM4AI 42->44, VOICE 44->46; unplaceable 54->42 each; all PASS). AI_READI stays refused (not UTF-8). at_risk_populations joins the single-valued class-range slots where the two crate arms differ on a mixed list, in _coerce's comment and the cross-arm test. Refs #2915 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…than place a person in committee_name (#2915) Every crate's dataGovernanceCommittee names a person, and CHORUS's adds an email address, so placing it in DataGovernance.committee_name wrote a person's name, and for CHORUS an address, into a slot that names a committee. Row 89 keeps its main mapping columns and is declared owner_question with that reason: 11 retargets and 43 declarations (11 out of scope, 32 owner questions). The three projects' outputs are regenerated from the table; no committed file gains an email address. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Oct 5, 2026
Open
Notes give the first commit's filled-row counts (33, 44, 46); the final reports say 32, 43, 45
#4412
Open
…ason, and the legend's counts (#4413) test_every_unplaceable_row_reported_is_one_the_table_declares compared the set of unplaceable paths and looked for "as the table declares: {reason}" anywhere in the row. A table edit that flipped a row's Unplaced kind and kept its reason left every committed report saying the old kind and the old legend counts, and the guard passed. A reason the table shortened was also found inside the longer one the stale report still carried. The guard now requires each declared row's cell to end with the kind and the whole reason, as map_crate words them, and the Outcome legend to end with the per-kind counts the table gives. A nested row refused only because its merge is undecided is recognised by the ending map_crate writes, not by a phrase anywhere in the row, since a declared reason may use the same words. TestTheGuardReadsWhatTheMapperWrites builds a report in memory, so it runs in the pull-request lane: it pins the three readings (the undecided-merge ending, the declared cell, the legend clause) to what map_crate and write_provenance write, with a declared reason in the undecided-merge note's own words. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ize and governance reasons (#4414, #4415, #4416) Line 90 (DatasetCollection.principal_investigator, exactMatch/none) was declared out_of_scope while its own reason described a decision still to be made. Retargeted to its one successor, Creator.principal_investigator under Dataset.creators, the row is refused in all three crates because the author row already fills creators and merging two crate properties into one object is not decided. #2915 asks for that decision. It is now owner_question, and its reason says so; its mapping columns are unchanged. The legend moves from 11 out of scope / 32 awaiting an owner's decision to 10 / 33, and of the 24 exactMatch/none rows still unplaced, 4 are out of scope and 20 owner questions. Line 13 (Dataset.bytes) said the record's size is total_size_bytes, read from the exact evi:totalContentSizeBytes. Only CM4AI's crate carries that property, so CHORUS's and VOICE's records carry no size; the reason now says so. Line 89 (dataGovernanceCommittee) said every crate names a person. That holds for the three crates the arm maps; AI_READI's, refused for its encoding (#3357), names a consortium. The reason now scopes the claim and names #4386. The three provenance reports are regenerated with `d4d rocrate map`: each changes in the legend and in the three rows. The three mapped records are byte-identical, and all three still validate (PASS, schema 3.0.0). notes/D4D_GENERATION_ARMS.md gets the same corrections and two more: the filled-row counts after the retargets were the first commit's (33, 44, 46); they are CM4AI 43 and VOICE 45, with CHORUS unchanged (#4412). And 52 of the 54 rows named a slot or class the schema does not have; the two FormatDialect rows name a class no Dataset slot ranges over, and each report row did give the schema's reason (#4418). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…re, and a legend that tells declared rows from the total (#4416, #4417, #4418, #4419) The guard's docstrings said a row that places but keeps its declaration "would be reported unplaced". map_crate reads a declaration only where placement fails, so it ignores one on a row that places: the record carries the value of a row the table says maps nowhere, and the report never shows the declaration. The docstrings now say that, and the test asserts it: the row is filled, carries no declaration, and its value is in the record (#4417). TestTableAgainstSchema's docstring said 54 rows named a slot or class the schema does not have, with nothing to say the table was wrong. 52 did; the two FormatDialect rows name a class no Dataset slot ranges over, and each report row gave the schema's reason. What was missing was a test (#4418). The governance docstring scopes its claim to the crates the arm maps (#4416). The legend test ran on a graph where every unplaceable row was declared, so writing the unplaceable total in place of the declared count passed it. It now maps the real table over a crate with one nested row refused only because its merge is undecided (the same property on two Dataset entities, #3270), and asserts 44 unplaceable rows, 43 of them declared. The small fixture's legend test adds a declared row beside an undecided one (#4419). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The sentence the previous commit added said the bytes row's value "is not carried elsewhere", which reads as true of every crate. CM4AI's exact size is carried, through evi:totalContentSizeBytes; only CHORUS's and VOICE's records carry none. Wording only; no test reads the note. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
d4d rocrate mapbuilds the registereddeterministic_oursarm (rocrate_static_map) by applyingdata/ro-crate_mapping/d4d_rocrate_interface_mapping.tsv. 54 of its 136 rows placed nowhere in aDatasetrecord at schema 3.0.0. 52 named a slot or class the schema does not have; the other two nameFormatDialect, a class that exists but that noDatasetslot ranges over. The arm dropped every value those rows read, includingexactMatch/nonerows. Each report row gave the schema's reason (for example "'vulnerable_populations' is not a slot on Dataset"), but no test failed on a table row the schema contradicts.This PR does four things:
Unplaced(10out_of_scope, 33owner_question) andUnplaced_Reason.map_crateadds the declaration to the row's detail, and the Outcome legend counts the declared rows.TestTableAgainstSchema. It checks every row's D4D path with the same placement logicmap_crateuses. It fails on a row that places nowhere and has no declaration, and on a row that places but still has one. A corpus-lane test also requires every committed report to match the current table's declarations: the same rows, each ending in the declared kind and reason, and the legend's per-kind counts.data/ro-crate_packages/{CHORUS,CM4AI,VOICE}/processed/*_crate_mapped_d4d.yamland*_crate_mapping_provenance.md. All three still validate (PASS, schema 3.0.0).Refs #2915. The issue stays open: 33 rows need an owner's ruling (listed below). #2915's Part 1 criterion "no row declaring
exactMatch+noneis unplaceable" is not yet met. 24 such rows remain: 4 are out of scope and 20 are owner questions.Retargets (11)
Dataset.vulnerable_populationsDataset.at_risk_populations(renamed in #129)rai:atRiskPopulations, which no crate carries, corrected tod4d:atRiskPopulations. That is the slot's ownslot_uri, it is what CM4AI carries, and the FAIRSCAPE converter already reads it. This is the prefix correctionnotes/D4D_GENERATION_ARMS.mddeferred until the row was retargeted.Dataset.external_resourceDataset.external_resourcesrelatedLink: no crate carries it (the row is nowempty)Dataset.machine_annotation_analysesDataset.machine_annotation_toolsrai:machineAnnotationTools: VOICEMachineAnnotation.tool_nameMachineAnnotationTools.toolssubsumedby line 77 in VOICEMaintenance.frequencyUpdatePlan.frequencyrai:dataReleaseMaintenancePlan;subsumedbyDataset.updatesin all threeSubset.is_data_split,Subset.is_sub_populationDataSubset.is_data_split,DataSubset.is_subpopulationN/A(unresolvable)Variable.name,Variable.typeVariableMetadata.variable_name,VariableMetadata.data_typeN/A(unresolvable)SamplingStrategy.strategy_type,SamplingStrategy.detailsSamplingStrategy.strategies,SamplingStrategy.descriptiond4d:samplingStrategy, whichmap_cratetreats as not a crate path (unresolvable)On each retargeted row, the
d4d:slot token inExchange_Layer_URIis renamed to match (line 81 also gets its crate token), and line 77'sD4D_Typeclass name is updated. The SKOS relation and information loss are unchanged on every row. A note inTransformation_Notesrecords each retarget.Before / after
All 43 remaining unplaceable rows per project carry the table's declaration ("the mapping table says why for 43 of them: 10 out of scope, 33 awaiting an owner's decision"). AI_READI is still refused, because its crate is not UTF-8 (#3357).
Before any change, the run reproduced every committed output exactly, apart from the verdict date.
Values the arm now writes:
at_risk_populations({name: "None — no human subjects involved; commercially sourced de-identified cell lines only."}).machine_annotation_tools(seven tool objects: OpenSMILE, Praat/parselmouth, TorchAudio pipelines, sparc, phonetic posteriorgram models, Whisper Large, b2aiprep).Declared out of scope (10)
These rows map nowhere by design.
File-level slots (8).bytes(thecontentSizerow),encoding,path,md5,sha256,hash,dialectandmedia_typeare slots onFile, which sits two levels belowDataset. A row places only one level deep.contentSizeis a rounded human string ('1.2 tb'), not a byte count, and is not parsed into one.total_size_bytesreads the exactevi:totalContentSizeBytes, but only CM4AI's crate carries it, so CHORUS's and VOICE's records carry no size, as on main. The row's reason says so.contentUrlalready fillsdownload_url.FormatDialect.delimiterandFormatDialect.header. These describe one file's CSV dialect, and their source is prose, not a crate path.Owner questions (33)
Each row's
Unplaced_Reasongives the specifics.EthicalReview.irb_id.DatasetCollection.data_governance_committee(mapping columns unchanged). The successor isDataset.data_governance(DataGovernance.committee_name, No governance or stewardship class exists; the one DAC contact slot is nested under ExportControlRegulatoryRestrictions #503). Butcommittee_nameis "the body responsible for access decisions", and in each crate the arm maps (CHORUS, CM4AI, VOICE)dataGovernanceCommitteenames a person (CHORUS's adds an email address). AI_READI's names a consortium, but the arm refuses that crate for its encoding (Decide whether the AI_READI crate is decoded as its declared windows-1252 or stays refused #3357). Whether to place it, incommittee_nameor as a contact, is the owner's decision (Interface table line 89: dataGovernanceCommittee names a person in every crate; where, if anywhere, should the static-map arm place it? #4386).DatasetCollection.principal_investigator(mapping columns unchanged). Its one slot isCreator.principal_investigator(rangePerson), underDataset.creators, which theauthorrow fills. Retargeted there, the row is refused in all three crates, because merging two crate properties into one object is not decided. Making aCreatororPersonof the free-text name (CHORUS's carries an email address) is the decision RO-Crate mapper fidelity: 54 of 136 interface-table rows cannot be placed, and nested *.description rows overwrite filled host slots #2915's proposal asks for, and RO-Crate mapper fidelity: 54 of 136 interface-table rows cannot be placed, and nested *.description rows overwrite filled host slots #2915 tracks it.DatasetCollection.parent_datasets→Dataset.parent_datasets. CHORUS's and CM4AI'sisPartOfname an organization and a project, which is FAIRSCAPE isPartOf names an organization and a project; the reverse converter now places them in parent_datasets #4047's open question.related_datasets.DatasetRelationshiprequires arelationship_typethatrelatedLinkdoes not state.Maintenance.versioning_strategy. Both candidate homes,UpdatePlan.update_detailsandVersionAccess.version_details, would be a judgement.EvidenceMetadata.total_content_size_bytes(Dataset.total_size_bytes), line 106EvidenceMetadata.formats(Dataset.distribution_formats) and line 110DatasetCollection.funding_and_acknowledgements(Dataset.funders). For each: remove it, or keep it as a documented duplicate.git log -Soversrc/data_sheets_schema/schema/): lines 61, 62, 65, 66, 73 and 74, thestep_type,pipeline_step,annotator_typeandevidence_typerows.HumanSubjectResearch.exemption(88)contact_email(93),data_sharing_agreement(94)EvidenceMetadatacounts (98–104)completeness(107),summary_statistics(108),quality_control(109),provenance_and_lineage(111)ValidationMetrics.validation_method(112)QualityControl.accuracy(113),QualityControl.data_quality_report(114),QualityControl.fda_compliant(115)Boundaries kept
Dataset.imputation_protocols,rai:imputationProtocol) and line 87 (EthicalReview.irb_id,rai:ethicalReview) wait for the owner's ruling on generate_interface_mapping.py (pinned) writes rai:prohibitedUses / rai:ethicalReview / rai:imputationProtocol, and the committed interface table disagrees with it #4043. Their mapping columns are byte-identical. Line 67 places, so it gains only two empty fields. Line 87 gains itsowner_questiondeclaration..claude/agents/scripts/generate_interface_mapping.pyis untouched. Regenerating from it would undo this PR and the earlier hand corrections, as generate_interface_mapping.py (pinned) writes rai:prohibitedUses / rai:ethicalReview / rai:imputationProtocol, and the committed interface table disagrees with it #4043 notes.data/d4d_concatenated/**file changes, and the published arm labels stay as they are. Deterministic crate mappers write resolver-URL DOIs: both deterministic arms fail the anchored doi pattern (#646) while crate-mapping provenance still reports PASS #2916 mints new deterministic labels.at_risk_populationsis now a slot both arms fill, and the arms agree on CM4AI's value. It joinsupdatesandhuman_subject_researchas a single-valued class-range slot where the arms differ on a list mixing values with lists (FAIRSCAPE converter: a single-valued class-range slot places nothing for text beside a nested list, which rocrate_map keeps #4194's class)._coerce's comment andtest_documented_exceptions_across_all_shared_pairs_and_value_kindsnow name it.machine_annotation_tools.dataGovernanceCommitteeis placed by neither arm. This one declares line 89 an owner question; the converter neither reads the property nor records it as dropped (Reverse FAIRSCAPE converter maps four RAI properties to slots other than those the schema declares, and leaves FAIRSCAPE properties with declared homes unread #4046).Follow-ups filed
Dataset.human_subject_researchreadsd4d:humanSubject, which no crate carries; the crates carryhumanSubjectResearch,humanSubjectsandhumanSubjectExemption.dataGovernanceCommittee, names a person in each crate the arm maps. Where, if anywhere, should the arm place it?Reviewing
git diff --ignore-space-at-eolon the TSV: 55 lines change (the header, 11 retargets, 43 declarations). Every other line gains only two empty trailing fields (the file keeps CRLF and its quoting; unchanged rows round-trip byte for byte throughcsv).data/ro-crate_packages/CM4AI/crate/re-extracted from the tracked zip (see that directory's README).DataGovernance.committee_name. The second reverts that before the PR opened: in each crate the arm maps the value names a person, and CHORUS's carries an email address, so the arm was writing a person, and for CHORUS an address, into a slot that names a committee. Line 89 keeps main's mapping columns and is declared an owner question. The last four are review round 1 (below). Review them together; the counts above are the final ones.1a3f13817says the report gave a bare "unplaceable" (it gave the schema's reason) and that every one of the 54 rows named a missing slot or class (52 did). The message of424ba2478says every crate'sdataGovernanceCommitteenames a person (true of the three crates the arm maps, not of AI_READI's). History is not rewritten; the table, the reports, the notes, the docstrings and this description carry the corrected statements.Tests
424ba2478.tests/test_rocrate,tests/test_fairscape_integration,tests/test_cli/test_rocrate_cli.py,tests/test_cli/test_render_gate.py,tests/test_method_split.py,tests/test_deterministic_arm_validity.pyandtests/test_rocrate_file_collection.py, corpus tests included: 647 passed, 6 skipped. The skips are environmental and existed before.0782e61ca. The final head,70c66ec2b, changes only one sentence ofnotes/D4D_GENERATION_ARMS.md, which no test reads. The round-0 files, plus four more that read the table, the reports or the package directory (tests/test_review_pack.py,tests/test_generation_specificity_skill.py,tests/test_release_inventory.pyandtests/test_download/test_audit_bundles.py), corpus tests included: 942 passed, 6 skipped. The six skips are environmental: linkml-map and the FAIRSCAPE integration are not installed on the development machine.-m "not corpus"lane was not run on the development machine, because it was under heavy load (load average about 80 in round 0 and above 100 in round 1). CI runs it.git archivecopy: 13 of 13 caught.DataGovernance.committee_nameagain (the first commit's version; caught bytest_the_retargeted_rows_place_the_crate_values_they_read, which findsdata_governancein the record).424ba2478.git archivecopy of0782e61ca: 8 of 8 caught.TestTheGuardReadsWhatTheMapperWrites.TestTheGuardReadsWhatTheMapperWrites); the legend's declared count replaced by the unplaceable total (Legend test cannot tell 'declared' from 'unplaceable': counting all unplaceable rows as declared survives every test #4419; caught by both legend tests and byTestTheGuardReadsWhatTheMapperWrites); a row that places given its declared kind, or its declared reason in its detail (Guard docstrings say a placed row keeping a declaration "would be reported unplaced"; the mapper ignores it #4417; each caught bytest_the_guard_fails_a_declaration_on_a_row_that_places).424ba2478: unmutated, the old and the new guard both pass. Under each of the three committed-report mutations, the old guard passes and the new one fails.Review round 1
Eight issues were filed from the round-1 review of
424ba2478. All eight are fixed here.notes/D4D_GENERATION_ARMS.mdgave the first commit's filled-row counts after the retargets (33, 44, 46). It now gives CM4AI 43 and VOICE 45, and says CHORUS's record is unchanged.map_cratewrites, not by a phrase anywhere in the row: line 90's new reason uses the same words, and the old filter would have dropped that row from the comparison.TestTheGuardReadsWhatTheMapperWritespins all three readings to what the mapper writes, and runs in the pull-request lane.principal_investigator) was declaredout_of_scope, but its own reason described a decision still to be made. It is nowowner_question, and its reason says why. The legend moves from 11 / 32 to 10 / 33, and theexactMatch+nonesplit from 5 / 19 to 4 / 20. The three reports are regenerated; the three mapped records are byte-identical.total_size_bytes, read from the exactevi:totalContentSizeBytes. Only CM4AI's crate carries that property. The reason now says CHORUS's and VOICE's records carry no size; the row stays out of scope.dataGovernanceCommitteenames a person" was false for the tracked AI_READI crate, which names a consortium. The table reason, the reports, the notes, the test docstring and this description now scope the claim to the crates the arm maps. The424ba2478commit message keeps the old wording.map_crateignores such a declaration. The docstrings now say so, andtest_the_guard_fails_a_declaration_on_a_row_that_placesasserts it: the row is filled, carries no declaration, and its value is in the record.FormatDialectrows, and the reports did give a reason for each row. The test docstring, the notes and this description now say 52 plus two, and that what was missing was a test. The1a3f13817commit message keeps the old wording.🤖 Generated with Claude Code