Skip to content

RO-Crate static-map arm: retarget the interface-table rows that no longer resolve, and test every row against the schema (#2915) - #4382

Open
realmarcin wants to merge 6 commits into
mainfrom
fix/2915-interface-table-stale-rows
Open

realmarcin wants to merge 6 commits into
mainfrom
fix/2915-interface-table-stale-rows

Conversation

@realmarcin

@realmarcin realmarcin commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

d4d rocrate map builds the registered deterministic_ours arm (rocrate_static_map) by applying data/ro-crate_mapping/d4d_rocrate_interface_mapping.tsv. 54 of its 136 rows placed nowhere in a Dataset record at schema 3.0.0. 52 named a slot or class the schema does not have; the other two name FormatDialect, a class that exists but that no Dataset slot ranges over. The arm dropped every value those rows read, including exactMatch/none rows. Each report row gave the schema's reason (for example "'vulnerable_populations' is not a slot on Dataset"), but no test failed on a table row the schema contradicts.

This PR does four things:

  • Retargets the 11 rows that have an obvious successor. For each one, the new D4D path was checked against the merged schema and the crate path against the CHORUS, CM4AI and VOICE crates.
  • Declares why the other 43 rows place nowhere. Two new columns at the end of each row hold this: Unplaced (10 out_of_scope, 33 owner_question) and Unplaced_Reason. map_crate adds the declaration to the row's detail, and the Outcome legend counts the declared rows.
  • Adds TestTableAgainstSchema. It checks every row's D4D path with the same placement logic map_crate uses. It fails on a row that places nowhere and has no declaration, and on a row that places but still has one. A corpus-lane test also requires every committed report to match the current table's declarations: the same rows, each ending in the declared kind and reason, and the legend's per-kind counts.
  • Regenerates data/ro-crate_packages/{CHORUS,CM4AI,VOICE}/processed/*_crate_mapped_d4d.yaml and *_crate_mapping_provenance.md. All three still validate (PASS, schema 3.0.0).

Refs #2915. The issue stays open: 33 rows need an owner's ruling (listed below). #2915's Part 1 criterion "no row declaring exactMatch + none is unplaceable" is not yet met. 24 such rows remain: 4 are out of scope and 20 are owner questions.

Retargets (11)

Line Was Now Crate side, checked
81 Dataset.vulnerable_populations Dataset.at_risk_populations (renamed in #129) rai:atRiskPopulations, which no crate carries, corrected to d4d:atRiskPopulations. That is the slot's own slot_uri, it is what CM4AI carries, and the FAIRSCAPE converter already reads it. This is the prefix correction notes/D4D_GENERATION_ARMS.md deferred until the row was retargeted.
28 Dataset.external_resource Dataset.external_resources relatedLink: no crate carries it (the row is now empty)
77 Dataset.machine_annotation_analyses Dataset.machine_annotation_tools rai:machineAnnotationTools: VOICE
78 MachineAnnotation.tool_name MachineAnnotationTools.tools same property; subsumed by line 77 in VOICE
96 Maintenance.frequency UpdatePlan.frequency rai:dataReleaseMaintenancePlan; subsumed by Dataset.updates in all three
129, 130 Subset.is_data_split, Subset.is_sub_population DataSubset.is_data_split, DataSubset.is_subpopulation N/A (unresolvable)
131, 132 Variable.name, Variable.type VariableMetadata.variable_name, VariableMetadata.data_type N/A (unresolvable)
133, 134 SamplingStrategy.strategy_type, SamplingStrategy.details SamplingStrategy.strategies, SamplingStrategy.description d4d:samplingStrategy, which map_crate treats as not a crate path (unresolvable)

On each retargeted row, the d4d: slot token in Exchange_Layer_URI is renamed to match (line 81 also gets its crate token), and line 77's D4D_Type class name is updated. The SKOS relation and information loss are unchanged on every row. A note in Transformation_Notes records each retarget.

Before / after

Project filled subsumed empty unresolvable unplaceable distinct slots
CHORUS 32 → 32 0 → 1 47 → 51 4 → 10 54 → 43 32 → 32
CM4AI 42 → 43 0 → 1 37 → 40 4 → 10 54 → 43 42 → 43
VOICE 44 → 45 4 → 6 31 → 33 4 → 10 54 → 43 44 → 45

All 43 remaining unplaceable rows per project carry the table's declaration ("the mapping table says why for 43 of them: 10 out of scope, 33 awaiting an owner's decision"). AI_READI is still refused, because its crate is not UTF-8 (#3357).

Before any change, the run reproduced every committed output exactly, apart from the verdict date.

Values the arm now writes:

  • CHORUS: none. Its mapped record is byte-identical to main's; only its report changes.
  • CM4AI: at_risk_populations ({name: "None — no human subjects involved; commercially sourced de-identified cell lines only."}).
  • VOICE: machine_annotation_tools (seven tool objects: OpenSMILE, Praat/parselmouth, TorchAudio pipelines, sparc, phonetic posteriorgram models, Whisper Large, b2aiprep).

Declared out of scope (10)

These rows map nowhere by design.

  • File-level slots (8). bytes (the contentSize row), encoding, path, md5, sha256, hash, dialect and media_type are slots on File, which sits two levels below Dataset. A row places only one level deep.
    • contentSize is a rounded human string ('1.2 tb'), not a byte count, and is not parsed into one. total_size_bytes reads the exact evi:totalContentSizeBytes, but only CM4AI's crate carries it, so CHORUS's and VOICE's records carry no size, as on main. The row's reason says so.
    • contentUrl already fills download_url.
  • FormatDialect.delimiter and FormatDialect.header. These describe one file's CSV dialect, and their source is prose, not a crate path.

Owner questions (33)

Each row's Unplaced_Reason gives the specifics.

Boundaries kept

Follow-ups filed

Reviewing

  • The table diff. Use git diff --ignore-space-at-eol on the TSV: 55 lines change (the header, 11 retargets, 43 declarations). Every other line gains only two empty trailing fields (the file keeps CRLF and its quoting; unchanged rows round-trip byte for byte through csv).
  • Regenerating CM4AI. It needs data/ro-crate_packages/CM4AI/crate/ re-extracted from the tracked zip (see that directory's README).
  • Six commits. The first retargeted line 89 to DataGovernance.committee_name. The second reverts that before the PR opened: in each crate the arm maps the value names a person, and CHORUS's carries an email address, so the arm was writing a person, and for CHORUS an address, into a slot that names a committee. Line 89 keeps main's mapping columns and is declared an owner question. The last four are review round 1 (below). Review them together; the counts above are the final ones.
  • Commit messages that the round-1 corrections supersede. The message of 1a3f13817 says the report gave a bare "unplaceable" (it gave the schema's reason) and that every one of the 54 rows named a missing slot or class (52 did). The message of 424ba2478 says every crate's dataGovernanceCommittee names a person (true of the three crates the arm maps, not of AI_READI's). History is not rewritten; the table, the reports, the notes, the docstrings and this description carry the corrected statements.

Tests

  • Round 0, at 424ba2478. tests/test_rocrate, tests/test_fairscape_integration, tests/test_cli/test_rocrate_cli.py, tests/test_cli/test_render_gate.py, tests/test_method_split.py, tests/test_deterministic_arm_validity.py and tests/test_rocrate_file_collection.py, corpus tests included: 647 passed, 6 skipped. The skips are environmental and existed before.
  • Round 1, at 0782e61ca. The final head, 70c66ec2b, changes only one sentence of notes/D4D_GENERATION_ARMS.md, which no test reads. The round-0 files, plus four more that read the table, the reports or the package directory (tests/test_review_pack.py, tests/test_generation_specificity_skill.py, tests/test_release_inventory.py and tests/test_download/test_audit_bundles.py), corpus tests included: 942 passed, 6 skipped. The six skips are environmental: linkml-map and the FAIRSCAPE integration are not installed on the development machine.
  • The full -m "not corpus" lane was not run on the development machine, because it was under heavy load (load average about 80 in round 0 and above 100 in round 1). CI runs it.
  • Mutation checks, round 0, run on a git archive copy: 13 of 13 caught.
    • Table mutations: the renamed slot restored; a declaration's kind cleared; a declaration on a row that places; line 81's crate side reverted; line 89 retargeted to DataGovernance.committee_name again (the first commit's version; caught by test_the_retargeted_rows_place_the_crate_values_they_read, which finds data_governance in the record).
    • Mapper mutations: the declaration not appended; the kind not recorded; the legend clause dropped.
    • Committed-report mutations: a declared reason removed; a declared row missing.
    • Guard-helper mutations: each of its three checks disabled.
    • The first twelve were run at the first commit; the line-89 mutation at 424ba2478.
  • Mutation checks, round 1, run on a git archive copy of 0782e61ca: 8 of 8 caught.

Review round 1

Eight issues were filed from the round-1 review of 424ba2478. All eight are fixed here.

🤖 Generated with Claude Code

realmarcin and others added 2 commits October 5, 2026 03:54
…nger resolve, and test every row against the schema (#2915)

54 of the 136 rows in data/ro-crate_mapping/d4d_rocrate_interface_mapping.tsv
named a slot or class the schema does not have, so `d4d rocrate map` dropped
every value they read and reported a bare "unplaceable".

Twelve rows had an obvious successor and now name it:
- Dataset.vulnerable_populations -> Dataset.at_risk_populations (renamed in
  #129), with the crate side corrected from rai:atRiskPopulations, which no
  crate carries, to d4d:atRiskPopulations, the slot's own slot_uri, which
  CM4AI carries and the FAIRSCAPE converter already reads
- Dataset.external_resource -> Dataset.external_resources
- Dataset.machine_annotation_analyses -> Dataset.machine_annotation_tools
- MachineAnnotation.tool_name -> MachineAnnotationTools.tools
- DatasetCollection.data_governance_committee -> DataGovernance.committee_name
- Maintenance.frequency -> UpdatePlan.frequency
- Subset.* -> DataSubset.*, Variable.* -> VariableMetadata.*
- SamplingStrategy.strategy_type/details -> strategies/description

The other 42 declare why they place nowhere in two new trailing columns,
Unplaced (out_of_scope: 11; owner_question: 31) and Unplaced_Reason, which the
provenance report now shows beside the schema's reason and counts in its
Outcome legend. Rows 67 and 87 keep their mapping columns as they are
(#4043); row 87 only gains its declaration.

TestTableAgainstSchema judges every row's D4D path by the placement map_crate
applies and fails on a row that places nowhere without a declaration, or that
places and keeps one. A corpus-lane test ties the committed reports to the
current table's declarations.

Regenerated data/ro-crate_packages/{CHORUS,CM4AI,VOICE}/processed/ (filled
rows CHORUS 32->33, CM4AI 42->44, VOICE 44->46; unplaceable 54->42 each; all
PASS). AI_READI stays refused (not UTF-8). at_risk_populations joins the
single-valued class-range slots where the two crate arms differ on a mixed
list, in _coerce's comment and the cross-arm test.

Refs #2915

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…than place a person in committee_name (#2915)

Every crate's dataGovernanceCommittee names a person, and CHORUS's adds an
email address, so placing it in DataGovernance.committee_name wrote a
person's name, and for CHORUS an address, into a slot that names a
committee. Row 89 keeps its main mapping columns and is declared
owner_question with that reason: 11 retargets and 43 declarations (11 out
of scope, 32 owner questions). The three projects' outputs are
regenerated from the table; no committed file gains an email address.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Oct 5, 2026
realmarcin and others added 4 commits October 5, 2026 08:43
…ason, and the legend's counts (#4413)

test_every_unplaceable_row_reported_is_one_the_table_declares compared the
set of unplaceable paths and looked for "as the table declares: {reason}"
anywhere in the row. A table edit that flipped a row's Unplaced kind and kept
its reason left every committed report saying the old kind and the old
legend counts, and the guard passed. A reason the table shortened was also
found inside the longer one the stale report still carried.

The guard now requires each declared row's cell to end with the kind and
the whole reason, as map_crate words them, and the Outcome legend to end
with the per-kind counts the table gives. A nested row refused only because
its merge is undecided is recognised by the ending map_crate writes, not by
a phrase anywhere in the row, since a declared reason may use the same
words.

TestTheGuardReadsWhatTheMapperWrites builds a report in memory, so it runs
in the pull-request lane: it pins the three readings (the undecided-merge
ending, the declared cell, the legend clause) to what map_crate and
write_provenance write, with a declared reason in the undecided-merge note's
own words.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ize and governance reasons (#4414, #4415, #4416)

Line 90 (DatasetCollection.principal_investigator, exactMatch/none) was
declared out_of_scope while its own reason described a decision still to be
made. Retargeted to its one successor, Creator.principal_investigator under
Dataset.creators, the row is refused in all three crates because the author
row already fills creators and merging two crate properties into one object
is not decided. #2915 asks for that decision. It is now owner_question, and
its reason says so; its mapping columns are unchanged. The legend moves from
11 out of scope / 32 awaiting an owner's decision to 10 / 33, and of the 24
exactMatch/none rows still unplaced, 4 are out of scope and 20 owner
questions.

Line 13 (Dataset.bytes) said the record's size is total_size_bytes, read
from the exact evi:totalContentSizeBytes. Only CM4AI's crate carries that
property, so CHORUS's and VOICE's records carry no size; the reason now says
so.

Line 89 (dataGovernanceCommittee) said every crate names a person. That holds
for the three crates the arm maps; AI_READI's, refused for its encoding
(#3357), names a consortium. The reason now scopes the claim and names
#4386.

The three provenance reports are regenerated with `d4d rocrate map`: each
changes in the legend and in the three rows. The three mapped records are
byte-identical, and all three still validate (PASS, schema 3.0.0).

notes/D4D_GENERATION_ARMS.md gets the same corrections and two more: the
filled-row counts after the retargets were the first commit's (33, 44, 46);
they are CM4AI 43 and VOICE 45, with CHORUS unchanged (#4412). And 52 of the
54 rows named a slot or class the schema does not have; the two FormatDialect
rows name a class no Dataset slot ranges over, and each report row did give
the schema's reason (#4418).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…re, and a legend that tells declared rows from the total (#4416, #4417, #4418, #4419)

The guard's docstrings said a row that places but keeps its declaration
"would be reported unplaced". map_crate reads a declaration only where
placement fails, so it ignores one on a row that places: the record carries
the value of a row the table says maps nowhere, and the report never shows
the declaration. The docstrings now say that, and the test asserts it: the
row is filled, carries no declaration, and its value is in the record
(#4417).

TestTableAgainstSchema's docstring said 54 rows named a slot or class the
schema does not have, with nothing to say the table was wrong. 52 did; the
two FormatDialect rows name a class no Dataset slot ranges over, and each
report row gave the schema's reason. What was missing was a test (#4418).
The governance docstring scopes its claim to the crates the arm maps
(#4416).

The legend test ran on a graph where every unplaceable row was declared, so
writing the unplaceable total in place of the declared count passed it. It
now maps the real table over a crate with one nested row refused only
because its merge is undecided (the same property on two Dataset entities,
#3270), and asserts 44 unplaceable rows, 43 of them declared. The small
fixture's legend test adds a declared row beside an undecided one (#4419).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The sentence the previous commit added said the bytes row's value "is not
carried elsewhere", which reads as true of every crate. CM4AI's exact size is
carried, through evi:totalContentSizeBytes; only CHORUS's and VOICE's records
carry none. Wording only; no test reads the note.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant