Skip to content

Latest commit

 

History

573 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

PURL Associator

This repository maintains canonical conda-forge package identity mappings. Each mapping record is keyed by conda-forge package name and carries identifiers that downstream security tooling can use:

  • a primary Package URL (PURL) and optional alternative PURLs
  • optional CPE 2.3 vendor/product prefixes for NVD matching
  • package context such as latest observed version, recipe/source URLs, summary, and download counts

CVE assignment, OpenVEX review state, AI CVE drafts, and SBOM-derived findings are intentionally out of scope. A downstream CVE project should consume the identity mapping payload produced here, enumerate conda-forge versions there, and join those versions with OSV/NVD affected-version data there.

Data model

The core object is a conda-forge package identity record:

{
  "name": "ncurses",
  "version": "6.5",
  "purl": "pkg:github/ThomasDickey/ncurses-snapshots",
  "type": "github",
  "namespace": "ThomasDickey",
  "pkg_name": "ncurses-snapshots",
  "alternative_purls": [],
  "cpes": [
    "cpe:2.3:a:gnu:ncurses",
    "cpe:2.3:a:invisible-island:ncurses"
  ],
  "status": "verified",
  "identities": [
    {
      "kind": "purl",
      "role": "primary",
      "value": "pkg:github/ThomasDickey/ncurses-snapshots",
      "provenance": {
        "availability": "available",
        "source": "manual",
        "review": {
          "status": "verified",
          "reviewer": "nichmor",
          "reviewed_at": "2026-05-25T00:00:00Z"
        }
      }
    },
    {
      "kind": "cpe",
      "role": "associated",
      "value": "cpe:2.3:a:gnu:ncurses",
      "provenance": {
        "availability": "available",
        "source": "manual",
        "review": {
          "status": "verified",
          "reviewer": "nichmor",
          "reviewed_at": "2026-05-25T00:00:00Z"
        }
      }
    },
    {
      "kind": "cpe",
      "role": "associated",
      "value": "cpe:2.3:a:invisible-island:ncurses",
      "provenance": {
        "availability": "available",
        "source": "manual",
        "review": {
          "status": "verified",
          "reviewer": "nichmor",
          "reviewed_at": "2026-05-25T00:00:00Z"
        }
      }
    }
  ]
}

PURLs identify source/package ecosystem coordinates. CPEs identify NVD vendor/product coordinates. CPE strings stored here are identity-level prefixes; this repository does not store CVE affected ranges or per-CVE version decisions.

Published identity provenance contract

The generated payload's identities array is the authoritative per-identity view for downstream consumers:

  • kind: "purl", role: "primary" identifies the primary PURL. Its provenance contains a required review status, a nullable reviewer, and an optional reviewed_at. Automatic primary identities additionally retain their available confidence and source evidence.
  • kind: "purl", role: "alternative" identifies an alternative PURL. Detailed alternatives retain their own source and confidence; bare string alternatives use availability: "unavailable". They never inherit the primary PURL's review object.
  • kind: "cpe", role: "associated" identifies a CPE prefix. Its provenance identifies the automatic or reviewed source layer that last replaced the package's CPE list, including that layer's review state and attribution.

Identity order is deterministic: primary PURL, alternatives in source order, then CPEs in source order. A missing primary PURL produces no primary identity.

Coverage is identity-based. A package has no identity only when identities is empty; a missing primary PURL can still leave alternative PURLs or CPEs. The legacy unmapped field records an intentional no-PURL decision, not the absence of every identity. Every published confidence is a finite JSON number; NaN and infinities are rejected. Contribution precedence compares timezone-aware timestamps as UTC instants, with the on-disk filename as the deterministic tie-breaker.

Sources and outputs

Path Purpose
mappings/auto.json automatically inferred PURL mappings
mappings/manual.json legacy/manual reviewed overrides
mappings/contributions/*.json PR-submitted mapping contributions, including CPE pipeline output
mappings/cpe_candidates/*.json audit output from CPE discovery
mappings/cpe_vet/*.json optional AI tiebreaker output for ambiguous CPE candidates
web/public/mappings.json generated full mapping bundle
web/public/mappings-index.json generated compact index for the web app
web/public/mapping_packages/*.json generated sharded package detail payloads

Payload versions and compatibility

PFX-1826 introduced identities in bundle schema 2, index schema 3, and detail schema 2. CPE provenance advances those payloads to bundle schema 3, index schema 4, and detail schema 3. Reviewed no-PURL reasons advance them to bundle schema 4, index schema 5, and detail schema 4. Legacy purl, alternative_purls, cpes, status, and attribution fields remain unchanged. New clients should reject unknown future schema versions rather than guessing. The validator and web decoder accept the earlier payload generations but require attributed CPE provenance on current schemas.

The payload deployed by the Pages workflow is canonical. The checked-in mappings-index.json and shard files are build inputs/cache for the app, not a supported raw-consumer contract on their own: a source change can leave them stale until mappings:merge regenerates the bundle, index, and all shards atomically. Consumers should use the deployed Pages payload rather than mixing checked-in generations. Routine changes should not commit the 257-file shard churn unless repository policy explicitly requires an atomic generated-data refresh.

PURL flow

flowchart TD
  A[conda-forge metadata] --> B[scripts.automap]
  B --> C[mappings/auto.json]
  C --> D[scripts.merge_mappings]
  E[mappings/manual.json] --> D
  F[mappings/contributions/*.json] --> D
  D --> G[web/public/mappings*.json]
  G --> H[PURL editing UI]
  H --> I[Worker POST /api/submit]
  I --> F
Loading

Useful commands:

pixi run purl:automap --only numpy,ripgrep,pandas
pixi run -e lite mappings:merge
pixi run -e lite mappings:validate
pixi run purl:test

Automap reuses successful cached entries while their package version and build are unchanged. Entries whose persisted note starts with fetch error: are retried on the next run even when those coordinates are unchanged, so transient processing failures cannot remain cached indefinitely.

Source URL inference includes registered pkg:bitbucket repository identities and case-sensitive pkg:cpan distribution identities. Because gitlab is not a registered PURL type, gitlab.com repositories use the registered host-neutral pkg:git/gitlab.com/<namespace>/<repository> form. MetaCPAN module (/pod/) pages alone are intentionally insufficient evidence because module and distribution names can differ; canonical distribution and release/archive URLs are accepted.

Primary-PURL coverage report

Report effective primary-PURL coverage after automatic mappings, manual overrides, and contributions have been merged. The reporter requires the current full bundle schema; it does not accept the compact index or raw auto.json.

Generate the merged mappings and browser report used by the GitHub Pages app:

pixi run -e lite mappings:report

To rebuild and report without changing the generated public payloads:

pixi run -e lite mappings:merge \
  --out .tmp/coverage/mappings.json \
  --index-out .tmp/coverage/mappings-index.json \
  --detail-dir .tmp/coverage/mapping_packages
pixi run -e lite python -m scripts.mappings_report \
  --input .tmp/coverage/mappings.json \
  --output .tmp/coverage/report.json

Render a bounded Markdown summary, optionally compared with an earlier JSON report:

pixi run -e lite python -m scripts.mappings_report \
  --input .tmp/coverage/mappings.json \
  --format markdown \
  --baseline-report /tmp/mappings-report-before.json

The scheduled automap workflow captures that baseline before refreshing mappings and appends the resulting coverage and diagnostic deltas to its PR body. The Markdown summary includes the top current unresolved source hosts but not the full missing-package list.

The JSON report separates primary_present, explicitly_unmapped, and unresolved; these three counts sum to total. primary_missing is the sum of the latter two states, while actionable_missing equals only unresolved. classified_unmapped and legacy_unmapped partition explicit decisions according to whether structured rationale was recorded. These are storage/reporting states, not security-negative findings. The legacy actionable_missing name describes only the existing unresolved queue; it does not mean deferred/legacy decisions need no further work.

Markdown reports and the inspector derive three review cohorts from the existing schema-2 report, without changing mapping decisions or the published wire format:

  • Packaging-only decisions (recorded evidence): the metapackage, selector, mutex, pinning, and compatibility-shim reason codes. Evidence is scoped to reviewed artifacts, not a guarantee about every version/platform.
  • CDT: distro/component mapping deferred: conda_cdt_repackage. These can contain identifiable upstream libraries and executables. They are not resolved security exclusions; distro backports and component/payload scope need review.
  • Legacy decisions needing review: reasonless explicit decisions. Do not invent rationale or assume an empty payload.

These cohorts sum to explicitly_unmapped. Each shows known download totals and packages with unknown download counts separately; downloads are not exposure. The inspector defaults to all missing-primary packages and provides individual cohort filters, keeping deferred and legacy work visible. Baseline reports remain compatible; no projected PR counts are added to the effective merged snapshot.

Alternative PURLs and CPEs do not count as primary PURLs. Each missing-primary package retains its recorded URLs, note, version, and download count for investigation.

Every unresolved package has one conservative diagnostic_reason, summarized in unresolved_by_diagnostic:

  • recorded_processing_error — the persisted note starts with fetch error:;
  • alternative_only — a canonical alternative PURL exists without a primary;
  • no_parseable_source_host — no hostname can be parsed from the persisted source, repository, or homepage URLs;
  • no_primary_from_url_evidence — URL-host evidence exists but no primary PURL was produced.

Diagnostics describe recorded evidence, not authoritative root causes. Notes can survive reviewed overrides, and missing evidence does not imply intentional exclusion. Only unmapped: true establishes an explicit no-PURL decision. The diagnostic counts are mutually exclusive and sum to unresolved.

Reviewed intentional no-PURL decisions may include an unmapped_reason with a stable code, rule ID, explanation, copied evidence, and review attribution. Generate deterministic review candidates without changing source mappings:

pixi run -e lite mappings:classify-unmapped \
  --input web/public/mappings.json \
  --output .tmp/unmapped-candidates.json

Classification requires explicit review and promotion through a normal mapping contribution. Rules never write mappings/auto.json, never overwrite an existing primary or alternative PURL, and do not classify from a broad host, suffix, or the word “metapackage” alone. Historical reviewed no-PURL decisions without structured rationale remain labeled as legacy rather than receiving fabricated explanations.

The read-only Generate no-PURL review candidates workflow runs after relevant changes on main, weekly, and on manual dispatch. Its job summary shows proposed reason counts and a bounded package preview; its 30-day artifact contains the full candidate evidence. The workflow never creates a contribution or changes a mapping, so downloading the artifact and promoting an exact reviewed list through a normal PR remain explicit human actions.

A CDT conda package is an architecture/toolchain wrapper assembled from a Linux distribution RPM, not a publication of that RPM in its native repository. A pkg:rpm identity would identify the source distribution artifact rather than the independently versioned conda wrapper, so the reviewed classification keeps the RPM URL as evidence without asserting it as the package's PURL. This policy does not establish absence of an upstream identity or vulnerability coverage. Source-component relationships and distribution version-release/backport semantics remain deferred; upstream CPE affected-version ranges cannot simply be applied to a distro-patched binary. This reporting change does not alter CPE discovery eligibility or authorize automated promotion of these packages. For reviewed metapackages/selectors, artifact evidence records payload-file, dependency, and constraint counts; zero payload is strong evidence for dependency-only packages, while conda-specific activation or pinning files do not create an independent upstream identity.

Each missing-primary row includes normalized, sorted source_hosts derived from its persisted source, repository, and homepage URLs. unresolved_by_source_host counts each host at most once per unresolved package and is sorted by package count, then hostname. Host groups are non-exclusive because one package can cite multiple hosts, so their counts do not sum to unresolved.

Reporting is offline and read-only. Package order is stable, and volatile bundle generation timestamps are omitted. Invalid contracts fail rather than producing apparently successful empty coverage.

CPE flow

CPE discovery is part of identity mapping, not CVE assignment. The retained CPE pipeline proposes NVD vendor/product prefixes and promotes accepted mappings as normal contribution files. Those CPEs then flow through scripts.merge_mappings into the public mapping payload.

Discovery also exposes a separate, human-only signal for split outputs: an identity-less package can be shown with reviewed CPE-backed siblings when all of them have the exact same stored source_url and version. Every reviewed anchor must agree on the complete CPE set, ignoring order. If anchors disagree, the audit emits a conflict and proposes nothing.

Shared source is evidence, not identity equivalence. These candidates do not enter the heuristic accept or ambiguous buckets, are not sent to the AI vetter, and are ignored by automated promotion. There is intentionally no URL normalization, mirror equivalence, package-name suffix stripping, homepage inference, or repository inference. A human must inspect the output payload and submit a normal reviewed contribution. This boundary prevents one component from inheriting another component's identity merely because both are built from the same archive—for example, libegl must not automatically inherit a GLX CPE from another libglvnd output.

CPE list behavior is intentionally unchanged: a contribution containing cpes replaces the prior list. Changing that to union semantics could make explicit UI/pipeline removals impossible, so PFX-1826 does not silently change it. A separate follow-up should decide whether CPE contributions need an explicit replace/union operation before any union behavior is introduced.

flowchart TD
  A[mappings/auto.json + reviewed mappings] --> B[scripts.cpe_discover]
  B --> C[mappings/cpe_candidates/*.json]
  C --> D[scripts.cpe_vet optional]
  D --> E[mappings/cpe_vet/*.json]
  C --> F[scripts.cpe_promote]
  E --> F
  C --> I[shared-source review queue]
  I --> J[human-reviewed contribution]
  F --> G[mappings/contributions/*--cpe-pipeline--*.json]
  G --> H[scripts.merge_mappings]
  J --> H
Loading

Useful commands:

pixi run cpe:discover --top 50
pixi run cpe:vet --dry-run
pixi run cpe:promote --dry-run
pixi run -e lite mappings:merge
pixi run -e lite mappings:validate

Frontend behavior

The GitHub Pages app provides two identity-mapping views:

  • index.html is the PURL/CPE mapping editor;
  • coverage.html is a read-only inspector for missing primary PURLs, with state, diagnostic, source-host, and text filters plus recorded evidence;
  • users can review, edit, approve, or mark PURL mappings as unmapped
  • staged identity edits are saved locally until submitted
  • submitted edits open PRs containing one new file under mappings/contributions/
  • CPEs can be reviewed and edited alongside PURL mappings
  • no CVE dashboard, OpenVEX review, AI CVE queue, or CVE deep-inspection routes are served from this repository

The Worker exposes only:

  • POST /exchange for GitHub OAuth code exchange
  • POST /api/submit for PURL mapping contribution PRs

Downstream CVE consumption

A downstream CVE project should consume web/public/mappings.json or the split mappings-index.json + mapping_packages/*.json payload. It should then:

  1. enumerate conda-forge package versions independently,
  2. use PURLs for OSV/package-ecosystem matching,
  3. use CPE prefixes for NVD matching,
  4. apply OSV/NVD affected-version logic in that downstream project, and
  5. store CVE assignment/review state outside this repository.

Verification

Run these checks before opening a PR:

pixi run -e lite mappings:merge
pixi run -e lite mappings:test
pixi run -e lite mappings:validate
pixi run app:check

For frontend/Worker-only checks:

cd web && npm run build
cd worker && npm run typecheck

About

PURL ↔ conda-forge mapping with auto-inference + edit-via-PR workflow

Resources

Stars

11 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages