From 0d64cedad069e95c60703ecb642289596f7a21ad Mon Sep 17 00:00:00 2001
From: Elena Del Pup <62478999+elenadelpup@users.noreply.github.com>
Date: Thu, 30 Apr 2026 16:00:04 +0200
Subject: [PATCH 1/2] Create a template for the multispecies tutorial
---
_tutorial/MultiSpecies.md | 10 ++++++++++
1 file changed, 10 insertions(+)
create mode 100644 _tutorial/MultiSpecies.md
diff --git a/_tutorial/MultiSpecies.md b/_tutorial/MultiSpecies.md
new file mode 100644
index 0000000..52a9e8e
--- /dev/null
+++ b/_tutorial/MultiSpecies.md
@@ -0,0 +1,10 @@
+---
+layout: docs
+title: "Modelling Multispecies pathways"
+order: 80
+---
+
+Goals:
+1. how is the multispecies modelling of the pathways done
+2. how to answer biological questions using the multispecies model
+
From 5f5c39ea21d8ac83aaf737e393291e0cc8eeb0fb Mon Sep 17 00:00:00 2001
From: Elena Del Pup <62478999+elenadelpup@users.noreply.github.com>
Date: Mon, 1 Jun 2026 23:49:14 +0200
Subject: [PATCH 2/2] Rewrite MultiSpecies tutorial; fix old graph URIs in
tutorials 3 and 4
MultiSpecies.md (full rewrite):
- Explains the new multi-graph structure (gpml-taxonomy-extra,
pathways, ncbitaxon) required for species queries
- 5 working example queries: list species, count pathways per species,
pathways by species name, proteins by taxon ID, species in a pathway
- Documents 6 unresolvable taxa / 54 affected nodes (0.2% of 25,141
annotated DataNodes); explains deprecated taxon and OBO scope issues
- Retains the UniProt federation approach as a secondary alternative
with BIND(IRI(REPLACE(...))) pattern for IRI conversion
- All queries verified against local Virtuoso
3.GeneClusterLinks.md / 4.FederatedQueries.md:
Replace 13 occurrences of outdated graph URI:
FROM
with the correct named graph:
FROM
---
_tutorial/3.GeneClusterLinks.md | 12 +-
_tutorial/4.FederatedQueries.md | 10 +-
_tutorial/MultiSpecies.md | 234 +++++++++++++++++++++++++++++++-
3 files changed, 242 insertions(+), 14 deletions(-)
diff --git a/_tutorial/3.GeneClusterLinks.md b/_tutorial/3.GeneClusterLinks.md
index 75bf736..0a99db8 100644
--- a/_tutorial/3.GeneClusterLinks.md
+++ b/_tutorial/3.GeneClusterLinks.md
@@ -23,7 +23,7 @@ https://plantmetwiki.bioinformatics.nl/sparql
**Graph used in all queries**
```sparql
-FROM
+FROM
```
@@ -51,7 +51,7 @@ To get an overview of how many pathway–BGC links exist, you can list all known
```
SELECT ?pathway ?bgc
-FROM
+FROM
WHERE {
?pathway ?predicate ?bgc .
FILTER(CONTAINS(STR(?bgc), "BGC"))
@@ -77,7 +77,7 @@ Example query (from plantiSMASHLinks.rq):
```
SELECT ?pathway ?plantiSMASH_BGC
-FROM
+FROM
WHERE {
?pathway ?p ?plantiSMASH_BGC .
FILTER(CONTAINS(STR(?plantiSMASH_BGC), "plantiSMASH"))
@@ -104,7 +104,7 @@ Example query (from MIBiGLinks.rq):
```
SELECT ?pathway ?mibig
-FROM
+FROM
WHERE {
?pathway ?p ?mibig .
FILTER(CONTAINS(STR(?mibig), "mibig"))
@@ -124,7 +124,7 @@ Some pathways link to both predicted and curated clusters.
You can retrieve all BGC-related links regardless of source:
```
SELECT ?pathway ?bgc
-FROM
+FROM
WHERE {
?pathway ?p ?bgc .
FILTER(
@@ -152,7 +152,7 @@ PREFIX dcterms:
# Retrieve thalianol pathway
SELECT DISTINCT ?pw (STR(?titleLit) AS ?title)
-FROM
+FROM
WHERE {
ro:0000051 ?gene .
diff --git a/_tutorial/4.FederatedQueries.md b/_tutorial/4.FederatedQueries.md
index 0b73ceb..8af24d5 100644
--- a/_tutorial/4.FederatedQueries.md
+++ b/_tutorial/4.FederatedQueries.md
@@ -19,7 +19,7 @@ https://plantmetwiki.bioinformatics.nl/sparql
**Graph used in all queries**
```sparql
-FROM
+FROM
```
## What is a federated SPARQL query?
@@ -78,7 +78,7 @@ Example (from WikidataTest.rq):
PREFIX gpml:
SELECT ?metabolite ?wikidataItem
-FROM
+FROM
WHERE {
?pathway gpml:hasDataNode ?metabolite .
@@ -101,7 +101,7 @@ Example (from WikidataInChiKeys.rq):
```
SELECT ?metabolite ?inchiKey ?wikidataItem
-FROM
+FROM
WHERE {
?metabolite ?p ?inchiKey .
FILTER(CONTAINS(STR(?p), "InChIKey"))
@@ -134,7 +134,7 @@ Using InChIKeys or ChEBI IDs, you can retrieve:
Example (from FederatedMetabolitesChEBI.rq):
```
SELECT ?metabolite ?chebi
-FROM
+FROM
WHERE {
?metabolite ?p ?chebi .
FILTER(CONTAINS(STR(?chebi), "CHEBI"))
@@ -162,7 +162,7 @@ Using federated queries, you can:
Example (from ListPubMedIDs.rq):
```
SELECT DISTINCT ?pmid
-FROM
+FROM
WHERE {
?pathway ?p ?pmid .
FILTER(CONTAINS(STR(?pmid), "pubmed"))
diff --git a/_tutorial/MultiSpecies.md b/_tutorial/MultiSpecies.md
index 52a9e8e..421e8da 100644
--- a/_tutorial/MultiSpecies.md
+++ b/_tutorial/MultiSpecies.md
@@ -4,7 +4,235 @@ title: "Modelling Multispecies pathways"
order: 80
---
-Goals:
-1. how is the multispecies modelling of the pathways done
-2. how to answer biological questions using the multispecies model
+PlantMetWiki integrates pathway data from **PlantCyc 17.0**, which covers biochemical reactions across **439 plant species** — from *Arabidopsis thaliana* to *Solanum tuberosum*, *Zea mays*, and many non-model species. This section explains how species information is stored and how to query it.
+
+Goals:
+1. Understand how multispecies annotation is modelled in PlantMetWiki
+2. Retrieve species-specific genes, proteins, and pathways using SPARQL
+
+---
+
+## How species annotation works
+
+In PlantMetWiki, **species annotations are at the DataNode level**, not the pathway level. Each gene or protein DataNode carries a `wp:organism` triple pointing to its NCBI Taxonomy IRI. Every pathway collectively covers Viridiplantae (the green plants clade), but individual genes and enzymes are annotated to the specific organism they come from.
+
+This information is stored across **three named graphs** that need to be joined:
+
+| Named graph | What it contains |
+|---|---|
+| `graph/pathways` | Core WikiPathways RDF: pathways, DataNodes, interactions |
+| `graph/gpml-taxonomy-extra` | Per-DataNode species annotations (`wp:organism`) |
+| `graph/ncbitaxon` | NCBITaxon ontology (OBO Foundry, CC0) for label resolution |
+
+**SPARQL endpoint**
+```
+https://sparql-plantmetwiki.bioinformatics.nl/sparql
+```
+
+---
+
+## Query 1 — List all species in PlantMetWiki
+
+Species names are resolved via the local NCBITaxon ontology graph (loaded from the OBO Foundry `ncbitaxon.owl`).
+
+```sparql
+PREFIX wp:
+PREFIX rdfs:
+PREFIX ncbi:
+
+SELECT DISTINCT ?taxonID ?species
+WHERE {
+ GRAPH {
+ ?node wp:organism ?taxonID .
+ FILTER(?taxonID != ncbi:33090) # exclude pathway-level Viridiplantae
+ }
+ GRAPH {
+ ?taxonID rdfs:label ?species .
+ }
+}
+ORDER BY ?species
+```
+
+---
+
+## Query 2 — Count pathways per species
+
+```sparql
+PREFIX wp:
+PREFIX rdfs:
+PREFIX dcterms:
+PREFIX ncbi:
+
+SELECT ?species (COUNT(DISTINCT ?pw) AS ?nPathways)
+WHERE {
+ GRAPH {
+ ?node wp:organism ?taxonID .
+ FILTER(?taxonID != ncbi:33090)
+ }
+ GRAPH {
+ ?node dcterms:isPartOf ?pw .
+ }
+ GRAPH {
+ ?taxonID rdfs:label ?species .
+ }
+}
+GROUP BY ?species
+ORDER BY DESC(?nPathways)
+```
+
+---
+
+## Query 3 — All pathways for a given species (by name)
+
+Replace `"Arabidopsis thaliana"` with any Latin species name.
+
+```sparql
+PREFIX wp:
+PREFIX rdfs:
+PREFIX dc:
+PREFIX dcterms:
+
+SELECT DISTINCT ?pwID (STR(?titleLit) AS ?title) (COUNT(DISTINCT ?dataNode) AS ?nDataNodes)
+WHERE {
+ GRAPH {
+ ?taxon rdfs:label "Arabidopsis thaliana" .
+ }
+ GRAPH {
+ ?dataNode wp:organism ?taxon .
+ }
+ GRAPH {
+ ?dataNode dcterms:isPartOf ?pw .
+ ?pw a wp:Pathway ;
+ dc:identifier ?pwID ;
+ dc:title ?titleLit .
+ }
+}
+GROUP BY ?pwID ?titleLit
+ORDER BY DESC(?nDataNodes)
+```
+
+---
+
+## Query 4 — Proteins in pathways for a given taxon ID
+
+Replace `ncbi:4113` with the NCBI Taxonomy ID of your species
+(4113 = *Solanum tuberosum* / potato).
+
+```sparql
+PREFIX wp:
+PREFIX dc:
+PREFIX dcterms:
+PREFIX ncbi:
+
+SELECT DISTINCT ?protein (STR(?titleLit) AS ?title)
+WHERE {
+ GRAPH {
+ ?protein wp:organism ncbi:4113 .
+ }
+ GRAPH {
+ ?protein a wp:Protein ;
+ dcterms:isPartOf ?pw .
+ ?pw dc:title ?titleLit .
+ }
+}
+ORDER BY ?title
+```
+
+---
+
+## Query 5 — Show all species in a specific pathway
+
+Replace `pmw:PC159` with the pathway ID of interest.
+
+```sparql
+PREFIX wp:
+PREFIX rdfs:
+PREFIX dc:
+PREFIX dcterms:
+PREFIX ncbi:
+PREFIX pmw:
+
+SELECT DISTINCT ?taxonID ?species ?title
+WHERE {
+ GRAPH {
+ ?pw dc:identifier pmw:PC159 ;
+ dc:title ?title .
+ ?dataNode dcterms:isPartOf ?pw .
+ }
+ GRAPH {
+ ?dataNode wp:organism ?taxonID .
+ FILTER(?taxonID != ncbi:33090)
+ }
+ GRAPH {
+ ?taxonID rdfs:label ?species .
+ }
+}
+ORDER BY ?species
+```
+
+---
+
+## Taxonomy coverage and known limitations
+
+PlantCyc 17.0 includes species annotations for **439 unique NCBI taxa** across 25,141 DataNodes. When using the local NCBITaxon ontology (OBO Foundry, CC0 licence), 6 taxa are absent from the OBO Foundry release, affecting 54 GeneProduct and Protein nodes (0.2% of annotated nodes). All affected nodes are gene or enzyme annotations — no metabolite nodes are impacted. One taxon (*Lycopersicon hirsutum*, NCBI:283673) is deprecated and merged into *Solanum habrochaites* (NCBI:62890); the remaining five are valid taxa not included in the OBO Foundry release scope.
+
+Use this query to inspect which taxa in your deployment cannot be resolved:
+
+```sparql
+PREFIX wp:
+PREFIX ncbi:
+PREFIX rdfs:
+
+SELECT ?taxon (STRAFTER(STR(?taxon), "NCBITaxon_") AS ?ncbiID)
+ (COUNT(DISTINCT ?node) AS ?nNodes)
+WHERE {
+ GRAPH {
+ ?node wp:organism ?taxon .
+ FILTER(?taxon != ncbi:33090)
+ }
+ FILTER NOT EXISTS {
+ GRAPH {
+ ?taxon rdfs:label ?label .
+ }
+ }
+}
+GROUP BY ?taxon
+ORDER BY DESC(?nNodes)
+```
+
+---
+
+## Alternative: taxonomy via UniProt federation
+
+If you prefer not to load the NCBITaxon ontology locally, you can federate to the UniProt SPARQL endpoint to resolve taxon labels. UniProt uses a parallel IRI scheme (`http://purl.uniprot.org/taxonomy/4113`) that maps 1:1 to OBO Foundry IRIs (`http://purl.obolibrary.org/obo/NCBITaxon_4113`).
+
+```sparql
+PREFIX wp:
+PREFIX up:
+PREFIX ncbi:
+
+SELECT ?taxon ?scientificName
+WHERE {
+ GRAPH {
+ ?node wp:organism ?taxon .
+ FILTER(?taxon != ncbi:33090)
+ }
+ BIND(IRI(REPLACE(STR(?taxon),
+ "http://purl.obolibrary.org/obo/NCBITaxon_",
+ "http://purl.uniprot.org/taxonomy/")) AS ?uniprotTaxon)
+
+ SERVICE {
+ ?uniprotTaxon up:scientificName ?scientificName .
+ }
+}
+LIMIT 20
+```
+
+**Advantage:** no local ontology required — enriches pathways with taxonomy labels without importing all taxonomy into your database.
+**Disadvantage:** depends on an external endpoint being available; slower for large result sets.
+
+This approach can also answer:
+- What scientific name corresponds to this taxon?
+- Is this species under Viridiplantae?
+- What is the parent taxon?