A knowledge graph build system for brain cell annotation transfer and reporting using the OBASK framework.
- Python 3.8+
- Docker and Docker Compose
- ROBOT (for OWL generation)
-
Create and activate virtual environment:
python3 -m venv .venv source .venv/bin/activate # On Windows: .venv\Scripts\activate
-
Install dependencies:
pip install -e .
docker compose upThis runs the full OBASK pipeline: fetches all source OWL/RDF files, loads them into the triplestore, and builds the Neo4j KG. Wait for all services to complete before proceeding.
Once the KG is loaded, apply post-build Cypher updates (e.g. adding Neo4j labels from taxonomy links):
make update-kgTo preview what would be executed without making changes:
make update-kg-dry-runUpdate statements live in src/cypher_updates/ and are executed in alphabetical order. These extend or modify the KG in ways not covered by the OWL import (e.g. attaching labels that require traversal of named taxonomy individuals).
make allThis will:
- Process ROBOT templates (
src/templates/*.tsv→owl/*.owl) - Generate CSV reports (
src/cypher/*.cypher→reports/*.csv)
- Template Generation:
make generate-templates(creates TSV templates from source data) - OWL Generation:
make owl(processes TSV templates to OWL files) - Report Generation:
make reports(requires Neo4j running)
-
Create Cypher Query: Add a
.cypherfile tosrc/cypher/// Example: src/cypher/my_report.cypher MATCH (n:Cell)-[:MAPS_TO]->(m:Cell) RETURN n.id, m.id, n.label, m.label
-
Run Report Generation:
make reports
This automatically creates
reports/my_report.csv
The Makefile automatically discovers all .cypher files and generates corresponding .csv files with the same name.
-
Create Source Data: Add folder under
src/source_data/my_dataset/src/source_data/my_dataset/ ├── data/ # Input CSV/TSV files └── code/ # Processing scripts └── generate.py # Script to create templates -
Generate Templates:
make generate-templates
-
Process to OWL:
make owl
The system includes comprehensive Whole Mouse Brain (WMB) cell cluster analysis with token parsing, knowledge graph mapping, and hierarchical reports.
Generate detailed WMB token mapping analysis:
# Generate comprehensive WMB token mapping reports
make wmb-token-mappingThis produces:
- Token usage analysis: Parse 6,905 WMB cell clusters into ~29,000 tokens
- Knowledge graph mapping: Map tokens to anatomical regions, genes, and cell types (98.4% success rate)
- Problem token analysis: Identify unmappable tokens for review
- Excel consolidation: Single file with all analyses for easy review
Generate advanced hierarchy and consistency reports:
# Generate most general terms and neurotransmission consistency reports
make wmb-additional-reportsThis produces:
- Most general terms report: For each anatomical/gene mapping, find the highest level in each WMB class branch
- Neurotransmission consistency report: Analyze consistency of neurotransmitter patterns across taxonomy levels (86.5% consistency rate)
Generate ROBOT templates for OWL integration:
# Generate ROBOT templates from WMB mapping results
make wmb-robot-templatesCreates templates linking:
- Cell types to anatomical regions via
CLM_0010001 - Cell types to genes via
CLM_0010003
Comprehensive analysis of MERFISH single-cell spatial transcriptomics data from the Allen Brain Cell Atlas.
Generate detailed cell count reports and regional distribution analysis:
# Generate cell count and proportion reports from CCF data
make cell-count-analysisThis produces:
- Cell count report: Total cells for every taxonomy node (neurotransmitter, class, subclass, supertype, cluster)
- Proportion reports: For each cell set, shows the proportion of cells in each brain structure
- Summary analysis: Overview with top cell sets and regional specificity metrics
Generate full cross-tabulation matrices:
# Generate complete taxonomy × brain region matrices
make taxonomy-matricesCreates comprehensive matrices showing cell counts for each taxonomy level across all brain regions with detailed documentation and usage examples.
- Automatic download: First run downloads ~1.5GB MERFISH dataset from Allen Brain Cell Atlas
- Smart caching: Data cached locally in
src/scripts/cell_counts/resources/aba_cache/ - Git-friendly: Cache directory excluded from version control
- Reusable: Same cached data used by both analysis pipelines
The OWL/RDF files loaded into the KG are listed in config/collectdata/vfb_fullontologies.txt. All sources are fetched from remote URLs at build time. Current sources:
| Source | Notes |
|---|---|
cl.owl |
Cell Ontology |
wmbo-full.owl |
Whole Mouse Brain Ontology — tracks releases/latest |
bgo-full.owl |
HMBA Basal Ganglia Ontology — not fetched by URL; pre-processed by make bgo-local into config/collectdata/local_ontologies/ (see note below) |
CCN20230722.rdf |
WMB taxonomy (named individuals) |
CS20250428.rdf |
BG consensus taxonomy (named individuals) |
CS202210140_non_neuronal.owl |
Human Brain Cell Atlas non-neuronal |
CS202210140_neurons.owl |
Human Brain Cell Atlas neurons |
BG2WMB_AT_map_template.owl |
Generated by this repo — fetched from main |
scFAIR_WHB2WMB_template.owl |
Generated by this repo — fetched from main |
| Location mapping OWLs (×4) | Generated by this repo — fetched from main |
Note on taxonomy vs ontology imports: Full ontology files (wmbo-full, bgo-full) do not currently expose named taxonomy individuals in a form usable by the
make update-kgCypher updates. Those Cypher statements traverse links that only exist in the raw taxonomy RDF files (CCN20230722.rdf, CS20250428.rdf), so both the ontology and taxonomy must be loaded separately. The taxonomy import can be dropped for a given ontology once its full OWL file exposes equivalent named individual links.
Note on duplicated BG cell sets —
bgo-full.owlneedsmake bgo-localbefore a rebuild: upstream asserts every HMBA BG cell set twice, under two ID bases with identical accessions:…/ontology/CS20250428/CS20250428_GROUP_0052(the CAS taxonomy export, which theBG:prefix maps to) and…/ontology/CCN20250428/CS20250428_GROUP_0052(the ontology build). Only the base differs, so the two copies load as separate nodes/subjects with the content split between them — the taxonomy metadata, marker gene symbols and this repo's BG2WMB mappings on the CS copy, and all 178has_exemplar_datalinks from the cell-type classes (the classes carrying theCLM:0010001soma locations andCLM:0010003marker sets) on the CCN copy. Loaded as-is, aBG:cell set cannot reach its locations or markers.
make bgo-localdownloads upstreambgo-full.owl, rewrites the CCN base onto the CS base withsrc/sparql/bgo_unify_id_base.ru(robot query --update), and writes the result toconfig/collectdata/local_ontologies/, which the collectdata container merges from its bind mount.bgo-full.owlis therefore not listed invfb_fullontologies.txt— if you rebuild without running the target, BG is missing from the KG entirely. Re-runmake bgo-local-refreshafter an upstream release. Verify a build withsrc/cypher/BG_duplicate_check.cypher.On the first rebuild after this change, start from clean volumes so previously-loaded CCN copies are purged — nothing in the pipeline deletes. The triplestore self-cleans (
rdf4j.txtdrops and recreates theobaskrepository each run), butupdatesolronly upserts into theontologycore (no delete-by-query) andupdateprodreplays node/relationship transactions without a wipe, so stale copies survive in Solr and in a reused KB container:make bgo-local docker compose down -v # drops obask_data, solr_data, triplestore_data and the KB's anonymous /data docker compose upThis is a workaround; the fix belongs upstream in
Cellular-Semantics/hmba_basal_ganglia_ontology. When it lands, delete the target, the.ruand the local file, and restore thebgo-full.owlURL tovfb_fullontologies.txt.
Note on generated OWL files: Because this repo's own OWL outputs are fetched from the remote
mainbranch, changes to templates or mapping scripts must be pushed and merged before a KG rebuild will pick them up.
All CURIE prefixes are managed in src/utils/prefixes.json (JSON-LD format) as the single source of truth.
Update Neo4j export configuration after modifying prefixes:
make update-neo4j-prefixesDetect missing CURIE prefixes (shown as ns{n}: patterns) in the knowledge graph:
# Generate report of missing namespaces
make detect-missing-namespaces
# Get prefix suggestions via prefix commons (requires internet)
make suggest-missing-prefixesThis produces reports/missing_namespaces_report.csv showing:
- Missing namespace prefixes and their frequency
- Example CURIEs and IRIs for each missing namespace
- Suggested base IRIs that could be used as prefixes
- Automatic suggestions from prefix commons (when available)
Neo4j connection settings (override with make VAR=value):
NEO4J_HOST=localhostNEO4J_PORT=7687NEO4J_USER=neo4jNEO4J_PASS=neo
├── src/
│ ├── cypher/ # Cypher query files (.cypher)
│ ├── templates/ # ROBOT template files (.tsv)
│ ├── utils/ # Shared utilities and tools
│ └── source_data/ # Source datasets with data/ and code/ subfolders
├── reports/ # Generated CSV reports
├── owl/ # Generated OWL files
└── config/ # OBASK configuration (DO NOT EDIT)
make helpFor more details, see CLAUDE.md for development guidelines.