This repository contains the automated workflow for combining cinema data for the Clusterflick project.
The combining workflow aggregates transformed data from all venues into a single unified dataset. This consolidated data serves as the primary data source for the Clusterflick application and other downstream consumers.
The workflow executes the combine command to merge all venue data:
npm run combineThis command:
- Reads the transformed data from all venues
- Merges the data into a unified format
- Resolves any data conflicts or duplicates
- Generates a combined dataset with consistent structure
- Publishes the combined data as a release
A second command then builds the departed-movies bundle:
npm run departedFilms the
seen registry knows
about that combine did not produce have finished their run, and this writes
them to departed-movies.json so the website can keep rendering pages that
would otherwise 404.
It is a separate artifact deliberately. Nothing reading combined-data.json —
the match stage, the listings, the payload every visitor downloads — should see
films that are not screening, so these never enter it.
The workflow is automatically triggered when the data caching workflow completes successfully. It can also be triggered manually via workflow dispatch if needed.
After successfully combining the data, this workflow triggers:
- data-matched - For enriching movie data with additional sources
- clusterflick.com - To update the main website with the latest data
- analysis.clusterflick.com - To update the analysis site with the latest data
The workflow requires API keys configured as GitHub secrets:
MOVIEDB_API_KEY- For fetching movie metadata from The Movie DatabasePAT- Personal Access Token for publishing releases and triggering downstream workflows
The code in this repository is licensed under the MIT licence.
The releases are not licensed at all. combined-data.json and
departed-movies.json are internal build artifacts — they exist to feed
clusterflick.com and the rest of the pipeline, and are published here only
because the pipeline runs in the open. They carry third-party metadata
Clusterflick cannot sublicense, they have no schema guarantees, and they change
without notice or a deprecation period.
For data you can use, see the data licence. The exact terms for this repository are in LICENSE-DATA.