Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
74 changes: 72 additions & 2 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,21 @@ on:
push:
tags:
- 'rel-*'
# Manual dispatch is a rehearsal by default: it assembles the tarball end to end, including
# downloading the three sibling JARs, but stops before publishing. Publishing is gated on
# dry_run, and the tag push path is unaffected. Without this the workflow -- and in
# particular the release-notes wiring below -- could only be exercised by cutting a real
# release. Follows the pattern in dp-desktop-app.
workflow_dispatch:
inputs:
version:
description: "Version to build, without the rel- prefix (e.g. 1.16.0). Must exist as rel-<version> in dp-grpc, dp-service, and dp-desktop-app."
required: true
dry_run:
description: "Build only; do not publish a GitHub Release"
required: false
type: boolean
default: true

env:
DP_SERVICE_REPO: osprey-dcs/dp-service
Expand All @@ -13,6 +28,12 @@ env:
jobs:
build-release:
runs-on: ubuntu-latest

env:
# A dry run assembles and verifies the tarball but never publishes a release. A rel-*
# tag push always publishes; a manual dispatch defaults to a dry run and must opt in.
DRY_RUN: ${{ github.event_name == 'workflow_dispatch' && inputs.dry_run }}

steps:

# Checkout data-platform repo
Expand All @@ -23,10 +44,54 @@ jobs:
# convention in CLAUDE.md. Dependabot (.github/dependabot.yml) keeps the pins current.
uses: actions/checkout@a37ce9120846195fa4ece8f58b268e6043cb2f26 # v3.7.0

# Extract version from tag
# On a tag push the version comes from the tag. On a manual dispatch GITHUB_REF_NAME is
# the branch, not a tag, so the version has to be supplied as an input -- it selects
# which rel-<version> of the sibling repos to download JARs from.
- name: Set VERSION env
shell: bash
run: |
set -euo pipefail
if [[ "${GITHUB_EVENT_NAME}" == "workflow_dispatch" ]]; then
VERSION="${{ inputs.version }}"
VERSION="${VERSION#rel-}"
else
TAG="${GITHUB_REF_NAME}"
if [[ "$TAG" != rel-* ]]; then
echo "::error::Tag does not start with 'rel-': $TAG"
exit 1
fi
VERSION="${TAG#rel-}"
fi
echo "VERSION=${VERSION}" >> $GITHUB_ENV
echo "Building version ${VERSION}"

# Checked up front rather than left to action-gh-release, which fails the job on a
# missing body_path only after the sibling-repo JARs have been downloaded and the
# tarball uploaded. Every release from 1.16.0 on ships notes; see
# doc/developer/release.md.
#
# The path is derived from VERSION rather than GITHUB_REF_NAME because on a manual
# dispatch GITHUB_REF_NAME is the branch, so it would look for doc/release-notes/main.md.
# On a tag push the two are identical.
- name: Verify release notes exist
shell: bash
run: |
echo "VERSION=${GITHUB_REF_NAME#rel-}" >> $GITHUB_ENV
set -euo pipefail
NOTES="doc/release-notes/rel-${VERSION}.md"
if [ ! -f "$NOTES" ]; then
# A dry run rehearses the build before the notes are written, so it warns rather
# than failing; the publishing path always fails. DRY_RUN is a job-level env and
# so always exported -- as the empty string on a tag push, which set -u would
# otherwise trip on.
if [[ "${DRY_RUN:-}" == "true" ]]; then
echo "::warning::Release notes not found at $NOTES -- a real run would fail here."
exit 0
fi
echo "::error::Release notes not found at $NOTES."
echo "::error::Add the notes for ${VERSION} and re-run, or retag once they are on the tagged commit."
exit 1
fi
echo "NOTES_PATH=$NOTES" >> $GITHUB_ENV

# Create directory structure
- name: Create installer directories
Expand Down Expand Up @@ -122,12 +187,17 @@ jobs:
sha256sum data-platform-${{ env.VERSION }}.tar.gz \
> data-platform-${{ env.VERSION }}.tar.gz.sha256

# Everything above this point -- downloading the sibling JARs, assembling the tarball,
# checksumming it, and verifying the notes -- runs in a dry run too, so a rehearsal
# exercises the whole job except the one step that writes.
- name: Publish GitHub Release
if: env.DRY_RUN != 'true'
uses: softprops/action-gh-release@3bb12739c298aeb8a4eeaf626c5b8d85266b0e65 # v2.6.2
with:
files: |
data-platform-${{ env.VERSION }}.tar.gz
data-platform-${{ env.VERSION }}.tar.gz.sha256
body_path: ${{ env.NOTES_PATH }}
overwrite_files: true
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
Expand Down
42 changes: 29 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ This document includes the following details:
- [Requirements and objectives](#requirements-and-objectives)
- [MLDP Project elements](#data-platform-project-elements)
- [Status and milestones](#status-and-milestones)
- [Todo and road map](#todo-and-road-map)
- [Todo and road map](#mldp-todo-and-road-map)
- [Additional documentation](#additional-documentation)
- [Installation and getting started](#installation-and-getting-started)

Expand Down Expand Up @@ -233,19 +233,19 @@ Performance benchmark applications were developed and utilized to evaluate candi

### Data Platform v1.0 (November 2023)

Version 1.0 of the Data Platform includes an initial Java implementation of the Ingestion Service providing a gRPC API and using MongoDB for storing time-series data. The initial ingestion service implementation focuses only on scalar data and with timestamps specified using the "sampling clock" mechanism with start time and sample period. It is accompanied by a performance benchmark application that is used at each stage of development to measure ingestion performance relative to the project goal. The initial implementation exceeds our goal by a comfortable margin, but this will continue to be a focus as the project evolves. [section "Data Platform API"](#data-platform-api) provides more information about the ingestion API.
Version 1.0 of the Data Platform includes an initial Java implementation of the Ingestion Service providing a gRPC API and using MongoDB for storing time-series data. The initial ingestion service implementation focuses only on scalar data and with timestamps specified using the "sampling clock" mechanism with start time and sample period. It is accompanied by a performance benchmark application that is used at each stage of development to measure ingestion performance relative to the project goal. The initial implementation exceeds our goal by a comfortable margin, but this will continue to be a focus as the project evolves. [section "gRPC API"](#grpc-api) provides more information about the ingestion API.

### v1.1 (January 2024)

Version 1.1 includes a Java implementation of the Query Service gRPC API, using the MongoDB database managed by the ingestion service to fulfill client query requests. A variety of API RPC methods for querying time-series data are provided to support the development of clients with varying performance requirements, ranging from streaming methods that return bucketed result data down to simple single response methods that return tabular data. See [section "Data Platform API"](#data-platform-api) for a detailed description of the query API.
Version 1.1 includes a Java implementation of the Query Service gRPC API, using the MongoDB database managed by the ingestion service to fulfill client query requests. A variety of API RPC methods for querying time-series data are provided to support the development of clients with varying performance requirements, ranging from streaming methods that return bucketed result data down to simple single response methods that return tabular data. See [section "gRPC API"](#grpc-api) for a detailed description of the query API.

### v1.2 (February 2024)

Version 1.2 saw changes to the "proto" files defining the gRPC API for the Data Platform to be more consistent and conventional, with corresponding changes to the Java service implementations.

### v1.3 (April 2024)

Version 1.3 provides an initial implementation of the annotation service for adding annotations to archived data and performing queries against those annotations. The primary focus for the initial annotation service implementation was on the data model for associating annotations with data in the archive. The only type of annotation currently supported is a simple user comment, but we will be adding many other types of annotations using the same underlying data model. See [section "Data Platform API"](#data-platform-api) for more details about the annotation data model.
Version 1.3 provides an initial implementation of the annotation service for adding annotations to archived data and performing queries against those annotations. The primary focus for the initial annotation service implementation was on the data model for associating annotations with data in the archive. The only type of annotation currently supported is a simple user comment, but we will be adding many other types of annotations using the same underlying data model. See [section "gRPC API"](#grpc-api) for more details about the annotation data model.

### v1.4 (July 2024)

Expand Down Expand Up @@ -291,31 +291,31 @@ The main focus of v1.13 is a new set of column-oriented data structures in the p

Version 1.14 includes three main new features. First, support for column-level metadata in the Ingestion Service: an optional `ColumnMetadata` field (carrying provenance, tags, and key/value attributes) has been added to all 16 column types in the gRPC API. When present, the metadata is persisted in MongoDB alongside the column data and is restored on query, so the retrieved column equals the original ingested column. A validation layer enforces limits on field lengths and collection sizes. Second, a new PV Metadata API is added to the Annotation Service for creating, querying, retrieving, and deleting `PvMetadata` records that describe the properties of archived PVs. As part of this work, the Query Service methods `queryPvMetadata()` and `queryProviderMetadata()` are renamed to `queryPvStats()` and `queryProviderStats()` to better reflect that they return archive ingestion statistics rather than user-defined metadata. Third, a new Machine Configuration API is added to the Annotation Service for managing named machine configuration records and time-bounded configuration activations. The API supports full CRUD operations for both `Configuration` and `ConfigurationActivation` records, enforces non-overlapping activation intervals per configuration and category, and provides a `getActiveConfigurations()` method for retrieving all configurations active at a given point in time.

### v1.15 (Auguest 2026)
### v1.15 (August 2026)

The two primary features included in version 1.15 are 1) a new "v2 query API", and 2) the initial release of the Python client API library. The v2 query API provides an interface that leverages archive metadata for filtering queries by time, PV metadata, temporal machine configuration, and sample status, a major update to the initial query API that supported only a list of PV names and a time range as search criteria. The initial implementation of the Python client library includes low-level wrappers for calling the MLDP PV metadata, machine configuration, and v2 query APIs, and provides a foundation for building higher-level features and conveniences to support data science applications. The v1.15 release also includes a number of performance improvements and bug fixes.


### v1.16 (September 2026)

Version 1.16 centers on two new API capabilities that span every repository in the ecosystem. First, a new Sample Status API assigns status codes to individual PV samples at specific timestamps, keyed by (PV name, timestamp, domain, layer), supporting automated data cleaning, quality assessment, and MLOps workflows; query methods can filter samples by status, and the API replaces the `DataValue.ValueStatus` field removed in this release, which was never queryable. Second, the DataSet and Annotation APIs — the oldest generation of the Annotation Service — are modernized to the CRUD conventions established by the PV metadata, machine configuration, and sample status APIs, gaining single-record get and delete methods, paging, audit fields, typed calculations columns, and column-level provenance. The release also includes substantial query performance work driven by SLAC deployment reports: the hours-long startup bucket scan is removed, query index bounds are now maintained per PV rather than sized to the longest bucket in the archive, and every bucket query is pinned to the compound index and bounded on both sides. A new metrics framework exports request rates, latency histograms, and per-stage query breakdowns from every service over a Prometheus endpoint, with a slow query log for diagnosing individual queries. The desktop application gains a deployment mode for running against remote services rather than only in-process demonstration services, along with new views for authoring and exploring curated metadata. This release is delivered through a new schema migration mechanism that migrates the database at first startup — see the [release notes](doc/release-notes/rel-1.16.0.md) for the upgrade procedure.

---
## MLDP TODO and Road Map

#### v1.16 Development Plans
* Sample Status API for marking the disposition of individual sample values for automated data cleaning tools (e.g, "suspect value").
* v2 modernization of the original Annotation API to follow conventions established for the new PV metadata and machine configuration APIs
* Python client library
* sample status API interface
* v2 annotation API interface
#### v1.17 Development Plans
* Python client library
* ingestion API interface
* bucket-oriented query interface
* improved metrics capture and observability
* tags / attribute usage API: new API and annotation service handling
* data subscription enhancements to support multiple ingestion servers

#### FY27 Development Priorities
* ingestion API and database schema improvements for handling timestamps with jitter (using shared timestamps between data buckets)
* configurable strategies for improved distribution of data in Mongo shards by ingestion service
* age-based archival of data to external storage while maintaining Mongo indexes
* enhance query API to support query by PV data value
* Python client library enhancements for building ML applications
* Modify data subscription mechanism to support multiple ingestion service instances (e.g., REDIS registry).
* support for large atomic data values that exceed the 16MB Mongo object limit
* tools for ingesting data from EPICS environment

Expand Down Expand Up @@ -362,3 +362,19 @@ current, since pinning otherwise trades supply-chain risk for silent staleness.
If you are adding or editing a workflow, the full convention — how to resolve a SHA correctly,
how to verify one, and why pinning is deliberately kept separate from upgrading — is in
[CLAUDE.md](CLAUDE.md#github-actions-pin-every-uses-to-a-commit-sha).

## release notes

Per-release notes live under [`doc/release-notes/`](doc/release-notes/), one document per
release, covering what changed across the whole ecosystem since the previous release and what
upgrading requires. These are the master notes: the releases in the `dp-grpc`, `dp-service`,
`dp-desktop-app`, and `dp-python-lib` repos point back here.

| Release | Notes |
|---|---|
| 1.16.0 | [rel-1.16.0](doc/release-notes/rel-1.16.0.md) |

Releases before 1.16.0 were documented on the
[GitHub release](https://github.com/osprey-dcs/data-platform/releases) itself, and summarized in
[status and milestones](#status-and-milestones) above. The `rel-*` tags remain the authority on
what any past release contained.
Loading
Loading