Skip to content

Repository files navigation

Machine Learning Data Platform

This is the primary repo for the Machine Learning Data Platform (MLDP), providing project background and including links to the various project elements.

This document includes the following details:

Brief MLDP overview

The MLDP provides tools for:

  1. building an annotated archive of PV time-series data, and
  2. using the archive to build and operate data science applications.

The MLDP features are organized in 4 service-oriented gRPC APIs, including:

  • PV time-series data ingestion
  • metadata and annotation
  • data retrieval
  • data / event subscription

The APIs are handled by scalable Java services using MongoDB to manage the archive.

The MLDP ecosystem includes the following elements:

  • Python and Java API clients for building data science applications
  • desktop and web GUI applications
  • administrative tools and configuration examples

Below is a brief overview of each of the MLDP APIs.

PV time-series data ingestion API

  • Provides optimized structures for ingesting a variety of data types including scalars, multi-dimensional arrays, structures, images, and serialized data.
  • Utilizes streaming API methods and supports static and dynamic load balancer configurations for maximum performance.
  • Includes mechanisms for specifying data provenance and ingestion metadata.
  • Supports both continuous and batch ingestion.

Metadata and annotation APIs

It is useful to think of the MLDP archive as a spreadsheet whose a column for each PV and a row for each timestamp, with a PV sample value in each cell. The metadata and annotation APIs provide a mechanism for attaching metadata at four different dimensions of that spreadsheet. Each is discussed in more detail below.

PV metadata API

The PV metadata API describes the properties of the columns in the conceptual archive spreadsheet. It can be used to answer questions like “what are the properties for PV name BPMS:GUNB:314:X?", which might return details like:

DEVICE=BPMS:GUNB:314 
ELEMENT=BPM1B          
TYPE=MONI
Z= -9.555017                     
S= 0.489650 
AREA=GUNB 
BEAMPATHS={SC_HXR, SC_DIAG0}

Temporal Machine Configuration API

The machine configuration API is a temporal annotation tool, describing the rows in the conceptual archive spreadsheet, intended to answer questions like “what was the machine configuration at 2-Feb-2026 18:04:01?", which might return details like:

PATH=CU_HXR                 
E=14.6 (GeV); 
RATE=10000 (Hz)             
MODE=09; 
DEST=CXI                           
EXP=CXI_3443

Annotation / Calculations API

The annotation API is used to describe one or more blocks of data (e.g., a set of PVs over a range of time) in the conceptual archive spreadsheet, and allows linking user-supplied calculations to archive data. This is useful for documenting details for a particular experiment or situation, and showing the provenance of derived data.

Sample Status API

The sample status API is primarily a data cleaning tool for indicating the quality or disposition of individual PV sample values (the cells in the conceptual archive), such as in the example shown below.

Time=141412342134.13412342 secs.nanos
PV=BPMS:GUNB:314:X
Value=2.4
Tag=SUSPECT
Domain=RMS fit

Data Retrieval and Export API

The MLDP provides a query interface that leverages archive metadata for filtering queries by time, PV metadata, temporal machine configuration, and sample status. It isn't natural language or SQL-like, but the following SQL analogy is helpful:

SELECT <PV time-series data>
FROM <the MLDP archive>
WHERE  
  <Time Range Criteria> AND
  <PV Selection Criteria> AND
  <Machine Configuration Criteria> AND
  <Sample Status Criteria>

This supports queries like the following:

"Get me all *good* BPM data from area L3 between time1 and time2 when we were running beam and the experiment was cxi-Q4r4234.”

The query service API incorporates paging, streaming, and multiple query result formats to support a wide range of use cases, and Includes utilities for exporting data to common file formats.

The MLDP Python and Java clients provide high level tools for building ML applications using the query API.

Data / Event Subscription API

The data and event subscription API enables clients to subscribe to live PV data from the ingestion stream, and Provides a mechanism for receiving notification of live data events from the ingestion stream with a configurable window of data around the trigger time. We are investigating approaches for running user code as a plugin, and mechanisms for publishing data using frameworks like Kafka.


Motivation

The Data Platform provides tools for managing the data captured in an experimental research facility, such as a particle accelerator. The data are used within control systems and analytics applications, and facilitate the creation of machine learning models for those applications.

The Data Platform is agnostic to the source and acquisition of the data. A project goal is to manage data captured from the EPICS "Experimental Physics and Industrial Control System", however, use of EPICS is not required. The Data Platform APIs are generic and can be used from essentially all programming languages and any type of application.

How is the Data Patform different from the Epics Archive Appliance?

This is a common question. The Data Platform is optimized for recalling thousands of signals at a single point in time. The Archive Appliance is not. It is good at recalling a small number of signals over a large period of time.

Data Provenance

The Data Platform is for managing data sets - annotating them, deleting them, and using them in the life cycle of the data. One of our use cases is managing experimental data, supporting scenarios like the one described below.

A scientist takes XRay data from some number of detectors, along with some scalar and vector data. The XRay data must be processed as these XRays are taken from different angles at different distances into some normalized coordinate data. The original data must be preserved for verification of published results especially in proton studies. The raw data set is stored in the archive noting important details about the data.

Data scientists normalize the coordinates and upload the normalized data to the archive, linking it to the RAW data set and including details about the code / version of the algorithm used to normalize the coordinates, the date it was run, and the person that performed the normalization.

This normalized data is then processed further to reconstruct the protein structure, creating a new data set that is uploaded to the archive and linked to the normalized data along with information about the data transformation.

Data provenance is a challenging problem and a key feature of the MLDP archive.

Data Cleaning Workflow

Using the same features as for tracking data provenance, the MLDP supports the MLOps data cleaning workflow with tools for identifying suspect data, annotating and marking up that data, downloading data for further processing, and uploading new datasets to the archive that are linked to the datasets from which they are derived.


Requirements and Objectives

  • Provide an API for ingestion of heterogeneous time-series data including scalar values, arrays, structures, and images.
  • Handle the data rates expected for an experimental research facility such as a particle accelerator. A baseline performance requirement is to handle 4,000 scalar data sources sampled at 1 KHz, or 4 million samples per second.
  • Provide an API for retrieval of ingested time-series data.
  • Provide mechanisms for adding post-ingestion annotations and calculations to the archive, and performing queries over those annotations.
  • Provide an API for exploring metadata for data sources available in the archive.
  • Provide mechanism for exporting data from the archive to common formats.

Data Platform Project Elements

The Data Platform includes the following technical components:

  • An API built upon the gRPC communication framework.
  • A suite of services built using the Java programming language, implementing the gRPC service APIs.
  • Utilities for deploying and managing the ecosystem.
  • High-level Python and Java client libraries for building applications.
  • A JavaFX desktop GUI application for navigating the data archive.
  • A JavaScript web application for exploring the data archive.
  • Benchmarks for comparing alternative technologies.

Each of these elements is described in more detail below.

gRPC API

The Data Platform API is built upon the gRPC open-source high-performance remote procedure call (RPC) framework. As described on Wikipedia, "this framework was originally developed by Google for use in connecting microservices. It uses HTTP/2 for transport, protocol buffers as the interface description languages, and provides features such as authentication and bidirectional streaming. It generates cross-platform client and server bindings for many languages."

We chose to use the gRPC framework for the Data Platform API because it can meet our performance requirements for data ingestion, and bindings are provided for virtually any programming language.

The API definition is managed separately from the service implementations so that it can be utilized for building client applications that are independent of other Data Platform technology. The Data Platform gRPC API is documented in the dp-grpc repo.

Service Implementations

The Data Platform Services are implemented as Java server applications. There are four independent server applications, providing ingestion, streaming / subscription, query, and annotation services, respectively. The MongoDB document-oriented database management system is used by the services for persistence. The dp-service repo provides more detail about the Java service implementations and the frameworks used to build them.

Python and Java Client API Library

A Python client API library for the MLDP is under development. It currently provides low-level wrappers around the API methods. It is a development priority to add additional high-level features and conveniences for building data science applications.

An initial Java client API library was created that provides interfaces to both the MLDP ingestion and query service APIs. The library requires additional development work to reflect newly added ingestion and query API methods, but this is not a current development priority.

Desktop GUI Application

Though not a primary project requirement, we decided it was useful to build a Java desktop GUI application to demonstrate the features of the MLDP. However, instead of making an application that can only be used as a demo, we decided to build a full-featured tool useful for navigating the MLDP data archive. It provides a user interface for navigating archive metadata and time-series data, viewing and creating annotations, and other tools for visualizing and exporting data. The application uses the MLDP gRPC API and provides a useful reference for calling those APIs from a Java client. The application is managed in the dp-desktop-app repo, which contains details for installing and using the GUI application.

Please note that the desktop GUI application requires development work to support APIs that have been added since it was created, but this is not a current development priority.

Web Application

The Data Platform Web Application is under development using the JavaScript React framework. It will provide similar features to the desktop GUI application. The dp-web-app repo contains the JavaScript code for the Data Platform Web Application, with documentation about the project.

The initial proof-of-concept project was completed, but, as with the desktop GUI application, the web application requires development work to support APIs that have been added since it was created, but this is not a current development priority.

The desktop GUI application offers a more complete set of functionality than the web application, so one option under consideration is to re-build the web application so that it more closely mirrors the features in the desktop app.

Installation and Deployment Support Tools

A set of utilities is provided to help manage the Data Platform ecosystem. There are scripts for managing infrastructure services including MongoDB and the Envoy proxy (used for deploying the web application), and a set of simple process-management utilities for managing the Data Platform server and benchmark applications.

There are also configuration files for running the Data Platform ecosystem via Docker (with statically configured Envoy Ingestion Service load balancer) and Kubernetes (with dynamic load balancing of all services).

The scripts and utilities for managing the components of the Data Platform ecosystem previously were managed in a separate "dp-support" repo, but have been moved to the data-platform repo in order to streamline project management.

Technology Benchmarks

The dp-benchmark repo is not currently active, but contains code developed for evaluating the performance of some candidate technologies considered for use in the Data Platform service technology stack. It includes an overview of the benchmark process with a summary of results.


status and milestones

"datastore" prototype (2022)

A prototype implementation of the Data Platform services was built focusing on the creation of a general API supporting ingestion and query of heterogeneous data types including scalar, array / table, structure, and image. Service implementations were created using Java for both the Ingestion and Query Services, as well as libraries for building client applications. The prototype technology stack included both InfluxDB (for time series data) and MongoDB (for metadata). This prototype successfully demonstrated the use of gRPC APIs for ingestion and retrieval of heterogeneous, but did not meet the baseline performance requirements.

datastore web application prototype (2022)

The datastore prototype included development of a web application using JavaScript React and Tailwind libraries. The prototype provided simple user interfaces for navigating metadata, as well as querying and displaying time-series data. It demonstrated calling gRPC APIs from a browser-based application using the gRPC Web JavaScript implementation of gRPC for browser clients.

technology performance benchmarking (September 2023)

Performance benchmark applications were developed and utilized to evaluate candidate technologies for use in the Data Platform implementation in light of the project performance goal stated above. Benchmarks focused on gRPC for API communication; InfluxDB, MongoDB and MariaDB for database storage; and writing JSON and HDF5 files to disk. The benchmark results showed that it was likely we could build service implementations meeting our performance requirements by using gRPC for communication and MongoDB for storing "buckets" of time series data.

Data Platform v1.0 (November 2023)

Version 1.0 of the Data Platform includes an initial Java implementation of the Ingestion Service providing a gRPC API and using MongoDB for storing time-series data. The initial ingestion service implementation focuses only on scalar data and with timestamps specified using the "sampling clock" mechanism with start time and sample period. It is accompanied by a performance benchmark application that is used at each stage of development to measure ingestion performance relative to the project goal. The initial implementation exceeds our goal by a comfortable margin, but this will continue to be a focus as the project evolves. section "gRPC API" provides more information about the ingestion API.

v1.1 (January 2024)

Version 1.1 includes a Java implementation of the Query Service gRPC API, using the MongoDB database managed by the ingestion service to fulfill client query requests. A variety of API RPC methods for querying time-series data are provided to support the development of clients with varying performance requirements, ranging from streaming methods that return bucketed result data down to simple single response methods that return tabular data. See section "gRPC API" for a detailed description of the query API.

v1.2 (February 2024)

Version 1.2 saw changes to the "proto" files defining the gRPC API for the Data Platform to be more consistent and conventional, with corresponding changes to the Java service implementations.

v1.3 (April 2024)

Version 1.3 provides an initial implementation of the annotation service for adding annotations to archived data and performing queries against those annotations. The primary focus for the initial annotation service implementation was on the data model for associating annotations with data in the archive. The only type of annotation currently supported is a simple user comment, but we will be adding many other types of annotations using the same underlying data model. See section "gRPC API" for more details about the annotation data model.

v1.4 (July 2024)

Version 1.4 adds Ingestion and Query Service support for all data types defined in the Data Platform API including scalars, multi-dimensional arrays, structures, and images. Support is also added to both services for ingesting and querying data with an explicit list of data timestamps to complement the existing support for specifying data timestamps using a SamplingClock (with start time, sample period, and number of samples). Both features utilize serialization of the protobuf DataColumn and DataTimestamps API objects as byte array fields of the MongoDB BucketDocument. This change improves ingestion performance significantly, while also reducing the MongoDB storage footprint and simplifying the codebase.

v1.5 (August 2024)

With version 1.5, we have now completed the implementation of the initial Data Platform API we defined at the outset for the Core Services. This version focuses on adding the remaining unimplemented Ingestion Service features including: unidirectional client-side streaming data ingestion API, API for registering providers, API for querying ingestion request status details, testing for handling of value status information, validation of data ingestion providers, as well as improvements to the performance benchmark framework. The java dp-grpc and dp-service projects are updated to use Java 21 and the latest versions of 3rd party libraries.

v1.6 (October 2024)

Version 1.6 includes a new Annotation Service API method for exporting time-series data from the archive to common file formats including HDF5, CSV, and XLSX (Excel). It provides an enhancement to the annotations query API method for filtering annotations by dataset id, in addition to the previously supported methods for filtering by owner and comment text field content. Support is also added for querying datasets and annotations by id.

v1.7 (January 2025)

Version 1.7 includes a new Ingestion Service API method for subscribing to data received in the ingestion stream, enabling downstream processing. It also includes a prototype data event monitoring framework, implemented as a new "Ingestion Stream Service". We decided to discontinue development on the new service for now. We envision that the functionality in this prototype will probably be divided between the new Ingestion Stream Service and a client application framework for building data event monitors and more general algorithm data processing. This partitioning will allow the user to create data event monitoring applications with computation that would be impossible to implement in a general way as a service. We intend to revisit the design and partitioning of functionality for the service and application framework in an upcoming release.

v1.8 (March 2025)

The primary focus of version 1.8 is an expanded API for creating and querying Annotations. The Annotation API is redesigned to support modular annotations including components for free-form text comments, linking of associated datasets and other annotations, user-defined calculations, and additional descriptive fields. This release also includes two new API methods for querying details and ingestion stats for data Providers, queryProviders() and queryProviderMetadata(). Behind the scenes changes include some bulk renaming of Java classes to follow a more consistent naming convention, and a more unified approach to the Java BSON document class framework used to store data in MongoDB for the application entities.

v1.9 (May 2025)

The main new features added in version 1.9 are 1) a mechanism for including user-defined Calculations alongside PV time-series data in the API for exporting data to CSV, XLSX, and HDF5 files, and 2) facilities for sending byte data in the ingestion, query, and subscription APIs for improved performance (on the order of 2-3x improvement for ingestion and query) by eliminating extra serialization operations in the gRPC communication framework and service implementations. The release also includes enhancements to the data query handling framework, and updates and testing to use MongoDB 8 as the official reference version for the Data Platform.

v1.10 (July 2025)

Version 1.10 includes the new Ingestion Stream Service providing a mechanism for subscribing to "data events". Using the subscribeDataEvent() API, a client registers one or more triggers each specifying a PV name, a condition (e.g., equal to, greater than, less than, etc.), and a trigger data value. When the condition is triggered by data in the ingestion stream for the specified PV, the client receives an Event notification that specifies the event time, condition that was triggered, and the data value that triggered the event. The client can optionally register to receive EventData for a list of PVs when an Event is triggered for a window of time offset from the event trigger time. This is useful for monitoring data conditions in "real-time", and building models and applications that respond to conditions in the data ingestion stream. The data event monitoring mechanism uses the Ingestion Service's data subscription API to receive data from the ingestion stream for specified PVs. The release includes improvements to data subscription handling in support of the new data event subscription API, as well as new Data Platform ecosystem scripts for managing the new Ingestion Stream Service.

v1.11 (September 2025)

Though not a primary project requirement, we decided it was useful to build a Java desktop GUI application to demonstrate the features of the MLDP. However, instead of making an application that can only be used as a demo, we decided to build a full-featured tool useful for navigating the MLDP data archive. Version 1.11 includes a new application that provides a user interface for navigating archive metadata and time-series data, viewing and creating annotations, and other tools for visualizing and exporting data. The application uses the MLDP gRPC API and provides a useful reference for calling those APIs from a Java client. The application is managed in the dp-desktop-app repo, which contains details for installing and using the GUI application.

v1.12 (January 2026)

Version 1.12 focused on improved deployment support. The contents of the dp-support repo, including tools for managing the MLDP ecosystem, are moved to the data-platform repo in order to streamline project management. Sample Docker scenarios are provided for running the full MLDP ecosystem via Docker (including MongoDB and a statically configured Envoy load balancer), and for running MongoDB in single- or multi-node replica set configurations. A configuration is provided as a staring place for running the full MLDP ecosystem via Kubernetes with dynamic horizontal scaling of all MLDP services. Each project repo provides Github Actions CI workflows for automatically running regression tests, building artifacts, and publishing releases in response to repo events like pull requests and tags. The MongoDB configuration for MLDP services is simplified to use a single URI parameter for connecting to the database.

v1.13 (February 2026)

The main focus of v1.13 is a new set of column-oriented data structures in the protobuf API for data ingestion, with the primary objective of reducing per-request memory allocation and garbage collection in the JVM. The original API definition included the DataColumn and DataValue messages, and was a simple mechanism for supporting heterogeneous data types in a single API data structure including scalars, arrays, structures, and images. The downside of this approach is that it leads to per-sample memory allocation in the Ingestion Service, so that handling an incoming ingestion request containing a bucket of 1000 samples creates a Java DataValue object for each sample, in addition to the other overhead of the request. In v1.13, we've added new column-oriented data structures like DoubleColumn and DoubleArrayColumn that are read in the Ingestion Service as a single primitive array of values, avoiding the per-sample JVM allocation. For a facility continuously ingesting data for 4000 PVs at a rate of 1 kHz, this avoids almost 20 TB of JVM memory allocation / garbage collection in a 24 hour period. An added benefit of the approach is that individual scalar data values are now visible in MongoDB BucketDocuments, where previously the data values were stored as a binary blob.

v1.14 (May 2026)

Version 1.14 includes three main new features. First, support for column-level metadata in the Ingestion Service: an optional ColumnMetadata field (carrying provenance, tags, and key/value attributes) has been added to all 16 column types in the gRPC API. When present, the metadata is persisted in MongoDB alongside the column data and is restored on query, so the retrieved column equals the original ingested column. A validation layer enforces limits on field lengths and collection sizes. Second, a new PV Metadata API is added to the Annotation Service for creating, querying, retrieving, and deleting PvMetadata records that describe the properties of archived PVs. As part of this work, the Query Service methods queryPvMetadata() and queryProviderMetadata() are renamed to queryPvStats() and queryProviderStats() to better reflect that they return archive ingestion statistics rather than user-defined metadata. Third, a new Machine Configuration API is added to the Annotation Service for managing named machine configuration records and time-bounded configuration activations. The API supports full CRUD operations for both Configuration and ConfigurationActivation records, enforces non-overlapping activation intervals per configuration and category, and provides a getActiveConfigurations() method for retrieving all configurations active at a given point in time.

v1.15 (August 2026)

The two primary features included in version 1.15 are 1) a new "v2 query API", and 2) the initial release of the Python client API library. The v2 query API provides an interface that leverages archive metadata for filtering queries by time, PV metadata, temporal machine configuration, and sample status, a major update to the initial query API that supported only a list of PV names and a time range as search criteria. The initial implementation of the Python client library includes low-level wrappers for calling the MLDP PV metadata, machine configuration, and v2 query APIs, and provides a foundation for building higher-level features and conveniences to support data science applications. The v1.15 release also includes a number of performance improvements and bug fixes.

v1.16 (September 2026)

Version 1.16 centers on two new API capabilities that span every repository in the ecosystem. First, a new Sample Status API assigns status codes to individual PV samples at specific timestamps, keyed by (PV name, timestamp, domain, layer), supporting automated data cleaning, quality assessment, and MLOps workflows; query methods can filter samples by status, and the API replaces the DataValue.ValueStatus field removed in this release, which was never queryable. Second, the DataSet and Annotation APIs — the oldest generation of the Annotation Service — are modernized to the CRUD conventions established by the PV metadata, machine configuration, and sample status APIs, gaining single-record get and delete methods, paging, audit fields, typed calculations columns, and column-level provenance. The release also includes substantial query performance work driven by SLAC deployment reports: the hours-long startup bucket scan is removed, query index bounds are now maintained per PV rather than sized to the longest bucket in the archive, and every bucket query is pinned to the compound index and bounded on both sides. A new metrics framework exports request rates, latency histograms, and per-stage query breakdowns from every service over a Prometheus endpoint, with a slow query log for diagnosing individual queries. The desktop application gains a deployment mode for running against remote services rather than only in-process demonstration services, along with new views for authoring and exploring curated metadata. This release is delivered through a new schema migration mechanism that migrates the database at first startup — see the release notes for the upgrade procedure.


MLDP TODO and Road Map

v1.17 Development Plans

  • Python client library
    • ingestion API interface
    • bucket-oriented query interface
  • tags / attribute usage API: new API and annotation service handling
  • data subscription enhancements to support multiple ingestion servers

FY27 Development Priorities

  • ingestion API and database schema improvements for handling timestamps with jitter (using shared timestamps between data buckets)
  • configurable strategies for improved distribution of data in Mongo shards by ingestion service
  • age-based archival of data to external storage while maintaining Mongo indexes
  • enhance query API to support query by PV data value
  • Python client library enhancements for building ML applications
  • support for large atomic data values that exceed the 16MB Mongo object limit
  • tools for ingesting data from EPICS environment

Longer Term Plans

  • JWT authentication / role-based authorization / LDAP integration
  • data ownership and sharing mechanisms
  • enhancements to desktop and web GUI applications
  • Java client library catch up to v2 API enhancements
  • Mechanism for executing user-defined code in ingestion stream.

Additional Documentation

Use the links below to learn more about the Data Platform project, or the links above to navigate to the other project repositories.

installation and getting started

project documents

developer notes

GitHub Actions are pinned to commit SHAs

Every uses: in every workflow, in every repo in this org, is pinned to a full 40-character commit SHA with a trailing # vX.Y.Z comment:

uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1

A tag is mutable — whoever controls an action's repo can repoint it at different code, and every workflow picks that up on its next run with no diff and no review. A commit SHA cannot be repointed. Each repo also carries a .github/dependabot.yml that keeps the pinned versions current, since pinning otherwise trades supply-chain risk for silent staleness.

If you are adding or editing a workflow, the full convention — how to resolve a SHA correctly, how to verify one, and why pinning is deliberately kept separate from upgrading — is in CLAUDE.md.

release notes

Per-release notes live under doc/release-notes/, one document per release, covering what changed across the whole ecosystem since the previous release and what upgrading requires. These are the master notes: the releases in the dp-grpc, dp-service, dp-desktop-app, and dp-python-lib repos point back here.

Release Notes
1.16.0 rel-1.16.0

Releases before 1.16.0 were documented on the GitHub release itself, and summarized in status and milestones above. The rel-* tags remain the authority on what any past release contained.

About

Parent repo for the data platform project

Resources

Stars

2 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages