From 35ce053b56c3638d6a6155b9d9f32b22532fdea5 Mon Sep 17 00:00:00 2001 From: jstirnaman <212227+jstirnaman@users.noreply.github.com> Date: Thu, 1 Oct 2026 06:20:45 +0000 Subject: [PATCH] sync(plugins): update InfluxDB 3 plugin documentation --- .../official/geo-enrichment.md | 58 ++- .../plugins-library/official/import.md | 374 +++++++++++---- .../official/prophet-forecasting.md | 39 +- .../official/schema-validator.md | 427 ++++++++++-------- .../plugins-library/official/signal-filter.md | 56 ++- .../official/system-metrics.md | 20 +- data/influxdb3_plugins.yml | 10 +- 7 files changed, 676 insertions(+), 308 deletions(-) diff --git a/content/shared/influxdb3-plugins/plugins-library/official/geo-enrichment.md b/content/shared/influxdb3-plugins/plugins-library/official/geo-enrichment.md index fa0f3f0ee9..cba539d00c 100644 --- a/content/shared/influxdb3-plugins/plugins-library/official/geo-enrichment.md +++ b/content/shared/influxdb3-plugins/plugins-library/official/geo-enrichment.md @@ -168,11 +168,13 @@ consumer GPS error. Raise it where exact edge behavior matters. ### HTTP body parameters Every parameter above may be given in the request body under the same name, -plus the backfill-only fields here. The body is the last layer: trigger -arguments, then the [TOML file](#toml-configuration) they name, then the body, -each overriding the one before. A trigger created without arguments is -configured by the body alone; one created with them holds the defaults each -request overrides only where it names them. +plus the backfill-only fields here. A request carries three layers of its own — +the body, then the headers, then the query string, each overriding the one +before — and all three override what the trigger holds: the +[environment](#environment-variables), the trigger arguments, then the +[TOML file](#toml-configuration) they name. A trigger created without arguments +is configured by the request alone; one created with them holds the defaults +each request overrides only where it names them. A body overrides a setting but cannot unset one: an empty string or `null` counts as absent, so the trigger's value stands, and a body naming only `start` @@ -205,18 +207,52 @@ Use `retry_unknown` after widening `max_radius_m`, and `force` after redrawing a zone — those rows already hold a resolved value, so `retry_unknown` would pass over them. The reference file is re-read on every HTTP call. +#### Headers and query parameters + +The same names reach the plugin as headers, spelled +`X-Influxdb3-Geo-Enrichment-` with underscores written as hyphens +(`X-Influxdb3-Geo-Enrichment-Max-Radius-M` sets `max_radius_m`), and as +query-string parameters, spelled exactly like the parameter +(`?retry_unknown=true`). Header names are matched regardless of casing, which +RFC 9110 makes meaningless. + +```bash +curl -X POST "http://localhost:8181/api/v3/engine/geo_backfill?retry_unknown=true" \ + -H "Authorization: Bearer $INFLUXDB3_AUTH_TOKEN" \ + -H "Content-Type: application/json" \ + -H "X-Influxdb3-Geo-Enrichment-Max-Radius-M: 5000" \ + -d '{"start": "2026-08-01T00:00:00Z", "end": "2026-08-29T00:00:00Z"}' +``` +A header the plugin does not ask for is ignored, since a client sends headers of +its own on every request; an unknown query parameter is a 400, like an unknown +body field, so a misspelled name cannot pass unnoticed. Neither layer can name +`config_file_path`, for the reason the body cannot. + +### Environment variables + +Every parameter can also come from an environment variable named +`INFLUXDB3_GEO_ENRICHMENT_` in upper case — for example, +`INFLUXDB3_GEO_ENRICHMENT_REFERENCE_FILE` sets `reference_file`. The +environment is the lowest layer: a trigger argument overrides it, the TOML file +overrides both, and on the HTTP trigger the request — its body, its headers and +its query string — overrides them all. +`INFLUXDB3_GEO_ENRICHMENT_CONFIG_FILE_PATH` names the TOML file when the trigger +doesn't carry a `config_file_path` argument. + ### TOML configuration | Parameter | Type | Default | Description | |--------------------|--------|-----------|-----------------------------------------------------| | `config_file_path` | string | *(empty)* | `.toml` file, relative to `PLUGIN_DIR` or absolute. | -A trigger argument on either trigger; the file's values override the other -trigger arguments, and on the HTTP trigger the request body overrides the file. -The path is never read from a request body. It names a layer rather than -setting a value: a body that could choose which file the trigger reads would -take the trigger's configuration out of the operator's hands, so the body -refuses it. +A trigger argument on either trigger, or the +`INFLUXDB3_GEO_ENRICHMENT_CONFIG_FILE_PATH` variable when the trigger leaves it +unset; the file's values override the other trigger arguments, and on the HTTP +trigger the request overrides the file. The path is never read from a request — +not from its body, its headers or its query string. It names a layer rather than +setting a value: a request that could choose which file the trigger reads would +take the trigger's configuration out of the operator's hands, so the body and +the query string refuse it and a header spelling it is dropped. On the HTTP trigger the file is how a long setup is named once instead of repeated in every backfill request; each call then carries only what differs — diff --git a/content/shared/influxdb3-plugins/plugins-library/official/import.md b/content/shared/influxdb3-plugins/plugins-library/official/import.md index 5f7667ca3e..5f5bca4d0a 100644 --- a/content/shared/influxdb3-plugins/plugins-library/official/import.md +++ b/content/shared/influxdb3-plugins/plugins-library/official/import.md @@ -3,7 +3,10 @@ > **Note:** This plugin requires {{% product-name %}}.8.2 or later. -The InfluxDB Import Plugin enables seamless data import from InfluxDB v1, v2, or v3 instances to {{% product-name %}}. It provides comprehensive import capabilities with pause/resume functionality, progress tracking, conflict detection, and robust error handling. The plugin operates via HTTP endpoints, allowing you to start, pause, resume, cancel, and monitor imports through simple HTTP requests. +The InfluxDB Import Plugin enables seamless data import from InfluxDB v1, v2, or v3 instances to {{% product-name %}} +Core/Enterprise. It provides comprehensive import capabilities with pause/resume functionality, progress tracking, +conflict detection, and robust error handling. The plugin operates via HTTP endpoints, allowing you to start, pause, +resume, cancel, and monitor imports through simple HTTP requests. Key features: - Import data from InfluxDB v1, v2, or v3 to {{% product-name %}} @@ -20,29 +23,36 @@ Key features: ## Configuration -Plugin parameters may be specified as key-value pairs in the `--trigger-arguments` flag (CLI), in the `trigger_arguments` field (API) when creating a trigger or via body of HTTP request. This plugin supports TOML configuration files, which can be specified using the `config_file_path` parameter. +Plugin parameters may be specified as key-value pairs in the `--trigger-arguments` flag (CLI), in the`trigger_arguments` +field (API) when creating a trigger or via body of HTTP request. This plugin supports TOML configuration files, which +can be specified using the `config_file_path` parameter. + +A key that is not a plugin parameter fails the request and is named in the error — in the trigger arguments, in the TOML +file and in the request body alike. ### Plugin metadata -This plugin includes a JSON metadata schema in its docstring that defines supported trigger types and configuration parameters. This metadata enables the [InfluxDB 3 Explorer](https://docs.influxdata.com/influxdb3/explorer/) UI to display and configure the plugin. +This plugin includes a JSON metadata schema in its docstring that defines supported trigger types and configuration +parameters. This metadata enables the [InfluxDB 3 Explorer](https://docs.influxdata.com/influxdb3/explorer/) UI to +display and configure the plugin. ### Required parameters -| Parameter | Type | Default | Description | -|---------------------|---------|----------|--------------------------------------------------------------------| -| `source_url` | string | required | Source InfluxDB URL (with optional port, for example, `http://localhost:8086`) | -| `influxdb_version` | integer | required | Source InfluxDB version: 1, 2, or 3 | -| `source_database` | string | required | Source database name to import from | +| Parameter | Type | Default | Description | +|--------------------|---------|----------|-------------------------------------------------------------------------| +| `source_url` | string | required | Source InfluxDB URL (with optional port, for example, `http://localhost:8086`) | +| `influxdb_version` | integer | required | Source InfluxDB version: 1, 2, or 3 | +| `source_database` | string | required | Source database name to import from | ### Authentication Credentials are passed via HTTP headers on each request. -| Header | Purpose | -|--------|---------| -| `Source-Token` | Bearer/API token authentication (InfluxDB v1/v2/v3) | -| `Source-Username` | Basic auth username (InfluxDB v1) | -| `Source-Password` | Basic auth password (InfluxDB v1) | +| Header | Purpose | +|-------------------|-----------------------------------------------------| +| `Source-Token` | Bearer/API token authentication (InfluxDB v1/v2/v3) | +| `Source-Username` | Basic auth username (InfluxDB v1) | +| `Source-Password` | Basic auth password (InfluxDB v1) | **Method 1: Token-based authentication** @@ -73,48 +83,55 @@ curl -X POST http://localhost:8181/api/v3/engine/import?action=start \ "source_database": "telegraf" }' ``` -> **Note**: Authentication errors from the source InfluxDB are returned directly. The plugin does not validate credentials upfront. +> **Note**: Authentication errors from the source InfluxDB are returned directly. The plugin does not validate +> credentials upfront. ### Optional parameters -| Parameter | Type | Default | Description | -|----------------------|---------|----------------|---------------------------------------------------------------------------------------------------------------| -| `dest_database` | string | none | Destination database name in {{% product-name %}} (if not specified, uses database where trigger was created) | -| `start_timestamp` | string | none | Import start time (datetime format). If not specified, starts from oldest data | -| `end_timestamp` | string | none | Import end time (datetime format). If not specified, imports to newest data | -| `query_interval_ms` | integer | 100 | Delay between queries in milliseconds to avoid overloading source database | -| `import_direction` | string | "oldest_first" | Import direction: "oldest_first" or "newest_first" | -| `target_batch_size` | integer | 2000 | Target number of rows per query batch | -| `table_filter` | string | none | Dot-separated list of tables to import (for example, "cpu.mem.disk"). If not specified, imports all tables | -| `dry_run` | boolean | false | If true, generates import plan without processing data (shows estimates, schema conflicts, and configuration) | +| Parameter | Type | Default | Description | +|---------------------|---------|----------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `dest_database` | string | none | Destination database name in {{% product-name %}} (if not specified, uses database where trigger was created) | +| `start_timestamp` | string | none | Import start time — RFC3339, Unix seconds or nanoseconds, or `YYYY-MM-DD`; a value naming no zone is read as UTC. If not specified, starts from oldest data | +| `end_timestamp` | string | none | Import end time, same formats as `start_timestamp`. If not specified, imports to newest data | +| `query_interval_ms` | integer | 100 | Delay between queries in milliseconds to avoid overloading source database (0 or greater) | +| `import_direction` | string | "oldest_first" | Import direction: "oldest_first" or "newest_first" | +| `target_batch_size` | integer | 2000 | Target number of rows per query batch (1 or greater) | +| `table_filter` | string | none | Dot-separated list of tables to import (for example, "cpu.mem.disk"). If not specified, imports all tables | +| `dry_run` | boolean | false | If true, generates import plan without processing data (shows estimates, schema conflicts, and configuration) | ### TOML configuration -| Parameter | Type | Default | Description | -|--------------------|--------|---------|----------------------------------------------------------------------------------| -| `config_file_path` | string | none | TOML config file path relative to `PLUGIN_DIR` (required for TOML configuration) | +| Parameter | Type | Default | Description | +|--------------------|--------|---------|-------------------------------------------------------------------------------------------------------------------------------| +| `config_file_path` | string | none | TOML config file path, absolute or relative to the plugin directory. Trigger arguments or `INFLUXDB3_IMPORT_CONFIG_FILE_PATH` | -*To use a TOML configuration file, set the `PLUGIN_DIR` environment variable and specify the `config_file_path` in the trigger arguments.* This is in addition to the `--plugin-dir` flag when starting {{% product-name %}}. +*A relative `config_file_path` is resolved against the plugin directory, taken from `PLUGIN_DIR`, then +`INFLUXDB3_PLUGIN_DIR`, then the parent of the Processing Engine virtual environment. The last of these is always set, +so a file placed next to the plugin is found without any environment variable.* The path comes from the trigger +arguments or from `INFLUXDB3_IMPORT_CONFIG_FILE_PATH`, the argument winning; a request body that carries +`config_file_path` is refused. #### Example TOML configuration [import_config.toml](https://github.com/influxdata/influxdb3_plugins/blob/master/influxdata/import/import_config.toml) -For more information on using TOML configuration files, see the Using TOML Configuration Files section in the [influxdb3_plugins/README.md](https://github.com/influxdata/influxdb3_plugins/blob/master/README.md). +For more information on using TOML configuration files, see the Using TOML Configuration Files section in +the [influxdb3_plugins/README.md](https://github.com/influxdata/influxdb3_plugins/blob/master/README.md). ## Software Requirements - **{{% product-name %}}**: with the Processing Engine enabled. - **Source InfluxDB instance**: InfluxDB v1.x or v2.x instance accessible via HTTP/HTTPS. - **Python packages**: - - `requests` (for HTTP communication with source InfluxDB) + - `influxdata-plugin-utils>=0.5.0` (shared configuration, validation and write helpers) + - `requests` (for HTTP communication with source InfluxDB) ### Installation steps -1. Start {{% product-name %}} with the Processing Engine and `PLUGIN_DIR` environment variable: +1. Start {{% product-name %}} with the Processing Engine: ```bash - PLUGIN_DIR=~/.plugins influxdb3 serve \ + influxdb3 serve \ --node-id node0 \ --object-store file \ --data-dir ~/.influxdb3 \ @@ -123,6 +140,7 @@ For more information on using TOML configuration files, see the Using TOML Confi 2. Install required Python packages: ```bash + influxdb3 install package "influxdata-plugin-utils>=0.5.0" influxdb3 install package requests ``` ## Trigger setup @@ -136,6 +154,7 @@ influxdb3 create trigger \ --database mydb \ --plugin-filename gh:influxdata/import/import.py \ --trigger-spec "request:import" \ + --run-asynchronous \ import_trigger ``` Enable the trigger: @@ -145,6 +164,18 @@ influxdb3 enable trigger --database mydb import_trigger ``` The endpoint is registered at `/api/v3/engine/import`. +### Why `--run-asynchronous` is required + +An import runs inside the HTTP request that starts it, and without this flag the +Processing Engine serves one invocation of a trigger at a time. A `pause`, +`status` or `cancel` request would then wait in line until the import it is +meant to control has already finished. With the flag, the engine runs +invocations of the trigger concurrently, and the control actions are answered +while the import is in progress. + +Without the flag the plugin still imports data correctly — only the control +actions become unusable. + ## HTTP Endpoint The import plugin provides the following type of requests: @@ -171,12 +202,21 @@ Start a new import from source InfluxDB to {{% product-name %}}. "table_filter": "cpu.mem.disk" } ``` +Any of these keys may arrive as a query parameter or as an +`X-Influxdb3-Import-` header instead, both of which override the +body. See "Configuration Priority and Loading" below. + ### Get Import Status Check the status and progress of a import. **Request**: `GET /api/v3/engine/import?action=status&import_id=` +The answer holds `summary` with the table counts and `total_rows_imported`, +`table_details` with one entry per table, and `config` with the settings the +import was started with. A table that failed to write some of its windows +reports them under `table_details[].errors`, and `summary.tables_with_errors` +counts such tables. ### Pause Import @@ -195,17 +235,28 @@ Resume a paused or interrupted import. **Headers**: - `Source-Token: my-token` (or `Source-Username` + `Source-Password`) -> **Note**: Credentials must be provided via headers when resuming. Authentication credentials are not stored for security reasons and must be provided when resuming. Returns error if import is not found, already cancelled, already completed, or actively running. +> **Note**: Credentials must be provided via headers when resuming. Authentication credentials are not stored for +> security reasons and must be provided when resuming. Returns error if import is not found, already cancelled, already +> completed, or actively running. + +The response field `errors` is the number of windows that failed across every run of the import, not just this one. The +failures themselves stay in the `errors` column of `import_state`. #### Crash recovery -If the plugin or the server crashes during an import, the import state may be left as "running" even though nothing is actually running. The resume action handles this with **stale import detection**: +If the plugin or the server crashes during an import, the import state may be left as "running" even though nothing is +actually running. The resume action handles this with **stale import detection**: - When the import state is "running", the plugin checks the timestamp of the last `import_state` record. - If the last update is older than **5 minutes**, the import is considered stale (crashed) and resume is allowed. -- If no `import_state` records exist at all (the import crashed before processing any tables), the import is restarted from the beginning. +- If no `import_state` records exist at all (the import crashed before processing any tables), the import is restarted + from the beginning. + +Each table then continues from its own checkpoint. A table that had written windows resumes after the last of them; a +table that had written none, including one that never started, is imported from the beginning. -If the plugin itself crashes (for example, source database becomes unavailable), it writes a paused state before exiting, so the import can be resumed with a regular resume call after fixing the issue. +If the plugin itself crashes (for example, source database becomes unavailable), it writes a paused state before exiting, so +the import can be resumed with a regular resume call after fixing the issue. ### Cancel Import @@ -245,7 +296,8 @@ Test connectivity to a URL and identify if it's an InfluxDB instance. Uses a 5-s "build": "" } ``` -> **Note**: InfluxDB v3 does not expose version headers without authentication. Detection uses the `cluster-uuid` header instead. +> **Note**: InfluxDB v3 does not expose version headers without authentication. Detection uses the `cluster-uuid` header +> instead. **Failure response** (not InfluxDB or unreachable): ```json @@ -261,7 +313,9 @@ Test connectivity to a URL and identify if it's an InfluxDB instance. Uses a 5-s "message": "Unable to determine InfluxDB version" } ``` -> **Note**: When InfluxDB returns 401/403 without version headers, the connection test cannot determine the version. This typically means authentication is required. The instance is likely InfluxDB, but version detection requires valid credentials. +> **Note**: When InfluxDB returns 401/403 without version headers, the connection test cannot determine the version. +> This typically means authentication is required. The instance is likely InfluxDB, but version detection requires valid +> credentials. ### List Databases @@ -274,6 +328,7 @@ Get list of databases from source InfluxDB instance. - `Content-Type: application/json` **Request body** (JSON): + ```json { "source_url": "http://localhost:8086", @@ -298,8 +353,6 @@ Get list of tables/measurements from a source database. "source_database": "telegraf" } ``` -> **Note**: For InfluxDB v2, include `source_org` in the request body. - ## Example usage ### Example 1: Basic import with token authentication @@ -312,6 +365,7 @@ influxdb3 create trigger \ --database mydb \ --plugin-filename import.py \ --trigger-spec "request:import" \ + --run-asynchronous \ import_trigger influxdb3 enable trigger --database mydb import_trigger @@ -419,7 +473,9 @@ curl -X POST http://localhost:8181/api/v3/engine/import?action=start \ ``` ### Expected results -With `dry_run: true`, the plugin generates a comprehensive import plan **without processing any data**. It only performs: +With `dry_run: true`, the plugin generates a comprehensive import plan **without processing any data**. It only +performs: + - Schema inspection (tags and fields) - Data sampling for time estimation - Conflict detection @@ -449,7 +505,13 @@ The response returns immediately with a detailed import plan: }, "tables": { "total": 5, - "list": ["cpu", "mem", "disk", "network", "processes"], + "list": [ + "cpu", + "mem", + "disk", + "network", + "processes" + ], "filtered": "all tables" }, "estimated_import": { @@ -475,14 +537,19 @@ The response returns immediately with a detailed import plan: { "measurement": "cpu", "type": "tag_field_conflict", - "conflicts": ["host", "region"], + "conflicts": [ + "host", + "region" + ], "resolution": "Tags will be renamed with '_tag' suffix: host -> host_tag, region -> region_tag" } ] } } ``` -**Note**: Dry run mode is fast and lightweight - it does not query or process any actual data points, only metadata. Use it to: +**Note**: Dry run mode is fast and lightweight - it does not query or process any actual data points, only metadata. Use +it to: + - Preview import scope and estimates - Identify schema conflicts before import - Validate configuration and connectivity @@ -492,16 +559,12 @@ The response returns immediately with a detailed import plan: This plugin supports using TOML configuration files to specify all plugin arguments. -### Important Requirements - -**To use TOML configuration files, you must set the `PLUGIN_DIR` environment variable in the {{% product-name %}} host environment.** - ### Setting Up TOML Configuration -1. **Start {{% product-name %}} with the PLUGIN_DIR environment variable set**: +1. **Start {{% product-name %}} with the Processing Engine**: ```bash - PLUGIN_DIR=~/.plugins influxdb3 serve \ + influxdb3 serve \ --node-id node0 \ --object-store file \ --data-dir ~/.influxdb3 \ @@ -536,6 +599,7 @@ This plugin supports using TOML configuration files to specify all plugin argume --plugin-filename import.py \ --trigger-spec "request:import" \ --trigger-arguments config_file_path=import_config.toml \ + --run-asynchronous \ import_trigger ``` 5. **Start import via HTTP** (config from TOML file will be used as defaults, can be overridden in request body): @@ -547,10 +611,30 @@ This plugin supports using TOML configuration files to specify all plugin argume The import plugin loads configuration from multiple sources with the following priority order (highest to lowest): -1. **HTTP Request Body** (highest priority) - JSON parameters in POST request body -2. **TOML Configuration File** - Parameters from file specified in `config_file_path` -3. **Trigger Arguments** - Parameters from `--trigger-arguments` when creating trigger -4. **Environment Variables** (lowest priority) - System environment variables +1. **Query Parameters** (highest priority) - Parameters in the query string of the `action=start` request +2. **Request Headers** - Parameters spelled `X-Influxdb3-Import-` +3. **HTTP Request Body** - JSON parameters in POST request body +4. **TOML Configuration File** - Parameters from file specified in `config_file_path` +5. **Trigger Arguments** - Parameters from `--trigger-arguments` when creating trigger +6. **Environment Variables** (lowest priority) - System environment variables + +Each source overrides only the keys it sets, so a TOML file may hold the stable +settings while the request supplies the time range for one import. The +`config_file_path` comes from the trigger arguments or from +`INFLUXDB3_IMPORT_CONFIG_FILE_PATH`, the argument winning; the request may not +point at a configuration file, and neither may the file itself. + +Only `action=start` assembles its configuration from these layers. The other +actions read less: + +- `action=resume` reads no settings from the request at all. It continues with + the configuration its import was started with, kept in the `import_config` + table. +- `action=test_connection`, `action=databases` and `action=tables` take their + source parameters from the JSON request body only — never from the + environment, the trigger arguments or the TOML file. Each of these actions is + documented above with the keys it needs. `databases` and `tables` also read + the credential headers; `test_connection` only probes the URL and sends none. ### Configuration Loading Process @@ -558,7 +642,7 @@ When a import starts, the plugin loads configuration in this order: ```python # 1. Start with environment variables (lowest priority) -IMPORT_SOURCE_URL, IMPORT_SOURCE_DATABASE, etc. +IMPORT_SOURCE_URL and its four companions, then INFLUXDB3_IMPORT_ * on top # 2. Override with trigger arguments (--trigger-arguments) config_file_path=import_config.toml, source_url=http://localhost:8086, etc. @@ -566,15 +650,71 @@ config_file_path=import_config.toml, source_url=http://localhost:8086, etc. # 3. Override with TOML file contents (if config_file_path specified) [from import_config.toml file] -# 4. Override with HTTP request body (highest priority) +# 4. Override with the HTTP request body { - "source_url": "http://localhost:8086", - ... + "source_url": "http://localhost:8086", + ... } + +# 5. Override with request headers +X-Influxdb3-Import-Source-Url: http://localhost:8086 + +# 6. Override with query parameters (highest priority) +?action=start&source_url=http://localhost:8086 +``` +### Request Headers Supported + +Every parameter except `config_file_path` can come from a header named +`X-Influxdb3-Import-` with underscores written as hyphens — for +example, `X-Influxdb3-Import-Source-Url` sets `source_url` and +`X-Influxdb3-Import-Target-Batch-Size` sets `target_batch_size`. Casing does not +matter, as RFC 9110 asks. A header the plugin does not ask for is ignored, so +the request still carries whatever else the client sends. + +Credentials are **not** among these. They keep their own unprefixed headers, and +`X-Influxdb3-Import-Source-Token` is ignored rather than read as a token: + +```bash +-H "Source-Token: my-super-secret-token" +-H "Source-Username: admin" -H "Source-Password: my-password" +``` +A header value must be ASCII. {{% product-name %}} closes the connection without a reply +when one is not, and the plugin is never invoked, so a `table_filter` naming a +measurement outside ASCII belongs in the request body or the query string. + +### Query Parameters Supported + +Every parameter except `config_file_path` can come from the query string of an +`action=start` request, spelled exactly as the parameter is. Unlike a header, +a query parameter is matched case-sensitively, so `Source_Url` is refused: + +```bash +curl -X POST "http://localhost:8181/api/v3/engine/import?action=start&source_url=http://localhost:8086&source_database=telegraf&influxdb_version=1" ``` +Each action accepts only the parameters it reads, and an unknown one is refused +and named along with what that action accepts — `?action=status&import_ids=abc` +answers `Query parameters may not set 'import_ids'; accepted keys: ['action', +'import_id']`. + +Three cautions specific to the query string: + +- A `+` in a timestamp is decoded as a space, so write the offset as `Z` or + percent-encode it: `start_timestamp=2024-01-01T00%3A00%3A00%2B03%3A00`. +- An empty value is ignored rather than treated as a value, so `?table_filter=` + does not clear a filter set by the TOML file. +- A value outside ASCII is carried correctly when percent-encoded as UTF-8, + which makes the query string the way to name such a measurement. + ### Environment Variables Supported -The following environment variables can be used: +Every parameter can come from an environment variable named +`INFLUXDB3_IMPORT_` in upper case — for example, +`INFLUXDB3_IMPORT_SOURCE_URL` sets `source_url` and +`INFLUXDB3_IMPORT_TARGET_BATCH_SIZE` sets `target_batch_size`. +`INFLUXDB3_IMPORT_CONFIG_FILE_PATH` names the TOML file when the trigger doesn't +carry a `config_file_path` argument. + +Five of the settings are read from a second name as well: - `IMPORT_SOURCE_URL` → `source_url` - `IMPORT_SOURCE_DATABASE` → `source_database` @@ -582,14 +722,18 @@ The following environment variables can be used: - `IMPORT_START_TIMESTAMP` → `start_timestamp` - `IMPORT_END_TIMESTAMP` → `end_timestamp` +When a setting is given under both spellings, `INFLUXDB3_IMPORT_*` wins. + ## Data Type Mismatch Handling -The plugin automatically handles data type mismatches that can occur in older InfluxDB versions where different nodes might have different field types for the same field name. +The plugin automatically handles data type mismatches that can occur in older InfluxDB versions where different nodes +might have different field types for the same field name. ### How It Works 1. **Schema Detection**: At import start, plugin queries source database for field types using `SHOW FIELD KEYS` -2. **Runtime Type Checking**: For each data point, plugin checks if the actual value type matches the expected field type +2. **Runtime Type Checking**: For each data point, plugin checks if the actual value type matches the expected field + type 3. **Automatic Field Creation**: If type mismatch is detected, plugin creates a new field with a type suffix ### Supported Type Suffixes @@ -633,7 +777,40 @@ Stores import configuration (credentials excluded for security). influxdb3 query --database mydb "SELECT * FROM import_config WHERE import_id = 'your-import-id'" ``` #### `import_state` -Tracks per-table import progress. + +Tracks per-table import progress. One row per query window, holding the table's +status, its running row count and the timestamp of the last row written, to the +nanosecond. That checkpoint is what a resume continues from, so an import +continues from its last written window whether it was paused or interrupted. +The row count, the checkpoint and the failures all carry across a resume, so the +latest row of a table always holds its totals. + +Every row also carries an `errors` column, holding the windows the table failed +to write, as JSON: + +```json +{ + "failed_windows": 240, + "errors": [ + { + "time_range": "2026-01-01 00:00:00+00:00 to 2026-01-01 02:00:00+00:00", + "error": "invalid column type for column 'room', expected iox::column_type::field::integer, got iox::column_type::tag" + } + ] +} +``` +A table that failed nothing records `{"failed_windows": 0, "errors": []}`. +`failed_windows` counts every window that failed, while `errors` holds only the +first 50 of them — a table that fails on every window repeats one reason, and +the whole list would outgrow a single point. Rows written while the import is +still running sample 3 of them instead of 50, to keep a row per window small; +`failed_windows` is exact on every row. + +`action=status` returns the same object for each table under +`table_details[].errors`, and counts the tables that failed anything as +`summary.tables_with_errors`. This is where to look when an import reports +`completed` with fewer rows than expected: the status answers long after the +entries in `system.processing_engine_logs` have been trimmed away. ```bash influxdb3 query --database mydb "SELECT * FROM import_state WHERE import_id = 'your-import-id' ORDER BY time DESC" @@ -648,11 +825,18 @@ influxdb3 query --database mydb "SELECT * FROM import_pause_state WHERE import_i #### `process_request(influxdb3_local, query_parameters, request_headers, request_body, args)` -HTTP request handler that routes to appropriate import actions based on the `action` query parameter. Extracts credentials from `request_headers` using `extract_credentials()` and passes them to action handlers. +HTTP request handler that routes to appropriate import actions based on the `action` query parameter. Each action +accepts only the query parameters it reads, so an unknown one is refused and named alongside the accepted keys. Reads +the credentials from the `Source-Token`, `Source-Username` and `Source-Password` headers and passes them to the action +handlers; a header that was not sent is absent from that dict. -#### `extract_credentials(request_headers)` +#### `load_import_settings(influxdb3_local, task_id, args, request_body, request_headers, query_settings)` -Extracts authentication credentials from HTTP headers. Returns a dict with keys `source_token`, `source_username`, `source_password` (values are `None` if header not present). +Assembles the import configuration from environment variables, trigger arguments, the TOML file, the request body, the +`X-Influxdb3-Import-*` headers and the query parameters, in that order of precedence, then validates it: required +parameters must be present, numeric and boolean parameters are coerced, and `influxdb_version` and `import_direction` +must be one of their supported values. The TOML file path comes from the trigger arguments or from +`INFLUXDB3_IMPORT_CONFIG_FILE_PATH`. Raises on the first problem, naming the parameter. #### `start_import(influxdb3_local, config, credentials, task_id)` @@ -676,16 +860,20 @@ Imports a single table: #### `resume_import(influxdb3_local, import_id, credentials, task_id)` Resumes an interrupted import: -1. Detects stale imports — if the import state is "running" but the last update is older than 5 minutes, treats it as crashed and allows resume +1. Detects stale imports — if the import state is "running" but the last update is older than 5 minutes, treats it as + crashed and allows resume 2. If no `import_state` records exist (crashed before processing any tables), restarts from the beginning 3. Loads saved import configuration 4. Identifies incomplete tables and their last checkpoint 5. Continues import from checkpoint positions -6. On any unhandled error, writes paused state so the import can be resumed again +6. Stops on the first table the user pauses or cancels, leaving the import resumable +7. On any unhandled error, writes paused state so the import can be resumed again #### `get_import_stats(influxdb3_local, import_id, task_id)` -Returns comprehensive statistics for a import including overall status, per-table progress, timing information, and configuration. +Returns comprehensive statistics for a import including overall status, per-table progress, timing information, and +configuration. Each entry of `table_details` carries the table's `errors` object, and `summary.tables_with_errors`counts +the tables that failed to write anything. #### `check_source_connection(body_data, session)` @@ -701,16 +889,17 @@ Tests connectivity to a URL and identifies if it's an InfluxDB instance (5-secon Lists databases from source InfluxDB instance: 1. Validates required parameters -2. For v1: Executes `SHOW DATABASES` query, filters out `_internal` -3. For v2: Queries `/api/v2/buckets` API, filters out system buckets (prefixed with `_`) -4. Returns sorted list of database names +2. For v1: Executes `SHOW DATABASES` through `/query`, filters out `_internal` +3. For v2: the same query, filtering out the system buckets, which InfluxDB 2 marks with a leading underscore +4. For v3: Queries `/api/v3/configure/database` +5. Returns sorted list of database names #### `get_source_tables_list(body_data, credentials, session)` Lists tables/measurements from a source database: 1. Validates required parameters including `source_database` -2. For v1: Executes `SHOW MEASUREMENTS` query -3. For v2: Executes Flux schema.measurements() query (requires `source_org`) +2. For v1 and v2: Executes `SHOW MEASUREMENTS` through `/query` +3. For v3: Executes `SHOW TABLES` through `/api/v3/query_sql` 4. Returns sorted list of table names ### Key algorithms @@ -720,7 +909,7 @@ Lists tables/measurements from a source database: The plugin samples data at different time intervals to determine optimal window size: ```python -# Test intervals: 1 second, 1 minute, 1 hour, 1 day +# Test intervals: 1 hour, 10 hours, 1 day, 5 days # Calculate rows per second from samples # Determine window size to achieve target_batch_size optimal_window = target_batch_size / avg_rows_per_second @@ -753,6 +942,18 @@ During import, the plugin saves checkpoints: ### Common issues +#### Issue: "Source query failed" and the import stops + +**Cause**: The source accepted the request and answered with a failed +statement — a point limit, a series limit, or a query it could not plan. +InfluxDB reports these with HTTP 200 and the reason inside the body, so the +plugin raises on them rather than reading an empty answer as "no data". + +**Solution**: the reason is in the message and in +`system.processing_engine_logs`. The table is left `paused` with its checkpoint, +so `action=resume` continues it once the source is able to answer. Lowering +`target_batch_size` helps when the source refuses a window for its size. + #### Issue: "Failed to connect to source database" error **Solution**: @@ -773,6 +974,21 @@ During import, the plugin saves checkpoints: 3. Verify credentials work directly against source InfluxDB 4. Check that headers are not being stripped by proxies +#### Issue: `pause`, `status` or `cancel` does not answer while an import runs + +**Solution**: The trigger was created without `--run-asynchronous`, so the +Processing Engine serves one invocation at a time and the control request waits +for the import to finish. Recreate the trigger with the flag: + +```bash +influxdb3 delete trigger --database mydb --force import_trigger +influxdb3 create trigger \ + --database mydb \ + --plugin-filename gh:influxdata/import/import.py \ + --trigger-spec "request:import" \ + --run-asynchronous \ + import_trigger +``` #### Issue: "Import is already running" after server crash **Solution**: @@ -781,10 +997,10 @@ During import, the plugin saves checkpoints: 2. Call the resume endpoint with credentials: ```bash curl -X POST "http://localhost:8181/api/v3/engine/import?action=resume&import_id=$IMPORT_ID" \ - -H "Content-Type: application/json" \ - -d '{"source_token": "my-token"}' + -H "Source-Token: my-token" ``` -3. The plugin detects the stale state and resumes from the last checkpoint (or restarts from the beginning if no checkpoint exists) +3. The plugin detects the stale state and resumes from the last checkpoint (or restarts from the beginning if no + checkpoint exists) #### Issue: "Import already completed" when trying to resume @@ -812,12 +1028,12 @@ During import, the plugin saves checkpoints: 3. Use table filtering to import tables in parallel using multiple triggers 4. Check network latency between source and destination - ### Performance considerations - **Network bandwidth**: Main bottleneck for large imports. Use local network when possible. - **Source database load**: The plugin includes rate limiting (`query_interval_ms`) to avoid overwhelming source. -- **Batch size optimization**: Plugin automatically samples data to determine optimal batch size, but you can override with `target_batch_size`. +- **Batch size optimization**: Plugin automatically samples data to determine optimal batch size, but you can override + with `target_batch_size`. - **Connection pooling**: Plugin uses HTTP session with connection pooling for better performance. - **Retry logic**: Built-in exponential backoff (1s → 2s → 4s → 8s → 16s) for transient errors. diff --git a/content/shared/influxdb3-plugins/plugins-library/official/prophet-forecasting.md b/content/shared/influxdb3-plugins/plugins-library/official/prophet-forecasting.md index b9b4f987b9..c9544917df 100644 --- a/content/shared/influxdb3-plugins/plugins-library/official/prophet-forecasting.md +++ b/content/shared/influxdb3-plugins/plugins-library/official/prophet-forecasting.md @@ -35,7 +35,7 @@ Set these parameters with `--trigger-arguments` when creating a scheduled trigge ### HTTP request parameters -Send these parameters as JSON in the HTTP POST request body. Trigger arguments are not used by the HTTP endpoint; a JSON `null` means "not set", so the default applies. +Send these parameters as JSON in the HTTP POST request body, which must be a JSON object of at most 10 MB. A request carries three layers of its own — the body, then the headers, then the query string, each overriding the one before — and all three override what the trigger holds: the [environment](#environment-variables), the trigger arguments, then the [TOML file](#toml-configuration) they name. A trigger created without arguments is configured by the request alone; one created with them holds the defaults each request overrides only where it names them. A value that arrives empty — a JSON `null` or a blank string — counts as not set, so the trigger's value or the default applies. A key outside the tables below is refused and named in the response, as is `config_file_path` whatever its value, so a misspelling is reported rather than silently dropped. | Parameter | Type | Default | Description | |----------------------|---------------|----------|----------------------------------------------------------------------------------------------------| @@ -49,6 +49,19 @@ Send these parameters as JSON in the HTTP POST request body. Trigger arguments a | `end_time` | string | required | Historical window end, ISO 8601 with timezone. Forecast points are written from this moment onward | | `save_mode` | boolean | false | When true, load the saved model for `unique_suffix`, or train and save it when no file exists | +#### Headers and query parameters + +The same names reach the plugin as headers, spelled `X-Influxdb3-Prophet-Forecasting-` with underscores written as hyphens (`X-Influxdb3-Prophet-Forecasting-Start-Time` sets `start_time`), and as query-string parameters, spelled exactly like the parameter (`?save_mode=true`). Header names are matched regardless of casing, which RFC 9110 makes meaningless. + +```bash +curl -X POST "http://localhost:8181/api/v3/engine/forecast?save_mode=true" \ + -H "Authorization: Bearer $INFLUXDB3_AUTH_TOKEN" \ + -H "Content-Type: application/json" \ + -H "X-Influxdb3-Prophet-Forecasting-Unique-Suffix: v2" \ + -d '{"start_time": "2026-08-01T00:00:00Z", "end_time": "2026-08-29T00:00:00Z"}' +``` +A header the plugin does not ask for is ignored, since a client sends headers of its own on every request; an unknown query parameter is refused and named in the response, like an unknown body field. Neither layer can name `config_file_path`, for the reason the body cannot. + ### Advanced parameters Available to both trigger types: @@ -83,15 +96,19 @@ Scheduled triggers only: Each channel listed in `senders` needs its own keys (`slack_webhook_url`, `discord_webhook_url`, `http_webhook_url`, `twilio_sid`, `twilio_token`, `twilio_from_number`, `twilio_to_number`, and the optional `*_headers`). See the [influxdata/notifier plugin](/influxdb3/version/plugins/library/official/notifier/). +### Environment variables + +Every parameter can also come from an environment variable named `INFLUXDB3_PROPHET_FORECASTING_` in upper case — for example, `INFLUXDB3_PROPHET_FORECASTING_MEASUREMENT` sets `measurement`, and `INFLUXDB3_PROPHET_FORECASTING_TWILIO_TOKEN` keeps a channel credential out of the trigger. `influxdb3_auth_token` keeps its own variable, `INFLUXDB3_AUTH_TOKEN`. The environment is the lowest layer: a trigger argument overrides it, the TOML file overrides both, and on the HTTP trigger the request — its body, its headers and its query string — overrides them all. `INFLUXDB3_PROPHET_FORECASTING_CONFIG_FILE_PATH` names the TOML file when the trigger doesn't carry a `config_file_path` argument. + ### TOML configuration | Parameter | Type | Default | Description | |--------------------|--------|---------|----------------------------------------------------------------------------------| -| `config_file_path` | string | none | TOML config file path relative to `PLUGIN_DIR` (required for TOML configuration) | +| `config_file_path` | string | none | `.toml` file, relative to `PLUGIN_DIR` or absolute. A trigger argument on either trigger | *To use a TOML configuration file, set the `PLUGIN_DIR` environment variable and specify the `config_file_path` in the trigger arguments.* This is in addition to the `--plugin-dir` flag when starting {{% product-name %}}. Relative paths are resolved against the first directory that is set: `PLUGIN_DIR`, then `INFLUXDB3_PLUGIN_DIR`, then the parent of `VIRTUAL_ENV`. Only that directory is used — the file is not looked up in the remaining ones. -When `config_file_path` is set, the TOML file provides the whole configuration and inline trigger arguments are ignored. `INFLUXDB3_AUTH_TOKEN` from the environment still applies when `influxdb3_auth_token` is not set in the file. In TOML, `tag_values`, `senders`, `changepoints`, `holiday_date_list`, `holiday_names` and `holiday_country_names` can use native structures (a table or a list) instead of the inline string formats, though the inline strings are also accepted. The HTTP endpoint ignores `config_file_path`. +The file's values override the other trigger arguments, and on the HTTP trigger the request overrides the file; an argument the file does not set still applies. `INFLUXDB3_AUTH_TOKEN` from the environment applies when `influxdb3_auth_token` is set neither as an argument nor in the file. In TOML, `tag_values`, `senders`, `changepoints`, `holiday_date_list`, `holiday_names` and `holiday_country_names` can use native structures (a table or a list) instead of the inline string formats, though the inline strings are also accepted. The path is never read from a request — not from its body, its headers or its query string. It names a layer rather than setting a value: a request that could choose which file the trigger reads would take the trigger's configuration out of the operator's hands, so the body and the query string refuse it and a header spelling it is dropped. #### Example TOML configuration @@ -103,7 +120,7 @@ For more information on using TOML configuration files, see the Using TOML Confi - **{{% product-name %}}**: with the Processing Engine enabled. - **Python packages**: - - `influxdata-plugin-utils>=0.3.0` (configuration loading, parsing, and writing) + - `influxdata-plugin-utils>=0.4.0` (configuration loading, parsing, and writing) - `pandas` (for data manipulation; 2.x and 3.x are both supported) - `requests` (for HTTP requests) - `prophet` (for time series forecasting) @@ -146,13 +163,23 @@ influxdb3 create trigger \ ``` ### HTTP trigger -Create a trigger for on-demand forecasting: +Create a trigger for on-demand forecasting. Without trigger arguments the request body carries the whole configuration; with them, they are the defaults the body overrides. See [TOML configuration](#toml-configuration) for naming a file on the trigger. + +```bash +influxdb3 create trigger \ + --database mydb \ + --path "gh:influxdata/prophet_forecasting/prophet_forecasting.py" \ + --trigger-spec "request:forecast" \ + prophet_forecast_http_trigger +``` +A trigger that fixes the source and destination leaves each request to name only its window and model version: ```bash influxdb3 create trigger \ --database mydb \ --path "gh:influxdata/prophet_forecasting/prophet_forecasting.py" \ --trigger-spec "request:forecast" \ + --trigger-arguments "measurement=temperature,field=value,forecast_horizont=7d,tag_values=region:us-west,target_measurement=temperature_forecast,target_database=mydb" \ prophet_forecast_http_trigger ``` ### Enable triggers @@ -324,7 +351,7 @@ Handles on-demand forecasts over an explicit window. The training window is `sta #### Issue: HTTP trigger issues -**Solution**: Verify the JSON request body matches the expected schema. Check authentication tokens and database permissions. Ensure `start_time` and `end_time` are valid ISO 8601 values with a timezone. +**Solution**: Verify the JSON request body matches the expected schema. The response message names the key or value that was refused. Check authentication tokens and database permissions. Ensure `start_time` and `end_time` are valid ISO 8601 values with a timezone. #### Issue: Forecast results are not in the expected database diff --git a/content/shared/influxdb3-plugins/plugins-library/official/schema-validator.md b/content/shared/influxdb3-plugins/plugins-library/official/schema-validator.md index 626dcf7337..1a9e48bf0f 100644 --- a/content/shared/influxdb3-plugins/plugins-library/official/schema-validator.md +++ b/content/shared/influxdb3-plugins/plugins-library/official/schema-validator.md @@ -2,86 +2,81 @@ > **Note:** This plugin requires {{% product-name %}}.8.2 or later. -An {{% product-name %}} Processing Engine plugin that validates incoming line protocol data against a user-defined JSON schema. Only data that conforms to the schema is written to a target database or table, enabling a clean data pipeline pattern. -## Use Case +The Schema Validator Plugin validates incoming line protocol against a user-defined JSON schema and forwards only conforming rows to a target database or table. It runs on every WAL flush of the tables named by its trigger specification: each row is checked against the measurement whitelist, the required tags and their allowed values, and the required fields, their types and their allowed values. Valid rows are stripped down to the schema-defined tags and fields and written to the target; rejected rows are optionally logged and recorded in a `_schema_rejections` measurement. -You have data coming into a "raw" database (for example, `raw_db`) from various sources. You want to ensure only properly-structured, validated data makes it into your "clean" database (for example, `clean_db`). This plugin sits on the WAL flush trigger and validates every incoming row against your schema definition before writing it to the target. +Typical pipelines: -**Common patterns:** -- `raw_db` -> validate -> `clean_db` (cross-database) -- `raw_table` -> validate -> `validated_table` (same database, different table) -- `source_table` -> validate -> `source_table_clean` (same database, with suffix) +- `raw_db` → validate → `clean_db` (cross-database) +- `raw_table` → validate → `validated_table` (same database, different table) +- `source_table` → validate → `source_table_clean` (same database, with a suffix) -This is a **single-file plugin** (`schema_validator.py`) and can be loaded from GitHub via `gh:` trigger paths or created in InfluxDB Explorer. +This is a single-file plugin (`schema_validator.py`) and can be loaded from GitHub via `gh:` trigger paths or created in InfluxDB 3 Explorer. -> **Note:** The JSON schema configuration file (`schema_validator_config.json`) must be manually uploaded to the plugin directory on the server. There is currently no API for uploading non-plugin files, so Explorer cannot upload it for you. You can use `scp`, `rsync`, or any other file transfer method to place the schema file in the plugin directory alongside the plugin. +> **Note:** The JSON schema file must be placed in the plugin directory on the server by hand — there is no API for uploading non-plugin files, so Explorer cannot upload it for you. Use `scp`, `rsync`, or any other file transfer method. ## Features -- **Measurement validation**: Define a whitelist of allowed measurement/table names -- **Tag validation**: Required/optional tags, allowed tag values -- **Field validation**: Required/optional fields, type checking (float, integer, string, boolean, uint64), allowed field values -- **Field stripping**: Extra tags/fields not defined in the schema are automatically stripped from the output -- **Flexible targeting**: Write to a different database, different table name, or add prefix/suffix -- **Per-table schemas**: Define different validation rules for each measurement -- **Rejection logging**: Optionally log rejected rows and/or write rejection details to a measurement -- **Cached config**: Schema file is cached for 5 minutes to avoid repeated file reads +- **Measurement whitelist**: `allowed_measurements` names the tables that are processed; a table outside the list is skipped +- **Tag validation**: each tag is required or optional and may carry a list of allowed values +- **Field validation**: each field is required or optional, is checked against its declared or inferred type, and may carry a list of allowed values +- **Field stripping**: tags and fields the schema does not define are dropped from the written row +- **Flexible targeting**: write to another database, to the name a table's `target_table` gives, or to the source name with a prefix or suffix +- **Per-table schemas**: every measurement carries its own tags, fields and target +- **Rejection logging**: a rejected row is logged with its reason and, optionally, recorded in the `_schema_rejections` measurement +- **Cached schema**: the JSON file is re-read at most once every five minutes -## Quick Start +## Configuration -### 1. Deploy the plugin files +Plugin parameters may be specified as key-value pairs in the `--trigger-arguments` flag (CLI) or in the `trigger_arguments` field (API) when creating a trigger. Some plugins support TOML configuration files, which can be specified using the plugin's `config_file_path` parameter. -The plugin code can be deployed via the InfluxDB CLI, Explorer, or GitHub (`gh:`) trigger paths. However, the **schema JSON configuration file must be manually placed** in the plugin directory on the server since there is no API for uploading non-plugin files. +### Plugin metadata -- `schema_validator.py` - the plugin code (can be uploaded via CLI/Explorer/GitHub) -- `schema_validator_config.json` - your schema definition (must be manually copied to the plugin directory) +This plugin includes a JSON metadata schema in its docstring that defines supported trigger types and configuration parameters. This metadata enables the [InfluxDB 3 Explorer](https://docs.influxdata.com/influxdb3/explorer/) UI to display and configure the plugin. -### 2. Create a trigger +### Required parameters -**Cross-database validation (raw_db -> clean_db):** -```bash -influxdb3 create trigger \ - --database raw_db \ - --plugin-filename schema_validator.py \ - --trigger-spec "all_tables" \ - --trigger-arguments schema_file=schema_validator_config.json,target_database=clean_db \ - schema_validator_trigger -``` -**Same database, different table (with suffix):** -```bash -influxdb3 create trigger \ - --database mydb \ - --plugin-filename schema_validator.py \ - --trigger-spec "table:weather" \ - --trigger-arguments schema_file=schema_validator_config.json,target_table_suffix=_clean \ - schema_validator_weather -``` -**Using a TOML config file:** -```bash -influxdb3 create trigger \ - --database raw_db \ - --plugin-filename schema_validator.py \ - --trigger-spec "all_tables" \ - --trigger-arguments config_file_path=schema_validator_trigger_config.toml \ - schema_validator_trigger -``` -### 3. Write data normally +| Parameter | Type | Default | Description | +|---------------|--------|----------|-----------------------------------------------------------------------------| +| `schema_file` | string | required | Path to the JSON schema file, absolute or relative to the plugin directory. Must end in `.json` | -Write to your raw database as usual. The plugin will automatically validate and forward conforming data. +### Data write trigger parameters -```bash -# This row has all required fields -> will be written to clean_db -influxdb3 write --database raw_db \ - "weather,location=us-east,station_id=ST001 temperature=72.5,humidity=45.2" +| Parameter | Type | Default | Description | +|-----------------------|---------|------------------|----------------------------------------------------------------------------------------------------------------------------------------------------| +| `target_database` | string | trigger's own DB | Database that validated rows and the rejection log are written to | +| `target_table_prefix` | string | `""` | Prefix added to the source measurement name in the target. Ignored for tables that define `target_table` | +| `target_table_suffix` | string | `""` | Suffix added to the source measurement name in the target. Ignored for tables that define `target_table` | +| `log_rejected` | boolean | `true` | Log one warning per rejected row, plus one message per skipped table. Accepts `true`/`false`, `1`/`0`, `yes`/`no`, `on`/`off` | +| `log_accepted` | boolean | `false` | Log one message per accepted row — noisy, use for debugging. Accepts `true`/`false`, `1`/`0`, `yes`/`no`, `on`/`off` | +| `write_rejection_log` | boolean | `false` | Write rejected row details to the `_schema_rejections` measurement in the target database. Accepts `true`/`false`, `1`/`0`, `yes`/`no`, `on`/`off` | -# This row is missing required tag 'station_id' -> will be rejected -influxdb3 write --database raw_db \ - "weather,location=us-east temperature=72.5,humidity=45.2" -``` -## Schema Configuration (JSON) +### Environment variables + +Every parameter can also come from an environment variable named +`INFLUXDB3_SCHEMA_VALIDATOR_` in upper case — for example, +`INFLUXDB3_SCHEMA_VALIDATOR_SCHEMA_FILE` sets `schema_file`. The environment is +the lowest layer: a trigger argument overrides it, and the TOML file overrides +both. `INFLUXDB3_SCHEMA_VALIDATOR_CONFIG_FILE_PATH` names the TOML file when the +trigger doesn't carry a `config_file_path` argument. + +### TOML configuration + +| Parameter | Type | Default | Description | +|--------------------|--------|---------|----------------------------------------------------------------------------------| +| `config_file_path` | string | none | Path to a TOML config file, absolute or relative to the plugin directory | + +To use a TOML configuration file, name it in `config_file_path` in the trigger arguments. A relative path — `config_file_path` and `schema_file` alike — is resolved against the plugin directory: `PLUGIN_DIR` when it is set, otherwise `INFLUXDB3_PLUGIN_DIR`, which the processing engine sets from `--plugin-dir`, otherwise the parent of `VIRTUAL_ENV`. An absolute path is used as written. + +The file accepts the same keys as inline arguments, and its values override them. A key outside the tables above is refused and named in the error, in the trigger arguments and in the TOML file alike, so a misspelling is reported rather than silently dropped. -The schema is defined in a JSON file. Here is the full structure: +#### Example TOML configuration + +- [schema_validator_trigger_config.toml](https://github.com/influxdata/influxdb3_plugins/blob/master/influxdata/schema_validator/schema_validator_trigger_config.toml) + +For more information on using TOML configuration files, see the Using TOML Configuration Files section in the [influxdb3_plugins/README.md](https://github.com/influxdata/influxdb3_plugins/blob/master/README.md). + +## Schema configuration (JSON) ```json { @@ -95,22 +90,12 @@ The schema is defined in a JSON file. Here is the full structure: "required": true, "allowed_values": ["us-east", "us-west", "eu-west"] }, - "station_id": { - "required": true - }, - "region": { - "required": false - } + "station_id": { "required": true }, + "region": { "required": false } }, "fields": { - "temperature": { - "required": true, - "type": "float" - }, - "humidity": { - "required": true, - "type": "float" - }, + "temperature": { "required": true, "type": "float" }, + "humidity": { "required": true, "type": "float" }, "condition": { "required": false, "type": "string", @@ -121,151 +106,231 @@ The schema is defined in a JSON file. Here is the full structure: } } ``` -### Schema Fields Reference +### Top-level -#### Top-level +| Field | Type | Description | +|------------------------|------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------| +| `allowed_measurements` | `list[str]` (optional) | Whitelist of measurement names. When omitted or empty, no measurement filter is applied and all measurements fall through to the `tables` rules | +| `tables` | `dict` (required) | Map of measurement name -> table definition. Must contain at least one entry; measurements without an entry are skipped | -| Field | Type | Description | -|---|---|---| -| `allowed_measurements` | `list[str]` (optional) | Whitelist of allowed measurement/table names. If omitted or empty (`[]`), no measurement filter is applied — all measurements fall through to the `tables` rules. | -| `tables` | `dict` (required) | Map of measurement name -> table schema definition. Must contain at least one entry. Only measurements with an entry here will be validated and written to the target. | +### Table definition -#### Table Schema +| Field | Type | Description | +|----------------|-------------------|----------------------------------------------------------------------------------------------| +| `target_table` | `str` (optional) | Target measurement name for this table. Takes precedence over `target_table_prefix`/`suffix` | +| `tags` | `dict` (optional) | Map of tag name -> tag definition | +| `fields` | `dict` (required) | Map of field name -> field definition. Must contain at least one entry | -| Field | Type | Description | -|---|---|---| -| `target_table` | `str` (optional) | Override the target measurement name. Takes precedence over prefix/suffix args. | -| `tags` | `dict` | Map of tag name -> tag definition. | -| `fields` | `dict` | Map of field name -> field definition. | -#### Tag Definition +### Tag definition -| Field | Type | Description | -|---|---|---| -| `required` | `bool` | If `true`, the tag must be present on every row. | -| `allowed_values` | `list` (optional) | Whitelist of allowed values for this tag. | +| Field | Type | Description | +|------------------|-------------------|-----------------------------------------------| +| `required` | `bool` | When `true`, the tag must be present on a row | +| `allowed_values` | `list` (optional) | Whitelist of values, compared as strings | -#### Field Definition +### Field definition -| Field | Type | Description | -|---|---|---| -| `required` | `bool` | If `true`, the field must be present on every row. | -| `type` | `str` (optional) | Expected data type: `"float"`, `"integer"`, `"string"`, `"boolean"`, `"uint64"`. | -| `allowed_values` | `list` (optional) | Whitelist of allowed values for this field. | +| Field | Type | Description | +|------------------|-------------------|------------------------------------------------------------------| +| `required` | `bool` | When `true`, the field must be present on a row | +| `type` | `str` (optional) | Expected type. When omitted, the type is inferred from the value | +| `allowed_values` | `list` (optional) | Whitelist of values, compared by value and as strings | -## Trigger Arguments +A tag or field written as a bare name instead of a definition object is treated as required with no value whitelist. -These can be passed via `--trigger-arguments` or in a TOML config file. +### Field types -| Argument | Required | Default | Description | -|---|---|---|---| -| `schema_file` | Yes | - | Path to the JSON schema file (relative to PLUGIN_DIR). | -| `target_database` | No | (same db) | Database to write validated data to. | -| `target_table_prefix` | No | `""` | Prefix added to measurement names in the target. | -| `target_table_suffix` | No | `""` | Suffix added to measurement names in the target. | -| `log_rejected` | No | `"true"` | Log info about rejected rows. | -| `log_accepted` | No | `"false"` | Log info about accepted rows. | -| `write_rejection_log` | No | `"false"` | Write rejection details to `_schema_rejections` measurement. | -| `config_file_path` | No | - | Path to TOML config file to override these arguments. | +| `type` | Accepted values | +|------------------------------|--------------------------------------------------------------------| +| `float`, `float64`, `double` | Integers and floats, excluding booleans; must be finite | +| `integer`, `int`, `int64` | Integers, excluding booleans; must fit into `int64` | +| `uint64`, `unsigned`, `uint` | Non-negative integers, excluding booleans; must fit into `uint64` | +| `string`, `str` | Strings | +| `boolean`, `bool` | Booleans | -## Validation Logic +An unknown type name in the schema is reported as an error and the plugin does not run. When a field definition has no `type`, the type is inferred from the value: booleans become `boolean`, integers `integer`, floats `float`, everything else `string`. -For each incoming row, the plugin checks (in order): +## Validation logic -1. **Measurement name**: Is the table name in `allowed_measurements`? (if defined) -2. **Table schema exists**: Is there a schema definition for this table in `tables`? If not, the table is skipped. -3. **Required tags**: Are all required tags present? -4. **Tag values**: Are tag values in the `allowed_values` list? (if defined) -5. **Required fields**: Are all required fields present? -6. **Field types**: Do field values match the expected type? (if defined) -7. **Field values**: Are field values in the `allowed_values` list? (if defined) -8. **Field stripping**: Any extra tags/fields not defined in the schema are stripped from the output. +For each row, in order: -If **any** check fails, the row is rejected and not written to the target. +1. **Required tags** — every tag marked `required` must be present. +2. **Tag values** — a tag with `allowed_values` must carry one of them. +3. **Required fields** — every field marked `required` must be present. +4. **Field types** — a present field must match its declared or inferred type. +5. **Field values** — a field with `allowed_values` must carry one of them. -## Examples +If any check fails, the row is rejected and nothing is written for it. A row is also rejected when it carries no schema-defined field, or when a value cannot be written as its type (a non-finite float, an integer outside the `int64`/`uint64` range). -### IoT Sensor Validation +Tags and fields absent from the schema are stripped from the output. Tables absent from a non-empty `allowed_measurements` list are skipped, as are tables with no `tables` entry. -Ensure sensor readings always have a device_id, valid sensor type, and a numeric value: +## Target resolution -```json -{ - "allowed_measurements": ["sensor_readings"], - "tables": { - "sensor_readings": { - "target_table": "sensors_validated", - "tags": { - "device_id": { "required": true }, - "sensor_type": { - "required": true, - "allowed_values": ["temperature", "pressure", "humidity"] - } - }, - "fields": { - "value": { "required": true, "type": "float" }, - "status": { - "required": false, - "type": "string", - "allowed_values": ["ok", "warning", "critical"] - } - } - } - } -} +Validated rows of a table are written to: + +1. the table's `target_table`, when set; otherwise +2. ``. + +When `target_database` is not set, rows go back into the trigger's own database, so a table must resolve to a measurement name other than its own — otherwise the trigger would feed itself. A flush carrying such a table is rejected, naming the offending table; tables that the trigger never receives are not checked. Set `target_database`, a prefix, a suffix, or a per-table `target_table`. + +The schema file is cached for 5 minutes; edits are picked up within that window. + +## Rejection log + +With `write_rejection_log` enabled, every rejected row adds one point to `_schema_rejections` in the target database: + +| Column | Kind | Description | +|----------------|-----------|------------------------------------------------------------| +| `source_table` | tag | Table the rejected row arrived on | +| `reason` | field | Rejection reason | +| `row_data` | field | The row rendered as a string, truncated to 1024 characters | +| `time` | timestamp | Time of validation | + +Entries are batched per table and written in one call. + +## Software Requirements + +- **{{% product-name %}}**: with the Processing Engine enabled +- **Python packages**: `influxdata-plugin-utils>=0.4.0` + +## Installation steps + +1. Start {{% product-name %}} with the Processing Engine enabled (`--plugin-dir /path/to/plugins`): + + ```bash + influxdb3 serve \ + --node-id node0 \ + --object-store file \ + --data-dir ~/.influxdb3 \ + --plugin-dir ~/.plugins + ``` +2. Install required Python packages: + + ```bash + influxdb3 install package "influxdata-plugin-utils>=0.4.0" + ``` +3. Copy the JSON schema file into the plugin directory: + + ```bash + scp schema_validator_config.json user@server:~/.plugins/ + ``` +## Trigger setup + +Cross-database validation (`raw_db` -> `clean_db`): + +```bash +influxdb3 create trigger \ + --database raw_db \ + --path "gh:influxdata/schema_validator/schema_validator.py" \ + --trigger-spec "all_tables" \ + --trigger-arguments "schema_file=schema_validator_config.json,target_database=clean_db" \ + --error-behavior log \ + schema_validator_trigger ``` -### Multi-table with Cross-database +Same database, different table: -Validate weather and cpu data from raw_db, writing clean data to clean_db: +```bash +influxdb3 create trigger \ + --database mydb \ + --path "gh:influxdata/schema_validator/schema_validator.py" \ + --trigger-spec "table:weather" \ + --trigger-arguments "schema_file=schema_validator_config.json,target_table_suffix=_clean" \ + --error-behavior log \ + schema_validator_weather +``` +Using a TOML config file: ```bash influxdb3 create trigger \ --database raw_db \ - --plugin-filename schema_validator.py \ + --path "gh:influxdata/schema_validator/schema_validator.py" \ --trigger-spec "all_tables" \ - --trigger-arguments schema_file=schema_validator_config.json,target_database=clean_db,write_rejection_log=true \ - schema_validator_all + --trigger-arguments "config_file_path=schema_validator_trigger_config.toml" \ + --error-behavior log \ + schema_validator_trigger ``` -### Monitoring Rejections +### Enable triggers -Query the rejection log to see what data is being rejected and why: +```bash +influxdb3 enable trigger --database raw_db schema_validator_trigger +``` +## Example usage + +### Example 1: Cross-database validation +```bash +# This row has all required tags and fields -> written to clean_db +influxdb3 write --database raw_db \ + "weather,location=us-east,station_id=ST001 temperature=72.5,humidity=45.2" + +# This row is missing the required tag 'station_id' -> rejected +influxdb3 write --database raw_db \ + "weather,location=us-east temperature=72.5,humidity=45.2" + +# Query the validated data +influxdb3 query --database clean_db "SELECT * FROM weather_clean ORDER BY time DESC LIMIT 10" +``` +### Example 2: Monitoring rejections + +```bash +influxdb3 create trigger \ + --database raw_db \ + --path "gh:influxdata/schema_validator/schema_validator.py" \ + --trigger-spec "all_tables" \ + --trigger-arguments "schema_file=schema_validator_config.json,target_database=clean_db,write_rejection_log=true" \ + --error-behavior log \ + schema_validator_all +``` ```sql -SELECT * FROM _schema_rejections +SELECT source_table, reason, row_data +FROM _schema_rejections WHERE time > now() - INTERVAL '1 hour' ORDER BY time DESC ``` -## File Structure +## Code overview -``` -schema_validator/ - schema_validator.py # Main plugin code - schema_validator_config.json # Example schema definition - schema_validator_trigger_config.toml # Example TOML trigger config - README.md # This file -``` -## Notes +### Files -- The schema JSON file is cached for 5 minutes. To force a reload, restart the trigger or wait for the cache to expire. -- Extra tags/fields not defined in the schema are silently stripped from the output (not written to the target). -- Measurements without an entry in `tables` are skipped entirely (no data written). -- The `target_table` property in a table schema takes precedence over `target_table_prefix`/`target_table_suffix`. -- The `_schema_rejections` table (if `write_rejection_log=true`) is written to the target database. -- Uses `write_sync` / `write_sync_to_db` with `no_sync=True` for optimal memory performance. -- Valid rows are batched per table and written in a single call for efficiency. +- `schema_validator.py`: The main plugin code containing `process_writes` +- `schema_validator_config.json`: Example JSON schema definition +- `schema_validator_trigger_config.toml`: Example TOML trigger configuration +- `test_schema_validator.py`: Pytest suite (55 tests, runs without a live {{% product-name %}} server) +- `requirements.txt`: Runtime dependencies (`influxdata-plugin-utils>=0.4.0`) -## Logging +### Logging -Logs are stored in the `_internal` database (or the database where the trigger is created) in the `system.processing_engine_logs` table. To view logs: +Logs are stored in the trigger's database in the `system.processing_engine_logs` table. To view logs: ```bash -influxdb3 query --database _internal "SELECT * FROM system.processing_engine_logs WHERE trigger_name = 'your_trigger_name'" +influxdb3 query --database YOUR_DATABASE "SELECT * FROM system.processing_engine_logs WHERE trigger_name = 'your_trigger_name'" ``` +Every log line is prefixed with a per-fire `task_id` (eight hex characters) so records from a single trigger fire can be correlated. + +### Main functions + +#### `process_writes(influxdb3_local, table_batches, args)` + +Loads the configuration and the cached schema, then processes each table batch independently: rows are validated, valid ones are collected as line protocol and written with `write_sync` (or `write_sync_to_db` when `target_database` is set), and rejection log entries are written in a second batch. Writes are not retried, so a WAL flush is never held by backoff; a failing write is reported with the number of rows that did not land, those rows are counted as dropped in the closing summary instead of accepted, and the remaining tables are still processed. + +## Troubleshooting + +### Common issues + +#### Issue: No data appears in the target + +**Solution**: Only measurements with an entry in `tables` are forwarded, and a non-empty `allowed_measurements` list filters them further. Check `system.processing_engine_logs` for `No schema defined for table` and `not in allowed_measurements` messages. + +#### Issue: Trigger reports "would be written back into itself" + +**Solution**: With no `target_database`, the target measurement name must differ from the source. Set `target_database`, `target_table_prefix`, `target_table_suffix`, or the table's `target_table`. + +#### Issue: Every row of a table is rejected + +**Solution**: Read the rejection reason in the logs or in `_schema_rejections`. Common causes are a `required` tag that arrives as a field in the source data (or the reverse), an `allowed_values` list that does not cover production values, and a `type` that does not match what is written (for example, `integer` against a value written as `72.5`). + +#### Issue: Schema edits do not take effect -Log columns: -- **event_time**: Timestamp of the log event -- **trigger_name**: Name of the trigger that generated the log -- **log_level**: Severity level (INFO, WARN, ERROR) -- **log_text**: Message describing the action or error +**Solution**: The schema file is cached for 5 minutes. Wait for the cache to expire or restart the trigger. ## Report an issue diff --git a/content/shared/influxdb3-plugins/plugins-library/official/signal-filter.md b/content/shared/influxdb3-plugins/plugins-library/official/signal-filter.md index 38946b7905..ac9186a228 100644 --- a/content/shared/influxdb3-plugins/plugins-library/official/signal-filter.md +++ b/content/shared/influxdb3-plugins/plugins-library/official/signal-filter.md @@ -25,8 +25,8 @@ plugin but works with any measurement carrying numeric fields. Plugin parameters may be specified as key-value pairs in the `--trigger-arguments` flag (`influxdb3 create trigger`) or in the `trigger_arguments` field of the API. -Values are strings; the plugin coerces them. Alternatively, supply every parameter -from a TOML file via `config_file_path` — see [TOML configuration](#toml-configuration). +Values are strings; the plugin coerces them. Parameters may also come from a TOML +file via `config_file_path` — see [TOML configuration](#toml-configuration). > **CLI limitation:** the `sos` argument is a JSON array containing commas, and > `influxdb3 create trigger --trigger-arguments` splits on every comma, so the value @@ -69,27 +69,37 @@ and configure the plugin. ### Output parameters -| Parameter | Type | Default | Description | -|---|---|---|---| -| `output_target_database` | string | *(trigger db)* | Database to write filtered output to. | -| `output_measurement` | string | *(source table)* | Measurement to write filtered output to. | -| `output_field` | string | *(source field)* | Base name override for the output field. Only valid when a single input field is configured. | -| `field_prefix` | string | *(empty)* | Prefix for the output field name. | -| `field_suffix` | string | `_filtered` | Suffix for the output field name. | -| `config_file_path` | string | — | Path to a TOML file supplying all parameters; mutually exclusive with inline arguments. Relative paths resolve against `PLUGIN_DIR`. | +| Parameter | Type | Default | Description | +|--------------------------|--------|------------------|--------------------------------------------------------------------------------------------------------------------------------------| +| `output_target_database` | string | *(trigger db)* | Database to write filtered output to. | +| `output_measurement` | string | *(source table)* | Measurement to write filtered output to. | +| `output_field` | string | *(source field)* | Base name override for the output field. Only valid when a single input field is configured. | +| `field_prefix` | string | *(empty)* | Prefix for the output field name. `none` means no prefix; an empty value counts as unset. | +| `field_suffix` | string | `_filtered` | Suffix for the output field name. `none` writes into the source field, replacing its samples; an empty value counts as unset. | +| `config_file_path` | string | — | Path to a TOML file supplying parameters; its values override inline arguments. Relative paths resolve against `PLUGIN_DIR`. | The final output field name is `{field_prefix}{output_field or source_field}{field_suffix}` — by default, `value` becomes `value_filtered`. The raw input field is never copied to the output. +### Environment variables + +Every parameter can also come from an environment variable named +`INFLUXDB3_SIGNAL_FILTER_` in upper case — for example, +`INFLUXDB3_SIGNAL_FILTER_FC` sets `fc`. The environment is the lowest layer: a +trigger argument overrides it, and the TOML file overrides both. +`INFLUXDB3_SIGNAL_FILTER_CONFIG_FILE_PATH` names the TOML file when the trigger +doesn't carry a `config_file_path` argument. + ### TOML configuration To use a TOML configuration file, set the `PLUGIN_DIR` environment variable and reference the file with the `config_file_path` trigger argument (relative paths resolve against `PLUGIN_DIR`, then `INFLUXDB3_PLUGIN_DIR`, then the parent of -`VIRTUAL_ENV`). The TOML file then supplies **all** parameters — it is mutually -exclusive with inline trigger arguments, so passing both is rejected. See -[`signal_filter_config_data_writes.toml`](signal_filter_config_data_writes.toml) +`VIRTUAL_ENV`). The file and the inline trigger arguments are layered, and the +file wins wherever both set a key, so a trigger can carry defaults that a file +overrides. A blank value counts as unset and leaves the layer below it standing. +See [`signal_filter_config_data_writes.toml`](signal_filter_config_data_writes.toml) for an annotated template. ## Data requirements @@ -289,12 +299,18 @@ Manual mode needs no sample rate — the coefficients are already digital. Writing the output into the source measurement re-fires this trigger. This is safe by default: re-fired rows carry only the output field, the input field is -null on them, and null values produce no samples. **However**, if your overrides -resolve the output field to the *same name* as the input field in the same -measurement and database (for example `field_suffix=""` with no `output_field`), -the output feeds the filter again and grows without bound. The plugin logs a -prominent warning in that configuration — change `field_suffix`, `output_field`, -`output_measurement`, or `output_target_database` to break the cycle. +null on them, and null values produce no samples. + +If your overrides resolve the output field to the *same name* as the input +field in the same measurement and database — `field_suffix=none` with no +`field_prefix`, or `input_fields=value_filtered` with `output_field=value` — +the filtered values land on the source field at the same timestamps and +**replace the source samples**, which cannot be undone. The re-fire that +follows is dropped by the out-of-order guard (the samples are at or before the +last processed timestamp), so it costs one extra empty invocation rather than +running away. The plugin logs a prominent warning in that configuration; set +`output_field`, `field_suffix`, `output_measurement`, or +`output_target_database` to write elsewhere. ## Code overview @@ -329,7 +345,7 @@ filtered points, and finally saves the advanced per-series state. Key operations: 1. Guards that `numpy`, `scipy`, and `influxdata-plugin-utils` are installed; logs an install command otherwise -2. Parses and validates trigger arguments (inline, or entirely from a TOML file) +2. Parses and validates the configuration (trigger arguments layered under a TOML file) 3. Warns on any write-loop hazard configuration 4. Groups rows into per-(field, series) samples, dropping null/non-numeric/non-finite values 5. Resolves the sample rate (explicit → frozen → inferred with warm-up) diff --git a/content/shared/influxdb3-plugins/plugins-library/official/system-metrics.md b/content/shared/influxdb3-plugins/plugins-library/official/system-metrics.md index 820bb0786e..257cbb7de4 100644 --- a/content/shared/influxdb3-plugins/plugins-library/official/system-metrics.md +++ b/content/shared/influxdb3-plugins/plugins-library/official/system-metrics.md @@ -27,6 +27,10 @@ This plugin includes a JSON metadata schema in its docstring that defines suppor Boolean parameters accept `true`/`false`, `1`/`0`, `yes`/`no`, and `on`/`off`. A value the plugin cannot interpret is reported in the logs and the run collects nothing, so fix the trigger arguments and the next run recovers. +### Environment variables + +Every parameter can also come from an environment variable named `INFLUXDB3_SYSTEM_METRICS_` in upper case — for example, `INFLUXDB3_SYSTEM_METRICS_HOSTNAME` sets `hostname`. The environment is the lowest layer: a trigger argument overrides it, and the TOML file overrides both. `INFLUXDB3_SYSTEM_METRICS_CONFIG_FILE_PATH` names the TOML file when the trigger doesn't carry a `config_file_path` argument. + ### TOML configuration | Parameter | Type | Default | Description | @@ -35,7 +39,7 @@ Boolean parameters accept `true`/`false`, `1`/`0`, `yes`/`no`, and `on`/`off`. A *To use a TOML configuration file, set the `PLUGIN_DIR` environment variable and specify the `config_file_path` in the trigger arguments.* This is in addition to the `--plugin-dir` flag when starting {{% product-name %}}. Relative paths are resolved against the first directory that is set: `PLUGIN_DIR`, then `INFLUXDB3_PLUGIN_DIR`, then the parent of `VIRTUAL_ENV`. Only that directory is used — the file is not looked up in the remaining ones. -Values in the TOML file override the inline trigger arguments. If the file cannot be read, the plugin logs an error and collects metrics using the inline arguments and defaults. +Values in the TOML file override the inline trigger arguments. A file that cannot be read, is not named `.toml`, or holds an invalid value stops the run with a configuration error rather than falling back to the inline arguments. #### Example TOML configuration @@ -46,7 +50,7 @@ For more information on using TOML configuration files, see the Using TOML Confi ## Software Requirements - **{{% product-name %}}**: with the Processing Engine enabled. -- **Python packages**: `influxdata-plugin-utils>=0.3.0`, `psutil` +- **Python packages**: `influxdata-plugin-utils>=0.4.0`, `psutil` ### Installation steps @@ -62,7 +66,7 @@ For more information on using TOML configuration files, see the Using TOML Confi 2. Install required Python packages: ```bash - influxdb3 install package "influxdata-plugin-utils>=0.3.0" + influxdb3 install package "influxdata-plugin-utils>=0.4.0" influxdb3 install package psutil ``` ## Trigger setup @@ -162,7 +166,11 @@ The main entry point for scheduled triggers. Loads the configuration, then runs ```python def process_scheduled_call(influxdb3_local, call_time, args=None): - config = _load_config(influxdb3_local, args, task_id) + config: Config = load_config( + parse_trigger_args(args, SETTINGS), + parse_toml(args.get("config_file_path"), SETTINGS), + validators=SETTING_VALIDATORS, + ) for config_key, metric_type, collect in _COLLECTORS: if not config[config_key]: @@ -256,7 +264,7 @@ Network interface statistics: **Solution**: Install the required packages: ```bash -influxdb3 install package "influxdata-plugin-utils>=0.3.0" +influxdb3 install package "influxdata-plugin-utils>=0.4.0" influxdb3 install package psutil ``` #### Issue: No `system_disk_performance` data, or CPU shares are missing @@ -269,7 +277,7 @@ influxdb3 install package psutil #### Issue: No metrics at all and a configuration error in the logs -**Solution**: An invalid parameter value stops the run before any collection. Look for `Failed to load configuration` in the logs, which names the offending value, and fix the trigger arguments. A TOML file that cannot be read is a separate case: it is logged as `Failed to apply config file` and collection continues with the inline arguments. +**Solution**: An invalid setting stops the run before any collection. Look for `Configuration error` in the logs, which names the offending key and value, and fix the trigger arguments or the TOML file. An unreadable, wrongly named or invalid config file is reported the same way, so check the path in `config_file_path` resolves under `PLUGIN_DIR` and names a `.toml` file. #### Issue: High CPU usage from plugin diff --git a/data/influxdb3_plugins.yml b/data/influxdb3_plugins.yml index 838f5947ae..8481c3c744 100644 --- a/data/influxdb3_plugins.yml +++ b/data/influxdb3_plugins.yml @@ -80,7 +80,7 @@ description: Enables seamless data import from InfluxDB v1, v2, or v3 instances to InfluxDB 3 Core/Enterprise tags: - HTTP request - introduced: v0.2.0 + introduced: v0.3.0 database_version: '>=3.8.2' repository: https://github.com/influxdata/influxdb3_plugins/tree/main/influxdata/import - name: influxdb_to_iceberg @@ -155,7 +155,7 @@ tags: - scheduled - HTTP request - introduced: v1.5.0 + introduced: v1.6.0 database_version: '>=3.0.0' repository: https://github.com/influxdata/influxdb3_plugins/tree/main/influxdata/prophet_forecasting - name: resampler @@ -203,7 +203,7 @@ description: Validates incoming line protocol data against a JSON schema. tags: - data-write - introduced: v0.2.0 + introduced: v0.3.0 database_version: '>=3.8.2' repository: https://github.com/influxdata/influxdb3_plugins/tree/main/influxdata/schema_validator - name: signal_filter @@ -211,7 +211,7 @@ description: Applies streaming digital IIR filters to numeric fields tags: - data-write - introduced: v0.2.0 + introduced: v0.3.0 database_version: '>=3.8.2' repository: https://github.com/influxdata/influxdb3_plugins/tree/main/influxdata/signal_filter - name: signal_generator @@ -269,7 +269,7 @@ description: Collects system-level metrics (CPU, memory, disk, and network) using psutil and writes them to InfluxDB. Designed to run on a schedule and provide observability into host-level performance. tags: - scheduled - introduced: v1.2.0 + introduced: v1.3.0 database_version: '>=3.0.0' repository: https://github.com/influxdata/influxdb3_plugins/tree/main/influxdata/system_metrics - name: threshold_deadman_checks