From e6def7ccf2f4d8b64a6bc97b82b605397d29ddfb Mon Sep 17 00:00:00 2001 From: "antoine.choimet" <12182686+achoimet@users.noreply.github.com.> Date: Wed, 30 Sep 2026 09:34:14 +0200 Subject: [PATCH 1/2] test: the platform test covers what the CLI manages, not only runs Templates, run properties, schedules, services and service profiles with the team token; environments, teams, access tokens, hubs and integrations with an admin token from STEADYBIT_E2E_ADMIN_TOKEN, skipped without it. What an admin creates is named cli-e2e-ci-*, is scoped to a team of its own, targets nothing, and is swept afterwards, also after a crashed run. The kill switch is only read and no invitation is sent. Tokens only go through the environment, so a failing check cannot show one, and new ones are masked in the job log. It found that execution property set sent a single value as a scalar, which the platform refuses for a list property; a list property now gets a list. --- .github/workflows/platform-e2e.yml | 5 +- CHANGELOG.md | 8 + CONTRIBUTING.md | 15 +- e2e/platform.sh | 308 +++++++++++++++++++++++++++ internal/execution/execution.go | 19 +- internal/execution/execution_test.go | 13 +- 6 files changed, 358 insertions(+), 10 deletions(-) diff --git a/.github/workflows/platform-e2e.yml b/.github/workflows/platform-e2e.yml index acdd4e0..b6984c6 100644 --- a/.github/workflows/platform-e2e.yml +++ b/.github/workflows/platform-e2e.yml @@ -1,8 +1,8 @@ name: Platform E2E # Uses the CLI against the dev platform the way a pipeline does (e2e/platform.sh), with -# the team token of the test team CLI-E2E. Only wait steps, and everything it creates is -# deleted afterwards. +# the team token of the test team CLI-E2E, then the rest of what the CLI covers, partly +# with an admin token. Only wait steps, and everything it creates is deleted afterwards. on: # Weekly, Tuesday evening, so a change on the platform side shows within a week. Cron # is UTC: 19:00 in Berlin in summer, 18:00 in winter. @@ -34,6 +34,7 @@ jobs: env: STEADYBIT_URL: https://platform.dev.steadybit.com STEADYBIT_TOKEN: ${{ secrets.STEADYBIT_E2E_TOKEN }} + STEADYBIT_E2E_ADMIN_TOKEN: ${{ secrets.STEADYBIT_E2E_ADMIN_TOKEN }} STEADYBIT_E2E_TEAM: CLI STEADYBIT_E2E_ENVIRONMENT: Global steps: diff --git a/CHANGELOG.md b/CHANGELOG.md index ec6fc01..0de5543 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,13 @@ # Changelog +## v6.2.2 + +- `execution property set` with a single `--value` sets a list property, such as a + `STRING_LIST`, to a list of that one value; the platform refused it with "must be a + list of strings". Found by the platform test, which now covers templates, run + properties, schedules, services, environments, teams, access tokens, hubs and + integrations. + ## v6.2.1 - `--expect-state RUNNING` (or `PREPARED`, `CREATED`) passes for a run that ended after it diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index eff2592..7603303 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -69,13 +69,18 @@ docker run --rm -v "$PWD/e2e:/e2e" --entrypoint sh steadybit/cli:under-test /e2e The platform test uses the released CLI the way a pipeline does, against the dev platform with the team token of the test team CLI-E2E (key `CLI`): experiments applied from files, run in parallel with a report, checked for drift, canceled by SIGTERM, -refused while another runs, and found by an `execution list` gate. It only uses wait -steps and deletes what it creates. CI runs it weekly, after each stable release, and on -pull requests that change it; a nightly job only checks that the token and the API -still work. To run it by hand, with `steadybit` on the `PATH`: +refused while another runs, and found by an `execution list` gate. Then the rest of +the platform the CLI covers: templates, run properties, schedules, services and service +profiles with the team token, and environments, teams, access tokens, hubs, +integrations and the audit log with an admin token. It only uses wait steps, and deletes +what it creates, which is named `cli-e2e-ci-*`. It only reads the kill switch and sends +no invitation. CI runs it weekly, after each stable release, and on pull requests that +change it; a nightly job only checks that the token and the API still work. To run it by +hand, with `steadybit` on the `PATH` (without the admin token, those checks are skipped): ```sh -STEADYBIT_URL=https://platform.dev.steadybit.com STEADYBIT_TOKEN= e2e/platform.sh +STEADYBIT_URL=https://platform.dev.steadybit.com STEADYBIT_TOKEN= \ + STEADYBIT_E2E_ADMIN_TOKEN= e2e/platform.sh ``` ### Output compatibility diff --git a/e2e/platform.sh b/e2e/platform.sh index 19ab716..963e693 100755 --- a/e2e/platform.sh +++ b/e2e/platform.sh @@ -7,8 +7,16 @@ # canceled job sends, refused while another runs, and found again by a gate. Everything # only waits, belongs to one test team, and is deleted afterwards, also when a check fails. # +# Then the rest of the platform the CLI covers: templates, run properties, schedules, +# services and their profiles with the team token, and environments, teams, access +# tokens, hubs, integrations and the audit log with an admin token. What an admin creates +# is named cli-e2e-ci-*, belongs to a team of its own when it has a scope, and targets +# nothing. The kill switch is only read, and no invitation is sent: both would reach +# beyond the test. +# # Needs STEADYBIT_TOKEN (a team token) and STEADYBIT_URL, and `steadybit` on the PATH. # STEADYBIT_E2E_TEAM and STEADYBIT_E2E_ENVIRONMENT name the team and its environment. +# STEADYBIT_E2E_ADMIN_TOKEN, an admin token, enables the checks that need one. # # Other test suites run experiments on the same platform at the same time, so every run # here allows running in parallel, except the one whose refusal is the point of its check. @@ -22,6 +30,11 @@ ENVIRONMENT=${STEADYBIT_E2E_ENVIRONMENT:-Global} # finds and deletes what a crashed one left behind. MARK=cli-e2e-ci RUN=${GITHUB_RUN_ID:-local-$$} +ADMIN_TOKEN=${STEADYBIT_E2E_ADMIN_TOKEN:-} +# The team the admin checks create; one run at a time, so its key can stay the same. +SCOPE_TEAM=CLIX +# A property of runs, made by the admin checks for the experiment they create. +PROPERTY=cliE2eCi work=$(mktemp -d) cd "$work" || exit 1 @@ -69,6 +82,18 @@ EOF } key_of() { sed -n 's/^key: //p' "$1"; } +id_of() { sed -n 's/^id: //p' "$1"; } + +# The CLI with the admin token. +admin() { STEADYBIT_TOKEN=$ADMIN_TOKEN steadybit "$@"; } + +# Runs a command and checks that its output, or the file it wrote, has a line. +prints() { # expected command... + expected=$1 + shift + "$@" >out.log 2>&1 || { echo " exit $? from: $*"; tail -n 20 out.log | sed 's/^/ /'; return 1; } + grep -qF -- "$expected" out.log || { echo " no '$expected' from: $*"; tail -n 20 out.log | sed 's/^/ /'; return 1; } +} # Waits until the experiment has a run in one of the states, for up to a minute. until_run_is() { # key state... @@ -99,8 +124,31 @@ cleanup() { done steadybit experiment get -k "$key" >/dev/null 2>&1 && echo " could not delete $key" || echo " deleted $key" done + [ -n "$ADMIN_TOKEN" ] && sweep_admin rm -rf "$work" } + +# Deletes what the admin checks create, this run's and any a crashed run left: each kind +# is listed and only what is named cli-e2e-ci-* is deleted. +sweep_admin() { + named() { admin "$@" --jq ".[] | select((.name // .hubName // .templateTitle // \"\") | startswith(\"$MARK-\")) | .id" 2>/dev/null; } + for kind in webhook preflight preflight-action; do + for id in $(named integration $kind list -t json); do admin integration $kind delete -i "$id" >/dev/null 2>&1 && echo " deleted $kind $id"; done + done + for id in $(named hub list -t json); do admin hub delete -i "$id" --yes >/dev/null 2>&1 && echo " deleted hub $id"; done + for id in $(named access-token list --output json); do admin access-token delete -i "$id" --yes >/dev/null 2>&1 && echo " deleted access token $id"; done + for id in $(named service list -t json); do admin service delete -i "$id" >/dev/null 2>&1 && echo " deleted service $id"; done + for id in $(named service-profile list -t json); do admin service-profile delete -i "$id" >/dev/null 2>&1 && echo " deleted service profile $id"; done + for id in $(admin property association list --key "$PROPERTY" --output json --jq '.[].id' 2>/dev/null); do + admin property association delete -i "$id" --delete-values --yes >/dev/null 2>&1 && echo " deleted property association $id" + done + admin property definition delete -k "$PROPERTY" --yes >/dev/null 2>&1 && echo " deleted property definition $PROPERTY" + for id in $(named template list -t json); do admin template delete -i "$id" --yes >/dev/null 2>&1 && echo " deleted template $id"; done + if admin team get -k "$SCOPE_TEAM" >/dev/null 2>&1; then + admin team delete -k "$SCOPE_TEAM" --purge-experiments --yes >/dev/null 2>&1 && echo " deleted team $SCOPE_TEAM" + fi + for id in $(named environment list -t json); do admin environment delete -i "$id" --yes >/dev/null 2>&1 && echo " deleted environment $id"; done +} trap cleanup EXIT echo "steadybit $(steadybit --version) against ${STEADYBIT_URL:-https://platform.steadybit.com}, team $TEAM" @@ -175,6 +223,266 @@ check "execution list prints the platform's runs as JSON" sh -c " [ \"\$(steadybit execution list --team $TEAM --limit 2 --jq length 2>/dev/null)\" -ge 1 ] " +# --- Templates, run properties, schedules, services ------------------------------------- +# +# Tokens only ever go through the environment, never into a command line, so that a +# failing check, which shows its command, cannot show one. + +check "target stats and actions read the platform" sh -c ' + steadybit target stats -t json >/dev/null && steadybit action list --kind BASIC -t json >/dev/null +' + +# The CLI with a token read from a file an access-token command wrote. +with_token() { # file command... + file=$1 + shift + STEADYBIT_TOKEN=$(token_of "$file") "$@" +} +token_of() { sed -n 's/^ *"token": "\(.*\)",*$/\1/p' "$1"; } +json_id() { sed -n 's/^ *"id": "*\([^",]*\)"*,*$/\1/p' "$1" | head -n 1; } +json_key() { sed -n 's/^ *"key": "\(.*\)",*$/\1/p' "$1" | head -n 1; } +# GitHub hides a masked value wherever it would show in the job log. +mask() { if [ -n "${GITHUB_ACTIONS:-}" ] && [ -n "$1" ]; then echo "::add-mask::$1"; fi; } + +applied_with_id() { # file command... + file=$1 + shift + "$@" -f "$file" >out.log 2>&1 && grep -q "^id: " "$file" && grep -q "^# Written by" "$file" || { tail -n 20 out.log | sed 's/^/ /'; return 1; } +} +# What get writes is what a repository keeps; diff has to find it unchanged. +round_trips() { # get-command... -- diff-command... + get=() + while [ "$1" != -- ]; do get+=("$1"); shift; done + shift + "${get[@]}" >out.log 2>&1 && "$@" >>out.log 2>&1 || { tail -n 20 out.log | sed 's/^/ /'; return 1; } +} +schedule_enabled() { steadybit schedule list --experiment "$FROM_TEMPLATE" --jq '.[0].enabled' 2>/dev/null; } + +if [ -n "$ADMIN_TOKEN" ]; then + cat >template.yml <property.yml <association.yml </dev/null) + check "schedule get and diff find no drift" round_trips steadybit schedule get -i "$SCHEDULE" -f schedule.yml -- steadybit schedule diff -f schedule.yml + check "schedule update changes it" exits_with 0 steadybit schedule update -i "$SCHEDULE" --cron "0 0 0 1 1 ? 2099" --timezone Europe/Berlin + check "schedule diff exits with 2 after the change" exits_with 2 steadybit schedule diff -f schedule.yml + check "schedule apply puts the file back" exits_with 0 steadybit schedule apply -f schedule.yml + steadybit schedule enable -i "$SCHEDULE" >/dev/null 2>&1 + check "schedule enable enables it" test "$(schedule_enabled)" = true + steadybit schedule disable -i "$SCHEDULE" >/dev/null 2>&1 + check "schedule disable disables it" test "$(schedule_enabled)" = false + check "schedule delete deletes it" exits_with 0 steadybit schedule delete -i "$SCHEDULE" + + cat >profile.yml <service.yml </dev/null 2>&1 + check "service variable set and get" prints "cliE2e: one" steadybit service variable get -i "$SERVICE" + check "service experiment unlink removes the experiment" exits_with 0 steadybit service experiment unlink -i "$SERVICE" -k "$FROM_TEMPLATE" + check "service delete deletes it" exits_with 0 steadybit service delete -i "$SERVICE" + check "service-profile delete deletes it" exits_with 0 admin service-profile delete -i "$(id_of profile.yml)" + + # --- Environments, teams, access tokens, hubs, integrations, audit log ---------------- + + cat >environment.yml </dev/null 2>&1 + check "environment variable set and get" prints "cliE2e: one" admin environment variable get -i "$ENV_ID" + + cat >team.yml </dev/null || date -u -v+1d +%F) + admin access-token create --name "$MARK-$RUN-token" --type TEAM --team "$SCOPE_TEAM" --expires-at "$tomorrow" -t json >token.json 2>out.log + mask "$(token_of token.json)" + check "access-token create makes a team token" test -n "$(token_of token.json)" + check "the new token works" prints "$SCOPE_TEAM" with_token token.json steadybit team get -k "$SCOPE_TEAM" --jq .key + admin access-token recreate -i "$(json_id token.json)" --yes -t json >recreated.json 2>out.log + mask "$(token_of recreated.json)" + check "access-token recreate makes a new token" test -n "$(token_of recreated.json)" + check "the replaced token stops working" exits_with 1 with_token token.json steadybit team get -k "$SCOPE_TEAM" + check "the recreated token works" prints "$SCOPE_TEAM" with_token recreated.json steadybit team get -k "$SCOPE_TEAM" --jq .key + check "access-token delete deletes it" exits_with 0 admin access-token delete -i "$(json_id recreated.json)" --yes + rm -f token.json recreated.json + + # A second entry for the public hub; nothing is imported from it. + cat >hub.yml <webhook.yml <preflight.yml <preflight-action.yml < Date: Wed, 30 Sep 2026 09:37:38 +0200 Subject: [PATCH 2/2] test: the platform test fetches the hub once, as the platform limits it --- e2e/platform.sh | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/e2e/platform.sh b/e2e/platform.sh index 963e693..29793a9 100755 --- a/e2e/platform.sh +++ b/e2e/platform.sh @@ -428,9 +428,10 @@ hubName: $MARK-$RUN-hub hubLink: https://hub.steadybit.com repositoryUrl: https://raw.githubusercontent.com/steadybit/reliability-hub-db/main/index.json EOF - check "hub apply --synchronize adds the hub and fetches its templates" prints "synchronized" admin hub apply -f hub.yml --synchronize + # The platform limits how often a hub's repository is fetched, so it is fetched once. + check "hub apply adds the hub" applied_with_id hub.yml admin hub apply check "hub diff finds no drift" exits_with 0 admin hub diff -f hub.yml - check "hub resync fetches it again" prints "synchronized" admin hub resync -i "$(id_of hub.yml)" + check "hub resync fetches its templates" prints "synchronized" admin hub resync -i "$(id_of hub.yml)" check "hub delete deletes it" exits_with 0 admin hub delete -i "$(id_of hub.yml)" --yes # For the new team only, which runs nothing, so none is ever called.