Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 5 additions & 4 deletions .github/workflows/platform-e2e.yml
Original file line number Diff line number Diff line change
Expand Up @@ -38,14 +38,15 @@ jobs:
STEADYBIT_E2E_ENVIRONMENT: Global
steps:
- uses: actions/checkout@v7
# The latest release through the action, as a pipeline installs it.
- if: ${{ !inputs.from-source }}
# The latest release through the action, as a pipeline installs it; a pull request
# tests its own code, which a change to the test may depend on.
- if: ${{ !inputs.from-source && github.event_name != 'pull_request' }}
uses: ./
- if: ${{ inputs.from-source }}
- if: ${{ inputs.from-source || github.event_name == 'pull_request' }}
uses: actions/setup-go@v7
with:
go-version-file: go.mod
- if: ${{ inputs.from-source }}
- if: ${{ inputs.from-source || github.event_name == 'pull_request' }}
run: |
go build -o "$RUNNER_TEMP/bin/steadybit" ./cmd/steadybit
echo "$RUNNER_TEMP/bin" >> "$GITHUB_PATH"
Expand Down
17 changes: 17 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,22 @@
# Changelog

## v6.2.0

- `experiment run` can do what the `steadybit/run-experiment` GitHub Action does, so that
the action can run on the CLI:
- `--expect-state` passes once the run reaches a state, which need not be its end, such
as `FAILED` for an experiment expected to find a weakness, or `RUNNING`, and fails when
it ends in another; `--expect-reason` also requires the run's reason.
- `--expectation-retries` and `--expectation-retry-interval` run the experiment again
when a run did not end as expected.
- `--busy-retries` and `--busy-retry-interval` wait and try again while another
experiment is running, whether the platform refuses the run or cancels it right after
accepting it, instead of asking, failing, or running in parallel as `--yes` would.
- `--external-id` without `--template` runs the experiment with that external id.
- The JSON report gives each run's `apiLocation`.
- With `--retries`, the last attempt at a run with validation errors is kept on the
platform, so the run shows what was wrong; the attempts before it are not.

## v6.1.0

- `experiment run --parallel N` runs up to N of the experiments given with `-f` at once,
Expand Down
19 changes: 12 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -327,13 +327,18 @@ platform fills in with defaults are not reported as differences.
still watches the run until it started, for up to 15 seconds, and fails when the platform
canceled or errored it before it ran. A few options make it fit pipelines:

| Option | Does |
| ----------------------------- | ------------------------------------------------------------------------ |
| `--report steadybit.xml` | A JUnit report, one test case per step; `.json` for JSON |
| `--timeout 30m` | Cancels the run and fails when it has not ended in time |
| `--show-steps` | Prints each step's state as it changes |
| `--keep-running-on-interrupt` | Leaves the run going when the job is cancelled; by default it is stopped |
| `--parallel 3` | Runs up to 3 of the experiments at once; all are reported |
| Option | Does |
| ----------------------------- | ------------------------------------------------------------------------- |
| `--report steadybit.xml` | A JUnit report, one test case per step; `.json` for JSON |
| `--timeout 30m` | Cancels the run and fails when it has not ended in time |
| `--show-steps` | Prints each step's state as it changes |
| `--keep-running-on-interrupt` | Leaves the run going when the job is cancelled; by default it is stopped |
| `--parallel 3` | Runs up to 3 of the experiments at once; all are reported |
| `--expect-state FAILED` | Passes once the run reaches this state, and fails when it ends in another |
| `--expect-reason "…"` | Also requires the run's reason to be exactly this |
| `--expectation-retries 2` | Runs the experiment again when a run did not end as expected |
| `--busy-retries 3` | Waits and tries again while another experiment runs, instead of failing |
| `--external-id shop-latency` | Runs the experiment with this external id, instead of `-k` |

In GitHub Actions a summary of every run is added to the job summary.

Expand Down
22 changes: 22 additions & 0 deletions e2e/platform.sh
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,7 @@ experiment() { # file name duration
cat >"$1" <<EOF
# Written by e2e/platform.sh; only waits.
name: $MARK-$RUN-$2
externalId: $MARK-$RUN-$2
team: $TEAM
environment: $ENVIRONMENT
lanes:
Expand Down Expand Up @@ -148,6 +149,27 @@ until_run_is "$LONG" RUNNING
check "--no-wait fails on a run the platform refused" exits_with 1 steadybit experiment run -k "$A" --yes --no-wait
check "execution list --fail-on-match gates on that canceled run" exits_with 1 \
steadybit execution list --key "$A" --state CANCELED --ended-from "$(date -u +%F)" --limit 1 --fail-on-match
# What the run-experiment action relies on: the experiment found by its external id, an
# expected state reached before the end, and waiting while another experiment runs.
check "the experiment is found by its external id and passes at the expected state" exits_with 0 \
steadybit experiment run --external-id "$MARK-$RUN-a" --yes --allowParallel --expect-state RUNNING --report expect.json
check "the report has the state reached and the run's API location" sh -c '
grep -Eq "\"state\": *\"RUNNING\"" expect.json && grep -q "\"apiLocation\"" expect.json || { cat expect.json; exit 1; }
'
until_run_is "$A" COMPLETED CANCELED
check "a run that ends otherwise than expected fails" sh -c "
steadybit experiment run -k $A --yes --allowParallel --expect-state FAILED >otherwise.log 2>&1
status=\$?
grep -q 'but failed was expected' otherwise.log && [ \$status -eq 1 ] || { tail -n 5 otherwise.log; exit 1; }
"
# The long run still goes, so both tries are refused: what is checked is that the CLI tries
# again instead of failing at once or, as --yes would otherwise do, running in parallel.
# Waiting for the platform to be free would depend on what other suites run at the time.
check "--busy-retries tries again while another experiment runs" sh -c "
steadybit experiment run -k $A --yes --busy-retries 1 --busy-retry-interval 5s >busy.log 2>&1
status=\$?
grep -q 'trying again in 5s (1/1)' busy.log && [ \$status -eq 1 ]
"
check "execution list prints the platform's runs as JSON" sh -c "
[ \"\$(steadybit execution list --team $TEAM --limit 2 --jq length 2>/dev/null)\" -ge 1 ]
"
Expand Down
9 changes: 9 additions & 0 deletions internal/cli/experiment.go
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ import (
"errors"
"fmt"
"strings"
"time"

"github.com/spf13/cobra"
"github.com/steadybit/cli/v6/internal/experiment"
Expand Down Expand Up @@ -75,6 +76,7 @@ func newExperimentRun() *cobra.Command {
"steadybit experiment run -f ./experiments -R --yes --timeout 30m --report steadybit.xml",
"steadybit experiment run -f ./experiments -R --yes --parallel 3 --report steadybit.xml",
"steadybit experiment run --template d7e65100-1d20-4980-be87-c351704910b8 --team ADM -p CLUSTER=prod",
"steadybit experiment run --external-id shop-latency --yes --expect-state FAILED --busy-retries 3",
),
RunE: withClient(func(ctx context.Context, c *platform.Client, _ []string) error {
o.Wait = !noWait
Expand All @@ -90,13 +92,20 @@ func newExperimentRun() *cobra.Command {
f.BoolVar(&o.AllowParallel, "allowParallel", false, "Skip the prompt warning about another experiment running and allow always parallel execution.")
f.IntVar(&o.Retries, "retries", 0, "Number of retries when the experiment fails validation (e.g., missing targets). 0 means no retry.")
f.IntVar(&o.RetryInterval, "retryInterval", 10, "Interval in seconds between retries.")
f.IntVar(&o.BusyRetries, "busy-retries", 0, "When another experiment is running and running in parallel is not allowed: try again this many times instead of asking, or failing in a pipeline.")
f.DurationVar(&o.BusyRetryInterval, "busy-retry-interval", 30*time.Second, "How long to wait before trying again while another experiment is running.")
f.StringVar(&o.ExpectState, "expect-state", "", "With waiting: pass once the run reaches this state, such as FAILED or RUNNING, and fail when it ends in another. (default: COMPLETED)")
f.StringVar(&o.ExpectReason, "expect-reason", "", "With waiting: also require the run's reason to be exactly this.")
f.IntVar(&o.ExpectationRetries, "expectation-retries", 0, "With waiting: run the experiment again this many times when a run does not end as expected.")
f.DurationVar(&o.ExpectationRetryInterval, "expectation-retry-interval", time.Minute, "How long to wait before running the experiment again after a run did not end as expected.")
f.IntVar(&o.Parallel, "parallel", 1, "How many of the experiments given with -f to run at once. Each waits for its own run; the command fails if any fails.")
f.DurationVar(&o.Timeout, "timeout", 0, `With waiting: cancel the run and fail when it has not ended after this long, e.g. "15m".`)
f.BoolVar(&o.KeepRunningOnInterrupt, "keep-running-on-interrupt", false, "With waiting: leave the run going when the CLI is interrupted, instead of cancelling it.")
f.BoolVar(&o.ShowSteps, "show-steps", false, "With waiting: print each step's state as it changes.")
f.StringVar(&o.Report, "report", "", `With waiting: write a JUnit report of the runs to this file, or JSON if it ends in ".json".`)
f.Var(executionVariables, "execution-variable", "With --template: a variable for this run only, overriding experiment and environment variables. Repeat for more.")
addTemplateFlags(cmd, &o.TemplateOptions)
cmd.Flags().Lookup("external-id").Usage = "Without --template: run the experiment with this external id. With --template: an identifier of your own; using the same one again updates the experiment it created before."
// --key with one --file updates that experiment from the file and runs it, as it did.
variadic(cmd, "file")
return cmd
Expand Down
Loading
Loading