Skip to content

[bot] Merge master/ab84a23b into rel/dev - #1856

Merged
yenkins-admin merged 1 commit into
rel/devfrom
snapshot-master-ab84a23b-to-rel/dev
Oct 8, 2026
Merged

yenkins-admin merged 1 commit into
rel/devfrom
snapshot-master-ab84a23b-to-rel/dev

Conversation

@yenkins-admin

Copy link
Copy Markdown
Contributor

🚀 Automated PR to perform merge from master into rel/dev with changes up to ab84a23 (created by https://github.com/gooddata/gooddata-python-sdk/actions/runs/37741394331).

#1831)

`run_agentic_what_if` already computed `pass_power_k`; the evaluator never read
it, asking `if not summary.pass_at_k` for the verdict. `--gate power` on a flaky
item therefore reported a pass on the strength of one good run out of K -- the
exact case the gate exists to catch, reporting the opposite of what happened.

The evaluator now takes `gate` like every other multi-run kind, asks
`gate_passed(...)`, and carries the gate's note into the assertion message, which
matters here because the message body describes the BEST run and under pass^K that
can be one that passed. It also stamps the gate on the run metadata and publishes
pass@K/pass^K/gate_passed, so a gated run is readable in Langfuse rather than only
in the exit code.

The default is unchanged: without --gate this is pass@K exactly as before, pinned
by a test alongside the power case. The power test fails against the previous
code.

Found by CodeRabbit on #1798, where the same gap was fixed for forecasting.
`agentic_anomaly_detection` has it too and is fixed on its own branch, #1801,
because that evaluator does not exist on master yet.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
@yenkins-admin
yenkins-admin merged commit 13f1555 into rel/dev Oct 8, 2026
1 check passed
@yenkins-admin
yenkins-admin deleted the snapshot-master-ab84a23b-to-rel/dev branch October 8, 2026 07:06
@coderabbitai

coderabbitai Bot commented Oct 8, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 49cac0b8-d690-4aac-b467-0ca505fb95b3

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Oct 8, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 84.19%. Comparing base (0f69236) to head (ab84a23).
⚠️ Report is 605 commits behind head on rel/dev.

Additional details and impacted files
@@           Coverage Diff            @@
##           rel/dev    #1856   +/-   ##
========================================
  Coverage    84.19%   84.19%           
========================================
  Files          333      333           
  Lines        23257    23261    +4     
========================================
+ Hits         19581    19585    +4     
  Misses        3676     3676           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants