Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 9 additions & 8 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
name: CI

on:
Expand All @@ -17,9 +17,10 @@
# 15-minute budget on it depending on runner speed — so whether a green
# commit stayed green was decided by which runner it drew, and `main`
# itself was cancelled at the cap. Splitting the *suite* rather than
# raising the number keeps a real bound: each shard is roughly a third of
# the work, so the margin grows back and stays there as the suite grows
# (add a shard).
# raising the number keeps a real bound: each shard is roughly a quarter
# of the work, so the margin grows back and stays there as the suite grows
# (add a shard). Three shards had reached 13 of their 15 minutes on `main`
# by #904, and #872's tests took one past the cap: a fourth was added.
#
# `conftest.py` assigns whole test *files* to shards, deterministically,
# from the collection and the measured seconds per file in
Expand All @@ -34,7 +35,7 @@
strategy:
fail-fast: false
matrix:
shard: [1, 2, 3]
shard: [1, 2, 3, 4]
runs-on: ubuntu-latest
timeout-minutes: 15

Expand Down Expand Up @@ -83,11 +84,11 @@
# there too, so keep it out of the coverage pass rather than doing the
# same AST sweep twice on every PR.
#
# No `--cov-fail-under` here: a shard covers a third of the suite, so
# No `--cov-fail-under` here: a shard covers a quarter of the suite, so
# its coverage is a fragment. The threshold is enforced once, on the
# combined data, in `coverage` below.
env:
SHIPGATE_TEST_SHARDS: 3
SHIPGATE_TEST_SHARDS: 4
SHIPGATE_TEST_SHARD: ${{ matrix.shard }}
run: python -m pytest -n auto -m "not perf" --ignore=tests/test_adapter_static_only.py --cov=agents_shipgate --cov-report=

Expand All @@ -107,7 +108,7 @@
coverage:
# The coverage gate, enforced once over the combined fragments.
#
# Splitting the suite split its coverage with it, and three partial
# Splitting the suite split its coverage with it, and four partial
# measurements each fail an 85% threshold that the whole run passes. This
# job is where the number is a number about the suite again.
needs: [suite]
Expand Down Expand Up @@ -148,7 +149,7 @@
# test pass.
#
# `coverage combine` consumes its inputs, so the fragments are copied
# to distinct names first — three files all called `.coverage` in
# to distinct names first — four files all called `.coverage` in
# separate directories would otherwise collide on the way in.
run: |
set -euo pipefail
Expand Down
16 changes: 16 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,22 @@
- **Submodules.** A gitlink under the scope is materialized as the empty directory a checkout without `--recurse-submodules` leaves; its content is never fetched. The same gitlink commit on both sides is named in `limits` as unchanged and does not make the comparison `partial`, since identical content cannot carry a change. An added, removed or moved gitlink is a coverage gap over its path on each side that has it, and over the whole scope when it is the selected scope itself, so the result is `partial`, never `compared`. A gitlink was refused (exit 2) before.
- **Google ADK.** `google.adk.Agent`, the package-root re-export of `google.adk.agents.llm_agent.Agent`, is read as an agent constructor by the ADK reader, for `from google.adk import Agent` and `import google.adk as adk` alike. CubeSandbox#1508 now establishes `cube_code_agent` and is `partial`, naming the tool it imports from a sibling module (#864), instead of `not_established`.
- **Re-measured.** At the root scope, main `a430e81a` refused CubeSandbox#1508, dlt#4417 and asterinas#3834; this change compares each, with its own limits (a Python census past `--max-python-files`, an unsupported framework, an unresolved import). No application comparison schema change; `application_comparison_schema_version` stays `"0.1"`.
- `diff --application` rows now name what a bound tool reaches: the endpoint, the request fields, the credential it sends and which arguments the model chooses. (#872, part of #868)
- **The problem.** tensorflow/tensorflow#128063 was the only one of 134 open SDK/ADK PRs whose rows were both correct and complete. Even so, each row named only a signature. `submit_pr_code_review` posts a pull-request review whose `event` can be `APPROVE`, using `GITHUB_TOKEN`; `get_pull_request_details` only reads. The existing assessment gave all three tools `write` with unknown evidence.
- **Reach.** Each side of a function-tool row carries `reach`: the outbound `requests`, `httpx`, `aiohttp` and `urllib` calls read from the tool's own code and its same-scope helpers, up to three calls deep. Every call the read cannot follow is a named limit with its location. Each call records:
- the method and the URL template (`{pr_number}`, `{env OWNER}`, `{…}`);
- literal request fields, and the literals a field is chosen from with the parameters that decide;
- `model_supplied`: which parameters flow where;
- `credential_sources`: environment variable names only, never a value;
- for GraphQL, query or mutation.
- **Effect evidence.** `effect_evidence` is the existing `assess_tool_semantics` over the tool, with the reach as one more structural source (`source_http_call`). A write or delete call supports that effect. `read` needs every call followed and every outbound call a read. Read-only method names pass only on plain data or on objects a known library returned. A decorator, a store into an unknown object, or a method on an object the read cannot name blocks `read`.
- **Secrets.** A hard-coded credential is `literal: true` and never printed. In a URL, a query value prints only when it is a short lowercase word or a number, and a token-shaped path piece is withheld. A field value prints only when it is a plain word and its name does not suggest a secret.
- **Two constructions of one ADK name.** Each row's `binding_location` names the first construction that lists the tool, and `construction_sites` lists every one. Before, every row pointed at the first construction.
- **Measured.** On the 134-PR corpus, rows, statuses and exit codes are unchanged, and run time is flat. Of 96 row sides in 19 PRs:
- 9 name outbound calls (3 PRs) and 5 name a credential;
- 7 carry structural effect evidence (3 read, 4 write);
- 65 name at least one limit, mostly database helpers, tracing and SDK clients.
- `reach`, `effect_evidence` and `construction_sites` are evidence outside the compared meaning. `application_comparison_schema_version` is `0.2`.
- Follow a tool an agent binds from another module in the repository. (#864)
- **The problem.** jpka/attest#3 added `memory_bank.remember_firm_finding` and `memory_bank.recall_firm_memory` to an ADK agent's `tools=[...]`, with the module brought in by `from . import memory_bank`; the reader stopped at the module boundary, so `diff --application` showed no row and `scan` catalogued only the unchanged local tools. The same gap left OpenAI Agents SDK tools such as `from ..tools.shop import add_to_cart` unresolved.
- **What resolves.** For Google ADK, a name imported from a sibling module or re-exported by a package, a module-qualified `module.function`, a plain `alias = function`, and `FunctionTool(imported_function)` / `LongRunningFunctionTool(...)`, including a wrapper built in the imported module; for the OpenAI Agents SDK, a name or `module.function` that reaches a definition carrying the SDK's `@function_tool`. The tool is the definition, with its own signature, location and implementation digest, so jpka/attest#3 now shows the two memory tools as `ADDED` and leaves its four unchanged bindings, `scorer.score_answer` included, alone. A definition reached by several spellings is one tool; same-named functions in different modules stay two.
Expand Down
Loading
Loading