From a837aa668026e3f0006816d64dfe453337c5ba6c Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Fri, 11 Sep 2026 12:01:13 +0200 Subject: [PATCH 01/25] test(flagd): adopt the OpenFeature Provider TCK Runs the conformance suite against the flagd provider in both resolver modes. The whole adoption is a shared abstract base and two subclasses that differ only in resolver and port: the TCK brings its own Gherkin, its own step definitions and its own Compose lifecycle, and works out which suite is running from the JUnit test plan, so a mode needs no registration and no build configuration. The Compose stack wraps the unmodified flagd-testbed image, which already serves both flagd and the launchpad control API that this TCK's control API contract was derived from. No host port bindings: the TCK discovers dynamically mapped ports after startup, so the suite runs in parallel and does not collide with a developer's local flagd. capabilities() is declarableExcept(NUMERIC_COERCION). Evaluating float-flag (0.5) through the integer API returns 0 with no error code rather than TYPE_MISMATCH with the code default -- the value is silently truncated. Coercion as such is permitted; it is the lossy case being accepted that is the defect, tracked as open-feature/flagd#1996. Both resolvers behave identically, which places it in the shared provider layer rather than in either transport, so it is declared once here. Delete the override when the defect is fixed. Split out of #1830 so that the suite and its first adopter are reviewed as separate questions: whether the TCK is the right contract, and whether flagd satisfies it. Signed-off-by: Simon Schrottner --- providers/flagd/pom.xml | 13 ++ .../flagd/e2e/AbstractFlagdTckTest.java | 122 ++++++++++++++++++ .../flagd/e2e/FlagdInProcessTckTest.java | 17 +++ .../providers/flagd/e2e/FlagdRpcTckTest.java | 17 +++ .../test/resources/tck/docker-compose.yaml | 15 +++ 5 files changed, 184 insertions(+) create mode 100644 providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java create mode 100644 providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdInProcessTckTest.java create mode 100644 providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdRpcTckTest.java create mode 100644 providers/flagd/src/test/resources/tck/docker-compose.yaml diff --git a/providers/flagd/pom.xml b/providers/flagd/pom.xml index 8980e8f090..15001f4bbc 100644 --- a/providers/flagd/pom.xml +++ b/providers/flagd/pom.xml @@ -22,6 +22,8 @@ 1.2.28 [2.0.0,3.0.0) + + [0.0.1,) flagd @@ -98,6 +100,17 @@ 5.14.3 test + + + dev.openfeature.contrib.tools + provider-tck + ${provider-tck.version} + test + org.testcontainers testcontainers diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java new file mode 100644 index 0000000000..a762a4afd9 --- /dev/null +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -0,0 +1,122 @@ +package dev.openfeature.contrib.providers.flagd.e2e; + +import dev.openfeature.contrib.providers.flagd.Config; +import dev.openfeature.contrib.providers.flagd.FlagdOptions; +import dev.openfeature.contrib.providers.flagd.FlagdProvider; +import dev.openfeature.contrib.tools.providertck.BackendEndpoint; +import dev.openfeature.contrib.tools.providertck.Capability; +import dev.openfeature.contrib.tools.providertck.ContainerizedProviderTckTest; +import dev.openfeature.sdk.FeatureProvider; +import java.io.File; +import java.util.Collections; +import java.util.List; +import java.util.Set; + +/** + * Shared configuration for running the OpenFeature Provider TCK against the flagd provider. + * + *

flagd resolves flags in two quite different ways, and both are worth conforming: RPC evaluates + * remotely over gRPC, while in-process syncs the ruleset and evaluates locally. They share a backend + * stack and differ only in resolver and port, so the modes are two small subclasses. + * + *

Each concrete subclass is its own JUnit suite and its own TCK harness; the TCK works out which + * one is running from the JUnit test plan, so adding a mode needs no registration or build + * configuration. + */ +abstract class AbstractFlagdTckTest extends ContainerizedProviderTckTest { + + /** + * A port nothing listens on, for the initialisation-failure scenarios. + * + *

Deliberately not a port on the Compose stack: the stack must stay up for the whole suite, + * and simulated outages belong to the control API. + */ + private static final int UNAVAILABLE_PORT = 9999; + + /** + * gRPC deadline for a provider that is expected to connect. + * + *

Generous on purpose. flagd derives its initialisation deadline from this value, and the + * in-process resolver must sync the entire ruleset before it reports ready — which intermittently + * takes longer than a deadline tuned for a single RPC round trip. + */ + private static final int CONNECTED_DEADLINE_MS = 5000; + + /** + * gRPC deadline for a provider pointed at a dead port. + * + *

Short on purpose, and deliberately not the same as {@link #CONNECTED_DEADLINE_MS}: the + * initialisation-failure scenarios assert that the failure is reported promptly, so a + * provider that takes as long to give up as it does to connect would defeat the point. + */ + private static final int UNAVAILABLE_DEADLINE_MS = 1000; + + /** The resolver under test. */ + protected abstract Config.Resolver resolver(); + + /** The container-internal port that resolver connects to. */ + protected abstract int backendPort(); + + @Override + public File composeFile() { + return new File("src/test/resources/tck/docker-compose.yaml"); + } + + @Override + public List backendPorts() { + return Collections.singletonList(backendPort()); + } + + @Override + public FeatureProvider createProvider(BackendEndpoint endpoint) { + return new FlagdProvider(baseOptions() + .deadline(CONNECTED_DEADLINE_MS) + .host(endpoint.host()) + .port(endpoint.port(backendPort())) + .build()); + } + + @Override + public FeatureProvider createUnavailableProvider() { + return new FlagdProvider(baseOptions() + .deadline(UNAVAILABLE_DEADLINE_MS) + .host("localhost") + .port(UNAVAILABLE_PORT) + .build()); + } + + /** + * {@inheritDoc} + * + *

Everything declarable except {@link Capability#NUMERIC_COERCION}. Evaluating + * {@code float-flag} (0.5) through the integer API returns {@code 0} with no error code + * rather than {@code TYPE_MISMATCH} with the code default — the value is silently truncated. + * Coercion as such is permitted, and the capability says so: the rule is that a lossless + * coercion must succeed and a lossy one must fail. It is the lossy case being accepted that is a + * defect to fix, not a design choice; this override should be deleted once it is. + * + *

Declared here rather than per mode because both resolvers behave identically, which places + * the defect in the shared provider layer rather than in either transport. Every other + * capability, including the full non-numeric type-mismatch matrix, holds in both modes. + * + *

That includes {@link Capability#LIFECYCLE}, and legitimately so: flagd reaches its backend + * during initialisation in both modes — an RPC round trip, or a full ruleset sync — so the + * lifecycle scenarios assert something real here rather than passing vacuously. + * + *

{@link Capability#declarableExcept} rather than {@code EnumSet.complementOf}, which is what + * this used to be. The complement of one capability is every other enum constant, + * including {@code @targeting} and {@code @caching} — reserved tags no scenario carries — so a + * report emitted from here claimed two capabilities nothing had examined. + */ + @Override + public Set capabilities() { + return Capability.declarableExcept(Capability.NUMERIC_COERCION); + } + + private FlagdOptions.FlagdOptionsBuilder baseOptions() { + return FlagdOptions.builder() + .resolverType(resolver()) + .retryGracePeriod(2) + .retryBackoffMs(500); + } +} diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdInProcessTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdInProcessTckTest.java new file mode 100644 index 0000000000..1ce5b1dc71 --- /dev/null +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdInProcessTckTest.java @@ -0,0 +1,17 @@ +package dev.openfeature.contrib.providers.flagd.e2e; + +import dev.openfeature.contrib.providers.flagd.Config; + +/** Runs the OpenFeature Provider TCK against the flagd provider in in-process mode. */ +public class FlagdInProcessTckTest extends AbstractFlagdTckTest { + + @Override + protected Config.Resolver resolver() { + return Config.Resolver.IN_PROCESS; + } + + @Override + protected int backendPort() { + return 8015; + } +} diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdRpcTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdRpcTckTest.java new file mode 100644 index 0000000000..30ee6db57c --- /dev/null +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdRpcTckTest.java @@ -0,0 +1,17 @@ +package dev.openfeature.contrib.providers.flagd.e2e; + +import dev.openfeature.contrib.providers.flagd.Config; + +/** Runs the OpenFeature Provider TCK against the flagd provider in RPC mode. */ +public class FlagdRpcTckTest extends AbstractFlagdTckTest { + + @Override + protected Config.Resolver resolver() { + return Config.Resolver.RPC; + } + + @Override + protected int backendPort() { + return 8013; + } +} diff --git a/providers/flagd/src/test/resources/tck/docker-compose.yaml b/providers/flagd/src/test/resources/tck/docker-compose.yaml new file mode 100644 index 0000000000..4cecaa1388 --- /dev/null +++ b/providers/flagd/src/test/resources/tck/docker-compose.yaml @@ -0,0 +1,15 @@ +# Backend stack for the OpenFeature Provider TCK, wrapping the unmodified flagd testbed image. +# +# The image already serves everything the TCK needs: flagd itself, and the "launchpad" control +# API on 8080 whose endpoints this TCK's control API contract was derived from. +# +# Note there are no host port bindings. The TCK requires dynamically mapped ports and discovers +# them after startup — a pinned host port would make the suite unrunnable in parallel and would +# collide with a developer's local flagd. +services: + backend: + image: ghcr.io/open-feature/flagd-testbed:v3.8.0 + ports: + - 8013 # flagd RPC evaluation (gRPC) + - 8015 # flagd in-process sync (gRPC) + - 8080 # launchpad control API From 1dfe32224615bf8d5fecfe91b74244cc1a491a55 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Fri, 11 Sep 2026 12:05:52 +0200 Subject: [PATCH 02/25] test(flagd): declare the numeric-coercion gap as a defect, not a choice capabilities() withholding @numeric-coercion reads, from the outside, exactly like a provider with no streaming transport declining @configuration-change: the scenarios are skipped either way and nothing in the run says which of the two happened. One is a limitation, the other is a bug, and a consumer comparing providers needs to be able to tell. knownDeviations() is the only place that can say so, because only the provider author knows. Tracked against open-feature/flagd#1996. The summary names the half of the coercion rule that is broken -- the lossy one -- because "flagd coerces numbers" on its own reads as intended behaviour rather than as a defect. Delete this and the capabilities() override together, once evaluating float-flag (0.5) through the integer API reports TYPE_MISMATCH. Signed-off-by: Simon Schrottner --- .../flagd/e2e/AbstractFlagdTckTest.java | 35 +++++++++++++++++-- 1 file changed, 33 insertions(+), 2 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index a762a4afd9..ac2a811c71 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -6,6 +6,7 @@ import dev.openfeature.contrib.tools.providertck.BackendEndpoint; import dev.openfeature.contrib.tools.providertck.Capability; import dev.openfeature.contrib.tools.providertck.ContainerizedProviderTckTest; +import dev.openfeature.contrib.tools.providertck.KnownDeviation; import dev.openfeature.sdk.FeatureProvider; import java.io.File; import java.util.Collections; @@ -105,14 +106,44 @@ public FeatureProvider createUnavailableProvider() { * *

{@link Capability#declarableExcept} rather than {@code EnumSet.complementOf}, which is what * this used to be. The complement of one capability is every other enum constant, - * including {@code @targeting} and {@code @caching} — reserved tags no scenario carries — so a - * report emitted from here claimed two capabilities nothing had examined. + * including {@code @targeting} and {@code @caching} — reserved tags no scenario carries — so + * declaring the complement claimed two capabilities nothing had examined. */ @Override public Set capabilities() { return Capability.declarableExcept(Capability.NUMERIC_COERCION); } + /** + * {@inheritDoc} + * + *

The withheld {@link Capability#NUMERIC_COERCION} is a defect, not a limitation, and + * something has to say so. In the results the two are indistinguishable: the scenario is skipped + * either way, and the declaration explains only that the capability was not claimed, + * never whether flagd chose not to claim it. A consumer comparing providers would otherwise read + * this exactly as it reads a provider with no streaming transport declining + * {@code @configuration-change}, which is a decision rather than a bug. + * + *

Tracked against flagd's numeric coercion ADR, which is where the rule this deviates from is + * settled: coercion is permitted when it is lossless and must fail with {@code TYPE_MISMATCH} + * only when information would be lost. The summary says which half is broken, because "flagd + * coerces numbers" on its own reads as a description of intended behaviour. Delete the entry — + * and the {@code capabilities()} override above — once the lossy case reports + * {@code TYPE_MISMATCH}. + */ + @Override + public List knownDeviations() { + return Collections.singletonList(KnownDeviation.tracked( + Capability.NUMERIC_COERCION, + "https://github.com/open-feature/flagd/issues/1996", + "The lossy half of the coercion rule is not enforced: evaluating float-flag (0.5) " + + "through the integer API returns 0 with no error code, rather than " + + "TYPE_MISMATCH with the code default, so the fractional part is discarded " + + "silently. Lossless coercion is permitted and is not the defect. Both " + + "resolvers behave identically, which places it in the shared provider layer " + + "rather than in either transport.")); + } + private FlagdOptions.FlagdOptionsBuilder baseOptions() { return FlagdOptions.builder() .resolverType(resolver()) From e42458ba8751b0676e9d277ee3499f1137c12309 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Fri, 11 Sep 2026 15:10:02 +0200 Subject: [PATCH 03/25] test(flagd): record what flagd-testbed does not serve, and why that is not a deviation Three of the flags the suite's assets added are absent from flagd-testbed v3.8.0: large-integer-flag, huge-integer-flag and integral-float-flag. Only the first is reached by a scenario that runs here -- huge-integer-flag is asked for solely under @large-integers, which is not applicable in Java, and integral-float-flag solely under @numeric-coercion, which this provider withholds -- so exactly one untagged scenario, the 32-bit precision one, fails FLAG_NOT_FOUND in both modes. open-feature/flagd-testbed#392 is open for it; the Compose tag gets bumped when it lands, which is why the note lives next to the tag as well as in the class. A missing flag is a gap in the stack, not in the provider, so it is documented rather than declared as a KnownDeviation. A deviation says the provider is wrong, and the provider was never given the flag to get wrong. The three falsy flags used to fail the same way and no longer do, which is worth writing down because the failure looked identical. The testbed's zero-flags.json already served boolean-zero-flag, integer-zero-flag and string-zero-flag with zero/non-zero variants, while the canonical set called them false-flag, zero-flag and empty-string-flag; the base moved the canonical names onto the testbed's rather than the other way round, so those three scenarios now resolve against flags that were always there. Also says why capabilities() calls declarableExcept rather than EnumSet.complementOf, which now matters more than it did: the complement would claim @large-integers as well as the two reserved tags, and the suite refuses that declaration at startup. Signed-off-by: Simon Schrottner --- .../flagd/e2e/AbstractFlagdTckTest.java | 26 ++++++++++++++++++- .../test/resources/tck/docker-compose.yaml | 8 ++++-- 2 files changed, 31 insertions(+), 3 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index ac2a811c71..ef0d6c141f 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -23,6 +23,26 @@ *

Each concrete subclass is its own JUnit suite and its own TCK harness; the TCK works out which * one is running from the JUnit test plan, so adding a mode needs no registration or build * configuration. + * + *

The testbed does not yet serve the whole canonical flag set. Three of the + * flags the suite's assets added are absent from {@code flagd-testbed} v3.8.0: + * {@code large-integer-flag}, {@code huge-integer-flag} and {@code integral-float-flag}. Only the + * first is reached — {@code huge-integer-flag} is asked for solely under {@code @large-integers}, + * which is not applicable in Java, and {@code integral-float-flag} solely under + * {@code @numeric-coercion}, which is withheld below — so exactly one untagged scenario, the 32-bit + * precision one, fails with {@code FLAG_NOT_FOUND} in both modes until + * open-feature/flagd-testbed#392 lands and the tag here is bumped. + * + *

The three falsy flags used to fail the same way and no longer do. The testbed's + * {@code zero-flags.json} already served {@code boolean-zero-flag}, {@code integer-zero-flag} and + * {@code string-zero-flag} with {@code zero}/{@code non-zero} variants, while the canonical set + * called them {@code false-flag}, {@code zero-flag} and {@code empty-string-flag}; spec ba002ce8 + * renamed the canonical flags to the testbed's names rather than the other way round, so those + * three scenarios now resolve against flags that were always there. + * + *

A missing flag is a gap in the stack, not in the provider, so it is recorded here rather than + * declared as a {@link KnownDeviation}: a deviation says the provider is wrong, and the provider was + * never given the flag to get wrong. */ abstract class AbstractFlagdTckTest extends ContainerizedProviderTckTest { @@ -107,7 +127,11 @@ public FeatureProvider createUnavailableProvider() { *

{@link Capability#declarableExcept} rather than {@code EnumSet.complementOf}, which is what * this used to be. The complement of one capability is every other enum constant, * including {@code @targeting} and {@code @caching} — reserved tags no scenario carries — so - * declaring the complement claimed two capabilities nothing had examined. + * declaring the complement claimed two capabilities nothing had examined. It would now also + * claim {@link Capability#LARGE_INTEGERS}, which no Java provider can have — the SDK's integer + * accessor is a 32-bit {@code Integer} — and the suite refuses such a declaration at startup. + * {@code declarableExcept} leaves the not-applicable tag out on its own, and its one scenario is + * reported as skipped with that reason on every run. */ @Override public Set capabilities() { diff --git a/providers/flagd/src/test/resources/tck/docker-compose.yaml b/providers/flagd/src/test/resources/tck/docker-compose.yaml index 4cecaa1388..9f06eab454 100644 --- a/providers/flagd/src/test/resources/tck/docker-compose.yaml +++ b/providers/flagd/src/test/resources/tck/docker-compose.yaml @@ -1,7 +1,11 @@ # Backend stack for the OpenFeature Provider TCK, wrapping the unmodified flagd testbed image. # -# The image already serves everything the TCK needs: flagd itself, and the "launchpad" control -# API on 8080 whose endpoints this TCK's control API contract was derived from. +# The image serves flagd itself and the "launchpad" control API on 8080, whose endpoints this +# TCK's control API contract was derived from. It does not yet serve the whole canonical flag +# set: v3.8.0 has none of large-integer-flag, huge-integer-flag or integral-float-flag. Only +# large-integer-flag is reached by a scenario that runs here -- the other two sit behind +# @large-integers and @numeric-coercion, which this provider does not declare -- so one untagged +# scenario fails FLAG_NOT_FOUND until open-feature/flagd-testbed#392 lands. Bump the tag then. # # Note there are no host port bindings. The TCK requires dynamically mapped ports and discovers # them after startup — a pinned host port would make the suite unrunnable in parallel and would From aeb7ed2c4345b942fe43ce61949bed650c68dd43 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Fri, 11 Sep 2026 15:11:11 +0200 Subject: [PATCH 04/25] test(flagd): withhold @reinitialization, and settle @stale by evidence Two declarations examined rather than asserted, one of which retracts a mistake. @reinitialization is withheld, and no KnownDeviation accompanies it. shutdown() sets the sync resources' isShutDown flag and never clears isInitialized (FlagdProvider.java:136-155, FlagdProviderSyncResources.java:27-28, 112-115), so a later initialize() returns at its first check without rebuilding the resolver, the gRPC channel it shutdownNow()'d, the retry scheduler it terminated or the final errorExecutor it tore down (FlagdProvider.java:121-125). A shut-down flagd provider is terminally shut down. That is permitted. Requirement 2.5.2 says a provider SHOULD revert to its uninitialized state after shutdown, and its supporting text says "some providers MAY allow reinitialization from this state". Reuse is an option, not an obligation, and declining it is one of the choices the requirement offers. An earlier version of this file recorded it as an untracked KnownDeviation against @lifecycle, which was wrong twice over: the scenario was mandatory only because the spec's assets had not yet gated it, and the entry asserted a defect against a provider behaving inside the requirement. Withholding the tag is the whole of what is owed; the one scenario it gates is now reported as skipped with that reason instead of failing in RPC mode. The lesson is more useful than the correction. Nothing had checked whether 2.5.2 requires reuse before the failure was written up as a defect -- the scenario failed, so a deviation was recorded. Find the numbered requirement first. This is the third rule in the suite found asserted more strongly than the spec states it. @lifecycle stays declared. flagd reaches its backend during initialisation in both modes, so the remaining lifecycle scenarios assert something real, and Java declaring it is what made the cross-language divergence visible in the first place -- Go and JavaScript withhold it and are being changed to match. @stale is declared for both resolvers and that is correct. PROVIDER_STALE is emitted from FlagdProvider.onError (FlagdProvider.java:258-264), which the shared onProviderEvent switch reaches on PROVIDER_ERROR from either resolver (FlagdProvider.java:197, 236), so the emit is in the provider layer rather than a transport -- and the scenario passes in RPC mode as well as in-process. Go's flagd provider withholds the tag for RPC; on this evidence that is a difference between the implementations, not a property of the transport. Signed-off-by: Simon Schrottner --- .../flagd/e2e/AbstractFlagdTckTest.java | 38 +++++++++++++++++-- 1 file changed, 35 insertions(+), 3 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index ef0d6c141f..aea26fb3fd 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -109,7 +109,9 @@ public FeatureProvider createUnavailableProvider() { /** * {@inheritDoc} * - *

Everything declarable except {@link Capability#NUMERIC_COERCION}. Evaluating + *

Everything declarable except {@link Capability#NUMERIC_COERCION} and + * {@link Capability#REINITIALIZATION}. The two omissions are different in kind: the first is a + * defect and is declared as one below, the second is a choice the specification offers. Evaluating * {@code float-flag} (0.5) through the integer API returns {@code 0} with no error code * rather than {@code TYPE_MISMATCH} with the code default — the value is silently truncated. * Coercion as such is permitted, and the capability says so: the rule is that a lossless @@ -122,7 +124,37 @@ public FeatureProvider createUnavailableProvider() { * *

That includes {@link Capability#LIFECYCLE}, and legitimately so: flagd reaches its backend * during initialisation in both modes — an RPC round trip, or a full ruleset sync — so the - * lifecycle scenarios assert something real here rather than passing vacuously. + * lifecycle scenarios assert something real here rather than passing vacuously. Worth stating + * because the Go and JavaScript flagd providers withhold it; Java declaring it is what made that + * divergence visible, and the other two are being changed to match rather than the reverse. + * + *

{@link Capability#REINITIALIZATION} is withheld, and that is a fact about the + * provider rather than a defect in it. {@code shutdown()} sets the sync resources' own + * {@code isShutDown} flag and never clears {@code isInitialized} + * (FlagdProvider.java:136-155, FlagdProviderSyncResources.java:27-28, 112-115), so a later + * {@code initialize()} returns at its first check without rebuilding anything + * (FlagdProvider.java:121-125): the resolver is shut down, the RPC channel was + * {@code shutdownNow()}'d, the retry scheduler is terminated and {@code errorExecutor} is a + * {@code final} field nothing re-creates. A shut-down flagd provider is terminally shut down. + * + *

Requirement 2.5.2 says a provider SHOULD revert to its uninitialized state after + * shutdown, and its supporting text says "some providers MAY allow reinitialization from + * this state" — so reuse is permitted, not required, and declining it is one of the options + * the requirement offers. An earlier version of this file recorded it as a {@code KnownDeviation} + * against {@code @lifecycle}, which was wrong twice over: the scenario was mandatory only because + * the spec's assets had not yet gated it, and the entry asserted a defect against a provider + * behaving within the requirement. Withholding the tag is the whole of what is owed here; the + * one scenario it gates is reported as skipped with this reason on every run. + * + *

{@link Capability#STALE} is declared for both resolvers, and the declaration is examined + * rather than inherited. {@code PROVIDER_STALE} is emitted from {@code FlagdProvider.onError} + * (FlagdProvider.java:258-264), which the shared {@code onProviderEvent} switch reaches on + * {@code PROVIDER_ERROR} from either resolver (FlagdProvider.java:197, 236), before the grace + * period turns it into {@code PROVIDER_ERROR} — so the emit sits in the provider layer, not in a + * transport, and the scenario "Losing the backend makes the provider stale, regaining it makes + * it ready again" passes in RPC mode as well as in-process. Worth stating because Go's flagd + * provider withholds the tag for its RPC resolver; on this evidence that is a difference between + * the two implementations, not a property of the transport. * *

{@link Capability#declarableExcept} rather than {@code EnumSet.complementOf}, which is what * this used to be. The complement of one capability is every other enum constant, @@ -135,7 +167,7 @@ public FeatureProvider createUnavailableProvider() { */ @Override public Set capabilities() { - return Capability.declarableExcept(Capability.NUMERIC_COERCION); + return Capability.declarableExcept(Capability.NUMERIC_COERCION, Capability.REINITIALIZATION); } /** From 085c1645a50750a296f981d174f23e2cd3ebcdb3 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Fri, 11 Sep 2026 18:53:17 +0200 Subject: [PATCH 05/25] test(flagd): withhold @large-integers explicitly The TCK no longer sets @large-integers apart as "not applicable in Java". It is an ordinary declarable capability, so declarableExcept(...) no longer leaves it out on its own and this suite has to say so. Nothing about the run changes. The scenario was skipped before and is skipped now, for the same reason in substance: the SDK's integer accessor is a 32-bit Integer and 2^53 - 1 has no room in it, so no Java provider can hold the tag. What changed is where that is written down -- Appendix F, once, rather than a field in every report -- and that the suite now expects the harness to withhold it rather than refusing to let it be declared. It stays out of knownDeviations, for the same reason the missing testbed flags do: flagd is not at fault for a value the accessor cannot carry. Signed-off-by: Simon Schrottner --- .../flagd/e2e/AbstractFlagdTckTest.java | 23 +++++++++++-------- 1 file changed, 14 insertions(+), 9 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index aea26fb3fd..c95e893bf4 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -27,9 +27,9 @@ *

The testbed does not yet serve the whole canonical flag set. Three of the * flags the suite's assets added are absent from {@code flagd-testbed} v3.8.0: * {@code large-integer-flag}, {@code huge-integer-flag} and {@code integral-float-flag}. Only the - * first is reached — {@code huge-integer-flag} is asked for solely under {@code @large-integers}, - * which is not applicable in Java, and {@code integral-float-flag} solely under - * {@code @numeric-coercion}, which is withheld below — so exactly one untagged scenario, the 32-bit + * first is reached — {@code huge-integer-flag} is asked for solely under {@code @large-integers} and + * {@code integral-float-flag} solely under {@code @numeric-coercion}, both withheld below — so + * exactly one untagged scenario, the 32-bit * precision one, fails with {@code FLAG_NOT_FOUND} in both modes until * open-feature/flagd-testbed#392 lands and the tag here is bumped. * @@ -156,18 +156,23 @@ public FeatureProvider createUnavailableProvider() { * provider withholds the tag for its RPC resolver; on this evidence that is a difference between * the two implementations, not a property of the transport. * + *

{@link Capability#LARGE_INTEGERS} is withheld, as it is by every Java provider. It asks for + * 2^53 − 1 and {@code Client.getIntegerDetails} is a 32-bit {@code Integer} with no room for it, + * so the limit is the SDK's rather than flagd's — Appendix F is where that is recorded, and it is + * not a {@link dev.openfeature.contrib.tools.providertck.KnownDeviation}, because flagd is not at + * fault for a value the accessor cannot carry. Its one scenario is reported as skipped for an + * undeclared capability, like any other. + * *

{@link Capability#declarableExcept} rather than {@code EnumSet.complementOf}, which is what * this used to be. The complement of one capability is every other enum constant, * including {@code @targeting} and {@code @caching} — reserved tags no scenario carries — so - * declaring the complement claimed two capabilities nothing had examined. It would now also - * claim {@link Capability#LARGE_INTEGERS}, which no Java provider can have — the SDK's integer - * accessor is a 32-bit {@code Integer} — and the suite refuses such a declaration at startup. - * {@code declarableExcept} leaves the not-applicable tag out on its own, and its one scenario is - * reported as skipped with that reason on every run. + * declaring the complement claimed two capabilities nothing had examined, and the suite refuses + * such a declaration at startup. */ @Override public Set capabilities() { - return Capability.declarableExcept(Capability.NUMERIC_COERCION, Capability.REINITIALIZATION); + return Capability.declarableExcept( + Capability.NUMERIC_COERCION, Capability.REINITIALIZATION, Capability.LARGE_INTEGERS); } /** From 7de6f99cebc770af0512bf8b125752ce9d552a26 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sat, 12 Sep 2026 08:37:48 +0200 Subject: [PATCH 06/25] docs(flagd): say what the new tags gate here, and what the testbed still owes @variants and @targeting arrived with the base's submodule bump, and declarableExcept picks both up without a line of this file changing. That is the right outcome and the reason nothing here said so, which is the problem: the declaration grew by two capabilities and the file that argues every other one either way was silent about them. Two runs against flagd-testbed v3.8.0, one per resolver. RPC: 45 pass, 5 skipped, 2 failed. In-process: the same, after a first attempt whose opening scenario timed out on a cold sync stream and passed on rerun -- a warm-up flake, not a result. @targeting's three scenarios resolve targeting-key-flag through flagd's own rule evaluation and pass in both modes on the image already pinned, so the tag cost no bump; the Compose header now says so, because the note beside that tag is where someone would otherwise go looking for a reason to bump it. Both failures are the one testbed gap, and the @variants outline reaches it a second time: once for the untagged precision scenario, once for the row asking for large-integer-flag's max-int32 variant. So the header's "one untagged scenario" is a scenario short. Withholding @variants would hide both -- and would be a claim about flagd made to accommodate a missing flag, which is the one thing a declaration must not be. The capabilities() javadoc also still called @targeting reserved, in the paragraph explaining why complementOf is the wrong call. It is declarable now, and the paragraph is stronger for losing it: the reserved set shrinks as the vocabulary fills up, so what protects the declaration is the form of the call rather than the size of the set. Signed-off-by: Simon Schrottner --- .../flagd/e2e/AbstractFlagdTckTest.java | 18 +++++++++++++++--- .../src/test/resources/tck/docker-compose.yaml | 11 ++++++++--- 2 files changed, 23 insertions(+), 6 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index c95e893bf4..520056c0a4 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -163,11 +163,23 @@ public FeatureProvider createUnavailableProvider() { * fault for a value the accessor cannot carry. Its one scenario is reported as skipped for an * undeclared capability, like any other. * + *

{@link Capability#VARIANTS} and {@link Capability#TARGETING} are both declared, and both on + * evidence rather than by inheriting the "everything except" default. flagd names the variant it + * served in every resolution, so seven of the {@code @variants} outline's eight rows pass in both + * modes; the eighth asks for {@code large-integer-flag}'s {@code max-int32} and is answered with + * no variant because testbed v3.8.0 does not serve that flag at all — the same gap that fails the + * untagged precision scenario, recorded next to the image tag in the Compose file rather than + * here. {@code @targeting} is the newer claim and the cheaper one to check: its three scenarios + * resolve {@code targeting-key-flag} through flagd's own rule evaluation and all three pass in + * both modes on the image already pinned, so nothing about it needed a testbed bump. + * *

{@link Capability#declarableExcept} rather than {@code EnumSet.complementOf}, which is what * this used to be. The complement of one capability is every other enum constant, - * including {@code @targeting} and {@code @caching} — reserved tags no scenario carries — so - * declaring the complement claimed two capabilities nothing had examined, and the suite refuses - * such a declaration at startup. + * including {@code @caching} — a reserved tag no scenario carries — so declaring the complement + * claimed a capability nothing had examined, and the suite refuses such a declaration at startup. + * It swept up {@code @targeting} the same way until that tag gated something, which is the point: + * the hazard shrinks as the vocabulary fills up and never disappears, so the form of the call is + * what protects the declaration, not the current size of the reserved set. */ @Override public Set capabilities() { diff --git a/providers/flagd/src/test/resources/tck/docker-compose.yaml b/providers/flagd/src/test/resources/tck/docker-compose.yaml index 9f06eab454..8feeec6c55 100644 --- a/providers/flagd/src/test/resources/tck/docker-compose.yaml +++ b/providers/flagd/src/test/resources/tck/docker-compose.yaml @@ -3,9 +3,14 @@ # The image serves flagd itself and the "launchpad" control API on 8080, whose endpoints this # TCK's control API contract was derived from. It does not yet serve the whole canonical flag # set: v3.8.0 has none of large-integer-flag, huge-integer-flag or integral-float-flag. Only -# large-integer-flag is reached by a scenario that runs here -- the other two sit behind -# @large-integers and @numeric-coercion, which this provider does not declare -- so one untagged -# scenario fails FLAG_NOT_FOUND until open-feature/flagd-testbed#392 lands. Bump the tag then. +# large-integer-flag is reached by scenarios that run here -- the other two sit behind +# @large-integers and @numeric-coercion, which this provider does not declare -- and it is +# reached twice: the untagged precision scenario, which gets the code default instead of +# 2147483647, and the @variants row that asks for its "max-int32" variant and is answered with +# none. Both fail until open-feature/flagd-testbed#392 lands. Bump the tag then. +# +# targeting-key-flag is served, which is why @targeting needs no bump: its three scenarios pass +# on this image in both modes. # # Note there are no host port bindings. The TCK requires dynamically mapped ports and discovers # them after startup — a pinned host port would make the suite unrunnable in parallel and would From 550106087d78392dc9ecb5007cce02d9625bfed2 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sat, 12 Sep 2026 12:36:37 +0200 Subject: [PATCH 07/25] test(flagd): declare @disabled-flags, on the evidence of both resolvers MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The tag arrives declared, because capabilities() is spelt "everything declarable except", and it is right to arrive that way here — but that is not the same as having been checked, so it was run. Both modes now execute 56 scenarios and pass 49 of them, with the same five skips and the same two testbed failures as before: all four rows of the new outline pass in RPC and in-process alike. The provider substitutes the caller's default for a flag whose state is DISABLED and reports no error code. Nothing was needed upstream. The four disabled-* flags are flagd-testbed's own, from flags/disabled-flags.json, present in the v3.8.0 image already pinned and combined into the served set by the launchpad; the canonical definition took the testbed's names and values rather than inventing its own, as it did for the falsy flags. Recorded next to the image tag, where the flags the testbed does *not* serve are recorded. Worth stating rather than assuming, because @disabled-flags is gated on architecture rather than on quality, and flagd is on the side of that line that can hold it: the in-process resolver evaluates locally, and the RPC resolver still decides locally what to do with a response carrying no value. Signed-off-by: Simon Schrottner --- .../flagd/e2e/AbstractFlagdTckTest.java | 18 ++++++++++++++++++ .../src/test/resources/tck/docker-compose.yaml | 5 ++++- 2 files changed, 22 insertions(+), 1 deletion(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index 520056c0a4..7d1d44b693 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -173,6 +173,24 @@ public FeatureProvider createUnavailableProvider() { * resolve {@code targeting-key-flag} through flagd's own rule evaluation and all three pass in * both modes on the image already pinned, so nothing about it needed a testbed bump. * + *

{@link Capability#DISABLED_FLAGS} is declared, and measured rather than assumed. Both + * resolvers substitute the caller's default for a flag whose state is {@code DISABLED} and report + * no error code, so all four rows of that outline pass in both modes — 56 scenarios, 49 passing, + * five skipped for the withheld tags above and the two testbed failures already described. + * Nothing needed bumping for it either: {@code disabled-boolean-flag}, + * {@code disabled-string-flag}, {@code disabled-integer-flag} and {@code disabled-float-flag} + * are flagd-testbed's own, from {@code flags/disabled-flags.json}, which the image already + * carries at the v3.8.0 pinned below and which the launchpad combines into the set it serves. The + * canonical definition took the testbed's names and values rather than inventing its own, exactly + * as it did for the falsy flags. + * + *

That the capability holds here is worth stating rather than assuming, because the tag is + * gated on architecture rather than on quality and flagd sits on the right side of that line + * twice over: the in-process resolver evaluates the ruleset locally, and the RPC resolver still + * decides locally what to do with a response that carries no value. A provider whose backend + * decides — one speaking OFREP — cannot hold it at all, which is the comparison the tag exists to + * make legible. + * *

{@link Capability#declarableExcept} rather than {@code EnumSet.complementOf}, which is what * this used to be. The complement of one capability is every other enum constant, * including {@code @caching} — a reserved tag no scenario carries — so declaring the complement diff --git a/providers/flagd/src/test/resources/tck/docker-compose.yaml b/providers/flagd/src/test/resources/tck/docker-compose.yaml index 8feeec6c55..65837206fe 100644 --- a/providers/flagd/src/test/resources/tck/docker-compose.yaml +++ b/providers/flagd/src/test/resources/tck/docker-compose.yaml @@ -10,7 +10,10 @@ # none. Both fail until open-feature/flagd-testbed#392 lands. Bump the tag then. # # targeting-key-flag is served, which is why @targeting needs no bump: its three scenarios pass -# on this image in both modes. +# on this image in both modes. The same goes for the four disabled-* flags behind +# @disabled-flags: they are this image's own, from its flags/disabled-flags.json, which the +# launchpad combines into the served set -- the canonical definition took the testbed's names and +# values rather than inventing its own. # # Note there are no host port bindings. The TCK requires dynamically mapped ports and discovers # them after startup — a pinned host port would make the suite unrunnable in parallel and would From f7245beccee45e66bfe537260f205371e3ff611a Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sat, 12 Sep 2026 14:26:15 +0200 Subject: [PATCH 08/25] test(flagd): follow the provider-tck -> tck rename Import-path and coordinate churn only: the artifact is dev.openfeature.contrib.tools:tck, the version range starts at 0.1.0, and the four imports come from dev.openfeature.contrib.tools.tck. No behavioural change, and no change to what is declared or withheld. Signed-off-by: Simon Schrottner --- providers/flagd/pom.xml | 8 ++++---- .../providers/flagd/e2e/AbstractFlagdTckTest.java | 10 +++++----- 2 files changed, 9 insertions(+), 9 deletions(-) diff --git a/providers/flagd/pom.xml b/providers/flagd/pom.xml index 15001f4bbc..fac28ba676 100644 --- a/providers/flagd/pom.xml +++ b/providers/flagd/pom.xml @@ -22,8 +22,8 @@ 1.2.28 [2.0.0,3.0.0) - - [0.0.1,) + + [0.1.0,) flagd @@ -107,8 +107,8 @@ --> dev.openfeature.contrib.tools - provider-tck - ${provider-tck.version} + tck + ${tck.version} test diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index 7d1d44b693..eb3034633a 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -3,10 +3,10 @@ import dev.openfeature.contrib.providers.flagd.Config; import dev.openfeature.contrib.providers.flagd.FlagdOptions; import dev.openfeature.contrib.providers.flagd.FlagdProvider; -import dev.openfeature.contrib.tools.providertck.BackendEndpoint; -import dev.openfeature.contrib.tools.providertck.Capability; -import dev.openfeature.contrib.tools.providertck.ContainerizedProviderTckTest; -import dev.openfeature.contrib.tools.providertck.KnownDeviation; +import dev.openfeature.contrib.tools.tck.BackendEndpoint; +import dev.openfeature.contrib.tools.tck.Capability; +import dev.openfeature.contrib.tools.tck.ContainerizedProviderTckTest; +import dev.openfeature.contrib.tools.tck.KnownDeviation; import dev.openfeature.sdk.FeatureProvider; import java.io.File; import java.util.Collections; @@ -159,7 +159,7 @@ public FeatureProvider createUnavailableProvider() { *

{@link Capability#LARGE_INTEGERS} is withheld, as it is by every Java provider. It asks for * 2^53 − 1 and {@code Client.getIntegerDetails} is a 32-bit {@code Integer} with no room for it, * so the limit is the SDK's rather than flagd's — Appendix F is where that is recorded, and it is - * not a {@link dev.openfeature.contrib.tools.providertck.KnownDeviation}, because flagd is not at + * not a {@link dev.openfeature.contrib.tools.tck.KnownDeviation}, because flagd is not at * fault for a value the accessor cannot carry. Its one scenario is reported as skipped for an * undeclared capability, like any other. * From 6d41ecd2fd9a9645aaaf5797a145a8436d444fad Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sat, 12 Sep 2026 15:21:14 +0200 Subject: [PATCH 09/25] test(flagd): give the in-process resolver a deadline its sync can meet The first two in-process scenarios failed reproducibly on a slower host with `Initialization timeout exceeded; did not complete within the 10000 ms deadline` out of FlagdProviderSyncResources.waitForInitialization. The in-process resolver syncs the whole ruleset before it reports ready, and the first scenario pays for a cold container on top of that; flagd doubles the configured deadline, so 5000 gave it 10s and that was not enough. 15000 gives it 30s and both modes are clean apart from the two flagd-testbed gaps already documented here -- 56 scenarios, 49 passing, 5 skipped, 2 failing, in RPC and in-process alike. Diagnosed rather than guessed, because this looked at first like fallout from the base dropping its 50ms post-command settle. It is not: the settle was restored locally at 50ms and at 3000ms and fixed nothing, and the failure is present on flagd-testbed v3.10.1 as well as on the pinned v3.8.0. What it covers is the provider's own initialisation, which belongs in a bound the scenario can see rather than in a sleep after an unrelated control call. UNAVAILABLE_DEADLINE_MS is untouched, so the initialisation-failure scenarios still assert promptness. Signed-off-by: Simon Schrottner --- .../flagd/e2e/AbstractFlagdTckTest.java | 21 +++++++++++++++---- 1 file changed, 17 insertions(+), 4 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index eb3034633a..3bb2ae56ac 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -57,11 +57,24 @@ abstract class AbstractFlagdTckTest extends ContainerizedProviderTckTest { /** * gRPC deadline for a provider that is expected to connect. * - *

Generous on purpose. flagd derives its initialisation deadline from this value, and the - * in-process resolver must sync the entire ruleset before it reports ready — which intermittently - * takes longer than a deadline tuned for a single RPC round trip. + *

Generous on purpose, and measured. flagd derives its initialisation deadline from this + * value — doubling it — and the in-process resolver must sync the entire ruleset before it + * reports ready, which takes longer than a deadline tuned for a single RPC round trip. + * + *

At 5000 the first two in-process scenarios failed reproducibly on a slower host + * with {@code Initialization timeout exceeded; did not complete within the 10000 ms deadline} + * out of {@code FlagdProviderSyncResources.waitForInitialization}, on both flagd-testbed v3.8.0 + * and v3.10.1; at 15000 the suite is clean apart from the two testbed gaps described above. The + * first scenario pays for a cold container as well as for the sync, which is why it is the first + * two rather than all of them. + * + *

Worth being explicit that this is not a post-command settle in disguise. + * A pause after the control call was tried at 50ms and at 3000ms and fixed nothing — the wait + * this covers is the provider's own initialisation, which is bounded here where the scenario can + * see it, rather than slept through where it cannot. {@link #UNAVAILABLE_DEADLINE_MS} stays + * short so the promptness assertions still mean something. */ - private static final int CONNECTED_DEADLINE_MS = 5000; + private static final int CONNECTED_DEADLINE_MS = 15000; /** * gRPC deadline for a provider pointed at a dead port. From ea4259868708da1423cb6fe444e5f0da56747e20 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sat, 12 Sep 2026 17:36:33 +0200 Subject: [PATCH 10/25] test(flagd): declare numeric coercion and let its failure show Switches the @numeric-coercion deviation from withheld-and-skipped to declared-and-failing, which is the shape the TCK's settled guidance prefers, and records why paying its cost is the honest report. The guidance says withdrawing a capability in order to turn a failing scenario into a skip is the failure mode the field exists to prevent, and that is exactly what the old shape did here. Measured on the pinned testbed, in both modes: of the tag's three scenarios, "An integer requested as a float is widened without loss" passes. flagd therefore does coerce, and gets the narrowing direction wrong - a skip cannot distinguish that from "flagd declines to coerce", and only the second reading was available before. The argument for the old shape was real and is recorded rather than dropped: declaring the tag also fails "An integral float requested as an integer is coerced without loss", because integral-float-flag is absent from flagd-testbed v3.8.0. That cost is accepted because it is not a new kind of cost - this adoption already carries two failures from the same missing flags and records them plainly - and because the alternative hides a real defect behind a stack gap. Measured result, both modes: 56 scenarios, 2 skipped (@reinitialization, @large-integers), 4 failing - one provider defect and three testbed gaps. Also records something the previous pass reported as fixed and which does not hold on a loaded host: the first scenario of errors.feature still errors in in-process mode with an initialisation timeout against the doubled 30000 ms deadline. Reproduced three times, and reproduced identically with @numeric-coercion withheld, so it is not a consequence of this change. Thirty seconds is not a plausible sync time for this ruleset and only the mode that must establish a sync stream after the first POST /start is affected, so it reads as stack-side readiness - the class of defect open-feature/flagd-testbed#394 closes. The deadline stays at 15000 rather than being raised again: a suite that sleeps instead of holding the control API to its promise stops being able to detect when the promise breaks. Signed-off-by: Simon Schrottner --- .../flagd/e2e/AbstractFlagdTckTest.java | 126 +++++++++++++----- .../test/resources/tck/docker-compose.yaml | 19 ++- 2 files changed, 102 insertions(+), 43 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index 3bb2ae56ac..c02b54f27c 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -26,12 +26,24 @@ * *

The testbed does not yet serve the whole canonical flag set. Three of the * flags the suite's assets added are absent from {@code flagd-testbed} v3.8.0: - * {@code large-integer-flag}, {@code huge-integer-flag} and {@code integral-float-flag}. Only the - * first is reached — {@code huge-integer-flag} is asked for solely under {@code @large-integers} and - * {@code integral-float-flag} solely under {@code @numeric-coercion}, both withheld below — so - * exactly one untagged scenario, the 32-bit - * precision one, fails with {@code FLAG_NOT_FOUND} in both modes until - * open-feature/flagd-testbed#392 lands and the tag here is bumped. + * {@code large-integer-flag}, {@code huge-integer-flag} and {@code integral-float-flag}. Two of the + * three are reached — {@code huge-integer-flag} is asked for solely under {@code @large-integers}, + * which is withheld below for a reason of its own — so three scenarios fail with + * {@code FLAG_NOT_FOUND} in both modes until open-feature/flagd-testbed#392 lands and the tag here + * is bumped: + * + *

    + *
  • the untagged 32-bit precision scenario, which gets the code default instead of + * {@code 2147483647}; + *
  • the {@code @variants} row asking for {@code large-integer-flag}'s {@code max-int32}, which + * is answered with no variant; + *
  • the {@code @numeric-coercion} scenario "An integral float requested as an integer is coerced + * without loss", which asks for {@code integral-float-flag}. + *
+ * + *

The third of those is new here, and it is the cost of declaring {@code @numeric-coercion} + * rather than withholding it — see {@link #knownDeviations()}, which explains why paying it is the + * honest report. * *

The three falsy flags used to fail the same way and no longer do. The testbed's * {@code zero-flags.json} already served {@code boolean-zero-flag}, {@code integer-zero-flag} and @@ -64,15 +76,29 @@ abstract class AbstractFlagdTckTest extends ContainerizedProviderTckTest { *

At 5000 the first two in-process scenarios failed reproducibly on a slower host * with {@code Initialization timeout exceeded; did not complete within the 10000 ms deadline} * out of {@code FlagdProviderSyncResources.waitForInitialization}, on both flagd-testbed v3.8.0 - * and v3.10.1; at 15000 the suite is clean apart from the two testbed gaps described above. The - * first scenario pays for a cold container as well as for the sync, which is why it is the first - * two rather than all of them. + * and v3.10.1. Raising it to 15000 cleared that. The first scenario pays for a cold container as + * well as for the sync, which is why it was the first two rather than all of them. * *

Worth being explicit that this is not a post-command settle in disguise. * A pause after the control call was tried at 50ms and at 3000ms and fixed nothing — the wait * this covers is the provider's own initialisation, which is bounded here where the scenario can * see it, rather than slept through where it cannot. {@link #UNAVAILABLE_DEADLINE_MS} stays * short so the promptness assertions still mean something. + * + *

Not fully solved, and the bound is not the thing to keep raising. On a + * loaded Docker-in-WSL host the first scenario of {@code errors.feature} still errors in + * in-process mode, with the same message against the doubled 30000 ms deadline after some 53 + * seconds of wall clock. Measured three times in a row, and — importantly — it reproduces with + * {@code @numeric-coercion} withheld exactly as it does with it declared, so it is not a + * consequence of what this suite declares. RPC mode never shows it. + * + *

That shape is stack-side readiness rather than provider slowness: a small ruleset does not + * take thirty seconds to sync, and only the mode that has to establish a sync stream and receive + * the whole ruleset after the first {@code POST /start} is affected. It is the class of defect + * open-feature/flagd-testbed#394 exists to close — a control endpoint returning before the + * backend is serving — and the TCK's own rule applies: a suite that sleeps instead of holding the + * control API to its promise stops being able to detect when the promise breaks. So the bound + * stays at 15000 and this is recorded rather than covered. */ private static final int CONNECTED_DEADLINE_MS = 15000; @@ -122,18 +148,28 @@ public FeatureProvider createUnavailableProvider() { /** * {@inheritDoc} * - *

Everything declarable except {@link Capability#NUMERIC_COERCION} and - * {@link Capability#REINITIALIZATION}. The two omissions are different in kind: the first is a - * defect and is declared as one below, the second is a choice the specification offers. Evaluating - * {@code float-flag} (0.5) through the integer API returns {@code 0} with no error code - * rather than {@code TYPE_MISMATCH} with the code default — the value is silently truncated. - * Coercion as such is permitted, and the capability says so: the rule is that a lossless - * coercion must succeed and a lossy one must fail. It is the lossy case being accepted that is a - * defect to fix, not a design choice; this override should be deleted once it is. + *

Everything declarable except {@link Capability#REINITIALIZATION} and + * {@link Capability#LARGE_INTEGERS}, both of which are facts about the provider or the SDK + * rather than defects, and both explained below. * - *

Declared here rather than per mode because both resolvers behave identically, which places - * the defect in the shared provider layer rather than in either transport. Every other - * capability, including the full non-numeric type-mismatch matrix, holds in both modes. + *

{@link Capability#NUMERIC_COERCION} is declared even though one of its scenarios + * fails, and that is deliberate — see {@link #knownDeviations()} for the reasoning. + * Evaluating {@code float-flag} (0.5) through the integer API returns {@code 0} with no + * error code rather than {@code TYPE_MISMATCH} with the code default, so the fractional part is + * discarded silently. Coercion as such is permitted, and the capability says so: the rule is + * that a lossless coercion must succeed and a lossy one must fail. It is the lossy case being + * accepted that is a defect to fix. + * + *

Measured rather than assumed, in both modes: of the tag's three scenarios, "An integer + * requested as a float is widened without loss" passes — which is the fact that settles + * the shape of the report, because a provider that performs the coercion and gets one direction + * wrong is not a provider that declines to coerce. The lossy scenario fails, and the remaining + * lossless one fails only because {@code integral-float-flag} is absent from the pinned testbed + * image. + * + *

flagd's numeric behaviour is identical across both resolvers, which places the defect in + * the shared provider layer rather than in either transport. Every other capability, including + * the full non-numeric type-mismatch matrix, holds in both modes. * *

That includes {@link Capability#LIFECYCLE}, and legitimately so: flagd reaches its backend * during initialisation in both modes — an RPC round trip, or a full ruleset sync — so the @@ -188,9 +224,11 @@ public FeatureProvider createUnavailableProvider() { * *

{@link Capability#DISABLED_FLAGS} is declared, and measured rather than assumed. Both * resolvers substitute the caller's default for a flag whose state is {@code DISABLED} and report - * no error code, so all four rows of that outline pass in both modes — 56 scenarios, 49 passing, - * five skipped for the withheld tags above and the two testbed failures already described. - * Nothing needed bumping for it either: {@code disabled-boolean-flag}, + * no error code, so all four rows of that outline pass in both modes. Measured on the pinned + * image: RPC is 56 scenarios, 50 passing, two skipped for the withheld tags above and four + * failing — the one real defect plus the three testbed gaps already described. In-process is the + * same four failures plus the cold-start initialisation error recorded on + * {@link #CONNECTED_DEADLINE_MS}, so 49 passing. Nothing needed bumping for it either: {@code disabled-boolean-flag}, * {@code disabled-string-flag}, {@code disabled-integer-flag} and {@code disabled-float-flag} * are flagd-testbed's own, from {@code flags/disabled-flags.json}, which the image already * carries at the v3.8.0 pinned below and which the launchpad combines into the set it serves. The @@ -214,26 +252,35 @@ public FeatureProvider createUnavailableProvider() { */ @Override public Set capabilities() { - return Capability.declarableExcept( - Capability.NUMERIC_COERCION, Capability.REINITIALIZATION, Capability.LARGE_INTEGERS); + return Capability.declarableExcept(Capability.REINITIALIZATION, Capability.LARGE_INTEGERS); } /** * {@inheritDoc} * - *

The withheld {@link Capability#NUMERIC_COERCION} is a defect, not a limitation, and - * something has to say so. In the results the two are indistinguishable: the scenario is skipped - * either way, and the declaration explains only that the capability was not claimed, - * never whether flagd chose not to claim it. A consumer comparing providers would otherwise read - * this exactly as it reads a provider with no streaming transport declining - * {@code @configuration-change}, which is a decision rather than a bug. + *

One entry, for {@link Capability#NUMERIC_COERCION}, and it takes the declared and + * failing shape rather than the withheld-and-skipped one. That is the shape the TCK's + * guidance prefers, and the reason it prefers it is exactly this case: flagd does + * attempt the coercion — the widening scenario passes — and gets the narrowing direction wrong, + * so withdrawing the capability would turn a real failure into a skip, which is the failure mode + * the field exists to prevent. The failure stays visible in the results and this entry says it is + * known and why. + * + *

An earlier version of this file withheld the tag instead, on the argument that the + * capability requires all three of its scenarios and one of them asks for + * {@code integral-float-flag}, which the pinned testbed image does not serve — so declaring it + * buys one honest failure and one that is the stack's fault. That cost is real and it is + * accepted, for two reasons. It is not a new kind of cost: this branch already carries two + * failures caused by the same missing flags and records them plainly rather than hiding them + * behind a withheld capability. And the alternative is worse, because a skip cannot distinguish + * "flagd declines to coerce" from "flagd coerces and gets it wrong", and only the second is true. * *

Tracked against flagd's numeric coercion ADR, which is where the rule this deviates from is * settled: coercion is permitted when it is lossless and must fail with {@code TYPE_MISMATCH} * only when information would be lost. The summary says which half is broken, because "flagd - * coerces numbers" on its own reads as a description of intended behaviour. Delete the entry — - * and the {@code capabilities()} override above — once the lossy case reports - * {@code TYPE_MISMATCH}. + * coerces numbers" on its own reads as a description of intended behaviour. Delete the entry once + * the lossy case reports {@code TYPE_MISMATCH}; the capability needs no change then, which is + * another small argument for this shape. */ @Override public List knownDeviations() { @@ -243,9 +290,14 @@ public List knownDeviations() { "The lossy half of the coercion rule is not enforced: evaluating float-flag (0.5) " + "through the integer API returns 0 with no error code, rather than " + "TYPE_MISMATCH with the code default, so the fractional part is discarded " - + "silently. Lossless coercion is permitted and is not the defect. Both " - + "resolvers behave identically, which places it in the shared provider layer " - + "rather than in either transport.")); + + "silently. Lossless coercion is permitted and is not the defect -- flagd " + + "does widen an integer to a float correctly, which is why the capability is " + + "declared and the scenario left to fail rather than the capability " + + "withheld. Both resolvers behave identically, which places it in the shared " + + "provider layer rather than in either transport. The tag's third scenario " + + "also fails, but for an unrelated reason that is not flagd's: " + + "integral-float-flag is absent from the pinned flagd-testbed image, " + + "open-feature/flagd-testbed#392.")); } private FlagdOptions.FlagdOptionsBuilder baseOptions() { diff --git a/providers/flagd/src/test/resources/tck/docker-compose.yaml b/providers/flagd/src/test/resources/tck/docker-compose.yaml index 65837206fe..b935014143 100644 --- a/providers/flagd/src/test/resources/tck/docker-compose.yaml +++ b/providers/flagd/src/test/resources/tck/docker-compose.yaml @@ -2,12 +2,19 @@ # # The image serves flagd itself and the "launchpad" control API on 8080, whose endpoints this # TCK's control API contract was derived from. It does not yet serve the whole canonical flag -# set: v3.8.0 has none of large-integer-flag, huge-integer-flag or integral-float-flag. Only -# large-integer-flag is reached by scenarios that run here -- the other two sit behind -# @large-integers and @numeric-coercion, which this provider does not declare -- and it is -# reached twice: the untagged precision scenario, which gets the code default instead of -# 2147483647, and the @variants row that asks for its "max-int32" variant and is answered with -# none. Both fail until open-feature/flagd-testbed#392 lands. Bump the tag then. +# set: v3.8.0 has none of large-integer-flag, huge-integer-flag or integral-float-flag. Two of +# the three are reached by scenarios that run here -- huge-integer-flag sits behind +# @large-integers, which this provider does not declare -- and between them they fail three +# scenarios, all until open-feature/flagd-testbed#392 lands. Bump the tag then. +# +# * large-integer-flag: the untagged precision scenario, which gets the code default instead +# of 2147483647, and the @variants row asking for its "max-int32" variant, answered with +# none. +# * integral-float-flag: the @numeric-coercion scenario "An integral float requested as an +# integer is coerced without loss". This provider declares @numeric-coercion because flagd +# does coerce and gets the lossy direction wrong, so the tag's failure is worth seeing -- +# see AbstractFlagdTckTest.knownDeviations(). This third failure is the price of that, and +# it is the same missing-flag problem as the two above rather than a new one. # # targeting-key-flag is served, which is why @targeting needs no bump: its three scenarios pass # on this image in both modes. The same goes for the four disabled-* flags behind From 96f9492892ddf268e83922ca802d9a0334ae5c7d Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sat, 12 Sep 2026 17:36:47 +0200 Subject: [PATCH 11/25] docs(flagd): say that the TCK suites are Docker-gated and hand-run The two TCK suites have been excluded from the default build since they were added, via **/e2e/*.java, and nothing said so. An exclusion nobody writes down is indistinguishable from an oversight - which is how providers/ofrep came to run a Docker-dependent suite in its default build unnoticed, the same mistake in the other direction. So the README now states the policy and its reasoning: the suites are Docker-gated, excluded from every job that exists, and run by hand by a maintainer before merging a change to resolution, event or lifecycle behaviour, with the result quoted in the pull request. A scheduled or path-filtered workflow was considered and declined. It also gives the two commands and says which failures are expected, so that a reader can tell a regression from the recorded testbed gaps and the recorded coercion defect. Signed-off-by: Simon Schrottner --- providers/flagd/README.md | 35 +++++++++++++++++++++++++++++++++++ 1 file changed, 35 insertions(+) diff --git a/providers/flagd/README.md b/providers/flagd/README.md index 3b88352791..422c7e9ff0 100644 --- a/providers/flagd/README.md +++ b/providers/flagd/README.md @@ -358,3 +358,38 @@ FlagdOptions options = FlagdOptions.builder() .resolverType(Config.Resolver.IN_PROCESS) .build(); ``` + +## Provider conformance (TCK) + +This provider adopts the [OpenFeature Provider TCK](../../tools/tck/README.md), once per resolver: +`FlagdRpcTckTest` and `FlagdInProcessTckTest`, both over the shared +`AbstractFlagdTckTest`. Read that class before changing either — it records which capabilities are +declared, which are withheld and why, and every known deviation, each against measured behaviour +rather than assumption. + +**The two suites are Docker-gated and excluded from the default build.** They live under +`src/test/java/.../e2e/`, which `**/e2e/*.java` keeps out of +`mvn verify`, so a machine or CI job without a Docker daemon is never asked to start a Compose +stack. That is a deliberate policy and not an omission: a default build that needs Docker fails in a +way that reads as a broken provider rather than as a missing prerequisite. + +The consequence is that **no CI job runs them**, so a maintainer runs them by hand before merging a +change that touches the provider's resolution, event or lifecycle behaviour, and quotes the result in +the pull request. A scheduled or path-filtered workflow was considered and declined: a suite whose +red is diagnosed by whoever happens to read the notification is worse than one whose red is +diagnosed by the person who caused it. + +```bash +# both resolvers +mvn -pl providers/flagd -am -DtestExclusions= -Dtest='Flagd*TckTest' \ + -Dsurefire.failIfNoSpecifiedTests=false test + +# one resolver +mvn -pl providers/flagd -am -DtestExclusions= -Dtest=FlagdInProcessTckTest \ + -Dsurefire.failIfNoSpecifiedTests=false test +``` + +Both suites are currently **expected to fail**, and the expected failures are enumerated in +`AbstractFlagdTckTest`: three come from flags that the pinned `flagd-testbed` image does not serve +(open-feature/flagd-testbed#392) and one is the real numeric-coercion defect, declared and left +visible rather than skipped (open-feature/flagd#1996). Anything else is a regression. From 417f99464e4b1d24139fffcb9979d3be92765ca2 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sat, 12 Sep 2026 17:47:12 +0200 Subject: [PATCH 12/25] test(flagd): keep the TCK suites out of CI, which the e2e profile was not doing Narrows the e2e profile's testExclusions from empty to **/e2e/*TckTest.java, so the exclusion the default build applies is no longer undone in the one job that matters. A previous pass reported that the TCK suites "are excluded from every job that exists". That was wrong, in two steps that have to be read together: providers/flagd's e2e profile sets , clearing the exclusion, and ci.yml's `main` job activates that profile on every push. Verified rather than re-read: $ mvn -Pe2e -pl providers/flagd help:evaluate -Dexpression=testExclusions (empty) $ mvn -pl providers/flagd help:evaluate -Dexpression=testExclusions **/e2e/*.java So FlagdRpcTckTest and FlagdInProcessTckTest were running in CI, on a runner that does have a Docker daemon, and they are expected to fail - three testbed gaps and one recorded coercion defect. Every unrelated pull request touching this module would have gone red for a reason that has nothing to do with it. Narrowing rather than clearing is what keeps both halves true: the legacy Run*Test suites over the test-harness submodule still run under -Pe2e exactly as they do on main, and only the two TCK suites stay out. The pattern matches AbstractFlagdTckTest.java as well, which is harmless - it is abstract and surefire would not select it - and testExclusions only filters what runs, never what compiles. Not runnable locally as a cross-check: providers/flagd/test-harness is an uninitialised submodule on this machine, so the legacy suites cannot be executed here to prove they still get selected. The pattern is checked against the directory's file names instead, which is deterministic: Run{File,InProcess,Rpc}Test do not end in TckTest. The README says all of this too, in the section added with the exclusion policy, because the narrowing is the half a reader would not guess from the POM alone. Signed-off-by: Simon Schrottner --- providers/flagd/README.md | 8 ++++++++ providers/flagd/pom.xml | 16 ++++++++++++++-- 2 files changed, 22 insertions(+), 2 deletions(-) diff --git a/providers/flagd/README.md b/providers/flagd/README.md index 422c7e9ff0..2cd52c6a79 100644 --- a/providers/flagd/README.md +++ b/providers/flagd/README.md @@ -373,6 +373,14 @@ rather than assumption. stack. That is a deliberate policy and not an omission: a default build that needs Docker fails in a way that reads as a broken provider rather than as a missing prerequisite. +**The `e2e` profile narrows that exclusion rather than clearing it**, to +`**/e2e/*TckTest.java`. This is the part that is easy to get wrong, so it is worth stating: the +`e2e` profile exists for the legacy `Run*Test` suites over the `test-harness` submodule, and +`ci.yml`'s `main` job activates it on every push. A profile that cleared the exclusion outright +would therefore run the TCK suites in CI, on a runner that does have a Docker daemon — and they are +expected to fail, so every unrelated pull request would go red for a reason that has nothing to do +with it. Narrowing keeps the legacy suites running exactly as before and the TCK suites out. + The consequence is that **no CI job runs them**, so a maintainer runs them by hand before merging a change that touches the provider's resolution, event or lifecycle behaviour, and quotes the result in the pull request. A scheduled or path-filtered workflow was considered and declined: a suite whose diff --git a/providers/flagd/pom.xml b/providers/flagd/pom.xml index fac28ba676..4ea4d40f19 100644 --- a/providers/flagd/pom.xml +++ b/providers/flagd/pom.xml @@ -263,8 +263,20 @@ e2e - - + + **/e2e/*TckTest.java From 4eb74ca7dea10919451eb28e678f474a3ed2334e Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sat, 12 Sep 2026 23:35:14 +0200 Subject: [PATCH 13/25] docs(flagd): point at Appendix F for the CI-exclusion reasoning Appendix F now carries "Running the suite in CI", promoted there because the same reasoning restated in four adoption READMEs is where it drifted. So this section keeps the mechanism -- the testExclusions property, what the e2e profile narrows it to and why the narrowing rather than the clearing, that the exclusion is Surefire's and not the compiler's, and the help:evaluate command that resolves it -- and links to the appendix instead of paraphrasing the argument. The resolved values are checked, not read: mvn -Pe2e -pl providers/flagd help:evaluate -Dexpression=testExclusions -> **/e2e/*TckTest.java mvn -pl providers/flagd help:evaluate -Dexpression=testExclusions -> **/e2e/*.java The locally decided part stays stated here rather than deferred: a scheduled or path-filtered workflow was considered and declined, and a maintainer quotes a hand-run result in the pull request. Signed-off-by: Simon Schrottner --- providers/flagd/README.md | 32 ++++++++++++++++++++------------ 1 file changed, 20 insertions(+), 12 deletions(-) diff --git a/providers/flagd/README.md b/providers/flagd/README.md index 2cd52c6a79..bdffb4bb4a 100644 --- a/providers/flagd/README.md +++ b/providers/flagd/README.md @@ -368,18 +368,26 @@ declared, which are withheld and why, and every known deviation, each against me rather than assumption. **The two suites are Docker-gated and excluded from the default build.** They live under -`src/test/java/.../e2e/`, which `**/e2e/*.java` keeps out of -`mvn verify`, so a machine or CI job without a Docker daemon is never asked to start a Compose -stack. That is a deliberate policy and not an omission: a default build that needs Docker fails in a -way that reads as a broken provider rather than as a missing prerequisite. - -**The `e2e` profile narrows that exclusion rather than clearing it**, to -`**/e2e/*TckTest.java`. This is the part that is easy to get wrong, so it is worth stating: the -`e2e` profile exists for the legacy `Run*Test` suites over the `test-harness` submodule, and -`ci.yml`'s `main` job activates it on every push. A profile that cleared the exclusion outright -would therefore run the TCK suites in CI, on a runner that does have a Docker daemon — and they are -expected to fail, so every unrelated pull request would go red for a reason that has nothing to do -with it. Narrowing keeps the legacy suites running exactly as before and the TCK suites out. +`src/test/java/.../e2e/`, which `**/e2e/*.java` in this module's +POM keeps out of `mvn verify`. Why an adoption suite is excluded rather than gating is written down +once for all four languages in +[Appendix F: Running the suite in CI](https://github.com/open-feature/spec/blob/main/specification/appendix-f-provider-conformance.md#running-the-suite-in-ci); +this section is only what that means in this module. + +**The `e2e` profile narrows that exclusion rather than clearing it**, to `**/e2e/*TckTest.java`. +The profile exists for the legacy `Run*Test` suites over the `test-harness` submodule, and +`ci.yml`'s `main` job activates it on every push — so a profile that cleared the exclusion outright, +which is what this one used to do, ran the TCK suites in CI on a runner that does have a Docker +daemon. They are expected to fail, so every unrelated pull request went red for a reason that had +nothing to do with it. That is the first of the two mistakes the appendix names, found here by +resolving the property rather than reading the POM: + +```bash +mvn -Pe2e -pl providers/flagd help:evaluate -Dexpression=testExclusions -DforceStdout +``` + +Narrowing keeps the legacy suites running exactly as before and the TCK suites out. The exclusion is +Surefire's, not the compiler's, so both suites still build against the harness in every job. The consequence is that **no CI job runs them**, so a maintainer runs them by hand before merging a change that touches the provider's resolution, event or lifecycle behaviour, and quotes the result in From c9cf6b495672d7efd543ffcfc36b9867142fb3b1 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sun, 13 Sep 2026 10:15:00 +0200 Subject: [PATCH 14/25] test(flagd): record that both resolvers report the standard reasons The TCK gained @standard-reasons, and this suite declares every declarable capability except @reinitialization and @large-integers -- so it picked the new tag up by default and started running reason.feature without anyone deciding that it should. That is the right answer here, but it was measured before it was written down rather than after. All nine scenarios pass in both modes, including the two that compose with @targeting and @disabled-flags: STATIC for the rule-less flags, TARGETING_MATCH and DEFAULT either side of targeting-key-flag's rule, DISABLED for a disabled flag, and ERROR beside FLAG_NOT_FOUND and TYPE_MISMATCH. So the claim the tag makes -- the standard vocabulary with the standard meanings -- holds for the RPC resolver and the in-process one alike. Each mode is now 65 scenarios, 59 passing, 2 skipped and 4 failing, up from 56 and 50. The four failures are the same four as before and none of them is new: the lossy numeric coercion that flagd#1996 tracks, and the three flags flagd-testbed v3.8.0 does not serve. The cold-start initialisation error recorded on CONNECTED_DEADLINE_MS did not reproduce in this run; it is intermittent and host-dependent, so the note stays and the count says which run it comes from. Signed-off-by: Simon Schrottner --- .../flagd/e2e/AbstractFlagdTckTest.java | 19 +++++++++++++++---- 1 file changed, 15 insertions(+), 4 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index c02b54f27c..6557001fff 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -225,10 +225,11 @@ public FeatureProvider createUnavailableProvider() { *

{@link Capability#DISABLED_FLAGS} is declared, and measured rather than assumed. Both * resolvers substitute the caller's default for a flag whose state is {@code DISABLED} and report * no error code, so all four rows of that outline pass in both modes. Measured on the pinned - * image: RPC is 56 scenarios, 50 passing, two skipped for the withheld tags above and four - * failing — the one real defect plus the three testbed gaps already described. In-process is the - * same four failures plus the cold-start initialisation error recorded on - * {@link #CONNECTED_DEADLINE_MS}, so 49 passing. Nothing needed bumping for it either: {@code disabled-boolean-flag}, + * image: 65 scenarios in each mode, 59 passing, two skipped for the withheld tags above and four + * failing — the one real defect plus the three testbed gaps already described. The cold-start + * initialisation error recorded on {@link #CONNECTED_DEADLINE_MS} did not reproduce in the run + * these numbers come from; it is intermittent and host-dependent, and when it appears in-process + * it costs one further scenario. Nothing needed bumping for it either: {@code disabled-boolean-flag}, * {@code disabled-string-flag}, {@code disabled-integer-flag} and {@code disabled-float-flag} * are flagd-testbed's own, from {@code flags/disabled-flags.json}, which the image already * carries at the v3.8.0 pinned below and which the launchpad combines into the set it serves. The @@ -242,6 +243,16 @@ public FeatureProvider createUnavailableProvider() { * decides — one speaking OFREP — cannot hold it at all, which is the comparison the tag exists to * make legible. * + *

{@link Capability#STANDARD_REASONS} is declared, and it arrived here by the + * {@code declarableExcept} default rather than by a decision — which is exactly why it was + * measured before this paragraph was written. All nine scenarios of {@code reason.feature} pass + * in both modes, including the two that compose with {@link Capability#TARGETING} and + * {@link Capability#DISABLED_FLAGS}: flagd reports {@code STATIC} for the rule-less flags, + * {@code TARGETING_MATCH} and {@code DEFAULT} either side of {@code targeting-key-flag}'s rule, + * {@code DISABLED} for a disabled flag, and {@code ERROR} beside {@code FLAG_NOT_FOUND} and + * {@code TYPE_MISMATCH}. So the claim the tag makes — the standard vocabulary with the standard + * meanings — holds for both resolvers, and the declaration is evidence rather than inheritance. + * *

{@link Capability#declarableExcept} rather than {@code EnumSet.complementOf}, which is what * this used to be. The complement of one capability is every other enum constant, * including {@code @caching} — a reserved tag no scenario carries — so declaring the complement From 2d6bb3517e0876793ed9e5aa15f4d9ce46fcba10 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sun, 13 Sep 2026 15:45:01 +0200 Subject: [PATCH 15/25] test(flagd): stop withholding a capability nobody in Java can hold capabilities() is now declarableExcept(REINITIALIZATION). @large-integers goes, and its going is the whole point: it was never a decision this suite took. flagd does not decline to resolve 2^53 - 1 -- Client.getIntegerDetails is a 32-bit Integer, so no Java provider can be asked, and the TCK refuses the capability centrally rather than having each adoption remember. The paragraph that used to explain it here was one of four saying the same thing about the same language. @reinitialization stays withheld, and stays this suite's call. That one is a fact about flagd: FlagdProviderSyncResources keeps isInitialized and isShutDown as separate flags and refuses initialize() when either is set, which Requirement 2.5.2 permits, so withholding is the honest report and no KnownDeviation is owed. Nothing about the SDK's accessor has any bearing on it, which is why one moved and the other did not. Measured, both resolvers, on the pinned testbed image: 65 scenarios, 59 passed, 2 skipped, 4 failed -- identical to the previous pass, because this changes why a scenario is skipped rather than whether it is. The two skips now read differently in the results, which is the observable part: Skipped: provider does not declare capability REINITIALIZATION (tag @reinitialization). Declared capabilities: [...] Skipped: the Java SDK cannot express capability LARGE_INTEGERS (tag @large-integers) - Client.getIntegerDetails takes and returns a 32-bit Integer ... not the provider under test declining A reader of the report can now tell which of the two absences describes flagd. The four failures are unchanged: flagd#1996's lossy coercion plus the three flags testbed v3.8.0 does not serve. Signed-off-by: Simon Schrottner --- .../flagd/e2e/AbstractFlagdTckTest.java | 28 +++++++++++-------- .../test/resources/tck/docker-compose.yaml | 6 ++-- 2 files changed, 20 insertions(+), 14 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index 6557001fff..1eef460df8 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -28,7 +28,8 @@ * flags the suite's assets added are absent from {@code flagd-testbed} v3.8.0: * {@code large-integer-flag}, {@code huge-integer-flag} and {@code integral-float-flag}. Two of the * three are reached — {@code huge-integer-flag} is asked for solely under {@code @large-integers}, - * which is withheld below for a reason of its own — so three scenarios fail with + * which no Java provider can declare because the SDK's integer accessor is 32 bits, so that + * scenario is skipped before the missing flag can matter — so three scenarios fail with * {@code FLAG_NOT_FOUND} in both modes until open-feature/flagd-testbed#392 lands and the tag here * is bumped: * @@ -148,9 +149,9 @@ public FeatureProvider createUnavailableProvider() { /** * {@inheritDoc} * - *

Everything declarable except {@link Capability#REINITIALIZATION} and - * {@link Capability#LARGE_INTEGERS}, both of which are facts about the provider or the SDK - * rather than defects, and both explained below. + *

Everything declarable except {@link Capability#REINITIALIZATION}, which is a fact about + * this provider rather than a defect and is explained below. Nothing else is withheld: what no + * Java provider can claim is no longer in {@link Capability#declarable()} to remove. * *

{@link Capability#NUMERIC_COERCION} is declared even though one of its scenarios * fails, and that is deliberate — see {@link #knownDeviations()} for the reasoning. @@ -205,12 +206,14 @@ public FeatureProvider createUnavailableProvider() { * provider withholds the tag for its RPC resolver; on this evidence that is a difference between * the two implementations, not a property of the transport. * - *

{@link Capability#LARGE_INTEGERS} is withheld, as it is by every Java provider. It asks for - * 2^53 − 1 and {@code Client.getIntegerDetails} is a 32-bit {@code Integer} with no room for it, - * so the limit is the SDK's rather than flagd's — Appendix F is where that is recorded, and it is - * not a {@link dev.openfeature.contrib.tools.tck.KnownDeviation}, because flagd is not at - * fault for a value the accessor cannot carry. Its one scenario is reported as skipped for an - * undeclared capability, like any other. + *

{@link Capability#LARGE_INTEGERS} is no longer named here, and that is the change rather + * than an omission. This suite used to withhold it with a paragraph explaining that + * {@code Client.getIntegerDetails} is a 32-bit {@code Integer} with no room for 2^53 − 1 — the + * same paragraph the OFREP suite and two suites inside {@code tools/tck} each carried, because + * every Java adopter was expected to know the fact and act on it. The TCK refuses the capability + * centrally now, so there is nothing for an adoption to decide and nothing here to get wrong. + * Its scenario is still skipped, with a reason naming the SDK rather than this provider, which + * is the part a report's reader needs: flagd declined nothing. * *

{@link Capability#VARIANTS} and {@link Capability#TARGETING} are both declared, and both on * evidence rather than by inheriting the "everything except" default. flagd names the variant it @@ -225,7 +228,8 @@ public FeatureProvider createUnavailableProvider() { *

{@link Capability#DISABLED_FLAGS} is declared, and measured rather than assumed. Both * resolvers substitute the caller's default for a flag whose state is {@code DISABLED} and report * no error code, so all four rows of that outline pass in both modes. Measured on the pinned - * image: 65 scenarios in each mode, 59 passing, two skipped for the withheld tags above and four + * image: 65 scenarios in each mode, 59 passing, two skipped — one for the withheld + * {@code @reinitialization}, one because the SDK cannot ask {@code @large-integers} — and four * failing — the one real defect plus the three testbed gaps already described. The cold-start * initialisation error recorded on {@link #CONNECTED_DEADLINE_MS} did not reproduce in the run * these numbers come from; it is intermittent and host-dependent, and when it appears in-process @@ -263,7 +267,7 @@ public FeatureProvider createUnavailableProvider() { */ @Override public Set capabilities() { - return Capability.declarableExcept(Capability.REINITIALIZATION, Capability.LARGE_INTEGERS); + return Capability.declarableExcept(Capability.REINITIALIZATION); } /** diff --git a/providers/flagd/src/test/resources/tck/docker-compose.yaml b/providers/flagd/src/test/resources/tck/docker-compose.yaml index b935014143..b287f8b097 100644 --- a/providers/flagd/src/test/resources/tck/docker-compose.yaml +++ b/providers/flagd/src/test/resources/tck/docker-compose.yaml @@ -4,8 +4,10 @@ # TCK's control API contract was derived from. It does not yet serve the whole canonical flag # set: v3.8.0 has none of large-integer-flag, huge-integer-flag or integral-float-flag. Two of # the three are reached by scenarios that run here -- huge-integer-flag sits behind -# @large-integers, which this provider does not declare -- and between them they fail three -# scenarios, all until open-feature/flagd-testbed#392 lands. Bump the tag then. +# @large-integers, which no Java provider can declare because Client.getIntegerDetails is a +# 32-bit Integer, so its scenario is skipped before the missing flag can matter -- and between +# them they fail three scenarios, all until open-feature/flagd-testbed#392 lands. Bump the tag +# then. # # * large-integer-flag: the untagged precision scenario, which gets the code default instead # of 2147483647, and the @variants row asking for its "max-int32" variant, answered with From 8d5d4eeec4667fc3fd48b12155ec54c21fda12b4 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sun, 13 Sep 2026 19:23:06 +0200 Subject: [PATCH 16/25] docs(flagd): cite the declaring rule this adoption argued out in longhand The javadoc beside knownDeviations() argues from first principles that declaring @numeric-coercion is right even though one of its three scenarios asks for a flag the pinned testbed does not serve. Appendix F now states that as a rule, so the argument belongs upstream and the citation belongs here. Both halves of it apply in order, which is worth spelling out because the first half is what keeps the rule from over-reaching: flagd is attempting the coercion -- the widening scenario passes, which is the evidence -- and two of the three scenarios can be put to it. A provider that does not coerce at all stops at the first half and withholds, as the SDK's in-memory provider does. Also names the appendix's other consequence as the reason the deviation summary names the testbed: a scenario that fails because the backend cannot serve its fixture is not a provider defect, and a summary that did not say so would leave the report accusing flagd of the stack's gap. Signed-off-by: Simon Schrottner --- .../providers/flagd/e2e/AbstractFlagdTckTest.java | 13 +++++++++++++ 1 file changed, 13 insertions(+) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java index 1eef460df8..d0805a3213 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java @@ -290,6 +290,19 @@ public Set capabilities() { * behind a withheld capability. And the alternative is worse, because a skip cannot distinguish * "flagd declines to coerce" from "flagd coerces and gets it wrong", and only the second is true. * + *

That argument is no longer this file's to make. + * Appendix + * F now states it as a declaring rule — once a provider is attempting a capability, the unit + * of the decision is the scenario rather than the tag, so it is declared when at least + * one scenario gating it can be put to the provider and withheld only when none can. Both halves + * apply here in order: flagd is attempting the coercion, which the widening scenario + * proves, and two of the tag's three scenarios can be put to it. A provider that simply does not + * coerce would stop at the first half and withhold, which is what the SDK's in-memory provider + * does and why the self-tests in {@code tools/tck} skip these scenarios. The appendix's other + * consequence is why the summary below names the testbed: a scenario that fails because the + * backend cannot serve its fixture is not a provider defect, and a deviation that did not say so + * would have the report accuse flagd of the stack's gap. + * *

Tracked against flagd's numeric coercion ADR, which is where the rule this deviates from is * settled: coercion is permitted when it is lossless and must fail with {@code TYPE_MISMATCH} * only when information would be lost. The summary says which half is broken, because "flagd From 7aa92a14effc94a14645a656fc74a4ca9729d958 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sun, 13 Sep 2026 20:56:48 +0200 Subject: [PATCH 17/25] test(flagd): run the conformance suites from a step of their own `mvn --projects providers/flagd -P e2e test` does not run the conformance suites. The `e2e` profile is the thing that keeps them out: it narrows the exclusion to **/e2e/*TckTest.java so the legacy Run*Test suites run and the TCK ones do not. Verified by running it rather than reading the POM - 788 tests in 12:18, RunFileTest, RunInProcessTest and RunRpcTest, and not one mention of either TckTest class anywhere in the log. Appendix F now asks for a step of its own rather than a corner of an existing e2e suite, because of what a red build says: `-Pe2e` red means the provider's own end-to-end suites regressed, while `-Ptck` red means conformance failed, and a conformance run carries failures by design wherever AbstractFlagdTckTest declares a knownDeviation. The new `tck` profile clears the exclusion and narrows Surefire's includes to **/e2e/*TckTest.java in the same breath, which is what makes it a conformance step and not a wider one: `mvn -Ptck -pl providers/flagd test` runs FlagdRpcTckTest and FlagdInProcessTckTest and nothing else - 130 tests, 65 scenarios per resolver, 59 passing, 2 skipped and 4 failing in both. Nothing activates the profile in CI, for the reason the `e2e` profile's comment already gives. The exclusion stays Surefire's and not the compiler's, so both suites still compile in the default build. The README says not to add `-am` to that run, and why: it pulls tools/tck and tools/flagd-core into the reactor and runs their suites too, so a failure in either comes out as a `-Ptck` failure. A one-off `install` is what `-am` was there for. Signed-off-by: Simon Schrottner --- providers/flagd/README.md | 34 +++++++++++++++++++++++++----- providers/flagd/pom.xml | 44 +++++++++++++++++++++++++++++++++++++++ 2 files changed, 73 insertions(+), 5 deletions(-) diff --git a/providers/flagd/README.md b/providers/flagd/README.md index bdffb4bb4a..4d197d7d0f 100644 --- a/providers/flagd/README.md +++ b/providers/flagd/README.md @@ -389,6 +389,23 @@ mvn -Pe2e -pl providers/flagd help:evaluate -Dexpression=testExclusions -DforceS Narrowing keeps the legacy suites running exactly as before and the TCK suites out. The exclusion is Surefire's, not the compiler's, so both suites still build against the harness in every job. +**So `-Pe2e` does not run the conformance suites — it is the profile that keeps them out.** Worth +stating plainly, because a command of the form `mvn -Pe2e -pl providers/flagd test` reads as if it +runs everything under `e2e/` and runs the legacy suites instead, in silence: 788 tests, no scenario +tally, and not one mention of either `TckTest` class in the log. + +**The conformance suites have a profile of their own, `tck`.** That is the separation +[Appendix F](https://github.com/open-feature/spec/blob/main/specification/appendix-f-provider-conformance.md#running-the-suite-in-ci) +asks for, and the reason is what a red build *says* rather than how long it takes. `-Pe2e` red means +the provider's own end-to-end suites regressed. `-Ptck` red means conformance failed — and a +conformance run carries failures by design, wherever `AbstractFlagdTckTest` declares a +`knownDeviation`. One signal shared between "you broke something" and "this is the known state" ends +with somebody silencing the informative half. + +The profile clears the exclusion and narrows Surefire's includes to `**/e2e/*TckTest.java` in the +same breath, so it runs the two conformance suites and nothing else — not the legacy `Run*Test` +suites and not the module's unit tests. Nothing activates it in CI. + The consequence is that **no CI job runs them**, so a maintainer runs them by hand before merging a change that touches the provider's resolution, event or lifecycle behaviour, and quotes the result in the pull request. A scheduled or path-filtered workflow was considered and declined: a suite whose @@ -396,16 +413,23 @@ red is diagnosed by whoever happens to read the notification is worse than one w diagnosed by the person who caused it. ```bash +# once, if tools/tck is not in your local repository yet +mvn -pl tools/tck -am -DskipTests install + # both resolvers -mvn -pl providers/flagd -am -DtestExclusions= -Dtest='Flagd*TckTest' \ - -Dsurefire.failIfNoSpecifiedTests=false test +mvn -Ptck -pl providers/flagd test # one resolver -mvn -pl providers/flagd -am -DtestExclusions= -Dtest=FlagdInProcessTckTest \ - -Dsurefire.failIfNoSpecifiedTests=false test +mvn -Ptck -pl providers/flagd -Dtest=FlagdInProcessTckTest test ``` +**Do not add `-am` to the run itself.** It pulls `tools/tck` and `tools/flagd-core` into the reactor +and runs their test suites too — 246 and 75 tests before the first scenario — so a failure anywhere +in either of them comes out as a `-Ptck` failure. That is the signal-mixing this step exists to +prevent, reintroduced by a flag. The separate `install` above is what `-am` was there for. + Both suites are currently **expected to fail**, and the expected failures are enumerated in `AbstractFlagdTckTest`: three come from flags that the pinned `flagd-testbed` image does not serve (open-feature/flagd-testbed#392) and one is the real numeric-coercion defect, declared and left -visible rather than skipped (open-feature/flagd#1996). Anything else is a regression. +visible rather than skipped (open-feature/flagd#1996). Anything else is a regression. Each resolver +is 65 scenarios: 59 passing, 2 skipped and 4 failing, the same in both. diff --git a/providers/flagd/pom.xml b/providers/flagd/pom.xml index 4ea4d40f19..6e1ae86b1c 100644 --- a/providers/flagd/pom.xml +++ b/providers/flagd/pom.xml @@ -325,6 +325,50 @@ + + + tck + + + + + + + + org.apache.maven.plugins + maven-surefire-plugin + + + + **/e2e/*TckTest.java + + + + + + From 418af516c4d7e236403a4ab157bc754cc679e843 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Sun, 13 Sep 2026 22:22:08 +0200 Subject: [PATCH 18/25] test(flagd): give the conformance adoption a package of its own The two conformance suites lived under e2e/ and were selected out of it by filename: `**/e2e/*TckTest.java`, in three places. They now live in a sibling package, `src/test/java/.../flagd/tck/`, and every selector names the directory instead. Nesting them under e2e/ said "this is a kind of e2e test", which is the conflation the separate `tck` step exists to undo. The two suites answer different questions -- e2e tests this provider against flagd's own test harness, tck tests it against the OpenFeature provider contract -- and they mean different things by a red run: an e2e suite is expected green, while a conformance suite fails scenarios by design wherever a knownDeviation is declared. A file is now in the conformance directory or it is not, and no convention about class names holds that line. So the names drop what the directory now says. FlagdRpcTckTest in package ...flagd.tck said "tck" twice and "flagd" twice; it is RpcTest, beside InProcessTest, over AbstractResolverTest. The `*Test` suffix stays because Surefire's default includes need it -- that is not the selector being removed. The three selectors, all directory-shaped now: * the module's default testExclusions gains tck, so both Docker-dependent packages stay out of `mvn verify`; * the `e2e` profile drops e2e from the exclusion and keeps tck, where it used to narrow to a filename pattern; * the `tck` profile is its mirror image -- drops tck, keeps e2e, and includes `**/tck/*.java`. Both halves of the `tck` profile are still needed, for the reason they always were: the include alone leaves the exclusion in force and runs nothing, and dropping the exclusion alone runs the module's unit tests alongside the suites. testExclusions is still a Surefire and not a compiler exclusion, so both packages still compile in the default build. RpcTest and InProcessTest now state their configuration() rather than deriving it. The default derivation reads the class name, so the rename would have filed their runs as "rpc" and "in-process" instead of "flagd-rpc" and "flagd-in-process" -- a report is read away from this repository, where the provider's name is the half that matters. The suites are otherwise unchanged. Unchanged too: 65 scenarios per resolver, 59 passed, 2 skipped, 4 failed, in both modes. Same scenarios, same results, a different directory. The profile comment documenting the run had kept `-am` on it, which the READMEs already warn against; it says the command that works now. Signed-off-by: Simon Schrottner --- providers/flagd/README.md | 60 +++++++++++-------- providers/flagd/pom.xml | 49 +++++++++------ .../flagd/e2e/FlagdInProcessTckTest.java | 17 ------ .../providers/flagd/e2e/FlagdRpcTckTest.java | 17 ------ .../AbstractResolverTest.java} | 4 +- .../providers/flagd/tck/InProcessTest.java | 28 +++++++++ .../contrib/providers/flagd/tck/RpcTest.java | 31 ++++++++++ .../test/resources/tck/docker-compose.yaml | 2 +- 8 files changed, 128 insertions(+), 80 deletions(-) delete mode 100644 providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdInProcessTckTest.java delete mode 100644 providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdRpcTckTest.java rename providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/{e2e/AbstractFlagdTckTest.java => tck/AbstractResolverTest.java} (99%) create mode 100644 providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/InProcessTest.java create mode 100644 providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/RpcTest.java diff --git a/providers/flagd/README.md b/providers/flagd/README.md index 4d197d7d0f..5992b612e0 100644 --- a/providers/flagd/README.md +++ b/providers/flagd/README.md @@ -362,49 +362,59 @@ FlagdOptions options = FlagdOptions.builder() ## Provider conformance (TCK) This provider adopts the [OpenFeature Provider TCK](../../tools/tck/README.md), once per resolver: -`FlagdRpcTckTest` and `FlagdInProcessTckTest`, both over the shared -`AbstractFlagdTckTest`. Read that class before changing either — it records which capabilities are -declared, which are withheld and why, and every known deviation, each against measured behaviour -rather than assumption. - -**The two suites are Docker-gated and excluded from the default build.** They live under -`src/test/java/.../e2e/`, which `**/e2e/*.java` in this module's -POM keeps out of `mvn verify`. Why an adoption suite is excluded rather than gating is written down +`RpcTest` and `InProcessTest`, both over the shared `AbstractResolverTest`. Read that class before +changing either — it records which capabilities are declared, which are withheld and why, and every +known deviation, each against measured behaviour rather than assumption. + +**The adoption has a source directory of its own**, `src/test/java/.../flagd/tck/`, beside the +legacy end-to-end suites in `.../flagd/e2e/` rather than inside them. The two answer different +questions — `e2e` tests this provider against flagd's own test harness, `tck` tests it against the +OpenFeature provider contract — and they mean different things by a red run: an e2e suite is +expected green, while a conformance suite fails scenarios by design wherever a `knownDeviation` is +declared. Everything below follows from that, including the class names: in a package called `tck`, +`RpcTest` needs no further label, and the fully-qualified name still reads +`...providers.flagd.tck.RpcTest`. + +**Both packages are Docker-gated and excluded from the default build**, by +`**/e2e/*.java,**/tck/*.java` in this module's POM, which is what +keeps them out of `mvn verify`. Why an adoption suite is excluded rather than gating is written down once for all four languages in [Appendix F: Running the suite in CI](https://github.com/open-feature/spec/blob/main/specification/appendix-f-provider-conformance.md#running-the-suite-in-ci); this section is only what that means in this module. -**The `e2e` profile narrows that exclusion rather than clearing it**, to `**/e2e/*TckTest.java`. -The profile exists for the legacy `Run*Test` suites over the `test-harness` submodule, and -`ci.yml`'s `main` job activates it on every push — so a profile that cleared the exclusion outright, -which is what this one used to do, ran the TCK suites in CI on a runner that does have a Docker -daemon. They are expected to fail, so every unrelated pull request went red for a reason that had -nothing to do with it. That is the first of the two mistakes the appendix names, found here by -resolving the property rather than reading the POM: +**The `e2e` profile drops `e2e` from that exclusion and keeps `tck`.** The profile exists for the +legacy `Run*Test` suites over the `test-harness` submodule, and `ci.yml`'s `main` job activates it +on every push — so a profile that cleared the property outright, which is what this one used to do, +ran the TCK suites in CI on a runner that does have a Docker daemon. They are expected to fail, so +every unrelated pull request went red for a reason that had nothing to do with it. That is the first +of the two mistakes the appendix names, found here by resolving the property rather than reading the +POM: ```bash mvn -Pe2e -pl providers/flagd help:evaluate -Dexpression=testExclusions -DforceStdout ``` -Narrowing keeps the legacy suites running exactly as before and the TCK suites out. The exclusion is -Surefire's, not the compiler's, so both suites still build against the harness in every job. +The legacy suites run exactly as before and the conformance suites stay out. The exclusion is +Surefire's, not the compiler's, so both packages still build against the harness in every job. **So `-Pe2e` does not run the conformance suites — it is the profile that keeps them out.** Worth stating plainly, because a command of the form `mvn -Pe2e -pl providers/flagd test` reads as if it -runs everything under `e2e/` and runs the legacy suites instead, in silence: 788 tests, no scenario -tally, and not one mention of either `TckTest` class in the log. +ran everything and runs the legacy suites instead, in silence: 788 tests, no scenario tally, and not +one mention of either conformance suite in the log. **The conformance suites have a profile of their own, `tck`.** That is the separation [Appendix F](https://github.com/open-feature/spec/blob/main/specification/appendix-f-provider-conformance.md#running-the-suite-in-ci) asks for, and the reason is what a red build *says* rather than how long it takes. `-Pe2e` red means the provider's own end-to-end suites regressed. `-Ptck` red means conformance failed — and a -conformance run carries failures by design, wherever `AbstractFlagdTckTest` declares a +conformance run carries failures by design, wherever `AbstractResolverTest` declares a `knownDeviation`. One signal shared between "you broke something" and "this is the known state" ends with somebody silencing the informative half. -The profile clears the exclusion and narrows Surefire's includes to `**/e2e/*TckTest.java` in the -same breath, so it runs the two conformance suites and nothing else — not the legacy `Run*Test` -suites and not the module's unit tests. Nothing activates it in CI. +The profile is the mirror image of `e2e`: it drops `tck` from the exclusion and narrows Surefire's +includes to `**/tck/*.java` in the same breath, so it runs the two conformance suites and nothing +else — not the legacy `Run*Test` suites and not the module's unit tests. Both halves are needed. The +include alone leaves the exclusion in force and runs nothing; dropping the exclusion alone runs the +module's unit tests alongside the suites. Nothing activates it in CI. The consequence is that **no CI job runs them**, so a maintainer runs them by hand before merging a change that touches the provider's resolution, event or lifecycle behaviour, and quotes the result in @@ -420,7 +430,7 @@ mvn -pl tools/tck -am -DskipTests install mvn -Ptck -pl providers/flagd test # one resolver -mvn -Ptck -pl providers/flagd -Dtest=FlagdInProcessTckTest test +mvn -Ptck -pl providers/flagd -Dtest=InProcessTest test ``` **Do not add `-am` to the run itself.** It pulls `tools/tck` and `tools/flagd-core` into the reactor @@ -429,7 +439,7 @@ in either of them comes out as a `-Ptck` failure. That is the signal-mixing this prevent, reintroduced by a flag. The separate `install` above is what `-am` was there for. Both suites are currently **expected to fail**, and the expected failures are enumerated in -`AbstractFlagdTckTest`: three come from flags that the pinned `flagd-testbed` image does not serve +`AbstractResolverTest`: three come from flags that the pinned `flagd-testbed` image does not serve (open-feature/flagd-testbed#392) and one is the real numeric-coercion defect, declared and left visible rather than skipped (open-feature/flagd#1996). Anything else is a regression. Each resolver is 65 scenarios: 59 passing, 2 skipped and 4 failing, the same in both. diff --git a/providers/flagd/pom.xml b/providers/flagd/pom.xml index 6e1ae86b1c..d81a4aaaf1 100644 --- a/providers/flagd/pom.xml +++ b/providers/flagd/pom.xml @@ -14,8 +14,13 @@ 0.14.2 - - **/e2e/*.java + + **/e2e/*.java,**/tck/*.java 1.82.0 3.25.6 @@ -102,8 +107,8 @@ dev.openfeature.contrib.tools @@ -265,18 +270,22 @@ - **/e2e/*TckTest.java + **/tck/*.java @@ -329,12 +338,16 @@ - + **/e2e/*.java @@ -360,9 +373,9 @@ maven-surefire-plugin - - **/e2e/*TckTest.java + + **/tck/*.java diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdInProcessTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdInProcessTckTest.java deleted file mode 100644 index 1ce5b1dc71..0000000000 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdInProcessTckTest.java +++ /dev/null @@ -1,17 +0,0 @@ -package dev.openfeature.contrib.providers.flagd.e2e; - -import dev.openfeature.contrib.providers.flagd.Config; - -/** Runs the OpenFeature Provider TCK against the flagd provider in in-process mode. */ -public class FlagdInProcessTckTest extends AbstractFlagdTckTest { - - @Override - protected Config.Resolver resolver() { - return Config.Resolver.IN_PROCESS; - } - - @Override - protected int backendPort() { - return 8015; - } -} diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdRpcTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdRpcTckTest.java deleted file mode 100644 index 30ee6db57c..0000000000 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/FlagdRpcTckTest.java +++ /dev/null @@ -1,17 +0,0 @@ -package dev.openfeature.contrib.providers.flagd.e2e; - -import dev.openfeature.contrib.providers.flagd.Config; - -/** Runs the OpenFeature Provider TCK against the flagd provider in RPC mode. */ -public class FlagdRpcTckTest extends AbstractFlagdTckTest { - - @Override - protected Config.Resolver resolver() { - return Config.Resolver.RPC; - } - - @Override - protected int backendPort() { - return 8013; - } -} diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java similarity index 99% rename from providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java rename to providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java index d0805a3213..2c5234e4f0 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/e2e/AbstractFlagdTckTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java @@ -1,4 +1,4 @@ -package dev.openfeature.contrib.providers.flagd.e2e; +package dev.openfeature.contrib.providers.flagd.tck; import dev.openfeature.contrib.providers.flagd.Config; import dev.openfeature.contrib.providers.flagd.FlagdOptions; @@ -57,7 +57,7 @@ * declared as a {@link KnownDeviation}: a deviation says the provider is wrong, and the provider was * never given the flag to get wrong. */ -abstract class AbstractFlagdTckTest extends ContainerizedProviderTckTest { +abstract class AbstractResolverTest extends ContainerizedProviderTckTest { /** * A port nothing listens on, for the initialisation-failure scenarios. diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/InProcessTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/InProcessTest.java new file mode 100644 index 0000000000..cfa7bce983 --- /dev/null +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/InProcessTest.java @@ -0,0 +1,28 @@ +package dev.openfeature.contrib.providers.flagd.tck; + +import dev.openfeature.contrib.providers.flagd.Config; + +/** Runs the OpenFeature Provider TCK against the flagd provider in in-process mode. */ +public class InProcessTest extends AbstractResolverTest { + + /** + * {@inheritDoc} + * + *

Stated rather than derived, for the reason {@link RpcTest#configuration()} gives: the + * class name no longer carries the provider, and a report is read away from this repository. + */ + @Override + public String configuration() { + return "flagd-in-process"; + } + + @Override + protected Config.Resolver resolver() { + return Config.Resolver.IN_PROCESS; + } + + @Override + protected int backendPort() { + return 8015; + } +} diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/RpcTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/RpcTest.java new file mode 100644 index 0000000000..cc18d4eae3 --- /dev/null +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/RpcTest.java @@ -0,0 +1,31 @@ +package dev.openfeature.contrib.providers.flagd.tck; + +import dev.openfeature.contrib.providers.flagd.Config; + +/** Runs the OpenFeature Provider TCK against the flagd provider in RPC mode. */ +public class RpcTest extends AbstractResolverTest { + + /** + * {@inheritDoc} + * + *

Stated rather than derived. The default derivation reads the suite's class name, which + * used to be {@code FlagdRpcTckTest} and now says only which resolver it is, because the + * package says the rest. A report is read away from this repository, where {@code rpc} on its + * own would not say whose RPC resolver it was, so the name a run is filed under is written here + * instead of moving with the class name. + */ + @Override + public String configuration() { + return "flagd-rpc"; + } + + @Override + protected Config.Resolver resolver() { + return Config.Resolver.RPC; + } + + @Override + protected int backendPort() { + return 8013; + } +} diff --git a/providers/flagd/src/test/resources/tck/docker-compose.yaml b/providers/flagd/src/test/resources/tck/docker-compose.yaml index b287f8b097..f1d7b4fadf 100644 --- a/providers/flagd/src/test/resources/tck/docker-compose.yaml +++ b/providers/flagd/src/test/resources/tck/docker-compose.yaml @@ -15,7 +15,7 @@ # * integral-float-flag: the @numeric-coercion scenario "An integral float requested as an # integer is coerced without loss". This provider declares @numeric-coercion because flagd # does coerce and gets the lossy direction wrong, so the tag's failure is worth seeing -- -# see AbstractFlagdTckTest.knownDeviations(). This third failure is the price of that, and +# see AbstractResolverTest.knownDeviations(). This third failure is the price of that, and # it is the same missing-flag problem as the two above rather than a new one. # # targeting-key-flag is served, which is why @targeting needs no bump: its three scenarios pass From 91efe843baa52932239d5e813f3339ffb0efc61b Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Mon, 14 Sep 2026 06:56:43 +0200 Subject: [PATCH 19/25] docs(flagd): cut the conformance section to what is this adoption's The TCK section was 5.5 KB of a 28 KB provider README, and most of it was the base README's or Appendix F's: why an adoption suite is excluded rather than gating, why it gets a step of its own, what a red conformance build says, why both halves of the tck profile are needed, and why the -am the command must not carry would mix signals. What is left answers the three questions an adoption README owes a reader. What this provider declares and why each absence is what it is - by pointing at AbstractResolverTest, where the reasoning sits next to the declaration and is measured rather than asserted. What the tally is and which failures are expected - 65 scenarios, 59/2/4 per resolver, three testbed gaps and one real defect. And the command that runs it. One local fact stays at length because it is a trap and is recorded nowhere else: -Pe2e does not run these suites, it is the profile that keeps them out, and it used to clear the exclusion outright and run them red in CI on every unrelated pull request. 2.4 KB, from 5.5. The adoption now adds 38 lines to this README rather than 85. Signed-off-by: Simon Schrottner --- providers/flagd/README.md | 97 ++++++++++----------------------------- 1 file changed, 25 insertions(+), 72 deletions(-) diff --git a/providers/flagd/README.md b/providers/flagd/README.md index 5992b612e0..8338780ec8 100644 --- a/providers/flagd/README.md +++ b/providers/flagd/README.md @@ -362,84 +362,37 @@ FlagdOptions options = FlagdOptions.builder() ## Provider conformance (TCK) This provider adopts the [OpenFeature Provider TCK](../../tools/tck/README.md), once per resolver: -`RpcTest` and `InProcessTest`, both over the shared `AbstractResolverTest`. Read that class before -changing either — it records which capabilities are declared, which are withheld and why, and every -known deviation, each against measured behaviour rather than assumption. - -**The adoption has a source directory of its own**, `src/test/java/.../flagd/tck/`, beside the -legacy end-to-end suites in `.../flagd/e2e/` rather than inside them. The two answer different -questions — `e2e` tests this provider against flagd's own test harness, `tck` tests it against the -OpenFeature provider contract — and they mean different things by a red run: an e2e suite is -expected green, while a conformance suite fails scenarios by design wherever a `knownDeviation` is -declared. Everything below follows from that, including the class names: in a package called `tck`, -`RpcTest` needs no further label, and the fully-qualified name still reads -`...providers.flagd.tck.RpcTest`. - -**Both packages are Docker-gated and excluded from the default build**, by -`**/e2e/*.java,**/tck/*.java` in this module's POM, which is what -keeps them out of `mvn verify`. Why an adoption suite is excluded rather than gating is written down -once for all four languages in -[Appendix F: Running the suite in CI](https://github.com/open-feature/spec/blob/main/specification/appendix-f-provider-conformance.md#running-the-suite-in-ci); -this section is only what that means in this module. - -**The `e2e` profile drops `e2e` from that exclusion and keeps `tck`.** The profile exists for the -legacy `Run*Test` suites over the `test-harness` submodule, and `ci.yml`'s `main` job activates it -on every push — so a profile that cleared the property outright, which is what this one used to do, -ran the TCK suites in CI on a runner that does have a Docker daemon. They are expected to fail, so -every unrelated pull request went red for a reason that had nothing to do with it. That is the first -of the two mistakes the appendix names, found here by resolving the property rather than reading the -POM: - -```bash -mvn -Pe2e -pl providers/flagd help:evaluate -Dexpression=testExclusions -DforceStdout -``` - -The legacy suites run exactly as before and the conformance suites stay out. The exclusion is -Surefire's, not the compiler's, so both packages still build against the harness in every job. - -**So `-Pe2e` does not run the conformance suites — it is the profile that keeps them out.** Worth -stating plainly, because a command of the form `mvn -Pe2e -pl providers/flagd test` reads as if it -ran everything and runs the legacy suites instead, in silence: 788 tests, no scenario tally, and not -one mention of either conformance suite in the log. - -**The conformance suites have a profile of their own, `tck`.** That is the separation -[Appendix F](https://github.com/open-feature/spec/blob/main/specification/appendix-f-provider-conformance.md#running-the-suite-in-ci) -asks for, and the reason is what a red build *says* rather than how long it takes. `-Pe2e` red means -the provider's own end-to-end suites regressed. `-Ptck` red means conformance failed — and a -conformance run carries failures by design, wherever `AbstractResolverTest` declares a -`knownDeviation`. One signal shared between "you broke something" and "this is the known state" ends -with somebody silencing the informative half. - -The profile is the mirror image of `e2e`: it drops `tck` from the exclusion and narrows Surefire's -includes to `**/tck/*.java` in the same breath, so it runs the two conformance suites and nothing -else — not the legacy `Run*Test` suites and not the module's unit tests. Both halves are needed. The -include alone leaves the exclusion in force and runs nothing; dropping the exclusion alone runs the -module's unit tests alongside the suites. Nothing activates it in CI. - -The consequence is that **no CI job runs them**, so a maintainer runs them by hand before merging a -change that touches the provider's resolution, event or lifecycle behaviour, and quotes the result in -the pull request. A scheduled or path-filtered workflow was considered and declined: a suite whose -red is diagnosed by whoever happens to read the notification is worse than one whose red is -diagnosed by the person who caused it. +`RpcTest` and `InProcessTest`, both over the shared `AbstractResolverTest` in +`src/test/java/.../flagd/tck/`. **Read that class before changing either.** It records, against +measured behaviour rather than assumption, which capabilities are declared, which are withheld and +why, and every known deviation — that reasoning is the most valuable thing about this adoption and it +lives next to the declaration rather than here. ```bash # once, if tools/tck is not in your local repository yet mvn -pl tools/tck -am -DskipTests install -# both resolvers -mvn -Ptck -pl providers/flagd test - -# one resolver -mvn -Ptck -pl providers/flagd -Dtest=InProcessTest test +mvn -Ptck -pl providers/flagd test # both resolvers +mvn -Ptck -pl providers/flagd -Dtest=InProcessTest test # one ``` -**Do not add `-am` to the run itself.** It pulls `tools/tck` and `tools/flagd-core` into the reactor -and runs their test suites too — 246 and 75 tests before the first scenario — so a failure anywhere -in either of them comes out as a `-Ptck` failure. That is the signal-mixing this step exists to -prevent, reintroduced by a flag. The separate `install` above is what `-am` was there for. +Do not add `-am` to the run itself; the TCK README says why. + +**`-Pe2e` does not run these suites — it is the profile that keeps them out.** Worth stating plainly, +because `mvn -Pe2e -pl providers/flagd test` reads as if it ran everything and instead runs the legacy +`Run*Test` suites in silence: 788 tests, no scenario tally, and no mention of either conformance +suite. This module excludes `**/e2e/*.java,**/tck/*.java` by default; the `e2e` profile drops only the +first, and the `tck` profile only the second. It used to clear the property outright, which — since +`ci.yml`'s `main` job activates `e2e` on every push, on a runner that has a Docker daemon — ran these +suites in CI, where they are expected to fail, and turned every unrelated pull request red. + +**No CI job runs them**, so a maintainer runs them by hand before merging a change to the provider's +resolution, event or lifecycle behaviour, and quotes the result in the pull request. A scheduled or +path-filtered workflow was considered and declined: a suite whose red is diagnosed by whoever happens +to read the notification is worse than one whose red is diagnosed by the person who caused it. -Both suites are currently **expected to fail**, and the expected failures are enumerated in -`AbstractResolverTest`: three come from flags that the pinned `flagd-testbed` image does not serve +Both suites are **expected to fail**, identically: 65 scenarios each, 59 passing, 2 skipped, 4 +failing. Three of the four failures come from flags the pinned `flagd-testbed` image does not serve (open-feature/flagd-testbed#392) and one is the real numeric-coercion defect, declared and left -visible rather than skipped (open-feature/flagd#1996). Anything else is a regression. Each resolver -is 65 scenarios: 59 passing, 2 skipped and 4 failing, the same in both. +visible rather than skipped (open-feature/flagd#1996). `AbstractResolverTest` enumerates them by name. +Anything else is a regression. From 046c6a5f4ab7c397aae8cafd9177d44c6a6ddb12 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Mon, 14 Sep 2026 08:35:56 +0200 Subject: [PATCH 20/25] test(flagd): one Compose file for both adoptions, and cut the prose around it Stripped of comments the two TCK Compose files in this repository were the same stack: same image, same service name, same absence of host bindings, differing only in which ports each listed. Two hand-maintained copies is the mechanism by which they drift onto different images while both claiming to ask the same questions, so there is now one file at tools/flagd-testbed/docker-compose.yaml exposing 8013, 8015, 8016 and 8080. Extra ports cost nothing: with no host bindings the TCK maps each container port dynamically and resolves only the ones a suite asks for, so the flagd suites never look up 8016. Placed outside both provider modules rather than in either, so neither reaches into the other's test tree and a third adoption against the same backend reads the same path. That costs a tools/ directory which is not a Maven module -- no pom, no CHANGELOG, nothing released -- and a composeFile() that climbs out of its own module, which both suites now explain rather than leaving to look like a mistake. The comment is down from 29 lines across two files to 13 in one, and keeps only what is a trap here: that it is NOT the legacy e2e stack at providers/flagd/test-harness, which bind-mounts ${FLAGS_DIR}, names its service "flagd" and runs an envoy sidecar. Which flags the image does not serve belongs to open-feature/flagd-testbed#392, and which scenarios that costs is in AbstractResolverTest. AbstractResolverTest loses the testbed-gap essay, the falsy-flag rename history and two paragraphs about what an earlier revision of the file declared. Every measured probe stays: the @numeric-coercion three-way result, the @reinitialization and @stale source-line evidence, the @disabled-flags and @standard-reasons runs, and both findings recorded on CONNECTED_DEADLINE_MS. Comments and YAML comments only. No behaviour change. Signed-off-by: Simon Schrottner --- .../flagd/tck/AbstractResolverTest.java | 293 ++++++------------ .../test/resources/tck/docker-compose.yaml | 36 --- tools/flagd-testbed/docker-compose.yaml | 21 ++ 3 files changed, 119 insertions(+), 231 deletions(-) delete mode 100644 providers/flagd/src/test/resources/tck/docker-compose.yaml create mode 100644 tools/flagd-testbed/docker-compose.yaml diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java index 2c5234e4f0..85f17b5726 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java @@ -24,38 +24,13 @@ * one is running from the JUnit test plan, so adding a mode needs no registration or build * configuration. * - *

The testbed does not yet serve the whole canonical flag set. Three of the - * flags the suite's assets added are absent from {@code flagd-testbed} v3.8.0: - * {@code large-integer-flag}, {@code huge-integer-flag} and {@code integral-float-flag}. Two of the - * three are reached — {@code huge-integer-flag} is asked for solely under {@code @large-integers}, - * which no Java provider can declare because the SDK's integer accessor is 32 bits, so that - * scenario is skipped before the missing flag can matter — so three scenarios fail with - * {@code FLAG_NOT_FOUND} in both modes until open-feature/flagd-testbed#392 lands and the tag here - * is bumped: - * - *

    - *
  • the untagged 32-bit precision scenario, which gets the code default instead of - * {@code 2147483647}; - *
  • the {@code @variants} row asking for {@code large-integer-flag}'s {@code max-int32}, which - * is answered with no variant; - *
  • the {@code @numeric-coercion} scenario "An integral float requested as an integer is coerced - * without loss", which asks for {@code integral-float-flag}. - *
- * - *

The third of those is new here, and it is the cost of declaring {@code @numeric-coercion} - * rather than withholding it — see {@link #knownDeviations()}, which explains why paying it is the - * honest report. - * - *

The three falsy flags used to fail the same way and no longer do. The testbed's - * {@code zero-flags.json} already served {@code boolean-zero-flag}, {@code integer-zero-flag} and - * {@code string-zero-flag} with {@code zero}/{@code non-zero} variants, while the canonical set - * called them {@code false-flag}, {@code zero-flag} and {@code empty-string-flag}; spec ba002ce8 - * renamed the canonical flags to the testbed's names rather than the other way round, so those - * three scenarios now resolve against flags that were always there. - * - *

A missing flag is a gap in the stack, not in the provider, so it is recorded here rather than - * declared as a {@link KnownDeviation}: a deviation says the provider is wrong, and the provider was - * never given the flag to get wrong. + *

Three scenarios fail in both modes on flags the pinned testbed image does not serve + * (open-feature/flagd-testbed#392 names them and says what each catches): the untagged 32-bit + * precision scenario and the {@code @variants} row asking for {@code large-integer-flag}'s + * {@code max-int32}, both answered as if the flag were absent, and the {@code @numeric-coercion} + * scenario "An integral float requested as an integer is coerced without loss". The third is the + * cost of declaring {@code @numeric-coercion} — see {@link #knownDeviations()}. None of the three is + * a {@link KnownDeviation}: the provider was never given the flag to get wrong. */ abstract class AbstractResolverTest extends ContainerizedProviderTckTest { @@ -70,36 +45,25 @@ abstract class AbstractResolverTest extends ContainerizedProviderTckTest { /** * gRPC deadline for a provider that is expected to connect. * - *

Generous on purpose, and measured. flagd derives its initialisation deadline from this - * value — doubling it — and the in-process resolver must sync the entire ruleset before it - * reports ready, which takes longer than a deadline tuned for a single RPC round trip. - * - *

At 5000 the first two in-process scenarios failed reproducibly on a slower host - * with {@code Initialization timeout exceeded; did not complete within the 10000 ms deadline} - * out of {@code FlagdProviderSyncResources.waitForInitialization}, on both flagd-testbed v3.8.0 - * and v3.10.1. Raising it to 15000 cleared that. The first scenario pays for a cold container as - * well as for the sync, which is why it was the first two rather than all of them. - * - *

Worth being explicit that this is not a post-command settle in disguise. - * A pause after the control call was tried at 50ms and at 3000ms and fixed nothing — the wait - * this covers is the provider's own initialisation, which is bounded here where the scenario can - * see it, rather than slept through where it cannot. {@link #UNAVAILABLE_DEADLINE_MS} stays - * short so the promptness assertions still mean something. + *

Generous on purpose, and measured. flagd doubles this value to get its + * initialisation deadline, and the in-process resolver must sync the entire ruleset before it + * reports ready. At 5000 the first two in-process scenarios failed reproducibly with + * {@code Initialization timeout exceeded; did not complete within the 10000 ms deadline} out of + * {@code FlagdProviderSyncResources.waitForInitialization}, on flagd-testbed v3.8.0 and v3.10.1 + * alike; 15000 cleared that. Not a post-command settle in disguise: a pause after the control + * call was tried at 50ms and at 3000ms and fixed nothing, because the wait this covers is the + * provider's own initialisation. {@link #UNAVAILABLE_DEADLINE_MS} stays short so the promptness + * assertions still mean something. * *

Not fully solved, and the bound is not the thing to keep raising. On a - * loaded Docker-in-WSL host the first scenario of {@code errors.feature} still errors in - * in-process mode, with the same message against the doubled 30000 ms deadline after some 53 - * seconds of wall clock. Measured three times in a row, and — importantly — it reproduces with - * {@code @numeric-coercion} withheld exactly as it does with it declared, so it is not a - * consequence of what this suite declares. RPC mode never shows it. - * - *

That shape is stack-side readiness rather than provider slowness: a small ruleset does not - * take thirty seconds to sync, and only the mode that has to establish a sync stream and receive - * the whole ruleset after the first {@code POST /start} is affected. It is the class of defect - * open-feature/flagd-testbed#394 exists to close — a control endpoint returning before the - * backend is serving — and the TCK's own rule applies: a suite that sleeps instead of holding the - * control API to its promise stops being able to detect when the promise breaks. So the bound - * stays at 15000 and this is recorded rather than covered. + * loaded Docker-in-WSL host the first scenario of {@code errors.feature} still errors + * in in-process mode against the doubled 30000 ms deadline, after some 53 seconds of wall clock. + * Measured three times in a row, and it reproduces with {@code @numeric-coercion} withheld + * exactly as with it declared, so it is not a consequence of what this suite declares. RPC mode + * never shows it. That shape is stack-side readiness rather than provider slowness — only the + * mode that has to receive the whole ruleset after the first {@code POST /start} is affected, + * which is the defect open-feature/flagd-testbed#394 exists to close. So the bound stays at + * 15000 and this is recorded rather than covered. */ private static final int CONNECTED_DEADLINE_MS = 15000; @@ -118,9 +82,16 @@ abstract class AbstractResolverTest extends ContainerizedProviderTckTest { /** The container-internal port that resolver connects to. */ protected abstract int backendPort(); + /** + * {@inheritDoc} + * + *

Outside this module on purpose, and not the idiomatic {@code src/test/resources} path: the + * OFREP adoption runs against the same stack and names the same file, so there is one image tag + * for both rather than two that can drift. Module-relative, like any other value here. + */ @Override public File composeFile() { - return new File("src/test/resources/tck/docker-compose.yaml"); + return new File("../../tools/flagd-testbed/docker-compose.yaml"); } @Override @@ -149,34 +120,29 @@ public FeatureProvider createUnavailableProvider() { /** * {@inheritDoc} * - *

Everything declarable except {@link Capability#REINITIALIZATION}, which is a fact about - * this provider rather than a defect and is explained below. Nothing else is withheld: what no - * Java provider can claim is no longer in {@link Capability#declarable()} to remove. + *

Everything declarable except {@link Capability#REINITIALIZATION}. Each declaration below is + * on evidence from a run or from the provider's source, not by inheriting the "everything + * except" default. Measured on the pinned image, both modes: 65 scenarios, 59 passing, 2 skipped + * — the withheld {@code @reinitialization}, and {@code @large-integers}, which the SDK cannot + * ask — and 4 failing, the one real defect plus the three testbed gaps above. The cold-start + * error on {@link #CONNECTED_DEADLINE_MS} did not reproduce in that run; when it appears + * in-process it costs one further scenario. * *

{@link Capability#NUMERIC_COERCION} is declared even though one of its scenarios - * fails, and that is deliberate — see {@link #knownDeviations()} for the reasoning. - * Evaluating {@code float-flag} (0.5) through the integer API returns {@code 0} with no - * error code rather than {@code TYPE_MISMATCH} with the code default, so the fractional part is - * discarded silently. Coercion as such is permitted, and the capability says so: the rule is - * that a lossless coercion must succeed and a lossy one must fail. It is the lossy case being - * accepted that is a defect to fix. - * - *

Measured rather than assumed, in both modes: of the tag's three scenarios, "An integer - * requested as a float is widened without loss" passes — which is the fact that settles - * the shape of the report, because a provider that performs the coercion and gets one direction - * wrong is not a provider that declines to coerce. The lossy scenario fails, and the remaining - * lossless one fails only because {@code integral-float-flag} is absent from the pinned testbed - * image. - * - *

flagd's numeric behaviour is identical across both resolvers, which places the defect in - * the shared provider layer rather than in either transport. Every other capability, including - * the full non-numeric type-mismatch matrix, holds in both modes. - * - *

That includes {@link Capability#LIFECYCLE}, and legitimately so: flagd reaches its backend - * during initialisation in both modes — an RPC round trip, or a full ruleset sync — so the - * lifecycle scenarios assert something real here rather than passing vacuously. Worth stating - * because the Go and JavaScript flagd providers withhold it; Java declaring it is what made that - * divergence visible, and the other two are being changed to match rather than the reverse. + * fails — see {@link #knownDeviations()}. Evaluating {@code float-flag} (0.5) through + * the integer API returns {@code 0} with no error code rather than + * {@code TYPE_MISMATCH} with the code default, so the fractional part is discarded silently. + * Measured in both modes: of the tag's three scenarios, "An integer requested as a float is + * widened without loss" passes, which is the fact that settles the shape of the report, + * because a provider that performs the coercion and gets one direction wrong is not one that + * declines to coerce. The lossy scenario fails; the remaining lossless one fails only on the + * absent {@code integral-float-flag}. Both resolvers behave identically, which places the defect + * in the shared provider layer rather than in either transport — as does every other capability + * here, including the full non-numeric type-mismatch matrix. + * + *

{@link Capability#LIFECYCLE} is declared because flagd reaches its backend during + * initialisation in both modes — an RPC round trip, or a full ruleset sync — so the lifecycle + * scenarios assert something real here rather than passing vacuously. * *

{@link Capability#REINITIALIZATION} is withheld, and that is a fact about the * provider rather than a defect in it. {@code shutdown()} sets the sync resources' own @@ -186,84 +152,42 @@ public FeatureProvider createUnavailableProvider() { * (FlagdProvider.java:121-125): the resolver is shut down, the RPC channel was * {@code shutdownNow()}'d, the retry scheduler is terminated and {@code errorExecutor} is a * {@code final} field nothing re-creates. A shut-down flagd provider is terminally shut down. + * Requirement 2.5.2 permits that, so withholding the tag is the whole of what is owed — see + * {@link Capability#REINITIALIZATION} — and the one scenario it gates is skipped with this + * reason on every run. * - *

Requirement 2.5.2 says a provider SHOULD revert to its uninitialized state after - * shutdown, and its supporting text says "some providers MAY allow reinitialization from - * this state" — so reuse is permitted, not required, and declining it is one of the options - * the requirement offers. An earlier version of this file recorded it as a {@code KnownDeviation} - * against {@code @lifecycle}, which was wrong twice over: the scenario was mandatory only because - * the spec's assets had not yet gated it, and the entry asserted a defect against a provider - * behaving within the requirement. Withholding the tag is the whole of what is owed here; the - * one scenario it gates is reported as skipped with this reason on every run. - * - *

{@link Capability#STALE} is declared for both resolvers, and the declaration is examined - * rather than inherited. {@code PROVIDER_STALE} is emitted from {@code FlagdProvider.onError} + *

{@link Capability#STALE} is declared for both resolvers, and examined rather than inherited. + * {@code PROVIDER_STALE} is emitted from {@code FlagdProvider.onError} * (FlagdProvider.java:258-264), which the shared {@code onProviderEvent} switch reaches on * {@code PROVIDER_ERROR} from either resolver (FlagdProvider.java:197, 236), before the grace - * period turns it into {@code PROVIDER_ERROR} — so the emit sits in the provider layer, not in a - * transport, and the scenario "Losing the backend makes the provider stale, regaining it makes - * it ready again" passes in RPC mode as well as in-process. Worth stating because Go's flagd - * provider withholds the tag for its RPC resolver; on this evidence that is a difference between - * the two implementations, not a property of the transport. - * - *

{@link Capability#LARGE_INTEGERS} is no longer named here, and that is the change rather - * than an omission. This suite used to withhold it with a paragraph explaining that - * {@code Client.getIntegerDetails} is a 32-bit {@code Integer} with no room for 2^53 − 1 — the - * same paragraph the OFREP suite and two suites inside {@code tools/tck} each carried, because - * every Java adopter was expected to know the fact and act on it. The TCK refuses the capability - * centrally now, so there is nothing for an adoption to decide and nothing here to get wrong. - * Its scenario is still skipped, with a reason naming the SDK rather than this provider, which - * is the part a report's reader needs: flagd declined nothing. - * - *

{@link Capability#VARIANTS} and {@link Capability#TARGETING} are both declared, and both on - * evidence rather than by inheriting the "everything except" default. flagd names the variant it - * served in every resolution, so seven of the {@code @variants} outline's eight rows pass in both - * modes; the eighth asks for {@code large-integer-flag}'s {@code max-int32} and is answered with - * no variant because testbed v3.8.0 does not serve that flag at all — the same gap that fails the - * untagged precision scenario, recorded next to the image tag in the Compose file rather than - * here. {@code @targeting} is the newer claim and the cheaper one to check: its three scenarios - * resolve {@code targeting-key-flag} through flagd's own rule evaluation and all three pass in - * both modes on the image already pinned, so nothing about it needed a testbed bump. - * - *

{@link Capability#DISABLED_FLAGS} is declared, and measured rather than assumed. Both - * resolvers substitute the caller's default for a flag whose state is {@code DISABLED} and report - * no error code, so all four rows of that outline pass in both modes. Measured on the pinned - * image: 65 scenarios in each mode, 59 passing, two skipped — one for the withheld - * {@code @reinitialization}, one because the SDK cannot ask {@code @large-integers} — and four - * failing — the one real defect plus the three testbed gaps already described. The cold-start - * initialisation error recorded on {@link #CONNECTED_DEADLINE_MS} did not reproduce in the run - * these numbers come from; it is intermittent and host-dependent, and when it appears in-process - * it costs one further scenario. Nothing needed bumping for it either: {@code disabled-boolean-flag}, - * {@code disabled-string-flag}, {@code disabled-integer-flag} and {@code disabled-float-flag} - * are flagd-testbed's own, from {@code flags/disabled-flags.json}, which the image already - * carries at the v3.8.0 pinned below and which the launchpad combines into the set it serves. The - * canonical definition took the testbed's names and values rather than inventing its own, exactly - * as it did for the falsy flags. - * - *

That the capability holds here is worth stating rather than assuming, because the tag is - * gated on architecture rather than on quality and flagd sits on the right side of that line - * twice over: the in-process resolver evaluates the ruleset locally, and the RPC resolver still - * decides locally what to do with a response that carries no value. A provider whose backend - * decides — one speaking OFREP — cannot hold it at all, which is the comparison the tag exists to - * make legible. - * - *

{@link Capability#STANDARD_REASONS} is declared, and it arrived here by the - * {@code declarableExcept} default rather than by a decision — which is exactly why it was - * measured before this paragraph was written. All nine scenarios of {@code reason.feature} pass - * in both modes, including the two that compose with {@link Capability#TARGETING} and - * {@link Capability#DISABLED_FLAGS}: flagd reports {@code STATIC} for the rule-less flags, - * {@code TARGETING_MATCH} and {@code DEFAULT} either side of {@code targeting-key-flag}'s rule, - * {@code DISABLED} for a disabled flag, and {@code ERROR} beside {@code FLAG_NOT_FOUND} and - * {@code TYPE_MISMATCH}. So the claim the tag makes — the standard vocabulary with the standard - * meanings — holds for both resolvers, and the declaration is evidence rather than inheritance. - * - *

{@link Capability#declarableExcept} rather than {@code EnumSet.complementOf}, which is what - * this used to be. The complement of one capability is every other enum constant, - * including {@code @caching} — a reserved tag no scenario carries — so declaring the complement - * claimed a capability nothing had examined, and the suite refuses such a declaration at startup. - * It swept up {@code @targeting} the same way until that tag gated something, which is the point: - * the hazard shrinks as the vocabulary fills up and never disappears, so the form of the call is - * what protects the declaration, not the current size of the reserved set. + * period turns it into {@code PROVIDER_ERROR} — so the emit sits in the provider layer and not in + * a transport, and "Losing the backend makes the provider stale, regaining it makes it ready + * again" passes in RPC mode as well as in-process. + * + *

{@link Capability#VARIANTS} and {@link Capability#TARGETING} are both declared. flagd names + * the variant it served in every resolution, so seven of the {@code @variants} outline's eight + * rows pass in both modes; the eighth is one of the testbed gaps above. {@code @targeting}'s + * three scenarios resolve {@code targeting-key-flag} through flagd's own rule evaluation and all + * three pass in both modes on the image already pinned. + * + *

{@link Capability#DISABLED_FLAGS} is declared, and measured: both resolvers substitute the + * caller's default for a flag whose state is {@code DISABLED} and report no error code, so all + * four rows of that outline pass in both modes. The tag is gated on architecture rather than on + * quality, and flagd sits on the right side of that line twice over — the in-process resolver + * evaluates the ruleset locally, and the RPC resolver still decides locally what to do with a + * response that carries no value. + * + *

{@link Capability#STANDARD_REASONS} arrived by the {@code declarableExcept} default rather + * than by a decision, which is why it was measured before being written down. All nine scenarios + * of {@code reason.feature} pass in both modes, including the two composing with + * {@link Capability#TARGETING} and {@link Capability#DISABLED_FLAGS}: {@code STATIC} for the + * rule-less flags, {@code TARGETING_MATCH} and {@code DEFAULT} either side of + * {@code targeting-key-flag}'s rule, {@code DISABLED} for a disabled flag, and {@code ERROR} + * beside {@code FLAG_NOT_FOUND} and {@code TYPE_MISMATCH}. + * + *

{@link Capability#declarableExcept} and not {@code EnumSet.complementOf}: the complement of + * one capability is every other enum constant, reserved tags included, and the suite + * refuses such a declaration at startup. */ @Override public Set capabilities() { @@ -273,42 +197,21 @@ public Set capabilities() { /** * {@inheritDoc} * - *

One entry, for {@link Capability#NUMERIC_COERCION}, and it takes the declared and - * failing shape rather than the withheld-and-skipped one. That is the shape the TCK's - * guidance prefers, and the reason it prefers it is exactly this case: flagd does - * attempt the coercion — the widening scenario passes — and gets the narrowing direction wrong, - * so withdrawing the capability would turn a real failure into a skip, which is the failure mode - * the field exists to prevent. The failure stays visible in the results and this entry says it is - * known and why. - * - *

An earlier version of this file withheld the tag instead, on the argument that the - * capability requires all three of its scenarios and one of them asks for - * {@code integral-float-flag}, which the pinned testbed image does not serve — so declaring it - * buys one honest failure and one that is the stack's fault. That cost is real and it is - * accepted, for two reasons. It is not a new kind of cost: this branch already carries two - * failures caused by the same missing flags and records them plainly rather than hiding them - * behind a withheld capability. And the alternative is worse, because a skip cannot distinguish - * "flagd declines to coerce" from "flagd coerces and gets it wrong", and only the second is true. + *

One entry, for {@link Capability#NUMERIC_COERCION}, taking the declared and + * failing shape. Appendix F's declaring rule decides it, and both halves apply here in + * order: flagd is attempting the coercion, which the widening scenario proves, and two + * of the tag's three scenarios can be put to it. A provider that simply does not coerce would + * stop at the first half and withhold, which is what the SDK's in-memory provider does and why + * the self-tests in {@code tools/tck} skip these scenarios. * - *

That argument is no longer this file's to make. - * Appendix - * F now states it as a declaring rule — once a provider is attempting a capability, the unit - * of the decision is the scenario rather than the tag, so it is declared when at least - * one scenario gating it can be put to the provider and withheld only when none can. Both halves - * apply here in order: flagd is attempting the coercion, which the widening scenario - * proves, and two of the tag's three scenarios can be put to it. A provider that simply does not - * coerce would stop at the first half and withhold, which is what the SDK's in-memory provider - * does and why the self-tests in {@code tools/tck} skip these scenarios. The appendix's other - * consequence is why the summary below names the testbed: a scenario that fails because the - * backend cannot serve its fixture is not a provider defect, and a deviation that did not say so - * would have the report accuse flagd of the stack's gap. + *

Declaring it costs one failure that is the stack's rather than flagd's, since + * {@code integral-float-flag} is absent from the pinned image. Accepted, and the summary below + * says so explicitly — a deviation that did not would have the report accuse flagd of the gap. * *

Tracked against flagd's numeric coercion ADR, which is where the rule this deviates from is - * settled: coercion is permitted when it is lossless and must fail with {@code TYPE_MISMATCH} - * only when information would be lost. The summary says which half is broken, because "flagd - * coerces numbers" on its own reads as a description of intended behaviour. Delete the entry once - * the lossy case reports {@code TYPE_MISMATCH}; the capability needs no change then, which is - * another small argument for this shape. + * settled. The summary names which half is broken, because "flagd coerces numbers" on its own + * reads as a description of intended behaviour. Delete the entry once the lossy case reports + * {@code TYPE_MISMATCH}; the capability needs no change then. */ @Override public List knownDeviations() { diff --git a/providers/flagd/src/test/resources/tck/docker-compose.yaml b/providers/flagd/src/test/resources/tck/docker-compose.yaml deleted file mode 100644 index f1d7b4fadf..0000000000 --- a/providers/flagd/src/test/resources/tck/docker-compose.yaml +++ /dev/null @@ -1,36 +0,0 @@ -# Backend stack for the OpenFeature Provider TCK, wrapping the unmodified flagd testbed image. -# -# The image serves flagd itself and the "launchpad" control API on 8080, whose endpoints this -# TCK's control API contract was derived from. It does not yet serve the whole canonical flag -# set: v3.8.0 has none of large-integer-flag, huge-integer-flag or integral-float-flag. Two of -# the three are reached by scenarios that run here -- huge-integer-flag sits behind -# @large-integers, which no Java provider can declare because Client.getIntegerDetails is a -# 32-bit Integer, so its scenario is skipped before the missing flag can matter -- and between -# them they fail three scenarios, all until open-feature/flagd-testbed#392 lands. Bump the tag -# then. -# -# * large-integer-flag: the untagged precision scenario, which gets the code default instead -# of 2147483647, and the @variants row asking for its "max-int32" variant, answered with -# none. -# * integral-float-flag: the @numeric-coercion scenario "An integral float requested as an -# integer is coerced without loss". This provider declares @numeric-coercion because flagd -# does coerce and gets the lossy direction wrong, so the tag's failure is worth seeing -- -# see AbstractResolverTest.knownDeviations(). This third failure is the price of that, and -# it is the same missing-flag problem as the two above rather than a new one. -# -# targeting-key-flag is served, which is why @targeting needs no bump: its three scenarios pass -# on this image in both modes. The same goes for the four disabled-* flags behind -# @disabled-flags: they are this image's own, from its flags/disabled-flags.json, which the -# launchpad combines into the served set -- the canonical definition took the testbed's names and -# values rather than inventing its own. -# -# Note there are no host port bindings. The TCK requires dynamically mapped ports and discovers -# them after startup — a pinned host port would make the suite unrunnable in parallel and would -# collide with a developer's local flagd. -services: - backend: - image: ghcr.io/open-feature/flagd-testbed:v3.8.0 - ports: - - 8013 # flagd RPC evaluation (gRPC) - - 8015 # flagd in-process sync (gRPC) - - 8080 # launchpad control API diff --git a/tools/flagd-testbed/docker-compose.yaml b/tools/flagd-testbed/docker-compose.yaml new file mode 100644 index 0000000000..ce39fe833c --- /dev/null +++ b/tools/flagd-testbed/docker-compose.yaml @@ -0,0 +1,21 @@ +# The flagd-testbed stack every Provider TCK adoption in this repository runs against, named from +# each suite's ContainerizedProviderTckTest.composeFile(). Shared so the adoptions cannot drift onto +# two backends while claiming to ask the same questions. +# +# NOT providers/flagd/test-harness/docker-compose.yaml, which the legacy e2e suites drive: that one +# bind-mounts ${FLAGS_DIR}, names its service "flagd" and runs an envoy sidecar (e2e/ContainerEntry). +# +# Every port either adoption uses is listed, and the extras are free -- with no host bindings the +# TCK maps each container port dynamically and resolves only the ones a suite asks for. Do not add +# a host binding; Appendix F says why. +# +# The image does not serve the whole canonical flag set yet, open-feature/flagd-testbed#392; what +# that costs is recorded in each adoption's suite class. +services: + backend: + image: ghcr.io/open-feature/flagd-testbed:v3.8.0 + ports: + - 8013 # flagd RPC evaluation (gRPC) -- providers/flagd + - 8015 # flagd in-process sync (gRPC) -- providers/flagd + - 8016 # flagd OFREP evaluation (HTTP) -- providers/ofrep + - 8080 # launchpad control API -- both From 950a4692a0f548ca82e34e4aa31de401f7c96640 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Mon, 14 Sep 2026 09:00:39 +0200 Subject: [PATCH 21/25] test(flagd): follow @disabled-flags' corrected gating question Capability.DISABLED_FLAGS no longer frames the gate as where evaluation happens, so the sentence here that echoed it says what was actually measured instead: both resolvers are told the flag is disabled. Comments only. Signed-off-by: Simon Schrottner --- .../contrib/providers/flagd/tck/AbstractResolverTest.java | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java index 85f17b5726..0c4764a045 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java @@ -172,10 +172,10 @@ public FeatureProvider createUnavailableProvider() { * *

{@link Capability#DISABLED_FLAGS} is declared, and measured: both resolvers substitute the * caller's default for a flag whose state is {@code DISABLED} and report no error code, so all - * four rows of that outline pass in both modes. The tag is gated on architecture rather than on - * quality, and flagd sits on the right side of that line twice over — the in-process resolver - * evaluates the ruleset locally, and the RPC resolver still decides locally what to do with a - * response that carries no value. + * four rows of that outline pass in both modes. Both resolvers are told the flag is disabled — + * the in-process one evaluates the ruleset locally, and the RPC one still decides locally what + * to do with a response that carries no value — so the question {@link Capability#DISABLED_FLAGS} + * gates on is answered here in both. * *

{@link Capability#STANDARD_REASONS} arrived by the {@code declarableExcept} default rather * than by a decision, which is why it was measured before being written down. All nine scenarios From 50305f8a1fee92b8db4ac3b0324245d6aa48ae0f Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Mon, 14 Sep 2026 09:24:15 +0200 Subject: [PATCH 22/25] test(flagd): record the intermittent testbed failure observed here too Measured this pass: an RPC run came back with a fifth failure, FLAG_NOT_FOUND on an evaluation.feature row expecting no error code, and the next run of the same tree was clean. Same shape the OFREP adoption already records against open-feature/flagd-testbed#394, so it is named here rather than left for the next reader to diagnose as a regression. Comments only. Signed-off-by: Simon Schrottner --- .../contrib/providers/flagd/tck/AbstractResolverTest.java | 7 +++++++ 1 file changed, 7 insertions(+) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java index 0c4764a045..37ad4c23e9 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java @@ -31,6 +31,13 @@ * scenario "An integral float requested as an integer is coerced without loss". The third is the * cost of declaring {@code @numeric-coercion} — see {@link #knownDeviations()}. None of the three is * a {@link KnownDeviation}: the provider was never given the flag to get wrong. + * + *

A run occasionally carries one failure beyond those, and it is the stack's. + * Observed in RPC mode as {@code FLAG_NOT_FOUND} on an {@code evaluation.feature} row that expected + * no error code; the next run of the same tree was clean. That is the flagd-testbed readiness window + * of open-feature/flagd-testbed#394, the same intermittency the OFREP adoption records, and it is + * not covered with a sleep here either. Repeat a run before treating an extra failure as a + * regression — anything that reproduces is one. */ abstract class AbstractResolverTest extends ContainerizedProviderTckTest { From 7b0500ee2f0e8129b3ddab916bcc33e84c2b1fc5 Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Mon, 14 Sep 2026 10:03:47 +0200 Subject: [PATCH 23/25] test(tck): run the conformance suite against flagd-testbed v3.10.1 Four languages pinned this image by hand and had drifted: JavaScript was already on v3.10.1 while Go, Java and Python sat on v3.8.0, so the suites whose results are only comparable if they asked the same backend were asking two. v3.10.1 is the current release, so aligning up rather than down. What this does NOT fix, stated because the tag is easy to mistake for a fix: open-feature/flagd-testbed#392 and #394 are both still open, so v3.10.1 carries neither the three missing precision flags nor the /start readiness fix. The fixture failures and the readiness race are unchanged. The one behavioural change in range is open-feature/flagd-testbed#390, which increases the simulated downtime -- and that is exactly the timing the @stale and @unavailable scenarios depend on. Tallies recorded against v3.8.0 have not been re-measured on this image, including the failure ranges the adoption READMEs cite by tag; those figures stay as they are because they are a record of what v3.8.0 did, and re-taking them is follow-up work rather than a rewrite. Signed-off-by: Simon Schrottner --- tools/flagd-testbed/docker-compose.yaml | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tools/flagd-testbed/docker-compose.yaml b/tools/flagd-testbed/docker-compose.yaml index ce39fe833c..da18c55d2c 100644 --- a/tools/flagd-testbed/docker-compose.yaml +++ b/tools/flagd-testbed/docker-compose.yaml @@ -13,7 +13,7 @@ # that costs is recorded in each adoption's suite class. services: backend: - image: ghcr.io/open-feature/flagd-testbed:v3.8.0 + image: ghcr.io/open-feature/flagd-testbed:v3.10.1 ports: - 8013 # flagd RPC evaluation (gRPC) -- providers/flagd - 8015 # flagd in-process sync (gRPC) -- providers/flagd From 4a66284fc45eebfb96e88c39aa2318740ff97baf Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Tue, 15 Sep 2026 22:27:06 +0200 Subject: [PATCH 24/25] test(flagd): declare @string-typing, on a run rather than on the default The four scenarios the new tag gates were mandatory and passing before it existed, so the "everything except @reinitialization" default already declares it. Recorded here anyway, because this file's standard is that each declaration rests on evidence from a run and not on inheriting the default. Measured in both resolver modes: 65 scenarios, 2 skipped, and the same four failures as before -- the lossy numeric coercion (open-feature/flagd#1996) and the three flags the pinned testbed image does not serve (open-feature/flagd-testbed#392). None of the four @string-typing scenarios is among them. flagd's flag definitions carry a JSON type per flag and both resolvers preserve it, so a non-string flag asked through the String accessor is a real mismatch here and is reported as one. RPC additionally showed the intermittent "half" variant failure this branch already records -- surefire's reruns had it pass four times in five, which is the testbed readiness window of open-feature/flagd-testbed#394 and not a property of any assertion. The clean-run tally is unchanged. Signed-off-by: Simon Schrottner --- .../providers/flagd/tck/AbstractResolverTest.java | 12 +++++++++++- 1 file changed, 11 insertions(+), 1 deletion(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java index 37ad4c23e9..d2109bd7d6 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java @@ -145,7 +145,17 @@ public FeatureProvider createUnavailableProvider() { * declines to coerce. The lossy scenario fails; the remaining lossless one fails only on the * absent {@code integral-float-flag}. Both resolvers behave identically, which places the defect * in the shared provider layer rather than in either transport — as does every other capability - * here, including the full non-numeric type-mismatch matrix. + * here, including the whole remaining type-mismatch matrix. + * + *

{@link Capability#STRING_TYPING} is declared, and all four of its scenarios + * pass in both modes. It gates what specification revision {@code d47a66eb} moved out + * of the mandatory matrix: {@code boolean-flag}, {@code integer-flag}, {@code float-flag} and + * {@code object-flag} asked through the String accessor. flagd's flag definitions carry a JSON + * type per flag and both resolvers preserve it, so a non-string flag requested as a string is a + * genuine mismatch here and is reported as one — which is exactly the position the capability + * exists to distinguish from a backend that stores every value as a string. Declared on the run + * rather than on the "everything except" default: these four were mandatory and passing before + * the tag existed, and the numbers below are unchanged by the move. * *

{@link Capability#LIFECYCLE} is declared because flagd reaches its backend during * initialisation in both modes — an RPC round trip, or a full ruleset sync — so the lifecycle From 5d439f9d8b0ca5be13184558c95cc2b5d3be329c Mon Sep 17 00:00:00 2001 From: Simon Schrottner Date: Wed, 16 Sep 2026 09:20:22 +0200 Subject: [PATCH 25/25] test(flagd): declare @fully-typed-values too, on a run in both modes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Appendix F split @string-typing at bda599f1: the outline keeps boolean-flag and integer-flag, and float-flag and object-flag moved behind a new @fully-typed-values that asks whether the store records a native type for those two as well. flagd is the case the split was not written for, and declaring both is how that shows. Its flag definitions carry a JSON type per flag and both resolvers preserve it, so all four questions have the same answer here. A partially typed backend declares the first tag and withholds the second; there is nothing partial about this one. declarableExcept(REINITIALIZATION) already picks the new tag up, so this is javadoc rather than a declaration change — but the tag was measured rather than inherited, which is this file's standing rule. Both modes after the re-pin: 65 scenarios, 59 passing, 2 skipped, 4 failing, identical to the run before it. The two skips are still the withheld @reinitialization and @large-integers; the surefire report has "A float flag is not returned as its string representation" and "A structured flag is not returned as its JSON text" as executed and passing, which is the part a tally alone would not have shown — an undeclared @fully-typed-values would have turned both into skips and left the failure count untouched. The four failures are unchanged and none is a string-typing scenario: the lossy numeric coercion (open-feature/flagd#1996) and the three flags the pinned testbed image does not serve (open-feature/flagd-testbed#392). Signed-off-by: Simon Schrottner --- .../flagd/tck/AbstractResolverTest.java | 26 ++++++++++++------- 1 file changed, 17 insertions(+), 9 deletions(-) diff --git a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java index d2109bd7d6..9633534e6a 100644 --- a/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java +++ b/providers/flagd/src/test/java/dev/openfeature/contrib/providers/flagd/tck/AbstractResolverTest.java @@ -147,15 +147,23 @@ public FeatureProvider createUnavailableProvider() { * in the shared provider layer rather than in either transport — as does every other capability * here, including the whole remaining type-mismatch matrix. * - *

{@link Capability#STRING_TYPING} is declared, and all four of its scenarios - * pass in both modes. It gates what specification revision {@code d47a66eb} moved out - * of the mandatory matrix: {@code boolean-flag}, {@code integer-flag}, {@code float-flag} and - * {@code object-flag} asked through the String accessor. flagd's flag definitions carry a JSON - * type per flag and both resolvers preserve it, so a non-string flag requested as a string is a - * genuine mismatch here and is reported as one — which is exactly the position the capability - * exists to distinguish from a backend that stores every value as a string. Declared on the run - * rather than on the "everything except" default: these four were mandatory and passing before - * the tag existed, and the numbers below are unchanged by the move. + *

{@link Capability#STRING_TYPING} and {@link Capability#FULLY_TYPED_VALUES} are both + * declared, and all four of their scenarios pass in both modes. Together they gate what + * specification revision {@code d47a66eb} moved out of the mandatory matrix: {@code boolean-flag} + * and {@code integer-flag} asked through the String accessor under the first tag, and + * {@code float-flag} and {@code object-flag} under both tags since {@code bda599f1} split them + * apart. flagd's flag definitions carry a JSON type per flag and both resolvers preserve it, so a + * non-string flag requested as a string is a genuine mismatch here and is reported as one — + * which is exactly the position these capabilities exist to distinguish from a backend that + * stores every value as a string. + * + *

flagd is the case the split was not written for, and declaring both is how that + * shows: the question {@code @fully-typed-values} asks separately — does the store record a + * native type for a float and for a structure — flagd answers yes to, just as it does for a + * boolean and an integer. A partially typed backend declares the first and withholds the second; + * there is nothing partial here. Declared on the run rather than on the "everything except" + * default, which would have swept the new tag up unexamined: the four scenarios were measured in + * both modes after the re-pin, and the numbers below are unchanged by the split. * *

{@link Capability#LIFECYCLE} is declared because flagd reaches its backend during * initialisation in both modes — an RPC round trip, or a full ruleset sync — so the lifecycle