Skip to content

[Bug][client] V5 ConnectionPolicy.connectionBackoff is ignored #26788

Description

@lhotari

Search before reporting

  • I searched in the issues and found nothing similar.

Read release policy

  • I understand that unsupported versions don't get bug fixes. I will attempt to reproduce the issue on a supported version of Pulsar client and Pulsar broker.

User environment

  • Pulsar master (5.0.0 development), Java V5 client (pulsar-client-v5)

Issue Description

ConnectionPolicy.Builder.connectionBackoff(BackoffPolicy) in the V5 client API is documented as the "Backoff strategy for broker reconnection attempts", but the value is silently ignored. PulsarClientBuilderV5.connectionPolicy(...) copies every other ConnectionPolicy field into the underlying v4 ClientConfigurationData but skips the backoff:

// PulsarClientBuilderV5.connectionPolicy(...)
        // BackoffPolicy adaptation will be implemented when the v4 client exposes
        // a public way to override the reconnection backoff.
        return this;

As a result, producers, consumers and the other connection handlers always reconnect with the v4 defaults (initial 100 ms, max 60 s), whatever the application configures.

The reason given in the comment no longer holds. The V5 builder writes directly to ClientConfigurationData, which already has initialBackoffIntervalNanos and maxBackoffIntervalNanos, and the v4 ClientBuilder exposes them publicly as startingBackoffInterval(...) and maxBackoffInterval(...). ProducerImpl, ConsumerImpl (via PulsarClientImpl) and TransactionMetaStoreHandler build their Backoff from those two fields.

Expected: BackoffPolicy.initialInterval() and maxInterval() are applied to the reconnection backoff.

BackoffPolicy also has multiplier() and jitterPercent(), which have no counterpart in ClientConfigurationData (the v4 Backoff doubles on each attempt and takes its jitter from the Backoff builder). Either those need to be carried through to where the Backoff instances are built, or it should be documented or validated that only the defaults are supported for now, rather than being silently dropped.

Error messages

None; the setting is silently ignored.

Reproducing the issue

  1. Build a V5 client with a non-default reconnection backoff:
    PulsarClient client = PulsarClient.builder()
            .serviceUrl(serviceUrl)
            .connectionPolicy(ConnectionPolicy.builder()
                    .connectionBackoff(BackoffPolicy.exponential(Duration.ofSeconds(5), Duration.ofSeconds(30)))
                    .build())
            .build();
  2. Create a consumer and have the broker disconnect it (for example by unloading the topic).
  3. ConnectionHandler.connectionClosed schedules the reconnect from backoff.next(), and that Backoff was built from the unchanged ClientConfigurationData defaults. So the client logs Closed connection - Will try again with a delaySec of about 0.1 (the 100 ms default), not 5.

Additional information

Found while fixing a flaky test, V5CumulativeAckTest.testRetryingAnAckThatFailedWhileDisconnectedAdvancesTheCursor. A longer reconnection backoff on a dedicated client would have been the simplest way to keep the consumer off the broker for the duration of an ack, but it had no effect because of this issue.

Are you willing to submit a PR?

  • I'm willing to submit a PR!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    type/bugThe PR fixed a bug or issue reported a bug

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions