Summary
QueryKillingIntegrationTest.testCpuBasedServerQueryKilling fails intermittently on the multi-stage engine path: the server kills the query for CPU time as expected, but the broker surfaces error code 503 (QUERY_CANCELLATION) instead of 245 (SERVER_RESOURCE_LIMIT_EXCEEDED).
Failure signature
java.lang.AssertionError: Unexpected error code: 503 from exception: {"message":"Received 1 error from stage 1 on Server_localhost_27001: Cancelled by sender with exception: Error block from stage 1 worker 0 on Server_localhost_27001. Msg: {SERVER_RESOURCE_LIMIT_EXCEEDED=CPU time based killed on SERVER ...
at org.apache.pinot.integration.tests.QueryKillingIntegrationTest.verifyCpuTimeKill(QueryKillingIntegrationTest.java:462)
at org.apache.pinot.integration.tests.QueryKillingIntegrationTest.testCpuBasedServerQueryKilling(QueryKillingIntegrationTest.java:282)
Occurrences
Scanning the 60 most recent failed Pinot Integration Tests runs (2026-09-06 to 2026-09-26) found only these two, both with the identical signature. Neither branch touches query killing, resource accounting, or the query runtime.
Likely cause
The resource-limit error block from the stage worker and the mailbox cancellation race on the way to the broker. GrpcSendingMailbox / InMemorySendingMailbox emit QueryErrorCode.QUERY_CANCELLATION with the message "Cancelled by sender with exception: " + msg, wrapping the worker's error block, so the original code 245 is lost and only survives inside the message text. Whichever path reaches the broker first determines the reported code.
Suggested direction
Preserve the wrapped error's QueryErrorCode when a sending mailbox cancels because of a downstream error block (or have the broker-side merge prefer the underlying resource-limit code over the cancellation wrapper), rather than loosening the test's assertion. Per the project's review principles, adding retries or relaxing the check is not the fix.
Summary
QueryKillingIntegrationTest.testCpuBasedServerQueryKillingfails intermittently on the multi-stage engine path: the server kills the query for CPU time as expected, but the broker surfaces error code 503 (QUERY_CANCELLATION) instead of 245 (SERVER_RESOURCE_LIMIT_EXCEEDED).Failure signature
Occurrences
cbo-4-planner-wiring, run 35118594082 (stage 2 worker)f9d64fb057, run 36208548223 attempt 1 (stage 1 worker)Scanning the 60 most recent failed
Pinot Integration Testsruns (2026-09-06 to 2026-09-26) found only these two, both with the identical signature. Neither branch touches query killing, resource accounting, or the query runtime.Likely cause
The resource-limit error block from the stage worker and the mailbox cancellation race on the way to the broker.
GrpcSendingMailbox/InMemorySendingMailboxemitQueryErrorCode.QUERY_CANCELLATIONwith the message"Cancelled by sender with exception: " + msg, wrapping the worker's error block, so the original code 245 is lost and only survives inside the message text. Whichever path reaches the broker first determines the reported code.Suggested direction
Preserve the wrapped error's
QueryErrorCodewhen a sending mailbox cancels because of a downstream error block (or have the broker-side merge prefer the underlying resource-limit code over the cancellation wrapper), rather than loosening the test's assertion. Per the project's review principles, adding retries or relaxing the check is not the fix.