Proposal
wait_for_workflow_completion and wait_for_workflow_start (sync and aio) should re-issue WaitForInstanceCompletion / WaitForInstanceStart when the server ends the call with CANCELLED and the caller's own timeout hasn't expired. Today that status reaches the caller as an error, even though the workflow is still running. This asks maintainers to decide on the behaviour first, because it reverses assertions added in #1112. It's a proposal, not a PR.
Why the server sends CANCELLED
When daprd shuts down or restarts mid-wait (rollout, pod restart, hot-reload, a config change that restarts the runtime), the actors router cancels every in-flight call it tracks. The wait then returns CANCELLED "context canceled", even though the caller is still waiting. The mechanism and a repro are in dapr/dapr#10566.
Reproduced against native daprd 1.18.0:
- a 60 s workflow with
wait_for_workflow_completion(id) and no timeout returns normally;
- the same run with daprd sent
SIGTERM about 10 s in fails with StatusCode.CANCELLED "context canceled".
Why retrying is safe
- A blocking unary call only gets
CANCELLED from the server. The sync client can't cancel it mid-flight, and in the aio client a caller's cancellation raises asyncio.CancelledError, not AioRpcError. A client-side deadline shows up as DEADLINE_EXCEEDED, which already maps to TimeoutError.
- The wait only reads state, so re-issuing it is idempotent. It returns straight away if the workflow already finished.
What would change
Relationship to the runtime fix
The proper fix is in daprd: return UNAVAILABLE in this case (dapr/dapr#10566). The SDK already retries UNAVAILABLE. This change would protect users on current and older runtimes until that ships, and it would still help afterwards with proxies that reset the stream.
Question for maintainers
Is it acceptable to treat a server-sent CANCELLED on these two waits as retryable, and change the #1112 tests? If yes, it's a small change in both clients plus tests.
Proposal
wait_for_workflow_completionandwait_for_workflow_start(sync and aio) should re-issueWaitForInstanceCompletion/WaitForInstanceStartwhen the server ends the call withCANCELLEDand the caller's own timeout hasn't expired. Today that status reaches the caller as an error, even though the workflow is still running. This asks maintainers to decide on the behaviour first, because it reverses assertions added in #1112. It's a proposal, not a PR.Why the server sends CANCELLED
When daprd shuts down or restarts mid-wait (rollout, pod restart, hot-reload, a config change that restarts the runtime), the actors router cancels every in-flight call it tracks. The wait then returns
CANCELLED"context canceled", even though the caller is still waiting. The mechanism and a repro are in dapr/dapr#10566.Reproduced against native daprd 1.18.0:
wait_for_workflow_completion(id)and no timeout returns normally;SIGTERMabout 10 s in fails withStatusCode.CANCELLED"context canceled".Why retrying is safe
CANCELLEDfrom the server. The sync client can't cancel it mid-flight, and in the aio client a caller's cancellation raisesasyncio.CancelledError, notAioRpcError. A client-side deadline shows up asDEADLINE_EXCEEDED, which already maps toTimeoutError.What would change
CANCELLEDto the retried codes, when the caller's timeout hasn't expired, in_durabletask/client.pyL335 and_durabletask/aio/client.pyL230._is_deadline_cancellation(L349) already turns aCANCELLEDafter the deadline intoTimeoutError.CANCELLEDreaches the caller, and would flip:test_cancelled_within_deadline_still_propagatestest_cancelled_unbounded_wait_still_propagatestest_cancelled_within_deadline_still_propagates(aio)Relationship to the runtime fix
The proper fix is in daprd: return
UNAVAILABLEin this case (dapr/dapr#10566). The SDK already retriesUNAVAILABLE. This change would protect users on current and older runtimes until that ships, and it would still help afterwards with proxies that reset the stream.Question for maintainers
Is it acceptable to treat a server-sent
CANCELLEDon these two waits as retryable, and change the #1112 tests? If yes, it's a small change in both clients plus tests.