Expected Behavior
A subscription made with subscribe_with_handler (sync or async) keeps delivering messages to the handler until the caller closes it. If the stream drops, it reconnects, and it keeps retrying while the sidecar is unavailable. An exception raised by the handler does not tear down the stream. The returned close function stops the handler before it returns.
Actual Behavior
Links are to main at 03eebe1.
Async client: dapr/aio/clients/grpc/client.py#L580-L599
- No reference is kept to the task. It calls
asyncio.create_task(stream_messages(subscription)) and discards the result. Python's docs warn that the event loop keeps only a weak reference to such a task, so it can be garbage-collected while it is still running. The subscription then stops with no error.
- Any error ends the subscription. The loop catches only
StreamInactiveError. If the handler raises, or the stream raises StreamCancelledError or a non-retryable gRPC error, the task dies. Nothing reconnects, and the error only appears later as an "exception was never retrieved" warning. The sync client reconnects in that case.
- The close function doesn't wait for the task.
close_subscription() closes the subscription but never awaits the task, so the handler can still be running after await close_fn() returns.
- No test covers it.
test_subscribe_topic_with_handler is commented out: tests/clients/test_dapr_grpc_client_async.py#L424-L477.
Sync client: dapr/clients/grpc/client.py#L628-L654
-
The "reconnect failed, back off and retry" branch never retries.
- On a stream error the loop calls
sub.reconnect_stream().
reconnect_stream() marks the stream inactive first. Then it waits for the sidecar and calls start().
- If that wait or
start() raises, the loop sleeps 5 seconds and runs continue.
- The next
for message in sub calls next_message(). The stream is still inactive, so that raises StreamInactiveError, and the loop treats it as a close and breaks.
So a sidecar outage longer than the health wait ends the subscription for good, with no error.
-
Handler errors reconnect a healthy stream. The same except Exception also catches exceptions raised by handler_fn. One failing handler call tears down and reconnects the stream, and the message is never acked, so it is delivered again.
Steps to Reproduce the Problem
- 5: stop the sidecar while a
subscribe_with_handler subscription is running. Keep it down for longer than DaprHealth.wait_for_sidecar waits, then start it again. The handler thread has exited, and no new messages reach the handler.
- 6: have the handler raise for one message. The log shows a stream reconnect, and the message is delivered again.
- 2: do the same with the async client. The task ends, and no further messages are handled.
Suggested fix
Build on #1231, which makes close() final: a subscription that has been closed can no longer be reopened by a reconnect.
- Sync (5): after a failed reconnect, retry the reconnect with the backoff. Leave the loop only when the subscription was closed by the caller.
- Both (6): treat a handler exception separately from a stream error. Log it and respond with retry, instead of reconnecting the stream.
- Async (1–3):
- Keep a reference to the task.
- Handle
StreamCancelledError and stream errors the way the sync client does.
- Have the close function await the task, with a timeout.
- Tests: restore the async handler test. Add tests for both clients: a reconnect that fails and then succeeds, and a handler that raises once.
Release Note
RELEASE NOTE: FIX subscribe_with_handler keeps retrying while the sidecar is unavailable, does not reconnect on handler errors, and the async version no longer stops silently.
Expected Behavior
A subscription made with
subscribe_with_handler(sync or async) keeps delivering messages to the handler until the caller closes it. If the stream drops, it reconnects, and it keeps retrying while the sidecar is unavailable. An exception raised by the handler does not tear down the stream. The returned close function stops the handler before it returns.Actual Behavior
Links are to
mainat03eebe1.Async client:
dapr/aio/clients/grpc/client.py#L580-L599asyncio.create_task(stream_messages(subscription))and discards the result. Python's docs warn that the event loop keeps only a weak reference to such a task, so it can be garbage-collected while it is still running. The subscription then stops with no error.StreamInactiveError. If the handler raises, or the stream raisesStreamCancelledErroror a non-retryable gRPC error, the task dies. Nothing reconnects, and the error only appears later as an "exception was never retrieved" warning. The sync client reconnects in that case.close_subscription()closes the subscription but never awaits the task, so the handler can still be running afterawait close_fn()returns.test_subscribe_topic_with_handleris commented out:tests/clients/test_dapr_grpc_client_async.py#L424-L477.Sync client:
dapr/clients/grpc/client.py#L628-L654The "reconnect failed, back off and retry" branch never retries.
sub.reconnect_stream().reconnect_stream()marks the stream inactive first. Then it waits for the sidecar and callsstart().start()raises, the loop sleeps 5 seconds and runscontinue.for message in subcallsnext_message(). The stream is still inactive, so that raisesStreamInactiveError, and the loop treats it as a close andbreaks.So a sidecar outage longer than the health wait ends the subscription for good, with no error.
Handler errors reconnect a healthy stream. The same
except Exceptionalso catches exceptions raised byhandler_fn. One failing handler call tears down and reconnects the stream, and the message is never acked, so it is delivered again.Steps to Reproduce the Problem
subscribe_with_handlersubscription is running. Keep it down for longer thanDaprHealth.wait_for_sidecarwaits, then start it again. The handler thread has exited, and no new messages reach the handler.Suggested fix
Build on #1231, which makes
close()final: a subscription that has been closed can no longer be reopened by a reconnect.StreamCancelledErrorand stream errors the way the sync client does.Release Note
RELEASE NOTE: FIX
subscribe_with_handlerkeeps retrying while the sidecar is unavailable, does not reconnect on handler errors, and the async version no longer stops silently.