Bug Report
1. Minimal reproduce step (Required)
Any MPP query in which an exchange consumer task legitimately finishes before
its upstream producer tasks have finished sending, while the producers still
have packets in flight toward it.
Known triggers:
The failure itself is a timing race: it requires a producer write to land
after the consumer has already closed. On an unmodified build it occurs with
low probability; it can be reproduced deterministically with a small
delay-injection patch that stretches the producer write window.
2. What did you expect to see? (Required)
The query succeeds. A consumer legitimately finishing early — because of
LIMIT, an empty build side, a filter that eliminates all rows, etc. — is a
fact that upstream producers must tolerate, not an error. In-flight packets
that arrive after the consumer has finished should be discarded silently,
and the producer tasks should finish normally.
3. What did you see instead (Required)
The query fails with err msg like:
ERROR 1105 (HY000): write to tunnel tunnel13+21 which is already closed, tunnel13+21: unexpectedWriteDone called
The correct query result is lost. Root cause:
- A consumer task legitimately finishes early and closes its
ExchangeReceiver / MPPTunnel while upstream producers are still sending.
- The MPP exchange layer has no graceful-close handshake: in the async gRPC
path the close reason (clean vs. erroneous) is not transported, so the
producer cannot distinguish a clean consumer finish from a real failure,
and its in-flight write fails with "tunnel is already closed" /
"unexpectedWriteDone called".
- The failed producer task reports error packets, which cascade to sibling
tasks and abort the whole MPP query, even though every task's own work
was correct.
Note this weakness is not introduced by #11001; it has existed in the
exchange layer since the beginning. #11001 merely added a new early-finisher
(the empty-build skip) that hits the window reliably.
4. What is your TiFlash version? (Required)
master (exchange-layer weakness predates #11001; reliably triggered by it)
Bug Report
1. Minimal reproduce step (Required)
Any MPP query in which an exchange consumer task legitimately finishes before
its upstream producer tasks have finished sending, while the producers still
have packets in flight toward it.
Known triggers:
and immediately closes its probe ExchangeReceiver, while the probe-side
producer tasks are still writing (the skip optimization from executor: skip probe for empty hash joins #11001 makes
this window wide and reliable, which is how the bug surfaced).
exits (MPPTask's earier exit may cause false positive error #7177). Not call abortMPPGather when mpp task met error #7969 removed the gather-level error cascade, but the
producer-side write failure and the task-level error-packet cascade remain.
The failure itself is a timing race: it requires a producer write to land
after the consumer has already closed. On an unmodified build it occurs with
low probability; it can be reproduced deterministically with a small
delay-injection patch that stretches the producer write window.
2. What did you expect to see? (Required)
The query succeeds. A consumer legitimately finishing early — because of
LIMIT, an empty build side, a filter that eliminates all rows, etc. — is a
fact that upstream producers must tolerate, not an error. In-flight packets
that arrive after the consumer has finished should be discarded silently,
and the producer tasks should finish normally.
3. What did you see instead (Required)
The query fails with err msg like:
ERROR 1105 (HY000): write to tunnel tunnel13+21 which is already closed, tunnel13+21: unexpectedWriteDone calledThe correct query result is lost. Root cause:
ExchangeReceiver / MPPTunnel while upstream producers are still sending.
path the close reason (clean vs. erroneous) is not transported, so the
producer cannot distinguish a clean consumer finish from a real failure,
and its in-flight write fails with "tunnel is already closed" /
"unexpectedWriteDone called".
tasks and abort the whole MPP query, even though every task's own work
was correct.
Note this weakness is not introduced by #11001; it has existed in the
exchange layer since the beginning. #11001 merely added a new early-finisher
(the empty-build skip) that hits the window reliably.
4. What is your TiFlash version? (Required)
master (exchange-layer weakness predates #11001; reliably triggered by it)