Bug Description
Cancelling a text-output task while it is waiting for TranscriptSynchronizer to rotate segments can cancel the shared rotation task. If this happens after the old segment is marked closed but before its replacement is created, the synchronizer remains attached to the closed segment.
Subsequent assistant audio frames still reach the downstream audio output, but assistant text is dropped with warnings such as:
WARNING: _SegmentSynchronizerImpl.push_text called after close
WARNING: _SegmentSynchronizerImpl.push_audio called after close
This is persistent loss across later turns, rather than truncation of the interrupted response alone. The standalone reproduction below delivers all four subsequent audio frames but no text on both livekit-agents==1.6.8 and 1.8.1.
Expected Behavior
Cancelling one forwarding task should cancel that caller promptly without cancelling session-owned segment rotation. Once rotation completes, subsequent turns should continue to deliver synchronized text and audio.
Reproduction Steps
- Save the standalone script below as
repro.py in any directory. It uses only the official SDK and Python's standard library. No LiveKit server, model provider, credentials, microphone, or network session is required. Package installation requires internet access.
- With
uv installed, run:
uv run --no-project --python 3.12 --with livekit-agents==1.8.1 repro.py
# Also reproduced with:
uv run --no-project --python 3.12 --with livekit-agents==1.6.8 repro.py
- Each invocation first runs a control without cancellation, then cancels a task awaiting the real text output during segment rotation. It waits until the old segment is closed to target the cancellation window; it does not replace or modify SDK methods in the default run. Private state is read only to select this interleaving and report diagnostics. The output sinks are in-memory implementations of the SDK's output interfaces, with synthetic silent PCM frames and synthetic text.
- Both unmodified versions exit with an assertion failure: the control delivers four responses, while the cancellation case delivers four audio frames and zero text. The probe observes cancellation from later rotations so it can continue through all four turns.
Standalone reproduction script
"""Offline reproduction; no server, model, credentials, or application code.
uv run --no-project --python 3.12 --with livekit-agents==1.6.8 repro.py
Append --shield to test the proposed barrier change in this process only.
"""
import argparse
import asyncio
import json
import logging
import time
from importlib.metadata import version
from livekit import rtc
from livekit.agents.voice import io
from livekit.agents.voice.transcription.synchronizer import TranscriptSynchronizer
class AudioSink(io.AudioOutput):
def __init__(self):
super().__init__(label="memory-audio", capabilities=io.AudioOutputCapabilities(pause=True))
self.frames = 0
async def capture_frame(self, frame):
await super().capture_frame(frame)
self.frames += 1
self.on_playback_started(created_at=time.time())
def flush(self):
super().flush()
def clear_buffer(self):
super().flush()
class TextSink(io.TextOutput):
def __init__(self):
super().__init__(label="memory-text", next_in_chain=None)
self.text = ""
async def capture_text(self, text):
self.text += text
def flush(self):
pass
async def shielded_barrier(self):
while (rotation := self._rotate_segment_atask) is not None and not rotation.done():
await asyncio.shield(rotation)
async def scenario(cancel_waiter):
audio, text = AudioSink(), TextSink()
sync = TranscriptSynchronizer(next_in_chain_audio=audio, next_in_chain_text=text)
try:
old_impl = sync._impl
sync.rotate_segment()
# The real text output waits on barrier before accepting this text.
waiter = asyncio.create_task(sync.text_output.capture_text(""))
async with asyncio.timeout(5):
while not old_impl.closed:
await asyncio.sleep(0)
assert not waiter.done(), "Missed the cancellation window"
assert not sync._rotate_segment_atask.done(), "Rotation already completed"
if cancel_waiter:
waiter.cancel()
# Observe the expected waiter cancellation without cancelling the driver.
outcome = (await asyncio.gather(waiter, return_exceptions=True))[0]
if cancel_waiter:
assert isinstance(outcome, asyncio.CancelledError)
else:
assert outcome is None
await sync.barrier()
rotation_cancelled = sync._rotate_segment_atask.cancelled()
later_rotation_cancellations = 0
message = "Next response."
for _ in range(4):
await sync.text_output.capture_text(message)
await sync.audio_output.capture_frame(rtc.AudioFrame(
data=bytes(960), sample_rate=24000, num_channels=1, samples_per_channel=480,
))
sync.text_output.flush()
sync.audio_output.flush()
audio.on_playback_finished(playback_position=0.02, interrupted=False)
# Keep probing after a failed rotation to demonstrate persistent loss.
result = (await asyncio.gather(sync.barrier(), return_exceptions=True))[0]
if isinstance(result, BaseException) and not isinstance(result, asyncio.CancelledError):
raise result
later_rotation_cancellations += sync._rotate_segment_atask.cancelled()
result = {
"cancel_waiter": cancel_waiter,
"rotation_cancelled": rotation_cancelled,
"same_closed_impl": sync._impl is old_impl and old_impl.closed,
"later_rotation_cancellations": later_rotation_cancellations,
"audio_frames": audio.frames,
"transcript": text.text,
}
print(json.dumps(result), flush=True)
return audio.frames == 4 and text.text == message * 4
finally:
await sync.aclose()
async def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--shield", action="store_true")
args = parser.parse_args()
if args.shield:
TranscriptSynchronizer.barrier = shielded_barrier
print(json.dumps({"livekit-agents": version("livekit-agents"),
"livekit": version("livekit"), "shield": args.shield}), flush=True)
async with asyncio.timeout(20):
control_ok = await scenario(cancel_waiter=False)
interrupted_ok = await scenario(cancel_waiter=True)
assert control_ok, "Control failed"
assert interrupted_ok, "Audio delivered, but later assistant transcripts were lost"
if __name__ == "__main__":
logging.basicConfig(level=logging.WARNING, format="%(levelname)s: %(message)s")
asyncio.run(main())
Representative output from 1.8.1 (repeated warnings and traceback paths omitted):
{"livekit-agents": "1.8.1", "livekit": "1.1.18", "shield": false}
{"cancel_waiter": false, "rotation_cancelled": false, "same_closed_impl": false, "later_rotation_cancellations": 0, "audio_frames": 4, "transcript": "Next response.Next response.Next response.Next response."}
WARNING: _SegmentSynchronizerImpl.push_text called after close
WARNING: _SegmentSynchronizerImpl.push_audio called after close
{"cancel_waiter": true, "rotation_cancelled": true, "same_closed_impl": true, "later_rotation_cancellations": 4, "audio_frames": 4, "transcript": ""}
AssertionError: Audio delivered, but later assistant transcripts were lost
Operating System
macOS 15.6.1, Apple Silicon (arm64); Python 3.12.9.
Models Used
None. This reproduction operates directly on the SDK transcript synchronizer with in-memory outputs; it does not instantiate an STT, LLM, TTS, or realtime model.
Package Versions
Two separate uv --no-project environments were tested:
Environment A:
livekit-agents==1.6.8
livekit==1.1.14
Environment B:
livekit-agents==1.8.1
livekit==1.1.18
Python==3.12.9
Session/Room/Call IDs
Not applicable. The reproduction runs offline without creating any session, room, or call.
Proposed Solution
A candidate fix is to shield the shared rotation task from cancellation of individual barrier() callers:
async def barrier(self) -> None:
while (rotation := self._rotate_segment_atask) is not None and not rotation.done():
await asyncio.shield(rotation)
The script's optional --shield flag applies only this change in the current process:
uv run --no-project --python 3.12 --with livekit-agents==1.8.1 repro.py --shield
uv run --no-project --python 3.12 --with livekit-agents==1.6.8 repro.py --shield
Both commands exit successfully: the caller is still cancelled, rotation completes, and all four subsequent audio frames and transcripts are delivered without the closed-segment warnings.
This demonstrates a preventative fix for the reproduced cancellation path. It does not recover an already-cancelled rotation task. Session shutdown ownership and any recovery for explicitly cancelled rotations should also be considered in an upstream fix; this component reproduction is not a full end-to-end validation of those behaviors.
Additional Context
The relevant implementation is in voice/transcription/synchronizer.py at tag livekit-agents@1.8.1:
barrier() directly awaits _rotate_segment_atask, so cancelling a waiter propagates cancellation to the shared task.
_rotate_segment_task() awaits old_impl.aclose() before assigning a new implementation. Cancellation during that await can leave self._impl pointing at the closed old implementation. except Exception does not catch asyncio.CancelledError on the tested Python version.
- Later rotations await the cancelled predecessor under
contextlib.suppress(Exception), which likewise does not suppress CancelledError; all four later rotations are cancelled in the reproduction.
- A later
barrier() call can return without raising because the rotation task is already done. Text then reaches the closed implementation and is discarded. The audio wrapper forwards frames to its downstream output before feeding the segment synchronizer, explaining why audio can continue while transcription is lost.
Verification results:
| SDK |
No-cancellation control |
After cancellation, unmodified SDK |
After cancellation, candidate shield change |
| 1.6.8 |
4 audio frames, 4 transcripts |
4 audio frames, 0 transcripts |
4 audio frames, 4 transcripts |
| 1.8.1 |
4 audio frames, 4 transcripts |
4 audio frames, 0 transcripts |
4 audio frames, 4 transcripts |
Related reports: #2265 concerns audio/text forwarding cancellation and closed-channel errors; #2216 reports missing assistant transcripts while audio continues. They may be related, but neither establishes this exact shared-rotation cancellation path, so I am not claiming they are duplicates.
Screenshots and Recordings
Not applicable. The standalone reproduction uses synthetic inputs and reports the downstream results directly.
Bug Description
Cancelling a text-output task while it is waiting for
TranscriptSynchronizerto rotate segments can cancel the shared rotation task. If this happens after the old segment is marked closed but before its replacement is created, the synchronizer remains attached to the closed segment.Subsequent assistant audio frames still reach the downstream audio output, but assistant text is dropped with warnings such as:
This is persistent loss across later turns, rather than truncation of the interrupted response alone. The standalone reproduction below delivers all four subsequent audio frames but no text on both
livekit-agents==1.6.8and1.8.1.Expected Behavior
Cancelling one forwarding task should cancel that caller promptly without cancelling session-owned segment rotation. Once rotation completes, subsequent turns should continue to deliver synchronized text and audio.
Reproduction Steps
repro.pyin any directory. It uses only the official SDK and Python's standard library. No LiveKit server, model provider, credentials, microphone, or network session is required. Package installation requires internet access.uvinstalled, run:uv run --no-project --python 3.12 --with livekit-agents==1.8.1 repro.py # Also reproduced with: uv run --no-project --python 3.12 --with livekit-agents==1.6.8 repro.pyStandalone reproduction script
Representative output from
1.8.1(repeated warnings and traceback paths omitted):Operating System
macOS 15.6.1, Apple Silicon (arm64); Python 3.12.9.
Models Used
None. This reproduction operates directly on the SDK transcript synchronizer with in-memory outputs; it does not instantiate an STT, LLM, TTS, or realtime model.
Package Versions
Two separate
uv --no-projectenvironments were tested:Session/Room/Call IDs
Not applicable. The reproduction runs offline without creating any session, room, or call.
Proposed Solution
A candidate fix is to shield the shared rotation task from cancellation of individual
barrier()callers:The script's optional
--shieldflag applies only this change in the current process:Both commands exit successfully: the caller is still cancelled, rotation completes, and all four subsequent audio frames and transcripts are delivered without the closed-segment warnings.
This demonstrates a preventative fix for the reproduced cancellation path. It does not recover an already-cancelled rotation task. Session shutdown ownership and any recovery for explicitly cancelled rotations should also be considered in an upstream fix; this component reproduction is not a full end-to-end validation of those behaviors.
Additional Context
The relevant implementation is in
voice/transcription/synchronizer.pyat taglivekit-agents@1.8.1:barrier()directly awaits_rotate_segment_atask, so cancelling a waiter propagates cancellation to the shared task._rotate_segment_task()awaitsold_impl.aclose()before assigning a new implementation. Cancellation during that await can leaveself._implpointing at the closed old implementation.except Exceptiondoes not catchasyncio.CancelledErroron the tested Python version.contextlib.suppress(Exception), which likewise does not suppressCancelledError; all four later rotations are cancelled in the reproduction.barrier()call can return without raising because the rotation task is already done. Text then reaches the closed implementation and is discarded. The audio wrapper forwards frames to its downstream output before feeding the segment synchronizer, explaining why audio can continue while transcription is lost.Verification results:
Related reports: #2265 concerns audio/text forwarding cancellation and closed-channel errors; #2216 reports missing assistant transcripts while audio continues. They may be related, but neither establishes this exact shared-rotation cancellation path, so I am not claiming they are duplicates.
Screenshots and Recordings
Not applicable. The standalone reproduction uses synthetic inputs and reports the downstream results directly.