Seen while benchmarking GLiClass (knowledgator/gliclass-large-v1.0) on one L4 with 16 requests in flight and mixed extract traffic: in each 4-minute window, 3 to 6 requests timed out at the client's 300 s limit.
All of them used separately encoded label groups (7 or 8 groups) over many long items. Their GPU work should take a few seconds, so the time seems to be spent waiting in the server's queue or batcher, not in inference. It happened the same way with CUDA graphs on (#389) and with the previous runner, so it is not graph-related.
To investigate:
- How separate-group requests are split into rows and scheduled. Can a large multi-group request starve, or be starved behind, smaller requests? Is there a fairness or ordering issue in the batcher?
- Whether any per-request timeout or cancellation surfaces the wait to the caller.
- Reproduce: mixed extract traffic at concurrency 16, with some requests of 7 or 8 separate label groups over 20+ long items.
Seen while benchmarking GLiClass (
knowledgator/gliclass-large-v1.0) on one L4 with 16 requests in flight and mixed extract traffic: in each 4-minute window, 3 to 6 requests timed out at the client's 300 s limit.All of them used separately encoded label groups (7 or 8 groups) over many long items. Their GPU work should take a few seconds, so the time seems to be spent waiting in the server's queue or batcher, not in inference. It happened the same way with CUDA graphs on (#389) and with the previous runner, so it is not graph-related.
To investigate: