Skip to content

GLiClass: large separate-label-group requests can wait minutes in the queue under mixed load #390

Description

@svonava

Seen while benchmarking GLiClass (knowledgator/gliclass-large-v1.0) on one L4 with 16 requests in flight and mixed extract traffic: in each 4-minute window, 3 to 6 requests timed out at the client's 300 s limit.

All of them used separately encoded label groups (7 or 8 groups) over many long items. Their GPU work should take a few seconds, so the time seems to be spent waiting in the server's queue or batcher, not in inference. It happened the same way with CUDA graphs on (#389) and with the previous runner, so it is not graph-related.

To investigate:

  • How separate-group requests are split into rows and scheduled. Can a large multi-group request starve, or be starved behind, smaller requests? Is there a fairness or ordering issue in the batcher?
  • Whether any per-request timeout or cancellation surfaces the wait to the caller.
  • Reproduce: mixed extract traffic at concurrency 16, with some requests of 7 or 8 separate label groups over 20+ long items.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions