Skip to content

Repository files navigation

rh-custom-models-plugin

A vLLM plugin for model definitions that are not carried in vLLM upstream. Install it next to vLLM and the models load like built-in ones.

Model families

Entry point Architectures Source
rh_bart BartForConditionalGeneration vllm-project/bart-plugin @ 4da3192
rh_florence2 Florence2ForConditionalGeneration (BART backbone from rh_bart) vllm-project/bart-plugin @ 4da3192
rh_gliner2 GLiNER2ForClassification: GLiNER2 zero-shot classification, such as fastino/GLiNER2.5-Decide (DeBERTa-v3) and GLiNER2.5-Decide-1B (ModernBERT) new; DeBERTa-v2 encoder from vllm-project/vllm#42094

Requires vLLM 0.30 or newer.

Status

  • BART runs on vLLM 0.30.0. It was ported from bart-plugin, which targets older vLLM, for three vLLM API changes: the removed MRV2 architecture allowlist, AutoWeightsLoader skip lists, and the multimodal processor hook.
  • GLiNER2 serves classification only; span extraction (entities, JSON structures, relations) is not ported. DeBERTa-v2/v3 and ModernBERT encoders are supported. The DeBERTa encoder computes its disentangled attention per sequence in PyTorch rather than through vLLM's attention backends, so it runs without CUDA graphs or torch.compile.
  • Florence-2 still uses the removed _call_hf_processor hook and does not load on vLLM 0.30 yet.
  • tests/bart/test_model_initialization.py is still written for the older vLLM API: its tests build the model outside a vLLM config context, and two of them use a small_model_name fixture that does not exist.

Install

uv pip install git+https://github.com/neuralmagic/rh-custom-models-plugin.git
# or, from a checkout
uv pip install -e ".[test,lint]"

vLLM loads every installed vllm.general_plugins entry point. To load only some families, list their entry points:

VLLM_PLUGINS=rh_bart vllm serve facebook/bart-large-cnn

Uninstall vllm-bart-plugin first if it is installed: both register the same architectures.

Usage

BART

BART is an encoder-decoder model. The encoder text goes in multi_modal_data:

from vllm import LLM, SamplingParams

llm = LLM(model="facebook/bart-large-cnn", max_model_len=1024)
outputs = llm.generate(
    {
        "encoder_prompt": {
            "prompt": "",
            "multi_modal_data": {"text": "The president of the United States is"},
        },
        "decoder_prompt": "<s>",
    },
    SamplingParams(temperature=0.0, max_tokens=20),
)

Plain /v1/completions prompts are routed to the encoder by the plugin. See examples/bart and examples/florence2.

GLiNER2

Serve with the /v1/classify endpoint. VLLM_PLUGINS=rh_gliner2 loads the model family and its endpoint, and --config-format gliner2 reads the checkpoint's encoder_config/:

VLLM_PLUGINS=rh_gliner2 uvx --python 3.12 --torch-backend auto \
  --from "vllm==0.30.0" \
  --with "rh-custom-models-plugin[gliner2] @ git+https://github.com/neuralmagic/rh-custom-models-plugin.git" \
  vllm serve fastino/GLiNER2.5-Decide --config-format gliner2

POST /v1/classify takes the arguments of AutoExtractor.classify_text and returns its answers, one per text:

curl localhost:8000/v1/classify -H 'content-type: application/json' -d '{
  "text": ["Refund my duplicate charge today.", "The app crashes on login."],
  "tasks": {
    "intent": ["refund", "bug_report", "other"],
    "topics": {"labels": ["billing", "login", "shipping"], "multi_label": true, "cls_threshold": 0.4}
  },
  "include_confidence": true
}'
# {"results": [{"intent": {"label": "refund", "confidence": ...}, "topics": [...]}, ...]}

Without the endpoint, the model returns one raw logit per candidate label from /pooling (token_classify). GLiNER2Client builds the prompt and decodes those logits with the gliner2 package's own code, which is what the endpoint does:

from rh_custom_models_plugin.gliner2.client import GLiNER2Client

client = GLiNER2Client("fastino/GLiNER2.5-Decide")
request = client.build(
    text, {"intent": ["refund", "cancel", "other"], "urgent": ["yes", "no"]}
)
(output,) = llm.encode(
    [{"prompt_token_ids": request.prompt_token_ids}], pooling_task="token_classify"
)
client.decode(request, output.outputs.data)  # {"intent": "refund", "urgent": "no"}

See examples/gliner2.

Adding a model family

  1. Add src/rh_custom_models_plugin/<family>/ with:
    • the model code, with imports under rh_custom_models_plugin.<family>;
    • an __init__.py exposing ARCHITECTURES, a map from architecture name to "module:Class", and a register() that calls ModelRegistry.register_model for each one, along with any config hooks the family needs.
  2. Add the entry point to pyproject.toml: rh_<family> = "rh_custom_models_plugin.<family>:register".
  3. Put tests in tests/<family>/ and mark tests that load weights or need a GPU with @pytest.mark.slow. Add a row to the table above.

Keep register() cheap and repeatable. vLLM runs it in the API server, the engine core and every worker, sometimes more than once per process. Register models as "module:Class" strings so the model code is only imported when that model is used. tests/test_entry_points.py checks these conventions for every family.

Tests

pytest tests -m "not slow"   # registration and CPU-only tests
pytest tests                 # everything; needs a GPU

CI runs ruff only.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages