A vLLM plugin for model definitions that are not carried in vLLM upstream. Install it next to vLLM and the models load like built-in ones.
| Entry point | Architectures | Source |
|---|---|---|
rh_bart |
BartForConditionalGeneration |
vllm-project/bart-plugin @ 4da3192 |
rh_florence2 |
Florence2ForConditionalGeneration (BART backbone from rh_bart) |
vllm-project/bart-plugin @ 4da3192 |
rh_gliner2 |
GLiNER2ForClassification: GLiNER2 zero-shot classification, such as fastino/GLiNER2.5-Decide (DeBERTa-v3) and GLiNER2.5-Decide-1B (ModernBERT) |
new; DeBERTa-v2 encoder from vllm-project/vllm#42094 |
Requires vLLM 0.30 or newer.
- BART runs on vLLM 0.30.0. It was ported from
bart-plugin, which targets older vLLM, for three vLLM API changes: the removed MRV2 architecture allowlist,AutoWeightsLoaderskip lists, and the multimodal processor hook. - GLiNER2 serves classification only; span extraction (entities, JSON structures, relations) is not ported. DeBERTa-v2/v3 and ModernBERT encoders are supported. The DeBERTa encoder computes its disentangled attention per sequence in PyTorch rather than through vLLM's attention backends, so it runs without CUDA graphs or
torch.compile. - Florence-2 still uses the removed
_call_hf_processorhook and does not load on vLLM 0.30 yet. tests/bart/test_model_initialization.pyis still written for the older vLLM API: its tests build the model outside a vLLM config context, and two of them use asmall_model_namefixture that does not exist.
uv pip install git+https://github.com/neuralmagic/rh-custom-models-plugin.git
# or, from a checkout
uv pip install -e ".[test,lint]"vLLM loads every installed vllm.general_plugins entry point. To load only some families, list their entry points:
VLLM_PLUGINS=rh_bart vllm serve facebook/bart-large-cnnUninstall vllm-bart-plugin first if it is installed: both register the same architectures.
BART is an encoder-decoder model. The encoder text goes in multi_modal_data:
from vllm import LLM, SamplingParams
llm = LLM(model="facebook/bart-large-cnn", max_model_len=1024)
outputs = llm.generate(
{
"encoder_prompt": {
"prompt": "",
"multi_modal_data": {"text": "The president of the United States is"},
},
"decoder_prompt": "<s>",
},
SamplingParams(temperature=0.0, max_tokens=20),
)Plain /v1/completions prompts are routed to the encoder by the plugin. See examples/bart and examples/florence2.
Serve with the /v1/classify endpoint. VLLM_PLUGINS=rh_gliner2 loads the model family and its endpoint, and --config-format gliner2 reads the checkpoint's encoder_config/:
VLLM_PLUGINS=rh_gliner2 uvx --python 3.12 --torch-backend auto \
--from "vllm==0.30.0" \
--with "rh-custom-models-plugin[gliner2] @ git+https://github.com/neuralmagic/rh-custom-models-plugin.git" \
vllm serve fastino/GLiNER2.5-Decide --config-format gliner2POST /v1/classify takes the arguments of AutoExtractor.classify_text and returns its answers, one per text:
curl localhost:8000/v1/classify -H 'content-type: application/json' -d '{
"text": ["Refund my duplicate charge today.", "The app crashes on login."],
"tasks": {
"intent": ["refund", "bug_report", "other"],
"topics": {"labels": ["billing", "login", "shipping"], "multi_label": true, "cls_threshold": 0.4}
},
"include_confidence": true
}'
# {"results": [{"intent": {"label": "refund", "confidence": ...}, "topics": [...]}, ...]}Without the endpoint, the model returns one raw logit per candidate label from /pooling (token_classify). GLiNER2Client builds the prompt and decodes those logits with the gliner2 package's own code, which is what the endpoint does:
from rh_custom_models_plugin.gliner2.client import GLiNER2Client
client = GLiNER2Client("fastino/GLiNER2.5-Decide")
request = client.build(
text, {"intent": ["refund", "cancel", "other"], "urgent": ["yes", "no"]}
)
(output,) = llm.encode(
[{"prompt_token_ids": request.prompt_token_ids}], pooling_task="token_classify"
)
client.decode(request, output.outputs.data) # {"intent": "refund", "urgent": "no"}See examples/gliner2.
- Add
src/rh_custom_models_plugin/<family>/with:- the model code, with imports under
rh_custom_models_plugin.<family>; - an
__init__.pyexposingARCHITECTURES, a map from architecture name to"module:Class", and aregister()that callsModelRegistry.register_modelfor each one, along with any config hooks the family needs.
- the model code, with imports under
- Add the entry point to
pyproject.toml:rh_<family> = "rh_custom_models_plugin.<family>:register". - Put tests in
tests/<family>/and mark tests that load weights or need a GPU with@pytest.mark.slow. Add a row to the table above.
Keep register() cheap and repeatable. vLLM runs it in the API server, the engine core and every worker, sometimes more than once per process. Register models as "module:Class" strings so the model code is only imported when that model is used. tests/test_entry_points.py checks these conventions for every family.
pytest tests -m "not slow" # registration and CPU-only tests
pytest tests # everything; needs a GPUCI runs ruff only.