Skip to content

Report each model's type in GET /api/inference/v1/models #37718

Description

@fmontes

Description

GET /api/inference/v1/models (added in #37431) lists every model the resolved site has configured, across the chat, embeddings and image sections of providerConfig, chat first. Each entry has only the four standard fields (id, object, created, owned_by), so a caller can't tell a chat model from an embedding or image model. Sending one to the wrong operation gets a 404 naming model. That makes callers keep their own table of model names.

This task adds a type field to every entry, taken from the providerConfig section the model was configured in:

{ "object": "list", "data": [
  { "id": "openai/gpt-5.6-luna",           "object": "model", "created": 1789000000, "owned_by": "dotcms", "type": "chat" },
  { "id": "openai/text-embedding-3-small", "object": "model", "created": 1789000000, "owned_by": "dotcms", "type": "embedding" },
  { "id": "gpt-image-2",                   "object": "model", "created": 1789000000, "owned_by": "dotcms", "type": "image" }
] }

This reverses a documented decision. ModelsResource#configuredModels and docs/backend/INFERENCE_API.md say that no capability field is added to the entries, because "nobody extends the standard model object that way". That claim is false. Several OpenAI-compatible gateways add fields to the objects /v1/models returns:

Service What it adds to each model entry
Together AI type (chat, language, code, image, embedding, moderation, rerank), context_length, display_name
Mistral capabilities (completion_chat, function_calling, vision, …), max_context_length, type
OpenRouter architecture.input_modalities / output_modalities, context_length, supported_parameters
Groq context_window, max_completion_tokens, active

The addition doesn't break the no-adapter promise. The four standard fields keep their current shape and values. The official OpenAI SDKs keep unknown fields on the model object instead of rejecting them, so a standard client still deserializes the listing into its own type. type reuses Together's vocabulary (chat, embedding, image), so clients that already read Together's field understand it.

Decisions

  • Only fields dotCMS can back from configuration. Tool/function-calling support is not reported. providerConfig doesn't record whether a model accepts tools, and the chat endpoint forwards tools to every chat model. Some models reject them upstream (e.g. deepseek-r1 on Bedrock), so a flag set for every chat model would be wrong for those models. An admin-declared per-section capabilities setting is the honest way to add this later, as a separate issue.
  • A name configured in more than one section is still listed once, and its type comes from the first section it appears in: chat, then embeddings, then image. id stays unique because clients key models by id.
  • Order is unchanged. The site's primary chat model stays the first entry, because "take the first entry" is the documented way to get "whatever this site runs".
  • Values are fixed: chat, embedding (singular, although the config section is named embeddings), image.

Acceptance Criteria

Happy path

  • Every entry in GET /api/inference/v1/models carries a type string field.
  • A model configured in the chat section (including every fallback-chain entry) reports "type": "chat".
  • A model configured in the embeddings section reports "type": "embedding".
  • A model configured in the image section reports "type": "image".
  • id, object, created and owned_by keep their current values, and the entry order stays chat first, in configured order.

Edge cases

  • A name configured in both chat and embeddings appears exactly once, with "type": "chat". A name configured in both embeddings and image appears once, with "type": "embedding".
  • A site with only an embeddings section lists only embedding entries, each with "type": "embedding".
  • A site with no dotAI configuration, on an instance with none at the system level either, still returns 200 with "data": [].
  • A malformed or unreadable section contributes no entries and does not suppress the other sections' entries or their type values (current behaviour kept).

Error paths (unchanged)

  • Non-bearer credentials still get a 401 in the InferenceErrorView shape, and a site the caller can't read still gets a 403.

Compatibility

  • The official OpenAI client used in InferenceClientConformanceTest still lists models without error, and the type field is present in the raw response.

Documentation and contract

  • ModelListView.ModelView declares type with an @Schema listing the allowed values, and openapi.yaml is regenerated and committed.
  • The Javadoc on ModelsResource / ModelListView and the "Model selection" and GET /models sections of docs/backend/INFERENCE_API.md describe the new field. The claim that no gateway extends the model object is removed.
  • InferenceModelsTest covers each type value, the duplicate-name rule, and the empty-configuration case.

Priority

Medium

Additional Context

  • Parent: Add OpenAI-compatible inference endpoints at /api/inference/v1 #37431 (OpenAI-compatible inference endpoints). The original spec is specs/37431-openai-compatible-inference/spec.md (FR-010, User Story 4).
  • Relevant code: dotCMS/src/main/java/com/dotcms/inference/rest/ModelsResource.java (LISTED_SECTIONS, configuredModels, toView), dotCMS/src/main/java/com/dotcms/inference/rest/view/ModelListView.java, LangChain4jAIClient#configuredModels.
  • Possible follow-ups, not in scope here: config-backed fields such as provider, dimensions for embeddings, the configured max_completion_tokens cap for chat, and an admin-declared capabilities setting for tool/vision support.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions