Summary
Add SDK support for listing available models so applications embedding the SDK can discover which models they can use without spinning up the full agent harness. Fabric is the first concrete consumer waiting on this.
Motivation
Some clients want to embed the SDK and use lightweight primitives (model discovery, inference) without loading the whole harness. Today the only path to model information requires standing up a full session, which is heavier than these clients need.
There are two distinct cohorts:
-
Harness customers that also want primitives. They run the harness and want lightweight primitives like model listing. This issue covers this cohort.
-
Non-harness customers A separate problem that needs its own discovery — how large the base is and what they require. (tracked separately.)
The direction is that the SDK is the front door: clients converge on the SDK rather than integrating against the underlying inference API directly.
Proposed behavior
-
A lightweight "list models" call available across the SDK languages that does not require creating a full session.
-
A bounded, documented model-list schema (we don't expose the full underlying model set).
-
Telemetry/headers that identify the hosting client, following the code-completions pattern for distinguishing embedding applications.
-
A scaling story for identity/auth that doesn't require minting a per-client integration ID.
Open questions
Next steps
- Nail down Fabric's actual requirements (must-have vs. nice-to-have).
Summary
Add SDK support for listing available models so applications embedding the SDK can discover which models they can use without spinning up the full agent harness. Fabric is the first concrete consumer waiting on this.
Motivation
Some clients want to embed the SDK and use lightweight primitives (model discovery, inference) without loading the whole harness. Today the only path to model information requires standing up a full session, which is heavier than these clients need.
There are two distinct cohorts:
Harness customers that also want primitives. They run the harness and want lightweight primitives like model listing. This issue covers this cohort.
Non-harness customers A separate problem that needs its own discovery — how large the base is and what they require. (tracked separately.)
The direction is that the SDK is the front door: clients converge on the SDK rather than integrating against the underlying inference API directly.
Proposed behavior
A lightweight "list models" call available across the SDK languages that does not require creating a full session.
A bounded, documented model-list schema (we don't expose the full underlying model set).
Telemetry/headers that identify the hosting client, following the code-completions pattern for distinguishing embedding applications.
A scaling story for identity/auth that doesn't require minting a per-client integration ID.
Open questions
Where does model listing live — the language-specific SDK or the core runtime?
What exactly do we expose for model access?
Next steps