Model Providers
Roster uses a model during participant resolution to interpret the workflow question, select relevant participants, and produce an auditable answer.
The deployment owns the model choice. Roster records model provider, model name, latency, token usage, and resolution metadata for observability.
Admins can store resolver agent metadata and cost inputs in Settings. Provider credentials and provider family defaults are still deployment runtime configuration.
Provider Model
Section titled “Provider Model”Roster runs one resolver model provider at a time.
Choose a Provider
Section titled “Choose a Provider”Choose the provider value that matches how your deployment reaches the model:
| Use case | ROSTER_MODEL_PROVIDER | Default API | Model name format |
|---|---|---|---|
| OpenAI | openai | responses | OpenAI model ID, for example gpt-6.1-sol |
| Mistral | mistral | chat-completions | Mistral model ID, for example mistral-large-2512 |
| Anthropic | anthropic | messages | Anthropic model ID, for example claude-sonnet-5-5 |
| OpenRouter or another OpenAI-compatible gateway | openai | responses | Gateway model ID, for example openai/gpt-5.4-mini |
Provider-specific *_BASE_URL values are optional. Set provider base URLs only
when routing through a provider proxy, a regional endpoint, or an approved model
gateway. The examples below show the default public provider
endpoints so operators know what is used when the variable is omitted.
OpenAI
Section titled “OpenAI”OpenAI is the default provider:
OPENAI_API_KEY=<openai-api-key>OPENAI_BASE_URL=https://api.openai.com/v1ROSTER_MODEL_PROVIDER=openaiROSTER_MODEL_NAME=gpt-6.1-solROSTER_MODEL_EFFORT=lowMistral
Section titled “Mistral”Mistral is supported through its chat-completions API:
MISTRAL_API_KEY=<mistral-api-key>MISTRAL_BASE_URL=https://api.mistral.ai/v1ROSTER_MODEL_PROVIDER=mistralROSTER_MODEL_NAME=mistral-large-2512Anthropic
Section titled “Anthropic”Anthropic is supported through its Messages API:
ANTHROPIC_API_KEY=<anthropic-api-key>ANTHROPIC_BASE_URL=https://api.anthropic.comROSTER_MODEL_PROVIDER=anthropicROSTER_MODEL_NAME=claude-sonnet-5-5OpenRouter Gateway
Section titled “OpenRouter Gateway”For OpenRouter, keep ROSTER_MODEL_PROVIDER=openai because OpenRouter exposes
an OpenAI-compatible API.
For OpenRouter-hosted OpenAI models, use the Responses API:
OPENAI_API_KEY=<openrouter-api-key>OPENAI_BASE_URL=https://openrouter.ai/api/v1ROSTER_MODEL_PROVIDER=openaiROSTER_MODEL_NAME=openai/gpt-5.4-miniROSTER_MODEL_EFFORT=lowFor gateway-backed third-party models, Chat Completions is the usual starting point because support for the newer Responses API varies by gateway model:
OPENAI_API_KEY=<openrouter-api-key>OPENAI_BASE_URL=https://openrouter.ai/api/v1ROSTER_MODEL_PROVIDER=openaiROSTER_MODEL_API=chat-completionsROSTER_MODEL_NAME=anthropic/claude-opus-4.8Kimi K3 is a supported exception: use the Responses API and disable reasoning effort:
OPENAI_API_KEY=<openrouter-api-key>OPENAI_BASE_URL=https://openrouter.ai/api/v1ROSTER_MODEL_PROVIDER=openaiROSTER_MODEL_API=responsesROSTER_MODEL_NAME=moonshotai/kimi-k3ROSTER_MODEL_EFFORT=noneDeepSeek V4 Pro is another supported exception. Use the Responses API with high reasoning effort and raise the gateway output ceiling to 8192 tokens. Lower effort settings are not recommended for this integration:
OPENAI_API_KEY=<openrouter-api-key>OPENAI_BASE_URL=https://openrouter.ai/api/v1ROSTER_MODEL_PROVIDER=openaiROSTER_MODEL_API=responsesROSTER_MODEL_NAME=deepseek/deepseek-v4-proROSTER_MODEL_EFFORT=highROSTER_MODEL_OPENAI_RESPONSES_MAX_OUTPUT_TOKENS=8192When OPENAI_BASE_URL is set, Roster defaults the OpenAI Responses output cap
to 2048. Override it only when the gateway, model, or key limit requires a
different bound:
ROSTER_MODEL_OPENAI_RESPONSES_MAX_OUTPUT_TOKENS=2048OpenRouter compatibility depends on the selected gateway model, model API, reasoning effort, key limits, and structured output behavior. Validate each gateway/provider/model pair with your own data before use.
Model API
Section titled “Model API”ROSTER_MODEL_API is optional. When unset, Roster uses the default API for the
selected provider:
ROSTER_MODEL_PROVIDER | Default ROSTER_MODEL_API |
|---|---|
openai | responses |
mistral | chat-completions |
anthropic | messages |
Set ROSTER_MODEL_API only when Roster documents another supported API for the
same provider. OpenAI-compatible gateways may use chat-completions with
ROSTER_MODEL_PROVIDER=openai.
Supported Runtime Providers
Section titled “Supported Runtime Providers”Roster is designed for bring-your-own-model deployments. Deployments can standardize on a direct provider or route through an approved model gateway.
- Use a direct provider when you want the simplest vendor-specific integration.
- Use an approved gateway when you need centralized routing, budgeting, or model access controls.
- Treat gateway compatibility tests separately from Resolve quality tests. A gateway model should still pass your Resolve eval suite before becoming a production recommendation.
Supported Resolve Models
Section titled “Supported Resolve Models”The models below have passed Roster’s full Resolve evaluation. Validate cost, latency, data-handling requirements, and account access before production rollout.
| Provider | Model | API | Recommended use | Operating note |
|---|---|---|---|---|
| OpenAI | gpt-6-astra | responses | Current quality-first option | Use ROSTER_MODEL_EFFORT=low. Validate latency and cost for your workload. |
| OpenAI | gpt-6.1-sol | responses | Current general-purpose recommendation | Use ROSTER_MODEL_EFFORT=low. Recommended starting point for new OpenAI deployments. |
| OpenAI | gpt-6-luna | responses | Current cost-sensitive option | Use ROSTER_MODEL_EFFORT=low. Canary with representative data before high-impact routing. |
| OpenAI | gpt-5.6-sol | responses | Previous-generation quality option | Use ROSTER_MODEL_EFFORT=low. Existing deployments can retain this validated configuration. |
| OpenAI | gpt-5.6-terra | responses | Balanced cost and quality option | Use ROSTER_MODEL_EFFORT=low. Lower-cost GPT-5.6 option for everyday production workloads. |
| OpenAI | gpt-5.6-luna | responses | Cost-sensitive OpenAI recommendation | Use ROSTER_MODEL_EFFORT=low. Canary on representative data before high-impact routing. |
| OpenAI | gpt-5.5 | responses | Previous-generation Full Resolve option | Use ROSTER_MODEL_EFFORT=low. Strong accuracy with moderate latency in the Resolve suite. |
| OpenAI | gpt-5.4 | responses | Previous-generation OpenAI alternative | Use ROSTER_MODEL_EFFORT=low. Useful when newer OpenAI models are not required or available. |
| OpenAI | gpt-5.4-mini | responses | Previous-generation cost option | Use ROSTER_MODEL_EFFORT=low. Canary on representative data before high-impact routing. |
| Mistral | mistral-large-2512 | chat-completions | Mistral Full Resolve recommendation | Prefer this over smaller Mistral models for complex multi-responsibility prompts. |
| Anthropic | claude-fable-5-1 | messages | Current Anthropic quality option | Requires Roster 1.2.3 or later. Use provider-default reasoning and the 4096-token output ceiling. |
| Anthropic | claude-opus-5-5 | messages | Current Anthropic alternative | Requires Roster 1.2.3 or later. Use provider-default reasoning; canary answer grounding on representative data. |
| Anthropic | claude-sonnet-5-5 | messages | Current Anthropic recommendation | Requires Roster 1.2.3 or later. Use provider-default reasoning and the 4096-token output ceiling. |
| Anthropic | claude-opus-4-8 | messages | Previous-generation Anthropic option | Confirm latency and cost fit before production rollout. |
| Anthropic | claude-fable-5 | messages | Anthropic Full Resolve recommendation | Confirm latency and cost fit; this is a direct Anthropic model, not an OpenRouter recommendation. |
Latest Releases to Evaluate
Section titled “Latest Releases to Evaluate”Release IDs below were checked and evaluated on October 5, 2026. These models remain candidates because they did not pass the full Resolve suite at the tested settings. Confirm tool calling, structured output, routing accuracy, and latency on your data before production use. A provider release does not inherit an older model’s evaluation result.
| Provider or gateway | Model | Evaluation API | Requested effort | Output ceiling | Full-suite result |
|---|---|---|---|---|---|
| OpenRouter | deepseek/deepseek-v4.1-flash | responses | high | 8192 tokens | 8/10; two invalid/truncated JSON responses |
| OpenRouter | z-ai/glm-5.3 | responses | high | 8192 tokens | 9/10; invalid JSON on delegation |
Repeated checks at the same settings still showed variability. Across three new attempts per previously failing case, Flash passed finance approval once and the unknown-record lookup twice; the other responses were truncated and could not be parsed. GLM passed the delegation case twice, but added unrelated approvers on the remaining attempt. Both configurations remain unvalidated candidates and are not recommended for production. These retries do not replace the original full-suite results above.
OpenAI’s current lineup is documented in its model catalog. Anthropic’s model overview lists the current Claude IDs. Starting with Roster 1.2.3, the new Claude releases use strict automatic tool selection and native structured JSON output, as required by Anthropic’s migration guide. Roster still checks required tool calls and rejects invalid or incomplete responses. Anthropic reasoning effort remains provider-controlled.
The same compatibility path also handles claude-mythos-5-1, which
rejects forced tool choices.
Adapter regression tests cover its tool selection and native JSON output; it has
not been evaluated through a full live Resolve suite and is not included among
the validated recommendations above.
For DeepSeek V4.1 Flash through OpenRouter, start evaluation with:
OPENAI_API_KEY=<openrouter-api-key>OPENAI_BASE_URL=https://openrouter.ai/api/v1ROSTER_MODEL_PROVIDER=openaiROSTER_MODEL_API=responsesROSTER_MODEL_NAME=deepseek/deepseek-v4.1-flashROSTER_MODEL_EFFORT=highROSTER_MODEL_OPENAI_RESPONSES_MAX_OUTPUT_TOKENS=8192DeepSeek V4.1 Flash and
V4 Pro 0813 are separate
gateway evaluation targets. V4 Pro 0813 passed 10/10 and is listed below.
V4.1 Flash failed two cases after exhausting the 8192-token output limit; it
is not recommended at these settings. The older V4 Flash result does not characterize
V4.1 Flash. On DeepSeek’s own API, the latest Flash model is called
deepseek-flash; that is a different endpoint and model ID from this
OpenRouter configuration. See the DeepSeek release log.
Among the already evaluated open model families,
Mistral Large 3 remains
mistral-large-2512, and
Kimi K3
remains moonshotai/kimi-k3 on OpenRouter. Continue using their validated
settings above and below. GLM 5.3 succeeds
the previously evaluated GLM 5.2 candidate. Its 9/10 run failed structured JSON
on a delegation case. Repeated checks also exposed incorrect extra approvers,
so it remains a candidate.
Choose a Cost-Sensitive Option
Section titled “Choose a Cost-Sensitive Option”Reasoning effort and model price are separate controls. A lower effort can reduce reasoning-token usage for models that support it, but it does not change the model’s per-token price. Do not lower effort below the configuration that passed the full Resolve suite merely to reduce cost.
Use the lowest-cost Full Resolve option that fits the required provider family:
- OpenAI: use
gpt-6-lunafor a current-generation cost-sensitive option, or compare current pricing with the validatedgpt-5.4-miniandgpt-5.6-lunamodels. - Anthropic: compare
claude-sonnet-5-5with the other validated Claude models using current account pricing. The new Sonnet integration requires Roster 1.2.3 or later. - Mistral: use
mistral-large-2512for complex multi-responsibility prompts. - Kimi through OpenRouter: use
moonshotai/kimi-k3with effortnone. - DeepSeek through OpenRouter: use
deepseek/deepseek-v4-prowith efforthighfor the previously validated configuration, or evaluate the newly validateddeepseek/deepseek-v4-pro-0813athighwith an 8192-token ceiling. V4.1 Flash did not pass the full suite at the tested settings.
Provider prices and gateway routing change independently of Roster. Confirm current pricing and run a canary with representative data before rollout.
Do not promote a different model to production until it passes the Resolve eval suite with your expected data shape and guardrails. Other provider-compatible models may work, but use them for testing or canaries until they pass your deployment’s Resolve evals.
OpenRouter Gateway Models
Section titled “OpenRouter Gateway Models”Treat OpenRouter as a gateway configuration for supported model families, not as a separate direct provider recommendation. The following gateway model IDs have passed a full Resolve suite. Retest representative cases on your deployment:
| Gateway model | API | Effort |
|---|---|---|
openai/gpt-5.5 |
responses |
low |
openai/gpt-5.4 |
responses |
low |
openai/gpt-5.4-mini |
responses |
low |
anthropic/claude-opus-4.8 |
chat-completions |
n/a |
moonshotai/kimi-k3 |
responses |
none |
deepseek/deepseek-v4-pro |
responses |
high |
deepseek/deepseek-v4-pro-0813 |
responses |
high |
For deepseek/deepseek-v4-pro and deepseek/deepseek-v4-pro-0813, also set
ROSTER_MODEL_OPENAI_RESPONSES_MAX_OUTPUT_TOKENS=8192. The default gateway
ceiling of 2048 is not the validated Full Resolve configuration.
The October 5 Kimi K3 retest passed nine cases; the remaining case hit an OpenRouter HTTP 402 credit reservation error during concurrent requests. The blocked case passed on an isolated retry. The original full run remains 9/10; its earlier full-suite pass remains historical evidence.
Treat OpenRouter model IDs as deployment-specific candidates. A one-case smoke
test can prove API compatibility, but the full Resolve suite is the minimum
bar for documenting a model as recommended. OpenRouter
mistralai/mistral-large-2512 remains a canary because it has not yet passed
that full-suite bar; this does not affect the direct Mistral recommendation
above. For Kimi K3 and DeepSeek V4 Pro, use the exact API, effort, and output
ceiling settings shown above rather than substituting another model or API
mode. Canary gateway models on representative data because pricing, routing,
and latency can vary by provider and region.
Reasoning Effort
Section titled “Reasoning Effort”ROSTER_MODEL_EFFORT controls the requested reasoning level for providers that
support effort controls. Roster accepts the value for all providers, but direct
Mistral and Anthropic requests currently ignore it.
noneminimallowmediumhighxhighmaxUse lower effort for high-volume routing and higher effort for complex,
high-impact workflows. Gateway models can interpret effort controls
differently from direct OpenAI models. If an OpenAI-compatible gateway model
returns empty outputs or unstable structured responses, validate the same model
with ROSTER_MODEL_EFFORT=none before promoting or rejecting it.
Production Checklist
Section titled “Production Checklist”- Choose the model provider or gateway before exposing resolution to users.
- Store model credentials in the deployment secret manager.
- Verify structured response behavior before approving a model outside the tested recommendations.
- Confirm data residency, retention, and logging requirements with the model provider or gateway.
- Monitor model runs for latency, cost, error rate, and resolution quality.
Use Model Runs to inspect recorded invocations and correlate provider/model choices with latency, token usage, cost, and errors.