OpenAI-compatible providers¶
Most of the model ecosystem speaks the OpenAI wire protocol. Managed
services (Groq, Together, OpenRouter, DeepSeek, Mistral, xAI, Fireworks,
Cerebras, Perplexity, NVIDIA NIM) and self-hosted servers (Ollama, vLLM,
LM Studio, llama.cpp, LiteLLM) differ only in base URL and which
environment variable holds the key — so Tulip reaches all of them through
the same OpenAIModel, with no extra client dependency.
Address one by prefix:
from tulip.agent import Agent
Agent(model="groq:llama-3.3-70b-versatile")
Agent(model="ollama:qwen3")
Agent(model="deepseek:deepseek-chat")
OpenRouter and Together.ai¶
Both are supported directly through named prefixes. Install the OpenAI extra, set the provider's own key, and use the exact model id shown in that provider's catalog:
The prefix selects the endpoint and credential variable; Tulip passes the part
after the colon through as the provider's model id. No OpenAI account or
OPENAI_API_KEY is required for these two routes. Tool calling, structured
output, vision, context size, pricing, and availability remain properties of
the selected model and provider.
Provider routes in this build¶
This table is generated from the provider registry in SDK 2.18.3 during the documentation build.
| Prefix | Provider | Endpoint | API key |
|---|---|---|---|
openai |
OpenAI | (default) | OPENAI_API_KEY |
anthropic |
Anthropic | (default) | ANTHROPIC_API_KEY |
bedrock |
Amazon Bedrock | (AWS region) | (boto3 credential chain) |
azure |
Azure OpenAI | AZURE_OPENAI_ENDPOINT |
AZURE_OPENAI_API_KEY |
ollama |
Ollama | http://localhost:11434/v1 |
OLLAMA_API_KEY (optional) |
vllm |
vLLM | http://localhost:8000/v1 |
VLLM_API_KEY (optional) |
lmstudio |
LM Studio | http://localhost:1234/v1 |
LMSTUDIO_API_KEY (optional) |
llamacpp |
llama.cpp server | http://localhost:8080/v1 |
LLAMACPP_API_KEY (optional) |
litellm |
LiteLLM gateway | http://localhost:4000/v1 |
LITELLM_API_KEY |
groq |
Groq | https://api.groq.com/openai/v1 |
GROQ_API_KEY |
together |
Together AI | https://api.together.xyz/v1 |
TOGETHER_API_KEY |
openrouter |
OpenRouter | https://openrouter.ai/api/v1 |
OPENROUTER_API_KEY |
deepseek |
DeepSeek | https://api.deepseek.com/v1 |
DEEPSEEK_API_KEY |
mistral |
Mistral AI | https://api.mistral.ai/v1 |
MISTRAL_API_KEY |
xai |
xAI (Grok) | https://api.x.ai/v1 |
XAI_API_KEY |
fireworks |
Fireworks AI | https://api.fireworks.ai/inference/v1 |
FIREWORKS_API_KEY |
cerebras |
Cerebras | https://api.cerebras.ai/v1 |
CEREBRAS_API_KEY |
perplexity |
Perplexity | https://api.perplexity.ai |
PERPLEXITY_API_KEY |
nvidia |
NVIDIA NIM | https://integrate.api.nvidia.com/v1 |
NVIDIA_API_KEY |
gemini |
Google Gemini | https://generativelanguage.googleapis.com/v1beta/openai/ |
GEMINI_API_KEY |
openai-compatible |
Any OpenAI-compatible endpoint | (supply base_url) |
OPENAI_COMPATIBLE_API_KEY (optional) |
Anything not listed is still reachable without a code change — give the base URL explicitly:
from tulip.models import get_model
model = get_model("openai-compatible:my-model", base_url="https://host/v1")
Agent(model=model)
Resolution order¶
For both the endpoint and the key, the first value found wins:
Endpoint — explicit base_url= → TULIP_<PREFIX>_BASE_URL → the
vendor's own variable where one exists (OLLAMA_BASE_URL, VLLM_BASE_URL,
LMSTUDIO_BASE_URL, LLAMACPP_BASE_URL, LITELLM_GATEWAY_URL) → the
default in the table.
Key — explicit api_key= → the provider's variable from the table. A
hosted provider with no key raises immediately, naming the variable to set.
Local servers need none.
# Point the ollama prefix at a GPU box instead of localhost
export TULIP_OLLAMA_BASE_URL=http://gpu-box:11434/v1
Passing configuration inline
AgentConfig rejects unknown keyword arguments, so
Agent(model="groq:x", api_key=...) raises. Build the model first when
the configuration is not in the environment:
What this does not change¶
- The Responses API is never auto-selected against a custom base URL.
api="auto"routes to/v1/responsesonly for model families that require it, and only againstapi.openai.comitself — a gateway serves chat-completions and would 404 on the Responses path. Setapi="responses"explicitly if your endpoint does serve it. - Capability still varies by model. A prefix makes an endpoint reachable; it does not promise that the model behind it supports tool calling, structured output, or vision. See Structured output for the fallbacks Tulip applies when a model cannot constrain its own decoding.
- Sampling defaults are sent.
temperatureandtop_pdefault to0.7/0.9on the model config and go out with chat-completions requests (model ids Tulip treats as reasoning or search-preview models are the exception).AgentConfig.temperaturedefaults toNoneand is forwarded on the agent's regular turns only when set, so it does not override them there. (The iteration-limit summary, empty-answer recovery and structured-output repair calls pass the agent'stemperatureexplicitly, so with the defaultNonethey sendtop_pbut notemperature.)AgentConfighas notop_pfield. To let a self-hosted server's owngeneration_config.jsondefaults apply, clear them when you build the model —get_model("vllm:my-model", temperature=None, top_p=None), orAgent(model="vllm:my-model", model_kwargs={"temperature": None, "top_p": None})— sinceNonemeans "do not send this parameter".