Skip to main content
Version: 1.4.0

Baseten

Baseten provides hosted inference through a shared model catalog and dedicated model deployments. With AISIX, applications call those models through one OpenAI-compatible API while the gateway manages credentials, model access, rate limits, and usage accounting.

Baseten serves models through two different surfaces, and the choice determines which endpoint you configure:

  • Model APIs — a shared, multi-tenant catalog of open-weight models at a single fixed endpoint. This page uses that surface.
  • Dedicated deployments — a per-deployment endpoint for a model you deployed into your own Baseten workspace. See Route to a Dedicated Deployment.

Prerequisites​

Before starting, prepare the following:

  • One AISIX setup:
    • For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the AISIX Cloud Quickstart. To request Hybrid Cloud access, contact API7.
    • For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the Open-Source AISIX Gateway Quickstart. Configure the gateway to load a declarative resources file.
  • A Baseten API key from the API keys page of your workspace. See Baseten API keys.
  • curl and jq.

Configure with AISIX Cloud​

Export the AISIX Cloud connection details:

# AISIX_CP is the Admin API base URL; include /api and omit a trailing slash
# The local On-Premises quickstart uses http://localhost:8080/api
export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL"
export AISIX_TOKEN="YOUR_ADMIN_TOKEN"
export ENV_ID="YOUR_ENVIRONMENT_ID"

Create a provider key, model alias, and caller API key for the Baseten-backed chat-completions route.

Because Baseten exposes an OpenAI-compatible API, AISIX connects through the openai adapter. Set api_base for the Baseten surface you use.

Create a Provider Key​

Create the provider key that stores the Baseten credential and API root:

# Replace with your value
export BASETEN_API_KEY="YOUR_PROVIDER_API_KEY"

PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \
-H "Authorization: Bearer $AISIX_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"display_name": "baseten-prod",
"provider": "baseten",
"api_key": "'"${BASETEN_API_KEY}"'",
"api_base": "https://inference.baseten.co/v1",
"allowed_environments": ["'"${ENV_ID}"'"]
}' | jq -r '.provider_key.id')

echo "$PROVIDER_KEY_ID"

❶ provider is baseten. The AISIX Cloud Admin API derives the adapter from the catalog provider; the adapter field is only accepted on BYO provider keys.

❷ api_key stores the Baseten API key. It follows the credential-handling behavior in Provider Keys.

❸ api_base is the Baseten Model APIs root and already includes the /v1 path. AISIX appends the endpoint path to it, so do not add /chat/completions. Baseten is one of the catalog providers for which the AISIX Cloud Admin API can fill this value in: if you omit api_base, the create still succeeds and resolves to https://inference.baseten.co/v1. Set it explicitly anyway — it is the only way to tell a Model APIs provider key apart from a dedicated-deployment provider key when you later read the resource back.

The command captures the returned provider key ID in PROVIDER_KEY_ID.

Create a Model​

Baseten Model APIs identifiers are namespaced as publisher/model-name, matching the upstream repository that published the weights. They are case-sensitive and the casing is not uniform across publishers: deepseek-ai/DeepSeek-V4-Pro and zai-org/GLM-5.2 use mixed case, while openai/gpt-oss-120b is entirely lowercase. Copy the identifier verbatim from the Baseten Model APIs overview rather than retyping it.

Create the model alias callers will send in requests:

MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \
-H "Authorization: Bearer $AISIX_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"display_name": "baseten-deepseek-v4-pro",
"model_name": "deepseek-ai/DeepSeek-V4-Pro",
"provider_key_id": "'"${PROVIDER_KEY_ID}"'"
}' | jq -r '.model.id')

echo "$MODEL_ID"

❶ display_name is the alias callers send in model. Because Baseten model identifiers contain a slash, a short alias also keeps client configuration readable.

❷ model_name is the Baseten model identifier, for example deepseek-ai/DeepSeek-V4-Pro, openai/gpt-oss-120b, or zai-org/GLM-5.2.

❸ provider_key_id attaches the alias to the Baseten provider key.

Create a Caller API Key​

Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the response — store it securely:

AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \
-H "Authorization: Bearer $AISIX_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"display_name": "baseten-caller",
"allowed_models": ["'"${MODEL_ID}"'"]
}' | jq -r '.plaintext')

echo "$AISIX_API_KEY"

The allowed_models value must reference the model ID captured in the previous step.

After the write, the configuration projects to attached gateways automatically.

Configure with the Open-Source AISIX Gateway​

Export the upstream credential and choose the caller API key that applications will send to the gateway:

export BASETEN_API_KEY="YOUR_PROVIDER_API_KEY"
export CALLER_API_KEY="YOUR_CALLER_API_KEY"

For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its current file, preserving its other resources:

resources.yaml
_format_version: "1"

provider_keys:
- display_name: "baseten-prod"
provider: "baseten"
adapter: "openai"
api_key: ${BASETEN_API_KEY}
api_base: "https://inference.baseten.co/v1"

models:
- display_name: "baseten-deepseek-v4-pro"
provider: "baseten"
model_name: "deepseek-ai/DeepSeek-V4-Pro"
provider_key: "baseten-prod"

api_keys:
- display_name: "baseten-caller"
key_env: CALLER_API_KEY
allowed_models:
- "baseten-deepseek-v4-pro"

If AISIX is installed locally, validate the file before loading it:

aisix validate --resources resources.yaml

After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment.

If you use Docker, adapt the validation and startup commands in the Open-Source AISIX Gateway Quickstart. Mount this resources.yaml file and pass every environment variable it references with -e in both commands. After the resources load, prepare the shared verification request below:

export AISIX_API_KEY="$CALLER_API_KEY"

Verify the Provider Connection​

Export the AISIX gateway origin:

# AISIX_PROXY has no trailing slash or endpoint path
# The local quickstarts use http://127.0.0.1:3000
export AISIX_PROXY="YOUR_AISIX_GATEWAY_URL"

Send a chat-completions request through the AISIX proxy:

curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \
-H "Authorization: Bearer ${AISIX_API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"model": "baseten-deepseek-v4-pro",
"messages": [
{
"role": "user",
"content": "Say hello from Baseten."
}
]
}'

The gateway returns an OpenAI-compatible response that echoes the caller-facing alias baseten-deepseek-v4-pro, not the upstream deepseek-ai/DeepSeek-V4-Pro identifier. If the request fails, check the provider key api_key, api_base, and the Baseten model identifier in model_name — a casing mismatch in model_name is the most common cause, because the identifier is case-sensitive upstream.

Route to a Dedicated Deployment​

A dedicated deployment does not answer on the shared Model APIs host. Baseten gives each deployment its own hostname derived from the model ID, and the OpenAI-compatible routes live under an environment-scoped path segment:

https://model-{model_id}.api.baseten.co/environments/production/sync/v1

See Call your model for the current URL form and the environment names available in your workspace.

Two consequences for gateway configuration:

  • Create a separate provider key for each dedicated deployment, with api_base set to that deployment's URL. The AISIX Cloud Admin API only fills in the shared Model APIs root, so api_base is effectively required here.
  • model_name on the alias is the model name the deployment itself serves, which is usually the original weights identifier rather than a Baseten Model APIs catalog identifier.

Keep the baseten provider value on a dedicated-deployment provider key so usage accounting, metrics, and access logs stay grouped with your other Baseten traffic. A dedicated deployment serves a model the pricing catalog does not cover, so set its rate as an organization pricing override rather than switching provider values. See Model Pricing and Cost Metadata.

Use Reasoning Models​

Several models on Baseten Model APIs emit reasoning traces. AISIX needs no extra configuration for them:

  • Reasoning output. Baseten returns reasoning in reasoning_content, which is already the field AISIX treats as canonical for both streaming and non-streaming responses. Unlike providers that stream reasoning under a vendor-specific delta path, Baseten needs no response.reasoning_field override on the provider key.
  • Reasoning controls. AISIX does not model every request parameter explicitly. Top-level fields it does not recognize are forwarded to the upstream verbatim through the openai adapter, so the per-model reasoning controls Baseten documents — a top-level reasoning_effort for some models, a top-level chat_template_args object for others — reach Baseten unchanged. Check Baseten reasoning for the control each model accepts, because it varies by model rather than by account.

No parameter renames are configured for Baseten, so token-limit fields also pass through as sent. Send the field name the target Baseten model documents.

Endpoint Support​

A Baseten-backed alias works on the normalized chat routes and on /v1/embeddings when the configured api_base serves an embeddings route. Baseten Embeddings Inference deployments expose an OpenAI-compatible /v1/embeddings route, so an embedding alias works when its provider key points at that deployment URL.

Anthropic-style clients calling /v1/messages are served through AISIX translation into the OpenAI request shape, because the alias resolves to the openai adapter. AISIX does not dispatch to a Baseten-native Anthropic-shaped route.

A Baseten model identifier that begins with openai/, such as openai/gpt-oss-120b, does not make the alias an OpenAI-provider model. The provider value is baseten, so routes that gate on the provider identity rather than the adapter — image generation, video generation, and rerank — reject the alias. See Provider Compatibility for the full route matrix.

Next Steps​

You have now connected AISIX to Baseten and verified the model alias. Continue with these guides: