# API7 Docs > Official documentation for AISIX AI Gateway, AISIX Cloud, Apache APISIX, API7 Gateway, API7 and APISIX Ingress Controllers, and gateway plugins. ## ai-gateway Route AI requests through a dedicated gateway layer with centralized provider keys, model aliases, routing, and traffic policy. - [AISIX AI Gateway](https://docs.api7.ai/ai-gateway.md): Route AI requests through a dedicated gateway layer with centralized provider keys, model aliases, routing, and traffic policy. ### reference #### admin-api Inspect resources, model status, and health through the authenticated, read-only Admin API of an open-source AISIX gateway. - [AISIX Admin API](https://docs.api7.ai/ai-gateway/reference/admin-api.md): Inspect resources, model status, and health through the authenticated, read-only Admin API of an open-source AISIX gateway. #### cloud-admin-api AISIX Cloud Admin API reference for organization-scoped automation in AISIX Cloud and AISIX Cloud On-Premises, generated from the AISIX control plane OpenAPI source. - [AISIX Cloud Admin API](https://docs.api7.ai/ai-gateway/reference/cloud-admin-api.md): AISIX Cloud Admin API reference for organization-scoped automation in AISIX Cloud and AISIX Cloud On-Premises, generated from the AISIX control plane OpenAPI source. #### cli Reference for AISIX gateway commands that validate declarative resources and export etcd configuration to a resources.yaml file. - [CLI Reference](https://docs.api7.ai/ai-gateway/reference/cli.md): Reference for AISIX gateway commands that validate declarative resources and export etcd configuration to a resources.yaml file. #### cloud-admin-api-changelog Endpoint and schema differences between released versions of the AISIX Cloud Admin API. - [AISIX Cloud Admin API Changelog](https://docs.api7.ai/ai-gateway/reference/cloud-admin-api-changelog.md): Endpoint and schema differences between released versions of the AISIX Cloud Admin API. #### config-status Reference for AISIX gateway configuration-status endpoints and metrics, including applied, rejected, stale, and partially compatible resources. - [Configuration Status](https://docs.api7.ai/ai-gateway/reference/config-status.md): Reference for AISIX gateway configuration-status endpoints and metrics, including applied, rejected, stale, and partially compatible resources. #### configuration-files Reference for AISIX gateway startup configuration formats, resource sources, connection settings, loading precedence, and common options. - [Startup Configuration Reference](https://docs.api7.ai/ai-gateway/reference/configuration-files.md): Reference for AISIX gateway startup configuration formats, resource sources, connection settings, loading precedence, and common options. #### environment-variables Reference for AISIX AI Gateway environment variables used for configuration files, startup overrides, AISIX gateways, and runtime tuning. - [Environment Variables](https://docs.api7.ai/ai-gateway/reference/environment-variables.md): Reference for AISIX AI Gateway environment variables used for configuration files, startup overrides, AISIX gateways, and runtime tuning. #### headers-and-error-codes Reference for AISIX gateway response headers, status codes, and error envelopes across proxy, MCP, A2A, and passthrough routes. - [Headers and Error Codes](https://docs.api7.ai/ai-gateway/reference/headers-and-error-codes.md): Reference for AISIX gateway response headers, status codes, and error envelopes across proxy, MCP, A2A, and passthrough routes. #### metrics Reference for AISIX AI Gateway Prometheus metrics covering traffic, latency, tokens, cost, caching, guardrails, and upstream health. - [Metrics Reference](https://docs.api7.ai/ai-gateway/reference/metrics.md): Reference for AISIX AI Gateway Prometheus metrics covering traffic, latency, tokens, cost, caching, guardrails, and upstream health. #### on-premises-configuration Reference for Docker Compose environment variables and Helm values used to run the AISIX Cloud control plane in your infrastructure. - [On-Premises Configuration](https://docs.api7.ai/ai-gateway/reference/on-premises-configuration.md): Reference for Docker Compose environment variables and Helm values used to run the AISIX Cloud control plane in your infrastructure. #### ports Reference for default AISIX gateway and AISIX Cloud control-plane ports, traffic direction, configuration, and recommended network exposure. - [Port Reference](https://docs.api7.ai/ai-gateway/reference/ports.md): Reference for default AISIX gateway and AISIX Cloud control-plane ports, traffic direction, configuration, and recommended network exposure. #### proxy-api Reference for AISIX gateway proxy routes, authentication, model discovery, routing behavior, and endpoint constraints. - [Proxy API Reference](https://docs.api7.ai/ai-gateway/reference/proxy-api.md): Reference for AISIX gateway proxy routes, authentication, model discovery, routing behavior, and endpoint constraints. #### resources-file Configure AISIX AI Gateway with resources.yaml using this reference for supported collections, fields, environment interpolation, and validation rules. - [Resources File Reference](https://docs.api7.ai/ai-gateway/reference/resources-file.md): Configure AISIX AI Gateway with resources.yaml using this reference for supported collections, fields, environment interpolation, and validation rules. ### agent-gateway #### agent-access-control Scope each AISIX caller API key to exact A2A agent names, name patterns, or every registered agent. A2A access is denied by default. - [Control Agent Access](https://docs.api7.ai/ai-gateway/agent-gateway/agent-access-control.md): Scope each AISIX caller API key to exact A2A agent names, name patterns, or every registered agent. A2A access is denied by default. #### observability Read the usage events and Prometheus metrics that A2A agent calls emit in AISIX AI Gateway, tagged so you can isolate A2A traffic from model traffic. - [Observability](https://docs.api7.ai/ai-gateway/agent-gateway/observability.md): Read the usage events and Prometheus metrics that A2A agent calls emit in AISIX AI Gateway, tagged so you can isolate A2A traffic from model traffic. #### overview Expose A2A agents through AISIX with caller access control, upstream authentication, rate limits, AISIX Cloud budgets, and telemetry. - [Agent Gateway Overview](https://docs.api7.ai/ai-gateway/agent-gateway/overview.md): Expose A2A agents through AISIX with caller access control, upstream authentication, rate limits, AISIX Cloud budgets, and telemetry. #### setup Register an A2A agent in AISIX Cloud or an open-source AISIX gateway, grant caller access, and verify it with the official A2A Go SDK client. - [Set Up Agent Gateway](https://docs.api7.ai/ai-gateway/agent-gateway/setup.md): Register an A2A agent in AISIX Cloud or an open-source AISIX gateway, grant caller access, and verify it with the official A2A Go SDK client. #### streaming-and-discovery Use the official A2A Go SDK client to test streaming through AISIX Agent Gateway and verify agent-card discovery through a public origin. - [A2A Streaming and Agent-Card Discovery](https://docs.api7.ai/ai-gateway/agent-gateway/streaming-and-discovery.md): Use the official A2A Go SDK client to test streaming through AISIX Agent Gateway and verify agent-card discovery through a public origin. #### traffic-controls Apply caller API key request and concurrency limits to A2A calls in AISIX, and use AISIX Cloud budgets across a caller's traffic. - [Rate Limits and Budgets](https://docs.api7.ai/ai-gateway/agent-gateway/traffic-controls.md): Apply caller API key request and concurrency limits to A2A calls in AISIX, and use AISIX Cloud budgets across a caller's traffic. #### upstream-authentication Configure how AISIX authenticates to an upstream A2A agent with no credential, a bearer token, or an API key held by the gateway. - [Upstream Authentication](https://docs.api7.ai/ai-gateway/agent-gateway/upstream-authentication.md): Configure how AISIX authenticates to an upstream A2A agent with no credential, a bearer token, or an API key held by the gateway. ### cloud #### admin-tokens Create and manage AISIX Cloud admin tokens for organization-level automation, including scopes, expiration, rotation, revocation, and SCIM access. - [Admin Tokens](https://docs.api7.ai/ai-gateway/cloud/admin-tokens.md): Create and manage AISIX Cloud admin tokens for organization-level automation, including scopes, expiration, rotation, revocation, and SCIM access. #### connect-a-gateway Add an AISIX gateway as the data plane for an AISIX Cloud environment with dashboard-issued mTLS credentials, deployment snippets, and connection verification. - [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md): Add an AISIX gateway as the data plane for an AISIX Cloud environment with dashboard-issued mTLS credentials, deployment snippets, and connection verification. #### custom-roles Understand AISIX Cloud owner, admin, member, and custom roles, then assign fine-grained permissions and environment-scoped access across the control plane. - [Roles and Custom Roles](https://docs.api7.ai/ai-gateway/cloud/custom-roles.md): Understand AISIX Cloud owner, admin, member, and custom roles, then assign fine-grained permissions and environment-scoped access across the control plane. #### high-availability Design a highly available AISIX Cloud traffic path with redundant AISIX gateways, resilient control-plane connectivity, and upstream failover. - [High Availability](https://docs.api7.ai/ai-gateway/cloud/high-availability.md): Design a highly available AISIX Cloud traffic path with redundant AISIX gateways, resilient control-plane connectivity, and upstream failover. #### kubernetes Deploy and scale AISIX gateways as AISIX Cloud data planes on Kubernetes using the api7/aisix Helm chart, a HorizontalPodAutoscaler, or KEDA. - [Deploy AISIX Gateways on Kubernetes](https://docs.api7.ai/ai-gateway/cloud/kubernetes.md): Deploy and scale AISIX gateways as AISIX Cloud data planes on Kubernetes using the api7/aisix Helm chart, a HorizontalPodAutoscaler, or KEDA. #### logging-and-auditing Understand how the AISIX Cloud control plane records gateway requests and resource changes. - [Logging and Auditing](https://docs.api7.ai/ai-gateway/cloud/logging-and-auditing.md): Understand how the AISIX Cloud control plane records gateway requests and resource changes. #### members Add organization members for AISIX Cloud control-plane access or API key ownership. - [Members](https://docs.api7.ai/ai-gateway/cloud/members.md): Add organization members for AISIX Cloud control-plane access or API key ownership. #### model-pricing Understand how the AISIX Cloud control plane derives per-request cost, and set prices for models the catalog does not cover. - [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md): Understand how the AISIX Cloud control plane derives per-request cost, and set prices for models the catalog does not cover. #### offline-resilience Understand how AISIX gateways serve cached configuration during control-plane outages, including restart, budget, telemetry, and recovery behavior. - [Offline Resilience](https://docs.api7.ai/ai-gateway/cloud/offline-resilience.md): Understand how AISIX gateways serve cached configuration during control-plane outages, including restart, budget, telemetry, and recovery behavior. #### organizations-and-environments Create and understand AISIX Cloud organizations and environments, including how environment resources reach connected AISIX gateways. - [Organizations and Environments](https://docs.api7.ai/ai-gateway/cloud/organizations-and-environments.md): Create and understand AISIX Cloud organizations and environments, including how environment resources reach connected AISIX gateways. #### overview Learn how AISIX Cloud adds centralized management while AISIX gateways serve AI traffic directly in your environment. - [AISIX Cloud](https://docs.api7.ai/ai-gateway/cloud/overview.md): Learn how AISIX Cloud adds centralized management while AISIX gateways serve AI traffic directly in your environment. #### playground Understand what AISIX Cloud playground requests can validate compared with gateway requests. - [Playground](https://docs.api7.ai/ai-gateway/cloud/playground.md): Understand what AISIX Cloud playground requests can validate compared with gateway requests. #### resource-projection Understand how AISIX Cloud projects environment resources, then verify publication, gateway application, and caller-visible behavior. - [Resource Projection](https://docs.api7.ai/ai-gateway/cloud/resource-projection.md): Understand how AISIX Cloud projects environment resources, then verify publication, gateway application, and caller-visible behavior. #### scim-directory-sync Provision AISIX organization members automatically from your identity provider with SCIM 2.0. - [SCIM Directory Sync](https://docs.api7.ai/ai-gateway/cloud/scim-directory-sync.md): Provision AISIX organization members automatically from your identity provider with SCIM 2.0. #### teams Group AISIX Cloud members into teams, bind caller API keys for attribution, and configure team budgets, rate limits, and MCP access. - [Teams](https://docs.api7.ai/ai-gateway/cloud/teams.md): Group AISIX Cloud members into teams, bind caller API keys for attribution, and configure team budgets, rate limits, and MCP access. #### usage-reporting Understand how the AISIX Cloud control plane reports gateway usage, spend, and budget-related signals. - [Usage Reporting](https://docs.api7.ai/ai-gateway/cloud/usage-reporting.md): Understand how the AISIX Cloud control plane reports gateway usage, spend, and budget-related signals. ### deployment #### configuration-propagation Understand how resource changes reach an AISIX gateway, reload a declarative resources file safely, and verify the applied configuration. - [Configuration Propagation](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md): Understand how resource changes reach an AISIX gateway, reload a declarative resources file safely, and verify the applied configuration. #### forward-proxy Deploy AISIX behind a TLS-terminating egress proxy for governed IDE AI traffic, using GitHub Copilot extensions and Copilot CLI as the worked example. - [Forward Proxy for IDE AI Traffic](https://docs.api7.ai/ai-gateway/deployment/forward-proxy.md): Deploy AISIX behind a TLS-terminating egress proxy for governed IDE AI traffic, using GitHub Copilot extensions and Copilot CLI as the worked example. #### health-checks Choose and interpret AISIX gateway health checks for process liveness, traffic readiness, model health, and configuration state. - [Health Checks](https://docs.api7.ai/ai-gateway/deployment/health-checks.md): Choose and interpret AISIX gateway health checks for process liveness, traffic readiness, model health, and configuration state. #### network-and-security Operate AISIX AI Gateway with correct listener exposure, configuration-source isolation, and credential handling. - [Network and Security](https://docs.api7.ai/ai-gateway/deployment/network-and-security.md): Operate AISIX AI Gateway with correct listener exposure, configuration-source isolation, and credential handling. #### performance-and-sizing Expected AISIX AI Gateway proxy overhead and throughput, and how to size CPU for your request volume. - [Performance and Sizing](https://docs.api7.ai/ai-gateway/deployment/performance-and-sizing.md): Expected AISIX AI Gateway proxy overhead and throughput, and how to size CPU for your request volume. #### production Prepare AISIX AI Gateway for production traffic by planning capacity, dependencies, network exposure, health checks, recovery, and shutdown. - [Production Readiness](https://docs.api7.ai/ai-gateway/deployment/production.md): Prepare AISIX AI Gateway for production traffic by planning capacity, dependencies, network exposure, health checks, recovery, and shutdown. #### startup-configuration Configure AISIX AI Gateway startup settings, including its resource source, listeners, runtime dependencies, observability, and AISIX Cloud connection. - [Startup Configuration](https://docs.api7.ai/ai-gateway/deployment/startup-configuration.md): Configure AISIX AI Gateway startup settings, including its resource source, listeners, runtime dependencies, observability, and AISIX Cloud connection. #### thread-per-core-workers Configure AISIX AI Gateway worker threads and thread-per-core serving, verify the active mode, and plan for per-worker upstream pools and shared listeners. - [Thread-per-Core Workers](https://docs.api7.ai/ai-gateway/deployment/thread-per-core-workers.md): Configure AISIX AI Gateway worker threads and thread-per-core serving, verify the active mode, and plan for per-worker upstream pools and shared listeners. #### tls-and-mtls Configure listener TLS, upstream trust, etcd mTLS, and AISIX Cloud control-plane mTLS for AISIX AI Gateway. - [TLS and mTLS](https://docs.api7.ai/ai-gateway/deployment/tls-and-mtls.md): Configure listener TLS, upstream trust, etcd mTLS, and AISIX Cloud control-plane mTLS for AISIX AI Gateway. #### troubleshooting Diagnose startup, configuration, caller access, policy, upstream, and AISIX gateway failures in AISIX AI Gateway. - [Troubleshooting](https://docs.api7.ai/ai-gateway/deployment/troubleshooting.md): Diagnose startup, configuration, caller access, policy, upstream, and AISIX gateway failures in AISIX AI Gateway. #### url-rewriting Map legacy or external URL shapes onto AISIX AI Gateway endpoints with entry-level rewrite rules, so existing clients migrate without configuration changes. - [URL Rewriting](https://docs.api7.ai/ai-gateway/deployment/url-rewriting.md): Map legacy or external URL shapes onto AISIX AI Gateway endpoints with entry-level rewrite rules, so existing clients migrate without configuration changes. ### endpoints #### anthropic-messages Understand how AISIX AI Gateway handles Anthropic-style Messages requests, including caller keys, model aliases, upstream translation, token counting, and errors. - [Anthropic-Style Messages API](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md): Understand how AISIX AI Gateway handles Anthropic-style Messages requests, including caller keys, model aliases, upstream translation, token counting, and errors. #### audio Learn how AISIX AI Gateway handles OpenAI-style audio transcription, translation, and speech endpoints. - [Speech and Audio](https://docs.api7.ai/ai-gateway/endpoints/audio.md): Learn how AISIX AI Gateway handles OpenAI-style audio transcription, translation, and speech endpoints. #### batch-files-fine-tuning Run the OpenAI-compatible Files, Batch, and Fine-tuning APIs through AISIX AI Gateway, with gateway-managed provider routing and batch usage attribution. - [Batch, Files, and Fine-Tuning](https://docs.api7.ai/ai-gateway/endpoints/batch-files-fine-tuning.md): Run the OpenAI-compatible Files, Batch, and Fine-tuning APIs through AISIX AI Gateway, with gateway-managed provider routing and batch usage attribution. #### chat-audio Send audio to an OpenAI-compatible chat model through AISIX and save generated audio from a non-streaming Chat Completions response. - [Audio Input and Output with Chat Completions](https://docs.api7.ai/ai-gateway/endpoints/chat-audio.md): Send audio to an OpenAI-compatible chat model through AISIX and save generated audio from a non-streaming Chat Completions response. #### embeddings Learn how AISIX AI Gateway handles the OpenAI-compatible embeddings endpoint, including request format and provider limits. - [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md): Learn how AISIX AI Gateway handles the OpenAI-compatible embeddings endpoint, including request format and provider limits. #### image-editing Edit images through the AISIX AI Gateway /v1/images/edits endpoint with multipart uploads, model aliases, prompt guardrails, and token usage accounting. - [Image Editing](https://docs.api7.ai/ai-gateway/endpoints/image-editing.md): Edit images through the AISIX AI Gateway /v1/images/edits endpoint with multipart uploads, model aliases, prompt guardrails, and token usage accounting. #### image-generation Learn how AISIX AI Gateway handles the OpenAI image generation endpoint and provider support. - [Image Generation](https://docs.api7.ai/ai-gateway/endpoints/image-generation.md): Learn how AISIX AI Gateway handles the OpenAI image generation endpoint and provider support. #### openai-client-to-anthropic Use an OpenAI-compatible client with an Anthropic-backed AISIX model alias and understand how the gateway translates requests and responses. - [OpenAI Client with Anthropic Upstream](https://docs.api7.ai/ai-gateway/endpoints/openai-client-to-anthropic.md): Use an OpenAI-compatible client with an Anthropic-backed AISIX model alias and understand how the gateway translates requests and responses. #### openai-compatible-chat Understand how AISIX AI Gateway handles OpenAI-compatible POST /v1/chat/completions requests, model aliases, authentication, provider translation, and errors. - [OpenAI-Compatible Chat Completions](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): Understand how AISIX AI Gateway handles OpenAI-compatible POST /v1/chat/completions requests, model aliases, authentication, provider translation, and errors. #### overview Compare AISIX caller-facing API families and proxy routes, including supported operations, model discovery, health checks, and shared gateway behavior. - [Supported Endpoints](https://docs.api7.ai/ai-gateway/endpoints/overview.md): Compare AISIX caller-facing API families and proxy routes, including supported operations, model discovery, health checks, and shared gateway behavior. #### provider-passthrough Relay provider-native APIs and forward-proxy traffic through AISIX passthrough routes with matching, authentication, credentials, controls, and telemetry. - [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md): Relay provider-native APIs and forward-proxy traffic through AISIX passthrough routes with matching, authentication, credentials, controls, and telemetry. #### realtime Connect OpenAI Realtime WebSocket clients through AISIX AI Gateway with gateway-managed authentication, policy enforcement, and session usage tracking. - [Realtime API](https://docs.api7.ai/ai-gateway/endpoints/realtime.md): Connect OpenAI Realtime WebSocket clients through AISIX AI Gateway with gateway-managed authentication, policy enforcement, and session usage tracking. #### request-lifecycle Understand how AISIX authenticates callers, resolves model aliases, applies controls, routes to providers, and records usage for each AI request. - [Request Lifecycle](https://docs.api7.ai/ai-gateway/endpoints/request-lifecycle.md): Understand how AISIX authenticates callers, resolves model aliases, applies controls, routes to providers, and records usage for each AI request. #### rerank Learn how AISIX AI Gateway proxies rerank requests through supported rerank providers. - [Rerank](https://docs.api7.ai/ai-gateway/endpoints/rerank.md): Learn how AISIX AI Gateway proxies rerank requests through supported rerank providers. #### responses-api Learn how AISIX AI Gateway handles the OpenAI Responses API and provider support. - [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md): Learn how AISIX AI Gateway handles the OpenAI Responses API and provider support. #### streaming Understand streaming behavior on AISIX AI Gateway, including OpenAI-style and Anthropic-style streaming paths. - [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md): Understand streaming behavior on AISIX AI Gateway, including OpenAI-style and Anthropic-style streaming paths. #### text-completions Learn how AISIX AI Gateway handles OpenAI-compatible text-completions requests. - [Text Completions](https://docs.api7.ai/ai-gateway/endpoints/text-completions.md): Learn how AISIX AI Gateway handles OpenAI-compatible text-completions requests. #### tool-calling Understand tool-calling behavior on AISIX AI Gateway, including OpenAI-compatible requests and Anthropic translation. - [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): Understand tool-calling behavior on AISIX AI Gateway, including OpenAI-compatible requests and Anthropic translation. #### video-generation Generate videos through the AISIX AI Gateway /v1/videos endpoint with asynchronous task submission, status polling, and video download. - [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md): Generate videos through the AISIX AI Gateway /v1/videos endpoint with asynchronous task submission, status polling, and video download. ### getting-started #### aisix-cloud-quickstart Deploy an On-Premises AISIX Cloud control plane with Docker Compose, configure a model, and send your first request through an attached AISIX gateway. - [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md): Deploy an On-Premises AISIX Cloud control plane with Docker Compose, configure a model, and send your first request through an attached AISIX gateway. #### anthropic-sdk Configure an Anthropic-compatible client to call AISIX AI Gateway through the /v1/messages proxy API. - [Anthropic SDK](https://docs.api7.ai/ai-gateway/getting-started/anthropic-sdk.md): Configure an Anthropic-compatible client to call AISIX AI Gateway through the /v1/messages proxy API. #### gateway-quickstart Run an open-source AISIX gateway in one container, declare provider keys, models, and caller API keys in a resources.yaml file, and send your first request. - [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md): Run an open-source AISIX gateway in one container, declare provider keys, models, and caller API keys in a resources.yaml file, and send your first request. #### openai-sdk Configure the official OpenAI SDK to call AISIX AI Gateway through the OpenAI-compatible proxy API. - [OpenAI SDK](https://docs.api7.ai/ai-gateway/getting-started/openai-sdk.md): Configure the official OpenAI SDK to call AISIX AI Gateway through the OpenAI-compatible proxy API. #### products-and-deployment-options Compare the open-source AISIX gateway with the On-Premises and Hybrid Cloud control-plane deployment options for AISIX Cloud. - [AISIX Products and Deployment Options](https://docs.api7.ai/ai-gateway/getting-started/products-and-deployment-options.md): Compare the open-source AISIX gateway with the On-Premises and Hybrid Cloud control-plane deployment options for AISIX Cloud. ### integrations Connect SDKs, coding agents, application frameworks, AI application platforms, and voice agent platforms to AISIX AI Gateway. - [Integrations](https://docs.api7.ai/ai-gateway/integrations.md): Connect SDKs, coding agents, application frameworks, AI application platforms, and voice agent platforms to AISIX AI Gateway. #### application-platforms Connect AI application platforms to AISIX AI Gateway for governed model access while the platform continues to run workflows, tools, RAG, and user interfaces. - [AI Application Platforms](https://docs.api7.ai/ai-gateway/integrations/application-platforms.md): Connect AI application platforms to AISIX AI Gateway for governed model access while the platform continues to run workflows, tools, RAG, and user interfaces. - [Dify](https://docs.api7.ai/ai-gateway/integrations/application-platforms/dify.md): Configure Dify's OpenAI-API-compatible model provider plugin to send application and agent model requests through AISIX AI Gateway. - [n8n](https://docs.api7.ai/ai-gateway/integrations/application-platforms/n8n.md): Configure the n8n OpenAI Chat Model to send AI Agent and chain requests through AISIX AI Gateway with a caller key and model alias. - [Open WebUI](https://docs.api7.ai/ai-gateway/integrations/application-platforms/open-webui.md): Connect Open WebUI to AISIX AI Gateway as an OpenAI-compatible endpoint with caller authentication, model discovery, streaming, and tool-call support. #### coding-agents Connect coding agents directly or through a forward proxy to govern model, tool, and official-service traffic with AISIX access controls and telemetry. - [Coding Agents](https://docs.api7.ai/ai-gateway/integrations/coding-agents.md): Connect coding agents directly or through a forward proxy to govern model, tool, and official-service traffic with AISIX access controls and telemetry. - [Claude Code](https://docs.api7.ai/ai-gateway/integrations/coding-agents/claude-code.md): Configure Claude Code to send Anthropic Messages traffic through AISIX AI Gateway. - [Cline](https://docs.api7.ai/ai-gateway/integrations/coding-agents/cline.md): Configure Cline to send OpenAI-compatible model traffic through AISIX AI Gateway. - [Codex](https://docs.api7.ai/ai-gateway/integrations/coding-agents/codex.md): Configure Codex to send Responses API traffic through AISIX AI Gateway. - [Cursor](https://docs.api7.ai/ai-gateway/integrations/coding-agents/cursor.md): Configure Cursor to send OpenAI-compatible chat traffic through AISIX AI Gateway. #### frameworks - [CrewAI](https://docs.api7.ai/ai-gateway/integrations/frameworks/crewai.md): Configure CrewAI to send OpenAI-compatible chat traffic through AISIX AI Gateway. - [Haystack](https://docs.api7.ai/ai-gateway/integrations/frameworks/haystack.md): Configure Haystack to send OpenAI Responses API traffic through AISIX AI Gateway. - [Instructor](https://docs.api7.ai/ai-gateway/integrations/frameworks/instructor.md): Configure Instructor to send OpenAI Responses API structured-output requests through AISIX AI Gateway. - [LangChain and LangGraph](https://docs.api7.ai/ai-gateway/integrations/frameworks/langchain.md): Configure LangChain and LangGraph to send OpenAI Responses API traffic through AISIX AI Gateway. - [LlamaIndex](https://docs.api7.ai/ai-gateway/integrations/frameworks/llamaindex.md): Configure LlamaIndex to send OpenAI Responses API traffic through AISIX AI Gateway. - [Microsoft Agent Framework](https://docs.api7.ai/ai-gateway/integrations/frameworks/microsoft-agent-framework.md): Configure Microsoft Agent Framework to send OpenAI Responses API traffic through AISIX AI Gateway. - [OpenAI Agents SDK](https://docs.api7.ai/ai-gateway/integrations/frameworks/openai-agents-sdk.md): Configure OpenAI Agents SDK to send OpenAI Responses API traffic through AISIX AI Gateway. - [Pydantic AI](https://docs.api7.ai/ai-gateway/integrations/frameworks/pydantic-ai.md): Configure Pydantic AI to send OpenAI Responses API traffic through AISIX AI Gateway. - [Vercel AI SDK](https://docs.api7.ai/ai-gateway/integrations/frameworks/vercel-ai-sdk.md): Configure Vercel AI SDK to send OpenAI Responses API traffic through AISIX AI Gateway. #### voice-agents Connect voice agent platforms to AISIX AI Gateway so their text model requests use gateway authentication, model aliases, routing, policy, and telemetry. - [Voice Agent Platforms](https://docs.api7.ai/ai-gateway/integrations/voice-agents.md): Connect voice agent platforms to AISIX AI Gateway so their text model requests use gateway authentication, model aliases, routing, policy, and telemetry. - [ElevenLabs Agents](https://docs.api7.ai/ai-gateway/integrations/voice-agents/elevenlabs.md): Configure an ElevenLabs Agent to use AISIX AI Gateway as its Custom LLM endpoint with a restricted caller key and model alias. - [LiveKit Agents](https://docs.api7.ai/ai-gateway/integrations/voice-agents/livekit.md): Configure the LiveKit Agents OpenAI plugin to stream Chat Completions requests through AISIX AI Gateway with a caller key and model alias. - [Pipecat](https://docs.api7.ai/ai-gateway/integrations/voice-agents/pipecat.md): Configure Pipecat OpenAILLMService to stream Chat Completions requests through AISIX AI Gateway while Pipecat operates the voice pipeline. - [Vapi](https://docs.api7.ai/ai-gateway/integrations/voice-agents/vapi.md): Configure a Vapi voice assistant to use AISIX AI Gateway as an authenticated OpenAI-compatible Custom LLM endpoint. ### mcp-gateway #### access-policies Manage MCP tool access across AISIX Cloud environments and teams, and combine those layers with each caller API key's own grant. - [Manage MCP Access with Policies](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md): Manage MCP tool access across AISIX Cloud environments and teams, and combine those layers with each caller API key's own grant. #### client-authentication Choose how MCP clients authenticate to AISIX AI Gateway — with a gateway API key, through OAuth sign-in, or anonymously from trusted networks. - [Client Authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/client-authentication.md): Choose how MCP clients authenticate to AISIX AI Gateway — with a gateway API key, through OAuth sign-in, or anonymously from trusted networks. #### cursor Connect Cursor to AISIX MCP Gateway with a caller API key, discover the permitted tools, and verify a complete tool call through AISIX. - [Connect Cursor to MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/cursor.md): Connect Cursor to AISIX MCP Gateway with a caller API key, discover the permitted tools, and verify a complete tool call through AISIX. #### guardrails Inspect MCP tool-call arguments and results with AISIX guardrails, blocking unsafe calls before dispatch or withholding unsafe results from clients. - [Guardrails](https://docs.api7.ai/ai-gateway/mcp-gateway/guardrails.md): Inspect MCP tool-call arguments and results with AISIX guardrails, blocking unsafe calls before dispatch or withholding unsafe results from clients. #### observability Read the usage events and Prometheus metrics that MCP tool calls emit in AISIX AI Gateway, tagged so you can isolate MCP traffic from model traffic. - [Observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md): Read the usage events and Prometheus metrics that MCP tool calls emit in AISIX AI Gateway, tagged so you can isolate MCP traffic from model traffic. #### openapi-servers Register a REST API with its OpenAPI document so AISIX AI Gateway generates MCP tools from its operations and executes tool calls as HTTP requests. - [Expose a REST API as MCP Tools](https://docs.api7.ai/ai-gateway/mcp-gateway/openapi-servers.md): Register a REST API with its OpenAPI document so AISIX AI Gateway generates MCP tools from its operations and executes tool calls as HTTP requests. #### overview Expose upstream MCP servers through AISIX AI Gateway with caller API key access control, upstream authentication, guardrails, telemetry, and automatic MCP protocol version negotiation. - [MCP Gateway Overview](https://docs.api7.ai/ai-gateway/mcp-gateway/overview.md): Expose upstream MCP servers through AISIX AI Gateway with caller API key access control, upstream authentication, guardrails, telemetry, and automatic MCP protocol version negotiation. #### server-review Review MCP server registrations and changes in AISIX Cloud before publishing them to gateways, and revoke approvals that no longer hold. - [Review and Approve MCP Servers](https://docs.api7.ai/ai-gateway/mcp-gateway/server-review.md): Review MCP server registrations and changes in AISIX Cloud before publishing them to gateways, and revoke approvals that no longer hold. #### setup Register an upstream MCP server in AISIX Cloud or an open-source AISIX gateway, grant one tool, and verify allowed and denied MCP calls. - [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md): Register an upstream MCP server in AISIX Cloud or an open-source AISIX gateway, grant one tool, and verify allowed and denied MCP calls. #### tool-access-control Scope each caller API key to the MCP tools it may list and call in AISIX AI Gateway, using exact names, per-server wildcards, or a global wildcard. - [Control Tool Access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md): Scope each caller API key to the MCP tools it may list and call in AISIX AI Gateway, using exact names, per-server wildcards, or a global wildcard. #### traffic-controls Configure MCP request and concurrency limits in AISIX AI Gateway, plus AISIX Cloud budgets, using the caller API key boundary. - [Rate Limits and Budgets](https://docs.api7.ai/ai-gateway/mcp-gateway/traffic-controls.md): Configure MCP request and concurrency limits in AISIX AI Gateway, plus AISIX Cloud budgets, using the caller API key boundary. #### upstream-authentication Configure how AISIX AI Gateway authenticates to each upstream MCP server with no credential, a bearer token, an API key, or OAuth 2.0 client credentials. - [Upstream Authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/upstream-authentication.md): Configure how AISIX AI Gateway authenticates to each upstream MCP server with no credential, a bearer token, an API key, or OAuth 2.0 client credentials. #### vscode Connect Visual Studio Code to AISIX MCP Gateway with a protected caller key, discover permitted tools, and verify a complete tool call. - [Connect VS Code to MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/vscode.md): Connect Visual Studio Code to AISIX MCP Gateway with a protected caller key, discover permitted tools, and verify a complete tool call. ### models #### model-aliases Configure direct and wildcard model aliases in AISIX gateway deployments, including retries, pricing, rate limits, and other model-level behavior. - [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): Configure direct and wildcard model aliases in AISIX gateway deployments, including retries, pricing, rate limits, and other model-level behavior. #### provider-key-rotation Rotate upstream provider credentials across AISIX gateway deployments with in-place and gradual workflows that keep caller API keys and model aliases stable. - [Provider Key Rotation](https://docs.api7.ai/ai-gateway/models/provider-key-rotation.md): Rotate upstream provider credentials across AISIX gateway deployments with in-place and gradual workflows that keep caller API keys and model aliases stable. #### provider-keys Configure provider keys across AISIX gateway deployments for upstream credentials, endpoints, adapters, request headers, compatibility overrides, and rotation. - [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md): Configure provider keys across AISIX gateway deployments for upstream credentials, endpoints, adapters, request headers, compatibility overrides, and rotation. #### reasoning-effort-mapping Rewrite reasoning-effort values for each direct model without changing client requests. - [Reasoning Effort Mapping](https://docs.api7.ai/ai-gateway/models/reasoning-effort-mapping.md): Rewrite reasoning-effort values for each direct model without changing client requests. #### resource-model Understand how provider keys, models, caller API keys, routing, traffic controls, and policies work together across AISIX gateway deployments. - [Resource Model](https://docs.api7.ai/ai-gateway/models/resource-model.md): Understand how provider keys, models, caller API keys, routing, traffic controls, and policies work together across AISIX gateway deployments. #### upstream-request-headers Forward approved caller headers to an upstream on any AISIX proxy face, inject request-context headers on a provider key, and understand which headers AISIX never relays. - [Upstream Request Headers](https://docs.api7.ai/ai-gateway/models/upstream-request-headers.md): Forward approved caller headers to an upstream on any AISIX proxy face, inject request-context headers on a provider key, and understand which headers AISIX never relays. ### observability #### exporters Send AISIX gateway request telemetry to OTLP collectors, object storage, Datadog, or Alibaba Cloud SLS through either configuration path. - [Observability Exporters](https://docs.api7.ai/ai-gateway/observability/exporters.md): Send AISIX gateway request telemetry to OTLP collectors, object storage, Datadog, or Alibaba Cloud SLS through either configuration path. #### load-logs-into-snowflake Export AISIX gateway request telemetry to Amazon S3 or Azure Blob, ingest the records with Snowpipe, and query them through a Snowflake view. - [Load Request Telemetry into Snowflake](https://docs.api7.ai/ai-gateway/observability/load-logs-into-snowflake.md): Export AISIX gateway request telemetry to Amazon S3 or Azure Blob, ingest the records with Snowpipe, and query them through a Snowflake view. #### metrics-and-logs Monitor AISIX gateway traffic with Prometheus metrics, structured access logs, response headers, and exportable per-attempt usage events. - [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md): Monitor AISIX gateway traffic with Prometheus metrics, structured access logs, response headers, and exportable per-attempt usage events. ### on-premises #### deployment Plan production resources and install the AISIX Cloud control plane in your infrastructure with Docker Compose, Helm, or an offline package. - [On-Premises Installation](https://docs.api7.ai/ai-gateway/on-premises/deployment.md): Plan production resources and install the AISIX Cloud control plane in your infrastructure with Docker Compose, Helm, or an offline package. #### external-database Provision external PostgreSQL and a least-privilege role for an AISIX Cloud control plane installed with Helm, without granting superuser access. - [External Database](https://docs.api7.ai/ai-gateway/on-premises/external-database.md): Provision external PostgreSQL and a least-privilege role for an AISIX Cloud control plane installed with Helm, without granting superuser access. ### providers #### adapters Understand the OpenAI, Anthropic, Amazon Bedrock, Google Vertex AI, and Azure OpenAI protocol adapter families available in AISIX gateway deployments. - [Adapter Protocol Families](https://docs.api7.ai/ai-gateway/providers/adapters.md): Understand the OpenAI, Anthropic, Amazon Bedrock, Google Vertex AI, and Azure OpenAI protocol adapter families available in AISIX gateway deployments. #### amazon-nova Connect the direct Amazon Nova API to AISIX gateway deployments. Configure API credentials, Nova model aliases, caller API keys, and model access controls. - [Amazon Nova API](https://docs.api7.ai/ai-gateway/providers/amazon-nova.md): Connect the direct Amazon Nova API to AISIX gateway deployments. Configure API credentials, Nova model aliases, caller API keys, and model access controls. #### anthropic Connect Anthropic Claude to AISIX gateway deployments through the native Messages or OpenAI-compatible API. Configure credentials, model aliases, and caller access. - [Anthropic](https://docs.api7.ai/ai-gateway/providers/anthropic.md): Connect Anthropic Claude to AISIX gateway deployments through the native Messages or OpenAI-compatible API. Configure credentials, model aliases, and caller access. #### aws-bedrock Connect AWS Bedrock to AISIX gateway deployments using SigV4. Configure regional endpoints, model or inference-profile aliases, caller keys, and access controls. - [AWS Bedrock](https://docs.api7.ai/ai-gateway/providers/aws-bedrock.md): Connect AWS Bedrock to AISIX gateway deployments using SigV4. Configure regional endpoints, model or inference-profile aliases, caller keys, and access controls. #### azure-openai Connect Azure OpenAI to AISIX gateway deployments using resource API keys or Microsoft Entra ID. Configure deployment aliases, caller keys, and access controls. - [Azure OpenAI](https://docs.api7.ai/ai-gateway/providers/azure-openai.md): Connect Azure OpenAI to AISIX gateway deployments using resource API keys or Microsoft Entra ID. Configure deployment aliases, caller keys, and access controls. #### baseten Connect Baseten Model APIs or dedicated endpoints to AISIX gateway deployments. Configure credentials, model aliases, caller keys, and access controls. - [Baseten](https://docs.api7.ai/ai-gateway/providers/baseten.md): Connect Baseten Model APIs or dedicated endpoints to AISIX gateway deployments. Configure credentials, model aliases, caller keys, and access controls. #### bring-your-own-endpoint Connect AISIX gateway deployments to private OpenAI-compatible endpoints such as vLLM, SGLang, or Ollama. Configure credentials, aliases, and custom token pricing. - [Bring Your Own Endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md): Connect AISIX gateway deployments to private OpenAI-compatible endpoints such as vLLM, SGLang, or Ollama. Configure credentials, aliases, and custom token pricing. #### cerebras Connect Cerebras Inference to AISIX gateway deployments through its OpenAI-compatible API. Configure credentials, model aliases, caller keys, and access controls. - [Cerebras](https://docs.api7.ai/ai-gateway/providers/cerebras.md): Connect Cerebras Inference to AISIX gateway deployments through its OpenAI-compatible API. Configure credentials, model aliases, caller keys, and access controls. #### cloudflare-workers-ai Connect Cloudflare Workers AI to AISIX gateway deployments. Configure account-scoped endpoints, API credentials, @cf model IDs, aliases, and caller access. - [Cloudflare Workers AI](https://docs.api7.ai/ai-gateway/providers/cloudflare-workers-ai.md): Connect Cloudflare Workers AI to AISIX gateway deployments. Configure account-scoped endpoints, API credentials, @cf model IDs, aliases, and caller access. #### cohere Connect Cohere to AISIX gateway deployments for Command chat models, embeddings, and reranking. Configure provider credentials, model aliases, and caller access. - [Cohere](https://docs.api7.ai/ai-gateway/providers/cohere.md): Connect Cohere to AISIX gateway deployments for Command chat models, embeddings, and reranking. Configure provider credentials, model aliases, and caller access. #### compatibility Compare provider protocol and proxy endpoint support across AISIX gateway deployments, including chat, embeddings, media, Realtime, and job APIs. - [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): Compare provider protocol and proxy endpoint support across AISIX gateway deployments, including chat, embeddings, media, Realtime, and job APIs. #### databricks Connect Databricks Model Serving to AISIX gateway deployments. Configure workspace endpoints, provider credentials, model aliases, caller keys, and access controls. - [Databricks](https://docs.api7.ai/ai-gateway/providers/databricks.md): Connect Databricks Model Serving to AISIX gateway deployments. Configure workspace endpoints, provider credentials, model aliases, caller keys, and access controls. #### deepinfra Connect DeepInfra to AISIX gateway deployments through its OpenAI-compatible API. Configure explicit endpoints, credentials, model aliases, and caller access. - [DeepInfra](https://docs.api7.ai/ai-gateway/providers/deepinfra.md): Connect DeepInfra to AISIX gateway deployments through its OpenAI-compatible API. Configure explicit endpoints, credentials, model aliases, and caller access. #### deepseek Connect DeepSeek models to AISIX gateway deployments through the OpenAI-compatible API. Configure provider credentials, model aliases, caller keys, and access controls. - [DeepSeek](https://docs.api7.ai/ai-gateway/providers/deepseek.md): Connect DeepSeek models to AISIX gateway deployments through the OpenAI-compatible API. Configure provider credentials, model aliases, caller keys, and access controls. #### digitalocean-gradient-ai Connect DigitalOcean Gradient AI to AISIX gateway deployments. Configure serverless-inference credentials, model aliases, caller API keys, and access controls. - [DigitalOcean Gradient AI](https://docs.api7.ai/ai-gateway/providers/digitalocean-gradient-ai.md): Connect DigitalOcean Gradient AI to AISIX gateway deployments. Configure serverless-inference credentials, model aliases, caller API keys, and access controls. #### fireworks-ai Connect Fireworks AI to AISIX gateway deployments through its OpenAI-compatible API. Configure provider credentials, model aliases, caller keys, and access controls. - [Fireworks AI](https://docs.api7.ai/ai-gateway/providers/fireworks-ai.md): Connect Fireworks AI to AISIX gateway deployments through its OpenAI-compatible API. Configure provider credentials, model aliases, caller keys, and access controls. #### gemini Connect Google Gemini to AISIX gateway deployments through the AI Studio OpenAI-compatible API. Configure credentials, model aliases, caller keys, and access controls. - [Gemini (Google AI Studio)](https://docs.api7.ai/ai-gateway/providers/gemini.md): Connect Google Gemini to AISIX gateway deployments through the AI Studio OpenAI-compatible API. Configure credentials, model aliases, caller keys, and access controls. #### google-vertex-ai Connect Google Vertex AI to AISIX. Configure service-account credentials, global or regional endpoints, model aliases, and understand adapter limits. - [Google Vertex AI](https://docs.api7.ai/ai-gateway/providers/google-vertex-ai.md): Connect Google Vertex AI to AISIX. Configure service-account credentials, global or regional endpoints, model aliases, and understand adapter limits. #### groq Connect Groq to AISIX. Configure credentials and model aliases, then understand chat, Responses, audio, batch, and provider-specific limits. - [Groq](https://docs.api7.ai/ai-gateway/providers/groq.md): Connect Groq to AISIX. Configure credentials and model aliases, then understand chat, Responses, audio, batch, and provider-specific limits. #### huggingface Connect Hugging Face Inference Providers to AISIX. Configure router credentials and model aliases, then understand route and provider boundaries. - [Hugging Face](https://docs.api7.ai/ai-gateway/providers/huggingface.md): Connect Hugging Face Inference Providers to AISIX. Configure router credentials and model aliases, then understand route and provider boundaries. #### jina Connect Jina AI embeddings and rerankers to AISIX gateway deployments, including model, route, passthrough, and accounting boundaries. - [Jina](https://docs.api7.ai/ai-gateway/providers/jina.md): Connect Jina AI embeddings and rerankers to AISIX gateway deployments, including model, route, passthrough, and accounting boundaries. #### meta-llama-api Connect the Meta Llama API to AISIX through its OpenAI-compatible endpoint, with the correct API base, example model, and route boundaries. - [Meta Llama API](https://docs.api7.ai/ai-gateway/providers/meta-llama-api.md): Connect the Meta Llama API to AISIX through its OpenAI-compatible endpoint, with the correct API base, example model, and route boundaries. #### minimax Connect MiniMax models to AISIX gateway deployments through the OpenAI-compatible API. Configure the explicit endpoint, credentials, aliases, and caller access. - [MiniMax](https://docs.api7.ai/ai-gateway/providers/minimax.md): Connect MiniMax models to AISIX gateway deployments through the OpenAI-compatible API. Configure the explicit endpoint, credentials, aliases, and caller access. #### mistral Connect Mistral AI to AISIX for chat and compatible endpoints, with model aliases, structured-response boundaries, and passthrough routes. - [Mistral AI](https://docs.api7.ai/ai-gateway/providers/mistral.md): Connect Mistral AI to AISIX for chat and compatible endpoints, with model aliases, structured-response boundaries, and passthrough routes. #### modelscope Connect ModelScope API-Inference to AISIX gateway deployments. Configure OpenAI-compatible endpoints, namespaced model IDs, provider keys, and caller access. - [ModelScope](https://docs.api7.ai/ai-gateway/providers/modelscope.md): Connect ModelScope API-Inference to AISIX gateway deployments. Configure OpenAI-compatible endpoints, namespaced model IDs, provider keys, and caller access. #### moonshotai Connect Moonshot AI and Kimi models to AISIX gateway deployments. Configure regional endpoints, credentials, model aliases, thinking mode, and caller access. - [Moonshot AI (Kimi)](https://docs.api7.ai/ai-gateway/providers/moonshotai.md): Connect Moonshot AI and Kimi models to AISIX gateway deployments. Configure regional endpoints, credentials, model aliases, thinking mode, and caller access. #### nebius-token-factory Connect Nebius Token Factory to AISIX through its OpenAI-compatible API, with model aliases, endpoint compatibility, and passthrough routes for native APIs. - [Nebius Token Factory](https://docs.api7.ai/ai-gateway/providers/nebius-token-factory.md): Connect Nebius Token Factory to AISIX through its OpenAI-compatible API, with model aliases, endpoint compatibility, and passthrough routes for native APIs. #### novita-ai Connect Novita AI to AISIX through its OpenAI-compatible API, with model aliases, endpoint compatibility, batch workflows, and passthrough routes. - [Novita AI](https://docs.api7.ai/ai-gateway/providers/novita-ai.md): Connect Novita AI to AISIX through its OpenAI-compatible API, with model aliases, endpoint compatibility, batch workflows, and passthrough routes. #### nvidia-nim Connect hosted or privately deployed NVIDIA NIM models to AISIX gateway deployments. Configure endpoint credentials, namespaced model aliases, and caller access. - [NVIDIA NIM](https://docs.api7.ai/ai-gateway/providers/nvidia-nim.md): Connect hosted or privately deployed NVIDIA NIM models to AISIX gateway deployments. Configure endpoint credentials, namespaced model aliases, and caller access. #### ollama Connect a local or private Ollama server to AISIX gateway deployments through its OpenAI-compatible API. Configure endpoints, model aliases, and caller access. - [Ollama](https://docs.api7.ai/ai-gateway/providers/ollama.md): Connect a local or private Ollama server to AISIX gateway deployments through its OpenAI-compatible API. Configure endpoints, model aliases, and caller access. #### openai Connect OpenAI chat models and Sora video generation to AISIX gateway deployments. Configure credentials, model aliases, caller API keys, access controls, rate limits, and request verification. - [OpenAI](https://docs.api7.ai/ai-gateway/providers/openai.md): Connect OpenAI chat models and Sora video generation to AISIX gateway deployments. Configure credentials, model aliases, caller API keys, access controls, rate limits, and request verification. #### openai-compatible-vendors Configure public OpenAI-compatible LLM providers in AISIX gateway deployments using explicit API endpoints, provider credentials, model aliases, and caller keys. - [Other OpenAI-Compatible Providers](https://docs.api7.ai/ai-gateway/providers/openai-compatible-vendors.md): Configure public OpenAI-compatible LLM providers in AISIX gateway deployments using explicit API endpoints, provider credentials, model aliases, and caller keys. #### openrouter Connect OpenRouter to AISIX gateway deployments through its OpenAI-compatible API. Configure credentials, namespaced model IDs, aliases, and caller access. - [OpenRouter](https://docs.api7.ai/ai-gateway/providers/openrouter.md): Connect OpenRouter to AISIX gateway deployments through its OpenAI-compatible API. Configure credentials, namespaced model IDs, aliases, and caller access. #### overview Compare hosted AI provider APIs and private model servers supported across AISIX gateway deployments, including native adapters, community providers, Ollama, and vLLM. - [Choose a Provider Upstream](https://docs.api7.ai/ai-gateway/providers/overview.md): Compare hosted AI provider APIs and private model servers supported across AISIX gateway deployments, including native adapters, community providers, Ollama, and vLLM. #### ovhcloud-ai-endpoints Connect OVHcloud AI Endpoints to AISIX gateway deployments through the OpenAI-compatible API. Configure credentials, model aliases, caller keys, and access controls. - [OVHcloud AI Endpoints](https://docs.api7.ai/ai-gateway/providers/ovhcloud-ai-endpoints.md): Connect OVHcloud AI Endpoints to AISIX gateway deployments through the OpenAI-compatible API. Configure credentials, model aliases, caller keys, and access controls. #### perplexity Connect Perplexity Sonar models to AISIX gateway deployments through the OpenAI-compatible API. Configure credentials, model aliases, caller keys, and access controls. - [Perplexity](https://docs.api7.ai/ai-gateway/providers/perplexity.md): Connect Perplexity Sonar models to AISIX gateway deployments through the OpenAI-compatible API. Configure credentials, model aliases, caller keys, and access controls. #### qwen Connect Alibaba Cloud Qwen chat models and Wan video generation to AISIX gateway deployments. Configure regional DashScope endpoints, credentials, model aliases, and caller access. - [Qwen (Alibaba Cloud)](https://docs.api7.ai/ai-gateway/providers/qwen.md): Connect Alibaba Cloud Qwen chat models and Wan video generation to AISIX gateway deployments. Configure regional DashScope endpoints, credentials, model aliases, and caller access. #### runwayml Connect RunwayML video generation to AISIX gateway deployments through the AISIX video API. Configure credentials, model aliases, caller keys, and access controls. - [RunwayML](https://docs.api7.ai/ai-gateway/providers/runwayml.md): Connect RunwayML video generation to AISIX gateway deployments through the AISIX video API. Configure credentials, model aliases, caller keys, and access controls. #### siliconflow Connect SiliconFlow to AISIX gateway deployments through its OpenAI-compatible API. Configure credentials, namespaced model IDs, aliases, and caller access. - [SiliconFlow](https://docs.api7.ai/ai-gateway/providers/siliconflow.md): Connect SiliconFlow to AISIX gateway deployments through its OpenAI-compatible API. Configure credentials, namespaced model IDs, aliases, and caller access. #### snowflake-cortex Connect Snowflake Cortex to AISIX gateway deployments. Configure account-specific OpenAI-compatible endpoints, credentials, model aliases, and caller access. - [Snowflake Cortex](https://docs.api7.ai/ai-gateway/providers/snowflake-cortex.md): Connect Snowflake Cortex to AISIX gateway deployments. Configure account-specific OpenAI-compatible endpoints, credentials, model aliases, and caller access. #### together Connect Together AI to AISIX gateway deployments through its OpenAI-compatible API. Configure provider credentials, model aliases, caller keys, and access controls. - [Together AI](https://docs.api7.ai/ai-gateway/providers/together.md): Connect Together AI to AISIX gateway deployments through its OpenAI-compatible API. Configure provider credentials, model aliases, caller keys, and access controls. #### vllm Connect a private vLLM OpenAI-compatible server to AISIX gateway deployments. Configure endpoint credentials, served model aliases, caller keys, and access controls. - [vLLM](https://docs.api7.ai/ai-gateway/providers/vllm.md): Connect a private vLLM OpenAI-compatible server to AISIX gateway deployments. Configure endpoint credentials, served model aliases, caller keys, and access controls. #### volcengine-ark Connect Volcengine Ark to AISIX gateway deployments for Doubao chat and Seedance video generation. Configure credentials, model aliases, and caller access. - [Volcengine Ark (Doubao)](https://docs.api7.ai/ai-gateway/providers/volcengine-ark.md): Connect Volcengine Ark to AISIX gateway deployments for Doubao chat and Seedance video generation. Configure credentials, model aliases, and caller access. #### wandb-inference Connect Weights & Biases Inference to AISIX gateway deployments through its OpenAI-compatible API. Configure credentials, model aliases, and caller access. - [Weights & Biases Inference](https://docs.api7.ai/ai-gateway/providers/wandb-inference.md): Connect Weights & Biases Inference to AISIX gateway deployments through its OpenAI-compatible API. Configure credentials, model aliases, and caller access. #### xai Connect xAI Grok models to AISIX gateway deployments through the OpenAI-compatible API. Configure regional endpoints, credentials, aliases, and caller access. - [xAI (Grok)](https://docs.api7.ai/ai-gateway/providers/xai.md): Connect xAI Grok models to AISIX gateway deployments through the OpenAI-compatible API. Configure regional endpoints, credentials, aliases, and caller access. #### zhipuai Connect Zhipu AI to AISIX gateway deployments for GLM chat and CogVideoX video generation. Configure API credentials, model aliases, and caller access. - [Zhipu AI (GLM)](https://docs.api7.ai/ai-gateway/providers/zhipuai.md): Connect Zhipu AI to AISIX gateway deployments for GLM chat and CogVideoX video generation. Configure API credentials, model aliases, and caller access. ### release-notes New features, improvements, and fixes in each AISIX AI Gateway release. - [Release Notes](https://docs.api7.ai/ai-gateway/release-notes.md): New features, improvements, and fixes in each AISIX AI Gateway release. ### routing #### ensemble-models Configure ensemble models in AISIX gateway deployments to fan out chat requests, synthesize panel responses, and manage latency, cost, and failures. - [Ensemble Models](https://docs.api7.ai/ai-gateway/routing/ensemble-models.md): Configure ensemble models in AISIX gateway deployments to fan out chat requests, synthesize panel responses, and manage latency, cost, and failures. #### proxy-errors-and-retries Understand how client applications should handle AISIX AI Gateway proxy errors, retry signals, and upstream failures. - [Proxy Errors and Retries](https://docs.api7.ai/ai-gateway/routing/proxy-errors-and-retries.md): Understand how client applications should handle AISIX AI Gateway proxy errors, retry signals, and upstream failures. #### routing-and-failover Configure and test multi-target routing and failover in AISIX gateway deployments, including selection strategies, retries, runtime filtering, and recovery. - [Multi-Target Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): Configure and test multi-target routing and failover in AISIX gateway deployments, including selection strategies, retries, runtime filtering, and recovery. #### semantic-routing Configure semantic routing in AISIX gateway deployments to select direct models by request meaning, tune similarity thresholds, and define safe fallbacks. - [Semantic Routing](https://docs.api7.ai/ai-gateway/routing/semantic-routing.md): Configure semantic routing in AISIX gateway deployments to select direct models by request meaning, tune similarity thresholds, and define safe fallbacks. ### traffic-controls #### budget-alerts Configure AISIX Cloud budget alerts to notify operators through webhooks or Slack when AI spending crosses a defined percentage of a budget limit. - [Budget Alerts and Notifications](https://docs.api7.ai/ai-gateway/traffic-controls/budget-alerts.md): Configure AISIX Cloud budget alerts to notify operators through webhooks or Slack when AI spending crosses a defined percentage of a budget limit. #### budgets Configure AISIX Cloud budgets to enforce AI spending limits across organizations, environments, caller API keys, provider keys, teams, and members. - [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md): Configure AISIX Cloud budgets to enforce AI spending limits across organizations, environments, caller API keys, provider keys, teams, and members. #### caching Configure memory or Redis response caching for eligible Chat Completions requests in AISIX Cloud and open-source AISIX gateway deployments. - [Response Caching](https://docs.api7.ai/ai-gateway/traffic-controls/caching.md): Configure memory or Redis response caching for eligible Chat Completions requests in AISIX Cloud and open-source AISIX gateway deployments. #### caller-api-keys Configure caller API keys for AISIX gateway deployments, including model access controls, expiration, disablement, rotation, authentication, and verification. - [Caller API Keys](https://docs.api7.ai/ai-gateway/traffic-controls/caller-api-keys.md): Configure caller API keys for AISIX gateway deployments, including model access controls, expiration, disablement, rotation, authentication, and verification. #### claim-mappings Map verified JWT claims in AISIX to caller API keys with priority-ordered rules, shared access controls, and per-identity usage attribution. - [JWT Claim Mappings](https://docs.api7.ai/ai-gateway/traffic-controls/claim-mappings.md): Map verified JWT claims in AISIX to caller API keys with priority-ordered rules, shared access controls, and per-identity usage attribution. #### guardrails - [Alibaba Cloud Content Moderation](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/alibaba-cloud-ai.md): Configure Alibaba Cloud Content Moderation through AISIX Cloud or a resources file, then verify TextModerationPlus risk-level enforcement. - [Alibaba Cloud AI Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/alibaba-cloud-ai-guardrails.md): Configure Alibaba Cloud AI Guardrails through AISIX Cloud or a resources file, then verify MultiModalGuard blocking and sensitive-data masking. - [AWS Bedrock Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/aws-bedrock.md): Configure AWS Bedrock Guardrails through AISIX Cloud or a resources file, then verify blocking and PII anonymization in AISIX AI Gateway. - [Azure AI Content Safety Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/azure-content-safety.md): Configure Azure AI Content Safety through AISIX Cloud or a resources file, then verify Prompt Shield and Text Moderation behavior in AISIX. - [Guardrail Behavior](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md): Understand AISIX guardrail hook points, enforcement modes, scope differences, streaming output controls, remote failures, and caller behavior. - [Custom Script Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/custom-script.md): Configure custom JavaScript guardrails through AISIX Cloud or a resources file to run screening logic, including calls to content-policy services. - [Built-in Keyword Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/keyword.md): Configure built-in AISIX keyword guardrails through AISIX Cloud or a resources file, then verify blocking, monitor mode, and scoped behavior. - [Lakera Guard](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/lakera.md): Configure Lakera Guard through AISIX Cloud or a resources file, then verify prompt-injection blocking and PII masking in AISIX AI Gateway. - [OpenAI Moderation Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/openai-moderation.md): Configure OpenAI Moderation through AISIX Cloud or a resources file, then verify blocking and tune per-category thresholds in AISIX AI Gateway. - [Choosing a Guardrail Provider](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/overview.md): Compare built-in and remote AISIX guardrail providers by evaluation location, detection coverage, enforcement actions, and configuration path. - [PII Detection and Redaction](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/pii.md): Configure built-in AISIX PII detection through AISIX Cloud or a resources file, then verify masking, blocking, and custom sensitive-data patterns. - [Presidio Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/presidio.md): Deploy the Presidio analyzer and anonymizer in your own infrastructure, connect them to AISIX as a guardrail, and verify PII anonymization and blocking. - [Qwen3Guard Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/qwen3guard.md): Run the open-source Qwen3Guard safety classifier in your own infrastructure and connect it to AISIX with a custom script guardrail that screens requests and model responses. - [Semantic Screening Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/semantic-screening.md): Configure and verify AISIX semantic screening guardrails for requests and responses, including deny policies, topic allow lists, and threshold setup. - [Calibrate Semantic Screening Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/semantic-screening-calibration.md): Calibrate AISIX semantic screening thresholds with sample and real-traffic scores, interpret request telemetry, and review upgraded policies. #### jwt-authentication Configure OIDC trust and JWT authentication in AISIX gateway deployments, including issuer discovery, claim validation, and external identity mapping to caller API keys. - [JWT Authentication](https://docs.api7.ai/ai-gateway/traffic-controls/jwt-authentication.md): Configure OIDC trust and JWT authentication in AISIX gateway deployments, including issuer discovery, claim validation, and external identity mapping to caller API keys. #### keycloak-integration A verified end-to-end walkthrough — configure a Keycloak realm so each user authenticates at the gateway with their own JWT, and map departments and groups to caller API keys with claim mappings. - [Keycloak Integration](https://docs.api7.ai/ai-gateway/traffic-controls/keycloak-integration.md): A verified end-to-end walkthrough — configure a Keycloak realm so each user authenticates at the gateway with their own JWT, and map departments and groups to caller API keys with claim mappings. #### overview Understand how caller identity, rate limits, AISIX Cloud budgets, guardrails, and caching govern requests through AISIX. - [Traffic Controls](https://docs.api7.ai/ai-gateway/traffic-controls/overview.md): Understand how caller identity, rate limits, AISIX Cloud budgets, guardrails, and caching govern requests through AISIX. #### prompt-caching Enable automatic Anthropic prompt caching in AISIX Cloud or an open-source AISIX gateway so repeated prompt prefixes receive provider cache discounts. - [Anthropic Prompt Caching](https://docs.api7.ai/ai-gateway/traffic-controls/prompt-caching.md): Enable automatic Anthropic prompt caching in AISIX Cloud or an open-source AISIX gateway so repeated prompt prefixes receive provider cache discounts. #### rate-limit-policies Configure AISIX rate limit policies with conditional traffic matching, independent quota buckets, classic single-scope rules, and suspension schedules. - [Rate Limit Policies](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limit-policies.md): Configure AISIX rate limit policies with conditional traffic matching, independent quota buckets, classic single-scope rules, and suspension schedules. #### rate-limits Configure and verify request, token, and concurrency limits for caller API keys and models in AISIX Cloud and the open-source AISIX gateway. - [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md): Configure and verify request, token, and concurrency limits for caller API keys and models in AISIX Cloud and the open-source AISIX gateway. #### semantic-caching Serve cached Chat Completions responses for semantically similar prompts using embedding similarity, with configurable thresholds, sharing scopes, and purging. - [Semantic Caching](https://docs.api7.ai/ai-gateway/traffic-controls/semantic-caching.md): Serve cached Chat Completions responses for semantically similar prompts using embedding similarity, with configurable thresholds, sharing scopes, and purging. ## api7-gateway ### reference Find API7 Gateway reference documentation for APIs, CLI tools, deployment configuration, security controls, and configuration syntax. - [Reference Overview](https://docs.api7.ai/api7-gateway/reference.md): Find API7 Gateway reference documentation for APIs, CLI tools, deployment configuration, security controls, and configuration syntax. #### admin-api API7 Enterprise Admin APIs are RESTful APIs that allow you to create and manage API7 resources. - [API7 Enterprise Admin APIs](https://docs.api7.ai/api7-gateway/reference/admin-api.md): API7 Enterprise Admin APIs are RESTful APIs that allow you to create and manage API7 resources. #### developer-portal-api API7 Enterprise Admin APIs are RESTful APIs that allow you to create and manage API7 resources. - [API7 Enterprise Developer Portal APIs](https://docs.api7.ai/api7-gateway/reference/developer-portal-api.md): API7 Enterprise Admin APIs are RESTful APIs that allow you to create and manage API7 resources. #### a7-cli Use the separately installed a7 CLI to manage API7 Gateway groups and resources interactively, in scripts, or through API7 Gateway Agent Skills. - [a7 CLI](https://docs.api7.ai/api7-gateway/reference/a7-cli.md): Use the separately installed a7 CLI to manage API7 Gateway groups and resources interactively, in scripts, or through API7 Gateway Agent Skills. #### adc Use API Declarative CLI (ADC) to manage API7 Gateway configuration declaratively. - [API Declarative CLI (ADC)](https://docs.api7.ai/api7-gateway/reference/adc.md): Use API Declarative CLI (ADC) to manage API7 Gateway configuration declaratively. #### alert-template Customize alert notifications with pre-defined variables in API7 Gateway, enabling dynamic content in alert messages and emails. - [Alert Variables and Templates](https://docs.api7.ai/api7-gateway/reference/alert-template.md): Customize alert notifications with pre-defined variables in API7 Gateway, enabling dynamic content in alert messages and emails. #### approval-variables Use API7 Gateway approval variables to customize API product subscription notification email content and webhook messages with current template values. - [Approval Notification Variables and Templates](https://docs.api7.ai/api7-gateway/reference/approval-variables.md): Use API7 Gateway approval variables to customize API product subscription notification email content and webhook messages with current template values. #### built-in-variables Discover the built-in variables available in API7 Gateway, including NGINX and APISIX variables, which can be utilized for route matching, log customization, and plugin configurations. - [Built-In Variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md): Discover the built-in variables available in API7 Gateway, including NGINX and APISIX variables, which can be utilized for route matching, log customization, and plugin configurations. #### configuration Understand the configuration files used in API7 Gateway, including default and user-defined files for managing settings effectively. - [Configuration Files](https://docs.api7.ai/api7-gateway/reference/configuration.md): Understand the configuration files used in API7 Gateway, including default and user-defined files for managing settings effectively. #### environment-variables Explore the use of environment variables in API7 Gateway for configuring consumer credentials, SSL certificates, and plugins. - [Environment Variables](https://docs.api7.ai/api7-gateway/reference/environment-variables.md): Explore the use of environment variables in API7 Gateway for configuring consumer credentials, SSL certificates, and plugins. #### expressions Understand how to use expressions in API7 Gateway for route matching, request filtering, and conditional logic in configurations. - [API7 Expressions](https://docs.api7.ai/api7-gateway/reference/expressions.md): Understand how to use expressions in API7 Gateway for route matching, request filtering, and conditional logic in configurations. #### hardening Learn about securing sensitive information in API7 Gateway, including storage, encryption, and communication practices to protect against threats. - [Security Hardening Reference](https://docs.api7.ai/api7-gateway/reference/hardening.md): Learn about securing sensitive information in API7 Gateway, including storage, encryption, and communication practices to protect against threats. #### helm-chart Learn where to find API7 Gateway Helm chart values and how Helm values are rendered into gateway configuration. - [Helm Chart](https://docs.api7.ai/api7-gateway/reference/helm-chart.md): Learn where to find API7 Gateway Helm chart values and how Helm values are rendered into gateway configuration. #### obtain-dashboard-token Create a token in the API7 Dashboard and use it for API7 Gateway Admin API and ADC authentication. - [Obtain a Token from the Dashboard](https://docs.api7.ai/api7-gateway/reference/obtain-dashboard-token.md): Create a token in the API7 Dashboard and use it for API7 Gateway Admin API and ADC authentication. #### permission-policy-action-and-resource Complete reference for all permission policy actions and ARN-style resources in API7 Gateway — organized by namespace (gateway, iam, portal) for building least-privilege policies. - [Permission Policy Actions and Resources](https://docs.api7.ai/api7-gateway/reference/permission-policy-action-and-resource.md): Complete reference for all permission policy actions and ARN-style resources in API7 Gateway — organized by namespace (gateway, iam, portal) for building least-privilege policies. #### permission-policy-examples Ready-to-adapt permission policy examples for API7 Gateway, organized by access pattern, service and plugin operations, IAM and governance, and portal management. - [Permission Policy Examples](https://docs.api7.ai/api7-gateway/reference/permission-policy-examples.md): Ready-to-adapt permission policy examples for API7 Gateway, organized by access pattern, service and plugin operations, IAM and governance, and portal management. ### 3.9.x #### ai-gateway - [Proxy Your First LLM Request in 5 Minutes](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/get-started.md): Set up API7 AI Gateway and proxy your first request to OpenAI in under 5 minutes. Step-by-step guide with code examples. - [Connect to Anthropic Claude](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/llm-providers/anthropic.md): Route Anthropic Claude API traffic through API7 Gateway for centralized security, rate limiting, and observability. - [Integrate Azure OpenAI Service](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/llm-providers/azure-openai.md): Manage Azure OpenAI deployments through API7 Gateway. Handle resource names, API versions, and auth centrally. - [Route Traffic to DeepSeek Models](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/llm-providers/deepseek.md): Proxy DeepSeek API requests through API7 Gateway. Manage authentication, enable failover, and monitor usage centrally. - [Integrate Google Gemini](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/llm-providers/google-gemini.md): Proxy Google Gemini API requests through API7 Gateway. Manage API keys and monitor AI traffic centrally. - [Route Traffic to OpenAI](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/llm-providers/openai.md): Proxy and secure OpenAI API requests through API7 Gateway. Centralize authentication, enable failover, and monitor usage. - [Connect Any OpenAI-Compatible LLM](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/llm-providers/openai-compatible.md): Proxy any OpenAI-compatible API through API7 Gateway. Connect self-hosted models, custom endpoints, or niche providers. - [Access Hundreds of LLMs via OpenRouter](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/llm-providers/openrouter.md): Use OpenRouter with API7 AI Gateway to access 200+ LLMs through one API while maintaining enterprise security controls. - [Route Enterprise AI Traffic to Vertex AI](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/llm-providers/vertex-ai.md): Securely proxy Google Cloud Vertex AI requests through API7 Gateway with service account auth and regional routing. - [Manage and Secure AI Traffic](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/overview.md): Centralize LLM access with API7 AI Gateway. Route traffic to multiple providers, enforce guardrails, and control costs from one platform. - [Monitor AI Traffic and Track LLM Costs](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/use-cases/ai-observability-and-cost-tracking.md): Gain visibility into LLM usage, token consumption, latency, and costs with API7 AI Gateway's observability features. - [Transform API Requests with AI-Powered Rewriting](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/use-cases/ai-request-transformation.md): Use LLMs to intelligently transform, enrich, or restructure API requests and responses at the gateway layer. - [Enforce AI Guardrails and Protect PII](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/use-cases/content-safety-and-guardrails.md): Block prompt injection, detect toxicity, and redact PII before requests reach LLMs using API7 AI Gateway guardrails. - [Expose REST APIs as MCP Tools for AI Agents](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/use-cases/expose-apis-as-mcp-tools.md): Convert existing OpenAPI services into MCP-compatible tools so AI agents can discover and invoke your APIs automatically. - [Manage API7 Enterprise from an AI Client with API7-MCP](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/use-cases/manage-api7-with-mcp.md): Deploy the API7-MCP server so an AI client such as Cursor, Claude Desktop, or Cline can read API7 Enterprise resources, check Prometheus metrics, manage RBAC, and send test traffic through the gateway. - [Set Up Multi-LLM Routing and Automatic Fallback](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/use-cases/multi-llm-routing-and-fallback.md): Route AI traffic across multiple LLM providers with weighted load balancing, automatic failover, and health checks. - [Implement Prompt Templates and Decorators](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/use-cases/prompt-engineering-and-templating.md): Standardize LLM interactions with reusable prompt templates and automatic system prompt injection using API7 AI Gateway. - [Convert Anthropic Messages to OpenAI Chat Completions](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/use-cases/protocol-conversion.md): Use API7 AI Gateway to transparently convert Anthropic Messages API requests to the OpenAI Chat Completions API format, enabling teams to use the Anthropic SDK with any OpenAI-compatible backend. - [Implement RAG at the Gateway Layer](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/use-cases/retrieval-augmented-generation.md): Enhance LLM responses with relevant context using Retrieval-Augmented Generation (RAG) built into API7 AI Gateway. - [Control AI Costs with Token-Based Rate Limiting](https://docs.api7.ai/api7-gateway/3.9.x/ai-gateway/use-cases/token-rate-limiting-and-quota-management.md): Implement token-based rate limits to prevent LLM abuse and control AI costs per route and model instance. #### configure-and-manage - [Run Benchmarks on AWS EKS](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/benchmark-on-aws-eks.md): Reproduce the published API7 Gateway performance benchmark on AWS EKS. Walkthrough covers EKS cluster setup, three isolated node groups, Helm install, NGINX upstream and wrk2 deployment, and running the full scenario suite. - [Configuration Reference for API7 Gateway Control Plane](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/configure-control-plane.md): Detailed configuration reference for the API7 Gateway Control Plane, covering the Dashboard and DP Manager configuration files. - [Configuration Reference for API7 Gateway Data Plane](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/configure-data-plane.md): Detailed configuration reference for API7 Gateway Data Plane, based on the config-default.yaml structure. - [Data Plane Resilience](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/data-plane-resilience.md): Configure fallback storage for API7 Gateway data plane nodes so they can restart and continue operating during extended control plane outages. - [Multiple Availability Zones Deployment of API7 Gateway](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/deployment-scenarios/multi-az-deployment.md): Configuration and architecture for deploying API7 Gateway across multiple availability zones for high availability and fault tolerance. - [Multi-Region Deployment Patterns](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/deployment-scenarios/multi-region-deployment.md): Architecture and considerations for deploying API7 Gateway across multiple geographic regions for global reach and disaster recovery. - [Data Plane High Availability](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/high-availability-data-plane.md): Design a highly available API7 Gateway data plane with multiple nodes, health checks, and load balancer failover. - [Labels](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/labels.md): Organize and filter API7 Gateway resources at scale with labels — key-value metadata attached to gateway groups, services, routes, consumers, and other entities for team, environment, and application-level segmentation. - [License Management](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/license-management.md): Manage your API7 Gateway license, understand core quotas and license states, and configure license file paths for automated deployment. - [Performance Benchmark](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/performance-benchmark.md): Published performance benchmark results for API7 Gateway (AWS EKS and single-host baselines), plus methodology and optimization guidance for running your own benchmarks accurately. - [Production Best Practices](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/production-best-practices.md): Operational best practices for managing API7 Gateway in production, including GitOps, change management, and disaster recovery. - [Running in Production](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/run-in-production.md): Pre-production checklist and deployment guide for API7 Gateway to ensure a stable and secure production environment. - [Scale Data Plane](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/scale-data-plane.md): Scale API7 Gateway data plane nodes horizontally to increase throughput and prepare for high-availability deployments. - [Shared Memory Sizing](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/shared-dict-sizing.md): Size API7 Gateway shared memory zones for metrics, service discovery, the Developer Portal, and tracing so they do not overflow at your deployment's scale. - [Optimize Telemetry Data Transfer](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/telemetry-opt-out.md): Configure telemetry data transfer between the data plane and control plane, including compression levels and how to disable telemetry. - [User Management](https://docs.api7.ai/api7-gateway/3.9.x/configure-and-manage/user-management.md): Manage users, roles, and permission policies in API7 Enterprise. #### developer-portal - [Configure the Developer Portal](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/deploy/configure-portal.md): Configure the Developer Portal settings, including public access, portal tokens, built-in authentication, and SCIM provisioning. - [Customize the Developer Portal](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/deploy/customize-portal.md): Customize the Developer Portal branding, theme, authentication providers, and functionality using the API7 Developer Portal Boilerplate. - [Deploy the Developer Portal](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/deploy/deploy-portal.md): Deploy the API7 Developer Portal with its two official images — the Portal API backend and the customer-facing frontend — on Docker Compose or Kubernetes, and connect them to your API7 control plane. - [Browse APIs](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/guides/browse-apis.md): Discover and explore available API products in the Developer Portal's API Hub. - [Create an Application](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/guides/create-application.md): Create an application in the Developer Portal to group your API subscriptions and credentials. - [Manage Credentials](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/guides/manage-credentials.md): Create, view, regenerate, and delete credentials in the Developer Portal for authenticating API requests. - [Manage Your Organization](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/guides/manage-organization.md): Manage your organization in the Developer Portal, including inviting members, assigning roles, and switching between organizations. - [Register and Log In](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/guides/register-and-login.md): Create a developer account and log in to the Developer Portal using email/password, SSO, or an organization invitation. - [Subscribe to an API](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/guides/subscribe-to-api.md): Subscribe your application to an API product to gain access to consume APIs through the Developer Portal. - [Try an API](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/guides/try-api.md): Test API endpoints directly from the Developer Portal using the built-in Try It Out feature. - [API Products](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/key-concepts/api-products.md): Understand API products in the Developer Portal, including product types, visibility settings, authentication options, and the publishing lifecycle. - [Applications](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/key-concepts/applications.md): Understand applications in the Developer Portal, which group subscriptions and credentials for a specific project or use case. - [Credentials](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/key-concepts/credentials.md): Understand credentials in the Developer Portal, including supported authentication types (key auth, basic auth, OAuth/DCR) and credential lifecycle management. - [Developers](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/key-concepts/developers.md): Understand developers in the Developer Portal, including registration methods, account states, and the difference between developers and consumers. - [Subscriptions](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/key-concepts/subscriptions.md): Understand subscriptions in the Developer Portal, including the approval workflow, status transitions, and auto-approval configuration. - [Configure Dynamic Client Registration (DCR)](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/manage/configure-dcr.md): Configure DCR providers to enable developers to register OAuth 2.0 clients through the Developer Portal. - [Configure SCIM Provisioning for a Custom Developer Portal with Okta](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/manage/configure-scim.md): Configure SCIM provisioning with Okta for a custom Developer Portal based on the API7 Developer Portal Boilerplate. - [Configure SSO for the Developer Portal](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/manage/configure-sso.md): Configure Single Sign-On (SSO) for the Developer Portal using OIDC, SAML, LDAP, or CAS identity providers. - [Manage API Products](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/manage/manage-api-products.md): Create, configure, publish, and manage API products in the Provider Portal for developer consumption through the Developer Portal. - [Manage Applications](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/manage/manage-applications.md): Manage developer applications in the Developer Portal, including lifecycle, structure, and how to call the Developer Portal backend programmatically. - [Manage Developers](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/manage/manage-developers.md): Manage developer accounts using the Provider Portal Admin API and the standalone Developer Portal backend, including listing, creating, approving registrations, and deleting developers. - [Manage Subscriptions](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/manage/manage-subscriptions.md): Manage API product subscriptions in the Provider Portal, including approving, rejecting, and cancelling subscription requests. - [Developer Portal Overview](https://docs.api7.ai/api7-gateway/3.9.x/developer-portal/overview.md): Learn about the API7 Developer Portal, a platform for API providers to publish API products and for developers to discover, subscribe to, and consume APIs. #### enterprise-features - [Alerts and Contact Points](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/alerts-and-contact-points.md): Explore the concept of alerts and contact points in API7 Gateway, which monitor exceptions and send timely notifications. - [Anonymous Consumers](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/anonymous-consumers.md): Explore the concept of anonymous consumers in API7 Gateway, allowing non-authenticated access to APIs while maintaining security. - [API Portal](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/api-portal.md): Explore the concept of the API portal in API7 Gateway, providing a centralized space for developers to access and manage APIs. - [Audit and Rollback](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/audit-and-rollback.md): Explore the concept of audit logging and rollback in API7 Gateway, tracking user actions and restoring previous configurations. - [Compliance](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/compliance.md): Explore the concept of compliance in API7 Gateway, helping organizations meet regulatory requirements and maintain security standards. - [Credentials](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/credentials.md): Explore the concept of credentials in API7 Gateway, authenticating users and ensuring secure access while facilitating management and rotation. - [Custom Plugins](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/custom-plugins.md): Explore the concept of custom plugins in API7 Gateway, enabling tailored extensions to meet specific business needs. - [Dashboard SSO Options](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/dashboard-sso.md): Explore the concept of Single Sign-On (SSO) in API7 Gateway, allowing users to authenticate with existing credentials for easy access. - [Enterprise Plugins](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/enterprise-plugins.md): Explore the concept of enterprise plugins in API7 Gateway, which extend functionality and enhance capabilities for API management. - [Gateway Groups](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/gateway-groups.md): Explore the concept of gateway groups in API7 Gateway, which manage multiple API gateway instances with shared configurations. - [High Availability](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/high-availability.md): Explore the concept of high availability in API7 Gateway, ensuring continuous service delivery for mission-critical applications. - [Organization and RBAC](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/organization-and-rbac.md): Explore the concept of organization management and RBAC in API7 Gateway, enabling fine-grained permission management. - [Enterprise Features Overview](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/overview.md): Discover the enterprise-grade features that distinguish API7 Gateway from open-source Apache APISIX, including centralized management, RBAC, audit logging, and professional support. - [Permission Policies and Boundaries](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/permission-policies-and-boundaries.md): Explore the concept of permission policies and boundaries in API7 Gateway, defining user access levels for enhanced security. - [Secret Providers](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/secret-providers.md): Explore the concept of secret providers in API7 Gateway, which enhance security by storing sensitive data using third-party tools. - [Security Hardening](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/security-hardening.md): Explore the concept of security hardening in API7 Gateway, designed to protect API infrastructure against threats and vulnerabilities. - [Service Release](https://docs.api7.ai/api7-gateway/3.9.x/enterprise-features/service-release.md): Explore the concept of service release in API7 Gateway, detailing the creation, configuration, and deployment of APIs effectively. #### getting-started - [Learn More About API7 Products and APISIX](https://docs.api7.ai/api7-gateway/3.9.x/getting-started/learn-more.md): Understand the API7 product family, the relationship between API7 Enterprise and Apache APISIX, and how they fit into your API management strategy. - [API7 Gateway Overview](https://docs.api7.ai/api7-gateway/3.9.x/getting-started/overview.md): Get started with API7 Gateway. Understand the platform components, how they work together, and choose the right path for your use case. - [Quick Start](https://docs.api7.ai/api7-gateway/3.9.x/getting-started/quick-start.md): Get API7 Gateway running locally with Docker Compose and proxy your first API request in under 10 minutes. - [Tutorial: Proxying and Managing API Requests via Plugins](https://docs.api7.ai/api7-gateway/3.9.x/getting-started/tutorial-proxying-api-requests.md): A hands-on tutorial that walks you through proxying API requests, adding authentication with Key Auth, and enabling rate limiting with the Limit Count plugin. #### how-to-guides - [Configure Basic Authentication](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/api-security/basic-auth.md): Learn how to secure your APIs by requiring clients to provide a standard username and password in the HTTP Authorization header. - [Configure Data Masking](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/api-security/data-masking.md): Use the data-mask plugin to redact, replace, or remove sensitive fields from request data before it is written to access logs and logger plugin output, helping you meet GDPR, HIPAA, and PCI-DSS requirements. - [Forward External Auth User Info to Upstream](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/api-security/forward-auth-user-info.md): Forward authenticated user information from OpenID Connect or SAML routes to upstream services as request headers or the consumer name. - [Configure HMAC Authentication](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/api-security/hmac-auth.md): Learn how to secure your APIs with Hash-based Message Authentication Code (HMAC) for request signing in API7 Enterprise. - [Configure JWT Authentication](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/api-security/jwt-auth.md): Learn how to secure your APIs using JSON Web Tokens (JWT) for stateless authentication in API7 Enterprise. - [Configure Key Authentication](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/api-security/key-auth.md): Learn how to secure your APIs by requiring clients to provide a unique API key in the request header or query string. - [Configure Readiness and Liveness Probes](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/ops/configure-readiness-probe.md): Configure readiness and liveness checks for API7 Gateway data planes in Kubernetes, Docker, and other non-Helm deployments. - [Create a Custom Role](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/ops/create-custom-role.md): Create a custom role in API7 Gateway by defining permission policies, attaching them to a role, and assigning the role to a user. - [Design a Custom Role System](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/ops/design-custom-role-system.md): Design a scalable custom role system in API7 Gateway by combining roles, permission policies, labels, and permission boundaries. - [Manage Gateway Groups](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/ops/multi-gateway-group.md): Learn how to create and manage multiple gateway groups to organize and isolate API traffic across environments or teams. - [Configure Secret Management](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/ops/secret-manager.md): Learn how to configure secret providers and reference external secrets in API7 Enterprise without hardcoding sensitive values in gateway resources. - [How-To Guides](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/overview.md): Practical how-to guides for common API7 Gateway tasks and workflows. - [Plugin Development Best Practices](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/plugin-development/best-practices.md): Write custom Lua plugins that are correct, fast, and maintainable by using the built-in core library instead of low-level OpenResty APIs, validating configuration with a schema, and choosing the right phase and priority. - [Develop Custom Lua Plugins](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/plugin-development/custom-lua-plugins.md): Learn how to develop, register, and test custom Lua plugins for API7 Gateway. - [Serverless Functions or Custom Plugins](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/plugin-development/serverless-or-custom-plugins.md): Compare the built-in serverless-function plugins with custom Lua plugins, and choose the right way to run your own Lua logic in API7 Gateway. - [Configure GraphQL Proxying](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/protocol-proxy/graphql-proxy.md): Learn how to proxy GraphQL APIs through API7 Gateway and when to add GraphQL-aware plugins for rate limiting and caching. - [Configure gRPC Proxying](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/protocol-proxy/grpc-proxy.md): Learn how to configure API7 Enterprise to proxy gRPC traffic, including REST-to-gRPC transcoding and gRPC-Web support for browser clients. - [Configure TCP/UDP Proxying](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/protocol-proxy/tcp-udp-proxy.md): Configure API7 Gateway to proxy TCP and UDP (Layer 4) traffic to upstream services such as databases, message queues, and custom protocols. - [Configure WebSocket Proxying](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/protocol-proxy/websocket-proxy.md): Learn how to enable and configure WebSocket proxying for your routes in API7 Enterprise to handle long-lived, bidirectional connections. - [Implement Blue-Green Deployment](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/blue-green-deployment.md): Implement blue-green deployments using API7 Gateway to switch traffic between two upstream environments with zero downtime. - [Implement Canary Release](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/canary-release.md): Learn how to gradually shift traffic to a new version of your backend service using the traffic-split plugin. - [Configure CORS](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/cors.md): Learn how to configure Cross-Origin Resource Sharing (CORS) for your APIs to allow or restrict access from different domains. - [Configure Fault Injection](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/fault-injection.md): Learn how to test the resilience of your application by injecting HTTP errors and response delays. - [Conditionally Disable Global Plugins](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/global-plugin-exemption.md): Conditionally skip global plugin execution for specific routes using route labels and the _meta.filter mechanism. - [Configure Upstream Health Checks](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/health-check.md): Learn how to configure active and passive health checks for upstream services to ensure high availability. - [Configure Proxy Cache](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/proxy-cache.md): Learn how to improve API performance and reduce upstream load by caching responses at the Gateway. - [Configure Proxy Mirror](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/proxy-mirror.md): Learn how to duplicate and send a percentage of real production traffic to a secondary service for testing and verification. - [Rewrite Proxy Requests](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/proxy-rewrite.md): Learn how to modify request URIs, methods, and headers before proxying them to upstream services. - [Configure Rate Limiting](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/rate-limiting.md): Learn how to configure simple and advanced rate limiting for your APIs to prevent abuse and ensure fair usage. - [Configure Response Rewrite](https://docs.api7.ai/api7-gateway/3.9.x/how-to-guides/traffic-management/response-rewrite.md): Learn how to modify response status codes, headers, and body content before they are returned to the client. #### install - [Deploy for High Availability](https://docs.api7.ai/api7-gateway/3.9.x/install/deploy-high-availability.md): Deploy API7 Gateway control plane and data plane in a high-availability configuration to eliminate single points of failure. - [Deploy on Kubernetes](https://docs.api7.ai/api7-gateway/3.9.x/install/deploy-on-kubernetes.md): Deploy API7 Enterprise on Kubernetes using Helm, including control plane setup, data plane configuration with mTLS, and cloud-specific guidance for AWS EKS, GCP GKE, and Azure AKS. - [Deploy on OpenShift](https://docs.api7.ai/api7-gateway/3.9.x/install/deploy-on-openshift.md): Deploy API7 Gateway on Red Hat OpenShift with proper Security Context Constraints (SCCs), service accounts, and Helm chart configuration. - [Deploy with Docker Compose](https://docs.api7.ai/api7-gateway/3.9.x/install/deploy-with-docker-compose.md): Deploy API7 Enterprise locally with Docker Compose. This development and testing setup brings up the control plane with PostgreSQL, Prometheus, Jaeger, the integrated Dashboard, and the DP Manager, then adds a data plane gateway with a dashboard-generated Docker command. - [Installation FAQ](https://docs.api7.ai/api7-gateway/3.9.x/install/installation-faq.md): Frequently asked questions and troubleshooting common installation issues. - [Installation Packages](https://docs.api7.ai/api7-gateway/3.9.x/install/installation-packages.md): Container images, Helm charts, and CLI tools used to install API7 Enterprise — including the official Docker Hub repositories for the gateway, DP manager, Developer Portal, and integrated control plane images. - [Install API7 Gateway On-Premises](https://docs.api7.ai/api7-gateway/3.9.x/install/overview.md): Overview of deployment options for API7 Enterprise on-premises infrastructure. - [Supported Versions and Interoperability](https://docs.api7.ai/api7-gateway/3.9.x/install/supported-versions-and-interoperability.md): Version compatibility matrix for API7 Enterprise components and infrastructure. - [System Requirements](https://docs.api7.ai/api7-gateway/3.9.x/install/system-requirements.md): Hardware, operating system, and network requirements for installing API7 Gateway. #### key-concepts - [Architecture](https://docs.api7.ai/api7-gateway/3.9.x/key-concepts/architecture.md): Deep dive into the decoupled control plane and data plane architecture of API7 Enterprise. - [Consumers and Credentials](https://docs.api7.ai/api7-gateway/3.9.x/key-concepts/consumers-and-credentials.md): Identity management and authentication using Consumers and Credentials. - [Gateway Groups](https://docs.api7.ai/api7-gateway/3.9.x/key-concepts/gateway-groups.md): Logical grouping of data plane instances for environment isolation and configuration management. - [Key Concepts Overview](https://docs.api7.ai/api7-gateway/3.9.x/key-concepts/overview.md): Introduction to the core abstractions and entities in API7 Enterprise. - [Plugins](https://docs.api7.ai/api7-gateway/3.9.x/key-concepts/plugins.md): Modular components to intercept and modify API traffic in API7 Gateway. - [Service Discovery](https://docs.api7.ai/api7-gateway/3.9.x/key-concepts/service-discovery.md): Dynamic upstream resolution in API7 Gateway — automatically discover backend endpoints from Kubernetes Services, Nacos, and Consul registries without hardcoding IPs. - [Services and Routes](https://docs.api7.ai/api7-gateway/3.9.x/key-concepts/services-and-routes.md): Understand how to group and route API traffic using Services and Routes. - [SSL Certificates](https://docs.api7.ai/api7-gateway/3.9.x/key-concepts/ssl-certificates.md): Understanding how API7 Gateway manages SSL/TLS certificates and terminates secure connections. - [Stream Routes](https://docs.api7.ai/api7-gateway/3.9.x/key-concepts/stream-routes.md): Understand stream routes in API7 Gateway for proxying TCP and UDP (Layer 4) traffic, including matching rules, supported plugins, and the relationship with stream services. - [Upstreams and Load Balancing](https://docs.api7.ai/api7-gateway/3.9.x/key-concepts/upstreams-and-load-balancing.md): Definitions of backend targets, including load balancing algorithms and health checks. #### observability - [Configure Alerts](https://docs.api7.ai/api7-gateway/3.9.x/observability/alerts.md): Configure alert policies and contact points in API7 Gateway to get notified by email or webhook when gateway instances go offline, certificates expire, or error rates breach a threshold. - [Include Consumer Labels in Access Logs](https://docs.api7.ai/api7-gateway/3.9.x/observability/consumer-label-based-logging.md): Include consumer labels in API7 Gateway access logs for per-consumer traffic tracking and analysis. - [Capture Request Traces with Debug Sessions](https://docs.api7.ai/api7-gateway/3.9.x/observability/debug-sessions.md): Find out exactly what the gateway did to a single request — which plugins ran, in what order, which one was slow, and what it logged. Capture on demand, or automatically when an alert fires. - [Use an Existing Prometheus](https://docs.api7.ai/api7-gateway/3.9.x/observability/external-prometheus.md): Point API7 Gateway at an existing external Prometheus and remove the bundled instance from a Docker Compose deployment, keeping the Dashboard Monitoring page working throughout. - [Send Kubernetes Error Logs to Splunk](https://docs.api7.ai/api7-gateway/3.9.x/observability/kubernetes-error-log-forwarding.md): Forward API7 Gateway error logs from Kubernetes to Splunk by using the Splunk OpenTelemetry Collector. - [Configure Centralized Logging](https://docs.api7.ai/api7-gateway/3.9.x/observability/logging.md): Understand the access and error logs that API7 Gateway produces, how to configure their format and verbosity, and how to forward them to a centralized log management system. - [Monitor Metrics](https://docs.api7.ai/api7-gateway/3.9.x/observability/metrics.md): Monitor API7 Gateway metrics via the built-in Dashboard page or by scraping the Prometheus endpoint exposed by the data plane. - [Send Access Logs to Splunk](https://docs.api7.ai/api7-gateway/3.9.x/observability/splunk-integration.md): Send API7 Gateway access logs to Splunk using the HTTP Event Collector (HEC) logging plugin. - [Configure Distributed Tracing](https://docs.api7.ai/api7-gateway/3.9.x/observability/tracing.md): Enable distributed tracing in API7 Gateway with the OpenTelemetry plugin. Configure the OTLP collector, sampling strategy, and per-route spans to visualize request flows and analyze latency bottlenecks. #### overview Discover API7 Gateway, a dynamic, high-performance API gateway for cloud-native environments, built on Apache APISIX. Learn about its architecture, features, and deployment models. - [API7 Gateway](https://docs.api7.ai/api7-gateway/3.9.x/overview.md): Discover API7 Gateway, a dynamic, high-performance API gateway for cloud-native environments, built on Apache APISIX. Learn about its architecture, features, and deployment models. #### reference Technical reference documentation for API7 Gateway, including configuration files, environment variables, CLI tools, and API specifications. - [Reference Overview](https://docs.api7.ai/api7-gateway/3.9.x/reference.md): Technical reference documentation for API7 Gateway, including configuration files, environment variables, CLI tools, and API specifications. - [API Declarative CLI (ADC)](https://docs.api7.ai/api7-gateway/3.9.x/reference/adc.md): Use API Declarative CLI (ADC) to manage API7 Gateway configuration declaratively. - [Alert Variables and Templates](https://docs.api7.ai/api7-gateway/3.9.x/reference/alert-template.md): Customize alert notifications with pre-defined variables in API7 Gateway, enabling dynamic content in alert messages and emails. - [Approval Notification Variables and Templates](https://docs.api7.ai/api7-gateway/3.9.x/reference/approval-variables.md): Use API7 Gateway approval variables to customize API product subscription notification email content and webhook messages with current template values. - [Built-In Variables](https://docs.api7.ai/api7-gateway/3.9.x/reference/built-in-variables.md): Discover the built-in variables available in API7 Gateway, including NGINX and APISIX variables, which can be utilized for route matching, log customization, and plugin configurations. - [Configuration Files](https://docs.api7.ai/api7-gateway/3.9.x/reference/configuration.md): Understand the configuration files used in API7 Gateway, including default and user-defined files for managing settings effectively. - [Environment Variables](https://docs.api7.ai/api7-gateway/3.9.x/reference/environment-variables.md): Explore the use of environment variables in API7 Gateway for configuring consumer credentials, SSL certificates, and plugins. - [API7 Expressions](https://docs.api7.ai/api7-gateway/3.9.x/reference/expressions.md): Understand how to use expressions in API7 Gateway for route matching, request filtering, and conditional logic in configurations. - [Security Hardening Reference](https://docs.api7.ai/api7-gateway/3.9.x/reference/hardening.md): Learn about securing sensitive information in API7 Gateway, including storage, encryption, and communication practices to protect against threats. - [Helm Chart](https://docs.api7.ai/api7-gateway/3.9.x/reference/helm-chart.md): Learn where to find API7 Gateway Helm chart values and how Helm values are rendered into gateway configuration. - [Obtain a Token from the Dashboard](https://docs.api7.ai/api7-gateway/3.9.x/reference/obtain-dashboard-token.md): Create a token in the API7 Dashboard and use it for API7 Gateway Admin API and ADC authentication. - [Permission Policy Actions and Resources](https://docs.api7.ai/api7-gateway/3.9.x/reference/permission-policy-action-and-resource.md): Complete reference for all permission policy actions and ARN-style resources in API7 Gateway — organized by namespace (gateway, iam, portal) for building least-privilege policies. - [Permission Policy Examples](https://docs.api7.ai/api7-gateway/3.9.x/reference/permission-policy-examples.md): Ready-to-adapt permission policy examples for API7 Gateway, organized by access pattern, service and plugin operations, IAM and governance, and portal management. #### release-notes Review the latest updates, features, improvements, and bug fixes for each version of API7 Gateway in release notes. - [API7 Gateway Release Notes](https://docs.api7.ai/api7-gateway/3.9.x/release-notes.md): Review the latest updates, features, improvements, and bug fixes for each version of API7 Gateway in release notes. #### scalability - [Autoscale Data Plane on Kubernetes](https://docs.api7.ai/api7-gateway/3.9.x/scalability/autoscale-on-kubernetes.md): Autoscale API7 Gateway data plane pods on Kubernetes with a Horizontal Pod Autoscaler (HPA) to handle changing traffic automatically. #### security-and-compliance - [Access Control Lists (ACLs)](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/access-control-lists.md): Restrict API access using the consumer-restriction plugin in API7 Gateway. Configure allowlists and denylists keyed by consumer name, consumer group ID, service ID, or route ID for granular security. - [Audit Logs](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/audit-logs.md): Track all administrative changes in the API7 Control Plane with detailed audit logs. Learn how to review, manage, and export audit trails for security and compliance. - [Client mTLS Authentication](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/authenticate/client-mtls.md): Configure mutual TLS (mTLS) on API7 Gateway to authenticate API clients with X.509 certificates before allowing access to your APIs. - [Mutual TLS between Control Plane and Data Plane](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/authenticate/mutual-tls-cp-dp.md): Learn how API7 Gateway uses a robust, PKI-based mutual TLS (mTLS) model to secure all communication between the Control Plane and Data Plane. - [SCIM Provisioning for Dashboard](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/authenticate/scim.md): Configure SCIM provisioning for API7 Gateway Dashboard with supported identity providers. - [SCIM Provisioning with Microsoft Entra ID](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/authenticate/scim-microsoft-entra-id.md): Configure SCIM provisioning with Microsoft Entra ID for API7 Gateway Dashboard. - [SCIM Provisioning with Okta](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/authenticate/scim-okta.md): Configure SCIM provisioning with Okta for API7 Gateway Dashboard. - [Upstream mTLS](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/authenticate/upstream-mtls.md): Configure mutual TLS (mTLS) between API7 Gateway and upstream services so the gateway presents a client certificate to the upstream and optionally validates the upstream's server certificate. - [IP Restrictions for Control Plane](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/ip-restrictions-control-plane.md): Restrict which client IP addresses can reach the API7 Control Plane (Dashboard and Admin API) using the Control Plane's built-in IP allow list, with complementary network-level controls for defence in depth. - [OAuth 2.0 and OIDC](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/oauth-oidc.md): Secure your APIs with modern token-based authentication. Learn how to integrate API7 Gateway with OAuth 2.0 and OpenID Connect (OIDC) providers like Okta and Keycloak. - [Open Source Licenses](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/open-source-licenses.md): API7 Gateway's commitment to open-source software and license compliance. Review the major open-source components used and their respective licenses. - [Security and Compliance Overview](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/overview.md): Learn how API7 Gateway provides comprehensive security and compliance features to protect your APIs, including authentication, authorization, encryption, and auditing capabilities. - [Permission Policies and Boundaries](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/permission-policies-and-boundaries.md): Author permission policies in API7 Gateway using the native JSON document format. Learn the statement structure, allowed actions, ARN-style resources, label-based conditions, and how permission boundaries enforce maximum allowable permissions. - [Role-Based Access Control (RBAC)](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/role-based-access-control.md): Manage user access to the API7 Gateway control plane with Role-Based Access Control. Assign users to roles, attach permission policies, and enforce least-privilege access across gateway groups. - [Secure Credentials Management](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/secure-credentials.md): Protect your sensitive information with API7 Gateway's secure credentials management. Learn how to manage SSL certificates and integrate with external secret managers like HashiCorp Vault. - [SSO for Dashboard](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/sso-dashboard.md): Centralize user access to the API7 Dashboard using Single Sign-On (SSO). Supports OIDC, SAML, LDAP, and CAS protocols with automatic role mapping. - [SSO with LDAP](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/sso-ldap.md): Configure Single Sign-On (SSO) for the API7 Dashboard using LDAP, allowing users to authenticate with their existing directory service credentials. - [SSO with OIDC](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/sso-oidc.md): Configure Single Sign-On (SSO) for the API7 Dashboard using OpenID Connect (OIDC), with provider-specific guidance for Keycloak, Microsoft Entra ID, and Auth0. - [SSO with SAML](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/sso-saml.md): Configure Single Sign-On (SSO) for the API7 Dashboard using SAML 2.0, with provider-specific guidance for Microsoft Entra ID and Okta. - [Trust Center](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/trust-center.md): The API7 Trust Center is your centralized resource for all security and compliance-related information, including certifications, security reports, and best practice guides. - [Verify Image Signatures](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/verify-image-signatures.md): Verify API7 Enterprise container image signatures with Cosign and keyless OIDC-based verification. Protect your supply chain by confirming every image was built and signed by API7.ai. - [Vulnerability Scanning](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/vulnerability-scanning.md): Learn about API7.ai's security testing and vulnerability scanning practices for API7 Gateway, including our CVE reporting and patching processes. - [Web Application Firewall (WAF)](https://docs.api7.ai/api7-gateway/3.9.x/security-and-compliance/web-application-firewall.md): Learn when to use a Web Application Firewall with API7 Gateway, how it fits into your security model, and where to find the current documented integration path. #### troubleshooting Diagnose and resolve common issues with API7 Gateway deployments, including connectivity problems, configuration errors, and performance degradation. - [Troubleshoot API7 Gateway](https://docs.api7.ai/api7-gateway/3.9.x/troubleshooting.md): Diagnose and resolve common issues with API7 Gateway deployments, including connectivity problems, configuration errors, and performance degradation. #### upgrade-guides - [Backup and Restoration](https://docs.api7.ai/api7-gateway/3.9.x/upgrade-guides/backup-and-restore.md): Learn how to back up and restore your API7 Gateway data with step-by-step instructions for database backup, declarative configuration backup, and data restoration procedures. - [API Gateway Cluster Migration](https://docs.api7.ai/api7-gateway/3.9.x/upgrade-guides/cluster-migration.md): A step-by-step guide for migrating your API7 Gateway deployment to a new cluster with zero downtime. - [Dual-Cluster Upgrade](https://docs.api7.ai/api7-gateway/3.9.x/upgrade-guides/dual-cluster.md): Step-by-step guide for performing a dual-cluster upgrade of API7 Gateway, covering new cluster deployment, traffic shifting, and rollback procedures. - [In-Place Upgrade](https://docs.api7.ai/api7-gateway/3.9.x/upgrade-guides/in-place.md): Step-by-step guide for performing an in-place upgrade of API7 Gateway Control Plane, covering database reuse, configuration updates, and verification procedures. - [Rolling Upgrade](https://docs.api7.ai/api7-gateway/3.9.x/upgrade-guides/rolling-upgrade.md): Step-by-step guide for performing a rolling upgrade of API7 Gateway Data Plane nodes, ensuring zero downtime while maintaining API request processing and service continuity. - [Upgrade API7 Gateway](https://docs.api7.ai/api7-gateway/3.9.x/upgrade-guides/upgrade.md): Comprehensive guide for upgrading API7 Gateway, covering in-place CP upgrade, rolling DP upgrade, data backup, and stability considerations. #### version-support-policy API7 Enterprise version support lifecycle, LTS policy, and upgrade guidance. Learn how long each release is supported and which versions are currently designated as LTS. - [Version Support Policy](https://docs.api7.ai/api7-gateway/3.9.x/version-support-policy.md): API7 Enterprise version support lifecycle, LTS policy, and upgrade guidance. Learn how long each release is supported and which versions are currently designated as LTS. ### ai-agent-skills Manage API7 Gateway with AI coding agents like Claude Code and Cursor. Agent skills that configure your API7 Enterprise Edition gateway from natural language. - [AI Agent Skills for API7 Gateway](https://docs.api7.ai/api7-gateway/ai-agent-skills.md): Manage API7 Gateway with AI coding agents like Claude Code and Cursor. Agent skills that configure your API7 Enterprise Edition gateway from natural language. #### a7-persona-developer Persona skill for API developers building and testing APIs on API7 Enterprise Edition (API7 EE) using the a7 CLI. Provides decision frameworks for servi… - [a7-persona-developer](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-persona-developer.md): Persona skill for API developers building and testing APIs on API7 Enterprise Edition (API7 EE) using the a7 CLI. Provides decision frameworks for servi… #### a7-persona-operator Persona skill for platform operators and DevOps engineers managing API7 Enterprise Edition (API7 EE) instances using the a7 CLI. Provides decision frame… - [a7-persona-operator](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-persona-operator.md): Persona skill for platform operators and DevOps engineers managing API7 Enterprise Edition (API7 EE) instances using the a7 CLI. Provides decision frame… #### a7-plugin-ai-content-moderation Skill for configuring API7 Enterprise Edition AI content moderation plugins via the a7 CLI. Covers both ai-aws-content-moderation (AWS Comprehend, reque… - [a7-plugin-ai-content-moderation](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-ai-content-moderation.md): Skill for configuring API7 Enterprise Edition AI content moderation plugins via the a7 CLI. Covers both ai-aws-content-moderation (AWS Comprehend, reque… #### a7-plugin-ai-prompt-decorator Skill for configuring the API7 Enterprise Edition ai-prompt-decorator plugin via the a7 CLI. Covers prepending and appending system/user/assistant messa… - [a7-plugin-ai-prompt-decorator](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-ai-prompt-decorator.md): Skill for configuring the API7 Enterprise Edition ai-prompt-decorator plugin via the a7 CLI. Covers prepending and appending system/user/assistant messa… #### a7-plugin-ai-prompt-template Skill for configuring the API7 Enterprise Edition ai-prompt-template plugin via the a7 CLI. Covers defining reusable prompt templates with variable plac… - [a7-plugin-ai-prompt-template](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-ai-prompt-template.md): Skill for configuring the API7 Enterprise Edition ai-prompt-template plugin via the a7 CLI. Covers defining reusable prompt templates with variable plac… #### a7-plugin-ai-proxy Skill for configuring the API7 Enterprise Edition ai-proxy plugin via the a7 CLI. Covers proxying requests to LLM providers (OpenAI, Azure OpenAI, DeepS… - [a7-plugin-ai-proxy](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-ai-proxy.md): Skill for configuring the API7 Enterprise Edition ai-proxy plugin via the a7 CLI. Covers proxying requests to LLM providers (OpenAI, Azure OpenAI, DeepS… #### a7-plugin-basic-auth Skill for configuring the API7 Enterprise Edition (API7 EE) basic-auth plugin via the a7 CLI. Covers HTTP Basic Authentication setup on routes, consumer… - [a7-plugin-basic-auth](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-basic-auth.md): Skill for configuring the API7 Enterprise Edition (API7 EE) basic-auth plugin via the a7 CLI. Covers HTTP Basic Authentication setup on routes, consumer… #### a7-plugin-consumer-restriction Skill for configuring the API7 Enterprise Edition consumer-restriction plugin via the a7 CLI. Covers restricting access by consumer name, service ID, or… - [a7-plugin-consumer-restriction](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-consumer-restriction.md): Skill for configuring the API7 Enterprise Edition consumer-restriction plugin via the a7 CLI. Covers restricting access by consumer name, service ID, or… #### a7-plugin-cors Skill for configuring the API7 Enterprise Edition (API7 EE) cors plugin via the a7 CLI. Covers Cross-Origin Resource Sharing setup on routes, allow_orig… - [a7-plugin-cors](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-cors.md): Skill for configuring the API7 Enterprise Edition (API7 EE) cors plugin via the a7 CLI. Covers Cross-Origin Resource Sharing setup on routes, allow_orig… #### a7-plugin-datadog Skill for configuring the API7 Enterprise Edition datadog plugin via the a7 CLI. Covers pushing custom metrics to Datadog via DogStatsD, metric tags, ba… - [a7-plugin-datadog](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-datadog.md): Skill for configuring the API7 Enterprise Edition datadog plugin via the a7 CLI. Covers pushing custom metrics to Datadog via DogStatsD, metric tags, ba… #### a7-plugin-ext-plugin Skill for configuring the API7 Enterprise Edition external plugin system (ext-plugin-pre-req, ext-plugin-post-req, ext-plugin-post-resp) via the a7 CLI.… - [a7-plugin-ext-plugin](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-ext-plugin.md): Skill for configuring the API7 Enterprise Edition external plugin system (ext-plugin-pre-req, ext-plugin-post-req, ext-plugin-post-resp) via the a7 CLI.… #### a7-plugin-fault-injection Skill for configuring the API7 Enterprise Edition fault-injection plugin via the a7 CLI. Covers injecting delays and HTTP aborts for chaos engineering,… - [a7-plugin-fault-injection](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-fault-injection.md): Skill for configuring the API7 Enterprise Edition fault-injection plugin via the a7 CLI. Covers injecting delays and HTTP aborts for chaos engineering,… #### a7-plugin-grpc-transcode Skill for configuring the API7 Enterprise Edition (API7 EE) grpc-transcode plugin via the a7 CLI. Covers converting RESTful HTTP requests to gRPC, proto… - [a7-plugin-grpc-transcode](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-grpc-transcode.md): Skill for configuring the API7 Enterprise Edition (API7 EE) grpc-transcode plugin via the a7 CLI. Covers converting RESTful HTTP requests to gRPC, proto… #### a7-plugin-hmac-auth Skill for configuring the API7 Enterprise Edition (API7 EE) hmac-auth plugin via the a7 CLI. Covers HMAC signature authentication, consumer credential b… - [a7-plugin-hmac-auth](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-hmac-auth.md): Skill for configuring the API7 Enterprise Edition (API7 EE) hmac-auth plugin via the a7 CLI. Covers HMAC signature authentication, consumer credential b… #### a7-plugin-http-logger Skill for configuring the API7 Enterprise Edition (API7 EE) http-logger plugin via the a7 CLI. Covers pushing access logs to HTTP/HTTPS endpoints in bat… - [a7-plugin-http-logger](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-http-logger.md): Skill for configuring the API7 Enterprise Edition (API7 EE) http-logger plugin via the a7 CLI. Covers pushing access logs to HTTP/HTTPS endpoints in bat… #### a7-plugin-ip-restriction Skill for configuring the API7 Enterprise Edition (API7 EE) ip-restriction plugin via the a7 CLI. Covers IP whitelist/blacklist setup on routes, CIDR ra… - [a7-plugin-ip-restriction](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-ip-restriction.md): Skill for configuring the API7 Enterprise Edition (API7 EE) ip-restriction plugin via the a7 CLI. Covers IP whitelist/blacklist setup on routes, CIDR ra… #### a7-plugin-jwt-auth Skill for configuring the API7 Enterprise Edition (API7 EE) jwt-auth plugin via the a7 CLI. Covers JWT token authentication, HS256/RS256 algorithm selec… - [a7-plugin-jwt-auth](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-jwt-auth.md): Skill for configuring the API7 Enterprise Edition (API7 EE) jwt-auth plugin via the a7 CLI. Covers JWT token authentication, HS256/RS256 algorithm selec… #### a7-plugin-kafka-logger Skill for configuring the API7 Enterprise Edition (API7 EE) kafka-logger plugin via the a7 CLI. Covers pushing access logs to Apache Kafka topics, broke… - [a7-plugin-kafka-logger](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-kafka-logger.md): Skill for configuring the API7 Enterprise Edition (API7 EE) kafka-logger plugin via the a7 CLI. Covers pushing access logs to Apache Kafka topics, broke… #### a7-plugin-key-auth Skill for configuring the API7 Enterprise Edition (API7 EE) key-auth plugin via the a7 CLI. Covers API key authentication setup on routes, consumer cred… - [a7-plugin-key-auth](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-key-auth.md): Skill for configuring the API7 Enterprise Edition (API7 EE) key-auth plugin via the a7 CLI. Covers API key authentication setup on routes, consumer cred… #### a7-plugin-limit-count Skill for configuring the API7 Enterprise Edition (API7 EE) limit-count plugin via the a7 CLI. Covers fixed-window rate limiting, count/time_window conf… - [a7-plugin-limit-count](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-limit-count.md): Skill for configuring the API7 Enterprise Edition (API7 EE) limit-count plugin via the a7 CLI. Covers fixed-window rate limiting, count/time_window conf… #### a7-plugin-limit-req Skill for configuring the API7 Enterprise Edition (API7 EE) limit-req plugin via the a7 CLI. Covers leaky-bucket rate limiting, rate/burst configuration… - [a7-plugin-limit-req](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-limit-req.md): Skill for configuring the API7 Enterprise Edition (API7 EE) limit-req plugin via the a7 CLI. Covers leaky-bucket rate limiting, rate/burst configuration… #### a7-plugin-openid-connect Skill for configuring the API7 Enterprise Edition (API7 EE) openid-connect plugin via the a7 CLI. Covers OIDC authorization code flow, bearer token vali… - [a7-plugin-openid-connect](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-openid-connect.md): Skill for configuring the API7 Enterprise Edition (API7 EE) openid-connect plugin via the a7 CLI. Covers OIDC authorization code flow, bearer token vali… #### a7-plugin-prometheus Skill for configuring the API7 Enterprise Edition (API7 EE) prometheus plugin via the a7 CLI. Covers enabling Prometheus metrics export on routes and gl… - [a7-plugin-prometheus](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-prometheus.md): Skill for configuring the API7 Enterprise Edition (API7 EE) prometheus plugin via the a7 CLI. Covers enabling Prometheus metrics export on routes and gl… #### a7-plugin-proxy-rewrite Skill for configuring the API7 Enterprise Edition (API7 EE) proxy-rewrite plugin via the a7 CLI. Covers rewriting request URI, host, method, headers, an… - [a7-plugin-proxy-rewrite](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-proxy-rewrite.md): Skill for configuring the API7 Enterprise Edition (API7 EE) proxy-rewrite plugin via the a7 CLI. Covers rewriting request URI, host, method, headers, an… #### a7-plugin-redirect Skill for configuring the API7 Enterprise Edition (API7 EE) redirect plugin via the a7 CLI. Covers URI redirects, HTTP-to-HTTPS redirection, regex-based… - [a7-plugin-redirect](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-redirect.md): Skill for configuring the API7 Enterprise Edition (API7 EE) redirect plugin via the a7 CLI. Covers URI redirects, HTTP-to-HTTPS redirection, regex-based… #### a7-plugin-response-rewrite Skill for configuring the API7 Enterprise Edition (API7 EE) response-rewrite plugin via the a7 CLI. Covers rewriting response status codes, headers, and… - [a7-plugin-response-rewrite](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-response-rewrite.md): Skill for configuring the API7 Enterprise Edition (API7 EE) response-rewrite plugin via the a7 CLI. Covers rewriting response status codes, headers, and… #### a7-plugin-serverless Skill for configuring the API7 Enterprise Edition serverless-pre-function and serverless-post-function plugins via the a7 CLI. Covers inline Lua functio… - [a7-plugin-serverless](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-serverless.md): Skill for configuring the API7 Enterprise Edition serverless-pre-function and serverless-post-function plugins via the a7 CLI. Covers inline Lua functio… #### a7-plugin-skywalking Skill for configuring the API7 Enterprise Edition (API7 EE) skywalking plugin via the a7 CLI. Covers distributed tracing with Apache SkyWalking OAP, sam… - [a7-plugin-skywalking](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-skywalking.md): Skill for configuring the API7 Enterprise Edition (API7 EE) skywalking plugin via the a7 CLI. Covers distributed tracing with Apache SkyWalking OAP, sam… #### a7-plugin-traffic-split Skill for configuring the API7 Enterprise Edition (API7 EE) traffic-split plugin via the a7 CLI. Covers weighted traffic splitting between upstreams wit… - [a7-plugin-traffic-split](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-traffic-split.md): Skill for configuring the API7 Enterprise Edition (API7 EE) traffic-split plugin via the a7 CLI. Covers weighted traffic splitting between upstreams wit… #### a7-plugin-wolf-rbac Skill for configuring the API7 Enterprise Edition (API7 EE) wolf-rbac plugin via the a7 CLI. Covers integration with the Wolf RBAC server for role-based… - [a7-plugin-wolf-rbac](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-wolf-rbac.md): Skill for configuring the API7 Enterprise Edition (API7 EE) wolf-rbac plugin via the a7 CLI. Covers integration with the Wolf RBAC server for role-based… #### a7-plugin-zipkin Skill for configuring the API7 Enterprise Edition (API7 EE) zipkin plugin via the a7 CLI. Covers distributed tracing with Zipkin, Jaeger, or any Zipkin-… - [a7-plugin-zipkin](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-plugin-zipkin.md): Skill for configuring the API7 Enterprise Edition (API7 EE) zipkin plugin via the a7 CLI. Covers distributed tracing with Zipkin, Jaeger, or any Zipkin-… #### a7-recipe-api-versioning Recipe skill for implementing API versioning strategies using API7 Enterprise Edition (API7 EE) and the a7 CLI. Covers URI path versioning, header-based… - [a7-recipe-api-versioning](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-recipe-api-versioning.md): Recipe skill for implementing API versioning strategies using API7 Enterprise Edition (API7 EE) and the a7 CLI. Covers URI path versioning, header-based… #### a7-recipe-blue-green Recipe skill for implementing blue-green deployments using the a7 CLI in API7 Enterprise Edition. Covers creating two service-backed environments, switc… - [a7-recipe-blue-green](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-recipe-blue-green.md): Recipe skill for implementing blue-green deployments using the a7 CLI in API7 Enterprise Edition. Covers creating two service-backed environments, switc… #### a7-recipe-canary Recipe skill for implementing canary releases using the a7 CLI in API7 Enterprise Edition. Covers gradual traffic shifting with the traffic-split plugin… - [a7-recipe-canary](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-recipe-canary.md): Recipe skill for implementing canary releases using the a7 CLI in API7 Enterprise Edition. Covers gradual traffic shifting with the traffic-split plugin… #### a7-recipe-circuit-breaker Recipe skill for implementing circuit breaker patterns using the a7 CLI in API7 Enterprise Edition. Covers the api-breaker plugin, unhealthy thresholds,… - [a7-recipe-circuit-breaker](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-recipe-circuit-breaker.md): Recipe skill for implementing circuit breaker patterns using the a7 CLI in API7 Enterprise Edition. Covers the api-breaker plugin, unhealthy thresholds,… #### a7-recipe-graphql-proxy Recipe skill for implementing GraphQL proxying patterns using API7 Enterprise Edition (API7 EE) and the a7 CLI. Covers operation-based routing, per-oper… - [a7-recipe-graphql-proxy](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-recipe-graphql-proxy.md): Recipe skill for implementing GraphQL proxying patterns using API7 Enterprise Edition (API7 EE) and the a7 CLI. Covers operation-based routing, per-oper… #### a7-recipe-health-check Recipe skill for configuring backend health checks using the a7 CLI in API7 Enterprise Edition. Covers active health checks, passive health checks, comb… - [a7-recipe-health-check](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-recipe-health-check.md): Recipe skill for configuring backend health checks using the a7 CLI in API7 Enterprise Edition. Covers active health checks, passive health checks, comb… #### a7-recipe-mtls Recipe skill for configuring mutual TLS (mTLS) using the a7 CLI in API7 Enterprise Edition. Covers SSL certificate management, upstream mTLS to backend… - [a7-recipe-mtls](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-recipe-mtls.md): Recipe skill for configuring mutual TLS (mTLS) using the a7 CLI in API7 Enterprise Edition. Covers SSL certificate management, upstream mTLS to backend… #### a7-recipe-multi-tenant Recipe skill for implementing multi-tenant patterns using API7 Enterprise Edition (API7 EE) and the a7 CLI. Covers gateway-group isolation, consumer pol… - [a7-recipe-multi-tenant](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-recipe-multi-tenant.md): Recipe skill for implementing multi-tenant patterns using API7 Enterprise Edition (API7 EE) and the a7 CLI. Covers gateway-group isolation, consumer pol… #### a7-shared Core skill for working with the a7 CLI — the command-line tool for API7 Enterprise Edition. Provides project conventions, command patterns, dual-API arc… - [a7 Shared Skill](https://docs.api7.ai/api7-gateway/ai-agent-skills/a7-shared.md): Core skill for working with the a7 CLI — the command-line tool for API7 Enterprise Edition. Provides project conventions, command patterns, dual-API arc… ### ai-gateway #### get-started Set up API7 AI Gateway and proxy your first request to OpenAI in under 5 minutes. Step-by-step guide with code examples. - [Proxy Your First LLM Request in 5 Minutes](https://docs.api7.ai/api7-gateway/ai-gateway/get-started.md): Set up API7 AI Gateway and proxy your first request to OpenAI in under 5 minutes. Step-by-step guide with code examples. #### llm-providers - [Connect to Anthropic Claude](https://docs.api7.ai/api7-gateway/ai-gateway/llm-providers/anthropic.md): Route Anthropic Claude API traffic through API7 Gateway for centralized security, rate limiting, and observability. - [Integrate Azure OpenAI Service](https://docs.api7.ai/api7-gateway/ai-gateway/llm-providers/azure-openai.md): Manage Azure OpenAI deployments through API7 Gateway. Handle resource names, API versions, and auth centrally. - [Route Traffic to DeepSeek Models](https://docs.api7.ai/api7-gateway/ai-gateway/llm-providers/deepseek.md): Proxy DeepSeek API requests through API7 Gateway. Manage authentication, enable failover, and monitor usage centrally. - [Integrate Google Gemini](https://docs.api7.ai/api7-gateway/ai-gateway/llm-providers/google-gemini.md): Proxy Google Gemini API requests through API7 Gateway. Manage API keys and monitor AI traffic centrally. - [Route Traffic to OpenAI](https://docs.api7.ai/api7-gateway/ai-gateway/llm-providers/openai.md): Proxy and secure OpenAI API requests through API7 Gateway. Centralize authentication, enable failover, and monitor usage. - [Connect Any OpenAI-Compatible LLM](https://docs.api7.ai/api7-gateway/ai-gateway/llm-providers/openai-compatible.md): Proxy any OpenAI-compatible API through API7 Gateway. Connect self-hosted models, custom endpoints, or niche providers. - [Access Hundreds of LLMs via OpenRouter](https://docs.api7.ai/api7-gateway/ai-gateway/llm-providers/openrouter.md): Use OpenRouter with API7 AI Gateway to access 200+ LLMs through one API while maintaining enterprise security controls. - [Route Enterprise AI Traffic to Vertex AI](https://docs.api7.ai/api7-gateway/ai-gateway/llm-providers/vertex-ai.md): Securely proxy Google Cloud Vertex AI requests through API7 Gateway with service account auth and regional routing. #### overview Centralize LLM access with API7 AI Gateway. Route traffic to multiple providers, enforce guardrails, and control costs from one platform. - [Manage and Secure AI Traffic](https://docs.api7.ai/api7-gateway/ai-gateway/overview.md): Centralize LLM access with API7 AI Gateway. Route traffic to multiple providers, enforce guardrails, and control costs from one platform. #### use-cases - [Monitor AI Traffic and Track LLM Costs](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/ai-observability-and-cost-tracking.md): Gain visibility into LLM usage, token consumption, latency, and costs with API7 AI Gateway's observability features. - [Transform API Requests with AI-Powered Rewriting](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/ai-request-transformation.md): Use LLMs to intelligently transform, enrich, or restructure API requests and responses at the gateway layer. - [Enforce AI Guardrails and Protect PII](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/content-safety-and-guardrails.md): Block prompt injection, detect toxicity, and redact PII before requests reach LLMs using API7 AI Gateway guardrails. - [Expose REST APIs as MCP Tools for AI Agents](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/expose-apis-as-mcp-tools.md): Convert existing OpenAPI services into MCP-compatible tools so AI agents can discover and invoke your APIs automatically. - [Manage API7 Enterprise from an AI Client with API7-MCP](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/manage-api7-with-mcp.md): Deploy the API7-MCP server so an AI client such as Cursor, Claude Desktop, or Cline can read API7 Enterprise resources, check Prometheus metrics, manage RBAC, and send test traffic through the gateway. - [Set Up Multi-LLM Routing and Automatic Fallback](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/multi-llm-routing-and-fallback.md): Route AI traffic across multiple LLM providers with weighted load balancing, automatic failover, and health checks. - [Implement Prompt Templates and Decorators](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/prompt-engineering-and-templating.md): Standardize LLM interactions with reusable prompt templates and automatic system prompt injection using API7 AI Gateway. - [Convert Anthropic Messages to OpenAI Chat Completions](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/protocol-conversion.md): Use API7 AI Gateway to transparently convert Anthropic Messages API requests to the OpenAI Chat Completions API format, enabling teams to use the Anthropic SDK with any OpenAI-compatible backend. - [Implement RAG at the Gateway Layer](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/retrieval-augmented-generation.md): Enhance LLM responses with relevant context using Retrieval-Augmented Generation (RAG) built into API7 AI Gateway. - [Control AI Costs with Token-Based Rate Limiting](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/token-rate-limiting-and-quota-management.md): Implement token-based rate limits to prevent LLM abuse and control AI costs per route and model instance. ### configure-and-manage #### benchmark-on-aws-eks Reproduce the published API7 Gateway performance benchmark on AWS EKS. Walkthrough covers EKS cluster setup, three isolated node groups, Helm install, NGINX upstream and wrk2 deployment, and running the full scenario suite. - [Run Benchmarks on AWS EKS](https://docs.api7.ai/api7-gateway/configure-and-manage/benchmark-on-aws-eks.md): Reproduce the published API7 Gateway performance benchmark on AWS EKS. Walkthrough covers EKS cluster setup, three isolated node groups, Helm install, NGINX upstream and wrk2 deployment, and running the full scenario suite. #### configure-control-plane Detailed configuration reference for the API7 Gateway Control Plane, covering the Dashboard and DP Manager configuration files. - [Configuration Reference for API7 Gateway Control Plane](https://docs.api7.ai/api7-gateway/configure-and-manage/configure-control-plane.md): Detailed configuration reference for the API7 Gateway Control Plane, covering the Dashboard and DP Manager configuration files. #### configure-data-plane Detailed configuration reference for API7 Gateway Data Plane, based on the config-default.yaml structure. - [Configuration Reference for API7 Gateway Data Plane](https://docs.api7.ai/api7-gateway/configure-and-manage/configure-data-plane.md): Detailed configuration reference for API7 Gateway Data Plane, based on the config-default.yaml structure. #### data-plane-resilience Configure fallback storage for API7 Gateway data plane nodes so they can restart and continue operating during extended control plane outages. - [Data Plane Resilience](https://docs.api7.ai/api7-gateway/configure-and-manage/data-plane-resilience.md): Configure fallback storage for API7 Gateway data plane nodes so they can restart and continue operating during extended control plane outages. #### deployment-scenarios - [Multiple Availability Zones Deployment of API7 Gateway](https://docs.api7.ai/api7-gateway/configure-and-manage/deployment-scenarios/multi-az-deployment.md): Configuration and architecture for deploying API7 Gateway across multiple availability zones for high availability and fault tolerance. - [Multi-Region Deployment Patterns](https://docs.api7.ai/api7-gateway/configure-and-manage/deployment-scenarios/multi-region-deployment.md): Architecture and considerations for deploying API7 Gateway across multiple geographic regions for global reach and disaster recovery. #### high-availability-data-plane Design a highly available API7 Gateway data plane with multiple nodes, health checks, and load balancer failover. - [Data Plane High Availability](https://docs.api7.ai/api7-gateway/configure-and-manage/high-availability-data-plane.md): Design a highly available API7 Gateway data plane with multiple nodes, health checks, and load balancer failover. #### labels Organize and filter API7 Gateway resources at scale with labels — key-value metadata attached to gateway groups, services, routes, consumers, and other entities for team, environment, and application-level segmentation. - [Labels](https://docs.api7.ai/api7-gateway/configure-and-manage/labels.md): Organize and filter API7 Gateway resources at scale with labels — key-value metadata attached to gateway groups, services, routes, consumers, and other entities for team, environment, and application-level segmentation. #### license-management Manage your API7 Gateway license, understand the production and non-production core quotas and license states, and configure license file paths for automated deployment. - [License Management](https://docs.api7.ai/api7-gateway/configure-and-manage/license-management.md): Manage your API7 Gateway license, understand the production and non-production core quotas and license states, and configure license file paths for automated deployment. #### performance-benchmark Published performance benchmark results for API7 Gateway (AWS EKS and single-host baselines), plus methodology and optimization guidance for running your own benchmarks accurately. - [Performance Benchmark](https://docs.api7.ai/api7-gateway/configure-and-manage/performance-benchmark.md): Published performance benchmark results for API7 Gateway (AWS EKS and single-host baselines), plus methodology and optimization guidance for running your own benchmarks accurately. #### production-best-practices Operational best practices for managing API7 Gateway in production, including GitOps, change management, and disaster recovery. - [Production Best Practices](https://docs.api7.ai/api7-gateway/configure-and-manage/production-best-practices.md): Operational best practices for managing API7 Gateway in production, including GitOps, change management, and disaster recovery. #### run-in-production Pre-production checklist and deployment guide for API7 Gateway to ensure a stable and secure production environment. - [Running in Production](https://docs.api7.ai/api7-gateway/configure-and-manage/run-in-production.md): Pre-production checklist and deployment guide for API7 Gateway to ensure a stable and secure production environment. #### scale-data-plane Scale API7 Gateway data plane nodes horizontally to increase throughput and prepare for high-availability deployments. - [Scale Data Plane](https://docs.api7.ai/api7-gateway/configure-and-manage/scale-data-plane.md): Scale API7 Gateway data plane nodes horizontally to increase throughput and prepare for high-availability deployments. #### shared-dict-sizing Size API7 Gateway shared memory zones for metrics, service discovery, the Developer Portal, and tracing so they do not overflow at your deployment's scale. - [Shared Memory Sizing](https://docs.api7.ai/api7-gateway/configure-and-manage/shared-dict-sizing.md): Size API7 Gateway shared memory zones for metrics, service discovery, the Developer Portal, and tracing so they do not overflow at your deployment's scale. #### telemetry-opt-out Configure telemetry data transfer between the data plane and control plane, including compression levels and how to disable telemetry. - [Optimize Telemetry Data Transfer](https://docs.api7.ai/api7-gateway/configure-and-manage/telemetry-opt-out.md): Configure telemetry data transfer between the data plane and control plane, including compression levels and how to disable telemetry. #### user-management Manage users, roles, and permission policies in API7 Enterprise. - [User Management](https://docs.api7.ai/api7-gateway/configure-and-manage/user-management.md): Manage users, roles, and permission policies in API7 Enterprise. ### developer-portal #### deploy - [Configure the Developer Portal](https://docs.api7.ai/api7-gateway/developer-portal/deploy/configure-portal.md): Configure the Developer Portal settings, including public access, portal tokens, built-in authentication, and SCIM provisioning. - [Customize the Developer Portal](https://docs.api7.ai/api7-gateway/developer-portal/deploy/customize-portal.md): Customize the Developer Portal branding, theme, authentication providers, and functionality using the API7 Developer Portal Boilerplate. - [Deploy the Developer Portal](https://docs.api7.ai/api7-gateway/developer-portal/deploy/deploy-portal.md): Deploy the API7 Developer Portal with its two official images — the Portal API backend and the customer-facing frontend — on Docker Compose or Kubernetes, and connect them to your API7 control plane. #### get-started Start a local API7 Developer Portal from the packaged Docker Compose deployment, register a developer, and verify access to the API Hub. - [Start a Local Developer Portal](https://docs.api7.ai/api7-gateway/developer-portal/get-started.md): Start a local API7 Developer Portal from the packaged Docker Compose deployment, register a developer, and verify access to the API Hub. #### guides - [Browse APIs](https://docs.api7.ai/api7-gateway/developer-portal/guides/browse-apis.md): Discover and explore available API products in the Developer Portal's API Hub. - [Create an Application](https://docs.api7.ai/api7-gateway/developer-portal/guides/create-application.md): Create an application in the Developer Portal to group your API subscriptions and credentials. - [Manage Credentials](https://docs.api7.ai/api7-gateway/developer-portal/guides/manage-credentials.md): Create, view, regenerate, and delete credentials in the Developer Portal for authenticating API requests. - [Manage Your Organization](https://docs.api7.ai/api7-gateway/developer-portal/guides/manage-organization.md): Manage your organization in the Developer Portal, including inviting members, assigning roles, and switching between organizations. - [Register and Log In](https://docs.api7.ai/api7-gateway/developer-portal/guides/register-and-login.md): Create a developer account and log in to the Developer Portal using email/password, SSO, or an organization invitation. - [Subscribe to an API](https://docs.api7.ai/api7-gateway/developer-portal/guides/subscribe-to-api.md): Subscribe your application to an API product to gain access to consume APIs through the Developer Portal. - [Try an API](https://docs.api7.ai/api7-gateway/developer-portal/guides/try-api.md): Test API endpoints directly from the Developer Portal using the built-in Try It Out feature. #### key-concepts - [API Products](https://docs.api7.ai/api7-gateway/developer-portal/key-concepts/api-products.md): Understand API products in the Developer Portal, including product types, visibility settings, authentication options, and the publishing lifecycle. - [Applications](https://docs.api7.ai/api7-gateway/developer-portal/key-concepts/applications.md): Understand applications in the Developer Portal, which group subscriptions and credentials for a specific project or use case. - [Credentials](https://docs.api7.ai/api7-gateway/developer-portal/key-concepts/credentials.md): Understand credentials in the Developer Portal, including supported authentication types (key auth, basic auth, OAuth/DCR) and credential lifecycle management. - [Developers](https://docs.api7.ai/api7-gateway/developer-portal/key-concepts/developers.md): Understand developers in the Developer Portal, including registration methods, account states, and the difference between developers and consumers. - [Subscriptions](https://docs.api7.ai/api7-gateway/developer-portal/key-concepts/subscriptions.md): Understand subscriptions in the Developer Portal, including the approval workflow, status transitions, and auto-approval configuration. #### manage - [Configure Dynamic Client Registration (DCR)](https://docs.api7.ai/api7-gateway/developer-portal/manage/configure-dcr.md): Configure DCR providers to enable developers to register OAuth 2.0 clients through the Developer Portal. - [Configure SCIM Provisioning for a Custom Developer Portal with Okta](https://docs.api7.ai/api7-gateway/developer-portal/manage/configure-scim.md): Configure SCIM provisioning with Okta for a custom Developer Portal based on the API7 Developer Portal Boilerplate. - [Configure SSO for the Developer Portal](https://docs.api7.ai/api7-gateway/developer-portal/manage/configure-sso.md): Configure Single Sign-On (SSO) for the Developer Portal using OIDC, SAML, LDAP, or CAS identity providers. - [Manage API Products](https://docs.api7.ai/api7-gateway/developer-portal/manage/manage-api-products.md): Create, configure, publish, and manage API products in the Provider Portal for developer consumption through the Developer Portal. - [Manage Applications](https://docs.api7.ai/api7-gateway/developer-portal/manage/manage-applications.md): Manage developer applications in the Developer Portal, including lifecycle, structure, and how to call the Developer Portal backend programmatically. - [Manage Developers](https://docs.api7.ai/api7-gateway/developer-portal/manage/manage-developers.md): Manage developer accounts using the Provider Portal Admin API and the standalone Developer Portal backend, including listing, creating, approving registrations, and deleting developers. - [Manage Subscriptions](https://docs.api7.ai/api7-gateway/developer-portal/manage/manage-subscriptions.md): Manage API product subscriptions in the Provider Portal, including approving, rejecting, and cancelling subscription requests. #### overview Learn about the API7 Developer Portal, a platform for API providers to publish API products and for developers to discover, subscribe to, and consume APIs. - [Developer Portal Overview](https://docs.api7.ai/api7-gateway/developer-portal/overview.md): Learn about the API7 Developer Portal, a platform for API providers to publish API products and for developers to discover, subscribe to, and consume APIs. ### enterprise-features #### alerts-and-contact-points Explore the concept of alerts and contact points in API7 Gateway, which monitor exceptions and send timely notifications. - [Alerts and Contact Points](https://docs.api7.ai/api7-gateway/enterprise-features/alerts-and-contact-points.md): Explore the concept of alerts and contact points in API7 Gateway, which monitor exceptions and send timely notifications. #### anonymous-consumers Explore the concept of anonymous consumers in API7 Gateway, allowing non-authenticated access to APIs while maintaining security. - [Anonymous Consumers](https://docs.api7.ai/api7-gateway/enterprise-features/anonymous-consumers.md): Explore the concept of anonymous consumers in API7 Gateway, allowing non-authenticated access to APIs while maintaining security. #### api-portal Explore the concept of the API portal in API7 Gateway, providing a centralized space for developers to access and manage APIs. - [API Portal](https://docs.api7.ai/api7-gateway/enterprise-features/api-portal.md): Explore the concept of the API portal in API7 Gateway, providing a centralized space for developers to access and manage APIs. #### audit-logging Explore audit logging in API7 Gateway, including how user actions and configuration changes are recorded. - [Audit Logging](https://docs.api7.ai/api7-gateway/enterprise-features/audit-logging.md): Explore audit logging in API7 Gateway, including how user actions and configuration changes are recorded. #### compliance Explore the concept of compliance in API7 Gateway, helping organizations meet regulatory requirements and maintain security standards. - [Compliance](https://docs.api7.ai/api7-gateway/enterprise-features/compliance.md): Explore the concept of compliance in API7 Gateway, helping organizations meet regulatory requirements and maintain security standards. #### credentials Explore the concept of credentials in API7 Gateway, authenticating users and ensuring secure access while facilitating management and rotation. - [Credentials](https://docs.api7.ai/api7-gateway/enterprise-features/credentials.md): Explore the concept of credentials in API7 Gateway, authenticating users and ensuring secure access while facilitating management and rotation. #### custom-plugins Explore the concept of custom plugins in API7 Gateway, enabling tailored extensions to meet specific business needs. - [Custom Plugins](https://docs.api7.ai/api7-gateway/enterprise-features/custom-plugins.md): Explore the concept of custom plugins in API7 Gateway, enabling tailored extensions to meet specific business needs. #### dashboard-sso Explore the concept of Single Sign-On (SSO) in API7 Gateway, allowing users to authenticate with existing credentials for easy access. - [Dashboard SSO Options](https://docs.api7.ai/api7-gateway/enterprise-features/dashboard-sso.md): Explore the concept of Single Sign-On (SSO) in API7 Gateway, allowing users to authenticate with existing credentials for easy access. #### gateway-groups Explore the concept of gateway groups in API7 Gateway, which manage multiple API gateway instances with shared configurations. - [Gateway Groups](https://docs.api7.ai/api7-gateway/enterprise-features/gateway-groups.md): Explore the concept of gateway groups in API7 Gateway, which manage multiple API gateway instances with shared configurations. #### high-availability Explore the concept of high availability in API7 Gateway, ensuring continuous service delivery for mission-critical applications. - [High Availability](https://docs.api7.ai/api7-gateway/enterprise-features/high-availability.md): Explore the concept of high availability in API7 Gateway, ensuring continuous service delivery for mission-critical applications. #### organization-and-rbac Explore the concept of organization management and RBAC in API7 Gateway, enabling fine-grained permission management. - [Organization and RBAC](https://docs.api7.ai/api7-gateway/enterprise-features/organization-and-rbac.md): Explore the concept of organization management and RBAC in API7 Gateway, enabling fine-grained permission management. #### overview Discover the enterprise-grade features that distinguish API7 Gateway from open-source Apache APISIX, including centralized management, RBAC, audit logging, and professional support. - [Enterprise Features Overview](https://docs.api7.ai/api7-gateway/enterprise-features/overview.md): Discover the enterprise-grade features that distinguish API7 Gateway from open-source Apache APISIX, including centralized management, RBAC, audit logging, and professional support. #### permission-policies-and-boundaries Explore the concept of permission policies and boundaries in API7 Gateway, defining user access levels for enhanced security. - [Permission Policies and Boundaries](https://docs.api7.ai/api7-gateway/enterprise-features/permission-policies-and-boundaries.md): Explore the concept of permission policies and boundaries in API7 Gateway, defining user access levels for enhanced security. #### secret-providers Explore the concept of secret providers in API7 Gateway, which enhance security by storing sensitive data using third-party tools. - [Secret Providers](https://docs.api7.ai/api7-gateway/enterprise-features/secret-providers.md): Explore the concept of secret providers in API7 Gateway, which enhance security by storing sensitive data using third-party tools. #### security-hardening Explore the concept of security hardening in API7 Gateway, designed to protect API infrastructure against threats and vulnerabilities. - [Security Hardening](https://docs.api7.ai/api7-gateway/enterprise-features/security-hardening.md): Explore the concept of security hardening in API7 Gateway, designed to protect API infrastructure against threats and vulnerabilities. ### getting-started #### learn-more Understand the API7 product family, the relationship between API7 Enterprise and Apache APISIX, and how they fit into your API management strategy. - [Learn More About API7 Products and APISIX](https://docs.api7.ai/api7-gateway/getting-started/learn-more.md): Understand the API7 product family, the relationship between API7 Enterprise and Apache APISIX, and how they fit into your API management strategy. #### management-options Compare the API7 Gateway Dashboard, Admin API, ADC, a7 CLI, and API7-MCP for interactive operations, automation, GitOps, and AI clients. - [Management Options](https://docs.api7.ai/api7-gateway/getting-started/management-options.md): Compare the API7 Gateway Dashboard, Admin API, ADC, a7 CLI, and API7-MCP for interactive operations, automation, GitOps, and AI clients. #### overview Get started with API7 Gateway. Understand the platform components, how they work together, and choose the right path for your use case. - [API7 Gateway Overview](https://docs.api7.ai/api7-gateway/getting-started/overview.md): Get started with API7 Gateway. Understand the platform components, how they work together, and choose the right path for your use case. #### quick-start Get API7 Gateway running locally with Docker Compose and proxy your first API request in under 10 minutes. - [Quick Start](https://docs.api7.ai/api7-gateway/getting-started/quick-start.md): Get API7 Gateway running locally with Docker Compose and proxy your first API request in under 10 minutes. #### tutorial-proxying-api-requests A hands-on tutorial that walks you through proxying API requests, adding authentication with Key Auth, and enabling rate limiting with the Limit Count plugin. - [Tutorial: Proxying and Managing API Requests via Plugins](https://docs.api7.ai/api7-gateway/getting-started/tutorial-proxying-api-requests.md): A hands-on tutorial that walks you through proxying API requests, adding authentication with Key Auth, and enabling rate limiting with the Limit Count plugin. ### how-to-guides #### api-security - [Configure Basic Authentication](https://docs.api7.ai/api7-gateway/how-to-guides/api-security/basic-auth.md): Learn how to secure your APIs by requiring clients to provide a standard username and password in the HTTP Authorization header. - [Configure Data Masking](https://docs.api7.ai/api7-gateway/how-to-guides/api-security/data-masking.md): Use the data-mask plugin to redact, replace, or remove sensitive fields from request data before it is written to access logs and logger plugin output, helping you meet GDPR, HIPAA, and PCI-DSS requirements. - [Forward External Auth User Info to Upstream](https://docs.api7.ai/api7-gateway/how-to-guides/api-security/forward-auth-user-info.md): Forward authenticated user information from OpenID Connect or SAML routes to upstream services as request headers or the consumer name. - [Configure HMAC Authentication](https://docs.api7.ai/api7-gateway/how-to-guides/api-security/hmac-auth.md): Learn how to secure your APIs with Hash-based Message Authentication Code (HMAC) for request signing in API7 Enterprise. - [Configure JWT Authentication](https://docs.api7.ai/api7-gateway/how-to-guides/api-security/jwt-auth.md): Learn how to secure your APIs using JSON Web Tokens (JWT) for stateless authentication in API7 Enterprise. - [Configure Key Authentication](https://docs.api7.ai/api7-gateway/how-to-guides/api-security/key-auth.md): Learn how to secure your APIs by requiring clients to provide a unique API key in the request header or query string. #### ops - [Configure Readiness and Liveness Probes](https://docs.api7.ai/api7-gateway/how-to-guides/ops/configure-readiness-probe.md): Configure readiness and liveness checks for API7 Gateway data planes in Kubernetes, Docker, and other non-Helm deployments. - [Create a Custom Role](https://docs.api7.ai/api7-gateway/how-to-guides/ops/create-custom-role.md): Create a custom role in API7 Gateway by defining permission policies, attaching them to a role, and assigning the role to a user. - [Design a Custom Role System](https://docs.api7.ai/api7-gateway/how-to-guides/ops/design-custom-role-system.md): Design a scalable custom role system in API7 Gateway by combining roles, permission policies, labels, and permission boundaries. - [Manage Gateway Groups](https://docs.api7.ai/api7-gateway/how-to-guides/ops/multi-gateway-group.md): Learn how to create and manage multiple gateway groups to organize and isolate API traffic across environments or teams. - [Configure Secret Management](https://docs.api7.ai/api7-gateway/how-to-guides/ops/secret-manager.md): Learn how to configure secret providers and reference external secrets in API7 Enterprise without hardcoding sensitive values in gateway resources. #### overview Practical how-to guides for common API7 Gateway tasks and workflows. - [How-To Guides](https://docs.api7.ai/api7-gateway/how-to-guides/overview.md): Practical how-to guides for common API7 Gateway tasks and workflows. #### plugin-development - [Plugin Development Best Practices](https://docs.api7.ai/api7-gateway/how-to-guides/plugin-development/best-practices.md): Write custom Lua plugins that are correct, fast, and maintainable by using the built-in core library instead of low-level OpenResty APIs, validating configuration with a schema, and choosing the right phase and priority. - [Develop Custom Lua Plugins](https://docs.api7.ai/api7-gateway/how-to-guides/plugin-development/custom-lua-plugins.md): Learn how to develop, register, and test custom Lua plugins for API7 Gateway. - [Serverless Functions or Custom Plugins](https://docs.api7.ai/api7-gateway/how-to-guides/plugin-development/serverless-or-custom-plugins.md): Compare the built-in serverless-function plugins with custom Lua plugins, and choose the right way to run your own Lua logic in API7 Gateway. #### protocol-proxy - [Configure GraphQL Proxying](https://docs.api7.ai/api7-gateway/how-to-guides/protocol-proxy/graphql-proxy.md): Learn how to proxy GraphQL APIs through API7 Gateway and when to add GraphQL-aware plugins for rate limiting and caching. - [Configure gRPC Proxying](https://docs.api7.ai/api7-gateway/how-to-guides/protocol-proxy/grpc-proxy.md): Learn how to configure API7 Enterprise to proxy gRPC traffic, including REST-to-gRPC transcoding and gRPC-Web support for browser clients. - [Configure TCP/UDP Proxying](https://docs.api7.ai/api7-gateway/how-to-guides/protocol-proxy/tcp-udp-proxy.md): Configure API7 Gateway to proxy TCP and UDP (Layer 4) traffic to upstream services such as databases, message queues, and custom protocols. - [Configure WebSocket Proxying](https://docs.api7.ai/api7-gateway/how-to-guides/protocol-proxy/websocket-proxy.md): Learn how to enable and configure WebSocket proxying for your routes in API7 Enterprise to handle long-lived, bidirectional connections. #### traffic-management - [Implement Blue-Green Deployment](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/blue-green-deployment.md): Implement blue-green deployments using API7 Gateway to switch traffic between two upstream environments with zero downtime. - [Implement Canary Release](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/canary-release.md): Learn how to gradually shift traffic to a new version of your backend service using the traffic-split plugin. - [Configure CORS](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/cors.md): Learn how to configure Cross-Origin Resource Sharing (CORS) for your APIs to allow or restrict access from different domains. - [Configure Fault Injection](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/fault-injection.md): Learn how to test the resilience of your application by injecting HTTP errors and response delays. - [Conditionally Disable Global Plugins](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/global-plugin-exemption.md): Conditionally skip global plugin execution for specific routes using route labels and the _meta.filter mechanism. - [Configure Upstream Health Checks](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/health-check.md): Learn how to configure active and passive health checks for upstream services to ensure high availability. - [Configure Proxy Cache](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/proxy-cache.md): Learn how to improve API performance and reduce upstream load by caching responses at the Gateway. - [Configure Proxy Mirror](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/proxy-mirror.md): Learn how to duplicate and send a percentage of real production traffic to a secondary service for testing and verification. - [Rewrite Proxy Requests](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/proxy-rewrite.md): Learn how to modify request URIs, methods, and headers before proxying them to upstream services. - [Configure Rate Limiting](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/rate-limiting.md): Learn how to configure simple and advanced rate limiting for your APIs to prevent abuse and ensure fair usage. - [Improve Rate Limiting Accuracy Across Gateway Instances](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/rate-limiting-accuracy.md): Understand where rate limiting error comes from when counter synchronization is batched across gateway instances, and why sliding windows keep the long-term error near zero. - [Configure Response Rewrite](https://docs.api7.ai/api7-gateway/how-to-guides/traffic-management/response-rewrite.md): Learn how to modify response status codes, headers, and body content before they are returned to the client. ### install #### deploy-high-availability Deploy API7 Gateway control plane and data plane in a high-availability configuration to eliminate single points of failure. - [Deploy for High Availability](https://docs.api7.ai/api7-gateway/install/deploy-high-availability.md): Deploy API7 Gateway control plane and data plane in a high-availability configuration to eliminate single points of failure. #### deploy-on-kubernetes Deploy API7 Enterprise on Kubernetes using Helm, including control plane setup, data plane configuration with mTLS, and cloud-specific guidance for AWS EKS, GCP GKE, and Azure AKS. - [Deploy on Kubernetes](https://docs.api7.ai/api7-gateway/install/deploy-on-kubernetes.md): Deploy API7 Enterprise on Kubernetes using Helm, including control plane setup, data plane configuration with mTLS, and cloud-specific guidance for AWS EKS, GCP GKE, and Azure AKS. #### deploy-on-openshift Deploy API7 Gateway on Red Hat OpenShift with proper Security Context Constraints (SCCs), service accounts, and Helm chart configuration. - [Deploy on OpenShift](https://docs.api7.ai/api7-gateway/install/deploy-on-openshift.md): Deploy API7 Gateway on Red Hat OpenShift with proper Security Context Constraints (SCCs), service accounts, and Helm chart configuration. #### deploy-with-docker-compose Install a complete API7 Gateway stack with Docker Compose using an online quickstart script, an offline bundle, or a customizable manual deployment. - [Deploy with Docker Compose](https://docs.api7.ai/api7-gateway/install/deploy-with-docker-compose.md): Install a complete API7 Gateway stack with Docker Compose using an online quickstart script, an offline bundle, or a customizable manual deployment. #### deploy-with-rpm Deploy API7 Enterprise on Red Hat-family systems (RHEL, Rocky Linux, AlmaLinux, and CentOS 8 and 9) from the offline RPM bundle, fully air-gapped and managed by systemd. - [Deploy with the Offline RPM Bundle](https://docs.api7.ai/api7-gateway/install/deploy-with-rpm.md): Deploy API7 Enterprise on Red Hat-family systems (RHEL, Rocky Linux, AlmaLinux, and CentOS 8 and 9) from the offline RPM bundle, fully air-gapped and managed by systemd. #### installation-faq Frequently asked questions and troubleshooting common installation issues. - [Installation FAQ](https://docs.api7.ai/api7-gateway/install/installation-faq.md): Frequently asked questions and troubleshooting common installation issues. #### installation-packages Container images, Helm charts, and CLI tools used to install API7 Enterprise — including the official Docker Hub repositories for the gateway, DP manager, Developer Portal, and integrated control plane images. - [Installation Packages](https://docs.api7.ai/api7-gateway/install/installation-packages.md): Container images, Helm charts, and CLI tools used to install API7 Enterprise — including the official Docker Hub repositories for the gateway, DP manager, Developer Portal, and integrated control plane images. #### overview Compare Kubernetes, Docker Compose, and package-based deployment options for API7 Gateway across production, evaluation, and air-gapped environments. - [Install API7 Gateway On-Premises](https://docs.api7.ai/api7-gateway/install/overview.md): Compare Kubernetes, Docker Compose, and package-based deployment options for API7 Gateway across production, evaluation, and air-gapped environments. #### rpm-from-docker When the Dashboard does not yet generate an RPM data-plane script, extract the connection parameters from the Docker script to onboard an api7-gateway RPM. - [Deploy an RPM Data Plane from a Docker Script](https://docs.api7.ai/api7-gateway/install/rpm-from-docker.md): When the Dashboard does not yet generate an RPM data-plane script, extract the connection parameters from the Docker script to onboard an api7-gateway RPM. #### rpm-quickstart Bring up an API7 Enterprise control plane on a single host in minutes with the bundled quickstart.sh for evaluation and PoC. - [Single-Host RPM Evaluation Quickstart](https://docs.api7.ai/api7-gateway/install/rpm-quickstart.md): Bring up an API7 Enterprise control plane on a single host in minutes with the bundled quickstart.sh for evaluation and PoC. #### supported-versions-and-interoperability Version compatibility matrix for API7 Enterprise components and infrastructure. - [Supported Versions and Interoperability](https://docs.api7.ai/api7-gateway/install/supported-versions-and-interoperability.md): Version compatibility matrix for API7 Enterprise components and infrastructure. #### system-requirements Hardware, operating system, and network requirements for installing API7 Gateway. - [System Requirements](https://docs.api7.ai/api7-gateway/install/system-requirements.md): Hardware, operating system, and network requirements for installing API7 Gateway. ### key-concepts #### architecture Deep dive into the decoupled control plane and data plane architecture of API7 Enterprise. - [Architecture](https://docs.api7.ai/api7-gateway/key-concepts/architecture.md): Deep dive into the decoupled control plane and data plane architecture of API7 Enterprise. #### consumers-and-credentials Identity management and authentication using Consumers and Credentials. - [Consumers and Credentials](https://docs.api7.ai/api7-gateway/key-concepts/consumers-and-credentials.md): Identity management and authentication using Consumers and Credentials. #### gateway-groups Logical grouping of data plane instances for environment isolation and configuration management. - [Gateway Groups](https://docs.api7.ai/api7-gateway/key-concepts/gateway-groups.md): Logical grouping of data plane instances for environment isolation and configuration management. #### overview Introduction to the core abstractions and entities in API7 Enterprise. - [Key Concepts Overview](https://docs.api7.ai/api7-gateway/key-concepts/overview.md): Introduction to the core abstractions and entities in API7 Enterprise. #### plugins Modular components to intercept and modify API traffic in API7 Gateway. - [Plugins](https://docs.api7.ai/api7-gateway/key-concepts/plugins.md): Modular components to intercept and modify API traffic in API7 Gateway. #### service-discovery Dynamic upstream resolution in API7 Gateway — automatically discover backend endpoints from Kubernetes Services, Nacos, and Consul registries without hardcoding IPs. - [Service Discovery](https://docs.api7.ai/api7-gateway/key-concepts/service-discovery.md): Dynamic upstream resolution in API7 Gateway — automatically discover backend endpoints from Kubernetes Services, Nacos, and Consul registries without hardcoding IPs. #### services-and-routes Understand how to group and route API traffic using Services and Routes. - [Services and Routes](https://docs.api7.ai/api7-gateway/key-concepts/services-and-routes.md): Understand how to group and route API traffic using Services and Routes. #### ssl-certificates Understanding how API7 Gateway manages SSL/TLS certificates and terminates secure connections. - [SSL Certificates](https://docs.api7.ai/api7-gateway/key-concepts/ssl-certificates.md): Understanding how API7 Gateway manages SSL/TLS certificates and terminates secure connections. #### stream-routes Understand stream routes in API7 Gateway for proxying TCP and UDP (Layer 4) traffic, including matching rules, supported plugins, and the relationship with stream services. - [Stream Routes](https://docs.api7.ai/api7-gateway/key-concepts/stream-routes.md): Understand stream routes in API7 Gateway for proxying TCP and UDP (Layer 4) traffic, including matching rules, supported plugins, and the relationship with stream services. #### upstreams-and-load-balancing Definitions of backend targets, including load balancing algorithms and health checks. - [Upstreams and Load Balancing](https://docs.api7.ai/api7-gateway/key-concepts/upstreams-and-load-balancing.md): Definitions of backend targets, including load balancing algorithms and health checks. ### observability #### alerts Configure alert policies and contact points in API7 Gateway to get notified by email or webhook when gateway instances go offline, certificates expire, or error rates breach a threshold. - [Configure Alerts](https://docs.api7.ai/api7-gateway/observability/alerts.md): Configure alert policies and contact points in API7 Gateway to get notified by email or webhook when gateway instances go offline, certificates expire, or error rates breach a threshold. #### consumer-label-based-logging Include consumer labels in API7 Gateway access logs for per-consumer traffic tracking and analysis. - [Include Consumer Labels in Access Logs](https://docs.api7.ai/api7-gateway/observability/consumer-label-based-logging.md): Include consumer labels in API7 Gateway access logs for per-consumer traffic tracking and analysis. #### debug-sessions Find out exactly what the gateway did to a single request — which plugins ran, in what order, which one was slow, and what it logged. Capture on demand, or automatically when an alert fires. - [Capture Request Traces with Debug Sessions](https://docs.api7.ai/api7-gateway/observability/debug-sessions.md): Find out exactly what the gateway did to a single request — which plugins ran, in what order, which one was slow, and what it logged. Capture on demand, or automatically when an alert fires. #### external-prometheus Point API7 Gateway at an existing external Prometheus and remove the bundled instance from a Docker Compose deployment, keeping the Dashboard Monitoring page working throughout. - [Use an Existing Prometheus](https://docs.api7.ai/api7-gateway/observability/external-prometheus.md): Point API7 Gateway at an existing external Prometheus and remove the bundled instance from a Docker Compose deployment, keeping the Dashboard Monitoring page working throughout. #### kubernetes-error-log-forwarding Forward API7 Gateway error logs from Kubernetes to Splunk by using the Splunk OpenTelemetry Collector. - [Send Kubernetes Error Logs to Splunk](https://docs.api7.ai/api7-gateway/observability/kubernetes-error-log-forwarding.md): Forward API7 Gateway error logs from Kubernetes to Splunk by using the Splunk OpenTelemetry Collector. #### kubernetes-log-collection Collect API7 Gateway access and error logs on Kubernetes with the OpenTelemetry Collector or Filebeat, either from container output or from log files written inside the pod. - [Collect Gateway Logs on Kubernetes](https://docs.api7.ai/api7-gateway/observability/kubernetes-log-collection.md): Collect API7 Gateway access and error logs on Kubernetes with the OpenTelemetry Collector or Filebeat, either from container output or from log files written inside the pod. #### logging Understand the access and error logs that API7 Gateway produces, how to configure their format and verbosity, and how to forward them to a centralized log management system. - [Configure Centralized Logging](https://docs.api7.ai/api7-gateway/observability/logging.md): Understand the access and error logs that API7 Gateway produces, how to configure their format and verbosity, and how to forward them to a centralized log management system. #### metrics Monitor API7 Gateway metrics via the built-in Dashboard page or by scraping the Prometheus endpoint exposed by the data plane. - [Monitor Metrics](https://docs.api7.ai/api7-gateway/observability/metrics.md): Monitor API7 Gateway metrics via the built-in Dashboard page or by scraping the Prometheus endpoint exposed by the data plane. #### splunk-integration Send API7 Gateway access logs to Splunk using the HTTP Event Collector (HEC) logging plugin. - [Send Access Logs to Splunk](https://docs.api7.ai/api7-gateway/observability/splunk-integration.md): Send API7 Gateway access logs to Splunk using the HTTP Event Collector (HEC) logging plugin. #### tracing Enable distributed tracing in API7 Gateway with the OpenTelemetry plugin. Configure the OTLP collector, sampling strategy, and per-route spans to visualize request flows and analyze latency bottlenecks. - [Configure Distributed Tracing](https://docs.api7.ai/api7-gateway/observability/tracing.md): Enable distributed tracing in API7 Gateway with the OpenTelemetry plugin. Configure the OTLP collector, sampling strategy, and per-route spans to visualize request flows and analyze latency bottlenecks. ### overview Discover API7 Gateway, a dynamic, high-performance API gateway for cloud-native environments, built on Apache APISIX. Learn about its architecture, features, and deployment models. - [API7 Gateway](https://docs.api7.ai/api7-gateway/overview.md): Discover API7 Gateway, a dynamic, high-performance API gateway for cloud-native environments, built on Apache APISIX. Learn about its architecture, features, and deployment models. ### release-notes Review the latest updates, features, improvements, and bug fixes for each version of API7 Gateway in release notes. - [API7 Gateway Release Notes](https://docs.api7.ai/api7-gateway/release-notes.md): Review the latest updates, features, improvements, and bug fixes for each version of API7 Gateway in release notes. ### scalability #### autoscale-on-kubernetes Autoscale API7 Gateway data plane pods on Kubernetes with a Horizontal Pod Autoscaler (HPA) to handle changing traffic automatically. - [Autoscale Data Plane on Kubernetes](https://docs.api7.ai/api7-gateway/scalability/autoscale-on-kubernetes.md): Autoscale API7 Gateway data plane pods on Kubernetes with a Horizontal Pod Autoscaler (HPA) to handle changing traffic automatically. ### security-and-compliance #### access-control-lists Restrict API access using the consumer-restriction plugin in API7 Gateway. Configure allowlists and denylists keyed by consumer name, consumer group ID, service ID, or route ID for granular security. - [Access Control Lists (ACLs)](https://docs.api7.ai/api7-gateway/security-and-compliance/access-control-lists.md): Restrict API access using the consumer-restriction plugin in API7 Gateway. Configure allowlists and denylists keyed by consumer name, consumer group ID, service ID, or route ID for granular security. #### audit-logs Track all administrative changes in the API7 Control Plane with detailed audit logs. Learn how to review, manage, and export audit trails for security and compliance. - [Audit Logs](https://docs.api7.ai/api7-gateway/security-and-compliance/audit-logs.md): Track all administrative changes in the API7 Control Plane with detailed audit logs. Learn how to review, manage, and export audit trails for security and compliance. #### authenticate - [Client mTLS Authentication](https://docs.api7.ai/api7-gateway/security-and-compliance/authenticate/client-mtls.md): Configure mutual TLS (mTLS) on API7 Gateway to authenticate API clients with X.509 certificates before allowing access to your APIs. - [Mutual TLS between Control Plane and Data Plane](https://docs.api7.ai/api7-gateway/security-and-compliance/authenticate/mutual-tls-cp-dp.md): Learn how API7 Gateway uses a robust, PKI-based mutual TLS (mTLS) model to secure all communication between the Control Plane and Data Plane. - [SCIM Provisioning for Dashboard](https://docs.api7.ai/api7-gateway/security-and-compliance/authenticate/scim.md): Configure SCIM provisioning for API7 Gateway Dashboard with supported identity providers. - [SCIM Provisioning with Microsoft Entra ID](https://docs.api7.ai/api7-gateway/security-and-compliance/authenticate/scim-microsoft-entra-id.md): Configure SCIM provisioning with Microsoft Entra ID for API7 Gateway Dashboard. - [SCIM Provisioning with Okta](https://docs.api7.ai/api7-gateway/security-and-compliance/authenticate/scim-okta.md): Configure SCIM provisioning with Okta for API7 Gateway Dashboard. - [Upstream mTLS](https://docs.api7.ai/api7-gateway/security-and-compliance/authenticate/upstream-mtls.md): Configure mutual TLS (mTLS) between API7 Gateway and upstream services so the gateway presents a client certificate to the upstream and optionally validates the upstream's server certificate. #### ip-restrictions-control-plane Restrict which client IP addresses can reach the API7 Control Plane (Dashboard and Admin API) using the Control Plane's built-in IP allow list, with complementary network-level controls for defence in depth. - [IP Restrictions for Control Plane](https://docs.api7.ai/api7-gateway/security-and-compliance/ip-restrictions-control-plane.md): Restrict which client IP addresses can reach the API7 Control Plane (Dashboard and Admin API) using the Control Plane's built-in IP allow list, with complementary network-level controls for defence in depth. #### oauth-oidc Secure your APIs with modern token-based authentication. Learn how to integrate API7 Gateway with OAuth 2.0 and OpenID Connect (OIDC) providers like Okta and Keycloak. - [OAuth 2.0 and OIDC](https://docs.api7.ai/api7-gateway/security-and-compliance/oauth-oidc.md): Secure your APIs with modern token-based authentication. Learn how to integrate API7 Gateway with OAuth 2.0 and OpenID Connect (OIDC) providers like Okta and Keycloak. #### open-source-licenses API7 Gateway's commitment to open-source software and license compliance. Review the major open-source components used and their respective licenses. - [Open Source Licenses](https://docs.api7.ai/api7-gateway/security-and-compliance/open-source-licenses.md): API7 Gateway's commitment to open-source software and license compliance. Review the major open-source components used and their respective licenses. #### overview Learn how API7 Gateway provides comprehensive security and compliance features to protect your APIs, including authentication, authorization, encryption, and auditing capabilities. - [Security and Compliance Overview](https://docs.api7.ai/api7-gateway/security-and-compliance/overview.md): Learn how API7 Gateway provides comprehensive security and compliance features to protect your APIs, including authentication, authorization, encryption, and auditing capabilities. #### permission-policies-and-boundaries Author permission policies in API7 Gateway using the native JSON document format. Learn the statement structure, allowed actions, ARN-style resources, label-based conditions, and how permission boundaries enforce maximum allowable permissions. - [Permission Policies and Boundaries](https://docs.api7.ai/api7-gateway/security-and-compliance/permission-policies-and-boundaries.md): Author permission policies in API7 Gateway using the native JSON document format. Learn the statement structure, allowed actions, ARN-style resources, label-based conditions, and how permission boundaries enforce maximum allowable permissions. #### role-based-access-control Manage user access to the API7 Gateway control plane with Role-Based Access Control. Assign users to roles, attach permission policies, and enforce least-privilege access across gateway groups. - [Role-Based Access Control (RBAC)](https://docs.api7.ai/api7-gateway/security-and-compliance/role-based-access-control.md): Manage user access to the API7 Gateway control plane with Role-Based Access Control. Assign users to roles, attach permission policies, and enforce least-privilege access across gateway groups. #### secure-credentials Protect your sensitive information with API7 Gateway's secure credentials management. Learn how to manage SSL certificates and integrate with external secret managers like HashiCorp Vault. - [Secure Credentials Management](https://docs.api7.ai/api7-gateway/security-and-compliance/secure-credentials.md): Protect your sensitive information with API7 Gateway's secure credentials management. Learn how to manage SSL certificates and integrate with external secret managers like HashiCorp Vault. #### sso-dashboard Centralize user access to the API7 Dashboard using Single Sign-On (SSO). Supports OIDC, SAML, LDAP, and CAS protocols with automatic role mapping. - [SSO for Dashboard](https://docs.api7.ai/api7-gateway/security-and-compliance/sso-dashboard.md): Centralize user access to the API7 Dashboard using Single Sign-On (SSO). Supports OIDC, SAML, LDAP, and CAS protocols with automatic role mapping. #### sso-ldap Configure Single Sign-On (SSO) for the API7 Dashboard using LDAP, allowing users to authenticate with their existing directory service credentials. - [SSO with LDAP](https://docs.api7.ai/api7-gateway/security-and-compliance/sso-ldap.md): Configure Single Sign-On (SSO) for the API7 Dashboard using LDAP, allowing users to authenticate with their existing directory service credentials. #### sso-oidc Configure Single Sign-On (SSO) for the API7 Dashboard using OpenID Connect (OIDC), with provider-specific guidance for Keycloak, Microsoft Entra ID, and Auth0. - [SSO with OIDC](https://docs.api7.ai/api7-gateway/security-and-compliance/sso-oidc.md): Configure Single Sign-On (SSO) for the API7 Dashboard using OpenID Connect (OIDC), with provider-specific guidance for Keycloak, Microsoft Entra ID, and Auth0. #### sso-saml Configure Single Sign-On (SSO) for the API7 Dashboard using SAML 2.0, with provider-specific guidance for Microsoft Entra ID and Okta. - [SSO with SAML](https://docs.api7.ai/api7-gateway/security-and-compliance/sso-saml.md): Configure Single Sign-On (SSO) for the API7 Dashboard using SAML 2.0, with provider-specific guidance for Microsoft Entra ID and Okta. #### trust-center The API7 Trust Center is your centralized resource for all security and compliance-related information, including certifications, security reports, and best practice guides. - [Trust Center](https://docs.api7.ai/api7-gateway/security-and-compliance/trust-center.md): The API7 Trust Center is your centralized resource for all security and compliance-related information, including certifications, security reports, and best practice guides. #### verify-image-signatures Verify API7 Enterprise container image signatures with Cosign and keyless OIDC-based verification. Protect your supply chain by confirming every image was built and signed by API7.ai. - [Verify Image Signatures](https://docs.api7.ai/api7-gateway/security-and-compliance/verify-image-signatures.md): Verify API7 Enterprise container image signatures with Cosign and keyless OIDC-based verification. Protect your supply chain by confirming every image was built and signed by API7.ai. #### vulnerability-scanning Learn about API7.ai's security testing and vulnerability scanning practices for API7 Gateway, including our CVE reporting and patching processes. - [Vulnerability Scanning](https://docs.api7.ai/api7-gateway/security-and-compliance/vulnerability-scanning.md): Learn about API7.ai's security testing and vulnerability scanning practices for API7 Gateway, including our CVE reporting and patching processes. #### web-application-firewall Learn when to use a Web Application Firewall with API7 Gateway, how it fits into your security model, and where to find the current documented integration path. - [Web Application Firewall (WAF)](https://docs.api7.ai/api7-gateway/security-and-compliance/web-application-firewall.md): Learn when to use a Web Application Firewall with API7 Gateway, how it fits into your security model, and where to find the current documented integration path. ### troubleshooting Diagnose and resolve common issues with API7 Gateway deployments, including connectivity problems, configuration errors, and performance degradation. - [Troubleshoot API7 Gateway](https://docs.api7.ai/api7-gateway/troubleshooting.md): Diagnose and resolve common issues with API7 Gateway deployments, including connectivity problems, configuration errors, and performance degradation. ### upgrade-guides #### backup-and-restore Back up and restore API7 Gateway CP data and gateway-group configuration with database-native tools and ADC. - [Backup and Restoration](https://docs.api7.ai/api7-gateway/upgrade-guides/backup-and-restore.md): Back up and restore API7 Gateway CP data and gateway-group configuration with database-native tools and ADC. #### cluster-migration A step-by-step guide for migrating your API7 Gateway deployment to a new cluster with zero downtime. - [API Gateway Cluster Migration](https://docs.api7.ai/api7-gateway/upgrade-guides/cluster-migration.md): A step-by-step guide for migrating your API7 Gateway deployment to a new cluster with zero downtime. #### dual-cluster Upgrade API7 Gateway with independent source and target clusters, controlled traffic shifting, write reconciliation, validation, and safe rollback. - [Dual-Cluster Upgrade](https://docs.api7.ai/api7-gateway/upgrade-guides/dual-cluster.md): Upgrade API7 Gateway with independent source and target clusters, controlled traffic shifting, write reconciliation, validation, and safe rollback. #### in-place Upgrade the API7 Gateway Control Plane against its existing database with a write freeze, source shutdown, target validation, and recoverable rollback. - [Control Plane In-Place Upgrade](https://docs.api7.ai/api7-gateway/upgrade-guides/in-place.md): Upgrade the API7 Gateway Control Plane against its existing database with a write freeze, source shutdown, target validation, and recoverable rollback. #### lts-upgrades Choose the supported API7 Gateway route and exact version anchors for upgrading from a predecessor LTS release to this LTS release. - [Choose an LTS Upgrade Path](https://docs.api7.ai/api7-gateway/upgrade-guides/lts-upgrades.md): Choose the supported API7 Gateway route and exact version anchors for upgrading from a predecessor LTS release to this LTS release. #### rolling-upgrade Replace API7 Gateway Data Plane nodes gradually with capacity planning, canary validation, add-before-drain rollout, traffic checks, and rollback. - [Data Plane Rolling Upgrade](https://docs.api7.ai/api7-gateway/upgrade-guides/rolling-upgrade.md): Replace API7 Gateway Data Plane nodes gradually with capacity planning, canary validation, add-before-drain rollout, traffic checks, and rollback. #### upgrade Plan an API7 Gateway upgrade by confirming the supported version path, selecting deployment strategies, preparing backups, and rehearsing rollback. - [Plan an API7 Gateway Upgrade](https://docs.api7.ai/api7-gateway/upgrade-guides/upgrade.md): Plan an API7 Gateway upgrade by confirming the supported version path, selecting deployment strategies, preparing backups, and rehearsing rollback. #### upgrade-3.8-to-3.10 Upgrade API7 Gateway from 3.8.23 to 3.10.6 on Helm and Kubernetes with external PostgreSQL, validation, and backup-based rollback. - [Upgrade from 3.8 LTS to 3.10 LTS](https://docs.api7.ai/api7-gateway/upgrade-guides/upgrade-3.8-to-3.10.md): Upgrade API7 Gateway from 3.8.23 to 3.10.6 on Helm and Kubernetes with external PostgreSQL, validation, and backup-based rollback. #### upgrade-3.9-to-3.10 Upgrade API7 Gateway directly from 3.9.19 to 3.10.6 on Helm and Kubernetes with external PostgreSQL 15.x, validation, and rollback. - [Upgrade from 3.9 LTS to 3.10 LTS](https://docs.api7.ai/api7-gateway/upgrade-guides/upgrade-3.9-to-3.10.md): Upgrade API7 Gateway directly from 3.9.19 to 3.10.6 on Helm and Kubernetes with external PostgreSQL 15.x, validation, and rollback. ### version-support-policy API7 Enterprise version support lifecycle, LTS policy, and upgrade guidance. Learn how long each release is supported and which versions are currently designated as LTS. - [Version Support Policy](https://docs.api7.ai/api7-gateway/version-support-policy.md): API7 Enterprise version support lifecycle, LTS policy, and upgrade guidance. Learn how long each release is supported and which versions are currently designated as LTS. ## apisix ### reference #### admin-api The Apache APISIX Admin API is a RESTful interface for managing all APISIX gateway resources — routes, upstreams, services, consumers, SSL certificates, plugins, and more. APISIX Admin API allows... - [Apache APISIX Admin API](https://docs.api7.ai/apisix/reference/admin-api.md): The Apache APISIX Admin API is a RESTful interface for managing all APISIX gateway resources — routes, upstreams, services, consumers, SSL certificates, plugins, and more. APISIX Admin API allows... #### control-api Use the APISIX Control API to inspect or control the runtime state of one APISIX instance. It is enabled by default at http://127.0.0.1:9090; configure apisix.enable control and apisix.control in... - [APISIX Control API](https://docs.api7.ai/apisix/reference/control-api.md): Use the APISIX Control API to inspect or control the runtime state of one APISIX instance. It is enabled by default at http://127.0.0.1:9090; configure apisix.enable control and apisix.control in... #### a6-cli Understand how the separately installed a6 CLI manages Apache APISIX through the Admin API and how it differs from other management tools. - [a6 CLI](https://docs.api7.ai/apisix/reference/a6-cli.md): Understand how the separately installed a6 CLI manages Apache APISIX through the Admin API and how it differs from other management tools. #### adc Use API Declarative CLI (ADC) to manage Apache APISIX configuration declaratively. - [API Declarative CLI (ADC)](https://docs.api7.ai/apisix/reference/adc.md): Use API Declarative CLI (ADC) to manage Apache APISIX configuration declaratively. #### api-standalone-usage Learn about api-driven standalone mode usage in Apache APISIX, which stores gateway configurations entirely in memory rather than in a configuration file. - [API-Driven Standalone Mode Usage](https://docs.api7.ai/apisix/reference/api-standalone-usage.md): Learn about api-driven standalone mode usage in Apache APISIX, which stores gateway configurations entirely in memory rather than in a configuration file. #### apisix-cli Discover the APISIX Command Line Interface (CLI), a tool designed for easy management and control of APISIX instances. - [APISIX CLI](https://docs.api7.ai/apisix/reference/apisix-cli.md): Discover the APISIX Command Line Interface (CLI), a tool designed for easy management and control of APISIX instances. #### apisix-expressions Learn about APISIX expressions, which combine variables and operators for route matching, request filtering, and other functionalities. - [APISIX Expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md): Learn about APISIX expressions, which combine variables and operators for route matching, request filtering, and other functionalities. #### apisix-mcp Use APISIX-MCP to expose Apache APISIX resource operations and gateway test requests to MCP-compatible AI clients and agents. - [APISIX Model Context Protocol (APISIX-MCP)](https://docs.api7.ai/apisix/reference/apisix-mcp.md): Use APISIX-MCP to expose Apache APISIX resource operations and gateway test requests to MCP-compatible AI clients and agents. #### batch-processor Understand how APISIX batches logging and telemetry entries, limits pending work, and controls memory use when a destination slows down or becomes unavailable. - [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md): Understand how APISIX batches logging and telemetry entries, limits pending work, and controls memory use when a destination slows down or becomes unavailable. #### built-in-variables Discover built-in variables in Apache APISIX that provide access to request-specific information for plugin configurations and routing. - [Built-In Variables](https://docs.api7.ai/apisix/reference/built-in-variables.md): Discover built-in variables in Apache APISIX that provide access to request-specific information for plugin configurations and routing. #### configuration-files Understand the configuration files in Apache APISIX, detailing how to customize parameters for various environments. - [Configuration Files](https://docs.api7.ai/apisix/reference/configuration-files.md): Understand the configuration files in Apache APISIX, detailing how to customize parameters for various environments. #### environment-variables Explore environment variables in Apache APISIX, which enable configurable settings during deployments for enhanced flexibility. - [Environment Variables](https://docs.api7.ai/apisix/reference/environment-variables.md): Explore environment variables in Apache APISIX, which enable configurable settings during deployments for enhanced flexibility. #### file-standalone-configurations Learn about file-driven standalone configurations in Apache APISIX, which allow the gateway to load gateway configurations from a YAML or JSON file. - [File-Driven Standalone Mode Configurations](https://docs.api7.ai/apisix/reference/file-standalone-configurations.md): Learn about file-driven standalone configurations in Apache APISIX, which allow the gateway to load gateway configurations from a YAML or JSON file. #### helm-chart Learn where to find Apache APISIX Helm chart values and how Helm values are rendered into gateway configuration. - [Helm Chart](https://docs.api7.ai/apisix/reference/helm-chart.md): Learn where to find Apache APISIX Helm chart values and how Helm values are rendered into gateway configuration. #### plugin-common-configurations Explore common plugin configurations in Apache APISIX, which enable universal settings for all plugins via meta attributes. - [Plugin Common Configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md): Explore common plugin configurations in Apache APISIX, which enable universal settings for all plugins via meta attributes. #### router-options Understand router options in Apache APISIX, which detail how to adjust routing behaviors for API request handling. - [Router Options](https://docs.api7.ai/apisix/reference/router-options.md): Understand router options in Apache APISIX, which detail how to adjust routing behaviors for API request handling. ### ai-agent-skills Manage Apache APISIX with AI coding agents like Claude Code and Cursor. Browse open-source agent skills that configure your API gateway from natural language. - [AI Agent Skills for Apache APISIX](https://docs.api7.ai/apisix/ai-agent-skills.md): Manage Apache APISIX with AI coding agents like Claude Code and Cursor. Browse open-source agent skills that configure your API gateway from natural language. #### a6-persona-developer Persona skill for API developers building and testing APIs on APISIX using the a6 CLI. Provides decision frameworks for API design, route configuration,… - [a6-persona-developer](https://docs.api7.ai/apisix/ai-agent-skills/a6-persona-developer.md): Persona skill for API developers building and testing APIs on APISIX using the a6 CLI. Provides decision frameworks for API design, route configuration,… #### a6-persona-operator Persona skill for platform operators and DevOps engineers managing APISIX instances using the a6 CLI. Provides decision frameworks for day-to-day operat… - [a6-persona-operator](https://docs.api7.ai/apisix/ai-agent-skills/a6-persona-operator.md): Persona skill for platform operators and DevOps engineers managing APISIX instances using the a6 CLI. Provides decision frameworks for day-to-day operat… #### a6-plugin-ai-content-moderation Skill for configuring APISIX AWS and Aliyun AI content moderation via the a6 CLI. Covers request and response checks, streaming, deny_code, and ai-proxy. - [a6-plugin-ai-content-moderation](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-ai-content-moderation.md): Skill for configuring APISIX AWS and Aliyun AI content moderation via the a6 CLI. Covers request and response checks, streaming, deny_code, and ai-proxy. #### a6-plugin-ai-prompt-decorator Skill for configuring the Apache APISIX ai-prompt-decorator plugin via the a6 CLI. Covers prepending and appending system/user/assistant messages to LLM… - [a6-plugin-ai-prompt-decorator](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-ai-prompt-decorator.md): Skill for configuring the Apache APISIX ai-prompt-decorator plugin via the a6 CLI. Covers prepending and appending system/user/assistant messages to LLM… #### a6-plugin-ai-prompt-template Skill for configuring the Apache APISIX ai-prompt-template plugin via the a6 CLI. Covers defining reusable prompt templates with variable placeholders,… - [a6-plugin-ai-prompt-template](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-ai-prompt-template.md): Skill for configuring the Apache APISIX ai-prompt-template plugin via the a6 CLI. Covers defining reusable prompt templates with variable placeholders,… #### a6-plugin-ai-proxy Skill for configuring the Apache APISIX ai-proxy plugin via the a6 CLI. Covers proxying requests to LLM providers (OpenAI, Azure OpenAI, DeepSeek, Anthr… - [a6-plugin-ai-proxy](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-ai-proxy.md): Skill for configuring the Apache APISIX ai-proxy plugin via the a6 CLI. Covers proxying requests to LLM providers (OpenAI, Azure OpenAI, DeepSeek, Anthr… #### a6-plugin-basic-auth Skill for configuring the Apache APISIX basic-auth plugin via the a6 CLI. Covers HTTP Basic Authentication setup on routes, consumer credential binding… - [a6-plugin-basic-auth](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-basic-auth.md): Skill for configuring the Apache APISIX basic-auth plugin via the a6 CLI. Covers HTTP Basic Authentication setup on routes, consumer credential binding… #### a6-plugin-consumer-restriction Skill for configuring the Apache APISIX consumer-restriction plugin via the a6 CLI. Covers restricting access by consumer name, consumer group ID, servi… - [a6-plugin-consumer-restriction](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-consumer-restriction.md): Skill for configuring the Apache APISIX consumer-restriction plugin via the a6 CLI. Covers restricting access by consumer name, consumer group ID, servi… #### a6-plugin-cors Skill for configuring the Apache APISIX cors plugin via the a6 CLI. Covers Cross-Origin Resource Sharing setup on routes, allow_origins, allow_methods,… - [a6-plugin-cors](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-cors.md): Skill for configuring the Apache APISIX cors plugin via the a6 CLI. Covers Cross-Origin Resource Sharing setup on routes, allow_origins, allow_methods,… #### a6-plugin-datadog Skill for configuring the Apache APISIX datadog plugin via the a6 CLI. Covers pushing custom metrics to Datadog via DogStatsD, metric tags, batching, pl… - [a6-plugin-datadog](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-datadog.md): Skill for configuring the Apache APISIX datadog plugin via the a6 CLI. Covers pushing custom metrics to Datadog via DogStatsD, metric tags, batching, pl… #### a6-plugin-ext-plugin Skill for configuring the Apache APISIX external plugin system (ext-plugin-pre-req, ext-plugin-post-req, ext-plugin-post-resp) via the a6 CLI. Covers Pl… - [a6-plugin-ext-plugin](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-ext-plugin.md): Skill for configuring the Apache APISIX external plugin system (ext-plugin-pre-req, ext-plugin-post-req, ext-plugin-post-resp) via the a6 CLI. Covers Pl… #### a6-plugin-fault-injection Skill for configuring the Apache APISIX fault-injection plugin via the a6 CLI. Covers injecting delays and HTTP aborts for chaos engineering, percentage… - [a6-plugin-fault-injection](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-fault-injection.md): Skill for configuring the Apache APISIX fault-injection plugin via the a6 CLI. Covers injecting delays and HTTP aborts for chaos engineering, percentage… #### a6-plugin-grpc-transcode Skill for configuring the Apache APISIX grpc-transcode plugin via the a6 CLI. Covers converting RESTful HTTP requests to gRPC, proto file management, pb… - [a6-plugin-grpc-transcode](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-grpc-transcode.md): Skill for configuring the Apache APISIX grpc-transcode plugin via the a6 CLI. Covers converting RESTful HTTP requests to gRPC, proto file management, pb… #### a6-plugin-hmac-auth Skill for configuring the Apache APISIX hmac-auth plugin via the a6 CLI. Covers HMAC signature authentication, consumer credential binding with key_id/s… - [a6-plugin-hmac-auth](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-hmac-auth.md): Skill for configuring the Apache APISIX hmac-auth plugin via the a6 CLI. Covers HMAC signature authentication, consumer credential binding with key_id/s… #### a6-plugin-http-logger Skill for configuring the Apache APISIX http-logger plugin via the a6 CLI. Covers pushing access logs to HTTP/HTTPS endpoints in batches, custom log for… - [a6-plugin-http-logger](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-http-logger.md): Skill for configuring the Apache APISIX http-logger plugin via the a6 CLI. Covers pushing access logs to HTTP/HTTPS endpoints in batches, custom log for… #### a6-plugin-ip-restriction Skill for configuring the Apache APISIX ip-restriction plugin via the a6 CLI. Covers IP whitelist/blacklist setup on routes, CIDR range support, IPv4/IP… - [a6-plugin-ip-restriction](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-ip-restriction.md): Skill for configuring the Apache APISIX ip-restriction plugin via the a6 CLI. Covers IP whitelist/blacklist setup on routes, CIDR range support, IPv4/IP… #### a6-plugin-jwt-auth Skill for configuring the Apache APISIX jwt-auth plugin via the a6 CLI. Covers JWT token authentication, HS256/RS256 algorithm selection, consumer crede… - [a6-plugin-jwt-auth](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-jwt-auth.md): Skill for configuring the Apache APISIX jwt-auth plugin via the a6 CLI. Covers JWT token authentication, HS256/RS256 algorithm selection, consumer crede… #### a6-plugin-kafka-logger Skill for configuring the Apache APISIX kafka-logger plugin via the a6 CLI. Covers pushing access logs to Apache Kafka topics, broker configuration, SAS… - [a6-plugin-kafka-logger](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-kafka-logger.md): Skill for configuring the Apache APISIX kafka-logger plugin via the a6 CLI. Covers pushing access logs to Apache Kafka topics, broker configuration, SAS… #### a6-plugin-key-auth Skill for configuring the Apache APISIX key-auth plugin via the a6 CLI. Covers API key authentication setup on routes, consumer credential binding, key… - [a6-plugin-key-auth](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-key-auth.md): Skill for configuring the Apache APISIX key-auth plugin via the a6 CLI. Covers API key authentication setup on routes, consumer credential binding, key… #### a6-plugin-limit-count Skill for configuring the APISIX limit-count plugin via the a6 CLI. Covers fixed and sliding windows, Redis Sentinel, delayed sync, and shared quotas. - [a6-plugin-limit-count](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-limit-count.md): Skill for configuring the APISIX limit-count plugin via the a6 CLI. Covers fixed and sliding windows, Redis Sentinel, delayed sync, and shared quotas. #### a6-plugin-limit-req Skill for configuring the Apache APISIX limit-req plugin via the a6 CLI. Covers leaky-bucket rate limiting, rate/burst configuration, nodelay behavior,… - [a6-plugin-limit-req](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-limit-req.md): Skill for configuring the Apache APISIX limit-req plugin via the a6 CLI. Covers leaky-bucket rate limiting, rate/burst configuration, nodelay behavior,… #### a6-plugin-openid-connect Skill for configuring the APISIX openid-connect plugin via the a6 CLI. Covers authorization-code and bearer flows, PAR, DPoP, and session validation. - [a6-plugin-openid-connect](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-openid-connect.md): Skill for configuring the APISIX openid-connect plugin via the a6 CLI. Covers authorization-code and bearer flows, PAR, DPoP, and session validation. #### a6-plugin-prometheus Skill for configuring APISIX prometheus via the a6 CLI. Covers HTTP, LLM, and AI cache metrics, latency type labels, and Grafana dashboards. - [a6-plugin-prometheus](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-prometheus.md): Skill for configuring APISIX prometheus via the a6 CLI. Covers HTTP, LLM, and AI cache metrics, latency type labels, and Grafana dashboards. #### a6-plugin-proxy-rewrite Skill for configuring the Apache APISIX proxy-rewrite plugin via the a6 CLI. Covers rewriting request URI, host, method, headers, and scheme before forw… - [a6-plugin-proxy-rewrite](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-proxy-rewrite.md): Skill for configuring the Apache APISIX proxy-rewrite plugin via the a6 CLI. Covers rewriting request URI, host, method, headers, and scheme before forw… #### a6-plugin-redirect Skill for configuring the Apache APISIX redirect plugin via the a6 CLI. Covers URI redirects, HTTP-to-HTTPS redirection, regex-based URI rewriting, quer… - [a6-plugin-redirect](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-redirect.md): Skill for configuring the Apache APISIX redirect plugin via the a6 CLI. Covers URI redirects, HTTP-to-HTTPS redirection, regex-based URI rewriting, quer… #### a6-plugin-response-rewrite Skill for configuring the Apache APISIX response-rewrite plugin via the a6 CLI. Covers rewriting response status codes, headers, and body before returni… - [a6-plugin-response-rewrite](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-response-rewrite.md): Skill for configuring the Apache APISIX response-rewrite plugin via the a6 CLI. Covers rewriting response status codes, headers, and body before returni… #### a6-plugin-serverless Skill for configuring the Apache APISIX serverless-pre-function and serverless-post-function plugins via the a6 CLI. Covers inline Lua function executio… - [a6-plugin-serverless](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-serverless.md): Skill for configuring the Apache APISIX serverless-pre-function and serverless-post-function plugins via the a6 CLI. Covers inline Lua function executio… #### a6-plugin-skywalking Skill for configuring the Apache APISIX skywalking plugin via the a6 CLI. Covers distributed tracing with Apache SkyWalking OAP, sampling configuration,… - [a6-plugin-skywalking](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-skywalking.md): Skill for configuring the Apache APISIX skywalking plugin via the a6 CLI. Covers distributed tracing with Apache SkyWalking OAP, sampling configuration,… #### a6-plugin-traffic-split Skill for configuring the Apache APISIX traffic-split plugin via the a6 CLI. Covers weighted traffic splitting between upstreams with conditional match… - [a6-plugin-traffic-split](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-traffic-split.md): Skill for configuring the Apache APISIX traffic-split plugin via the a6 CLI. Covers weighted traffic splitting between upstreams with conditional match… #### a6-plugin-wolf-rbac Skill for configuring the Apache APISIX wolf-rbac plugin via the a6 CLI. Covers integration with the Wolf RBAC server for role-based access control, tok… - [a6-plugin-wolf-rbac](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-wolf-rbac.md): Skill for configuring the Apache APISIX wolf-rbac plugin via the a6 CLI. Covers integration with the Wolf RBAC server for role-based access control, tok… #### a6-plugin-zipkin Skill for configuring the Apache APISIX zipkin plugin via the a6 CLI. Covers distributed tracing with Zipkin, Jaeger, or any Zipkin-compatible collector… - [a6-plugin-zipkin](https://docs.api7.ai/apisix/ai-agent-skills/a6-plugin-zipkin.md): Skill for configuring the Apache APISIX zipkin plugin via the a6 CLI. Covers distributed tracing with Zipkin, Jaeger, or any Zipkin-compatible collector… #### a6-recipe-api-versioning Recipe skill for implementing API versioning strategies using the a6 CLI. Covers URI path versioning with proxy-rewrite, header-based versioning with tr… - [a6-recipe-api-versioning](https://docs.api7.ai/apisix/ai-agent-skills/a6-recipe-api-versioning.md): Recipe skill for implementing API versioning strategies using the a6 CLI. Covers URI path versioning with proxy-rewrite, header-based versioning with tr… #### a6-recipe-blue-green Recipe skill for implementing blue-green deployments using the a6 CLI. Covers creating two upstream environments, switching traffic instantly via route… - [a6-recipe-blue-green](https://docs.api7.ai/apisix/ai-agent-skills/a6-recipe-blue-green.md): Recipe skill for implementing blue-green deployments using the a6 CLI. Covers creating two upstream environments, switching traffic instantly via route… #### a6-recipe-canary Recipe skill for implementing canary releases using the a6 CLI. Covers gradual traffic shifting with the traffic-split plugin, header-based canary routi… - [a6-recipe-canary](https://docs.api7.ai/apisix/ai-agent-skills/a6-recipe-canary.md): Recipe skill for implementing canary releases using the a6 CLI. Covers gradual traffic shifting with the traffic-split plugin, header-based canary routi… #### a6-recipe-circuit-breaker Recipe skill for implementing circuit breaker patterns using the a6 CLI. Covers the api-breaker plugin for automatic upstream circuit breaking, configur… - [a6-recipe-circuit-breaker](https://docs.api7.ai/apisix/ai-agent-skills/a6-recipe-circuit-breaker.md): Recipe skill for implementing circuit breaker patterns using the a6 CLI. Covers the api-breaker plugin for automatic upstream circuit breaking, configur… #### a6-recipe-graphql-proxy Recipe skill for implementing GraphQL proxying patterns using the a6 CLI. Covers operation-based routing with built-in GraphQL variables, per-operation… - [a6-recipe-graphql-proxy](https://docs.api7.ai/apisix/ai-agent-skills/a6-recipe-graphql-proxy.md): Recipe skill for implementing GraphQL proxying patterns using the a6 CLI. Covers operation-based routing with built-in GraphQL variables, per-operation… #### a6-recipe-health-check Recipe skill for configuring upstream health checks using the a6 CLI. Covers active health checks (HTTP probing), passive health checks (response analys… - [a6-recipe-health-check](https://docs.api7.ai/apisix/ai-agent-skills/a6-recipe-health-check.md): Recipe skill for configuring upstream health checks using the a6 CLI. Covers active health checks (HTTP probing), passive health checks (response analys… #### a6-recipe-mtls Recipe skill for configuring mutual TLS (mTLS) using the a6 CLI. Covers SSL certificate management, upstream mTLS to backend services, client certificat… - [a6-recipe-mtls](https://docs.api7.ai/apisix/ai-agent-skills/a6-recipe-mtls.md): Recipe skill for configuring mutual TLS (mTLS) using the a6 CLI. Covers SSL certificate management, upstream mTLS to backend services, client certificat… #### a6-recipe-multi-tenant Recipe skill for implementing tenant-aware policies on a shared APISIX gateway using the a6 CLI. Covers shared policies through Consumer Groups, host/pa… - [Build Tenant-Aware Policies on a Shared Gateway](https://docs.api7.ai/apisix/ai-agent-skills/a6-recipe-multi-tenant.md): Recipe skill for implementing tenant-aware policies on a shared APISIX gateway using the a6 CLI. Covers shared policies through Consumer Groups, host/pa… #### a6-shared Core skill for working with the a6 CLI — the Apache APISIX command-line tool. Provides project conventions, command patterns, architecture overview, and… - [a6 Shared Skill](https://docs.api7.ai/apisix/ai-agent-skills/a6-shared.md): Core skill for working with the a6 CLI — the Apache APISIX command-line tool. Provides project conventions, command patterns, architecture overview, and… ### documentation Explore the comprehensive Apache APISIX documentation, covering setup, guides, plugins, and key concepts for effective API management and gateway functionality. - [Apache APISIX Documentation](https://docs.api7.ai/apisix/documentation.md): Explore the comprehensive Apache APISIX documentation, covering setup, guides, plugins, and key concepts for effective API management and gateway functionality. ### getting-started Learn how to quickly install and set up Apache APISIX, a dynamic and high-performance API gateway, to streamline your API lifecycle management. - [Get APISIX](https://docs.api7.ai/apisix/getting-started.md): Learn how to quickly install and set up Apache APISIX, a dynamic and high-performance API gateway, to streamline your API lifecycle management. #### configure-routes Learn how to define routes in Apache APISIX to manage traffic, match client requests, and forward them to upstream services effectively. - [Configure Routes](https://docs.api7.ai/apisix/getting-started/configure-routes.md): Learn how to define routes in Apache APISIX to manage traffic, match client requests, and forward them to upstream services effectively. #### key-authentication Explore how to configure key authentication in Apache APISIX, allowing secure access to your APIs by managing consumer credentials effectively. - [Key Authentication](https://docs.api7.ai/apisix/getting-started/key-authentication.md): Explore how to configure key authentication in Apache APISIX, allowing secure access to your APIs by managing consumer credentials effectively. #### load-balancing Understand how to implement load balancing in Apache APISIX, utilizing various algorithms to distribute incoming requests across multiple upstream services. - [Load Balancing](https://docs.api7.ai/apisix/getting-started/load-balancing.md): Understand how to implement load balancing in Apache APISIX, utilizing various algorithms to distribute incoming requests across multiple upstream services. #### management-options Compare the Admin API, ADC, a6 CLI, and APISIX-MCP to choose an APISIX resource management workflow for interactive use or automation. - [Management Options](https://docs.api7.ai/apisix/getting-started/management-options.md): Compare the Admin API, ADC, a6 CLI, and APISIX-MCP to choose an APISIX resource management workflow for interactive use or automation. #### rate-limiting Implement rate limiting in Apache APISIX to control traffic flow, protect your APIs from misuse, and ensure fair usage by setting request limits. - [Rate Limiting](https://docs.api7.ai/apisix/getting-started/rate-limiting.md): Implement rate limiting in Apache APISIX to control traffic flow, protect your APIs from misuse, and ensure fair usage by setting request limits. ### how-apisix-works Understand how Apache APISIX handles a request end-to-end, from configuration store to upstream selection, and how its router and hot-reload design let it scale to thousands of routes without a restart. - [How APISIX Works](https://docs.api7.ai/apisix/how-apisix-works.md): Understand how Apache APISIX handles a request end-to-end, from configuration store to upstream selection, and how its router and hot-reload design let it scale to thousands of routes without a restart. ### how-to-guide #### ai-gateway - [Configure Prompt Decorators](https://docs.api7.ai/apisix/how-to-guide/ai-gateway/configure-prompt-decorators.md): Learn how to configure APISIX prompt decorators to prepend and append instructions to OpenAI Chat Completions and Responses API requests at the gateway. - [Implement Prompt Guardrails](https://docs.api7.ai/apisix/how-to-guide/ai-gateway/implement-prompt-guardrails.md): Discover how to implement prompt guardrails in Apache APISIX to protect user privacy and discourage unintended model behaviors when using large language models (LLMs). - [Pre-Define Prompt Templates](https://docs.api7.ai/apisix/how-to-guide/ai-gateway/pre-define-prompt-templates.md): Learn how to configure reusable Chat Completions and Responses API prompt templates in APISIX using client-supplied values, OpenAI web search, and tool use. - [Proxy Amazon Bedrock Requests](https://docs.api7.ai/apisix/how-to-guide/ai-gateway/proxy-amazon-bedrock-requests.md): Learn how to configure APISIX to authenticate with Amazon Bedrock and proxy Converse and ConverseStream requests with the ai-proxy plugin. - [Proxy Anthropic Requests](https://docs.api7.ai/apisix/how-to-guide/ai-gateway/proxy-anthropic-requests.md): Configure Apache APISIX to proxy OpenAI-compatible and native Anthropic Messages requests to Claude models with the ai-proxy plugin. - [Proxy Azure OpenAI Requests](https://docs.api7.ai/apisix/how-to-guide/ai-gateway/proxy-azure-openai-requests.md): Learn how to configure APISIX to authenticate with Azure OpenAI and proxy streaming and non-streaming Chat Completions and Responses API requests. - [Proxy Gemini Requests](https://docs.api7.ai/apisix/how-to-guide/ai-gateway/proxy-gemini-requests.md): Learn how to configure Apache APISIX to proxy requests to Google Gemini with the ai-proxy plugin, enabling access to Gemini models via an OpenAI-compatible API without specifying a custom endpoint. - [Proxy OpenAI Requests](https://docs.api7.ai/apisix/how-to-guide/ai-gateway/proxy-openai-requests.md): Learn how to configure APISIX to authenticate with OpenAI and proxy Chat Completions, Responses API, and Embeddings requests with the ai-proxy plugin. - [Proxy OpenRouter Requests](https://docs.api7.ai/apisix/how-to-guide/ai-gateway/proxy-openrouter-requests.md): Learn how to configure Apache APISIX to proxy requests to OpenRouter with the ai-proxy plugin, enabling access to many model providers via an OpenAI-compatible API without specifying a custom endpoint. - [Proxy Vertex AI Requests](https://docs.api7.ai/apisix/how-to-guide/ai-gateway/proxy-vertex-ai-requests.md): Learn how to configure Apache APISIX to proxy requests to Google Vertex AI with the ai-proxy plugin, enabling access to Gemini models through an OpenAI-compatible API. #### authentication - [Implement Basic Authentication](https://docs.api7.ai/apisix/how-to-guide/authentication/implement-basic-auth.md): Learn how to set up basic authentication in Apache APISIX, allowing clients to authenticate using a username and password combination securely. - [Implement HMAC Authentication](https://docs.api7.ai/apisix/how-to-guide/authentication/implement-hmac-auth.md): Explore the process of setting up HMAC authentication in Apache APISIX, ensuring secure API requests through cryptographic signatures using shared secret keys. - [Implement JWT Authentication](https://docs.api7.ai/apisix/how-to-guide/authentication/implement-jwt-auth.md): Understand how to implement JWT authentication in Apache APISIX, allowing secure and stateless authentication of API clients using JSON Web Tokens. - [Implement Key Authentication](https://docs.api7.ai/apisix/how-to-guide/authentication/implement-key-auth.md): Learn how to set up key authentication in Apache APISIX, allowing you to issue unique API keys to consumers for effective access control to your APIs. - [Secure OIDC with PAR and DPoP](https://docs.api7.ai/apisix/how-to-guide/authentication/secure-oidc-with-par-and-dpop.md): Configure APISIX and Keycloak to secure an OIDC authorization code flow with PAR, PKCE, DPoP-bound tokens, and private-key JWT authentication. - [Secure WebSocket Traffic](https://docs.api7.ai/apisix/how-to-guide/authentication/secure-websocket-traffic.md): Learn how to secure WebSocket traffic in Apache APISIX, implementing authentication mechanisms during the initial handshake to protect WebSocket connections. - [Set Up SSO with Amazon Cognito](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-amazon-cognito.md): Learn how to set up single sign-on (SSO) with Amazon Cognito in Apache APISIX, facilitating secure user authentication and access management. - [Set Up SSO with Auth0](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-auth0.md): Discover how to configure Apache APISIX to use Auth0 for single sign-on (SSO), enabling secure user authentication through various identity management features. - [Set Up SSO with Microsoft Entra ID (Azure AD)](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-azure-ad.md): Understand how to set up single sign-on (SSO) in Apache APISIX using Microsoft Entra ID (Azure AD) for secure authentication via OpenID Connect. - [Set Up SSO with Google](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-google.md): Explore how to integrate Apache APISIX with Google for single sign-on (SSO) using OpenID Connect, facilitating secure user authentication for your APIs. - [Set Up SSO with Keycloak](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-keycloak.md): Learn how to integrate Apache APISIX with Keycloak to implement single sign-on (SSO) using OpenID Connect for secure authentication processes. - [Set Up SSO with Okta](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-okta.md): Discover how to configure Apache APISIX to work with Okta for single sign-on (SSO), enabling secure authentication processes using OpenID Connect. #### custom-plugins - [Create a Custom Plugin in Lua](https://docs.api7.ai/apisix/how-to-guide/custom-plugins/create-plugin-in-lua.md): Discover how to develop custom Lua plugins for Apache APISIX to extend its functionality. - [Use Wasm Plugins in APISIX](https://docs.api7.ai/apisix/how-to-guide/custom-plugins/wasm-plugins.md): Learn how to implement WebAssembly (Wasm) plugins in Apache APISIX, leveraging Proxy-Wasm specifications for enhanced capabilities. #### observability - [Log Consumer Label in Access Log](https://docs.api7.ai/apisix/how-to-guide/observability/log-consumer-label-in-access-log.md): Learn how to configure Apache APISIX to log consumer labels in access log, enhancing API management and security. - [Log with ClickHouse](https://docs.api7.ai/apisix/how-to-guide/observability/log-with-clickhouse.md): Learn how to configure Apache APISIX to log access information to ClickHouse, facilitating efficient log management and analysis. - [Log with Elasticsearch](https://docs.api7.ai/apisix/how-to-guide/observability/log-with-elasticsearch.md): Understand how to integrate Apache APISIX with Elasticsearch to collect and index logs, providing powerful search and visualization capabilities through the ELK stack. - [Monitor APISIX Metrics with Datadog](https://docs.api7.ai/apisix/how-to-guide/observability/monitor-apisix-with-datadog.md): Explore the process of integrating Datadog with Apache APISIX to monitor metrics, enabling enhanced observability and alerting capabilities. - [Monitor APISIX Metrics with Prometheus](https://docs.api7.ai/apisix/how-to-guide/observability/monitor-apisix-with-prometheus.md): Understand how to enable Prometheus in Apache APISIX for metrics collection, allowing you to monitor system performance and health effectively. - [Trace Requests with Zipkin](https://docs.api7.ai/apisix/how-to-guide/observability/trace-with-zipkin.md): Discover how to implement request tracing in Apache APISIX with Zipkin, enabling detailed monitoring of request flows and performance diagnostics. #### security - [Manage Secrets in AWS Secrets Manager](https://docs.api7.ai/apisix/how-to-guide/security/secrets-management/manage-secrets-in-aws.md): Discover how to use AWS Secrets Manager with Apache APISIX to securely store and manage sensitive credentials, ensuring automated rotation and secure access. - [Manage Secrets in GCP Secret Manager](https://docs.api7.ai/apisix/how-to-guide/security/secrets-management/manage-secrets-in-gcp-secret-manager.md): Understand how to integrate GCP Secret Manager with Apache APISIX for centralized management of secrets like API keys and passwords, with secure retrieval mechanisms. - [Manage Secrets in HashiCorp Vault](https://docs.api7.ai/apisix/how-to-guide/security/secrets-management/manage-secrets-in-hashicorp-vault.md): Learn how to integrate HashiCorp Vault with Apache APISIX for secure management of sensitive information, including API keys and passwords, and how to retrieve these secrets within your application. - [Integrate with Coraza](https://docs.api7.ai/apisix/how-to-guide/security/waf/integrate-with-coraza.md): Explore the integration of Coraza Web Application Firewall (WAF) with Apache APISIX to enhance security, providing robust protection against various cyber attacks. #### service-discovery - [Integrate with HashiCorp Consul](https://docs.api7.ai/apisix/how-to-guide/service-discovery/consul-integration.md): Learn how to set up HashiCorp Consul for service discovery and integrate it with Apache APISIX to dynamically route and load balance traffic across microservices. - [Integrate with Netflix Eureka](https://docs.api7.ai/apisix/how-to-guide/service-discovery/eureka-integration.md): Discover how to configure Netflix Eureka for service discovery and integrate it with Apache APISIX to manage service registration and routing seamlessly. - [Integrate with Kubernetes Service Discovery](https://docs.api7.ai/apisix/how-to-guide/service-discovery/kubernetes-service-discovery.md): Configure Apache APISIX to discover Kubernetes Endpoints or EndpointSlices securely and route requests to services in one or more clusters. #### traffic-management - [Implement API Versioning](https://docs.api7.ai/apisix/how-to-guide/traffic-management/api-versioning.md): Learn to implement API versioning in Apache APISIX using path, query, and header strategies for better API management. - [Manage Traffic Conditionally](https://docs.api7.ai/apisix/how-to-guide/traffic-management/conditional-traffic-management.md): Explore how to implement conditional traffic management in Apache APISIX, enabling dynamic routing and actions based on request characteristics such as headers or parameters. - [Configure Upstream Health Checks](https://docs.api7.ai/apisix/how-to-guide/traffic-management/health-check.md): Learn how to configure both active and passive health checks for upstream services in Apache APISIX, ensuring that requests are only forwarded to healthy services. - [Configure HTTP/3 QUIC Between Client and APISIX](https://docs.api7.ai/apisix/how-to-guide/traffic-management/http3-quic.md): Learn how to configure HTTP/3 connections in Apache APISIX to leverage the benefits of QUIC for improved performance and reduced latency. - [Proxy Transport Layer (L4) Traffic](https://docs.api7.ai/apisix/how-to-guide/traffic-management/proxy-transport-layer-l4-traffic.md): Explore how to configure Apache APISIX to handle transport layer (L4) TCP and UDP traffic, allowing for efficient proxying of various types of network traffic. - [Proxy WebSocket Connections](https://docs.api7.ai/apisix/how-to-guide/traffic-management/proxy-websocket.md): Discover how to enable Apache APISIX to proxy WebSocket connections, facilitating real-time, bidirectional communication between clients and servers. - [Configure Rate Limiting](https://docs.api7.ai/apisix/how-to-guide/traffic-management/rate-limiting.md): Understand how to set up rate limiting in Apache APISIX using various plugins to control access to your APIs and protect them from excessive requests. - [Configure HTTPS Between Client and APISIX](https://docs.api7.ai/apisix/how-to-guide/traffic-management/tls-and-mtls/configure-https-between-client-and-apisix.md): Discover how to configure HTTPS between clients and Apache APISIX to enhance API security. - [Configure mTLS Between APISIX and Upstream](https://docs.api7.ai/apisix/how-to-guide/traffic-management/tls-and-mtls/configure-mtls-between-apisix-and-upstream.md): Discover how to configure mutual TLS between Apache APISIX and upstream services for enhanced API security. - [Configure mTLS Between Client and APISIX](https://docs.api7.ai/apisix/how-to-guide/traffic-management/tls-and-mtls/configure-mtls-between-client-and-apisix.md): Discover how to configure mutual TLS between clients and Apache APISIX for enhanced API security. - [Configure Upstream HTTPS](https://docs.api7.ai/apisix/how-to-guide/traffic-management/tls-and-mtls/configure-upstream-https.md): Discover how to connect to upstream services on HTTPS ports in Apache APISIX, to secure communication and enhance API security. - [Implement Traffic Mirroring](https://docs.api7.ai/apisix/how-to-guide/traffic-management/traffic-mirroring.md): Understand how to set up traffic mirroring in Apache APISIX, allowing you to duplicate incoming traffic to a secondary service for testing or analysis without affecting the primary service. #### transformation - [Convert JSON to XML](https://docs.api7.ai/apisix/how-to-guide/transformation/convert-json-to-xml.md): Discover how to use the body-transformer plugin in Apache APISIX to convert JSON data to XML format and vice versa, facilitating data interchange between different systems. - [Transcode HTTP to gRPC](https://docs.api7.ai/apisix/how-to-guide/transformation/transcode-http-to-grpc.md): Learn how to use the grpc-transcode plugin in Apache APISIX to convert between RESTful HTTP requests and gRPC requests, enabling seamless integration of gRPC services. ### install #### docker Learn how to install Apache APISIX using Docker, providing a straightforward method for deploying and managing your API gateway in a containerized environment. - [Install APISIX with Docker](https://docs.api7.ai/apisix/install/docker.md): Learn how to install Apache APISIX using Docker, providing a straightforward method for deploying and managing your API gateway in a containerized environment. - [Build Your Own Docker Images](https://docs.api7.ai/apisix/install/docker/build-custom-images.md): Learn how to build custom Docker images for Apache APISIX, allowing for tailored configurations to meet your specific deployment needs. #### kubernetes - [Install APISIX on ROSA](https://docs.api7.ai/apisix/install/kubernetes/rosa.md): Follow the steps to install Apache APISIX on Red Hat OpenShift Service on AWS (ROSA), a fully managed service for deploying OpenShift clusters. ### key-concepts #### consumer-groups Understand the concept of consumer groups in Apache APISIX, which allow for the management of multiple consumers with shared configurations. - [Consumer Groups](https://docs.api7.ai/apisix/key-concepts/consumer-groups.md): Understand the concept of consumer groups in Apache APISIX, which allow for the management of multiple consumers with shared configurations. #### consumers Understand the concept of consumers in Apache APISIX, which represent users or applications that interact with the API gateway and its services. - [Consumers](https://docs.api7.ai/apisix/key-concepts/consumers.md): Understand the concept of consumers in Apache APISIX, which represent users or applications that interact with the API gateway and its services. #### credentials Understand the concept of credentials in Apache APISIX, which manage authentication configurations for consumers to enhance security and access control. - [Credentials](https://docs.api7.ai/apisix/key-concepts/credentials.md): Understand the concept of credentials in Apache APISIX, which manage authentication configurations for consumers to enhance security and access control. #### plugin-configs Understand the concept of plugin configs in Apache APISIX, which centralize plugin configurations to enhance efficiency in API management. - [Plugin Configs](https://docs.api7.ai/apisix/key-concepts/plugin-configs.md): Understand the concept of plugin configs in Apache APISIX, which centralize plugin configurations to enhance efficiency in API management. #### plugin-global-rules Understand the concept of global rules in Apache APISIX, which allow plugins to be executed on every incoming request for consistent API behavior. - [Plugin Global Rules](https://docs.api7.ai/apisix/key-concepts/plugin-global-rules.md): Understand the concept of global rules in Apache APISIX, which allow plugins to be executed on every incoming request for consistent API behavior. #### plugin-metadata Understand the concept of plugin metadata in Apache APISIX, which manages shared configurations for plugins, ensuring consistency across multiple instances. - [Plugin Metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md): Understand the concept of plugin metadata in Apache APISIX, which manages shared configurations for plugins, ensuring consistency across multiple instances. #### plugins Understand the concept of plugins in Apache APISIX, which extend functionality for traffic management, security, and observability in API operations. - [Plugins](https://docs.api7.ai/apisix/key-concepts/plugins.md): Understand the concept of plugins in Apache APISIX, which extend functionality for traffic management, security, and observability in API operations. #### protos Understand the concept of protos in Apache APISIX, which facilitates efficient data serialization and communication between services. - [Protos](https://docs.api7.ai/apisix/key-concepts/protos.md): Understand the concept of protos in Apache APISIX, which facilitates efficient data serialization and communication between services. #### routes Understand the concept of routes in Apache APISIX, which define paths to upstream services and facilitate effective traffic management. - [Routes](https://docs.api7.ai/apisix/key-concepts/routes.md): Understand the concept of routes in Apache APISIX, which define paths to upstream services and facilitate effective traffic management. #### secrets Understand the concept of secrets in Apache APISIX, which enable secure storage and management of sensitive information like API keys. - [Secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md): Understand the concept of secrets in Apache APISIX, which enable secure storage and management of sensitive information like API keys. #### services Understand the concept of services in Apache APISIX, which represent backend applications and streamline API management by reducing configuration redundancies. - [Services](https://docs.api7.ai/apisix/key-concepts/services.md): Understand the concept of services in Apache APISIX, which represent backend applications and streamline API management by reducing configuration redundancies. #### ssl-certificates Understand the concept of SSL certificates in Apache APISIX, which ensure secure communication between clients and the API gateway. - [SSL Certificates](https://docs.api7.ai/apisix/key-concepts/ssl-certificates.md): Understand the concept of SSL certificates in Apache APISIX, which ensure secure communication between clients and the API gateway. #### stream-routes Understand the concept of stream routes in Apache APISIX, which manage TCP/UDP traffic, enhancing the gateway's capabilities for various protocols. - [Stream Routes](https://docs.api7.ai/apisix/key-concepts/stream-routes.md): Understand the concept of stream routes in Apache APISIX, which manage TCP/UDP traffic, enhancing the gateway's capabilities for various protocols. #### upstreams Understand the concept of upstreams in Apache APISIX, which manage backend service addresses for efficient load balancing and service discovery. - [Upstreams](https://docs.api7.ai/apisix/key-concepts/upstreams.md): Understand the concept of upstreams in Apache APISIX, which manage backend service addresses for efficient load balancing and service discovery. ### migration #### nginx-to-apisix Follow the guide for migrating from NGINX to Apache APISIX, ensuring a smooth transition while leveraging the benefits of the API gateway. - [Migrate from NGINX to APISIX](https://docs.api7.ai/apisix/migration/nginx-to-apisix.md): Follow the guide for migrating from NGINX to Apache APISIX, ensuring a smooth transition while leveraging the benefits of the API gateway. ### networking #### port-reference Explore the default port configurations for Apache APISIX, detailing the ports used for various protocols and services within the API gateway. - [Port Reference](https://docs.api7.ai/apisix/networking/port-reference.md): Explore the default port configurations for Apache APISIX, detailing the ports used for various protocols and services within the API gateway. ### production #### deployment-modes Understand the various deployment modes of Apache APISIX, including traditional, decoupled, and standalone modes, to optimize your API gateway deployment strategy. - [Deployment Modes](https://docs.api7.ai/apisix/production/deployment-modes.md): Understand the various deployment modes of Apache APISIX, including traditional, decoupled, and standalone modes, to optimize your API gateway deployment strategy. #### performance - [Performance Testing Benchmarks](https://docs.api7.ai/apisix/production/performance/performance-testing.md): Explore methods for conducting performance testing on Apache APISIX, ensuring your API gateway can handle expected traffic loads effectively. #### recovery - [Back Up and Restore etcd](https://docs.api7.ai/apisix/production/recovery/etcd-backup-restore.md): Discover best practices for backing up and restoring etcd in Apache APISIX, ensuring data integrity and availability for your API configurations. #### scaling - [Autoscale APISIX Gateway (AWS EC2)](https://docs.api7.ai/apisix/production/scaling/autoscale-apisix-gateway-aws.md): Learn how to autoscale APISIX Gateway on AWS EC2 using Auto Scaling Group (ASG) to maintain consistent API performance under varying traffic loads. - [Autoscale APISIX Gateway (K8s)](https://docs.api7.ai/apisix/production/scaling/autoscale-apisix-gateway-k8s.md): Learn how to autoscale APISIX Gateway on Kubernetes using a Horizontal Pod Autoscaler (HPA) to maintain consistent API performance under varying traffic loads. #### security - [Admin API Key](https://docs.api7.ai/apisix/production/security/admin-api-key.md): Learn about the importance of configuring Admin API keys in Apache APISIX, ensuring secure access to Admin API endpoints and managing permissions effectively. - [Data Encryption with Keyring](https://docs.api7.ai/apisix/production/security/data-encryption-with-keyring.md): Understand the significance of data encryption in Apache APISIX and how to secure sensitive information using keyrings for enhanced protection. - [IP Restriction](https://docs.api7.ai/apisix/production/security/ip-restriction.md): Discover how to implement IP restriction in Apache APISIX to control access to resources, enhancing security by allowing only authorized IP addresses. - [Configure mTLS between APISIX and etcd](https://docs.api7.ai/apisix/production/security/mtls/configure-mtls-between-apisix-and-etcd.md): Learn how to configure mutual TLS between Apache APISIX and etcd, ensuring secure communication and authentication between these components. - [Configure mTLS between Client and APISIX Admin API](https://docs.api7.ai/apisix/production/security/mtls/configure-mtls-between-client-and-admin-api.md): Explore the steps to set up mutual TLS between clients and the APISIX Admin API, enhancing API security through two-way authentication. #### serve-static-resources Understand how to configure Apache APISIX to serve static resources efficiently, improving performance and resource management for your APIs. - [Serve Static Resources](https://docs.api7.ai/apisix/production/serve-static-resources.md): Understand how to configure Apache APISIX to serve static resources efficiently, improving performance and resource management for your APIs. #### upgrade - [Canary Deployment](https://docs.api7.ai/apisix/production/upgrade/canary-deployment.md): Learn about the canary deployment strategy in Apache APISIX, allowing for gradual rollouts of new features while minimizing risk. - [Review Changes Before Upgrade](https://docs.api7.ai/apisix/production/upgrade/review-changes-before-upgrade.md): Review APISIX behavior changes that can affect existing routes, plugins, metrics, and authentication before you upgrade into this version. ### troubleshooting #### debug-mode Enable and configure debug mode in Apache APISIX to inspect runtime behavior, log details, and troubleshoot issues effectively. - [Use Debug Mode](https://docs.api7.ai/apisix/troubleshooting/debug-mode.md): Enable and configure debug mode in Apache APISIX to inspect runtime behavior, log details, and troubleshoot issues effectively. #### set-breakpoints Learn how to use the APISIX inspect plugin to capture local variables and function-scoped captured values at any line of Lua source in a running worker, without restarting APISIX or modifying source code. - [Set Breakpoints](https://docs.api7.ai/apisix/troubleshooting/set-breakpoints.md): Learn how to use the APISIX inspect plugin to capture local variables and function-scoped captured values at any line of Lua source in a running worker, without restarting APISIX or modifying source code. ## hub Explore plugin documentation for Apache APISIX and API7 Gateway, including configuration references and examples for gateway and Ingress Controller workflows. - [Welcome to API Gateway Plugin Hub](https://docs.api7.ai/hub.md): Explore plugin documentation for Apache APISIX and API7 Gateway, including configuration references and examples for gateway and Ingress Controller workflows. ### acl The acl plugin controls access to upstream resources by verifying if the user is on the access control lists, enhancing API management. - [acl](https://docs.api7.ai/hub/acl.md): The acl plugin controls access to upstream resources by verifying if the user is on the access control lists, enhancing API management. #### configuration Parameters - [ACL](https://docs.api7.ai/hub/acl/configuration.md): Parameters ### ai-aliyun-content-moderation The ai-aliyun-content-moderation plugin uses Aliyun to evaluate selected request roles and LLM responses, including streaming responses, against a risk threshold. - [ai-aliyun-content-moderation](https://docs.api7.ai/hub/ai-aliyun-content-moderation.md): The ai-aliyun-content-moderation plugin uses Aliyun to evaluate selected request roles and LLM responses, including streaming responses, against a risk threshold. #### configuration Parameters - [AI Aliyun Content Moderation](https://docs.api7.ai/hub/ai-aliyun-content-moderation/configuration.md): Parameters ### ai-aws-content-moderation The ai-aws-content-moderation plugin uses Amazon Comprehend to detect toxicity in selected request roles and LLM responses, including streaming responses. - [ai-aws-content-moderation](https://docs.api7.ai/hub/ai-aws-content-moderation.md): The ai-aws-content-moderation plugin uses Amazon Comprehend to detect toxicity in selected request roles and LLM responses, including streaming responses. #### configuration Parameters - [AI AWS Content Moderation](https://docs.api7.ai/hub/ai-aws-content-moderation/configuration.md): Parameters ### ai-cache The ai-cache plugin stores exact and semantically similar LLM responses in Redis, reducing response latency and repeated upstream model usage. - [ai-cache](https://docs.api7.ai/hub/ai-cache.md): The ai-cache plugin stores exact and semantically similar LLM responses in Redis, reducing response latency and repeated upstream model usage. #### configuration Parameters - [AI Cache](https://docs.api7.ai/hub/ai-cache/configuration.md): Parameters ### ai-lakera-guard The ai-lakera-guard plugin screens AI traffic through the Lakera Guard API to detect prompt injection and other unsafe content in requests and LLM responses. - [ai-lakera-guard](https://docs.api7.ai/hub/ai-lakera-guard.md): The ai-lakera-guard plugin screens AI traffic through the Lakera Guard API to detect prompt injection and other unsafe content in requests and LLM responses. #### configuration Parameters - [AI Lakera Guard](https://docs.api7.ai/hub/ai-lakera-guard/configuration.md): Parameters ### ai-prompt-decorator The ai-prompt-decorator plugin decorates user prompts to LLMs by prefixing and appending pre-engineered prompts, streamlining API operation and content generation. - [ai-prompt-decorator](https://docs.api7.ai/hub/ai-prompt-decorator.md): The ai-prompt-decorator plugin decorates user prompts to LLMs by prefixing and appending pre-engineered prompts, streamlining API operation and content generation. #### configuration Parameters - [AI Prompt Decorator](https://docs.api7.ai/hub/ai-prompt-decorator/configuration.md): Parameters ### ai-prompt-guard The ai-prompt-guard plugin safeguards prompts to LLM using allow/deny patterns, ensuring only approved inputs pass. It can check the latest message or full history. - [ai-prompt-guard](https://docs.api7.ai/hub/ai-prompt-guard.md): The ai-prompt-guard plugin safeguards prompts to LLM using allow/deny patterns, ensuring only approved inputs pass. It can check the latest message or full history. #### configuration Parameters - [AI Prompt Guard](https://docs.api7.ai/hub/ai-prompt-guard/configuration.md): Parameters ### ai-prompt-template The ai-prompt-template plugin supports pre-configured templates for user inputs to LLMs in a "fill in the blank" fashion, streamlining API management. - [ai-prompt-template](https://docs.api7.ai/hub/ai-prompt-template.md): The ai-prompt-template plugin supports pre-configured templates for user inputs to LLMs in a "fill in the blank" fashion, streamlining API management. #### configuration Parameters - [AI Prompt Template](https://docs.api7.ai/hub/ai-prompt-template/configuration.md): Parameters ### ai-proxy The ai-proxy plugin simplifies access to LLM and embedding models providers by converting plugin configurations into the required request format for OpenAI, DeepSeek, Anthropic, and other OpenAI-compatible APIs. - [ai-proxy](https://docs.api7.ai/hub/ai-proxy.md): The ai-proxy plugin simplifies access to LLM and embedding models providers by converting plugin configurations into the required request format for OpenAI, DeepSeek, Anthropic, and other OpenAI-compatible APIs. #### configuration Static Configurations - [AI Proxy](https://docs.api7.ai/hub/ai-proxy/configuration.md): Static Configurations ### ai-proxy-multi The ai-proxy-multi plugin extends the capabilities of ai-proxy with load balancing, retries, fallbacks, and health checks, simplifying the integration with OpenAI, DeepSeek, and other OpenAI-compatible APIs. - [ai-proxy-multi](https://docs.api7.ai/hub/ai-proxy-multi.md): The ai-proxy-multi plugin extends the capabilities of ai-proxy with load balancing, retries, fallbacks, and health checks, simplifying the integration with OpenAI, DeepSeek, and other OpenAI-compatible APIs. #### configuration Static Configurations - [AI Proxy Multi](https://docs.api7.ai/hub/ai-proxy-multi/configuration.md): Static Configurations ### ai-proxy-protocol-reference Understand how ai-proxy and ai-proxy-multi detect client request protocols and convert Anthropic Messages requests to OpenAI Chat Completions. - [Protocol Reference](https://docs.api7.ai/hub/ai-proxy-protocol-reference.md): Understand how ai-proxy and ai-proxy-multi detect client request protocols and convert Anthropic Messages requests to OpenAI Chat Completions. ### ai-rag The ai-rag plugin retrieves context with Azure OpenAI embeddings and Azure AI Search before an LLM request is proxied. - [ai-rag](https://docs.api7.ai/hub/ai-rag.md): The ai-rag plugin retrieves context with Azure OpenAI embeddings and Azure AI Search before an LLM request is proxied. #### configuration Plugin Parameters - [AI RAG](https://docs.api7.ai/hub/ai-rag/configuration.md): Plugin Parameters ### ai-rate-limiting The ai-rate-limiting plugin enforces token-based rate limiting for LLM service requests, preventing overuse, optimizing API consumption, and ensuring efficient resource allocation. - [ai-rate-limiting](https://docs.api7.ai/hub/ai-rate-limiting.md): The ai-rate-limiting plugin enforces token-based rate limiting for LLM service requests, preventing overuse, optimizing API consumption, and ensuring efficient resource allocation. #### configuration Parameters - [AI Rate Limiting](https://docs.api7.ai/hub/ai-rate-limiting/configuration.md): Parameters ### ai-request-rewrite The ai-request-rewrite plugin forwards client requests to LLM services for processing before sending them upstream, enabling AI-driven redaction, enrichment, and reformatting. - [ai-request-rewrite](https://docs.api7.ai/hub/ai-request-rewrite.md): The ai-request-rewrite plugin forwards client requests to LLM services for processing before sending them upstream, enabling AI-driven redaction, enrichment, and reformatting. #### configuration Parameters - [AI Request Rewrite](https://docs.api7.ai/hub/ai-request-rewrite/configuration.md): Parameters ### attach-consumer-label The attach-consumer-label plugin attaches custom consumer labels to authenticated requests, for upstream services to implement additional business logics. - [attach-consumer-label](https://docs.api7.ai/hub/attach-consumer-label.md): The attach-consumer-label plugin attaches custom consumer labels to authenticated requests, for upstream services to implement additional business logics. #### configuration Parameters - [Attach Consumer Label](https://docs.api7.ai/hub/attach-consumer-label/configuration.md): Parameters ### authz-keycloak The authz-keycloak plugin integrates with Keycloak for user authentication and authorization, enhancing API security and management. - [authz-keycloak](https://docs.api7.ai/hub/authz-keycloak.md): The authz-keycloak plugin integrates with Keycloak for user authentication and authorization, enhancing API security and management. #### configuration Parameters - [Authz Keycloak](https://docs.api7.ai/hub/authz-keycloak/configuration.md): Parameters ### aws-lambda The aws-lambda plugin simplifies APISIX integration with AWS Lambda and Amazon API gateway, supporting authentication via IAM user credentials and API keys. - [aws-lambda](https://docs.api7.ai/hub/aws-lambda.md): The aws-lambda plugin simplifies APISIX integration with AWS Lambda and Amazon API gateway, supporting authentication via IAM user credentials and API keys. #### configuration Attributes - [AWS Lambda](https://docs.api7.ai/hub/aws-lambda/configuration.md): Attributes ### basic-auth The basic-auth plugin provides basic access authentication, requiring clients to authenticate before accessing upstream resources, enhancing API security. - [basic-auth](https://docs.api7.ai/hub/basic-auth.md): The basic-auth plugin provides basic access authentication, requiring clients to authenticate before accessing upstream resources, enhancing API security. #### configuration Parameters - [Basic Auth](https://docs.api7.ai/hub/basic-auth/configuration.md): Parameters ### body-transformer The body-transformer plugin converts request and response bodies between formats, such as JSON to XML, facilitating seamless data exchange. - [body-transformer](https://docs.api7.ai/hub/body-transformer.md): The body-transformer plugin converts request and response bodies between formats, such as JSON to XML, facilitating seamless data exchange. #### configuration Parameters - [Body Transformer](https://docs.api7.ai/hub/body-transformer/configuration.md): Parameters ### chaitin-waf The chaitin-waf plugin integrates with Chaitin WAF (SafeLine) to detect and block web threats, strengthening application security and protecting user data. - [chaitin-waf](https://docs.api7.ai/hub/chaitin-waf.md): The chaitin-waf plugin integrates with Chaitin WAF (SafeLine) to detect and block web threats, strengthening application security and protecting user data. #### configuration Parameters - [Chaitin WAF](https://docs.api7.ai/hub/chaitin-waf/configuration.md): Parameters ### clickhouse-logger The clickhouse-logger plugin pushes request and response logs to ClickHouse databases in batches, allowing for customizable log formats to enhance data management. - [clickhouse-logger](https://docs.api7.ai/hub/clickhouse-logger.md): The clickhouse-logger plugin pushes request and response logs to ClickHouse databases in batches, allowing for customizable log formats to enhance data management. #### configuration Parameters - [ClickHouse Logger](https://docs.api7.ai/hub/clickhouse-logger/configuration.md): Parameters ### consumer-restriction The consumer-restriction plugin implements access controls based on consumer name, route ID, service ID, or consumer group ID, enhancing API security. - [consumer-restriction](https://docs.api7.ai/hub/consumer-restriction.md): The consumer-restriction plugin implements access controls based on consumer name, route ID, service ID, or consumer group ID, enhancing API security. #### configuration Parameters - [Consumer Restriction](https://docs.api7.ai/hub/consumer-restriction/configuration.md): Parameters ### cors The cors plugin enables cross-origin resource sharing, allowing servers to specify permitted origins and instructing browsers to load resources from those origins, enhancing API accessibility. - [cors](https://docs.api7.ai/hub/cors.md): The cors plugin enables cross-origin resource sharing, allowing servers to specify permitted origins and instructing browsers to load resources from those origins, enhancing API accessibility. #### configuration Parameters - [CORS](https://docs.api7.ai/hub/cors/configuration.md): Parameters ### data-mask The data-mask plugin removes or replaces sensitive information in request headers, bodies, and URL queries for logging purposes, enhancing data privacy and security. - [data-mask](https://docs.api7.ai/hub/data-mask.md): The data-mask plugin removes or replaces sensitive information in request headers, bodies, and URL queries for logging purposes, enhancing data privacy and security. #### configuration Parameters - [Data Mask](https://docs.api7.ai/hub/data-mask/configuration.md): Parameters ### datadog The datadog plugin integrates with Datadog, sending metrics to DogStatsD in batches to improve API monitoring and API performance tracking. - [datadog](https://docs.api7.ai/hub/datadog.md): The datadog plugin integrates with Datadog, sending metrics to DogStatsD in batches to improve API monitoring and API performance tracking. #### configuration Parameters - [Datadog](https://docs.api7.ai/hub/datadog/configuration.md): Parameters ### degraphql The degraphql plugin enables communication with upstream GraphQL services through standard HTTP requests by mapping GraphQL queries to HTTP endpoints, simplifying API integration. - [degraphql](https://docs.api7.ai/hub/degraphql.md): The degraphql plugin enables communication with upstream GraphQL services through standard HTTP requests by mapping GraphQL queries to HTTP endpoints, simplifying API integration. #### configuration Parameters - [degraphql](https://docs.api7.ai/hub/degraphql/configuration.md): Parameters ### elasticsearch-logger The elasticsearch-logger plugin pushes request and response logs in batches to Elasticsearch, allowing for customizable log formats to enhance data management. - [elasticsearch-logger](https://docs.api7.ai/hub/elasticsearch-logger.md): The elasticsearch-logger plugin pushes request and response logs in batches to Elasticsearch, allowing for customizable log formats to enhance data management. #### configuration Parameters - [Elasticsearch Logger](https://docs.api7.ai/hub/elasticsearch-logger/configuration.md): Parameters ### error-log-collect The error-log-collect plugin captures the error logs produced while processing selected requests, including lower-severity entries that the configured log level would otherwise discard, and writes them to the gateway error log for targeted debugging. - [error-log-collect Enterprise](https://docs.api7.ai/hub/error-log-collect.md): The error-log-collect plugin captures the error logs produced while processing selected requests, including lower-severity entries that the configured log level would otherwise discard, and writes them to the gateway error log for targeted debugging. #### configuration Parameters - [Error Log Collect](https://docs.api7.ai/hub/error-log-collect/configuration.md): Parameters ### error-log-logger The error-log-logger plugin pushes APISIX's error logs to TCP, Apache SkyWalking, Apache Kafka, or ClickHouse servers, in batches. You can specify the severity level of which the plugin should send the corresponding logs. - [error-log-logger](https://docs.api7.ai/hub/error-log-logger.md): The error-log-logger plugin pushes APISIX's error logs to TCP, Apache SkyWalking, Apache Kafka, or ClickHouse servers, in batches. You can specify the severity level of which the plugin should send the corresponding logs. #### configuration Parameters - [Error Log Logger](https://docs.api7.ai/hub/error-log-logger/configuration.md): Parameters ### error-page The error-page plugin customizes gateway-generated 404, 500, 502, and 503 responses without modifying responses returned by upstream services. - [error-page](https://docs.api7.ai/hub/error-page.md): The error-page plugin customizes gateway-generated 404, 500, 502, and 503 responses without modifying responses returned by upstream services. #### configuration Parameters - [Error Page](https://docs.api7.ai/hub/error-page/configuration.md): Parameters ### exit-transformer The exit-transformer plugin customizes responses generated by gateway plugins or missing routes before APISIX sends them to clients. - [exit-transformer](https://docs.api7.ai/hub/exit-transformer.md): The exit-transformer plugin customizes responses generated by gateway plugins or missing routes before APISIX sends them to clients. #### configuration Parameters - [Exit Transformer](https://docs.api7.ai/hub/exit-transformer/configuration.md): Parameters ### fault-injection The fault-injection plugin tests application resiliency by simulating controlled faults or delays, making it ideal for chaos engineering and failure condition analysis. - [fault-injection](https://docs.api7.ai/hub/fault-injection.md): The fault-injection plugin tests application resiliency by simulating controlled faults or delays, making it ideal for chaos engineering and failure condition analysis. #### configuration Parameters - [Fault Injection](https://docs.api7.ai/hub/fault-injection/configuration.md): Parameters ### forward-auth The forward-auth plugin integrates with external authorization services, enhancing API security and access control. - [forward-auth](https://docs.api7.ai/hub/forward-auth.md): The forward-auth plugin integrates with external authorization services, enhancing API security and access control. #### configuration Parameters - [Forward Auth](https://docs.api7.ai/hub/forward-auth/configuration.md): Parameters ### google-cloud-logging The google-cloud-logging plugin pushes request and response logs in batches to Google Cloud Logging Service and supports the customization of log formats. - [google-cloud-logging](https://docs.api7.ai/hub/google-cloud-logging.md): The google-cloud-logging plugin pushes request and response logs in batches to Google Cloud Logging Service and supports the customization of log formats. #### configuration Parameters - [Google Cloud Logging](https://docs.api7.ai/hub/google-cloud-logging/configuration.md): Parameters ### graphql-limit-count The graphql-limit-count plugin uses fixed windows to limit accumulated GraphQL document cost, with selection depth as the default measure. - [graphql-limit-count](https://docs.api7.ai/hub/graphql-limit-count.md): The graphql-limit-count plugin uses fixed windows to limit accumulated GraphQL document cost, with selection depth as the default measure. #### configuration Parameters - [GraphQL Limit Count](https://docs.api7.ai/hub/graphql-limit-count/configuration.md): Parameters ### graphql-proxy-cache The graphql-proxy-cache plugin enables caching of responses for GraphQL queries, improving API performance. - [graphql-proxy-cache](https://docs.api7.ai/hub/graphql-proxy-cache.md): The graphql-proxy-cache plugin enables caching of responses for GraphQL queries, improving API performance. #### configuration Static Configurations - [GraphQL Proxy Cache](https://docs.api7.ai/hub/graphql-proxy-cache/configuration.md): Static Configurations ### grpc-transcode The grpc-transcode plugin converts between HTTP and gRPC requests and responses, facilitating seamless communication between different API protocols. - [grpc-transcode](https://docs.api7.ai/hub/grpc-transcode.md): The grpc-transcode plugin converts between HTTP and gRPC requests and responses, facilitating seamless communication between different API protocols. #### configuration Parameters - [gRPC Transcode](https://docs.api7.ai/hub/grpc-transcode/configuration.md): Parameters ### grpc-web The grpc-web plugin enables the gateway to handle gRPC-Web requests from browsers and JavaScript clients by translating them into standard gRPC calls and forwarding them to upstream gRPC services. - [grpc-web](https://docs.api7.ai/hub/grpc-web.md): The grpc-web plugin enables the gateway to handle gRPC-Web requests from browsers and JavaScript clients by translating them into standard gRPC calls and forwarding them to upstream gRPC services. #### configuration Parameters - [gRPC Web](https://docs.api7.ai/hub/grpc-web/configuration.md): Parameters ### hmac-auth The hmac-auth plugin supports HMAC authentication to ensure request integrity, preventing modifications during transmission and enhancing API security. - [hmac-auth](https://docs.api7.ai/hub/hmac-auth.md): The hmac-auth plugin supports HMAC authentication to ensure request integrity, preventing modifications during transmission and enhancing API security. #### configuration Parameters - [HMAC Auth](https://docs.api7.ai/hub/hmac-auth/configuration.md): Parameters ### http-logger The http-logger plugin pushes request and response logs as JSON objects to HTTP(S) servers in batches, allowing for customizable log formats to enhance data management. - [http-logger](https://docs.api7.ai/hub/http-logger.md): The http-logger plugin pushes request and response logs as JSON objects to HTTP(S) servers in batches, allowing for customizable log formats to enhance data management. #### configuration Parameters - [HTTP Logger](https://docs.api7.ai/hub/http-logger/configuration.md): Parameters ### ip-restriction The ip-restriction plugin restricts access to upstream resources based on an IP address whitelist or blacklist, improving API security. - [ip-restriction](https://docs.api7.ai/hub/ip-restriction.md): The ip-restriction plugin restricts access to upstream resources based on an IP address whitelist or blacklist, improving API security. #### configuration Parameters - [IP Restriction](https://docs.api7.ai/hub/ip-restriction/configuration.md): Parameters ### jwe-decrypt The jwe-decrypt plugin decrypts its supported five-part compact token format and forwards the plaintext in a configured request header. - [jwe-decrypt](https://docs.api7.ai/hub/jwe-decrypt.md): The jwe-decrypt plugin decrypts its supported five-part compact token format and forwards the plaintext in a configured request header. #### configuration Parameters - [JWE Decrypt](https://docs.api7.ai/hub/jwe-decrypt/configuration.md): Parameters ### jwt-auth The jwt-auth plugin supports the use of JSON Web Token (JWT) for client authentication before accessing upstream resources, enhancing API security measures. - [jwt-auth](https://docs.api7.ai/hub/jwt-auth.md): The jwt-auth plugin supports the use of JSON Web Token (JWT) for client authentication before accessing upstream resources, enhancing API security measures. #### configuration Parameters - [JWT Auth](https://docs.api7.ai/hub/jwt-auth/configuration.md): Parameters ### kafka-logger The kafka-logger plugin pushes request and response logs as JSON objects to Apache Kafka clusters in batches, allowing for customizable log formats to enhance data management. - [kafka-logger](https://docs.api7.ai/hub/kafka-logger.md): The kafka-logger plugin pushes request and response logs as JSON objects to Apache Kafka clusters in batches, allowing for customizable log formats to enhance data management. #### configuration Parameters - [Kafka Logger](https://docs.api7.ai/hub/kafka-logger/configuration.md): Parameters ### key-auth The key-auth plugin allows clients to authenticate using an authentication key before accessing upstream resources, enhancing API security measures. - [key-auth](https://docs.api7.ai/hub/key-auth.md): The key-auth plugin allows clients to authenticate using an authentication key before accessing upstream resources, enhancing API security measures. #### configuration Parameters - [Key Auth](https://docs.api7.ai/hub/key-auth/configuration.md): Parameters ### ldap-auth-advanced The ldap-auth-advanced plugin authenticates clients against an LDAP directory and maps the authenticated user onto a consumer, so directory identities can be used with per-consumer plugins, rate limits, and analytics. - [ldap-auth-advanced](https://docs.api7.ai/hub/ldap-auth-advanced.md): The ldap-auth-advanced plugin authenticates clients against an LDAP directory and maps the authenticated user onto a consumer, so directory identities can be used with per-consumer plugins, rate limits, and analytics. #### configuration Parameters - [LDAP Auth Advanced](https://docs.api7.ai/hub/ldap-auth-advanced/configuration.md): Parameters ### limit-conn The limit-conn plugin restricts the rate of requests by managing concurrent connections. Requests exceeding the threshold may be delayed or rejected, ensuring controlled API usage and preventing overload. - [limit-conn](https://docs.api7.ai/hub/limit-conn.md): The limit-conn plugin restricts the rate of requests by managing concurrent connections. Requests exceeding the threshold may be delayed or rejected, ensuring controlled API usage and preventing overload. #### configuration Parameters - [Limit Conn](https://docs.api7.ai/hub/limit-conn/configuration.md): Parameters ### limit-count The limit-count plugin enforces API rate limiting with a fixed window algorithm, restricting requests within a time interval. Requests over the quota are rejected. - [limit-count](https://docs.api7.ai/hub/limit-count.md): The limit-count plugin enforces API rate limiting with a fixed window algorithm, restricting requests within a time interval. Requests over the quota are rejected. #### configuration Parameters - [Limit Count](https://docs.api7.ai/hub/limit-count/configuration.md): Parameters ### limit-count-advanced The limit-count-advanced plugin enforces API rate limiting with a fixed window or sliding window algorithm, restricting requests within a time window. Requests over the quota are rejected. - [limit-count-advanced Enterprise](https://docs.api7.ai/hub/limit-count-advanced.md): The limit-count-advanced plugin enforces API rate limiting with a fixed window or sliding window algorithm, restricting requests within a time window. Requests over the quota are rejected. #### configuration Parameters - [Limit Count Advanced](https://docs.api7.ai/hub/limit-count-advanced/configuration.md): Parameters ### limit-req The limit-req plugin enforces API rate limiting with a leaky bucket algorithm to rate limit requests, enabling effective throttling to manage traffic flow. - [limit-req](https://docs.api7.ai/hub/limit-req.md): The limit-req plugin enforces API rate limiting with a leaky bucket algorithm to rate limit requests, enabling effective throttling to manage traffic flow. #### configuration Parameters - [Limit Req](https://docs.api7.ai/hub/limit-req/configuration.md): Parameters ### loki-logger The loki-logger plugin sends request and response logs as JSON objects to Grafana Loki in batches via the Loki HTTP API, allowing for customizable log formats to enhance data management. - [loki-logger](https://docs.api7.ai/hub/loki-logger.md): The loki-logger plugin sends request and response logs as JSON objects to Grafana Loki in batches via the Loki HTTP API, allowing for customizable log formats to enhance data management. #### configuration Parameters - [Loki Logger](https://docs.api7.ai/hub/loki-logger/configuration.md): Parameters ### mcp-tools-acl The mcp-tools-acl plugin provides per-consumer access control for MCP tool calls on routes powered by openapi-to-mcp, supporting rule-based allowlist and denylist modes with optional expression conditions. - [mcp-tools-acl Enterprise](https://docs.api7.ai/hub/mcp-tools-acl.md): The mcp-tools-acl plugin provides per-consumer access control for MCP tool calls on routes powered by openapi-to-mcp, supporting rule-based allowlist and denylist modes with optional expression conditions. #### configuration Parameters - [MCP Tools ACL](https://docs.api7.ai/hub/mcp-tools-acl/configuration.md): Parameters ### mocking The mocking plugin simulates API responses without forwarding requests to upstream services, offering customization of status codes, response bodies, headers, and more for API testing and development. - [mocking](https://docs.api7.ai/hub/mocking.md): The mocking plugin simulates API responses without forwarding requests to upstream services, offering customization of status codes, response bodies, headers, and more for API testing and development. #### configuration Parameters - [Mocking](https://docs.api7.ai/hub/mocking/configuration.md): Parameters ### mqtt-proxy The mqtt-proxy plugin supports proxying and load balancing MQTT requests to MQTT servers, enhancing API operation and management. - [mqtt-proxy](https://docs.api7.ai/hub/mqtt-proxy.md): The mqtt-proxy plugin supports proxying and load balancing MQTT requests to MQTT servers, enhancing API operation and management. #### configuration Parameters - [MQTT Proxy](https://docs.api7.ai/hub/mqtt-proxy/configuration.md): Parameters ### multi-auth The multi-auth plugin enables consumers using diverse authentication methods to share the same route or service, streamlining API lifecycle management. - [multi-auth](https://docs.api7.ai/hub/multi-auth.md): The multi-auth plugin enables consumers using diverse authentication methods to share the same route or service, streamlining API lifecycle management. #### configuration Parameters - [Multi Auth](https://docs.api7.ai/hub/multi-auth/configuration.md): Parameters ### oas-validator The oas-validator plugin checks incoming HTTP requests against an OpenAPI specification before they are forwarded to upstream services. - [oas-validator](https://docs.api7.ai/hub/oas-validator.md): The oas-validator plugin checks incoming HTTP requests against an OpenAPI specification before they are forwarded to upstream services. #### configuration Parameters - [OAS Validator](https://docs.api7.ai/hub/oas-validator/configuration.md): Parameters ### opa The OPA plugin integrates with Open Policy Agent, enabling unified policy definition and enforcement for authorization in API operations. - [OPA](https://docs.api7.ai/hub/opa.md): The OPA plugin integrates with Open Policy Agent, enabling unified policy definition and enforcement for authorization in API operations. #### configuration Parameters - [OPA](https://docs.api7.ai/hub/opa/configuration.md): Parameters ### openapi-to-mcp The openapi-to-mcp plugin lets API7 expose OpenAPI services through MCP, proxy requests with custom headers, and support real-time SSE streaming. - [openapi-to-mcp Enterprise](https://docs.api7.ai/hub/openapi-to-mcp.md): The openapi-to-mcp plugin lets API7 expose OpenAPI services through MCP, proxy requests with custom headers, and support real-time SSE streaming. #### configuration Static Configurations - [OpenAPI to MCP](https://docs.api7.ai/hub/openapi-to-mcp/configuration.md): Static Configurations ### openid-connect The openid-connect plugin integrates with OIDC providers like Keycloak and Auth0, simplifying user authentication in API management. - [openid-connect](https://docs.api7.ai/hub/openid-connect.md): The openid-connect plugin integrates with OIDC providers like Keycloak and Auth0, simplifying user authentication in API management. #### configuration Parameters - [OpenID Connect](https://docs.api7.ai/hub/openid-connect/configuration.md): Parameters ### opentelemetry The opentelemetry plugin instruments the API gateway, sending traces to the OpenTelemetry collector for monitoring API operations per OpenTelemetry specs. - [OpenTelemetry](https://docs.api7.ai/hub/opentelemetry.md): The opentelemetry plugin instruments the API gateway, sending traces to the OpenTelemetry collector for monitoring API operations per OpenTelemetry specs. #### configuration Parameters - [OpenTelemetry](https://docs.api7.ai/hub/opentelemetry/configuration.md): Parameters ### prometheus The Prometheus plugin integrates with Prometheus for metric collection and continuous monitoring, enhancing API observability. - [Prometheus](https://docs.api7.ai/hub/prometheus.md): The Prometheus plugin integrates with Prometheus for metric collection and continuous monitoring, enhancing API observability. #### configuration Static Configurations - [Prometheus](https://docs.api7.ai/hub/prometheus/configuration.md): Static Configurations ### proxy-buffering The proxy-buffering plugin dynamically disables NGINX proxy buffering, optimizing performance with SSE and other streaming upstream services in API gateway. - [proxy-buffering](https://docs.api7.ai/hub/proxy-buffering.md): The proxy-buffering plugin dynamically disables NGINX proxy buffering, optimizing performance with SSE and other streaming upstream services in API gateway. #### configuration Parameters - [Proxy Buffering](https://docs.api7.ai/hub/proxy-buffering/configuration.md): Parameters ### proxy-cache The proxy-cache plugin caches responses based on keys, supporting disk and memory caching for GET, POST, and HEAD requests, enhancing API performance. - [proxy-cache](https://docs.api7.ai/hub/proxy-cache.md): The proxy-cache plugin caches responses based on keys, supporting disk and memory caching for GET, POST, and HEAD requests, enhancing API performance. #### configuration Static Configurations - [Proxy Cache](https://docs.api7.ai/hub/proxy-cache/configuration.md): Static Configurations ### proxy-mirror The proxy-mirror plugin duplicates ingress traffic to API gateway, forwarding it to a designated upstream while keeping regular services uninterrupted. - [proxy-mirror](https://docs.api7.ai/hub/proxy-mirror.md): The proxy-mirror plugin duplicates ingress traffic to API gateway, forwarding it to a designated upstream while keeping regular services uninterrupted. #### configuration Static Configurations - [Proxy Mirror](https://docs.api7.ai/hub/proxy-mirror/configuration.md): Static Configurations ### proxy-rewrite The proxy-rewrite plugin offers flexible options to rewrite requests that API gateway forwards to upstream services, enhancing API management. - [proxy-rewrite](https://docs.api7.ai/hub/proxy-rewrite.md): The proxy-rewrite plugin offers flexible options to rewrite requests that API gateway forwards to upstream services, enhancing API management. #### configuration Parameters - [Proxy Rewrite](https://docs.api7.ai/hub/proxy-rewrite/configuration.md): Parameters ### public-api The public-api plugin exposes internal API endpoints, allowing external access while maintaining control over API management and security. - [public-api](https://docs.api7.ai/hub/public-api.md): The public-api plugin exposes internal API endpoints, allowing external access while maintaining control over API management and security. #### configuration Parameters - [Public API](https://docs.api7.ai/hub/public-api/configuration.md): Parameters ### real-ip The real-ip plugin enables the API gateway to fetch the client's real IP using the IP address from the HTTP header or query string, improving data quality. - [real-ip](https://docs.api7.ai/hub/real-ip.md): The real-ip plugin enables the API gateway to fetch the client's real IP using the IP address from the HTTP header or query string, improving data quality. #### configuration Parameters - [Real IP](https://docs.api7.ai/hub/real-ip/configuration.md): Parameters ### request-id The request-id plugin adds a unique ID to each request proxied through the API gateway, facilitating effective tracking of API requests for better API management. - [request-id](https://docs.api7.ai/hub/request-id.md): The request-id plugin adds a unique ID to each request proxied through the API gateway, facilitating effective tracking of API requests for better API management. #### configuration Parameters - [Request ID](https://docs.api7.ai/hub/request-id/configuration.md): Parameters ### request-validation The request-validation plugin checks requests for compliance before forwarding them to upstream services, enhancing security in API operations. - [request-validation](https://docs.api7.ai/hub/request-validation.md): The request-validation plugin checks requests for compliance before forwarding them to upstream services, enhancing security in API operations. #### configuration Parameters - [Request Validation](https://docs.api7.ai/hub/request-validation/configuration.md): Parameters ### response-rewrite The response-rewrite plugin allows rewriting of responses from API gateway and upstream services, providing flexibility in API responses. - [response-rewrite](https://docs.api7.ai/hub/response-rewrite.md): The response-rewrite plugin allows rewriting of responses from API gateway and upstream services, providing flexibility in API responses. #### configuration Parameters - [Response Rewrite](https://docs.api7.ai/hub/response-rewrite/configuration.md): Parameters ### rocketmq-logger The rocketmq-logger plugin pushes request and response logs as JSON objects to RocketMQ clusters in batches, allowing for customizable log formats to enhance data management. - [rocketmq-logger](https://docs.api7.ai/hub/rocketmq-logger.md): The rocketmq-logger plugin pushes request and response logs as JSON objects to RocketMQ clusters in batches, allowing for customizable log formats to enhance data management. #### configuration Parameters - [RocketMQ Logger](https://docs.api7.ai/hub/rocketmq-logger/configuration.md): Parameters ### saml-auth The saml-auth plugin enables user authentication via SAML 2.0 in the API gateway by interacting with identity providers (IdP), enhancing API security. - [saml-auth](https://docs.api7.ai/hub/saml-auth.md): The saml-auth plugin enables user authentication via SAML 2.0 in the API gateway by interacting with identity providers (IdP), enhancing API security. #### configuration Parameters - [SAML Auth](https://docs.api7.ai/hub/saml-auth/configuration.md): Parameters ### serverless-functions The serverless function plugins (pre-function and post-function) allow execution of user-defined logic at the start or end of specified execution phases in API gateway. - [Serverless Functions](https://docs.api7.ai/hub/serverless-functions.md): The serverless function plugins (pre-function and post-function) allow execution of user-defined logic at the start or end of specified execution phases in API gateway. #### configuration Parameters - [Serverless Functions](https://docs.api7.ai/hub/serverless-functions/configuration.md): Parameters ### skywalking The skywalking plugin integrates with Apache SkyWalking for effective request tracing, enhancing API observability. - [SkyWalking](https://docs.api7.ai/hub/skywalking.md): The skywalking plugin integrates with Apache SkyWalking for effective request tracing, enhancing API observability. #### configuration Static Configurations - [skywalking](https://docs.api7.ai/hub/skywalking/configuration.md): Static Configurations ### skywalking-logger The skywalking-logger pushes request and response logs as JSON objects to SkyWalking OAP server in batches, allowing for customizable log formats to enhance data management. - [skywalking-logger](https://docs.api7.ai/hub/skywalking-logger.md): The skywalking-logger pushes request and response logs as JSON objects to SkyWalking OAP server in batches, allowing for customizable log formats to enhance data management. #### configuration Parameters - [SkyWalking Logger](https://docs.api7.ai/hub/skywalking-logger/configuration.md): Parameters ### soap The soap plugin simplifies transformation between RESTful HTTP requests and SOAP requests, including their corresponding responses, for better API interoperability. - [soap Enterprise](https://docs.api7.ai/hub/soap.md): The soap plugin simplifies transformation between RESTful HTTP requests and SOAP requests, including their corresponding responses, for better API interoperability. #### configuration Static Configurations - [SOAP](https://docs.api7.ai/hub/soap/configuration.md): Static Configurations ### splunk-hec-logging The splunk-hec-logging plugin serializes request and response context information to Splunk Event Data format and push to your Splunk HTTP Event Collector (HEC) in batches, allowing for customizable log formats to enhance data management. - [splunk-hec-logging](https://docs.api7.ai/hub/splunk-hec-logging.md): The splunk-hec-logging plugin serializes request and response context information to Splunk Event Data format and push to your Splunk HTTP Event Collector (HEC) in batches, allowing for customizable log formats to enhance data management. #### configuration Parameters - [splunk-hec-logging](https://docs.api7.ai/hub/splunk-hec-logging/configuration.md): Parameters ### syslog The syslog plugin pushes request and response logs as JSON objects to syslog servers in batches, allowing for customizable log formats to enhance data management. - [syslog](https://docs.api7.ai/hub/syslog.md): The syslog plugin pushes request and response logs as JSON objects to syslog servers in batches, allowing for customizable log formats to enhance data management. #### configuration Parameters - [syslog](https://docs.api7.ai/hub/syslog/configuration.md): Parameters ### traffic-label The traffic-label plugin evaluates request expressions and applies weighted request-header changes for conditional traffic management. - [traffic-label](https://docs.api7.ai/hub/traffic-label.md): The traffic-label plugin evaluates request expressions and applies weighted request-header changes for conditional traffic management. #### configuration Parameters - [Traffic Label](https://docs.api7.ai/hub/traffic-label/configuration.md): Parameters ### traffic-split The traffic-split plugin directs traffic to multiple upstream services based on conditions or weights, providing a flexible approach for API release strategies and traffic management. - [traffic-split](https://docs.api7.ai/hub/traffic-split.md): The traffic-split plugin directs traffic to multiple upstream services based on conditions or weights, providing a flexible approach for API release strategies and traffic management. #### configuration Parameters - [Traffic Split](https://docs.api7.ai/hub/traffic-split/configuration.md): Parameters ### ua-restriction The ua-restriction plugin restricts access to upstream resources using an allowlist or denylist of user agents, preventing overload from web crawlers and enhancing API security. - [ua-restriction](https://docs.api7.ai/hub/ua-restriction.md): The ua-restriction plugin restricts access to upstream resources using an allowlist or denylist of user agents, preventing overload from web crawlers and enhancing API security. #### configuration Parameters - [UA Restriction](https://docs.api7.ai/hub/ua-restriction/configuration.md): Parameters ### workflow The workflow plugin enables conditional execution of user-defined actions on client traffic based on specific rules, allowing granular API traffic management. - [workflow](https://docs.api7.ai/hub/workflow.md): The workflow plugin enables conditional execution of user-defined actions on client traffic based on specific rules, allowing granular API traffic management. #### configuration Parameters - [Workflow](https://docs.api7.ai/hub/workflow/configuration.md): Parameters ### zipkin The zipkin plugin instruments the API gateway to send traces to Zipkin or compatible collectors like Jaeger and Apache SkyWalking, enhancing request tracing capabilities. - [zipkin](https://docs.api7.ai/hub/zipkin.md): The zipkin plugin instruments the API gateway to send traces to Zipkin or compatible collectors like Jaeger and Apache SkyWalking, enhancing request tracing capabilities. #### configuration Static Configurations - [Zipkin](https://docs.api7.ai/hub/zipkin/configuration.md): Static Configurations ## ingress-controller ### apply-plugins-to-l4-routes Learn how to attach APISIX stream plugins to Gateway API TCPRoute, UDPRoute, and TLSRoute resources by configuring and verifying an L4RoutePolicy. - [Apply Plugins to L4 Routes](https://docs.api7.ai/ingress-controller/apply-plugins-to-l4-routes.md): Learn how to attach APISIX stream plugins to Gateway API TCPRoute, UDPRoute, and TLSRoute resources by configuring and verifying an L4RoutePolicy. ### canary-releases Plan canary releases with APISIX or API7 Ingress Controller by coordinating workloads, weighted routes, analysis, promotion, and rollback. - [Canary Releases](https://docs.api7.ai/ingress-controller/canary-releases.md): Plan canary releases with APISIX or API7 Ingress Controller by coordinating workloads, weighted routes, analysis, promotion, and rollback. #### argo-rollouts Configure Argo Rollouts to shift APISIX or API7 Gateway traffic with HTTPRoute or ApisixRoute, then promote, analyze, and abort releases. - [Canary Releases with Argo Rollouts](https://docs.api7.ai/ingress-controller/canary-releases/argo-rollouts.md): Configure Argo Rollouts to shift APISIX or API7 Gateway traffic with HTTPRoute or ApisixRoute, then promote, analyze, and abort releases. #### flagger Use Flagger and APISIX metrics to automate canary analysis, promotion, and rollback through Gateway API HTTPRoute or ApisixRoute resources. - [Automated Canary Releases with Flagger](https://docs.api7.ai/ingress-controller/canary-releases/flagger.md): Use Flagger and APISIX metrics to automate canary analysis, promotion, and rollback through Gateway API HTTPRoute or ApisixRoute resources. ### common-use-cases Explore common use cases enabled by the gateway’s plugin ecosystem and learn how to configure these plugins using the Ingress Controller. - [Common Use Cases](https://docs.api7.ai/ingress-controller/common-use-cases.md): Explore common use cases enabled by the gateway’s plugin ecosystem and learn how to configure these plugins using the Ingress Controller. ### configure-upstream-health-checks Learn how to configure APISIX or API7 Ingress Controller to configure upstream health checks. - [Configure Upstream Health Checks](https://docs.api7.ai/ingress-controller/configure-upstream-health-checks.md): Learn how to configure APISIX or API7 Ingress Controller to configure upstream health checks. ### custom-plugins #### lua Learn how to load Lua custom plugins into gateways in a Kubernetes environment and apply them to routes using APISIX or API7 Ingress Controller. - [Deploy Lua Custom Plugins](https://docs.api7.ai/ingress-controller/custom-plugins/lua.md): Learn how to load Lua custom plugins into gateways in a Kubernetes environment and apply them to routes using APISIX or API7 Ingress Controller. #### wasm Learn how to load Wasm custom plugins into gateways in a Kubernetes environment and apply them to routes using APISIX Ingress Controller. - [Deploy Wasm Plugins](https://docs.api7.ai/ingress-controller/custom-plugins/wasm.md): Learn how to load Wasm custom plugins into gateways in a Kubernetes environment and apply them to routes using APISIX Ingress Controller. ### detect-upstream-protocol-appprotocol Learn how to use APISIX or API7 Ingress Controller to automatically configure upstream protocols based on the values of appProtocol. - [Detect Upstream Protocol with appProtocol](https://docs.api7.ai/ingress-controller/detect-upstream-protocol-appprotocol.md): Learn how to use APISIX or API7 Ingress Controller to automatically configure upstream protocols based on the values of appProtocol. ### documentation Explore API7 and APISIX Ingress Controller docs, covering installation, how-tos, troubleshooting, and reference for dynamic traffic routing via Ingress, Gateway API, and APISIX CRDs. - [Ingress Controller Documentation](https://docs.api7.ai/ingress-controller/documentation.md): Explore API7 and APISIX Ingress Controller docs, covering installation, how-tos, troubleshooting, and reference for dynamic traffic routing via Ingress, Gateway API, and APISIX CRDs. ### high-availability Configure APISIX or API7 Ingress Controller replicas, placement, leader election, and failover validation for a highly available deployment. - [High Availability](https://docs.api7.ai/ingress-controller/high-availability.md): Configure APISIX or API7 Ingress Controller replicas, placement, leader election, and failover validation for a highly available deployment. ### installation #### gitops Plan a GitOps deployment of APISIX or API7 Ingress Controller, including repository structure, resource ownership, CRDs, and credentials. - [Prepare for GitOps](https://docs.api7.ai/ingress-controller/installation/gitops.md): Plan a GitOps deployment of APISIX or API7 Ingress Controller, including repository structure, resource ownership, CRDs, and credentials. - [Manage with Argo CD](https://docs.api7.ai/ingress-controller/installation/gitops/argo-cd.md): Install APISIX or API7 Ingress Controller with Argo CD using stable webhook certificates, explicit CRD ownership, and safe reconciliation. - [Manage with Flux](https://docs.api7.ai/ingress-controller/installation/gitops/flux.md): Install APISIX or API7 Ingress Controller with Flux using HelmRelease resources, explicit CRD policies, drift detection, and verification. #### openshift Learn how to install API7 Ingress Controller on an OpenShift cluster, including prerequisites, configuration, and verification steps. - [Install API7 Ingress Controller on OpenShift](https://docs.api7.ai/ingress-controller/installation/openshift.md): Learn how to install API7 Ingress Controller on an OpenShift cluster, including prerequisites, configuration, and verification steps. ### production #### cross-namespace Configure secure cross-namespace references for Routes, backends, TLS Secrets, and Consumer credentials with APISIX or API7 Ingress Controller. - [Configure Cross-Namespace References](https://docs.api7.ai/ingress-controller/production/cross-namespace.md): Configure secure cross-namespace references for Routes, backends, TLS Secrets, and Consumer credentials with APISIX or API7 Ingress Controller. #### gateway-api-access-control Delegate Gateway API resource management safely between platform and application teams with Kubernetes RBAC for APISIX or API7 Ingress Controller. - [Delegate Gateway API Access with Kubernetes RBAC](https://docs.api7.ai/ingress-controller/production/gateway-api-access-control.md): Delegate Gateway API resource management safely between platform and application teams with Kubernetes RBAC for APISIX or API7 Ingress Controller. #### upgrade Upgrade APISIX or API7 Ingress Controller from 2.1.0 to 2.2.0 by updating CRDs, reviewing compatibility changes, and verifying traffic. - [Upgrade Ingress Controller](https://docs.api7.ai/ingress-controller/production/upgrade.md): Upgrade APISIX or API7 Ingress Controller from 2.1.0 to 2.2.0 by updating CRDs, reviewing compatibility changes, and verifying traffic. ### proxy-grpc-traffic Learn how to use APISIX or API7 Ingress Controller to configure routes to proxy gRPC traffic. - [Proxy gRPC Traffic](https://docs.api7.ai/ingress-controller/proxy-grpc-traffic.md): Learn how to use APISIX or API7 Ingress Controller to configure routes to proxy gRPC traffic. ### proxy-requests-to-a-service Learn how to create a route with the Ingress Controller to proxy requests to a sample HTTP upstream service and verify routing, enabling efficient API traffic management. - [Proxy Requests to a Service](https://docs.api7.ai/ingress-controller/proxy-requests-to-a-service.md): Learn how to create a route with the Ingress Controller to proxy requests to a sample HTTP upstream service and verify routing, enabling efficient API traffic management. ### proxy-tcp-traffic Learn how to configure APISIX or API7 Ingress Controller to proxy TCP traffic by port. - [Proxy TCP Traffic by Port](https://docs.api7.ai/ingress-controller/proxy-tcp-traffic.md): Learn how to configure APISIX or API7 Ingress Controller to proxy TCP traffic by port. ### proxy-tcp-traffic-over-tls Learn how to use APISIX or API7 Ingress Controller to configure routes to proxy TCP traffic over TLS by SNI. - [Proxy TCP Traffic over TLS by SNI](https://docs.api7.ai/ingress-controller/proxy-tcp-traffic-over-tls.md): Learn how to use APISIX or API7 Ingress Controller to configure routes to proxy TCP traffic over TLS by SNI. ### proxy-to-external-services Learn how to configure APISIX or API7 Ingress Controller to proxy requests to external services hosted outside your Kubernetes cluster. - [Proxy Requests to External Services](https://docs.api7.ai/ingress-controller/proxy-to-external-services.md): Learn how to configure APISIX or API7 Ingress Controller to proxy requests to external services hosted outside your Kubernetes cluster. ### proxy-to-weighted-backends Learn how to configure weighted routing to distribute traffic across multiple upstream services using APISIX or API7 Ingress Controller. - [Proxy Requests to Weighted Backends](https://docs.api7.ai/ingress-controller/proxy-to-weighted-backends.md): Learn how to configure weighted routing to distribute traffic across multiple upstream services using APISIX or API7 Ingress Controller. ### proxy-udp-traffic Learn how to configure APISIX or API7 Ingress Controller to proxy UDP traffic by port. - [Proxy UDP Traffic by Port](https://docs.api7.ai/ingress-controller/proxy-udp-traffic.md): Learn how to configure APISIX or API7 Ingress Controller to proxy UDP traffic by port. ### proxy-websocket-connection Learn how to configure APISIX or API7 Ingress Controller to configure routes to proxy WebSocket connections. - [Proxy WebSocket Connection](https://docs.api7.ai/ingress-controller/proxy-websocket-connection.md): Learn how to configure APISIX or API7 Ingress Controller to configure routes to proxy WebSocket connections. ### reference #### annotations Learn how annotations extend the functionality of Kubernetes Ingress and IngressClass resource in the Ingress Controller to configure routing, security, and gateway behaviors. - [Annotations](https://docs.api7.ai/ingress-controller/reference/annotations.md): Learn how annotations extend the functionality of Kubernetes Ingress and IngressClass resource in the Ingress Controller to configure routing, security, and gateway behaviors. #### configuration-file Configure APISIX or API7 Ingress Controller logging, leader election, metrics, synchronization, Gateway API, and webhook settings. - [Configuration File](https://docs.api7.ai/ingress-controller/reference/configuration-file.md): Configure APISIX or API7 Ingress Controller logging, leader election, metrics, synchronization, Gateway API, and webhook settings. #### crd-reference Explore detailed reference documentation for the custom resource definitions (CRDs) supported by the Ingress Controller. - [Custom Resource Definitions API Reference](https://docs.api7.ai/ingress-controller/reference/crd-reference.md): Explore detailed reference documentation for the custom resource definitions (CRDs) supported by the Ingress Controller. #### examples Discover various examples showcasing the Ingress Controller resource configurations to help you effectively tailor settings for your environment. - [Configuration Examples](https://docs.api7.ai/ingress-controller/reference/examples.md): Discover various examples showcasing the Ingress Controller resource configurations to help you effectively tailor settings for your environment. #### helm-charts Learn about the Helm charts for deploying APISIX and API7 Ingress Controllers, including how they function and references for configurable chart values. - [Helm Charts](https://docs.api7.ai/ingress-controller/reference/helm-charts.md): Learn about the Helm charts for deploying APISIX and API7 Ingress Controllers, including how they function and references for configurable chart values. #### ingress-and-gateway-api-support Learn about the Gateway API and Ingress resources supported by the Ingress Controller and their current capabilities. - [Ingress and Gateway API Support](https://docs.api7.ai/ingress-controller/reference/ingress-and-gateway-api-support.md): Learn about the Gateway API and Ingress resources supported by the Ingress Controller and their current capabilities. ### release-notes Review shared and product-specific changes, upgrade requirements, compatibility, and important fixes for APISIX and API7 Ingress Controller releases. - [Release Notes](https://docs.api7.ai/ingress-controller/release-notes.md): Review shared and product-specific changes, upgrade requirements, compatibility, and important fixes for APISIX and API7 Ingress Controller releases. ### set-up-ingress-controller-and-gateway Learn how to quickly deploy and configure API7 Ingress Controller or APISIX Ingress Controller for managing Kubernetes ingress traffic. - [Set Up Ingress Controller and Gateway](https://docs.api7.ai/ingress-controller/set-up-ingress-controller-and-gateway.md): Learn how to quickly deploy and configure API7 Ingress Controller or APISIX Ingress Controller for managing Kubernetes ingress traffic. ### tls-and-mtls #### configure-downstream-https Learn how to use APISIX or API7 Ingress Controller to configure the gateway to accept HTTPS traffic from clients. - [Configure HTTPS Between Client and Gateway](https://docs.api7.ai/ingress-controller/tls-and-mtls/configure-downstream-https.md): Learn how to use APISIX or API7 Ingress Controller to configure the gateway to accept HTTPS traffic from clients. #### configure-downstream-mtls Learn how to use APISIX or API7 Ingress Controller to configure the gateway to require mutual TLS (mTLS) from clients. - [Configure mTLS Between Client and Gateway](https://docs.api7.ai/ingress-controller/tls-and-mtls/configure-downstream-mtls.md): Learn how to use APISIX or API7 Ingress Controller to configure the gateway to require mutual TLS (mTLS) from clients. #### configure-upstream-mtls Learn how to use APISIX or API7 Ingress Controller to configure the gateway to forward traffic to upstream services over mutual TLS (mTLS). - [Configure mTLS Between Gateway and Upstream](https://docs.api7.ai/ingress-controller/tls-and-mtls/configure-upstream-mtls.md): Learn how to use APISIX or API7 Ingress Controller to configure the gateway to forward traffic to upstream services over mutual TLS (mTLS). #### proxy-to-https-upstream Learn how to use APISIX or API7 Ingress Controller to configure the gateway to forward traffic to upstream services over HTTPS. - [Proxy Requests to HTTPS Upstream Services](https://docs.api7.ai/ingress-controller/tls-and-mtls/proxy-to-https-upstream.md): Learn how to use APISIX or API7 Ingress Controller to configure the gateway to forward traffic to upstream services over HTTPS. ### troubleshooting #### admission-webhook Learn about the admission webhook, the error and warning messages it may return when applying Ingress Controller resources, and how to resolve them. - [Understand the Admission Webhook](https://docs.api7.ai/ingress-controller/troubleshooting/admission-webhook.md): Learn about the admission webhook, the error and warning messages it may return when applying Ingress Controller resources, and how to resolve them. #### common-issues Learn how to identify and resolve common issues in APISIX or API7 Ingress Controller with practical guidance for effective troubleshooting. - [Common Issues and Solutions](https://docs.api7.ai/ingress-controller/troubleshooting/common-issues.md): Learn how to identify and resolve common issues in APISIX or API7 Ingress Controller with practical guidance for effective troubleshooting. #### configuration-synchronization Learn how to inspect and troubleshoot configuration translation and synchronization in APISIX or API7 Ingress Controller. - [Troubleshoot Manifest Translation and Synchronization](https://docs.api7.ai/ingress-controller/troubleshooting/configuration-synchronization.md): Learn how to inspect and troubleshoot configuration translation and synchronization in APISIX or API7 Ingress Controller. #### gateway-debug-mode Learn how to enable gateway debug mode in a Kubernetes environment to troubleshoot and monitor the gateway’s runtime behavior effectively. - [Enable Gateway Debug Mode](https://docs.api7.ai/ingress-controller/troubleshooting/gateway-debug-mode.md): Learn how to enable gateway debug mode in a Kubernetes environment to troubleshoot and monitor the gateway’s runtime behavior effectively. --- # Full Documentation Content [Skip to main content](#__docusaurus_skipToContent_fallback) [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) ProductsSolutions[Customers](https://api7.ai/customers) Pricing Resources[Blog](https://api7.ai/blog) [Login](https://console.api7.cloud)Get a DemoStart for Free [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) * Products [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/enterprise) [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/api7-enterprise-vs-apisix) [Apache APISIX vs API7](https://api7.ai/api7-enterprise-vs-apisix)[- ](https://api7.ai/portal) [API7 API Portal](https://api7.ai/portal) [Apache APISIX](https://api7.ai/apisix)[- ](https://api7.ai/apisix) [What's Apache APISIX?](https://api7.ai/apisix)[- ](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway) [Why Apache APISIX?](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway)[- ](https://api7.ai/apache-apisix-enterprise-support) [APISIX Commercial Support](https://api7.ai/apache-apisix-enterprise-support) [AISIX AI Gateway](https://api7.ai/ai-gateway)[- ](https://api7.ai/ai-gateway) [AISIX AI Gateway](https://api7.ai/ai-gateway) * Solutions [Developer](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/monolith-to-microservices) [Monolith to Microservices](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/on-prem-to-hybrid-cloud) [On-Prem to Hybrid Cloud](https://api7.ai/solutions/on-prem-to-hybrid-cloud)[- ](https://api7.ai/solutions/observability) [Observability](https://api7.ai/solutions/observability) [- ](https://api7.ai/solutions/vm-to-kubernetes) [VM to Kubernetes](https://api7.ai/solutions/vm-to-kubernetes)[- ](https://api7.ai/solutions/zero-trust-security) [Zero Trust Security](https://api7.ai/solutions/zero-trust-security) [Industry](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/financial-services) [Financial Services](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/iot) [IoT and Automotive](https://api7.ai/solutions/iot)[- ](https://api7.ai/solutions/blockchain) [Blockchain](https://api7.ai/solutions/blockchain) [- ](https://api7.ai/solutions/manufacturing) [Manufacturing](https://api7.ai/solutions/manufacturing) * [Customers](https://api7.ai/customers) * Pricing [- ](https://api7.ai/pricing) [API Gateway](https://api7.ai/pricing)[- ](https://api7.ai/ai-gateway/pricing) [AI Gateway](https://api7.ai/ai-gateway/pricing) * Resources [Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/ai-gateway/.md) [AISIX Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/api7-gateway) [API7 Gateway](https://docs.api7.ai/api7-gateway)[- ](https://docs.api7.ai/api7-gateway/ai-agent-skills.md) [API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[- ](https://docs.api7.ai/apisix) [Apache APISIX](https://docs.api7.ai/apisix)[- ](https://docs.api7.ai/apisix/ai-agent-skills.md) [APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md) [Compare](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-kong) [Apache APISIX vs Kong](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-nginx) [Apache APISIX vs NGINX](https://api7.ai/apisix-vs-nginx)[- ](https://api7.ai/api-gateway-comparison) [2026 Top API Gateway Comparison](https://api7.ai/api-gateway-comparison)[- ](https://api7.ai/ai-gateway-comparison) [AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) [Learn](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/openresty) [OpenResty (NGINX + Lua)](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/api-gateway-guide) [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[- ](https://api7.ai/learning-center/ai-gateway-guide) [AI Gateway Guide](https://api7.ai/learning-center/ai-gateway-guide)[- ](https://api7.ai/learning-center/api-infrastructure-guide) [API Infrastructure Guide](https://api7.ai/learning-center/api-infrastructure-guide) [Explore](https://api7.ai/demos)[- ](https://api7.ai/demos) [Demo Hub](https://api7.ai/demos)[- ](https://docs.api7.ai/hub.md) [Plugin Hub](https://docs.api7.ai/hub.md)[- ](https://api7.ai/category/usercase) [Case Studies](https://api7.ai/category/usercase) * [Blog](https://api7.ai/blog) Get a DemoStart for Free [](https://docs.api7.ai/)[Apache APISIX](https://docs.api7.ai/apisix/documentation.md)[API7 Gateway](https://docs.api7.ai/api7-gateway/overview.md)[Ingress Controller](https://docs.api7.ai/ingress-controller/documentation.md)[AISIX AI Gateway](https://docs.api7.ai/ai-gateway/.md)[Plugin Hub](https://docs.api7.ai/hub.md) Search API version1.0.0 (/ai-gateway/reference/admin-api) # AISIX Admin API The AISIX Admin API is the read-only operational surface of an open-source AISIX gateway: list and inspect the loaded models, caller API keys, provider credentials, guardrails, MCP servers, A2A... ## Health * [`GET` Get Gateway Health `/admin/v1/health`](#fallback-operation-0-0) Get model health levels and configuration-watch freshness for this gateway. * [`GET` Get Liveness Status `/livez`](#fallback-operation-0-1) Process liveness: should this instance be restarted? Answers 200 whenever it answers at all, including throughout a graceful drain — draining is deliberate work, and restarting an instance that is finishing the requests it accepted would kill exactly those. Use /readyz for... * [`GET` Get Readiness Status `/readyz`](#fallback-operation-0-2) Traffic eligibility (readiness): 200 when the instance can serve, 503 while draining or before the first config apply. Stays 200 for as long as the instance keeps serving that configuration, however long ago the last config event was — see /admin/v1/health and /status/config... ## OpenAPI * [`GET` Open Scalar UI `/admin/openapi-scalar`](#fallback-operation-1-0) Open the browser UI that loads /admin/openapi.json from the admin listener. * [`GET` Get OpenAPI Document `/admin/openapi.json`](#fallback-operation-1-1) Get the machine-readable OpenAPI 3.1 document served by this gateway process. ## Models * [`GET` List Models `/admin/v1/models`](#fallback-operation-2-0) List all configured model resources. * [`GET` List Model Runtime Status `/admin/v1/models/status`](#fallback-operation-2-1) Returns runtime routing and exclusion state for every model. * [`GET` Get Model by ID `/admin/v1/models/{id}`](#fallback-operation-2-2) Get a model resource by ID. ## Caller API Keys * [`GET` List Caller API Keys `/admin/v1/api_keys`](#fallback-operation-3-0) List caller API keys with plaintext credentials redacted. * [`GET` Get Caller API Key by ID `/admin/v1/api_keys/{id}`](#fallback-operation-3-1) Get a caller API key by ID with plaintext credentials redacted. * [`GET` List Caller API Keys (alternate path) `/admin/v1/apikeys`](#fallback-operation-3-2) Alternate spelling of /admin/v1/api keys. Requests and responses are identical on both paths. List caller API keys with plaintext credentials redacted. * [`GET` Get Caller API Key by ID (alternate path) `/admin/v1/apikeys/{id}`](#fallback-operation-3-3) Alternate spelling of /admin/v1/api keys/{id}. Requests and responses are identical on both paths. Get a caller API key by ID with plaintext credentials redacted. ## Provider Keys * [`GET` List Provider Keys `/admin/v1/provider_keys`](#fallback-operation-4-0) List all configured provider key resources. * [`GET` Get Provider Key by ID `/admin/v1/provider_keys/{id}`](#fallback-operation-4-1) Get a provider key resource by ID. ## MCP Servers * [`GET` List MCP Servers `/admin/v1/mcp_servers`](#fallback-operation-5-0) List registered upstream MCP server resources. * [`GET` Get MCP Server by ID `/admin/v1/mcp_servers/{id}`](#fallback-operation-5-1) Get an upstream MCP server resource by ID. ## A2A Agents * [`GET` List A2A Agents `/admin/v1/a2a_agents`](#fallback-operation-6-0) List registered upstream A2A agent resources. * [`GET` Get A2A Agent by ID `/admin/v1/a2a_agents/{id}`](#fallback-operation-6-1) Get an upstream A2A agent resource by ID. ## Passthrough Routes * [`GET` List Passthrough Routes `/admin/v1/passthrough_routes`](#fallback-operation-7-0) List explicit passthrough route resources. * [`GET` Get Passthrough Route by ID `/admin/v1/passthrough_routes/{id}`](#fallback-operation-7-1) Get an explicit passthrough route resource by ID. ## Guardrails * [`GET` List Guardrails `/admin/v1/guardrails`](#fallback-operation-8-0) List all configured guardrail resources. * [`GET` Get Guardrail by ID `/admin/v1/guardrails/{id}`](#fallback-operation-8-1) Get a guardrail resource by ID. ## Cache Policies * [`GET` List Cache Policies `/admin/v1/cache_policies`](#fallback-operation-9-0) List all configured cache policy resources. * [`GET` Get Cache Policy by ID `/admin/v1/cache_policies/{id}`](#fallback-operation-9-1) Get a cache policy resource by ID. ## Observability Exporters * [`GET` List Observability Exporters `/admin/v1/observability_exporters`](#fallback-operation-10-0) List all configured observability exporter resources. * [`GET` Get Observability Exporter by ID `/admin/v1/observability_exporters/{id}`](#fallback-operation-10-1) Get an observability exporter resource by ID. ## Playground * [`POST` Create Playground Chat Completion `/playground/chat/completions`](#fallback-operation-11-0) Forwards a chat completion through the proxy path for local playground testing. Use a proxy API key, not an admin key. ![API7.ai Logo](https://static.api7.ai/uploads/2025/03/02/api7.ai-white.avif) The digital world is connected by APIs,
API7.ai exists to make APIs more efficient, reliable, and secure. Sign up for API7 newsletter [Email address]()Subscribe Product [API7 Gateway](https://api7.ai/enterprise)[AISIX AI Gateway](https://api7.ai/ai-gateway)[API7 API Portal](https://api7.ai/portal) Learn [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[Plugin Hub](https://docs.api7.ai/hub.md)[API Gateway Comparison](https://api7.ai/api-gateway-comparison)[Customers](https://api7.ai/customers) Resources [API Gateway Docs](https://docs.api7.ai/apisix/documentation.md)[APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md)[API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[Blog](https://api7.ai/blog)[Demo Hub](https://api7.ai/demos)[APISIX vs Kong](https://api7.ai/apisix-vs-kong)[AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) Company [About](https://api7.ai/about)[Contact](https://api7.ai/contact)[Partners](https://api7.ai/partners)[Compliance Standards](https://api7.ai/compliance)[Brand Assets](https://api7.ai/branding)[Terms & Privacy](https://api7.ai/terms) *** [![SOC2 Type II](https://static.api7.ai/uploads/2025/03/02/fMrK6JR5_21972-312_SOC_NonCPA.avif)](https://api7.ai/compliance) [![ISO 27001](https://static.api7.ai/uploads/2025/03/02/kStGFFd2_iso-27001.avif)](https://api7.ai/compliance) [![HIPAA](https://static.api7.ai/uploads/2025/03/02/PSOVypp6_hipaa.avif)](https://api7.ai/compliance) [![GDPR](https://static.api7.ai/uploads/2025/03/02/6cv5RTfR_gdpr.avif)](https://api7.ai/compliance) [![Red Herring](https://static.api7.ai/uploads/2025/03/02/6385ad60e1f6a.avif)](https://api7.ai/blog/among-2022-red-herring-top-100-global) Copyright © APISEVEN PTE. LTD 2019 – 2026. Apache, Apache APISIX, APISIX, and associated open source project names are trademarks of the [Apache Software Foundation](https://www.apache.org/) [](https://www.linkedin.com/company/api7-ai/)[](https://github.com/api7)[](https://twitter.com/api7_ai) --- [Skip to main content](#__docusaurus_skipToContent_fallback) [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) ProductsSolutions[Customers](https://api7.ai/customers) Pricing Resources[Blog](https://api7.ai/blog) [Login](https://console.api7.cloud)Get a DemoStart for Free [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) * Products [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/enterprise) [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/api7-enterprise-vs-apisix) [Apache APISIX vs API7](https://api7.ai/api7-enterprise-vs-apisix)[- ](https://api7.ai/portal) [API7 API Portal](https://api7.ai/portal) [Apache APISIX](https://api7.ai/apisix)[- ](https://api7.ai/apisix) [What's Apache APISIX?](https://api7.ai/apisix)[- ](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway) [Why Apache APISIX?](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway)[- ](https://api7.ai/apache-apisix-enterprise-support) [APISIX Commercial Support](https://api7.ai/apache-apisix-enterprise-support) [AISIX AI Gateway](https://api7.ai/ai-gateway)[- ](https://api7.ai/ai-gateway) [AISIX AI Gateway](https://api7.ai/ai-gateway) * Solutions [Developer](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/monolith-to-microservices) [Monolith to Microservices](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/on-prem-to-hybrid-cloud) [On-Prem to Hybrid Cloud](https://api7.ai/solutions/on-prem-to-hybrid-cloud)[- ](https://api7.ai/solutions/observability) [Observability](https://api7.ai/solutions/observability) [- ](https://api7.ai/solutions/vm-to-kubernetes) [VM to Kubernetes](https://api7.ai/solutions/vm-to-kubernetes)[- ](https://api7.ai/solutions/zero-trust-security) [Zero Trust Security](https://api7.ai/solutions/zero-trust-security) [Industry](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/financial-services) [Financial Services](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/iot) [IoT and Automotive](https://api7.ai/solutions/iot)[- ](https://api7.ai/solutions/blockchain) [Blockchain](https://api7.ai/solutions/blockchain) [- ](https://api7.ai/solutions/manufacturing) [Manufacturing](https://api7.ai/solutions/manufacturing) * [Customers](https://api7.ai/customers) * Pricing [- ](https://api7.ai/pricing) [API Gateway](https://api7.ai/pricing)[- ](https://api7.ai/ai-gateway/pricing) [AI Gateway](https://api7.ai/ai-gateway/pricing) * Resources [Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/ai-gateway/.md) [AISIX Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/api7-gateway) [API7 Gateway](https://docs.api7.ai/api7-gateway)[- ](https://docs.api7.ai/api7-gateway/ai-agent-skills.md) [API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[- ](https://docs.api7.ai/apisix) [Apache APISIX](https://docs.api7.ai/apisix)[- ](https://docs.api7.ai/apisix/ai-agent-skills.md) [APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md) [Compare](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-kong) [Apache APISIX vs Kong](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-nginx) [Apache APISIX vs NGINX](https://api7.ai/apisix-vs-nginx)[- ](https://api7.ai/api-gateway-comparison) [2026 Top API Gateway Comparison](https://api7.ai/api-gateway-comparison)[- ](https://api7.ai/ai-gateway-comparison) [AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) [Learn](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/openresty) [OpenResty (NGINX + Lua)](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/api-gateway-guide) [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[- ](https://api7.ai/learning-center/ai-gateway-guide) [AI Gateway Guide](https://api7.ai/learning-center/ai-gateway-guide)[- ](https://api7.ai/learning-center/api-infrastructure-guide) [API Infrastructure Guide](https://api7.ai/learning-center/api-infrastructure-guide) [Explore](https://api7.ai/demos)[- ](https://api7.ai/demos) [Demo Hub](https://api7.ai/demos)[- ](https://docs.api7.ai/hub.md) [Plugin Hub](https://docs.api7.ai/hub.md)[- ](https://api7.ai/category/usercase) [Case Studies](https://api7.ai/category/usercase) * [Blog](https://api7.ai/blog) Get a DemoStart for Free [](https://docs.api7.ai/)[Apache APISIX](https://docs.api7.ai/apisix/documentation.md)[API7 Gateway](https://docs.api7.ai/api7-gateway/overview.md)[Ingress Controller](https://docs.api7.ai/ingress-controller/documentation.md)[AISIX AI Gateway](https://docs.api7.ai/ai-gateway/.md)[Plugin Hub](https://docs.api7.ai/hub.md) Search API versionCurrent release (1.0.0) (/ai-gateway/reference/cloud-admin-api)[Compare versions](https://docs.api7.ai/ai-gateway/reference/cloud-admin-api-changelog.md) # AISIX Cloud Admin API The AISIX Cloud Admin API is the stable, customer-facing automation contract for AISIX Cloud across its control-plane deployment options. It lets operators manage organization-scoped environments... ## Caller API Keys * [`GET` List Caller API Keys `/environments/{env_id}/api_keys`](#fallback-operation-0-0) List caller API keys in an environment. Pagination is opt-in: omit page size to get the full key list in one response. page is only meaningful together with page size and is rejected without it. * [`POST` Create Caller API Key `/environments/{env_id}/api_keys`](#fallback-operation-0-1) Create a caller credential for an environment. The plaintext bearer is returned once in the create response and cannot be recovered later. * [`DELETE` Delete Caller API Key `/environments/{env_id}/api_keys/{api_key_id}`](#fallback-operation-0-2) Deletes the caller credential. Any caller still using the plaintext bearer receives 401 Unauthorized on subsequent gateway requests. Refused with 409 while something resolves an identity to this key — a passthrough route using it as its anonymous principal, or a claim mapping... * [`PATCH` Update Caller API Key `/environments/{env_id}/api_keys/{api_key_id}`](#fallback-operation-0-3) Update selected caller API key fields. Nullable fields can be cleared with an explicit null. Changing the underlying bearer is not part of this operation — use the rotate operation instead. * [`POST` Rotate Caller API Key `/environments/{env_id}/api_keys/{api_key_id}/rotate`](#fallback-operation-0-4) Replace the key's underlying bearer with a freshly generated value in one operation. The key resource is preserved — name, allowed models, rate limit, bindings, expiry deadline, and disabled state carry over; only the credential changes. The old plaintext stops authenticating... ## Environments * [`GET` List Environments `/environments`](#fallback-operation-1-0) Return every environment in the authenticated organization. The response is not paginated. * [`POST` Create Environment `/environments`](#fallback-operation-1-1) Create an environment in the authenticated organization. Environment names are case-sensitive and must be unique within the organization. * [`GET` Get Environment by ID `/environments/{env_id}`](#fallback-operation-1-2) Return one environment in the authenticated organization. An ID that is missing or belongs to another organization returns 404. * [`DELETE` Delete Environment `/environments/{env_id}`](#fallback-operation-1-3) Deletes an environment after revoking its active AISIX gateway certificates. If certificate revocation fails, the environment remains intact and the delete operation can be retried. * [`PATCH` Update Environment `/environments/{env_id}`](#fallback-operation-1-4) Update an environment's display name and MCP OAuth discovery settings. Fields that are omitted keep their current value. ## Models * [`GET` List Models `/environments/{env_id}/models`](#fallback-operation-2-0) Return every direct, routing, ensemble, semantic, and embedding model configured in the environment. The response is not paginated. * [`POST` Create Model `/environments/{env_id}/models`](#fallback-operation-2-1) Create a model alias in an environment. The AISIX Cloud control plane creates a direct model unless another kind is selected. * [`GET` Get Model by ID `/environments/{env_id}/models/{model_id}`](#fallback-operation-2-2) Return one model and the configuration block for its model kind. The model must belong to the environment in the request path. * [`DELETE` Delete Model `/environments/{env_id}/models/{model_id}`](#fallback-operation-2-3) Delete a model. If another model still references it, rebind or delete the dependent model first. Refused with 409 while a semantic cache policy or a kind: semantic guardrail uses it as their embedding model. The guardrail case is not cosmetic: a guardrail defaults to fail... * [`PATCH` Update Model `/environments/{env_id}/models/{model_id}`](#fallback-operation-2-4) Update selected fields on a model. The model kind is fixed at creation, and editable fields depend on the current kind. Renaming a model rebinds everything that references it. The one rename that is refused with 400 is a rename into a wildcard alias — display name or model... ## Provider Keys * [`GET` List Provider Keys `/provider_keys`](#fallback-operation-3-0) List provider keys in the authenticated organization. Fetch a single key to see its endpoint override. Pagination is opt-in: omit page size to get the full key list in one response. page is only meaningful together with page size and is rejected without it. * [`POST` Create Provider Key `/provider_keys`](#fallback-operation-3-1) Create an upstream provider credential. The plaintext API key is encrypted before storage and is never returned by read endpoints. * [`GET` Get Provider Key by ID `/provider_keys/{provider_key_id}`](#fallback-operation-3-2) Get a provider key, including its endpoint override. * [`DELETE` Delete Provider Key `/provider_keys/{provider_key_id}`](#fallback-operation-3-3) Deletes the provider key from every environment where it was allowed. Models that still reference the key cannot dispatch successfully, so rebind affected models before deleting the key. * [`PATCH` Update Provider Key `/provider_keys/{provider_key_id}`](#fallback-operation-3-4) Update selected provider key fields, including the upstream secret. Send api key (or config, for a multi-field credential) to rotate the secret in place: every model that references this provider key picks up the new credential, with no model or caller change. Omit those... ## MCP Servers * [`GET` List MCP Servers `/mcp_servers`](#fallback-operation-4-0) List MCP servers in the authenticated organization. Stored bearer secrets are never included in read responses. * [`POST` Create MCP Server `/mcp_servers`](#fallback-operation-4-1) Register an upstream MCP server and expose it to the allowed environments. * [`GET` Get MCP Server by ID `/mcp_servers/{mcp_server_id}`](#fallback-operation-4-2) Get an MCP server. Stored bearer secrets are never returned. * [`DELETE` Delete MCP Server `/mcp_servers/{mcp_server_id}`](#fallback-operation-4-3) Deletes the MCP server from every environment where it was allowed. * [`PATCH` Update MCP Server `/mcp_servers/{mcp_server_id}`](#fallback-operation-4-4) Update selected MCP server fields. A new secret rotates a bearer credential. This caller holds the permission that approves servers, so an approved server keeps serving across the patch and the new configuration is published immediately; the edit is recorded as its review\... * [`GET` List Generated MCP Tools `/mcp_servers/{mcp_server_id}/tools`](#fallback-operation-4-5) List the MCP tools an OpenAPI-backed server generates, derived from its stored OpenAPI document by the same walk that validated it at write time. Only servers of type: openapi can be listed: an upstream MCP server's tool set lives on the upstream, which the control plane never... * [`POST` Approve MCP Server `/mcp_servers/{mcp_server_id}/approve`](#fallback-operation-4-6) Publishes a reviewed MCP server: it is projected to the environments in allowed environments and becomes discoverable and callable by gateway clients. A previously rejected server can be approved. Approving one that is already approved returns 400. On a live server carrying a... * [`POST` Reject MCP Server `/mcp_servers/{mcp_server_id}/reject`](#fallback-operation-4-7) Refuses a submitted MCP server, and is also how an approval is revoked: rejecting an approved server withdraws it from every environment it was serving in, so gateway clients stop seeing it. Rejecting one that is already rejected returns 400. On a live server carrying a... * [`POST` Submit MCP Server for Review `/mcp_server_submissions`](#fallback-operation-4-8) Proposes an upstream MCP server without publishing it. The server is registered with approval status set to pending review and is not projected to any environment, so it cannot be discovered or called until a reviewer approves it. This is the entry point for roles that may... * [`PATCH` Revise a Submitted MCP Server `/mcp_server_submissions/{mcp_server_id}`](#fallback-operation-4-9) The proposer's write path, on a permission that cannot publish. On a server that is not live it corrects the submission in place and leaves it in the review queue — how a rejected server is fixed and resubmitted. On a server that IS live it changes nothing: the patch is stored... ## A2A Agents * [`GET` List A2A Agents `/a2a_agents`](#fallback-operation-5-0) List A2A agents in the authenticated organization. Stored credentials are never included in read responses. * [`POST` Create A2A Agent `/a2a_agents`](#fallback-operation-5-1) Register an upstream A2A agent and expose it to the allowed environments. * [`GET` Get A2A Agent by ID `/a2a_agents/{a2a_agent_id}`](#fallback-operation-5-2) Get an A2A agent. Stored credentials are never returned. * [`DELETE` Delete A2A Agent `/a2a_agents/{a2a_agent_id}`](#fallback-operation-5-3) Deletes the A2A agent from every environment where it was allowed. * [`PATCH` Update A2A Agent `/a2a_agents/{a2a_agent_id}`](#fallback-operation-5-4) Update selected A2A agent fields. A new secret rotates the stored credential. ## Members * [`GET` List Members `/members`](#fallback-operation-6-0) List organization members. Pagination is opt-in: omit page size to get the full member list in one response. page is only meaningful together with page size and is rejected without it. * [`POST` Create Member `/members`](#fallback-operation-6-1) Create a login-less member directly, bypassing the email-invitation handshake. The created principal can be added to teams and issued caller API keys immediately, but carries no dashboard credentials and can never sign in. The role is fixed to member. * [`DELETE` Remove Organization Member `/members/{user_id}`](#fallback-operation-6-2) Remove a member from the organization. Only an organization owner can do this, and the last owner cannot be removed. The member's environment role bindings and team memberships go with them. Caller API keys they own are not deleted and keep working — the user id already... ## Invitations * [`GET` List Invitations `/invitations`](#fallback-operation-7-0) List the organization's invitations. By default only invitations that can still be redeemed are returned; pass include stale=true to also see accepted, revoked and expired ones. * [`POST` Create Invitation `/invitations`](#fallback-operation-7-1) Invite an email address to join the organization with a given role. The response carries the invitation token in clear text, once. plaintext and the invite url built from it are never recoverable afterwards — anyone holding either can redeem the invitation, so deliver it over... * [`DELETE` Revoke Invitation `/invitations/{invitation_id}`](#fallback-operation-7-2) Revoke a pending invitation so its link can no longer be redeemed. An invitation that is already accepted, revoked or expired is reported as 404 — there is nothing left to revoke. ## Roles * [`PATCH` Update Member Organization Role `/members/{user_id}`](#fallback-operation-8-0) Replace a member's organization-wide role. Only an organization owner can assign roles. Assigning owner clears the member's environment role bindings because owners already have full access. The last owner cannot be demoted. * [`GET` List Member Environment Role Bindings `/members/{member_id}/role_bindings`](#fallback-operation-8-1) List the roles granted to one member in individual environments. Each binding adds its role to the member's organization role for resources inside that environment. It does not reduce the member's organization-level access or grant control over the environment object itself. * [`PUT` Replace Member Environment Role Bindings `/members/{member_id}/role_bindings`](#fallback-operation-8-2) Replace the member's complete set of environment-scoped role bindings. Send an empty array to remove every binding. Only an organization owner can change bindings. The target member cannot be an owner, and each environment can appear at most once. Changes can take up to 30... * [`GET` List Roles `/roles`](#fallback-operation-8-3) List the built-in owner, admin, and member roles together with every custom role defined in the authenticated organization. * [`POST` Create Custom Role `/roles`](#fallback-operation-8-4) Create an organization-scoped custom role. Its name is permanent because member and directory-sync assignments reference it by name. A custom role can grant only permission pairs available to the built-in admin role and cannot grant role management. * [`DELETE` Delete Custom Role `/roles/{role_name}`](#fallback-operation-8-5) Delete a custom role. Built-in roles cannot be deleted. A custom role must first be removed from members, pending invitations, directory-sync settings, and environment role bindings. Member assignments and environment bindings can be cleared through this API. Pending... * [`PATCH` Update Custom Role `/roles/{role_name}`](#fallback-operation-8-6) Update a custom role's description or replace its permissions. The role name cannot change. Built-in roles cannot be updated. Permission changes can take up to 30 seconds to propagate across control-plane replicas. ## Teams * [`GET` List Teams `/teams`](#fallback-operation-9-0) List the organization's teams, each with its current member count. Pagination is opt-in: omit page size to get every team in one response. page is only meaningful together with page size and is rejected without it. * [`POST` Create Team `/teams`](#fallback-operation-9-1) Create a team. Teams group organization members so a limit, budget or MCP entitlement can be written once and apply to everyone on the team. display name must be unique within the organization. * [`GET` Get Team `/teams/{team_id}`](#fallback-operation-9-2) Return one team with its current member count. * [`DELETE` Delete Team `/teams/{team_id}`](#fallback-operation-9-3) Delete the team. Its entitlements are cleared first, so caller API keys bound to it stop carrying the team's MCP layer; the keys themselves are not deleted. * [`PATCH` Update Team `/teams/{team_id}`](#fallback-operation-9-4) Update the team's name or description. Omitted fields are left unchanged. * [`GET` List Team Members `/teams/{team_id}/members`](#fallback-operation-9-5) List everyone on the team, with the organization email and display name resolved alongside the team role. * [`POST` Add Team Member `/teams/{team_id}/members`](#fallback-operation-9-6) Put an existing organization member on the team. The user must already belong to the organization — this endpoint does not invite anyone. role defaults to member. * [`DELETE` Remove Team Member `/teams/{team_id}/members/{user_id}`](#fallback-operation-9-7) Take a member off the team. The organization membership is untouched. A team must keep at least one lead, so removing the last one is refused. * [`PATCH` Update Team Member Role `/teams/{team_id}/members/{user_id}`](#fallback-operation-9-8) Change a member's role within the team. A team must keep at least one lead: demoting the last one is refused, so promote another member first. ## Guardrails * [`GET` List Guardrails `/environments/{env_id}/guardrails`](#fallback-operation-10-0) Return every guardrail definition in the environment. Scope attachments are listed separately through the attachments endpoint. * [`POST` Create Guardrail `/environments/{env_id}/guardrails`](#fallback-operation-10-1) Create a guardrail in an environment. The per-kind config shape is validated by the server against the guardrail catalog (GET /guardrails/schema); this spec models the stable envelope and treats config as an open object. * [`POST` Test Guardrail Connection `/environments/{env_id}/guardrails/test-connection`](#fallback-operation-10-2) Probe a remote-API guardrail provider (Bedrock, Azure Content Safety, Aliyun, Lakera, OpenAI Moderation, …) with the supplied credentials before saving the guardrail. The request mirrors the create body's kind + config; the exact config shape is provider-specific and validated... * [`GET` Get Guardrail by ID `/environments/{env_id}/guardrails/{guardrail_id}`](#fallback-operation-10-3) Return one guardrail definition in the environment. Provider credentials in the kind-specific configuration are redacted. * [`DELETE` Delete Guardrail `/environments/{env_id}/guardrails/{guardrail_id}`](#fallback-operation-10-4) Delete a guardrail definition from the environment and remove it from the configuration distributed to connected gateways. * [`PATCH` Update Guardrail `/environments/{env_id}/guardrails/{guardrail_id}`](#fallback-operation-10-5) Update selected fields on a guardrail. The kind is fixed at creation; config (when present) is validated by the server against the guardrail catalog for that kind. * [`GET` List Guardrail Attachments `/environments/{env_id}/guardrails/{guardrail_id}/attachments`](#fallback-operation-10-6) Return the scope attachments for one guardrail. Each attachment determines whether the guardrail applies to an environment, model, caller API key, or team. * [`POST` Attach Guardrail to a Scope `/environments/{env_id}/guardrails/{guardrail_id}/attachments`](#fallback-operation-10-7) Attach a guardrail to an env, model, api key, or team scope. scope id is required for every scope except env (which must omit it). * [`DELETE` Detach Guardrail from a Scope `/environments/{env_id}/guardrails/{guardrail_id}/attachments/{attachment_id}`](#fallback-operation-10-8) Delete one scope attachment so it no longer makes the guardrail applicable through that scope. Other attachments are unchanged. * [`GET` List Guardrail Providers `/guardrails/providers`](#fallback-operation-10-9) Catalog of guardrail providers and their kinds. Read-only metadata used by the dashboard's guardrail form; the response shape is catalog-driven. * [`GET` Get Guardrail Config Schema `/guardrails/schema`](#fallback-operation-10-10) Per-kind JSON schema for the guardrail config blob, consumed by the dashboard form. Read-only; the response is catalog-driven. Not every kind publishes a dynamic schema: the ones the dashboard renders with a built-in form (see has schema on GET /guardrails/providers) answer... ## Cache Policies * [`GET` List Cache Policies `/environments/{env_id}/cache_policies`](#fallback-operation-11-0) Return every prompt-response cache policy configured in the environment. The response is not paginated. * [`POST` Create Cache Policy `/environments/{env_id}/cache_policies`](#fallback-operation-11-1) Create a prompt-response cache rule in an environment. Policy names are unique within the environment, and the name and backend cannot be changed after creation. * [`GET` Get Cache Policy by ID `/environments/{env_id}/cache_policies/{cache_policy_id}`](#fallback-operation-11-2) Return one cache policy. The policy must belong to the environment in the request path. * [`DELETE` Delete Cache Policy `/environments/{env_id}/cache_policies/{cache_policy_id}`](#fallback-operation-11-3) Deletes the cache policy. The gateway stops serving cached responses for the traffic the policy covered. * [`PATCH` Update Cache Policy `/environments/{env_id}/cache_policies/{cache_policy_id}`](#fallback-operation-11-4) Update selected fields on a cache policy. The policy name and backend are fixed at creation — delete and recreate the policy to change them. * [`POST` Purge Cache Policy Entries `/environments/{env_id}/cache_policies/{cache_policy_id}/purge`](#fallback-operation-11-5) Invalidate every entry cached under this policy, across both exact and semantic matching and on every gateway instance. The operation increments the policy's purge generation; gateways pick the new generation up through configuration propagation (typically within seconds) and... ## Observability Exporters * [`GET` List Observability Exporters `/environments/{env_id}/observability_exporters`](#fallback-operation-12-0) Return every telemetry exporter configured in the environment. Secret header values and credential material are not returned. * [`POST` Create Observability Exporter `/environments/{env_id}/observability_exporters`](#fallback-operation-12-1) Create a telemetry exporter in an environment. The kind value selects which configuration fields apply; fields that belong to other kinds are ignored. Exporter names are unique within the environment. Secrets are never part of this request: aliyun sls, object store, and... * [`GET` Get Observability Exporter by ID `/environments/{env_id}/observability_exporters/{exporter_id}`](#fallback-operation-12-2) Return one telemetry exporter in the environment. Secret header values and credential material are not returned. * [`DELETE` Delete Observability Exporter `/environments/{env_id}/observability_exporters/{exporter_id}`](#fallback-operation-12-3) Deletes the exporter. The gateway stops shipping telemetry to the target. * [`PATCH` Update Observability Exporter `/environments/{env_id}/observability_exporters/{exporter_id}`](#fallback-operation-12-4) Update selected fields on an exporter. The exporter name and kind are fixed at creation, and a field that does not belong to the exporter's kind is rejected. Configuration changes are validated against the same rules as create; an explicit empty string clears optional fields... ## Rate Limit Policies * [`GET` List Rate Limit Policies `/environments/{env_id}/rate_limits`](#fallback-operation-13-0) List rate limit policies in an environment. Pagination is opt-in: omit page size to get the full policy list in one response. page is only meaningful together with page size and is rejected without it. * [`POST` Create Rate Limit Policy `/environments/{env_id}/rate_limits`](#fallback-operation-13-1) Create a rate limit policy in an environment. Each policy pins one scope and scope ref pair, and a second policy for the same pair is rejected. At least one of max requests or max tokens must be set, and max tokens is only accepted with window: minute or window: day — the... * [`GET` Get Rate Limit Policy by ID `/environments/{env_id}/rate_limits/{rate_limit_id}`](#fallback-operation-13-2) Return one rate-limit policy. The policy must belong to the environment in the request path. * [`DELETE` Delete Rate Limit Policy `/environments/{env_id}/rate_limits/{rate_limit_id}`](#fallback-operation-13-3) Deletes the policy. The gateway stops enforcing the limit. * [`PATCH` Update Rate Limit Policy `/environments/{env_id}/rate_limits/{rate_limit_id}`](#fallback-operation-13-4) Update selected fields on a rate limit policy. The scope and scope ref pair is fixed at creation. The limit fields are independently clearable: an explicit null clears one limit while keeping the other, and an update that would leave the policy with neither limit is rejected. ## Data Plane Nodes * [`GET` List Data Plane Nodes `/environments/{env_id}/dp_nodes`](#fallback-operation-14-0) List the data plane nodes that have connected to the environment. A node appears after its first status report; a gateway certificate that was issued but never used to connect is not listed. Each entry reflects the node's most recent report, including which configuration... ## Rejected Resources * [`GET` List Rejected Resources `/environments/{env_id}/rejected_resources`](#fallback-operation-15-0) List configuration resources in the environment that at least one data plane node is currently refusing to apply. A save can succeed at the API and still be rejected at a gateway — for example when an older gateway version does not recognize a newer field. A rejected resource... ## MCP Access Policies * [`GET` Get Effective Permissions `/environments/{env_id}/api_keys/{api_key_id}/effective_permissions`](#fallback-operation-16-0) Resolve the MCP tool access a caller API key ends up with once the environment layer, the key team's layer, and the key's own mcp access block are intersected. Every allow and deny pattern in the answer carries its source, and layers names the layers that constrain the key —... * [`GET` Get MCP Access Policy `/environments/{env_id}/mcp_policy`](#fallback-operation-16-1) Fetch the environment layer of the MCP tool ACL. It applies to every caller API key in the environment, intersected with the key's team layer and the key's own mcp access block. The layer is optional: mcp policy is null when the environment configures none, the same way the... * [`PUT` Set MCP Access Policy `/environments/{env_id}/mcp_policy`](#fallback-operation-16-2) Create or replace the environment layer of the MCP tool ACL. allow: \[" "] covers every tool on every MCP server, including servers and tools registered later — choosing it is always an explicit decision, never a default. The layer narrows what keys can reach but never widens... * [`DELETE` Delete MCP Access Policy `/environments/{env_id}/mcp_policy`](#fallback-operation-16-3) Remove the environment layer of the MCP tool ACL. Keys with no team layer and no mcp access block of their own are then left with no layer at all, which means no MCP tool access; keys that configure their own layer keep it. * [`GET` Get Team Entitlements `/teams/{team_id}/entitlements`](#fallback-operation-16-4) Fetch the team's entitlements. The mcp block, when present, is the MCP ACL layer caller API keys bound to this team carry in every environment of the organization; it is intersected with the environment layer for those keys and can only narrow it. Absent means the team adds no... * [`PUT` Set Team Entitlements `/teams/{team_id}/entitlements`](#fallback-operation-16-5) Create, replace, or clear the team's entitlements. Setting the mcp block applies it to the team's caller API keys in every environment of the organization — identity-provider group changes synced to the team propagate automatically, with no per-key edits. Sending "mcp": null... ## OIDC Providers * [`GET` List OIDC Providers `/environments/{env_id}/oidc_providers`](#fallback-operation-17-0) Return every OIDC provider configured for JWT authentication in the environment. The response is not paginated. * [`POST` Create OIDC Provider `/environments/{env_id}/oidc_providers`](#fallback-operation-17-1) Register an identity provider the gateway trusts for JWT authentication in this environment. Once at least one enabled provider exists, requests may authenticate with a JWT issued by it instead of an API key: the token's issuer selects the provider, its signature and claims... * [`GET` Get OIDC Provider by ID `/environments/{env_id}/oidc_providers/{oidc_provider_id}`](#fallback-operation-17-2) Return one OIDC provider. The provider must belong to the environment in the request path. * [`DELETE` Delete OIDC Provider `/environments/{env_id}/oidc_providers/{oidc_provider_id}`](#fallback-operation-17-3) Deletes the OIDC provider. Tokens issued by it stop authenticating as soon as the gateway picks up the change; API keys and their jwt subject bindings are unaffected. * [`PATCH` Update OIDC Provider `/environments/{env_id}/oidc_providers/{oidc_provider_id}`](#fallback-operation-17-4) Update selected fields on an OIDC provider. The provider name is fixed at creation — delete and recreate the provider to change it. Changes take effect on new requests without a gateway restart. ## Claim Mappings * [`GET` List Claim Mappings `/environments/{env_id}/claim_mappings`](#fallback-operation-18-0) Return every claim mapping in the environment. The response is not paginated. * [`POST` Create Claim Mapping `/environments/{env_id}/claim_mappings`](#fallback-operation-18-1) Create a rule that resolves verified JWT claims to an existing caller API key. When a token passes an OIDC provider's verification and no key binds its subject via jwt subject, the enabled mappings naming that provider are evaluated in priority order (lower first, ties broken... * [`GET` Get Claim Mapping by ID `/environments/{env_id}/claim_mappings/{claim_mapping_id}`](#fallback-operation-18-2) Return one claim mapping. The mapping must belong to the environment in the request path. * [`DELETE` Delete Claim Mapping `/environments/{env_id}/claim_mappings/{claim_mapping_id}`](#fallback-operation-18-3) Deletes the claim mapping. Identities it admitted stop authenticating as soon as the gateway picks up the change; API keys bound directly via jwt subject are unaffected. * [`PATCH` Update Claim Mapping `/environments/{env_id}/claim_mappings/{claim_mapping_id}`](#fallback-operation-18-4) Update selected fields on a claim mapping. The mapping name is fixed at creation — delete and recreate the mapping to change it. Changes take effect on new requests without a gateway restart. ## Passthrough Routes * [`GET` List Passthrough Routes `/environments/{env_id}/passthrough_routes`](#fallback-operation-19-0) Return every passthrough route in the environment. The response is not paginated. * [`POST` Create Passthrough Route `/environments/{env_id}/passthrough_routes`](#fallback-operation-19-1) Create an explicit passthrough route: a binding from a gateway entry — a path prefix on the gateway's own URL space, an inbound hosts allowlist (forward-proxy traffic delivered with its original Host), or both — to one upstream target, forwarded without protocol translation... * [`GET` Get Passthrough Route by ID `/environments/{env_id}/passthrough_routes/{passthrough_route_id}`](#fallback-operation-19-2) Return one passthrough route. The route must belong to the environment in the request path. * [`DELETE` Delete Passthrough Route `/environments/{env_id}/passthrough_routes/{passthrough_route_id}`](#fallback-operation-19-3) Deletes the passthrough route. Traffic it served answers 404 (or the tunnel-namespace 410) as soon as the gateway picks up the change. Guardrail attachments scoped to the route are deleted with it; API keys keep any now-dangling allowed routes patterns, which simply grant... * [`PATCH` Update Passthrough Route `/environments/{env_id}/passthrough_routes/{passthrough_route_id}`](#fallback-operation-19-4) Update selected fields on a passthrough route. The route name is fixed at creation — delete and recreate the route to change it. The create-time coupling rules apply to the PATCHED result: the route must keep at least one match dimension, exactly one target shape, and each... ## Budgets * [`GET` List Budgets `/budgets`](#fallback-operation-20-0) List every budget in the organization, across all scopes, each with its current-period spend state. Budgets whose spend is not tracked as a single total (team member) return a zero-seeded state: the limit applies to each member of the team separately. * [`POST` Create Budget `/budgets`](#fallback-operation-20-1) Create a spending cap. Each target — identified by the scope + scope ref pair — can hold at most one budget; creating a second one for the same target is rejected with 409. A hard stop budget makes the gateway reject matching traffic with 429 budget exceeded once the period's... * [`GET` Get Budget `/budgets/{budget_id}`](#fallback-operation-20-2) Return one budget together with its current-period spend state. The state starts as a zero seed at creation and updates as spend is aggregated. * [`DELETE` Delete Budget `/budgets/{budget_id}`](#fallback-operation-20-3) Remove a budget. Spend tracking continues; only the cap is removed. Enforcement stops within a few seconds — in-flight traffic checked against a cached decision may still be rejected briefly. * [`PATCH` Update Budget `/budgets/{budget_id}`](#fallback-operation-20-4) Update a budget's name, limit, period, or enforcement mode. Fields left out keep their current values. The budget's scope and scope ref are fixed at creation — to cap a different target, create a new budget. ## Model Pricing * [`GET` List Model Prices `/model_pricing`](#fallback-operation-21-0) List the prices this organization is billed at: one row per (provider, model), with the organization's own override taking the place of the catalog default where one exists. source says which you are looking at — user for an override, models.dev or snapshot for the catalog... * [`PUT` Set Model Price Override `/model_pricing`](#fallback-operation-21-1) Set this organization's price for one (provider, model) pair, creating the override or replacing the existing one. Every rate is replaced, not merged: a rate you omit is stored as 0, which means the token class bills at the prompt or completion rate rather than keeping... * [`DELETE` Delete Model Price Override `/model_pricing/{id}`](#fallback-operation-21-2) Drop the organization's override so the catalog price takes over again. Catalog rows are not deletable and report 404, the same as an id that does not exist. ## Usage * [`GET` List Usage Events `/environments/{env_id}/usage_events`](#fallback-operation-22-0) Page through the environment's request telemetry, newest first. One row is one upstream attempt , not one request: a request that retried or failed over emits several rows sharing a request id, ordered by attempt index. Aggregate by request id when you need per-request... * [`GET` Export Usage Events `/environments/{env_id}/usage_events/export`](#fallback-operation-22-1) Download the rows listUsageEvents would return for the same filters, as one file. Takes the identical filter set; limit and page are ignored — an export is the whole match, capped at 50,000 rows. Each row carries the model, caller API key and member names alongside their ids... * [`GET` Get Usage Summary `/environments/{env_id}/usage_summary`](#fallback-operation-22-2) Roll the environment's usage up into one bucket per day, model or caller API key over the given window. Counts are request-level: request count is the number of distinct requests in the bucket, not the number of upstream attempts, so a request that retried or failed over... * [`GET` Get Usage Metrics `/environments/{env_id}/usage_metrics`](#fallback-operation-22-3) Request-level counters and exact latency percentiles for the window. These are the figures to use for success rate and latency over a whole window: usage summary also counts distinct requests, but it does so per bucket, so a request whose attempts straddle a bucket boundary is... ## Notification Channels * [`GET` List Notification Channels `/notification_channels`](#fallback-operation-23-0) List the organization's outbound notification channels. Channel URLs are masked — the full URL is write-only. * [`POST` Create Notification Channel `/notification_channels`](#fallback-operation-23-1) Create an outbound notification channel. webhook channels receive alert events as JSON POSTs; slack channels expect a Slack incoming-webhook URL and receive a rendered text message. Enabled channels receive every alert raised in the organization (budget threshold alerts... * [`GET` Get Notification Channel `/notification_channels/{channel_id}`](#fallback-operation-23-2) Return one notification channel in the authenticated organization. The destination URL is masked in the response. * [`DELETE` Delete Notification Channel `/notification_channels/{channel_id}`](#fallback-operation-23-3) Remove a channel. Its delivery history is kept as an audit trail; pending deliveries to it are marked failed. * [`PATCH` Update Notification Channel `/notification_channels/{channel_id}`](#fallback-operation-23-4) Update a channel's name, type, URL, or enabled state. Fields left out keep their current values. Disabling a channel stops future deliveries; already-queued deliveries to it are marked failed rather than parked. Reads mask url, so a read-modify-write sends the mask back under... * [`POST` Test Notification Channel `/notification_channels/{channel_id}/test`](#fallback-operation-23-5) Synchronously send a clearly-labeled test notification through the channel and report the outcome. Always returns 200; the body carries the verdict. ## Notification Deliveries * [`GET` List Notification Deliveries `/notification_deliveries`](#fallback-operation-24-0) Read the delivery log for the organization's notification channels, newest first — what was sent, to which channel, and whether it landed. Paginate with the cursor rather than an offset: pass the next before id from the previous response as before id. A response whose next... ![API7.ai Logo](https://static.api7.ai/uploads/2025/03/02/api7.ai-white.avif) The digital world is connected by APIs,
API7.ai exists to make APIs more efficient, reliable, and secure. Sign up for API7 newsletter [Email address]()Subscribe Product [API7 Gateway](https://api7.ai/enterprise)[AISIX AI Gateway](https://api7.ai/ai-gateway)[API7 API Portal](https://api7.ai/portal) Learn [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[Plugin Hub](https://docs.api7.ai/hub.md)[API Gateway Comparison](https://api7.ai/api-gateway-comparison)[Customers](https://api7.ai/customers) Resources [API Gateway Docs](https://docs.api7.ai/apisix/documentation.md)[APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md)[API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[Blog](https://api7.ai/blog)[Demo Hub](https://api7.ai/demos)[APISIX vs Kong](https://api7.ai/apisix-vs-kong)[AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) Company [About](https://api7.ai/about)[Contact](https://api7.ai/contact)[Partners](https://api7.ai/partners)[Compliance Standards](https://api7.ai/compliance)[Brand Assets](https://api7.ai/branding)[Terms & Privacy](https://api7.ai/terms) *** [![SOC2 Type II](https://static.api7.ai/uploads/2025/03/02/fMrK6JR5_21972-312_SOC_NonCPA.avif)](https://api7.ai/compliance) [![ISO 27001](https://static.api7.ai/uploads/2025/03/02/kStGFFd2_iso-27001.avif)](https://api7.ai/compliance) [![HIPAA](https://static.api7.ai/uploads/2025/03/02/PSOVypp6_hipaa.avif)](https://api7.ai/compliance) [![GDPR](https://static.api7.ai/uploads/2025/03/02/6cv5RTfR_gdpr.avif)](https://api7.ai/compliance) [![Red Herring](https://static.api7.ai/uploads/2025/03/02/6385ad60e1f6a.avif)](https://api7.ai/blog/among-2022-red-herring-top-100-global) Copyright © APISEVEN PTE. LTD 2019 – 2026. Apache, Apache APISIX, APISIX, and associated open source project names are trademarks of the [Apache Software Foundation](https://www.apache.org/) [](https://www.linkedin.com/company/api7-ai/)[](https://github.com/api7)[](https://twitter.com/api7_ai) --- [Skip to main content](#__docusaurus_skipToContent_fallback) [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) ProductsSolutions[Customers](https://api7.ai/customers) Pricing Resources[Blog](https://api7.ai/blog) [Login](https://console.api7.cloud)Get a DemoStart for Free [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) * Products [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/enterprise) [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/api7-enterprise-vs-apisix) [Apache APISIX vs API7](https://api7.ai/api7-enterprise-vs-apisix)[- ](https://api7.ai/portal) [API7 API Portal](https://api7.ai/portal) [Apache APISIX](https://api7.ai/apisix)[- ](https://api7.ai/apisix) [What's Apache APISIX?](https://api7.ai/apisix)[- ](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway) [Why Apache APISIX?](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway)[- ](https://api7.ai/apache-apisix-enterprise-support) [APISIX Commercial Support](https://api7.ai/apache-apisix-enterprise-support) [AISIX AI Gateway](https://api7.ai/ai-gateway)[- ](https://api7.ai/ai-gateway) [AISIX AI Gateway](https://api7.ai/ai-gateway) * Solutions [Developer](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/monolith-to-microservices) [Monolith to Microservices](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/on-prem-to-hybrid-cloud) [On-Prem to Hybrid Cloud](https://api7.ai/solutions/on-prem-to-hybrid-cloud)[- ](https://api7.ai/solutions/observability) [Observability](https://api7.ai/solutions/observability) [- ](https://api7.ai/solutions/vm-to-kubernetes) [VM to Kubernetes](https://api7.ai/solutions/vm-to-kubernetes)[- ](https://api7.ai/solutions/zero-trust-security) [Zero Trust Security](https://api7.ai/solutions/zero-trust-security) [Industry](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/financial-services) [Financial Services](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/iot) [IoT and Automotive](https://api7.ai/solutions/iot)[- ](https://api7.ai/solutions/blockchain) [Blockchain](https://api7.ai/solutions/blockchain) [- ](https://api7.ai/solutions/manufacturing) [Manufacturing](https://api7.ai/solutions/manufacturing) * [Customers](https://api7.ai/customers) * Pricing [- ](https://api7.ai/pricing) [API Gateway](https://api7.ai/pricing)[- ](https://api7.ai/ai-gateway/pricing) [AI Gateway](https://api7.ai/ai-gateway/pricing) * Resources [Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/ai-gateway/.md) [AISIX Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/api7-gateway) [API7 Gateway](https://docs.api7.ai/api7-gateway)[- ](https://docs.api7.ai/api7-gateway/ai-agent-skills.md) [API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[- ](https://docs.api7.ai/apisix) [Apache APISIX](https://docs.api7.ai/apisix)[- ](https://docs.api7.ai/apisix/ai-agent-skills.md) [APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md) [Compare](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-kong) [Apache APISIX vs Kong](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-nginx) [Apache APISIX vs NGINX](https://api7.ai/apisix-vs-nginx)[- ](https://api7.ai/api-gateway-comparison) [2026 Top API Gateway Comparison](https://api7.ai/api-gateway-comparison)[- ](https://api7.ai/ai-gateway-comparison) [AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) [Learn](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/openresty) [OpenResty (NGINX + Lua)](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/api-gateway-guide) [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[- ](https://api7.ai/learning-center/ai-gateway-guide) [AI Gateway Guide](https://api7.ai/learning-center/ai-gateway-guide)[- ](https://api7.ai/learning-center/api-infrastructure-guide) [API Infrastructure Guide](https://api7.ai/learning-center/api-infrastructure-guide) [Explore](https://api7.ai/demos)[- ](https://api7.ai/demos) [Demo Hub](https://api7.ai/demos)[- ](https://docs.api7.ai/hub.md) [Plugin Hub](https://docs.api7.ai/hub.md)[- ](https://api7.ai/category/usercase) [Case Studies](https://api7.ai/category/usercase) * [Blog](https://api7.ai/blog) Get a DemoStart for Free [](https://docs.api7.ai/)[Apache APISIX](https://docs.api7.ai/apisix/documentation.md)[API7 Gateway](https://docs.api7.ai/api7-gateway/overview.md)[Ingress Controller](https://docs.api7.ai/ingress-controller/documentation.md)[AISIX AI Gateway](https://docs.api7.ai/ai-gateway/.md)[Plugin Hub](https://docs.api7.ai/hub.md) Search # API7 Enterprise Admin APIs API7 Enterprise Admin APIs are RESTful APIs that allow you to create, configure, and manage all API7 Enterprise resources programmatically. These APIs power the API7 Dashboard and can be used... ## Service * [`GET` List all services on a gateway group `/apisix/admin/services`](#fallback-operation-0-0) List services through an APISIX Admin API compatible endpoint under /apisix/admin/. Use this to browse APISIX-formatted service objects in a gateway group. Required IAM Permission: Action gateway:GetPublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`POST` Create a service on a gateway group `/apisix/admin/services`](#fallback-operation-0-1) Create a service through an APISIX Admin API compatible endpoint under /apisix/admin/. The payload follows APISIX conventions while operating on the same underlying service resource managed by dashboard APIs. Required IAM Permission: Action gateway:CreatePublishedService... * [`GET` Get a service on a gateway group `/apisix/admin/services/{service_id}`](#fallback-operation-0-2) Get one service through an APISIX Admin API compatible endpoint under /apisix/admin/. The response keeps APISIX field conventions for migration and interoperability scenarios. Required IAM Permission: Action gateway:GetPublishedService, Resource... * [`PUT` Update a service directly `/apisix/admin/services/{service_id}`](#fallback-operation-0-3) Fully update a service through an APISIX Admin API compatible endpoint under /apisix/admin/. This replaces the stored service configuration. Required IAM Permission: Action gateway:UpdatePublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`DELETE` Delete a service on a gateway group `/apisix/admin/services/{service_id}`](#fallback-operation-0-4) Delete a service through an APISIX Admin API compatible endpoint under /apisix/admin/. Removing this object affects all APISIX-compatible references to the service in that gateway group. Required IAM Permission: Action gateway:DeletePublishedService, Resource... * [`PATCH` Patch a service on a gateway group `/apisix/admin/services/{service_id}`](#fallback-operation-0-5) Partially update a service via JSON Patch (RFC 6902) through an APISIX Admin API compatible endpoint under /apisix/admin/. Use this for targeted field changes. Required IAM Permission: Action gateway:UpdatePublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`GET` List all GraphQL cost decorations in a service on a gateway group `/apisix/admin/services/{service_id}/graphql_cost_decorations`](#fallback-operation-0-6) List the GraphQL cost decorations attached to a service through an APISIX Admin API compatible endpoint under /apisix/admin/. Together they are the service's GraphQL cost model. Required IAM Permission: Action gateway:GetPublishedService, Resource... * [`POST` Create a GraphQL cost decoration in a service on a gateway group `/apisix/admin/services/{service_id}/graphql_cost_decorations`](#fallback-operation-0-7) Create a GraphQL cost decoration within a service through an APISIX Admin API compatible endpoint under /apisix/admin/. A decoration gives one position in the service's GraphQL schema graph a weight, which the graphql-limit-count plugin uses to compute a query's cost. A field... * [`GET` Get a GraphQL cost decoration in a service on a gateway group `/apisix/admin/services/{service_id}/graphql_cost_decorations/{graphql_cost_decoration_id}`](#fallback-operation-0-8) Get one GraphQL cost decoration in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. Required IAM Permission: Action gateway:GetPublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`PUT` Update a GraphQL cost decoration in a service on a gateway group `/apisix/admin/services/{service_id}/graphql_cost_decorations/{graphql_cost_decoration_id}`](#fallback-operation-0-9) Fully update a GraphQL cost decoration through an APISIX Admin API compatible endpoint under /apisix/admin/. Submit the complete decoration object to replace the existing one. Required IAM Permission: Action gateway:UpdatePublishedService, Resource... * [`DELETE` Delete a GraphQL cost decoration in a service on a gateway group `/apisix/admin/services/{service_id}/graphql_cost_decorations/{graphql_cost_decoration_id}`](#fallback-operation-0-10) Delete a GraphQL cost decoration from a service through an APISIX Admin API compatible endpoint under /apisix/admin/. Deleting the service reclaims its decorations automatically. Required IAM Permission: Action gateway:UpdatePublishedService, Resource... ## Route * [`GET` List all routes in a service `/apisix/admin/routes`](#fallback-operation-1-0) List routes in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. Use pagination and filters to inspect APISIX-compatible route entries. Required IAM Permission: Action gateway:GetPublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`POST` Create a route in a service on a gateway group `/apisix/admin/routes`](#fallback-operation-1-1) Create a route in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. This adds APISIX-formatted HTTP routing rules on shared dashboard resources. Required IAM Permission: Action gateway:UpdatePublishedService, Resource... * [`GET` Get a route in a service on a gateway group `/apisix/admin/routes/{route_id}`](#fallback-operation-1-2) Get one route in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. This returns the route with APISIX field format. Required IAM Permission: Action gateway:GetPublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`PUT` Update a route in a service on a gateway group `/apisix/admin/routes/{route_id}`](#fallback-operation-1-3) Fully update a route in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. Use this when replacing the entire route object. Required IAM Permission: Action gateway:UpdatePublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`DELETE` Delete a route in a service on a gateway group `/apisix/admin/routes/{route_id}`](#fallback-operation-1-4) Delete a route from a service through an APISIX Admin API compatible endpoint under /apisix/admin/. Requests that matched this rule will no longer be routed by it. Required IAM Permission: Action gateway:UpdatePublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`PATCH` Patch a route in a service on a gateway group `/apisix/admin/routes/{route_id}`](#fallback-operation-1-5) Partially update a route in a service via JSON Patch (RFC 6902) through an APISIX Admin API compatible endpoint under /apisix/admin/. Required IAM Permission: Action gateway:UpdatePublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s ## Stream Route * [`GET` List all stream routes in a service on a gateway group `/apisix/admin/stream_routes`](#fallback-operation-2-0) List stream routes in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. Use this for visibility into current L4 routing rules. Required IAM Permission: Action gateway:GetPublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`POST` Create a stream route in a service on a gateway group `/apisix/admin/stream_routes`](#fallback-operation-2-1) Create a stream route in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. This configures APISIX-style TCP/UDP traffic matching rules. Required IAM Permission: Action gateway:UpdatePublishedService, Resource... * [`GET` Get a stream route in a service on a gateway group `/apisix/admin/stream_routes/{stream_route_id}`](#fallback-operation-2-2) Get one stream route in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. The returned object matches APISIX stream-route conventions. Required IAM Permission: Action gateway:GetPublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`PUT` Update a stream route in a service on a gateway group `/apisix/admin/stream_routes/{stream_route_id}`](#fallback-operation-2-3) Fully update a stream route in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. This replaces the existing stream-route configuration. Required IAM Permission: Action gateway:UpdatePublishedService, Resource... * [`DELETE` Delete a stream route in a service on a gateway group `/apisix/admin/stream_routes/{stream_route_id}`](#fallback-operation-2-4) Delete a stream route from a service through an APISIX Admin API compatible endpoint under /apisix/admin/. This removes a specific L4 route while preserving other service objects. Required IAM Permission: Action gateway:UpdatePublishedService, Resource... ## Upstream * [`GET` List all upstreams in a service on a gateway group `/apisix/admin/services/{service_id}/upstreams`](#fallback-operation-3-0) List upstreams attached to a service through an APISIX Admin API compatible endpoint under /apisix/admin/. Use this to inspect backend pools and their current settings. Required IAM Permission: Action gateway:GetPublishedService, Resource... * [`POST` Create an upstream in a service on a gateway group `/apisix/admin/services/{service_id}/upstreams`](#fallback-operation-3-1) Create an upstream within a service through an APISIX Admin API compatible endpoint under /apisix/admin/. This adds backend target configuration in APISIX format. Required IAM Permission: Action gateway:UpdatePublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`GET` Get an upstream in a service on a gateway group `/apisix/admin/services/{service_id}/upstreams/{upstream_id}`](#fallback-operation-3-2) Get one upstream in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. The response uses APISIX-style upstream structure. Required IAM Permission: Action gateway:GetPublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`PUT` Update an upstream in a service on a gateway group `/apisix/admin/services/{service_id}/upstreams/{upstream_id}`](#fallback-operation-3-3) Fully update an upstream in a service through an APISIX Admin API compatible endpoint under /apisix/admin/. Submit the complete upstream object to replace existing configuration. Required IAM Permission: Action gateway:UpdatePublishedService, Resource... * [`DELETE` Delete an upstream in a service on a gateway group `/apisix/admin/services/{service_id}/upstreams/{upstream_id}`](#fallback-operation-3-4) Delete an upstream from a service through an APISIX Admin API compatible endpoint under /apisix/admin/. This updates service backend routing targets without deleting the service itself. Required IAM Permission: Action gateway:UpdatePublishedService, Resource... * [`PATCH` Patch an upstream in a service on a gateway group `/apisix/admin/services/{service_id}/upstreams/{upstream_id}`](#fallback-operation-3-5) Partially update an upstream in a service via JSON Patch (RFC 6902) through an APISIX Admin API compatible endpoint under /apisix/admin/. Required IAM Permission: Action gateway:UpdatePublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`GET` Get healthcheck status for the upstream of a service on a gateway group, if upstream\_id is not provided, get healthcheck status for default upstream of this service `/api/gateway_groups/{gateway_group_id}/services/{apisix_service_id}/healthcheck`](#fallback-operation-3-6) Retrieve upstream node health check results for a service in a gateway group. If no upstream ID is provided, the status of the service's default upstream is returned. Required IAM Permission: Action gateway:GetPublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s ## OpenAPI * [`GET` Get the OpenAPI Specification of a service `/api/gateway_groups/{gateway_group_id}/services/{apisix_service_id}/oas`](#fallback-operation-4-0) Get the OAS document for a service. Required IAM Permission: Action gateway:GetPublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`PUT` Update the OpenAPI Specification of a service `/api/gateway_groups/{gateway_group_id}/services/{apisix_service_id}/oas`](#fallback-operation-4-1) Update the OAS document for a service. Required IAM Permission: Action gateway:UpdatePublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/%s * [`POST` Generate an OpenAPI specification from services in a gateway group `/api/gateway_groups/{gateway_group_id}/services/export`](#fallback-operation-4-2) Export service definitions from a specific gateway group as an OpenAPI 3.0 specification document. Required IAM Permission: Action gateway:GetPublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/ * [`POST` Import services based on OpenAPI Specification `/api/import/services`](#fallback-operation-4-3) Import an OpenAPI specification directly into services for a gateway group. This operation creates runtime service resources scoped to the target gateway group. Required IAM Permission: Action gateway:CreatePublishedService, Resource arn:api7:gateway:gatewaygroup/%s/service/ * [`PUT` Convert OpenAPI Specification to service and route resources `/api/openapi/convert`](#fallback-operation-4-4) Convert a given OpenAPI Specification into service and route resource structures without creating those resources. Use this endpoint for preview, validation, and transformation workflows before import. ## Consumer * [`GET` List all consumers on a gateway group `/apisix/admin/consumers`](#fallback-operation-5-0) IAM Action: gateway:GetConsumer, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/%s * [`POST` Create a consumer on a gateway group `/apisix/admin/consumers`](#fallback-operation-5-1) IAM Action: gateway:CreateConsumer, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/ * [`GET` Get a consumer on a gateway group `/apisix/admin/consumers/{username}`](#fallback-operation-5-2) IAM Action: gateway:GetConsumer, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/%s * [`PUT` Update a consumer on a gateway group `/apisix/admin/consumers/{username}`](#fallback-operation-5-3) IAM Action: gateway:UpdateConsumer, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/%s * [`DELETE` Delete a consumer `/apisix/admin/consumers/{username}`](#fallback-operation-5-4) IAM Action: gateway:DeleteConsumer, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/%s * [`PATCH` Update a consumer on a gateway group `/apisix/admin/consumers/{username}`](#fallback-operation-5-5) IAM Action: gateway:UpdateConsumer, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/%s * [`GET` List all consumer credentials on a gateway group `/apisix/admin/consumers/{username}/credentials`](#fallback-operation-5-6) IAM Action: gateway:GetConsumerCredential, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/%s * [`POST` Create a consumer credential on a gateway group `/apisix/admin/consumers/{username}/credentials`](#fallback-operation-5-7) IAM Action: gateway:CreateConsumerCredential, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/%s * [`GET` Get a consumer credential on a gateway group `/apisix/admin/consumers/{username}/credentials/{credential_id}`](#fallback-operation-5-8) IAM Action: gateway:GetConsumerCredential, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/%s * [`PUT` Update a consumer credential on a gateway group `/apisix/admin/consumers/{username}/credentials/{credential_id}`](#fallback-operation-5-9) IAM Action: gateway:UpdateConsumerCredential, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/%s * [`DELETE` Delete a consumer credential `/apisix/admin/consumers/{username}/credentials/{credential_id}`](#fallback-operation-5-10) IAM Action: gateway:DeleteConsumerCredential, Resource: arn:api7:gateway:gatewaygroup/%s/consumer/%s ## Gateway Group * [`POST` Check service route conflicts in a gateway group `/api/gateway_groups/{gateway_group_id}/services/conflict_check`](#fallback-operation-6-0) Check for duplicate or overlapping routes among services within a gateway group. * [`GET` List all gateway groups `/api/gateway_groups`](#fallback-operation-6-1) IAM Action: gateway:GetGatewayGroup, Resource: arn:api7:gateway:gatewaygroup/%s * [`POST` Create a gateway group `/api/gateway_groups`](#fallback-operation-6-2) IAM Action: gateway:CreateGatewayGroup, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Count resources by type in each gateway group `/api/gateway_groups/count/{resource_type}`](#fallback-operation-6-3) * [`GET` List SSL Usage in a gateway group `/api/gateway_groups/{gateway_group_id}/ssls/{ssl_id}/usage`](#fallback-operation-6-4) IAM Action: gateway:GetSSLCertificate, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List SNI Usage in a gateway group `/api/gateway_groups/{gateway_group_id}/snis/{sni_id}/usage`](#fallback-operation-6-5) IAM Action: gateway:GetSNI, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List Certificate Usage in a gateway group `/api/gateway_groups/{gateway_group_id}/certificates/{certificate_id}/usage`](#fallback-operation-6-6) IAM Action: gateway:GetCertificate, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List CA Certificate Usage in a gateway group `/api/gateway_groups/{gateway_group_id}/ca_certificates/{ca_certificate_id}/usage`](#fallback-operation-6-7) IAM Action: gateway:GetCACertificate, Resource: arn:api7:gateway:gatewaygroup/%s * [`POST` Check if a certificate exists in a gateway group `/api/gateway_groups/{gateway_group_id}/certificates/exists`](#fallback-operation-6-8) IAM Action: gateway:GetCACertificate, Resource: arn:api7:gateway:gatewaygroup/%s * [`POST` Check if a CA certificate exists in a gateway group `/api/gateway_groups/{gateway_group_id}/ca_certificates/exists`](#fallback-operation-6-9) IAM Action: gateway:GetCACertificate, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List Secret Provider Usage in a gateway group `/api/gateway_groups/{gateway_group_id}/secret_providers/{secret_provider}/{secret_provider_id}/usage`](#fallback-operation-6-10) IAM Action: gateway:GetSecretProvider, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Get a gateway group `/api/gateway_groups/{gateway_group_id}`](#fallback-operation-6-11) IAM Action: gateway:GetGatewayGroup, Resource: arn:api7:gateway:gatewaygroup/%s * [`PUT` Update a gateway group `/api/gateway_groups/{gateway_group_id}`](#fallback-operation-6-12) IAM Action: gateway:UpdateGatewayGroup, Resource: arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete a gateway group `/api/gateway_groups/{gateway_group_id}`](#fallback-operation-6-13) IAM Action: gateway:DeleteGatewayGroup, Resource: arn:api7:gateway:gatewaygroup/%s * [`POST` Get resource paths `/api/gateway_groups/{gateway_group_id}/resource_paths`](#fallback-operation-6-14) Resolve resource IDs to ordered business resource paths on demand. IAM Action: gateway:GetGatewayGroup, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List monitored shared dictionaries of a gateway group `/api/gateway_groups/{gateway_group_id}/shared_dict_names`](#fallback-operation-6-15) List the shared dictionaries (shared memory zones) worth monitoring for a gateway group, i.e. every shared dict the data plane reports metrics for minus the ones on the alert denylist (lock dicts and LRU caches). * [`GET` Get the admin key for a gateway group. `/api/gateway_groups/{gateway_group_id}/admin_key`](#fallback-operation-6-16) IAM Action: gateway:GetAdminKey, Resource: arn:api7:gateway:gatewaygroup/%s * [`PUT` Generate an admin key for a gateway group `/api/gateway_groups/{gateway_group_id}/admin_key`](#fallback-operation-6-17) IAM Action: gateway:UpdateGatewayGroup, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Generate a script to install gateway API resources for ingress gateway group `/api/gateway_groups/{gateway_group_id}/ingress/script`](#fallback-operation-6-18) IAM Action: gateway:GetAdminKey, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Generate script to install the gateway api resources for ingress gateway group. `/api/gateway_groups/{gateway_group_id}/ingress/step1`](#fallback-operation-6-19) ## Gateway Instance * [`GET` List all gateway instances of all gateway groups `/api/instances`](#fallback-operation-7-0) IAM Action: gateway:GetGatewayInstance, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List all gateway instances on a gateway group `/api/gateway_groups/{gateway_group_id}/instances`](#fallback-operation-7-1) IAM Action: gateway:GetGatewayInstance, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Count the number of gateway instances by status in a gateway group `/api/instances/count/{field}`](#fallback-operation-7-2) IAM Action: gateway:GetGatewayInstance, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List gateway instances cores of all gateway groups `/api/instances/cores`](#fallback-operation-7-3) IAM Action: gateway:GetGatewayInstanceCore, Resource: arn:api7:gateway:gatewaygroup/ * [`GET` Export the gateway instance core usage `/api/instances/cores_usages/export`](#fallback-operation-7-4) The gateway instance’s core usage is exported hourly within the specified time interval. IAM Action: gateway:GetGatewayInstanceCore, Resource: arn:api7:gateway:gatewaygroup/ * [`GET` Generate a script to install the gateway instance by Docker `/api/gateway_groups/{gateway_group_id}/deployment/docker`](#fallback-operation-7-5) IAM Action: gateway:CreateGatewayInstance, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Generate a script to install the gateway instance by Docker Compose `/api/gateway_groups/{gateway_group_id}/deployment/docker-compose`](#fallback-operation-7-6) IAM Action: gateway:CreateGatewayInstance, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Generate a script to install the gateway instance by Helm in Kubernetes `/api/gateway_groups/{gateway_group_id}/deployment/helm/script`](#fallback-operation-7-7) IAM Action: gateway:CreateGatewayInstance, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Generate a values file for the gateway's Kubernetes Helm chart `/api/gateway_groups/{gateway_group_id}/deployment/helm/yaml`](#fallback-operation-7-8) IAM Action: gateway:CreateGatewayInstance, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Generate a script to install the gateway instance by RPM `/api/gateway_groups/{gateway_group_id}/deployment/rpm`](#fallback-operation-7-9) IAM Action: gateway:CreateGatewayInstance, Resource: arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete gateway instance `/api/gateway_groups/{gateway_group_id}/instances/{gateway_instance_id}`](#fallback-operation-7-10) IAM Action: gateway:DeleteGatewayInstance, Resource: arn:api7:gateway:gatewaygroup/%s * [`POST` Issue a data plane certificate on a gateway group `/api/gateway_groups/{gateway_group_id}/dp_client_certificates`](#fallback-operation-7-11) Issue a client TLS certificate for data plane instances in the specified gateway group to authenticate with the control plane. Use this during gateway bootstrap or certificate rotation. Required IAM Permission: Action gateway:CreateGatewayInstance, Resource... * [`POST` Create a token for all gateway instances in a gateway group `/api/gateway_groups/{gateway_group_id}/instance_token`](#fallback-operation-7-12) IAM Action: gateway:CreateGatewayInstance, Resource: arn:api7:gateway:gatewaygroup/%s ## SSL * [`GET` List all SSL certificates on a gateway group `/apisix/admin/ssls`](#fallback-operation-8-0) IAM Action: gateway:GetSSLCertificate, Resource: arn:api7:gateway:gatewaygroup/%s * [`POST` Create an SSL certificate `/apisix/admin/ssls`](#fallback-operation-8-1) IAM Action: gateway:CreateSSLCertificate, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Get an SSL certificate on a gateway group `/apisix/admin/ssls/{ssl_id}`](#fallback-operation-8-2) IAM Action: gateway:GetSSLCertificate, Resource: arn:api7:gateway:gatewaygroup/%s * [`PUT` Update an SSL certificate on a gateway group `/apisix/admin/ssls/{ssl_id}`](#fallback-operation-8-3) IAM Action: gateway:UpdateSSLCertificate, Resource: arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete an SSL certificate on a gateway group `/apisix/admin/ssls/{ssl_id}`](#fallback-operation-8-4) IAM Action: gateway:DeleteSSLCertificate, Resource: arn:api7:gateway:gatewaygroup/%s * [`PUT` Parse an SSL certificate `/api/parse_certificate`](#fallback-operation-8-5) * [`PUT` Validate an SSL certificate and key `/api/validate_cert_key`](#fallback-operation-8-6) ## Certificate * [`GET` List all certificates on a gateway group `/apisix/admin/certificates`](#fallback-operation-9-0) List TLS server certificates configured in the gateway group. Use filters such as labels, related SNI, expiration time, and search keywords to find certificates for rotation or troubleshooting. Required IAM Permission: Action gateway:GetCertificate, Resource... * [`POST` Create a certificate `/apisix/admin/certificates`](#fallback-operation-9-1) Create a TLS server certificate for the gateway group, including the certificate chain and private key used for HTTPS termination. Required IAM Permission: Action gateway:CreateCertificate, Resource arn:api7:gateway:gatewaygroup/%s * [`GET` Get a certificate on a gateway group `/apisix/admin/certificates/{certificate_id}`](#fallback-operation-9-2) Retrieve details of a specific TLS server certificate in the gateway group, including its configured metadata and bindings. Required IAM Permission: Action gateway:GetCertificate, Resource arn:api7:gateway:gatewaygroup/%s * [`PUT` Update a certificate on a gateway group `/apisix/admin/certificates/{certificate_id}`](#fallback-operation-9-3) Replace the full configuration of an existing TLS server certificate in the gateway group. This operation may impact HTTPS traffic using this certificate. Required IAM Permission: Action gateway:UpdateCertificate, Resource arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete a certificate on a gateway group `/apisix/admin/certificates/{certificate_id}`](#fallback-operation-9-4) Delete a TLS server certificate from the gateway group. Ensure no active SNI or route still depends on this certificate before removal. Required IAM Permission: Action gateway:DeleteCertificate, Resource arn:api7:gateway:gatewaygroup/%s * [`PATCH` Patch a certificate on a gateway group `/apisix/admin/certificates/{certificate_id}`](#fallback-operation-9-5) Partially update fields of a TLS server certificate using JSON Patch (RFC 6902). Use this when adjusting selected attributes without replacing the entire certificate object. Required IAM Permission: Action gateway:UpdateCertificate, Resource arn:api7:gateway:gatewaygroup/%s ## CACertificate * [`GET` List all CA certificates on a gateway group `/apisix/admin/ca_certificates`](#fallback-operation-10-0) List CA certificates configured for mTLS client certificate verification in the gateway group. Use query filters to locate certificates by labels, expiration, or associated SNI. Required IAM Permission: Action gateway:GetCertificate, Resource arn:api7:gateway:gatewaygroup/%s * [`POST` Create a CA certificate `/apisix/admin/ca_certificates`](#fallback-operation-10-1) Create a CA certificate used to verify client certificates during mTLS handshakes in the gateway group. Required IAM Permission: Action gateway:CreateCertificate, Resource arn:api7:gateway:gatewaygroup/%s * [`GET` Get a CA certificate on a gateway group `/apisix/admin/ca_certificates/{ca_certificate_id}`](#fallback-operation-10-2) Retrieve details of a specific CA certificate configured in the gateway group for mTLS validation. Required IAM Permission: Action gateway:GetCertificate, Resource arn:api7:gateway:gatewaygroup/%s * [`PUT` Update a CA certificate on a gateway group `/apisix/admin/ca_certificates/{ca_certificate_id}`](#fallback-operation-10-3) Replace the full configuration of an existing CA certificate used for client certificate verification. Changes can affect mTLS authentication for related traffic. Required IAM Permission: Action gateway:UpdateCertificate, Resource arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete a CA certificate on a gateway group `/apisix/admin/ca_certificates/{ca_certificate_id}`](#fallback-operation-10-4) Delete a CA certificate from the gateway group. Confirm no active mTLS SNI or client verification flow still relies on this certificate. Required IAM Permission: Action gateway:DeleteCertificate, Resource arn:api7:gateway:gatewaygroup/%s * [`PATCH` Patch a CA certificate on a gateway group `/apisix/admin/ca_certificates/{ca_certificate_id}`](#fallback-operation-10-5) Partially update a CA certificate with JSON Patch (RFC 6902) operations. Use patching for targeted changes without sending the full CA certificate payload. Required IAM Permission: Action gateway:UpdateCertificate, Resource arn:api7:gateway:gatewaygroup/%s ## SNI * [`GET` List all SNIs on a gateway group `/apisix/admin/snis`](#fallback-operation-11-0) List SNI configurations for the gateway group, including hostname and mTLS related settings. Use filters to locate entries by domain, labels, or mTLS enablement. Required IAM Permission: Action gateway:GetSNI, Resource arn:api7:gateway:gatewaygroup/%s * [`POST` Create an SNI `/apisix/admin/snis`](#fallback-operation-11-1) Create an SNI entry that maps one or more hostnames to TLS certificates in the gateway group. This allows a single gateway to serve multiple domains over HTTPS. Required IAM Permission: Action gateway:CreateSNI, Resource arn:api7:gateway:gatewaygroup/%s * [`GET` Get an SNI on a gateway group `/apisix/admin/snis/{sni_id}`](#fallback-operation-11-2) Retrieve the configuration of a specific SNI entry in the gateway group, including bound domains and certificate references. Required IAM Permission: Action gateway:GetSNI, Resource arn:api7:gateway:gatewaygroup/%s * [`PUT` Update an SNI on a gateway group `/apisix/admin/snis/{sni_id}`](#fallback-operation-11-3) Replace an existing SNI configuration in full, such as domain mappings, certificate bindings, or mTLS options. Updates take effect on TLS handshakes for matching hostnames. Required IAM Permission: Action gateway:UpdateSNI, Resource arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete an SNI on a gateway group `/apisix/admin/snis/{sni_id}`](#fallback-operation-11-4) Delete an SNI entry from the gateway group to remove its hostname-to- certificate mapping. Verify that traffic for those domains has a replacement SNI before deletion. Required IAM Permission: Action gateway:DeleteSNI, Resource arn:api7:gateway:gatewaygroup/%s * [`PATCH` Patch an SNI on a gateway group `/apisix/admin/snis/{sni_id}`](#fallback-operation-11-5) Partially update selected SNI fields using JSON Patch (RFC 6902). This is useful for incremental hostname or certificate adjustments without replacing the entire SNI object. Required IAM Permission: Action gateway:UpdateSNI, Resource arn:api7:gateway:gatewaygroup/%s ## Global Rule * [`GET` List all global rules on a gateway group `/apisix/admin/global_rules`](#fallback-operation-12-0) List global plugin rules configured for the gateway group. This helps audit request-wide policies and understand which plugins are enforced universally. Required IAM Permission: Action gateway:GetGlobalPluginRule, Resource arn:api7:gateway:gatewaygroup/%s * [`POST` Create a global rule on a gateway group `/apisix/admin/global_rules`](#fallback-operation-12-1) Create a global rule that applies plugin configuration to all requests in the gateway group. Use this for cross-cutting behavior such as global authentication, logging, or rate controls. Required IAM Permission: Action gateway:CreateGlobalPluginRule, Resource... * [`GET` Get a global rule on a gateway group `/apisix/admin/global_rules/{global_rule_id}`](#fallback-operation-12-2) Retrieve one global rule in the gateway group, including its plugin configuration and execution settings. Required IAM Permission: Action gateway:GetGlobalPluginRule, Resource arn:api7:gateway:gatewaygroup/%s * [`PUT` Update a global rule on a gateway group `/apisix/admin/global_rules/{global_rule_id}`](#fallback-operation-12-3) Replace an existing global rule configuration in full for the gateway group. Changes immediately affect all matching traffic because global rules are applied gateway-wide. Required IAM Permission: Action gateway:UpdateGlobalPluginRule, Resource arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete a global rule on a gateway group `/apisix/admin/global_rules/{global_rule_id}`](#fallback-operation-12-4) Delete a global rule from the gateway group, removing its plugin behavior from all requests. Review related route or service plugin configuration if equivalent controls are still required. Required IAM Permission: Action gateway:DeleteGlobalPluginRule, Resource... ## Plugin * [`GET` Get all plugin schemas and priorities `/apisix/admin/plugins`](#fallback-operation-13-0) * [`GET` List all plugin names `/apisix/admin/plugins/list`](#fallback-operation-13-1) * [`GET` Get schema definition of a plugin `/apisix/admin/plugins/{plugin_name}`](#fallback-operation-13-2) Get schema definition of a plugin, including plugin meta properties and plugin properties. The endpoint returns the same response as the /apisix/admin/schema/plugins/{plugin name} endpoint when scope is not set. * [`GET` Get all plugin details `/api/plugins`](#fallback-operation-13-3) * [`GET` List all plugin catalogs `/api/plugins/catalogs`](#fallback-operation-13-4) * [`GET` Get the usage of a plugin in a gateway group `/api/gateway_groups/{gateway_group_id}/plugins/{plugin_name}/usage`](#fallback-operation-13-5) List the resources of the gateway group that reference the plugin. A custom plugin belongs to a gateway group, so its usage is a question about one group. Required IAM Permission: Action gateway:GetCustomPlugin, Resource arn:api7:gateway:gatewaygroup/{gateway group id} ## Plugin Metadata * [`GET` List all plugin metadata on a gateway group `/apisix/admin/plugin_metadata`](#fallback-operation-14-0) List plugin metadata objects configured in the gateway group. Plugin metadata provides shared settings for a plugin type across all its instances. Required IAM Permission: Action gateway:GetPluginMetadata, Resource arn:api7:gateway:gatewaygroup/%s * [`GET` Get a plugin metadata on a gateway group `/apisix/admin/plugin_metadata/{plugin_name}`](#fallback-operation-14-1) Retrieve metadata for a specific plugin type in the gateway group. You can optionally request default metadata values for comparison. Required IAM Permission: Action gateway:GetPluginMetadata, Resource arn:api7:gateway:gatewaygroup/%s * [`PUT` Update a plugin metadata on a gateway group `/apisix/admin/plugin_metadata/{plugin_name}`](#fallback-operation-14-2) Update the shared metadata configuration for a specific plugin type in the gateway group. The change affects behavior of all plugin instances that consume this metadata. Required IAM Permission: Action gateway:UpdatePluginMetadata, Resource arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete a plugin metadata on a gateway group `/apisix/admin/plugin_metadata/{plugin_name}`](#fallback-operation-14-3) Delete metadata for a specific plugin type in the gateway group and revert to plugin defaults where applicable. Validate downstream plugin behavior after removal. Required IAM Permission: Action gateway:DeletePluginMetadata, Resource arn:api7:gateway:gatewaygroup/%s * [`GET` Get the default value of a plugin metadata `/apisix/admin/plugin_metadata/{plugin_name}/default`](#fallback-operation-14-4) ## Custom Plugin * [`GET` List the custom plugins of a gateway group `/api/gateway_groups/{gateway_group_id}/custom_plugins`](#fallback-operation-15-0) List the custom plugins the gateway group runs. Use pagination and search parameters to locate plugins by name. Required IAM Permission: Action gateway:GetCustomPlugin, Resource arn:api7:gateway:gatewaygroup/{gateway group id} * [`GET` Get a custom plugin `/api/gateway_groups/{gateway_group_id}/custom_plugins/{custom_plugin_name}`](#fallback-operation-15-1) Retrieve the code and metadata of the custom plugin the gateway group runs under this name. Required IAM Permission: Action gateway:GetCustomPlugin, Resource arn:api7:gateway:gatewaygroup/{gateway group id} * [`PUT` Create or replace a custom plugin `/api/gateway_groups/{gateway_group_id}/custom_plugins/{custom_plugin_name}`](#fallback-operation-15-2) Upload a custom plugin to the gateway group, creating it or replacing the code the group is running under this name. The plugin name declared by the uploaded code must match the name in the path. Only this gateway group is affected: the same plugin name in another gateway... * [`DELETE` Delete a custom plugin `/api/gateway_groups/{gateway_group_id}/custom_plugins/{custom_plugin_name}`](#fallback-operation-15-3) Remove the custom plugin from the gateway group. The plugin must no longer be referenced by the resources of this gateway group. Required IAM Permission: Action gateway:DeleteCustomPlugin, Resource arn:api7:gateway:gatewaygroup/{gateway group id} * [`PUT` Parse custom plugin code `/api/gateway_groups/{gateway_group_id}/custom_plugins/code/parse`](#fallback-operation-15-4) Parse and validate custom plugin source payload to extract plugin metadata and detect structural issues before uploading a plugin. Required IAM Permission: Action gateway:UpdateCustomPlugin, Resource arn:api7:gateway:gatewaygroup/{gateway group id} ## Secret Provider * [`GET` List all secret providers on a gateway group `/apisix/admin/secret_providers`](#fallback-operation-16-0) List secret provider integrations configured for the gateway group, such as Vault or cloud secret managers. Use this view to audit external secret backends available for plugin and route configurations. Required IAM Permission: Action gateway:GetSecretProvider, Resource... * [`GET` Get a secret provider on a gateway group `/apisix/admin/secret_providers/{secret_provider}/{secret_provider_id}`](#fallback-operation-16-1) Retrieve details of one secret provider integration in the gateway group, including provider-specific connection settings. Required IAM Permission: Action gateway:GetSecretProvider, Resource arn:api7:gateway:gatewaygroup/%s/secret provider/%s * [`PUT` Update a secret provider on a gateway group `/apisix/admin/secret_providers/{secret_provider}/{secret_provider_id}`](#fallback-operation-16-2) Create or replace the configuration of a secret provider integration in the gateway group. This controls how the gateway resolves externally managed secrets referenced by runtime configs. Required IAM Permission: Action gateway:PutSecretProvider, Resource... * [`DELETE` Delete a secret provider on a gateway group `/apisix/admin/secret_providers/{secret_provider}/{secret_provider_id}`](#fallback-operation-16-3) Delete a secret provider integration from the gateway group. Ensure no plugin or resource still references secrets from this provider before removal. Required IAM Permission: Action gateway:DeleteSecretProvider, Resource arn:api7:gateway:gatewaygroup/%s/secret provider/%s ## Proto * [`GET` List all protos on a gateway group `/apisix/admin/protos`](#fallback-operation-17-0) List Protocol Buffers definitions stored in the gateway group. These proto files are used by gRPC-transcode related configurations to map REST calls to gRPC methods. Required IAM Permission: Action gateway:GetProto, Resource arn:api7:gateway:gatewaygroup/%s * [`POST` Create a proto on a gateway group `/apisix/admin/protos`](#fallback-operation-17-1) Upload a new .proto definition to the gateway group for gRPC transcoding scenarios. Ensure package and service definitions align with upstream gRPC services. Required IAM Permission: Action gateway:CreateProto, Resource arn:api7:gateway:gatewaygroup/%s * [`GET` Get a proto on a gateway group `/apisix/admin/protos/{proto_id}`](#fallback-operation-17-2) Retrieve one proto definition from the gateway group, including its current content and metadata. Required IAM Permission: Action gateway:GetProto, Resource arn:api7:gateway:gatewaygroup/%s * [`PUT` Update a proto on a gateway group `/apisix/admin/protos/{proto_id}`](#fallback-operation-17-3) Replace an existing proto definition in the gateway group. After updates, verify dependent gRPC-transcode routes still match the revised service and method signatures. Required IAM Permission: Action gateway:UpdateProto, Resource arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete a proto on a gateway group `/apisix/admin/protos/{proto_id}`](#fallback-operation-17-4) Delete a proto definition from the gateway group. Check for any gRPC-transcode plugin configurations that still reference this proto before deletion. Required IAM Permission: Action gateway:DeleteProto, Resource arn:api7:gateway:gatewaygroup/%s ## Service Registry * [`GET` List all service registry connections on a gateway group `/api/gateway_groups/{gateway_group_id}/service_registries`](#fallback-operation-18-0) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`POST` Create a service registry connection on a gateway group `/api/gateway_groups/{gateway_group_id}/service_registries`](#fallback-operation-18-1) IAM Action: gateway:ConnectServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Get a service registry connection on a gateway group `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}`](#fallback-operation-18-2) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`PUT` Update a service registry connection on a gateway group `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}`](#fallback-operation-18-3) IAM Action: gateway:UpdateServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete a service registry connection on a gateway group `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}`](#fallback-operation-18-4) IAM Action: gateway:DisconnectServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List all services connected to a service registry `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}/connected_services`](#fallback-operation-18-5) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List all internal services in a Kubernetes service registry `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}/kubernetes/internal_services`](#fallback-operation-18-6) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List all namespaces in a Nacos service registry `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}/nacos/namespaces`](#fallback-operation-18-7) * [`GET` List all groups in a Nacos namespace `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}/nacos/namespaces/{nacos_namespace}/groups`](#fallback-operation-18-8) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List all internal services in a Nacos group `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}/nacos/namespaces/{nacos_namespace}/groups/{nacos_group}/services`](#fallback-operation-18-9) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Get all instance metadata of a Nacos services registry `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}/nacos/namespaces/{nacos_namespace}/groups/{nacos_group}/services/{nacos_service}/instances_metadata`](#fallback-operation-18-10) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List all datacenters in a Consul service registry `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}/consul/datacenters`](#fallback-operation-18-11) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List all services in a Consul datacenter `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}/consul/datacenters/{consul_datacenter}/services`](#fallback-operation-18-12) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Get all instance metadata of a Consul service `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}/consul/datacenters/{consul_datacenter}/services/{consul_service}/instances_metadata`](#fallback-operation-18-13) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Get health check history of a service registry connection on a gateway group `/api/gateway_groups/{gateway_group_id}/service_registries/{service_registry_id}/health_check_history`](#fallback-operation-18-14) IAM Action: gateway:GetServiceRegistry, Resource: arn:api7:gateway:gatewaygroup/%s ## User * [`GET` List all users `/api/users`](#fallback-operation-19-0) List dashboard user accounts with pagination and role-based filtering. Use this to audit who can authenticate by local credentials or configured SSO methods. Required IAM Permission: Action iam:GetUser, Resource arn:api7:iam:user/%s * [`GET` Get a user `/api/users/{user_id}`](#fallback-operation-19-1) Get one dashboard user account by ID to inspect account details and current state. Required IAM Permission: Action iam:GetUser, Resource arn:api7:iam:user/%s * [`PUT` Update basic attributes of a user `/api/users/{user_id}`](#fallback-operation-19-2) Update basic attributes of a specified dashboard user account. This is intended for administrative account maintenance. Required IAM Permission: Action iam:UpdateUser, Resource arn:api7:iam:user/%s * [`DELETE` Delete a user `/api/users/{user_id}`](#fallback-operation-19-3) Delete a dashboard user account and revoke its login access. Verify downstream ownership and role dependencies before removal. Required IAM Permission: Action iam:DeleteUser, Resource arn:api7:iam:user/%s * [`PUT` Update the user's permission boundaries `/api/users/{user_id}/boundaries`](#fallback-operation-19-4) Replace a user's permission boundaries using permission policy IDs. Boundaries constrain the maximum effective permissions a user can obtain. Required IAM Permission: Action iam:UpdateUserBoundary, Resource arn:api7:iam:user/%s * [`PUT` Reset the password to specific value `/api/users/{user_id}/password_reset`](#fallback-operation-19-5) Reset a specified user's password to an administrator-provided value. Use this for recovery or emergency credential rotation. Required IAM Permission: Action iam:ResetPassword, Resource arn:api7:iam:user/%s * [`PUT` Reset (disable) a user's two-factor authentication `/api/users/{user_id}/2fa_reset`](#fallback-operation-19-6) Administratively disable a specified user's two-factor authentication, for example when the user has lost their authenticator device. Required IAM Permission: Action iam:ResetTwoFactor, Resource arn:api7:iam:user/%s * [`GET` Get my user detail `/api/me`](#fallback-operation-19-7) Get details of the currently authenticated dashboard user. * [`PUT` Update my user profile `/api/me`](#fallback-operation-19-8) Update profile attributes of the currently authenticated dashboard user without changing role assignments. * [`PUT` Update the user email `/api/me/email`](#fallback-operation-19-9) Update the email address for the currently authenticated user account. * [`DELETE` Delete the user email `/api/me/email`](#fallback-operation-19-10) Remove the email address currently bound to the authenticated user account. * [`POST` Start two-factor (2FA) enrollment for the current user `/api/me/2fa/setup`](#fallback-operation-19-11) Generate a new TOTP secret for the currently authenticated user and return it together with a provisioning URI and QR code. The secret is not active until confirmed via the enable endpoint. * [`POST` Confirm and enable two-factor (2FA) for the current user `/api/me/2fa/enable`](#fallback-operation-19-12) Confirm a pending TOTP secret with a valid code and enable two-factor authentication. Returns one-time recovery codes that are shown only once. * [`POST` Disable two-factor (2FA) for the current user `/api/me/2fa/disable`](#fallback-operation-19-13) Disable two-factor authentication for the currently authenticated user after validating a current TOTP or recovery code. * [`POST` Invite a user `/api/invites`](#fallback-operation-19-14) Invite a new dashboard user account. The invitation initiates onboarding so the user can later sign in with supported authentication methods. Required IAM Permission: Action iam:InviteUser, Resource arn:api7:iam:user/ * [`PUT` Update my user password `/api/password`](#fallback-operation-19-15) Change the password for the currently authenticated user account. * [`POST` Log in to API7 Enterprise using the built-in username and password `/api/login`](#fallback-operation-19-16) Authenticate a dashboard user with built-in username and password credentials and create a session. * [`POST` Log out from API7 Enterprise using the built-in username and password `/api/logout`](#fallback-operation-19-17) Log out the current built-in authentication session and invalidate related session state. * [`PUT` Update assigned roles for a user `/api/users/{user_id}/assigned_roles`](#fallback-operation-19-18) Update role assignments for a user. Assigned roles determine the permission policies and effective access granted to that account. Required IAM Permission: Action iam:UpdateUserRole, Resource arn:api7:iam:user/%s * [`POST` Check if a user has permissions on specific resources `/api/allow_access`](#fallback-operation-19-19) Evaluate whether a user is allowed to perform specified actions on given resources based on current RBAC and policy configuration. * [`POST` Log in to API7 Enterprise using the LDAP username and password `/api/ldap/{login_option_id}/login`](#fallback-operation-19-20) Authenticate a user through the specified LDAP login option and create a dashboard session. * [`POST` Log out from API7 Enterprise using the LDAP username and password `/api/ldap/{login_option_id}/logout`](#fallback-operation-19-21) Log out a session established via LDAP authentication for the specified login option. * [`GET` Log in using the CAS provider `/api/cas/{login_option_id}/login`](#fallback-operation-19-22) Start or complete CAS login for the selected login option, including CAS ticket processing. * [`GET` Log out using the CAS provider `/api/cas/{login_option_id}/logout`](#fallback-operation-19-23) Start CAS logout for the selected login option and redirect to the configured post-logout target. * [`GET` Log in using the OIDC provider `/api/oidc/{login_option_id}/login`](#fallback-operation-19-24) Start OIDC authentication by redirecting to the configured OpenID Connect provider for the selected login option. * [`GET` Log in using the OIDC provider `/api/oidc/{login_option_id}/callback`](#fallback-operation-19-25) Process OIDC callback parameters and complete authentication for the selected login option. * [`GET` Log out using the OIDC provider `/api/oidc/{login_option_id}/logout`](#fallback-operation-19-26) Start OIDC logout flow and redirect through the provider logout endpoint for the selected login option. * [`GET` SAML login (redirect to IdP and call back to Dashboard) `/api/saml/{login_option_id}/login`](#fallback-operation-19-27) Start SAML 2.0 login by redirecting to the identity provider and preparing for ACS callback handling. * [`GET` SAML Logout (redirect to IdP and call back to Dashboard) `/api/saml/{login_option_id}/logout`](#fallback-operation-19-28) Start SAML 2.0 logout by redirecting through the identity provider for the selected login option. * [`POST` SAML ACS/SLO callback (from IdP to Dashboard) `/api/saml/{login_option_id}/acs`](#fallback-operation-19-29) Handle SAML ACS/SLO callback payloads posted by the identity provider for the selected login option. * [`GET` SAML ACS/SLO callback (from IdP to Dashboard) `/api/saml/{login_option_id}/slo`](#fallback-operation-19-30) Handle SAML ACS/SLO callback query parameters from the identity provider and finalize SAML sign-in or sign-out. * [`POST` SAML ACS/SLO callback (from IdP to Dashboard) `/api/saml/{login_option_id}/slo`](#fallback-operation-19-31) Handle SAML ACS/SLO callback payloads posted by the identity provider for the selected login option. * [`GET` SAML SP metadata `/api/saml/{login_option_id}/metadata`](#fallback-operation-19-32) Get SAML service provider metadata XML for the specified login option to help configure identity-provider trust. ## Role * [`GET` List all roles `/api/roles`](#fallback-operation-20-0) List RBAC roles with pagination and filtering. Results include built-in roles (Super Admin, Admin, Viewer) and custom roles. Required IAM Permission: Action iam:GetRole, Resource arn:api7:iam:role/%s * [`POST` Create a role `/api/roles`](#fallback-operation-20-1) Create a custom RBAC role used to group permission policies. The new role can then be assigned to users. Required IAM Permission: Action iam:CreateRole, Resource arn:api7:iam:role/ * [`GET` Get a role `/api/roles/{role_id}`](#fallback-operation-20-2) Get detailed information for a specific RBAC role. This includes role metadata used to grant permissions through user-role assignments. Required IAM Permission: Action iam:GetRole, Resource arn:api7:iam:role/%s * [`PUT` Update a role `/api/roles/{role_id}`](#fallback-operation-20-3) Update a role definition by ID. Changes affect all users currently assigned to the role. Required IAM Permission: Action iam:UpdateRole, Resource arn:api7:iam:role/%s * [`DELETE` Delete a role `/api/roles/{role_id}`](#fallback-operation-20-4) Delete a custom RBAC role from the organization. Review user assignments first to avoid unintended access loss. Required IAM Permission: Action iam:DeleteRole, Resource arn:api7:iam:role/%s * [`GET` List all permission policies attached to a role `/api/roles/{role_id}/permission_policies`](#fallback-operation-20-5) List permission policies attached to a role. This reveals the policy statements that determine the role's effective access. Required IAM Permission: Action iam:GetPermissionPolicy, Resource arn:api7:iam:permissionpolicy/%s * [`POST` Attach permission policies to a role `/api/roles/{role_id}/attach_permission_policies`](#fallback-operation-20-6) Attach permission policies to a role. Attached policies immediately affect all users assigned to that role. Required IAM Permission: Action iam:UpdateRole, Resource arn:api7:iam:role/%s * [`POST` Detach permission policies of a role `/api/roles/{role_id}/detach_permission_policies`](#fallback-operation-20-7) Detach permission policies from a role. Removing policies may reduce or revoke access for assigned users. Required IAM Permission: Action iam:UpdateRole, Resource arn:api7:iam:role/%s ## Permission Policy * [`GET` List all permission policies `/api/permission_policies`](#fallback-operation-21-0) List available permission policies in the organization with pagination and filters. Use this to choose policies for roles and boundaries. Required IAM Permission: Action iam:GetPermissionPolicy, Resource arn:api7:iam:permissionpolicy/%s * [`POST` Create a permission policy `/api/permission_policies`](#fallback-operation-21-1) Create a fine-grained permission policy with allow/deny statements over specific IAM actions and resource ARNs. Required IAM Permission: Action iam:CreatePermissionPolicy, Resource arn:api7:iam:permissionpolicy/ * [`GET` Get the permission policy authoring catalog `/api/permission_policies/metadata`](#fallback-operation-21-2) Return the static catalog used by the console Visual Editor to author permission policies: every supported resource type with its ARN templates, available list endpoint, condition keys, and the ordered list of actions classified by access level. The payload is purely... * [`GET` Get a permission policy `/api/permission_policies/{permission_policy_id}`](#fallback-operation-21-3) Get a permission policy by ID, including all statements and metadata. Required IAM Permission: Action iam:GetPermissionPolicy, Resource arn:api7:iam:permissionpolicy/%s * [`PUT` Update a permission policy `/api/permission_policies/{permission_policy_id}`](#fallback-operation-21-4) Update an existing permission policy definition. Changes apply to all roles or users that reference this policy. Required IAM Permission: Action iam:UpdatePermissionPolicy, Resource arn:api7:iam:permissionpolicy/%s * [`DELETE` Delete a permission policy `/api/permission_policies/{permission_policy_id}`](#fallback-operation-21-5) Delete a permission policy. Ensure references are reviewed first to prevent accidental permission breakage. Required IAM Permission: Action iam:DeletePermissionPolicy, Resource arn:api7:iam:permissionpolicy/%s * [`GET` List the Roles or Users that directly reference the Permission Policy `/api/permission_policies/{permission_policy_id}/references`](#fallback-operation-21-6) List roles and users that directly reference the specified permission policy. This helps estimate impact before policy changes. Required IAM Permission: Action iam:GetPermissionPolicy, Resource arn:api7:iam:permissionpolicy/%s ## Token * [`GET` List all tokens `/api/tokens`](#fallback-operation-22-0) List API access tokens created for programmatic dashboard API access, with pagination, ordering, and search filters. * [`POST` Create a token `/api/tokens`](#fallback-operation-22-1) Create a new API access token for programmatic, non-interactive access to dashboard APIs. You can configure token metadata and optional expiration at creation time. * [`GET` Get a token `/api/tokens/{token_id}`](#fallback-operation-22-2) Get details of an API access token by token ID. * [`PUT` Update a token `/api/tokens/{token_id}`](#fallback-operation-22-3) Update mutable properties of an API access token, such as expiration and metadata fields. * [`DELETE` Delete a token `/api/tokens/{token_id}`](#fallback-operation-22-4) Delete an API access token and immediately revoke its ability to authenticate API requests. * [`PUT` Regenerate a token `/api/tokens/{token_id}/regenerate`](#fallback-operation-22-5) Regenerate the secret value of an existing API access token. Clients must replace stored credentials after regeneration. ## Dashboard Login Option * [`GET` Get a login option `/api/login_options/{login_option_id}`](#fallback-operation-23-0) Get a login option configuration by ID. This returns settings for one external authentication integration such as OIDC, LDAP, SAML 2.0, or CAS. Required IAM Permission: Action iam:GetLoginOption, Resource arn:api7:iam:organization/ * [`PUT` Update a login option `/api/login_options/{login_option_id}`](#fallback-operation-23-1) Fully update a login option configuration with a complete protocol-specific payload. Required IAM Permission: Action iam:UpdateLoginOption, Resource arn:api7:iam:organization/ * [`DELETE` Delete a login option `/api/login_options/{login_option_id}`](#fallback-operation-23-2) Delete a login option and remove that authentication method from available sign-in choices. Required IAM Permission: Action iam:DeleteLoginOption, Resource arn:api7:iam:organization/ * [`PATCH` Patch a login option `/api/login_options/{login_option_id}`](#fallback-operation-23-3) Partially update a login option using JSON Patch (RFC 6902). Use this for targeted changes without replacing full configuration. Required IAM Permission: Action iam:UpdateLoginOption, Resource arn:api7:iam:organization/ * [`GET` List all login options `/api/login_options`](#fallback-operation-23-4) List all configured login options for the organization, including protocol-specific provider settings. Required IAM Permission: Action iam:GetLoginOption, Resource arn:api7:iam:organization/ * [`POST` Create a login option `/api/login_options`](#fallback-operation-23-5) Create a new login option for external authentication integration. Supported protocols include OIDC, LDAP, SAML 2.0, and CAS. Required IAM Permission: Action iam:CreateLoginOption, Resource arn:api7:iam:organization/ * [`GET` List all login options (public) `/api/login_options_for_login`](#fallback-operation-23-6) List login options. No authentication is required, and provider or policy/role details are not included in the response. ## License * [`GET` Get API7 Enterprise license details `/api/license`](#fallback-operation-24-0) Retrieve current API7 Enterprise license information, including validity and licensed capabilities. Use this endpoint to inspect license status for operations and troubleshooting. * [`PUT` Import or update the API7 Enterprise license `/api/license`](#fallback-operation-24-1) Import a new license payload or update the existing enterprise license for the organization. You can optionally use dry-run mode to validate license content before applying it. Required IAM Permission: Action iam:UpdateLicense, Resource arn:api7:iam:organization/ ## Label * [`GET` Get all labels of a resource type `/api/labels/{resource_type}`](#fallback-operation-25-0) List available key-value labels for the specified resource type. Use these labels to organize resources and apply label-based filtering in list queries. ## Alert * [`GET` List all alert policies `/api/alert/policies`](#fallback-operation-26-0) List alert policies configured for monitoring conditions such as error rates, latency thresholds, and availability checks. Use query filters to narrow results by severity, status, labels, and search terms. Required IAM Permission: Action gateway:GetAlertPolicy, Resource... * [`POST` Create an alert policy `/api/alert/policies`](#fallback-operation-26-1) Create a new alert policy that defines trigger conditions, evaluation behavior, and notification routing. After creation, the policy can begin generating alert history entries when conditions are met. Required IAM Permission: Action gateway:CreateAlertPolicy, Resource... * [`GET` Get an alert policy `/api/alert/policies/{alert_policy_id}`](#fallback-operation-26-2) Retrieve the full configuration of a specific alert policy, including its trigger rules and notification settings. Use this endpoint before updating or troubleshooting policy behavior. Required IAM Permission: Action gateway:GetAlertPolicy, Resource arn:api7:gateway:alert/%s * [`PUT` Update an alert policy `/api/alert/policies/{alert_policy_id}`](#fallback-operation-26-3) Replace the full configuration of an alert policy with the provided payload. Use this operation when you want to update all policy fields in a single request. Required IAM Permission: Action gateway:UpdateAlertPolicy, Resource arn:api7:gateway:alert/%s * [`DELETE` Delete an alert policy `/api/alert/policies/{alert_policy_id}`](#fallback-operation-26-4) Delete an existing alert policy so it no longer evaluates conditions or sends notifications. This action affects future alerting only and does not remove historical alert records. Required IAM Permission: Action gateway:DeleteAlertPolicy, Resource arn:api7:gateway:alert/%s * [`PATCH` Patch an alert policy `/api/alert/policies/{alert_policy_id}`](#fallback-operation-26-5) Partially update an alert policy using JSON Patch (RFC 6902) operations. This is useful for targeted edits without resubmitting the full policy definition. Required IAM Permission: Action gateway:UpdateAlertPolicy, Resource arn:api7:gateway:alert/%s * [`GET` List all alert histories `/api/alert/policies/histories`](#fallback-operation-26-6) List historical alert events triggered by alert policies, including occurrence time, severity, and related gateway group context. Use time-range and policy filters to investigate incidents and alert trends. Required IAM Permission: Action gateway:GetAlertPolicy, Resource... ## Contact Point * [`GET` List Contact Points `/api/contact_points`](#fallback-operation-27-0) List contact points available for alert notification delivery. Use filters and pagination to locate channels by type, labels, or search terms. Required IAM Permission: Action iam:GetContactPoint, Resource arn:api7:iam:contactpoint/%s * [`POST` Create a contact point `/api/contact_points`](#fallback-operation-27-1) Create a new contact point to deliver alert notifications through channels such as email, webhook, Slack, or DingTalk. The created contact point can then be referenced by alert policies. Required IAM Permission: Action iam:CreateContactPoint, Resource arn:api7:iam:contactpoint/ * [`GET` Get a contact point `/api/contact_points/{contact_point_id}`](#fallback-operation-27-2) Retrieve details of a specific contact point, including its configuration and metadata. Use this before updating or validating notification settings. Required IAM Permission: Action iam:GetContactPoint, Resource arn:api7:iam:contactpoint/%s * [`PUT` Update a contact point `/api/contact_points/{contact_point_id}`](#fallback-operation-27-3) Update a contact point configuration, such as destination address, authentication details, or labels. Changes take effect for subsequent alert notifications. Required IAM Permission: Action iam:UpdateContactPoint, Resource arn:api7:iam:contactpoint/%s * [`DELETE` Delete a contact point `/api/contact_points/{contact_point_id}`](#fallback-operation-27-4) Delete a contact point so it can no longer be used as an alert notification channel. Ensure dependent alert policies are updated to avoid delivery failures. Required IAM Permission: Action iam:DeleteContactPoint, Resource arn:api7:iam:contactpoint/%s * [`GET` List notification logs of a contact point `/api/contact_points/{contact_point_id}/notification_logs`](#fallback-operation-27-5) List notification delivery logs for a contact point, including status, target resource type, and timestamps. Use this to troubleshoot failed or delayed alert notifications. Required IAM Permission: Action iam:GetContactPoint, Resource arn:api7:iam:contactpoint/%s * [`GET` List a contact point usages `/api/contact_points/{contact_point_id}/usages`](#fallback-operation-27-6) List resources that currently reference the specified contact point. Check usage before deletion or major edits to understand downstream impact. Required IAM Permission: Action iam:GetContactPoint, Resource arn:api7:iam:contactpoint/%s ## Audit Logs * [`GET` List all audit logs `/api/audit_logs`](#fallback-operation-28-0) Retrieve immutable audit log records for administrative and configuration actions, including who performed each action and when. Use filters such as event type, operator, resource, gateway group, and time range to narrow the result set. Required IAM Permission: Action... * [`GET` List all event types of audit logs `/api/audit_logs/event_types`](#fallback-operation-28-1) List supported audit event types that can be used for filtering and analysis when querying audit logs. No IAM permission required. * [`GET` Export all audit logs `/api/audit_logs/export`](#fallback-operation-28-2) Export audit logs that match the specified filters into a downloadable file format for compliance, archival, or external analysis. Apply event, operator, resource, and time-range filters before export to control data scope. Required IAM Permission: Action iam:ExportAudits... ## Audit Config * [`GET` Get the configuration for audit logs. `/api/audit_logs/config`](#fallback-operation-29-0) Get the current audit logging configuration for the organization. Use this to verify how audit events are captured and retained. No IAM permission required. ## Monitoring * [`GET` Get provider portal monitoring data at a single point in time `/api/portal/monitor/query`](#fallback-operation-30-0) See Prometheus instant queries for more information. * [`GET` Get provider portal monitoring data over a range of time `/api/portal/monitor/query_range`](#fallback-operation-30-1) See Prometheus range queries for more information. * [`GET` Get data from Prometheus `/api/control_plane/prometheus/{prometheus_path}`](#fallback-operation-30-2) * [`POST` Get data from Prometheus `/api/control_plane/prometheus/{prometheus_path}`](#fallback-operation-30-3) ## Debugger * [`GET` List debug sessions `/api/gateway_groups/{gateway_group_id}/debug_sessions`](#fallback-operation-31-0) IAM Action: gateway:GetDebugSession, Resource: arn:api7:gateway:gatewaygroup/%s * [`POST` Create a debug session `/api/gateway_groups/{gateway_group_id}/debug_sessions`](#fallback-operation-31-1) IAM Action: gateway:CreateDebugSession, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Get a debug session `/api/gateway_groups/{gateway_group_id}/debug_sessions/{debug_session_id}`](#fallback-operation-31-2) IAM Action: gateway:GetDebugSession, Resource: arn:api7:gateway:gatewaygroup/%s * [`DELETE` Delete a debug session `/api/gateway_groups/{gateway_group_id}/debug_sessions/{debug_session_id}`](#fallback-operation-31-3) IAM Action: gateway:DeleteDebugSession, Resource: arn:api7:gateway:gatewaygroup/%s * [`POST` Stop a debug session `/api/gateway_groups/{gateway_group_id}/debug_sessions/{debug_session_id}/stop`](#fallback-operation-31-4) IAM Action: gateway:StopDebugSession, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` List traces of a debug session `/api/gateway_groups/{gateway_group_id}/debug_sessions/{debug_session_id}/traces`](#fallback-operation-31-5) IAM Action: GetDebugSession, Resource: arn:api7:gateway:gatewaygroup/%s * [`GET` Get a trace detail `/api/gateway_groups/{gateway_group_id}/debug_sessions/{debug_session_id}/traces/{trace_id}`](#fallback-operation-31-6) IAM Action: GetDebugSession, Resource: arn:api7:gateway:gatewaygroup/%s * [`POST` Download a trace `/api/gateway_groups/{gateway_group_id}/debug_sessions/{debug_session_id}/traces/{trace_id}/download`](#fallback-operation-31-7) IAM Action: gateway:ExportDebugSession, Resource: arn:api7:gateway:gatewaygroup/%s ## Schema * [`GET` Get schema by resource name `/apisix/admin/schema/{resource_name}`](#fallback-operation-32-0) Get the schema definition for a specific resource type by name. This is useful when rendering dynamic forms or validating resource-specific request payloads. * [`GET` Get schema definition of a plugin `/apisix/admin/schema/plugins/{plugin_name}`](#fallback-operation-32-1) Get schema definition of a plugin, including plugin meta properties and plugin properties. * [`GET` Get OpenAPI schema `/api/openapi/request_body_schema`](#fallback-operation-32-2) The endpoint returns the request body schema of PUT/POST requests for users to understand how to structure a request. * [`GET` Get core resources schema `/api/schema/core`](#fallback-operation-32-3) Retrieve schema definitions for core gateway resources. Use these schemas to validate payload structures and build configuration tooling. ## Configuration * [`POST` Validate batch configuration `/apisix/admin/configs/validate`](#fallback-operation-33-0) Validate a batch of APISIX declarative configurations including routes, services, consumers, upstreams, etc. Performs resource-level JSON Schema validation, plugin check schema advanced validation, and duplicate ID detection. Returns all validation errors at once. ## Variables * [`GET` Get all variables `/apisix/admin/variables`](#fallback-operation-34-0) List all APISIX variables (including built-in NGINX variables) that can be used in route matching conditions and plugin configurations. ## System Settings * [`GET` Get deployment settings `/api/system_settings`](#fallback-operation-35-0) Retrieve current global deployment settings for the control plane. Use this endpoint to inspect active system-level configuration before making changes. * [`PUT` Update deployment settings `/api/system_settings`](#fallback-operation-35-1) Update global system settings that control dashboard deployment behavior and platform-wide defaults. This operation affects configuration used across managed gateway resources. Required IAM Permission: Action gateway:UpdateDeploymentSetting, Resource... * [`GET` Get SCIM settings `/api/system_settings/scim`](#fallback-operation-35-2) Retrieve the current SCIM provisioning configuration for the organization. Use this to validate synchronization endpoints and provisioning status. Required IAM Permission: Action iam:GetSCIMProvisioning, Resource arn:api7:iam:organization/ * [`PUT` Update SCIM settings `/api/system_settings/scim`](#fallback-operation-35-3) Update SCIM provisioning settings used for automated identity and user lifecycle synchronization. Changes here affect how external identity providers integrate with organization users. Required IAM Permission: Action iam:UpdateSCIMProvisioning, Resource arn:api7:iam:organization/ * [`PUT` Generate SCIM Token `/api/system_settings/scim/token`](#fallback-operation-35-4) Generate or rotate the SCIM access token used by external identity providers to call SCIM provisioning APIs. Rotating this token may require updating the provider configuration. Required IAM Permission: Action iam:UpdateSCIMProvisioning, Resource arn:api7:iam:organization/ * [`GET` Get SMTP server settings `/api/system_settings/smtp_server`](#fallback-operation-35-5) Retrieve the current SMTP server settings configured for outbound email delivery. Use this endpoint when auditing email configuration or debugging mail issues. Required IAM Permission: Action iam:GetSMTPServer, Resource arn:api7:iam:organization/ * [`PUT` Update SMTP server settings `/api/system_settings/smtp_server`](#fallback-operation-35-6) Update SMTP server configuration used for system email delivery, including notifications and verification emails. Ensure credentials and host settings are valid to avoid delivery failures. Required IAM Permission: Action iam:UpdateSMTPServer, Resource arn:api7:iam:organization/ * [`GET` Get SMTP server settings status `/api/system_settings/smtp_server_status`](#fallback-operation-35-7) Get the health and readiness status of the configured SMTP server settings. This endpoint helps verify whether current configuration can be used for email sending. * [`GET` Get login failure restriction settings `/api/system_settings/login_failure_restriction`](#fallback-operation-35-8) Retrieve the current policy for temporarily banning built-in users after too many consecutive failed login attempts. Required IAM Permission: Action iam:GetLoginFailureRestriction, Resource arn:api7:iam:organization/ * [`PUT` Update login failure restriction settings `/api/system_settings/login_failure_restriction`](#fallback-operation-35-9) Update the policy that temporarily bans built-in users after too many consecutive failed login attempts, including the failure threshold and ban duration. Required IAM Permission: Action iam:UpdateLoginFailureRestriction, Resource arn:api7:iam:organization/ ## Developer Portal Settings * [`GET` Get developer portal public access `/api/portal/system_settings/public_access`](#fallback-operation-36-0) Get the current developer portal public access configuration for the specified portal. Use this to verify login and public accessibility settings. Required IAM Permission: Action portal:GetDeveloperPortalPublicAccess, Resource arn:api7:portal:portal/%s/loginsetting/ * [`PUT` Update developer portal public access `/api/portal/system_settings/public_access`](#fallback-operation-36-1) Update public access and login behavior for a developer portal, such as how external users can access sign-in or registration flows. This setting controls portal exposure and authentication entry points. Required IAM Permission: Action portal:UpdateDeveloperPortalPublicAccess... ## System Infos * [`GET` Get all system infos `/api/system_infos`](#fallback-operation-37-0) Retrieve system metadata for the dashboard environment, such as version details, license state, and runtime information used for administration and diagnostics. ## Email * [`GET` Check if an email is verified `/api/email_verified`](#fallback-operation-38-0) Check whether the specified email address has completed verification. Use this to gate workflows that require verified contact information. * [`GET` Get email verification `/api/verify_email`](#fallback-operation-38-1) Verify an email address using the verification token and redirect to the corresponding result page. This endpoint is typically called from links sent in verification emails. ## Common * [`GET` Get Dashboard Version `/api/version`](#fallback-operation-39-0) Return the current dashboard version information. Use this endpoint for compatibility checks and operational diagnostics. * [`POST` Get resource names `/api/resource_names`](#fallback-operation-39-1) Query resource names by conditions in the request payload to support selectors, autocomplete, or dependency checks in UI workflows. ## File * [`POST` Upload a file `/api/files`](#fallback-operation-40-0) Upload a file to the control plane. The file is stored compressed in the database. Returns the file ID which can be used to construct a dp-manager URL for plugin configuration. * [`GET` Download file content `/api/files/{file_id}`](#fallback-operation-40-1) Download the original content of an uploaded file. ## Provider Portal - DCR Provider * [`GET` List all DCR providers `/api/dcr_providers`](#fallback-operation-41-0) IAM Action: portal:GetDCRProvider, Resource: arn:api7:portal:dcrprovider/%s * [`POST` Create an DCR provider `/api/dcr_providers`](#fallback-operation-41-1) IAM Action: portal:CreateDCRProvider, Resource: arn:api7:portal:dcrprovider/ * [`GET` Get an DCR provider `/api/dcr_providers/{dcr_provider_id}`](#fallback-operation-41-2) IAM Action: portal:GetDCRProvider, Resource: arn:api7:portal:dcrprovider/%s * [`PUT` Update an DCR Provider `/api/dcr_providers/{dcr_provider_id}`](#fallback-operation-41-3) IAM Action: portal:UpdateDCRProvider, Resource: arn:api7:portal:dcrprovider/%s * [`DELETE` Delete an DCR Provider `/api/dcr_providers/{dcr_provider_id}`](#fallback-operation-41-4) IAM Action: portal:DeleteDCRProvider, Resource: arn:api7:portal:dcrprovider/%s ## Provider Portal - API Product * [`GET` List all API products in Provider Portal `/api/api_products`](#fallback-operation-42-0) IAM Action: portal:GetAPIProduct, Resource: arn:api7:portal:portal/%s/apiproduct/%s * [`POST` Create an API product in Provider Portal `/api/api_products`](#fallback-operation-42-1) IAM Action: portal:CreateAPIProduct, Resource: arn:api7:portal:portal/%s/apiproduct/ * [`GET` Get an API product in Provider Portal `/api/api_products/{api_product_id}`](#fallback-operation-42-2) IAM Action: portal:GetAPIProduct, Resource: arn:api7:portal:portal/%s/apiproduct/%s * [`PUT` Update an API product in Provider Portal `/api/api_products/{api_product_id}`](#fallback-operation-42-3) IAM Action: portal:UpdateAPIProduct, Resource: arn:api7:portal:portal/%s/apiproduct/%s * [`DELETE` Delete an API product in Provider Portal `/api/api_products/{api_product_id}`](#fallback-operation-42-4) IAM Action: portal:DeleteAPIProduct, Resource: arn:api7:portal:portal/%s/apiproduct/%s * [`PATCH` Patch an API product in Provider Portal `/api/api_products/{api_product_id}`](#fallback-operation-42-5) IAM Action: portal:UpdateAPIProduct, Resource: arn:api7:portal:portal/%s/apiproduct/%s * [`GET` List all subscriptions in Provider Portal for an API product `/api/api_products/{api_product_id}/subscriptions`](#fallback-operation-42-6) IAM Action: portal:GetAPIProduct, Resource: arn:api7:portal:portal/%s/apiproduct/%s * [`DELETE` Cancel a subscription in Provider Portal for an API product `/api/api_products/{api_product_id}/subscriptions/{subscription_id}`](#fallback-operation-42-7) IAM Action: portal:UpdateAPIProduct, Resource: arn:api7:portal:portal/%s/apiproduct/%s * [`GET` List all notification histories for an API product `/api/api_products/{api_product_id}/notification_histories`](#fallback-operation-42-8) IAM Action: portal:GetAPIProduct, Resource: arn:api7:portal:portal/%s/apiproduct/%s ## Provider Portal - Portal Instance * [`GET` List all portal instances `/api/portals`](#fallback-operation-43-0) IAM Action: portal:ListPortals, Resource: arn:api7:portal:portal/ * [`POST` Create a portal instance `/api/portals`](#fallback-operation-43-1) IAM Action: portal:CreatePortal, Resource: arn:api7:portal:portal/ * [`GET` Get a portal instance `/api/portals/{portal_id}`](#fallback-operation-43-2) IAM Action: portal:GetPortal, Resource: arn:api7:portal:portal/%s * [`PUT` Update a portal instance `/api/portals/{portal_id}`](#fallback-operation-43-3) IAM Action: portal:UpdatePortal, Resource: arn:api7:portal:portal/%s * [`DELETE` Delete a portal instance `/api/portals/{portal_id}`](#fallback-operation-43-4) IAM Action: portal:DeletePortal, Resource: arn:api7:portal:portal/%s ## Provider Portal - Portal Token * [`GET` List all portal tokens `/api/portal/tokens`](#fallback-operation-44-0) IAM Action: portal:GetPortalToken, Resource: arn:api7:portal:portal/%s/token/ * [`POST` Create a portal token `/api/portal/tokens`](#fallback-operation-44-1) IAM Action: portal:CreatePortalToken, Resource: arn:api7:portal:portal/%s/token/ * [`GET` Get a portal token `/api/portal/tokens/{portal_token_id}`](#fallback-operation-44-2) IAM Action: portal:GetPortalToken, Resource: arn:api7:portal:portal/%s/token/ * [`PUT` Update a portal token `/api/portal/tokens/{portal_token_id}`](#fallback-operation-44-3) IAM Action: portal:UpdatePortalToken, Resource: arn:api7:portal:portal/%s/token/ * [`DELETE` Delete a portal token `/api/portal/tokens/{portal_token_id}`](#fallback-operation-44-4) IAM Action: portal:DeletePortalToken, Resource: arn:api7:portal:portal/%s/token/ * [`PUT` Regenerate a portal token `/api/portal/tokens/{portal_token_id}/regenerate`](#fallback-operation-44-5) IAM Action: portal:UpdatePortalToken, Resource: arn:api7:portal:portal/%s/token/ ## Developer * [`GET` List Developers `/api/developers`](#fallback-operation-45-0) IAM Action: portal:GetDeveloper, Resource: arn:api7:portal:portal/%s/developer/%s * [`DELETE` Delete a developer `/api/developers/{developer_external_id}`](#fallback-operation-45-1) IAM Action: portal:DeleteDeveloper, Resource: arn:api7:portal:portal/%s/developer/%s ## Approval * [`GET` List approvals `/api/approvals`](#fallback-operation-46-0) List pending and processed approval workflow items for developer portal operations, such as API product subscriptions and developer registrations. Use filters to review approvals by status, event type, operator, applicant, and resource. Required IAM Permission: For event type... * [`POST` Accept an approval request `/api/approvals/{approval_id}/accept`](#fallback-operation-46-1) Approve a specific workflow request and apply the corresponding portal-side change, such as granting an API product subscription or accepting a developer sign-up. This operation advances the approval lifecycle to an accepted state. Required IAM Permission: For event type api... * [`POST` Reject an approval request `/api/approvals/{approval_id}/reject`](#fallback-operation-46-2) Reject a specific workflow approval request so the requested action is not applied. Use this endpoint to explicitly deny subscription or registration requests that do not meet review criteria. Required IAM Permission: For event type api product subscription: Action... ## AI Gateway Group * [`GET` List all AI Gateway groups `/api/ai_gateway_groups`](#fallback-operation-47-0) Returns a paginated list of all AI Gateway groups (AISIX clusters). * [`POST` Create an AI Gateway group `/api/ai_gateway_groups`](#fallback-operation-47-1) IAM Action: ai gateway:CreateAIGatewayGroup, Resource: arn:api7:ai gateway:aigatewaygroup/ Creates a new AI Gateway group (AISIX cluster). * [`GET` Get an AI Gateway group `/api/ai_gateway_groups/{ai_gateway_group_id}`](#fallback-operation-47-2) IAM Action: ai gateway:GetAIGatewayGroup, Resource: arn:api7:ai gateway:aigatewaygroup/%s Returns the specified AI Gateway group by ID. * [`PUT` Update an AI Gateway group `/api/ai_gateway_groups/{ai_gateway_group_id}`](#fallback-operation-47-3) IAM Action: ai gateway:UpdateAIGatewayGroup, Resource: arn:api7:ai gateway:aigatewaygroup/%s Updates the name, description, or configuration of the specified AI Gateway group. * [`DELETE` Delete an AI Gateway group `/api/ai_gateway_groups/{ai_gateway_group_id}`](#fallback-operation-47-4) IAM Action: ai gateway:DeleteAIGatewayGroup, Resource: arn:api7:ai gateway:aigatewaygroup/%s Deletes the specified AI Gateway group and all its associated instances. ## AI Gateway Instance * [`GET` List AI Gateway instances in a group `/api/ai_gateway_groups/{ai_gateway_group_id}/instances`](#fallback-operation-48-0) IAM Action: ai gateway:GetAIGatewayInstance, Resource: arn:api7:ai gateway:aigatewaygroup/%s Returns all AI Gateway instances (AISIX nodes) registered in the specified group. Instance status is computed dynamically from the last heartbeat time: Healthy within 60 seconds... * [`DELETE` Delete a single AI Gateway instance `/api/ai_gateway_groups/{ai_gateway_group_id}/instances/{ai_gateway_instance_id}`](#fallback-operation-48-1) IAM Action: ai gateway:DeleteAIGatewayInstance, Resource: arn:api7:ai gateway:aigatewaygroup/%s Removes the specified AI Gateway instance from its group. * [`GET` Generate a Docker run command to install an AI Gateway instance `/api/ai_gateway_groups/{ai_gateway_group_id}/deployment/docker`](#fallback-operation-48-2) IAM Action: ai gateway:GetAIGatewayGroup, Resource: arn:api7:ai gateway:aigatewaygroup/%s Generates a docker run command that configures and starts an AISIX AI Gateway instance connected to the specified AI Gateway group. * [`GET` Generate a Docker Compose file to install an AI Gateway instance `/api/ai_gateway_groups/{ai_gateway_group_id}/deployment/docker-compose`](#fallback-operation-48-3) IAM Action: ai gateway:GetAIGatewayGroup, Resource: arn:api7:ai gateway:aigatewaygroup/%s Generates a Docker Compose configuration that starts an AISIX AI Gateway instance connected to the specified AI Gateway group. * [`GET` Generate a Helm install script for an AI Gateway instance `/api/ai_gateway_groups/{ai_gateway_group_id}/deployment/helm/script`](#fallback-operation-48-4) IAM Action: ai gateway:GetAIGatewayGroup, Resource: arn:api7:ai gateway:aigatewaygroup/%s Generates a shell script that uses Helm to install an AISIX AI Gateway instance in Kubernetes, connected to the specified AI Gateway group. * [`GET` Generate a Helm values YAML for an AI Gateway instance `/api/ai_gateway_groups/{ai_gateway_group_id}/deployment/helm/yaml`](#fallback-operation-48-5) IAM Action: ai gateway:GetAIGatewayGroup, Resource: arn:api7:ai gateway:aigatewaygroup/%s Generates a Helm values YAML file for deploying an AISIX AI Gateway instance in Kubernetes, connected to the specified AI Gateway group. * [`POST` Issue an AI data plane certificate `/api/ai_gateway_groups/{ai_gateway_group_id}/dp_client_certificates`](#fallback-operation-48-6) Issues a client mTLS certificate for an aisix-ee AI Gateway instance to authenticate with the control plane (dp-manager). Use this during aisix-ee bootstrap or certificate rotation. ## AI Provider * [`GET` List AI providers in a group `/aisix/admin/providers`](#fallback-operation-49-0) IAM Action: ai gateway:ListAIProviders, Resource: arn:api7:ai gateway:aigatewaygroup/%s Returns all AI provider configurations in the specified AI Gateway group. * [`POST` Create an AI provider `/aisix/admin/providers`](#fallback-operation-49-1) IAM Action: ai gateway:CreateAIProvider, Resource: arn:api7:ai gateway:aigatewaygroup/%s Creates a new AI provider configuration in the specified AI Gateway group. The provider will be synced to the aisix data plane via etcd. * [`GET` Get an AI provider `/aisix/admin/providers/{provider_id}`](#fallback-operation-49-2) IAM Action: ai gateway:GetAIProvider, Resource: arn:api7:ai gateway:aigatewaygroup/%s Returns the specified AI provider configuration by ID. * [`PUT` Update an AI provider `/aisix/admin/providers/{provider_id}`](#fallback-operation-49-3) IAM Action: ai gateway:UpdateAIProvider, Resource: arn:api7:ai gateway:aigatewaygroup/%s Updates the specified AI provider configuration. Changes are synced to the aisix data plane via etcd. * [`DELETE` Delete an AI provider `/aisix/admin/providers/{provider_id}`](#fallback-operation-49-4) IAM Action: ai gateway:DeleteAIProvider, Resource: arn:api7:ai gateway:aigatewaygroup/%s Deletes the specified AI provider configuration and removes it from etcd. Returns 409 Conflict if any AI models reference this provider. ## AI Model * [`GET` List AI models in a group `/aisix/admin/models`](#fallback-operation-50-0) IAM Action: ai gateway:ListAIModels, Resource: arn:api7:ai gateway:aigatewaygroup/%s Returns all AI model configurations in the specified AI Gateway group. * [`POST` Create an AI model `/aisix/admin/models`](#fallback-operation-50-1) IAM Action: ai gateway:CreateAIModel, Resource: arn:api7:ai gateway:aigatewaygroup/%s Creates a new AI model configuration in the specified AI Gateway group. The model will be synced to the aisix data plane via etcd. * [`GET` Get an AI model `/aisix/admin/models/{model_id}`](#fallback-operation-50-2) IAM Action: ai gateway:GetAIModel, Resource: arn:api7:ai gateway:aigatewaygroup/%s Returns the specified AI model configuration by ID. * [`PUT` Update an AI model `/aisix/admin/models/{model_id}`](#fallback-operation-50-3) IAM Action: ai gateway:UpdateAIModel, Resource: arn:api7:ai gateway:aigatewaygroup/%s Updates the specified AI model configuration. Changes are synced to the aisix data plane via etcd. * [`DELETE` Delete an AI model `/aisix/admin/models/{model_id}`](#fallback-operation-50-4) IAM Action: ai gateway:DeleteAIModel, Resource: arn:api7:ai gateway:aigatewaygroup/%s Deletes the specified AI model configuration and removes it from etcd. ## AI API Key * [`GET` List AI API keys in a group `/aisix/admin/apikeys`](#fallback-operation-51-0) IAM Action: ai gateway:ListAIAPIKeys, Resource: arn:api7:ai gateway:aigatewaygroup/%s Returns all AI API key configurations in the specified AI Gateway group. * [`POST` Create an AI API key `/aisix/admin/apikeys`](#fallback-operation-51-1) IAM Action: ai gateway:CreateAIAPIKey, Resource: arn:api7:ai gateway:aigatewaygroup/%s Creates a new AI API key in the specified AI Gateway group. The key will be synced to the aisix data plane via etcd. * [`GET` Get an AI API key `/aisix/admin/apikeys/{api_key_id}`](#fallback-operation-51-2) IAM Action: ai gateway:GetAIAPIKey, Resource: arn:api7:ai gateway:aigatewaygroup/%s Returns the specified AI API key configuration by ID. * [`PUT` Update an AI API key `/aisix/admin/apikeys/{api_key_id}`](#fallback-operation-51-3) IAM Action: ai gateway:UpdateAIAPIKey, Resource: arn:api7:ai gateway:aigatewaygroup/%s Updates the specified AI API key configuration. Changes are synced to the aisix data plane via etcd. * [`DELETE` Delete an AI API key `/aisix/admin/apikeys/{api_key_id}`](#fallback-operation-51-4) IAM Action: ai gateway:DeleteAIAPIKey, Resource: arn:api7:ai gateway:aigatewaygroup/%s Deletes the specified AI API key configuration and removes it from etcd. ## AI Gateway Log * [`GET` List AI gateway span logs `/api/ai_gateway_logs`](#fallback-operation-52-0) IAM Action: ai gateway:ListAIGatewayLogs, Resource: arn:api7:ai gateway:aigatewaygroup/ Returns a paginated list of AI/LLM span logs extracted from OTLP traces sent by AI gateway dataplanes. ## Model Catalog * [`GET` List model catalog entries `/api/model_catalogs`](#fallback-operation-53-0) IAM Action: ai gateway:ListModelCatalogs, Resource: arn:api7:ai gateway:modelcatalog/ Returns globally synchronized model catalog entries that can be used to populate AI model selections. Set compact=true to receive the reduced dropdown-oriented item shape. * [`POST` Trigger a manual model catalog sync `/api/model_catalogs/sync`](#fallback-operation-53-1) IAM Action: ai gateway:SyncModelCatalog, Resource: arn:api7:ai gateway:modelcatalog/ Triggers an immediate synchronization against the configured upstream model catalog endpoint. Remote fetch failures are returned as HTTP 200 with success=false; HTTP 409 is reserved for... * [`GET` List model catalog sync logs `/api/model_sync_logs`](#fallback-operation-53-2) IAM Action: ai gateway:ListModelSyncLogs, Resource: arn:api7:ai gateway:modelcatalog/ Returns synchronization history entries, including execution status and structured diff information. ![API7.ai Logo](https://static.api7.ai/uploads/2025/03/02/api7.ai-white.avif) The digital world is connected by APIs,
API7.ai exists to make APIs more efficient, reliable, and secure. Sign up for API7 newsletter [Email address]()Subscribe Product [API7 Gateway](https://api7.ai/enterprise)[AISIX AI Gateway](https://api7.ai/ai-gateway)[API7 API Portal](https://api7.ai/portal) Learn [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[Plugin Hub](https://docs.api7.ai/hub.md)[API Gateway Comparison](https://api7.ai/api-gateway-comparison)[Customers](https://api7.ai/customers) Resources [API Gateway Docs](https://docs.api7.ai/apisix/documentation.md)[APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md)[API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[Blog](https://api7.ai/blog)[Demo Hub](https://api7.ai/demos)[APISIX vs Kong](https://api7.ai/apisix-vs-kong)[AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) Company [About](https://api7.ai/about)[Contact](https://api7.ai/contact)[Partners](https://api7.ai/partners)[Compliance Standards](https://api7.ai/compliance)[Brand Assets](https://api7.ai/branding)[Terms & Privacy](https://api7.ai/terms) *** [![SOC2 Type II](https://static.api7.ai/uploads/2025/03/02/fMrK6JR5_21972-312_SOC_NonCPA.avif)](https://api7.ai/compliance) [![ISO 27001](https://static.api7.ai/uploads/2025/03/02/kStGFFd2_iso-27001.avif)](https://api7.ai/compliance) [![HIPAA](https://static.api7.ai/uploads/2025/03/02/PSOVypp6_hipaa.avif)](https://api7.ai/compliance) [![GDPR](https://static.api7.ai/uploads/2025/03/02/6cv5RTfR_gdpr.avif)](https://api7.ai/compliance) [![Red Herring](https://static.api7.ai/uploads/2025/03/02/6385ad60e1f6a.avif)](https://api7.ai/blog/among-2022-red-herring-top-100-global) Copyright © APISEVEN PTE. LTD 2019 – 2026. Apache, Apache APISIX, APISIX, and associated open source project names are trademarks of the [Apache Software Foundation](https://www.apache.org/) [](https://www.linkedin.com/company/api7-ai/)[](https://github.com/api7)[](https://twitter.com/api7_ai) --- [Skip to main content](#__docusaurus_skipToContent_fallback) [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) ProductsSolutions[Customers](https://api7.ai/customers) Pricing Resources[Blog](https://api7.ai/blog) [Login](https://console.api7.cloud)Get a DemoStart for Free [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) * Products [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/enterprise) [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/api7-enterprise-vs-apisix) [Apache APISIX vs API7](https://api7.ai/api7-enterprise-vs-apisix)[- ](https://api7.ai/portal) [API7 API Portal](https://api7.ai/portal) [Apache APISIX](https://api7.ai/apisix)[- ](https://api7.ai/apisix) [What's Apache APISIX?](https://api7.ai/apisix)[- ](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway) [Why Apache APISIX?](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway)[- ](https://api7.ai/apache-apisix-enterprise-support) [APISIX Commercial Support](https://api7.ai/apache-apisix-enterprise-support) [AISIX AI Gateway](https://api7.ai/ai-gateway)[- ](https://api7.ai/ai-gateway) [AISIX AI Gateway](https://api7.ai/ai-gateway) * Solutions [Developer](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/monolith-to-microservices) [Monolith to Microservices](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/on-prem-to-hybrid-cloud) [On-Prem to Hybrid Cloud](https://api7.ai/solutions/on-prem-to-hybrid-cloud)[- ](https://api7.ai/solutions/observability) [Observability](https://api7.ai/solutions/observability) [- ](https://api7.ai/solutions/vm-to-kubernetes) [VM to Kubernetes](https://api7.ai/solutions/vm-to-kubernetes)[- ](https://api7.ai/solutions/zero-trust-security) [Zero Trust Security](https://api7.ai/solutions/zero-trust-security) [Industry](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/financial-services) [Financial Services](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/iot) [IoT and Automotive](https://api7.ai/solutions/iot)[- ](https://api7.ai/solutions/blockchain) [Blockchain](https://api7.ai/solutions/blockchain) [- ](https://api7.ai/solutions/manufacturing) [Manufacturing](https://api7.ai/solutions/manufacturing) * [Customers](https://api7.ai/customers) * Pricing [- ](https://api7.ai/pricing) [API Gateway](https://api7.ai/pricing)[- ](https://api7.ai/ai-gateway/pricing) [AI Gateway](https://api7.ai/ai-gateway/pricing) * Resources [Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/ai-gateway/.md) [AISIX Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/api7-gateway) [API7 Gateway](https://docs.api7.ai/api7-gateway)[- ](https://docs.api7.ai/api7-gateway/ai-agent-skills.md) [API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[- ](https://docs.api7.ai/apisix) [Apache APISIX](https://docs.api7.ai/apisix)[- ](https://docs.api7.ai/apisix/ai-agent-skills.md) [APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md) [Compare](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-kong) [Apache APISIX vs Kong](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-nginx) [Apache APISIX vs NGINX](https://api7.ai/apisix-vs-nginx)[- ](https://api7.ai/api-gateway-comparison) [2026 Top API Gateway Comparison](https://api7.ai/api-gateway-comparison)[- ](https://api7.ai/ai-gateway-comparison) [AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) [Learn](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/openresty) [OpenResty (NGINX + Lua)](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/api-gateway-guide) [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[- ](https://api7.ai/learning-center/ai-gateway-guide) [AI Gateway Guide](https://api7.ai/learning-center/ai-gateway-guide)[- ](https://api7.ai/learning-center/api-infrastructure-guide) [API Infrastructure Guide](https://api7.ai/learning-center/api-infrastructure-guide) [Explore](https://api7.ai/demos)[- ](https://api7.ai/demos) [Demo Hub](https://api7.ai/demos)[- ](https://docs.api7.ai/hub.md) [Plugin Hub](https://docs.api7.ai/hub.md)[- ](https://api7.ai/category/usercase) [Case Studies](https://api7.ai/category/usercase) * [Blog](https://api7.ai/blog) Get a DemoStart for Free [](https://docs.api7.ai/)[Apache APISIX](https://docs.api7.ai/apisix/documentation.md)[API7 Gateway](https://docs.api7.ai/api7-gateway/overview.md)[Ingress Controller](https://docs.api7.ai/ingress-controller/documentation.md)[AISIX AI Gateway](https://docs.api7.ai/ai-gateway/.md)[Plugin Hub](https://docs.api7.ai/hub.md) Search # API7 Enterprise Developer Portal APIs API7 Enterprise Developer Portal APIs are RESTful APIs that allow you to create and manage developer portal resources. Base URL All API paths are relative to your API7 Developer Portal address... ## Developer * [`GET` List Developers `/api/developers`](#fallback-operation-0-0) * [`POST` Create a developer `/api/developers`](#fallback-operation-0-1) * [`DELETE` Delete a developer `/api/developers/{developer_id}`](#fallback-operation-0-2) ## API Product * [`GET` List all API products `/api/api_products`](#fallback-operation-1-0) * [`GET` Get an API Product for Developer Portal `/api/api_products/{api_product_id}`](#fallback-operation-1-1) * [`POST` Create a subscription for an API Product. `/api/api_products/{api_product_id}/subscriptions`](#fallback-operation-1-2) ## Subscription * [`GET` List subscriptions. `/api/subscriptions`](#fallback-operation-2-0) * [`POST` Create a subscription for an API Product `/api/subscriptions`](#fallback-operation-2-1) * [`DELETE` Unsubscribe an API product for the given application `/api/subscriptions/{subscription_id}`](#fallback-operation-2-2) ## Application * [`GET` List all applications for the logged in developer. `/api/applications`](#fallback-operation-3-0) * [`POST` Create an application by the logged in developer. `/api/applications`](#fallback-operation-3-1) * [`GET` Get an application for the logged in developer. `/api/applications/{application_id}`](#fallback-operation-3-2) * [`PUT` Update an application basic information by the logged in developer. `/api/applications/{application_id}`](#fallback-operation-3-3) * [`DELETE` Delete an application by the logged in developer. `/api/applications/{application_id}`](#fallback-operation-3-4) ## Credential * [`GET` List all credentials for the logged in developer `/api/applications/{application_id}/credentials`](#fallback-operation-4-0) * [`POST` Create an application credential `/api/applications/{application_id}/credentials`](#fallback-operation-4-1) * [`GET` Get an application credential for Developer Portal `/api/applications/{application_id}/credentials/{credential_id}`](#fallback-operation-4-2) * [`PUT` Update an application credential for Developer Portal `/api/applications/{application_id}/credentials/{credential_id}`](#fallback-operation-4-3) * [`DELETE` Delete an application credential for Developer Portal `/api/applications/{application_id}/credentials/{credential_id}`](#fallback-operation-4-4) * [`PUT` Regenerate an application credential auth conf `/api/applications/{application_id}/credentials/{credential_id}/regenerate`](#fallback-operation-4-5) * [`GET` List all Credentials `/api/credentials`](#fallback-operation-4-6) ## API Calls * [`GET` Get API Calls `/api/applications/api_calls`](#fallback-operation-5-0) Retrieve a list of API calls made by the user. ## Label * [`GET` Get all labels of a resource type `/api/labels/{resource_type}`](#fallback-operation-6-0) ## System Settings * [`GET` Get SMTP server settings status `/api/system_settings/smtp_server_status`](#fallback-operation-7-0) Get SMTP server settings status. * [`GET` Get public access settings `/api/system_settings/public_access`](#fallback-operation-7-1) ## DCR Provider * [`GET` List all DCR providers `/api/dcr_providers`](#fallback-operation-8-0) ## Approval * [`GET` List approvals `/api/approvals`](#fallback-operation-9-0) List pending and processed approval workflow items for the current portal, such as API product subscriptions and developer registrations. Results are scoped to the portal of the request token. * [`POST` Accept an approval request `/api/approvals/{approval_id}/accept`](#fallback-operation-9-1) Approve a specific workflow request and apply the corresponding portal-side change, such as granting an API product subscription or accepting a developer sign-up. * [`POST` Reject an approval request `/api/approvals/{approval_id}/reject`](#fallback-operation-9-2) Reject a specific workflow approval request so the requested action is not applied. ![API7.ai Logo](https://static.api7.ai/uploads/2025/03/02/api7.ai-white.avif) The digital world is connected by APIs,
API7.ai exists to make APIs more efficient, reliable, and secure. Sign up for API7 newsletter [Email address]()Subscribe Product [API7 Gateway](https://api7.ai/enterprise)[AISIX AI Gateway](https://api7.ai/ai-gateway)[API7 API Portal](https://api7.ai/portal) Learn [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[Plugin Hub](https://docs.api7.ai/hub.md)[API Gateway Comparison](https://api7.ai/api-gateway-comparison)[Customers](https://api7.ai/customers) Resources [API Gateway Docs](https://docs.api7.ai/apisix/documentation.md)[APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md)[API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[Blog](https://api7.ai/blog)[Demo Hub](https://api7.ai/demos)[APISIX vs Kong](https://api7.ai/apisix-vs-kong)[AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) Company [About](https://api7.ai/about)[Contact](https://api7.ai/contact)[Partners](https://api7.ai/partners)[Compliance Standards](https://api7.ai/compliance)[Brand Assets](https://api7.ai/branding)[Terms & Privacy](https://api7.ai/terms) *** [![SOC2 Type II](https://static.api7.ai/uploads/2025/03/02/fMrK6JR5_21972-312_SOC_NonCPA.avif)](https://api7.ai/compliance) [![ISO 27001](https://static.api7.ai/uploads/2025/03/02/kStGFFd2_iso-27001.avif)](https://api7.ai/compliance) [![HIPAA](https://static.api7.ai/uploads/2025/03/02/PSOVypp6_hipaa.avif)](https://api7.ai/compliance) [![GDPR](https://static.api7.ai/uploads/2025/03/02/6cv5RTfR_gdpr.avif)](https://api7.ai/compliance) [![Red Herring](https://static.api7.ai/uploads/2025/03/02/6385ad60e1f6a.avif)](https://api7.ai/blog/among-2022-red-herring-top-100-global) Copyright © APISEVEN PTE. LTD 2019 – 2026. Apache, Apache APISIX, APISIX, and associated open source project names are trademarks of the [Apache Software Foundation](https://www.apache.org/) [](https://www.linkedin.com/company/api7-ai/)[](https://github.com/api7)[](https://twitter.com/api7_ai) --- [Skip to main content](#__docusaurus_skipToContent_fallback) [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) ProductsSolutions[Customers](https://api7.ai/customers) Pricing Resources[Blog](https://api7.ai/blog) [Login](https://console.api7.cloud)Get a DemoStart for Free [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) * Products [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/enterprise) [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/api7-enterprise-vs-apisix) [Apache APISIX vs API7](https://api7.ai/api7-enterprise-vs-apisix)[- ](https://api7.ai/portal) [API7 API Portal](https://api7.ai/portal) [Apache APISIX](https://api7.ai/apisix)[- ](https://api7.ai/apisix) [What's Apache APISIX?](https://api7.ai/apisix)[- ](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway) [Why Apache APISIX?](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway)[- ](https://api7.ai/apache-apisix-enterprise-support) [APISIX Commercial Support](https://api7.ai/apache-apisix-enterprise-support) [AISIX AI Gateway](https://api7.ai/ai-gateway)[- ](https://api7.ai/ai-gateway) [AISIX AI Gateway](https://api7.ai/ai-gateway) * Solutions [Developer](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/monolith-to-microservices) [Monolith to Microservices](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/on-prem-to-hybrid-cloud) [On-Prem to Hybrid Cloud](https://api7.ai/solutions/on-prem-to-hybrid-cloud)[- ](https://api7.ai/solutions/observability) [Observability](https://api7.ai/solutions/observability) [- ](https://api7.ai/solutions/vm-to-kubernetes) [VM to Kubernetes](https://api7.ai/solutions/vm-to-kubernetes)[- ](https://api7.ai/solutions/zero-trust-security) [Zero Trust Security](https://api7.ai/solutions/zero-trust-security) [Industry](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/financial-services) [Financial Services](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/iot) [IoT and Automotive](https://api7.ai/solutions/iot)[- ](https://api7.ai/solutions/blockchain) [Blockchain](https://api7.ai/solutions/blockchain) [- ](https://api7.ai/solutions/manufacturing) [Manufacturing](https://api7.ai/solutions/manufacturing) * [Customers](https://api7.ai/customers) * Pricing [- ](https://api7.ai/pricing) [API Gateway](https://api7.ai/pricing)[- ](https://api7.ai/ai-gateway/pricing) [AI Gateway](https://api7.ai/ai-gateway/pricing) * Resources [Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/ai-gateway/.md) [AISIX Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/api7-gateway) [API7 Gateway](https://docs.api7.ai/api7-gateway)[- ](https://docs.api7.ai/api7-gateway/ai-agent-skills.md) [API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[- ](https://docs.api7.ai/apisix) [Apache APISIX](https://docs.api7.ai/apisix)[- ](https://docs.api7.ai/apisix/ai-agent-skills.md) [APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md) [Compare](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-kong) [Apache APISIX vs Kong](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-nginx) [Apache APISIX vs NGINX](https://api7.ai/apisix-vs-nginx)[- ](https://api7.ai/api-gateway-comparison) [2026 Top API Gateway Comparison](https://api7.ai/api-gateway-comparison)[- ](https://api7.ai/ai-gateway-comparison) [AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) [Learn](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/openresty) [OpenResty (NGINX + Lua)](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/api-gateway-guide) [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[- ](https://api7.ai/learning-center/ai-gateway-guide) [AI Gateway Guide](https://api7.ai/learning-center/ai-gateway-guide)[- ](https://api7.ai/learning-center/api-infrastructure-guide) [API Infrastructure Guide](https://api7.ai/learning-center/api-infrastructure-guide) [Explore](https://api7.ai/demos)[- ](https://api7.ai/demos) [Demo Hub](https://api7.ai/demos)[- ](https://docs.api7.ai/hub.md) [Plugin Hub](https://docs.api7.ai/hub.md)[- ](https://api7.ai/category/usercase) [Case Studies](https://api7.ai/category/usercase) * [Blog](https://api7.ai/blog) Get a DemoStart for Free [](https://docs.api7.ai/)[Apache APISIX](https://docs.api7.ai/apisix/documentation.md)[API7 Gateway](https://docs.api7.ai/api7-gateway/overview.md)[Ingress Controller](https://docs.api7.ai/ingress-controller/documentation.md)[AISIX AI Gateway](https://docs.api7.ai/ai-gateway/.md)[Plugin Hub](https://docs.api7.ai/hub.md) Search # Apache APISIX Admin API The Apache APISIX Admin API is a RESTful interface for managing all APISIX gateway resources — routes, upstreams, services, consumers, SSL certificates, plugins, and more. APISIX Admin API allows... ## Routes * [`GET` List All Routes `/apisix/admin/routes`](#fallback-operation-0-0) Retrieve all configured routes. Supports filtering by name, label, uri, or filter expression. Results include full route configuration and metadata (creation/update timestamps). * [`POST` Create a Route `/apisix/admin/routes`](#fallback-operation-0-1) Create a new route with an auto-generated ID. A route must include at least a uri and one of: upstream (inline), upstream id, service id, or plugins. * [`GET` Get a Route `/apisix/admin/routes/{id}`](#fallback-operation-0-2) Retrieve a single route by its ID. * [`PUT` Create or Replace a Route `/apisix/admin/routes/{id}`](#fallback-operation-0-3) Create a new route with a specified ID, or fully replace an existing route. Note: This is a full replacement — all fields must be provided. Use PATCH for partial updates. * [`DELETE` Delete a Route `/apisix/admin/routes/{id}`](#fallback-operation-0-4) Delete a route by its ID. This action is irreversible. * [`PATCH` Update a Route (Partial) `/apisix/admin/routes/{id}`](#fallback-operation-0-5) Partially update a route's configuration. Only the fields included in the request body are modified; all other fields remain unchanged. Tip: Use this to update a single field (e.g., change the upstream) without resending the entire route configuration. * [`PATCH` Replace a single nested field of a route `/apisix/admin/routes/{id}/{sub_path}`](#fallback-operation-0-6) Replaces the value at sub path within the route document with the request body. The body is any valid JSON value, not necessarily an object. The updated complete resource must pass validation. ## Upstreams * [`GET` List All Upstreams `/apisix/admin/upstreams`](#fallback-operation-1-0) Retrieve all configured upstreams. * [`POST` Create an Upstream `/apisix/admin/upstreams`](#fallback-operation-1-1) Create a new upstream with an auto-generated ID. * [`GET` Get an Upstream `/apisix/admin/upstreams/{id}`](#fallback-operation-1-2) Retrieve a single upstream by its ID. * [`PUT` Create or Replace an Upstream `/apisix/admin/upstreams/{id}`](#fallback-operation-1-3) Create a new upstream with a specified ID, or fully replace an existing upstream. * [`DELETE` Delete an Upstream `/apisix/admin/upstreams/{id}`](#fallback-operation-1-4) Delete an upstream by its ID. Fails if the upstream is still referenced by a route or service. * [`PATCH` Update an Upstream (Partial) `/apisix/admin/upstreams/{id}`](#fallback-operation-1-5) Partially update an upstream's configuration. * [`PATCH` Replace a single nested field of an upstream `/apisix/admin/upstreams/{id}/{sub_path}`](#fallback-operation-1-6) Replaces the value at sub path within the upstream document with the request body. The body is any valid JSON value, not necessarily an object. The updated complete resource must pass validation. ## Services * [`GET` List All Services `/apisix/admin/services`](#fallback-operation-2-0) Retrieve all configured services. * [`POST` Create a Service `/apisix/admin/services`](#fallback-operation-2-1) Create a new service with an auto-generated ID. * [`GET` Get a Service `/apisix/admin/services/{id}`](#fallback-operation-2-2) Retrieve a single service by its ID. * [`PUT` Create or Replace a Service `/apisix/admin/services/{id}`](#fallback-operation-2-3) Create a new service with a specified ID, or fully replace an existing service. * [`DELETE` Delete a Service `/apisix/admin/services/{id}`](#fallback-operation-2-4) Delete a service by its ID. * [`PATCH` Update a Service (Partial) `/apisix/admin/services/{id}`](#fallback-operation-2-5) Partially update a service's configuration. * [`PATCH` Replace a single nested field of a service `/apisix/admin/services/{id}/{sub_path}`](#fallback-operation-2-6) Replaces the value at sub path within the service document with the request body. The body is any valid JSON value, not necessarily an object. The updated complete resource must pass validation. ## Consumers * [`GET` List All Consumers `/apisix/admin/consumers`](#fallback-operation-3-0) Retrieve all configured consumers. * [`PUT` Create or Update a Consumer `/apisix/admin/consumers`](#fallback-operation-3-1) Create a new consumer or update an existing one. The consumer is identified by username in the request body. * [`GET` Get a Consumer `/apisix/admin/consumers/{username}`](#fallback-operation-3-2) Retrieve a single consumer by username. * [`DELETE` Delete a Consumer `/apisix/admin/consumers/{username}`](#fallback-operation-3-3) Delete a consumer by username. This also deletes all associated credentials. ## Consumer Groups * [`GET` List All Consumer Groups `/apisix/admin/consumer_groups`](#fallback-operation-4-0) Retrieve all configured consumer groups. * [`GET` Get a Consumer Group `/apisix/admin/consumer_groups/{id}`](#fallback-operation-4-1) Retrieve a single consumer group by its ID. * [`PUT` Create or Replace a Consumer Group `/apisix/admin/consumer_groups/{id}`](#fallback-operation-4-2) Create a new consumer group with a specified ID, or fully replace an existing one. * [`DELETE` Delete a Consumer Group `/apisix/admin/consumer_groups/{id}`](#fallback-operation-4-3) Delete a consumer group by its ID. * [`PATCH` Update a Consumer Group (Partial) `/apisix/admin/consumer_groups/{id}`](#fallback-operation-4-4) Partially update a consumer group's configuration. * [`PATCH` Replace a single nested field of a consumer group `/apisix/admin/consumer_groups/{id}/{sub_path}`](#fallback-operation-4-5) Replaces the value at sub path within the consumer group document with the request body. The body is any valid JSON value, not necessarily an object. The updated complete resource must pass validation. ## Credentials * [`GET` List Consumer Credentials `/apisix/admin/consumers/{username}/credentials`](#fallback-operation-5-0) Retrieve all credentials for a specific consumer. * [`GET` Get a Credential `/apisix/admin/consumers/{username}/credentials/{id}`](#fallback-operation-5-1) Retrieve a specific credential for a consumer. * [`PUT` Create or Replace a Credential `/apisix/admin/consumers/{username}/credentials/{id}`](#fallback-operation-5-2) Create a new credential for a consumer with a specified ID, or replace an existing one. Each credential holds exactly one authentication plugin. * [`DELETE` Delete a Credential `/apisix/admin/consumers/{username}/credentials/{id}`](#fallback-operation-5-3) Delete a specific credential from a consumer. ## Plugins * [`GET` List Plugin Names `/apisix/admin/plugins/list`](#fallback-operation-6-0) Get a list of all available plugin names, optionally filtered by subsystem (http or stream). * [`GET` Get Plugin JSON Schema `/apisix/admin/plugins/{plugin_name}`](#fallback-operation-6-1) Retrieve the JSON Schema definition for a specific plugin. This is useful for validating plugin configurations before applying them. * [`GET` List Plugin Attributes `/apisix/admin/plugins`](#fallback-operation-6-2) Get attributes of all plugins configured on the APISIX instance. Deprecated: This endpoint is being replaced. Use GET /plugins/list and GET /plugins/{plugin name} instead. * [`PUT` Hot-Reload Plugins `/apisix/admin/plugins/reload`](#fallback-operation-6-3) Trigger a hot reload of all plugins. This re-reads plugin configurations from etcd without restarting APISIX. Important: Only the PUT method is accepted. ## Plugin Configs * [`GET` List All Plugin Configs `/apisix/admin/plugin_configs`](#fallback-operation-7-0) Retrieve all configured plugin configs. * [`GET` Get a Plugin Config `/apisix/admin/plugin_configs/{id}`](#fallback-operation-7-1) Retrieve a single plugin config by its ID. * [`PUT` Create or Replace a Plugin Config `/apisix/admin/plugin_configs/{id}`](#fallback-operation-7-2) Create a new plugin config with a specified ID, or fully replace an existing one. * [`DELETE` Delete a Plugin Config `/apisix/admin/plugin_configs/{id}`](#fallback-operation-7-3) Delete a plugin config by its ID. * [`PATCH` Update a Plugin Config (Partial) `/apisix/admin/plugin_configs/{id}`](#fallback-operation-7-4) Partially update a plugin config's configuration. * [`PATCH` Replace a single nested field of a plugin config `/apisix/admin/plugin_configs/{id}/{sub_path}`](#fallback-operation-7-5) Replaces the value at sub path within the plugin config document with the request body. The body is any valid JSON value, not necessarily an object. The updated complete resource must pass validation. ## Plugin Metadata * [`GET` List All Plugin Metadata `/apisix/admin/plugin_metadata`](#fallback-operation-8-0) Retrieve metadata for all plugins that have metadata configured. * [`GET` Get Plugin Metadata `/apisix/admin/plugin_metadata/{plugin_name}`](#fallback-operation-8-1) Retrieve metadata for a specific plugin by plugin name. * [`PUT` Create or Replace Plugin Metadata `/apisix/admin/plugin_metadata/{plugin_name}`](#fallback-operation-8-2) Set metadata for a specific plugin. This applies to all instances of the plugin across the cluster. * [`DELETE` Delete Plugin Metadata `/apisix/admin/plugin_metadata/{plugin_name}`](#fallback-operation-8-3) Delete metadata for a specific plugin. ## Global Rules * [`GET` List All Global Rules `/apisix/admin/global_rules`](#fallback-operation-9-0) Retrieve all configured global rules. * [`GET` Get a Global Rule `/apisix/admin/global_rules/{id}`](#fallback-operation-9-1) Retrieve a single global rule by its ID. * [`PUT` Create or Replace a Global Rule `/apisix/admin/global_rules/{id}`](#fallback-operation-9-2) Create a new global rule with a specified ID, or fully replace an existing one. * [`DELETE` Delete a Global Rule `/apisix/admin/global_rules/{id}`](#fallback-operation-9-3) Delete a global rule by its ID. * [`PATCH` Update a Global Rule (Partial) `/apisix/admin/global_rules/{id}`](#fallback-operation-9-4) Partially update a global rule's configuration. * [`PATCH` Replace a single nested field of a global rule `/apisix/admin/global_rules/{id}/{sub_path}`](#fallback-operation-9-5) Replaces the value at sub path within the global rule document with the request body. The body is any valid JSON value, not necessarily an object. The updated complete resource must pass validation. ## Stream Routes * [`GET` List All Stream Routes `/apisix/admin/stream_routes`](#fallback-operation-10-0) Retrieve all configured stream (L4) routes. Prerequisite: Stream mode must be enabled in config.yaml. * [`POST` Create a Stream Route `/apisix/admin/stream_routes`](#fallback-operation-10-1) Create a new stream route with an auto-generated ID. * [`GET` Get Stream Route by ID `/apisix/admin/stream_routes/{id}`](#fallback-operation-10-2) Get stream route by ID. * [`PUT` Create Stream Route by ID `/apisix/admin/stream_routes/{id}`](#fallback-operation-10-3) Create a stream route with a specified ID. * [`DELETE` Delete Stream Route by ID `/apisix/admin/stream_routes/{id}`](#fallback-operation-10-4) Delete a stream route by ID. ## SSL Certificates * [`GET` List All SSL Certificates `/apisix/admin/ssls`](#fallback-operation-11-0) Retrieve all configured SSL certificates. Private keys are not included in the response for security. * [`POST` Create an SSL Certificate `/apisix/admin/ssls`](#fallback-operation-11-1) Upload a new SSL certificate with an auto-generated ID. * [`GET` Get an SSL Certificate `/apisix/admin/ssls/{id}`](#fallback-operation-11-2) Retrieve a single SSL certificate by its ID. The private key is not included in the response. * [`PUT` Create or Replace an SSL Certificate `/apisix/admin/ssls/{id}`](#fallback-operation-11-3) Upload a new SSL certificate with a specified ID, or replace an existing one. * [`DELETE` Delete an SSL Certificate `/apisix/admin/ssls/{id}`](#fallback-operation-11-4) Delete an SSL certificate by its ID. * [`PATCH` Update an SSL Certificate (Partial) `/apisix/admin/ssls/{id}`](#fallback-operation-11-5) Partially update an SSL certificate's configuration. Only the fields included in the request body are modified; all other fields remain unchanged. * [`PATCH` Replace a single nested field of an SSL certificate `/apisix/admin/ssls/{id}/{sub_path}`](#fallback-operation-11-6) Replaces the value at sub path within the SSL certificate document with the request body. The body is any valid JSON value, not necessarily an object. The updated complete resource must pass validation. ## Secrets * [`GET` List All Secrets `/apisix/admin/secrets`](#fallback-operation-12-0) Retrieve all configured secret manager integrations. * [`GET` Get a Secret `/apisix/admin/secrets/{secret_type}/{id}`](#fallback-operation-12-1) Retrieve a specific secret manager configuration. * [`PUT` Create or Replace a Secret `/apisix/admin/secrets/{secret_type}/{id}`](#fallback-operation-12-2) Configure a secret manager integration with a specified ID. * [`DELETE` Delete a Secret `/apisix/admin/secrets/{secret_type}/{id}`](#fallback-operation-12-3) Delete a secret manager configuration. * [`PATCH` Update a Secret (Partial) `/apisix/admin/secrets/{secret_type}/{id}`](#fallback-operation-12-4) Partially update a secret manager configuration. * [`GET` List secrets of a single secret-manager type `/apisix/admin/secrets/{secret_type}`](#fallback-operation-12-5) Returns secrets registered under the given manager (vault, aws, or gcp). For listing across all managers, use GET /apisix/admin/secrets. * [`PATCH` Replace a single nested field of a secret `/apisix/admin/secrets/{secret_type}/{id}/{sub_path}`](#fallback-operation-12-6) Replaces the value at sub path within the secret document with the request body. The body is any valid JSON value, not necessarily an object. The updated complete resource must pass validation. ## Protos * [`GET` Get a Proto `/apisix/admin/protos/{id}`](#fallback-operation-13-0) Retrieve a single protobuf definition by its ID. * [`PUT` Create or Replace a Proto `/apisix/admin/protos/{id}`](#fallback-operation-13-1) Upload a protobuf definition with a specified ID, or replace an existing one. * [`DELETE` Delete a Proto `/apisix/admin/protos/{id}`](#fallback-operation-13-2) Delete a protobuf definition by its ID. * [`GET` List All Protos `/apisix/admin/protos`](#fallback-operation-13-3) Retrieve all stored protobuf definitions. * [`POST` Create a Proto `/apisix/admin/protos`](#fallback-operation-13-4) Upload a new protobuf definition with an auto-generated ID. ## Schema Validation * [`POST` Validate Resource Configuration `/apisix/admin/schema/validate/{resource}`](#fallback-operation-14-0) Validate a resource configuration against APISIX's JSON Schema without creating the resource. Useful for dry-run validation in CI/CD pipelines. The {resource} path parameter specifies the resource type (e.g., routes, upstreams, services, etc.). * [`POST` Validate a Configuration Without Applying It `/apisix/admin/configs/validate`](#fallback-operation-14-1) Validates a declarative APISIX configuration when the deployment uses either the etcd or yaml configuration provider. The endpoint accepts JSON and YAML request bodies up to 1.5 MiB and uses the normal Admin API authentication. Validation covers known resource arrays, numeric... * [`GET` Get the JSON Schema of a built-in resource `/apisix/admin/schema/{resource}`](#fallback-operation-14-2) Returns the JSON Schema document used internally by APISIX to validate the named resource type. * [`GET` Get a plugin's JSON Schema (alternative path) `/apisix/admin/schema/plugins/{plugin_name}`](#fallback-operation-14-3) Equivalent to GET /apisix/admin/plugins/{plugin name}. Provided for symmetry with GET /apisix/admin/schema/{resource}. ## Health * [`HEAD` Liveness probe `/apisix/admin`](#fallback-operation-15-0) Returns 200 OK if the Admin API is reachable. Useful as a Kubernetes liveness probe target. This endpoint does not require authentication. ## Standalone * [`GET` Get the standalone configuration `/apisix/admin/configs`](#fallback-operation-16-0) Available only when deployment.role data plane.config provider is yaml. Returns the entire current configuration document. * [`PUT` Replace the standalone configuration `/apisix/admin/configs`](#fallback-operation-16-1) Replaces the entire current configuration. Available only when deployment.role data plane.config provider is yaml. * [`HEAD` Probe standalone configuration availability `/apisix/admin/configs`](#fallback-operation-16-2) Returns 200 OK if the deployment is in YAML provider mode and the configuration endpoint is reachable. ![API7.ai Logo](https://static.api7.ai/uploads/2025/03/02/api7.ai-white.avif) The digital world is connected by APIs,
API7.ai exists to make APIs more efficient, reliable, and secure. Sign up for API7 newsletter [Email address]()Subscribe Product [API7 Gateway](https://api7.ai/enterprise)[AISIX AI Gateway](https://api7.ai/ai-gateway)[API7 API Portal](https://api7.ai/portal) Learn [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[Plugin Hub](https://docs.api7.ai/hub.md)[API Gateway Comparison](https://api7.ai/api-gateway-comparison)[Customers](https://api7.ai/customers) Resources [API Gateway Docs](https://docs.api7.ai/apisix/documentation.md)[APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md)[API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[Blog](https://api7.ai/blog)[Demo Hub](https://api7.ai/demos)[APISIX vs Kong](https://api7.ai/apisix-vs-kong)[AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) Company [About](https://api7.ai/about)[Contact](https://api7.ai/contact)[Partners](https://api7.ai/partners)[Compliance Standards](https://api7.ai/compliance)[Brand Assets](https://api7.ai/branding)[Terms & Privacy](https://api7.ai/terms) *** [![SOC2 Type II](https://static.api7.ai/uploads/2025/03/02/fMrK6JR5_21972-312_SOC_NonCPA.avif)](https://api7.ai/compliance) [![ISO 27001](https://static.api7.ai/uploads/2025/03/02/kStGFFd2_iso-27001.avif)](https://api7.ai/compliance) [![HIPAA](https://static.api7.ai/uploads/2025/03/02/PSOVypp6_hipaa.avif)](https://api7.ai/compliance) [![GDPR](https://static.api7.ai/uploads/2025/03/02/6cv5RTfR_gdpr.avif)](https://api7.ai/compliance) [![Red Herring](https://static.api7.ai/uploads/2025/03/02/6385ad60e1f6a.avif)](https://api7.ai/blog/among-2022-red-herring-top-100-global) Copyright © APISEVEN PTE. LTD 2019 – 2026. Apache, Apache APISIX, APISIX, and associated open source project names are trademarks of the [Apache Software Foundation](https://www.apache.org/) [](https://www.linkedin.com/company/api7-ai/)[](https://github.com/api7)[](https://twitter.com/api7_ai) --- [Skip to main content](#__docusaurus_skipToContent_fallback) [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) ProductsSolutions[Customers](https://api7.ai/customers) Pricing Resources[Blog](https://api7.ai/blog) [Login](https://console.api7.cloud)Get a DemoStart for Free [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) * Products [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/enterprise) [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/api7-enterprise-vs-apisix) [Apache APISIX vs API7](https://api7.ai/api7-enterprise-vs-apisix)[- ](https://api7.ai/portal) [API7 API Portal](https://api7.ai/portal) [Apache APISIX](https://api7.ai/apisix)[- ](https://api7.ai/apisix) [What's Apache APISIX?](https://api7.ai/apisix)[- ](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway) [Why Apache APISIX?](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway)[- ](https://api7.ai/apache-apisix-enterprise-support) [APISIX Commercial Support](https://api7.ai/apache-apisix-enterprise-support) [AISIX AI Gateway](https://api7.ai/ai-gateway)[- ](https://api7.ai/ai-gateway) [AISIX AI Gateway](https://api7.ai/ai-gateway) * Solutions [Developer](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/monolith-to-microservices) [Monolith to Microservices](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/on-prem-to-hybrid-cloud) [On-Prem to Hybrid Cloud](https://api7.ai/solutions/on-prem-to-hybrid-cloud)[- ](https://api7.ai/solutions/observability) [Observability](https://api7.ai/solutions/observability) [- ](https://api7.ai/solutions/vm-to-kubernetes) [VM to Kubernetes](https://api7.ai/solutions/vm-to-kubernetes)[- ](https://api7.ai/solutions/zero-trust-security) [Zero Trust Security](https://api7.ai/solutions/zero-trust-security) [Industry](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/financial-services) [Financial Services](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/iot) [IoT and Automotive](https://api7.ai/solutions/iot)[- ](https://api7.ai/solutions/blockchain) [Blockchain](https://api7.ai/solutions/blockchain) [- ](https://api7.ai/solutions/manufacturing) [Manufacturing](https://api7.ai/solutions/manufacturing) * [Customers](https://api7.ai/customers) * Pricing [- ](https://api7.ai/pricing) [API Gateway](https://api7.ai/pricing)[- ](https://api7.ai/ai-gateway/pricing) [AI Gateway](https://api7.ai/ai-gateway/pricing) * Resources [Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/ai-gateway/.md) [AISIX Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/api7-gateway) [API7 Gateway](https://docs.api7.ai/api7-gateway)[- ](https://docs.api7.ai/api7-gateway/ai-agent-skills.md) [API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[- ](https://docs.api7.ai/apisix) [Apache APISIX](https://docs.api7.ai/apisix)[- ](https://docs.api7.ai/apisix/ai-agent-skills.md) [APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md) [Compare](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-kong) [Apache APISIX vs Kong](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-nginx) [Apache APISIX vs NGINX](https://api7.ai/apisix-vs-nginx)[- ](https://api7.ai/api-gateway-comparison) [2026 Top API Gateway Comparison](https://api7.ai/api-gateway-comparison)[- ](https://api7.ai/ai-gateway-comparison) [AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) [Learn](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/openresty) [OpenResty (NGINX + Lua)](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/api-gateway-guide) [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[- ](https://api7.ai/learning-center/ai-gateway-guide) [AI Gateway Guide](https://api7.ai/learning-center/ai-gateway-guide)[- ](https://api7.ai/learning-center/api-infrastructure-guide) [API Infrastructure Guide](https://api7.ai/learning-center/api-infrastructure-guide) [Explore](https://api7.ai/demos)[- ](https://api7.ai/demos) [Demo Hub](https://api7.ai/demos)[- ](https://docs.api7.ai/hub.md) [Plugin Hub](https://docs.api7.ai/hub.md)[- ](https://api7.ai/category/usercase) [Case Studies](https://api7.ai/category/usercase) * [Blog](https://api7.ai/blog) Get a DemoStart for Free [](https://docs.api7.ai/)[Apache APISIX](https://docs.api7.ai/apisix/documentation.md)[API7 Gateway](https://docs.api7.ai/api7-gateway/overview.md)[Ingress Controller](https://docs.api7.ai/ingress-controller/documentation.md)[AISIX AI Gateway](https://docs.api7.ai/ai-gateway/.md)[Plugin Hub](https://docs.api7.ai/hub.md) Search # APISIX Control API Use the APISIX Control API to inspect or control the runtime state of one APISIX instance. It is enabled by default at http\://127.0.0.1:9090; configure apisix.enable control and apisix.control in... ## Schema * [`GET` Get JSON Schema `/v1/schema`](#fallback-operation-0-0) Get the JSON schema used by the gateway instance. Only loaded plugins are included. ## Health Check * [`GET` Get Health Check Information `/v1/healthcheck`](#fallback-operation-1-0) Get health check information of the APISIX instance. You need to initiate a request to the route to generate Control API health check information. * [`GET` Get Health Status By Type and ID `/v1/healthcheck/{src_type}/{src_id}`](#fallback-operation-1-1) Get health status of a specified resource. ## Garbage Collection * [`POST` Trigger Garbage Collection `/v1/gc`](#fallback-operation-2-0) Trigger a full garbage collection (GC) in the HTTP subsystem. Note that a request to this endpoint would not trigger a garbage collection in the stream subsystem because the subsystems are run in the separate Lua VM. ## Route * [`GET` Get All Routes `/v1/routes`](#fallback-operation-3-0) Get all configured routes. * [`GET` Get Route by ID `/v1/route/{route_id}`](#fallback-operation-3-1) Get a route by ID. ## Upstream * [`GET` Get All Upstreams `/v1/upstreams`](#fallback-operation-4-0) Get all configured upstreams. * [`GET` Get Upstream by ID `/v1/upstream/{upstream_id}`](#fallback-operation-4-1) Get an upstream by ID. ## Service * [`GET` Get All Services `/v1/services`](#fallback-operation-5-0) Get all configured services. * [`GET` Get Service by ID `/v1/service/{service_id}`](#fallback-operation-5-1) Get a service by ID. ## Plugin Metadata * [`GET` Get All Plugin Metadata `/v1/plugin_metadatas`](#fallback-operation-6-0) Get all plugin metadata. * [`GET` Get Plugin Metadata by Name `/v1/plugin_metadata/{plugin_name}`](#fallback-operation-6-1) Get plugin metadata by the plugin name. ## Plugin * [`PUT` Reload All Plugins `/v1/plugins/reload`](#fallback-operation-7-0) Hot reload plugins for changes to the plugin list in configuration files or plugin source files to take effect. ## Service Discovery * [`GET` Get Memory Dump `/v1/discovery/{service}/dump`](#fallback-operation-8-0) Get the in-memory service and configuration data for a configured service discovery module. This endpoint exists only when the module implements runtime data dumping. * [`GET` Get Dump File `/v1/discovery/{service}/show_dump_file`](#fallback-operation-8-1) Read the persisted service discovery dump. This endpoint is currently provided by compatible service discovery modules such as Consul and exists only when the module is configured. ## Server Information * [`GET` Get Server Information `/v1/server_info`](#fallback-operation-9-0) Get information about the local APISIX instance. This endpoint is available only when the server-info plugin is enabled. The plugin is deprecated in APISIX 3.17.0. ![API7.ai Logo](https://static.api7.ai/uploads/2025/03/02/api7.ai-white.avif) The digital world is connected by APIs,
API7.ai exists to make APIs more efficient, reliable, and secure. Sign up for API7 newsletter [Email address]()Subscribe Product [API7 Gateway](https://api7.ai/enterprise)[AISIX AI Gateway](https://api7.ai/ai-gateway)[API7 API Portal](https://api7.ai/portal) Learn [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[Plugin Hub](https://docs.api7.ai/hub.md)[API Gateway Comparison](https://api7.ai/api-gateway-comparison)[Customers](https://api7.ai/customers) Resources [API Gateway Docs](https://docs.api7.ai/apisix/documentation.md)[APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md)[API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[Blog](https://api7.ai/blog)[Demo Hub](https://api7.ai/demos)[APISIX vs Kong](https://api7.ai/apisix-vs-kong)[AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) Company [About](https://api7.ai/about)[Contact](https://api7.ai/contact)[Partners](https://api7.ai/partners)[Compliance Standards](https://api7.ai/compliance)[Brand Assets](https://api7.ai/branding)[Terms & Privacy](https://api7.ai/terms) *** [![SOC2 Type II](https://static.api7.ai/uploads/2025/03/02/fMrK6JR5_21972-312_SOC_NonCPA.avif)](https://api7.ai/compliance) [![ISO 27001](https://static.api7.ai/uploads/2025/03/02/kStGFFd2_iso-27001.avif)](https://api7.ai/compliance) [![HIPAA](https://static.api7.ai/uploads/2025/03/02/PSOVypp6_hipaa.avif)](https://api7.ai/compliance) [![GDPR](https://static.api7.ai/uploads/2025/03/02/6cv5RTfR_gdpr.avif)](https://api7.ai/compliance) [![Red Herring](https://static.api7.ai/uploads/2025/03/02/6385ad60e1f6a.avif)](https://api7.ai/blog/among-2022-red-herring-top-100-global) Copyright © APISEVEN PTE. LTD 2019 – 2026. Apache, Apache APISIX, APISIX, and associated open source project names are trademarks of the [Apache Software Foundation](https://www.apache.org/) [](https://www.linkedin.com/company/api7-ai/)[](https://github.com/api7)[](https://twitter.com/api7_ai) --- [Skip to main content](#__docusaurus_skipToContent_fallback) [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) ProductsSolutions[Customers](https://api7.ai/customers) Pricing Resources[Blog](https://api7.ai/blog) [Login](https://console.api7.cloud)Get a DemoStart for Free [![API7.ai Logo](https://static.api7.ai/2022/10/02/63398bceeeac7.webp?imageMogr2/thumbnail/240x)](https://api7.ai/) * Products [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/enterprise) [API7 Gateway](https://api7.ai/enterprise)[- ](https://api7.ai/api7-enterprise-vs-apisix) [Apache APISIX vs API7](https://api7.ai/api7-enterprise-vs-apisix)[- ](https://api7.ai/portal) [API7 API Portal](https://api7.ai/portal) [Apache APISIX](https://api7.ai/apisix)[- ](https://api7.ai/apisix) [What's Apache APISIX?](https://api7.ai/apisix)[- ](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway) [Why Apache APISIX?](https://api7.ai/blog/why-is-apache-apisix-the-best-api-gateway)[- ](https://api7.ai/apache-apisix-enterprise-support) [APISIX Commercial Support](https://api7.ai/apache-apisix-enterprise-support) [AISIX AI Gateway](https://api7.ai/ai-gateway)[- ](https://api7.ai/ai-gateway) [AISIX AI Gateway](https://api7.ai/ai-gateway) * Solutions [Developer](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/monolith-to-microservices) [Monolith to Microservices](https://api7.ai/solutions/monolith-to-microservices)[- ](https://api7.ai/solutions/on-prem-to-hybrid-cloud) [On-Prem to Hybrid Cloud](https://api7.ai/solutions/on-prem-to-hybrid-cloud)[- ](https://api7.ai/solutions/observability) [Observability](https://api7.ai/solutions/observability) [- ](https://api7.ai/solutions/vm-to-kubernetes) [VM to Kubernetes](https://api7.ai/solutions/vm-to-kubernetes)[- ](https://api7.ai/solutions/zero-trust-security) [Zero Trust Security](https://api7.ai/solutions/zero-trust-security) [Industry](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/financial-services) [Financial Services](https://api7.ai/solutions/financial-services)[- ](https://api7.ai/solutions/iot) [IoT and Automotive](https://api7.ai/solutions/iot)[- ](https://api7.ai/solutions/blockchain) [Blockchain](https://api7.ai/solutions/blockchain) [- ](https://api7.ai/solutions/manufacturing) [Manufacturing](https://api7.ai/solutions/manufacturing) * [Customers](https://api7.ai/customers) * Pricing [- ](https://api7.ai/pricing) [API Gateway](https://api7.ai/pricing)[- ](https://api7.ai/ai-gateway/pricing) [AI Gateway](https://api7.ai/ai-gateway/pricing) * Resources [Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/ai-gateway/.md) [AISIX Docs](https://docs.api7.ai/ai-gateway/.md)[- ](https://docs.api7.ai/api7-gateway) [API7 Gateway](https://docs.api7.ai/api7-gateway)[- ](https://docs.api7.ai/api7-gateway/ai-agent-skills.md) [API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[- ](https://docs.api7.ai/apisix) [Apache APISIX](https://docs.api7.ai/apisix)[- ](https://docs.api7.ai/apisix/ai-agent-skills.md) [APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md) [Compare](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-kong) [Apache APISIX vs Kong](https://api7.ai/apisix-vs-kong)[- ](https://api7.ai/apisix-vs-nginx) [Apache APISIX vs NGINX](https://api7.ai/apisix-vs-nginx)[- ](https://api7.ai/api-gateway-comparison) [2026 Top API Gateway Comparison](https://api7.ai/api-gateway-comparison)[- ](https://api7.ai/ai-gateway-comparison) [AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) [Learn](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/openresty) [OpenResty (NGINX + Lua)](https://api7.ai/learning-center/openresty)[- ](https://api7.ai/learning-center/api-gateway-guide) [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[- ](https://api7.ai/learning-center/ai-gateway-guide) [AI Gateway Guide](https://api7.ai/learning-center/ai-gateway-guide)[- ](https://api7.ai/learning-center/api-infrastructure-guide) [API Infrastructure Guide](https://api7.ai/learning-center/api-infrastructure-guide) [Explore](https://api7.ai/demos)[- ](https://api7.ai/demos) [Demo Hub](https://api7.ai/demos)[- ](https://docs.api7.ai/hub.md) [Plugin Hub](https://docs.api7.ai/hub.md)[- ](https://api7.ai/category/usercase) [Case Studies](https://api7.ai/category/usercase) * [Blog](https://api7.ai/blog) Get a DemoStart for Free [](https://docs.api7.ai/)[Apache APISIX](https://docs.api7.ai/apisix/documentation.md)[API7 Gateway](https://docs.api7.ai/api7-gateway/overview.md)[Ingress Controller](https://docs.api7.ai/ingress-controller/documentation.md)[AISIX AI Gateway](https://docs.api7.ai/ai-gateway/.md)[Plugin Hub](https://docs.api7.ai/hub.md) Search ![](https://static.api7.ai/uploads/2025/03/10/CYPjJmOl_bg-title-bg.avif) # Welcome to API Gateway Plugin Hub ![](https://static.api7.ai/uploads/2025/03/13/BrEQKYAn_plugin.avif) Discover powerful plugins to extend and enhance Apache APISIX’s capabilities for seamless API management [Plugin Overview](https://docs.api7.ai/apisix/key-concepts/plugins.md)|[Common Configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) ## AI[#](#ai) [![AI Aliyun Content Moderation](https://static.api7.ai/uploads/2025/03/13/tcMjtNjI_ai-aliyun-content-moderation.png)](https://docs.api7.ai/hub/ai-aliyun-content-moderation.md) #### [AI Aliyun Content Moderation](https://docs.api7.ai/hub/ai-aliyun-content-moderation.md) [The ai-aliyun-content-moderation plugin uses Aliyun to evaluate selected request roles and LLM responses, including streaming responses, against a risk threshold.](https://docs.api7.ai/hub/ai-aliyun-content-moderation.md) [![AI AWS Content Moderation](https://static.api7.ai/uploads/2025/03/13/86ShT776_ai-aws-content-moderation.png)](https://docs.api7.ai/hub/ai-aws-content-moderation.md) #### [AI AWS Content Moderation](https://docs.api7.ai/hub/ai-aws-content-moderation.md) [The ai-aws-content-moderation plugin uses Amazon Comprehend to detect toxicity in selected request roles and LLM responses, including streaming responses.](https://docs.api7.ai/hub/ai-aws-content-moderation.md) [![AI Cache](https://static.api7.ai/uploads/2024/03/11/ZWCmMDim_proxy-cache.png)](https://docs.api7.ai/hub/ai-cache.md) #### [AI Cache](https://docs.api7.ai/hub/ai-cache.md) [The ai-cache plugin stores exact and semantically similar LLM responses in Redis, reducing response latency and repeated upstream model usage.](https://docs.api7.ai/hub/ai-cache.md) [![AI Lakera Guard](https://static.api7.ai/uploads/2025/03/13/8v3RKxEA_ai-prompt-guard.png)](https://docs.api7.ai/hub/ai-lakera-guard.md) #### [AI Lakera Guard](https://docs.api7.ai/hub/ai-lakera-guard.md) [The ai-lakera-guard plugin screens AI traffic through the Lakera Guard API to detect prompt injection and other unsafe content in requests and LLM responses.](https://docs.api7.ai/hub/ai-lakera-guard.md) [![AI Prompt Decorator](https://static.api7.ai/uploads/2024/10/10/D9oM81AC_ai-prompt-decorator.png)](https://docs.api7.ai/hub/ai-prompt-decorator.md) #### [AI Prompt Decorator](https://docs.api7.ai/hub/ai-prompt-decorator.md) [The ai-prompt-decorator plugin decorates user prompts to LLMs by prefixing and appending pre-engineered prompts, streamlining API operation and content generation.](https://docs.api7.ai/hub/ai-prompt-decorator.md) [![AI Prompt Template](https://static.api7.ai/uploads/2024/10/09/nmOsFCYK_ai-prompt-template.png)](https://docs.api7.ai/hub/ai-prompt-template.md) #### [AI Prompt Template](https://docs.api7.ai/hub/ai-prompt-template.md) [The ai-prompt-template plugin supports pre-configured templates for user inputs to LLMs in a "fill in the blank" fashion, streamlining API management.](https://docs.api7.ai/hub/ai-prompt-template.md) [![AI RAG](https://static.api7.ai/uploads/2024/10/24/qoiBcXyS_ai-rag.png)](https://docs.api7.ai/hub/ai-rag.md) #### [AI RAG](https://docs.api7.ai/hub/ai-rag.md) [The ai-rag plugin retrieves context with Azure OpenAI embeddings and Azure AI Search before an LLM request is proxied.](https://docs.api7.ai/hub/ai-rag.md) [![AI Proxy](https://static.api7.ai/uploads/2024/10/10/hj8UBb3W_ai-proxy.png)](https://docs.api7.ai/hub/ai-proxy.md) #### [AI Proxy](https://docs.api7.ai/hub/ai-proxy.md) [The ai-proxy plugin simplifies access to LLM and embedding models providers by converting plugin configurations into the required request format for OpenAI, DeepSeek, Anthropic, and other OpenAI-compatible APIs.](https://docs.api7.ai/hub/ai-proxy.md) [![AI Prompt Guard](https://static.api7.ai/uploads/2025/03/13/8v3RKxEA_ai-prompt-guard.png)](https://docs.api7.ai/hub/ai-prompt-guard.md) #### [AI Prompt Guard](https://docs.api7.ai/hub/ai-prompt-guard.md) [The ai-prompt-guard plugin safeguards prompts to LLM using allow/deny patterns, ensuring only approved inputs pass. It can check the latest message or full history.](https://docs.api7.ai/hub/ai-prompt-guard.md) [![AI Proxy Multi](https://static.api7.ai/uploads/2025/03/05/0Dzb0DI3_ai-proxy-multi.png)](https://docs.api7.ai/hub/ai-proxy-multi.md) #### [AI Proxy Multi](https://docs.api7.ai/hub/ai-proxy-multi.md) [The ai-proxy-multi plugin extends the capabilities of ai-proxy with load balancing, retries, fallbacks, and health checks, simplifying the integration with OpenAI, DeepSeek, and other OpenAI-compatible APIs.](https://docs.api7.ai/hub/ai-proxy-multi.md) [![AI Rate Limiting](https://static.api7.ai/uploads/2025/03/05/CJqBH5Rp_ai-rate-limiting.png)](https://docs.api7.ai/hub/ai-rate-limiting.md) #### [AI Rate Limiting](https://docs.api7.ai/hub/ai-rate-limiting.md) [The ai-rate-limiting plugin enforces token-based rate limiting for LLM service requests, preventing overuse, optimizing API consumption, and ensuring efficient resource allocation.](https://docs.api7.ai/hub/ai-rate-limiting.md) [![AI Request Rewrite](https://static.api7.ai/uploads/2025/04/07/bbcNos7x_ai-request-rewrite.png)](https://docs.api7.ai/hub/ai-request-rewrite.md) #### [AI Request Rewrite](https://docs.api7.ai/hub/ai-request-rewrite.md) [The ai-request-rewrite plugin forwards client requests to LLM services for processing before sending them upstream, enabling AI-driven redaction, enrichment, and reformatting.](https://docs.api7.ai/hub/ai-request-rewrite.md) [![OpenAPI to MCP](https://static.api7.ai/uploads/2025/09/22/rxyJgPnj_tmp-icon.png)Enterprise](https://docs.api7.ai/hub/openapi-to-mcp.md) #### [OpenAPI to MCP](https://docs.api7.ai/hub/openapi-to-mcp.md) [The openapi-to-mcp plugin lets API7 expose OpenAPI services through MCP, proxy requests with custom headers, and support real-time SSE streaming.](https://docs.api7.ai/hub/openapi-to-mcp.md) ## Traffic Management[#](#traffic-management) [![GraphQL Limit Count](https://static.api7.ai/uploads/2023/10/18/4qPNgRwS_test_2.8.png)](https://docs.api7.ai/hub/graphql-limit-count.md) #### [GraphQL Limit Count](https://docs.api7.ai/hub/graphql-limit-count.md) [The graphql-limit-count plugin uses fixed windows to limit accumulated GraphQL document cost, with selection depth as the default measure.](https://docs.api7.ai/hub/graphql-limit-count.md) [![GraphQL Proxy Cache](https://static.api7.ai/uploads/2023/10/18/T9HdQS9U_test_2.6.png)](https://docs.api7.ai/hub/graphql-proxy-cache.md) #### [GraphQL Proxy Cache](https://docs.api7.ai/hub/graphql-proxy-cache.md) [The graphql-proxy-cache plugin enables caching of responses for GraphQL queries, improving API performance.](https://docs.api7.ai/hub/graphql-proxy-cache.md) [![Limit Count Advanced](https://static.api7.ai/uploads/2024/10/31/QdYZeJh7_limit-count-advanced.png)Enterprise](https://docs.api7.ai/hub/limit-count-advanced.md) #### [Limit Count Advanced](https://docs.api7.ai/hub/limit-count-advanced.md) [The limit-count-advanced plugin enforces API rate limiting with a fixed window or sliding window algorithm, restricting requests within a time window. Requests over the quota are rejected.](https://docs.api7.ai/hub/limit-count-advanced.md) [![Limit Req](https://static.api7.ai/uploads/2023/11/01/2CZt13jD_limit-req.png)](https://docs.api7.ai/hub/limit-req.md) #### [Limit Req](https://docs.api7.ai/hub/limit-req.md) [The limit-req plugin enforces API rate limiting with a leaky bucket algorithm to rate limit requests, enabling effective throttling to manage traffic flow.](https://docs.api7.ai/hub/limit-req.md) [![Limit Conn](https://static.api7.ai/uploads/2024/12/12/2OJrqIVM_limit-conn.png)](https://docs.api7.ai/hub/limit-conn.md) #### [Limit Conn](https://docs.api7.ai/hub/limit-conn.md) [The limit-conn plugin restricts the rate of requests by managing concurrent connections. Requests exceeding the threshold may be delayed or rejected, ensuring controlled API usage and preventing overload.](https://docs.api7.ai/hub/limit-conn.md) [![Limit Count](https://static.api7.ai/uploads/2023/10/27/wLA7vu2k_limit-count.png)](https://docs.api7.ai/hub/limit-count.md) #### [Limit Count](https://docs.api7.ai/hub/limit-count.md) [The limit-count plugin enforces API rate limiting with a fixed window algorithm, restricting requests within a time interval. Requests over the quota are rejected.](https://docs.api7.ai/hub/limit-count.md) [![OAS Validator](https://static.api7.ai/uploads/2023/11/08/Fu3KAzpq_oas-validator.png)](https://docs.api7.ai/hub/oas-validator.md) #### [OAS Validator](https://docs.api7.ai/hub/oas-validator.md) [The oas-validator plugin checks incoming HTTP requests against an OpenAPI specification before they are forwarded to upstream services.](https://docs.api7.ai/hub/oas-validator.md) [![Proxy Buffering](https://static.api7.ai/uploads/2023/10/25/03goqJor_proxy_buffering.png)](https://docs.api7.ai/hub/proxy-buffering.md) #### [Proxy Buffering](https://docs.api7.ai/hub/proxy-buffering.md) [The proxy-buffering plugin dynamically disables NGINX proxy buffering, optimizing performance with SSE and other streaming upstream services in API gateway.](https://docs.api7.ai/hub/proxy-buffering.md) [![Proxy Cache](https://static.api7.ai/uploads/2024/03/11/ZWCmMDim_proxy-cache.png)](https://docs.api7.ai/hub/proxy-cache.md) #### [Proxy Cache](https://docs.api7.ai/hub/proxy-cache.md) [The proxy-cache plugin caches responses based on keys, supporting disk and memory caching for GET, POST, and HEAD requests, enhancing API performance.](https://docs.api7.ai/hub/proxy-cache.md) [![Proxy Mirror](https://static.api7.ai/uploads/2024/01/29/Hxr9HkkD_proxy-mirror.png)](https://docs.api7.ai/hub/proxy-mirror.md) #### [Proxy Mirror](https://docs.api7.ai/hub/proxy-mirror.md) [The proxy-mirror plugin duplicates ingress traffic to API gateway, forwarding it to a designated upstream while keeping regular services uninterrupted.](https://docs.api7.ai/hub/proxy-mirror.md) [![Request ID](https://static.api7.ai/uploads/2024/12/09/1U8hV11N_request-id.png)](https://docs.api7.ai/hub/request-id.md) #### [Request ID](https://docs.api7.ai/hub/request-id.md) [The request-id plugin adds a unique ID to each request proxied through the API gateway, facilitating effective tracking of API requests for better API management.](https://docs.api7.ai/hub/request-id.md) [![Request Validation](https://static.api7.ai/uploads/2023/12/07/bzStlCLK_request-validation.png)](https://docs.api7.ai/hub/request-validation.md) #### [Request Validation](https://docs.api7.ai/hub/request-validation.md) [The request-validation plugin checks requests for compliance before forwarding them to upstream services, enhancing security in API operations.](https://docs.api7.ai/hub/request-validation.md) [![Traffic Label](https://static.api7.ai/uploads/2023/10/18/IRDVY9sk_test_2.10.png)](https://docs.api7.ai/hub/traffic-label.md) #### [Traffic Label](https://docs.api7.ai/hub/traffic-label.md) [The traffic-label plugin labels traffic based on user-defined rules, enabling actions based on labels and associated weights for improved API traffic management.](https://docs.api7.ai/hub/traffic-label.md) [![Workflow](https://static.api7.ai/uploads/2024/03/05/CTQ2O8NF_workflow.png)](https://docs.api7.ai/hub/workflow.md) #### [Workflow](https://docs.api7.ai/hub/workflow.md) [The workflow plugin enables conditional execution of user-defined actions on client traffic based on specific rules, allowing granular API traffic management.](https://docs.api7.ai/hub/workflow.md) [![Traffic Split](https://static.api7.ai/uploads/2023/11/01/tYpqD1OB_traffic-split.png)](https://docs.api7.ai/hub/traffic-split.md) #### [Traffic Split](https://docs.api7.ai/hub/traffic-split.md) [The traffic-split plugin directs traffic to multiple upstream services based on conditions or weights, providing a flexible approach for API release strategies and traffic management.](https://docs.api7.ai/hub/traffic-split.md) ## Transformation[#](#transformation) [![Attach Consumer Label](https://static.api7.ai/uploads/2024/09/14/dNILru0T_update-icon-bcg.jpeg)](https://docs.api7.ai/hub/attach-consumer-label.md) #### [Attach Consumer Label](https://docs.api7.ai/hub/attach-consumer-label.md) [The attach-consumer-label plugin attaches custom consumer labels to authenticated requests, for upstream services to implement additional business logics.](https://docs.api7.ai/hub/attach-consumer-label.md) [![Body Transformer](https://static.api7.ai/uploads/2023/11/07/e7rKZFn9_body-transformer.png)](https://docs.api7.ai/hub/body-transformer.md) #### [Body Transformer](https://docs.api7.ai/hub/body-transformer.md) [The body-transformer plugin converts request and response bodies between formats, such as JSON to XML, facilitating seamless data exchange.](https://docs.api7.ai/hub/body-transformer.md) [![degraphql](https://static.api7.ai/uploads/2023/12/06/kms6zvcY_degraphql.png)](https://docs.api7.ai/hub/degraphql.md) #### [degraphql](https://docs.api7.ai/hub/degraphql.md) [The degraphql plugin enables communication with upstream GraphQL services through standard HTTP requests by mapping GraphQL queries to HTTP endpoints, simplifying API integration.](https://docs.api7.ai/hub/degraphql.md) [![Exit transformer](https://static.api7.ai/uploads/2024/10/25/R6XnlT7J_exit-transformer.png)](https://docs.api7.ai/hub/exit-transformer.md) #### [Exit transformer](https://docs.api7.ai/hub/exit-transformer.md) [The exit-transformer plugin customizes responses generated by gateway plugins or missing routes before APISIX sends them to clients.](https://docs.api7.ai/hub/exit-transformer.md) [![Fault Injection](https://static.api7.ai/uploads/2025/01/22/m6b1DCku_fault-injection.png)](https://docs.api7.ai/hub/fault-injection.md) #### [Fault Injection](https://docs.api7.ai/hub/fault-injection.md) [The fault-injection plugin tests application resiliency by simulating controlled faults or delays, making it ideal for chaos engineering and failure condition analysis.](https://docs.api7.ai/hub/fault-injection.md) [![gRPC Transcode](https://static.api7.ai/uploads/2024/02/21/3uzsq5N4_grpc-transcode.png)](https://docs.api7.ai/hub/grpc-transcode.md) #### [gRPC Transcode](https://docs.api7.ai/hub/grpc-transcode.md) [The grpc-transcode plugin converts between HTTP and gRPC requests and responses, facilitating seamless communication between different API protocols.](https://docs.api7.ai/hub/grpc-transcode.md) [![gRPC Web](https://static.api7.ai/uploads/2026/01/06/al39Dafe_gRPC-web.png)](https://docs.api7.ai/hub/grpc-web.md) #### [gRPC Web](https://docs.api7.ai/hub/grpc-web.md) [The grpc-web plugin enables the gateway to handle gRPC-Web requests from browsers and JavaScript clients by translating them into standard gRPC calls and forwarding them to upstream gRPC services.](https://docs.api7.ai/hub/grpc-web.md) [![Mocking](https://static.api7.ai/uploads/2025/01/24/WttCR3KD_mocking.png)](https://docs.api7.ai/hub/mocking.md) #### [Mocking](https://docs.api7.ai/hub/mocking.md) [The mocking plugin simulates API responses without forwarding requests to upstream services, offering customization of status codes, response bodies, headers, and more for API testing and development.](https://docs.api7.ai/hub/mocking.md) [![Proxy Rewrite](https://static.api7.ai/uploads/2023/10/18/yCccWhP0_test_2.11.png)](https://docs.api7.ai/hub/proxy-rewrite.md) #### [Proxy Rewrite](https://docs.api7.ai/hub/proxy-rewrite.md) [The proxy-rewrite plugin offers flexible options to rewrite requests that API gateway forwards to upstream services, enhancing API management.](https://docs.api7.ai/hub/proxy-rewrite.md) [![Response Rewrite](https://static.api7.ai/uploads/2023/10/25/Nm2E852J_rr.png)](https://docs.api7.ai/hub/response-rewrite.md) #### [Response Rewrite](https://docs.api7.ai/hub/response-rewrite.md) [The response-rewrite plugin allows rewriting of responses from API gateway and upstream services, providing flexibility in API responses.](https://docs.api7.ai/hub/response-rewrite.md) [![SOAP](https://static.api7.ai/uploads/2023/10/18/V3vWj7A6_test_2.2.png)Enterprise](https://docs.api7.ai/hub/soap.md) #### [SOAP](https://docs.api7.ai/hub/soap.md) [The soap plugin simplifies transformation between RESTful HTTP requests and SOAP requests, including their corresponding responses, for better API interoperability.](https://docs.api7.ai/hub/soap.md) ## Authentication[#](#authentication) [![Authz Keycloak](https://static.api7.ai/uploads/2023/12/08/pqqgJ2YO_keycloak-authz.png)](https://docs.api7.ai/hub/authz-keycloak.md) #### [Authz Keycloak](https://docs.api7.ai/hub/authz-keycloak.md) [The authz-keycloak plugin integrates with Keycloak for user authentication and authorization, enhancing API security and management.](https://docs.api7.ai/hub/authz-keycloak.md) [![Basic Auth](https://static.api7.ai/uploads/2024/08/23/YwSkGnhY_basic-auth.png)](https://docs.api7.ai/hub/basic-auth.md) #### [Basic Auth](https://docs.api7.ai/hub/basic-auth.md) [The basic-auth plugin provides basic access authentication, requiring clients to authenticate before accessing upstream resources, enhancing API security.](https://docs.api7.ai/hub/basic-auth.md) [![Forward Auth](https://static.api7.ai/uploads/2024/04/29/UNSKKKqr_forward-auth.png)](https://docs.api7.ai/hub/forward-auth.md) #### [Forward Auth](https://docs.api7.ai/hub/forward-auth.md) [The forward-auth plugin integrates with external authorization services, enhancing API security and access control.](https://docs.api7.ai/hub/forward-auth.md) [![JWT Auth](https://static.api7.ai/uploads/2024/03/20/3Dw978og_jwt-auth.png)](https://docs.api7.ai/hub/jwt-auth.md) #### [JWT Auth](https://docs.api7.ai/hub/jwt-auth.md) [The jwt-auth plugin supports the use of JSON Web Token (JWT) for client authentication before accessing upstream resources, enhancing API security measures.](https://docs.api7.ai/hub/jwt-auth.md) [![JWE Decrypt](https://static.api7.ai/uploads/2024/01/15/8AkaEKui_jwe-icon.png)](https://docs.api7.ai/hub/jwe-decrypt.md) #### [JWE Decrypt](https://docs.api7.ai/hub/jwe-decrypt.md) [The jwe-decrypt plugin decrypts its supported five-part compact token format and forwards the plaintext in a configured request header.](https://docs.api7.ai/hub/jwe-decrypt.md) [![Key Auth](https://static.api7.ai/uploads/2024/03/20/DRWFmK4D_key-auth.png)](https://docs.api7.ai/hub/key-auth.md) #### [Key Auth](https://docs.api7.ai/hub/key-auth.md) [The key-auth plugin allows clients to authenticate using an authentication key before accessing upstream resources, enhancing API security measures.](https://docs.api7.ai/hub/key-auth.md) [![HMAC Auth](https://static.api7.ai/uploads/2024/09/03/xG9Vqxl5_hmac-auth.png)](https://docs.api7.ai/hub/hmac-auth.md) #### [HMAC Auth](https://docs.api7.ai/hub/hmac-auth.md) [The hmac-auth plugin supports HMAC authentication to ensure request integrity, preventing modifications during transmission and enhancing API security.](https://docs.api7.ai/hub/hmac-auth.md) [![LDAP Auth Advanced](https://static.api7.ai/uploads/2024/08/23/YwSkGnhY_basic-auth.png)](https://docs.api7.ai/hub/ldap-auth-advanced.md) #### [LDAP Auth Advanced](https://docs.api7.ai/hub/ldap-auth-advanced.md) [The ldap-auth-advanced plugin authenticates clients against an LDAP directory and maps the authenticated user onto a consumer, so directory identities can be used with per-consumer plugins, rate limits, and analytics.](https://docs.api7.ai/hub/ldap-auth-advanced.md) [![Multi Auth](https://static.api7.ai/uploads/2024/01/15/d4Lo3cix_multi-auth.png)](https://docs.api7.ai/hub/multi-auth.md) #### [Multi Auth](https://docs.api7.ai/hub/multi-auth.md) [The multi-auth plugin enables consumers using diverse authentication methods to share the same route or service, streamlining API lifecycle management.](https://docs.api7.ai/hub/multi-auth.md) [![OpenID Connect](https://static.api7.ai/uploads/2023/11/15/TdzRQl8n_oidc.png)](https://docs.api7.ai/hub/openid-connect.md) #### [OpenID Connect](https://docs.api7.ai/hub/openid-connect.md) [The openid-connect plugin integrates with OIDC providers like Keycloak and Auth0, simplifying user authentication in API management.](https://docs.api7.ai/hub/openid-connect.md) [![SAML Auth](https://static.api7.ai/uploads/2024/08/23/pYpD0ncp_saml-auth.png)](https://docs.api7.ai/hub/saml-auth.md) #### [SAML Auth](https://docs.api7.ai/hub/saml-auth.md) [The saml-auth plugin enables user authentication via SAML 2.0 in the API gateway by interacting with identity providers (IdP), enhancing API security.](https://docs.api7.ai/hub/saml-auth.md) [![OPA](https://static.api7.ai/uploads/2024/08/23/BdRXak7x_opa.png)](https://docs.api7.ai/hub/opa.md) #### [OPA](https://docs.api7.ai/hub/opa.md) [The opa plugin integrates with Open Policy Agent, enabling unified policy definition and enforcement for authorization in API operations.](https://docs.api7.ai/hub/opa.md) ## Security[#](#security) [![ACL](https://static.api7.ai/uploads/2024/02/04/dnlZbDhq_acl.png)](https://docs.api7.ai/hub/acl.md) #### [ACL](https://docs.api7.ai/hub/acl.md) [The acl plugin controls access to upstream resources by verifying if the user is on the access control lists, enhancing API management.](https://docs.api7.ai/hub/acl.md) [![Chaitin WAF](https://static.api7.ai/uploads/2025/09/08/8mi7erpW_chaitin-waf.png)](https://docs.api7.ai/hub/chaitin-waf.md) #### [Chaitin WAF](https://docs.api7.ai/hub/chaitin-waf.md) [The chaitin-waf plugin integrates with Chaitin WAF (SafeLine) to detect and block web threats, strengthening application security and protecting user data.](https://docs.api7.ai/hub/chaitin-waf.md) [![CORS](https://static.api7.ai/uploads/2024/02/01/VGYw3KjS_20240201-095424.jpeg)](https://docs.api7.ai/hub/cors.md) #### [CORS](https://docs.api7.ai/hub/cors.md) [The cors plugin enables cross-origin resource sharing, allowing servers to specify permitted origins and instructing browsers to load resources from those origins, enhancing API accessibility.](https://docs.api7.ai/hub/cors.md) [![Data Mask](https://static.api7.ai/uploads/2024/04/01/2JlB10Xi_data-mask.png)](https://docs.api7.ai/hub/data-mask.md) #### [Data Mask](https://docs.api7.ai/hub/data-mask.md) [The data-mask plugin removes or replaces sensitive information in request headers, bodies, and URL queries for logging purposes, enhancing data privacy and security.](https://docs.api7.ai/hub/data-mask.md) [![Consumer Restriction](https://static.api7.ai/uploads/2024/03/23/2Or0rpNg_consumer-restriction.png)](https://docs.api7.ai/hub/consumer-restriction.md) #### [Consumer Restriction](https://docs.api7.ai/hub/consumer-restriction.md) [The consumer-restriction plugin implements access controls based on consumer name, route ID, service ID, or consumer group ID, enhancing API security.](https://docs.api7.ai/hub/consumer-restriction.md) [![IP Restriction](https://static.api7.ai/uploads/2024/02/28/0WSO7GZL_ip-res.png)](https://docs.api7.ai/hub/ip-restriction.md) #### [IP Restriction](https://docs.api7.ai/hub/ip-restriction.md) [The ip-restriction plugin restricts access to upstream resources based on an IP address whitelist or blacklist, improving API security.](https://docs.api7.ai/hub/ip-restriction.md) [![MCP Tools ACL](https://static.api7.ai/uploads/2024/02/04/dnlZbDhq_acl.png)Enterprise](https://docs.api7.ai/hub/mcp-tools-acl.md) #### [MCP Tools ACL](https://docs.api7.ai/hub/mcp-tools-acl.md) [The mcp-tools-acl plugin provides per-consumer access control for MCP tool calls on routes powered by openapi-to-mcp, supporting rule-based allowlist and denylist modes with optional expression conditions.](https://docs.api7.ai/hub/mcp-tools-acl.md) [![UA Restriction](https://static.api7.ai/uploads/2024/02/29/Y0qfz6CT_ua-restriction.png)](https://docs.api7.ai/hub/ua-restriction.md) #### [UA Restriction](https://docs.api7.ai/hub/ua-restriction.md) [The ua-restriction plugin restricts access to upstream resources using an allowlist or denylist of user agents, preventing overload from web crawlers and enhancing API security.](https://docs.api7.ai/hub/ua-restriction.md) ## Observability[#](#observability) [![ClickHouse Logger](https://static.api7.ai/uploads/2023/12/19/HYDN4Dmw_clickhouse.png)](https://docs.api7.ai/hub/clickhouse-logger.md) #### [ClickHouse Logger](https://docs.api7.ai/hub/clickhouse-logger.md) [The clickhouse-logger plugin pushes request and response logs to ClickHouse databases in batches, allowing for customizable log formats to enhance data management.](https://docs.api7.ai/hub/clickhouse-logger.md) [![Datadog](https://static.api7.ai/uploads/2024/01/18/s7gmob0S_datadog.png)](https://docs.api7.ai/hub/datadog.md) #### [Datadog](https://docs.api7.ai/hub/datadog.md) [The datadog plugin integrates with Datadog, sending metrics to DogStatsD in batches to improve API monitoring and API performance tracking.](https://docs.api7.ai/hub/datadog.md) [![Elasticsearch Logger](https://static.api7.ai/uploads/2025/01/13/9VsL3UBH_sukBLfk2_elasticsearch-logger.png)](https://docs.api7.ai/hub/elasticsearch-logger.md) #### [Elasticsearch Logger](https://docs.api7.ai/hub/elasticsearch-logger.md) [The elasticsearch-logger plugin pushes request and response logs in batches to Elasticsearch, allowing for customizable log formats to enhance data management.](https://docs.api7.ai/hub/elasticsearch-logger.md) [![Error Log Collect](/img/plugins/error-log-collect.png)Enterprise](https://docs.api7.ai/hub/error-log-collect.md) #### [Error Log Collect](https://docs.api7.ai/hub/error-log-collect.md) [The error-log-collect plugin captures the error logs produced while processing selected requests, including lower-severity entries that the configured log level would otherwise discard, and writes them to the gateway error log for targeted debugging.](https://docs.api7.ai/hub/error-log-collect.md) [![Error Log Logger](https://static.api7.ai/uploads/2025/01/26/9gRLjv8N_error-log-logger.png)](https://docs.api7.ai/hub/error-log-logger.md) #### [Error Log Logger](https://docs.api7.ai/hub/error-log-logger.md) [The error-log-logger plugin pushes APISIX's error logs to TCP, Apache SkyWalking, Apache Kafka, or ClickHouse servers, in batches. You can specify the severity level of which the plugin should send the corresponding logs.](https://docs.api7.ai/hub/error-log-logger.md) [![Google Cloud Logging](https://static.api7.ai/uploads/2025/02/06/4qDGkFhw_google-cloud.png)](https://docs.api7.ai/hub/google-cloud-logging.md) #### [Google Cloud Logging](https://docs.api7.ai/hub/google-cloud-logging.md) [The google-cloud-logging plugin pushes request and response logs in batches to Google Cloud Logging Service and supports the customization of log formats.](https://docs.api7.ai/hub/google-cloud-logging.md) [![HTTP Logger](https://static.api7.ai/uploads/2025/01/20/XhSZolpi_http-logger.png)](https://docs.api7.ai/hub/http-logger.md) #### [HTTP Logger](https://docs.api7.ai/hub/http-logger.md) [The http-logger plugin pushes request and response logs as JSON objects to HTTP(S) servers in batches, allowing for customizable log formats to enhance data management.](https://docs.api7.ai/hub/http-logger.md) [![Kafka Logger](https://static.api7.ai/uploads/2025/01/20/ipzxNV6T_kafka-logger.png)](https://docs.api7.ai/hub/kafka-logger.md) #### [Kafka Logger](https://docs.api7.ai/hub/kafka-logger.md) [The kafka-logger plugin pushes request and response logs as JSON objects to Apache Kafka clusters in batches, allowing for customizable log formats to enhance data management.](https://docs.api7.ai/hub/kafka-logger.md) [![Loki Logger](https://static.api7.ai/uploads/2024/12/27/YXMsoZns_output-2.png)](https://docs.api7.ai/hub/loki-logger.md) #### [Loki Logger](https://docs.api7.ai/hub/loki-logger.md) [The loki-logger plugin sends request and response logs as JSON objects to Grafana Loki in batches via the Loki HTTP API, allowing for customizable log formats to enhance data management.](https://docs.api7.ai/hub/loki-logger.md) [![Prometheus](https://static.api7.ai/uploads/2023/10/18/Ep8zNIyd_test_2.5.png)](https://docs.api7.ai/hub/prometheus.md) #### [Prometheus](https://docs.api7.ai/hub/prometheus.md) [The Prometheus plugin integrates with Prometheus for metric collection and continuous monitoring, enhancing API observability.](https://docs.api7.ai/hub/prometheus.md) [![OpenTelemetry](https://static.api7.ai/uploads/2024/02/17/TIoDf61O_otel-icon.png)](https://docs.api7.ai/hub/opentelemetry.md) #### [OpenTelemetry](https://docs.api7.ai/hub/opentelemetry.md) [The opentelemetry plugin instruments the API gateway, sending traces to the OpenTelemetry collector for monitoring API operations per OpenTelemetry specs.](https://docs.api7.ai/hub/opentelemetry.md) [![RocketMQ Logger](https://static.api7.ai/uploads/2023/12/18/JKMT7TOw_rocketmq.png)](https://docs.api7.ai/hub/rocketmq-logger.md) #### [RocketMQ Logger](https://docs.api7.ai/hub/rocketmq-logger.md) [The rocketmq-logger plugin pushes request and response logs as JSON objects to RocketMQ clusters in batches, allowing for customizable log formats to enhance data management.](https://docs.api7.ai/hub/rocketmq-logger.md) [![SkyWalking Logger](https://static.api7.ai/uploads/2025/01/20/bRD2lASg_SkyWalking-logger.png)](https://docs.api7.ai/hub/skywalking-logger.md) #### [SkyWalking Logger](https://docs.api7.ai/hub/skywalking-logger.md) [The skywalking-logger pushes request and response logs as JSON objects to SkyWalking OAP server in batches, allowing for customizable log formats to enhance data management.](https://docs.api7.ai/hub/skywalking-logger.md) [![syslog](https://static.api7.ai/uploads/2024/03/01/wq96rSoX_Syslog.png)](https://docs.api7.ai/hub/syslog.md) #### [syslog](https://docs.api7.ai/hub/syslog.md) [The syslog plugin pushes request and response logs as JSON objects to syslog servers in batches, allowing for customizable log formats to enhance data management.](https://docs.api7.ai/hub/syslog.md) [![SkyWalking](https://static.api7.ai/uploads/2025/01/15/JGYsNvMp_SkyWalking.png)](https://docs.api7.ai/hub/skywalking.md) #### [SkyWalking](https://docs.api7.ai/hub/skywalking.md) [The skywalking plugin integrates with Apache SkyWalking for effective request tracing, enhancing API observability.](https://docs.api7.ai/hub/skywalking.md) [![Splunk HEC Logging](https://static.api7.ai/uploads/2024/11/28/wFKN9wIR_output.png)](https://docs.api7.ai/hub/splunk-hec-logging.md) #### [Splunk HEC Logging](https://docs.api7.ai/hub/splunk-hec-logging.md) [The splunk-hec-logging plugin serializes request and response context information to Splunk Event Data format and push to your Splunk HTTP Event Collector (HEC) in batches, allowing for customizable log formats to enhance data management.](https://docs.api7.ai/hub/splunk-hec-logging.md) [![Zipkin](https://static.api7.ai/uploads/2024/01/23/K5z0SiCS_zipkin-white.png)](https://docs.api7.ai/hub/zipkin.md) #### [Zipkin](https://docs.api7.ai/hub/zipkin.md) [The zipkin plugin instruments the API gateway to send traces to Zipkin or compatible collectors like Jaeger and Apache SkyWalking, enhancing request tracing capabilities.](https://docs.api7.ai/hub/zipkin.md) ## General[#](#general) [![Error Page](https://static.api7.ai/uploads/2024/03/15/yh6VviFP_error-page.png)](https://docs.api7.ai/hub/error-page.md) #### [Error Page](https://docs.api7.ai/hub/error-page.md) [The error-page plugin customizes gateway-generated 404, 500, 502, and 503 responses without modifying responses returned by upstream services.](https://docs.api7.ai/hub/error-page.md) [![Public API](https://static.api7.ai/uploads/2024/01/26/OJVqobOZ_public-api.png)](https://docs.api7.ai/hub/public-api.md) #### [Public API](https://docs.api7.ai/hub/public-api.md) [The public-api plugin exposes internal API endpoints, allowing external access while maintaining control over API management and security.](https://docs.api7.ai/hub/public-api.md) [![Real IP](https://static.api7.ai/uploads/2023/11/01/fi20fydE_real-ip.png)](https://docs.api7.ai/hub/real-ip.md) #### [Real IP](https://docs.api7.ai/hub/real-ip.md) [The real-ip plugin enables the API gateway to fetch the client's real IP using the IP address from the HTTP header or query string, improving data quality.](https://docs.api7.ai/hub/real-ip.md) ## Serverless[#](#serverless) [![AWS Lambda](https://static.api7.ai/uploads/2024/04/26/FjXGfhOO_aws-lambda.png)](https://docs.api7.ai/hub/aws-lambda.md) #### [AWS Lambda](https://docs.api7.ai/hub/aws-lambda.md) [The aws-lambda plugin simplifies APISIX integration with AWS Lambda and Amazon API gateway, supporting authentication via IAM user credentials and API keys.](https://docs.api7.ai/hub/aws-lambda.md) [![Serverless Functions](https://static.api7.ai/uploads/2024/05/09/2NzjIiNI_serverless-funcs.png)](https://docs.api7.ai/hub/serverless-functions.md) #### [Serverless Functions](https://docs.api7.ai/hub/serverless-functions.md) [The serverless function plugins (pre-function and post-function) allow execution of user-defined logic at the start or end of specified execution phases in API gateway.](https://docs.api7.ai/hub/serverless-functions.md) ## Other Protocols[#](#other-protocols) [![MQTT Proxy](https://static.api7.ai/uploads/2024/09/19/4S2PT5Bc_mqtt.png)](https://docs.api7.ai/hub/mqtt-proxy.md) #### [MQTT Proxy](https://docs.api7.ai/hub/mqtt-proxy.md) [The mqtt-proxy plugin supports proxying and load balancing MQTT requests to MQTT servers, enhancing API operation and management.](https://docs.api7.ai/hub/mqtt-proxy.md) ![API7.ai Logo](https://static.api7.ai/uploads/2025/03/02/api7.ai-white.avif) The digital world is connected by APIs,
API7.ai exists to make APIs more efficient, reliable, and secure. Sign up for API7 newsletter [Email address]()Subscribe Product [API7 Gateway](https://api7.ai/enterprise)[AISIX AI Gateway](https://api7.ai/ai-gateway)[API7 API Portal](https://api7.ai/portal) Learn [API Gateway Guide](https://api7.ai/learning-center/api-gateway-guide)[Plugin Hub](https://docs.api7.ai/hub.md)[API Gateway Comparison](https://api7.ai/api-gateway-comparison)[Customers](https://api7.ai/customers) Resources [API Gateway Docs](https://docs.api7.ai/apisix/documentation.md)[APISIX AI Agent Skills](https://docs.api7.ai/apisix/ai-agent-skills.md)[API7 AI Agent Skills](https://docs.api7.ai/api7-gateway/ai-agent-skills.md)[Blog](https://api7.ai/blog)[Demo Hub](https://api7.ai/demos)[APISIX vs Kong](https://api7.ai/apisix-vs-kong)[AI Gateway Comparison](https://api7.ai/ai-gateway-comparison) Company [About](https://api7.ai/about)[Contact](https://api7.ai/contact)[Partners](https://api7.ai/partners)[Compliance Standards](https://api7.ai/compliance)[Brand Assets](https://api7.ai/branding)[Terms & Privacy](https://api7.ai/terms) *** [![SOC2 Type II](https://static.api7.ai/uploads/2025/03/02/fMrK6JR5_21972-312_SOC_NonCPA.avif)](https://api7.ai/compliance) [![ISO 27001](https://static.api7.ai/uploads/2025/03/02/kStGFFd2_iso-27001.avif)](https://api7.ai/compliance) [![HIPAA](https://static.api7.ai/uploads/2025/03/02/PSOVypp6_hipaa.avif)](https://api7.ai/compliance) [![GDPR](https://static.api7.ai/uploads/2025/03/02/6cv5RTfR_gdpr.avif)](https://api7.ai/compliance) [![Red Herring](https://static.api7.ai/uploads/2025/03/02/6385ad60e1f6a.avif)](https://api7.ai/blog/among-2022-red-herring-top-100-global) Copyright © APISEVEN PTE. LTD 2019 – 2026. Apache, Apache APISIX, APISIX, and associated open source project names are trademarks of the [Apache Software Foundation](https://www.apache.org/) [](https://www.linkedin.com/company/api7-ai/)[](https://github.com/api7)[](https://twitter.com/api7_ai) --- # acl The `acl` plugin allows or denies request access to upstream resources by verifying whether the user initiating the request is in the access control lists. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can use the `acl` plugin for different scenarios. ### Control Access by Examining Consumer Labels[​](#control-access-by-examining-consumer-labels "Direct link to Control Access by Examining Consumer Labels") The following example demonstrates how to control consumer access based on consumer labels, upon a successful authentication. * Admin API * ADC * Ingress Controller Create two consumers, `john` and `jane`, each with their own labels for organizations and projects: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john", "labels": { "org": "[\"opensource\",\"apache\"]", "project": "[\"tomcat\",\"web-server\",\"http,server\"]" } }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jane", "labels": { "org": "apache", "project": "gateway,apisix,web-server" } }' ``` Create `key-auth` credentials for `john` and `jane`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jane/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-key-auth", "plugins": { "key-auth": { "key": "jane-key" } } }' ``` tip Consumer labels can be configured with either of the two approaches: 1. comma-separated string value, such as `{"project": "gateway,apisix"}` 2. character escaped string array,such as `{"project": "[\"gateway\",\"apisix\"]"}` Create a route with `key-auth` enabled, and configure the `acl` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "acl-route", "uri": "/get", "plugins": { "key-auth": {}, "acl": { "allow_labels": { "org": ["opensource"] } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` ❶ Allow only consumers with `org` label value `opensource` to access the upstream resource. Create two consumers, `john` and `jane`, each with their own labels, credentials, and a route with `key-auth` and `acl` plugins configured: adc.yaml ``` consumers: - username: john labels: org: "[\"opensource\",\"apache\"]" project: "[\"tomcat\",\"web-server\",\"http,server\"]" credentials: - name: cred-john-key-auth type: key-auth config: key: john-key - username: jane labels: org: "apache" project: "gateway,apisix,web-server" credentials: - name: cred-jane-key-auth type: key-auth config: key: jane-key services: - name: acl-service routes: - name: acl-route uris: - /get plugins: key-auth: {} acl: allow_labels: org: - opensource upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ Consumer labels can be configured with either of the two approaches: comma-separated string value, such as `"apache"`, or character escaped string array, such as `"[\"opensource\",\"apache\"]"`. ❷ Allow only consumers with `org` label value `opensource` to access the upstream resource. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create two labeled consumers and attach the `acl` plugin through a `PluginConfig` referenced by `HTTPRoute`: acl-gateway-api.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john labels: org: opensource spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jane labels: org: apache spec: ingressClassName: apisix authParameter: keyAuth: value: key: jane-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: acl-plugin-config spec: plugins: - name: key-auth config: {} - name: acl config: allow_labels: org: - opensource --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: acl-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: acl-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f acl-gateway-api.yaml ``` Create two labeled consumers and apply the `acl` plugin through an `ApisixRoute`: acl-apisix-crd.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john labels: org: opensource spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jane labels: org: apache spec: ingressClassName: apisix authParameter: keyAuth: value: key: jane-key --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: acl-route spec: ingressClassName: apisix http: - name: acl match: paths: - /get backends: - serviceName: httpbin servicePort: 80 plugins: - name: key-auth config: {} - name: acl config: allow_labels: org: - opensource ``` Apply the configuration to your cluster: ``` kubectl apply -f acl-apisix-crd.yaml ``` Send a request to the route as consumer `jane`: ``` curl -i "http://127.0.0.1:9080/get" -H 'apikey: jane-key' ``` You should see an `HTTP/1.1 403 Forbidden` response, as consumer `jane` was not configured with the required label to access the route. Send a request to the route as consumer `john`: ``` curl -i "http://127.0.0.1:9080/get" -H 'apikey: john-key' ``` You should see an `HTTP/1.1 200 OK` response, as consumer `john` was configured with the required label to access the route. ### Control Access by Examining User Information from External Identity Provider[​](#control-access-by-examining-user-information-from-external-identity-provider "Direct link to Control Access by Examining User Information from External Identity Provider") The following example demonstrates how to control user access based on user labels, upon a successful authentication with an external identity provider. Specifically, the example uses Keycloak and user groups as labels. Follow the steps in [set up SSO with Keycloak how-to guide](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-keycloak.md) to create a realm, a client, and a user. Go to **Groups** and create two new groups, `apisix` and `opensource`: ![Keycloak User Groups page showing apisix and open-source groups in the tree, with the New button highlighted](https://static.api7.ai/uploads/2024/02/05/n4M3Id6L_new-group.png) To add the user to group memberships, click into the user and go to the **Groups** tab. Select each group in turn and click **join**: ![Keycloak admin UI: adding a user to one or more groups](https://static.api7.ai/uploads/2024/02/05/vaBSVk5U_add-membership.png) To include the group membership when user info is requested from Keycloak, go to the client and go to the **Mappers** tab. Create a new mapper: ![Keycloak Clients Mappers tab for apisix-quickstart-client with no mappers configured and the Create button highlighted](https://static.api7.ai/uploads/2024/02/05/GTGLRwc9_tHUb4QIw_create-mapper.png) Fill in the name for the protocol mapper, select **Group Membership** as the mapper type, use `groups` as the token claim name, and click **Save**: ![Keycloak Create Protocol Mapper form with Name set to apisix-acl, Mapper Type set to Group Membership, and Token Claim Name set to groups](https://static.api7.ai/uploads/2024/02/05/A5ABZSmE_protocol-mapper.png) To verify if the attribute will be visible when requesting user info, first obtain an access token from Keycloak: ``` OIDC_USER=quickstart-user OIDC_PASSWORD=quickstart-user-pass OIDC_CLIENT_ID=apisix-quickstart-client OIDC_CLIENT_SECRET=bi9NFscFT4k0ljaRzQWlJWthrlygUn3x # replace with your client secret curl "http://$KEYCLOAK_IP:8080/realms/quickstart-realm/protocol/openid-connect/token" -X POST \ -d 'grant_type=password' \ -d 'client_id='$OIDC_CLIENT_ID'' \ -d 'client_secret='$OIDC_CLIENT_SECRET'' \ -d 'username='$OIDC_USER'' \ -d 'password='$OIDC_PASSWORD'' ``` Save the access token to an environment variable called `ACCESS_TOKEN` and send a request to the Keycloak user info endpoint with the token: ``` curl "http://$KEYCLOAK_IP:8080/realms/quickstart-realm/protocol/openid-connect/userinfo" -H "Authorization: Bearer $ACCESS_TOKEN" ``` You should see a response similar to the following: ``` { "sub":"4310e97c-d4c3-479b-bbbd-8c66120e6cee", "email_verified":false, "groups":["/apisix", "/opensource"], "preferred_username":"quickstart-user" } ``` Suppose you would like to only allow users with `/apisix` value in the `groups` attribute to access upstream resources. Create a route with [`openid-connect`](https://docs.api7.ai/hub/openid-connect.md) and `acl` plugins as such: * Admin API * ADC * Ingress Controller ``` KEYCLOAK_IP=192.168.1.81 # replace with your host IP OIDC_DISCOVERY=http://${KEYCLOAK_IP}:8080/realms/quickstart-realm/.well-known/openid-configuration curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- <= 23.0.0`. Next, request user access token and user information from `/userinfo` endpoint, similar to the [last example](#control-access-by-examining-user-information-from-external-identity-provider). You should see Keycloak returning user information similar to the following: ``` { "sub": "f62086ef-29e1-4401-8609-451a2d724bd7", "email_verified": false, "acl_labels": { "nested": { "groups": [ "/apisix", "/opensource" ] } }, "preferred_username": "quickstart-user" } ``` In API7, create a route with [`openid-connect`](https://docs.api7.ai/hub/openid-connect.md) to authenticate with Keycloak and configure `acl` plugins as such: * Admin API * ADC * Ingress Controller ``` KEYCLOAK_IP=192.168.1.81 # replace with your host IP OIDC_DISCOVERY=http://${KEYCLOAK_IP}:8080/realms/quickstart-realm/.well-known/openid-configuration curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < ## Demo[​](#demo "Direct link to Demo") The following demo demonstrates the [moderate request content toxicity example](#moderate-request-content-toxicity) in API7 Enterprise using the Dashboard, where you can moderate request content for toxicity and customize the rejection code and message. ## Behavior by Request Format[​](#behavior-by-request-format "Direct link to Behavior by Request Format") The plugin moderates Chat Completions, Responses API, Embeddings, Anthropic Messages, and Bedrock Converse requests using each protocol's native content structure. The gateway identifies each request by checking URI-specific rules before body-only rules: * Bedrock Converse requires a URI ending in `/converse` and a `messages` array. * Anthropic Messages requires a URI ending in `/v1/messages`. * Responses API requires a URI ending in `/v1/responses` and an `input` field. * Chat Completions uses a `messages` array. * Embeddings uses `input` after the earlier rules do not match. * Other non-empty JSON objects use passthrough after none of the earlier rules match. | Request format | Content available for moderation | | ------------------------ | --------------------------------------------- | | Bedrock Converse | Text from `system` and `messages`. | | Anthropic Messages | Text from `messages`. | | Responses API | Text from `instructions` and `input`. | | Chat Completions | Text from `messages`. | | Embeddings | A string or an array of strings in `input`. | | Other JSON (passthrough) | No request-format-specific text is extracted. | APISIX moderates all extracted content shown in the table. It moderates the latest user turn by default. Use `request_check_roles` to select user, tool, or system content and `request_check_mode` to select the latest or all matching turns. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. With role-aware selection, the Anthropic top-level `system` prompt is available when the system role is selected. Tool-result moderation applies when the request format represents tool output as a distinct tool role or item. Anthropic Messages and Bedrock Converse nest tool results in user messages, so they are not extracted when only the tool role is selected. To moderate the broadest supported request content, configure all available roles and all turns in the plugin configuration: ``` { "request_check_roles": ["user", "tool", "system"], "request_check_mode": "all" } ``` This configuration scans all text that the detected protocol exposes for those roles. It does not reproduce raw-body moderation: nested Anthropic and Bedrock tool results remain part of user content rather than distinct `tool` messages, and unsupported non-AI structures follow `fail_mode`. If Responses content is rejected, the plugin returns the configured message in Responses API format. Streaming requests receive typed server-sent events ending with `response.completed`. Embeddings has no conversation roles or turns. With role-aware selection, its `input` is selected by the user role. Rejected requests receive an OpenAI-style error response. If an Aliyun moderation request fails, the plugin logs the error and allows the content without a moderation verdict. The `fail_mode` setting governs unsupported or non-AI request formats; it does not make Aliyun service failures fail closed. Monitor moderation errors and Aliyun availability when this plugin is an enforcement control. ## Examples[​](#examples "Direct link to Examples") The following examples will be using OpenAI as the upstream service provider. Before proceeding, create an [OpenAI account](https://openai.com) and obtain an [API key](https://openai.com/blog/openai-api). If you are working with other LLM providers, please refer to the provider's documentation to obtain an API key. Additionally, create an [Aliyun account](https://www.aliyun.com), enable Machine-Assisted Moderation Plus, and obtain the endpoint, region ID, access key ID, and access key secret. You can optionally save these information to environment variables: ``` # replace with your data export OPENAI_API_KEY=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 export ALIYUN_ENDPOINT=https://green-cip.cn-shanghai.aliyuncs.com export ALIYUN_REGION_ID=cn-shanghai export ALIYUN_ACCESS_KEY_ID=LTAI5yXKZP77gR3BQQM9WJnA export ALIYUN_ACCESS_KEY_SECRET=hT2YpkqLs9FIjh3dyznBw7RMux5OKv ``` ### Moderate Request Content Toxicity[​](#moderate-request-content-toxicity "Direct link to Moderate Request Content Toxicity") The following example demonstrates how you can use the plugin to moderate content toxicity in requests and customize rejection code and message. * Admin API * ADC * Ingress Controller Create a route to the LLM chat completion endpoint using the [`ai-proxy`](https://docs.api7.ai/hub/ai-proxy.md) plugin and configure the integration details as well as the deny code and message in the `ai-aliyun-content-moderation` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- <
## Behavior by Request Format[​](#behavior-by-request-format "Direct link to Behavior by Request Format") The plugin identifies the request format by checking URI-specific rules before body-only rules: * Bedrock Converse requires a URI ending in `/converse` and a `messages` array. * Anthropic Messages requires a URI ending in `/v1/messages`. * Responses API requires a URI ending in `/v1/responses` and an `input` field. * Chat Completions uses a `messages` array. * Embeddings uses `input` after the earlier rules do not match. * Other non-empty JSON objects use passthrough after none of the earlier rules match. It can extract the following content for moderation: | Request format | Text moderated | | ------------------------ | ------------------------------------------------------- | | Bedrock Converse | Text from `system` and `messages`. | | Anthropic Messages | Text from the top-level `system` prompt and `messages`. | | Responses API | Text from `input` and `instructions`. | | Chat Completions | Text from all entries in `messages`. | | Embeddings | A string or an array of strings in `input`. | | Other JSON (passthrough) | No request-format-specific text is extracted. | Introduced in API7 Enterprise 3.9.16 and 3.10.3. By default, the plugin moderates every extracted `user`, `assistant`, `system`, and `tool` message. Use `request_check_roles` to select roles and `request_check_mode` to moderate all selected turn messages or only the latest consecutive block. A selected `system` role also covers OpenAI `developer` messages and is always checked; it is not limited by `request_check_mode`. Tool-result moderation applies to OpenAI-compatible formats where the tool output is a distinct `tool` role or item. Anthropic Messages and Bedrock Converse nest tool results inside user messages, so selecting only the `tool` role does not extract those nested results. When request content exceeds a threshold, the plugin returns a denial in the detected AI protocol using `deny_code` and `deny_message`. The default status is `200` so AI SDKs can parse the provider-compatible refusal; set a `4xx` value when clients should treat moderation as an HTTP error. Streaming Chat Completions, Responses API, and Anthropic Messages use protocol-specific SSE denial events. Bedrock ConverseStream uses the non-streaming Converse denial body rather than AWS event-stream framing. Set `check_response` to `true` to moderate LLM responses. A non-streaming response is buffered and fails closed with HTTP 500 if Comprehend cannot score it. In `final_packet` streaming mode, earlier chunks have already reached the client, so the plugin annotates the final data event with `risk_level` instead of retracting content. In `realtime` mode, a flagged batch replaces the remainder of the stream with the denial message. A Comprehend failure after streaming begins is logged, and the remaining stream passes without moderation. The plugin sets `$llm_content_risk_level` to `high` when content exceeds a threshold and to `none` after a clean score. Request-side Comprehend failures return HTTP 500 rather than forwarding the request without moderation. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. ## Examples[​](#examples "Direct link to Examples") The following examples will be using OpenAI as the upstream service provider. Before proceeding, create an [OpenAI account](https://openai.com) and obtain an [API key](https://openai.com/blog/openai-api). If you are working with other LLM providers, please refer to the provider's documentation to obtain an API key. Additionally, create [AWS IAM user access keys](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_credentials_access-keys.html) for APISIX to access [AWS Comprehend](https://aws.amazon.com/comprehend/). You can optionally save these keys to environment variables: ``` # replace with your keys export OPENAI_API_KEY=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 export AWS_ACCESS_KEY=AKIARK7HKSJVSHWLD6OS export AWS_SECRET_ACCESS_KEY=4ehUfCPoQmC+AKpG5/5ZaHlzFxFziZ88AylyPerj ``` ### Moderate Profanity[​](#moderate-profanity "Direct link to Moderate Profanity") The following example demonstrates how you can use the plugin to moderate the level of profanity in prompts. * Admin API * ADC * Ingress Controller Create a route to the LLM chat completion endpoint using the [`ai-proxy`](https://docs.api7.ai/hub/ai-proxy.md) plugin and configure the allowed profanity level in `ai-aws-content-moderation`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- <", "object": "chat.completion", "model": "gpt-4", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "request body exceeds PROFANITY threshold" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0 } } ``` Send another request to the route with a typical question in the request body: ``` curl -i "http://127.0.0.1:9080/post" -X POST \ -H "Content-Type: application/json" \ -d '{ "messages": [ { "role": "system", "content": "You are a mathematician" }, { "role": "user", "content": "What is 1+1?" } ] }' ``` You should receive an `HTTP/1.1 200 OK` response with the model output: ``` { ..., "model": "gpt-4-0613", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "1+1 equals 2.", "refusal": null }, "logprobs": null, "finish_reason": "stop" } ], ... } ``` ### Moderate Overall Toxicity[​](#moderate-overall-toxicity "Direct link to Moderate Overall Toxicity") The following example demonstrates how you can use the plugin to moderate the overall toxicity level in prompts, in addition to moderating individual categories. * Admin API * ADC * Ingress Controller Create a route to the LLM chat completion endpoint using the [`ai-proxy`](https://docs.api7.ai/hub/ai-proxy.md) plugin and configure the allowed profanity and overall toxicity levels in `ai-aws-content-moderation`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- <", "object": "chat.completion", "model": "gpt-4", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "request body exceeds toxicity threshold" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0 } } ``` Send another request to the route without any profane word in the request body: ``` curl -i "http://127.0.0.1:9080/post" -X POST \ -H "Content-Type: application/json" \ -d '{ "messages": [ { "role": "system", "content": "You are a mathematician" }, { "role": "user", "content": "What is 1+1?" } ] }' ``` You should receive an `HTTP/1.1 200 OK` response with the model output: ``` { ..., "model": "gpt-4-0613", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "1+1 equals 2.", "refusal": null }, "logprobs": null, "finish_reason": "stop" } ], ... } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. This plugin supports referencing sensitive parameter values from environment variables using the `env://` prefix, or from a secret manager, such as HashiCorp Vault’s [KV secrets engine](https://developer.hashicorp.com/vault/docs/secrets/kv), using the `secret://` prefix. For more information, see [environment variables in plugin](https://docs.api7.ai/apisix/reference/environment-variables.md#plugins) and [secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md). * comprehend object required *** [AWS Comprehend](https://aws.amazon.com/comprehend) configurations. * access\_key\_id string required *** AWS access key ID. * secret\_access\_key string required *** AWS secret access key. The value is encrypted with AES before being stored in etcd. * region string required *** AWS region. * endpoint string *** AWS Comprehend service endpoint. If not set, defaults to `https://comprehend.{region}.amazonaws.com`. * ssl\_verify boolean default: `true` *** If true, enable TLS certificate verification. * moderation\_categories object *** Key-value pairs of moderation category and their corresponding threshold. In each pair, the key should be one of the `PROFANITY`, `HATE_SPEECH`, `INSULT`, `HARASSMENT_OR_ABUSE`, `SEXUAL`, or `VIOLENCE_OR_THREAT`; and the threshold value should be between 0 and 1 (inclusive). * moderation\_threshold number default: `0.5` vaild vaule: between 0 and 1 inclusive *** Overall toxicity threshold. A higher value means more toxic content allowed. This option differs from the individual category thresholds in `moderation_categories`. For example, if `moderation_categories` is set with a `PROFANITY` threshold of `0.5`, and a request has a `PROFANITY` score of `0.1`, the request will not exceed the category threshold. However, if the request has other categories like `SEXUAL` or `VIOLENCE_OR_THREAT` exceeding the `moderation_threshold`, the request will be rejected. * check\_request boolean default: `true` *** If true, moderate request content. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. * deny\_code integer default: `200` vaild vaule: between 200 and 599 inclusive *** HTTP status code returned when flagged traffic is denied before response headers are sent. The default `200` returns a provider-compatible refusal; set a `4xx` value to expose moderation as an HTTP error. After streaming starts, the status cannot be changed. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. * deny\_message string *** Message returned when request or response content is denied. If unset, the plugin returns the threshold failure reason. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. * fail\_mode string default: `skip` vaild vaule: `skip`, `warn`, or `error` *** Behavior when the plugin receives a request it cannot moderate, such as non-AI traffic on a Consumer binding or a request that did not pass through AI Proxy. With `skip`, the request passes unchecked. With `warn`, it passes unchecked and a warning is logged. With `error`, the plugin rejects it with the applicable HTTP 400 or 500 response. None of these outcomes means moderation succeeded. Introduced in API7 Enterprise 3.9.14 and APISIX 3.18.0. * request\_check\_roles array\[string] default: `["user", "tool", "system", "assistant"]` vaild vaule: `user`, `assistant`, `system`, or `tool` *** Message roles to moderate on the request side. `user`, `tool`, and `assistant` follow `request_check_mode`; `system` is checked on every request because system content can be affected by malicious tool-call arguments. `assistant` messages are client-supplied conversation history, so they are moderated by default as well. Selecting `system` also covers `developer` messages, which is the role OpenAI uses in place of `system` on newer models and on the Responses API. There is no separate `developer` entry. Tool-result moderation applies to OpenAI-compatible formats where tool output is represented as a distinct `tool` role or item. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. * request\_check\_mode string default: `all` vaild vaule: `all` or `last` *** Which messages of the selected roles are moderated. With `all`, every message of a selected role is checked. With `last`, only the latest consecutive block of selected-role messages is checked. The `system` role is unaffected and is always checked when selected. Selecting `assistant` together with `last` widens the block that is considered latest, because assistant turns no longer end it. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. * request\_check\_length\_limit integer default: `1000` vaild vaule: between 4 and 1024 inclusive *** Maximum number of bytes of request content per Amazon Comprehend text segment. Longer content is split across several segments so that it is moderated in full instead of being truncated. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. * check\_response boolean default: `false` *** If true, moderate the content of the LLM response in addition to the request. A non-streaming response is moderated before it is returned and fails closed with HTTP 500 if Comprehend cannot score it. A streaming response is moderated according to `stream_check_mode`; after bytes are sent, provider failures are logged and the remaining stream passes without a verdict. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. * response\_check\_length\_limit integer default: `1000` vaild vaule: between 4 and 1024 inclusive *** Maximum number of bytes of response content per Amazon Comprehend text segment. Longer content is split across several segments. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. * stream\_check\_mode string default: `final_packet` vaild vaule: `final_packet` or `realtime` *** How a streaming response is moderated when `check_response` is enabled. With `final_packet`, the assembled response is moderated once and the last chunk is annotated with its risk level. With `realtime`, batches are moderated while the response streams, and the remainder of the stream is replaced with the denial message as soon as a batch is flagged. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. * stream\_check\_cache\_size integer default: `128` vaild vaule: greater than or equal to 1 *** Maximum number of characters accumulated per moderation batch in `realtime` mode. A smaller value detects harmful content earlier at the cost of more moderation calls. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. * stream\_check\_interval number default: `3` vaild vaule: greater than or equal to 0.1 *** Number of seconds between batch checks in `realtime` mode. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. * timeout integer default: `10000` vaild vaule: greater than or equal to 1 *** Timeout in milliseconds for a request to Amazon Comprehend. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. * keepalive boolean default: `true` *** If true, keep the connection to Amazon Comprehend alive so that it is reused across the moderation calls of a request instead of being reopened for each of them. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. * keepalive\_timeout integer default: `60000` vaild vaule: greater than or equal to 1000 *** Idle time in milliseconds after which a pooled connection to Amazon Comprehend is closed. Introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. --- # ai-cache The `ai-cache` plugin caches responses from LLM services so that repeated requests are served from the cache instead of calling the upstream model again. This reduces response latency and upstream token usage for repeated prompts. The plugin supports exact-match caching, where a response is reused only when the normalized request is identical to a previously cached one. It can also use a semantic cache layer that compares prompt embeddings through RediSearch after an exact miss. ## Behavior by Request Format[​](#behavior-by-request-format "Direct link to Behavior by Request Format") The plugin keeps each detected request format in separate cache entries. The gateway identifies each request by checking URI-specific rules before body-only rules: * Bedrock Converse requires a URI ending in `/converse` and a `messages` array. * Anthropic Messages requires a URI ending in `/v1/messages`. * Responses API requires a URI ending in `/v1/responses` and an `input` field. * Chat Completions uses a `messages` array. * Embeddings uses `input` after the earlier rules do not match. * Other non-empty JSON objects use passthrough after none of the earlier rules match. | Request format | Exact-match cache | Semantic cache | | ------------------------ | ----------------- | -------------- | | Bedrock Converse | Supported | Bypassed | | Anthropic Messages | Supported | Bypassed | | Responses API | Supported | Bypassed | | Chat Completions | Supported | Supported | | Embeddings | Supported | Bypassed | | Other JSON (passthrough) | Supported | Bypassed | Exact-match caching was introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.18.0. Semantic and streaming response caching were introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. ## How It Works[​](#how-it-works "Direct link to How It Works") The `ai-cache` plugin must be used together with the [`ai-proxy`](https://docs.api7.ai/hub/ai-proxy.md) or [`ai-proxy-multi`](https://docs.api7.ai/hub/ai-proxy-multi.md) plugin on the same route, because it caches the LLM traffic those plugins proxy. On each request, the plugin computes a cache key from the detected request format, the request body, and the selected AI instance's configuration. From API7 Enterprise version 3.9.20, the key under the `passthrough` protocol also includes the client's request method, path, and query string. That protocol proxies all three verbatim, so they select the upstream endpoint. The key is scoped as configured by `cache_key`. Exact cache entries are stored in Redis with a configurable time-to-live. For Chat Completions requests, semantic caching runs after an exact cache miss. The plugin embeds the configured prompt window and queries a RediSearch vector index for a sufficiently similar cached response. The plugin sets the `X-AI-Cache-Status` response header to one of the following: * `HIT` - a valid cached response was found and is returned directly, without calling the upstream. The `X-AI-Cache-Age` header reports the age of the cached entry in seconds. Semantic hits also return `X-AI-Cache-Similarity`. * `MISS` - no cached response was found. The request is proxied to the upstream, and a successful (`HTTP 200`) response within `max_cache_body_size` is cached for future requests. * `BYPASS` - caching is skipped for this request, for example because it matches a `bypass_on` rule, no AI instance was selected, or the response cannot be safely captured. Complete SSE streaming responses can be cached and replayed with their streaming content type. A stream is cached only after the plugin receives the client protocol's terminal event, such as `[DONE]` for OpenAI Chat Completions, `message_stop` for Anthropic Messages, or `response.completed` for the OpenAI Responses API. Interrupted or limit-truncated streams are not cached. Streaming and non-streaming requests use separate entries, and a cached stream is replayed immediately rather than with its original token timing. Streams that use another framing format, such as Bedrock ConverseStream's AWS event-stream format, bypass the cache. caution Cached prompts and responses can contain sensitive data. Restrict access to Redis, choose an appropriate cache TTL, and use `cache_key.include_consumer` or `cache_key.include_vars` when cached responses should not be shared across consumers or request contexts. ## Examples[​](#examples "Direct link to Examples") The following example uses OpenAI as the upstream LLM service and a Redis instance to store the cache. Before proceeding, create an [OpenAI account](https://openai.com) and an [API key](https://openai.com/blog/openai-api), and make sure a Redis instance is reachable from the gateway. You can optionally save the key to an environment variable: ``` export OPENAI_API_KEY=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 # replace with your API key ``` If you are working with other LLM providers, please refer to the provider's documentation to obtain an API key. ### Cache LLM Responses[​](#cache-llm-responses "Direct link to Cache LLM Responses") The following example demonstrates how to configure `ai-cache` together with `ai-proxy` so that repeated, identical requests are served from Redis. * Admin API * ADC * Ingress Controller Create a route that proxies to OpenAI with `ai-proxy` and caches responses with `ai-cache`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < ## Demo[​](#demo "Direct link to Demo") The following demo demonstrates the [implement allow and deny patterns example](#implement-allow-and-deny-patterns) in API7 Enterprise using the Dashboard, where you can validate user prompts by defining both allow and deny patterns and understand how the allow pattern takes precedence. ## Behavior by Request Format[​](#behavior-by-request-format "Direct link to Behavior by Request Format") The plugin checks Chat Completions, Responses API, Embeddings, Anthropic Messages, and Bedrock Converse requests using each protocol's native content structure. The gateway identifies each request by checking URI-specific rules before body-only rules: * Bedrock Converse requires a URI ending in `/converse` and a `messages` array. * Anthropic Messages requires a URI ending in `/v1/messages`. * Responses API requires a URI ending in `/v1/responses` and an `input` field. * Chat Completions uses a `messages` array. * Embeddings uses `input` after the earlier rules do not match. * Other non-empty JSON objects use passthrough after none of the earlier rules match. By default, the plugin checks the latest user content. Use `match_all_roles` to include other roles and `match_all_conversation_history` to include earlier messages. | Request format | Content checked | | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Bedrock Converse | Content in `system` and `messages`, subject to the role and conversation-history settings. | | Anthropic Messages | Content in the top-level `system` prompt and `messages`, subject to the role and conversation-history settings. | | Responses API | User content in `input`. With `match_all_roles` enabled, it also checks `instructions` and content assigned other roles in `input`. The conversation-history setting does not apply because `instructions` and `input` are parallel fields. | | Chat Completions | Content in `messages`, subject to the role and conversation-history settings. | | Embeddings | A string in `input`, treated as user content. An array of input strings is not inspected. | | Other JSON (passthrough) | Not inspected as a supported AI request format. | ## Examples[​](#examples "Direct link to Examples") The following examples will be using OpenAI as the upstream service provider. Before proceeding, create an [OpenAI account](https://openai.com) and an [API key](https://openai.com/blog/openai-api). You can optionally save the key to an environment variable as such: ``` export OPENAI_API_KEY=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 # replace with your API key ``` If you are working with other LLM providers, please refer to the provider's documentation to obtain an API key. ### Implement Allow and Deny Patterns[​](#implement-allow-and-deny-patterns "Direct link to Implement Allow and Deny Patterns") The following example demonstrates how to use the `ai-prompt-guard` plugin to validate user prompts by defining both allow and deny patterns and understand how the allow pattern takes precedence. Define the allow and deny patterns. You can optionally save them to environment variables for easier escape: ``` # allow US dollar amount export ALLOW_PATTERN_1='\\$?\\(?\\d{1,3}(,\\d{3})*(\\.\\d{1,2})?\\)?' # deny phone number in US number format export DENY_PATTERN_1='(\\([0-9]{3}\\)|[0-9]{3}-)[0-9]{3}-[0-9]{4}' ``` * Admin API * ADC * Ingress Controller Create a route that uses `ai-proxy` to proxy to OpenAI and `ai-prompt-guard` to inspect input prompts: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- </` format. Create a route with the `ai-proxy` plugin configured as such: adc.yaml ``` services: - name: vertex-ai-service routes: - name: vertex-ai-route uris: - /anything methods: - POST plugins: ai-proxy: provider: vertex-ai auth: gcp: service_account_json: "${GCP_SA_JSON}" provider_conf: project_id: api7-vertex region: us-central1 options: model: google/gemini-2.5-flash ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` ❶ Specify the provider to be `vertex-ai`. ❷ Replace with your JSON credentials. Ensure that it is a JSON-escaped string. ❸ Replace with your Vertex AI project ID and region. ❹ Specify the name of the Gemini model through Vertex AI in the `/` format. * Gateway API * APISIX CRD vertex-ai-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: ai-proxy-plugin-config spec: plugins: - name: ai-proxy config: provider: vertex-ai auth: gcp: service_account_json: '{"type":"service_account","project_id":"api7-vertex","private_key_id":"...","private_key":"-----BEGIN PRIVATE KEY-----\n...\n-----END PRIVATE KEY-----","client_email":"api7-docs@api7-vertex.iam.gserviceaccount.com","client_id":"...","auth_uri":"https://accounts.google.com/o/oauth2/auth","token_uri":"https://oauth2.googleapis.com/token","auth_provider_x509_cert_url":"https://www.googleapis.com/oauth2/v1/certs","client_x509_cert_url":"https://www.googleapis.com/robot/v1/metadata/x509/api7-docs%40api7-vertex.iam.gserviceaccount.com","universe_domain":"googleapis.com"}' provider_conf: project_id: api7-vertex region: us-central1 options: model: google/gemini-2.5-flash --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: vertex-ai-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: ai-proxy-plugin-config ``` Apply the configuration to your cluster: ``` kubectl apply -f vertex-ai-ic.yaml ``` ❶ Specify the provider to be `vertex-ai`. ❷ Replace with your JSON credentials. Ensure that it is a JSON-escaped string. ❸ Replace with your Vertex AI project ID and region. ❹ Specify the name of the Gemini model through Vertex AI in the `/` format. vertex-ai-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: vertex-ai-route spec: ingressClassName: apisix http: - name: vertex-ai-route match: paths: - /anything methods: - POST plugins: - name: ai-proxy enable: true config: provider: vertex-ai auth: gcp: service_account_json: '{"type":"service_account","project_id":"api7-vertex","private_key_id":"...","private_key":"-----BEGIN PRIVATE KEY-----\n...\n-----END PRIVATE KEY-----","client_email":"api7-docs@api7-vertex.iam.gserviceaccount.com","client_id":"...","auth_uri":"https://accounts.google.com/o/oauth2/auth","token_uri":"https://oauth2.googleapis.com/token","auth_provider_x509_cert_url":"https://www.googleapis.com/oauth2/v1/certs","client_x509_cert_url":"https://www.googleapis.com/robot/v1/metadata/x509/api7-docs%40api7-vertex.iam.gserviceaccount.com","universe_domain":"googleapis.com"}' provider_conf: project_id: api7-vertex region: us-central1 options: model: google/gemini-2.5-flash ``` Apply the configuration to your cluster: ``` kubectl apply -f vertex-ai-ic.yaml ``` ❶ Specify the provider to be `vertex-ai`. ❷ Replace with your JSON credentials. Ensure that it is a JSON-escaped string. ❸ Replace with your Vertex AI project ID and region. ❹ Specify the name of the Gemini model through Vertex AI in the `/` format. Send a POST request to the route with a system prompt and a sample user question in the request body: ``` curl "http://127.0.0.1:9080/anything" -X POST \ -H "Content-Type: application/json" \ -d '{ "messages": [ { "role": "system", "content": "You are a mathematician" }, { "role": "user", "content": "What is 1+1?" } ] }' ``` You should receive a response similar to the following: ``` { "choices": [ { "message": { "role": "assistant", "content": "1 + 1 = 2\n" }, "index": 0, "logprobs": null, "finish_reason": "stop" } ], "usage": { "completion_tokens": 8, "extra_properties": { "google": { "traffic_type": "ON_DEMAND" } }, "total_tokens": 19, "prompt_tokens": 11 }, "object": "chat.completion", "model": "google/gemini-2.5-flash", ... } ``` ### Proxy to Vertex AI Embedding Models[​](#proxy-to-vertex-ai-embedding-models "Direct link to Proxy to Vertex AI Embedding Models") The following example demonstrates how you can configure the `ai-proxy` plugin to proxy requests to Vertex AI embedding models using GCP service account authentication. This example applies to API7 Enterprise from version 3.9.2 and APISIX from version 3.17.0. Before proceeding: * [Enable Vertex AI](https://docs.cloud.google.com/vertex-ai/docs/featurestore/setup) and billing for your GCP project. * Follow the [service account credentials](https://developers.google.com/workspace/guides/create-credentials#service-account) section to create a service account in GCP, assign the account with the "Vertex AI User" role, and obtain the account credentials in JSON. Your credentials file should look similar to the following: credentials.json ``` { "type": "service_account", "project_id": "api7-vertex", "private_key_id": "...", "private_key": "-----BEGIN PRIVATE KEY-----\n...\n-----END PRIVATE KEY-----\n", "client_email": "api7-docs@api7-vertex.iam.gserviceaccount.com", "client_id": "....", "auth_uri": "https://accounts.google.com/o/oauth2/auth", "token_uri": "https://oauth2.googleapis.com/token", "auth_provider_x509_cert_url": "https://www.googleapis.com/oauth2/v1/certs", "client_x509_cert_url": "https://www.googleapis.com/robot/v1/metadata/x509/api7-docs%40api7-vertex.iam.gserviceaccount.com", "universe_domain": "googleapis.com" } ``` Optionally save the JSON to an environment variable: ``` export GCP_SA_JSON="$(cat credentials.json)" ``` * Admin API * ADC * Ingress Controller Create a route and configure the `ai-proxy` plugin as such: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < export AWS_SECRET_ACCESS_KEY= ``` Create a route with the `ai-proxy` plugin configured for Bedrock: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < -n -f values.yaml ``` Now if you create a route following the [Proxy to OpenAI example](#proxy-to-openai). Send a request like this: ``` curl "http://127.0.0.1:9080/anything" -X POST \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-3.5", "messages": [ { "role": "system", "content": "You are a mathematician" }, { "role": "user", "content": "What is 1+1?" } ] }' ``` Since the model in `ai-proxy` is `gpt-4`, the request will be forwarded to GPT-4 model and you will receive a response similar to the following: ``` { ..., "model": "gpt-4-0613", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "1+1 equals 2.", "refusal": null, "annotations": [] }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 23, "completion_tokens": 8, "total_tokens": 31, "prompt_tokens_details": { "cached_tokens": 0, "audio_tokens": 0 }, ... }, "service_tier": "default", "system_fingerprint": null } ``` In the gateway's access log, you should see a log entry similar to the following: ``` 192.168.215.1 - - [29/Aug/2025:09:54:16 +0000] 127.0.0.1:9080 "POST /anything HTTP/1.1" 200 808 2.670 "-" "curl/8.6.0" - - 2670 "http://127.0.0.1:9080" "6526bf5c961b6e6bb8cfcb66486f02dc" "ai_chat" "2670" "gpt-4" "gpt-3.5" "23" "8" "31" "false" "false" "0" "" "0" "0" "0" ``` The access log entry shows an upstream response time and time to first token of `2670` milliseconds. The request uses the `ai_chat` type, requests `gpt-3.5`, and is forwarded to `gpt-4`. It uses 23 prompt tokens, 8 completion tokens, and 31 total tokens. The remaining values show a non-streaming request with no tool calls, provided tools, end-user identifier, prompt-cache tokens, or reasoning tokens. ### Send Request Log to Logger[​](#send-request-log-to-logger "Direct link to Send Request Log to Logger") The following example demonstrates how you can log request and request information, including LLM model, token, and payload, and push them to a logger. Before proceeding, you should first set up a logger, such as Kafka. See [`kafka-logger`](https://docs.api7.ai/hub/kafka-logger.md) for more information. * Admin API * ADC * Ingress Controller Create a route to your LLM service and configure logging details as such: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < ## Demo[​](#demo "Direct link to Demo") The following demo demonstrates the [configure instance priority and rate limiting example](#configure-instance-priority-and-rate-limiting). It shows how you can configure two models with different priorities and apply rate limiting on the instance with a higher priority in API7 Enterprise using the Dashboard. In the case where `fallback_strategy` is set to `["rate_limiting"]`, the plugin should continue to forward requests to the low priority instance once the high priority instance's rate limiting quota is fully consumed. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `ai-proxy-multi` for different scenarios. ### Load Balance between Instances[​](#load-balance-between-instances "Direct link to Load Balance between Instances") The following example configures two models for load balancing, forwarding 80% of the traffic to one instance and 20% to the other. It also retries one additional instance when the first attempt returns `429` or `5xx` within two seconds. `max_retries` and `retry_on_failure_within_ms` are available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. For demonstration and easier differentiation, you will be configuring one OpenAI instance and one DeepSeek instance as the upstream LLM services. Create a route as such and update with your LLM providers, models, API keys, and endpoints if applicable: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- <" options: model: gpt-4.1 - name: openai-responses-secondary provider: openai weight: 1 auth: header: Authorization: "Bearer " options: model: gpt-4.1-mini --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: ai-proxy-multi-responses-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /v1/responses method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: ai-proxy-multi-responses-plugin-config ``` Apply the configuration to your cluster: ``` kubectl apply -f ai-proxy-multi-ic.yaml ``` ai-proxy-multi-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: ai-proxy-multi-responses-route spec: ingressClassName: apisix http: - name: ai-proxy-multi-responses-route match: paths: - /v1/responses methods: - POST plugins: - name: ai-proxy-multi enable: true config: instances: - name: openai-responses-primary provider: openai weight: 1 auth: header: Authorization: "Bearer " options: model: gpt-4.1 - name: openai-responses-secondary provider: openai weight: 1 auth: header: Authorization: "Bearer " options: model: gpt-4.1-mini ``` Apply the configuration to your cluster: ``` kubectl apply -f ai-proxy-multi-ic.yaml ``` Send a request using the OpenAI Responses API format: ``` curl "http://127.0.0.1:9080/v1/responses" -X POST \ -H "Content-Type: application/json" \ -d '{ "input": "Write one sentence about API gateways." }' ``` The request is forwarded to one of the configured OpenAI instances and the response is returned in the Responses API format. ### Route by Semantic Similarity[​](#route-by-semantic-similarity "Direct link to Route by Semantic Similarity") The `semantic` balancer selects an instance by comparing the request prompt with example utterances assigned to each instance. It is available in API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0. Export API keys for the embedding request and LLM requests: ``` export EMBEDDING_API_KEY="" export LLM_API_KEY="" ``` Create a route with instances for programming, translation, and general prompts. The general instance is also the fallback when no score reaches the configured threshold or the embedding request fails: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- <` format. ❹ Configure the provider to be `vertex-ai` for Vertex AI Gemini access. ❺ Replace with your JSON credentials. Ensure that it is a JSON-escaped string. ❻ Replace with your Vertex AI project ID and region. ❼ Specify the name of the Gemini model through Vertex AI in the `/` format. adc.yaml ``` services: - name: ai-proxy-multi-service routes: - name: ai-proxy-multi-route uris: - /anything methods: - POST plugins: ai-proxy-multi: fallback_strategy: - rate_limiting instances: - name: gemini-instance provider: gemini weight: 7 auth: header: Authorization: "Bearer ${GEMINI_API_KEY}" options: model: gemini-2.5-flash - name: vertex-ai-instance provider: vertex-ai weight: 3 auth: gcp: service_account_json: "${GCP_SA_JSON}" provider_conf: project_id: api7-vertex region: us-central1 options: model: google/gemini-2.5-flash ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` ❶ Configure the provider to be `gemini` for Google AI Studio Gemini access. ❷ Replace with your Gemini API key in the `Authorization` header. ❸ Specify the name of the Gemini model through Google AI Studio in the `` format. ❹ Configure the provider to be `vertex-ai` for Vertex AI Gemini access. ❺ Replace with your JSON credentials. Ensure that it is a JSON-escaped string. ❻ Replace with your Vertex AI project ID and region. ❼ Specify the name of the Gemini model through Vertex AI in the `/` format. * Gateway API * APISIX CRD ai-proxy-multi-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: ai-proxy-multi-plugin-config spec: plugins: - name: ai-proxy-multi config: fallback_strategy: - rate_limiting instances: - name: gemini-instance provider: gemini weight: 7 auth: header: Authorization: "Bearer AIzaSyDUMZbZmHCmJ5BNNLl0KfQk" options: model: gemini-2.5-flash - name: vertex-ai-instance provider: vertex-ai weight: 3 auth: gcp: service_account_json: '{"type":"service_account","project_id":"api7-vertex","private_key_id":"...","private_key":"-----BEGIN PRIVATE KEY-----\n...\n-----END PRIVATE KEY-----","client_email":"api7-docs@api7-vertex.iam.gserviceaccount.com","client_id":"...","auth_uri":"https://accounts.google.com/o/oauth2/auth","token_uri":"https://oauth2.googleapis.com/token","auth_provider_x509_cert_url":"https://www.googleapis.com/oauth2/v1/certs","client_x509_cert_url":"https://www.googleapis.com/robot/v1/metadata/x509/api7-docs%40api7-vertex.iam.gserviceaccount.com","universe_domain":"googleapis.com"}' provider_conf: project_id: api7-vertex region: us-central1 options: model: google/gemini-2.5-flash --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: ai-proxy-multi-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: ai-proxy-multi-plugin-config ``` Apply the configuration to your cluster: ``` kubectl apply -f ai-proxy-multi-ic.yaml ``` ❶ Configure the provider to be `gemini` for Google AI Studio Gemini access. ❷ Replace with your Gemini API key in the `Authorization` header. ❸ Specify the name of the Gemini model through Google AI Studio in the `` format. ❹ Configure the provider to be `vertex-ai` for Vertex AI Gemini access. ❺ Replace with your JSON credentials. Ensure that it is a JSON-escaped string. ❻ Replace with your Vertex AI project ID and region. ❼ Specify the name of the Gemini model through Vertex AI in the `/` format. ai-proxy-multi-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: ai-proxy-multi-route spec: ingressClassName: apisix http: - name: ai-proxy-multi-route match: paths: - /anything methods: - POST plugins: - name: ai-proxy-multi enable: true config: fallback_strategy: - rate_limiting instances: - name: gemini-instance provider: gemini weight: 7 auth: header: Authorization: "Bearer AIzaSyDUMZbZmHCmJ5BNNLl0KfQk" options: model: gemini-2.5-flash - name: vertex-ai-instance provider: vertex-ai weight: 3 auth: gcp: service_account_json: '{"type":"service_account","project_id":"api7-vertex","private_key_id":"...","private_key":"-----BEGIN PRIVATE KEY-----\n...\n-----END PRIVATE KEY-----","client_email":"api7-docs@api7-vertex.iam.gserviceaccount.com","client_id":"...","auth_uri":"https://accounts.google.com/o/oauth2/auth","token_uri":"https://oauth2.googleapis.com/token","auth_provider_x509_cert_url":"https://www.googleapis.com/oauth2/v1/certs","client_x509_cert_url":"https://www.googleapis.com/robot/v1/metadata/x509/api7-docs%40api7-vertex.iam.gserviceaccount.com","universe_domain":"googleapis.com"}' provider_conf: project_id: api7-vertex region: us-central1 options: model: google/gemini-2.5-flash ``` Apply the configuration to your cluster: ``` kubectl apply -f ai-proxy-multi-ic.yaml ``` ❶ Configure the provider to be `gemini` for Google AI Studio Gemini access. ❷ Replace with your Gemini API key in the `Authorization` header. ❸ Specify the name of the Gemini model through Google AI Studio in the `` format. ❹ Configure the provider to be `vertex-ai` for Vertex AI Gemini access. ❺ Replace with your JSON credentials. Ensure that it is a JSON-escaped string. ❻ Replace with your Vertex AI project ID and region. ❼ Specify the name of the Gemini model through Vertex AI in the `/` format. Send 10 POST requests to the route to see the load balancing distribution: ``` studio_count=0 vertex_count=0 for i in {1..10}; do model=$(curl -s "http://127.0.0.1:9080/anything" -X POST \ -H "Content-Type: application/json" \ -d '{ "messages": [ { "role": "system", "content": "You are a mathematician" }, { "role": "user", "content": "What is 1+1?" } ] }' | jq -r '.model') if [[ "$model" == "gemini-2.5-flash" ]]; then ((studio_count++)) elif [[ "$model" == "google/gemini-2.5-flash" ]]; then ((vertex_count++)) fi done echo "Google AI Studio Gemini responses: $studio_count" echo "Vertex AI Gemini responses: $vertex_count" ``` You should see a response similar to the following: ``` Google AI Studio Gemini responses: 7 Vertex AI Gemini responses: 3 ``` ### Configure Instance Priority and Rate Limiting[​](#configure-instance-priority-and-rate-limiting "Direct link to Configure Instance Priority and Rate Limiting") The following example demonstrates how you can configure two models with different priorities and apply rate limiting on the instance with a higher priority. In the case where `fallback_strategy` is set to `["rate_limiting"]`, the plugin should continue to forward requests to the low priority instance once the high priority instance's rate limiting quota is fully consumed. Create a route as such and update with your LLM providers, models, API keys, and endpoints if applicable: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < -n -f values.yaml ``` Next, create a route with the `ai-proxy-multi` plugin following the previous examples and send a request. For instance, if you send a request like this: ``` curl "http://127.0.0.1:9080/anything" -X POST \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-3.5", "messages": [ { "role": "system", "content": "You are a mathematician" }, { "role": "user", "content": "What is 1+1?" } ] }' ``` If the LLM instance model in `ai-proxy-multi` is `gpt-4`, then the request will be forwarded to GPT-4 model and you will receive a response similar to the following: ``` { ..., "model": "gpt-4-0613", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "1+1 equals 2.", "refusal": null, "annotations": [] }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 23, "completion_tokens": 8, "total_tokens": 31, "prompt_tokens_details": { "cached_tokens": 0, "audio_tokens": 0 }, ... }, "service_tier": "default", "system_fingerprint": null } ``` In the gateway's access log, you should see a log entry similar to the following: ``` 192.168.215.1 - - [29/Aug/2025:09:54:16 +0000] 127.0.0.1:9080 "POST /anything HTTP/1.1" 200 808 2.670 "-" "curl/8.6.0" - - 2670 "http://127.0.0.1:9080" "6526bf5c961b6e6bb8cfcb66486f02dc" "ai_chat" "2670" "gpt-4" "gpt-3.5" "23" "8" "31" "false" "false" "0" "" "0" "0" "0" ``` The access log entry shows an upstream response time and time to first token of `2670` milliseconds. The request uses the `ai_chat` type, requests `gpt-3.5`, and is forwarded to `gpt-4`. It uses 23 prompt tokens, 8 completion tokens, and 31 total tokens. The remaining values show a non-streaming request with no tool calls, provided tools, end-user identifier, prompt-cache tokens, or reasoning tokens. ### Send Request Log to Logger[​](#send-request-log-to-logger "Direct link to Send Request Log to Logger") The following example demonstrates how you can log request and request information, including LLM model, token, and payload, and push them to a logger. Before proceeding, you should first set up a logger, such as Kafka. See [`kafka-logger`](https://docs.api7.ai/hub/kafka-logger.md) for more information. Create a route to your LLM services and configure logging details as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < -n --all -o yaml > values.yaml ``` Add or update the following value: values.yaml ``` apisix: pluginAttrs: ai-proxy: http_client: lua-resty-http ``` Then apply the values file with the chart used for this APISIX release: ``` helm upgrade apisix/apisix -n -f values.yaml ``` ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * fallback\_strategy string or array vaild vaule: string: `instance_health_and_rate_limiting`, `http_429`, or `http_5xx`
array: Any combination of `rate_limiting`, `http_429`, and `http_5xx` *** Fallback strategy. The option `instance_health_and_rate_limiting` is kept for backward compatibility and is functionally the same as `rate_limiting`. With `rate_limiting` or `instance_health_and_rate_limiting`, when the current instance's quota is exhausted, the request is forwarded to the next instance regardless of priority. With `http_429`, if an instance returns status code 429, the request is retried with other instances. With `http_5xx`, if an instance returns a 5xx status code, the request is retried with other instances. If all instances fail, the plugin returns the last upstream status, response body, and `Content-Type`. When not set, the plugin will not forward the request to low priority instances when tokens of the high priority instance are exhausted. * max\_retries integer vaild vaule: greater than or equal to 0 *** Maximum number of fallback retries after the initial request fails. This bounds how many additional instances a single request tries, so it does not exhaust every configured instance. Only takes effect together with `fallback_strategy`. When not set, there is no explicit cap and the plugin retries until an instance succeeds or all instances have been tried. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * retry\_on\_failure\_within\_ms integer vaild vaule: greater than or equal to 1 *** Only fall back to another instance when the upstream fails within this many milliseconds. Fast failures (such as connection errors and quick 429 or 5xx responses) are retried, while a slow failure that takes longer than this is returned to the client directly to avoid doubling the total wait time. Only takes effect together with `fallback_strategy`. When not set, the plugin retries regardless of how long the failed attempt took. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * fallback\_http\_statuses array\[integer] vaild vaule: each between 400 and 599, no duplicates *** Additional upstream HTTP status codes that make the request fall back to another instance, on top of the `http_429` and `http_5xx` entries of `fallback_strategy`. Use it for statuses that mean the instance's own credential or quota is the problem rather than the request, such as `401` for an expired key or `403` for a disabled account, so the request is retried elsewhere instead of being returned to the client. Only takes effect together with `fallback_strategy`. Available in API7 Enterprise from version 3.9.19 on the 3.9 line and from version 3.10.6 on the 3.10 line. * balancer object *** Load balancing configurations. * algorithm string default: `roundrobin` vaild vaule: `roundrobin`, `chash`, or `semantic` *** Load balancing algorithm. When set to `roundrobin`, weighted round robin algorithm is used. When set to `chash`, consistent hashing algorithm is used. When set to `semantic`, the instance whose `examples` are semantically closest to the prompt is used, configured under `semantic_opts`. The `semantic` algorithm does not participate in health checks, `fallback_strategy`, or `max_retries`. An upstream failure on the selected instance is returned to the client; the algorithm falls back only when no instance clears its threshold or the embedding request fails. The `semantic` algorithm is available in API7 Enterprise from version 3.9.18 on the 3.9 line and from version 3.10.5 on the 3.10 line, and in APISIX from version 3.18.0. * hash\_on string default: `vars` vaild vaule: `vars`, `header`, `cookie`, `consumer`, or `vars_combinations` *** Used when `type` is `chash`. Support hashing on [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md), header, cookie, consumer, or a combination of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). * key string *** Used when `type` is `chash`. When `hash_on` is set to `header` or `cookie`, `key` is required. When `hash_on` is set to `consumer`, `key` is not required as the consumer name will be used as the key automatically. * semantic\_opts object *** Configurations for the `semantic` balancer algorithm. Required when `balancer.algorithm` is `semantic`, and ignored otherwise. Available in API7 Enterprise from version 3.9.18 on the 3.9 line and from version 3.10.5 on the 3.10 line, and in APISIX from version 3.18.0. * embeddings object required *** Embedding service used to turn the prompt and each instance's `examples` into vectors. The prompt is embedded on every request, so this service is on the request path. * provider string required vaild vaule: `openai` or `azure-openai` *** Embedding service provider. * model string required *** Name of the embedding model, such as `text-embedding-3-small`. * endpoint string *** Embedding API endpoint. Optional for `openai`, which defaults to the public API. Required for `azure-openai`, where it has to be the full URL, such as `https://{resource}.openai.azure.com/openai/deployments/{deployment}/embeddings?api-version={version}`. * auth object required *** Authentication for the embedding service, carried either as headers or as query parameters. * header object *** Key-value pairs sent as request headers to the embedding service. * query object *** Key-value pairs sent as query parameters to the embedding service. * timeout integer default: `3000` vaild vaule: greater than or equal to 1 *** Timeout in milliseconds for an embedding request. Because the prompt is embedded synchronously, this bounds the latency added to each request when the embedding service is slow. On a timeout the request is routed to the fallback instance rather than failed. * ssl\_verify boolean default: `true` *** If true, verify the embedding service's TLS certificate. * threshold number default: `0` vaild vaule: between -1 and 1 inclusive *** Global minimum cosine similarity an instance has to reach to be selected. An instance's own `threshold` overrides this value. The default of `0` admits almost any prompt, so the fallback instance is only ever reached once a threshold above `0` is set. * fallback string *** Name of the instance to route to when no instance clears its threshold or the embedding request fails. It is otherwise a normally ranked instance and needs its own `examples`. Defaults to the first instance when unset. * debugging boolean default: `false` *** If true, return the per-instance similarity scores and the routing decision in the `X-AI-Semantic-Scores` and `X-AI-Semantic-Picked-Instance` response headers. Intended for tuning `examples` and thresholds, not for production traffic. * instances array\[object] required *** LLM instance configurations. * name string required *** Name of the LLM service instance. * examples array\[string] vaild vaule: between 1 and 64 items *** Example utterances representing what this instance handles. Each one is embedded into its own reference vector, and the semantic balancer routes a request to the instance whose closest example is most similar to the prompt. Required for every instance when `balancer.algorithm` is `semantic`, including the instance named by `semantic_opts.fallback`. Ignored by the other algorithms. Available in API7 Enterprise from version 3.9.18 on the 3.9 line and from version 3.10.5 on the 3.10 line, and in APISIX from version 3.18.0. * threshold number vaild vaule: between -1 and 1 inclusive *** Minimum cosine similarity a prompt has to reach for this instance to be selected by the semantic balancer. Overrides `semantic_opts.threshold` for this instance. Available in API7 Enterprise from version 3.9.18 on the 3.9 line and from version 3.10.5 on the 3.10 line, and in APISIX from version 3.18.0. * provider string required vaild vaule: `openai`, `deepseek`, `azure-openai`, `aimlapi`, `gemini`, `vertex-ai`, `anthropic`, `openrouter`, `bedrock`, `openai-compatible` *** LLM service provider. When set to `openai`, the plugin sends detected Chat Completions, Responses API, and Embeddings requests to their corresponding OpenAI endpoints. When set to `deepseek`, the plugin will proxy requests to `https://api.deepseek.com/chat/completions`. When set to `gemini` (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests to `https://generativelanguage.googleapis.com/v1beta/openai/chat/completions`. If you are proxying requests to an embedding model, you should configure the embedding model endpoint in the `override`. When set to `vertex-ai` (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin proxies requests to Google Cloud Vertex AI. For chat completions, the plugin will proxy requests to `https://{region}-aiplatform.googleapis.com/v1beta1/projects/{project_id}/locations/{region}/endpoints/openapi/chat/completions`. For embeddings, the plugin will proxy requests to `https://{region}-aiplatform.googleapis.com/v1/projects/{project_id}/locations/{region}/publishers/google/models/{model}:predict`. These require configuring `provider_conf` with `project_id` and `region`. Alternatively, you can configure `override` for a custom endpoint. When set to `anthropic` (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin sends detected Chat Completions requests to `https://api.anthropic.com/v1/chat/completions` and native Anthropic Messages requests to `https://api.anthropic.com/v1/messages`. When set to `openrouter` (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests to `https://openrouter.ai/api/v1/chat/completions`. When set to `bedrock` (available from API7 Enterprise 3.9.12 and APISIX 3.17.0), the plugin proxies requests to AWS Bedrock using the Converse API. When set to `aimlapi` (available from APISIX 3.14.0 and Enterprise 3.8.17), the plugin uses the OpenAI-compatible driver and proxies the request to `https://api.aimlapi.com/v1/chat/completions`. When set to `openai-compatible`, the plugin proxies requests to the custom endpoint configured in `override`. When set to `azure-openai`, the plugin also proxies requests to the custom endpoint configured in `override` and additionally removes the `model` parameter from user requests. * priority integer default: `0` *** Priority of the LLM instance in load balancing. `priority` takes precedence over `weight`. * weight integer required vaild vaule: greater than or equal to 0 *** Weight of the LLM instance in load balancing. * auth object required *** Authentication configurations. * header object *** Authentication headers. At least one of the `header` and `query` should be configured. You can configure additional custom headers that will be forwarded to the upstream LLM service. * query object *** Authentication query parameters. At least one of the `header` and `query` should be configured. * gcp object *** GCP service account authentication for Vertex AI. Available in API7 Enterprise from 3.9.2 and APISIX from version 3.17.0. * service\_account\_json string *** GCP service account JSON content used for authentication. This can be configured using this parameter or by setting the `GCP_SERVICE_ACCOUNT` environment variable. * max\_ttl integer *** Maximum TTL for GCP access token caching, in seconds. * expire\_early\_secs integer default: `60` *** Number of seconds to expire the access token before its actual expiration time. This prevents edge cases where tokens expire during active requests. * aws object *** AWS IAM credentials for SigV4 signing. Required when `provider` is `bedrock` (for Bedrock, `auth.aws` is sufficient and `auth.header`/`auth.query` are not required). Available in API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0. * access\_key\_id string required *** AWS IAM access key ID. * secret\_access\_key string required *** AWS IAM secret access key. * session\_token string *** AWS session token for temporary credentials (e.g. from STS AssumeRole). * options object *** Model configurations. In addition to `model`, you can configure additional parameters and they will be forwarded to the upstream LLM service in the request body. For instance, if you are working with OpenAI or DeepSeek, you can configure additional parameters such as `max_tokens`, `temperature`, `top_p`, and `stream`. See your LLM provider's API documentation for more available options. * model string *** Name of the LLM model, such as `gpt-4` or `gpt-3.5`. See your LLM provider's API documentation for more available models. * provider\_conf object *** Provider-specific configuration. Required when `provider` is `bedrock`. When `provider` is `vertex-ai`, configure either `provider_conf` or `override.endpoint`. Available in API7 Enterprise from 3.9.2 and APISIX from version 3.17.0. * project\_id string *** Google Cloud Project ID for Vertex AI. * region string required *** Cloud region. For `vertex-ai`, this is the GCP region. For `bedrock`, this is the AWS region (e.g. `us-east-1`). * override object *** Override setting. * endpoint string *** LLM provider endpoint to replace the endpoint selected for the detected request protocol. * llm\_options object *** Provider-aware LLM option overrides. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. * max\_tokens integer *** Maximum number of output tokens. The gateway automatically maps this to the correct field name for the target provider, such as `max_completion_tokens` for OpenAI Chat or `max_output_tokens` for OpenAI Responses API, and overwrites the client value. * request\_body object *** Per target-protocol request body overrides. Keys are target protocol names, such as `openai-chat`, `openai-responses`, `openai-embeddings`, `anthropic-messages`, `bedrock-converse`, and `passthrough`. Values are partial request bodies that are deep-merged into the outgoing body. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. * request\_body\_force\_override boolean default: `false` *** When `false` (default), client request body fields take priority and `request_body` override values only fill in missing fields. When `true`, `request_body` override values overwrite client fields. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. * checks object *** Health check configurations. Note that at the moment, OpenAI and DeepSeek do not provide an official health check endpoint. Other LLM services that you can configure under `openai-compatible` provider may have available health check endpoints. * active object required *** Active health check configurations. * type string default: `http` vaild vaule: `http`, `https`, or `tcp` *** Type of health check connection. * timeout number default: `1` *** Health check timeout in seconds. * concurrency integer default: `10` *** Number of upstream nodes to be checked at the same time. * host string *** HTTP host. * port integer vaild vaule: between 1 and 65535 inclusive *** HTTP port. * http\_path string default: `/` *** Path for HTTP probing requests. * http\_method string default: `GET` vaild vaule: `CONNECT`, `DELETE`, `GET`, `HEAD`, `OPTIONS`, `PATCH`, `POST`, `PURGE`, `PUT`, or `TRACE` *** HTTP method for active health check probing requests. Available in API7 Enterprise and in APISIX from version 3.18.0. * http\_req\_body string *** Request body to send in active health check probing requests. This is useful when `http_method` is set to `POST`. Defaults to empty string. Available in API7 Enterprise and in APISIX from version 3.18.0. * https\_verify\_certificate boolean default: `true` *** If true, verify the node's TLS certificate. * healthy object *** Healthy check configurations. * interval integer default: `1` *** Time interval of checking healthy nodes, in seconds. * http\_statuses array\[integer] default: `[200,302]` vaild vaule: status code between 200 and 599 inclusive *** An array of HTTP status codes that defines a healthy node. * successes integer default: `2` vaild vaule: between 1 and 254 inclusive *** Number of successful probes to define a healthy node. * req\_headers array\[string] *** List of additional HTTP headers to send in health check probing requests, in `"Header: Value"` format. * unhealthy object *** Unhealthy check configurations. * interval integer default: `1` *** Time interval of checking unhealthy nodes, in seconds. * http\_statuses array\[integer] default: `[429,404,500,501,502,503,504,505]` vaild vaule: status code between 200 and 599 inclusive *** An array of HTTP status codes that defines an unhealthy node. * http\_failures integer default: `5` vaild vaule: between 1 and 254 inclusive *** Number of HTTP failures to define an unhealthy node. * tcp\_failures integer default: `2` vaild vaule: between 1 and 254 inclusive *** Number of TCP failures to define an unhealthy node. * timeouts integer default: `3` vaild vaule: between 1 and 254 inclusive *** Number of probe timeouts to define an unhealthy node. * logging object *** Logging configurations. These configurations apply to access logs and logs sent to logging plugins, and do not affect the error log. * summaries boolean default: `false` *** If true, add an `llm_summary` object to logger entries with model, latency, and token usage. In API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0, the summary also includes stream status, tool count and usage, end-user ID, cache read and creation tokens, reasoning tokens, and content risk level when available. * payloads boolean default: `false` *** If true, log request and response payload. * timeout integer default: `30000` vaild vaule: between 1 and 600000 inclusive *** Timeout in milliseconds for each connect, send, or blocking read operation to the LLM service. It does not limit the total duration of a streaming response; use `max_stream_duration_ms` for that limit. * max\_req\_body\_size integer default: `67108864` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes that the plugin reads into memory. Larger requests are rejected with HTTP 413. This prevents unbounded memory buffering of large request bodies. The default is 67108864 bytes (64 MiB). Available in API7 Enterprise from versions 3.9.14 and 3.10.1 in their respective release lines, and APISIX from version 3.17.0. * max\_stream\_duration\_ms integer vaild vaule: greater than or equal to 1 *** Maximum wall-clock duration, in milliseconds, for a streaming AI response. The limit is optional. When reached, the gateway closes the connection; if output has already started, the stream ends without a protocol terminator such as `[DONE]`, `message_stop`, or `response.completed`. Enforcement occurs between upstream reads, so the final chunk can exceed the configured duration. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. * max\_response\_bytes integer vaild vaule: greater than or equal to 1 *** Maximum total bytes read from the upstream for one streaming or non-streaming AI response. The limit is optional and checked between upstream reads, so the final chunk can exceed it. If the limit is exceeded before output starts, the gateway returns `502 Bad Gateway`; after output starts, the gateway closes the stream without a protocol terminator. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. * streaming\_flush\_interval\_ms integer default: `10` vaild vaule: greater than or equal to 0 *** Background flush interval in milliseconds for streaming responses. A positive value periodically flushes buffered output to bound client latency when the upstream sends tokens in bursts. Set to 0 to flush each chunk synchronously. Available in API7 Enterprise from version 3.9.13 and APISIX from version 3.17.0. * keepalive boolean default: `true` *** If true, keep the connection alive when requesting the LLM service. * keepalive\_timeout integer default: `60000` vaild vaule: greater than or equal to 1000 *** Keepalive timeout in milliseconds when requesting the LLM service. * keepalive\_pool integer default: `30` vaild vaule: greater than or equal to 1 *** Keepalive pool size for when connecting with the LLM service. * ssl\_verify boolean default: `true` *** If true, verify the LLM service's certificate. --- # Protocol Reference The `ai-proxy` and `ai-proxy-multi` plugins share the same request protocol detection and conversion pipeline. The pipeline identifies the client format before routing to the configured provider or selected instance. For plugin-specific configuration, see [`ai-proxy`](https://docs.api7.ai/hub/ai-proxy.md) and [`ai-proxy-multi`](https://docs.api7.ai/hub/ai-proxy-multi.md). ## Request Protocol Detection[​](#request-protocol-detection "Direct link to Request Protocol Detection") The plugins identify the client protocol before matching it to a protocol supported by the selected provider or instance. The following detection rules apply to both plugins. ### Request Requirements[​](#request-requirements "Direct link to Request Requirements") Requests that include `Content-Type` must use `application/json`. If the header is omitted, the plugins treat the body as JSON. They reject unsupported content types and invalid bodies before selecting a provider or instance. The request body cannot exceed `max_req_body_size`, which defaults to 67,108,864 bytes. A request that exceeds this limit receives an HTTP 413 response. In API7 Gateway, this setting is available from 3.9.14 in the 3.9.x line and from 3.10.1 in the 3.10.x line. It is available in APISIX 3.17.0 and later. ### Detection Order[​](#detection-order "Direct link to Detection Order") The plugins check the following rules in order: | Client protocol | Body signal | Path requirement | | ----------------------- | --------------------------------------------------------------- | -------------------------------------------------------------- | | Bedrock Converse | The request body contains a `messages` array. | The path ends in `/converse`; custom prefixes are allowed. | | Anthropic Messages | The request body is a JSON object. | The path ends in `/v1/messages`; custom prefixes are allowed. | | OpenAI Responses | The request body contains `input`. | The path ends in `/v1/responses`; custom prefixes are allowed. | | OpenAI Chat Completions | The request body contains a `messages` array. | Any path matched by the route. | | OpenAI Embeddings | The request body contains `input`, and no earlier rule matched. | Any path matched by the route. | The path-specific rules run before the body-only rules. This prevents Bedrock Converse and Anthropic Messages requests containing `messages` from being identified as Chat Completions. Responses and Embeddings requests both use `input`, so a request containing `input` but not `messages` is identified as Embeddings unless its path ends in `/v1/responses`. Any other non-empty JSON object is treated as passthrough. This mode keeps the original request path and can reuse the original body when no request transformation changes it. Provider authentication and `override.endpoint` still apply. Empty or invalid request bodies are rejected. Passthrough does not provide an AI protocol model to downstream AI-aware plugins. Usage extraction, prompt decoration or templating, content moderation text extraction, and protocol conversion therefore do not run for the passthrough body. ### After Detection[​](#after-detection "Direct link to After Detection") For a named protocol, the plugin uses the detected protocol without conversion when the selected provider supports it. Otherwise, the plugin looks for a registered converter to a protocol the provider supports. The request is rejected if neither native support nor a compatible converter is available. The detected protocol and request body determine whether `request_type` is recorded as `ai_stream` or `ai_chat` for logging plugins. Response parsing uses the upstream response's content type to distinguish streaming from non-streaming responses. ## Request Override Precedence[​](#request-override-precedence "Direct link to Request Override Precedence") The final plugin configuration uses `override.llm_options.max_tokens` for provider-aware token-limit mapping. The provider maps that value to the field expected by the target protocol, then applies the matching `override.request_body` object. Request-body objects are merged recursively. Arrays and scalar values replace the existing value rather than being combined. With `request_body_force_override: false`, client fields win and the override fills only missing values; with `true`, override values replace matching client fields. The request-body key names the target protocol after any conversion, such as `openai-chat`, `anthropic-messages`, or `bedrock-converse`. ## Failure Responses[​](#failure-responses "Direct link to Failure Responses") | Condition | Client-visible behavior | | -------------------------------------------------------------------------------- | ------------------------------ | | Request body exceeds `max_req_body_size` | `413 Request Entity Too Large` | | LLM connection or read times out before a response is available | `504 Gateway Timeout` | | A streaming converter receives a response it cannot parse in the selected format | `502 Bad Gateway` | | `ai-request-rewrite` receives a request without a body | `400 Bad Request` | If a response limit is reached after streaming output has begun, the gateway closes the downstream stream. It cannot replace bytes already sent with a new HTTP error response. ## Anthropic-to-OpenAI Conversion[​](#anthropic-to-openai-conversion "Direct link to Anthropic-to-OpenAI Conversion") An Anthropic Messages client can send requests through `ai-proxy` or `ai-proxy-multi` to a backend that supports OpenAI Chat Completions. The plugins convert the client request to OpenAI format and convert the backend response to Anthropic format. This conversion supports a subset of the Anthropic Messages API: it preserves some fields, transforms others, and discards unsupported fields. ### When Conversion Applies[​](#when-conversion-applies "Direct link to When Conversion Applies") The plugins identify an Anthropic Messages request when the request path ends in `/v1/messages` and the body is a JSON object. If the selected provider supports Anthropic Messages, the request uses that protocol without conversion. If the provider supports OpenAI Chat Completions instead, the shared converter translates the request and response. The reverse client/backend pairing is not supported: an OpenAI Chat Completions client cannot use this converter to call an Anthropic Messages backend. For `ai-proxy-multi`, the selected instance determines whether conversion is required. The [conversion configuration example](https://docs.api7.ai/hub/ai-proxy.md#convert-anthropic-requests-to-openai-compatible-backend) uses `ai-proxy`; for multi-instance configuration, see [`ai-proxy-multi`](https://docs.api7.ai/hub/ai-proxy-multi.md). ### Request Conversion[​](#request-conversion "Direct link to Request Conversion") The converted request body is built from an allowlist. The following tables summarize the Anthropic inputs the converter reads. Unrecognized fields are discarded before the request reaches the backend. #### Request Fields[​](#request-fields "Direct link to Request Fields") | Anthropic field | OpenAI field | Behavior | | --------------------------------------- | --------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `model` | `model` | Forwarded, unless the route pins a model with `options.model`. | | `max_tokens` | `max_completion_tokens` | Renamed. | | `stop_sequences` | `stop` | Renamed. | | `temperature`, `top_p` | Same names | Forwarded. | | `stream` | `stream`, plus `stream_options.include_usage` | The plugin sets `stream_options.include_usage` so that the stream carries usage. | | `system` | A leading message with `role: system` | Text blocks are concatenated into one string. | | `tools[]` (custom tools) | `tools[].function` | A tool name containing characters outside `[a-zA-Z0-9_-]`, or longer than 64 characters, is rewritten to satisfy the OpenAI naming rules. The original name is restored in the response. | | `tool_choice` | `tool_choice` | Converted. `{"type": "auto"}` becomes `"auto"`, `{"type": "any"}` becomes `"required"`, `{"type": "none"}` becomes `"none"`, and `{"type": "tool", "name": "..."}` becomes an object naming that function. | | `tool_choice.disable_parallel_tool_use` | `parallel_tool_calls: false` | Converted. | | `thinking` | `reasoning_effort` | Approximated. A continuous `budget_tokens` value is mapped to one discrete effort level. The thresholds are release-dependent. | | `output_config.effort` | `reasoning_effort` | Used when `thinking.type` is `adaptive`. Release-dependent; see [Release Compatibility](#release-compatibility). | | `output_format`, `output_config.format` | `response_format` | Release-dependent; see [Release Compatibility](#release-compatibility). | | `metadata.user_id` | `user` | Renamed. | | `service_tier` | `service_tier` | Forwarded. | #### Message Content[​](#message-content "Direct link to Message Content") | Anthropic content | OpenAI equivalent | Behavior | | -------------------------------- | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | `messages[].content` as a string | `messages[].content` as the same string | Forwarded. | | `text` | A text content part, or a plain string | The shape depends on the release; see [Release Compatibility](#release-compatibility). | | `image` with a base64 source | `image_url` with a `data:` URL | Converted. Whether the backend model accepts image input varies by model. | | `image` with a URL source | `image_url` with the same URL | Forwarded unchanged. | | `document` with a base64 source | `image_url` with a `data:` URL | Approximation. The document bytes are placed in a field the OpenAI schema defines for images, so whether a backend accepts them is outside that schema. | | `tool_use` | An assistant message with `tool_calls` | Converted. Tool-name handling in message history is release-dependent; see [Release Compatibility](#release-compatibility). | | `tool_result` | A message with `role: tool` | Converted. Its ordering relative to ordinary text is release-dependent; see [Release Compatibility](#release-compatibility). | #### Dropped Fields[​](#dropped-fields "Direct link to Dropped Fields") The backend does not receive the following fields, and the response carries no signal that they were removed: | Anthropic field | Why it is dropped | | ------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------- | | `top_k` | OpenAI Chat Completions has no equivalent parameter. | | `cache_control` | No equivalent. The converted request carries no caching directive. | | `citations` | The converter does not map citations in either direction. | | `thinking` and `redacted_thinking` blocks in message history | OpenAI Chat Completions has no equivalent field. Ordinary text in the same assistant message is kept. | | Anthropic built-in tools (`computer_`, `bash_`, `text_editor_`, `web_search`, `code_execution_`) | The converter has no mapping for these tools. | #### Request Headers[​](#request-headers "Direct link to Request Headers") If a request has an `x-api-key` header but no `Authorization` header, the converter sends the key as a bearer token in `Authorization`. It removes the original `x-api-key` header and headers whose names start with `anthropic-` or `x-stainless-`. ### Response Conversion[​](#response-conversion "Direct link to Response Conversion") The converter reads completion fields only from `choices[0]`; any additional OpenAI choices are discarded. It maps top-level `usage` and `error` fields separately. | OpenAI response field | Anthropic response | Behavior | | -------------------------------------------------- | ----------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `message.content` | A `text` content block | Converted. | | `message.reasoning_content` or `message.reasoning` | A `thinking` content block | Converted when the backend returns a non-empty string. A non-streaming block has an empty signature; see [Known Limitations](#known-limitations). | | `message.tool_calls` | `tool_use` content blocks | Converted. Sanitized tool names are restored when the original name is available. | | `finish_reason` | `stop_reason` | `stop` and `content_filter` become `end_turn`; `length` becomes `max_tokens`; `tool_calls` and `function_call` become `tool_use`. Any other value defaults to `end_turn`. | | `usage` | `input_tokens`, `output_tokens`, and available cache-token fields | `prompt_tokens` becomes `input_tokens`, and `completion_tokens` becomes `output_tokens`. When cache details are available, cached prompt tokens are removed from `input_tokens` and reported as `cache_read_input_tokens`; `cache_creation_input_tokens` is also included when provided. | | `error` | An Anthropic error object | Converted when a normally parsed upstream response body contains an error object. HTTP 429, 5xx, and transport errors can bypass this conversion. | For streaming responses, the converter emits Anthropic message and content-block events for OpenAI text, reasoning, and tool-call deltas. The initial `message_start` usage values are zero; final token usage is emitted in `message_delta`. Available cache-token fields can be included when the backend supplies them in a supported usage chunk. Clients that report streaming usage should read the final event and should not assume that cache-token fields are present. ### Release Compatibility[​](#release-compatibility "Direct link to Release Compatibility") API7 Gateway 3.9.x and 3.10.x receive fixes independently. Check the column for the release line you run. | Behavior | API7 Gateway 3.9.x | API7 Gateway 3.10.x | APISIX | | --------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ | ------------------- | ---------------- | | An orphaned `tool_choice` is removed when every tool was dropped | 3.9.16 and later | 3.10.2 and later | 3.18.0 and later | | A malformed backend tool call degrades instead of failing the response | 3.9.16 and later | 3.10.2 and later | 3.18.0 and later | | `message_start.content` is serialized as an array | 3.9.16 and later | 3.10.3 and later | 3.18.0 and later | | Anthropic's current structured-output shape is recognized | 3.9.16 and later | 3.10.3 and later | 3.18.0 and later | | `thinking.type: adaptive` uses `output_config.effort` | 3.9.16 and later | 3.10.3 and later | 3.18.0 and later | | `thinking.budget_tokens` uses the four-level mapping described below | 3.9.16 and later | 3.10.3 and later | 3.18.0 and later | | A user message with a single text block is sent as a content array, and an assistant message with several text blocks is concatenated into a string | 3.9.16 and later | 3.10.3 and later | 3.18.0 and later | | Tool names in message history are rewritten consistently with declared tools | 3.9.16 and later | 3.10.3 and later | 3.18.0 and later | | `tool_result` messages are placed before ordinary text from the same user message | 3.9.16 and later | 3.10.3 and later | 3.18.0 and later | | Media is retained when a user message also contains `tool_result` | 3.9.16 and later | 3.10.3 and later | 3.18.0 and later | These version differences affect structured output, message content, thinking effort, and tool history as follows: #### Structured Output[​](#structured-output "Direct link to Structured Output") API7 Gateway 3.9.15 and earlier, API7 Gateway 3.10.0 through 3.10.2, and APISIX 3.17.0 recognize only the converter's legacy expected shape. That shape is `output_config` or `output_format` carrying `type: json_schema` together with a `json_schema` field, or carrying `type: json` or `type: json_object`. Anthropic's current shape carries the schema in `output_format.schema` or in `output_config.format`. In the 3.9.x line, API7 Gateway 3.9.16 and later recognize this shape. In the 3.10.x line, 3.10.3 and later recognize it. APISIX 3.18.0 and later recognize it as well. These releases normalize the schema and send `response_format` with strict mode enabled. On earlier releases, the backend receives no `response_format`, and the client receives no error. #### Message Content Shape[​](#message-content-shape "Direct link to Message Content Shape") On a release that predates the change, a user message carrying a single text block is sent as a plain string. An assistant message carrying several text blocks is sent as a content array. A backend that accepts only one of these shapes behaves differently across an upgrade. #### Thinking Effort[​](#thinking-effort "Direct link to Thinking Effort") The `budget_tokens` mapping changes across releases. The earlier mapping applies to APISIX 3.17.0 and API7 Gateway releases before 3.9.16 or 3.10.3 in their respective lines. The later mapping applies from APISIX 3.18.0 and from API7 Gateway 3.9.16 and 3.10.3 in their respective lines. | `budget_tokens` | Earlier releases | Later releases | | ------------------ | ---------------- | -------------- | | Below 1024 | `low` | `minimal` | | 1024 through 2047 | `low` | `low` | | 2048 through 4095 | `low` | `medium` | | 4096 through 16383 | `medium` | `high` | | 16384 or higher | `high` | `high` | | Not provided | `medium` | `minimal` | #### Tool History[​](#tool-history "Direct link to Tool History") On an earlier release, a `tool_use` name in message history is not rewritten with the corresponding declared tool name. Ordinary text can also be sent before `tool_result` messages from the same user message, and media in that message is discarded. A strict backend may reject the name or ordering mismatch. APISIX 3.18.0 and later, API7 Gateway 3.9.16 and later, and API7 Gateway 3.10.3 and later rewrite history names consistently, place tool messages first, and retain media. ### Known Limitations[​](#known-limitations "Direct link to Known Limitations") The following limitations can affect converted requests and responses across the supported releases. #### Streaming Can End Without a Terminating Event[​](#streaming-can-end-without-a-terminating-event "Direct link to Streaming Can End Without a Terminating Event") When a backend closes the stream without a properly delimited final frame, the plugins do not emit the closing `message_delta` and `message_stop` events. A client that relies on `message_stop` may wait indefinitely or treat the stream as incomplete. This affects every release listed above. Set a client-side timeout and treat an unexpected end of the stream as a failure. #### Converted `thinking` Blocks Do Not Carry a Valid Signature[​](#converted-thinking-blocks-do-not-carry-a-valid-signature "Direct link to converted-thinking-blocks-do-not-carry-a-valid-signature") For a non-streaming response, the plugins set the block's `signature` to an empty string. For a streaming response, they emit thinking deltas without a signature. Clients that require a valid signature cannot replay either converted form as a signed Anthropic thinking block. A backend that instead embeds reasoning in ordinary message content produces a response where the reasoning appears as visible text. #### Error Responses Are Not Consistently Anthropic-Shaped[​](#error-responses-are-not-consistently-anthropic-shaped "Direct link to Error Responses Are Not Consistently Anthropic-Shaped") HTTP 429, 5xx, and transport timeout responses can bypass response conversion. Clients should be prepared to receive an upstream or gateway error body that does not follow the Anthropic error schema. #### Backend Capabilities Are Not Validated[​](#backend-capabilities-are-not-validated "Direct link to Backend Capabilities Are Not Validated") The plugins convert the request but do not check whether the backend model supports the result. A backend can return HTTP 200 while silently ignoring a capability, such as dropping an image, ignoring `response_format`, or returning no tool call. Backend behavior varies by model and between dated snapshots of the same model name. Validate the specific models you plan to use rather than generalizing from a backend. ### Validate Backend Compatibility[​](#validate-backend-compatibility "Direct link to Validate Backend Compatibility") Test the following converted inputs against each backend model because a backend can reject them even when the original Anthropic request is valid: * **A named `tool_choice`.** The converter emits an object naming the function, or `"required"`. Some backends accept only `"auto"` while the model is reasoning. If a backend rejects the converted form, send `{"type": "auto"}` from the client. * **`thinking` together with a small `max_tokens`.** `thinking` becomes `reasoning_effort`, which can make a backend reserve a reasoning budget. When that budget exceeds the converted `max_completion_tokens`, the backend rejects the request. Raise `max_tokens` when you enable `thinking`. * **A `document` block.** The converter can only offer it to the backend as an image. A model that cannot read it may answer with invented content instead of reporting an error. * **Mixed text, media, and `tool_result` content.** Earlier releases can put ordinary text before the converted tool message and discard media from the same user message. Test this shape if the backend validates tool-message ordering, or upgrade to a release that places tool messages first and retains media. ## Related Configuration[​](#related-configuration "Direct link to Related Configuration") * [Configure `ai-proxy` to convert Anthropic requests](https://docs.api7.ai/hub/ai-proxy.md#convert-anthropic-requests-to-openai-compatible-backend). * [Configure native Anthropic Messages pass-through](https://docs.api7.ai/hub/ai-proxy.md#native-anthropic-messages-api-pass-through). * [Configure `ai-proxy-multi`](https://docs.api7.ai/hub/ai-proxy-multi.md). * [Convert Anthropic Messages with API7 Gateway](https://docs.api7.ai/api7-gateway/ai-gateway/use-cases/protocol-conversion.md). --- ## Static Configurations[​](#static-configurations "Direct link to Static Configurations") APISIX 3.18.0 uses `ngx_http_ffi_client` by default for upstream requests from `ai-proxy`, `ai-proxy-multi`, and `ai-request-rewrite`. Set `http_client` to `lua-resty-http` to use the Lua client instead. API7 Gateway 3.9 and 3.10 use the Lua client and do not expose this setting. * Host or Docker * Kubernetes (Helm) To use the Lua client in an APISIX host or Docker deployment, configure the following setting: config.yaml ``` plugin_attr: ai-proxy: http_client: lua-resty-http ``` Then reload APISIX for the change to take effect. Export the full effective values for the installed APISIX release: ``` helm get values -n --all -o yaml > values.yaml ``` Add or update the following value: values.yaml ``` apisix: pluginAttrs: ai-proxy: http_client: lua-resty-http ``` Then apply the values file with the chart used for this APISIX release: ``` helm upgrade apisix/apisix -n -f values.yaml ``` ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * provider string required vaild vaule: `openai`, `deepseek`, `azure-openai`, `aimlapi`, `gemini`, `vertex-ai`, `anthropic`, `openrouter`, `bedrock`, `openai-compatible` *** LLM service provider. When set to `openai`, the plugin sends detected Chat Completions, Responses API, and Embeddings requests to their corresponding OpenAI endpoints. When set to `deepseek`, the plugin will proxy requests to `https://api.deepseek.com/chat/completions`. When set to `gemini` (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests to `https://generativelanguage.googleapis.com/v1beta/openai/chat/completions`. If you are proxying requests to an embedding model, you should configure the embedding model endpoint in the `override`. When set to `vertex-ai` (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin proxies requests to Google Cloud Vertex AI. For chat completions, the plugin will proxy requests to `https://{region}-aiplatform.googleapis.com/v1beta1/projects/{project_id}/locations/{region}/endpoints/openapi/chat/completions`. For embeddings, the plugin will proxy requests to `https://{region}-aiplatform.googleapis.com/v1/projects/{project_id}/locations/{region}/publishers/google/models/{model}:predict`. These require configuring `provider_conf` with `project_id` and `region`. Alternatively, you can configure `override` for a custom endpoint. When set to `anthropic` (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin sends detected Chat Completions requests to `https://api.anthropic.com/v1/chat/completions` and native Anthropic Messages requests to `https://api.anthropic.com/v1/messages`. When set to `openrouter` (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests to `https://openrouter.ai/api/v1/chat/completions`. When set to `bedrock` (available from API7 Enterprise 3.9.12 and APISIX 3.17.0), the plugin proxies requests to AWS Bedrock using the [Converse API](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_Converse.html). Requires configuring `auth.aws` with IAM credentials and `provider_conf.region` with the AWS region. Supports both non-streaming and streaming (ConverseStream) when `stream` is set to `true` in the request body. When set to `aimlapi` (available from APISIX 3.14.0 and Enterprise 3.8.17), the plugin uses the OpenAI-compatible driver and proxies the request to `https://api.aimlapi.com/v1/chat/completions`. When set to `openai-compatible`, the plugin proxies requests to the custom endpoint configured in `override`. When set to `azure-openai`, the plugin also proxies requests to the custom endpoint configured in `override` and additionally removes the `model` parameter from user requests. * auth object required *** Authentication configurations. * header object *** Authentication headers. * query object *** Authentication query parameters. * gcp object *** GCP service account authentication for Vertex AI. Available in API7 Enterprise from 3.9.2 and APISIX from version 3.17.0. * service\_account\_json string *** GCP service account JSON content used for authentication. This can be configured using this parameter or by setting the `GCP_SERVICE_ACCOUNT` environment variable. * max\_ttl integer *** Maximum TTL for GCP access token caching, in seconds. * expire\_early\_secs integer default: `60` *** Number of seconds to expire the access token before its actual expiration time. This prevents edge cases where tokens expire during active requests. * aws object *** AWS IAM credentials for SigV4 signing. Required when `provider` is `bedrock` (for Bedrock, `auth.aws` is sufficient and `auth.header`/`auth.query` are not required). Available in API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0. * access\_key\_id string required *** AWS IAM access key ID. * secret\_access\_key string required *** AWS IAM secret access key. * session\_token string *** AWS session token for temporary credentials (e.g. from STS AssumeRole). * options object *** Model configurations. In addition to `model`, you can configure additional parameters and they will be forwarded to the upstream LLM service in the request body. For instance, if you are working with OpenAI, you can configure additional parameters such as `temperature`, `top_p`, and `stream`. See your LLM provider's API documentation for more available options. * model string *** Name of the LLM model, such as `gpt-4` or `gpt-3.5`. See your LLM provider's API documentation for more available models. * provider\_conf object *** Provider-specific configuration. Required when `provider` is `bedrock`. When `provider` is `vertex-ai`, configure either `provider_conf` or `override.endpoint`. Available in API7 Enterprise from 3.9.2 and APISIX from version 3.17.0. * project\_id string *** Google Cloud Project ID. Required when `provider` is `vertex-ai`. * region string required *** Cloud region. For `vertex-ai`, this is the GCP region. For `bedrock`, this is the AWS region (e.g. `us-east-1`). * override object *** Override setting. * endpoint string *** LLM provider endpoint. Required when `provider` is `openai-compatible`. * llm\_options object *** Provider-aware LLM option overrides. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. * max\_tokens integer *** Maximum number of output tokens. The gateway automatically maps this to the correct field name for the target provider (e.g. `max_completion_tokens` for OpenAI Chat, `max_output_tokens` for OpenAI Responses API). Always force-overwrites the client value. * request\_body object *** Per target-protocol request body overrides. Keys are target protocol names (`openai-chat`, `openai-responses`, `openai-embeddings`, `anthropic-messages`, `bedrock-converse`, `passthrough`); values are partial request bodies that are deep-merged into the outgoing body (objects merged recursively, arrays and scalars replaced wholesale). Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. * request\_body\_force\_override boolean default: `false` *** When `false` (default), client request body fields take priority and `request_body` override values only fill in missing fields. When `true`, `request_body` override values forcefully overwrite client fields. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. * logging object *** Logging configurations. These configurations apply to access logs and logs sent to logging plugins, and do not affect the error log. * summaries boolean default: `false` *** If true, add an `llm_summary` object to logger entries with model, latency, and token usage. In API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0, the summary also includes stream status, tool count and usage, end-user ID, cache read and creation tokens, reasoning tokens, and content risk level when available. * payloads boolean default: `false` *** If true, log request and response payload. * timeout integer default: `30000` vaild vaule: between 1 and 600000 inclusive *** Timeout in milliseconds for each connect, send, or blocking read operation to the LLM service. It does not limit the total duration of a streaming response; use `max_stream_duration_ms` for that limit. * max\_req\_body\_size integer default: `67108864` *** Maximum request body size in bytes that the plugin reads into memory (default 67108864 bytes, which is 64 MiB). Requests with a body larger than this limit are rejected with HTTP 413. This prevents unbounded memory buffering of large request bodies. Available in API7 Enterprise from versions 3.9.14 and 3.10.1 in their respective release lines, and APISIX from version 3.17.0. * keepalive boolean default: `true` *** If true, keep the connection alive when requesting the LLM service. * keepalive\_timeout integer default: `60000` vaild vaule: greater than or equal to 1000 *** Keepalive timeout in milliseconds when requesting the LLM service. * keepalive\_pool integer default: `30` vaild vaule: greater than or equal to 1 *** Keepalive pool size for when connecting with the LLM service. * ssl\_verify boolean default: `true` *** If true, verify the LLM service's certificate. * max\_stream\_duration\_ms integer *** Maximum wall-clock duration, in milliseconds, for a streaming AI response. The limit is optional. When reached, the gateway closes the connection; if output has already started, the stream ends without a protocol terminator such as `[DONE]`, `message_stop`, or `response.completed`. Enforcement occurs between upstream reads, so the final chunk can exceed the configured duration. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. * max\_response\_bytes integer *** Maximum total bytes read from the upstream for one streaming or non-streaming AI response. The limit is optional and checked between upstream reads, so the final chunk can exceed it. If the limit is exceeded before output starts, the gateway returns `502 Bad Gateway`; after output starts, the gateway closes the stream without a protocol terminator. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. * streaming\_flush\_interval\_ms integer default: `10` *** Background flush interval in milliseconds for streaming responses. A positive value starts a background thread that flushes output periodically to bound client latency when upstreams burst multiple tokens at once. Set to 0 to flush each chunk synchronously inline. Available in API7 Enterprise from version 3.9.13 and APISIX from version 3.17.0. --- # ai-rag The `ai-rag` plugin implements the retrieval step of a Retrieval-Augmented Generation (RAG) request flow. It generates an embedding from the request and performs a vector search. It then adds the retrieved content to the protocol-specific LLM input and removes the `ai_rag` object before proxying the request. The current implementation supports [Azure OpenAI](https://azure.microsoft.com/en-us/products/ai-services/openai-service) for embeddings and [Azure AI Search](https://azure.microsoft.com/en-us/products/ai-services/ai-search) for vector search. Use the [`ai-proxy`](https://docs.api7.ai/hub/ai-proxy.md) plugin in the same request flow to proxy the augmented request to the LLM provider. The plugin does not create or populate a search index; prepare the index and its content before sending requests through APISIX.
## Behavior by Request Format[​](#behavior-by-request-format "Direct link to Behavior by Request Format") The plugin enriches Chat Completions, Responses API, Anthropic Messages, and Bedrock Converse requests using each protocol's native prompt structure. The gateway identifies each request by checking URI-specific rules before body-only rules: * Bedrock Converse requires a URI ending in `/converse` and a `messages` array. * Anthropic Messages requires a URI ending in `/v1/messages`. * Responses API requires a URI ending in `/v1/responses` and an `input` field. * Chat Completions uses a `messages` array. * Embeddings uses `input` after the earlier rules do not match. * Other non-empty JSON objects use passthrough after none of the earlier rules match. | Request format | Context enrichment | | ------------------------ | -------------------------------------------------------------------------------------------------------------------------------- | | Bedrock Converse | Appends the retrieved context as a user message in `messages`. | | Anthropic Messages | Appends the retrieved context as a user message in `messages`. | | Responses API | Appends the retrieved context to `input`. | | Chat Completions | Appends the retrieved context as a user message in `messages`. | | Embeddings | Does not enrich the request. The nested `ai_rag.embeddings` object configures the embedding input used internally for retrieval. | | Other JSON (passthrough) | Does not enrich the request. | ## Verify Upstream TLS[​](#verify-upstream-tls "Direct link to Verify Upstream TLS") In API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0, `ssl_verify` defaults to `true` for calls to the embedding and vector-search services. Configure a trusted certificate chain for both endpoints before enabling the plugin in production. Setting `ssl_verify` to `false` preserves connectivity to an endpoint with an untrusted certificate, but should be limited to temporary migration or development use. ## Example[​](#example "Direct link to Example") To follow along the example, create an [Azure account](https://portal.azure.com) and complete the following steps: * In [Azure AI Foundry](https://oai.azure.com/portal), deploy a generative chat model, such as `gpt-4o`, and an embedding model, such as `text-embedding-3-large`. Obtain the API key and model endpoints. * Follow [Azure's example](https://github.com/Azure/azure-search-vector-samples/blob/main/demo-python/code/basic-vector-workflow/azure-search-vector-python-sample.ipynb) to prepare for a vector search in [Azure AI Search](https://azure.microsoft.com/en-us/products/ai-services/ai-search) using Python. The example will create a search index called `vectest` with the desired schema and upload the [sample data](https://github.com/Azure/azure-search-vector-samples/blob/main/data/text-sample.json) which contains 108 descriptions of various Azure services, for embeddings `titleVector` and `contentVector` to be generated based on `title` and `content`. Complete all the setups before performing vector searches in Python. * In [Azure AI Search](https://azure.microsoft.com/en-us/products/ai-services/ai-search), [obtain the Azure vector search API key and the search service endpoint](https://learn.microsoft.com/en-us/azure/search/search-get-started-vector?tabs=api-key#retrieve-resource-information). Save the API keys and endpoints to environment variables: ``` # replace with your values export AZ_OPENAI_DOMAIN=https://your-openai-resource.openai.azure.com export AZ_OPENAI_API_KEY=your-azure-openai-api-key export AZ_CHAT_ENDPOINT=${AZ_OPENAI_DOMAIN}/openai/deployments/gpt-4o/chat/completions?api-version=2024-02-15-preview export AZ_EMBEDDING_MODEL=text-embedding-3-large export AZ_EMBEDDINGS_ENDPOINT=${AZ_OPENAI_DOMAIN}/openai/deployments/${AZ_EMBEDDING_MODEL}/embeddings?api-version=2023-05-15 export AZ_AI_SEARCH_SVC_DOMAIN=https://your-search-service.search.windows.net export AZ_AI_SEARCH_KEY=your-azure-ai-search-api-key export AZ_AI_SEARCH_INDEX=vectest export AZ_AI_SEARCH_ENDPOINT=${AZ_AI_SEARCH_SVC_DOMAIN}/indexes/${AZ_AI_SEARCH_INDEX}/docs/search?api-version=2024-07-01 ``` ### Integrate with Azure for RAG-Enhanced Responses[​](#integrate-with-azure-for-rag-enhanced-responses "Direct link to Integrate with Azure for RAG-Enhanced Responses") The following example demonstrates how you can use the [`ai-proxy`](https://docs.api7.ai/hub/ai-proxy.md) plugin to proxy requests to Azure OpenAI LLM and use the `ai-rag` plugin to generate embeddings and perform vector search to enhance LLM responses. * Admin API * ADC * Ingress Controller Create a route as such: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- <>> Check for open slots... >>> Check slots coverage... [OK] All 16384 slots covered. ``` 4. Verify cluster nodes: ``` docker exec -it redis-node-7000 redis-cli -c -a redis-cluster-password -p 7000 cluster nodes ``` The expected output should be similar to the following: ``` node-id-1 172.XX.0.2:7000@17000 myself,master - 0 0 1 connected 0-5460 node-id-2 172.XX.0.3:7001@17001 master - 0 0 2 connected 5461-10922 node-id-3 172.XX.0.4:7002@17002 master - 0 0 3 connected 10923-16383 node-id-4 172.XX.0.5:7003@17003 slave node-id-1 0 0 1 connected node-id-5 172.XX.0.6:7004@17004 slave node-id-2 0 0 2 connected node-id-6 172.XX.0.7:7005@17005 slave node-id-3 0 0 3 connected ``` 5. Check cluster health (optional): ``` docker exec redis-node-7000 redis-cli -c -a redis-cluster-password -p 7000 cluster info ``` You should see the following response: ``` cluster_state:ok cluster_slots_assigned:16384 cluster_slots_ok:16384 cluster_known_nodes:6 cluster_size:3 ... ``` Create a Kubernetes manifest for the Redis cluster: redis-cluster.yaml ``` apiVersion: apps/v1 kind: StatefulSet metadata: namespace: aic name: redis-cluster spec: serviceName: redis-cluster replicas: 6 selector: matchLabels: app: redis-cluster template: metadata: labels: app: redis-cluster spec: containers: - name: redis image: redis:7.2-alpine ports: - containerPort: 6379 name: client - containerPort: 16379 name: gossip command: - redis-server - --cluster-enabled - "yes" - --cluster-config-file - nodes.conf - --cluster-node-timeout - "5000" - --appendonly - "yes" - --requirepass - redis-cluster-password - --masterauth - redis-cluster-password volumeMounts: - name: data mountPath: /data volumeClaimTemplates: - metadata: name: data spec: accessModes: ["ReadWriteOnce"] resources: requests: storage: 1Gi --- apiVersion: v1 kind: Service metadata: namespace: aic name: redis-cluster spec: clusterIP: None selector: app: redis-cluster ports: - port: 6379 name: client - port: 16379 name: gossip ``` Apply the manifest: ``` kubectl apply -f redis-cluster.yaml ``` Wait for all pods to be ready, then initialize the cluster: ``` kubectl exec -n aic redis-cluster-0 -- redis-cli \ --cluster create \ $(for i in 0 1 2 3 4 5; do \ echo -n "$(kubectl get pod -n aic redis-cluster-$i -o jsonpath='{.status.podIP}'):6379 "; \ done) \ --cluster-replicas 1 \ --cluster-yes \ -a redis-cluster-password ``` #### Create Route and Configure Rate Limiting[​](#create-route-and-configure-rate-limiting-1 "Direct link to Create Route and Configure Rate Limiting") Create a route with the following configurations in the gateway group: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- </dev/null | grep -v "^$" || echo "No related keys found" done ``` You should see output similar to the following: ``` Checking node redis-node-7000: No related keys found Checking node redis-node-7001: No related keys found Checking node redis-node-7002: plugin-ai-rate-limitingroute&service&::: Checking node redis-node-7003: plugin-ai-rate-limitingroute&service&::: Checking node redis-node-7004: No related keys found Checking node redis-node-7005: No related keys found ``` ### Share Quota Among Gateway Nodes with a Redis Sentinel[​](#share-quota-among-gateway-nodes-with-a-redis-sentinel "Direct link to Share Quota Among Gateway Nodes with a Redis Sentinel") This authenticated Redis Sentinel example applies to API7 Enterprise version 3.10.5 and later, and to APISIX version 3.18.0 and later. Use Redis Sentinel when you need automatic failover and high availability but do not require data partitioning. This pattern is simpler to manage and suitable for most high-availability requirements. Ensure that your Redis instances are running in [Sentinel mode](https://redis.io/docs/latest/operate/oss_and_stack/management/sentinel/). #### Prerequisites[​](#prerequisites-2 "Direct link to Prerequisites") * Docker * Kubernetes 1. Create a Docker network: ``` docker network create redis-sentinel-network ``` Ensure that your gateway instance is running within the same network as your Redis Sentinel cluster. 2. Start a Redis master node: ``` docker run -d --name redis-master --network redis-sentinel-network \ -p 6379:6379 \ redis:7.2-alpine \ redis-server --requirepass StrongP@ss123 --appendonly yes ``` 3. Start Sentinel replica nodes: ``` for i in 1 2; do PORT=$((6380 + i - 1)) docker run -d --name redis-slave-$i --network redis-sentinel-network \ -p $PORT:6379 \ redis:7.2-alpine \ redis-server --slaveof redis-master 6379 \ --requirepass StrongP@ss123 \ --masterauth StrongP@ss123 \ --appendonly yes done ``` 4. Get master node IP address for next step: ``` MASTER_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' redis-master) echo "Redis master node IP: $MASTER_IP" ``` 5. Start Sentinel cluster and replace `$MASTER_IP` with your master node IP: ``` for i in 1 2 3; do docker run -d --name redis-sentinel-$i --network redis-sentinel-network -p $((26378+i-1)):26379 \ redis:7.2-alpine \ sh -c " cat << 'EOF' > /sentinel.conf port 26379 sentinel monitor mymaster $MASTER_IP 6379 2 sentinel auth-pass mymaster StrongP@ss123 requirepass admin-password sentinel down-after-milliseconds mymaster 5000 sentinel failover-timeout mymaster 10000 sentinel parallel-syncs mymaster 1 protected-mode no EOF redis-sentinel /sentinel.conf " done echo "✅ Sentinel cluster started successfully." ``` You can see the following response: ``` Starting redis-sentinel-1 (port:26379)... eb9efacb629d0cfdfaa48856f42ba8c67642baa79f1589df5b251c11d3ec6e1a Starting redis-sentinel-2 (port:26380)... 7f23f4b6e63c9b6be4c5e1903a244f078d481952a1465a9650c743ea2ee4600f Starting redis-sentinel-3 (port:26381)... 1df087502124e3903df7ae665ef597bf735669c5ce3f9d87696c4acd82526626 ✅ Sentinel cluster started successfully. ``` 6. Confirm the Sentinel environment is running correctly: ``` echo "Waiting for Sentinel cluster establishment (10 seconds)..." sleep 10 echo -e "\nVerifying Sentinel cluster status:" for i in 1 2 3; do echo "--- Sentinel $i status ---" if docker ps | grep -q "redis-sentinel-$i"; then echo "Container: ✅ Running" docker exec redis-sentinel-$i redis-cli -p 26379 SENTINEL master mymaster 2>&1 | grep -E "(flags|num-slaves|num-other-sentinels)" else echo "Container: ❌ Not running (run 'docker logs redis-sentinel-$i' to check)" fi echo "" done ``` You can see the following response: ``` Verifying Sentinel cluster status: --- Sentinel 1 status --- Container: ✅ Running --- Sentinel 2 status --- Container: ✅ Running --- Sentinel 3 status --- Container: ✅ Running ``` 7. Get Sentinel IP addresses for plugin configuration: ``` echo -e "Getting Sentinel container IP addresses:" for i in 1 2 3; do IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' redis-sentinel-$i) echo " redis-sentinel-$i : $IP" done ``` You can see the following response: ``` Getting Sentinel container IP addresses: redis-sentinel-1 : 172.22.0.4 redis-sentinel-2 : 172.22.0.5 redis-sentinel-3 : 172.22.0.6 ``` 8. Conduct detailed status check: ``` echo "Checking detailed Sentinel cluster status..." for i in 1 2 3; do echo "=== Sentinel $i details ===" docker exec redis-sentinel-$i redis-cli -p 26379 SENTINEL master mymaster echo "" done ``` You can see the following response: ``` Checking detailed Sentinel cluster status... === Sentinel 1 Details === name: mymaster ip: 172.22.0.2 port: 6379 runid: ${YOUR_RUN_ID} flags: master link-pending-commands: 0 link-refcount: 1 last-ping-sent: 0 last-ok-ping-reply: 113 last-ping-reply: 113 down-after-milliseconds: 5000 info-refresh: 6979 role-reported: master role-reported-time: 107360 config-epoch: 0 num-slaves: 1 num-other-sentinels: 2 quorum: 2 failover-timeout: 10000 parallel-syncs: 1 ... ``` Create a Kubernetes manifest for the Redis master, replicas, and Sentinel cluster: redis-sentinel.yaml ``` apiVersion: v1 kind: ConfigMap metadata: namespace: aic name: redis-sentinel-config data: sentinel.conf: | port 26379 sentinel monitor mymaster redis-master.aic.svc 6379 2 sentinel auth-pass mymaster StrongP@ss123 requirepass admin-password sentinel down-after-milliseconds mymaster 5000 sentinel failover-timeout mymaster 10000 sentinel parallel-syncs mymaster 1 protected-mode no --- apiVersion: apps/v1 kind: StatefulSet metadata: namespace: aic name: redis-master spec: serviceName: redis-master replicas: 1 selector: matchLabels: app: redis-master template: metadata: labels: app: redis-master spec: containers: - name: redis image: redis:7.2-alpine ports: - containerPort: 6379 command: - redis-server - --requirepass - StrongP@ss123 - --appendonly - "yes" --- apiVersion: v1 kind: Service metadata: namespace: aic name: redis-master spec: clusterIP: None selector: app: redis-master ports: - port: 6379 --- apiVersion: apps/v1 kind: StatefulSet metadata: namespace: aic name: redis-replica spec: serviceName: redis-replica replicas: 2 selector: matchLabels: app: redis-replica template: metadata: labels: app: redis-replica spec: containers: - name: redis image: redis:7.2-alpine ports: - containerPort: 6379 command: - redis-server - --slaveof - redis-master.aic.svc - "6379" - --requirepass - StrongP@ss123 - --masterauth - StrongP@ss123 - --appendonly - "yes" --- apiVersion: v1 kind: Service metadata: namespace: aic name: redis-replica spec: clusterIP: None selector: app: redis-replica ports: - port: 6379 --- apiVersion: apps/v1 kind: StatefulSet metadata: namespace: aic name: redis-sentinel spec: serviceName: redis-sentinel replicas: 3 selector: matchLabels: app: redis-sentinel template: metadata: labels: app: redis-sentinel spec: containers: - name: sentinel image: redis:7.2-alpine ports: - containerPort: 26379 command: - redis-sentinel - /etc/sentinel/sentinel.conf volumeMounts: - name: sentinel-config mountPath: /etc/sentinel volumes: - name: sentinel-config configMap: name: redis-sentinel-config --- apiVersion: v1 kind: Service metadata: namespace: aic name: redis-sentinel spec: clusterIP: None selector: app: redis-sentinel ports: - port: 26379 ``` Apply the manifest: ``` kubectl apply -f redis-sentinel.yaml ``` Wait for all pods to be ready: ``` kubectl wait --for=condition=Ready pod -l app=redis-sentinel -n aic --timeout=120s ``` #### Create Route and Configure Rate Limiting[​](#create-route-and-configure-rate-limiting-2 "Direct link to Create Route and Configure Rate Limiting") Create a route with the following configurations in the gateway group: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- <When `rules` is set, the headers use a prefix instead. See `rules.header_prefix` for details. - limit\_strategy string default: `total_tokens` vaild vaule: `total_tokens`, `prompt_tokens`, `completion_tokens`, or `expression` *** Type of token to apply rate limiting. `total_tokens`, `prompt_tokens`, and `completion_tokens` values are returned in each model response, where `total_tokens` is the sum of `prompt_tokens` and `completion_tokens`. When set to `expression`, rate limiting cost is calculated using a custom Lua arithmetic expression defined in `cost_expr`. Available in API7 Enterprise from version 3.9.8 and APISIX from version 3.17.0. - cost\_expr string vaild vaule: any non-empty string (must be a valid Lua arithmetic expression) *** Lua arithmetic expression for dynamic token cost calculation. Variables are injected from the LLM provider's raw usage response fields (e.g., `input_tokens`, `output_tokens`, `cache_creation_input_tokens`). Missing variables default to `0`. Only math functions (`abs`, `ceil`, `floor`, `max`, `min`) and arithmetic operators are allowed. Expression syntax is validated when `limit_strategy` is `expression`, where this field is required. Example: `input_tokens + cache_creation_input_tokens` computes cost from Anthropic Claude's cache-aware token usage. Available in API7 Enterprise from version 3.9.8 and APISIX from version 3.17.0. - instances array\[object] *** LLM instance rate limiting configurations. * name string required *** Name of the LLM service instance. * limit integer | string required vaild vaule: greater than 0 *** The maximum number of tokens allowed to consume within a given time interval. In API7 Enterprise (from 3.8.17) and in APISIX (from 3.16.0), this parameter also supports the string data type and allows the use of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) prefixed with a dollar sign (`$`). In earlier APISIX versions, only the integer type is supported. * time\_window integer | string required vaild vaule: greater than 0 *** The time interval corresponding to the rate limiting `limit` in seconds. In API7 Enterprise (from 3.8.17) and in APISIX (from 3.16.0), this parameter also supports the string data type and allows the use of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) prefixed with a dollar sign (`$`). In earlier APISIX versions, only the integer type is supported. - rejected\_code integer default: `503` vaild vaule: between 200 and 599 inclusive *** The HTTP status code returned when a request exceeding the quota is rejected. - rejected\_msg string vaild vaule: any non-empty string *** The response body returned when a request exceeding the quota is rejected. - policy string default: `local` vaild vaule: `local`, `redis`, `redis-cluster`, or `redis-sentinel` *** The policy for rate limiting counters. API7 Gateway requires this field; use `local` for configurations that do not use Redis. APISIX uses `local` when the field is omitted. Redis-backed policies are available in API7 Enterprise from version 3.8.19 and in APISIX from version 3.18.0. Set to `local` to store the counter in memory locally. Set to `redis` to store the counter on a Redis instance. Set to `redis-cluster` to store the counter in a Redis cluster. Set to `redis-sentinel` to store the counter on the Redis primary node managed by Redis Sentinel, which ensures high availability by automatically promoting a replica to primary in case of failure. Redis Sentinel provides high availability for Redis when not using Redis Cluster. - redis\_host string *** The address of the Redis node. Required when `policy` is `redis`. - redis\_port integer default: `6379` vaild vaule: greater than or equal to 1 *** The port of the Redis node when `policy` is `redis`. - redis\_username string *** The username for Redis if Redis ACL is used. If you use the legacy authentication method `requirepass`, configure only `redis_password`. Used when `policy` is `redis`, and with `redis-sentinel` in API7 Enterprise 3.10.5 and APISIX 3.18.0. - redis\_password string *** Password of the Redis node when `policy` is `redis` or `redis-cluster`, and with `redis-sentinel` in API7 Enterprise 3.10.5 and APISIX 3.18.0. In API7 Gateway 3.10.2 or later in the 3.10 release series, and 3.9.16 or later in the 3.9 release series, the value is encrypted with AES256 before being saved to the database. In APISIX 3.18.0 or later, the value is encrypted with AES before being stored in etcd. - redis\_database integer default: `0` vaild vaule: greater than or equal to 0 *** The database number in Redis when `policy` is `redis` or `redis-sentinel`. - redis\_ssl boolean default: `false` *** If true, use SSL to connect to Redis when `policy` is `redis`. - redis\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis`. - redis\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** The Redis timeout value in milliseconds when `policy` is `redis` or `redis-cluster`. - redis\_cluster\_nodes array\[string] *** List of Redis cluster nodes with at least one address. Required when `policy` is `redis-cluster`. - redis\_cluster\_name string *** The name of the Redis cluster. Required when `policy` is `redis-cluster`. - redis\_cluster\_ssl boolean default: `false` *** If true, use SSL to connect to Redis cluster when `policy` is `redis-cluster`. - redis\_cluster\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis-cluster`. - redis\_sentinels array\[object] *** An array of Redis Sentinel nodes (host and port). Required when `policy` is `redis-sentinel`. - redis\_master\_name string *** The name of the Redis master group that Sentinels are monitoring. Required when `policy` is `redis-sentinel`. - redis\_role string default: `master` vaild vaule: `master` or `slave` *** The Redis node role to connect to. Configurable when `policy` is `redis-sentinel`. Set to `master` to connect to the current Redis master, and set to `slave` to connect to a Redis replica. - redis\_connect\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** Timeout in milliseconds for establishing a connection to a Redis node. Configurable when `policy` is `redis-sentinel`. - redis\_read\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** Timeout in milliseconds for reading data from a Redis node. Configurable when `policy` is `redis-sentinel`. - redis\_keepalive\_timeout integer default: `` `10000` for `redis` or `redis-cluster`; `60000` for `redis-sentinel` `` vaild vaule: `redis` and `redis-cluster`: greater than or equal to 1000; `redis-sentinel`: greater than or equal to 1 *** Time in milliseconds that an idle Redis connection is kept alive in the connection pool before being closed. Used by all Redis-backed policies in APISIX 3.18.0. In API7 Enterprise, it is used by `redis-sentinel`; `redis` and `redis-cluster` support it from version 3.9.16 on the 3.9 release series and version 3.10.3 on the 3.10 release series. - redis\_keepalive\_pool integer default: `100` vaild vaule: greater than or equal to 1 *** Maximum number of idle Redis connections in the keepalive pool. Used when `policy` is `redis` or `redis-cluster`. Available in API7 Enterprise from version 3.9.16 on the 3.9 release series and version 3.10.3 on the 3.10 release series, and in APISIX from version 3.18.0. - sentinel\_username string *** Username used to authenticate with the Redis Sentinel instance. Configurable when `policy` is `redis-sentinel`. - sentinel\_password string *** Password used to authenticate with the Redis Sentinel instance. Configurable when `policy` is `redis-sentinel`. In API7 Gateway 3.10.2 or later in the 3.10 release series, and 3.9.16 or later in the 3.9 release series, the value is encrypted with AES256 before being saved to the database. In APISIX 3.18.0 or later, the value is encrypted with AES before being stored in etcd. - allow\_degradation boolean default: `false` *** If true, allow the gateway to continue handling requests without the plugin when the plugin or its dependencies become unavailable. Available in API7 Enterprise from version 3.8.19 and in APISIX from version 3.18.0. - rules array\[object] *** An array of rate-limiting rules that are applied sequentially. Available in API7 Enterprise from 3.8.17 and in APISIX from 3.16.0. * count integer | string required vaild vaule: greater than 0 *** The maximum number of tokens allowed to consume within a given time interval. This parameter also supports the string data type and allows the use of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) prefixed with a dollar sign (`$`). * time\_window integer | string required vaild vaule: greater than 0 *** The time interval corresponding to the rate limiting `count` in seconds. This parameter also supports the string data type and allows the use of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) prefixed with a dollar sign (`$`). * key string required *** The key to count requests by. If the configured key does not exist, the rule will not be executed. The `key` is interpreted as a variable. The variable does not need to be prefixed by a dollar sign (`$`). See [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) for available variables. * header\_prefix string *** Prefix for all rate limiting response headers. Available in API7 Enterprise from version 3.8.19 and in APISIX from version 3.17.0. When configured, the prefix is inserted after `X-AI-` in the header name. For example, with `header_prefix` set to `test`, the headers become `X-AI-Test-RateLimit-Limit`, `X-AI-Test-RateLimit-Remaining`, and `X-AI-Test-RateLimit-Reset`. When not configured, the index of the rule in the rules array is used as the prefix. For example, headers for the first rule will be `X-AI-1-RateLimit-Limit`, `X-AI-1-RateLimit-Remaining`, and `X-AI-1-RateLimit-Reset`. --- # ai-request-rewrite The `ai-request-rewrite` plugin processes client requests by forwarding them to LLM services for transformation before relaying them to upstream services. This enables LLM-powered modifications such as data redaction, content enrichment, or reformatting. The plugin supports the integration with OpenAI, DeepSeek, Gemini, Vertex AI, Anthropic, OpenRouter, and other OpenAI-compatible APIs. The LLM call used for rewriting is separate from the client's request format. With the `openai` provider, the plugin builds a non-streaming Chat Completions request from the configured prompt and the original request body. It does not classify or proxy the client's request as an AI protocol. That internal request carries the plugin's configured provider credentials and does not forward client headers. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `ai-request-rewrite` for different scenarios. The examples will use OpenAI as the LLM service. To follow along, obtain the OpenAI [API key](https://openai.com/blog/openai-api) and save it to an environment variable: ``` export OPENAI_API_KEY=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 # replace with your API key ``` ### Redact Sensitive Information[​](#redact-sensitive-information "Direct link to Redact Sensitive Information") The following example demonstrates how you can use the `ai-request-rewrite` plugin to redact sensitive information before the request reaches the upstream service. * Admin API * ADC * Ingress Controller Create a route and configure the `ai-request-rewrite` plugin as such: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < Authorization Scopes tab with the Create button highlighted](https://static.api7.ai/uploads/2024/01/06/bVHhiALe_auth-scope.png) Create the resource `httpbin-anything` with URI `/anything` and scope `access`: ![Keycloak client Authorization > Resources tab listing the Default Resource, with the Create button highlighted](https://static.api7.ai/uploads/2024/01/06/15DJ9HAU_create-resource.png) Create the client scope policy `access-client-scope-policy` that requires `httpbin-access`: ![Keycloak client Authorization > Policies tab listing the Default Policy, with the Create Policy dropdown highlighted](https://static.api7.ai/uploads/2024/01/06/7UtT3cF6_create-policy.png) Create the scope-based permission `access-scope-perm` that uses the `access` scope and `access-client-scope-policy`: ![Adding a scope-based permission in Keycloak](https://static.api7.ai/uploads/2024/01/12/Y0vlk1Tj_add-scope-permission.png) Add `httpbin-access` to the default client scopes of `apisix-quickstart-client`: ![Attaching a client scope to the Keycloak client](https://static.api7.ai/uploads/2024/01/06/sJKUMUcP_add-client-scope.png) Create a user named `quickstart-user`: ![Saving a new Keycloak user](https://static.api7.ai/uploads/2024/01/12/3fUQOFWg_save-user.png) Set the password to `quickstart-user-pass` and turn off **Temporary**: ![Setting the password for the Keycloak user](https://static.api7.ai/uploads/2024/01/12/aoabcBbC_set-password.png) Save the client secret from **Clients** > `apisix-quickstart-client` > **Credentials**: ![Keycloak client credentials tab showing the generated client secret](https://static.api7.ai/uploads/2024/01/12/3VqiXdf9_client-secret.png) Save the OIDC client ID and secret to environment variables: ``` OIDC_CLIENT_ID=apisix-quickstart-client OIDC_CLIENT_SECRET=bSaIN3MV1YynmtXvU8lKkfeY0iwpr9cH # replace with your value ``` tip If APISIX runs in Kubernetes, use the same Keycloak hostname consistently in both the plugin configuration and the token request. Otherwise, Keycloak may reject the bearer token because the token issuer does not match the configured authorization endpoints. #### Request Access Token[​](#request-access-token "Direct link to Request Access Token") Request an access token from Keycloak and save it to `ACCESS_TOKEN`: * Docker * Kubernetes ``` ACCESS_TOKEN=$(curl -sS "$KEYCLOAK_URL/realms/quickstart-realm/protocol/openid-connect/token" \ -d 'grant_type=client_credentials' \ -d 'client_id='$OIDC_CLIENT_ID'' \ -d 'client_secret='$OIDC_CLIENT_SECRET'' | jq -r '.access_token') ``` Run the token request inside the Keycloak pod and save the result to `ACCESS_TOKEN`: ``` ACCESS_TOKEN=$(kubectl exec -n aic deploy/keycloak -- env OIDC_CLIENT_SECRET="$OIDC_CLIENT_SECRET" sh -lc 'curl -sS "http://keycloak.aic.svc.cluster.local:8080/realms/quickstart-realm/protocol/openid-connect/token" \ -d grant_type=client_credentials \ -d client_id=apisix-quickstart-client \ -d client_secret="$OIDC_CLIENT_SECRET"' | jq -r '.access_token') ``` ### Use Lazy Load Path and Resource Registration Endpoint[​](#use-lazy-load-path-and-resource-registration-endpoint "Direct link to Use Lazy Load Path and Resource Registration Endpoint") The examples below demonstrate how you can configure `authz-keycloak` to dynamically resolve the request URI to one or more resources using the resource registration endpoint instead of static permissions. * Admin API * ADC * Ingress Controller Create a route with `authz-keycloak-route` as follows: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < info Amazon API Gateway supports HTTP APIs and REST APIs. API key support is available only for REST APIs, which is why this example uses a REST API trigger. You should now be redirected back to the Lambda interface. To find the API key and gateway API endpoint, go to the **Configuration** tab of the Lambda function and under **Triggers**, you can find the details of the API gateway: ![AWS console showing the API Gateway endpoint URL and the generated API key](https://static.api7.ai/uploads/2024/04/25/6bjpeNIb_api-gateway-info.png) Finally, create a route in APISIX with your gateway endpoint and API key: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "aws-lambda-route", "uri": "/aws-lambda", "plugins": { "aws-lambda": { "function_uri": "https://your-api-id.execute-api.us-west-2.amazonaws.com/default/your-resource", "authorization": { "apikey": "YOUR_API_GATEWAY_API_KEY" }, "ssl_verify": false } } }' ``` adc.yaml ``` services: - name: aws-lambda-service routes: - name: aws-lambda-route uris: - /aws-lambda plugins: aws-lambda: function_uri: https://your-api-id.execute-api.us-west-2.amazonaws.com/default/your-resource authorization: apikey: YOUR_API_GATEWAY_API_KEY ssl_verify: false ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD aws-lambda-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: aws-lambda-plugin-config spec: plugins: - name: aws-lambda config: function_uri: https://your-api-id.execute-api.us-west-2.amazonaws.com/default/your-resource authorization: apikey: YOUR_API_GATEWAY_API_KEY ssl_verify: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: aws-lambda-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /aws-lambda filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: aws-lambda-plugin-config ``` aws-lambda-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: aws-lambda-route spec: ingressClassName: apisix http: - name: aws-lambda-route match: paths: - /aws-lambda plugins: - name: aws-lambda enable: true config: function_uri: https://your-api-id.execute-api.us-west-2.amazonaws.com/default/your-resource authorization: apikey: YOUR_API_GATEWAY_API_KEY ssl_verify: false ``` Apply the configuration: ``` kubectl apply -f aws-lambda-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/aws-lambda" ``` You should receive an `HTTP/1.1 200 OK` response with the following message: ``` "Hello from Lambda!" ``` If your API key is invalid, you should receive an `HTTP/1.1 403 Forbidden` response. ### Forward Requests to Amazon API Gateway Sub-Paths[​](#forward-requests-to-amazon-api-gateway-sub-paths "Direct link to Forward Requests to Amazon API Gateway Sub-Paths") The following example demonstrates how you can forward requests to a sub-path of the Amazon API gateway API and configure the API to trigger the execution of Lambda function. Please follow the [previous example](#integrate-with-amazon-api-gateway-securely-with-api-key) to set up an API gateway first. To create a sub-path, go to the **Configuration** tab of the Lambda function and under **Triggers**, click into the API gateway: ![AWS Lambda console: opening the API Gateway integration for the function](https://static.api7.ai/uploads/2024/04/26/5Twffgyr_click-into-adjusted.png) Next, select **Create resource** to create a sub-path: ![AWS API Gateway Resources page with the Create resource button highlighted](https://static.api7.ai/uploads/2024/04/26/hXlnuVwk_create-resource.png) Enter the sub-path information and complete creation: ![complete resource creation](https://static.api7.ai/uploads/2024/04/26/7t1yiWjl_create-resource-2.png) Once redirected back to the main gateway console, you should see the newly created path. Select **Create method** to configure HTTP methods for the path and the associated action: ![AWS API Gateway Resources page showing the /api7-docs resource with the Create method button highlighted](https://static.api7.ai/uploads/2024/04/26/3rZZJy3e_create-method.png) Select the allowed HTTP method in the dropdown. For the purpose of demonstration, this example continues to use the same Lambda function as the triggered action when the path is requested: ![create method and lambda function](https://static.api7.ai/uploads/2024/04/26/vni7yS2q_create%20method%202.png) Finish the method creation. Once redirected back to the main gateway console, click on **Deploy API** to deploy the path and method changes: ![AWS API Gateway method execution view (Client to Method/Integration request to Lambda integration) with the Deploy API button highlighted](https://static.api7.ai/uploads/2024/04/26/2vrqnVPB_deploy-api.png) Finally, create a route in APISIX with your gateway endpoint and API key: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "aws-lambda-route", "uri": "/aws-lambda/*", "plugins": { "aws-lambda": { "function_uri": "https://your-api-id.execute-api.us-west-2.amazonaws.com/default", "authorization": { "apikey": "YOUR_API_GATEWAY_API_KEY" }, "ssl_verify": false } } }' ``` adc.yaml ``` services: - name: aws-lambda-service routes: - name: aws-lambda-route uris: - /aws-lambda/* plugins: aws-lambda: function_uri: https://your-api-id.execute-api.us-west-2.amazonaws.com/default authorization: apikey: YOUR_API_GATEWAY_API_KEY ssl_verify: false ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD aws-lambda-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: aws-lambda-plugin-config spec: plugins: - name: aws-lambda config: function_uri: https://your-api-id.execute-api.us-west-2.amazonaws.com/default authorization: apikey: YOUR_API_GATEWAY_API_KEY ssl_verify: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: aws-lambda-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /aws-lambda/ filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: aws-lambda-plugin-config ``` aws-lambda-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: aws-lambda-route spec: ingressClassName: apisix http: - name: aws-lambda-route match: paths: - /aws-lambda/* plugins: - name: aws-lambda enable: true config: function_uri: https://your-api-id.execute-api.us-west-2.amazonaws.com/default authorization: apikey: YOUR_API_GATEWAY_API_KEY ssl_verify: false ``` Apply the configuration: ``` kubectl apply -f aws-lambda-ic.yaml ``` ❶ match all sub-paths of `/aws-lambda/` ❷ For Admin API, ADC, and APISIX CRD examples, the sub-paths matched by the wildcard `*` will be appended to the end of the `function_uri`. In the Gateway API example, `PathPrefix` matches requests under `/aws-lambda/`, so the forwarded request path continues after the configured `function_uri` prefix. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/aws-lambda/api7-docs" ``` APISIX will forward the request to `https://your-api-id.execute-api.us-west-2.amazonaws.com/default/api7-docs` and you should receive an `HTTP/1.1 200 OK` response with the following message: ``` "Hello from Lambda!" ``` If your API key is invalid or if the requested path is not associated with any method, you should receive an `HTTP/1.1 403 Forbidden` response. --- ## Attributes[​](#attributes "Direct link to Attributes") ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * function\_uri string required *** AWS Lambda function URL or AWS API gateway endpoint that triggers the Lambda function. * authorization object *** Credentials used in authentication and authorization on AWS to invoke Lambda function. * apikey string *** API key for the REST API gateway when API key is selected as the security mechanism. API7 Gateway encrypts the value with AES at rest. APISIX encrypts it before etcd storage when `apisix.data_encryption.enable_encrypt_fields` is enabled. * iam object *** IAM credentials to be authenticated using [AWS Signature Version 4](https://docs.aws.amazon.com/AmazonS3/latest/API/sig-v4-authenticating-requests.html) and authorized. * accesskey string *** IAM user access key. API7 Gateway encrypts the value with AES at rest. APISIX encrypts it before etcd storage when `apisix.data_encryption.enable_encrypt_fields` is enabled. * secretkey string *** IAM user secret access key. API7 Gateway encrypts the value with AES at rest. APISIX encrypts it before etcd storage when `apisix.data_encryption.enable_encrypt_fields` is enabled. * aws\_region string default: `us-east-1` *** AWS region. * service string default: `execute-api` *** Service receiving the request. To integrate with AWS API gateway for API execution, set the service to `execute-api` for HTTP trigger. To integrate with Lambda function directly, set the service to `lambda`. * timeout integer default: `3000` vaild vaule: greater than or equal to 100 *** Proxy request timeout in milliseconds. * ssl\_verify boolean default: `true` *** If true, perform SSL verification. * keepalive boolean default: `true` *** If true, keep the connection alive for reuse. * keepalive\_pool integer default: `5` vaild vaule: greater than or equal to 1 *** If true, keep the connection alive for reuse. * keepalive\_timeout integer default: `60000` vaild vaule: greater than or equal to 1000 *** Time for connection to remain idle without closing in milliseconds. * max\_req\_body\_size integer default: `67108864` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes read before the request is sent to AWS Lambda. A larger body is rejected with `400 Bad Request`. Introduced in API7 Enterprise 3.9.17 and 3.10.4, and APISIX 3.18.0. --- # basic-auth The `basic-auth` plugin adds [basic access authentication](https://en.wikipedia.org/wiki/Basic_access_authentication) for [consumers](https://docs.api7.ai/apisix/key-concepts/consumers.md) to authenticate themselves before being able to access upstream resources. When a consumer is successfully authenticated, APISIX adds additional headers, such as `X-Consumer-Username`, `X-Credential-Identifier`, and other consumer custom headers if configured, to the request, before proxying it to the upstream service. The upstream service will be able to differentiate between consumers and implement additional logics as needed. If any of these values is not available, the corresponding header will not be added. About X-Consumer-Username When consumers are configured using the Ingress Controller, the consumer name is generated in the format `namespace_consumername`. As a result, the `X-Consumer-Username` header will also follow this format instead of just `consumername`. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can work with the `basic-auth` plugin for different scenarios. ### Implement Basic Authentication on Route[​](#implement-basic-authentication-on-route "Direct link to Implement Basic Authentication on Route") The following example demonstrates how to implement basic authentication on a route. * Admin API * ADC * Ingress Controller Create a consumer `johndoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "johndoe" }' ``` Create `basic-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/johndoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-basic-auth", "plugins": { "basic-auth": { "username": "johndoe", "password": "john-key" } } }' ``` Create a route with `basic-auth`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "basic-auth-route", "uri": "/anything", "plugins": { "basic-auth": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `basic-auth` credential and a route with `basic-auth` plugin configured: adc.yaml ``` consumers: - username: johndoe credentials: - name: basic-auth type: basic-auth config: username: johndoe password: john-key services: - name: basic-auth-service routes: - name: basic-auth-route uris: - /anything plugins: basic-auth: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create a consumer with `basic-auth` credential and a route with `basic-auth` plugin configured: * Gateway API * APISIX CRD basic-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: johndoe spec: gatewayRef: name: apisix credentials: - type: basic-auth name: primary-cred config: username: johndoe password: john-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: basic-auth-plugin-config spec: plugins: - name: basic-auth config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: basic-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: basic-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f basic-auth-ic.yaml ``` basic-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: johndoe spec: ingressClassName: apisix authParameter: basicAuth: value: username: johndoe password: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: basic-auth-route spec: ingressClassName: apisix http: - name: basic-auth-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: basic-auth enable: true ``` Apply the configuration to your cluster: ``` kubectl apply -f basic-auth-ic.yaml ``` #### Verify with Valid Credentials[​](#verify-with-valid-credentials "Direct link to Verify with Valid Credentials") Send a request to the route with valid credentials: ``` curl -i "http://127.0.0.1:9080/anything" -u johndoe:john-key ``` You should see an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "headers": { "Accept": "*/*", "Authorization": "Basic am9obmRvZTpqb2huLWtleQ==", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-66e5107c-5bb3e24f2de5baf733aec1cc", "X-Consumer-Username": "johndoe", "X-Credential-Identifier": "cred-john-basic-auth", "X-Forwarded-Host": "127.0.0.1" }, "origin": "192.168.65.1, 205.198.122.37", "url": "http://127.0.0.1/anything" } ``` #### Verify with Invalid Credentials[​](#verify-with-invalid-credentials "Direct link to Verify with Invalid Credentials") Send a request with invalid credentials: ``` curl -i "http://127.0.0.1:9080/anything" -u johndoe:invalid-password ``` You should see an `HTTP/1.1 401 Unauthorized` response with the following: ``` {"message":"Invalid user authorization"} ``` #### Verify without Credentials[​](#verify-without-credentials "Direct link to Verify without Credentials") Send a request without credentials: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should see an `HTTP/1.1 401 Unauthorized` response with the following: ``` {"message":"Missing authorization in request"} ``` ### Hide Authentication Information From Upstream[​](#hide-authentication-information-from-upstream "Direct link to Hide Authentication Information From Upstream") The following example demonstrates how to prevent the client's credentials (the `Authorization` header) from being sent to the upstream services by configuring `hide_credentials`. If you are using APISIX, the `Authorization` header containing the client's credentials is forwarded to the upstream services by default, which might lead to security risks in some circumstances and you should consider updating `hide_credentials` as shown in this example. * Admin API * ADC * Ingress Controller Create a consumer `johndoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "johndoe" }' ``` Create `basic-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/johndoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-basic-auth", "plugins": { "basic-auth": { "username": "johndoe", "password": "john-key" } } }' ``` #### Without Hiding Credentials[​](#without-hiding-credentials "Direct link to Without Hiding Credentials") Create a route with `basic-auth` and configure `hide_credentials` to `false`, which is the default configuration: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "basic-auth-route", "uri": "/anything", "plugins": { "basic-auth": { "hide_credentials": false } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `basic-auth` credential and a route with `basic-auth` plugin configured: adc.yaml ``` consumers: - username: johndoe credentials: - name: basic-auth type: basic-auth config: username: johndoe password: john-key services: - name: basic-auth-service routes: - name: basic-auth-route uris: - /anything plugins: basic-auth: hide_credentials: false upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create a consumer with `basic-auth` credential and a route with `basic-auth` plugin configured: * Gateway API * APISIX CRD basic-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: johndoe spec: gatewayRef: name: apisix credentials: - type: basic-auth name: primary-cred config: username: johndoe password: john-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: basic-auth-plugin-config spec: plugins: - name: basic-auth config: _meta: disable: false hide_credentials: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: basic-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: basic-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f basic-auth-ic.yaml ``` basic-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: johndoe spec: ingressClassName: apisix authParameter: basicAuth: value: username: johndoe password: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: basic-auth-route spec: ingressClassName: apisix http: - name: basic-auth-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: basic-auth enable: true config: hide_credentials: false ``` Apply the configuration to your cluster: ``` kubectl apply -f basic-auth-ic.yaml ``` Send a request with the valid key: ``` curl -i "http://127.0.0.1:9080/anything" -u johndoe:john-key ``` You should see an `HTTP/1.1 200 OK` response with the following: ``` { "args": {}, "data": "", "files": {}, "form": {}, "headers": { "Accept": "*/*", "Authorization": "Basic am9obmRvZTpqb2huLWtleQ==", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-66cc2195-22bd5f401b13480e63c498c6", "X-Consumer-Username": "johndoe", "X-Credential-Identifier": "cred-john-basic-auth", "X-Forwarded-Host": "127.0.0.1" }, "json": null, "method": "GET", "origin": "192.168.65.1, 43.228.226.23", "url": "http://127.0.0.1/anything" } ``` Note that the credentials are visible to the upstream service in base64-encoded format. tip You can also pass the base64-encoded credentials in the request using the `Authorization` header as such: ``` curl -i "http://127.0.0.1:9080/anything" -H "Authorization: Basic am9obmRvZTpqb2huLWtleQ==" ``` #### Hide Credentials[​](#hide-credentials "Direct link to Hide Credentials") * Admin API * ADC * Ingress Controller Update the plugin's `hide_credentials` to `true`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes/basic-auth-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "basic-auth": { "hide_credentials": true } } }' ``` Update the route configuration: adc.yaml ``` # other configs # ... services: - name: basic-auth-service routes: - name: basic-auth-route uris: - /anything plugins: basic-auth: hide_credentials: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update the PluginConfig to set `hide_credentials` to `true`: basic-auth-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: basic-auth-plugin-config spec: plugins: - name: basic-auth config: _meta: disable: false hide_credentials: true ``` Apply the configuration to your cluster: ``` kubectl apply -f basic-auth-ic.yaml ``` Update the ApisixRoute to set `hide_credentials` to `true`: basic-auth-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: basic-auth-route spec: ingressClassName: apisix http: - name: basic-auth-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: basic-auth enable: true config: hide_credentials: true ``` Apply the configuration to your cluster: ``` kubectl apply -f basic-auth-ic.yaml ``` Send a request with the valid key: ``` curl -i "http://127.0.0.1:9080/anything" -u johndoe:john-key ``` You should see an `HTTP/1.1 200 OK` response with the following: ``` { "args": {}, "data": "", "files": {}, "form": {}, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-66cc21a7-4f6ac87946e25f325167d53a", "X-Consumer-Username": "johndoe", "X-Credential-Identifier": "cred-john-basic-auth", "X-Forwarded-Host": "127.0.0.1" }, "json": null, "method": "GET", "origin": "192.168.65.1, 43.228.226.23", "url": "http://127.0.0.1/anything" } ``` Note that the credentials are no longer visible to the upstream service. ### Add Consumer Custom ID to Header[​](#add-consumer-custom-id-to-header "Direct link to Add Consumer Custom ID to Header") The following example demonstrates how you can attach a consumer custom ID to authenticated request in the `Consumer-Custom-Id` header, which can be used to implement additional logics as needed. * Admin API * ADC * Ingress Controller Create a consumer `johndoe` with a custom ID label: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "johndoe", "labels": { "custom_id": "495aec6a" } }' ``` Create `basic-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/johndoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-basic-auth", "plugins": { "basic-auth": { "username": "johndoe", "password": "john-key" } } }' ``` Create a route with `basic-auth`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "basic-auth-route", "uri": "/anything", "plugins": { "basic-auth": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `basic-auth` credential and a route with `basic-auth` plugin enabled: adc.yaml ``` consumers: - username: johndoe labels: custom_id: "495aec6a" credentials: - name: basic-auth type: basic-auth config: username: johndoe password: john-key services: - name: basic-auth-service routes: - name: basic-auth-route uris: - /anything plugins: basic-auth: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create a consumer with `basic-auth` credential and a route with `basic-auth` plugin enabled: * Gateway API * APISIX CRD basic-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: johndoe labels: custom_id: "495aec6a" spec: gatewayRef: name: apisix credentials: - type: basic-auth name: primary-key config: username: johndoe password: john-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: basic-auth-plugin-config spec: plugins: - name: basic-auth config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: basic-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: basic-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f basic-auth-ic.yaml ``` basic-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: johndoe labels: custom_id: "495aec6a" spec: ingressClassName: apisix authParameter: basicAuth: value: username: johndoe password: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: basic-auth-route spec: ingressClassName: apisix http: - name: basic-auth-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: basic-auth enable: true config: _meta: disable: false ``` Apply the configuration to your cluster: ``` kubectl apply -f basic-auth-ic.yaml ``` To verify, send a request to the route with the valid key: ``` curl -i "http://127.0.0.1:9080/anything" -u johndoe:john-key ``` You should see an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "data": "", "files": {}, "form": {}, "headers": { "Accept": "*/*", "Authorization": "Basic am9obmRvZTpqb2huLWtleQ==", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-66ea8d64-33df89052ae198a706e18c2a", "X-Consumer-Username": "aic_johndoe", "X-Consumer-Custom-Id": "495aec6a", "X-Forwarded-Host": "127.0.0.1" }, "json": null, "method": "GET", "origin": "192.168.65.1, 205.198.122.37", "url": "http://127.0.0.1/anything" } ``` If you would like to attach more consumer custom headers to authenticated requests, see the [`attach-consumer-label`](https://docs.api7.ai/hub/attach-consumer-label.md) plugin. ### Rate Limit with Anonymous Consumer[​](#rate-limit-with-anonymous-consumer "Direct link to Rate Limit with Anonymous Consumer") The following example demonstrates how you can configure different rate limiting policies by regular and anonymous consumers, where the anonymous consumer does not need to authenticate and has less quota. * Admin API * ADC * Ingress Controller Create a regular consumer `johndoe` and configure the `limit-count` plugin to allow for a quota of 3 within a 30-second window: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "johndoe", "plugins": { "limit-count": { "count": 3, "time_window": 30, "rejected_code": 429, "policy": "local" } } }' ``` Create the `basic-auth` credential for the consumer `johndoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/johndoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-basic-auth", "plugins": { "basic-auth": { "username": "johndoe", "password": "john-key" } } }' ``` Create an anonymous user `anonymous` and configure the `limit-count` plugin to allow for a quota of 1 within a 30-second window: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "anonymous", "plugins": { "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "policy": "local" } } }' ``` Create a route and configure the `basic-auth` plugin to accept anonymous consumer `anonymous` from bypassing the authentication: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "basic-auth-route", "uri": "/anything", "plugins": { "basic-auth": { "anonymous_consumer": "anonymous" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Configure consumers with different rate limits and a route that accepts anonymous users: adc.yaml ``` consumers: - username: johndoe plugins: limit-count: count: 3 time_window: 30 rejected_code: 429 policy: local credentials: - name: basic-auth type: basic-auth config: username: johndoe password: john-key - username: anonymous plugins: limit-count: count: 1 time_window: 30 rejected_code: 429 policy: local services: - name: anonymous-rate-limit-service routes: - name: basic-auth-route uris: - /anything plugins: basic-auth: anonymous_consumer: anonymous upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Configure consumers with different rate limits and a route that accepts anonymous users: basic-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: johndoe spec: gatewayRef: name: apisix credentials: - type: basic-auth name: primary-key config: username: johndoe password: john-key plugins: - name: limit-count config: count: 3 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: anonymous spec: gatewayRef: name: apisix plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 policy: local --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: basic-auth-plugin-config spec: plugins: - name: basic-auth config: anonymous_consumer: aic_anonymous # namespace_consumername --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: basic-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: basic-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f basic-auth-ic.yaml ``` Configure consumers with different rate limits and a route that accepts anonymous users: basic-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: johndoe spec: ingressClassName: apisix authParameter: basicAuth: value: username: johndoe password: john-key plugins: - name: limit-count enable: true config: count: 3 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: anonymous spec: ingressClassName: apisix plugins: - name: limit-count enable: true config: count: 1 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: basic-auth-route spec: ingressClassName: apisix http: - name: basic-auth-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: basic-auth enable: true config: anonymous_consumer: aic_anonymous ``` Apply the configuration to your cluster: ``` kubectl apply -f basic-auth-ic.yaml ``` To verify, send five consecutive requests with `johndoe`'s credentials: ``` resp=$(seq 5 | xargs -I{} curl "http://127.0.0.1:9080/anything" -u johndoe:john-key -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that out of the 5 requests, 3 requests were successful (status code 200) while the others were rejected (status code 429). ``` 200: 3, 429: 2 ``` Send five anonymous requests: ``` resp=$(seq 5 | xargs -I{} curl "http://127.0.0.1:9080/anything" -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that only one request was successful: ``` 200: 1, 429: 4 ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. ### Credentials[​](#credentials "Direct link to Credentials") The following are plugin attributes available for configurations on [credentials](https://docs.api7.ai/apisix/key-concepts/credentials.md). * username string required *** Unique basic auth username for a consumer. * password string required *** Basic auth password for the consumer. In API7 Enterprise from version 3.9.20, the password must not be empty, and a password containing colons is accepted. Following RFC 7617, everything after the first colon of the decoded credentials is taken as the password. Whitespace is still stripped from both halves of the decoded credentials before they are compared, so a password containing spaces cannot be used. The password is encrypted with AES before being stored in etcd. You can also store it in an environment variable and reference it using the `$env://` prefix, or in a secret manager such as HashiCorp Vault's [KV secrets engine](https://developer.hashicorp.com/vault/docs/secrets/kv), and reference it using the `$secret://` prefix. For more information, see [secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md). A `$env://` or `$secret://` reference that resolves to an empty value still passes configuration validation, because references are resolved at request time. From API7 Enterprise version 3.9.20, the gateway fails closed in that case, rejecting every request for the consumer with `HTTP 401` and logging a warning. ### Routes or Services[​](#routes-or-services "Direct link to Routes or Services") The following are plugin attributes available for configurations on [routes](https://docs.api7.ai/apisix/key-concepts/routes.md) or [services](https://docs.api7.ai/apisix/key-concepts/services.md). * hide\_credentials boolean default: `false` *** If true, do not pass the authorization request header to upstream services. * anonymous\_consumer string *** Anonymous consumer name. If configured, allow anonymous users to bypass the authentication. See [Rate Limit with Anonymous Consumer](https://docs.api7.ai/hub/basic-auth.md#rate-limit-with-anonymous-consumer) for more details. * realm string default: `basic` *** Realm in the [`WWW-Authenticate`](https://datatracker.ietf.org/doc/html/rfc7235#section-4.1) response header returned with a `401 Unauthorized` response due to authentication failure. For example: * If `realm` is set to `basic-auth`, the 401 response will include the following header: ``` WWW-Authenticate: Basic realm="basic-auth" ``` * If `realm` is not configured, the 401 response will include the following header: ``` WWW-Authenticate: Basic realm="basic" ``` This parameter is available in API7 Enterprise version 3.9.2 and later, and in Apache APISIX version 3.15.0 and later. --- # body-transformer The `body-transformer` plugin performs template-based transformations to transform the request and/or response bodies from one format to another. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `body-transformer` for different scenarios. The transformation template uses [lua-resty-template](https://github.com/bungle/lua-resty-template) syntax. See the [template syntax](https://github.com/bungle/lua-resty-template#template-syntax) to learn more. You can also use auxiliary functions `_escape_json()` and `_escape_xml()` to escape special characters such as double quotes, `_body` to access request body, and `_ctx` to access context variables. In all cases, you should ensure that the transformation template is a valid JSON string. ### Transform between JSON and XML SOAP[​](#transform-between-json-and-xml-soap "Direct link to Transform between JSON and XML SOAP") The following example demonstrates how to transform the request body from JSON to XML and the response body from XML to JSON when working with a SOAP upstream service. Start the sample SOAP service: ``` cd /tmp git clone https://github.com/spring-guides/gs-producing-web-service.git cd gs-producing-web-service/complete ./mvnw spring-boot:run ``` Create the request and response transformation templates: ``` req_template=$(cat < {{_escape_xml(name)}} EOF ) rsp_template=$(cat < {{_escape_xml(name)}} input_format: json response: template: | {% if Envelope.Body.Fault == nil then %} { "status":"{{_ctx.var.status}}", "currency":"{{Envelope.Body.getCountryResponse.country.currency}}", "population":{{Envelope.Body.getCountryResponse.country.population}}, "capital":"{{Envelope.Body.getCountryResponse.country.capital}}", "name":"{{Envelope.Body.getCountryResponse.country.name}}" } {% else %} { "message":{*_escape_json(Envelope.Body.Fault.faultstring[1])*}, "code":"{{Envelope.Body.Fault.faultcode}}" {% if Envelope.Body.Fault.faultactor ~= nil then %} , "actor":"{{Envelope.Body.Fault.faultactor}}" {% end %} } {% end %} input_format: xml proxy-rewrite: headers: set: Content-Type: text/xml upstream: type: roundrobin nodes: - host: host.docker.internal port: 8080 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create a Kubernetes manifest file of a route with the `body-transformer` plugin: soap-route.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: json-xml-plugin spec: plugins: - name: body-transformer config: request: template: | {{_escape_xml(name)}} input_format: json response: template: | {% if Envelope.Body.Fault == nil then %} { "status":"{{_ctx.var.status}}", "currency":"{{Envelope.Body.getCountryResponse.country.currency}}", "population":{{Envelope.Body.getCountryResponse.country.population}}, "capital":"{{Envelope.Body.getCountryResponse.country.capital}}", "name":"{{Envelope.Body.getCountryResponse.country.name}}" } {% else %} { "message":{*_escape_json(Envelope.Body.Fault.faultstring[1])*}, "code":"{{Envelope.Body.Fault.faultcode}}" {% if Envelope.Body.Fault.faultactor ~= nil then %} , "actor":"{{Envelope.Body.Fault.faultactor}}" {% end %} } {% end %} input_format: xml - name: proxy-rewrite config: headers: set: Content-Type: text/xml --- apiVersion: v1 kind: Service metadata: namespace: aic name: ws-external-domain spec: type: ExternalName externalName: host.docker.internal --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: body-transformer-route spec: parentRefs: - name: apisix rules: - matches: - method: POST path: type: Exact value: /services filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: json-xml-plugin backendRefs: - name: ws-external-domain port: 8080 ``` Create a Kubernetes manifest file of a route with the `body-transformer` plugin: soap-route.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: body-transformer-route spec: ingressClassName: apisix http: - name: body-transformer-route match: paths: - /services methods: - POST plugins: - name: body-transformer enable: true config: request: template: | {{_escape_xml(name)}} input_format: json response: template: | {% if Envelope.Body.Fault == nil then %} { "status":"{{_ctx.var.status}}", "currency":"{{Envelope.Body.getCountryResponse.country.currency}}", "population":{{Envelope.Body.getCountryResponse.country.population}}, "capital":"{{Envelope.Body.getCountryResponse.country.capital}}", "name":"{{Envelope.Body.getCountryResponse.country.name}}" } {% else %} { "message":{*_escape_json(Envelope.Body.Fault.faultstring[1])*}, "code":"{{Envelope.Body.Fault.faultcode}}" {% if Envelope.Body.Fault.faultactor ~= nil then %} , "actor":"{{Envelope.Body.Fault.faultactor}}" {% end %} } {% end %} input_format: xml - name: proxy-rewrite enable: true config: headers: set: Content-Type: text/xml upstreams: - name: ws-external-domain --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: ws-external-domain spec: externalNodes: - type: Domain name: host.docker.internal port: 8080 ``` Apply the configuration to your cluster: ``` kubectl apply -f soap-route.yaml ``` ❶ Set the request input format as JSON, so that the plugin will apply the JSON decoder internally. ❷ Set the response input format as XML, so that the plugin will apply the XML decoder internally. ❸ Set the `Content-Type` header to `text/xml` for the upstream SOAP service to respond properly. ❹ The address of the SOAP service. `host.docker.internal` resolves to the host machine when APISIX runs in Docker. Replace with the actual address if you are running APISIX differently. tip If it is cumbersome to adjust complex text files to be valid transformation templates, you can use the base64 utility to encode the files, such as the following: ``` "body-transformer": { "request": { "template": "'"$(base64 -w0 /path/to/request_template_file)"'" }, "response": { "template": "'"$(base64 -w0 /path/to/response_template_file)"'" } } ``` Send a request with a valid JSON body: ``` curl "http://127.0.0.1:9080/services" -X POST -d '{"name": "Spain"}' ``` The JSON body sent in the request will be transformed into XML before being forwarded to the upstream SOAP service, and the response body will be transformed back from XML to JSON. You should see a response similar to the following: ``` { "status": "200", "currency": "EUR", "population": 46704314, "capital": "Madrid", "name": "Spain" } ``` ### Modify Request Body[​](#modify-request-body "Direct link to Modify Request Body") The following example demonstrates how to dynamically modify the request body. * Admin API * ADC * Ingress Controller Create a route with `body-transformer`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "body-transformer-route", "uri": "/anything", "plugins": { "body-transformer": { "request": { "template": "{\"foo\":\"{{name .. \" world\"}}\",\"bar\":{{age+10}}}" } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a route with `body-transformer`: adc.yaml ``` services: - name: body-transformer-service routes: - name: body-transformer-route uris: - /anything plugins: body-transformer: request: template: "{\"foo\":\"{{name .. \" world\"}}\",\"bar\":{{age+10}}}" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create a Kubernetes manifest file of a route with the `body-transformer` plugin: body-transformer-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: body-transformer-plugin-config spec: plugins: - name: body-transformer config: request: template: "{\"foo\":\"{{name .. \" world\"}}\",\"bar\":{{age+10}}}" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: body-transformer-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: body-transformer-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Create a Kubernetes manifest file of a route with the `body-transformer` plugin: body-transformer-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: body-transformer-route spec: ingressClassName: apisix http: - name: body-transformer-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: body-transformer enable: true config: request: template: "{\"foo\":\"{{name .. \" world\"}}\",\"bar\":{{age+10}}}" ``` Apply the configuration to your cluster: ``` kubectl apply -f body-transformer-ic.yaml ``` ❶ Set a template that appends "world" to the name and adds 10 to the age and set them as values to "foo" and "bar" respectively. Send a request to the route: ``` curl "http://127.0.0.1:9080/anything" -X POST \ -H "Content-Type: application/json" \ -d '{"name":"hello","age":20}' \ -i ``` You should see a response of the following: ``` { "args": {}, "data": "{\"foo\":\"hello world\",\"bar\":30}", ... "json": { "bar": 30, "foo": "hello world" }, "method": "POST", ... } ``` ### Generate Request Body Using Variables[​](#generate-request-body-using-variables "Direct link to Generate Request Body Using Variables") The following example demonstrates how to generate request body dynamically using the `ctx` context variables. * Admin API * ADC * Ingress Controller Create a route with `body-transformer`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "body-transformer-route", "uri": "/anything", "plugins": { "body-transformer": { "request": { "template": "{\"foo\":\"{{_ctx.var.arg_name .. \" world\"}}\"}" } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a route with `body-transformer`: adc.yaml ``` services: - name: body-transformer-service routes: - name: body-transformer-route uris: - /anything plugins: body-transformer: request: template: "{\"foo\":\"{{_ctx.var.arg_name .. \" world\"}}\"}" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create a Kubernetes manifest file of a route with the `body-transformer` plugin: body-transformer-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: body-transformer-plugin-config spec: plugins: - name: body-transformer config: request: template: "{\"foo\":\"{{_ctx.var.arg_name .. \" world\"}}\"}" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: body-transformer-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: body-transformer-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Create a Kubernetes manifest file of a route with the `body-transformer` plugin: body-transformer-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: body-transformer-route spec: ingressClassName: apisix http: - name: body-transformer-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: body-transformer enable: true config: request: template: "{\"foo\":\"{{_ctx.var.arg_name .. \" world\"}}\"}" ``` Apply the configuration to your cluster: ``` kubectl apply -f body-transformer-ic.yaml ``` ❶ Set a template which accesses the request argument using the [NGINX variable](https://docs.api7.ai/apisix/reference/built-in-variables.md#nginx-variables) `arg_name`. Send a request to the route with `name` argument: ``` curl -i "http://127.0.0.1:9080/anything?name=hello" ``` You should see a response like this: ``` { "args": { "name": "hello" }, ..., "json": { "foo": "hello world" }, ... } ``` ### Transform Body from YAML to JSON[​](#transform-body-from-yaml-to-json "Direct link to Transform Body from YAML to JSON") The following example demonstrates how to transform request body from YAML to JSON. Create the request transformation template: ``` req_template=$(cat < 18 then context._multipart:set_simple("status", "adult") else context._multipart:set_simple("status", "minor") end local body = context._multipart:tostring() %}{* body *} EOF ) ``` * Admin API * ADC * Ingress Controller Create a route with `body-transformer` as such: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < 18 then context._multipart:set_simple("status", "adult") else context._multipart:set_simple("status", "minor") end local body = context._multipart:tostring() %}{* body *} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create a Kubernetes manifest file of a route with the `body-transformer` plugin: body-transformer-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: body-transformer-plugin-config spec: plugins: - name: body-transformer config: request: input_format: multipart template: | {% if tonumber(context.age) > 18 then context._multipart:set_simple("status", "adult") else context._multipart:set_simple("status", "minor") end local body = context._multipart:tostring() %}{* body *} --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: body-transformer-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: body-transformer-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Create a Kubernetes manifest file of a route with the `body-transformer` plugin: body-transformer-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: body-transformer-route spec: ingressClassName: apisix http: - name: body-transformer-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: body-transformer enable: true config: request: input_format: multipart template: | {% if tonumber(context.age) > 18 then context._multipart:set_simple("status", "adult") else context._multipart:set_simple("status", "minor") end local body = context._multipart:tostring() %}{* body *} ``` Apply the configuration to your cluster: ``` kubectl apply -f body-transformer-ic.yaml ``` ❶ Set the `input_format` to `multipart`. ❷ Set to the previously created request template. Send a multipart POST request to the route: ``` curl -X POST \ -F "name=john" \ -F "age=10" \ "http://127.0.0.1:9080/anything" ``` You should see a response similar to the following: ``` { "args": {}, "data": "", "files": {}, "form": { "age": "10", "name": "john", "status": "minor" }, "headers": { "Accept": "*/*", "Content-Length": "361", "Content-Type": "multipart/form-data; boundary=------------------------qtPjk4c8ZjmGOXNKzhqnOP", ... }, ... } ``` ### Transform Response Body Based on Consumer Identity[​](#transform-response-body-based-on-consumer-identity "Direct link to Transform Response Body Based on Consumer Identity") The following example demonstrates how to customize response body transformations based on different consumer identities. The example shows how to return different response formats to different consumers while filtering sensitive fields and renaming properties. Create the response transformation template that applies different transformations based on the consumer identity: ``` rsp_template=$(cat < ## Response Headers[​](#response-headers "Direct link to Response Headers") The plugin can add the following response headers, depending on the configuration of `append_waf_resp_header` and `append_waf_debug_header`: | Header | Description | | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `X-APISIX-CHAITIN-WAF` | Indicates whether APISIX forwarded the request to the WAF server.
• `yes`: Request was forwarded to the WAF server.
• `no`: Request was not forwarded to the WAF server.
• `unhealthy`: Request matches the configured rules, but no WAF service is available.
• `err`: An error occurred during plugin execution. The `X-APISIX-CHAITIN-WAF-ERROR` header is also included with details.
• `waf-err`: Error while interacting with the WAF server. The `X-APISIX-CHAITIN-WAF-ERROR` header is also included with details.
• `timeout`: Request to the WAF server timed out. | | `X-APISIX-CHAITIN-WAF-TIME` | Round-trip time (RTT) in milliseconds for the request to the Chaitin WAF server, including both network latency and WAF server processing. | | `X-APISIX-CHAITIN-WAF-STATUS` | Status code returned to APISIX by the WAF server. | | `X-APISIX-CHAITIN-WAF-ACTION` | Action returned to APISIX by the WAF server.
• `pass`: Request was allowed by the WAF service.
• `reject`: Request was blocked by the WAF service. | | `X-APISIX-CHAITIN-WAF-ERROR` | Debug header. Contains WAF error message. | | `X-APISIX-CHAITIN-WAF-SERVER` | Debug header. Indicates which WAF server was selected. | ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `chaitin-waf` plugin for different scenarios. Before proceeding, make sure you have installed [Chaitin WAF (SafeLine)](https://docs.waf.chaitin.com/en/GetStarted/Deploy). ### Block Malicious Requests on a Route[​](#block-malicious-requests-on-a-route "Direct link to Block Malicious Requests on a Route") The following example demonstrates how to integrate with Chaitin WAF to protect traffic on a route, rejecting malicious requests immediately. * Admin API * ADC * Ingress Controller Configure the Chaitin WAF connection details using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) (update the address accordingly): ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/chaitin-waf" -X PUT \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "nodes": [ { "host": "172.22.222.5", "port": 8000 } ] }' ``` Create a route and enable `chaitin-waf` on the route to block requests identified to be malicious: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "chaitin-waf-route", "uri": "/anything", "plugins": { "chaitin-waf": { "mode": "block", "append_waf_resp_header": true, "append_waf_debug_header": true } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` ❶ Set `mode` to `block` to block requests identified to be malicious. ❷ Set `append_waf_resp_header` to `true` to include WAF-related standard response headers. ❸ Set `append_waf_debug_header` to `true` to include WAF-related debugging response headers. adc.yaml ``` plugin_metadata: chaitin-waf: nodes: - host: "172.22.222.5" port: 8000 services: - name: chaitin-waf-service routes: - name: chaitin-waf-route uris: - /anything plugins: chaitin-waf: mode: block append_waf_resp_header: true append_waf_debug_header: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` ❶ `nodes`: List of Chaitin WAF service addresses. Update the `host` and `port` to match your Chaitin WAF (SafeLine) deployment. ❷ Set `mode` to `block` to block requests identified to be malicious. ❸ Set `append_waf_resp_header` to `true` to include WAF-related standard response headers. ❹ Set `append_waf_debug_header` to `true` to include WAF-related debugging response headers. Update your GatewayProxy manifest to configure the plugin metadata. If Chaitin WAF is installed on a host machine outside the cluster, use the host's IP address. In most deployments, expose Chaitin WAF as a reachable Service, IP address, or DNS name inside the cluster, or use the node/LoadBalancer address instead. gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: chaitin-waf: nodes: - host: "172.22.222.5" port: 8000 ``` ❶ `nodes`: List of Chaitin WAF service addresses. Update the `host` and `port` to match your Chaitin WAF (SafeLine) deployment. * Gateway API * APISIX CRD Create a route with `chaitin-waf` to block malicious requests: chaitin-waf-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: chaitin-waf-plugin-config spec: plugins: - name: chaitin-waf config: mode: block append_waf_resp_header: true append_waf_debug_header: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: chaitin-waf-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: chaitin-waf-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f gatewayproxy.yaml -f chaitin-waf-ic.yaml ``` Create a route with `chaitin-waf` to block malicious requests: chaitin-waf-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: chaitin-waf-route spec: ingressClassName: apisix http: - name: chaitin-waf-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: chaitin-waf enable: true config: mode: block append_waf_resp_header: true append_waf_debug_header: true ``` Apply the configuration to your cluster: ``` kubectl apply -f gatewayproxy.yaml -f chaitin-waf-ic.yaml ``` ❷ Set `mode` to `block` to block requests identified to be malicious. ❸ Set `append_waf_resp_header` to `true` to include WAF-related standard response headers. ❹ Set `append_waf_debug_header` to `true` to include WAF-related debugging response headers. Send a standard request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Send a request with SQL injection to the route: ``` curl -i "http://127.0.0.1:9080/anything" -d 'a=1 and 1=1' ``` You should see an `HTTP/1.1 403 Forbidden` response similar to the following: ``` ... X-APISIX-CHAITIN-WAF-STATUS: 403 X-APISIX-CHAITIN-WAF-ACTION: reject X-APISIX-CHAITIN-WAF-SERVER: 172.22.222.5 X-APISIX-CHAITIN-WAF: yes X-APISIX-CHAITIN-WAF-TIME: 3 ... {"code": 403, "success":false, "message": "blocked by Chaitin SafeLine Web Application Firewall", "event_id": "276be6457d8447a4bf1f792501dfba6c"} ``` ### Monitor Requests for Malicious Intent[​](#monitor-requests-for-malicious-intent "Direct link to Monitor Requests for Malicious Intent") This example shows how to integrate with Chaitin WAF to monitor all routes with `chaitin-waf` without rejection, and to reject potentially malicious requests on a specific route. * Admin API * ADC * Ingress Controller Configure the Chaitin WAF connection details using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) (update the address accordingly) and configure the mode: ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/chaitin-waf" -X PUT \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "nodes": [ { "host": "172.22.222.5", "port": 8000 } ], "mode": "monitor" }' ``` ❶ Set `mode` to `monitor` in the plugin metadata. This applies to all `chaitin-waf` plugin instances if `mode` is not specified on a route. Create a route and enable `chaitin-waf` without any configuration on the route: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "chaitin-waf-route", "uri": "/anything", "plugins": { "chaitin-waf": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` plugin_metadata: chaitin-waf: nodes: - host: "172.22.222.5" port: 8000 mode: monitor services: - name: chaitin-waf-service routes: - name: chaitin-waf-route uris: - /anything plugins: chaitin-waf: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` ❶ `nodes`: List of Chaitin WAF service addresses. Update the `host` and `port` to match your Chaitin WAF (SafeLine) deployment. ❷ Set `mode` to `monitor` in the plugin metadata. This applies to all `chaitin-waf` plugin instances if `mode` is not specified on a route. Update your GatewayProxy manifest to configure the plugin metadata. If Chaitin WAF is installed on a host machine outside the cluster, use the host's IP address. In most deployments, expose Chaitin WAF as a reachable Service, IP address, or DNS name inside the cluster, or use the node/LoadBalancer address instead. gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: chaitin-waf: nodes: - host: "172.22.222.5" port: 8000 mode: monitor ``` ❶ Set `mode` to `monitor` in the plugin metadata. This applies to all `chaitin-waf` plugin instances if `mode` is not specified on a route. * Gateway API * APISIX CRD Create a route with `chaitin-waf` enabled without any plugin-level configuration, so it inherits the `monitor` mode from the plugin metadata: chaitin-waf-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: chaitin-waf-plugin-config spec: plugins: - name: chaitin-waf config: {} --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: chaitin-waf-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: chaitin-waf-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f gatewayproxy.yaml -f chaitin-waf-ic.yaml ``` To override the `monitor` mode and block malicious requests on the route, update the PluginConfig to set `mode: block`: chaitin-waf-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: chaitin-waf-plugin-config spec: plugins: - name: chaitin-waf config: mode: block ``` Apply the updated configuration to your cluster: ``` kubectl apply -f chaitin-waf-ic.yaml ``` Create a route with `chaitin-waf` enabled without any plugin-level configuration, so it inherits the `monitor` mode from the plugin metadata: chaitin-waf-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: chaitin-waf-route spec: ingressClassName: apisix http: - name: chaitin-waf-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: chaitin-waf enable: true config: {} ``` Apply the configuration to your cluster: ``` kubectl apply -f gatewayproxy.yaml -f chaitin-waf-ic.yaml ``` To override the `monitor` mode and block malicious requests on the route, update the ApisixRoute to set `mode: block`: chaitin-waf-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: chaitin-waf-route spec: ingressClassName: apisix http: - name: chaitin-waf-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: chaitin-waf enable: true config: mode: block ``` Apply the updated configuration to your cluster: ``` kubectl apply -f chaitin-waf-ic.yaml ``` Send a standard request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Send a request with SQL injection to the route: ``` curl -i "http://127.0.0.1:9080/anything" -d 'a=1 and 1=1' ``` You should also receive an `HTTP/1.1 200 OK` response as the request is not blocked in the `monitor` mode, but observe the following in the log entry: ``` 2025/09/09 11:44:08 [warn] 115#115: *31683 [lua] chaitin-waf.lua:385: do_access(): chaitin-waf monitor mode: request would have been rejected, event_id: 49bed20603e242f9be5ba6f1744bba4b, client: 172.20.0.1, server: _, request: "POST /anything HTTP/1.1", host: "127.0.0.1:9080" ``` If you explicitly configure the `mode` on a route, it will take precedence over the configuration in the plugin metadata. For instance, if you create a route like this: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "chaitin-waf-route", "uri": "/anything", "plugins": { "chaitin-waf": { "mode": "block" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` plugin_metadata: chaitin-waf: nodes: - host: "172.22.222.5" port: 8000 mode: monitor services: - name: chaitin-waf-service routes: - name: chaitin-waf-route uris: - /anything plugins: chaitin-waf: mode: block upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update the PluginConfig to set `mode: block`: chaitin-waf-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: chaitin-waf-plugin-config spec: plugins: - name: chaitin-waf config: mode: block ``` Apply the updated configuration to your cluster: ``` kubectl apply -f chaitin-waf-ic.yaml ``` Update the ApisixRoute to set `mode: block`: chaitin-waf-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: chaitin-waf-route spec: ingressClassName: apisix http: - name: chaitin-waf-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: chaitin-waf enable: true config: mode: block ``` Apply the updated configuration to your cluster: ``` kubectl apply -f chaitin-waf-ic.yaml ``` Send a standard request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Send a request with SQL injection to the route: ``` curl -i "http://127.0.0.1:9080/anything" -d 'a=1 and 1=1' ``` You should see an `HTTP/1.1 403 Forbidden` response similar to the following: ``` ... X-APISIX-CHAITIN-WAF-STATUS: 403 X-APISIX-CHAITIN-WAF-ACTION: reject X-APISIX-CHAITIN-WAF: yes X-APISIX-CHAITIN-WAF-TIME: 3 ... {"code": 403, "success":false, "message": "blocked by Chaitin SafeLine Web Application Firewall", "event_id": "c3eb25eaa7ae4c0d82eb8ceebf3600d0"} ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * mode string default: `block` vaild vaule: `off`, `monitor`, or `block` *** Mode to determine how the plugin behaves for matched requests. In `off` mode, WAF checks are skipped. In `monitor` mode, requests with potential threats are logged but not blocked. In `block` mode, requests with threats are blocked as determined by the WAF service. * match array\[object] *** An array of matching rules. The plugin uses these rules to decide whether to perform a WAF check on a request. If the list is empty, all requests are processed. * vars array\[array] *** An array of one or more matching conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md) to conditionally execute the plugin. * append\_waf\_resp\_header boolean default: `true` *** If true, add response headers `X-APISIX-CHAITIN-WAF`, `X-APISIX-CHAITIN-WAF-TIME`, `X-APISIX-CHAITIN-WAF-ACTION`, and `X-APISIX-CHAITIN-WAF-STATUS`. * append\_waf\_debug\_header boolean default: `false` *** If true, add debugging headers `X-APISIX-CHAITIN-WAF-ERROR` and `X-APISIX-CHAITIN-WAF-SERVER` to the response. Effective only when `append_waf_resp_header` is `true`. * config object *** Chaitin WAF service configurations. These settings override the corresponding metadata defaults when specified. * connect\_timeout integer default: `1000` *** The connection timeout to the WAF service, in milliseconds. * send\_timeout integer default: `1000` *** The sending timeout for transmitting data to the WAF service, in milliseconds. * read\_timeout integer default: `1000` *** The reading timeout for receiving data from the WAF service, in milliseconds. * req\_body\_size integer default: `1024` *** The maximum allowed request body size, in KB. * keepalive\_size integer default: `256` *** The maximum number of idle connections to the WAF detection service that can be maintained concurrently. * keepalive\_timeout integer default: `60000` *** The idle connection timeout for the WAF service, in milliseconds. * real\_client\_ip boolean default: `true` *** If true, use the client IP already resolved by the gateway, including any trusted-proxy or real-IP configuration. If false, use the direct peer address from the connection. The plugin does not read client-supplied forwarded headers directly. * log\_resp boolean *** If true, report the response to the WAF detection service after it has been delivered to the client, in addition to the request. The report is advisory and never blocks or modifies the response. Available in API7 Enterprise from version 3.9.20. * resp\_body\_size integer *** The maximum amount of the response body to report, in KB. Set to `0` to report only the response headers. Effective only when `log_resp` is true. Available in API7 Enterprise from version 3.9.20. * extra\_ignored\_content\_types string *** A comma-separated list of additional response content types to skip, on top of the built-in ignored list. A response whose content type matches is not reported to the WAF detection service at all, headers included. Effective only when `log_resp` is true. Available in API7 Enterprise from version 3.9.20. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * nodes array\[object] required *** An array of addresses for the Chaitin WAF service. * host string required *** Address of Chaitin WAF service. Supports IPv4, IPv6, Unix Socket, etc. * port integer default: `80` *** Port of Chaitin WAF service. * mode string default: `block` *** Mode to determine how the plugin behaves for matched requests. In `off` mode, WAF checks are skipped. In `monitor` mode, requests with potential threats are logged but not blocked. In `block` mode, requests with threats are blocked as determined by the WAF service. * config object *** Chaitin WAF service configurations. * connect\_timeout integer default: `1000` *** The connection timeout to the WAF service, in milliseconds. * send\_timeout integer default: `1000` *** The sending timeout for transmitting data to the WAF service, in milliseconds. * read\_timeout integer default: `1000` *** The reading timeout for receiving data from the WAF service, in milliseconds. * req\_body\_size integer default: `1024` *** The maximum allowed request body size, in KB. * keepalive\_size integer default: `256` *** The maximum number of idle connections to the WAF detection service that can be maintained concurrently. * keepalive\_timeout integer default: `60000` *** The idle connection timeout for the WAF service, in milliseconds. * real\_client\_ip boolean default: `true` *** If true, use the client IP already resolved by the gateway, including any trusted-proxy or real-IP configuration. If false, use the direct peer address from the connection. The plugin does not read client-supplied forwarded headers directly. * log\_resp boolean default: `false` *** If true, report the response to the WAF detection service after it has been delivered to the client, in addition to the request. The report is advisory and never blocks or modifies the response. Available in API7 Enterprise from version 3.9.20. * resp\_body\_size integer default: `4` *** The maximum amount of the response body to report, in KB. Set to `0` to report only the response headers. Effective only when `log_resp` is true. Available in API7 Enterprise from version 3.9.20. * extra\_ignored\_content\_types string *** A comma-separated list of additional response content types to skip, on top of the built-in ignored list. A response whose content type matches is not reported to the WAF detection service at all, headers included. Effective only when `log_resp` is true. Available in API7 Enterprise from version 3.9.20. --- # clickhouse-logger The `clickhouse-logger` plugin pushes request and response logs to ClickHouse database in batches and supports the customization of log formats. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `clickhouse-logger` plugin for different scenarios. To follow along the examples, start a sample ClickHouse server with user `default` and empty password: * Docker * Kubernetes ``` docker run -d -p 8123:8123 -p 9000:9000 -p 9009:9009 --name clickhouse-server clickhouse/clickhouse-server ``` Create a Kubernetes manifest file for the ClickHouse deployment: clickhouse-deployment.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: clickhouse-server spec: replicas: 1 selector: matchLabels: app: clickhouse-server template: metadata: labels: app: clickhouse-server spec: containers: - name: clickhouse-server image: clickhouse/clickhouse-server ports: - containerPort: 8123 - containerPort: 9000 - containerPort: 9009 ``` Create a Kubernetes manifest file for the ClickHouse service: clickhouse-service.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: clickhouse-server spec: selector: app: clickhouse-server ports: - name: http port: 8123 targetPort: 8123 - name: native port: 9000 targetPort: 9000 type: ClusterIP ``` Apply the manifests: ``` kubectl apply -f clickhouse-deployment.yaml -f clickhouse-service.yaml ``` ### Log in the Default Log Formats[​](#log-in-the-default-log-formats "Direct link to Log in the Default Log Formats") The following example demonstrates how you can log in the default request body. Create a table named `default_logs` in your ClickHouse database with columns corresponding to your log format: * Docker * Kubernetes ``` curl "http://127.0.0.1:8123" -X POST -d ' CREATE TABLE default.default_logs ( host String, client_ip String, route_id String, service_id String, start_time String, latency String, upstream_latency String, apisix_latency String, consumer String, request String, response String, server String, PRIMARY KEY(`start_time`) ) ENGINE = MergeTree() ' --user default: ``` ``` kubectl exec -n aic deploy/clickhouse-server -- clickhouse-client --query " CREATE TABLE default.default_logs ( host String, client_ip String, route_id String, service_id String, start_time String, latency String, upstream_latency String, apisix_latency String, consumer String, request String, response String, server String, PRIMARY KEY(start_time) ) ENGINE = MergeTree() " ``` Create a route with `clickhouse-logger` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "clickhouse-logger-route", "uri": "/get", "plugins": { "clickhouse-logger": { "user": "default", "password": "", "database": "default", "logtable": "default_logs", "endpoint_addrs": ["http://127.0.0.1:8123"] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: clickhouse-logger-route plugins: clickhouse-logger: user: default password: "" database: default logtable: default_logs endpoint_addrs: - "http://127.0.0.1:8123" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD clickhouse-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: clickhouse-logger-plugin-config spec: plugins: - name: clickhouse-logger config: user: default password: "" database: default logtable: default_logs endpoint_addrs: - "http://clickhouse-server.aic.svc:8123" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: clickhouse-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: clickhouse-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` clickhouse-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: clickhouse-logger-route spec: ingressClassName: apisix http: - name: clickhouse-logger-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: clickhouse-logger config: user: default password: "" database: default logtable: default_logs endpoint_addrs: - "http://clickhouse-server.aic.svc:8123" ``` Apply the configuration: ``` kubectl apply -f clickhouse-logger-ic.yaml ``` Send a request to the route to generate a log entry: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response. Send a request to ClickHouse to see the log entries: ``` echo 'SELECT * FROM default.default_logs FORMAT Pretty' | curl "http://127.0.0.1:8123/?" -d @- ``` You should see a log entry similar to the following: ``` ┏━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓ ┃ host ┃ client_ip ┃ route_id ┃ service_id ┃ start_time ┃ latency ┃ upstream_latency ┃ apisix_latency ┃ consumer ┃ request ┃ response ┃ server ┃ ┡━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩ │ │ 172.19.0.1 │ clickhouse-logger-route │ │ 1703026935235 │ 481.00018501282 │ 473 │ 8.0001850128174 │ │ {"method":"GET","uri":"/get","headers":{"host":"127.0.0.1:9080","user-agent":"curl/7.29.0","accept":"*/*"},"url":"http://127.0.0.1:9080/get","querystring":{},"size":81} │ {"headers":{"access-control-allow-credentials":"true","access-control-allow-origin":"*","content-type":"application/json","content-length":"299","date":"Tue,19 Dec 2023 23:02:15 GMT","connection":"close","server":"APISIX/3.8.0"},"status":200,"size":526} │ {"hostname":"85cf6f06914e","version":"3.8.0"} │ └──────┴────────────┴─────────────────────────┴────────────┴───────────────┴─────────────────┴──────────────────┴─────────────────┴──────────┴──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────────────────────────────────────┘ ``` ### Customize Log Format With Plugin Metadata[​](#customize-log-format-with-plugin-metadata "Direct link to Customize Log Format With Plugin Metadata") The following example demonstrates how you can customize log format using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md). Create a table named `custom_logs` in your ClickHouse database with columns corresponding to your customized log format: * Docker * Kubernetes ``` curl "http://127.0.0.1:8123" -X POST -d ' CREATE TABLE default.custom_logs ( host String, client_ip String, route_id String, service_id String, `@timestamp` String, PRIMARY KEY(`@timestamp`) ) ENGINE = MergeTree() ' --user default: ``` ``` kubectl exec -n aic deploy/clickhouse-server -- clickhouse-client --query " CREATE TABLE default.custom_logs ( host String, client_ip String, route_id String, service_id String, \`@timestamp\` String, PRIMARY KEY(\`@timestamp\`) ) ENGINE = MergeTree() " ``` Create a route with the `clickhouse-logger` plugin that is used to forward logs in the specified format to ClickHouse: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "clickhouse-logger-route", "uri": "/get", "plugins": { "clickhouse-logger": { "user": "default", "password": "", "database": "default", "logtable": "custom_logs", "endpoint_addrs": ["http://127.0.0.1:8123"] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: clickhouse-logger-route plugins: clickhouse-logger: user: default password: "" database: default logtable: custom_logs endpoint_addrs: - "http://127.0.0.1:8123" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD clickhouse-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: clickhouse-logger-plugin-config spec: plugins: - name: clickhouse-logger config: user: default password: "" database: default logtable: custom_logs endpoint_addrs: - "http://clickhouse-server.aic.svc:8123" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: clickhouse-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: clickhouse-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` clickhouse-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: clickhouse-logger-route spec: ingressClassName: apisix http: - name: clickhouse-logger-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: clickhouse-logger config: user: default password: "" database: default logtable: custom_logs endpoint_addrs: - "http://clickhouse-server.aic.svc:8123" ``` Apply the configuration: ``` kubectl apply -f clickhouse-logger-ic.yaml ``` Configure plugin metadata for `clickhouse-logger`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/clickhouse-logger" -X PUT \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "log_format": { "host": "$host", "client_ip": "$remote_addr", "route_id": "$route_id", "service_id": "$service_id", "@timestamp": "$time_iso8601" } }' ``` adc.yaml ``` plugin_metadata: - name: clickhouse-logger log_format: host: "$host" client_ip: "$remote_addr" route_id: "$route_id" service_id: "$service_id" "@timestamp": "$time_iso8601" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` clickhouse-logger-metadata.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: service: name: apisix-admin port: 9180 auth: type: AdminKey adminKey: value: edd1c9f034335f136f87ad84b625c8f1 pluginMetadata: clickhouse-logger: log_format: host: "$host" client_ip: "$remote_addr" route_id: "$route_id" service_id: "$service_id" "@timestamp": "$time_iso8601" ``` Apply the configuration: ``` kubectl apply -f clickhouse-logger-metadata.yaml ``` Send a request to the route to generate a log entry: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response. Send a request to ClickHouse to see the log entries: ``` echo 'SELECT * FROM default.custom_logs FORMAT Pretty' | curl "http://127.0.0.1:8123/?" -d @- ``` You should see a log entry similar to the following: ``` ┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━┓ ┃ host ┃ client_ip ┃ route_id ┃ service_id ┃ @timestamp ┃ ┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━┩ │ 127.0.0.1 │ 172.19.0.1 │ clickhouse-logger-route │ │ 2023-12-19T23:25:43+00:00 │ └───────────┴────────────┴─────────────────────────┴────────────┴───────────────────────────┘ ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * endpoint\_addr string *** Deprecated. Use `endpoint_addrs` instead. ClickHouse endpoint. Configure either `endpoint_addr` or `endpoint_addrs`. * endpoint\_addrs array *** ClickHouse endpoints. Configure either `endpoint_addrs` or the deprecated `endpoint_addr`. * database string required *** Name of the database to store the logs. * logtable string required *** Name of the table that stores the logs. * user string required *** ClickHouse username. From APISIX 3.16.0, supports referencing values from environment variables using the `$ENV://` prefix or from a secret manager using the `$secret://` prefix. For more information, see [secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md). * password string required *** ClickHouse password. The value is encrypted with AES before being stored in etcd. From APISIX 3.16.0, supports referencing values from environment variables using the `$ENV://` prefix or from a secret manager using the `$secret://` prefix. For more information, see [secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md). * timeout integer default: `3` vaild vaule: greater than 0 *** Time to keep the connection alive for after sending a request. * ssl\_verify boolean default: `true` *** If true, verify SSL. * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. You can also configure log format on a global scale using the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md), which configures the log format for all `clickhouse-logger` plugin instances. If the log format configured on the individual plugin instance differs from the log format configured on plugin metadata, the log format configured on the individual plugin instance takes precedence. See the [example](https://docs.api7.ai/hub/clickhouse-logger.md#customize-log-format-with-plugin-metadata) for more details. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * include\_req\_body boolean default: `false` *** If true, include the request body in the log. Note that if the request body is too big to be kept in the memory, it can not be logged due to NGINX's limitations. * include\_req\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_req_body` is true. Request body would only be logged when the expressions configured here evaluate to true. * include\_resp\_body boolean default: `false` *** If true, include the response body in the log. * include\_resp\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_resp_body` is true. Response body would only be logged when the expressions configured here evaluate to true. * max\_req\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes to include in the log. If the request body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * max\_resp\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum response body size in bytes to include in the log. If the response body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * name string default: `clickhouse-logger` *** Unique identifier of the plugin for the batch processor. If you use [Prometheus](https://docs.api7.ai/hub/prometheus.md) to monitor APISIX metrics, the name is exported in `apisix_batch_process_entries`. * batch\_max\_size integer default: `1000` vaild vaule: greater than 0 *** The number of log entries allowed in one batch. Once reached, the batch will be sent to the logging service. Setting this parameter to 1 means immediate processing. * inactive\_timeout integer default: `5` vaild vaule: greater than 0 *** The maximum time in seconds to wait for new logs before sending the batch to the logging service. The value should be smaller than `buffer_duration`. * buffer\_duration integer default: `60` vaild vaule: greater than 0 *** The maximum time in seconds from the earliest entry allowed before sending the batch to the logging service. * retry\_delay integer default: `1` vaild vaule: greater than or equal to 0 *** The time interval in seconds to retry sending the batch to the logging service if the batch was not successfully sent. * max\_retry\_count integer default: `0` vaild vaule: greater than or equal to 0 *** The maximum number of unsuccessful retries allowed before dropping the log entries. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * max\_pending\_entries integer default: `` `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 and 3.10.6; none in API7 Enterprise 3.9.18 and 3.10.5 `` vaild vaule: greater than or equal to 1 *** Maximum number of entries waiting in the batch processor. New entries are discarded when the backlog reaches the limit. Introduced in API7 Enterprise 3.8.17 and APISIX 3.15.0. The default changed to `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line. In API7 Enterprise 3.9.18 and 3.10.5, and in earlier APISIX versions, omitting the parameter leaves the backlog unlimited. See [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md#configure-the-pending-entry-limit) for sizing and verification guidance. --- # consumer-restriction The `consumer-restriction` plugin enables access controls based on consumer name, route ID, service ID, or consumer group ID. The plugin needs to work with authentication plugins, such as [`key-auth`](https://docs.api7.ai/hub/key-auth.md) and [`jwt-auth`](https://docs.api7.ai/hub/jwt-auth.md), which means you should always create at least one [consumer](https://docs.api7.ai/apisix/key-concepts/consumers.md) in your use case. See examples below for more details. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `consumer-restriction` plugin for different scenarios. While the examples use [`key-auth`](https://docs.api7.ai/hub/key-auth.md) as the authentication method, you can easily adjust to other authentication plugins based on your needs. ### Restrict Access by Consumers[​](#restrict-access-by-consumers "Direct link to Restrict Access by Consumers") The example below demonstrates how you can use the `consumer-restriction` plugin on a route to restrict consumer access by consumer names, where consumers are authenticated with [`key-auth`](https://docs.api7.ai/hub/key-auth.md). * Admin API * ADC * Ingress Controller Create a consumer `JohnDoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "JohnDoe" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/JohnDoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Create a second consumer `JaneDoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "JaneDoe" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/JaneDoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-key-auth", "plugins": { "key-auth": { "key": "jane-key" } } }' ``` Next, create a route with key authentication enabled, and configure `consumer-restriction` to allow only consumer `JaneDoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "consumer-restricted-route", "uri": "/get", "plugins": { "key-auth": {}, "consumer-restriction": { "whitelist": ["JaneDoe"] } }, "upstream" : { "nodes": { "httpbin.org":1 } } }' ``` adc.yaml ``` consumers: - username: JohnDoe credentials: - name: cred-john-key-auth type: key-auth config: key: john-key - username: JaneDoe credentials: - name: cred-jane-key-auth type: key-auth config: key: jane-key services: - name: consumer-restriction-service routes: - name: consumer-restricted-route uris: - /get plugins: key-auth: {} consumer-restriction: whitelist: - "JaneDoe" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Consumer Name Format in Ingress Controller When consumers are configured using the Ingress Controller, the consumer name is generated in the format `namespace_consumername`. For example, a consumer named `janedoe` in the `aic` namespace becomes `aic_janedoe`. Use this format in the `whitelist` or `blacklist` of `consumer-restriction`. * Gateway API * APISIX CRD consumer-restriction-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: johndoe spec: gatewayRef: name: apisix credentials: - type: key-auth name: john-key-auth config: key: john-key --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: janedoe spec: gatewayRef: name: apisix credentials: - type: key-auth name: jane-key-auth config: key: jane-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: consumer-restriction-plugin-config spec: plugins: - name: key-auth config: _meta: disable: false - name: consumer-restriction config: whitelist: - "aic_janedoe" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: consumer-restriction-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: consumer-restriction-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` consumer-restriction-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: johndoe spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: janedoe spec: ingressClassName: apisix authParameter: keyAuth: value: key: jane-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: consumer-restriction-route spec: ingressClassName: apisix http: - name: consumer-restriction-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: key-auth enable: true - name: consumer-restriction enable: true config: whitelist: - "aic_janedoe" ``` Apply the configuration to your cluster: ``` kubectl apply -f consumer-restriction-ic.yaml ``` Send a request to the route as consumer `JohnDoe`: ``` curl -i "http://127.0.0.1:9080/get" -H 'apikey: john-key' ``` You should receive an `HTTP/1.1 403 Forbidden` response with the following message: ``` {"message":"The consumer_name is forbidden."} ``` Send another request to the route as consumer `JaneDoe`: ``` curl -i "http://127.0.0.1:9080/get" -H 'apikey: jane-key' ``` You should receive an `HTTP/1.1 200 OK` response, showing the consumer access is permitted. ### Restrict Access by Consumers and HTTP Methods[​](#restrict-access-by-consumers-and-http-methods "Direct link to Restrict Access by Consumers and HTTP Methods") The example below demonstrates how you can use the `consumer-restriction` plugin on a route to restrict consumer access by consumer name and HTTP methods, where consumers are authenticated with [`key-auth`](https://docs.api7.ai/hub/key-auth.md). * Admin API * ADC * Ingress Controller Create a consumer `JohnDoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "JohnDoe" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/JohnDoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Create a second consumer `JaneDoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "JaneDoe" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/JaneDoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-key-auth", "plugins": { "key-auth": { "key": "jane-key" } } }' ``` Next, create a route with key authentication enabled, and use `consumer-restriction` to allow only the configured HTTP methods by consumers: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "consumer-restricted-route", "uri": "/anything", "plugins": { "key-auth": {}, "consumer-restriction": { "allowed_by_methods":[ { "user": "JohnDoe", "methods": ["GET"] }, { "user": "JaneDoe", "methods": ["POST"] } ] } }, "upstream" : { "nodes": { "httpbin.org":1 } } }' ``` adc.yaml ``` consumers: - username: JohnDoe credentials: - name: cred-john-key-auth type: key-auth config: key: john-key - username: JaneDoe credentials: - name: cred-jane-key-auth type: key-auth config: key: jane-key services: - name: consumer-restriction-service routes: - name: consumer-restricted-route uris: - /anything plugins: key-auth: {} consumer-restriction: allowed_by_methods: - user: "JohnDoe" methods: - "GET" - user: "JaneDoe" methods: - "POST" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD consumer-restriction-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: johndoe spec: gatewayRef: name: apisix credentials: - type: key-auth name: john-key-auth config: key: john-key --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: janedoe spec: gatewayRef: name: apisix credentials: - type: key-auth name: jane-key-auth config: key: jane-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: consumer-restriction-methods-config spec: plugins: - name: key-auth config: _meta: disable: false - name: consumer-restriction config: allowed_by_methods: - user: "aic_johndoe" methods: - "GET" - user: "aic_janedoe" methods: - "POST" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: consumer-restriction-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: consumer-restriction-methods-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f consumer-restriction-ic.yaml ``` consumer-restriction-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: johndoe spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: janedoe spec: ingressClassName: apisix authParameter: keyAuth: value: key: jane-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: consumer-restriction-route spec: ingressClassName: apisix http: - name: consumer-restriction-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: key-auth enable: true - name: consumer-restriction enable: true config: allowed_by_methods: - user: "aic_johndoe" methods: - "GET" - user: "aic_janedoe" methods: - "POST" ``` Apply the configuration to your cluster: ``` kubectl apply -f consumer-restriction-ic.yaml ``` Send a POST request to the route as consumer `JohnDoe`: ``` curl -i "http://127.0.0.1:9080/anything" -X POST -H 'apikey: john-key' ``` You should receive an `HTTP/1.1 403 Forbidden` response with the following message: ``` {"message":"The consumer_name is forbidden."} ``` Now, send a GET request to the route as consumer `JohnDoe`: ``` curl -i "http://127.0.0.1:9080/anything" -X GET -H 'apikey: john-key' ``` You should receive an `HTTP/1.1 200 OK` response, showing the consumer access is permitted. You can also verify the configurations by sending requests as consumer `JaneDoe` and observe the behaviours match up to what was configured in the `consumer-restriction` plugin on the route. ### Restricting by Service ID[​](#restricting-by-service-id "Direct link to Restricting by Service ID") The example below demonstrates how you can use the `consumer-restriction` plugin to restrict consumer access by service ID, where the consumer is authenticated with [`key-auth`](https://docs.api7.ai/hub/key-auth.md). * Admin API * ADC * Ingress Controller Create two sample services: ``` curl "http://127.0.0.1:9180/apisix/admin/services" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "srv-1", "upstream": { "type": "roundrobin", "nodes": { "httpbin.org":1 } } }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/services" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "srv-2", "upstream": { "type": "roundrobin", "nodes": { "mock.api7.ai":1 } } }' ``` Next, create a consumer with `key-auth` and configure `consumer-restriction` to allow only `srv-1` service: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "JohnDoe", "plugins": { "key-auth": { "key": "john-key" }, "consumer-restriction": { "type": "service_id", "whitelist": ["srv-1"] } } }' ``` Finally, create two routes, with each belonging to one of the services created earlier: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "srv-1-route", "uri": "/anything", "service_id": "srv-1" }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "srv-2-route", "uri": "/srv-2", "service_id": "srv-2" }' ``` adc.yaml ``` consumers: - username: JohnDoe plugins: key-auth: key: john-key consumer-restriction: type: service_id whitelist: - "srv-1" services: - name: srv-1 routes: - name: srv-1-route uris: - /anything upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 - name: srv-2 routes: - name: srv-2-route uris: - /srv-2 upstream: type: roundrobin nodes: - host: mock.api7.ai port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Service ID Format in Ingress Controller When routes are configured using the Ingress Controller, APISIX service IDs are auto-generated as the hash of `{namespace}_{routeName}_{ruleIndex}`. These IDs cannot be easily predetermined. Consider using consumer name-based restriction. Send a request to the route in the `srv-1` service: ``` curl -i "http://127.0.0.1:9080/anything" -H 'apikey: john-key' ``` You should receive an `HTTP/1.1 200 OK` response, showing the consumer access is permitted. Send a request to the route in the `srv-2` service: ``` curl -i "http://127.0.0.1:9080/srv-2" -H 'apikey: john-key' ``` You should receive an `HTTP/1.1 401 Unauthorized` response with the following message: ``` {"message":"The request is rejected, please check the service_id for this request"} ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * type string default: `consumer_name` vaild vaule: `consumer_name`, `consumer_group_id`, `service_id`, or `route_id` *** Key type to restrict by. * whitelist array\[string] *** List of objects to whitelist. At least one of `whitelist`, `blacklist`, and `allowed_by_methods` should be configured. If all are configured, the precedence is `blacklist` > `whitelist` > `allowed_by_methods`. * blacklist array\[string] *** List of objects to blacklist. At least one of `whitelist`, `blacklist`, and `allowed_by_methods` should be configured. If all are configured, the precedence is `blacklist` > `whitelist` > `allowed_by_methods`. * allowed\_by\_methods array\[object] *** List of key-value pairs of consumer name and their corresponding HTTP methods allowed. At least one of `whitelist`, `blacklist`, and `allowed_by_methods` should be configured. If all are configured, the precedence is `blacklist` > `whitelist` > `allowed_by_methods`. * user string *** Consumer username. * methods array\[string] vaild vaule: Any combination of the `GET`, `POST`, `PUT`, `DELETE`, `PATCH`, `HEAD`, `OPTIONS`, `CONNECT`, `TRACE`, `PURGE` methods *** List of allowed HTTP methods for the consumer. * rejected\_code integer default: `403` vaild vaule: greater than or equal to 200 *** HTTP status code to return when the request is rejected. * rejected\_msg string *** Error message to return when the request is rejected. --- # cors The `cors` plugin allows you to enable [Cross-Origin Resource Sharing (CORS)](https://developer.mozilla.org/en-US/docs/Web/HTTP/CORS). CORS is an HTTP-header based mechanism which allows a server to specify any origins (domain, scheme, or port) other than its own, and instructs browsers to allow the loading of resources from those origins. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure routes using the `cors` plugin for different scenarios. ### Enable CORS for a Route[​](#enable-cors-for-a-route "Direct link to Enable CORS for a Route") The following example demonstrates how to enable CORS on a route to allow resource loading from a list of origins. * Admin API * ADC * Ingress Controller Create a route with the `cors` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cors-route", "uri": "/anything", "plugins": { "cors": { "allow_origins": "http://sub.domain.com,http://sub2.domain.com", "allow_methods": "GET,POST", "allow_headers": "headr1,headr2", "expose_headers": "ex-headr1,ex-headr2", "max_age": 50, "allow_credential": true } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: cors-service routes: - name: cors-route uris: - /anything plugins: cors: allow_origins: "http://sub.domain.com,http://sub2.domain.com" allow_methods: "GET,POST" allow_headers: "headr1,headr2" expose_headers: "ex-headr1,ex-headr2" max_age: 50 allow_credential: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD cors-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: cors-plugin-config spec: plugins: - name: cors config: allow_origins: "http://sub.domain.com,http://sub2.domain.com" allow_methods: "GET,POST" allow_headers: "headr1,headr2" expose_headers: "ex-headr1,ex-headr2" max_age: 50 allow_credential: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: cors-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: cors-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f cors-ic.yaml ``` cors-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: cors-route spec: ingressClassName: apisix http: - name: cors-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: cors enable: true config: allow_origins: "http://sub.domain.com,http://sub2.domain.com" allow_methods: "GET,POST" allow_headers: "headr1,headr2" expose_headers: "ex-headr1,ex-headr2" max_age: 50 allow_credential: true ``` Apply the configuration to your cluster: ``` kubectl apply -f cors-ic.yaml ``` ❶ `allow_origins`: configure allowed origins, comma-separated. To allow all origins, set it to `*`. ❷ `max_age`: configure the maximum time the result is cached in seconds. ❸ `allow_credential`: set to `true` to allow credentials (cookies, HTTP authentication, and client-side SSL certificates) to be sent with the request. If you set this to true, you cannot use `*` for other cors attributes. Send a head request to the route with an allowed origin: ``` curl "http://127.0.0.1:9080/anything" -H "Origin: http://sub2.domain.com" -I ``` You should receive an `HTTP/1.1 200 OK` response and observe CORS headers: ``` ... Access-Control-Allow-Origin: http://sub2.domain.com Access-Control-Allow-Credentials: true Server: APISIX/3.8.0 Vary: Origin Access-Control-Allow-Methods: GET,POST Access-Control-Max-Age: 50 Access-Control-Expose-Headers: ex-headr1,ex-headr2 Access-Control-Allow-Headers: headr1,headr2 ``` Send a head request to the route with an origin that is not allowed: ``` curl "http://127.0.0.1:9080/anything" -H "Origin: http://sub3.domain.com" -I ``` You should receive an `HTTP/1.1 200 OK` response without any CORS header: ``` ... Server: APISIX/3.8.0 Vary: Origin ``` ### Use RegEx to Match Origin[​](#use-regex-to-match-origin "Direct link to Use RegEx to Match Origin") The following example demonstrates how to use RegEx to match the origin in `allow_origins` using the `allow_origins_by_regex` field. * Admin API * ADC * Ingress Controller Create a route with the `cors` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cors-route", "uri": "/anything", "plugins": { "cors": { "allow_methods": "GET,POST", "allow_headers": "headr1,headr2", "expose_headers": "ex-headr1,ex-headr2", "max_age": 50, "allow_origins_by_regex": [ ".*\\.test.com$" ] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: cors-service routes: - name: cors-route uris: - /anything plugins: cors: allow_methods: "GET,POST" allow_headers: "headr1,headr2" expose_headers: "ex-headr1,ex-headr2" max_age: 50 allow_origins_by_regex: - ".*\\.test.com$" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD cors-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: cors-regex-plugin-config spec: plugins: - name: cors config: allow_methods: "GET,POST" allow_headers: "headr1,headr2" expose_headers: "ex-headr1,ex-headr2" max_age: 50 allow_origins_by_regex: - ".*\\.test.com$" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: cors-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: cors-regex-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f cors-ic.yaml ``` cors-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: cors-route spec: ingressClassName: apisix http: - name: cors-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: cors enable: true config: allow_methods: "GET,POST" allow_headers: "headr1,headr2" expose_headers: "ex-headr1,ex-headr2" max_age: 50 allow_origins_by_regex: - ".*\\.test.com$" ``` Apply the configuration to your cluster: ``` kubectl apply -f cors-ic.yaml ``` ❶ `allow_origins_by_regex`: allow origins using RegEx. If used together with `allow_origins`, then `allow_origins` will be ignored. Send a head request to the route with an allowed origin: ``` curl "http://127.0.0.1:9080/anything" -H "Origin: http://a.test.com" -I ``` You should receive an `HTTP/1.1 200 OK` response and observe CORS headers: ``` ... Access-Control-Allow-Origin: http://a.test.com Access-Control-Allow-Credentials: true Server: APISIX/3.8.0 Access-Control-Allow-Methods: GET,POST Access-Control-Max-Age: 50 Access-Control-Expose-Headers: ex-headr1,ex-headr2 Access-Control-Allow-Headers: headr1,headr2 ``` You can also try to make a request with an invalid origin: ``` curl "http://127.0.0.1:9080/anything" -H "Origin: http://a.test2.com" -I ``` You should receive an `HTTP/1.1 200 OK` response without any CORS header: ``` ... Server: APISIX/3.8.0 Vary: Origin ``` ### Configure Origins in Plugin Metadata[​](#configure-origins-in-plugin-metadata "Direct link to Configure Origins in Plugin Metadata") The following example demonstrates how to configure origins in [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) and reference them as the allowed origins in the `cors` plugin. * Admin API * ADC * Ingress Controller Configure plugin metadata for the `cors` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/cors" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "allow_origins": { "key_1": "https://domain.com", "key_2": "https://sub.domain.com,https://sub2.domain.com", "key_3": "*" } }' ``` ❶ `allow_origins` : a map of keys and allowed origins. The key will be used to match the origin in the route. Create a route with the `cors` plugin using `allow_origins_by_metadata`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cors-route", "uri": "/anything", "plugins": { "cors": { "allow_methods": "GET,POST", "allow_headers": "headr1,headr2", "expose_headers": "ex-headr1,ex-headr2", "max_age": 50, "allow_origins_by_metadata": ["key_1"] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` ❶ `allow_origins_by_metadata`: keys in the metadata to match the origin. adc.yaml ``` plugin_metadata: cors: allow_origins: key_1: "https://domain.com" key_2: "https://sub.domain.com,https://sub2.domain.com" key_3: "*" services: - name: cors-service routes: - name: cors-route uris: - /anything plugins: cors: allow_methods: "GET,POST" allow_headers: "headr1,headr2" expose_headers: "ex-headr1,ex-headr2" max_age: 50 allow_origins_by_metadata: - "key_1" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Update your GatewayProxy manifest to configure the plugin metadata: gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: cors: allow_origins: key_1: "https://domain.com" key_2: "https://sub.domain.com,https://sub2.domain.com" key_3: "*" ``` ❶ `allow_origins`: a map of keys and allowed origins. The key will be used to match the origin in the route. * Gateway API * APISIX CRD Create the route with `allow_origins_by_metadata`: cors-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: cors-metadata-plugin-config spec: plugins: - name: cors config: allow_methods: "GET,POST" allow_headers: "headr1,headr2" expose_headers: "ex-headr1,ex-headr2" max_age: 50 allow_origins_by_metadata: - "key_1" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: cors-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: cors-metadata-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Create the route with `allow_origins_by_metadata`: cors-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: cors-route spec: ingressClassName: apisix http: - name: cors-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: cors enable: true config: allow_methods: "GET,POST" allow_headers: "headr1,headr2" expose_headers: "ex-headr1,ex-headr2" max_age: 50 allow_origins_by_metadata: - "key_1" ``` ❶ `allow_origins_by_metadata`: keys in the metadata to match the origin. Apply the configuration to your cluster: ``` kubectl apply -f gatewayproxy.yaml -f cors-ic.yaml ``` Send a head request to the route with an allowed origin: ``` curl "http://127.0.0.1:9080/anything" -H "Origin: https://domain.com" -I ``` You should receive an `HTTP/1.1 200 OK` response and observe CORS headers: ``` ... Access-Control-Allow-Origin: https://domain.com Access-Control-Allow-Credentials: true Server: APISIX/3.8.0 Access-Control-Allow-Methods: GET,POST Access-Control-Max-Age: 50 Access-Control-Expose-Headers: ex-headr1,ex-headr2 Access-Control-Allow-Headers: headr1,headr2 ``` Send another request with an invalid origin: ``` curl "http://127.0.0.1:9080/anything" -H "Origin: http://a.test2.com" -I ``` You should receive an `HTTP/1.1 200 OK` response without any CORS header: ``` ... Server: APISIX/3.8.0 Vary: Origin ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * allow\_origins string default: `*` *** Comma-separated string of origins to allow CORS. If `allow_credential` is set to `true`, you can forcefully allow CORS on all origins by configuring the field to `**` but sensitive data, such as authentication tokens or cookies, can get exposed to any malicious website. You can also configure allow origins on a global scale using the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md), which configures the allow origins for all `cors` plugin instances. See the [example](https://docs.api7.ai/hub/cors.md#use-metadata-to-match-origin) for more details. * allow\_methods string default: `*` *** Comma-separated string of HTTP request methods to allow CORS. If `allow_credential` is set to `true`, you can forcefully allow CORS on all methods by configuring the field to `**`, but a malicious actor can use HTTP methods, such as `PUT` or `DELETE`, to make unexpected modifications to shared resource and pose a security threat. * allow\_headers string default: `*` *** Comma-separated string of HTTP headers allowed in requests. If `allow_credential` is set to `true`, you can forcefully allow CORS on all request headers by configuring the field to `**`, but it can potentially allow malicious headers to be sent to the server. * expose\_headers string *** Comma-separated string of HTTP headers that should be made available in response to a cross-origin request. * max\_age integer default: `5` *** Maximum time in seconds for which the results of a [preflight request](https://developer.mozilla.org/en-US/docs/Glossary/Preflight_request) can be cached. If the time is within this limit, the browser will check the cached result. To disable caching, set `max_age` to `-1`. Note that the maximum value allowed is browser-dependent. See [`Access-Control-Max-Age`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Access-Control-Max-Age#Directives) for more details. * allow\_credential boolean *** If true, allow requests to include credentials, such as cookies. According to CORS specification, when `allow_credential` is set to true, you cannot use `*` for other CORS attributes. To allow all origins, set the field to `**`. This can potentially allow sensitive user data, such as authentication tokens or cookies, to be exposed to malicious actors. * allow\_origins\_by\_regex array\[string] *** RegEx to match origins that allow CORS. When configured, only domains in this range will be allowed and any configuration in `allow_origins` will be ignored. For example, `['.*\.test.com$']` can match all subdomains of `test.com`. * allow\_origins\_by\_metadata array\[string] *** Origins to enable CORS referenced from `allow_origins` set in the plugin metadata. For example, if `allow_origins: {'EXAMPLE': 'https://example.com'}` is set in the plugin metadata, then `['EXAMPLE']` can be used to allow CORS on the origin `https://example.com`. * timing\_allow\_origins string *** Comma-separated string of origins to allow to access the resource timing information. See [`Timing-Allow-Origin`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Timing-Allow-Origin) for more details. * timing\_allow\_origins\_by\_regex array\[string] *** RegEx to match with origin for enabling access to the resource timing information. When configured, only domains matching the RegEx will be allowed and any configuration in `timing_allow_origins` will be ignored. For example, `['.*\.test.com']` can match all subdomain of `test.com`. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * allow\_origins object *** A map of named origins to allow for CORS, where each key is an identifier referenced by `allow_origins_by_metadata` and each value is the corresponding origin string. For example, `{'EXAMPLE': 'https://example.com'}` defines the key `EXAMPLE` for the origin `https://example.com`. If `allow_credential` is set to `true`, you can forcefully allow CORS on all origins by setting a map value to `**`, but sensitive data, such as authentication tokens or cookies, can get exposed to any malicious website. --- # data-mask The `data-mask` plugin masks sensitive information in request headers, bodies, and URL queries when using logging plugins. Note that it does not modify the actual request or response traffic. To mask sensitive information in the gateway's access log, see [Mask Sensitive Data in Access Log](https://docs.api7.ai/api7-gateway/how-to-guides/api-security/data-masking.md). About Plugin Execution Order The plugin can be configured on routes, services, or as a global plugin. However, be aware that [global plugins are always executed before route- or service-level plugins](https://docs.api7.ai/apisix/key-concepts/plugins.md#plugins-execution-order), so data masking may occur after logging. For instance, if a logging plugin is configured globally while `data-mask` is applied at the route level, requests will be logged before masking occurs, and sensitive data will appear in plaintext. To ensure the intended behavior, it is recommended to configure both plugins at the same level: 1. Both at the global level (recommended if suitable for your use case) 2. Both at the route or service level ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can use the `data-mask` plugin for different scenarios. While all examples use the `file-logger` plugin for logging, the plugin is used only to demonstrate the results of data masking. Select the logging plugin that best suits your environment. ### Mask Sensitive Information in URL Query[​](#mask-sensitive-information-in-url-query "Direct link to Mask Sensitive Information in URL Query") The following example demonstrates how you can mask sensitive information in the request URL queries, before the request is logged to a local file by the `file-logger` plugin. Create a route with the `file-logger` plugin to log requests and the `data-mask` plugin with three data masking rules: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "data-mask-route", "uri": "/anything", "plugins": { "data-mask": { "request": [ { "action": "remove", "name": "password", "type": "query" }, { "action": "replace", "name": "token", "type": "query", "value": "*****" }, { "action": "regex", "name": "card", "regex": "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)", "type": "query", "value": "$1-****-****-$2" } ] }, "file-logger": { "path": "/tmp/mask-query.log" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: data-mask-service routes: - name: data-mask-route uris: - /anything plugins: data-mask: request: - action: remove name: password type: query - action: replace name: token type: query value: "*****" - action: regex name: card regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: query value: "$1-****-****-$2" file-logger: path: /tmp/mask-query.log upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD data-mask-query-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: data-mask-query-plugin-config spec: plugins: - name: data-mask config: request: - action: remove name: password type: query - action: replace name: token type: query value: "*****" - action: regex name: card regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: query value: "$1-****-****-$2" - name: file-logger config: path: /tmp/mask-query.log --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: data-mask-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: data-mask-query-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f data-mask-query-ic.yaml ``` data-mask-query-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: data-mask-route spec: ingressClassName: apisix http: - name: data-mask-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: data-mask enable: true config: request: - action: remove name: password type: query - action: replace name: token type: query value: "*****" - action: regex name: card regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: query value: "$1-****-****-$2" - name: file-logger enable: true config: path: /tmp/mask-query.log ``` Apply the configuration to your cluster: ``` kubectl apply -f data-mask-query-ic.yaml ``` ❶ Configure the data masking rule to remove `password` URL query from the request. ❷ Configure the data masking rule to replace the value of `token` URL query with `*****`. ❸ Configure the data masking rule that matches card number in the URL query with RegEx and mask the middle portion of the card number. ❹ path to the log file on the filesystem where logs should be saved. Send a request to the route with sensitive information in URL queries: ``` curl -i "http://127.0.0.1:9080/anything?password=abc&token=xyz&card=1234-1234-1234-1234" ``` You should receive an `HTTP/1.1 200 OK` response. Navigating to the `/tmp/mask-query.log` file and examining the log content, you should see a log entry similar to the following: ``` { "request": { "uri": "/anything?token=*****&card=1234-****-****-1234", "method": "GET", "url": "http://127.0.0.1:9080/anything?token=*****&card=1234-****-****-1234", "querystring": { "token": "*****", "card": "1234-****-****-1234" } } } ``` ### Mask Sensitive Information in Request Headers[​](#mask-sensitive-information-in-request-headers "Direct link to Mask Sensitive Information in Request Headers") The following example demonstrates how you can mask sensitive information in request headers, before the request is logged to a local file by the `file-logger` plugin. Create a route with the `file-logger` plugin to log requests and the `data-mask` plugin with three data masking rules: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "data-mask-route", "uri": "/anything", "plugins": { "data-mask": { "request": [ { "action": "remove", "name": "password", "type": "header" }, { "action": "replace", "name": "token", "type": "header", "value": "*****" }, { "action": "regex", "name": "card", "regex": "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)", "type": "header", "value": "$1-****-****-$2" } ] }, "file-logger": { "path": "/tmp/mask-header.log" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: data-mask-service routes: - name: data-mask-route uris: - /anything plugins: data-mask: request: - action: remove name: password type: header - action: replace name: token type: header value: "*****" - action: regex name: card regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: header value: "$1-****-****-$2" file-logger: path: /tmp/mask-header.log upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD data-mask-header-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: data-mask-header-plugin-config spec: plugins: - name: data-mask config: request: - action: remove name: password type: header - action: replace name: token type: header value: "*****" - action: regex name: card regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: header value: "$1-****-****-$2" - name: file-logger config: path: /tmp/mask-header.log --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: data-mask-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: data-mask-header-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f data-mask-header-ic.yaml ``` data-mask-header-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: data-mask-route spec: ingressClassName: apisix http: - name: data-mask-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: data-mask enable: true config: request: - action: remove name: password type: header - action: replace name: token type: header value: "*****" - action: regex name: card regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: header value: "$1-****-****-$2" - name: file-logger enable: true config: path: /tmp/mask-header.log ``` Apply the configuration to your cluster: ``` kubectl apply -f data-mask-header-ic.yaml ``` ❶ Configure the data masking rule to remove `password` header from the request. ❷ Configure the data masking rule to replace the value of `token` request header with `*****`. ❸ Configure the data masking rule that matches card number in the request header with RegEx and mask the middle portion of the card number. ❹ path to the log file on the filesystem where logs should be saved. Send a POST request to the route with sensitive information in headers: ``` curl -i "http://127.0.0.1:9080/anything" -X POST \ -H "password: abc" \ -H "token: xyz" \ -H "card: 1234-1234-1234-1234" ``` You should receive an `HTTP/1.1 200 OK` response. Navigating to the `/tmp/mask-header.log` file and examining the log content, you should see a log entry similar to the following: ``` { "request": { "uri": "/anything", "method": "GET", "url": "http://127.0.0.1:9080/anything", "headers": { "user-agent": "curl/8.6.0", "token": "*****", "card": "1234-****-****-1234" } } } ``` ### Mask Sensitive Information in URL-Encoded Request Bodies[​](#mask-sensitive-information-in-url-encoded-request-bodies "Direct link to Mask Sensitive Information in URL-Encoded Request Bodies") The following example demonstrates how you can mask sensitive information in URL-encoded request bodies, before the request is logged to a local file by the `file-logger` plugin. Create a route with the `file-logger` plugin to log requests and the `data-mask` plugin with three data masking rules: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "data-mask-route", "uri": "/anything", "plugins": { "data-mask": { "request": [ { "action": "remove", "body_format": "urlencoded", "name": "password", "type": "body" }, { "action": "replace", "body_format": "urlencoded", "name": "token", "type": "body", "value": "*****" }, { "action": "regex", "body_format": "urlencoded", "name": "card", "regex": "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)", "type": "body", "value": "$1-****-****-$2" } ] }, "file-logger": { "include_req_body": true, "path": "/tmp/mask-urlencoded-body.log" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: data-mask-service routes: - name: data-mask-route uris: - /anything plugins: data-mask: request: - action: remove body_format: urlencoded name: password type: body - action: replace body_format: urlencoded name: token type: body value: "*****" - action: regex body_format: urlencoded name: card regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: body value: "$1-****-****-$2" file-logger: include_req_body: true path: /tmp/mask-urlencoded-body.log upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD data-mask-urlencoded-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: data-mask-urlencoded-plugin-config spec: plugins: - name: data-mask config: request: - action: remove body_format: urlencoded name: password type: body - action: replace body_format: urlencoded name: token type: body value: "*****" - action: regex body_format: urlencoded name: card regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: body value: "$1-****-****-$2" - name: file-logger config: include_req_body: true path: /tmp/mask-urlencoded-body.log --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: data-mask-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: data-mask-urlencoded-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f data-mask-urlencoded-ic.yaml ``` data-mask-urlencoded-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: data-mask-route spec: ingressClassName: apisix http: - name: data-mask-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: data-mask enable: true config: request: - action: remove body_format: urlencoded name: password type: body - action: replace body_format: urlencoded name: token type: body value: "*****" - action: regex body_format: urlencoded name: card regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: body value: "$1-****-****-$2" - name: file-logger enable: true config: include_req_body: true path: /tmp/mask-urlencoded-body.log ``` Apply the configuration to your cluster: ``` kubectl apply -f data-mask-urlencoded-ic.yaml ``` ❶ Configure the data masking rule to remove `password` information from the request body. ❷ Configure the data masking rule to replace `token` information in the request body with `*****`. ❸ Configure the data masking rule that matches card number in the request body with RegEx and mask the middle portion of the card number. ❹ Include the request body in the log. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" \ --data-urlencode "password=abc" \ --data-urlencode "token=xyz" \ --data-urlencode "card=1234-1234-1234-1234" ``` You should receive an `HTTP/1.1 200 OK` response. Navigating to the `/tmp/mask-urlencoded-body.log` file and examining the log content, you should see a log entry similar to the following: ``` { "request": { "uri": "/anything", "body": "token=*****&card=1234-****-****-1234", "method": "POST", "url": "http://127.0.0.1:9080/anything" } } ``` ### Mask Sensitive Information in JSON-Encoded Request Bodies[​](#mask-sensitive-information-in-json-encoded-request-bodies "Direct link to Mask Sensitive Information in JSON-Encoded Request Bodies") The following example demonstrates how you can mask sensitive information in JSON-encoded request bodies using [JSON path](https://goessner.net/articles/JsonPath) syntax in the plugin to look for the target field, before the request is logged to a local file by the `file-logger` plugin. Create a route with the `file-logger` plugin to log requests and the `data-mask` plugin with three data masking rules: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "data-mask-route", "uri": "/anything", "plugins": { "data-mask": { "request": [ { "action": "remove", "body_format": "json", "name": "$.password", "type": "body" }, { "action": "replace", "body_format": "json", "name": "users[*].token", "type": "body", "value": "*****" }, { "action": "regex", "body_format": "json", "name": "$.users[*].credit.card", "regex": "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)", "type": "body", "value": "$1-****-****-$2" } ] }, "file-logger": { "include_req_body": true, "path": "/tmp/mask-json-body.log" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: data-mask-service routes: - name: data-mask-route uris: - /anything plugins: data-mask: request: - action: remove body_format: json name: "$.password" type: body - action: replace body_format: json name: "users[*].token" type: body value: "*****" - action: regex body_format: json name: "$.users[*].credit.card" regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: body value: "$1-****-****-$2" file-logger: include_req_body: true path: /tmp/mask-json-body.log upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD data-mask-json-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: data-mask-json-plugin-config spec: plugins: - name: data-mask config: request: - action: remove body_format: json name: "$.password" type: body - action: replace body_format: json name: "users[*].token" type: body value: "*****" - action: regex body_format: json name: "$.users[*].credit.card" regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: body value: "$1-****-****-$2" - name: file-logger config: include_req_body: true path: /tmp/mask-json-body.log --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: data-mask-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: data-mask-json-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f data-mask-json-ic.yaml ``` data-mask-json-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: data-mask-route spec: ingressClassName: apisix http: - name: data-mask-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: data-mask enable: true config: request: - action: remove body_format: json name: "$.password" type: body - action: replace body_format: json name: "users[*].token" type: body value: "*****" - action: regex body_format: json name: "$.users[*].credit.card" regex: "(\\d+)\\-\\d+\\-\\d+\\-(\\d+)" type: body value: "$1-****-****-$2" - name: file-logger enable: true config: include_req_body: true path: /tmp/mask-json-body.log ``` Apply the configuration to your cluster: ``` kubectl apply -f data-mask-json-ic.yaml ``` ❶ Configure the data masking rule to remove `password` information from the request body. ❷ Configure the data masking rule to replace `token` information in the request body with `*****`. ❸ Configure the data masking rule that matches card number in the request body with RegEx and mask the middle portion of the card number. ❹ Include the request body in the log. Send a request to the route with sensitive information in the request body: ``` curl -i "http://127.0.0.1:9080/anything" -X POST -d ' { "password": "abc", "users": [ { "token": "xyz", "credit": { "card": "1234-1234-1234-1234" } }, { "token": "xyz", "credit": { "card": "1234-1234-1234-1234" } } ] }' ``` You should receive an `HTTP/1.1 200 OK` response. Navigating to the `/tmp/mask-json-body.log` file and examining the log content, you should see a log entry similar to the following: ``` { "request": { "uri": "/anything", "body": "{\"users\":[{\"token\":\"*****\",\"credit\":{\"card\":\"1234-****-****-1234\"}},{\"token\":\"*****\",\"credit\":{\"card\":\"1234-****-****-1234\"}}]}", "method": "POST", "url": "http://127.0.0.1:9080/anything" } } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * request array\[object] *** An array of actions to mask sensitive information in the request. * type string required vaild vaule: `query`, `header`, or `body` *** Location where sensitive information should be masked. * body\_format string vaild vaule: `json` or `urlencoded` *** Encoding of the request body. Required when `type` is `body`. * name string required *** Name of the information field that contains sensitive data. For JSON body, you can use [JSONPath](https://goessner.net/articles/JsonPath) syntax. * action string required vaild vaule: `regex`, `replace`, or `remove` *** Action to mask the sensitive data. * regex string *** Regular expressions to match the sensitive data. Required when `action` is `regex`. * value string *** Value to replace the sensitive data with. Required when `action` is `regex` or `replace`. * max\_body\_size integer default: `1048576` *** Maximum body size allowed in bytes. If a request's body size exceeds the configured value, data masking rules will be ignored. * max\_req\_post\_args integer default: `100` vaild vaule: greater than or equal to 0 *** Maximum number of URL-encoded form fields to parse when masking request body data with `body_format` set to `urlencoded`. --- # datadog The `datadog` plugin supports the integration with [Datadog](https://www.datadoghq.com), one of the most used observability service for cloud applications. When enabled, the plugin pushes metrics to [DogStatsD](https://docs.datadoghq.com/developers/dogstatsd/?tab=hostagent) server, which comes bundled with the [Datadog agent](https://docs.datadoghq.com/agent), over UDP protocol. ## Metrics[​](#metrics "Direct link to Metrics") The plugin exports the following metrics by default. All metrics will be prefixed by the `namespace` configured in metadata. For example, if the `namespace` is configured to be `apisix`, you will see the `request.counter` metric exported as `apisix.request.counter` in Datadog. | Name | Type | Description | | ---------------- | --------- | ----------------------------------------------------------------------------------------------------- | | request.counter | counter | Number of requests received. | | request.latency | histogram | Time taken to process the request, in milliseconds. | | upstream.latency | histogram | Time taken to proxy the request to the upstream server until a response is received, in milliseconds. | | apisix.latency | histogram | Time taken by APISIX agent to process the request, in milliseconds. | | ingress.size | timer | Request body size in bytes. | | egress.size | timer | Response body size in bytes. | ## Tags[​](#tags "Direct link to Tags") The plugin exports metrics with the following [tags](https://docs.datadoghq.com/getting_started/tagging). When there are no suitable values for any particular tag, the tag will be omitted. | Name | Description | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | route\_name | Name of the route. If not present or if the attribute `prefer_name` is set to false, fall back to the route ID. | | service\_name | Name of the service. If not present or if the attribute `prefer_name` is set to false, fall back to the service ID. | | consumer | Username of the consumer if the route is connected to a consumer. | | balancer\_ip | IP address of the upstream balancer that processes the current request. | | response\_status | HTTP response status code, such as `201`, `404`, or `503`. | | response\_status\_class | HTTP response status code class, such as `2xx`, `4xx`, or `5xx`. Available in APISIX from version 3.14.0 and API7 Enterprise from version 3.9.0. | | scheme | Request scheme, such as HTTP and gRPC. | | path | HTTP path pattern. Only available if the parameter `include_path` is set to `true`. Available in APISIX from version 3.14.0 and API7 Enterprise from version 3.9.0. | | method | HTTP method. Only available if the attribute `include_method` is set to true. Available in APISIX from version 3.14.0 and API7 Enterprise from version 3.9.0. | ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `datadog` plugin for different scenarios. Before proceeding, please make sure you have installed [Datadog agent](https://docs.datadoghq.com/agent) which collects events and metrics from monitored objects and sends them to Datadog. Start the Datadog agent: * Docker * Kubernetes ``` docker run -d \ --name dogstatsd-agent \ -e DD_API_KEY=35ebe12345678dec56218930b79fdb4cf \ -e DD_SITE="us5.datadoghq.com" \ -e DD_HOSTNAME=apisix.quickstart \ -e DD_DOGSTATSD_NON_LOCAL_TRAFFIC=true \ -p 8125:8125/udp \ datadog/dogstatsd:latest ``` ❶ `DD_API_KEY`: replace with your API key. ❷ `DD_SITE`: replace with your Datadog site. ❸ `DD_HOSTNAME`: replace with your hostname. ❹ `DD_DOGSTATSD_NON_LOCAL_TRAFFIC`: set to true to listen to DogStatsD packets from other containers. Create a Kubernetes manifest file for the Datadog DogStatsD agent: dogstatsd-deployment.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: dogstatsd-agent spec: replicas: 1 selector: matchLabels: app: dogstatsd-agent template: metadata: labels: app: dogstatsd-agent spec: containers: - name: dogstatsd-agent image: datadog/dogstatsd:latest env: - name: DD_API_KEY value: "35ebe12345678dec56218930b79fdb4cf" - name: DD_SITE value: "us5.datadoghq.com" - name: DD_HOSTNAME value: "apisix.quickstart" - name: DD_DOGSTATSD_NON_LOCAL_TRAFFIC value: "true" ports: - containerPort: 8125 protocol: UDP ``` ❶ `DD_API_KEY`: replace with your API key. ❷ `DD_SITE`: replace with your Datadog site. ❸ `DD_HOSTNAME`: replace with your hostname. ❹ `DD_DOGSTATSD_NON_LOCAL_TRAFFIC`: set to true to listen to DogStatsD packets from other containers. Create a Kubernetes manifest file for the DogStatsD service: dogstatsd-service.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: dogstatsd-agent spec: selector: app: dogstatsd-agent ports: - name: dogstatsd port: 8125 targetPort: 8125 protocol: UDP type: ClusterIP ``` Apply the manifests: ``` kubectl apply -f dogstatsd-deployment.yaml -f dogstatsd-service.yaml ``` You can configure most options in the agent’s main configuration file `datadog.yaml` through environment variables, prefixed with `DD_`. For more information, see [agent environment variables](https://docs.datadoghq.com/agent/guide/environment-variables). ### Update Datadog Agent Address and Other Metadata[​](#update-datadog-agent-address-and-other-metadata "Direct link to Update Datadog Agent Address and Other Metadata") By default, the plugin expects the DogStatsD server to be available at `127.0.0.1:8125`. To customize the address and other metadata, update the [plugin metadata](https://docs.api7.ai/hub/datadog/configuration.md#metadata) as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/datadog" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "host": "192.168.0.90", "port": 8125, "namespace": "apisix", "constant_tags": [ "source:apisix", "service:custom" ] }' ``` ❶ Replace with your private IP address. If you are running the Datadog agent in Kubernetes, use the service DNS name (e.g., `dogstatsd-agent.aic.svc`). ❷ Set to Datadog agent listening port. ❸ Set namespace which prefixes all metrics. ❹ Configure constant tags. To reset to default configuration, send a request to the `datadog` plugin metadata with an empty body: ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/datadog" -X PUT -d '{}' ``` adc.yaml ``` plugin_metadata: - name: datadog host: "192.168.0.90" port: 8125 namespace: apisix constant_tags: - "source:apisix" - "service:custom" ``` ❶ Replace with your private IP address. If you are running the Datadog agent in Kubernetes, use the service DNS name (e.g., `dogstatsd-agent.aic.svc`). ❷ Set to Datadog agent listening port. ❸ Set namespace which prefixes all metrics. ❹ Configure constant tags. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` datadog-metadata.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: service: name: apisix-admin port: 9180 auth: type: AdminKey adminKey: value: edd1c9f034335f136f87ad84b625c8f1 pluginMetadata: datadog: host: "dogstatsd-agent.aic.svc" port: 8125 namespace: apisix constant_tags: - "source:apisix" - "service:custom" ``` ❶ Set to the Datadog DogStatsD agent service DNS name in Kubernetes. ❷ Set to Datadog agent listening port. ❸ Set namespace which prefixes all metrics. ❹ Configure constant tags. Apply the configuration: ``` kubectl apply -f datadog-metadata.yaml ``` ### Monitor Route Metrics[​](#monitor-route-metrics "Direct link to Monitor Route Metrics") The example below shows how you can send the metrics of a particular route to Datadog. Create a route with the `datadog` plugin and a few optional configuration options: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "datadog-route", "uri": "/anything", "plugins": { "datadog": { "batch_max_size" : 1, "max_retry_count": 0 } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: datadog-route plugins: datadog: batch_max_size: 1 max_retry_count: 0 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD datadog-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: datadog-plugin-config spec: plugins: - name: datadog config: batch_max_size: 1 max_retry_count: 0 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: datadog-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: datadog-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` datadog-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: datadog-route spec: ingressClassName: apisix http: - name: datadog-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: datadog config: batch_max_size: 1 max_retry_count: 0 ``` Apply the configuration: ``` kubectl apply -f datadog-ic.yaml ``` ❶ `batch_max_size`: set to 1 to send the metric immediately. ❷ `max_retry_count`: set to 0 to disallow retries if metrics were unsuccessfully sent. Generate a few requests to the previously created route: ``` curl "http://127.0.0.1:9080/anything" ``` In Datadog, Select **Metrics** from the left menu and go to **Explorer**. Select `apisix.ingress.size.count` as the metric. You should see the count reflecting the number of requests generated: ![Datadog Metrics Explorer with an apisix.ingress.size.count query and the resulting line chart](https://static.api7.ai/uploads/2024/01/17/Y0uHlIeS_dd-count.png) --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * prefer\_name boolean default: `true` *** If true, export route/service name instead of their ID in metric tags. * include\_path boolean default: `false` *** If true, include the path pattern in metric tags. This option is available in APISIX but not yet supported in API7 Enterprise. * include\_method boolean default: `false` *** If true, include the HTTP method in metric tags. This option is available in APISIX but not yet supported in API7 Enterprise. * constant\_tags array\[string] *** Static key-value tags that are attached to all metrics. These tags can be used to add metadata such as team ownership or environment, enabling easier filtering, aggregation, and alerting across related endpoints. Available in APISIX from version 3.14.0 and API7 Enterprise from version 3.9.0. * name string default: `datadog` *** Unique identifier of the plugin for the batch processor. If you use [Prometheus](https://docs.api7.ai/hub/prometheus.md) to monitor APISIX metrics, the name is exported in `apisix_batch_process_entries`. * batch\_max\_size integer default: `1000` vaild vaule: greater than 0 *** The number of log entries allowed in one batch. Once reached, the batch will be sent to Datadog agent. Setting this parameter to 1 means immediate processing. * inactive\_timeout integer default: `5` vaild vaule: greater than 0 *** The maximum time in seconds to wait for new logs before sending the batch to the logging service. The value should be smaller than `buffer_duration`. * buffer\_duration integer default: `60` vaild vaule: greater than 0 *** The maximum time in seconds from the earliest entry allowed before sending the batch to the logging service. * retry\_delay integer default: `1` vaild vaule: greater than or equal to 0 *** The time interval in seconds to retry sending the batch to the logging service if the batch was not successfully sent. * max\_retry\_count integer default: `0` vaild vaule: greater than or equal to 0 *** The maximum number of unsuccessful retries allowed before dropping the log entries. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * host string default: `127.0.0.1` *** DogStatsD server host address. * port integer default: `8125` *** DogStatsD server port. * namespace string default: `apisix` *** Prefix for all metrics. For example, if the namespace is configured as `apisix`, you should see the `request.counter` metric exported as `apisix.request.counter` to Datadog. * constant\_tags array\[string] default: `[source:apisix]` *** Metric [tags](https://docs.datadoghq.com/getting_started/tagging). * max\_pending\_entries integer default: `` `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 and 3.10.6 `` vaild vaule: greater than or equal to 1 *** Maximum number of entries waiting in the batch processor. New entries are discarded when the backlog reaches the limit. Introduced in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line, and in APISIX 3.18.0. See [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md#configure-the-pending-entry-limit) for sizing and verification guidance. --- # degraphql The `degraphql` plugin supports communicating with upstream GraphQL services over regular HTTP requests by mapping GraphQL queries to HTTP endpoints. ## Examples[​](#examples "Direct link to Examples") The examples below use [Pokemon GraphQL API](https://graphql-pokemon.js.org/) as the upstream GraphQL server and demonstrate how you can configure `degraphql` to transform different types of GraphQL queries. ### Transform a Basic Query[​](#transform-a-basic-query "Direct link to Transform a Basic Query") The following example demonstrates how you can transform a simple query below: ``` query { getAllPokemon { key color } } ``` * Admin API * ADC * Ingress Controller Create a route with the `degraphql` plugin as follows: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "degraphql-route", "methods": ["POST"], "uri": "/v8", "upstream": { "type": "roundrobin", "nodes": { "graphqlpokemon.favware.tech": 1 }, "scheme": "https", "pass_host": "node" }, "plugins": { "degraphql": { "query": "{\n getAllPokemon {\n key\n color\n }\n}" } } }' ``` Create a route with the `degraphql` plugin as follows: adc.yaml ``` services: - name: degraphql-service routes: - name: degraphql-route methods: - POST uris: - /v8 plugins: degraphql: query: | { getAllPokemon { key color } } upstream: type: roundrobin nodes: - host: graphqlpokemon.favware.tech port: 443 weight: 1 scheme: https ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD degraphql-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: graphql-pokemon spec: type: ExternalName externalName: graphqlpokemon.favware.tech --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: graphql-pokemon-https spec: targetRefs: - name: graphql-pokemon kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: degraphql-plugin-config spec: plugins: - name: degraphql config: query: | { getAllPokemon { key color } } --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: degraphql-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /v8 method: - POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: degraphql-plugin-config backendRefs: - name: graphql-pokemon port: 443 ``` degraphql-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: degraphql-route spec: ingressClassName: apisix http: - name: degraphql-route match: paths: - /v8 methods: - POST upstreams: - name: graphql-pokemon plugins: - name: degraphql enable: true config: query: | { getAllPokemon { key color } } --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: graphql-pokemon spec: ingressClassName: apisix externalNodes: - type: Domain name: graphqlpokemon.favware.tech port: 443 scheme: https passHost: node ``` Apply the configuration to your cluster: ``` kubectl apply -f degraphql-ic.yaml ``` Send a request to the route to verify: ``` curl "http://127.0.0.1:9080/v8" -X POST ``` You should see a response similar to the following: ``` { "data": { "getAllPokemon": [ { "key": "pokestarsmeargle", "color": "White" }, { "key": "pokestarufo", "color": "White" }, { "key": "pokestarufo2", "color": "White" }, ... { "key": "terapagosstellar", "color": "Blue" }, { "key": "pecharunt", "color": "Purple" } ] } } ``` ### Transform a Query with Variables[​](#transform-a-query-with-variables "Direct link to Transform a Query with Variables") The following example demonstrates how you can transform the query below, with a variable: ``` query ($pokemon: PokemonEnum!) { getPokemon( pokemon: $pokemon ) { color species } } variable: { "pokemon": "pikachu" } ``` * Admin API * ADC * Ingress Controller Create a route with the `degraphql` plugin as follows: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "degraphql-route", "uri": "/v8", "upstream": { "type": "roundrobin", "nodes": { "graphqlpokemon.favware.tech": 1 }, "scheme": "https", "pass_host": "node" }, "plugins": { "degraphql": { "query": "query ($pokemon: PokemonEnum!) {\n getPokemon(\n pokemon: $pokemon\n ) {\n color\n species\n }\n}\n", "variables": ["pokemon"] } } }' ``` Create a route with the `degraphql` plugin as follows: adc.yaml ``` services: - name: degraphql-service routes: - name: degraphql-route uris: - /v8 plugins: degraphql: query: | query ($pokemon: PokemonEnum!) { getPokemon( pokemon: $pokemon ) { color species } } variables: - pokemon upstream: type: roundrobin nodes: - host: graphqlpokemon.favware.tech port: 443 weight: 1 scheme: https ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD degraphql-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: graphql-pokemon spec: type: ExternalName externalName: graphqlpokemon.favware.tech --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: graphql-pokemon-https spec: targetRefs: - name: graphql-pokemon kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: degraphql-plugin-config spec: plugins: - name: degraphql config: query: | query ($pokemon: PokemonEnum!) { getPokemon( pokemon: $pokemon ) { color species } } variables: - pokemon --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: degraphql-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /v8 filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: degraphql-plugin-config backendRefs: - name: graphql-pokemon port: 443 ``` degraphql-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: degraphql-route spec: ingressClassName: apisix http: - name: degraphql-route match: paths: - /v8 upstreams: - name: graphql-pokemon plugins: - name: degraphql enable: true config: query: | query ($pokemon: PokemonEnum!) { getPokemon( pokemon: $pokemon ) { color species } } variables: - pokemon --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: graphql-pokemon spec: ingressClassName: apisix externalNodes: - type: Domain name: graphqlpokemon.favware.tech port: 443 scheme: https passHost: node ``` Apply the configuration to your cluster: ``` kubectl apply -f degraphql-ic.yaml ``` Send a request to the route to verify: ``` curl "http://127.0.0.1:9080/v8" -X POST \ -d '{ "pokemon": "pikachu" }' ``` You should see a response similar to the following: ``` { "data": { "getPokemon": { "color": "Yellow", "species": "pikachu" } } } ``` Alternatively, you can also pass the variable in the URL query string of a GET request: ``` curl "http://127.0.0.1:9080/v8?pokemon=pikachu" -H "x-apollo-operation-name: GET" ``` You should see the same response as the previous. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * query string required *** GraphQL query to be sent to the upstream. * operation\_name string *** Name of the operation. Required if multiple operations are included in the query. * variables array\[string] *** Variables used in GraphQL queries. * max\_req\_body\_size integer default: `67108864` vaild vaule: greater than or equal to 1 *** Maximum POST request body size in bytes read when extracting configured GraphQL `variables`. A larger body returns `503 Service Unavailable`. The request body is not used when `variables` is not configured. Introduced in API7 Enterprise 3.9.17 and 3.10.4, and APISIX 3.18.0. --- # elasticsearch-logger The `elasticsearch-logger` plugin pushes request and response logs in batches to [Elasticsearch](https://www.elastic.co) and supports the customization of log formats. When enabled, the plugin will serialize the request context information to [Elasticsearch Bulk format](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-bulk.html#docs-bulk) and add them to the queue, before they are pushed to Elasticsearch. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `elasticsearch-logger` plugin for different scenarios. To follow along the examples, start Elasticsearch and Kibana: * Docker * Kubernetes Start an Elasticsearch instance: ``` docker run -d \ --name elasticsearch \ --network apisix-quickstart-net \ -v elasticsearch_vol:/usr/share/elasticsearch/data/ \ -p 9200:9200 \ -p 9300:9300 \ -e ES_JAVA_OPTS="-Xms512m -Xmx512m" \ -e discovery.type=single-node \ -e xpack.security.enabled=false \ docker.elastic.co/elasticsearch/elasticsearch:7.17.29 ``` Start a Kibana instance to visualize the indexed data in Elasticsearch: ``` docker run -d \ --name kibana \ --network apisix-quickstart-net \ -p 5601:5601 \ -e ELASTICSEARCH_HOSTS="http://elasticsearch:9200" \ docker.elastic.co/kibana/kibana:7.17.29 ``` Create a Kubernetes manifest file for the Elasticsearch deployment: elasticsearch-deployment.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: elasticsearch spec: replicas: 1 selector: matchLabels: app: elasticsearch template: metadata: labels: app: elasticsearch spec: containers: - name: elasticsearch image: docker.elastic.co/elasticsearch/elasticsearch:7.17.29 env: - name: ES_JAVA_OPTS value: "-Xms512m -Xmx512m" - name: discovery.type value: single-node - name: xpack.security.enabled value: "false" ports: - containerPort: 9200 - containerPort: 9300 ``` Create a Kubernetes manifest file for the Elasticsearch service: elasticsearch-service.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: elasticsearch spec: selector: app: elasticsearch ports: - name: http port: 9200 targetPort: 9200 - name: transport port: 9300 targetPort: 9300 type: ClusterIP ``` Create a Kubernetes manifest file for the Kibana deployment: kibana-deployment.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: kibana spec: replicas: 1 selector: matchLabels: app: kibana template: metadata: labels: app: kibana spec: containers: - name: kibana image: docker.elastic.co/kibana/kibana:7.17.29 env: - name: ELASTICSEARCH_HOSTS value: "http://elasticsearch.aic.svc:9200" ports: - containerPort: 5601 ``` Create a Kubernetes manifest file for the Kibana service: kibana-service.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: kibana spec: selector: app: kibana ports: - name: http port: 5601 targetPort: 5601 type: ClusterIP ``` Apply the manifests: ``` kubectl apply -f elasticsearch-deployment.yaml -f elasticsearch-service.yaml -f kibana-deployment.yaml -f kibana-service.yaml ``` To access Kibana, forward the service port: ``` kubectl port-forward -n aic svc/kibana 5601:5601 ``` If successful, you should see the Kibana dashboard on [localhost:5601](http://localhost:5601). ### Log in the Default Log Format[​](#log-in-the-default-log-format "Direct link to Log in the Default Log Format") The following example demonstrates how you can enable the `elasticsearch-logger` plugin on a route, which logs client requests and responses, as well as pushing logs to Elasticsearch. Create a route with `elasticsearch-logger` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "elasticsearch-logger-route", "uri": "/anything", "plugins": { "elasticsearch-logger": { "endpoint_addrs": ["http://elasticsearch:9200"], "field": { "index": "gateway" } } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: elasticsearch-logger-route plugins: elasticsearch-logger: endpoint_addrs: - "http://elasticsearch:9200" field: index: gateway upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD elasticsearch-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: elasticsearch-logger-plugin-config spec: plugins: - name: elasticsearch-logger config: endpoint_addrs: - "http://elasticsearch.aic.svc:9200" field: index: gateway --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: elasticsearch-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: elasticsearch-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` elasticsearch-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: elasticsearch-logger-route spec: ingressClassName: apisix http: - name: elasticsearch-logger-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: elasticsearch-logger config: endpoint_addrs: - "http://elasticsearch.aic.svc:9200" field: index: gateway ``` Apply the configuration: ``` kubectl apply -f elasticsearch-logger-ic.yaml ``` ❶ Configure the endpoint address to Elasticsearch. ❷ Configure the `index` field as `gateway`. Send a request to the route to generate a log entry: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to the Kibana dashboard on [localhost:5601](http://localhost:5601) and under **Discover** tab, create a new index pattern `gateway` to fetch the data from Elasticsearch. Once configured, navigate back to the **Discover** tab and you should see a log generated, similar to the following: ``` { "_index": "gateway", "_id": "CE-JL5QBOkdYRG7kEjTJ", "_version": 1, "_score": 1, "_source": { "request": { "headers": { "host": "127.0.0.1:9080", "accept": "*/*", "user-agent": "curl/8.6.0" }, "size": 85, "querystring": {}, "method": "GET", "url": "http://127.0.0.1:9080/anything", "uri": "/anything" }, "response": { "headers": { "content-type": "application/json", "access-control-allow-credentials": "true", "server": "APISIX/3.13.0", "content-length": "390", "access-control-allow-origin": "*", "connection": "close", "date": "Mon, 13 Jan 2025 10:18:14 GMT" }, "status": 200, "size": 618 }, "route_id": "elasticsearch-logger-route", "latency": 585.00003814697, "apisix_latency": 18.000038146973, "upstream_latency": 567, "upstream": "50.19.58.113:80", "server": { "hostname": "0b9a772e68f8", "version": "3.13.0" }, "service_id": "", "client_ip": "192.168.65.1" }, "fields": { ... } } ``` ### Customize Log Format With Plugin Metadata[​](#customize-log-format-with-plugin-metadata "Direct link to Customize Log Format With Plugin Metadata") The following example demonstrates how you can customize log format using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) and [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) to log specific headers from request and response. In APISIX, [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) is used to configure the common metadata fields of all plugin instances of the same plugin. It is useful when a plugin is enabled across multiple resources and requires a universal update to their metadata fields. First, create a route with `elasticsearch-logger` as follows: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "elasticsearch-logger-route", "uri": "/anything", "plugins": { "elasticsearch-logger": { "endpoint_addrs": ["http://elasticsearch:9200"], "field": { "index": "gateway" } } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` Next, configure the plugin metadata for `elasticsearch-logger`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/elasticsearch-logger" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "log_format": { "host": "$host", "@timestamp": "$time_iso8601", "client_ip": "$remote_addr", "env": "$http_env", "resp_content_type": "$sent_http_Content_Type" } }' ``` adc.yaml ``` plugin_metadata: - name: elasticsearch-logger log_format: host: "$host" "@timestamp": "$time_iso8601" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` elasticsearch-logger-metadata.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: service: name: apisix-admin port: 9180 auth: type: AdminKey adminKey: value: edd1c9f034335f136f87ad84b625c8f1 pluginMetadata: elasticsearch-logger: log_format: host: "$host" "@timestamp": "$time_iso8601" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` Apply the configuration: ``` kubectl apply -f elasticsearch-logger-metadata.yaml ``` ❶ log the custom request header `env`. ❷ log the response header `Content-Type`. Send a request to the route with the `env` header: ``` curl -i "http://127.0.0.1:9080/anything" -H "env: dev" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to the Kibana dashboard on [localhost:5601](http://localhost:5601) and under **Discover** tab, create a new index pattern `gateway` to fetch the data from Elasticsearch, if you have not done so already. Once configured, navigate back to the **Discover** tab and you should see a log generated, similar to the following: ``` { "_index": "gateway", "_id": "Ck-WL5QBOkdYRG7kODS0", "_version": 1, "_score": 1, "_source": { "client_ip": "192.168.65.1", "route_id": "elasticsearch-logger-route", "@timestamp": "2025-01-06T10:32:36+00:00", "host": "127.0.0.1", "resp_content_type": "application/json" }, "fields": { ... } } ``` ### Log Request Bodies Conditionally[​](#log-request-bodies-conditionally "Direct link to Log Request Bodies Conditionally") The following example demonstrates how you can conditionally log request body. Create a route with `elasticsearch-logger` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "elasticsearch-logger": { "endpoint_addrs": ["http://elasticsearch:9200"], "field": { "index": "gateway" }, "include_req_body": true, "include_req_body_expr": [["arg_log_body", "==", "yes"]] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" }, "uri": "/anything", "id": "elasticsearch-logger-route" }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: elasticsearch-logger-route plugins: elasticsearch-logger: endpoint_addrs: - "http://elasticsearch:9200" field: index: gateway include_req_body: true include_req_body_expr: - - arg_log_body - "==" - "yes" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD elasticsearch-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: elasticsearch-logger-plugin-config spec: plugins: - name: elasticsearch-logger config: endpoint_addrs: - "http://elasticsearch.aic.svc:9200" field: index: gateway include_req_body: true include_req_body_expr: - - arg_log_body - "==" - "yes" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: elasticsearch-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: elasticsearch-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` elasticsearch-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: elasticsearch-logger-route spec: ingressClassName: apisix http: - name: elasticsearch-logger-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: elasticsearch-logger config: endpoint_addrs: - "http://elasticsearch.aic.svc:9200" field: index: gateway include_req_body: true include_req_body_expr: - - arg_log_body - "==" - "yes" ``` Apply the configuration: ``` kubectl apply -f elasticsearch-logger-ic.yaml ``` ❶ `include_req_body`: set to true to include request body. ❷ `include_req_body_expr`: only include request body if the URL query string `log_body` is `true`. Send a request to the route with an URL query string satisfying the condition: ``` curl -i "http://127.0.0.1:9080/anything?log_body=yes" -X POST -d '{"env": "dev"}' ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to the Kibana dashboard on [localhost:5601](http://localhost:5601) and under **Discover** tab, create a new index pattern `gateway` to fetch the data from Elasticsearch, if you have not done so already. Once configured, navigate back to the **Discover** tab and you should see a log generated, similar to the following: ``` { "_index": "gateway", "_id": "Dk-cL5QBOkdYRG7k7DSW", "_version": 1, "_score": 1, "_source": { "request": { "headers": { "user-agent": "curl/8.6.0", "accept": "*/*", "content-length": "14", "host": "127.0.0.1:9080", "content-type": "application/x-www-form-urlencoded" }, "size": 182, "querystring": { "log_body": "yes" }, "body": "{\"env\": \"dev\"}", "method": "POST", "url": "http://127.0.0.1:9080/anything?log_body=yes", "uri": "/anything?log_body=yes" }, "start_time": 1735965595203, "response": { "headers": { "content-type": "application/json", "server": "APISIX/3.13.0", "access-control-allow-credentials": "true", "content-length": "548", "access-control-allow-origin": "*", "connection": "close", "date": "Mon, 13 Jan 2025 11:02:32 GMT" }, "status": 200, "size": 776 }, "route_id": "elasticsearch-logger-route", "latency": 703.9999961853, "apisix_latency": 34.999996185303, "upstream_latency": 669, "upstream": "34.197.122.172:80", "server": { "hostname": "0b9a772e68f8", "version": "3.13.0" }, "service_id": "", "client_ip": "192.168.65.1" }, "fields": { ... } } ``` Send a request to the route without any URL query string: ``` curl -i "http://127.0.0.1:9080/anything" -X POST -d '{"env": "dev"}' ``` Navigate to the Kibana dashboard **Discover** tab and you should see a log generated, but without the request body: ``` { "_index": "gateway", "_id": "EU-eL5QBOkdYRG7kUDST", "_version": 1, "_score": 1, "_source": { "request": { "headers": { "content-type": "application/x-www-form-urlencoded", "accept": "*/*", "content-length": "14", "host": "127.0.0.1:9080", "user-agent": "curl/8.6.0" }, "size": 169, "querystring": {}, "method": "POST", "url": "http://127.0.0.1:9080/anything", "uri": "/anything" }, "start_time": 1735965686363, "response": { "headers": { "content-type": "application/json", "access-control-allow-credentials": "true", "server": "APISIX/3.13.0", "content-length": "510", "access-control-allow-origin": "*", "connection": "close", "date": "Mon, 13 Jan 2025 11:15:54 GMT" }, "status": 200, "size": 738 }, "route_id": "elasticsearch-logger-route", "latency": 680.99999427795, "apisix_latency": 4.9999942779541, "upstream_latency": 676, "upstream": "34.197.122.172:80", "server": { "hostname": "0b9a772e68f8", "version": "3.13.0" }, "service_id": "", "client_ip": "192.168.65.1" }, "fields": { ... } } ``` info If you have customized the `log_format` in addition to setting `include_req_body` or `include_resp_body` to `true`, the plugin would not include the bodies in the logs. As a workaround, you may be able to use the NGINX variable `$request_body` in the log format, such as: ``` { "elasticsearch-logger": { ..., "log_format": {"body": "$request_body"} } } ``` ### Include Request Date in Elasticsearch Index[​](#include-request-date-in-elasticsearch-index "Direct link to Include Request Date in Elasticsearch Index") The following example demonstrates how you can configure the `elasticsearch-logger` plugin to include the request date in Elasticsearch index. Create a route with `elasticsearch-logger` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "elasticsearch-logger-route", "uri": "/anything", "plugins": { "elasticsearch-logger": { "endpoint_addrs": ["http://elasticsearch:9200"], "field": { "index": "api7-{%Y.%m.%d}" } } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: elasticsearch-logger-route plugins: elasticsearch-logger: endpoint_addrs: - "http://elasticsearch:9200" field: index: "api7-{%Y.%m.%d}" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD elasticsearch-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: elasticsearch-logger-plugin-config spec: plugins: - name: elasticsearch-logger config: endpoint_addrs: - "http://elasticsearch.aic.svc:9200" field: index: "api7-{%Y.%m.%d}" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: elasticsearch-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: elasticsearch-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` elasticsearch-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: elasticsearch-logger-route spec: ingressClassName: apisix http: - name: elasticsearch-logger-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: elasticsearch-logger config: endpoint_addrs: - "http://elasticsearch.aic.svc:9200" field: index: "api7-{%Y.%m.%d}" ``` Apply the configuration: ``` kubectl apply -f elasticsearch-logger-ic.yaml ``` ❶ Configure the endpoint address to Elasticsearch. ❷ Configure the `index` field to use the current year, month, and date. Send a request to the route to generate a log entry: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to the Kibana dashboard on [localhost:5601](http://localhost:5601) and under **Discover** tab, create a new index pattern `api7*` to fetch the data from Elasticsearch. Once configured, navigate back to the **Discover** tab and you should see a log generated, similar to the following: ``` { "_index": "api7-2025.3.10", "_id": "CE-KL5QB0kdYRG7dEiTJ", "_version": 1, "_score": 1, "_source": { "request": { ... }, "response": { ... }, "status": 200, "size": 618 }, ... } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * endpoint\_addr string *** Deprecated. Use `endpoint_addrs` instead. Elasticsearch API endpoint address. Configure either `endpoint_addr` or `endpoint_addrs`. * endpoint\_addrs array\[string] *** Elasticsearch API endpoint addresses. If multiple endpoints are configured, one is selected randomly for each write. Configure either `endpoint_addrs` or the deprecated `endpoint_addr`. * field object required *** Elasticsearch field configurations. * index string required *** Elasticsearch [`_index`](https://www.elastic.co/guide/en/elasticsearch/reference/current/mapping-index-field.html#mapping-index-field) field. In API7 Enterprise from version 3.8.0 and APISIX from version 3.17.0, `index` supports the configuration of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) and a [lua time format](https://www.lua.org/pil/22.1.html) in curly brackets to include the current date, such as `service-$host-{%Y-%m-%d}`. * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. You can also configure log format on a global scale using the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md), which configures the log format for all `elasticsearch-logger` plugin instances. If the log format configured on the individual plugin instance differs from the log format configured on plugin metadata, the log format configured on the individual plugin instance takes precedence. See the [example](https://docs.api7.ai/hub/elasticsearch-logger.md#log-request-and-response-headers-with-plugin-metadata) for more details. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * auth object *** Elasticsearch user authentication configurations. * username string *** Elasticsearch authentication username. * password string *** Elasticsearch authentication password. The value is encrypted before being stored. * headers object *** Custom HTTP request headers to include in requests sent to Elasticsearch, as key-value pairs. They can complement or replace `auth` for authentication and other purposes. Introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.16.0. Header-value encryption was introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.18.0. The values are [encrypted before storage](https://docs.api7.ai/apisix/production/security/data-encryption-with-keyring.md) when data encryption is enabled. Authorized Admin API `GET` requests return the complete decrypted header object; encryption at rest does not redact header names or values from API responses. * ssl\_verify boolean default: `true` *** If true, perform SSL verification. * timeout integer default: `10` *** Elasticsearch send data timeout in seconds. * include\_req\_body boolean default: `false` *** If true, include the request body in the log. Note that if the request body is too big to be kept in the memory, it can not be logged due to NGINX's limitations. * include\_req\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_req_body` is true. Request body would only be logged when the expressions configured here evaluate to true. * include\_resp\_body boolean default: `false` *** If true, include the response body in the log. * include\_resp\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_resp_body` is true. Response body would only be logged when the expressions configured here evaluate to true. * max\_req\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes to include in the log. If the request body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * max\_resp\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum response body size in bytes to include in the log. If the response body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * name string default: `elasticsearch-logger` *** Unique identifier of the plugin for the batch processor. If you use [Prometheus](https://docs.api7.ai/hub/prometheus.md) to monitor APISIX metrics, the name is exported in `apisix_batch_process_entries`. * batch\_max\_size integer default: `1000` vaild vaule: greater than 0 *** Number of log entries allowed in one batch. Once reached, the batch is sent to Elasticsearch. Setting this parameter to 1 enables immediate processing. * inactive\_timeout integer default: `5` vaild vaule: greater than 0 *** Maximum time in seconds to wait for new logs before sending the batch. The value should be smaller than `buffer_duration`. * buffer\_duration integer default: `60` vaild vaule: greater than 0 *** Maximum time in seconds from the earliest entry before sending the batch. * retry\_delay integer default: `1` vaild vaule: greater than or equal to 0 *** Time in seconds to wait before retrying a failed batch. * max\_retry\_count integer default: `0` vaild vaule: greater than or equal to 0 *** Maximum number of unsuccessful retries before dropping the log entries. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * max\_pending\_entries integer default: `` `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 and 3.10.6; none in API7 Enterprise 3.9.18 and 3.10.5 `` vaild vaule: greater than or equal to 1 *** Maximum number of entries waiting in the batch processor. New entries are discarded when the backlog reaches the limit. Introduced in API7 Enterprise 3.8.17 and APISIX 3.15.0. The default changed to `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line. In API7 Enterprise 3.9.18 and 3.10.5, and in earlier APISIX versions, omitting the parameter leaves the backlog unlimited. See [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md#configure-the-pending-entry-limit) for sizing and verification guidance. --- # error-log-collect [Enterprise](https://api7.ai/enterprise) The `error-log-collect` plugin captures the error logs that the gateway produces while processing matched requests and writes them to the gateway error log (`error.log`). The captured entries include lower-severity logs, such as `INFO` and `DEBUG`, that the configured error log level would normally discard. This lets you collect detailed, request-scoped diagnostics for a targeted subset of traffic, without lowering the global error log level for all requests. Each captured entry is written at `error` level, prefixed with `[error-log-collect]`, and tagged with the request ID, so you can filter and correlate the entries in the gateway log. Use `vars` to restrict collection to requests that match a condition, and `sample_ratio` to capture only a fraction of requests on high-traffic routes. This plugin is configured on routes or services, and is available in API7 Enterprise from version 3.10.0. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `error-log-collect` in different scenarios. ### Collect Error Logs on a Route[​](#collect-error-logs-on-a-route "Direct link to Collect Error Logs on a Route") The following example demonstrates how to enable the plugin on a route and view the collected logs. Create a route to httpbin.org with the `error-log-collect` plugin enabled: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "error-log-collect-route", "uri": "/anything", "plugins": { "error-log-collect": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: error-log-collect-service routes: - name: error-log-collect-route uris: - /anything plugins: error-log-collect: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` error-log-collect.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: error-log-collect-route spec: ingressClassName: apisix http: - name: anything match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: error-log-collect config: {} ``` Apply the configuration to your cluster: ``` kubectl apply -f error-log-collect.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. In the gateway log, look for entries prefixed with `[error-log-collect]`, each tagged with the request ID. The plugin re-emits the internal logs generated while handling the request. These include `INFO`-level entries, such as DNS resolution and upstream selection, that the default `warn` log level would normally omit: ``` 2026/06/26 09:21:31 [error] 47#47: 1750901491123#0 [error-log-collect] 2026-06-26 09:21:31 b9f8c1d2e3a4f5061728394a5b6c7d8e parse_domain():118: dns resolve httpbin.org, context: ngx.timer ``` note The plugin buffers the captured logs in memory per worker process, up to `buffer_max_size` entries. It flushes them to `error.log` when a request matches `vars`, or on every request when `vars` is not set. The buffer is shared by all requests that a worker handles, so a flush can also surface buffered logs from other recent requests on the same worker. This helps capture the context leading up to a matched event. ### Collect Logs Only for Matching Requests[​](#collect-logs-only-for-matching-requests "Direct link to Collect Logs Only for Matching Requests") To collect logs only for requests that meet a condition, set `vars` to one or more [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). For example, the following configuration collects logs only when the request carries an `X-Debug: true` header: ``` { "plugins": { "error-log-collect": { "vars": [ ["http_x_debug", "==", "true"] ] } } } ``` Requests that do not match the condition are not flushed to the error log on their own. On high-traffic routes, set `sample_ratio` below `1` to collect logs for a random sample of requests and keep the log volume manageable. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * vars array\[array] *** An array of one or more matching conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). The buffered logs are flushed to the error log only when all expressions evaluate to true. When unset, the logs are collected for every request. * sample\_ratio number default: `1` vaild vaule: between 0.00001 and 1 inclusive *** Probability of collecting the logs for a request. The default value of 1 collects the logs for all requests. Set to a value below 1 to collect the logs for a random sample of requests. * buffer\_max\_size integer default: `1000` vaild vaule: greater than or equal to 1 *** Maximum number of log entries held in the per-worker buffer. When the number of buffered logs exceeds this value, the oldest entries are overwritten. --- # error-log-logger The `error-log-logger` plugin pushes APISIX's error logs (`error.log`) to TCP, Apache SkyWalking, Apache Kafka, or ClickHouse servers, in batches. You can specify the severity level of which the plugin should send the corresponding logs. The plugin is disabled by default. Once enabled, it will automatically start pushing error logs to remote servers. You should configure remote server details in [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) only, instead of on other resources, such as routes. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `error-log-logger` plugin for different scenarios. The APISIX and API7 Gateway runtime configurations do not load `error-log-logger` by default. Enable it in the gateway static configuration before configuring plugin metadata. * Host or Docker * Kubernetes (Helm) Keep the existing plugin list in `config.yaml` and add `error-log-logger`: config.yaml ``` plugins: # Keep the complete plugin list used by your gateway. - error-log-logger ``` Reload the gateway for changes to take effect. For the APISIX Helm chart, `apisix.plugins` replaces the loaded plugin list. Start from the complete plugin list used by your gateway and add `error-log-logger`: values.yaml ``` apisix: plugins: # Keep the complete plugin list used by your gateway. - error-log-logger ``` For API7 Gateway Helm deployments, continue with the plugin metadata configuration after confirming that `error-log-logger` is loaded in the gateway plugin list. The current chart does not expose a dedicated `values.yaml` field for adding `error-log-logger` to the loaded plugin list. Check the [API7 Gateway Helm chart reference](https://docs.api7.ai/api7-gateway/reference/helm-chart.md) for the latest supported plugin-list configuration. Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` ### Send Logs to TCP Server[​](#send-logs-to-tcp-server "Direct link to Send Logs to TCP Server") The following example demonstrates how you can configure the `error-log-logger` plugin to send error logs to a TCP server. Start a TCP server listening on port `19000`: * Docker * Kubernetes ``` nc -l 19000 ``` Create a Kubernetes manifest for a TCP server deployment using `socat`: tcp-server.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: tcp-server spec: replicas: 1 selector: matchLabels: app: tcp-server template: metadata: labels: app: tcp-server spec: containers: - name: tcp-server image: alpine/socat args: ["TCP-LISTEN:19000,fork,reuseaddr", "STDOUT"] ports: - containerPort: 19000 --- apiVersion: v1 kind: Service metadata: namespace: aic name: tcp-server spec: selector: app: tcp-server ports: - name: tcp port: 19000 targetPort: 19000 type: ClusterIP ``` Apply the manifests: ``` kubectl apply -f tcp-server.yaml ``` Configure the plugin metadata for `error-log-logger`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/error-log-logger" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "tcp": { "host": "192.168.2.103", "port": 19000 }, "level": "INFO" }' ``` adc.yaml ``` plugin_metadata: - name: error-log-logger tcp: host: "192.168.2.103" port: 19000 level: INFO ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` error-log-logger-metadata.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: service: name: apisix-admin port: 9180 auth: type: AdminKey adminKey: value: edd1c9f034335f136f87ad84b625c8f1 pluginMetadata: error-log-logger: tcp: host: "tcp-server.aic.svc" port: 19000 level: INFO ``` Apply the configuration: ``` kubectl apply -f error-log-logger-metadata.yaml ``` ❶ Configure the host to the TCP server address. ❷ Configure the port to your TCP server listening port. ❸ Configure the severity level to `INFO` so most logs would be sent, for easier verification. To verify, you can manually generate a log at `warn` level by [reloading APISIX](https://docs.api7.ai/apisix/reference/apisix-cli.md#apisix-reload). If you are using Docker, in the terminal session where netcat is listening, you should see a log entry. If you are using Kubernetes, check the tcp-server pod logs: ``` kubectl logs -n aic -l app=tcp-server ``` You should see a log entry similar to the following: ``` 2025/01/26 20:15:29 [warn] 211#211: *35552 [lua] plugin.lua:205: load(): new plugins: {"cas-auth":true,"real-ip":true,"ai":true,"client-control":true,"proxy-control":true,"request-id":true,"zipkin":true,"ext-plugin-pre-req":true,"fault-injection":true,"mocking":true,"serverless-pre-function":true,"cors":true,"ip-restriction":true,"ua-restriction":true,"referer-restriction":true,"csrf":true,"uri-blocker":true,"request-validation":true,"chaitin-waf":true,"multi-auth":true,"openid-connect":true,"authz-casbin":true,"authz-casdoor":true,"wolf-rbac":true,"ldap-auth":true,"hmac-auth":true,"basic-auth":true,"jwt-auth":true,"redirect":true,"key-auth":true,"consumer-restriction":true,"attach-consumer-label":true,"authz-keycloak":true,"proxy-cache":true,"body-transformer":true,"ai-prompt-template":true,"ai-prompt-decorator":true,"proxy-mirror":true,"proxy-rewrite":true,"workflow":true,"api-breaker":true,"ai-proxy":true,"limit-conn":true,"limit-count":true,"limit-req":true,"gzip":true,"server-info":true,"traffic-split":true,"response-rewrite":true,"degraphql":true,"kafka-proxy":true,"grpc-transcode":true,"grpc-web":true,"http-dubbo":true,"public-api":true,"prometheus":true,"datadog":true,"loki-logger":true,"elasticsearch-logger":true,"echo":true,"loggly":true,"http-logger":true,"splunk-hec-logging":true,"skywalking-logger":true,"google-cloud-logging":true,"sls-logger":true,"tcp-logger":true,"kafka-logger":true,"rocketmq-logger":true,"syslog":true,"udp-logger":true,"file-logger":true,"clickhouse-logger":true,"tencent-cloud-cls":true,"inspect":true,"example-plugin":true,"aws-lambda":true,"azure-functions":true,"openwhisk":true,"openfunction":true,"error-log-logger":true,"ext-plugin-post-req":true,"ext-plugin-post-resp":true,"serverless-post-function":true,"opa":true,"forward-auth":true,"jwe-decrypt":true}, context: init_worker_by_lua* ``` ### Send Logs to SkyWalking[​](#send-logs-to-skywalking "Direct link to Send Logs to SkyWalking") The following example demonstrates how you can configure the `error-log-logger` plugin to send error logs to SkyWalking. Set up SkyWalking OAP server: * Docker * Kubernetes Start a SkyWalking storage, OAP and Booster UI with Docker Compose, following [Skywalking's documentation](https://skywalking.apache.org/docs/main/next/en/setup/backend/backend-docker/). Once set up, the OAP server should be listening on `12800` and you should be able to access the UI at . Create a Kubernetes manifest for the SkyWalking OAP server: skywalking-oap.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: skywalking-oap spec: replicas: 1 selector: matchLabels: app: skywalking-oap template: metadata: labels: app: skywalking-oap spec: containers: - name: skywalking-oap image: apache/skywalking-oap-server:10.1.0 env: - name: SW_STORAGE value: H2 ports: - containerPort: 11800 - containerPort: 12800 --- apiVersion: v1 kind: Service metadata: namespace: aic name: skywalking-oap spec: selector: app: skywalking-oap ports: - name: grpc port: 11800 targetPort: 11800 - name: http port: 12800 targetPort: 12800 type: ClusterIP ``` Apply the manifests: ``` kubectl apply -f skywalking-oap.yaml ``` Wait for the OAP server to become ready: ``` kubectl wait --for=condition=available --timeout=120s -n aic deployment/skywalking-oap ``` Configure the plugin metadata for `error-log-logger`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/error-log-logger" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "skywalking": { "endpoint_addr": "http://192.168.2.103:12800/v3/logs" }, "level": "INFO" }' ``` adc.yaml ``` plugin_metadata: - name: error-log-logger skywalking: endpoint_addr: "http://192.168.2.103:12800/v3/logs" level: INFO ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` error-log-logger-metadata.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: service: name: apisix-admin port: 9180 auth: type: AdminKey adminKey: value: edd1c9f034335f136f87ad84b625c8f1 pluginMetadata: error-log-logger: skywalking: endpoint_addr: "http://skywalking-oap.aic.svc:12800/v3/logs" level: INFO ``` Apply the configuration: ``` kubectl apply -f error-log-logger-metadata.yaml ``` ❶ Configure the endpoint address to the SkyWalking server. ❷ Configure the severity level to `INFO` so most logs would be sent, for easier verification. To verify, you can manually generate a log at `warn` level by [reloading APISIX](https://docs.api7.ai/apisix/reference/apisix-cli.md#apisix-reload). In [Skywalking UI](http://localhost:8080), navigate to **General Service** > **Services**. You should see a service called `APISIX` with the following log entry: ``` 2025/01/27 07:40:06 [warn] 211#211: *35552 [lua] plugin.lua:205: load(): new plugins: {"cas-auth":true,"real-ip":true,"ai":true,"client-control":true,"proxy-control":true,"request-id":true,"zipkin":true,"ext-plugin-pre-req":true,"fault-injection":true,"mocking":true,"serverless-pre-function":true,"cors":true,"ip-restriction":true,"ua-restriction":true,"referer-restriction":true,"csrf":true,"uri-blocker":true,"request-validation":true,"chaitin-waf":true,"multi-auth":true,"openid-connect":true,"authz-casbin":true,"authz-casdoor":true,"wolf-rbac":true,"ldap-auth":true,"hmac-auth":true,"basic-auth":true,"jwt-auth":true,"redirect":true,"key-auth":true,"consumer-restriction":true,"attach-consumer-label":true,"authz-keycloak":true,"proxy-cache":true,"body-transformer":true,"ai-prompt-template":true,"ai-prompt-decorator":true,"proxy-mirror":true,"proxy-rewrite":true,"workflow":true,"api-breaker":true,"ai-proxy":true,"limit-conn":true,"limit-count":true,"limit-req":true,"gzip":true,"server-info":true,"traffic-split":true,"response-rewrite":true,"degraphql":true,"kafka-proxy":true,"grpc-transcode":true,"grpc-web":true,"http-dubbo":true,"public-api":true,"prometheus":true,"datadog":true,"loki-logger":true,"elasticsearch-logger":true,"echo":true,"loggly":true,"http-logger":true,"splunk-hec-logging":true,"skywalking-logger":true,"google-cloud-logging":true,"sls-logger":true,"tcp-logger":true,"kafka-logger":true,"rocketmq-logger":true,"syslog":true,"udp-logger":true,"file-logger":true,"clickhouse-logger":true,"tencent-cloud-cls":true,"inspect":true,"example-plugin":true,"aws-lambda":true,"azure-functions":true,"openwhisk":true,"openfunction":true,"error-log-logger":true,"ext-plugin-post-req":true,"ext-plugin-post-resp":true,"serverless-post-function":true,"opa":true,"forward-auth":true,"jwe-decrypt":true}, context: init_worker_by_lua* ``` You should also observe logs at other severity levels, such as `error`, `emerg`, and `info`, when they are generated. ### Send Logs to Kafka over TLS[​](#send-logs-to-kafka-over-tls "Direct link to Send Logs to Kafka over TLS") The following example sends error-level gateway logs to a TLS-enabled Kafka broker. Complete the trusted-CA setup in [Send Logs to a TLS-Enabled Broker](https://docs.api7.ai/hub/kafka-logger.md#send-logs-to-a-tls-enabled-broker), then set the broker address and topic: ``` export KAFKA_TLS_HOST="kafka-tls" export KAFKA_TLS_PORT="9093" export KAFKA_ERROR_TOPIC="apisix-error-logs" ``` Create the dedicated error-log topic: ``` docker exec kafka-tls /opt/kafka/bin/kafka-topics.sh \ --bootstrap-server kafka-tls:9093 \ --command-config /etc/kafka/secrets/client.properties \ --create \ --if-not-exists \ --topic "${KAFKA_ERROR_TOPIC}" \ --partitions 1 \ --replication-factor 1 ``` Configure the plugin metadata: ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/error-log-logger" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- <\n \n 404\n \n \n
\n

404 not found

\n
\n
\n
Gateway
\n \n", "content_type": "text/html" } }' ``` To demonstrate the function of the plugin, create a route with the [`serverless-post-function`](https://docs.api7.ai/hub/serverless-functions.md) plugin, which returns a 404 error code from gateways for all requests to the route: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "uri": "/*", "id": "error-page-route", "plugins": { "serverless-post-function": { "functions": [ "return function (conf, ctx) local core = require(\"apisix.core\") core.response.exit(404) end" ] }, "error-page": {} }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` ❶ Return a 404 status code for all requests to the route. ❷ Enable `error-page` to return the customized error page. Configure the plugin metadata and create a route with the `serverless-post-function` and `error-page` plugins: adc.yaml ``` plugin_metadata: error-page: enable: true error_404: body: "\n \n 404\n \n \n
\n

404 not found

\n
\n
\n
Gateway
\n \n" content_type: text/html services: - name: error-page-service routes: - name: error-page-route uris: - /* plugins: serverless-post-function: functions: - | return function (conf, ctx) local core = require("apisix.core") core.response.exit(404) end error-page: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ Return a 404 status code for all requests to the route. ❷ Enable `error-page` to return the customized error page. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create a route with the `serverless-post-function` and `error-page` plugins: error-page-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: error-page-plugin-config spec: plugins: - name: serverless-post-function config: functions: - | return function (conf, ctx) local core = require("apisix.core") core.response.exit(404) end - name: error-page config: {} --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: error-page-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: / filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: error-page-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` ❶ Return a 404 status code for all requests to the route. ❷ Enable `error-page` to return the customized error page. Create a route with the `serverless-post-function` and `error-page` plugins: error-page-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: error-page-route spec: ingressClassName: apisix http: - name: error-page-route match: paths: - /* upstreams: - name: httpbin-external-domain plugins: - name: serverless-post-function enable: true config: functions: - | return function (conf, ctx) local core = require("apisix.core") core.response.exit(404) end - name: error-page enable: true config: {} ``` ❶ Return a 404 status code for all requests to the route. ❷ Enable `error-page` to return the customized error page. Update your GatewayProxy manifest to configure the `error-page` plugin metadata: gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: error-page: enable: true error_404: body: "\n \n 404\n \n \n
\n

404 not found

\n
\n
\n
Gateway
\n \n" content_type: text/html ``` Apply the configuration to your cluster: ``` kubectl apply -f gatewayproxy.yaml -f error-page-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 404 Not Found` response with the following response body: ``` 404

404 not found


Gateway
``` --- ## Parameters[​](#parameters "Direct link to Parameters") This plugin has no configurable parameters when configured on routes and services. All configuration options should be configured using [plugin metadata](#plugin-metadata). See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * enable boolean default: `false` *** If true, enable the plugin. * error\_404 object *** Error page to return when APISIX returns 404 status codes. * body string default: `Product-specific HTML response` *** Response body. The default identifies Apache APISIX or API7 Enterprise Edition, depending on the gateway distribution. * content\_type string default: `text/html` *** Response content type. * error\_500 object *** Error page to return when APISIX returns 500 status codes. * body string default: `Product-specific HTML response` *** Response body. The default identifies Apache APISIX or API7 Enterprise Edition, depending on the gateway distribution. * content\_type string default: `text/html` *** Response content type. * error\_502 object *** Error page to return when APISIX returns 502 status codes. * body string default: `Product-specific HTML response` *** Response body. The default identifies Apache APISIX or API7 Enterprise Edition, depending on the gateway distribution. * content\_type string default: `text/html` *** Response content type. * error\_503 object *** Error page to return when APISIX returns 503 status codes. * body string default: `Product-specific HTML response` *** Response body. The default identifies Apache APISIX or API7 Enterprise Edition, depending on the gateway distribution. * content\_type string default: `text/html` *** Response content type. --- # exit-transformer The `exit-transformer` plugin customizes responses produced through APISIX's response-exit path, including plugin rejections and route-not-found responses. It does not transform ordinary responses returned by upstream services. The transformation logics are defined in the plugin using Lua functions, following the syntax: ``` return (function(code, body, header) if {{ condition }} then return {{ modified_resp }} end return code, body, header end)(...) ``` ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can use `exit-transformer` for different scenarios. ### Enable `exit-transformer` Plugin[​](#enable-exit-transformer-plugin "Direct link to enable-exit-transformer-plugin") For APISIX deployments, load `exit-transformer` in the gateway static configuration before configuring global rules or routes that use it. API7 Gateway users can configure the plugin through the Dashboard or Admin API when it is enabled in the installed gateway release. * Host or Docker * Kubernetes (Helm) For APISIX host or Docker deployments, keep the existing plugin list in `config.yaml` and add `exit-transformer`: config.yaml ``` plugins: # Keep the existing plugin list. - exit-transformer ``` Reload the gateway for changes to take effect. For the APISIX Helm chart, `apisix.plugins` replaces the loaded plugin list. Start from the complete plugin list used by your gateway, add `exit-transformer`, and keep the rest of the list unchanged: values.yaml ``` apisix: plugins: # Keep the complete plugin list used by your gateway. - exit-transformer ``` Apply the values file with the APISIX Helm chart: ``` helm upgrade -n -f values.yaml ``` ### Modify 404 Route Not Found Response[​](#modify-404-route-not-found-response "Direct link to Modify 404 Route Not Found Response") The following example demonstrates how you can use the plugin to update the `404 Not Found` response code and header when the route does not exist. In this case, the plugin needs to be configured as a global rule plugin. Create a global rule with the `exit-transformer` plugin, in which the function updates the response status code to `405` and adds a custom `X-Custom-Header` header if the original status code was `404`: * Admin API * ADC * Ingress Controller ``` curl -i "http://127.0.0.1:9180/apisix/admin/global_rules" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "transform-404-not-found", "plugins": { "exit-transformer": { "functions": ["return (function(code, body, header) header = header or {} if code == 404 then header[\"X-Custom-Header\"] = \"Modified\" return 405, body, header end return code, body, header end)(...)"] } } }' ``` adc.yaml ``` global_rules: - id: transform-404-not-found plugins: exit-transformer: functions: - "return (function(code, body, header) header = header or {} if code == 404 then header[\"X-Custom-Header\"] = \"Modified\" return 405, body, header end return code, body, header end)(...)" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # add your control plane connection configuration here # .... plugins: - name: exit-transformer enabled: true config: functions: - "return (function(code, body, header) header = header or {} if code == 404 then header[\"X-Custom-Header\"] = \"Modified\" return 405, body, header end return code, body, header end)(...)" ``` Apply the configuration: ``` kubectl apply -f gatewayproxy.yaml ``` exit-transformer-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixGlobalRule metadata: namespace: aic name: transform-404-not-found spec: ingressClassName: apisix plugins: - name: exit-transformer enable: true config: functions: - "return (function(code, body, header) header = header or {} if code == 404 then header[\"X-Custom-Header\"] = \"Modified\" return 405, body, header end return code, body, header end)(...)" ``` Apply the configuration: ``` kubectl apply -f exit-transformer-ic.yaml ``` Send a request to a route that does not exist: ``` curl -i "http://127.0.0.1:9080/non-existent" ``` You should receive an `HTTP/1.1 405 Not Allowed` response and observe the `X-Custom-Header: Modified` header. ### Modify 401 Unauthorized Response for Failed Authentication[​](#modify-401-unauthorized-response-for-failed-authentication "Direct link to Modify 401 Unauthorized Response for Failed Authentication") The following example demonstrates how you can use the plugin to update the `401 Unauthorized` response when authentication fails. * Admin API * ADC * Ingress Controller Create a route with the `exit-transformer` plugin, in which the function updates the response status code to `402` if the original status code was `401`; and enable `key-auth`: ``` curl -i "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "transform-auth-route", "uri": "/get", "plugins": { "exit-transformer": { "functions": ["return (function(code, body, header) if code == 401 then return 402, body, header end return code, body, header end)(...)"] }, "key-auth":{} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john" }' ``` Configure `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Create a consumer with `key-auth` credential and a route with `exit-transformer` and `key-auth` plugins configured as such: adc.yaml ``` consumers: - username: john credentials: - name: key-auth type: key-auth config: key: john-key services: - name: httpbin routes: - name: transform-auth-route uris: - /get plugins: exit-transformer: functions: - "return (function(code, body, header) if code == 401 then return 402, body, header end return code, body, header end)(...)" key-auth: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create a consumer with `key-auth` credential and a route with `exit-transformer` and `key-auth` plugins configured as such: * Gateway API * APISIX CRD exit-transformer-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-cred config: key: john-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: exit-transformer-plugin-config spec: plugins: - name: exit-transformer config: functions: - "return (function(code, body, header) if code == 401 then return 402, body, header end return code, body, header end)(...)" - name: key-auth config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: transform-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: exit-transformer-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` exit-transformer-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: transform-auth-route spec: ingressClassName: apisix http: - name: transform-auth-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: exit-transformer enable: true config: functions: - "return (function(code, body, header) if code == 401 then return 402, body, header end return code, body, header end)(...)" - name: key-auth enable: true ``` Apply the configuration: ``` kubectl apply -f exit-transformer-ic.yaml ``` Send a request to the route without the credential: ``` curl -i "http://127.0.0.1:9080/get" ``` You should receive an `HTTP/1.1 402 Payment Required` response for unauthorized access, where the response status code has been modified. ### Modify Response Conditionally on Request Headers[​](#modify-response-conditionally-on-request-headers "Direct link to Modify Response Conditionally on Request Headers") The following example demonstrates how you can use the plugin to conditionally modify responses based on request headers. Create a route with the `exit-transformer` plugin, in which the function updates the response status code by `Content-Type` header. If the header value is `application/json` and the original status code is `404`, update the response status code to `405`. Print warning messages inside and outside the condition evaluation for demonstration purposes. * Admin API * ADC * Ingress Controller ``` curl -i "http://127.0.0.1:9180/apisix/admin/global_rules" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "transform-by-header-condition", "plugins": { "exit-transformer": { "functions": [ "return (function(code, body, header) local core = require(\"apisix.core\") local ct = ngx.req.get_headers()[\"Content-Type\"] core.log.warn(\"exit transformer logics running outside the condition\") if ct == \"application/json\" and code == 404 then core.log.warn(\"exit transformer logics running inside the condition\") return 405 end return code, body, header end) (...)" ] } } }' ``` adc.yaml ``` global_rules: - id: transform-by-header-condition plugins: exit-transformer: functions: - "return (function(code, body, header) local core = require(\"apisix.core\") local ct = ngx.req.get_headers()[\"Content-Type\"] core.log.warn(\"exit transformer logics running outside the condition\") if ct == \"application/json\" and code == 404 then core.log.warn(\"exit transformer logics running inside the condition\") return 405 end return code, body, header end)(...)" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # add your control plane connection configuration here # .... plugins: - name: exit-transformer config: functions: - "return (function(code, body, header) local core = require(\"apisix.core\") local ct = ngx.req.get_headers()[\"Content-Type\"] core.log.warn(\"exit transformer logics running outside the condition\") if ct == \"application/json\" and code == 404 then core.log.warn(\"exit transformer logics running inside the condition\") return 405 end return code, body, header end)(...)" ``` Apply the configuration: ``` kubectl apply -f gatewayproxy.yaml ``` exit-transformer-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixGlobalRule metadata: namespace: aic name: transform-by-header-condition spec: ingressClassName: apisix plugins: - name: exit-transformer enable: true config: functions: - "return (function(code, body, header) local core = require(\"apisix.core\") local ct = ngx.req.get_headers()[\"Content-Type\"] core.log.warn(\"exit transformer logics running outside the condition\") if ct == \"application/json\" and code == 404 then core.log.warn(\"exit transformer logics running inside the condition\") return 405 end return code, body, header end)(...)" ``` Apply the configuration: ``` kubectl apply -f exit-transformer-ic.yaml ``` Send a request to a route that does not exist, without any header: ``` curl -i "http://127.0.0.1:9080/non-existent" ``` You should receive an `HTTP/1.1 404 Not Found` response and see the following message in the log: ``` exit transformer logics running outside the condition ``` Send a request to the non-existent route with the JSON `Content-Type` header: ``` curl -i "http://127.0.0.1:9080/non-existent" -H "Content-Type: application/json" ``` You should receive an `HTTP/1.1 405 Not Allowed` response, where the response status code has been modified, and see the following message in the log: ``` exit transformer logics running outside the condition exit transformer logics running inside the condition ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * functions array\[string] required *** Exit transformation Lua function. The functions do not allow the use of `coroutine`, `math`, `os`, `string`, or `table` for security reasons. --- # fault-injection The `fault-injection` plugin is designed to test your application's resiliency by simulating controlled faults or delays. It executes before other configured plugins, ensuring that faults are applied consistently. This makes it ideal for scenarios like chaos engineering, where the behavior of your system under failure conditions is analyzed. The plugin supports two key actions: `abort`, which immediately terminates a request with a specified HTTP status code (e.g., `503 Service Unavailable`), skipping all subsequent plugins; and `delay`, which introduces a specified delay before processing the request further. These features allow you to simulate scenarios such as service outages or latency, helping you validate error-handling logic and improve system reliability. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure the `fault-injection` plugin for different scenarios. ### Inject Faults[​](#inject-faults "Direct link to Inject Faults") The following example demonstrate how you can configure the `fault-injection` plugin on a route to intercept further request sending, and respond with a specific HTTP code. Create a route using the `fault-injection` plugin with the `abort` action to respond any request with `404` and the specified response body: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "fault-injection-route", "uri": "/anything", "plugins": { "fault-injection": { "abort": { "http_status": 404, "body": "APISIX Fault Injection" } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: fault-injection-route uris: - /anything plugins: fault-injection: abort: http_status: 404 body: "APISIX Fault Injection" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD fault-injection-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: fault-injection-plugin-config spec: plugins: - name: fault-injection config: abort: http_status: 404 body: "APISIX Fault Injection" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: fault-injection-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: fault-injection-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f fault-injection-ic.yaml ``` fault-injection-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: fault-injection-route spec: ingressClassName: apisix http: - name: fault-injection-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: fault-injection enable: true config: abort: http_status: 404 body: "APISIX Fault Injection" ``` Apply the configuration to your cluster: ``` kubectl apply -f fault-injection-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 404 Not Found` response and see the following response body, without the request being forwarded to the upstream service: ``` APISIX Fault Injection ``` ### Inject Latencies[​](#inject-latencies "Direct link to Inject Latencies") The following example demonstrate how you can configure the `fault-injection` plugin on a route to inject request latencies. Create a route using the `fault-injection` plugin with the `delay` action to delay response sending for 3 seconds: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "fault-injection-route", "uri": "/anything", "plugins": { "fault-injection": { "delay": { "duration": 3 } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: fault-injection-route uris: - /anything plugins: fault-injection: delay: duration: 3 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD fault-injection-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: fault-injection-plugin-config spec: plugins: - name: fault-injection config: delay: duration: 3 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: fault-injection-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: fault-injection-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f fault-injection-ic.yaml ``` fault-injection-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: fault-injection-route spec: ingressClassName: apisix http: - name: fault-injection-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: fault-injection enable: true config: delay: duration: 3 ``` Apply the configuration to your cluster: ``` kubectl apply -f fault-injection-ic.yaml ``` Send a request to the route and use `time` command to summarize of how long the request took to complete: ``` time curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response from the upstream service and see the following timing summary: ``` 0.01s user 0.01s system 0% cpu 3.685 total ``` ### Inject Faults Conditionally[​](#inject-faults-conditionally "Direct link to Inject Faults Conditionally") The following example demonstrate how you can configure the `fault-injection` plugin on a route to intercept further request sending, and respond with a specific HTTP code. Create a route using the `fault-injection` plugin with the `abort` action as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "fault-injection-route", "uri": "/anything", "plugins": { "fault-injection": { "abort": { "http_status": 404, "body": "APISIX Fault Injection", "headers": { "X-APISIX-Remote-Addr": "$remote_addr" }, "vars": [ [ [ "arg_name","==","john" ] ] ] } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: fault-injection-route uris: - /anything plugins: fault-injection: abort: http_status: 404 body: "APISIX Fault Injection" headers: X-APISIX-Remote-Addr: $remote_addr vars: - - - arg_name - "==" - john upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD fault-injection-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: fault-injection-plugin-config spec: plugins: - name: fault-injection config: abort: http_status: 404 body: "APISIX Fault Injection" headers: X-APISIX-Remote-Addr: $remote_addr vars: - - - arg_name - "==" - john --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: fault-injection-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: fault-injection-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f fault-injection-ic.yaml ``` fault-injection-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: fault-injection-route spec: ingressClassName: apisix http: - name: fault-injection-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: fault-injection enable: true config: abort: http_status: 404 body: "APISIX Fault Injection" headers: X-APISIX-Remote-Addr: $remote_addr vars: - - - arg_name - "==" - john ``` Apply the configuration to your cluster: ``` kubectl apply -f fault-injection-ic.yaml ``` ❶ Respond requests with HTTP status code `404`. ❷ Respond requests with `APISIX Fault Injection` as the body. ❸ Respond requests with header `X-APISIX-Remote-Addr` and the IP the request originates from. ❹ Respond requests with the above specifications only if the URL parameter `name` value is `john`. Send a request to the route with the URL parameter `name` being `john`: ``` curl -i "http://127.0.0.1:9080/anything?name=john" ``` You should receive an `HTTP/1.1 404 Not Found` response: ``` HTTP/1.1 404 Not Found ... X-APISIX-Remote-Addr: 192.168.65.1 APISIX Fault Injection ``` Send a request to the route with the URL parameter `name` being a different value: ``` curl -i "http://127.0.0.1:9080/anything?name=jane" ``` You should receive an `HTTP/1.1 200 OK` response from the upstream service without the injections. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * abort object *** Abort action configurations. * http\_status integer required vaild vaule: greater than or equal to 200 *** Response HTTP status code to return to client. * body string *** Body of the response returned to the client. Support the use of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) in the body. * headers object *** Headers of the response returned to the client. Support the use of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) in headers. * percentage integer vaild vaule: between 0 and 100 inclusive *** Percentage of requests to be aborted. * vars array\[array] *** An array of one or more matching conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md) to conditionally execute the plugin. * delay object *** Delay action configurations. * duration number required *** Delay duration in seconds. * percentage integer vaild vaule: between 0 and 100 inclusive *** Percentage of requests to be delayed. * vars array\[array] *** An array of one or more matching conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md) to conditionally execute the plugin. --- # forward-auth The `forward-auth` plugin supports the integration with an external authorization service for authentication and authorization. If the authentication fails, a customizable error message will be returned to the client. If the authentication succeeds, the request will be forwarded to the upstream service along with the following request headers that APISIX added: * `X-Forwarded-Proto`: scheme * `X-Forwarded-Method`: HTTP method * `X-Forwarded-Host`: host * `X-Forwarded-Uri`: URI * `X-Forwarded-For`: source IP ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can use `forward-auth` for different scenarios. To follow along the first two examples, please have your external authorization service set up, or create a mock auth service using the [serverless function plugin](https://docs.api7.ai/hub/serverless-functions.md) as shown below: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -H 'Content-Type: application/json' \ -d '{ "id": "auth-mock", "uri": "/auth", "plugins": { "serverless-pre-function": { "phase": "rewrite", "functions": [ "return function (conf, ctx) local core = require(\"apisix.core\"); local authorization = core.request.header(ctx, \"Authorization\"); if authorization == \"123\" then core.response.exit(200); elseif authorization == \"321\" then core.response.set_header(\"X-User-ID\", \"i-am-user\"); core.response.exit(200); else core.response.set_header(\"X-Forward-Auth\", \"Fail\"); core.response.exit(403); end end" ] } } }' ``` ❶ If the `Authorization` header has a value of `123`, respond with `200 OK`; ❷ If the `Authorization` header has a value of `321`, set a header `X-User-ID: i-am-user` and respond with `200 OK`; ❸ Otherwise, set a header `X-Forward-Auth: Fail` and respond with `403 Forbidden`. adc-auth-mock.yaml ``` services: - name: auth-mock-service routes: - name: auth-mock-route uris: - /auth plugins: serverless-pre-function: phase: rewrite functions: - | return function(conf, ctx) local core = require("apisix.core") local authorization = core.request.header(ctx, "Authorization") if authorization == "123" then core.response.exit(200) elseif authorization == "321" then core.response.set_header("X-User-ID", "i-am-user") core.response.exit(200) else core.response.set_header("X-Forward-Auth", "Fail") core.response.exit(403) end end upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ If the `Authorization` header has a value of `123`, respond with `200 OK`; ❷ If the `Authorization` header has a value of `321`, set a header `X-User-ID: i-am-user` and respond with `200 OK`; ❸ Otherwise, set a header `X-Forward-Auth: Fail` and respond with `403 Forbidden`. Synchronize the configuration to the gateway: ``` adc sync -f adc-auth-mock.yaml ``` * Gateway API * APISIX CRD forward-auth-mock-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: auth-mock-plugin-config spec: plugins: - name: serverless-pre-function config: phase: rewrite functions: - | return function(conf, ctx) local core = require("apisix.core") local authorization = core.request.header(ctx, "Authorization") if authorization == "123" then core.response.exit(200) elseif authorization == "321" then core.response.set_header("X-User-ID", "i-am-user") core.response.exit(200) else core.response.set_header("X-Forward-Auth", "Fail") core.response.exit(403) end end --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: auth-mock-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /auth filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: auth-mock-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` ❶ If the `Authorization` header has a value of `123`, respond with `200 OK`; ❷ If the `Authorization` header has a value of `321`, set a header `X-User-ID: i-am-user` and respond with `200 OK`; ❸ Otherwise, set a header `X-Forward-Auth: Fail` and respond with `403 Forbidden`. Apply the configuration to your cluster: ``` kubectl apply -f forward-auth-mock-ic.yaml ``` forward-auth-mock-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: auth-mock-route spec: ingressClassName: apisix http: - name: auth-mock-route match: paths: - /auth upstreams: - name: httpbin-external-domain plugins: - name: serverless-pre-function enable: true config: phase: rewrite functions: - | return function(conf, ctx) local core = require("apisix.core") local authorization = core.request.header(ctx, "Authorization") if authorization == "123" then core.response.exit(200) elseif authorization == "321" then core.response.set_header("X-User-ID", "i-am-user") core.response.exit(200) else core.response.set_header("X-Forward-Auth", "Fail") core.response.exit(403) end end ``` ❶ If the `Authorization` header has a value of `123`, respond with `200 OK`; ❷ If the `Authorization` header has a value of `321`, set a header `X-User-ID: i-am-user` and respond with `200 OK`; ❸ Otherwise, set a header `X-Forward-Auth: Fail` and respond with `403 Forbidden`. Apply the configuration to your cluster: ``` kubectl apply -f forward-auth-mock-ic.yaml ``` ### Forward Designated Headers to Upstream Resource[​](#forward-designated-headers-to-upstream-resource "Direct link to Forward Designated Headers to Upstream Resource") The following example demonstrates how to set up `forward-auth` on a route to regulate client access to the resources upstream based on a value in the request header. It also allows passing a specific header from the authorization service to the upstream resource. Create a route with the `forward-auth` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "forward-auth-route", "uri": "/headers", "plugins": { "forward-auth": { "uri": "http://127.0.0.1:9080/auth", "request_headers": ["Authorization"], "upstream_headers": ["X-User-ID"] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` ❶ The URI of the authorization service. ❷ The request header that should be forwarded to the authorization service. ❸ The request header set by the authorization service that should be forwarded to the upstream resource when the authorization succeeds. adc.yaml ``` services: - name: forward-auth-service routes: - name: forward-auth-route uris: - /headers plugins: forward-auth: uri: http://127.0.0.1:9080/auth request_headers: - Authorization upstream_headers: - X-User-ID upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ The URI of the authorization service. ❷ The request header that should be forwarded to the authorization service. ❸ The request header set by the authorization service that should be forwarded to the upstream resource when the authorization succeeds. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD forward-auth-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: forward-auth-plugin-config spec: plugins: - name: forward-auth config: uri: http://apisix-gateway.aic.svc.cluster.local/auth request_headers: - Authorization upstream_headers: - X-User-ID --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: forward-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: forward-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` ❶ The URI of the authorization service. When using the Ingress Controller, reference the mock auth service using its Kubernetes service address. ❷ The request header that should be forwarded to the authorization service. ❸ The request header set by the authorization service that should be forwarded to the upstream resource when the authorization succeeds. Apply the configuration to your cluster: ``` kubectl apply -f forward-auth-ic.yaml ``` forward-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: forward-auth-route spec: ingressClassName: apisix http: - name: forward-auth-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: forward-auth enable: true config: uri: http://apisix-gateway.aic.svc.cluster.local/auth request_headers: - Authorization upstream_headers: - X-User-ID ``` ❶ The URI of the authorization service. When using the Ingress Controller, reference the mock auth service using its Kubernetes service address. ❷ The request header that should be forwarded to the authorization service. ❸ The request header set by the authorization service that should be forwarded to the upstream resource when the authorization succeeds. Apply the configuration to your cluster: ``` kubectl apply -f forward-auth-ic.yaml ``` Send a request to the route with authorization detail in the header: ``` curl "http://127.0.0.1:9080/headers" -H 'Authorization: 123' ``` You should see an `HTTP/1.1 200 OK` response of the following: ``` { "headers": { "Accept": "*/*", "Authorization": "123", ... } } ``` To verify if the `X-User-ID` header set by the authorization service will be forwarded to the upstream service, send a request to the route with the corresponding authorization detail: ``` curl "http://127.0.0.1:9080/headers" -H 'Authorization: 321' ``` You should see an `HTTP/1.1 200 OK` response of the following, showing the header is forwarded to the upstream: ``` { "headers": { "Accept": "*/*", "Authorization": "123", "X-User-ID": "i-am-user", ... } } ``` ### Return Designated Headers to Clients on Authentication Failure[​](#return-designated-headers-to-clients-on-authentication-failure "Direct link to Return Designated Headers to Clients on Authentication Failure") The following example demonstrates how you can configure `forward-auth` on a route to regulate client access to the upstream resources. It also passes a specific header returned by the authorization service to the client when the authentication fails. Create a route with the `forward-auth` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "forward-auth-route", "uri": "/headers", "plugins": { "forward-auth": { "uri": "http://127.0.0.1:9080/auth", "request_headers": ["Authorization"], "client_headers": ["X-Forward-Auth"] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` ❶ Pass the `X-Forward-Auth` header from the authorization service back to the client when authentication fails. adc.yaml ``` services: - name: forward-auth-service routes: - name: forward-auth-route uris: - /headers plugins: forward-auth: uri: http://127.0.0.1:9080/auth request_headers: - Authorization client_headers: - X-Forward-Auth upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ Pass the `X-Forward-Auth` header from the authorization service back to the client when authentication fails. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD forward-auth-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: forward-auth-plugin-config spec: plugins: - name: forward-auth config: uri: http://apisix-gateway.aic.svc.cluster.local/auth request_headers: - Authorization client_headers: - X-Forward-Auth --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: forward-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: forward-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` ❶ Pass the `X-Forward-Auth` header from the authorization service back to the client when authentication fails. Apply the configuration to your cluster: ``` kubectl apply -f forward-auth-ic.yaml ``` forward-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: forward-auth-route spec: ingressClassName: apisix http: - name: forward-auth-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: forward-auth enable: true config: uri: http://apisix-gateway.aic.svc.cluster.local/auth request_headers: - Authorization client_headers: - X-Forward-Auth ``` ❶ Pass the `X-Forward-Auth` header from the authorization service back to the client when authentication fails. Apply the configuration to your cluster: ``` kubectl apply -f forward-auth-ic.yaml ``` Send a request without any authentication information: ``` curl -i "http://127.0.0.1:9080/headers" ``` You should receive an `HTTP/1.1 403 Forbidden` response: ``` ... X-Forward-Auth: Fail Server: APISIX/3.x.x 403 Forbidden

403 Forbidden


openresty

Powered by APISIX.

``` ### Authorize Based on POST Body[​](#authorize-based-on-post-body "Direct link to Authorize Based on POST Body") This example demonstrates how to configure the `forward-auth` plugin to control access based on POST body data, pass values as headers to the authorization service, and reject the request when authorization failed per the body data. The example uses the built-in variable `$post_arg.*` to read a parameter from the request body. APISIX resolves `$post_arg.*` against `application/x-www-form-urlencoded`, `application/json`, and `multipart/form-data` bodies, so the client must send the correct `Content-Type` header for the body it actually sends. See [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) for details. Please have your external authorization service set up, or create a mock auth service using the [serverless function plugin](https://docs.api7.ai/hub/serverless-functions.md). The function checks if the `tenant_id` header is `123` and returns `200 OK` if it is, otherwise it returns 403 with an error message. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -H 'Content-Type: application/json' \ -d '{ "id": "auth-mock", "uri": "/auth", "plugins": { "serverless-pre-function": { "phase": "rewrite", "functions": [ "return function(conf, ctx) local core = require(\"apisix.core\") local tenant_id = core.request.header(ctx, \"tenant_id\") if tenant_id == \"123\" then core.response.exit(200); else core.response.exit(403, \"tenant_id is \"..tenant_id .. \" but expecting 123\"); end end" ] } } }' ``` adc-auth-mock.yaml ``` services: - name: auth-mock-service routes: - name: auth-mock-route uris: - /auth plugins: serverless-pre-function: phase: rewrite functions: - | return function(conf, ctx) local core = require("apisix.core") local tenant_id = core.request.header(ctx, "tenant_id") if tenant_id == "123" then core.response.exit(200) else core.response.exit(403, "tenant_id is " .. tenant_id .. " but expecting 123") end end upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc-auth-mock.yaml ``` * Gateway API * APISIX CRD forward-auth-post-mock-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: auth-mock-plugin-config spec: plugins: - name: serverless-pre-function config: phase: rewrite functions: - | return function(conf, ctx) local core = require("apisix.core") local tenant_id = core.request.header(ctx, "tenant_id") if tenant_id == "123" then core.response.exit(200) else core.response.exit(403, "tenant_id is " .. tenant_id .. " but expecting 123") end end --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: auth-mock-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /auth filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: auth-mock-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f forward-auth-post-mock-ic.yaml ``` forward-auth-post-mock-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: auth-mock-route spec: ingressClassName: apisix http: - name: auth-mock-route match: paths: - /auth upstreams: - name: httpbin-external-domain plugins: - name: serverless-pre-function enable: true config: phase: rewrite functions: - | return function(conf, ctx) local core = require("apisix.core") local tenant_id = core.request.header(ctx, "tenant_id") if tenant_id == "123" then core.response.exit(200) else core.response.exit(403, "tenant_id is " .. tenant_id .. " but expecting 123") end end ``` Apply the configuration to your cluster: ``` kubectl apply -f forward-auth-post-mock-ic.yaml ``` Create a route with the `forward-auth` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "forward-auth-route", "uri": "/post", "methods": ["POST"], "plugins": { "forward-auth": { "uri": "http://127.0.0.1:9080/auth", "request_method": "GET", "extra_headers": {"tenant_id": "$post_arg.tenant_id"} } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` ❶ Set an extra header `tenant_id` using the value from the POST parameter `tenant_id`. adc.yaml ``` services: - name: forward-auth-service routes: - name: forward-auth-route uris: - /post methods: - POST plugins: forward-auth: uri: http://127.0.0.1:9080/auth request_method: GET extra_headers: tenant_id: "$post_arg.tenant_id" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ Set an extra header `tenant_id` using the value from the POST parameter `tenant_id`. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD forward-auth-post-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: forward-auth-post-plugin-config spec: plugins: - name: forward-auth config: uri: http://apisix-gateway.aic.svc.cluster.local/auth request_method: GET extra_headers: tenant_id: "$post_arg.tenant_id" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: forward-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /post method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: forward-auth-post-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` ❶ Set an extra header `tenant_id` using the value from the POST parameter `tenant_id`. Apply the configuration to your cluster: ``` kubectl apply -f forward-auth-post-ic.yaml ``` forward-auth-post-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: forward-auth-route spec: ingressClassName: apisix http: - name: forward-auth-route match: paths: - /post methods: - POST upstreams: - name: httpbin-external-domain plugins: - name: forward-auth enable: true config: uri: http://apisix-gateway.aic.svc.cluster.local/auth request_method: GET extra_headers: tenant_id: "$post_arg.tenant_id" ``` ❶ Set an extra header `tenant_id` using the value from the POST parameter `tenant_id`. Apply the configuration to your cluster: ``` kubectl apply -f forward-auth-post-ic.yaml ``` Send a POST request with `tenant_id` in a JSON body: ``` curl -i "http://127.0.0.1:9080/post" -X POST \ -H 'Content-Type: application/json' \ -d '{"tenant_id": "123"}' ``` You should receive an `HTTP/1.1 200 OK` response. Send a POST request with a different `tenant_id`: ``` curl -i "http://127.0.0.1:9080/post" -X POST \ -H 'Content-Type: application/json' \ -d '{"tenant_id": "000"}' ``` You should receive an `HTTP/1.1 403 Forbidden` response of the following: ``` tenant_id is 000 but expecting 123 ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * uri string required *** URI of the external authorization service. * ssl\_verify boolean default: `true` *** If true, verify the authorization service's SSL certificate. * request\_method string default: `GET` vaild vaule: `GET` or `POST` *** HTTP method APISIX uses to send requests to the external authorization service. By default, APISIX sends GET requests to the external authorization service. When set to `POST`, APISIX will send POST requests along with the request body to the external authorization service. This is, however, not recommended. If the authorization decision depends on request parameters from a POST body, it is recommended to extract the necessary fields using `$post_arg.*` and pass them via the `extra_headers` field. This approach avoids sending the full request body, reduces overhead, and keeps the authorization service focused on headers for decision-making. * max\_req\_body\_size integer default: `67108864` *** Maximum request body size in bytes that is buffered and forwarded to the external authorization service when `request_method` is `POST` (default 67108864 bytes, which is 64 MB). Requests with a body larger than this limit are rejected with HTTP 413. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * request\_headers array\[string] default: `[]` *** Client request headers that should be forwarded to the external authorization service. If not configured, only headers added by APISIX are forwarded, such as `X-Forwarded-*`. * upstream\_headers array\[string] default: `[]` *** External authorization response headers controlled by the plugin before the request is forwarded upstream. The gateway forwards a configured header when the authorization service returns it and clears any client-supplied value when the authorization response omits it. If not configured, no authorization response headers are forwarded. * client\_headers array\[string] default: `[]` *** External authorization service response headers that should be forwarded to the client when authentication fails. If not configured, no headers are forwarded to the client. * extra\_headers object *** Additional headers to send to the authorization service. Support [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) in values. * timeout integer default: `3000` vaild vaule: between 1 and 60000 inclusive *** Timeout for the external authorization service HTTP call in milliseconds. * keepalive boolean default: `true` *** If true, keep the connections open for multiple requests. * keepalive\_timeout integer default: `60000` vaild vaule: greater than or equal to 1000 *** Idle time after which the established HTTP connections will be closed. * keepalive\_pool integer default: `5` vaild vaule: greater than or equal to 1 *** Maximum number of connections in the connection pool. * allow\_degradation boolean default: `false` *** If true, allow APISIX to continue handling requests without the plugin when the plugin or its dependencies become unavailable. * status\_on\_error integer default: `403` vaild vaule: between 200 and 599 inclusive *** HTTP status code to return to the client when there is a network error with the external authorization service. --- # google-cloud-logging The `google-cloud-logging` plugin pushes request and response logs in batches to [Google Cloud Logging Service](https://cloud.google.com/logging?hl=en) and supports the customization of log formats. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `google-cloud-logging` plugin for different scenarios. To follow along with the examples, you should have a GCP account with active billing. You should also first obtain authentication credentials in GCP by completing the following steps: * Visit **IAM & Admin** to create a service account. * Assign the service account with the **Logs Writer** role, which assigns the account with `logging.logEntries.create` and `logging.logEntries.route` permissions. * Create a private key for the service account and download the credentials in JSON format. The credentials JSON file content should look similar to the following: ``` { "type": "service_account", "project_id": "api7ai-docs", "private_key_id": "6330a8c37b15a26d3fb4e9e3986f04c004826d1a", "private_key": "-----BEGIN PRIVATE KEY-----\nMIIEvwIBADANBgkqhkiG9w0BAQEFAASCBKkwggSlAgEAAoIBAQDYTl1QKxgClpgq\n1FyZNKZTq4os9AoXU+h/1gdngtc681xqMIWlwycrJ7Bo69L//7REyUKnuIOPgHU6\nPCp4rGFokxdXzBJC0+WsxwZ/FZoaqLAD5Fbs4BpZ9q2F8fKz07l9Da+Ul2lLlQq6\nEgij2NOh9ytBvFiYEAnMY5DDWFyoXWBB0OXfGEE6486+DcfG8gMWQ7rXKVbKNyA1\nJdbS63cDJNERLb6z8QsaZOqYZwaqIn6apEv9aadnNEU+4HrXrjxsoDtk7zLmsbtp\nUOpYVVSiYz2uYbUz3XRJjW+NAeyeVBK8tePbe1n5WHM4Sg1Mp1wYtaJknS5gmOXe\nxglMt4vTAgMBAAECggEAHzGZ6mRJ56GmcH1vRywyalw8JoR2ahZ7L+hX6VkTR0ND\nn2VqTf/pR6Nxy4fAG5QEKsFS1VOE1tk3I/6mP1XYtwHeEBbJcWK+kLP5CghoULzl\nTq0LeMikHu+uY6w8OUlVTS/UQtC+SxwVMbstlEGyhWERxjdu0VwL\nY/jb6DA123cqjHteEwOFuipG+GELKJGIjgNhzyRimowOsY6F+3WrDHZrf2sM7AlD\nLbjrA3MdvIe6rNC8zy7zf/didygjryrJpjiHkKsLIPIPbu0l5xENHd3TNWuVAg48\nhf4nRwyZ7q1RXgRYnp/SfPH1YB0p4+7D0xLQUd2OEQKBgQDxvOED6IQ3zxipW+uX\nX4c+6QxwnOCTY/oQOtCwmgPSvzIMSyoNCH0YY3sdoUmygSP0hmBFIaP\nBH6A5d3A06iMTUiAwEOp5JDQImqVTN+Sz/JBBOxCpjuW/dmG72MFlZBL161lY0g6\n79ku2xatxvncdJvcpEWqB4UBEQKBgQDlEV/Tapm950M+PYTtYHry1AYxGum+Eb2+\nNg9u5kWbgl6aWSgR/XsKQPTcsYX0gFSkrYhFrVwdruDeG9JYSCckH6FtCoa8yv5s\nMB+QR7VWJoa3ej7Hc0O6VUjwUfUkXuQRoFCEl8lFCZzugsjSw93xTeo6w3s9oaCB\neY9RXGn+owKBgQCMU/Tba/K04weR6MZOTSoZnveVt7u2U+cp3LqgigeGI29OK6Px\nhOf5bGZfwO0jLlJAVJin5tdtgK1FfUDPbPByqv2bnkLNj19zPikJSqG18QSmPsXa\nV9RtYgo0doNJF3tbFUQKTdRB8qW5oXSgofMVfCEiJ8uL6jVAVCwMk+jlwQKBgQCD\ntE6lbwhAcORvt81i8nMehRueRjwYpXi0Eb8j41AoTnf4RMTOOzDwP1LKRWOgpdyE\n5qWQclGhW3g9HD//tFSU537YBBJeIFTSfYTYXvJ7OyGAAtBvuu05CGosiuLo64o0\nPDmvUtpNUG6jkBzJWgaVBFhlOxnz4Kc5alwlyn3DAwKBgQCwNJsqb4pOjwjaJl/m\nePXpeX7YdVyFnBDbSQ1BFxDYGU12yTKRYqQVIB+VIIGN28acta1EPI8tF2ODG5az\nCBmgH5amLRHHCDYRKwrP+BTA39lK0pQEUP47RSzOdY82KQB13BW1uEZTcifjS9HN\niZPoV+OYHG5iJiiWEQi9/Q1AfQ==\n-----END PRIVATE KEY-----\n", "client_email": "api7-docs-log@api7ai-docs.iam.gserviceaccount.com", "client_id": "100920913890704420895", "auth_uri": "https://accounts.google.com/o/oauth2/auth", "token_uri": "https://oauth2.googleapis.com/token", "auth_provider_x509_cert_url": "https://www.googleapis.com/oauth2/v1/certs", "client_x509_cert_url": "https://www.googleapis.com/robot/v1/metadata/x509/api7-docs-log%40api7ai-docs.iam.gserviceaccount.com", "universe_domain": "googleapis.com" } ``` ### Configure Authentication Using `auth_config`[​](#configure-authentication-using-auth_config "Direct link to configure-authentication-using-auth_config") The following example demonstrates how you can configure the `google-cloud-logging` plugin on a route, which logs client requests and responses, as well as pushing logs to Google Cloud Logging. You will be using the `auth_config` option to configure GCP authentication details. Create a route with `google-cloud-logging` as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "google-cloud-logging-route", "uri": "/anything", "plugins": { "google-cloud-logging": { "auth_config": { "client_email": "api7-docs-logging@api7ai-docs.iam.gserviceaccount.com", "project_id": "api7ai-docs", "private_key": "-----BEGIN PRIVATE KEY-----\nMIIEvwIBADANBgkqhkiG9w0BAQEFAASCBKkwggSlAgEAAoIBAQDYTl1QKxgClpgq\n1FyZNKZTq4os9AoXU+h/1gdngtc681xqMIWlwycrJ7Bo69L//7REyUKnuIOPgHU6\nPCp4rGFokxdXzBJC0+WsxwZ/FZoaqLAD5Fbs4BpZ9q2F8fKz07l9Da+Ul2lLlQq6\nEgij2NOh9ytBvFiYEAnMY5DDWFyoXWBB0OXfGEE6486+DcfG8gMWQ7rXKVbKNyA1\nJdbS63cDJNERLb6z8QsaZOqYZwaqIn6apEv9aadnNEU+4HrXrjxsoDtk7zLmsbtp\nUOpYVVSiYz2uYbUz3XRJjW+NAeyeVBK8tePbe1n5WHM4SnS5gmOXe\nxglMt4vTAgMBAAECggEAHzGZ6mRJ56GmcH1vRywyalw8JoR2ahZ7L+hX6VkTR0ND\nn2VqTf/pR6Nxy4fAG5QEKsFS1VOE1tk3I/6mP1XYtwHeEBbJcWK+kLP5CghoULzl\nTq0LeMikHuI19FxH3HVwSV+uY6w8OUlVTS/UQtC+SxwVMbstlEGyhWERxjdu0VwL\nY/jb6DA123cqjHteEwOFuipG+GELKJGIjgNhzyRimowOsY6F+3WrDHZrf2sM7AlD\nLbjrA3MdvIe6rNC8zy7zf/didygjryrJpjiHkKsLIPIPbu0l5xENHd3TNWuVAg48\nhf4nRwyZ7q1RXgRYnp/SfPH1YB0p4+7D0xLQUd2xvOED6IQ3zxipW+uX\nX4c+6QxwnOCTY/oQOtCwmgPSvzIMSyoNCH0YY3sdoUmygS40v30OV8vP0hmBFIaP\nBH6A5d3A06iMTUiAwEOp5JDQImqVTN+Sz/JBBOxCpjuW/dmG72MFlZBL161lY0g6\n79ku2xatxvncdJvcpEWqB4UBEQKBgQDlEV/Tapm950M+PYTtYHry1AYxGum+Eb2+\nNg9u5kWbgl6aWSgR/XsKQPTcsYX0gFSkrYhFrVwdruDeG9JYSCckH6FtCoa8yv5s\nMB+QR7VWJoa3ej7Hc0O6VUjwUfUkXuQRoFCEl8lFCZzugsjSw93xTeo6w3s9oaCB\neY9RXGn+owKBgQCMU/Tba/K04weR6MZOTSoZnveVt7u2U+cp3LqgigeGI29OK6Px\nhOf5bGZfwO0jLlJAVJin5tdtgK1FfUDPbPByqv2bnkLNj19zPikJSqG18QSmPsXa\nV9RtYgo0doNJF3tbFUQKTdRB8qW5oXSgofMVfCEiJ8uL6jVAVCwMk+jlwQKBgQCD\ntE6lbwhAcORvt81i8nMehRueRjwYpXi0Eb8j41AoTnf4RMTOOzDwP1LKRWOgpdyE\n5qWQclGhW3g9HD//tFSU537YBBJeIFTSfYTYXvJ7OyGAAtBvuu05CGosiuLo64o0\nPDmvUtpNUG6jkBzJWgaVBFhlOxnz4Kc5alwlyn3DAwKBgQCwNJsqb4pOjwjaJl/m\nePXpeX7YdVyFnBDbSQ1BFxDYGU12yTKRYqQVIB+VIIGN28acta1EPI8tF2ODG5az\nCBmgH5amLRHHCDYRKwrP+BTA39lK0pQEUP47RSzOdY82KQB13BW1uEZTcifjS9HN\niZPoV+OYHG5iJiiWEQi9/Q1AfQ==\n-----END PRIVATE KEY-----\n", "token_uri": "https://oauth2.googleapis.com/token" } } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: google-cloud-logging-route plugins: google-cloud-logging: auth_config: client_email: "api7-docs-logging@api7ai-docs.iam.gserviceaccount.com" project_id: "api7ai-docs" private_key: | -----BEGIN PRIVATE KEY----- ... -----END PRIVATE KEY----- token_uri: "https://oauth2.googleapis.com/token" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD google-cloud-logging-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: google-cloud-logging-plugin-config spec: plugins: - name: google-cloud-logging config: auth_config: client_email: "api7-docs-logging@api7ai-docs.iam.gserviceaccount.com" project_id: "api7ai-docs" private_key: | -----BEGIN PRIVATE KEY----- ... -----END PRIVATE KEY----- token_uri: "https://oauth2.googleapis.com/token" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: google-cloud-logging-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: google-cloud-logging-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` google-cloud-logging-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: google-cloud-logging-route spec: ingressClassName: apisix http: - name: google-cloud-logging-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: google-cloud-logging config: auth_config: client_email: "api7-docs-logging@api7ai-docs.iam.gserviceaccount.com" project_id: "api7ai-docs" private_key: | -----BEGIN PRIVATE KEY----- ... -----END PRIVATE KEY----- token_uri: "https://oauth2.googleapis.com/token" ``` Apply the configuration: ``` kubectl apply -f google-cloud-logging-ic.yaml ``` ❶ Replace with your service account. ❷ Replace with your project ID. ❸ Replace with your private key. ❹ Replace with your token URI. Send a request to the route to generate a log entry: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to Google Cloud Logs Explorer, you should see a log entry corresponding to your request, similar to the following: ``` { "insertId": "5400340ea330b35f2d557da2cbb9e88d", "jsonPayload": { "service_id": "", "route_id": "google-cloud-logging-route" }, "httpRequest": { "requestMethod": "GET", "requestUrl": "http://127.0.0.1:9080/anything", "requestSize": "85", "status": 200, "responseSize": "615", "userAgent": "curl/8.6.0", "remoteIp": "192.168.107.1", "serverIp": "54.86.137.185:80", "latency": "1.083s" }, "resource": { "type": "global", "labels": { "project_id": "api7ai-docs" } }, "timestamp": "2025-02-07T07:39:51.859Z", "labels": { "source": "apache-apisix-google-cloud-logging" }, "logName": "projects/api7ai-docs/logs/apisix.apache.org%2Flogs", "receiveTimestamp": "2025-02-07T07:39:58.012811475Z" } ``` ### Configure Authentication Using `auth_file`[​](#configure-authentication-using-auth_file "Direct link to configure-authentication-using-auth_file") The following example demonstrates how you can configure the `google-cloud-logging` plugin on a route, which logs client requests and responses, as well as pushing logs to Google Cloud Logging. You will be using the `auth_file` option to configure GCP authentication details. Copy the previously downloaded GCP service account credentials JSON file to a location accessible for APISIX. If you are running APISIX in Docker, you should copy the file into the container, for instance, to `/usr/local/apisix/conf/gcp-logging-auth.json`. Create a route with `google-cloud-logging` as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "google-cloud-logging-route", "uri": "/anything", "plugins": { "google-cloud-logging": { "auth_file": "/usr/local/apisix/conf/gcp-logging-auth.json" } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: google-cloud-logging-route plugins: google-cloud-logging: auth_file: "/usr/local/apisix/conf/gcp-logging-auth.json" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD google-cloud-logging-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: google-cloud-logging-plugin-config spec: plugins: - name: google-cloud-logging config: auth_file: "/usr/local/apisix/conf/gcp-logging-auth.json" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: google-cloud-logging-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: google-cloud-logging-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` google-cloud-logging-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: google-cloud-logging-route spec: ingressClassName: apisix http: - name: google-cloud-logging-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: google-cloud-logging config: auth_file: "/usr/local/apisix/conf/gcp-logging-auth.json" ``` Apply the configuration: ``` kubectl apply -f google-cloud-logging-ic.yaml ``` ❶ Replace with your GCP service account credentials JSON file path. Send a request to the route to generate a log entry: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to Google Cloud Logs Explorer, you should see a log entry corresponding to your request, similar to the following: ``` { "insertId": "5400340ea330b35f2d557da2cbb9e88d", "jsonPayload": { "service_id": "", "route_id": "google-cloud-logging-route" }, "httpRequest": { "requestMethod": "GET", "requestUrl": "http://127.0.0.1:9080/anything", "requestSize": "85", "status": 200, "responseSize": "615", "userAgent": "curl/8.6.0", "remoteIp": "192.168.107.1", "serverIp": "54.86.137.185:80", "latency": "1.083s" }, "resource": { "type": "global", "labels": { "project_id": "api7ai-docs" } }, "timestamp": "2025-02-07T08:25:11.325Z", "labels": { "source": "apache-apisix-google-cloud-logging" }, "logName": "projects/api7ai-docs/logs/apisix.apache.org%2Flogs", "receiveTimestamp": "2025-02-07T08:25:11.423190575Z" } ``` ### Customize Log Format With Plugin Metadata[​](#customize-log-format-with-plugin-metadata "Direct link to Customize Log Format With Plugin Metadata") The following example demonstrates how you can customize log format using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) and [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) to log specific headers from request and response. In APISIX, [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) is used to configure the common metadata fields of all plugin instances of the same plugin. It is useful when a plugin is enabled across multiple resources and requires a universal update to their metadata fields. First, create a route with `google-cloud-logging` as such, and replace with your credentials: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "google-cloud-logging-route", "uri": "/anything", "plugins": { "google-cloud-logging": { "auth_config": { "client_email": "api7-docs-logging@api7ai-docs.iam.gserviceaccount.com", "project_id": "api7ai-docs", "private_key": "-----BEGIN PRIVATE KEY-----\nMIIEvwIBADANBgkqhkiG9w0BAQEFAASCBKkwggSlAgEAAoIBAQDYTl1QKxgClpgq\n1FyZNKZTq4os9AoXU+h/1gdngtc681xqMIWlwycrJ7Bo69L//7REyUKnuIOPgHU6\nPCp4rGFokxdXzBJC0+WsxwZ/FZoaqLAD5Fbs4BpZ9q2F8fKz07l9Da+Ul2lLlQq6\nEgij2NOh9ytBvFiYEAnMY5DDWFyoXWBB0OXfGEE6486+DcfG8gMWQ7rXKVbKNyA1\nJdbS63cDJNERLb6z8QsaZOqYZwaqIn6apEv9aadnNEU+4HrXrjxsoDtk7zLmsbtp\nUOpYVVSiYz2uYbUz3XRJjW+NAeyeVBK8tePbe1n5WHM4SnS5gmOXe\nxglMt4vTAgMBAAECggEAHzGZ6mRJ56GmcH1vRywyalw8JoR2ahZ7L+hX6VkTR0ND\nn2VqTf/pR6Nxy4fAG5QEKsFS1VOE1tk3I/6mP1XYtwHeEBbJcWK+kLP5CghoULzl\nTq0LeMikHuI19FxH3HVwSV+uY6w8OUlVTS/UQtC+SxwVMbstlEGyhWERxjdu0VwL\nY/jb6DA123cqjHteEwOFuipG+GELKJGIjgNhzyRimowOsY6F+3WrDHZrf2sM7AlD\nLbjrA3MdvIe6rNC8zy7zf/didygjryrJpjiHkKsLIPIPbu0l5xENHd3TNWuVAg48\nhf4nRwyZ7q1RXgRYnp/SfPH1YB0p4+7D0xLQUd2xvOED6IQ3zxipW+uX\nX4c+6QxwnOCTY/oQOtCwmgPSvzIMSyoNCH0YY3sdoUmygS40v30OV8vP0hmBFIaP\nBH6A5d3A06iMTUiAwEOp5JDQImqVTN+Sz/JBBOxCpjuW/dmG72MFlZBL161lY0g6\n79ku2xatxvncdJvcpEWqB4UBEQKBgQDlEV/Tapm950M+PYTtYHry1AYxGum+Eb2+\nNg9u5kWbgl6aWSgR/XsKQPTcsYX0gFSkrYhFrVwdruDeG9JYSCckH6FtCoa8yv5s\nMB+QR7VWJoa3ej7Hc0O6VUjwUfUkXuQRoFCEl8lFCZzugsjSw93xTeo6w3s9oaCB\neY9RXGn+owKBgQCMU/Tba/K04weR6MZOTSoZnveVt7u2U+cp3LqgigeGI29OK6Px\nhOf5bGZfwO0jLlJAVJin5tdtgK1FfUDPbPByqv2bnkLNj19zPikJSqG18QSmPsXa\nV9RtYgo0doNJF3tbFUQKTdRB8qW5oXSgofMVfCEiJ8uL6jVAVCwMk+jlwQKBgQCD\ntE6lbwhAcORvt81i8nMehRueRjwYpXi0Eb8j41AoTnf4RMTOOzDwP1LKRWOgpdyE\n5qWQclGhW3g9HD//tFSU537YBBJeIFTSfYTYXvJ7OyGAAtBvuu05CGosiuLo64o0\nPDmvUtpNUG6jkBzJWgaVBFhlOxnz4Kc5alwlyn3DAwKBgQCwNJsqb4pOjwjaJl/m\nePXpeX7YdVyFnBDbSQ1BFxDYGU12yTKRYqQVIB+VIIGN28acta1EPI8tF2ODG5az\nCBmgH5amLRHHCDYRKwrP+BTA39lK0pQEUP47RSzOdY82KQB13BW1uEZTcifjS9HN\niZPoV+OYHG5iJiiWEQi9/Q1AfQ==\n-----END PRIVATE KEY-----\n", "token_uri": "https://oauth2.googleapis.com/token" } } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` Next, configure the plugin metadata for `google-cloud-logging`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/google-cloud-logging" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "log_format": { "host": "$host", "@timestamp": "$time_iso8601", "client_ip": "$remote_addr", } }' ``` adc.yaml ``` plugin_metadata: - name: google-cloud-logging log_format: host: "$host" "@timestamp": "$time_iso8601" client_ip: "$remote_addr" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` google-cloud-logging-metadata.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: service: name: apisix-admin port: 9180 auth: type: AdminKey adminKey: value: edd1c9f034335f136f87ad84b625c8f1 pluginMetadata: google-cloud-logging: log_format: host: "$host" "@timestamp": "$time_iso8601" client_ip: "$remote_addr" ``` Apply the configuration: ``` kubectl apply -f google-cloud-logging-metadata.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to Google Cloud Logs Explorer, you should see a log entry corresponding to your request, similar to the following: ``` { "@timestamp":"2025-02-07T09:10:42+00:00", "client_ip":"192.168.107.1", "host":"127.0.0.1", "route_id":"google-cloud-logging-route" } ``` The log format configured in plugin metadata is effective for all instances of `google-cloud-logging` if the log format is not specifically specified on the individual instance. If you specifically configure the log format in the `google-cloud-logging` plugin on the route: ``` curl "http://127.0.0.1:9180/apisix/admin/routes/google-cloud-logging-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "google-cloud-logging": { "log_format": { "host": "$host", "@timestamp": "$time_iso8601", "client_ip": "$remote_addr", "env": "$http_env", "resp_content_type": "$sent_http_Content_Type" } } } }' ``` ❶ log the custom request header `env`. ❷ log the response header `Content-Type`. Send a request to the route with the `env` header: ``` curl -i "http://127.0.0.1:9080/anything" -H "env: dev" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to Google Cloud Logs Explorer, you should see a log entry corresponding to your request, similar to the following: ``` { "@timestamp":"2025-02-07T09:38:55+00:00", "client_ip":"192.168.107.1", "host":"127.0.0.1", "env":"dev", "resp_content_type":"application/json", "route_id":"google-cloud-logging-route" } ``` The configuration of log format on the route has taken precedence over the log format configured on the `google-cloud-logging` plugin metadata. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * auth\_config object *** Authentication configurations. At least one of the `auth_config` and `auth_file` should be provided. * client\_email string required *** Email address of the Google Cloud service account. * private\_key string required *** Private key of the Google Cloud service account. The value is encrypted with AES before being stored in etcd. * project\_id string required *** Project ID in the Google Cloud service account. * token\_uri string required default: `https://oauth2.googleapis.com/token` *** Token URI of the Google Cloud service account. * entries\_uri string default: `https://logging.googleapis.com/v2/entries:write` *** Google Cloud Logging Service API. * scope array\[string] default: `["https://www.googleapis.com/auth/logging.read", "https://www.googleapis.com/auth/logging.write", "https://www.googleapis.com/auth/logging.admin", "https://www.googleapis.com/auth/cloud-platform"]` *** Access scopes of the Google Cloud service account. See [OAuth 2.0 Scopes for Google APIs](https://developers.google.com/identity/protocols/oauth2/scopes#logging). Can also be specified as `scopes`. * auth\_file string *** Path to the Google Cloud service account authentication JSON file. At least one of the `auth_config` and `auth_file` should be provided. * ssl\_verify boolean default: `true` *** If true, verify the server's SSL certificate. * resource object default: `{"type": "global"}` *** Google monitored resource composed of `type` and optionally `labels`, for example: ```json { "type": "gce_instance", "labels": { "project_id": "my-project", "instance_id": "12345678901234", "zone": "us-central1-a" } } ``` See [MonitoredResource](https://cloud.google.com/logging/docs/reference/v2/rest/v2/MonitoredResource) for more details. * log\_id string default: `apisix.apache.org%2Flogs` *** Google Cloud logging ID. See [LogEntry](https://cloud.google.com/logging/docs/reference/v2/rest/v2/LogEntry) for details. * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. You can also configure log format on a global scale using the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md), which configures the log format for all `google-cloud-logging` plugin instances. If the log format configured on the individual plugin instance differs from the log format configured on plugin metadata, the log format configured on the individual plugin instance takes precedence. See the [example](https://docs.api7.ai/hub/google-cloud-logging.md#customize-log-format-with-plugin-metadata) for more details. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * name string default: `google-cloud-logging` *** Unique identifier of the plugin for the batch processor. If you use [Prometheus](https://docs.api7.ai/hub/prometheus.md) to monitor APISIX metrics, the name is exported in `apisix_batch_process_entries`. * batch\_max\_size integer default: `1000` vaild vaule: greater than 0 *** The number of log entries allowed in one batch. Once reached, the batch will be sent to the logging service. Setting this parameter to 1 means immediate processing. * inactive\_timeout integer default: `5` vaild vaule: greater than 0 *** The maximum time in seconds to wait for new logs before sending the batch to the logging service. The value should be smaller than `buffer_duration`. * buffer\_duration integer default: `60` vaild vaule: greater than 0 *** The maximum time in seconds from the earliest entry allowed before sending the batch to the logging service. * retry\_delay integer default: `1` vaild vaule: greater than or equal to 0 *** The time interval in seconds to retry sending the batch to the logging service if the batch was not successfully sent. * max\_retry\_count integer default: `0` vaild vaule: greater than or equal to 0 *** The maximum number of unsuccessful retries allowed before dropping the log entries. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * max\_pending\_entries integer default: `` `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 and 3.10.6; none in API7 Enterprise 3.9.18 and 3.10.5 `` vaild vaule: greater than or equal to 1 *** Maximum number of entries waiting in the batch processor. New entries are discarded when the backlog reaches the limit. Introduced in API7 Enterprise 3.8.17 and APISIX 3.15.0. The default changed to `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line. In API7 Enterprise 3.9.18 and 3.10.5, and in earlier APISIX versions, omitting the parameter leaves the backlog unlimited. See [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md#configure-the-pending-entry-limit) for sizing and verification guidance. --- # graphql-limit-count The `graphql-limit-count` plugin uses fixed windows to limit the accumulated cost of GraphQL [queries](https://graphql.org/learn/queries/) and [mutations](https://graphql.org/learn/queries#mutations). Query depth is the default cost, preserving the plugin's original behavior. API7 Enterprise also provides `complexity` and `node_quantifier` strategies that account for the work a document requests. In GraphQL, the depth refers to the number of nesting levels in a query or mutation. The following is an example query with a depth of 3: ``` { a { b { c } } } ``` With the default depth strategy, the plugin consumes a quota of depth within each time interval. For example, if the quota is 4 in a 30-second interval, a request with depth 3 is allowed and leaves 1. A request with depth 2 during the same interval is rejected. The plugin accepts `POST` requests with either a JSON body containing a `query` field or an `application/graphql` body containing the GraphQL document. Fragments contribute to the calculated query depth. Unsupported methods return `405 Method Not Allowed`; unreadable, malformed, or invalid GraphQL requests return `400 Bad Request`. APISIX reads up to 1 MiB of GraphQL request data by default. To change this limit, configure `graphql.max_size` in `config.yaml` and reload APISIX: config.yaml ``` graphql: max_size: 1048576 ``` ## Local vs Redis Rate Limiting[​](#local-vs-redis-rate-limiting "Direct link to Local vs Redis Rate Limiting") The `graphql-limit-count` plugin supports two modes of rate limiting: * **Local rate limiting**: Limits are enforced independently on each gateway instance. Each instance maintains its own counters, so the effective limit is roughly (limit × number of instances) when traffic is spread across instances. This is the default when no `policy` is set or when `policy` is `local`. * **Redis-based rate limiting**: Limits are shared across all gateway instances through Redis. All instances share the same quota, so the configured limit applies to all gateway instances. ## Query Cost[​](#query-cost "Direct link to Query Cost") The `complexity` and `node_quantifier` strategies and their supporting fields were introduced in API7 Enterprise 3.10.6. By default, a request is charged the depth of its query. `cost_strategy` selects a different cost model, so a request consumes quota in proportion to how much work it asks the upstream for: * `depth` charges the selection nesting depth. This is what the plugin has always done and remains the default, so an existing configuration keeps its behavior after an upgrade. * `complexity` computes the raw score from the nodes the query resolves. Each node contributes `(sum of its children) × mul + add`, where `add` and `mul` default to `1`. * `node_quantifier` computes the raw score only from nodes whose matching cost decoration names a usable quantifier in `mul_arguments`. For example, a decoration with `mul_arguments: ["first"]` carries `first: 10` as the multiplier for deeper quantified nodes. If no node has both a matching decoration and a usable quantifier, the document's raw score is `0`. The default `score_factor` produces a charged cost of `1`; a factor greater than `100` increases it after the `0.01` adjustment. The plugin turns the raw strategy score into the integer charged against the quota. For `complexity` and `node_quantifier`, it adds `0.01` to the raw score before applying `score_factor`, then rounds the result up. With the default factor of `1`, an integer raw score of `3` is therefore charged as `4`. The `depth` strategy skips the `0.01` adjustment but still applies the factor and rounds up. `max_cost` rejects a query whose charged cost exceeds the configured value with `403 Forbidden` before it reaches the upstream. The plugin charges the quota before applying this check, so a rejected over-cost query still consumes its computed amount. When `show_limit_quota_header` is enabled, `X-Graphql-Query-Cost` reports that amount. With `resolve_variables` enabled, which is the default, the plugin resolves supplied GraphQL variables, variable defaults declared by the operation, and argument defaults from the upstream schema before computing cost. Turning it off treats `first: $n` like an absent argument and can assign too little cost to a query whose quantifier is supplied through a variable. Matching a cost decoration against the query requires the upstream schema. Each gateway worker introspects a Service with decorations on the first applicable request and caches the schema until the plugin reloads. A route with no decorations is never introspected. In that case, `complexity` counts each node with the default weight, while `node_quantifier` produces a raw score of `0`; its charged cost follows the adjustment and scaling described above. Set `introspection_endpoint` when the introspection endpoint is not the upstream itself, and `introspection_headers` when it requires credentials. The credentials come from the configuration rather than from the request because each cached schema is reused by callers handled by that worker. ### Cost Decorations[​](#cost-decorations "Direct link to Cost Decorations") A decoration adjusts what one position in the upstream schema contributes to the cost. Decorations are managed on the Service as `graphql_cost_decorations`, so they are shared by every route under it and can be changed without editing the routes that carry the plugin. A decoration names a `field_path`, which can identify a GraphQL type such as `Product`, a type and field such as `Product.name`, or a chain such as `Query.products.nodes`. It adjusts the node it matches: | Field | Effect | | --------------- | ---------------------------------------------------------------------------------------------------------------------------------- | | `add_value` | Added to the node's own cost. | | `mul_value` | Multiplies the cost of the node's children. | | `add_arguments` | Names arguments whose values are added to the node's own cost. | | `mul_arguments` | Names arguments whose values multiply descendant cost. Under `node_quantifier`, the multiplier carries to deeper quantified nodes. | A `field_path` can only be decorated once per Service. ## Examples[​](#examples "Direct link to Examples") The examples below use [GitHub GraphQL API](https://docs.github.com/en/graphql) endpoint as an upstream and demonstrate how you can configure `graphql-limit-count` for different scenarios. To follow along, create a GitHub [personal access token](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens) with the appropriate scopes for the resources you want to interact with. ### Apply Rate Limiting by Remote Address[​](#apply-rate-limiting-by-remote-address "Direct link to Apply Rate Limiting by Remote Address") The following example demonstrates the rate limiting of GraphQL requests by a single variable, `remote_addr`. Create a route with `graphql-limit-count` plugin that allows for a quota of depth 2 within a 30-second window per remote address: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-limit-count-route", "uri": "/graphql", "plugins": { "graphql-limit-count": { "count": 2, "time_window": 30, "rejected_code": 429, "key_type": "var", "key": "remote_addr", "policy": "local" } }, "upstream": { "type": "roundrobin", "pass_host": "node", "scheme": "https", "nodes": { "api.github.com:443": 1 } } }' ``` adc.yaml ``` services: - name: graphql-service routes: - uris: - /graphql name: graphql-limit-count-route plugins: graphql-limit-count: count: 2 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local upstream: type: roundrobin scheme: https nodes: - host: api.github.com port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD graphql-limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: github-graphql-external-domain spec: type: ExternalName externalName: api.github.com --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: github-graphql-https spec: targetRefs: - name: github-graphql-external-domain kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: graphql-limit-count-plugin-config spec: plugins: - name: graphql-limit-count config: count: 2 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: graphql-limit-count-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /graphql filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: graphql-limit-count-plugin-config backendRefs: - name: github-graphql-external-domain port: 443 ``` graphql-limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: github-graphql-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: api.github.com port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: graphql-limit-count-route spec: ingressClassName: apisix http: - name: graphql-limit-count-route match: paths: - /graphql upstreams: - name: github-graphql-external-domain plugins: - name: graphql-limit-count enable: true config: count: 2 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local ``` Apply the configuration to your cluster: ``` kubectl apply -f graphql-limit-count-ic.yaml ``` #### Verify with GraphQL Query[​](#verify-with-graphql-query "Direct link to Verify with GraphQL Query") Send a request with a GraphQL query of depth 2 to verify: ``` curl -i "http://127.0.0.1:9080/graphql" -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${GH_ACCESS_TOKEN}" \ -d '{"query": "query {viewer{login}}"}' ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body. The request has consumed all the quota allowed for the time window. If you send the request again within the same 30-second time interval, you should receive an `HTTP/1.1 429 Too Many Requests` response, indicating the request surpasses the quota threshold. #### Verify with GraphQL Mutation[​](#verify-with-graphql-mutation "Direct link to Verify with GraphQL Mutation") You can also send a request with a GraphQL mutation of depth 3 to verify: ``` curl -i "http://127.0.0.1:9080/graphql" -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${GH_ACCESS_TOKEN}" \ -d '{"query": "mutation AddReactionToIssue {addReaction(input:{subjectId:\"MDU6SXNzdWUyMzEzOTE1NTE=\",content:HOORAY}) {reaction {content} subject {id}}}"}' ``` You should see an `HTTP/1.1 429 Too Many Requests` response at any time, as depth 3 always surpasses the quota of depth 2. ### Apply Rate Limiting by Remote Address and Consumer Name[​](#apply-rate-limiting-by-remote-address-and-consumer-name "Direct link to Apply Rate Limiting by Remote Address and Consumer Name") The following example demonstrates the rate limiting of GraphQL requests by a combination of variables, `remote_addr` and `consumer_name`. It allows for a quota of depth 2 within a 30-second window per remote address and for each [consumer](https://docs.api7.ai/apisix/key-concepts/consumers.md). * Admin API * ADC * Ingress Controller Create a consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Create a second consumer `jane`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jane" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jane/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-key-auth", "plugins": { "key-auth": { "key": "jane-key" } } }' ``` Create a route with `key-auth` and `graphql-limit-count` plugins: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-limit-count-route", "uri": "/graphql", "plugins": { "key-auth": {}, "graphql-limit-count": { "count": 2, "time_window": 30, "rejected_code": 429, "policy": "local", "key_type": "var_combination", "key": "$remote_addr $consumer_name" } }, "upstream": { "type": "roundrobin", "pass_host": "node", "scheme": "https", "nodes": { "api.github.com:443": 1 } } }' ``` Create two consumers and a route that enables rate limiting by consumers: adc.yaml ``` consumers: - username: john credentials: - name: key-auth type: key-auth config: key: john-key - username: jane credentials: - name: key-auth type: key-auth config: key: jane-key services: - name: graphql-limit-service routes: - name: graphql-limit-count-route uris: - /graphql plugins: key-auth: {} graphql-limit-count: count: 2 time_window: 30 rejected_code: 429 policy: local key_type: var_combination key: "$remote_addr $consumer_name" upstream: type: roundrobin scheme: https nodes: - host: api.github.com port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create two consumers and a route that enables rate limiting by consumers: * Gateway API * APISIX CRD graphql-limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: john-key --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jane spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: jane-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: github-graphql-external-domain spec: type: ExternalName externalName: api.github.com --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: github-graphql-https spec: targetRefs: - name: github-graphql-external-domain kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: graphql-limit-count-plugin-config spec: plugins: - name: key-auth config: _meta: disable: false - name: graphql-limit-count config: count: 2 time_window: 30 rejected_code: 429 policy: local key_type: var_combination key: "$remote_addr $consumer_name" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: graphql-limit-count-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /graphql filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: graphql-limit-count-plugin-config backendRefs: - name: github-graphql-external-domain port: 443 ``` graphql-limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jane spec: ingressClassName: apisix authParameter: keyAuth: value: key: jane-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: github-graphql-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: api.github.com port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: graphql-limit-count-route spec: ingressClassName: apisix http: - name: graphql-limit-count-route match: paths: - /graphql upstreams: - name: github-graphql-external-domain plugins: - name: key-auth enable: true - name: graphql-limit-count enable: true config: count: 2 time_window: 30 rejected_code: 429 policy: local key_type: var_combination key: "$remote_addr $consumer_name" ``` Apply the configuration to your cluster: ``` kubectl apply -f graphql-limit-count-ic.yaml ``` ❶ `key-auth`: enable key authentication on the route. ❷ `key_type`: set to `var_combination` to interpret the `key` as a combination of variables. ❸ `key`: set to `$remote_addr $consumer_name` to apply rate limiting quota by remote address and consumer. Send a request with a GraphQL query of depth 2 as the consumer `jane`: ``` curl -i "http://127.0.0.1:9080/graphql" -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${GH_ACCESS_TOKEN}" \ -H 'apikey: jane-key' \ -d '{"query": "query {viewer{login}}"}' ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body. This request has consumed all the quota set for the time window. If you send the same request as the consumer `jane` within the same 30-second time interval, you should receive an `HTTP/1.1 429 Too Many Requests` response, indicating the request surpasses the quota threshold. Send the same request as the consumer `john` within the same 30-second time interval: ``` curl -i "http://127.0.0.1:9080/graphql" -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${GH_ACCESS_TOKEN}" \ -H 'apikey: john-key' \ -d '{"query": "query {viewer{login}}"}' ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body, indicating the request is not rate limited. Send the same request as the consumer `john` again within the same 30-second time interval, you should receive an `HTTP/1.1 429 Too Many Requests` response. This verifies the plugin rate limits by the combination of variables, `remote_addr` and `consumer_name`. ### Share Quota among Routes[​](#share-quota-among-routes "Direct link to Share Quota among Routes") The following example demonstrates the sharing of GraphQL rate limiting quota among multiple routes by configuring the `group` of the `graphql-limit-count` plugin. Note that the configurations of the `graphql-limit-count` plugin of the same `group` should be identical. To avoid update anomalies and repetitive configurations, you can create a [service](https://docs.api7.ai/apisix/key-concepts/services.md) with `graphql-limit-count` plugin and upstream for routes to connect to. * Admin API * ADC * Ingress Controller Create a service: ``` curl "http://127.0.0.1:9180/apisix/admin/services" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-limit-count-service", "plugins": { "graphql-limit-count": { "count": 2, "time_window": 30, "rejected_code": 429, "policy": "local", "group": "srv1" } }, "upstream": { "type": "roundrobin", "pass_host": "node", "scheme": "https", "nodes": { "api.github.com:443": 1 } } }' ``` Create two routes and configure their `service_id` to be `graphql-limit-count-service`, so that they share the same configurations for the plugin and upstream: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-limit-count-route-1", "service_id": "graphql-limit-count-service", "uri": "/graphql1", "plugins": { "proxy-rewrite": { "uri": "/graphql" } } }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-limit-count-route-2", "service_id": "graphql-limit-count-service", "uri": "/graphql2", "plugins": { "proxy-rewrite": { "uri": "/graphql" } } }' ``` Create a service with two routes that share the same rate limiting quota: adc.yaml ``` services: - name: graphql-limit-count-service plugins: graphql-limit-count: count: 2 time_window: 30 rejected_code: 429 policy: local group: srv1 routes: - name: graphql-limit-count-route-1 uris: - /graphql1 plugins: proxy-rewrite: uri: /graphql - name: graphql-limit-count-route-2 uris: - /graphql2 plugins: proxy-rewrite: uri: /graphql upstream: type: roundrobin scheme: https nodes: - host: api.github.com port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create two HTTPRoute resources that reference the same PluginConfig to share quota: graphql-limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: github-graphql-external-domain spec: type: ExternalName externalName: api.github.com --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: github-graphql-https spec: targetRefs: - name: github-graphql-external-domain kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: graphql-limit-count-plugin-config spec: plugins: - name: graphql-limit-count config: count: 2 time_window: 30 rejected_code: 429 policy: local group: srv1 - name: proxy-rewrite config: uri: /graphql --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: graphql-limit-count-route-1 spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /graphql1 filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: graphql-limit-count-plugin-config backendRefs: - name: github-graphql-external-domain port: 443 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: graphql-limit-count-route-2 spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /graphql2 filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: graphql-limit-count-plugin-config backendRefs: - name: github-graphql-external-domain port: 443 ``` Create an ApisixRoute with multiple paths that share the same plugin configuration: graphql-limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: github-graphql-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: api.github.com port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: graphql-limit-count-shared-route spec: ingressClassName: apisix http: - name: graphql-limit-count-shared match: paths: - /graphql1 - /graphql2 upstreams: - name: github-graphql-external-domain plugins: - name: proxy-rewrite enable: true config: uri: /graphql - name: graphql-limit-count enable: true config: count: 2 time_window: 30 rejected_code: 429 policy: local group: srv1 ``` Apply the configuration to your cluster: ``` kubectl apply -f graphql-limit-count-ic.yaml ``` note The [`proxy-rewrite`](https://docs.api7.ai/hub/proxy-rewrite.md) plugin is used to rewrite the URI to `/graphql` so that requests are forwarded to the correct endpoint. Send a request with a GraphQL query of depth 2 to route `/graphql1`: ``` curl -i "http://127.0.0.1:9080/graphql1" -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${GH_ACCESS_TOKEN}" \ -d '{"query": "query {viewer{login}}"}' ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body. Send the same query of depth 2 to route `/graphql2` within the same 30-second time interval: ``` curl -i "http://127.0.0.1:9080/graphql2" -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${GH_ACCESS_TOKEN}" \ -d '{"query": "query {viewer{login}}"}' ``` You should receive an `HTTP/1.1 429 Too Many Requests` response, which verifies the two routes share the same rate limiting quota. ### Share Quota Among Gateway Nodes with a Redis Server[​](#share-quota-among-gateway-nodes-with-a-redis-server "Direct link to Share Quota Among Gateway Nodes with a Redis Server") The following example demonstrates the rate limiting of GraphQL requests across multiple gateway nodes with a Redis server, such that different gateway nodes share the same rate limiting quota. * Admin API * ADC * Ingress Controller Create a route with the following configurations in the gateway group: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-limit-count-route", "uri": "/graphql", "plugins": { "graphql-limit-count": { "count": 2, "time_window": 30, "rejected_code": 429, "key": "remote_addr", "policy": "redis", "redis_host": "192.168.xxx.xxx", "redis_port": 6379, "redis_password": "p@ssw0rd", "redis_database": 1 } }, "upstream": { "type": "roundrobin", "pass_host": "node", "scheme": "https", "nodes": { "api.github.com:443": 1 } } }' ``` Create a route with Redis-based rate limiting: adc.yaml ``` services: - name: graphql-redis-limit-service routes: - name: graphql-redis-limit-route uris: - /graphql plugins: graphql-limit-count: count: 2 time_window: 30 rejected_code: 429 key: remote_addr policy: redis redis_host: "192.168.xxx.xxx" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 upstream: type: roundrobin scheme: https nodes: - host: api.github.com port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD graphql-limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: github-graphql-external-domain spec: type: ExternalName externalName: api.github.com --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: github-graphql-https spec: targetRefs: - name: github-graphql-external-domain kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: graphql-limit-count-redis-plugin-config spec: plugins: - name: graphql-limit-count config: count: 2 time_window: 30 rejected_code: 429 key: remote_addr policy: redis redis_host: "redis-service.aic.svc" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: graphql-redis-limit-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /graphql filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: graphql-limit-count-redis-plugin-config backendRefs: - name: github-graphql-external-domain port: 443 ``` graphql-limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: github-graphql-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: api.github.com port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: graphql-redis-limit-route spec: ingressClassName: apisix http: - name: graphql-redis-limit-route match: paths: - /graphql upstreams: - name: github-graphql-external-domain plugins: - name: graphql-limit-count enable: true config: count: 2 time_window: 30 rejected_code: 429 key: remote_addr policy: redis redis_host: "redis-service.aic.svc" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 ``` Apply the configuration to your cluster: ``` kubectl apply -f graphql-limit-count-ic.yaml ``` ❶ `policy`: set to `redis` to use a Redis instance for rate limiting. ❷ `redis_host`: set to Redis instance IP address. ❸ `redis_port`: set to Redis instance listening port. ❹ `redis_password`: set to the password of the Redis instance, if any. ❺ `redis_database`: set to the database number in the Redis instance. Send a request with a GraphQL query of depth 2 to a gateway instance: ``` curl -i "http://127.0.0.1:9080/graphql" -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${GH_ACCESS_TOKEN}" \ -d '{"query": "query {viewer{login}}"}' ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body. Send the same request to a different gateway instance within the same 30-second time interval, you should receive an `HTTP/1.1 429 Too Many Requests` response, verifying routes configured in different gateway nodes share the same quota. ### Share Quota Among Gateway Nodes with a Redis Cluster[​](#share-quota-among-gateway-nodes-with-a-redis-cluster "Direct link to Share Quota Among Gateway Nodes with a Redis Cluster") You can also use a Redis cluster to apply the same quota across multiple gateway nodes, such that different gateway nodes share the same rate limiting quota. Ensure that your Redis instances are running in [cluster mode](https://redis.io/docs/management/scaling/#create-and-use-a-redis-cluster). A minimum of two nodes are required for the `graphql-limit-count` plugin configurations. * Admin API * ADC * Ingress Controller Create a route with the following configurations in the gateway group: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-limit-count-route", "uri": "/graphql", "plugins": { "graphql-limit-count": { "count": 2, "time_window": 30, "rejected_code": 429, "key": "remote_addr", "policy": "redis-cluster", "redis_cluster_nodes": [ "192.168.xxx.xxx:6379", "192.168.xxx.xxx:16379" ], "redis_password": "p@ssw0rd", "redis_cluster_name": "redis-cluster-1", "redis_cluster_ssl": true } }, "upstream": { "type": "roundrobin", "pass_host": "node", "scheme": "https", "nodes": { "api.github.com:443": 1 } } }' ``` Create a route with Redis cluster-based rate limiting: adc.yaml ``` services: - name: graphql-redis-cluster-limit-service routes: - name: graphql-redis-cluster-limit-route uris: - /graphql plugins: graphql-limit-count: count: 2 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-cluster redis_cluster_nodes: - "192.168.xxx.xxx:6379" - "192.168.xxx.xxx:16379" redis_password: "p@ssw0rd" redis_cluster_name: redis-cluster-1 redis_cluster_ssl: true upstream: type: roundrobin scheme: https nodes: - host: api.github.com port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD graphql-limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: github-graphql-external-domain spec: type: ExternalName externalName: api.github.com --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: github-graphql-https spec: targetRefs: - name: github-graphql-external-domain kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: graphql-limit-count-redis-cluster-plugin-config spec: plugins: - name: graphql-limit-count config: count: 2 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-cluster redis_cluster_nodes: - "redis-cluster-0.redis-cluster.aic.svc:6379" - "redis-cluster-1.redis-cluster.aic.svc:6379" redis_password: "p@ssw0rd" redis_cluster_name: redis-cluster-1 redis_cluster_ssl: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: graphql-redis-cluster-limit-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /graphql filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: graphql-limit-count-redis-cluster-plugin-config backendRefs: - name: github-graphql-external-domain port: 443 ``` graphql-limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: github-graphql-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: api.github.com port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: graphql-redis-cluster-limit-route spec: ingressClassName: apisix http: - name: graphql-redis-cluster-limit-route match: paths: - /graphql upstreams: - name: github-graphql-external-domain plugins: - name: graphql-limit-count enable: true config: count: 2 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-cluster redis_cluster_nodes: - "redis-cluster-0.redis-cluster.aic.svc:6379" - "redis-cluster-1.redis-cluster.aic.svc:6379" redis_password: "p@ssw0rd" redis_cluster_name: redis-cluster-1 redis_cluster_ssl: true ``` Apply the configuration to your cluster: ``` kubectl apply -f graphql-limit-count-ic.yaml ``` ❶ `policy`: set to `redis-cluster` to use a Redis cluster for rate limiting. ❷ `redis_cluster_nodes`: set to Redis node addresses in the Redis cluster. ❸ `redis_password`: set to the password of the Redis cluster, if any. ❹ `redis_cluster_name`: set to the Redis cluster name. ➎ `redis_cluster_ssl`: enable SSL/TLS communication with Redis cluster. Send a request with a GraphQL query of depth 2 to a gateway instance: ``` curl -i "http://127.0.0.1:9080/graphql" -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${GH_ACCESS_TOKEN}" \ -d '{"query": "query {viewer{login}}"}' ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body. Send the same request to a different gateway instance within the same 30-second time interval, you should receive an `HTTP/1.1 429 Too Many Requests` response, verifying routes configured in different gateway nodes share the same quota. ### Rate Limit by Query Complexity[​](#rate-limit-by-query-complexity "Direct link to Rate Limit by Query Complexity") The following example calculates the raw score from the nodes a query resolves instead of its depth, then rejects a query whose charged cost exceeds a fixed budget. This example applies to API7 Enterprise 3.10.6 and later. Create a route with the `graphql-limit-count` plugin that charges the `complexity` cost. It allows a quota of 100 within a 30-second window per remote address, and rejects any single query costing more than 20: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-cost-route", "uri": "/graphql", "plugins": { "graphql-limit-count": { "count": 100, "time_window": 30, "rejected_code": 429, "key_type": "var", "key": "remote_addr", "policy": "local", "show_limit_quota_header": true, "cost_strategy": "complexity", "max_cost": 20 } }, "upstream": { "type": "roundrobin", "pass_host": "node", "scheme": "https", "nodes": { "api.github.com:443": 1 } } }' ``` adc.yaml ``` services: - name: graphql-service routes: - uris: - /graphql name: graphql-cost-route plugins: graphql-limit-count: count: 100 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local show_limit_quota_header: true cost_strategy: complexity max_cost: 20 upstream: type: roundrobin scheme: https nodes: - host: api.github.com port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD graphql-cost-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: github-graphql-external-domain spec: type: ExternalName externalName: api.github.com --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: github-graphql-https spec: targetRefs: - name: github-graphql-external-domain kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: graphql-cost-plugin-config spec: plugins: - name: graphql-limit-count config: count: 100 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local show_limit_quota_header: true cost_strategy: complexity max_cost: 20 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: graphql-cost-route spec: parentRefs: - name: api7ee3-apisix-gateway rules: - matches: - path: type: Exact value: /graphql filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: graphql-cost-plugin-config backendRefs: - name: github-graphql-external-domain port: 443 ``` Apply the configuration to your cluster: ``` kubectl apply -f graphql-cost-ic.yaml ``` graphql-cost-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: github-graphql-external-domain spec: type: ExternalName externalName: api.github.com --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: github-graphql-external-domain spec: ingressClassName: apisix passHost: node scheme: https --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: graphql-cost-route spec: ingressClassName: apisix http: - name: graphql-cost-route match: paths: - /graphql backends: - serviceName: github-graphql-external-domain servicePort: 443 plugins: - name: graphql-limit-count enable: true config: count: 100 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local show_limit_quota_header: true cost_strategy: complexity max_cost: 20 ``` Apply the configuration to your cluster: ``` kubectl apply -f graphql-cost-ic.yaml ``` Send a small query to a gateway instance: ``` curl -i "http://127.0.0.1:9080/graphql" -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${GH_ACCESS_TOKEN}" \ -d '{"query": "query {viewer{login}}"}' ``` You should see an `HTTP/1.1 200 OK` response carrying the cost it was charged: ``` X-Graphql-Query-Cost: 4 ``` Send a query whose charged cost exceeds the budget: ``` curl -i "http://127.0.0.1:9080/graphql" -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${GH_ACCESS_TOKEN}" \ -d '{"query": "query {viewer{login name email location company bio websiteUrl twitterUsername createdAt updatedAt databaseId url avatarUrl isHireable isViewer isEmployee isSiteAdmin pronouns}}"}' ``` You should see an `HTTP/1.1 403 Forbidden` response, and the query never reaches the upstream: ``` {"message":"Invalid graphql request: query cost 21 exceeds max_cost 20"} ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * count integer | string vaild vaule: greater than 0 *** The maximum accumulated GraphQL query cost allowed within a given time interval. Query cost is depth under the default strategy. Required when `rules` is not configured. The value can use [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) when configured as a string. * time\_window integer | string vaild vaule: greater than 0 *** The time interval corresponding to the rate limiting `count` in seconds. Required when `rules` is not configured. The value can use [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) when configured as a string. * rules array\[object] *** An array of rate-limiting rules that are applied sequentially. Configure either `rules` or the top-level `count` and `time_window`, but not both. * count integer | string required vaild vaule: greater than 0 *** The maximum accumulated GraphQL query cost allowed within the rule's `time_window`. Query cost is depth under the default strategy. The value can use [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) when configured as a string. * time\_window integer | string required vaild vaule: greater than 0 *** The time interval corresponding to the rule's `count` in seconds. The value can use [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) when configured as a string. * key string required *** The key to count requests by. Supports combinations of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md), with each variable prefixed by a dollar sign (`$`). If the key cannot be resolved, the rule is not applied. * header\_prefix string *** Prefix inserted into the rate limiting response headers for this rule. For example, `foo` produces `X-foo-RateLimit-Limit`, `X-foo-RateLimit-Remaining`, and `X-foo-RateLimit-Reset`. * cost\_strategy string default: `depth` vaild vaule: `depth`, `complexity`, or `node_quantifier` *** How the raw cost of a GraphQL document is computed. `depth` counts the selection nesting depth, which is what this plugin has always done. `complexity` scores the nodes the query resolves. `node_quantifier` scores only nodes whose matching cost decoration resolves an argument listed in `mul_arguments`. If no node has both a matching decoration and a usable quantifier, the document's raw score is `0`. For `complexity` and `node_quantifier`, the plugin adds `0.01`, applies `score_factor`, and rounds up before charging the quota. The default factor produces a charged cost of `1` from a zero raw score; a factor greater than `100` increases it. Introduced in API7 Enterprise 3.10.6. * max\_cost number default: `0` vaild vaule: greater than or equal to 0 *** Reject a document whose charged cost exceeds this value with `403 Forbidden` before it reaches the upstream. The request consumes quota before this check. `0` disables the check and lets the quota alone decide. Introduced in API7 Enterprise 3.10.6. * score\_factor number default: `1` vaild vaule: greater than 0 *** Scaling applied before the cost is rounded up, charged against the quota, and compared with `max_cost`. Introduced in API7 Enterprise 3.10.6. * resolve\_variables boolean default: `true` *** If true, resolve supplied GraphQL variables, variable defaults declared by the operation, and argument defaults from the upstream schema when computing cost. The default ensures variable-based quantifiers contribute their resolved value. Introduced in API7 Enterprise 3.10.6. * introspection\_endpoint string vaild vaule: starts with `http://` or `https://` *** Endpoint used to introspect the upstream GraphQL schema, which the `complexity` and `node_quantifier` strategies need in order to match cost decorations. Derived from the upstream when unset. The result is cached per worker and Service. Introduced in API7 Enterprise 3.10.6. * introspection\_headers object *** Headers sent on the schema introspection request for an upstream whose introspection endpoint requires credentials. The headers come from the configuration rather than from the request because the introspected schema is cached per worker and Service. When Data Plane data encryption is enabled, this field is encrypted at rest. Introduced in API7 Enterprise 3.10.6. * key\_type string default: `var` vaild vaule: `var`, `var_combination`, or `constant` *** The type of key. If the `key_type` is `var`, the `key` is interpreted as a variable. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. If the `key_type` is `constant`, the `key` is interpreted as a constant. * key string default: `remote_addr` *** The key to count requests by. If the `key_type` is `var`, the `key` is interpreted as a variable. The variable does not need to be prefixed by a dollar sign (`$`). See [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) for available variables. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. All variables should be prefixed by dollar signs (`$`). For example, to configure the `key` to use a combination of two request headers `custom-a` and `custom-b`, the `key` should be configured as `$http_custom_a $http_custom_b`. If the `key_type` is `constant`, the `key` is interpreted as a constant value. * rejected\_code integer default: `503` vaild vaule: between 200 and 599 inclusive *** The HTTP status code returned when a request is rejected for exceeding the threshold. * rejected\_msg string vaild vaule: any non-empty string *** The response body returned when a request is rejected for exceeding the threshold. * policy string default: `local` vaild vaule: `local`, `redis`, or `redis-cluster` *** The policy for rate limiting counter. If it is `local`, the counter is stored in memory locally. If it is `redis`, the counter is stored on a Redis instance. If it is `redis-cluster`, the counter is stored in a Redis cluster. * allow\_degradation boolean default: `false` *** If true, allow the gateway to continue handling requests without the plugin when the plugin or its dependencies become unavailable. * show\_limit\_quota\_header boolean default: `true` *** If true, includes the rate limiting response headers. Specifically: * `X-RateLimit-Limit` shows the total quota. * `X-RateLimit-Remaining` shows the remaining quota. * `X-RateLimit-Reset` shows the number of seconds until the counter resets. * group string vaild vaule: non-empty *** The `group` ID for the plugin, such that routes of the same `group` can share the same rate limiting counter. * redis\_host string *** The address of the Redis node. Required when `policy` is `redis`. * redis\_port integer default: `6379` vaild vaule: greater than or equal to 1 *** The port of the Redis node when `policy` is `redis`. * redis\_username string *** The username for Redis if Redis ACL is used. If you use the legacy authentication method `requirepass`, configure only the `redis_password`. Used when `policy` is `redis`. * redis\_password string *** The password of the Redis node when `policy` is `redis` or `redis-cluster`. * redis\_database integer default: `0` vaild vaule: greater than or equal to 0 *** The database number in Redis when `policy` is `redis`. * redis\_ssl boolean default: `false` *** If true, use SSL to connect to Redis when `policy` is `redis`. * redis\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis`. * redis\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** The Redis timeout value in milliseconds when `policy` is `redis` or `redis-cluster`. * redis\_keepalive\_timeout integer default: `10000` vaild vaule: greater than or equal to 1000 *** Keepalive timeout in milliseconds for Redis when `policy` is `redis` or `redis-cluster`. This parameter is available in API7 Enterprise from version 3.9.16 on the 3.9 line and from version 3.10.3 on the 3.10 line, and in APISIX from version 3.17.0. * redis\_keepalive\_pool integer default: `100` vaild vaule: greater than or equal to 1 *** Keepalive pool size for Redis when `policy` is `redis` or `redis-cluster`. This parameter is available in API7 Enterprise from version 3.9.16 on the 3.9 line and from version 3.10.3 on the 3.10 line, and in APISIX from version 3.17.0. * redis\_cluster\_nodes array\[string] *** The list of Redis cluster nodes with at least two addresses. Required when `policy` is `redis-cluster`. * redis\_cluster\_name string *** The name of the Redis cluster. Required when `policy` is `redis-cluster`. * redis\_cluster\_ssl boolean default: `false` *** If true, use SSL to connect to Redis cluster when `policy` is `redis-cluster`. * redis\_cluster\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis-cluster`. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * limit\_header string default: `X-RateLimit-Limit` *** Default response header name for the total rate limit quota. Available in API7 Enterprise from version 3.10.6. * remaining\_header string default: `X-RateLimit-Remaining` *** Default response header name for the remaining rate limit quota. Available in API7 Enterprise from version 3.10.6. * reset\_header string default: `X-RateLimit-Reset` *** Default response header name for the number of seconds until the rate limit counter resets. Available in API7 Enterprise from version 3.10.6. --- # graphql-proxy-cache The `graphql-proxy-cache` plugin caches GraphQL query responses using disk-based or in-memory caching. It supports GraphQL [GET](https://graphql.org/learn/serving-over-http/#get-request) and [POST](https://graphql.org/learn/serving-over-http/#post-request) requests. The plugin generates an MD5 cache key from the plugin configuration version, request host, route ID, service ID, authenticated consumer identity, and complete GraphQL request body. Consumer identity is included by default when APISIX resolves the request to a consumer or remote user. If a request contains a [mutation](https://graphql.org/learn/queries#mutations) operation, the plugin will not cache the data. Instead, it adds an `Apisix-Cache-Status: BYPASS` header to the response to show that the request bypasses the caching mechanism. For GET requests, provide the GraphQL document in the `query` query parameter. POST requests can use a JSON body containing a `query` field or an `application/graphql` body. Unsupported methods return `405 Method Not Allowed`; unreadable, malformed, or invalid GraphQL requests return `400 Bad Request`. APISIX reads up to 1 MiB of GraphQL request data by default. To change this limit, configure `graphql.max_size` in `config.yaml` and reload APISIX: config.yaml ``` graphql: max_size: 1048576 ``` ## Examples[​](#examples "Direct link to Examples") The examples below use the public [Countries GraphQL API](https://countries.trevorblades.com/) as an upstream and demonstrate how you can configure `graphql-proxy-cache` for different scenarios. ### Cache Data on Disk[​](#cache-data-on-disk "Direct link to Cache Data on Disk") On-disk caching strategy offers the advantages of data persistency when system restarts and having larger storage capacity compared to in-memory cache. It is suitable for applications that prioritize durability and can tolerate slightly larger cache access latency. The following example demonstrates how you can use `graphql-proxy-cache` plugin on a route to cache data on disk. Create a route with the `graphql-proxy-cache` plugin with the default configuration to cache data on disk: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-proxy-cache-route", "uri": "/graphql", "plugins": { "graphql-proxy-cache": {} }, "upstream": { "type": "roundrobin", "pass_host": "node", "scheme": "https", "nodes": { "countries.trevorblades.com:443": 1 } } }' ``` adc.yaml ``` services: - name: graphql-service routes: - uris: - /graphql name: graphql-proxy-cache-route plugins: graphql-proxy-cache: {} upstream: type: roundrobin scheme: https nodes: - host: countries.trevorblades.com port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD graphql-proxy-cache-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: countries-graphql-external-domain spec: type: ExternalName externalName: countries.trevorblades.com --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: countries-graphql-https spec: targetRefs: - name: countries-graphql-external-domain kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: graphql-proxy-cache-plugin-config spec: plugins: - name: graphql-proxy-cache config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: graphql-proxy-cache-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /graphql filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: graphql-proxy-cache-plugin-config backendRefs: - name: countries-graphql-external-domain port: 443 ``` graphql-proxy-cache-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: countries-graphql-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: countries.trevorblades.com port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: graphql-proxy-cache-route spec: ingressClassName: apisix http: - name: graphql-proxy-cache-route match: paths: - /graphql upstreams: - name: countries-graphql-external-domain plugins: - name: graphql-proxy-cache enable: true ``` Apply the configuration to your cluster: ``` kubectl apply -f graphql-proxy-cache-ic.yaml ``` Send a request with a GraphQL query to verify: ``` curl -i "http://127.0.0.1:9080/graphql" -X POST \ -H "Content-Type: application/json" \ -d '{"query": "query { country(code: \"US\") { name capital } }"}' ``` You should see an `HTTP/1.1 200 OK` response with the following headers, showing the plugin is successfully enabled: ``` APISIX-Cache-Key: 5908e74856ea02835af198678b879a71 Apisix-Cache-Status: MISS ``` As there is no cache available before the first response, `Apisix-Cache-Status: MISS` is shown. The exact cache key depends on your configuration. Send the same request again within the cache TTL window. You should see an `HTTP/1.1 200 OK` response with the following headers, showing the cache is hit: ``` APISIX-Cache-Key: 5908e74856ea02835af198678b879a71 Apisix-Cache-Status: HIT ``` Wait for the cache to expire after the TTL and send the same request again. You should see an `HTTP/1.1 200 OK` response with the following headers, showing the cache has expired: ``` APISIX-Cache-Key: 5908e74856ea02835af198678b879a71 Apisix-Cache-Status: EXPIRED ``` ### Cache Data in Memory[​](#cache-data-in-memory "Direct link to Cache Data in Memory") In-memory caching strategy offers the advantage of low-latency access to the cached data, as retrieving data from RAM is faster than retrieving data from disk storage. It also works well for storing temporary data that does not need to be persisted long-term, allowing for efficient caching of frequently changing data. The following example demonstrates how you can use `graphql-proxy-cache` plugin on a route to cache data in memory. Create a route with `graphql-proxy-cache` enabled and configure it to use memory-based caching: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-proxy-cache-route", "uri": "/graphql", "plugins": { "graphql-proxy-cache": { "cache_strategy": "memory", "cache_zone": "memory_cache", "cache_ttl": 10 } }, "upstream": { "type": "roundrobin", "pass_host": "node", "scheme": "https", "nodes": { "countries.trevorblades.com:443": 1 } } }' ``` adc.yaml ``` services: - name: graphql-service routes: - uris: - /graphql name: graphql-proxy-cache-route plugins: graphql-proxy-cache: cache_strategy: memory cache_zone: memory_cache cache_ttl: 10 upstream: type: roundrobin scheme: https nodes: - host: countries.trevorblades.com port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD graphql-proxy-cache-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: countries-graphql-external-domain spec: type: ExternalName externalName: countries.trevorblades.com --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: countries-graphql-https spec: targetRefs: - name: countries-graphql-external-domain kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: graphql-proxy-cache-plugin-config spec: plugins: - name: graphql-proxy-cache config: cache_strategy: memory cache_zone: memory_cache cache_ttl: 10 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: graphql-proxy-cache-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /graphql filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: graphql-proxy-cache-plugin-config backendRefs: - name: countries-graphql-external-domain port: 443 ``` graphql-proxy-cache-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: countries-graphql-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: countries.trevorblades.com port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: graphql-proxy-cache-route spec: ingressClassName: apisix http: - name: graphql-proxy-cache-route match: paths: - /graphql upstreams: - name: countries-graphql-external-domain plugins: - name: graphql-proxy-cache enable: true config: cache_strategy: memory cache_zone: memory_cache cache_ttl: 10 ``` Apply the configuration to your cluster: ``` kubectl apply -f graphql-proxy-cache-ic.yaml ``` ❶ `cache_strategy`: set to `memory` for in-memory setting. ❷ `cache_zone`: set to the name of an in-memory cache zone. ❸ `cache_ttl`: set the time to live for the in-memory cache. Send a request with a GraphQL query to verify: ``` curl "http://127.0.0.1:9080/graphql" -i -X POST \ -H "Content-Type: application/json" \ -d '{"query": "query { country(code: \"US\") { name capital } }"}' ``` You should see an `HTTP/1.1 200 OK` response with the following headers, showing the plugin is successfully enabled: ``` APISIX-Cache-Key: a661316c4b1b70ae2db5347743dec6b6 Apisix-Cache-Status: MISS ``` As there is no cache available before the first response, `Apisix-Cache-Status: MISS` is shown. The exact cache key depends on your configuration. Send the same request again within the cache TTL window. You should see an `HTTP/1.1 200 OK` response with the following headers, showing the cache is hit: ``` APISIX-Cache-Key: a661316c4b1b70ae2db5347743dec6b6 Apisix-Cache-Status: HIT ``` ### Remove Cache Manually[​](#remove-cache-manually "Direct link to Remove Cache Manually") While most of the time it is not necessary, there may be situations where you would want to manually remove cached data. The following example demonstrates how you can use the `public-api` plugin to expose the `/apisix/plugin/graphql-proxy-cache/{cache_strategy}/{route_id}/{key}` endpoint created by the `graphql-proxy-cache` plugin. The example also enables `key-auth` so that only an authenticated operator can purge cached responses. Create a consumer with a `key-auth` credential and a route that matches the URI `/apisix/plugin/graphql-proxy-cache/*`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "cache-operator" }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/cache-operator/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cache-operator-key-auth", "plugins": { "key-auth": { "key": "purge-key" } } }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "graphql-cache-purge", "uri": "/apisix/plugin/graphql-proxy-cache/*", "plugins": { "key-auth": {}, "public-api": {} } }' ``` adc.yaml ``` consumers: - username: cache-operator credentials: - name: cache-operator-key-auth type: key-auth config: key: purge-key services: - name: graphql-cache-purge-service routes: - name: graphql-cache-purge-route uris: - /apisix/plugin/graphql-proxy-cache/* plugins: key-auth: {} public-api: {} ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD graphql-proxy-cache-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: cache-operator spec: gatewayRef: name: apisix credentials: - type: key-auth name: cache-operator-key-auth config: key: purge-key --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: graphql-cache-purge-plugin-config spec: plugins: - name: key-auth config: _meta: disable: false - name: public-api config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: graphql-cache-purge-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /apisix/plugin/graphql-proxy-cache/ filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: graphql-cache-purge-plugin-config ``` graphql-proxy-cache-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: cache-operator spec: ingressClassName: apisix authParameter: keyAuth: value: key: purge-key --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: graphql-cache-purge-route spec: ingressClassName: apisix http: - name: graphql-cache-purge-route match: paths: - /apisix/plugin/graphql-proxy-cache/* plugins: - name: key-auth enable: true - name: public-api enable: true ``` Apply the configuration to your cluster: ``` kubectl apply -f graphql-proxy-cache-ic.yaml ``` Send the disk-cache request and save the generated `APISIX-Cache-Key` response header: ``` CACHE_KEY=$(curl -sS -D - -o /dev/null "http://127.0.0.1:9080/graphql" -X POST \ -H "Content-Type: application/json" \ -d '{"query": "query { country(code: \"US\") { name capital } }"}' | \ awk 'tolower($1) == "apisix-cache-key:" {gsub("\\r", "", $2); print $2}') ``` Send a PURGE request using that value and the ID of the route containing `graphql-proxy-cache`: ``` curl -i "http://127.0.0.1:9080/apisix/plugin/graphql-proxy-cache/disk/graphql-proxy-cache-route/${CACHE_KEY}" -X PURGE \ -H "apikey: purge-key" ``` The Admin API and ADC examples use `graphql-proxy-cache-route` as the route ID. For an Ingress Controller deployment, replace it with the generated APISIX route ID. An `HTTP/1.1 200 OK` response verifies that the cache corresponding to the key is successfully removed. If you send the same PURGE request again, you should see an `HTTP/1.1 404 Not Found` response, showing there is no cache on disk with this cache key after the cache removal. Responses with a `Vary` header can produce multiple cache variants. A successful PURGE response confirms removal of the targeted cache entry, but does not guarantee that every variant was removed. --- ## Static Configurations[​](#static-configurations "Direct link to Static Configurations") The gateway default configuration includes proxy cache settings for disk caching and cache zones. The file to update depends on how the gateway is deployed: * Host or Docker * Kubernetes (Helm) For host or Docker deployments, configure the following settings: config.yaml ``` apisix: proxy_cache: cache_ttl: 10s # for caching on disk zones: - name: disk_cache_one memory_size: 50m disk_size: 1G disk_path: /tmp/disk_cache_one cache_levels: 1:2 # - name: disk_cache_two # memory_size: 50m # disk_size: 1G # disk_path: "/tmp/disk_cache_two" # cache_levels: "1:2" - name: memory_cache memory_size: 50m ``` Then reload the gateway for changes to take effect. In the APISIX Helm chart version `2.16.0` or later and the API7 Gateway Helm chart version `3.10.3` or later, set `apisix.proxyCache`. Each chart renders this value as `apisix.proxy_cache` in the gateway configuration: values.yaml ``` apisix: proxyCache: cacheTtl: 10s zones: - name: disk_cache_one memory_size: 50m disk_size: 1G disk_path: /tmp/disk_cache_one cache_levels: 1:2 - name: memory_cache memory_size: 50m ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * cache\_strategy string default: `disk` vaild vaule: `disk` or `memory` *** Caching strategy. Cache on disk or in memory. * cache\_zone string default: `disk_cache_one` *** Cache zone used with the caching strategy. The value should match one of the cache zones defined in the [configuration files](https://docs.api7.ai/hub/graphql-proxy-cache/configuration.md#static-configurations) and should correspond to the caching strategy. For example, when using the in-memory caching strategy, you should use an in-memory cache zone. * cache\_ttl integer default: `300` vaild vaule: greater than or equal to 1 *** Cache time to live (TTL) in seconds when caching in memory. To adjust the TTL when caching on disk, update `cache_ttl` in the [configuration files](https://docs.api7.ai/hub/proxy-cache/configuration.md#static-configurations). Note that the TTL value is only used when the response headers [`Cache-Control`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control) and [`Expires`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Expires) are both absent. * consumer\_isolation boolean default: `true` *** If true, prepend the authenticated consumer identity to the effective cache key when the request resolves to an APISIX consumer or remote user. Arbitrary upstream credentials, such as a forwarded bearer token, are not used as the identity. If a route forwards user-specific credentials without using an APISIX authentication plugin, different users could receive the same cached response. Enable an APISIX authentication plugin or disable caching for that route. Available in API7 Enterprise from version 3.9.13 and APISIX from version 3.17.0. * cache\_set\_cookie boolean default: `false` *** If true, allow the in-memory strategy to cache responses that include a `Set-Cookie` header. By default, such responses are not cached. Available in API7 Enterprise from version 3.9.13 and APISIX from version 3.17.0. --- # grpc-transcode The `grpc-transcode` plugin converts between HTTP requests and gRPC requests, as well as their corresponding responses. With this plugin enabled, APISIX accepts an HTTP request from the client, converts it, and forwards it to an upstream gRPC service. When APISIX receives the gRPC response, it transforms the response back to an HTTP response and sends it to the client. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure the `grpc-transcode` plugin for different scenarios. To follow along the examples, start an [example gRPC server](https://github.com/api7/grpc_server_example): ``` docker run -d \ --name grpc-example-server \ -p 50051:50051 \ api7/grpc-server-example:1.0.2 ``` ### Transform between HTTP and gRPC Requests[​](#transform-between-http-and-grpc-requests "Direct link to Transform between HTTP and gRPC Requests") The following example demonstrates how to configure protobuf in APISIX and transform between HTTP and gRPC Requests using the `grpc-transcode` plugin. * Admin API * ADC * Ingress Controller Create a proto resource to store the protobuf: ``` curl "http://127.0.0.1:9180/apisix/admin/protos" -X PUT -d ' { "id": "echo-proto", "content": "syntax = \"proto3\"; package echo; service EchoService { rpc Echo (EchoMsg) returns (EchoMsg); } message EchoMsg { string msg = 1; }" }' ``` Create a route with the `grpc-transcode` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT -d ' { "id": "grpc-transcode-route", "methods": ["GET"], "uri": "/echo", "plugins": { "grpc-transcode": { "proto_id": "echo-proto", "service": "echo.EchoService", "method": "Echo" } }, "upstream": { "scheme": "grpc", "type": "roundrobin", "nodes": { "grpc-example-server:50051": 1 } } }' ``` ❶ `proto_id`: ID of the proto object which defines gRPC services ❷ `service`: gRPC service to interact with ❸ `method`: gRPC method to use To verify, send an HTTP request to the route with parameters defined in `EchoMsg`: ``` curl "http://127.0.0.1:9080/echo?msg=Hello" ``` You should receive the following response: ``` {"msg":"Hello"} ``` caution ADC currently does not support the configuration of proto resources. This example cannot be completed using ADC alone. caution The Ingress Controller currently does not support the configuration of proto resources. This example cannot be completed using Ingress Controller resources. ### Configure Protobuf with `.pb` File[​](#configure-protobuf-with-pb-file "Direct link to configure-protobuf-with-pb-file") The following example demonstrates how to configure protobuf with `.pb` file in APISIX and transform between HTTP and gRPC Requests using the `grpc-transcode` plugin. If your proto file contains imports, or if you want to combine multiple proto files, you can generate a `.pb` file using the [protoc](https://google.github.io/proto-lens/installing-protoc.html) utility and use it in APISIX, following the below steps. * Admin API * ADC * Ingress Controller Save the protocol buffer definition to a file called `echo.proto`: echo.proto ``` syntax = "proto3"; package echo; service EchoService { rpc Echo (EchoMsg) returns (EchoMsg); } message EchoMsg { string msg = 1; } ``` Generate the `.pb` file with the [protoc](https://google.github.io/proto-lens/installing-protoc.html) utility and output it to a new file called `echo_proto.pb`: ``` protoc --include_imports --descriptor_set_out=echo_proto.pb echo.proto ``` Convert the `.pb` file from binary to base64 and configure it in APISIX: ``` curl "http://127.0.0.1:9180/apisix/admin/protos" -X PUT --data-binary @- <
## Request Handling[​](#request-handling "Direct link to Request Handling") The `grpc-web` plugin processes client requests with specific HTTP methods, content types, and CORS rules. ### Supported HTTP Methods[​](#supported-http-methods "Direct link to Supported HTTP Methods") The plugin supports: * `POST` for gRPC-Web requests * `OPTIONS` for CORS preflight checks See [CORS support](https://github.com/grpc/grpc-web/blob/master/doc/browser-features.md#cors-support) for details. ### Supported Content Types[​](#supported-content-types "Direct link to Supported Content Types") The plugin recognizes the following content types: * `application/grpc-web` * `application/grpc-web-text` * `application/grpc-web+proto` * `application/grpc-web-text+proto` It automatically decodes messages in binary or base64 text format and translates them into standard gRPC for the upstream server. See [Protocol differences vs gRPC over HTTP2](https://github.com/grpc/grpc/blob/master/doc/PROTOCOL-WEB.md#protocol-differences-vs-grpc-over-http2) for more details. ### CORS Handling[​](#cors-handling "Direct link to CORS Handling") The plugin automatically handles cross-origin requests. By default: * All origins (`*`) are allowed * `POST` requests are permitted * Accepted request headers: `content-type`, `x-grpc-web`, `x-user-agent` * Exposed response headers: `grpc-status`, `grpc-message` ## Examples[​](#examples "Direct link to Examples") The following examples demonstrate how to configure and use the `grpc-web` plugin with a gRPC-Web client. ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before proceeding with the examples, complete the following preliminary steps to set up an upstream server and gRPC-Web client. #### Start an Upstream Server[​](#start-an-upstream-server "Direct link to Start an Upstream Server") Start a [grpcbin server](https://github.com/moul/grpcbin) to serve as the example upstream: * Docker * Kubernetes ``` docker run -d \ --name grpcbin \ -p 9000:9000 \ moul/grpcbin ``` Create a Kubernetes manifest file for the deployment of the grpcbin server: grpcbin.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: grpcbin spec: replicas: 1 selector: matchLabels: app: grpcbin template: metadata: labels: app: grpcbin spec: containers: - name: grpcbin image: moul/grpcbin ports: - containerPort: 9000 --- apiVersion: v1 kind: Service metadata: namespace: aic name: grpcbin spec: selector: app: grpcbin ports: - protocol: TCP port: 9000 targetPort: 9000 type: ClusterIP ``` Apply the configuration to your cluster: ``` kubectl apply -f grpcbin.yaml ``` #### Generate gRPC-Web client code[​](#generate-grpc-web-client-code "Direct link to Generate gRPC-Web client code") Download the protocol buffer definition `hello.proto`: ``` curl -O https://raw.githubusercontent.com/moul/pb/refs/heads/master/hello/hello.proto ``` Install [`protobuf`](https://github.com/protocolbuffers/protobuf/releases) and [`protoc-gen-grpc-web`](https://github.com/grpc/grpc-web/releases). Generate the gRPC-Web client code from `hello.proto`: ``` protoc \ --js_out=import_style=commonjs:. \ --grpc-web_out=import_style=commonjs,mode=grpcwebtext:. \ hello.proto ``` You should see two files generated in the current directory: `hello_pb.js` for protocol buffers message classes and `hello_grpc_web_pb.js` for gRPC-Web client stubs. #### Create a Client[​](#create-a-client "Direct link to Create a Client") Create a Node.js project and install the required dependencies: ``` npm init -y npm install xhr2 grpc-web google-protobuf ``` Create a client file: client.js ``` const XMLHttpRequest = require('xhr2'); const { HelloServiceClient } = require('./hello_grpc_web_pb'); const { HelloRequest } = require('./hello_pb'); global.XMLHttpRequest = XMLHttpRequest; function sayHello(){ const client = new HelloServiceClient('http://127.0.0.1:9080/grpc/web', null, { format: 'text', }); const req = new HelloRequest(); req.setGreeting('jack'); const call = client.sayHello(req, {}, (err, resp) => { if (err) { console.error('grpc error:', err.code, err.message); } else { console.log('reply:', resp.getReply()); } }); call.on('metadata', (metadata) => { console.log('Response headers:', metadata); }); } function lotsOfReplies() { const client = new HelloServiceClient('http://127.0.0.1:9080/grpc/web', null, { format: 'text', }); const req = new HelloRequest(); req.setGreeting('rep'); const stream = client.lotsOfReplies(req, {}); stream.on('metadata', (metadata) => { console.log('Response headers:', metadata); }); } lotsOfReplies() sayHello() ``` You can later run the client with `node client.js` to send both unary and server-streaming requests to your gRPC server via the gateway. ### Proxy gRPC-Web (Prefix Match Route)[​](#proxy-grpc-web-prefix-match-route "Direct link to Proxy gRPC-Web (Prefix Match Route)") The following examples demonstrate how to configure and use the `grpc-web` plugin with the gRPC-Web client set up previously. Create a route with the `grpc-web` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT -d ' { "id": "grpc-web-route", "uri": "/grpc/web/*", "plugins": { "grpc-web": {} }, "upstream": { "scheme": "grpc", "type": "roundrobin", "nodes": { "192.168.10.103:9000": 1 } } }' ``` ❶ Configure the `uri` to prefix-match the route requested in `client.js`. ❷ Enable the `grpc-web` plugin. ❸ Set the upstream scheme to `grpc`. ❹ Replace with your upstream server address. adc.yaml ``` services: - name: grpcbin routes: - name: grpc-web-route uris: - /grpc/web/* plugins: grpc-web: {} upstream: scheme: grpc type: roundrobin nodes: - host: grpcbin port: 9000 weight: 1 ``` ❶ Configure the `uri` to prefix-match the route requested in `client.js`. ❷ Enable the `grpc-web` plugin. ❸ Set the upstream scheme to `grpc`. ❹ Replace with your upstream server address. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Gateway API and gRPC-Web The Gateway API `GRPCRoute` matches requests by gRPC service and method names, not by HTTP path prefix. To proxy gRPC-Web traffic with the `grpc-web` plugin, configure a `GRPCRoute` with no method matches so it accepts all incoming gRPC requests. When using this configuration, update the `HelloServiceClient` base URL in `client.js` to `http://127.0.0.1:9080` (without the `/grpc/web` prefix). grpc-web-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: grpc-web-plugin-config spec: plugins: - name: grpc-web config: {} --- apiVersion: gateway.networking.k8s.io/v1 kind: GRPCRoute metadata: namespace: aic name: grpc-web-route spec: parentRefs: - name: apisix rules: - filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: grpc-web-plugin-config backendRefs: - name: grpcbin port: 9000 ``` grpc-web-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: grpcbin spec: ingressClassName: apisix scheme: grpc --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: grpc-web-route spec: ingressClassName: apisix http: - name: grpc-web-route match: paths: - /grpc/web/* backends: - serviceName: grpcbin servicePort: 9000 plugins: - name: grpc-web enable: true config: {} ``` ❶ Set the upstream scheme to `grpc`. ❷ Configure the route path to prefix-match the route requested in `client.js`. ❸ Replace with your upstream service name and port. ❹ Enable the `grpc-web` plugin. Apply the configuration to your cluster: ``` kubectl apply -f grpc-web-ic.yaml ``` understand the uri In APISIX versions prior to 3.15.0 and API7 Enterprise versions prior to 3.8.21, the route URI must use a prefix match because gRPC-Web clients include the package name, service name, and method name in the request URI. Using an absolute URI match in these versions will prevent the request from matching the route. [Absolute URI routes](#proxy-grpc-web-absolute-uri) are supported in later versions. In this example, the route URI must be configured as `/grpc/web/*` to correctly match client requests such as `/grpc/web/hello.HelloService/SayHello`. Using a broader prefix like `/grpc/*` would prevent the gateway from correctly extracting the full service path, resulting in errors such as `unknown service web/hello.HelloService`. Run the client to send requests to the gateway route: ``` node client.js ``` You should see a reply from the upstream gRPC server: ``` Response headers: { ... 'access-control-allow-origin': '*', 'access-control-expose-headers': 'grpc-message,grpc-status' } Response headers: { ... 'access-control-allow-origin': '*', 'access-control-expose-headers': 'grpc-message,grpc-status' } reply: hello jack ``` ### Proxy gRPC-Web (Absolute URI)[​](#proxy-grpc-web-absolute-uri "Direct link to Proxy gRPC-Web (Absolute URI)") This example applies to APISIX 3.15.0 and later, and API7 Enterprise 3.8.21 and later. When an absolute URI is used, the gateway does not automatically strip the URI path prefix. To forward requests correctly to the upstream gRPC server, use the `proxy-rewrite` plugin to adjust the request path. Create a route with the `grpc-web` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT -d ' { "id": "grpc-web-route", "uri": "/grpc/web/hello.HelloService/SayHello", "plugins": { "grpc-web": {}, "proxy-rewrite": { "uri": "/hello.HelloService/SayHello", "set_ngx_uri": "true" } }, "upstream": { "scheme": "grpc", "type": "roundrobin", "nodes": { "192.168.10.103:9000": 1 } } }' ``` ❶ Configure the `uri` to use the absolute path, including the base prefix and the gRPC full service path. ❷ Configure `proxy-rewrite` to rewrite the path and strip the route prefix. ❸ Set `set_ngx_uri` to `true` to update the requested path to the URI defined in the `proxy-rewrite` plugin. Without this setting, the gateway will not correctly forward the request to the upstream, resulting in errors such as `unknown service grpc/web/hello.HelloService`. ❹ Set the upstream scheme to `grpc`. ❺ Replace with your upstream server address. adc.yaml ``` services: - name: grpcbin routes: - name: grpc-web-route uris: - /grpc/web/hello.HelloService/SayHello plugins: grpc-web: {} proxy-rewrite: uri: /hello.HelloService/SayHello set_ngx_uri: "true" upstream: scheme: grpc type: roundrobin nodes: - host: grpcbin port: 9000 weight: 1 ``` ❶ Configure the `uri` to use the absolute path, including the base prefix and the gRPC full service path. ❷ Configure `proxy-rewrite` to rewrite the path and strip the route prefix. ❸ Set `set_ngx_uri` to `true` to update the requested path to the URI defined in the `proxy-rewrite` plugin. Without this setting, the gateway will not correctly forward the request to the upstream, resulting in errors such as `unknown service grpc/web/hello.HelloService`. ❹ Set the upstream scheme to `grpc`. ❺ Replace with your upstream server address. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Gateway API and gRPC-Web The Gateway API `GRPCRoute` matches requests by gRPC service and method names, not by HTTP path prefix. A `GRPCRoute` with `service: hello.HelloService` and `method: SayHello` creates an exact route at `/hello.HelloService/SayHello`. No `proxy-rewrite` is needed because the route path already matches the gRPC method path. When using this configuration, update the `HelloServiceClient` base URL in `client.js` to `http://127.0.0.1:9080` (without the `/grpc/web` prefix). grpc-web-absolute-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: grpc-web-absolute-plugin-config spec: plugins: - name: grpc-web config: {} --- apiVersion: gateway.networking.k8s.io/v1 kind: GRPCRoute metadata: namespace: aic name: grpc-web-absolute-route spec: parentRefs: - name: apisix rules: - matches: - method: service: hello.HelloService method: SayHello filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: grpc-web-absolute-plugin-config backendRefs: - name: grpcbin port: 9000 ``` grpc-web-absolute-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: grpcbin spec: ingressClassName: apisix scheme: grpc --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: grpc-web-absolute-route spec: ingressClassName: apisix http: - name: grpc-web-absolute-route match: paths: - /grpc/web/hello.HelloService/SayHello backends: - serviceName: grpcbin servicePort: 9000 plugins: - name: grpc-web enable: true config: {} - name: proxy-rewrite enable: true config: uri: /hello.HelloService/SayHello set_ngx_uri: "true" ``` ❶ Set the upstream scheme to `grpc`. ❷ Configure the route path to use the absolute URI, including the base prefix and the gRPC full service path. ❸ Replace with your upstream service name and port. ❹ Configure `proxy-rewrite` to rewrite the path and strip the route prefix. ❺ Set `set_ngx_uri` to `true` to update the requested path to the URI defined in the `proxy-rewrite` plugin. Without this setting, the gateway will not correctly forward the request to the upstream, resulting in errors such as `unknown service grpc/web/hello.HelloService`. Apply the configuration to your cluster: ``` kubectl apply -f grpc-web-absolute-ic.yaml ``` Run the client to send requests to the gateway route: ``` node client.js ``` You should see a reply from the upstream gRPC server: ``` Response headers: { ... 'access-control-allow-origin': '*', 'access-control-expose-headers': 'grpc-message,grpc-status' } Response headers: { ... 'access-control-allow-origin': '*', 'access-control-expose-headers': 'grpc-message,grpc-status' } reply: hello jack ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * cors\_allow\_headers string default: `content-type,x-grpc-web,x-user-agent` *** Comma-separated list of request headers allowed for cross-origin requests. * max\_req\_body\_size integer default: `67108864` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes read for gRPC-Web processing. A larger body is rejected with `400 Bad Request`. Introduced in API7 Enterprise 3.9.17 and 3.10.4, and APISIX 3.18.0. --- # hmac-auth The `hmac-auth` plugin supports HMAC (Hash-based Message Authentication Code) authentication as a mechanism to ensure the integrity of requests, preventing them from being modified during transmissions. To use the plugin, you would configure HMAC secret keys on [consumers](https://docs.api7.ai/apisix/key-concepts/consumers.md) and enable the plugin on routes or services. When a consumer is successfully authenticated, APISIX adds additional headers, such as `X-Consumer-Username`, `X-Credential-Identifier`, and other consumer custom headers if configured, to the request, before proxying it to the upstream service. The upstream service will be able to differentiate between consumers and implement additional logic as needed. If any of these values is not available, the corresponding header will not be added. About X-Consumer-Username When consumers are configured using the Ingress Controller, the consumer name is generated in the format `namespace_consumername`. As a result, the `X-Consumer-Username` header will also follow this format instead of just `consumername`. ## Implementation[​](#implementation "Direct link to Implementation") Once enabled, the plugin verifies the HMAC signature in the request's `Authorization` header and checks that incoming requests are from trusted sources. Specifically, when APISIX receives an HMAC-signed request, the key ID is extracted from the `Authorization` header. APISIX then retrieves the corresponding consumer configuration, including the secret key. If the key ID is valid and exists, APISIX generates an HMAC signature using the request's `Date` header and the secret key. If this generated signature matches the signature provided in the `Authorization` header, the request is authenticated and forwarded to upstream services. The plugin implementation is based on [draft-cavage-http-signatures](https://www.ietf.org/archive/id/draft-cavage-http-signatures-12.txt). ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can work with the `hmac-auth` plugin for different scenarios. ### Implement HMAC Authentication on a Route[​](#implement-hmac-authentication-on-a-route "Direct link to Implement HMAC Authentication on a Route") The following example demonstrates how to implement HMAC authentication on a route. You will also attach a consumer custom ID to authenticated requests in the `X-Consumer-Custom-Id` header, which can be used to implement additional logics as needed. * Admin API * ADC * Ingress Controller Create a consumer `john` with a custom ID label: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john", "labels": { "custom_id": "495aec6a" } }' ``` Create `hmac-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-hmac-auth", "plugins": { "hmac-auth": { "key_id": "john-key", "secret_key": "john-secret-key" } } }' ``` Create a route with the `hmac-auth` plugin using its default configurations: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "hmac-auth-route", "uri": "/get", "methods": ["GET"], "plugins": { "hmac-auth": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `hmac-auth` credential and a route with `hmac-auth` plugin configured as such: adc.yaml ``` consumers: - username: john labels: custom_id: "495aec6a" credentials: - name: hmac-auth type: hmac-auth config: key_id: john-key secret_key: john-secret-key services: - name: hmac-auth-service routes: - name: hmac-auth-route uris: - /get methods: - GET plugins: hmac-auth: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create a consumer with `hmac-auth` credential and a route with `hmac-auth` plugin configured as such: * Gateway API * APISIX CRD hmac-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john labels: custom_id: "495aec6a" spec: gatewayRef: name: apisix credentials: - type: hmac-auth name: primary-cred config: key_id: john-key secret_key: john-secret-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: hmac-auth-plugin-config spec: plugins: - name: hmac-auth config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: hmac-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get method: GET filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: hmac-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f hmac-auth-ic.yaml ``` hmac-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john labels: custom_id: "495aec6a" spec: ingressClassName: apisix authParameter: hmacAuth: value: key_id: john-key secret_key: john-secret-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: hmac-auth-route spec: ingressClassName: apisix http: - name: hmac-auth-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: hmac-auth enable: true ``` Apply the configuration to your cluster: ``` kubectl apply -f hmac-auth-ic.yaml ``` Generate a signature. You can use the below Python snippet or other stack of your choice: hmac-sig-header-gen.py ``` import hmac import hashlib import base64 from datetime import datetime, timezone key_id = "john-key" # key id secret_key = b"john-secret-key" # secret key request_method = "GET" # HTTP method request_path = "/get" # route URI algorithm= "hmac-sha256" # can use other algorithms in allowed_algorithms # get current datetime in GMT # note: the signature will become invalid after the clock skew (default 300s) # you can regenerate the signature after it becomes invalid, or increase the clock # skew to prolong the validity within the advised security boundary gmt_time = datetime.now(timezone.utc).strftime('%a, %d %b %Y %H:%M:%S GMT') # construct the signing string (ordered) # the date and any subsequent custom headers should be lowercased and separated by a # single space character, i.e. `:` # https://datatracker.ietf.org/doc/html/draft-cavage-http-signatures-12#section-2.1.6 signing_string = ( f"{key_id}\n" f"{request_method} {request_path}\n" f"date: {gmt_time}\n" ) # create signature signature = hmac.new(secret_key, signing_string.encode('utf-8'), hashlib.sha256).digest() signature_base64 = base64.b64encode(signature).decode('utf-8') # construct the request headers headers = { "Date": gmt_time, "Authorization": ( f'Signature keyId="{key_id}",algorithm="{algorithm}",' f'headers="@request-target date",' f'signature="{signature_base64}"' ) } # print headers print(headers) ``` Run the script: ``` python3 hmac-sig-header-gen.py ``` You should see the request headers printed: ``` {'Date': 'Fri, 06 Sep 2024 06:41:29 GMT', 'Authorization': 'Signature keyId="john-key",algorithm="hmac-sha256",headers="@request-target date",signature="wWfKQvPDr0wHQ4IHdluB4IzeNZcj0bGJs2wvoCOT5rM="'} ``` Using the headers generated, send a request to the route: ``` curl -X GET "http://127.0.0.1:9080/get" \ -H "Date: Fri, 06 Sep 2024 06:41:29 GMT" \ -H 'Authorization: Signature keyId="john-key",algorithm="hmac-sha256",headers="@request-target date",signature="wWfKQvPDr0wHQ4IHdluB4IzeNZcj0bGJs2wvoCOT5rM="' ``` You should see an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "headers": { "Accept": "*/*", "Authorization": "Signature keyId=\"john-key\",algorithm=\"hmac-sha256\",headers=\"@request-target date\",signature=\"wWfKQvPDr0wHQ4IHdluB4IzeNZcj0bGJs2wvoCOT5rM=\"", "Date": "Fri, 06 Sep 2024 06:41:29 GMT", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-66d96513-2e52d4f35c9b6a2772d667ea", "X-Consumer-Username": "john", "X-Credential-Identifier": "cred-john-hmac-auth", "X-Consumer-Custom-Id": "495aec6a", "X-Forwarded-Host": "127.0.0.1" }, "origin": "192.168.65.1, 34.0.34.160", "url": "http://127.0.0.1/get" } ``` If you would like to attach more consumer custom headers to authenticated requests, see the [`attach-consumer-label`](https://docs.api7.ai/hub/attach-consumer-label.md) plugin. ### Hide Authorization Information From Upstream[​](#hide-authorization-information-from-upstream "Direct link to Hide Authorization Information From Upstream") As seen in the previous example, the `Authorization` header passed to the upstream includes the signature and all other details. This could potentially introduce security risks. This example continues from the [previous example](#implement-hmac-authentication-on-a-route) to demonstrate how to prevent this information from being sent to the upstream service. * Admin API * ADC * Ingress Controller Update the plugin configuration to set `hide_credentials` to `true`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes/hmac-auth-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "hmac-auth": { "hide_credentials": true } } }' ``` Update the plugin configuration as such: adc.yaml ``` consumers: - username: john labels: custom_id: "495aec6a" credentials: - name: hmac-auth type: hmac-auth config: key_id: john-key secret_key: john-secret-key services: - name: hmac-auth-service routes: - name: hmac-auth-route uris: - /get methods: - GET plugins: hmac-auth: hide_credentials: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update the PluginConfig to set `hide_credentials` to `true`: hmac-auth-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: hmac-auth-plugin-config spec: plugins: - name: hmac-auth config: _meta: disable: false hide_credentials: true ``` Apply the configuration to your cluster: ``` kubectl apply -f hmac-auth-ic.yaml ``` Update the ApisixRoute to set `hide_credentials` to `true`: hmac-auth-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: hmac-auth-route spec: ingressClassName: apisix http: - name: hmac-auth-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: hmac-auth enable: true config: hide_credentials: true ``` Apply the configuration to your cluster: ``` kubectl apply -f hmac-auth-ic.yaml ``` Send a request to the route: ``` curl -X GET "http://127.0.0.1:9080/get" \ -H "Date: Fri, 06 Sep 2024 06:41:29 GMT" \ -H 'Authorization: Signature keyId="john-key",algorithm="hmac-sha256",headers="@request-target date",signature="wWfKQvPDr0wHQ4IHdluB4IzeNZcj0bGJs2wvoCOT5rM="' ``` You should see an `HTTP/1.1 200 OK` response and notice the `Authorization` header is entirely removed: ``` { "args": {}, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-66d96513-2e52d4f35c9b6a2772d667ea", "X-Consumer-Username": "john", "X-Credential-Identifier": "cred-john-hmac-auth", "X-Forwarded-Host": "127.0.0.1" }, "origin": "192.168.65.1, 34.0.34.160", "url": "http://127.0.0.1/get" } ``` ### Enable Body Validation[​](#enable-body-validation "Direct link to Enable Body Validation") The following example demonstrates how to validate the request body and bind that digest to the HMAC signature. When `validate_request_body` is true, APISIX compares the `Digest` header with a SHA-256 digest of the request body. A missing or mismatched `Digest` header is rejected. That comparison does not bind the digest to the HMAC signature, so a client can still replace both the body and the `Digest` header together. Sign `digest` as well, and list it in `signed_headers`, so the digest cannot change without invalidating the signature. * Admin API * ADC * Ingress Controller Create a consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john" }' ``` Create `hmac-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-hmac-auth", "plugins": { "hmac-auth": { "key_id": "john-key", "secret_key": "john-secret-key" } } }' ``` Create a route with the `hmac-auth` plugin as such: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "hmac-auth-route", "uri": "/post", "methods": ["POST"], "plugins": { "hmac-auth": { "signed_headers": ["date", "digest"], "validate_request_body": true } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `hmac-auth` credential and a route with `hmac-auth` plugin configured as such: adc.yaml ``` consumers: - username: john credentials: - name: hmac-auth type: hmac-auth config: key_id: john-key secret_key: john-secret-key services: - name: hmac-auth-service routes: - name: hmac-auth-route uris: - /post methods: - POST plugins: hmac-auth: signed_headers: - date - digest validate_request_body: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create a consumer with `hmac-auth` credential and a route with `hmac-auth` plugin configured as such: * Gateway API * APISIX CRD hmac-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: hmac-auth name: primary-cred config: key_id: john-key secret_key: john-secret-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: hmac-auth-plugin-config spec: plugins: - name: hmac-auth config: _meta: disable: false signed_headers: - date - digest validate_request_body: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: hmac-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /post method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: hmac-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f hmac-auth-ic.yaml ``` hmac-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: hmacAuth: value: key_id: john-key secret_key: john-secret-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: hmac-auth-route spec: ingressClassName: apisix http: - name: hmac-auth-route match: paths: - /post methods: - POST upstreams: - name: httpbin-external-domain plugins: - name: hmac-auth enable: true config: signed_headers: - date - digest validate_request_body: true ``` Apply the configuration to your cluster: ``` kubectl apply -f hmac-auth-ic.yaml ``` Generate a signature. You can use the below Python snippet or other stack of your choice: hmac-sig-digest-header-gen.py ``` import hmac import hashlib import base64 from datetime import datetime, timezone key_id = "john-key" # key id secret_key = b"john-secret-key" # secret key request_method = "POST" # HTTP method request_path = "/post" # route URI algorithm= "hmac-sha256" # can use other algorithms in allowed_algorithms body = '{"name": "world"}' # example request body # get current datetime in GMT # note: the signature will become invalid after the clock skew (default 300s). # you can regenerate the signature after it becomes invalid, or increase the clock # skew to prolong the validity within the advised security boundary gmt_time = datetime.now(timezone.utc).strftime('%a, %d %b %Y %H:%M:%S GMT') # create the SHA-256 digest of the request body and include it in the signing string body_digest = hashlib.sha256(body.encode('utf-8')).digest() digest_header = "SHA-256=" + base64.b64encode(body_digest).decode('utf-8') # construct the signing string (ordered) # the date and any subsequent custom headers should be lowercased and separated by a # single space character, i.e. `:` # https://datatracker.ietf.org/doc/html/draft-cavage-http-signatures-12#section-2.1.6 signing_string = ( f"{key_id}\n" f"{request_method} {request_path}\n" f"date: {gmt_time}\n" f"digest: {digest_header}\n" ) # create signature signature = hmac.new(secret_key, signing_string.encode('utf-8'), hashlib.sha256).digest() signature_base64 = base64.b64encode(signature).decode('utf-8') # construct the request headers headers = { "Date": gmt_time, "Digest": digest_header, "Authorization": ( f'Signature keyId="{key_id}",algorithm="hmac-sha256",' f'headers="@request-target date digest",' f'signature="{signature_base64}"' ) } # print headers print(headers) ``` Run the script: ``` python3 hmac-sig-digest-header-gen.py ``` You should see the request headers printed: ``` {'Date': 'Thu, 20 Aug 2026 09:40:53 GMT', 'Digest': 'SHA-256=78qzJuLwSpZ8HacsTdFCQJWxzPMOf8bYctRk2ySLpS8=', 'Authorization': 'Signature keyId="john-key",algorithm="hmac-sha256",headers="@request-target date digest",signature="GD+WVdC2hIzLCgMfLlEx5gYqGC4wUQl59kl2XMqKbKM="'} ``` Using the headers generated, send a request to the route: ``` curl "http://127.0.0.1:9080/post" -X POST \ -H "Date: Thu, 20 Aug 2026 09:40:53 GMT" \ -H "Digest: SHA-256=78qzJuLwSpZ8HacsTdFCQJWxzPMOf8bYctRk2ySLpS8=" \ -H 'Authorization: Signature keyId="john-key",algorithm="hmac-sha256",headers="@request-target date digest",signature="GD+WVdC2hIzLCgMfLlEx5gYqGC4wUQl59kl2XMqKbKM="' \ -d '{"name": "world"}' ``` You should see an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "data": "", "files": {}, "form": { "{\"name\": \"world\"}": "" }, "headers": { "Accept": "*/*", "Authorization": "Signature keyId=\"john-key\",algorithm=\"hmac-sha256\",headers=\"@request-target date digest\",signature=\"GD+WVdC2hIzLCgMfLlEx5gYqGC4wUQl59kl2XMqKbKM=\"", "Content-Length": "17", "Content-Type": "application/x-www-form-urlencoded", "Date": "Thu, 20 Aug 2026 09:40:53 GMT", "Digest": "SHA-256=78qzJuLwSpZ8HacsTdFCQJWxzPMOf8bYctRk2ySLpS8=", "Host": "127.0.0.1:9080", "User-Agent": "curl/8.7.1", "X-Consumer-Username": "john", "X-Credential-Identifier": "cred-john-hmac-auth", "X-Forwarded-For": "192.168.117.1", "X-Forwarded-Host": "127.0.0.1:9080", "X-Forwarded-Port": "9080", "X-Forwarded-Proto": "http", "X-Real-IP": "192.168.117.1" }, "json": null, "origin": "192.168.117.3", "url": "http://127.0.0.1/post" } ``` If you send the same signed headers with a different body, the `Digest` header no longer matches: ``` curl "http://127.0.0.1:9080/post" -X POST \ -H "Date: Thu, 20 Aug 2026 09:40:53 GMT" \ -H "Digest: SHA-256=78qzJuLwSpZ8HacsTdFCQJWxzPMOf8bYctRk2ySLpS8=" \ -H 'Authorization: Signature keyId="john-key",algorithm="hmac-sha256",headers="@request-target date digest",signature="GD+WVdC2hIzLCgMfLlEx5gYqGC4wUQl59kl2XMqKbKM="' \ -d '{"name": "tampered"}' ``` You should see an `HTTP/1.1 401 Unauthorized` response with the following message: ``` {"message":"client request can't be validated"} ``` ### Mandate Signed Headers[​](#mandate-signed-headers "Direct link to Mandate Signed Headers") The following example demonstrates how you can mandate certain headers to be signed in the request's HMAC signature. * Admin API * ADC * Ingress Controller Create a consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john" }' ``` Create `hmac-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-hmac-auth", "plugins": { "hmac-auth": { "key_id": "john-key", "secret_key": "john-secret-key" } } }' ``` Create a route with the `hmac-auth` plugin which requires three headers to be present in the HMAC signature: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "hmac-auth-route", "uri": "/get", "methods": ["GET"], "plugins": { "hmac-auth": { "signed_headers": ["date","x-custom-header-a", "x-custom-header-b"] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `hmac-auth` credential and a route with `hmac-auth` plugin configured as such: adc.yaml ``` consumers: - username: john credentials: - name: hmac-auth type: hmac-auth config: key_id: john-key secret_key: john-secret-key services: - name: hmac-auth-service routes: - name: hmac-auth-route uris: - /get methods: - GET plugins: hmac-auth: signed_headers: - date - x-custom-header-a - x-custom-header-b upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create a consumer with `hmac-auth` credential and a route with `hmac-auth` plugin configured as such: * Gateway API * APISIX CRD hmac-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: hmac-auth name: primary-cred config: key_id: john-key secret_key: john-secret-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: hmac-auth-plugin-config spec: plugins: - name: hmac-auth config: _meta: disable: false signed_headers: - date - x-custom-header-a - x-custom-header-b --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: hmac-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get method: GET filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: hmac-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f hmac-auth-ic.yaml ``` hmac-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: hmacAuth: value: key_id: john-key secret_key: john-secret-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: hmac-auth-route spec: ingressClassName: apisix http: - name: hmac-auth-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: hmac-auth enable: true config: signed_headers: - date - x-custom-header-a - x-custom-header-b ``` Apply the configuration to your cluster: ``` kubectl apply -f hmac-auth-ic.yaml ``` Generate a signature. You can use the below Python snippet or other stack of your choice: hmac-sig-req-header-gen.py ``` import hmac import hashlib import base64 from datetime import datetime, timezone key_id = "john-key" # key id secret_key = b"john-secret-key" # secret key request_method = "GET" # HTTP method request_path = "/get" # route URI algorithm= "hmac-sha256" # can use other algorithms in allowed_algorithms custom_header_a = "hello123" # required custom header custom_header_b = "world456" # required custom header # get current datetime in GMT # note: the signature will become invalid after the clock skew (default 300s) # you can regenerate the signature after it becomes invalid, or increase the clock # skew to prolong the validity within the advised security boundary gmt_time = datetime.now(timezone.utc).strftime('%a, %d %b %Y %H:%M:%S GMT') # construct the signing string (ordered) # the date and any subsequent custom headers should be lowercased and separated by a # single space character, i.e. `:` # https://datatracker.ietf.org/doc/html/draft-cavage-http-signatures-12#section-2.1.6 signing_string = ( f"{key_id}\n" f"{request_method} {request_path}\n" f"date: {gmt_time}\n" f"x-custom-header-a: {custom_header_a}\n" f"x-custom-header-b: {custom_header_b}\n" ) # create signature signature = hmac.new(secret_key, signing_string.encode('utf-8'), hashlib.sha256).digest() signature_base64 = base64.b64encode(signature).decode('utf-8') # construct the request headers headers = { "Date": gmt_time, "Authorization": ( f'Signature keyId="{key_id}",algorithm="hmac-sha256",' f'headers="@request-target date x-custom-header-a x-custom-header-b",' f'signature="{signature_base64}"' ), "x-custom-header-a": custom_header_a, "x-custom-header-b": custom_header_b } # print headers print(headers) ``` Run the script: ``` python3 hmac-sig-req-header-gen.py ``` You should see the request headers printed: ``` {'Date': 'Fri, 06 Sep 2024 09:58:49 GMT', 'Authorization': 'Signature keyId="john-key",algorithm="hmac-sha256",headers="@request-target date x-custom-header-a x-custom-header-b",signature="MwJR8JOhhRLIyaHlJ3Snbrf5hv0XwdeeRiijvX3A3yE="', 'x-custom-header-a': 'hello123', 'x-custom-header-b': 'world456'} ``` Using the headers generated, send a request to the route: ``` curl -X GET "http://127.0.0.1:9080/get" \ -H "Date: Fri, 06 Sep 2024 09:58:49 GMT" \ -H 'Authorization: Signature keyId="john-key",algorithm="hmac-sha256",headers="@request-target date x-custom-header-a x-custom-header-b",signature="MwJR8JOhhRLIyaHlJ3Snbrf5hv0XwdeeRiijvX3A3yE="' \ -H "x-custom-header-a: hello123" \ -H "x-custom-header-b: world456" ``` You should see an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "headers": { "Accept": "*/*", "Authorization": "Signature keyId=\"john-key\",algorithm=\"hmac-sha256\",headers=\"@request-target date x-custom-header-a x-custom-header-b\",signature=\"MwJR8JOhhRLIyaHlJ3Snbrf5hv0XwdeeRiijvX3A3yE=\"", "Date": "Fri, 06 Sep 2024 09:58:49 GMT", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-66d98196-64a58db25ece71c077999ecd", "X-Consumer-Username": "john", "X-Credential-Identifier": "cred-john-hmac-auth", "X-Custom-Header-A": "hello123", "X-Custom-Header-B": "world456", "X-Forwarded-Host": "127.0.0.1" }, "origin": "192.168.65.1, 103.97.2.206", "url": "http://127.0.0.1/get" } ``` ### Rate Limit with Anonymous Consumer[​](#rate-limit-with-anonymous-consumer "Direct link to Rate Limit with Anonymous Consumer") The following example demonstrates how you can configure different rate limiting policies by regular and anonymous consumers, where the anonymous consumer does not need to authenticate and has less quota. * Admin API * ADC * Ingress Controller Create a regular consumer `john` and configure the `limit-count` plugin to allow for a quota of 3 within a 30-second window: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john", "plugins": { "limit-count": { "count": 3, "time_window": 30, "rejected_code": 429, "policy": "local" } } }' ``` Create the `hmac-auth` credential for the consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-hmac-auth", "plugins": { "hmac-auth": { "key_id": "john-key", "secret_key": "john-secret-key" } } }' ``` Create an anonymous user `anonymous` and configure the `limit-count` plugin to allow for a quota of 1 within a 30-second window: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "anonymous", "plugins": { "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "policy": "local" } } }' ``` Create a route and configure the `hmac-auth` plugin to accept anonymous consumer `anonymous` from bypassing the authentication: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "hmac-auth-route", "uri": "/get", "methods": ["GET"], "plugins": { "hmac-auth": { "anonymous_consumer": "anonymous" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Configure consumers with different rate limits and a route that accepts anonymous users: adc.yaml ``` consumers: - username: john plugins: limit-count: count: 3 time_window: 30 rejected_code: 429 policy: local credentials: - name: hmac-auth type: hmac-auth config: key_id: john-key secret_key: john-secret-key - username: anonymous plugins: limit-count: count: 1 time_window: 30 rejected_code: 429 policy: local services: - name: anonymous-rate-limit-service routes: - name: hmac-auth-route uris: - /get methods: - GET plugins: hmac-auth: anonymous_consumer: anonymous upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Configure consumers with different rate limits and a route that accepts anonymous users: hmac-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: hmac-auth name: primary-cred config: key_id: john-key secret_key: john-secret-key plugins: - name: limit-count config: count: 3 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: anonymous spec: gatewayRef: name: apisix plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 policy: local --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: hmac-auth-plugin-config spec: plugins: - name: hmac-auth config: anonymous_consumer: aic_anonymous # namespace_consumername --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: hmac-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get method: GET filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: hmac-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f hmac-auth-ic.yaml ``` hmac-auth-apisix-crd.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: hmacAuth: value: key_id: john-key secret_key: john-secret-key plugins: - name: limit-count config: count: 3 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: anonymous spec: ingressClassName: apisix plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: hmac-auth-route spec: ingressClassName: apisix http: - name: hmac-auth-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: hmac-auth config: anonymous_consumer: aic_anonymous ``` Apply the configuration to your cluster: ``` kubectl apply -f hmac-auth-apisix-crd.yaml ``` Generate a signature. You can use the below Python snippet or other stack of your choice: hmac-sig-header-gen.py ``` import hmac import hashlib import base64 from datetime import datetime, timezone key_id = "john-key" # key id secret_key = b"john-secret-key" # secret key request_method = "GET" # HTTP method request_path = "/get" # route URI algorithm= "hmac-sha256" # can use other algorithms in allowed_algorithms # get current datetime in GMT # note: the signature will become invalid after the clock skew (default 300s) # you can regenerate the signature after it becomes invalid, or increase the clock # skew to prolong the validity within the advised security boundary gmt_time = datetime.now(timezone.utc).strftime('%a, %d %b %Y %H:%M:%S GMT') # construct the signing string (ordered) # the date and any subsequent custom headers should be lowercased and separated by a # single space character, i.e. `:` # https://datatracker.ietf.org/doc/html/draft-cavage-http-signatures-12#section-2.1.6 signing_string = ( f"{key_id}\n" f"{request_method} {request_path}\n" f"date: {gmt_time}\n" ) # create signature signature = hmac.new(secret_key, signing_string.encode('utf-8'), hashlib.sha256).digest() signature_base64 = base64.b64encode(signature).decode('utf-8') # construct the request headers headers = { "Date": gmt_time, "Authorization": ( f'Signature keyId="{key_id}",algorithm="{algorithm}",' f'headers="@request-target date",' f'signature="{signature_base64}"' ) } # print headers print(headers) ``` Run the script: ``` python3 hmac-sig-header-gen.py ``` You should see the request headers printed: ``` {'Date': 'Mon, 21 Oct 2024 17:31:18 GMT', 'Authorization': 'Signature keyId="john-key",algorithm="hmac-sha256",headers="@request-target date",signature="ztFfl9w7LmCrIuPjRC/DWSF4gN6Bt8dBBz4y+u1pzt8="'} ``` To verify, send five consecutive requests with the generated headers: ``` resp=$(seq 5 | xargs -I{} curl "http://127.0.0.1:9080/get" -H "Date: Mon, 21 Oct 2024 17:31:18 GMT" -H 'Authorization: Signature keyId="john-key",algorithm="hmac-sha256",headers="@request-target date",signature="ztFfl9w7LmCrIuPjRC/DWSF4gN6Bt8dBBz4y+u1pzt8="' -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that out of the 5 requests, 3 requests were successful (status code 200) while the others were rejected (status code 429). ``` 200: 3, 429: 2 ``` Send five anonymous requests: ``` resp=$(seq 5 | xargs -I{} curl "http://127.0.0.1:9080/get" -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that only one request was successful: ``` 200: 1, 429: 4 ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. ### Credentials[​](#credentials "Direct link to Credentials") The following are plugin attributes available for configurations on [credentials](https://docs.api7.ai/apisix/key-concepts/credentials.md). * key\_id string required *** Unique identifier for the consumer, which identifies the associated configurations such as the secret key. * secret\_key string required *** Secret key used to generate an HMAC. The key is encrypted with AES before being stored in etcd. You can also store it in an environment variable and reference it using the `env://` prefix, or in a secret manager such as HashiCorp Vault's [KV secrets engine](https://developer.hashicorp.com/vault/docs/secrets/kv), and reference it using the `secret://` prefix. For more information, see [secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md). ### Routes or Services[​](#routes-or-services "Direct link to Routes or Services") The following are plugin attributes available for configurations on [routes](https://docs.api7.ai/apisix/key-concepts/routes.md) or [services](https://docs.api7.ai/apisix/key-concepts/services.md). * allowed\_algorithms array\[string] default: `["hmac-sha1", "hmac-sha256", "hmac-sha512"]` *** The list of HMAC algorithms allowed. * clock\_skew integer default: `300` vaild vaule: greater than or equal to 1 *** Maximum allowable time difference in seconds between the client request's timestamp and APISIX server's current time. This helps account for discrepancies in time synchronization between the client’s and server’s clocks and protect against replay attacks. The timestamp in the `Date` header (must be in GMT format) will be used for the calculation. * signed\_headers array\[string] default: `["date"]` *** The list of headers whose values must be included in the client request's HMAC signature. In API7 Enterprise from version 3.10.0 and APISIX from version 3.17.0, this defaults to `["date"]`, so the `Date` header must be signed unless you override this field. If you enable `validate_request_body`, also include `digest` so the body digest is covered by the signature. * validate\_request\_body boolean *** If true, compare the request body with the `Digest` header. The plugin computes a SHA-256 digest of the body, base64-encodes it, and expects `Digest` to be `SHA-256=`. A missing or mismatched `Digest` header fails validation. This check does not bind the digest to the HMAC signature. Include `digest` in the signed headers if you want the body digest covered by the signature. * max\_req\_body\_size integer default: `524288 in API7 Enterprise 3.9.14, 3.10.0, and 3.10.1; 67108864 in API7 Enterprise from 3.9.15 and APISIX from 3.17.0` *** Maximum size in bytes of the request body that the plugin reads when validate*request*body is true. A request whose body exceeds this size is rejected (with HTTP 413 in API7 Enterprise from version 3.9.15 and APISIX from version 3.17.0). The default is 524288 bytes (512 KiB) in API7 Enterprise 3.9.14, 3.10.0, and 3.10.1, and 67108864 bytes (64 MiB) in API7 Enterprise from version 3.9.15 and APISIX from version 3.17.0. * hide\_credentials boolean default: `false` *** If true, do not pass the authorization request header to upstream services. * anonymous\_consumer string *** Anonymous consumer name. If configured, allow anonymous users to bypass the authentication. See [Rate Limit with Anonymous Consumer](https://docs.api7.ai/hub/hmac-auth.md#rate-limit-with-anonymous-consumer) for more details. * realm string default: `hmac` *** Realm in the [`WWW-Authenticate`](https://datatracker.ietf.org/doc/html/rfc7235#section-4.1) response header returned with a `401 Unauthorized` response due to authentication failure. For example: * If `realm` is set to `hmac-auth`, the 401 response will include the following header: ``` WWW-Authenticate: hmac realm="hmac-auth" ``` * If `realm` is not configured, the 401 response will include the following header: ``` WWW-Authenticate: hmac realm="hmac" ``` This parameter is available in API7 Enterprise version 3.9.2 and later, and in Apache APISIX version 3.15.0 and later. --- # http-logger The `http-logger` plugin pushes request and response logs as JSON objects to HTTP(S) servers in batches and supports the customization of log formats. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `http-logger` plugin for different scenarios. To follow along the examples, start a mock HTTP logging endpoint using [mockbin](https://mockbin.io) and note down the mockbin URL. ### Log Requests in Default Log Format[​](#log-requests-in-default-log-format "Direct link to Log Requests in Default Log Format") The following example demonstrates how you can configure the `http-logger` plugin on a route to log information of requests hitting the route. Create a route with the `http-logger` plugin and configure the plugin with your server URI: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "http-logger-route", "uri": "/anything", "plugins": { "http-logger": { "uri": "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: http-logger-route plugins: http-logger: uri: "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD http-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: http-logger-plugin-config spec: plugins: - name: http-logger config: uri: "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: http-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: http-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` http-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: http-logger-route spec: ingressClassName: apisix http: - name: http-logger-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: http-logger config: uri: "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" ``` Apply the configuration: ``` kubectl apply -f http-logger-ic.yaml ``` Send a request to the route: ``` curl "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. In your mockbin, you should see a log entry similar to the following: ``` [ { "upstream": "3.213.1.197:80", "server": { "hostname": "7d8d831179d4", "version": "3.9.0" }, "start_time": 1718291190508, "client_ip": "192.168.65.1", "response": { "status": 200, "headers": { "server": "APISIX/3.9.0", "content-length": "390", "access-control-allow-credentials": "true", "connection": "close", "date": "Thu, 13 Jun 2024 15:06:31 GMT", "access-control-allow-origin": "*", "content-type": "application/json" }, "size": 617 }, "latency": 1200.0000476837, "upstream_latency": 1133, "apisix_latency": 67.000047683716, "request": { "url": "http://127.0.0.1:9080/anything", "querystring": {}, "method": "GET", "uri": "/anything", "headers": { "accept": "*/*", "user-agent": "curl/8.6.0", "host": "127.0.0.1:9080" }, "size": 85 }, "service_id": "", "route_id": "http-logger-route" } ] ``` ### Log Request and Response Headers With Plugin Metadata[​](#log-request-and-response-headers-with-plugin-metadata "Direct link to Log Request and Response Headers With Plugin Metadata") The following example demonstrates how you can customize log format using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) and [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) to log specific headers from request and response. In APISIX, [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) is used to configure the common metadata fields of all plugin instances of the same plugin. It is useful when a plugin is enabled across multiple resources and requires a universal update to their metadata fields. First, create a route with the `http-logger` plugin and configure the plugin with your server URI: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "http-logger-route", "uri": "/anything", "plugins": { "http-logger": { "uri": "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: http-logger-route plugins: http-logger: uri: "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD http-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: http-logger-plugin-config spec: plugins: - name: http-logger config: uri: "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: http-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: http-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` http-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: http-logger-route spec: ingressClassName: apisix http: - name: http-logger-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: http-logger config: uri: "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" ``` Apply the configuration: ``` kubectl apply -f http-logger-ic.yaml ``` Next, configure the plugin metadata for `http-logger`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/http-logger" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "log_format": { "host": "$host", "@timestamp": "$time_iso8601", "client_ip": "$remote_addr", "env": "$http_env", "resp_content_type": "$sent_http_Content_Type" } }' ``` adc.yaml ``` plugin_metadata: - name: http-logger log_format: host: "$host" "@timestamp": "$time_iso8601" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: http-logger: log_format: host: "$host" "@timestamp": "$time_iso8601" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` Apply the configuration: ``` kubectl apply -f gatewayproxy.yaml ``` ❶ log the custom request header `env`. ❷ log the response header `Content-Type`. Send a request to the route with the `env` header: ``` curl "http://127.0.0.1:9080/anything" -H "env: dev" ``` You should receive an `HTTP/1.1 200 OK` response. In your mockbin, you should see a log entry similar to the following: ``` [ { "route_id": "http-logger-route", "client_ip": "192.168.65.1", "@timestamp": "2024-06-13T15:19:34+00:00", "host": "127.0.0.1", "env": "dev", "resp_content_type": "application/json" } ] ``` ### Log Request Bodies Conditionally[​](#log-request-bodies-conditionally "Direct link to Log Request Bodies Conditionally") The following example demonstrates how you can conditionally log request body. Create a route with the `http-logger` plugin as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "http-logger-route", "uri": "/anything", "plugins": { "http-logger": { "uri": "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/", "include_req_body": true, "include_req_body_expr": [["arg_log_body", "==", "yes"]] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: http-logger-route plugins: http-logger: uri: "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" include_req_body: true include_req_body_expr: - - "arg_log_body" - "==" - "yes" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD http-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: http-logger-plugin-config spec: plugins: - name: http-logger config: uri: "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" include_req_body: true include_req_body_expr: - - "arg_log_body" - "==" - "yes" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: http-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: http-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` http-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: http-logger-route spec: ingressClassName: apisix http: - name: http-logger-route match: paths: - /anything methods: - GET - POST upstreams: - name: httpbin-external-domain plugins: - name: http-logger config: uri: "https://669f05eb10ca49f18763e023312c3d77.api.mockbin.io/" include_req_body: true include_req_body_expr: - - "arg_log_body" - "==" - "yes" ``` Apply the configuration: ``` kubectl apply -f http-logger-ic.yaml ``` ❶ `include_req_body`: set to true to include request body. ❷ `include_req_body_expr`: only include request body if the URL query string `log_body` is `yes`. Send a request to the route with a URL query string satisfying the condition: ``` curl -i "http://127.0.0.1:9080/anything?log_body=yes" -X POST -d '{"env": "dev"}' ``` You should see the request body logged: ``` [ { "request": { "url": "http://127.0.0.1:9080/anything?log_body=yes", "querystring": { "log_body": "yes" }, "uri": "/anything?log_body=yes", ..., "body": "{\"env\": \"dev\"}", }, ... } ] ``` Send a request to the route without any URL query string: ``` curl -i "http://127.0.0.1:9080/anything" -X POST -d '{"env": "dev"}' ``` You should not observe the request body in the log. info If you have customized the `log_format` in addition to setting `include_req_body` or `include_resp_body` to `true`, the plugin would not include the bodies in the logs. As a workaround, you may be able to use the NGINX variable `$request_body` in the log format, such as: ``` { "http-logger": { ..., "log_format": {"body": "$request_body"} } } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * uri string required *** URI of the HTTP(S) server. * auth\_header string *** Authorization headers, if required by the HTTP(S) server. The value is encrypted with AES before being stored in etcd. * timeout integer default: `3` vaild vaule: greater than 0 *** Time to keep the connection alive after sending a request. * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. You can also configure log format on a global scale using the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md), which configures the log format for all `http-logger` plugin instances. If the log format configured on the individual plugin instance differs from the log format configured on plugin metadata, the log format configured on the individual plugin instance takes precedence. See the [example](https://docs.api7.ai/hub/http-logger.md#log-request-and-response-headers-with-plugin-metadata) for more details. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * include\_req\_body boolean default: `false` *** If true, include the request body in the log. Note that if the request body is too big to be kept in the memory, it can not be logged due to NGINX's limitations. * include\_req\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_req_body` is true. Request body would only be logged when the expressions configured here evaluate to true. * include\_resp\_body boolean default: `false` *** If true, include the response body in the log. * include\_resp\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_resp_body` is true. Response body would only be logged when the expressions configured here evaluate to true. * max\_req\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes to include in the log. If the request body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * max\_resp\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum response body size in bytes to include in the log. If the response body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * concat\_method string default: `json` vaild vaule: `json` or `new_line` *** Method to concatenate logs. When set to `json`, use `json.encode` for all pending logs. When set to `new_line`, also use `json.encode` but use the newline character ``to concatenate lines. * ssl\_verify boolean default: `false` *** If true, verify the server's SSL certificate. * name string default: `http logger` *** Unique identifier of the plugin for the batch processor. If you use [Prometheus](https://docs.api7.ai/hub/prometheus.md) to monitor APISIX metrics, the name is exported in `apisix_batch_process_entries`. * batch\_max\_size integer default: `1000` vaild vaule: greater than 0 *** Number of log entries allowed in one batch. Once reached, the batch is sent to the logging service. Setting this parameter to 1 enables immediate processing. * inactive\_timeout integer default: `5` vaild vaule: greater than 0 *** Maximum time in seconds to wait for new logs before sending the batch. The value should be smaller than `buffer_duration`. * buffer\_duration integer default: `60` vaild vaule: greater than 0 *** Maximum time in seconds from the earliest entry before sending the batch. * retry\_delay integer default: `1` vaild vaule: greater than or equal to 0 *** Time in seconds to wait before retrying a failed batch. * max\_retry\_count integer default: `0` vaild vaule: greater than or equal to 0 *** Maximum number of unsuccessful retries before dropping the log entries. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * max\_pending\_entries integer default: `` `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 and 3.10.6; none in API7 Enterprise 3.9.18 and 3.10.5 `` vaild vaule: greater than or equal to 1 *** Maximum number of entries waiting in the batch processor. New entries are discarded when the backlog reaches the limit. Introduced in API7 Enterprise 3.8.17 and APISIX 3.15.0. The default changed to `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line. In API7 Enterprise 3.9.18 and 3.10.5, and in earlier APISIX versions, omitting the parameter leaves the backlog unlimited. See [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md#configure-the-pending-entry-limit) for sizing and verification guidance. --- # ip-restriction The `ip-restriction` plugin supports restricting access to upstream resources by IP addresses, through either configuring a whitelist or blacklist of IP addresses. Restricting IP to resources helps prevent unauthorized access and harden API security. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure the `ip-restriction` plugin for different scenarios. ### Restrict Access by Whitelisting[​](#restrict-access-by-whitelisting "Direct link to Restrict Access by Whitelisting") The following example demonstrates how you can whitelist a list of IP addresses that should have access to the upstream resource and customize the error message for access denial. * Admin API * ADC * Ingress Controller Create a route with the `ip-restriction` plugin as such: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "ip-restriction-route", "uri": "/anything", "plugins": { "ip-restriction": { "whitelist": [ "192.168.0.1/24" ], "message": "Access denied" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: ip-restriction-service routes: - name: ip-restriction-route uris: - /anything plugins: ip-restriction: whitelist: - "192.168.0.1/24" message: "Access denied" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Client IP with kubectl port-forward When testing locally with `kubectl port-forward`, APISIX sees `127.0.0.1` as the client IP regardless of your machine's actual IP address. Make sure your whitelist or blacklist includes `127.0.0.1` when testing in this setup. In production with a NodePort or LoadBalancer service, APISIX receives the actual client IP. * Gateway API * APISIX CRD ip-restriction-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: ip-restriction-plugin-config spec: plugins: - name: ip-restriction config: whitelist: - "192.168.0.1/24" message: "Access denied" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: ip-restriction-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: ip-restriction-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f ip-restriction-ic.yaml ``` ip-restriction-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: ip-restriction-route spec: ingressClassName: apisix http: - name: ip-restriction-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: ip-restriction enable: true config: whitelist: - "192.168.0.1/24" message: "Access denied" ``` Apply the configuration to your cluster: ``` kubectl apply -f ip-restriction-ic.yaml ``` ❶ Replace with the IP addresses you would like to whitelist. ❷ Customize the error message for when the access is denied. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` If your IP is allowed, you should receive an `HTTP/1.1 200 OK` response. If not, you should receive an `HTTP/1.1 403 Forbidden` response with the following error message: ``` {"message":"Access denied"} ``` ### Restrict Access Using Modified IP[​](#restrict-access-using-modified-ip "Direct link to Restrict Access Using Modified IP") The following example demonstrates how you can modify the IP used for IP restriction, using the `real-ip` plugin. This is particularly useful if APISIX is behind a reverse proxy and the real client IP is not available to APISIX. * Admin API * ADC * Ingress Controller Create a route as follows: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "ip-restriction-route", "uri": "/anything", "plugins": { "ip-restriction": { "whitelist": [ "192.168.1.241" ] }, "real-ip": { "source": "arg_realip" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` ❶ Obtain client IP address from the URL parameter `realip` using the [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). adc.yaml ``` services: - name: ip-restriction-service routes: - name: ip-restriction-route uris: - /anything plugins: ip-restriction: whitelist: - "192.168.1.241" real-ip: source: arg_realip upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ Obtain client IP address from the URL parameter `realip` using the [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD ip-restriction-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: ip-restriction-realip-plugin-config spec: plugins: - name: ip-restriction config: whitelist: - "192.168.1.241" - name: real-ip config: source: arg_realip --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: ip-restriction-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: ip-restriction-realip-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` ip-restriction-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: ip-restriction-route spec: ingressClassName: apisix http: - name: ip-restriction-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: ip-restriction enable: true config: whitelist: - "192.168.1.241" - name: real-ip enable: true config: source: arg_realip ``` ❶ Obtain client IP address from the URL parameter `realip` using the [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). Apply the configuration to your cluster: ``` kubectl apply -f ip-restriction-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything?realip=192.168.1.241" ``` You should receive an `HTTP/1.1 200 OK` response. Send another request with a different IP address: ``` curl -i "http://127.0.0.1:9080/anything?realip=192.168.10.24" ``` You should receive an `HTTP/1.1 403 Forbidden` response. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * whitelist array\[string] *** List of IPs to whitelist. Support IPv4, IPv6, and CIDR notations. At least one of the `whitelist` and `blacklist` should be configured, but they cannot be configured at the same time. * blacklist array\[string] *** List of IPs to blacklist. Support IPv4, IPv6, and CIDR notations. At least one of the `whitelist` and `blacklist` should be configured, but they cannot be configured at the same time. * message string default: `Your IP address is not allowed` *** Message returned when the IP address is denied access. * response\_code integer default: `403` vaild vaule: between 403 and 404 inclusive *** HTTP response code returned when the request is rejected due to IP address restriction. Available in API7 Enterprise from version 3.10.3. --- # jwe-decrypt The `jwe-decrypt` plugin reads a five-part compact token from a request header. It selects a [consumer](https://docs.api7.ai/apisix/key-concepts/consumers.md) by the token's `kid` and decrypts the encrypted payload with AES-256-GCM. Before proxying the request, it writes the plaintext to a configured header. You can enable the plugin on APISIX [routes](https://docs.api7.ai/apisix/key-concepts/routes.md) or [services](https://docs.api7.ai/apisix/key-concepts/services.md). Configure a 32-byte decryption secret on the consumer. The token resembles [JWE Compact Serialization](https://datatracker.ietf.org/doc/html/rfc7516#section-3.1), but it is a plugin-specific format. The implementation reads `kid` from the decoded header; it does not validate the `alg` or `enc` fields, and it does not use the protected-header segment as AES-GCM additional authenticated data (AAD). Standard RFC 7516 JWE libraries are therefore not directly interoperable. Generate tokens with the exact format described below, use a fixed trusted token generator, and do not treat header fields as authenticated. caution The decrypted plaintext is forwarded in a request header. For sensitive plaintext, do not rely on an APISIX HTTPS upstream alone: APISIX does not verify the upstream server certificate when proxying to standard HTTPS upstreams. Send the request over an authenticated, protected network path, such as through a proxy or service mesh that validates the upstream server's identity. Restrict access to the upstream and avoid logging the configured forwarding header. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can work with the `jwe-decrypt` plugin for different scenarios. ### Decrypt Data from the Plugin Token[​](#decrypt-data-from-the-plugin-token "Direct link to Decrypt Data from the Plugin Token") The following example demonstrates how to decrypt a plugin token. Generate tokens outside APISIX, configure the matching decryption key on a consumer, and create a route with `jwe-decrypt` to decrypt the authorization header. * Admin API * ADC * Ingress Controller Create a consumer with `jwe-decrypt` and configure the decryption key: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jack", "plugins": { "jwe-decrypt": { "key": "jack-key", "secret": "key-length-should-be-32-chars123" } } }' ``` Create a route with `jwe-decrypt` to decrypt the authorization header: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "jwe-decrypt-route", "uri": "/anything/jwe", "plugins": { "jwe-decrypt": { "header": "Authorization", "forward_header": "Authorization" } }, "upstream": { "type": "roundrobin", "scheme": "https", "nodes": { "httpbin.org:443": 1 } } }' ``` adc.yaml ``` consumers: - username: jack plugins: jwe-decrypt: key: jack-key secret: key-length-should-be-32-chars123 services: - name: jwe-decrypt-service routes: - name: jwe-decrypt-route uris: - /anything/jwe plugins: jwe-decrypt: header: Authorization forward_header: Authorization upstream: type: roundrobin scheme: https nodes: - host: httpbin.org port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` The following Ingress Controller configurations use public HTTPBin only with the non-sensitive demonstration payload shown on this page. Before forwarding real decrypted data, replace it with a controlled upstream and use an authenticated, protected network path. APISIX does not verify the upstream server certificate when proxying to standard HTTPS upstreams; use a proxy or service mesh that validates the upstream server's identity. * Gateway API * APISIX CRD jwe-decrypt-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jack spec: gatewayRef: name: apisix plugins: - name: jwe-decrypt config: key: jack-key secret: key-length-should-be-32-chars123 --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: jwe-decrypt-plugin-config spec: plugins: - name: jwe-decrypt config: header: Authorization forward_header: Authorization --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: jwe-decrypt-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything/jwe filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: jwe-decrypt-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f jwe-decrypt-ic.yaml ``` jwe-decrypt-apisix-crd.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jack spec: ingressClassName: apisix plugins: - name: jwe-decrypt config: key: jack-key secret: key-length-should-be-32-chars123 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: jwe-decrypt-route spec: ingressClassName: apisix http: - name: jwe-decrypt-route match: paths: - /anything/jwe upstreams: - name: httpbin-external-domain plugins: - name: jwe-decrypt config: header: Authorization forward_header: Authorization ``` Apply the configuration to your cluster: ``` kubectl apply -f jwe-decrypt-apisix-crd.yaml ``` Generate plugin tokens outside APISIX by encrypting the payload with AES-256-GCM without protected-header AAD and using the consumer secret as the key. Standard RFC 7516 libraries normally authenticate the protected header as AAD and are not directly interoperable with this plugin. Use this exact token structure: ``` base64url(header)..base64url(iv).base64url(ciphertext).base64url(tag) ``` where the header is `{"alg":"dir","enc":"A256GCM","kid":""}`. These fields describe the intended algorithm and identify the consumer, but the current plugin does not authenticate or validate them. Use a unique, randomly generated IV for each token; never reuse an IV with the same key. APISIX decrypts the encrypted payload and authentication tag directly with AES-256-GCM. It does not pass the protected header as AAD. A token generated with standard protected-header AAD is rejected with `failed to decrypt JWE token`. Send a request to the route with the encrypted plugin token in the `Authorization` header. For example, the following token encrypts the payload `{"uid":10000,"uname":"test"}` for the consumer key `jack-key` with the secret configured above: ``` curl "http://127.0.0.1:9080/anything/jwe" -H 'Authorization: eyJraWQiOiJqYWNrLWtleSIsImFsZyI6ImRpciIsImVuYyI6IkEyNTZHQ00ifQ..vi29KBCQKcVmPwTT.VToyPMFbq-ZY05MIpntP1N3AmYeq3zELQ0B6iQ.vuTPG2ODc-DjUTjNCzfA2A' ``` You should see a response similar to the following, where the `Authorization` header shows the plaintext of the payload: ``` { "args": {}, "data": "", "files": {}, "form": {}, "headers": { "Accept": "*/*", "Authorization": "{\"uid\":10000,\"uname\":\"test\"}", "Host": "127.0.0.1", "User-Agent": "curl/8.1.2", "X-Amzn-Trace-Id": "Root=1-6510f2c3-1586ec011a22b5094dbe1896", "X-Forwarded-Host": "127.0.0.1" }, "json": null, "method": "GET", "origin": "127.0.0.1, 119.143.79.94", "url": "http://127.0.0.1/anything/jwe" } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. ### Consumers[​](#consumers "Direct link to Consumers") The following are plugin attributes available for configurations on [consumers](https://docs.api7.ai/apisix/key-concepts/consumers.md). * key string required *** A unique key that identifies the credential for a consumer. * secret string required vaild vaule: 32 bytes *** A shared symmetric key. Use a [secret reference](https://docs.api7.ai/apisix/key-concepts/secrets.md), such as `$env://...` or `$secret://...`. * is\_base64\_encoded boolean default: `false` *** Set to true if the secret is base64url encoded. The decoded secret must still be 32 bytes. ### Routes or Services[​](#routes-or-services "Direct link to Routes or Services") The following are plugin attributes available for configurations on [routes](https://docs.api7.ai/apisix/key-concepts/routes.md) or [services](https://docs.api7.ai/apisix/key-concepts/services.md). * header string required default: `Authorization` *** The header to get the token from. * forward\_header string required default: `Authorization` *** Name of the header that passes the plaintext to the upstream. * strict boolean default: `true` *** If true, return a 403 error when the encrypted plugin token is missing. If false, continue when the token is not found. --- # jwt-auth The `jwt-auth` plugin supports the use of [JSON Web Token (JWT)](https://jwt.io/) as a mechanism for clients to authenticate themselves before accessing upstream resources. Once enabled, JWT credentials are configured on [consumers](https://docs.api7.ai/apisix/key-concepts/consumers.md), and clients carry a signed token to identify themselves to APISIX. The token can be included in the request URL query string, request header, or cookie. APISIX will then verify the token to determine if a request should be allowed or denied to access upstream resources. When a consumer is successfully authenticated, APISIX adds additional headers, such as `X-Consumer-Username`, `X-Credential-Identifier`, and other consumer custom headers if configured, to the request, before proxying it to the upstream service. The upstream service will be able to differentiate between consumers and implement additional logics as needed. If any of these values is not available, the corresponding header will not be added. About X-Consumer-Username When consumers are configured using the Ingress Controller, the consumer name is generated in the format `namespace_consumername`. As a result, the `X-Consumer-Username` header will also follow this format instead of just `consumername`. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can work with the `jwt-auth` plugin for different scenarios. ### Use JWT for Consumer Authentication[​](#use-jwt-for-consumer-authentication "Direct link to Use JWT for Consumer Authentication") The following example demonstrates how to implement JWT for consumer key authentication. * Admin API * ADC * Ingress Controller Create a consumer `jack`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jack" }' ``` Create `jwt-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jack/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jack-jwt-auth", "plugins": { "jwt-auth": { "key": "jack-key", "secret": "jack-hs256-secret-that-is-very-long" } } }' ``` Create a route with `jwt-auth` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "jwt-route", "uri": "/headers", "plugins": { "jwt-auth": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `jwt-auth` credential and a route with `jwt-auth` plugin configured as such: adc.yaml ``` consumers: - username: jack credentials: - name: jwt-auth type: jwt-auth config: key: jack-key secret: jack-hs256-secret-that-is-very-long services: - name: jwt-auth-service routes: - name: jwt-route uris: - /headers plugins: jwt-auth: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create a consumer with `jwt-auth` credential and a route with `jwt-auth` plugin configured as such: jwt-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jack spec: gatewayRef: name: apisix credentials: - type: jwt-auth name: primary-cred config: key: jack-key secret: jack-hs256-secret-that-is-very-long --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: jwt-auth-plugin-config spec: plugins: - name: jwt-auth config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: jwt-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: jwt-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f jwt-auth-ic.yaml ``` Create a consumer with `jwt-auth` credential using HS256 and a route with `jwt-auth` plugin enabled as such: jwt-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jack spec: ingressClassName: apisix authParameter: jwtAuth: value: key: jack-key secret: jack-hs256-secret-that-is-very-long --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: jwt-route spec: ingressClassName: apisix http: - name: jwt-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: jwt-auth enable: true config: _meta: disable: false ``` Apply the configuration to your cluster: ``` kubectl apply -f jwt-auth-ic.yaml ``` To issue a JWT for `jack`, you could use [JWT.io's JWT encoder](https://jwt.io) or other utilities. If you are using [JWT.io's JWT encoder](https://jwt.io), do the following: * Fill in `HS256` as the algorithm. * Update the secret in the **Valid secret** section to be `jack-hs256-secret-that-is-very-long`. * Update payload with consumer key `jack-key`; and add `exp` or `nbf` in UNIX timestamp. note When `claims_to_verify` is a nonempty list, every listed claim is required and validated. When it is unset or empty, `exp` and `nbf` are validated when present, but neither claim is required. Your payload should look similar to the following: ``` { "key": "jack-key", "nbf": 1729132271 } ``` Copy the generated JWT and save to a variable: ``` export jwt_token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqYWNrLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.UEPXy5jpid624T1XpfjM0PLY73LZPjV3Qt8yZ92kVuU ``` Send a request to the route with the JWT in the `Authorization` header: ``` curl -i "http://127.0.0.1:9080/headers" -H "Authorization: ${jwt_token}" ``` You should receive an `HTTP/1.1 200 OK` response similar to the following: ``` { "headers": { "Accept": "*/*", "Authorization": "eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJleHAiOjE3MjY2NDk2NDAsImtleSI6ImphY2sta2V5In0.kdhumNWrZFxjUvYzWLt4lFr546PNsr9TXuf0Az5opoM", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-66ea951a-4d740d724bd2a44f174d4daf", "X-Consumer-Username": "jack", "X-Credential-Identifier": "cred-jack-jwt-auth", "X-Forwarded-Host": "127.0.0.1" } } ``` Send a request with an invalid token: ``` curl -i "http://127.0.0.1:9080/headers" -H "Authorization: eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJleHAiOjE3MjY2NDk2NDAsImtleSI6ImphY2sta2V5In0.kdhumNWrZFxjU_random_random" ``` You should receive an `HTTP/1.1 401 Unauthorized` response similar to the following: ``` {"message":"failed to verify jwt"} ``` ### Carry JWT in Request Header, Query String, or Cookie[​](#carry-jwt-in-request-header-query-string-or-cookie "Direct link to Carry JWT in Request Header, Query String, or Cookie") The following example demonstrates how to accept JWT in specified header, query string, and cookie. * Admin API * ADC * Ingress Controller Create a consumer `jack`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jack" }' ``` Create `jwt-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jack/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jack-jwt-auth", "plugins": { "jwt-auth": { "key": "jack-key", "secret": "jack-hs256-secret-that-is-very-long" } } }' ``` Create a route with `jwt-auth` plugin, and specify the request parameters carrying the token: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "jwt-route", "uri": "/get", "plugins": { "jwt-auth": { "header": "jwt-auth-header", "query": "jwt-query", "cookie": "jwt-cookie" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `jwt-auth` credential and a route with `jwt-auth` plugin configured as such: adc.yaml ``` consumers: - username: jack credentials: - name: jwt-auth type: jwt-auth config: key: jack-key secret: jack-hs256-secret-that-is-very-long services: - name: jwt-auth-service routes: - name: jwt-route uris: - /get plugins: jwt-auth: header: jwt-auth-header query: jwt-query cookie: jwt-cookie upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create a consumer with `jwt-auth` credential and a route with `jwt-auth` plugin configured as such: jwt-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jack spec: gatewayRef: name: apisix credentials: - type: jwt-auth name: primary-cred config: key: jack-key secret: jack-hs256-secret-that-is-very-long --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: jwt-auth-plugin-config spec: plugins: - name: jwt-auth config: _meta: disable: false header: jwt-auth-header query: jwt-query cookie: jwt-cookie --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: jwt-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: jwt-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f jwt-auth-ic.yaml ``` Create a consumer with `jwt-auth` credential and a route with `jwt-auth` plugin configured as such: jwt-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jack spec: ingressClassName: apisix authParameter: jwtAuth: value: key: jack-key secret: jack-hs256-secret-that-is-very-long --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: jwt-route spec: ingressClassName: apisix http: - name: jwt-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: jwt-auth enable: true config: header: jwt-auth-header query: jwt-query cookie: jwt-cookie ``` Apply the configuration to your cluster: ``` kubectl apply -f jwt-auth-ic.yaml ``` To issue a JWT for `jack`, you could use [JWT.io's JWT encoder](https://jwt.io) or other utilities. If you are using [JWT.io's JWT encoder](https://jwt.io), do the following: * Fill in `HS256` as the algorithm. * Update the secret in the **Valid secret** section to be `jack-hs256-secret-that-is-very-long`. * Update payload with consumer key `jack-key`; and add `exp` or `nbf` in UNIX timestamp. note When `claims_to_verify` is a nonempty list, every listed claim is required and validated. When it is unset or empty, `exp` and `nbf` are validated when present, but neither claim is required. Your payload should look similar to the following: ``` { "key": "jack-key", "nbf": 1729132271 } ``` Copy the generated JWT and save to a variable: ``` export jwt_token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqYWNrLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.UEPXy5jpid624T1XpfjM0PLY73LZPjV3Qt8yZ92kVuU ``` #### Verify With JWT in Header[​](#verify-with-jwt-in-header "Direct link to Verify With JWT in Header") Sending request with JWT in the header: ``` curl -i "http://127.0.0.1:9080/get" -H "jwt-auth-header: ${jwt_token}" ``` You should receive an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "Jwt-Auth-Header": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqYWNrLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.UEPXy5jpid624T1XpfjM0PLY73LZPjV3Qt8yZ92kVuU", ... }, ... } ``` #### Verify With JWT in Query String[​](#verify-with-jwt-in-query-string "Direct link to Verify With JWT in Query String") Sending request with JWT in the query string: ``` curl -i "http://127.0.0.1:9080/get?jwt-query=${jwt_token}" ``` You should receive an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": { "jwt-query": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqYWNrLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.UEPXy5jpid624T1XpfjM0PLY73LZPjV3Qt8yZ92kVuU" }, "headers": { "Accept": "*/*", ... }, "origin": "127.0.0.1, 183.17.233.107", "url": "http://127.0.0.1/get?jwt-query=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJrZXkiOiJ1c2VyLWtleSIsImV4cCI6MTY5NTEyOTA0NH0.EiktFX7di_tBbspbjmqDKoWAD9JG39Wo_CAQ1LZ9voQ" } ``` #### Verify With JWT in Cookie[​](#verify-with-jwt-in-cookie "Direct link to Verify With JWT in Cookie") Sending request with JWT in the cookie: ``` curl -i "http://127.0.0.1:9080/get" --cookie jwt-cookie=${jwt_token} ``` You should receive an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "headers": { "Accept": "*/*", "Cookie": "jwt-cookie=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqYWNrLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.UEPXy5jpid624T1XpfjM0PLY73LZPjV3Qt8yZ92kVuU", ... }, ... } ``` ### Manage Secrets in Environment Variables[​](#manage-secrets-in-environment-variables "Direct link to Manage Secrets in Environment Variables") The following example demonstrates how to save `jwt-auth` consumer key to an environment variable and reference it in configuration. APISIX supports referencing system and user environment variables configured through the [NGINX `env` directive](https://nginx.org/en/docs/ngx_core_module.html#env). Save the key to an environment variable: ``` export JACK_JWT_SECRET=jack-hs256-secret-that-is-very-long ``` tip If you are running APISIX in Docker, you should set the environment variable using the `-e` flag when starting the container. If you are running APISIX on Kubernetes, see the Ingress Controller tab for more details. * Admin API * ADC Create a consumer `jack`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jack" }' ``` Create `jwt-auth` credential for the consumer and reference the environment variable: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jack/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jack-jwt-auth", "plugins": { "jwt-auth": { "key": "jack-key", "secret": "$env://JACK_JWT_SECRET" } } }' ``` Create a route with `jwt-auth` enabled: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "jwt-route", "uri": "/get", "plugins": { "jwt-auth": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `jwt-auth` credential referencing an environment variable and a route with `jwt-auth` plugin enabled as such: adc.yaml ``` consumers: - username: jack credentials: - name: jwt-auth type: jwt-auth config: key: jack-key secret: $env://JACK_JWT_SECRET services: - name: jwt-auth-service routes: - name: jwt-route uris: - /get plugins: jwt-auth: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` To issue a JWT for `jack`, you could use [JWT.io's JWT encoder](https://jwt.io) or other utilities. If you are using [JWT.io's JWT encoder](https://jwt.io), do the following: * Fill in `HS256` as the algorithm. * Update the secret in the **Valid secret** section to be `jack-hs256-secret-that-is-very-long`. * Update payload with consumer key `jack-key`; and add `exp` or `nbf` in UNIX timestamp. note When `claims_to_verify` is a nonempty list, every listed claim is required and validated. When it is unset or empty, `exp` and `nbf` are validated when present, but neither claim is required. Your payload should look similar to the following: ``` { "key": "jack-key", "nbf": 1729132271 } ``` Copy the generated JWT and save to a variable: ``` export jwt_token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqYWNrLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.UEPXy5jpid624T1XpfjM0PLY73LZPjV3Qt8yZ92kVuU ``` Sending request with JWT in the header: ``` curl -i "http://127.0.0.1:9080/get" -H "Authorization: ${jwt_token}" ``` You should receive an `HTTP/1.1 200 OK` response. ### Manage Secrets in Secret Manager[​](#manage-secrets-in-secret-manager "Direct link to Manage Secrets in Secret Manager") The following example demonstrates how to manage `jwt-auth` consumer key in [HashiCorp Vault](https://www.vaultproject.io) and reference it in plugin configuration. Start a Vault development server in Docker: ``` docker run -d \ --name vault \ -p 8200:8200 \ --cap-add IPC_LOCK \ -e VAULT_DEV_ROOT_TOKEN_ID=root \ -e VAULT_DEV_LISTEN_ADDRESS=0.0.0.0:8200 \ vault:1.9.0 \ vault server -dev ``` APISIX currently supports [Vault KV engine version 1](https://developer.hashicorp.com/vault/docs/secrets/kv#kv-version-1). Enable it in Vault: ``` docker exec -i vault sh -c "VAULT_TOKEN='root' VAULT_ADDR='http://0.0.0.0:8200' vault secrets enable -path=kv -version=1 kv" ``` You should see a response similar to the following: ``` Success! Enabled the kv secrets engine at: kv/ ``` * Admin API * ADC Create a [secret](https://docs.api7.ai/apisix/key-concepts/secrets.md) and configure the Vault address and other connection information: ``` curl "http://127.0.0.1:9180/apisix/admin/secrets/vault/jwt" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "uri": "http://127.0.0.1:8200", "prefix": "kv/apisix", "token": "root" }' ``` ❶ Adjust the Vault address accordingly. Create a consumer `jack`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jack" }' ``` Create `jwt-auth` credential for the consumer and reference the secret: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jack/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jack-jwt-auth", "plugins": { "jwt-auth": { "key": "jwt-vault-key", "secret": "$secret://vault/jwt/jack/jwt-secret" } } }' ``` Create a route with `jwt-auth` enabled: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "jwt-route", "uri": "/get", "plugins": { "jwt-auth": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a secret and configure the Vault address: adc.yaml ``` secrets: - name: vault-jwt vault: url: http://127.0.0.1:8200 prefix: kv/apisix token: root consumers: - username: jack credentials: - name: jwt-auth type: jwt-auth config: key: jwt-vault-key secret: $secret://vault-jwt/jack/jwt-secret services: - name: jwt-auth-service routes: - name: jwt-route uris: - /get plugins: jwt-auth: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ Adjust the Vault address accordingly. ❷ Reference the secret in the secret manager. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Set `jwt-auth` key value to be `vault-hs256-secret-that-is-very-long` in Vault: ``` docker exec -i vault sh -c "VAULT_TOKEN='root' VAULT_ADDR='http://0.0.0.0:8200' vault kv put kv/apisix/jack jwt-secret=vault-hs256-secret-that-is-very-long" ``` You should see a response similar to the following: ``` Success! Data written to: kv/apisix/jack ``` To issue a JWT, you could use [JWT.io's JWT encoder](https://jwt.io) or other utilities. If you are using [JWT.io's JWT encoder](https://jwt.io), do the following: * Fill in `HS256` as the algorithm. * Update the secret in the **Valid secret** section to be `vault-hs256-secret-that-is-very-long`. * Update payload with consumer key `jwt-vault-key`; and add `exp` or `nbf` in UNIX timestamp. note When `claims_to_verify` is a nonempty list, every listed claim is required and validated. When it is unset or empty, `exp` and `nbf` are validated when present, but neither claim is required. Your payload should look similar to the following: ``` { "key": "jwt-vault-key", "nbf": 1729132271 } ``` Copy the generated JWT and save to a variable: ``` export jwt_token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqd3QtdmF1bHQta2V5IiwibmJmIjoxNzI5MTMyMjcxfQ.i2pLj7QcQvnlSjB7iV5V522tIV43boQRtee7L0rwlkQ ``` Send a request with the token in the header: ``` curl -i "http://127.0.0.1:9080/get" -H "Authorization: ${jwt_token}" ``` You should receive an `HTTP/1.1 200 OK` response. ### Sign JWT with RS256 Algorithm[​](#sign-jwt-with-rs256-algorithm "Direct link to Sign JWT with RS256 Algorithm") The following example demonstrates how you can use asymmetric algorithms, such as RS256, to sign and validate JWT when implementing JWT for consumer authentication. You will be generating RSA key pairs using [openssl](https://openssl-library.org/source/) and generating JWT using [JWT.io](https://jwt.io) to better understand the composition of JWT. Generate a 2048-bit RSA private key and extract the corresponding public key in PEM format: ``` openssl genrsa -out jwt-rsa256-private.pem 2048 openssl rsa -in jwt-rsa256-private.pem -pubout -out jwt-rsa256-public.pem ``` You should see `jwt-rsa256-private.pem` and `jwt-rsa256-public.pem` generated in your current working directory. Visit [JWT.io's JWT encoder](https://jwt.io) and do the following: * Fill in `RS256` as the algorithm. * Copy and paste the private key content into the **SIGN JWT: PRIVATE KEY** section. * Update payload with consumer key `jack-key`; and add `exp` or `nbf` in UNIX timestamp. Your payload should look similar to the following: ``` { "key": "jack-key", "nbf": 1729132271 } ``` note When `claims_to_verify` is a nonempty list, every listed claim is required and validated. When it is unset or empty, `exp` and `nbf` are validated when present, but neither claim is required. Copy the generated JWT and save to a variable: ``` export jwt_token=eyJhbGciOiJSUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqYWNrLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.K-I13em84kAcyH1jfIJl7ls_4jlwg1GzEzo5_xrDu-3wt3Xa3irS6naUsWpxX-a-hmcZZxRa9zqunqQjUP4kvn5e3xg2f_KyCR-_ZbwqYEPk3bXeFV1l4iypv6z5L7W1Niharun-dpMU03b1Tz64vhFx6UwxNL5UIZ7bunDAo_BXZ7Xe8rFhNHvIHyBFsDEXIBgx8lNYMq8QJk3iKxZhZZ5Om7lgYjOOKRgew4WkhBAY0v1AkO77nTlvSK0OEeeiwhkROyntggyx-S-U222ykMQ6mBLxkP4Cq5qHwXD8AUcLk5mhEij-3QhboYnt7yhKeZ3wDSpcjDvvL2aasC25ng ``` * Admin API * ADC * Ingress Controller Create a consumer `jack`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jack" }' ``` Create `jwt-auth` credential for the consumer and configure the RSA keys: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jack/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jack-jwt-auth", "plugins": { "jwt-auth": { "key": "jack-key", "algorithm": "RS256", "public_key": "-----BEGIN PUBLIC KEY-----\nMIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAoTxe7ZPycrEP0SK4OBA2\n0OUQsDN9gSFSHVvx/t++nZNrFxzZnV6q6/TRsihNXUIgwaOu5icFlIcxPL9Mf9UJ\na5/XCQExp1TxpuSmjkhIFAJ/x5zXrC8SGTztP3SjkhYnQO9PKVXI6ljwgakVCfpl\numuTYqI+ev7e45NdK8gJoJxPp8bPMdf8/nHfLXZuqhO/btrDg1x+j7frDNrEw+6B\nCK2SsuypmYN+LwHfaH4Of7MQFk3LNIxyBz0mdbsKJBzp360rbWnQeauWtDymZxLT\nATRNBVyl3nCNsURRTkc7eyknLaDt2N5xTIoUGHTUFYSdE68QWmukYMVGcEHEEPkp\naQIDAQAB\n-----END PUBLIC KEY-----" } } }' ``` ❶ Configure the consumer key to be `jack-key`. ❷ Configure the JWT signing algorithm to be `RS256`. ❸ Configure the RSA public key. tip You should add a newline character after the opening line and before the closing line, for example `-----BEGIN PUBLIC KEY-----\n......\n-----END PUBLIC KEY-----`. The key content can be directly concatenated. Create a route with the `jwt-auth` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "jwt-route", "uri": "/headers", "plugins": { "jwt-auth": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `jwt-auth` credential using RS256 algorithm and a route with `jwt-auth` plugin enabled as such: adc.yaml ``` consumers: - username: jack credentials: - name: jwt-auth type: jwt-auth config: key: jack-key algorithm: RS256 public_key: | -----BEGIN PUBLIC KEY----- MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAoTxe7ZPycrEP0SK4OBA2 0OUQsDN9gSFSHVvx/t++nZNrFxzZnV6q6/TRsihNXUIgwaOu5icFlIcxPL9Mf9UJ a5/XCQExp1TxpuSmjkhIFAJ/x5zXrC8SGTztP3SjkhYnQO9PKVXI6ljwgakVCfpl umuTYqI+ev7e45NdK8gJoJxPp8bPMdf8/nHfLXZuqhO/btrDg1x+j7frDNrEw+6B CK2SsuypmYN+LwHfaH4Of7MQFk3LNIxyBz0mdbsKJBzp360rbWnQeauWtDymZxLT ATRNBVyl3nCNsURRTkc7eyknLaDt2N5xTIoUGHTUFYSdE68QWmukYMVGcEHEEPkp aQIDAQAB -----END PUBLIC KEY----- services: - name: jwt-auth-service routes: - name: jwt-route uris: - /headers plugins: jwt-auth: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create a consumer with `jwt-auth` credential using RS256 algorithm and a route with `jwt-auth` plugin enabled as such: jwt-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jack spec: gatewayRef: name: apisix credentials: - type: jwt-auth name: primary-cred config: key: jack-key algorithm: RS256 public_key: | -----BEGIN PUBLIC KEY----- MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAoTxe7ZPycrEP0SK4OBA2 0OUQsDN9gSFSHVvx/t++nZNrFxzZnV6q6/TRsihNXUIgwaOu5icFlIcxPL9Mf9UJ a5/XCQExp1TxpuSmjkhIFAJ/x5zXrC8SGTztP3SjkhYnQO9PKVXI6ljwgakVCfpl umuTYqI+ev7e45NdK8gJoJxPp8bPMdf8/nHfLXZuqhO/btrDg1x+j7frDNrEw+6B CK2SsuypmYN+LwHfaH4Of7MQFk3LNIxyBz0mdbsKJBzp360rbWnQeauWtDymZxLT ATRNBVyl3nCNsURRTkc7eyknLaDt2N5xTIoUGHTUFYSdE68QWmukYMVGcEHEEPkp aQIDAQAB -----END PUBLIC KEY----- --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: jwt-auth-plugin-config spec: plugins: - name: jwt-auth config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: jwt-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: jwt-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f jwt-auth-ic.yaml ``` Create a consumer with `jwt-auth` credential using RS256 algorithm and a route with `jwt-auth` plugin enabled as such: jwt-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jack spec: ingressClassName: apisix authParameter: jwtAuth: value: key: jack-key algorithm: RS256 public_key: | -----BEGIN PUBLIC KEY----- MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAyBhBBT5u2BtQs3+s2nnq IXq9DRD8rWrmuk9lTI+rvzELaPZYzT7YxhBGuRJmbW+RnrIIB6dG6v9Kpn18qsvi 3u6UfsXKtoXckdk2tTCXSweNg1rzR9Szf/TxLSoi3KqA/0b/l9DqO9LYiWacEGgS mqs0bCKtvxq+0TGQfuPHJiapvzgPTT1CYAp84CYDvyIo6d4NJOiPPSTEb1jxagSq eLGZ3LVLZjSOC1kP4rbZP5U2VBMbkAtPtdFB1rOTCLykOQrH5eJxYxMkgiaDe9Da ZilQ3vhGBTeqPL07NwOoiK0/iuBojMCdCKOdZfqgsBpEPP7qxqM3GNgPjAY0ah8x awIDAQAB -----END PUBLIC KEY----- --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: jwt-route spec: ingressClassName: apisix http: - name: jwt-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: jwt-auth enable: true config: _meta: disable: false ``` Apply the configuration to your cluster: ``` kubectl apply -f jwt-auth-ic.yaml ``` To verify, send a request to the route with the JWT in the `Authorization` header: ``` curl -i "http://127.0.0.1:9080/headers" -H "Authorization: ${jwt_token}" ``` You should receive an `HTTP/1.1 200 OK` response. ### Add Consumer Custom ID to Header[​](#add-consumer-custom-id-to-header "Direct link to Add Consumer Custom ID to Header") The following example demonstrates how you can attach a consumer custom ID to authenticated request in the `Consumer-Custom-Id` header, which can be used to implement additional logics as needed. * Admin API * ADC * Ingress Controller Create a consumer `jack` with a custom ID label: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jack", "labels": { "custom_id": "495aec6a" } }' ``` Create `jwt-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jack/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jack-jwt-auth", "plugins": { "jwt-auth": { "key": "jack-key", "secret": "jack-hs256-secret-that-is-very-long" } } }' ``` Create a route with `jwt-auth`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "jwt-auth-route", "uri": "/anything", "plugins": { "jwt-auth": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `jwt-auth` credential and a route with `jwt-auth` plugin enabled as such: adc.yaml ``` consumers: - username: jack labels: custom_id: "495aec6a" credentials: - name: jwt-auth type: jwt-auth config: key: jack-key secret: jack-hs256-secret-that-is-very-long services: - name: jwt-auth-service routes: - name: jwt-auth-route uris: - /anything plugins: jwt-auth: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create a consumer with `jwt-auth` credential and a route with `jwt-auth` plugin enabled: * Gateway API * APISIX CRD jwt-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jack labels: custom_id: "495aec6a" spec: gatewayRef: name: apisix credentials: - type: jwt-auth name: primary-cred config: key: jack-key secret: jack-hs256-secret-that-is-very-long --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: jwt-auth-plugin-config spec: plugins: - name: jwt-auth config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: jwt-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: jwt-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f jwt-auth-ic.yaml ``` jwt-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jack labels: custom_id: "495aec6a" spec: ingressClassName: apisix authParameter: jwtAuth: value: key: jack-key secret: jack-hs256-secret-that-is-very-long --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: jwt-auth-route spec: ingressClassName: apisix http: - name: jwt-auth-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: jwt-auth enable: true config: _meta: disable: false ``` Apply the configuration to your cluster: ``` kubectl apply -f jwt-auth-ic.yaml ``` To issue a JWT for `jack`, you could use [JWT.io's JWT encoder](https://jwt.io) or other utilities. If you are using [JWT.io's JWT encoder](https://jwt.io), do the following: * Fill in `HS256` as the algorithm. * Update the secret in the **Valid secret** section to be `jack-hs256-secret-that-is-very-long`. * Update payload with consumer key `jack-key`; and add `exp` or `nbf` in UNIX timestamp. note When `claims_to_verify` is a nonempty list, every listed claim is required and validated. When it is unset or empty, `exp` and `nbf` are validated when present, but neither claim is required. Your payload should look similar to the following: ``` { "key": "jack-key", "nbf": 1729132271 } ``` Copy the generated JWT and save to a variable: ``` export jwt_token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqYWNrLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.UEPXy5jpid624T1XpfjM0PLY73LZPjV3Qt8yZ92kVuU ``` To verify, send a request to the route with the JWT in the `Authorization` header: ``` curl -i "http://127.0.0.1:9080/anything" -H "Authorization: ${jwt_token}" ``` You should see an `HTTP/1.1 200 OK` response similar to the following: ``` { "headers": { "Accept": "*/*", "Authorization": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqYWNrLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.UEPXy5jpid624T1XpfjM0PLY73LZPjV3Qt8yZ92kVuU", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-6873b19d-329331db76e5e7194c942b47", "X-Consumer-Custom-Id": "495aec6a", "X-Consumer-Username": "aic_jack", "X-Forwarded-Host": "127.0.0.1" }, "url": "http://127.0.0.1/anything" } ``` If you would like to attach more consumer custom headers to authenticated requests, see the [`attach-consumer-label`](https://docs.api7.ai/hub/attach-consumer-label.md) plugin. ### Rate Limit with Anonymous Consumer[​](#rate-limit-with-anonymous-consumer "Direct link to Rate Limit with Anonymous Consumer") The following example demonstrates how you can configure different rate limiting policies by regular and anonymous consumers, where the anonymous consumer does not need to authenticate and has less quota. * Admin API * ADC * Ingress Controller Create a regular consumer `jack` and configure the `limit-count` plugin to allow for a quota of 3 within a 30-second window: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jack", "plugins": { "limit-count": { "count": 3, "time_window": 30, "rejected_code": 429, "policy": "local" } } }' ``` Create the `jwt-auth` credential for the consumer `jack`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jack/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jack-jwt-auth", "plugins": { "jwt-auth": { "key": "jack-key", "secret": "jack-hs256-secret-that-is-very-long" } } }' ``` Create an anonymous user `anonymous` and configure the `limit-count` plugin to allow for a quota of 1 within a 30-second window: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "anonymous", "plugins": { "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "policy": "local" } } }' ``` Create a route and configure the `jwt-auth` plugin to accept anonymous consumer `anonymous` from bypassing the authentication: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "jwt-auth-route", "uri": "/anything", "plugins": { "jwt-auth": { "anonymous_consumer": "anonymous" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Configure consumers with different rate limits and a route that accepts anonymous users: adc.yaml ``` consumers: - username: jack plugins: limit-count: count: 3 time_window: 30 rejected_code: 429 policy: local credentials: - name: jwt-auth type: jwt-auth config: key: jack-key secret: jack-hs256-secret-that-is-very-long - username: anonymous plugins: limit-count: count: 1 time_window: 30 rejected_code: 429 policy: local services: - name: anonymous-rate-limit-service routes: - name: jwt-auth-route uris: - /anything plugins: jwt-auth: anonymous_consumer: anonymous upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Configure consumers with different rate limits and a route that accepts anonymous users: jwt-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jack spec: gatewayRef: name: apisix credentials: - type: jwt-auth name: primary-key config: key: jack-key secret: jack-hs256-secret-that-is-very-long plugins: - name: limit-count config: count: 3 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: anonymous spec: gatewayRef: name: apisix plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 policy: local --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: jwt-auth-plugin-config spec: plugins: - name: jwt-auth config: anonymous_consumer: aic_anonymous # namespace_consumername --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: jwt-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: jwt-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f jwt-auth-ic.yaml ``` Configure consumers with different rate limits and a route that accepts anonymous users: jwt-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jack spec: ingressClassName: apisix authParameter: jwtAuth: value: key: jack-key secret: jack-hs256-secret-that-is-very-long plugins: - name: limit-count enable: true config: count: 3 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: anonymous spec: ingressClassName: apisix plugins: - name: limit-count enable: true config: count: 1 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: jwt-auth-route spec: ingressClassName: apisix http: - name: jwt-auth-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: jwt-auth enable: true config: anonymous_consumer: aic_anonymous ``` Apply the configuration to your cluster: ``` kubectl apply -f jwt-auth-ic.yaml ``` To issue a JWT for `jack`, you could use [JWT.io's JWT encoder](https://jwt.io) or other utilities. If you are using [JWT.io's JWT encoder](https://jwt.io), do the following: * Fill in `HS256` as the algorithm. * Update the secret in the **Valid secret** section to be `jack-hs256-secret-that-is-very-long`. * Update payload with consumer key `jack-key`; and add `exp` or `nbf` in UNIX timestamp. note When `claims_to_verify` is a nonempty list, every listed claim is required and validated. When it is unset or empty, `exp` and `nbf` are validated when present, but neither claim is required. Your payload should look similar to the following: ``` { "key": "jack-key", "nbf": 1729132271 } ``` Copy the generated JWT and save to a variable: ``` export jwt_token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJrZXkiOiJqYWNrLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.UEPXy5jpid624T1XpfjM0PLY73LZPjV3Qt8yZ92kVuU ``` To verify the rate limiting, send five consecutive requests with `jack`'s JWT: ``` resp=$(seq 5 | xargs -I{} curl "http://127.0.0.1:9080/anything" -H "Authorization: ${jwt_token}" -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that out of the 5 requests, 3 requests were successful (status code 200) while the others were rejected (status code 429). ``` 200: 3, 429: 2 ``` Send five anonymous requests: ``` resp=$(seq 5 | xargs -I{} curl "http://127.0.0.1:9080/anything" -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that only one request was successful: ``` 200: 1, 429: 4 ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. ### Credentials[​](#credentials "Direct link to Credentials") The following are plugin attributes available for configurations on [credentials](https://docs.api7.ai/apisix/key-concepts/credentials.md). * key string required vaild vaule: non-empty *** A unique key that identifies the credential for a consumer. * secret string vaild vaule: non-empty *** Shared key used to sign and verify the JWT when the algorithm is symmetric. Required when using `HS256`, `HS384`, or `HS512` as the algorithm. The secret is encrypted with AES before being stored in etcd. You can also store it in an environment variable and reference it using the `env://` prefix, or in a secret manager such as HashiCorp Vault's [KV secrets engine](https://developer.hashicorp.com/vault/docs/secrets/kv), and reference it using the `secret://` prefix. For more information, see [secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md). * public\_key string *** RSA or ECDSA public key. Required if the `algorithm` is `RS256`, `ES256`, `RS384`, `RS512`, `ES256`, `ES384`, `ES512`, `PS256`, `PS384`, `PS512`, or `EdDSA`. * algorithm string default: `HS256` vaild vaule: `HS256`, `HS384`, `HS512`, `RS256`, `RS384`, `RS512`, `ES256`, `ES384`, `ES512`, `PS256`, `PS384`, `PS512`, `EdDSA` *** Algorithm used to sign and verify the token. The `alg` value in the JWT header must exactly match this configured value; a mismatch is rejected with `401 Unauthorized`. * exp integer default: `86400` vaild vaule: greater than or equal to 1 *** Expiry time of the token in seconds. If you are not using APISIX to sign the JWT, this parameter is ignored and you should specify the expiration in the payload when signing the JWT. * base64\_secret boolean default: `false` *** Set to true if the secret is base64 encoded. * lifetime\_grace\_period integer default: `0` vaild vaule: greater than or equal to 0 *** Grace period in seconds. Used to account for clock skew between the server generating the JWT and the server validating the JWT. ### Routes or Services[​](#routes-or-services "Direct link to Routes or Services") The following are plugin attributes available for configurations on [routes](https://docs.api7.ai/apisix/key-concepts/routes.md) or [services](https://docs.api7.ai/apisix/key-concepts/services.md). * header string default: `authorization` *** The header to get the token from. * query string default: `jwt` *** The query string to get the token from. Lower priority than header. * cookie string default: `jwt` *** The cookie to get the token from. Lower priority than query. * hide\_credentials boolean default: `false` *** If true, do not pass the header, query, or cookie with JWT to upstream services. * anonymous\_consumer string *** Anonymous consumer name. If configured, allow anonymous users to bypass the authentication. See [Rate Limit with Anonymous Consumer](https://docs.api7.ai/hub/jwt-auth.md#rate-limit-with-anonymous-consumer) for more details. * claims\_to\_verify array\[string] vaild vaule: combination of `exp` and `nbf` *** Claims used to verify that the token is within its allowed time window. A nonempty list makes every listed claim required. A token missing a configured claim is rejected. When this option is unset or empty, `exp` and `nbf` are validated whenever they are present, but neither claim is required. These validation rules were introduced in API7 Enterprise 3.9.14 and 3.10.1, and in APISIX 3.17.0. * key\_claim\_name string default: `key` *** The claim in the JWT payload that identifies the associated secret, such as `iss`. * store\_in\_ctx boolean default: `false` *** If true, store JWT payload in the request context variable `ctx.jwt_auth_payload`. This allows plugins executed after `jwt-auth` on the same request to retrieve and use the payload information. For instance, to retrieve the key in the payload, you can use `ctx.jwt_auth_payload.key`. Supported in APISIX and from Enterprise 3.8.9. * realm string default: `jwt` *** Realm in the [`WWW-Authenticate`](https://datatracker.ietf.org/doc/html/rfc7235#section-4.1) response header returned with a `401 Unauthorized` response due to authentication failure. For example: * If `realm` is set to `jwt-auth`, the 401 response will include the following header: ``` WWW-Authenticate: Bearer realm="jwt-auth" ``` * If `realm` is not configured, the 401 response will include the following header: ``` WWW-Authenticate: Bearer realm="jwt" ``` This parameter is available in API7 Enterprise version 3.9.2 and later, and in Apache APISIX version 3.15.0 and later. --- # kafka-logger The `kafka-logger` plugin pushes request and response logs as JSON objects to Apache Kafka clusters in batches and supports the customization of log formats. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `kafka-logger` plugin for different scenarios. To follow along the examples, start a sample Kafka cluster. The Docker example attaches the broker to the `apisix-quickstart-net` network created by the [APISIX Docker quickstart](https://docs.api7.ai/apisix/getting-started/.md), so the gateway can resolve `notkafka:29092`. * Docker * Kubernetes docker-compose.yml ``` services: zookeeper: image: confluentinc/cp-zookeeper:7.8.0 container_name: zookeeper environment: ZOOKEEPER_CLIENT_PORT: 2181 ZOOKEEPER_TICK_TIME: 2000 networks: - apisix-quickstart-net notkafka: image: confluentinc/cp-kafka:7.8.0 container_name: notkafka depends_on: - zookeeper environment: KAFKA_BROKER_ID: 1 KAFKA_ZOOKEEPER_CONNECT: zookeeper:2181 KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: PLAINTEXT:PLAINTEXT,PLAINTEXT_HOST:PLAINTEXT KAFKA_ADVERTISED_LISTENERS: PLAINTEXT://notkafka:29092,PLAINTEXT_HOST://127.0.0.1:9092 KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR: 1 KAFKA_AUTO_CREATE_TOPICS_ENABLE: "true" ports: - "9092:9092" networks: - apisix-quickstart-net networks: apisix-quickstart-net: external: true ``` Start containers: ``` docker compose up -d ``` Create a Kubernetes manifest file for the Zookeeper and Kafka deployments: kafka-deployment.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: zookeeper spec: replicas: 1 selector: matchLabels: app: zookeeper template: metadata: labels: app: zookeeper spec: containers: - name: zookeeper image: confluentinc/cp-zookeeper:7.8.0 env: - name: ZOOKEEPER_CLIENT_PORT value: "2181" - name: ZOOKEEPER_TICK_TIME value: "2000" ports: - containerPort: 2181 --- apiVersion: v1 kind: Service metadata: namespace: aic name: zookeeper spec: selector: app: zookeeper ports: - port: 2181 targetPort: 2181 type: ClusterIP --- apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: kafka-server spec: replicas: 1 selector: matchLabels: app: kafka-server template: metadata: labels: app: kafka-server spec: containers: - name: kafka-server image: confluentinc/cp-kafka:7.8.0 env: - name: KAFKA_BROKER_ID value: "1" - name: KAFKA_ZOOKEEPER_CONNECT value: "zookeeper:2181" - name: KAFKA_LISTENER_SECURITY_PROTOCOL_MAP value: "PLAINTEXT:PLAINTEXT" - name: KAFKA_ADVERTISED_LISTENERS value: "PLAINTEXT://kafka-server.aic.svc:9092" - name: KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR value: "1" - name: KAFKA_AUTO_CREATE_TOPICS_ENABLE value: "true" ports: - containerPort: 9092 --- apiVersion: v1 kind: Service metadata: namespace: aic name: kafka-server spec: selector: app: kafka-server ports: - port: 9092 targetPort: 9092 type: ClusterIP ``` Apply the manifests: ``` kubectl apply -f kafka-deployment.yaml ``` Wait for messages in the configured Kafka topic: * Docker * Kubernetes ``` docker exec -it notkafka kafka-console-consumer --bootstrap-server localhost:9092 --topic test2 --from-beginning ``` ``` kubectl exec -n aic deploy/kafka-server -- kafka-console-consumer --bootstrap-server kafka-server.aic.svc:9092 --topic test2 --from-beginning ``` Open a new terminal session for the following steps working with APISIX. ### Log in Different Meta Log Formats[​](#log-in-different-meta-log-formats "Direct link to Log in Different Meta Log Formats") The following example demonstrates how you can enable the `kafka-logger` plugin on a route, which logs client requests to the route and pushes logs to Kafka. You will also understand the differences between the `default` and `origin` meta log formats. Create a route with `kafka-logger` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "kafka-logger-route", "uri": "/get", "plugins": { "kafka-logger": { "meta_format": "default", "brokers": [ { "host": "notkafka", "port": 29092 } ], "kafka_topic": "test2", "key": "key1", "batch_max_size": 1 } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: kafka-logger-route plugins: kafka-logger: meta_format: "default" brokers: - host: "notkafka" port: 29092 kafka_topic: "test2" key: "key1" batch_max_size: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD kafka-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: kafka-logger-plugin-config spec: plugins: - name: kafka-logger config: meta_format: "default" brokers: - host: "kafka-server.aic.svc" port: 9092 kafka_topic: "test2" key: "key1" batch_max_size: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: kafka-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: kafka-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` kafka-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: kafka-logger-route spec: ingressClassName: apisix http: - name: kafka-logger-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: kafka-logger config: meta_format: "default" brokers: - host: "kafka-server.aic.svc" port: 9092 kafka_topic: "test2" key: "key1" batch_max_size: 1 ``` Apply the configuration: ``` kubectl apply -f kafka-logger-ic.yaml ``` ❶ `meta_format`: set to the `default` log format. ❷ `batch_max_size`: set to 1 to send the log entry immediately. Send a request to the route to generate a log entry: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response. You should see a log entry in the Kafka topic similar to the following: ``` { "latency": 411.00001335144, "request": { "querystring": {}, "headers": { "host": "127.0.0.1:9080", "user-agent": "curl/8.7.1", "accept": "*/*", "x-forwarded-proto": "http", "x-forwarded-host": "127.0.0.1", "x-forwarded-port": "9080" }, "method": "GET", "size": 83, "uri": "/get", "url": "http://127.0.0.1:9080/get" }, "response": { "headers": { "content-length": "233", "access-control-allow-credentials": "true", "content-type": "application/json", "connection": "close", "access-control-allow-origin": "*", "date": "Fri, 10 Nov 2023 06:02:44 GMT", "server": "APISIX/3.16.0" }, "status": 200, "size": 475 }, "route_id": "kafka-logger-route", "client_ip": "127.0.0.1", "server": { "hostname": "apisix", "version": "3.16.0" }, "apisix_latency": 18.00001335144, "service_id": "", "upstream_latency": 393, "start_time": 1699596164550, "upstream": "54.90.18.68:80" } ``` Update the meta log format to `origin`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes/kafka-logger-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "kafka-logger": { "meta_format": "origin" } } }' ``` Update `adc.yaml` to set `meta_format` to `origin`: adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: kafka-logger-route plugins: kafka-logger: meta_format: "origin" brokers: - host: "notkafka" port: 29092 kafka_topic: "test2" key: "key1" batch_max_size: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update `kafka-logger-ic.yaml` to set `meta_format` to `origin`: kafka-logger-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: kafka-logger-plugin-config spec: plugins: - name: kafka-logger config: meta_format: "origin" brokers: - host: "kafka-server.aic.svc" port: 9092 kafka_topic: "test2" key: "key1" batch_max_size: 1 ``` Update `kafka-logger-ic.yaml` to set `meta_format` to `origin`: kafka-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: kafka-logger-route spec: ingressClassName: apisix http: - name: kafka-logger-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: kafka-logger config: meta_format: "origin" brokers: - host: "kafka-server.aic.svc" port: 9092 kafka_topic: "test2" key: "key1" batch_max_size: 1 ``` Apply the updated configuration: ``` kubectl apply -f kafka-logger-ic.yaml ``` Send a request to the route again to generate a new log entry: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response. You should see a log entry in the Kafka topic similar to the following: ``` GET /get HTTP/1.1 x-forwarded-proto: http x-forwarded-host: 127.0.0.1 user-agent: curl/8.7.1 x-forwarded-port: 9080 host: 127.0.0.1:9080 accept: */* ``` ### Send Logs to a TLS-Enabled Broker[​](#send-logs-to-a-tls-enabled-broker "Direct link to Send Logs to a TLS-Enabled Broker") The following Docker example starts a local Kafka broker with a CA-signed TLS certificate on the `apisix-quickstart-net` network. It then connects with certificate verification and uses Produce API version `2` so Kafka records the message timestamp. Install OpenSSL and the Java `keytool` command before continuing. Generate a sample CA, a broker certificate for the `kafka-tls` container hostname, and the Java key store and trust store files required by Kafka: ``` mkdir -p kafka-tls-certs openssl req -x509 -newkey rsa:2048 -nodes -days 365 \ -subj "/CN=kafka-example-ca" \ -keyout kafka-tls-certs/ca.key \ -out kafka-tls-certs/ca.crt openssl req -newkey rsa:2048 -nodes \ -subj "/CN=kafka-tls" \ -keyout kafka-tls-certs/server.key \ -out kafka-tls-certs/server.csr printf "subjectAltName=DNS:kafka-tls\n" > kafka-tls-certs/server-ext.cnf openssl x509 -req -days 365 \ -in kafka-tls-certs/server.csr \ -CA kafka-tls-certs/ca.crt \ -CAkey kafka-tls-certs/ca.key \ -CAcreateserial \ -extfile kafka-tls-certs/server-ext.cnf \ -out kafka-tls-certs/server.crt openssl pkcs12 -export \ -name kafka-tls \ -in kafka-tls-certs/server.crt \ -inkey kafka-tls-certs/server.key \ -certfile kafka-tls-certs/ca.crt \ -out kafka-tls-certs/kafka.keystore.p12 \ -passout pass:changeit keytool -importkeystore -noprompt \ -srckeystore kafka-tls-certs/kafka.keystore.p12 \ -srcstoretype PKCS12 \ -srcstorepass changeit \ -destkeystore kafka-tls-certs/kafka.keystore.jks \ -deststoretype JKS \ -deststorepass changeit \ -destkeypass changeit keytool -importcert -noprompt \ -alias kafka-example-ca \ -file kafka-tls-certs/ca.crt \ -keystore kafka-tls-certs/kafka.truststore.jks \ -storepass changeit printf "changeit\n" > kafka-tls-certs/kafka_keystore_creds printf "changeit\n" > kafka-tls-certs/kafka_ssl_key_creds ``` Create the client configuration used later to verify the record: kafka-tls-certs/client.properties ``` security.protocol=SSL ssl.truststore.location=/etc/kafka/secrets/kafka.truststore.jks ssl.truststore.password=changeit ssl.endpoint.identification.algorithm=https ``` Start the TLS-enabled broker on the APISIX quickstart network: ``` docker run -d \ --name kafka-tls \ --hostname kafka-tls \ --network apisix-quickstart-net \ -v "${PWD}/kafka-tls-certs:/etc/kafka/secrets:ro" \ -e KAFKA_NODE_ID=1 \ -e KAFKA_PROCESS_ROLES=broker,controller \ -e KAFKA_LISTENER_SECURITY_PROTOCOL_MAP="SSL:SSL,CONTROLLER:PLAINTEXT" \ -e KAFKA_ADVERTISED_LISTENERS="SSL://kafka-tls:9093" \ -e KAFKA_LISTENERS="SSL://:9093,CONTROLLER://:29093" \ -e KAFKA_CONTROLLER_QUORUM_VOTERS="1@kafka-tls:29093" \ -e KAFKA_CONTROLLER_LISTENER_NAMES=CONTROLLER \ -e KAFKA_INTER_BROKER_LISTENER_NAME=SSL \ -e KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR=1 \ -e KAFKA_GROUP_INITIAL_REBALANCE_DELAY_MS=0 \ -e KAFKA_TRANSACTION_STATE_LOG_MIN_ISR=1 \ -e KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR=1 \ -e KAFKA_SSL_KEYSTORE_FILENAME=kafka.keystore.jks \ -e KAFKA_SSL_KEYSTORE_CREDENTIALS=kafka_keystore_creds \ -e KAFKA_SSL_KEY_CREDENTIALS=kafka_ssl_key_creds \ -e KAFKA_SSL_TRUSTSTORE_LOCATION=/etc/kafka/secrets/kafka.truststore.jks \ -e KAFKA_SSL_TRUSTSTORE_PASSWORD=changeit \ -e KAFKA_SSL_CLIENT_AUTH=none \ -e CLUSTER_ID="4L6g3nShT-eMCtK--X86sw" \ apache/kafka:4.1.0 ``` After the broker log contains `Transition from STARTING to STARTED`, create the topic: ``` docker exec kafka-tls /opt/kafka/bin/kafka-topics.sh \ --bootstrap-server kafka-tls:9093 \ --command-config /etc/kafka/secrets/client.properties \ --create \ --if-not-exists \ --topic apisix-logs \ --partitions 1 \ --replication-factor 1 ``` Set the broker address and topic for the route configuration: ``` export KAFKA_TLS_HOST="kafka-tls" export KAFKA_TLS_PORT="9093" export KAFKA_TOPIC="apisix-logs" ``` Copy the generated CA certificate into the quickstart gateway. Append it to the existing system trust bundle so that the gateway continues to trust the public CA certificates already installed in the container: ``` docker cp kafka-tls-certs/ca.crt \ apisix-quickstart:/usr/local/apisix/conf/kafka-example-ca.crt docker exec apisix-quickstart sh -c ' cat /etc/ssl/certs/ca-certificates.crt \ /usr/local/apisix/conf/kafka-example-ca.crt \ > /usr/local/apisix/conf/combined-ca-bundle.pem ' ``` Update the trust bundle path in the quickstart configuration and reload the gateway: ``` docker exec apisix-quickstart sh -c ' config=/usr/local/apisix/conf/config.yaml certificate=/usr/local/apisix/conf/combined-ca-bundle.pem if grep -q "^ ssl_trusted_certificate:" "$config"; then sed "s#ssl_trusted_certificate:.*#ssl_trusted_certificate: $certificate#" "$config" elif grep -q "^ ssl:$" "$config"; then sed "/^ ssl:$/a\\ ssl_trusted_certificate: $certificate" "$config" else sed "/^apisix:$/a\\ ssl:\\ ssl_trusted_certificate: $certificate" "$config" fi > /tmp/config.yaml cat /tmp/config.yaml > /usr/local/apisix/conf/config.yaml apisix reload ' ``` For a multi-instance deployment, distribute the combined trust bundle and configuration change to every gateway instance. The `kafka-tls` hostname matches the DNS name in the sample broker certificate. Create a route that sends each log entry immediately to the TLS listener: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- <`; with `basic`, they send an ordinary `Authorization: Basic ` header, which lets existing HTTP Basic clients authenticate unchanged. The plugin returns `401 Unauthorized` for missing, malformed, or rejected user credentials, ambiguous user matches, and a missing consumer when `consumer_required` is enabled. LDAP transport, TLS, protocol, server, search-bind, and other directory failures return `500 Internal Server Error` so an outage is not presented as an invalid user credential. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can work with the `ldap-auth-advanced` plugin for different scenarios. ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") The examples assume an LDAP directory reachable at `192.168.1.10:389` that holds user entries under `ou=users,dc=example,dc=org`, with a user whose `uid` is `johndoe` and whose password is `john-secret`. Searches are performed as `cn=admin,dc=example,dc=org`. Adjust these values for your own directory. ### Authenticate Against an LDAP Directory[​](#authenticate-against-an-ldap-directory "Direct link to Authenticate Against an LDAP Directory") The following example demonstrates how to authenticate clients against an LDAP directory without mapping them onto consumers, by setting `consumer_required` to `false`. * Admin API * ADC * Ingress Controller Create a route with `ldap-auth-advanced`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "ldap-auth-route", "uri": "/anything", "plugins": { "ldap-auth-advanced": { "ldap_uri": "192.168.1.10:389", "base_dn": "ou=users,dc=example,dc=org", "attribute": "uid", "bind_dn": "cn=admin,dc=example,dc=org", "ldap_password": "admin-secret", "consumer_required": false } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a route with the `ldap-auth-advanced` plugin configured: adc.yaml ``` services: - name: ldap-auth-service routes: - name: ldap-auth-route uris: - /anything plugins: ldap-auth-advanced: ldap_uri: 192.168.1.10:389 base_dn: ou=users,dc=example,dc=org attribute: uid bind_dn: cn=admin,dc=example,dc=org ldap_password: admin-secret consumer_required: false upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create a route with the `ldap-auth-advanced` plugin configured: * Gateway API * APISIX CRD ldap-auth-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: ldap-auth-plugin-config spec: plugins: - name: ldap-auth-advanced config: ldap_uri: 192.168.1.10:389 base_dn: ou=users,dc=example,dc=org attribute: uid bind_dn: cn=admin,dc=example,dc=org ldap_password: admin-secret consumer_required: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: ldap-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: ldap-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f ldap-auth-ic.yaml ``` ldap-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: ldap-auth-route spec: ingressClassName: apisix http: - name: ldap-auth-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: ldap-auth-advanced enable: true config: ldap_uri: 192.168.1.10:389 base_dn: ou=users,dc=example,dc=org attribute: uid bind_dn: cn=admin,dc=example,dc=org ldap_password: admin-secret consumer_required: false ``` Apply the configuration to your cluster: ``` kubectl apply -f ldap-auth-ic.yaml ``` #### Verify with Valid Credentials[​](#verify-with-valid-credentials "Direct link to Verify with Valid Credentials") Send a request to the route with the directory user's credentials, base64 encoded and presented in the `ldap` scheme: ``` curl -i "http://127.0.0.1:9080/anything" \ -H "Authorization: ldap $(printf '%s' 'johndoe:john-secret' | base64)" ``` You should receive an `HTTP/1.1 200 OK` response. #### Verify with Invalid Credentials[​](#verify-with-invalid-credentials "Direct link to Verify with Invalid Credentials") Send a request to the route with an incorrect password: ``` curl -i "http://127.0.0.1:9080/anything" \ -H "Authorization: ldap $(printf '%s' 'johndoe:wrong-password' | base64)" ``` You should receive an `HTTP/1.1 401 Unauthorized` response: ``` WWW-Authenticate: ldap realm="ldap" ``` ``` {"message":"Authorization required"} ``` #### Verify without Credentials[​](#verify-without-credentials "Direct link to Verify without Credentials") Send a request to the route without any credentials: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 401 Unauthorized` response. ### Map LDAP Users to Consumers[​](#map-ldap-users-to-consumers "Direct link to Map LDAP Users to Consumers") The following example demonstrates how to map a directory user onto a consumer, so that consumer-scoped configurations apply to their traffic. The consumer's credential records the user's full DN, which is the value the plugin resolves through its directory search. info Consumer credentials of type `ldap-auth-advanced` are created through the Admin API or the Dashboard. ADC and the Ingress Controller currently support only `key-auth`, `basic-auth`, `jwt-auth`, and `hmac-auth` credentials. Create a consumer `johndoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "johndoe" }' ``` Create an `ldap-auth-advanced` credential for the consumer, recording the DN of the directory entry: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/johndoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-ldap-auth", "plugins": { "ldap-auth-advanced": { "user_dn": "uid=johndoe,ou=users,dc=example,dc=org" } } }' ``` Create a route with `ldap-auth-advanced`, leaving `consumer_required` at its default of `true`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "ldap-auth-consumer-route", "uri": "/anything", "plugins": { "ldap-auth-advanced": { "ldap_uri": "192.168.1.10:389", "base_dn": "ou=users,dc=example,dc=org", "attribute": "uid", "bind_dn": "cn=admin,dc=example,dc=org", "ldap_password": "admin-secret" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Send a request to the route with the directory user's credentials: ``` curl -i "http://127.0.0.1:9080/anything" \ -H "Authorization: ldap $(printf '%s' 'johndoe:john-secret' | base64)" ``` You should receive an `HTTP/1.1 200 OK` response, and the upstream should see the consumer headers: ``` { "headers": { "X-Consumer-Username": "johndoe", "X-Credential-Identifier": "cred-john-ldap-auth", ... }, ... } ``` A directory user who authenticates successfully but whose DN is not recorded on any consumer credential receives `HTTP/1.1 401 Unauthorized`. ### Accept the HTTP Basic Authentication Scheme[​](#accept-the-http-basic-authentication-scheme "Direct link to Accept the HTTP Basic Authentication Scheme") The following example demonstrates how to accept credentials in the standard HTTP Basic scheme instead of the `ldap` scheme, so that existing Basic clients work unchanged. Set `header_type` to `basic` on the route: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "ldap-auth-basic-route", "uri": "/anything", "plugins": { "ldap-auth-advanced": { "ldap_uri": "192.168.1.10:389", "base_dn": "ou=users,dc=example,dc=org", "attribute": "uid", "bind_dn": "cn=admin,dc=example,dc=org", "ldap_password": "admin-secret", "header_type": "basic", "consumer_required": false } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Send a request using an ordinary Basic credential: ``` curl -i "http://127.0.0.1:9080/anything" -u johndoe:john-secret ``` You should receive an `HTTP/1.1 200 OK` response. An unauthenticated request is now challenged with the Basic scheme: ``` WWW-Authenticate: Basic realm="ldap" ``` ### Connect to the Directory Over TLS[​](#connect-to-the-directory-over-tls "Direct link to Connect to the Directory Over TLS") The following example demonstrates how to reach the directory over LDAPS. Set `use_ldaps` and point `ldap_uri` at the LDAPS port; when the port is omitted, `636` is used under LDAPS and `389` otherwise. Use `use_starttls` instead to upgrade a plaintext connection on port `389`. The two options are mutually exclusive. ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "ldaps-auth-route", "uri": "/anything", "plugins": { "ldap-auth-advanced": { "ldap_uri": "192.168.1.10:636", "use_ldaps": true, "base_dn": "ou=users,dc=example,dc=org", "attribute": "uid", "bind_dn": "cn=admin,dc=example,dc=org", "ldap_password": "admin-secret", "consumer_required": false } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Certificate verification is controlled by `ssl_verify`, which is enabled by default. Leave it enabled and make sure the directory's issuing CA is trusted by the gateway. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. This plugin supports referencing sensitive parameter values from environment variables using the `env://` prefix, or from a secret manager, such as HashiCorp Vault’s [KV secrets engine](https://developer.hashicorp.com/vault/docs/secrets/kv), using the `secret://` prefix. For more information, see [environment variables in plugin](https://docs.api7.ai/apisix/reference/environment-variables.md#plugins) and [secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md). ### Credentials[​](#credentials "Direct link to Credentials") The following are plugin attributes available for configurations on [credentials](https://docs.api7.ai/apisix/key-concepts/credentials.md). * user\_dn string required vaild vaule: between 1 and 4096 characters *** Distinguished name of the directory entry this consumer represents, such as `uid=johndoe,ou=users,dc=example,dc=org`. It has to match the DN the plugin resolves through its directory search. ### Routes or Services[​](#routes-or-services "Direct link to Routes or Services") The following are plugin attributes available for configurations on [routes](https://docs.api7.ai/apisix/key-concepts/routes.md) or [services](https://docs.api7.ai/apisix/key-concepts/services.md). * ldap\_uri string required vaild vaule: between 1 and 256 characters *** Address of the LDAP directory, as `host` or `host:port`. When the port is omitted, `636` is used if `use_ldaps` is enabled and `389` otherwise. * base\_dn string required vaild vaule: between 1 and 4096 characters *** Distinguished name of the subtree the plugin searches to resolve the user, such as `ou=users,dc=example,dc=org`. * attribute string default: `cn` vaild vaule: an RFC 4512 attribute description, at most 256 characters *** Attribute matched against the user name supplied by the client. The search filter is `(attribute=username)`, so `uid` suits most OpenLDAP directories and `sAMAccountName` suits Active Directory. * bind\_dn string vaild vaule: between 1 and 4096 characters *** Distinguished name the plugin binds as to perform the search. When unset, the search is performed anonymously. Setting it requires `ldap_password`. * ldap\_password string vaild vaule: between 1 and 4096 characters *** Password for `bind_dn`. Required when `bind_dn` is set. When Data Plane data encryption is enabled, this field is encrypted at rest. * use\_ldaps boolean default: `false` *** If true, connect to the directory over LDAPS. Mutually exclusive with `use_starttls`. * use\_starttls boolean default: `false` *** If true, upgrade a plaintext connection to TLS with StartTLS. Mutually exclusive with `use_ldaps`. * ssl\_verify boolean default: `true` *** If true, verify the directory's TLS certificate when connecting over LDAPS or StartTLS. * timeout integer default: `10000` vaild vaule: between 1 and 60000 inclusive *** Timeout in milliseconds for the connection to the directory. * size\_limit integer default: `2` vaild vaule: greater than or equal to 2 *** Maximum number of entries the directory returns for the search. The default of `2` is enough to detect an ambiguous user name, which the plugin rejects rather than binding as an arbitrary match. * time\_limit integer default: `5` vaild vaule: greater than or equal to 0 *** Time limit in seconds the directory applies to the search. Set to `0` to use the directory's own default. * consumer\_required boolean default: `true` *** If true, the authenticated user has to map onto a consumer whose credential records their distinguished name, and a user without such a credential is rejected. Set to `false` to authenticate against the directory without involving consumers. * header\_type string default: `ldap` vaild vaule: `ldap` or `basic` *** Authentication scheme the plugin accepts in the `Authorization` or `Proxy-Authorization` header, and the scheme it names in the `WWW-Authenticate` challenge. In both cases the credential itself is `base64(username:password)`, so `basic` produces an ordinary HTTP Basic exchange. * hide\_credentials boolean default: `false` *** If true, remove the header carrying the directory credentials once the client has been authenticated, so the username and password are not forwarded to the upstream. Available in API7 Enterprise from version 3.10.6. * realm string default: `ldap` *** Realm reported in the `WWW-Authenticate` header of the challenge returned to unauthenticated clients. * keepalive boolean default: `true` *** If true, keep connections to the directory alive so that they are reused across requests. * keepalive\_timeout integer default: `60000` vaild vaule: greater than or equal to 1000 *** Idle time in milliseconds after which a pooled connection to the directory is closed. * keepalive\_pool\_size integer default: `5` vaild vaule: greater than or equal to 1 *** Maximum number of pooled connections to the directory per worker. * keepalive\_pool\_name string vaild vaule: between 1 and 256 characters *** Name of the connection pool. Set it to keep the connections of different plugin configurations in separate pools. --- # limit-conn The `limit-conn` plugin limits the rate of requests by the number of concurrent connections. Requests exceeding the threshold will be delayed or rejected based on the configuration, ensuring controlled resource usage and preventing overload. ## Local vs Redis Rate Limiting[​](#local-vs-redis-rate-limiting "Direct link to Local vs Redis Rate Limiting") The `limit-conn` plugin supports two modes of rate limiting: * **Local rate limiting**: Limits are enforced independently on each gateway instance. Each instance maintains its own counters, so the effective limit is roughly (limit × number of instances) when traffic is spread across instances. This is the default when no `policy` is set or when `policy` is `local`. * **Redis-based rate limiting**: Limits are shared across all gateway instances through Redis. All instances share the same quota, so the configured limit applies to all gateway instances. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `limit-conn` in different scenarios. ### Apply Rate Limiting by Remote Address[​](#apply-rate-limiting-by-remote-address "Direct link to Apply Rate Limiting by Remote Address") The following example demonstrates how to use `limit-conn` to rate limit requests by `remote_addr`, with example connection and burst thresholds. Create a route with `limit-conn` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-conn-route", "uri": "/get", "plugins": { "limit-conn": { "conn": 2, "burst": 1, "default_conn_delay": 0.1, "key_type": "var", "key": "remote_addr", "policy": "local", "rejected_code": 429 } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-conn-route plugins: limit-conn: conn: 2 burst: 1 default_conn_delay: 0.1 key_type: var key: remote_addr policy: local rejected_code: 429 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-conn-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-conn-plugin-config spec: plugins: - name: limit-conn config: conn: 2 burst: 1 default_conn_delay: 0.1 key_type: var key: remote_addr policy: local rejected_code: 429 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-conn-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-conn-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-conn-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-conn-route spec: ingressClassName: apisix http: - name: limit-conn-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-conn config: conn: 2 burst: 1 default_conn_delay: 0.1 key_type: var key: remote_addr policy: local rejected_code: 429 ``` Apply the configuration: ``` kubectl apply -f limit-conn-ic.yaml ``` ❶ `conn`: allow 2 concurrent requests. ❷ `burst`: allow 1 excessive concurrent request. ❸ `default_conn_delay`: Allow 0.1 second of processing latency for concurrent requests between `conn` and `conn + burst`. ❹ `key_type`: set to `var` to interpret `key` as a variable. ❺ `key`: calculate rate limiting count by request's `remote_addr`. ❻ `policy`: use the local counter in memory. ❼ `rejected_code`: set the rejection status code to `429`. Send five concurrent requests to the route: ``` seq 1 5 | xargs -n1 -P5 bash -c 'curl -s -o /dev/null -w "Response: %{http_code}\n" "http://127.0.0.1:9080/get"' ``` You should see responses similar to the following, where excessive requests are rejected: ``` Response: 200 Response: 200 Response: 200 Response: 429 Response: 429 ``` ### Apply Rate Limiting by Remote Address and Consumer Name[​](#apply-rate-limiting-by-remote-address-and-consumer-name "Direct link to Apply Rate Limiting by Remote Address and Consumer Name") The following example demonstrates how to use `limit-conn` to rate limit requests by a combination of variables, `remote_addr` and `consumer_name`. * Admin API * ADC * Ingress Controller Create consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Create a second consumer `jane`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jane" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jane/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-key-auth", "plugins": { "key-auth": { "key": "jane-key" } } }' ``` Create a route with `key-auth` and `limit-conn` plugins: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-conn-route", "uri": "/get", "plugins": { "key-auth": {}, "limit-conn": { "conn": 2, "burst": 1, "default_conn_delay": 0.1, "rejected_code": 429, "policy": "local", "key_type": "var_combination", "key": "$remote_addr $consumer_name" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create two consumers and a route that enables rate limiting by consumers: adc.yaml ``` consumers: - username: john credentials: - name: key-auth type: key-auth config: key: john-key - username: jane credentials: - name: key-auth type: key-auth config: key: jane-key services: - name: limit-conn-service routes: - name: limit-conn-route uris: - /get plugins: key-auth: {} limit-conn: conn: 2 burst: 1 default_conn_delay: 0.1 rejected_code: 429 policy: local key_type: var_combination key: "$remote_addr $consumer_name" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create two consumers and a route that enables rate limiting by consumers: * Gateway API * APISIX CRD limit-conn-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: john-key --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jane spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: jane-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-conn-plugin-config spec: plugins: - name: key-auth config: _meta: disable: false - name: limit-conn config: conn: 2 burst: 1 default_conn_delay: 0.1 rejected_code: 429 policy: local key_type: var_combination key: "$remote_addr $consumer_name" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-conn-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-conn-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-conn-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jane spec: ingressClassName: apisix authParameter: keyAuth: value: key: jane-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-conn-route spec: ingressClassName: apisix http: - name: limit-conn-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: key-auth config: _meta: disable: false - name: limit-conn config: conn: 2 burst: 1 default_conn_delay: 0.1 rejected_code: 429 policy: local key_type: var_combination key: "$remote_addr $consumer_name" ``` Apply the configuration: ``` kubectl apply -f limit-conn-ic.yaml ``` ❶ `key-auth`: enable key authentication on the route. ❷ `key_type`: set to `var_combination` to interpret the `key` as a combination of variables. ❸ `key`: set to `$remote_addr $consumer_name` to apply rate limiting quota by remote address and consumer. Send five concurrent requests as the consumer `john`: ``` seq 1 5 | xargs -n1 -P5 bash -c 'curl -s -o /dev/null -w "Response: %{http_code}\n" "http://127.0.0.1:9080/get" -H "apikey: john-key"' ``` You should see responses similar to the following, where excessive requests are rejected: ``` Response: 200 Response: 200 Response: 200 Response: 429 Response: 429 ``` Immediately send five concurrent requests as the consumer `jane`: ``` seq 1 5 | xargs -n1 -P5 bash -c 'curl -s -o /dev/null -w "Response: %{http_code}\n" "http://127.0.0.1:9080/get" -H "apikey: jane-key"' ``` You should also see responses similar to the following, where excessive requests are rejected: ``` Response: 200 Response: 200 Response: 200 Response: 429 Response: 429 ``` In this case, the plugin rate limits by the combination of variables `remote_addr` and `consumer_name`, which means each consumer's quota is independent. ### Rate Limit WebSocket Connections[​](#rate-limit-websocket-connections "Direct link to Rate Limit WebSocket Connections") The following example demonstrates how you can use the `limit-conn` plugin to limit the number of concurrent WebSocket connections. Start a [sample upstream WebSocket server](https://hub.docker.com/r/jmalloc/echo-server): * Docker * Kubernetes ``` docker run -d \ -p 8080:8080 \ --name websocket-server \ --network=apisix-quickstart-net \ jmalloc/echo-server ``` Create a Kubernetes manifest file for the deployment of WebSocket server: ws-deployment.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: websocket-server spec: replicas: 1 selector: matchLabels: app: websocket-server template: metadata: labels: app: websocket-server spec: containers: - name: echo-server image: jmalloc/echo-server ports: - containerPort: 8080 ``` Create another Kubernetes manifest file for the WebSocket service: ws-service.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: websocket-server spec: selector: app: websocket-server ports: - protocol: TCP port: 8080 targetPort: 8080 appProtocol: kubernetes.io/ws type: ClusterIP ``` Gateway API and WebSocket For Gateway API, WebSocket support is enabled through the Service's `appProtocol` field (`kubernetes.io/ws` or `kubernetes.io/wss`). Unlike ApisixRoute, there is no direct `websocket` field or annotation support in HTTPRoute. Ensure that your Service is configured with `appProtocol` if you are working with Gateway API resources. See [Detect Upstream Protocol with appProtocol](https://docs.api7.ai/ingress-controller/detect-upstream-protocol-appprotocol.md) for more information. The server has a WebSocket endpoint at `/.ws` that echoes back any message received. Create a route to the server WebSocket endpoint and enable WebSocket for the route: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT -d ' { "id": "ws-route", "uri": "/.ws", "plugins": { "limit-conn": { "conn": 2, "burst": 1, "default_conn_delay": 0.1, "key_type": "var", "key": "remote_addr", "rejected_code": 429, "policy": "local" } }, "enable_websocket": true, "upstream": { "type": "roundrobin", "nodes": { "websocket-server:8080": 1 } } }' ``` ❶ Enable WebSocket for the route. ❷ Replace with your WebSocket server address. adc.yaml ``` services: - name: websocket-service routes: - name: ws-route uris: - /.ws enable_websocket: true plugins: limit-conn: conn: 2 burst: 1 default_conn_delay: 0.1 key_type: var key: remote_addr rejected_code: 429 policy: local upstream: type: roundrobin nodes: - host: websocket-server port: 8080 weight: 1 ``` ❶ Enable WebSocket for the route. ❷ Replace with your WebSocket server address. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-conn-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-conn-plugin-config spec: plugins: - name: limit-conn config: conn: 2 burst: 1 default_conn_delay: 0.1 key_type: var key: remote_addr rejected_code: 429 policy: local --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: ws-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /.ws filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-conn-plugin-config backendRefs: - name: websocket-server port: 8080 ``` limit-conn-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: ws-route spec: ingressClassName: apisix http: - name: ws-route match: paths: - /.ws methods: - GET websocket: true backends: - serviceName: websocket-server servicePort: 8080 plugins: - name: limit-conn config: conn: 2 burst: 1 default_conn_delay: 0.1 key_type: var key: remote_addr rejected_code: 429 policy: local ``` Apply the configuration: ``` kubectl apply -f limit-conn-ic.yaml ``` Install a WebSocket client, such as [websocat](https://github.com/vi/websocat), if you have not already. Establish connection with the WebSocket server through the route: ``` websocat "ws://127.0.0.1:9080/.ws" ``` Send a "hello" message in the terminal, you should see the WebSocket server echoes back the same message: ``` Request served by 1cd244052136 hello hello ``` Open three more terminal sessions and run: ``` websocat "ws://127.0.0.1:9080/.ws" ``` You should see the last terminal session prints `429 Too Many Requests` when you try to establish a WebSocket connection with the server, due to the rate limiting effect. ### Share Quota Among APISIX Nodes with a Redis Server[​](#share-quota-among-apisix-nodes-with-a-redis-server "Direct link to Share Quota Among APISIX Nodes with a Redis Server") The following example demonstrates the rate limiting of requests across multiple APISIX nodes with a Redis server, such that different APISIX nodes share the same rate limiting quota. On each APISIX instance, create a route with the following configurations. Adjust the configuration details accordingly. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-conn-route", "uri": "/get", "plugins": { "limit-conn": { "conn": 1, "burst": 1, "default_conn_delay": 0.1, "rejected_code": 429, "key_type": "var", "key": "remote_addr", "policy": "redis", "redis_host": "192.168.xxx.xxx", "redis_port": 6379, "redis_password": "p@ssw0rd", "redis_database": 1 } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-conn-route plugins: limit-conn: conn: 1 burst: 1 default_conn_delay: 0.1 rejected_code: 429 key_type: var key: remote_addr policy: redis redis_host: "192.168.xxx.xxx" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-conn-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-conn-plugin-config spec: plugins: - name: limit-conn config: conn: 1 burst: 1 default_conn_delay: 0.1 rejected_code: 429 key_type: var key: remote_addr policy: redis redis_host: "redis-service.aic.svc" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-conn-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-conn-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-conn-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-conn-route spec: ingressClassName: apisix http: - name: limit-conn-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-conn config: conn: 1 burst: 1 default_conn_delay: 0.1 rejected_code: 429 key_type: var key: remote_addr policy: redis redis_host: "redis-service.aic.svc" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 ``` Apply the configuration: ``` kubectl apply -f limit-conn-ic.yaml ``` ❶ `policy`: set to `redis` to use a Redis instance for rate limiting. ❷ `redis_host`: set to Redis instance IP address. ❸ `redis_port`: set to Redis instance listening port. ❹ `redis_password`: set to the password of the Redis instance, if any. ❺ `redis_database`: set to the database number in the Redis instance. Send five concurrent requests to the route: ``` seq 1 5 | xargs -n1 -P5 bash -c 'curl -s -o /dev/null -w "Response: %{http_code}\n" "http://127.0.0.1:9080/get"' ``` You should see responses similar to the following, where excessive requests are rejected: ``` Response: 200 Response: 200 Response: 429 Response: 429 Response: 429 ``` This shows the two routes configured in different APISIX instances share the same quota. ### Share Quota Among APISIX Nodes with a Redis Cluster[​](#share-quota-among-apisix-nodes-with-a-redis-cluster "Direct link to Share Quota Among APISIX Nodes with a Redis Cluster") You can also use a Redis cluster to apply the same quota across multiple APISIX nodes, such that different APISIX nodes share the same rate limiting quota. Ensure that your Redis instances are running in [cluster mode](https://redis.io/docs/management/scaling/#create-and-use-a-redis-cluster). Configure `redis_cluster_name` and one or more node addresses in `redis_cluster_nodes` for the `limit-conn` plugin. On each APISIX instance, create a route with the following configurations. Adjust the configuration details accordingly. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-conn-route", "uri": "/get", "plugins": { "limit-conn": { "conn": 1, "burst": 1, "default_conn_delay": 0.1, "rejected_code": 429, "key_type": "var", "key": "remote_addr", "policy": "redis-cluster", "redis_cluster_nodes": [ "192.168.xxx.xxx:6379", "192.168.xxx.xxx:16379" ], "redis_password": "p@ssw0rd", "redis_cluster_name": "redis-cluster", "redis_cluster_ssl": true } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-conn-route plugins: limit-conn: conn: 1 burst: 1 default_conn_delay: 0.1 rejected_code: 429 key_type: var key: remote_addr policy: redis-cluster redis_cluster_nodes: - "192.168.xxx.xxx:6379" - "192.168.xxx.xxx:16379" redis_password: "p@ssw0rd" redis_cluster_name: "redis-cluster" redis_cluster_ssl: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-conn-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-conn-plugin-config spec: plugins: - name: limit-conn config: conn: 1 burst: 1 default_conn_delay: 0.1 rejected_code: 429 key_type: var key: remote_addr policy: redis-cluster redis_cluster_nodes: - "redis-cluster-0.redis-cluster.aic.svc:6379" - "redis-cluster-1.redis-cluster.aic.svc:6379" redis_password: "p@ssw0rd" redis_cluster_name: "redis-cluster" redis_cluster_ssl: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-conn-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-conn-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-conn-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-conn-route spec: ingressClassName: apisix http: - name: limit-conn-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-conn config: conn: 1 burst: 1 default_conn_delay: 0.1 rejected_code: 429 key_type: var key: remote_addr policy: redis-cluster redis_cluster_nodes: - "redis-cluster-0.redis-cluster.aic.svc:6379" - "redis-cluster-1.redis-cluster.aic.svc:6379" redis_password: "p@ssw0rd" redis_cluster_name: "redis-cluster" redis_cluster_ssl: true ``` Apply the configuration: ``` kubectl apply -f limit-conn-ic.yaml ``` ❶ `policy`: set to `redis-cluster` to use a Redis cluster for rate limiting. ❷ `redis_cluster_nodes`: set to Redis node addresses in the Redis cluster. ❸ `redis_password`: set to the password of the Redis cluster, if any. ❹ `redis_cluster_name`: set to the Redis cluster name. ➎ `redis_cluster_ssl`: enable SSL/TLS communication with Redis cluster. Send five concurrent requests to the route: ``` seq 1 5 | xargs -n1 -P5 bash -c 'curl -s -o /dev/null -w "Response: %{http_code}\n" "http://127.0.0.1:9080/get"' ``` You should see responses similar to the following, where excessive requests are rejected: ``` Response: 200 Response: 200 Response: 429 Response: 429 Response: 429 ``` This shows the two routes configured in different APISIX instances share the same quota. ### Rate Limit by Rules[​](#rate-limit-by-rules "Direct link to Rate Limit by Rules") The following example demonstrates how you can configure `limit-conn` to apply different rate-limiting rules (available from API7 Enterprise 3.8.17) based on request attributes. In this example, rate limits are applied based on HTTP header values that represent the caller’s access tier. Note that all rules are applied sequentially. If a configured key does not exist, the corresponding rule will be skipped. tip In addition to HTTP headers, you can also base rules on other [built-in variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md) to implement more flexible and fine-grained rate-limiting strategies. Create a route with the `limit-conn` plugin that applies different rate limits based on request headers, allowing requests to be rate limited per subscription (`X-Subscription-ID`) and enforcing a stricter limit for trial users (`X-Trial-ID`): * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-conn-rules-route", "uri": "/get", "plugins": { "limit-conn": { "rejected_code": 429, "default_conn_delay": 0.1, "policy": "local", "rules": [ { "key": "${http_x_subscription_id}", "conn": "${http_x_custom_conn ?? 5}", "burst": 1 }, { "key": "${http_x_trial_id}", "conn": 1, "burst": 1 } ] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-conn-rules-route plugins: limit-conn: rejected_code: 429 default_conn_delay: 0.1 policy: local rules: - key: "${http_x_subscription_id}" conn: "${http_x_custom_conn ?? 5}" burst: 1 - key: "${http_x_trial_id}" conn: 1 burst: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-conn-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-conn-plugin-config spec: plugins: - name: limit-conn config: rejected_code: 429 default_conn_delay: 0.1 policy: local rules: - key: "${http_x_subscription_id}" conn: "${http_x_custom_conn ?? 5}" burst: 1 - key: "${http_x_trial_id}" conn: 1 burst: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-conn-rules-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-conn-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-conn-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-conn-rules-route spec: ingressClassName: apisix http: - name: limit-conn-rules-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-conn config: rejected_code: 429 default_conn_delay: 0.1 policy: local rules: - key: "${http_x_subscription_id}" conn: "${http_x_custom_conn ?? 5}" burst: 1 - key: "${http_x_trial_id}" conn: 1 burst: 1 ``` Apply the configuration: ``` kubectl apply -f limit-conn-ic.yaml ``` ❶ Use the value of the `X-Subscription-ID` request header as the rate-limiting key. ❷ Set the request connection dynamically based on the `X-Custom-Conn` header. If the header is not provided, a default concurrent connection count of 5 is applied. ❸ Use the value of the `X-Trial-ID` request header as the rate-limiting key. To verify rate limiting, send 7 concurrent requests to the route with the same subscription ID: ``` seq 1 7 | xargs -n1 -P7 bash -c 'curl -s -o /dev/null -w "Response: %{http_code}\n" "http://127.0.0.1:9080/get" -H "X-Subscription-ID: sub-123456789"' ``` You should see the following response, which shows that the default concurrent connection limit of 5 with a burst of 1 is applied when the `X-Custom-Conn` header is not provided: ``` Response: 429 Response: 200 Response: 200 Response: 200 Response: 200 Response: 200 Response: 200 ``` Send 5 concurrent requests to the route with the same subscription ID and set the `X-Custom-Conn` header to 1: ``` seq 1 5 | xargs -n1 -P5 bash -c 'curl -s -o /dev/null -w "Response: %{http_code}\n" "http://127.0.0.1:9080/get" -H "X-Subscription-ID: sub-123456789" -H "X-Custom-Conn: 1"' ``` You should see the following response, which shows that the concurrent connection limit of 1 with a burst of 1 is applied: ``` Response: 429 Response: 429 Response: 429 Response: 200 Response: 200 ``` Finally, generate 5 requests to the route with the trial ID header: ``` seq 1 5 | xargs -n1 -P5 bash -c 'curl -s -o /dev/null -w "Response: %{http_code}\n" "http://127.0.0.1:9080/get" -H "X-Trial-ID: trial-123456789"' ``` You should see the following response, which shows that the concurrent connection limit of 1 with a burst of 1 is applied: ``` Response: 429 Response: 429 Response: 429 Response: 200 Response: 200 ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. note In API7 Enterprise (from 3.8.17) and in APISIX (from 3.16.0), you should configure one of the following parameter sets, but not both: * `conn`, `burst`, `default_conn_delay`, `key` * `rules`, `default_conn_delay` - conn integer | string required vaild vaule: greater than 0 *** The maximum number of concurrent requests allowed. Requests exceeding the configured limit and below `conn + burst` will be delayed. A string value can reference a [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md) by prefixing the variable name with a dollar sign (`$`). A resolved string must be a positive integer no greater than `9007199254740991`. If resolution fails or produces an invalid value, the gateway returns `500 Internal Server Error` unless degradation is enabled. String-value support was introduced in API7 Enterprise 3.8.17 and APISIX 3.16.0. The validation requirements were introduced in API7 Enterprise 3.9.14 and 3.10.1, and in APISIX 3.17.0. Earlier APISIX versions accept only integer values. - burst integer | string required vaild vaule: greater than or equal to 0 *** The number of excessive concurrent requests allowed to be delayed. Requests exceeding `conn + burst` will be rejected immediately. A string value can reference a [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md) by prefixing the variable name with a dollar sign (`$`). A resolved string must be a non-negative integer no greater than `9007199254740991`. If resolution fails or produces an invalid value, the gateway returns `500 Internal Server Error` unless degradation is enabled. String-value support was introduced in API7 Enterprise 3.8.17 and APISIX 3.16.0. The validation requirements were introduced in API7 Enterprise 3.9.14 and 3.10.1, and in APISIX 3.17.0. Earlier APISIX versions accept only integer values. - default\_conn\_delay number required vaild vaule: greater than 0 *** Processing latency allowed in seconds for concurrent requests exceeding `conn` and up to `conn + burst`, which can be dynamically adjusted based on `only_use_default_delay` setting. - only\_use\_default\_delay boolean default: `false` *** If false, delay requests proportionally based on how much they exceed the `conn` limit. The delay grows larger as congestion increases. For instance, with `conn` being `5`, `burst` being `3`, and `default_conn_delay` being `1`, 6 concurrent requests would result in a 1-second delay, 7 requests a 2-second delay, 8 requests a 3-second delay, and so on, until the total limit of `conn + burst` is reached, beyond which requests are rejected. If true, use `default_conn_delay` to delay all excessive requests within the `burst` range. Requests beyond `conn + burst` are rejected immediately. For instance, with `conn` being `5`, `burst` being `3`, and `default_conn_delay` being `1`, 6, 7, or 8 concurrent requests are all delayed by exactly 1 second each. - key\_type string default: `var` vaild vaule: `var` or `var_combination` *** The type of key. If the `key_type` is `var`, the `key` is interpreted as a variable. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. - key string required *** The key to count requests by. If the `key_type` is `var`, the `key` is interpreted as a variable. The variable does not need to be prefixed by a dollar sign (`$`). See [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) for available variables. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. All variables should be prefixed by dollar signs (`$`). For example, to configure the `key` to use a combination of two request headers `custom-a` and `custom-b`, the `key` should be configured as `$http_custom_a $http_custom_b`. - rejected\_code integer default: `503` vaild vaule: between 200 and 599 inclusive *** The HTTP status code returned when a request is rejected for exceeding the threshold. - rejected\_msg string vaild vaule: any non-empty string *** The response body returned when a request is rejected for exceeding the threshold. - allow\_degradation boolean default: `false` *** If true, allow the gateway to continue handling requests without the plugin when the plugin or its dependencies become unavailable. - rules array\[object] *** An array of rate-limiting rules that are applied sequentially. Rule support was introduced in API7 Enterprise 3.8.17 and APISIX 3.16.0. * conn integer | string required vaild vaule: greater than 0 *** The maximum number of concurrent requests allowed. Requests exceeding the configured limit and below `conn + burst` will be delayed. A string value can reference a [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md) by prefixing the variable name with a dollar sign (`$`). A resolved string must be a positive integer no greater than `9007199254740991`. String-value validation was introduced in API7 Enterprise 3.9.14 and 3.10.1, and in APISIX 3.17.0. * burst integer | string required vaild vaule: greater than or equal to 0 *** The number of excessive concurrent requests allowed to be delayed. Requests exceeding `conn + burst` will be rejected immediately. A string value can reference a [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md) by prefixing the variable name with a dollar sign (`$`). A resolved string must be a non-negative integer no greater than `9007199254740991`. String-value validation was introduced in API7 Enterprise 3.9.14 and 3.10.1, and in APISIX 3.17.0. * key string required *** The key to count requests by. If the configured key does not exist, the rule will not be executed. If the `key_type` is `var`, the `key` is interpreted as a variable. The variable does not need to be prefixed by a dollar sign (`$`). See [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) for available variables. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. All variables should be prefixed by dollar signs (`$`). For example, to configure the `key` to use a combination of two request headers `custom-a` and `custom-b`, the `key` should be configured as `$http_custom_a $http_custom_b`. - policy string default: `local` vaild vaule: `local`, `redis`, or `redis-cluster` *** The policy for rate limiting counter. Required for API7 Enterprise (from 3.9.0) and optional for APISIX. When set to `local`, the counter is stored in memory locally. When set to `redis`, the counter is stored on a Redis instance. When set to `redis-cluster`, the counter is stored in a Redis cluster. - redis\_host string *** The address of the Redis node. Required when `policy` is `redis`. - redis\_port integer default: `6379` vaild vaule: greater than or equal to 1 *** The port of the Redis node when `policy` is `redis`. - redis\_username string *** The username for Redis if Redis ACL is used. If you use the legacy authentication method `requirepass`, configure only the `redis_password`. Used when `policy` is `redis`. - redis\_password string *** The password of the Redis node when `policy` is `redis` or `redis-cluster`. The password is encrypted at rest in API7 Enterprise. In APISIX, [enable data encryption](https://docs.api7.ai/apisix/production/security/data-encryption-with-keyring.md) to encrypt it before etcd storage. Encryption was introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.18.0. - redis\_database integer default: `0` vaild vaule: greater than or equal to 0 *** The database number in Redis when `policy` is `redis`. - redis\_ssl boolean default: `false` *** If true, use SSL to connect to Redis when `policy` is `redis`. - redis\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis`. - redis\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** The Redis timeout value in milliseconds when `policy` is `redis` or `redis-cluster`. - redis\_keepalive\_timeout integer default: `10000` vaild vaule: greater than or equal to 1000 *** Keepalive timeout in milliseconds for Redis when `policy` is `redis` or `redis-cluster`. This parameter is available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line, and in APISIX from version 3.15.0. - redis\_keepalive\_pool integer default: `100` vaild vaule: greater than or equal to 1 *** Keepalive pool size for Redis when `policy` is `redis` or `redis-cluster`. This parameter is available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line, and in APISIX from version 3.15.0. - key\_ttl integer default: `3600` *** TTL of the Redis key in seconds when `policy` is `redis` or `redis-cluster`. Available in API7 Enterprise from version 3.9.4 and in APISIX from version 3.15.0. - redis\_cluster\_nodes array\[string] *** The list of Redis cluster nodes with at least one address. Required when `policy` is `redis-cluster`. - redis\_cluster\_name string *** The name of the Redis cluster. Required when `policy` is `redis-cluster`. - redis\_cluster\_ssl boolean default: `false` *** If true, use SSL to connect to Redis cluster when `policy` is `redis-cluster`. - redis\_cluster\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis-cluster`. --- # limit-count The `limit-count` plugin uses a fixed window algorithm to limit the rate of requests by the number of requests within a given time interval. Requests exceeding the configured quota will be rejected. When `show_limit_quota_header` is `true` (the default), responses include quota headers: * `X-RateLimit-Limit`: the total quota * `X-RateLimit-Remaining`: the remaining quota * `X-RateLimit-Reset`: number of seconds left for the counter to reset You can rename these headers with [plugin metadata](#customize-rate-limiting-headers). If you configure multiple `rules`, each rule inserts its `header_prefix` (or the rule index, starting at 1) before `RateLimit-`, so clients can tell which rule produced each header. A prefix of `foo` becomes `X-foo-RateLimit-Limit`; omitting the prefix on the first rule becomes `X-1-RateLimit-Limit`. Limit, Remaining, and Reset keep the same meaning. ## Local vs Redis Rate Limiting[​](#local-vs-redis-rate-limiting "Direct link to Local vs Redis Rate Limiting") The `limit-count` plugin supports two modes of rate limiting: * **Local rate limiting**: Limits are enforced independently on each gateway instance. Each instance maintains its own counters, so the effective limit is roughly (limit × number of instances) when traffic is spread across instances. This is the default when no `policy` is set or when `policy` is `local`. * **Redis-based rate limiting**: Limits are shared across all gateway instances through Redis. All instances share the same quota, so the configured limit applies to all gateway instances. Set `policy` to `redis` for a single Redis instance, `redis-cluster` for a Redis cluster, or `redis-sentinel` for Redis nodes managed by Redis Sentinel. note Redis Sentinel, sliding windows, and delayed Redis synchronization are built directly into `limit-count`. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. Multiple rules and variable-resolved limits are also supported. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.16.0. API7 Enterprise retains the separate `limit-count-advanced` plugin for backward compatibility. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `limit-count` in different scenarios. ### Apply Rate Limiting by Remote Address[​](#apply-rate-limiting-by-remote-address "Direct link to Apply Rate Limiting by Remote Address") The following example demonstrates the rate limiting of requests by a single variable, `remote_addr`. Create a route with `limit-count` plugin that allows for a quota of 1 within a 30-second window per remote address: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route", "uri": "/get", "plugins": { "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "key_type": "var", "key": "remote_addr", "policy": "local" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-count-route plugins: limit-count: count: 1 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-plugin-config spec: plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-route spec: ingressClassName: apisix http: - name: limit-count-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: limit-count enable: true config: count: 1 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local ``` Apply the configuration to your cluster: ``` kubectl apply -f limit-count-ic.yaml ``` Send a request to verify: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response. The request has consumed all the quota allowed for the time window. If you send the request again within the same 30-second time interval, you should receive an `HTTP/1.1 429 Too Many Requests` response, indicating the request surpasses the quota threshold. ### Apply Rate Limiting with a Sliding Window[​](#apply-rate-limiting-with-a-sliding-window "Direct link to Apply Rate Limiting with a Sliding Window") Set `window_type` to `sliding` to smooth traffic bursts at window boundaries by weighting the previous window's count. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. Create a route with the `limit-count` plugin that allows a quota of 10 requests within a 30-second sliding window per remote address: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route", "uri": "/get", "plugins": { "limit-count": { "count": 10, "time_window": 30, "rejected_code": 429, "key_type": "var", "key": "remote_addr", "policy": "local", "window_type": "sliding" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: limit-count-route uris: - /get plugins: limit-count: count: 10 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local window_type: sliding upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-plugin-config spec: plugins: - name: limit-count config: count: 10 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local window_type: sliding --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-route spec: ingressClassName: apisix http: - name: limit-count-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: limit-count enable: true config: count: 10 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local window_type: sliding ``` Apply the configuration to your cluster: ``` kubectl apply -f limit-count-ic.yaml ``` Send a request to verify: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response. As you keep sending requests, the sliding window enforces the quota more evenly across the boundary between consecutive windows than the fixed window does. ### Apply Rate Limiting by Remote Address and Consumer Name[​](#apply-rate-limiting-by-remote-address-and-consumer-name "Direct link to Apply Rate Limiting by Remote Address and Consumer Name") The following example demonstrates the rate limiting of requests by a combination of variables, `remote_addr` and `consumer_name`. It allows for a quota of 1 within a 30-second window per remote address and for each [consumer](https://docs.api7.ai/apisix/key-concepts/consumers.md). * Admin API * ADC * Ingress Controller Create consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john" }' ``` Create `key-auth` credential for consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Create consumer `jane`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jane" }' ``` Create `key-auth` credential for consumer `jane`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jane/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-key-auth", "plugins": { "key-auth": { "key": "jane-key" } } }' ``` Create a route with `key-auth` and `limit-count` plugins: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route", "uri": "/get", "plugins": { "key-auth": {}, "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "key_type": "var_combination", "key": "$remote_addr $consumer_name", "policy": "local" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create two consumers and a route that enables rate limiting by consumers: adc.yaml ``` consumers: - username: john credentials: - name: key-auth type: key-auth config: key: john-key - username: jane credentials: - name: key-auth type: key-auth config: key: jane-key services: - name: limit-count-service routes: - name: limit-count-route uris: - /get plugins: key-auth: {} limit-count: count: 1 time_window: 30 rejected_code: 429 key_type: var_combination key: "$remote_addr $consumer_name" policy: local upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create two consumers and a route that enables rate limiting by consumers: * Gateway API * APISIX CRD limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: john-key --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jane spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: jane-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-plugin-config spec: plugins: - name: key-auth config: _meta: disable: false - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 key_type: var_combination key: "$remote_addr $consumer_name" policy: local --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jane spec: ingressClassName: apisix authParameter: keyAuth: value: key: jane-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-route spec: ingressClassName: apisix http: - name: limit-count-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: key-auth enable: true - name: limit-count enable: true config: count: 1 time_window: 30 rejected_code: 429 key_type: var_combination key: "$remote_addr $consumer_name" policy: local ``` Apply the configuration to your cluster: ``` kubectl apply -f limit-count-ic.yaml ``` ❶ `key-auth`: enable key authentication on the route. ❷ `key_type`: set to `var_combination` to interpret the `key` as a combination of variables. ❸ `key`: set to `$remote_addr $consumer_name` to apply rate limiting quota by remote address and for each consumer. Send a request as the consumer `jane`: ``` curl -i "http://127.0.0.1:9080/get" -H 'apikey: jane-key' ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body. This request has consumed all the quota set for the time window. If you send the same request as the consumer `jane` within the same 30-second time interval, you should receive an `HTTP/1.1 429 Too Many Requests` response, indicating the request surpasses the quota threshold. Send the same request as the consumer `john` within the same 30-second time interval: ``` curl -i "http://127.0.0.1:9080/get" -H 'apikey: john-key' ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body, indicating the request is not rate limited. Send the same request as the consumer `john` again within the same 30-second time interval, you should receive an `HTTP/1.1 429 Too Many Requests` response. This verifies the plugin rate limits by the combination of variables, `remote_addr` and `consumer_name`. ### Share Quota among Routes[​](#share-quota-among-routes "Direct link to Share Quota among Routes") The following example demonstrates the sharing of rate limiting quota among multiple routes by configuring the `group` of the `limit-count` plugin. Note that the configurations of the `limit-count` plugin of the same `group` should be identical. To avoid update anomalies and repetitive configurations, you can create a [service](https://docs.api7.ai/apisix/key-concepts/services.md) with `limit-count` plugin and upstream for routes to connect to. * Admin API * ADC * Ingress Controller Create a service with rate limiting group: ``` curl "http://127.0.0.1:9180/apisix/admin/services" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-service", "plugins": { "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "policy": "local", "group": "srv1" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create two routes that use the same service to share the quota: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route-1", "service_id": "limit-count-service", "uri": "/get1", "plugins": { "proxy-rewrite": { "uri": "/get" } } }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route-2", "service_id": "limit-count-service", "uri": "/get2", "plugins": { "proxy-rewrite": { "uri": "/get" } } }' ``` Create a service with two routes that share the same rate limiting quota: adc.yaml ``` services: - name: limit-count-service plugins: limit-count: count: 1 time_window: 30 rejected_code: 429 policy: local group: srv1 routes: - name: limit-count-route-1 uris: - /get1 plugins: proxy-rewrite: uri: /get - name: limit-count-route-2 uris: - /get2 plugins: proxy-rewrite: uri: /get upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create two HTTPRoute resources that reference the same PluginConfig to share quota: limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-plugin-config spec: plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 policy: local group: srv1 - name: proxy-rewrite config: uri: /get --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-route-1 spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get1 filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-plugin-config backendRefs: - name: httpbin-external-domain port: 80 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-route-2 spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get2 filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Create an ApisixRoute with multiple paths that share the same plugin configuration: limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-shared-route spec: ingressClassName: apisix http: - name: limit-count-shared match: paths: - /get1 - /get2 upstreams: - name: httpbin-external-domain plugins: - name: proxy-rewrite enable: true config: uri: /get - name: limit-count enable: true config: count: 1 time_window: 30 rejected_code: 429 policy: local group: srv1 ``` Apply the configuration to your cluster: ``` kubectl apply -f limit-count-ic.yaml ``` note The [`proxy-rewrite`](https://docs.api7.ai/hub/proxy-rewrite.md) plugin is used to rewrite the URI to `/get` so that requests are forwarded to the correct endpoint. Send a request to route `/get1`: ``` curl -i "http://127.0.0.1:9080/get1" ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body. Send the same request to route `/get2` within the same 30-second time interval: ``` curl -i "http://127.0.0.1:9080/get2" ``` You should receive an `HTTP/1.1 429 Too Many Requests` response, which verifies the two routes share the same rate limiting quota. ### Share Quota Among APISIX Nodes with a Redis Server[​](#share-quota-among-apisix-nodes-with-a-redis-server "Direct link to Share Quota Among APISIX Nodes with a Redis Server") The following example demonstrates rate limiting across multiple APISIX nodes with a Redis server, so different APISIX nodes share the same rate limiting quota. Configure the route with Redis connection details. Adjust the Redis host, credentials, and database for your environment. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route", "uri": "/get", "plugins": { "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "key": "remote_addr", "policy": "redis", "redis_host": "192.168.xxx.xxx", "redis_port": 6379, "redis_password": "p@ssw0rd", "redis_database": 1 } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a route with Redis-based rate limiting: adc.yaml ``` services: - name: redis-limit-service routes: - name: redis-limit-route uris: - /get plugins: limit-count: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis redis_host: "192.168.xxx.xxx" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-redis-plugin-config spec: plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis redis_host: "redis-service.aic.svc" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: redis-limit-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-redis-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: redis-limit-route spec: ingressClassName: apisix http: - name: redis-limit-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: limit-count enable: true config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis redis_host: "redis-service.aic.svc" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 ``` Apply the configuration to your cluster: ``` kubectl apply -f limit-count-ic.yaml ``` ❶ `policy`: set to `redis` to use a Redis instance for rate limiting. ❷ `redis_host`: set to Redis instance IP address. ❸ `redis_port`: set to Redis instance listening port. ❹ `redis_password`: set to the password of the Redis instance, if any. ❺ `redis_database`: set to the database number in the Redis instance. Send a request to an APISIX instance: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body. Send the same request to a different APISIX instance within the same 30-second interval. You should receive an `HTTP/1.1 429 Too Many Requests` response, verifying that routes on different APISIX nodes share the same quota. ### Reduce Redis Round Trips with Delayed Synchronization[​](#reduce-redis-round-trips-with-delayed-synchronization "Direct link to Reduce Redis Round Trips with Delayed Synchronization") By default, Redis-based policies synchronize the counter on every request. You can instead accumulate increments locally and synchronize them with Redis at a configured interval. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. To enable delayed synchronization in the complete Redis example above, add `sync_interval: 1` to the `limit-count` configuration. The value is in seconds, must be at least `0.1`, and must be smaller than a numeric `time_window`. This option also works with the `redis-cluster` and `redis-sentinel` policies. Set it to `-1`, the default, to synchronize every request. Delayed synchronization reduces Redis round trips, but counters on different workers or gateway nodes can temporarily diverge before their local increments are synchronized. As a result, the shared quota is not exact during the synchronization interval and concurrent requests can exceed the configured count. Use direct synchronization when exact cross-node enforcement is more important than reducing Redis traffic. ### Share Quota Among APISIX Nodes with a Redis Cluster[​](#share-quota-among-apisix-nodes-with-a-redis-cluster "Direct link to Share Quota Among APISIX Nodes with a Redis Cluster") You can also use a Redis cluster to apply the same quota across multiple APISIX nodes, such that different APISIX nodes share the same rate limiting quota. Ensure that your Redis instances are running in [cluster mode](https://redis.io/docs/management/scaling/#create-and-use-a-redis-cluster). A minimum of two nodes are required for the `limit-count` plugin configurations. Configure the route with Redis cluster details. Adjust the cluster nodes, credentials, and cluster name for your environment. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route", "uri": "/get", "plugins": { "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "key": "remote_addr", "policy": "redis-cluster", "redis_cluster_nodes": [ "192.168.xxx.xxx:6379", "192.168.xxx.xxx:16379" ], "redis_password": "p@ssw0rd", "redis_cluster_name": "redis-cluster", "redis_cluster_ssl": true } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a route with Redis cluster-based rate limiting: adc.yaml ``` services: - name: redis-cluster-limit-service routes: - name: redis-cluster-limit-route uris: - /get plugins: limit-count: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-cluster redis_cluster_nodes: - "192.168.xxx.xxx:6379" - "192.168.xxx.xxx:16379" redis_password: "p@ssw0rd" redis_cluster_name: redis-cluster redis_cluster_ssl: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-redis-cluster-plugin-config spec: plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-cluster redis_cluster_nodes: - "redis-cluster-0.redis-cluster.aic.svc:6379" - "redis-cluster-1.redis-cluster.aic.svc:6379" redis_password: "p@ssw0rd" redis_cluster_name: redis-cluster redis_cluster_ssl: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: redis-cluster-limit-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-redis-cluster-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: redis-cluster-limit-route spec: ingressClassName: apisix http: - name: redis-cluster-limit-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: limit-count enable: true config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-cluster redis_cluster_nodes: - "redis-cluster-0.redis-cluster.aic.svc:6379" - "redis-cluster-1.redis-cluster.aic.svc:6379" redis_password: "p@ssw0rd" redis_cluster_name: redis-cluster redis_cluster_ssl: true ``` Apply the configuration to your cluster: ``` kubectl apply -f limit-count-ic.yaml ``` ❶ `policy`: set to `redis-cluster` to use a Redis cluster for rate limiting. ❷ `redis_cluster_nodes`: set to Redis node addresses in the Redis cluster. ❸ `redis_password`: set to the password of the Redis cluster, if any. ❹ `redis_cluster_name`: set to the Redis cluster name. ➎ `redis_cluster_ssl`: enable SSL/TLS communication with Redis cluster. Send a request to an APISIX instance: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body. Send the same request to a different APISIX instance within the same 30-second interval. You should receive an `HTTP/1.1 429 Too Many Requests` response, verifying that routes on different APISIX nodes share the same quota. ### Share Quota Among APISIX Nodes with Redis Sentinel[​](#share-quota-among-apisix-nodes-with-redis-sentinel "Direct link to Share Quota Among APISIX Nodes with Redis Sentinel") Redis Sentinel provides a shared counter backend and discovers a promoted Redis master after failover. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. Configure the route with Redis Sentinel details. Adjust the Sentinel nodes, credentials, and monitored master name for your environment. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route", "uri": "/get", "plugins": { "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "key": "remote_addr", "policy": "redis-sentinel", "redis_sentinels": [ { "host": "192.168.xxx.xxx", "port": 26379 }, { "host": "192.168.xxx.xxx", "port": 26380 } ], "redis_master_name": "mymaster", "sentinel_username": "sentinel-user", "sentinel_password": "p@ssw0rd", "redis_database": 1 } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: limit-count-route uris: - /get plugins: limit-count: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-sentinel redis_sentinels: - host: 192.168.xxx.xxx port: 26379 - host: 192.168.xxx.xxx port: 26380 redis_master_name: mymaster sentinel_username: sentinel-user sentinel_password: p@ssw0rd redis_database: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-plugin-config spec: plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-sentinel redis_sentinels: - host: 192.168.xxx.xxx port: 26379 - host: 192.168.xxx.xxx port: 26380 redis_master_name: mymaster sentinel_username: sentinel-user sentinel_password: p@ssw0rd redis_database: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-route spec: ingressClassName: apisix http: - name: limit-count-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: limit-count enable: true config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-sentinel redis_sentinels: - host: 192.168.xxx.xxx port: 26379 - host: 192.168.xxx.xxx port: 26380 redis_master_name: mymaster sentinel_username: sentinel-user sentinel_password: p@ssw0rd redis_database: 1 ``` Apply the configuration to your cluster: ``` kubectl apply -f limit-count-ic.yaml ``` Where: * `policy` is set to `redis-sentinel` to use Redis nodes managed by Redis Sentinel. * `redis_sentinels` lists the Sentinel nodes, each with a `host` and `port`. * `redis_master_name` is the name of the master that Sentinel monitors. * `sentinel_username` and `sentinel_password` authenticate with the Sentinel nodes, if required. Send a request to an APISIX instance: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response with the corresponding response body. Send the same request to a different APISIX instance within the same 30-second interval. You should receive an `HTTP/1.1 429 Too Many Requests` response, verifying that routes on different APISIX nodes share the same quota. ### Rate Limit with Anonymous Consumer[​](#rate-limit-with-anonymous-consumer "Direct link to Rate Limit with Anonymous Consumer") The following example demonstrates how you can configure different rate limiting policies by regular and anonymous consumers, where the anonymous consumer does not need to authenticate and has less quota. While this example uses [`key-auth`](https://docs.api7.ai/hub/key-auth.md) for authentication, the anonymous consumer can also be configured with [`basic-auth`](https://docs.api7.ai/hub/basic-auth.md), [`jwt-auth`](https://docs.api7.ai/hub/jwt-auth.md), and [`hmac-auth`](https://docs.api7.ai/hub/hmac-auth.md). * Admin API * ADC * Ingress Controller Create consumer `john` with a quota of 3: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john", "plugins": { "limit-count": { "count": 3, "time_window": 30, "rejected_code": 429, "policy": "local" } } }' ``` Create `key-auth` credential for consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Create an anonymous user `anonymous` with a quota of 1: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "anonymous", "plugins": { "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "policy": "local" } } }' ``` Create a route with `key-auth` plugin that accepts anonymous consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "key-auth-route", "uri": "/anything", "plugins": { "key-auth": { "anonymous_consumer": "anonymous" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Configure consumers with different rate limits and a route that accepts anonymous users: adc.yaml ``` consumers: - username: john plugins: limit-count: count: 3 time_window: 30 rejected_code: 429 policy: local credentials: - name: key-auth type: key-auth config: key: john-key - username: anonymous plugins: limit-count: count: 1 time_window: 30 rejected_code: 429 policy: local services: - name: anonymous-rate-limit-service routes: - name: key-auth-route uris: - /anything plugins: key-auth: anonymous_consumer: anonymous upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Configure consumers with different rate limits and a route that accepts anonymous users: limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: john-key plugins: - name: limit-count config: count: 3 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: anonymous spec: gatewayRef: name: apisix plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 policy: local --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: key-auth-plugin-config spec: plugins: - name: key-auth config: anonymous_consumer: aic_anonymous # namespace_consumername --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: key-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: key-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f limit-count-ic.yaml ``` limit-count-apisix-crd.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key plugins: - name: limit-count config: count: 3 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: anonymous spec: ingressClassName: apisix plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 policy: local --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: key-auth-route spec: ingressClassName: apisix http: - name: anything match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: key-auth config: anonymous_consumer: aic_anonymous ``` Apply the configuration to your cluster: ``` kubectl apply -f limit-count-apisix-crd.yaml ``` To verify, send five consecutive requests with `john`'s key: ``` resp=$(seq 5 | xargs -I{} curl "http://127.0.0.1:9080/anything" -H 'apikey: john-key' -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that out of the 5 requests, 3 requests were successful (status code 200) while the others were rejected (status code 429). ``` 200: 3, 429: 2 ``` Send five anonymous requests: ``` resp=$(seq 5 | xargs -I{} curl "http://127.0.0.1:9080/anything" -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that only one request was successful: ``` 200: 1, 429: 4 ``` ### Customize Rate Limiting Headers[​](#customize-rate-limiting-headers "Direct link to Customize Rate Limiting Headers") The following example demonstrates how you can use plugin metadata to customize the rate limiting response header names, which are by default `X-RateLimit-Limit`, `X-RateLimit-Remaining`, and `X-RateLimit-Reset`. * Admin API * ADC * Ingress Controller Configure plugin metadata to customize rate limiting headers: ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/limit-count" -X PUT -d ' { "limit_header": "X-Custom-RateLimit-Limit", "remaining_header": "X-Custom-RateLimit-Remaining", "reset_header": "X-Custom-RateLimit-Reset" }' ``` Create a route with `limit-count` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route", "uri": "/get", "plugins": { "limit-count": { "count": 1, "time_window": 30, "rejected_code": 429, "key_type": "var", "key": "remote_addr", "policy": "local" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Configure plugin metadata and create a route with rate limiting: adc.yaml ``` plugin_metadata: limit-count: limit_header: X-Custom-RateLimit-Limit remaining_header: X-Custom-RateLimit-Remaining reset_header: X-Custom-RateLimit-Reset services: - name: limit-count-service routes: - name: limit-count-route uris: - /get plugins: limit-count: count: 1 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Update your GatewayProxy manifest for the plugin metadata: gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: limit-count: limit_header: X-Custom-RateLimit-Limit remaining_header: X-Custom-RateLimit-Remaining reset_header: X-Custom-RateLimit-Reset ``` * Gateway API * APISIX CRD Create a route with the plugin enabled: limit-count-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-plugin-config spec: plugins: - name: limit-count config: count: 1 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Create a route with the plugin enabled: limit-count-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-route spec: ingressClassName: apisix http: - name: limit-count-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: limit-count enable: true config: count: 1 time_window: 30 rejected_code: 429 key_type: var key: remote_addr policy: local ``` Apply the configuration to your cluster: ``` kubectl apply -f gatewayproxy.yaml -f limit-count-ic.yaml ``` Send a request to verify: ``` curl -i "http://127.0.0.1:9080/get" ``` You should receive an `HTTP/1.1 200 OK` response and see the following headers: ``` X-Custom-RateLimit-Limit: 1 X-Custom-RateLimit-Remaining: 0 X-Custom-RateLimit-Reset: 28 ``` --- # limit-count-advanced [Enterprise](https://api7.ai/enterprise) The `limit-count-advanced` plugin uses a fixed or sliding window algorithm to limit the rate of requests by the number of requests within a given time interval. Requests exceeding the configured quota will be rejected. Specifically: * Fixed window algorithm tracks requests in non-overlapping time intervals. If the request count exceeds the quota in any interval, excess requests are immediately rejected until the next time window begins. * Sliding window algorithm tracks requests in overlapping intervals, smoothing out the rate limit by counting recent requests within the last configured time period, regardless of when the interval began. This method reduces traffic spikes and is more effective at evenly distributing requests over time. Additionally, you may also see the following rate limiting response headers, the name of which can be customized using [plugin metadata](#customize-rate-limiting-headers): * `X-RateLimit-Limit`: the total quota * `X-RateLimit-Remaining`: the remaining quota * `X-RateLimit-Reset`: number of seconds left for the counter to reset Occasionally, you might observe a small negative value for the `X-RateLimit-Remaining`. This is acceptable as the sliding window algorithm is an approximation. ## Local vs Redis Rate Limiting[​](#local-vs-redis-rate-limiting "Direct link to Local vs Redis Rate Limiting") The `limit-count-advanced` plugin supports two modes of rate limiting: * **Local rate limiting**: Limits are enforced independently on each gateway instance. Each instance maintains its own counters, so the effective limit is roughly (limit × number of instances) when traffic is spread across instances. This is the default when no `policy` is set or when `policy` is `local`. * **Redis-based rate limiting**: Limits are shared across all gateway instances through Redis. All instances share the same quota, so the configured limit applies to all gateway instances. ## Examples[​](#examples "Direct link to Examples") The plugin supports sliding window algorithm in addition to the [`limit-count`](https://docs.api7.ai/hub/limit-count.md) plugin features. Please refer to [`limit-count`](https://docs.api7.ai/hub/limit-count.md#examples) plugin for fixed window examples, which can also be configured in `limit-count-advanced`. The examples below demonstrate how you can use `limit-count-advanced` to rate limit using sliding window algorithm. ### Rate Limit with Local Counters[​](#rate-limit-with-local-counters "Direct link to Rate Limit with Local Counters") The following example demonstrates how you can configure `limit-count-advanced` to use the sliding window algorithm for rate limiting on a route, using the counter in the gateway. Note that each gateway instance has its own counter and independent quota. If you have multiple gateway instances that need to share the same quota, please see [share quota among gateways with a Redis server](#share-quota-among-apisix-nodes-with-a-redis-server). Create a route with `limit-count-advanced` plugin that allows for a quota of 5 within a 10-second sliding window per remote address: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-sliding-route", "uri": "/get", "plugins": { "limit-count-advanced": { "policy": "local", "count": 5, "time_window": 10, "rejected_code": 429, "key_type": "var", "key": "remote_addr", "window_type": "sliding" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-count-sliding-route plugins: limit-count-advanced: policy: local count: 5 time_window: 10 rejected_code: 429 key_type: var key: remote_addr window_type: sliding upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-count-advanced-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-advanced-plugin-config spec: plugins: - name: limit-count-advanced config: policy: local count: 5 time_window: 10 rejected_code: 429 key_type: var key: remote_addr window_type: sliding --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-sliding-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-advanced-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-advanced-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-sliding-route spec: ingressClassName: apisix http: - name: limit-count-sliding-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-count-advanced config: policy: local count: 5 time_window: 10 rejected_code: 429 key_type: var key: remote_addr window_type: sliding ``` Apply the configuration: ``` kubectl apply -f limit-count-advanced-ic.yaml ``` Generate 7 requests to the route every other second: ``` for i in $(seq 7); do (curl -I "http://127.0.0.1:9080/get" &) sleep 1 done ``` You should receive `HTTP/1.1 200 OK` responses for most requests, with the remainder being `HTTP 429 Too Many Requests` responses. The specific number rejected depends on when the first request is sent. ### Share Quota Among Gateways with a Redis Server[​](#share-quota-among-gateways-with-a-redis-server "Direct link to Share Quota Among Gateways with a Redis Server") The following example demonstrates the rate limiting of requests across multiple gateway nodes with a Redis server using the sliding window algorithm, such that different gateway nodes share the same rate limiting quota. Create a route with the following configurations in the gateway group: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-sliding-route", "uri": "/get", "plugins": { "limit-count-advanced": { "count": 1, "time_window": 30, "rejected_code": 429, "key": "remote_addr", "policy": "redis", "redis_host": "192.168.xxx.xxx", "redis_port": 6379, "redis_password": "p@ssw0rd", "redis_database": 1, "window_type": "sliding", "sync_interval": 0.2 } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-count-sliding-route plugins: limit-count-advanced: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis redis_host: "192.168.xxx.xxx" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 window_type: sliding sync_interval: 0.2 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-count-advanced-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-advanced-plugin-config spec: plugins: - name: limit-count-advanced config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis redis_host: "redis-service.aic.svc" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 window_type: sliding sync_interval: 0.2 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-sliding-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-advanced-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-advanced-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-sliding-route spec: ingressClassName: apisix http: - name: limit-count-sliding-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-count-advanced config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis redis_host: "redis-service.aic.svc" redis_port: 6379 redis_password: "p@ssw0rd" redis_database: 1 window_type: sliding sync_interval: 0.2 ``` Apply the configuration: ``` kubectl apply -f limit-count-advanced-ic.yaml ``` ❶ `policy`: Set to `redis` to use a Redis instance for rate limiting. ❷ `redis_host`: Set to Redis instance IP address. ❸ `redis_port`: Set to Redis instance listening port. ❹ `redis_password`: Set to the password of the Redis instance, if any. ❺ `redis_database`: Set to the database number in the Redis instance. ❻ `window_type`: Set the window type to sliding window. ❼ `sync_interval`: Set the synchronization interval (optional). Generate 7 requests to the route every other second: ``` for i in $(seq 7); do (curl -I "http://127.0.0.1:9080/get" &) sleep 1 done ``` You should receive `HTTP/1.1 200 OK` responses for most requests, with the remainder being `HTTP 429 Too Many Requests` responses. The specific number rejected depends on when the first request is sent. This verifies routes configured in different gateway nodes share the same quota. ### Share Quota Among Gateway Nodes with a Redis Cluster[​](#share-quota-among-gateway-nodes-with-a-redis-cluster "Direct link to Share Quota Among Gateway Nodes with a Redis Cluster") The following example demonstrates how you can configure `limit-count-advanced` to use the sliding window algorithm and apply the same quota across multiple gateway nodes, such that different gateway nodes share the same rate limiting quota. Ensure that your Redis instances are running in [cluster mode](https://redis.io/docs/management/scaling/#create-and-use-a-redis-cluster). A minimum of two nodes are required for the `limit-count-advanced` plugin configurations. Create a route with the following configurations in the gateway group: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route", "uri": "/get", "plugins": { "limit-count-advanced": { "count": 1, "time_window": 30, "rejected_code": 429, "key": "remote_addr", "policy": "redis-cluster", "redis_cluster_nodes": [ "192.168.xxx.xxx:6379", "192.168.xxx.xxx:16379" ], "redis_password": "p@ssw0rd", "redis_cluster_name": "redis-cluster", "redis_cluster_ssl": true, "window_type": "sliding", "sync_interval": 0.2 } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-count-route plugins: limit-count-advanced: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-cluster redis_cluster_nodes: - "192.168.xxx.xxx:6379" - "192.168.xxx.xxx:16379" redis_password: "p@ssw0rd" redis_cluster_name: "redis-cluster" redis_cluster_ssl: true window_type: sliding sync_interval: 0.2 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-count-advanced-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-advanced-plugin-config spec: plugins: - name: limit-count-advanced config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-cluster redis_cluster_nodes: - "redis-cluster-0.redis-cluster.aic.svc:6379" - "redis-cluster-1.redis-cluster.aic.svc:6379" redis_password: "p@ssw0rd" redis_cluster_name: "redis-cluster" redis_cluster_ssl: true window_type: sliding sync_interval: 0.2 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-advanced-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-advanced-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-route spec: ingressClassName: apisix http: - name: limit-count-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-count-advanced config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-cluster redis_cluster_nodes: - "redis-cluster-0.redis-cluster.aic.svc:6379" - "redis-cluster-1.redis-cluster.aic.svc:6379" redis_password: "p@ssw0rd" redis_cluster_name: "redis-cluster" redis_cluster_ssl: true window_type: sliding sync_interval: 0.2 ``` Apply the configuration: ``` kubectl apply -f limit-count-advanced-ic.yaml ``` ❶ `policy`: Set to `redis-cluster` to use a Redis cluster for rate limiting. ❷ `redis_cluster_nodes`: Set to Redis node addresses in the Redis cluster. ❸ `redis_password`: Set to the password of the Redis cluster, if any. ❹ `redis_cluster_name`: Set to the Redis cluster name. ➎ `redis_cluster_ssl`: Enable SSL/TLS communication with Redis cluster. ❻ `window_type`: Set the window type to sliding window. ❼ `sync_interval`: Set the synchronization interval (optional). Generate 7 requests to the route every other second: ``` for i in $(seq 7); do (curl -I "http://127.0.0.1:9080/get" &) sleep 1 done ``` You should receive `HTTP/1.1 200 OK` responses for most requests, with the remainder being `HTTP 429 Too Many Requests` responses. The specific number rejected depends on when the first request is sent. This verifies routes configured in different gateway nodes share the same quota. ### Customize Rate Limiting Headers[​](#customize-rate-limiting-headers "Direct link to Customize Rate Limiting Headers") The following example demonstrates how you can use plugin metadata to customize the rate limiting response header names, which are by default `X-RateLimit-Limit`, `X-RateLimit-Remaining`, and `X-RateLimit-Reset`. This example assumes a route `/get` with the `limit-count-advanced` plugin already exists. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/limit-count-advanced" -X PUT -d ' { "log_format": { "limit_header": "X-Custom-RateLimit-Limit", "remaining_header": "X-Custom-RateLimit-Remaining", "reset_header": "X-Custom-RateLimit-Reset" } }' ``` adc.yaml ``` ... # other ADC configs plugin_metadata: limit-count-advanced: limit_header: X-Custom-RateLimit-Limit remaining_header: X-Custom-RateLimit-Remaining reset_header: X-Custom-RateLimit-Reset ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update your GatewayProxy manifest for the plugin metadata: gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # add your control plane connection configuration here # .... pluginMetadata: limit-count-advanced: log_format: limit_header: X-Custom-RateLimit-Limit remaining_header: X-Custom-RateLimit-Remaining reset_header: X-Custom-RateLimit-Reset ``` Apply the configuration: ``` kubectl apply -f gatewayproxy.yaml ``` Send a request to verify: ``` curl -i "http://127.0.0.1:9080/get" ``` You should receive an `HTTP/1.1 200 OK` response and see the following headers: ``` X-Custom-RateLimit-Limit: 1 X-Custom-RateLimit-Remaining: 0 X-Custom-RateLimit-Reset: 28 ``` ### Share Quota Among Gateway Nodes with Redis Sentinel[​](#share-quota-among-gateway-nodes-with-redis-sentinel "Direct link to Share Quota Among Gateway Nodes with Redis Sentinel") The following example demonstrates how you can use `limit-count-advanced` plugin with Redis Sentinel policy for rate limiting. Ensure that your Redis instances are running in [Sentinel mode](https://redis.io/docs/latest/operate/oss_and_stack/management/sentinel/). Create a route with the following configurations in the gateway group: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-route", "uri": "/get", "plugins": { "limit-count-advanced": { "count": 1, "time_window": 30, "rejected_code": 429, "key": "remote_addr", "policy": "redis-sentinel", "redis_sentinels": [ {"host": "127.0.0.1", "port": 26379}, {"host": "127.0.10.1", "port": 26379}, {"host": "127.0.101.1", "port": 26379} ], "redis_master_name": "mymaster", "redis_role": "master", "sentinel_username": "admin", "sentinel_password": "admin-password" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-count-route plugins: limit-count-advanced: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-sentinel redis_sentinels: - host: "127.0.0.1" port: 26379 - host: "127.0.10.1" port: 26379 - host: "127.0.101.1" port: 26379 redis_master_name: "mymaster" redis_role: "master" sentinel_username: "admin" sentinel_password: "admin-password" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-count-advanced-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-advanced-plugin-config spec: plugins: - name: limit-count-advanced config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-sentinel redis_sentinels: - host: "redis-sentinel-0.redis-sentinel.aic.svc" port: 26379 - host: "redis-sentinel-1.redis-sentinel.aic.svc" port: 26379 - host: "redis-sentinel-2.redis-sentinel.aic.svc" port: 26379 redis_master_name: "mymaster" redis_role: "master" sentinel_username: "admin" sentinel_password: "admin-password" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-advanced-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-advanced-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-route spec: ingressClassName: apisix http: - name: limit-count-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-count-advanced config: count: 1 time_window: 30 rejected_code: 429 key: remote_addr policy: redis-sentinel redis_sentinels: - host: "redis-sentinel-0.redis-sentinel.aic.svc" port: 26379 - host: "redis-sentinel-1.redis-sentinel.aic.svc" port: 26379 - host: "redis-sentinel-2.redis-sentinel.aic.svc" port: 26379 redis_master_name: "mymaster" redis_role: "master" sentinel_username: "admin" sentinel_password: "admin-password" ``` Apply the configuration: ``` kubectl apply -f limit-count-advanced-ic.yaml ``` ❶ `policy`: Set to `redis-sentinel` to use a Redis in sentinel mode for rate limiting. ❷ `redis_sentinels`: Configure a list of Sentinel node addresses (host and port). ❸ `redis_master_name`: Configure the name of the Redis master group that Sentinels are monitoring. ❹ `redis_role`: Set to `master` to connect to the current Redis master. ❺ `sentinel_username`: Configure the username used to authenticate with Redis Sentinel. ❻ `sentinel_password`: Configure the password used to authenticate with Redis Sentinel. Generate 5 requests to the route every other second: ``` for i in $(seq 5); do (curl -I "http://127.0.0.1:9080/get" &) sleep 1 done ``` You should receive an `HTTP/1.1 200 OK` response for one request in a 30-second window, while the rest being `HTTP 429 Too Many Requests` responses. ### Rate Limit by Rules[​](#rate-limit-by-rules "Direct link to Rate Limit by Rules") The following example demonstrates how you can configure `limit-count-advanced` to apply different rate-limiting rules (available from API7 Enterprise 3.8.17) based on request attributes. In this example, rate limits are applied based on HTTP header values that represent the caller’s access tier. Note that all rules are applied sequentially. If a configured key does not exist, the corresponding rule will be skipped. tip In addition to HTTP headers, you can also base rules on other [built-in variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md) to implement more flexible and fine-grained rate-limiting strategies. Create a route with the `limit-count-advanced` plugin that applies different rate limits based on request headers, allowing requests to be rate limited per subscription (`X-Subscription-ID`) and enforcing a stricter limit for trial users (`X-Trial-ID`): * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-count-rules-route", "uri": "/get", "plugins": { "limit-count-advanced": { "policy": "local", "rejected_code": 429, "rules": [ { "key": "${http_x_subscription_id}", "count": "${http_x_custom_count ?? 5}", "time_window": 60 }, { "key": "${http_x_trial_id}", "count": 1, "time_window": 60 } ] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-count-rules-route plugins: limit-count-advanced: policy: local rejected_code: 429 rules: - key: "${http_x_subscription_id}" count: "${http_x_custom_count ?? 5}" time_window: 60 - key: "${http_x_trial_id}" count: 1 time_window: 60 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-count-advanced-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-count-advanced-plugin-config spec: plugins: - name: limit-count-advanced config: policy: local rejected_code: 429 rules: - key: "${http_x_subscription_id}" count: "${http_x_custom_count ?? 5}" time_window: 60 - key: "${http_x_trial_id}" count: 1 time_window: 60 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-count-rules-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-count-advanced-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-count-advanced-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-count-rules-route spec: ingressClassName: apisix http: - name: limit-count-rules-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-count-advanced config: policy: local rejected_code: 429 rules: - key: "${http_x_subscription_id}" count: "${http_x_custom_count ?? 5}" time_window: 60 - key: "${http_x_trial_id}" count: 1 time_window: 60 ``` Apply the configuration: ``` kubectl apply -f limit-count-advanced-ic.yaml ``` ❶ Use the value of the `X-Subscription-ID` request header as the rate-limiting key. ❷ Set the request limit dynamically based on the `X-Custom-Count` header. If the header is not provided, a default count of 5 requests is applied. ❸ Use the value of the `X-Trial-ID` request header as the rate-limiting key. To verify rate limiting, generate 7 requests to the route with the same subscription ID: ``` resp=$(seq 7 | xargs -I{} curl "http://127.0.0.1:9080/get" -H "X-Subscription-ID: sub-123456789" -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that the default count of 5 requests is applied when `X-Custom-Count` header is not provided: ``` 200: 5, 429: 2 ``` Wait for the time window to reset. Generate 5 requests to the route with the same subscription ID and set the `X-Custom-Count` header to 3: ``` resp=$(seq 5 | xargs -I{} curl "http://127.0.0.1:9080/get" -H "X-Subscription-ID: sub-123456789" -H "X-Custom-Count: 3" -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that the count of 3 requests from the `X-Custom-Count` header is applied: ``` 200: 3, 429: 2 ``` Finally, generate 3 requests to the route with the same trial ID: ``` resp=$(seq 3 | xargs -I{} curl "http://127.0.0.1:9080/get" -H "X-Trial-ID: trial-123456789" -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that the count of 1 request from the second rule is applied: ``` 200: 1, 429: 2 ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. This plugin supports referencing sensitive parameter values from environment variables using the `env://` prefix, or from a secret manager, such as HashiCorp Vault’s [KV secrets engine](https://developer.hashicorp.com/vault/docs/secrets/kv), using the `secret://` prefix. For more information, see [environment variables in plugin](https://docs.api7.ai/apisix/reference/environment-variables.md#plugins) and [secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md). note In API7 Enterprise (from 3.8.17), you should configure one of the following parameter sets, but not both: * `count`, `time_window` * `rules` - count integer | string required vaild vaule: greater than 0 *** The maximum number of requests allowed within a given time interval. In API7 Enterprise (from 3.8.17), this parameter also supports the string data type and allows the use of [built-in variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md) prefixed with a dollar sign (`$`). - time\_window integer | string required vaild vaule: greater than 0 *** The time interval corresponding to the rate limiting `count` in seconds. In API7 Enterprise (from 3.8.17), this parameter also supports the string data type and allows the use of [built-in variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md) prefixed with a dollar sign (`$`). - window\_type string default: `fixed` vaild vaule: `fixed` or `sliding` *** Rate limiting algorithm, fixed window or sliding window. - key\_type string default: `var` vaild vaule: `var`, `var_combination`, or `constant` *** The type of key. If the `key_type` is `var`, the `key` is interpreted as a variable. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. If the `key_type` is `constant`, the `key` is interpreted as a constant. - key string default: `remote_addr` *** The key to count requests by. If the `key_type` is `var`, the `key` is interpreted as a variable. The variable does not need to be prefixed by a dollar sign (`$`). See [built-in variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md) for available variables. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. All variables should be prefixed by dollar signs (`$`). For example, to configure the `key` to use a combination of two request headers `custom-a` and `custom-b`, the `key` should be configured as `$http_custom_a $http_custom_b`. If the `key_type` is `constant`, the `key` is interpreted as a constant value. - rejected\_code integer default: `503` vaild vaule: between 200 and 599 inclusive *** The HTTP status code returned when a request is rejected for exceeding the threshold. - rejected\_msg string vaild vaule: any non-empty string *** The response body returned when a request is rejected for exceeding the threshold. - policy string default: `local` vaild vaule: `local`, `redis`, `redis-cluster`, or `redis-sentinel` *** The policy for rate limiting counter. Set to `local` to store the counter in memory locally. Set to `redis` to store the counter on a Redis instance. Set to `redis-cluster` to store the counter in a Redis cluster. Set to `redis-sentinel` to store the counter on the Redis primary node managed by Redis Sentinel, which ensures high availability by automatically promoting a replica to primary in case of failure. Redis Sentinel provides high availability for Redis when not using Redis Cluster. - redis\_sentinels array\[object] *** An array of Redis Sentinel nodes (host and port). Required when `policy` is `redis-sentinel`. - redis\_master\_name string *** The name of the Redis master group that Sentinels are monitoring. Required when `policy` is `redis-sentinel`. - redis\_role string default: `master` vaild vaule: `master` or `slave` *** The Redis node role to connect to. Configurable when `policy` is `redis-sentinel`. Set to `master` to connect to the current Redis master, and set to `slave` to connect to a Redis replica. - redis\_connect\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** Timeout in milliseconds for establishing a connection to a Redis node. Configurable when `policy` is `redis-sentinel`. - redis\_read\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** Timeout in milliseconds for reading data from a Redis node. Configurable when `policy` is `redis-sentinel`. - sentinel\_username string *** Username used to authenticate with the Redis Sentinel instance. Configurable when `policy` is `redis-sentinel`. - sentinel\_password string *** Password used to authenticate with the Redis Sentinel instance. Configurable when `policy` is `redis-sentinel`. - allow\_degradation boolean default: `false` *** If true, allow the gateway to continue handling requests without the plugin when the plugin or its dependencies become unavailable. - rules array\[object] *** An array of rate-limiting rules that are applied sequentially. Available in API7 Enterprise from 3.8.17. * count integer | string required vaild vaule: greater than 0 *** The maximum number of requests allowed within a given time interval. This parameter also supports the string data type and allows the use of [built-in variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md) prefixed with a dollar sign (`$`). * time\_window integer | string required vaild vaule: greater than 0 *** The time interval corresponding to the rate limiting `count` in seconds. This parameter also supports the string data type and allows the use of [built-in variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md) prefixed with a dollar sign (`$`). * key string required *** The key to count requests by. If the configured key does not exist, the rule will not be executed. The `key` is interpreted as a combination of variables, for example, `$http_custom_a $http_custom_b`. * header\_prefix string *** Prefix for all rate limiting response headers. Available in API7 Enterprise from version 3.8.19. When configured, the prefix is inserted after `X-` in the header name. For example, with `header_prefix` set to `test`, the headers become `X-Test-RateLimit-Limit`, `X-Test-RateLimit-Remaining`, and `X-Test-RateLimit-Reset`. When not configured, the index of the rule in the rules array is used as the prefix. For example, headers for the first rule will be `X-1-RateLimit-Limit`, `X-1-RateLimit-Remaining`, and `X-1-RateLimit-Reset`. - show\_limit\_quota\_header boolean default: `true` *** If true, includes the rate limiting response headers. Specifically, if `rules` is not set, the headers are: * `X-RateLimit-Limit` shows the total quota. * `X-RateLimit-Remaining` shows the remaining quota. * `X-RateLimit-Reset` shows the number of seconds until the counter resets.
When `rules` is set, a prefix (followed by a hyphen) is inserted after `X-`. See `rules.header_prefix` for details. - group string vaild vaule: non-empty *** The `group` ID for the plugin, such that routes of the same `group` can share the same rate limiting counter. - redis\_host string *** The address of the Redis node. Required when `policy` is `redis`. - redis\_port integer default: `6379` vaild vaule: greater than or equal to 1 *** The port of the Redis node when `policy` is `redis`. - redis\_username string *** The username for Redis if Redis ACL is used. If you use the legacy authentication method `requirepass`, configure only the `redis_password`. Used when `policy` is `redis`. - redis\_password string *** The password of the Redis node when `policy` is `redis`, or `redis-cluster`. - redis\_database integer default: `0` vaild vaule: greater than or equal to 0 *** The database number in Redis when `policy` is `redis` or `redis-sentinel`. - redis\_ssl boolean default: `false` *** If true, use SSL to connect to Redis when `policy` is `redis`. - redis\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis`. - redis\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** The Redis timeout value in milliseconds when `policy` is `redis` or `redis-cluster`. - redis\_keepalive\_timeout integer vaild vaule: greater than or equal to 1000 for `redis` and `redis-cluster`; greater than or equal to 1 for `redis-sentinel` *** Time in milliseconds that an idle Redis connection is kept alive in the connection pool before being closed. When `policy` is `redis` or `redis-cluster`, the default is `10000`. When `policy` is `redis-sentinel`, the default is `60000`. For `redis` and `redis-cluster`, this parameter was introduced in API7 Enterprise 3.9.16 and 3.10.3. - redis\_keepalive\_pool integer default: `100` vaild vaule: greater than or equal to 1 *** Maximum number of idle Redis connections in the keepalive pool. Used when `policy` is `redis` or `redis-cluster`. Introduced in API7 Enterprise 3.9.16 and 3.10.3. - redis\_cluster\_nodes array\[string] *** The list of Redis cluster nodes with at least two addresses. Required when `policy` is `redis-cluster`. - redis\_cluster\_name string *** The name of the Redis cluster. Required when `policy` is `redis-cluster`. - redis\_cluster\_ssl boolean default: `false` *** If true, use SSL to connect to Redis cluster when `policy` is `redis-cluster`. - redis\_cluster\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis-cluster`. - sync\_interval number default: `-1` vaild vaule: greater than or equal to 0.1, or the default -1 *** The frequency of synchronizing counter data to Redis. Available only in Enterprise. The `sync_interval` value should be smaller than `time_window`. A value of `1` results in synchronizing counter data every second. A value of `-1` yields no change in synchronizing behaviour, i.e. counter data will be synchronized for each request. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. This plugin supports referencing parameter values from environment variables using the `env://` prefix, or from a secret manager, such as HashiCorp Vault’s [KV secrets engine](https://developer.hashicorp.com/vault/docs/secrets/kv), using the `secret://` prefix. For more information, see [environment variables in plugin](https://docs.api7.ai/apisix/reference/environment-variables.md#plugins) and [secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md). * count integer | string vaild vaule: greater than 0 *** The maximum number of requests allowed within a given time interval. A string value can reference a [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md) by prefixing the variable with a dollar sign (`$`). Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.16.0. Earlier versions accept only integer values. Required with `time_window` when `rules` is not configured. Do not configure `count` or `time_window` together with `rules`. A string value must resolve to a positive integer no greater than `9007199254740991`. An invalid value returns `500 Internal Server Error` unless `allow_degradation` is `true`. Introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.18.0. * time\_window integer | string vaild vaule: greater than 0 *** The time interval corresponding to the rate limiting `count` in seconds. A string value can reference a [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md) by prefixing the variable with a dollar sign (`$`). Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.16.0. Earlier versions accept only integer values. Required with `count` when `rules` is not configured. Do not configure `count` or `time_window` together with `rules`. A string value must resolve to a positive integer no greater than `9007199254740991`. An invalid value returns `500 Internal Server Error` unless `allow_degradation` is `true`. Introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.18.0. * key\_type string default: `var` vaild vaule: `var`, `var_combination`, or `constant` *** The type of key. If the `key_type` is `var`, the `key` is interpreted as a variable. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. If the `key_type` is `constant`, the `key` is interpreted as a constant. * key string default: `remote_addr` *** The key to count requests by. If the `key_type` is `var`, the `key` is interpreted as a variable. The variable does not need to be prefixed by a dollar sign (`$`). See [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) for available variables. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. All variables should be prefixed by dollar signs (`$`). For example, to configure the `key` to use a combination of two request headers `custom-a` and `custom-b`, the `key` should be configured as `$http_custom_a $http_custom_b`. If the `key_type` is `constant`, the `key` is interpreted as a constant value. * rejected\_code integer default: `503` vaild vaule: between 200 and 599 inclusive *** The HTTP status code returned when a request is rejected for exceeding the threshold. * rejected\_msg string vaild vaule: any non-empty string *** The response body returned when a request is rejected for exceeding the threshold. * policy string default: `local` vaild vaule: `local`, `redis`, `redis-cluster`, or `redis-sentinel` *** The policy for the rate limiting counter. Required for API7 Enterprise and optional for APISIX. When set to `local`, the counter is stored in memory locally. When set to `redis`, the counter is stored on a Redis instance. When set to `redis-cluster`, the counter is stored in a Redis cluster. When set to `redis-sentinel`, the counter is stored on Redis nodes managed by Redis Sentinel. The `redis-sentinel` value adds high availability through Sentinel-managed failover. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * window\_type string default: `fixed` vaild vaule: `fixed` or `sliding` *** The rate limiting window algorithm. When set to `fixed`, a fixed window algorithm is used, where each time window enforces the quota independently. When set to `sliding`, a sliding window algorithm is used, which smooths out bursts at window boundaries by weighting the previous window when calculating the current count. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * sync\_interval number default: `-1` vaild vaule: `-1` or greater than or equal to `0.1`; must be smaller than a numeric top-level `time_window` *** The interval in seconds at which the local counter is synchronized to the shared store (Redis). Only takes effect when `policy` is `redis`, `redis-cluster`, or `redis-sentinel`. A value of `-1` disables delayed synchronization, so each request synchronizes directly. When enabled, the value should not be smaller than `0.1` and should be smaller than `time_window`. Enabling delayed synchronization reduces the number of round trips to the shared store at the cost of slightly looser enforcement. At runtime, if the applicable `time_window` is less than or equal to `sync_interval`, the gateway falls back to direct synchronization for that request. This can occur with variable-resolved values or values inside `rules`, which are evaluated at request time. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * allow\_degradation boolean default: `false` *** If true, continue handling requests without rate limiting when the counter backend fails or a variable-resolved `count` or `time_window` is invalid. If false, these failures return `500 Internal Server Error`. * show\_limit\_quota\_header boolean default: `true` *** If true, include quota headers on the response. With the default names, `X-RateLimit-Limit` is the total quota, `X-RateLimit-Remaining` is how many requests remain in the window, and `X-RateLimit-Reset` is the number of seconds until the counter resets. You can rename those headers with plugin metadata. When `rules` is configured, each rule inserts its `header_prefix` (or the rule index if omitted) before `RateLimit-` so Limit, Remaining, and Reset stay distinct per rule. See `rules.header_prefix`. * group string vaild vaule: non-empty *** The `group` ID for the plugin, such that routes of the same `group` can share the same rate limiting counter. * redis\_host string *** The address of the Redis node. Required when `policy` is `redis`. * redis\_port integer default: `6379` vaild vaule: greater than or equal to 1 *** The port of the Redis node when `policy` is `redis`. * redis\_username string *** The username for Redis if Redis ACL is used. If you use the legacy authentication method `requirepass`, configure only the `redis_password`. Used when `policy` is `redis` or `redis-sentinel`. * redis\_password string *** The password of the Redis node when `policy` is `redis`, `redis-cluster`, or `redis-sentinel`. The password is encrypted at rest in API7 Enterprise. In APISIX, [enable data encryption](https://docs.api7.ai/apisix/production/security/data-encryption-with-keyring.md) to encrypt it before etcd storage. Encryption was introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.18.0. * redis\_database integer default: `0` vaild vaule: greater than or equal to 0 *** The database number in Redis when `policy` is `redis` or `redis-sentinel`. * redis\_ssl boolean default: `false` *** If true, use SSL to connect to Redis when `policy` is `redis`. * redis\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis`. * redis\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** The Redis timeout value in milliseconds when `policy` is `redis` or `redis-cluster`. * redis\_keepalive\_timeout integer vaild vaule: greater than or equal to 1000 for `redis` and `redis-cluster`; greater than or equal to 1 for `redis-sentinel` *** Keepalive timeout in milliseconds for Redis connections. When `policy` is `redis` or `redis-cluster`, the default is `10000` and the minimum is `1000`. When `policy` is `redis-sentinel`, the default is `60000` and the minimum is `1`. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.15.0. Sentinel defaults were introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * redis\_keepalive\_pool integer default: `100` vaild vaule: greater than or equal to 1 *** Keepalive pool size for Redis when `policy` is `redis` or `redis-cluster`. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.15.0. * redis\_cluster\_nodes array\[string] *** The list of Redis cluster nodes with at least one address. Required when `policy` is `redis-cluster`. * redis\_cluster\_name string *** The name of the Redis cluster. Required when `policy` is `redis-cluster`. * redis\_cluster\_ssl boolean default: `false` *** If true, use SSL to connect to Redis cluster when `policy` is `redis-cluster`. * redis\_cluster\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis-cluster`. * redis\_sentinels array\[object] *** The list of Redis Sentinel nodes, with at least one node. Required when `policy` is `redis-sentinel`. Each node is an object with `host` (string) and `port` (integer between 1 and 65535). Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * redis\_master\_name string *** The name of the Redis master monitored by Sentinel. Required when `policy` is `redis-sentinel`. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * redis\_role string default: `master` vaild vaule: `master` or `slave` *** The role of the Redis node to connect to when `policy` is `redis-sentinel`. Use `master` for read and write operations, or `slave` for read-only replicas. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * redis\_connect\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** The connection timeout in milliseconds when `policy` is `redis-sentinel`. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * redis\_read\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** The read timeout in milliseconds when `policy` is `redis-sentinel`. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * sentinel\_username string *** The username used to authenticate with the Redis Sentinel nodes when `policy` is `redis-sentinel`. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * sentinel\_password string *** The password used to authenticate with the Redis Sentinel nodes when `policy` is `redis-sentinel`. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. The password is encrypted at rest in API7 Enterprise. In APISIX, [enable data encryption](https://docs.api7.ai/apisix/production/security/data-encryption-with-keyring.md) to encrypt it before etcd storage. Encryption was introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.18.0. * rules array\[object] *** An array of rate-limiting rules that are applied sequentially. Do not configure `rules` together with the top-level `count`, `time_window`, or `group` fields. The top-level `key` and `key_type` are not used in rules mode. Rule keys must be unique. A rule whose `key` variable is absent from a request is skipped. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.16.0. * count integer | string required vaild vaule: greater than 0 *** The maximum number of requests allowed within the given `time_window`. This parameter also supports the string data type and allows the use of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) prefixed with a dollar sign (`$`). A string value must resolve to a positive integer no greater than `9007199254740991`. If the rule applies and the value is invalid, the request returns `500 Internal Server Error` unless `allow_degradation` is `true`. Introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.18.0. * time\_window integer | string required vaild vaule: greater than 0 *** The time interval in seconds for the rate limiting `count`. This parameter also supports the string data type and allows the use of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) prefixed with a dollar sign (`$`). A string value must resolve to a positive integer no greater than `9007199254740991`. If the rule applies and the value is invalid, the request returns `500 Internal Server Error` unless `allow_degradation` is `true`. Introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.18.0. * key string required *** A variable expression that resolves to the key used to count requests for this rule. Prefix every APISIX [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md) or NGINX variable with a dollar sign (`$`), for example `$remote_addr` or `$remote_addr $http_x_tenant`. The top-level `key_type` does not apply to rules. A rule with no resolvable variable is skipped for that request. * header\_prefix string *** A prefix inserted before `RateLimit-` in this rule's quota headers so each rule stays distinguishable. With the default names, `foo` produces `X-foo-RateLimit-Limit`, `X-foo-RateLimit-Remaining`, and `X-foo-RateLimit-Reset`. Those headers still mean total quota, remaining quota, and seconds until reset. If omitted, the rule's array index is used, so the first rule becomes `X-1-RateLimit-Limit`. Only sent when `show_limit_quota_header` is `true`. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * limit\_header string default: `X-RateLimit-Limit` *** Default response header name for the total rate limit quota. * remaining\_header string default: `X-RateLimit-Remaining` *** Default response header name for the remaining rate limit quota. * reset\_header string default: `X-RateLimit-Reset` *** Default response header name for the number of seconds until the rate limit counter resets. --- # limit-req The `limit-req` plugin uses the [leaky bucket](https://en.wikipedia.org/wiki/Leaky_bucket) algorithm to rate limit the number of the requests and allow for throttling. ## Local vs Redis Rate Limiting[​](#local-vs-redis-rate-limiting "Direct link to Local vs Redis Rate Limiting") The `limit-req` plugin supports two modes of rate limiting: * **Local rate limiting**: Limits are enforced independently on each gateway instance. Each instance maintains its own counters, so the effective limit is roughly (limit × number of instances) when traffic is spread across instances. This is the default when no `policy` is set or when `policy` is `local`. * **Redis-based rate limiting**: Limits are shared across all gateway instances through Redis. All instances share the same quota, so the configured limit applies to all gateway instances. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `limit-req` in different scenarios. ### Apply Rate Limiting by Remote Address[​](#apply-rate-limiting-by-remote-address "Direct link to Apply Rate Limiting by Remote Address") The following example demonstrates the rate limiting of HTTP requests by a single variable, `remote_addr`. Create a route with `limit-req` plugin that allows for 1 QPS per remote address: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d ' { "id": "limit-req-route", "uri": "/get", "plugins": { "limit-req": { "rate": 1, "burst": 0, "key": "remote_addr", "key_type": "var", "rejected_code": 429, "policy": "local", "nodelay": true } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-req-route plugins: limit-req: rate: 1 burst: 0 key: remote_addr key_type: var rejected_code: 429 policy: local nodelay: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-req-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-req-plugin-config spec: plugins: - name: limit-req config: rate: 1 burst: 0 key: remote_addr key_type: var rejected_code: 429 policy: local nodelay: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-req-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-req-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-req-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-req-route spec: ingressClassName: apisix http: - name: limit-req-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-req config: rate: 1 burst: 0 key: remote_addr key_type: var rejected_code: 429 policy: local nodelay: true ``` Apply the configuration: ``` kubectl apply -f limit-req-ic.yaml ``` ❶ `rate`: limit the QPS to 1. ❷ `key`: set to `remote_addr` to apply rate limiting quota by remote address and consumer. ❸ `key_type`: set to `var` to interpret the `key` as a variable. Send a request to verify: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response. The request has consumed all the quota allowed for the time window. If you send the request again within the same second, you should receive an `HTTP/1.1 429 Too Many Requests` response, indicating the request surpasses the quota threshold. ### Implement API Throttling[​](#implement-api-throttling "Direct link to Implement API Throttling") The following example demonstrates how to configure `burst` to allow overrun of the rate limiting threshold by the configured value and achieve request throttling. You will also see a comparison against when throttling is not implemented. Create a route with `limit-req` plugin that allows for 1 QPS per remote address, with a `burst` of 1: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-req-route", "uri": "/get", "plugins": { "limit-req": { "rate": 1, "burst": 1, "key": "remote_addr", "rejected_code": 429, "policy": "local" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-req-route plugins: limit-req: rate: 1 burst: 1 key: remote_addr rejected_code: 429 policy: local upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD limit-req-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-req-plugin-config spec: plugins: - name: limit-req config: rate: 1 burst: 1 key: remote_addr rejected_code: 429 policy: local --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-req-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-req-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-req-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-req-route spec: ingressClassName: apisix http: - name: limit-req-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-req config: rate: 1 burst: 1 key: remote_addr rejected_code: 429 policy: local ``` Apply the configuration: ``` kubectl apply -f limit-req-ic.yaml ``` ❶ `burst`: allow for 1 request exceeding the `rate` to be delayed for processing. Generate three requests to the route: ``` resp=$(seq 3 | xargs -I{} curl -i "http://127.0.0.1:9080/get" -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200 responses: $count_200 ; 429 responses: $count_429" ``` You are likely to see that all three requests are successful: ``` 200 responses: 3 ; 429 responses: 0 ``` To see the effect without `burst`, update `burst` to 0 or set `nodelay` to `true` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes/limit-req-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "limit-req": { "nodelay": true } } }' ``` Update the ADC YAML with `nodelay: true`: adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: limit-req-route plugins: limit-req: rate: 1 burst: 1 # alternatively, set burst to 0 key: remote_addr rejected_code: 429 policy: local nodelay: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration with updated plugin settings: ``` adc sync -f adc.yaml ``` Update the manifest file as such: * Gateway API * APISIX CRD limit-req-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-req-plugin-config spec: plugins: - name: limit-req config: rate: 1 burst: 1 # alternatively, set burst to 0 key: remote_addr rejected_code: 429 policy: local nodelay: true ``` limit-req-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-req-route spec: ingressClassName: apisix http: - name: limit-req-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: limit-req config: rate: 1 burst: 1 # alternatively, set burst to 0 key: remote_addr rejected_code: 429 policy: local nodelay: true ``` Apply the updated configuration: ``` kubectl apply -f limit-req-ic.yaml ``` Generate three requests to the route again: ``` resp=$(seq 3 | xargs -I{} curl -i "http://127.0.0.1:9080/get" -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200 responses: $count_200 ; 429 responses: $count_429" ``` You should see a response similar to the following, showing requests surpassing the rate have been rejected: ``` 200 responses: 1 ; 429 responses: 2 ``` ### Apply Rate Limiting by Remote Address and Consumer Name[​](#apply-rate-limiting-by-remote-address-and-consumer-name "Direct link to Apply Rate Limiting by Remote Address and Consumer Name") The following example demonstrates the rate limiting of requests by a combination of variables, `remote_addr` and `consumer_name`. * Admin API * ADC * Ingress Controller Create a consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Create a second consumer `jane`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jane" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jane/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-key-auth", "plugins": { "key-auth": { "key": "jane-key" } } }' ``` Create a route with `key-auth` and `limit-req` plugins: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "limit-req-route", "uri": "/get", "plugins": { "key-auth": {}, "limit-req": { "rate": 1, "burst": 0, "key": "$remote_addr $consumer_name", "key_type": "var_combination", "rejected_code": 429, "policy": "local" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create two consumers and a route that enables rate limiting by consumers: adc.yaml ``` consumers: - username: john credentials: - name: key-auth type: key-auth config: key: john-key - username: jane credentials: - name: key-auth type: key-auth config: key: jane-key services: - name: limit-req-service routes: - name: limit-req-route uris: - /get plugins: key-auth: {} limit-req: rate: 1 burst: 0 key: "$remote_addr $consumer_name" key_type: var_combination rejected_code: 429 policy: local upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create two consumers and a route that enables rate limiting by consumers: * Gateway API * APISIX CRD limit-req-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: john-key --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jane spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: jane-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: limit-req-plugin-config spec: plugins: - name: key-auth config: _meta: disable: false - name: limit-req config: rate: 1 burst: 0 key: "$remote_addr $consumer_name" key_type: var_combination rejected_code: 429 policy: local --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: limit-req-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: limit-req-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` limit-req-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jane spec: ingressClassName: apisix authParameter: keyAuth: value: key: jane-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: limit-req-route spec: ingressClassName: apisix http: - name: limit-req-route match: paths: - /get methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: key-auth config: _meta: disable: false - name: limit-req config: rate: 1 burst: 0 key: "$remote_addr $consumer_name" key_type: var_combination rejected_code: 429 policy: local ``` Apply the configuration: ``` kubectl apply -f limit-req-ic.yaml ``` ❶ `key-auth`: enable key authentication on the route. ❷ `key`: set to `$remote_addr $consumer_name` to apply rate limiting quota by remote address and consumer. ❸ `key_type`: set to `var_combination` to interpret the `key` as a combination of variables. Send two requests simultaneously, each for one consumer: ``` curl -i "http://127.0.0.1:9080/get" -H 'apikey: jane-key' & \ curl -i "http://127.0.0.1:9080/get" -H 'apikey: john-key' & ``` You should receive `HTTP/1.1 200 OK` for both requests, indicating the request has not exceeded the threshold for each consumer. If you send more requests as either consumer within the same second, you should receive an `HTTP/1.1 429 Too Many Requests` response. This verifies the plugin rate limits by the combination of variables, `remote_addr` and `consumer_name`. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * rate number required vaild vaule: greater than 0 *** The maximum number of requests allowed per second. Requests exceeding the rate and below burst will be delayed. * burst number required vaild vaule: greater than or equal to 0 *** The number of requests allowed to be delayed per second for throttling. Requests exceeding the rate and burst will get rejected. * key\_type string default: `var` vaild vaule: `var` or `var_combination` *** The type of key. If the `key_type` is `var`, the `key` is interpreted as a variable. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. * key string required *** The key to count requests by. If the `key_type` is `var`, the `key` is interpreted as a variable. The variable does not need to be prefixed by a dollar sign (`$`). See [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) for available variables. If the `key_type` is `var_combination`, the `key` is interpreted as a combination of variables. All variables should be prefixed by dollar signs (`$`). For example, to configure the `key` to use a combination of two request headers `custom-a` and `custom-b`, the `key` should be configured as `$http_custom_a $http_custom_b`. * rejected\_code integer default: `503` vaild vaule: between 200 and 599 inclusive *** The HTTP status code returned when a request is rejected for exceeding the threshold. * rejected\_msg string vaild vaule: any non-empty string *** The response body returned when a request is rejected for exceeding the threshold. * nodelay boolean default: `false` *** If true, do not delay requests within the burst threshold. * allow\_degradation boolean default: `false` *** If true, allow the gateway to continue handling requests without the plugin when the plugin or its dependencies become unavailable. * policy string default: `local` vaild vaule: `local`, `redis`, or `redis-cluster` *** The policy for rate limiting counter. Required for API7 Enterprise (from 3.9.0) and optional for APISIX. When set to `local`, the counter is stored in memory locally. When set to `redis`, the counter is stored on a Redis instance. When set to `redis-cluster`, the counter is stored in a Redis cluster. * redis\_host string *** The address of the Redis node. Required when `policy` is `redis`. * redis\_port integer default: `6379` vaild vaule: greater than or equal to 1 *** The port of the Redis node when `policy` is `redis`. * redis\_username string *** The username for Redis if Redis ACL is used. If you use the legacy authentication method `requirepass`, configure only the `redis_password`. Used when `policy` is `redis`. * redis\_password string *** The password of the Redis node when `policy` is `redis` or `redis-cluster`. The password is encrypted at rest in API7 Enterprise. In APISIX, [enable data encryption](https://docs.api7.ai/apisix/production/security/data-encryption-with-keyring.md) to encrypt it before etcd storage. Encryption was introduced in API7 Enterprise 3.9.16 and 3.10.2, and APISIX 3.18.0. * redis\_database integer default: `0` vaild vaule: greater than or equal to 0 *** The database number in Redis when `policy` is `redis`. * redis\_ssl boolean default: `false` *** If true, use SSL to connect to Redis when `policy` is `redis`. * redis\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis`. * redis\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** The Redis timeout value in milliseconds when `policy` is `redis` or `redis-cluster`. * redis\_keepalive\_timeout integer default: `10000` vaild vaule: greater than or equal to 1000 *** Keepalive timeout in milliseconds for Redis when `policy` is `redis` or `redis-cluster`. This parameter is available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line, and in APISIX from version 3.15.0. * redis\_keepalive\_pool integer default: `100` vaild vaule: greater than or equal to 1 *** Keepalive pool size for Redis when `policy` is `redis` or `redis-cluster`. This parameter is available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line, and in APISIX from version 3.15.0. * redis\_cluster\_nodes array\[string] *** The list of Redis cluster nodes with at least one address. Required when `policy` is `redis-cluster`. * redis\_cluster\_name string *** The name of the Redis cluster. Required when `policy` is `redis-cluster`. * redis\_cluster\_ssl boolean default: `false` *** If true, use SSL to connect to Redis cluster when `policy` is `redis-cluster`. * redis\_cluster\_ssl\_verify boolean default: `false` *** If true, verify the server SSL certificate when `policy` is `redis-cluster`. --- # loki-logger The `loki-logger` plugin pushes request and response logs in batches to [Grafana Loki](https://grafana.com/oss/loki/), via the [Loki HTTP API](https://grafana.com/docs/loki/latest/reference/loki-http-api/#loki-http-api) `/loki/api/v1/push`. When enabled, the plugin will serialize the request context information to [JSON objects](https://grafana.com/docs/loki/latest/api/#push-log-entries-to-loki) and add them to the queue, before they are pushed to Loki. The plugin also supports the customization of log formats. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `loki-logger` plugin for different scenarios. To follow along the examples, start a sample Loki instance: * Docker * Kubernetes ``` wget https://raw.githubusercontent.com/grafana/loki/v3.0.0/cmd/loki/loki-local-config.yaml -O loki-config.yaml docker run --name loki -d -v $(pwd):/mnt/config -p 3100:3100 grafana/loki:3.2.1 -config.file=/mnt/config/loki-config.yaml ``` Additionally, start a Grafana instance to view and visualize the logs: ``` docker run -d --name=apisix-quickstart-grafana \ -p 3000:3000 \ grafana/grafana-oss ``` To connect Loki and Grafana, visit Grafana at [`http://localhost:3000`](http://localhost:3000). Under **Connections > Data sources**, add a new data source and select Loki. Your connection URL should follow the format of `http://{your_ip_address}:3100`. When saving the new data source, Grafana should also test the connection, and you are expected to see Grafana notifying the data source is successfully connected. Create a Kubernetes manifest file for the Loki deployment: loki-deployment.yaml ``` apiVersion: v1 kind: ConfigMap metadata: namespace: aic name: loki-config data: loki-config.yaml: | auth_enabled: false server: http_listen_port: 3100 grpc_listen_port: 9096 common: instance_addr: 127.0.0.1 path_prefix: /tmp/loki storage: filesystem: chunks_directory: /tmp/loki/chunks rules_directory: /tmp/loki/rules replication_factor: 1 ring: kvstore: store: inmemory schema_config: configs: - from: 2020-10-24 store: tsdb object_store: filesystem schema: v13 index: prefix: index_ period: 24h ruler: alertmanager_url: http://localhost:9093 --- apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: loki spec: replicas: 1 selector: matchLabels: app: loki template: metadata: labels: app: loki spec: containers: - name: loki image: grafana/loki:3.2.1 args: - -config.file=/mnt/config/loki-config.yaml ports: - containerPort: 3100 volumeMounts: - name: config mountPath: /mnt/config volumes: - name: config configMap: name: loki-config --- apiVersion: v1 kind: Service metadata: namespace: aic name: loki spec: selector: app: loki ports: - port: 3100 targetPort: 3100 type: ClusterIP ``` Apply the manifest to your cluster: ``` kubectl apply -f loki-deployment.yaml ``` ### Log Requests and Responses in Default Log Format[​](#log-requests-and-responses-in-default-log-format "Direct link to Log Requests and Responses in Default Log Format") The following example demonstrates how you can configure the `loki-logger` plugin on a route to log requests and responses going through the route. Create a route with the `loki-logger` plugin and configure the address of Loki: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "loki-logger-route", "uri": "/anything", "plugins": { "loki-logger": { "endpoint_addrs": ["http://192.168.1.5:3100"] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: loki-logger-route plugins: loki-logger: endpoint_addrs: - "http://192.168.1.5:3100" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD loki-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: loki-logger-plugin-config spec: plugins: - name: loki-logger config: endpoint_addrs: - "http://loki.aic.svc:3100" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: loki-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: loki-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` loki-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: loki-logger-route spec: ingressClassName: apisix http: - name: loki-logger-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: loki-logger config: endpoint_addrs: - "http://loki.aic.svc:3100" ``` Apply the configuration: ``` kubectl apply -f loki-logger-ic.yaml ``` ❶ Replace with your Loki address. For Kubernetes, use the in-cluster service address `http://loki.aic.svc:3100`. Send a few requests to the route to generate log entries: ``` curl "http://127.0.0.1:9080/anything" ``` You should receive `HTTP/1.1 200 OK` responses for all requests. Navigate to the [Grafana explore view](http://localhost:3000/explore) and run a query `job = apisix`. You should see a number of logs corresponding to your requests, such as the following: ``` { "route_id": "loki-logger-route", "response": { "status": 200, "headers": { "date": "Fri, 03 Jan 2025 03:54:26 GMT", "server": "APISIX/3.13.0", "access-control-allow-credentials": "true", "content-length": "391", "access-control-allow-origin": "*", "content-type": "application/json", "connection": "close" }, "size": 619 }, "start_time": 1735876466, "client_ip": "192.168.65.1", "service_id": "", "apisix_latency": 5.0000038146973, "upstream": "34.197.122.172:80", "upstream_latency": 666, "server": { "hostname": "0b9a772e68f8", "version": "3.13.0" }, "request": { "headers": { "user-agent": "curl/8.6.0", "accept": "*/*", "host": "127.0.0.1:9080" }, "size": 85, "method": "GET", "url": "http://127.0.0.1:9080/anything", "querystring": {}, "uri": "/anything" }, "latency": 671.0000038147 } ``` This verifies that Loki has been receiving logs from APISIX. You may also create dashboards in Grafana to further visualize and analyze the logs. ### Customize Log Format with Plugin Metadata[​](#customize-log-format-with-plugin-metadata "Direct link to Customize Log Format with Plugin Metadata") The following example demonstrates how you can customize log format using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md). Create a route with the `loki-logger` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "loki-logger-route", "uri": "/anything", "plugins": { "loki-logger": { "endpoint_addrs": ["http://192.168.1.5:3100"] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: loki-logger-route plugins: loki-logger: endpoint_addrs: - "http://192.168.1.5:3100" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD loki-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: loki-logger-plugin-config spec: plugins: - name: loki-logger config: endpoint_addrs: - "http://loki.aic.svc:3100" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: loki-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: loki-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` loki-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: loki-logger-route spec: ingressClassName: apisix http: - name: loki-logger-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: loki-logger config: endpoint_addrs: - "http://loki.aic.svc:3100" ``` Apply the configuration: ``` kubectl apply -f loki-logger-ic.yaml ``` Configure plugin metadata for `loki-logger`, which will update the log format for all routes of which requests would be logged: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/loki-logger" -X PUT \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "log_format": { "host": "$host", "client_ip": "$remote_addr", "route_id": "$route_id", "@timestamp": "$time_iso8601" } }' ``` adc.yaml ``` plugin_metadata: - name: loki-logger log_format: host: "$host" client_ip: "$remote_addr" route_id: "$route_id" "@timestamp": "$time_iso8601" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: loki-logger: log_format: host: "$host" client_ip: "$remote_addr" route_id: "$route_id" "@timestamp": "$time_iso8601" ``` Apply the configuration: ``` kubectl apply -f gatewayproxy.yaml ``` Send a request to the route to generate a new log entry: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to the [Grafana explore view](http://localhost:3000/explore) and run a query `job = apisix`. You should see a log entry corresponding to your request, similar to the following: ``` { "@timestamp":"2025-01-03T21:11:34+00:00", "client_ip":"192.168.65.1", "route_id":"loki-logger-route", "host":"127.0.0.1" } ``` If the plugin on a route specifies a specific log format, it will take precedence over the log format specified in the plugin metadata. For instance, update the plugin on the previous route as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes/loki-logger-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "loki-logger": { "log_format": { "route_id": "$route_id", "client_ip": "$remote_addr", "@timestamp": "$time_iso8601" } } } }' ``` Update `adc.yaml` to add a per-route `log_format` to the `loki-logger` plugin: adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: loki-logger-route plugins: loki-logger: endpoint_addrs: - "http://192.168.1.5:3100" log_format: route_id: "$route_id" client_ip: "$remote_addr" "@timestamp": "$time_iso8601" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update `loki-logger-ic.yaml` to add a per-route `log_format` to the `PluginConfig`: loki-logger-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: loki-logger-plugin-config spec: plugins: - name: loki-logger config: endpoint_addrs: - "http://loki.aic.svc:3100" log_format: route_id: "$route_id" client_ip: "$remote_addr" "@timestamp": "$time_iso8601" ``` Update `loki-logger-ic.yaml` to add a per-route `log_format` to the `ApisixRoute`: loki-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: loki-logger-route spec: ingressClassName: apisix http: - name: loki-logger-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: loki-logger config: endpoint_addrs: - "http://loki.aic.svc:3100" log_format: route_id: "$route_id" client_ip: "$remote_addr" "@timestamp": "$time_iso8601" ``` Apply the updated configuration: ``` kubectl apply -f loki-logger-ic.yaml ``` Send a request to the route to generate a new log entry: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to the [Grafana explore view](http://localhost:3000/explore) and re-run the query `job = apisix`. You should see a log entry corresponding to your request, consistent with the format configured on the route, similar to the following: ``` { "client_ip":"192.168.65.1", "route_id":"loki-logger-route", "@timestamp":"2025-01-03T21:19:45+00:00" } ``` ### Log Request Bodies Conditionally[​](#log-request-bodies-conditionally "Direct link to Log Request Bodies Conditionally") The following example demonstrates how you can conditionally log request body. Create a route with `loki-logger` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "loki-logger-route", "uri": "/anything", "plugins": { "loki-logger": { "endpoint_addrs": ["http://192.168.1.5:3100"], "include_req_body": true, "include_req_body_expr": [["arg_log_body", "==", "yes"]] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: loki-logger-route plugins: loki-logger: endpoint_addrs: - "http://192.168.1.5:3100" include_req_body: true include_req_body_expr: - - "arg_log_body" - "==" - "yes" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD loki-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: loki-logger-plugin-config spec: plugins: - name: loki-logger config: endpoint_addrs: - "http://loki.aic.svc:3100" include_req_body: true include_req_body_expr: - - "arg_log_body" - "==" - "yes" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: loki-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: loki-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` loki-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: loki-logger-route spec: ingressClassName: apisix http: - name: loki-logger-route match: paths: - /anything methods: - GET - POST upstreams: - name: httpbin-external-domain plugins: - name: loki-logger config: endpoint_addrs: - "http://loki.aic.svc:3100" include_req_body: true include_req_body_expr: - - "arg_log_body" - "==" - "yes" ``` Apply the configuration: ``` kubectl apply -f loki-logger-ic.yaml ``` ❶ `include_req_body`: set to true to include request body. ❷ `include_req_body_expr`: only include request body if the URL query string `log_body` is `yes`. Send a request to the route with a URL query string satisfying the condition: ``` curl -i "http://127.0.0.1:9080/anything?log_body=yes" -X POST -d '{"env": "dev"}' ``` Navigate to the [Grafana explore view](http://localhost:3000/explore) and run the query `job = apisix`. You should see a log entry corresponding to your request, where the request body is logged: ``` { "route_id": "loki-logger-route", ..., "request": { "headers": { ... }, "body": "{\"env\": \"dev\"}", "size": 182, "method": "POST", "url": "http://127.0.0.1:9080/anything?log_body=yes", "querystring": { "log_body": "yes" }, "uri": "/anything?log_body=yes" }, "latency": 809.99994277954 } ``` Send a request to the route without any URL query string: ``` curl -i "http://127.0.0.1:9080/anything" -X POST -d '{"env": "dev"}' ``` Navigate to the [Grafana explore view](http://localhost:3000/explore) and run the query `job = apisix`. You should see a log entry corresponding to your request, where the request body is not logged: ``` { "route_id": "loki-logger-route", ..., "request": { "headers": { ... }, "size": 169, "method": "POST", "url": "http://127.0.0.1:9080/anything", "querystring": {}, "uri": "/anything" }, "latency": 557.00016021729 } ``` info If you have customized the `log_format` in addition to setting `include_req_body` or `include_resp_body` to `true`, the plugin would not include the bodies in the logs. As a workaround, you may be able to use the NGINX variable `$request_body` in the log format, such as: ``` { "loki-logger": { ..., "log_format": {"body": "$request_body"} } } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * endpoint\_addrs array\[string] required *** Loki API base URLs, such as `http://127.0.0.1:3100`. If multiple endpoints are configured, the log will be pushed to a randomly determined endpoint from the list. * endpoint\_uri string default: `/loki/api/v1/push` *** URI path to the Loki ingest endpoint. * tenant\_id string default: `fake` *** Loki tenant ID. According to Loki's [multi-tenancy documentation](https://grafana.com/docs/loki/latest/operations/multi-tenancy/#multi-tenancy), the default value is set to `fake` under single-tenancy. * headers object *** Key-value pairs of request headers, such as the `Authorization` header for non-local Loki services. The parameter cannot set `X-Scope-OrgID` or `Content-Type`. Introduced in API7 Enterprise 3.9.0. Header-value encryption was introduced in API7 Enterprise 3.9.18 and 3.10.5, and APISIX 3.18.0. The values are [encrypted before storage](https://docs.api7.ai/apisix/production/security/data-encryption-with-keyring.md) when data encryption is enabled. Authorized Admin API `GET` requests return the complete decrypted header object; encryption at rest does not redact header names or values from API responses. * log\_labels object default: `{job = "apisix"}` *** Loki log label. Support [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) and constant strings in values. Variables should be prefixed with a `$` sign. For example, the label can be `{"origin" = "apisix"}` or `{"origin" = "$remote_addr"}`. * ssl\_verify boolean default: `false` *** If true, verify Loki's SSL certificates. * timeout integer default: `3000` vaild vaule: between 1 and 60000 inclusive *** Timeout for the Loki service HTTP call in milliseconds. * keepalive boolean default: `true` *** If true, keep the connection alive for multiple requests. * keepalive\_timeout integer default: `60000` vaild vaule: greater than or equal to 1000 *** Keepalive timeout in milliseconds. * keepalive\_pool integer default: `5` vaild vaule: greater than or equal to 1 *** Maximum number of connections in the connection pool. * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. You can also configure log format on a global scale using the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md), which configures the log format for all `loki-logger` plugin instances. If the log format configured on the individual plugin instance differs from the log format configured on plugin metadata, the log format configured on the individual plugin instance takes precedence. See this [example](https://docs.api7.ai/hub/loki-logger.md#customize-log-format-with-plugin-metadata) for more details. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * name string default: `loki logger` *** Unique identifier of the plugin for the batch processor. If you use [Prometheus](https://docs.api7.ai/hub/prometheus.md) to monitor APISIX metrics, the name is exported in `apisix_batch_process_entries`. * include\_req\_body boolean default: `false` *** If true, include the request body in the log. Note that if the request body is too big to be kept in the memory, it can not be logged due to NGINX's limitations. * include\_req\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_req_body` is true. Request body would only be logged when the expressions configured here evaluate to true. * include\_resp\_body boolean default: `false` *** If true, include the response body in the log. * include\_resp\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_resp_body` is true. Response body would only be logged when the expressions configured here evaluate to true. * max\_req\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes to include in the log. If the request body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * max\_resp\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum response body size in bytes to include in the log. If the response body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * batch\_max\_size integer default: `1000` vaild vaule: greater than 0 *** The number of log entries allowed in one batch. Once reached, the batch will be sent to the logging service. Setting this parameter to 1 means immediate processing. * inactive\_timeout integer default: `5` vaild vaule: greater than 0 *** The maximum time in seconds to wait for new logs before sending the batch to the logging service. The value should be smaller than `buffer_duration`. * buffer\_duration integer default: `60` vaild vaule: greater than 0 *** The maximum time in seconds from the earliest entry allowed before sending the batch to the logging service. * retry\_delay integer default: `1` vaild vaule: greater than or equal to 0 *** The time interval in seconds to retry sending the batch to the logging service if the batch was not successfully sent. * max\_retry\_count integer default: `0` vaild vaule: greater than or equal to 0 *** The maximum number of unsuccessful retries allowed before dropping the log entries. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * max\_pending\_entries integer default: `` `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 and 3.10.6; none in API7 Enterprise 3.9.18 and 3.10.5 `` vaild vaule: greater than or equal to 1 *** Maximum number of entries waiting in the batch processor. New entries are discarded when the backlog reaches the limit. Introduced in API7 Enterprise 3.8.17 and APISIX 3.15.0. The default changed to `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line. In API7 Enterprise 3.9.18 and 3.10.5, and in earlier APISIX versions, omitting the parameter leaves the backlog unlimited. See [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md#configure-the-pending-entry-limit) for sizing and verification guidance. --- # mcp-tools-acl [Enterprise](https://api7.ai/enterprise) The `mcp-tools-acl` plugin controls which MCP tools each consumer can call or discover on a route powered by [`openapi-to-mcp`](https://docs.api7.ai/hub/openapi-to-mcp.md). It enforces two kinds of restrictions: * **`tools/call` blocking** — rejects calls to denied tools with an HTTP error response. * **`tools/list` filtering** — strips denied tools from the list returned to the client, so disallowed tools are invisible to MCP clients. This applies to both JSON responses and SSE (Server-Sent Events) streaming responses. The plugin uses a **rules-based** configuration where each rule specifies either an allowlist or a denylist, and optionally an expression condition. Rules are evaluated in order — the first rule whose conditions match is applied and the rest are skipped. This plugin is available in API7 Enterprise from version 3.9.8. ## Examples[​](#examples "Direct link to Examples") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before using this plugin, ensure that: 1. The [`openapi-to-mcp`](https://docs.api7.ai/hub/openapi-to-mcp.md) plugin is enabled on the same route. 2. An authentication plugin (e.g., [`key-auth`](https://docs.api7.ai/hub/key-auth.md)) is configured on the route. Without an authenticated consumer on the request, `mcp-tools-acl` passes all traffic unchanged. The examples below demonstrate how you can use the `mcp-tools-acl` plugin for different scenarios. ### Restrict Tool Access with an Allowlist[​](#restrict-tool-access-with-an-allowlist "Direct link to Restrict Tool Access with an Allowlist") The following example demonstrates how to configure an allowlist so that a consumer can only call specific MCP tools. * Admin API * ADC * Ingress Controller Create a consumer `alice`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "alice" }' ``` Create a `key-auth` credential for `alice`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/alice/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-alice-key-auth", "plugins": { "key-auth": { "key": "alice-key" } } }' ``` Configure the `mcp-tools-acl` plugin on consumer `alice`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "alice", "plugins": { "mcp-tools-acl": { "rules": [ { "allow_tools": ["getPetById", "getUserByName"] } ] } } }' ``` ❶ Allow consumer `alice` to call only `getPetById` and `getUserByName`. All other tools are blocked and will not appear in the tool listing. tip Configure `mcp-tools-acl` on the **consumer** (or **consumer group**) to enable per-consumer tool access control. The route only needs `openapi-to-mcp` and an auth plugin. When the plugin is configured on both the consumer and the route, the **consumer configuration takes priority** and the route-level configuration is ignored for that consumer. If a consumer has no `mcp-tools-acl` configuration, the route-level configuration applies as a fallback. Create a route with `openapi-to-mcp` and `key-auth` enabled: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "mcp-tools-acl-route", "uri": "/mcp", "methods": ["GET", "POST"], "plugins": { "openapi-to-mcp": { "transport": "streamable_http", "openapi_url": "https://petstore3.swagger.io/api/v3/openapi.json", "base_url": "https://petstore3.swagger.io/api/v3", "headers": { "Authorization": "special-key" } }, "key-auth": {} }, "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "petstore3.swagger.io:443": 1 } } }' ``` Send a `tools/list` request as consumer `alice`: ``` curl -s "http://127.0.0.1:9080/mcp" \ -H "apikey: alice-key" \ -H "Accept: application/json, text/event-stream" \ -H "Content-Type: application/json" \ -d '{ "jsonrpc": "2.0", "id": 1, "method": "tools/list" }' ``` You should see only `getPetById` and `getUserByName` in the response — all other tools have been filtered out. Call an allowed tool: ``` curl -i "http://127.0.0.1:9080/mcp" \ -H "apikey: alice-key" \ -H "Accept: application/json, text/event-stream" \ -H "Content-Type: application/json" \ -d '{ "jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": { "name": "getPetById", "arguments": { "pathParameters": { "petId": 1 } } } }' ``` You should see an `HTTP/1.1 200 OK` response with the Petstore pet payload. Call a tool that is not on the allowlist: ``` curl -i "http://127.0.0.1:9080/mcp" \ -H "apikey: alice-key" \ -H "Accept: application/json, text/event-stream" \ -H "Content-Type: application/json" \ -d '{ "jsonrpc": "2.0", "id": 3, "method": "tools/call", "params": { "name": "deletePet", "arguments": { "pathParameters": { "petId": 1 } } } }' ``` You should see an `HTTP/1.1 403 Forbidden` response, indicating the tool call was rejected. Create a consumer with `key-auth` credential and `mcp-tools-acl` configured, and create a route with `openapi-to-mcp` and `key-auth` enabled: adc.yaml ``` consumers: - username: alice credentials: - name: cred-alice-key-auth type: key-auth config: key: alice-key plugins: mcp-tools-acl: rules: - allow_tools: - getPetById - getUserByName services: - name: mcp-tools-acl-service upstream: type: roundrobin scheme: https pass_host: node nodes: - host: petstore3.swagger.io port: 443 weight: 1 routes: - name: mcp-tools-acl-route uris: - /mcp methods: - GET - POST plugins: key-auth: header: apikey openapi-to-mcp: transport: streamable_http base_url: https://petstore3.swagger.io/api/v3 headers: Authorization: special-key openapi_url: https://petstore3.swagger.io/api/v3/openapi.json ``` ❶ Allow consumer `alice` to call only `getPetById` and `getUserByName`. All other tools are blocked and will not appear in the tool listing. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create a consumer and a route with `openapi-to-mcp`, `key-auth`, and `mcp-tools-acl` configured as such: mcp-tools-acl-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: alice spec: gatewayRef: name: apisix credentials: - type: key-auth name: cred-alice-key-auth config: key: alice-key plugins: - name: mcp-tools-acl config: rules: - allow_tools: - getPetById - getUserByName --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: mcp-tools-acl-plugin-config spec: plugins: - name: key-auth config: header: apikey - name: openapi-to-mcp config: transport: streamable_http base_url: https://petstore3.swagger.io/api/v3 headers: Authorization: special-key openapi_url: https://petstore3.swagger.io/api/v3/openapi.json --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: mcp-tools-acl-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /mcp method: GET - path: type: Exact value: /mcp method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: mcp-tools-acl-plugin-config backendRefs: - name: petstore-external-domain port: 443 --- apiVersion: v1 kind: Service metadata: namespace: aic name: petstore-external-domain spec: type: ExternalName externalName: petstore3.swagger.io --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: petstore-external-domain spec: targetRefs: - group: "" kind: Service name: petstore-external-domain passHost: node scheme: https ``` ❶ Allow consumer `alice` to call only `getPetById` and `getUserByName`. Create a consumer and a route with `openapi-to-mcp`, `key-auth`, and `mcp-tools-acl` configured as such: mcp-tools-acl-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: alice spec: ingressClassName: apisix authParameter: keyAuth: value: key: alice-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: petstore-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: petstore3.swagger.io port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: mcp-tools-acl-route spec: ingressClassName: apisix http: - name: mcp-tools-acl-route match: paths: - /mcp methods: - GET - POST upstreams: - name: petstore-external-domain plugins: - name: key-auth enable: true config: header: apikey - name: openapi-to-mcp enable: true config: transport: streamable_http base_url: https://petstore3.swagger.io/api/v3 headers: Authorization: special-key openapi_url: https://petstore3.swagger.io/api/v3/openapi.json - name: mcp-tools-acl enable: true config: rules: - allow_tools: - getPetById - getUserByName ``` ❶ Allow only `getPetById` and `getUserByName` on this route. Apply the configuration to your cluster: ``` kubectl apply -f mcp-tools-acl-ic.yaml ``` ### Restrict Tool Access with a Denylist[​](#restrict-tool-access-with-a-denylist "Direct link to Restrict Tool Access with a Denylist") The following example demonstrates how to block a specific tool while leaving all other tools accessible. * Admin API * ADC * Ingress Controller Building on the previous example, create consumer `bob` with a denylist: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "bob", "plugins": { "mcp-tools-acl": { "rules": [ { "deny_tools": ["deletePet"] } ] } } }' ``` ❶ Deny consumer `bob` from calling `deletePet`. All other tools remain accessible. Create a `key-auth` credential for `bob`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/bob/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-bob-key-auth", "plugins": { "key-auth": { "key": "bob-key" } } }' ``` Send a `tools/list` request as consumer `bob`: ``` curl -s "http://127.0.0.1:9080/mcp" \ -H "apikey: bob-key" \ -H "Accept: application/json, text/event-stream" \ -H "Content-Type: application/json" \ -d '{ "jsonrpc": "2.0", "id": 1, "method": "tools/list" }' ``` You should see all tools listed except `deletePet`. Call the denied tool: ``` curl -i "http://127.0.0.1:9080/mcp" \ -H "apikey: bob-key" \ -H "Accept: application/json, text/event-stream" \ -H "Content-Type: application/json" \ -d '{ "jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": { "name": "deletePet", "arguments": { "pathParameters": { "petId": 1 } } } }' ``` You should see an `HTTP/1.1 403 Forbidden` response. Update the consumer configuration in `adc.yaml` to change the ACL rule: adc.yaml ``` consumers: - username: bob credentials: - name: cred-bob-key-auth type: key-auth config: key: bob-key plugins: mcp-tools-acl: rules: - deny_tools: - deletePet services: - name: mcp-tools-acl-service upstream: type: roundrobin scheme: https pass_host: node nodes: - host: petstore3.swagger.io port: 443 weight: 1 routes: - name: mcp-tools-acl-route uris: - /mcp methods: - GET - POST plugins: key-auth: header: apikey openapi-to-mcp: transport: streamable_http base_url: https://petstore3.swagger.io/api/v3 headers: Authorization: special-key openapi_url: https://petstore3.swagger.io/api/v3/openapi.json ``` ❶ Deny consumer `bob` from calling `deletePet`. All other tools remain accessible. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update the ACL rule in your consumer or route configuration: mcp-tools-acl-ic.yaml ``` # Other Configs # --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: bob spec: gatewayRef: name: apisix credentials: - type: key-auth name: cred-bob-key-auth config: key: bob-key plugins: - name: mcp-tools-acl config: rules: - deny_tools: - deletePet ``` ❶ Deny consumer `bob` from calling `deletePet`. Update the ACL rule in your consumer or route configuration: mcp-tools-acl-ic.yaml ``` # Other Configs # --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: mcp-tools-acl-route spec: ingressClassName: apisix http: - name: mcp-tools-acl-route match: paths: - /mcp methods: - GET - POST upstreams: - name: petstore-external-domain plugins: - name: key-auth enable: true config: header: apikey - name: openapi-to-mcp enable: true config: transport: streamable_http base_url: https://petstore3.swagger.io/api/v3 headers: Authorization: special-key openapi_url: https://petstore3.swagger.io/api/v3/openapi.json - name: mcp-tools-acl enable: true config: rules: - deny_tools: - deletePet ``` ❶ Deny calls to `deletePet` while allowing other tools on the route. Apply the configuration to your cluster: ``` kubectl apply -f mcp-tools-acl-ic.yaml ``` ### Apply Different Rules Based on Route Conditions[​](#apply-different-rules-based-on-route-conditions "Direct link to Apply Different Rules Based on Route Conditions") The following example demonstrates how to use expression conditions (`expr`) to apply different ACL rules depending on request context, such as the route being accessed. * Admin API * ADC * Ingress Controller This example applies to API7 Enterprise version 3.9.8 and later. Create a consumer `grace` with conditional rules: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "grace", "plugins": { "mcp-tools-acl": { "rules": [ { "expr": [["route_id", "==", "route-pets"]], "allow_tools": ["getPetById"] }, { "allow_tools": ["getUserByName"] } ] } } }' ``` ❶ Rule 1: When the request hits the route `route-pets`, only `getPetById` is allowed. ❷ Rule 2: A catch-all rule (no `expr`) — for all other routes, only `getUserByName` is allowed. This rule is reached only if Rule 1 does not match. Create a `key-auth` credential for `grace`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/grace/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-grace-key-auth", "plugins": { "key-auth": { "key": "grace-key" } } }' ``` Create two routes with `openapi-to-mcp` and `key-auth`: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "route-pets", "uri": "/mcp", "methods": ["GET", "POST"], "plugins": { "openapi-to-mcp": { "transport": "streamable_http", "openapi_url": "https://petstore3.swagger.io/api/v3/openapi.json", "base_url": "https://petstore3.swagger.io/api/v3", "headers": { "Authorization": "special-key" } }, "key-auth": {} }, "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "petstore3.swagger.io:443": 1 } } }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "route-inventory", "uri": "/mcp2", "methods": ["GET", "POST"], "plugins": { "openapi-to-mcp": { "transport": "streamable_http", "openapi_url": "https://petstore3.swagger.io/api/v3/openapi.json", "base_url": "https://petstore3.swagger.io/api/v3", "headers": { "Authorization": "special-key" } }, "key-auth": {} }, "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "petstore3.swagger.io:443": 1 } } }' ``` When `grace` calls `route-pets`, Rule 1 matches and only `getPetById` is allowed: ``` curl -i "http://127.0.0.1:9080/mcp" \ -H "apikey: grace-key" \ -H "Accept: application/json, text/event-stream" \ -H "Content-Type: application/json" \ -d '{ "jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": { "name": "getPetById", "arguments": { "pathParameters": { "petId": 1 } } } }' ``` You should see an `HTTP/1.1 200 OK` response with the Petstore pet payload. When `grace` calls `route-inventory`, Rule 1 does not match (wrong `route_id`), so the catch-all Rule 2 applies — only `getUserByName` is allowed: ``` curl -i "http://127.0.0.1:9080/mcp2" \ -H "apikey: grace-key" \ -H "Accept: application/json, text/event-stream" \ -H "Content-Type: application/json" \ -d '{ "jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": { "name": "getUserByName", "arguments": { "pathParameters": { "username": "user1" } } } }' ``` You should see an `HTTP/1.1 200 OK` response with the Petstore user payload. note Rules are evaluated top to bottom. The first matching rule takes effect and all subsequent rules are skipped. Place more specific rules (with `expr`) before broader catch-all rules (without `expr`). If no rule matches (all rules have `expr` conditions and none evaluate to true), the plugin does not enforce any access control — all tools are passed through. Create a consumer with conditional rules and two routes: adc.yaml ``` consumers: - username: grace credentials: - name: cred-grace-key-auth type: key-auth config: key: grace-key plugins: mcp-tools-acl: rules: - expr: - - route_name - == - pets-route allow_tools: - getPetById - allow_tools: - getUserByName services: - name: pets-mcp-service upstream: type: roundrobin scheme: https pass_host: node nodes: - host: petstore3.swagger.io port: 443 weight: 1 routes: - name: pets-route uris: - /mcp methods: - GET - POST plugins: key-auth: header: apikey openapi-to-mcp: transport: streamable_http base_url: https://petstore3.swagger.io/api/v3 headers: Authorization: special-key openapi_url: https://petstore3.swagger.io/api/v3/openapi.json - name: inventory-mcp-service upstream: type: roundrobin scheme: https pass_host: node nodes: - host: petstore3.swagger.io port: 443 weight: 1 routes: - name: inventory-route uris: - /mcp2 methods: - GET - POST plugins: key-auth: header: apikey openapi-to-mcp: transport: streamable_http base_url: https://petstore3.swagger.io/api/v3 headers: Authorization: special-key openapi_url: https://petstore3.swagger.io/api/v3/openapi.json ``` ❶ When the request hits `pets-route`, only `getPetById` is allowed. ❷ A catch-all rule for all other routes. Here only `getUserByName` is allowed. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Configure conditional rules with separate MCP routes: mcp-tools-acl-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: grace spec: gatewayRef: name: apisix credentials: - type: key-auth name: cred-grace-key-auth config: key: grace-key plugins: - name: mcp-tools-acl config: rules: - expr: - - route_name - == - pets-route allow_tools: - getPetById - allow_tools: - getUserByName --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: mcp-tools-acl-plugin-config spec: plugins: - name: key-auth config: header: apikey - name: openapi-to-mcp config: transport: streamable_http base_url: https://petstore3.swagger.io/api/v3 headers: Authorization: special-key openapi_url: https://petstore3.swagger.io/api/v3/openapi.json --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: pets-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /mcp method: GET - path: type: Exact value: /mcp method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: mcp-tools-acl-plugin-config backendRefs: - name: petstore-external-domain port: 443 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: inventory-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /mcp2 method: GET - path: type: Exact value: /mcp2 method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: mcp-tools-acl-plugin-config backendRefs: - name: petstore-external-domain port: 443 --- apiVersion: v1 kind: Service metadata: namespace: aic name: petstore-external-domain spec: type: ExternalName externalName: petstore3.swagger.io --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: petstore-external-domain spec: targetRefs: - group: "" kind: Service name: petstore-external-domain passHost: node scheme: https ``` ❶ Apply `getPetById` only when the request matches `pets-route`. ❷ Apply `getUserByName` to all other matching routes. Configure conditional rules with separate MCP routes: mcp-tools-acl-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: grace spec: ingressClassName: apisix authParameter: keyAuth: value: key: grace-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: petstore-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: petstore3.swagger.io port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: mcp-tools-acl-routes spec: ingressClassName: apisix http: - name: pets-route match: paths: - /mcp methods: - GET - POST upstreams: - name: petstore-external-domain plugins: - name: key-auth enable: true config: header: apikey - name: openapi-to-mcp enable: true config: transport: streamable_http base_url: https://petstore3.swagger.io/api/v3 headers: Authorization: special-key openapi_url: https://petstore3.swagger.io/api/v3/openapi.json - name: mcp-tools-acl enable: true config: rules: - expr: - - uri - == - /mcp allow_tools: - getPetById - allow_tools: - getUserByName - name: inventory-route match: paths: - /mcp2 methods: - GET - POST upstreams: - name: petstore-external-domain plugins: - name: key-auth enable: true config: header: apikey - name: openapi-to-mcp enable: true config: transport: streamable_http base_url: https://petstore3.swagger.io/api/v3 headers: Authorization: special-key openapi_url: https://petstore3.swagger.io/api/v3/openapi.json - name: mcp-tools-acl enable: true config: rules: - expr: - - uri - == - /mcp allow_tools: - getPetById - allow_tools: - getUserByName ``` ❶ Apply `getPetById` only when the request URI is `/mcp`. ❷ Use the catch-all rule to allow only `getUserByName` for the `/mcp2` route. Apply the configuration to your cluster: ``` kubectl apply -f mcp-tools-acl-ic.yaml ``` ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") **Plugin has no effect** Check that `openapi-to-mcp` is enabled on the same route and that an authentication plugin is configured. Without an authenticated consumer on the request, `mcp-tools-acl` passes all traffic unchanged by design. **`tools/call` returns HTTP 400** The request body is valid JSON but `params` is missing or `params.name` is not a string. The plugin returns `{"message": "Invalid MCP tools/call request"}` with HTTP 400. This is distinct from `rejected_code` (which applies to denied tools) and indicates a malformed request from the MCP client. **`allow_tools: []` blocks all tools** An empty allowlist is schema-valid and denies all tools. Every `tools/call` will be rejected and `tools/list` will return an empty list. If you want consumers to access any tools, ensure the `allow_tools` array is non-empty. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * rules array\[object] required *** An array of access control rules evaluated in order. The first rule whose `expr` conditions are all met (or that has no `expr`) is applied; remaining rules are skipped. Each rule must contain exactly one of `allow_tools` or `deny_tools`. * allow\_tools array\[string] *** Allowlist of MCP tool names the consumer is permitted to call and see in `tools/list`. Matching is exact and case-sensitive. An empty array (`[]`) denies all tools. Exactly one of `allow_tools` or `deny_tools` must be configured per rule; they cannot be used together in the same rule. * deny\_tools array\[string] *** Blocklist of MCP tool names the consumer is not permitted to call. Denied tools are also hidden from `tools/list`. Matching is exact and case-sensitive. Exactly one of `allow_tools` or `deny_tools` must be configured per rule; they cannot be used together in the same rule. * rejected\_code integer default: `403` vaild vaule: 200 to 599 *** HTTP status code returned when a `tools/call` request is rejected by this rule. * rejected\_msg string default: `MCP tool is not allowed` vaild vaule: non-empty string *** Message returned in the response body when a `tools/call` request is rejected by this rule. * expr array *** An array of one or more matching conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). The rule is applied only when all expressions evaluate to true. If omitted, the rule matches unconditionally (catch-all). * max\_resp\_body\_size integer default: `67108864` *** Maximum response body size in bytes buffered into memory for tool filtering. Larger responses are truncated. Available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line. --- # mocking The `mocking` plugin allows you to simulate API responses without forwarding requests to upstream services. The plugin supports the customization of the response status code, body, headers, and more. This is particularly useful during development, testing, or debugging phases, where the actual upstream service might be unavailable, under maintenance, or expensive to call. By providing mock responses in a predefined format, the plugin enables you to test client-side integrations, validate request handling, and debug issues without relying on the upstream infrastructure. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `mocking` plugin for different scenarios. ### Generate Specific Mock Responses[​](#generate-specific-mock-responses "Direct link to Generate Specific Mock Responses") The following example demonstrates how to configure the plugin to generate a specific mock response and response status code without forwarding the request to the upstream service. Create a route using the `mocking` plugin and define a response body for the expected mock responses: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "mocking-route", "uri": "/anything", "plugins": { "mocking": { "response_status":201, "response_example":"{\"Lastname\":\"Brown\",\"Age\":56}" } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: mocking-route uris: - /anything plugins: mocking: response_status: 201 response_example: '{"Lastname":"Brown","Age":56}' upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD mocking-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: mocking-plugin-config spec: plugins: - name: mocking config: response_status: 201 response_example: '{"Lastname":"Brown","Age":56}' --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: mocking-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: mocking-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f mocking-ic.yaml ``` mocking-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: mocking-route spec: ingressClassName: apisix http: - name: mocking-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: mocking enable: true config: response_status: 201 response_example: '{"Lastname":"Brown","Age":56}' ``` Apply the configuration to your cluster: ``` kubectl apply -f mocking-ic.yaml ``` ❶ Configure the expected mock response status code to be `201`. ❷ Configure the expected mock response body to be `{"Lastname":"Brown","Age":56}`. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 201 Created` mock response and see the following response body: ``` {"Lastname":"Brown","Age":56} ``` ### Generate Mock Response Headers[​](#generate-mock-response-headers "Direct link to Generate Mock Response Headers") The following example demonstrates how to configure the plugin to generate mock response headers and use a [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md) in the response body. Create a route using the `mocking` plugin, define response headers, and response body for the expected mock responses: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "mocking-route", "uri": "/anything", "plugins": { "mocking": { "response_headers": { "X-User-Id": 100, "X-Product-Id": "apac-398-472" }, "response_example":"Client IP: $remote_addr" } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: mocking-route uris: - /anything plugins: mocking: response_headers: X-User-Id: 100 X-Product-Id: apac-398-472 response_example: "Client IP: $remote_addr" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD mocking-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: mocking-plugin-config spec: plugins: - name: mocking config: response_headers: X-User-Id: 100 X-Product-Id: apac-398-472 response_example: "Client IP: $remote_addr" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: mocking-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: mocking-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f mocking-ic.yaml ``` mocking-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: mocking-route spec: ingressClassName: apisix http: - name: mocking-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: mocking enable: true config: response_headers: X-User-Id: 100 X-Product-Id: apac-398-472 response_example: "Client IP: $remote_addr" ``` Apply the configuration to your cluster: ``` kubectl apply -f mocking-ic.yaml ``` ❶ Configure the expected mock response header `X-User-Id: 100`. ❷ Configure the expected mock response header `X-Product-Id: apac-398-472`. ❸ Configure the expected mock response body to display the client IP address. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive a response similar to the following: ``` HTTP/1.1 200 OK ... X-Product-Id: apac-398-472 X-User-Id: 100 Client IP: 192.168.65.1 ``` ### Generate Mock Responses using JSON Schema[​](#generate-mock-responses-using-json-schema "Direct link to Generate Mock Responses using JSON Schema") The following example demonstrates how to configure the plugin to generate mock responses following a specific [JSON schema](https://json-schema.org). Create a route using the `mocking` plugin and define a JSON schema for the expected mock responses: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "mocking-route", "uri": "/anything", "plugins": { "mocking": { "response_schema": { "type": "object", "properties": { "id": { "type": "string", "example": "abcd" }, "ip": { "type": "number", "example": 192.168.0.10 }, "random_str_arr": { "type": "array", "items": { "type": "string" } }, "nested_obj": { "type": "object", "properties": { "random_str": { "type": "string" }, "child_nested_obj": { "type": "object", "properties": { "random_bool": { "type": "boolean", "example": true }, "random_int_arr": { "type": "array", "items": { "type": "integer", "example": 155 } } } } } } } } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: mocking-route uris: - /anything plugins: mocking: response_schema: type: object properties: id: type: string example: abcd ip: type: number example: 192.168.0.10 random_str_arr: type: array items: type: string nested_obj: type: object properties: random_str: type: string child_nested_obj: type: object properties: random_bool: type: boolean example: true random_int_arr: type: array items: type: integer example: 155 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD mocking-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: mocking-plugin-config spec: plugins: - name: mocking config: response_schema: type: object properties: id: type: string example: abcd ip: type: number example: 192.168.0.10 random_str_arr: type: array items: type: string nested_obj: type: object properties: random_str: type: string child_nested_obj: type: object properties: random_bool: type: boolean example: true random_int_arr: type: array items: type: integer example: 155 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: mocking-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: mocking-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f mocking-ic.yaml ``` mocking-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: mocking-route spec: ingressClassName: apisix http: - name: mocking-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: mocking enable: true config: response_schema: type: object properties: id: type: string example: abcd ip: type: number example: 192.168.0.10 random_str_arr: type: array items: type: string nested_obj: type: object properties: random_str: type: string child_nested_obj: type: object properties: random_bool: type: boolean example: true random_int_arr: type: array items: type: integer example: 155 ``` Apply the configuration to your cluster: ``` kubectl apply -f mocking-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should see a mock response similar to the following, without the actual response from the upstream service: ``` { "ip":192.168.0.10, "random_str_arr":[ "fb","lyquibkwc","r" ], "id":"abcd", "nested_obj":{ "random_str":"bzbb", "child_nested_obj":{ "random_bool":true, "random_int_arr":[155,155,155] } } } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * delay integer default: `0` *** Mock response delay in seconds. * response\_status integer default: `200` *** HTTP status code of the mock response. * content\_type string default: `application/json;charset=utf8` *** `Content-Type` header of the mock response. * response\_example string *** Response body of the mock response. Support the use of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) in the body. At least one of the `response_example` and `response_schema` should be configured. * response\_schema object *** Response [JSON Schema](https://json-schema.org). The JSON schema supports `string`, `number`, `integer`, `boolean`, `object`, and `array` in data types. At least one of the `response_example` and `response_schema` should be configured. `response_schema` is only effective when `response_example` is not configured. * with\_mock\_header boolean default: `true` *** If true, add a response header `x-mock-by` with APISIX version. * response\_headers object *** Headers to be added in the mock response. --- # mqtt-proxy The `mqtt-proxy` plugin is an L4 plugin that supports proxying and load balancing MQTT requests to MQTT servers. It supports MQTT versions 3.1.x and 5.0. The plugin must be configured on a [stream route](https://docs.api7.ai/apisix/key-concepts/stream-routes.md), and APISIX should enable L4 traffic proxying. ## Examples[​](#examples "Direct link to Examples") By default, APISIX only proxies L7 traffic. Enable L4 traffic proxying before continuing with the examples. * Host or Docker * Kubernetes (Helm) Add or update this section in the gateway configuration file: config.yaml ``` apisix: proxy_mode: http&stream # Enable both L4 & L7 proxies stream_proxy: # Configure L4 proxy tcp: - 9100 # Set TCP proxy listening port ``` Reload the gateway for changes to take effect. The gateway should now start listening for L4 traffic on port `9100`. For Helm deployments, update the chart values that render stream proxy listeners. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` service: stream: enabled: true tcp: - 9100 ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` gateway: stream: enabled: true tcp: - addr: 9100 ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` The examples below use an MQTT client from the Mosquitto project to publish and subscribe messages. You can download it [here](https://mosquitto.org/download/) or use any other MQTT client of your choice. ### Proxy to a MQTT Broker[​](#proxy-to-a-mqtt-broker "Direct link to Proxy to a MQTT Broker") The following example demonstrates how you can configure a stream route to proxy traffic to a hosted MQTT server and verify the APISIX can proxy MQTT messages successfully. Create a stream route to the MQTT server and configure the `mqtt-proxy` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/stream_routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "mqtt-route", "plugins": { "mqtt-proxy": { "protocol_name": "MQTT", "protocol_level": 4 } }, "upstream": { "type": "roundrobin", "nodes": { "test.mosquitto.org:1883": 1 } } }' ``` adc.yaml ``` services: - name: mqtt-service upstream: name: default scheme: tcp nodes: - host: test.mosquitto.org port: 1883 weight: 1 stream_routes: - name: mqtt-route server_port: 9100 plugins: mqtt-proxy: protocol_name: MQTT protocol_level: 4 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD info Attaching L4 plugins is currently not supported with Gateway API. At the moment, this example cannot be completed with Gateway API. Use APISIX CRD to attach the `mqtt-proxy` plugin to the stream route: mqtt-proxy-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: mqtt-broker spec: type: ExternalName externalName: test.mosquitto.org ports: - name: mqtt port: 1883 targetPort: 1883 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: mqtt-route spec: ingressClassName: apisix stream: - name: mqtt-route protocol: TCP match: ingressPort: 9100 backend: serviceName: mqtt-broker servicePort: 1883 plugins: - name: mqtt-proxy enable: true config: protocol_name: MQTT protocol_level: 4 ``` Apply the configuration: ``` kubectl apply -f mqtt-proxy-ic.yaml ``` Open two terminal sessions. In the first one, subscribe to the test topic: ``` mosquitto_sub -h test.mosquitto.org -p 1883 -t "test/apisix" ``` In the other one, publish a sample message to the created route: ``` mosquitto_pub -h 127.0.0.1 -p 9100 -t "test/apisix" -m "Hello APISIX" ``` You should see the message `Hello APISIX` in the first terminal. ### Load Balance MQTT Traffic[​](#load-balance-mqtt-traffic "Direct link to Load Balance MQTT Traffic") The following example demonstrates how you can configure a stream route to load balance MQTT traffic to different MQTT servers. When the plugin is enabled, it registers a variable `mqtt_client_id` which can be used for load balancing. MQTT connections with different client ID will be forwarded to different upstream nodes based on the consistent hash algorithm. If the client ID is missing, client IP will be used instead. Create a stream route to two MQTT servers and configure the `mqtt-proxy` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/stream_routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "mqtt-route", "plugins": { "mqtt-proxy": { "protocol_name": "MQTT", "protocol_level": 4 } }, "upstream": { "type": "chash", "key": "mqtt_client_id", "nodes": [ { "host": "test.mosquitto.org", "port": 1883, "weight": 1 }, { "host": "broker.mqtt.cool", "port": 1883, "weight": 1 } ] } }' ``` adc.yaml ``` services: - name: mqtt-service upstream: name: default scheme: tcp type: chash key: mqtt_client_id nodes: - host: test.mosquitto.org port: 1883 weight: 1 - host: broker.mqtt.cool port: 1883 weight: 1 stream_routes: - name: mqtt-route server_port: 9100 plugins: mqtt-proxy: protocol_name: MQTT protocol_level: 4 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD info Attaching L4 plugins is currently not supported with Gateway API. At the moment, this example cannot be completed with Gateway API. mqtt-proxy-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: mqtt-brokers spec: ports: - name: mqtt port: 1883 protocol: TCP --- apiVersion: discovery.k8s.io/v1 kind: EndpointSlice metadata: namespace: aic name: mqtt-brokers-1 labels: kubernetes.io/service-name: mqtt-brokers addressType: FQDN ports: - name: mqtt protocol: TCP port: 1883 endpoints: - addresses: - test.mosquitto.org - addresses: - broker.mqtt.cool --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: mqtt-brokers spec: ingressClassName: apisix loadbalancer: type: chash key: mqtt_client_id hashOn: vars --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: mqtt-route spec: ingressClassName: apisix stream: - name: mqtt-route protocol: TCP match: ingressPort: 9100 backend: serviceName: mqtt-brokers servicePort: 1883 plugins: - name: mqtt-proxy enable: true config: protocol_name: MQTT protocol_level: 4 ``` Apply the configuration: ``` kubectl apply -f mqtt-proxy-ic.yaml ``` For the Admin API and ADC examples, open three terminal sessions. In the first one, subscribe to the test topic in the first MQTT broker: ``` mosquitto_sub -h test.mosquitto.org -p 1883 -t "test/apisix" ``` In the second terminal, subscribe to the same topic in the second MQTT broker: ``` mosquitto_sub -h broker.mqtt.cool -p 1883 -t "test/apisix" ``` In the third terminal, publish messages with different MQTT client IDs to the created route: ``` mosquitto_pub -h 127.0.0.1 -p 9100 -i publisher-1 -t "test/apisix" -m "Hello from publisher-1" mosquitto_pub -h 127.0.0.1 -p 9100 -i publisher-2 -t "test/apisix" -m "Hello from publisher-2" ``` You should see the published messages in the subscriber terminals, verifying that different `mqtt_client_id` values can be steered to different upstream brokers. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * protocol\_name string default: `MQTT` *** Protocol name, generally `MQTT`. In API7 Gateway, this parameter is required and does not have a default. * protocol\_level integer required *** Protocol version. It should be set to `4` for MQTT 3.1.x and `5` for MQTT 5.0. --- # multi-auth The `multi-auth` plugin allows consumers using different authentication methods to share the same route or service. It supports the configuration of multiple authentication plugins, so that a request would be allowed through if it authenticates successfully against any configured authentication method. When every configured method fails, the plugin returns `401 Unauthorized` and writes one warning for each failed method. Each warning identifies the authentication plugin, its status code, and its returned error message, which helps distinguish an invalid key from another authentication failure. This diagnostic behavior is available in API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0. ## Examples[​](#examples "Direct link to Examples") ### Allow Different Authentications on the Same Route[​](#allow-different-authentications-on-the-same-route "Direct link to Allow Different Authentications on the Same Route") The following example demonstrates how to have one consumer using basic authentication, while another consumer using key authentication, both sharing the same route. * Admin API * ADC * Ingress Controller Create two consumers: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username":"consumer1" }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username":"consumer2" }' ``` Configure basic authentication credential for `consumer1`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/consumer1/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-key-auth", "plugins": { "basic-auth": { "username":"consumer1", "password":"consumer1_pwd" } } }' ``` Configure key authentication credential for `consumer2`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/consumer2/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-key-auth", "plugins": { "key-auth": { "key":"consumer2_pwd" } } }' ``` Create a route with `multi-auth` and configure the two authentication plugins that consumers use: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "multi-auth-route", "uri": "/anything", "plugins": { "multi-auth":{ "auth_plugins":[ { "basic-auth":{} }, { "key-auth":{ "hide_credentials":true, "header":"apikey", "query":"apikey" } } ] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` Create two consumers with their respective credentials and a route with `multi-auth`: adc.yaml ``` consumers: - username: consumer1 credentials: - name: cred-consumer1-basic-auth type: basic-auth config: username: consumer1 password: consumer1_pwd - username: consumer2 credentials: - name: cred-consumer2-key-auth type: key-auth config: key: consumer2_pwd services: - name: multi-auth-service routes: - name: multi-auth-route uris: - /anything plugins: multi-auth: auth_plugins: - basic-auth: {} - key-auth: hide_credentials: true header: apikey query: apikey upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD multi-auth-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: consumer1 spec: gatewayRef: name: apisix credentials: - type: basic-auth name: cred-consumer1-basic-auth config: username: consumer1 password: consumer1_pwd --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: consumer2 spec: gatewayRef: name: apisix credentials: - type: key-auth name: cred-consumer2-key-auth config: key: consumer2_pwd --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: multi-auth-plugin-config spec: plugins: - name: multi-auth config: auth_plugins: - basic-auth: {} - key-auth: hide_credentials: true header: apikey query: apikey --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: multi-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: multi-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f multi-auth-ic.yaml ``` multi-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: consumer1 spec: ingressClassName: apisix authParameter: basicAuth: value: username: consumer1 password: consumer1_pwd --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: consumer2 spec: ingressClassName: apisix authParameter: keyAuth: value: key: consumer2_pwd --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: multi-auth-route spec: ingressClassName: apisix http: - name: multi-auth-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: multi-auth enable: true config: auth_plugins: - basic-auth: {} - key-auth: hide_credentials: true header: apikey query: apikey ``` Apply the configuration to your cluster: ``` kubectl apply -f multi-auth-ic.yaml ``` Send a request to the route with `consumer1` basic authentication credentials: ``` curl -i "http://127.0.0.1:9080/anything" -u consumer1:consumer1_pwd ``` You should receive an `HTTP/1.1 200 OK` response. Send another request to the route with `consumer2` key authentication credential: ``` curl -i "http://127.0.0.1:9080/anything" -H 'apikey: consumer2_pwd' ``` You should again receive an `HTTP/1.1 200 OK` response. Send a request to the route without any credential: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 401 Unauthorized` response. This shows that consumers using different authentication methods are able to authenticate and access the resource behind the same route. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. ### Routes or Services[​](#routes-or-services "Direct link to Routes or Services") The following are plugin attributes available for configurations on [routes](https://docs.api7.ai/apisix/key-concepts/routes.md) or [services](https://docs.api7.ai/apisix/key-concepts/services.md). * auth\_plugins array\[object] required *** An array of at least two authentication plugins. --- # oas-validator The `oas-validator` plugin validates incoming HTTP requests against an OpenAPI 3 specification before forwarding them upstream. It does not validate upstream responses. OpenAPI 3.1 is supported in API7 Enterprise from version 3.9.8 and APISIX from version 3.17.0. This includes numeric `exclusiveMinimum` and `exclusiveMaximum`, conditional schemas, types that include `null`, `const`, `patternProperties`, `prefixItems`, and JSON Schema dynamic references. ## Examples[​](#examples "Direct link to Examples") Before proceeding, fetch the Swagger Petstore [Open API spec](https://petstore3.swagger.io/api/v3/openapi.json), which will be used in the following examples. * Admin API * ADC * Ingress Controller ``` export OPEN_API_SPEC=$(curl -s "https://petstore3.swagger.io/api/v3/openapi.json" | sed 's/"/\\"/g') ``` ``` export OPEN_API_SPEC=$(curl -s "https://petstore3.swagger.io/api/v3/openapi.json") ``` ``` curl -s "https://petstore3.swagger.io/api/v3/openapi.json" ``` You will replace `` with the actual OpenAPI JSON in your manifest. ### Validate Request Body[​](#validate-request-body "Direct link to Validate Request Body") This example demonstrates validation of request body against a given specification. Create the following route with the OAS validator plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: body-validation spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: / filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: oas-validator-plugin-config backendRefs: - name: petstore-external-domain port: 443 --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: petstore-external-domain spec: targetRefs: - group: "" kind: Service name: petstore-external-domain passHost: node scheme: https ``` oas-validator-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: petstore-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: petstore3.swagger.io port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: body-validation spec: ingressClassName: apisix http: - name: body-validation match: paths: - /* upstreams: - name: petstore-external-domain plugins: - name: oas-validator config: spec: ``` Apply the configuration: ``` kubectl apply -f oas-validator-ic.yaml ``` #### Failed Validation[​](#failed-validation "Direct link to Failed Validation") Send a request to the above route with a request body that does not satisfy the defined Open API Spec: ``` curl -i "http://127.0.0.1:9080/api/v3/pet" -X POST \ -H 'accept: application/json' \ -H 'Content-Type: application/json' \ -d '{"invalid-body": "this is an invalid body"}' ``` You should see an `HTTP/1.1 400 Bad Request` response with the response body similar to the following: ``` {"message":"failed to validate request."} ``` #### Successful Validation[​](#successful-validation "Direct link to Successful Validation") Send a request to the route with a valid request body: ``` curl -i "http://127.0.0.1:9080/api/v3/pet" -X POST \ -H 'accept: application/json' \ -H 'Content-Type: application/json' \ -d '{ "id": 1, "name": "doggie", "category": { "id": 1, "name": "Dogs" }, "photoUrls": ["string"], "tags": [{ "id": 1, "name": "tag1" }], "status": "available" }' ``` You should see an `HTTP/1.1 200 OK` response with the response body similar to the following: ``` { "id": 1, "category": { "id": 1, "name": "Dogs" }, "name": "doggie", "photoUrls": ["string"], "tags": [{ "id": 1, "name": "tag1" }], "status": "available" } ``` ### Get Verbose Error Response[​](#get-verbose-error-response "Direct link to Get Verbose Error Response") This example demonstrates getting verbose error response when the validation fails. Create a route with the OAS Validator plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < verbose_errors: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: body-validation spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: / filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: oas-validator-plugin-config backendRefs: - name: petstore-external-domain port: 443 --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: petstore-external-domain spec: targetRefs: - group: "" kind: Service name: petstore-external-domain passHost: node scheme: https ``` oas-validator-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: petstore-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: petstore3.swagger.io port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: body-validation spec: ingressClassName: apisix http: - name: body-validation match: paths: - /* upstreams: - name: petstore-external-domain plugins: - name: oas-validator config: spec: verbose_errors: true ``` Apply the configuration: ``` kubectl apply -f oas-validator-ic.yaml ``` Send a request to the route created above with an invalid request body: ``` curl -i "http://127.0.0.1:9080/api/v3/pet" -X POST \ -H 'accept: application/json' \ -H 'Content-Type: application/json' \ -d '{"invalid-body": "this is an invalid body"}' ``` You should see an `HTTP/1.1 400 Bad Request` response with the response body similar to the following: ``` doesn't match schema #/components/schemas/Pet: Error at "/name": property "name" is missing Schema: { "properties": { "category": { "$ref": "#/components/schemas/Category" }, "id": { "example": 10, "format": "int64", "type": "integer" }, ... } Value: { "invalid-body": "this is an invalid body" } | Error at "/photoUrls": property "photoUrls" is missing Schema: { "properties": { "category": { "$ref": "#/components/schemas/Category" }, ... } Value: { "invalid-body": "this is an invalid body" } ``` ### Monitor Violations Without Blocking Traffic[​](#monitor-violations-without-blocking-traffic "Direct link to Monitor Violations Without Blocking Traffic") Use `reject_if_not_match` to control whether non-compliant requests are blocked or allowed through. This option is available in API7 Enterprise from version 3.9.6 and APISIX from version 3.17.0. #### Reject Non-Compliant Requests[​](#reject-non-compliant-requests "Direct link to Reject Non-Compliant Requests") When `reject_if_not_match` is set to `true` (default), requests that fail OAS validation are blocked with a `HTTP/1.1 400 Bad Request` response. Create a route with the OAS Validator plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < reject_if_not_match: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: body-validation spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: / filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: oas-validator-plugin-config backendRefs: - name: petstore-external-domain port: 443 ``` oas-validator-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: petstore-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: petstore3.swagger.io port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: body-validation spec: ingressClassName: apisix http: - name: body-validation match: paths: - /* upstreams: - name: petstore-external-domain plugins: - name: oas-validator config: spec: reject_if_not_match: true ``` Apply the configuration: ``` kubectl apply -f oas-validator-ic.yaml ``` Send a request with an invalid body: ``` curl -i "http://127.0.0.1:9080/api/v3/pet" -X POST \ -H 'accept: application/json' \ -H 'Content-Type: application/json' \ -d '{"invalid-body": "this is an invalid body"}' ``` You should see an `HTTP/1.1 400 Bad Request` response with the response body similar to the following: ``` {"message":"failed to validate request."} ``` #### Allow Non-Compliant Requests to Pass[​](#allow-non-compliant-requests-to-pass "Direct link to Allow Non-Compliant Requests to Pass") When `reject_if_not_match` is set to `false`, non-compliant requests are forwarded to the upstream instead of being blocked, and validation errors are recorded in the error logs. Update the route to set `reject_if_not_match` to `false`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ --data-binary @- < reject_if_not_match: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: body-validation spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: / filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: oas-validator-plugin-config backendRefs: - name: petstore-external-domain port: 443 ``` oas-validator-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: petstore-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: petstore3.swagger.io port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: body-validation spec: ingressClassName: apisix http: - name: body-validation match: paths: - /* upstreams: - name: petstore-external-domain plugins: - name: oas-validator config: spec: reject_if_not_match: false ``` Apply the configuration: ``` kubectl apply -f oas-validator-ic.yaml ``` Send a request with an invalid body: ``` curl -i "http://127.0.0.1:9080/api/v3/pet" -X POST \ -H 'accept: application/json' \ -H 'Content-Type: application/json' \ -d '{"invalid-body": "this is an invalid body"}' ``` You should see an `HTTP/1.1 500 Internal Server Error` response since `petstore3.swagger.io` does not correctly handle an invalid body. However, you can see that the request got passed successfully to the upstream. You should also see an error log capturing the method, URI, and validation error similar to this: ``` [error] error occurred while validating request [POST /api/v3/pet], err: ... ``` This lets you audit non-compliant traffic in your logs without breaking existing clients. ### Validate Using Remote Spec URL[​](#validate-using-remote-spec-url "Direct link to Validate Using Remote Spec URL") This example applies to API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0. When your OpenAPI specification is too large to embed inline (the `spec` field has a 2 MB size limit), or you want the spec to auto-refresh periodically, use `spec_url` to load it from a remote URL. The plugin caches the compiled specification. After the cache expires, the stale entry continues serving requests while APISIX refreshes it in the background. If the initial fetch or compilation fails and no cached specification is available, the plugin returns `500 Internal Server Error`. TLS certificate verification for `spec_url` is disabled by default for compatibility. Set `ssl_verify` to `true` and configure a trusted certificate chain in production. Create a route that fetches the spec from a URL: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "url_validation", "uri": "/*", "plugins": { "oas-validator": { "spec_url": "https://petstore3.swagger.io/api/v3/openapi.json", "timeout": 5000 } }, "upstream": { "type": "roundrobin", "nodes": { "petstore3.swagger.io:443": 1 }, "scheme": "https", "pass_host": "node" } }' ``` adc.yaml ``` services: - name: petstore routes: - name: url-validation uris: - /* plugins: oas-validator: spec_url: https://petstore3.swagger.io/api/v3/openapi.json timeout: 5000 upstream: type: roundrobin nodes: - host: petstore3.swagger.io port: 443 weight: 1 scheme: https pass_host: node ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD oas-validator-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: petstore-external-domain spec: type: ExternalName externalName: petstore3.swagger.io --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: oas-validator-plugin-config spec: plugins: - name: oas-validator config: spec_url: https://petstore3.swagger.io/api/v3/openapi.json timeout: 5000 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: url-validation spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: / filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: oas-validator-plugin-config backendRefs: - name: petstore-external-domain port: 443 --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: petstore-external-domain spec: targetRefs: - group: "" kind: Service name: petstore-external-domain passHost: node scheme: https ``` oas-validator-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: petstore-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: petstore3.swagger.io port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: url-validation spec: ingressClassName: apisix http: - name: url-validation match: paths: - /* upstreams: - name: petstore-external-domain plugins: - name: oas-validator enable: true config: spec_url: https://petstore3.swagger.io/api/v3/openapi.json timeout: 5000 ``` Apply the configuration: ``` kubectl apply -f oas-validator-ic.yaml ``` If the spec endpoint requires authentication, use the same configuration shape and replace the placeholder URL, backend, and token with values from your environment: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "url_validation", "uri": "/*", "plugins": { "oas-validator": { "spec_url": "https://internal-api.example.com/openapi.json", "spec_url_request_headers": { "Authorization": "Bearer " }, "timeout": 5000 } }, "upstream": { "type": "roundrobin", "nodes": { "internal-api.example.com:443": 1 }, "scheme": "https", "pass_host": "node" } }' ``` adc.yaml ``` services: - name: internal-api routes: - name: url-validation uris: - /* plugins: oas-validator: spec_url: https://internal-api.example.com/openapi.json spec_url_request_headers: Authorization: Bearer timeout: 5000 upstream: type: roundrobin nodes: - host: internal-api.example.com port: 443 weight: 1 scheme: https pass_host: node ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD oas-validator-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: internal-api-external-domain spec: type: ExternalName externalName: internal-api.example.com --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: oas-validator-plugin-config spec: plugins: - name: oas-validator config: spec_url: https://internal-api.example.com/openapi.json spec_url_request_headers: Authorization: Bearer timeout: 5000 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: url-validation spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: / filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: oas-validator-plugin-config backendRefs: - name: internal-api-external-domain port: 443 --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: internal-api-external-domain spec: targetRefs: - group: "" kind: Service name: internal-api-external-domain passHost: node scheme: https ``` oas-validator-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: internal-api-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: internal-api.example.com port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: url-validation spec: ingressClassName: apisix http: - name: url-validation match: paths: - /* upstreams: - name: internal-api-external-domain plugins: - name: oas-validator enable: true config: spec_url: https://internal-api.example.com/openapi.json spec_url_request_headers: Authorization: Bearer timeout: 5000 ``` Apply the configuration: ``` kubectl apply -f oas-validator-ic.yaml ``` To configure the cache TTL for the fetched spec (default is 3600 seconds), set plugin metadata: * Admin API * ADC ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/oas-validator" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "spec_url_ttl": 1800 }' ``` adc.yaml ``` plugin_metadata: oas-validator: spec_url_ttl: 1800 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * spec string *** String containing the OpenAPI spec. Mutually exclusive with `spec_url`. The inline spec has a 2 MB size limit imposed by the control plane. If your OpenAPI specification exceeds this limit, use `spec_url` instead to load the spec from a remote URL. * spec\_url string *** URL to fetch the OpenAPI spec from (must start with `http://` or `https://`). Mutually exclusive with `spec`. The fetched spec is cached with a configurable TTL (see plugin metadata `spec_url_ttl`), and stale entries continue serving requests while the spec refreshes in the background. Available in API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0. * spec\_url\_request\_headers object *** Custom HTTP headers to include when fetching `spec_url` (e.g. for authentication). Available in API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0. * ssl\_verify boolean default: `false` *** Whether to verify the SSL certificate when fetching `spec_url`. Available in API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0. * timeout integer default: `10000` vaild vaule: 1000–60000 *** HTTP request timeout in milliseconds when fetching `spec_url`. Available in API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0. * verbose\_errors boolean default: `false` *** If true, respond with detailed error if the validation fails. * skip\_request\_body\_validation boolean default: `false` *** If true, skip the validation of request body. * skip\_request\_header\_validation boolean default: `false` *** If true, skip the validation of request header. * skip\_query\_param\_validation boolean default: `false` *** If true, skip the validation of query parameters. * skip\_path\_params\_validation boolean default: `false` *** If true, skip the validation of path parameters. * reject\_if\_not\_match boolean default: `true` *** If false, requests that fail OAS validation are logged as error but the request is still forwarded to the upstream service. Available in API7 Enterprise from 3.9.6 and APISIX from version 3.17.0. * rejection\_status\_code integer default: `400` vaild vaule: 400–599 *** HTTP status code to return when request validation fails. For example, set to `422` to distinguish semantic validation errors (Unprocessable Entity) from malformed request syntax (`400` Bad Request). Only effective when `reject_if_not_match` is `true`. Available in API7 Enterprise from version 3.9.8 and APISIX from version 3.17.0. * max\_req\_body\_size integer default: `67108864` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes read for OpenAPI validation. A larger body returns `500 Internal Server Error`. This field has no effect when `skip_request_body_validation` is `true`. Introduced in API7 Enterprise 3.9.17 and 3.10.4, and APISIX 3.18.0. note One of `spec` or `spec_url` must be configured. They are mutually exclusive. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * spec\_url\_ttl integer default: `3600` *** TTL in seconds for cached specs fetched from `spec_url`. After expiry, stale entries continue serving requests while the spec refreshes asynchronously in the background. Available in API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0. --- # OPA The `opa` plugin supports the integration with [Open Policy Agent (OPA)](https://www.openpolicyagent.org), a unified policy engine and framework that helps define and enforce authorization policies. Authorization logic is defined in [Rego](https://www.openpolicyagent.org/docs/latest/policy-language/) and stored in OPA. Once configured, the OPA engine will evaluate the client request to a protected route to determine whether the request should have access to the upstream resource based on the defined policies. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can work with the `opa` plugin for different scenarios. Before proceeding, you should have a running OPA server. Start one using Docker or deploy it to Kubernetes: * Docker * Kubernetes ``` docker run -d --name opa-server -p 8181:8181 openpolicyagent/opa:1.6.0 run --server --addr :8181 --log-level debug ``` * `run -s` starts OPA as a server. * `--log-level debug` prints debug information to examine the data APISIX pushes to OPA. To verify that the OPA server is installed and port is exposed properly, run: ``` curl "http://127.0.0.1:8181" | grep Version ``` You should see a response similar to the following: ``` Version: 1.6.0 ``` Create a Deployment and Service for OPA in your cluster: opa-server.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: opa spec: replicas: 1 selector: matchLabels: app: opa template: metadata: labels: app: opa spec: containers: - name: opa image: openpolicyagent/opa:1.6.0 args: - run - --server - --addr=:8181 - --log-level=debug ports: - containerPort: 8181 --- apiVersion: v1 kind: Service metadata: namespace: aic name: opa spec: selector: app: opa ports: - port: 8181 targetPort: 8181 ``` Apply the configuration to your cluster: ``` kubectl apply -f opa-server.yaml ``` Wait for the OPA pod to be ready. Once ready, the OPA server will be available within the cluster at `http://opa.aic.svc.cluster.local:8181`. To push policies to it from outside the cluster, set up a port-forward: ``` kubectl port-forward -n aic svc/opa 8181:8181 & ``` ### Implement a Basic Policy[​](#implement-a-basic-policy "Direct link to Implement a Basic Policy") The following example implements a basic authorization policy in OPA to allow only GET requests. Create an OPA policy that only allows HTTP GET requests: ``` curl "http://127.0.0.1:8181/v1/policies/getonly" -X PUT \ -H "Content-Type: text/plain" \ -d ' package getonly default allow = false allow if { input.request.method == "GET" }' ``` Create a route with the `opa` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "opa-route", "uri": "/anything", "plugins": { "opa": { "host": "http://192.168.2.104:8181", "policy": "getonly" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` ❶ Configure the OPA server address. Replace with your IP address. ❷ Set the authorization policy to be `getonly`. adc.yaml ``` services: - name: opa-service routes: - name: opa-route uris: - /anything plugins: opa: host: "http://192.168.2.104:8181" policy: getonly upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ Configure the OPA server address. Replace with your IP address. ❷ Set the authorization policy to be `getonly`. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD opa-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: opa-plugin-config spec: plugins: - name: opa config: host: "http://opa.aic.svc.cluster.local:8181" policy: getonly --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: opa-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: opa-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` ❶ Configure the OPA server address. ❷ Set the authorization policy to be `getonly`. Apply the configuration to your cluster: ``` kubectl apply -f opa-ic.yaml ``` opa-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: opa-route spec: ingressClassName: apisix http: - name: opa-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: opa enable: true config: host: "http://opa.aic.svc.cluster.local:8181" policy: getonly ``` ❶ Configure the OPA server address. ❷ Set the authorization policy to be `getonly`. Apply the configuration to your cluster: ``` kubectl apply -f opa-ic.yaml ``` To verify the policy, send a GET request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Send another request to the route using PUT: ``` curl -i "http://127.0.0.1:9080/anything" -X PUT ``` You should receive an `HTTP/1.1 403 Forbidden` response. ### Understand Data Format[​](#understand-data-format "Direct link to Understand Data Format") The following example helps you understand the data and the format APISIX pushes to OPA to support authorization logic writing. The example continues with the policy and the route in the [last example](#implement-a-basic-policy) Suppose your OPA server is started with `--log-level debug` and you have completed the verification steps in the [last example](#implement-a-basic-policy) sending requests to the sample route. Navigate to the OPA server log. You should see an entry similar to the following: ``` { "client_addr": "192.168.215.1:58467", "level": "info", "msg": "Received request.", "req_body": "{\"input\":{\"type\":\"http\",\"var\":{\"server_port\":\"9080\",\"timestamp\":1752400020,\"server_addr\":\"192.168.107.3\",\"remote_port\":\"58544\",\"remote_addr\":\"192.168.107.1\"},\"request\":{\"host\":\"127.0.0.1\",\"path\":\"/anything\",\"headers\":{\"host\":\"127.0.0.1:9080\",\"accept\":\"*/*\",\"user-agent\":\"curl/8.6.0\"},\"query\":{},\"port\":9080,\"scheme\":\"http\",\"method\":\"PUT\"}}}", "req_id": 12, "req_method": "POST", "req_params": {}, "req_path": "/v1/data/getonly", "time": "2025-07-14T15:07:00Z" } ``` where the `req_body` shows the data APISIX pushed: ``` { "input": { "type": "http", "var": { "server_port": "9080", "timestamp": 1752400020, "server_addr": "192.168.107.3", "remote_port": "58544", "remote_addr": "192.168.107.1" }, "request": { "host": "127.0.0.1", "path": "/anything", "headers": { "host": "127.0.0.1:9080", "accept": "*/*", "user-agent": "curl/8.6.0" }, "query": {}, "port": 9080, "scheme": "http", "method": "PUT" } } } ``` Now, update the plugin on the [previously created route](#implement-a-basic-policy) to include route information: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes/opa-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "opa": { "with_route": true } } }' ``` Update `adc.yaml` to add `with_route: true`: adc.yaml ``` services: - name: opa-service routes: - name: opa-route uris: - /anything plugins: opa: host: "http://192.168.2.104:8181" policy: getonly with_route: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update `opa-ic.yaml` to add `with_route: true`: opa-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: opa-plugin-config spec: plugins: - name: opa config: host: "http://opa.aic.svc.cluster.local:8181" policy: getonly with_route: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: opa-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: opa-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the updated configuration to your cluster: ``` kubectl apply -f opa-ic.yaml ``` Update `opa-ic.yaml` to add `with_route: true`: opa-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: opa-route spec: ingressClassName: apisix http: - name: opa-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: opa enable: true config: host: "http://opa.aic.svc.cluster.local:8181" policy: getonly with_route: true ``` Apply the updated configuration to your cluster: ``` kubectl apply -f opa-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` In the OPA server log, you should see a new entry: ``` { "client_addr": "192.168.215.1:43706", "level": "info", "msg": "Received request.", "req_body": "{\"input\":{\"route\":{\"id\":\"opa-route\",\"uri\":\"/anything\",\"update_time\":1752395758,\"plugins\":{\"opa\":{\"keepalive_pool\":5,\"keepalive_timeout\":60000,\"host\":\"http://172.17.1.196:8181\",\"ssl_verify\":true,\"with_route\":true,\"with_service\":false,\"with_consumer\":false,\"timeout\":3000,\"keepalive\":true,\"policy\":\"getonly\"}},\"priority\":0,\"status\":1,\"create_time\":1752393063},\"type\":\"http\",\"var\":{\"server_port\":\"9080\",\"timestamp\":1752396233,\"server_addr\":\"192.168.107.3\",\"remote_port\":\"47838\",\"remote_addr\":\"192.168.107.1\"},\"request\":{\"host\":\"127.0.0.1\",\"path\":\"/anything\",\"headers\":{\"host\":\"127.0.0.1:9080\",\"accept\":\"*/*\",\"user-agent\":\"curl/8.6.0\"},\"query\":{},\"port\":9080,\"scheme\":\"http\",\"method\":\"GET\"}}}", "req_id": 14, "req_method": "POST", "req_params": {}, "req_path": "/v1/data/getonly", "time": "2025-07-13T08:43:53Z" } ``` The `req_body` now includes route information: ``` { "input": { "route": { "id": "opa-route", "uri": "/anything", "update_time": 1752395758, "plugins": { "opa": { "keepalive_pool": 5, "keepalive_timeout": 60000, "host": "http://172.17.1.196:8181", "ssl_verify": true, "with_route": true, "with_service": false, "with_consumer": false, "timeout": 3000, "keepalive": true, "policy": "getonly" } }, "priority": 0, "status": 1, "create_time": 1752393063 }, "type": "http", "var": { "server_port": "9080", "timestamp": 1752396233, "server_addr": "192.168.107.3", "remote_port": "47838", "remote_addr": "192.168.107.1" }, "request": { "host": "127.0.0.1", "path": "/anything", "headers": { "host": "127.0.0.1:9080", "accept": "*/*", "user-agent": "curl/8.6.0" }, "query": {}, "port": 9080, "scheme": "http", "method": "GET" } } } ``` ### Return Custom Response[​](#return-custom-response "Direct link to Return Custom Response") The following example demonstrates how you can return custom response code and message when the request is unauthorized. Create an OPA policy that only allows HTTP GET requests and return `302` with a custom message the request is unauthorized: ``` curl "http://127.0.0.1:8181/v1/policies/customresp" -X PUT \ -H "Content-Type: text/plain" \ -d ' package customresp default allow = false allow if { input.request.method == "GET" } reason := "The resource has temporarily moved. Please follow the new URL." if { not allow } headers := { "Location": "http://example.com/auth" } if { not allow } status_code := 302 if { not allow } ' ``` Create a route with the `opa` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "opa-route", "uri": "/anything", "plugins": { "opa": { "host": "http://192.168.2.104:8181", "policy": "customresp" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` ❶ Configure the OPA server address. Replace with your IP address. ❷ Set the authorization policy to be `customresp`. adc.yaml ``` services: - name: opa-service routes: - name: opa-route uris: - /anything plugins: opa: host: "http://192.168.2.104:8181" policy: customresp upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ Configure the OPA server address. Replace with your IP address. ❷ Set the authorization policy to be `customresp`. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD opa-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: opa-customresp-plugin-config spec: plugins: - name: opa config: host: "http://opa.aic.svc.cluster.local:8181" policy: customresp --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: opa-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: opa-customresp-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` ❶ Configure the OPA server address. ❷ Set the authorization policy to be `customresp`. Apply the configuration to your cluster: ``` kubectl apply -f opa-ic.yaml ``` opa-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: opa-route spec: ingressClassName: apisix http: - name: opa-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: opa enable: true config: host: "http://opa.aic.svc.cluster.local:8181" policy: customresp ``` ❶ Configure the OPA server address. ❷ Set the authorization policy to be `customresp`. Apply the configuration to your cluster: ``` kubectl apply -f opa-ic.yaml ``` Send a GET request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Send a POST request to the route: ``` curl -i "http://127.0.0.1:9080/anything" -X POST ``` You should receive an `HTTP/1.1 302 Moved Temporarily` response: ``` HTTP/1.1 302 Moved Temporarily ... Location: http://example.com/auth The resource has temporarily moved. Please follow the new URL. ``` ### Implement RBAC[​](#implement-rbac "Direct link to Implement RBAC") The following example demonstrates how to implement authentication and RBAC using the [`jwt-auth`](https://docs.api7.ai/hub/jwt-auth.md) and `opa` plugins. You will be implementing RBAC logics such that: * An `user` role can only read the upstream resources. * An `admin` role can read and write the upstream resources. Create an OPA policy for RBAC of two example consumers, where `john` has the `user` role and `jane` has the `admin` role: ``` curl "http://127.0.0.1:8181/v1/policies/rbac" -X PUT \ -H "Content-Type: text/plain" \ -d ' package rbac # Assign roles to users user_roles := { "john": ["user"], "jane": ["admin"] } # Map permissions to HTTP methods permission_methods := { "read": "GET", "write": "POST" } # Assign role permissions role_permissions := { "user": ["read"], "admin": ["read", "write"] } # Get JWT authorization token bearer_token := t if { t := input.request.headers.authorization } # Decode the token to get role and permission token := {"payload": payload} if { [_, payload, _] := io.jwt.decode(bearer_token) } # Normalize permission to a list normalized_permissions := ps if { ps := token.payload.permission not is_string(ps) } normalized_permissions := [ps] if { ps := token.payload.permission is_string(ps) } # Implement RBAC logic default result := {"allow": false} result := {"allow": true} if { # Look up the list of roles for the user roles := user_roles[input.consumer.username] # For each role in that list r := roles[_] # Look up the permissions list for the role permissions := role_permissions[r] # For each permission p := permissions[_] # Check if the permission matches the request method permission_methods[p] == input.request.method # Check if the normalized permissions include the permission p in normalized_permissions } ' ``` Create two consumers `john` and `jane` in APISIX and configure their `jwt-auth` credentials: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" \ -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "username": "john" }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" \ -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "username": "jane" }' ``` Configure the `jwt-auth` credentials for the consumers, using the default algorithm `HS256`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-jwt-auth", "plugins": { "jwt-auth": { "key": "john-key", "secret": "john-hs256-secret-that-is-very-long" } } }' ``` ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jane/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-jwt-auth", "plugins": { "jwt-auth": { "key": "jane-key", "secret": "jane-hs256-secret-that-is-very-long" } } }' ``` adc.yaml ``` consumers: - username: john credentials: - name: cred-john-jwt-auth type: jwt-auth config: key: john-key secret: john-hs256-secret-that-is-very-long - username: jane credentials: - name: cred-jane-jwt-auth type: jwt-auth config: key: jane-key secret: jane-hs256-secret-that-is-very-long ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD opa-consumers-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: jwt-auth name: cred-john-jwt-auth config: key: john-key secret: john-hs256-secret-that-is-very-long --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jane spec: gatewayRef: name: apisix credentials: - type: jwt-auth name: cred-jane-jwt-auth config: key: jane-key secret: jane-hs256-secret-that-is-very-long ``` Apply the configuration to your cluster: ``` kubectl apply -f opa-consumers-ic.yaml ``` opa-consumers-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: jwtAuth: value: key: john-key secret: john-hs256-secret-that-is-very-long --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jane spec: ingressClassName: apisix authParameter: jwtAuth: value: key: jane-key secret: jane-hs256-secret-that-is-very-long ``` Apply the configuration to your cluster: ``` kubectl apply -f opa-consumers-ic.yaml ``` Create a route and configure the `jwt-auth` and `opa` plugins as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "opa-route", "methods": ["GET", "POST"], "uris": ["/get","/post"], "plugins": { "jwt-auth": {}, "opa": { "host": "http://192.168.2.104:8181", "policy": "rbac/result", "with_consumer": true } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` ❶ Enable the `jwt-auth` plugin on the route. ❷ Configure the OPA server address. Replace with your IP address. ❸ Set the authorization policy to be `rbac/result`. ❹ Set `with_consumer` to true to send consumer information. Update `adc.yaml` to add the route with `jwt-auth` and `opa` plugins: adc.yaml ``` consumers: - username: john credentials: - name: cred-john-jwt-auth type: jwt-auth config: key: john-key secret: john-hs256-secret-that-is-very-long - username: jane credentials: - name: cred-jane-jwt-auth type: jwt-auth config: key: jane-key secret: jane-hs256-secret-that-is-very-long services: - name: opa-service routes: - name: opa-route uris: - /get - /post methods: - GET - POST plugins: jwt-auth: {} opa: host: "http://192.168.2.104:8181" policy: rbac/result with_consumer: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` ❶ Enable the `jwt-auth` plugin on the route. ❷ Configure the OPA server address. Replace with your IP address. ❸ Set the authorization policy to be `rbac/result`. ❹ Set `with_consumer` to true to send consumer information. Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD opa-route-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: opa-rbac-plugin-config spec: plugins: - name: jwt-auth config: _meta: disable: false - name: opa config: host: "http://opa.aic.svc.cluster.local:8181" policy: rbac/result with_consumer: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: opa-rbac-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get method: GET filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: opa-rbac-plugin-config backendRefs: - name: httpbin-external-domain port: 80 - matches: - path: type: Exact value: /post method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: opa-rbac-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` ❶ Configure the OPA server address. ❷ Set the authorization policy to be `rbac/result`. ❸ Set `with_consumer` to true to send consumer information. Apply the configuration to your cluster: ``` kubectl apply -f opa-route-ic.yaml ``` note When using the Ingress Controller, APISIX prefixes consumer names with the Kubernetes namespace. For example, a consumer named `john` in the `aic` namespace becomes `aic_john`. Update the OPA RBAC policy to use the prefixed names: ``` curl "http://127.0.0.1:8181/v1/policies/rbac" -X PUT \ -H "Content-Type: text/plain" \ -d ' package rbac # Assign roles to users user_roles := { "aic_john": ["user"], "aic_jane": ["admin"] } # Map permissions to HTTP methods permission_methods := { "read": "GET", "write": "POST" } # Assign role permissions role_permissions := { "user": ["read"], "admin": ["read", "write"] } # Get JWT authorization token bearer_token := t if { t := input.request.headers.authorization } # Decode the token to get role and permission token := {"payload": payload} if { [_, payload, _] := io.jwt.decode(bearer_token) } # Normalize permission to a list normalized_permissions := ps if { ps := token.payload.permission not is_string(ps) } normalized_permissions := [ps] if { ps := token.payload.permission is_string(ps) } # Implement RBAC logic default result := {"allow": false} result := {"allow": true} if { # Look up the list of roles for the user roles := user_roles[input.consumer.username] # For each role in that list r := roles[_] # Look up the permissions list for the role permissions := role_permissions[r] # For each permission p := permissions[_] # Check if the permission matches the request method permission_methods[p] == input.request.method # Check if the normalized permissions include the permission p in normalized_permissions } ' ``` opa-route-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: opa-rbac-route spec: ingressClassName: apisix http: - name: get-route match: methods: - GET paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: jwt-auth enable: true config: _meta: disable: false - name: opa enable: true config: host: "http://opa.aic.svc.cluster.local:8181" policy: rbac/result with_consumer: true - name: post-route match: methods: - POST paths: - /post upstreams: - name: httpbin-external-domain plugins: - name: jwt-auth enable: true config: _meta: disable: false - name: opa enable: true config: host: "http://opa.aic.svc.cluster.local:8181" policy: rbac/result with_consumer: true ``` Apply the configuration to your cluster: ``` kubectl apply -f opa-route-ic.yaml ``` note When using the Ingress Controller, APISIX prefixes consumer names with the Kubernetes namespace. For example, a consumer named `john` in the `aic` namespace becomes `aic_john`. Update the OPA RBAC policy to use the prefixed names: ``` curl "http://127.0.0.1:8181/v1/policies/rbac" -X PUT \ -H "Content-Type: text/plain" \ -d ' package rbac user_roles := { "aic_john": ["user"], "aic_jane": ["admin"] } permission_methods := { "read": "GET", "write": "POST" } role_permissions := { "user": ["read"], "admin": ["read", "write"] } bearer_token := t if { t := input.request.headers.authorization } token := {"payload": payload} if { [_, payload, _] := io.jwt.decode(bearer_token) } normalized_permissions := ps if { ps := token.payload.permission not is_string(ps) } normalized_permissions := [ps] if { ps := token.payload.permission is_string(ps) } default result := {"allow": false} result := {"allow": true} if { roles := user_roles[input.consumer.username] r := roles[_] permissions := role_permissions[r] p := permissions[_] permission_methods[p] == input.request.method p in normalized_permissions } ' ``` #### Verify as `john`[​](#verify-as-john "Direct link to verify-as-john") To issue a JWT for `john`, you could use [JWT.io's JWT encoder](https://jwt.io) or other utilities. If you are using [JWT.io's JWT encoder](https://jwt.io), do the following: * Fill in `HS256` as the algorithm. * Update the secret in the **Valid secret** section to be `john-hs256-secret-that-is-very-long`. * Update payload with role `user`, permission `read`, and consumer key `john-key`; as well as `exp` or `nbf` in UNIX timestamp. note When `claims_to_verify` is a nonempty list, every listed claim is required and validated. When it is unset or empty, `exp` and `nbf` are validated when present, but neither claim is required. Your payload should look similar to the following: ``` { "role": "user", "permission": "read", "key": "john-key", "nbf": 1729132271 } ``` Copy the generated JWT and save to a variable: ``` export john_jwt_token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJyb2xlIjoidXNlciIsInBlcm1pc3Npb24iOiJyZWFkIiwia2V5Ijoiam9obi1rZXkiLCJuYmYiOjE3MjkxMzIyNzF9.rAHMTQfnnGFnKYc3am_lpE9pZ9E8EaOT_NBQ5Ss8pk4 ``` Send a GET request to the route with the JWT of `john`: ``` curl -i "http://127.0.0.1:9080/get" -H "Authorization: ${john_jwt_token}" ``` You should receive an `HTTP/1.1 200 OK` response. Send a POST request to the route with the same JWT: ``` curl -i "http://127.0.0.1:9080/post" -X POST -H "Authorization: ${john_jwt_token}" ``` You should receive an `HTTP/1.1 403 Forbidden` response. #### Verify as `jane`[​](#verify-as-jane "Direct link to verify-as-jane") Similarly, to issue a JWT for `jane`, you could use [JWT.io's JWT encoder](https://jwt.io) or other utilities. If you are using [JWT.io's JWT encoder](https://jwt.io), do the following: * Fill in `HS256` as the algorithm. * Update the secret in the **Valid secret** section to be `jane-hs256-secret-that-is-very-long`. * Update payload with role `admin`, permission `["read","write"]`, and consumer key `jane-key`; as well as `exp` or `nbf` in UNIX timestamp. note When `claims_to_verify` is a nonempty list, every listed claim is required and validated. When it is unset or empty, `exp` and `nbf` are validated when present, but neither claim is required. Your payload should look similar to the following: ``` { "role": "admin", "permission": ["read","write"], "key": "jane-key", "nbf": 1729132271 } ``` Copy the generated JWT and save to a variable: ``` export jane_jwt_token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJyb2xlIjoiYWRtaW4iLCJwZXJtaXNzaW9uIjpbInJlYWQiLCJ3cml0ZSJdLCJrZXkiOiJqYW5lLWtleSIsIm5iZiI6MTcyOTEzMjI3MX0.meZ-AaGHUPwN_GvVOE3IkKuAJ1wqlCguaXf3gm3Ww8s ``` Send a GET request to the route with the JWT of `jane`: ``` curl -i "http://127.0.0.1:9080/get" -H "Authorization: ${jane_jwt_token}" ``` You should receive an `HTTP/1.1 200 OK` response. Send a POST request to the route with the same JWT: ``` curl -i "http://127.0.0.1:9080/post" -X POST -H "Authorization: ${jane_jwt_token}" ``` You should also receive an `HTTP/1.1 200 OK` response. tip To examine whether the authorization decision comes from OPA, you should observe the following log in the OPA server if you have set `--log-level debug`: ``` { "result":{ "allow": true, "bearer_token": "eyJ...", ... } } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * host string required *** Address of the OPA server. * policy string required *** Policy to evaluate. For example, if you would like to evaluate all rules in a package called `rbac`, configure the policy to be `rbac`. If you would like to evaluate specific rule(s) in a package, you can specify the rule name behind the package, such as `rbac/allow`. * ssl\_verify boolean default: `true` *** If true, verify the OPA server's SSL certificate. * timeout integer default: `3000` vaild vaule: between 1 and 60000 inclusive *** Timeout for the HTTP call in milliseconds. * keepalive boolean default: `true` *** If true, keep the connection alive for multiple requests. * keepalive\_timeout integer default: `60000` vaild vaule: greater than or equal to 1000 *** Idle time in milliseconds after which the connection is closed. * keepalive\_pool integer default: `5` vaild vaule: greater than or equal to 1 *** The number of idle connections. * with\_route boolean default: `false` *** If true, send information of the current route. * with\_service boolean default: `false` *** If true, send information of the current service. * with\_consumer boolean default: `false` *** If true, send information of the current consumer. Note that the consumer information may include sensitive information such as the API key. Only set this option to `true` if you are sure it is safe to do so. * send\_headers\_upstream array\[string] *** Header names controlled by the OPA response when the request is allowed. The gateway forwards a configured header when OPA returns it and clears any client-supplied value when OPA omits it. --- # openapi-to-mcp [Enterprise](https://api7.ai/enterprise) The `openapi-to-mcp` plugin enables the gateway to act as a bridge between OpenAPI specifications and MCP (Model Context Protocol) servers. With this plugin, you can expose your existing OpenAPI-based services through an MCP interface, making them accessible to AI models and clients. The plugin works by converting your OpenAPI specification into the MCP format and serving it through an MCP server interface. Requests from AI clients are then proxied to your upstream services, with support for custom headers and two transport methods for streaming responses: streamable HTTP and Server-Sent Events (SSE), allowing flexible and reliable real-time communication. The following diagram illustrates the interaction between the MCP client, API7 Gateway, and an upstream OpenAPI service. The path and data are example values for demonstration.
## Demo[​](#demo "Direct link to Demo") The following example demonstrates how to [enable MCP access to Petstore APIs](#enable-mcp-access-to-petstore-apis), allowing AI models and clients to interact with the Petstore service. When configured correctly, the AI client should immediately see available Petstore tools; if tools aren't appearing, verify the OpenAPI specification URL is accessible and the gateway address is reachable from your AI client environment.
## Deployment and Compatibility[​](#deployment-and-compatibility "Direct link to Deployment and Compatibility") ### Deployment Prerequisite[​](#deployment-prerequisite "Direct link to Deployment Prerequisite") Starting from API7 Enterprise **3.9.10**, the OpenAPI-to-MCP service is no longer bundled inside the gateway image and must be deployed alongside the gateway in the same network namespace. * **Kubernetes (Helm)**: set `openapiToMcp.enabled: true` in the gateway chart values to run the service as a sidecar. * **Docker / bare metal**: run the `api7/openapi-to-mcp` image as a standalone container and share the network namespace with the gateway (for example, `--network=container:` or host networking) so the plugin can reach it at `127.0.0.1:` (default `3000`). The plugin hard-codes `127.0.0.1` as the target, so the two containers must share a network namespace; a shared Docker bridge network is not sufficient. If the service is unreachable, the plugin will fail with a 503. This applies equally to the `mcp-tools-acl` plugin. To run the service on a different port, see [Static Configurations](https://docs.api7.ai/hub/openapi-to-mcp/configuration.md#static-configurations). ### Sidecar Image Tag and Gateway Version Compatibility[​](#sidecar-image-tag-and-gateway-version-compatibility "Direct link to Sidecar Image Tag and Gateway Version Compatibility") The plugin and the OpenAPI-to-MCP service communicate over a stable internal contract that rarely changes, so a single sidecar image tag is compatible with a wide range of gateway versions. The table below tracks which sidecar tag to use for each gateway version range — only update the sidecar when the protocol between the two changes (i.e. when a new row is added below). This guidance is intended for **Docker and bare-metal** deployments where you choose the sidecar image tag yourself. The Helm chart already pins a known-good sidecar tag for the bundled gateway version, so Helm users do not need to consult this table. | Gateway version | Sidecar image tag (`api7/openapi-to-mcp`) | | ------------------ | ----------------------------------------- | | `3.9.10` and later | `1.0.2` or later | ### OpenAPI Document Caching[​](#openapi-document-caching "Direct link to OpenAPI Document Caching") The OpenAPI-to-MCP service caches the parsed OpenAPI document and the tools generated from it, keyed by the `openapi_url` string. The cache is enabled by default and each entry expires after 3600 seconds. If the document changes but `openapi_url` stays the same, MCP clients see the change within about an hour. The cache is not invalidated by document content or HTTP caching headers. To make a document change take effect immediately, change the `openapi_url` value — for example, add or update a query parameter such as `?v=2` — and save the plugin. The service treats the new string as a new document and fetches it on the next request. For **Docker and bare-metal** deployments you can tune the cache through environment variables on the `api7/openapi-to-mcp` container: | Variable | Default | Description | | --------------- | ------- | ----------------------------------------------------------------------------------------------------------------------- | | `CACHE_ENABLED` | `true` | Set to `false` to disable caching. Every request then downloads and parses the document again, which increases latency. | | `CACHE_TTL` | `3600` | Seconds an entry stays in the cache. | The Helm chart does not currently expose these variables for the sidecar; Helm deployments use the defaults. ## Deploy with Docker Compose[​](#deploy-with-docker-compose "Direct link to Deploy with Docker Compose") The following `docker-compose.yaml` runs the API7 Enterprise gateway and the OpenAPI-to-MCP service together. The MCP service joins the gateway's network namespace so the plugin can reach it at `127.0.0.1:3000`. When you add a gateway instance in the Dashboard, a ready-to-use `docker-compose.yaml` is generated automatically. To enable the `openapi-to-mcp` plugin, add the `openapi-to-mcp` service shown below to that generated file: docker-compose.yaml ``` services: gateway: image: api7/api7-ee-3-gateway:3.9.12 container_name: gateway hostname: gateway restart: always ports: - "9080:9080" - "9443:9443" environment: API7_DP_MANAGER_ENDPOINTS: '["https://:7943"]' API7_GATEWAY_GROUP_SHORT_ID: "" API7_DP_MANAGER_CERT: | -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- API7_DP_MANAGER_KEY: | -----BEGIN PRIVATE KEY----- ... -----END PRIVATE KEY----- API7_CONTROL_PLANE_CA: | -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- openapi-to-mcp: image: api7/openapi-to-mcp:1.0.2 network_mode: "service:gateway" restart: always ``` ❶ The DP Manager endpoint(s) — replace with the address provided in the Dashboard. ❷ The gateway group short ID — replace with the value shown in the Dashboard. ❸ The TLS client certificate, private key, and CA certificate. These are populated automatically when you copy the generated compose file from the Dashboard. ❹ `network_mode: "service:gateway"` makes the MCP container share the gateway's network stack, so the plugin can reach the MCP service at `127.0.0.1:3000`. A shared Docker bridge network is **not** sufficient — the plugin hard-codes `127.0.0.1` as the target. Start the services: ``` docker compose up -d ``` ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `openapi-to-mcp` plugin for different scenarios. ### Enable MCP Access to Petstore APIs[​](#enable-mcp-access-to-petstore-apis "Direct link to Enable MCP Access to Petstore APIs") The following example demonstrates how to expose the Petstore APIs through the MCP protocol, allowing AI models and clients to interact with the Petstore service. * Admin API * ADC * Ingress Controller Create a route with the `openapi-to-mcp` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "openapi-to-mcp-route", "uri": "/mcp", "methods": ["GET", "POST"], "plugins": { "openapi-to-mcp": { "transport": "streamable_http", "base_url": "https://petstore3.swagger.io/api/v3", "headers": { "Authorization": "special-key" }, "openapi_url": "https://petstore3.swagger.io/api/v3/openapi.json" } } }' ``` ❶ Configure the route to allow GET and POST methods. The GET method enables the tool discovery and response streaming (SSE), while the POST method enables the execution and action capabilities (messages). ❷ Configure the transport method to be `streamable_http` (recommended for production). ❸ Configure the Petstore API address. ❹ Configure the Petstore API credential. ❺ Configure the Petstore OpenAPI document URL. Create a route with the `openapi-to-mcp` plugin configured as follows: adc.yaml ``` services: - name: openapi-to-mcp-service upstream: type: roundrobin scheme: https pass_host: node nodes: - host: petstore3.swagger.io port: 443 weight: 1 routes: - name: openapi-to-mcp-route uris: - /mcp methods: - GET - POST plugins: openapi-to-mcp: transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` ❶ Configure the route to allow GET and POST methods. The GET method enables the tool discovery and response streaming (SSE), while the POST method enables the execution and action capabilities (messages). ❷ Configure the transport method to be `streamable_http` (recommended for production). ❸ Configure the Petstore API address. ❹ Configure the Petstore API credential. ❺ Configure the Petstore OpenAPI document URL. * Gateway API * APISIX CRD Create a route with the `openapi-to-mcp` plugin configured as follows: openapi-to-mcp-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: openapi-to-mcp-plugin-config spec: plugins: - name: openapi-to-mcp config: transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: openapi-to-mcp-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /mcp method: GET - path: type: Exact value: /mcp method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: openapi-to-mcp-plugin-config backendRefs: - name: petstore-external-domain port: 443 --- apiVersion: v1 kind: Service metadata: namespace: aic name: petstore-external-domain spec: type: ExternalName externalName: petstore3.swagger.io --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: petstore-external-domain spec: targetRefs: - group: "" kind: Service name: petstore-external-domain passHost: node scheme: https ``` Apply the configuration to your cluster: ``` kubectl apply -f openapi-to-mcp-ic.yaml ``` ❶ Configure the transport method to be `streamable_http` (recommended for production). ❷ Configure the Petstore API address. ❸ Configure the Petstore API credential. ❹ Configure the Petstore OpenAPI document URL. ❺ Configure the route to allow GET and POST methods. The GET method enables the tool discovery and response streaming (SSE), while the POST method enables the execution and action capabilities (messages). Create a route with the `openapi-to-mcp` plugin configured as follows: openapi-to-mcp-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: openapi-to-mcp-route spec: ingressClassName: apisix http: - name: openapi-to-mcp-route match: paths: - /mcp methods: - GET - POST plugins: - name: openapi-to-mcp enable: true config: transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" ``` Apply the configuration to your cluster: ``` kubectl apply -f openapi-to-mcp-ic.yaml ``` ❶ Configure the route to allow GET and POST methods. The GET method enables the tool discovery and response streaming (SSE), while the POST method enables the execution and action capabilities (messages). ❷ Configure the transport method to be `streamable_http` (recommended for production). ❸ Configure the Petstore API address. ❹ Configure the Petstore API credential. ❺ Configure the Petstore OpenAPI document URL. After applying the Admin API, ADC, or APISIX CRD configuration, update your MCP settings with the API7 Gateway address and append the previously created route path. For instance: mcp.json ``` { "mcpServers": { "api7-petstore-mcp": { "url": "http://123.123.123.123:9080/mcp" } } } ``` If the configuration is successful, you should see the available tools (external functions or services exposed to AI clients through MCP). You can now interact with the Petstore service directly from the chat window of your AI client. For example, try asking: "Show me pet 1 from the petstore." ![AI client interaction with Petstore](https://static.api7.ai/uploads/2025/09/22/6TE6DgXy_oet.png) ### Configure Authentication for MCP Routes[​](#configure-authentication-for-mcp-routes "Direct link to Configure Authentication for MCP Routes") The following example demonstrates how to expose Petstore APIs through the MCP protocol when the route is protected by an authentication method such as `key-auth`. * Admin API * ADC * Ingress Controller Create a consumer `johndoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "johndoe" }' ``` Configure the `key-auth` credential for `johndoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/johndoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Create a route with the `openapi-to-mcp` and `key-auth` plugins: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "openapi-to-mcp-route", "uri": "/mcp", "methods": ["GET", "POST"], "plugins": { "openapi-to-mcp": { "transport": "streamable_http", "base_url": "https://petstore3.swagger.io/api/v3", "headers": { "Authorization": "special-key" }, "openapi_url": "https://petstore3.swagger.io/api/v3/openapi.json" }, "key-auth": { "header": "apikey" } } }' ``` Create a consumer and a route with the `openapi-to-mcp` and `key-auth` plugins configured as follows: adc.yaml ``` consumers: - username: johndoe credentials: - name: primary-key type: key-auth config: key: john-key services: - name: openapi-to-mcp-service upstream: type: roundrobin scheme: https pass_host: node nodes: - host: petstore3.swagger.io port: 443 weight: 1 routes: - name: openapi-to-mcp-route uris: - /mcp methods: - GET - POST plugins: key-auth: header: apikey openapi-to-mcp: transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Create a consumer and a route with the `openapi-to-mcp` and `key-auth` plugins configured as follows: openapi-to-mcp-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: johndoe spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: john-key --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: openapi-to-mcp-plugin-config spec: plugins: - name: key-auth config: header: apikey - name: openapi-to-mcp config: transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: openapi-to-mcp-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /mcp method: GET - path: type: Exact value: /mcp method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: openapi-to-mcp-plugin-config backendRefs: - name: petstore-external-domain port: 443 --- apiVersion: v1 kind: Service metadata: namespace: aic name: petstore-external-domain spec: type: ExternalName externalName: petstore3.swagger.io --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: petstore-external-domain spec: targetRefs: - group: "" kind: Service name: petstore-external-domain passHost: node scheme: https ``` Create an ApisixConsumer and a route with the `openapi-to-mcp` and `key-auth` plugins configured as follows: openapi-to-mcp-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: johndoe spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: petstore-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: petstore3.swagger.io port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: openapi-to-mcp-route spec: ingressClassName: apisix http: - name: openapi-to-mcp-route match: paths: - /mcp methods: - GET - POST upstreams: - name: petstore-external-domain plugins: - name: key-auth enable: true config: header: apikey - name: openapi-to-mcp enable: true config: transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" ``` Apply the configuration to your cluster: ``` kubectl apply -f openapi-to-mcp-ic.yaml ``` When an MCP server requires authentication, you can specify headers in the `mcp.json` configuration. Refer to the documentation of your AI client to verify whether headers are supported. #### If Headers Are Supported[​](#if-headers-are-supported "Direct link to If Headers Are Supported") For example, in Cursor you can update the MCP settings with your API7 Gateway address, append the previously created route path, and include the header required for `key-auth` after applying the Admin API, ADC, or APISIX CRD configuration: mcp.json ``` { "mcpServers": { "api7-petstore-mcp": { "url": "http://123.123.123.123:9080/mcp", "headers": { "apikey": "john-key" } } } } ``` The configured headers will be added to both GET and POST requests. If the configuration is successful, you should see the available tools (external functions or services exposed to AI clients through MCP). You can then interact with Petstore directly from the chat window of your AI client. If the authentication header is not configured in `mcp.json`, the AI client will be unable to load tools from the MCP server. #### If Headers Are Not Supported[​](#if-headers-are-not-supported "Direct link to If Headers Are Not Supported") If your AI client does not support configuring headers in `mcp.json`, you can include the authentication credential in the MCP URL query, since `key-auth` supports obtaining credential from the URL query. Update the `key-auth` configuration on the route as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes/openapi-to-mcp-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "key-auth": { "_meta": { "filter": [ [ "request_method", "==", "GET" ] ] }, "query": "apikey" } } }' ``` adc.yaml ``` # other config # ... services: - name: openapi-to-mcp-service upstream: type: roundrobin scheme: https pass_host: node nodes: - host: petstore3.swagger.io port: 443 weight: 1 routes: - name: openapi-to-mcp-route uris: - /mcp methods: - GET - POST plugins: key-auth: _meta: filter: - - request_method - "==" - GET query: apikey openapi-to-mcp: transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update the PluginConfig: openapi-to-mcp-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: openapi-to-mcp-plugin-config spec: plugins: - name: key-auth config: _meta: filter: - - request_method - "==" - GET query: apikey - name: openapi-to-mcp config: transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" ``` Apply the updated configuration to your cluster: ``` kubectl apply -f openapi-to-mcp-ic.yaml ``` Update the ApisixRoute: openapi-to-mcp-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: openapi-to-mcp-route spec: ingressClassName: apisix http: - name: openapi-to-mcp-route match: paths: - /mcp methods: - GET - POST plugins: - name: key-auth enable: true config: _meta: filter: - - request_method - "==" - GET query: apikey - name: openapi-to-mcp enable: true config: transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" ``` Apply the updated configuration to your cluster: ``` kubectl apply -f openapi-to-mcp-ic.yaml ``` ❶ Only apply the `key-auth` on GET requests. This is because the `apikey` configured in the query parameter is only sent with the GET request to the SSE endpoint. It is not included in the subsequent POST message requests. As a result, the message requests will be blocked by the `key-auth` plugin if the filter is not applied. ❷ Configure the plugin to obtain the authentication key from the query. After applying the Admin API, ADC, or APISIX CRD configuration, include the credential in the API7 Gateway address query parameter: mcp.json ``` { "mcpServers": { "api7-petstore-mcp": { "url": "http://123.123.123.123:9080/mcp?apikey=john-key" } } } ``` If the configuration is successful, you should see the available tools (external functions or services exposed to AI clients through MCP). You can then interact with Petstore directly from the chat window of your AI client. If the authentication credential is not configured in the MCP server URL query, the AI client will be unable to load tools from the MCP server. ### Pass Dynamic Headers to Upstream[​](#pass-dynamic-headers-to-upstream "Direct link to Pass Dynamic Headers to Upstream") When the upstream API requires per-request credentials or context that differ between MCP clients — such as user-specific API tokens, tenant identifiers, or session IDs — you can pass them dynamically from the MCP client to the upstream using the `x-openapi2mcp-header-*` convention. Any HTTP header sent to the gateway that matches the pattern `x-openapi2mcp-header-{name}` is extracted by the OpenAPI-to-MCP sidecar, stripped of the prefix, and forwarded to the upstream API as the header `{name}`. For example, a header `x-openapi2mcp-header-my-token: abc123` in the client request becomes `my-token: abc123` in the upstream API request. #### How It Works[​](#how-it-works "Direct link to How It Works") The gateway plugin and the OpenAPI-to-MCP sidecar work together to forward headers: 1. **Plugin-level headers**: Headers configured in the plugin `headers` field are resolved at the gateway and forwarded as `x-openapi2mcp-header-{name}` to the sidecar. These headers are shared across all clients, but their values can vary per request when using [built-in variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md). 2. **Client-level headers** (dynamic): Headers set by the MCP client in `mcp.json` using the `x-openapi2mcp-header-*` prefix are passed through the gateway to the sidecar, then forwarded to the upstream. These can vary per client. When both static plugin headers and dynamic client headers are present, they are merged. If a client header has the same name as a plugin header, the plugin header takes precedence and the client value is ignored. #### Transport-Specific Behavior[​](#transport-specific-behavior "Direct link to Transport-Specific Behavior") The behavior of dynamic headers depends on the transport method configured in the plugin: * **`streamable_http` (recommended)**: Every MCP request is independent and stateless. The sidecar reads `x-openapi2mcp-header-*` headers on each request, so dynamic headers are truly per-request. This is the recommended transport for dynamic header passthrough. * **`sse`**: The `x-openapi2mcp-header-*` headers are only read during the initial `GET` request that establishes the SSE connection. Subsequent POST requests within the same session do not re-read these headers. As a result, dynamic headers are fixed for the entire session and cannot be changed mid-session. If your use case requires different header values across requests (for example, per-user tokens that change), use `streamable_http` transport. #### Configure Client Headers[​](#configure-client-headers "Direct link to Configure Client Headers") If your MCP client supports custom headers (such as Cursor or Claude Desktop), add `x-openapi2mcp-header-*` entries to the `headers` field in `mcp.json`: mcp.json ``` { "mcpServers": { "my-api-mcp": { "url": "http://123.123.123.123:9080/mcp", "headers": { "x-openapi2mcp-header-authorization": "Bearer ", "x-openapi2mcp-header-x-tenant-id": "tenant-42" } } } } ``` When the MCP client sends a `tools/call` request, the sidecar extracts these headers and forwards them to the upstream API as: ``` authorization: Bearer x-tenant-id: tenant-42 ``` #### Header Name Mapping[​](#header-name-mapping "Direct link to Header Name Mapping") HTTP infrastructure (such as Nginx and Fastify) normalizes header names to lowercase. As a result, the header name extracted after the `x-openapi2mcp-header-` prefix is always lowercase in the upstream request. The following table summarizes the mapping: | Client header | Upstream header | | ------------------------------------ | --------------- | | `x-openapi2mcp-header-authorization` | `authorization` | | `x-openapi2mcp-header-x-api-key` | `x-api-key` | | `x-openapi2mcp-header-my-token` | `my-token` | note The `x-openapi2mcp-header-*` headers are consumed by the sidecar and are not forwarded to the upstream as-is. Only the extracted header names and values are sent to the upstream. #### Security Considerations[​](#security-considerations "Direct link to Security Considerations") Any `x-openapi2mcp-header-*` header sent by the MCP client is forwarded to the upstream API after prefix stripping. This means clients can inject arbitrary headers into upstream requests. To mitigate risks: * Use gateway-level authentication plugins (such as `key-auth` or `jwt-auth`) to restrict access to the MCP route, ensuring only authorized clients can send requests. * If the upstream API relies on specific headers for authentication or authorization, ensure those headers are set in the plugin-level `headers` configuration rather than relying on client-provided values, since plugin-level headers take precedence over client-level headers. ### Flatten Tool Schema Parameters[​](#flatten-tool-schema-parameters "Direct link to Flatten Tool Schema Parameters") The following example demonstrates how `flatten_parameters` affects the structure of query and path parameters in the generated MCP tool input schema. Complete the [previous example](#enable-mcp-access-to-petstore-apis) using Admin API, ADC, or APISIX CRD to set up MCP access to the Petstore APIs. Although the configuration does not explicitly set `flatten_parameters`, the parameter defaults to `false`. In your AI client, such as Cursor, inspect the tool input schema. You should see parameters nested under `pathParameters` and `queryParameters`: ``` { "operations": { ..., "getPetById": { "method": "GET", "path": "/pet/{petId}", "pathParameters": { "type": "object", "required": ["petId"], "properties": { "petId": { "type": "integer", "description": "ID of pet to return" } }, "additionalProperties": false } }, "findPetsByStatus": { "method": "GET", "path": "/pet/findByStatus", "queryParameters": { "type": "object", "properties": { "status": { "type": "string", "enum": ["available", "pending", "sold"], "description": "Status values that need to be considered for filter", "default": "available" } }, "additionalProperties": false } } } } ``` Update the plugin to flatten query and path parameters: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes/openapi-to-mcp-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "openapi-to-mcp": { "flatten_parameters": true } } }' ``` adc.yaml ``` services: - name: openapi-to-mcp-service upstream: type: roundrobin scheme: https pass_host: node nodes: - host: petstore3.swagger.io port: 443 weight: 1 routes: - name: openapi-to-mcp-route uris: - /mcp methods: - GET - POST plugins: openapi-to-mcp: flatten_parameters: true transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update the PluginConfig: openapi-to-mcp-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: openapi-to-mcp-plugin-config spec: plugins: - name: openapi-to-mcp config: flatten_parameters: true transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" ``` Apply the updated configuration to your cluster: ``` kubectl apply -f openapi-to-mcp-ic.yaml ``` Update the ApisixRoute: openapi-to-mcp-ic.yaml ``` # other configs # --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: openapi-to-mcp-route spec: ingressClassName: apisix http: - name: openapi-to-mcp-route match: paths: - /mcp methods: - GET - POST plugins: - name: openapi-to-mcp enable: true config: flatten_parameters: true transport: streamable_http base_url: "https://petstore3.swagger.io/api/v3" headers: Authorization: "special-key" openapi_url: "https://petstore3.swagger.io/api/v3/openapi.json" ``` Apply the updated configuration to your cluster: ``` kubectl apply -f openapi-to-mcp-ic.yaml ``` In your AI client, such as Cursor, inspect the tool input schema. You should see that parameters like `status` are no longer nested under `pathParameters` or `queryParameters`: ``` { "operations": { ..., "getPetById": { "parameters": { "type": "object", "required": ["petId"], "properties": { "petId": { "type": "integer", "description": "ID of pet to return" } }, "additionalProperties": false } }, "findPetsByStatus": { "parameters": { "type": "object", "properties": { "status": { "type": "string", "enum": ["available", "pending", "sold"], "description": "Status values that need to be considered for filter", "default": "available" } }, "additionalProperties": false } } } } ``` ### Customize MCP Tool Annotations[​](#customize-mcp-tool-annotations "Direct link to Customize MCP Tool Annotations") Availability MCP tool annotations are available from API7 Enterprise version 3.9.7. The following example demonstrates how to add MCP tool annotations to OpenAPI operations exposed by the `openapi-to-mcp` plugin. Without these annotations, AI clients only receive the generated tool name, description, and input schema. They cannot reliably tell whether a tool is read-only, destructive, or idempotent, which makes it harder to rank tools correctly and use them safely. This is implemented in the bundled OpenAPI-to-MCP converter in two ways: 1. It infers default tool behavior from the HTTP method. 2. It reads explicit operation-level configuration from the OpenAPI vendor extension `x-mcp-annotations`. When both are present, explicit `x-mcp-annotations` values override the inferred defaults. Complete the [previous example](#enable-mcp-access-to-petstore-apis) using Admin API, ADC, or APISIX CRD to expose an OpenAPI document through the `openapi-to-mcp` plugin, then add annotations to the OpenAPI operations: The previous Petstore example uses a public OpenAPI document that you cannot edit directly. To apply `x-mcp-annotations`, host your own OpenAPI document and update the `openapi_url` field in the `openapi-to-mcp` plugin configuration to point to that hosted document. openapi.yaml ``` paths: /users/{id}: get: operationId: getUser summary: Get user information x-mcp-annotations: title: Get User readOnlyHint: true openWorldHint: false delete: operationId: deleteUser summary: Delete a user x-mcp-annotations: title: Delete User destructiveHint: true ``` Supported annotation fields: * `title` * `readOnlyHint` * `destructiveHint` * `idempotentHint` * `openWorldHint` If `x-mcp-annotations` is not configured, the converter still applies default inference rules: * `GET`, `HEAD`, and `OPTIONS` map to `readOnlyHint: true` * `DELETE` maps to `destructiveHint: true` and `idempotentHint: true` * `PUT` maps to `idempotentHint: true` After updating the hosted OpenAPI document, ask your MCP client to list tools. The following snippet shows the `result.tools` portion of the `tools/list` response: ``` { "tools": [ { "name": "getUser", "annotations": { "title": "Get User", "readOnlyHint": true, "openWorldHint": false } }, { "name": "deleteUser", "annotations": { "title": "Delete User", "destructiveHint": true, "idempotentHint": true } } ] } ``` Notes: * Only operation-level `x-mcp-annotations` is supported. * Invalid values and unsupported fields are ignored. * `summary` and `description` still control the generated tool description. * `title` is only read from `x-mcp-annotations.title`. ### Enable MCP Access to API7 Enterprise APIs[​](#enable-mcp-access-to-api7-enterprise-apis "Direct link to Enable MCP Access to API7 Enterprise APIs") The following example demonstrates how to expose API7 Enterprise APIs through the MCP protocol, enabling AI models and clients to interact with your API7 Enterprise configuration. Create a route with the `openapi-to-mcp` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "openapi-to-mcp-route", "uri": "/mcp", "methods": ["GET", "POST"], "plugins": { "openapi-to-mcp": { "transport": "streamable_http", "base_url": "https://your-dashboard.com", "headers": { "X-API-KEY": "" }, "openapi_url": "https://run.api7.ai/api7-ee/openapi-latest.json" } } }' ``` adc.yaml ``` services: - name: openapi-to-mcp-service upstream: type: roundrobin scheme: https pass_host: node nodes: - host: your-dashboard.com port: 443 weight: 1 routes: - name: openapi-to-mcp-route uris: - /mcp methods: - GET - POST plugins: openapi-to-mcp: transport: streamable_http base_url: "https://your-dashboard.com" headers: X-API-KEY: "" openapi_url: "https://run.api7.ai/api7-ee/openapi-latest.json" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD openapi-to-mcp-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: openapi-to-mcp-plugin-config spec: plugins: - name: openapi-to-mcp config: transport: streamable_http base_url: "https://your-dashboard.com" headers: X-API-KEY: "" openapi_url: "https://run.api7.ai/api7-ee/openapi-latest.json" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: openapi-to-mcp-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /mcp method: GET - path: type: Exact value: /mcp method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: openapi-to-mcp-plugin-config backendRefs: - name: api7-enterprise-external-domain port: 443 --- apiVersion: v1 kind: Service metadata: namespace: aic name: api7-enterprise-external-domain spec: type: ExternalName externalName: your-dashboard.com --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: api7-enterprise-external-domain spec: targetRefs: - group: "" kind: Service name: api7-enterprise-external-domain passHost: node scheme: https ``` Apply the configuration to your cluster: ``` kubectl apply -f openapi-to-mcp-ic.yaml ``` openapi-to-mcp-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: openapi-to-mcp-route spec: ingressClassName: apisix http: - name: openapi-to-mcp-route match: paths: - /mcp methods: - GET - POST plugins: - name: openapi-to-mcp enable: true config: transport: streamable_http base_url: "https://your-dashboard.com" headers: X-API-KEY: "" openapi_url: "https://run.api7.ai/api7-ee/openapi-latest.json" upstreams: - name: api7-enterprise-external-domain --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: api7-enterprise-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: your-dashboard.com port: 443 passHost: node scheme: https ``` Apply the configuration to your cluster: ``` kubectl apply -f openapi-to-mcp-ic.yaml ``` ❶ Configure the route to allow GET and POST methods. The GET method enables the tool discovery and response streaming (SSE), while the POST method enables the execution and action capabilities (messages). ❷ Configure the transport method to be `streamable_http` (recommended for production). ❸ Replace with your API7 Enterprise address, where requests will be forwarded. ❹ Replace with your credential in the `X-API-KEY` header for API7 Enterprise authentication. ❺ Configure the URL of the API7 Enterprise OpenAPI document. In your AI client, such as Cursor, update the MCP settings with your API7 Gateway address and append the previously created route path. For instance: mcp.json ``` { "mcpServers": { "api7-enterprise-mcp": { "url": "http://123.123.123.123:9080/mcp" } } } ``` If the configuration is successful, you should see the available tools (external functions or services exposed to AI clients through MCP). You can now interact with API7 Enterprise directly from the chat window of your AI client. For example, try asking: "How many gateway groups are there in API7 Enterprise?" ![AI client interaction with API7 Enterprise](https://static.api7.ai/uploads/2025/09/19/vUBryakn_cursor.png) ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") To diagnose issues, check the `openapi-to-mcp` error log at `/usr/local/openapi2mcp/error.log` in your gateway container or pod. Note that this log is separate from the gateway’s error log. ### Known Issues[​](#known-issues "Direct link to Known Issues") 1. The error `Cannot use 'in' operator to search for '$ref' in undefined` typically occurs when an OpenAPI v2 document is used in `openapi_url`. The plugin only supports OpenAPI v3 document in `openapi_url`. 2. The plugin has a known parsing issue when handling `oneOf` schemas in OpenAPI v3 document retrieved from `openapi_url`. In this case, the MCP client will be stuck at tool loading. --- ## Static Configurations[​](#static-configurations "Direct link to Static Configurations") By default, the plugin proxies MCP traffic to the OpenAPI-to-MCP service at `127.0.0.1:3000`. The file to update depends on how the gateway is deployed: * Host or Docker * Kubernetes (Helm) For host or Docker deployments, configure the following settings: config.yaml ``` plugin_attr: openapi-to-mcp: port: 4000 ``` Then reload the gateway for static configuration changes to take effect. For Helm deployments, set the following values in the API7 Gateway Helm chart. Keep the plugin attribute port and the chart-managed sidecar port the same. values.yaml ``` openapiToMcp: enabled: true port: 4000 pluginAttrs: openapi-to-mcp: port: 4000 ``` Then apply the values file to the existing gateway release: ``` helm upgrade api7/gateway -n -f values.yaml ``` When changing this value outside Helm, you must also update the OpenAPI-to-MCP service to listen on the same port, otherwise the plugin will fail with a 503. ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * transport string default: `sse` vaild vaule: `sse` or `streamable_http` *** Transport method for client-server communication. The `streamable_http` method is recommended for production deployments, as it supports stateless communication suitable for multiple gateway instances. The `sse` method is stateful and may exhibit unexpected behavior when multiple gateways are deployed. The `streamable_http` transport is available from API7 Enterprise version 3.8.15. * openapi\_url string required *** URL of the OpenAPI specification document that defines the API structure to be exposed through MCP. Note that the plugin supports only OpenAPI Specification (OAS) version 3. OpenAPI v2 (Swagger) is not supported. Additionally, the plugin has a known parsing issue when handling `oneOf` schemas in OpenAPI v3 document retrieved from `openapi_url`. In this case, the MCP client will be stuck at tool loading. The OpenAPI-to-MCP service caches the document by this URL string for 3600 seconds by default. To pick up a changed document immediately, change the URL (for example, a query parameter). See [OpenAPI Document Caching](https://docs.api7.ai/hub/openapi-to-mcp.md#openapi-document-caching). * base\_url string required *** Base URL of the API service where requests will be forwarded. Support [built-in variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md) in values (available from API7 Enterprise version 3.8.19), for example, `https://${http_baseurl}.swagger.io`. * allowed\_hosts array\[string] vaild vaule: Exact host names or wildcard host names such as `api.example.com` and `*.example.com` *** Optional allow-list of hosts that the resolved `base_url` may target. When set, requests whose resolved host is not in the list are rejected with HTTP 400. Available in API7 Enterprise from version 3.9.13. Not available in APISIX yet. * headers object *** Headers to include in requests to the upstream service. Support [built-in variables](https://docs.api7.ai/api7-gateway/reference/built-in-variables.md) in values, for example, `$arg_username-$http_apikey`. * flatten\_parameters boolean default: `false` *** Whether to flatten parameters in the tool schema. Flattening of query and path parameters is available from API7 Enterprise version 3.8.21. Support for header parameters defined in the OpenAPI spec (`in: header`) is available from version 3.9.8. Not available in APISIX yet. If set to `false`, query parameters are nested under `queryParameters` and path parameters are nested under `pathParameters`. From API7 Enterprise version 3.9.8, header parameters are nested under `headerParameters`. If set to `true`, query and path parameters are placed directly under `properties`, and header parameters are also placed directly under `properties` from version 3.9.8. Setting the parameter to `true` simplifies AI model interaction by reducing schema complexity. Keep the parameter at `false` when query, path, and header parameters share the same names, to avoid conflicts. --- # openid-connect The `openid-connect` plugin supports the integration with [OpenID Connect (OIDC)](https://openid.net/connect/) identity providers, such as [Keycloak](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-keycloak.md), [Auth0](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-auth0.md), [Microsoft Entra ID](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-azure-ad.md), [Google](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-google.md), [Amazon Cognito](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-amazon-cognito.md), and [Okta](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-okta.md). It allows APISIX to authenticate clients and obtain their information from the identity provider before allowing or denying their access to upstream protected resources. ## Examples[​](#examples "Direct link to Examples") ### Authorization Code Flow[​](#authorization-code-flow "Direct link to Authorization Code Flow") The authorization code flow is defined in [RFC 6749, Section 4.1](https://datatracker.ietf.org/doc/html/rfc6749#section-4.1). It involves exchanging a temporary authorization code for an access token, and is typically used by confidential and public clients. The following diagram illustrates the interaction between different entities when you implement the authorization code flow:
When an incoming request does not contain an access token in its header nor in an appropriate session cookie, the plugin acts as a relying party and redirects to the authorization server to continue the authorization code flow. After successful authentication, the plugin keeps the token in the session cookie, and subsequent requests will use the token stored in the cookie. See [Implement Authorization Code Grant](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-keycloak.md#implement-authorization-code-grant) for an example to use the `openid-connect` plugin to integrate with Keycloak using the authorization code flow. See [Secure OIDC with PAR and DPoP](https://docs.api7.ai/apisix/how-to-guide/authentication/secure-oidc-with-par-and-dpop.md) for an example to use the `openid-connect` plugin to integrate with Keycloak using PAR, DPoP, PKCE, and `private_key_jwt` client authentication. ### Proof Key for Code Exchange (PKCE)[​](#proof-key-for-code-exchange-pkce "Direct link to Proof Key for Code Exchange (PKCE)") The Proof Key for Code Exchange (PKCE) is defined in [RFC 7636](https://datatracker.ietf.org/doc/html/rfc7636). PKCE enhances the authorization code flow by adding a code challenge and verifier to prevent authorization code interception attacks. The following diagram illustrates the interaction between different entities when you implement the authorization code flow with PKCE:
See [Implement Authorization Code Grant](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-keycloak.md#implement-authorization-code-grant) for an example to use the `openid-connect` plugin to integrate with Keycloak using the authorization code flow with PKCE. ### Authorization Code Flow with PAR and DPoP[​](#authorization-code-flow-with-par-and-dpop "Direct link to Authorization Code Flow with PAR and DPoP") PAR, PKCE, `private_key_jwt`, and DPoP can be combined in one authorization code flow. Configure separate signing keys for client authentication and DPoP. This workflow was introduced in API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0.
In this workflow, the gateway is the DPoP client for token and user-info calls to the authorization server. It does not validate DPoP proofs from external clients calling the protected route. For an APISIX deployment, see [Secure OIDC with PAR and DPoP](https://docs.api7.ai/apisix/how-to-guide/authentication/secure-oidc-with-par-and-dpop.md) for a tested Keycloak example with key generation, configuration, and verification. ### Client Credential Flow[​](#client-credential-flow "Direct link to Client Credential Flow") The client credential flow is defined in [RFC 6749, Section 4.4](https://datatracker.ietf.org/doc/html/rfc6749#section-4.4). It involves clients requesting an access token with its own credentials to access protected resources, typically used in machine to machine authentication and is not on behalf of a specific user. The following diagram illustrates the interaction between different entities when you implement the client credential flow with local JWT verification, such as by configuring `public_key` or `use_jwks`:
See [Implement Client Credentials Grant](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-keycloak.md#implement-client-credentials-grant) for an example to use the `openid-connect` plugin to integrate with Keycloak using the client credentials flow. ### Introspection Flow[​](#introspection-flow "Direct link to Introspection Flow") The introspection flow is defined in [RFC 7662](https://datatracker.ietf.org/doc/html/rfc7662). It involves verifying the validity and details of an access token by querying an authorization server’s introspection endpoint. In this flow, when a client presents an access token to the resource server, the resource server sends a request to the authorization server’s introspection endpoint, which responds with token details if the token is active, including information like token expiration, associated scopes, and the user or client it belongs to. The following diagram illustrates the interaction between different entities when you implement the authorization code flow with token introspection:
See [Implement Client Credentials Grant](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-keycloak.md#implement-client-credentials-grant) for an example to use the `openid-connect` plugin to integrate with Keycloak using the client credentials flow with token introspection. ### Password Flow[​](#password-flow "Direct link to Password Flow") The password flow is defined in [RFC 6749, Section 4.3](https://datatracker.ietf.org/doc/html/rfc6749#section-4.3). It is designed for trusted applications, allowing them to obtain an access token directly using a user’s username and password. In this grant type, the client app sends the user’s credentials along with its own client ID and secret to the authorization server, which then authenticates the user and, if valid, issues an access token. Though efficient, this flow is intended for highly trusted, first-party applications only, as it requires the app to handle sensitive user credentials directly, posing significant security risks if used in third-party contexts. The following diagram illustrates the interaction between different entities when you implement the password flow:
See [Implement Password Grant](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-keycloak.md#implement-password-grant) for an example to use the `openid-connect` plugin to integrate with Keycloak using the password flow. ### Refresh Token Grant[​](#refresh-token-grant "Direct link to Refresh Token Grant") The refresh token grant is defined in [RFC 6749, Section 6](https://datatracker.ietf.org/doc/html/rfc6749#section-6). It enables clients to request a new access token without requiring the user to re-authenticate, using a previously issued refresh token. This flow is typically used when an access token expires, allowing the client to maintain continuous access to resources without user intervention. Refresh tokens are issued along with access tokens in certain OAuth flows and their lifespan and security requirements depend on the authorization server’s configuration. The following diagram illustrates the interaction between different entities when implementing password flow with refresh token flow:
See [Refresh Token](https://docs.api7.ai/apisix/how-to-guide/authentication/set-up-sso-with-keycloak.md#refresh-token) for an example to use the `openid-connect` plugin to integrate with Keycloak using the password flow with token refreshes. ### User Info[​](#user-info "Direct link to User Info") The UserInfo endpoint in OpenID Connect (OIDC) is defined in [OpenID Connect Core 1.0, Section 5.3](https://openid.net/specs/openid-connect-core-1_0.html#UserInfo). It enables clients to retrieve additional claims about an authenticated user by presenting a valid access token. This endpoint is particularly useful for obtaining user profile information, such as name, email, and other attributes, after the user has been authenticated. The data returned by the UserInfo endpoint depends on the scope of the access token and the claims configured by the authorization server. The following diagram illustrates the interaction between different entities when APISIX verifies the user info:
See [Control Access by Examining User Information from External Identity Provider](https://docs.api7.ai/hub/acl.md#control-access-by-examining-user-information-from-external-identity-provider) for an example to use the `openid-connect` plugin to integrate with Keycloak, and implement access control with the Enterprise `acl` plugin based on the user info. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") This section covers a few commonly seen issues when working with this plugin to help you troubleshoot. ### APISIX Cannot Connect to OpenID provider[​](#apisix-cannot-connect-to-openid-provider "Direct link to APISIX Cannot Connect to OpenID provider") If APISIX fails to resolve or cannot connect to the OpenID provider, double check the DNS settings in your configuration file `config.yaml` and modify as needed. ### State Mismatch in an Authorization Callback[​](#state-mismatch-in-an-authorization-callback "Direct link to State Mismatch in an Authorization Callback") A callback can arrive after its authorization state was completed, replayed, or removed. In API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0, a stale `GET` callback outside Multi Auth is redirected to the originally requested URL. This starts a fresh authentication flow. If the identity provider still has an SSO session, the flow can complete without another prompt. Callbacks that use another HTTP method, have no recoverable target URL, or run inside Multi Auth still fail instead of redirecting. ### Identity Provider Temporarily Unavailable[​](#identity-provider-temporarily-unavailable "Direct link to Identity Provider Temporarily Unavailable") In APISIX 3.18.0, an authorization callback containing the OAuth error `temporarily_unavailable` is redirected to the originally requested URL when its state can be verified. This restarts the authentication flow instead of returning `500 Internal Server Error`. Other OAuth errors, a missing or invalid state, and non-`GET` callbacks are not retried automatically. ### No Session State Found[​](#no-session-state-found "Direct link to No Session State Found") If you encounter a `500 internal server error` with the following message in the log when working with [authorization code flow](#authorization-code-flow), there could be a number of reasons. ``` the error request to the redirect_uri path, but there's no session state found ``` #### 1. Incorrect Redirection URI[​](#1-incorrect-redirection-uri "Direct link to 1. Incorrect Redirection URI") A common configuration error is to set `redirect_uri` to the same URI as the route. When a user requests the protected resource, the request directly reaches the redirection URI without a session cookie, which produces the `no session state found` error. Configure `redirect_uri` as a fully qualified URI with a path that matches the route without being identical to the protected request path. For example, if the route `uri` is `/api/v1/*`, set `redirect_uri` to `https://gateway.example.com/api/v1/redirect`. Configure the same URI as an allowed redirect URI in the OpenID provider. If `redirect_uri` is not configured or is a root-relative path beginning with `/`, the gateway constructs the URI from the request scheme and host. A fully qualified URI avoids relying on this request-derived origin. #### 2. Missing Session Secret[​](#2-missing-session-secret "Direct link to 2. Missing Session Secret") When `bearer_only` is `false`, explicitly configure `session.secret`. APISIX rejects the plugin configuration when this field is missing, whether configuration is stored in etcd or loaded from [standalone YAML](https://docs.api7.ai/apisix/production/deployment-modes.md#standalone-mode). Use at least 16 characters. In multi-instance deployments, use the same secret on every gateway instance that needs to read the encrypted session cookies. #### 3. Cookie Not Sent or Absent[​](#3-cookie-not-sent-or-absent "Direct link to 3. Cookie Not Sent or Absent") Check if the [`SameSite`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Set-Cookie#samesitesamesite-value) cookie attribute is properly set (i.e. if your application needs to send the cookie cross sites) to see if this could be a factor that prevents the cookie being saved to the browser's cookie jar or being sent from the browser. #### 4. Upstream Sent Too Big Header[​](#4-upstream-sent-too-big-header "Direct link to 4. Upstream Sent Too Big Header") If you have NGINX sitting in front of APISIX to proxy client traffic, see if you observe the following error in NGINX's `error.log`: ``` upstream sent too big header while reading response header from upstream ``` If so, try adjusting `proxy_buffers`, `proxy_buffer_size`, and `proxy_busy_buffers_size` to larger values. Alternatively, adjust the plugin's `session_contents` parameter to include only the necessary information. For instance, to include only the access token and refresh token, you can configure the plugin as such: ``` { ... "plugins": { "openid-connect": { ..., "session_contents": { "access_token": true } } } } ``` Available options are `id_token`, `user`, `enc_id_token`, and `access_token` (which includes the refresh token). When this field is not configured, everything is included in the session. #### 5. Invalid Client Secret[​](#5-invalid-client-secret "Direct link to 5. Invalid Client Secret") Verify `client_secret` for flows that authenticate to the provider with a shared secret, such as token introspection or an authorization code flow without PKCE. The field is optional for bearer-only local JWT/JWKS validation, non-bearer PKCE, and applicable `private_key_jwt` modes. In a secret-bearing flow, an invalid value causes authentication to fail and no token is stored in the session. The default `introspection_endpoint_auth_method` is `client_secret_basic`, which sends the client credentials in the `Authorization` header. If the provider expects them in the introspection request body, set the method to `client_secret_post`. #### 6. PKCE IdP Configuration[​](#6-pkce-idp-configuration "Direct link to 6. PKCE IdP Configuration") If you are enabling PKCE with the authorization code flow, make sure you have configured the IdP client to use PKCE. For example, in Keycloak, you should configure the PKCE challenge method in the client's advanced settings: ![PKCE keycloak configuration](https://static.api7.ai/uploads/2024/11/04/xvnCNb20_pkce-keycloak-revised.jpeg) --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. This plugin supports referencing parameter values from environment variables using the `env://` prefix, or from a secret manager, such as HashiCorp Vault’s [KV secrets engine](https://developer.hashicorp.com/vault/docs/secrets/kv), using the `secret://` prefix. For more information, see [environment variables in plugin](https://docs.api7.ai/apisix/reference/environment-variables.md#plugins) and [secrets](https://docs.api7.ai/apisix/key-concepts/secrets.md). * client\_id string required *** Client ID. * client\_secret string *** Client secret. The value is encrypted with AES before being stored in etcd. In API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0, the client secret is optional for local JWT verification modes that do not contact the OpenID provider, such as `bearer_only` combined with `public_key` or `use_jwks`. It is also optional when client authentication uses `private_key_jwt`, or when the authorization code flow uses PKCE. It remains required for flows that authenticate to the provider with a client secret, such as token introspection or the authorization code flow without PKCE. * discovery string required *** URL to the well-known discovery document of the OpenID provider, which contains a list of [OP API endpoints](https://samples.auth0.com/.well-known/openid-configuration). The plugin can directly utilize the endpoints from the discovery document. You can also configure these endpoints individually, which takes precedence over the endpoints supplied in the discovery document. * scope string default: `openid` *** OIDC scope that corresponds to information that should be returned about the authenticated user, also known as [claims](https://openid.net/specs/openid-connect-core-1_0.html#StandardClaims). This is used to authorize users with proper permission. The default value is `openid`, the required scope for OIDC to return a `sub` claim that uniquely identifies the authenticated user. Additional scopes can be appended and delimited by spaces, such as `openid email profile`. * required\_scopes array\[string] *** Scopes required for authorization. If any required scope is missing, the plugin rejects the request with `403 Forbidden`. With bearer introspection, the scopes are read from the introspection response. In APISIX 3.18.0, authorization code sessions are also checked: scopes are read from the access token and then the ID token, and the session is rejected if its granted scopes cannot be determined. API7 Enterprise 3.9.18 and 3.10.5 enforce this field on bearer introspection only. * realm string default: `apisix` *** Realm in the[`WWW-Authenticate`](https://www.rfc-editor.org/rfc/rfc6750#section-3) response header returned with a `401 Unauthorized` response due to authentication failure. For example: * If `realm` is set to `apisix-oidc`, the 401 response will include the following header: ``` WWW-Authenticate: Bearer realm="apisix-oidc" ``` * If `realm` is not configured, the 401 response will include the following header: ``` WWW-Authenticate: Bearer realm="apisix" ``` * claim\_validator object *** JWT claim validation configurations. * issuer object *** Claim issuer validation configurations. * valid\_issuers array\[string] *** An array of trusted JWT issuers. If unconfigured, the issuer from the discovery document is used. In APISIX 3.18.0, bearer JWT verification fails closed while discovery is unavailable because no trusted issuer can be established. API7 Enterprise 3.9.18 and 3.10.5 skip issuer validation in that failure case unless `valid_issuers` is configured explicitly. * audience object *** Audience claim validation configurations. * claim string default: `aud` *** Name of the claim that contains the audience. * required boolean default: `false` *** If true, audience claim is required and the name of the claim will be the name defined in `claim`. For instance, suppose `claim_validator` is configured to be the following: ```json { "audience": { "claim": "custom_claim", "required": true } } ``` If the claim `custom_claim` is not present in the request, you will receive a `required audience claim not present` error. * match\_with\_client\_id boolean default: `false` *** If true, require the audience to match the client ID. If the audience is a string, it must exactly match the client ID. If the audience is an array of strings, at least one value must match. In APISIX 3.18.0, this option also rejects a token that omits the audience claim. API7 Enterprise 3.9.18 and 3.10.5 perform the match only when the claim is present; set `required` to `true` there to reject a missing claim. This requirement is stated in the [OpenID Connect specification](https://openid.net/specs/openid-connect-core-1_0-final.html) to ensure that the token is intended for the specific client. * claim\_schema object *** JSON Schema used to validate the claims returned in the OIDC response. For instance, the schema `{"type":"object","properties":{"access_token":{"type":"string"}},"required":["access_token"]}` ensures that the response includes a required string field named `access_token`. Available in APISIX from 3.14.0 and API7 Enterprise from 3.9.2 * bearer\_only boolean default: `false` *** If true, strictly require bearer access token in requests for authentication. * logout\_path string default: `/logout` *** Path to activate the logout. * post\_logout\_redirect\_uri string *** URL to redirect users to after the `logout_path` receive a request to log out. * redirect\_uri string default: `` `${ngx.var.request_uri}/.apisix/redirect` `` *** URI to redirect to after authentication with the OpenID provider. Configure a fully qualified URI with a scheme and host. The path should match the route without being identical to the protected request path. For example, if the route `uri` is `/api/v1/*`, set `redirect_uri` to `https://gateway.example.com/api/v1/redirect`. If `redirect_uri` is not configured or is a root-relative path beginning with `/`, the gateway constructs the URI from the request scheme and host. It preserves an explicit request port and uses forwarded origin headers only when the immediate proxy is included in `apisix.trusted_addresses`. A fully qualified URI avoids relying on a request-derived origin. Configure the same URI as an allowed redirect URI in the OpenID provider. * timeout integer default: `3` vaild vaule: greater than 0 *** Request timeout in seconds. * ssl\_verify boolean default: `true` *** If true, verify the OpenID provider's SSL certificates. The default value changed from `false` to `true` in APISIX 3.16.0 and API7 Enterprise 3.9.8. This is a breaking change. * introspection\_endpoint string *** URL of the [token introspection](https://datatracker.ietf.org/doc/html/rfc7662) endpoint for the OpenID provider used to introspect access tokens. If this is unset, the introspection endpoint presented in the well-known discovery document is used as a fallback. * introspection\_endpoint\_auth\_method string default: `client_secret_basic` *** Authentication method for the token introspection endpoint. The value should be one of the authentication methods specified in the `introspection_endpoint_auth_methods_supported` [authorization server metadata](https://www.rfc-editor.org/rfc/rfc8414.html) as seen in the well-known discovery document, such as `client_secret_basic`, `client_secret_post`, `private_key_jwt`, and `client_secret_jwt`. With the default `client_secret_basic`, client credentials are sent only in the `Authorization` header. Set this field to `client_secret_post` if the identity provider expects them in the introspection request body. * token\_endpoint\_auth\_method string default: `client_secret_basic` *** Authentication method for the token endpoint. The value should be one of the authentication methods specified in the `token_endpoint_auth_methods_supported` [authorization server metadata](https://www.rfc-editor.org/rfc/rfc8414.html) as seen in the well-known discovery document, such as `client_secret_basic`, `client_secret_post`, `private_key_jwt`, and `client_secret_jwt`. If the configured method is not supported by the plugin, it is ignored and the plugin uses the first usable method advertised by the OpenID provider. If the configured method is supported by the plugin, and `token_endpoint_auth_methods_supported` is present but does not include it, token endpoint authentication fails. * client\_rsa\_private\_key string *** Private key used to sign a client assertion JWT. Required when `private_key_jwt` is selected for the token, introspection, or PAR endpoint. The key type must match `client_jwt_assertion_alg`: RSA for `RS*`, or an EC key on the appropriate curve for `ES*`. The value is encrypted with AES before being stored in etcd. * client\_rsa\_private\_key\_id string *** Optional key ID used in the signed client assertion JWT when `private_key_jwt` is selected for an endpoint. * client\_jwt\_assertion\_expires\_in integer default: `60` *** Lifetime in seconds of a client assertion JWT used with `private_key_jwt` or `client_secret_jwt` on the token, introspection, or PAR endpoint. * public\_key string *** Public key used to verify JWT signature id asymmetric algorithm is used. Providing this value to perform token verification will skip token introspection in client credentials flow. You can pass the public key in `-----BEGIN PUBLIC KEY-----\n……\n-----END PUBLIC KEY-----` format. * token\_signing\_alg\_values\_expected string *** Algorithm used for signing JWT, such as `RS256`. * set\_access\_token\_header boolean default: `true` *** If true, set the access token used by the authenticated request in a request header. By default, the `X-Access-Token` header is used. The gateway clears any client-supplied value before setting the token. * access\_token\_in\_authorization\_header boolean default: `false` *** If true and if `set_access_token_header` is also true, set the access token in the `Authorization` header. * accept\_none\_alg boolean default: `false` *** Set to true if the OpenID provider does not sign its ID token, such as when the signature algorithm is set to `none`. * use\_jwks boolean default: `false` *** If true and if `public_key` is not set, use the JWKS to verify JWT signature and skip token introspection in client credentials flow. The JWKS endpoint is parsed from the discovery document. * jwk\_expires\_in integer default: `86400` *** Expiration time for JWK cache in seconds. * jwt\_verification\_cache\_ignore boolean default: `false` *** If true, force re-verification for a bearer token and ignore any existing cached verification results. * cache\_segment string *** Optional name of a cache segment, used to separate and differentiate caches used by token introspection or JWT verification. * use\_pkce boolean default: `false` *** If true, use the Proof Key for Code Exchange (PKCE) for Authorization Code Flow as defined in [RFC 7636](https://datatracker.ietf.org/doc/html/rfc7636). * set\_id\_token\_header boolean default: `true` *** If true and if a validated ID token is available, set base64-encoded decoded claims in the `X-ID-Token` request header. This value is not the signed JWT and cannot be verified against the provider's JWKS. The gateway clears any client-supplied value first. * set\_userinfo\_header boolean default: `true` *** If true and if user info data is available, set the value in the `X-Userinfo` request header. The gateway clears any client-supplied value first. * set\_raw\_id\_token\_header boolean default: `false` *** If true and a raw ID token is available, add the original signed JWT issued by the identity provider to the `X-Raw-ID-Token` request header. The token is persisted in the session so an upstream service can verify it against the provider's JWKS. Available in API7 Enterprise 3.9.17 and 3.10.4, and in APISIX 3.18.0. The raw ID token is a bearer credential. Enable this only for upstreams you trust, and keep the header out of access logs and out of any response returned to the client. * set\_refresh\_token\_header boolean default: `false` *** If true and if a refresh token obtained from the identity provider is available, set the value in the `X-Refresh-Token` request header. The gateway clears any client-supplied value first. * session object *** Session configuration used when `bearer_only` is `false` and the plugin uses Authorization Code flow. * secret string vaild vaule: 16 or more characters *** Key used for session encryption and HMAC operation when `bearer_only` is `false`. When `bearer_only` is `false`, this field is required. This requirement was introduced in API7 Enterprise 3.9.2 and APISIX 3.14.0. API7 Gateway encrypts the value with AES at rest. APISIX encrypts it before etcd storage when `apisix.data_encryption.enable_encrypt_fields` is enabled. * cookie\_name string *** Name of the session cookie. Maps to the lua-resty-session `cookie_name` option. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * cookie\_path string *** Path scope of the session cookie. Maps to the lua-resty-session `cookie_path` option. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * cookie\_domain string *** Domain scope of the session cookie. Maps to the lua-resty-session `cookie_domain` option. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * cookie\_secure boolean *** If true, set the `Secure` attribute on the session cookie. Maps to the lua-resty-session `cookie_secure` option. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * cookie\_http\_only boolean *** If true, set the `HttpOnly` attribute on the session cookie. Maps to the lua-resty-session `cookie_http_only` option. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * cookie\_same\_site string vaild vaule: `Strict`, `Lax`, `None`, or `Default` *** SameSite attribute of the session cookie. Maps to the lua-resty-session `cookie_same_site` option. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * idling\_timeout integer *** Idling timeout in seconds, after which an idle session is regenerated. Maps to the lua-resty-session `idling_timeout` option. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * rolling\_timeout integer *** Rolling timeout in seconds, after which the session is renewed. Maps to the lua-resty-session `rolling_timeout` option. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * absolute\_timeout integer *** Absolute session lifetime in seconds, after which the session expires regardless of activity. Maps to the lua-resty-session `absolute_timeout` option. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. * cookie object *** Cookie configurations. Deprecated and kept for backward compatibility with the lua-resty-session 3.x schema. Use the flat `session.*` options such as `cookie_name` and `absolute_timeout` instead. * lifetime integer default: `3600` *** Cookie lifetime in seconds. Deprecated. Mapped to `absolute_timeout` at runtime when `absolute_timeout` is not set. * storage string default: `cookie` vaild vaule: `cookie` or `redis` *** Session storage backend. When set to `redis`, sessions are stored in Redis instead of cookies. Available in API7 Enterprise from version 3.9.15 and APISIX from version 3.16.0. * redis object *** Redis connection configurations. Required when `storage` is `redis`. Available in API7 Enterprise from version 3.9.15 and APISIX from version 3.16.0. * host string default: `127.0.0.1` *** Redis host. * port integer default: `6379` vaild vaule: greater than or equal to 1 *** Redis port. * username string *** Redis username. * password string *** Redis password. The value is encrypted with AES before being stored in etcd. * database integer default: `0` vaild vaule: greater than or equal to 0 *** Redis database index. * prefix string default: `sessions` *** Prefix for Redis session keys. * ssl boolean default: `false` *** If true, use SSL for the Redis connection. * ssl\_verify boolean default: `true` *** If true, verify the Redis server SSL certificate. * server\_name string *** Server name for TLS SNI when connecting to Redis. * connect\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** Redis connection timeout in milliseconds. * send\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** Redis send timeout in milliseconds. * read\_timeout integer default: `1000` vaild vaule: greater than or equal to 1 *** Redis read timeout in milliseconds. * keepalive\_timeout integer default: `10000` vaild vaule: greater than or equal to 1000 *** Redis keepalive timeout in milliseconds. * unauth\_action string default: `auth` vaild vaule: `auth`, `deny`, or `pass` *** Action for unauthenticated requests. When set to `auth`, redirect to the authentication endpoint of the OpenID provider. When set to `pass`, allow the request without authentication. When set to `deny`, return 401 unauthenticated responses rather than start the authorization code grant flow. * proxy\_opts object *** Configurations for the proxy server that the OpenID provider is behind. * http\_proxy string *** Proxy server address for HTTP requests, such as `http://:`. * https\_proxy string *** Proxy server address for HTTPS requests, such as `http://:`. * http\_proxy\_authorization string *** Default `Proxy-Authorization` header value to be used with `http_proxy`. Can be overridden with custom `Proxy-Authorization` request header. * https\_proxy\_authorization string *** Default `Proxy-Authorization` header value to be used with `https_proxy`. Cannot be overridden with custom `Proxy-Authorization` request header since with HTTPS, the authorization is completed when connecting. * no\_proxy string *** Comma separated list of hosts that should not be proxied. * authorization\_params object *** Additional parameters to send in the request to the authorization endpoint. * renew\_access\_token\_on\_expiry boolean default: `true` *** If true, attempt to silently renew the access token when it expires or if a refresh token is available. If the token fails to renew, redirect user for re-authentication. * access\_token\_expires\_in integer default: `3600` *** Lifetime of the access token in seconds if no `expires_in` attribute is present in the token endpoint response. * refresh\_session\_interval integer *** Time interval to refresh user ID token without re-authentication. In APISIX, when not set, the plugin will not attempt to silently renew. In API7 Gateway, the default value is `900`. * iat\_slack integer default: `120` *** Tolerance of clock skew in seconds with the `iat` claim in an ID token. * introspection\_expiry\_claim string default: `exp` *** Name of the expiry claim, which controls the TTL of the cached and introspected access token. * introspection\_interval integer *** TTL of the cached and introspected access token in seconds. The default value is 0, which means this option is not used and the plugin defaults to use the TTL passed by expiry claim defined in `introspection_expiry_claim`. If `introspection_interval` is larger than 0 and less than the TTL passed by expiry claim defined in `introspection_expiry_claim`, use `introspection_interval`. * introspection\_addon\_headers array\[string] *** Used to append additional header values to the introspection HTTP request. If the specified header does not exist in the original request, header value will not be appended. * accept\_unsupported\_alg boolean default: `true` *** If an ID token uses an expected signing algorithm that the gateway does not support, setting this to true continues without verifying the signature. Setting it to false rejects the token. Set this to false in security-sensitive deployments unless you explicitly accept the risk of an unverified ID token signature. Setting it to false does not add support for additional signing algorithms. In APISIX 3.18.0 and API7 Gateway 3.9.18 and 3.10.5, the ID-token verification path supports `RS256`, `RS512`, `HS256`, and `HS512`. It does not verify `PS*`, `ES*`, or `EdDSA` signatures. * access\_token\_expires\_leeway integer *** Expiration leeway in seconds for access token renewal. When set to a value greater than 0, token renewal will take place the set amount of time before token expiration. This avoids errors in case the access token just expires when arriving to the resource server. * force\_reauthorize boolean default: `false` *** If true, execute the authorization flow even when a token has been cached. * use\_nonce boolean default: `false` *** If true, enable nonce parameter in authorization request. * revoke\_tokens\_on\_logout boolean default: `false` *** If true, notify the authorization server a previously obtained refresh or access token is no longer needed at the revocation endpoint. * session\_contents object *** Content that should be stored in the session, used to minimize the size of the session data. When not set, everything is included in the session. * id\_token boolean *** If true, store the ID token in session. * access\_token boolean *** If true, store the access token and refresh token in session. * enc\_id\_token boolean *** If true, store the encrypted ID token in session. * user boolean *** If true, store the user info in session. * par object *** Pushed Authorization Request (PAR) configuration, as defined in [RFC 9126](https://datatracker.ietf.org/doc/html/rfc9126). With PAR, the gateway sends the authorization request parameters to the identity provider over a back channel and redirects the user agent with only the returned `request_uri`, so the parameters never travel through the browser. Configure PAR through this nested object. Flat options `use_par`, `pushed_authorization_request_endpoint`, and `pushed_authorization_request_endpoint_auth_method` are rejected so the plugin validates the endpoint and authentication method. Available in API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0. * enabled boolean default: `false` *** If true, push the authorization request to the PAR endpoint instead of sending its parameters in the redirect to the authorization endpoint. * endpoint string *** URL of the identity provider's PAR endpoint. When unset, the endpoint advertised by the provider's discovery document is used. * endpoint\_auth\_method string vaild vaule: `client_secret_basic`, `client_secret_post`, `client_secret_jwt`, or `private_key_jwt` *** Client authentication method used on the PAR endpoint. When unset, the method configured for the token endpoint is used. `private_key_jwt` requires `client_rsa_private_key`, and `client_secret_jwt` requires `client_secret`. The PAR request fails if the selected method cannot be used. * dpop object *** Demonstrating Proof-of-Possession (DPoP) configuration, as defined in [RFC 9449](https://datatracker.ietf.org/doc/html/rfc9449). The gateway signs a proof JWT for each token request, binding issued access tokens to the configured key. A stolen token cannot be replayed at a DPoP-protected resource endpoint without that key. When `par.enabled` is also set, the key thumbprint is sent as `dpop_jkt` on the pushed request. The gateway acts as the DPoP client for token and user-info calls to the identity provider. This configuration does not validate DPoP proofs on incoming requests from external API clients. The gateway rejects a token response whose `token_type` is not `DPoP`. It retries the token request once after a `400` or `401` response with a `DPoP-Nonce` header, and the user-info request once after a `401` response with that header. Configure DPoP through this nested object. Flat options `use_dpop`, `dpop_signing_alg`, `dpop_private_key`, and `dpop_public_jwk` are rejected so the plugin's DPoP validation and encrypted-field handling apply. Available in API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0. * enabled boolean default: `false` *** If true, send a DPoP proof JWT with token requests. When enabled, `private_key` and `public_jwk` are both required. * private\_key string *** PEM-encoded private key used to sign the DPoP proof JWT. When Data Plane data encryption is enabled, this field is encrypted at rest. * public\_jwk object *** Public JWK matching `private_key`, embedded in the proof JWT header. It must not contain private key parameters. * signing\_alg string default: `ES256` vaild vaule: `ES256`, `RS256`, or `PS256` *** Algorithm used to sign the DPoP proof JWT. It must match the type of the configured key and, when the discovery document lists supported DPoP algorithms, must be accepted by the identity provider. * client\_jwt\_assertion\_alg string vaild vaule: `HS256`, `HS512`, `RS256`, `RS512`, `ES256`, or `ES512` *** Algorithm used to sign the client assertion JWT when the endpoint authentication method is `client_secret_jwt` or `private_key_jwt`. Use an `HS*` algorithm with `client_secret_jwt` and an `RS*` or `ES*` algorithm with `private_key_jwt`. The algorithm must match `client_rsa_private_key` and, when the discovery document lists supported client assertion algorithms, must be accepted by the identity provider. One configured algorithm is used for all endpoints. Available in API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0. * client\_jwt\_assertion\_audience string *** Audience claim of the client assertion JWT. When unset, the URL of the endpoint being called is used. Configure this when the gateway reaches an internal endpoint URL but the identity provider expects its external URL as the audience. Available in API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0. --- # OpenTelemetry The `opentelemetry` plugin instruments APISIX and sends traces to OpenTelemetry collector based on the [OpenTelemetry specification](https://opentelemetry.io/docs/reference/specification/), in binary-encoded [OTLP over HTTP](https://opentelemetry.io/docs/reference/specification/protocol/otlp/#otlphttp). ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can work with the `opentelemetry` plugin for different scenarios. ### Enable `opentelemetry` Plugin[​](#enable-opentelemetry-plugin "Direct link to enable-opentelemetry-plugin") In API7 Gateway, `opentelemetry` is available in Dashboard and Admin API by default. For APISIX deployments, load the plugin in the gateway static configuration before configuring routes that use it. * Host or Docker * Kubernetes (Helm) For APISIX host or Docker deployments, keep the existing plugin list in `config.yaml` and add `opentelemetry`: config.yaml ``` plugins: # Keep the complete plugin list used by your gateway. - opentelemetry ``` Reload the gateway for changes to take effect. For the APISIX Helm chart, `apisix.plugins` replaces the loaded plugin list. Start from the complete plugin list used by your gateway and add `opentelemetry`: values.yaml ``` apisix: plugins: # Keep the complete plugin list used by your gateway. - opentelemetry ``` API7 Gateway Helm deployments do not require a Helm values change in this section. Continue with the plugin metadata and route configuration. Apply the values file with the APISIX Helm chart: ``` helm upgrade -n -f values.yaml ``` ### Send Traces to OpenTelemetry[​](#send-traces-to-opentelemetry "Direct link to Send Traces to OpenTelemetry") The following example demonstrates how to trace requests to a route and send traces to OpenTelemetry. Start an OpenTelemetry collector instance: * Docker * Kubernetes ``` docker run -d --name otel-collector -p 4318:4318 otel/opentelemetry-collector-contrib ``` otel-collector.yaml ``` apiVersion: v1 kind: ConfigMap metadata: namespace: aic name: otel-collector-config data: config.yaml: | receivers: otlp: protocols: http: endpoint: 0.0.0.0:4318 exporters: debug: verbosity: detailed service: pipelines: traces: receivers: [otlp] exporters: [debug] --- apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: otel-collector spec: replicas: 1 selector: matchLabels: app: otel-collector template: metadata: labels: app: otel-collector spec: containers: - name: otel-collector image: otel/opentelemetry-collector-contrib args: - "--config=/conf/config.yaml" ports: - containerPort: 4318 volumeMounts: - name: config mountPath: /conf volumes: - name: config configMap: name: otel-collector-config --- apiVersion: v1 kind: Service metadata: namespace: aic name: otel-collector spec: selector: app: otel-collector ports: - name: otlp-http port: 4318 targetPort: 4318 type: ClusterIP ``` Apply the manifest: ``` kubectl apply -f otel-collector.yaml ``` The collector should start listening on `127.0.0.1:4318` (Docker) or `otel-collector.aic.svc.cluster.local:4318` (Kubernetes). Configure the plugin metadata to set the collector address: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/opentelemetry" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "collector": { "address": "127.0.0.1:4318" } }' ``` adc.yaml ``` plugin_metadata: - name: opentelemetry collector: address: "127.0.0.1:4318" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Update the `pluginMetadata` field in your existing `GatewayProxy` resource: gateway-proxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # your control plane connection configuration # .... pluginMetadata: opentelemetry: collector: address: "otel-collector.aic.svc.cluster.local:4318" ``` Apply the configuration to your cluster: ``` kubectl apply -f gateway-proxy.yaml ``` Create a route with `opentelemetry` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "otel-tracing-route", "uri": "/anything", "plugins": { "opentelemetry": { "sampler": { "name": "always_on" } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: otel-tracing-route plugins: opentelemetry: sampler: name: always_on upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD otel-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: otel-plugin-config spec: plugins: - name: opentelemetry config: sampler: name: always_on --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: otel-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: otel-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` otel-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: otel-route spec: ingressClassName: apisix http: - name: otel-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: opentelemetry enable: true config: sampler: name: always_on ``` Apply the configuration to your cluster: ``` kubectl apply -f otel-ic.yaml ``` Send a request to the route: ``` curl "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. In OpenTelemetry collector's log, you should see information similar to the following: ``` 2024-02-18T17:14:03.825Z info ResourceSpans #0 Resource SchemaURL: Resource attributes: -> telemetry.sdk.language: Str(lua) -> telemetry.sdk.name: Str(opentelemetry-lua) -> telemetry.sdk.version: Str(0.1.1) -> hostname: Str(e34673e24631) -> service.name: Str(APISIX) ScopeSpans #0 ScopeSpans SchemaURL: InstrumentationScope opentelemetry-lua Span #0 Trace ID : fbd0a38d4ea4a128ff1a688197bc58b0 Parent ID : ID : af3dc7642104748a Name : GET /anything Kind : Server Start time : 2024-02-18 17:14:03.763244032 +0000 UTC End time : 2024-02-18 17:14:03.920229888 +0000 UTC Status code : Unset Status message : Attributes: -> net.host.name: Str(127.0.0.1) -> http.method: Str(GET) -> http.scheme: Str(http) -> http.target: Str(/anything) -> http.user_agent: Str(curl/7.64.1) -> apisix.route_id: Str(otel-tracing-route) -> apisix.route_name: Empty() -> apisix.response_source: Str(upstream) -> http.route: Str(/anything) -> http.status_code: Int(200) {"kind": "exporter", "data_type": "traces", "name": "debug"} ``` To visualize these traces, you can export your telemetry to backend services, such as Zipkin and Prometheus. See [exporters](https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/exporter) for more details. In API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0, each request span includes an `apisix.response_source` attribute that classifies the origin of the HTTP response: * `apisix` — the response was generated by APISIX itself, such as a plugin rejection, authentication failure, or route-not-found error. * `nginx` — the response was generated by the NGINX proxy layer, such as a connection refused or upstream timeout error. * `upstream` — the response came from the actual upstream service. This attribute enables more precise error attribution in trace analysis, for example, distinguishing gateway-side rejections from real upstream errors. ### Using Trace Variables in Logging[​](#using-trace-variables-in-logging "Direct link to Using Trace Variables in Logging") The following example demonstrates how to configure the `opentelemetry` plugin to set the following built-in variables, which can be used in logger plugins or access logs: * `opentelemetry_context_traceparent`: [trace parent](https://www.w3.org/TR/trace-context/#trace-context-http-headers-format) ID * `opentelemetry_trace_id`: trace ID of the current span * `opentelemetry_span_id`: span ID of the current span Configure the plugin metadata to set `set_ngx_var` as true: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/opentelemetry" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "set_ngx_var": true }' ``` adc.yaml ``` plugin_metadata: - name: opentelemetry set_ngx_var: true ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Update the `pluginMetadata` field in your existing `GatewayProxy` resource and keep the collector configuration: gateway-proxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # your control plane connection configuration # .... pluginMetadata: opentelemetry: collector: address: "otel-collector.aic.svc.cluster.local:4318" set_ngx_var: true ``` Apply the configuration to your cluster: ``` kubectl apply -f gateway-proxy.yaml ``` After the OpenTelemetry collector is available, configure the gateway according to how it was deployed. * Host or Docker * Kubernetes (Helm) Add or update this section in the gateway configuration file to use the `opentelemetry` plugin variables: config.yaml ``` nginx_config: http: enable_access_log: true access_log_format: '{"time": "$time_iso8601","opentelemetry_context_traceparent": "$opentelemetry_context_traceparent","opentelemetry_trace_id": "$opentelemetry_trace_id","opentelemetry_span_id": "$opentelemetry_span_id","remote_addr": "$remote_addr"}' access_log_format_escape: json ``` ❶ `access_log_format`: customize the access log format to use the `opentelemetry` plugin variables. Reload the gateway for configuration changes to take effect. For Helm deployments, update the values that render the gateway access log format. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` apisix: nginx: logs: enableAccessLog: true accessLogFormat: '{"time": "$time_iso8601","opentelemetry_context_traceparent": "$opentelemetry_context_traceparent","opentelemetry_trace_id": "$opentelemetry_trace_id","opentelemetry_span_id": "$opentelemetry_span_id","remote_addr": "$remote_addr"}' accessLogFormatEscape: json ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` logs: enableAccessLog: true accessLogFormat: '{"time": "$time_iso8601","opentelemetry_context_traceparent": "$opentelemetry_context_traceparent","opentelemetry_trace_id": "$opentelemetry_trace_id","opentelemetry_span_id": "$opentelemetry_span_id","remote_addr": "$remote_addr"}' accessLogFormatEscape: json ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` You should see access log entries similar to the following when you generate requests: ``` {"time": "18/Feb/2024:15:09:00 +0000","opentelemetry_context_traceparent": "00-fbd0a38d4ea4a128ff1a688197bc58b0-8f4b9d9970a02629-01","opentelemetry_trace_id": "fbd0a38d4ea4a128ff1a688197bc58b0","opentelemetry_span_id": "af3dc7642104748a","remote_addr": "172.10.0.1"} ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * sampler object *** Sampling configuration. * name string default: `always_off` vaild vaule: `always_on`, `always_off`, `trace_id_ratio`, or `parent_base` *** Sampling strategy. To always sample, use `always_on`. To never sample, use `always_off`. To randomly sample based on a given ratio, use `trace_id_ratio`. To use to sampling decision of the span’s parent, use `parent_base`. If there is no parent, use the root sampler. * options object *** Parameters for sampling strategy. * fraction number default: `0` vaild vaule: between 0 and 1 inclusive *** Sampling ratio when the sampling strategy is `trace_id_ratio`. * root object *** Root sampler when the sampling strategy is `parent_base` strategy. * name string default: `always_off` vaild vaule: `always_on`, `always_off`, or `trace_id_ratio` *** Root sampling strategy. * options object *** Root sampling strategy parameters. * fraction number default: `0` vaild vaule: between 0 and 1 inclusive *** Root sampling ratio when the sampling strategy is `trace_id_ratio`. * additional\_attributes array\[string] *** Names of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) to append to the trace span as string attributes. Values are resolved in the log phase, so variables populated late in request processing are available. Numeric and boolean values are converted to strings; boolean `false` is retained. * additional\_header\_prefix\_attributes array\[string] *** Headers or header prefixes appended to the trace span as string attributes in the log phase. For example, use `x-my-header` or `x-my-headers-*` to include all headers with the prefix `x-my-headers-`. Multiple values for one header are joined with `,`. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * tracing boolean default: `true` *** Global flag to enable or disable tracing. When set to `false`, no spans will be emitted regardless of per-route sampler configuration. * trace\_id\_source string default: `random` vaild vaule: `x-request-id` or `random` *** Source of the trace ID. When set to `x-request-id`, the value of the `x-request-id` header will be used as the trace ID. * resource object *** Additional resource to append to the trace, for example, `{"service_name": "APISIX"}`. * collector object *** Collector configurations. * address string default: `127.0.0.1:4318` *** Address of the OpenTelemetry collector to send traces to. * request\_timeout integer default: `3` *** Request timeout to OpenTelemetry collector in seconds. * request\_headers object *** Request header to include in requests to OpenTelemetry collector, such as `{"Authorization": "token"}`. * batch\_span\_processor object *** Batch span processor configurations. * drop\_on\_queue\_full boolean *** If true, drop span when the queue is full, otherwise force process batches. * max\_queue\_size integer *** Maximum queue size to buffer spans for delayed processing. * batch\_timeout number *** Timeout for span batches to wait in the export queue before being sent, in seconds. * inactive\_timeout number *** Timeout for spans to wait in the export queue before being sent, if the queue is not full, in seconds. * max\_export\_batch\_size integer *** Maximum number of spans to include in a single batch sent to the OpenTelemetry collector. * set\_ngx\_var boolean default: `false` *** Export `opentelemetry` variables to [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). --- # Prometheus The `prometheus` plugin provides the capability to integrate APISIX with [Prometheus](https://prometheus.io). After enabling the plugin, APISIX will start collecting relevant metrics, such as API requests and latencies, and exporting them in a [text-based exposition format](https://prometheus.io/docs/instrumenting/exposition_formats/#exposition-formats) to Prometheus. You can then create event monitoring and alerting in Prometheus to monitor the health of your API gateway and APIs. ## Metrics[​](#metrics "Direct link to Metrics") There are different types of metrics in Prometheus. To understand their differences, see [metrics types](https://prometheus.io/docs/concepts/metric_types/). The following metrics are exported by the `prometheus` plugin by default. See [get APISIX metrics](#get-apisix-metrics) for an example. Note that some metrics, such as `apisix_batch_process_entries`, are not readily visible if there are no data. | Name | Type | Description | | ----------------------------------------- | --------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | apisix\_bandwidth | counter | Total amount of traffic flowing through APISIX in bytes. | | apisix\_etcd\_modify\_indexes | gauge | Number of changes to etcd by APISIX keys. | | apisix\_batch\_process\_entries | gauge | Number of remaining entries in a batch when sending data in batches, such as with `http logger`, and other logging plugins. | | apisix\_etcd\_reachable | gauge | Whether APISIX can reach etcd. A value of `1` represents reachable and `0` represents unreachable. | | apisix\_http\_status | counter | HTTP status codes returned to clients. This is the status the client receives after plugins and proxying, which can differ from the upstream status. | | apisix\_http\_requests\_total | gauge | Number of HTTP requests from clients. | | apisix\_nginx\_http\_current\_connections | gauge | Number of current connections with clients. | | apisix\_nginx\_metric\_errors\_total | counter | Total number of `nginx-lua-prometheus` errors. | | apisix\_http\_latency | histogram | HTTP request latency in milliseconds. | | apisix\_node\_info | gauge | Information about the APISIX node, such as the host name and APISIX version. | | apisix\_shared\_dict\_capacity\_bytes | gauge | The total capacity of an [NGINX shared dictionary](https://github.com/openresty/lua-nginx-module#ngxshareddict). | | apisix\_shared\_dict\_free\_space\_bytes | gauge | The remaining space in an [NGINX shared dictionary](https://github.com/openresty/lua-nginx-module#ngxshareddict). | | apisix\_upstream\_status | gauge | Health check status of upstream nodes, available if health checks are configured on the upstream. A value of `1` represents healthy and `0` represents unhealthy. | | apisix\_stream\_connection\_total | counter | Total number of connections handled per stream route. | | apisix\_stream\_active\_connections | gauge | Number of active TCP connections and UDP sessions for each stream listening address. Requires APISIX-Runtime. Introduced in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line, and in APISIX 3.18.0. | | apisix\_stream\_status | counter | Number of completed stream sessions by termination status, listening address, and upstream node. Introduced in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line, and in APISIX 3.18.0. | | apisix\_stream\_bandwidth | counter | Number of bytes proxied by the stream subsystem by listening address, direction, and side. Requires APISIX-Runtime. Introduced in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line, and in APISIX 3.18.0. | | apisix\_llm\_prompt\_tokens | counter | Number of prompt tokens. Only exported for AI request types. Introduced in API7 Enterprise 3.9.7 and APISIX 3.17.0. | | apisix\_llm\_completion\_tokens | counter | Number of completion tokens. Only exported for AI request types. Introduced in API7 Enterprise 3.9.7 and APISIX 3.17.0. | | apisix\_llm\_latency | histogram | LLM request latency in milliseconds. Only exported for AI request types. Introduced in API7 Enterprise 3.9.7 and APISIX 3.17.0. The `type` label distinguishes total response latency (`total`) from time to first token (`ttft`). Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. Queries that omit `type` match both observations; select `type="total"` to keep the previous total-latency meaning. Each streaming request records one `total` sample and one `ttft` sample. | | apisix\_llm\_active\_connections | gauge | Number of in-flight LLM upstream requests. Only exported for AI request types. Introduced in API7 Enterprise 3.9.7 and APISIX 3.17.0. | | apisix\_llm\_prompt\_tokens\_dist | histogram | Distribution of prompt tokens per request. Only exported for AI request types. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. | | apisix\_llm\_completion\_tokens\_dist | histogram | Distribution of completion tokens per request. Only exported for AI request types. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. | | apisix\_ai\_cache\_hits\_total | counter | Number of requests served from AI Cache, separated by exact or semantic cache layer. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. | | apisix\_ai\_cache\_misses\_total | counter | Number of AI Cache lookups that did not return a cached response. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. | | apisix\_ai\_cache\_bypasses\_total | counter | Number of requests that bypassed AI Cache lookup. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. | | apisix\_ai\_cache\_embedding\_latency | histogram | Latency in milliseconds of embedding-provider calls made by the semantic cache layer. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. | note LLM metrics (`apisix_llm_prompt_tokens`, `apisix_llm_completion_tokens`, `apisix_llm_latency`, `apisix_llm_prompt_tokens_dist`, and `apisix_llm_completion_tokens_dist`) are only exported when the request is processed by an AI plugin (such as [AI Proxy](https://docs.api7.ai/hub/ai-proxy.md)). Routes without AI plugins do not produce these metrics. `apisix_llm_active_connections` is managed directly by AI plugins and is also only present for AI-enabled routes. The gauge increases when an LLM upstream attempt starts and decreases once in the request log phase. For a single attempt, it tracks active LLM upstream requests. With `ai-proxy-multi` fallback retries, each retry increments a new per-instance series, and the request is decremented only once using the final instance labels. A failed instance series can therefore stay above the true active count until the metric expires or the storage is reset. To reduce high-cardinality labels on LLM metrics, use `disabled_labels` in [plugin metadata](https://docs.api7.ai/hub/prometheus/configuration.md#plugin-metadata) to selectively disable labels such as `consumer` or `node`. The prompt- and completion-token histograms use configurable buckets. See the [static plugin attributes](https://docs.api7.ai/hub/prometheus/configuration.md#static-configurations) for their defaults and configuration. ## Labels[​](#labels "Direct link to Labels") [Labels](https://prometheus.io/docs/practices/naming/#labels) are attributes of metrics that are used to differentiate metrics. For example, the `apisix_http_status` metric can be labeled with `route` information to identify which route the HTTP status originates from. The following are labels for a non-exhaustive list of APISIX metrics and their descriptions. ### Labels for `apisix_http_status`[​](#labels-for-apisix_http_status "Direct link to labels-for-apisix_http_status") The following labels are used to differentiate `apisix_http_status` metrics. | Name | Description | | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | code | HTTP response code returned to the client after plugins and proxying. | | route | ID of the route that the HTTP status originates from when `prefer_name` is `false` (default), and name of the route when `prefer_name` to `true`. Default to an empty string if a request does not match any route. | | route\_id | Available only in API7 Enterprise. ID of the route that the HTTP status originates from regardless of the `prefer_name` setting. | | matched\_uri | URI of the route that matches the request. Default to an empty string if a request does not match any route. | | matched\_host | Host of the route that matches the request. Default to an empty string if a request does not match any route, or host is not configured on the route. | | service | ID of the service that the HTTP status originates from when `prefer_name` is `false` (default), and name of the service when `prefer_name` to `true`. Default to the configured value of host on the route if the matched route does not belong to any service. | | service\_id | Available only in API7 Enterprise. ID of the service that the HTTP status originates from regardless of the `prefer_name` setting. | | consumer | Name of the consumer associated with a request. Default to an empty string if no consumer is associated with the request. | | node | IP address of the upstream node. For AI routes, this is the LLM instance name from `ai-proxy` or `ai-proxy-multi`. | | gateway\_group\_id | ID of the gateway group that the HTTP status originates from. Available only in API7 Enterprise. | | instance\_id | ID of the gateway instance that the HTTP status originates from. Available only in API7 Enterprise. | | api\_product\_id | Product ID that the HTTP status originates from. Available only in API7 Enterprise. | | request\_type | Request type associated with the HTTP status. | | request\_llm\_model | Model name sent in the client request. Empty for ordinary HTTP traffic. | | llm\_model | Model the gateway actually targets. If the AI instance configures a model, that value is used; otherwise the client's requested model is used. Empty for ordinary HTTP traffic. | | response\_source | Source of the HTTP response: `apisix` (generated by APISIX, such as plugin rejections or route-not-found), `nginx` (NGINX proxy errors such as connection refused or upstream timeout), or `upstream` (real response from the upstream service). Introduced in API7 Enterprise 3.9.10 and APISIX 3.17.0. | | mcp\_request\_type | MCP request type, such as `tools/list` or `tools/call`. Empty for non-MCP requests. Introduced in API7 Enterprise 3.9.14. | | mcp\_tool\_name | MCP tool name for `tools/call` requests. Empty for other requests. Introduced in API7 Enterprise 3.9.14. | The `response_source` label is mandatory on `apisix_http_status` and can add up to three series per existing status-label combination. Update PromQL joins, recording rules, alerts, and dashboards that match the complete label set, and review the resulting cardinality before rollout. ### Labels for `apisix_bandwidth`[​](#labels-for-apisix_bandwidth "Direct link to labels-for-apisix_bandwidth") The following labels are used to differentiate `apisix_bandwidth` metrics. | Name | Description | | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | type | Type of traffic, `egress` or `ingress`. | | route | ID of the route that bandwidth corresponds to when `prefer_name` is `false` (default), and name of the route when `prefer_name` to `true`. Default to an empty string if a request does not match any route. | | route\_id | Available only in API7 Enterprise. ID of the route that bandwidth corresponds to regardless of the `prefer_name` setting. | | service | ID of the service that bandwidth corresponds to when `prefer_name` is `false` (default), and name of the service when `prefer_name` to `true`. Default to the configured value of host on the route if the matched route does not belong to any service. | | service\_id | Available only in API7 Enterprise. ID of the service that bandwidth corresponds to regardless of the `prefer_name` setting. | | consumer | Name of the consumer associated with a request. Default to an empty string if no consumer is associated with the request. | | node | IP address of the upstream node. For AI routes, this is the LLM instance name from `ai-proxy` or `ai-proxy-multi`. | | gateway\_group\_id | Available only in API7 Enterprise. ID of the gateway group that bandwidth corresponds to. | | instance\_id | Available only in API7 Enterprise. ID of the gateway instance that bandwidth corresponds to. | | api\_product\_id | Available only in API7 Enterprise. Product ID that bandwidth corresponds to. | | request\_type | Request type that bandwidth corresponds to. | | request\_llm\_model | Model name sent in the client request. Empty for ordinary HTTP traffic. | | llm\_model | Model the gateway actually targets. If the AI instance configures a model, that value is used; otherwise the client's requested model is used. Empty for ordinary HTTP traffic. | | mcp\_request\_type | MCP request type, such as `tools/list` or `tools/call`. Empty for non-MCP requests. Introduced in API7 Enterprise 3.9.14 and 3.10.1. | | mcp\_tool\_name | MCP tool name for `tools/call` requests. Empty for other requests. Introduced in API7 Enterprise 3.9.14 and 3.10.1. | ### Labels for `apisix_http_latency`[​](#labels-for-apisix_http_latency "Direct link to labels-for-apisix_http_latency") The following labels are used to differentiate `apisix_http_latency` metrics. | Name | Description | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | type | Type of latencies. See [latency types](#latency-types) for details. | | route | ID of the route that latencies correspond to when `prefer_name` is `false` (default), and name of the route when `prefer_name` to `true`. Default to an empty string if a request does not match any route. | | route\_id | Available only in API7 Enterprise. ID of the route that latencies correspond to regardless of the `prefer_name` setting. | | service | ID of the service that latencies correspond to when `prefer_name` is `false` (default), and name of the service when `prefer_name` to `true`. Default to the configured value of host on the route if the matched route does not belong to any service. | | service\_id | Available only in API7 Enterprise. ID of the service that latencies correspond to regardless of the `prefer_name` setting. | | consumer | Name of the consumer associated with latencies. Default to an empty string if no consumer is associated with the request. | | node | IP address of the upstream node associated with latencies. For AI routes, this is the LLM instance name from `ai-proxy` or `ai-proxy-multi`. | | gateway\_group\_id | Available only in API7 Enterprise. ID of the gateway group that latencies correspond to. | | instance\_id | Available only in API7 Enterprise. ID of the gateway instance that latencies correspond to. | | api\_product\_id | Available only in API7 Enterprise. Product ID that latencies correspond to. | | request\_type | Request type that latencies correspond to. | | request\_llm\_model | Model name sent in the client request. Empty for ordinary HTTP traffic. | | llm\_model | Model the gateway actually targets. If the AI instance configures a model, that value is used; otherwise the client's requested model is used. Empty for ordinary HTTP traffic. | | mcp\_request\_type | MCP request type, such as `tools/list` or `tools/call`. Empty for non-MCP requests. Introduced in API7 Enterprise 3.9.14 and 3.10.1. | | mcp\_tool\_name | MCP tool name for `tools/call` requests. Empty for other requests. Introduced in API7 Enterprise 3.9.14 and 3.10.1. | #### Latency Types[​](#latency-types "Direct link to Latency Types") `apisix_http_latency` can be labeled with one of the three types: * `request` represents the time elapsed between the first byte was read from the client and the log write after the last byte was sent to the client. * `upstream` represents the time elapsed waiting on responses from the upstream service. * `apisix` represents the difference between the `request` latency and `upstream` latency. In other words, the APISIX latency is not only attributed to the Lua processing. It should be understood as follows: ``` APISIX latency = downstream request time - upstream response time = downstream traffic latency + NGINX latency ``` ### Labels for `apisix_upstream_status`[​](#labels-for-apisix_upstream_status "Direct link to labels-for-apisix_upstream_status") The following labels are used to differentiate `apisix_upstream_status` metrics. | Name | Description | | ---- | ------------------------------------------------------------------------------------------------------------------------------ | | name | Resource ID corresponding to the upstream configured with health checks, such as `/apisix/routes/1` and `/apisix/upstreams/1`. | | ip | IP address of the upstream node. | | port | Port number of the node. | ### Labels for Stream Metrics[​](#labels-for-stream-metrics "Direct link to Labels for Stream Metrics") The `apisix_stream_active_connections` metric uses the stream listening address as a label: | Name | Description | | ------------- | ------------------------------------------------------------------------------------------- | | `listen_addr` | Address on which APISIX accepted the TCP connection or UDP session, such as `0.0.0.0:9100`. | The `apisix_stream_status` counter records each stream session when it ends: | Name | Description | | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `code` | How the session ended. `200` is a normal close. `400` is a client-side problem, such as a reset or invalid data from the client. `403` is an access-rule rejection. `500` is an internal error. `502` is an upstream or transport problem, such as a connect failure, reset, or idle timeout. `503` is a connection-limit rejection. | | `listen_addr` | Address on which APISIX accepted the session. | | `node` | Selected upstream address in `IP:port` form. Empty when the session ends before APISIX selects an upstream. | NGINX reports Stream `$status` as 200 for some failures that happen after the upstream connection is established. This metric uses the session termination reason to distinguish a clean close from a later timeout or reset. `200` is also used for worker shutdown and when no recognized reason is recorded. The `apisix_stream_bandwidth` counter updates while a session remains open: | Name | Description | | ------------- | --------------------------------------------- | | `listen_addr` | Address on which APISIX accepted the session. | | `type` | Traffic direction: `ingress` or `egress`. | | `side` | Connection side: `downstream` or `upstream`. | ### Labels for `apisix_llm_latency`[​](#labels-for-apisix_llm_latency "Direct link to labels-for-apisix_llm_latency") The following labels are used to differentiate `apisix_llm_latency` metrics. | Name | Description | | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | type | LLM latency type: `total` for full response latency, or `ttft` for time to first token on streaming requests. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. Dashboards, alerts, and recording rules that previously treated every `apisix_llm_latency` sample as total latency must add `type="total"`. | | route | ID of the route that the HTTP status originates from when `prefer_name` is `false` (default), and name of the route when `prefer_name` to `true`. Default to an empty string if a request does not match any route. | | route\_id | ID of the route that the HTTP status originates from regardless of the `prefer_name` setting. | | service | ID of the service that the HTTP status originates from when `prefer_name` is `false` (default), and name of the service when `prefer_name` to `true`. Default to the configured value of host on the route if the matched route does not belong to any service. | | service\_id | ID of the service that the HTTP status originates from regardless of the `prefer_name` setting. | | consumer | Name of the consumer associated with a request. Default to an empty string if no consumer is associated with the request. | | node | Name of the selected LLM instance from `ai-proxy` or `ai-proxy-multi`, not the upstream IP. | | gateway\_group\_id | ID of the gateway group that the HTTP status originates from. | | instance\_id | ID of the gateway instance that the HTTP status originates from. | | api\_product\_id | Product ID that the HTTP status originates from. | | request\_type | Request type that the HTTP status originates from. | | request\_llm\_model | Model name sent in the client request. Introduced in API7 Enterprise 3.9.7 and APISIX 3.17.0. | | llm\_model | Model the gateway actually targets. If the AI instance configures a model, that value is used; otherwise the client's requested model is used. | ### Labels for Other LLM Metrics[​](#labels-for-other-llm-metrics "Direct link to Labels for Other LLM Metrics") The following labels are used to differentiate `apisix_llm_prompt_tokens`, `apisix_llm_completion_tokens`, `apisix_llm_active_connections`, `apisix_llm_prompt_tokens_dist`, and `apisix_llm_completion_tokens_dist` metrics. | Name | Description | | ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | route | ID of the route that the HTTP status originates from when `prefer_name` is `false` (default), and name of the route when `prefer_name` to `true`. Default to an empty string if a request does not match any route. | | route\_id | ID of the route that the HTTP status originates from regardless of the `prefer_name` setting. | | matched\_uri | URI of the route that matches the request. Default to an empty string if a request does not match any route. | | matched\_host | Host of the route that matches the request. Default to an empty string if a request does not match any route, or host is not configured on the route. | | service | ID of the service that the HTTP status originates from when `prefer_name` is `false` (default), and name of the service when `prefer_name` to `true`. Default to the configured value of host on the route if the matched route does not belong to any service. | | service\_id | ID of the service that the HTTP status originates from regardless of the `prefer_name` setting. | | consumer | Name of the consumer associated with a request. Default to an empty string if no consumer is associated with the request. | | node | Name of the selected LLM instance from `ai-proxy` or `ai-proxy-multi`, not the upstream IP. | | gateway\_group\_id | ID of the gateway group that the HTTP status originates from. | | instance\_id | ID of the gateway instance that the HTTP status originates from. | | api\_product\_id | Product ID that the HTTP status originates from. | | request\_type | Request type that the HTTP status originates from. | | request\_llm\_model | Model name sent in the client request. Introduced in API7 Enterprise 3.9.7 and APISIX 3.17.0. | | llm\_model | Model the gateway actually targets. If the AI instance configures a model, that value is used; otherwise the client's requested model is used. | The prompt- and completion-token counters and distribution histograms use different label sets by product. APISIX exports `route_id`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`. API7 Enterprise additionally exports `route`, `matched_uri`, `matched_host`, and `service`, as well as its gateway-instance and API-product labels. The `request_llm_model` and `llm_model` values originate from client and provider data. Their values are limited to 128 bytes. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. Disable the labels when per-model series are not required. ### Labels for AI Cache Metrics[​](#labels-for-ai-cache-metrics "Direct link to Labels for AI Cache Metrics") The four AI Cache metrics share the following labels. Only `apisix_ai_cache_hits_total` includes `layer`. | Name | Description | | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `layer` | Cache layer that served a hit: `exact` or `semantic`. This structural label cannot be disabled. | | `route` | Route name, or an empty string if the route has no name. | | `route_id` | Route ID. | | `service` | Service name, or an empty string if the route does not reference a service. | | `service_id` | Service ID, or an empty string if the route does not reference a service. | | `consumer` | Consumer name, or an empty string when the request has no Consumer. | | `node` | Name of the selected LLM instance from `ai-proxy` or `ai-proxy-multi`, not the upstream IP. | | `request_type` | Request type, such as `ai_chat` or `ai_stream`. | | `request_llm_model` | Model name sent in the client request. | | `llm_model` | Model the gateway actually targets. If the AI instance configures a model, that value is used; otherwise the client's requested model is used. Empty on cache hits, which are served without reaching the LLM. | API7 Enterprise also exports `matched_uri`, `matched_host`, and its gateway-instance and API-product labels for these metrics. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can work with the `prometheus` plugin for different scenarios. ### Get APISIX Metrics[​](#get-apisix-metrics "Direct link to Get APISIX Metrics") The following example demonstrates how you can get metrics from APISIX. The default Prometheus metrics endpoint and other Prometheus related configurations can be found in the [static configuration](https://docs.api7.ai/hub/prometheus/configuration.md#static-configurations). If you would like to customize these configurations, see [configuration files](https://docs.api7.ai/apisix/reference/configuration-files.md#configyaml-and-configyamlexample). If you deploy the gateway in a containerized environment and would like to access the Prometheus metrics endpoint externally, update the Prometheus export address in the gateway static configuration: * Host or Docker * Kubernetes (Helm) Add or update this section in the gateway configuration file: config.yaml ``` plugin_attr: prometheus: export_addr: ip: 0.0.0.0 ``` Reload the gateway for changes to take effect. For the APISIX Helm chart, enable Prometheus in the chart values. This renders `plugin_attr.prometheus.export_addr.ip` as `0.0.0.0`: values.yaml ``` apisix: prometheus: enabled: true ``` For the API7 Gateway Helm chart, update the Prometheus plugin attributes: values.yaml ``` pluginAttrs: prometheus: export_addr: ip: 0.0.0.0 port: 9091 ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` Send a request to the APISIX Prometheus metrics endpoint: ``` curl "http://127.0.0.1:9091/apisix/prometheus/metrics" ``` You should see an output similar to the following: ``` # HELP apisix_bandwidth Total bandwidth in bytes consumed per service in Apisix # TYPE apisix_bandwidth counter apisix_bandwidth{type="egress",route="",service="",consumer="",node=""} 8417 apisix_bandwidth{type="egress",route="1",service="",consumer="",node="127.0.0.1"} 1420 apisix_bandwidth{type="egress",route="2",service="",consumer="",node="127.0.0.1"} 1420 apisix_bandwidth{type="ingress",route="",service="",consumer="",node=""} 189 apisix_bandwidth{type="ingress",route="1",service="",consumer="",node="127.0.0.1"} 332 apisix_bandwidth{type="ingress",route="2",service="",consumer="",node="127.0.0.1"} 332 # HELP apisix_etcd_modify_indexes Etcd modify index for APISIX keys # TYPE apisix_etcd_modify_indexes gauge apisix_etcd_modify_indexes{key="consumers"} 0 apisix_etcd_modify_indexes{key="global_rules"} 0 ... ``` ### Reduce Metric Cardinality by Disabling Labels[​](#reduce-metric-cardinality-by-disabling-labels "Direct link to Reduce Metric Cardinality by Disabling Labels") Plugin metadata can collapse selected label values to an empty string, reducing the number of time series while preserving the metric's label schema. APISIX and API7 Enterprise use different metadata keys for the HTTP status and latency metrics. The accepted labels also differ where API7 Enterprise exports additional dimensions: | Metric metadata key | APISIX labels that can be disabled | API7 Enterprise differences | | ------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | `http_status` | `route`, `matched_uri`, `matched_host`, `service`, `consumer`, `node`, `request_type`, `request_llm_model`, `llm_model`, `response_source` | Use the key `status`. Adds `route_id`, `service_id`, `mcp_request_type`, and `mcp_tool_name`, but does not allow `response_source` to be disabled. | | `http_latency` | `route`, `service`, `consumer`, `node`, `request_type`, `request_llm_model`, `llm_model` | Use the key `latency`. Adds `route_id`, `service_id`, `mcp_request_type`, and `mcp_tool_name`. | | `bandwidth` | `route`, `service`, `consumer`, `node`, `request_type`, `request_llm_model`, `llm_model` | Adds `route_id`, `service_id`, `mcp_request_type`, and `mcp_tool_name`. | | `llm_latency` | `route_id`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, `llm_model` | Also allows `route` and `service`. | | `llm_prompt_tokens`, `llm_completion_tokens`, `llm_prompt_tokens_dist`, `llm_completion_tokens_dist` | `route_id`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, `llm_model` | Also allows `route`, `matched_uri`, `matched_host`, and `service`. | | `llm_active_connections` | `route`, `route_id`, `matched_uri`, `matched_host`, `service`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, `llm_model` | Same labels. | | `ai_cache_hits_total`, `ai_cache_misses_total`, `ai_cache_bypasses_total`, `ai_cache_embedding_latency` | `route`, `route_id`, `service`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, `llm_model` | Also allows `matched_uri` and `matched_host`. | Structural labels are excluded from this table because they cannot be disabled. In API7 Enterprise, the `stream_status` metadata key disables the `node` label on `apisix_stream_status`. Its `code` and `listen_addr` labels are structural. Introduced in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line. In APISIX, disable the `node` label on HTTP status and latency metrics: ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/prometheus" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "disabled_labels": { "http_status": ["node"], "http_latency": ["node"] } }' ``` For API7 Enterprise, use `status` and `latency` instead: ``` { "disabled_labels": { "status": ["node"], "latency": ["node"] } } ``` Send a request through a route with the plugin enabled, then fetch the metrics endpoint. The affected series should keep the `node` label with an empty value: ``` apisix_http_status{code="200",route="1",matched_uri="/get",matched_host="",service="",consumer="",node="",request_type="traditional_http",request_llm_model="",llm_model="",response_source="upstream"} 1 ``` The schema rejects structural labels that distinguish different measurements, including `code` for HTTP status, `type` for HTTP latency, bandwidth, and LLM latency, and `layer` for AI Cache hits. The exact optional-label set differs between APISIX and API7 Enterprise; use the [plugin metadata reference](https://docs.api7.ai/hub/prometheus/configuration.md#plugin-metadata) for the gateway you are configuring. ### Expose APISIX Metrics on Public API Endpoint[​](#expose-apisix-metrics-on-public-api-endpoint "Direct link to Expose APISIX Metrics on Public API Endpoint") The following example demonstrates how you can disable the Prometheus export server that, by default, exposes an endpoint on port `9091`, and expose APISIX Prometheus metrics on a new public API endpoint on port `9080`, which APISIX uses to listen to other client requests. caution If a large quantity of metrics are being collected, the plugin could take up a significant amount of CPU resources for metric computations and negatively impact the processing of regular requests. To address this issue, APISIX uses [privileged agent](https://github.com/openresty/lua-resty-core/blob/master/lib/ngx/process.md#enable_privileged_agent) and offloads the metric computations to a separate process. This optimization applies automatically if you use the metric endpoint configured in the configuration files, as demonstrated [above](#get-apisix-metrics). If you expose the metric endpoint with the `public-api` plugin, you will not benefit from this optimization. To expose metrics through `public-api`, first disable the default Prometheus export server: * Host or Docker * Kubernetes (Helm) Add or update this section in the gateway configuration file: config.yaml ``` plugin_attr: prometheus: enable_export_server: false ``` Reload the gateway for changes to take effect. For Helm deployments, update the chart values that render `plugin_attr.prometheus`. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` apisix: pluginAttrs: prometheus: enable_export_server: false ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` pluginAttrs: prometheus: enable_export_server: false ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` Next, create a route with [`public-api`](https://docs.api7.ai/hub/public-api.md) plugin and expose a public API endpoint for APISIX metrics: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "prometheus-metrics", "uri": "/prometheus_metrics", "plugins": { "public-api": { "uri": "/apisix/prometheus/metrics" } } }' ``` adc.yaml ``` routes: - uri: /prometheus_metrics name: prometheus-metrics plugins: public-api: uri: /apisix/prometheus/metrics ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD prometheus-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: prometheus-public-api-config spec: plugins: - name: public-api config: uri: /apisix/prometheus/metrics --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: prometheus-metrics-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /prometheus_metrics filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: prometheus-public-api-config ``` prometheus-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: prometheus-metrics-route spec: ingressClassName: apisix http: - name: prometheus-metrics-route match: paths: - /prometheus_metrics plugins: - name: public-api enable: true config: uri: /apisix/prometheus/metrics ``` Apply the configuration to your cluster: ``` kubectl apply -f prometheus-ic.yaml ``` Send a request to the new metrics endpoint to verify: ``` curl "http://127.0.0.1:9080/prometheus_metrics" ``` You should see an output similar to the following: ``` # HELP apisix_http_requests_total The total number of client requests since APISIX started # TYPE apisix_http_requests_total gauge apisix_http_requests_total 1 # HELP apisix_nginx_http_current_connections Number of HTTP connections # TYPE apisix_nginx_http_current_connections gauge apisix_nginx_http_current_connections{state="accepted"} 1 apisix_nginx_http_current_connections{state="active"} 1 apisix_nginx_http_current_connections{state="handled"} 1 apisix_nginx_http_current_connections{state="reading"} 0 apisix_nginx_http_current_connections{state="waiting"} 0 apisix_nginx_http_current_connections{state="writing"} 1 ... ``` ### Integrate APISIX with Prometheus and Grafana[​](#integrate-apisix-with-prometheus-and-grafana "Direct link to Integrate APISIX with Prometheus and Grafana") To learn about how to collect APISIX metrics with Prometheus and visualize them in Grafana, see [how-to guide](https://docs.api7.ai/apisix/how-to-guide/observability/monitor-apisix-with-prometheus.md). ### Monitor Upstream Health Statuses[​](#monitor-upstream-health-statuses "Direct link to Monitor Upstream Health Statuses") The following example demonstrates how to monitor the health status of upstream nodes. Create a route with the `prometheus` plugin and configure upstream active health checks: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "prometheus-route", "uri": "/get", "plugins": { "prometheus": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1, "127.0.0.1:20001": 1 }, "checks": { "active": { "timeout": 5, "http_path": "/status", "healthy": { "interval": 2, "successes": 1 }, "unhealthy": { "interval": 1, "http_failures": 2 } }, "passive": { "healthy": { "http_statuses": [200, 201], "successes": 3 }, "unhealthy": { "http_statuses": [500], "http_failures": 3, "tcp_failures": 3 } } } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: prometheus-route plugins: prometheus: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 - host: 127.0.0.1 port: 20001 weight: 1 checks: active: timeout: 5 http_path: /status healthy: interval: 2 successes: 1 unhealthy: interval: 1 http_failures: 2 passive: healthy: http_statuses: - 200 - 201 successes: 3 unhealthy: http_statuses: - 500 http_failures: 3 tcp_failures: 3 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD prometheus-health-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: healthy-httpbin spec: type: ExternalName externalName: httpbin.org --- apiVersion: v1 kind: Service metadata: namespace: aic name: unhealthy-httpbin spec: type: ExternalName externalName: example.com --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: healthy-httpbin-health spec: targetRefs: - group: "" kind: Service name: healthy-httpbin healthCheck: active: type: http httpPath: /status/200 timeout: 5s healthy: interval: 2s successes: 1 unhealthy: interval: 1s httpFailures: 2 passive: type: http healthy: httpCodes: - 200 - 201 successes: 3 unhealthy: httpCodes: - 500 httpFailures: 3 tcpFailures: 3 --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: unhealthy-httpbin-health spec: targetRefs: - group: "" kind: Service name: unhealthy-httpbin healthCheck: active: type: http httpPath: /status/200 timeout: 5s healthy: interval: 2s successes: 1 unhealthy: interval: 1s httpFailures: 2 passive: type: http healthy: httpCodes: - 200 - 201 successes: 3 unhealthy: httpCodes: - 500 httpFailures: 3 tcpFailures: 3 --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: prometheus-plugin-config spec: plugins: - name: prometheus config: {} --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: prometheus-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /status/200 filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: prometheus-plugin-config backendRefs: - name: healthy-httpbin port: 80 weight: 1 - name: unhealthy-httpbin port: 80 weight: 1 ``` prometheus-health-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Service name: healthy-httpbin port: 80 - type: Service name: unhealthy-httpbin port: 80 healthCheck: active: type: http httpPath: /status timeout: 5 healthy: interval: 2s successes: 1 unhealthy: interval: 1s httpFailures: 2 passive: type: http healthy: httpCodes: - 200 - 201 successes: 3 unhealthy: httpCodes: - 500 httpFailures: 3 tcpFailures: 3 --- apiVersion: v1 kind: Service metadata: namespace: aic name: healthy-httpbin spec: type: ExternalName externalName: httpbin.org --- apiVersion: v1 kind: Service metadata: namespace: aic name: unhealthy-httpbin spec: type: ExternalName externalName: example.com --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: prometheus-route spec: ingressClassName: apisix http: - name: prometheus-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: prometheus enable: true config: {} ``` Apply the configuration to your cluster: ``` kubectl apply -f prometheus-health-ic.yaml ``` Send a request to the APISIX Prometheus metrics endpoint: ``` curl "http://127.0.0.1:9091/apisix/prometheus/metrics" ``` You should see an output similar to the following: ``` # HELP apisix_upstream_status upstream status from health check # TYPE apisix_upstream_status gauge apisix_upstream_status{name="/upstreams/",ip="",port="80"} 1 apisix_upstream_status{name="/upstreams/",ip="",port="80"} 0 ``` In that sample output, one upstream node is healthy and another upstream node is unhealthy. To learn more about how to configure active and passive health checks, see [health checks](https://docs.api7.ai/apisix/how-to-guide/traffic-management/health-check.md). ### Add Extra Labels for Metrics[​](#add-extra-labels-for-metrics "Direct link to Add Extra Labels for Metrics") The following example demonstrates how to add additional labels to metrics and use [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) in label values. Currently, extra labels are supported on: * `apisix_http_status` * `apisix_http_latency` * `apisix_bandwidth` * all `apisix_llm_*` metrics listed above * all four `apisix_ai_cache_*` metrics Both APISIX and API7 Gateway apply extra labels the same way. Add extra labels to the Prometheus static configuration: * Host or Docker * Kubernetes (Helm) Add or update this section in the gateway configuration file: config.yaml ``` plugin_attr: prometheus: # Plugin: prometheus metrics: # Create extra labels from built-in variables. http_status: extra_labels: # Set the extra labels for http_status metrics. - upstream_addr: $upstream_addr # Add an extra upstream_addr label with value being the NGINX variable $upstream_addr. - route_name: $route_name # Add an extra route_name label with value being the APISIX variable $route_name. ``` Reload the gateway for changes to take effect. For Helm deployments, update the chart values that render `plugin_attr.prometheus.metrics`. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` apisix: pluginAttrs: prometheus: metrics: http_status: extra_labels: - upstream_addr: $upstream_addr - route_name: $route_name ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` pluginAttrs: prometheus: metrics: http_status: extra_labels: - upstream_addr: $upstream_addr - route_name: $route_name ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` Note that if you define a variable in the label value but it does not correspond to any existing [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md), the label value will default to an empty string. Create a route with the `prometheus` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "prometheus-route", "uri": "/get", "name": "extra-label", "plugins": { "prometheus": {} }, "upstream": { "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: extra-label plugins: prometheus: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD prometheus-labels-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: prometheus-plugin-config spec: plugins: - name: prometheus config: {} --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: prometheus-route spec: parentRefs: - name: apisix hostnames: - "prometheus.example.com" rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: prometheus-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` prometheus-labels-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: prometheus-route spec: ingressClassName: apisix http: - name: prometheus-route match: hosts: - "prometheus.example.com" paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: prometheus enable: true config: {} ``` Apply the configuration to your cluster: ``` kubectl apply -f prometheus-labels-ic.yaml ``` Send a request to the route to verify: ``` curl -i "http://127.0.0.1:9080/get" ``` You should see an `HTTP/1.1 200 OK` response. Send a request to the APISIX Prometheus metrics endpoint: ``` curl "http://127.0.0.1:9091/apisix/prometheus/metrics" ``` You should see an output similar to the following: ``` # HELP apisix_http_status HTTP status codes per service in APISIX # TYPE apisix_http_status counter apisix_http_status{code="200",route="1",matched_uri="/get",matched_host="",service="",consumer="",node="54.237.103.220",request_type="traditional_http",request_llm_model="",llm_model="",response_source="upstream",upstream_addr="54.237.103.220:80",route_name="extra-label"} 1 ``` ### Monitor TCP/UDP Traffic with Prometheus[​](#monitor-tcpudp-traffic-with-prometheus "Direct link to Monitor TCP/UDP Traffic with Prometheus") The following example demonstrates how to collect TCP/UDP traffic metrics in APISIX. To collect TCP/UDP metrics, enable stream proxy and add `prometheus` to the existing stream plugin list. Preserve the other stream plugins used by the deployment; the host/Docker example below shows the minimal list for this walkthrough. * Host or Docker * Kubernetes (Helm) Add or update this section in the gateway configuration file: config.yaml ``` apisix: proxy_mode: http&stream # Enable both L4 & L7 proxies stream_proxy: # Configure L4 proxy tcp: - 9100 # Set TCP proxy listening port udp: - 9200 # Set UDP proxy listening port stream_plugins: - prometheus # Enable prometheus for stream proxy ``` Reload the gateway for changes to take effect. For the APISIX Helm chart, configure stream listeners and include the complete stream plugin list you want the gateway to load. The following example keeps the default stream plugins and adds `prometheus`: values.yaml ``` service: stream: enabled: true tcp: - 9100 udp: - 9200 apisix: stream_plugins: - ip-restriction - limit-conn - mqtt-proxy - prometheus - syslog ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` gateway: stream: enabled: true tcp: - addr: 9100 udp: - addr: 9200 ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` In API7 Enterprise the stream plugin list is delivered by the Control Plane, so the values above are all that is needed and the `stream_plugins` entry shown for the other deployment forms is not required. Create a [stream route](https://docs.api7.ai/apisix/key-concepts/stream-routes.md) with the `prometheus` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/stream_routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "prometheus-route", "plugins": { "prometheus":{} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Send a request to the stream route to verify: ``` curl -i "http://127.0.0.1:9100" ``` You should see an `HTTP/1.1 200 OK` response. Send a request to the APISIX Prometheus metrics endpoint: ``` curl "http://127.0.0.1:9091/apisix/prometheus/metrics" ``` You should see an output similar to the following: ``` # HELP apisix_stream_connection_total Total number of connections handled per stream route in APISIX # TYPE apisix_stream_connection_total counter apisix_stream_connection_total{route="prometheus-route"} 1 ``` APISIX also exports the termination status. When APISIX-Runtime provides the stream-metrics module, the scrape includes active connections and bandwidth: ``` # HELP apisix_stream_active_connections Number of stream sessions currently being proxied per listening address # TYPE apisix_stream_active_connections gauge apisix_stream_active_connections{listen_addr="0.0.0.0:9100"} 0 # HELP apisix_stream_status Stream sessions per termination status in APISIX # TYPE apisix_stream_status counter apisix_stream_status{code="200",listen_addr="0.0.0.0:9100",node="54.237.103.220:80"} 1 # HELP apisix_stream_bandwidth Total bandwidth in bytes proxied by the stream subsystem in APISIX # TYPE apisix_stream_bandwidth counter apisix_stream_bandwidth{listen_addr="0.0.0.0:9100",type="ingress",side="downstream"} 78 apisix_stream_bandwidth{listen_addr="0.0.0.0:9100",type="egress",side="downstream"} 219 apisix_stream_bandwidth{listen_addr="0.0.0.0:9100",type="egress",side="upstream"} 78 apisix_stream_bandwidth{listen_addr="0.0.0.0:9100",type="ingress",side="upstream"} 219 ``` The exact upstream address and byte counts depend on the request. The active-connections gauge is `0` above because the request completed before the scrape; scrape while a connection remains open to observe a positive value. The active-connection and bandwidth metrics use a shared memory zone that defaults to `1m`. Increase the zone when a gateway exposes many stream listening addresses: config.yaml ``` nginx_config: stream: metrics_zone_size: 2m ``` Reload APISIX after changing the zone size. On a runtime without the stream-metrics module, APISIX continues to export the connection-total and status metrics but does not publish active-connection or bandwidth metrics. --- ## Static Configurations[​](#static-configurations "Direct link to Static Configurations") By default, `prometheus` configurations are pre-configured in the [default configuration](https://github.com/apache/apisix/blob/master/apisix/cli/config.lua). The prompt- and completion-token histograms default to buckets at `1`, `10`, `50`, `100`, `200`, `500`, `1000`, `2000`, `5000`, `10000`, `20000`, `50000`, `100000`, `200000`, `500000`, and `1000000` tokens. Configure `llm_prompt_tokens_buckets` and `llm_completion_tokens_buckets` when different boundaries better fit the workload. The file to update depends on how the gateway is deployed: * Host or Docker * Kubernetes (Helm) For host or Docker deployments, configure the following settings: config.yaml ``` plugin_attr: prometheus: # Plugin: prometheus attributes export_uri: /apisix/prometheus/metrics # Set the URI for the Prometheus metrics endpoint. metric_prefix: apisix_ # Set the prefix for Prometheus metrics generated by APISIX. enable_export_server: true # Enable the Prometheus export server. export_addr: # Set the address for the Prometheus export server. ip: 127.0.0.1 # Set the IP. port: 9091 # Set the port. refresh_interval: 15 # Only available in APISIX. # Set the interval for refreshing cached metric data, in seconds. fetch_metric_timeout: 5 # Only available in API7 Enterprise. # Timeout for fetching metrics in seconds. If exceeded, the API only returns # basic metrics, including nginx_http_current_connections, http_requests_total, # etcd_reachable, prometheus_disable, node_info, etcd_modify_indexes, # shared_dict_capacity_bytes, and shared_dict_free_space_bytes. allow_degradation: false # Only available in API7 Enterprise. # If true, allow degradation when shared memory is insufficient. degradation_pause_steps: [ 60 ] # Only available in API7 Enterprise. # Time to skip the execution of the plugin in seconds when the plugin is in # degradation while reclaiming the shared memory used by the plugin. # metrics: # Create extra labels for metrics. # http_status: # These metrics will be prefixed with `apisix_`. # extra_labels: # Set the extra labels for http_status metrics. # - upstream_addr: $upstream_addr # - status: $upstream_status # expire: 0 # The expiration time of metrics in seconds. # 0 means the metrics will not expire. # http_latency: # extra_labels: # Set the extra labels for http_latency metrics. # - upstream_addr: $upstream_addr # expire: 0 # The expiration time of metrics in seconds. # 0 means the metrics will not expire. # bandwidth: # extra_labels: # Set the extra labels for bandwidth metrics. # - upstream_addr: $upstream_addr # expire: 0 # The expiration time of metrics in seconds. # 0 means the metrics will not expire. # default_buckets: # Built-in `http_latency` histogram buckets in milliseconds when this key is omitted. # Uncomment the list only to override those defaults. # - 1 # - 2 # - 5 # - 10 # - 20 # - 50 # - 100 # - 200 # - 500 # - 1000 # - 2000 # - 5000 # - 10000 # - 30000 # - 60000 # llm_latency_buckets: # Set buckets for `apisix_llm_latency`, in milliseconds. # # Introduced in API7 Enterprise 3.9.7 and APISIX 3.17.0. # # Applies to both `type=total` and `type=ttft`. # # Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. # - 100 # - 500 # - 1000 # - 5000 # llm_prompt_tokens_buckets: # Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. # # Set the buckets for `apisix_llm_prompt_tokens_dist` histogram, in tokens. # - 100 # - 1000 # - 10000 # llm_completion_tokens_buckets: # Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. # # Set the buckets for `apisix_llm_completion_tokens_dist` histogram, in tokens. # - 100 # - 1000 # - 10000 ``` Then reload the gateway for changes to take effect. For Helm deployments, update the chart values that render `plugin_attr.prometheus`. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` apisix: pluginAttrs: prometheus: export_uri: /apisix/prometheus/metrics metric_prefix: apisix_ enable_export_server: true export_addr: ip: 127.0.0.1 port: 9091 refresh_interval: 15 # Add the remaining prometheus plugin_attr fields here. ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` pluginAttrs: prometheus: export_uri: /apisix/prometheus/metrics metric_prefix: apisix_ enable_export_server: true export_addr: ip: 127.0.0.1 port: 9091 fetch_metric_timeout: 5 allow_degradation: false degradation_pause_steps: [ 60 ] # Add the remaining prometheus plugin_attr fields here. ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` You can use [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) to create `extra_labels`. See [add extra labels](https://docs.api7.ai/hub/prometheus.md#add-extra-labels-for-metrics) for more details. ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * prefer\_name boolean default: `false` *** If true, export route/service name instead of their ID in Prometheus metrics. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") APISIX and API7 Enterprise both support plugin metadata. The available fields, metadata keys, and version boundaries are described in the table below. Plugin metadata is configured through the Admin API or declarative configuration. It is separate from the Helm values that render `plugin_attr` in `config.yaml`. * disabled\_labels object *** Labels to disable to reduce the number of metrics and prevent resource bottlenecks. APISIX uses `http_status` and `http_latency` as the metadata keys for the corresponding metrics, while API7 Enterprise uses `status` and `latency`. Labels that define a metric's identity cannot be disabled, because collapsing them would merge distinct measurements into a single series. These are `code` on the HTTP status metric, `type` on the HTTP latency, bandwidth, and LLM latency metrics, and `layer` on `ai_cache_hits_total`. Structural-label validation was introduced in API7 Enterprise 3.9.17 and 3.10.4, and APISIX 3.18.0. See [Reduce Metric Cardinality by Disabling Labels](https://docs.api7.ai/hub/prometheus.md#reduce-metric-cardinality-by-disabling-labels) for product-specific keys and label sets. * http\_status array\[string] vaild vaule: Any combination of `route`, `matched_uri`, `matched_host`, `service`, `consumer`, `node`, `request_type`, `request_llm_model`, `llm_model`, and `response_source` *** Labels to disable for `apisix_http_status`. API7 Enterprise uses `status` instead. Introduced in APISIX 3.18.0. * http\_latency array\[string] vaild vaule: Any combination of `route`, `service`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model` *** Labels to disable for `apisix_http_latency`. API7 Enterprise uses `latency` instead. Introduced in APISIX 3.18.0. * status array\[string] vaild vaule: Any combination of `route`, `route_id`, `matched_uri`, `matched_host`, `service`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, `llm_model`, `mcp_request_type`, and `mcp_tool_name` *** Labels to disable for `apisix_http_status` metrics. Available in API7 Enterprise; APISIX uses `http_status` instead. The `request_type`, `request_llm_model`, and `llm_model` labels were introduced in API7 Enterprise 3.9.7. The `mcp_request_type` and `mcp_tool_name` labels were introduced in API7 Enterprise 3.9.14. * latency array\[string] vaild vaule: Any combination of `route`, `route_id`, `service`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, `llm_model`, `mcp_request_type`, and `mcp_tool_name` *** Labels to disable for `apisix_http_latency` metrics. Available in API7 Enterprise; APISIX uses `http_latency` instead. The `request_type`, `request_llm_model`, and `llm_model` labels were introduced in API7 Enterprise 3.9.7. The `mcp_request_type` and `mcp_tool_name` labels were introduced in API7 Enterprise 3.9.14. * bandwidth array\[string] vaild vaule: APISIX: Any combination of `route`, `service`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`
API7 Enterprise: APISIX values plus `route_id`, `service_id`, `mcp_request_type`, and `mcp_tool_name` *** Labels to disable for `apisix_bandwidth` metrics. Available in API7 Enterprise and introduced in APISIX 3.18.0. The `request_type`, `request_llm_model`, and `llm_model` labels were introduced in API7 Enterprise 3.9.7. The `mcp_request_type` and `mcp_tool_name` labels were introduced in API7 Enterprise 3.9.14. * stream\_status array\[string] vaild vaule: `node` *** Labels to disable for `apisix_stream_status` metrics. The `code` and `listen_addr` labels define the metric's identity and cannot be disabled. Available in API7 Enterprise from version 3.9.19 on the 3.9 line and from version 3.10.6 on the 3.10 line. * llm\_latency array\[string] vaild vaule: APISIX: Any combination of `route_id`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`
API7 Enterprise: APISIX values plus `route` and `service` *** Labels to disable for `apisix_llm_latency` metrics. Introduced in API7 Enterprise 3.9.7 and APISIX 3.18.0. * llm\_prompt\_tokens array\[string] vaild vaule: APISIX: Any combination of `route_id`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`
API7 Enterprise: APISIX values plus `route`, `matched_uri`, `matched_host`, and `service` *** Labels to disable for `apisix_llm_prompt_tokens` metrics. Introduced in API7 Enterprise 3.9.7 and APISIX 3.18.0. * llm\_completion\_tokens array\[string] vaild vaule: APISIX: Any combination of `route_id`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`
API7 Enterprise: APISIX values plus `route`, `matched_uri`, `matched_host`, and `service` *** Labels to disable for `apisix_llm_completion_tokens` metrics. Introduced in API7 Enterprise 3.9.7 and APISIX 3.18.0. * llm\_active\_connections array\[string] vaild vaule: Any combination of `route`, `route_id`, `matched_uri`, `matched_host`, `service`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model` *** Labels to disable for `apisix_llm_active_connections` metrics. Introduced in API7 Enterprise 3.9.7 and APISIX 3.18.0. * llm\_prompt\_tokens\_dist array\[string] vaild vaule: APISIX: Any combination of `route_id`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`
API7 Enterprise: APISIX values plus `route`, `matched_uri`, `matched_host`, and `service` *** Labels to disable for `apisix_llm_prompt_tokens_dist` metrics. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * llm\_completion\_tokens\_dist array\[string] vaild vaule: APISIX: Any combination of `route_id`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`
API7 Enterprise: APISIX values plus `route`, `matched_uri`, `matched_host`, and `service` *** Labels to disable for `apisix_llm_completion_tokens_dist` metrics. Introduced in API7 Enterprise 3.9.14 and 3.10.1, and APISIX 3.18.0. * ai\_cache\_hits\_total array\[string] vaild vaule: APISIX: Any combination of `route`, `route_id`, `service`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`
API7 Enterprise: APISIX values plus `matched_uri` and `matched_host` *** Labels to disable for `apisix_ai_cache_hits_total` metrics. The `layer` label is structural and cannot be disabled, because collapsing it would merge exact-match and semantic cache hits into a single series. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. * ai\_cache\_misses\_total array\[string] vaild vaule: APISIX: Any combination of `route`, `route_id`, `service`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`
API7 Enterprise: APISIX values plus `matched_uri` and `matched_host` *** Labels to disable for `apisix_ai_cache_misses_total` metrics. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. * ai\_cache\_bypasses\_total array\[string] vaild vaule: APISIX: Any combination of `route`, `route_id`, `service`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`
API7 Enterprise: APISIX values plus `matched_uri` and `matched_host` *** Labels to disable for `apisix_ai_cache_bypasses_total` metrics. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. * ai\_cache\_embedding\_latency array\[string] vaild vaule: APISIX: Any combination of `route`, `route_id`, `service`, `service_id`, `consumer`, `node`, `request_type`, `request_llm_model`, and `llm_model`
API7 Enterprise: APISIX values plus `matched_uri` and `matched_host` *** Labels to disable for `apisix_ai_cache_embedding_latency` metrics. Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0. --- # proxy-buffering The `proxy-buffering` plugin dynamically disables the NGINX [`proxy_buffering`](http://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_buffering) directive. Disable buffering for [server-sent events (SSE)](https://en.wikipedia.org/wiki/Server-sent_events) upstream services and other services that send incremental or chunked responses, such as etcd watch events. ## Examples[​](#examples "Direct link to Examples") ### Configure with SSE Upstream[​](#configure-with-sse-upstream "Direct link to Configure with SSE Upstream") The following example demonstrates how to disable `proxy_buffering` on a route with an SSE upstream service. Start a [sample upstream service](https://hub.docker.com/r/jmalloc/echo-server) for SSE: * Docker * Kubernetes ``` docker run -d -p 8080:8080 jmalloc/echo-server ``` Create a Kubernetes manifest file for the deployment of the SSE server: sse-deployment.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: sse-server spec: replicas: 1 selector: matchLabels: app: sse-server template: metadata: labels: app: sse-server spec: containers: - name: echo-server image: jmalloc/echo-server ports: - containerPort: 8080 ``` Create another Kubernetes manifest file for the SSE service: sse-service.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: sse-service spec: selector: app: sse-server ports: - protocol: TCP port: 8080 targetPort: 8080 type: ClusterIP ``` Apply the manifests: ``` kubectl apply -f sse-deployment.yaml -f sse-service.yaml ``` Create a route to the upstream and configure `proxy-buffering`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-buffering-route", "uri": "/.sse", "plugins": { "proxy-buffering": { "disable_proxy_buffering": true } }, "upstream": { "type": "roundrobin", "nodes": { "127.0.0.1:8080": 1 } } }' ``` adc.yaml ``` services: - name: sse-service routes: - uris: - /.sse name: proxy-buffering-route plugins: proxy-buffering: disable_proxy_buffering: true upstream: type: roundrobin nodes: - host: 127.0.0.1 port: 8080 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-buffering-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-buffering-plugin-config spec: plugins: - name: proxy-buffering config: disable_proxy_buffering: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-buffering-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /.sse filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-buffering-plugin-config backendRefs: - name: sse-service port: 8080 ``` proxy-buffering-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-buffering-route spec: ingressClassName: apisix http: - name: proxy-buffering-route match: paths: - /.sse upstreams: - serviceName: sse-service servicePort: 8080 plugins: - name: proxy-buffering config: disable_proxy_buffering: true ``` Apply the configuration: ``` kubectl apply -f proxy-buffering-ic.yaml ``` Send a request to the route: ``` curl "http://127.0.0.1:9080/.sse" -H "Accept: text/event-stream" ``` You should receive a `HTTP/1.1 200 OK` response and see continuous event stream similar to the following: ``` event: server data: 162291b28f55 id: 1 event: request data: GET /.sse HTTP/1.1 data: data: Host: 127.0.0.1:9080 data: Accept: text/event-stream data: User-Agent: curl/7.74.0 data: X-Forwarded-For: 172.19.0.1 data: X-Forwarded-Host: 127.0.0.1 data: X-Forwarded-Port: 9080 data: X-Forwarded-Proto: http data: X-Real-Ip: 172.19.0.1 data: id: 2 event: time data: 2023-10-19T02:13:53Z id: 3 event: time data: 2023-10-19T02:13:54Z id: 4 event: time data: 2023-10-19T02:13:55Z id: 5 ... ``` #### (Optional) See Buffering in Effect[​](#optional-see-buffering-in-effect "Direct link to (Optional) See Buffering in Effect") You will see the proxy buffering effect on SSE when `proxy_buffering` is not turned off in this section. For demonstration, the proxy buffer size will be adjusted to a larger value. Add the following snippet to the gateway static configuration: * Host or Docker * Kubernetes (Helm) config.yaml ``` nginx_config: http_configuration_snippet: | server { listen 9080; location /sse { proxy_buffering on; proxy_buffers 4 2m; } } ``` For the APISIX Helm chart, set the following values: values.yaml ``` apisix: nginx: configurationSnippet: httpStart: | server { listen 9080; location /sse { proxy_buffering on; proxy_buffers 4 2m; } } ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` configurationSnippet: httpStart: | server { listen 9080; location /sse { proxy_buffering on; proxy_buffers 4 2m; } } ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` ❶ Though explicitly configured, `proxy_buffering` is `on` [by default](https://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_buffering). ❷ Configure `proxy_buffers` to use 4 buffers, each with a size of 2 MB. Reload the gateway for changes to take effect. Recreate the route without `proxy-buffering`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-buffering-route", "uri": "/.sse", "upstream": { "type": "roundrobin", "nodes": { "127.0.0.1:8080": 1 } } }' ``` adc.yaml ``` services: - name: sse-service routes: - uris: - /.sse name: proxy-buffering-route upstream: type: roundrobin nodes: - host: 127.0.0.1 port: 8080 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-buffering-ic.yaml ``` apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-buffering-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /.sse backendRefs: - name: sse-service port: 8080 ``` proxy-buffering-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-buffering-route spec: ingressClassName: apisix http: - name: proxy-buffering-route match: paths: - /.sse upstreams: - serviceName: sse-service servicePort: 8080 ``` Apply the configuration: ``` kubectl apply -f proxy-buffering-ic.yaml ``` Send a request to the route: ``` curl "http://127.0.0.1:9080/.sse" -H "Accept: text/event-stream" ``` You should receive a `HTTP/1.1 200 OK` response and see the same event stream. However, note that events are not received at a regular interval due to the effect of the buffer, which is undesirable when working with an SSE upstream. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * disable\_proxy\_buffering boolean default: `false` *** If true, set the NGINX `proxy_buffering` directive to `off`. --- # proxy-cache The `proxy-cache` plugin provides the capability to cache responses based on a cache key. The plugin supports both disk-based and memory-based caching options to cache for [GET](https://anything.org/learn/serving-over-http/#get-request), [POST](https://anything.org/learn/serving-over-http/#post-request), and [HEAD](https://developer.mozilla.org/en-US/docs/Web/HTTP/Methods/HEAD) requests. Responses can be conditionally cached based on request HTTP methods, response status codes, request header values, and more. Authenticated consumers use separate effective cache keys by default. This isolation was introduced in API7 Enterprise 3.9.13 and APISIX 3.17.0. The in-memory strategy does not cache responses containing `Set-Cookie` unless `cache_set_cookie` is enabled, and it never caches responses marked `private`, `no-store`, or `no-cache` by the upstream. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `proxy-cache` for different scenarios. ### Cache Data on Disk[​](#cache-data-on-disk "Direct link to Cache Data on Disk") On-disk caching strategy offers the advantages of data persistency when system restarts and having larger storage capacity compared to in-memory cache. It is suitable for applications that prioritize durability and can tolerate slightly larger cache access latency. The following example demonstrates how you can use `proxy-cache` plugin on a route to cache data on disk. When using the on-disk caching strategy, the cache TTL is determined by the `Expires` or `Cache-Control` response header. If neither header is present, or if APISIX returns `502 Bad Gateway` or `504 Gateway Timeout` due to unavailable upstreams, the cache TTL defaults to the value configured in the [configuration files](https://docs.api7.ai/hub/proxy-cache/configuration.md#static-configurations). Create a route with the `proxy-cache` plugin to cache data on disk: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-cache-route", "uri": "/anything", "plugins": { "proxy-cache": { "cache_strategy": "disk" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: proxy-cache-service routes: - name: proxy-cache-route uris: - /anything plugins: proxy-cache: cache_strategy: disk upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-cache-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-cache-plugin-config spec: plugins: - name: proxy-cache config: cache_strategy: disk --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-cache-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-cache-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` proxy-cache-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-cache-route spec: ingressClassName: apisix http: - name: proxy-cache-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: proxy-cache enable: true config: cache_strategy: disk ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-cache-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should see an `HTTP/1.1 200 OK` response with the following header, showing the plugin is successfully enabled: ``` Apisix-Cache-Status: MISS ``` As there is no cache available before the first response, `Apisix-Cache-Status: MISS` is shown. Send the same request again within the cache TTL window. You should see an `HTTP/1.1 200 OK` response with the following headers, showing the cache is hit: ``` Apisix-Cache-Status: HIT ``` Wait for the cache to expire after the TTL and send the same request again. You should see an `HTTP/1.1 200 OK` response with the following headers, showing the cache has expired: ``` Apisix-Cache-Status: EXPIRED ``` ### Cache Data in Memory[​](#cache-data-in-memory "Direct link to Cache Data in Memory") In-memory caching strategy offers the advantage of low-latency access to the cached data, as retrieving data from RAM is faster than retrieving data from disk storage. It also works well for storing temporary data that does not need to be persisted long-term, allowing for efficient caching of frequently changing data. The in-memory strategy caches responses separately for each variant computed from the request headers listed in the upstream `Vary` response header. Responses with `Vary: *` are not cached. This behavior was introduced in API7 Enterprise 3.9.14 and 3.10.0, and in APISIX 3.17.0. The following example demonstrates how you can use `proxy-cache` plugin on a route to cache data in memory. Create a route with `proxy-cache` and configure it to use memory-based caching: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-cache-route", "uri": "/anything", "plugins": { "proxy-cache": { "cache_strategy": "memory", "cache_zone": "memory_cache", "cache_ttl": 10 } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: proxy-cache-service routes: - name: proxy-cache-route uris: - /anything plugins: proxy-cache: cache_strategy: memory cache_zone: memory_cache cache_ttl: 10 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-cache-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-cache-plugin-config spec: plugins: - name: proxy-cache config: cache_strategy: memory cache_zone: memory_cache cache_ttl: 10 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-cache-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-cache-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` proxy-cache-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-cache-route spec: ingressClassName: apisix http: - name: proxy-cache-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: proxy-cache enable: true config: cache_strategy: memory cache_zone: memory_cache cache_ttl: 10 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-cache-ic.yaml ``` ❶ `cache_strategy`: set to `memory` for in-memory setting. ❷ `cache_zone`: set to the name of an in-memory cache zone. ❸ `cache_ttl`: set the time to live for the in-memory cache to 10 seconds. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should see an `HTTP/1.1 200 OK` response with the following header, showing the plugin is successfully enabled: ``` Apisix-Cache-Status: MISS ``` As there is no cache available before the first response, `Apisix-Cache-Status: MISS` is shown. Send the same request again within the cache TTL window. You should see an `HTTP/1.1 200 OK` response with the following headers, showing the cache is hit: ``` Apisix-Cache-Status: HIT ``` ### Remove Cache Manually[​](#remove-cache-manually "Direct link to Remove Cache Manually") While cached responses normally expire based on their TTL, you might need to remove cached data before it expires. The following example demonstrates how you can use the `PURGE` method to remove data cached on disk. `PURGE` also supports in-memory caching; to test it, use the memory cache configuration from the previous example. Send the `PURGE` request to the same route URI as the cached request. The plugin derives the cache key for a `PURGE` request in the same way as for a request that populates the cache. Use the same host, URI, query parameters, and any other values referenced by `cache_key`. If the cache key is isolated by consumer, send the request as the same consumer. `PURGE` removes every in-memory `Vary` variant indexed under the effective base cache key. This behavior was introduced in API7 Enterprise 3.9.14 and 3.10.0, and in APISIX 3.17.0. Create a route that caches responses on disk: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-cache-route", "uri": "/anything", "plugins": { "proxy-cache": { "cache_strategy": "disk" } }, "upstream": { "type": "roundrobin", "pass_host": "node", "scheme": "https", "nodes": { "httpbingo.org:443": 1 } } }' ``` adc.yaml ``` services: - name: proxy-cache-service routes: - name: proxy-cache-route uris: - /anything plugins: proxy-cache: cache_strategy: disk upstream: type: roundrobin pass_host: node scheme: https nodes: - host: httpbingo.org port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-cache-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbingo-external-domain spec: type: ExternalName externalName: httpbingo.org --- apiVersion: apisix.apache.org/v1alpha1 kind: BackendTrafficPolicy metadata: namespace: aic name: httpbingo-https spec: targetRefs: - name: httpbingo-external-domain kind: Service group: "" passHost: node scheme: https --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-cache-plugin-config spec: plugins: - name: proxy-cache config: cache_strategy: disk --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-cache-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-cache-plugin-config backendRefs: - name: httpbingo-external-domain port: 443 ``` proxy-cache-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbingo-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: httpbingo.org port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-cache-route spec: ingressClassName: apisix http: - name: proxy-cache-route match: paths: - /anything upstreams: - name: httpbingo-external-domain plugins: - name: proxy-cache enable: true config: cache_strategy: disk ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-cache-ic.yaml ``` Send a request to populate the cache: ``` curl -i "http://127.0.0.1:9080/anything" ``` Send the same request again and verify that the response contains `Apisix-Cache-Status: HIT`. Send a `PURGE` request to the same URI: ``` curl -i "http://127.0.0.1:9080/anything" -X PURGE ``` You should see an `HTTP/1.1 200 OK` response, showing that the cached response was removed. If no cached response matches the cache key, the plugin returns `HTTP/1.1 404 Not Found`. Send a `GET` request to the route again: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should see the following header, showing that the previous cached response is no longer available: ``` Apisix-Cache-Status: MISS ``` ### Cache Responses Conditionally[​](#cache-responses-conditionally "Direct link to Cache Responses Conditionally") The following example demonstrates how you can configure the `proxy-cache` plugin to conditionally cache responses. Create a route with the `proxy-cache` plugin and configure the `no_cache` attribute: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-cache-route", "uri": "/anything", "plugins": { "proxy-cache": { "no_cache": ["$arg_no_cache", "$http_no_cache"] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: proxy-cache-service routes: - name: proxy-cache-route uris: - /anything plugins: proxy-cache: no_cache: - $arg_no_cache - $http_no_cache upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-cache-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-cache-plugin-config spec: plugins: - name: proxy-cache config: no_cache: - $arg_no_cache - $http_no_cache --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-cache-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-cache-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` proxy-cache-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-cache-route spec: ingressClassName: apisix http: - name: proxy-cache-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: proxy-cache enable: true config: no_cache: - $arg_no_cache - $http_no_cache ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-cache-ic.yaml ``` ❶ `no_cache`: If at least one of the values of the URL parameter `no_cache` and header `no_cache` is not empty and is not equal to `0`, the response will not be cached. Send a few requests to the route with the URL parameter `no_cache` value indicating cache bypass: ``` curl -i "http://127.0.0.1:9080/anything?no_cache=1" ``` You should receive `HTTP/1.1 200 OK` responses for all requests and observe the following header every time: ``` Apisix-Cache-Status: EXPIRED ``` Send a few other requests to the route with the URL parameter `no_cache` value being zero: ``` curl -i "http://127.0.0.1:9080/anything?no_cache=0" ``` You should receive `HTTP/1.1 200 OK` responses for all requests and start seeing the cache being hit: ``` Apisix-Cache-Status: HIT ``` You can also specify the value in the `no_cache` header as such: ``` curl -i "http://127.0.0.1:9080/anything" -H "no_cache: 1" ``` The response should not be cached: ``` Apisix-Cache-Status: EXPIRED ``` ### Retrieve Responses from Cache Conditionally[​](#retrieve-responses-from-cache-conditionally "Direct link to Retrieve Responses from Cache Conditionally") The following example demonstrates how you can configure the `proxy-cache` plugin to conditionally retrieve responses from cache. Create a route with the `proxy-cache` plugin and configure the `cache_bypass` attribute: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-cache-route", "uri": "/anything", "plugins": { "proxy-cache": { "cache_bypass": ["$arg_bypass", "$http_bypass"] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: proxy-cache-service routes: - name: proxy-cache-route uris: - /anything plugins: proxy-cache: cache_bypass: - $arg_bypass - $http_bypass upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-cache-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-cache-plugin-config spec: plugins: - name: proxy-cache config: cache_bypass: - $arg_bypass - $http_bypass --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-cache-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-cache-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` proxy-cache-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-cache-route spec: ingressClassName: apisix http: - name: proxy-cache-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: proxy-cache enable: true config: cache_bypass: - $arg_bypass - $http_bypass ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-cache-ic.yaml ``` ❶ `cache_bypass`: If at least one of the values of the URL parameter `bypass` and header `bypass` is not empty and is not equal to `0`, the response will not be retrieved from the cache. Send a request to the route with the URL parameter `bypass` value indicating cache bypass: ``` curl -i "http://127.0.0.1:9080/anything?bypass=1" ``` You should see an `HTTP/1.1 200 OK` response with the following header: ``` Apisix-Cache-Status: BYPASS ``` Send another request to the route with the URL parameter `bypass` value being zero: ``` curl -i "http://127.0.0.1:9080/anything?bypass=0" ``` You should see an `HTTP/1.1 200 OK` response with the following header: ``` Apisix-Cache-Status: MISS ``` You can also specify the value in the `bypass` header as such: ``` curl -i "http://127.0.0.1:9080/anything" -H "bypass: 1" ``` The cache should be bypassed: ``` Apisix-Cache-Status: BYPASS ``` ### Cache for 502 and 504 Error Response Code[​](#cache-for-502-and-504-error-response-code "Direct link to Cache for 502 and 504 Error Response Code") When the upstream services return server errors in the 500 range, `proxy-cache` plugin will cache the responses if and only if the returned status is `502 Bad Gateway` or `504 Gateway Timeout`. The following example demonstrates the behavior of `proxy-cache` plugin when the upstream service returns `504 Gateway Timeout`. Create a route with the `proxy-cache` plugin and configure a dummy upstream service: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-cache-route", "uri": "/timeout", "plugins": { "proxy-cache": { } }, "upstream": { "type": "roundrobin", "nodes": { "12.34.56.78": 1 } } }' ``` adc.yaml ``` services: - name: proxy-cache-service routes: - name: proxy-cache-route uris: - /timeout plugins: proxy-cache: {} upstream: type: roundrobin nodes: - host: 12.34.56.78 port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-cache-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: dummy-upstream spec: type: ExternalName externalName: dummy.example.com --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-cache-plugin-config spec: plugins: - name: proxy-cache config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-cache-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /timeout filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-cache-plugin-config backendRefs: - name: dummy-upstream port: 80 ``` proxy-cache-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: dummy-upstream spec: ingressClassName: apisix externalNodes: - type: Domain name: dummy.example.com --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-cache-route spec: ingressClassName: apisix http: - name: proxy-cache-route match: paths: - /timeout upstreams: - name: dummy-upstream plugins: - name: proxy-cache enable: true ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-cache-ic.yaml ``` Generate a few requests to the route: ``` seq 4 | xargs -I{} curl -I "http://127.0.0.1:9080/timeout" ``` You should see a response similar to the following: ``` HTTP/1.1 504 Gateway Time-out ... Apisix-Cache-Status: MISS HTTP/1.1 504 Gateway Time-out ... Apisix-Cache-Status: HIT HTTP/1.1 504 Gateway Time-out ... Apisix-Cache-Status: HIT HTTP/1.1 504 Gateway Time-out ... Apisix-Cache-Status: HIT ``` However, if the upstream services returns `503 Service Temporarily Unavailable`, the response will not be cached. --- ## Static Configurations[​](#static-configurations "Direct link to Static Configurations") The gateway [default configuration](https://github.com/apache/apisix/blob/master/apisix/cli/config.lua) includes proxy cache settings for disk caching and cache zones. The file to update depends on how the gateway is deployed: * Host or Docker * Kubernetes (Helm) For host or Docker deployments, configure the following settings: config.yaml ``` apisix: proxy_cache: cache_ttl: 10s # default cache TTL used when caching on disk, only if none of the `Expires` # and `Cache-Control` response headers is present, or if APISIX returns # `502 Bad Gateway` or `504 Gateway Timeout` due to unavailable upstreams zones: - name: disk_cache_one memory_size: 50m disk_size: 1G disk_path: /tmp/disk_cache_one cache_levels: 1:2 # - name: disk_cache_two # memory_size: 50m # disk_size: 1G # disk_path: "/tmp/disk_cache_two" # cache_levels: "1:2" - name: memory_cache memory_size: 50m ``` Then [reload APISIX](https://docs.api7.ai/apisix/reference/apisix-cli.md#apisix-reload) for changes to take effect. In the APISIX Helm chart version `2.16.0` or later and the API7 Gateway Helm chart version `3.10.3` or later, set `apisix.proxyCache`. Each chart renders this value as `apisix.proxy_cache` in the gateway configuration: values.yaml ``` apisix: proxyCache: cacheTtl: 10s zones: - name: disk_cache_one memory_size: 50m disk_size: 1G disk_path: /tmp/disk_cache_one cache_levels: 1:2 - name: memory_cache memory_size: 50m ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * cache\_strategy string default: `disk` vaild vaule: `disk` or `memory` *** Caching strategy. Cache on disk or in memory. * cache\_zone string default: `disk_cache_one` *** Cache zone used with the caching strategy. The value should match one of the cache zones defined in the [configuration files](https://docs.api7.ai/hub/proxy-cache/configuration.md#static-configurations) and should correspond to the caching strategy. For example, when using the in-memory caching strategy, you should use an in-memory cache zone. * cache\_key array\[string] default: `["$host", "$request_uri"]` *** Key to use for caching. Support [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) and constant strings in values. Variables should be prefixed with a `$` sign. * cache\_bypass array\[string] *** One or more parameters to parse value from, such that if any of the values is not empty and is not equal to `0`, response will not be retrieved from cache. Support [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) and constant strings in values. Variables should be prefixed with a `$` sign. * cache\_method array\[string] default: `["GET", "HEAD"]` vaild vaule: Any combination of methods from "GET", "POST", and "HEAD" *** Request methods of which the response should be cached. * cache\_http\_status array\[integer] default: `[200, 301, 404]` vaild vaule: Any combination of integer values from 200 to 599 inclusive *** Response HTTP status codes of which the response should be cached. * hide\_cache\_headers boolean default: `false` *** If true, hide `Expires` and `Cache-Control` response headers. * cache\_control boolean default: `false` *** If true, the in-memory strategy honors supported request `Cache-Control` directives and derives the cache TTL from the upstream response's `s-maxage`, `max-age`, or `Expires` value. Regardless of this setting, responses containing `Cache-Control: private`, `no-store`, or `no-cache` are not cached in memory. * no\_cache array\[string] *** One or more parameters to parse value from, such that if any of the values is not empty and is not equal to `0`, response will not be cached. Support [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) and constant strings in values. Variables should be prefixed with a `$` sign. * cache\_ttl integer default: `300` vaild vaule: greater than or equal to 1 *** Cache time to live (TTL) in seconds when caching in memory. To adjust the TTL when caching on disk, update `cache_ttl` in the [configuration files](https://docs.api7.ai/hub/proxy-cache/configuration.md#static-configurations). The TTL value is evaluated in conjunction with the values in the response headers [`Cache-Control`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control) and [`Expires`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Expires) received from the upstream service. * consumer\_isolation boolean default: `true` *** If true, prepend the authenticated consumer identity to the effective cache key when the request resolves to a consumer or remote user. This is skipped when `cache_key` already contains an identity-bearing variable such as `$consumer_name`, `$consumer_group_id`, `$remote_user`, or `$http_authorization`. Available in API7 Enterprise from version 3.9.13 and APISIX from version 3.17.0. * cache\_set\_cookie boolean default: `false` *** If true, allow the in-memory strategy to cache responses that include a `Set-Cookie` header. By default, such responses are not cached. Available in API7 Enterprise from version 3.9.13 and APISIX from version 3.17.0. * max\_resp\_body\_size integer default: `67108864` vaild vaule: greater than or equal to 1 *** Maximum response body size in bytes buffered by the memory cache strategy. A response that reaches or exceeds this size is streamed to the client without being cached. The chunk that crosses the threshold is buffered before the limit is enforced, so transient memory use can exceed the configured size. This field does not apply to disk caching. Introduced in API7 Enterprise 3.9.17 and 3.10.4, and APISIX 3.18.0. --- # proxy-mirror The `proxy-mirror` plugin duplicates ingress traffic to APISIX and forwards them to a designated upstream, without interrupting the regular services. You can configure the plugin to mirror all traffic or only a portion. The mechanism benefits a few use cases, including troubleshooting, security inspection, analytics, and more. Note that APISIX ignores any response from the upstream host receiving mirrored traffic. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how to configure `proxy-mirror` for different scenarios. ### Mirror Partial Traffic[​](#mirror-partial-traffic "Direct link to Mirror Partial Traffic") The following example demonstrates how you can configure `proxy-mirror` to mirror 50% of the traffic to a route and forward them to another upstream service. Start a sample NGINX server for receiving mirrored traffic: * Docker * Kubernetes ``` docker run -p 8081:80 --name nginx nginx ``` You should see NGINX access log and error log on the terminal session. Create a Kubernetes manifest file for the NGINX deployment: nginx-deployment.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: nginx spec: replicas: 1 selector: matchLabels: app: nginx template: metadata: labels: app: nginx spec: containers: - name: nginx image: nginx ports: - containerPort: 80 --- apiVersion: v1 kind: Service metadata: namespace: aic name: nginx spec: selector: app: nginx ports: - protocol: TCP port: 80 targetPort: 80 type: ClusterIP ``` Apply the manifest to your cluster: ``` kubectl apply -f nginx-deployment.yaml ``` Open a new terminal session and create a route with `proxy-mirror`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "traffic-mirror-route", "uri": "/get", "plugins": { "proxy-mirror": { "host": "http://127.0.0.1:8081", "sample_ratio": 0.5 } }, "upstream": { "nodes": { "httpbin.org": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: proxy-mirror-service routes: - name: traffic-mirror-route uris: - /get plugins: proxy-mirror: host: "http://127.0.0.1:8081" sample_ratio: 0.5 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-mirror-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-mirror-plugin-config spec: plugins: - name: proxy-mirror config: host: "http://nginx.aic.svc" sample_ratio: 0.5 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-mirror-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-mirror-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` proxy-mirror-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-mirror-route spec: ingressClassName: apisix http: - name: traffic-mirror-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: proxy-mirror enable: true config: host: "http://nginx.aic.svc" sample_ratio: 0.5 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-mirror-ic.yaml ``` ❶ `host`: configure the scheme and host address to forward the mirrored traffic to. ❷ `sample_ratio`: configure the sampling ratio to 0.5 to mirror 50% of the traffic. Send a few requests to the route: ``` curl -i "http://127.0.0.1:9080/get" ``` You should receive `HTTP/1.1 200 OK` responses for all requests. Navigating back to the NGINX terminal session, you should see a number of access log entries, roughly half the number of requests generated: ``` 172.17.0.1 - - [29/Jan/2024:23:11:01 +0000] "GET /get HTTP/1.1" 404 153 "-" "curl/7.64.1" "-" ``` This suggests APISIX has mirrored the request to the NGINX server. Here, the HTTP response status is `404` since the sample NGINX server does not implement the route. ### Configure Mirroring Timeouts[​](#configure-mirroring-timeouts "Direct link to Configure Mirroring Timeouts") The following example demonstrates how you can update the default connect, read, and send timeouts for the plugin. This could be useful when mirroring traffic to a very slow backend service. As the request mirroring was implemented as sub-requests, excessive delays in the sub-requests could lead to the blocking of the original requests. By default, the connect, read, and send timeouts are set to 60 seconds. Configure the gateway static settings to change these defaults: * Host or Docker * Kubernetes (Helm) Add or update this section in the gateway configuration file: config.yaml ``` plugin_attr: proxy-mirror: timeout: connect: 2000ms read: 2000ms send: 2000ms ``` Reload the gateway for changes to take effect. For Helm deployments, update the chart values that render `plugin_attr.proxy-mirror`. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` apisix: pluginAttrs: proxy-mirror: timeout: connect: 2000ms read: 2000ms send: 2000ms ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` pluginAttrs: proxy-mirror: timeout: connect: 2000ms read: 2000ms send: 2000ms ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` --- ## Static Configurations[​](#static-configurations "Direct link to Static Configurations") By default, timeout values for the plugin are pre-configured in the [default configuration](https://github.com/apache/apisix/blob/master/apisix/cli/config.lua). The file to update depends on how the gateway is deployed: * Host or Docker * Kubernetes (Helm) For host or Docker deployments, configure the following settings: config.yaml ``` plugin_attr: proxy-mirror: timeout: connect: 60s read: 60s send: 60s ``` Then reload the gateway for changes to take effect. For Helm deployments, update the chart values that render `plugin_attr.proxy-mirror`. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` apisix: pluginAttrs: proxy-mirror: timeout: connect: 60s read: 60s send: 60s ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` pluginAttrs: proxy-mirror: timeout: connect: 60s read: 60s send: 60s ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * host string required *** Address of the host to forward the mirrored traffic to. The address should contain the scheme but without the path, such as `http://127.0.0.1:8081`. * path string *** Request path to use for mirrored HTTP traffic. If unspecified, the current route URI path is used. For native gRPC and grpc-web traffic, this field is ignored and the gateway preserves the effective gRPC method path after access-phase rewrites. * path\_concat\_mode string default: `replace` vaild vaule: `replace` or `prefix` *** Concatenation mode when `path` is specified. When set to `replace`, the configured path replaces the request path. When set to `prefix`, the request path is appended to the configured path. For native gRPC and grpc-web traffic, this field is ignored and the gateway preserves the effective gRPC method path. * sample\_ratio number default: `1` vaild vaule: between 0.00001 and 1 inclusive *** Ratio of the requests that will be mirrored. By default, all traffic are mirrored. --- # proxy-rewrite The `proxy-rewrite` plugin offers options to rewrite requests that APISIX forwards to upstream services. With the plugin, you can modify the HTTP methods, request destination upstream addresses, request headers, and more. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `proxy-rewrite` on a route in different scenarios. ### Rewrite Host Header[​](#rewrite-host-header "Direct link to Rewrite Host Header") The following example demonstrates how you can modify the `Host` header in a request. Note that you should not use `headers.set` to set the `Host` header. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-rewrite-route", "methods": ["GET"], "uri": "/headers", "plugins": { "proxy-rewrite": { "host": "myapisix.demo" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /headers name: proxy-rewrite-route methods: - GET plugins: proxy-rewrite: host: myapisix.demo upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-rewrite-plugin-config spec: plugins: - name: proxy-rewrite config: host: myapisix.demo --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` proxy-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-rewrite-route spec: ingressClassName: apisix http: - name: proxy-rewrite-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: proxy-rewrite enable: true config: host: myapisix.demo ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` Send a request to `/headers` to check all the request headers sent to upstream: ``` curl "http://127.0.0.1:9080/headers" ``` You should see a response similar to the following: ``` { "headers": { "Accept": "*/*", "Host": "myapisix.demo", "User-Agent": "curl/8.2.1", "X-Amzn-Trace-Id": "Root=1-64fef198-29da0970383150175bd2d76d", "X-Forwarded-Host": "127.0.0.1" } } ``` ### Rewrite URI And Set Headers[​](#rewrite-uri-and-set-headers "Direct link to Rewrite URI And Set Headers") The following example rewrites the upstream URI and sets request headers. A scalar value sets one header line, while an array sets repeated header lines in array order. Here, `X-Api-Engine` remains a scalar and `X-Api-Version` replaces any client value with two values. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-rewrite-route", "methods": ["GET"], "uri": "/", "plugins": { "proxy-rewrite": { "uri": "/anything", "headers": { "set": { "X-Api-Version": [ "v1", "v2" ], "X-Api-Engine": "apisix" } } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbingo.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbingo routes: - uris: - / name: proxy-rewrite-route methods: - GET plugins: proxy-rewrite: uri: /anything headers: set: X-Api-Version: - v1 - v2 X-Api-Engine: apisix upstream: type: roundrobin nodes: - host: httpbingo.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbingo-external-domain spec: type: ExternalName externalName: httpbingo.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-rewrite-plugin-config spec: plugins: - name: proxy-rewrite config: uri: /anything headers: set: X-Api-Version: - v1 - v2 X-Api-Engine: apisix --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: / filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-rewrite-plugin-config backendRefs: - name: httpbingo-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` proxy-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbingo-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbingo.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-rewrite-route spec: ingressClassName: apisix http: - name: proxy-rewrite-route match: paths: - / upstreams: - name: httpbingo-external-domain plugins: - name: proxy-rewrite enable: true config: uri: /anything headers: set: X-Api-Version: - v1 - v2 X-Api-Engine: apisix ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` Send a request to verify: ``` curl "http://127.0.0.1:9080/" -H 'X-Api-Version: client' ``` You should see a response similar to the following: ``` { "headers": { "X-Api-Engine": [ "apisix" ], "X-Api-Version": [ "v1", "v2" ] } } ``` The plugin sends two `X-Api-Version` header lines in the configured order. The incoming `client` value is absent because `headers.set` replaces existing values. The scalar `X-Api-Engine` configuration continues to set one header line. ### Rewrite URI And Append Headers[​](#rewrite-uri-and-append-headers "Direct link to Rewrite URI And Append Headers") The following example demonstrates how you can rewrite the request upstream URI and append additional header values. If the same headers present in the client request, their headers values will append to the configured header values in the plugin. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-rewrite-route", "methods": ["GET"], "uri": "/", "plugins": { "proxy-rewrite": { "uri": "/headers", "headers": { "add": { "X-Api-Version": "v1", "X-Api-Engine": "apisix" } } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - / name: proxy-rewrite-route methods: - GET plugins: proxy-rewrite: uri: /headers headers: add: X-Api-Version: v1 X-Api-Engine: apisix upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-rewrite-plugin-config spec: plugins: - name: proxy-rewrite config: uri: /headers headers: add: X-Api-Version: v1 X-Api-Engine: apisix --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: / filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` proxy-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-rewrite-route spec: ingressClassName: apisix http: - name: proxy-rewrite-route match: paths: - / upstreams: - name: httpbin-external-domain plugins: - name: proxy-rewrite enable: true config: uri: /headers headers: add: X-Api-Version: v1 X-Api-Engine: apisix ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` Send a request to verify: ``` curl "http://127.0.0.1:9080/" -H 'X-Api-Version: v2' ``` You should see a response similar to the following: ``` { "headers": { "Accept": "*/*", "Host": "httpbin.org", "User-Agent": "curl/8.2.1", "X-Amzn-Trace-Id": "Root=1-64fed73a-59cd3bd640d76ab16c97f1f1", "X-Api-Engine": "apisix", "X-Api-Version": "v2,v1", "X-Forwarded-Host": "127.0.0.1" } } ``` Note that both headers present and the header value of `X-Api-Version` passed in the request is preserved, with the plugin's configured value appended after it. ### Remove Existing Header[​](#remove-existing-header "Direct link to Remove Existing Header") The following example demonstrates how you can remove an existing header `User-Agent`. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-rewrite-route", "methods": ["GET"], "uri": "/headers", "plugins": { "proxy-rewrite": { "headers": { "remove":[ "User-Agent" ] } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /headers name: proxy-rewrite-route methods: - GET plugins: proxy-rewrite: headers: remove: - User-Agent upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-rewrite-plugin-config spec: plugins: - name: proxy-rewrite config: headers: remove: - User-Agent --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` proxy-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-rewrite-route spec: ingressClassName: apisix http: - name: proxy-rewrite-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: proxy-rewrite enable: true config: headers: remove: - User-Agent ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` Send a request to verify if the specified header is removed: ``` curl "http://127.0.0.1:9080/headers" ``` You should see a response similar to the following, where the `User-Agent` header is not present: ``` { "headers": { "Accept": "*/*", "Host": "httpbin.org", "X-Amzn-Trace-Id": "Root=1-64fef302-07f2b13e0eb006ba776ad91d", "X-Forwarded-Host": "127.0.0.1" } } ``` ### Rewrite URI Using RegEx[​](#rewrite-uri-using-regex "Direct link to Rewrite URI Using RegEx") The following example demonstrates how you can parse text from the original upstream URI path and use them to compose a new upstream URI path. In this example, APISIX is configured to forward all requests from `/test/user/agent` to `/user-agent`. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-rewrite-route", "uri": "/test/*", "plugins": { "proxy-rewrite": { "regex_uri": ["^/test/(.*)/(.*)", "/$1-$2"] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /test/* name: proxy-rewrite-route plugins: proxy-rewrite: regex_uri: - ^/test/(.*)/(.*) - /$1-$2 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-rewrite-plugin-config spec: plugins: - name: proxy-rewrite config: regex_uri: - ^/test/(.*)/(.*) - /$1-$2 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /test/ filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` proxy-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-rewrite-route spec: ingressClassName: apisix http: - name: proxy-rewrite-route match: paths: - /test/* upstreams: - name: httpbin-external-domain plugins: - name: proxy-rewrite enable: true config: regex_uri: - ^/test/(.*)/(.*) - /$1-$2 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` Send a request to `/test/user/agent` to check if it is redirected to `/user-agent`: ``` curl "http://127.0.0.1:9080/test/user/agent" ``` You should see a response similar to the following: ``` { "user-agent": "curl/8.2.1" } ``` ### Add URL Parameters[​](#add-url-parameters "Direct link to Add URL Parameters") The following example demonstrates how you can add URL parameters to the request. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-rewrite-route", "methods": ["GET"], "uri": "/get", "plugins": { "proxy-rewrite": { "uri": "/get?arg1=apisix&arg2=plugin" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: proxy-rewrite-route methods: - GET plugins: proxy-rewrite: uri: /get?arg1=apisix&arg2=plugin upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-rewrite-plugin-config spec: plugins: - name: proxy-rewrite config: uri: /get?arg1=apisix&arg2=plugin --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` proxy-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-rewrite-route spec: ingressClassName: apisix http: - name: proxy-rewrite-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: proxy-rewrite enable: true config: uri: /get?arg1=apisix&arg2=plugin ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` Send a request to verify if the URL parameters are also forwarded to upstream: ``` curl "http://127.0.0.1:9080/get" ``` You should see a response similar to the following: ``` { "args": { "arg1": "apisix", "arg2": "plugin" }, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/8.2.1", "X-Amzn-Trace-Id": "Root=1-64fef6dc-2b0e09591db7353a275cdae4", "X-Forwarded-Host": "127.0.0.1" }, "origin": "127.0.0.1, 103.248.35.148", "url": "http://127.0.0.1/get?arg1=apisix&arg2=plugin" } ``` ### Rewrite HTTP Method[​](#rewrite-http-method "Direct link to Rewrite HTTP Method") The following example demonstrates how you can rewrite a GET request into a POST request. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "proxy-rewrite-route", "methods": ["GET"], "uri": "/get", "plugins": { "proxy-rewrite": { "uri": "/anything", "method":"POST" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: proxy-rewrite-route methods: - GET plugins: proxy-rewrite: uri: /anything method: POST upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-rewrite-plugin-config spec: plugins: - name: proxy-rewrite config: uri: /anything method: POST --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: proxy-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` proxy-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: proxy-rewrite-route spec: ingressClassName: apisix http: - name: proxy-rewrite-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: proxy-rewrite enable: true config: uri: /anything method: POST ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` Send a GET request to `/get` to verify if it is transformed into a POST request to `/anything`: ``` curl "http://127.0.0.1:9080/get" ``` You should see a response similar to the following: ``` { "args": {}, "data": "", "files": {}, "form": {}, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/8.2.1", "X-Amzn-Trace-Id": "Root=1-64fef7de-0c63387645353998196317f2", "X-Forwarded-Host": "127.0.0.1" }, "json": null, "method": "POST", "origin": "::1, 103.248.35.179", "url": "http://localhost/anything" } ``` ### Forward Consumer Names to Upstream[​](#forward-consumer-names-to-upstream "Direct link to Forward Consumer Names to Upstream") The following example demonstrates how you can forward the name of consumers who authenticates successfully to upstream services. As an example, you will be using [`key-auth`](https://docs.api7.ai/hub/key-auth.md) as the authentication method. * Admin API * ADC * Ingress Controller Create a consumer `JohnDoe`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "JohnDoe" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/JohnDoe/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Next, create a route with key authentication enabled, and configure `proxy-rewrite` to add consumer name to the header: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "consumer-restricted-route", "uri": "/get", "plugins": { "key-auth": {}, "proxy-rewrite": { "headers": { "set": { "X-Apisix-Consumer": "$consumer_name" }, "remove": [ "Apikey" ] } } }, "upstream" : { "nodes": { "httpbin.org":1 } } }' ``` Create a consumer with `key-auth` credential and a route with `key-auth` and `proxy-rewrite` plugins configured as such: adc.yaml ``` consumers: - username: JohnDoe credentials: - name: cred-john-key-auth type: key-auth config: key: john-key services: - name: httpbin routes: - name: consumer-restricted-route uris: - /get plugins: key-auth: {} proxy-rewrite: headers: set: X-Apisix-Consumer: $consumer_name remove: - Apikey upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: johndoe spec: gatewayRef: name: apisix credentials: - type: key-auth name: cred-john-key-auth config: key: john-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-rewrite-plugin-config spec: plugins: - name: key-auth config: _meta: disable: false - name: proxy-rewrite config: headers: set: X-Apisix-Consumer: $consumer_name remove: - Apikey --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: consumer-restricted-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` proxy-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: johndoe spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: consumer-restricted-route spec: ingressClassName: apisix http: - name: consumer-restricted-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: key-auth enable: true - name: proxy-rewrite enable: true config: headers: set: X-Apisix-Consumer: $consumer_name remove: - Apikey ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` ❶ Add the consumer name to the header `X-Apisix-Consumer` using the [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md#apisix-variables). ❷ Remove the authentication key so that it is not visible to the upstream service. Send a request to the route as consumer `JohnDoe`: ``` curl -i "http://127.0.0.1:9080/get" -H 'apikey: john-key' ``` You should receive an `HTTP/1.1 200 OK` response with the following body: ``` { "args": {}, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/8.4.0", "X-Amzn-Trace-Id": "Root=1-664b01a6-2163c0156ed4bff51d87d877", "X-Apisix-Consumer": "JohnDoe", "X-Forwarded-Host": "127.0.0.1" }, "origin": "172.19.0.1, 203.12.12.12", "url": "http://127.0.0.1/get" } ``` info When using the Ingress Controller, the consumer name is prefixed with the namespace. For example, a consumer named `JohnDoe` in namespace `aic` will appear as `aic_johndoe` in the `X-Apisix-Consumer` header. Send another request to the route without the valid credential: ``` curl -i "http://127.0.0.1:9080/get" ``` You should receive an `HTTP/1.1 401 Unauthorized` response. ### Dynamically Forward Requests in `radixtree_uri_with_parameter` Router Mode[​](#dynamically-forward-requests-in-radixtree_uri_with_parameter-router-mode "Direct link to dynamically-forward-requests-in-radixtree_uri_with_parameter-router-mode") The following example demonstrates how to extract part of the URL path using the `uri_param_*` variable and forward the value to the upstream service in a new header. This example assumes that APISIX is operating in the `radixtree_uri_with_parameter` [router mode](https://docs.api7.ai/apisix/reference/router-options.md). Create a route as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "httpbin", "uri": "/anything/user/:user_id/profile", "plugins":{ "proxy-rewrite": { "headers": { "set": { "X-User-ID": "$uri_param_user_id" } } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: httpbin uris: - /anything/user/:user_id/profile plugins: proxy-rewrite: headers: set: X-User-ID: $uri_param_user_id upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD proxy-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: proxy-rewrite-plugin-config spec: plugins: - name: proxy-rewrite config: headers: set: X-User-ID: $uri_param_user_id --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: httpbin spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything/user/ filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: proxy-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` proxy-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: httpbin spec: ingressClassName: apisix http: - name: httpbin match: paths: - /anything/user/*/profile upstreams: - name: httpbin-external-domain plugins: - name: proxy-rewrite enable: true config: headers: set: X-User-ID: $uri_param_user_id ``` Apply the configuration to your cluster: ``` kubectl apply -f proxy-rewrite-ic.yaml ``` ❶ Match requests to `/anything/user/:user_id/profile` where `user_id` is a parameter. ❷ Assign the `user_id` parameter value to a new header `X-User-ID`. Send a request to the route: ``` curl "http://127.0.0.1:9080/anything/user/123/profile" ``` You should see the following response: ``` { "args": {}, "data": "", "files": {}, "form": {}, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-68873cf5-7248f64d19d607ea50aa9735", "X-Forwarded-Host": "127.0.0.1", "X-User-Id": "123" }, ... } ``` The route parameter can also accept URL-encoded string. For instance, if you send a request as such: ``` curl -i "http://127.0.0.1:9080/anything/user/123%20456/profile" ``` The user ID would be extracted as `123 456`: ``` { "args": {}, "data": "", "files": {}, "form": {}, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-68873d37-7634825b20d05dee3a852cb9", "X-Forwarded-Host": "127.0.0.1", "X-User-Id": "123 456" }, ... } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * uri string *** New upstream URI path. Value could be a [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md). * method string vaild vaule: `GET`, `POST`, `PUT`, `HEAD`, `DELETE`, `OPTIONS`, `MKCOL`, `COPY`, `MOVE`, `PROPFIND`, `LOCK`, `UNLOCK`, `PATCH`, or `TRACE` *** HTTP method to rewrite requests to use. * regex\_uri array\[string] *** Regular expressions used to match the URI path from client requests and compose a new upstream URI path. When both `uri` and `regex_uri` are configured, `uri` has a higher priority. Provide one or more pairs in a flat array. Odd-positioned elements are the match pattern and even-positioned elements are the replacement. The plugin evaluates pairs in order and applies the first match. For example, with `["^/test/(.*)/(.*)", "/$1-$2", "^/other/(.*)", "/other"]`, a request to `/test/user/agent` is rewritten to `/user-agent`, and a request to `/other/hello` is rewritten to `/other`. * set\_ngx\_uri boolean default: `false` *** The parameter is currently only available in API7 Enterprise and will be updated to APISIX soon. If false, the value of `ngx.var.uri` will remain unchanged, preserving the original route `uri`. If true, `ngx.var.uri` will be updated to the `uri` value specified in `proxy-rewrite`. * host string *** Set [`Host`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Host) request header. * headers object *** Header actions to be executed. Can be set to objects of action verbs `add`, `remove`, and/or `set`; or an object consisting of headers to be `set`. When multiple action verbs are configured, actions are executed in the order of `add`, `remove`, and `set`. * add object *** Headers to append to requests. If a header is already present in the request, the header value will be appended. Header value could be set to a constant, one or more [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md), or the matched result of `regex_uri` using variables such as `$1-$2-$3`. A header value can also be an array. Each value is resolved independently and appended as a separate header line in array order. Introduced in API7 Enterprise 3.9.17 and 3.10.4, and APISIX 3.18.0. * set object *** Headers to set to requests. If a header is already present in the request, the header value will be overwritten. Header value could be set to a constant, one or more [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md), or the matched result of `regex_uri` using variables such as `$1-$2-$3`. A header value can also be an array. Each value is resolved independently, and the array replaces any existing value with separate header lines in array order. Introduced in API7 Enterprise 3.9.17 and 3.10.4, and APISIX 3.18.0. Should not be used to set `Host`. * remove array\[string] *** Headers to remove from requests. * use\_real\_request\_uri\_unsafe boolean default: `false` *** If true, bypass URI normalization and allow for the full original request URI. Enabling this option is considered unsafe. --- # public-api The `public-api` plugin exposes an internal API endpoint, making it publicly accessible. One of the primary use cases of this plugin is to expose internal endpoints created by other plugins. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `public-api` in different scenarios. ### Expose Prometheus Metrics at Custom Endpoint[​](#expose-prometheus-metrics-at-custom-endpoint "Direct link to Expose Prometheus Metrics at Custom Endpoint") The following example demonstrates how you can disable the Prometheus export server that, by default, exposes an endpoint on port `9091`, and expose APISIX Prometheus metrics on a new public API endpoint on port `9080`, which APISIX uses to listen to other client requests. You will also configure the route such that the internal endpoint `/apisix/prometheus/metrics` is exposed at a custom endpoint. caution If a large quantity of metrics is being collected, the plugin could take up a significant amount of CPU resources for metric computations and negatively impact the processing of regular requests. To address this issue, APISIX uses [privileged agent](https://github.com/openresty/lua-resty-core/blob/master/lib/ngx/process.md#enable_privileged_agent) and offloads the metric computations to a separate process. This optimization applies automatically if you use the metric endpoint configured under `plugin_attr.prometheus.export_addr` in the configuration file. If you expose the metric endpoint with the `public-api` plugin, you will not benefit from this optimization. To expose metrics through `public-api`, first disable the default Prometheus export server: * Host or Docker * Kubernetes (Helm) Add or update this section in the gateway configuration file: config.yaml ``` plugin_attr: prometheus: enable_export_server: false ``` Reload the gateway for changes to take effect. For Helm deployments, update the chart values that render `plugin_attr.prometheus`. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` apisix: pluginAttrs: prometheus: enable_export_server: false ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` pluginAttrs: prometheus: enable_export_server: false ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` Next, create a route with `public-api` plugin and expose a public API endpoint for APISIX metrics: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "prometheus-metrics", "uri": "/prometheus_metrics", "plugins": { "public-api": { "uri": "/apisix/prometheus/metrics" } } }' ``` adc.yaml ``` services: - name: public-api-metrics-service routes: - name: prometheus-metrics uris: - /prometheus_metrics plugins: public-api: uri: /apisix/prometheus/metrics ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD public-api-ic.yaml ``` apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: prometheus-metrics spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /prometheus_metrics filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: public-api-metrics-plugin-config --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: public-api-metrics-plugin-config spec: plugins: - name: public-api config: uri: /apisix/prometheus/metrics ``` public-api-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: prometheus-metrics spec: ingressClassName: apisix http: - name: prometheus-metrics match: paths: - /prometheus_metrics plugins: - name: public-api enable: true config: uri: /apisix/prometheus/metrics ``` Apply the configuration: ``` kubectl apply -f public-api-ic.yaml ``` ❶ Set the route `uri` to the custom endpoint path. ❷ Set the plugin `uri` to the internal endpoint to be exposed. Send a request to the custom metrics endpoint: ``` curl "http://127.0.0.1:9080/prometheus_metrics" ``` You should see an output similar to the following: ``` # HELP apisix_http_requests_total The total number of client requests since APISIX started # TYPE apisix_http_requests_total gauge apisix_http_requests_total 1 # HELP apisix_nginx_http_current_connections Number of HTTP connections # TYPE apisix_nginx_http_current_connections gauge apisix_nginx_http_current_connections{state="accepted"} 1 apisix_nginx_http_current_connections{state="active"} 1 apisix_nginx_http_current_connections{state="handled"} 1 apisix_nginx_http_current_connections{state="reading"} 0 apisix_nginx_http_current_connections{state="waiting"} 0 apisix_nginx_http_current_connections{state="writing"} 1 ... ``` ### Expose Batch Requests Endpoint[​](#expose-batch-requests-endpoint "Direct link to Expose Batch Requests Endpoint") The following example demonstrates how you can use the `public-api` plugin to expose an endpoint for the `batch-requests` plugin, which is used for assembling multiple requests into one single request before sending them to the gateway. Create a sample route to httpbin's `/anything` endpoint for verification purpose: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "httpbin-anything", "uri": "/anything", "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: httpbin-anything uris: - /anything upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD public-api-httpbin-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: httpbin-anything spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything backendRefs: - name: httpbin-external-domain port: 80 ``` public-api-httpbin-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: httpbin-anything spec: ingressClassName: apisix http: - name: httpbin-anything match: paths: - /anything upstreams: - name: httpbin-external-domain ``` Apply the configuration: ``` kubectl apply -f public-api-httpbin-ic.yaml ``` Create a route with `public-api` plugin and set the route `uri` to the internal endpoint to be exposed: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "batch-requests", "uri": "/apisix/batch-requests", "plugins": { "public-api": {} } }' ``` adc.yaml ``` services: - name: public-api-batch-service routes: - name: batch-requests uris: - /apisix/batch-requests plugins: public-api: {} ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD public-api-batch-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: public-api-batch-plugin-config spec: plugins: - name: public-api config: {} --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: batch-requests spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /apisix/batch-requests filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: public-api-batch-plugin-config ``` public-api-batch-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: batch-requests spec: ingressClassName: apisix http: - name: batch-requests match: paths: - /apisix/batch-requests plugins: - name: public-api enable: true ``` Apply the configuration: ``` kubectl apply -f public-api-batch-ic.yaml ``` Send a pipelined request consisting of a GET and a POST request to the exposed batch requests endpoint: ``` curl "http://127.0.0.1:9080/apisix/batch-requests" -X POST -d ' { "pipeline": [ { "method": "GET", "path": "/anything" }, { "method": "POST", "path": "/anything", "body": "a post request" } ] }' ``` You should receive responses from both requests, similar to the following: ``` [ { "reason": "OK", "body": "{\n \"args\": {}, \n \"data\": \"\", \n \"files\": {}, \n \"form\": {}, \n \"headers\": {\n \"Accept\": \"*/*\", \n \"Host\": \"127.0.0.1\", \n \"User-Agent\": \"curl/8.6.0\", \n \"X-Amzn-Trace-Id\": \"Root=1-67b6e33b-5a30174f5534287928c54ca9\", \n \"X-Forwarded-Host\": \"127.0.0.1\"\n }, \n \"json\": null, \n \"method\": \"GET\", \n \"origin\": \"192.168.107.1, 43.252.208.84\", \n \"url\": \"http://127.0.0.1/anything\"\n}\n", "headers": { ... }, "status": 200 }, { "reason": "OK", "body": "{\n \"args\": {}, \n \"data\": \"a post request\", \n \"files\": {}, \n \"form\": {}, \n \"headers\": {\n \"Accept\": \"*/*\", \n \"Content-Length\": \"14\", \n \"Host\": \"127.0.0.1\", \n \"User-Agent\": \"curl/8.6.0\", \n \"X-Amzn-Trace-Id\": \"Root=1-67b6e33b-0eddcec07f154dac0d77876f\", \n \"X-Forwarded-Host\": \"127.0.0.1\"\n }, \n \"json\": null, \n \"method\": \"POST\", \n \"origin\": \"192.168.107.1, 43.252.208.84\", \n \"url\": \"http://127.0.0.1/anything\"\n}\n", "headers": { ... }, "status": 200 } ] ``` If you would like to expose the batch requests endpoint at a custom endpoint, create a route with `public-api` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "batch-requests", "uri": "/batch-requests", "plugins": { "public-api": { "uri": "/apisix/batch-requests" } } }' ``` adc.yaml ``` services: - name: public-api-batch-service routes: - name: batch-requests uris: - /batch-requests plugins: public-api: uri: /apisix/batch-requests ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD public-api-batch-ic.yaml ``` apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: batch-requests spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /batch-requests filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: public-api-batch-plugin-config --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: public-api-batch-plugin-config spec: plugins: - name: public-api config: uri: /apisix/batch-requests ``` public-api-batch-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: batch-requests spec: ingressClassName: apisix http: - name: batch-requests match: paths: - /batch-requests plugins: - name: public-api enable: true config: uri: /apisix/batch-requests ``` Apply the configuration: ``` kubectl apply -f public-api-batch-ic.yaml ``` ❶ Set the route `uri` to the custom endpoint path. ❷ Set the plugin `uri` to the internal endpoint to be exposed. The batch requests endpoint should now be exposed as `/batch-requests`, instead of `/apisix/batch-requests`. Send a pipelined request consisting of a GET and a POST request to the exposed batch requests endpoint: ``` curl "http://127.0.0.1:9080/batch-requests" -X POST -d ' { "pipeline": [ { "method": "GET", "path": "/anything" }, { "method": "POST", "path": "/anything", "body": "a post request" } ] }' ``` You should receive responses from both requests, similar to the following: ``` [ { "reason": "OK", "body": "{\n \"args\": {}, \n \"data\": \"\", \n \"files\": {}, \n \"form\": {}, \n \"headers\": {\n \"Accept\": \"*/*\", \n \"Host\": \"127.0.0.1\", \n \"User-Agent\": \"curl/8.6.0\", \n \"X-Amzn-Trace-Id\": \"Root=1-67b6e33b-5a30174f5534287928c54ca9\", \n \"X-Forwarded-Host\": \"127.0.0.1\"\n }, \n \"json\": null, \n \"method\": \"GET\", \n \"origin\": \"192.168.107.1, 43.252.208.84\", \n \"url\": \"http://127.0.0.1/anything\"\n}\n", "headers": { ... }, "status": 200 }, { "reason": "OK", "body": "{\n \"args\": {}, \n \"data\": \"a post request\", \n \"files\": {}, \n \"form\": {}, \n \"headers\": {\n \"Accept\": \"*/*\", \n \"Content-Length\": \"14\", \n \"Host\": \"127.0.0.1\", \n \"User-Agent\": \"curl/8.6.0\", \n \"X-Amzn-Trace-Id\": \"Root=1-67b6e33b-0eddcec07f154dac0d77876f\", \n \"X-Forwarded-Host\": \"127.0.0.1\"\n }, \n \"json\": null, \n \"method\": \"POST\", \n \"origin\": \"192.168.107.1, 43.252.208.84\", \n \"url\": \"http://127.0.0.1/anything\"\n}\n", "headers": { ... }, "status": 200 } ] ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * uri string *** Internal endpoint to expose. If not configured, expose the route URI. --- # real-ip The `real-ip` plugin allows APISIX to set the client's real IP by IP address passed in the HTTP header or HTTP query string. This is particularly useful when APISIX is behind a reverse proxy, since the proxy could act as the request originating client otherwise. The plugin is functionally similar to NGINX's [ngx\_http\_realip\_module](https://nginx.org/en/docs/http/ngx_http_realip_module.html) but offers more flexibilities. ## Trust Forwarded Client Information[​](#trust-forwarded-client-information "Direct link to Trust Forwarded Client Information") Two trust settings apply at different scopes: * Global `apisix.trusted_addresses` identifies immediate reverse proxies whose `Forwarded` and `X-Forwarded-*` headers the gateway can retain. Features such as automatic OpenID Connect redirect URI construction use forwarded scheme, host, and port values only for requests from these addresses. This setting does not itself rewrite the gateway's client IP. * The plugin field `trusted_addresses` limits which immediate peers can supply the value configured by `source` on a route. When the peer is trusted, the plugin rewrites the gateway's client IP from that source. Configure only proxy addresses you operate. Trusting arbitrary addresses allows a client to forge its apparent IP address, scheme, host, or port. When a peer fails the applicable trust check, the gateway retains the direct connection address and replaces or clears untrusted forwarded origin headers. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `real-ip` in different scenarios. ### Obtain Real Client Address From URI Parameter[​](#obtain-real-client-address-from-uri-parameter "Direct link to Obtain Real Client Address From URI Parameter") The following example demonstrates how to update client IP address with an URI parameter. Create a route as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "real-ip-route", "uri": "/get", "plugins": { "real-ip": { "source": "arg_realip", "trusted_addresses": ["127.0.0.0/24"] }, "response-rewrite": { "headers": { "remote_addr": "$remote_addr", "remote_port": "$remote_port" } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: real-ip-route uris: - /get plugins: real-ip: source: arg_realip trusted_addresses: - 127.0.0.0/24 response-rewrite: headers: remote_addr: $remote_addr remote_port: $remote_port upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD real-ip-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: real-ip-plugin-config spec: plugins: - name: real-ip config: source: arg_realip trusted_addresses: - 127.0.0.0/24 - name: response-rewrite config: headers: remote_addr: $remote_addr remote_port: $remote_port --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: real-ip-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: real-ip-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` real-ip-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: real-ip-route spec: ingressClassName: apisix http: - name: real-ip-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: real-ip config: source: arg_realip trusted_addresses: - 127.0.0.0/24 - name: response-rewrite config: headers: remote_addr: $remote_addr remote_port: $remote_port ``` Apply the configuration: ``` kubectl apply -f real-ip-ic.yaml ``` ❶ Configure `source` to obtain value from the URL parameter `realip` using the [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). ❷ Use the `response-rewrite` plugin to set response headers to verify if the client IP and port were actually updated. Send a request to the route with real IP and port in the URL parameter: ``` curl -i "http://127.0.0.1:9080/get?realip=1.2.3.4:9080" ``` You should see the response includes the following headers: ``` remote_addr: 1.2.3.4 remote_port: 9080 ``` ### Obtain Real Client Address From Header[​](#obtain-real-client-address-from-header "Direct link to Obtain Real Client Address From Header") The following example shows how to set the real client IP when APISIX is behind a reverse proxy, such as a load balancer, when the proxy exposes the real client IP in the [`X-Forwarded-For`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/X-Forwarded-For) header. Create a route as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "real-ip-route", "uri": "/get", "plugins": { "real-ip": { "source": "http_x_forwarded_for", "trusted_addresses": ["127.0.0.0/24"] }, "response-rewrite": { "headers": { "remote_addr": "$remote_addr" } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: real-ip-route uris: - /get plugins: real-ip: source: http_x_forwarded_for trusted_addresses: - 127.0.0.0/24 response-rewrite: headers: remote_addr: $remote_addr upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD real-ip-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: real-ip-plugin-config spec: plugins: - name: real-ip config: source: http_x_forwarded_for trusted_addresses: - 127.0.0.0/24 - name: response-rewrite config: headers: remote_addr: $remote_addr --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: real-ip-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: real-ip-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` real-ip-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: real-ip-route spec: ingressClassName: apisix http: - name: real-ip-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: real-ip config: source: http_x_forwarded_for trusted_addresses: - 127.0.0.0/24 - name: response-rewrite config: headers: remote_addr: $remote_addr ``` Apply the configuration: ``` kubectl apply -f real-ip-ic.yaml ``` ❶ Configure `source` to obtain value from the request header `X-Forwarded-For` using the [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). ❷ Use the `response-rewrite` plugin to set a response header to verify if the client IP was actually updated. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/get" \ -H "X-Forwarded-For: 10.26.3.19" ``` You should see a response including the following header: ``` remote_addr: 10.26.3.19 ``` The IP address should correspond to the IP address of the request originating client. ### Obtain Real Client Address Behind Multiple Proxies[​](#obtain-real-client-address-behind-multiple-proxies "Direct link to Obtain Real Client Address Behind Multiple Proxies") The following example shows how to get the real client IP when APISIX is behind multiple proxies, which causes [`X-Forwarded-For`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/X-Forwarded-For) header to include a list of proxy IP addresses. Create a route as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "real-ip-route", "uri": "/get", "plugins": { "real-ip": { "source": "http_x_forwarded_for", "recursive": true, "trusted_addresses": ["192.128.0.0/16", "127.0.0.1/32"] }, "response-rewrite": { "headers": { "remote_addr": "$remote_addr" } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: real-ip-route uris: - /get plugins: real-ip: source: http_x_forwarded_for recursive: true trusted_addresses: - 192.128.0.0/16 - 127.0.0.1/32 response-rewrite: headers: remote_addr: $remote_addr upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD real-ip-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: real-ip-plugin-config spec: plugins: - name: real-ip config: source: http_x_forwarded_for recursive: true trusted_addresses: - 192.128.0.0/16 - 127.0.0.1/32 - name: response-rewrite config: headers: remote_addr: $remote_addr --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: real-ip-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: real-ip-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` real-ip-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: real-ip-route spec: ingressClassName: apisix http: - name: real-ip-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: real-ip config: source: http_x_forwarded_for recursive: true trusted_addresses: - 192.128.0.0/16 - 127.0.0.1/32 - name: response-rewrite config: headers: remote_addr: $remote_addr ``` Apply the configuration: ``` kubectl apply -f real-ip-ic.yaml ``` ❶ Configure `source` to obtain value from the request header `X-Forwarded-For` using the [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). ❷ Set `recursive` to `true` so that the original client address that matches one of the trusted addresses is replaced by the last non-trusted address sent in the configured `source`. ❸ Use the `response-rewrite` plugin to set a response header to verify if the client IP was actually updated. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/get" \ -H "X-Forwarded-For: 127.0.0.2, 192.128.1.1, 127.0.0.1" ``` You should see a response including the following header: ``` remote_addr: 127.0.0.2 ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * source string required *** A [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md), such as `http_x_forwarded_for` or `arg_realip`. The variable value should be a valid IP address that represents the client's real IP address, with an optional port. * trusted\_addresses array\[string] vaild vaule: array of IPv4 or IPv6 addresses (CIDR notation acceptable) *** Trusted addresses that are known to send correct replacement addresses. This configuration sets the [`set_real_ip_from`](https://nginx.org/en/docs/http/ngx_http_realip_module.html#set_real_ip_from) directive. * recursive boolean default: `false` *** If false, replace the original client address that matches one of the trusted addresses by the last address sent in the configured `source`. If true, replace the original client address that matches one of the trusted addresses by the last non-trusted address sent in the configured `source`. --- # request-id The `request-id` plugin assigns a unique ID to each request proxied through the gateway, which can be used for request tracking and debugging. If a request already includes an ID in the header specified by `header_name`, the plugin uses that value instead of generating a new one. The request ID is included in gateway logs by default. When the plugin is enabled, it is also added to the response header. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `request-id` in different scenarios. ### Understand Request ID in Gateway Logs[​](#understand-request-id-in-gateway-logs "Direct link to Understand Request ID in Gateway Logs") The request ID is included in both access and error logs in APISIX from version 3.15.0 and in API7 Enterprise from version 3.3.0, regardless of whether the plugin is enabled. * When the plugin is disabled, the request ID defaults to Nginx's built-in `$request_id`. * When the plugin is enabled, the request ID is set to the unique ID generated by the plugin. This ensures request tracing is always available, with enhanced functionality when the plugin is enabled. The following example demonstrates how the request ID appears in the gateway logs when the plugin is disabled and when it is enabled. Create a route without the `request-id` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-id-route", "uri": "/anything", "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: request-id-service routes: - name: request-id-route uris: - /anything upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD request-id-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-id-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything backendRefs: - name: httpbin-external-domain port: 80 ``` request-id-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-id-route spec: ingressClassName: apisix http: - name: request-id-route match: paths: - /anything upstreams: - name: httpbin-external-domain ``` Apply the configuration to your cluster: ``` kubectl apply -f request-id-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. In the gateway log, you should see an entry similar to the following, where the last value is the request ID from Nginx's built-in `$request_id`: ``` 192.168.215.1 - - [30/Jan/2026:07:21:31 +0000] localhost:9080 "GET /anything HTTP/1.1" 200 391 1.657 "-" "curl/8.6.0" 3.210.41.225:80 200 1.608 "http://localhost:9080" "8a14012e5d0414aff4f15f04b0bd8cb9" ``` Update the route with the `request-id` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-id-route", "uri": "/anything", "plugins": { "request-id": {} }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: request-id-service routes: - name: request-id-route uris: - /anything plugins: request-id: {} upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD request-id-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: request-id-plugin-config spec: plugins: - name: request-id config: _meta: disable: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-id-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: request-id-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` request-id-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-id-route spec: ingressClassName: apisix http: - name: request-id-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: request-id enable: true ``` Apply the configuration to your cluster: ``` kubectl apply -f request-id-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. In the gateway log, you should see an entry similar to the following, where the last value is the request ID generated by the plugin: ``` 192.168.215.1 - - [30/Jan/2026:07:36:24 +0000] localhost:9080 "GET /anything HTTP/1.1" 200 391 0.685 "-" "curl/8.6.0" 52.20.30.6:80 200 0.653 "http://localhost:9080" "8c0ac818-f9d6-4160-be60-8fc74e76be73" ``` ### Attach Request ID to Default Response Header[​](#attach-request-id-to-default-response-header "Direct link to Attach Request ID to Default Response Header") The following example demonstrates how to configure `request-id` on a route which attaches a generated request ID to the default `X-Request-Id` response header, if the header value is not passed in the request. When the `X-Request-Id` header is set in the request, the plugin will take the value in the request header as the request ID. Create a route with the `request-id` plugin using its default configurations (explicitly defined): * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-id-route", "uri": "/anything", "plugins": { "request-id": { "header_name": "X-Request-Id", "include_in_response": true, "algorithm": "uuid" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: request-id-service routes: - name: request-id-route uris: - /anything plugins: request-id: header_name: X-Request-Id include_in_response: true algorithm: uuid upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD request-id-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: request-id-plugin-config spec: plugins: - name: request-id config: header_name: X-Request-Id include_in_response: true algorithm: uuid --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-id-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: request-id-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` request-id-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-id-route spec: ingressClassName: apisix http: - name: request-id-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: request-id enable: true config: header_name: X-Request-Id include_in_response: true algorithm: uuid ``` Apply the configuration to your cluster: ``` kubectl apply -f request-id-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response and see the response includes the `X-Request-Id` header with a generated ID: ``` X-Request-Id: b9b2c0d4-d058-46fa-bafc-dd91a0ccf441 ``` Send a request to the route with a custom request ID in the header: ``` curl -i "http://127.0.0.1:9080/anything" -H 'X-Request-Id: some-custom-request-id' ``` You should receive an `HTTP/1.1 200 OK` response and see the response includes the `X-Request-Id` header with the custom request ID: ``` X-Request-Id: some-custom-request-id ``` ### Attach Request ID to Custom Response Header[​](#attach-request-id-to-custom-response-header "Direct link to Attach Request ID to Custom Response Header") The following example demonstrates how to configure `request-id` on a route which attaches a generated request ID to a specified header. Create a route with the `request-id` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-id-route", "uri": "/anything", "plugins": { "request-id": { "header_name": "X-Req-Identifier", "include_in_response": true } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: request-id-service routes: - name: request-id-route uris: - /anything plugins: request-id: header_name: X-Req-Identifier include_in_response: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD request-id-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: request-id-plugin-config spec: plugins: - name: request-id config: header_name: X-Req-Identifier include_in_response: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-id-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: request-id-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` request-id-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-id-route spec: ingressClassName: apisix http: - name: request-id-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: request-id enable: true config: header_name: X-Req-Identifier include_in_response: true ``` Apply the configuration to your cluster: ``` kubectl apply -f request-id-ic.yaml ``` ❶ Define a custom header that carries the request ID. ❷ Include the request ID in the response header. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response and see the response includes the `X-Req-Identifier` header with a generated ID: ``` X-Req-Identifier: 1c42ff59-ee4c-4103-a980-8359f4135b21 ``` ### Hide Request ID in Response Header[​](#hide-request-id-in-response-header "Direct link to Hide Request ID in Response Header") The following example demonstrates how to configure `request-id` on a route which attaches a generated request ID to a specified header. The header containing the request ID should be forwarded to the upstream service but not returned in the response header. Create a route with the `request-id` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-id-route", "uri": "/anything", "plugins": { "request-id": { "header_name": "X-Req-Identifier", "include_in_response": false } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: request-id-service routes: - name: request-id-route uris: - /anything plugins: request-id: header_name: X-Req-Identifier include_in_response: false upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD request-id-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: request-id-plugin-config spec: plugins: - name: request-id config: header_name: X-Req-Identifier include_in_response: false --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-id-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: request-id-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` request-id-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-id-route spec: ingressClassName: apisix http: - name: request-id-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: request-id enable: true config: header_name: X-Req-Identifier include_in_response: false ``` Apply the configuration to your cluster: ``` kubectl apply -f request-id-ic.yaml ``` ❶ Define a custom header that carries the request ID. ❷ Do not include the request ID in the response header. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response not and see `X-Req-Identifier` header among the response headers. In the response body, you should see: ``` { "args": {}, "data": "", "files": {}, "form": {}, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/8.6.0", "X-Amzn-Trace-Id": "Root=1-6752748c-7d364f48564508db1e8c9ea8", "X-Forwarded-Host": "127.0.0.1", "X-Req-Identifier": "268092bc-15e1-4461-b277-bf7775f2856f" }, ... } ``` This shows the request ID is forwarded to the upstream service but not returned in the response header. ### Generate Time-Ordered UUID v7 IDs[​](#generate-time-ordered-uuid-v7-ids "Direct link to Generate Time-Ordered UUID v7 IDs") The following example configures `request-id` to generate RFC 9562 UUID v7 values. A millisecond timestamp and worker-local sequence make the IDs lexicographically sortable and monotonic within one worker, but ordering is not coordinated across workers or gateway instances. UUID v7 support was introduced in API7 Enterprise 3.9.8 and APISIX 3.17.0. To generate compact URL-safe IDs without time ordering instead, set `algorithm` to `nanoid`. Create a route with the `request-id` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-id-route", "uri": "/anything", "plugins": { "request-id": { "algorithm": "uuidv7" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: request-id-service routes: - name: request-id-route uris: - /anything plugins: request-id: algorithm: uuidv7 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD request-id-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: request-id-plugin-config spec: plugins: - name: request-id config: algorithm: uuidv7 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-id-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: request-id-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` request-id-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-id-route spec: ingressClassName: apisix http: - name: request-id-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: request-id enable: true config: algorithm: uuidv7 ``` Apply the configuration to your cluster: ``` kubectl apply -f request-id-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response and see that the response includes the `X-Request-Id` header with a UUID v7 value: ``` X-Request-Id: 0194f4d8-e8d7-7d39-8a6b-6f18f5c95b46 ``` To use `nanoid` instead, change `algorithm` to `nanoid` in the same route configuration, apply the update, and send the request again: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `X-Request-Id` header containing a compact, URL-safe ID similar to the following: ``` X-Request-Id: kepgHWCH2ycQ6JknQKrX2 ``` ### Attach Request ID Globally and on a Route[​](#attach-request-id-globally-and-on-a-route "Direct link to Attach Request ID Globally and on a Route") The following example demonstrates how to configure `request-id` as a global plugin and on a route to attach two IDs. * Admin API * ADC * Ingress Controller Create a global rule for the `request-id` plugin which adds request ID to a custom header: ``` curl -i "http://127.0.0.1:9180/apisix/admin/global_rules" -X PUT -d '{ "id": "rule-for-request-id", "plugins": { "request-id": { "header_name": "Global-Request-ID" } } }' ``` Create a route with the `request-id` plugin which adds request ID to a different custom header: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-id-route", "uri": "/anything", "plugins": { "request-id": { "header_name": "Route-Request-ID" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Configure the `request-id` plugin as a global rule and on a route: adc.yaml ``` global_rules: - id: rule-for-request-id plugins: request-id: header_name: Global-Request-ID services: - name: request-id-service routes: - name: request-id-route uris: - /anything plugins: request-id: header_name: Route-Request-ID upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update your GatewayProxy manifest to enable `request-id` as a global plugin: gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration plugins: - name: request-id enabled: true config: header_name: Global-Request-ID ``` Create a route with the `request-id` plugin which adds request ID to a different custom header: request-id-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: request-id-plugin-config spec: plugins: - name: request-id config: header_name: Route-Request-ID --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-id-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: request-id-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f gatewayproxy.yaml -f request-id-ic.yaml ``` Create a Kubernetes manifest file for a global `request-id` plugin: global-request-id.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixGlobalRule metadata: namespace: aic name: global-request-id spec: ingressClassName: apisix plugins: - name: request-id enable: true config: header_name: Global-Request-ID ``` Create a route with the `request-id` plugin which adds request ID to a different custom header: request-id-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-id-route spec: ingressClassName: apisix http: - name: request-id-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: request-id enable: true config: header_name: Route-Request-ID ``` Apply the configuration to your cluster: ``` kubectl apply -f global-request-id.yaml -f request-id-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response and see the response includes the following headers: ``` Global-Request-ID: 2e9b99c1-08ed-4a74-b347-49c0891b07ad Route-Request-ID: d755666b-732c-4f0e-a30e-a7a71ace4e26 ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * header\_name string default: `X-Request-Id` *** Specify the name of the header that carries the request unique ID. Note that if a request carries an ID in the `header_name` header, the plugin will use the header value as the unique ID and will not overwrite it with the generated ID. * include\_in\_response boolean default: `true` *** If true, include the generated request ID in the response header, where the name of the header is the `header_name` value. * algorithm string default: `uuid` vaild vaule: `uuid`, `nanoid`, `range_id`, `ksuid`, or `uuidv7` *** Specify the algorithm used for generating the unique ID. When set to `uuid`, the plugin generates a universally unique identifier. When set to `nanoid`, the plugin generates a compact, URL-safe ID. When set to `range_id`, the plugin generates a sequential ID with specific parameters. When set to `ksuid`, the plugin generates a time-sortable, globally unique ID. KSUID support was introduced in API7 Enterprise 3.9.0 and APISIX 3.14.0. When set to `uuidv7`, the plugin generates an RFC 9562 UUID v7 identifier with a millisecond timestamp, a worker-local sequence, and random bits. UUID v7 values are monotonic within one worker, but ordering is not coordinated across workers or gateway instances. UUID v7 support was introduced in API7 Enterprise 3.9.8 and APISIX 3.17.0. * range\_id object *** Define the configuration for generating a request ID using the `range_id` algorithm. * char\_set string default: `abcdefghijklmnopqrstuvwxyzABCDEFGHIGKLMNOPQRSTUVWXYZ0123456789` vaild vaule: minimum length 6 *** Specify the character set used for the `range_id` algorithm. * length integer default: `16` vaild vaule: greater than or equal to 6 *** Set the length of the generated ID for the `range_id` algorithm. --- # request-validation The `request-validation` plugin validates requests before forwarding them to upstream services. This plugin uses [JSON Schema](https://github.com/api7/jsonschema) for validation and can validate headers and body of a request. See [JSON schema specification](https://json-schema.org/specification) to learn more about the syntax. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `request-validation` for different scenarios. ### Validate Request Header[​](#validate-request-header "Direct link to Validate Request Header") The following example demonstrates how to validate request headers against a defined JSON schema. Create a route with `request-validation` plugin as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-validation-route", "uri": "/get", "plugins": { "request-validation": { "header_schema": { "type": "object", "required": ["User-Agent", "Host"], "properties": { "User-Agent": { "type": "string", "pattern": "^curl\/" }, "Host": { "type": "string", "enum": ["httpbin.org", "httpbin"] } } } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: request-validation-service routes: - name: request-validation-route uris: - /get plugins: request-validation: header_schema: type: object required: - User-Agent - Host properties: User-Agent: type: string pattern: "^curl/" Host: type: string enum: - httpbin.org - httpbin upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD request-validation-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: request-validation-plugin-config spec: plugins: - name: request-validation config: header_schema: type: object required: - User-Agent - Host properties: User-Agent: type: string pattern: "^curl/" Host: type: string enum: - httpbin.org - httpbin --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-validation-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: request-validation-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` request-validation-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-validation-route spec: ingressClassName: apisix http: - name: request-validation-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: request-validation enable: true config: header_schema: type: object required: - User-Agent - Host properties: User-Agent: type: string pattern: "^curl/" Host: type: string enum: - httpbin.org - httpbin ``` Apply the configuration to your cluster: ``` kubectl apply -f request-validation-ic.yaml ``` ❶ `required`: require requests to include the specified headers. ❷ `properties`: require headers to conform to the specified requirements. #### Verify with Request Conforming to the Schema[​](#verify-with-request-conforming-to-the-schema "Direct link to Verify with Request Conforming to the Schema") Send a request with header `Host: httpbin`, which complies with the schema: ``` curl -i "http://127.0.0.1:9080/get" -H "Host: httpbin" ``` You should receive an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "headers": { "Accept": "*/*", "Host": "httpbin", "User-Agent": "curl/7.74.0", "X-Amzn-Trace-Id": "Root=1-6509ae35-63d1e0fd3934e3f221a95dd8", "X-Forwarded-Host": "httpbin" }, "origin": "127.0.0.1, 183.17.233.107", "url": "http://httpbin/get" } ``` #### Verify with Request Not Conforming to the Schema[​](#verify-with-request-not-conforming-to-the-schema "Direct link to Verify with Request Not Conforming to the Schema") Send a request without any header: ``` curl -i "http://127.0.0.1:9080/get" ``` You should receive an `HTTP/1.1 400 Bad Request` response, showing that the request fails to pass validation: ``` property "Host" validation failed: matches none of the enum value ``` Send a request with the required headers but with non-conformant header value: ``` curl -i "http://127.0.0.1:9080/get" -H "Host: httpbin" -H "User-Agent: cli-mock" ``` You should receive an `HTTP/1.1 400 Bad Request` response showing the `User-Agent` header value does not match the expected pattern: ``` property "User-Agent" validation failed: failed to match pattern "^curl/" with "cli-mock" ``` ### Customize Rejection Message and Status Code[​](#customize-rejection-message-and-status-code "Direct link to Customize Rejection Message and Status Code") The following example demonstrates how to customize response status and message when the validation fails. Configure the route with `request-validation` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-validation-route", "uri": "/get", "plugins": { "request-validation": { "header_schema": { "type": "object", "required": ["Host"], "properties": { "Host": { "type": "string", "enum": ["httpbin.org", "httpbin"] } } }, "rejected_code": 403, "rejected_msg": "Request header validation failed." } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: request-validation-service routes: - name: request-validation-route uris: - /get plugins: request-validation: header_schema: type: object required: - Host properties: Host: type: string enum: - httpbin.org - httpbin rejected_code: 403 rejected_msg: "Request header validation failed." upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD request-validation-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: request-validation-plugin-config spec: plugins: - name: request-validation config: header_schema: type: object required: - Host properties: Host: type: string enum: - httpbin.org - httpbin rejected_code: 403 rejected_msg: "Request header validation failed." --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-validation-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: request-validation-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` request-validation-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-validation-route spec: ingressClassName: apisix http: - name: request-validation-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: request-validation enable: true config: header_schema: type: object required: - Host properties: Host: type: string enum: - httpbin.org - httpbin rejected_code: 403 rejected_msg: "Request header validation failed." ``` Apply the configuration to your cluster: ``` kubectl apply -f request-validation-ic.yaml ``` ❶ `rejected_code`: customize rejection code. ❷ `rejected_msg`: customize rejection message. Send a request with a misconfigured `Host` in the header: ``` curl -i "http://127.0.0.1:9080/get" -H "Host: httpbin2" ``` You should receive an `HTTP/1.1 403 Forbidden` response with the custom message: ``` Request header validation failed. ``` ### Validate Request Body[​](#validate-request-body "Direct link to Validate Request Body") The following example demonstrates how to validate request body against a defined JSON schema. The `request-validation` plugin supports validation of two types of media types: * `application/json` * `application/x-www-form-urlencoded` #### Validate JSON Request Body[​](#validate-json-request-body "Direct link to Validate JSON Request Body") Create a route with `request-validation` plugin as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-validation-route", "uri": "/post", "plugins": { "request-validation": { "header_schema": { "type": "object", "required": ["Content-Type"], "properties": { "Content-Type": { "type": "string", "pattern": "^application\/json$" } } }, "body_schema": { "type": "object", "required": ["required_payload"], "properties": { "required_payload": {"type": "string"}, "boolean_payload": {"type": "boolean"}, "array_payload": { "type": "array", "minItems": 1, "items": { "type": "integer", "minimum": 200, "maximum": 599 }, "uniqueItems": true, "default": [200] } } } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: request-validation-service routes: - name: request-validation-route uris: - /post plugins: request-validation: header_schema: type: object required: - Content-Type properties: Content-Type: type: string pattern: "^application/json$" body_schema: type: object required: - required_payload properties: required_payload: type: string boolean_payload: type: boolean array_payload: type: array minItems: 1 items: type: integer minimum: 200 maximum: 599 uniqueItems: true default: - 200 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD request-validation-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: request-validation-plugin-config spec: plugins: - name: request-validation config: header_schema: type: object required: - Content-Type properties: Content-Type: type: string pattern: "^application/json$" body_schema: type: object required: - required_payload properties: required_payload: type: string boolean_payload: type: boolean array_payload: type: array minItems: 1 items: type: integer minimum: 200 maximum: 599 uniqueItems: true default: - 200 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-validation-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /post filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: request-validation-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` request-validation-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-validation-route spec: ingressClassName: apisix http: - name: request-validation-route match: paths: - /post upstreams: - name: httpbin-external-domain plugins: - name: request-validation enable: true config: header_schema: type: object required: - Content-Type properties: Content-Type: type: string pattern: "^application/json$" body_schema: type: object required: - required_payload properties: required_payload: type: string boolean_payload: type: boolean array_payload: type: array minItems: 1 items: type: integer minimum: 200 maximum: 599 uniqueItems: true default: - 200 ``` Apply the configuration to your cluster: ``` kubectl apply -f request-validation-ic.yaml ``` Send a request with JSON body that conforms to the schema to verify: ``` curl -i "http://127.0.0.1:9080/post" -X POST \ -H "Content-Type: application/json" \ -d '{"required_payload":"hello", "array_payload":[301]}' ``` You should receive an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "data": "{\"array_payload\":[301],\"required_payload\":\"hello\"}", "files": {}, "form": {}, "headers": { ... }, "json": { "array_payload": [ 301 ], "required_payload": "hello" }, "origin": "127.0.0.1, 183.17.233.107", "url": "http://127.0.0.1/post" } ``` If you send a request without specifying `Content-Type: application/json`: ``` curl -i "http://127.0.0.1:9080/post" -X POST \ -d '{"required_payload":"hello,world"}' ``` You should receive an `HTTP/1.1 400 Bad Request` response similar to the following: ``` property "Content-Type" validation failed: failed to match pattern "^application/json$" with "application/x-www-form-urlencoded" ``` Similarly, if you send a request without the required JSON field `required_payload`: ``` curl -i "http://127.0.0.1:9080/post" -X POST \ -H "Content-Type: application/json" \ -d '{}' ``` You should receive an `HTTP/1.1 400 Bad Request` response: ``` property "required_payload" is required ``` #### Validate URL-Encoded Form Body[​](#validate-url-encoded-form-body "Direct link to Validate URL-Encoded Form Body") Create a route with `request-validation` plugin as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "request-validation-route", "uri": "/post", "plugins": { "request-validation": { "header_schema": { "type": "object", "required": ["Content-Type"], "properties": { "Content-Type": { "type": "string", "pattern": "^application\/x-www-form-urlencoded$" } } }, "body_schema": { "type": "object", "required": ["required_payload","enum_payload"], "properties": { "required_payload": {"type": "string"}, "enum_payload": { "type": "string", "enum": ["enum_string_1", "enum_string_2"], "default": "enum_string_1" } } } } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: request-validation-service routes: - name: request-validation-route uris: - /post plugins: request-validation: header_schema: type: object required: - Content-Type properties: Content-Type: type: string pattern: "^application/x-www-form-urlencoded$" body_schema: type: object required: - required_payload - enum_payload properties: required_payload: type: string enum_payload: type: string enum: - enum_string_1 - enum_string_2 default: enum_string_1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD request-validation-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: request-validation-plugin-config spec: plugins: - name: request-validation config: header_schema: type: object required: - Content-Type properties: Content-Type: type: string pattern: "^application/x-www-form-urlencoded$" body_schema: type: object required: - required_payload - enum_payload properties: required_payload: type: string enum_payload: type: string enum: - enum_string_1 - enum_string_2 default: enum_string_1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: request-validation-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /post filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: request-validation-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` request-validation-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: request-validation-route spec: ingressClassName: apisix http: - name: request-validation-route match: paths: - /post upstreams: - name: httpbin-external-domain plugins: - name: request-validation enable: true config: header_schema: type: object required: - Content-Type properties: Content-Type: type: string pattern: "^application/x-www-form-urlencoded$" body_schema: type: object required: - required_payload - enum_payload properties: required_payload: type: string enum_payload: type: string enum: - enum_string_1 - enum_string_2 default: enum_string_1 ``` Apply the configuration to your cluster: ``` kubectl apply -f request-validation-ic.yaml ``` Send a request with URL-encoded form data to verify: ``` curl -i "http://127.0.0.1:9080/post" -X POST \ -H "Content-Type: application/x-www-form-urlencoded" \ -d "required_payload=hello&enum_payload=enum_string_1" ``` You should receive an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "data": "", "files": {}, "form": { "enum_payload": "enum_string_1", "required_payload": "hello" }, "headers": { ... }, "json": null, "origin": "127.0.0.1, 183.17.233.107", "url": "http://127.0.0.1/post" } ``` Send a request without the URL-encoded field `enum_payload`: ``` curl -i "http://127.0.0.1:9080/post" -X POST \ -H "Content-Type: application/x-www-form-urlencoded" \ -d "required_payload=hello" ``` You should receive an `HTTP/1.1 400 Bad Request` of the following: ``` property "enum_payload" is required ``` ## Appendix: JSON Schema[​](#appendix-json-schema "Direct link to Appendix: JSON Schema") The following section provides boilerplate JSON schema for you to adjust, combine, and use with this plugin. For a complete reference, see [JSON schema specification](https://json-schema.org/specification). ### Enumerated Values[​](#enumerated-values "Direct link to Enumerated Values") ``` { "body_schema": { "type": "object", "required": ["enum_payload"], "properties": { "enum_payload": { "type": "string", "enum": ["enum_string_1", "enum_string_2"], "default": "enum_string_1" } } } } ``` ### Boolean Values[​](#boolean-values "Direct link to Boolean Values") ``` { "body_schema": { "type": "object", "required": ["bool_payload"], "properties": { "bool_payload": { "type": "boolean", "default": true } } } } ``` ### Numeric Values[​](#numeric-values "Direct link to Numeric Values") ``` { "body_schema": { "type": "object", "required": ["integer_payload"], "properties": { "integer_payload": { "type": "integer", "minimum": 1, "maximum": 65535 } } } } ``` ### Strings[​](#strings "Direct link to Strings") ``` { "body_schema": { "type": "object", "required": ["string_payload"], "properties": { "string_payload": { "type": "string", "minLength": 1, "maxLength": 32 } } } } ``` ### RegEx for Strings[​](#regex-for-strings "Direct link to RegEx for Strings") ``` { "body_schema": { "type": "object", "required": ["regex_payload"], "properties": { "regex_payload": { "type": "string", "minLength": 1, "maxLength": 32, "pattern": "[[^[a-zA-Z0-9_]+$]]" } } } } ``` ### Arrays[​](#arrays "Direct link to Arrays") ``` { "body_schema": { "type": "object", "required": ["array_payload"], "properties": { "array_payload": { "type": "array", "minItems": 1, "items": { "type": "integer", "minimum": 200, "maximum": 599 }, "uniqueItems": true, "default": [200, 302] } } } } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * header\_schema object *** Schema for request header. At lease one of the `header_schema` and `body_schema` should be configured. * body\_schema object *** Schema for request body. At lease one of the `header_schema` and `body_schema` should be configured. * rejected\_code integer default: `400` vaild vaule: between 200 and 599 inclusive *** Status code to return when rejecting requests. * rejected\_msg string *** Message to return when rejecting requests. * max\_req\_body\_size integer default: `67108864` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes read when `body_schema` is configured. A larger body is rejected with `rejected_code`, which defaults to `400`. This field has no effect when only `header_schema` is configured. Introduced in API7 Enterprise 3.9.17 and 3.10.4, and APISIX 3.18.0. --- # response-rewrite The `response-rewrite` plugin offers options to rewrite responses that APISIX and its upstream services return to clients. With the plugin, you can modify HTTP status codes, response headers, response body, and more. For instance, you can use this plugin to: * Support [CORS](https://developer.mozilla.org/en-US/docs/Web/HTTP/CORS) by setting `Access-Control-Allow-*` headers. * Indicate redirection by setting HTTP status codes and `Location` header. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `response-rewrite` on a route in different scenarios. ### Rewrite Header and Body[​](#rewrite-header-and-body "Direct link to Rewrite Header and Body") The following example demonstrates how to add response body and headers, only to responses with `200` HTTP status codes. Create a route with the `response-rewrite` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "response-rewrite-route", "methods": ["GET"], "uri": "/headers", "plugins": { "response-rewrite": { "body": "{\"code\":\"ok\",\"message\":\"new json body\"}", "headers": { "set": { "X-Server-id": 3, "X-Server-status": "on", "X-Server-balancer-addr": "$balancer_ip:$balancer_port" } }, "vars": [ [ "status","==",200 ] ] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /headers name: response-rewrite-route methods: - GET plugins: response-rewrite: body: '{"code":"ok","message":"new json body"}' headers: set: X-Server-id: 3 X-Server-status: "on" X-Server-balancer-addr: "$balancer_ip:$balancer_port" vars: - - status - "==" - 200 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD response-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: response-rewrite-plugin-config spec: plugins: - name: response-rewrite config: body: '{"code":"ok","message":"new json body"}' headers: set: X-Server-id: 3 X-Server-status: "on" X-Server-balancer-addr: "$balancer_ip:$balancer_port" vars: - - status - "==" - 200 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: response-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: response-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f response-rewrite-ic.yaml ``` response-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: response-rewrite-route spec: ingressClassName: apisix http: - name: response-rewrite-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: response-rewrite enable: true config: body: '{"code":"ok","message":"new json body"}' headers: set: X-Server-id: 3 X-Server-status: "on" X-Server-balancer-addr: "$balancer_ip:$balancer_port" vars: - - status - "==" - 200 ``` Apply the configuration to your cluster: ``` kubectl apply -f response-rewrite-ic.yaml ``` Send a request to verify: ``` curl -i "http://127.0.0.1:9080/headers" ``` You should receive a `HTTP/1.1 200 OK` response similar to the following: ``` ... X-Server-id: 3 X-Server-status: on X-Server-balancer-addr: 50.237.103.220:80 {"code":"ok","message":"new json body"} ``` ### Rewrite Header With RegEx Filter[​](#rewrite-header-with-regex-filter "Direct link to Rewrite Header With RegEx Filter") The following example demonstrates how to use RegEx filter matching to replace `X-Amzn-Trace-Id` for responses. Create a route with the `response-rewrite` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "response-rewrite-route", "methods": ["GET"], "uri": "/headers", "plugins":{ "response-rewrite":{ "filters":[ { "regex":"X-Amzn-Trace-Id", "scope":"global", "replace":"X-Amzn-Trace-Id-Replace" } ] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /headers name: response-rewrite-route methods: - GET plugins: response-rewrite: filters: - regex: X-Amzn-Trace-Id scope: global replace: X-Amzn-Trace-Id-Replace upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD response-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: response-rewrite-plugin-config spec: plugins: - name: response-rewrite config: filters: - regex: X-Amzn-Trace-Id scope: global replace: X-Amzn-Trace-Id-Replace --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: response-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: response-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f response-rewrite-ic.yaml ``` response-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: response-rewrite-route spec: ingressClassName: apisix http: - name: response-rewrite-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: response-rewrite enable: true config: filters: - regex: X-Amzn-Trace-Id scope: global replace: X-Amzn-Trace-Id-Replace ``` Apply the configuration to your cluster: ``` kubectl apply -f response-rewrite-ic.yaml ``` Send a request to verify: ``` curl -i "http://127.0.0.1:9080/headers" ``` You should see a response similar to the following: ``` { "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/8.2.1", "X-Amzn-Trace-Id-Replace": "Root=1-6500095d-1041b05e2ba9c6b37232dbc7", "X-Forwarded-Host": "127.0.0.1" } } ``` ### Decode Body from Base64[​](#decode-body-from-base64 "Direct link to Decode Body from Base64") The following example demonstrates how to Decode Body from Base64 format. Create a route with the `response-rewrite` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "response-rewrite-route", "methods": ["GET"], "uri": "/get", "plugins":{ "response-rewrite": { "body": "SGVsbG8gV29ybGQ=", "body_base64": true } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /get name: response-rewrite-route methods: - GET plugins: response-rewrite: body: SGVsbG8gV29ybGQ= body_base64: true upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD response-rewrite-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: response-rewrite-plugin-config spec: plugins: - name: response-rewrite config: body: SGVsbG8gV29ybGQ= body_base64: true --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: response-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: response-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f response-rewrite-ic.yaml ``` response-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: response-rewrite-route spec: ingressClassName: apisix http: - name: response-rewrite-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: response-rewrite enable: true config: body: SGVsbG8gV29ybGQ= body_base64: true ``` Apply the configuration to your cluster: ``` kubectl apply -f response-rewrite-ic.yaml ``` Send a request to verify: ``` curl "http://127.0.0.1:9080/get" ``` You should see a response of the following: ``` Hello World ``` ### Rewrite Response and Its Connection with Execution Phases[​](#rewrite-response-and-its-connection-with-execution-phases "Direct link to Rewrite Response and Its Connection with Execution Phases") The following example demonstrates the connection between the `response-rewrite` plugin and [execution phases](https://docs.api7.ai/apisix/key-concepts/plugins.md#plugins-execution-lifecycle) by configuring the plugin with the `key-auth` plugin, and see how the response is still rewritten to `200 OK` in the case of an unauthenticated request. * Admin API * ADC * Ingress Controller Create a consumer `jack`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jack" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jack/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jack-key-auth", "plugins": { "key-auth": { "key": "jack-key" } } }' ``` Create a route with `key-auth` and configure `response-rewrite` to rewrite the response status code and body: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "response-rewrite-route", "uri": "/get", "plugins": { "key-auth": {}, "response-rewrite": { "status_code": 200, "body": "{\"code\": 200, \"msg\": \"success\"}" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a consumer with `key-auth` credential and a route with `key-auth` and `response-rewrite` plugins configured as such: adc.yaml ``` consumers: - username: jack credentials: - name: cred-jack-key-auth type: key-auth config: key: jack-key services: - name: httpbin routes: - name: response-rewrite-route uris: - /get plugins: key-auth: {} response-rewrite: status_code: 200 body: '{"code": 200, "msg": "success"}' upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD response-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jack spec: gatewayRef: name: apisix credentials: - type: key-auth name: cred-jack-key-auth config: key: jack-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: response-rewrite-plugin-config spec: plugins: - name: key-auth config: _meta: disable: false - name: response-rewrite config: status_code: 200 body: '{"code": 200, "msg": "success"}' --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: response-rewrite-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: response-rewrite-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f response-rewrite-ic.yaml ``` response-rewrite-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jack spec: ingressClassName: apisix authParameter: keyAuth: value: key: jack-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: response-rewrite-route spec: ingressClassName: apisix http: - name: response-rewrite-route match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: key-auth enable: true - name: response-rewrite enable: true config: status_code: 200 body: '{"code": 200, "msg": "success"}' ``` Apply the configuration to your cluster: ``` kubectl apply -f response-rewrite-ic.yaml ``` Send a request to the route with the valid key: ``` curl -i "http://127.0.0.1:9080/get" -H 'apikey: jack-key' ``` You should receive an `HTTP/1.1 200 OK` response of the following: ``` {"code": 200, "msg": "success"} ``` Send a request to the route without any key: ``` curl -i "http://127.0.0.1:9080/get" ``` You should still receive an `HTTP/1.1 200 OK` response of the same, instead of `HTTP/1.1 401 Unauthorized` from the `key-auth` plugin. This shows that the `response-rewrite` plugin still rewrites the response. This is because **header\_filter** and **body\_filter** phase logics of the `response-rewrite` plugin will continue to run after [`ngx.exit`](https://openresty-reference.readthedocs.io/en/latest/Lua_Nginx_API/#ngxexit) in the **access** or **rewrite** phases from other plugins. The following table summarizes the impact of `ngx.exit` on execution phases. | Phase | rewrite | access | header\_filter | body\_filter | | ------------------ | -------- | -------- | -------------- | ------------ | | **rewrite** | ngx.exit | | | | | **access** | × | ngx.exit | | | | **header\_filter** | ✓ | ✓ | ngx.exit | | | **body\_filter** | ✓ | ✓ | × | ngx.exit | For example, if `ngx.exit` takes places in the **rewrite** phase, it will interrupt the execution of **access** phase but not interfere with **header\_filter** and **body\_filter** phases. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * status\_code integer vaild vaule: between 200 and 598 inclusive *** New HTTP status code in the response. * body string *** New response body. The `Content-Length` header would also be reset. Should not be configured with `filters`. * body\_base64 boolean default: `false` *** If true, decode the response body configured in `body` before sending to client, which is useful for image and protobuf decoding. Note that this configuration cannot be used to decode upstream response. * headers object *** Actions to be executed in the order of `add`, `remove`, and `set`. * add array\[string] *** Headers to append to responses. If a header is already present in the response, the header value will be appended. Header value could be set to a constant, or one or more [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). * set object *** Headers to set in responses. If a header is already present in the response, the header value will be overwritten. Header value could be set to a constant, or one or more [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). * remove array\[string] *** Headers to remove from responses. * vars array\[array] *** An array of one or more matching conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md) to conditionally execute the plugin. * filters array\[object] *** List of filters that modify the response body by replacing one specified string with another. Should not be configured with `body`. * regex string required *** RegEx pattern to match on the response body. * scope string default: `once` vaild vaule: `once` or `global` *** Scope of substitution. `once` substitutes the first matched instance and `global` substitutes globally. * replace string required *** Content to substitute with. * options string default: `jo` *** RegEx options to control how the match operation should be performed. See [Lua NGINX module](https://github.com/openresty/lua-nginx-module#ngxrematch) for the available options. * max\_resp\_body\_size integer default: `67108864` vaild vaule: greater than or equal to 1 *** Maximum response body size in bytes buffered when `filters` are configured. A larger response is truncated to this size before the filters run. The chunk that crosses the threshold is buffered before the limit is enforced, so transient memory use can exceed the configured size. This field has no effect when `filters` is not configured. Introduced in API7 Enterprise 3.9.17 and 3.10.4, and APISIX 3.18.0. --- # rocketmq-logger The `rocketmq-logger` plugin pushes request and response logs as JSON objects to your RocketMQ clusters in batches and supports the customization of log formats. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `rocketmq-logger` plugin for different scenarios. To follow along the examples, start a sample RocketMQ cluster: * Docker * Kubernetes docker-compose.yml ``` version: "3" services: rocketmq_namesrv: image: apacherocketmq/rocketmq:4.6.0 container_name: rmqnamesrv restart: unless-stopped ports: - "9876:9876" command: sh mqnamesrv networks: rocketmq_net: rocketmq_broker: image: apacherocketmq/rocketmq:4.6.0 container_name: rmqbroker restart: unless-stopped ports: - "10909:10909" - "10911:10911" - "10912:10912" depends_on: - rocketmq_namesrv command: sh mqbroker -n rmqnamesrv:9876 -c ../conf/broker.conf networks: rocketmq_net: networks: rocketmq_net: ``` Start containers: ``` docker compose up -d ``` In a few seconds, the name server and broker should start. Create the `TopicTest` topic: ``` docker exec -i rmqnamesrv rm /home/rocketmq/rocketmq-4.6.0/conf/tools.yml docker exec -i rmqnamesrv /home/rocketmq/rocketmq-4.6.0/bin/mqadmin updateTopic -n rmqnamesrv:9876 -t TopicTest -c DefaultCluster ``` Create a Kubernetes manifest file for the RocketMQ name server and broker deployments: rocketmq-deployment.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: rocketmq-namesrv spec: replicas: 1 selector: matchLabels: app: rocketmq-namesrv template: metadata: labels: app: rocketmq-namesrv spec: containers: - name: rocketmq-namesrv image: apacherocketmq/rocketmq:4.6.0 command: ["sh", "mqnamesrv"] ports: - containerPort: 9876 --- apiVersion: v1 kind: Service metadata: namespace: aic name: rocketmq-namesrv spec: selector: app: rocketmq-namesrv ports: - port: 9876 targetPort: 9876 type: ClusterIP --- apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: rocketmq-broker spec: replicas: 1 selector: matchLabels: app: rocketmq-broker template: metadata: labels: app: rocketmq-broker spec: containers: - name: rocketmq-broker image: apacherocketmq/rocketmq:4.6.0 command: ["sh", "mqbroker", "-n", "rocketmq-namesrv:9876", "-c", "../conf/broker.conf"] ports: - containerPort: 10909 - containerPort: 10911 - containerPort: 10912 --- apiVersion: v1 kind: Service metadata: namespace: aic name: rocketmq-broker spec: selector: app: rocketmq-broker ports: - name: fastlisten port: 10909 targetPort: 10909 - name: listen port: 10911 targetPort: 10911 - name: haservice port: 10912 targetPort: 10912 type: ClusterIP ``` Apply the manifests: ``` kubectl apply -f rocketmq-deployment.yaml ``` Once the pods are running, create the `TopicTest` topic: ``` kubectl exec -n aic deploy/rocketmq-namesrv -- sh -c \ "rm -f /home/rocketmq/rocketmq-4.6.0/conf/tools.yml && \ /home/rocketmq/rocketmq-4.6.0/bin/mqadmin updateTopic \ -n rocketmq-namesrv:9876 -t TopicTest -c DefaultCluster" ``` Wait for messages in the configured RocketMQ topic: * Docker * Kubernetes ``` docker run -it --name rockemq_consumer -e NAMESRV_ADDR=localhost:9876 --net host apacherocketmq/rocketmq:4.6.0 sh tools.sh org.apache.rocketmq.example.quickstart.Consumer ``` In a few seconds, the consumer should start and listen for messages from APISIX: ``` 01:32:17.823 [main] DEBUG i.n.u.i.l.InternalLoggerFactory - Using SLF4J as the default logging framework Consumer Started. ``` Open a new terminal session for the following steps working with APISIX. After sending requests to APISIX in the following examples, run the following command to print messages from the `TopicTest` topic: ``` kubectl exec -n aic deploy/rocketmq-namesrv -- sh -c \ "/home/rocketmq/rocketmq-4.6.0/bin/mqadmin printMsg \ -n rocketmq-namesrv:9876 -t TopicTest" ``` ### Log in Different Meta Log Formats[​](#log-in-different-meta-log-formats "Direct link to Log in Different Meta Log Formats") The following example demonstrates how you can enable the `rocketmq-logger` plugin on a route, which logs client requests to the route and pushes logs to RocketMQ. You will also understand the differences between the `default` and `origin` meta log formats. Create a route with `rocketmq-logger` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "rocketmq-logger-route", "uri": "/anything", "plugins": { "rocketmq-logger": { "nameserver_list": [ "127.0.0.1:9876" ], "topic": "TopicTest", "key": "key1", "timeout": 30, "meta_format": "default", "batch_max_size": 1 } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: rocketmq-logger-route plugins: rocketmq-logger: nameserver_list: - "127.0.0.1:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "default" batch_max_size: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD rocketmq-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: rocketmq-logger-plugin-config spec: plugins: - name: rocketmq-logger config: nameserver_list: - "rocketmq-namesrv.aic.svc:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "default" batch_max_size: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: rocketmq-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: rocketmq-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` rocketmq-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: rocketmq-logger-route spec: ingressClassName: apisix http: - name: rocketmq-logger-route match: paths: - /anything* upstreams: - name: httpbin-external-domain plugins: - name: rocketmq-logger enable: true config: nameserver_list: - "rocketmq-namesrv.aic.svc:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "default" batch_max_size: 1 ``` Apply the configuration: ``` kubectl apply -f rocketmq-logger-ic.yaml ``` ❶ `meta_format`: set to the `default` log format. ❷ `batch_max_size`: set to 1 to send the log entry immediately. Send a request to the route to generate a log entry: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should see a log entry similar to the following: ``` { "client_ip": "127.0.0.1", "upstream": "34.197.122.172:80", "start_time": 1744727400000, "request": { "headers": { "host": "127.0.0.1:9080", "accept": "*/*", "user-agent": "curl/8.6.0" }, "querystring": {}, "size": 86, "uri": "/anything", "url": "http://127.0.0.1:9080/anything", "method": "GET" }, "route_id": "rocketmq-logger-route", "apisix_latency": 8.9998455047607, "upstream_latency": 503, "latency": 511.99984550476, "response": { "size": 617, "headers": { "content-length": "391", "connection": "close", "date": "Tue, 15 Apr 2025 14:30:00 GMT", "server": "APISIX/3.15.0", "content-type": "application/json" }, "status": 200 }, "server": { "hostname": "apisix", "version": "3.15.0" }, "service_id": "" } ``` Update the `rocketmq-logger` meta log format to `origin`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes/rocketmq-logger-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "rocketmq-logger": { "meta_format": "origin" } } }' ``` Update `adc.yaml` to change `meta_format` to `origin`: adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: rocketmq-logger-route plugins: rocketmq-logger: nameserver_list: - "127.0.0.1:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "origin" batch_max_size: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update `rocketmq-logger-ic.yaml` to change `meta_format` to `origin` in the `PluginConfig`: rocketmq-logger-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: rocketmq-logger-plugin-config spec: plugins: - name: rocketmq-logger config: nameserver_list: - "rocketmq-namesrv.aic.svc:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "origin" batch_max_size: 1 ``` Apply the updated configuration: ``` kubectl apply -f rocketmq-logger-ic.yaml ``` Update `rocketmq-logger-ic.yaml` to change `meta_format` to `origin` in the `ApisixRoute`: rocketmq-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: rocketmq-logger-route spec: ingressClassName: apisix http: - name: rocketmq-logger-route match: paths: - /anything* upstreams: - name: httpbin-external-domain plugins: - name: rocketmq-logger enable: true config: nameserver_list: - "rocketmq-namesrv.aic.svc:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "origin" batch_max_size: 1 ``` Apply the updated configuration: ``` kubectl apply -f rocketmq-logger-ic.yaml ``` Send a request to the route again to generate a new log entry: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should see a log entry in the raw HTTP request format: ``` GET /anything HTTP/1.1 host: 127.0.0.1:9080 user-agent: curl/8.6.0 accept: */* ``` ### Log Request and Response Headers With Plugin Metadata[​](#log-request-and-response-headers-with-plugin-metadata "Direct link to Log Request and Response Headers With Plugin Metadata") The following example demonstrates how you can customize log format using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) and [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) to log specific headers from request and response. In APISIX, [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) is used to configure the common metadata fields of all plugin instances of the same plugin. It is useful when a plugin is enabled across multiple resources and requires a universal update to their metadata fields. First, create a route with `rocketmq-logger` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "rocketmq-logger-route", "uri": "/anything", "plugins": { "rocketmq-logger": { "nameserver_list": [ "127.0.0.1:9876" ], "topic": "TopicTest", "key": "key1", "timeout": 30, "meta_format": "default", "batch_max_size": 1 } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: rocketmq-logger-route plugins: rocketmq-logger: nameserver_list: - "127.0.0.1:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "default" batch_max_size: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD rocketmq-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: rocketmq-logger-plugin-config spec: plugins: - name: rocketmq-logger config: nameserver_list: - "rocketmq-namesrv.aic.svc:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "default" batch_max_size: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: rocketmq-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: rocketmq-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` rocketmq-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: rocketmq-logger-route spec: ingressClassName: apisix http: - name: rocketmq-logger-route match: paths: - /anything* upstreams: - name: httpbin-external-domain plugins: - name: rocketmq-logger enable: true config: nameserver_list: - "rocketmq-namesrv.aic.svc:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "default" batch_max_size: 1 ``` Apply the configuration: ``` kubectl apply -f rocketmq-logger-ic.yaml ``` ❶ `meta_format`: set to the `default` log format. It is important to note that this is mandatory if you would like to customize log format with plugin metadata. If `meta_format` is set to `origin`, the log entries will remain in `origin` format. ❷ `batch_max_size`: set to 1 to send the log entry immediately. Next, configure the plugin metadata for `rocketmq-logger`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/rocketmq-logger" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "log_format": { "host": "$host", "@timestamp": "$time_iso8601", "client_ip": "$remote_addr", "env": "$http_env", "resp_content_type": "$sent_http_Content_Type" } }' ``` adc.yaml ``` plugin_metadata: - name: rocketmq-logger log_format: host: "$host" "@timestamp": "$time_iso8601" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: rocketmq-logger: log_format: host: "$host" "@timestamp": "$time_iso8601" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` Apply the configuration: ``` kubectl apply -f gatewayproxy.yaml ``` ❶ log the custom request header `env`. ❷ log the response header `Content-Type`. Send a request to the route with the `env` header: ``` curl -i "http://127.0.0.1:9080/anything" -H "env: dev" ``` You should see a log entry similar to the following: ``` { "host": "127.0.0.1", "client_ip": "127.0.0.1", "resp_content_type": "application/json", "route_id": "rocketmq-logger-route", "env": "dev", "@timestamp": "2025-04-15T14:30:00+00:00" } ``` ### Log Request Bodies Conditionally[​](#log-request-bodies-conditionally "Direct link to Log Request Bodies Conditionally") The following example demonstrates how you can conditionally log request body. Create a route with `rocketmq-logger` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "rocketmq-logger": { "nameserver_list": [ "127.0.0.1:9876" ], "topic": "TopicTest", "key": "key1", "timeout": 30, "meta_format": "default", "batch_max_size": 1, "include_req_body": true, "include_req_body_expr": [["arg_log_body", "==", "yes"]] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" }, "uri": "/anything", "id": "rocketmq-logger-route" }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: rocketmq-logger-route plugins: rocketmq-logger: nameserver_list: - "127.0.0.1:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "default" batch_max_size: 1 include_req_body: true include_req_body_expr: - - "arg_log_body" - "==" - "yes" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD rocketmq-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: rocketmq-logger-plugin-config spec: plugins: - name: rocketmq-logger config: nameserver_list: - "rocketmq-namesrv.aic.svc:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "default" batch_max_size: 1 include_req_body: true include_req_body_expr: - - "arg_log_body" - "==" - "yes" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: rocketmq-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: rocketmq-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` rocketmq-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: rocketmq-logger-route spec: ingressClassName: apisix http: - name: rocketmq-logger-route match: paths: - /anything* upstreams: - name: httpbin-external-domain plugins: - name: rocketmq-logger enable: true config: nameserver_list: - "rocketmq-namesrv.aic.svc:9876" topic: "TopicTest" key: "key1" timeout: 30 meta_format: "default" batch_max_size: 1 include_req_body: true include_req_body_expr: - - "arg_log_body" - "==" - "yes" ``` Apply the configuration: ``` kubectl apply -f rocketmq-logger-ic.yaml ``` ❶ `include_req_body`: set to true to include request body. ❷ `include_req_body_expr`: only include request body if the URL query string `log_body` is `yes`. Send a request to the route with an URL query string satisfying the condition: ``` curl -i "http://127.0.0.1:9080/anything?log_body=yes" -X POST -d '{"env": "dev"}' ``` You should see the request body logged: ``` { ..., "method": "POST", "body": "{\"env\": \"dev\"}", "size": 183 } } ``` Send a request to the route without any URL query string: ``` curl -i "http://127.0.0.1:9080/anything" -X POST -d '{"env": "dev"}' ``` You should not observe the request body in the log. info If you have customized the `log_format` in addition to setting `include_req_body` or `include_resp_body` to `true`, the plugin would not include the bodies in the logs. As a workaround, you may be able to use the NGINX variable `$request_body` in the log format, such as: ``` { "rocketmq-logger": { ..., "log_format": {"body": "$request_body"} } } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * nameserver\_list array\[string] required *** List of RocketMQ nameservers. * topic string required *** Target topic to push the data to. * key string *** Key of the message. * tag string *** Tag of the message. * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. You can also configure log format on a global scale using the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md), which configures the log format for all `rocketmq-logger` plugin instances. If the log format configured on the individual plugin instance differs from the log format configured on plugin metadata, the log format configured on the individual plugin instance takes precedence. See the [example](https://docs.api7.ai/hub/rocketmq-logger.md#log-request-and-response-headers-with-plugin-metadata) for more details. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * timeout integer default: `3` *** Timeout for the upstream to send data. * use\_tls boolean default: `false` *** If true, verify SSL. * access\_key string *** Access key for ACL. Setting to an empty string will disable the ACL. * secret\_key string *** Secret key for ACL. The value is encrypted with AES before being stored in etcd. * name string default: `rocketmq logger` *** Unique identifier of the plugin for the batch processor. If you use [Prometheus](https://docs.api7.ai/hub/prometheus.md) to monitor APISIX metrics, the name is exported in `apisix_batch_process_entries`. * meta\_format string default: `default` vaild vaule: `default` or `origin` *** Format to collect the request information. Setting to `default` collects the information in JSON format and `origin` collects the information with the original HTTP request. See the [example](https://docs.api7.ai/hub/kafka-logger.md#log-in-different-meta-log-formats) for more details. * include\_req\_body boolean default: `false` *** If true, include the request body in the log. Note that if the request body is too big to be kept in the memory, it can not be logged due to NGINX's limitations. * include\_req\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_req_body` is true. Request body would only be logged when the expressions configured here evaluate to true. * include\_resp\_body boolean default: `false` *** If true, include the response body in the log. * include\_resp\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_resp_body` is true. Response body would only be logged when the expressions configured here evaluate to true. * max\_req\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes to include in the log. If the request body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * max\_resp\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum response body size in bytes to include in the log. If the response body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * batch\_max\_size integer default: `1000` vaild vaule: greater than 0 *** The number of log entries allowed in one batch. Once reached, the batch will be sent to the logging service. Setting this parameter to 1 means immediate processing. * inactive\_timeout integer default: `5` vaild vaule: greater than 0 *** The maximum time in seconds to wait for new logs before sending the batch to the logging service. The value should be smaller than `buffer_duration`. * buffer\_duration integer default: `60` vaild vaule: greater than 0 *** The maximum time in seconds from the earliest entry allowed before sending the batch to the logging service. * retry\_delay integer default: `1` vaild vaule: greater than or equal to 0 *** The time interval in seconds to retry sending the batch to the logging service if the batch was not successfully sent. * max\_retry\_count integer default: `0` vaild vaule: greater than or equal to 0 *** The maximum number of unsuccessful retries allowed before dropping the log entries. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * max\_pending\_entries integer default: `` `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 and 3.10.6; none in API7 Enterprise 3.9.18 and 3.10.5 `` vaild vaule: greater than or equal to 1 *** Maximum number of entries waiting in the batch processor. New entries are discarded when the backlog reaches the limit. Introduced in API7 Enterprise 3.8.17 and APISIX 3.15.0. The default changed to `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line. In API7 Enterprise 3.9.18 and 3.10.5, and in earlier APISIX versions, omitting the parameter leaves the backlog unlimited. See [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md#configure-the-pending-entry-limit) for sizing and verification guidance. --- # saml-auth The `saml-auth` plugin enables APISIX or API7 Gateway to act as a service provider (SP) and authenticate users through a [SAML 2.0](https://en.wikipedia.org/wiki/SAML_2.0) identity provider (IdP). ## Example[​](#example "Direct link to Example") ### Integrate with Keycloak[​](#integrate-with-keycloak "Direct link to Integrate with Keycloak") The following example assumes that you have a locally accessible gateway and demonstrates how to set up SAML single sign-on (SSO) with Keycloak. #### Start a Keycloak Server[​](#start-a-keycloak-server "Direct link to Start a Keycloak Server") Start a Keycloak instance with admin username `admin` and admin password `admin-pass`: * Docker * Kubernetes ``` docker run -d --name keycloak \ -e 'KEYCLOAK_ADMIN=admin' \ -e 'KEYCLOAK_ADMIN_PASSWORD=admin-pass' \ -p 8080:8080 \ quay.io/keycloak/keycloak:25.0.4 start-dev ``` Once started, visit [`http://localhost:8080`](http://localhost:8080) in your browser to access the Keycloak admin console. Log in with the admin username and password. Create a Kubernetes manifest file for the Keycloak deployment and service: keycloak.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: keycloak spec: replicas: 1 selector: matchLabels: app: keycloak template: metadata: labels: app: keycloak spec: containers: - name: keycloak image: quay.io/keycloak/keycloak:25.0.4 args: - start-dev env: - name: KEYCLOAK_ADMIN value: admin - name: KEYCLOAK_ADMIN_PASSWORD value: admin-pass - name: KC_HTTP_PORT value: "8080" ports: - containerPort: 8080 --- apiVersion: v1 kind: Service metadata: namespace: aic name: keycloak spec: selector: app: keycloak ports: - port: 8080 targetPort: 8080 type: ClusterIP ``` Apply the manifest: ``` kubectl apply -f keycloak.yaml ``` Wait for the pod to be ready. Once ready, forward the Keycloak port to your local machine so you can access the admin console and so that the browser can reach the IdP during the SAML flow: ``` kubectl port-forward -n aic service/keycloak 8080:8080 & ``` Once the port-forward is active, visit [`http://localhost:8080`](http://localhost:8080) in your browser to access the Keycloak admin console. Log in with the admin username and password. #### Create a Client[​](#create-a-client "Direct link to Create a Client") Create a new client in Keycloak and configure the following: * Configure the **Client type** as `SAML`. * Enter the service provider's (SP) name as the **Client ID**, for example, `api7`. * Note that this value should be consistent with the `sp_issuer` parameter value, which you will configure later in the plugin. * Add `http://127.0.0.1:9080/anything/login_callback` to **Valid redirect URIs**. * Add `http://127.0.0.1:9080/anything/logout_callback` to **Valid post logout redirect URIs**. * Set **Force POST binding** option to `Off`. * Make sure the **Sign documents** option is `On`. Otherwise, `SigAlg` and `Signature` will be missing from the SAML response. #### Find the Realm's SAML Metadata[​](#find-the-realms-saml-metadata "Direct link to Find the Realm's SAML Metadata") In this step, you will find the `idp_cert` and `idp_uri` from the realm's SAML metadata file. Select **Realm Settings** and under the **General** tab, you should find **SAML 2.0 Identity Provider Metadata** in **Endpoints**. The metadata file should look similar to the following: ``` wDDsXcgLGAZwZgpSb_jlBRf5MF8FoTcOYs0DgZ30Xcc MIICmzCCAYMCBgGRl7njKjANBgkqhkiG9w0BAQsFADARMQ8wDQYDVQQDDAZtYXN0ZXIwHhcNMjQwODI4MDY0MjA3WhcNMzQwODI4MDY0MzQ3WjARMQ8wDQYDVQQDDAZtYXN0ZXIwggEiMA0GCSqGSIb3DQEBAQUAA4IBDwAwggEKAoIBAQCdYPYSFoX2MADSIgfLYQ5oZcLNE+qB+qsO8sNpiebMQE3RmI5+MmZC/aozRzkzxcY+AoM50qfHrM1yM99A9ZxZt6fW/MuIv6IP5zWLDl0XWGVeOH0HIH4/xBxQetBxm1HdOYpCQg5Wm9hmYfebmN7NfW8HjnORjfUuUGgs5eCiHVqfiCfphLF5w+DcIcnjIwyF+xVH/7fRWgo5inBSeIavZh/LEv7LzBeRleGgoZ/+q7cVQiL2e0b8rsslqUOZJmwdPU3VSS0vW1bmXsZsfaZD0bgakFvSj0ARzwIbxc74eEQYKflHGS0zkrpm+TsO5KUn59SCPOhGNgGYpKKv6cY1AgMBAAEwDQYJKoZIhvcNAQELBQADggEBABN21PoEiTaZ20qQUdKD03m+bySlF4jRX2AeZqCedBaW+nHrbefaJdEnE9AcXBENCWVr6ntdeREaL9dW6KpV1hT4BmnXO2aiFotZe4Vc2W6cv7nDpjil6Q5/isbT5sriYhcU9oXBAaLf9dlg7K/X1l1+zcy9Pd1uKUfrC+5ds/Zv+xHiiK4h55o8shcmBmQ7bsanzNmjIQNnyF+lNRciGRvgJp59TR7AWpiBQDTNW1KK3XjO9lmN8nCEPbpdNGi77TDX0OZVrbbPy3vL4n8Gi3oQptHhmV7xou4fTEn9TCrdW82OLOduBCMk9t0tFFNB8Hlxq5XsLVLYW7O9GGcjDmI= urn:oasis:names:tc:SAML:2.0:nameid-format:persistent urn:oasis:names:tc:SAML:2.0:nameid-format:transient urn:oasis:names:tc:SAML:1.1:nameid-format:unspecified urn:oasis:names:tc:SAML:1.1:nameid-format:emailAddress ``` Note down the certificate content in `X509Certificate` for plugin configuration `idp_cert`. The `SingleSignOnService.Location` URL will also be used in the `idp_uri` plugin configuration later, where the `localhost` should be replace with your private IP address. For example, the `idp_uri` would look similar to `http://192.168.2.101:8080/realms/master/protocol/saml`. #### Create Service Provider (SP) Certificate and Key[​](#create-service-provider-sp-certificate-and-key "Direct link to Create Service Provider (SP) Certificate and Key") There are two approaches to create the service provider's certificate and key, and configure the certificate in Keycloak: 1. Generate the certificate and private key locally using openssl, and import the certificate to Keycloak client; or 2. Generate the certificate and private key in Keycloak, which automatically configures the certificate in Keycloak. You will save the certificate and private key for plugin configuration later in API7. With the first approach, generate the certificate and private key using the openssl utility: ``` # Generate Private Key openssl genrsa -out sp_private_key.pem 2048 # Generate Certificate Signing Request (CSR) openssl req -new -key sp_private_key.pem -out sp_csr.pem -subj "/CN=API7" # Generate Self-Signed Certificate openssl x509 -req -days 365 -in sp_csr.pem -signkey sp_private_key.pem -out sp_cert.pem ``` In Keycloak, go to the client, and under **Keys** tab where you see the **Client signature required** option is set on **On**, you should see an option to **Import Key** for the **Certificate**. Choose **Certificate PEM** as the **Archive format** and import `sp_cert.pem`. Alternatively if you wish to use the second approach, to generate the certificate and key in Keycloak, you can click **Regenerate** under **Certificate**. This will update the certificate configured in the client and download the private key to your host. #### Create a Route with `saml-auth` Plugin[​](#create-a-route-with-saml-auth-plugin "Direct link to create-a-route-with-saml-auth-plugin") tip Replace the `secret`, `idp_uri` IP address, `idp_cert`, `sp_cert`, and `sp_private_key` with your own values. * Admin API * ADC * Ingress Controller Create a route with the `saml-auth` plugin: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "saml-auth-route", "uri": "/anything/*", "plugins": { "saml-auth": { "secret": "my_secret_key", "sp_issuer": "api7", "idp_uri": "http://192.168.2.101:8080/realms/master/protocol/saml", "login_callback_uri": "/anything/login_callback", "logout_callback_uri": "/anything/logout_callback", "logout_uri": "/anything/logout", "logout_redirect_uri": "/anything/logout_ok", "idp_cert": "-----BEGIN CERTIFICATE-----\nMIICmzCCAYMCBgGRl7njKjANBgkqhkiG9w0BAQsFADARMQ8wDQYDVQQDDAZtYXN0\nZXIwHhcNMjQwODI4MDY0MjA3WhcNMzQwODI4MDY0MzQ3WjARMQ8wDQYDVQQDDAZt\nYXN0ZXIwggEiMA0GCSqGSIb3DQEBAQUAA4IBDwAwggEKAoIBAQCdYPYSFoX2MADS\nIgfLYQ5oZcLNE+qB+qsO8sNpiebMQE3RmI5+MmZC/aozRzkzxcY+AoM50qfHrM1y\nM99A9ZxZt6fW/MuIv6IP5zWLDl0XWGVeOH0HIH4/xBxQetBxm1HdOYpCQg5Wm9hm\nYfebmN7NfW8HjnORjfUuUGgs5eCiHVqfiCfphLF5w+DcIcnjIwyF+xVH/7fRWgo5\ninBSeIavZh/LEv7LzBeRleGgoZ/+q7cVQiL2e0b8rsslqUOZJmwdPU3VSS0vW1bm\nXsZsfaZD0bgakFvSj0ARzwIbxc74eEQYKflHGS0zkrpm+TsO5KUn59SCPOhGNgGY\npKKv6cY1AgMBAAEwDQYJKoZIhvcNAQELBQADggEBABN21PoEiTaZ20qQUdKD03m+\nbySlF4jRX2AeZqCedBaW+nHrbefaJdEnE9AcXBENCWVr6ntdeREaL9dW6KpV1hT4\nBmnXO2aiFotZe4Vc2W6cv7nDpjil6Q5/isbT5sriYhcU9oXBAaLf9dlg7K/X1l1+\nzcy9Pd1uKUfrC+5ds/Zv+xHiiK4h55o8shcmBmQ7bsanzNmjIQNnyF+lNRciGRvg\nJp59TR7AWpiBQDTNW1KK3XjO9lmN8nCEPbpdNGi77TDX0OZVrbbPy3vL4n8Gi3oQ\nptHhmV7xou4fTEn9TCrdW82OLOduBCMk9t0tFFNB8Hlxq5XsLVLYW7O9GGcjDmI=\n-----END CERTIFICATE-----", "sp_cert": "-----BEGIN CERTIFICATE-----\nMIIC0TCCAbmgAwIBAgIUAT7h3zLAul/3S1F9Ms9w7JjpoJ0wDQYJKoZIhvcNAQEL\nBQAwETEPMA0GA1UEAwwGQVBJU0lYMB4XDTI0MDgyNzA5MDk1NloXDTI1MDgyNzA5\nMDk1NlowETEPMA0GA1UEAwwGQVBJU0lYMIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8A\nMIIBCgKCAQEAvJBsuNfgvxe+xrBPr9+OCwD4dk3M9ua+14l9tlQHFgtqGXEq7nYc\n0ic9wqim+kdxpJWfiwG0mClklO0nELNsgBVrC06FqrcSe2CGEh91UBkGEOzvOgm7\nEBJOB/5Nc4tE/3NXM0ocfRgFXNEvGMkH9M+odGk7ZQraI/hfazwYgjOty1LrvSMp\nKhCfx0DpKlLuX0w2P9CfLuSgZ0ZTdN3Yr4icEuEs0i3ptCd/bip2fccKkRWEguIe\nywoDl/2fjubJFc5sFhl7Rtf+CeFKgqeByNPX2+UCix136L1r+VIlA+3ClInPWZUY\nWCbs/envBO6omUsnqPPCU2zVdYW0Qb+rrQIDAQABoyEwHzAdBgNVHQ4EFgQUvGIj\nuvPoHC74lhKSlOJAwrdq4WwwDQYJKoZIhvcNAQELBQADggEBAJU0+aKCUSYvN6oe\n7PHYD0ZvE13wItzKq/7DQQe1zA/kDoCvSyC8+gB+FZmdHmkGGNdNqXsQgHEnP7Y0\nx7gDqA3s0blXEkECfmmRcVxcS3rb8CVVFqiKdyRO91opdir5J9vbmiF7RK1ajFTy\nyemhK0xxFpPM+gTdetEj7AoVMrlRoOLC+L62GaSi/gpQmKPR91FLyj33vCfVrDCo\nQXYMPQmSbBCwlHrHWa/Px7F7aQ3fuwmY6jgObxewl3HUSCfV1TT4/uYV9GsrVx4p\np9LcyuVBuJroIlCJrk5Q/ozGhuiRoKApaTeUSjy5opziBRC2bF+TIxbO9Mkibtbh\nxvXZ4WE=\n-----END CERTIFICATE-----", "sp_private_key": "-----BEGIN PRIVATE KEY-----\nMIIEvgIBADANBgkqhkiG9w0BAQEFAASCBKgwggSkAgEAAoIBAQC8kGy41+C/F77G\nsE+v344LAPh2Tcz25r7XiX22VAcWC2oZcSrudhzSJz3CqKb6R3GklZ+LAbSYKWSU\n7ScQs2yAFWsLToWqtxJ7YIYSH3VQGQYQ7O86CbsQEk4H/k1zi0T/c1czShx9GAVc\n0S8YyQf0z6h0aTtlCtoj+F9rPBiCM63LUuu9IykqEJ/HQOkqUu5fTDY/0J8u5KBn\nRlN03diviJwS4SzSLem0J39uKnZ9xwqRFYSC4h7LCgOX/Z+O5skVzmwWGXtG1/4J\n4UqCp4HI09fb5QKLHXfovWv5UiUD7cKUic9ZlRhYJuz96e8E7qiZSyeo88JTbNV1\nhbRBv6utAgMBAAECggEAFPBeulnykZXD8BFVD/0dq1gkvxJdn884wvt4E76Z+Nc0\npXWdJFTGV4nXAF41CJbVZkbdLBT45mq2ShlZnK+n7UMzm1JRYocozL2Htcx7fPUC\naO++ku3QsXSu6JFTLXD6LPm0ZbQlnLiFo+xws+pi8Ur79E1ZNJuzZIooomJOgGqm\nz/0aTCw1JbMXAI7x0ygCYarfhqX4/M6qokV0Nt64hHxHxtrIWzVac+1QdR4WLaFL\nbdrb6QQeeCw5rWUrZfqmF6+NwCCeP5k/HMeVSwXsI+WrEVCjQBB2qpFqgiDNyfz2\n2i7UYXBP0PUmHEPsctWCYlWwqskBxLZnJdDKTmBCKwKBgQDpXpUpNgaI1LOrhxEQ\n5v1iXDSJweV8Kcdth+e6IGFLtxBgvhDNCijBhwKaFe90SFRldGQgZrz4tBKcxdEw\nslGbbNSSmVZ7nSMpZQoV74Uyrk2i7vxq6A9+ZCMWFpFIwoFBz4SUpnwEe+TEe/l3\nAMOz8BdFa3J0XzUhL7k5X+KaewKBgQDO2Yvi84JhwcmgRzlhv5o6gD40C5x1dv5w\nRqnXxnZGigVwtSBS6CoayBtL8MYNdTB5oM6qoF/FiVHYxbnwgD3d4net/BbUYz9H\nkONxwuEM0a35uSf2FHCaRn7BDPjmbdNmrkWr0bHyjlNAd1CQiwmeLxnaTwjf3C+M\nTdI+p08t9wKBgDms84Zk4MaOcv0we2pG/FaD3UQylInUNYJ/dSjN+d3hl32hW7uh\nCCOUP3NfenetrJYKZviPC6MXtgXi6el0GLEl+39jwDj6xAbl/tEfCjdVVsCu+dle\nEv40t2stFqj50UI3jFfEsZ/WEtrwnN3pZXSiIM46WOYj5ZiXF9rzNKjjAoGBAJcv\nwKPf8fM7rgA9Lr64SaTqqQxnVDMzByPPMkKpJzfFl9ZaPMb8NBIhInpuAIRDnGu5\n0nQ6BeYeyTjUxGP5h76O0YTUVWdlJxJK30L9+nnhI/T7lS6yn97TGcBGmAHsUfCh\n/gBoo1SzHDxpOPR8+0moCZBb5hOhHwvAsaPjq+bfAoGBALV/t+smBLogJOBXpfKU\n9LCXnnIqG804vyobSNCVoJm832gBTM7fVcTZa5I+0O+l1emEETIgKU+5ioP/qwou\nU3a/7jXX4hewCpmPVhvvlHgjs+UOBS6hXQMnq52h6mPhiikGOQ6YnqHtxyFORIlo\ntUlwMjanVlxRKyGJlBYQtADk\n-----END PRIVATE KEY-----" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` Create a route with the `saml-auth` plugin: adc.yaml ``` services: - name: saml-auth-service routes: - name: saml-auth-route uris: - /anything/* plugins: saml-auth: secret: my_secret_key sp_issuer: api7 idp_uri: "http://192.168.2.101:8080/realms/master/protocol/saml" login_callback_uri: /anything/login_callback logout_callback_uri: /anything/logout_callback logout_uri: /anything/logout logout_redirect_uri: /anything/logout_ok idp_cert: "-----BEGIN CERTIFICATE-----\nMIICmzCCAYMCBgGRl7njKjANBgkqhkiG9w0BAQsFADARMQ8wDQYDVQQDDAZtYXN0\nZXIwHhcNMjQwODI4MDY0MjA3WhcNMzQwODI4MDY0MzQ3WjARMQ8wDQYDVQQDDAZt\nYXN0ZXIwggEiMA0GCSqGSIb3DQEBAQUAA4IBDwAwggEKAoIBAQCdYPYSFoX2MADS\nIgfLYQ5oZcLNE+qB+qsO8sNpiebMQE3RmI5+MmZC/aozRzkzxcY+AoM50qfHrM1y\nM99A9ZxZt6fW/MuIv6IP5zWLDl0XWGVeOH0HIH4/xBxQetBxm1HdOYpCQg5Wm9hm\nYfebmN7NfW8HjnORjfUuUGgs5eCiHVqfiCfphLF5w+DcIcnjIwyF+xVH/7fRWgo5\ninBSeIavZh/LEv7LzBeRleGgoZ/+q7cVQiL2e0b8rsslqUOZJmwdPU3VSS0vW1bm\nXsZsfaZD0bgakFvSj0ARzwIbxc74eEQYKflHGS0zkrpm+TsO5KUn59SCPOhGNgGY\npKKv6cY1AgMBAAEwDQYJKoZIhvcNAQELBQADggEBABN21PoEiTaZ20qQUdKD03m+\nbySlF4jRX2AeZqCedBaW+nHrbefaJdEnE9AcXBENCWVr6ntdeREaL9dW6KpV1hT4\nBmnXO2aiFotZe4Vc2W6cv7nDpjil6Q5/isbT5sriYhcU9oXBAaLf9dlg7K/X1l1+\nzcy9Pd1uKUfrC+5ds/Zv+xHiiK4h55o8shcmBmQ7bsanzNmjIQNnyF+lNRciGRvg\nJp59TR7AWpiBQDTNW1KK3XjO9lmN8nCEPbpdNGi77TDX0OZVrbbPy3vL4n8Gi3oQ\nptHhmV7xou4fTEn9TCrdW82OLOduBCMk9t0tFFNB8Hlxq5XsLVLYW7O9GGcjDmI=\n-----END CERTIFICATE-----" sp_cert: "-----BEGIN CERTIFICATE-----\nMIIC0TCCAbmgAwIBAgIUAT7h3zLAul/3S1F9Ms9w7JjpoJ0wDQYJKoZIhvcNAQEL\nBQAwETEPMA0GA1UEAwwGQVBJU0lYMB4XDTI0MDgyNzA5MDk1NloXDTI1MDgyNzA5\nMDk1NlowETEPMA0GA1UEAwwGQVBJU0lYMIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8A\nMIIBCgKCAQEAvJBsuNfgvxe+xrBPr9+OCwD4dk3M9ua+14l9tlQHFgtqGXEq7nYc\n0ic9wqim+kdxpJWfiwG0mClklO0nELNsgBVrC06FqrcSe2CGEh91UBkGEOzvOgm7\nEBJOB/5Nc4tE/3NXM0ocfRgFXNEvGMkH9M+odGk7ZQraI/hfazwYgjOty1LrvSMp\nKhCfx0DpKlLuX0w2P9CfLuSgZ0ZTdN3Yr4icEuEs0i3ptCd/bip2fccKkRWEguIe\nywoDl/2fjubJFc5sFhl7Rtf+CeFKgqeByNPX2+UCix136L1r+VIlA+3ClInPWZUY\nWCbs/envBO6omUsnqPPCU2zVdYW0Qb+rrQIDAQABoyEwHzAdBgNVHQ4EFgQUvGIj\nuvPoHC74lhKSlOJAwrdq4WwwDQYJKoZIhvcNAQELBQADggEBAJU0+aKCUSYvN6oe\n7PHYD0ZvE13wItzKq/7DQQe1zA/kDoCvSyC8+gB+FZmdHmkGGNdNqXsQgHEnP7Y0\nx7gDqA3s0blXEkECfmmRcVxcS3rb8CVVFqiKdyRO91opdir5J9vbmiF7RK1ajFTy\nyemhK0xxFpPM+gTdetEj7AoVMrlRoOLC+L62GaSi/gpQmKPR91FLyj33vCfVrDCo\nQXYMPQmSbBCwlHrHWa/Px7F7aQ3fuwmY6jgObxewl3HUSCfV1TT4/uYV9GsrVx4p\np9LcyuVBuJroIlCJrk5Q/ozGhuiRoKApaTeUSjy5opziBRC2bF+TIxbO9Mkibtbh\nxvXZ4WE=\n-----END CERTIFICATE-----" sp_private_key: "-----BEGIN PRIVATE KEY-----\nMIIEvgIBADANBgkqhkiG9w0BAQEFAASCBKgwggSkAgEAAoIBAQC8kGy41+C/F77G\nsE+v344LAPh2Tcz25r7XiX22VAcWC2oZcSrudhzSJz3CqKb6R3GklZ+LAbSYKWSU\n7ScQs2yAFWsLToWqtxJ7YIYSH3VQGQYQ7O86CbsQEk4H/k1zi0T/c1czShx9GAVc\n0S8YyQf0z6h0aTtlCtoj+F9rPBiCM63LUuu9IykqEJ/HQOkqUu5fTDY/0J8u5KBn\nRlN03diviJwS4SzSLem0J39uKnZ9xwqRFYSC4h7LCgOX/Z+O5skVzmwWGXtG1/4J\n4UqCp4HI09fb5QKLHXfovWv5UiUD7cKUic9ZlRhYJuz96e8E7qiZSyeo88JTbNV1\nhbRBv6utAgMBAAECggEAFPBeulnykZXD8BFVD/0dq1gkvxJdn884wvt4E76Z+Nc0\npXWdJFTGV4nXAF41CJbVZkbdLBT45mq2ShlZnK+n7UMzm1JRYocozL2Htcx7fPUC\naO++ku3QsXSu6JFTLXD6LPm0ZbQlnLiFo+xws+pi8Ur79E1ZNJuzZIooomJOgGqm\nz/0aTCw1JbMXAI7x0ygCYarfhqX4/M6qokV0Nt64hHxHxtrIWzVac+1QdR4WLaFL\nbdrb6QQeeCw5rWUrZfqmF6+NwCCeP5k/HMeVSwXsI+WrEVCjQBB2qpFqgiDNyfz2\n2i7UYXBP0PUmHEPsctWCYlWwqskBxLZnJdDKTmBCKwKBgQDpXpUpNgaI1LOrhxEQ\n5v1iXDSJweV8Kcdth+e6IGFLtxBgvhDNCijBhwKaFe90SFRldGQgZrz4tBKcxdEw\nslGbbNSSmVZ7nSMpZQoV74Uyrk2i7vxq6A9+ZCMWFpFIwoFBz4SUpnwEe+TEe/l3\nAMOz8BdFa3J0XzUhL7k5X+KaewKBgQDO2Yvi84JhwcmgRzlhv5o6gD40C5x1dv5w\nRqnXxnZGigVwtSBS6CoayBtL8MYNdTB5oM6qoF/FiVHYxbnwgD3d4net/BbUYz9H\nkONxwuEM0a35uSf2FHCaRn7BDPjmbdNmrkWr0bHyjlNAd1CQiwmeLxnaTwjf3C+M\nTdI+p08t9wKBgDms84Zk4MaOcv0we2pG/FaD3UQylInUNYJ/dSjN+d3hl32hW7uh\nCCOUP3NfenetrJYKZviPC6MXtgXi6el0GLEl+39jwDj6xAbl/tEfCjdVVsCu+dle\nEv40t2stFqj50UI3jFfEsZ/WEtrwnN3pZXSiIM46WOYj5ZiXF9rzNKjjAoGBAJcv\nwKPf8fM7rgA9Lr64SaTqqQxnVDMzByPPMkKpJzfFl9ZaPMb8NBIhInpuAIRDnGu5\n0nQ6BeYeyTjUxGP5h76O0YTUVWdlJxJK30L9+nnhI/T7lS6yn97TGcBGmAHsUfCh\n/gBoo1SzHDxpOPR8+0moCZBb5hOhHwvAsaPjq+bfAoGBALV/t+smBLogJOBXpfKU\n9LCXnnIqG804vyobSNCVoJm832gBTM7fVcTZa5I+0O+l1emEETIgKU+5ioP/qwou\nU3a/7jXX4hewCpmPVhvvlHgjs+UOBS6hXQMnq52h6mPhiikGOQ6YnqHtxyFORIlo\ntUlwMjanVlxRKyGJlBYQtADk\n-----END PRIVATE KEY-----" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD saml-auth-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: saml-auth-plugin-config spec: plugins: - name: saml-auth config: secret: my_secret_key sp_issuer: api7 idp_uri: "http://192.168.2.101:8080/realms/master/protocol/saml" login_callback_uri: /anything/login_callback logout_callback_uri: /anything/logout_callback logout_uri: /anything/logout logout_redirect_uri: /anything/logout_ok idp_cert: "-----BEGIN CERTIFICATE-----\n...\n-----END CERTIFICATE-----" sp_cert: "-----BEGIN CERTIFICATE-----\n...\n-----END CERTIFICATE-----" sp_private_key: "-----BEGIN PRIVATE KEY-----\n...\n-----END PRIVATE KEY-----" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: saml-auth-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything/ filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: saml-auth-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f saml-auth-ic.yaml ``` saml-auth-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: saml-auth-route spec: ingressClassName: apisix http: - name: saml-auth-route match: paths: - /anything/* upstreams: - name: httpbin-external-domain plugins: - name: saml-auth enable: true config: secret: my_secret_key sp_issuer: api7 idp_uri: "http://192.168.2.101:8080/realms/master/protocol/saml" login_callback_uri: /anything/login_callback logout_callback_uri: /anything/logout_callback logout_uri: /anything/logout logout_redirect_uri: /anything/logout_ok idp_cert: "-----BEGIN CERTIFICATE-----\n...\n-----END CERTIFICATE-----" sp_cert: "-----BEGIN CERTIFICATE-----\n...\n-----END CERTIFICATE-----" sp_private_key: "-----BEGIN PRIVATE KEY-----\n...\n-----END PRIVATE KEY-----" ``` Apply the configuration to your cluster: ``` kubectl apply -f saml-auth-ic.yaml ``` #### Verify[​](#verify "Direct link to Verify") Navigate to [`http://127.0.0.1:9080/anything/saml-test`](http://127.0.0.1:9080/anything/saml-test) in your browser and log in with your Keycloak credentials. If successful, you should be redirected and see a response similar to the following in your browser: ``` { "args": {}, "data": "", "files": {}, "form": {}, "headers": { "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "Accept-Encoding": "gzip, deflate", "Accept-Language": "en-CA,en-US;q=0.9,en;q=0.8", "Cookie": "saml_session=90f84a61-cb03-4f8c-8202-5e7b5267bda6", "Host": "127.0.0.1", "Sec-Fetch-Dest": "document", "Sec-Fetch-Mode": "navigate", "Sec-Fetch-Site": "none", "Upgrade-Insecure-Requests": "1", "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.5 Safari/605.1.15", "X-Amzn-Trace-Id": "Root=1-66cf4d36-18bbacc80af8987b77b1f5c4", "X-Forwarded-Host": "127.0.0.1" }, "json": null, "method": "GET", "origin": "192.168.65.1, 203.91.85.123", "url": "http://127.0.0.1/anything/saml-test" } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * sp\_issuer string required *** The unique identifier the service provider (SP) uses when communicating with the identity provider (IdP) in the SAML authentication process. * idp\_uri string required *** The URL of the identity provider (IdP) where the service provider (SP) sends authentication requests to initiate the SAML authentication process. * idp\_cert string required *** The X.509 certificate provided by the identity provider (IdP), used by the service provider (SP) to verify the authenticity and integrity of SAML assertions and responses. * login\_callback\_uri string required *** The endpoint on the service provider (SP) where the identity provider (IdP) will send the SAML response after a user successfully authenticates. The login callback URI should be a sub-path of the route URI. For example, if the route `uri` is `/anything/*`, the login callback URI can be `/anything/login_callback`. * logout\_uri string required *** The URI path to trigger the SAML logout process. The logout URI should be a sub-path of the route URI. For example, if the route `uri` is `/anything/*`, the logout URI can be `/anything/logout`. * logout\_callback\_uri string required *** The endpoint on the service provider (SP) that receives the SAML logout response from the identity provider (IdP) after the logout process is completed. The logout callback URI should be a sub-path of the route URI. For example, if the route `uri` is `/anything/*`, the logout callback URI can be `/anything/logout_callback`. * logout\_redirect\_uri string required *** The URI where the user is redirected after the logout process is completed, usually back to the Service Provider's (SP) application or a specified landing page. The logout redirect URI should be a sub-path of the route URI. For example, if the route `uri` is `/anything/*`, the logout redirect URI can be `/anything/logout_ok`. * sp\_cert string required *** The X.509 certificate used by the service provider (SP) to sign SAML requests and assertions, ensuring secure communication with the identity provider (IdP). * sp\_private\_key string required *** The private key corresponding to the Service Provider's (SP) certificate `sp_cert`, used to sign SAML requests and decrypt SAML assertions. The value is encrypted with AES before being stored in the database. * secret string required vaild vaule: 8 to 32 characters *** A cryptographic secret used to derive encryption keys for securing SAML session data and tokens. The secret should be a strong, random string for security. This ensures that sensitive authentication information is encrypted and tamper-resistant. The value is encrypted with AES before being stored in the database. Available in API7 Enterprise from version 3.9.3 and APISIX from version 3.17.0. * auth\_protocol\_binding\_method string default: `HTTP-Redirect` vaild vaule: `HTTP-Redirect` or `HTTP-POST` *** Binding method for the authentication protocol. Available in API7 Enterprise from version 3.9.3 and APISIX from version 3.17.0. When the binding method is `HTTP-Redirect`, the plugin uses browser redirects via GET requests. The plugin does not explicitly configure cookie attributes for this binding; cookies follow the defaults of the browser or underlying HTTP stack (for example, `SameSite` typically defaults to `Lax`, and the `Secure` attribute may be omitted depending on the environment). When the binding method is `HTTP-POST`, the plugin sends SAML messages via POST requests. Cookies are explicitly configured with `SameSite=None` and the `Secure` attribute enabled to support cross-origin authentication over HTTPS. * secret\_fallbacks array\[string] *** An array of alternative secrets used during key rotation. The value is encrypted with AES before being stored in the database. Available in API7 Enterprise from version 3.9.3 and APISIX from version 3.17.0. --- # Serverless Functions The serverless functions consist of two plugins, `serverless-pre-function` and `serverless-post-function`. These plugins enable the execution of user-defined logic at the beginning and end of the [execution phases](https://docs.api7.ai/apisix/key-concepts/plugins.md#plugins-execution-lifecycle) the functions hook to. ## Tips for Writing Functions[​](#tips-for-writing-functions "Direct link to Tips for Writing Functions") Only Lua functions are allowed in the serverless plugins and not other Lua code. For example, anonymous functions are legal: ``` return function() ngx.log(ngx.ERR, 'one') end ``` Closures are also legal: ``` local count = 1 return function() count = count + 1 ngx.say(count) end ``` But code other than functions are illegal: ``` local count = 1 ngx.say(count) ``` ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure the `serverless-pre-function` and `serverless-post-function` plugins for different scenarios. ### Log Information before and after a Phase[​](#log-information-before-and-after-a-phase "Direct link to Log Information before and after a Phase") The example below demonstrates how you can configure the serverless plugins to execute custom logics to log information to error logs before and after the `rewrite` [phase](https://docs.api7.ai/apisix/key-concepts/plugins.md#plugins-execution-lifecycle). Create a route as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "id": "serverless-pre-route", "uri": "/anything", "plugins": { "serverless-pre-function": { "phase": "rewrite", "functions" : [ "return function() ngx.log(ngx.ERR, \"serverless pre function\"); end" ] }, "serverless-post-function": { "phase": "rewrite", "functions" : [ "return function(conf, ctx) ngx.log(ngx.ERR, \"match uri \", ctx.curr_req_matched and ctx.curr_req_matched._path); end" ] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: serverless-pre-route uris: - /anything plugins: serverless-pre-function: phase: rewrite functions: - | return function() ngx.log(ngx.ERR, "serverless pre function") end serverless-post-function: phase: rewrite functions: - | return function(conf, ctx) ngx.log(ngx.ERR, "match uri ", ctx.curr_req_matched and ctx.curr_req_matched._path) end upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD serverless-functions-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: serverless-functions-plugin-config spec: plugins: - name: serverless-pre-function config: phase: rewrite functions: - | return function() ngx.log(ngx.ERR, "serverless pre function") end - name: serverless-post-function config: phase: rewrite functions: - | return function(conf, ctx) ngx.log(ngx.ERR, "match uri ", ctx.curr_req_matched and ctx.curr_req_matched._path) end --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: serverless-pre-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: serverless-functions-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` serverless-functions-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: serverless-pre-route spec: ingressClassName: apisix http: - name: serverless-pre-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: serverless-pre-function config: phase: rewrite functions: - | return function() ngx.log(ngx.ERR, "serverless pre function") end - name: serverless-post-function config: phase: rewrite functions: - | return function(conf, ctx) ngx.log(ngx.ERR, "match uri ", ctx.curr_req_matched and ctx.curr_req_matched._path) end ``` Apply the configuration: ``` kubectl apply -f serverless-functions-ic.yaml ``` ❶ Hook the serverless pre-function logic to the `rewrite` [phase](https://docs.api7.ai/apisix/key-concepts/plugins.md#plugins-execution-lifecycle). ❷ Define a Lua function that logs a message of `serverless pre function` in the error log. ❸ Hook the serverless post-function logic to the `rewrite` [phase](https://docs.api7.ai/apisix/key-concepts/plugins.md#plugins-execution-lifecycle). ❹ Define a Lua function that logs the matched URI in the error log. `conf` and `ctx` can be passed as the first two arguments like other plugins, where `conf` is the plugin configurations and `ctx` is the request context. Send the request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response and see the following entries in the error log: ``` 2024/05/09 15:07:09 [error] 51#51: *3963 [lua] [string "return function() ngx.log(ngx.ERR, "serverles..."]:1: func(): serverless pre function, client: 172.21.0.1, server: _, request: "GET /anything HTTP/1.1", host: "127.0.0.1:9080" 2024/05/09 15:16:58 [error] 50#50: *9343 [lua] [string "return function(conf, ctx) ngx.log(ngx.ERR, "..."]:1: func(): match uri /anything, client: 172.21.0.1, server: _, request: "GET /anything HTTP/1.1", host: "127.0.0.1:9080" ``` The first entry is added by the pre-function and the second entry is added by the post-function. ### Register Custom Variables[​](#register-custom-variables "Direct link to Register Custom Variables") The example below demonstrates how you can register [custom built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) using the serverless plugins and use the newly created variable in logs. info This example cannot be completed with the Ingress Controller because it does not support configuring route labels. Start an example rsyslog server: ``` docker run -d -p 514:514 --name example-rsyslog-server rsyslog/syslog_appliance_alpine ``` Create a [service](https://docs.api7.ai/apisix/key-concepts/services.md) with a serverless function to register a custom variable `a6_route_labels`, enable a logging plugin to later log the custom variable, and configure an upstream: * Admin API * ADC ``` curl "http://127.0.0.1:9180/apisix/admin/services" -X PUT \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "id":"srv_custom_var", "plugins": { "serverless-pre-function": { "phase": "rewrite", "functions": [ "return function() local core = require \"apisix.core\" core.ctx.register_var(\"a6_route_labels\", function(ctx) local route = ctx.matched_route and ctx.matched_route.value if route and route.labels then return route.labels end return nil end); end" ] }, "syslog": { "host" : "172.0.0.1", "port" : 514, "flush_limit" : 1 } }, "upstream": { "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: srv-custom-var plugins: serverless-pre-function: phase: rewrite functions: - | return function() local core = require("apisix.core") core.ctx.register_var("a6_route_labels", function(ctx) local route = ctx.matched_route and ctx.matched_route.value if route and route.labels then return route.labels end return nil end) end syslog: host: 172.0.0.1 port: 514 flush_limit: 1 upstream: nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` ❶ `functions`: register a custom variable `a6_route_labels` and fetch the variable value from the matched route's `labels` property. ❷ `host` and `port`: replace with the address of your syslog server. ❸ `flush_limit`: set to 1 to push log to the syslog server immediately. Next, update the log format for all `syslog` instances with the new variable by configuring the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md): * Admin API * ADC ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/syslog" -X PUT \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "log_format": { "host": "$host", "client_ip": "$remote_addr", "labels": "$a6_route_labels" } }' ``` adc.yaml ``` plugin_metadata: syslog: log_format: host: "$host" client_ip: "$remote_addr" labels: "$a6_route_labels" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` ❶ `$host` and `$remote_addr`: NGINX variables. ❷ `$a6_route_labels`: custom variable. Finally, create a route: * Admin API * ADC ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "id":"route_custom_var", "uri":"/get", "service_id": "srv_custom_var", "labels": { "key": "test_a6_route_labels" } }' ``` adc.yaml ``` # Other Configs services: - name: srv-custom-var routes: - name: route-custom-var uris: - /get labels: key: test_a6_route_labels ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` ❶ In the Admin API example, set `service_id` to associate the route with the existing service. In ADC, the route is nested under the service definition. ❷ Add route `labels` so the custom variable can log them. To verify the variable registration, send a request to the route: ``` curl "http://127.0.0.1:9080/get" ``` You should see a log entry in your syslog server similar to the following: ``` { "host":"127.0.0.1", "route_id":"route_custom_var", "client_ip":"172.19.0.1", "labels":{ "key":"test_a6_route_labels" }, "service_id":"srv_custom_var" } ``` This verifies the custom variable was registered and it logs the `labels` information in a route successfully. ### Modify a Specific Field in Response Body[​](#modify-a-specific-field-in-response-body "Direct link to Modify a Specific Field in Response Body") The example below demonstrates how you can use the serverless plugins to remove a specific field from a JSON response body. Before proceeding with the removal, first configure a route as follows to see the unmodified response: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "id":"serverless-remove-body-info", "uri": "/get", "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: serverless-remove-body-info uris: - /get upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD serverless-remove-body-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: serverless-remove-body-info spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get backendRefs: - name: httpbin-external-domain port: 80 ``` serverless-remove-body-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: serverless-remove-body-info spec: ingressClassName: apisix http: - name: serverless-remove-body-info match: paths: - /get upstreams: - name: httpbin-external-domain ``` Apply the configuration: ``` kubectl apply -f serverless-remove-body-ic.yaml ``` Send a request to the route: ``` curl "http://127.0.0.1:9080/get" ``` You should see a response similar to the following with your host and proxy's IP information: ``` { "args": {}, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/8.4.0", "X-Amzn-Trace-Id": "Root=1-663db30f-51448a1b635f2f4338a4fcfc", "X-Forwarded-Host": "127.0.0.1" }, "origin": "172.19.0.1, 43.252.208.84", "url": "http://127.0.0.1/get" } ``` To remove the `origin` field from the response, update the route with serverless plugins: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes/serverless-remove-body-info" -X PATCH \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "plugins": { "serverless-pre-function": { "phase": "header_filter", "functions" : [ "return function(conf, ctx) local core = require(\"apisix.core\") core.response.clear_header_as_body_modified() end" ] }, "serverless-post-function": { "phase": "body_filter", "functions" : [ "return function(conf, ctx) local cjson = require(\"cjson\") local core = require(\"apisix.core\") local body = core.response.hold_body_chunk(ctx) if not body then return end body = cjson.decode(body) body.origin = nil body = cjson.encode(body) ngx.arg[1] = body end" ] } } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: serverless-remove-body-info uris: - /get plugins: serverless-pre-function: phase: header_filter functions: - | return function(conf, ctx) local core = require("apisix.core") core.response.clear_header_as_body_modified() end serverless-post-function: phase: body_filter functions: - | return function(conf, ctx) local cjson = require("cjson") local core = require("apisix.core") local body = core.response.hold_body_chunk(ctx) if not body then return end body = cjson.decode(body) body.origin = nil body = cjson.encode(body) ngx.arg[1] = body end upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD serverless-remove-body-ic.yaml ``` # Other Configs # --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: serverless-remove-body-plugin-config spec: plugins: - name: serverless-pre-function config: phase: header_filter functions: - | return function(conf, ctx) local core = require("apisix.core") core.response.clear_header_as_body_modified() end - name: serverless-post-function config: phase: body_filter functions: - | return function(conf, ctx) local cjson = require("cjson") local core = require("apisix.core") local body = core.response.hold_body_chunk(ctx) if not body then return end body = cjson.decode(body) body.origin = nil body = cjson.encode(body) ngx.arg[1] = body end --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: serverless-remove-body-info spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /get filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: serverless-remove-body-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` serverless-remove-body-ic.yaml ``` # Other Configs # --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: serverless-remove-body-info spec: ingressClassName: apisix http: - name: serverless-remove-body-info match: paths: - /get upstreams: - name: httpbin-external-domain plugins: - name: serverless-pre-function config: phase: header_filter functions: - | return function(conf, ctx) local core = require("apisix.core") core.response.clear_header_as_body_modified() end - name: serverless-post-function config: phase: body_filter functions: - | return function(conf, ctx) local cjson = require("cjson") local core = require("apisix.core") local body = core.response.hold_body_chunk(ctx) if not body then return end body = cjson.decode(body) body.origin = nil body = cjson.encode(body) ngx.arg[1] = body end ``` Apply the configuration: ``` kubectl apply -f serverless-remove-body-ic.yaml ``` ❶ Execute a pre-function in the `header_filter` [phase](https://docs.api7.ai/apisix/key-concepts/plugins.md#plugins-execution-lifecycle). ❷ Execute a post-function in the `body_filter` [phase](https://docs.api7.ai/apisix/key-concepts/plugins.md#plugins-execution-lifecycle). The pre-function calls `clear_header_as_body_modified` to clear body-related response headers such as `Content-Length`. The post-function collects the response body with `hold_body_chunk`, decodes the JSON payload, removes the `origin` field, and writes the updated body back to the response. Send another request to the route: ``` curl "http://127.0.0.1:9080/get" ``` You should see a response without the `origin` information: ``` { "url":"http://127.0.0.1/get", "args":{}, "headers":{ "X-Forwarded-Host":"127.0.0.1", "Host":"127.0.0.1", "Accept":"*/*", "User-Agent":"curl/8.4.0", "X-Amzn-Trace-Id":"Root=1-663db276-1c15276864294d963c6e1755" } } ``` For simpler response modifications, such as modifying HTTP status codes, request headers, or the entire response body, please use the [`response-rewrite`](https://docs.api7.ai/hub/response-rewrite.md) plugin. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * phase string default: `access` vaild vaule: `rewrite`, `access`, `header_filter`, `body_filter`, `log`, and `before_proxy` *** Phase before which the serverless function is executed. * functions array\[string] required *** List of functions that are executed sequentially. Only Lua functions are allowed and not other Lua code. For example, anonymous functions and closures are legal, whereas other code not in a function will not be allowed. See [tips for writing functions](https://docs.api7.ai/hub/serverless-functions.md#tips-for-writing-functions) for more details. --- # SkyWalking The `skywalking` plugin supports integration with [Apache SkyWalking](https://skywalking.apache.org) for request tracing. SkyWalking uses its native Nginx Lua tracer to provide tracing, topology analysis, and metrics from both service and URI perspectives. The gateway communicates with the SkyWalking server over HTTP. Tracing adds work to each sampled request. Use `sample_ratio` to balance trace coverage against that overhead, and use a lower ratio on high-throughput routes when full sampling is unnecessary. Requests that are not selected for sampling skip trace creation, but the performance effect depends on the workload and collector configuration. ## Example[​](#example "Direct link to Example") To follow along the example, start a SkyWalking storage, OAP server, and Booster UI: * Docker * Kubernetes Start a storage, OAP and Booster UI with Docker Compose, following [SkyWalking's documentation](https://skywalking.apache.org/docs/main/next/en/setup/backend/backend-docker/). Once set up, the OAP server should be listening on `12800` and you should be able to access the UI at . Deploy SkyWalking OAP server and UI to your Kubernetes cluster following [SkyWalking's documentation](https://skywalking.apache.org/docs/main/next/en/setup/backend/backend-k8s/). For a quick start, you can use the [SkyWalking Helm chart](https://skywalking.apache.org/docs/main/next/en/setup/backend/backend-k8s/#use-helm-to-install). Once deployed, the OAP server is typically accessible at `skywalking-oap.skywalking.svc.cluster.local:12800` within the cluster. After the SkyWalking OAP server is available, configure the gateway according to how it was deployed. In API7 Gateway, `skywalking` is available in Dashboard and Admin API by default. For APISIX deployments, load `skywalking` in the gateway plugin list before setting the endpoint address for the SkyWalking OAP server. * Host or Docker * Kubernetes (Helm) For APISIX host or Docker deployments, keep the existing plugin list in `config.yaml`, add `skywalking`, and update `plugin_attr.skywalking`: config.yaml ``` plugins: # Keep the complete plugin list used by your gateway. - skywalking plugin_attr: skywalking: report_interval: 3 service_name: APISIX service_instance_name: APISIX Instance endpoint_addr: http://192.168.2.103:12800 ``` Reload the gateway for configuration changes to take effect. For Helm deployments, update the values that render the SkyWalking plugin attributes. For APISIX, also update the value that renders the gateway plugin list. Keep the rest of your values file unchanged. For the APISIX Helm chart, `apisix.plugins` replaces the loaded plugin list. Start from the complete plugin list used by your gateway, add `skywalking`, and configure the plugin attributes under `apisix.pluginAttrs`: values.yaml ``` apisix: plugins: # Keep the complete plugin list used by your gateway. - skywalking pluginAttrs: skywalking: report_interval: 3 service_name: APISIX service_instance_name: APISIX Instance endpoint_addr: http://192.168.2.103:12800 ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` pluginAttrs: skywalking: report_interval: 3 service_name: APISIX service_instance_name: APISIX Instance endpoint_addr: http://192.168.2.103:12800 ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` ### Trace All Requests[​](#trace-all-requests "Direct link to Trace All Requests") The following example demonstrates how you can trace all requests passing through a route. Create a route with `skywalking` and configure the sampling ratio to be 1 to trace all requests: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "skywalking-route", "uri": "/anything", "plugins": { "skywalking": { "sample_ratio": 1 } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: skywalking-route plugins: skywalking: sample_ratio: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD skywalking-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: skywalking-plugin-config spec: plugins: - name: skywalking config: sample_ratio: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: skywalking-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: skywalking-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` skywalking-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: skywalking-route spec: ingressClassName: apisix http: - name: skywalking-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: skywalking enable: true config: sample_ratio: 1 ``` Apply the configuration to your cluster: ``` kubectl apply -f skywalking-ic.yaml ``` Send a few requests to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive `HTTP/1.1 200 OK` responses. In [SkyWalking UI](http://localhost:8080), navigate to **General Service** > **Services**. You should see a service called `APISIX` with traces corresponding to your requests: ![SkyWalking APISIX traces](https://static.api7.ai/uploads/2025/01/15/UdwiO8NJ_skywalking-traces.png) ### Associate Traces with Logs[​](#associate-traces-with-logs "Direct link to Associate Traces with Logs") The following example demonstrates how you can configure the `skywalking-logger` plugin on a route to log information of requests hitting the route. Create a route with the `skywalking-logger` plugin and configure the plugin with your OAP server URI: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "skywalking-logger-route", "uri": "/anything", "plugins": { "skywalking": { "sample_ratio": 1 }, "skywalking-logger": { "endpoint_addr": "http://192.168.2.103:12800" } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: skywalking-logger-route plugins: skywalking: sample_ratio: 1 skywalking-logger: endpoint_addr: "http://192.168.2.103:12800" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD skywalking-logs-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: skywalking-logs-config spec: plugins: - name: skywalking config: sample_ratio: 1 - name: skywalking-logger config: endpoint_addr: "http://skywalking-oap.skywalking.svc.cluster.local:12800" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: skywalking-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: skywalking-logs-config backendRefs: - name: httpbin-external-domain port: 80 ``` skywalking-logs-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: skywalking-route spec: ingressClassName: apisix http: - name: skywalking-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: skywalking enable: true config: sample_ratio: 1 - name: skywalking-logger enable: true config: endpoint_addr: "http://skywalking-oap.skywalking.svc.cluster.local:12800" ``` Apply the configuration to your cluster: ``` kubectl apply -f skywalking-logs-ic.yaml ``` Generate a few requests to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive `HTTP/1.1 200 OK` responses. In [SkyWalking UI](http://localhost:8080), navigate to **General Service** > **Services**. You should see a service called `APISIX` with a trace corresponding to your request, where you can view the associated logs: ![SkyWalking UI showing a trace timeline with the View Logs button highlighted](https://static.api7.ai/uploads/2025/01/16/soUpXm6b_trace-view-logs.png) ![SkyWalking UI showing the log entries associated with the selected trace](https://static.api7.ai/uploads/2025/01/16/XD934LvU_associated-logs.png) --- # skywalking-logger The `skywalking-logger` plugin pushes request and response logs as JSON objects to SkyWalking OAP server in batches and supports the customization of log formats. If there is an existing tracing context, it sets up the trace-log correlation automatically and relies on [SkyWalking Cross Process Propagation Headers Protocol](https://skywalking.apache.org/docs/main/next/en/api/x-process-propagation-headers-v3/). ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `skywalking-logger` plugin for different scenarios. * Docker * Kubernetes To follow along the example, start a storage, OAP and Booster UI with Docker Compose, following [SkyWalking's documentation](https://skywalking.apache.org/docs/main/next/en/setup/backend/backend-docker/). Once set up, the OAP server should be listening on `12800` and you should be able to access the UI at . To follow along the example, deploy SkyWalking OAP server and UI to your Kubernetes cluster following [SkyWalking's documentation](https://skywalking.apache.org/docs/main/next/en/setup/backend/backend-k8s/). For a quick start, you can use the [SkyWalking Helm chart](https://skywalking.apache.org/docs/main/next/en/setup/backend/backend-k8s/#use-helm-to-install). Once deployed, the OAP server is typically accessible at `skywalking-oap.skywalking.svc.cluster.local:12800` within the cluster. ### Log Requests in Default Log Format[​](#log-requests-in-default-log-format "Direct link to Log Requests in Default Log Format") The following example demonstrates how you can configure the `skywalking-logger` plugin on a route to log information of requests hitting the route. Create a route with the `skywalking-logger` plugin and configure the plugin with your OAP server URI: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "skywalking-logger-route", "uri": "/anything", "plugins": { "skywalking-logger": { "endpoint_addr": "http://192.168.2.103:12800" } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: skywalking-logger-route plugins: skywalking-logger: endpoint_addr: "http://192.168.2.103:12800" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD skywalking-logger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: skywalking-logger-plugin-config spec: plugins: - name: skywalking-logger config: endpoint_addr: "http://skywalking-oap.skywalking.svc.cluster.local:12800" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: skywalking-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: skywalking-logger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` skywalking-logger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: skywalking-logger-route spec: ingressClassName: apisix http: - name: skywalking-logger-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: skywalking-logger enable: true config: endpoint_addr: "http://skywalking-oap.skywalking.svc.cluster.local:12800" ``` Apply the configuration to your cluster: ``` kubectl apply -f skywalking-logger-ic.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. In [SkyWalking UI](http://localhost:8080), navigate to **General Service** > **Services**. You should see a service called `APISIX` with a log entry corresponding to your request: ``` { "upstream_latency": 674, "request": { "method": "GET", "headers": { "user-agent": "curl/8.6.0", "host": "127.0.0.1:9080", "accept": "*/*" }, "url": "http://127.0.0.1:9080/anything", "size": 85, "querystring": {}, "uri": "/anything" }, "client_ip": "192.168.65.1", "route_id": "skywalking-logger-route", "start_time": 1736945107345, "upstream": "3.210.94.60:80", "server": { "version": "3.13.0", "hostname": "7edbcebe8eb3" }, "service_id": "", "response": { "size": 619, "status": 200, "headers": { "content-type": "application/json", "date": "Thu, 16 Jan 2025 12:45:08 GMT", "server": "APISIX/3.13.0", "access-control-allow-origin": "*", "connection": "close", "access-control-allow-credentials": "true", "content-length": "391" } }, "latency": 764.9998664856, "apisix_latency": 90.999866485596 } ``` ### Log Request and Response Headers With Plugin Metadata[​](#log-request-and-response-headers-with-plugin-metadata "Direct link to Log Request and Response Headers With Plugin Metadata") The following example demonstrates how you can customize log format using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) and [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) to log specific headers from request and response. In APISIX, [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) is used to configure the common metadata fields of all plugin instances of the same plugin. It is useful when a plugin is enabled across multiple resources and requires a universal update to their metadata fields. First, create a route with the `skywalking-logger` plugin and configure the plugin with your OAP server URI (same as [Log Requests in Default Log Format](#log-requests-in-default-log-format)). Next, configure the plugin metadata for `skywalking-logger`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/skywalking-logger" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "log_format": { "host": "$host", "@timestamp": "$time_iso8601", "client_ip": "$remote_addr", "env": "$http_env", "resp_content_type": "$sent_http_Content_Type" } }' ``` adc.yaml ``` plugin_metadata: - name: skywalking-logger log_format: host: "$host" "@timestamp": "$time_iso8601" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Update the `pluginMetadata` field in your existing `GatewayProxy` resource: gateway-proxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # your control plane connection configuration # .... pluginMetadata: skywalking-logger: log_format: host: "$host" "@timestamp": "$time_iso8601" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` Apply the configuration to your cluster: ``` kubectl apply -f gateway-proxy.yaml ``` ❶ Log the custom request header `env`. ❷ Log the response header `Content-Type`. Send a request to the route with the `env` header: ``` curl -i "http://127.0.0.1:9080/anything" -H "env: dev" ``` You should receive an `HTTP/1.1 200 OK` response. In [SkyWalking UI](http://localhost:8080), navigate to **General Service** > **Services**. You should see a service called `APISIX` with a log entry corresponding to your request: ``` [ { "route_id": "skywalking-logger-route", "client_ip": "192.168.65.1", "@timestamp": "2025-01-16T12:51:53+00:00", "host": "127.0.0.1", "env": "dev", "resp_content_type": "application/json" } ] ``` ### Log Request Bodies Conditionally[​](#log-request-bodies-conditionally "Direct link to Log Request Bodies Conditionally") The following example demonstrates how you can conditionally log request body. Create a route with the `skywalking-logger` plugin as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "skywalking-logger-route", "uri": "/anything", "plugins": { "skywalking-logger": { "endpoint_addr": "http://192.168.2.103:12800", "include_req_body": true, "include_req_body_expr": [["arg_log_body", "==", "yes"]] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: skywalking-logger-route plugins: skywalking-logger: endpoint_addr: "http://192.168.2.103:12800" include_req_body: true include_req_body_expr: - ["arg_log_body", "==", "yes"] upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD skywalking-logger-body-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: skywalking-logger-body-config spec: plugins: - name: skywalking-logger config: endpoint_addr: "http://skywalking-oap.skywalking.svc.cluster.local:12800" include_req_body: true include_req_body_expr: - ["arg_log_body", "==", "yes"] --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: skywalking-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: skywalking-logger-body-config backendRefs: - name: httpbin-external-domain port: 80 ``` skywalking-logger-body-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: skywalking-logger-route spec: ingressClassName: apisix http: - name: skywalking-logger-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: skywalking-logger enable: true config: endpoint_addr: "http://skywalking-oap.skywalking.svc.cluster.local:12800" include_req_body: true include_req_body_expr: - ["arg_log_body", "==", "yes"] ``` Apply the configuration to your cluster: ``` kubectl apply -f skywalking-logger-body-ic.yaml ``` ❶ `include_req_body`: Set to true to include request body. ❷ `include_req_body_expr`: Only include request body if the URL query string `log_body` is `yes`. Send a request to the route with a URL query string satisfying the condition: ``` curl -i "http://127.0.0.1:9080/anything?log_body=yes" -X POST -d '{"env": "dev"}' ``` You should receive an `HTTP/1.1 200 OK` response. In [SkyWalking UI](http://localhost:8080), navigate to **General Service** > **Services**. You should see a service called `APISIX` with a log entry corresponding to your request, with the request body logged: ``` [ { "request": { "url": "http://127.0.0.1:9080/anything?log_body=yes", "querystring": { "log_body": "yes" }, "uri": "/anything?log_body=yes", ..., "body": "{\"env\": \"dev\"}", }, ... } ] ``` Send a request to the route without any URL query string: ``` curl -i "http://127.0.0.1:9080/anything" -X POST -d '{"env": "dev"}' ``` You should not observe a log entry without the request body. info If you have customized the `log_format` in addition to setting `include_req_body` or `include_resp_body` to `true`, the plugin would not include the bodies in the logs. As a workaround, you may be able to use the NGINX variable `$request_body` in the log format, such as: ``` { "skywalking-logger": { ..., "log_format": {"body": "$request_body"} } } ``` ### Associate Traces with Logs[​](#associate-traces-with-logs "Direct link to Associate Traces with Logs") The following example demonstrates how you can configure the `skywalking-logger` plugin on a route to log information of requests hitting the route. SkyWalking setup This example also requires the `skywalking` plugin to be enabled globally and configured with a reachable OAP endpoint address. For Helm deployments, see the [SkyWalking plugin setup](https://docs.api7.ai/hub/skywalking.md#example). Create a route with the `skywalking-logger` plugin and configure the plugin with your OAP server URI: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "skywalking-logger-route", "uri": "/anything", "plugins": { "skywalking": { "sample_ratio": 1 }, "skywalking-logger": { "endpoint_addr": "http://192.168.2.103:12800" } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: skywalking-logger-route plugins: skywalking: sample_ratio: 1 skywalking-logger: endpoint_addr: "http://192.168.2.103:12800" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD skywalking-logger-trace-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: skywalking-logger-trace-config spec: plugins: - name: skywalking config: sample_ratio: 1 - name: skywalking-logger config: endpoint_addr: "http://skywalking-oap.skywalking.svc.cluster.local:12800" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: skywalking-logger-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: skywalking-logger-trace-config backendRefs: - name: httpbin-external-domain port: 80 ``` skywalking-logger-trace-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: skywalking-logger-route spec: ingressClassName: apisix http: - name: skywalking-logger-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: skywalking enable: true config: sample_ratio: 1 - name: skywalking-logger enable: true config: endpoint_addr: "http://skywalking-oap.skywalking.svc.cluster.local:12800" ``` Apply the configuration to your cluster: ``` kubectl apply -f skywalking-logger-trace-ic.yaml ``` Generate a few requests to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive `HTTP/1.1 200 OK` responses. In [SkyWalking UI](http://localhost:8080), navigate to **General Service** > **Services**. You should see a service called `APISIX` with a trace corresponding to your request, where you can view the associated logs: ![SkyWalking UI showing a trace timeline with the View Logs button highlighted](https://static.api7.ai/uploads/2025/01/16/soUpXm6b_trace-view-logs.png) ![SkyWalking UI showing the log entries associated with the selected trace](https://static.api7.ai/uploads/2025/01/16/XD934LvU_associated-logs.png) --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * endpoint\_addr string required *** URI of the SkyWalking OAP server. * service\_name string default: `APISIX` *** Service name for the SkyWalking reporter. * service\_instance\_name string default: `APISIX Instance Name` *** Service instance name for the SkyWalking reporter. Set to `$hostname` to get the local hostname. * timeout integer default: `3` vaild vaule: greater than 0 *** Time to keep the connection alive after sending a request. * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. You can also configure log format on a global scale using the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md), which configures the log format for all `skywalking-logger` plugin instances. If the log format configured on the individual plugin instance differs from the log format configured on plugin metadata, the log format configured on the individual plugin instance takes precedence. See the [example](https://docs.api7.ai/hub/skywalking-logger.md#log-request-and-response-headers-with-plugin-metadata) for more details. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * include\_req\_body boolean default: `false` *** If true, include the request body in the log. Note that if the request body is too big to be kept in the memory, it can not be logged due to NGINX's limitations. * include\_req\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_req_body` is true. Request body would only be logged when the expressions configured here evaluate to true. * include\_resp\_body boolean default: `false` *** If true, include the response body in the log. * include\_resp\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_resp_body` is true. Response body would only be logged when the expressions configured here evaluate to true. * max\_req\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes to include in the log. If the request body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * max\_resp\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum response body size in bytes to include in the log. If the response body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * name string default: `skywalking logger` *** Unique identifier of the plugin for the batch processor. If you use [Prometheus](https://docs.api7.ai/hub/prometheus.md) to monitor APISIX metrics, the name is exported in `apisix_batch_process_entries`. * batch\_max\_size integer default: `1000` vaild vaule: greater than 0 *** The number of log entries allowed in one batch. Once reached, the batch will be sent to the logging service. Setting this parameter to 1 means immediate processing. * inactive\_timeout integer default: `5` vaild vaule: greater than 0 *** The maximum time in seconds to wait for new logs before sending the batch to the logging service. The value should be smaller than `buffer_duration`. * buffer\_duration integer default: `60` vaild vaule: greater than 0 *** The maximum time in seconds from the earliest entry allowed before sending the batch to the logging service. * retry\_delay integer default: `1` vaild vaule: greater than or equal to 0 *** The time interval in seconds to retry sending the batch to the logging service if the batch was not successfully sent. * max\_retry\_count integer default: `0` vaild vaule: greater than or equal to 0 *** The maximum number of unsuccessful retries allowed before dropping the log entries. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * max\_pending\_entries integer default: `` `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 and 3.10.6; none in API7 Enterprise 3.9.18 and 3.10.5 `` vaild vaule: greater than or equal to 1 *** Maximum number of entries waiting in the batch processor. New entries are discarded when the backlog reaches the limit. Introduced in API7 Enterprise 3.8.17 and APISIX 3.15.0. The default changed to `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line. In API7 Enterprise 3.9.18 and 3.10.5, and in earlier APISIX versions, omitting the parameter leaves the backlog unlimited. See [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md#configure-the-pending-entry-limit) for sizing and verification guidance. --- ## Static Configurations[​](#static-configurations "Direct link to Static Configurations") By default, service names and endpoint address for the plugin are pre-configured in the [default configuration](https://github.com/apache/apisix/blob/master/apisix/cli/config.lua). The file to update depends on how the gateway is deployed: * Host or Docker * Kubernetes (Helm) For host or Docker deployments, configure the following settings: config.yaml ``` plugin_attr: skywalking: report_interval: 3 # Reporting interval time in seconds. service_name: APISIX # Service name for SkyWalking reporter. service_instance_name: "APISIX Instance Name" # Service instance name for SkyWalking reporter. # Set to $hostname to get the local hostname. endpoint_addr: http://127.0.0.1:12800 # SkyWalking HTTP endpoint. ``` Then reload the gateway for changes to take effect. For Helm deployments, update the chart values that render `plugin_attr.skywalking`. For APISIX, also confirm that the plugin is loaded in the gateway plugin list before applying them. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` apisix: pluginAttrs: skywalking: report_interval: 3 service_name: APISIX service_instance_name: "APISIX Instance Name" endpoint_addr: http://127.0.0.1:12800 ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` pluginAttrs: skywalking: report_interval: 3 service_name: APISIX service_instance_name: "APISIX Instance Name" endpoint_addr: http://127.0.0.1:12800 ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * sample\_ratio number default: `1` vaild vaule: between 0.00001 and 1 inclusive *** Frequency of request sampling. Setting the sample ratio to `1` means to sample all requests. --- # soap [Enterprise](https://api7.ai/enterprise) The `soap` plugin provides a convenient approach to transform between RESTful HTTP requests and SOAP requests, as well as their corresponding responses. With a single URL to the [WSDL](https://en.wikipedia.org/wiki/Web_Services_Description_Language) file, API7 automatically parses the file content and generates conversion logics to allow for the protocol transcoding. ## Examples[​](#examples "Direct link to Examples") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") The `soap` plugin depends on `soap-proxy`, a separate service that handles WSDL parsing and JSON-to-SOAP transcoding. For Docker deployments, start `soap-proxy` before using the plugin. For Helm deployments, enable the chart-managed `soap-proxy` sidecar when configuring the gateway values. #### Start SOAP Proxy[​](#start-soap-proxy "Direct link to Start SOAP Proxy") * Docker * Kubernetes (Helm) Start the `soap-proxy` container on the same network as your APISIX instance (adjust accordingly for your environment): ``` docker run -d \ --name soap-proxy \ --network=apisix-quickstart-net \ -p 5000:5000 \ api7/soap-proxy:1.0.0 ``` For Helm deployments, enable the chart-managed `soap-proxy` sidecar in the gateway values. Keep the rest of your values file unchanged. values.yaml ``` soapProxy: enabled: true ``` Then apply the values file to the existing gateway release: ``` helm upgrade api7/gateway -n -f values.yaml ``` #### Configure the Gateway[​](#configure-the-gateway "Direct link to Configure the Gateway") By default, the `soap` plugin expects `soap-proxy` to be available at `http://127.0.0.1:5000`. Update the gateway static configuration to point the plugin to the proxy service address used by your deployment: * Docker * Kubernetes (Helm) Add the following to your `config.yaml`: config.yaml ``` plugin_attr: soap: endpoint: http://soap-proxy:5000 timeout: 3000 ``` Reload the gateway for the changes to take effect: ``` apisix reload ``` For API7 Gateway Helm deployments, enable the chart-managed SOAP proxy sidecar and update `pluginAttrs.soap` in the chart values. The chart renders this value to `plugin_attr.soap` in the generated gateway configuration. Keep the rest of your values file unchanged. values.yaml ``` soapProxy: enabled: true pluginAttrs: soap: endpoint: http://127.0.0.1:5000 timeout: 3000 ``` Then apply the values file to the existing gateway release: ``` helm upgrade api7/gateway -n -f values.yaml ``` ### Invoke an Operation[​](#invoke-an-operation "Direct link to Invoke an Operation") The following example demonstrates how you can configure the plugin on a route and invoke an operation available on the upstream server as specified in the WSDL file. Create a route with the `soap` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H 'X-API-KEY: ${ADMIN_API_KEY}' \ -d '{ "id": "soap-hello", "uri": "/SayHello", "methods": ["POST"], "plugins": { "soap": { "wsdl_url": "https://apps.learnwebservices.com/services/hello?wsdl" } } }' ``` adc.yaml ``` services: - name: soap-service routes: - name: soap-hello uris: - /SayHello methods: - POST plugins: soap: wsdl_url: "https://apps.learnwebservices.com/services/hello?wsdl" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD soap-ic.yaml ``` apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: soap-hello spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /SayHello method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: soap-plugin-config --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: soap-plugin-config spec: plugins: - name: soap config: wsdl_url: "https://apps.learnwebservices.com/services/hello?wsdl" ``` soap-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: soap-hello spec: ingressClassName: apisix http: - name: soap-hello match: paths: - /SayHello methods: - POST plugins: - name: soap enable: true config: wsdl_url: "https://apps.learnwebservices.com/services/hello?wsdl" ``` Apply the configuration: ``` kubectl apply -f soap-ic.yaml ``` ❶ Set the URI to the name of the operation in the WSDL file. ❷ Allow only for the POST request method. ❸ Set the URL path to the WSDL file. Send a request to the route to verify: ``` curl "http://127.0.0.1:9080/SayHello" -X POST -d '{"Name": "John Doe"}' ``` You should see an `HTTP/1.1 200 OK` response with the following: ``` "Hello John Doe!" ``` --- ## Static Configurations[​](#static-configurations "Direct link to Static Configurations") By default, values such as SOAP proxy `endpoint` and `timeout` are pre-configured in the default configuration file `config-default.yaml`. The file to update depends on how the gateway is deployed: * Host or Docker * Kubernetes (Helm) For host or Docker deployments, configure the following settings: config.yaml ``` plugin_attr: soap: endpoint: http://127.0.0.1:5000 timeout: 3000 # in milliseconds ``` Then [reload the gateway](https://docs.api7.ai/apisix/reference/apisix-cli.md#apisix-reload) for changes to take effect. For Helm deployments, set the following values in the API7 Gateway Helm chart. Keep the rest of your values file unchanged. values.yaml ``` soapProxy: enabled: true pluginAttrs: soap: endpoint: http://127.0.0.1:5000 timeout: 3000 ``` Then apply the values file to the existing gateway release: ``` helm upgrade api7/gateway -n -f values.yaml ``` ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * wsdl\_url string required *** URL to the WSDL file. * max\_req\_body\_size integer default: `67108864` *** Maximum request body size in bytes buffered into memory. Larger request bodies are rejected. Available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line. --- # splunk-hec-logging The `splunk-hec-logging` plugin serializes request and response context information to [Splunk Event Data format](https://docs.splunk.com/Documentation/Splunk/latest/Data/FormateventsforHTTPEventCollector#Event_metadata) and push to your [Splunk HTTP Event Collector (HEC)](https://docs.splunk.com/Documentation/Splunk/latest/Data/UsetheHTTPEventCollector) in batches. The plugin also supports the customization of log formats. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `splunk-hec-logging` plugin for different scenarios. To follow along the examples, set up a Splunk HEC endpoint: * Local Splunk * Kubernetes Complete the following steps to set up Splunk: 1. Install [Splunk](https://www.splunk.com/en_us/download.html). Splunk Web should be running at `localhost:8000` by default. 2. See [set up and use HTTP Event Collector in Splunk Web](https://docs.splunk.com/Documentation/Splunk/latest/Data/UsetheHTTPEventCollector) to create an HTTP Event Collector. 3. Navigate to **Settings > Data Inputs** and note down the token value. 4. In **HTTP Event Collector > Global Settings**, enable all tokens and note down the collector port, which defaults to `8088`. To verify the setup, execute the following command with your token: ``` curl "http://localhost:8088/services/collector/event" \ -H "Authorization: Splunk " \ -d '{"event": "hello world"}' ``` You should see a `success` response. Create a Kubernetes manifest to deploy Splunk with HEC enabled: splunk-hec-server.yaml ``` apiVersion: v1 kind: ConfigMap metadata: namespace: aic name: splunk-defaults data: default.yml: | splunk: hec: enable: True ssl: False token: apisix-hec-token --- apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: splunk spec: replicas: 1 selector: matchLabels: app: splunk template: metadata: labels: app: splunk spec: enableServiceLinks: false containers: - name: splunk image: splunk/splunk:9.4 env: - name: SPLUNK_START_ARGS value: "--accept-license" # Accept Splunk General Terms: https://www.splunk.com/en_us/legal/splunk-general-terms.html - name: SPLUNK_GENERAL_TERMS value: "--accept-sgt-current-at-splunk-com" - name: SPLUNK_PASSWORD value: "Splunk@1234" ports: - containerPort: 8088 - containerPort: 8000 volumeMounts: - name: defaults mountPath: /tmp/defaults readinessProbe: httpGet: path: /services/collector/health port: 8088 initialDelaySeconds: 60 periodSeconds: 10 failureThreshold: 10 volumes: - name: defaults configMap: name: splunk-defaults --- apiVersion: v1 kind: Service metadata: namespace: aic name: splunk-hec spec: selector: app: splunk ports: - name: hec port: 8088 targetPort: 8088 - name: web port: 8000 targetPort: 8000 type: ClusterIP ``` Apply the manifest: ``` kubectl apply -f splunk-hec-server.yaml ``` Port forward the Splunk Web port to your local machine: ``` kubectl port-forward -n aic svc/splunk-hec 8000:8000 ``` Then open `http://localhost:8000` and log in with username `admin` and password `Splunk@1234`. ### Push Log to Splunk[​](#push-log-to-splunk "Direct link to Push Log to Splunk") The following example demonstrates how you can enable the `splunk-hec-logging` plugin on a route, which logs client requests and pushes logs to Splunk. Create a route as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "splunk-route", "uri": "/anything", "plugins": { "splunk-hec-logging": { "endpoint": { "uri": "http://127.0.0.1:8088/services/collector/event", "token": "example-splunk-hec-token" } } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: splunk-route uris: - /anything plugins: splunk-hec-logging: endpoint: uri: http://127.0.0.1:8088/services/collector/event token: example-splunk-hec-token upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD splunk-hec-logging-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: splunk-hec-logging-plugin-config spec: plugins: - name: splunk-hec-logging config: endpoint: uri: http://splunk-hec.aic.svc.cluster.local:8088/services/collector/event token: apisix-hec-token --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: splunk-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: splunk-hec-logging-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` splunk-hec-logging-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: splunk-route spec: ingressClassName: apisix http: - name: splunk-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: splunk-hec-logging enable: true config: endpoint: uri: http://splunk-hec.aic.svc.cluster.local:8088/services/collector/event token: apisix-hec-token ``` Apply the configuration: ``` kubectl apply -f splunk-hec-logging-ic.yaml ``` ❶ Configure the Splunk HTTP collector endpoint. For Kubernetes, use the in-cluster Service address such as `http://splunk-hec.aic.svc.cluster.local:8088/services/collector/event`. ❷ Replace with your collector token. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to Splunk Web and select **Search & Reporting** in the left menu. In the search box, enter `source="apache-apisix-splunk-hec-logging"` and search for events from APISIX. ### Log Request and Response Headers With Plugin Metadata[​](#log-request-and-response-headers-with-plugin-metadata "Direct link to Log Request and Response Headers With Plugin Metadata") The following example demonstrates how you can customize log format using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) and [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) to log specific headers from request and response. In APISIX, [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md) is used to configure the common metadata fields of all plugin instances of the same plugin. It is useful when a plugin is enabled across multiple resources and requires a universal update to their metadata fields. Create a route as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "splunk-route", "uri": "/anything", "plugins": { "splunk-hec-logging": { "endpoint": { "uri": "http://127.0.0.1:8088/services/collector/event", "token": "example-splunk-hec-token" } } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: splunk-route uris: - /anything plugins: splunk-hec-logging: endpoint: uri: http://127.0.0.1:8088/services/collector/event token: example-splunk-hec-token upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD splunk-hec-logging-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: splunk-hec-logging-plugin-config spec: plugins: - name: splunk-hec-logging config: endpoint: uri: http://splunk-hec.aic.svc.cluster.local:8088/services/collector/event token: apisix-hec-token --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: splunk-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: splunk-hec-logging-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` splunk-hec-logging-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: splunk-route spec: ingressClassName: apisix http: - name: splunk-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: splunk-hec-logging enable: true config: endpoint: uri: http://splunk-hec.aic.svc.cluster.local:8088/services/collector/event token: apisix-hec-token ``` Apply the configuration: ``` kubectl apply -f splunk-hec-logging-ic.yaml ``` Configure the plugin metadata for `splunk-hec-logging`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/splunk-hec-logging" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "log_format": { "host": "$host", "@timestamp": "$time_iso8601", "route_id": "$route_id", "client_ip": "$remote_addr", "env": "$http_env", "resp_content_type": "$sent_http_Content_Type" } }' ``` adc.yaml ``` plugin_metadata: - name: splunk-hec-logging log_format: host: "$host" "@timestamp": "$time_iso8601" route_id: "$route_id" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: splunk-hec-logging: log_format: host: "$host" "@timestamp": "$time_iso8601" route_id: "$route_id" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: splunk-hec-logging: log_format: host: "$host" "@timestamp": "$time_iso8601" route_id: "$route_id" client_ip: "$remote_addr" env: "$http_env" resp_content_type: "$sent_http_Content_Type" ``` Apply the configuration: ``` kubectl apply -f gatewayproxy.yaml ``` ❶ log the custom request header `env`. ❷ log the response header `Content-Type`. Send a request to the route with the `env` header: ``` curl -i "http://127.0.0.1:9080/anything" -H "env: dev" ``` Navigate to Splunk Web and select **Search & Reporting** in the left menu. In the search box, enter `source="apache-apisix-splunk-hec-logging"` and search for events. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * endpoint object\[object] required *** Splunk HEC endpoint configurations. * uri string required *** Splunk HEC event collector API endpoint. * token string required *** Splunk HEC authentication token. API7 Gateway encrypts the value with AES at rest. APISIX encrypts it before etcd storage when `apisix.data_encryption.enable_encrypt_fields` is enabled. * channel string *** Splunk HEC send data channel identifier. For more information, see [About HTTP Event Collector Indexer Acknowledgment](https://docs.splunk.com/Documentation/Splunk/latest/Data/AboutHECIDXAck). * timeout integer default: `10` *** Splunk HEC send data timeout in seconds. * keepalive\_timeout integer default: `60000` vaild vaule: greater than or equal to 1000 *** Keepalive timeout in milliseconds. * ssl\_verify boolean default: `true` *** * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. You can also configure log format on a global scale using the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md), which configures the log format for all `splunk-hec-logging` plugin instances. If the log format configured on the individual plugin instance differs from the log format configured on plugin metadata, the log format configured on the individual plugin instance takes precedence. See the [example](https://docs.api7.ai/hub/splunk-hec-logging.md#log-request-and-response-headers-with-plugin-metadata) for more details. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * name string default: `splunk-hec-logging` *** Unique identifier of the plugin for the batch processor. * batch\_max\_size integer default: `1000` vaild vaule: greater than 0 *** The number of log entries allowed in one batch. Once reached, the batch will be sent to Splunk HEC. Setting this parameter to 1 means immediate processing. * inactive\_timeout integer default: `5` vaild vaule: greater than 0 *** The maximum time in seconds to wait for new logs before sending the batch to the logging service. The value should be smaller than `buffer_duration`. * buffer\_duration integer default: `60` vaild vaule: greater than 0 *** The maximum time in seconds from the earliest entry allowed before sending the batch to the logging service. * retry\_delay integer default: `1` vaild vaule: greater than or equal to 0 *** The time interval in seconds to retry sending the batch to the logging service if the batch was not successfully sent. * max\_retry\_count integer default: `0` vaild vaule: greater than or equal to 0 *** The maximum number of unsuccessful retries allowed before dropping the log entries. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * max\_pending\_entries integer default: `` `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 and 3.10.6; none in API7 Enterprise 3.9.18 and 3.10.5 `` vaild vaule: greater than or equal to 1 *** Maximum number of entries waiting in the batch processor. New entries are discarded when the backlog reaches the limit. Introduced in API7 Enterprise 3.8.17 and APISIX 3.15.0. The default changed to `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line. In API7 Enterprise 3.9.18 and 3.10.5, and in earlier APISIX versions, omitting the parameter leaves the backlog unlimited. See [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md#configure-the-pending-entry-limit) for sizing and verification guidance. --- # syslog The `syslog` plugin pushes request and response logs as JSON objects to syslog servers in batches and supports the customization of log formats. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `syslog` plugin for different scenarios. To follow along the examples, prepare a syslog server: * Docker * Kubernetes ``` docker run -d -p 514:514/tcp --name example-rsyslog-server rsyslog/syslog_appliance_alpine ``` To view the logs received by the server: ``` docker logs -f example-rsyslog-server ``` Create a Kubernetes manifest for a sample TCP syslog receiver: syslog-server.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: example-rsyslog-server spec: replicas: 1 selector: matchLabels: app: example-rsyslog-server template: metadata: labels: app: example-rsyslog-server spec: containers: - name: tcp-syslog image: alpine/socat args: - -v - TCP-LISTEN:514,reuseaddr,fork - STDOUT ports: - containerPort: 514 protocol: TCP --- apiVersion: v1 kind: Service metadata: namespace: aic name: example-rsyslog-server spec: selector: app: example-rsyslog-server ports: - name: tcp-syslog port: 514 targetPort: 514 protocol: TCP type: ClusterIP ``` Apply the manifest: ``` kubectl apply -f syslog-server.yaml ``` To view the logs received by the server: ``` kubectl logs -n aic deploy/example-rsyslog-server -f ``` ### Push Log to Syslog Server[​](#push-log-to-syslog-server "Direct link to Push Log to Syslog Server") The following example demonstrates how you can enable the `syslog` plugin on a route, which logs client requests to the route and pushes logs to a syslog server. Create a route with `syslog` as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "syslog-route", "uri": "/anything", "plugins": { "syslog": { "host": "127.0.0.1", "port": 514, "flush_limit": 1 } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: syslog-route uris: - /anything plugins: syslog: host: 127.0.0.1 port: 514 flush_limit: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD syslog-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: syslog-plugin-config spec: plugins: - name: syslog config: host: example-rsyslog-server.aic.svc.cluster.local port: 514 flush_limit: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: syslog-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: syslog-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` syslog-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: syslog-route spec: ingressClassName: apisix http: - name: syslog-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: syslog enable: true config: host: example-rsyslog-server.aic.svc.cluster.local port: 514 flush_limit: 1 ``` Apply the configuration: ``` kubectl apply -f syslog-ic.yaml ``` ❶ `host`: replace with the address of your syslog server. For Kubernetes, use the in-cluster Service address such as `example-rsyslog-server.aic.svc.cluster.local`. ❷ `port`: replace with the port of your syslog server. ❸ `flush_limit`: set to `1` to push logs to the syslog server immediately. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. In the syslog server, you should see a log entry similar to the following: ``` { "response": { "status": 200, "headers": { "access-control-allow-credentials": "true", "connection": "close", "date": "Fri, 17 Apr 2026 05:39:46 GMT", "access-control-allow-origin": "*", "server": "APISIX/3.16.0", "content-type": "application/json", "content-length": "387" }, "size": 614 }, "service_id": "", "client_ip": "172.19.0.1", "server": { "hostname": "eff61bf7be4d", "version": "3.16.0" }, "upstream": "35.171.123.176:80", "apisix_latency": 13.999900817871, "request": { "method": "GET", "url": "http://127.0.0.1:9080/anything", "querystring": {}, "size": 86, "uri": "/anything", "headers": { "host": "127.0.0.1:9080", "accept": "*/*", "user-agent": "curl/7.29.0" } }, "route_id": "syslog-route", "upstream_latency": 165, "latency": 178.99990081787, "start_time": 1709334859598 } ``` ### Customize Log Format With Plugin Metadata[​](#customize-log-format-with-plugin-metadata "Direct link to Customize Log Format With Plugin Metadata") The following example demonstrates how you can customize log format using [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md). The log format configured in plugin metadata will apply to all `syslog` plugin instances. Create a route with the `syslog` plugin: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "syslog-route", "uri": "/anything", "plugins": { "syslog": { "host": "127.0.0.1", "port": 514, "flush_limit": 1 } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: syslog-route uris: - /anything plugins: syslog: host: 127.0.0.1 port: 514 flush_limit: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD syslog-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: syslog-plugin-config spec: plugins: - name: syslog config: host: example-rsyslog-server.aic.svc.cluster.local port: 514 flush_limit: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: syslog-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: syslog-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` syslog-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: syslog-route spec: ingressClassName: apisix http: - name: syslog-route match: paths: - /anything methods: - GET upstreams: - name: httpbin-external-domain plugins: - name: syslog enable: true config: host: example-rsyslog-server.aic.svc.cluster.local port: 514 flush_limit: 1 ``` Apply the configuration: ``` kubectl apply -f syslog-ic.yaml ``` Configure plugin metadata for `syslog`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/plugin_metadata/syslog" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "log_format": { "host": "$host", "@timestamp": "$time_iso8601", "route_id": "$route_id", "client_ip": "$remote_addr", "resp_content_type": "$sent_http_Content_Type" } }' ``` adc.yaml ``` plugin_metadata: - name: syslog log_format: host: "$host" "@timestamp": "$time_iso8601" route_id: "$route_id" client_ip: "$remote_addr" resp_content_type: "$sent_http_Content_Type" ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: syslog: log_format: host: "$host" "@timestamp": "$time_iso8601" route_id: "$route_id" client_ip: "$remote_addr" resp_content_type: "$sent_http_Content_Type" ``` gatewayproxy.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: GatewayProxy metadata: namespace: aic name: apisix-config spec: provider: type: ControlPlane controlPlane: # ... # your control plane connection configuration pluginMetadata: syslog: log_format: host: "$host" "@timestamp": "$time_iso8601" route_id: "$route_id" client_ip: "$remote_addr" resp_content_type: "$sent_http_Content_Type" ``` Apply the configuration: ``` kubectl apply -f gatewayproxy.yaml ``` Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` In the syslog server, you should see a log entry similar to the following: ``` { "@timestamp": "2026-04-17T05:39:46+00:00", "resp_content_type": "application/json", "host": "127.0.0.1", "route_id": "syslog-route", "client_ip": "172.19.0.1" } ``` ### Log Request Bodies Conditionally[​](#log-request-bodies-conditionally "Direct link to Log Request Bodies Conditionally") The following example demonstrates how you can conditionally log request body. Create a route with the `syslog` plugin as follows: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "syslog-route", "uri": "/anything", "plugins": { "syslog": { "host": "127.0.0.1", "port": 514, "flush_limit": 1, "include_req_body": true, "include_req_body_expr": [["arg_log_body", "==", "yes"]] } }, "upstream": { "nodes": { "httpbin.org:80": 1 }, "type": "roundrobin" } }' ``` adc.yaml ``` services: - name: httpbin routes: - name: syslog-route uris: - /anything plugins: syslog: host: 127.0.0.1 port: 514 flush_limit: 1 include_req_body: true include_req_body_expr: - - arg_log_body - == - "yes" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD syslog-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: syslog-plugin-config spec: plugins: - name: syslog config: host: example-rsyslog-server.aic.svc.cluster.local port: 514 flush_limit: 1 include_req_body: true include_req_body_expr: - - arg_log_body - == - "yes" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: syslog-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: syslog-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` syslog-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: syslog-route spec: ingressClassName: apisix http: - name: syslog-route match: paths: - /anything methods: - POST upstreams: - name: httpbin-external-domain plugins: - name: syslog enable: true config: host: example-rsyslog-server.aic.svc.cluster.local port: 514 flush_limit: 1 include_req_body: true include_req_body_expr: - - arg_log_body - == - "yes" ``` Apply the configuration: ``` kubectl apply -f syslog-ic.yaml ``` ❶ `include_req_body`: set to `true` to include request body. ❷ `include_req_body_expr`: only include request body if the URL query string `log_body` is `yes`. In YAML-based configurations, quote `"yes"` to avoid YAML converting it to a boolean. Send a request to the route with a URL query string satisfying the condition: ``` curl -i "http://127.0.0.1:9080/anything?log_body=yes" -X POST \ -H "Content-Type: application/json" \ -d '{"env":"dev"}' ``` You should see the request body logged: ``` { "response": { "status": 200, "headers": { "connection": "close", "server": "APISIX/3.16.0", "date": "Fri, 17 Apr 2026 05:55:06 GMT", "access-control-allow-origin": "*", "access-control-allow-credentials": "true", "content-type": "application/json", "content-length": "531" }, "size": 759 }, "service_id": "", "client_ip": "172.19.0.1", "server": { "hostname": "eff61bf7be4d", "version": "3.16.0" }, "upstream": "35.171.123.176:80", "apisix_latency": 0, "request": { "method": "POST", "url": "http://127.0.0.1:9080/anything?log_body=yes", "querystring": { "log_body": "yes" }, "size": 164, "body": "{\"env\":\"dev\"}", "uri": "/anything?log_body=yes", "headers": { "accept": "*/*", "user-agent": "curl/7.29.0", "host": "127.0.0.1:9080", "content-type": "application/json", "content-length": "13" } }, "route_id": "syslog-route", "upstream_latency": 892, "latency": 1011.0001564026, "start_time": 1709340364390 } ``` Send a request to the route without any URL query string: ``` curl -i "http://127.0.0.1:9080/anything" -X POST \ -H "Content-Type: application/json" \ -d '{"env":"dev"}' ``` You should not observe the request body in the log. info If you have customized the `log_format` in addition to setting `include_req_body` or `include_resp_body` to `true`, the plugin would not include the bodies in the logs. As a workaround, you may be able to use the NGINX variable `$request_body` in the log format, such as: ``` { "syslog": { ..., "log_format": {"body": "$request_body"} } } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * host string required *** IP address or hostname of the syslog server. * port integer required *** Target port of the syslog server. * timeout integer default: `3000` vaild vaule: greater than 0 *** Timeout for the upstream to send data, in milliseconds. * tls boolean default: `false` *** If true, verify TLS. * flush\_limit integer default: `4096` vaild vaule: greater than 0 *** Maximum size of the buffer and the current message in KB before the logs are pushed to the syslog server. * drop\_limit integer default: `1048576` vaild vaule: greater than 0 *** Maximum size of the buffer and the current message allowed in KB before the logs are dropped. * sock\_type string default: `tcp` vaild vaule: `tcp` or `udp` *** Transport layer protocol to use. * pool\_size integer default: `5` vaild vaule: greater than or equal to 5 *** Keep-alive pool size used by `sock:keepalive`. * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. You can also configure log format on a global scale using the [plugin metadata](https://docs.api7.ai/apisix/key-concepts/plugin-metadata.md), which configures the log format for all `syslog` plugin instances. If the log format configured on the individual plugin instance differs from the log format configured on plugin metadata, the log format configured on the individual plugin instance takes precedence. See the [example](https://docs.api7.ai/hub/syslog.md#customize-log-format-with-plugin-metadata) for more details. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * include\_req\_body boolean default: `false` *** If true, include the request body in the log. Note that if the request body is too big to be kept in the memory, it can not be logged due to NGINX's limitations. * include\_req\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_req_body` is true. Request body would only be logged when the expressions configured here evaluate to true. * include\_resp\_body boolean default: `false` *** If true, include the response body in the log. * include\_resp\_body\_expr array\[array] *** An array of one or more conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). Used when the `include_resp_body` is true. Response body would only be logged when the expressions configured here evaluate to true. * max\_req\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum request body size in bytes to include in the log. If the request body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * max\_resp\_body\_bytes integer default: `524288` vaild vaule: greater than or equal to 1 *** Maximum response body size in bytes to include in the log. If the response body exceeds this value, it will be truncated. Available in APISIX from 3.16.0. * name string default: `sys logger` *** Unique identifier of the plugin for the batch processor. If you use [Prometheus](https://docs.api7.ai/hub/prometheus.md) to monitor APISIX metrics, the name is exported in `apisix_batch_process_entries`. * batch\_max\_size integer default: `1000` vaild vaule: greater than 0 *** The number of log entries allowed in one batch. Once reached, the batch will be sent to the logging service. Setting this parameter to 1 means immediate processing. * inactive\_timeout integer default: `5` vaild vaule: greater than 0 *** The maximum time in seconds to wait for new logs before sending the batch to the logging service. The value should be smaller than `buffer_duration`. * buffer\_duration integer default: `60` vaild vaule: greater than 0 *** The maximum time in seconds from the earliest entry allowed before sending the batch to the logging service. * retry\_delay integer default: `1` vaild vaule: greater than or equal to 0 *** The time interval in seconds to retry sending the batch to the logging service if the batch was not successfully sent. * max\_retry\_count integer default: `0` vaild vaule: greater than or equal to 0 *** The maximum number of unsuccessful retries allowed before dropping the log entries. ## Plugin Metadata[​](#plugin-metadata "Direct link to Plugin Metadata") * log\_format object *** Custom log format using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). In APISIX from 3.15.0, log format nested structures are supported up to five levels deep. In API7 Enterprise, only flat key-value structures are supported; nested structures are not yet supported. * log\_format\_extra object *** Additional fields to add to the default log entry, using key-value pairs in JSON format. Values can reference [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). A configured field does not overwrite an existing default field. A plugin instance takes precedence over plugin metadata; setting an empty object on the instance disables the metadata value. When `log_format` is configured, `log_format_extra` is ignored. Introduced in API7 Enterprise 3.9.15 and 3.10.2, and APISIX 3.18.0. * max\_pending\_entries integer default: `` `8192` in APISIX 3.18.0 and in API7 Enterprise 3.9.19 and 3.10.6 `` vaild vaule: greater than or equal to 1 *** Maximum number of entries waiting in the batch processor. New entries are discarded when the backlog reaches the limit. Introduced in API7 Enterprise 3.9.19 on the 3.9 line and 3.10.6 on the 3.10 line, and in APISIX 3.18.0. See [Batch Processor](https://docs.api7.ai/apisix/reference/batch-processor.md#configure-the-pending-entry-limit) for sizing and verification guidance. --- # traffic-label The `traffic-label` plugin evaluates request expressions and selects a weighted action from the first matching rule that produces an action. In this version, the supported action adds or replaces request headers before the request is forwarded upstream. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `traffic-label` on a route in different scenarios. ### Define a Single Matching Condition[​](#define-a-single-matching-condition "Direct link to Define a Single Matching Condition") The following example demonstrates a simple rule with one matching condition and one associated action. If the URI of the request is `/headers`, the plugin will add the header `"X-Server-Id": "100"` to the request. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "traffic-label-route", "uri":"/headers", "plugins":{ "traffic-label": { "rules": [ { "match": [ ["uri", "==", "/headers"] ], "actions": [ { "set_headers": { "X-Server-Id": 100 } } ] } ] } }, "upstream":{ "type":"roundrobin", "nodes":{ "httpbin.org:80":1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /headers name: traffic-label-route plugins: traffic-label: rules: - match: - - uri - "==" - /headers actions: - set_headers: X-Server-Id: 100 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD traffic-label-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: traffic-label-plugin-config spec: plugins: - name: traffic-label config: rules: - match: - - uri - "==" - /headers actions: - set_headers: X-Server-Id: 100 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-label-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: traffic-label-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` traffic-label-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-label-route spec: ingressClassName: apisix http: - name: traffic-label-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: traffic-label config: rules: - match: - - uri - "==" - /headers actions: - set_headers: X-Server-Id: 100 ``` Apply the configuration: ``` kubectl apply -f traffic-label-ic.yaml ``` Send a request to verify: ``` curl "http://127.0.0.1:9080/headers" ``` You should see a response similar to the following: ``` { "headers": { "Accept": "*/*", ... "X-Server-Id": "100" } } ``` ### Define Multiple Matching Conditions with Logical Operators[​](#define-multiple-matching-conditions-with-logical-operators "Direct link to Define Multiple Matching Conditions with Logical Operators") You can build more complex matching conditions with [logical operators](https://docs.api7.ai/apisix/reference/apisix-expressions.md#logical-operators). The following example demonstrates a rule with two matching conditions logically grouped by `OR` and one associated action. If one of the conditions is met, the plugin will add the header `"X-Server-Id": "100"` to the request. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "traffic-label-route", "uri":"/headers", "plugins":{ "traffic-label": { "rules": [ { "match": [ "OR", ["arg_version", "==", "v1"], ["arg_env", "==", "dev"] ], "actions": [ { "set_headers": { "X-Server-Id": 100 } } ] } ] } }, "upstream":{ "type":"roundrobin", "nodes":{ "httpbin.org:80":1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /headers name: traffic-label-route plugins: traffic-label: rules: - match: - OR - - arg_version - "==" - v1 - - arg_env - "==" - dev actions: - set_headers: X-Server-Id: 100 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD traffic-label-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: traffic-label-plugin-config spec: plugins: - name: traffic-label config: rules: - match: - OR - - arg_version - "==" - v1 - - arg_env - "==" - dev actions: - set_headers: X-Server-Id: 100 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-label-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: traffic-label-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` traffic-label-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-label-route spec: ingressClassName: apisix http: - name: traffic-label-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: traffic-label config: rules: - match: - OR - - arg_version - "==" - v1 - - arg_env - "==" - dev actions: - set_headers: X-Server-Id: 100 ``` Apply the configuration: ``` kubectl apply -f traffic-label-ic.yaml ``` Send a request to verify: ``` curl "http://127.0.0.1:9080/headers?env=dev" ``` You should see a response similar to the following: ``` { "headers": { "Accept": "*/*", ... "X-Server-Id": "100" } } ``` If you send a request that does not match any of the conditions, you will not see `"X-Server-Id": "100"` added to the request header. ### Create Weighted Actions[​](#create-weighted-actions "Direct link to Create Weighted Actions") The following example demonstrates a rule with one matching condition and multiple weighted actions, where incoming requests are distributed proportionally based on the weights. If a `weight` is not associated with any action, this portion of the requests will not have any action performed on them. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "traffic-label-route", "uri":"/headers", "plugins":{ "traffic-label": { "rules": [ { "match": [ ["uri", "==", "/headers"] ], "actions": [ { "set_headers": { "X-Server-Id": 100 }, "weight": 3 }, { "set_headers": { "X-API-Version": "v2" }, "weight": 2 }, { "weight": 5 } ] } ] } }, "upstream":{ "type":"roundrobin", "nodes":{ "httpbin.org:80":1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /headers name: traffic-label-route plugins: traffic-label: rules: - match: - - uri - "==" - /headers actions: - set_headers: X-Server-Id: 100 weight: 3 - set_headers: X-API-Version: v2 weight: 2 - weight: 5 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD traffic-label-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: traffic-label-plugin-config spec: plugins: - name: traffic-label config: rules: - match: - - uri - "==" - /headers actions: - set_headers: X-Server-Id: 100 weight: 3 - set_headers: X-API-Version: v2 weight: 2 - weight: 5 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-label-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: traffic-label-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` traffic-label-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-label-route spec: ingressClassName: apisix http: - name: traffic-label-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: traffic-label config: rules: - match: - - uri - "==" - /headers actions: - set_headers: X-Server-Id: 100 weight: 3 - set_headers: X-API-Version: v2 weight: 2 - weight: 5 ``` Apply the configuration: ``` kubectl apply -f traffic-label-ic.yaml ``` The proportion of times each action is executed is determined by the weight of the action relative to the total weight of all actions listed under the `actions` field. Here, the total weight is calculated as the sum of all action weights: 3 + 2 + 5 = 10. Therefore: ❶ 30% of the requests should have the `X-Server-Id: 100` request header. ❷ 20% of the requests should have the `X-API-Version: v2` request header. ❸ 50% of the requests should not have any action performed on them. Generate 50 consecutive requests to verify the weighted actions: ``` resp=$(seq 50 | xargs -I{} curl "http://127.0.0.1:9080/headers" -sL) && \ count_w3=$(echo "$resp" | grep -i "X-Server-Id" | wc -l) && \ count_w2=$(echo "$resp" | grep -i "X-API-Version" | wc -l) && \ echo X-Server-Id: $count_w3, X-API-Version: $count_w2 ``` The response shows that headers are added to requests in a weighted manner: ``` X-Server-Id: 15, X-API-Version: 10 ``` ### Define Multiple Matching Rules[​](#define-multiple-matching-rules "Direct link to Define Multiple Matching Rules") The following example demonstrates the use of multiple rules, each with their matching condition and action. * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "traffic-label-route", "uri":"/headers", "plugins":{ "traffic-label": { "rules": [ { "match": [ ["arg_version", "==", "v1"] ], "actions": [ { "set_headers": { "X-Server-Id": 100 } } ] }, { "match": [ ["arg_version", "==", "v2"] ], "actions": [ { "set_headers": { "X-Server-Id": 200 } } ] } ] } }, "upstream":{ "type":"roundrobin", "nodes":{ "httpbin.org:80":1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /headers name: traffic-label-route plugins: traffic-label: rules: - match: - - arg_version - "==" - v1 actions: - set_headers: X-Server-Id: 100 - match: - - arg_version - "==" - v2 actions: - set_headers: X-Server-Id: 200 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD traffic-label-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: traffic-label-plugin-config spec: plugins: - name: traffic-label config: rules: - match: - - arg_version - "==" - v1 actions: - set_headers: X-Server-Id: 100 - match: - - arg_version - "==" - v2 actions: - set_headers: X-Server-Id: 200 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-label-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: traffic-label-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` traffic-label-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-label-route spec: ingressClassName: apisix http: - name: traffic-label-route match: paths: - /headers upstreams: - name: httpbin-external-domain plugins: - name: traffic-label config: rules: - match: - - arg_version - "==" - v1 actions: - set_headers: X-Server-Id: 100 - match: - - arg_version - "==" - v2 actions: - set_headers: X-Server-Id: 200 ``` Apply the configuration: ``` kubectl apply -f traffic-label-ic.yaml ``` Send a request to `/headers?version=v1` to verify: ``` curl "http://127.0.0.1:9080/headers?version=v1" ``` You should see a response similar to the following: ``` { "headers": { "Accept": "*/*", ... "X-Server-Id": "100" } } ``` Send a request to `/headers?version=v2` to verify: ``` curl "http://127.0.0.1:9080/headers?version=v2" ``` You should see a response similar to the following: ``` { "headers": { "Accept": "*/*", ... "X-Server-Id": "200" } } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * rules array\[object] required *** An array of one or more pairs of matching conditions and actions to be executed. Rules are evaluated sequentially. For a matching rule, the plugin selects one action by weight. When the selected action contains `set_headers`, it is applied and evaluation stops. If the selected object contains only `weight`, evaluation continues to the next rule. * match array\[array] required *** An array of one or more matching conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). * actions array\[object] required *** An array of one or more actions to be executed when a condition is successfully matched. * set\_headers object *** One or more request headers to apply to requests in the format of `{"name": "value", ...}`, where `value` could be a [built-in variable](https://docs.api7.ai/apisix/reference/built-in-variables.md). If a header of the same name already exists, it will be overwritten. * weight integer default: `1` *** The weight of action distribution. --- # traffic-split The `traffic-split` plugin directs traffic to various upstream services based on conditions and/or weights. It provides a dynamic and flexible approach to implement release strategies and manage traffic. ## Examples[​](#examples "Direct link to Examples") The examples below shows different use cases for using the `traffic-split` plugin. ### Implement Canary Release[​](#implement-canary-release "Direct link to Implement Canary Release") The following example demonstrates how to implement canary release with this plugin. Canary release is a gradual deployment in which an increasing percentage of traffic is directed to a new release, allowing for a controlled and monitored rollout. This method ensures that any potential issues or bugs in the new release can be identified and addressed early on, before fully redirecting all traffic. Create a route and configure `traffic-split` plugin with the following rules: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "uri": "/headers", "id": "traffic-split-route", "plugins": { "traffic-split": { "rules": [ { "weighted_upstreams": [ { "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "httpbin.org:443":1 } }, "weight": 3 }, { "weight": 2 } ] } ] } }, "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "mock.api7.ai:443":1 } } }' ``` adc.yaml ``` services: - name: traffic-split-service routes: - uris: - /headers name: traffic-split-route plugins: traffic-split: rules: - weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 3 - weight: 2 upstream: type: roundrobin scheme: https pass_host: node nodes: - host: mock.api7.ai port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD traffic-split-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: v1 kind: Service metadata: namespace: aic name: mockapi7-external-domain spec: type: ExternalName externalName: mock.api7.ai --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: traffic-split-plugin-config spec: plugins: - name: traffic-split config: rules: - weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 3 - weight: 2 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-split-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: traffic-split-plugin-config backendRefs: - name: mockapi7-external-domain port: 443 ``` traffic-split-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: httpbin.org port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: mockapi7-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: mock.api7.ai port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-split-route spec: ingressClassName: apisix http: - name: traffic-split-route match: paths: - /headers upstreams: - name: mockapi7-external-domain plugins: - name: traffic-split enable: true config: rules: - weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 3 - weight: 2 ``` Apply the configuration to your cluster: ``` kubectl apply -f traffic-split-ic.yaml ``` The proportion of traffic to each upstream is determined by the weight of the upstream relative to the total weight of all upstreams. Here, the total weight is calculated as: 3 + 2 = 5. Therefore: ❶ 60% of the traffic are expected to be forwarded to `httpbin.org`. ❷ 40% of the traffic are expected to be forwarded to `mock.api7.ai`. Send 10 consecutive requests to the route to verify: ``` resp=$(seq 10 | xargs -I{} curl "http://127.0.0.1:9080/headers" -sL) && \ count_httpbin=$(echo "$resp" | grep "httpbin.org" | wc -l) && \ count_mockapi7=$(echo "$resp" | grep "mock.api7.ai" | wc -l) && \ echo httpbin.org: $count_httpbin, mock.api7.ai: $count_mockapi7 ``` You should see a response similar to the following: ``` httpbin.org: 6, mock.api7.ai: 4 ``` Adjust the upstream weights accordingly to complete the canary release. ### Implement Blue-Green Deployment[​](#implement-blue-green-deployment "Direct link to Implement Blue-Green Deployment") The following example demonstrates how to implement blue-green deployment with this plugin. Blue-green deployment is a deployment strategy that involves maintaining two identical environments: the *blue* and the *green*. The blue environment refers to the current production deployment and the green environment refers to the new deployment. Once the green environment is tested to be ready for production, traffic will be routed to the green environment, making it the new production deployment. Create a route and configure `traffic-split` plugin with the following rules: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "uri": "/headers", "id": "traffic-split-route", "plugins": { "traffic-split": { "rules": [ { "match": [ { "vars": [ ["http_release","==","new_release"] ] } ], "weighted_upstreams": [ { "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "httpbin.org:443":1 } } } ] } ] } }, "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "mock.api7.ai:443":1 } } }' ``` adc.yaml ``` services: - name: traffic-split-service routes: - uris: - /headers name: traffic-split-route plugins: traffic-split: rules: - match: - vars: - ["http_release", "==", "new_release"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 upstream: type: roundrobin scheme: https pass_host: node nodes: - host: mock.api7.ai port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD traffic-split-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: v1 kind: Service metadata: namespace: aic name: mockapi7-external-domain spec: type: ExternalName externalName: mock.api7.ai --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: traffic-split-plugin-config spec: plugins: - name: traffic-split config: rules: - match: - vars: - ["http_release", "==", "new_release"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-split-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: traffic-split-plugin-config backendRefs: - name: mockapi7-external-domain port: 443 ``` traffic-split-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: httpbin.org port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: mockapi7-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: mock.api7.ai port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-split-route spec: ingressClassName: apisix http: - name: traffic-split-route match: paths: - /headers upstreams: - name: mockapi7-external-domain plugins: - name: traffic-split enable: true config: rules: - match: - vars: - ["http_release", "==", "new_release"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 ``` Apply the configuration to your cluster: ``` kubectl apply -f traffic-split-ic.yaml ``` ❶ Execute the plugin to redirect traffic only when the request contains a header `release: new_release`. Send a request to the route with the `release` header: ``` curl "http://127.0.0.1:9080/headers" -H 'release: new_release' ``` You should see a response similar to the following: ``` { "headers": { "Accept": "*/*", "Host": "httpbin.org", ... } } ``` Send a request to the route without any additional header: ``` curl "http://127.0.0.1:9080/headers" ``` You should see a response similar to the following: ``` { "headers": { "accept": "*/*", "host": "mock.api7.ai", ... } } ``` ### Define Matching Condition for POST Request With APISIX Expressions[​](#define-matching-condition-for-post-request-with-apisix-expressions "Direct link to Define Matching Condition for POST Request With APISIX Expressions") The following example demonstrates how to use [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md) in rules to conditionally execute the plugin when certain condition of a POST request is satisfied. Create a route and configure `traffic-split` plugin with the following rules: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "uri": "/post", "methods": ["POST"], "id": "traffic-split-route", "plugins": { "traffic-split": { "rules": [ { "match": [ { "vars": [ ["post_arg_id", "==", "1"] ] } ], "weighted_upstreams": [ { "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "httpbin.org:443":1 } } } ] } ] } }, "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "mock.api7.ai:443":1 } } }' ``` adc.yaml ``` services: - name: traffic-split-service routes: - uris: - /post methods: - POST name: traffic-split-route plugins: traffic-split: rules: - match: - vars: - ["post_arg_id", "==", "1"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 upstream: type: roundrobin scheme: https pass_host: node nodes: - host: mock.api7.ai port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD traffic-split-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: v1 kind: Service metadata: namespace: aic name: mockapi7-external-domain spec: type: ExternalName externalName: mock.api7.ai --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: traffic-split-plugin-config spec: plugins: - name: traffic-split config: rules: - match: - vars: - ["post_arg_id", "==", "1"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-split-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /post method: POST filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: traffic-split-plugin-config backendRefs: - name: mockapi7-external-domain port: 443 ``` traffic-split-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: httpbin.org port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: mockapi7-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: mock.api7.ai port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-split-route spec: ingressClassName: apisix http: - name: traffic-split-route match: paths: - /post methods: - POST upstreams: - name: mockapi7-external-domain plugins: - name: traffic-split enable: true config: rules: - match: - vars: - ["post_arg_id", "==", "1"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 ``` Apply the configuration to your cluster: ``` kubectl apply -f traffic-split-ic.yaml ``` Send a POST request with body `id=1`: ``` curl "http://127.0.0.1:9080/post" -X POST \ -H 'Content-Type: application/x-www-form-urlencoded' \ -d 'id=1' ``` ❶ You can specify charset in the `Content-Type` as well, such as `Content-Type: application/x-www-form-urlencoded;charset=UTF-8`. You should see a response similar to the following: ``` { "args": {}, "data": "", "files": {}, "form": { "id": "1" }, "headers": { "Accept": "*/*", "Content-Length": "4", "Content-Type": "application/x-www-form-urlencoded", "Host": "httpbin.org", ... }, ... } ``` Send a POST request without `id=1` in the body: ``` curl "http://127.0.0.1:9080/post" -X POST \ -H 'Content-Type: application/x-www-form-urlencoded' \ -d 'random=string' ``` You should see that the request was forwarded to `mock.api7.ai`. ### Define AND Matching Conditions With APISIX Expressions[​](#define-and-matching-conditions-with-apisix-expressions "Direct link to Define AND Matching Conditions With APISIX Expressions") The following example demonstrates how to use [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md) in rules to conditionally execute the plugin when multiple conditions are satisfied. Create a route and configure `traffic-split` plugin with the following matching rules: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "uri": "/headers", "id": "traffic-split-route", "plugins": { "traffic-split": { "rules": [ { "match": [ { "vars": [ ["arg_name","==","jack"], ["http_user-id",">","23"], ["http_apisix-key","~~","[a-z]+"] ] } ], "weighted_upstreams": [ { "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "httpbin.org:443":1 } }, "weight": 3 }, { "weight": 2 } ] } ] } }, "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "mock.api7.ai:443":1 } } }' ``` adc.yaml ``` services: - name: traffic-split-service routes: - uris: - /headers name: traffic-split-route plugins: traffic-split: rules: - match: - vars: - ["arg_name", "==", "jack"] - ["http_user-id", ">", "23"] - ["http_apisix-key", "~~", "[a-z]+"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 3 - weight: 2 upstream: type: roundrobin scheme: https pass_host: node nodes: - host: mock.api7.ai port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD traffic-split-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: v1 kind: Service metadata: namespace: aic name: mockapi7-external-domain spec: type: ExternalName externalName: mock.api7.ai --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: traffic-split-plugin-config spec: plugins: - name: traffic-split config: rules: - match: - vars: - ["arg_name", "==", "jack"] - ["http_user-id", ">", "23"] - ["http_apisix-key", "~~", "[a-z]+"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 3 - weight: 2 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-split-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: traffic-split-plugin-config backendRefs: - name: mockapi7-external-domain port: 443 ``` traffic-split-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: httpbin.org port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: mockapi7-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: mock.api7.ai port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-split-route spec: ingressClassName: apisix http: - name: traffic-split-route match: paths: - /headers upstreams: - name: mockapi7-external-domain plugins: - name: traffic-split enable: true config: rules: - match: - vars: - ["arg_name", "==", "jack"] - ["http_user-id", ">", "23"] - ["http_apisix-key", "~~", "[a-z]+"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 3 - weight: 2 ``` Apply the configuration to your cluster: ``` kubectl apply -f traffic-split-ic.yaml ``` ❶ Execute the plugin to redirect traffic only when all three conditions are satisfied. If conditions are satisfied, 60% of the traffic should be directed to `httpbin.org` and the other 40% should be directed to `mock.api7.ai`. If conditions are not satisfied, all traffic should be directed to `mock.api7.ai`. Send 10 consecutive requests that satisfy all conditions to verify: ``` resp=$(seq 10 | xargs -I{} curl "http://127.0.0.1:9080/headers?name=jack" -H 'user-id: 30' -H 'apisix-key: helloapisix' -sL) && \ count_httpbin=$(echo "$resp" | grep "httpbin.org" | wc -l) && \ count_mockapi7=$(echo "$resp" | grep "mock.api7.ai" | wc -l) && \ echo httpbin.org: $count_httpbin, mock.api7.ai: $count_mockapi7 ``` You should see a response similar to the following: ``` httpbin.org: 6, mock.api7.ai: 4 ``` Send 10 consecutive requests that do not satisfy the conditions to verify: ``` resp=$(seq 10 | xargs -I{} curl "http://127.0.0.1:9080/headers?name=random" -sL) && \ count_httpbin=$(echo "$resp" | grep "httpbin.org" | wc -l) && \ count_mockapi7=$(echo "$resp" | grep "mock.api7.ai" | wc -l) && \ echo httpbin.org: $count_httpbin, mock.api7.ai: $count_mockapi7 ``` You should see a response similar to the following: ``` httpbin.org: 0, mock.api7.ai: 10 ``` ### Define OR Matching Conditions With APISIX Expressions[​](#define-or-matching-conditions-with-apisix-expressions "Direct link to Define OR Matching Conditions With APISIX Expressions") The following example demonstrates how to use [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md) in rules to conditionally execute the plugin when either set of the condition is satisfied. Create a route and configure `traffic-split` plugin with the following matching rules: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "uri": "/headers", "id": "traffic-split-route", "plugins": { "traffic-split": { "rules": [ { "match": [ { "vars": [ ["arg_name","==","jack"], ["http_user-id",">","23"], ["http_apisix-key","~~","[a-z]+"] ] }, { "vars": [ ["arg_name2","==","rose"], ["http_user-id2","!",">","33"], ["http_apisix-key2","~~","[a-z]+"] ] } ], "weighted_upstreams": [ { "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "httpbin.org:443":1 } }, "weight": 3 }, { "weight": 2 } ] } ] } }, "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "mock.api7.ai:443":1 } } }' ``` adc.yaml ``` services: - name: traffic-split-service routes: - uris: - /headers name: traffic-split-route plugins: traffic-split: rules: - match: - vars: - ["arg_name", "==", "jack"] - ["http_user-id", ">", "23"] - ["http_apisix-key", "~~", "[a-z]+"] - vars: - ["arg_name2", "==", "rose"] - ["http_user-id2", "!", ">", "33"] - ["http_apisix-key2", "~~", "[a-z]+"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 3 - weight: 2 upstream: type: roundrobin scheme: https pass_host: node nodes: - host: mock.api7.ai port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD traffic-split-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: v1 kind: Service metadata: namespace: aic name: mockapi7-external-domain spec: type: ExternalName externalName: mock.api7.ai --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: traffic-split-plugin-config spec: plugins: - name: traffic-split config: rules: - match: - vars: - ["arg_name", "==", "jack"] - ["http_user-id", ">", "23"] - ["http_apisix-key", "~~", "[a-z]+"] - vars: - ["arg_name2", "==", "rose"] - ["http_user-id2", "!", ">", "33"] - ["http_apisix-key2", "~~", "[a-z]+"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 3 - weight: 2 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-split-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: traffic-split-plugin-config backendRefs: - name: mockapi7-external-domain port: 443 ``` traffic-split-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: httpbin.org port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: mockapi7-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: mock.api7.ai port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-split-route spec: ingressClassName: apisix http: - name: traffic-split-route match: paths: - /headers upstreams: - name: mockapi7-external-domain plugins: - name: traffic-split enable: true config: rules: - match: - vars: - ["arg_name", "==", "jack"] - ["http_user-id", ">", "23"] - ["http_apisix-key", "~~", "[a-z]+"] - vars: - ["arg_name2", "==", "rose"] - ["http_user-id2", "!", ">", "33"] - ["http_apisix-key2", "~~", "[a-z]+"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 3 - weight: 2 ``` Apply the configuration to your cluster: ``` kubectl apply -f traffic-split-ic.yaml ``` ❶ and ❷: Execute the plugin to redirect traffic when either set of the conditions are satisfied. Alternatively, you can also use the OR operator in the [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md#logical-operators) for these conditions. If conditions are satisfied, 60% of the traffic should be directed to `httpbin.org` and the other 40% should be directed to `mock.api7.ai`. If conditions are not satisfied, all traffic should be directed to `mock.api7.ai`. Send 10 consecutive requests that satisfy the second set of conditions to verify: ``` resp=$(seq 10 | xargs -I{} curl "http://127.0.0.1:9080/headers?name2=rose" -H 'user-id:30' -H 'apisix-key2: helloapisix' -sL) && \ count_httpbin=$(echo "$resp" | grep "httpbin.org" | wc -l) && \ count_mockapi7=$(echo "$resp" | grep "mock.api7.ai" | wc -l) && \ echo httpbin.org: $count_httpbin, mock.api7.ai: $count_mockapi7 ``` You should see a response similar to the following: ``` httpbin.org: 6, mock.api7.ai: 4 ``` Send 10 consecutive requests that do not satisfy any set of conditions to verify: ``` resp=$(seq 10 | xargs -I{} curl "http://127.0.0.1:9080/headers?name=random" -sL) && \ count_httpbin=$(echo "$resp" | grep "httpbin.org" | wc -l) && \ count_mockapi7=$(echo "$resp" | grep "mock.api7.ai" | wc -l) && \ echo httpbin.org: $count_httpbin, mock.api7.ai: $count_mockapi7 ``` You should see a response similar to the following: ``` httpbin.org: 0, mock.api7.ai: 10 ``` ### Configure Different Rules for Different Upstreams[​](#configure-different-rules-for-different-upstreams "Direct link to Configure Different Rules for Different Upstreams") The following example demonstrates how to set one-to-one mapping between rule sets and upstreams. Create a route and configure `traffic-split` plugin with the following matching rules: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "uri": "/headers", "id": "traffic-split-route", "plugins": { "traffic-split": { "rules": [ { "match": [ { "vars": [ ["http_x-api-id","==","1"] ] } ], "weighted_upstreams": [ { "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "httpbin.org:443":1 } }, "weight": 1 } ] }, { "match": [ { "vars": [ ["http_x-api-id","==","2"] ] } ], "weighted_upstreams": [ { "upstream": { "type": "roundrobin", "scheme": "https", "pass_host": "node", "nodes": { "mock.api7.ai:443":1 } }, "weight": 1 } ] } ] } }, "upstream": { "type": "roundrobin", "nodes": { "postman-echo.com:443": 1 }, "scheme": "https", "pass_host": "node" } }' ``` adc.yaml ``` services: - name: traffic-split-service routes: - uris: - /headers name: traffic-split-route plugins: traffic-split: rules: - match: - vars: - ["http_x-api-id", "==", "1"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 1 - match: - vars: - ["http_x-api-id", "==", "2"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: mock.api7.ai port: 443 weight: 1 weight: 1 upstream: type: roundrobin scheme: https pass_host: node nodes: - host: postman-echo.com port: 443 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD traffic-split-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: v1 kind: Service metadata: namespace: aic name: mockapi7-external-domain spec: type: ExternalName externalName: mock.api7.ai --- apiVersion: v1 kind: Service metadata: namespace: aic name: postman-echo-external-domain spec: type: ExternalName externalName: postman-echo.com --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: traffic-split-plugin-config spec: plugins: - name: traffic-split config: rules: - match: - vars: - ["http_x-api-id", "==", "1"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 1 - match: - vars: - ["http_x-api-id", "==", "2"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: mock.api7.ai port: 443 weight: 1 weight: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: traffic-split-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /headers filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: traffic-split-plugin-config backendRefs: - name: postman-echo-external-domain port: 443 ``` traffic-split-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: httpbin.org port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: mockapi7-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: mock.api7.ai port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: postman-echo-external-domain spec: ingressClassName: apisix scheme: https passHost: node externalNodes: - type: Domain name: postman-echo.com port: 443 --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: traffic-split-route spec: ingressClassName: apisix http: - name: traffic-split-route match: paths: - /headers upstreams: - name: postman-echo-external-domain plugins: - name: traffic-split enable: true config: rules: - match: - vars: - ["http_x-api-id", "==", "1"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: httpbin.org port: 443 weight: 1 weight: 1 - match: - vars: - ["http_x-api-id", "==", "2"] weighted_upstreams: - upstream: type: roundrobin scheme: https pass_host: node nodes: - host: mock.api7.ai port: 443 weight: 1 weight: 1 ``` Apply the configuration to your cluster: ``` kubectl apply -f traffic-split-ic.yaml ``` ❶ Execute the plugin to redirect traffic only when the request contains a header `x-api-id: 1`. ❷ Execute the plugin to redirect traffic only when the request contains a header `x-api-id: 2`. Send a request with header `x-api-id: 1`: ``` curl "http://127.0.0.1:9080/headers" -H 'x-api-id: 1' ``` You should see an `HTTP/1.1 200 OK` response similar to the following: ``` { "headers": { "Accept": "*/*", "Host": "httpbin.org", ... } } ``` Send a request with header `x-api-id: 2`: ``` curl "http://127.0.0.1:9080/headers" -H 'x-api-id: 2' ``` You should see an `HTTP/1.1 200 OK` response similar to the following: ``` { "headers": { "accept": "*/*", "host": "mock.api7.ai", ... } } ``` Send a request without any additional header: ``` curl "http://127.0.0.1:9080/headers" ``` You should see a response similar to the following: ``` { "headers": { "accept": "*/*", "host": "postman-echo.com", ... } } ``` --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * rules array\[object] *** An array of one or more pairs of matching conditions and actions to be executed. * match array\[object] *** Rules to match for conditional traffic split. * vars array\[array] *** An array of one or more matching conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md) to conditionally execute the plugin. * weighted\_upstreams array\[object] *** List of upstream configurations. * upstream\_id string or integer *** ID of the configured upstream object. * weight integer default: `1` *** Weight for each upstream. * upstream object *** Configuration of the upstream. Certain configuration options of [upstream](https://docs.api7.ai/apisix/reference/admin-api/.md#tag/Upstream/paths/~1apisix~1admin~1upstreams/post) are not supported here. These fields are `service_name`, `discovery_type`, `checks`, `retries`, `retry_timeout`, `desc`, and `labels`. As a workaround, you can create an upstream object and configure it in `upstream_id`. * type string default: `roundrobin` vaild vaule: `roundrobin`, `chash`, `ewma`, or `least_conn` *** Algorithm for traffic splitting. Use `roundrobin` for weighted round robin, `chash` for consistent hashing, `ewma` for exponential weighted moving average, and `least_conn` for least connections. * hash\_on string default: `vars` vaild vaule: `vars`, `header`, `cookie`, `consumer`, or `vars_combinations` *** Used when `type` is `chash`. Support hashing on [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md), header, cookie, consumer, or a combination of [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md). * key string *** Used when `type` is `chash`. When `hash_on` is set to `header` or `cookie`, `key` is required. When `hash_on` is set to `consumer`, `key` is not required as the consumer name will be used as the key automatically. * nodes object *** Addresses of the upstream nodes. * timeout object *** Timeout in seconds for connecting, sending and receiving messages. * pass\_host string default: `pass` vaild vaule: `pass`, `node`, or `rewrite` *** Mode deciding how the host name is passed. `pass` passes the client's host name to the upstream. `node` passes the host configured in the node of the upstream. `rewrite` passes the value configured in `upstream_host`. * upstream\_host string *** Used when `pass_host` is `rewrite`. Host name of the upstream. * name string *** Identifier for the upstream for specifying service name, usage scenarios, and so on. --- # ua-restriction The `ua-restriction` plugin supports restricting access to upstream resources through either configuring an allowlist or denylist of user agents. A common use case is to prevent web crawlers from overloading the upstream resources and causing service degradation. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrate how you can configure `ua-restriction` for different scenarios. ### Reject Web Crawlers and Customize Error Message[​](#reject-web-crawlers-and-customize-error-message "Direct link to Reject Web Crawlers and Customize Error Message") The following example demonstrates how you can configure the plugin to fend off unwanted web crawlers and customize the rejection message. * Admin API * ADC * Ingress Controller Create a route as such: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "ua-restriction-route", "uri": "/anything", "plugins": { "ua-restriction": { "bypass_missing": false, "denylist": [ "(Baiduspider)/(\\d+)\\.(\\d+)", "bad-bot-1" ], "message": "Access denied" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: ua-restriction-service routes: - name: ua-restriction-route uris: - /anything plugins: ua-restriction: bypass_missing: false denylist: - "(Baiduspider)/(\\d+)\\.(\\d+)" - "bad-bot-1" message: "Access denied" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD ua-restriction-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: ua-restriction-plugin-config spec: plugins: - name: ua-restriction config: bypass_missing: false denylist: - "(Baiduspider)/(\\d+)\\.(\\d+)" - "bad-bot-1" message: "Access denied" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: ua-restriction-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: ua-restriction-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f ua-restriction-ic.yaml ``` ua-restriction-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: ua-restriction-route spec: ingressClassName: apisix http: - name: ua-restriction-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: ua-restriction enable: true config: bypass_missing: false denylist: - "(Baiduspider)/(\\d+)\\.(\\d+)" - "bad-bot-1" message: "Access denied" ``` Apply the configuration to your cluster: ``` kubectl apply -f ua-restriction-ic.yaml ``` ❶ Do not allow bypassing UA restriction rules. ❷ Configure user agents that should not be able to access the upstream resource. ❸ Customize the error message for when the access is denied. Send a request to the route: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Send another request to the route with a disallowed user agent: ``` curl -i "http://127.0.0.1:9080/anything" -H 'User-Agent: Baiduspider/5.0' ``` You should receive an `HTTP/1.1 403 Forbidden` response with the following message: ``` {"message":"Access denied"} ``` ### Bypass UA Restriction Checks[​](#bypass-ua-restriction-checks "Direct link to Bypass UA Restriction Checks") The following example demonstrates how to configure the plugin to allow requests of a specific user agent to bypass the UA restriction. * Admin API * ADC * Ingress Controller Create a route as such: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "ua-restriction-route", "uri": "/anything", "plugins": { "ua-restriction": { "bypass_missing": true, "allowlist": [ "good-bot-1" ], "message": "Access denied" } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org:80": 1 } } }' ``` adc.yaml ``` services: - name: ua-restriction-service routes: - name: ua-restriction-route uris: - /anything plugins: ua-restriction: bypass_missing: true allowlist: - "good-bot-1" message: "Access denied" upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD ua-restriction-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: ua-restriction-allowlist-plugin-config spec: plugins: - name: ua-restriction config: bypass_missing: true allowlist: - "good-bot-1" message: "Access denied" --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: ua-restriction-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: ua-restriction-allowlist-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Apply the configuration to your cluster: ``` kubectl apply -f ua-restriction-ic.yaml ``` ua-restriction-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: ua-restriction-route spec: ingressClassName: apisix http: - name: ua-restriction-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: ua-restriction enable: true config: bypass_missing: true allowlist: - "good-bot-1" message: "Access denied" ``` Apply the configuration to your cluster: ``` kubectl apply -f ua-restriction-ic.yaml ``` ❶ Allow bypassing UA restriction rules. ❷ Configure user agents that should be allowed to access the upstream resource. Send a request to the route without modifying the user agent: ``` curl -i "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 403 Forbidden` response with the following message: ``` {"message":"Access denied"} ``` Send another request to the route with an empty user agent: ``` curl -i "http://127.0.0.1:9080/anything" -H 'User-Agent: ' ``` You should receive an `HTTP/1.1 200 OK` response. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * bypass\_missing boolean default: `false` *** If true, bypass the user agent restriction check when the `User-Agent` header is missing. * allowlist array\[string] *** List of user agents to allow. Support regular expressions. At least one of the `allowlist` and `denylist` should be configured, but they cannot be configured at the same time. * denylist array\[string] *** List of user agents to deny. Support regular expressions. At least one of the `allowlist` and `denylist` should be configured, but they cannot be configured at the same time. * message string default: `Not allowed` *** Message returned when the user agent is denied access. --- # workflow The `workflow` plugin supports the conditional execution of user-defined actions to client traffic based on a given set of rules, defined using [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). This provides a granular approach to traffic management. If you would like to apply more complex matching conditions and actions, see the [`traffic-label`](https://docs.api7.ai/hub/traffic-label.md) plugin. ## Examples[​](#examples "Direct link to Examples") The examples below demonstrates how you can use the `workflow` plugin for different scenarios. ### Return Response HTTP Status Code Conditionally[​](#return-response-http-status-code-conditionally "Direct link to Return Response HTTP Status Code Conditionally") The following example demonstrates a simple rule with one matching condition and one associated action to return HTTP status code conditionally. Create a route with the `workflow` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "workflow-route", "uri": "/anything/*", "plugins": { "workflow":{ "rules":[ { "case":[ ["uri", "==", "/anything/rejected"] ], "actions":[ [ "return", {"code": 403} ] ] } ] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything/* name: workflow-route plugins: workflow: rules: - case: - ["uri", "==", "/anything/rejected"] actions: - - return - code: 403 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD workflow-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: workflow-plugin-config spec: plugins: - name: workflow config: rules: - case: - ["uri", "==", "/anything/rejected"] actions: - - return - code: 403 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: workflow-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything/ filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: workflow-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` workflow-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: workflow-route spec: ingressClassName: apisix http: - name: workflow-route match: paths: - /anything/* upstreams: - name: httpbin-external-domain plugins: - name: workflow enable: true config: rules: - case: - ["uri", "==", "/anything/rejected"] actions: - - return - code: 403 ``` Apply the configuration to your cluster: ``` kubectl apply -f workflow-ic.yaml ``` ❶ Trigger the action only when the request's URI path is `/anything/rejected`. ❷ Return HTTP status code 403 when the rule is matched. Send a request that matches none of the rules: ``` curl -i "http://127.0.0.1:9080/anything/anything" ``` You should receive an `HTTP/1.1 200 OK` response. Send a request that matches the configured rule: ``` curl -i "http://127.0.0.1:9080/anything/rejected" ``` You should receive an `HTTP/1.1 403 Forbidden` response of following: ``` {"error_msg":"rejected by workflow"} ``` ### Apply Rate Limiting Conditionally by URI and Query Parameter[​](#apply-rate-limiting-conditionally-by-uri-and-query-parameter "Direct link to Apply Rate Limiting Conditionally by URI and Query Parameter") The following example demonstrates a rule with two matching conditions and one associated action to rate limit requests conditionally. Create a route with the `workflow` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "workflow-route", "uri": "/anything/*", "plugins":{ "workflow":{ "rules":[ { "case":[ ["uri", "==", "/anything/rate-limit"], ["arg_env", "==", "v1"] ], "actions":[ [ "limit-count", { "count":1, "time_window":60, "rejected_code":429 } ] ] } ] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything/* name: workflow-route plugins: workflow: rules: - case: - ["uri", "==", "/anything/rate-limit"] - ["arg_env", "==", "v1"] actions: - - limit-count - count: 1 time_window: 60 rejected_code: 429 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD workflow-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: workflow-plugin-config spec: plugins: - name: workflow config: rules: - case: - ["uri", "==", "/anything/rate-limit"] - ["arg_env", "==", "v1"] actions: - - limit-count - count: 1 time_window: 60 rejected_code: 429 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: workflow-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything/ filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: workflow-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` workflow-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: workflow-route spec: ingressClassName: apisix http: - name: workflow-route match: paths: - /anything/* upstreams: - name: httpbin-external-domain plugins: - name: workflow enable: true config: rules: - case: - ["uri", "==", "/anything/rate-limit"] - ["arg_env", "==", "v1"] actions: - - limit-count - count: 1 time_window: 60 rejected_code: 429 ``` Apply the configuration to your cluster: ``` kubectl apply -f workflow-ic.yaml ``` ❶ Match URI path `/anything/rate-limit`. ❷ Match query parameter `env` whose value being `v1`. See [built-in variables](https://docs.api7.ai/apisix/reference/built-in-variables.md) for more variables to help construct conditions. ❸ Apply rate limiting when both of the conditions are matched. Generate two consecutive requests that matches the second rule: ``` curl -i "http://127.0.0.1:9080/anything/rate-limit?env=v1" ``` You should receive an `HTTP/1.1 200 OK` response and an `HTTP 429 Too Many Requests` response. Generate requests that do not match the condition: ``` curl -i "http://127.0.0.1:9080/anything/anything?env=v1" ``` You should receive `HTTP/1.1 200 OK` responses for all requests, as they are not rate limited. ### Apply Rate Limiting Conditionally by Consumers[​](#apply-rate-limiting-conditionally-by-consumers "Direct link to Apply Rate Limiting Conditionally by Consumers") The following example demonstrates how to configure the plugin to perform rate limiting based on the following specifications: * consumer `john` should have a quota of 5 requests within a 30-second window * consumer `jane` should have a quota of 3 requests within a 30-second window * all other consumers should have a quota of 2 requests within a 30-second window While this example will be using [`key-auth`](https://docs.api7.ai/hub/key-auth.md), you can easily replace it with other authentication plugins. * Admin API * ADC * Ingress Controller Create a consumer `john`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "john" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/john/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-john-key-auth", "plugins": { "key-auth": { "key": "john-key" } } }' ``` Create a second consumer `jane`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jane" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jane/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jane-key-auth", "plugins": { "key-auth": { "key": "jane-key" } } }' ``` Create a third consumer `jimmy`: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "username": "jimmy" }' ``` Create `key-auth` credential for the consumer: ``` curl "http://127.0.0.1:9180/apisix/admin/consumers/jimmy/credentials" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "cred-jimmy-key-auth", "plugins": { "key-auth": { "key": "jimmy-key" } } }' ``` Create a route with the `workflow` plugin as such: ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "workflow-route", "uri": "/anything", "plugins":{ "key-auth": {}, "workflow":{ "rules":[ { "actions": [ [ "limit-count", { "count": 5, "key": "consumer_john", "key_type": "constant", "rejected_code": 429, "time_window": 30, "policy": "local" } ] ], "case": [ [ "consumer_name", "==", "john" ] ] }, { "actions": [ [ "limit-count", { "count": 3, "key": "consumer_jane", "key_type": "constant", "rejected_code": 429, "time_window": 30, "policy": "local" } ] ], "case": [ [ "consumer_name", "==", "jane" ] ] }, { "actions": [ [ "limit-count", { "count": 2, "key": "$consumer_name", "key_type": "var", "rejected_code": 429, "time_window": 30, "policy": "local" } ] ] } ] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` Create three consumers and a route that enables per-consumer rate limiting: adc.yaml ``` consumers: - username: john credentials: - name: key-auth type: key-auth config: key: john-key - username: jane credentials: - name: key-auth type: key-auth config: key: jane-key - username: jimmy credentials: - name: key-auth type: key-auth config: key: jimmy-key services: - name: httpbin routes: - uris: - /anything name: workflow-route plugins: key-auth: {} workflow: rules: - case: - ["consumer_name", "==", "john"] actions: - - limit-count - count: 5 key: consumer_john key_type: constant rejected_code: 429 time_window: 30 policy: local - case: - ["consumer_name", "==", "jane"] actions: - - limit-count - count: 3 key: consumer_jane key_type: constant rejected_code: 429 time_window: 30 policy: local - actions: - - limit-count - count: 2 key: "$consumer_name" key_type: var rejected_code: 429 time_window: 30 policy: local upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` Create three consumers and a route that enables per-consumer rate limiting: About Consumer Name When consumers are configured using the Ingress Controller, the consumer name is generated in the format `namespace_consumername`. As a result, the `consumer_name` logic in the `workflow` plugin should match the consumer name in this format. * Gateway API * APISIX CRD workflow-ic.yaml ``` apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: john spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: john-key --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jane spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: jane-key --- apiVersion: apisix.apache.org/v1alpha1 kind: Consumer metadata: namespace: aic name: jimmy spec: gatewayRef: name: apisix credentials: - type: key-auth name: primary-key config: key: jimmy-key --- apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: workflow-plugin-config spec: plugins: - name: key-auth config: _meta: disable: false - name: workflow config: rules: - case: - ["consumer_name", "==", "aic_john"] actions: - - limit-count - count: 5 key: consumer_john key_type: constant rejected_code: 429 time_window: 30 policy: local - case: - ["consumer_name", "==", "aic_jane"] actions: - - limit-count - count: 3 key: consumer_jane key_type: constant rejected_code: 429 time_window: 30 policy: local - actions: - - limit-count - count: 2 key: "$consumer_name" key_type: var rejected_code: 429 time_window: 30 policy: local --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: workflow-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: workflow-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` workflow-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: john spec: ingressClassName: apisix authParameter: keyAuth: value: key: john-key --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jane spec: ingressClassName: apisix authParameter: keyAuth: value: key: jane-key --- apiVersion: apisix.apache.org/v2 kind: ApisixConsumer metadata: namespace: aic name: jimmy spec: ingressClassName: apisix authParameter: keyAuth: value: key: jimmy-key --- apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: workflow-route spec: ingressClassName: apisix http: - name: workflow-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: key-auth enable: true - name: workflow enable: true config: rules: - case: - ["consumer_name", "==", "aic_john"] actions: - - limit-count - count: 5 key: consumer_john key_type: constant rejected_code: 429 time_window: 30 policy: local - case: - ["consumer_name", "==", "aic_jane"] actions: - - limit-count - count: 3 key: consumer_jane key_type: constant rejected_code: 429 time_window: 30 policy: local - actions: - - limit-count - count: 2 key: "$consumer_name" key_type: var rejected_code: 429 time_window: 30 policy: local ``` Apply the configuration to your cluster: ``` kubectl apply -f workflow-ic.yaml ``` ❶ Enable `key-auth` on the route. ❷ Match consumer `john` and apply a rate limiting quota of 5 requests within a 30-second window. ❸ Match consumer `jane` and apply a rate limiting quota of 3 requests within a 30-second window. ❹ Match all other consumers and apply a rate limiting quota of 2 requests within a 30-second window, per consumer. To verify, send 6 consecutive requests with `john`'s key: ``` resp=$(seq 6 | xargs -I{} curl "http://127.0.0.1:9080/anything" -H 'apikey: john-key' -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that out of the 6 requests, 5 requests were successful (status code 200) while the others were rejected (status code 429). ``` 200: 5, 429: 1 ``` Send 6 consecutive requests with `jane`'s key: ``` resp=$(seq 6 | xargs -I{} curl "http://127.0.0.1:9080/anything" -H 'apikey: jane-key' -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that out of the 6 requests, 3 requests were successful (status code 200) while the others were rejected (status code 429). ``` 200: 3, 429: 3 ``` Send 3 consecutive requests with `jimmy`'s key: ``` resp=$(seq 3 | xargs -I{} curl "http://127.0.0.1:9080/anything" -H 'apikey: jimmy-key' -o /dev/null -s -w "%{http_code}\n") && \ count_200=$(echo "$resp" | grep "200" | wc -l) && \ count_429=$(echo "$resp" | grep "429" | wc -l) && \ echo "200": $count_200, "429": $count_429 ``` You should see the following response, showing that out of the 3 requests, 2 requests were successful (status code 200) while the others were rejected (status code 429). ``` 200: 2, 429: 1 ``` ### Apply Advanced Rate Limiting with Sliding Window[​](#apply-advanced-rate-limiting-with-sliding-window "Direct link to Apply Advanced Rate Limiting with Sliding Window") The following example demonstrates how to configure `workflow` with the Enterprise `limit-count-advanced` plugin to perform rate limiting conditionally, using the sliding window algorithm. Create a route with the `workflow` plugin as such: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "workflow-route", "uri": "/anything/*", "plugins":{ "workflow":{ "rules":[ { "case": [ ["uri", "==", "/anything/rate-limit-advanced"] ], "actions": [ [ "limit-count-advanced", { "count": 5, "time_window": 10, "rejected_code": 429, "policy": "local", "key_type": "var", "key": "remote_addr", "window_type": "sliding" } ] ] } ] } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything/* name: workflow-route plugins: workflow: rules: - case: - ["uri", "==", "/anything/rate-limit-advanced"] actions: - - limit-count-advanced - count: 5 time_window: 10 rejected_code: 429 policy: local key_type: var key: remote_addr window_type: sliding upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD workflow-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: workflow-plugin-config spec: plugins: - name: workflow config: rules: - case: - ["uri", "==", "/anything/rate-limit-advanced"] actions: - - limit-count-advanced - count: 5 time_window: 10 rejected_code: 429 policy: local key_type: var key: remote_addr window_type: sliding --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: workflow-route spec: parentRefs: - name: apisix rules: - matches: - path: type: PathPrefix value: /anything/ filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: workflow-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` workflow-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: workflow-route spec: ingressClassName: apisix http: - name: workflow-route match: paths: - /anything/* upstreams: - name: httpbin-external-domain plugins: - name: workflow enable: true config: rules: - case: - ["uri", "==", "/anything/rate-limit-advanced"] actions: - - limit-count-advanced - count: 5 time_window: 10 rejected_code: 429 policy: local key_type: var key: remote_addr window_type: sliding ``` Apply the configuration to your cluster: ``` kubectl apply -f workflow-ic.yaml ``` ❶ Match URI path `/anything/rate-limit-advanced`. ❷ Apply rate limiting when the condition is matched. ❸ Set the rate limiting algorithm to sliding window. Generate 7 requests to the route that matches the condition every other second: ``` for i in $(seq 7); do (curl -I "http://127.0.0.1:9080/anything/rate-limit-advanced" &) sleep 1 done ``` You should receive `HTTP/1.1 200 OK` responses for most requests, with the remainder being `HTTP 429 Too Many Requests responses`. If you send requests to the route with other paths, such as: ``` curl -i "http://127.0.0.1:9080/anything/else" ``` You should not observe any rate limiting in effect as the condition is not matched. --- ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * rules array\[object] required *** An array of one or more pairs of matching conditions and actions to be executed. * case array\[array] *** An array of one or more matching conditions in the form of [APISIX expressions](https://docs.api7.ai/apisix/reference/apisix-expressions.md). * actions array\[array] required *** An array of actions to be executed when a condition is successfully matched. Currently, the array only supports one action, and it should be either `return`, `limit-count`, or `limit-conn`. If you are using API7 Enterprise, you can also use `limit-count-advanced` as the action. When the action is set to `return`, you can configure an HTTP status code to return to the client when the condition is matched. When the action is set to `limit-count`, you can configure all options of the [`limit-count`](https://docs.api7.ai/hub/limit-count/configuration.md) plugin, except for `group`. When the action is set to `limit-count-advanced`, you can configure all options of the [`limit-count-advanced`](https://docs.api7.ai/hub/limit-count-advanced/configuration.md) plugin, except for `group`. When the action is set to `limit-conn`, you can configure all options of the [`limit-conn`](https://docs.api7.ai/hub/limit-conn/configuration.md) plugin. --- # zipkin [Zipkin](https://github.com/openzipkin/zipkin) is an open-source distributed tracing system. The `zipkin` plugin instruments APISIX and sends traces to Zipkin based on the [Zipkin API specification](https://zipkin.io/pages/instrumenting.html). The plugin can also send traces to other compatible collectors, such as [Jaeger](https://www.jaegertracing.io/docs/1.51/getting-started/#migrating-from-zipkin) and [Apache SkyWalking](https://skywalking.apache.org/docs/main/latest/en/setup/backend/zipkin-trace/#zipkin-receiver), both of which support Zipkin [v1](https://zipkin.io/zipkin-api/zipkin-api.yaml) and [v2](https://zipkin.io/zipkin-api/zipkin2-api.yaml) APIs. Tracing adds work to each sampled request. Use `sample_ratio` to balance trace coverage against that overhead, and use a lower ratio on high-throughput routes when full sampling is unnecessary. Requests excluded by sampling do not build span tags. Measure the effect with your traffic and collector configuration rather than assuming a fixed performance improvement. ## Examples[​](#examples "Direct link to Examples") The examples below show different use cases for the `zipkin` plugin. ### Send Traces to Zipkin[​](#send-traces-to-zipkin "Direct link to Send Traces to Zipkin") The following example demonstrates how to trace requests to a route and send traces to Zipkin using [Zipkin API v2](https://zipkin.io/zipkin-api/zipkin2-api.yaml). You will also understand the differences between span version 2 and span version 1. Start a Zipkin instance: * Docker * Kubernetes ``` docker run -d --name zipkin -p 9411:9411 openzipkin/zipkin ``` zipkin-server.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: zipkin spec: replicas: 1 selector: matchLabels: app: zipkin template: metadata: labels: app: zipkin spec: containers: - name: zipkin image: openzipkin/zipkin ports: - containerPort: 9411 --- apiVersion: v1 kind: Service metadata: namespace: aic name: zipkin spec: selector: app: zipkin ports: - port: 9411 targetPort: 9411 type: ClusterIP ``` Apply the manifest: ``` kubectl apply -f zipkin-server.yaml ``` Create a route with `zipkin` and use the default span version 2: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "zipkin-tracing-route", "uri": "/anything", "plugins": { "zipkin": { "endpoint": "http://127.0.0.1:9411/api/v2/spans", "sample_ratio": 1, "span_version": 2 } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: zipkin-tracing-route plugins: zipkin: endpoint: "http://127.0.0.1:9411/api/v2/spans" sample_ratio: 1 span_version: 2 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD zipkin-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: zipkin-plugin-config spec: plugins: - name: zipkin config: endpoint: "http://zipkin.aic.svc.cluster.local:9411/api/v2/spans" sample_ratio: 1 span_version: 2 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: zipkin-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: zipkin-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` zipkin-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: zipkin-route spec: ingressClassName: apisix http: - name: zipkin-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: zipkin enable: true config: endpoint: "http://zipkin.aic.svc.cluster.local:9411/api/v2/spans" sample_ratio: 1 span_version: 2 ``` Apply the configuration to your cluster: ``` kubectl apply -f zipkin-ic.yaml ``` ❶ Adjust the endpoint URL as needed. ❷ Configure the sample ratio to 1 to trace every request. ❸ Set span version to 2. Send a request to the route: ``` curl "http://127.0.0.1:9080/anything" ``` You should receive an `HTTP/1.1 200 OK` response similar to the following: ``` { "args": {}, "data": "", "files": {}, "form": {}, "headers": { "Accept": "*/*", "Host": "127.0.0.1", "User-Agent": "curl/7.64.1", "X-Amzn-Trace-Id": "Root=1-65af2926-497590027bcdb09e34752b78", "X-B3-Parentspanid": "347dddedf73ec176", "X-B3-Sampled": "1", "X-B3-Spanid": "429afa01d0b0067c", "X-B3-Traceid": "aea58f4b490766eccb08275acd52a13a", "X-Forwarded-Host": "127.0.0.1" }, ... } ``` Navigate to the Zipkin web UI at and click **Run Query**, you should see a trace corresponding to the request: ![Zipkin UI showing a list of traces matching the search query](https://static.api7.ai/uploads/2024/01/23/MaXhacYO_zipkin-run-query.png) Click **Show** to see more tracing details: ![Zipkin trace detail view showing spans for a single request](https://static.api7.ai/uploads/2024/01/23/3SmfFq9f_trace-details.png) Note that with span version 2, every traced request creates the following spans: ``` request ├── proxy └── response ``` where `proxy` represents the time from the beginning of the request to the beginning of `header_filter`, and `response` represents the time from the beginning of `header_filter` to the beginning of `log`. The request span includes an `apisix.response_source` tag. It classifies the response origin as `apisix` (generated by APISIX, such as plugin rejections), `nginx` (NGINX proxy errors), or `upstream` (real response from the upstream service). Introduced in API7 Enterprise 3.9.10 and APISIX 3.17.0. Now, update the plugin on the route to use span version 1: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes/zipkin-tracing-route" -X PATCH \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "plugins": { "zipkin": { "span_version": 1 } } }' ``` Update `adc.yaml` to set `span_version` to 1: adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: zipkin-tracing-route plugins: zipkin: endpoint: "http://127.0.0.1:9411/api/v2/spans" sample_ratio: 1 span_version: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD Update `zipkin-ic.yaml` to set `span_version` to 1: zipkin-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: zipkin-plugin-config spec: plugins: - name: zipkin config: endpoint: "http://zipkin.aic.svc.cluster.local:9411/api/v2/spans" sample_ratio: 1 span_version: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: zipkin-route spec: parentRefs: - name: apisix rules: - matches: - path: type: Exact value: /anything filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: zipkin-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` Update `zipkin-ic.yaml` to set `span_version` to 1: zipkin-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: zipkin-route spec: ingressClassName: apisix http: - name: zipkin-route match: paths: - /anything upstreams: - name: httpbin-external-domain plugins: - name: zipkin enable: true config: endpoint: "http://zipkin.aic.svc.cluster.local:9411/api/v2/spans" sample_ratio: 1 span_version: 1 ``` Reapply the configuration: ``` kubectl apply -f zipkin-ic.yaml ``` Send another request to the route: ``` curl "http://127.0.0.1:9080/anything" ``` In the Zipkin web UI, you should see a new trace with details similar to the following: ![Zipkin v1 trace detail view showing spans for a single request](https://static.api7.ai/uploads/2024/01/23/OPw2sTPa_v1-trace-spans.png) Note that with the older span version 1, every traced request creates the following spans: ``` request ├── rewrite ├── access └── proxy └── body_filter ``` ### Send Traces to Jaeger[​](#send-traces-to-jaeger "Direct link to Send Traces to Jaeger") The following example demonstrates how to trace requests to a route and send traces to Jaeger. Start a Jaeger instance: * Docker * Kubernetes ``` docker run -d --name jaeger \ -e COLLECTOR_ZIPKIN_HOST_PORT=9411 \ -p 16686:16686 \ -p 9411:9411 \ jaegertracing/all-in-one ``` jaeger-server.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: namespace: aic name: jaeger spec: replicas: 1 selector: matchLabels: app: jaeger template: metadata: labels: app: jaeger spec: containers: - name: jaeger image: jaegertracing/all-in-one env: - name: COLLECTOR_ZIPKIN_HOST_PORT value: "9411" ports: - containerPort: 16686 - containerPort: 9411 --- apiVersion: v1 kind: Service metadata: namespace: aic name: jaeger spec: selector: app: jaeger ports: - name: ui port: 16686 targetPort: 16686 - name: zipkin port: 9411 targetPort: 9411 type: ClusterIP ``` Apply the manifest: ``` kubectl apply -f jaeger-server.yaml ``` Create a route with `zipkin`: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \ -H "X-API-KEY: ${ADMIN_API_KEY}" \ -d '{ "id": "zipkin-tracing-route", "uri": "/anything", "plugins": { "zipkin": { "endpoint": "http://127.0.0.1:9411/api/v2/spans", "sample_ratio": 1 } }, "upstream": { "type": "roundrobin", "nodes": { "httpbin.org": 1 } } }' ``` adc.yaml ``` services: - name: httpbin routes: - uris: - /anything name: zipkin-tracing-route plugins: zipkin: endpoint: "http://127.0.0.1:9411/api/v2/spans" sample_ratio: 1 upstream: type: roundrobin nodes: - host: httpbin.org port: 80 weight: 1 ``` Synchronize the configuration to the gateway: ``` adc sync -f adc.yaml ``` * Gateway API * APISIX CRD zipkin-jaeger-ic.yaml ``` apiVersion: v1 kind: Service metadata: namespace: aic name: httpbin-external-domain spec: type: ExternalName externalName: httpbin.org --- apiVersion: apisix.apache.org/v1alpha1 kind: PluginConfig metadata: namespace: aic name: zipkin-jaeger-plugin-config spec: plugins: - name: zipkin config: endpoint: "http://jaeger.aic.svc.cluster.local:9411/api/v2/spans" sample_ratio: 1 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: namespace: aic name: zipkin-jaeger-route spec: parentRefs: - name: apisix hostnames: - "jaeger.example.com" rules: - matches: - path: type: PathPrefix value: / filters: - type: ExtensionRef extensionRef: group: apisix.apache.org kind: PluginConfig name: zipkin-jaeger-plugin-config backendRefs: - name: httpbin-external-domain port: 80 ``` zipkin-jaeger-ic.yaml ``` apiVersion: apisix.apache.org/v2 kind: ApisixUpstream metadata: namespace: aic name: httpbin-external-domain spec: ingressClassName: apisix externalNodes: - type: Domain name: httpbin.org --- apiVersion: apisix.apache.org/v2 kind: ApisixRoute metadata: namespace: aic name: zipkin-jaeger-route spec: ingressClassName: apisix http: - name: zipkin-jaeger-route match: hosts: - "jaeger.example.com" paths: - /* upstreams: - name: httpbin-external-domain plugins: - name: zipkin enable: true config: endpoint: "http://jaeger.aic.svc.cluster.local:9411/api/v2/spans" sample_ratio: 1 ``` Apply the configuration to your cluster: ``` kubectl apply -f zipkin-jaeger-ic.yaml ``` ❶ Adjust the endpoint URL as needed. ❷ Configure the sample ratio to 1 to trace every request. Send a request to the route: * Admin API * ADC * Ingress Controller ``` curl "http://127.0.0.1:9080/anything" ``` ``` curl "http://127.0.0.1:9080/anything" ``` ``` curl "http://127.0.0.1:9080/anything" -H "Host: jaeger.example.com" ``` You should receive an `HTTP/1.1 200 OK` response. Navigate to the Jaeger web UI at , select APISIX as the service, and click **Find Traces**, you should see a trace corresponding to the request: ![Jaeger UI showing traces forwarded by Zipkin](https://static.api7.ai/uploads/2024/01/23/X6QdLN3l_jaeger.png) Similarly, you should find more span details once you click into a trace: ![Jaeger trace detail view for a request forwarded from Zipkin](https://static.api7.ai/uploads/2024/01/23/iP9fXI2A_jaeger-details.png) ### Using Trace Variables in Logging[​](#using-trace-variables-in-logging "Direct link to Using Trace Variables in Logging") The following example demonstrates how to configure the `zipkin` plugin to set the following built-in variables, which can be used in logger plugins or access logs: * `zipkin_context_traceparent`: [trace parent](https://www.w3.org/TR/trace-context/#trace-context-http-headers-format) ID * `zipkin_trace_id`: trace ID of the current span * `zipkin_span_id`: span ID of the current span Enable access log output for these variables and allow the plugin to set NGINX variables: * Host or Docker * Kubernetes (Helm) Add or update this section in the gateway configuration file: config.yaml ``` nginx_config: http: enable_access_log: true access_log_format: '{"time": "$time_iso8601","zipkin_context_traceparent": "$zipkin_context_traceparent","zipkin_trace_id": "$zipkin_trace_id","zipkin_span_id": "$zipkin_span_id","remote_addr": "$remote_addr"}' access_log_format_escape: json plugin_attr: zipkin: set_ngx_var: true ``` ❶ `access_log_format`: customize the access log format to use the `zipkin` plugin variables. ❷ `set_ngx_var`: set `zipkin` variables. Reload the gateway for configuration changes to take effect. For Helm deployments, update the values that render the access log format and `plugin_attr.zipkin`. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` apisix: nginx: logs: enableAccessLog: true accessLogFormat: '{"time": "$time_iso8601","zipkin_context_traceparent": "$zipkin_context_traceparent","zipkin_trace_id": "$zipkin_trace_id","zipkin_span_id": "$zipkin_span_id","remote_addr": "$remote_addr"}' accessLogFormatEscape: json pluginAttrs: zipkin: set_ngx_var: true ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` logs: enableAccessLog: true accessLogFormat: '{"time": "$time_iso8601","zipkin_context_traceparent": "$zipkin_context_traceparent","zipkin_trace_id": "$zipkin_trace_id","zipkin_span_id": "$zipkin_span_id","remote_addr": "$remote_addr"}' accessLogFormatEscape: json pluginAttrs: zipkin: set_ngx_var: true ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` You should see access log entries similar to the following when you generate requests: ``` {"time": "23/Jan/2024:06:28:00 +0000","zipkin_context_traceparent": "00-61bce33055c56f5b9bec75227befd142-13ff3c7370b29925-01","zipkin_trace_id": "61bce33055c56f5b9bec75227befd142","zipkin_span_id": "13ff3c7370b29925","remote_addr": "172.28.0.1"} ``` --- ## Static Configurations[​](#static-configurations "Direct link to Static Configurations") By default, `zipkin` plugin NGINX variables configuration is set to false in the [default configuration](https://github.com/apache/apisix/blob/master/apisix/cli/config.lua). The file to update depends on how the gateway is deployed: * Host or Docker * Kubernetes (Helm) For host or Docker deployments, configure the following settings: config.yaml ``` plugin_attr: zipkin: set_ngx_var: true ``` Then reload the gateway for changes to take effect. For Helm deployments, update the chart values that render `plugin_attr.zipkin`. Keep the rest of your values file unchanged. For the APISIX Helm chart, set the following values: values.yaml ``` apisix: pluginAttrs: zipkin: set_ngx_var: true ``` For the API7 Gateway Helm chart, set the following values: values.yaml ``` pluginAttrs: zipkin: set_ngx_var: true ``` Then apply the values file with the chart used for this gateway release: ``` helm upgrade -n -f values.yaml ``` ## Parameters[​](#parameters "Direct link to Parameters") See plugin [common configurations](https://docs.api7.ai/apisix/reference/plugin-common-configurations.md) for configuration options available to all plugins. * endpoint string required *** Zipkin span endpoint to POST to, such as `http://127.0.0.1:9411/api/v2/spans`. * sample\_ratio number required vaild vaule: between 0.00001 and 1 inclusive *** Frequency to sample requests. Setting to 1 means sampling every request. * service\_name string default: `APISIX` *** Service name for the Zipkin reporter to be displayed in Zipkin. * server\_addr string default: `the value of $server_addr` vaild vaule: IPv4 address *** IPv4 address for the Zipkin reporter. For example, you can set this to your external IP address. * span\_version integer default: `2` vaild vaule: 1 or 2 *** Version of the span type. --- AI traffic gateway # AISIX AI Gateway Put one stable API contract in front of your AI providers AISIX AI Gateway is a Rust-native gateway for LLM and AI-agent traffic. The [open-source gateway](https://github.com/api7/aisix) runs as a single static binary and can operate standalone, or you can use AISIX Cloud for centralized management. In either case, applications call stable model aliases while the AI platform team controls provider credentials, routing, failover, rate limits, caching, guardrails, and observability. AISIX Cloud adds centralized usage and budget management. [Get started](https://docs.api7.ai/ai-gateway/getting-started/products-and-deployment-options.md)[View supported endpoints](https://docs.api7.ai/ai-gateway/endpoints/overview.md) Start an on-premises AISIX Cloud deploymentlocalhost:8080 ``` curl -fSL https://run.api7.ai/aisix-self-hosted/aisix-self-hosted-1.0.0.tar.gz -o aisix-self-hosted-1.0.0.tar.gz tar -xzf aisix-self-hosted-1.0.0.tar.gz ./aisix-self-hosted/run.sh # Starts the control plane and dashboard # Dashboard: http://localhost:8080 ``` ## How requests flow Applications call AISIX first. AISIX authenticates the caller, resolves the model alias, applies policy, and forwards the request to the selected provider. Applications**Apps, agents, and services**Send OpenAI-compatible requests with gateway-issued caller API keys. -> AISIX gateway boundary**Stable contract, controlled provider access** AISIX keeps client traffic on one API shape while resolving the provider-side target for each request. Authenticate callerResolve model aliasApply limits, cache, and guardrailsSelect provider route -> Providers**OpenAI, Anthropic, Bedrock, Vertex, Azure**Receive provider-authenticated requests from the gateway. Provider responses return through AISIX using the same client-facing contract. ## What changes when AI traffic goes through AISIX Application teams keep calling a familiar API. AI platform teams move provider credentials, model aliases, routing, and policy into one operator-managed gateway. AISIX Cloud adds shared usage and budget controls. *Access***Caller API keys and model allowlists**Authenticate applications and decide which model aliases each key can use. *Providers***Centralized upstream credentials**Keep provider keys, base URLs, and adapter details out of application code. *Routing***Stable aliases with failover**Expose one model name while AISIX selects target models behind it. *Policy***Limits, cache, guardrails, telemetry**Apply AI-specific controls and usage visibility before requests leave your gateway boundary. ## Why AISIX if you already use APISIX or API7 Gateway? APISIX and API7 Gateway can add AI behavior to regular gateway routes through AI plugins. That path fits teams whose main workload is API traffic and whose AI calls are part of a broader API gateway deployment. AISIX is for teams whose main workload is AI traffic. It manages provider keys, model aliases, caller API keys, routing, policy, and AI usage telemetry as first-class resources instead of route-level plugin settings. AISIX Cloud adds centralized usage views and budgets. *APISIX* **Add AI features to existing API routes.**Use AI plugins when your existing API gateway route is the natural place to call or transform AI services. *AISIX* **Run AI traffic through a dedicated gateway domain.**Use AISIX when applications should call stable model names while an AI platform team owns provider keys, routing, policy, and usage visibility. Read next ## Choose Your Next Path Start by choosing how to run AISIX. From there, explore AISIX Cloud, connect an application, or prepare a gateway for production. [*Get started*](https://docs.api7.ai/ai-gateway/getting-started/products-and-deployment-options.md) [**Choose a product and deployment option**Run the open-source gateway standalone, or use AISIX Cloud and decide who hosts the control plane.](https://docs.api7.ai/ai-gateway/getting-started/products-and-deployment-options.md) [*AISIX Cloud*](https://docs.api7.ai/ai-gateway/cloud/overview.md) [**Manage AISIX gateways with a control plane**Understand the control-plane workflow, centralized usage reporting, and budget controls.](https://docs.api7.ai/ai-gateway/cloud/overview.md) [*Integrations*](https://docs.api7.ai/ai-gateway/integrations.md) [**Connect applications and AI tools**Point client SDKs, coding agents, and application frameworks at an AISIX gateway.](https://docs.api7.ai/ai-gateway/integrations.md) [*Production*](https://docs.api7.ai/ai-gateway/deployment/production.md) [**Prepare the runtime for production**Review deployment, security, health, metrics, and troubleshooting guidance for production gateways.](https://docs.api7.ai/ai-gateway/deployment/production.md) --- # Control Agent Access For Agent-to-Agent (A2A) traffic, each caller API key defines which registered agents that caller can reach. AISIX denies agent calls and agent-card discovery unless the key's `allowed_agents` grant covers the agent's registered name. Use separate grants when clients sharing a gateway need access to different upstream agents. This guide explains how AISIX matches agent names and patterns, how to update grants through AISIX Cloud or `resources.yaml`, and what happens when a request falls outside the grant. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Complete [Set Up Agent Gateway](https://docs.api7.ai/ai-gateway/agent-gateway/setup.md) for AISIX Cloud or the open-source AISIX gateway. Keep the shell, gateway, and test agent running if you want to verify a changed grant. ## How Agent Access Works[​](#how-agent-access-works "Direct link to How Agent Access Works") `allowed_agents` is a list of agent-name patterns on the caller API key. When the field is omitted, `null`, or an empty list, the key has no A2A agent access. Each pattern is matched against the agent's registered `name`: | Entry | Grants | Example | | ------------ | ------------------------------------------ | ----------------------------------------------------------------- | | Exact name | One agent. | `invoice-processor` grants only that agent. | | Name pattern | Agents whose names match one `*` wildcard. | `invoice-*` grants every agent whose name starts with `invoice-`. | | `*` | Every registered agent. | `*` grants current and future agents. | Choose exact names for the narrowest access. Use `*` only for a caller that may reach every current and future agent. The matcher is the same single-`*` glob matcher used by [MCP tool patterns](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md). An entry can contain at most one `*`. ## Configure Agent Access[​](#configure-agent-access "Direct link to Configure Agent Access") Set the complete list of agents that the caller should keep. AISIX Cloud stores the grant on the API key resource; an open-source AISIX gateway reads it from `resources.yaml`. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Update the caller API key created in the setup guide: ``` curl -fsS -X PATCH \ "$AISIX_CP/environments/$ENV_ID/api_keys/$API_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "allowed_agents": ["echo-agent", "invoice-*"] }' | jq ``` This partial update replaces the previous agent grant while retaining the key's model, MCP, and other settings. The control plane projects the change to attached gateways automatically. To revoke all A2A agent access, send `"allowed_agents": []` or `"allowed_agents": null`. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Replace the existing `quickstart-caller` entry with the updated entry below to set `allowed_agents`. Preserve every other entry in `api_keys` and all unrelated collections, and do not create a second top-level `api_keys` key: resources.yaml (agent access) ``` api_keys: - display_name: quickstart-caller key_env: CALLER_API_KEY allowed_models: - gpt-4o-mini allowed_agents: - echo-agent - invoice-* ``` Validate the complete file and reload it: ``` docker exec aisix-quickstart \ aisix validate --resources /etc/aisix/resources.yaml docker kill --signal HUP aisix-quickstart ``` Remove `allowed_agents` or set it to an empty list to revoke all A2A agent access for this caller. ## How Enforcement Works[​](#how-enforcement-works "Direct link to How Enforcement Works") The gateway checks the grant after confirming that the agent exists and is enabled. It performs the check before contacting the upstream agent for both A2A calls and agent-card discovery: * JSON-RPC calls to `/a2a/`. * Agent-card requests to `/a2a//.well-known/agent-card.json`. The caller-visible outcome depends on the agent and key state: * An unknown or disabled agent returns `404`. * A known agent outside the key's grant returns `403` before AISIX contacts the upstream agent. * A known agent covered by the key's grant is forwarded upstream. Because the existence check runs first, an authenticated caller can distinguish an agent it cannot access from an agent that does not exist. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now scoped which agents a caller API key may reach. Use these guides to complete the caller and upstream control path: * [Upstream Authentication](https://docs.api7.ai/ai-gateway/agent-gateway/upstream-authentication.md): configure how AISIX authenticates to each upstream agent. * [Rate Limits and Budgets](https://docs.api7.ai/ai-gateway/agent-gateway/traffic-controls.md): apply request and concurrency limits, and use AISIX Cloud budgets. * [Caller API Keys](https://docs.api7.ai/ai-gateway/traffic-controls/caller-api-keys.md): review the shared key settings for model, MCP, and A2A traffic. --- # Observability Agent-to-Agent (A2A) traffic uses the same telemetry pipelines as model traffic. Usage-event fields identify the caller, agent, method, and outcome, while Prometheus labels let you separate agent traffic from model traffic. Use these signals to measure agent call volume and monitor failures. Rate limits and upstream errors appear in the same observability tools you already use for model traffic. In AISIX Cloud, budget rejections appear there as well. ## Usage Events[​](#usage-events "Direct link to Usage Events") AISIX emits an A2A usage event once the gateway can attribute the request to an enabled agent and caller. This includes calls rejected by unsupported upstream authentication, rate limits, upstream failures, or budgets configured through AISIX Cloud. Malformed request bodies can be rejected before a usage event is emitted. The event goes into the same sink as model usage and identifies the caller, agent, method, outcome, and timing: | Field | Value | | ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `inbound_protocol` | `a2a` | | `a2a_agent_name` | The registered agent that was called. | | `a2a_method` | The JSON-RPC method exactly as the caller wrote it, such as `message/send` or its 1.0 spelling `SendMessage`, when AISIX can read it from the request body. | | `a2a_operation` | The canonical operation `a2a_method` names. Group by this rather than by the raw method: A2A 0.3 and 1.0 spell one operation two ways, and an unrecognized method becomes `unknown`. | | `a2a_protocol_version` | The wire version AISIX announced to the agent, `0.3` or `1.0`. | | `a2a_task_id` | The task the call created or acted on. Use it to gather every request that touched one task across `message/send`, `tasks/get`, and `tasks/resubscribe`. Empty when the call named no task. | | `a2a_context_id` | The conversation the task belongs to, which ties a multi-turn exchange's tasks together. | | `a2a_task_state` | The last state the agent reported, normalized to `submitted`, `working`, `input-required`, `auth-required`, `completed`, `canceled`, `failed`, `rejected`, or `unknown`. Empty when no response carried a state. | | `a2a_stream_event_count` | Events relayed to the caller on a streamed call. 0 on a non-streaming call. | | `upstream_ttft_ms` | Time to the agent's first streamed event. Read with `upstream_latency_ms` it separates a stream that said nothing for a long time from one that said plenty. | | `prompt_tokens`, `completion_tokens` | Token counts for the message text, counted by the gateway. See [Token counts](#token-counts). | | `api_key_id` | The caller API key that made the call. | | `status_code` | The call's outcome status. | | `upstream_latency_ms` | Time spent on the upstream agent call. | | `downstream_latency_ms` | Total time the caller waited for the agent call. | | `request_id`, `occurred_at` | Correlation id and timestamp. | ### Token Counts[​](#token-counts "Direct link to Token Counts") An A2A agent reports no usage of its own, because the protocol has no usage block. AISIX therefore counts the message text itself, using the gateway's own tokenizer over the text parts of the caller's message and of the agent's reply. Those counts land in `prompt_tokens` and `completion_tokens`. The event also carries `usage_estimated: true`, which says the gateway produced the counts rather than an upstream reporting them. Filter on that flag when you need provider-billed exactness. Only `message/send` and `message/stream` are counted. A read such as `tasks/get` returns an answer that was already counted when the agent produced it, so counting it again would report one answer once per poll. `cost_usd` remains zero. What an agent charges is not something the gateway can observe. These counts are reported only, never charged: they do not add to the caller key's `tpm` / `tpd` token windows. Budgets are a separate mechanism — they limit USD spend, and a zero cost adds none. See [Traffic Controls](https://docs.api7.ai/ai-gateway/agent-gateway/traffic-controls.md). File and data parts are not counted. Their bytes are not language, and including a base64 blob would distort the estimate. Agent Gateway currently makes a single upstream attempt that spans the request, so `upstream_latency_ms` and `downstream_latency_ms` report the same duration. The separate fields keep A2A records consistent with other gateway traffic, where retries and gateway processing can make the two values differ. The event does not include request or response message content. Message text reaches an [observability exporter](https://docs.api7.ai/ai-gateway/observability/exporters.md) only when that exporter is configured for full content capture, and it never reaches the AISIX Cloud control plane. A2A usage events follow the same delivery paths as model usage events. Any configured observability exporter receives them, so A2A traffic appears alongside the rest of your gateway traffic. In AISIX Cloud, they also flow to the control plane's usage sink. ## Metrics[​](#metrics "Direct link to Metrics") A2A requests appear in the gateway's Prometheus metrics with labels that distinguish them from model traffic. Use the labels below to filter the relevant metric series: | Goal | Metric | Filter | | -------------------------------- | ----------------------------------- | ---------------------------------------------- | | Count A2A requests by outcome. | `aisix_requests_total` | `provider="a2a"` and `model="a2a"` | | Track active A2A requests. | `aisix_proxy_in_flight_requests` | `endpoint="/a2a"` and `inbound_protocol="a2a"` | | Measure the agent-call lifetime. | `aisix_request_e2e_latency_seconds` | `endpoint="/a2a"` | | Check A2A usage-event emission. | `aisix_usage_events_emitted_total` | `handler="a2a"` | The `aisix_a2a_*` family carries the dimension the shared families cannot: which agent was reached and which operation was invoked. | Metric | Labels | What it answers | | ------------------------------- | ------------------------------ | ----------------------------------------------------------------------------------------------------- | | `aisix_a2a_requests_total` | `agent`, `operation`, `status` | Call volume and failure rate for one agent's operation. | | `aisix_a2a_ttfb_seconds` | `agent`, `operation` | How long an agent takes to send its first streamed event. | | `aisix_a2a_stream_events_total` | `agent`, `operation` | Events relayed. Divided by the request count over the streaming operations, it gives events per call. | | `aisix_a2a_task_state_total` | `agent`, `state` | The rate at which calls end on each state, such as a rising share of `failed`. | `aisix_a2a_requests_total` does not agree with `aisix_proxy_requests_total{endpoint="/a2a"}`. Two differences are deliberate. A call refused before its agent is resolved has no agent to file under, so it is counted in the proxy family only. That covers a bad key, a denied agent, and an unknown one. A stream the caller abandons is a `4xx` here but a `2xx` there, because the response really did begin as a 200. Read this family for agent health and the proxy family for route traffic. The end-to-end histogram also begins after early request validation. Calls rejected before A2A accounting are absent, while a pre-dispatch rejection that reaches accounting records zero duration. For dispatched calls, the histogram covers the agent-call lifetime through stream completion. Task ids, context ids, and JSON-RPC request ids are never metric labels. They are what makes a call traceable, and that is exactly what makes them unusable as label values; they appear in usage events and traces instead. For usage-event emission, filter by handler. The emission counter keeps its protocol label bounded and groups A2A events under `other`. This label behavior applies only to the Prometheus emission counter. The delivered usage event still identifies the traffic as A2A and includes the agent name and method. Metrics are exposed on `GET /metrics` through the dedicated metrics listener. For the full metric catalog and label semantics, see [Metrics Reference](https://docs.api7.ai/ai-gateway/reference/metrics.md). ## Verify Metrics[​](#verify-metrics "Direct link to Verify Metrics") To verify that A2A metrics are emitted, send one A2A call through the gateway, then scrape the dedicated metrics listener. The example below uses the default listener address and path. If your startup configuration sets a different `observability.metrics.prometheus.addr`, use that address instead. Metric families register on first observation, so the A2A series appears only after a call is recorded: ``` curl -sS "http://127.0.0.1:9090/metrics" \ | grep -E 'aisix_a2a_|aisix_request_e2e_latency_seconds.*endpoint="/a2a"|handler="a2a"|provider="a2a"' ``` The output should include shared and A2A-specific metric samples: | Metric | Label | | ----------------------------------- | --------------------------------------------------- | | `aisix_usage_events_emitted_total` | `handler="a2a"` | | `aisix_requests_total` | `provider="a2a"` | | `aisix_a2a_requests_total` | `agent` and `operation` identify the resolved call. | | `aisix_request_e2e_latency_seconds` | `endpoint="/a2a"` | ## Next Steps[​](#next-steps "Direct link to Next Steps") You now know where A2A calls appear in usage events and metrics. Use these guides to review the full metric catalog or adjust the traffic that produces those signals: * [Metrics Reference](https://docs.api7.ai/ai-gateway/reference/metrics.md): review the full metric catalog and label semantics. * [Rate limits and budgets](https://docs.api7.ai/ai-gateway/agent-gateway/traffic-controls.md): apply request limits, concurrency limits, and budgets to A2A calls. * [Control agent access](https://docs.api7.ai/ai-gateway/agent-gateway/agent-access-control.md): scope caller API keys to specific agents or every agent. --- # Agent Gateway Overview AISIX exposes registered Agent-to-Agent (A2A) agents at `/a2a/`, giving A2A clients and other agents one authenticated path to reach them. Callers present an AISIX caller API key; the gateway checks whether the key may reach the target and forwards the A2A JSON-RPC request without exposing the upstream credential. Agent traffic therefore uses the same authentication, access-control, traffic-control, and telemetry boundary as model and MCP traffic. The same caller API key can govern the models, MCP tools, and A2A agents a caller may use. Each A2A agent resource represents one upstream agent that speaks the [A2A protocol](https://a2a-protocol.org/) over HTTP with JSON-RPC 2.0. ## How the Agent Gateway Works[​](#how-the-agent-gateway-works "Direct link to How the Agent Gateway Works") Each upstream agent has a `name`. AISIX exposes the agent at `/a2a/` on the proxy listener. In AISIX Cloud, the agent is registered at the organization level and exposed to selected environments. In an open-source AISIX gateway, it is normally declared in `resources.yaml`. AISIX forwards the request body unchanged. Each agent resource is pinned to A2A `1.0` or `0.3`, and the caller must use the format configured for that agent. AISIX does not translate between protocol versions. For an agent call that passes authentication and access checks, the gateway applies request and concurrency limits, contacts the agent with its configured upstream credential, and records A2A usage telemetry. AISIX Cloud can also apply budgets that cover the caller API key. ## Get Started[​](#get-started "Direct link to Get Started") Follow [Set Up Agent Gateway](https://docs.api7.ai/ai-gateway/agent-gateway/setup.md) to run an A2A echo agent, register it through AISIX Cloud or `resources.yaml`, and grant caller access. The official A2A Go SDK client then sends an A2A 1.0 message through the gateway. Continue to [A2A Streaming and Agent-Card Discovery](https://docs.api7.ai/ai-gateway/agent-gateway/streaming-and-discovery.md) to consume a streamed task. After the gateway has a public origin, the same guide verifies resolution of the client-facing card. Both management paths configure the same gateway runtime and A2A endpoint. ## Client Connection[​](#client-connection "Direct link to Client Connection") A2A clients call the registered agent on the AISIX proxy listener: | Setting | Value | | -------------- | --------------------------------------------------------------- | | Agent URL | `/a2a/` | | Agent card URL | `/a2a//.well-known/agent-card.json` | | Protocol | A2A 1.0 or 0.3 over HTTP with JSON-RPC 2.0 | | Request header | `Authorization: Bearer ` | The caller API key controls which agents the client may reach. The client connects only to AISIX; it does not receive the upstream credential. The endpoint accepts the JSON-RPC methods defined by the agent's configured A2A version, including message, task, streaming, and push-notification configuration methods. Streaming methods return `text/event-stream`, and AISIX relays each event as the upstream agent emits it. ## Agent Cards[​](#agent-cards "Direct link to Agent Cards") Clients can request a registered agent's discovery document through AISIX: ``` GET /a2a//.well-known/agent-card.json ``` The request uses the same caller authentication and agent-access check as an A2A call. AISIX fetches the upstream card with the agent's configured upstream credential, then rewrites every advertised service URL to the gateway path. Other card fields are preserved. AISIX derives the advertised scheme from `X-Forwarded-Proto` and the authority from `Host`. Configure a trusted reverse proxy to set these headers to the gateway's public address. When `X-Forwarded-Proto` is absent, AISIX uses `https`. The upstream card currently must include the top-level `url` field used by A2A 0.3. For an A2A 1.0 agent, publish a compatibility card with both `url` and `supportedInterfaces`. AISIX rewrites every service URL in the card. AISIX first checks `agent-card.json` under the registered path and then under its origin. If neither location returns a usable card, AISIX repeats the checks for the earlier `agent.json` filename. ## Govern A2A Calls[​](#govern-a2a-calls "Direct link to Govern A2A Calls") A2A calls use the same caller API key boundary as model requests. You do not configure a separate policy stack for agent traffic. Use these guides to refine the A2A path: * [Upstream Authentication](https://docs.api7.ai/ai-gateway/agent-gateway/upstream-authentication.md): configure how AISIX authenticates to an upstream agent. * [Control Agent Access](https://docs.api7.ai/ai-gateway/agent-gateway/agent-access-control.md): scope each caller API key to exact agent names, name patterns, or every agent. * [Rate Limits and Budgets](https://docs.api7.ai/ai-gateway/agent-gateway/traffic-controls.md): apply caller request and concurrency limits, and use AISIX Cloud budgets. * [Guardrail Behavior](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md): screen A2A message text on the input hook. Environment, caller API key, and team attachments can cover A2A calls; model and MCP-server attachments cannot. * [Observability](https://docs.api7.ai/ai-gateway/agent-gateway/observability.md): review the usage events and metrics emitted by A2A calls. ## Current Limitations[​](#current-limitations "Direct link to Current Limitations") The following limitations affect the agent interfaces and controls that AISIX exposes: * AISIX serves A2A JSON-RPC calls through `/a2a/`. It does not expose A2A REST path-based endpoints such as `POST .../v1/message:send` or `GET .../v1/tasks/{id}`. * OAuth 2.0 upstream authentication is not available. Use `none`, `bearer`, or `api_key` as described in [Upstream Authentication](https://docs.api7.ai/ai-gateway/agent-gateway/upstream-authentication.md). * A2A guardrails run on the input hook only. They do not inspect the upstream agent's response. * Register agents by their A2A HTTP URL. Direct registration for cloud agent runtime resources such as Amazon Bedrock AgentCore, Azure AI Foundry, or Vertex AI Agent Engine is not available. ## Troubleshoot Agent Calls[​](#troubleshoot-agent-calls "Direct link to Troubleshoot Agent Calls") If a caller cannot reach an agent, check these items: * The A2A agent resource is enabled and reachable from the gateway. * In AISIX Cloud, `allowed_environments` includes the caller key's environment. * The caller API key's `allowed_agents` grant covers the registered agent name. * The request body and method use the agent's configured A2A protocol version. * The configured upstream authentication matches what the agent expects. A missing or invalid caller API key returns `401`. A known agent outside the key's grant returns `403`, while an unknown or disabled agent returns `404`. For a JSON-RPC call, an unreachable upstream or an unsuccessful upstream HTTP status returns `502` with a JSON-RPC error envelope; AISIX does not expose the upstream response body. An unsuccessful agent-card fetch also returns `502`, but as an ordinary HTTP error rather than a JSON-RPC envelope. For endpoint-level behavior, see [Proxy API Reference](https://docs.api7.ai/ai-gateway/reference/proxy-api.md#a2a-gateway). For error response details, see [Headers and Error Codes](https://docs.api7.ai/ai-gateway/reference/headers-and-error-codes.md#a2a-errors). ## Next Steps[​](#next-steps "Direct link to Next Steps") Use these guides to configure how A2A traffic is handled: * [Set Up Agent Gateway](https://docs.api7.ai/ai-gateway/agent-gateway/setup.md): register and call an agent through AISIX Cloud or an open-source AISIX gateway. * [A2A Streaming and Agent-Card Discovery](https://docs.api7.ai/ai-gateway/agent-gateway/streaming-and-discovery.md): consume streamed task events and resolve a gateway-rewritten agent card. * [Upstream Authentication](https://docs.api7.ai/ai-gateway/agent-gateway/upstream-authentication.md): keep an upstream bearer token or API key gateway-side. * [Proxy API Reference](https://docs.api7.ai/ai-gateway/reference/proxy-api.md#a2a-gateway): review the A2A endpoint contract and constraints. --- # Set Up Agent Gateway AISIX exposes registered Agent-to-Agent (A2A) agents through caller-authenticated gateway endpoints at `/a2a/`. This setup connects an official A2A echo agent to an existing AISIX gateway and grants the quickstart caller access. The official A2A Go SDK client then sends an A2A 1.0 message through the gateway. Follow either the AISIX Cloud or open-source configuration path. The registration steps differ, but both paths configure the same gateway behavior. After either path, use the same agent-card readiness check and SDK client call to verify the result. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * For AISIX Cloud, complete the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), then keep its shell and `aisix-dp` gateway running. * For the open-source AISIX gateway, complete the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md), then remain in its shell and working directory with the `aisix-quickstart` gateway running. * [Docker](https://docs.docker.com/get-docker/), [cURL](https://curl.se/), and [jq](https://jqlang.github.io/jq/). ## Start an A2A Test Agent[​](#start-an-a2a-test-agent "Direct link to Start an A2A Test Agent") The example runs the [echo server](https://github.com/a2aproject/a2a-go/blob/v2.5.0/cmd/README.md#--echo---echo-mode) provided by the official A2A Go SDK in Docker. The agent and your existing gateway join a temporary Docker network, so the gateway can reach the agent without a restart. Use the echo server only for local testing: it has no authentication and returns the caller's message text. The server publishes a dual A2A 0.3 and 1.0 compatibility card so you can also verify client-facing discovery through AISIX. The echo server starts without upstream credentials; later client checks pass the disposable quickstart caller key into this test container. Stop and remove it after finishing the guide. Set the gateway container name for your configuration path. For AISIX Cloud: ``` export AISIX_GATEWAY_CONTAINER="aisix-dp" ``` For the open-source AISIX gateway: ``` export AISIX_GATEWAY_CONTAINER="aisix-quickstart" ``` Create a temporary network and connect the running gateway to it: ``` docker network create aisix-a2a docker network connect aisix-a2a "$AISIX_GATEWAY_CONTAINER" ``` Start the echo agent on the same network. The pinned Go SDK is downloaded and compiled inside the container, so the first startup can take about a minute: ``` docker run -d --name aisix-a2a-echo \ --network aisix-a2a \ golang:1.25-alpine \ sh -c 'go run github.com/a2aproject/a2a-go/v2/cmd/a2a@v2.5.0 \ serve --echo --card-compat --host 0.0.0.0 --port 8080 \ --transport jsonrpc --protocol latest --name "AISIX A2A Echo"' ``` Wait until the agent is ready: ``` for attempt in $(seq 1 120); do docker logs aisix-a2a-echo 2>&1 | grep -q "Listening on" && break sleep 1 done docker logs aisix-a2a-echo 2>&1 | grep "Listening on" ``` The final command prints the listener address. From the gateway container, the agent is available at `http://aisix-a2a-echo:8080`. ## Register the Agent and Grant Access[​](#register-the-agent-and-grant-access "Direct link to Register the Agent and Grant Access") Register the agent using the management path for your deployment. Both paths name the agent `echo-agent`, pin it to A2A 1.0, and grant it to the existing quickstart caller. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Register the test agent and expose it to the quickstart environment: ``` A2A_AGENT_RESPONSE=$(curl -fsS -X POST "$AISIX_CP/a2a_agents" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @- </dev/null 2>&1 && break sleep 2 done curl -fsS \ "$AISIX_PROXY/a2a/echo-agent/.well-known/agent-card.json" \ -H "Authorization: Bearer $AISIX_A2A_KEY" \ >/dev/null ``` Set the agent URL that the client container can reach. Both AISIX quickstart guides use port `3000` inside the gateway container: ``` export AISIX_A2A_URL="http://${AISIX_GATEWAY_CONTAINER}:3000/a2a/echo-agent" ``` Run the command-line client from the official A2A Go SDK inside the echo-agent container. `--transport jsonrpc` connects directly to the registered AISIX endpoint instead of resolving an agent card first: ``` docker exec aisix-a2a-echo \ go run github.com/a2aproject/a2a-go/v2/cmd/a2a@v2.5.0 \ send "$AISIX_A2A_URL" "Hello through AISIX" \ --transport jsonrpc \ --auth "Bearer $AISIX_A2A_KEY" \ --output json | jq -e \ '.artifacts[].parts[] | select(.text == "Hello through AISIX")' ``` The command prints the matching artifact part. AISIX authenticated the caller, checked its agent grant, and forwarded the SDK client's A2A 1.0 request to the echo agent. ## Adapt the Setup for Your Agent[​](#adapt-the-setup-for-your-agent "Direct link to Adapt the Setup for Your Agent") Replace the test URL with an A2A JSON-RPC endpoint reachable from the gateway. Set `protocol_version` to the wire format the agent supports, and send requests in that format. AISIX supports `"1.0"` and `"0.3"` but does not translate between them. Configure a required upstream credential with `auth_type` and `secret`. See [Upstream Authentication](https://docs.api7.ai/ai-gateway/agent-gateway/upstream-authentication.md). For an open-source AISIX gateway, validate the complete resources file and send `SIGHUP` when the running gateway already has every referenced environment variable. If you add or change an environment variable, recreate the container with the new value. See [Reload a Resources File](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md#reload-a-resources-file). ## Clean Up[​](#clean-up "Direct link to Clean Up") Keep the agent and caller grant if you plan to continue with the other Agent Gateway guides. Otherwise, remove the resources added through your management path. For AISIX Cloud, clear the caller's agent grant and delete the agent: ``` curl -fsS -X PATCH \ "$AISIX_CP/environments/$ENV_ID/api_keys/$API_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"allowed_agents":[]}' | jq curl -fsS -X DELETE "$AISIX_CP/a2a_agents/$A2A_AGENT_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" ``` For an open-source AISIX gateway, remove the `allowed_agents` and `a2a_agents` additions from `resources.yaml`, validate the file, and send `SIGHUP` again. Remove the test agent and temporary network: ``` docker rm -f aisix-a2a-echo docker network disconnect aisix-a2a "$AISIX_GATEWAY_CONTAINER" docker network rm aisix-a2a ``` ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now registered an A2A agent, granted caller access, and sent an A2A 1.0 message through AISIX with the official SDK client. Use these guides to extend the setup: * [A2A Streaming and Agent-Card Discovery](https://docs.api7.ai/ai-gateway/agent-gateway/streaming-and-discovery.md): consume a streamed task and verify the client-facing agent card through a public gateway origin. * [Configure upstream authentication](https://docs.api7.ai/ai-gateway/agent-gateway/upstream-authentication.md): use a bearer token or API key for the upstream agent. * [Control agent access](https://docs.api7.ai/ai-gateway/agent-gateway/agent-access-control.md): grant exact agent names, name patterns, or every registered agent. * [Apply rate limits and budgets](https://docs.api7.ai/ai-gateway/agent-gateway/traffic-controls.md): govern A2A calls with caller rate limits and AISIX Cloud budgets. * [Observability](https://docs.api7.ai/ai-gateway/agent-gateway/observability.md): review the usage events and metrics emitted by A2A calls. --- # A2A Streaming and Agent-Card Discovery The [Agent Gateway setup](https://docs.api7.ai/ai-gateway/agent-gateway/setup.md) uses the official A2A Go SDK client to send a non-streaming message through AISIX. This guide extends that verified path by consuming a streaming task and resolving the agent card that AISIX rewrites for clients. Both tests exercise client behavior beyond a direct JSON-RPC request. The SDK interprets each streamed task event and, during discovery, selects the gateway service URL from the returned card without learning the upstream agent URL or credential. The streaming test works with the local quickstart network. End-to-end agent-card resolution requires a client-reachable gateway origin; this guide uses a public HTTPS origin. If that origin is not configured yet, complete the streaming test and return to the discovery section after deployment. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, complete [Set Up Agent Gateway](https://docs.api7.ai/ai-gateway/agent-gateway/setup.md) and keep these resources available: * The AISIX gateway, `aisix-a2a-echo` container, and `aisix-a2a` Docker network. * The `AISIX_GATEWAY_CONTAINER`, `AISIX_A2A_URL`, and `AISIX_A2A_KEY` environment variables. * The `echo-agent` registration and caller grant. The setup downloads and compiles the pinned SDK release in the echo-agent container. The client commands reuse the same release and container build cache. ## Verify Streaming[​](#verify-streaming "Direct link to Verify Streaming") The setup guide retains `AISIX_A2A_URL`, which addresses the gateway from the client container. Send a streaming message through that endpoint: ``` docker exec aisix-a2a-echo \ go run github.com/a2aproject/a2a-go/v2/cmd/a2a@v2.5.0 \ send "$AISIX_A2A_URL" "Stream through AISIX" \ --transport jsonrpc \ --stream \ --auth "Bearer $AISIX_A2A_KEY" \ --output json ``` The client prints events as they arrive. The echo agent produces this sequence: 1. A submitted task. 2. A working status update. 3. An artifact update containing `Stream through AISIX`. 4. A completed status update. This verifies that AISIX relays the A2A event stream instead of buffering it into one response. ## Verify Agent-Card Discovery[​](#verify-agent-card-discovery "Direct link to Verify Agent-Card Discovery") The echo server started by the setup guide publishes a dual A2A 0.3 and 1.0 compatibility card. AISIX serves its client-facing form at: ``` /a2a/echo-agent/.well-known/agent-card.json ``` AISIX derives the advertised authority from `Host` and the scheme from `X-Forwarded-Proto`, defaulting to `https` when the forwarded protocol is absent. When a trusted reverse proxy terminates TLS, configure it to supply the gateway's public values. The returned top-level `url` and every entry in `supportedInterfaces` should point back to the client-reachable AISIX `/a2a/echo-agent` endpoint. Once the gateway has a client-reachable public origin, run discovery with the complete card URL: ``` export AISIX_A2A_CARD_URL="https://gateway.example.com/a2a/echo-agent/.well-known/agent-card.json" docker exec aisix-a2a-echo \ go run github.com/a2aproject/a2a-go/v2/cmd/a2a@v2.5.0 \ discover "$AISIX_A2A_CARD_URL" \ --auth "Bearer $AISIX_A2A_KEY" \ --output json ``` Provide the complete nested card URL. The A2A Go SDK treats a URL with a non-root path as the complete card URL, so passing only the agent service URL sends discovery to `/a2a/echo-agent` and returns `405`. Passing only the gateway origin makes the SDK request `/.well-known/agent-card.json`, which returns `404` because AISIX exposes the card under the registered agent's path. To exercise card resolution and the resulting service URL together, pass the same card URL to `send` and omit `--transport`: ``` docker exec aisix-a2a-echo \ go run github.com/a2aproject/a2a-go/v2/cmd/a2a@v2.5.0 \ send "$AISIX_A2A_CARD_URL" "Discover and call through AISIX" \ --auth "Bearer $AISIX_A2A_KEY" \ --output json ``` The client fetches the card through AISIX, selects its JSON-RPC 1.0 interface, and sends the message to the rewritten gateway URL. It never receives the upstream agent URL or credential. ## Troubleshoot Streaming and Discovery[​](#troubleshoot-streaming-and-discovery "Direct link to Troubleshoot Streaming and Discovery") | Symptom | Check | | --------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | The direct client returns `401` | Confirm `--auth` contains `Bearer` followed by the AISIX caller API key. | | The direct client returns `403` | Confirm the caller key's `allowed_agents` grant includes `echo-agent`. | | Discovery returns `404` or `405` | Pass the complete `/a2a/echo-agent/.well-known/agent-card.json` URL. An origin-only URL resolves to the unavailable origin-level card path and returns `404`; the agent service URL is fetched as though it were the card and returns `405`. | | The card advertises the wrong URL, or the client bypasses AISIX or cannot connect | Inspect the top-level `url` and `supportedInterfaces[].url`; each should use the client-reachable AISIX origin and `/a2a/echo-agent` path. If a trusted reverse proxy terminates TLS, configure it to set the public `Host` and `X-Forwarded-Proto` values. AISIX defaults the scheme to `https` when the forwarded protocol is absent. | | Streaming fails after a non-streaming message succeeds | Confirm the upstream card advertises streaming and inspect the gateway and upstream agent logs for the `SendStreamingMessage` request. | ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Control Agent Access](https://docs.api7.ai/ai-gateway/agent-gateway/agent-access-control.md): grant exact agent names, name patterns, or every registered agent. * [Upstream Authentication](https://docs.api7.ai/ai-gateway/agent-gateway/upstream-authentication.md): keep upstream bearer tokens and API keys in AISIX. * [Observability](https://docs.api7.ai/ai-gateway/agent-gateway/observability.md): inspect A2A requests, task outcomes, and stream failures. --- # Rate Limits and Budgets For Agent-to-Agent (A2A) traffic, caller-level limits are configured on the caller API key. A2A calls share the key's request and concurrency limits with model traffic. AISIX Cloud budgets that cover the key also apply, but A2A calls report zero cost and do not increase tracked spend. There is no separate A2A traffic-control resource. Configure rate limits on the caller API key and, in AISIX Cloud, budgets at a scope that covers the key. Then verify the behavior on `/a2a/`. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Complete [Set Up Agent Gateway](https://docs.api7.ai/ai-gateway/agent-gateway/setup.md) for AISIX Cloud or the open-source AISIX gateway. Keep the shell, gateway, and test agent running to verify the limit. ## Where Controls Apply[​](#where-controls-apply "Direct link to Where Controls Apply") The gateway applies traffic controls to each JSON-RPC call at `/a2a/`. Agent-card discovery is not rate-limited. A throttled caller can still fetch the card of an agent its key may access, but it cannot invoke the agent again until the window resets. When a rate limit or budget rejects a call, AISIX returns before contacting the upstream agent and records the rejected call as a [usage event](https://docs.api7.ai/ai-gateway/agent-gateway/observability.md). Because an A2A call resolves no model, model-scoped rate-limit policies do not apply. Request-level controls on the caller API key do apply. ## Applicable Rate Limits[​](#applicable-rate-limits "Direct link to Applicable Rate Limits") A caller API key's `rate_limit` object supports request-rate, token-rate, and concurrency limits. Only request-rate and concurrency limits directly meter A2A calls: | Limit | Applies to A2A calls | Behavior | | -------------------------- | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `rps`, `rpm`, `rph`, `rpd` | Yes | Each `/a2a/` call counts as one request in the matching window. | | `concurrency` | Yes | Each in-flight call holds one permit until it returns. | | `tpm`, `tpd` | No | Estimated A2A token counts are reported but are not added to token windows. An A2A call can still be rejected if the key's model traffic already exhausted a shared token window. | Use a request-rate limit or `concurrency` to control A2A call volume. A token limit alone does not cap A2A calls. For the complete field reference and counter-storage options, see [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md). ## Set a Caller Rate Limit[​](#set-a-caller-rate-limit "Direct link to Set a Caller Rate Limit") The following examples limit the setup guide's caller API key to one request per minute. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Update the existing caller API key: ``` curl -fsS -X PATCH \ "$AISIX_CP/environments/$ENV_ID/api_keys/$API_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "rate_limit": { "rpm": 1 } }' | jq ``` The partial update retains the key's model, MCP, and agent grants. The control plane projects the change to attached gateways automatically. Send `"rate_limit": null` to clear the inline limit. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Add `rate_limit` to the caller entry in `resources.yaml`: resources.yaml (caller rate limit) ``` api_keys: - display_name: quickstart-caller key_env: CALLER_API_KEY allowed_models: - gpt-4o-mini allowed_agents: - echo-agent rate_limit: rpm: 1 ``` Validate the complete file and reload it: ``` docker exec aisix-quickstart \ aisix validate --resources /etc/aisix/resources.yaml docker kill --signal HUP aisix-quickstart ``` Remove `rate_limit` to clear the inline limit. ## Verify the Rate Limit[​](#verify-the-rate-limit "Direct link to Verify the Rate Limit") After the configuration is active, call the agent twice within one minute: ``` for call in 1 2; do curl -sS -o "/tmp/a2a-rate-limit-${call}.json" \ -w "call ${call}: HTTP %{http_code}\n" \ -X POST "$AISIX_PROXY/a2a/echo-agent" \ -H "Authorization: Bearer $AISIX_A2A_KEY" \ -H "Content-Type: application/json" \ -d '{ "jsonrpc": "2.0", "id": "req-rate-limit", "method": "SendMessage", "params": { "message": { "messageId": "msg-rate-limit", "role": "ROLE_USER", "parts": [{"text": "Test the caller rate limit"}] } } }' done ``` The command prints HTTP `200` for the first call and HTTP `429` for the second. The rejected call does not reach the upstream agent. ## Apply AISIX Cloud Budgets[​](#apply-aisix-cloud-budgets "Direct link to Apply AISIX Cloud Budgets") Budgets are configured in AISIX Cloud and enforced by its attached gateways. When a budget that covers the caller API key is exhausted, AISIX rejects that key's A2A calls with a `budget_exceeded` error before contacting the upstream agent. The A2A protocol has no provider-usage block, so AISIX estimates text tokens for telemetry and records `cost_usd` as zero. A2A calls therefore do not add spend to a budget. A budget exhausted by the caller's model traffic still blocks that caller's A2A calls because both use the same caller API key. The open-source AISIX gateway has no budget-management service. Use caller API key request and concurrency limits to control A2A call volume. See [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md) for AISIX Cloud budget targets, rejection behavior, and caching. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now applied caller API key controls to A2A traffic. Use these guides to observe or refine the result: * [Observability](https://docs.api7.ai/ai-gateway/agent-gateway/observability.md): review the usage events and metrics emitted by A2A calls. * [Control Agent Access](https://docs.api7.ai/ai-gateway/agent-gateway/agent-access-control.md): scope caller API keys to specific agents or patterns. * [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md): review all rate-limit fields. * [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md): configure AISIX Cloud budget targets and rejection behavior. --- # Upstream Authentication Each registered Agent-to-Agent (A2A) agent defines whether and how AISIX authenticates to its upstream. AISIX presents the configured credential, if any, when fetching the agent card or forwarding a JSON-RPC call. Clients authenticate to AISIX with a caller API key, which is never forwarded upstream. This separation lets each upstream use the authentication scheme it requires while callers keep one AISIX credential. AISIX Cloud and the open-source AISIX gateway support the same modes through different management paths. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Complete [Set Up Agent Gateway](https://docs.api7.ai/ai-gateway/agent-gateway/setup.md) for AISIX Cloud or the open-source AISIX gateway. To run the optional end-to-end verification, keep the same shell, gateway, test agent, and temporary Docker network running. ## Authentication Modes[​](#authentication-modes "Direct link to Authentication Modes") Choose the mode that matches what the upstream agent expects: | `auth_type` | Required field | Upstream header | | ----------- | -------------- | -------------------------------- | | `none` | None | No credential is sent. | | `bearer` | `secret` | `Authorization: Bearer ` | | `api_key` | `secret` | `x-api-key: ` | `auth_type` defaults to `none`. Leave `secret` unset in that mode. A nonempty `secret` is required for `bearer` and `api_key`. Use HTTPS for a credentialed upstream. When a bearer token or API key is configured on an `http://` URL, the gateway logs a warning because the credential crosses the network in plaintext. ## Configure Upstream Authentication[​](#configure-upstream-authentication "Direct link to Configure Upstream Authentication") The following examples add bearer authentication to the setup guide's registered `echo-agent`. When adapting that entry to a credentialed upstream, also replace its URL. Use `api_key` instead when the upstream expects `x-api-key`. The echo agent itself does not validate credentials. To test credential forwarding locally, skip these examples and continue with [Optional: Verify with a Local Test Proxy](#optional-verify-with-a-local-test-proxy), which supplies its own token and final resource update. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Export the credential, then update the registered agent: ``` export A2A_AGENT_TOKEN="YOUR_UPSTREAM_TOKEN" curl -fsS -X PATCH "$AISIX_CP/a2a_agents/$A2A_AGENT_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @- </dev/null 2>&1 || true docker run -d --name aisix-a2a-auth \ --network aisix-a2a \ -e UPSTREAM_TOKEN="$1" \ caddy:2.11.4-alpine \ sh -c 'caddy run --config /dev/stdin --adapter caddyfile <&1 | grep -q "serving initial configuration" && return sleep 1 done docker logs aisix-a2a-auth >&2 return 1 } start_a2a_auth_proxy "$A2A_AGENT_TOKEN" ``` ### Configure AISIX to Use the Proxy[​](#configure-aisix-to-use-the-proxy "Direct link to Configure AISIX to Use the Proxy") Update the existing `echo-agent` to use `http://aisix-a2a-auth:8082`, bearer authentication, and `A2A_AGENT_TOKEN` as its secret. For AISIX Cloud, update the agent created by the setup guide: ``` curl -fsS -X PATCH "$AISIX_CP/a2a_agents/$A2A_AGENT_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @- <&2 exit 1 fi ``` The request prints HTTP `502`, the `jq` command prints `true`, and the credential check produces no output. AISIX includes the upstream status in a JSON-RPC error but does not proxy the upstream response body. Restore the expected token and remove the temporary response file: ``` start_a2a_auth_proxy "$A2A_AGENT_TOKEN" rm /tmp/aisix-a2a-auth-failure.json ``` Keep `aisix-a2a-auth` running while you use this authenticated local setup. Remove it before running the cleanup commands in the setup guide: ``` docker rm -f aisix-a2a-auth ``` ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now configured how AISIX authenticates to an upstream agent. Use these guides to control callers and traffic: * [Control Agent Access](https://docs.api7.ai/ai-gateway/agent-gateway/agent-access-control.md): scope caller API keys to specific agents or patterns. * [Rate Limits and Budgets](https://docs.api7.ai/ai-gateway/agent-gateway/traffic-controls.md): apply request and concurrency limits, and use AISIX Cloud budgets. * [Observability](https://docs.api7.ai/ai-gateway/agent-gateway/observability.md): review the usage events and metrics emitted by A2A calls. --- # Admin Tokens Admin tokens let automation call the AISIX Cloud Admin API without a browser session. They are intended for organization-level workflows such as CI pipelines, infrastructure automation, or internal tools that manage control-plane resources. They are not caller API keys for gateway traffic or upstream provider keys for model access. An admin token belongs to one organization. The control plane uses that organization binding when it authenticates the request, so API calls do not need a separate organization header. ## Create an Admin Token[​](#create-an-admin-token "Direct link to Create an Admin Token") Only organization owners can create or revoke admin tokens. 1. Open **Admin tokens**. 2. Select **New token**. 3. Enter a unique name. 4. Choose an expiration period, or select **Never**. 5. Select one or more scopes. 6. Create the token and copy the plaintext value before leaving the one-time view. Generated admin tokens use the `aisix_pat_` prefix. After displaying the plaintext value once, the control plane stores only its SHA-256 digest. If you lose the plaintext value, revoke the token and create a replacement. ## Choose Scopes[​](#choose-scopes "Direct link to Choose Scopes") Admin tokens support the following scopes: | Scope | Access | | ------- | ---------------------------------------------------------------------------------------------------------------------------------- | | `read` | Read-only access to AISIX Cloud Admin API routes. | | `write` | Read and write access to AISIX Cloud Admin API routes. | | `scim` | Access only to the [SCIM directory sync](https://docs.api7.ai/ai-gateway/cloud/scim-directory-sync.md) endpoints under `/scim/v2`. | Use `read` for inventory or reporting jobs. Use `write` only when automation needs to create, update, or delete control-plane resources. SCIM tokens follow a separate creation flow: generate them in **Settings**, under **Directory sync (SCIM)**. The `scim` scope cannot be combined with other scopes, and these tokens are rejected on non-SCIM routes. This limits the identity provider credential to directory provisioning. ## Use an Admin Token[​](#use-an-admin-token "Direct link to Use an Admin Token") Set the AISIX Cloud Admin API base URL and send the token in the `Authorization` header as a bearer token: ``` # AISIX_CP includes /api and has no trailing slash. # The local On-Premises quickstart uses http://localhost:8080/api. export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_BASE_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" curl -sS "${AISIX_CP}/environments" \ -H "Authorization: Bearer ${AISIX_TOKEN}" ``` Use the [AISIX Cloud Admin API Reference](https://docs.api7.ai/ai-gateway/reference/cloud-admin-api/.md) for routes covered by the public OpenAPI specification. Workflow guides may also provide task-specific examples. For example, [Members](https://docs.api7.ai/ai-gateway/cloud/members.md#use-the-api) shows how to create a member for API key ownership. ## Rotate or Revoke Tokens[​](#rotate-or-revoke-tokens "Direct link to Rotate or Revoke Tokens") To rotate an admin token, create a replacement, update the automation that uses it, verify that the automation succeeds, and then revoke the old token. Revocation takes effect immediately. Requests that use the revoked token receive authentication errors. ## Next Steps[​](#next-steps "Direct link to Next Steps") After creating an admin token with the `write` scope, continue with [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md) to attach the runtime that will receive environment resources and serve AI traffic. To configure organization access instead, continue with [Members](https://docs.api7.ai/ai-gateway/cloud/members.md). For identity-provider provisioning, see [SCIM Directory Sync](https://docs.api7.ai/ai-gateway/cloud/scim-directory-sync.md). --- # Connect an AISIX Gateway In an AISIX Cloud deployment, the AISIX gateway runs as the data plane in your infrastructure. Add a gateway to an environment so it can receive configuration, report heartbeats and telemetry, and request budget decisions. Live AI requests go directly to the gateway and do not pass through the control plane. The gateway initiates outbound management connections authenticated with a certificate bundle issued for one AISIX Cloud environment. ## How the Connection Works[​](#how-the-connection-works "Direct link to How the Connection Works") The certificate bundle contains a client certificate, private key, and CA certificate. The client certificate identifies the gateway and the environment it serves. AISIX uses the same certificate identity for the configuration watch and the gateway management APIs. The configured management endpoint is also the base for heartbeat, telemetry, and budget-check requests. The gateway derives the configuration-store endpoint from that URL unless the deployment supplies a separate endpoint. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before connecting the gateway, prepare: * An AISIX Cloud environment for the gateway. * Outbound network access from the gateway host or cluster to the gateway management endpoint. * A secret store for the client certificate and private key. * Writable local storage for gateway identity, materialized mTLS files, and the configuration snapshot cache. AISIX stores these under `/var/lib/aisix` by default. For a container or Kubernetes deployment, mount `/var/lib/aisix` as persistent storage when a recreated gateway must reuse its identity and latest accepted configuration without first reconnecting to the control plane. Give each gateway instance its own writable volume. ## Issue a Gateway Certificate[​](#issue-a-gateway-certificate "Direct link to Issue a Gateway Certificate") 1. In the AISIX Cloud dashboard, select the environment the gateway should serve. 2. Open **Data planes**. 3. Select the certificate validity period and optionally enter a hostname. 4. Select **Issue certificate**. 5. Copy the generated install snippet for the target deployment. caution The private key appears only in the certificate-issuance response. Copy the bundle immediately, store it in your deployment secret system, and do not commit or share the generated snippet. The dashboard provides Docker, Docker Compose, Kubernetes (Helm), and systemd installation tabs. The systemd tab generates the service and credential configuration, but it does not install the executable. ## Deploy on Kubernetes[​](#deploy-on-kubernetes "Direct link to Deploy on Kubernetes") Use the **Kubernetes (Helm)** tab to install the `api7/aisix` chart. The chart is the only Kubernetes path the dashboard generates, because it carries the settings a gateway needs to stop without dropping traffic — a `preStop` pause and a termination grace period sized to cover the drain — which a hand-written manifest leaves at the Kubernetes defaults. For a chart-managed production deployment with autoscaling, disruption budgets, and Prometheus integration, follow [Deploy AISIX Gateways on Kubernetes](https://docs.api7.ai/ai-gateway/cloud/kubernetes.md). ## Start the Gateway[​](#start-the-gateway "Direct link to Start the Gateway") Run the generated instructions in the target runtime environment. Before using the systemd instructions, confirm that a compatible `aisix-dp` executable is installed at `/usr/local/bin/aisix-dp`, as expected by the generated unit. The generated configuration binds the proxy listener to port `3000` and the metrics and status listener to port `9090` by default. See the [Port Reference](https://docs.api7.ai/ai-gateway/reference/ports.md) for listener exposure and outbound control-plane connectivity. Make `/var/lib/aisix` writable by the same user that runs the gateway. Mounting only `/var/lib/aisix/mtls` does not preserve the gateway identity file and snapshot cache. ## Verify the Connection[​](#verify-the-connection "Direct link to Verify the Connection") Verify the initial connection before configuring traffic: 1. Confirm that the process starts without certificate, trust-chain, or configuration-store connection errors. 2. In the environment's **Data planes** view, confirm that the gateway appears with a recent heartbeat. 3. Confirm that the reported hostname and certificate ID identify the intended deployment and environment. After configuring a provider, model alias, and caller API key, verify the end-to-end data path. Confirm that projected resources reach the gateway, a live request succeeds, and its usage or telemetry appears in the control plane. If the gateway does not appear, check the management endpoint, certificate bundle, trust root, file permissions, state directory, and outbound network access. A healthy heartbeat confirms the management API path, but it does not prove that a resource change has reached every gateway instance. Use [Resource Projection](https://docs.api7.ai/ai-gateway/cloud/resource-projection.md) after saving the first provider key, model, and caller API key. ## AISIX Cloud Connection Configuration[​](#aisix-cloud-connection-configuration "Direct link to AISIX Cloud Connection Configuration") Provide the certificate, key, and CA together. Use file-path variables when the bundle is mounted as files, or inline variables when the deployment system injects PEM content. Do not configure both forms for the same certificate role. | Configure | Use | | ------------------------------------- | ------------------------------------ | | Gateway management base URL | `AISIX_MANAGED__CP_BASE_URL` | | Separate configuration-store endpoint | `AISIX_MANAGED__CP_ETCD_ENDPOINT` | | Certificate file | `AISIX_MANAGED__CP_CERT_FILE` | | Private key file | `AISIX_MANAGED__CP_KEY_FILE` | | CA certificate file | `AISIX_MANAGED__CP_CA_FILE` | | Inline certificate | `AISIX_MANAGED__CP_CERT_PEM` | | Inline private key | `AISIX_MANAGED__CP_KEY_PEM` | | Inline CA certificate | `AISIX_MANAGED__CP_CA_PEM` | | Materialized mTLS directory | `AISIX_MANAGED__MTLS_DIR` | | Gateway identity file | `AISIX_MANAGED__DP_ID_FILE` | | Snapshot cache file | `AISIX_MANAGED__SNAPSHOT_CACHE_PATH` | Set `AISIX_MANAGED__CP_ETCD_ENDPOINT` only when the control plane provides a configuration-store endpoint that differs from the gateway management base URL. Specify this endpoint as `host:port` without a URL scheme. The default state paths are `/var/lib/aisix/mtls`, `/var/lib/aisix/dp_id`, and `/var/lib/aisix/config_cache.json`. A single per-instance mount at `/var/lib/aisix` covers all three. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Choose a Provider Upstream](https://docs.api7.ai/ai-gateway/providers/overview.md). Each provider guide creates a provider key, model alias, and caller API key, then verifies the configuration with a live request through the connected gateway. If a saved resource does not affect live traffic as expected, use [Resource Projection](https://docs.api7.ai/ai-gateway/cloud/resource-projection.md) to trace how environment configuration reaches the gateway. --- # Roles and Custom Roles Every organization member has one role that controls what they can do in the AISIX Cloud dashboard and AISIX Cloud Admin API. AISIX Cloud provides three built-in roles for common access patterns. Organizations that need finer control can define custom roles as named permission sets over the same resource vocabulary enforced by the API. An organization role sets the member's baseline across every environment. When a member needs additional access in one environment, an owner can add an environment access grant without changing that baseline elsewhere. Roles govern who can configure the gateway through the control plane. API key settings separately determine which models and tools a key can use and which budgets apply to its traffic. ## Built-in Roles[​](#built-in-roles "Direct link to Built-in Roles") | Role | Access | | -------- | ------------------------------------------------------------------------------------------------------------------------- | | `owner` | Full control, including billing, member role changes, member removal, and admin token management. | | `admin` | Read and write access to every resource. Member role changes, member removal, and admin token creation remain owner-only. | | `member` | Read-only access to the organization's resources. Audit events are not readable. | The **Roles** page lists each role and its enforced permissions. The page uses the same permission catalog that the API checks for every request. ## Custom Roles[​](#custom-roles "Direct link to Custom Roles") A custom role is a named set of `read` and `write` permissions on resources such as environments, models, API keys, guardrails, budgets, and audit events. It replaces the member's built-in baseline instead of extending it. A role that grants only `read` access to `environments`, for example, cannot list teams or view usage. Typical uses include: * An `auditor` role that reads the audit trail and usage but configures nothing. * A `gateway-operator` role that manages models, provider keys, and guardrails but cannot change members or billing. * A read-mostly role with write access to one resource type, such as budgets. ### Create and Assign[​](#create-and-assign "Direct link to Create and Assign") 1. As an organization admin or owner, open **Roles** and select **New role**. 2. Enter a permanent lowercase name such as `auditor`. Members reference the role by name, so create a new role when you need a different name. 3. Select the permissions the role grants and save it. 4. As an organization owner, assign the role on the **Members** page. For directory-managed access, a custom role can serve as the default role or the target of a SCIM group-to-role mapping. See [directory sync](https://docs.api7.ai/ai-gateway/cloud/scim-directory-sync.md). To change a custom role's description or permissions, open **Roles** and select **Edit**. To remove it, select **Delete** after clearing the references described in [Rules and Limits](#rules-and-limits). Built-in roles cannot be edited or deleted. ## Environment-Scoped Access[​](#environment-scoped-access "Direct link to Environment-Scoped Access") A member's organization role applies across the whole organization. An environment access grant adds another role inside one environment. It can extend but cannot narrow the member's organization role. For example, a member can keep read-only organization access while receiving an `admin` grant for the production environment. The grant adds write access to resources inside production; it does not remove the member's organization-level access elsewhere. Use a custom organization role when you need to reduce the baseline permissions that apply across the organization. * Effective access to resources inside an environment combines the permissions from the organization role and any grant for that environment. * Grants have no effect outside their environment. The organization role alone governs resources in every other environment and organization-level resources such as members, teams, settings, billing, and custom roles. * Grants cover resources inside the environment, not the environment object itself. An environment-scoped admin cannot rename or delete the environment. * Grants can target `admin`, `member`, or any custom role. They cannot target `owner`. * Owners already have full access, so they cannot receive grants. Only owners can edit a member's grants. * Deleting an environment removes the grants that pointed into it. Manage grants on the **Members** page: expand **Environment access** on a member row, add, change, or remove environment and role pairs, and save. ## Use the API[​](#use-the-api "Direct link to Use the API") Use an [admin token](https://docs.api7.ai/ai-gateway/cloud/admin-tokens.md) with `read` scope to list roles and environment access grants. A token with `write` scope can create, update, or delete roles. Only an organization owner can assign roles or replace environment access grants. Set the AISIX Cloud Admin API base URL and token before running the examples: ``` # AISIX_CP includes /api and has no trailing slash. # The local On-Premises quickstart uses http://localhost:8080/api. export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_BASE_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" ``` ### Create and Assign a Custom Role[​](#create-and-assign-a-custom-role "Direct link to Create and Assign a Custom Role") Create an organization-scoped role with the permissions it should grant: ``` curl -sS -X POST "${AISIX_CP}/roles" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{ "name": "auditor", "description": "Read-only audit access", "permissions": [ { "action": "read", "resource": "audit" }, { "action": "read", "resource": "usage" } ] }' ``` An owner can assign the role to a member. Use the member's `user_id` from `GET /members`: ``` export USER_ID="2c7d6e5f-4a3b-4c2d-8e1f-9a0b1c2d3e4f" curl -sS -X PATCH "${AISIX_CP}/members/${USER_ID}" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{"role": "auditor"}' ``` ### Update or Delete a Custom Role[​](#update-or-delete-a-custom-role "Direct link to Update or Delete a Custom Role") Supplying `permissions` replaces the role's complete permission set: ``` curl -sS -X PATCH "${AISIX_CP}/roles/auditor" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{ "permissions": [ { "action": "read", "resource": "audit" } ] }' ``` After clearing every reference described in [Rules and Limits](#rules-and-limits), delete the role: ``` curl -sS -X DELETE "${AISIX_CP}/roles/auditor" \ -H "Authorization: Bearer ${AISIX_TOKEN}" ``` Deleting a role that is still referenced returns `409 ROLE_IN_USE`. Member assignments and environment grants can be cleared through the AISIX Cloud Admin API. Pending invitations and directory sync references must currently be cleared in the dashboard. ### Manage Environment Access Grants[​](#manage-environment-access-grants "Direct link to Manage Environment Access Grants") Environment access routes use the membership `id` returned by `GET /members`, not the member's `user_id`. List the member's current grants before replacing them: ``` export MEMBER_ID="8f3b2a1c-9d4e-4f6a-b7c8-1e2d3f4a5b6c" curl -sS "${AISIX_CP}/members/${MEMBER_ID}/role_bindings" \ -H "Authorization: Bearer ${AISIX_TOKEN}" ``` An owner can replace the complete grant set. Each environment can appear at most once: ``` curl -sS -X PUT "${AISIX_CP}/members/${MEMBER_ID}/role_bindings" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{ "bindings": [ { "env_id": "6b1c2c1e-0000-4000-8000-000000000002", "role": "admin" } ] }' ``` Send an empty `bindings` array to remove every environment access grant. Changes may take up to 30 seconds to propagate across control-plane replicas. See the [AISIX Cloud Admin API Reference](https://docs.api7.ai/ai-gateway/reference/cloud-admin-api/.md) for response schemas and error details. ## Rules and Limits[​](#rules-and-limits "Direct link to Rules and Limits") * Custom role definitions can be created, edited, and deleted by admins and owners. Assigning any role to a member stays owner-only. * A custom role can never grant more than the built-in `admin` role holds: owner-only operations cannot be granted, and directory sync can never assign `owner`. * A custom role cannot be deleted while it is assigned to a member or pending invitation, used as the directory sync default role or the target of a group-to-role mapping, or used in an environment access grant. * Before deleting a custom role, reassign members, revoke pending invitations, clear any references under **Default role** and **Group → role mappings** in **Directory sync (SCIM)**, and remove the role from **Environment access** grants. * Permission changes take effect for every member holding the role, but may take up to 30 seconds to propagate across control-plane replicas. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [SCIM Directory Sync](https://docs.api7.ai/ai-gateway/cloud/scim-directory-sync.md) when your identity provider should manage member and role assignments. To connect the gateway that serves environment resources, see [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md). --- # High Availability In AISIX Cloud, applications send requests through AISIX gateways in your runtime environment while a separate control plane manages those gateways. You operate the control plane in On-Premises; API7 operates it in Hybrid Cloud. This separation allows the traffic and management layers to use independent availability strategies. A highly available deployment must account for the application traffic path, AISIX gateways, upstream services, and each gateway's connection to the AISIX Cloud control plane. ## High-Availability Architecture[​](#high-availability-architecture "Direct link to High-Availability Architecture") The diagram below illustrates an active-active reference pattern that keeps live traffic independent from the AISIX Cloud control-plane path. The AISIX Cloud control plane is not a network hop between applications and upstream services. Choose the number of deployments, gateway instances, and traffic distribution layers according to your availability requirements. In this pattern, AISIX gateways run in two active deployments across independent failure domains. A global load balancer steers traffic between the deployments, and a deployment load balancer distributes traffic across multiple gateway instances. Both deployments receive configuration from the same AISIX Cloud environment. Each deployment uses its own gateway certificate, and instances that share a certificate report distinct runtime instances to the control plane. ![AISIX high-availability architecture with active gateway deployments, gateway-initiated management connections, and a shared AISIX Cloud control plane](https://static.api7.ai/uploads/2026/07/23/CRaBZgpH_aisix-cloud-ha.svg) The control plane exposes separate operator and gateway management endpoints. In this reference pattern, the Control Plane API, dashboard, and Data Plane Manager run as redundant service replicas. Shared state is provided by a replicated PostgreSQL deployment behind a stable endpoint. These elements describe an HA deployment pattern, not the exact topology of the API7-hosted AISIX Cloud control plane. The diagram does not prescribe a particular replication or failover implementation. ## AISIX Cloud Control-Plane Components[​](#aisix-cloud-control-plane-components "Direct link to AISIX Cloud Control-Plane Components") The AISIX Cloud control plane separates user management, gateway management, and shared state. | Component | Role in the architecture | | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | AISIX Cloud control-plane endpoint | Provides the stable public entry point and routes requests to the Control Plane API. | | Control Plane API | Handles AISIX Cloud Admin API operations and reverse-proxies browser requests to the dashboard. It manages organizations, environments, resources, certificates, usage, and budgets. | | Dashboard | Provides the browser interface and authentication workflows behind the Control Plane API. | | Gateway management endpoint | Accepts gateway-initiated mTLS connections without requiring inbound access to gateway hosts. | | Data Plane Manager | Serves projected configuration and receives heartbeats, usage telemetry, and AISIX Cloud budget checks from gateways. | | PostgreSQL HA cluster | Stores shared control-plane state. In this reference topology, replication and failover operate behind a stable service endpoint. | The responsibility boundary depends on the control-plane deployment option. For an On-Premises HA control plane, use Helm. Run redundant replicas of the Control Plane API, Data Plane Manager, and dashboard across failure domains, and connect them to an external HA PostgreSQL endpoint. The bundled PostgreSQL chart does not provide the replicated database shown in this reference pattern by default. The packaged Docker Compose deployments are single-host and do not support this topology. In Hybrid Cloud, API7 operates the control plane. See [On-Premises Configuration](https://docs.api7.ai/ai-gateway/reference/on-premises-configuration.md). ## Availability Across the Request Path[​](#availability-across-the-request-path "Direct link to Availability Across the Request Path") High availability spans the runtime environment, the AISIX Cloud control plane, and upstream services. | Area | Availability requirement | | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Application traffic | Provide a stable gateway endpoint and health-check each deployment from the global traffic layer. Use the proxy listener's [`/readyz` endpoint](https://docs.api7.ai/ai-gateway/deployment/health-checks.md#traffic-readiness) for traffic eligibility: it withdraws an instance that is draining or has no configuration yet, and keeps a running instance eligible for as long as it can serve the configuration it holds. | | AISIX gateways | Operate redundant instances across independent failure domains. Persist each instance's state directory when it must recover cached configuration after a restart. | | AISIX Cloud control plane | For On-Premises, run redundant control-plane services across failure domains and use an external highly available PostgreSQL database. API7 operates the management services in Hybrid Cloud. | | Upstream services | Configure model routes with retries and multiple targets when requests must survive a model or provider outage. Deploy each MCP server and A2A agent behind a resilient service endpoint because AISIX does not automatically select an alternate registered service. | ## Failure Behavior[​](#failure-behavior "Direct link to Failure Behavior") The live traffic path and management path fail independently. | Failure | Expected behavior | | ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | One gateway instance fails | The load-balancing layer directs new requests to healthy instances in that deployment. | | One gateway deployment fails | The global load balancer directs new requests to another healthy deployment. | | Gateway-to-control-plane connectivity is interrupted | A running gateway can continue from the latest accepted configuration held in memory if traffic continues to reach it. A gateway that restarts during the outage can recover from a valid persisted snapshot; without one, it cannot receive traffic until it obtains configuration. New resource projections stop, and heartbeat reporting resumes after connectivity recovers. Alert on configuration freshness through `/status/config` or the `aisix_config_*` metrics — an interrupted watch affects every instance at once, so it is an operator signal rather than a load-balancer one. AISIX Cloud budget checks follow their configured failure behavior, and failed telemetry batches sent to the control plane can be dropped. | | A model service fails | AISIX can retry or use another model target only when routing or failover has been configured. | | An MCP server or A2A agent fails | Requests to that service fail until its configured endpoint recovers. AISIX does not automatically select another registered MCP server or A2A agent. | | An external guardrail service fails | AISIX follows the configured input and output failure behavior. Fail-open policies allow unscanned traffic, while fail-closed policies block it. | ## Deployment Guidance[​](#deployment-guidance "Direct link to Deployment Guidance") For the active-active pattern shown above, use the following practices for the customer-operated traffic layer: * Run at least two active gateway deployments behind a global load balancer with active health checks. * Place each deployment in a separate failure domain. Use a redundant deployment load balancer to distribute traffic across multiple gateway instances. Use `/readyz` as the traffic-eligibility probe: it prevents traffic before the gateway has usable configuration and allows an instance restored from a valid cached snapshot to serve while reconnecting to the control plane. Track configuration freshness as an operator signal instead of a load-balancer one, since a stalled watch affects every instance at once. Configure probe intervals, failure and recovery thresholds, and connection draining to avoid traffic flapping and interrupted in-flight requests. * Issue a separate gateway certificate for each deployment so the deployments have independent credential lifecycles. Instances within a deployment can share that deployment's certificate bundle. Protect each private key through your deployment secret system. * Give each instance its own persistent state directory when it must recover its certificate bundle, deployment identity, and latest accepted configuration after a restart. Do not share a writable state directory between instances. The [`api7/aisix` Helm chart](https://docs.api7.ai/ai-gateway/cloud/kubernetes.md) uses per-pod ephemeral state by default, so a restarted chart-managed pod must reconnect and download its configuration before receiving traffic. * Ensure every gateway instance can initiate mTLS connections to the AISIX Cloud control-plane endpoint. * Use Redis for rate-limit counters and cache entries that must be shared across instances. In-memory counters and cache entries remain local to one gateway process. When Redis is an availability dependency, use Redis Cluster, Redis Sentinel, or a managed Redis service exposed through a compatible endpoint. * Configure model routing and failover separately. Gateway redundancy does not make a single model or provider highly available. Deploy MCP servers and A2A agents behind resilient endpoints because AISIX does not fail over between registered services. * Monitor `/readyz`, gateway heartbeat freshness, applied configuration status, rejected resources, and exporter health. A successful live request confirms the traffic path, but it does not prove that management or telemetry paths are healthy. ## Network and Security Boundaries[​](#network-and-security-boundaries "Direct link to Network and Security Boundaries") Keep the traffic and management paths distinct: * Expose the gateway load balancer to applications over HTTPS. * Preserve TLS from the gateway to upstream services according to each provider or service configuration. * Permit gateway-initiated mTLS connections from every gateway instance to the AISIX Cloud control plane. Gateway hosts do not require inbound control-plane access. * Restrict metrics and health endpoints to the monitoring network used by the load balancer and platform operators. Live AI requests pass through AISIX gateways in your runtime environment; they do not pass through the AISIX Cloud control plane. Usage telemetry sent to the AISIX Cloud control plane excludes prompt and response bodies, but it includes request, caller, routing, usage, and error metadata. External observability exporters can include request and response content when content capture is enabled, so apply the same data-handling controls used for other logging and tracing systems. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Offline Resilience](https://docs.api7.ai/ai-gateway/cloud/offline-resilience.md) for behavior during temporary control-plane connectivity loss. For upstream continuity, see [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md). --- # Deploy AISIX Gateways on Kubernetes Use the `api7/aisix` Helm chart to deploy, expose, and scale AISIX gateways in your Kubernetes cluster. The gateways serve live AI traffic in your environment and connect to an existing AISIX Cloud control plane for configuration. The chart requires the data-plane manager endpoint and a gateway certificate bundle issued for the target environment. Once connected, the gateways receive models, caller API keys, and policies from the control plane. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A Kubernetes cluster with `kubectl` configured to access it and Helm 3 installed. * Access to an AISIX Cloud environment where you can issue a gateway certificate. * Network access from the cluster to the AISIX Cloud data-plane manager endpoint. If you use On-Premises, complete [On-Premises Installation](https://docs.api7.ai/ai-gateway/on-premises/deployment.md) first. ## Install the Chart[​](#install-the-chart "Direct link to Install the Chart") In the console, open the target environment's **Data planes** view and issue a gateway certificate. The **Kubernetes (Helm)** tab provides the data-plane manager endpoint and certificate bundle used below. The example also pins the initial deployment to one replica. The private key is shown only once. Store the bundle in a Secret so the private key stays out of your values file: ``` kubectl create namespace aisix kubectl -n aisix create secret generic aisix-gateway-certificate \ --from-file=cert.pem=./cert.pem \ --from-file=key.pem=./key.pem \ --from-file=ca.pem=./ca.pem ``` Install the chart against the data-plane manager endpoint from the same view: ``` helm repo add api7 https://charts.api7.ai helm repo update helm install aisix api7/aisix --namespace aisix --version 1.0.0 \ --set controlPlane.baseURL=https://dp-manager.example.com:7944 \ --set controlPlane.certificate.existingSecret=aisix-gateway-certificate \ --set replicaCount=1 ``` The gateway appears in the environment's **Data planes** view once its first heartbeat lands. The chart defaults to two replicas, but this initial command starts one because rate-limit counters use per-replica memory until you configure shared Redis. The console Helm snippet omits `replicaCount`, so installing that snippet unchanged starts two replicas. Add `--set replicaCount=1` until Redis is configured, as shown above. [Share Rate-Limit Counters across Replicas](#share-rate-limit-counters-across-replicas) before increasing the replica count or enabling an autoscaler. Each replica registers as its own instance, and they all share the one certificate. To see every value the chart accepts: ``` helm show values api7/aisix --version 1.0.0 ``` The chart source and package are published in the [`aisix-1.0.0` Helm chart release](https://github.com/api7/api7-helm-chart/releases/tag/aisix-1.0.0). ## Expose the Gateway[​](#expose-the-gateway "Direct link to Expose the Gateway") The chart creates a `ClusterIP` Service for the proxy by default. To publish it through a cloud load balancer: ``` service: type: LoadBalancer port: 80 # Preserve the client source IP, which per-model IP allowlists match on. externalTrafficPolicy: Local ``` The gateway binds port 3000 inside the container by default. A Service can expose port `80` or `443` without making the process bind a privileged container port. Bind a port below `1024` inside the container only when the gateway must listen on that port directly. The published image runs as non-root UID `10001`, and the gateway binary carries the effective `CAP_NET_BIND_SERVICE` file capability: ``` containerPorts: proxy: 80 ``` If you customize the rendered Pod to drop all capabilities, add `NET_BIND_SERVICE` back to the AISIX container: ``` spec: containers: - name: aisix securityContext: capabilities: drop: ["ALL"] add: ["NET_BIND_SERVICE"] ``` The Kubernetes Restricted Pod Security Standard allows this capability. Because the binary's file capability has the effective bit set, the container can fail with `exec: Operation not permitted` when the runtime prevents the capability from being granted. Use `hostNetwork` or `hostPort` only when the gateway must bind directly on a node and the cluster policy allows it. These options introduce node-port conflicts and reduce network isolation; the Baseline and Restricted Pod Security Standards also disallow them. See [Network and Security](https://docs.api7.ai/ai-gateway/deployment/network-and-security.md) for the full exposure and credential model. ## Share Rate-Limit Counters across Replicas[​](#share-rate-limit-counters-across-replicas "Direct link to Share Rate-Limit Counters across Replicas") Rate-limit counters live in each gateway's own memory by default, so *N* replicas enforce *N* times every configured request, token, and concurrency limit. Before running more than one replica — including any replica an autoscaler adds — point every replica at one Redis: ``` rateLimit: backend: redis redis: url: redis://redis.default.svc:6379 ``` Use `rateLimit.redis.existingSecret` instead when the connection URL carries a password. After every replica points to the same Redis deployment, increase `replicaCount` or enable one of the autoscaling options below. If the deployment does not use request, token, or concurrency limits, you can deliberately accept the per-replica memory backend instead. ## Scale on CPU or Memory[​](#scale-on-cpu-or-memory "Direct link to Scale on CPU or Memory") `autoscaling` creates a `HorizontalPodAutoscaler` for the gateway Deployment: ``` autoscaling: enabled: true minReplicas: 2 maxReplicas: 20 targetCPUUtilizationPercentage: 70 ``` The targets are percentages of the pod's resource *requests*, so the chart sets a CPU request by default. It deliberately sets no CPU limit: throttling adds tail latency to a proxy and suppresses the signal the autoscaler reads. Scaling on CPU requires `metrics-server` in the cluster. On Linux, `proxy.workers` defaults to the CPU parallelism available to the gateway process. A Kubernetes CPU request does not constrain that value, but a CPU limit does. Because the chart sets no CPU limit, set `AISIX_PROXY__WORKERS` explicitly when each replica needs a stable worker count, and keep it within the CPU capacity planned for the replica. See [Thread-per-Core Workers](https://docs.api7.ai/ai-gateway/deployment/thread-per-core-workers.md) for worker configuration and sizing considerations. When autoscaling is enabled, the Deployment omits `spec.replicas` so that a later `helm upgrade` cannot reset the replica count the autoscaler chose. The `replicaCount` value is ignored from then on. Pass any `behavior` policy through unchanged, for example to make scale-down gentler than the Kubernetes default: ``` autoscaling: behavior: scaleDown: stabilizationWindowSeconds: 300 policies: - type: Pods value: 1 periodSeconds: 60 ``` Use `autoscaling.extraMetrics` for Pods, Object, or External metrics such as a series exposed through a Prometheus adapter. ## Scale on Request Load with KEDA[​](#scale-on-request-load-with-keda "Direct link to Scale on Request Load with KEDA") CPU is a proxy for load. To scale on the gateway's own traffic instead, use [KEDA](https://keda.sh) and a Prometheus query over the [gateway metrics](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md). Before enabling the example, prepare these cluster components: * Install KEDA, including the `ScaledObject` CRD and controller. * Provide a Prometheus server that can query the gateway metrics. * If the chart should create the scrape configuration shown below, install Prometheus Operator or another controller that consumes `ServiceMonitor` objects. Installing only the CRD lets Helm create the object, but nothing turns it into scrape configuration. Set `metrics.serviceMonitor.labels` when the Prometheus instance selects ServiceMonitors by label. If Prometheus discovers the metrics Service another way, leave `metrics.serviceMonitor.enabled` set to `false` and omit that block. Confirm the CRDs required by the values you enable: ``` kubectl get crd scaledobjects.keda.sh # Required only when metrics.serviceMonitor.enabled is true. kubectl get crd servicemonitors.monitoring.coreos.com ``` Configure the chart after those prerequisites are available: ``` metrics: serviceMonitor: enabled: true keda: enabled: true minReplicas: 2 maxReplicas: 20 pollingInterval: 15 cooldownPeriod: 300 triggers: - type: prometheus metadata: serverAddress: http://prometheus.monitoring.svc:9090 query: sum(rate(aisix_llm_requests_total[2m])) threshold: "100" ``` `aisix_llm_requests_total` counts model-inference requests, such as `/v1/chat/completions`. It does not include MCP or A2A calls. Use [`aisix_proxy_requests_total`](https://docs.api7.ai/ai-gateway/reference/metrics.md#request-metrics) if that is the load you want to scale on. `autoscaling` and `keda` are mutually exclusive. Enabling both fails the Helm render rather than letting two controllers write `spec.replicas`. ## What Happens during a Scaling Event[​](#what-happens-during-a-scaling-event "Direct link to What Happens during a Scaling Event") A replica the autoscaler adds does not receive traffic before it can serve it. The chart's readiness probe reports `503` until the gateway has applied configuration from the control plane, so Kubernetes keeps the new pod out of the Service endpoints until then. See [Health Checks](https://docs.api7.ai/ai-gateway/deployment/health-checks.md) for the readiness contract. When a replica scales down, Kubernetes starts removing it from Service endpoints and terminating the pod concurrently. Two chart values protect in-flight requests while that happens. `preStopSleepSeconds` pauses before the container receives `SIGTERM`, so endpoint removal reaches every node before the gateway begins shutting down. That covers a balancer watching the Kubernetes API; one that polls a health check instead sees a still-ready pod throughout the pause and is covered by the gateway's own drain window, `shutdown.min_drain_secs`. `terminationGracePeriodSeconds` caps the whole sequence — the pause, the drain window, and the in-flight drain that follows — because the gateway drains without a deadline of its own. The Kubernetes default of 30 seconds would cut a streaming response, so the chart ships a longer one. Both values ship with defaults sized for that sequence; `helm show values api7/aisix --version 1.0.0` reports the defaults for this release. Raise `terminationGracePeriodSeconds` if your workloads stream for longer than the budget it leaves once the pause and the drain window have run. To watch scaling decisions and the metrics behind them: ``` kubectl -n aisix get hpa aisix --watch kubectl -n aisix describe hpa aisix ``` ## Survive Node Disruption[​](#survive-node-disruption "Direct link to Survive Node Disruption") A `PodDisruptionBudget` keeps voluntary disruptions — node drains, cluster upgrades — from taking every gateway down at once. Spread constraints keep replicas out of a single failure domain: ``` podDisruptionBudget: enabled: true minAvailable: 50% topologySpreadConstraints: - maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway labelSelector: matchLabels: app.kubernetes.io/name: aisix ``` The chart uses an `emptyDir` volume for `/var/lib/aisix`, so its state is ephemeral and scoped to one pod. A restarted pod therefore re-registers and downloads its configuration from the control plane before receiving traffic. Keep multiple replicas across failure domains so an existing replica can continue serving during a restart. If a gateway must restart from cached configuration while the control plane is unavailable, use a customized workload that gives each replica its own persistent state directory. See [High Availability](https://docs.api7.ai/ai-gateway/cloud/high-availability.md) for the wider deployment pattern. ## Set Any Other Gateway Configuration[​](#set-any-other-gateway-configuration "Direct link to Set Any Other Gateway Configuration") The control plane owns dynamic resources, and the chart owns the pod. For [startup configuration](https://docs.api7.ai/ai-gateway/deployment/startup-configuration.md) the chart does not surface as a value, set the environment variable directly — every configuration field is reachable as `AISIX_
__`: ``` extraEnvVars: - name: AISIX_OBSERVABILITY__LOG_LEVEL value: "debug" - name: AISIX_UPSTREAM__POOL_MAX_IDLE_PER_HOST value: "32" ``` See [Environment Variables](https://docs.api7.ai/ai-gateway/reference/environment-variables.md) for the naming rules. --- # Logging and Auditing The AISIX Cloud control plane gives operators two evidence trails: request logs for AISIX gateway traffic and an audit log for control-plane state changes. Together, they help teams understand what happened to a request and who changed the resources that affect traffic. Use request logs to investigate a specific AISIX gateway request. Use the audit log to investigate configuration, access, and other control-plane changes. ## Request Logs[​](#request-logs "Direct link to Request Logs") Request logs are built from AISIX gateway telemetry. They show individual request outcomes, including request time, status, requested model, caller API key, latency, token counts, and attempt details when available. Latency is reported from two angles — what the caller waited for and what the upstream spent — so a slow request can be attributed without guesswork. Expanding a row shows the **Request ID** AISIX assigned, and — when the call reached a provider that returned one — a **Provider request ID** beside it. They are different things. The first is the ID AISIX returns to the caller in the `x-aisix-request-id` response header, and the one a caller reporting a problem will normally have. The second is the ID the upstream provider returned in its own response, such as an OpenAI `chat.completion.id` or an Anthropic message `id`, which is what the provider's console and support channel index the call by. Look the request up by the first, read the second from it, and take that to the provider. The provider request ID is not shown when the call produced none: a response served from cache, a request rejected before AISIX reached the provider, and endpoints whose provider response carries no ID at all, such as embeddings, audio, and image generation. A retried or failed-over request records one per attempt that got a response, so expand the attempt that actually served the caller. Rows with an `estimated` badge contain one or more token counts calculated locally because the upstream response omitted those values or reported zero. This can occur with OpenAI-compatible relays, client disconnects mid-stream, and upstream errors after a partial response. Estimated counts are included in spend and budget calculations alongside provider-reported counts. The badge distinguishes estimated values when investigating usage or spend. See [Usage Reporting](https://docs.api7.ai/ai-gateway/cloud/usage-reporting.md#how-usage-is-reported) for the supported endpoints and estimation behavior. ![AISIX Cloud control plane Request Logs showing filters, request status, token usage, latency, and expanded request details](https://static.api7.ai/uploads/2026/06/25/vZnxzxum_log.png) Freshly completed requests can take a few seconds to appear because data planes flush telemetry in batches. Each row is timestamped with its local date and time, so a window that spans several days stays readable. Hover over a timestamp to see the full date with its time zone. Start with request logs when you need to verify a live request, inspect an upstream error, confirm a policy rejection, or check how routing and failover resolved a request. ### Read the Latency Figures[​](#read-the-latency-figures "Direct link to Read the Latency Figures") A request log row shows what the caller waited for. Expand the row to see the measurement broken out, because a single number cannot answer whether a slow request was the provider's fault or the gateway's: | Field | Measures | Scope | | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------- | | **Caller latency** | Gateway received the request until it completed a non-streaming response or started sending a streamed one. | The whole request, including every retry and failover attempt. | | **Upstream latency** | How long the attempt that served the request spent talking to the provider. | One attempt. | | **Upstream TTFT** | How long that attempt waited for the provider's first streamed frame, whatever it carries — metadata openers such as `response.created` or `message_start` included. Streaming only. | One attempt. | Caller latency is the figure to quote in an SLO discussion, and it is what the **Latency p50 / p99** card on the Overview page reports. Upstream TTFT stops at the first frame of the stream, not the first visible token. This is the convention proxies and gateways in front of AISIX use, so the figure compares directly with what they log. A reasoning model that streams nothing visible while it thinks — for example, a `/v1/responses` upstream that opens the stream immediately but emits no reasoning summaries — therefore shows a small TTFT even though the answer text starts much later; that thinking wait is part of upstream latency, not TTFT. For a streamed response, caller latency deliberately stops when the stream begins rather than when it ends. A stream's total duration grows with however many tokens the model generated, which says little about the experience; the wait before output starts appearing is what a user notices. Comparing the fields locates the delay. When caller latency far exceeds upstream TTFT on a request that did not retry, the time went to gateway-side work. The usual cause is an output guardrail that masks responses: it must buffer the whole stream before releasing any of it. When the two figures are close, the provider was the slow part. On a request that retried or failed over, the gap also contains the earlier attempts, so check the attempt rows first. The **Latency p50 / p99** card counts successful requests only. A rejected request is usually fast precisely because little happened, and including those would pull the percentiles down and mask real slowness. note Requests recorded by a data plane older than 0.7 carry a single latency value and no caller-facing figure, so they do not appear in the latency percentiles. The split applies to traffic recorded after the upgrade. ### See What a Semantic Guardrail Measured[​](#see-what-a-semantic-guardrail-measured "Direct link to See What a Semantic Guardrail Measured") Expanding a request also shows **Semantic guardrail scores** when an embedding-similarity guardrail screened it. Each entry names the guardrail and the hook that ran, which example list it scored against, the similarity it measured against the threshold it was compared to, the embedding model that produced the score, and the line number of the closest example in that list. These are recorded on requests the guardrail allowed as well as on ones it refused, and in `monitor` mode as well as `block`. That is what makes them useful for tuning: a monitor hit appears only when the guardrail would have blocked, so a row tuned just short of firing — a deny threshold slightly too high, or an allow threshold slightly too low — produces no monitor hit at all and looks identical to a guardrail that is not running. The scores are recorded either way. A2A and rerank requests can report these scores, but both run only the input hook. A2A resolves no model or MCP server, so only an environment, caller API key, or team attachment reaches it. A screened rerank request emits a usage event even when the upstream returns no readable usage block. Guardrail attribution alone is enough to create the row. An expanded request showing no scores has several causes: no semantic guardrail screened it, an older gateway recorded it, another guardrail refused the request first (the chain stops at the first block, so a guardrail attached at a higher priority means the semantic one downstream of it never runs), the row is a superseded attempt of a retried, failed-over or ensemble request (only the request's terminal row carries the scores, which is not always the last one listed), on an output-hook row the streamed reply outgrew `max_buffer_bytes` and was refused before the guardrail ran, the row covers an input-only endpoint but the guardrail is output-only, there was nothing to screen (under the default `text_source: user_messages`, a request carrying only an image reaches no embedding call), or the embedding call failed. That last one is the one to rule out, because screening stops at the failure before anything is scored, so in this section a failing embedding model looks identical to a quiet guardrail. A failed embedding call does leave a signal elsewhere on the same row, and which field depends on the row's mode. A `block` guardrail failing closed, the default, refuses the request and records `guardrail_enforced_hits` with the action `blocked_unavailable`, shown under **Enforced hits** as **check unavailable**. A `monitor` guardrail failing closed serves the request and records `guardrail_monitor_hits` with the action `would_block`, shown under **Monitor hits** as **would block** — neither of the other two fields is written, which makes this the easiest case to misread. Either mode failing open served the request unscreened by that guardrail — the rest of the chain still ran — and records a **Bypass reason** (`guardrail_bypassed_reason`). Check all three before reading an empty section as "nothing to report." Scores from different embedding models are not comparable, which is why the model is shown beside every score. Neither the screened text nor the example text is recorded — the example is identified only by `top_example_index`, a zero-based index into that direction's list, which the console renders as a line number counting from one. See [Calibrate Semantic Screening Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/semantic-screening-calibration.md) for how to use these values to choose a threshold. ### Filter by Request Kind[​](#filter-by-request-kind "Direct link to Filter by Request Kind") Rows carry the kind of work the request asked for — a chat completion, an image generation, a video submission, a tool call — and the **Operation** filter narrows the feed to one of them. Nothing else on a row answers that question: every OpenAI-compatible endpoint reports the same inbound protocol, so a text conversation and an image generation look alike, and a model name is no substitute because one model serves several endpoints and callers address models through aliases and model groups. Rows are marked with the operation except for the four conversational ones — `chat`, `messages`, `responses` and `completions` — which are the common case; expand a row to see the value itself. Requests recorded before the gateway was upgraded carry no operation: they are neither marked nor returned by the filter. The filter applies on every tab, since MCP, A2A and passthrough traffic are operations too, and it stays applied when you switch tabs. It narrows within the tab rather than overriding it, so a combination with nothing in it — `mcp` on the LLM tab, say — returns an empty list rather than the MCP traffic. The values and what each one covers are listed in [Tell Request Kinds Apart](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md#tell-request-kinds-apart), which also describes how the same field reaches an external exporter. Two things are worth knowing before reading a count. The operation describes the **request**, so a failed or policy-rejected row carries the same value a successful one would — which is what lets you see which endpoint a rejection came from. And a request that retries or fails over contributes one row per attempt, so counting rows counts attempts rather than requests. ### Search Requests[​](#search-requests "Direct link to Search Requests") The search box above the filters looks for the text you type anywhere in a request's text fields, ignoring case. It covers the error message and error class, the requested and resolved model names, and the provider and provider key labels. It also covers the client user agent and source IP, the finish reason, the request ID, and the provider request ID — so an ID quoted by a provider's support channel finds its request here. Use it when all you have is part of an error a client reported and you do not know which field carries it. Search combines with the filters rather than replacing them. For example, a search for `rate limit` together with the **5xx** status filter returns only the server errors whose text mentions a rate limit. The dedicated filters remain the precise way to narrow by a known request ID, model, provider key, or caller API key. In the **Requested model / group** filter, typing searches aliases by case-insensitive substring, while choosing a suggested model or group matches the complete alias exactly, including case. Exports preserve whether the active filter in the request list was typed or selected. ### Export Requests[​](#export-requests "Direct link to Export Requests") **Export** downloads every request matching the current filters and search, not just the page on screen. Choose **CSV** for spreadsheet review or **JSON** for a downstream pipeline; JSON returns the same fields as the control-plane API, wrapped in a `data` array with a `total` count. Both formats carry the request time, request ID and attempt details, and the status and error text. They also carry the operation, the requested and resolved model names, caller API key name, token counts, latency, and cost. Both the caller-facing and upstream latency figures are included as separate columns. CSV files are UTF-8 with a byte-order mark, so spreadsheet applications read non-ASCII error messages correctly. Guardrail evidence travels with the row as well: `guardrail_scores` carries the similarity each semantic guardrail measured, as a column in CSV and as an array in JSON. One export returns at most 50,000 requests, newest first. When the filters match more than that, the control plane exports the newest 50,000 and reports the full match count. Narrow the time range or the filters to export the remainder. ## Audit Log[​](#audit-log "Direct link to Audit Log") The audit log records control-plane state changes for compliance review and operational investigation. It shows who created, updated, or deleted resources such as environments, models, API keys, provider keys, budgets, policies, and admin tokens. Use the audit log when traffic behavior changes after a configuration update. Request logs can show the request outcome, while the audit log can show whether a control-plane resource changed before that outcome. Audit log access is restricted to organization owners and admins. Entries are listed newest first, and the footer below the trail reports how many entries the current filters match in total. Page numbers refer to that whole filtered set rather than to what has been loaded so far, so a review can be resumed at a known position. ### Filter and Search the Trail[​](#filter-and-search-the-trail "Direct link to Filter and Search the Trail") The filters above the trail narrow it by resource type, actor, and time. The resource-type list offers only the types this organization has actually recorded. The time range offers presets and a **Custom range**, which takes an explicit start and end time so a review can be pinned to the exact window an incident covers. The search box looks for the text you type anywhere in an entry's readable fields, ignoring case. It covers the recorded before and after state, which is where a resource's display name lives. It also covers the resource type and identifier, the action, the actor identifier, the client IP address, and the user agent. Use it when a ticket names the resource but not which change touched it. Search combines with the filters rather than replacing them. For example, a search for a model name together with an actor returns only that person's changes mentioning the model. The dedicated filters remain the precise way to narrow by a known resource type or actor. ### Export the Trail[​](#export-the-trail "Direct link to Export the Trail") **Export** downloads every entry matching the current filters and search, not just the page on screen. Choose **CSV** for spreadsheet review, or **JSON** for a downstream pipeline or an evidence archive. JSON returns the same fields as the control-plane API, wrapped in a `data` array with a `total` count. Both formats carry the entry time and identifier, the action, the resource type and identifier, the client IP address and user agent, and the complete before and after state. They also carry the actor's email address alongside the actor identifier, so an exported file stays readable for a reviewer who does not have dashboard access. CSV files are UTF-8 with a byte-order mark, so spreadsheet applications read non-ASCII resource names correctly. One export returns at most 50,000 entries, newest first. When the filters match more than that, the control plane exports the newest 50,000 and reports the full match count. Narrow the time range or the filters to export the remainder. ## Investigate a Request Outcome[​](#investigate-a-request-outcome "Direct link to Investigate a Request Outcome") Send the request through the AISIX gateway endpoint with a caller API key and model alias that belong to the same environment as the gateway. After the request completes, check request logs for a matching request time, status, requested model, and caller API key. If the request used routing or failover, inspect the resolved model or attempt details when they are available. An upstream authentication, quota, or provider-side error can still prove that the AISIX gateway path is working. In this case, the request reached AISIX. AISIX selected the configured model and provider key, then the upstream provider returned an error. Do not treat a provider error as a resource-projection failure unless the log shows the wrong model, provider key, or environment. ## Investigate Policy Rejections[​](#investigate-policy-rejections "Direct link to Investigate Policy Rejections") AISIX Cloud policies can reject traffic before AISIX calls the upstream provider. Budget hard stops return a budget-related error, and rate-limit policies return a rate-limit error. Guardrails can reject unsafe content before or after the provider call depending on the guardrail hook. When a request is rejected, first identify whether the response came from AISIX or from the upstream provider. Then use request logs to check the status and request identity. For budget rejections, compare the returned budget scope with the Budgets view. For rate-limit rejections, check the caller API key, model, team, or member policy that matches the request. For guardrail rejections, check the guardrail scope and the model or caller identity that triggered it. ## If Traffic Does Not Appear in Request Logs[​](#if-traffic-does-not-appear-in-request-logs "Direct link to If Traffic Does Not Appear in Request Logs") If request logs do not show the expected record, check the request path first. Request logs are scoped to AISIX gateway traffic for the selected environment, so requests sent to another gateway endpoint or another environment appear elsewhere. If the request path is correct, check that the AISIX gateway has a recent heartbeat and can reach the control-plane telemetry endpoint. Other control-plane signals can help narrow the cause: | Signal | What it shows | | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | | Data planes heartbeat | Whether the AISIX gateway is connected to the control plane and reporting from the expected environment. | | Usage | Whether the control plane received AISIX Cloud gateway telemetry for aggregate usage and budget workflows. | | Observability exporter health | Whether the gateway has applied exporter configuration and is reporting delivery status for external telemetry destinations. | ## External Exporters[​](#external-exporters "Direct link to External Exporters") Request logs show the control plane's view of AISIX Cloud gateway telemetry. Observability exporters are configured from the environment's **Observability** view and send usage events from the AISIX gateway to destinations you control. Use exporters when you need to send request telemetry to an external tracing, logging, storage, or accounting system. Exporter delivery happens from the AISIX gateway directly to the destination. A request log can exist even when an external destination has a credential, network, or receiver-path issue. ## Data Retention[​](#data-retention "Direct link to Data Retention") The control plane keeps AISIX Cloud gateway telemetry for a per-organization retention window. This includes the data behind both the **Request Logs** and **Usage** views. By default, records are kept for 30 days, and the control plane automatically removes older records each day. Organization owners and admins set the retention window in **Settings**, under **Usage log retention**, to any value from 1 to 3650 days. A longer window preserves more history for investigation and reporting. A shorter window reduces how much data is stored. A change applies going forward and takes effect on the next daily cleanup. Retention is why older traffic eventually stops appearing in Request Logs and Usage. To keep request telemetry beyond the retention window, configure an [external exporter](#external-exporters) to deliver usage events to a destination you control before the records are removed. ## Next Steps[​](#next-steps "Direct link to Next Steps") Use the [AISIX Cloud Admin API Reference](https://docs.api7.ai/ai-gateway/reference/cloud-admin-api/.md) for supported automation operations. To deliver gateway telemetry to systems you operate, continue with [Observability Exporters](https://docs.api7.ai/ai-gateway/observability/exporters.md). --- # Members Members belong to an organization in the AISIX Cloud control plane. They represent the people or service owners responsible for control-plane administration, API keys, usage, and limits. The AISIX Cloud control plane supports two member onboarding paths: * Invite a member when they need to sign in to the dashboard. * Create a member directly when you need an API key owner that does not sign in to the dashboard. Both paths create organization members. The difference is whether the member receives an invitation and can use the dashboard. ## Invite a Member[​](#invite-a-member "Direct link to Invite a Member") Use an invitation when the person needs dashboard access to manage resources, view usage, or administer the organization. 1. Open **Members** and select **Invite member**. 2. Enter the member's email address and choose a role. 3. Send the invitation and share the one-time invitation link. Opening the link shows the invitee who invited them, which organization they are joining, and with which role. An invitee who does not have an account yet creates one from that page — the email address is filled in and cannot be changed — and comes back to the invitation. Joining always takes an explicit **Accept invitation**: opening the link never changes membership on its own, which also means an existing user who is signed in to another organization is not moved by following it. The invitation is bound to the address you entered. Only an account using that address can accept it; anyone else sees which address the invitation was issued to and is offered a way to sign in as that person. Until the invitee accepts, the invitation stays pending on the **Pending invitations** tab. You cannot invite an address that already belongs to a member of the organization. Accepting an invitation never changes an existing member's role, so such an invitation would have nothing to do — change the role from the members list instead. ### Invitation Lifetime[​](#invitation-lifetime "Direct link to Invitation Lifetime") An invitation expires 7 days after it is sent. Until then it stays on the **Pending invitations** tab, where you can revoke it to invalidate the link immediately. That tab lists live invitations only. Select **Show stale invitations** to also see accepted, revoked, and expired ones. An expired invitation can be revoked from there when you want to retire the record. An expired invitation does not keep the email address reserved: inviting the same person again issues a fresh link and retires the lapsed invitation, which stays visible under **Show stale invitations**. Only an invitation that is still pending and valid blocks a second invitation to the same address. ## Create a Member Directly[​](#create-a-member-directly "Direct link to Create a Member Directly") Create a member directly when the member only needs to own API keys and does not need dashboard access. Typical cases include services, applications, or developers in a private deployment where dashboard access is restricted. A directly created member: * Becomes active immediately, with no invitation link or confirmation step. * Can be added to teams and assigned API keys. * Can be governed with rate limits and budgets. * Has no password, so it cannot sign in to the dashboard. The member still has a name and email address. Use values that identify the responsible person, service, or application owner so usage and limits can be attributed correctly. ### In the Dashboard[​](#in-the-dashboard "Direct link to In the Dashboard") 1. Open **Members** and select **Create user**. 2. Enter a **Name** and an **Email** that identify the responsible owner. 3. Select **Create user**. The member appears in the list right away. You can then [add the member to a team](https://docs.api7.ai/ai-gateway/cloud/teams.md#add-and-manage-members) and issue API keys from the environment that serves its traffic. ### Use the API[​](#use-the-api "Direct link to Use the API") Use the API when you need to provision members from automation. Authenticate with an [organization admin token](https://docs.api7.ai/ai-gateway/cloud/admin-tokens.md) that has write scope. Admin tokens are scoped to one organization, so the request does not need a separate organization header. ``` # AISIX_CP includes /api and has no trailing slash. # The local On-Premises quickstart uses http://localhost:8080/api. export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_BASE_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" ``` Create the member: ``` curl -X POST "${AISIX_CP}/members" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{ "name": "Responsible Person", "email": "svc-payments@example.com" }' ``` A successful call returns `201 Created` with the new member: ``` { "member": { "id": "8f3b2a1c-9d4e-4f6a-b7c8-1e2d3f4a5b6c", "user_id": "2c7d6e5f-4a3b-4c2d-8e1f-9a0b1c2d3e4f", "email": "svc-payments@example.com", "display_name": "Responsible Person", "role": "member" } } ``` The email address must be valid and unique across the deployment. For an application owner, you can use a synthetic address such as `svc-payments@example.com`. Reusing an existing email address returns `409 Conflict`. Directly created members always use the `member` role. ## Browse and Search Members[​](#browse-and-search-members "Direct link to Browse and Search Members") The **Members** list supports search and pagination, so large organizations stay manageable. * Use the search box to filter by name or email. The search runs server-side, so it matches across every page, not only the rows currently in view. * Use the controls below the list to change the page size or move between pages. ### List Members with the API[​](#list-members-with-the-api "Direct link to List Members with the API") Reuse the `AISIX_CP` and `AISIX_TOKEN` values exported above. A token with read scope is sufficient for this request. ``` curl "${AISIX_CP}/members?page=1&page_size=20&q=payments" \ -H "Authorization: Bearer ${AISIX_TOKEN}" ``` The query parameters are optional: * `q`: case-insensitive match against member name and email. * `page`: 1-based page number. Requires `page_size`; setting `page` on its own returns `400`. * `page_size`: page size, up to `200`. Omit both `page` and `page_size` to disable paging and return every member in a single response. The response wraps the members in a pagination envelope: ``` { "data": [], "total": 128, "page": 1, "page_size": 20, "owner_count": 2 } ``` `total` is the number of members that match the filter across all pages, and `owner_count` is the number of owners in the organization. ## Remove a Member[​](#remove-a-member "Direct link to Remove a Member") Select the remove action on a member row. Removing a member takes their organization membership, their environment-scoped role bindings, and their place on every team. It also **disables every API key that member owns**, and the change reaches the gateway, so those credentials stop authenticating there and not only in the dashboard. This is the same deprovisioning your identity provider triggers over [SCIM](https://docs.api7.ai/ai-gateway/cloud/scim-directory-sync.md) — removing someone through the dashboard and removing them through the directory now leave the organization in the same state. The keys are disabled rather than deleted, so their usage history stays available for reporting. A key an operator had already disabled by hand keeps that provenance. If the person's credentials should keep serving traffic after they leave — an application key created under their account, say — rebind or recreate them before removing the member. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Teams](https://docs.api7.ai/ai-gateway/cloud/teams.md) to group members for attribution and shared controls, or use [Roles and Custom Roles](https://docs.api7.ai/ai-gateway/cloud/custom-roles.md) to control member permissions. Use [SCIM Directory Sync](https://docs.api7.ai/ai-gateway/cloud/scim-directory-sync.md) when your identity provider should manage membership. --- # Model Pricing The AISIX Cloud control plane calculates request cost from the usage it observes and the price associated with the upstream provider and model. Most models are priced on token counts; models billed by duration are priced on audio length. Configure a pricing override when a model has no catalog price or when your organization uses a different rate. Pricing affects the spend shown in Usage and the totals evaluated by AISIX Cloud budgets. Prompt and completion rates can also affect target ordering for [`least_cost` routing groups](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md#route-by-cost-latency-or-load). Pricing does not change what the upstream provider charges. ## How AISIX Resolves Prices[​](#how-aisix-resolves-prices "Direct link to How AISIX Resolves Prices") AISIX applies the configured rates without hidden multipliers. Each usage event resolves a price by the exact `(provider, model name)` pair: 1. An organization override wins when it matches the provider and model name. 2. Otherwise, AISIX uses the matching catalog default from [models.dev](https://models.dev). 3. If neither source matches, the usage event retains its reported usage but records `$0.00` spend. Catalog refresh behavior depends on how pricing synchronization is configured. In online mode, the control plane refreshes the catalog at startup and every 24 hours. By default, the On-Premises Docker Compose package loads the bundled catalog snapshot at startup, even when it has internet access. To enable online synchronization for that package, see [Pricing Catalog](https://docs.api7.ai/ai-gateway/reference/on-premises-configuration.md#pricing-catalog). Within a matching price, AISIX calculates token and duration cost independently, then adds the two terms. Token pricing uses the prompt and completion rates, with the cache and reasoning fallbacks described below. Duration contributes only when the event reports audio length and **Audio per minute** is nonzero. Cost is calculated and stored when the usage event is recorded. A pricing change applies to later events and does not re-price historical usage. ## Identify a Missing Price[​](#identify-a-missing-price "Direct link to Identify a Missing Price") Start with an expanded event in [Request Logs](https://docs.api7.ai/ai-gateway/cloud/logging-and-auditing.md#request-logs), which shows the cost of each usage event or routing attempt to six decimal places. The aggregate **Usage** view rounds spend to two decimal places and shows only the top 10 environment/model rows. Therefore, `$0.00` spend or an absent model in Usage does not prove that pricing is missing. When checking a very low-priced request, generate enough token usage to make the expected cost visible before concluding that a price is missing. A catalog price can be missing when: * The model runs on a self-hosted or reseller upstream that models.dev does not price. * The configured model name is a custom alias. * The provider offers a newer model that the catalog does not yet include under that provider. Before adding an override, note the provider and model name exactly as they are configured. A pricing row that differs from either value does not match the usage event. ## Add a Pricing Override[​](#add-a-pricing-override "Direct link to Add a Pricing Override") 1. Open **Model pricing** from the organization navigation. 2. Select **Add pricing**. 3. Enter the **Provider** and **Model name** exactly as configured on the model. The Provider field suggests known catalog providers but also accepts a custom value. 4. Configure the basis the provider charges: * For a token-priced model, enter the **Prompt** and **Completion** rates in USD per 1M tokens and leave **Audio per minute** at `0`. * For a duration-priced speech-to-text model, leave the token rates at `0` and enter the **Audio per minute** rate in USD per minute of audio, such as `0.006`. 5. For a token-priced model, optionally enter distinct **Cache read**, **Cache write**, and **Reasoning** rates. 6. Save the override. A cache or reasoning rate of `0` means that no distinct rate is set. Cache tokens fall back to the prompt rate, and reasoning tokens fall back to the completion rate. ## Price a Model Billed by Audio Length[​](#price-a-model-billed-by-audio-length "Direct link to Price a Model Billed by Audio Length") Some speech-to-text models are billed per minute of audio rather than per token and report no token counts. Such a request records its audio length instead, and the **Audio per minute** rate turns that length into spend. A model left at `0` for this rate contributes no duration-based cost. The current AISIX catalog does not supply duration rates, so there is no default to inherit: create the row with the provider and model name exactly as configured. For example, OpenAI currently lists [`whisper-1`](https://developers.openai.com/api/docs/models/whisper-1) at `$0.006` per minute. At that rate, a 60-second transcription costs `$0.006`. caution Token and duration costs are additive. If an event contains both token usage and audio duration, and the matching price has nonzero rates for both, AISIX adds both terms. Configure only the usage basis that the provider charges. ## Verify the Override[​](#verify-the-override "Direct link to Verify the Override") For the simplest verification, send a new request directly to the matching model, then expand the event in **Request Logs**. For a token-priced model, generate enough token usage to produce a visible cost. For a duration-priced model, use an audio sample with a known or sufficiently long duration. Confirm that the event shows the expected model and event cost. Request Logs does not currently display the measured audio duration, so calculate the expected duration cost from the sample length. If you verify a routed or ensemble request instead, find the event or attempt rows that share its request ID. Add their costs before comparing the total with the expected cost. Request Logs shows six decimal places. If the expected event cost rounds to `$0.000000`, generate enough token usage to make the cost visible at that precision. After sufficient traffic, you can also confirm aggregate spend in **Usage**. Usage rounds spend to two decimal places and shows only the top 10 model rows, so it is less reliable for verifying a single low-cost request. Historical rows do not change after the override is saved. Compare an event created after the change rather than expecting an earlier `$0.00` event to be recalculated. ## Edit or Reset a Price[​](#edit-or-reset-a-price "Direct link to Edit or Reset a Price") Edit a row to change the rates used for later events. **Reset to catalog default** removes the organization override. If a catalog price exists, it takes effect again on the next usage event. A manually created duration-pricing row usually has no catalog default. Resetting that row removes its effective price, so later matching events record `$0.00` until another override is configured. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md) to enforce spending limits, then configure [Budget Alerts and Notifications](https://docs.api7.ai/ai-gateway/traffic-controls/budget-alerts.md). Use [Logging and Auditing](https://docs.api7.ai/ai-gateway/cloud/logging-and-auditing.md) to investigate individual request outcomes. --- # Offline Resilience Temporary control-plane connectivity loss separates the live traffic path from management workflows. A running AISIX gateway can keep serving from its latest accepted configuration while it reconnects, but it cannot receive new resources and some AISIX Cloud services remain connectivity-dependent. Offline resilience protects traffic that depends on configuration already held by the gateway. It does not make the control plane optional or make external upstreams available during their own outages. ## Continue Serving Accepted Configuration[​](#continue-serving-accepted-configuration "Direct link to Continue Serving Accepted Configuration") After a gateway has applied valid projected configuration, it keeps that snapshot in memory and continues serving it while the configuration connection is unavailable. AISIX does not withdraw the gateway from traffic merely because no newer configuration event has arrived. Use configuration freshness as an operator alert instead of a load-balancer health condition. No resource changes reach the gateway until connectivity returns. Changes that the control plane accepts during the outage wait for projection and do not affect the snapshot already in service. ## Restart from Cached Configuration[​](#restart-from-cached-configuration "Direct link to Restart from Cached Configuration") In managed mode, AISIX caches the latest accepted snapshot on disk after successful configuration applies. The default path is `/var/lib/aisix/config_cache.json`; setting `managed.snapshot_cache_path` to an empty string disables this cache. See [AISIX Cloud startup configuration](https://docs.api7.ai/ai-gateway/reference/configuration-files.md#aisix-cloud). A restarted gateway can serve while reconnecting only when that cache is enabled and contains a valid snapshot. AISIX ignores a missing, unreadable, corrupt, or incompatible cache. Without a usable snapshot, `/readyz` and `/status/ready` return `503` until the gateway connects and applies valid configuration. Persist the gateway state directory when a restarted instance must recover its snapshot. Give each gateway instance its own writable state directory rather than sharing one between replicas. The [`api7/aisix` Helm chart](https://docs.api7.ai/ai-gateway/cloud/kubernetes.md) uses ephemeral state by default, so a replacement pod must reconnect and obtain configuration before receiving traffic. ## Connectivity-Dependent Workflows[​](#connectivity-dependent-workflows "Direct link to Connectivity-Dependent Workflows") The gateway handles control-plane-dependent workflows separately: | Workflow | Behavior while disconnected | | ----------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Resource projection | New and updated resources do not reach the gateway. The latest accepted snapshot remains in service. | | AISIX Cloud budgets | AISIX reuses the last decision for up to `AISIX_DP_BUDGET_STALE_MAX_SECONDS`, which defaults to `600`. Without a cached decision, it denies the request. After the stale window, it applies the failure mode returned with the last decision. See [Budget Availability and Caching](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md#availability-and-caching). | | AISIX Cloud usage telemetry | Failed batches are dropped. AISIX does not persist, retry, or replay them after connectivity returns. Live requests are not failed or delayed to preserve telemetry. | | External observability exporters | The gateway sends directly to each configured destination, so control-plane loss alone does not interrupt them. Export still depends on connectivity from the gateway to that destination. | | Heartbeats and certificate rotation | Heartbeat status stops updating. Certificate rotation requires control-plane connectivity; the gateway continues using its current certificate while it remains valid. | ## Detect and Recover[​](#detect-and-recover "Direct link to Detect and Recover") A running gateway that is serving a snapshot remains traffic-ready. Check management-path health separately: * `GET /status/config` reports `source.connected: false` and retains the applied revision and configuration hash. * `GET /status/ready` remains `200` after a valid configuration has been applied. It is `503` when a restarted gateway has no usable snapshot. * The `aisix_config_*` metrics expose connection state, applied configuration, and load failures for alerting. * AISIX Cloud gateway status becomes stale when heartbeats stop. See [Configuration Status](https://docs.api7.ai/ai-gateway/reference/config-status.md) for the complete status fields, metrics, and alert examples. Keep the metrics and status listener private because it is unauthenticated. To recover, restore DNS, network, and TLS connectivity from the gateway to the AISIX Cloud control-plane endpoint. Confirm that heartbeats resume, compare the published and applied revisions, and send a caller-visible request through the recovered gateway. A successful live request confirms the traffic path, but it does not recover telemetry batches that were dropped during the outage. ## Next Steps[​](#next-steps "Direct link to Next Steps") Use [High Availability](https://docs.api7.ai/ai-gateway/cloud/high-availability.md) for the broader failure-domain design. Use [Resource Projection](https://docs.api7.ai/ai-gateway/cloud/resource-projection.md) to verify that a recovered gateway applied the expected revision. --- # Organizations and Environments Organizations and environments define the two primary scopes in the AISIX Cloud control plane. The organization establishes ownership and shared administration. The environment establishes which AISIX gateways receive projected configuration. This separation keeps account administration separate from traffic behavior. Members, billing, and provider credentials belong to the organization. Models, caller API keys, policies, and cache policies belong to the environment that serves traffic. For live traffic, the environment is usually the scope that matters first. A saved environment resource affects live traffic only when it belongs to the environment attached to the target AISIX gateway. Account scope**Organization** Members, billing, shared administration, and provider credentials. MembersBillingProvider keys Environment**Development** ModelsCaller API keysPoliciesCache policies Projects to**Development gateway** Environment**Production** ModelsCaller API keysPoliciesCache policies Projects to**Production gateway** ## Scope Model[​](#scope-model "Direct link to Scope Model") The AISIX Cloud control plane can manage more than one gateway and more than one set of gateway resources. Organizations and environments make it clear which team owns the deployment and which AISIX gateway receives each environment-scoped resource. ## Organization Scope[​](#organization-scope "Direct link to Organization Scope") An organization is the top-level ownership boundary. It is the scope for members, billing, shared administration, and provider credentials. Use organization scope to reason about account ownership and shared administration. It is usually not the first scope to check when a saved gateway resource does not affect live traffic. ## Environment Scope[​](#environment-scope "Direct link to Environment Scope") An environment is the unit that groups gateway resources and projects them to the attached AISIX gateway or gateways. Environment scope turns saved control-plane state into live gateway behavior. A resource can exist in the control plane and still not affect the gateway you are testing if it belongs to another environment. When a resource does not affect traffic as expected, confirm first that the resource belongs to the same environment as the target AISIX gateway. For provider keys, also confirm that the provider key is allowed in that environment. ## Create or Select an Environment[​](#create-or-select-an-environment "Direct link to Create or Select an Environment") Open **Environments** in the dashboard. If your account does not belong to an organization yet, the dashboard first prompts you to create one. To create an environment, select **New environment**, enter a display name, and select **Create environment**. The confirmation shows the environment ID and links to its **Data planes** view. Copy the ID for provider setup examples, which use it as `ENV_ID`. To use an existing environment, select it from the environment list. The environment ID appears below its display name. ## Rename an Environment[​](#rename-an-environment "Direct link to Rename an Environment") Select the pencil next to an environment in the list, type the new name, and save. Names are unique within the organization, so a name another environment already uses is rejected. The environment ID does not change, and neither does anything the gateway is serving. Attached AISIX gateways identify their environment by ID, so a rename is a control-plane label change only — running data planes are not reconfigured and no traffic is interrupted. Scripts and setup examples that reference `ENV_ID` keep working. A member whose access comes from an environment-scoped role binding cannot rename the environment. Renaming requires write access to environments at the organization level, the same permission that creates and deletes them. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Admin Tokens](https://docs.api7.ai/ai-gateway/cloud/admin-tokens.md) to create the organization-scoped credential used by the provider setup examples. After creating a token, [connect an AISIX gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md) to the environment that should receive the configuration. --- # AISIX Cloud AISIX Cloud is API7's commercial offering for centrally managing AISIX gateways across teams and environments. Its management layer stays separate from the live AI traffic path, so organizations can govern their gateway fleets without routing model traffic through the control plane. AISIX gateways run in your environment and connect to the control plane for configuration delivery and operational reporting. Live AI traffic passes through those gateways rather than through the control plane or API7. ## Deployment Options[​](#deployment-options "Direct link to Deployment Options") AISIX Cloud has two deployment options: On-Premises and Hybrid Cloud. Both use the same resource model and gateway workflow. In both options, you configure gateway resources through the dashboard or AISIX Cloud Admin API and operate the AISIX gateways in your infrastructure. The options differ in who deploys and operates the control-plane services and where control-plane data is stored. * **On-Premises:** You deploy and operate the control-plane services in your infrastructure, including in fully air-gapped environments. Try it locally with the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). * **Hybrid Cloud:** API7 manages the deployment and operation of the control-plane services. Hybrid Cloud is not currently available through public self-service registration. [Contact API7](https://api7.ai/contact) to request a trial or demo. AISIX Cloud is also available for purchase through [AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-o7ltvkj4qjnr2). Marketplace changes procurement and billing, not the division of operational responsibilities. ## How AISIX Cloud Works[​](#how-aisix-cloud-works "Direct link to How AISIX Cloud Works") The control plane provides centralized resource management, certificate issuance, access governance, usage and cost controls, and visibility into gateway health. AISIX gateways serve as the data plane and handle application traffic directly. ![AISIX Cloud management capabilities and live AI traffic through AISIX gateways](https://static.api7.ai/uploads/2026/07/27/Sa2n147n_aisix-cloud-overview.svg) Each gateway initiates outbound mTLS connections to the control plane to receive configuration updates, report health and usage, and request budget decisions. During temporary connectivity loss, the gateway continues serving its latest accepted configuration. For model traffic, the gateway applies routing, reliability controls, policy enforcement, and observability before forwarding requests to model APIs. The gateway can also proxy MCP and A2A traffic. ## Manage AISIX Cloud[​](#manage-aisix-cloud "Direct link to Manage AISIX Cloud") The dashboard is the browser interface for the control plane. Use it for interactive tasks such as managing organizations and environments, issuing gateway certificates, and reviewing usage. For automation, use the operations documented in the [AISIX Cloud Admin API Reference](https://docs.api7.ai/ai-gateway/reference/cloud-admin-api/.md). The reference is the supported public API contract; a dashboard workflow does not necessarily have a corresponding public API operation. Changes to gateway resources made through either supported interface are stored by the control plane and projected to gateways attached to the affected environment. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Organizations and Environments](https://docs.api7.ai/ai-gateway/cloud/organizations-and-environments.md) to understand how AISIX Cloud scopes resources, access, and gateways. --- # Playground The playground lets you test a single-target model from the AISIX Cloud control plane. Use it to confirm that the model can reach its upstream provider and return a chat response before sending application traffic through the gateway. ![AISIX Cloud control-plane playground showing a successful model test](https://static.api7.ai/uploads/2026/06/23/HssKkly4_cloud-playground.png) ## Testing Models in the Playground[​](#testing-models-in-the-playground "Direct link to Testing Models in the Playground") Playground requests run from the control plane to the upstream provider. Use them to validate the provider key, upstream base URL, model name, and basic prompt response from the control plane UI. The playground sends OpenAI-compatible chat-completions requests. Use it with single-target models whose provider key points to an OpenAI-compatible upstream base URL, including Anthropic-backed models exposed through an OpenAI-compatible endpoint. Playground requests bypass the AISIX gateway, so they are not recorded in Request Logs, Usage, AISIX gateway metrics, or external telemetry exporter output. Budgets and rate limits configured in AISIX Cloud do not apply to playground requests. Each run still calls the upstream provider with the model's provider key, so it consumes provider quota and is billed by the provider like any other API call. AISIX gateway requests run through the runtime that applications use. Send traffic through the AISIX gateway to validate caller API keys, model aliases, routing rules and failover, cache behavior, guardrails, budgets, rate limits, streaming behavior, logs, usage reporting, and metrics. Use AISIX gateway traffic when you need to validate endpoint families beyond chat completions or provider-specific protocol behavior. ## Reach Private-Network Endpoints On Premises[​](#reach-private-network-endpoints-on-premises "Direct link to Reach Private-Network Endpoints On Premises") The playground proxy refuses to connect to private, internal, or loopback IP addresses by default. This server-side request forgery (SSRF) protection prevents the control plane from reaching internal network services through a crafted model endpoint. On-Premises may need to reach an LLM endpoint on an internal network, while Hybrid Cloud expects public upstream endpoints. A blocked request fails with the error code `UPSTREAM_PRIVATE_IP_BLOCKED`. To let the playground reach a private-network endpoint, set `AISIX_PLAYGROUND_ALLOW_PRIVATE_IPS=1` on the control-plane API service and restart it: * **Offline package (Docker Compose):** uncomment `AISIX_PLAYGROUND_ALLOW_PRIVATE_IPS=1` in `.env`, then run `docker compose up -d api`. * **Helm chart:** set `api.playgroundAllowPrivateIPs=true`. Leave the default in place when your model endpoints are public. This setting affects only the playground proxy; the data plane always reaches upstream providers directly and is unaffected. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Usage Reporting](https://docs.api7.ai/ai-gateway/cloud/usage-reporting.md) to understand gateway usage, spend, and budget-related signals. --- # Resource Projection AISIX Cloud publishes saved resource changes to gateways connected to an environment. Each gateway validates a new revision and applies the accepted configuration to new requests. Publication is asynchronous, so gateway instances can temporarily report different revisions. A successful dashboard or API response confirms only that the control plane accepted the change. To verify that it took effect, trace it from the control-plane revision to the gateway snapshot and then to a live request. ## Resource Scope[​](#resource-scope "Direct link to Resource Scope") Resource ownership determines which environments receive a resource. Its attachments, references, or conditions then determine which requests use it. | Resource family | Ownership and projection | Request scope | | -------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Models and caller API keys | Belong to one environment and project to every gateway connected to it. | The requested model and resolved caller API key select the applicable entries. | | Provider keys | Belong to the organization and project only to their allowed environments. | Models reference provider keys. Passthrough routes reference one when they inject upstream credentials. | | MCP and OpenAPI-backed servers | Belong to the organization. Approved servers project only to their allowed environments; an OpenAPI-backed server is an MCP server with `type: openapi`. | Caller tool grants and applicable guardrail attachments control access and inspection. | | A2A agents | Belong to the organization and project only to their allowed environments. | The caller API key must allow the agent. | | Guardrails and attachments | Belong to one environment and project together. An unattached guardrail is not enforced. | Attachments scope a guardrail to an environment, model, caller API key, team, MCP server, or passthrough route. | | Cache policies | Belong to one environment and project to its gateways. | A policy can cover the environment, a model alias, or a caller API key. | | Rate-limit policies | Belong to one environment and project to its gateways. | Conditions can select traffic by team, member, caller API key, model, model name, or provider. | | OIDC providers and claim mappings | Belong to one environment and project to its gateways. | They authenticate JWT callers and resolve eligible claims to caller API keys. | | Passthrough routes | Belong to one environment and project to its gateways. | Route matching and the caller API key's route grant determine access. | | Observability exporters | Belong to one environment and project to its gateways. | Each enabled exporter receives eligible gateway telemetry according to its kind and content settings. | | MCP environment access and authentication settings | Belong to one environment and project to its gateways. | The environment policy defines a default tool-access layer. Authentication settings define the API key, OAuth, or anonymous behavior of `/mcp`. | | MCP team access policies | Belong to an organization through their team and project to every environment in that organization. | A team policy applies to caller API keys bound to that team and intersects with the environment and key-level access layers. | | Budgets | Remain in the AISIX Cloud control plane instead of entering the projected snapshot. | The gateway requests a decision for the resolved caller API key and can temporarily use a cached decision during an outage. See [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md#availability-and-caching). | Changing an organization-owned resource's allowed environments adds or removes that resource from the corresponding snapshots. A successful write confirms only that the control plane accepted the resource and its target environments; it does not confirm that every gateway has received it. ## Verify a Projected Change[​](#verify-a-projected-change "Direct link to Verify a Projected Change") Check the control plane, gateway, and live request in order. ### 1. Compare Published and Applied Revisions[​](#1-compare-published-and-applied-revisions "Direct link to 1. Compare Published and Applied Revisions") Set the control-plane API URL, a read-scoped admin token, and the environment ID: ``` # AISIX_CP includes /api and has no trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Compare `store_revision` with each node's `applied_revision`: ``` curl -sS "${AISIX_CP}/environments/${ENV_ID}/dp_nodes" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ | jq '{ store_revision, nodes: [ .data[] | { hostname, last_heartbeat_at, applied_revision, config_hash } ] }' ``` A node is current when its `applied_revision` is equal to or greater than `store_revision`. Matching `config_hash` values indicate that gateway instances have accepted the same configuration. Optional revision fields The control plane can omit `store_revision`, `applied_revision`, or `config_hash` when tracking is unavailable or a gateway has not reported them. In that case, use the gateway status endpoint in the next step. Heartbeats are periodic, so direct gateway status can also be newer than the control-plane view. If a gateway has no recent heartbeat, troubleshoot its management connection before checking projection. See [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md). ### 2. Inspect the Gateway Snapshot[​](#2-inspect-the-gateway-snapshot "Direct link to 2. Inspect the Gateway Snapshot") Query the metrics and status listener on the gateway instance handling traffic: ``` # AISIX_STATUS_URL is the metrics and status listener origin without a trailing slash or endpoint path # The open-source quickstart uses http://127.0.0.1:9090 export AISIX_STATUS_URL="YOUR_AISIX_STATUS_LISTENER_URL" curl -sS "${AISIX_STATUS_URL}/status/config" \ | jq '{ state, source, applied, rejected, last_failure }' ``` The change is fully applied when `state` is `synced`, `source.connected` is `true`, `applied.applied_revision` reflects the expected revision, and `rejected` is empty. If `source.connected` is `false`, restore the configuration connection. The gateway can continue serving its last accepted snapshot during an outage, but it cannot receive new changes. See [Offline Resilience](https://docs.api7.ai/ai-gateway/cloud/offline-resilience.md). If `state` is `degraded` or `out_of_sync`, use `rejected` and `last_failure` to identify the invalid resource. See [Configuration Status](https://docs.api7.ai/ai-gateway/reference/config-status.md) for the state definitions, complete response, and Prometheus metrics. Keep the status listener private The status listener is unauthenticated. Keep port `9090` private to the monitoring network. ### 3. Verify Caller-Visible Behavior[​](#3-verify-caller-visible-behavior "Direct link to 3. Verify Caller-Visible Behavior") After the gateway applies the change, verify it through the same gateway endpoint and caller identity the application uses. For a model or caller-access change, first query model discovery: ``` # AISIX_PROXY is the gateway origin without a trailing slash or endpoint path # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_URL" export AISIX_API_KEY="YOUR_CALLER_API_KEY" curl -sS "${AISIX_PROXY}/v1/models" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ | jq -r '.data[].id' ``` Confirm that the expected model alias appears, then send a request from the relevant provider or feature guide. If the result is unexpected after revisions match, first confirm that the request reached the inspected gateway. Then check the environment, model alias, caller API key, provider-key access, and policy scope. ## Next Steps[​](#next-steps "Direct link to Next Steps") If you have not sent live traffic yet, [choose a provider upstream](https://docs.api7.ai/ai-gateway/providers/overview.md) and follow its setup guide. Before placing multiple gateway instances behind a production load balancer, continue with [High Availability](https://docs.api7.ai/ai-gateway/cloud/high-availability.md). --- # SCIM Directory Sync SCIM directory sync connects an AISIX organization to your identity provider (IdP), so member accounts can be created, updated, and deactivated from the directory. AISIX implements the SCIM 2.0 protocol (RFC 7643/7644), which Okta, Microsoft Entra ID, and other enterprise IdPs support natively. Directory sync is built for organizations that attribute gateway usage to many people, including hundreds or thousands of members, without dashboard logins for each of them. Members provisioned this way are the same login-less members you can [create directly](https://docs.api7.ai/ai-gateway/cloud/members.md#create-a-member-directly): they own API keys, carry rate limits and budgets, and appear in usage reporting. They cannot sign in to the dashboard. When a person is deactivated or removed in the IdP, AISIX deactivates the member and disables every API key that member owns. The change propagates to the gateway data plane, so the credentials stop working at the gateway itself, not only in the dashboard. ## Enable Directory Sync[​](#enable-directory-sync "Direct link to Enable Directory Sync") 1. Open **Settings** and find the **Directory sync (SCIM)** card. 2. Turn on **Enable SCIM provisioning** (organization admins and owners only). 3. Copy the **SCIM base URL** shown on the card, for example: ``` https:///scim/v2 ``` 4. Select **Generate SCIM token** (organization owners only) and copy the plaintext value before leaving the page. It is shown exactly once. Export the values when you want to test the SCIM endpoints from the shell: ``` export AISIX_SCIM_BASE_URL="https:///scim/v2" export AISIX_SCIM_TOKEN="YOUR_SCIM_TOKEN" ``` Configure the SCIM connector in your IdP with the same base URL and token as the bearer credential. The token is an [admin token](https://docs.api7.ai/ai-gateway/cloud/admin-tokens.md) with the exclusive `scim` scope. It can call only the SCIM endpoints and is rejected on every other AISIX Cloud Admin API route, so a leaked IdP credential cannot read or change gateway resources. To rotate the credential, generate a new SCIM token, update the IdP connector, and revoke the old token on the **Admin tokens** page. ## How Users Map to Members[​](#how-users-map-to-members "Direct link to How Users Map to Members") | SCIM attribute | AISIX member field | | ----------------------------- | ----------------------------------------------------------------- | | `userName` / `emails[].value` | Email address (the identity key) | | `displayName` or `name` | Display name | | `externalId` | Stable IdP identifier, unique per organization | | `active` | `false` deactivates the member and disables its API keys | | `groups` | Team membership (read-only; managed through the Groups endpoints) | Creating a user whose email already belongs to a member of the organization returns a `409` uniqueness error. Members created in the dashboard stay dashboard-managed. They are visible to the IdP in list responses, but SCIM cannot modify, deactivate, or delete them. Directory sync can never modify an organization owner. Deleting a user over SCIM disables the member's API keys first, removes the member from every team, and then removes the membership. Usage history is preserved for reporting. The email becomes available again, so a later re-provision creates a fresh member. ### Deactivate vs. Delete[​](#deactivate-vs-delete "Direct link to Deactivate vs. Delete") * `active: false` keeps the member and its keys, but the keys are disabled at the gateway. Setting `active: true` again re-enables only the keys that directory sync disabled. Keys an operator disabled by hand stay disabled. * `DELETE` removes the membership permanently. Keys stay disabled and keep their usage history. ## Map Groups to Teams and Roles[​](#map-groups-to-teams-and-roles "Direct link to Map Groups to Teams and Roles") SCIM groups sync to AISIX [teams](https://docs.api7.ai/ai-gateway/cloud/teams.md): pushing a group creates a team with the same name, and group membership changes add or remove team members. The **Directory sync (SCIM)** card controls the role mapping: * **Default role**: the organization role a synced member gets when they belong to no mapped group — `member`, `admin`, or a [custom role](https://docs.api7.ai/ai-gateway/cloud/custom-roles.md). * **Group → role mappings**: each mapping assigns a role to the members of one directory group, matched by group display name. Any assignable role works as a target, including custom roles. When a member belongs to several mapped groups, an `admin` mapping wins; otherwise the mapping whose group name sorts first applies. Roles of directory-synced members are recomputed on every group change, including group renames, and whenever the mappings or the default role change. A role changed by hand in the dashboard is overwritten by the next sync, so treat the directory as the source of truth for synced members. Directory sync can never assign the `owner` role. ## Use Synced Members for Attribution[​](#use-synced-members-for-attribution "Direct link to Use Synced Members for Attribution") Directory sync creates the members; attribution works the same as for any member: 1. Create or update an API key and set its **owner** to the synced member. 2. Usage and [budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md) are attributed to that member, and member-scoped [rate limit policies](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limit-policies.md) apply. When the IdP deactivates the person, the keys stop authenticating at the gateway within seconds. ## Supported Endpoints[​](#supported-endpoints "Direct link to Supported Endpoints") The SCIM surface lives under `/scim/v2` and answers in `application/scim+json`: | Endpoint | Methods | | ------------------------------------------------------ | ------------------------------------------------------------------------------ | | `/Users` | `GET` (list, `eq` filters on `userName`, `emails.value`, `externalId`), `POST` | | `/Users/{id}` | `GET`, `PUT`, `PATCH`, `DELETE` | | `/Groups` | `GET` (list, `eq` filters on `displayName`, `externalId`), `POST` | | `/Groups/{id}` | `GET`, `PUT`, `PATCH`, `DELETE` | | `/ServiceProviderConfig`, `/ResourceTypes`, `/Schemas` | `GET` | `PATCH` follows the SCIM PatchOp shape, including the path forms Okta and Microsoft Entra ID emit (`active`, `members[value eq "..."]`, and operations without a `path`). List responses page with `startIndex` and `count` (up to 200 per page). Attributes with no AISIX mapping, such as phone numbers or addresses, are accepted and ignored, so IdP attribute mappings do not need trimming. Example: provision a user with `curl`: ``` curl -sS "${AISIX_SCIM_BASE_URL}/Users" \ -H "Authorization: Bearer ${AISIX_SCIM_TOKEN}" \ -H "Content-Type: application/scim+json" \ -d '{ "schemas": ["urn:ietf:params:scim:schemas:core:2.0:User"], "userName": "dev@example.com", "displayName": "Developer One", "emails": [ { "value": "dev@example.com", "primary": true } ] }' ``` ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md) to attach a gateway to an environment. To investigate directory-driven changes later, use [Logging and Auditing](https://docs.api7.ai/ai-gateway/cloud/logging-and-auditing.md). --- # Teams Teams group organization members for ownership, attribution, and shared governance. Budgets, rate limit policies, and MCP access can use the team identity to control traffic from team-bound caller API keys. Team membership and caller API key binding serve different purposes. The roster identifies who belongs to the team. Traffic counts toward team controls only when the caller API key used for the request is explicitly bound to that team. ## Create a Team[​](#create-a-team "Direct link to Create a Team") 1. Open **Teams** and select **New team**. 2. Enter a **Display name** and an optional **Description**. 3. Select **Create team**. The team appears with its member count and **Team ID**. Select the team name or **Manage** to open its detail page. Copy the Team ID when another guide requires `TEAM_ID`. The same UUID appears at the end of the team detail page URL. ``` export TEAM_ID="YOUR_TEAM_ID" ``` ## Add and Manage Members[​](#add-and-manage-members "Direct link to Add and Manage Members") Team members come from the organization member pool. [Create or invite the member](https://docs.api7.ai/ai-gateway/cloud/members.md) before adding them to a team. 1. Open the team detail page. 2. Under **Members**, select **Add member**. 3. Select the **Org member** and assign the team role **Member** or **Lead**. 4. Select **Add member**. The Member and Lead labels belong to the team roster. They do not replace the member's organization role. When a team has one lead, promote another member before demoting or removing the only lead. Use the role selector on a member row to change their team role. Select the remove action to remove the member from the team without removing their organization membership. ## Bind Caller API Keys to the Team[​](#bind-caller-api-keys-to-the-team "Direct link to Bind Caller API Keys to the Team") Team membership alone does not attribute traffic to the team. Bind each caller API key that should use team budgets, rate limits, or MCP access: 1. Open the environment and select **API keys**. 2. Create a key with **New API key**, or edit an existing key. 3. Select the team in **Team**. 4. Optionally select an **Owner**. After choosing a team, the owner list contains only members of that team. 5. Create the key or save the changes. The API key row shows its team binding. A member's personal key does not count toward a team merely because that member appears on the roster. ## Configure Team Controls[​](#configure-team-controls "Direct link to Configure Team Controls") The team detail page provides **Team budget (shared)**, **Per-member budget**, and **MCP entitlement** controls. Related guides cover the complete behavior: * [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md) explains shared team budgets and separate allowances for each member in a team. * [Rate Limit Policies](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limit-policies.md) shows how to match a team and divide its quota into per-member buckets. * [MCP Access Policies](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md#grant-a-team-policy) configures the MCP tools reachable by team-bound keys. Use the dashboard to create, edit, delete, and manage membership for teams. The public AISIX Cloud Admin API supports team entitlements, but team lifecycle and membership operations are not currently part of its public contract. ## Sync Teams from an Identity Provider[​](#sync-teams-from-an-identity-provider "Direct link to Sync Teams from an Identity Provider") [SCIM directory sync](https://docs.api7.ai/ai-gateway/cloud/scim-directory-sync.md) represents an identity provider group as a team with the same name. Group membership changes update the team roster automatically. Manage the roster of a directory-synced team in the identity provider so later synchronization does not overwrite dashboard changes. Directory-synced teams can use the same caller API key bindings, budgets, rate limit policies, and MCP entitlements as teams created in the dashboard. ## Edit or Delete a Team[​](#edit-or-delete-a-team "Direct link to Edit or Delete a Team") Open the team detail page and select **Edit** to change its display name or description. Select **Delete** to remove the team and its roster. The people on that roster remain organization members. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Caller API Keys](https://docs.api7.ai/ai-gateway/traffic-controls/caller-api-keys.md) to create team-bound credentials. Then configure [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md), [Rate Limit Policies](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limit-policies.md), or [MCP Access Policies](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md) for the team. --- # Usage Reporting Use the **Usage** view to compare request volume, token consumption, and spend across AISIX gateway environments and models. It also summarizes organization-wide semantic-cache savings and helps you investigate unexpected spend or confirm that usage records are reaching the control plane. ## How Usage Is Reported[​](#how-usage-is-reported "Direct link to How Usage Is Reported") An AISIX gateway serves AI traffic in your runtime environment and reports usage events to the control plane. It derives the telemetry endpoint from the control-plane URL and sends usage data to the fixed `/dp/telemetry` path. Operators do not configure a separate destination for control-plane usage reporting. Each event can include request status, latency, token usage, and cost. It also distinguishes the model alias requested by the caller from the resolved model that served an attempt, which helps explain routed and ensemble traffic. Streaming chat requests can report time to first token. AISIX uses nonzero provider-reported token counts when available. For streaming requests, AISIX also requests usage data from OpenAI-compatible upstreams in the final stream chunk. The Chat Completions, Completions, Messages, Responses, and Embeddings endpoints may return responses that omit some or all usage data. Examples include an OpenAI-compatible relay that never reports usage, a client that disconnects mid-stream, or an upstream error after a partial response. In those cases, AISIX estimates missing or zero token fields with a local tokenizer. It estimates input tokens from the request sent upstream and output tokens from content delivered to the caller while retaining nonzero provider-reported values. AISIX emits the usage event with the estimated counts, so the request remains included in telemetry, spend, budgets, and token rate-limit accounting. Locally counted events are marked as estimated, and the [Request Logs](https://docs.api7.ai/ai-gateway/cloud/logging-and-auditing.md#request-logs) view shows an `estimated` badge on those rows. Estimated counts closely match OpenAI-family models and are approximations for models with proprietary tokenizers. Estimation does not change the provider response returned to the caller and does not apply to other endpoint families or passthrough routes. ## Interpret Usage[​](#interpret-usage "Direct link to Interpret Usage") The Usage view summarizes AISIX gateway traffic across environments over a rolling 30-day window. All totals and tables use this same window. ![AISIX Cloud control plane Usage view showing spend, request count, token totals, and usage by environment](https://static.api7.ai/uploads/2026/06/25/9wPdvgfS_usage-screen.png) | Signal | What It Helps Explain | | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Request count | Traffic volume and changes in demand. | | Token totals | Input and output consumption reported by providers. | | Spend | Cost calculated by applying matching model rates to token counts and, for duration-priced transcription or translation requests, measured audio length. | | Cache savings | Upstream input and output token consumption avoided by semantic-cache hits. | | Usage by environment | Which deployment environment generated traffic and spend. | | Top models | Up to 10 environment and caller-requested model alias combinations, ranked by spend. Older records without a requested-model value fall back to the resolved model. | | Top API keys | Up to 10 environment and caller API key combinations, ranked by spend. | Where [local token estimation](#how-usage-is-reported) is supported, missing or zero token counts are estimated by the gateway. Audio endpoints do not estimate missing token counts; transcription and translation requests can instead contribute duration-based spend. Spend also requires a price that matches the event's `(provider, model name)` pair and a nonzero rate for the applicable usage basis. Usage rounds spend to two decimal places, so low-volume traffic can show `$0.00` even when a matching price exists. Use [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md) to add or inspect a price, and use [Request Logs](https://docs.api7.ai/ai-gateway/cloud/logging-and-auditing.md#request-logs) to inspect event or attempt cost at six decimal places. Pricing changes apply only to events recorded after the change. ## Usage and Budget Enforcement[​](#usage-and-budget-enforcement "Direct link to Usage and Budget Enforcement") The control plane uses AISIX gateway usage records to evaluate budgets. Each budget defines its scope, spending limit, period, and enforcement mode. The rolling 30-day Usage window is independent of a budget's configured period, so their totals can cover different time ranges. When a hard-stop budget is exceeded, the AISIX gateway can reject matching requests with HTTP `429`. Warn-only budgets surface the over-budget state without blocking traffic. If a budget rejection is unexpected, check the returned budget scope, the configured limit, and the caller API key bindings that determine which team or member budgets apply. For the complete enforcement path, see [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md). ## Troubleshoot Missing or Unexpected Usage[​](#troubleshoot-missing-or-unexpected-usage "Direct link to Troubleshoot Missing or Unexpected Usage") Check the reporting path in order: 1. Confirm that the request used the AISIX gateway attached to the expected environment. Traffic through another environment in the same organization appears under that environment. Traffic in another organization appears in that organization's Usage view. 2. Check that the request completed and appears in Request Logs. A newly completed request can take a few seconds to appear because gateways flush telemetry in batches. 3. Confirm that the AISIX gateway has a recent heartbeat. 4. Confirm that the gateway can reach the control-plane telemetry endpoint. 5. If spend is zero or unexpected, check Model Pricing for an exact provider and model-name match, then confirm that the applicable token or **Audio per minute** rate is nonzero. During a temporary control-plane outage, live traffic can continue with the latest projected configuration. Telemetry batches that fail during the outage are dropped instead of retried, so affected usage records can remain missing after connectivity returns. New records resume after the connection recovers. Exporter health, heartbeat, and fresh budget decisions also require control-plane connectivity. Budget checks can temporarily reuse a cached decision before applying the configured fail mode; see [Availability and Caching](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md#availability-and-caching). ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md) when spend is zero or differs from the provider's pricing basis. To investigate individual requests or control-plane resource changes, see [Logging and Auditing](https://docs.api7.ai/ai-gateway/cloud/logging-and-auditing.md). --- # Configuration Propagation AISIX separates configuration updates from proxy request handling. Every gateway serves requests from its latest applied snapshot, regardless of where its dynamic resources originate. An accepted resource update and proxy readiness are therefore not the same state. Verify important changes through the same caller-facing path the application uses. ## How Updates Reach the Gateway[​](#how-updates-reach-the-gateway "Direct link to How Updates Reach the Gateway") The update trigger depends on the configured resource source: | Resource Source | How an Update Reaches AISIX | | --------------------------------- | ----------------------------------------------------------------------- | | Declarative `resources.yaml` file | The gateway loads the file at startup and re-reads it after `SIGHUP`. | | etcd | The gateway watches its configured keyspace for resource changes. | | AISIX Cloud | The control plane projects environment resources to connected gateways. | Each source feeds the same snapshot-application path: AISIX replaces the loaded configuration atomically after a successful application. New requests use the current snapshot. A request that started before the replacement can continue using the previous snapshot. Invalid updates do not silently replace a valid snapshot. Depending on the source and failure, AISIX either applies the accepted subset and reports rejected resources, or continues serving the last-known-good configuration. ## Reload a Resources File[​](#reload-a-resources-file "Direct link to Reload a Resources File") An open-source AISIX gateway does not watch `resources.yaml` for changes. Validate an edited file before sending `SIGHUP`, then confirm that the gateway applied the new snapshot. Begin with the complete resources file mounted into the running gateway. Merge additions or replacements into that file, preserve unrelated entries and collections, and validate the assembled result. AISIX does not merge a smaller file with the active snapshot; after a successful reload, omitted resources are no longer active. The commands below use the container name and resources path from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md#update-and-reload-the-configuration). Adapt them if your gateway uses a different container name or path. Edit the resources file mounted into the running container. The quickstart provides a complete example that adds a second model and grants the existing caller API key access to it. ### Add New Environment Variables[​](#add-new-environment-variables "Direct link to Add New Environment Variables") A running container cannot inherit environment variables that you export later on the host. If an edited resources file introduces a new `${VAR}` reference, first validate the file in a short-lived container with every required variable: ``` docker run --rm \ -v "$(pwd):/etc/aisix:ro" \ -e OPENAI_API_KEY \ -e CALLER_API_KEY \ -e PROVIDER_VARIABLE_1 \ -e PROVIDER_VARIABLE_2 \ --entrypoint /usr/local/bin/aisix \ ghcr.io/api7/aisix:1.0.0 \ validate --resources /etc/aisix/resources.yaml ``` Replace the provider-variable names with those used in the resources file. If the provider needs only one credential variable, remove the `-e PROVIDER_VARIABLE_2` line from both commands. Re-export `OPENAI_API_KEY` and `CALLER_API_KEY` first if they are not available in the current shell. After validation succeeds, recreate the quickstart container with the same variables: ``` docker rm -f aisix-quickstart docker run -d --name aisix-quickstart \ -v "$(pwd):/etc/aisix:ro" \ -e OPENAI_API_KEY \ -e CALLER_API_KEY \ -e PROVIDER_VARIABLE_1 \ -e PROVIDER_VARIABLE_2 \ -p 3000:3000 -p 9090:9090 \ ghcr.io/api7/aisix:1.0.0 ``` The mounted working directory preserves `resources.yaml` when the old container is removed. The replacement loads the validated file at startup, so skip `SIGHUP` and continue with [Confirm the Applied Configuration](#confirm-the-applied-configuration). ### Reload with Existing Environment Variables[​](#reload-with-existing-environment-variables "Direct link to Reload with Existing Environment Variables") If the edited file does not introduce any new environment variables, validate it inside the running container: ``` docker exec aisix-quickstart \ /usr/local/bin/aisix validate --resources /etc/aisix/resources.yaml ``` This command reuses the running container and its environment. Validation uses the same file-loading pipeline as startup and reload, including environment-variable interpolation, name-reference resolution, and schema validation. An invalid file exits non-zero with the full error report. Successful validation reports that the file loaded and shows its resource count: ``` OK: /etc/aisix/resources.yaml loaded resource(s) ``` Send `SIGHUP` to reload the file: ``` docker kill --signal=HUP aisix-quickstart ``` ### Confirm the Applied Configuration[​](#confirm-the-applied-configuration "Direct link to Confirm the Applied Configuration") Confirm that the new configuration was applied: ``` curl -sS "http://127.0.0.1:9090/status/config" ``` A successful application reports `"state": "synced"` and resource counts that reflect the edit. After a `SIGHUP` reload, `apply_seq` is greater than its previous value. If a reload fails, the gateway keeps serving the last valid configuration, reports `out_of_sync`, and identifies the rejected entries. ## Apply Related Resources in Order[​](#apply-related-resources-in-order "Direct link to Apply Related Resources in Order") Dynamic resources can depend on one another: a model can reference a provider key, and a caller API key can allow that model. When a source delivers resources individually, one accepted resource can become visible before another during a multi-resource change. Apply related resources in dependency order: 1. Create or update the provider key. 2. Create or update the model that references it. 3. Create or update the caller API key that can use the model. 4. Verify the resulting model and request path. This order reduces temporary reference failures, but the final caller-facing check remains the readiness signal for the complete change. ## Verify a Configuration Change[​](#verify-a-configuration-change "Direct link to Verify a Configuration Change") For model access changes, query model discovery with the same caller API key the application will use: ``` AISIX_API_KEY="YOUR_CALLER_API_KEY" curl -sS "http://127.0.0.1:3000/v1/models" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ | jq -r '.data[].id' ``` When automation must wait for a change, poll for the expected model alias instead of sleeping for a fixed interval. After the alias appears, send a request through the exact endpoint and model whose behavior changed. A caller-facing probe confirms more than configuration acceptance. It verifies that the gateway serving the application loaded the resource relationship and can resolve the caller's access. ## Inspect a Delayed Change[​](#inspect-a-delayed-change "Direct link to Inspect a Delayed Change") If the expected behavior does not appear, inspect the configuration state on the affected gateway: ``` curl -sS "http://127.0.0.1:9090/status/config" ``` Use the result to locate the delay: * `source` shows the latest configuration observation and store connectivity where applicable. * `applied` shows the snapshot AISIX is serving. * `rejected` identifies resources that failed validation. * `last_failure` records the latest load error. A `degraded` state means AISIX is serving an accepted subset while reporting rejected resources. An `out_of_sync` state means the latest observation was rejected as a whole and AISIX continues serving the last valid configuration when one is available. For the complete response schema, state meanings, metrics, and alerts, see [Configuration Status](https://docs.api7.ai/ai-gateway/reference/config-status.md). ## Separate Propagation from Request Failures[​](#separate-propagation-from-request-failures "Direct link to Separate Propagation from Request Failures") Use the point of failure to avoid repeating an update that already propagated: * If `/status/config` reports a rejection, correct the rejected resource or source document. * If the source revision does not advance during an expected etcd change, check store connectivity and the watched prefix. * If a model is missing from `GET /v1/models`, check its alias, type, caller access, environment, and applied snapshot. * If the model appears but the provider request fails, troubleshoot the provider credential, endpoint, quota, or network path. * If one gateway differs from its peers, compare the resource source and applied snapshot on each instance. Repeating the same write does not repair a gateway that is not receiving configuration. Establish whether the source, application step, resource relationship, or request path is failing before changing the resource again. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Health Checks](https://docs.api7.ai/ai-gateway/deployment/health-checks.md) to choose probes for process, traffic, configuration, and model health. --- # Forward Proxy for IDE AI Traffic Organizations can place AISIX behind a TLS-terminating egress device to govern traffic from IDEs and coding agents that must keep their official service endpoints. This guide uses GitHub Copilot IDE extensions and Copilot CLI as the worked example. AISIX receives plaintext HTTP from the device. It can apply access control, audit, content inspection, and request limits before relaying each request to the official upstream with the employee's credential. AISIX does not intercept TLS or issue a certificate authority. Clients keep their official service endpoints. Traffic can reach the egress device through explicit proxy configuration or transparent interception; clients must trust the device's certificate authority when it terminates TLS. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Prepare the following: * One AISIX deployment: * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. * For the open-source AISIX gateway, a gateway configured to load a declarative resources file. * A TLS-terminating egress device that can preserve the original `Host` header and inject an HTTP header. * Permission to configure proxy and certificate trust on the client machines. * The current [GitHub Copilot allowlist](https://docs.github.com/en/copilot/reference/copilot-allowlist-reference) and [Copilot network settings](https://docs.github.com/en/copilot/concepts/network-settings). * `curl` and `jq`. The local validation uses `pipx` to install [mitmproxy](https://mitmproxy.org/). ## Traffic Topology[​](#traffic-topology "Direct link to Traffic Topology") The client sends HTTPS traffic to the egress device. The device terminates TLS, preserves the original `Host`, injects the gateway key and employee identity, and sends decrypted HTTP to AISIX. AISIX strips gateway-only headers and relays the request with the employee's upstream credential over HTTPS. The gateway accepts origin-form requests carrying the original `Host`, as used by transparent redirection and proxy chaining. It also accepts absolute-form request targets from a chained proxy by reading the URI authority when `Host` is absent. A matching `hosts` route runs before the gateway's own typed routes, so an upstream path such as `/v1/messages` is relayed instead of being handled as the gateway's Messages endpoint. ## Choose the Hosts to Inspect[​](#route-the-copilot-hosts "Direct link to Choose the Hosts to Inspect") The route examples use the following values in the route's `hosts` field to inspect Copilot inference, suggestions, and selected GitHub API traffic: resources.yaml (route hosts field) ``` hosts: - api.githubcopilot.com - "*.individual.githubcopilot.com" - "*.business.githubcopilot.com" - "*.enterprise.githubcopilot.com" - copilot-proxy.githubusercontent.com - origin-tracker.githubusercontent.com - api.github.com ``` This is not a complete Copilot network allowlist. Authentication, assets, telemetry, experimentation, and editor-specific services may use other hosts. Reconcile the egress policy with GitHub's current allowlist, then decide which hosts the device sends through AISIX and which it permits directly. `api.github.com` is shared by Copilot and other GitHub clients. If the device diverts that host, every request to it from clients using the proxy can match this route. Remove it from the route or narrow diversion at the device when only selected GitHub API paths should pass through AISIX. ## Configure the Copilot Route[​](#configure-the-copilot-route "Direct link to Configure the Copilot Route") The route uses `preserve_host` so one allowlist can relay several official hosts. `header_key` consumes a gateway credential injected by the device, while `forward_client` leaves the employee's upstream `Authorization` available for GitHub. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP includes /api and has no trailing slash. export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_BASE_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create the route: ``` ROUTE_RESPONSE=$(curl --fail-with-body -sS -X POST \ "$AISIX_CP/environments/$ENV_ID/passthrough_routes" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "copilot", "hosts": [ "api.githubcopilot.com", "*.individual.githubcopilot.com", "*.business.githubcopilot.com", "*.enterprise.githubcopilot.com", "copilot-proxy.githubusercontent.com", "origin-tracker.githubusercontent.com", "api.github.com" ], "preserve_host": true, "auth_mode": "header_key", "auth_header_name": "x-aisix-api-key", "credential_mode": "forward_client", "identity_header": "x-aisix-user" }') export ROUTE_ID=$(printf '%s' "$ROUTE_RESPONSE" | jq -er '.passthrough_route.id') printf '%s' "$ROUTE_RESPONSE" | jq '.warnings // []' ``` Keep `ROUTE_ID` if you plan to attach a guardrail specifically to this route. Create a dedicated caller key for the egress device. Its plaintext is returned once: ``` CALLER_RESPONSE=$(curl --fail-with-body -sS -X POST \ "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "Copilot egress device", "allowed_models": [], "allowed_routes": ["copilot"] }') export EGRESS_DEVICE_KEY=$(printf '%s' "$CALLER_RESPONSE" | jq -er '.plaintext') printf '%s' "$CALLER_RESPONSE" | jq '.warnings // []' ``` The same workflow is available in the dashboard under an environment's **Passthrough Routes** page and the caller key's **Passthrough route access** section. Review any returned compatibility warnings before rollout. Warnings are advisory, so verify traffic through each gateway. Route-scoped guardrails can produce their own warning when attached. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Choose the key that the egress device will inject: ``` export EGRESS_DEVICE_KEY="YOUR_GATEWAY_CALLER_KEY" ``` Start with the [complete resources file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely) the gateway currently uses. Add the route and caller key to the matching collections, and keep unrelated resources unchanged: resources.yaml (Copilot route and caller key) ``` passthrough_routes: - name: copilot hosts: - api.githubcopilot.com - "*.individual.githubcopilot.com" - "*.business.githubcopilot.com" - "*.enterprise.githubcopilot.com" - copilot-proxy.githubusercontent.com - origin-tracker.githubusercontent.com - api.github.com preserve_host: true auth_mode: header_key auth_header_name: x-aisix-api-key credential_mode: forward_client identity_header: x-aisix-user api_keys: - display_name: egress-device key_env: EGRESS_DEVICE_KEY allowed_models: [] allowed_routes: [copilot] ``` Validate the assembled complete file: ``` aisix validate --resources resources.yaml ``` Start or recreate the gateway with `EGRESS_DEVICE_KEY` in its process environment. A reload cannot add an environment variable to an already-running process. ## Understand Authentication and Attribution[​](#gateway-authentication-options "Direct link to Understand Authentication and Attribution") The employee's upstream credential remains in `Authorization`, so the gateway credential needs a different channel: * **`header_key`**, shown above, reads the gateway key from `x-aisix-api-key` and strips that header before forwarding. The employee's `Authorization` remains available to GitHub. * **`anonymous`** can be used when the device cannot inject a gateway key. Bind the route to a dedicated caller-key principal and restrict `source_cidrs` to the addresses AISIX resolves for these requests: normally the device addresses, or the original client networks when real-client-IP resolution is enabled. In both modes, the resolved caller key must grant `copilot` in `allowed_routes`. ### Per-Employee Attribution[​](#per-employee-attribution "Direct link to Per-Employee Attribution") AISIX does not authenticate the value in `identity_header`. The egress device must authenticate the employee, remove any client-supplied `x-aisix-user`, and overwrite it with the trusted identity. AISIX records the bounded value as `client_identity` and removes the header before forwarding. Without that header, the resolved source IP normally identifies the egress device rather than the employee. To recover an original client address, configure [real-client-IP resolution](https://docs.api7.ai/ai-gateway/deployment/network-and-security.md) only for the exact trusted device ranges and forwarded header. For an anonymous route, also allow those resolved client networks in `source_cidrs`. ## Apply Audit, Guardrails, and Limits[​](#audit-dlp-and-limits "Direct link to Apply Audit, Guardrails, and Limits") The route does not own rate limits or budgets. Controls resolve from the authenticated caller and, where supported, other attached scopes. ### Request Limits[​](#request-limits "Direct link to Request Limits") Caller API-key, team, and member request limits apply before dispatch. Passthrough routes have no rate-limit field or route policy scope. Use separate caller keys when different host groups need different request or concurrency limits. Token usage from recognized envelopes is recorded on usage events, but it does not increment `tpm` or `tpd` counters. A shared token counter that is already exhausted can reject a request, but passthrough traffic does not advance it. For SSE, a concurrency reservation is released when AISIX returns the streaming response, not when the stream ends. ### Budgets[​](#budgets "Direct link to Budgets") In AISIX Cloud, budgets that already apply to the resolved caller are checked before dispatch. Passthrough usage currently has no model ID, is recorded with zero cost, and does not add spend to the budget ledger. The open-source AISIX gateway has no local budget resource. ### Guardrails[​](#guardrails "Direct link to Guardrails") In AISIX Cloud, attach a guardrail with the **Passthrough routes** scope to inspect only this route. Environment, caller-key, and team scopes can also apply. The open-source resources file declares attachments in its `guardrail_attachments` collection, so a file-defined guardrail reaches this route's traffic only if an attachment scopes it there — `scope_type: env`, or `scope_type: passthrough_route` naming the route. A request block returns `422` before the upstream call. Buffered responses are checked before delivery. After an SSE response starts, a block ends the stream with an SSE `content_filter` error frame. Hold-back guardrails may delay frames while they are inspected. ### Content and Usage Export[​](#content-and-usage-export "Direct link to Content and Usage Export") For successfully relayed traffic, an observability exporter with `content_mode: full` receives the request body as string content rather than a normalized copy of the provider envelope. Buffered responses record extracted text when the response matches a supported extraction shape and otherwise record the body as text. Streamed responses record accumulated extracted text; opaque data payloads are retained as text. All capture remains subject to the configured limits. Captured content is sent only to content-capable exporters. Usage events carry the matched route, caller, recorded token counts, and `client_identity`. The current AISIX Cloud Request Logs UI shows caller and token metadata, but it does not display `passthrough_route_name` or `client_identity`. ## Understand Copilot CLI Traffic[​](#github-copilot-cli-agent "Direct link to Understand Copilot CLI Traffic") GitHub Copilot CLI is an agent that can edit files, run shell commands, select models, and use its built-in GitHub MCP server. Its exact network endpoints and envelope choices are version-sensitive and are not part of this AISIX configuration contract. AISIX detects recognized `messages`, `input`, and `prompt` request shapes for guardrail and usage extraction. Other traffic, including JSON-RPC and ordinary GitHub API calls, is treated as opaque and relayed without body-schema translation. The selected host route therefore does not need a protocol field. GitHub documents `/model`, `/mcp`, `/usage`, and `/context` as CLI commands, but whether a specific command sends network traffic can change with the client version. Validate the current client behavior in your environment rather than relying on a fixed endpoint inventory. ## Validate an Open-Source Gateway with mitmproxy[​](#validate-without-the-production-device "Direct link to Validate an Open-Source Gateway with mitmproxy") The following local exercise uses mitmproxy as the TLS-terminating device. The local quickstart gateway listens on `127.0.0.1:3000`, and mitmproxy listens on `127.0.0.1:8888`. 1. Start the gateway with the Copilot route, caller key, and `EGRESS_DEVICE_KEY` configured above. 2. Install and verify mitmproxy: ``` pipx install mitmproxy mitmdump --version ``` 3. Save the following device script as `mitm_to_aisix.py`. It diverts the selected hosts, restores the original `Host`, injects the gateway key and user identity, and keeps SSE streaming: mitm\_to\_aisix.py ``` import os from mitmproxy import http AISIX_HOST, AISIX_PORT = "127.0.0.1", 3000 GATEWAY_KEY = os.environ["EGRESS_DEVICE_KEY"] IDENTITY = os.environ.get("AISIX_CLIENT_IDENTITY", "alice@example.com") COPILOT_HOSTS = { "api.githubcopilot.com", "copilot-proxy.githubusercontent.com", "origin-tracker.githubusercontent.com", "api.github.com", } COPILOT_SUFFIXES = ( ".individual.githubcopilot.com", ".business.githubcopilot.com", ".enterprise.githubcopilot.com", ) def _matches_one_label(host: str, suffix: str) -> bool: if not host.endswith(suffix): return False label = host[: -len(suffix)] return bool(label) and "." not in label def _diverted(host: str) -> bool: return host in COPILOT_HOSTS or any( _matches_one_label(host, suffix) for suffix in COPILOT_SUFFIXES ) def request(flow: http.HTTPFlow) -> None: host = flow.request.pretty_host if not _diverted(host): return flow.request.host = AISIX_HOST flow.request.port = AISIX_PORT flow.request.scheme = "http" flow.request.headers["host"] = host flow.request.headers["x-aisix-api-key"] = GATEWAY_KEY flow.request.headers["x-aisix-user"] = IDENTITY def responseheaders(flow: http.HTTPFlow) -> None: if "text/event-stream" in flow.response.headers.get("content-type", ""): flow.response.stream = True ``` Setting `flow.request.host` rewrites the `Host` header, so the script restores the upstream host afterward. Without that line, the request matches no host route. The response hook prevents mitmproxy from buffering SSE into one delayed response. 4. Start the proxy with the gateway key in its environment: ``` export AISIX_CLIENT_IDENTITY="alice@example.com" mitmdump -s mitm_to_aisix.py --listen-port 8888 ``` On its first start, mitmproxy creates the local certificate authority used in the next steps. 5. In another terminal, smoke-test proxy diversion, route matching, and relay to GitHub with a credential authorized for the test request: ``` curl -x "http://127.0.0.1:8888" \ --cacert ~/.mitmproxy/mitmproxy-ca-cert.pem \ -H "Authorization: Bearer YOUR_GITHUB_TOKEN" \ "https://api.github.com/user" ``` A `401` naming `x-aisix-api-key` means the device key did not arrive. A `403` means the key does not grant `copilot`. An empty `404` usually means the original host did not match the route. 6. Configure a Copilot client to use mitmproxy and trust its certificate authority. Copilot checks standard proxy variables and `NODE_EXTRA_CA_CERTS`: ``` export HTTPS_PROXY="http://127.0.0.1:8888" export HTTP_PROXY="http://127.0.0.1:8888" export NODE_EXTRA_CA_CERTS="$HOME/.mitmproxy/mitmproxy-ca-cert.pem" copilot ``` For an editor plugin, configure its HTTP proxy setting and make the same certificate authority available to the editor process. Follow GitHub's current network-settings documentation for the client in use. 7. Configure an OTLP/HTTP exporter if needed, send a normal request, and inspect the exported span. Confirm it records `aisix.passthrough.route_name: copilot` and the injected identity as `aisix.client_identity`. For an open-source DLP check, declare a keyword guardrail and attach it — `scope_type: env` to cover the whole gateway, or `scope_type: passthrough_route` naming this route. In AISIX Cloud, attach the guardrail to the route the same way before testing a blocked prompt. ## Scope and Limits[​](#scope-and-limits "Direct link to Scope and Limits") * Passthrough routes do not relay WebSocket upgrades; exclude those hosts or paths from device diversion. * `preserve_host` targets `https://` on port 443. * AISIX performs no TLS interception and includes no certificate-authority tooling. * Copilot hosts and client behavior change independently of AISIX. Recheck GitHub's allowlist and client network documentation when rolling out or upgrading the integration. ## Next Steps[​](#next-steps "Direct link to Next Steps") * Review all route fields and error behavior: [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). * Configure content controls: [Guardrail Behavior](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md). * Export request and response content: [Observability Exporters](https://docs.api7.ai/ai-gateway/observability/exporters.md). --- # Health Checks AISIX exposes separate health and status endpoints for process, traffic, configuration, and model state. The same endpoints apply to the open-source AISIX gateway and to AISIX gateways connected to AISIX Cloud. Use the endpoint that matches the condition you want to monitor. To verify the complete caller-to-provider path, send a test request through the gateway. | Question | Endpoint | Listener | | -------------------------------------------------- | -------------------- | -------------- | | Should this instance be restarted? | `GET /livez` | Proxy | | Should this instance receive proxy traffic? | `GET /readyz` | Proxy | | Has this instance applied any valid configuration? | `GET /status/ready` | Metrics/status | | What configuration is this instance serving? | `GET /status/config` | Metrics/status | | Is a model available for routing? | `GET /status/models` | Metrics/status | These endpoints do not require authentication. The metrics/status endpoints are available when Prometheus metrics are enabled, as they are by default. Keep that listener private to monitoring and operations systems. ## Proxy Liveness[​](#proxy-liveness "Direct link to Proxy Liveness") Use `/livez` for process liveness: ``` curl -i "http://127.0.0.1:3000/livez" ``` A healthy process returns `200 OK` with an `ok` body. A draining instance answers `200` as well. Liveness decides whether to **restart** the instance, and draining is deliberate work: an instance that has been told to shut down is finishing the requests it already accepted, and restarting it would kill exactly those. It is `/readyz` that reports the drain, so that traffic is withdrawn without the instance being replaced under its own in-flight work. Append `?verbose=1` when investigating manually. Do not make an automated probe depend on the verbose response body. Liveness is intentionally narrow. It does not prove that the gateway has loaded configuration, that a model is available, or that a provider request can succeed. ## Traffic Readiness[​](#traffic-readiness "Direct link to Traffic Readiness") Use `/readyz` to decide whether an instance should receive traffic: ``` curl -i "http://127.0.0.1:3000/readyz" ``` It returns `503 Service Unavailable` while the instance is draining or before it applies its first configuration. After a valid configuration is available, the gateway remains ready for as long as it can serve that configuration. A control-plane or configuration-store interruption does not make a running gateway unready merely because no recent update arrived. A stalled source commonly affects every instance, so withdrawing them would remove the traffic path rather than shift traffic to a healthy peer. Monitor configuration freshness separately. Append `?verbose=1` when investigating why an instance is not ready. Do not make an automated probe depend on the verbose response body. For Kubernetes, point the liveness and readiness probes at `/livez` and `/readyz` on the proxy listener. Give the gateway enough termination time to drain in-flight and streaming requests. ## Shutdown and Draining[​](#shutdown-and-draining "Direct link to Shutdown and Draining") A load balancer learns that an instance is withdrawing on its next health check, not the moment the instance decides to. Between those two points it keeps routing new connections. Closing the listener as soon as the shutdown signal arrives would refuse every connection routed inside that interval. Callers then see gateway errors during an ordinary rolling update or scale-down. So the gateway separates the two events. On `SIGTERM` or `SIGINT` it: 1. Answers `/readyz` with `503 Service Unavailable` immediately, so the next health check withdraws it. `/livez` stays `200`: the process is healthy and must not be restarted while it drains. 2. Keeps accepting new connections for at least `shutdown.min_drain_secs`, which defaults to 30 seconds. 3. Adds `Connection: close` to every HTTP/1.1 response, so a client that pools connections retires them as it uses them instead of holding idle ones open. HTTP/2 forbids that header, so an HTTP/2 client is sent a `GOAWAY` frame instead, at the moment the drain starts. `GOAWAY` asks the peer to finish the streams it has already opened and to start no new ones; it does not close the connection or interrupt anything in flight. 4. After the minimum window, waits without a deadline of its own for the in-flight count to reach zero, then stops accepting new connections and exits. Set `min_drain_secs` above the detection latency of whatever load-balances the instance. A Kubernetes readiness probe needs `periodSeconds x failureThreshold`. An external load balancer needs its own check interval multiplied by its retry count. Setting it too low closes the listener while traffic is still arriving. Setting it too high only delays the exit. config.yaml ``` shutdown: min_drain_secs: 30 ``` The window is a minimum, not a deadline. After it elapses the gateway still waits for the in-flight count to reach zero, so a load balancer slower than configured cannot make it close under live traffic. `0` drops the window entirely and is only correct when nothing routes to the instance by health check. Because the wait for in-flight work to reach zero is unbounded, the deployment platform sets the real limit on the whole sequence: `terminationGracePeriodSeconds` in Kubernetes, `TimeoutStopSec` under systemd. Size it above `shutdown.min_drain_secs` plus your longest request or stream, with operational margin. Add the duration of any `preStop` hook because it consumes the same Kubernetes termination budget. note Point the health check of an external load balancer at `/readyz` rather than at a bare TCP connect. A TCP check cannot observe readiness, so the only signal it ever receives is the listener closing — the very event the drain window exists to avoid. ### Watch a Drain in the Logs[​](#watch-a-drain-in-the-logs "Direct link to Watch a Drain in the Logs") The drain logs its own progress, because the record a request normally leaves is written when the request completes — and a request still running when the platform's grace period expires never gets there. These lines are what separate work the gateway had already taken on from traffic still being routed at it after the signal: | Message | Fields | Written | | --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ | ------------------------------------------------------ | | `draining — /readyz now reports 503, still accepting new connections` | `min_drain_secs`, `in_flight`, `open_connections` | Once, when the signal arrives. | | `still draining in-flight requests` | `in_flight`, `open_connections` | Periodically, while the drain runs. | | `request arrived while draining` | `method`, `path`, plus the request's own `request_id`, `peer`, and `downstream_request_id` | Per request that reaches the gateway after the signal. | | `accepted a new downstream connection while draining` | `peer`, `open_connections` | Per connection accepted after the signal. | The two counts answer different questions and neither is derivable from the other. `in_flight` counts requests being served, and a streaming response holds its slot for as long as bytes may still flow. `open_connections` counts downstream connections open on the proxy listener, including pooled ones sitting idle with no request on them. So a drain that will not end while `in_flight` stays up is the gateway finishing work it already accepted; one where `in_flight` has reached zero while `open_connections` stays up is clients holding connections they are not using; one where `open_connections` keeps climbing is something still routing new connections here. Together, the two per-event lines explain a request that failed during a rolling update. An arrival line with a matching accept line for the same `peer` means the connection itself was routed here after the gateway had already asked to be withdrawn — a load balancer that has not caught up, which is the case `min_drain_secs` exists to absorb. An arrival line with no matching accept line means the connection predates the signal and the client reused it out of its pool, which is what the `Connection: close` header retires. The accept line cannot tell your callers from your platform, though. `/livez` and `/readyz` are served on the proxy listener too, and probes keep arriving on their own connections throughout the drain, so they raise `open_connections` and produce accept lines like any other connection. The arrival line is where that judgement is made: it knows the path, and deliberately excludes those two endpoints so the handful of real arrivals is not buried under a probe every few seconds. Read the accept line for what the arrival line cannot see — a connection opened and never used. `peer` and `downstream_request_id` are the connection identity described in [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md#collect-access-logs). Matching them against the records of whatever fronts the gateway is how a gateway line and a balancer line are shown to be about the same connection. ## Configuration State[​](#configuration-state "Direct link to Configuration State") The metrics/status listener provides two configuration checks. `GET /status/ready` is a configuration-only startup gate: * `503 Service Unavailable` before the first valid configuration is applied; * `200 OK` after a valid configuration is available, including while AISIX serves a last-known-good snapshot after a later update fails. `GET /status/config` explains what AISIX observed and applied: ``` curl -sS "http://127.0.0.1:9090/status/config" ``` Use it when a resource update does not appear in proxy behavior. Compare the source and applied state, then inspect rejected resources and the latest load failure. See [Configuration Status](https://docs.api7.ai/ai-gateway/reference/config-status.md) for the complete response fields, state meanings, Prometheus metrics, and alert examples. Configuration status does not replace caller-facing verification. After the expected snapshot applies, query `GET /v1/models` and send the request whose behavior changed. See [Configuration Propagation](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md). ## Per-Model Runtime Health[​](#per-model-runtime-health "Direct link to Per-Model Runtime Health") Use `GET /status/models` when the gateway is ready but routing avoids a model or reports no eligible target: ``` curl -sS "http://127.0.0.1:9090/status/models" ``` Each configured model reports one of these high-level states: * `healthy`: available for routing; * `cooldown`: temporarily removed after recent upstream failures; only a model that enables [cooldown](https://docs.api7.ai/ai-gateway/reference/resources-file.md#direct-models) enters this state; * `unhealthy`: excluded after failed background model checks; * `not_applicable`: a virtual model whose availability derives from its targets. The status view helps identify cooldown and background-check failures, but it does not validate caller access or provider credentials. A model can report healthy while a caller API key, provider key, or upstream response still prevents a request from succeeding. See [Configuration Status](https://docs.api7.ai/ai-gateway/reference/config-status.md#get-statusmodels) for the full response fields. ## Combined Health Signals[​](#combined-health-signals "Direct link to Combined Health Signals") Work from the earliest failing layer: | Signal | Next Check | | ----------------------------------------------------- | ------------------------------------------------------------------------------ | | `/livez` fails. | Process state, listener binding, and listener TLS | | `/readyz` fails. | Drain state and initial configuration availability | | `/status/ready` fails. | Initial configuration source and load errors | | `/status/config` is `degraded` or `out_of_sync`. | Rejected resources, source connectivity, and `last_failure` | | `/status/models` reports `cooldown` or `unhealthy`. | Provider credential, provider availability, model checks, and outbound network | | Every health endpoint succeeds but the request fails. | Caller access, provider path, policy enforcement, and upstream response | Finish with the same path the application uses: ``` AISIX_API_KEY="YOUR_CALLER_API_KEY" curl -sS "http://127.0.0.1:3000/v1/models" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` Then send a real request through the required endpoint and model. That final probe verifies conditions that runtime health endpoints deliberately do not evaluate. ## Next Steps[​](#next-steps "Direct link to Next Steps") Use [Troubleshooting](https://docs.api7.ai/ai-gateway/deployment/troubleshooting.md) to narrow a failed health or request-path check to configuration, caller policy, AISIX Cloud projection, or the upstream provider. --- # Network and Security Treat the proxy listener, metrics/status listener, etcd, and AISIX Cloud control-plane connection as separate trust zones. ## Network Surfaces[​](#network-surfaces "Direct link to Network Surfaces") Use this map to decide what can be reachable from each network: | Surface | What it carries | Exposure | | ------------------------------------ | --------------------------------------------------------------------------- | ----------------------------------------------- | | Proxy listener | Caller-facing AI traffic, such as `/v1/chat/completions` | Intended callers or the ingress tier | | Metrics/status listener | Prometheus metrics plus configuration and model status | Trusted monitoring network | | Configuration store | Dynamic resources for open-source AISIX gateways configured through etcd | AISIX and configuration-management systems only | | AISIX Cloud control-plane connection | mTLS-authenticated gateway communication with the AISIX Cloud control plane | Outbound mTLS path to the control plane | Expose only the proxy listener to callers. Keep metrics/status and etcd on private networks. The metrics path and `/status/*` routes are unauthenticated and served on the dedicated metrics/status listener. Model status includes model IDs and display names, so do not expose this listener publicly. Use TLS or mTLS where network placement requires encrypted or mutually authenticated transport. When etcd mTLS is enabled, the startup configuration must point to the CA, client certificate, and client key files that the gateway process can read. ## Expose the Proxy Listener[​](#expose-the-proxy-listener "Direct link to Expose the Proxy Listener") Keep the gateway on a non-privileged container port unless it must bind directly below `1024`. Publish the caller-facing port through a load balancer, ingress, or service instead of widening access to the other listeners. The published AISIX image runs as non-root UID `10001`. Its gateway binary carries the effective `CAP_NET_BIND_SERVICE` file capability, so it can bind ports such as `80` and `443` without running as root. A runtime policy that prevents the capability from being granted can make the container fail at startup with `exec: Operation not permitted`. Terminate TLS either on the AISIX proxy listener or at the trusted ingress tier in front of it. When another proxy terminates or forwards traffic, configure real-client-IP resolution only for the exact trusted proxy ranges and forwarded header. For a TLS-terminating egress path that sends GitHub Copilot traffic to AISIX while clients keep their official service endpoints, see [Forward Proxy for IDE AI Traffic](https://docs.api7.ai/ai-gateway/deployment/forward-proxy.md). For AISIX gateways that run on Kubernetes and connect to AISIX Cloud, see [Deploy AISIX Gateways on Kubernetes](https://docs.api7.ai/ai-gateway/cloud/kubernetes.md) for Service exposure, container capabilities, autoscaling, and disruption handling. ## Credential Boundaries[​](#credential-boundaries "Direct link to Credential Boundaries") Caller credentials and upstream provider credentials have different storage and forwarding rules: | Credential or Secret | How AISIX Uses It | Protect | | --------------------------------- | ----------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | | Caller API key | Stored as a hash; plaintext is held by the calling application | Application secret storage and API key rotation workflow | | OIDC-issued JWT | Signature and claims are verified, then the identity maps to a caller API key | Application token storage, OIDC trust-provider policy, and identity-provider availability | | Provider key | Used to authenticate upstream provider requests | Resources file or configuration store, environment variables, and backups | | Observability exporter credential | Sent to or resolved for the configured telemetry destination | Dynamic resource store and observability configuration access | | AISIX Cloud certificate bundle | Authenticates an AISIX gateway to the control plane | Runtime state directory, trust root, and control-plane bootstrap process | For open-source AISIX gateways, provider credentials can be referenced from environment variables in a resources file or stored in the configuration store. OTLP HTTP exporter headers can be stored in dynamic configuration. Object-storage, Aliyun SLS, and Datadog exporters use credential references that the gateway resolves locally. Treat resources files, etcd, environment variables, and backups as secret-bearing surfaces. In AISIX Cloud deployments, the control plane handles provider keys and projects runtime configuration to the AISIX gateway. AISIX authenticates each caller with a caller API key or an OIDC-issued JWT and authenticates upstream requests with provider credentials. A verified JWT maps to a caller API key before AISIX applies access and traffic controls. Passthrough strips sensitive inbound headers by default. Provider-key strip settings can affect what is forwarded, so keep credential-bearing headers such as `authorization`, `cookie`, and `x-api-key` stripped unless a specific upstream integration requires otherwise. ## Security Baseline[​](#security-baseline "Direct link to Security Baseline") Confirm these controls before routing production traffic: * The proxy listener is reachable only from intended callers or ingress. * Metrics endpoints are private. * etcd is private, persistent, backed up, and access-controlled when it stores open-source AISIX gateway resources. * Provider-key secrets and exporter credentials are treated as sensitive operational data. * OIDC trust providers pin the expected issuer and audiences, and their discovery and JWKS endpoints are reviewed as outbound trust destinations. * Listener TLS, etcd mTLS, or AISIX Cloud mTLS is configured where the deployment requires encrypted or mutually authenticated transport. * For gateways connected to AISIX Cloud, gateway identity is validated through the certificate-based bootstrap path. * Backups and observability pipelines do not expose dynamic resource payloads or provider credentials. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [TLS and mTLS](https://docs.api7.ai/ai-gateway/deployment/tls-and-mtls.md) to configure transport security for each connection in your deployment. --- # Performance and Sizing The AISIX gateway runtime is a compiled Rust proxy, so it adds very little latency of its own to a request. This page reports measured proxy latency and throughput on a reference machine. It also gives a formula to size CPU for your traffic. These figures follow the benchmark pattern that Kong, LiteLLM, and TensorZero publish, so you can compare AISIX against their published numbers. AISIX was pinned to **4 vCPUs** on a dedicated AWS `c7i.4xlarge`, proxying OpenAI-compatible `/v1/chat/completions` requests with a \~1000-token prompt. The upstream was a near-zero-latency mock, and no traffic-control policies were attached. The gateway used the shared runtime (`proxy.thread_per_core: false`). As a result, the reported latency is essentially the gateway's own overhead. ## Measured Gateway Latency[​](#measured-gateway-latency "Direct link to Measured Gateway Latency") The gateway's own processing stays in the **sub-millisecond** range at low-to-moderate load and grows as CPU fills. Because the test upstream returns in \~0.07 ms, the gateway latency below is effectively the overhead AISIX adds on top of a real upstream: | Offered load | Throughput (req/s) | Gateway latency p50 / p95 / p99 | | -------------- | ------------------ | ------------------------------- | | Light (20%) | 5,700 | 0.31 / 0.51 / 0.59 ms | | Moderate (40%) | 11,300 | 0.52 / 0.89 / 1.04 ms | | Busy (60%) | 17,000 | 0.82 / 1.37 / 1.69 ms | | Heavy (80%) | 22,600 | 1.12 / 2.14 / 2.54 ms | | Saturated | 28,300 | — | Gateway overhead can also be isolated by subtracting direct-to-upstream latency at the same rate. With that calculation, p50 overhead is **0.24 ms** at light load and **0.99 ms** at heavy load. A real LLM call takes hundreds of milliseconds to several seconds, so this gateway overhead is negligible end to end. ## Throughput and CPU[​](#throughput-and-cpu "Direct link to Throughput and CPU") On 4 vCPUs, a single AISIX instance sustained about **28,300 req/s** for this workload. CPU usage scales almost perfectly linearly with request rate: ``` CPU% ≈ 14 + 0.0144 × (req/s) # per instance, in % of one vCPU ``` That is roughly **0.14 ms of one CPU core per request** plus a small fixed runtime cost. This linear fit holds across the tested 20-80% range. Near saturation, the curve flattens against the 4-vCPU ceiling. The 28,300 req/s peak drew about 383% CPU, not the \~421% that a naive extrapolation implies. This flattening is one reason to size below saturation. Throughput scales horizontally, either by adding vCPUs to an instance or by adding replicas behind a load balancer. How an instance turns its vCPUs into throughput is governed by its worker pool; see [Thread-per-Core Workers](https://docs.api7.ai/ai-gateway/deployment/thread-per-core-workers.md). ## Streaming Responses[​](#streaming-responses "Direct link to Streaming Responses") In the no-policy benchmark, AISIX relayed server-sent event (SSE) tokens as they arrived from the upstream rather than buffering the whole response. Time-to-first-token overhead was about **0.65 ms**, and the total stream duration matched the upstream. Output guardrails can instead hold streamed content in windows or buffer the full response before releasing it. See [Streaming Output](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md#streaming-output). ## Sizing Your Deployment[​](#sizing-your-deployment "Direct link to Sizing Your Deployment") For the shared-runtime benchmark above, estimate the vCPUs a single instance needs for a target request rate `Q` (req/s): ``` vCPUs ≈ (14 + 0.0144 × Q) / 100 ``` | Target throughput | vCPUs (approx) | | ----------------- | -------------- | | 5,000 req/s | \~0.9 | | 10,000 req/s | \~1.6 | | 25,000 req/s | \~3.7 | | 50,000 req/s | \~7.3 | | 100,000 req/s | \~14.5 | Use the estimate with the following deployment guidance: * Leave headroom. Size for roughly 70-80% of saturation, not 100%, so latency stays low under bursts. * Scale out for high availability and aggregate throughput. Run multiple replicas behind a load balancer; per-instance overhead stays flat as you add replicas. * Account for traffic controls. These figures are a pure-proxy baseline. Authentication, rate limiting, guardrails, caching, and request logging each add per-request work. Measure with your policy set enabled. * Benchmark representative request shapes. Larger prompts and response bodies cost more per request; streaming and non-streaming differ. Benchmark with representative traffic before committing capacity. ## Benchmark Setup[​](#benchmark-setup "Direct link to Benchmark Setup") The reference machine was a dedicated AWS `c7i.4xlarge` (16 vCPU). AISIX ran pinned to 4 vCPUs, with a load generator and a canned near-zero-latency mock upstream on separate cores. That isolation keeps the upstream and the load generator from ever becoming the bottleneck. Reported latency is the gateway's own overhead at a given rate, with no policies attached. Your results will vary with hardware, request shape, and enabled policies. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Startup Configuration](https://docs.api7.ai/ai-gateway/deployment/startup-configuration.md) to configure resource sources, listeners, shared runtime state, and process observability. --- # Production Readiness Prepare each AISIX gateway for production before sending application traffic through it. The same runtime checks apply whether the gateway loads standalone resources or receives configuration from AISIX Cloud; differences in configuration and operational ownership are called out where they matter. This page provides the production baseline. The remaining pages in this section explain capacity, startup configuration, network security, transport security, configuration updates, and health checks in detail. ## Plan Capacity and Availability[​](#plan-capacity-and-availability "Direct link to Plan Capacity and Availability") Estimate the capacity of each gateway instance from representative request and response sizes, streaming behavior, and enabled policies. Leave headroom for bursts instead of sizing at measured saturation. See [Performance and Sizing](https://docs.api7.ai/ai-gateway/deployment/performance-and-sizing.md) for the reference benchmark and sizing formula. Run multiple instances when the traffic path must survive a process or host failure. Put them behind a load balancer, spread them across the failure domains that matter to your service, and ensure every instance receives the same dynamic resources. Memory-backed rate-limit counters and cache entries are local to one gateway process. Use a shared Redis deployment when limits or cached responses must remain consistent across instances. For gateways connected to AISIX Cloud, see [High Availability](https://docs.api7.ai/ai-gateway/cloud/high-availability.md) for the wider traffic and management topology. ## Prepare Configuration and Dependencies[​](#prepare-configuration-and-dependencies "Direct link to Prepare Configuration and Dependencies") Confirm how the gateway receives dynamic resources before deploying replicas: * An open-source AISIX gateway can load a declarative `resources.yaml` file and reload it on `SIGHUP`. * An open-source AISIX gateway configured through etcd watches an etcd keyspace. * A gateway connected to AISIX Cloud receives environment resources projected by the control plane. Keep process settings such as listeners, resource source, Redis connections, and observability in [startup configuration](https://docs.api7.ai/ai-gateway/deployment/startup-configuration.md). Do not configure multiple dynamic-resource sources on the same gateway. Protect every secret-bearing configuration source. A resources file should reference credentials through environment variables rather than contain plaintext values. A standalone etcd deployment needs persistence, backups, access control, and a private network path. Gateways connected to AISIX Cloud need their certificate bundle and runtime state directory protected and available across restarts. Treat Redis as a runtime dependency when cache policies or rate limits select it. Use a topology that matches the availability requirement, and verify that every gateway instance points to the intended Redis deployment. ## Secure Runtime Surfaces[​](#secure-runtime-surfaces "Direct link to Secure Runtime Surfaces") Expose the proxy listener only to intended callers or the ingress tier in front of the gateway. Keep the following surfaces on private networks: * the metrics/status listener and its unauthenticated `/metrics` and `/status/*` routes; * standalone etcd endpoints; * files and directories containing provider credentials or AISIX Cloud certificates. Configure listener TLS when AISIX terminates HTTPS. Configure etcd mTLS or AISIX Cloud mTLS for the corresponding management connection. Review [Network and Security](https://docs.api7.ai/ai-gateway/deployment/network-and-security.md) and [TLS and mTLS](https://docs.api7.ai/ai-gateway/deployment/tls-and-mtls.md) before exposing the gateway. ## Verify before Routing Traffic[​](#verify-before-routing-traffic "Direct link to Verify before Routing Traffic") Liveness proves that the process is running; it does not prove that an authorized request can reach a provider. Verify the runtime from the narrowest signal to the complete request path. Set the proxy and metrics/status URLs: ``` PROXY_URL="https://gateway.example.com" METRICS_URL="http://gateway.internal.example.com:9090" ``` Check process and traffic readiness: ``` curl -i "${PROXY_URL}/livez" curl -i "${PROXY_URL}/readyz" ``` Inspect configuration and model status: ``` curl -sS "${METRICS_URL}/status/config" curl -sS "${METRICS_URL}/status/models" ``` Then exercise behavior the health endpoints cannot prove: * Call `GET /v1/models` with the caller API key the application will use. * Send a provider-backed request through every endpoint family you plan to expose. * Find the request in logs, metrics, usage reporting, or the configured exporter. * Send an invalid caller API key and confirm that AISIX rejects it. Use [Health Checks](https://docs.api7.ai/ai-gateway/deployment/health-checks.md) to choose and interpret operational probes. If a configuration change has not reached the proxy path, see [Configuration Propagation](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md). ## Plan Shutdown and Recovery[​](#plan-shutdown-and-recovery "Direct link to Plan Shutdown and Recovery") AISIX handles `SIGINT` and `SIGTERM` as graceful shutdown signals. When draining starts, `/readyz` immediately returns `503` so the traffic layer can withdraw the instance, while `/livez` remains healthy to prevent a restart during the drain. The proxy continues accepting new connections for at least `shutdown.min_drain_secs`, then stops accepting after that minimum window has elapsed and the in-flight request count reaches zero. See [Shutdown and Draining](https://docs.api7.ai/ai-gateway/deployment/health-checks.md#shutdown-and-draining) for the complete sequence and load-balancer timing guidance. Give the load balancer or orchestration platform enough time to stop assigning connections before the process exits. Set the termination grace period longer than `shutdown.min_drain_secs` plus the longest request or stream you intend to preserve, with additional margin for shutdown and load-balancer timing. Document how each deployment recovers its configuration: * retain and validate the standalone resources file; * back up and restore standalone etcd; * preserve the AISIX Cloud certificate bundle, gateway identity, and latest accepted configuration according to the deployment design. Test restart and replacement behavior before depending on it during an outage. ## Production Checklist[​](#production-checklist "Direct link to Production Checklist") Before widening traffic, confirm that: * Capacity includes headroom and the intended number of gateway instances. * Every instance uses the intended startup configuration and dynamic-resource source. * Shared Redis is configured wherever cross-instance limits or cache entries are required. * Proxy, metrics/status, and configuration-store surfaces have the intended network exposure. * TLS, mTLS, certificate paths, and runtime-state permissions are valid. * At least one provider key, model alias, and caller API key are available to the gateway. * The model alias appears through the caller-facing proxy path. * A real provider-backed request succeeds. * Logs, metrics, usage reporting, or exporters contain the verification request. * Shutdown, replacement, and configuration recovery have been tested. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Performance and Sizing](https://docs.api7.ai/ai-gateway/deployment/performance-and-sizing.md) to plan gateway capacity, then configure the runtime with [Startup Configuration](https://docs.api7.ai/ai-gateway/deployment/startup-configuration.md). --- # Startup Configuration Startup configuration defines the process-level settings AISIX needs before it can serve traffic. It controls how the gateway receives dynamic resources, which listeners it binds, which shared backends it uses, and how it connects to AISIX Cloud. Dynamic resources are configured separately. Models, provider keys, caller API keys, guardrails, cache policies, rate-limit policies, and observability exporters come from a resources file, a configuration store, or the AISIX Cloud control plane. ## Understand the Configuration Sources[​](#understand-the-configuration-sources "Direct link to Understand the Configuration Sources") AISIX uses different configuration sources for different responsibilities: | Configuration | Applies To | Source | | ---------------------------------- | ---------------------------------- | ----------------------------------------------------------------------- | | Process settings | Every gateway | Startup configuration file with optional environment-variable overrides | | Dynamic gateway resources | Open-source AISIX gateway | Declarative `resources.yaml` file or etcd | | Dynamic gateway resources | Gateway connected to AISIX Cloud | AISIX Cloud control plane | | On-Premises control-plane settings | On-Premises AISIX Cloud deployment | Docker Compose environment variables or Helm values | This page covers gateway startup configuration. For On-Premises control-plane settings, see the [On-Premises Configuration Reference](https://docs.api7.ai/ai-gateway/reference/on-premises-configuration.md). ## Choose the Dynamic-Resource Source[​](#choose-the-dynamic-resource-source "Direct link to Choose the Dynamic-Resource Source") A gateway reads dynamic resources from one source. Choose that source in startup configuration before configuring listeners and runtime dependencies. ### Resources File[​](#resources-file "Direct link to Resources File") Use a declarative resources file for the normal open-source gateway workflow: config.yaml ``` resources_file: /etc/aisix/resources.yaml proxy: addr: "0.0.0.0:3000" admin: enabled: false observability: metrics: prometheus: enabled: true path: "/metrics" addr: "0.0.0.0:9090" ``` The gateway loads the file at startup and reloads it on `SIGHUP`. Validate resource changes before applying them: ``` aisix validate --resources /etc/aisix/resources.yaml ``` See the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) for the complete workflow and the [Resources File Reference](https://docs.api7.ai/ai-gateway/reference/resources-file.md) for supported resource kinds. ### Configuration Store[​](#configuration-store "Direct link to Configuration Store") Use etcd when an existing automation system manages open-source AISIX gateway resources through a shared store: config.yaml ``` etcd: endpoints: - "http://127.0.0.1:2379" prefix: "/aisix" ``` Keep the prefix stable across gateway instances that should receive the same resources. Use `env_id` only when the configuration system writes environment-scoped keys. Configure `etcd.tls` when the store requires mTLS. Do not configure both `resources_file` and `etcd`; AISIX rejects that combination at startup. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") The generated gateway installation snippet sets `managed.enabled` to `true` and supplies the control-plane endpoint, certificate bundle, runtime state directory, and gateway identity path. The gateway then receives environment resources from the control plane. Do not combine `managed.enabled: true` with `resources_file`. Follow [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md) for certificate issuance and the complete connection workflow. ## Configure Shared Runtime State[​](#configure-shared-runtime-state "Direct link to Configure Shared Runtime State") AISIX always has an in-process response cache. Configure Redis when cache policies should share entries across gateway instances: config.yaml ``` cache: redis: mode: single url: "redis://127.0.0.1:6379" ``` Configuring `cache.redis` makes Redis available to cache policies. Each cache policy selects memory or Redis. Rate-limit counters use process memory by default. Configure a shared Redis backend when request, token, or concurrency limits must apply across multiple instances: config.yaml ``` ratelimit: backend: redis redis: mode: single url: "redis://127.0.0.1:6379" ``` Cache and rate-limit Redis connections support single-node, cluster, and Sentinel deployments. Use an available Redis topology when either feature is part of the production traffic path. ## Configure Runtime Listeners[​](#configure-runtime-listeners "Direct link to Configure Runtime Listeners") AISIX separates caller traffic from operational status: | Listener | Purpose | Exposure | | -------------- | ------------------------------------------------ | -------------------------------- | | Proxy | Caller-facing AI APIs and `/livez` and `/readyz` | Intended callers or ingress tier | | Metrics/status | Prometheus metrics and `/status/*` routes | Trusted monitoring network | Set an explicit proxy address. The metrics/status routes do not require application authentication, so never expose that listener publicly. Use listener TLS only when AISIX should terminate HTTPS directly. Resolve forwarded client addresses only when the gateway runs behind a trusted load balancer or ingress. See [Network and Security](https://docs.api7.ai/ai-gateway/deployment/network-and-security.md) for the exposure model and [TLS and mTLS](https://docs.api7.ai/ai-gateway/deployment/tls-and-mtls.md) for the independent TLS contexts. ## Configure Process Observability[​](#configure-process-observability "Direct link to Configure Process Observability") Startup observability settings control process logging and the metrics/status listener: config.yaml ``` observability: service_name: "aisix" log_level: "info" metrics: prometheus: enabled: true path: "/metrics" addr: "0.0.0.0:9090" ``` These settings are different from dynamic observability exporters. Startup settings control the process and local Prometheus listener. Configure runtime telemetry delivery through [Observability Exporters](https://docs.api7.ai/ai-gateway/observability/exporters.md). ## Configure Shutdown Behavior[​](#configure-shutdown-behavior "Direct link to Configure Shutdown Behavior") On `SIGTERM` the gateway reports itself unready straight away. It keeps accepting new connections for a further window, so the load balancer in front has time to withdraw it before the listener closes: config.yaml ``` shutdown: min_drain_secs: 30 ``` Set the window above the detection latency of whatever load-balances the instance, and give the platform enough termination time for the in-flight drain that follows. See [Shutdown and Draining](https://docs.api7.ai/ai-gateway/deployment/health-checks.md#shutdown-and-draining). ## Load and Verify the Configuration[​](#load-and-verify-the-configuration "Direct link to Load and Verify the Configuration") AISIX applies built-in defaults, then the startup configuration file, then `AISIX_` environment-variable overrides. Use `__` between nested field names: ``` export AISIX_PROXY__ADDR="0.0.0.0:3000" ``` See the [Startup Configuration Reference](https://docs.api7.ai/ai-gateway/reference/configuration-files.md) for file formats, field behavior, file selection, and loading precedence. See [Environment Variables](https://docs.api7.ai/ai-gateway/reference/environment-variables.md) for override syntax. After starting the gateway, check the proxy listener: ``` curl -i "http://127.0.0.1:3000/livez" curl -i "http://127.0.0.1:3000/readyz" ``` When the metrics/status listener is enabled, confirm that AISIX has applied a valid configuration: ``` curl -sS "http://127.0.0.1:9090/status/config" ``` If the proxy starts but resources do not appear, check the configured resource source rather than changing listener settings. Use [Configuration Propagation](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md) to trace a change from its source to the caller-facing path. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Thread-per-Core Workers](https://docs.api7.ai/ai-gateway/deployment/thread-per-core-workers.md) to size the worker pool behind the proxy listener, then [Network and Security](https://docs.api7.ai/ai-gateway/deployment/network-and-security.md) to protect the listeners, stores, and credentials configured here. --- # Thread-per-Core Workers The AISIX gateway serves traffic from a pool of worker threads. Two startup settings control that pool: `proxy.workers` sets how many worker threads serve traffic, and `proxy.thread_per_core` selects how those workers share the proxy listener and the upstream connections. This page explains both settings and their defaults, shows how to verify which mode a running gateway uses, and covers the behaviors that follow from serving with independent workers. ## Understand the Serving Modes[​](#understand-the-serving-modes "Direct link to Understand the Serving Modes") In thread-per-core serving, each worker is an independent thread with its own runtime, its own listener on `proxy.addr`, and its own pool of upstream connections. The workers share the proxy address through the `SO_REUSEPORT` socket option, and the kernel assigns each new client connection to one worker by hashing the connection's addresses and ports. A request is then accepted, processed, and answered on a single thread, and the upstream call it makes stays on that thread as well. With `proxy.thread_per_core: false`, the gateway serves from one shared runtime instead: a single listener accepts connections, and the worker threads balance tasks among themselves through work stealing. A request can then migrate between threads during processing, for example when an upstream response arrives on a different thread from the one that started the request. Each migration costs a thread wake-up and a context switch. Thread-per-core serving keeps a request on one thread end to end, so those handoffs do not occur. Thread-per-core serving is **on by default on Linux** and off by default elsewhere, because the connection spreading it relies on is Linux kernel behavior. In the validation benchmarks for this mode, it sustained 54-88% more requests per second than the shared runtime on a 4-worker x86 virtual machine, and 28-42% more on Arm (AWS Graviton) hardware. The p99 latency under load roughly halved in both cases. The difference is gateway throughput capacity; provider latency dominates end-to-end request time in either mode. ## Configure the Worker Pool[​](#configure-the-worker-pool "Direct link to Configure the Worker Pool") Both settings live in the `proxy` block of the startup configuration: config.yaml ``` proxy: addr: "0.0.0.0:3000" thread_per_core: true workers: 4 ``` | Field | Default | Description | | ----------------------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ | | `proxy.thread_per_core` | `true` on Linux, `false` elsewhere | Serve from independent workers with per-worker listeners and upstream pools. Set `false` to serve from one shared runtime on any platform. | | `proxy.workers` | Parallelism available to the process | Number of worker threads that serve traffic, in either mode. Must be at least `1`. | Both settings are read once at process start. Changing either one takes effect at the next gateway restart. Leave `proxy.workers` unset in most deployments. The default follows the parallelism actually available to the process, so a container CPU limit, a `cgroup` quota, or a `taskset` affinity mask sizes the pool without restating the count in configuration. A gateway limited to 4 vCPUs starts 4 workers even on a 16-core host. Set an explicit count only when the gateway should use fewer threads than it may run on, for example when it shares its CPU allowance with a sidecar. AISIX rejects `proxy.workers: 0` at startup, because zero workers would bind no listener at all. Omit the field to use the default instead. Both fields follow the standard environment-override form, with double underscores between nested field names: ``` export AISIX_PROXY__THREAD_PER_CORE=false export AISIX_PROXY__WORKERS=8 ``` Deployments that inject all startup configuration through environment variables, such as Kubernetes installs, set the fields this way. See [Environment Variables](https://docs.api7.ai/ai-gateway/reference/environment-variables.md) for the override mechanism. ## Verify the Active Mode[​](#verify-the-active-mode "Direct link to Verify the Active Mode") The serving mode is visible without any dedicated endpoint. At startup, a gateway in thread-per-core mode logs one `aisix listening (http, thread-per-core)` line per worker, each carrying its worker index, or `(https, thread-per-core)` when the proxy listener terminates TLS. The shared runtime logs a single `aisix listening (http)` line. On a running process, list its threads: ``` ps -T -p "$(pgrep -x aisix)" ``` In thread-per-core mode, the worker threads are named `tpc-0` through `tpc-`, one per worker: ``` PID SPID TTY TIME CMD 23110 23110 ? 00:00:00 aisix 23110 23111 ? 00:00:00 tokio-runtime-w 23110 23112 ? 00:00:00 tokio-runtime-w 23110 23115 ? 00:00:41 tpc-0 23110 23116 ? 00:00:40 tpc-1 23110 23117 ? 00:00:41 tpc-2 23110 23118 ? 00:00:39 tpc-3 ``` In this mode, a small control runtime carries the metrics listener, signal handling, and background export work. Its `tokio-runtime-w` threads appear alongside the workers and do not count toward `proxy.workers`. With `thread_per_core: false`, no `tpc-` threads exist, and all serving threads carry the default runtime thread name `tokio-runtime-w`. ## Account for Per-Worker Connection Pools[​](#account-for-per-worker-connection-pools "Direct link to Account for Per-Worker Connection Pools") In thread-per-core serving, each worker keeps its own pool of connections to upstream hosts. Pooled connections are never handed between workers: a worker that holds an idle connection to a provider reuses it, and a worker that does not opens its own. Two sizing consequences follow: * `upstream.pool_max_idle_per_host` applies **per worker**. A gateway with 8 workers and `pool_max_idle_per_host: 32` can hold up to 256 idle connections to one provider host. To bound the process as a whole, divide the intended total by the worker count. * The idle upstream connections a process holds scale with the worker count. Account for that when a provider, a NAT gateway, or a corporate egress proxy limits connections per client. The pool timeouts keep their meaning unchanged; see [Tune the Upstream Connection Layer](https://docs.api7.ai/ai-gateway/reference/configuration-files.md#tune-the-upstream-connection-layer) for the pool settings themselves. ## Plan for Low-Concurrency Traffic[​](#plan-for-low-concurrency-traffic "Direct link to Plan for Low-Concurrency Traffic") The kernel spreads client connections across workers, not individual requests. That spreading is even across many connections and uneven across few. Below about **four client connections per worker**, some workers can sit idle while their siblings carry several connections each. In that range, throughput can fall below the shared runtime, which balances individual requests instead of connections. What counts is the number of connections the gateway itself accepts, not the number of end clients. Direct traffic from many clients sits far above the threshold. However, an L7 load balancer or ingress in front of the gateway, as [Network and Security](https://docs.api7.ai/ai-gateway/deployment/network-and-security.md) recommends for the caller-facing port, can pool requests from many clients onto a small number of gateway-side connections. Benchmarks that drive the gateway from a handful of persistent connections sit in the same range: a load test with 8 connections against 8 workers measures connection placement, not gateway capacity. If the gateway receives its traffic over only a few long-lived connections, raise the number of connections the tier in front of it keeps toward the gateway, lower `proxy.workers` so that each worker still receives several connections, or set `proxy.thread_per_core: false`. ## Understand Listener Sharing[​](#understand-listener-sharing "Direct link to Understand Listener Sharing") Thread-per-core workers share one proxy address through the `SO_REUSEPORT` socket option. Two properties of that option are worth knowing when operating or auditing a gateway host. **Port conflicts still fail loudly at startup.** Before binding its workers' listeners, the gateway probes the address with a regular bind. Starting a gateway while another process holds the address, including a second gateway, therefore fails at startup, exactly as a single-listener process does. The probe leaves one theoretical gap: two gateways started in the same instant can both pass their probe and bind the same address. Sequential restarts, as orchestrators perform them, never hit this window; avoid deliberately starting two gateways on one address at the same moment. **A same-user process can join the listener while the gateway runs.** While the gateway is serving, any process running under the same effective user ID that itself sets `SO_REUSEPORT` can bind the proxy address and receive a share of new connections. The kernel restricts joining to the same effective user ID, so this is not a cross-user risk. It is a property of the mode worth remembering when auditing what runs on a gateway host: if some connections appear to bypass the gateway, check for another process bound to the proxy port, for example with `ss -tlpn`. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Network and Security](https://docs.api7.ai/ai-gateway/deployment/network-and-security.md) to protect the proxy and metrics listeners, or return to [Performance and Sizing](https://docs.api7.ai/ai-gateway/deployment/performance-and-sizing.md) to size CPU for your request volume. --- # TLS and mTLS AISIX AI Gateway uses TLS in four different places. Configure each area that applies to the connections in your deployment. | Connection | Configuration Area | Purpose | | ------------------------------ | ------------------------------------- | ---------------------------------------------------------------- | | Caller to AISIX | `proxy.tls` | Terminates HTTPS on the proxy listener | | AISIX to an upstream | `upstream.tls`, `provider_key.tls` | Decides which certificates AISIX trusts when it calls out | | AISIX to etcd | `etcd.tls` | Verifies the etcd server and presents the AISIX client identity | | AISIX gateway to control plane | `managed.*` certificate bundle fields | Authenticates the AISIX gateway to the AISIX Cloud control plane | These settings are independent. Enabling HTTPS on the proxy listener does not configure etcd mTLS, trusting a private certificate authority for an upstream does not affect etcd, and the AISIX Cloud certificate bundle does not replace listener TLS. ## Configure Listener TLS[​](#configure-listener-tls "Direct link to Configure Listener TLS") Use listener TLS when AISIX should terminate HTTPS directly on the proxy listener. Configure TLS on the proxy listener: config.yaml ``` proxy: addr: "0.0.0.0:3000" tls: cert_file: "/etc/aisix/tls/proxy.crt" key_file: "/etc/aisix/tls/proxy.key" ``` Listener TLS protects inbound traffic to that listener. It does not prove that AISIX can connect to etcd, reach the AISIX Cloud control plane, or authenticate to an upstream provider. ## Trust an Upstream Behind a Private Certificate Authority[​](#trust-an-upstream-behind-a-private-certificate-authority "Direct link to Trust an Upstream Behind a Private Certificate Authority") AISIX verifies outbound connections against the platform's certificate authorities. A self-hosted endpoint whose certificate is signed by your own authority is not among them, so requests to it fail: ``` transport error: error sending request for url (https://internal-llm.example:8443/v1/chat/completions): client error (Connect): invalid peer certificate: UnknownIssuer ``` The outbound transport determines which deployment-wide `upstream.tls` fields it can apply: | Outbound transport | `ca_file` | Client certificate and key | `verify: false` | | -------------------------------------------------------------------------------------------------------------------- | --------- | -------------------------- | ------------------------------------------------------- | | HTTP providers, HTTP guardrails, MCP and A2A upstreams, OIDC and JWKS requests, and OTLP, SLS, and Datadog exporters | Applies | Applies | Applies | | Realtime WebSocket | Applies | Ignored | Applies | | Amazon Bedrock models and guardrails | Applies | Ignored | Ignored; Bedrock always verifies the server certificate | | Object-store exporters | Applies | Ignored | Applies | Per-provider-key `tls` settings have a narrower scope. They apply to HTTP provider dispatch, including compatible REST endpoints and passthrough requests, but not to Amazon Bedrock or the Realtime WebSocket. For Realtime, use the deployment-wide CA and verification settings. For Bedrock, only the deployment-wide `ca_file` setting applies. Point `upstream.tls.ca_file` at the authority's certificate to fix this for the whole deployment: config.yaml ``` upstream: tls: ca_file: "/etc/aisix/tls/private-ca.pem" ``` The file is PEM-encoded and may contain several certificates, so a full chain in one bundle works. These certificates are trusted **in addition to** the platform's own, so adding a private authority never makes a public provider unreachable. If AISIX cannot read the file, or the file contains no certificate, startup fails with the path in the message rather than the connection failing later on every request. ### Present a Client Certificate[​](#present-a-client-certificate "Direct link to Present a Client Certificate") Some HTTP upstreams require the caller to authenticate with a certificate as well. Set both fields together: config.yaml ``` upstream: tls: ca_file: "/etc/aisix/tls/private-ca.pem" client_cert_file: "/etc/aisix/tls/client.crt" client_key_file: "/etc/aisix/tls/client.key" ``` Setting only one of the two is rejected at startup. Realtime, Bedrock, and object-store exporter connections do not present this client certificate. Do not use these fields as an mTLS control for those transports. ### Trust a Different Authority per Endpoint[​](#trust-a-different-authority-per-endpoint "Direct link to Trust a Different Authority per Endpoint") Add this provider key entry to trust a different certificate authority for one endpoint. The certificate stays with the resource rather than the gateway configuration: resources.yaml (provider TLS) ``` provider_keys: - display_name: internal-llm provider: openai api_key: "sk-..." api_base: "https://internal-llm.example:8443/v1" tls: ca_cert: | -----BEGIN CERTIFICATE----- MIIB... -----END CERTIFICATE----- ``` In AISIX Cloud, the same settings are on the provider key's **Endpoint TLS** section in the dashboard. On supported HTTP paths, a certificate configured on one provider key applies only to that key's endpoint. Keys are not pooled into one trust store, so two endpoints signed by two different authorities each need their own. Bedrock and Realtime accept a provider key for credentials and endpoint selection, but their dispatch transports do not consume its `tls` block. Use the applicable deployment-wide setting from the table above instead. ### Skip Verification in a Test Environment[​](#skip-verification-in-a-test-environment "Direct link to Skip Verification in a Test Environment") Certificate verification can be turned off, deployment-wide or for one provider key: config.yaml ``` upstream: tls: verify: false ``` danger This accepts any certificate for the affected connections, including an expired one, one issued for a different host, and one presented by an interceptor. Anyone able to intercept the connection can read and rewrite every prompt, response, and upstream API key that crosses it. Use `ca_file` or `ca_cert` in any environment where that matters. `upstream.tls.verify: false` applies to HTTP request paths, the Realtime WebSocket, and object-store exporters, but not to Amazon Bedrock. The AWS SDK exposes additional trust roots but does not let AISIX disable certificate verification. Bedrock therefore continues to verify the server certificate, and the gateway logs a warning when it first builds the Bedrock client. The per-provider-key `tls.verify` field also does not apply to Bedrock or Realtime. ### Use the Environment Instead[​](#use-the-environment-instead "Direct link to Use the Environment Instead") `SSL_CERT_FILE` and `SSL_CERT_DIR` are honoured, and are additive to the platform's certificate authorities. They remain a valid way to trust a private authority without changing the configuration file. They apply to the whole process, so they cannot express "this one endpoint, this one authority." Use `upstream.tls.ca_file` for an explicit deployment-wide authority, or provider-key `tls.ca_cert` for one supported HTTP provider endpoint. note Installing the certificate into the container's system trust store with `update-ca-certificates` does **not** work. The image runs as the unprivileged `aisix` user, so the command fails with a permission error, leaves the bundle unchanged, and the gateway still rejects the certificate — while looking as though the authority was installed. ### Configure a Redis Backend[​](#configure-a-redis-backend "Direct link to Configure a Redis Backend") The shared cache and rate-limit backend keeps its own trust settings, because it usually sits inside your deployment and is issued by a different authority than the model endpoints. The fields match `upstream.tls`, and apply only to a `rediss://` URL: config.yaml ``` ratelimit: backend: redis redis: mode: single url: "rediss://redis.internal:6379" tls: ca_file: "/etc/aisix/tls/redis-ca.pem" ``` In Sentinel mode, `ca_file` does not apply. The client library accepts no custom trust roots for the master it discovers, so put the certificate in the system trust store or point `SSL_CERT_FILE` at it instead. The gateway logs a warning at startup when `ca_file` is set in this mode. `verify` does apply in Sentinel mode. ## Configure etcd mTLS[​](#configure-etcd-mtls "Direct link to Configure etcd mTLS") Use `etcd.tls` when the configuration store requires mTLS. AISIX expects the CA certificate, client certificate, and client key together. Configure etcd trust and client identity: config.yaml ``` etcd: endpoints: - "https://etcd.internal.example.com:2379" prefix: "/aisix" tls: ca_cert_file: "/etc/aisix/etcd/ca.crt" client_cert_file: "/etc/aisix/etcd/client.crt" client_key_file: "/etc/aisix/etcd/client.key" ``` AISIX uses the CA file to verify the etcd server certificate and presents the client certificate and key to etcd. All three files must be readable by the AISIX process at startup. When the etcd certificate uses a server name that differs from the endpoint hostname, set `domain_name` explicitly: config.yaml ``` etcd: endpoints: - "https://10.0.0.10:2379" tls: ca_cert_file: "/etc/aisix/etcd/ca.crt" client_cert_file: "/etc/aisix/etcd/client.crt" client_key_file: "/etc/aisix/etcd/client.key" domain_name: "etcd.internal.example.com" ``` If `domain_name` is omitted, AISIX derives it from the first etcd endpoint. ## Configure AISIX Cloud mTLS[​](#configure-aisix-cloud-mtls "Direct link to Configure AISIX Cloud mTLS") AISIX gateways authenticate to the control plane with a certificate bundle. This is separate from listener TLS and etcd mTLS for an open-source AISIX gateway. Set `managed.enabled` to `true`, provide the control-plane connection settings, and supply the certificate bundle: config.managed.yaml ``` managed: enabled: true cp_base_url: "https://dpm.example.com:7944" mtls_dir: "/var/lib/aisix/mtls" dp_id_file: "/var/lib/aisix/dp_id" cp_cert_file: "/etc/aisix/mtls/client.crt" cp_key_file: "/etc/aisix/mtls/client.key" cp_ca_file: "/etc/aisix/mtls/ca.crt" ``` The AISIX Cloud certificate bundle must include a certificate, private key, and CA bundle. The example uses file paths; AISIX also accepts inline PEM values. Provide all three values through the same style, and do not set both the inline and file-path variant for the same certificate, key, or CA role. Most AISIX gateways derive the control-plane etcd endpoint from `cp_base_url`. Set `cp_etcd_endpoint` only when the control-plane deployment exposes a separate, known etcd endpoint. AISIX materializes the certificate bundle into `mtls_dir` and reuses the persisted bundle on restart. The runtime state directory must be writable by the gateway process. For the full AISIX Cloud connection flow, see [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md). ## Check the Right Connection[​](#check-the-right-connection "Direct link to Check the Right Connection") Start with the failing connection and check the matching configuration area. If HTTPS caller traffic fails while the process is running, check `proxy.tls`, certificate and key readability, and the client-facing hostname. If startup fails while connecting to etcd, check `etcd.tls`, etcd network reachability, and certificate trust. If expected configuration changes stop applying after startup, keep the focus on the etcd connection and configuration watch health. If a request to a provider fails with `invalid peer certificate: UnknownIssuer`, the endpoint's certificate is signed by an authority AISIX does not trust — check `upstream.tls.ca_file`, or, for an HTTP provider endpoint, the provider key's own `tls.ca_cert` when only that endpoint is affected. If AISIX Cloud heartbeat, telemetry, budget checks, or certificate rotation fail, check the certificate bundle, trust root, runtime state directory, and `managed.cp_base_url`. Each TLS area uses its own certificate context. Listener certificates, upstream trust settings, etcd client certificates, and AISIX Cloud control-plane certificates are configured and validated separately. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Configuration Propagation](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md) to understand how validated updates become active gateway snapshots. --- # Troubleshooting Narrow an AISIX failure to startup, configuration, caller policy, the provider path, or AISIX Cloud projection before changing resources. Work through these checks in order until the failing layer is clear, then use the matching section to decide the next action. ## Fast triage[​](#fast-triage "Direct link to Fast triage") Start with the runtime path before changing configuration. Check listener health first: ``` curl -i "http://127.0.0.1:3000/livez" ``` When Prometheus metrics are enabled, check the latest observed and applied configuration on the private metrics/status listener: ``` curl -sS "http://127.0.0.1:9090/status/config" ``` Verify model discovery with the same caller API key the application uses: ``` AISIX_API_KEY="YOUR_CALLER_API_KEY" curl -sS "http://127.0.0.1:3000/v1/models" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` Then send one real request to the endpoint that failed. For example, check the OpenAI-compatible chat path with the caller API key: ``` curl -sS "http://127.0.0.1:3000/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-4o-mini", "messages": [ { "role": "user", "content": "Say hello." } ] }' ``` Use the response status, error type, and correlation header to choose the next check. | Signal | What to Check Next | | --------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | | Listener health fails. | Process, listener binding, startup config, and TLS. | | Configuration status is `degraded` or `out_of_sync`. | Rejected resources, the last load failure, and the configuration source. | | Configuration status does not reflect an expected direct etcd change. | etcd reachability, the watched prefix, and configuration source state. | | Model discovery does not show the expected alias. | Caller API key access, model type, and configuration visibility. | | Real proxy request fails before upstream dispatch. | Caller authentication, model access, guardrails, rate limits, or budgets. | | Real proxy request reaches the provider and fails. | Provider key, base URL, upstream model ID, quota, outage, or outbound network path. | For request correlation, use `x-aisix-request-id`, which AISIX adds to proxy responses across endpoint families. Successful Chat Completions responses also include `x-aisix-call-id`. For exact header scope, see [Headers and Error Codes](https://docs.api7.ai/ai-gateway/reference/headers-and-error-codes.md#proxy-response-headers). ## Verify Configuration Visibility[​](#verify-configuration-visibility "Direct link to Verify Configuration Visibility") Use this step when resources were created or updated recently, or when the proxy behaves as if it has older configuration. | Signal | Check | Action | | --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Process fails during startup. | Dynamic-resource source, network reachability, TLS certificate paths, file permissions, and startup config syntax. | Fix the startup configuration or resource source, then restart the gateway. | | Configuration status is `never_loaded`, `degraded`, or `out_of_sync`. | `/status/config` fields `source`, `applied`, `rejected`, and `last_failure`. | Restore the source or correct rejected configuration, then confirm that the next load is `synced`. | | A direct etcd change does not appear in proxy traffic. | Applied gateway configuration snapshot, store connectivity, and the watched prefix. | Verify the final proxy path after the snapshot updates; see [Configuration Propagation](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md). | | Model discovery omits a new alias. | Caller API key allow list, model type, and snapshot freshness. | Correct the model or caller API key, then query model discovery again. | | Error mentions a missing provider key or unknown resource. | Provider key, model, and caller API key references. | Create or correct dependent resources in order, then send a real proxy request. | For a store-backed gateway, etcd provides the dynamic-resource source. If etcd becomes unavailable after a snapshot has loaded, the gateway can continue using that snapshot, but it cannot receive new or updated dynamic resources until the configuration connection recovers. ## Verify Caller Access and Policy[​](#verify-caller-access-and-policy "Direct link to Verify Caller Access and Policy") Use this step when AISIX rejects the request before a provider call, or when model discovery differs between caller API keys. | Signal | Check | Action | | --------------------------------- | ------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------- | | Authentication error. | The application sends the plaintext caller API key, not the stored hash or upstream provider key. | Update the application secret or authorization header. | | Permission or model-access error. | The requested model alias is allowed by the caller API key. | Add the alias to the caller API key or request an allowed model. | | Content-policy error. | Enabled guardrails, the stage where each guardrail runs, and the triggering prompt or response content. | Adjust the prompt, guardrail rules, or fail-open behavior where appropriate. | | Rate-limit or budget error. | Retry hint, API-key limit, model limit, shared policy, AISIX Cloud budget state, and replica-local counters. | Wait for the retry window, increase the limit, or adjust the matching policy. | For exact proxy error envelopes, status codes, and retry headers, see [Proxy Errors and Retries](https://docs.api7.ai/ai-gateway/routing/proxy-errors-and-retries.md). When rate limits differ across gateway instances, identify the counter backend before changing a policy. Memory counters are local to each process. Redis shares counters, but a runtime Redis failure falls back to process-local enforcement. Outage-time counts are not reconciled into Redis after recovery, so restore Redis and let the active windows roll over before assuming the counters are aligned again. ## Verify the Provider Path[​](#verify-the-provider-path "Direct link to Verify the Provider Path") Use this step after AISIX authenticates the caller, resolves the model alias, and starts dispatching to the configured provider. | Signal | Check | Action | | -------------------------------------- | ------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------- | | `502` or `upstream_error`. | Provider key secret, base URL, upstream model ID, provider quota, provider outage, and outbound network path. | Fix the provider path, then send the same proxy request again. | | `503` with provider unavailable. | Provider adapter availability and whether the resolved adapter supports the requested route. | Use a supported provider, adapter, endpoint, or model. | | `503` with all candidates unavailable. | Multi-target model health, cooldown state, and routing filters. | Restore a healthy target or adjust routing behavior. | | Model health is degraded or down. | Recent upstream failure streak, provider outage, quota, outbound network path, and credential validity. | Restore provider reachability or route traffic to a healthy target. | Compare the failing route with the matching provider upstream guide when the problem is provider-specific. ### Read a Transport Error[​](#read-a-transport-error "Direct link to Read a Transport Error") A transport error means the request never completed at the HTTP layer, so there is no upstream status code to interpret. Gateway logs record the cause chain alongside the message, which is what distinguishes the possible faults: | Cause in the log line | Meaning | Where to look | | --------------------------------------------------------------------- | --------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `dns error: failed to lookup address information` | The provider hostname did not resolve. | DNS configuration, the `api_base` hostname, and egress DNS policy. | | `tcp connect error: Connection refused` | Nothing accepted the connection on that host and port. | The `api_base` port, and whether an intermediate proxy is listening. | | `tcp connect error: Connection timed out` | The connection attempt was black-holed. | Firewall, security group, and egress rules on the path. | | `connection closed before message completed` or a reset while sending | A pooled connection was closed by the far end before the request completed. | `upstream.pool_idle_timeout_secs`; see [Tune the Upstream Connection Layer](https://docs.api7.ai/ai-gateway/reference/configuration-files.md#tune-the-upstream-connection-layer). | | A TLS or certificate failure | The handshake with the provider or an intercepting proxy failed. | Trust roots, and whether a TLS-terminating proxy re-signs traffic. | Errors that appear only intermittently against a provider that is otherwise healthy usually point to connection reuse rather than the provider. Lower `pool_idle_timeout_secs` below the shortest idle timeout on the path and re-check. The same failure has a mirror image on the inbound side. If a gateway or load balancer **in front** of AISIX reports intermittent resets or 502s against a healthy AISIX, check whether `downstream.idle_timeout_secs` is set below that node's own pool idle timeout — the gateway would then be closing connections the node in front still considers usable. Leaving it at `0` (the default) rules this out. See [Tune the Downstream Connection Layer](https://docs.api7.ai/ai-gateway/reference/configuration-files.md#tune-the-downstream-connection-layer). ### Attribute an Upstream Error to the Right Hop[​](#attribute-an-upstream-error-to-the-right-hop "Direct link to Attribute an Upstream Error to the Right Hop") When the recorded message is `upstream returned HTTP : `, the status and body came from whatever answered at `api_base` — which is the provider only when nothing sits in front of it. Response bodies that mention a proxy's own vocabulary indicate an intermediate hop generated the error rather than the model provider: * `upstream connect error or disconnect/reset before headers` with a `reset reason`, or `exceeded request buffer limit while retrying upstream`, are emitted by Envoy-based proxies, service meshes, and API gateways. * A generic HTML error page is emitted by a reverse proxy or load balancer, not by a provider's JSON API. On modeled routes, the gateway records these bodies for diagnosis but does not return them to the caller. Upstream `5xx` responses reach the application as a generic `502` envelope. [Passthrough routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md#errors) instead forward the upstream provider's native status and body. Use the per-attempt records to identify the failing `provider_key_id` and target model, then compare them with access logs from the upstream endpoint configured in `api_base`. ## Verify AISIX Cloud Projection[​](#verify-aisix-cloud-projection "Direct link to Verify AISIX Cloud Projection") Use this step when control-plane state and live gateway behavior do not match, or when an AISIX gateway cannot receive projected configuration. | Signal | Check | Action | | -------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | AISIX Cloud heartbeat fails. | Certificate identity, trust roots, runtime state, control-plane URL, and outbound network path. | Restore AISIX Cloud connectivity before investigating resource projection; see [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md). | | The control plane shows a resource but live traffic does not use it. | Environment scope and projection status for the gateway handling traffic. | Move the resource to the correct environment or wait for projection to complete. | | AISIX Cloud budget checks fail or appear unavailable. | Control-plane connectivity, budget policy target, and budget-check response details. | Restore budget-check connectivity or correct the AISIX Cloud policy. | | Playground succeeds but live traffic differs. | Live gateway, environment, model alias, caller key, and provider target. | Send the live request through the intended AISIX gateway and environment. | Use the relevant feature guide or reference page after identifying the failing layer. Re-run the caller-facing request after each correction to verify the complete path rather than only the component that failed. --- # URL Rewriting AISIX AI Gateway can rewrite the request path before routing. An ordered list of rewrite rules runs at the entry of the proxy listener: the first rule whose `match` regex matches the request path rewrites it, and the request then flows through the normal endpoint — authentication, access control, rate limits, and telemetry apply exactly as if the client had sent the rewritten path. Use URL rewriting when existing clients are configured with URL shapes AISIX does not serve natively — for example, gateways that expose one URL per MCP server, or internal conventions like a health path an external monitor already probes. ## Configure Rewrite Rules[​](#configure-rewrite-rules "Direct link to Configure Rewrite Rules") Rewrite rules are startup configuration in `config.yaml`, under the `proxy` block: ``` proxy: addr: "0.0.0.0:3000" url_rewrites: - name: per-server-mcp-compat match: "^/mcp-servers/([^/]+)/mcp$" rewrite: "/mcp/$1" ``` | Field | Required | Description | | --------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `name` | No | Label used in gateway logs when the rule fires. | | `match` | Yes | Regex tested against the raw, percent-encoded request path (never the query string), with no decoding or normalization. Anchor with `^` and `$` to match the whole path. | | `rewrite` | Yes | Replacement for the matched portion of the path. `$1` … expand numbered capture groups, and `${server}` expands a named group defined as `(?P…)` in `match`. Use the braced form (`${1}text`) when literal text follows a reference. Must not contain `?`, `#`, or whitespace — the query string is preserved automatically. | Changing rules requires a gateway restart, like any other startup configuration. Startup validation rejects broken rules with an error naming the rule — invalid regexes, references to capture groups the pattern does not define, forbidden characters in the template, and patterns that would match the empty string — so a typo surfaces immediately instead of as silently mis-routed traffic. When the gateway is configured purely through environment variables, such as a Helm-managed deployment, set the whole list as one JSON array: ``` export AISIX_PROXY__URL_REWRITES='[{"name":"per-server-mcp-compat","match":"^/mcp-servers/([^/]+)/mcp$","rewrite":"/mcp/$1"}]' ``` ## How Rules Are Applied[​](#how-rules-are-applied "Direct link to How Rules Are Applied") * Rules run in declaration order; the **first** match wins and is applied **once** — a rewritten path is never fed back through the rule list. * `rewrite` replaces the matched portion of the path. With an unanchored `match`, the unmatched prefix and suffix are preserved. * The query string is preserved as sent. * A request no rule matches is passed through untouched, and canonical paths keep working alongside the rewritten shapes. * Rewriting applies to every request on the proxy listener. The admin and metrics listeners are not affected. Rewriting selects which gateway endpoint serves the request; it never bypasses governance. The rewritten request is authenticated and authorized by its target endpoint, and metrics and access logs record the rewritten route. ## Example: Serve Per-Server MCP URLs[​](#example-serve-per-server-mcp-urls "Direct link to Example: Serve Per-Server MCP URLs") Some gateways expose each MCP server at its own URL, such as `/mcp-servers/github/mcp`, and clients call tools by their original names. AISIX serves that contract natively on its [per-server MCP endpoint](https://docs.api7.ai/ai-gateway/mcp-gateway/overview.md#per-server-endpoints) `/mcp/{server}` — one rewrite rule connects the legacy URL shape to it: ``` proxy: url_rewrites: - name: per-server-mcp-compat match: "^/mcp-servers/([^/]+)/mcp$" rewrite: "/mcp/$1" ``` A client configured with `https://gateway.example.com/mcp-servers/github/mcp` now reaches `/mcp/github`. AISIX normally lists the `github` server's tools under their original names and accepts calls with those names, so no client changes are needed. If an original name would be ambiguous with a registered server prefix, AISIX advertises the namespaced spelling that remains callable. The caller API key's tool access, rate limits, and guardrails apply unchanged. Any matching AISIX Cloud budgets also continue to apply. Only the URL shapes you declare are served: with the rule above, `/mcp-servers/github/sse` matches nothing and returns 404 rather than being silently routed. ## Example: Alias an Arbitrary Path[​](#example-alias-an-arbitrary-path "Direct link to Example: Alias an Arbitrary Path") Rules are not MCP-specific. Any path can map onto any proxy endpoint: ``` proxy: url_rewrites: - name: legacy-health match: "^/healthz-compat$" rewrite: "/livez" ``` --- # Anthropic-Style Messages API Anthropic-style proxy routes fit applications that already send Anthropic Messages requests and need AISIX to manage gateway-side authentication, model aliases, routing, and policy. The client keeps the Anthropic-style request and response format. AISIX becomes the endpoint the client calls, and the upstream provider can be Anthropic or another supported provider family. This guide describes the proxy API behavior. For a runnable client integration, see [Anthropic SDK](https://docs.api7.ai/ai-gateway/getting-started/anthropic-sdk.md). ## What the Client Sends[​](#what-the-client-sends "Direct link to What the Client Sends") The client sends three AISIX-owned values: * The base URL is the AISIX gateway origin, without a trailing slash or endpoint path. * The API key is an AISIX caller API key. * The model value is an AISIX model alias, such as `claude-prod`. The request body keeps the Anthropic Messages format, including `messages`, `max_tokens`, `tools`, and `stream`. The caller uses the AISIX caller API key, not the upstream Anthropic provider key. AISIX also accepts a `system` role inside `messages[]` for clients that send that shape. When the upstream path needs Anthropic's native format, AISIX maps leading system messages to Anthropic's top-level `system` field. Anthropic SDKs send the caller API key as `x-api-key`. AISIX also accepts bearer tokens for direct HTTP clients. Export the gateway connection and request values used below: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="claude-prod" ``` ## Messages Request[​](#messages-request "Direct link to Messages Request") Send a Messages request through AISIX: ``` curl -sS -X POST "${AISIX_PROXY}/v1/messages" \ -H "x-api-key: ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "max_tokens": 128, "messages": [ { "role": "user", "content": "Say hello from AISIX." } ] }' ``` A successful response uses the Anthropic Messages format. The `model` value in the request is the AISIX model alias, not necessarily the upstream provider model ID. ## Choose an Upstream Path[​](#choose-an-upstream-path "Direct link to Choose an Upstream Path") Like other AISIX proxy APIs, `/v1/messages` lets the client request format stay stable while the upstream provider changes. For Anthropic-style requests, the upstream choice matters because a native Anthropic-protocol route preserves more Anthropic-specific behavior than a translated upstream. | Upstream path | What it gives you | Use it when | | ------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | | Native Anthropic-protocol route | Native Anthropic request and response behavior, with AISIX handling the caller key, provider key, and model alias. This includes provider keys that use the `anthropic` adapter and keys that declare [`apis.messages`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#declare-the-api-surfaces). | The application depends on Anthropic-specific behavior such as thinking blocks, image blocks, cache control, or exact tool-use semantics. | | Translated upstream | An Anthropic-style client edge with a non-Anthropic upstream behind AISIX. | The application needs to keep an Anthropic-style client while AISIX routes traffic to another supported provider family. | The translated path supports text, vision, and tool-calling flows end to end. `text`, `image` (base64 and URL), and `document` blocks translate to multimodal content parts on the upstream wire. Assistant `tool_use` history becomes upstream tool calls, and `tool_result` blocks become the tool-response turns the upstream expects, so multi-turn tool loops keep their history across providers. `thinking` and `redacted_thinking` history blocks are dropped on translation because another vendor cannot replay Anthropic's signed reasoning blocks. Prefer a native Anthropic-protocol route when the application depends on those reasoning blocks or other provider-specific request fields. ### Request Fields on the Translated Path[​](#request-fields-on-the-translated-path "Direct link to Request Fields on the Translated Path") When the provider key does not select a native Anthropic-protocol route, AISIX rewrites the Anthropic request fields into the shape the upstream family expects instead of forwarding them verbatim: | Anthropic field | Translated behavior | | ------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `tools`, `tool_choice` | Translated to the OpenAI tool-calling shape. | | `stop_sequences` | Sent as `stop`. | | `metadata.user_id` | Sent as `user`. | | `thinking`, `output_config.effort` | Mapped to `reasoning_effort`. See [Reasoning Effort](#reasoning-effort) for the resolution order and the tiers AISIX forwards. | | `output_format`, `output_config.format` | A `json_schema` block is sent as `response_format` in the OpenAI `json_schema` form with strict mode enabled. Strict mode requires every object in the schema to close over its properties, so AISIX adds `additionalProperties: false` and lists each declared property in `required` at every object level. Any other shape is dropped. When a request carries both fields, `output_format` wins. | | `context_management`, `top_k`, `mcp_servers`, `container`, `service_tier`, `betas`, and other Anthropic-only fields | Dropped. Forwarding them would fail OpenAI-compatible upstreams with unknown-parameter errors, so requests from newer Anthropic SDKs keep working as new fields appear. | Native Anthropic-protocol routes do not go through this translation. Apart from resolving the `model` alias, AISIX forwards the request fields unmodified, so every Anthropic field keeps its native behavior on that path. #### Reasoning Effort[​](#reasoning-effort "Direct link to Reasoning Effort") An Anthropic request can carry the reasoning depth in two places. `output_config.effort` is the current control, and the only one that Claude Opus 4.7 and later models accept. `thinking.budget_tokens` is the older one, deprecated on Claude Opus 4.6. An OpenAI-compatible upstream has a single `reasoning_effort` field for both, so AISIX resolves them in this order: 1. `thinking.type: disabled` sends `reasoning_effort: none`. An explicit opt-out is a stronger instruction than a depth tier, so an `output_config.effort` alongside it does not override it. 2. `output_config.effort` is sent as `reasoning_effort`, with the tier unchanged. 3. `thinking.type: enabled` maps `budget_tokens` to a tier: `0-1023` minimal, `1024-2047` low, `2048-4095` medium, `4096` and above high. 4. `thinking.type: adaptive` with no `output_config.effort` sends `reasoning_effort: high`, the tier Anthropic itself applies when a request omits an effort. AISIX forwards the tier the request asked for and does not check it against the upstream model. Anthropic models accept tiers such as `max` and `xhigh` that many OpenAI-compatible models do not, and an upstream that does not accept a tier rejects the request. That rejection is deliberate. Substituting a tier the upstream happens to accept would quietly change the reasoning depth the application asked for, and that is far harder to notice than an error. Pick a tier the upstream model supports, and check the provider's own documentation for its range. Upstream models with no reasoning support at all reject `reasoning_effort` itself, so send thinking and effort fields only to reasoning-capable models. Streaming and non-streaming requests take the same translation. ## Route Behavior[​](#route-behavior "Direct link to Route Behavior") `/v1/messages` can use direct and routing model aliases. Non-streaming requests can fail over to the next target on retryable upstream failures. Streaming requests can fail over before AISIX sends response bytes to the client. AISIX does not switch targets after the client-visible stream starts. For general streaming behavior, see [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md). `POST /v1/messages/count_tokens` uses the same AISIX caller API key and accepts the Anthropic token-counting request format. This route only uses targets whose provider key uses the `anthropic` adapter or declares `apis.messages`. A declaration covers both Messages routes, so confirm that the upstream also implements `/v1/messages/count_tokens`; otherwise, use a translated model for `/v1/messages` or a passthrough route for the provider's native Messages API. If no native Anthropic-protocol target is available, AISIX rejects the request. ## Handle Errors[​](#handle-errors "Direct link to Handle Errors") Messages routes return errors in an Anthropic-style envelope. The error type follows Anthropic SDK-compatible status mappings, so Anthropic clients can parse gateway-generated errors with the same error-handling path they use for provider errors. Native Anthropic upstream errors can include `request_id`. AISIX does not add that field to gateway-generated Anthropic-style errors. When an output guardrail holds a stream back, a buffered frame that AISIX cannot parse is dropped rather than released unscanned. A response left with nothing to return ends with a terminal SSE `error` event. That event uses the Anthropic envelope like every other error on this route. It therefore carries an Anthropic-legal `error.type` and no `code` field, rather than the `content_filter` an OpenAI-style route would return. The reason is named in the message instead. See [Frames AISIX Cannot Scan](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md#frames-aisix-cannot-scan) and [Guardrail Refusals](https://docs.api7.ai/ai-gateway/reference/headers-and-error-codes.md#guardrail-refusals). For the full error and header reference, see [Headers and Error Codes](https://docs.api7.ai/ai-gateway/reference/headers-and-error-codes.md). For provider-defined error types, see Anthropic's [Errors documentation](https://docs.anthropic.com/en/api/errors). ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen how Anthropic-style clients call AISIX. Continue with [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md), [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md), or [Proxy Errors and Retries](https://docs.api7.ai/ai-gateway/routing/proxy-errors-and-retries.md) when your application depends on those behaviors. --- # Speech and Audio Applications use audio to transcribe or translate recorded speech, generate spoken output, add audio to a model turn, or sustain a live conversation. These workflows need different request and delivery models. AISIX gives each workflow its own interface. Across them, the gateway authenticates callers, resolves model aliases, applies the access restrictions and rate limits supported by that interface, and records request telemetry. ## Choose an Audio Interface[​](#choose-an-audio-interface "Direct link to Choose an Audio Interface") Choose the endpoint based on how audio participates in the application: | Goal | Endpoint | | ------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | | Transcribe or translate a recorded file | `POST /v1/audio/transcriptions` or `POST /v1/audio/translations` | | Generate speech from text as a standalone task | `POST /v1/audio/speech` | | Send audio in a chat message or receive audio in a complete chat response | [`POST /v1/chat/completions`](https://docs.api7.ai/ai-gateway/endpoints/chat-audio.md) | | Exchange audio over an interactive WebSocket session | [`GET /v1/realtime`](https://docs.api7.ai/ai-gateway/endpoints/realtime.md) | The rest of this guide covers the standalone `/v1/audio/*` routes for transcription, translation, and speech generation. AISIX resolves the caller-facing model alias, applies access checks and supported text guardrails, and forwards the request to an upstream that supports the same audio route. It preserves the endpoint-specific request and response shapes instead of converting them into a chat-style format. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias backed by a provider and model that support the audio route you want to call. Export the gateway connection and caller API key used by the examples: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" ``` ## Send a Transcription Request[​](#send-a-transcription-request "Direct link to Send a Transcription Request") Transcription is a `multipart/form-data` upload, not a JSON body. Send the audio file in the `file` field and the AISIX model alias in the `model` field: ``` curl -sS -X POST "${AISIX_PROXY}/v1/audio/transcriptions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -F "file=@meeting.wav" \ -F "model=transcribe-prod" ``` For a successful request, the upstream returns the transcript: ``` { "text": "The quick brown fox jumps over the lazy dog." } ``` AISIX rebuilds the multipart form with the upstream model ID and preserves the remaining fields. A configured input guardrail can block or mask a text-bearing `prompt` field before AISIX sends the form upstream. To translate non-English speech into English text instead, send the same form to `/v1/audio/translations`. ### Choose a Response Format[​](#choose-a-response-format "Direct link to Choose a Response Format") The optional `response_format` field selects the transcript representation. For a successful request, AISIX preserves the upstream response body and content type unless a configured output guardrail blocks or masks the transcript. The requested representation therefore remains intact when no guardrail changes it: | `response_format` | Response content type | Body | | ----------------- | --------------------- | --------------------------------------------------------------- | | `json` (default) | `application/json` | `{"text": "..."}` | | `verbose_json` | `application/json` | Transcript plus `duration`, `language`, and per-segment timings | | `text` | `text/plain` | The transcript alone | | `srt` | `text/plain` | SubRip subtitles with cue timings | | `vtt` | `text/plain` | WebVTT subtitles with cue timings | Handle transcription responses by content type rather than assuming JSON. Only `json` and `verbose_json` produce a JSON body: ``` curl -sS -X POST "${AISIX_PROXY}/v1/audio/transcriptions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -F "file=@meeting.wav" \ -F "model=transcribe-prod" \ -F "response_format=srt" ``` Format support is a property of the upstream model, not of AISIX. If a model rejects a format, inspect the error returned by AISIX and the provider's model documentation. Do not rely on the error body being a byte-for-byte copy of the upstream response. ### Transcription Streaming[​](#transcription-streaming "Direct link to Transcription Streaming") Some transcription models accept `stream=true` and return server-sent events. AISIX relays those events as the upstream produces them, so a client receives the transcript incrementally instead of waiting for the whole request to finish. The events pass through unchanged, and AISIX reads usage from the terminal event as they do. An output guardrail that can block or mask a transcript is the exception. Such a guardrail has to inspect the whole transcript before any of it reaches the caller. AISIX therefore holds the response back, scans it, and then releases or blocks it — the same protection a non-streamed request gets. A guardrail in monitor mode never blocks, so it does not hold the response back; it observes the transcript once the stream ends. Streaming support is model-specific. For example, OpenAI's [file-transcription guide](https://developers.openai.com/api/docs/guides/speech-to-text#streaming-transcriptions) uses `gpt-transcribe` models for streaming, while the [official SDK specification](https://github.com/openai/openai-python/blob/main/src/openai/types/audio/transcription_create_params.py) states that `whisper-1` ignores `stream`. Check the upstream model documentation before relying on streamed transcription events. ## Send a Speech Request[​](#send-a-speech-request "Direct link to Send a Speech Request") Send a speech-generation request through the gateway proxy with the AISIX model alias in the request body: ``` curl -sS -X POST "${AISIX_PROXY}/v1/audio/speech" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "tts-prod", "input": "Hello from AISIX.", "voice": "alloy" }' \ --output aisix-speech.mp3 ``` For a successful request, the output file should contain the audio bytes returned by the upstream provider. Handle the response as a binary file, not as a chat-style JSON response. AISIX forwards the audio as the provider produces it, so a client that plays the response can start on the first bytes instead of waiting for the whole file. Check that the file was written as audio output: ``` file aisix-speech.mp3 ``` You should see output that identifies the file as audio. The exact wording depends on the operating system and the upstream response format: ``` aisix-speech.mp3: MPEG ADTS, layer III, v2, 160 kbps, 24 kHz, Monaural ``` ## Audio Endpoint Behavior[​](#audio-endpoint-behavior "Direct link to Audio Endpoint Behavior") Audio endpoints do not all use the same request or response shape: | Endpoint | Request body | Response body | | -------------- | ------------------------------------------------- | ----------------------------------------------------------------- | | Transcriptions | Multipart form with an audio file and model alias | Upstream transcription result, in the requested `response_format` | | Translations | Multipart form with an audio file and model alias | Upstream translation result, in the requested `response_format` | | Speech | JSON body with model alias, text input, and voice | Binary audio bytes | For transcription and translation requests, AISIX rebuilds the multipart form with the upstream model ID before forwarding it. It preserves the other form fields, including the uploaded file name and content type when they are present. For speech requests, AISIX rewrites the model field in the JSON body and forwards the remaining request fields to the upstream provider. For successful requests, the gateway preserves the upstream response body and content type unless a transcript guardrail changes or blocks the output. Clients should handle transcription and translation according to the requested `response_format`, and speech as binary audio output. ## Provider Support[​](#provider-support "Direct link to Provider Support") Audio support depends on the resolved provider and model. AISIX does not translate audio formats across provider families. Use these routes with upstreams that expose matching OpenAI-style audio endpoints. If the upstream does not support the requested audio route, the failure is a provider capability or base-URL issue, not a caller-authentication issue. ## Usage and Guardrail Behavior[​](#usage-and-guardrail-behavior "Direct link to Usage and Guardrail Behavior") Successful audio requests are attributed in gateway usage events. Token counts are populated only when the upstream response includes recognized token usage. Transcription and translation requests can also report the audio length. AISIX reads the duration from a supported upstream response or, when needed, measures the uploaded file. This allows AISIX Cloud to price models billed by duration rather than by tokens, including requests that use a non-JSON `response_format`. Set the rate in [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md). Speech requests report no audio duration, so AISIX duration pricing does not apply to them. Input guardrails can inspect, block, or mask the `input` text on speech requests and the optional `prompt` field on transcription and translation requests before AISIX calls the provider. Output guardrails can inspect, block, or mask transcript text. If an output guardrail blocks a transcript, AISIX still records the billed usage because the upstream has already processed the audio. Uploaded audio bytes and generated speech bytes are not scanned as text. ## Troubleshoot Audio Requests[​](#troubleshoot-audio-requests "Direct link to Troubleshoot Audio Requests") If a speech request succeeds but the client expects JSON, adjust the response handling. Speech returns audio bytes. If a transcription or translation request returns 400 from AISIX or the upstream, check the multipart form construction. The request must include a model field and the expected audio file field. If a speech guardrail does not block a request, check the request text. Speech guardrails inspect the input text, not the generated audio bytes. If a requested transcription format fails, confirm that the resolved provider and model support it. If streamed transcription events arrive only after the request finishes, check whether the model is attached to an output guardrail that blocks or masks transcripts. Such a guardrail holds the response back by design. Also confirm that the upstream model supports `stream=true`. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen how AISIX forwards OpenAI-style audio requests and where audio response handling differs from JSON proxy routes. Continue with [Audio Input and Output with Chat Completions](https://docs.api7.ai/ai-gateway/endpoints/chat-audio.md) for audio inside a chat turn or [Realtime API](https://docs.api7.ai/ai-gateway/endpoints/realtime.md) for an interactive session. Use [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for a provider-native route that AISIX does not model directly. --- # Batch, Files, and Fine-Tuning AISIX AI Gateway exposes the OpenAI-compatible Files, Batch, and Fine-tuning APIs as first-class proxy routes. Clients keep the standard OpenAI SDK calls while the gateway manages caller authentication, provider credentials, and usage attribution for completed batches. In this guide, you will upload a batch input file, create and track a batch, and see how the gateway routes each call to the right provider. ## Supported Proxy Routes[​](#supported-proxy-routes "Direct link to Supported Proxy Routes") | Family | Routes | | ----------- | --------------------------------------------------------------------------------------------------------------------------------- | | Files | `POST /v1/files`, `GET /v1/files`, `GET /v1/files/{id}`, `DELETE /v1/files/{id}`, `GET /v1/files/{id}/content` | | Batch | `POST /v1/batches`, `GET /v1/batches`, `GET /v1/batches/{id}`, `POST /v1/batches/{id}/cancel` | | Fine-tuning | `POST /v1/fine_tuning/jobs`, `GET /v1/fine_tuning/jobs`, `GET /v1/fine_tuning/jobs/{id}`, `POST /v1/fine_tuning/jobs/{id}/cancel` | Provider support covers OpenAI-compatible providers (adapter `openai`, including custom `api_base` deployments) and Azure OpenAI (adapter `azure-openai`, resource-scoped routes with `api-key` authentication). Vertex AI, Bedrock, and Anthropic-native batch flows use different wire and storage models and are not served by these routes yet. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias backed by an OpenAI-compatible or Azure OpenAI provider. Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` ## How AISIX Routes Files and Jobs[​](#how-aisix-routes-files-and-jobs "Direct link to How AISIX Routes Files and Jobs") A batch create request references a previously uploaded file id. There is no `model` field in the body for the gateway to route on, so AISIX uses gateway-encoded resource ids: 1. When you upload a file, name the routing model once: a `model` multipart field, a `?model=` query parameter, or an `x-aisix-model` header. 2. The file id returned by the gateway (`aisix-…`) embeds that model. Any later call that references the id routes automatically, including batch create, file retrieve or download, and fine-tuning `training_file`. 3. Ids returned by create and retrieve responses (batch ids, output file ids, fine-tuning job ids) are encoded the same way, so follow-up calls need no extra hints. Raw provider ids keep working: the gateway falls back to an explicit `model` query parameter or header, and then to the first OpenAI-compatible model the caller key can access. Pass an explicit model for deterministic routing. ## Upload a File and Run a Batch[​](#upload-a-file-and-run-a-batch "Direct link to Upload a File and Run a Batch") Upload the batch input file through the gateway, naming the routing model on the request: ``` curl -sS -X POST "${AISIX_PROXY}/v1/files" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "x-aisix-model: ${AISIX_MODEL}" \ -F purpose=batch \ -F file=@batch-input.jsonl ``` The response `id` starts with `aisix-` and embeds the routing model. Create the batch with that id: ``` curl -sS -X POST "${AISIX_PROXY}/v1/batches" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "input_file_id": "aisix-…", "endpoint": "/v1/chat/completions", "completion_window": "24h" }' ``` Track the batch with the batch ID returned by the create request: ``` curl -sS "${AISIX_PROXY}/v1/batches/aisix-…" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` When the batch completes, download the result file with the `output_file_id` from the batch response: ``` curl -sS "${AISIX_PROXY}/v1/files/aisix-…/content" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` The official OpenAI SDKs work unchanged. Point `base_url` at the gateway and pass the routing hint as a header on `files.create`. ## Fine-Tuning Jobs[​](#fine-tuning-jobs "Direct link to Fine-Tuning Jobs") Fine-tuning jobs route through the encoded `training_file` id. The `model` field in the job body is the provider's base model to fine-tune and is forwarded verbatim; it is not a gateway model alias. The returned job object also names a provider base model rather than a gateway alias: the gateway does not translate this field in either direction, so the job object reports whatever the provider records for the model you submitted: ``` curl -sS -X POST "${AISIX_PROXY}/v1/fine_tuning/jobs" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-4o-mini-2024-07-18", "training_file": "aisix-…" }' ``` ## Usage and Cost Attribution[​](#usage-and-cost-attribution "Direct link to Usage and Cost Attribution") File and job management calls, such as upload, create, list, and cancel, record zero-token usage events for visibility in logs. When a batch retrieve first observes `status: "completed"`, the gateway downloads the batch output file and aggregates per-line token usage by provider-billed model. AISIX then emits usage events with real token counts. These events carry deterministic request IDs for each batch and model slice, so repeated attribution after a gateway restart retains the same identities. ## Gateway Policy Behavior[​](#gateway-policy-behavior "Direct link to Gateway Policy Behavior") Caller API key authentication and model access lists, per-model client IP restrictions, and rate limits apply to every route in this family. Matching AISIX Cloud budgets also apply. Guardrail coverage differs. Input and output scans run on `/v1/batches` and `/v1/fine_tuning/jobs`, over the JSON body the caller submits and the JSON the provider returns, not over the file a job references. The five Files routes run no guardrail check in either direction. Uploaded and downloaded files are relayed unscreened, and this route family never screens the JSONL records that a batch job processes. See [The Files Routes Are Not Screened](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md#the-files-routes-are-not-screened). Upstream calls honor the model's request timeout. When cooldown and its transport-error trigger are enabled for the model, connection, request-timeout, and response-read failures can take it out of rotation. Upstream HTTP status responses on these routes do not trigger cooldown. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now run file uploads, batches, and fine-tuning jobs through the gateway. Next, review [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for provider-native endpoints outside this surface. --- # Audio Input and Output with Chat Completions Audio-capable chat models can accept recorded audio inside a message and return generated audio in the same Chat Completions response. This fits turn-based applications that need a model to understand or answer with audio while keeping the OpenAI-compatible `POST /v1/chat/completions` request shape. For standalone transcription, translation, or speech generation, use [Speech and Audio](https://docs.api7.ai/ai-gateway/endpoints/audio.md). For interactive, bidirectional sessions, use the [Realtime API](https://docs.api7.ai/ai-gateway/endpoints/realtime.md). AISIX preserves the OpenAI chat-audio request and response fields when the selected upstream uses a compatible OpenAI-shaped API. It does not translate these fields into another provider's native audio protocol. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias backed by an audio-capable Chat Completions model. Its provider key must use the `openai` or `azure-openai` adapter, and the upstream must implement the OpenAI chat-audio request and response shape. * A local WAV recording named `question.wav` for the audio-input example. * `curl`, `jq`, Python 3, and the `base64`, `file`, and `tr` command-line utilities for the examples. If you have not configured the upstream yet, follow [OpenAI](https://docs.api7.ai/ai-gateway/providers/openai.md), [Azure OpenAI](https://docs.api7.ai/ai-gateway/providers/azure-openai.md), or [Bring Your Own Endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md). Choose a currently supported upstream audio model and create an AISIX alias for it. For a routing alias, every eligible target must use one of these adapters and support the same chat-audio fields. A fallback to a text-only or differently shaped provider can lose the audio content or fail upstream. Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="audio-chat-prod" ``` ## Generate an Audio Response[​](#generate-an-audio-response "Direct link to Generate an Audio Response") Request both text and audio output, then save the complete response: ``` curl -sS -X POST "${AISIX_PROXY}/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "modalities": ["text", "audio"], "audio": { "voice": "alloy", "format": "wav" }, "messages": [ { "role": "user", "content": "Say: Your AISIX audio route is working." } ] }' > chat-audio-response.json ``` AISIX replaces the alias with the upstream model ID before dispatch. The upstream chooses which voices and output formats it accepts. Inspect the returned audio metadata without printing the base64 payload: ``` jq '{ model, audio: (.choices[0].message.audio | { id, transcript, expires_at, encoded_characters: (.data | length) }) }' chat-audio-response.json ``` The response keeps `model` set to the AISIX alias. The `audio` object contains the provider's audio identifier, base64 data, transcript, and expiration timestamp when the upstream returns those fields. Decode the generated WAV file with the Python standard library: ``` python3 - <<'PY' import base64 import json with open("chat-audio-response.json", encoding="utf-8") as response_file: response = json.load(response_file) audio = response["choices"][0]["message"]["audio"] with open("aisix-chat-audio.wav", "wb") as audio_file: audio_file.write(base64.b64decode(audio["data"])) print(audio.get("transcript", "")) PY file aisix-chat-audio.wav ``` The final command should identify a WAV file. If you request another format, use a matching filename and media player. ## Send Audio in a Message[​](#send-audio-in-a-message "Direct link to Send Audio in a Message") Encode a local recording and place it in an `input_audio` content block. This example asks the model to answer with audio so you can use the same response handling as the previous request: ``` base64 < question.wav | tr -d '\n' | \ jq -Rs \ --arg model "$AISIX_MODEL" \ '{ model: $model, modalities: ["text", "audio"], audio: { voice: "alloy", format: "wav" }, messages: [ { role: "user", content: [ { type: "text", text: "Answer the question in this recording." }, { type: "input_audio", input_audio: { data: ., format: "wav" } } ] } ] }' > chat-audio-input.json curl -sS -X POST "${AISIX_PROXY}/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ --data @chat-audio-input.json \ > chat-audio-response.json ``` AISIX forwards the typed content blocks unchanged through an OpenAI-compatible adapter. It does not decode the recording or convert it to another provider's audio-input shape. ## Understand Current Behavior[​](#understand-current-behavior "Direct link to Understand Current Behavior") Chat audio uses the ordinary Chat Completions authentication, model access, routing, retry, and request telemetry path. The audio-specific behavior has these boundaries: * **Use a non-streaming request.** Omit `stream` or set it to `false`. AISIX preserves `message.audio` on a complete response, but it does not currently return streamed `delta.audio` chunks. * **Keep the provider protocol compatible.** The `openai` and `azure-openai` adapters preserve the audio request fields and the non-streaming `message.audio` object. Other adapters can reduce typed message content to text and do not translate the audio fields into a provider-native voice API. * **Apply guardrails to text separately from audio.** Input guardrails can inspect the text content blocks, and output guardrails can inspect ordinary returned message text. They do not inspect the bytes in `input_audio`, the generated audio data, or the transcript nested inside `message.audio`. * **Treat Cloud audio cost as an upstream detail.** AISIX records the normalized prompt and completion totals the upstream reports, but it does not retain separate audio-token counts. AISIX Cloud pricing therefore cannot apply distinct text-token and audio-token rates to the same Chat Completions request. Because audio is base64-encoded inside JSON, request and response bodies are larger than the underlying binary files. Account for that expansion when setting client, proxy, or load-balancer body limits. ## Troubleshoot Chat Audio[​](#troubleshoot-chat-audio "Direct link to Troubleshoot Chat Audio") If the response succeeds but has no `message.audio`, check that the request includes the audio modality and an `audio` object, and that the upstream model supports audio output through Chat Completions. A text-only model can accept the HTTP request but reject or ignore unsupported audio fields upstream. If AISIX returns an upstream decode or provider error, call the same upstream model directly with the provider's documented request shape. Confirm the model ID, voice, format, and audio-input encoding before changing gateway policy. If the request works without streaming but produces no audio when `stream` is enabled, keep the request non-streaming. Use the [Realtime API](https://docs.api7.ai/ai-gateway/endpoints/realtime.md) when the application needs incremental, bidirectional audio. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now sent and received audio inside an OpenAI-compatible chat request. Continue with [Speech and Audio](https://docs.api7.ai/ai-gateway/endpoints/audio.md) for standalone transcription or speech synthesis, or [Realtime API](https://docs.api7.ai/ai-gateway/endpoints/realtime.md) for interactive voice sessions. --- # Embeddings Embeddings convert text into vectors that applications can use for semantic search, retrieval pipelines, clustering, and similarity checks. AISIX AI Gateway lets embedding clients keep the OpenAI-compatible request and response format while the gateway manages caller authentication, model aliases, upstream credentials, and policy. In this guide, you will send an embeddings request through AISIX and review the provider behavior that matters for this endpoint. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias backed by a provider and model that support embeddings. Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="text-embedding-prod" ``` ## Send an Embeddings Request[​](#send-an-embeddings-request "Direct link to Send an Embeddings Request") Send the embeddings request through the gateway proxy with the AISIX model alias in the request body: ``` curl -sS -X POST "${AISIX_PROXY}/v1/embeddings" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "input": [ "hello", "world" ] }' ``` AISIX resolves the model alias, checks the caller API key, rewrites the upstream model ID, and forwards the embeddings request to the provider. The response keeps the OpenAI-compatible embeddings format: ``` { "object": "list", "data": [ { "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, 0.0789] }, { "object": "embedding", "index": 1, "embedding": [0.0234, -0.0567, 0.0891] } ], "model": "text-embedding-prod", "usage": { "prompt_tokens": 2, "total_tokens": 2 } } ``` The embedding value may be a float array or a base64 string, depending on the requested encoding format and upstream response. ## Provider and Request Behavior[​](#provider-and-request-behavior "Direct link to Provider and Request Behavior") The embeddings route accepts the OpenAI-compatible request shape and translates it to the provider's native embeddings API where needed: | Provider family | Upstream API | Notes | | ------------------------------------- | ---------------------------- | --------------------------------------------------------------------------------------------------------------------------- | | OpenAI-compatible (adapter `openai`) | `POST {api_base}/embeddings` | Forwarded in OpenAI shape, including custom `api_base` deployments. | | Vertex AI / Gemini (adapter `vertex`) | google-publisher `:predict` | `dimensions` maps to `outputDimensionality`; per-input token counts sum into usage. | | Bedrock (adapter `bedrock`) | `InvokeModel` | `amazon.titan-embed-*` (one call per input; `dimensions` on V2 models only) and `cohere.embed-*` (whole batch in one call). | | Anthropic | No embeddings API | Requests return 501. | AISIX accepts embedding input as a single string or an array of strings. It preserves the caller's input shape when it forwards the request upstream. Callers do not need separate client-side logic to switch between a single input and a batch input. Input guardrails can inspect text in those supported input forms before AISIX calls the provider. Token-array inputs are not currently supported by this gateway route. The gateway records usage when the upstream returns token usage. Embeddings do not use completion tokens, response caching, streaming, or output guardrails on this proxy path. ## Troubleshoot Embeddings[​](#troubleshoot-embeddings "Direct link to Troubleshoot Embeddings") If AISIX returns 501, the resolved provider has no embeddings API (for example, Anthropic). Use a model backed by a provider family from the table above. For batch requests, the response should return one embedding entry per input item. If fewer vectors come back, inspect the upstream response and gateway logs for provider-specific batch handling. If an input guardrail does not block the request, check whether the input contains inspectable text. AISIX can scan a single string or an array of strings. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen how AISIX proxies embeddings requests and where provider support can differ. Next, continue with [Rerank](https://docs.api7.ai/ai-gateway/endpoints/rerank.md) when your application needs ranking behavior for retrieved documents. --- # Image Editing Image editing lets applications send image-plus-prompt editing requests through AISIX while keeping caller authentication, model aliases, upstream credentials, and request-side policy in one gateway path. Editing models such as `gpt-image-2` take the source image or images, an optional mask, the prompt, and every tuning parameter in one `multipart/form-data` body. AISIX reads the form, resolves the caller-facing model alias, rewrites the `model` field to the upstream model ID, rebuilds the form with every other part byte-for-byte intact — aside from `prompt` text a mask-action guardrail rule rewrites, described below — and forwards it to the upstream image-editing endpoint. In this guide, you will edit an image through AISIX and review the request shape and provider requirement for this endpoint. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias whose configured provider is OpenAI and whose upstream `model_name` is an image-editing model such as `gpt-image-2`. The examples use the alias `image-edit-prod`. * A source image file to edit. The example uses `original.png`. Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="image-edit-prod" ``` ## Send an Image Edit Request[​](#send-an-image-edit-request "Direct link to Send an Image Edit Request") Send the edit request through the gateway proxy as a multipart form with the AISIX model alias in the `model` field: ``` curl -sS -X POST "${AISIX_PROXY}/v1/images/edits" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -F "model=${AISIX_MODEL}" \ -F "image=@original.png" \ -F "prompt=Add a red hat to the subject" \ -F "size=1024x1024" \ -o aisix-image-edit-response.json ``` AISIX resolves the model alias, checks the caller API key, runs supported input policy checks on the prompt, rewrites the `model` form field to the upstream model ID, and forwards the rebuilt form — the image bytes, filenames, and every other field unchanged unless a mask-action guardrail rule rewrote the prompt — to the upstream image-editing endpoint. The response keeps the OpenAI image format. Editing models return base64 image data and a token usage block: ``` { "created": 1710000000, "data": [ { "b64_json": "..." } ], "usage": { "input_tokens": 50, "output_tokens": 1056, "total_tokens": 1106 } } ``` Check that the response includes one image item: ``` jq '.data | length' aisix-image-edit-response.json ``` The command should print: ``` 1 ``` ## Request Fields[​](#request-fields "Direct link to Request Fields") The route accepts `multipart/form-data` only. A JSON body returns `400` in the gateway's error envelope. | Field | Required by | Meaning | | -------- | ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `model` | Gateway | The AISIX model alias. The only field AISIX always rewrites; a configured mask-action guardrail rule can also rewrite `prompt`. | | `image` | Provider | The source image file. The field can repeat for models that accept multiple input images; AISIX forwards every part in order with its bytes and filename intact. | | `prompt` | Provider | The edit instruction. Input guardrails inspect and can mask this text before the request goes upstream. | | `mask` | — | An optional mask image whose transparent areas mark the region to edit. Forwarded intact. | Every other form field — `n`, `size`, `quality`, `background`, `input_fidelity`, and any parameter the provider adds later — is forwarded verbatim, so new upstream parameters do not require a gateway upgrade. Unset fields are simply absent from the upstream request. note `stream=true` is not supported on this route. Editing models can stream partial images as server-sent events, but AISIX does not yet relay that stream and rejects the request with `400` instead of silently buffering it. ## OpenAI Provider Requirement[​](#openai-provider-requirement "Direct link to OpenAI Provider Requirement") The image-editing route is provider-specific. AISIX accepts it only when the resolved model is configured with the OpenAI provider. This is stricter than using the OpenAI-compatible adapter. An OpenAI-compatible vendor can work on the chat-completions route and still be rejected on the image-editing route because its configured provider is not OpenAI. When the resolved model is not configured with the OpenAI provider, AISIX returns `400` before sending anything upstream. The upstream URL derives from the provider key the same way as the other OpenAI routes: an unset `api_base` resolves to the standard OpenAI API, and a bare-host `api_base` gets `/v1` appended before the endpoint path. ## Image-Editing Behavior[​](#image-editing-behavior "Direct link to Image-Editing Behavior") Input guardrails inspect every `prompt` form field before AISIX calls the provider, and mask-action rules rewrite the prompt text in place. A blocked prompt returns `422` before any upstream call and does not consume the model's rate-limit capacity. Image and mask bytes are not scannable text, and generated image bytes are not scanned by output guardrails. Submissions count against both the caller API key layers and the model's rate limits. When the upstream response includes a token usage block — `gpt-image` models return one — AISIX records those tokens and counts them toward token-based limits; a response without one records zero tokens. Per-image cost details such as image count, size, and quality are not inferred by this proxy path. ## Errors[​](#errors "Direct link to Errors") Failures return the gateway's JSON error envelope. A `4xx` from the provider — for example, a `size` value the model rejects — is relayed with the provider's own status and message and is not listed below; the table covers the statuses AISIX generates itself. | Status | When | | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------ | | `400` | The body is not valid `multipart/form-data`, the `model` field is missing, `stream=true` is set, or the resolved model's provider is not OpenAI. | | `401` | The caller API key is missing or invalid. | | `403` | The caller API key is not allowed to use the model alias, or the request comes from a client IP outside the model's allowlist. | | `404` | The model alias does not resolve. | | `413` | The request body exceeds the configured request-body limit. | | `422` | An input guardrail blocked the prompt. No provider request is sent. | | `429` | A rate limit rejected the request, or an AISIX Cloud budget rejected it. | | `502` | The provider returned a server error or a response that is not valid JSON, or was unreachable. | | `504` | The provider did not answer within the model's request timeout. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen how AISIX proxies OpenAI image-editing requests and how the multipart form travels through the gateway. Continue with [Image Generation](https://docs.api7.ai/ai-gateway/endpoints/image-generation.md) for the prompt-to-image route, or [Speech and Audio](https://docs.api7.ai/ai-gateway/endpoints/audio.md) for the other multipart surfaces. --- # Image Generation Image generation lets applications send prompt-to-image requests through AISIX while keeping caller authentication, model aliases, upstream credentials, and request-side policy in one gateway path. AISIX exposes the OpenAI image-generation route for model aliases whose configured provider is OpenAI. It resolves the caller-facing model alias, rewrites only the model field to the upstream model ID, and returns the provider's JSON image response. In this guide, you will send an image-generation request through AISIX and review the provider requirement for this endpoint. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias whose configured provider is OpenAI and whose adapter supports image generation. Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="image-prod" ``` ## Send an Image Request[​](#send-an-image-request "Direct link to Send an Image Request") Send the image-generation request through the gateway proxy with the AISIX model alias in the request body: ``` curl -sS -X POST "${AISIX_PROXY}/v1/images/generations" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "prompt": "A minimal illustration of an AI gateway" }' \ -o aisix-image-response.json ``` AISIX resolves the model alias, checks the caller API key, runs supported input policy checks on the prompt, rewrites only the model field to the upstream model ID, and forwards the request to the upstream image-generation endpoint. The response keeps the OpenAI image-generation format: ``` { "created": 1710000000, "data": [ { "url": "https://example.com/generated-image.png" } ] } ``` Some OpenAI image models return base64 image data instead of a URL, depending on the request and upstream model behavior. Check that the response includes one image item: ``` jq '.data | length' aisix-image-response.json ``` The command should print: ``` 1 ``` note `stream: true` is not supported on this route. Image models can stream partial images as server-sent events, but AISIX does not yet relay that stream and rejects the request with `400` before contacting the provider, instead of silently buffering it: ``` { "error": { "message": "request payload is invalid: `stream` is not supported on /v1/images/generations", "type": "invalid_request_error" } } ``` Requests that set `stream: false`, or that omit `stream`, are unaffected and are forwarded as before. ## OpenAI Provider Requirement[​](#openai-provider-requirement "Direct link to OpenAI Provider Requirement") The image-generation route is provider-specific. AISIX accepts it only when the resolved model is configured with the OpenAI provider. This is stricter than using the OpenAI-compatible adapter. An OpenAI-compatible vendor can work on the chat-completions route and still be rejected on the image-generation route because its configured provider is not OpenAI. When the resolved model is not configured with the OpenAI provider, AISIX returns 400 before sending the request upstream. If the resolved OpenAI-provider bridge does not implement image generation, AISIX returns 501. This is a provider capability issue, not a caller-authentication issue. ## Image-Generation Behavior[​](#image-generation-behavior "Direct link to Image-Generation Behavior") Input guardrails can inspect the prompt before AISIX calls the provider. Output guardrails do not scan generated image bytes. When the upstream image response includes recognized token usage, AISIX records it. Some image models do not return token usage. Those successful requests remain visible with zero token counts, but per-image cost details such as image count, size, and quality are not inferred by this proxy path. If a guardrail does not block a request, check whether the prompt text contains content the configured guardrail can inspect. Generated image bytes are not inspected by output guardrails. If a request returns 400, check whether the request set `stream: true`, and whether the resolved model is configured with the OpenAI provider. If a request returns 501, check whether the resolved OpenAI-provider bridge supports image generation. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen how AISIX proxies OpenAI image-generation requests and where provider support is intentionally narrow. Next, continue with [Image Editing](https://docs.api7.ai/ai-gateway/endpoints/image-editing.md) to edit images through the same gateway path, or [Speech and Audio](https://docs.api7.ai/ai-gateway/endpoints/audio.md) to review audio request behavior through AISIX. --- # OpenAI Client with Anthropic Upstream AISIX lets an application keep the OpenAI Chat Completions request shape while the gateway calls an Anthropic upstream model. Use this pattern when application code is already built around an OpenAI-compatible SDK, but the platform team wants to route that traffic to Claude. AISIX resolves the model alias, translates the request to Anthropic Messages, calls Anthropic with the stored provider credential, and translates the response back into an OpenAI-compatible chat completion. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that your application can reach. * An Anthropic-backed model alias that accepts OpenAI-compatible Chat Completions requests. * A caller API key allowed to use that model alias. * Node.js 20 LTS or newer with `npm` for the SDK example, or `curl` for the HTTP example. If you have not configured the model alias and caller API key, follow [Anthropic](https://docs.api7.ai/ai-gateway/providers/anthropic.md) for either AISIX Cloud or the open-source AISIX gateway. ## Request Flow[​](#request-flow "Direct link to Request Flow") The application keeps the OpenAI-compatible client contract. Provider selection and protocol translation stay in the gateway. The application sends the model alias and caller API key to AISIX. The gateway resolves the upstream model, supplies the stored Anthropic credential, and translates both sides of the exchange. The application continues to send and receive OpenAI-compatible data. ## Call the Alias[​](#call-the-alias "Direct link to Call the Alias") Export the caller API key and model alias used by both request examples: ``` export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="claude-sonnet-prod" ``` ### OpenAI SDK[​](#openai-sdk "Direct link to OpenAI SDK") Install the OpenAI SDK: ``` npm install openai ``` Set the OpenAI-compatible base URL. The OpenAI SDK requires the `/v1` path: ``` # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" ``` Create a minimal Chat Completions client: anthropic-via-openai-sdk.mjs ``` import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.AISIX_API_KEY, baseURL: process.env.AISIX_BASE_URL, }); const completion = await client.chat.completions.create({ model: process.env.AISIX_MODEL, messages: [{ role: "user", content: "Say hello from AISIX." }], }); console.log(completion.choices[0]?.message.content); console.log(completion.usage); ``` Run the example from the shell where the AISIX values are set: ``` node anthropic-via-openai-sdk.mjs ``` ### HTTP[​](#http "Direct link to HTTP") To inspect the response without an SDK, export the gateway origin and send the same request with curl: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send the request: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "'"$AISIX_MODEL"'", "messages": [{"role":"user","content":"Say hello from AISIX."}] }' ``` Both examples return the OpenAI-compatible Chat Completions shape. The caller does not receive Anthropic-shaped content blocks: ``` { "object": "chat.completion", "model": "claude-sonnet-prod", "choices": [ { "message": { "role": "assistant", "content": "Hello from AISIX." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 9, "completion_tokens": 5, "total_tokens": 14 } } ``` ## Translation Behavior[​](#translation-behavior "Direct link to Translation Behavior") The translation preserves the parts of the OpenAI-compatible contract that common chat applications depend on: | Behavior | What AISIX does | | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Model and authentication | Resolves the model alias to the configured Anthropic model, authenticates the caller, and uses the stored provider credential for the upstream request. | | Messages and tools | Maps leading system messages, user and assistant messages, function tools, tool calls, and tool results to Anthropic Messages structures. | | Responses | Converts Anthropic text and tool-use blocks, stop reasons, token usage, and streaming events back into OpenAI-compatible fields. | | Output limit | Supplies `max_tokens: 4096` when the OpenAI-compatible request omits an output limit, because Anthropic requires one. | | Reasoning effort | Sends `reasoning_effort` as `output_config.effort`, Anthropic's current depth control. `minimal` maps to `low`, Anthropic's floor, and `none` becomes `thinking: {"type": "disabled"}` instead, because Anthropic has no `none` tier. AISIX adds no `thinking` block in the other cases: the request asked for a depth, not a thinking mode, and current Anthropic models apply their own. A `reasoning_effort` value outside that set is dropped entirely: it corresponds to no known Anthropic tier, and the original field cannot be forwarded in its place either, because `/v1/messages` rejects unknown top-level fields. An `output_config` or `thinking` the request supplies itself is left as written and takes precedence. | caution When an OpenAI Chat Completions message uses typed content parts, the Anthropic translation preserves text parts but drops non-text parts such as images and audio. Use the Anthropic-style `/v1/messages` route when image or document content must reach an Anthropic upstream. For a complete tool loop, see [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md). AISIX can also add Anthropic prompt-cache markers to eligible Chat Completions requests; see [Anthropic Prompt Caching](https://docs.api7.ai/ai-gateway/traffic-controls/prompt-caching.md). Use the Anthropic-style `/v1/messages` route when the application must keep the Anthropic request and response shape, especially for provider-specific content or thinking blocks. See [Anthropic-Style Messages API](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md) for the native client contract and its compatibility boundaries. ### Token Usage[​](#token-usage "Direct link to Token Usage") Anthropic and OpenAI count prompt-cache tokens differently, so AISIX converts the counts instead of passing them through. Anthropic reports `input_tokens` as the **non-cached** input, with `cache_creation_input_tokens` and `cache_read_input_tokens` as separate counters beside it. OpenAI accounting has one `prompt_tokens` that **already includes** the cached part, named under `prompt_tokens_details.cached_tokens`. OpenAI has no cache-write concept at all, so AISIX folds the write into `prompt_tokens` — it is billed input — and reports it beside the hit. An upstream response reporting: ``` { "usage": { "input_tokens": 40, "output_tokens": 10, "cache_creation_input_tokens": 30, "cache_read_input_tokens": 70 } } ``` reaches an OpenAI-compatible caller as: ``` { "usage": { "prompt_tokens": 140, "completion_tokens": 10, "total_tokens": 150, "prompt_tokens_details": { "cached_tokens": 70, "cache_creation_tokens": 30 } } } ``` The rules a caller can rely on: * `prompt_tokens` is the full input the model read, cache reads and cache writes included. * `total_tokens` is `prompt_tokens + completion_tokens`. * `cached_tokens` is a subset of `prompt_tokens`, and counts only cache **reads**. * `cache_creation_tokens` is the cache **write**, also a subset of `prompt_tokens`. It is billed input but is not a cache hit, so it is reported separately rather than inside `cached_tokens`. OpenAI has no cache-write concept, so this field appears only when the upstream reported one — which matters because a provider typically bills a write above the plain input rate. On the first turn of a cached conversation, a write with no read, it is the only signal that a cache was involved at all. The same conversion applies to streaming responses and to the [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over an Anthropic upstream. Logs, metrics, and spend reporting are not converted: they keep Anthropic's own counters, so a call costs the same whichever protocol addressed it. See [Anthropic Prompt Caching](https://docs.api7.ai/ai-gateway/traffic-controls/prompt-caching.md) for the recorded view. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now routed an OpenAI-compatible client to an Anthropic upstream. See [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md) for caller-facing route behavior. Use [Anthropic Messages](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md) when you want the Anthropic request and response shape end to end, or review [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for endpoint and provider support boundaries. --- # OpenAI-Compatible Chat Completions Applications that use the OpenAI Chat Completions format can send the same supported request shape to an AISIX gateway at `POST /v1/chat/completions`. The gateway authenticates the caller, resolves the model alias, applies gateway policy, and dispatches the request through the selected provider adapter. The compatibility is client-facing: the upstream can be OpenAI or another supported provider. The gateway returns a supported OpenAI-compatible response shape while the adapter handles provider translation. This guide describes the Chat Completions proxy behavior. For a runnable SDK setup, see [OpenAI SDK](https://docs.api7.ai/ai-gateway/getting-started/openai-sdk.md). ## What the Client Sends[​](#what-the-client-sends "Direct link to What the Client Sends") The client sends three gateway-facing values: * The base URL is the AISIX proxy API root, which is the gateway origin followed by `/v1`. * The API key is an AISIX caller API key. * The model value is an AISIX model alias, such as `gpt-4o-prod`. The request body keeps the OpenAI-compatible format, including messages, tools, streaming options, and supported multimodal fields. Provider credentials, upstream model IDs, routing policy, rate limits, guardrails, and other gateway policy stay in AISIX. For audio content blocks and generated audio, see [Audio Input and Output with Chat Completions](https://docs.api7.ai/ai-gateway/endpoints/chat-audio.md). Send the caller API key with the standard bearer token format: ``` Authorization: Bearer YOUR_CALLER_API_KEY ``` AISIX also accepts `x-api-key: YOUR_CALLER_API_KEY` for compatibility. Use the bearer token format for OpenAI-compatible clients when the client supports it. Export the gateway connection and request values used below: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` ## Send a Chat Completions Request[​](#send-a-chat-completions-request "Direct link to Send a Chat Completions Request") Use `POST /v1/chat/completions` as the default route for OpenAI-compatible chat clients. Send a chat-completions request through AISIX: ``` curl -sS -X POST "${AISIX_PROXY}/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "messages": [ {"role": "user", "content": "Hello from AISIX."} ] }' ``` A successful response uses the OpenAI-compatible chat-completions format. The response keeps `model` set to the caller-facing alias from the request. ## Discover Available Models[​](#discover-available-models "Direct link to Discover Available Models") `GET /v1/models` returns every concrete model alias visible to the caller API key, including direct, routing, semantic, and ensemble aliases. Wildcard aliases are patterns rather than concrete model names, so they are not listed. A key that allows every model sees every concrete alias, while a restricted key sees only the aliases its allowlist permits. List the model aliases visible to the caller API key: ``` curl -sS "${AISIX_PROXY}/v1/models" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` ## Handle Errors[​](#handle-errors "Direct link to Handle Errors") The Chat Completions route returns errors in an OpenAI-style envelope. Use the error type before the status code when you need to distinguish caller authentication, model access, policy blocks, rate limits, and upstream failures. For the full error and header reference, see [Headers and Error Codes](https://docs.api7.ai/ai-gateway/reference/headers-and-error-codes.md). ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen how OpenAI-compatible clients call AISIX. Continue with [Audio Input and Output with Chat Completions](https://docs.api7.ai/ai-gateway/endpoints/chat-audio.md), [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md), or [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md). Use [OpenAI Client with Anthropic Upstream](https://docs.api7.ai/ai-gateway/endpoints/openai-client-to-anthropic.md) when the application should keep the OpenAI-compatible client shape while AISIX calls Anthropic upstream. --- # Supported Endpoints AISIX exposes proxy endpoints for the request formats application teams already use. Start here when you need to know which client API shape to call before choosing a provider, model alias, or traffic policy. For coding tools and application frameworks that need client-specific setup, see [Integrations](https://docs.api7.ai/ai-gateway/integrations.md). ## Main API Families[​](#main-api-families "Direct link to Main API Families") | API family | Routes | Use for | | ----------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [Chat completions](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md) | `POST /v1/chat/completions` | OpenAI-compatible chat requests, [audio input and output](https://docs.api7.ai/ai-gateway/endpoints/chat-audio.md), streaming text, tool calling, routed models, and ensemble models. | | [Anthropic messages](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md) | `POST /v1/messages`, `POST /v1/messages/count_tokens` | Anthropic-style messages requests and token counting for Anthropic-backed models. | | [Responses](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) | `POST /v1/responses` | Responses API clients and agent-style response flows across supported providers. | | [Text completions](https://docs.api7.ai/ai-gateway/endpoints/text-completions.md) | `POST /v1/completions` | Legacy OpenAI-compatible text completion clients. | | [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md) | `POST /v1/embeddings` | Vector embeddings through supported providers. | | [Rerank](https://docs.api7.ai/ai-gateway/endpoints/rerank.md) | `POST /v1/rerank` | Document reranking through supported providers. | | [Image generation](https://docs.api7.ai/ai-gateway/endpoints/image-generation.md) | `POST /v1/images/generations` | Text-to-image generation requests. | | [Image editing](https://docs.api7.ai/ai-gateway/endpoints/image-editing.md) | `POST /v1/images/edits` | Multipart image-plus-prompt editing requests. | | [Video generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md) | `POST /v1/videos`, `GET /v1/videos/{video_id}`, `GET /v1/videos/{video_id}/content` | Asynchronous text-to-video tasks: submit, poll status, and download the result. | | [Speech and audio](https://docs.api7.ai/ai-gateway/endpoints/audio.md) | `POST /v1/audio/transcriptions`, `POST /v1/audio/translations`, `POST /v1/audio/speech` | Speech-to-text, translation, and text-to-speech requests. | | [Batch, files, and fine-tuning](https://docs.api7.ai/ai-gateway/endpoints/batch-files-fine-tuning.md) | `POST /v1/files`, `POST /v1/batches`, `POST /v1/fine_tuning/jobs`, and their retrieve/list/cancel/content routes | OpenAI-compatible batch processing, file management, and fine-tuning jobs with gateway-managed provider routing. | | [Realtime](https://docs.api7.ai/ai-gateway/endpoints/realtime.md) | `GET /v1/realtime` (WebSocket) | OpenAI Realtime clients relayed over WebSocket with session usage tracking. | | [Passthrough routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) | Configured path prefixes and hosts | Provider-specific calls relayed through explicitly configured routes, with AISIX authentication and policy but without AISIX normalizing the request body. | | [MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/overview.md) | `ANY /mcp`, `ANY /mcp/{server}` | Agent tool calls through all permitted MCP servers or one named server, governed by caller API key tool access. | | [Agent Gateway](https://docs.api7.ai/ai-gateway/agent-gateway/overview.md) | `POST /a2a/{agent}`, `GET /a2a/{agent}/.well-known/agent-card.json` | A2A JSON-RPC calls and caller-authenticated agent-card discovery for registered upstream agents. | ## Discovery and Health[​](#discovery-and-health "Direct link to Discovery and Health") | Endpoint | Use for | | ---------------- | -------------------------------------------------------------------------------------------------------------------------- | | `GET /v1/models` | Return the model aliases the caller API key can access. Use it when a client needs to discover gateway-facing model names. | | `GET /livez` | Check whether the proxy listener is alive. Use it for proxy listener health checks, not for model or provider readiness. | | `GET /readyz` | Check whether the instance should receive traffic. Returns 503 while draining or before the first configuration apply. | ## Gateway Behavior[​](#gateway-behavior "Direct link to Gateway Behavior") Modeled proxy routes share the same core gateway behavior. AISIX authenticates the caller API key, checks model access, resolves the requested model alias, and applies configured controls. AISIX then dispatches to the selected upstream provider and records usage and telemetry when the route can be attributed to a model. Some behavior is route-specific. For example, response caching applies to chat completions when a matching cache policy is configured, and ensemble models are supported on chat completions. MCP tool calls use caller API key tool access, while token counting is limited to Anthropic-backed models. For provider and route constraints, see [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md). For exact request and response details, see the [Proxy API Reference](https://docs.api7.ai/ai-gateway/reference/proxy-api.md). --- # Passthrough Routes A passthrough route forwards matching requests to one upstream target without translating the request body. It is useful for provider-native endpoints that AISIX does not model as first-class routes and for forward-proxy traffic delivered with its original `Host` header. Each route defines how traffic matches, where it goes, how AISIX authenticates the caller, and whether AISIX injects a provider credential or forwards the caller's credential. AISIX can still apply caller access controls, request limits, guardrails, and telemetry around the relay. Explicit route required The former implicit `/passthrough//...` tunnel has been removed. An unclaimed `/passthrough/*` path follows the ordinary empty-body `404` path. Create and verify the replacement route before moving client traffic. Passthrough routes do not rewrite model identifiers. If a provider-native request names a model in its body, path, query, or headers, send the identifier expected by that provider. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Prepare the following: * One AISIX deployment: * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. * For the open-source AISIX gateway, a gateway configured to load a declarative resources file. * An upstream provider credential for an `inject` route. The examples use OpenAI. A `forward_client` route instead relays the caller's upstream credential. * `curl` and `jq`. ## Understand the Passthrough Flow[​](#understand-the-passthrough-flow "Direct link to Understand the Passthrough Flow") AISIX preserves the request body while handling gateway authentication, target construction, intentional header filtering, guardrails, and telemetry: Route matching runs in two stages: 1. A request whose inbound `Host` matches a route's `hosts` allowlist is dispatched before the gateway's typed routes. This lets a forward proxy relay an upstream path such as `/v1/messages` without the gateway treating it as its own endpoint. 2. Path-prefix matching runs after the typed routes, so a path-only route cannot shadow the gateway's `/v1`, `/mcp`, or `/a2a` endpoints. When several routes match, host matches beat path-only matches, and a longer matching prefix beats a shorter one. ## Configure a Provider-Native Route[​](#create-a-passthrough-route "Direct link to Configure a Provider-Native Route") The following examples expose OpenAI's native model-list endpoint at `/passthrough/openai/v1/models`. AISIX injects the configured OpenAI credential and requires a caller key that grants `openai-tunnel`. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Export the AISIX Cloud connection details and the provider credential: ``` # AISIX_CP includes /api and has no trailing slash. export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_BASE_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" export OPENAI_API_KEY="YOUR_OPENAI_API_KEY" ``` Create an OpenAI provider key that is available to the environment: ``` PROVIDER_KEY_ID=$(curl --fail-with-body -sS -X POST \ "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "OpenAI passthrough", "provider": "openai", "api_key": "'"${OPENAI_API_KEY}"'", "api_base": "https://api.openai.com/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id') ``` Create the route with the provider key ID: ``` ROUTE_RESPONSE=$(curl --fail-with-body -sS -X POST \ "$AISIX_CP/environments/$ENV_ID/passthrough_routes" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "openai-tunnel", "path_prefix": "/passthrough/openai", "target_url": "https://api.openai.com/v1", "credential_mode": "inject", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }') export ROUTE_ID=$(printf '%s' "$ROUTE_RESPONSE" | jq -er '.passthrough_route.id') printf '%s' "$ROUTE_RESPONSE" | jq '.warnings // []' ``` Keep `ROUTE_ID` for route updates or route-scoped guardrail attachments. Create a dedicated caller key and grant the route in the same request. The plaintext is returned once: ``` CALLER_RESPONSE=$(curl --fail-with-body -sS -X POST \ "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "OpenAI passthrough caller", "allowed_models": [], "allowed_routes": ["openai-tunnel"] }') export AISIX_API_KEY=$(printf '%s' "$CALLER_RESPONSE" | jq -er '.plaintext') printf '%s' "$CALLER_RESPONSE" | jq '.warnings // []' ``` The control plane projects the provider key, route, and caller grant to attached gateways. Review any returned compatibility warnings before rollout. Warnings are advisory, so verify traffic through each gateway. In the dashboard, the same workflow is available under **Provider keys**, an environment's **Passthrough Routes**, and the caller key's **Passthrough route access** section. If you grant an existing caller instead, include every route grant it should keep. The AISIX Cloud Admin API replaces the complete `allowed_routes` list when that field is patched. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Prepare an `openai-prod` provider key in the [complete resources file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), then choose a dedicated caller credential: ``` export PASSTHROUGH_CALLER_KEY="YOUR_CALLER_API_KEY" ``` Add the route and caller entries to the matching collections. Keep the provider key and other resources unchanged: resources.yaml (route and caller key) ``` passthrough_routes: - name: openai-tunnel path_prefix: /passthrough/openai target_url: https://api.openai.com/v1 provider_key: openai-prod api_keys: - display_name: passthrough-caller key_env: PASSTHROUGH_CALLER_KEY allowed_models: [] allowed_routes: [openai-tunnel] ``` The `provider_key` name is resolved to the provider key's derived ID when AISIX loads the file. An unknown name fails validation. Exact entries in `allowed_routes` are also checked against the routes defined in the file; wildcard patterns are allowed. Validate the assembled complete file before loading it: ``` aisix validate --resources resources.yaml ``` Because this example introduces `PASSTHROUGH_CALLER_KEY`, start or recreate the gateway with that variable in its process environment. After it loads, use the same value for the verification request: ``` export AISIX_API_KEY="$PASSTHROUGH_CALLER_KEY" ``` ## Verify the Route[​](#verify-the-route "Direct link to Verify the Route") Export the gateway origin without a trailing slash: ``` # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Request OpenAI's native model list through AISIX: ``` curl --fail-with-body -sS \ "$AISIX_PROXY/passthrough/openai/v1/models" \ -H "Authorization: Bearer $AISIX_API_KEY" ``` The route strips its `path_prefix`, avoids duplicating the `/v1` segment already present in `target_url`, injects the OpenAI provider credential, and relays the upstream response. ## Compare Cloud and Resources-File References[​](#compare-cloud-and-resources-file-references "Direct link to Compare Cloud and Resources-File References") Most route fields use the same names in both management paths. References to provider and caller credentials differ: | Purpose | AISIX Cloud Admin API | Recommended resources-file field | Explicit ID accepted in a resources file | | ---------------------------- | ------------------------- | -------------------------------- | ---------------------------------------- | | Injected provider credential | `provider_key_id` (UUID) | `provider_key` (`display_name`) | `provider_key_id` | | Anonymous caller principal | `anonymous_key_id` (UUID) | `anonymous_key` (`display_name`) | `anonymous_key_id` | Name references in a resources file receive stronger load-time checking and produce candidate names when a reference is unknown. Prefer them over explicit IDs. AISIX Cloud validates that the provider key is visible to the environment and that the anonymous caller key belongs to it. The route `name` is fixed after creation in AISIX Cloud. In a resources file, changing the name changes the route identity, so update every caller's `allowed_routes` entry at the same time. ## Configuration Reference[​](#configuration-reference "Direct link to Configuration Reference") | Field | Behavior | | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `name` | Route identity referenced by callers' `allowed_routes` patterns and recorded on usage events. | | `path_prefix` | Gateway path prefix, matched on segment boundaries. A `target_url` route strips it before joining the remaining path to the target. A `preserve_host` route keeps the complete path. A path-only route cannot claim `/v1`, `/mcp`, `/a2a`, `/admin`, `/livez`, `/readyz`, or `/metrics`; a route that also matches `hosts` may use those upstream-owned paths. | | `hosts` | Case-insensitive inbound `Host` allowlist; a port is ignored. A leading `*.` wildcard matches one additional label and must retain at least two literal labels. | | `target_url` | Explicit upstream base URL. Configure exactly one target shape: `target_url`, or `preserve_host: true`. | | `preserve_host` | Derive `https://` as the target. It is accepted only with `hosts`, which bounds the derived destination. | | `auth_mode` | `gateway_key` (default), `header_key`, or `anonymous`. | | `auth_header_name` | Lowercase side-channel header containing the gateway credential in `header_key` mode. AISIX strips it before forwarding, unless `forward_client_headers` names it in full; a glob such as `x-*` does not reach it. `authorization`, `proxy-authorization`, `cookie`, `set-cookie`, and `x-api-key` are rejected. | | `source_cidrs` | Client source allowlist. It is required and nonempty for `anonymous`; for other modes it is optional hardening. | | `credential_mode` | `inject` (default) or `forward_client`. | | `forward_client_headers` | Inbound client headers relayed upstream even when this route would otherwise strip them, as exact names or single-`*` globs, matched case-insensitively. Empty by default. A patch replaces the stored list; send `null` to clear it. In the dashboard it is **Forward client headers** under **Advanced**, one entry per line. See [Upstream Request Headers](https://docs.api7.ai/ai-gateway/models/upstream-request-headers.md#forward-client-headers). | | `identity_header` | Optional lowercase device-injected identity header. AISIX records a bounded value as `client_identity` and strips the header before forwarding, unless `forward_client_headers` names it in full; a glob such as `x-*` does not reach it. Configure it only behind a trusted device that removes and replaces any client-supplied value. `authorization`, `proxy-authorization`, `cookie`, `set-cookie`, and `x-api-key` are rejected. | | `timeout_ms` | Bounds upstream response headers and non-SSE body reads. It does not bound a healthy SSE relay. | | `enabled` | A disabled route matches nothing. Default: `true`. | At least one match dimension, `path_prefix` or `hosts`, is required. When both are present, the request must satisfy both. ## Gateway Authentication Modes[​](#gateway-authentication-modes "Direct link to Gateway Authentication Modes") * **`gateway_key`** reads the standard gateway credential from `Authorization: Bearer` or `x-api-key`. * **`header_key`** reads the gateway credential from `auth_header_name`, leaving `Authorization` available for the caller's upstream credential. This is the standard forward-proxy pairing with `forward_client`. * **`anonymous`** accepts no gateway credential. The request runs as the configured caller-key principal and must originate within `source_cidrs`. Every resolved principal still needs the route name granted by `allowed_routes`; `*` grants every route. A valid key without a matching grant receives `403`. In AISIX Cloud, budgets that already apply to the resolved caller are checked before dispatch. Passthrough usage currently carries no model ID, is recorded with zero cost, and does not add spend to those budgets. The open-source AISIX gateway has no local budget resource. ## Upstream Credential Modes[​](#upstream-credential-modes "Direct link to Upstream Credential Modes") * **`inject`** strips inbound credential headers and injects the configured provider key. AISIX uses `x-api-key` plus `anthropic-version` for Anthropic and `Authorization: Bearer` for other providers. The provider key's configured header-strip and TLS settings also apply. * **`forward_client`** forwards the caller's upstream credential when the gateway credential arrived through `header_key`, or when the route is anonymous. With `gateway_key`, AISIX removes `Authorization` and `x-api-key` because either may contain the gateway credential. A route relays the caller's other headers by default and strips a small set: hop-by-hop and transport headers, `host`, `content-length`, the `x-aisix-*` namespace, `proxy-authorization`, and the caller's W3C trace headers. AISIX consumes valid W3C context for its own [OTLP trace](https://docs.api7.ai/ai-gateway/observability/exporters.md#understand-otlp-trace-structure) and sends a gateway request ID upstream. `forward_client_headers` overrides that strip, and is the one way to put a header back that the route would otherwise remove. Naming `authorization` under `auth_mode: gateway_key` therefore relays the caller's own credential — the very header the gateway consumed to authenticate them — in place of the injected one, which is what lets an internal service that authorizes on the end user keep doing so. `host`, `content-length`, hop-by-hop headers, and `x-aisix-*` are stripped whatever the patterns say, and a credential or trace-context header must be [named exactly](https://docs.api7.ai/ai-gateway/models/upstream-request-headers.md#headers-that-must-be-named-exactly) rather than matched by a wildcard. The route's own `auth_header_name` and `identity_header` are read the same way: AISIX consumes both, so a glob does not reach either and naming one in full is what forwards it. For every other name the route strips, a glob is enough — including a provider key `strip_headers` entry, though three of that list's four defaults (`authorization`, `cookie`, `x-api-key`) are credential slots that still need their own entry, leaving `set-cookie` the only default a glob restores. There is no credential fallback: an `inject` route without a resolvable provider key fails closed, and a `forward_client` route cannot carry a provider key reference. ## Envelope Detection and Usage[​](#envelope-detection-and-token-usage "Direct link to Envelope Detection and Usage") AISIX detects the request shape for extraction only; detection does not change the relayed body. If several recognized carrier fields appear, detection uses this order: 1. `messages` for OpenAI-compatible chat or Anthropic Messages traffic. 2. `input` for the OpenAI Responses shape. 3. `prompt` for legacy completions or fill-in-the-middle traffic. 4. Opaque handling for every other body, including JSON-RPC, REST, non-JSON, and empty bodies. Detection selects the text presented to guardrails and the token fields recorded on usage events. If a detected shape yields no text, AISIX scans the complete body instead. Request and response bodies are still relayed without schema translation. Opaque buffered responses do not receive speculative token extraction. An opaque SSE stream can report usage through a top-level `usage` object or a flat token report on an `event: usage` or `event: token_usage` frame. ## Rate Limits[​](#rate-limits "Direct link to Rate Limits") Caller API-key, team, and member request limits apply before dispatch. On an `inject` route, a top-level JSON `model` that resolves to a configured AISIX model of the same provider also reserves that model's request limits. `forward_client` routes do not perform this model lookup. Request-count dimensions (`rps`, `rpm`, `rph`, and `rpd`) are enforced. AISIX also rejects a request when an applicable `tpm` or `tpd` counter is already exhausted, but passthrough token usage does not increment those counters. Use recorded usage for telemetry, not passthrough token-quota enforcement. Concurrency is checked before upstream dispatch. For SSE, the reservation is released when AISIX returns the streaming response, not when the stream ends. Passthrough routes have no rate-limit field or policy scope. To apply different request limits to different routes, grant them to separate caller keys and configure limits for those caller identities. ## Guardrails and Streaming[​](#guardrails-and-streaming "Direct link to Guardrails and Streaming") In AISIX Cloud, a guardrail can be attached to one passthrough route by selecting the **Passthrough routes** scope. The attachment uses the route UUID. Environment, caller-key, and team guardrails can also apply. The open-source resources file declares attachments in its `guardrail_attachments` collection, so a file-defined guardrail reaches passthrough traffic only if an attachment scopes it there — `scope_type: env`, or `scope_type: passthrough_route` naming the route. Input guardrails run before upstream dispatch. A block returns `422` without contacting the upstream. Buffered responses are checked before delivery. Once an SSE response has started, a block ends the stream with an SSE `content_filter` error frame; it cannot change the HTTP status to `422`. SSE responses relay incrementally unless a hold-back guardrail buffers frames for inspection. AISIX does not rewrite provider-native bodies to apply redaction. For the built-in `pii` guardrail, a mask-action-only match is forwarded without masking; use a block action when the matched content must not reach the upstream. Guardrail kinds such as Presidio and Lakera instead block a maskable result when passthrough has no write-back channel. ## Audit Capture[​](#audit-capture "Direct link to Audit Capture") For successfully relayed traffic, an observability exporter with `content_mode: full` receives the request body as a string, subject to the exporter's content cap. Buffered responses record extracted text when the response matches a supported extraction shape and otherwise record the body as text. Streamed responses record accumulated extracted text; opaque data payloads are retained as text. Captured content is exporter-only and is never sent through the AISIX Cloud telemetry path. Usage events carry the route name, caller, recorded token counts, and `client_identity`. External exporters can expose these values. The current AISIX Cloud Request Logs UI shows caller and token metadata, but it does not display `passthrough_route_name` or `client_identity`. ## Migrate from the Removed Implicit Tunnel[​](#migrate-from-the-implicit-tunnel "Direct link to Migrate from the Removed Implicit Tunnel") The removed tunnel selected a provider key indirectly through an accessible model. Explicit routes replace that ambiguous selection with a fixed target and credential binding. For each provider prefix that clients still use: 1. Create an explicit route with the old path prefix and the intended upstream target. 2. Bind the intended provider key in `inject` mode. 3. Grant the route name to every caller that should retain access. 4. Verify the route before moving or restarting clients. Clients can keep their existing `/passthrough//...` URL when the explicit route claims the same prefix. Requests answer ordinary `404` until that route is active. Two former tunnel behaviors do not carry over: passthrough routes make one upstream attempt without transport retries, and their failures do not mark configured models for cooldown. ## Errors[​](#errors "Direct link to Errors") | Status or signal | Meaning | | -------------------------- | ------------------------------------------------------------------------------------------------------- | | `401` | Missing or invalid gateway credential for the route's authentication mode. | | `403` | The resolved caller key does not grant the route, or the client source is outside `source_cidrs`. | | `404` | An unclaimed `/passthrough/*` path reaches the ordinary empty-body not-found path. | | `422` | A guardrail blocked the request before dispatch or blocked a buffered response before delivery. | | SSE `content_filter` frame | A guardrail blocked content after a stream had started. | | `429` | A gateway request limit or budget check rejected the request, or the upstream returned a relayed `429`. | | Other upstream status | AISIX relays the upstream status and body after filtering response headers. | The unmatched `404` has an empty body. Other AISIX-generated failures use the gateway error envelope; upstream error statuses and bodies are relayed after response-header filtering. ## Next Steps[​](#next-steps "Direct link to Next Steps") * Receive TLS-terminated IDE traffic with host-matched routes: [Forward Proxy for IDE AI Traffic](https://docs.api7.ai/ai-gateway/deployment/forward-proxy.md). * Configure guardrail behavior: [Guardrail Behavior](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md). * Export request and response content: [Observability Exporters](https://docs.api7.ai/ai-gateway/observability/exporters.md). --- # Realtime API AISIX AI Gateway relays the OpenAI Realtime API over WebSocket at `GET /v1/realtime`. Before accepting the connection, the gateway authenticates the client, resolves the model alias, and enforces access and rate-limit policy. It then relays events between the client and the provider in both directions. In this guide, you will connect a Realtime client through the gateway and review the session behavior that matters for this endpoint. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias backed by a provider that serves the OpenAI Realtime protocol. Export the gateway connection and request values for the server-side example: ``` # AISIX_PROXY uses http or https and has no trailing slash or endpoint path. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="realtime-prod" ``` Provider support covers OpenAI-compatible providers (adapter `openai`, including custom `api_base` deployments) and Azure OpenAI (adapter `azure-openai`, `/openai/realtime` with `api-key` authentication). Gemini Live and Bedrock use different realtime event protocols and are not served by this endpoint. ## Connect a Client[​](#connect-a-client "Direct link to Connect a Client") Select the model with the `model` query parameter. Server-side clients authenticate with the standard headers: ``` import WebSocket from "ws"; const realtimeUrl = new URL("/v1/realtime", process.env.AISIX_PROXY); realtimeUrl.protocol = realtimeUrl.protocol === "https:" ? "wss:" : "ws:"; realtimeUrl.searchParams.set("model", process.env.AISIX_MODEL); const ws = new WebSocket(realtimeUrl, { headers: { Authorization: `Bearer ${process.env.AISIX_API_KEY}` }, }); ``` Browser clients cannot set WebSocket headers. Pass the caller API key as a subprotocol item instead, matching the flow used by OpenAI browser examples. The gateway echoes the `realtime` subprotocol back: ``` const aisixProxy = "YOUR_AISIX_GATEWAY_ORIGIN"; const aisixApiKey = "YOUR_CALLER_API_KEY"; const aisixModel = "realtime-prod"; const realtimeUrl = new URL("/v1/realtime", aisixProxy); realtimeUrl.protocol = realtimeUrl.protocol === "https:" ? "wss:" : "ws:"; realtimeUrl.searchParams.set("model", aisixModel); const ws = new WebSocket(realtimeUrl, [ "realtime", `openai-insecure-api-key.${aisixApiKey}`, ]); ``` After the connection is established, send and receive Realtime events exactly as you would against the provider directly. AISIX relays events such as `session.update`, audio buffers, `response.create`, and server events without changing their shape. The one value AISIX rewrites is the model name inside a session object. The `session.created` and `session.updated` events name the model alias the client connected with, not the provider's own model ID. A `session.update` that sends that alias back is translated to the provider's model ID on the way upstream, so a client can echo the session object it was given. ## Authentication and Policy[​](#authentication-and-policy "Direct link to Authentication and Policy") Authentication, model access checks, client IP restrictions, and rate limits run **before** the WebSocket upgrade completes. AISIX Cloud budget checks run at the same point. A request that fails any of these checks is rejected at the HTTP handshake (401, 403, or 429), which clients observe as a failed connection attempt. During the session, configured guardrails scan text events in both directions. A blocked event produces an OpenAI-shaped `error` event followed by connection close. ## Usage Tracking[​](#usage-tracking "Direct link to Usage Tracking") The gateway harvests usage from the provider's `response.done` events (and transcription-completed events for transcription sessions) and records one aggregated usage event per session, including cached-token counts. Total session tokens count toward token-based rate limits. ## Session Limits[​](#session-limits "Direct link to Session Limits") AISIX applies an idle cap between events in either direction. It resolves the cap from the direct model's `stream_timeout`, then its `timeout`, and then the deployment-wide `upstream.stream_timeout_ms` or `upstream.timeout_ms` defaults. The default deployment cap is 6000 seconds. A silent session past the resolved deadline closes with code `1001` and reason `idle timeout`. To opt a model out of the deployment backstop, set `timeout: 0` and make sure the model does not set a nonzero `stream_timeout`. Deployment operators can instead set both upstream timeout defaults to `0`. See [How the Timeouts Relate](https://docs.api7.ai/ai-gateway/reference/configuration-files.md#how-the-timeouts-relate) for the complete precedence rules. Upstream connection failures close the session with code `1011`, and count toward the model's cooldown when the model enables it. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected a Realtime client through the gateway. Next, review [Speech and Audio](https://docs.api7.ai/ai-gateway/endpoints/audio.md) for the non-Realtime audio endpoints. --- # Request Lifecycle AISIX sits between applications and AI providers. Applications send requests to the proxy API with a caller credential and a model alias. AISIX uses that information to apply access control, resolve the upstream target, enforce AI traffic policy, and record what happened. Each request moves through the following stages: Caller**Application request** AISIX gateway **Caller Authentication**Caller API key by hash, or a JWT mapped to a caller key **Model Resolution**Alias becomes a direct, routing, semantic, or ensemble shape **Request Controls**Input guardrails, AISIX Cloud budget, rate limits, cache lookup — a rejection or cache hit exits early **Provider Dispatch**Provider credential, upstream model name, provider adapter Upstream · outside the gateway**Provider model API**Retries and failover **Response Handling**Output guardrails, usage and telemetry Caller**Response returned** ## Caller Authentication[​](#caller-authentication "Direct link to Caller Authentication") Each proxy request presents either a caller API key or a JWT issued by a configured OIDC provider. AISIX looks up a plaintext key by its hash. For a JWT, AISIX verifies the issuer, signature, and required claims, then maps the external identity to its bound caller API key. After authentication, both paths use the resolved caller API key as the authorization identity. The key controls which model aliases the caller can use and which traffic controls apply, so application teams do not need direct provider credentials. See [Caller API Keys](https://docs.api7.ai/ai-gateway/traffic-controls/caller-api-keys.md) for plaintext gateway credentials and [JWT Authentication](https://docs.api7.ai/ai-gateway/traffic-controls/jwt-authentication.md) for external identity-provider credentials. ## Model Resolution[​](#model-resolution "Direct link to Model Resolution") The model value in the request is the caller-facing alias. AISIX resolves that alias to one of four dispatch shapes: * A direct model, which points to one upstream model through one provider credential. A direct model can also include embedding metadata when a semantic router uses it for similarity comparisons. * A routing model, which lets AISIX choose one target model by failover, weighted round-robin, consistent hashing, cost, latency, or load. * A semantic model, which selects a target model by comparing request text with configured route examples. * An ensemble model, which sends a chat request to panel models and uses a judge model to synthesize the response. Only one dispatch shape can be configured on a model resource. For the complete resource relationships, see [Models and Providers](https://docs.api7.ai/ai-gateway/models/resource-model.md). ## Request Controls[​](#request-controls "Direct link to Request Controls") AISIX can stop a request before it reaches a provider. Input guardrails can inspect and reject unsafe content, AISIX Cloud deployments can enforce request budgets, and caller API keys and model aliases can carry rate limits. Response caching can return a stored chat completion before an upstream call. For a stage-by-stage view of where each control acts, see the [Traffic Controls](https://docs.api7.ai/ai-gateway/traffic-controls/overview.md) request-path diagram. ## Provider Dispatch[​](#provider-dispatch "Direct link to Provider Dispatch") After the request is allowed, AISIX dispatches it to the selected provider using the provider key and adapter configured by the operator. Applications keep their gateway-facing API shape while AISIX handles provider credentials, upstream model names, base URLs, and provider-specific request handling. ## Response Handling[​](#response-handling "Direct link to Response Handling") Provider responses return through AISIX. Output guardrails can inspect generated text before the response reaches the caller. AISIX records usage and telemetry once the caller is authenticated and the request parses, whether it succeeds or fails. Operators can review requested aliases, resolved models, provider attempts, token usage, latency, and errors. ## Deployment Boundary[​](#deployment-boundary "Direct link to Deployment Boundary") For the open-source AISIX gateway, operators can declare gateway resources in a `resources.yaml` file. With AISIX Cloud, the control plane owns resource management and projects accepted configuration to the AISIX gateway. The proxy request lifecycle remains the same from the caller's perspective: applications call the proxy API, and AISIX applies the configured model access, routing, controls, and observability behavior. --- # Rerank Rerank requests reorder candidate documents for a query before an application uses those documents in search, retrieval, or RAG workflows. AISIX AI Gateway exposes `POST /v1/rerank` so rerank traffic can use the same caller API keys, model aliases, upstream credentials, and request-side policy as the rest of the gateway traffic path. In this guide, you will send a rerank request through AISIX and review the provider requirement for this endpoint. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias whose configured provider label is OpenAI, Cohere, or Jina. Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="rerank-prod" ``` ## Send a Rerank Request[​](#send-a-rerank-request "Direct link to Send a Rerank Request") Send the rerank request through the gateway proxy with the AISIX model alias in the request body: ``` curl -sS -X POST "${AISIX_PROXY}/v1/rerank" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "query": "gateway docs", "documents": ["doc a", "doc b", "doc c"] }' \ -o aisix-rerank-response.json ``` AISIX resolves the model alias, checks the caller API key, runs supported input policy checks on the query and document text, rewrites only the model field to the upstream model ID, and forwards the request to the upstream rerank endpoint. The response keeps the upstream rerank response shape: ``` { "results": [ { "index": 1, "relevance_score": 0.95 }, { "index": 0, "relevance_score": 0.42 }, { "index": 2, "relevance_score": 0.18 } ] } ``` Some providers include additional fields such as a response ID, model name, usage, or metadata. Check that the response contains ranked results: ``` jq '.results | length' aisix-rerank-response.json ``` The command should print the number of returned results: ``` 3 ``` ## Provider Requirement[​](#provider-requirement "Direct link to Provider Requirement") AISIX accepts rerank requests only when the resolved model is configured with the OpenAI, Cohere, or Jina provider label. These provider paths share the common rerank fields for model, query, and documents. Optional rerank fields are forwarded unchanged and are not normalized across providers. When the resolved model uses another provider label, AISIX returns 400 before sending the request upstream. This prevents a rerank request from being sent to a provider route that does not use the expected rerank format. Voyage AI also exposes a rerank API, but its request and response fields differ from the supported rerank format. AISIX needs a dedicated adapter before treating it as compatible. For Cohere and [Jina](https://docs.api7.ai/ai-gateway/providers/jina.md#add-a-rerank-model), configure the provider key base URL for the API root in the provider's reference. AISIX appends the rerank path and avoids duplicating a common API-version segment when the base URL already ends in a version prefix. ## Rerank Behavior[​](#rerank-behavior "Direct link to Rerank Behavior") AISIX forwards the request body with only the model field rewritten. It does not add chat-completion fields or translate provider-specific rerank parameters. Input guardrails can inspect the query and document text before AISIX calls the provider. Output guardrails do not inspect reranked response content on this path because the response contains ranking results, not generated text. Successful rerank responses are returned as upstream bytes with the upstream content type, with one exception: when the upstream response carries a top-level `model` field, AISIX rewrites it to the model alias the request addressed, so the response names the same model the caller asked for rather than the upstream model ID. A response that carries no top-level `model` field does not gain one, and a model name nested elsewhere in the response is left as the provider wrote it. AISIX parses usage only as a best-effort telemetry step. If usage is missing or uses an unrecognized shape, AISIX still returns the rest of the upstream response unchanged. If a guardrail does not block a request, check whether the configured guardrail can inspect the query or document text. If the request returns 400 before reaching the provider, check the resolved model's provider label. If the upstream returns 404, check the provider key base URL. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen how AISIX proxies rerank requests and where provider support is intentionally narrow. Next, continue with [Image Generation](https://docs.api7.ai/ai-gateway/endpoints/image-generation.md) to review another provider-specific endpoint family. --- # Responses API The Responses API is OpenAI's response-generation endpoint for applications that use its input/output item format instead of the chat-completions message format. AISIX AI Gateway exposes this route for Responses API clients while keeping caller authentication, model aliases, upstream credentials, and gateway policy in the gateway. Use this route when an application or tool already speaks the Responses API. When the model's upstream serves the Responses API itself, AISIX forwards the request there. Otherwise AISIX translates the request through the provider adapter and returns a Responses API result to the caller. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias backed by a provider that can handle the translated request shape. Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` ## Send a Responses Request[​](#send-a-responses-request "Direct link to Send a Responses Request") Send the request through the gateway proxy with the AISIX model alias in the request body: ``` curl -sS -X POST "${AISIX_PROXY}/v1/responses" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "input": "Say hello from AISIX." }' ``` AISIX resolves the model alias and chooses the provider path for the selected model. OpenAI-backed models are forwarded to the upstream Responses API without body translation. Other providers use the cross-provider bridge. The response body is returned in the upstream Responses API format: ``` { "id": "resp_***", "object": "response", "model": "gpt-4o-prod", "output": [ { "type": "message", "role": "assistant", "content": [ { "type": "output_text", "text": "Hello from AISIX." } ] } ], "usage": { "input_tokens": 12, "output_tokens": 5, "total_tokens": 17 } } ``` ## Provider Behavior[​](#provider-behavior "Direct link to Provider Behavior") AISIX handles Responses requests in two ways: | Upstream | Behavior | | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Serves the Responses API | AISIX rewrites the request model to the upstream model ID and forwards the request to the upstream Responses API. Upstream-specific Responses features pass through when the upstream supports them. | | Does not serve it | AISIX translates supported Responses fields into the gateway's chat format, dispatches through the provider adapter, and returns a Responses API result. | Which one applies is decided per provider key. Without a surface declaration, AISIX forwards natively for the `openai` provider and translates for every other provider. That default is wrong in both directions for some endpoints: an OpenAI-compatible endpoint reached through the `openai` provider with a custom `api_base` may have no `/v1/responses` route, and an endpoint on another provider may serve one. Declare the key's API surfaces to say which it is — see [Declare the API Surfaces](https://docs.api7.ai/ai-gateway/models/provider-keys.md#declare-the-api-surfaces). The bridge supports common text, tool call, tool result, sampling, and streaming fields. It carries `reasoning.effort` into the canonical chat request as `reasoning_effort`; provider adapters that support an effort control then translate it to their wire field. For example, Anthropic receives `output_config.effort`. You can also rewrite the value per direct model with [Reasoning Effort Mapping](https://docs.api7.ai/ai-gateway/models/reasoning-effort-mapping.md). OpenAI-only Responses features that do not have a provider-neutral chat equivalent are not forwarded on the bridged path. The following fields are ignored on translated requests: * `reasoning` members other than `effort`, such as `summary` * `store` * `previous_response_id` * hosted tools such as `web_search`, `file_search`, and `code_interpreter` * `text`, `metadata`, `service_tier`, and other OpenAI-specific controls ## Policy and Usage Behavior[​](#policy-and-usage-behavior "Direct link to Policy and Usage Behavior") Input guardrails can inspect request text before AISIX calls the provider. Output guardrails can inspect non-streaming responses before content reaches the caller. If an output guardrail blocks the response, AISIX returns a content-policy error to the caller and records the blocked request for observability. For streaming requests, AISIX preserves the Responses SSE shape. On OpenAI-backed models, AISIX can pass through upstream SSE. On bridged providers, AISIX encodes provider stream chunks into Responses events. If output guardrails are enabled, AISIX buffers the stream for policy inspection before returning it or blocking it. A buffered frame that AISIX cannot parse is dropped rather than released unscanned, and a response left with nothing to return is refused with `422`. See [Frames AISIX Cannot Scan](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md#frames-aisix-cannot-scan). For successful responses, the gateway records usage when the upstream response includes token usage. Streaming usage is emitted when AISIX receives terminal usage information from the stream. On bridged providers whose token accounting differs from OpenAI's, AISIX converts the counts into the Responses API shape: `input_tokens` is the full input, `input_tokens_details.cached_tokens` is the cache-read subset of it, `input_tokens_details.cache_creation_tokens` is the cache-write subset when the upstream reported one, and `total_tokens` is `input_tokens + output_tokens`. See [Token Usage](https://docs.api7.ai/ai-gateway/endpoints/openai-client-to-anthropic.md#token-usage) for a worked example with an Anthropic upstream — it is written in Chat Completions field names, which map to these one for one (`prompt_tokens` to `input_tokens`, `completion_tokens` to `output_tokens`, `prompt_tokens_details` to `input_tokens_details`, `completion_tokens_details` to `output_tokens_details`). One shape difference: `output_tokens_details.reasoning_tokens` is always present on a Responses reply, including when it is `0`, whereas Chat Completions omits the block entirely. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen when to use Responses API through AISIX and how provider handling differs between direct forwarding and bridging. Next, continue with [Text Completions](https://docs.api7.ai/ai-gateway/endpoints/text-completions.md) for the legacy completions route, or [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md) when your Responses API client depends on SSE behavior. --- # Streaming AISIX AI Gateway can stream proxy responses to clients that expect server-sent events. Streaming keeps the gateway responsibilities in place: AISIX still authenticates the caller API key, resolves the model alias, applies supported policy, and forwards the request to the selected upstream provider. Streaming endpoints preserve the client-facing stream format for each route, but gateway policy can affect delivery. In this guide, you will send an OpenAI-compatible streaming request, then review where streaming behavior differs across endpoint families. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias backed by a provider and model that support streaming. Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` ## Send a Streaming Request[​](#send-a-streaming-request "Direct link to Send a Streaming Request") The request keeps the model value as the AISIX model alias and asks the upstream to stream the response: ``` curl -sS -N -X POST "${AISIX_PROXY}/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "stream": true, "messages": [ {"role": "user", "content": "Stream a short greeting."} ] }' ``` The response is an OpenAI-style SSE stream. A direct HTTP client sees `data:` frames similar to the following: ``` data: {"id":"***","object":"chat.completion.chunk","choices":[{"delta":{"content":"Hello"}}]} data: [DONE] ``` An OpenAI-compatible SDK reads the same stream through its normal streaming API. ## Choose a Streaming Path[​](#choose-a-streaming-path "Direct link to Choose a Streaming Path") Choose the proxy endpoint that matches the client response format: | Client format | Proxy path | Behavior | | ---------------------- | ---------------------- | ---------------------------------------------------------------------------------------------------------------------------- | | OpenAI-compatible chat | `/v1/chat/completions` | Returns OpenAI-style SSE chunks for OpenAI-compatible SDKs and direct SSE consumers. | | Anthropic Messages | `/v1/messages` | Returns Anthropic-style SSE events. Upstreams that serve this route stream natively; the rest stream through translation. | | OpenAI Responses API | `/v1/responses` | Returns Responses SSE events. Upstreams that serve this route stream natively; the rest stream through the Responses bridge. | Chat Completions audio output is an exception to the general streaming path. AISIX does not currently preserve upstream `delta.audio` chunks. Keep [Chat Completions audio](https://docs.api7.ai/ai-gateway/endpoints/chat-audio.md) non-streaming, or use the [Realtime API](https://docs.api7.ai/ai-gateway/endpoints/realtime.md) for incremental, bidirectional audio. Which upstreams stream natively is decided per provider key — see [Declare the API Surfaces](https://docs.api7.ai/ai-gateway/models/provider-keys.md#declare-the-api-surfaces). Use the client format as the deciding factor. Do not switch to `/v1/messages` or `/v1/responses` only because the upstream provider changes. ## Review Streaming Behavior[​](#review-streaming-behavior "Direct link to Review Streaming Behavior") Streaming starts after AISIX has accepted the request and selected the target. If a client aborts a stream mid-response, the gateway remains healthy and continues serving later requests. Streaming chat completions are not cached. Each streaming request dispatches upstream, even when a cache policy exists. For multi-target models, AISIX can retry or fail over before any stream bytes are sent to the client. Once bytes have reached the client, the selected upstream remains responsible for the stream and AISIX does not switch targets. If output guardrails are enabled, AISIX may hold, scan, or terminate streamed output depending on the endpoint and guardrail policy. For chat-completions and Messages streams, a blocked output can be signaled with a terminal SSE error event instead of normal stream completion. For Responses API streaming, AISIX can buffer the stream for policy inspection before returning it or blocking it. When the upstream disconnects mid-stream, treat the partial stream as incomplete unless the endpoint-specific client contract says otherwise. If chunks do not arrive, confirm the request includes the streaming flag shown in the example and that the client reads server-sent events. For Responses API streaming, check whether the selected provider supports the translated request shape. If a stream ends early, check the upstream provider status, gateway logs, and any configured stream timeout. While the model has produced nothing yet, the gateway emits an SSE comment every `downstream.sse_keepalive_interval_secs` (15 seconds by default) so a proxy between the client and the gateway does not treat a model that is slow to its first token as an abandoned connection. Conforming SSE clients ignore these comments. See [Tune the Downstream Connection Layer](https://docs.api7.ai/ai-gateway/reference/configuration-files.md#tune-the-downstream-connection-layer). ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen how streaming behavior differs across AISIX endpoint families. For the main OpenAI-style path, see [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). For Anthropic-style events, see [Anthropic-Style Messages API](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md). For stream errors and response headers, see [Headers and Error Codes](https://docs.api7.ai/ai-gateway/reference/headers-and-error-codes.md). --- # Text Completions Some applications send a single prompt and expect a text-completions response instead of a chat message response. AISIX AI Gateway supports that OpenAI-compatible request shape through the completions proxy route. Text completions are mainly useful for existing prompt-based clients. For new conversational applications, the chat-completions route usually gives broader provider support and chat-oriented client features. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias backed by a provider and model that support text completions. Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="text-prod" ``` ## Send a Text-Completions Request[​](#send-a-text-completions-request "Direct link to Send a Text-Completions Request") Send the request through the gateway proxy with the AISIX model alias in the request body: ``` curl -sS -X POST "${AISIX_PROXY}/v1/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "prompt": "Write one sentence about API gateways.", "max_tokens": 40 }' ``` AISIX resolves the model alias, checks the caller API key, rewrites the upstream model ID, and forwards the remaining request body to the provider's completions endpoint. The response keeps the OpenAI-compatible text-completions format: ``` { "id": "cmpl-***", "object": "text_completion", "model": "text-prod", "choices": [ { "text": "API gateways help manage, secure, and observe traffic between clients and services.", "index": 0, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 7, "completion_tokens": 13, "total_tokens": 20 } } ``` ## Provider and Gateway Behavior[​](#provider-and-gateway-behavior "Direct link to Provider and Gateway Behavior") The text-completions route is an OpenAI-compatible proxy path. Providers that support completions can receive the request through their configured adapter. Providers that do not support completions return 501 with error type `not_implemented`. The route does not stream. A request that sets `stream: true` is rejected with `400` before AISIX contacts the provider, so the provider never generates or bills for a response: ``` { "error": { "message": "request payload is invalid: `stream` is not supported on /v1/completions; use /v1/chat/completions for streaming", "type": "invalid_request_error" } } ``` Requests that set `stream: false`, or that omit `stream`, are unaffected and are forwarded as before. To stream, use [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). Input guardrails can inspect string prompts before AISIX calls the provider. Prompt arrays that contain strings are also inspectable. Token-ID prompts are forwarded, but they do not contain text for guardrails to scan. Output guardrails scan the completion text before AISIX returns it, the same as the chat-completions path. If an output guardrail blocks the response, AISIX returns `422`. The provider has already generated and billed for the response at that point, so the request's token usage is still recorded. Successful responses keep the OpenAI-compatible completions response format. When the upstream response includes token usage, AISIX records usage for the request. ## Text-Completions Behavior[​](#text-completions-behavior "Direct link to Text-Completions Behavior") Use chat completions when the application can send role-based messages, use tools, stream assistant output, or work across the broadest set of configured provider backends. If a request returns 501, the resolved provider path does not support text completions. Use a model backed by a provider that supports completions, or move the application to chat completions. If an input guardrail does not block the prompt, check whether the prompt contains inspectable text. String prompts and arrays of strings can be scanned. Token-ID prompt arrays do not expose text to input guardrails. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen when to use text completions through AISIX and why chat completions are usually the better default for new integrations. For the main chat path, see [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). --- # Tool Calling Tool calling allows a model to ask the application to run a named function, such as looking up weather, querying a database, calling an internal service, or running business logic. The model does not run the tool. It returns a structured tool call, and the application parses the arguments, runs the function, and sends the result back so the model can continue the conversation. AISIX AI Gateway carries these tool-calling requests through the OpenAI-compatible chat-completions path and includes targeted translation between OpenAI-style and Anthropic-style tool formats. Applications can keep provider credentials and model routing behind AISIX while preserving the tool loop their SDK or agent framework expects. In this guide, you will send a tool definition through AISIX, review the follow-up tool loop, and choose the request path that matches your client format. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the model alias. * A model alias backed by a provider and model that support tool calling. Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` ## Send a Tool-Calling Request[​](#send-a-tool-calling-request "Direct link to Send a Tool-Calling Request") The example below uses the OpenAI-compatible chat-completions path. It sends a function definition and asks the model to call that function: ``` curl -sS -X POST "${AISIX_PROXY}/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "messages": [ {"role": "user", "content": "What is the weather in Paris? Use the tool if needed."} ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Get weather for a city.", "parameters": { "type": "object", "properties": { "city": {"type": "string"} }, "required": ["city"] } } } ], "tool_choice": { "type": "function", "function": {"name": "get_weather"} } }' ``` The response remains OpenAI-compatible. You should see a tool call in the assistant message: ``` { "choices": [ { "message": { "role": "assistant", "tool_calls": [ { "id": "call_***", "type": "function", "function": { "name": "get_weather", "arguments": "{\"city\":\"Paris\"}" } } ] } } ] } ``` If the model returns plain text instead, first confirm that the selected upstream model supports tool calling and that the request asks for a tool call. ## Continue the Tool Loop[​](#continue-the-tool-loop "Direct link to Continue the Tool Loop") After the model returns a tool call, the application parses the tool arguments, runs the function, and sends the result back through the same chat-completions route. Use the `tool_call_id` returned by the assistant message so the model can connect the tool result to the original call. Send the tool result as a follow-up message: ``` curl -sS -X POST "${AISIX_PROXY}/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "messages": [ {"role": "user", "content": "What is the weather in Paris? Use the tool if needed."}, { "role": "assistant", "content": null, "tool_calls": [ { "id": "call_***", "type": "function", "function": { "name": "get_weather", "arguments": "{\"city\":\"Paris\"}" } } ] }, { "role": "tool", "tool_call_id": "call_***", "content": "Sunny, 21 C." } ] }' ``` The gateway keeps the same caller authentication and model-alias behavior for the follow-up request. Provider credentials and upstream model IDs stay inside AISIX. ## Choose a Tool-Calling Path[​](#choose-a-tool-calling-path "Direct link to Choose a Tool-Calling Path") Use the path that matches the client format your application already uses: | Client format | Proxy path | Behavior | | -------------------------------------------- | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | OpenAI-compatible tools | `/v1/chat/completions` | Use for OpenAI SDKs, OpenAI-style agent frameworks, and applications that expect OpenAI-style tool calls and follow-up tool messages. | | Anthropic-style tools | `/v1/messages` | Fits clients already built around Anthropic Messages. Anthropic upstreams preserve more native behavior than translated upstreams. | | Provider-native tools outside modeled routes | Passthrough routes | Use only when a first-class AISIX route does not model the provider endpoint you need. | For production tool loops, prefer provider-native tool-calling paths when possible and validate the exact client, provider, model, stream mode, and tool schema you plan to run. ## Review Tool-Calling Behavior[​](#review-tool-calling-behavior "Direct link to Review Tool-Calling Behavior") Tool-calling behavior is most predictable when the client format and upstream provider format already match. Anthropic-native upstreams preserve richer Anthropic behavior than translated upstreams, while OpenAI-compatible upstreams preserve the OpenAI-style tool loop directly. OpenAI-style requests to Anthropic-backed models can translate function tools, tool choice, assistant tool calls, and follow-up tool messages into Anthropic Messages API structures. Anthropic-style requests to non-Anthropic upstreams can translate top-level tools and tool choice into OpenAI-style function tools. When a non-Anthropic upstream returns tool calls, AISIX can render them back to Anthropic-style tool-use blocks. Cross-provider translation can work for common tool definitions, tool choice, assistant tool calls, and follow-up tool results, but it does not guarantee full provider parity. Streaming tool calls can also arrive as partial argument fragments, so validate the exact stream behavior before relying on translated tool calling in production. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now seen how AISIX handles tool-calling requests and translated tool definitions. For the default OpenAI-style chat path, see [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). For Anthropic-style requests, see [Anthropic-Style Messages API](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md). --- # Video Generation Video generation lets applications submit prompt-to-video tasks through AISIX while keeping caller authentication, model aliases, upstream credentials, rate limits, and content guardrails in one gateway path. AISIX exposes an OpenAI-compatible video surface with three routes that mirror the provider-side asynchronous workflow: submit a task, poll its status, and download the result. The gateway holds no task state — the returned video ID encodes everything AISIX needs to route later status and download calls to the right provider. In this guide, you will generate a video through AISIX using an Alibaba Model Studio video model and follow the task to a downloadable result. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A running AISIX gateway that can serve proxy requests. * A caller API key that can access the video model alias. * A model alias whose configured provider is one of the supported video providers (see [Endpoint Behavior](#endpoint-behavior)). Every provider except OpenAI needs a provider key whose `api_base` reaches the provider's API — there is no built-in default base URL for them. An OpenAI model falls back to the standard OpenAI base URL when `api_base` is unset. The examples below use an Alibaba Model Studio model. The examples use a model alias configured like the following. The upstream model name is a text-to-video model from the provider's catalog: ``` { "display_name": "wan-video-prod", "model_name": "wan2.7-t2v", "provider_key_id": "YOUR_PROVIDER_KEY_ID" } ``` Export the gateway connection and request values: ``` # AISIX_PROXY has no trailing slash or endpoint path such as /v1. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="wan-video-prod" ``` ## Create a Video Generation Task[​](#create-a-video-generation-task "Direct link to Create a Video Generation Task") Submit the task with the model alias, a prompt, and optionally a duration in seconds: ``` curl -sS -X POST "${AISIX_PROXY}/v1/videos" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${AISIX_MODEL}"'", "prompt": "A miniature city built from cardboard comes alive at night.", "seconds": 5 }' ``` AISIX resolves the alias, runs input guardrails on the prompt, reserves rate-limit capacity, and submits the task to the provider asynchronously. The response is a video job object: ``` { "id": "bW9kZWwtaWQtMTpkMkZ1TFhacFpHVnZMWEJ5YjJROnRhc2stMDE", "object": "video", "model": "wan-video-prod", "status": "queued", "progress": 0, "created_at": 1753257600, "seconds": "5" } ``` The `id` value is an opaque gateway-issued video ID. Store it — the status and download routes take it as the path parameter. Request fields: | Field | Required | Meaning | | --------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `model` | Yes | The AISIX model alias. | | `prompt` | Yes | The text prompt for the video. | | `seconds` | No | Video duration in seconds, as an integer or a numeric string. Forwarded as the provider's own duration parameter — see [Parameter Mapping](#parameter-mapping). | | `size` | No | Pixel dimensions as `WIDTHxHEIGHT`, for example `1280x720`. Each provider expresses output dimensions differently and validates the value against its own per-model list, so check [Parameter Mapping](#parameter-mapping) and the provider's model documentation before setting it. | Unset optional fields are omitted from the upstream request entirely. ## Poll the Task Status[​](#poll-the-task-status "Direct link to Poll the Task Status") Poll the task with the video ID until the status reaches a terminal value: ``` curl -sS "${AISIX_PROXY}/v1/videos/YOUR_VIDEO_ID" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` A finished task reports `completed` and, when the provider states it, the actual video duration: ``` { "id": "bW9kZWwtaWQtMTpkMkZ1TFhacFpHVnZMWEJ5YjJROnRhc2stMDE", "object": "video", "model": "wan-video-prod", "status": "completed", "progress": 100, "created_at": 0, "seconds": "5" } ``` The `status` field is a four-value enum: | Status | Meaning | | ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `queued` | The provider accepted the task and has not started it. | | `in_progress` | The provider is generating the video. | | `completed` | The video is ready to download. | | `failed` | Generation failed, was canceled, or the provider no longer knows the task (for example, an expired task). The response carries an `error` object with the provider's `code` and `message` when available. | AISIX normalizes each provider's own task states onto that enum. Some providers report no distinct queued state, so a submission there starts directly at `in_progress`: | `provider` value | `queued` | `in_progress` | `completed` | `failed` | | ---------------- | --------------------------------------------- | ------------- | ----------- | ---------------------------------------------------- | | `alibaba` | `PENDING` | `RUNNING` | `SUCCEEDED` | `FAILED`, `CANCELED`, `UNKNOWN`, or any other state | | `zhipuai` | Not reported — a task starts at `in_progress` | `PROCESSING` | `SUCCESS` | `FAIL` or any other state | | `volcengine` | `queued` | `running` | `succeeded` | `failed`, `cancelled`, `expired`, or any other state | | `runwayml` | `PENDING`, `THROTTLED` | `RUNNING` | `SUCCEEDED` | `FAILED`, `CANCELLED`, or any other state | | `openai` | `queued` | `in_progress` | `completed` | `failed` or any other state | `progress` reports a real completion percentage for providers that expose one (OpenAI Sora). For providers that do not, it reports `0` until the task completes and `100` afterward. Because the gateway stores no task state, `created_at` is populated on the submit response only; poll responses report `0`. ## Download the Video[​](#download-the-video "Direct link to Download the Video") When the status is `completed`, request the content route. Use `curl -L` so the command works for every provider — AISIX either redirects to the provider's download URL or streams the video itself, depending on how the provider delivers finished files: ``` curl -sS -L -o video.mp4 \ "${AISIX_PROXY}/v1/videos/YOUR_VIDEO_ID/content" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` Both paths save the same MP4. The difference matters when you script around the response: | Delivery | Providers | What the content route returns | | -------------- | -------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Redirect | Alibaba, Zhipu, Volcengine Ark, Runway | `302` with a `Location` header pointing at the provider's signed download URL. The transfer goes directly from the provider's storage to the client and does not pass through the gateway. AISIX only redirects to absolute `http` or `https` URLs. | | Gateway stream | OpenAI | `200` with the MP4 bytes, the provider's `Content-Type` (normally `video/mp4`), and a `Content-Disposition` attachment header. The provider requires its own credential to download the file, so AISIX fetches it with the configured provider key and streams the bytes through. The provider credential is never exposed to the caller. | Streamed responses pass through the gateway chunk by chunk rather than being held in memory, so a large file does not grow the gateway's memory use. Per-chunk reads are bounded by the model's stream timeout: if a slow upstream stalls, the transfer is cut mid-body. When the provider declares a `Content-Length`, the gateway relays it, so an interrupted transfer surfaces to the client as a short read against that length — retry the content request. To inspect which path a provider uses, ask curl for the status without following redirects: ``` curl -sS -o /dev/null -w "%{http_code} %{redirect_url}\n" \ "${AISIX_PROXY}/v1/videos/YOUR_VIDEO_ID/content" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` A redirect provider prints the redirect status and the provider-hosted URL: ``` 302 https://provider-cdn.example.com/videos/task-01/out.mp4 ``` A gateway-stream provider prints `200` with an empty redirect URL. On that path the probe transfers the whole file to `/dev/null`, so run it against a small task: ``` 200 ``` If the task is not finished, the content route returns `400` with a message telling the caller to keep polling. If the task failed, it returns `400` with the provider's failure detail. An error from the provider's download endpoint is always returned as a JSON error envelope, never as a truncated video body. ## Rate Limits and Guardrails[​](#rate-limits-and-guardrails "Direct link to Rate Limits and Guardrails") Before a task reaches the provider, the submit route reserves the caller API key layers and the model's limits, as it does for other modeled routes. Model limits include the inline `rate_limit` and any model-scope rate limit policies. See [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md) and [Rate Limit Policies](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limit-policies.md). The status and content routes reserve only the caller API key layers. Task polling is deliberately exempt from model-level limits: a client that submits a task and then hits the model's submission cap can still poll that task to completion. For gateway-stream providers this also means the video bytes transit the gateway without counting against the model's limits — size the gateway's egress accordingly. Input [guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md) resolved for the request — whether attached to the model, the caller API key, the team, or the environment — scan the prompt before submission. A blocked prompt is rejected before any provider task is created and does not consume the model's rate-limit capacity. ## Endpoint Behavior[​](#endpoint-behavior "Direct link to Endpoint Behavior") * If the provider key's `api_base` ends with the provider's OpenAI-compatible or versioned suffix (`/compatible-mode/v1`, `/api/v1`, or `/v1` for Alibaba; `/api/paas/v4` for Zhipu; `/api/v3` for Volcengine Ark), AISIX derives the provider root automatically — an existing key configured for chat traffic works unchanged. Runway's documented base is the bare host, and OpenAI accepts either the bare host or a `/v1` base. * The submit route requires a JSON media type, such as `application/json`. The video create helper in the OpenAI Python SDK sends `multipart/form-data` on every call, even without a reference asset, so it cannot drive this route — submit with an ordinary HTTP client. The two GET routes are plain GETs and have no such constraint. * OpenAI is the only video provider with a built-in default base URL: an OpenAI model with no `api_base` resolves to the standard OpenAI API. Every other provider requires `api_base` on the provider key. * Each submission that reaches a provider is recorded in usage logs with zero tokens. A request rejected before dispatch — malformed JSON, or a provider outside the allowlist — emits no usage event. Duration-based cost accounting for video tasks is not yet applied to AISIX Cloud budgets. * Submission is not at-most-once. AISIX retries a send-phase transport failure or an upstream `5xx`, and whether the first attempt reached the provider is unknowable, so a retry can start a second billable task whose ID the caller never sees. AISIX deliberately does not retry once the provider has responded and only the body failed to read or decode. ### Supported Providers[​](#supported-providers "Direct link to Supported Providers") AISIX dispatches the video routes on the model alias's own `provider` value, not on the upstream model name; the provider key attached to the alias supplies the credential and `api_base`. The routes accept direct aliases only — a routing or ensemble alias returns `400`. AISIX keeps no allowlist of model IDs. It forwards the configured upstream model name through the fixed mapping below, so a generation still depends on that model accepting the resulting request shape. Check the provider's current catalog, and that model's own parameter rules, before creating an alias. | `provider` value | Provider setup | Video models | Delivery | | ----------------------------------- | ----------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | -------------- | | `alibaba` | [Qwen (Alibaba Cloud)](https://docs.api7.ai/ai-gateway/providers/qwen.md) | Model Studio Wan and HappyHorse text-to-video, such as `wan2.7-t2v`, `wan2.2-t2v-plus`, and `happyhorse-1.1-t2v` | Redirect | | `zhipuai` (`zhipu` also accepted) | [Zhipu AI](https://docs.api7.ai/ai-gateway/providers/zhipuai.md) | CogVideoX, such as `cogvideox-3` | Redirect | | `volcengine` | [Volcengine Ark](https://docs.api7.ai/ai-gateway/providers/volcengine-ark.md) | Ark Seedance, such as `doubao-seedance-2-0-260128` | Redirect | | `runwayml` (`runway` also accepted) | [RunwayML](https://docs.api7.ai/ai-gateway/providers/runwayml.md) | Runway Gen and Runway-hosted models on the text-to-video endpoint, such as `gen4.5`, `veo3.1`, and `seedance2` | Redirect | | `openai` | [OpenAI](https://docs.api7.ai/ai-gateway/providers/openai.md) | Sora: `sora-2` and `sora-2-pro` | Gateway stream | caution OpenAI [deprecated the Videos API and the Sora 2 models](https://developers.openai.com/api/docs/deprecations) on March 24, 2026 and will remove them from the API on September 24, 2026. Aliases backed by `sora-2` or `sora-2-pro` stop working on that date. A model alias whose provider is not in this table returns `501 not_implemented` on submit. Reach those providers' native video APIs through a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) instead. Two provider values that serve chat traffic are deliberately outside this list: `alibaba-cn` and `zai`. Both reach a different API root than their video-capable counterpart, so a video alias must use `alibaba` or `zhipuai`. ### Parameter Mapping[​](#parameter-mapping "Direct link to Parameter Mapping") AISIX maps the unified `seconds` and `size` fields onto each provider's own parameters. The mapping is per provider, not per model: AISIX does not inspect which model family the alias names, so where a provider's newer models expect different parameters, omitting the unified field is the caller's job. A field the provider cannot express at all is dropped rather than translated into a different parameter, and the provider validates whatever it receives against its own per-model list. | `provider` value | `seconds` maps to | `size` maps to | | ---------------- | ---------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `alibaba` | `parameters.duration` | `parameters.size` as `WIDTH*HEIGHT`. This matches the Wan 2.6 and earlier request protocol. Wan 2.7 models replaced `size` with `resolution` and `ratio` tiers; AISIX still forwards whatever you send, so omit `size` yourself on a Wan 2.7 alias and let the provider default apply. | | `zhipuai` | `duration` | `size` as `WIDTHxHEIGHT`, forwarded verbatim | | `volcengine` | `duration` | Not forwarded. Ark expresses output dimensions as `resolution` and `ratio` quality tiers, which cannot carry an arbitrary `WIDTHxHEIGHT`. AISIX validates the format and then drops the value, so the provider default applies. The submit response still echoes the `size` you sent, which does not mean Ark used it; poll responses omit it. | | `runwayml` | `duration` | `ratio` as `WIDTH:HEIGHT`. AISIX only swaps the separator; Runway validates the result against its per-model resolution list. | | `openai` | `seconds`, rendered as a string. OpenAI's create-video schema accepts `4`, `8`, or `12`. | `size` as `WIDTHxHEIGHT`, forwarded verbatim. The schema lists `720x1280`, `1280x720`, `1024x1792`, and `1792x1024`, but OpenAI publishes a narrower set per model, so check the model's own page. AISIX validates only the `WIDTHxHEIGHT` syntax. | ### Request Fields the Routes Do Not Model[​](#request-fields-the-routes-do-not-model "Direct link to Request Fields the Routes Do Not Model") The modeled routes cover text-to-video generation. Fields other than `model`, `prompt`, `seconds`, and `size` are ignored rather than rejected, so a request carrying them still generates a video — from the prompt alone. * `input_reference` is ignored. Image-to-video and video-to-video generation are not modeled on these routes. * Provider-native generation controls, such as a negative prompt, a seed, a reference image, or Wan 2.7's `resolution` and `ratio` tiers, have no unified field and are not forwarded. * `POST /v1/videos`, `GET /v1/videos/{video_id}`, and `GET /v1/videos/{video_id}/content` are the only modeled video routes AISIX serves. Provider routes that list, delete, remix, edit, or extend a video are not modeled, and neither is any provider-specific route beyond generation. Configure a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) when a caller needs any of these — it reaches the provider's native video API with the gateway still holding the credential and the caller key's access rules. ### Errors[​](#errors "Direct link to Errors") The table lists the statuses AISIX itself produces. A provider's own `4xx` keeps its status and is relayed in an upstream error envelope; a provider `5xx`, a transport failure, or a response AISIX cannot decode becomes `502`; an upstream timeout becomes `504`; a request body over the size limit becomes `413`. How a failed download surfaces depends on the delivery mode. On gateway-stream delivery, AISIX checks the provider's content endpoint before it starts the transfer, so a failure there is a JSON envelope; a stream cut after the response headers are sent is not, and the caller sees a short read against the declared `Content-Length` when the provider supplied one. On redirect delivery, AISIX never fetches the file: it returns the absolute `http` or `https` URL the provider reported, without verifying it, so anything that fails after the client follows the redirect comes from the provider or its CDN, in whatever format they use. | Status | When | | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `400` | The request is malformed: `seconds` is not a positive integer, or `size` is not `WIDTHxHEIGHT`. The content route also returns `400` when the task is still running, with a message telling the caller to keep polling, and when the task failed, with the provider's failure detail. | | `401` | The caller API key is missing or invalid. | | `403` | On submit, the caller API key is not allowed to use the model alias, or the request comes from a client IP outside the model's allowlist. The error type is `permission_denied`; an IP rejection also carries the `ip_restricted` code. | | `404` | On submit, the `model` names no alias the gateway knows; the error type is `model_not_found`. On the GET routes, the video ID is unknown or malformed, or it belongs to a model the caller key cannot access; the error type is `video_not_found`. The GET routes fold both an ACL denial and a not-implemented provider into `404`, so an ID cannot be used to probe which models exist. | | `422` | An input guardrail blocked the prompt. No provider task is created and no rate-limit capacity is consumed. | | `429` | A rate limit rejected the request, or an AISIX Cloud budget rejected it. Submissions count against both the caller key layers and the model's limits; status and content calls count against the caller key layers only. AISIX Cloud checks the caller's budget on all three routes, so an over-budget caller cannot poll or download a task it already paid to submit. | | `501` | The model alias resolves to a provider outside the video route's allowlist. The error type is `not_implemented`. | A task that the provider reports as failed is not an HTTP error on the status route: `GET /v1/videos/{video_id}` returns `200` with `"status": "failed"` and an `error` object carrying the provider's `code` and `message`. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now generated a video through the gateway's modeled video surface. For provider video APIs that AISIX has not modeled yet, configure a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) — on inject-mode routes, model-level rate limits apply there as well when the request body names a configured model. --- # AISIX Cloud Quickstart In this quickstart, you evaluate AISIX Cloud on one machine using Docker Compose, the bundled PostgreSQL database, and local endpoints. You start the control plane and dashboard, attach an AISIX gateway, configure an OpenAI model, and send a request through the gateway. The local setup runs the control-plane services, dashboard, and PostgreSQL as one management stack. The AISIX gateway runs separately, giving configuration and live traffic distinct paths: This quickstart configures resources through the AISIX Cloud Admin API so the workflow is reproducible and prepares the environment for subsequent guides. AISIX Cloud sends those resources to the gateway. Client requests then travel through the gateway to OpenAI without crossing the control plane, while the dashboard provides gateway status, logs, and usage. Licensing The AISIX Cloud control plane and dashboard are commercial software. Certain features are free for development, testing, and evaluation, but production use requires a commercial license. To run the control plane in production, [contact API7](https://api7.ai/contact) or email . ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Install [Docker](https://docs.docker.com/get-docker/) with Docker Compose V2. * Install [cURL](https://curl.se/), [jq](https://jqlang.github.io/jq/), `tar`, and OpenSSL. * Make sure the installation host can reach `run.api7.ai`, Docker Hub, and OpenAI. * Make sure ports `5432`, `8080`, `7944`, and `3000` are available on the installation host. * Use a browser that can reach port `8080` on the installation host. * Prepare an OpenAI API key for the model configured by this quickstart. ## Start the Control Plane[​](#step-1--start-aisix-cloud "Direct link to Start the Control Plane") On a host with Docker and internet access, run: ``` curl -fSL "https://run.api7.ai/aisix-self-hosted/aisix-self-hosted-1.0.0.tar.gz" \ -o aisix-self-hosted-1.0.0.tar.gz tar -xzf aisix-self-hosted-1.0.0.tar.gz ./aisix-self-hosted/run.sh ``` The commands download the AISIX 1.0.0 on-premises package into `./aisix-self-hosted`, generate a `.env` file with fresh secrets, pull the container images, and start the stack. The default dashboard URL is `http://localhost:8080`. For this single-host quickstart, open `./aisix-self-hosted/.env` and set the data-plane manager URL to: ``` AISIX_CLOUD_DPMGR_BASE_URL=https://host.docker.internal:7944 ``` Recreate the `api` and `dpm` Docker Compose services so the dashboard uses the updated endpoint and `dp-manager` issues its TLS certificate for it: ``` cd aisix-self-hosted docker compose up -d api dpm ``` Check the control-plane health endpoint: ``` curl -fsS "http://127.0.0.1:8080/healthz" ``` The command should return `{"status":"ok"}`. Keep this terminal in the `aisix-self-hosted` directory for the remaining shell commands. ## Create an Admin Token[​](#step-2--create-an-admin-token "Direct link to Create an Admin Token") The AISIX Cloud Admin API authenticates with an organization-scoped admin token. Create one in the dashboard: 1. Open `http://localhost:8080` in a browser and select **Create an account**. 2. Register the first user, accept the user agreement, and create your first organization. 3. In the organization navigation, open **Admin tokens** and select **New token**. 4. Enter `quickstart-admin` as the name, choose an expiration period, and enable the **write** scope. 5. Create the token and copy its plaintext value before leaving the page. The value is shown only once. Export the control-plane API URL and token in the terminal: ``` export AISIX_CP="http://localhost:8080/api" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" ``` For token scope, expiration, and rotation details, see [Admin Tokens](https://docs.api7.ai/ai-gateway/cloud/admin-tokens.md). ## Create Gateway Resources[​](#step-3--create-an-environment-a-model-and-a-caller-api-key "Direct link to Create Gateway Resources") Create the environment, provider key, model, and caller API key through the Admin API. The dashboard can create the same resources with the corresponding fields, but this quickstart uses API requests to provide one copyable workflow and retain the IDs used by later guides. Create the `prod` environment: ``` ENV_RESPONSE=$(curl -fsS -X POST "$AISIX_CP/environments" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{"display_name": "prod"}') export ENV_ID=$(echo "$ENV_RESPONSE" | jq -er '.environment.id') echo "$ENV_RESPONSE" | jq ``` Create a provider key that stores your OpenAI credential and allow it in the environment: ``` export OPENAI_API_KEY="YOUR_OPENAI_API_KEY" PROVIDER_KEY_RESPONSE=$(curl -fsS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{ "provider": "openai", "display_name": "OpenAI", "api_key": "'"${OPENAI_API_KEY}"'", "api_base": "https://api.openai.com/v1", "allowed_environments": ["'"${ENV_ID}"'"] }') export PROVIDER_KEY_ID=$(echo "$PROVIDER_KEY_RESPONSE" | jq -er '.provider_key.id') echo "$PROVIDER_KEY_RESPONSE" | jq ``` Create a `gpt-4o-mini` model backed by the provider key: ``` MODEL_RESPONSE=$(curl -fsS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{ "display_name": "gpt-4o-mini", "model_name": "gpt-4o-mini", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }') export MODEL_ID=$(echo "$MODEL_RESPONSE" | jq -er '.model.id') echo "$MODEL_RESPONSE" | jq ``` Create a caller API key that can use the model. Save the plaintext value because it is returned only once: ``` API_KEY_RESPONSE=$(curl -fsS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{ "display_name": "quickstart-caller", "allowed_models": ["'"${MODEL_ID}"'"] }') export API_KEY_ID=$(echo "$API_KEY_RESPONSE" | jq -er '.api_key.id') export AISIX_API_KEY=$(echo "$API_KEY_RESPONSE" | jq -er '.plaintext') echo "$API_KEY_RESPONSE" | jq ``` Each command saves the resource ID needed by the next step or by subsequent guides. If `curl` or `jq` reports an error, stop and correct it before continuing; a missing ID causes later commands to fail. ## Attach an AISIX Gateway[​](#step-4--attach-a-gateway "Direct link to Attach an AISIX Gateway") The control plane manages gateway configuration but does not serve AI traffic. Attach a gateway to the `prod` environment: 1. In the dashboard, open the `prod` environment, select **Data planes**, and then select **Issue certificate**. 2. Open the **Docker** tab and copy the generated snippet. The snippet contains a gateway certificate and private key, so handle it as a secret. 3. On Linux, add `--add-host host.docker.internal:host-gateway` to the generated `docker run` command. Docker Desktop resolves `host.docker.internal` automatically. 4. Run the snippet. It starts a container named `aisix-dp`, publishes the proxy on port `3000`, and follows the connection logs. When the logs show `etcd connected`, press **Ctrl+C**; the gateway continues running in the background. 5. Return to **Data planes**, refresh the page, and confirm that it reports one connected gateway instance. For certificate handling, generated deployment commands, networking, and connection troubleshooting, see [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md). ## Send and Verify a Request[​](#step-5--send-a-request "Direct link to Send and Verify a Request") Export the local gateway origin, then check that it is live: ``` export AISIX_PROXY="http://127.0.0.1:3000" ``` Check the proxy listener: ``` curl -fsS "$AISIX_PROXY/livez" ``` The command should return `ok`. Then confirm that the configured model reached the gateway: ``` curl -fsS "$AISIX_PROXY/v1/models" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` The `data` array should include `gpt-4o-mini`. Resource projection is asynchronous, so if the model is not listed yet, wait a few seconds and rerun the command. Do not continue until the model appears. Then send a chat request: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-4o-mini", "messages": [ {"role": "user", "content": "Say hello from AISIX AI Gateway."} ] }' ``` You should receive an OpenAI-compatible response with an assistant message under `choices[0].message`. In the dashboard, open **Logs** in the `prod` environment to inspect the request, and open **Usage** in the organization navigation to review its usage data. ## Clean Up[​](#clean-up "Direct link to Clean Up") Keep the example resources and admin token if you plan to continue to other AISIX Cloud guides. Otherwise, delete the resources in dependency order: ``` curl -fsS -X DELETE \ "$AISIX_CP/environments/$ENV_ID/api_keys/$API_KEY_ID" \ -H "Authorization: Bearer ${AISIX_TOKEN}" | jq curl -fsS -X DELETE \ "$AISIX_CP/environments/$ENV_ID/models/$MODEL_ID" \ -H "Authorization: Bearer ${AISIX_TOKEN}" | jq curl -fsS -X DELETE \ "$AISIX_CP/provider_keys/$PROVIDER_KEY_ID" \ -H "Authorization: Bearer ${AISIX_TOKEN}" | jq ``` Revoke `quickstart-admin` in the dashboard after the API cleanup if you no longer need it. Then remove the gateway container: ``` docker rm -f aisix-dp ``` Stop and remove the control-plane containers: ``` ./run.sh down ``` This preserves the PostgreSQL data volume. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now sent a request through a gateway connected to your local control plane. From here: * Follow [On-Premises Installation](https://docs.api7.ai/ai-gateway/on-premises/deployment.md) to choose an installation method and prepare a persistent environment. * Use the [AISIX Cloud Admin API Reference](https://docs.api7.ai/ai-gateway/reference/cloud-admin-api/.md) to automate control-plane operations with the admin token created in this quickstart. * Read [Resource Model](https://docs.api7.ai/ai-gateway/models/resource-model.md) to see how provider keys, models, and caller API keys fit together. * Call the gateway from application code with the [OpenAI SDK](https://docs.api7.ai/ai-gateway/getting-started/openai-sdk.md) or [Anthropic SDK](https://docs.api7.ai/ai-gateway/getting-started/anthropic-sdk.md) guide. --- # Anthropic SDK Point the Anthropic Python SDK at AISIX AI Gateway to keep the Anthropic Messages request format while the gateway manages caller authentication, model aliases, routing, and policy. The SDK needs only a gateway base URL, an AISIX caller API key, and a model alias available through `/v1/messages`. The alias can use Anthropic directly or another supported provider through AISIX translation. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running AISIX gateway that your application can reach. * A configured model alias available through `/v1/messages`. * A caller API key allowed to use that model alias. * Python 3.9 or later. If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. To configure an Anthropic-backed model alias for either product, follow [Anthropic](https://docs.api7.ai/ai-gateway/providers/anthropic.md). The SDK client configuration is the same for AISIX Cloud and the open-source AISIX gateway. ## Request Flow[​](#request-flow "Direct link to Request Flow") Keep the Anthropic SDK client, but send requests through the gateway instead of calling the upstream provider directly:
The application sends the caller API key and AISIX model alias. AISIX authorizes the caller, resolves the alias, and supplies the stored provider credential when it calls the upstream model. ## Configure the SDK[​](#configure-the-sdk "Direct link to Configure the SDK") Set the caller API key, model alias, and gateway base URL: ``` # AISIX_BASE_URL is the gateway origin without a trailing slash or /v1 # The local quickstarts use http://127.0.0.1:3000 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="claude-sonnet-prod" ``` The Anthropic SDK appends `/v1/messages` to the base URL. Do not include `/v1` in `AISIX_BASE_URL`. ### Install the Anthropic SDK[​](#install-the-anthropic-sdk "Direct link to Install the Anthropic SDK") Create and activate a Python virtual environment: ``` python3 -m venv .venv . .venv/bin/activate ``` Install the Anthropic SDK: ``` python -m pip install anthropic ``` ### Create a Client Example[​](#create-a-client-example "Direct link to Create a Client Example") Create the following client: anthropic-sdk-example.py ``` import os from anthropic import Anthropic client = Anthropic( api_key=os.environ["AISIX_API_KEY"], base_url=os.environ["AISIX_BASE_URL"], ) message = client.messages.create( model=os.environ["AISIX_MODEL"], max_tokens=128, messages=[{"role": "user", "content": "Say hello from AISIX."}], ) print(message.content[0].text) ``` Run the example: ``` python anthropic-sdk-example.py ``` You should see a short assistant response. The exact text depends on the upstream model. The SDK sends `POST /v1/messages` with the AISIX model alias. AISIX authenticates the caller API key, checks the model allowlist, resolves the alias, and returns an Anthropic-style message response. ## Compatibility Boundaries[​](#compatibility-boundaries "Direct link to Compatibility Boundaries") `POST /v1/messages` can resolve both Anthropic-backed and non-Anthropic-backed model aliases. Anthropic-backed aliases preserve Anthropic-specific request and response behavior most directly. Non-Anthropic translation is useful when you need a stable Anthropic-style client edge, but it is not feature-identical to native Anthropic behavior. If your application depends on tool-result round trips, thinking blocks, image blocks, or other Anthropic-specific content blocks, prefer an Anthropic-backed alias and validate the exact flow. For the full endpoint behavior, see [Anthropic-Style Messages API](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md). ## If the SDK Request Fails[​](#if-the-sdk-request-fails "Direct link to If the SDK Request Fails") First send a direct request to `/v1/messages` with the same caller API key and model alias. If the direct request succeeds, confirm that `AISIX_BASE_URL` points to the gateway root and does not end in `/v1`. If AISIX returns `404`, the requested model alias is not configured. If AISIX returns `403`, the caller API key exists but is not allowed to use that alias. For an upstream authentication or model error, verify the provider configuration rather than replacing the caller API key. ## Clean Up[​](#clean-up "Direct link to Clean Up") Remove the Python environment created here: ``` deactivate rm -rf .venv ``` The gateway resources this example uses were created outside this page. With AISIX Cloud, delete the caller API key, model alias, and provider key from the dashboard. With the open-source AISIX gateway, remove their entries from `resources.yaml` and reload. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now called AISIX from an Anthropic SDK client. Continue with [Anthropic-Style Messages API](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md) for endpoint behavior, [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md) for streamed responses, or [Anthropic](https://docs.api7.ai/ai-gateway/providers/anthropic.md) for upstream configuration. --- # Open-Source AISIX Gateway Quickstart Use this quickstart to run the open-source AISIX gateway in a single Docker container and send your first AI request through it. You will declare the required provider key, model, and caller API key in one `resources.yaml` file, start the gateway, and verify the request through its OpenAI-compatible API. This setup does not require a dashboard, control plane, or separate configuration store, making it the fastest way to evaluate the gateway locally. The example uses OpenAI as the upstream provider. Your client authenticates to AISIX with a caller API key, while the gateway uses a separate provider key to authenticate to OpenAI. The request follows this path: When the client requests the AISIX model name `gpt-4o-mini`, the gateway authenticates the request with the caller API key and uses the stored provider key to call OpenAI. The upstream OpenAI key is never exposed to the client. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Install [Docker](https://docs.docker.com/get-docker/) to run the AISIX AI Gateway container. * Install [cURL](https://curl.se/) to send requests to the gateway. * Prepare an OpenAI API key for an account with access to `gpt-4o-mini` and available quota. ## Create the Resources File[​](#create-the-resources-file "Direct link to Create the Resources File") First, create a working directory: ``` mkdir aisix-quickstart cd aisix-quickstart ``` Provider keys, models, and caller API keys are declared together in one `resources.yaml` file. Create it: resources.yaml ``` _format_version: "1" provider_keys: - display_name: openai-main provider: openai adapter: openai api_key: ${OPENAI_API_KEY} api_base: https://api.openai.com/v1 models: - display_name: gpt-4o-mini provider: openai model_name: gpt-4o-mini provider_key: openai-main api_keys: - display_name: quickstart-caller key_env: CALLER_API_KEY allowed_models: - gpt-4o-mini ``` ❶ `_format_version: "1"` is mandatory and must be a quoted string. It pins the file format so a future revision can never silently misread this file. ❷ `${OPENAI_API_KEY}` is resolved from the gateway's environment when the file loads, so the file itself never contains the upstream credential. A referenced variable that is unset or empty fails the load. ❸ `provider_key` references the provider key above by its `display_name`. A reference to an undefined name fails the load, so a typo can never become a silent runtime failure. ❹ `key_env` names an environment variable that holds the plaintext caller API key. The gateway hashes the value at load time and stores only the hash. Do not start the variable name with `AISIX_`, because that prefix is reserved for [startup-configuration overrides](https://docs.api7.ai/ai-gateway/reference/environment-variables.md). To supply a precomputed SHA-256 hash instead, set `key_hash` in place of `key_env`. The gateway derives stable IDs from resource names, so each name must be unique within its collection. ## Create the Startup Configuration[​](#create-the-startup-configuration "Direct link to Create the Startup Configuration") Create a `config.yaml` file that points the gateway at the resources file: config.yaml ``` resources_file: /etc/aisix/resources.yaml proxy: addr: "0.0.0.0:3000" admin: enabled: false ``` ❶ `resources_file` selects the file as the gateway's resource source. With this setting, the gateway does not use an external configuration store, so the `etcd` section is omitted. The two are mutually exclusive. ❷ `proxy.addr` listens for gateway traffic on port `3000` across all container interfaces. The Docker command below publishes this port at `http://127.0.0.1:3000` on the host. For the remaining startup options, see the [Startup Configuration Reference](https://docs.api7.ai/ai-gateway/reference/configuration-files.md). ## Start AISIX AI Gateway[​](#start-aisix-ai-gateway "Direct link to Start AISIX AI Gateway") Export the two values referenced by `resources.yaml` and the local gateway origin: ``` # Replace with your OpenAI API key. export OPENAI_API_KEY="YOUR_PROVIDER_API_KEY" # Choose a caller API key for client requests. export CALLER_API_KEY="YOUR_CALLER_API_KEY" # This quickstart publishes the gateway at this origin. export AISIX_PROXY="http://127.0.0.1:3000" ``` Before starting the gateway, validate `resources.yaml` in a short-lived container. This catches interpolation, reference, and schema errors without starting a listener: ``` docker run --rm \ -v "$(pwd):/etc/aisix:ro" \ -e OPENAI_API_KEY \ -e CALLER_API_KEY \ --entrypoint /usr/local/bin/aisix \ ghcr.io/api7/aisix:1.0.0 \ validate --resources /etc/aisix/resources.yaml ``` The command should report that the file loaded three resources. Then start the gateway, mounting the working directory so later edits remain visible inside the running container: ``` docker run -d --name aisix-quickstart \ -v "$(pwd):/etc/aisix:ro" \ -e OPENAI_API_KEY \ -e CALLER_API_KEY \ -p 3000:3000 -p 9090:9090 \ ghcr.io/api7/aisix:1.0.0 ``` If any resource entry is invalid, the container exits at startup. `docker logs aisix-quickstart` reports every invalid kind, entry, and field rather than stopping at the first error. ## Verify the Gateway[​](#verify-the-gateway "Direct link to Verify the Gateway") Check that the proxy listener is live: ``` curl -sS "$AISIX_PROXY/livez" ``` The command should return `ok`. List the models visible to the caller API key: ``` curl -sS "$AISIX_PROXY/v1/models" \ -H "Authorization: Bearer ${CALLER_API_KEY}" ``` The `data` array should include `gpt-4o-mini`. Send a chat request through the gateway: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${CALLER_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-4o-mini", "messages": [ {"role": "user", "content": "Say hello from AISIX AI Gateway."} ] }' ``` The response should use the OpenAI chat-completions format and contain an assistant message under `choices[0].message`. ## Update and Reload the Configuration[​](#update-and-reload-the-configuration "Direct link to Update and Reload the Configuration") The running gateway does not watch `resources.yaml`. To apply a change without restarting the container, edit the mounted file, validate it, and then send `SIGHUP` to the gateway process. For example, update `resources.yaml` to add a second model and allow the existing caller API key to use it: resources.yaml ``` _format_version: "1" provider_keys: - display_name: openai-main provider: openai adapter: openai api_key: ${OPENAI_API_KEY} api_base: https://api.openai.com/v1 models: - display_name: gpt-4o-mini provider: openai model_name: gpt-4o-mini provider_key: openai-main - display_name: gpt-4o provider: openai model_name: gpt-4o provider_key: openai-main api_keys: - display_name: quickstart-caller key_env: CALLER_API_KEY allowed_models: - gpt-4o-mini - gpt-4o ``` Validate the edited file inside the running container: ``` docker exec aisix-quickstart \ /usr/local/bin/aisix validate --resources /etc/aisix/resources.yaml ``` The command should report that the file loaded four resources. If validation fails, correct the reported entries before continuing. Reload the file: ``` docker kill --signal=HUP aisix-quickstart ``` Confirm that the caller can see both models: ``` curl -sS "$AISIX_PROXY/v1/models" \ -H "Authorization: Bearer ${CALLER_API_KEY}" ``` The `data` array should include both `gpt-4o-mini` and `gpt-4o`. For configuration states, rejected-resource details, and other resource sources, see [Configuration Propagation](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md). ## Clean Up[​](#clean-up "Direct link to Clean Up") Stop and remove the quickstart gateway when you are done: ``` docker rm -f aisix-quickstart ``` The `config.yaml` and `resources.yaml` files in your working directory are untouched. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now run an open-source AISIX gateway from one declarative file and sent provider-backed requests through it. From here: * Read [Resource Model](https://docs.api7.ai/ai-gateway/models/resource-model.md) to see how provider keys, models, and caller API keys fit together. * Check resource files in CI with the [CLI Reference](https://docs.api7.ai/ai-gateway/reference/cli.md). * Call the same gateway from application code with the [OpenAI SDK](https://docs.api7.ai/ai-gateway/getting-started/openai-sdk.md) or [Anthropic SDK](https://docs.api7.ai/ai-gateway/getting-started/anthropic-sdk.md) guide. * Review [Production Readiness](https://docs.api7.ai/ai-gateway/deployment/production.md) before deploying the gateway for production traffic. --- # OpenAI SDK In this guide, you will point the official OpenAI SDK at AISIX AI Gateway instead of sending requests directly to an upstream provider. This flow fits applications that already use the OpenAI-compatible chat-completions format and need to keep that client code mostly unchanged. The example authenticates to AISIX with a caller API key, sends requests to the AISIX proxy API root, uses an AISIX model alias, and receives OpenAI-compatible chat-completions responses. The upstream provider can still be changed behind the gateway when the configured model alias and provider support that request shape. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running AISIX gateway that your application can reach. * A configured model alias that accepts OpenAI-compatible chat-completions requests. * A caller API key allowed to use that model alias. * Install Node.js 20 LTS or newer with `npm`. If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure the SDK[​](#configure-the-sdk "Direct link to Configure the SDK") Create a small Node.js project, install the SDK, and point it at the AISIX proxy API root. ### Install the SDK[​](#install-the-sdk "Direct link to Install the SDK") Create a small demo project: ``` mkdir aisix-openai-demo && cd aisix-openai-demo npm init -y ``` Install the OpenAI SDK in the demo project: ``` npm install openai ``` Set the caller API key, model alias, and gateway base URL: ``` # AISIX_BASE_URL ends in /v1 and has no trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-mini" ``` ### Create the Chat Example[​](#create-the-chat-example "Direct link to Create the Chat Example") Use the `.mjs` extension so Node treats top-level `await` and `import` as ES modules without extra configuration. openai-sdk-example.mjs ``` import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.AISIX_API_KEY, baseURL: process.env.AISIX_BASE_URL, }); const response = await client.chat.completions.create({ model: process.env.AISIX_MODEL ?? "gpt-4o-mini", messages: [{ role: "user", content: "Say hello from AISIX." }], }); console.log(response.choices[0]?.message.content); ``` ### Run the Example[​](#run-the-example "Direct link to Run the Example") Run the chat example from the demo project: ``` node openai-sdk-example.mjs ``` You should see a short assistant response. The exact text depends on the upstream model. When the gateway can resolve `gpt-4o-mini` and the upstream provider is reachable, the SDK returns a standard OpenAI chat-completions object with an assistant message. AISIX resolves the model alias and injects the stored provider credential before calling the upstream provider. ## Streaming Responses[​](#streaming-responses "Direct link to Streaming Responses") The same `baseURL` works for streaming. openai-sdk-streaming.mjs ``` import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.AISIX_API_KEY, baseURL: process.env.AISIX_BASE_URL, maxRetries: 0, }); const stream = await client.chat.completions.create({ model: process.env.AISIX_MODEL ?? "gpt-4o-mini", messages: [{ role: "user", content: "Stream a short greeting." }], stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); } ``` Run the streaming example from the demo project: ``` node openai-sdk-streaming.mjs ``` You should see streamed text printed to the terminal. ## Production Setup Pattern[​](#production-setup-pattern "Direct link to Production Setup Pattern") In most deployments, application code needs only the gateway base URL, caller API key, and AISIX model alias. Upstream credentials, provider base URLs, upstream model IDs, routing policies, rate limits, guardrails, and observability hooks stay behind the gateway. This separation lets you rotate provider credentials, change upstream model IDs, or add gateway policy without changing the SDK call site. ## If the SDK Request Fails[​](#if-the-sdk-request-fails "Direct link to If the SDK Request Fails") First check the caller API key, gateway base URL, and AISIX model alias with a direct request to the gateway. Then use the same values in the SDK. If the SDK still sends traffic directly to OpenAI, check `baseURL`. It must point to the AISIX proxy API root, not to the upstream OpenAI API URL. If AISIX returns `404`, the requested model alias is not configured. If AISIX returns `403`, the caller API key exists but is not allowed to use that alias. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now called AISIX from an OpenAI SDK client. From here, review [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md) for proxy behavior. For SSE responses and tool definitions, see [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md) and [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md). If your application expects Claude-style requests, continue with [Anthropic SDK](https://docs.api7.ai/ai-gateway/getting-started/anthropic-sdk.md). --- # AISIX Products and Deployment Options AISIX is available as the open-source AISIX gateway or AISIX Cloud, API7's commercial offering. The open-source gateway runs standalone without a control plane. AISIX Cloud adds a control plane for managing AISIX gateways. You can host the control plane in your infrastructure, or API7 can host it. In every case, you run the AISIX gateway in your own environment. This page compares the open-source AISIX gateway with the two AISIX Cloud control-plane deployment options so you can choose the right product and deployment option. ## Compare Your Options[​](#compare-your-options "Direct link to Compare Your Options") | Product | Control-plane deployment option | Control-plane host | Gateway location | | ----------------------------- | ------------------------------- | ------------------ | ------------------- | | **Open-source AISIX gateway** | Not applicable | None | Your infrastructure | | **AISIX Cloud** | On-Premises | You | Your infrastructure | | **AISIX Cloud** | Hybrid Cloud | API7 | Your infrastructure | ## Compare Management Capabilities[​](#compare-management-capabilities "Direct link to Compare Management Capabilities") The AISIX gateway handles the same traffic-processing path in every option. AISIX Cloud adds shared management and governance around the gateways. | Area | Open-source AISIX gateway | AISIX Cloud | | --------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | | Resource management | Declarative `resources.yaml` file and automation you operate | Dashboard and AISIX Cloud Admin API with environment-scoped configuration delivery | | Organization and access | No organization or membership management layer; integrate your own access workflows | Organizations, environments, members, teams, and roles | | Credential lifecycle | Provider and caller credentials managed through your configuration and secret systems | Encrypted, write-only provider secrets and managed caller API key lifecycles | | Gateway connection and visibility | You deploy, secure, monitor, and upgrade gateway instances using your own tooling | You still operate the gateways; the control plane adds certificates, registration, configuration delivery, heartbeat, and status visibility | | Usage, cost, and governance | Export logs and metrics to systems you operate | Centralized request logs, usage and cost views, model pricing, budgets, and audit history | These control-plane capabilities do not place the control plane in the live AI traffic path. Applications call the AISIX gateway directly in every option. ## Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") The open-source AISIX gateway is available under the Apache License 2.0 and runs without a control plane. You declare providers, models, caller API keys, and policies in a [`resources.yaml` file](https://docs.api7.ai/ai-gateway/reference/resources-file.md) that the gateway loads at startup and reloads on `SIGHUP`. * **Choose this if** you want the open-source gateway, you will build or integrate your own management and automation, and you do not need AISIX Cloud features such as organizations, usage reporting, or budgets. * **You operate** the gateway process, its resources file, and upgrades. * **Start with** the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) to declare a `resources.yaml` file, start the gateway, and send a first request. ## AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") On-Premises and Hybrid Cloud are the two AISIX Cloud control-plane deployment options. They use the same dashboard, AISIX Cloud Admin API, resource model, and gateway workflow. The difference is who hosts the control plane and where its data is stored. Specific commercial capabilities can vary by deployment option and release. ### On-Premises[​](#on-premises "Direct link to On-Premises") You host the AISIX Cloud control plane in your own infrastructure, including fully air-gapped environments. You manage resources through the same dashboard or AISIX Cloud Admin API as Hybrid Cloud. * **Choose this if** data residency, network isolation, compliance, or air-gap requirements mean the control plane must also run in your environment. * **You operate** the complete private deployment, including the control plane and gateways. * **Start with** the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md) to bring up the control plane, create an environment, connect a gateway, and send a first request. ### Hybrid Cloud[​](#hybrid-cloud "Direct link to Hybrid Cloud") API7 hosts the AISIX Cloud control plane, while you deploy the gateway in your own environment and connect it to the control plane, which projects environment resources to it. You manage resources through the dashboard or the AISIX Cloud Admin API. * **Choose this if** you want centralized management, usage reporting, budgets, and governance without operating the control plane. Your AI traffic remains in your environment. * **You operate** the gateway; API7 operates the control plane. * **Get access:** Hybrid Cloud is not currently available through public self-service registration. [Contact API7](https://api7.ai/contact) to request a trial or demo. AISIX Cloud is also available for purchase through [AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-o7ltvkj4qjnr2). Marketplace changes procurement and billing, not the Hybrid Cloud architecture or operational responsibilities. After API7 provides control-plane access, continue with the [AISIX Cloud overview](https://docs.api7.ai/ai-gateway/cloud/overview.md) to create or select an environment, connect a gateway, configure a provider, and send a first request. ## Plan a Migration[​](#plan-a-migration "Direct link to Plan a Migration") The gateway runtime and the proxy API are the same across every choice. When you change how the gateway is managed, applications keep the same AISIX endpoint conventions and model aliases. The management path does change. The open-source gateway loads dynamic resources from a `resources.yaml` file, while a gateway connected to AISIX Cloud receives resources projected from the control plane. Plan to recreate gateway resources when moving to or from an AISIX Cloud control plane rather than treating it as an in-place switch of the startup configuration. --- # Integrations AISIX integrates with developer tools, application frameworks, AI application platforms, and voice agent platforms through direct client configuration or an operator-managed forward proxy. In both patterns, the client keeps its native request shape while AISIX applies access control, policy, and telemetry. Most integrations point the client at an AISIX API endpoint and replace the provider API key and model name with an AISIX caller key and model alias. This direct pattern applies to clients that support a custom endpoint for OpenAI-compatible Chat Completions, the OpenAI Responses API, or Anthropic-compatible traffic. Some tools must keep their official service endpoints and upstream credentials. For those tools, a TLS-terminating egress device can deliver selected traffic to a host-matched AISIX passthrough route. [Forward Proxy for IDE AI Traffic](https://docs.api7.ai/ai-gateway/deployment/forward-proxy.md) demonstrates this pattern with GitHub Copilot. Both patterns let teams adopt AISIX without rewriting application code or changing the coding tool itself. The client keeps using its familiar SDK, CLI, or framework while gateway policy and observability move into AISIX. Direct integrations also move routing and provider credentials behind AISIX model aliases. For new OpenAI-family integrations, prefer the client's current recommended API surface when it can still target a custom AISIX base URL. Use the Responses API when the framework provides a first-class Responses client. Use Chat Completions when the client only documents an OpenAI-compatible provider path, or when an existing workflow depends on Chat Completions compatibility. ## How Integrations Use AISIX[​](#how-integrations-use-aisix "Direct link to How Integrations Use AISIX") Direct integrations use the same three AISIX values: * The proxy API base URL. * A caller API key. * A model alias that the caller API key can access. Many clients describe the third value as a custom model, model ID, model name, or public model name. In AISIX, use the model alias for that field. AISIX authenticates the caller key, resolves the alias to an upstream provider configuration, applies configured policies, and records gateway telemetry. In the GitHub Copilot forward-proxy pattern, a passthrough route resolves traffic from the trusted egress device to a dedicated caller-key principal. In the primary `header_key` configuration, the device presents a gateway credential. An anonymous route can instead bind the principal and restrict traffic by source CIDR. The client keeps its official endpoint and upstream credential; the device can also attach a trusted employee identity for attribution. ## Integration Types[​](#integration-types "Direct link to Integration Types") | Client type | Use when | Guides | | ------------------------ | ----------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Client SDKs | Existing application code uses an OpenAI- or Anthropic-compatible client. | [OpenAI SDK](https://docs.api7.ai/ai-gateway/getting-started/openai-sdk.md), [Anthropic SDK](https://docs.api7.ai/ai-gateway/getting-started/anthropic-sdk.md) | | Coding agents | A local editor, AI CLI, remote development environment, or automation job should send requests through AISIX. | [Coding Agents](https://docs.api7.ai/ai-gateway/integrations/coding-agents.md) | | Frameworks and libraries | An application framework, agent framework, or AI library should call AISIX from application code. | [LangChain and LangGraph](https://docs.api7.ai/ai-gateway/integrations/frameworks/langchain.md), [LlamaIndex](https://docs.api7.ai/ai-gateway/integrations/frameworks/llamaindex.md), [Haystack](https://docs.api7.ai/ai-gateway/integrations/frameworks/haystack.md), [Vercel AI SDK](https://docs.api7.ai/ai-gateway/integrations/frameworks/vercel-ai-sdk.md), [Pydantic AI](https://docs.api7.ai/ai-gateway/integrations/frameworks/pydantic-ai.md), [Instructor](https://docs.api7.ai/ai-gateway/integrations/frameworks/instructor.md), [OpenAI Agents SDK](https://docs.api7.ai/ai-gateway/integrations/frameworks/openai-agents-sdk.md), [CrewAI](https://docs.api7.ai/ai-gateway/integrations/frameworks/crewai.md), [Microsoft Agent Framework](https://docs.api7.ai/ai-gateway/integrations/frameworks/microsoft-agent-framework.md) | | AI application platforms | A visual application, workflow, or chat platform should send its model requests through AISIX. | [AI Application Platforms](https://docs.api7.ai/ai-gateway/integrations/application-platforms.md) | | Voice agent platforms | A hosted platform or voice framework should use AISIX for the text LLM step while keeping its own audio and session pipeline. | [Voice Agent Platforms](https://docs.api7.ai/ai-gateway/integrations/voice-agents.md) | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Direct-client guides assume the gateway already has the resources needed for the client path: * A caller API key for the application, developer, or automation profile. * A model alias that the caller API key can access. * A proxy API route compatible with the client, such as OpenAI-compatible Chat Completions, the OpenAI Responses API, or Anthropic Messages. If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. The forward-proxy guide has different prerequisites, including a TLS-terminating egress device, permission to configure client proxy and certificate trust, and the current upstream service allowlist. Provider guides under [Models and Providers](https://docs.api7.ai/ai-gateway/providers/overview.md) cover upstream-specific configuration. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [OpenAI SDK](https://docs.api7.ai/ai-gateway/getting-started/openai-sdk.md): call an OpenAI-compatible gateway endpoint from application code. * [Anthropic SDK](https://docs.api7.ai/ai-gateway/getting-started/anthropic-sdk.md): call the Anthropic Messages endpoint from application code. * [Coding Agents](https://docs.api7.ai/ai-gateway/integrations/coding-agents.md): route AI coding tools through AISIX. * [AI Application Platforms](https://docs.api7.ai/ai-gateway/integrations/application-platforms.md): govern model requests from visual workflows and AI user interfaces. * [Voice Agent Platforms](https://docs.api7.ai/ai-gateway/integrations/voice-agents.md): govern the LLM step in realtime voice applications. * [Forward Proxy for IDE AI Traffic](https://docs.api7.ai/ai-gateway/deployment/forward-proxy.md): govern selected Copilot traffic while clients keep their official service endpoints. * [Supported Endpoints](https://docs.api7.ai/ai-gateway/endpoints/overview.md): review the proxy API families available to clients. --- # AI Application Platforms AI application platforms combine model access with user interfaces, visual workflows, agents, tools, retrieval, and application state. They let teams build and operate AI experiences without assembling every component directly in application code. When a platform accepts an OpenAI-compatible endpoint, its language-model requests can pass through AISIX for caller authentication, model aliases, routing, policy, and telemetry. The platform continues to run the application workflow and user experience, while AISIX governs the request path from the platform to the configured model provider. ## Choose a Platform Guide[​](#choose-a-platform-guide "Direct link to Choose a Platform Guide") Each platform exposes the AISIX connection through a different configuration surface: | Platform | Integration surface | Guide | | ---------- | ------------------------------------------- | ---------------------------------------------------------------------------------------------- | | Open WebUI | OpenAI-compatible connection | [Open WebUI](https://docs.api7.ai/ai-gateway/integrations/application-platforms/open-webui.md) | | n8n | OpenAI Chat Model sub-node | [n8n](https://docs.api7.ai/ai-gateway/integrations/application-platforms/n8n.md) | | Dify | OpenAI-API-compatible model provider plugin | [Dify](https://docs.api7.ai/ai-gateway/integrations/application-platforms/dify.md) | These platforms are AISIX clients, not upstream model providers. Configure provider credentials and upstream model IDs in AISIX. Give the platform an AISIX caller API key and model alias instead. ## Prepare AISIX[​](#prepare-aisix "Direct link to Prepare AISIX") Each platform needs these gateway-facing values: * An AISIX proxy URL the platform deployment can reach. * An AISIX caller API key dedicated to the platform or application. * A chat-capable model alias that the caller key can access through `POST /v1/chat/completions`. If AISIX is already deployed for your organization, obtain these values from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. Use a caller key scoped to only the aliases the platform needs. A shared platform connection normally identifies the platform or application to AISIX, not each person using its interface. Keep end-user authorization and audit context in the application platform unless you have separately designed and verified per-user identity forwarding. ## Understand the Boundary[​](#understand-the-boundary "Direct link to Understand the Boundary") AISIX receives model requests after the application platform has assembled prompts, retrieved context, or selected tools. AISIX does not replace the platform's workflow engine, user accounts, conversation state, knowledge base, vector store, or tool execution. Gateway policies apply to the supported content that reaches AISIX. A platform may perform additional model calls for titles, summaries, memory, retrieval, or agent planning. Confirm which calls use the configured AISIX connection before relying on gateway telemetry or limits as a complete record of the platform's activity. AISIX model discovery returns every alias the caller key can access, including aliases intended for non-chat endpoints. Select an alias configured for Chat Completions in each application platform. Tool calling also crosses both product boundaries. AISIX must preserve the model's OpenAI-compatible tool-call response, and the application platform must execute the tool and send the result in a follow-up request. Test a complete tool round trip instead of treating a successful text response as proof of tool compatibility. ## Verify an Integration[​](#verify-an-integration "Direct link to Verify an Integration") Use one short platform workflow to confirm the complete path: 1. The platform sends a prompt through the configured AISIX connection. 2. AISIX records `POST /v1/chat/completions` for the expected caller key and model alias. 3. The platform renders or processes the model response. After the first success, verify streaming, tools, retrieval, and background model calls that the application depends on. ## Related Guides[​](#related-guides "Direct link to Related Guides") Choose a platform guide above for setup. These guides explain the gateway behavior shared by the integrations: * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): review the gateway-facing request format. * [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md): understand streamed response behavior. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): verify complete tool-call loops. * [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md): limit platform traffic. --- # Dify [Dify](https://docs.dify.ai/) is an AI application platform for building chat applications, workflows, agents, retrieval pipelines, tools, and model-powered APIs. Its official OpenAI-API-compatible model provider plugin lets a workspace add a model served through a custom API endpoint. Configure that plugin with an AISIX proxy URL, caller API key, and model alias. Dify then uses the alias from application and workflow model nodes, while AISIX manages the upstream provider credential, routing, policy, and telemetry. Dify is the AISIX client in this integration. It is not an upstream model provider, and AISIX does not replace Dify's workflow engine, application state, knowledge bases, tools, or published application APIs. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Dify workspace where you can install plugins and configure model providers. * An AISIX proxy URL the Dify deployment can reach. * An AISIX caller API key dedicated to Dify or the application. * A chat-capable model alias the caller key can access through the [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). * Context-window and maximum-output-token limits supported by every eligible target behind the AISIX alias. ## Install the Model Provider Plugin[​](#install-the-model-provider-plugin "Direct link to Install the Model Provider Plugin") In the Dify Marketplace, find and install the official [OpenAI-API-compatible model provider plugin](https://marketplace.dify.ai/plugins/langgenius/openai_api_compatible) published by `langgenius`. The plugin is versioned independently from the Dify application. When testing or troubleshooting the integration, record both the Dify version and plugin version so a later plugin change is not mistaken for an AISIX behavior change. ## Add the AISIX Model[​](#add-the-aisix-model "Direct link to Add the AISIX Model") Open **Integrations → Model Provider** in the workspace, select **OpenAI-API-compatible**, and add a model with these values: | Field | Value | | --------------------------- | ------------------------------------------------------------------------------------ | | Model Type | **LLM** | | Model Name | The AISIX model alias, for example `support-agent-prod` | | API Key | The AISIX caller API key | | API Base URL | The AISIX proxy API root through `/v1`, for example `https://gateway.example.com/v1` | | model name for API endpoint | Leave empty when **Model Name** is already the AISIX alias | | Completion mode | **Chat** | | API Type | **Chat Completions API (/chat/completions)** | | Model context size | The lowest context limit across every eligible target behind the alias | | Upper bound for max tokens | The lowest maximum output across every eligible target behind the alias | | Function Call Type | **Not Support** for the initial text test | Save the model configuration. The plugin uses the API key as a bearer credential and sends the configured model name to the OpenAI-compatible endpoint. Do not add `/chat/completions` to **API Base URL**. ## Verify an Application Request[​](#verify-an-application-request "Direct link to Verify an Application Request") Create or open a Dify chat application or workflow, select the configured AISIX model, and run a preview with a short prompt. Confirm these results: * Dify displays the model response. * AISIX records a successful `POST /v1/chat/completions` request. * The recorded caller key and model match the Dify model-provider configuration. Dify applications may issue additional model calls for agent planning, question classification, parameter extraction, or other workflow nodes. Configure every model node that should use AISIX, and inspect the complete workflow before assuming all model traffic uses the same connection. ## Enable Tool Calls[​](#enable-tool-calls "Direct link to Enable Tool Calls") After the text request succeeds, edit the model-provider configuration and set **Function Call Type** to **Tool Call** if the Dify application uses agent tools. Set **Stream function calling** to **Support** only when the selected model alias and upstream model support streamed OpenAI-compatible tool calls. Run a complete agent tool test and confirm that Dify executes the expected tool, returns the tool result to the model, and produces a final answer. A successful text request does not verify these additional requests or the tool-call stream. If tool behavior is inconsistent, return **Function Call Type** to **Not Support** while you inspect the Dify plugin, AISIX, and upstream model logs. Do not label a model as tool-capable only because the configuration field is available. ## Troubleshoot Dify Requests[​](#troubleshoot-dify-requests "Direct link to Troubleshoot Dify Requests") | Symptom | Check | | ------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | | Model validation fails | Confirm **API Base URL** reaches AISIX and uses the `/v1` API root without `/chat/completions`. | | Request returns `401` | Confirm **API Key** contains the AISIX caller key, not an upstream provider key. | | Request returns `403` | Confirm the caller key can access the alias entered as **Model Name**. | | Request uses the wrong model | Leave **model name for API endpoint** empty, or set it to the same AISIX alias deliberately. | | Agent does not call a tool | Confirm **Function Call Type** is **Tool Call** and verify that the selected upstream model supports the Dify tool schema. | | Streamed tool calls fail but non-streaming tool calls succeed | Set **Stream function calling** to **Not Support**, then compare the plugin's streamed request and response with the AISIX tool-call contract. | ## Next Steps[​](#next-steps "Direct link to Next Steps") * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): review the endpoint Dify calls. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): verify the model and gateway tool-call contract. * [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md): inspect Dify model traffic in AISIX. --- # n8n [n8n](https://docs.n8n.io/) is a workflow automation platform with visual AI nodes for agents, chains, models, memory, tools, and data sources. Its OpenAI Chat Model sub-node can use a custom OpenAI-compatible endpoint instead of calling OpenAI directly. Connect that model sub-node to AISIX when an n8n AI Agent or chain should use gateway-managed credentials, model aliases, routing, policy, and telemetry. n8n continues to run the workflow and execute tools; AISIX governs the model requests emitted by the configured sub-node. This integration covers the **OpenAI Chat Model** sub-node. It does not claim compatibility for every operation in n8n's general OpenAI application node, which also calls provider-specific file, image, audio, assistant, and other APIs. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * An n8n project where you can create credentials and edit workflows. * An AISIX proxy URL the n8n deployment can reach. * An AISIX caller API key dedicated to the workflow or n8n project. * A chat-capable model alias the caller key can access through the [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). ## Create the OpenAI Credential[​](#create-the-openai-credential "Direct link to Create the OpenAI Credential") Create a credential for the AISIX connection from an OpenAI Chat Model sub-node: 1. Add or open an **OpenAI Chat Model** sub-node in the workflow. 2. Under **Credential to connect with**, create an **OpenAI** credential. 3. Set **API Key** to the AISIX caller API key. 4. Set **Base URL** to the AISIX proxy API root, including `/v1`, for example `https://gateway.example.com/v1`. 5. Leave **Organization ID** empty. 6. Save the credential. n8n validates the credential with `GET /v1/models`. AISIX authenticates the caller key and returns the aliases it can access. ## Configure the Chat Model[​](#configure-the-chat-model "Direct link to Configure the Chat Model") Select the AISIX connection and Chat Completions path: 1. Select the AISIX-backed OpenAI credential. 2. Under **Model**, select a chat-capable AISIX model alias returned from the gateway. If it does not appear, choose the model ID entry mode and enter the alias exactly. 3. Turn off **Use Responses API** so the sub-node sends `POST /v1/chat/completions`. 4. Connect the model output to the **Chat Model** input of the AI Agent or chain that should use AISIX. Set **Use Responses API** explicitly rather than relying on a version-dependent default. It must be off for the Chat Completions path documented here. Do not enable OpenAI-hosted built-in tools such as Web Search, File Search, or Code Interpreter for this configuration. Those tools belong to OpenAI's Responses service rather than the n8n workflow tool loop. The model value is the AISIX alias, not the upstream provider model ID. n8n sends prompts and tool definitions to AISIX, while the workflow continues to own node execution, memory, branching, retries, and data movement. ## Test the Model Connection[​](#test-the-model-connection "Direct link to Test the Model Connection") For a minimal interactive test, connect a **Chat Trigger** to an **AI Agent**, attach the configured **OpenAI Chat Model**, and run the workflow chat. To verify incremental output, set the Chat Trigger's response mode to **Streaming**. Send a short prompt and confirm these results: * The n8n chat displays the model response. * AISIX records a successful `POST /v1/chat/completions` request. * The recorded caller key and model match the n8n credential and selected alias. When the workflow runs from a webhook, schedule, or another trigger, verify that execution separately. Test and production executions may use different credentials or workflow versions. ## Verify an Agent Tool[​](#verify-an-agent-tool "Direct link to Verify an Agent Tool") Attach a deterministic n8n tool, such as a Calculator, to the AI Agent and ask a question that requires it. Confirm that the agent calls the tool, receives the result, and returns a final answer. One tool-using turn normally produces at least two model requests: the first returns the tool call, and the next includes the tool result. Account for that request fan-out when setting AISIX rate limits and reviewing usage. With AISIX Cloud, include it when setting budgets. ## Troubleshoot the Model[​](#troubleshoot-the-model "Direct link to Troubleshoot the Model") | Symptom | Check | | ----------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- | | Credential validation fails | Confirm **Base URL** ends in `/v1` and `GET /v1/models` accepts the caller key. | | Model list is empty | Confirm the caller key can access at least one AISIX alias, or enter the alias through the model ID mode. | | Request reaches `/v1/responses` | Turn off **Use Responses API** on the OpenAI Chat Model sub-node. | | Request returns `403` | Confirm the n8n credential's caller key can access the selected alias. | | Agent answers without using a tool | Confirm the tool is attached to the AI Agent and the selected upstream model supports OpenAI-compatible tool calling. | | General OpenAI node operation fails | Use this integration only for the OpenAI Chat Model sub-node unless the specific AISIX endpoint has been verified separately. | ## Next Steps[​](#next-steps "Direct link to Next Steps") * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): review the request path used by the model sub-node. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): verify agent tool definitions and results. * [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md): account for multi-request agent workflows with AISIX Cloud. --- # Open WebUI [Open WebUI](https://docs.openwebui.com/) is a web interface for chatting with models, managing conversations, attaching knowledge, and using tools. It connects to model services through protocol-based connections, including OpenAI-compatible APIs. An Open WebUI administrator can add AISIX as an OpenAI-compatible connection. Open WebUI then discovers the model aliases available to the AISIX caller key and sends chat requests through the gateway. AISIX governs the model request path, while Open WebUI continues to manage users, chats, knowledge, tools, and interface behavior. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * An Open WebUI account with administrator access. * An AISIX proxy URL the Open WebUI deployment can reach. * An AISIX caller API key dedicated to Open WebUI. * A chat-capable model alias the caller key can access through the [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). If Open WebUI runs in a container, an AISIX gateway on the container host is not reachable through `localhost`. Use a hostname that resolves from the Open WebUI container, such as the gateway service name on a shared container network. ## Add the AISIX Connection[​](#add-the-aisix-connection "Direct link to Add the AISIX Connection") Configure AISIX as a standard OpenAI-compatible connection: 1. In Open WebUI, open **Settings → Admin → Connections**. 2. Under **Manage OpenAI API Connections**, select **Add Connection**. 3. Set **URL** to the AISIX proxy API root, including `/v1`, for example `https://gateway.example.com/v1`. 4. Keep **Auth** set to **Bearer** and set **API Key** to the AISIX caller API key. 5. Keep **API Type** set to **Chat Completions** so Open WebUI sends requests to `/v1/chat/completions`. 6. Under **Advanced**, leave **Provider** set to **Default**. 7. Leave **Model IDs** empty so Open WebUI discovers the aliases that the caller key can access. 8. Select **Save**. Open WebUI requests `GET /v1/models` with the caller key. AISIX returns only the model aliases that key can use, so the Open WebUI model selector follows the gateway's access policy. If you want to expose only part of that returned list, add the selected AISIX aliases under **Model IDs**. This filter narrows what Open WebUI displays; it does not grant access that the AISIX caller key does not already have. ## Verify Chat and Streaming[​](#verify-chat-and-streaming "Direct link to Verify Chat and Streaming") Start a new chat, select a chat-capable AISIX model alias, and send a short prompt such as “Reply with one sentence about AI gateways.” Confirm these results: * Open WebUI displays the streamed response. * AISIX records a successful `POST /v1/chat/completions` request. * The recorded caller key and model match the Open WebUI connection and selected alias. Open WebUI may use models for background tasks in addition to the visible chat. If you configure separate models or endpoints for tasks, embeddings, image generation, speech, or reranking, verify those paths independently before treating all Open WebUI model traffic as governed by AISIX. ## Verify Tools Separately[​](#verify-tools-separately "Direct link to Verify Tools Separately") Open WebUI's native tool path expects OpenAI-compatible streamed tool-call fragments. Each fragment must retain its tool-call `index` so Open WebUI can assemble the function name and arguments before executing the tool. AISIX preserves the indexed tool-call stream on the Chat Completions path. The upstream model must also support the requested tool, and Open WebUI must enable and authorize it. Run one real tool-using turn and confirm these outcomes: * Open WebUI executes the expected tool. * AISIX records the model request that produces the tool call and the follow-up request containing the tool result. * The final answer appears in the chat. A successful plain-text chat does not prove that this multi-request tool loop works. ## Troubleshoot the Connection[​](#troubleshoot-the-connection "Direct link to Troubleshoot the Connection") | Symptom | Check | | ----------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | No models appear | Confirm the URL ends in `/v1`, the caller key is valid, and `GET /v1/models` returns the expected aliases. Add an alias to **Model IDs** only after confirming the request path. | | Connection returns `401` | Confirm **API Key** contains the AISIX caller key, not an upstream provider key. | | Chat returns `403` | Confirm the caller key can access the selected AISIX model alias. | | Chat returns `404` | Remove `/chat/completions` from **URL**. Open WebUI appends that path to the `/v1` API root. | | Container deployment cannot connect | Use a gateway hostname reachable from the Open WebUI container instead of `localhost`. | | A tool turn returns an empty reply | Verify the selected model emits indexed OpenAI-compatible tool-call stream fragments, then inspect the AISIX and Open WebUI logs for the tool request. | ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md): review gateway stream behavior. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): understand the OpenAI-compatible tool loop. * [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md): monitor requests from Open WebUI. --- # Coding Agents Coding agents can generate frequent model and tool requests or call official AI services from development environments. Route that traffic through AISIX when platform teams need governed access, policy, and telemetry for developer AI tools. These guides cover coding tools that connect directly to AISIX and tools that retain their official service endpoints behind an operator-managed forward proxy. The client may run in a local editor, CLI, remote development environment, or automation job. ## Why Route Coding Agents Through AISIX[​](#why-route-coding-agents-through-aisix "Direct link to Why Route Coding Agents Through AISIX") A common rollout is for a platform team to make AISIX the approved gateway for developer AI tools. For direct integrations, developers configure Codex, Claude Code, Cline, Cursor, or another client with an AISIX proxy URL, caller API key, and custom model value instead of raw provider credentials. The custom model value is the AISIX model alias. For a tool such as GitHub Copilot that keeps its official service endpoints, an operator-managed egress device can terminate TLS and deliver selected traffic to an AISIX passthrough route. The route resolves a dedicated caller-key principal. In the primary `header_key` configuration, the device presents a gateway credential; an anonymous route can instead bind the principal and restrict traffic by source CIDR. The client keeps its upstream credential, and the device can attach a trusted employee identity when configured. In both patterns, the coding agent continues to use its native configuration and request format. AISIX enforces access, policy, and telemetry before relaying the selected traffic. Direct model and MCP paths use model or tool authorization. The forward-proxy path uses its caller-key principal and route grant; employee identity is recorded only when a trusted identity header is configured. This setup is useful when teams want to: * keep provider credentials out of directly configured editor and CLI clients. * authenticate each developer, project, automation job, or shared tool profile with a caller API key, or bind forward-proxy traffic to a dedicated caller-key principal. * use model aliases to control which upstream models a directly configured client can reach. * enforce request limits and guardrails, record usage across supported coding-agent paths, and apply matching AISIX Cloud budgets to model calls, MCP tool calls, and passthrough requests attributed to a caller key. * change the upstream provider or model behind an alias without asking users of direct integrations to reconfigure their tools. ## Sensitive Code and Credentials[​](#sensitive-code-and-credentials "Direct link to Sensitive Code and Credentials") Coding agents may send source code, configuration snippets, stack traces, or terminal output as part of a task. That context can contain API keys, customer data, personal information, or internal identifiers. Because AISIX sits between the coding agent and the upstream model, MCP server, or official service, teams can inspect and control that traffic at the gateway layer. On routes that support redaction write-back, use [PII Detection and Redaction](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/pii.md) when sensitive values should be masked or blocked. Passthrough routes do not rewrite provider-native bodies, so configure a block action when matched content must not leave the network. Use [Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/overview.md) for broader request and response policy checks. ## Choose a Client Guide[​](#choose-a-client-guide "Direct link to Choose a Client Guide") Start with the guide for the coding tool you want to route through AISIX: | Client | Primary AISIX path | Guide | | --------------- | ------------------------------------------------------ | ----------------------------------------------------------------------------------------------- | | Codex | OpenAI Responses API | [Codex](https://docs.api7.ai/ai-gateway/integrations/coding-agents/codex.md) | | Claude Code | Anthropic Messages | [Claude Code](https://docs.api7.ai/ai-gateway/integrations/coding-agents/claude-code.md) | | Cline | OpenAI-compatible API | [Cline](https://docs.api7.ai/ai-gateway/integrations/coding-agents/cline.md) | | Cursor Ask mode | OpenAI-compatible API | [Cursor](https://docs.api7.ai/ai-gateway/integrations/coding-agents/cursor.md) | | GitHub Copilot | Host-matched passthrough route through an egress proxy | [Forward Proxy for IDE AI Traffic](https://docs.api7.ai/ai-gateway/deployment/forward-proxy.md) | The direct-client guides assume that the AISIX gateway already has a model alias and caller API key for the route the client will use. In the client UI or config file, that alias may appear as a custom model, model ID, or model name. The GitHub Copilot guide lists its forward-proxy prerequisites separately. For a direct integration, obtain the gateway URL, model alias, and caller API key from the team that manages AISIX. If AISIX is not yet deployed, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. The forward-proxy guide lists its egress, network, and route prerequisites separately. ## How Coding Agents Use AISIX[​](#how-coding-agents-use-aisix "Direct link to How Coding Agents Use AISIX") Coding agents connect to AISIX through direct model and tool APIs or through a trusted forward proxy: | Client path | AISIX route | Use when | | ---------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | | Model requests | OpenAI-compatible routes for Chat Completions or Responses API, or Anthropic Messages | The agent should call models through an AISIX caller API key and model alias. | | Tool requests | MCP Gateway at `/mcp` | The agent should discover and call upstream MCP tools through AISIX tool access control. | | Forward-proxy requests | Host-matched passthrough route | The tool must keep its official service endpoint while an egress device sends selected traffic through AISIX. | All three paths use AISIX as the boundary between the coding tool and the upstream service. Each client keeps its native request format. Direct integrations use an AISIX caller key and model or tool authorization. A forward-proxy integration resolves a caller-key principal from a device credential or source-restricted anonymous binding, then applies the route grant. Traffic controls and guardrails apply according to the selected route and caller-key principal. Observability can additionally include an employee identity supplied by the trusted device. The direct model and tool paths look like this: For MCP-specific governance, see [MCP Gateway Overview](https://docs.api7.ai/ai-gateway/mcp-gateway/overview.md). For the available proxy API families, see [Supported Endpoints](https://docs.api7.ai/ai-gateway/endpoints/overview.md). ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Codex](https://docs.api7.ai/ai-gateway/integrations/coding-agents/codex.md): route Codex Responses API traffic through AISIX. * [Claude Code](https://docs.api7.ai/ai-gateway/integrations/coding-agents/claude-code.md): route Claude Code Anthropic Messages traffic through AISIX. * [Cline](https://docs.api7.ai/ai-gateway/integrations/coding-agents/cline.md): route Cline OpenAI-compatible traffic through AISIX. * [Cursor](https://docs.api7.ai/ai-gateway/integrations/coding-agents/cursor.md): route Cursor Ask mode traffic through AISIX. * [Forward Proxy for IDE AI Traffic](https://docs.api7.ai/ai-gateway/deployment/forward-proxy.md): route selected Copilot traffic through a trusted egress device and AISIX. * [MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/overview.md): connect coding agents to upstream MCP tools through AISIX. * [Supported Endpoints](https://docs.api7.ai/ai-gateway/endpoints/overview.md): review the caller-facing API families AISIX exposes. --- # Claude Code Claude Code can read endpoint and authentication settings from environment variables or settings files. Use those settings when Claude Code should call AISIX as an Anthropic-compatible gateway instead of calling Anthropic directly. In this guide, you will point Claude Code at AISIX and authenticate with an AISIX caller API key. You will also select an AISIX model alias that is available on the Anthropic Messages route. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * Claude Code installed. * A running AISIX gateway that Claude Code can reach. * An AISIX caller API key. * An AISIX model alias that the caller API key can access through [Anthropic Messages](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md). Use an alias backed by a native Anthropic-protocol route to preserve Anthropic-specific request and response behavior. If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure Claude Code[​](#configure-claude-code "Direct link to Configure Claude Code") Set the endpoint, caller key, and model alias before launching Claude Code: ``` # ANTHROPIC_BASE_URL is the gateway origin without a trailing slash or /v1 # The local quickstarts use http://127.0.0.1:3000 export ANTHROPIC_BASE_URL="YOUR_AISIX_GATEWAY_URL" export ANTHROPIC_AUTH_TOKEN="YOUR_CALLER_API_KEY" export ANTHROPIC_MODEL="claude-sonnet-prod" export ANTHROPIC_DEFAULT_HAIKU_MODEL="claude-sonnet-prod" claude ``` `ANTHROPIC_BASE_URL` points Claude Code at AISIX. Claude Code sends `ANTHROPIC_AUTH_TOKEN` as a bearer token, so AISIX receives the value as the caller API key. `ANTHROPIC_MODEL` and `ANTHROPIC_DEFAULT_HAIKU_MODEL` should be AISIX model aliases, not upstream provider model IDs. You can point both variables to the same alias, or use a separate fast alias for Claude Code background behavior. To apply the same settings every time Claude Code starts, put them in the `env` block of a Claude Code settings file: \~/.claude/settings.json ``` { "model": "claude-sonnet-prod", "env": { "ANTHROPIC_BASE_URL": "YOUR_AISIX_GATEWAY_URL", "ANTHROPIC_AUTH_TOKEN": "YOUR_CALLER_API_KEY", "ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-sonnet-prod" } } ``` ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Start Claude Code and send a short prompt. When the request succeeds, verify the following results: * Claude Code prints a model response. * AISIX records a successful `POST /v1/messages` request for the selected model alias. AISIX authenticates the caller API key, checks model access, resolves the model alias, applies configured policies, and dispatches to the upstream provider behind the alias. Verify the request in the AISIX gateway logs. You can also use configured metrics or upstream provider logs. ## Troubleshoot Claude Code Requests[​](#troubleshoot-claude-code-requests "Direct link to Troubleshoot Claude Code Requests") Claude Code can send Anthropic-specific fields and beta headers. Use an AISIX alias backed by a native Anthropic-protocol route for direct protocol compatibility, and validate the Claude Code workflow your team plans to deploy. If Claude Code reports an endpoint, authentication, or model error, check these items: | Symptom | Check | | -------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | Authentication fails | Confirm `ANTHROPIC_AUTH_TOKEN` is the AISIX caller API key and that the key can access the selected model alias. | | Model is not found | Confirm `ANTHROPIC_MODEL` and `ANTHROPIC_DEFAULT_HAIKU_MODEL` are AISIX model aliases available to the caller API key. | | Requests reach AISIX but fail upstream | Confirm the alias points to a compatible upstream model and provider key. | | Tool or beta behavior fails | Validate whether the Anthropic-specific feature is supported by the selected AISIX route and upstream provider. | For endpoint behavior, see [Anthropic Messages](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md). ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Anthropic Messages](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md): review route behavior and compatibility boundaries. * [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md): confirm gateway-side request metrics and logs. * [Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/overview.md): add request or response policy checks. --- # Cline Cline supports OpenAI-compatible providers with a custom base URL, API key, and model ID. Use that provider type when Cline should send model requests through AISIX. In this guide, you will configure Cline to call the AISIX OpenAI-compatible proxy API with an AISIX caller API key. The Cline model ID will be an AISIX model alias. This configuration keeps the request endpoint, API key, and model alias under AISIX control. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * Cline installed in your editor. * A running AISIX gateway with the proxy listener available. * An AISIX caller API key. * A model alias the caller API key can access through the [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure Cline[​](#configure-cline "Direct link to Configure Cline") Open Cline settings and choose the OpenAI-compatible provider. Then set these values: | Cline setting | AISIX value | Description | | ------------- | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | | API provider | OpenAI Compatible | Uses Cline's configurable OpenAI-compatible client. | | Base URL | `YOUR_AISIX_GATEWAY_URL/v1` | AISIX proxy API root, including `/v1` and without a trailing slash. For a local quickstart deployment, use `http://127.0.0.1:3000/v1`. | | API key | `YOUR_CALLER_API_KEY` | AISIX caller API key. | | Model ID | `gpt-4o-prod` | AISIX model alias, not the upstream provider model ID. Replace the example with an alias the caller API key can access. | Cline also documents OpenAI-specific provider options. Select `OpenAI Compatible` for gateway-routed traffic. OpenAI account and OAuth-based provider paths can bypass AISIX policy. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Use Cline's provider verification action to confirm the endpoint, key, and model alias. Then send a short prompt from Cline. When the request succeeds, verify the following results: * Cline prints a model response. * AISIX records a successful `POST /v1/chat/completions` request for the selected model alias. AISIX authenticates the caller API key, checks model access, resolves the model alias, applies configured policies, and dispatches to the upstream provider behind the alias. If the request fails, check these items: | Symptom | Check | | ------------------------------------- | ------------------------------------------------------------------------------------- | | Authentication fails | Confirm the API key in Cline is an AISIX caller API key. | | Model is not found | Confirm the model ID in Cline is an AISIX model alias visible to that caller API key. | | Cline reports route or payload errors | Confirm the selected model alias supports OpenAI-compatible chat requests. | | Requests bypass AISIX | Confirm the Base URL points to the AISIX proxy API root and includes `/v1`. | ## Add Gateway Policy[​](#add-gateway-policy "Direct link to Add Gateway Policy") After Cline traffic reaches AISIX, apply the same controls you use for other OpenAI-compatible clients: * Use [Caller API Keys](https://docs.api7.ai/ai-gateway/traffic-controls/caller-api-keys.md) to isolate developers, teams, projects, or automation profiles. * Use [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md) to control high-volume coding-agent traffic. AISIX Cloud deployments can also enforce [budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md). * Use [Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/overview.md) when prompts or responses need content checks. * Use [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md) to inspect request volume, latency, and status. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): review gateway-facing request behavior. * [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md): confirm streaming behavior if your Cline workflow depends on streamed output. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): review OpenAI-compatible tool-call behavior through AISIX. --- # Codex Codex supports custom model providers in its configuration file. Use a custom provider when Codex should send Responses API requests to AISIX instead of calling a model provider directly. In this guide, you will point Codex at the AISIX proxy API root and authenticate with an AISIX caller API key. The Codex custom model value will be an AISIX model alias. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * Codex installed and able to read `~/.codex/config.toml`. * A running AISIX gateway with the proxy listener available. * An AISIX caller API key. * A model alias that the caller API key can access and that works with the [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure Codex[​](#configure-codex "Direct link to Configure Codex") Set the caller API key in the environment where Codex runs: ``` # Replace with your values export AISIX_API_KEY="YOUR_CALLER_API_KEY" ``` Add an AISIX model provider to `~/.codex/config.toml`: \~/.codex/config.toml ``` model = "gpt-4o-prod" # Use an AISIX model alias, not the upstream provider model ID. model_provider = "aisix" [model_providers.aisix] name = "AISIX AI Gateway" # Include /v1 and omit the trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 base_url = "YOUR_AISIX_GATEWAY_URL/v1" env_key = "AISIX_API_KEY" # Read the AISIX caller API key from the environment. wire_api = "responses" ``` ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Start Codex from the shell where `AISIX_API_KEY` is set: ``` codex ``` Ask a short question that requires a model response. When the request succeeds, verify the following results: * Codex prints a model response. * AISIX records a successful `POST /v1/responses` request for the selected model alias. AISIX authenticates the caller API key, checks model access, resolves the model alias, applies configured policies, and dispatches to the upstream provider behind the alias. ## Troubleshoot Codex Requests[​](#troubleshoot-codex-requests "Direct link to Troubleshoot Codex Requests") If Codex cannot reach AISIX, check these items: | Symptom | Check | | ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | | Authentication fails | Confirm `AISIX_API_KEY` is set in the same shell or environment where Codex runs. | | Model is not found | Confirm `model` in `config.toml` is an AISIX model alias visible to the caller API key. | | Responses request fails | Confirm the selected model alias works with the AISIX Responses API. Some provider-backed aliases do not support every OpenAI-specific Responses feature. | | Requests bypass AISIX | Confirm `model_provider = "aisix"` is in the active Codex configuration layer. | For endpoint behavior and provider boundaries, see [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md): review how AISIX handles Responses API requests. * [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md): confirm gateway-side request metrics and logs. * [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md): add caller or model limits for coding-agent traffic. --- # Cursor Cursor supports custom API keys, an OpenAI base URL override, and custom model names for standard chat models. Configure these settings to send Cursor Ask mode requests through AISIX. In this guide, you will configure Cursor with an AISIX caller API key and use an AISIX model alias as the custom model name. AISIX then authenticates the caller, applies gateway policies, and forwards requests to the upstream model configured for the alias. info This configuration applies to Ask mode. Agent mode, nested agents, background agents, Tab Completion, inline edits, and other features that use Cursor's specialized models follow Cursor-managed paths instead of the AISIX base URL. To route Agent-mode MCP tool calls through AISIX, see [Connect Cursor to MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/cursor.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * Cursor Pro or a higher plan. The free plan saves the API key, base URL, and custom model, but requires an upgrade before the custom model can be selected in Chat. * A running AISIX gateway with a proxy URL that Cursor can reach. * An AISIX caller API key. * A model alias the caller API key can access through the [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. Set the gateway origin and caller API key, then confirm that AISIX exposes the alias before configuring Cursor: ``` # AISIX_PROXY has no trailing slash or endpoint path # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_URL" export AISIX_API_KEY="YOUR_CALLER_API_KEY" curl -sS "${AISIX_PROXY}/v1/models" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` ## Configure Cursor[​](#configure-cursor "Direct link to Configure Cursor") 1. Open **Cursor Settings** and select **Models**. 2. In the **OpenAI API Key** field, enter your AISIX caller API key. 3. In **Override OpenAI Base URL**, enter the AISIX proxy API root, including `/v1`. For example: ``` YOUR_AISIX_GATEWAY_URL/v1 ``` 4. In **Add or search model**, enter the AISIX model alias and press **Enter**. For example, enter `aisix_cursor`. 5. Enable the custom model. 6. Open Cursor Chat, select Ask mode, and select the custom model. Use a neutral model alias that does not contain or resemble a provider model ID. For example, use `aisix_cursor` rather than `gpt-4o-prod`. Cursor can otherwise interpret the alias as a built-in model and route or format the request differently. The custom model name sent by Cursor must match the AISIX model alias, not the upstream provider's model name. The resulting configuration maps Cursor settings to AISIX as follows: | Cursor setting | AISIX value | | ------------------------ | -------------------------- | | OpenAI API Key | AISIX caller API key | | Override OpenAI Base URL | AISIX proxy URL with `/v1` | | Custom model | AISIX model alias | ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Open Ask mode with `Cmd+L` on macOS or `Ctrl+L` on Windows and Linux. Select the custom model, then send a short prompt. For example, ask the model to respond with a specific word or short phrase. When the request succeeds, verify the following results: * Cursor Chat displays the expected response. * AISIX records a successful `POST /v1/chat/completions` request. * The AISIX request record identifies the expected caller API key and model alias. AISIX resolves the alias to the configured upstream provider and model. Cursor does not need the upstream provider credential or upstream model name. ## Troubleshoot Cursor Requests[​](#troubleshoot-cursor-requests "Direct link to Troubleshoot Cursor Requests") If the request fails, check these items: | Symptom | Check | | -------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Cursor prompts you to upgrade when you open the model selector | Confirm that you are using Cursor Pro or a higher plan. The free plan saves the configuration but does not let you select the custom model. | | Authentication fails | Confirm the value in **OpenAI API Key** is an AISIX caller API key, not the upstream provider key. | | Model is not found | Confirm the custom model name exactly matches an AISIX model alias available to the caller. | | Cursor reports a connection or endpoint error | Confirm the base URL is reachable from Cursor, uses HTTPS when required by your environment, and ends in `/v1`. | | Requests use an unexpected model or payload | Use a neutral alias that does not contain a provider model ID, then explicitly select it in Ask mode. | | Built-in models return errors | Clear **Override OpenAI Base URL** when switching back to built-in models. The override applies globally to OpenAI-compatible requests. | | Requests do not appear in AISIX logs | Confirm **Override OpenAI Base URL** is configured and the custom model is selected in Ask mode. Agent mode, Tab Completion, and other specialized features do not use this path. | ## Next Steps[​](#next-steps "Direct link to Next Steps") * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): review the gateway-facing request format. * [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md): confirm gateway-side request metrics and logs. * [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md): add caller or model limits for Cursor chat traffic. * [Cursor API Keys](https://docs.cursor.com/settings/api-keys): review which Cursor models and features support custom API keys. --- # CrewAI [CrewAI](https://docs.crewai.com/) is a framework for building multi-agent systems with agents, tasks, crews, flows, tools, and memory. AISIX fits at the model request boundary, while CrewAI continues to own agent coordination and task execution. CrewAI's documented gateway-friendly path is its `LLM` configuration with a custom `base_url`. Use that setting when a CrewAI application should send OpenAI-compatible chat requests through AISIX. This guide uses CrewAI with a custom `LLM`. You will configure the LLM with an AISIX proxy URL, caller API key, and model alias. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Python environment supported by [CrewAI](https://docs.crewai.com/). * A running AISIX gateway that the application can reach. * An AISIX caller API key. * A model alias the caller API key can access through the [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure CrewAI[​](#configure-crewai "Direct link to Configure CrewAI") Install CrewAI if the application does not already include it: ``` pip install crewai ``` Set the values that the CrewAI application will use: ``` # AISIX_BASE_URL ends in /v1 and has no trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` Create a CrewAI `LLM` with the AISIX values and pass it to agents: ``` import os from crewai import Agent, Crew, LLM, Task llm = LLM( model=os.environ["AISIX_MODEL"], provider="openai", api_key=os.environ["AISIX_API_KEY"], base_url=os.environ["AISIX_BASE_URL"], ) writer = Agent( role="Writer", goal="Write concise explanations.", backstory="You explain technical systems clearly.", llm=llm, ) task = Task( description="Write one sentence about governed multi-agent systems.", expected_output="One concise sentence.", agent=writer, ) crew = Crew( agents=[writer], tasks=[task], ) print(crew.kickoff()) ``` The model value is the AISIX model alias, not the upstream provider model ID. The `provider` value tells CrewAI to use its OpenAI-compatible client for the custom endpoint. AISIX governs the model request path, while CrewAI still owns agents, tasks, crews, flows, memory, and tool orchestration. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Run the script from the shell where the AISIX environment variables are set. When the request succeeds, verify the following results: * The script prints the crew output. * AISIX records one or more successful `POST /v1/chat/completions` requests for the selected model alias. Verify the request in the AISIX gateway logs. You can also use configured metrics or upstream provider logs. If the request fails, first confirm that the caller API key can access the selected model alias and that `base_url` points to the AISIX proxy API root with `/v1`. Multi-agent workflows can issue multiple model calls for planning, delegation, tool use, memory, and task execution. Configure every agent or crew-level LLM that should use AISIX, and account for the additional requests when configuring rate limits and reviewing usage. With AISIX Cloud, include that request fan-out when setting budgets. CrewAI does not currently document a stable native OpenAI Responses API configuration equivalent to the `base_url` path shown here. Keep the OpenAI-compatible route for AISIX unless your CrewAI version exposes and documents a Responses API client path. If it does, validate the exact agent, tool, and memory workflow before switching production traffic. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): review gateway-facing request behavior. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): confirm tool-call behavior before enabling CrewAI tools. * [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md): account for multi-agent request fan-out with AISIX Cloud spend controls. --- # Haystack [Haystack](https://docs.haystack.deepset.ai/) is a framework for building LLM applications with pipelines, retrieval, indexing, document processing, agents, and generators. AISIX fits at the model request boundary, while Haystack continues to own pipeline execution and retrieval. The Haystack `OpenAIResponsesChatGenerator` component supports custom OpenAI-compatible deployments through `api_base_url`. Use that setting when a Haystack pipeline should send Responses API requests through AISIX. This guide uses Haystack with `OpenAIResponsesChatGenerator`. You will configure the generator with an AISIX proxy URL, caller API key, and model alias. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Python environment supported by [Haystack](https://docs.haystack.deepset.ai/docs/installation). * A running AISIX gateway that the application can reach. * An AISIX caller API key. * A model alias the caller API key can access through the [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure Haystack[​](#configure-haystack "Direct link to Configure Haystack") Install Haystack if the application does not already include it: ``` pip install haystack-ai ``` Set the values that the Haystack application will use: ``` # AISIX_BASE_URL ends in /v1 and has no trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` Create an `OpenAIResponsesChatGenerator` with the AISIX values: ``` import os from haystack.components.generators.chat import OpenAIResponsesChatGenerator from haystack.dataclasses import ChatMessage from haystack.utils import Secret generator = OpenAIResponsesChatGenerator( model=os.environ["AISIX_MODEL"], api_base_url=os.environ["AISIX_BASE_URL"], api_key=Secret.from_token(os.environ["AISIX_API_KEY"]), ) response = generator.run([ ChatMessage.from_user("Write one sentence about retrieval pipelines."), ]) print(response["replies"][0].text) ``` The model value is the AISIX model alias, not the upstream provider model ID. AISIX governs the model request path, while Haystack still owns pipeline components, prompt builders, retrievers, ranking components, and response handling. Haystack also provides [`OpenAIChatGenerator`](https://docs.haystack.deepset.ai/docs/openaichatgenerator) for Chat Completions. Use that component when a pipeline needs broad OpenAI-compatible chat support instead of Responses-specific features. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Run the script from the shell where the AISIX environment variables are set. When the request succeeds, verify the following results: * The script prints a chat response. * AISIX records a successful `POST /v1/responses` request for the selected model alias. Verify the request in the AISIX gateway logs. You can also use configured metrics or upstream provider logs. If the request fails, first confirm that the caller API key can access the selected model alias and that `api_base_url` points to the AISIX proxy API root with `/v1`. Haystack applications often combine generators with retrievers, ranking components, tools, and structured output. Validate the exact pipeline before relying on gateway policy or telemetry for streaming, tool calls, or structured output. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md): review gateway-facing Responses behavior. * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): use Chat Completions when a workflow needs broad OpenAI-compatible support. * [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md): configure embedding traffic if your Haystack workflow uses AISIX for embeddings. * [Rerank](https://docs.api7.ai/ai-gateway/endpoints/rerank.md): configure rerank traffic if your Haystack pipeline uses rerank models through AISIX. --- # Instructor [Instructor](https://python.useinstructor.com/) is a Python library for extracting structured data from LLM responses with Pydantic models. It adds response-model validation and retry behavior on top of provider clients. Instructor can wrap an OpenAI-compatible client that already points to AISIX. Use this setup when an application should keep Instructor's response-model workflow while routing model traffic through the gateway. This guide uses Instructor with the OpenAI Responses API. You will configure the OpenAI Python client with an AISIX proxy URL, caller API key, and model alias, then pass it to Instructor. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Python environment supported by [Instructor](https://python.useinstructor.com/). * A running AISIX gateway that the application can reach. * An AISIX caller API key. * A model alias the caller API key can access through the [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure Instructor[​](#configure-instructor "Direct link to Configure Instructor") Install the Python packages if the application does not already include them: ``` pip install instructor openai "pydantic>=2" ``` Set the values that the application will use: ``` # AISIX_BASE_URL ends in /v1 and has no trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` Create an OpenAI client for AISIX, then pass it to Instructor in Responses mode: ``` import os import instructor from instructor import Mode from openai import OpenAI from pydantic import BaseModel class Ticket(BaseModel): category: str priority: str openai_client = OpenAI( base_url=os.environ["AISIX_BASE_URL"], api_key=os.environ["AISIX_API_KEY"], ) client = instructor.from_openai( openai_client, mode=Mode.RESPONSES_TOOLS, ) ticket = client.responses.create( model=os.environ["AISIX_MODEL"], response_model=Ticket, input="Classify this ticket: production login failures for all users.", ) print(ticket.model_dump_json(indent=2)) ``` The model value is the AISIX model alias, not the upstream provider model ID. AISIX governs the model request path: it authenticates the caller API key, resolves the alias, applies configured policies, records telemetry, and dispatches the request to the provider behind the alias. Instructor still owns schema validation and retry behavior. Instructor also supports [Chat Completions modes](https://python.useinstructor.com/integrations/openai/) through `client.chat.completions.create(...)`. Use those modes when a workflow needs broad OpenAI-compatible chat support. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Run the script from the shell where the AISIX environment variables are set. When the request succeeds, verify the following results: * The script prints a validated Pydantic object. * AISIX records one or more successful `POST /v1/responses` requests for the selected model alias. Verify the request in the AISIX gateway logs. You can also use configured metrics or upstream provider logs. If the request fails, first confirm that the caller API key can access the selected model alias and that the OpenAI client uses the AISIX `base_url`. Instructor may retry requests when validation fails. Account for those retries when configuring rate limits and reviewing usage. With AISIX Cloud, include them when setting budgets. If validation keeps retrying, check whether the upstream model can satisfy the schema and whether the selected Instructor mode is supported by the AISIX route. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md): review gateway-facing Responses behavior. * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): use Chat Completions when a workflow needs broad OpenAI-compatible support. * [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md): account for validation retries with AISIX Cloud spend controls. --- # LangChain and LangGraph [LangChain](https://docs.langchain.com/) is a framework for building LLM applications, including chains, agents, retrieval workflows, and tool-calling flows. [LangGraph](https://docs.langchain.com/oss/python/langgraph/overview) builds on the LangChain ecosystem for stateful agents and graph-based workflows. In both cases, the gateway integration point is usually the model client rather than the rest of the application. The LangChain OpenAI package provides `ChatOpenAI`, a chat model client that can call a custom base URL and opt in to the OpenAI Responses API. Use those settings when a LangChain application should send Responses API requests through AISIX. This guide uses LangChain Python with `langchain-openai`. You will configure `ChatOpenAI` with an AISIX proxy URL, caller API key, model alias, and `use_responses_api=True`. LangGraph applications can reuse the same `ChatOpenAI` client inside graph nodes or agent workflows. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Python environment supported by [LangChain](https://python.langchain.com/docs/how_to/installation/). * A running AISIX gateway that the application can reach. * An AISIX caller API key. * A model alias the caller API key can access through the [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure LangChain[​](#configure-langchain "Direct link to Configure LangChain") Install the LangChain OpenAI integration if the application does not already include it: ``` pip install langchain-openai langgraph ``` Set the values that the LangChain application will use: ``` # AISIX_BASE_URL ends in /v1 and has no trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` Create a `ChatOpenAI` client with the AISIX values: ``` import os from langchain_openai import ChatOpenAI llm = ChatOpenAI( model=os.environ["AISIX_MODEL"], api_key=os.environ["AISIX_API_KEY"], base_url=os.environ["AISIX_BASE_URL"], use_responses_api=True, ) response = llm.invoke("Write one sentence about why gateways help AI teams.") print(response.content) ``` The model value is the AISIX model alias, not the upstream provider model ID. AISIX governs the model request path: it authenticates the caller API key, resolves the alias, applies configured policies, records telemetry, and dispatches the request to the provider behind the alias. The LangChain application still owns chains, prompts, tools, retrieval, and response handling. ## Use the Client in LangGraph[​](#use-the-client-in-langgraph "Direct link to Use the Client in LangGraph") Pass the same `ChatOpenAI` client into LangGraph nodes or graph state transitions: ``` import os from langchain_openai import ChatOpenAI from typing_extensions import TypedDict from langgraph.graph import END, START, StateGraph llm = ChatOpenAI( model=os.environ["AISIX_MODEL"], api_key=os.environ["AISIX_API_KEY"], base_url=os.environ["AISIX_BASE_URL"], use_responses_api=True, ) class State(TypedDict): prompt: str answer: str def call_model(state: State): response = llm.invoke(state["prompt"]) return {"answer": response.content} graph = StateGraph(State) graph.add_node("call_model", call_model) graph.add_edge(START, "call_model") graph.add_edge("call_model", END) app = graph.compile() result = app.invoke({ "prompt": "Write one sentence about governed agents.", }) print(result["answer"]) ``` LangGraph owns state management, graph execution, retries, and tool orchestration. AISIX only sees the model requests emitted by the configured `ChatOpenAI` client. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Run the script from the shell where the AISIX environment variables are set. When the request succeeds, verify the following results: * The script prints a chat response. * AISIX records a successful `POST /v1/responses` request for the selected model alias. Verify the request in the AISIX gateway logs. You can also use configured metrics or upstream provider logs. If the request fails, first confirm that the caller API key can access the selected model alias and that `base_url` points to the AISIX proxy API root with `/v1`. LangChain can also use Chat Completions through the same `ChatOpenAI` class. Use that path only when an existing chain or dependency requires Chat Completions compatibility; for new OpenAI-family LangChain integrations, prefer the Responses API path shown here. LangChain may send optional OpenAI fields when you enable streaming, tool calling, structured output, built-in tools, or reasoning options. Validate the exact chain or agent workflow before relying on gateway policy or telemetry for that path. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md): review gateway-facing request behavior. * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): use the Chat Completions route when an existing LangChain workflow requires it. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): confirm tool-call behavior before enabling LangChain tools or agents. * [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md): confirm gateway-side request metrics and logs. --- # LlamaIndex [LlamaIndex](https://docs.llamaindex.ai/) is a framework for building applications that combine LLM calls with data connectors, indexes, retrieval, and query workflows. AISIX fits at the model request boundary, while LlamaIndex continues to own document loading, indexing, retrieval, and response assembly. The LlamaIndex OpenAI integration provides `OpenAIResponses`, an LLM adapter for the OpenAI Responses API with a custom API base. Use that adapter when a LlamaIndex application should send Responses API requests through AISIX. This guide uses LlamaIndex Python with `llama-index-llms-openai`. You will configure `OpenAIResponses` with an AISIX proxy URL, caller API key, and model alias. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Python environment supported by [LlamaIndex](https://docs.llamaindex.ai/en/stable/getting_started/installation/). * A running AISIX gateway that the application can reach. * An AISIX caller API key. * A model alias the caller API key can access through the [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure LlamaIndex[​](#configure-llamaindex "Direct link to Configure LlamaIndex") Install the OpenAI LLM integration if the application does not already include it: ``` pip install llama-index-llms-openai ``` Set the values that the LlamaIndex application will use: ``` # AISIX_BASE_URL ends in /v1 and has no trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" export AISIX_CONTEXT_WINDOW="128000" ``` Create an `OpenAIResponses` client with the AISIX values: ``` import os from llama_index.core.llms import ChatMessage from llama_index.llms.openai import OpenAIResponses llm = OpenAIResponses( model=os.environ["AISIX_MODEL"], api_base=os.environ["AISIX_BASE_URL"], api_key=os.environ["AISIX_API_KEY"], context_window=int(os.environ["AISIX_CONTEXT_WINDOW"]), ) response = llm.chat([ ChatMessage(role="user", content="Write one sentence about retrieval systems."), ]) print(response.message.content) ``` The model value is the AISIX model alias, not the upstream provider model ID. Set `AISIX_CONTEXT_WINDOW` to the context size of the model behind the alias so LlamaIndex does not need to infer it from the alias name. AISIX governs the model request path, while the LlamaIndex application continues to own data loading, indexing, retrieval, and response assembly. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Run the script from the shell where the AISIX environment variables are set. When the request succeeds, verify the following results: * The script prints a chat response. * AISIX records a successful `POST /v1/responses` request for the selected model alias. Verify the request in the AISIX gateway logs. You can also use configured metrics or upstream provider logs. If the request fails, first confirm that the caller API key can access the selected model alias and that `api_base` points to the AISIX proxy API root with `/v1`. LlamaIndex also provides OpenAI-compatible chat adapters, including `OpenAILike`, for applications that still require Chat Completions compatibility. Use those adapters only when an existing workflow or integration depends on the Chat Completions route; for new OpenAI-family LlamaIndex integrations, prefer `OpenAIResponses`. LlamaIndex applications often combine retrieval, chat, tool calling, and structured response parsing. Validate the exact workflow before relying on gateway policy or telemetry for that path. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md): review gateway-facing request behavior. * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): use the Chat Completions route when an existing LlamaIndex workflow requires it. * [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md): configure embedding traffic if your LlamaIndex workflow uses AISIX for embeddings. * [Response Caching](https://docs.api7.ai/ai-gateway/traffic-controls/caching.md): cache eligible non-streaming chat responses. --- # Microsoft Agent Framework [Microsoft Agent Framework](https://learn.microsoft.com/en-us/agent-framework/) supports Python, C#, and Go workflows for building agents, tools, context providers, middleware, and AI applications. AISIX fits at the model request boundary when the application uses an OpenAI-compatible client with a custom endpoint. This guide uses the Python `OpenAIChatClient`, which uses the Responses API and supports OpenAI-compatible endpoints through `base_url`. You will configure the client with an AISIX proxy URL, caller API key, and model alias. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Python environment supported by [Microsoft Agent Framework](https://learn.microsoft.com/en-us/agent-framework/). * A running AISIX gateway that the application can reach. * An AISIX caller API key. * A model alias the caller API key can access through the [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure Microsoft Agent Framework[​](#configure-microsoft-agent-framework "Direct link to Configure Microsoft Agent Framework") Install the OpenAI provider package if the application does not already include it: ``` pip install agent-framework-openai ``` Set the values that the agent application will use: ``` # AISIX_BASE_URL ends in /v1 and has no trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` Create an `OpenAIChatClient` with the AISIX values and convert it to an agent: ``` import asyncio import os from agent_framework.openai import OpenAIChatClient client = OpenAIChatClient( base_url=os.environ["AISIX_BASE_URL"], api_key=os.environ["AISIX_API_KEY"], model=os.environ["AISIX_MODEL"], ) agent = client.as_agent( name="GatewayAgent", instructions="You answer in one concise sentence.", ) async def main(): response = await agent.run("Why do gateways help AI teams?") print(response) asyncio.run(main()) ``` Microsoft documents `OpenAIChatCompletionClient` as the [Chat Completions fallback](https://learn.microsoft.com/en-us/agent-framework/agents/providers/openai#chat-completion-client) for broad model compatibility or existing Chat Completions integrations. If your application still uses Semantic Kernel, configure its OpenAI chat completion connector with the AISIX endpoint instead: ``` #pragma warning disable SKEXP0010 builder.AddOpenAIChatCompletion( modelId: Environment.GetEnvironmentVariable("AISIX_MODEL")!, endpoint: new Uri(Environment.GetEnvironmentVariable("AISIX_BASE_URL")!), apiKey: Environment.GetEnvironmentVariable("AISIX_API_KEY")! ); #pragma warning restore SKEXP0010 ``` The model value is the AISIX model alias, not the upstream provider model ID. AISIX governs the model request path, while the Microsoft framework code still owns agents, tools, context providers, middleware, prompts, and workflow behavior. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Run the application from the shell where the AISIX environment variables are set. When the request succeeds, verify the following results: * The application prints an agent response. * AISIX records a successful `POST /v1/responses` request for the `OpenAIChatClient` path, or a successful `POST /v1/chat/completions` request for the Semantic Kernel fallback. Verify the request in the AISIX gateway logs. You can also use configured metrics or upstream provider logs. If the request fails, first confirm that the caller API key can access the selected model alias and that `base_url` points to the AISIX proxy API root with `/v1`. Microsoft Agent Framework supports both [Responses and Chat Completions clients](https://learn.microsoft.com/en-us/agent-framework/agents/providers/openai). Use Chat Completions when you need broad OpenAI-compatible model support through AISIX. Responses-hosted tools and provider-specific features may not map to every upstream provider. Validate Responses workflows before relying on gateway policy or telemetry for streaming, tools, or Responses-based hosted-tool behavior. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md): review gateway-facing Responses behavior. * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): use Chat Completions when a workflow needs broad OpenAI-compatible support. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): confirm tool-call behavior before enabling agent tools. --- # OpenAI Agents SDK [OpenAI Agents SDK](https://openai.github.io/openai-agents-python/) is a Python framework for building agent applications with agents, tools, handoffs, guardrails, sessions, and tracing. It uses the Responses API by default and can use a custom OpenAI-compatible endpoint by configuring a custom `AsyncOpenAI` client. Use this setup when an Agents SDK application should route model requests through AISIX while keeping the agent workflow. This guide uses the default Responses API model path in OpenAI Agents SDK. You will configure `AsyncOpenAI` with an AISIX proxy URL, caller API key, and model alias. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Python environment supported by [OpenAI Agents SDK](https://openai.github.io/openai-agents-python/). * A running AISIX gateway that the application can reach. * An AISIX caller API key. * A model alias the caller API key can access through the [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure OpenAI Agents SDK[​](#configure-openai-agents-sdk "Direct link to Configure OpenAI Agents SDK") Install OpenAI Agents SDK if the application does not already include it: ``` pip install openai-agents ``` Set the values that the agent application will use: ``` # AISIX_BASE_URL ends in /v1 and has no trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` Create an `AsyncOpenAI` client for AISIX and make it the SDK default client: ``` import asyncio import os from agents import ( Agent, AsyncOpenAI, Runner, set_default_openai_client, set_tracing_disabled, ) set_tracing_disabled(True) client = AsyncOpenAI( base_url=os.environ["AISIX_BASE_URL"], api_key=os.environ["AISIX_API_KEY"], ) set_default_openai_client(client, use_for_tracing=False) agent = Agent( name="GatewayAgent", instructions="You answer in one concise sentence.", model=os.environ["AISIX_MODEL"], ) async def main(): result = await Runner.run(agent, "Why do gateways help AI teams?") print(result.final_output) asyncio.run(main()) ``` The model value is the AISIX model alias, not the upstream provider model ID. AISIX governs the model request path, while OpenAI Agents SDK still owns agent instructions, tools, handoffs, guardrails, and run orchestration. This example disables the Agents SDK default trace export. If you need tracing, configure the SDK trace exporter separately so trace data follows your organization's policy. OpenAI Agents SDK also supports `OpenAIChatCompletionsModel` for the [Chat Completions API](https://openai.github.io/openai-agents-python/models/). Use that model shape when the workflow needs broad OpenAI-compatible provider support or already depends on Chat Completions behavior. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Run the script from the shell where the AISIX environment variables are set. When the request succeeds, verify the following results: * The script prints an agent response. * AISIX records a successful `POST /v1/responses` request for the selected model alias. Verify the request in the AISIX gateway logs. You can also use configured metrics or upstream provider logs. If the request fails, first confirm that the caller API key can access the selected model alias and that `base_url` points to the AISIX proxy API root with `/v1`. Hosted OpenAI tools and provider-specific Responses features may not map to every upstream provider. Validate Responses workflows before relying on gateway policy or telemetry for hosted tools, streaming, or provider-specific response items. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md): review gateway-facing Responses behavior. * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): use Chat Completions when a workflow needs broad OpenAI-compatible support. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): confirm tool-call behavior before enabling agent tools. --- # Pydantic AI [Pydantic AI](https://ai.pydantic.dev/) is a Python agent framework for building typed LLM applications with tools, dependency injection, structured outputs, and validation. It can use the OpenAI Responses API through `OpenAIResponsesModel` and `OpenAIProvider`. Use this setup when a Pydantic AI application should keep its agent and validation workflow while routing model requests through AISIX. This guide uses Pydantic AI with the OpenAI Responses model. You will configure `OpenAIProvider` with an AISIX proxy URL, caller API key, and model alias. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Python environment supported by [Pydantic AI](https://ai.pydantic.dev/install/). * A running AISIX gateway that the application can reach. * An AISIX caller API key. * A model alias the caller API key can access through the [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure Pydantic AI[​](#configure-pydantic-ai "Direct link to Configure Pydantic AI") Install Pydantic AI if the application does not already include it: ``` pip install pydantic-ai ``` Set the values that the Pydantic AI application will use: ``` # AISIX_BASE_URL ends in /v1 and has no trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` Create an OpenAI Responses model with the AISIX values: ``` import os from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIResponsesModel from pydantic_ai.providers.openai import OpenAIProvider model = OpenAIResponsesModel( os.environ["AISIX_MODEL"], provider=OpenAIProvider( base_url=os.environ["AISIX_BASE_URL"], api_key=os.environ["AISIX_API_KEY"], ), ) agent = Agent(model) response = agent.run_sync("Write one sentence about typed AI applications.") print(response.output) ``` The model value is the AISIX model alias, not the upstream provider model ID. AISIX governs the model request path, while Pydantic AI still owns agent execution, tool calls, dependency injection, structured outputs, and validation. Pydantic AI also provides [`OpenAIChatModel`](https://ai.pydantic.dev/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) for Chat Completions. Use that model when a workflow needs broad OpenAI-compatible chat support. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Run the script from the shell where the AISIX environment variables are set. When the request succeeds, verify the following results: * The script prints an agent response. * AISIX records a successful `POST /v1/responses` request for the selected model alias. Verify the request in the AISIX gateway logs. You can also use configured metrics or upstream provider logs. If the request fails, first confirm that the caller API key can access the selected model alias and that `base_url` points to the AISIX proxy API root with `/v1`. Pydantic AI supports multiple OpenAI model surfaces. Use `OpenAIResponsesModel` for Responses workflows and `OpenAIChatModel` for broad OpenAI-compatible Chat Completions support through AISIX. Validate structured outputs, tool definitions, and retries with the exact alias before relying on gateway policy or telemetry for that workflow. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md): review gateway-facing Responses behavior. * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): use Chat Completions when a workflow needs broad OpenAI-compatible support. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): confirm tool-call behavior before enabling Pydantic AI tools. * [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md): account for retries or validation loops with AISIX Cloud spend controls. --- # Vercel AI SDK [Vercel AI SDK](https://ai-sdk.dev/docs/introduction) is a TypeScript toolkit for building AI application flows, including text generation, streaming UI responses, tool calls, and structured output. In AISIX deployments, the application can keep using AI SDK helpers while the model provider configuration points to the gateway. The OpenAI provider package lets Vercel AI SDK call a custom base URL with an API key. Use that provider when a TypeScript application should send Responses API requests through AISIX. This guide uses the AI SDK core package with `@ai-sdk/openai`. You will create an OpenAI provider that points at AISIX and use an AISIX model alias from application code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A TypeScript project for the application. * A running AISIX gateway that the application can reach. * An AISIX caller API key. * A model alias the caller API key can access through the [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). If AISIX is already deployed for your organization, obtain the gateway URL, model alias, and caller API key from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. ## Configure Vercel AI SDK[​](#configure-vercel-ai-sdk "Direct link to Configure Vercel AI SDK") Install the AI SDK packages if the application does not already include them: ``` npm install ai @ai-sdk/openai ``` Set the values that the application will use: ``` # AISIX_BASE_URL ends in /v1 and has no trailing slash # The local quickstarts use http://127.0.0.1:3000/v1 export AISIX_BASE_URL="YOUR_AISIX_GATEWAY_URL/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="gpt-4o-prod" ``` Create an OpenAI provider for AISIX: ``` import { createOpenAI } from "@ai-sdk/openai"; import { generateText } from "ai"; const aisix = createOpenAI({ apiKey: process.env.AISIX_API_KEY, baseURL: process.env.AISIX_BASE_URL, }); async function main() { const { text } = await generateText({ model: aisix.responses(process.env.AISIX_MODEL ?? "gpt-4o-prod"), prompt: "Write one sentence about model governance.", }); console.log(text); } main(); ``` The model value is the AISIX model alias, not the upstream provider model ID. AISIX governs the model request path: it authenticates the caller API key, resolves the alias, applies configured policies, records telemetry, and dispatches the request to the provider behind the alias. The application still owns UI streaming, tool orchestration, and response handling. The OpenAI provider in AI SDK also supports Chat Completions through `aisix.chat(...)`, and the [`@ai-sdk/openai-compatible`](https://ai-sdk.dev/providers/ai-sdk-providers/openai-compatible) provider remains useful for broad OpenAI-compatible chat integrations. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Run the application from the shell where the AISIX environment variables are set. When the request succeeds, verify the following results: * The application prints generated text. * AISIX records a successful `POST /v1/responses` request for the selected model alias. Verify the request in the AISIX gateway logs. You can also use configured metrics or upstream provider logs. If the request fails, first confirm that the caller API key can access the selected model alias and that `baseURL` points to the AISIX proxy API root with `/v1`. This guide uses Responses-backed text generation. Validate any other AI SDK helper, such as streaming, tool calling, or structured output, before relying on gateway policy or telemetry for that path. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md): review gateway-facing Responses behavior. * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): use Chat Completions when a workflow needs broad OpenAI-compatible support. * [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md): confirm streaming behavior before using streamed UI responses. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): review OpenAI-compatible tool-call behavior through AISIX. --- # Voice Agent Platforms Voice applications bring together several latency-sensitive stages: audio transport, speech recognition, turn detection, language-model inference, tool execution, and speech synthesis. Voice agent platforms coordinate those stages so applications can hold realtime conversations over the web, in an application, or on a telephone call. The language model is only one part of that runtime, but it is often the part that organizations need to govern centrally. When a voice platform accepts an OpenAI-compatible LLM endpoint, its text model requests can pass through AISIX for caller authentication, model aliases, routing, policy, and telemetry. Teams can then manage model access in AISIX without replacing the platform that delivers the voice experience. The voice application or hosted platform remains responsible for the media session and converts speech into conversation content before calling AISIX. It sends that content and any tool definitions to the gateway. AISIX then authenticates the caller, resolves the model alias, applies configured policy, records telemetry, and dispatches the request to the selected model provider. ## Choose a Platform Guide[​](#choose-a-platform-guide "Direct link to Choose a Platform Guide") All guides use the same AISIX values, but each platform exposes them differently: | Platform | Integration surface | Guide | | ----------------- | --------------------------------------- | -------------------------------------------------------------------------------------------- | | ElevenLabs Agents | Hosted Custom LLM configuration | [ElevenLabs Agents](https://docs.api7.ai/ai-gateway/integrations/voice-agents/elevenlabs.md) | | LiveKit Agents | Python OpenAI LLM plugin | [LiveKit Agents](https://docs.api7.ai/ai-gateway/integrations/voice-agents/livekit.md) | | Pipecat | Python `OpenAILLMService` | [Pipecat](https://docs.api7.ai/ai-gateway/integrations/voice-agents/pipecat.md) | | Vapi | Hosted `custom-llm` model configuration | [Vapi](https://docs.api7.ai/ai-gateway/integrations/voice-agents/vapi.md) | ElevenLabs Agents and Vapi host the voice runtime. LiveKit Agents and Pipecat are application frameworks, so your application also configures its transport, speech-to-text service, text-to-speech service, and turn detection. ## Prepare AISIX[​](#prepare-aisix "Direct link to Prepare AISIX") Each platform needs these gateway-facing values: * An HTTPS AISIX proxy URL the platform or application can reach. * An AISIX caller API key dedicated to the voice application. * A model alias that the caller key can access through `POST /v1/chat/completions`. If AISIX is already deployed for your organization, obtain these values from the team that manages it. Otherwise, follow the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) or [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), or [contact API7](https://api7.ai/contact) for Hybrid Cloud access. Use a caller key scoped to only the aliases the voice application needs. Configure request and token limits for expected conversation volume, and set an expiration time for temporary evaluations. ## Understand the Boundary[​](#understand-the-boundary "Direct link to Understand the Boundary") AISIX receives the model request after the voice platform has converted speech into conversation content. AISIX therefore does not replace the platform's audio transport, speech recognition, speech synthesis, voice selection, interruption handling, or telephony. Gateway guardrails can inspect supported request and response text on the Chat Completions path. They do not inspect audio that remains inside the voice platform. See [Guardrail Behavior](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md) for exact endpoint and content coverage, and review each platform's recording, transcript, retention, and regional processing settings separately. Voice applications normally stream the model response so speech synthesis can begin before the complete answer is available. Confirm the model alias supports Chat Completions streaming, and account for network distance between the voice runtime, AISIX gateway, and model provider when measuring response latency. ## Verify an Integration[​](#verify-an-integration "Direct link to Verify an Integration") Use one short conversation to confirm the complete path: 1. The voice platform accepts a spoken or text test message. 2. AISIX records `POST /v1/chat/completions` for the expected caller key and model alias. 3. The platform receives the streamed model text and returns it to the user as audio or text. After the first success, test interruption behavior and any function tools the application depends on. Tool calling requires the selected upstream model and the platform integration to preserve the OpenAI-compatible tool-call stream. ## Related Guides[​](#related-guides "Direct link to Related Guides") Choose a platform guide above for setup. These guides explain the gateway behavior shared by the integrations: * [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md): review the gateway-facing request format. * [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md): understand stream delivery and failover boundaries. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): verify tool definitions and streamed tool calls. * [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md): limit voice-application traffic. --- # ElevenLabs Agents [ElevenLabs Agents](https://elevenlabs.io/docs/eleven-agents/overview) is a hosted platform for building and deploying conversational voice agents. It coordinates speech recognition, voice synthesis, turn handling, tools, and conversation delivery. The agent can listen, decide what to do, and respond without the application assembling each media component itself. Organizations may still want the agent's language-model traffic to use the same access controls, routing policy, and telemetry as their other AI applications. ElevenLabs supports a Custom LLM for this purpose. The agent continues to run its voice session in ElevenLabs while its OpenAI-compatible Chat Completions requests pass through AISIX to a configured model provider. In this integration, ElevenLabs stores a dedicated AISIX caller API key as a secret, sends an AISIX model alias in `model`, and streams requests to the gateway's `/v1/chat/completions` endpoint. AISIX governs the model request and response; ElevenLabs continues to handle the audio and conversation experience. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * An ElevenLabs account with access to ElevenLabs Agents. * An ElevenLabs agent you can edit. * A publicly reachable HTTPS AISIX proxy URL. * An AISIX caller API key dedicated to the agent. * A model alias the caller key can access through the [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). Do not store an upstream provider key in ElevenLabs. The Custom LLM secret should contain only the AISIX caller key. ## Configure the Custom LLM[​](#configure-the-custom-llm "Direct link to Configure the Custom LLM") Open the agent in the ElevenLabs dashboard, then configure the model: 1. On the **Agent** page, open **LLM**. 2. Open the primary model selector and choose **Custom LLM**. 3. Keep **Chat Completions** as the API format. 4. Set **Server URL** to the AISIX proxy API root, including `/v1`, for example `https://gateway.example.com/v1`. ElevenLabs appends `/chat/completions`. 5. Set **Model ID** to the AISIX model alias, for example `voice-agent-prod`. 6. Under **API Key**, create or select a secret containing the AISIX caller key. 7. Select **Test Connection**. The connection test should receive a successful OpenAI-compatible stream. If it fails, confirm the URL is reachable from the public internet and that the caller key can access the alias. Close the LLM settings, then publish the agent when the connection test succeeds. ## Verify a Conversation[​](#verify-a-conversation "Direct link to Verify a Conversation") Open **Preview** and start one short conversation. Speak a simple request, such as “Reply with one short sentence.” Confirm these results: * The agent returns a spoken response. * AISIX records a successful `POST /v1/chat/completions` request. * The recorded caller identity and model match the dedicated ElevenLabs key and AISIX alias. ElevenLabs handles the speech-to-text result, generated voice, recording settings, and conversation transcript. AISIX sees the OpenAI-compatible model request and response text, not the original audio stream. ## Use Tools Carefully[​](#use-tools-carefully "Direct link to Use Tools Carefully") ElevenLabs system tools and custom tools require the Custom LLM to produce compatible function calls. Before enabling a tool in production, verify that the selected model alias returns the tool name and arguments through AISIX and that ElevenLabs executes the tool as expected. A successful plain-text connection test does not prove tool compatibility. Test every system or custom tool the agent uses, including any multi-tool turn, before production. ## Troubleshoot the Connection[​](#troubleshoot-the-connection "Direct link to Troubleshoot the Connection") | Symptom | Check | | -------------------------------- | ------------------------------------------------------------------------------------------------------------- | | Server URL is invalid | Enter the HTTPS AISIX base through `/v1`, without adding `/chat/completions`; the dashboard adds that suffix. | | Connection test returns `401` | Confirm the selected ElevenLabs secret contains the AISIX caller key. | | Connection test returns `403` | Confirm the caller key allows the selected model alias. | | Connection test returns `404` | Confirm the model ID is an AISIX model alias and the AISIX proxy route is reachable. | | The agent waits without speaking | Check the AISIX stream and upstream latency, then verify the ElevenLabs voice and turn settings. | ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md): review gateway stream behavior. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): verify function calls through AISIX. * [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md): monitor model latency and request outcomes. --- # LiveKit Agents [LiveKit Agents](https://docs.livekit.io/agents/) is a framework for building realtime voice and media agents. It connects an agent's session logic with audio or video transport, speech recognition, language models, speech synthesis, turn detection, and tools. Together, these components allow the application to respond during a live interaction. The framework keeps those components modular. A LiveKit application can continue using its existing room or transport and speech services while sending only the text language-model step through AISIX. This gives the application AISIX model aliases, access control, routing, and telemetry without moving the media session into the gateway. The integration uses LiveKit's `openai.LLM` client with an AISIX proxy URL, caller API key, and model alias. It does not use the OpenAI Realtime model plugin; LiveKit continues to coordinate the audio pipeline around the text LLM. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Python project using LiveKit Agents. * A running AISIX gateway the LiveKit worker can reach. * An AISIX caller API key. * A model alias the caller key can access through the [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). The worker also needs its existing LiveKit, speech-to-text, and text-to-speech configuration. Those credentials remain separate from the AISIX caller key. ## Configure the LLM Plugin[​](#configure-the-llm-plugin "Direct link to Configure the LLM Plugin") Install the LiveKit OpenAI plugin if the project does not already include it: ``` pip install "livekit-agents[openai]~=1.5" ``` Set the gateway values in the worker environment: ``` # AISIX_BASE_URL includes /v1 and has no trailing slash. export AISIX_BASE_URL="https://gateway.example.com/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="voice-agent-prod" ``` Create the LLM plugin with those values: ``` import os from livekit.plugins import openai llm = openai.LLM( api_key=os.environ["AISIX_API_KEY"], base_url=os.environ["AISIX_BASE_URL"], model=os.environ["AISIX_MODEL"], ) ``` Pass `llm` to the application's `AgentSession`. Keep the existing transport, STT, TTS, and turn-detection components unchanged. ## Test the LLM Connection[​](#test-the-llm-connection "Direct link to Test the LLM Connection") Before joining a LiveKit room, test the same plugin as a standalone streaming client: ``` import asyncio import os from livekit.agents import ChatContext from livekit.plugins import openai async def main(): context = ChatContext() context.add_message( role="user", content="Reply with one short sentence about voice gateways.", ) llm = openai.LLM( api_key=os.environ["AISIX_API_KEY"], base_url=os.environ["AISIX_BASE_URL"], model=os.environ["AISIX_MODEL"], ) stream = llm.chat(chat_ctx=context) async for text in stream.to_str_iterable(): print(text, end="", flush=True) await llm.aclose() asyncio.run(main()) ``` The script should print the streamed model response. AISIX should record `POST /v1/chat/completions` for the selected caller key and alias. After this test passes, run the normal LiveKit worker and verify one short voice turn. If text generation succeeds but no audio is returned to the participant, troubleshoot the LiveKit TTS and room pipeline rather than changing the AISIX model endpoint. ## Understand the API Choice[​](#understand-the-api-choice "Direct link to Understand the API Choice") LiveKit documents the Responses API plugin as its preferred path for direct OpenAI usage. This guide uses `openai.LLM` because LiveKit assigns that client to OpenAI-compatible Chat Completions endpoints, which is the broadly compatible AISIX integration path. Do not point the LiveKit OpenAI Realtime plugin at AISIX unless the application specifically uses the gateway's [Realtime API](https://docs.api7.ai/ai-gateway/endpoints/realtime.md) and has verified the complete audio event protocol. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md): understand stream completion and failover behavior. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): verify tools used by the LiveKit agent. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): configure resilient model targets for voice traffic. --- # Pipecat [Pipecat](https://docs.pipecat.ai/) is an open-source Python framework for realtime voice and multimodal agents. It represents a conversation as a pipeline of frames and services. Applications can compose transport, speech recognition, context management, language-model inference, speech synthesis, and other processors around their own interaction logic. Because the services are modular, an existing Pipecat pipeline can route its text language-model requests through AISIX without changing the transport or speech components. AISIX then provides the model alias, access controls, routing, and telemetry for that stage, while Pipecat continues to move audio, transcripts, model text, and synthesized speech through the application pipeline. Pipecat's `OpenAILLMService` accepts a custom OpenAI base URL and API key. Set them to the AISIX proxy API root and caller key, and use an AISIX model alias in the service settings. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Python project using Pipecat. * A running AISIX gateway the Pipecat process can reach. * An AISIX caller API key. * A model alias the caller key can access through the [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). The pipeline also needs its normal transport, speech-to-text, and text-to-speech services. Their credentials are not replaced by the AISIX caller key. ## Configure OpenAILLMService[​](#configure-openaillmservice "Direct link to Configure OpenAILLMService") Install Pipecat's OpenAI integration if it is not already available: ``` pip install "pipecat-ai[openai]" ``` Set the gateway values: ``` # AISIX_BASE_URL includes /v1 and has no trailing slash. export AISIX_BASE_URL="https://gateway.example.com/v1" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export AISIX_MODEL="voice-agent-prod" ``` Create the LLM service with the current settings-based configuration: ``` import os from pipecat.services.openai.llm import OpenAILLMService llm = OpenAILLMService( api_key=os.environ["AISIX_API_KEY"], base_url=os.environ["AISIX_BASE_URL"], settings=OpenAILLMService.Settings( model=os.environ["AISIX_MODEL"], ), ) ``` Place `llm` in the existing Pipecat pipeline between the user context aggregator and the TTS service. The model setting uses the AISIX alias, not the upstream provider's model ID. ## Test the LLM Connection[​](#test-the-llm-connection "Direct link to Test the LLM Connection") Test streaming before starting the complete media pipeline: ``` import asyncio import os from pipecat.processors.aggregators.llm_context import LLMContext from pipecat.services.openai.llm import OpenAILLMService async def main(): llm = OpenAILLMService( api_key=os.environ["AISIX_API_KEY"], base_url=os.environ["AISIX_BASE_URL"], settings=OpenAILLMService.Settings( model=os.environ["AISIX_MODEL"], ), ) context = LLMContext( messages=[ { "role": "user", "content": "Reply with one short sentence about voice gateways.", } ] ) stream = await llm.get_chat_completions(context) async for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) await llm.cleanup() asyncio.run(main()) ``` The script should print the streamed response. AISIX should record `POST /v1/chat/completions` for the configured alias. The final cleanup line runs the service's public cleanup hook for this standalone smoke test. In a normal Pipecat application, the pipeline lifecycle manages service startup and cleanup. ## Verify the Voice Pipeline[​](#verify-the-voice-pipeline "Direct link to Verify the Voice Pipeline") After the LLM test passes, start the normal Pipecat application and complete one voice turn. Use Pipecat metrics to separate speech recognition, LLM, and speech synthesis latency, and compare the LLM segment with AISIX request metrics. If the standalone LLM test succeeds but the full pipeline does not speak, inspect the context aggregator, TTS service, and output transport. AISIX returns model text and tool calls; it does not create Pipecat audio frames. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md): review the gateway stream path. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): validate Pipecat function handlers with AISIX model output. * [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md): compare gateway latency with Pipecat service metrics. --- # Vapi [Vapi](https://docs.vapi.ai/) is a hosted platform for building and operating voice assistants for web and telephone conversations. It coordinates call transport, speech recognition, language-model turns, voice generation, tools, and assistant orchestration so applications can deliver a realtime voice experience through a configured assistant. Vapi's Custom LLM provider separates that orchestration from the model endpoint. A Vapi assistant can keep its speech services, call handling, and tools in Vapi while its OpenAI-compatible language-model requests pass through AISIX. This allows teams to apply AISIX model aliases, caller authorization, routing, and telemetry to the LLM stage without treating AISIX as the voice runtime. Configure the Vapi model with the AISIX OpenAI-compatible API root, an AISIX model alias, and a credential containing a dedicated AISIX caller API key. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A Vapi account and an assistant you can edit. * A publicly reachable HTTPS AISIX proxy URL. * An AISIX caller API key dedicated to the assistant. * A model alias the caller key can access through the [OpenAI-Compatible API](https://docs.api7.ai/ai-gateway/endpoints/openai-compatible-chat.md). Keep the upstream provider credential in AISIX. Vapi should store only the restricted AISIX caller key. ## Configure a Custom LLM[​](#configure-a-custom-llm "Direct link to Configure a Custom LLM") First, store the AISIX caller key as the organization-level Custom LLM credential: 1. In the Vapi dashboard, open **Settings → Integrations**. 2. Under **Model Providers**, select **Configure Custom LLM**. 3. Enter the AISIX caller key in **API Key**. Leave the optional OAuth2 fields empty. 4. Select **Save**. Next, configure the assistant to use AISIX: 1. Open **Assistants**, then select the assistant you want to configure. 2. Select the **Model** card and choose **Custom LLM**. 3. Set **Model** to the AISIX model alias, for example `voice-agent-prod`. 4. Set **Custom LLM URL** to the AISIX API root, including `/v1`, for example `https://gateway.example.com/v1`. 5. After Vapi saves the draft, select **Publish** and review the model changes. 6. Select **Next**, enter a version name under **Publish Description**, and select **Publish** again. Vapi uses the URL as an OpenAI client base URL and appends `/chat/completions`. It sends the credential as a bearer token in the `Authorization` header when it calls the custom endpoint. AISIX authenticates that value as the caller key. The AISIX alias and its upstream model must support streamed Chat Completions. Vapi can begin synthesizing the answer as text deltas arrive. ## Verify the Integration[​](#verify-the-integration "Direct link to Verify the Integration") Run a short web or telephone test and ask for a one-sentence response. Confirm these results: * Vapi returns the response as speech. * AISIX records a successful `POST /v1/chat/completions` request. * The caller key and model alias match the dedicated Vapi configuration. Vapi handles call recording settings, transcripts, speech services, and telephony events. AISIX governs only the LLM exchange that Vapi sends through the custom endpoint. ## Verify Tools Separately[​](#verify-tools-separately "Direct link to Verify Tools Separately") Vapi can include OpenAI-compatible tool definitions in Custom LLM requests. If the assistant uses tools, verify the streamed function name, arguments, and tool-call ID through AISIX before relying on the integration in production. Vapi tool execution can also involve assistant, tool, or account-level server URLs. Those webhook endpoints are separate from the Custom LLM URL and are not replaced by AISIX. ## Troubleshoot Vapi Requests[​](#troubleshoot-vapi-requests "Direct link to Troubleshoot Vapi Requests") | Symptom | Check | | ------------------------------ | ----------------------------------------------------------------------------------------------- | | Authentication fails | Confirm the Custom LLM credential contains the AISIX caller key, not the upstream provider key. | | Model is not found | Confirm the Vapi model value is an AISIX alias visible to the caller key. | | Response is delayed | Check the Vapi transcriber and voice latency separately from AISIX upstream latency. | | Text succeeds but a tool fails | Verify Vapi's tool server configuration and the model's streamed function-call shape. | ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Streaming](https://docs.api7.ai/ai-gateway/endpoints/streaming.md): understand streamed response delivery. * [Tool Calling](https://docs.api7.ai/ai-gateway/endpoints/tool-calling.md): verify the gateway tool-call contract. * [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md): limit assistant traffic. --- # Manage MCP Access with Policies [AISIX Cloud Only](https://docs.api7.ai/ai-gateway/cloud/overview.md)Available with AISIX Cloud Granting MCP tools on each caller API key works well for a small number of callers. At larger scale, registering a new MCP server or tool can require updating hundreds of keys. MCP access policies move the shared part of the grant to the environment or team level, where it applies to every key those layers cover. This guide explains how the layers combine, how to configure the environment and team grants, and how to inspect a key's effective access. ## How the Layers Combine[​](#how-the-layers-combine "Direct link to How the Layers Combine") Up to three layers apply to a caller API key. They all carry the same two fields — `allow` and `deny` — and none of them overrides another: 1. **Environment policy:** applies to every key in the environment. 2. **Team policy:** applies to the keys bound to one team, in every environment of the organization. 3. **Key `mcp_access` block:** the key's own layer. A key's effective access is resolved per request: ``` effective = (every present layer's allow, intersected) − (every present layer's deny, unioned) ``` Three rules keep the model predictable: * **Allow intersects.** A tool is available only when every layer that is present allows it, so any layer can narrow the result and none can widen it. A team policy cannot grant a tool the environment policy withholds, and neither can a key. * **Deny always wins.** A deny pattern on any layer removes the tool, whatever the other layers allow. * **A missing layer imposes no constraint, but no layer at all grants nothing.** A key with no `mcp_access` block, a team with no policy, and a disabled policy each simply drop out of the intersection. When none of the three is configured, the key has no MCP tool access — access is granted explicitly, never by the absence of configuration. Because `allow` is required on every layer, the two edge cases are always spelled out rather than implied: | `allow` | Meaning | | ------- | ---------------------------------------------------------------------------------- | | `[]` | This layer allows nothing, so every key it covers loses MCP access. | | `["*"]` | This layer narrows nothing. Use it for a layer that only subtracts through `deny`. | All patterns use the `__` naming and single-`*` glob matching described in [Control Tool Access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") The examples on this page use the AISIX Cloud Admin API. Prepare the control-plane URL, environment ID, and write-scoped admin token for your organization. Install `curl` and `jq` to run the commands. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). Then set: ``` # AISIX_CP includes /api and has no trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` ## Set the Environment Policy[​](#set-the-environment-policy "Direct link to Set the Environment Policy") Save the environment layer: ``` curl -sS -X PUT "$AISIX_CP/environments/$ENV_ID/mcp_policy" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "allow": ["github__*"], "deny": ["github__delete_repository"] }' ``` ❶ The environment allows every `github` tool and nothing else. Send `["*"]` only when every tool on every current and future server should be reachable, and `[]` to block MCP access for the whole environment. ❷ Deny patterns apply to every key in the environment, however the team and key layers are configured. The dashboard provides the same editor under **Environment → MCP Access**. ## Give a Key Its Own Layer[​](#give-a-key-its-own-layer "Direct link to Give a Key Its Own Layer") A key that adds no layer of its own takes whatever the environment and team layers leave: ``` API_KEY_RESPONSE=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "inheriting-caller", "allowed_models": [] }') export API_KEY_ID=$(printf '%s\n' "$API_KEY_RESPONSE" | jq -er '.api_key.id') export AISIX_API_KEY=$(printf '%s\n' "$API_KEY_RESPONSE" | jq -er '.plaintext') printf '%s\n' "$API_KEY_RESPONSE" | jq ``` `API_KEY_ID` and `AISIX_API_KEY` refer to the same caller. AISIX Cloud returns the plaintext only in this create response, so retain both variables for setup and verification. Add an `mcp_access` block when the key should reach less than its environment and team allow: ``` "mcp_access": { "allow": ["github__*"], "deny": ["github__delete_repository"] } ``` The block is the key's whole allow side, so `"allow": []` leaves the key no MCP access at all, and `"allow": ["*"]` narrows nothing — useful when the key only means to subtract through its own `deny`. Send `"mcp_access": null` on an update to remove the key's layer, returning it to whatever the other layers leave. ## Verify Effective Access[​](#verify-effective-access "Direct link to Verify Effective Access") Inspect the key to see which layers constrain it and where each pattern came from: ``` curl -sS "$AISIX_CP/environments/$ENV_ID/api_keys/$API_KEY_ID/effective_permissions" \ -H "Authorization: Bearer $AISIX_TOKEN" | jq ``` The key above is constrained by the environment layer alone: ``` { "effective_permissions": { "mcp": { "layers": [ { "source": "env_policy", "policy_id": "a5065729-3049-407a-90d4-357a6ab214c2" } ], "all_tools": false, "allow": [ { "pattern": "github__*", "source": "env_policy" } ], "deny": [ { "pattern": "github__delete_repository", "source": "env_policy" } ] } } } ``` An empty `layers` array means nothing is configured anywhere, which is why such a key has no MCP tool access. The dashboard shows the same view from the key's row on the **API Keys** page. ## Grant a Team Policy[​](#grant-a-team-policy "Direct link to Grant a Team Policy") A team policy is configured once per team and applies to caller API keys explicitly bound to that team, in every environment of the organization. [SCIM directory sync](https://docs.api7.ai/ai-gateway/cloud/scim-directory-sync.md) updates the team's member roster when identity provider group membership changes, but it does not bind or rebind caller API keys. A caller key receives the team policy only after it is [bound to the team](https://docs.api7.ai/ai-gateway/cloud/teams.md#bind-caller-api-keys-to-the-team). Copy the Team ID from the [Teams](https://docs.api7.ai/ai-gateway/cloud/teams.md#create-a-team) page, then export it for the request: ``` export TEAM_ID="YOUR_TEAM_ID" curl -sS -X PUT "$AISIX_CP/teams/$TEAM_ID/entitlements" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "mcp": { "allow": ["postgres__*"] } }' ``` The team layer intersects with the environment layer rather than replacing it: team-bound keys reach `postgres__*` only if the environment policy allows it too. Send `"mcp": null` to remove the layer, leaving those keys with the environment layer and any key-level `mcp_access` block that remains. ## Operational Notes[​](#operational-notes "Direct link to Operational Notes") * A policy can carry `"enabled": false`. A disabled policy is not a layer at all: it neither grants nor denies, and drops out of the intersection. * Deleting the environment policy leaves keys that have no team policy and no `mcp_access` block with no layer at all, and therefore no MCP access. The resolution is fail-closed, never fail-open. * Unauthorized tools are filtered from `tools/list` and rejected on `tools/call` before any upstream routing, with the same neutral error. * Every policy write and entitlement change is recorded in the audit log. ## Next Steps[​](#next-steps "Direct link to Next Steps") If the upstream server is not registered yet, use the registration workflow in [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md#register-the-server-and-grant-a-tool). Before verifying `everything__echo`, make sure every access layer that applies to the caller allows it. The setup workflow updates the key-level grant; update any applicable environment or team policy as well. Then add [rate limits and budgets](https://docs.api7.ai/ai-gateway/mcp-gateway/traffic-controls.md), [guardrails](https://docs.api7.ai/ai-gateway/mcp-gateway/guardrails.md), or [observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md) to the tool-call path. --- # Client Authentication An MCP client reaches AISIX at one of two entries: the aggregated `/mcp` endpoint, which serves every registered server's tools under `__` names, and `/mcp/{server}`, which serves one server's tools under their original names. This guide covers how a caller proves who it is at those entries. Client authentication is independent of [upstream authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/upstream-authentication.md): the credential a client sends to AISIX identifies the caller to AISIX, and the credential AISIX sends to an upstream MCP server is held gateway-side and never forwarded. Three modes are available, and an environment can combine them: | Mode | The client sends | Use it for | | --------------- | ------------------------------------------- | --------------------------------------------------------------------- | | Gateway API key | `Authorization: Bearer ` | The default. Machine-to-machine callers and agents you issue keys to. | | OAuth sign-in | An access token from your identity provider | Standard MCP clients that discover the sign-in flow themselves. | | Anonymous | Nothing | Clients on trusted networks that cannot present a credential. | Whatever the mode, the caller resolves to an API key principal. Its tool grant determines what it can list and call. Configured rate limits and guardrails can govern tool calls, while usage remains attributed to the principal. In AISIX Cloud, matching budgets also apply. Anonymous callers therefore remain subject to the same controls as authenticated callers. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * Complete [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md) for AISIX Cloud or an open-source AISIX gateway, and retain the `AISIX_PROXY` and `AISIX_MCP_KEY` values used at the end of that guide. * For the AISIX Cloud configuration examples, also retain the `AISIX_CP`, `AISIX_TOKEN`, and `ENV_ID` values from the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * [cURL](https://curl.se/) for the request examples. ## Gateway API Key[​](#gateway-api-key "Direct link to Gateway API Key") This is the default and needs no configuration. A client sends its key on every request: ``` curl -sS -X POST "$AISIX_PROXY/mcp" \ -H "Authorization: Bearer $AISIX_MCP_KEY" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' ``` The key's grant decides which tools the caller sees and can call. See [Control Tool Access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md) for scoping a key to specific tools, whole servers, or every tool, and [MCP Access Policies](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md) for granting at the environment or team level. `x-api-key: ` is accepted as an alternative header. ## OAuth Sign-In[​](#oauth-sign-in "Direct link to OAuth Sign-In") Standard MCP clients such as desktop assistants can sign a user in instead of asking them to paste a key. They do it by reading the `WWW-Authenticate` header on a `401`, fetching the protected resource metadata it points at, and running the OAuth flow against the authorization server named there. AISIX publishes that metadata once an environment has both: * a canonical MCP resource URL — the public URL clients use to reach this environment's `/mcp` endpoint; and * at least one enabled OIDC trust provider, which is the authorization server tokens must come from. With both configured, `GET /.well-known/oauth-protected-resource` (and its `/.well-known/oauth-protected-resource/mcp` sibling) returns the resource identity, the issuers tokens may come from, and the scopes they must carry. Without them the routes return `404` and `401` responses carry no challenge, exactly as before the feature existed. Access tokens must include the resource URL in their audience claim. This is the most common configuration mistake: a token minted for a different audience is rejected at the gateway even though sign-in succeeded. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Set the resource URL on the environment: ``` curl -sS -X PATCH "$AISIX_CP/environments/$ENV_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "mcp_resource_url": "https://gateway.example.com/mcp" }' ``` The URL must be an absolute `http` or `https` URL whose path is exactly `/mcp`, with no query or fragment, and no embedded credentials — it is published on an unauthenticated endpoint. Send `null` to clear it and turn discovery off. In the dashboard, the same setting lives on the environment's **MCP Access** page, which also warns when an enabled provider's audiences do not include the URL. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Set `resource_url` on the singleton `mcp_auth_settings` entry, creating it if absent, and keep any `anonymous` settings unchanged. Add `corp-sso` to `oidc_providers`: resources.yaml (OAuth discovery) ``` mcp_auth_settings: - resource_url: https://gateway.example.com/mcp oidc_providers: - name: corp-sso issuer: https://sso.example.com/realms/agents audiences: - https://gateway.example.com/mcp required_scopes: - mcp:tools ``` At most one `mcp_auth_settings` entry may exist. A second one is rejected at load, and if a duplicate ever reaches a running gateway the discovery surface stays off rather than picking one. ## Anonymous Access[​](#anonymous-access "Direct link to Anonymous Access") Anonymous access lets clients that present no credential reach entries you open for them. It exists for fleets migrating from a gateway that never required a credential, where changing every client is not practical. An anonymous request still resolves to an API key principal you choose. Its tool grants determine what it can list and call. Configured rate limits and guardrails can govern tool calls, while usage remains attributed to the principal. In AISIX Cloud, matching budgets also apply. Only the credential check changes. caution Anyone who can reach the gateway from the allowed networks can call the permitted tools without a credential, and the usage counts against the environment. Treat the source network allowlist as the access control it is, and keep the principal's tool grant as narrow as the clients actually need. ### What Anonymous Is Not[​](#what-anonymous-is-not "Direct link to What Anonymous Is Not") **It is not a downgrade path.** A request that presents a credential is authenticated normally, and an invalid, expired, disabled or malformed one is rejected with `401`. Only a request that presents nothing at all takes the anonymous path. A wrong authentication scheme or an empty header value counts as presenting something, so a client that tries to authenticate and gets it wrong fails rather than quietly succeeding with a different identity. **It is not visible to callers who are not allowed in.** Every refusal — source outside the allowlist, a server not offered anonymously, the principal deleted or disabled, anonymous access off — answers the same `401` the entry gives without anonymous access at all. Callers cannot tell those apart, nor tell a registered MCP server from one that does not exist. Operators see the reason on the gateway's `aisix_auth_decisions_total` metric. ### Configuration[​](#configuration "Direct link to Configuration") Anonymous access is configured per environment: | Field | Meaning | | ----------------- | --------------------------------------------------------------------------------------------- | | `api_key_id` | The API key anonymous traffic runs as. | | `source_cidrs` | Client source networks allowed in. Required and non-empty. | | `servers` | MCP servers anonymous callers may reach. Required and non-empty. | | `aggregate_entry` | Whether the aggregated `/mcp` endpoint also serves anonymous callers. Off by default. | | `enabled` | Set to `false` to close anonymous access while keeping the configuration. Defaults to `true`. | Two of these deserve more than a one-line description. #### The Server List Is a Ceiling[​](#the-server-list-is-a-ceiling "Direct link to The Server List Is a Ceiling") `servers` is not only the list of `/mcp/{server}` entries to open — it is the limit on what the principal can reach anywhere, including through the aggregated endpoint. Without that, a principal whose own tool grant is wider than the list could name `__` on the aggregated endpoint and reach a server whose per-server entry is closed. So the effective grant of an anonymous caller is the listed servers' tools **intersected with** the principal's own grant. Both `tools/list` and `tools/call` follow it, which is why an anonymous caller never sees a tool it could not call. A newly registered MCP server is never anonymous by default. Reaching anonymous callers is always a name added to this list. #### The Principal Needs Its Own Grant[​](#the-principal-needs-its-own-grant "Direct link to The Principal Needs Its Own Grant") In AISIX Cloud, the principal must carry its own `mcp_access` block. A key without one is rejected, because it would take whatever the environment and team layers leave — so a later policy change could widen anonymous access without anyone revisiting this setting. A key with its own block is bounded by its own `allow` list whatever those layers do. In an open-source gateway configured with a resources file the same block is the principal's only layer, since the file has no policy collection. The `servers` list remains an additional ceiling on it. ### Configure in AISIX Cloud[​](#configure-in-aisix-cloud "Direct link to Configure in AISIX Cloud") Export the ID of the API key that anonymous traffic should run as. The key must belong to the environment and carry the MCP grant described above: ``` export ANON_KEY_ID="YOUR_ANONYMOUS_PRINCIPAL_API_KEY_ID" ``` Set the block on the environment: ``` curl -sS -X PATCH "$AISIX_CP/environments/$ENV_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "mcp_anonymous": { "api_key_id": "'"$ANON_KEY_ID"'", "source_cidrs": ["10.0.0.0/8"], "servers": ["everything"], "aggregate_entry": false } }' ``` ❶ The API key principal that anonymous traffic runs as. It must belong to this environment and carry its own `mcp_access` block. ❷ Matched against the source address AISIX resolves through its real-IP configuration, never a header the client supplies. Use `10.0.0.1/32` for a single address. ❸ Approved servers exposed to this environment. Anonymous callers reach `/mcp/everything` and nothing else. ❹ Leave the aggregated endpoint on gateway credentials. See [Anonymous Access and OAuth Sign-In](#anonymous-access-and-oauth-sign-in) before turning it on. Send `"mcp_anonymous": null` to turn anonymous access off. The change reaches running gateways without a restart. In the dashboard, the same settings live on the environment's **MCP Access** page, where enabling anonymous access requires an explicit risk acknowledgement. ### Configure in an Open-Source Gateway[​](#configure-in-an-open-source-gateway "Direct link to Configure in an Open-Source Gateway") Add `anonymous-mcp` to `api_keys`. In the singleton `mcp_auth_settings` entry, add `anonymous`; create the entry if it is absent. Keep `resource_url`, the `everything` server, and the other resources unchanged. The ID below is derived from `anonymous-mcp`: resources.yaml (anonymous MCP access) ``` api_keys: - display_name: anonymous-mcp key_env: ANONYMOUS_MCP_KEY allowed_models: [] mcp_access: allow: - everything__* mcp_auth_settings: - anonymous: api_key_id: d6869ae7-741a-598e-8213-16672e922546 source_cidrs: - 10.0.0.0/8 servers: - everything aggregate_entry: false ``` Set `ANONYMOUS_MCP_KEY` in the gateway process environment before loading the file. If you use a different API key display name, replace `api_key_id` with that entry's [deterministic derived ID](https://docs.api7.ai/ai-gateway/reference/resources-file.md#identity-and-derived-ids). The `resource_url` field for OAuth discovery lives on the same `mcp_auth_settings` entry; the two settings are independent, and either can be configured without the other. ### Anonymous Access and OAuth Sign-In[​](#anonymous-access-and-oauth-sign-in "Direct link to Anonymous Access and OAuth Sign-In") An environment can run both, and for most deployments the natural split is: existing clients that cannot present a credential use the per-server entries anonymously, while standard MCP clients sign in through the aggregated `/mcp`. Turning on `aggregate_entry` in an environment that also publishes OAuth discovery changes that. A no-credential request to `/mcp` then succeeds instead of returning the `401` that carries the discovery hint, so OAuth-capable clients never start the sign-in flow and stay on the anonymous grant. The per-server entries are unaffected. ## Verify[​](#verify "Direct link to Verify") Confirm the mode you configured behaves as intended. An anonymous call to a listed entry succeeds with no credential: ``` curl -sS -X POST "$AISIX_PROXY/mcp/everything" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' ``` The response lists the tools the principal's grant allows on that server, under their original names. A bad credential is still rejected, rather than served anonymously: ``` curl -sS -o /dev/null -w '%{http_code}\n' -X POST "$AISIX_PROXY/mcp/everything" \ -H "Authorization: Bearer not-a-real-key" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' ``` The gateway returns `401`. ## Observability[​](#observability "Direct link to Observability") Anonymous traffic is attributable. Usage events carry the principal's API key id like any other request, plus `auth_type: anonymous`, which distinguishes traffic that inherited the principal from an entry from traffic that presented that key's own credential. Authentication decisions, including refusals and their reasons, are counted on `aisix_auth_decisions_total`. ## Limitations[​](#limitations "Direct link to Limitations") Anonymous access is designed for trusted networks. Two capacity protections that authenticated deployments can rely on are not yet in place for it: * Per-source-IP rate limiting is not available. A single anonymous client can consume the principal's whole quota, since all anonymous traffic shares one principal. * The `initialize`, `ping` and `tools/list` methods are not metered. Only `tools/call` passes the rate-limit gate and checks applicable AISIX Cloud budgets. The mandatory source network allowlist is what keeps this bounded. Do not expose anonymous entries to untrusted networks. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Control tool access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md): scope a key — including an anonymous principal — to specific tools or whole servers. * [Rate limits and budgets](https://docs.api7.ai/ai-gateway/mcp-gateway/traffic-controls.md): apply caller rate limits and use AISIX Cloud budgets that cover the caller API key. * [Guardrails](https://docs.api7.ai/ai-gateway/mcp-gateway/guardrails.md): inspect MCP tool arguments and results. * [Observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md): find MCP traffic in logs, metrics and usage. --- # Connect Cursor to MCP Gateway Cursor can connect to the AISIX MCP endpoint as a remote Streamable HTTP client. It sends an AISIX caller API key, discovers only the tools that key may use, and invokes those tools through the gateway without receiving upstream server credentials. Cursor remains responsible for selecting a model, deciding when to request a tool, obtaining any required approval, and presenting the result. This connection sends MCP tool traffic through AISIX; it does not route Cursor's model requests through the gateway. To route supported Ask-mode model requests separately, see the [Cursor model integration](https://docs.api7.ai/ai-gateway/integrations/coding-agents/cursor.md). This guide connects Cursor to the aggregated `/mcp` endpoint. You can continue from [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md), which grants the caller only `everything__echo`, or use an existing AISIX environment and a safe tool the caller may access. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * Either complete [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md) and retain `AISIX_PROXY` and `AISIX_MCP_KEY`, or obtain an AISIX proxy origin and caller API key from the team that operates the gateway. For an existing environment, export them with those variable names and choose one safe permitted tool for verification. * Install [Cursor](https://www.cursor.com/) with an agent-capable model configured. * Make sure Cursor can reach the AISIX proxy URL. Cursor on the gateway host can use the quickstart address; a remote development environment needs an address it can reach. The example uses workspace configuration so the connection can be reviewed with the project. Do not commit the caller API key. Cursor reads it from the environment instead. ## Configure the Connection[​](#configure-the-connection "Direct link to Configure the Connection") Launch Cursor from the shell where `AISIX_MCP_KEY` is exported. If Cursor is already running and did not inherit this variable, close and relaunch it from that environment before testing the connection. Create `.cursor/mcp.json` in the workspace. Replace the example URL with `$AISIX_PROXY/mcp`: .cursor/mcp.json ``` { "mcpServers": { "aisix": { "url": "https://gateway.example.com/mcp", "headers": { "Authorization": "Bearer ${env:AISIX_MCP_KEY}" } } } } ``` For a connection available in every workspace, add the same object to `~/.cursor/mcp.json` instead. Keep the environment variable available to every Cursor process that uses the configuration. Open **Customize → MCPs**, select `aisix`, and enable the workspace source if it is disabled. Confirm that the local environment is connected and the selected permitted tool appears. For the Everything fixture, only `everything__echo` should appear. ## Verify a Tool Call[​](#verify-a-tool-call "Direct link to Verify a Tool Call") The prompt below uses the Everything fixture from the setup guide. For an existing MCP server, substitute a permitted tool name, valid arguments, and an expected result. In Cursor agent chat, ask it to use the exact tool instead of relying on automatic tool selection: ``` Use the MCP tool everything__echo with the message "hello through AISIX". Return the tool result exactly. ``` Review and approve the invocation if Cursor requests confirmation. With the Everything fixture, the result should be: ``` Echo: hello through AISIX ``` Confirm the complete path: * Cursor shows the permitted tool and does not show tools excluded by the caller's effective grant. For the fixture, only `everything__echo` should appear. * [AISIX MCP observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md) records a successful `tools/call` for the expected caller API key and server. * Cursor displays the tool result returned through AISIX. Tool discovery proves that the connection and caller grant work. It does not prove that the model will select a tool or that Cursor's approval policy permits execution, so retain the explicit tool-call test. ## Troubleshoot Cursor[​](#troubleshoot-cursor "Direct link to Troubleshoot Cursor") | Symptom | Check | | ----------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Cursor is disconnected | Confirm the URL ends in `/mcp`, the gateway is reachable from Cursor, and `AISIX_MCP_KEY` is present in the environment that launched it. | | The workspace source is disabled | Open **Customize → MCPs**, select `aisix`, and enable the workspace source. | | Cursor returns `401` | Confirm `AISIX_MCP_KEY` contains the AISIX caller API key, not an upstream MCP credential. | | The connection succeeds but no tools appear | Follow [Troubleshoot Tool Access](https://docs.api7.ai/ai-gateway/mcp-gateway/overview.md#troubleshoot-tool-access) to check the server and effective grant. Then reload the server in Cursor. | | The tool appears but the agent does not call it | Name the selected tool explicitly, enable it in the MCP tool list, and review Cursor's tool-approval policy. For the fixture, select `everything__echo`. | ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Client Authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/client-authentication.md): use gateway API keys, OAuth sign-in, or anonymous access for suitable trusted networks. * [Control Tool Access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md): grant exact tools, server patterns, or every registered tool. * [Observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md): inspect MCP calls by caller, server, tool, and outcome. --- # Guardrails AISIX AI Gateway can inspect both sides of MCP tool execution. An input guardrail runs before the upstream call, while an output guardrail runs before the tool result returns to the client. An input block prevents the call from reaching the upstream MCP server. An output block occurs after the upstream call but withholds its result from the client. Both return a failed tool result. MCP tool calls use the same guardrail chain that governs model traffic. Create guardrails in the shared traffic-control section, then attach them at a scope that can apply to MCP traffic. For shared scope and enforcement-mode behavior, see [Guardrail Behavior](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md). ## How Guardrails Apply to MCP[​](#how-guardrails-apply-to-mcp "Direct link to How Guardrails Apply to MCP") The gateway runs guardrails only on `tools/call` requests. The MCP handshake and `tools/list` carry no tool content and are not scanned. For each tool call, AISIX resolves the guardrail chain once and runs both directions through it: * Input: the tool-call arguments are scanned before the call. If a guardrail blocks, AISIX rejects the call and never contacts the upstream server. * Output: the tool result is scanned before it returns to the client. If a guardrail blocks, AISIX withholds the result and returns a failed tool result instead. An MCP tool call has no model, so model-scoped guardrails do not apply to MCP. Environment, MCP server, caller API key, and team scopes can match MCP calls. In both AISIX Cloud and the open-source AISIX gateway, attachments determine a guardrail's scope. An unattached guardrail inspects no traffic, including MCP calls. When no guardrail matches the caller, the tool call runs with no added guardrail latency. ## What Gets Scanned[​](#what-gets-scanned "Direct link to What Gets Scanned") * Input: the arguments object of the `tools/call` request. AISIX feeds the arguments to the same input hook the model path uses. * Output: the decoded text exposed by standard content blocks. This includes each block's `text`, a resource link's `title` and `description`, and an embedded resource's `resource.text`. Resource names and URIs are identifiers rather than content, and base64 `blob` values are not decoded. * Output: every string value in the result's `structuredContent`. A tool can return this machine-readable field alongside `content` without mirroring it into a text block. AISIX scans values rather than field names, which describe the output schema rather than its data. If a result uses a nonstandard shape and none of these fields yields text, guardrails that evaluate one combined text scan the serialized result as a fallback. Guardrails that moderate individual fields instead inspect only the fields listed above. This second group includes Alibaba Cloud AI Guardrails, Bedrock, Lakera, Presidio, and custom scripts. A protocol-level error result with no result payload has no tool output to scan and is allowed through. The gateway inspects MCP tool arguments and results in flight. It does not store them; content capture is a separate surface from guardrail inspection. ## Block Response[​](#block-response "Direct link to Block Response") When a guardrail blocks a tool call or tool result, AISIX returns HTTP `200` with a tool result marked `isError`, not the model path's HTTP `422`: ``` { "jsonrpc": "2.0", "id": 1, "result": { "content": [{ "type": "text", "text": "tool call blocked by content policy (guardrail 'block-secrets')" }], "isError": true } } ``` MCP separates a request that was not valid, which is a JSON-RPC protocol error, from a call that did not succeed, which is `isError` on the result. A policy rejection is the second kind: the request was well formed, so the rejection reaches the calling agent as tool output it can read and act on rather than as a transport failure. The message names the firing guardrail and whether the input arguments or output result was blocked. It never repeats the content that matched. The gateway records the blocked call as a [usage event](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md) with its guardrail-blocked flag set. If AISIX cannot parse a tool result as the expected JSON response, the output failure policy decides whether it can be returned. One output-reading guardrail that fails closed is enough to withhold the result. AISIX returns it only when nothing in the chain reads output or every output reader is fail-open; the latter case records `guardrail_bypassed_reason: unscannable_body`. For the full MCP error format and other MCP status behavior, see [Headers and Error Codes](https://docs.api7.ai/ai-gateway/reference/headers-and-error-codes.md#mcp-errors). ## Scope a Guardrail to One MCP Server[​](#scope-a-guardrail-to-one-mcp-server "Direct link to Scope a Guardrail to One MCP Server") Attach the guardrail with the `mcp_server` scope to inspect only tool calls routed to one registered server. In AISIX Cloud, set `scope_id` to the server ID. In `resources.yaml`, use the server name. AISIX scans both the arguments sent to that server and the results it returns. `mcp_server` is the only scope that selects traffic by destination server. To narrow coverage by caller instead, use an `api_key` attachment for one caller or a `team` attachment for callers on one team. These caller-based scopes also apply to model traffic from the same API key or team. A model scope never applies to MCP because a tool call resolves no model. The server must be available to the environment where the guardrail applies. AISIX Cloud rejects an attachment to a server that is not exposed there, while a resources file fails to load if the attachment names an undefined server. Guarding several servers takes one attachment each; the guardrail still runs once per request. ## Verify Guardrail Blocking[​](#verify-guardrail-blocking "Direct link to Verify Guardrail Blocking") Create a keyword guardrail with a term you can trigger, such as `secret`. Attach it to the caller API key you will use for the test. In AISIX Cloud, use the caller API key ID as the attachment's `scope_id`. In `resources.yaml`, use the key's `display_name`. See [Built-in Keyword Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/keyword.md) for the configuration workflow. Then connect an MCP client with a caller API key that allows a tool, and call that tool with an argument containing the blocked term. The tool call should return HTTP `200` with a result whose `isError` is `true`, and the upstream MCP server is never contacted. A call whose arguments and result are both clean returns normally. Switching the guardrail to monitor mode lets the same call through while still recording the match; see [Use Monitor Mode](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/keyword.md#use-monitor-mode). ## AISIX Cloud Control Plane[​](#aisix-cloud-control-plane "Direct link to AISIX Cloud Control Plane") In AISIX Cloud, create and attach guardrails through the control plane instead of declaring them in a resources file. To inspect MCP tool calls, use a scope that can apply to non-model traffic: the whole environment, a specific MCP server, the caller API key, or the team. Do not use a model-specific scope when the guardrail should inspect MCP traffic. An MCP tool call has no model, so a model-scoped guardrail never runs on it. A gateway that predates MCP-server scoping does not understand the scope and discards the attachment, so the scope you saved does not run there. What that gateway does instead depends on how old it is: one predating attachment-only scoping applies the guardrail environment-wide, running the rule on more traffic than intended, while a current gateway leaves it inspecting nothing. AISIX warns you at save time when any data plane in the environment runs such a version; upgrade those gateways before relying on the narrower scope. ## Next Steps[​](#next-steps "Direct link to Next Steps") You now understand how guardrails inspect MCP tool arguments and results. Use these guides to create guardrails or observe the blocked calls: * [Built-in Keyword Guardrails](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/keyword.md): create a guardrail and choose its enforcement mode. * [Guardrail Behavior](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md): review scope matching, enforcement modes, and remote-guardrail failure handling. * [Observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md): see blocked tool calls in usage events and metrics. --- # Observability MCP tool calls use the same telemetry pipelines as model traffic. Usage-event fields identify the caller, upstream server, tool, and outcome, while Prometheus labels let you separate MCP traffic from model traffic. Use these signals to measure tool-call volume and monitor failures. Rate limits and guardrail blocks appear in the same observability tools you already use for model traffic. In AISIX Cloud, budget rejections appear there as well. ## Usage Events[​](#usage-events "Direct link to Usage Events") AISIX emits one usage event only after a request enters `tools/call` usage accounting. Requests rejected before that point do not produce an event. These include authentication failures, an unknown scoped server, an unreadable or oversized body, invalid JSON or a `params` shape AISIX cannot parse, and an unsupported protocol version. Once a request enters that accounting path, AISIX records the call even if the `params` object or tool name is missing. It also records calls rejected later by access control, a rate limit, a guardrail, an AISIX Cloud budget, or the MCP protocol handler. A failure while AISIX later reads or buffers the MCP response body can return before the event is emitted. Usage telemetry therefore covers the listed recognized outcomes, not every recognized tool call. It does not cover MCP handshake or discovery methods, or requests rejected before tool-call usage accounting. The event identifies the caller, server, tool, outcome, and timing: | Field | Value | | --------------------------- | ----------------------------------------------------------- | | `inbound_protocol` | `mcp` | | `mcp_server_name` | The registered server the tool belongs to. | | `mcp_tool_name` | The upstream tool that was called. | | `api_key_id` | The caller API key that made the call. | | `status_code` | The call's outcome status. | | `upstream_latency_ms` | Time spent on the upstream tool call. | | `downstream_latency_ms` | Total time the caller waited for the tool call. | | `guardrail_blocked` | `true` when a guardrail blocked the call's input or output. | | `request_id`, `occurred_at` | Correlation id and timestamp. | MCP tool calls do not carry model tokens, so token and cost fields remain zero. Use `mcp_server_name` and `mcp_tool_name` for per-tool call-volume attribution rather than token or spend analytics. MCP currently makes a single upstream attempt that spans the request, so `upstream_latency_ms` and `downstream_latency_ms` report the same duration. The separate fields keep MCP records consistent with other gateway traffic, where retries and gateway processing can make the two values differ. The event records the server name, tool name, and outcome. It does not include tool arguments or tool results. MCP content capture is a separate surface from usage telemetry. MCP usage events follow the same delivery paths as model usage events. Any configured observability exporter receives them, so MCP traffic appears alongside the rest of your gateway traffic. In AISIX Cloud, they also flow to the control plane's usage sink. ## Metrics[​](#metrics "Direct link to Metrics") MCP requests appear in the gateway's Prometheus metrics with labels that distinguish them from model traffic. Use the labels below to filter the relevant metric series: | Goal | Metric | Filter | | ------------------------------- | ---------------------------------- | -------------------------------------------- | | Track active MCP requests. | `aisix_proxy_in_flight_requests` | `inbound_protocol="mcp"` | | Check MCP usage-event emission. | `aisix_usage_events_emitted_total` | `handler="mcp"` and `inbound_protocol="mcp"` | Metrics are exposed on `GET /metrics` through the dedicated metrics listener. For the full metric catalog and label semantics, see [Metrics Reference](https://docs.api7.ai/ai-gateway/reference/metrics.md). ## Verify Metrics[​](#verify-metrics "Direct link to Verify Metrics") To verify that MCP metrics are emitted, send one MCP tool call through the gateway, then scrape the dedicated metrics listener. The example below uses the default listener address and path. If your startup configuration sets a different `observability.metrics.prometheus.addr`, use that address instead. Metric families register on first observation, so the MCP series appears only after a tool call is recorded: ``` curl -sS "http://127.0.0.1:9090/metrics" | grep 'inbound_protocol="mcp"' ``` The output should include metric samples with these labels: | Metric | Label | | ---------------------------------- | -------------------------------------------- | | `aisix_proxy_in_flight_requests` | `inbound_protocol="mcp"` | | `aisix_usage_events_emitted_total` | `handler="mcp"` and `inbound_protocol="mcp"` | ## Next Steps[​](#next-steps "Direct link to Next Steps") You now know where MCP tool calls appear in usage events and metrics. Use these guides to review the full metric catalog or adjust the traffic that produces those signals: * [Metrics Reference](https://docs.api7.ai/ai-gateway/reference/metrics.md): review the full metric catalog and label semantics. * [Rate limits and budgets](https://docs.api7.ai/ai-gateway/mcp-gateway/traffic-controls.md): apply request and concurrency limits, and configure AISIX Cloud budgets for MCP tool calls. * [Guardrails](https://docs.api7.ai/ai-gateway/mcp-gateway/guardrails.md): inspect MCP tool arguments and results. --- # Expose a REST API as MCP Tools An MCP server registry entry can be backed by a plain REST API instead of an upstream MCP server. Register the API's OpenAPI 3.x document with `type` set to `openapi`, and AISIX generates one MCP tool per operation. A `tools/call` executes as an HTTP request against the API's base URL, with the gateway-held credential attached when configured. The MCP caller never receives the credential, and the API needs no MCP server of its own. This turns existing internal services, such as an ERP, inventory system, or payroll API, into agent-callable tools. [Tool access control](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md), [traffic controls](https://docs.api7.ai/ai-gateway/mcp-gateway/traffic-controls.md), [guardrails](https://docs.api7.ai/ai-gateway/mcp-gateway/guardrails.md), and [observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md) apply with both management paths. AISIX Cloud also provides [server review](https://docs.api7.ai/ai-gateway/mcp-gateway/server-review.md) and shared [access policies](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * For AISIX Cloud, complete the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), then keep its shell and attached gateway running. The workflow reuses its environment, admin token, caller API key, and gateway URL. To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For an open-source AISIX gateway, complete [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md), then remain in its shell and working directory without cleaning up its temporary Docker network. * [cURL](https://curl.se/) and [jq](https://jqlang.org/) for the AISIX Cloud example. ## How Tools Are Generated[​](#how-tools-are-generated "Direct link to How Tools Are Generated") AISIX walks the document's `paths` and generates one tool per operation for the `get`, `post`, `put`, `delete`, and `patch` methods: * Tool name: the operation's `operationId`, lowercased, with any character outside `a-z`, `0-9`, `_`, and `-` replaced by `_`, capped at 128 characters. An operation without an `operationId` is named `_` under the same rules. Tools are exposed to callers as `__`, like every MCP tool. * Input schema: each `path` and `query` parameter becomes a property with its type, description, `enum` values, and `required` flag. A JSON request body becomes a single `body` object property, marked required when the spec says so. Local `$ref`s, including referenced component schemas inside the body, are resolved so the agent sees the real shape. Header and cookie parameters are not exposed because upstream headers belong to the gateway, not the caller. * Skipped operations: an operation whose request body has no `application/json` variant, such as a `multipart/form-data` file upload, is skipped rather than generating a tool that cannot succeed. Validation timing differs by management path. AISIX Cloud validates the document during registration. It rejects a document that cannot be parsed, a Swagger 2.0 document, a document with no tool-generatable operations, or `operationId`s that collide after normalization. The create response returns the generated names in `tool_names`. With `resources.yaml`, `aisix validate` checks the resource shape, including that `spec` is a mapping and is not a Swagger document. The gateway generates the tools when a client lists or calls them. If the document has no usable `paths` object, that server contributes no tools to the aggregated list and the gateway logs the error. Normalized name collisions receive `_2`, `_3`, and subsequent suffixes so every operation remains reachable. List tools after loading the file to verify the generated surface. ## Register a REST API[​](#register-a-rest-api "Direct link to Register a REST API") The registration shape differs between AISIX Cloud and an open-source AISIX gateway. In both cases, `url` is the REST API base URL, and generated calls are issued against ``. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Provide the document in one of two ways: * `spec_content`: the document itself, as JSON or YAML text. Use this when the control plane cannot reach the API's network. * `spec_url`: a URL the control plane fetches once during registration. The fetched document is validated, normalized, and stored; the data plane never re-fetches it, so the tool set only changes when you update the registry entry. By default, URLs that resolve to non-public addresses are refused. An On-Premises deployment can allow them by setting `AISIX_CLOUD_MCP_SPEC_ALLOW_PRIVATE_URLS=true` on the control plane. Otherwise, paste the document instead. AISIX Cloud caps OpenAPI documents at 1 MiB. Export the control-plane address, admin token, and environment ID: ``` # AISIX_CP includes /api and has no trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Register an HTTPBin endpoint and its OpenAPI document. This public endpoint makes the generated tool call reproducible without an internal API or upstream credential: ``` MCP_SERVER_RESPONSE=$(curl -fsS -X POST "$AISIX_CP/mcp_servers" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "httpbin", "type": "openapi", "url": "https://httpbin.org", "spec_content": "{\"openapi\":\"3.0.0\",\"info\":{\"title\":\"HTTPBin\",\"version\":\"1.0.0\"},\"paths\":{\"/anything\":{\"get\":{\"operationId\":\"inspectRequest\",\"responses\":{\"200\":{\"description\":\"OK\"}}}}}}", "auth_type": "none", "allowed_environments": ["'$ENV_ID'"] }') echo "$MCP_SERVER_RESPONSE" | jq export MCP_SERVER_ID=$(echo "$MCP_SERVER_RESPONSE" | jq -er '.mcp_server.id') ``` The response includes the generated `tool_names`: ``` { "mcp_server": { "id": "6f64f080-17d7-44d9-b995-6a353e71f6bc", "name": "httpbin", "type": "openapi", "url": "https://httpbin.org", "tool_names": ["inspectrequest"], "approval_status": "approved" } } ``` Grant the generated tool to the caller API key created by the quickstart: ``` curl -fsS -X PATCH \ "$AISIX_CP/environments/$ENV_ID/api_keys/$API_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"mcp_access":{"allow":["httpbin__inspectrequest"]}}' | jq export AISIX_MCP_KEY="$AISIX_API_KEY" ``` The partial update retains the key's model access. If environment or team MCP access policies also apply to the key, those layers must allow `httpbin__inspectrequest` as well. Send the `initialize` request and `notifications/initialized` notification from [Verify MCP Tool Access with HTTP](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md#verify-mcp-tool-access-with-http). Then poll until the projected tool appears: ``` for attempt in $(seq 1 45); do TOOLS_RESPONSE=$(curl -fsS -X POST "$AISIX_PROXY/mcp" \ -H "Authorization: Bearer $AISIX_MCP_KEY" \ -H "MCP-Protocol-Version: 2025-11-25" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{ "jsonrpc": "2.0", "id": 2, "method": "tools/list", "params": {} }') || true if echo "$TOOLS_RESPONSE" | jq -e \ '.result.tools | any(.name == "httpbin__inspectrequest")' >/dev/null 2>&1; then break fi sleep 2 done echo "$TOOLS_RESPONSE" | jq -e \ '.result.tools | any(.name == "httpbin__inspectrequest")' ``` The final command prints `true`. Call the generated tool and verify the URL HTTPBin received: ``` curl -fsS -X POST "$AISIX_PROXY/mcp" \ -H "Authorization: Bearer $AISIX_MCP_KEY" \ -H "MCP-Protocol-Version: 2025-11-25" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{ "jsonrpc": "2.0", "id": 3, "method": "tools/call", "params": { "name": "httpbin__inspectrequest", "arguments": {} } }' | jq -e \ '.result.content[] | select(.text | fromjson | .url == "https://httpbin.org/anything")' ``` The command prints the matching tool-result content block. In the dashboard, choose **REST API (OpenAPI)** when registering an MCP server. Paste the document, provide its URL, or pick a local `.json` or `.yaml` file with **Choose file…**. A picked file is read into the editor, so you can review and adjust it before saving. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Set `type: openapi` on an `mcp_servers` entry and provide the document as a nested mapping under `spec`. The resources file does not accept the AISIX Cloud write fields `spec_content` or `spec_url`. Start an HTTP fixture on the temporary Docker network created in [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md). The server exposes an empty directory over HTTP and requires no packages on the host: ``` docker run -d --name aisix-openapi-fixture \ --network aisix-mcp \ python:3.13-alpine \ python3 -m http.server 8081 --bind 0.0.0.0 --directory /tmp ``` In the complete file from [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md), replace `quickstart-caller` and add `fixture` to `mcp_servers`. Keep the other resources unchanged: resources.yaml (API key and MCP server) ``` api_keys: - display_name: quickstart-caller key_env: CALLER_API_KEY allowed_models: - gpt-4o-mini mcp_access: { allow: ["fixture__list_directory"] } mcp_servers: - name: fixture type: openapi url: http://aisix-openapi-fixture:8081 auth_type: none spec: openapi: 3.0.0 info: title: Local directory API version: 1.0.0 paths: /: get: operationId: list_directory summary: List the fixture directory responses: "200": description: Directory listing returned successfully ``` This example reuses `CALLER_API_KEY`, which is already present in the running gateway from the open-source quickstart. Validate and reload the complete resources file. Use the initialization request from [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md#verify-mcp-tool-access-with-http), then call the generated tool: ``` curl -sS -X POST "$AISIX_PROXY/mcp" \ -H "Authorization: Bearer $AISIX_MCP_KEY" \ -H "MCP-Protocol-Version: 2025-11-25" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{ "jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": { "name": "fixture__list_directory", "arguments": {} } }' | jq -e \ '.result.content[] | select(.text | contains("Directory listing for /"))' ``` The command prints the matching tool-result content block. Remove the fixture container when you finish: ``` docker rm -f aisix-openapi-fixture ``` ## Review the Generated Tools in AISIX Cloud[​](#review-the-generated-tools-in-aisix-cloud "Direct link to Review the Generated Tools in AISIX Cloud") Each registered server's card lists the first few generated tool names and links to a **tools page** for that server, which lists every tool with: * its `__` name, which is the form to copy into an API key's [`mcp_access` block](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md) or an [access policy](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md) pattern; * the HTTP operation it calls, such as `GET /items/{id}`; * its description. The same listing is available from the API: ``` curl -sS "$AISIX_CP/mcp_servers/$MCP_SERVER_ID/tools" \ -H "Authorization: Bearer $AISIX_TOKEN" ``` ``` { "data": [ { "name": "inspectrequest", "namespaced_name": "httpbin__inspectrequest", "method": "GET", "path": "/anything", "description": "GET /anything" } ] } ``` The listing is derived from the stored document, so it always matches the tools the gateway serves. It applies to `type: openapi` servers only. An upstream MCP server's tools live on the upstream, and the endpoint returns `400` for one. ## Authenticate to the REST API[​](#authenticate-to-the-rest-api "Direct link to Authenticate to the REST API") The [upstream authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/upstream-authentication.md) modes apply as-is; the credential is attached to every generated tool call: * `bearer`: `Authorization: Bearer `. * `api_key`: the key from `secret`, sent as the `x-api-key` header by default. REST APIs often expect a custom header: set `api_key_header` (for example `X-ERP-Key`) to override the header name. This field exists only on `openapi` servers with `auth_type: api_key`. * `oauth2`: AISIX mints an access token from the configured client credentials and sends it as a bearer, with the same token caching as MCP upstreams. Redirects are never followed on generated tool calls, so the credential cannot be re-sent to a host you did not configure. ## Call Results and Errors[​](#call-results-and-errors "Direct link to Call Results and Errors") A successful response's body is returned as the tool result text. A non-2xx response returns a tool-level error result (`isError: true`) carrying `HTTP ` and the response body, so the agent can see and react to the failure. A missing required path parameter is reported the same readable way. AISIX also rejects a path value that contains `/` or `\`, or that equals `.` or `..`, to keep the request on the configured path. ## Update the Document[​](#update-the-document "Direct link to Update the Document") Replacing the document regenerates the tool list. For an open-source AISIX gateway, replace the nested `spec`, validate the complete resources file, and reload the gateway. A rejected reload leaves the previous tool surface active. After a successful reload, list the tools to verify that the updated document produces the expected surface. In AISIX Cloud, provide `spec_content` or `spec_url` on an update call, or use the **Replace OpenAPI document** editor in the dashboard. The control plane begins projecting the new tool surface when the call returns. The caller holds the permission that approves servers, so the replacement counts as its review and updates the review timestamp. User-session actions also record the reviewing user; admin-token actions do not carry a user ID. A role holding only `write` on `mcp_server_submissions` stages the replacement instead, and the current tools keep serving until a reviewer approves it. Changing `api_key_header` works the same way. Re-uploading a document that normalizes to the identical stored version is not treated as a change at all. In AISIX Cloud, the server's `type` is fixed at creation: to switch between an MCP upstream and an OpenAPI backing, delete the entry and register a new one. In `resources.yaml`, change the entry and reload the file; the newly validated configuration replaces the previous runtime entry. ## Next Steps[​](#next-steps "Direct link to Next Steps") You can now expose a REST API as MCP tools through either management path. Use these guides to secure and govern the generated tools: * [Upstream Authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/upstream-authentication.md): configure the credential AISIX sends to the REST API. * [Control Tool Access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md): choose which generated tools each caller API key may list and call. * [Review and Approve MCP Servers](https://docs.api7.ai/ai-gateway/mcp-gateway/server-review.md): review OpenAPI-backed servers and document changes before publishing them through AISIX Cloud. --- # MCP Gateway Overview The AISIX gateway fronts registered Model Context Protocol (MCP) tool sources through an aggregated `/mcp` endpoint. A source can be an upstream MCP server that uses Streamable HTTP or a REST API described by an OpenAPI document. See [Expose a REST API as MCP Tools](https://docs.api7.ai/ai-gateway/mcp-gateway/openapi-servers.md) for the REST API path. MCP clients and agents connect to `/mcp` with an AISIX caller API key, discover the tools that key can use, and call those tools without receiving the upstream MCP server credential. This gives tool traffic the same authentication, access-control, and telemetry boundary as model and A2A traffic. One caller API key can govern the models a caller may use, the MCP tools it may call, and the A2A agents it may reach. AISIX authenticates every MCP request and filters tool discovery by the caller's effective tool grant. For `tools/call`, AISIX also applies rate limits and guardrails, checks applicable AISIX Cloud budgets, routes the call with the configured upstream credential, and records usage telemetry. ## How the MCP Gateway Works[​](#how-the-mcp-gateway-works "Direct link to How the MCP Gateway Works") Each MCP server has a `name`. In AISIX Cloud, the server is registered at the organization level and exposed to selected environments. In an open-source AISIX gateway, it is normally declared in `resources.yaml`. AISIX aggregates tools from enabled servers and exposes each tool under a prefixed name. AISIX uses two underscores to separate the registered server name from the upstream tool name. For example, `github__create_issue` routes to the registered MCP server named `github` and calls its upstream tool named `create_issue`. Across both management paths, a server name cannot contain the reserved `__` separator or end with an underscore. A single underscore can appear inside a name, such as `internal_tools`. AISIX Cloud additionally limits names to 56 letters, digits, underscores, dots, or hyphens, with a letter or digit at each end. ## Get Started[​](#get-started "Direct link to Get Started") Follow [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md) to register an upstream server through AISIX Cloud or `resources.yaml` and grant one tool to a caller API key. The guide isolates the gateway path with allowed and denied HTTP checks. Then connect [Cursor](https://docs.api7.ai/ai-gateway/mcp-gateway/cursor.md) or [VS Code](https://docs.api7.ai/ai-gateway/mcp-gateway/vscode.md) to verify discovery, any required approval, and execution in a real MCP client. Both management paths configure the same gateway runtime and MCP endpoints. ## Client Connection[​](#client-connection "Direct link to Client Connection") MCP clients connect to the AISIX proxy listener over Streamable HTTP. Use `/mcp` for the aggregated tool surface: | Setting | Value | | -------------- | ---------------------------------------- | | Server URL | `/mcp` | | Transport | Streamable HTTP | | Request header | `Authorization: Bearer ` | The caller API key controls which tools the client can discover and call. The client connects only to AISIX; it does not receive upstream server URLs or credentials. ## Protocol Version Support[​](#protocol-version-support "Direct link to Protocol Version Support") The AISIX gateway serves both current MCP protocol generations on `/mcp` and `/mcp/{server}` and negotiates with each client automatically, so no AISIX-specific client configuration is needed: | MCP protocol revision | Client support | Notes | | --------------------- | -------------- | --------------------------------------------------------------------------------------------------------- | | `2026-07-28` | Supported | Stateless revision: handshake-free startup through `server/discover`, with per-request protocol metadata. | | `2025-11-25` | Supported | `initialize` handshake. | | `2025-06-18` | Supported | `initialize` handshake. | | `2025-03-26` | Supported | `initialize` handshake. | | `2024-11-05` | Not supported | HTTP+SSE transport generation; the MCP endpoints serve Streamable HTTP only. | Two version signals are involved, and the gateway handles each on its own terms. During the `initialize` handshake, the gateway echoes a supported `protocolVersion` from the request and answers `2025-11-25` for an unsupported one. Separately, the `MCP-Protocol-Version` HTTP header is optional on requests outside the handshake: an absent header is accepted (and treated as `2025-03-26`, per the specification's compatibility rule), while a header naming an unsupported revision is rejected with HTTP `400` and a JSON-RPC error envelope that lists the supported revisions. A client on the `2026-07-28` lifecycle can start with `server/discover` instead of a handshake. Serving is stateless on every generation: the gateway issues no `Mcp-Session-Id`, so MCP requests need no session affinity across gateway replicas. ### Upstream Protocol Selection[​](#upstream-protocol-selection "Direct link to Upstream Protocol Selection") The client-facing protocol and the upstream session are independent: the gateway terminates the client's protocol at the MCP endpoint and opens its own session to each registered server of `type: mcp`. (An `openapi` source has no upstream MCP session, so this setting does not apply to it.) The revision of that upstream session is a per-server setting, `protocol_version`: | `protocol_version` | Upstream session behavior | | ------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Not set (default) | The gateway opens the session with the `initialize` handshake, which negotiates a supported 2025 Streamable HTTP revision. This also works with `2026-07-28` servers that continue to answer `initialize`. | | `"2026-07-28"` | The gateway opens the session with handshake-free `server/discover`. Required for servers that no longer answer `initialize`. | In `resources.yaml`, omit `protocol_version` to keep the default lifecycle. In AISIX Cloud, an existing pin is removed by setting `protocol_version` to `null` — omitting the field in an update leaves the pin in place. For the configuration steps on both management paths, see [Pin the MCP Protocol Revision](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md#pin-the-mcp-protocol-revision). The selection is explicit: the gateway uses only the configured lifecycle and never probes or silently downgrades across protocol generations. When the configured lifecycle is incompatible with the server, `tools/list` logs the failure and omits that server's tools from the aggregated list, and a `tools/call` to it returns an upstream failure. The downstream and upstream selections stay independent of each other. The gateway does not forward the caller's `MCP-Protocol-Version` or `Mcp-Session-Id` headers to the upstream server, and forwards the caller's `Authorization` only when the server's `forward_client_headers` names it exactly. It opens an independent upstream session using the server's registered authentication and protocol settings; a stateful upstream can mint its own session identifier for that session, and tool names, arguments, and results cross the boundary as the call requires. Conformance The gateway's continuous integration runs the applicable tools-surface scenarios from the official MCP conformance suite against the served gateway and bridge chain, as a merge-blocking check. ## Govern MCP Tool Calls[​](#govern-mcp-tool-calls "Direct link to Govern MCP Tool Calls") MCP tool calls use the same caller API key identity and telemetry pipeline as model requests. They share the key's request and concurrency limits with model traffic. Applicable guardrails and, in AISIX Cloud, budgets that cover the caller also apply. MCP-specific controls include per-server rate limits and tool authorization: each key can carry its own grant, and AISIX Cloud can add environment and team layers. Use these guides to refine the MCP path: * [Control Tool Access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md): scope each caller API key to the tools it may list and call. * [Rate Limits and Budgets](https://docs.api7.ai/ai-gateway/mcp-gateway/traffic-controls.md): apply caller API key request and concurrency limits, and use AISIX Cloud budgets for `tools/call` requests. * [Guardrails](https://docs.api7.ai/ai-gateway/mcp-gateway/guardrails.md): inspect MCP tool arguments and results. * [Observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md): review the usage events and metrics emitted by MCP tool calls. AISIX Cloud also provides shared [MCP access policies](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md) and a [server review and approval workflow](https://docs.api7.ai/ai-gateway/mcp-gateway/server-review.md). ## Per-Server Endpoints[​](#per-server-endpoints "Direct link to Per-Server Endpoints") Use `/mcp/{server}` when a client expects a separate URL for each registered MCP server. The endpoint presents only the named server. `initialize` reports its registered name, and `tools/list` normally returns permitted tools under their original upstream names. `tools/call` accepts both an original name such as `create_issue` and its aggregated form such as `github__create_issue`. Authentication, tool access, rate limits, guardrails, usage telemetry, and AISIX Cloud budgets work the same way as on `/mcp`. Access grants and per-server rate limits keep their `{server}__{tool}` identity across both endpoints, so using both URL forms does not create a second allowance. An unknown or disabled server returns 404 after caller authentication. For tool-name collision behavior and endpoint errors, see [Proxy API Reference](https://docs.api7.ai/ai-gateway/reference/proxy-api.md#mcp-gateway). If an existing client uses another URL shape, such as `/mcp-servers/{server}/mcp`, map it to `/mcp/{server}` with [URL rewriting](https://docs.api7.ai/ai-gateway/deployment/url-rewriting.md#example-serve-per-server-mcp-urls). ## Troubleshoot Tool Access[​](#troubleshoot-tool-access "Direct link to Troubleshoot Tool Access") If a client cannot see or call a tool, check these items: * The MCP server resource is `enabled`. * In AISIX Cloud, the server is `approved` and its `allowed_environments` includes the caller key's environment. * The upstream MCP server is reachable from the AISIX gateway, and the configured upstream authentication is valid. See [Upstream Authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/upstream-authentication.md). * The caller API key's effective tool grant covers the prefixed tool name. That grant is the intersection of every layer that applies: the key's own `mcp_access` block and, in AISIX Cloud, the environment and team [MCP access policies](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md). * The MCP client sends the caller API key to AISIX, not the upstream MCP credential. For MCP error response behavior, see [Headers and Error Codes](https://docs.api7.ai/ai-gateway/reference/headers-and-error-codes.md#mcp-errors). For endpoint-level behavior, see [Proxy API Reference](https://docs.api7.ai/ai-gateway/reference/proxy-api.md#mcp-gateway). ## Next Steps[​](#next-steps "Direct link to Next Steps") Use these guides to configure how MCP traffic is handled: * [Connect Cursor to MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/cursor.md): configure Cursor and verify a complete tool call. * [Connect VS Code to MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/vscode.md): configure VS Code and verify a complete tool call. * [Expose a REST API as MCP Tools](https://docs.api7.ai/ai-gateway/mcp-gateway/openapi-servers.md): generate tools from an OpenAPI document. * [Upstream Authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/upstream-authentication.md): configure how AISIX authenticates to upstream servers. * [Proxy API Reference](https://docs.api7.ai/ai-gateway/reference/proxy-api.md#mcp-gateway): review the MCP endpoint contract and constraints. --- # Review and Approve MCP Servers [AISIX Cloud Only](https://docs.api7.ai/ai-gateway/cloud/overview.md)Available with AISIX Cloud An MCP server registry entry gives agents a gateway-managed tool surface backed by an upstream MCP server or REST API. Publishing the entry lets agents invoke remote operations with caller-supplied arguments. AISIX Cloud therefore separates registration from publication: every MCP server carries an `approval_status`, and only an approved server is sent to the gateways. A server that is waiting for review, or that was rejected, has no presence on the gateway at all. It does not appear in `tools/list` and its tools cannot be called, even by a caller API key that explicitly allowlists them. Approval is required for publication; the `enabled` flag then controls whether the approved server can serve tools. This separation lets a member propose an MCP server without being able to publish it. It also prevents a different configuration from inheriting an approval: a change either comes from someone who can approve it, or waits for review. A live server keeps serving its approved configuration while a proposed change waits, so the review causes no downtime. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * An AISIX Cloud organization and environment. * An admin token with write scope for the review and approval examples. * [cURL](https://curl.se/) and [jq](https://jqlang.github.io/jq/). For On-Premises, the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md) creates the organization, environment, and write-scoped admin token. To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). Export the values used by the examples: ``` # AISIX_CP includes /api and has no trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` ## How the Review Workflow Works[​](#how-the-review-workflow-works "Direct link to How the Review Workflow Works") | `approval_status` | On the gateways | Reached by | | ----------------- | ----------------------------------- | --------------------------------------------------------- | | `pending_review` | Not projected | Submitting a server, or revising a proposal. | | `approved` | Projected to `allowed_environments` | Approving it, or registering it with `POST /mcp_servers`. | | `rejected` | Not projected | Rejecting it. | A change proposed to a server that is already live does not move it out of `approved`. The change waits in a separate `pending_change` field while every other field keeps describing what the gateways are serving. See [Submit a Change to a Live Server](#submit-a-change-to-a-live-server). Servers registered before this feature was introduced are approved, so upgrading does not interrupt tool traffic. ## Configure Reviewer Permissions[​](#configure-reviewer-permissions "Direct link to Configure Reviewer Permissions") Approving is the same permission as registering: `write` on `mcp_servers`. Owners and admins hold it. A user with this permission can already publish a server directly, so a separate review would not add control. This also lets an organization with a single administrator approve its own submissions. The separable half is a second permission, `write` on `mcp_server_submissions`. A [custom role](https://docs.api7.ai/ai-gateway/cloud/custom-roles.md) that holds it without `write` on `mcp_servers` can propose servers, revise proposals, and propose changes to servers that are already live, but cannot publish anything. That is the role to give a team that integrates its own tooling. ## Submit a Server for Review[​](#submit-a-server-for-review "Direct link to Submit a Server for Review") Submit the server with `POST /mcp_server_submissions`. The body is the same as `POST /mcp_servers`: ``` curl -sS -X POST "$AISIX_CP/mcp_server_submissions" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "runbooks", "url": "https://mcp.example.com/runbooks", "auth_type": "none", "allowed_environments": ["'"$ENV_ID"'"] }' ``` The server is registered but unpublished: ``` { "mcp_server": { "id": "6f64f080-17d7-44d9-b995-6a353e71f6bc", "name": "runbooks", "url": "https://mcp.example.com/runbooks", "enabled": true, "allowed_environments": ["YOUR_ENVIRONMENT_ID"], "approval_status": "pending_review", "submitted_at": "2026-07-29T09:30:00Z" } } ``` ❶ `enabled` and `allowed_environments` describe the configuration that will take effect after approval. ❷ `pending_review` keeps the server unpublished, so no gateway can reach it yet. In the dashboard, the MCP servers page shows a **Pending review** badge on the row and a band at the top of the page listing how many servers are waiting. ## Review a Pending Server[​](#review-a-pending-server "Direct link to Review a Pending Server") List what is waiting by reading `approval_status` from `GET /mcp_servers`: ``` curl -sS "$AISIX_CP/mcp_servers" \ -H "Authorization: Bearer $AISIX_TOKEN" \ | jq '.data[] | select(.approval_status == "pending_review") | {id, name, url}' ``` Review the server before approving it: * Confirm that the name is not a near-copy of an existing server's name and that the URL identifies the intended upstream. * Check `allowed_environments` and `enabled` to confirm where the server will be available after approval. * Check the authentication mode and verify the credential's provenance through your submission process. Stored credentials are write-only and cannot be inspected from the API or dashboard; replace the credential through the approver route if its provenance is uncertain. * For an OpenAPI-backed server, review the generated tool names and the stored document. See [Review the Generated Tools in AISIX Cloud](https://docs.api7.ai/ai-gateway/mcp-gateway/openapi-servers.md#review-the-generated-tools-in-aisix-cloud). Approve it, and it is published to the environments in `allowed_environments`: ``` export MCP_SERVER_ID="YOUR_MCP_SERVER_ID" curl -sS -X POST "$AISIX_CP/mcp_servers/$MCP_SERVER_ID/approve" \ -H "Authorization: Bearer $AISIX_TOKEN" ``` Reject it with a reason instead. The note is returned on the server, so whoever proposed it can see what to change: ``` curl -sS -X POST "$AISIX_CP/mcp_servers/$MCP_SERVER_ID/reject" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"review_notes": "Not on the trusted registry. Use the internal mirror."}' ``` Approving an already-approved server, or rejecting an already-rejected one, returns `400`. A server carrying a staged change is the exception; see [Submit a Change to a Live Server](#submit-a-change-to-a-live-server). ## Revise a Pending Submission[​](#revise-a-pending-submission "Direct link to Revise a Pending Submission") `PATCH /mcp_server_submissions/{id}` corrects a submission and puts it back in the queue: ``` curl -sS -X PATCH "$AISIX_CP/mcp_server_submissions/$MCP_SERVER_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"url": "https://mcp.internal.example.com/runbooks"}' ``` On a server that is not published yet, the change is applied to the row and the server stays in the queue. On a server that is already live, the same call stages the change instead of applying it. ## Submit a Change to a Live Server[​](#submit-a-change-to-a-live-server "Direct link to Submit a Change to a Live Server") A role holding only `write` on `mcp_server_submissions` cannot publish, so its patch to a live server cannot take effect on its own. It must not take the server down while it waits either, so `PATCH /mcp_server_submissions/{id}` stages it: ``` curl -sS -X PATCH "$AISIX_CP/mcp_server_submissions/$MCP_SERVER_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"url": "https://mcp.internal.example.com/runbooks"}' ``` The response is the server as it still stands, with the proposal alongside it: ``` { "mcp_server": { "id": "6f64f080-17d7-44d9-b995-6a353e71f6bc", "name": "runbooks", "url": "https://mcp.example.com/runbooks", "approval_status": "approved", "pending_change": { "changes": { "url": "https://mcp.internal.example.com/runbooks" }, "submitted_by": "YOUR_USER_ID", "submitted_at": "2026-07-30T09:30:00Z" } } } ``` ❶ Fields outside `pending_change` describe what the gateways are serving. Agents keep listing and calling the server's current tools throughout the review window. ❷ `pending_change` contains the proposed replacement values and submission metadata. A proposal is validated when it is made, so an invalid change is refused at that point rather than at approval time. A proposed credential is encrypted at rest exactly like the live one and is never returned; `pending_change.secret_set` reports that the proposal sets one. Only one proposal is staged at a time, so patching again replaces it. ## Review a Change to a Live Server[​](#review-a-change-to-a-live-server "Direct link to Review a Change to a Live Server") Approving applies the staged change and publishes it in a single step: ``` curl -sS -X POST "$AISIX_CP/mcp_servers/$MCP_SERVER_ID/approve" \ -H "Authorization: Bearer $AISIX_TOKEN" ``` The environments a proposal names are checked again at that point, so a change proposed for an environment that was deleted while it waited is refused instead of partly applied. Rejecting discards the proposal. The server stays `approved` on the configuration it was already serving, and nothing changes on the gateways: ``` curl -sS -X POST "$AISIX_CP/mcp_servers/$MCP_SERVER_ID/reject" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"review_notes": "Point it at the internal mirror instead."}' ``` In the dashboard, a row with a staged change shows which fields the change touches. The status filter's **Awaiting review** option lists both unpublished submissions and live servers with a change waiting. ## Update a Live Server as an Approver[​](#update-a-live-server-as-an-approver "Direct link to Update a Live Server as an Approver") `PATCH /mcp_servers/{id}` takes `write` on `mcp_servers`, the same permission that approves servers. An edit made through it is a publication in its own right. The server stays approved, and the control plane projects the new configuration without taking the existing one offline for another review. This lets an approver rotate an upstream credential without first withdrawing the server. Which fields the edit changes determines whether it counts as a new review. Changing any of the following updates `reviewed_at`. Requests made through a user session also set `reviewed_by` to that user; this field is absent for admin-token actions. * `name`: the namespace agents address the tools by * `url`: the upstream itself * `transport` * `auth_type`, `secret`, `client_id`, `token_url`, `scopes`: the credential and where it is presented * `allowed_environments`: where the server is exposed * `spec_content` / `spec_url` and `api_key_header` on [OpenAPI-backed servers](https://docs.api7.ai/ai-gateway/mcp-gateway/openapi-servers.md): the tool surface `enabled` and `timeout_ms` are operational knobs inside an already-reviewed configuration. Changing them is not a new review, so they leave `reviewed_by` and `reviewed_at` alone. Narrowing `allowed_environments` withdraws the server from every environment it no longer names as the asynchronous projection reaches each gateway. A holder of `write` on `mcp_server_submissions` alone cannot use this route; its change is staged for review instead. See [Submit a Change to a Live Server](#submit-a-change-to-a-live-server). To take a live server off the gateways, reject it. ## Revoke an Approval[​](#revoke-an-approval "Direct link to Revoke an Approval") Rejecting an approved server that carries no staged change revokes the server itself: it is withdrawn from every environment it was serving, and callers lose its tools. Use it when an upstream stops being trustworthy. ``` curl -sS -X POST "$AISIX_CP/mcp_servers/$MCP_SERVER_ID/reject" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"review_notes": "Upstream credential compromised."}' ``` Revocation is projected to the gateways asynchronously and needs no gateway restart. To remove the registry entry entirely, delete the server instead. Rejecting a server that does carry a staged change discards the change first, as described above. Revoking such a server takes two calls: the first discards the proposal, the second withdraws the server. ## Audit Trail[​](#audit-trail "Direct link to Audit Trail") Each mutation in this workflow is recorded in the organization audit log with the time and, where applicable, the server's before and after state. User-session actions include the acting user; admin-token actions do not carry a user ID. | Action | Recorded when | | --------- | ------------------------------------------------------------------------------ | | `submit` | A server is proposed for review, or a change is proposed to a live server. | | `create` | A server is registered directly and published. | | `approve` | A server is approved, or a staged change is applied. | | `reject` | A server is rejected, an approval is revoked, or a staged change is discarded. | | `update` | A server's configuration is changed. | | `delete` | A server is removed. | Read them from the audit log in the dashboard, or with `GET /audit_events?resource_type=mcp_server`. ## Next Steps[​](#next-steps "Direct link to Next Steps") You now understand how AISIX Cloud separates MCP server submission from publication. Use these guides to register servers directly or govern caller access: * [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md): register a server directly and verify MCP tool access. * [Control Tool Access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md): choose which tools a caller API key may call. * [Manage MCP Access with Policies](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md): configure environment- and team-level tool grants. --- # Set Up MCP Gateway In this guide, you will register an upstream MCP server and grant one of its tools to a caller API key. You will then verify that the caller can use only the granted tool through `/mcp`. AISIX Cloud and the open-source AISIX gateway configure the same runtime behavior through different management paths. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * For AISIX Cloud, complete the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md), then keep its shell and `aisix-dp` gateway running. * For the open-source AISIX gateway, complete the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md), then remain in its shell and working directory with the `aisix-quickstart` gateway running. * [Docker](https://docs.docker.com/get-docker/), [cURL](https://curl.se/), and [jq](https://jqlang.github.io/jq/). ## Start an MCP Test Server[​](#start-an-mcp-test-server "Direct link to Start an MCP Test Server") The example runs the official MCP [Everything test server](https://github.com/modelcontextprotocol/servers/tree/main/src/everything) in Docker. The server and your existing gateway join a temporary Docker network, so the gateway can reach the server without a restart. Use the Everything server only for local testing: it has no authentication and includes a diagnostic tool that can return its process environment. The container used here receives no secrets. Stop and remove it after finishing the guide. Set the gateway container name for your configuration path. For AISIX Cloud: ``` export AISIX_GATEWAY_CONTAINER="aisix-dp" ``` For the open-source AISIX gateway: ``` export AISIX_GATEWAY_CONTAINER="aisix-quickstart" ``` Create a temporary network and connect the running gateway to it: ``` docker network create aisix-mcp docker network connect aisix-mcp "$AISIX_GATEWAY_CONTAINER" ``` Start the Everything server on the same network: ``` docker run -d --name aisix-mcp-everything \ --network aisix-mcp \ node:22-alpine \ sh -c 'npx -y @modelcontextprotocol/server-everything@2026.7.4 streamableHttp' ``` Wait until the server is ready: ``` for attempt in $(seq 1 120); do docker logs aisix-mcp-everything 2>&1 | grep -q "listening on port 3001" && break sleep 1 done docker logs aisix-mcp-everything 2>&1 | grep "listening on port 3001" ``` The final command prints a log line confirming that port `3001` is ready. From the gateway container, the MCP endpoint is available at `http://aisix-mcp-everything:3001/mcp`. ## Register the Server and Grant a Tool[​](#register-the-server-and-grant-a-tool "Direct link to Register the Server and Grant a Tool") Register the server using the management path for your deployment. Both paths name the server `everything` and grant only its `echo` tool to the existing quickstart caller. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Register the test server and allow it in the quickstart environment: ``` MCP_SERVER_RESPONSE=$(curl -fsS -X POST "$AISIX_CP/mcp_servers" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @- </dev/null 2>&1; then break fi sleep 2 done echo "$TOOLS_RESPONSE" | jq -e \ '.result.tools | map(.name) == ["everything__echo"]' ``` The final command prints `true`. Call the permitted tool: ``` curl -fsS -X POST "$AISIX_PROXY/mcp" \ -H "Authorization: Bearer $AISIX_MCP_KEY" \ -H "MCP-Protocol-Version: 2025-11-25" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{ "jsonrpc": "2.0", "id": 3, "method": "tools/call", "params": { "name": "everything__echo", "arguments": {"message": "hello through AISIX"} } }' | jq -e \ '.result.content[] | select(.text == "Echo: hello through AISIX")' ``` The command prints the matching tool-result content block. To verify that the allowlist is enforced, attempt to call another tool from the same upstream server: ``` curl -fsS -X POST "$AISIX_PROXY/mcp" \ -H "Authorization: Bearer $AISIX_MCP_KEY" \ -H "MCP-Protocol-Version: 2025-11-25" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{ "jsonrpc": "2.0", "id": 4, "method": "tools/call", "params": { "name": "everything__get-sum", "arguments": {"a": 1, "b": 2} } }' | jq -e \ '.error.message == "tool '\''everything__get-sum'\'' is not available"' ``` The command prints `true`. AISIX rejects the call without sending it upstream. ## Adapt the Setup for Your MCP Server[​](#adapt-the-setup-for-your-mcp-server "Direct link to Adapt the Setup for Your MCP Server") Replace the test server URL with a Streamable HTTP endpoint reachable from the gateway. Keep `type: mcp`, and configure the upstream credential with `auth_type` and its related fields. See [Upstream Authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/upstream-authentication.md). For an open-source AISIX gateway, validate the complete resources file and send `SIGHUP` when the running gateway already has every referenced environment variable. If you add or change an environment variable, recreate the container with the new value. See [Reload a Resources File](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md#reload-a-resources-file). AISIX Cloud stores upstream credentials in the control plane and projects the approved configuration to attached gateways. To let members submit servers without publishing them directly, use [Review and Approve MCP Servers](https://docs.api7.ai/ai-gateway/mcp-gateway/server-review.md). ### Pin the MCP Protocol Revision[​](#pin-the-mcp-protocol-revision "Direct link to Pin the MCP Protocol Revision") The gateway opens the upstream session with the MCP `initialize` handshake, which negotiates the protocol revision with the server. Keep that default for most servers, including servers that implement the stateless `2026-07-28` revision while remaining backward compatible. Set `protocol_version` to `2026-07-28` only for a server that requires that revision, which replaces the handshake with an on-demand `server/discover` call. A pinned server never falls back: if the upstream does not support the pinned revision, the connection fails instead of quietly negotiating an older one. The setting applies to `type: mcp` servers only, since an `openapi` server has no upstream MCP session. For AISIX Cloud, pin the registered server: ``` curl -fsS -X PATCH "$AISIX_CP/mcp_servers/$MCP_SERVER_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"protocol_version":"2026-07-28"}' | jq ``` Send `"protocol_version": null` to remove the pin and return the server to the handshake. For an open-source AISIX gateway, add the field to the server entry and reload: resources.yaml (MCP protocol version) ``` mcp_servers: - name: everything type: mcp url: http://aisix-mcp-everything:3001/mcp auth_type: none protocol_version: "2026-07-28" ``` ## Clean Up[​](#clean-up "Direct link to Clean Up") Keep the server and caller grant if you plan to continue with the other MCP Gateway guides. Otherwise, remove the resources added through your management path. For AISIX Cloud, clear the caller's tool grant and delete the server: ``` curl -fsS -X PATCH \ "$AISIX_CP/environments/$ENV_ID/api_keys/$API_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"mcp_access":{"allow":[]}}' | jq curl -fsS -X DELETE "$AISIX_CP/mcp_servers/$MCP_SERVER_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" ``` For an open-source AISIX gateway, remove the `mcp_access` and `mcp_servers` additions from `resources.yaml`, validate the file, and send `SIGHUP` again. Remove the test server and temporary network: ``` docker rm -f aisix-mcp-everything docker network disconnect aisix-mcp "$AISIX_GATEWAY_CONTAINER" docker network rm aisix-mcp ``` ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now registered an MCP server, granted one tool, and verified both allowed and denied tool calls. Use these guides to extend the setup: * [Connect Cursor to MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/cursor.md): configure Cursor and verify a complete tool call. * [Connect VS Code to MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/vscode.md): configure VS Code and verify a complete tool call. * [Expose a REST API as MCP tools](https://docs.api7.ai/ai-gateway/mcp-gateway/openapi-servers.md): generate tools from an OpenAPI document. * [Configure upstream authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/upstream-authentication.md): use a bearer token, API key, or OAuth client credentials. * [Control tool access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md): grant exact tool names, one server's tools, or every registered tool. * [Apply rate limits and budgets](https://docs.api7.ai/ai-gateway/mcp-gateway/traffic-controls.md): govern MCP tool calls with caller and server limits. * [Configure guardrails](https://docs.api7.ai/ai-gateway/mcp-gateway/guardrails.md): inspect tool arguments and results. --- # Control Tool Access For MCP traffic, the caller API key is the tool-access boundary. A key cannot list or call an MCP tool until access is granted explicitly. Configure tool access when different clients should reach different upstream tools through the same gateway. This guide explains how AISIX names aggregated tools, how to create or update a key with tool access, and how enforcement behaves when a caller lists or calls tools. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * For AISIX Cloud, an environment, a registered MCP server, and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For an open-source AISIX gateway, a caller API key and registered MCP server in `resources.yaml`. [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md) provides a working configuration and the validation and reload workflow. * [cURL](https://curl.se/) and [jq](https://jqlang.org/) for the AISIX Cloud examples. ## How Tool Access Works[​](#how-tool-access-works "Direct link to How Tool Access Works") Tool access is stored on the caller API key in `mcp_access`, an object with an `allow` list and an optional `deny` list. The block is the key's own layer of the tool ACL: in AISIX Cloud it is intersected with the environment and team [MCP access policies](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md), so the key can narrow what those layers allow but never widen it. When the block is omitted the key adds no constraint of its own. That means it takes whatever the policy layers leave — and with no policy configured anywhere, no MCP tool access at all. Access is always granted explicitly. AISIX names each exposed tool in the `__` form. The `` segment is the registered MCP server's `name`; the `` segment is the upstream tool name. For example, `github__create_issue` calls the upstream `create_issue` tool on the registered `github` server. Each entry in `allow` is matched against the prefixed tool name: | Entry | Grants | Example | | ------------- | -------------------------------------- | --------------------------------------------------------------------------------------------- | | Exact name | One specific tool. | `github__create_issue` allows only that tool. | | `__*` | Every tool on one registered server. | `github__*` allows `github__create_issue`, `github__list_repos`, and any other `github` tool. | | `*` | Every tool on every registered server. | `*` allows all current and future tools. | Choose exact names for the narrowest access. Use a per-server wildcard when a caller may use every tool on one server, and use `*` only for a key that should narrow nothing — which is what you send alongside `deny` when the key only means to subtract tools. An empty `allow` list leaves the key no MCP access at all. Entries are single-asterisk globs, so the wildcard can appear outside the trailing per-server form. For example, `*__search` grants a tool named `search` on every registered server. Prefer per-server or exact grants unless you specifically need a cross-server pattern. `deny` uses the same patterns and always wins: a tool matched there is unavailable however the key or any policy layer allows it. Managing many keys? A per-key block is the narrowest control. In AISIX Cloud, use [Manage MCP Access with Policies](https://docs.api7.ai/ai-gateway/mcp-gateway/access-policies.md) to grant access at the environment or team level. Keys that add no block of their own then follow the shared layer, including tools registered later when a matching wildcard covers them. ## Configure Tool Access[​](#configure-tool-access "Direct link to Configure Tool Access") Configure tool access using the management path for your deployment. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Use the Admin API to create a caller API key with tool access or update the tool grant on an existing key. #### Create a Key[​](#create-a-key "Direct link to Create a Key") Export the control-plane URL, admin token, and environment ID: ``` # AISIX_CP includes /api and has no trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create the caller API key that the MCP client uses. This example grants every tool on one server plus one specific tool on another server: ``` curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "mcp-caller", "allowed_models": [], "mcp_access": { "allow": ["github__*", "runbooks__search"] } }' > api_key.json export API_KEY_ID=$(jq -r '.api_key.id' api_key.json) export CALLER_KEY=$(jq -r '.plaintext' api_key.json) ``` ❶ Use an empty model allowlist when the key is only for MCP traffic. ❷ This layer allows every tool on the `github` server plus the single `runbooks__search` tool, and nothing else. In AISIX Cloud the key still only reaches what the environment and team layers also allow. The response returns the plaintext bearer in `plaintext` exactly once. Store it on the client side immediately; subsequent reads return only key metadata. The MCP client sends this value as `Authorization: Bearer ` on gateway requests. #### Update Tool Access[​](#update-tool-access "Direct link to Update Tool Access") Update the caller API key when its tool grant changes. The update is partial: only the fields you send are changed, and `mcp_access` is replaced as a whole block: ``` curl -sS -X PATCH "$AISIX_CP/environments/$ENV_ID/api_keys/$API_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "mcp_access": { "allow": ["github__create_issue"] } }' ``` This update replaces the previous block with one allowing `github__create_issue` only. The key's model access and other settings are untouched because the request does not include them. To revoke all MCP tool access for a key, send `"mcp_access": {"allow": []}`. The key keeps its model access and other settings but can no longer list or call any MCP tool, whatever the policy layers allow. Sending `"mcp_access": null` instead removes the key's layer, which is different: the key then follows the environment and team layers again. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Set `mcp_access` on the caller entry in `resources.yaml`: resources.yaml (MCP access) ``` api_keys: - display_name: mcp-caller key_env: MCP_CALLER_KEY allowed_models: [] mcp_access: allow: - github__* - runbooks__search ``` Set `MCP_CALLER_KEY` in the gateway process environment. This caller can use every tool registered under `github` and only the `search` tool registered under `runbooks`. Other caller settings, including model and A2A access, remain on the same entry. A resources file has no policy layers, so the key's own block is the only one — a caller entry without `mcp_access` therefore has no MCP tool access. To change the grant, edit the complete block, validate `resources.yaml`, and reload the gateway. Set `allow: []` to revoke all MCP tool access while retaining the key's other permissions. See [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md#open-source-aisix-gateway) for the validation and reload commands. ## How Enforcement Works[​](#how-enforcement-works "Direct link to How Enforcement Works") When an MCP client lists tools, AISIX aggregates tools from every enabled server, then filters the list by the caller's effective grant — the intersection of every layer that applies to the key. The client sees only permitted tools. When the client calls a tool, AISIX checks the effective grant again before contacting the upstream server. A tool the caller cannot access is rejected with a neutral MCP error and is never routed upstream. The rejection does not reveal whether the tool or server exists. The same grant applies at a [per-server endpoint](https://docs.api7.ai/ai-gateway/mcp-gateway/overview.md#per-server-endpoints). AISIX evaluates each tool in namespaced `<server>__<tool>` form, then normally presents permitted tools under their original names. For example, a grant for `github__create_issue` presents `create_issue` at `/mcp/github` but grants nothing at another server's endpoint. ## AISIX Cloud Control Plane[​](#aisix-cloud-control-plane "Direct link to AISIX Cloud Control Plane") Tool access can also be configured from the control plane user interface instead of the API calls shown above. The same grant model applies: a caller API key can be scoped to individual tools, to every tool on a server, or to every MCP tool available in the environment. This workflow is available with both AISIX Cloud control-plane deployment options: [On-Premises](https://docs.api7.ai/ai-gateway/on-premises/deployment.md) and [Hybrid Cloud](https://docs.api7.ai/ai-gateway/cloud/overview.md). ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now scoped which MCP tools each caller API key can list and call. Use these guides to add runtime controls or review the shared caller-key settings: * [Rate limits and budgets](https://docs.api7.ai/ai-gateway/mcp-gateway/traffic-controls.md): apply request and concurrency limits, and configure AISIX Cloud budgets for MCP tool calls. * [Guardrails](https://docs.api7.ai/ai-gateway/mcp-gateway/guardrails.md): inspect MCP tool arguments and results. * [Caller API Keys](https://docs.api7.ai/ai-gateway/traffic-controls/caller-api-keys.md): review the shared key settings that also govern model and A2A traffic. --- # Rate Limits and Budgets For MCP traffic, the caller API key remains the traffic-control boundary. MCP tool calls share the key's request and concurrency limits with model traffic, but they do not consume model tokens. In AISIX Cloud, they are also subject to budgets that cover the same caller API key. A caller API key can additionally carry a limit for each MCP server it reaches. An agent looping on one server then cannot exhaust the allowance the same key needs for the others. Configure these limits in `resources.yaml` or through AISIX Cloud, then verify the behavior on the MCP path. Budgets are available only through AISIX Cloud, remain key-wide, and have no per-server setting. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * For AISIX Cloud, an environment, a caller API key, and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For an open-source AISIX gateway, a caller API key and registered MCP servers in `resources.yaml`. [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md) provides a working configuration and the validation and reload workflow. * [cURL](https://curl.se/) for the AISIX Cloud examples. ## Where Controls Apply[​](#where-controls-apply "Direct link to Where Controls Apply") The gateway applies rate-limit and budget checks only to `tools/call` requests. The MCP handshake, including `initialize`, and tool discovery through `tools/list` are not rate-limited. A throttled caller can still connect and list the tools its key allows, but cannot invoke another tool until the window resets. When a rate limit or budget rejects a tool call, AISIX returns before contacting the upstream MCP server and still records the rejected call as a [usage event](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md). ## Applicable Rate Limits[​](#applicable-rate-limits "Direct link to Applicable Rate Limits") A caller API key's `rate_limit` object can define request-rate, token-rate, and concurrency limits. On the MCP path, only request-rate and concurrency limits directly meter tool calls: | Limit | Applies to MCP tool calls | Notes | | -------------------------- | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `rps`, `rpm`, `rph`, `rpd` | Yes | Each `tools/call` counts as one request in the matching window. | | `concurrency` | Yes | Each in-flight tool call holds one concurrency permit until it returns. | | `tpm`, `tpd` | Indirectly | MCP tool calls carry no model tokens, so they do not add to token windows. If the caller key's model traffic already exhausted a token window, the key's tool calls are still rejected with HTTP `429` until the window resets. | Each field is optional. When a field is omitted, AISIX does not enforce that limit. Use request-rate limits or `concurrency` to throttle MCP call volume. A token limit alone does not cap MCP tool calls. Use `mcp_rate_limits` to bound each MCP server separately instead of the caller as a whole. For the full rate-limit field reference and counter-storage options, see [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md). ## Configure Rate Limits[​](#configure-rate-limits "Direct link to Configure Rate Limits") Configure rate limits using the management path for your deployment. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Use the Admin API to configure caller-wide limits and optional limits for individual MCP servers. #### Set a Caller Rate Limit[​](#set-a-caller-rate-limit "Direct link to Set a Caller Rate Limit") Set the control-plane URL, admin token, and environment ID for your AISIX Cloud organization: ``` # AISIX_CP includes /api and has no trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Update the caller API key with a `rate_limit`. The following example keeps the key MCP-only and limits it to one tool call per minute. The `PATCH` request changes only the fields it sends; other key fields keep their current values. Replace `YOUR_API_KEY_ID` with the caller API key ID from its create response. ``` curl -sS -X PATCH "$AISIX_CP/environments/$ENV_ID/api_keys/YOUR_API_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "allowed_models": [], "mcp_access": { "allow": ["github__*"] }, "rate_limit": { "rpm": 1 } }' ``` ❶ Use an empty model allowlist when the key is only for MCP traffic. Preserve existing model access if the same key also needs to call models. ❷ `rpm: 1` limits this caller API key to one request per minute. The updated limit projects to attached gateways automatically. To verify the limit, connect an MCP client with this caller API key and call a permitted tool twice within the same minute. The first `tools/call` succeeds. The second `tools/call` is rejected with HTTP `429` before AISIX contacts the upstream MCP server. The handshake and `tools/list` continue to work while the caller is throttled. #### Limit a Caller per MCP Server[​](#limit-a-caller-per-mcp-server "Direct link to Limit a Caller per MCP Server") The `rate_limit` above is one ceiling over everything the key does. To give each MCP server its own ceiling, add `mcp_rate_limits` to the caller API key. It maps an MCP server name to that key's limits for it: ``` curl -sS -X PATCH "$AISIX_CP/environments/$ENV_ID/api_keys/YOUR_API_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "mcp_rate_limits": { "github": { "rpm": 100, "concurrency": 5 }, "payments": { "rpm": 10 } } }' ``` ❶ Keys the limit by the MCP server's registered name, the same name that prefixes its tools as `github__create_issue`. ❷ Servers can carry different ceilings. A server the map does not name is bounded only by the key's own `rate_limit`. Each named server counts on its own, so the caller exhausting `payments` keeps its full `github` allowance, and another caller's traffic to `payments` is unaffected. Every limit that matches a tool call must pass: a call to `github` counts against both the `github` entry and the key's `rate_limit`. `mcp_rate_limits` accepts the same request-rate and concurrency fields as `rate_limit`: `rps`, `rpm`, `rph`, `rpd`, and `concurrency`. It has no token fields, because MCP tool calls carry no tokens to meter. A `PATCH` replaces the whole map, so send every server you want limited in one request. Send `{}` or `null` to remove all per-server limits. A name that does not match a registered MCP server yet is accepted and takes effect once a server is registered under it. Renaming an MCP server carries its limits with it. Every caller API key that capped the server keeps capping it under the new name. Deleting the server leaves those entries in place: they bind nothing while no server answers to that name, and they apply again if you register one under it. Tool grants behave the same way, so a server recreated under its old name comes back with both its permissions and its caps. Tool grants follow a rename in the same way. Patterns that named the old server are rewritten in each caller key's `mcp_access` block and in the environment- and team-level access policies. Patterns that match a shape rather than one server, such as a bare `*`, are left alone. Deleting a server does not remove its grants. A grant naming a server that does not exist is supported, which lets a key be provisioned before its server is. An entry naming no registered server keeps waiting for one, which is what makes it possible to limit a key before its server exists. The dashboard marks such an entry and offers a **Remove** control for it. To verify, call a permitted tool on the capped server past its limit within the window. Those calls are rejected with HTTP `429`, while tool calls to your other MCP servers keep succeeding. In the dashboard, open **API keys**, create or edit a key, and expand **Per-MCP-server limits** to set the same values per registered server. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Set the general and per-server limits on the caller entry in `resources.yaml`: resources.yaml (MCP rate limits) ``` api_keys: - display_name: mcp-caller key_env: MCP_CALLER_KEY allowed_models: [] mcp_access: { allow: ["github__*", "payments__*"] } rate_limit: rpm: 120 concurrency: 10 mcp_rate_limits: github: rpm: 100 concurrency: 5 payments: rpm: 10 ``` Every tool call must pass the caller's general limit and the matching server limit. In this example, calls to `github` count against both `rate_limit` and `mcp_rate_limits.github`. A server omitted from `mcp_rate_limits` is governed only by the caller's general limit. Edit the complete maps, validate the resources file, and reload the gateway to change these limits. See [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md#open-source-aisix-gateway) for the commands. ## Apply AISIX Cloud Budgets[​](#apply-aisix-cloud-budgets "Direct link to Apply AISIX Cloud Budgets") Budgets are configured in the AISIX Cloud control plane and enforced by the AISIX gateway. When a budget that covers the caller API key is exhausted, the gateway rejects that key's `tools/call` requests with a `budget_exceeded` error before contacting the upstream MCP server. This is the same check the model path runs. MCP tool calls carry no token cost, so they do not add to token-based spend themselves. A budget exhausted by the caller's model traffic still blocks that caller's MCP tool calls when the same budget covers both. See [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md) for budget targets, the rejection response, and caching behavior. For an open-source AISIX gateway, use caller API key rate limits to govern MCP tool-call volume. For AISIX Cloud budget configuration details, see [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md). ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now applied caller API key controls to MCP traffic. Use these guides to observe the result or refine the shared limit and budget settings: * [Observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md): review the usage events and metrics emitted by MCP tool calls. * [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md): review the full rate-limit field reference and counter-storage options. * [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md): configure budget targets, rejection behavior, and cache settings. --- # Upstream Authentication Each registered MCP server can define how AISIX authenticates to the upstream server. AISIX holds any upstream credential gateway-side and presents it when listing tools or forwarding a tool call. The caller API key an MCP client sends to AISIX authenticates the caller to AISIX, and by default no caller header reaches the upstream server. A server that must see one — including the caller's own credential — opts in with [`forward_client_headers`](#forward-caller-headers-to-an-upstream-server). Set the credential with the `auth_type` field and the fields required by that authentication mode. AISIX Cloud and the open-source AISIX gateway support the same modes through different management paths. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * For AISIX Cloud, an environment and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * Complete [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md) for AISIX Cloud or the open-source AISIX gateway. To run the optional end-to-end verification, keep the same shell, gateway, Everything test server, and temporary Docker network running. * [cURL](https://curl.se/) for the AISIX Cloud examples. ## Authentication Modes[​](#authentication-modes "Direct link to Authentication Modes") Choose the mode that matches how the upstream MCP server expects AISIX to authenticate: | `auth_type` | Upstream credential | How AISIX presents it | | ----------- | ------------------------------------ | --------------------------------------------------------------------------------- | | `none` | None | No credential is sent. | | `bearer` | Bearer token in `secret` | `Authorization: Bearer <secret>` | | `api_key` | API key in `secret` | `x-api-key: <secret>` | | `oauth2` | `client_id` + `token_url` + `secret` | AISIX obtains an access token, then sends `Authorization: Bearer <access_token>`. | The `secret` holds the plaintext credential AISIX presents upstream. It is used only gateway-side and is never sent to the calling client. To rotate a credential, update the resource with a new `secret`. Use HTTPS for a credentialed upstream. When a bearer token, API key, or OAuth credential is configured on an `http://` URL, the gateway logs a warning because the credential crosses the network in plaintext. For OAuth, the gateway also warns when `token_url` uses `http://` because the client secret is sent to that endpoint. ## Configure Upstream Authentication[​](#configure-upstream-authentication "Direct link to Configure Upstream Authentication") Configure upstream authentication using the management path for your deployment. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") The AISIX Cloud Admin API creates the server and its upstream credential together. `allowed_environments` lists the environments that receive the server. A server with an empty or missing list is exposed to no environment, so every example below includes the target environment. Export the control-plane connection values: ``` # AISIX_CP includes /api and has no trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` #### No Authentication[​](#no-authentication "Direct link to No Authentication") Use `none` when the upstream MCP server does not require a credential, such as a server reachable only on a trusted internal network. The following example creates a server resource without an upstream credential: ``` curl -sS -X POST "$AISIX_CP/mcp_servers" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "runbooks", "url": "https://runbooks.internal/mcp", "auth_type": "none", "allowed_environments": ["'"$ENV_ID"'"] }' ``` `auth_type` defaults to `none`, so you can also omit it. Leave `secret`, `client_id`, `token_url`, and `scopes` unset for a `none` server. #### Bearer Token[​](#bearer-token "Direct link to Bearer Token") Use `bearer` when the upstream server expects a static token in the `Authorization` header. The following example creates a server resource with the bearer token AISIX should send upstream: ``` curl -sS -X POST "$AISIX_CP/mcp_servers" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "github", "url": "https://mcp.example.com/mcp", "auth_type": "bearer", "secret": "YOUR_UPSTREAM_MCP_TOKEN", "allowed_environments": ["'"$ENV_ID"'"] }' ``` ❶ `bearer` sends `Authorization: Bearer <secret>` on every request to this upstream. ❷ `secret` is required and must be non-empty. #### API Key[​](#api-key "Direct link to API Key") Use `api_key` when the upstream server expects a key in the `x-api-key` header. The following example creates a server resource with the API key AISIX should send upstream: ``` curl -sS -X POST "$AISIX_CP/mcp_servers" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "catalog", "url": "https://catalog.example.com/mcp", "auth_type": "api_key", "secret": "YOUR_UPSTREAM_API_KEY", "allowed_environments": ["'"$ENV_ID"'"] }' ``` ❶ `api_key` sends `x-api-key: <secret>` on every request to this upstream. ❷ `secret` is required and must be non-empty. #### OAuth 2.0 Client Credentials[​](#oauth-20-client-credentials "Direct link to OAuth 2.0 Client Credentials") Use `oauth2` when the upstream server accepts OAuth 2.0 access tokens and you have machine-to-machine client credentials for it. AISIX exchanges the client credentials at the token endpoint, sends the access token upstream, and reuses it until shortly before it expires. If the upstream server rejects a token as unauthorized, AISIX discards the cached token and obtains a fresh one on the next call. The following example creates a server resource with the OAuth client credentials AISIX should use for this upstream: ``` curl -sS -X POST "$AISIX_CP/mcp_servers" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "orders", "url": "https://orders.example.com/mcp", "auth_type": "oauth2", "client_id": "aisix-gateway", "token_url": "https://auth.example.com/oauth/token", "secret": "YOUR_OAUTH_CLIENT_SECRET", "scopes": ["mcp.read", "mcp.write"], "allowed_environments": ["'"$ENV_ID"'"] }' ``` ❶ `oauth2` uses the OAuth 2.0 client credentials grant for this upstream. ❷ `client_id`, `token_url`, and `secret` are required for an `oauth2` server. ❸ `scopes` is optional. AISIX joins the values with spaces into the token request's `scope` parameter. AISIX keeps the access token gateway-side and never returns it to the caller. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Add this authenticated `github` entry to `mcp_servers`; replace it if that name already exists. Keep the other resources unchanged, and supply the secret through environment interpolation: resources.yaml (MCP server authentication) ``` mcp_servers: - name: github type: mcp url: https://mcp.example.com/mcp auth_type: bearer secret: ${GITHUB_MCP_TOKEN} ``` Set `GITHUB_MCP_TOKEN` in the gateway process environment. For `none`, omit `secret`. For `api_key`, the same `secret` field is sent as `x-api-key`. For `oauth2`, add `client_id`, `token_url`, and `secret`, with optional `scopes`. Environment variables belong to the gateway process. A running process cannot receive a variable that you add or change in the host shell. After adding or rotating an environment-supplied credential, validate the resources file with the new value and restart the process. For a container, recreate it with the new environment value. Use a reload only when the resource file changes and the running process already has every referenced variable with the intended value. If validation or reload fails, the gateway never applies the invalid entry; on reload it keeps serving the last valid configuration. See the [CLI Reference](https://docs.api7.ai/ai-gateway/reference/cli.md#validate-a-resources-file) and [Configuration Status](https://docs.api7.ai/ai-gateway/reference/config-status.md). ## Credential Handling[​](#credential-handling "Direct link to Credential Handling") AISIX validates credential fields against `auth_type` when you create or update a server. A server with invalid credential fields is rejected at write time. If a credential later stops working, such as after a secret is rotated or revoked, only that server's tools become unavailable. `tools/list` omits that server's tools, while a direct permitted call returns a generic JSON-RPC internal error. Other registered MCP servers keep working. AISIX logs the detailed failure and does not expose credential details to the calling agent. ## Forward Caller Headers to an Upstream Server[​](#forward-caller-headers-to-an-upstream-server "Direct link to Forward Caller Headers to an Upstream Server") The credential above is the gateway's. An internal MCP server that authorizes on the end user instead needs the caller's own header, and `forward_client_headers` is how it gets one. It is an array of header-name patterns, empty by default, available in both management paths, and it applies to both `type: mcp` and `type: openapi` — so a [REST API exposed as tools](https://docs.api7.ai/ai-gateway/mcp-gateway/openapi-servers.md) receives the headers on every tool call. Through the AISIX Cloud Admin API, set it when you create the server or patch it later. A patch replaces the stored list; send an empty array to clear it: ``` curl -sS -X PATCH "$AISIX_CP/mcp_servers/$MCP_SERVER_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "forward_client_headers": ["authorization", "x-trace-*"] }' ``` In the dashboard, the same setting is the **Forward client headers** box under **Advanced** on the server form, one header name or glob per line. In the resources file an open-source AISIX gateway loads: resources.yaml (forward the caller's credential to an internal server) ``` mcp_servers: - name: runbooks type: mcp url: https://runbooks.internal/mcp auth_type: none forward_client_headers: - authorization - x-trace-* ``` Each entry is an exact header name or a name with a single `*` wildcard, matched case-insensitively. A forwarded header reaches the server whatever AISIX would otherwise have done with it. Naming a credential slot hands the server the caller's credential **in place of** the gateway's, never both. `authorization` is the slot `bearer` and `oauth2` fill; for `api_key` the slot is whatever `api_key_header` names, which is `x-api-key` unless a `type: openapi` server overrides it. Naming that slot on a server that also configures `auth_type` means the caller's value wins for callers who send it. A server that validates the `aud` claim rejects a token minted for the gateway. A [credential slot](https://docs.api7.ai/ai-gateway/models/upstream-request-headers.md#headers-that-must-be-named-exactly) and the trace-context headers `traceparent` and `tracestate` are forwarded only when a pattern names them exactly. The credential slots include the AWS SigV4 headers `x-amz-security-token`, `x-amz-date`, and `x-amz-content-sha256`. A glob such as `*`, `x-*`, or `x-amz-*` never matches any of them. Forwarding a credential or a trace context is an explicit act, not something a broad pattern should sweep up. The linked table is the whole of it on this face. A renamed `api_key_header` is not on it: a `type: openapi` server whose slot is `x-mcp-token` has a name `["x-*"]` matches, so a caller sending that header supplies its own upstream credential in place of the gateway's. The MCP session slots `mcp-session-id`, `mcp-protocol-version`, and `last-event-id` are never forwarded. They name the session the caller holds with AISIX, not the one AISIX opens upstream, and an upstream server refuses a session id it never issued. The rest of what no pattern can reach — `host`, hop-by-hop headers, the `x-aisix-*` namespace, and the headers describing a body AISIX re-serializes — is listed under [Caller Headers AISIX Never Forwards](https://docs.api7.ai/ai-gateway/models/upstream-request-headers.md#caller-headers-aisix-never-forwards). ## Optional: Verify with a Local Test Proxy[​](#optional-verify-with-a-local-test-proxy "Direct link to Optional: Verify with a Local Test Proxy") The configuration above defines the credential AISIX should send, but the Everything server from the setup guide accepts requests without authenticating them. It therefore cannot prove that the gateway presented the expected token. For an end-to-end local check, put a small bearer-authenticated reverse proxy in front of the server and verify both accepted and rejected requests. The proxy is for local testing only: it uses plain HTTP and removes the bearer token before forwarding an accepted request to the Everything server. The helper below contains the local proxy setup. Copy it as-is; the remaining steps configure and test AISIX. Start the local bearer-checking proxy Export a distinct upstream token, then define and start the test proxy on the existing Docker network: ``` export MCP_UPSTREAM_TOKEN="mcp-upstream-test-token" start_mcp_auth_proxy() { docker rm -f aisix-mcp-auth >/dev/null 2>&1 || true docker run -d --name aisix-mcp-auth \ --network aisix-mcp \ -e UPSTREAM_TOKEN="$1" \ caddy:2.11.4-alpine \ sh -c 'caddy run --config /dev/stdin --adapter caddyfile <<EOF :3002 { @authorized header Authorization "Bearer $UPSTREAM_TOKEN" handle @authorized { reverse_proxy aisix-mcp-everything:3001 { header_up -Authorization header_up Host {upstream_hostport} } } handle { respond "upstream authentication failed" 401 } } EOF' for attempt in $(seq 1 30); do docker logs aisix-mcp-auth 2>&1 | grep -q "serving initial configuration" && return sleep 1 done docker logs aisix-mcp-auth >&2 return 1 } start_mcp_auth_proxy "$MCP_UPSTREAM_TOKEN" ``` ### Configure AISIX to Use the Proxy[​](#configure-aisix-to-use-the-proxy "Direct link to Configure AISIX to Use the Proxy") Update the existing `everything` server to use `http://aisix-mcp-auth:3002/mcp`, `auth_type: bearer`, and `MCP_UPSTREAM_TOKEN` as its secret. For AISIX Cloud, update the server created by the setup guide: ``` curl -fsS -X PATCH "$AISIX_CP/mcp_servers/$MCP_SERVER_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @- <<EOF | jq { "url": "http://aisix-mcp-auth:3002/mcp", "auth_type": "bearer", "secret": "${MCP_UPSTREAM_TOKEN}" } EOF ``` For an open-source AISIX gateway, replace the existing `everything` entry and preserve the other resources: resources.yaml (authenticated Everything server) ``` mcp_servers: - name: everything type: mcp url: http://aisix-mcp-auth:3002/mcp auth_type: bearer secret: ${MCP_UPSTREAM_TOKEN} ``` Add `MCP_UPSTREAM_TOKEN` to the gateway container environment. Because the setup guide did not start the container with this variable, validate the complete file with the variable and recreate the container. Keep the mounts, ports, and other environment variables from the quickstart. See [Reload a Resources File](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md#reload-a-resources-file). For the open-source path, recreating the gateway container disconnects it from the temporary MCP network. Connect the replacement container before sending a test request: ``` docker network connect aisix-mcp "$AISIX_GATEWAY_CONTAINER" ``` ### Verify a Valid Credential[​](#verify-a-valid-credential "Direct link to Verify a Valid Credential") After AISIX Cloud projects the update or the open-source gateway restarts, call the permitted tool: ``` MCP_RESPONSE=$(curl -fsS -X POST "$AISIX_PROXY/mcp" \ -H "Authorization: Bearer $AISIX_MCP_KEY" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{ "jsonrpc": "2.0", "id": 10, "method": "tools/call", "params": { "name": "everything__echo", "arguments": {"message": "authenticated through AISIX"} } }') echo "$MCP_RESPONSE" | jq -e \ '.result.content[] | select(.text == "Echo: authenticated through AISIX")' ``` The command prints the matching result. The proxy accepts a request only when it carries the gateway-held token, so this response confirms that AISIX replaced the caller's `Authorization` header with the upstream credential. The proxy removes that credential before forwarding to the Everything server. ### Verify a Rejected Credential[​](#verify-a-rejected-credential "Direct link to Verify a Rejected Credential") Make the proxy expect a different token without changing the credential in AISIX: ``` start_mcp_auth_proxy "deliberately-wrong-token" MCP_FAILURE=$(curl -fsS -X POST "$AISIX_PROXY/mcp" \ -H "Authorization: Bearer $AISIX_MCP_KEY" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{ "jsonrpc": "2.0", "id": 11, "method": "tools/call", "params": { "name": "everything__echo", "arguments": {"message": "this call should fail"} } }') echo "$MCP_FAILURE" | jq -e \ '.error.code == -32603 and .error.message == "upstream MCP server '\''everything'\'' failed to call tool"' if echo "$MCP_FAILURE" | grep -qF "$MCP_UPSTREAM_TOKEN" || \ echo "$MCP_FAILURE" | grep -qF "$AISIX_MCP_KEY"; then echo "credential found in client response" >&2 exit 1 fi ``` The `jq` command prints `true`, and the credential check produces no output. This upstream-authentication failure uses HTTP `200`, so inspect the JSON-RPC `error` object rather than the HTTP status. A failed upstream also contributes no tools to `tools/list`; other reachable servers still contribute theirs. Restore the expected token before continuing: ``` start_mcp_auth_proxy "$MCP_UPSTREAM_TOKEN" ``` Keep `aisix-mcp-auth` running while you use this authenticated local setup. Remove it before running the cleanup commands in the setup guide: ``` docker rm -f aisix-mcp-auth ``` ## Next Steps[​](#next-steps "Direct link to Next Steps") You now understand how AISIX authenticates to upstream MCP servers. Use these guides to control which callers can reach those tools and how their traffic is governed: * [Control tool access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md): scope caller API keys to specific tools, whole servers, or every tool. * [Rate limits and budgets](https://docs.api7.ai/ai-gateway/mcp-gateway/traffic-controls.md): apply request and concurrency limits, and configure AISIX Cloud budgets for MCP tool calls. * [Guardrails](https://docs.api7.ai/ai-gateway/mcp-gateway/guardrails.md): inspect MCP tool arguments and results. * [Upstream request headers](https://docs.api7.ai/ai-gateway/models/upstream-request-headers.md): the same `forward_client_headers` field on the other two faces AISIX proxies through, and the full list of headers no pattern can reach. --- # Connect VS Code to MCP Gateway Visual Studio Code can connect to the AISIX MCP endpoint as a remote Streamable HTTP client. It sends an AISIX caller API key, discovers only the tools that key may use, and invokes those tools through the gateway without receiving upstream server credentials. VS Code remains responsible for selecting a model, deciding when to request a tool, obtaining any required approval, and presenting the result. This connection sends MCP tool traffic through AISIX; it does not route model requests through the gateway. Where VS Code supports a compatible model override, configure that path separately if model traffic should also use AISIX. This guide connects VS Code to the aggregated `/mcp` endpoint. You can continue from [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md), which grants the caller only `everything__echo`, or use an existing AISIX environment and a safe tool the caller may access. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * Either complete [Set Up MCP Gateway](https://docs.api7.ai/ai-gateway/mcp-gateway/setup.md) and retain `AISIX_PROXY` and `AISIX_MCP_KEY`, or obtain an AISIX proxy origin and caller API key from the team that operates the gateway. For an existing environment, export them with those variable names and choose one safe permitted tool for verification. * Install [Visual Studio Code](https://code.visualstudio.com/) with an agent-capable chat provider configured. * Make sure VS Code can reach the AISIX proxy URL. VS Code on the gateway host can use the quickstart address; a remote development environment needs an address it can reach. The example uses workspace configuration and a protected input, so the connection can be reviewed with the project without committing the caller API key. ## Configure the Connection[​](#configure-the-connection "Direct link to Configure the Connection") Create `.vscode/mcp.json` in the workspace. Replace the example URL with `$AISIX_PROXY/mcp`: .vscode/mcp.json ``` { "inputs": [ { "type": "promptString", "id": "aisix-mcp-key", "description": "AISIX caller API key", "password": true } ], "servers": { "aisix": { "type": "http", "url": "https://gateway.example.com/mcp", "headers": { "Authorization": "Bearer ${input:aisix-mcp-key}" } } } } ``` This configuration applies to chat running in the local VS Code extension host. VS Code does not forward servers that require interactive inputs such as `${input:aisix-mcp-key}` to Agent Host sessions. For an Agent Host session, use its portable workspace `.mcp.json` or user-level `~/.copilot/mcp-config.json` path and a supported non-interactive secret source. Open the Command Palette and run **MCP: List Servers**. Select `aisix`, then select **Start Server**. If VS Code asks whether you trust the workspace or server configuration, review the file before approving it. When prompted, enter the caller API key value stored in `AISIX_MCP_KEY`, not the variable name. VS Code stores the protected input separately from the workspace file. Open the chat tool picker with **Configure Tools** and confirm that the selected permitted tool appears under the AISIX server. For the Everything fixture, only `everything__echo` should appear, and the MCP output log should report that one tool was discovered. ## Verify a Tool Call[​](#verify-a-tool-call "Direct link to Verify a Tool Call") The prompt below uses the Everything fixture from the setup guide. For an existing MCP server, substitute a permitted tool name, valid arguments, and an expected result. In VS Code agent chat, ask it to use the exact tool instead of relying on automatic tool selection: ``` Use the MCP tool everything__echo with the message "hello through AISIX". Return the tool result exactly. ``` Review and approve the invocation if VS Code requests confirmation. With the Everything fixture, the result should be: ``` Echo: hello through AISIX ``` Confirm the complete path: * VS Code shows the permitted tool and does not show tools excluded by the caller's effective grant. For the fixture, only `everything__echo` should appear. * [AISIX MCP observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md) records a successful `tools/call` for the expected caller API key and server. * VS Code displays the tool result returned through AISIX. Tool discovery proves that the connection and caller grant work. It does not prove that the model will select a tool or that VS Code's approval policy permits execution, so retain the explicit tool-call test. ## Troubleshoot VS Code[​](#troubleshoot-vs-code "Direct link to Troubleshoot VS Code") | Symptom | Check | | ----------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | VS Code does not start the server | Trust the intended workspace, run **MCP: List Servers**, and start `aisix`. Use **Show Output** to inspect the MCP connection log. | | VS Code returns `401` | Confirm the protected input contains the AISIX caller API key, not an upstream MCP credential. | | The connection succeeds but no tools appear | Follow [Troubleshoot Tool Access](https://docs.api7.ai/ai-gateway/mcp-gateway/overview.md#troubleshoot-tool-access) to check the server and effective grant. After changing the grant, run **MCP: Reset Cached Tools**. | | The tool appears but the agent does not call it | Name the selected tool explicitly, enable it with **Configure Tools**, and review the tool-approval policy. For the fixture, select `everything__echo`. | | Agent Host cannot use the server | Interactive `${input:...}` values are not forwarded to Agent Host. Use its portable MCP configuration and a non-interactive secret source. | ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Client Authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/client-authentication.md): use gateway API keys, OAuth sign-in, or anonymous access for suitable trusted networks. * [Control Tool Access](https://docs.api7.ai/ai-gateway/mcp-gateway/tool-access-control.md): grant exact tools, server patterns, or every registered tool. * [Observability](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md): inspect MCP calls by caller, server, tool, and outcome. --- # Model Aliases Model aliases give callers stable names while AISIX controls how their requests reach upstream models. Each AISIX model resource defines a caller-facing alias in `display_name` and the dispatch path behind it. A direct model maps the alias to one upstream model through one provider key. Routing, semantic, and ensemble models are virtual model aliases that resolve to one or more direct models at request time. Start with direct models for the upstreams AISIX can call. Add a virtual model only when the caller-facing alias needs target selection or response synthesis. ## Choose a Model Shape[​](#choose-a-model-shape "Direct link to Choose a Model Shape") AISIX supports the following dispatch shapes. Select a shape to open its configuration instructions. An embedding model dispatches like a direct model and adds vector metadata; it is not a virtual dispatch shape. | Model shape | How AISIX handles a request | Use when | | ---------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------- | | [Direct](#create-a-direct-model) | Calls one upstream model through one provider key. | The alias has one upstream destination. | | [Embedding](https://docs.api7.ai/ai-gateway/routing/semantic-routing.md#configure-a-semantic-router) | Calls an embedding-capable upstream and records vector metadata for semantic comparisons. | A semantic router needs an embedding model. | | [Routing](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md) | Selects one direct target by failover, round robin, weights, cost, latency, or load. | One stable alias should distribute traffic or survive target failure. | | [Semantic](https://docs.api7.ai/ai-gateway/routing/semantic-routing.md) | Embeds the latest user message and selects one direct target by meaning. | Different topics should reach different models without caller-side routing. | | [Ensemble](https://docs.api7.ai/ai-gateway/routing/ensemble-models.md) | Calls several direct panel models and asks a direct judge model to synthesize one response. | One answer should incorporate multiple model responses. | A model resource contains exactly one dispatch shape: direct upstream fields, a `routing` block, a `semantic` block, or an `ensemble` block. AISIX Cloud references provider keys and other models by resource ID. The open-source AISIX gateway references them by `display_name` in `resources.yaml`. Both management paths reject resources that mix these shapes. A direct model stores its upstream model in `model_name`. It can serve `/v1/embeddings` without embedding metadata. Add an `embedding` block only when the model will support semantic routing; [Semantic Routing](https://docs.api7.ai/ai-gateway/routing/semantic-routing.md) owns that configuration workflow. For supported providers and caller request formats, see [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * A [provider key](https://docs.api7.ai/ai-gateway/models/provider-keys.md) for each upstream credential the models will use. * For AISIX Cloud, access to an environment, an attached gateway, and a write-scoped admin token. The provider key must allow the target environment, and you need its resource ID. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, a gateway that loads a declarative resources file containing the provider key. Models reference that key by its `display_name`. ## Create a Direct Model[​](#create-a-direct-model "Direct link to Create a Direct Model") A direct model maps one caller-facing alias to one upstream model. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Export the AISIX Cloud connection details and provider key ID: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash. # The local On-Premises quickstart uses http://localhost:8080/api. export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" export PROVIDER_KEY_ID="YOUR_PROVIDER_KEY_ID" ``` Create a direct model with the provider key ID you prepared: ``` curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "gpt-4o-prod", "model_name": "gpt-4o", "provider_key_id": "'"$PROVIDER_KEY_ID"'" }' ``` Every successful model creation request returns the created resource in the same response envelope. The following example shows the response for the direct model: ``` { "model": { "id": "677c847f-d92d-4f0e-b445-8b449764f06a", "env_id": "YOUR_ENVIRONMENT_ID", "kind": "direct", "display_name": "gpt-4o-prod", "model_name": "gpt-4o", "provider_key_id": "YOUR_PROVIDER_KEY_ID", "created_at": "2026-06-24T12:18:39Z", "updated_at": "2026-06-24T12:18:39Z" } } ``` Copy the highlighted `id`. Routing, semantic, and ensemble models reference other models by this ID, and you also need it to update, inspect, or delete the model later. The examples for the other model shapes omit this common response. The `display_name` is the name callers send in `model`. The `model_name` is the upstream model ID or deployment name AISIX sends to the provider. These values can be the same, but they do not have to be. The upstream provider comes from the referenced provider key. On the Dashboard, the **Upstream model id** field suggests models published in the catalog for the selected provider key. Open the suggestions with the arrow in the field, or type to filter them. The field still accepts any value. Enter the ID verbatim for a preview model, a private deployment, or another model the catalog does not list. A [Bring Your Own Endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md) provider key has no catalog entry, so the field stays a plain text input for it. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Add the model to the [`models` collection](https://docs.api7.ai/ai-gateway/reference/resources-file.md#models) and reference the provider key by name: resources.yaml (model alias) ``` models: - display_name: gpt-4o-prod provider: openai model_name: gpt-4o provider_key: openai-prod ``` The `display_name` remains the alias callers send in `model`, while `model_name` is the upstream model ID. The model entry declares `provider` explicitly and references the provider key by name. Validate and reload the complete resources file to apply the model. ## Match Model Names with a Wildcard[​](#match-model-names-with-a-wildcard "Direct link to Match Model Names with a Wildcard") A model alias whose `display_name` contains one `*` matches every requested model name that fits the pattern. One alias can therefore front many upstream models without a separate resource for each name. A wildcard alias is a direct model. Set `model_name` to `*` to forward the matched portion upstream, or use a fixed value to send every match to one upstream model. For AISIX Cloud, create the wildcard model with the provider key prepared earlier: ``` curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "openai/*", "provider_key_id": "'"$PROVIDER_KEY_ID"'", "model_name": "*" }' ``` With this alias, a request for `openai/gpt-4o` uses upstream model `gpt-4o`, while a request for `openai/o3-mini` uses `o3-mini`. An exact alias always wins over a wildcard, and the most specific wildcard wins when several match. Add the returned model ID to the caller API key's `allowed_models` list. Wildcard aliases do not appear in `GET /v1/models`, because they are patterns rather than concrete model names. For the open-source gateway, choose a distinct caller key value and make it available to the gateway process: ``` export WILDCARD_CALLER_KEY="YOUR_CALLER_API_KEY" ``` Add `openai/*` to the existing `models` collection and add `wildcard-caller` to `api_keys`. Preserve unrelated entries and collections: resources.yaml (wildcard model access) ``` models: - display_name: openai/* provider: openai provider_key: openai-prod model_name: "*" api_keys: - display_name: wildcard-caller key_env: WILDCARD_CALLER_KEY allowed_models: - openai/* ``` Validate the assembled complete file before loading it. If `WILDCARD_CALLER_KEY` is not already available to the running gateway process, restart or recreate the gateway with the new variable instead of reloading it in place. ## Configure Optional Model Behavior[​](#configure-optional-model-behavior "Direct link to Configure Optional Model Behavior") Most direct models only need a caller-facing alias, upstream model name, and provider key reference. Add optional fields only when the behavior is part of your traffic plan. Common optional fields include: * `timeout`, when a provider request should have a stricter per-request timeout. * `stream_timeout`, when a streaming request should have a separate per-chunk read timeout. * `retries`, when this model should be re-attempted a specific number of times after a retryable upstream failure. See [Retry Budget](#retry-budget). * `allowed_cidrs`, when only callers from specific client IP ranges should use the model alias. * `background_model_check`, when AISIX should probe a direct model outside the request path and mark it unhealthy after failed probes. * `cooldown`, when real request failures should temporarily exclude a direct model from routing. * [`effort_mapping`](https://docs.api7.ai/ai-gateway/models/reasoning-effort-mapping.md), when callers and an upstream model use different reasoning-effort values. * `rate_limit`, when the limit should apply to one model alias. For details, see [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md). Routing models use the selected target's provider settings, timeout, health, and cooldown behavior. Semantic routers use the settings on the selected target and their embedding model, and honor a target's own gates at selection: a route target the caller's IP cannot reach falls through to the default target with the `x-aisix-route` header cleared, a router whose route target and default are both excluded returns the same `403` as a directly addressed model, and a target in cooldown or marked unhealthy by its background check is displaced by an available default. The embedding sub-call resolves its deadline from the router's `embedding_timeout_ms`, then the embedding model's own `timeout`, then the deployment default. A semantic router also carries its own `retries` as the group slot of the retry chain — see [Retry Budget](#retry-budget). Ensemble models use the settings on their panel and judge models. Configure provider settings, health, and cooldown behavior on the referenced direct models, not on the virtual model alias. `allowed_cidrs` applies wherever the model is used, including when it serves as a target of a [routing model](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md). A caller outside the ranges cannot reach the model directly or through the routing model. AISIX resolves the client IP from the immediate peer unless `proxy.real_ip` is configured to trust forwarded headers from your load balancer or ingress. ## Retry Budget[​](#retry-budget "Direct link to Retry Budget") `retries` is the number of extra attempts AISIX makes against a model after a retryable upstream failure, such as a `5xx` response or a transport error. AISIX uses this budget for Chat Completions, Text Completions, Messages, Count Tokens, and Responses. It also applies to Embeddings, Rerank, Audio, Image Generation, Image Editing, video submission and status polling, and the direct model calls made by an ensemble. The setting applies to a model used on its own as well as one used as a routing target. The retry budget does not apply to Realtime, [passthrough routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md), or the Files, Batches, and Fine-tuning APIs. A passthrough route relays each request as a single upstream attempt. It can surface an upstream failure, but the failure does not mark a configured model for cooldown or repeat the request from this model setting. AISIX resolves the budget for each attempt in this order: | Source | Applies When | | ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `retries` on the model | The model sets it. This wins even when the model is a target of a routing model that sets its own `retries`. | | `routing.retries` on the routing model | The target does not set `retries`. Acts as the group-wide default for every target. | | `retries` on the semantic router | The dispatched route target does not set `retries` and the request came through a semantic router. The router's top-level value fills the same group slot as `routing.retries` does on a model group. | | `upstream.retries` in the gateway configuration file | None of the above is set. Defaults to `2`. See the [Startup Configuration Reference](https://docs.api7.ai/ai-gateway/reference/configuration-files.md). | Set `retries: 0` to turn retrying off for a model. A `0` is an explicit setting at every level, never shorthand for leaving the field unset. Two rules apply to the deployment-wide default only. A budget you configure explicitly, at either the model or the routing level, is always applied as written. * **Another target available.** When a routing model has a further target to try and no `retries` is configured, AISIX moves to that target instead of repeating the current one. Repeating a failing target delays the failover without improving the outcome. The last target in the list has nothing to move to, so the default applies there. * **Request timeouts.** A budget you did not configure is not spent on a `timeout`. A `timeout` is a deliberate limit on how long to wait for a model, and repeating the attempt multiplies that wait, usually with the same result. Configure `retries` on the model when a timeout should be retried. Timeouts still trigger failover to another target in both cases. Each retry re-sends the full request body, and AISIX retries in addition to any retry the provider's own edge performs. Account for the combined number of upstream attempts before raising the budget on a high-volume model. On the video submit call, a non-idempotent request is never retried once the provider has already returned its response status. A lost or invalid response body at that point means the operation committed upstream, so replaying it would duplicate the write. Failures during connection or send remain retryable everywhere. ## Cost Metadata[​](#cost-metadata "Direct link to Cost Metadata") Cost metadata affects `least_cost` routing and populates `cost_usd` in usage events for Realtime sessions and completed Batch jobs. It does not affect provider routing or access control. Other inference paths send token counts without a locally calculated cost. In AISIX Cloud, the control plane ignores any gateway-supplied cost and calculates usage and budget totals from its pricing catalog and organization-level overrides. To set a price for a model the catalog does not cover, see [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md). For the open-source AISIX gateway, record pricing metadata with the `cost` field on the model in `resources.yaml`. The field records input and output cost in USD per 1,000 tokens: ``` models: - display_name: gpt-4o-prod provider: openai model_name: gpt-4o provider_key: openai-prod cost: input_per_1k: 0.0025 output_per_1k: 0.01 ``` For running a gateway from a `resources.yaml` file, see the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now configured the caller-facing model alias. Continue with [Caller API Keys](https://docs.api7.ai/ai-gateway/traffic-controls/caller-api-keys.md) to allow applications to use the alias and verify it through the proxy. --- # Provider Key Rotation Provider key rotation replaces the credential AISIX uses to reach an upstream provider. It does not require callers to change their caller API keys or model aliases. Rotate a provider key in place when every dependent model can switch to the new credential together. Create a replacement provider key when you want to migrate models gradually and keep the old credential available for rollback. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * The replacement upstream credential. * An inventory of every model that uses the provider key. A provider key is a shared dependency, so changing it in place affects every dependent model. * For AISIX Cloud, permission to manage provider keys and models. * For the open-source AISIX gateway, access to the gateway process environment and the declarative resources file. * If the replacement also changes the upstream endpoint or protocol, review [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md). ## Choose a Rotation Strategy[​](#choose-a-rotation-strategy "Direct link to Choose a Rotation Strategy") | Strategy | Effect | Use When | | ------------------------ | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ | | In-place rotation | Keeps the provider key reference and switches every dependent model together. | You can validate the replacement credential quickly and accept one coordinated transition. | | Replacement provider key | Keeps both credentials available while models move to a new provider key. | You need gradual rollout, per-model verification, or a straightforward rollback path. | ## Rotate a Provider Key in AISIX Cloud[​](#rotate-a-provider-key-in-aisix-cloud "Direct link to Rotate a Provider Key in AISIX Cloud") AISIX Cloud stores provider credentials as write-only secrets. Read operations never return the stored credential. Updating the secret on an existing provider key preserves its ID, so dependent models do not need to change. Callers do not need new caller API keys, and applications do not need to change model aliases as long as those aliases keep the same names. ### Rotate in Place[​](#rotate-in-place "Direct link to Rotate in Place") In the dashboard: 1. Open **Provider keys** and select **Edit** on the key. 2. Enter the replacement value in **Upstream API key**. The field is blank because the control plane does not return the stored secret; leaving it blank keeps the current value. For a provider whose credential has several fields, supply every required field and select one credential method when alternatives are available. The credential is replaced as a whole, not field by field. 3. Select **Save**. The control plane encrypts the replacement credential and projects it to every environment allowed to use the key. [Resource Projection](https://docs.api7.ai/ai-gateway/cloud/resource-projection.md) explains how saved configuration becomes active gateway configuration. 4. Send requests through each affected model and confirm the upstream accepts the new credential. To automate the same operation, use an [admin token](https://docs.api7.ai/ai-gateway/cloud/admin-tokens.md) with write access. The [AISIX Cloud Admin API Reference](https://docs.api7.ai/ai-gateway/reference/cloud-admin-api/.md) defines the complete request and response schemas. Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash. # The local On-Premises quickstart uses http://localhost:8080/api. export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export PROVIDER_KEY_ID="YOUR_PROVIDER_KEY_ID" ``` Send the replacement plaintext secret in `api_key`: ``` curl -sS -X PATCH "$AISIX_CP/provider_keys/$PROVIDER_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "api_key": "YOUR_NEW_UPSTREAM_API_KEY" }' ``` Omitting `api_key` leaves the stored secret unchanged. An empty value is rejected. For a provider whose credential has several fields, send the complete replacement credential in `config`. This example rotates an Amazon Bedrock credential: ``` curl -sS -X PATCH "$AISIX_CP/provider_keys/$PROVIDER_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "config": { "access_key_id": "YOUR_NEW_ACCESS_KEY_ID", "secret_access_key": "YOUR_NEW_SECRET_ACCESS_KEY", "region": "us-west-2" } }' ``` The whole structured credential is replaced, so include every field the provider requires. For example, Google Vertex AI requires `project`, `region`, and exactly one of `access_token` or `service_account_json`. caution An in-place rotation switches every dependent model to the replacement credential. If the credential is invalid, those models fail until you provide a working value. Use a replacement provider key when you need to keep the old credential available during verification. ### Rotate with a Replacement Provider Key[​](#rotate-with-a-replacement-provider-key "Direct link to Rotate with a Replacement Provider Key") In the dashboard: 1. Create a provider key with the same provider and endpoint settings as the old key, but supply the replacement credential. For a bring-your-own upstream, select the same adapter. 2. Allow the replacement key in every affected environment. 3. Edit each affected model and select the replacement provider key. The dropdown lists only provider keys allowed in that environment. If the replacement key is missing, update its allowed environments first. 4. After the configuration reaches the gateways, send live requests through the migrated models. See [Resource Projection](https://docs.api7.ai/ai-gateway/cloud/resource-projection.md) if a saved change has not reached a gateway. 5. Delete the old provider key only after every dependent model has moved and the replacement credential has been verified. The following API example migrates models in one environment. If the old provider key is shared across environments, allow the replacement key in every affected environment and migrate all dependent models before deleting the old key. Include direct models as well as embedding models used by semantic routers. Export the AISIX Cloud connection details and the affected resource IDs: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash. # The local On-Premises quickstart uses http://localhost:8080/api. export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" export MODEL_IDS="YOUR_MODEL_ID_1 YOUR_MODEL_ID_2" export OLD_PROVIDER_KEY_ID="YOUR_OLD_PROVIDER_KEY_ID" ``` Create the replacement provider key and copy its returned ID: ``` curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @- <<EOF { "provider": "openai", "display_name": "OpenAI replacement", "api_key": "YOUR_NEW_UPSTREAM_API_KEY", "allowed_environments": ["${ENV_ID}"] } EOF ``` For a bring-your-own upstream, include the endpoint and adapter required by that provider key type. Export the replacement ID, then update each affected model: ``` export REPLACEMENT_PROVIDER_KEY_ID="YOUR_REPLACEMENT_PROVIDER_KEY_ID" for MODEL_ID in $MODEL_IDS; do curl -sS -X PATCH "$AISIX_CP/environments/$ENV_ID/models/$MODEL_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @- <<EOF { "provider_key_id": "${REPLACEMENT_PROVIDER_KEY_ID}" } EOF done ``` After every affected model succeeds with the replacement credential, remove the old provider key: ``` curl -sS -X DELETE "$AISIX_CP/provider_keys/$OLD_PROVIDER_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" ``` Deleting a provider key that models still reference prevents those models from dispatching successfully. Rebind every dependent model before deletion. ## Rotate a Provider Key in the Open-Source AISIX Gateway[​](#rotate-a-provider-key-in-the-open-source-aisix-gateway "Direct link to Rotate a Provider Key in the Open-Source AISIX Gateway") The resources file references provider keys by `display_name`. The gateway resolves environment variables when it loads the file, so changing a variable in your shell does not update an already-running gateway process. ### Rotate in Place[​](#rotate-in-place-1 "Direct link to Rotate in Place") Inspect the existing provider key entry and leave it unchanged. The rotation happens by replacing the environment value that the entry references: resources.yaml (provider key) ``` provider_keys: - display_name: openai-prod provider: openai adapter: openai api_key: ${OPENAI_API_KEY} api_base: https://api.openai.com/v1 ``` Replace `OPENAI_API_KEY` in the environment used to start the gateway, then restart or recreate the gateway so the process receives the new value. Because the provider key name stays `openai-prod`, no model references need to change. ### Rotate with a Replacement Provider Key[​](#rotate-with-a-replacement-provider-key-1 "Direct link to Rotate with a Replacement Provider Key") Make both credentials available to the gateway process. Add `openai-replacement` to `provider_keys`, then replace the `gpt-4o-prod` model entry. Keep `openai-prod` and the other resources unchanged: resources.yaml (replacement provider and model) ``` provider_keys: - display_name: openai-replacement provider: openai adapter: openai api_key: ${OPENAI_API_KEY_NEW} api_base: https://api.openai.com/v1 models: - display_name: gpt-4o-prod provider: openai model_name: gpt-4o provider_key: openai-replacement ``` The replacement is a separate provider key. Copy every applicable non-secret field from the old entry, including `provider`, `adapter`, `api_base`, `strip_headers`, `telemetry_tags`, `request`, and `response`. Change only the `display_name` and credential reference unless another setting is intentionally changing. Restart or recreate the gateway if the replacement environment variable was not already available to its process. To migrate models in stages, keep both provider keys in the file, change selected model references to `openai-replacement`, and validate and reload the file after each cohort. Remove `openai-prod` only after no model references it and live requests succeed through every migrated model. ## Verify the Rotation[​](#verify-the-rotation "Direct link to Verify the Rotation") The gateway request used to verify a model is independent of its management path. For each affected model: 1. Send a representative live request through its caller-facing alias with an authorized caller API key. 2. Use the endpoint and request shape that the model normally serves. For example, verify an embedding model through `/v1/embeddings`, not the chat-completions endpoint. 3. Confirm that the upstream accepts the replacement credential and returns the expected response. When using a replacement provider key, confirm that every affected model references it before deleting the old key. In AISIX Cloud, use [Request Logs](https://docs.api7.ai/ai-gateway/cloud/logging-and-auditing.md#request-logs) to inspect the verification requests and [Resource Projection](https://docs.api7.ai/ai-gateway/cloud/resource-projection.md) if a saved change has not reached a gateway. For the open-source gateway, check [Configuration Status](https://docs.api7.ai/ai-gateway/reference/config-status.md) after reloading the resources file. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now rotated a provider credential without changing caller-facing access. Continue with [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) to compare supported request paths and adapter behavior before moving a model to a different upstream. --- # Provider Keys Create a provider key to store the upstream credential and endpoint settings AISIX uses after resolving a model alias. In AISIX Cloud, models reference provider keys by ID. For the open-source AISIX gateway, models in `resources.yaml` reference them by `display_name`. Both paths keep upstream credentials out of application code and allow one credential to be reused across multiple aliases. The creation examples cover both management paths. The field and behavior sections explain differences between the AISIX Cloud Admin API and the declarative resources file where they matter. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * An upstream provider credential. * For AISIX Cloud, access to an environment and an admin token with write scope. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, a gateway that loads a declarative resources file. The [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) provides a working configuration and reload workflow. ## Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Configure the credential using the management path for your deployment. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Create the provider key and save its returned ID for a model configuration. Provider keys are organization-scoped. `allowed_environments` lists the environments that may reference the key when creating models, so include the environment where the models will live. Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash. # The local On-Premises quickstart uses http://localhost:8080/api. export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` The following example creates an OpenAI provider key: ``` # Replace with your value export OPENAI_API_KEY="YOUR_OPENAI_API_KEY" curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "openai-prod", "provider": "openai", "api_key": "'"${OPENAI_API_KEY}"'", "api_base": "https://api.openai.com/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' ``` You should see a response similar to the following: ``` { "provider_key": { "id": "db8613ea-2ecd-40e4-91aa-08197119f766", "org_id": "3f1c2b6a-9d4e-4c1f-8a2b-5e6d7c8f9a0b", "provider": "openai", "display_name": "openai-prod", "allowed_environments": ["9a8b7c6d-5e4f-4a3b-2c1d-0e9f8a7b6c5d"], "strip_headers": null, "telemetry_label": "openai-prod", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" } } ``` The create response does not include `api_base`; fetch the key with `GET $AISIX_CP/provider_keys/{id}` to see the endpoint override. Copy the highlighted `id` and export it. You will use it as `provider_key_id` when creating [models](https://docs.api7.ai/ai-gateway/models/model-aliases.md): ``` export PROVIDER_KEY_ID="YOUR_PROVIDER_KEY_ID" ``` This creates the upstream credential resource only. To send traffic through AISIX, attach the provider key to a model, allow that model on a caller API key, and send a proxy request with the caller API key. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Add the provider key to the [`provider_keys` collection](https://docs.api7.ai/ai-gateway/reference/resources-file.md#provider-keys) in your resources file. Supply the credential through an environment variable instead of writing the plaintext secret into YAML: ``` export OPENAI_API_KEY="YOUR_OPENAI_API_KEY" ``` Add this entry to `provider_keys` in the complete resources file: resources.yaml (provider key) ``` provider_keys: - display_name: openai-prod provider: openai adapter: openai api_key: ${OPENAI_API_KEY} api_base: https://api.openai.com/v1 ``` Models reference this key by its `display_name`, `openai-prod`. Validate the complete resources file before applying it. If `OPENAI_API_KEY` was already available to the running gateway process, follow [Reload a Resources File](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md#reload-a-resources-file). If you introduced the variable now, adapt the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md#start-aisix-ai-gateway) startup command to pass it with `-e`, then recreate the gateway so the process can resolve it. ## Set Provider and Adapter[​](#set-provider-and-adapter "Direct link to Set Provider and Adapter") Provider keys separate the upstream identity from the upstream API format. In AISIX Cloud, a provider identifies the upstream vendor or endpoint. It is not an open string: `provider` must be either an AISIX provider-catalog ID, such as `openai`, `anthropic`, or `deepseek`, or the sentinel value `byo` for a bring-your-own endpoint. The catalog includes native AISIX integrations and community entries sourced from [models.dev](https://models.dev). For the open-source AISIX gateway, `provider` in `resources.yaml` is an open label and `adapter` selects the implemented protocol family. An adapter identifies the upstream API format AISIX should use. It is a closed value because AISIX can only encode implemented protocol families, such as `openai`, `anthropic`, `bedrock`, `vertex`, and `azure-openai`. For an AISIX Cloud catalog provider, the control plane derives the adapter from the catalog entry; sending an `adapter` field returns a 400 error. The OpenAI example above sets only `provider: "openai"`, and the control plane derives the OpenAI adapter. A catalog provider that exposes an OpenAI-compatible API, such as DeepSeek, works the same way — set `provider` and, when needed, `api_base`: ``` { "provider": "deepseek", "api_base": "https://api.deepseek.com" } ``` For a private or OpenAI-compatible endpoint that is not in the AISIX Cloud catalog, set `provider` to `byo`, choose the `adapter` explicitly, and configure `api_base`, which is required for BYO keys: ``` { "provider": "byo", "adapter": "openai", "api_base": "https://api.example.com/v1" } ``` The AISIX Cloud Admin API does not allow changing the adapter of an existing BYO provider key. Create a new provider key with the required adapter and update dependent models to reference it. For the open-source AISIX gateway, change the adapter in the provider key entry in `resources.yaml` and reload the configuration. For adapter selection details, see [Adapter Protocol Families](https://docs.api7.ai/ai-gateway/providers/adapters.md). ## Configure the Base URL[​](#configure-the-base-url "Direct link to Configure the Base URL") `api_base` controls where AISIX sends upstream requests by default. Configure it in the shape expected by the selected adapter. An upstream that serves a second protocol on another path can name that path separately — see [Declare the API Surfaces](#declare-the-api-surfaces). AISIX Cloud can supply a catalog default when one is available, while the open-source gateway derives an endpoint only for the provider-specific cases identified in the setup guides. Common examples are: | Upstream API | Adapter | Base URL | | ---------------------------- | -------------- | --------------------------------------------------------- | | OpenAI | `openai` | `https://api.openai.com/v1` | | DeepSeek | `openai` | `https://api.deepseek.com` | | Gemini OpenAI-compatible API | `openai` | `https://generativelanguage.googleapis.com/v1beta/openai` | | Anthropic | `anthropic` | `https://api.anthropic.com` | | Azure OpenAI | `azure-openai` | `https://<resource>.openai.azure.com` | | AWS Bedrock | `bedrock` | `https://bedrock-runtime.<region>.amazonaws.com` | | Google Vertex AI | `vertex` | `https://<region>-aiplatform.googleapis.com` | For Bedrock, the AISIX Cloud Admin API requires the regional runtime endpoint as `api_base`. The open-source AISIX gateway can omit it for standard AWS and let the AWS SDK derive the endpoint from `region`. AISIX normalizes common copy-paste mistakes, such as trailing slashes and full endpoint paths. It does not guess arbitrary provider URL layouts. For private model services, corporate proxies, or bring-your-own endpoints, configure `api_base` explicitly. ## Declare the API Surfaces[​](#declare-the-api-surfaces "Direct link to Declare the API Surfaces") `api_base` names one endpoint, and the adapter names one protocol. Two common upstreams do not fit that shape: * **One account, two protocols.** Some providers, including DeepSeek, serve an OpenAI-compatible path and an Anthropic-compatible path on the same host with the same credential. Reaching only the path `api_base` names means every request in the other protocol is translated, which drops what that protocol carries and the chat format does not, including Anthropic prompt-cache breakpoints and thinking blocks. * **An OpenAI-compatible endpoint without the Responses API.** Many self-hosted and relay endpoints implement `/v1/chat/completions` and nothing else. The Responses API adds to chat completions rather than renaming it, so the adapter cannot tell you whether the route exists. A Codex request forwarded to a route the endpoint does not implement returns an upstream 404. Use `apis` to state which surfaces the endpoint serves natively and where. Each entry may carry its own `base`, subject to the same shape rules as `api_base`; omit `base` when the surface is served at `api_base`. ``` { "provider": "deepseek", "api_base": "https://api.deepseek.com", "apis": { "responses": {}, "messages": { "base": "https://api.deepseek.com/anthropic" } } } ``` `responses` is listed because this endpoint serves it at `api_base`. Leaving it out would not mean "unchanged" — it would state that the endpoint has no such route, and Responses requests would be translated instead. With that key, a caller on `/v1/chat/completions` reaches `https://api.deepseek.com/chat/completions`, a caller on `/v1/messages` reaches `https://api.deepseek.com/anthropic/v1/messages` untranslated, a caller on `/v1/responses` reaches `https://api.deepseek.com/v1/responses`, and all three address the same model by the same name. Two surfaces can be declared: | Surface | Route | Resolution | | ----------- | ------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `responses` | `/v1/responses` | Declared only. Once `apis` is present, AISIX forwards a Responses request natively only when this surface is listed. Leaving it out is how you tell AISIX the endpoint has no such route, and the request is translated to chat completions instead. | | `messages` | `/v1/messages`, `/v1/messages/count_tokens` | Additive. A key whose adapter is `anthropic` serves this surface whether `apis` lists it or not. Listing it adds both routes to a key whose adapter is something else. AISIX sends the provider credential as `x-api-key` and adds `anthropic-version` on both routes. | Because `responses` is decided by the declaration, an existing OpenAI key that adds `apis` for any reason must list `responses` to keep forwarding Responses requests natively. The AISIX Cloud dashboard fills in the surfaces a key already serves when you switch the section on, so editing from the current state is the default. Declare `messages` only when the upstream accepts the Anthropic authentication headers and implements both routes at the same base. If its native Messages API requires bearer authentication or does not provide `/v1/messages/count_tokens`, use an authenticated [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for that API instead. Some catalog providers ship a declaration of their own, verified against the vendor's live endpoints — DeepSeek is one — and a key on that provider's own base gets it without you configuring anything. The key carries it as a value like any other: the dashboard opens the section already filled in, and the provider key's detail response returns `apis` together with `apis_source: catalog`, so what AISIX declared on your behalf is visible rather than implied. Setting `apis` yourself replaces it and reports `apis_source: operator`, which is what lets AISIX correct a curated entry later without touching a declaration you chose. Clearing the field with `null` returns the key to the curated declaration rather than to no declaration at all. If you repoint `api_base` at your own endpoint the declaration drops away, because it describes the vendor's paths rather than yours. Omit `apis` entirely, on a provider that ships none, to leave every surface resolving from the provider and adapter, which is how every key without the field behaves. Surfaces the field does not cover — embeddings, audio, images, video, files, batches, fine-tuning, and rerank — always use `api_base`. The `bedrock`, `vertex`, and `azure-openai` adapters do not accept `apis`. Those platforms serve their own routes rather than these, so every request reaches them translated. ## Credential Handling[​](#credential-handling "Direct link to Credential Handling") Provider keys store sensitive upstream credentials. Assign ownership for the upstream credential before reusing one provider key across multiple models. In AISIX Cloud, `api_key` is write-only. The plaintext value is encrypted before storage and is never returned by read endpoints. Bedrock and Vertex AI use a structured `config` object because their credentials have multiple fields. When creating these provider keys, set the required `api_key` field to an empty string and supply the credential through `config`. When updating `config`, omit `api_key`; update requests reject an empty `api_key`. For the open-source AISIX gateway, reference credentials in `resources.yaml` through environment variables instead of storing literal secrets in YAML. Structured credentials are serialized as JSON strings and supplied through the provider key's `api_key` field. Provider keys are shared dependencies. Rotating one in place affects every model that references it. Use [Provider Key Rotation](https://docs.api7.ai/ai-gateway/models/provider-key-rotation.md) to choose between an in-place update and a gradual replacement workflow for either management path. ## Configure Provider-Specific Overrides[​](#configure-provider-specific-overrides "Direct link to Configure Provider-Specific Overrides") Provider-key overrides adapt an upstream API that differs slightly from the selected adapter. Every model that references the provider key inherits its overrides, so configure them only when needed. The following AISIX Cloud Admin API example configures request and response compatibility for a custom OpenAI-compatible upstream: ``` export COMPAT_API_KEY="YOUR_UPSTREAM_API_KEY" curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "custom-openai-prod", "provider": "byo", "adapter": "openai", "api_key": "'"${COMPAT_API_KEY}"'", "api_base": "https://api.example.com/v1", "allowed_environments": ["'"${ENV_ID}"'"], "request": { "param_renames": { "max_completion_tokens": "max_tokens" } }, "response": { "reasoning_field": "delta.thinking" } }' ``` ❶ `request.param_renames` renames a top-level parameter when an upstream expects a different name. If a request contains both names, AISIX uses the value from the original caller-facing name. ❷ `response.reasoning_field` maps reasoning from a nonstandard streaming `delta` path to `delta.reasoning_content`. This override applies to the `openai` and `azure-openai` adapters. For the open-source AISIX gateway, add the same `request` and `response` blocks to the provider key entry in the resources file. Support varies by adapter, request path, and management path. The [AISIX Cloud Admin API Reference](https://docs.api7.ai/ai-gateway/reference/cloud-admin-api/.md) defines the overrides accepted by the control plane. The [Resources File Reference](https://docs.api7.ai/ai-gateway/reference/resources-file.md#provider-keys) defines the complete open-source field catalog. The `request` object can also control the headers AISIX sends upstream. Both management paths support `request.default_headers` for generated values such as the calling team and `request.forward_client_headers` for relaying selected caller headers. See [Upstream Request Headers](https://docs.api7.ai/ai-gateway/models/upstream-request-headers.md). ## Verify the Provider Key[​](#verify-the-provider-key "Direct link to Verify the Provider Key") Send a representative request through a model that uses the provider key and confirm that the upstream accepts it. If you configured a custom `reasoning_field`, send a streaming chat-completions request and confirm that reasoning appears in `delta.reasoning_content`. Test provider-key overrides with a non-production alias before reusing the key across models. An incorrect override affects every model that references the provider key. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md) to attach a new provider key to a caller-facing alias. To replace a credential used by existing models, follow [Provider Key Rotation](https://docs.api7.ai/ai-gateway/models/provider-key-rotation.md). --- # Reasoning Effort Mapping Reasoning-capable models do not always use the same effort vocabulary. A client might send `medium`, while a selected upstream model accepts only `high` or `max`. Configure `effort_mapping` on a direct model to rewrite those values at the gateway. The mapping is disabled by default. AISIX changes a request only when it contains an effort string that exactly matches a configured key. ## How Mapping Works[​](#how-mapping-works "Direct link to How Mapping Works") AISIX reads the effort from the normalized endpoint the caller uses: | Endpoint | Request field | | -------------------------------- | ---------------------- | | `POST /v1/chat/completions` | `reasoning_effort` | | `POST /v1/responses` | `reasoning.effort` | | `POST /v1/messages` | `output_config.effort` | | `POST /v1/messages/count_tokens` | `output_config.effort` | It applies one exact, case-sensitive lookup after selecting the final direct model and before serializing the provider request. The same behavior applies to streaming and non-streaming requests and when AISIX translates between OpenAI and Anthropic request formats. For this mapping: ``` { "medium": "high", "high": "max" } ``` * `medium` becomes `high`. AISIX does not look up the resulting `high` again, so it does not become `max`. * `high` becomes `max`. * An unmapped value such as `low` remains `low`. * A request with no effort field remains unchanged; the mapping does not add one. The mapping belongs to the direct target, not the alias the caller addressed. When a routing model or semantic router selects a direct model, that target's mapping applies. Each ensemble panel member uses its own mapping, and the direct judge model uses its own mapping for the synthesis request. Configure no mapping on routing, semantic, ensemble, or embedding model resources; AISIX rejects it on those kinds. Provider and model support still applies after the rewrite. Use only target values that the selected upstream model accepts. A provider adapter can translate, omit, or reject an unsupported value according to that provider's normalized endpoint behavior. [Passthrough routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) do not use model resources and do not apply effort mapping. ## Configure AISIX Cloud[​](#configure-aisix-cloud "Direct link to Configure AISIX Cloud") In the Dashboard, create or edit a direct model, expand **Reasoning effort mapping**, and add each requested value and upstream value pair. Remove every row and save to disable the mapping. You can also set the mapping with the Admin API. The following request creates a direct model that changes `medium` to `high` and `high` to `max`: ``` curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "kind": "direct", "display_name": "reasoning-prod", "model_name": "YOUR_UPSTREAM_MODEL", "provider_key_id": "'"$PROVIDER_KEY_ID"'", "effort_mapping": { "medium": "high", "high": "max" } }' ``` On update, omit `effort_mapping` to leave it unchanged. Send `null` or an empty object to clear it: ``` curl -sS -X PATCH "$AISIX_CP/environments/$ENV_ID/models/$MODEL_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"effort_mapping": null}' ``` ## Configure the Open-Source Gateway[​](#configure-the-open-source-gateway "Direct link to Configure the Open-Source Gateway") Add `effort_mapping` to a direct model in the complete [`resources.yaml`](https://docs.api7.ai/ai-gateway/reference/resources-file.md) snapshot: resources.yaml (direct model) ``` models: - display_name: reasoning-prod provider: openai model_name: YOUR_UPSTREAM_MODEL provider_key: openai-prod effort_mapping: medium: high high: max ``` Omit `effort_mapping`, or set it to an empty object, to disable the rewrite. ## Send a Request[​](#send-a-request "Direct link to Send a Request") Call the model normally. Applications do not need provider-specific effort settings: ``` export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" curl -sS "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "reasoning-prod", "messages": [{"role": "user", "content": "Solve this problem."}], "reasoning_effort": "medium" }' ``` AISIX sends `high` to the selected direct model for this request. --- # Resource Model AISIX gateways use a small set of resources to turn a caller request into an authenticated upstream provider request. The main resources are caller API keys, models, and provider keys. Virtual model shapes and policy resources build on that path when you need target selection, response synthesis, rate limits, guardrails, caching, observability, or AISIX Cloud budget checks. ## Core Traffic Resources[​](#core-traffic-resources "Direct link to Core Traffic Resources") Most AISIX traffic starts with three resources: a caller API key, a model, and a provider key. Together, they decide who can call the gateway, which model name the caller can use, and which upstream provider AISIX calls. For a single-target model, the same relationship applies to AISIX Cloud and the open-source AISIX gateway: The products manage and reference these resources differently: | Product | Management relationship | | ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | AISIX Cloud | Caller API keys and models belong to an environment and reference related resources by ID. Provider keys belong to the organization and must be allowed in the environment. The control plane projects the resulting configuration to attached gateways. | | Open-source AISIX gateway | The resources are declared in `resources.yaml`. Caller API keys reference model `display_name` values, and models reference provider key `display_name` values. | In AISIX Cloud, create or update each resource through the dashboard or Admin API, then use [Resource Projection](https://docs.api7.ai/ai-gateway/cloud/resource-projection.md) to verify that the change reached the gateways. In the open-source gateway, validate the complete resources file and reload it as one configuration snapshot. A failed reload keeps the last valid snapshot active. In the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md), the caller API key is `YOUR_CALLER_API_KEY`, the model alias is `gpt-4o-mini`, the provider key stores the OpenAI credential, and the upstream model is also `gpt-4o-mini`. In production, the alias and upstream model do not need to match. For example, an application can keep sending `prod-chat` while the gateway changes the upstream model, provider key, or routing policy behind that alias. ### Caller API Key[​](#caller-api-key "Direct link to Caller API Key") A caller API key is the authorization identity AISIX uses for an application request. An application can present the key's plaintext value directly or present an OIDC-issued JWT whose external identity is bound to the key. The resolved key controls which model aliases the caller may use and which key-scoped traffic controls apply. For key hashing, rotation, and the model allowlist, see [Caller API Keys](https://docs.api7.ai/ai-gateway/traffic-controls/caller-api-keys.md). For binding an external identity to a key, see [JWT Authentication](https://docs.api7.ai/ai-gateway/traffic-controls/jwt-authentication.md). ### Model[​](#model "Direct link to Model") A model is the gateway-facing model name callers send in the request body. For a direct model, the caller-facing alias can differ from the upstream provider's model ID. The model also points to the provider key AISIX should use for the upstream call. Virtual models resolve the alias through a routing, semantic, or ensemble decision before AISIX reaches a direct model. For each model shape and its configuration, see [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md). ### Provider Key[​](#provider-key "Direct link to Provider Key") A provider key stores the upstream credential and connection settings AISIX uses after it resolves a model. Provider keys keep upstream secrets out of application code and let multiple models reuse the same upstream account, base URL, and adapter family. For credential fields, base URL behavior, provider labels, and adapters, see [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md). For changing a shared upstream credential, see [Provider Key Rotation](https://docs.api7.ai/ai-gateway/models/provider-key-rotation.md). In the open-source gateway, each provider key declares its adapter. AISIX Cloud derives the adapter for catalog providers and accepts an explicit adapter for bring-your-own providers. ## Model Dispatch Shapes[​](#model-dispatch-shapes "Direct link to Model Dispatch Shapes") Every model has a `display_name`, which is the alias used by callers. In AISIX Cloud, references between model resources use resource IDs. For the open-source AISIX gateway, references in `resources.yaml` use `display_name` values. The remaining fields define one dispatch shape: | Shape | Resource relationship | | -------- | ----------------------------------------------------------------------------------------------------------------- | | Direct | Stores an upstream `model_name` and references one provider key. | | Routing | References several direct targets and selects one by its routing strategy. | | Semantic | References an embedding-capable direct model, a default direct model, and direct targets for its semantic routes. | | Ensemble | References direct panel models and a direct judge model. | The four shapes are mutually exclusive. A resource cannot combine direct upstream fields with a `routing`, `semantic`, or `ensemble` block. An embedding model remains a direct model. Its `embedding` block records vector dimensions and normalization behavior so the model can serve embeddings and support semantic routing. Callers can keep one stable alias while operators change the direct upstream, routing targets, semantic routes, or ensemble members behind it. The caller API key must allow the alias named in the request; AISIX selects or invokes referenced models internally. For detailed behavior, see [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md), [Semantic Routing](https://docs.api7.ai/ai-gateway/routing/semantic-routing.md), and [Ensemble Models](https://docs.api7.ai/ai-gateway/routing/ensemble-models.md). ## Policy Resources[​](#policy-resources "Direct link to Policy Resources") Policy resources add gateway behavior around the caller API key, model, and provider key path. Rate limits control request rate and concurrency. Guardrails check request or response content. Caching can reuse chat-completion responses. Observability exporters send gateway request telemetry to an OTLP/HTTP, object storage, Aliyun SLS, or Datadog destination. Budget checks are enforced through AISIX Cloud control-plane policy. The open-source AISIX gateway does not expose a local budget resource. --- # Upstream Request Headers `forward_client_headers` names the inbound headers a caller sends that AISIX relays to an upstream. It is one field with one meaning on all three faces AISIX proxies through, and it is an array of header-name patterns that is empty by default. The reason to set it is an upstream that needs something only the caller can supply: a provider beta flag, an application routing hint, a trace correlation header, or — for an internal service that authorizes on the end user rather than on the gateway — the caller's own credential. Where it lives depends on which face reaches the upstream: | Field | Applies to | | --------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `provider_key.request.forward_client_headers` | The standard protocol endpoints (`/v1/chat/completions`, `/v1/completions`, `/v1/messages`, `/v1/responses`, `/v1/realtime`, embeddings, rerank, audio, images, videos, and the files, batches, and fine-tuning surfaces) that use this provider key. `/v1/realtime` carries this field but not `default_headers`; see [below](#forward-client-headers). | | `passthrough_route.forward_client_headers` | Requests served by that [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). | | `mcp_server.forward_client_headers` | Tool calls to that [MCP server](https://docs.api7.ai/ai-gateway/mcp-gateway/upstream-authentication.md#forward-caller-headers-to-an-upstream-server), for both `type: mcp` and `type: openapi`. | Both management paths configure the field on all three resources. A gateway loading a declarative resources file supports everything on this page, and so do the AISIX Cloud Admin API and dashboard — including an entry that names a credential slot exactly, which is what an internal upstream reading the end user's own credential depends on. The field behaves the same way everywhere — a header a pattern names reaches the upstream — but the default it departs from is not the same on every face, and that changes what setting the field does: * On the standard endpoints and on MCP, AISIX builds the upstream request from scratch and relays no caller header. The list is an **allowlist**: it is the only way a caller header reaches the upstream. * On a passthrough route, AISIX relays the caller's headers by default and strips a small set. The list is an **override of that strip set**: it only matters for the headers the route would otherwise remove, including the credential slot the gateway just consumed to authenticate the caller. A provider key has a second, unrelated header setting: `request.default_headers` adds headers AISIX generates, with values that can reference the current request. Both live on the provider key, so every model that references it inherits them. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * An upstream provider credential and endpoint. The examples configure these settings on a provider key; see [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md) for the rest of that resource's fields. * For AISIX Cloud, access to an environment, an attached gateway, and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, a gateway that loads a declarative resources file. ## Configure Upstream Request Headers[​](#configure-upstream-request-headers "Direct link to Configure Upstream Request Headers") Configure the provider key using the management path for your deployment. The settings have the same runtime behavior in both paths. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Export the AISIX Cloud connection details and upstream credential: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash. # The local On-Premises quickstart uses http://localhost:8080/api. export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" export UPSTREAM_API_KEY="YOUR_UPSTREAM_API_KEY" ``` Create a provider key that injects gateway context and relays approved client headers. You can configure either setting independently; this example shows how they can coexist on one key: ``` curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "internal-vllm", "provider": "byo", "adapter": "openai", "api_key": "'"${UPSTREAM_API_KEY}"'", "api_base": "https://models.internal.example.com/v1", "allowed_environments": ["'"${ENV_ID}"'"], "request": { "default_headers": { "x-tenant-id": "${request.api_key.team_id}", "x-audit-context": "key=${request.api_key.id};model=${model.name}", "x-correlation-id": "${request.id}", "x-upstream-tier": "premium" }, "forward_client_headers": [ "anthropic-beta", "x-routing-hint", "x-trace-*" ] } }' ``` ❶ The team that owns the calling API key. This can identify the tenant for upstream quota or routing decisions. ❷ A value that combines literal text with the calling API key's identifier and the caller-facing model name. ❸ This request's correlation ID, the same value AISIX sends as `x-aisix-request-id` and reports in its own logs. ❹ A literal value, sent unchanged on every request. ❺ An allowlist of inbound header names and patterns. AISIX still removes the [headers it never forwards](#caller-headers-aisix-never-forwards) before sending the request upstream. A credential or trace-context header would have to be [named exactly](#headers-that-must-be-named-exactly); the `x-trace-*` pattern here does not match `traceparent`. The response includes the provider key ID. Save it when you need to change the header settings later: ``` export PK_ID="YOUR_PROVIDER_KEY_ID" ``` #### Change the Configuration Later[​](#change-the-configuration-later "Direct link to Change the Configuration Later") `PATCH /provider_keys/{id}` replaces the complete `request` block, and the change reaches every model that references the key. Read the current block first, then send every field you want to retain with the edited values: ``` curl -sS "$AISIX_CP/provider_keys/$PK_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ | jq '.provider_key.request' curl -sS -X PATCH "$AISIX_CP/provider_keys/$PK_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "request": { "default_headers": { "x-tenant-id": "${request.api_key.team_id}" }, "forward_client_headers": ["anthropic-beta", "x-trace-*"] } }' ``` An empty `"request": {}` clears the stored override. Catalog provider keys return to the provider's bundled defaults; BYO keys have no request override after the clear. Omitting `request` leaves the stored block unchanged. In the dashboard, the same block is under **Advanced wire-shape overrides** on the provider key's edit form and is pre-filled with the stored values. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Add the `request` block to the provider key entry in `resources.yaml`. Keep the entry's existing credential fields and preserve unrelated entries and collections. The example uses `UPSTREAM_API_KEY`, which must be available to the gateway process: resources.yaml (provider key request headers) ``` provider_keys: - display_name: internal-vllm provider: byo adapter: openai api_key: ${UPSTREAM_API_KEY} api_base: https://models.internal.example.com/v1 request: default_headers: x-tenant-id: $${request.api_key.team_id} x-audit-context: key=$${request.api_key.id};model=$${model.name} x-correlation-id: $${request.id} forward_client_headers: - anthropic-beta - x-trace-* ``` In a resources file, every `${NAME}` without an escape is substituted from the environment when the file loads. Escape the dollar sign in request-context references as `$${...}`. The loader converts `$$` to a literal `$`, so `$${request.id}` reaches the gateway runtime as `${request.id}`. An escaped name that is not listed in [Available Variables](#available-variables) does not resolve at runtime, and AISIX drops that header. Validate and reload the complete resources file after adding or changing the block. See [Reload a Resources File](https://docs.api7.ai/ai-gateway/deployment/configuration-propagation.md#reload-a-resources-file) for the complete workflow and the [Resources File Reference](https://docs.api7.ai/ai-gateway/reference/resources-file.md#provider-keys) for the full field catalog. ## Request-Context Header Values[​](#request-context-header-values "Direct link to Request-Context Header Values") Use `request.default_headers` when the upstream needs to know who is calling. An internal model service can apply per-tenant quotas, attribute cost to a team, or join its access log to the AISIX request ID. These headers avoid requiring a separate upstream credential per tenant. A value can be a literal string, a `${...}` reference, or a combination of literal text and multiple references. AISIX resolves each reference after authenticating the caller and selecting the model. ### Available Variables[​](#available-variables "Direct link to Available Variables") | Variable | Value | | ------------------------- | ---------------------------------------------------------------------- | | `request.id` | Correlation id for this request. | | `request.api_key.id` | Identifier of the calling API key. | | `request.api_key.name` | Name of the calling API key. | | `request.api_key.team_id` | Team that owns the calling API key. | | `request.api_key.user_id` | Organization member who owns the calling API key. | | `model.id` | Identifier of the resolved model. | | `model.name` | Caller-facing name of the resolved model, not the upstream model name. | | `provider_key.id` | Identifier of this provider key. | | `provider_key.name` | Name of this provider key. | Only these variables resolve at request time. The AISIX Cloud Admin API rejects a value that references any other name when you save the provider key. This makes a typo fail immediately instead of becoming a header that never arrives. A resources file has a separate load-time interpolation step, so escape request-context references as described under [Open-Source AISIX Gateway](#open-source-aisix-gateway). No variable exposes a secret. Caller API keys, upstream credentials, and signing material are not part of the request context a header value can read. A header whose variables do not all have a value for a request is dropped from that request rather than sent empty. If the calling API key belongs to no team, a request through the example above carries `x-audit-context` and `x-correlation-id`, but no `x-tenant-id`. An empty `x-tenant-id` would tell the upstream that the tenant is the empty string. ### Protected Header Names[​](#protected-header-names "Direct link to Protected Header Names") `request.default_headers` cannot set the names no configuration puts on an upstream request: `host`, the hop-by-hop set, and the gateway's own `x-aisix-*` namespace. Both management paths agree on that boundary — AISIX Cloud rejects these names when you save the provider key, and a gateway reading a resources file drops them at dispatch time. They are the **first** group under [Caller Headers AISIX Never Forwards](#caller-headers-aisix-never-forwards), which binds `default_headers` whatever the header's source. The **second** group there applies only to headers received from the caller, so it places no restriction on `default_headers`. A credential name is not among them. Naming `authorization`, `x-api-key`, or another credential slot in `default_headers` is how an upstream that reads a second, static credential gets one — on a provider whose own credential goes elsewhere. What such an entry cannot do is displace a header AISIX already set. A default header fills a slot the selected provider bridge left empty; it never replaces the upstream credential or `content-type`. ## Forward Client Headers[​](#forward-client-headers "Direct link to Forward Client Headers") List each header as an exact name or a name containing a single `*` wildcard. Matching is case-insensitive, so `X-Trace-*` and `x-trace-*` are the same pattern and both match `x-trace-id`. An empty or absent list — the default — forwards nothing and overrides no stripping. Callers cannot opt themselves in. The list is operator configuration on the upstream-facing resource, and a header a caller sends that no entry names is handled exactly as it would be without the field. On the standard endpoints and MCP, a caller that sends the same header more than once has its first value forwarded and the rest dropped, so the upstream receives one well-formed header rather than a list the gateway never interpreted. A passthrough route relays the caller's headers as they arrived, repeats included. Some requests have no caller behind them at all — a background poll of an asynchronous job, or the embedding lookup a semantic router makes. Those forward nothing regardless of configuration, because there is no inbound request to take a header from. A `/v1/realtime` WebSocket forwards the headers its provider key's list names, like the other standard endpoints. Three things about that face are its own: * It refuses the five handshake slots it owns — `sec-websocket-accept`, `sec-websocket-extensions`, `sec-websocket-key`, `sec-websocket-protocol`, and `sec-websocket-version` — on top of the names [no pattern reaches anywhere](#caller-headers-aisix-never-forwards). They describe the handshake the caller opened to AISIX, not the one AISIX opens upstream. * No pattern reaches those five, including one that names a header in full: this face checks its own refusals before it consults the list. `sec-websocket-protocol` is the one that matters most, because the browser flow carries the caller's own AISIX key as an item in that list. * `request.default_headers` does **not** apply here. One half of the provider key's `request` block takes effect on this face and the other half does not. A forwarded value that is not ASCII is dropped from the handshake rather than failing the session, because this face writes its upstream handshake as text. Every other face forwards the same value byte for byte. ### Configure It on a Route or an MCP Server[​](#configure-it-on-a-route-or-an-mcp-server "Direct link to Configure It on a Route or an MCP Server") On a passthrough route and an MCP server the field is top-level rather than inside a `request` block, and it takes the same patterns. Through the AISIX Cloud Admin API, patch the resource with the complete list you want it to have. `ROUTE_ID` and `SERVER_ID` are the IDs returned when the route and the server were created: ``` curl -sS -X PATCH "$AISIX_CP/environments/$ENV_ID/passthrough_routes/$ROUTE_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"forward_client_headers": ["authorization", "x-trace-*"]}' curl -sS -X PATCH "$AISIX_CP/mcp_servers/$SERVER_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{"forward_client_headers": ["authorization"]}' ``` A patch replaces the stored list rather than adding to it. An empty array clears it on either resource, and a passthrough route additionally accepts `null`, which an MCP server does not. Omitting the field leaves the stored list unchanged on both. In the dashboard, the same setting is a **Forward client headers** box under **Advanced** on the passthrough route and MCP server forms, one header name or glob per line. Configure it in the resources file an open-source AISIX gateway loads: resources.yaml (the same field on the other two faces) ``` passthrough_routes: - name: internal-copilot path_prefix: /internal target_url: https://models.internal.example.com auth_mode: gateway_key # `inject` is the default and would require a provider key to inject. credential_mode: forward_client forward_client_headers: - authorization - x-trace-* mcp_servers: - name: runbooks type: mcp url: https://runbooks.internal/mcp # No gateway credential: the caller's own is what this server reads. auth_type: none forward_client_headers: - authorization ``` ### Headers That Must Be Named Exactly[​](#headers-that-must-be-named-exactly "Direct link to Headers That Must Be Named Exactly") A wildcard does not sweep up a header AISIX itself consumes as a credential, nor a W3C trace-context header. Forwarding one of those is an explicit act, and naming it in full is what forwards it. The headers that need their own entry on every face are: | Headers | Kind | | ------------------------------------------------------------------------------------------ | ------------------------------------- | | `authorization`, `proxy-authorization`, `x-api-key`, `api-key`, `x-goog-api-key`, `cookie` | Credential slots | | `x-amz-security-token`, `x-amz-date`, `x-amz-content-sha256` | Credential slots (AWS SigV4 material) | | `traceparent`, `tracestate` | W3C trace context | The three `x-amz-*` names travel together as one caller identity. `x-amz-security-token` is a live AWS session credential; the other two are the timestamp and body hash the signature covers. No glob reaches any of them, `x-amz-*` included. Relaying a caller's AWS credentials therefore takes all three names written out in full, together with `authorization`. That header carries the signature itself, and needs its own entry on the same terms. A Bedrock provider key applies a different rule to these same names — it [drops them](#caller-headers-aisix-never-forwards) so its own signer owns them. Every other rule on this page treats all three as credential slots, displacement included. A passthrough route adds the two slots it names for itself. AISIX consumes both, and the shared list cannot know the name a given route chose for either: | Headers | Kind | | ------------------------------ | ----------------------------------------------------- | | The route's `auth_header_name` | The gateway credential, under `auth_mode: header_key` | | The route's `identity_header` | The end-user identity the route records and strips | Naming either in full forwards it, and a glob does not reach it. Without that, a route configured with `["x-*"]` would relay the very header AISIX had just consumed to authenticate the caller, or the identity value the route promises to strip. Writing `*` or `x-*` is a statement about your own headers. It is not consent to hand a third-party provider the caller's credential, nor to graft the caller's trace onto that provider's telemetry — both of which a broad glob would otherwise do the moment a caller happened to send the header. Naming the header in full is that consent, and is all that is required: `"forward_client_headers": ["authorization"]` forwards the caller's `Authorization` on every face. This also keeps an existing broad pattern meaning what it meant when it was written. A provider key configured with `["x-*"]` does not begin relaying the caller's `x-api-key` — which on `/v1/*` is the caller's own AISIX gateway key — because the gateway was upgraded. ### Forwarding the Caller's Credential[​](#forwarding-the-callers-credential "Direct link to Forwarding the Caller's Credential") Naming a credential slot is how an internal upstream that already authorizes on the end user keeps doing so with AISIX in front of it. The caller's value **takes** that slot from the credential AISIX would otherwise have injected there: * On a provider key, in place of the key's own credential. * On an MCP server, in place of the credential `auth_type` would fill — `authorization` for `bearer` and `oauth2`, and for `api_key` the header `api_key_header` names (`x-api-key` unless a `type: openapi` server overrides it). * On a passthrough route, in place of the injected provider credential, and instead of the strip that `gateway_key` authentication would otherwise apply to `authorization` and `x-api-key`. An MCP server's `api_key_header` is reachable by a glob unless it happens to be one of the exact-name headers. Its default, `x-api-key`, is a credential slot and so needs its own entry; a `type: openapi` server that renames the slot — to `x-mcp-token`, say — gives it a name that `["x-*"]` matches, and a caller sending that header then supplies its own upstream credential. The slot is single-valued either way: AISIX replaces rather than appends, so the upstream receives one credential and never chooses between two. AISIX still authenticates the caller as usual first — this setting changes only what the upstream sees, never who AISIX believes is calling. The forwarded credential is whatever the caller put in that slot — not necessarily an end-user identity token. On `/v1/*`, on a `gateway_key` passthrough route, and for every MCP client, the caller's `Authorization` is the AISIX caller API key itself, so naming `authorization` sends a live gateway credential upstream. Name the slot only for an upstream you would trust with that value. The upstream must also be one that accepts it: an upstream that validates an audience claim rejects a token minted for the gateway. Do not name a credential slot on a public model provider. Only a credential slot displaces something AISIX already set. Any other header AISIX put on the request, it put there to make the exchange work — a provider's asynchronous-mode flag or API-version selector — so a forwarded header of that name is dropped rather than allowed to break the call. ### W3C Trace Context[​](#w3c-trace-context "Direct link to W3C Trace Context") AISIX reads the caller's `traceparent` and `tracestate` as telemetry input. One valid inbound `traceparent` makes the gateway's HTTP SERVER span a child of the caller's span. A malformed value, or more than one `traceparent`, starts a new local trace instead of failing the request. AISIX keeps `tracestate` only when `traceparent` is valid. See [OTLP Trace Structure](https://docs.api7.ai/ai-gateway/observability/exporters.md#understand-otlp-trace-structure) for the spans AISIX exports. Reading the context is separate from relaying it. By default AISIX does not send the caller's trace headers upstream, on any face. Name `traceparent` or `tracestate` exactly in `forward_client_headers` to relay them as well — appropriate for an internal upstream that reports into the same tracing backend, and not for a third-party provider. ## Caller Headers AISIX Never Forwards[​](#caller-headers-aisix-never-forwards "Direct link to Caller Headers AISIX Never Forwards") Some headers no pattern can reach. These restrictions exist because forwarding the header would break the exchange rather than change who it comes from, so they apply regardless of configuration. They also apply only to caller-supplied values. AISIX still sends its own value for several of these names: it sets `content-type` for the body it builds and stamps `x-aisix-request-id`. The first group applies on every face: | Header group | Headers | Why | | ------------- | ------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Host | `host` | Selects which server the request reaches at all. | | Hop-by-hop | `connection`, `keep-alive`, `te`, `trailer`, `transfer-encoding`, `upgrade`, `proxy-authenticate` | These describe the connection the caller opened to AISIX, not the one AISIX opens to the upstream. | | Gateway-owned | `x-aisix-*` | Assertions AISIX makes about a request it handled. A caller's copy would forge them upstream and lose the assertion itself. A header named in `proxy.request_id.accept_headers` — `x-aisix-request-id` by default — is an exception only in that AISIX [reads the request ID](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md#reuse-your-own-request-id) from it; it then emits that value as its own header, so the upstream receives a single value. | The standard endpoints and MCP rebuild the outbound message, so they exclude a second group as well: | Header group | Headers | Why | | -------------------- | ------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Body and negotiation | `content-type`, `content-length`, `content-encoding`, `accept`, `accept-encoding`, `expect` | They describe a body AISIX re-serializes and a response shape it parses, so the caller's copies would describe the wrong message. | | Response-only | `set-cookie` | A response header with no meaning on a request. | | Provider wire format | `anthropic-version` | Selects the wire format an Anthropic-shaped upstream answers in, which AISIX then decodes. A caller's value breaks the decode. | | Client SDK | `x-stainless-*` | Version headers the caller's SDK sends about itself. Relaying them to a provider that reads the same headers for its own SDK breaks the call. | Three faces exclude a little more of their own: * **MCP** also never forwards `mcp-session-id`, `mcp-protocol-version`, and `last-event-id`. They name the session the caller holds with AISIX, not the one AISIX opens upstream, and an upstream MCP server refuses a session id it never issued. * **Passthrough routes** relay the body verbatim, so the second group above does not apply to them — `content-type` survives. They exclude `content-length` on top of the first group, because the outbound client derives the length from the body it is handed and a relayed value is a request-framing bug. * **`/v1/realtime`** also never forwards `sec-websocket-accept`, `sec-websocket-extensions`, `sec-websocket-key`, `sec-websocket-protocol`, and `sec-websocket-version`. They describe the handshake the caller opened to AISIX, not the one AISIX opens upstream, and no pattern reaches them even when it names one in full. One provider-specific exception: on an AWS Bedrock provider key, the headers AWS SigV4 derives from the request it signs — `authorization`, `x-amz-date`, `x-amz-content-sha256`, `x-amz-security-token`, `x-amz-target`, and `x-amzn-bedrock-accept` — are dropped from both `forward_client_headers` and `default_headers`. A supplied value there would break the signature rather than authenticate anyone. That drop belongs to Bedrock's own signer. It is a separate rule from the [exact-name requirement](#headers-that-must-be-named-exactly). Every other upstream can receive those three `x-amz-*` names, but only when a pattern spells each one out. ## Precedence[​](#precedence "Direct link to Precedence") When the same header name comes from more than one place, AISIX resolves it as follows: 1. `request.default_headers` beats a header relayed by `request.forward_client_headers` — both are operator configuration and the static one is the more specific statement of intent — for every name except a [credential slot](#headers-that-must-be-named-exactly). A forwarded value takes a credential slot from a `default_headers` entry as readily as from the gateway's own. 2. A `default_headers` entry never replaces a header AISIX set itself, including the upstream credential, `content-type`, and `x-aisix-request-id`. 3. A forwarded caller header replaces a header AISIX set itself **only** when the name is a [credential slot](#headers-that-must-be-named-exactly). For every other name, the value AISIX set stands. 4. On a passthrough route, `forward_client_headers` beats the provider key's `strip_headers`. A stripped name is not permanently barred: under `credential_mode: inject`, a route whose list is `["x-*"]` puts a stripped `x-` header back on the upstream request, because a glob is all an ordinary header name needs. The exact-name rule still holds for a credential or trace-context header and for the route's own two slots, and [the headers AISIX never forwards](#caller-headers-aisix-never-forwards) stay out whatever the patterns say. A header is single-valued on the wire throughout: AISIX replaces rather than appends, so an upstream never receives both an operator value and a caller value for the same name. ## Verify[​](#verify "Direct link to Verify") Export a caller API key and model alias that can use a model referencing the provider key. Then send a request with a header the allowlist names: ``` # AISIX_PROXY is the gateway origin; omit a trailing slash and endpoint path. # The local quickstarts use http://127.0.0.1:3000. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export MODEL_ALIAS="YOUR_MODEL_ALIAS" curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -H "x-trace-id: trace-0001" \ -d '{ "model": "'"${MODEL_ALIAS}"'", "messages": [{"role": "user", "content": "ping"}] }' ``` Confirm the expected headers in the upstream service's access log or equivalent request telemetry. If you point `api_base` at a request-inspection service instead, remember that AISIX sends the provider key's credential and the request body to whatever that `api_base` names. Use an endpoint you control, a throwaway credential on the test provider key, and a prompt that carries no real data. If an expected header is missing, check the following: * The header name is listed in `forward_client_headers`, or matches one of its wildcard entries. * The header is not in [Caller Headers AISIX Never Forwards](#caller-headers-aisix-never-forwards). * A credential or trace-context header is [named exactly](#headers-that-must-be-named-exactly), not matched by a wildcard entry. On a passthrough route, the same holds for the route's own `auth_header_name` and `identity_header`. * On `/v1/realtime`, the header is not one of the handshake slots that face refuses, and it is not a `default_headers` entry — that half of the `request` block does not apply there. * For a `default_headers` value with a variable, the calling API key actually has that attribute. A key with no team drops a header that references `request.api_key.team_id`. * In a resources file, each request-context reference escapes the load-time environment interpolation as `$${...}`. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md) for the rest of the provider-key configuration, including compatibility overrides. * [Resources File Reference](https://docs.api7.ai/ai-gateway/reference/resources-file.md#provider-keys) for the declarative field catalog. * [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) and [MCP Upstream Authentication](https://docs.api7.ai/ai-gateway/mcp-gateway/upstream-authentication.md) for the other two faces this field appears on. --- # Observability Exporters Observability exporters send usage events from the gateway to destinations used for tracing, logging, storage, or accounting. This guide configures an OTLP/HTTP exporter using AISIX Cloud or the open-source AISIX gateway, then explains destination choices, content capture, and delivery behavior. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One of these configuration paths: <!-- --> * AISIX Cloud with an environment, an attached gateway, and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * An open-source AISIX gateway that loads a declarative [`resources.yaml`](https://docs.api7.ai/ai-gateway/reference/resources-file.md) file. * A telemetry destination for the exporter you plan to use. * A working model alias and caller API key for delivery verification. Export the gateway origin without a trailing slash or endpoint path, together with the request values used for verification: ``` export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export MODEL_ALIAS="YOUR_MODEL_ALIAS" ``` In the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md), `AISIX_PROXY` is `http://127.0.0.1:3000`. In other deployments, use the address through which your client reaches the gateway. For the AISIX Cloud examples, export the Admin API base URL, admin token, and environment ID: ``` export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_BASE_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` `AISIX_CP` includes `/api` and has no trailing slash. The local On-Premises quickstart uses `http://localhost:8080/api`; use your control plane's reachable Admin API URL in other deployments. ## Choose an Exporter Kind[​](#choose-an-exporter-kind "Direct link to Choose an Exporter Kind") All exporter kinds are available through both configuration paths. AISIX Cloud API requests and resource-file entries select a kind with the `kind` value, and the dashboard shows a label for the same underlying kind: | `kind` Value | Dashboard Label | Use When | | -------------- | ----------------- | ---------------------------------------------------------------------------------------------------------------- | | `otlp_http` | OTLP/HTTP | You already collect traces through an OTLP/HTTP collector or vendor endpoint. | | `object_store` | Object storage | You want batched NDJSON request events in Amazon S3, S3-compatible storage, Google Cloud Storage, or Azure Blob. | | `datadog` | Datadog | You use Datadog Logs HTTP intake. | | `aliyun_sls` | Alibaba Cloud SLS | You use Alibaba Cloud Simple Log Service as the log destination. | The gateway sends telemetry directly to the selected destination. ## Configure an OTLP Exporter[​](#configure-an-otlp-exporter "Direct link to Configure an OTLP Exporter") Set the collector endpoint and authorization header used by the example: ``` # Replace with your values export OTLP_ENDPOINT="https://collector.example.com/v1/traces" export OTLP_AUTH_HEADER="Bearer YOUR_COLLECTOR_TOKEN" ``` The endpoint must be reachable from the gateway process. A receiver running beside a gateway process on the same host can use `http://localhost:4318/v1/traces`; a containerized gateway needs the receiver's container-network or host address instead. Remove the `headers` block from the exporter when the receiver does not require authentication. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Create the exporter and capture its ID: ``` EXPORTER_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/observability_exporters" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "prod-otlp", "kind": "otlp_http", "endpoint": "'"${OTLP_ENDPOINT}"'", "headers": { "Authorization": "'"${OTLP_AUTH_HEADER}"'" }, "sample_rate": 1 }' | jq -r '.observability_exporter.id') ``` ❶ Set static headers only when the OTLP destination requires them. Header values are encrypted at rest and never returned by read operations. ❷ `sample_rate: 1` keeps the delivery check deterministic. It is equivalent to omitting the field, which exports every request trace. After verification, lower the value when you need to reduce span volume. Retrieve the exporter to confirm the stored configuration: ``` curl -sS "$AISIX_CP/environments/$ENV_ID/observability_exporters/$EXPORTER_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" ``` You should see a response similar to the following. Reads expose the configured header names through `header_keys` and confirm stored values through `headers_set`; the header values themselves are never returned. ``` { "observability_exporter": { "id": "b46a9f5d-6a4d-4bb1-ae2c-4ab1b22f5e80", "env_id": "0f2f6a1e-9d33-4a8f-9a6e-2a7b6a9c1d2e", "name": "prod-otlp", "enabled": true, "kind": "otlp_http", "endpoint": "https://collector.example.com/v1/traces", "header_keys": ["Authorization"], "headers_set": true, "sample_rate": 1, "created_at": "2026-07-22T08:30:00Z", "updated_at": "2026-07-22T08:30:00Z" } } ``` Keep `$EXPORTER_ID` for later updates or deletion. Exporters are enabled by default. Set `enabled: false` to save the resource without sending telemetry. ### Open-Source AISIX Gateway[​](#open-source-aisix-gateway "Direct link to Open-Source AISIX Gateway") Add this exporter to `observability_exporters` in the complete resources file. Environment interpolation keeps the collector token out of the file: resources.yaml (OTLP exporter) ``` observability_exporters: - name: prod-otlp kind: otlp_http endpoint: ${OTLP_ENDPOINT} headers: Authorization: ${OTLP_AUTH_HEADER} sample_rate: 1 ``` Validate the assembled complete file: ``` aisix validate --resources resources.yaml ``` Start or restart the gateway with `OTLP_ENDPOINT` and `OTLP_AUTH_HEADER` in its process environment. Reload a running gateway only if those variables are already available to the process. Exporters are enabled by default; add `enabled: false` to keep the entry without sending telemetry. ## Verify OTLP Delivery[​](#verify-otlp-delivery "Direct link to Verify OTLP Delivery") A saved exporter confirms only that AISIX accepted its configuration. Verify the OTLP exporter configured above by capturing the request ID that AISIX returns, finding that ID at the destination, and checking the gateway's delivery signal. If you are checking an existing OTLP exporter with `sample_rate` below `1`, temporarily set the rate to `1` so this request cannot be sampled out. Send one successful request through a working model alias and capture the ID AISIX returns. This remains accurate whether AISIX accepts caller-supplied IDs or generates its own: ``` export EXPORTER_NAME="prod-otlp" if TEST_REQUEST_ID=$( set -o pipefail curl -fsS -D - -o /dev/null \ -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${MODEL_ALIAS}"'", "messages": [{"role": "user", "content": "Reply with: exporter check"}] }' \ | awk 'tolower($1) == "x-aisix-request-id:" { id = $2; sub(/\r$/, "", id) } END { print id }' ) && test -n "$TEST_REQUEST_ID"; then export TEST_REQUEST_ID printf 'request ID: %s\n' "$TEST_REQUEST_ID" else unset TEST_REQUEST_ID printf 'verification request failed or returned no request ID\n' >&2 false fi ``` Wait for the exporter to flush the batch, then search your OTLP backend for a span whose `aisix.request_id` attribute equals `$TEST_REQUEST_ID`. Finding the span confirms end-to-end delivery: the gateway loaded the exporter, authenticated to the destination when required, and the destination accepted the batch. For AISIX Cloud, the exporter row in the environment's Observability view reports whether delivery is healthy and how many batches online gateways have shipped. You can inspect the same per-gateway heartbeat data through the Admin API: ``` curl -sS "$AISIX_CP/environments/$ENV_ID/dp_nodes" \ -H "Authorization: Bearer $AISIX_TOKEN" \ | jq --arg exporter "$EXPORTER_NAME" \ '[.data[] | select(.status != "offline") | .exporter_health[]? | select(.name == $exporter)]' ``` After the next heartbeat, at least one online gateway should report `delivered_batches` greater than zero and `last_error: null`. The counters reset when a gateway restarts or the exporter configuration changes. A non-null `last_error` means delivery is currently failing. The historical fields `failed_batches` and `last_failure_unix` remain populated after a later success clears `last_error`. For an open-source AISIX gateway, confirm success at the destination. If the record does not arrive, check gateway logs for `sink delivery failed; retrying` or `sink delivery dropped after retries`. ## Understand OTLP Trace Structure[​](#understand-otlp-trace-structure "Direct link to Understand OTLP Trace Structure") AISIX exports one trace hierarchy for each request that produces OTLP telemetry, rather than one unrelated span for each usage event. For a model request that reaches an upstream, the hierarchy is: | Span | OTLP Kind | Scope | | ----------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | Inbound request | SERVER | Covers the request from arrival until the response finishes or the caller disconnects. A valid inbound `traceparent` becomes this span's remote parent. | | Logical operation | CLIENT | Covers the gateway's upstream operation across retries and failover attempts. Child of the SERVER span. | | Upstream attempt | CLIENT | Represents one provider dispatch. Each attempt is a child of the logical operation and carries `aisix.attempt_index`. | Retries and failover therefore appear as sibling attempt spans under one logical operation. For each usage event, AISIX puts the complete attributes and captured content on the most specific span present: the attempt, otherwise the logical operation, otherwise the SERVER span. The other spans carry the correlation fields needed to join the hierarchy without repeating the whole event. Not every request has all three levels: * A cache hit, input guardrail block, or other pre-dispatch result has only the SERVER span because AISIX made no upstream call. * MCP, A2A, Realtime, job, and passthrough calls that do not use per-attempt tracking have the SERVER span and one CLIENT span for the upstream operation. Use span kind and parentage rather than the span name alone when counting exported traces. Filter for SERVER spans to count sampled request traces. When `sample_rate` is below `1`, requests that were not selected never appear. Some authentication and malformed-input paths reject requests before a usage event or OTLP span exists, so SERVER spans are not a complete inbound-request count. Use [request metrics](https://docs.api7.ai/ai-gateway/reference/metrics.md#request-metrics) for gateway request volume. Use CLIENT spans carrying `aisix.attempt_index` when analyzing individual model attempts. ### Continue an Inbound W3C Trace[​](#continue-an-inbound-w3c-trace "Direct link to Continue an Inbound W3C Trace") When a request contains exactly one valid `traceparent`, AISIX continues that trace and makes its SERVER span a child of the caller's span. A malformed value or multiple `traceparent` headers are ignored, and AISIX starts a local trace instead of rejecting the request. An accompanying `tracestate` that passes the gateway's character and length checks is recorded on the SERVER span. AISIX does not forward the caller's `traceparent` or `tracestate` to model providers or passthrough targets unless `forward_client_headers` names one of them exactly; a wildcard pattern never matches them. By default the context establishes the caller-to-gateway relationship only. See [Upstream Request Headers](https://docs.api7.ai/ai-gateway/models/upstream-request-headers.md#w3c-trace-context) for the forwarding boundary. ## Configure Other Exporters[​](#configure-other-exporters "Direct link to Configure Other Exporters") Use another exporter kind when telemetry should go to object storage or a logging service instead of an OTLP trace backend. With AISIX Cloud, send one of the following objects as the body of `POST $AISIX_CP/environments/$ENV_ID/observability_exporters`. ### Object Storage[​](#object-storage "Direct link to Object Storage") For object storage, choose the storage provider, bucket, and key prefix. The default authentication mode uses a credential reference resolved by the gateway: ``` { "name": "request-events-s3", "kind": "object_store", "provider": "s3", "bucket": "acme-aisix-events", "prefix": "ai-gateway", "region": "us-east-1", "credential_ref": "acme_s3" } ``` Object storage supports Amazon S3, Google Cloud Storage, Azure Blob, and S3-compatible targets. Cloud identity is supported only for S3 or GCS when the gateway has an attached identity that can write to the bucket. Use `credential_ref` for Azure Blob and for S3-compatible targets that require static credentials. For an S3-compatible target such as MinIO, Cloudflare R2, or Alibaba Cloud OSS, set the target endpoint explicitly. Without an endpoint, the exporter uses the native AWS S3 endpoint. ### Alibaba Cloud SLS[​](#alibaba-cloud-sls "Direct link to Alibaba Cloud SLS") Configure the endpoint host, project, log store, and credential reference: ``` { "name": "request-events-sls", "kind": "aliyun_sls", "endpoint": "ap-southeast-3.log.aliyuncs.com", "project": "acme-observability", "logstore": "ai-gateway", "credential_ref": "acme_sls" } ``` ### Datadog[​](#datadog "Direct link to Datadog") Configure the Datadog site, service name, tags, and credential reference: ``` { "name": "request-events-datadog", "kind": "datadog", "site": "datadoghq.com", "service": "ai-gateway", "tags": ["team:platform", "tier:prod"], "credential_ref": "acme_datadog" } ``` For the resources-file path, add these entries to `observability_exporters`: resources.yaml (request exporters) ``` observability_exporters: - name: request-events-s3 kind: object_store provider: s3 bucket: acme-aisix-events prefix: ai-gateway region: us-east-1 credential_ref: acme_s3 - name: request-events-sls kind: aliyun_sls endpoint: ap-southeast-3.log.aliyuncs.com project: acme-observability logstore: ai-gateway credential_ref: acme_sls - name: request-events-datadog kind: datadog site: datadoghq.com service: ai-gateway tags: ["team:platform", "tier:prod"] credential_ref: acme_datadog ``` Keep only the destinations you use, then validate the resources file and start or reload the gateway as described for the OTLP example. Verify these exporters with the same request-ID approach used for OTLP: search object-storage and SLS records by `request_id`, or Datadog logs by `aisix.request_id`. The [Snowflake guide](https://docs.api7.ai/ai-gateway/observability/load-logs-into-snowflake.md) shows how to inspect object-storage output directly. SLS, Datadog, and object-storage exporters keep destination credentials out of the exporter resource by using credential references or cloud identity. The gateway resolves those credentials locally when it sends telemetry. ## Configure Content Capture[​](#configure-content-capture "Direct link to Configure Content Capture") Exporters include request status, token counts, model and provider identifiers, request IDs, finish reason, and timing by default. Prompt and response bodies remain excluded unless full content capture is enabled on an OTLP/HTTP, SLS, or Datadog exporter. ### Enable Full Content Capture[​](#enable-full-content-capture "Direct link to Enable Full Content Capture") Add `content_mode` and `content_max_bytes` when creating an exporter. In AISIX Cloud, patch an existing exporter with the same fields: ``` curl -sS -X PATCH "$AISIX_CP/environments/$ENV_ID/observability_exporters/$EXPORTER_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "content_mode": "full", "content_max_bytes": 131072 }' ``` In the dashboard, set **Content mode** to **Full**, then adjust **Max content bytes** in the exporter form. For the resources-file path, add the fields to the exporter entry, then validate and reload the file: ``` observability_exporters: - name: prod-otlp kind: otlp_http endpoint: ${OTLP_ENDPOINT} headers: Authorization: ${OTLP_AUTH_HEADER} content_mode: full content_max_bytes: 131072 ``` Use full content capture only when the destination is approved to receive end-user prompt and response text. AISIX applies the configured byte cap independently to captured prompt and response fields. ### Understand Truncated Records[​](#understand-truncated-records "Direct link to Understand Truncated Records") When valid JSON exceeds `content_max_bytes`, AISIX first tries to reduce the value structurally so the exported field remains valid JSON. If the reduced value still cannot fit the configured cap, AISIX falls back to a UTF-8-safe byte cut. The captured prompt is the serialized request body and follows the same behavior. | Content | Truncation Behavior | | -------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | Long string | Keeps a prefix and adds the inline marker `...[aisix: truncated, N bytes total]`. | | Base64 data URI | Replaces the encoded data with a size placeholder. | | Long array | Keeps head and tail samples around an `{"_aisix_truncated": true, "omitted_items": N}` element that accounts for every omitted item. | | Non-JSON content, or JSON that cannot fit after structural reduction | Cuts the value at a UTF-8 character boundary. | When truncation occurs, OTLP records include `aisix.content_truncated: true`. Datadog and SLS records include `content_truncated: true`. ### Capture Passthrough Content[​](#capture-passthrough-content "Direct link to Capture Passthrough Content") For successfully relayed [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) traffic, full content capture records the request body as a string, subject to the exporter's content cap and JSON-aware structural truncation. Buffered responses record extracted text when the response matches a supported extraction shape and otherwise record the body as text. Streamed responses record accumulated extracted text; opaque data payloads are retained as text. ### Capture Failed Requests[​](#capture-failed-requests "Direct link to Capture Failed Requests") Full content capture applies to the supported AI proxy endpoint types summarized below. A2A capture records message content rather than JSON-RPC envelopes. Full content capture does not apply to MCP, Realtime, or job and batch telemetry. Passthrough requests rejected during authentication, authorization, input guardrails, or rate limiting do not include captured content. On `/v1/chat/completions`, `/v1/messages`, and `/v1/responses`, a failure that produces a usage event after request parsing records the request body in the `prompt` field, except for `401` and `403` responses. This includes input guardrail blocks (`422`), upstream failures other than `401` and `403`, and validation failures after parsing, such as an empty messages array. When data masking runs, AISIX captures the post-mask request body. Malformed JSON is rejected before a usage event is created and is not captured. The following boundaries also apply: * Caller authentication failures are rejected before a usage event is created. * When every routing target fails, the last attempt's record carries the prompt. * A response-side guardrail block never captures the blocked output. These boundaries keep rejected credentials and blocked output out of exported content while retaining the request context needed to investigate other failures. ### Review Captured Content by Endpoint[​](#review-captured-content-by-endpoint "Direct link to Review Captured Content by Endpoint") | Endpoint Type | Captured Content | | ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | Text generation | Response text. | | A2A | Text from request and reply message parts on `message/send` and `message/stream`. File parts, data parts, and JSON-RPC envelopes are not captured. | | Embeddings, rerank, and image generation | Full response JSON. | | Audio transcription | Returned transcript. Uploaded audio is not captured; its SHA-256 checksum is recorded alongside the request's text fields. | | Text to speech | Binary speech response is not captured. | ## Model Fields in Exported Telemetry[​](#model-fields-in-exported-telemetry "Direct link to Model Fields in Exported Telemetry") Usage telemetry records both the model alias requested by the caller and the model that served an attempt when those values differ. This distinction matters for routed and ensemble traffic, where one caller-facing alias can resolve to several target calls. Destinations render these values in their own telemetry format. OTLP traces use the following fields: * `gen_ai.request.model` contains the caller-requested alias. * `gen_ai.response.model` contains the concrete response model version when the provider reports one. * `aisix.model_id` identifies the resolved model resource. ### Gateway-Protocol Spans[​](#gateway-protocol-spans "Direct link to Gateway-Protocol Spans") An A2A or MCP call is not a model inference, so its span is not encoded as one. A2A exports as `invoke_agent <agent>` and MCP as `execute_tool <tool>`, following the OpenTelemetry generative-AI semantic conventions, and carries the protocol's own detail: * A2A: `gen_ai.agent.name`, `gen_ai.conversation.id` (the A2A context), and `aisix.a2a.operation`, `aisix.a2a.method`, `aisix.a2a.protocol_version`, `aisix.a2a.task_id`, `aisix.a2a.task_state`, `aisix.a2a.stream_event_count`. * MCP: `gen_ai.tool.name` and `aisix.mcp.server_name`. * Passthrough routes: the span exports as `passthrough <route>` with `aisix.passthrough.route_name`, and `aisix.client_identity` carries the end-user identity when the route's `identity_header` extracted one. Model traffic keeps the `chat.completions` span name and `gen_ai.operation.name: chat`. A trace query written against those values used to match agent and tool calls as well. It now matches model traffic only, which is what it was meant to describe. Datadog logs use `aisix.requested_model` for the caller-requested alias, `gen_ai.response.model` for the concrete response model version, and `aisix.model_id` for the resolved model resource. Object storage and Alibaba Cloud SLS retain usage-event field names such as `requested_model`, `model_id`, and `provider_model_version`. ## Request Kind in Exported Telemetry[​](#request-kind-in-exported-telemetry "Direct link to Request Kind in Exported Telemetry") Every record also carries the kind of work the request asked for — a chat completion, an image generation, a video submission, a tool call. See [Tell Request Kinds Apart](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md#tell-request-kinds-apart) for the values and what they mean. Destinations name it differently: * Object storage and Alibaba Cloud SLS keep the usage-event name, `operation`. * Datadog maps it to `aisix.operation`. * OTLP carries it as the `aisix.operation` span attribute, on every span of the request's hierarchy so a trace can be filtered by kind at its root. OTLP spans also carry the OpenTelemetry `gen_ai.operation.name` attribute, and the two answer different questions. That one uses OpenTelemetry's own vocabulary, which distinguishes model inference from agent and tool calls. It has a single value for every OpenAI-compatible endpoint, so it cannot separate a conversation from an image or a video. Filter on `aisix.operation` when the endpoint matters. ## AISIX Cloud Control Plane[​](#aisix-cloud-control-plane "Direct link to AISIX Cloud Control Plane") The dashboard manages the same exporter resources from the target environment's Observability view. The exporter form collects the destination fields, content mode, and any kind-specific options. Whether an exporter is saved through the API or the dashboard, the control plane projects the configuration to the AISIX gateways attached to that environment. The following behavior applies to exporters projected to AISIX gateways: * The AISIX gateway sends request telemetry directly to your destination. The control plane does not proxy exported telemetry. * Prompt and response content stays on the gateway unless you enable full content capture on an exporter. * Credential references are resolved by the AISIX gateway. When a destination needs runtime credentials, the dashboard shows the environment variables to configure on the gateway. * Delivery health comes from gateway heartbeat data and shows whether batches are being shipped or whether the gateway is reporting a delivery error. For OTLP/HTTP exporters, the dashboard provides presets for Langfuse, Honeycomb, and Grafana Cloud Tempo, and accepts custom OTLP endpoints. The trace UI URL template is optional. Use it when the **Request Logs** view should link a request record to an external trace UI. The template must include `{request_id}` so the control plane can replace it with the request ID from the log record. ## Next Steps[​](#next-steps "Direct link to Next Steps") Continue with [Load Request Telemetry into Snowflake](https://docs.api7.ai/ai-gateway/observability/load-logs-into-snowflake.md) to query object-storage telemetry in Snowflake. Use [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md) to correlate exported records with gateway metrics, access logs, and response headers. --- # Load Request Telemetry into Snowflake AISIX can write request telemetry to object storage through an observability exporter. Use this guide after [configuring an exporter](https://docs.api7.ai/ai-gateway/observability/exporters.md) when finance, business intelligence, or audit teams need to query gateway usage, cost, latency, cache outcomes, and guardrail outcomes in Snowflake. The gateway writes batched NDJSON objects to a bucket you own. Snowpipe ingests those objects into a Snowflake table, and a SQL view turns each request-attempt record into columns that can be queried. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One of these configuration paths: <!-- --> * AISIX Cloud with an environment, an attached gateway, and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * An open-source AISIX gateway that loads a declarative [`resources.yaml`](https://docs.api7.ai/ai-gateway/reference/resources-file.md) file. * A working model alias and caller API key that can send chat-completions requests. * One Amazon S3 bucket or Azure Blob container for staged request telemetry. * A Snowflake account and a role that can create storage integrations, notification integrations, stages, pipes, tables, and views. Export the gateway URL and request values used to generate test telemetry: ``` export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" export AISIX_API_KEY="YOUR_CALLER_API_KEY" export MODEL_ALIAS="YOUR_MODEL_ALIAS" ``` `AISIX_PROXY` has no trailing slash or endpoint path. The [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) uses `http://127.0.0.1:3000`; in other deployments, use the address through which your client reaches the gateway. For the AISIX Cloud examples, also export the Admin API base URL, admin token, and environment ID: ``` export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_BASE_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` `AISIX_CP` includes `/api` and has no trailing slash. The local On-Premises quickstart uses `http://localhost:8080/api`; use your control plane's reachable Admin API URL in other deployments. This guide uses `acme-aisix-events` as the S3 bucket name and `ai-gateway` as the object prefix. Replace them with values for your environment. ## Telemetry Ingestion Flow[​](#telemetry-ingestion-flow "Direct link to Telemetry Ingestion Flow") The `object_store` exporter is an object-storage staging path. AISIX does not push request telemetry directly to Snowflake from the request path. AISIX emits usage telemetry for each request attempt. A request that retries or fails over produces multiple records with the same request ID. The exporter writes gzipped NDJSON objects to Amazon S3 or Azure Blob Storage. Snowpipe loads each record into a Snowflake landing table, and a SQL view exposes request, attempt, cost, latency, cache, and guardrail columns for analysis. Exporter delivery is metadata-oriented by default. It includes fields such as request ID, requested model alias, resolved model ID, status, token counts, cost, latency, cache status, and guardrail outcome. It does not include prompt or response text unless a supported exporter is explicitly configured for content capture. Give each gateway deployment its own bucket prefix. AISIX partitions objects by date and hour under the configured prefix, but the object path does not add a deployment segment for you. ## Configure Object Storage Export[​](#configure-object-storage-export "Direct link to Configure Object Storage Export") Create the exporter first and confirm files land in object storage before wiring Snowflake. This isolates the gateway side from the warehouse ingestion side. ### Amazon S3[​](#amazon-s3 "Direct link to Amazon S3") With AISIX Cloud, create an exporter through the Admin API: ``` EXPORTER_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/observability_exporters" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "snowflake-staging", "kind": "object_store", "provider": "s3", "bucket": "acme-aisix-events", "prefix": "ai-gateway", "region": "us-east-1", "credential_ref": "acme_s3_prod" }' | jq -r '.observability_exporter.id') ``` ❶ Use `s3` for Amazon S3 and S3-compatible storage targets. ❷ `bucket` and `prefix` define where AISIX writes Snowflake staging files. Use a distinct prefix for each gateway deployment. ❸ `credential_ref` maps this exporter to gateway environment variables. The suffix comes from the reference value, converted to uppercase. For the open-source AISIX gateway, add this entry to `observability_exporters` in the complete resources file: resources.yaml (S3 exporter) ``` observability_exporters: - name: snowflake-staging kind: object_store provider: s3 bucket: acme-aisix-events prefix: ai-gateway region: us-east-1 credential_ref: acme_s3_prod ``` Configure the matching credentials in the environment of every gateway instance that sends telemetry. For a gateway launched from a shell, export these values before starting it: ``` # Replace with your values export OBJSTORE_CRED_ACME_S3_PROD_AWS_ACCESS_KEY_ID="YOUR_AWS_ACCESS_KEY_ID" export OBJSTORE_CRED_ACME_S3_PROD_AWS_SECRET_ACCESS_KEY="YOUR_AWS_SECRET_ACCESS_KEY" ``` For the resources-file path, validate the assembled complete file: ``` aisix validate --resources resources.yaml ``` After validation, start or restart the gateway with the credential variables in its process environment. Reload a running gateway only if those variables are already available to the process. Generate a few requests so the exporter has telemetry to flush: ``` for i in 1 2 3; do curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${MODEL_ALIAS}"'", "messages": [{"role": "user", "content": "ping"}] }' > /dev/null done sleep 8 ``` List the staged objects: ``` aws s3 ls "s3://acme-aisix-events/ai-gateway/" --recursive ``` The bucket should contain one or more `.ndjson.gz` objects under the configured prefix, date partition, and hour partition: ``` 2026-06-09 14:03:11 742 ai-gateway/dt=2026-06-09/hh=14/9f86d081884c7d659a2feaa0c55ad015.ndjson.gz ``` Inspect one object before configuring Snowpipe: ``` aws s3 cp "s3://acme-aisix-events/ai-gateway/dt=2026-06-09/hh=14/OBJECT.ndjson.gz" - | gunzip ``` The output contains one request-attempt record per line. This example shows one formatted record: ``` { "schema_version": "1.0", "request_id": "742c6f5e-7b97-4bb1-9f5f-8cb42b4c93e1", "occurred_at": "2026-06-09T14:03:10Z", "requested_model": "gpt-4o-prod", "model_id": "b7c8e4f2-2e4d-4776-a8d7-09b4eb0cb2b1", "attempt_index": 0, "attempt_kind": "initial", "prompt_tokens": 8, "completion_tokens": 12, "upstream_latency_ms": 548, "downstream_latency_ms": 612, "status_code": 200, "cost_usd": 0.00021, "cache_status": "miss", "guardrail_blocked": false } ``` ### Azure Blob[​](#azure-blob "Direct link to Azure Blob") With AISIX Cloud, create an exporter through the Admin API: ``` EXPORTER_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/observability_exporters" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "snowflake-staging", "kind": "object_store", "provider": "azure_blob", "bucket": "ai-gateway", "prefix": "ai-gateway", "credential_ref": "acme_az_prod" }' | jq -r '.observability_exporter.id') ``` ❶ Use `azure_blob` when the staging destination is an Azure Blob container. ❷ For Azure Blob, `bucket` is the container name. The prefix controls the folder-like path under that container. ❸ `credential_ref` maps this exporter to gateway environment variables. The suffix comes from the reference value, converted to uppercase. For the open-source AISIX gateway, add this entry to `observability_exporters` in the complete resources file: resources.yaml (Azure Blob exporter) ``` observability_exporters: - name: snowflake-staging kind: object_store provider: azure_blob bucket: ai-gateway prefix: ai-gateway credential_ref: acme_az_prod ``` Configure the matching credentials in the environment of every gateway instance that sends telemetry. For a gateway launched from a shell, export these values before starting it: ``` # Replace with your values export OBJSTORE_CRED_ACME_AZ_PROD_AZURE_ACCOUNT="YOUR_STORAGE_ACCOUNT" export OBJSTORE_CRED_ACME_AZ_PROD_AZURE_ACCESS_KEY="YOUR_STORAGE_ACCESS_KEY" ``` For the resources-file path, validate the assembled complete file: ``` aisix validate --resources resources.yaml ``` After validation, start or restart the gateway with the credential variables in its process environment. Reload a running gateway only if those variables are already available to the process. Generate traffic and list the staged blobs: ``` for i in 1 2 3; do curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "'"${MODEL_ALIAS}"'", "messages": [{"role": "user", "content": "ping"}] }' > /dev/null done sleep 8 az storage blob list \ --account-name "YOUR_STORAGE_ACCOUNT" \ --account-key "YOUR_STORAGE_ACCESS_KEY" \ --container-name "ai-gateway" \ --prefix "ai-gateway/" \ --query "[].name" \ -o tsv ``` The container should contain gzipped NDJSON objects under the same prefix, date partition, and hour partition layout. ## Create a Snowflake Landing Table[​](#create-a-snowflake-landing-table "Direct link to Create a Snowflake Landing Table") Create a landing table with a `VARIANT` column so Snowflake can ingest each NDJSON record without losing fields: ``` CREATE DATABASE IF NOT EXISTS aisix; CREATE SCHEMA IF NOT EXISTS aisix.gateway; USE SCHEMA aisix.gateway; CREATE TABLE IF NOT EXISTS gateway_events ( record VARIANT, source_file STRING, loaded_at TIMESTAMP_LTZ DEFAULT CURRENT_TIMESTAMP() ); ``` Snowflake detects gzip when `COMPRESSION = AUTO` is used in the pipe file format. ## Connect Snowflake to S3[​](#connect-snowflake-to-s3 "Direct link to Connect Snowflake to S3") For S3, Snowflake reads the bucket through a storage integration and receives object-created events through the SQS queue associated with the pipe. Create the storage integration: ``` CREATE STORAGE INTEGRATION aisix_s3_int TYPE = EXTERNAL_STAGE STORAGE_PROVIDER = 'S3' ENABLED = TRUE STORAGE_AWS_ROLE_ARN = 'arn:aws:iam::123456789012:role/aisix-snowflake-read' STORAGE_ALLOWED_LOCATIONS = ('s3://acme-aisix-events/ai-gateway/'); DESC INTEGRATION aisix_s3_int; ``` Use the `STORAGE_AWS_IAM_USER_ARN` and `STORAGE_AWS_EXTERNAL_ID` values returned by `DESC INTEGRATION` to configure the IAM role trust policy. Grant the role permission to list the bucket and read objects under the configured prefix. Create the stage and pipe: ``` CREATE STAGE aisix_stage URL = 's3://acme-aisix-events/ai-gateway/' STORAGE_INTEGRATION = aisix_s3_int; CREATE PIPE aisix_events_pipe AUTO_INGEST = TRUE AS COPY INTO gateway_events (record, source_file) FROM (SELECT $1, METADATA$FILENAME FROM @aisix_stage) FILE_FORMAT = (TYPE = JSON COMPRESSION = AUTO); SHOW PIPES; ``` Use the `notification_channel` value from `SHOW PIPES` to add an S3 event notification for object-created events: ``` aws s3api put-bucket-notification-configuration \ --bucket "acme-aisix-events" \ --notification-configuration '{ "QueueConfigurations": [{ "QueueArn": "arn:aws:sqs:us-east-1:NNNN:sf-snowpipe-example", "Events": ["s3:ObjectCreated:*"], "Filter": {"Key": {"FilterRules": [{"Name": "prefix", "Value": "ai-gateway/"}]}} }] }' ``` ## Connect Snowflake to Azure Blob[​](#connect-snowflake-to-azure-blob "Direct link to Connect Snowflake to Azure Blob") For Azure Blob, Snowflake reads the container through a storage integration and receives object-created events through a storage queue. Create a storage queue and Event Grid subscription: ``` az storage queue create \ --name "aisix-snowpipe" \ --account-name "YOUR_STORAGE_ACCOUNT" az eventgrid event-subscription create \ --source-resource-id "/subscriptions/YOUR_SUBSCRIPTION_ID/resourceGroups/YOUR_RESOURCE_GROUP/providers/Microsoft.Storage/storageAccounts/YOUR_STORAGE_ACCOUNT" \ --name "aisix-snowpipe-sub" \ --endpoint-type storagequeue \ --endpoint "/subscriptions/YOUR_SUBSCRIPTION_ID/resourceGroups/YOUR_RESOURCE_GROUP/providers/Microsoft.Storage/storageAccounts/YOUR_STORAGE_ACCOUNT/queueServices/default/queues/aisix-snowpipe" \ --advanced-filter data.api stringin CopyBlob PutBlob PutBlockList FlushWithClose ``` Create the notification integration: ``` CREATE NOTIFICATION INTEGRATION aisix_az_notif ENABLED = TRUE TYPE = QUEUE NOTIFICATION_PROVIDER = AZURE_STORAGE_QUEUE AZURE_STORAGE_QUEUE_PRIMARY_URI = 'https://YOUR_STORAGE_ACCOUNT.queue.core.windows.net/aisix-snowpipe' AZURE_TENANT_ID = 'YOUR_TENANT_ID'; DESC NOTIFICATION INTEGRATION aisix_az_notif; ``` Grant the Snowflake service principal returned by `DESC NOTIFICATION INTEGRATION` access to the storage queue. Create the storage integration, stage, and pipe: ``` CREATE STORAGE INTEGRATION aisix_az_int TYPE = EXTERNAL_STAGE STORAGE_PROVIDER = 'AZURE' ENABLED = TRUE AZURE_TENANT_ID = 'YOUR_TENANT_ID' STORAGE_ALLOWED_LOCATIONS = ('azure://YOUR_STORAGE_ACCOUNT.blob.core.windows.net/ai-gateway/ai-gateway/'); DESC INTEGRATION aisix_az_int; CREATE STAGE aisix_stage URL = 'azure://YOUR_STORAGE_ACCOUNT.blob.core.windows.net/ai-gateway/ai-gateway/' STORAGE_INTEGRATION = aisix_az_int; CREATE PIPE aisix_events_pipe AUTO_INGEST = TRUE INTEGRATION = 'AISIX_AZ_NOTIF' AS COPY INTO gateway_events (record, source_file) FROM (SELECT $1, METADATA$FILENAME FROM @aisix_stage) FILE_FORMAT = (TYPE = JSON COMPRESSION = AUTO); ``` Grant the Snowflake service principal returned by `DESC INTEGRATION` read access to the blob container. ## Query Request Telemetry[​](#query-request-telemetry "Direct link to Query Request Telemetry") After sending more traffic through AISIX, check that Snowpipe is receiving files: ``` SELECT SYSTEM$PIPE_STATUS('aisix_events_pipe'); ``` Confirm rows arrived and create a view over the raw records: ``` SELECT COUNT(*) FROM gateway_events; CREATE OR REPLACE VIEW gateway_attempts AS SELECT record:request_id::string AS request_id, record:occurred_at::timestamp_tz AS occurred_at, record:requested_model::string AS requested_model, record:model_id::string AS model_id, record:attempt_index::number AS attempt_index, record:attempt_kind::string AS attempt_kind, record:status_code::number AS status_code, record:prompt_tokens::number AS prompt_tokens, record:completion_tokens::number AS completion_tokens, record:cost_usd::float AS cost_usd, record:upstream_latency_ms::number AS upstream_latency_ms, record:upstream_ttft_ms::number AS upstream_ttft_ms, record:downstream_latency_ms::number AS downstream_latency_ms, record:cache_status::string AS cache_status, record:guardrail_blocked::boolean AS guardrail_blocked, record:finish_reason::string AS finish_reason, source_file FROM gateway_events; ``` Use `downstream_latency_ms` for the time experienced by the caller; it appears only on the terminal attempt. Use `upstream_latency_ms` to analyze each provider attempt, and `upstream_ttft_ms` for the time until that attempt streamed its first frame. The time-to-first-frame field is omitted or zero for non-streaming requests, errors, and cache hits. Query recent attempts: ``` SELECT requested_model, status_code, prompt_tokens, completion_tokens, cost_usd, cache_status FROM gateway_attempts ORDER BY occurred_at DESC LIMIT 10; ``` Run a simple cost and token summary: ``` SELECT requested_model, COUNT(DISTINCT request_id) AS requests, COUNT(*) AS attempts, SUM(prompt_tokens + completion_tokens) AS total_tokens, ROUND(SUM(cost_usd), 5) AS total_cost_usd FROM gateway_attempts GROUP BY requested_model ORDER BY total_cost_usd DESC; ``` ## Cleanup[​](#cleanup "Direct link to Cleanup") Drop the Snowflake objects when you finish testing: ``` DROP PIPE IF EXISTS aisix_events_pipe; DROP STAGE IF EXISTS aisix_stage; DROP VIEW IF EXISTS gateway_attempts; DROP TABLE IF EXISTS gateway_events; DROP STORAGE INTEGRATION IF EXISTS aisix_s3_int; DROP STORAGE INTEGRATION IF EXISTS aisix_az_int; DROP NOTIFICATION INTEGRATION IF EXISTS aisix_az_notif; ``` Remove the bucket notification or Event Grid subscription you created. With AISIX Cloud, delete the exporter through the Admin API: ``` curl -sS -X DELETE "$AISIX_CP/environments/$ENV_ID/observability_exporters/$EXPORTER_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" ``` For the open-source AISIX gateway, remove the exporter entry from `resources.yaml`, then validate and reload the file. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now loaded AISIX request telemetry into Snowflake. Continue with [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md) when you need to correlate exported usage events with runtime metrics, access logs, and response headers. --- # Metrics and Logs AISIX AI Gateway exposes aggregate metrics, per-request logs, response headers, and exportable usage events. Together, these signals show service health and connect a caller-visible response to the model, route, and policy outcome that produced it. ## Choose a Telemetry Source[​](#choose-a-telemetry-source "Direct link to Choose a Telemetry Source") Start with the source that matches the operational question, then correlate signals by request, model, or provider when an issue needs deeper investigation. | Operational Question | Start With | What It Provides | | ---------------------------------------------------------------------- | ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------- | | Is traffic healthy across gateway instances? | Prometheus metrics | Request rates, latency distributions, token and cost counters, policy outcomes, routing health, cache behavior, and exporter delivery health. | | What happened to one request? | Access logs | Structured request fields, including status, latency, model, provider, request ID, and routing outcome when available. | | What can the calling application observe? | Response headers | Request correlation, cache outcome, retry timing, and selected-target hints on supported routes. | | Where can request records be stored, analyzed, or used for accounting? | Usage events | Per-attempt outcome and consumption records delivered through an observability exporter. | ## Scrape Prometheus Metrics[​](#scrape-prometheus-metrics "Direct link to Scrape Prometheus Metrics") Prometheus metrics are suited for monitoring traffic trends, latency, token counters, cost counters, rate-limit outcomes, cache behavior, and exporter delivery health. AISIX serves Prometheus metrics on the dedicated metrics listener at `/metrics` by default. Change the path or disable the endpoint through the startup observability settings. This endpoint is unauthenticated by design. Keep the dedicated metrics listener private. Configure Prometheus exposure in the startup configuration: config.yaml ``` observability: metrics: prometheus: enabled: true path: "/metrics" ``` The dedicated listener binds to `0.0.0.0:9090` by default. Set a different listener address when Prometheus should scrape another interface or port. Scrape the default metrics endpoint: ``` curl -sS "http://127.0.0.1:9090/metrics" ``` Traffic metrics appear after activity AISIX publishes configuration status on every scrape. Other metric families are registered on first observation, so traffic metrics might not appear immediately after startup. Send one model request, then check again for series such as `aisix_requests_total` and `aisix_tokens_consumed_total`. AISIX emits native metric names with the `aisix_` prefix. Use the histogram series for latency percentiles across gateway instances and the request counters for success-rate and routing analysis. For exact metric names, label scope, and PromQL examples, see [Metrics Reference](https://docs.api7.ai/ai-gateway/reference/metrics.md). ## Count Requests and Attempts Separately[​](#count-requests-and-attempts-separately "Direct link to Count Requests and Attempts Separately") A request and an upstream attempt are different units, and mixing them is the most common source of apparently contradictory numbers. One client request can place several upstream calls: a same-target retry, or a failover to the next target in a Model Group. AISIX counts both units, in different places. | Unit | Where it is counted | What one sample or row means | | ---------------- | ------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- | | Request | `aisix_proxy_requests_total`, `aisix_llm_requests_total` | One client request, with the status the caller received. | | Attempt | `aisix_deployment_requests_total` and the deployment families | One upstream call to one target model. | | Emission attempt | `aisix_usage_events_emitted_total` | One attempt to enqueue a usage event, counted before the delivery queue accepts or rejects it. | | Attempt | Usage-event records, and the usage log built from them | One attempt, joined to its siblings by `request_id` and ordered by `attempt_index`. | The consequence worth internalizing: **a failed attempt that a fallback rescued is invisible in the request counters.** The caller received a 200, so the request counter records `status="200"` and nothing else. The 502 the first target returned lives in the deployment counters and in the usage log, as its own attempt. So a gateway fronting one broken target in an otherwise healthy Model Group can legitimately show tens of thousands of `5xx` rows in the usage log while `sum(increase(aisix_proxy_requests_total{status=~"5.."}[24h]))` returns a few hundred. The log is counting attempts that failed; the metric is counting callers who saw a failure. Neither is wrong, and they are not expected to converge. Use these queries for the questions they actually answer: ``` # Requests that ended in a server error — what callers experienced. sum(increase(aisix_proxy_requests_total{status=~"5.."}[24h])) # Requests that ended in a server error even after failing over. sum(increase(aisix_proxy_requests_total{status=~"5..", is_fallback="true"}[24h])) # Upstream attempts that failed, by target — including the ones a fallback # rescued. This is the attempt-level view of a 5xx usage-log row count, # restricted to the endpoints that dispatch through a Model Group. sum(increase(aisix_deployment_failure_responses_total[24h])) by (model) # Fallbacks that rescued a request, by group and by the target reached. sum(increase(aisix_routing_successful_fallbacks_total[24h])) by (model, fallback_model) # Usage-event emission attempts, and those the handoff queue rejected. sum(increase(aisix_usage_events_emitted_total{status_code="5xx"}[24h])) sum(increase(aisix_usage_event_drops_total[24h])) # One member's rate-limit rejections. `status` carries the raw HTTP code, # so a single failure mode is addressable without scanning a whole family. sum(increase(aisix_usage_events_emitted_total{user_id="<member-id>", status="429"}[24h])) # Whose queue handoffs were rejected. Both counters carry the same model and # provider-key labels, so the difference holds per model, not only in total. sum(increase(aisix_usage_event_drops_total[24h])) by (model, provider_key_name) ``` Before concluding that two sources disagree, confirm they cover the same population: * **Scrape coverage.** Request counters carry no environment label, so a per-environment figure from usage records is not comparable with a gateway-wide PromQL result unless every gateway instance serving that environment is scraped. `sum by (job, instance) (...)` shows which instances contributed; compare that against the instances actually running. * **Window coverage.** `increase(...[24h])` measures only the samples present in the range. If instances restarted partway through the window, or Prometheus retention is shorter than the window, the result covers less time than the usage-log filter does. Graphing the raw counter over the same range makes both gaps visible. * **Delivery.** `aisix_usage_event_drops_total` measures events rejected at the producer-to-worker queue handoff. After the queue accepts an event, delivery can still fail without incrementing this counter, so it does not bound the difference between emission attempts and records in final storage. Group queue drops by `model`. For later failures, monitor `telemetry batch failed (events dropped)` for control-plane delivery and `sink delivery dropped after retries` for exporter delivery. * **Attempts that never left the gateway.** The deployment families count upstream calls, so an attempt that produced no upstream call is deliberately absent from them: one the target's own rate limits refused, and one rejected while the request was still being assembled — an unusable credential, a missing `model_name`, a missing or malformed `api_base`. Those attempts still appear in the usage log and in `aisix_usage_events_emitted_total`. This is therefore one of the reasons a deployment attempt count runs below a usage-log attempt count, not the whole of it: the scope noted above is another (the deployment families cover only endpoints dispatched through a Model Group, while usage events also cover direct models), as are the scrape, window, and delivery gaps in this list. Isolate one contributor at a time rather than reading the whole difference as gateway-side rejections. The exclusion does have one unambiguous signature: a misconfigured target produces failed *requests* with no deployment failures at all, because the provider was never asked anything. ## Collect Access Logs[​](#collect-access-logs "Direct link to Collect Access Logs") Access logs describe an individual proxy request. AISIX writes them through the process logger to the standard error stream, so they appear in the container or process logs collected by your runtime. Configure process logging in the startup configuration: config.yaml ``` observability: log_level: "info" ``` The `RUST_LOG` environment variable can override the configured log level. One access log entry is written per proxy request, whether it succeeded, failed, or stopped early. It includes fields such as method, path, status, latency, provider, model, API key ID, request ID, token counts, and routing outcome when those values are available. Its timing differs by response type, which determines which fields it can carry. A non-streamed request is written when the request ends, so its entry carries everything the gateway resolved. A streamed one is written when the response is opened, before the stream is consumed — so its entry has no token counts and no provider response ID, because neither exists yet. Usage events, described below, carry the streamed figures. When the gateway has it by the time the entry is written, the entry also carries `provider_request_id`: the response object ID the provider returned, such as an OpenAI `chat.completion.id`, an Anthropic message `id`, or a Responses API `resp_…`. This is the ID a provider's own console and support channel index a call by. The field is omitted rather than blank whenever there is none to record, so you can filter on its presence. Besides the streamed case above, that covers a rejection before dispatch such as a guardrail block, a response served from cache, and the endpoints whose provider response carries no ID at all, including embeddings, audio, and image generation. For the calls whose ID cannot reach the access log entry, the gateway emits a separate `provider call completed` entry, carrying `request_id`, `attempt_index`, `attempt_kind`, and `provider_request_id`. One is written per provider call that returned an ID — so a streamed response produces one, and a request that failed over part-way through a stream produces one per provider call it made. Join them to the access log entry by `request_id`, and tell one call apart from another within a retried or failed-over request by `attempt_index`. Two further fields describe the connection a request arrived on rather than the request itself. They are attached to the request as a whole, so they appear on the access log entry, on the `provider call completed` entry, and on every diagnostic line the request emits in between: | Field | What it records | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `peer` | The remote end of the accepted downstream connection, as `<ip>:<port>`. | | `downstream_request_id` | The correlation ID that arrived in the `x-request-id` request header — what reverse proxies, ingress controllers, and service meshes stamp on the requests they forward. | Both are recorded rather than declared: a request that carried neither logs neither field instead of an empty one, so you can filter on presence. `peer` is not the caller's IP address, and substituting one for the other loses the point of both. The gateway resolves the caller's address separately, following `proxy.real_ip` when it sits behind trusted proxies, and uses it for access control and usage records; that value carries no port. `peer` is the remote end of the TCP connection the request physically arrived on, and its port is the whole of its value. Behind a layer-4 load balancer, with the gateway on the host network, `peer` is the only field that joins a gateway log line to the fronting proxy's own record of the same connection. A streamed response is the one case where that attachment could quietly lapse: its body is written after the point at which a request's log context normally ends. The streaming paths carry that context forward explicitly, so the lines written at the end of a stream — `provider call completed` among them — carry the same `request_id`, `peer`, and `downstream_request_id` as the rest of the request. A caller that disconnects before the gateway sends response headers also produces an entry, with status `499` and `error_kind="client_disconnected"`. It carries the model and provider the request had resolved by then, and no token fields — the request never reached a response to count tokens from. A request abandoned earlier still, before the model was resolved, carries neither. Requests recorded this way are also counted by [`aisix_proxy_client_cancelled_requests_total`](https://docs.api7.ai/ai-gateway/reference/metrics.md#request-metrics). A caller that disconnects mid-response is not recorded this way, because that request already has a normal status and a usage event. A request refused for exceeding `proxy.request_body_limit_bytes` produces a second entry on the `aisix::body_limit` target, carrying `declared_content_length`, `configured_limit_bytes`, `drained_bytes`, and `drain_outcome`. Join it to the access log entry by `request_id`. `drain_outcome` reports how the gateway finished reading the body it refused, which determines what the caller saw. `completed` means the caller sent everything it declared and could read the `413`. `cap_reached`, `timeout`, and `client_read_error` mean the read stopped early, so the caller usually sees a closed connection instead. The first is logged at `info`; the other three are logged at `warn` and are rate limited to one entry per outcome per second, so use [`aisix_proxy_request_body_limit_rejections_total`](https://docs.api7.ai/ai-gateway/reference/metrics.md#request-metrics) for volume. The `access_log` field is currently reserved and has no effect. Proxy handlers still emit structured access logs, and there is no separate access-log format or sink setting. Collect the standard error stream with your runtime log pipeline when logs need to leave the gateway host. ## Correlate Responses with Telemetry[​](#correlate-responses-with-telemetry "Direct link to Correlate Responses with Telemetry") Response headers provide caller-visible correlation and routing hints. They can identify the request, cache outcome, retry timing, or selected target on supported paths. Use the request ID and other supported response headers to join a caller-visible response to access logs or exported records. For the header scope on each proxy route, see [Headers and Error Codes](https://docs.api7.ai/ai-gateway/reference/headers-and-error-codes.md#proxy-response-headers). Up to three IDs travel with a call, and none of them substitutes for another: | ID | Assigned by | Use it to | | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `request_id` | The gateway, one per request — or the caller, when it supplies one and AISIX adopts it (see below). Returned in the `x-aisix-request-id` response header. | Correlate the response with any gateway access logs and usage events the request produces, including exported records or entries on the dashboard's Logs page. It means nothing to the provider. | | `downstream_request_id` | Whatever sits in front of the gateway, in the `x-request-id` request header. Recorded on the request's log lines when one arrived; never returned to the caller. | Find the same request in the logs of your ingress controller, reverse proxy, service mesh, or CDN. | | `provider_request_id` | The provider, in its response body, when it sends one. Recorded per attempt on the usage event for that attempt, and on the log entries described above. | Locate the same call in the provider's own console or quote it to provider support. | A caller reporting a problem normally has `request_id`, which the gateway returns on every response. Look that request up, read `provider_request_id` from the attempt that served it, and take that to the provider. A caller who kept the whole response body may be able to read the provider's ID out of it directly, but only on the endpoints that return one. `downstream_request_id` is **recorded, never adopted.** Whether a caller-supplied ID becomes the gateway's own `request_id` is a separate decision, governed by [`proxy.request_id.accept_headers`](#choose-the-accepted-headers) — which by default accepts only `x-aisix-request-id`. So a gateway behind an ingress normally logs two IDs that mean different things, and neither can be mistaken for the other. It is screened by the same rule as an adopted ID (see [Accepted Values](#accepted-values)) and omitted when the header is absent or unusable. ### Reuse Your Own Request ID[​](#reuse-your-own-request-id "Direct link to Reuse Your Own Request ID") If your service already generates a request ID for the business call, send it and AISIX adopts it instead of generating one. That ID then becomes the request's identity everywhere: the `x-aisix-request-id` response header, the `request_id` on the access log and on every usage event the request produces, and the `x-aisix-request-id` the provider receives. Every trail you use to investigate a request is then keyed by an ID your own application logs already carry. That covers the gateway access log, exported usage events, and AISIX Cloud [Request Logs](https://docs.api7.ai/ai-gateway/cloud/logging-and-auditing.md#request-logs), with no second mapping to maintain. Send it in `x-aisix-request-id`: ``` # AISIX_PROXY is the gateway origin; omit a trailing slash and endpoint path. export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" curl "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -H "x-aisix-request-id: req_abc123-orders-svc" \ -d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Hello"}]}' -i ``` The response echoes the same value: ``` HTTP/1.1 200 OK x-aisix-request-id: req_abc123-orders-svc ``` A request that retries or fails over emits one usage event per attempt. All carry your ID, so filtering on it returns the whole chain rather than only the attempt that succeeded. #### Accepted Values[​](#accepted-values "Direct link to Accepted Values") An ID is used as sent when it is 1–256 bytes and contains only visible ASCII characters (`!` through `~`): no spaces, no control characters, no non-ASCII. Qualifying shapes include a UUID, a ULID, a prefixed ID such as `req_abc123`, and the hexadecimal ID nginx puts in `$request_id`. A value outside that range is ignored and AISIX generates a UUID instead, which is the behavior when no ID is sent at all. The request itself is never rejected over its correlation ID. AISIX does not require IDs to be unique. Sending the same ID for two different requests makes them indistinguishable in every trail that keys on it, so generate a fresh one per request. #### Choose the Accepted Headers[​](#choose-the-accepted-headers "Direct link to Choose the Accepted Headers") By default AISIX reads only its own `x-aisix-request-id`. Set `proxy.request_id.accept_headers` to change that: config.yaml ``` proxy: request_id: accept_headers: ["x-aisix-request-id", "x-request-id"] ``` Headers are consulted in the order listed, and the first acceptable value wins, so the list is also a priority order. Environment-only deployments set the list comma-separated as `AISIX_PROXY__REQUEST_ID__ACCEPT_HEADERS`. `x-request-id` is not accepted by default on purpose. Reverse proxies, ingress controllers, and load balancers stamp that header on every request they forward. Enabling it behind one of those means the correlation ID comes from your infrastructure rather than from the calling service. Add it when AISIX is the first hop, or when the ID your proxy assigns is the one you want to trace by. Accepting it is not what makes it visible, though. An acceptable `x-request-id` is recorded as `downstream_request_id` on the request's log lines either way, so the infrastructure's own ID is available to search on without changing this setting. Listing it here does something further: it makes that value the request's identity everywhere — the response header, the access log's `request_id`, and every usage event — at which point `request_id` and `downstream_request_id` hold the same value, which is the honest report of what happened. Set `accept_headers: []` to ignore caller-supplied IDs entirely and always generate one. A header name that is not a valid HTTP header name fails startup rather than being silently skipped. ## Export Usage Events[​](#export-usage-events "Direct link to Export Usage Events") Usage events are per-attempt records emitted by supported proxy paths. A request that retries or fails over emits multiple events with the same `request_id`, ordered by `attempt_index`. Counting these records therefore counts attempts, not requests — see [Count Requests and Attempts Separately](#count-requests-and-attempts-separately) before comparing a record count against a request metric. Each event includes its outcome, consumption details, the model alias the caller requested, and the resolved model that served the attempt when the gateway can observe those values. A response served from the response cache records the matching layer in `cache_hit_layer` (`exact` or `semantic`) and, for semantic hits, the matched similarity in `cache_similarity`. Latency fields distinguish provider time from caller-visible time: | Field | Scope | | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `upstream_latency_ms` | Time spent on one upstream attempt. It excludes request parsing, guardrails, routing, retry delays, and earlier attempts. | | `upstream_ttft_ms` | Time from the start of an upstream attempt until its first streamed frame of any type — metadata openers such as `response.created` or a role-only chat delta included, matching what caller-side proxies measure. It is omitted or zero for non-streaming requests, errors, and cache hits. | | `downstream_latency_ms` | Total time the caller waited for the request. It includes gateway processing, retries and retry delays, and any output holdback, and appears only on the terminal attempt. | A caller that disconnects part-way through a streamed response still produces a usage event, because the upstream did the work and may have charged for it. That event carries status `499` rather than `200`, and its token counts cover only what arrived before the disconnect. Count streamed successes with `status_code = 200` to keep abandoned responses out of the total. ### Tell Request Kinds Apart[​](#tell-request-kinds-apart "Direct link to Tell Request Kinds Apart") Every event carries `operation`, the kind of work the request asked for. Without it, one field says which protocol addressed the gateway and another says which model answered. Neither says whether the call was a conversation, an image, or a video. Every OpenAI-compatible endpoint reports the same `inbound_protocol`, so a chat completion, an image generation, and a video submission arrive as one undifferentiated stream. The value comes from the endpoint the request matched, so it is a small fixed set safe to index, group, and chart: | Value | Endpoint | | --------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | | `chat` | `/v1/chat/completions` | | `messages` | `/v1/messages` | | `count_tokens` | `/v1/messages/count_tokens` | | `responses` | `/v1/responses` | | `completions` | `/v1/completions` | | `embeddings` | `/v1/embeddings` | | `rerank` | `/v1/rerank` | | `image_generation` | `/v1/images/generations` | | `image_edit` | `/v1/images/edits` | | `transcription` | `/v1/audio/transcriptions` | | `translation` | `/v1/audio/translations` | | `speech` | `/v1/audio/speech` | | `video_generation` | `POST /v1/videos` | | `realtime` | `/v1/realtime` | | `files`, `batches`, `fine_tuning` | The file, batch, and fine-tuning management endpoints | | `batch_completion` | The gateway's own accounting of a finished batch job, recorded when the job completes rather than when it was submitted | | `mcp`, `a2a` | The MCP and A2A gateways | | `passthrough` | A passthrough route | Three points are worth knowing before you write a query against it: * **It describes the request, not the outcome.** A request that failed, or that a guardrail refused, carries the same value a successful one would. It is the only field on such a record that names the endpoint at all. * **It is request-scoped.** A request that retries or fails over emits one event per attempt, and every attempt carries the same value, so counting events per operation counts attempts. See [Count Requests and Attempts Separately](#count-requests-and-attempts-separately). * **Polling a video job is not video generation.** Only the submission (`POST /v1/videos`) produces a usage event; retrieving the job's status or downloading its result does not. A count of `video_generation` is therefore a count of videos asked for, not of requests made about them. The field is part of the record's metadata, so it is present under both content modes. An exporter configured for `metadata_only` — which never receives a prompt — can still separate traffic by kind, which content inspection could not do there at all. In Alibaba Cloud SLS, the field arrives as its own column and needs no parsing. Enable analytics for the fields a query names, since SQL analysis reads indexed fields only: ``` * | SELECT operation, COUNT(*) AS calls, SUM(prompt_tokens + completion_tokens) AS tokens GROUP BY operation ORDER BY calls DESC ``` To pull one kind of traffic, filter on it directly — `operation: video_generation` — instead of matching a model name or searching the prompt. A model name is a poor substitute here. One model can serve several endpoints, and callers address models through aliases and model groups, so `requested_model` answers "which configured entry," not "what kind of call." ### Attribute Events to a Member[​](#attribute-events-to-a-member "Direct link to Attribute Events to a Member") Each event carries `user_id`: the organization member who owned the API key the request authenticated with, assigned on the key itself. It is absent for keys that belong to no member — ownership is set explicitly when a key is created or edited, not inferred from whoever created it. One member usually holds several credentials, and this field is what makes them one identity. A JWT-authenticated request runs as the API key its identity resolves to, so a member calling through both an API key and an OIDC token produces events under two different `api_key_id` values and the same `user_id`. Filtering on the member returns the whole picture; filtering per key returns a slice of it. The value is a snapshot taken when the request ran, not a lookup performed when you read the event. Reassigning a key to a different member changes who later requests are attributed to and leaves earlier events attributed to the previous owner, which is what makes a historical query answerable. Deleting the key does not erase the attribution of the events it already produced. Events written before a data plane that records this field carry no member, so a member filter covers traffic from that upgrade forward. In the dashboard, **Logs** offers this as the **Member** filter, alongside a **Status** filter that accepts a family (`4xx`), an exact code (`429`), or a range (`500-599`). Combining the two answers questions like "which of this member's requests were rate-limited in the last 24 hours" in one query. The CSV export carries `user_id` and the member's name. Usage events are consumed through a sink rather than read from a local endpoint. Configure an observability exporter to deliver them to OTLP/HTTP, object storage, Alibaba Cloud SLS, or Datadog. ## Next Steps[​](#next-steps "Direct link to Next Steps") Configure [Observability Exporters](https://docs.api7.ai/ai-gateway/observability/exporters.md) to send usage events to an external collector, log destination, object store, or warehouse workflow. Use the [Metrics Reference](https://docs.api7.ai/ai-gateway/reference/metrics.md) when building Prometheus dashboards or alerts. --- # On-Premises Installation Install the AISIX Cloud control plane in infrastructure you operate using Docker Compose, Helm, or an offline package. This guide helps you choose an installation method, configure the control-plane endpoints, and verify a persistent environment. The installation provides the same [AISIX Cloud control-plane workflows](https://docs.api7.ai/ai-gateway/cloud/overview.md) as Hybrid Cloud, including resource management, gateway certificate issuance, usage reporting, and budget enforcement. The control-plane services and data remain in your infrastructure, including in a fully air-gapped environment. For a local evaluation that continues through creating gateway resources and sending a first AI request, use the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md) instead. Licensing Production use of the AISIX Cloud control plane and dashboard requires a commercial license. Before deploying the control plane in production, [contact API7](https://api7.ai/contact) or email <support@api7.ai>. ## Plan the Installation[​](#plan-the-installation "Direct link to Plan the Installation") Choose an installation method based on the target infrastructure and its network access: | Installation method | Target | Requirements | | ------------------- | ------------------------------------- | ------------------------------------------------------------------------------------------------------------- | | Docker Compose | Internet-connected host | Docker with Docker Compose V2, cURL, `tar`, OpenSSL, and access to Docker Hub | | Helm | Internet-connected Kubernetes cluster | A working cluster, Helm, `kubectl`, OpenSSL, and access to Docker Hub | | Offline package | Air-gapped host | Docker with Docker Compose V2, `tar`, OpenSSL, and a separate machine with cURL that can download the package | Before installing with Helm, decide whether to use the bundled PostgreSQL database or an [external database](https://docs.api7.ai/ai-gateway/on-premises/external-database.md). Docker Compose and offline package installations use the bundled database. Also choose the public dashboard origin and the data-plane manager endpoint that gateways will reach. You can start with local endpoints, but configure externally reachable endpoints before exposing the dashboard or attaching gateways on other hosts. Docker Compose installations use host ports `5432`, `8080`, and `7944` by default. Make sure they are available or configure different host bindings. See the [Port Reference](https://docs.api7.ai/ai-gateway/reference/ports.md#on-premises-control-plane-ports) for each component, traffic direction, and recommended exposure. ## Control Plane Components[​](#control-plane-components "Direct link to Control Plane Components") Every installation method uses these components: | Service | Role | | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cp-api` | Manages organizations, environments, resources, and billing | | `dp-manager` | Issues mTLS certificates and delivers configuration to data planes | | `dashboard` | Web console | | PostgreSQL database | Shared datastore. Docker Compose packages include it; Helm can deploy it or connect to an [external database](https://docs.api7.ai/ai-gateway/on-premises/external-database.md). | The online and offline Docker Compose packages require the bundled PostgreSQL 16 service. The Helm chart can deploy a bundled PostgreSQL instance or connect to an external database. AISIX gateways run separately as data planes. They connect outbound to `dp-manager` over mTLS, so the control plane does not need inbound network access to gateway hosts. ## Production Resource Baseline[​](#production-resource-baseline "Direct link to Production Resource Baseline") The following specifications are deployment-planning starting points, not Helm chart defaults, benchmark-derived capacity guarantees, or fixed minimums for every workload. Benchmark and adjust them for your request rate, request size, enabled traffic controls, number of gateways, and data retention periods. ### Docker Compose or Offline Package on Hosts[​](#docker-compose-or-offline-package-on-hosts "Direct link to Docker Compose or Offline Package on Hosts") The online and offline Docker Compose packages provide a single-host topology with one instance of each control-plane service and a bundled PostgreSQL database. They do not support an external database or expose the multi-host service configuration required for control-plane high availability. For an HA production topology, use Helm and the instance counts below. For a package-based evaluation or non-HA deployment, use the CPU, memory, and storage columns to assess host capacity. CPU, memory, and host storage are per component instance. Host storage is a capacity allowance for the operating system, container images, platform-managed logs, and local component state; it is not equivalent to an application persistent-volume requirement in Kubernetes. | Component | CPU | Memory | Host storage | Starting instances | Sizing guidance | | --------------------- | --------- | ------- | ---------------- | ----------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | AISIX gateway | 4 cores | 8–16 GB | ≥100 GB | 3 minimum; 4 recommended per gateway deployment | Scale horizontally with traffic. Start with 16 GB for high concurrency or guardrail workloads. Increase host storage for longer local log retention. | | `cp-api` | 2 cores | 4 GB | ≥32 GB | 2 | Stateless; place instances in separate failure domains. | | `dp-manager` | 2–4 cores | 4–8 GB | ≥32 GB | 2 | Delivers configuration and processes gateway heartbeats and usage telemetry. Place instances in separate failure domains. | | `dashboard` | 1–2 cores | 2 GB | ≥20 GB | 2 | Stateless; place instances in separate failure domains. | | PostgreSQL | 4–8 cores | 16 GB | ≥500 GB NVMe SSD | 3 | Critical stateful component. For production HA, use Helm with an external PostgreSQL deployment, such as a primary, synchronous standby, and asynchronous standby. Adjust storage for the usage-event retention period. | | Prometheus (optional) | 4 cores | 8–16 GB | ≥200 GB | 1–2 | Stores monitoring metrics. Reuse an existing Prometheus deployment when available, and size storage for metric cardinality and retention. | For smaller deployments, `cp-api`, `dp-manager`, `dashboard`, and Prometheus can share hosts if you reserve CPU and memory for each component and provide enough shared host capacity for container images and logs. Keep redundant instances in separate failure domains. ### Helm on Kubernetes[​](#helm-on-kubernetes "Direct link to Helm on Kubernetes") For Kubernetes, plan compute per replica and scale the AISIX gateway independently using [Performance and Sizing](https://docs.api7.ai/ai-gateway/deployment/performance-and-sizing.md). The following values are planning recommendations rather than chart defaults: | Component | CPU per replica | Memory per replica | Starting replicas | Sizing guidance | | --------------------- | --------------- | ------------------ | ----------------------------------------------- | ---------------------------------------------------------------------------------------------- | | AISIX gateway | 4 cores | 8–16 GB | 3 minimum; 4 recommended per gateway deployment | Scale horizontally with traffic. Start with 16 GB for high concurrency or guardrail workloads. | | `cp-api` | 2 cores | 4 GB | 2 | Stateless; place replicas in separate failure domains. | | `dp-manager` | 2–4 cores | 4–8 GB | 2 | Place replicas in separate failure domains. | | `dashboard` | 1–2 cores | 2 GB | 2 | Stateless; place replicas in separate failure domains. | | PostgreSQL | 4–8 cores | 16 GB | 3 for production HA | Use an external HA deployment rather than the bundled single-instance database. | | Prometheus (optional) | 4 cores | 8–16 GB | 1–2 | Reuse an existing Prometheus deployment when available. | The default control-plane Helm chart does not assign application persistent volumes to `cp-api`, `dp-manager`, or `dashboard`. Account for container images and platform-managed logs as Kubernetes node capacity instead of per-pod storage. If you operate PostgreSQL or Prometheus in the cluster, provision persistent storage separately and size it for usage-event or metric retention. As starting points, allocate at least 500 GB of NVMe SSD storage per PostgreSQL instance and 200 GB per Prometheus instance. Do not place PostgreSQL replicas in the same failure domain or on shared storage. The bundled PostgreSQL chart does not provide a replicated database by default; see [High Availability](https://docs.api7.ai/ai-gateway/cloud/high-availability.md) and [External Database](https://docs.api7.ai/ai-gateway/on-premises/external-database.md). These specifications cover AISIX components only, not the compute resources required by upstream model inference. AISIX itself does not require a GPU. ## Install Online[​](#install-online "Direct link to Install Online") Use an online installation when the host or Kubernetes cluster can pull container images from Docker Hub. ### Docker Compose on a Host[​](#docker-compose-on-a-host "Direct link to Docker Compose on a Host") Use Docker Compose for an internet-connected host with Docker and Docker Compose. Download the package for this release: ``` curl -fSL "https://run.api7.ai/aisix-self-hosted/aisix-self-hosted-1.0.0.tar.gz" \ -o aisix-self-hosted-1.0.0.tar.gz tar -xzf aisix-self-hosted-1.0.0.tar.gz ./aisix-self-hosted/run.sh ``` The startup script generates a `.env` file with fresh secrets, pulls images from Docker Hub, and starts the stack. The package uses its bundled PostgreSQL service and does not support an external database. Use the [Helm installation](#helm-on-kubernetes) when an external PostgreSQL database is required. When startup finishes, the script prints the dashboard URL. The default URL is `http://localhost:8080`. Open the dashboard and create the first admin account. Manage the stack from `./aisix-self-hosted`: ``` ./aisix-self-hosted/run.sh logs # tail logs ./aisix-self-hosted/run.sh stop # stop containers ./aisix-self-hosted/run.sh down # remove containers (keeps the data volume) ``` ### Helm on Kubernetes[​](#helm-on-kubernetes-1 "Direct link to Helm on Kubernetes") For Kubernetes, install the chart from the API7 Helm repository: ``` helm repo add api7 https://charts.api7.ai helm repo update helm install aisix-cp api7/aisix-cp --version 1.0.0 \ --set secrets.masterKey="$(openssl rand -base64 32)" \ --set secrets.betterAuthSecret="$(openssl rand -base64 48)" \ --set postgresql.auth.password="$(openssl rand -hex 24)" \ --set postgresql.auth.postgresPassword="$(openssl rand -hex 24)" ``` The chart deploys the core API, data-plane manager, dashboard, and a bundled PostgreSQL instance by default. To reach the dashboard before configuring external access, forward the `cp-api` service: ``` kubectl port-forward svc/aisix-cp-api 8080:8080 ``` Open `http://localhost:8080` while the port-forward is running. To use an existing database, first [provision the external database and role](https://docs.api7.ai/ai-gateway/on-premises/external-database.md). Then disable the bundled PostgreSQL chart with `postgresql.builtin=false` and configure the top-level `externalDatabase.*` values. To inspect the default chart values locally, run: ``` helm show values api7/aisix-cp --version 1.0.0 ``` The chart source and package are published in the [`aisix-cp-1.0.0` Helm chart release](https://github.com/api7/api7-helm-chart/releases/tag/aisix-cp-1.0.0). warning Use URL-safe database passwords, such as values generated with `openssl rand -hex 24`. The database password is embedded in a `postgres://` connection URL, so characters such as `+`, `/`, and `=` from `openssl rand -base64` can break the URL. ## Install in an Air-Gapped Environment[​](#install-in-an-air-gapped-environment "Direct link to Install in an Air-Gapped Environment") For a host with no registry access, use the offline package. It includes every required container image. The offline package URL is pinned to AISIX 1.0.0. On a machine with internet access, download the package: ``` curl -fSL "https://run.api7.ai/aisix-self-hosted/aisix-self-hosted-offline-1.0.0.tar.gz" \ -o aisix-self-hosted-offline-1.0.0.tar.gz ``` Transfer the package to the air-gapped host, then start the stack: ``` tar -xzf aisix-self-hosted-offline-1.0.0.tar.gz cd aisix-self-hosted ./run.sh ``` The startup script: * loads the bundled container images * generates a `.env` file with fresh secrets * starts the stack without internet access * prints the dashboard URL when startup finishes The default dashboard URL is `http://localhost:8080`. The `cp-api` image included in the package contains a model-pricing snapshot so usage and budget calculations can initialize without reaching `models.dev`. At startup, the control plane loads the model-pricing catalog from that snapshot. To use online pricing instead, set `AISIX_CLOUD_PRICESYNC_SNAPSHOT_PATH=` in `.env` and recreate the `api` service. See [On-Premises Configuration](https://docs.api7.ai/ai-gateway/reference/on-premises-configuration.md#pricing-catalog) for the pricing catalog settings. ## Configure External Access[​](#configure-external-access "Direct link to Configure External Access") Before exposing the control plane outside the local host or cluster, configure the public dashboard origin and the data-plane manager endpoint. For Docker Compose, edit `.env`. For Kubernetes, update your Helm values: | Docker Compose setting | Helm value | Purpose | | ----------------------------- | ------------------- | ------------------------------------------------------------------------------------------------------------------ | | `AISIX_CLOUD_PUBLIC_BASE_URL` | `api.publicBaseURL` | Browser-facing origin, such as `https://aisix.example.com`. Login validates the session issuer against this value. | | `AISIX_CLOUD_DPMGR_BASE_URL` | `api.dpmgrBaseURL` | `dp-manager` mTLS endpoint that data-plane hosts connect to. A DNS name or an IP address. | The Helm chart exposes the `cp-api`, `dp-manager`, and dashboard services as `ClusterIP` by default. Expose the `cp-api` and `dp-manager` services through network endpoints appropriate for your cluster, then set the two public values above to those endpoints. `dp-manager` receives the data-plane manager endpoint too, and issues its TLS server certificate for that host. Data planes verify the certificate against the address they dial, so the value must match the endpoint in the generated gateway install command — whether that is a DNS name or an IP address. After updating these settings, recreate the affected services with `docker compose up -d` or apply the changes with `helm upgrade`. ### Trusted Sign-In Origins[​](#trusted-sign-in-origins "Direct link to Trusted Sign-In Origins") Sign-in accepts requests only from trusted browser origins. The public base URL origin is trusted automatically, along with its loopback twin. The twin is the same origin with the hostname swapped between `localhost` and `127.0.0.1`, keeping the same scheme and port as the base URL. For example, a base URL of `http://localhost:8080` also trusts `http://127.0.0.1:8080` — but not a different port or scheme — so a local install works from either address. You do not need to list the base URL again anywhere. Set additional origins only when the console is reached through more than one hostname, such as a second domain or a reverse proxy. Provide them as a comma-separated list: | Docker Compose setting | Helm value | Purpose | | ----------------------- | ----------------------------------------------- | -------------------------------------------------------------------------------------------------------- | | `AISIX_TRUSTED_ORIGINS` | `ui.extraEnvVars` (add `AISIX_TRUSTED_ORIGINS`) | Extra sign-in origins, comma-separated, such as `https://console.example.com,https://admin.example.com`. | A sign-in attempt from an untrusted origin fails with a message stating that the address is not allowed. Add the origin here to resolve it. For more Docker Compose environment variables and Helm values, see [On-Premises Configuration](https://docs.api7.ai/ai-gateway/reference/on-premises-configuration.md). ## Verify the Installation[​](#verify-the-installation "Direct link to Verify the Installation") Open the dashboard at its configured public base URL. For a new installation, create the first admin account, then sign in. The version of the running control plane appears at the bottom of the dashboard's left navigation, in one of two forms: * `v1.0.0` — a released version. Quote this number when you check the [release notes](https://docs.api7.ai/ai-gateway/release-notes.md) or report an issue. * `dev · 2f9c1ab` — a build from a development commit rather than a release, identified by the abbreviated commit it was built from. The same identity is available without signing in, which is useful in scripts and support bundles. Replace the base URL below with your own public base URL if the control plane is not reached at the default local address: ``` AISIX_CP_URL="http://localhost:8080" curl -sS "${AISIX_CP_URL}/api/config/public" | jq '{cp_version, cp_commit}' ``` ``` { "cp_version": "1.0.0", "cp_commit": "a22f131" } ``` The API reports the released version without the `v` prefix the dashboard displays. `cp_version` is empty on a build that was not cut from a release; `cp_commit` identifies the commit in either case. Gateway versions are reported separately, per instance, in the environment's **Data planes** view — the control plane and your gateways are upgraded independently. ## Next Steps[​](#next-steps "Direct link to Next Steps") Create or select the target [organization and environment](https://docs.api7.ai/ai-gateway/cloud/organizations-and-environments.md), then [connect an AISIX gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md) to issue a gateway certificate and attach the runtime that will serve traffic. Use the [On-Premises Configuration Reference](https://docs.api7.ai/ai-gateway/reference/on-premises-configuration.md) to review Docker Compose environment variables and Helm values. --- # External Database An [AISIX Cloud control plane](https://docs.api7.ai/ai-gateway/on-premises/deployment.md) installed with Helm can use the chart's bundled PostgreSQL instance or an external PostgreSQL database managed by your DBA team. The external database connection role needs specific database privileges, but not `SUPERUSER`. Other database engines are not supported. Docker Compose requires bundled PostgreSQL The online and offline Docker Compose packages do not support an external database. Their Compose files set the control-plane services' database URLs to the bundled `postgres` service and do not expose overrides in `.env`. Use Helm when the control plane must connect to an external PostgreSQL database. The control-plane services share one database: * `cp-api`: Control-plane API server. It runs the schema migration on startup. * `dp-manager`: Data-plane manager. It connects to the same database but does not run the control-plane application migration. Its embedded Kine backend manages its own storage schema. On a fresh database, `dp-manager` retries missing application-schema errors for up to about two minutes while `cp-api` completes the migration. * Dashboard: Control-plane user interface. It stores authentication data in the same database. ## PostgreSQL Requirements[​](#postgresql-requirements "Direct link to PostgreSQL Requirements") * PostgreSQL 14 or later. * One empty database owned by the role the control plane connects as. * One login role with `CREATEROLE` and `BYPASSRLS`, but not `SUPERUSER`. * No PostgreSQL extensions. UUIDs use the built-in `gen_random_uuid()`. ## Why the Control Plane Needs These Privileges[​](#why-the-control-plane-needs-these-privileges "Direct link to Why the Control Plane Needs These Privileges") On its first start, `cp-api` prepares the database schema and creates an internal low-privilege role. The connection role needs enough privilege for that bootstrap work and for shared control-plane operations. Each privilege has a specific purpose: * **Database ownership** lets `cp-api` create the `public` and `auth` schemas, tables, indexes, functions, and Row-Level Security policies. Under default PostgreSQL privileges, the database owner can create objects in `public`, so a standard cluster needs no separate schema grant. * **`CREATEROLE`** lets `cp-api` create and configure `cp_api_app`, the internal role used for tenant-scoped queries with Row-Level Security. You do not create this role yourself. * **`BYPASSRLS`** lets shared control-plane operations read rows across organizations when required. These paths include caller-token authentication, billing webhooks, the background budget aggregator, and `dp-manager` operations. Tenant-scoped requests still switch to `cp_api_app`, which has neither `SUPERUSER` nor `BYPASSRLS`. The control-plane connection role therefore needs specific PostgreSQL capabilities, but not superuser privileges. ## Create the Database and Role[​](#create-the-database-and-role "Direct link to Create the Database and Role") Run this once as a database administrator. Use an account that can create roles, create databases, and grant `BYPASSRLS`. The control-plane connection role itself does not need to be a superuser. ``` -- 1) A dedicated login role for the control plane. Not a superuser. CREATE ROLE aisix LOGIN PASSWORD 'change-me-to-a-strong-password' NOSUPERUSER CREATEROLE BYPASSRLS; -- 2) A dedicated database owned by that role. CREATE DATABASE aisix_cloud OWNER aisix; ``` Choose your own role name, password, and database name. Ownership is what lets `aisix` create objects in `public` and add the `auth` schema on first boot. Use a URL-safe password because the password is embedded in a PostgreSQL connection URL. Characters such as `+`, `/`, and `=` can break the DSN unless they are percent-encoded. Generate a URL-safe value instead, for example with `openssl rand -hex 24`. See [On-Premises Installation](https://docs.api7.ai/ai-gateway/on-premises/deployment.md#helm-on-kubernetes). If your DBA has hardened the `public` schema, for example by revoking the default `CREATE` privilege, also grant schema access explicitly while connected to the new database: ``` \c aisix_cloud GRANT USAGE, CREATE ON SCHEMA public TO aisix; ``` Do not pre-create the `cp_api_app` role. The control plane creates and configures it during migration and asserts its attributes. Pre-creating it, especially with different attributes, can make the migration fail its safety checks. ## Connect the Control Plane to the Database[​](#connect-the-control-plane-to-the-database "Direct link to Connect the Control Plane to the Database") With Helm, set the `externalDatabase.*` values (and `postgresql.builtin=false`) as described in [On-Premises Configuration](https://docs.api7.ai/ai-gateway/reference/on-premises-configuration.md#postgresql). The chart builds the connection URLs for the control-plane services from those values. Set `sslmode=require` (or stricter, such as `verify-full`) whenever the database is reached over a network you do not fully control. ## What the Control Plane Creates[​](#what-the-control-plane-creates "Direct link to What the Control Plane Creates") On the first successful start, `cp-api` builds the full schema in the database you provisioned. On later starts, it applies pending migrations and skips one-time migrations that have already completed: * the `public` schema for control-plane tables, such as organizations, environments, models, caller API keys, and budgets, owned by your role; * the `auth` schema for authentication tables, such as users and sessions; * the `cp_api_app` role, with `NOSUPERUSER NOBYPASSRLS NOINHERIT NOLOGIN`, granted only CRUD on control-plane tables and read/write on a few non-secret identity columns; * Row-Level Security policies that scope each tenant table to a single organization. The connection role is automatically made a member of `cp_api_app` so it can switch to that role per request. You do not grant this yourself. ## Database Privilege Checklist[​](#database-privilege-checklist "Direct link to Database Privilege Checklist") Use this checklist when reviewing the database role with your DBA team: | Requirement | Purpose | | ------------------ | ------------------------------------------------------------------------------------------------------------------- | | `LOGIN` | Allows control-plane services to connect as this role. | | Database ownership | Allows `cp-api` to create and alter schemas, tables, indexes, functions, and RLS policies. | | `CREATEROLE` | Allows `cp-api` to create and manage the internal `cp_api_app` role. | | `BYPASSRLS` | Allows cross-organization control-plane paths and `dp-manager` to read the rows they need on the shared connection. | | Not `SUPERUSER` | Keeps the role below PostgreSQL superuser privilege. | ## Verify the Database Setup[​](#verify-the-database-setup "Direct link to Verify the Database Setup") After the control plane starts, connect as an administrator and confirm the bootstrap looks right. The internal role exists and is correctly de-privileged: ``` SELECT rolname, rolsuper, rolbypassrls, rolcanlogin FROM pg_roles WHERE rolname = 'cp_api_app'; -- expect: cp_api_app | f | f | f ``` Both schemas were created: ``` SELECT nspname FROM pg_namespace WHERE nspname IN ('public', 'auth'); -- expect: two rows ``` ## Troubleshoot Startup Errors[​](#troubleshoot-startup-errors "Direct link to Troubleshoot Startup Errors") Use the startup error message to find the missing database capability. **`cp-api database role must be SUPERUSER or have BYPASSRLS ...`** The connection role is missing `BYPASSRLS`. Grant it as a database administrator: ``` ALTER ROLE aisix BYPASSRLS; ``` **`permission denied for schema public`** The role cannot create objects in the `public` schema. Confirm database ownership. If the schema is locked down, grant schema access explicitly: ``` \c aisix_cloud GRANT USAGE, CREATE ON SCHEMA public TO aisix; ``` **`permission denied to create role`** The connection role is missing `CREATEROLE`. Grant it as a database administrator: ``` ALTER ROLE aisix CREATEROLE; ``` **`permission denied to set role "cp_api_app"`** The migration that grants `cp_api_app` membership likely did not finish. Check the `cp-api` startup logs for an earlier migration error. ## Next Steps[​](#next-steps "Direct link to Next Steps") After the database and role are ready, return to [On-Premises Installation](https://docs.api7.ai/ai-gateway/on-premises/deployment.md) and configure the control plane to use the external database. For the full set of Helm values, see [On-Premises Configuration](https://docs.api7.ai/ai-gateway/reference/on-premises-configuration.md#helm-values). --- # Adapter Protocol Families An adapter is the upstream protocol family AISIX uses after a model alias resolves to a provider key. It controls upstream authentication, request encoding, and provider-specific request handling. For route-level provider support, see [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md). ## Choose an Adapter[​](#choose-an-adapter "Direct link to Choose an Adapter") How you select an adapter depends on how you manage the gateway: * For an AISIX Cloud catalog provider, select the provider. The control plane derives the adapter, and the AISIX Cloud Admin API rejects an explicit `adapter` field. * For a bring-your-own endpoint in AISIX Cloud, set `provider` to `byo` and select the adapter explicitly. * For the open-source AISIX gateway, set both `provider` and `adapter` in `resources.yaml`. The provider is an open string, so you can use a catalog ID or a descriptive name for a private endpoint. In every case, the provider identifies the upstream vendor or endpoint, while the adapter identifies the protocol family AISIX uses to communicate with it. Adapter values form a closed set because AISIX can only encode implemented protocols. For example, this AISIX Cloud provider key connects a private OpenAI-compatible endpoint through the OpenAI adapter: ``` { "provider": "byo", "adapter": "openai", "api_base": "https://llm.private.example/v1" } ``` AISIX then uses the OpenAI-compatible protocol for upstream requests to the configured `api_base`. ## Adapter Values[​](#adapter-values "Direct link to Adapter Values") | Upstream API format | `adapter` | Examples | | --------------------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------ | | OpenAI-compatible APIs | `openai` | OpenAI, DeepSeek, Groq, Mistral, Together.ai, Fireworks, Perplexity, vLLM, SGLang, Ollama, private OpenAI-compatible endpoints | | Anthropic Messages | `anthropic` | Anthropic native Messages API | | AWS Bedrock Runtime | `bedrock` | Anthropic Claude on Bedrock, Bedrock Converse publishers | | Google Vertex AI publisher routes | `vertex` | Gemini and supported Vertex AI publisher routes | | Azure OpenAI Service | `azure-openai` | Azure OpenAI deployments with API-key or Entra ID authentication | ## Adapter Behavior[​](#adapter-behavior "Direct link to Adapter Behavior") | `adapter` | Behavior | | -------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `openai` | Uses OpenAI-compatible request and response formats. This adapter covers OpenAI, public OpenAI-compatible vendors, and private OpenAI-compatible endpoints. | | `anthropic` | Uses Anthropic Messages API requests. Upstream authentication uses `x-api-key` and `anthropic-version`. | | `bedrock` | Uses AWS Bedrock Runtime. AISIX signs outbound requests with AWS SigV4. Anthropic Claude models use Bedrock invoke requests, and other supported publishers use Bedrock Converse. | | `vertex` | Uses Google Vertex AI publisher routes. AISIX authenticates with a GCP OAuth2 Bearer token and calls publisher-specific Vertex endpoints. | | `azure-openai` | Uses Azure OpenAI Service deployment routes. AISIX builds Azure URLs from the provider key's resource host and the model's upstream deployment name. | ## Request Handling[​](#request-handling "Direct link to Request Handling") A direct [model](https://docs.api7.ai/ai-gateway/models/model-aliases.md) references a [provider key](https://docs.api7.ai/ai-gateway/models/provider-keys.md). AISIX Cloud stores the reference in `provider_key_id`. For the open-source AISIX gateway, `resources.yaml` uses `provider_key` with the provider key's `display_name`. The provider key supplies the provider value and adapter. AISIX uses those fields with the model's upstream model ID to build the provider request. AISIX selects upstream request handling in two steps: 1. Check whether the provider value has provider-specific request handling. 2. If it does not, use the adapter family. If neither path is available, the request fails before reaching the upstream provider. The model's `display_name` is the caller-facing model alias. The model's `model_name` is the upstream model ID. The response model field echoes the caller-facing alias on chat completions, the Responses API, text completions, embeddings, Anthropic Messages, video job objects, rerank responses that carry a top-level `model` field, and the Realtime `session.created` and `session.updated` events. Passthrough routes are the exception: they relay the provider response unchanged. Adapters describe the upstream protocol family. They do not guarantee that every proxy endpoint supports every provider. ## Catalog and Custom Providers[​](#catalog-and-custom-providers "Direct link to Catalog and Custom Providers") In AISIX Cloud, users select a catalog provider through the dashboard or AISIX Cloud Admin API. The control plane maps that provider to an adapter family, supplies a default base URL when the catalog defines one, and delivers the provider key configuration to AISIX gateways. For the open-source AISIX gateway, set `provider`, `adapter`, `api_base`, and the upstream credential (`api_key`, or its `secret` alias) directly on each provider key in the declarative resources file. Runtime behavior depends on the model alias, provider key, provider value, adapter, and connection settings. --- # Amazon Nova API [Amazon Nova](https://nova.amazon.com/dev/documentation) is Amazon's family of foundation models, available through a direct API as well as Amazon Bedrock. This guide connects the direct API to AISIX so applications can call Nova through the gateway's OpenAI-compatible API. This setup is for the API-key-based endpoint at `api.nova.amazon.com`. It is different from Amazon Bedrock, which uses AWS credentials, a region, and SigV4 request signing. To call Nova through Bedrock, use [Amazon Bedrock](https://docs.api7.ai/ai-gateway/providers/aws-bedrock.md) instead. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An API key issued for the direct Amazon Nova API. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Export the Nova API key: ``` export NOVA_API_KEY="YOUR_NOVA_API_KEY" ``` Create the provider key: ``` PROVIDER_KEY_ID=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "nova-prod", "provider": "nova", "api_key": "'"${NOVA_API_KEY}"'", "api_base": "https://api.nova.amazon.com/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id' ) echo "$PROVIDER_KEY_ID" ``` AISIX accepts `nova` through the community catalog and derives the `openai` adapter with bearer authentication. Do not send `adapter` on this catalog provider key. AISIX appends `/chat/completions` to the configured API base. Keep `/v1` in the value; a bare `https://api.nova.amazon.com` host points at the wrong route. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Create an alias for Nova 2 Lite: ``` MODEL_ID=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "nova-lite-prod", "model_name": "nova-2-lite-v1", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -er '.model.id' ) echo "$MODEL_ID" ``` Use the exact model ID published in the [Amazon Nova developer documentation](https://nova.amazon.com/dev/documentation). The direct Nova API and Bedrock can use different identifiers for access to the same model family, so do not copy a Bedrock model or inference-profile ARN into this field. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create a caller key limited to the Nova alias: ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "nova-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export NOVA_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "nova-prod" provider: "nova" adapter: "openai" api_key: ${NOVA_API_KEY} api_base: "https://api.nova.amazon.com/v1" models: - display_name: "nova-lite-prod" provider: "nova" model_name: "nova-2-lite-v1" provider_key: "nova-prod" api_keys: - display_name: "nova-caller" key_env: CALLER_API_KEY allowed_models: - "nova-lite-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "nova-lite-prod", "messages": [ { "role": "user", "content": "Say hello from Amazon Nova." } ] }' ``` AISIX sends `nova-2-lite-v1` to `https://api.nova.amazon.com/v1/chat/completions` with the Nova API key in the bearer header. ## Choose Between the Nova API and Bedrock[​](#choose-between-the-nova-api-and-bedrock "Direct link to Choose Between the Nova API and Bedrock") | Account path | AISIX provider | Credential and adapter | | --------------- | ---------------- | ------------------------------------------------- | | Direct Nova API | `nova` | Nova API key with the `openai` adapter | | Amazon Bedrock | `amazon-bedrock` | AWS access credentials with the `bedrock` adapter | Keep the two paths in separate provider keys. This makes credential rotation, account attribution, and failover behavior explicit even when both aliases represent Nova models. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") Use `/v1/chat/completions` for the direct Nova API. AISIX can bridge compatible Responses and Anthropic Messages requests through the chat adapter. Other normalized routes work only when the direct Nova API implements the corresponding OpenAI-shaped endpoint and AISIX accepts the `nova` provider on that route. Image generation, video generation, and rerank do not accept the `nova` provider value. See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md). ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Check | | ----------------------------------- | ----------------------------------------------------------------------- | | Upstream authentication error | Confirm the credential is a direct Nova API key, not an AWS access key. | | Upstream `404` | Keep `/v1` in `api_base` and use a direct Nova model ID. | | SigV4 or AWS region is required | You are using a Bedrock endpoint; follow the Bedrock setup instead. | | Provider key creation returns `400` | Use `provider: "nova"` and omit `adapter`. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to the direct Amazon Nova API and verified the model alias. Continue with these guides: * [AWS Bedrock](https://docs.api7.ai/ai-gateway/providers/aws-bedrock.md): configure a Bedrock-hosted Nova model with AWS credentials instead. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between the direct Nova API and Bedrock. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Anthropic [Anthropic](https://docs.anthropic.com/) develops the Claude family of models and provides the native Messages API for accessing them. AISIX lets applications use either the native Messages format or the gateway's OpenAI-compatible API while managing the Anthropic credential, caller access, rate limits, and usage accounting. This guide connects AISIX directly to Anthropic. To reach Claude through AWS instead, use [AWS Bedrock](https://docs.api7.ai/ai-gateway/providers/aws-bedrock.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An Anthropic API key from the [Anthropic Console](https://console.anthropic.com/settings/keys). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Anthropic-backed route. AISIX connects to Anthropic through the native `anthropic` adapter and authenticates upstream requests with Anthropic's `x-api-key` header. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Anthropic credential and allow it into the environment: ``` # Replace with your value export ANTHROPIC_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "anthropic-prod", "provider": "anthropic", "api_key": "'"${ANTHROPIC_API_KEY}"'", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') ``` ❶ `provider` is `anthropic`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the adapter field is only accepted on BYO provider keys. AISIX sends the credential as the `x-api-key` header and adds `anthropic-version: 2023-06-01` on outbound calls. ❷ `api_key` stores the Anthropic API key. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). You can omit `api_base` for Anthropic; it falls back to the Anthropic default endpoint. The response returns the provider key ID, captured above as `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Claude model IDs from the 4.6 generation onward are pinned snapshots, not evergreen pointers. Check the [Anthropic models reference](https://docs.anthropic.com/en/docs/about-claude/models) for the current IDs. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "claude-sonnet-prod", "model_name": "claude-sonnet-5", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Claude model ID, for example `claude-sonnet-5`, `claude-opus-4-8`, or `claude-haiku-4-5`. ❸ `provider_key_id` attaches the alias to the Anthropic provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the create response: ``` export AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "claude-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') ``` The `allowed_models` value references the model by its ID, captured above as `MODEL_ID`. Store the plaintext key securely; it is not retrievable later. Callers authenticate to AISIX with a caller API key in the `Authorization: Bearer` header. AISIX supplies the Anthropic `x-api-key` and `anthropic-version` headers to the upstream; clients do not send them. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export ANTHROPIC_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "anthropic-prod" provider: "anthropic" adapter: "anthropic" api_key: ${ANTHROPIC_API_KEY} api_base: "https://api.anthropic.com" models: - display_name: "claude-sonnet-prod" provider: "anthropic" model_name: "claude-sonnet-5" provider_key: "anthropic-prod" api_keys: - display_name: "claude-caller" key_env: CALLER_API_KEY allowed_models: - "claude-sonnet-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a native Messages request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/messages" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-sonnet-prod", "max_tokens": 64, "messages": [ { "role": "user", "content": "Say hello from Claude." } ] }' ``` The gateway returns an Anthropic Messages response that echoes the caller-facing alias `claude-sonnet-prod`. If the request fails with an upstream authentication error, check the provider key `api_key`. The same alias also works on `/v1/chat/completions` for OpenAI-shaped clients. This translation drops non-text content blocks, so call `/v1/messages` directly for image or document inputs. ## Clean Up[​](#clean-up "Direct link to Clean Up") When you no longer need the example resources, delete the `claude-caller` caller API key, the `claude-sonnet-prod` model, and the `anthropic-prod` provider key from the dashboard, in that order. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Anthropic and verified the model alias. Continue with these guides: * [Anthropic SDK](https://docs.api7.ai/ai-gateway/getting-started/anthropic-sdk.md): call this alias from an Anthropic SDK at `/v1/messages`. * [AWS Bedrock](https://docs.api7.ai/ai-gateway/providers/aws-bedrock.md): configure a Bedrock-hosted Claude model instead. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # AWS Bedrock [Amazon Bedrock](https://docs.aws.amazon.com/bedrock/) is an AWS service for accessing foundation models from Amazon and other providers through a managed API. AISIX gives applications one OpenAI-compatible interface for Bedrock-hosted Claude, Llama, Mistral, Amazon Nova, Cohere, and other models. This configuration is for Bedrock-hosted models that should use AISIX authentication, model allowlists, rate limits, and usage accounting. AISIX signs outbound Bedrock calls with AWS SigV4. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An AWS access key ID and secret access key with `bedrock:InvokeModel` permission for the target model. Prepare an STS session token when using temporary credentials. * For cross-Region inference, `bedrock:InvokeModel` permission on the inference-profile ARN and on the foundation-model ARN in the source Region and every destination Region. See the requirements for [geographic](https://docs.aws.amazon.com/bedrock/latest/userguide/geographic-cross-region-inference.html) and [global](https://docs.aws.amazon.com/bedrock/latest/userguide/global-cross-region-inference.html) profiles. * Access to the target Bedrock model in the selected region, and its model ID or inference profile ID. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create the provider key with the Bedrock catalog provider and a structured `config` credential: ``` PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "bedrock-prod", "provider": "amazon-bedrock", "api_key": "", "api_base": "https://bedrock-runtime.us-west-2.amazonaws.com", "config": { "access_key_id": "YOUR_AWS_ACCESS_KEY_ID", "secret_access_key": "YOUR_AWS_SECRET_ACCESS_KEY", "region": "us-west-2" }, "allowed_environments": ["'"$ENV_ID"'"] }' | jq -r '.provider_key.id') ``` The empty `api_key` is intentional. Create requests require the field even when `config` supplies the structured Bedrock credential. When updating `config`, omit `api_key` instead of sending another empty string. The open-source gateway can derive the standard AWS endpoint, but AISIX Cloud currently requires an explicit `api_base` for Bedrock. Set the region in the hostname and `config` to the same value. The dashboard exposes the access key ID, secret access key, and region fields. If you use temporary STS credentials through the AISIX Cloud Admin API, also add `session_token` to `config`. Create an alias for [Claude Sonnet 5](https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-anthropic-claude-sonnet-5.html). In `us-west-2`, in-Region inference is unavailable for this model, so the example uses its US geographic inference profile: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "claude-bedrock", "model_name": "us.anthropic.claude-sonnet-5", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') ``` Create a caller API key that can access the model: ``` BEDROCK_CALLER_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "bedrock-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the AWS credential and choose the caller API key that applications will send to AISIX: ``` # Replace with your values export BEDROCK_CREDENTIALS='{"access_key_id":"YOUR_AWS_ACCESS_KEY_ID","secret_access_key":"YOUR_AWS_SECRET_ACCESS_KEY","region":"us-west-2"}' export BEDROCK_CALLER_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge these entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources. resources.yaml ``` _format_version: "1" provider_keys: - display_name: bedrock-prod provider: amazon-bedrock adapter: bedrock api_key: ${BEDROCK_CREDENTIALS} models: - display_name: claude-bedrock provider: amazon-bedrock model_name: us.anthropic.claude-sonnet-5 provider_key: bedrock-prod api_keys: - display_name: bedrock-caller key_env: BEDROCK_CALLER_KEY allowed_models: ["claude-bedrock"] ``` ❶ `provider` labels the upstream. ❷ `adapter` selects Bedrock. ❸ `api_key` is a JSON string with `access_key_id`, `secret_access_key`, and `region`. Bedrock's endpoint is region-keyed, for example `bedrock-runtime.us-west-2.amazonaws.com`, so the region is required. Leave `api_base` unset for standard AWS, or set it to a private Bedrock endpoint if you use one. ❹ `model_name` is the Bedrock model ID or full inference profile ID. The example's `us.` profile is available from `us-west-2` and keeps inference within the United States and Canada. ❺ `provider_key` attaches the model to the credential by the provider key `display_name`. The model's `provider` uses the same upstream label as the provider key. To use Meta Llama in the example file, replace the `claude-bedrock` model entry and update `bedrock-caller` to allow the new alias. Keep `bedrock-prod` and the other resources unchanged: resources.yaml (Meta Llama model access) ``` models: - display_name: llama-bedrock provider: amazon-bedrock model_name: us.meta.llama3-3-70b-instruct-v1:0 provider_key: bedrock-prod api_keys: - display_name: bedrock-caller key_env: BEDROCK_CALLER_KEY allowed_models: ["llama-bedrock"] ``` For Amazon Nova, use a Bedrock model or inference profile ID, such as the Nova 2 Lite profile `us.amazon.nova-2-lite-v1:0` from `us-west-2`. As in the Claude and Llama examples, the `us.` prefix selects a US geographic inference profile. The gateway reads the plaintext caller key from `BEDROCK_CALLER_KEY` and stores only a hash. The verification below uses `claude-bedrock`. If you applied the Meta Llama block, send `llama-bedrock` instead and expect that alias in the response. Include `session_token` in the credential JSON when you use temporary STS credentials. Omit it for long-lived static keys. Provider key secrets follow the credential-handling behavior described in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ### Validate and Load the Configuration[​](#validate-and-load-the-configuration "Direct link to Validate and Load the Configuration") If AISIX is installed locally, validate the complete file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${BEDROCK_CALLER_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-bedrock", "messages": [ { "role": "user", "content": "Say hello from Bedrock." } ] }' ``` The gateway returns an OpenAI-compatible response with the caller-facing alias: ``` { "id": "msg_01example", "object": "chat.completion", "model": "claude-bedrock", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Hello from Bedrock!" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 9, "completion_tokens": 5, "total_tokens": 14 } } ``` Check Bedrock invocation metrics, CloudTrail, or provider-side logs for the test request. If AISIX returns an upstream authentication or authorization error, check the AWS credential, region, IAM permissions, and Bedrock model access. ## Prepare for Production[​](#prepare-for-production "Direct link to Prepare for Production") If applications use streaming, add `bedrock:InvokeModelWithResponseStream` to the AWS permissions and confirm streaming behavior with the target model. For non-Claude Bedrock models, send at least one user or assistant message. Bedrock Converse does not accept system-only requests, so AISIX rejects them before calling the provider. Upstream error detail from AWS is redacted in the caller-visible error to avoid leaking AWS identifiers such as ARNs, region, and account ID. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to AWS Bedrock and verified the model alias. Continue with these guides: * [Anthropic](https://docs.api7.ai/ai-gateway/providers/anthropic.md): configure Claude through the Anthropic API instead. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Azure OpenAI [Azure OpenAI Service](https://learn.microsoft.com/en-us/azure/ai-services/openai/) provides Azure-hosted deployments of OpenAI models behind resource-specific endpoints. AISIX lets applications reach those deployments through a single OpenAI-compatible gateway endpoint. This configuration is for Azure OpenAI deployments that should use AISIX authentication, model allowlists, rate limits, and usage accounting. AISIX can authenticate upstream with a resource API key or an Entra ID client credential. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An Azure OpenAI resource with a deployment. * Either a resource API key or an Entra ID app registration with `tenant_id`, `client_id`, and `client_secret` granted access to the resource. * The Azure OpenAI resource host and deployment name. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create the provider key with the Azure catalog provider. The control plane derives the `azure-openai` adapter, so do not send `adapter`: ``` # Replace with your values export AZURE_OPENAI_API_KEY="YOUR_AZURE_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "azure-prod", "provider": "azure", "api_key": "'"$AZURE_OPENAI_API_KEY"'", "api_base": "https://acme-west.openai.azure.com", "allowed_environments": ["'"$ENV_ID"'"] }' | jq -r '.provider_key.id') ``` The dashboard's Azure provider form accepts a resource API key. The AISIX Cloud Admin API also accepts the Entra ID credential JSON described above as the `api_key` value. The API does not accept that credential under `config`. Create the model with the Azure deployment name as `model_name`: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "gpt-4o-azure", "model_name": "gpt4o-prod", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') ``` Create a caller API key that can access the model: ``` AZURE_CALLER_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "azure-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Choose the authentication scheme that matches how your Azure OpenAI resource is managed. Both examples below are complete resources files for a new gateway and use the same model and caller-key configuration. For an existing gateway, add the entries from one example to the matching collections in its current [resources file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely). Create a missing collection once and preserve unrelated resources. Choose the caller API key that applications will send to AISIX before configuring either authentication scheme: ``` # Replace with your value export AZURE_CALLER_KEY="YOUR_CALLER_API_KEY" ``` In both authentication alternatives, `provider` labels the upstream, `adapter` selects Azure OpenAI, and `api_base` points to the Azure OpenAI resource host. AISIX also accepts the bare resource name, such as `acme-west`. The model's `provider_key` must match the selected provider key's `display_name`, and `model_name` is the Azure deployment name rather than the underlying model ID. The caller key's `allowed_models` value must match the model alias. The gateway reads the plaintext caller key from `AZURE_CALLER_KEY` and stores only a hash. Provider key secrets follow the credential-handling behavior described in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ### Use Resource API Key Authentication[​](#use-resource-api-key-authentication "Direct link to Use Resource API Key Authentication") Use this option when your Azure OpenAI resource is managed with a resource API key. Export the key so the loader can interpolate it: ``` # Replace with your value export AZURE_OPENAI_API_KEY="YOUR_AZURE_API_KEY" ``` Use the exported provider key and caller key in the complete resources file: resources.yaml ``` _format_version: "1" provider_keys: - display_name: azure-prod provider: azure adapter: azure-openai api_key: ${AZURE_OPENAI_API_KEY} api_base: https://acme-west.openai.azure.com models: - display_name: gpt-4o-azure provider: azure model_name: gpt4o-prod provider_key: azure-prod api_keys: - display_name: azure-caller key_env: AZURE_CALLER_KEY allowed_models: ["gpt-4o-azure"] ``` The provider key interpolates the Azure OpenAI resource API key from the environment. Never place the literal secret in the file. ### Use Entra ID Authentication[​](#use-entra-id-authentication "Direct link to Use Entra ID Authentication") Use this option when your Azure OpenAI resource should be accessed through an Entra ID app registration. Export the client credential as a JSON value: ``` # Replace with your values export AZURE_ENTRA_CREDENTIAL='{"tenant_id":"YOUR_TENANT_ID","client_id":"YOUR_CLIENT_ID","client_secret":"YOUR_CLIENT_SECRET"}' ``` Use the exported client credential and caller key in the complete resources file: resources.yaml ``` _format_version: "1" provider_keys: - display_name: azure-aad-prod provider: azure adapter: azure-openai api_key: ${AZURE_ENTRA_CREDENTIAL} api_base: https://acme-west.openai.azure.com models: - display_name: gpt-4o-azure provider: azure model_name: gpt4o-prod provider_key: azure-aad-prod api_keys: - display_name: azure-caller key_env: AZURE_CALLER_KEY allowed_models: ["gpt-4o-azure"] ``` The JSON credential must include `tenant_id`, `client_id`, and `client_secret`. The `client_secret` value is the secret value, not the secret ID. For national or sovereign clouds, add `authority_host` to the JSON credential. Omit it for public Azure. The value must be a bare HTTP(S) origin, such as `https://login.microsoftonline.us`. AISIX builds the outbound chat-completions URL from the provider key's `api_base` and the model's `model_name`: ``` https://<resource>.openai.azure.com/openai/deployments/<deployment>/chat/completions?api-version=2024-10-21 ``` ### Validate and Load the Resources[​](#validate-and-load-the-resources "Direct link to Validate and Load the Resources") If AISIX is installed locally, validate the complete file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy. The request is identical regardless of which upstream authentication scheme the provider key uses. ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AZURE_CALLER_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-4o-azure", "messages": [ { "role": "user", "content": "Say hello from Azure OpenAI." } ] }' ``` The gateway returns an OpenAI-compatible response with the caller-facing alias: ``` { "id": "cmpl_azure_example", "object": "chat.completion", "model": "gpt-4o-azure", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Hello from Azure OpenAI!" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 7, "completion_tokens": 5, "total_tokens": 12 } } ``` For a resource API-key provider key, AISIX sends the Azure `api-key` header. For an Entra ID provider key, AISIX sends `Authorization: Bearer <token>`. Check Azure OpenAI metrics, logs, or quota usage for the test request. If AISIX returns an upstream authentication error, check the resource API key or Entra ID credential. If it returns an upstream route error, check `api_base`, the deployment name in `model_name`, and the Azure API version supported by your deployment. ## Prepare for Production[​](#prepare-for-production "Direct link to Prepare for Production") AISIX currently sends Azure OpenAI requests with `api-version=2024-10-21`. Confirm that this API version is supported by your Azure OpenAI deployment and track Azure's [API version deprecation schedule](https://learn.microsoft.com/en-us/azure/ai-services/openai/api-version-deprecation). Azure may attach `prompt_filter_results` and `content_filter_results` to successful responses. AISIX accepts these Azure extension fields and returns the standard OpenAI-compatible response to the caller. For a corporate proxy, private endpoint, or test endpoint, set `api_base` to the exact host AISIX should call. AISIX appends the Azure deployment path and rejects query strings, fragments, and embedded user information. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Azure OpenAI and verified the model alias. Continue with these guides: * [OpenAI](https://docs.api7.ai/ai-gateway/providers/openai.md): configure a model through the OpenAI API instead of an Azure deployment. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Baseten [Baseten](https://docs.baseten.co/inference/model-apis/overview) provides hosted inference through a shared model catalog and dedicated model deployments. With AISIX, applications call those models through one OpenAI-compatible API while the gateway manages credentials, model access, rate limits, and usage accounting. Baseten serves models through two different surfaces, and the choice determines which endpoint you configure: * **Model APIs** — a shared, multi-tenant catalog of open-weight models at a single fixed endpoint. This page uses that surface. * **Dedicated deployments** — a per-deployment endpoint for a model you deployed into your own Baseten workspace. See [Route to a Dedicated Deployment](#route-to-a-dedicated-deployment). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Baseten API key from the API keys page of your workspace. See [Baseten API keys](https://docs.baseten.co/organization/api-keys). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Baseten-backed chat-completions route. Because Baseten exposes an OpenAI-compatible API, AISIX connects through the `openai` adapter. Set `api_base` for the Baseten surface you use. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Baseten credential and API root: ``` # Replace with your value export BASETEN_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "baseten-prod", "provider": "baseten", "api_key": "'"${BASETEN_API_KEY}"'", "api_base": "https://inference.baseten.co/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `baseten`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the Baseten API key. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is the Baseten Model APIs root and already includes the `/v1` path. AISIX appends the endpoint path to it, so do not add `/chat/completions`. Baseten is one of the catalog providers for which the AISIX Cloud Admin API can fill this value in: if you omit `api_base`, the create still succeeds and resolves to `https://inference.baseten.co/v1`. Set it explicitly anyway — it is the only way to tell a Model APIs provider key apart from a dedicated-deployment provider key when you later read the resource back. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Baseten Model APIs identifiers are namespaced as `publisher/model-name`, matching the upstream repository that published the weights. They are case-sensitive and the casing is not uniform across publishers: `deepseek-ai/DeepSeek-V4-Pro` and `zai-org/GLM-5.2` use mixed case, while `openai/gpt-oss-120b` is entirely lowercase. Copy the identifier verbatim from the [Baseten Model APIs overview](https://docs.baseten.co/inference/model-apis/overview) rather than retyping it. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "baseten-deepseek-v4-pro", "model_name": "deepseek-ai/DeepSeek-V4-Pro", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. Because Baseten model identifiers contain a slash, a short alias also keeps client configuration readable. ❷ `model_name` is the Baseten model identifier, for example `deepseek-ai/DeepSeek-V4-Pro`, `openai/gpt-oss-120b`, or `zai-org/GLM-5.2`. ❸ `provider_key_id` attaches the alias to the Baseten provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the response — store it securely: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "baseten-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export BASETEN_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "baseten-prod" provider: "baseten" adapter: "openai" api_key: ${BASETEN_API_KEY} api_base: "https://inference.baseten.co/v1" models: - display_name: "baseten-deepseek-v4-pro" provider: "baseten" model_name: "deepseek-ai/DeepSeek-V4-Pro" provider_key: "baseten-prod" api_keys: - display_name: "baseten-caller" key_env: CALLER_API_KEY allowed_models: - "baseten-deepseek-v4-pro" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "baseten-deepseek-v4-pro", "messages": [ { "role": "user", "content": "Say hello from Baseten." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `baseten-deepseek-v4-pro`, not the upstream `deepseek-ai/DeepSeek-V4-Pro` identifier. If the request fails, check the provider key `api_key`, `api_base`, and the Baseten model identifier in `model_name` — a casing mismatch in `model_name` is the most common cause, because the identifier is case-sensitive upstream. ## Route to a Dedicated Deployment[​](#route-to-a-dedicated-deployment "Direct link to Route to a Dedicated Deployment") A dedicated deployment does not answer on the shared Model APIs host. Baseten gives each deployment its own hostname derived from the model ID, and the OpenAI-compatible routes live under an environment-scoped path segment: ``` https://model-{model_id}.api.baseten.co/environments/production/sync/v1 ``` See [Call your model](https://docs.baseten.co/inference/calling-your-model) for the current URL form and the environment names available in your workspace. Two consequences for gateway configuration: * Create a **separate provider key** for each dedicated deployment, with `api_base` set to that deployment's URL. The AISIX Cloud Admin API only fills in the shared Model APIs root, so `api_base` is effectively required here. * `model_name` on the alias is the model name the deployment itself serves, which is usually the original weights identifier rather than a Baseten Model APIs catalog identifier. Keep the `baseten` provider value on a dedicated-deployment provider key so usage accounting, metrics, and access logs stay grouped with your other Baseten traffic. A dedicated deployment serves a model the pricing catalog does not cover, so set its rate as an organization pricing override rather than switching provider values. See [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md) and [Cost Metadata](https://docs.api7.ai/ai-gateway/models/model-aliases.md#cost-metadata). ## Use Reasoning Models[​](#use-reasoning-models "Direct link to Use Reasoning Models") Several models on Baseten Model APIs emit reasoning traces. AISIX needs no extra configuration for them: * **Reasoning output.** Baseten returns reasoning in `reasoning_content`, which is already the field AISIX treats as canonical for both streaming and non-streaming responses. Unlike providers that stream reasoning under a vendor-specific `delta` path, Baseten needs no [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) override on the provider key. * **Reasoning controls.** AISIX does not model every request parameter explicitly. Top-level fields it does not recognize are forwarded to the upstream verbatim through the `openai` adapter, so the per-model reasoning controls Baseten documents — a top-level `reasoning_effort` for some models, a top-level `chat_template_args` object for others — reach Baseten unchanged. Check [Baseten reasoning](https://docs.baseten.co/inference/model-apis/reasoning) for the control each model accepts, because it varies by model rather than by account. No parameter renames are configured for Baseten, so token-limit fields also pass through as sent. Send the field name the target Baseten model documents. ## Endpoint Support[​](#endpoint-support "Direct link to Endpoint Support") A Baseten-backed alias works on the normalized chat routes and on `/v1/embeddings` when the configured `api_base` serves an embeddings route. Baseten Embeddings Inference deployments expose an OpenAI-compatible `/v1/embeddings` route, so an embedding alias works when its provider key points at that deployment URL. Anthropic-style clients calling `/v1/messages` are served through AISIX translation into the OpenAI request shape, because the alias resolves to the `openai` adapter. AISIX does not dispatch to a Baseten-native Anthropic-shaped route. A Baseten model identifier that begins with `openai/`, such as `openai/gpt-oss-120b`, does not make the alias an OpenAI-provider model. The provider value is `baseten`, so routes that gate on the provider identity rather than the adapter — image generation, video generation, and rerank — reject the alias. See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md#endpoint-compatibility) for the full route matrix. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Baseten and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between a Baseten dedicated deployment and the shared Model APIs catalog. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Bring Your Own Endpoint Private model servers such as [vLLM](https://docs.vllm.ai/), [SGLang](https://docs.sglang.ai/), and [Ollama](https://ollama.com/) can expose an OpenAI-compatible API from infrastructure you control. AISIX can route model traffic to these servers or to a private proxy in front of your own models. Use a BYO endpoint when applications should keep calling AISIX with the OpenAI-compatible API while AISIX forwards traffic to a private or air-gapped model service. The endpoint must accept OpenAI-compatible chat-completions requests. For AISIX Cloud steps tested against each engine's documented API shape, use the dedicated [Ollama](https://docs.api7.ai/ai-gateway/providers/ollama.md) or [vLLM](https://docs.api7.ai/ai-gateway/providers/vllm.md) guide. For another private OpenAI-compatible server such as SGLang, use the generic AISIX Cloud resource shape on this page. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A reachable OpenAI-compatible endpoint and its served model name. The examples use vLLM serving `meta-llama/Llama-3.1-8B-Instruct`. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Use `byo` as the provider identifier, select the `openai` adapter, and set the endpoint explicitly: ``` PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "vllm-private", "provider": "byo", "adapter": "openai", "api_key": "not-used-by-vllm", "api_base": "http://10.0.0.5:8000/v1", "allowed_environments": ["'"$ENV_ID"'"] }' | jq -r '.provider_key.id') ``` The AISIX Cloud Admin API uses `provider: "byo"` so usage data can distinguish custom endpoints from catalog providers. For the open-source AISIX gateway, `provider` in `resources.yaml` can be a descriptive label such as `vllm`. Both paths require a non-empty `api_key`; use a placeholder only when the endpoint ignores authentication. Create the model: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "llama-3-private", "model_name": "meta-llama/Llama-3.1-8B-Instruct", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') ``` Create a caller API key that can access the model: ``` BYO_CALLER_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "byo-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') ``` Configure BYO pricing under [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md). The `cost` block used by the open-source gateway is not accepted in an AISIX Cloud model request. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Many privately operated inference servers do not require an API key. For an unauthenticated endpoint, use a non-empty placeholder in the provider key; AISIX sends it as the bearer token, and your server can ignore it. Use the endpoint root expected by the server, such as `http://host:8000/v1` for vLLM, `http://host:30000/v1` for SGLang, or `http://host:11434/v1` for Ollama. The example below uses `http://10.0.0.5:8000/v1`. Choose the caller API key that applications will send to AISIX: ``` # Replace with your value export BYO_CALLER_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge these entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources. resources.yaml ``` _format_version: "1" provider_keys: - display_name: vllm-private provider: vllm adapter: openai api_key: not-used-by-vllm api_base: http://10.0.0.5:8000/v1 models: - display_name: llama-3-private provider: vllm model_name: meta-llama/Llama-3.1-8B-Instruct provider_key: vllm-private cost: input_per_1k: 0.0 output_per_1k: 0.0 api_keys: - display_name: byo-caller key_env: BYO_CALLER_KEY allowed_models: ["llama-3-private"] ``` * `provider` is any short label that makes sense for your environment. * `adapter` selects the OpenAI-compatible upstream format. * `api_key` is a non-empty placeholder for unauthenticated endpoints. For an authenticated endpoint, reference the real credential from an environment variable, such as `api_key: ${VLLM_API_KEY}`, instead of writing a literal secret into the file. * `api_base` is the endpoint root. Include `/v1` when that is part of the server's route. * In the model entry, `display_name` is the alias callers send in `model`. * `model_name` is the upstream ID your endpoint expects. For vLLM and SGLang, use the served model name. For Ollama, use the local model tag, such as `llama3.1:8b`. * `provider_key` attaches the model alias to the provider key by its `display_name`. * `cost` is optional. It provides the pricing metadata described below. The caller key's `allowed_models` value must match the model alias. The gateway reads the plaintext caller key from `BYO_CALLER_KEY` and stores only a hash. Provider key secrets follow the credential-handling behavior described in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ### Add Pricing Metadata[​](#add-pricing-metadata "Direct link to Add Pricing Metadata") Catalog providers carry pricing from the models.dev catalog. A BYO endpoint is not in that catalog, so set pricing metadata yourself if you need token-cost accounting. Replace the placeholder zero cost values in the model entry with your real per-1K rates: resources.yaml (model cost) ``` cost: input_per_1k: 0.10 output_per_1k: 0.30 ``` Both values are in USD per 1,000 tokens. `input_per_1k` applies to prompt tokens and `output_per_1k` to completion tokens. Both fields are required when the `cost` block is present. The open-source gateway uses this metadata for usage events and `least_cost` routing, but does not enforce budgets from it. AISIX Cloud does not consume the `cost` block from `resources.yaml`; configure BYO pricing separately through [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md). See [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md#cost-metadata) for the resources-file fields. ### Validate and Load the Configuration[​](#validate-and-load-the-configuration "Direct link to Validate and Load the Configuration") If AISIX is installed locally, validate the complete file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. If those variables are already available to a locally installed gateway process, send it a `SIGHUP` to reload the file: ``` kill -HUP "$(pgrep -x aisix)" ``` If you introduced a variable or changed its value, restart the gateway with the updated process environment instead. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a request through the proxy with the caller API key and model alias you created: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${BYO_CALLER_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "llama-3-private", "messages": [ { "role": "user", "content": "Say hello from the private model." } ] }' ``` The response should be an OpenAI-compatible chat-completions response that echoes the caller-facing alias. Check the endpoint access log for a `POST /v1/chat/completions` entry from AISIX. If AISIX returns an upstream route or connection error, check `api_base`, the served model name, and endpoint reachability. ## Support Additional Endpoints[​](#support-additional-endpoints "Direct link to Support Additional Endpoints") The private endpoint must implement each OpenAI-compatible route that applications call through AISIX. AISIX can forward embedding requests when the endpoint provides a compatible embeddings route. Other routes have additional provider requirements; see [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md). For small differences in request or response shape, configure [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) on the provider key. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected a private OpenAI-compatible endpoint to AISIX. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md): use the pricing metadata for AISIX Cloud budget enforcement. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Cerebras [Cerebras Inference](https://inference-docs.cerebras.ai/) hosts open-weight models on Cerebras inference systems. AISIX places those models behind one OpenAI-compatible API and centralizes upstream credentials, caller access, rate limits, and usage accounting. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Cerebras API key from the [Cerebras Cloud console](https://cloud.cerebras.ai/). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Cerebras-backed chat-completions route. Because the Cerebras API is OpenAI-compatible, AISIX connects through the `openai` adapter and uses the Cerebras API root as `api_base`. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Cerebras credential and API root: ``` # Replace with your value export CEREBRAS_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cerebras-prod", "provider": "cerebras", "api_key": "'"${CEREBRAS_API_KEY}"'", "api_base": "https://api.cerebras.ai/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `cerebras`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the Cerebras API key. Cerebras authenticates with HTTP bearer authentication, which is what the `openai` adapter already sends. The value follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is `https://api.cerebras.ai/v1`, the same root that Cerebras documents as the `baseURL` for [OpenAI client libraries](https://inference-docs.cerebras.ai/resources/openai). It already includes the `/v1` path, so AISIX appends the endpoint path, such as `/chat/completions`, directly to it. For the `cerebras` catalog provider the field is optional — the AISIX Cloud Admin API fills in this same value when you omit it — but the examples set it explicitly so the root each key targets stays visible in the configuration. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Cerebras model IDs are bare identifiers with no vendor prefix, even when the weights come from another vendor. The OpenAI open-weight model is `gpt-oss-120b` on Cerebras, and the Google model is `gemma-4-31b`. Check the [Cerebras model catalog](https://inference-docs.cerebras.ai/models/overview) for the current list before you create an alias; the catalog is short and rotates as models are added and retired. Current IDs include `gpt-oss-120b`, `gemma-4-31b`, and `zai-glm-4.7`. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cerebras-gptoss-prod", "model_name": "gpt-oss-120b", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Cerebras model ID, for example `gpt-oss-120b` or `gemma-4-31b`. Do not carry over a prefixed ID such as `openai/gpt-oss-120b` from another provider that hosts the same weights. ❸ `provider_key_id` attaches the alias to the Cerebras provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cerebras-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export CEREBRAS_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "cerebras-prod" provider: "cerebras" adapter: "openai" api_key: ${CEREBRAS_API_KEY} api_base: "https://api.cerebras.ai/v1" models: - display_name: "cerebras-gptoss-prod" provider: "cerebras" model_name: "gpt-oss-120b" provider_key: "cerebras-prod" api_keys: - display_name: "cerebras-caller" key_env: CALLER_API_KEY allowed_models: - "cerebras-gptoss-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "cerebras-gptoss-prod", "messages": [ { "role": "user", "content": "Say hello from Cerebras." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `cerebras-gptoss-prod`. If the request fails, check the provider key `api_key`, `api_base`, and the Cerebras model ID in `model_name`. ## Send Token Limits[​](#send-token-limits "Direct link to Send Token Limits") Cerebras follows current OpenAI parameter naming and documents `max_completion_tokens` for the generated-token cap in [Chat Completions](https://inference-docs.cerebras.ai/api-reference/chat-completions). The Cerebras entry in the AISIX provider catalog therefore configures no parameter renames, and AISIX forwards `max_completion_tokens` to Cerebras exactly as the caller sent it. Some other OpenAI-compatible upstreams still expect the older `max_tokens` name and carry a rename for it, so a request body that works against Cerebras is not automatically portable to every openai-adapter provider. If an existing client sends the older `max_tokens` name and you need it delivered under the current name, add a rename on the provider key: ``` { "request": { "param_renames": { "max_tokens": "max_completion_tokens" } } } ``` The rename applies to every model that references the provider key. See [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides). ## Control Reasoning Effort[​](#control-reasoning-effort "Direct link to Control Reasoning Effort") Cerebras accepts the standard OpenAI `reasoning_effort` parameter at the top level of the chat-completions body. AISIX does not strip unrecognized top-level parameters, so the value reaches the upstream unchanged: ``` { "model": "cerebras-gptoss-prod", "messages": [ { "role": "user", "content": "Plan a three-step migration." } ], "reasoning_effort": "low" } ``` Accepted values depend on the model: | Model | `reasoning_effort` values | Default | | -------------- | ------------------------------- | ----------------- | | `gpt-oss-120b` | `low`, `medium`, `high` | `medium` | | `gemma-4-31b` | `none`, `low`, `medium`, `high` | `none` | | `zai-glm-4.7` | `none` to disable reasoning | Reasoning enabled | Confirm the values for the model you configured in the [Cerebras reasoning documentation](https://inference-docs.cerebras.ai/capabilities/reasoning), because a value one model accepts can be rejected by another. The Cerebras catalog entry sets no reasoning-field override, so AISIX preserves reasoning that the upstream already returns in the canonical `reasoning_content` field for streaming and non-streaming responses. If a Cerebras model streams reasoning at a different `delta` path, set [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) on the provider key. ## Route Latency-Sensitive Traffic[​](#route-latency-sensitive-traffic "Direct link to Route Latency-Sensitive Traffic") A Cerebras alias is a useful target in a routing model that ranks targets by observed latency. Create a routing model with `strategy` set to `least_latency` and list the Cerebras alias alongside a fallback target. AISIX ranks targets by a moving average of recent upstream latency, using time to first token for streaming requests, and probes targets that have no samples yet before ranking them. See [Route by Cost, Latency, or Load](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md#route-by-cost-latency-or-load). ## Endpoint Support[​](#endpoint-support "Direct link to Endpoint Support") Cerebras is an inference-only upstream, so only part of the proxy surface applies to a Cerebras-backed alias. | Route | Behavior with a Cerebras alias | | ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including `stream: true`. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over the chat adapter path. OpenAI-specific Responses fields without a chat equivalent are ignored. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation. Token counting at `/v1/messages/count_tokens` requires an Anthropic-backed model. | | `/v1/embeddings` | Not usable. Cerebras serves language models and does not publish embedding models, so route embeddings to a separate provider. See [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md). | | `/v1/images/generations` | Rejected. The route accepts only models whose provider is `openai`. | | `/v1/rerank` | Rejected. The route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/v1/videos` | Rejected. The route accepts only its own provider allowlist, which does not include `cerebras`. | | `/passthrough/cerebras/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming this prefix with the `https://api.cerebras.ai/v1` root as its `target_url`; grant the route on the caller key's `allowed_routes`. Provider-native paths relay with limited gateway normalization. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Cerebras and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Cerebras and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Cloudflare Workers AI [Cloudflare Workers AI](https://developers.cloudflare.com/workers-ai/) provides serverless inference for models hosted on Cloudflare's network. AISIX gives applications an OpenAI-compatible API for those models while keeping the Cloudflare token, caller access, rate limits, and usage accounting at the gateway. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Cloudflare account ID and a Workers AI API token. In the Cloudflare dashboard, open the Workers AI page and select **Use REST API** to create the token and copy the account ID, as described in [Get started with the REST API](https://developers.cloudflare.com/workers-ai/get-started/rest-api/). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Workers AI-backed chat-completions route. Cloudflare Workers AI is a community catalog provider. AISIX accepts `cloudflare-workers-ai` as the provider value and selects the `openai` adapter with bearer authentication. AISIX does not supply a curated base URL or provider-specific request and response rewrites, so you must configure the account-scoped `api_base`. The dashboard labels this provider as a community entry whose wire format is unverified. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Cloudflare credential and the account-scoped API root: ``` # Replace with your values export CLOUDFLARE_API_TOKEN="YOUR_PROVIDER_API_KEY" export CLOUDFLARE_ACCOUNT_ID="YOUR_CLOUDFLARE_ACCOUNT_ID" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cloudflare-workers-ai-prod", "provider": "cloudflare-workers-ai", "api_key": "'"${CLOUDFLARE_API_TOKEN}"'", "api_base": "https://api.cloudflare.com/client/v4/accounts/'"${CLOUDFLARE_ACCOUNT_ID}"'/ai/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `cloudflare-workers-ai`, the catalog ID. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys, so do not set it here. ❷ `api_key` stores the Workers AI API token. Cloudflare authenticates the REST API with an `Authorization: Bearer` header, which is what the `openai` adapter already sends. The value follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is the account-scoped root `https://api.cloudflare.com/client/v4/accounts/<ACCOUNT_ID>/ai/v1`. Cloudflare documents the full OpenAI-compatible endpoint as `https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1/chat/completions` in [OpenAI compatible API endpoints](https://developers.cloudflare.com/workers-ai/configuration/open-ai-compatibility/). AISIX appends the endpoint path such as `/chat/completions` to `api_base`, so the value must stop at `/ai/v1`. If you paste the full endpoint URL by mistake, AISIX strips the recognized suffix and the trailing slash, but the shorter root is the value to store. caution Always set `api_base` on a Cloudflare Workers AI provider key. For a community catalog provider, AISIX falls back to the base URL published in the public model catalog when you omit `api_base`. For Cloudflare Workers AI, that published value is `https://api.cloudflare.com/client/v4/accounts/${CLOUDFLARE_ACCOUNT_ID}/ai/v1` — a template, not a resolved URL, because every Cloudflare account has its own endpoint. AISIX performs no placeholder substitution, and the fallback value is stored as written rather than re-validated. The create request still succeeds. The stored root then keeps the literal `${CLOUDFLARE_ACCOUNT_ID}` text in place of an account ID, and the first request fails at the upstream. There is no shared default that could work for this provider. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Workers AI model IDs always begin with `@cf/`, followed by the publisher and the model name: `@cf/<publisher>/<model>`. Carry the entire string into `model_name`, including the `@cf/` prefix and any precision or variant suffix such as `-fp8-fast`. Dropping the prefix, or reusing a bare ID that another host publishes for the same weights, produces an upstream model error. The following IDs are current in the Cloudflare model catalog: | Cloudflare model ID | Notes | | ---------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- | | [`@cf/openai/gpt-oss-120b`](https://developers.cloudflare.com/workers-ai/models/gpt-oss-120b/) | Open-weight reasoning model with a selectable reasoning effort. | | [`@cf/meta/llama-3.3-70b-instruct-fp8-fast`](https://developers.cloudflare.com/workers-ai/models/llama-3.3-70b-instruct-fp8-fast/) | Llama 3.3 70B quantized to fp8 for faster inference. | | [`@cf/qwen/qwen3-30b-a3b-fp8`](https://developers.cloudflare.com/workers-ai/models/qwen3-30b-a3b-fp8/) | Qwen3 instruction model for multilingual chat, reasoning, and tool use. | Check the [Workers AI model catalog](https://developers.cloudflare.com/workers-ai/models/) for the current list before you create an alias. Confirm that the model you pick is a text-generation model rather than an embedding, image, or speech model. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cloudflare-gptoss-prod", "model_name": "@cf/openai/gpt-oss-120b", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the full Workers AI model ID, for example `@cf/openai/gpt-oss-120b`. ❸ `provider_key_id` attaches the alias to the Cloudflare Workers AI provider key. To attach cost metadata for budget accounting and usage reports, see [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md#cost-metadata). ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cloudflare-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export CLOUDFLARE_API_TOKEN="YOUR_PROVIDER_API_KEY" export CLOUDFLARE_ACCOUNT_ID="YOUR_CLOUDFLARE_ACCOUNT_ID" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "cloudflare-workers-ai-prod" provider: "cloudflare-workers-ai" adapter: "openai" api_key: ${CLOUDFLARE_API_TOKEN} api_base: "https://api.cloudflare.com/client/v4/accounts/${CLOUDFLARE_ACCOUNT_ID}/ai/v1" models: - display_name: "cloudflare-gptoss-prod" provider: "cloudflare-workers-ai" model_name: "@cf/openai/gpt-oss-120b" provider_key: "cloudflare-workers-ai-prod" api_keys: - display_name: "cloudflare-caller" key_env: CALLER_API_KEY allowed_models: - "cloudflare-gptoss-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "cloudflare-gptoss-prod", "messages": [ { "role": "user", "content": "Say hello from Cloudflare Workers AI." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `cloudflare-gptoss-prod`. If the request fails, work through the three causes in order: 1. An upstream authentication failure points at the `api_key` value or at a token that lacks Workers AI permissions. 2. An upstream route or account error points at `api_base`. Confirm that the account ID segment is a real account ID and that the value ends at `/ai/v1`. 3. An upstream model error points at `model_name`. Confirm that the ID still carries the `@cf/` prefix. ## Supply What the Community Catalog Does Not[​](#supply-what-the-community-catalog-does-not "Direct link to Supply What the Community Catalog Does Not") Providers with an AISIX-curated adapter mapping ship request and response adjustments alongside the adapter. A caller-facing parameter that the upstream spells differently is renamed before it leaves the gateway. Cloudflare Workers AI has no such mapping, so **no parameter rename and no reasoning-field mapping are registered for it**. AISIX sends the caller's chat-completions body to the `/ai/v1` root with the field names the caller used, and reads the reply as standard OpenAI chat-completions JSON. That default is correct as long as the Cloudflare surface matches the OpenAI shape. When it does not, configure the difference explicitly on the provider key rather than reshaping every client: ``` { "request": { "param_renames": { "max_completion_tokens": "max_tokens" } }, "response": { "reasoning_field": "delta.reasoning" } } ``` * `request.param_renames` renames a top-level parameter on the way out. Use it when a Workers AI model rejects the name your clients already send. If a request carries both names, AISIX uses the value from the original caller-facing name. * `response.reasoning_field` maps reasoning from a nonstandard streaming `delta` path onto the canonical `delta.reasoning_content` field. Use it when a model streams its reasoning somewhere other than `reasoning_content`. An override applies to every model that references the provider key, so validate it against a non-production alias first. See [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) for the full field catalog. AISIX does not strip top-level parameters it does not model natively. A parameter that Cloudflare supports and AISIX has no typed field for still reaches the upstream unchanged. A rename is therefore needed only when the parameter name differs, not when the parameter is simply unfamiliar to the gateway. ## Control Reasoning Effort[​](#control-reasoning-effort "Direct link to Control Reasoning Effort") The `@cf/openai/gpt-oss-120b` model exposes a reasoning-effort control that accepts `low`, `medium`, and `high`. Because unrecognized top-level parameters pass through, send it at the top level of the chat-completions body: ``` { "model": "cloudflare-gptoss-prod", "messages": [ { "role": "user", "content": "Plan a three-step migration." } ], "reasoning_effort": "low" } ``` Reasoning support is per model on Workers AI, and several models expose no reasoning control at all. Confirm the parameter on the [model page](https://developers.cloudflare.com/workers-ai/models/) for the specific model before you send it, because a model that does not accept the parameter can reject the request. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") Cloudflare Workers AI serves text generation and text embeddings over its OpenAI-compatible root, so only part of the AISIX proxy surface applies to a Workers AI-backed alias. | Route | Behavior with a Cloudflare Workers AI alias | | -------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including `stream: true`. | | `/v1/embeddings` | Supported. Cloudflare implements OpenAI-compatible embeddings under the same `/ai/v1` root, so create a second alias on the same provider key whose `model_name` is an embedding model such as `@cf/baai/bge-m3`. See [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md). | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over the chat adapter path rather than passed through verbatim. OpenAI-specific Responses fields without a chat equivalent are ignored. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation. Token counting at `/v1/messages/count_tokens` requires an Anthropic-backed model. | | `/v1/images/generations` | Rejected. The route accepts only models whose provider is `openai`. | | `/v1/rerank` | Rejected. The route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/v1/videos` | Returns a not-implemented error. The route dispatches on a fixed provider set that does not include `cloudflare-workers-ai`. | | `/passthrough/cloudflare-workers-ai/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming this prefix with the account-scoped `/ai/v1` root as its `target_url`; grant the route on the caller key's `allowed_routes`. Such a route reaches paths under `/ai/v1`, so Cloudflare's native `/ai/run/@cf/...` endpoint sits outside the root and is not reachable through it. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Cloudflare Workers AI and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Cloudflare Workers AI and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Cohere [Cohere](https://docs.cohere.com/) provides Command models and APIs for generation, embeddings, and reranking. AISIX places these capabilities behind gateway-managed credentials, caller access, rate limits, and usage accounting. Cohere publishes a native API and a [compatibility API](https://docs.cohere.com/docs/compatibility-api) for OpenAI-shaped requests. This guide uses the compatibility API for chat completions and embeddings, then configures the native API separately for reranking. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Cohere API key from the [Cohere dashboard](https://dashboard.cohere.com/api-keys). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Cohere-backed chat-completions route. AISIX connects to Cohere's compatibility API through the `openai` adapter. Chat completions and embeddings need no request translation. Cohere rerank uses the native API and requires the separate provider key configured in [Add a Rerank Model](#add-a-rerank-model). ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Cohere credential and API root: ``` # Replace with your value export COHERE_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cohere-prod", "provider": "cohere", "api_key": "'"${COHERE_API_KEY}"'", "api_base": "https://api.cohere.ai/compatibility/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `cohere`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the Cohere API key. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` points at Cohere's compatibility API. The path carries a `/compatibility` segment before the `/v1` version segment, because the OpenAI-shaped routes are served on a separate prefix from Cohere's native API. AISIX appends `/chat/completions` to this base. For the `cohere` catalog provider the field is optional — the AISIX Cloud Admin API fills in the same value when you omit it — but the examples set it explicitly so the surface each key targets stays visible in the configuration. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Cohere model IDs combine a family name with a release date: `command-a-03-2025` is the March 2025 snapshot of the `command-a` family, and capability variants add a suffix before the date, as in `command-a-reasoning-08-2025` or `command-a-vision-07-2025`. Because the date is part of the ID, a model alias created from a catalog ID pins one snapshot. Check the [Cohere models list](https://docs.cohere.com/docs/models) for current IDs. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cohere-command-a-prod", "model_name": "command-a-03-2025", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Cohere model ID, for example `command-a-03-2025`, `command-a-plus-05-2026`, or `command-a-reasoning-08-2025`. ❸ `provider_key_id` attaches the alias to the Cohere provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the response — store it securely: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cohere-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export COHERE_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "cohere-prod" provider: "cohere" adapter: "openai" api_key: ${COHERE_API_KEY} api_base: "https://api.cohere.ai/compatibility/v1" models: - display_name: "cohere-command-a-prod" provider: "cohere" model_name: "command-a-03-2025" provider_key: "cohere-prod" api_keys: - display_name: "cohere-caller" key_env: CALLER_API_KEY allowed_models: - "cohere-command-a-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "cohere-command-a-prod", "messages": [ { "role": "user", "content": "Say hello from Cohere." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `cohere-command-a-prod`. If the request fails, check the provider key `api_key`, `api_base`, and the Cohere model ID in `model_name`. A 404 from the upstream usually means `api_base` points at Cohere's native API root instead of the compatibility path. ## Control Reasoning Effort[​](#control-reasoning-effort "Direct link to Control Reasoning Effort") Cohere's reasoning-capable Command models, such as `command-a-reasoning-08-2025` and `command-a-plus-05-2026`, expose thinking differently on each API surface. Because AISIX routes to the compatibility API, callers control it with the OpenAI-shaped `reasoning_effort` field: ``` { "reasoning_effort": "none" } ``` Cohere's compatibility API accepts `none` and `high` on this field, which map to disabled and enabled thinking. The native API's token-budget control is not documented on this surface, so a request cannot cap the reasoning token count through the compatibility route. See Cohere's [Reasoning](https://docs.cohere.com/docs/reasoning) documentation for the current field behavior. AISIX preserves `reasoning_content` in responses and normalizes `reasoning` to that canonical field. The Cohere provider key carries no reasoning-field override. If your model streams reasoning under a different `delta` path, set [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) on the provider key. ## Add a Rerank Model[​](#add-a-rerank-model "Direct link to Add a Rerank Model") The `/v1/rerank` route accepts a model whose provider value is `openai`, `cohere`, or `jina`, so a Cohere-backed alias is admitted on this route. Two details decide whether the request reaches Cohere: * Cohere rerank is not part of the compatibility API. It is published on the native API host. * The rerank route appends `/rerank` to the provider key base and inserts a `/v1` segment when the base does not already end in `/v1`. A key that points at `https://api.cohere.ai/compatibility/v1` therefore builds `https://api.cohere.ai/compatibility/v1/rerank`, which is not a rerank endpoint. In AISIX Cloud, create a second Cohere provider key for the native API root: ``` RERANK_PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cohere-rerank", "provider": "cohere", "api_key": "'"${COHERE_API_KEY}"'", "api_base": "https://api.cohere.com", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$RERANK_PROVIDER_KEY_ID" ``` ❶ `api_base` is the native API root without a version segment. The rerank route adds the version itself and resolves the request to Cohere's [v1 rerank endpoint](https://docs.cohere.com/v1/reference/rerank). caution The rerank route inserts the `/v1` segment whenever the base does not already end in `/v1`, so Cohere's [v2 rerank](https://docs.cohere.com/reference/rerank) path is not reachable through `/v1/rerank`. Setting `api_base` to `https://api.cohere.com/v2` builds `https://api.cohere.com/v2/v1/rerank` and fails upstream. Create the rerank model alias and a caller key scoped to it: ``` RERANK_MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cohere-rerank-prod", "model_name": "rerank-v3.5", "provider_key_id": "'"${RERANK_PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') RERANK_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "cohere-rerank-caller", "allowed_models": ["'"${RERANK_MODEL_ID}"'"] }' | jq -r '.plaintext') ``` For the open-source AISIX gateway, add `cohere-rerank` to `provider_keys` and `cohere-rerank-prod` to `models`. Replace the existing `cohere-caller` entry with the updated entry below so it allows both model aliases. Preserve unrelated entries and collections: resources.yaml (rerank resources) ``` provider_keys: - display_name: "cohere-rerank" provider: "cohere" adapter: "openai" api_key: ${COHERE_API_KEY} api_base: "https://api.cohere.com" models: - display_name: "cohere-rerank-prod" provider: "cohere" model_name: "rerank-v3.5" provider_key: "cohere-rerank" api_keys: - display_name: "cohere-caller" key_env: CALLER_API_KEY allowed_models: - "cohere-command-a-prod" - "cohere-rerank-prod" ``` Validate and reload or restart the declarative resources file as described above, then use the existing caller key for the rerank request: ``` export RERANK_API_KEY="$CALLER_API_KEY" ``` Rerank model IDs follow their own naming, separate from the Command families: `rerank-v3.5`, and the two Rerank 4.0 variants `rerank-v4.0-fast` and `rerank-v4.0-pro`. Because this route resolves to Cohere's v1 rerank endpoint, confirm on the [rerank model overview](https://docs.cohere.com/docs/rerank-overview) that the ID you pick is served there before you point a production alias at it. Send a rerank request through the proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/rerank" \ -H "Authorization: Bearer ${RERANK_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "cohere-rerank-prod", "query": "How do I rotate a provider credential?", "documents": [ "Provider keys store the upstream credential.", "Caller API keys authorize model access.", "Rate limits apply per caller key." ], "top_n": 2 }' ``` AISIX rewrites only the `model` field to `rerank-v3.5` and forwards the body unchanged, so Cohere's `top_n` parameter reaches the upstream as written. The response keeps Cohere's rerank shape: a `results` array ordered by `relevance_score`, plus a `meta` object. AISIX reads `meta.billed_units.input_tokens` from that response for usage accounting, so rerank traffic appears in gateway logs alongside chat traffic. In AISIX Cloud, that traffic contributes to budget totals when matching model pricing is configured. ## Endpoint Support for Cohere[​](#endpoint-support-for-cohere "Direct link to Endpoint Support for Cohere") The routes below behave differently depending on which Cohere provider key backs the model alias. | Route | Behavior with a Cohere-backed model | | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported through the compatibility API base. | | `/v1/embeddings` | Supported. The `openai` adapter appends `/embeddings` to the compatibility base, which is Cohere's OpenAI-shaped embeddings route. Use a Cohere [embedding model](https://docs.cohere.com/docs/cohere-embed) ID such as `embed-v4.0` in `model_name`. | | `/v1/responses` | Bridged through the chat adapter. AISIX returns a Responses-shaped result rather than forwarding to a native Responses API, and ignores OpenAI-specific Responses fields without a chat equivalent. See [Responses API](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). | | `/v1/rerank` | Supported with a second provider key on the native API root. See [Add a Rerank Model](#add-a-rerank-model). | | `/v1/images/generations` | Not supported. The route accepts only models whose provider value is `openai`. | | `/v1/videos` | Not supported. Cohere is not in the video route's provider allowlist. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation. Token counting at `/v1/messages/count_tokens` requires an Anthropic-backed model. | | `/passthrough/cohere/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md); grant the route on the caller key's `allowed_routes`. A route binds one fixed `target_url` and provider key, so reaching both the compatibility base and the native root takes two routes under distinct prefixes, one per Cohere provider key. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Cohere, verified the model alias, and added a rerank route. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for these aliases. * [Rerank](https://docs.api7.ai/ai-gateway/endpoints/rerank.md): review the rerank request contract and its provider requirement. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Cohere and a second provider. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Provider Compatibility Provider compatibility depends on both the caller-facing endpoint and the upstream provider configuration behind the model alias. A model can work on a broad chat endpoint and still be rejected on a provider-specific endpoint. AISIX evaluates compatibility in two layers. The [adapter family](https://docs.api7.ai/ai-gateway/providers/adapters.md) determines how chat-style requests are encoded for the upstream provider. The endpoint's rules determine whether the selected proxy route accepts the model's provider or adapter family. ## Endpoint Compatibility[​](#endpoint-compatibility "Direct link to Endpoint Compatibility") Use the caller's API format and provider support requirements to choose a proxy route. | Need | Route | Provider support | | ------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Broad chat compatibility | `/v1/chat/completions` | OpenAI, Anthropic, Bedrock, Vertex AI, Azure OpenAI, and OpenAI-compatible providers through their configured adapter. | | [Text completions](https://docs.api7.ai/ai-gateway/endpoints/text-completions.md) | `/v1/completions` | Models whose provider key uses the `openai` adapter when the configured upstream implements the legacy `/completions` route. Other adapters return `501 not_implemented`. | | Anthropic-style clients | `/v1/messages` | Provider keys that use the `anthropic` adapter or declare `apis.messages` forward natively; other supported upstreams use translation. Text has the broadest support. Image, document, and tool-calling support depends on the selected provider adapter. Signed thinking history remains Anthropic-specific. | | Anthropic token counting | `/v1/messages/count_tokens` | Targets whose provider key uses the `anthropic` adapter or declares `apis.messages`. The upstream must implement the token-counting route. | | Streaming text chat | `/v1/chat/completions` or `/v1/messages` with `stream: true` | Same provider support as the chosen endpoint. A routing model can fail over before AISIX sends response bytes, but it cannot switch targets after the response stream has started. Chat Completions audio output is an exception: AISIX does not currently preserve `delta.audio`. | | Embeddings | `/v1/embeddings` | OpenAI-compatible upstreams, Amazon Titan and Cohere embedding models on Bedrock, and Google-publisher embedding models on Vertex AI. Other provider and model combinations return `501 not_implemented`. | | OpenAI Responses API | `/v1/responses` | Provider keys that declare `apis.responses`, and OpenAI keys without an `apis` declaration, forward natively. Other upstreams use the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) when the provider adapter supports the translated request shape. On the bridged path, OpenAI-specific fields without a chat equivalent are ignored. | | [Chat audio](https://docs.api7.ai/ai-gateway/endpoints/chat-audio.md) | `/v1/chat/completions` | Non-streaming audio input and output through an `openai` or `azure-openai` adapter when the upstream implements the OpenAI chat-audio shape. Every eligible routing target must meet the same requirement. AISIX preserves `input_audio`, `modalities`, `audio`, and `message.audio`; it does not translate them across provider protocols. | | Image generation | `/v1/images/generations` | Models whose configured provider is OpenAI. | | [Image editing](https://docs.api7.ai/ai-gateway/endpoints/image-editing.md) | `/v1/images/edits` | Models whose configured provider is OpenAI. Requests are `multipart/form-data`; the gateway rewrites only the `model` field and forwards the form intact. | | [Video generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md) | `/v1/videos` and its status and content routes | Models whose configured provider is `alibaba` (Wan), `zhipuai` or `zhipu` (CogVideoX), `volcengine` (Ark Seedance), `runwayml` or `runway` (Runway Gen and Runway-hosted models), or `openai` (Sora). Text-to-video only. Other providers return `501 not_implemented`. See [Video Generation Support](#video-generation-support). | | Audio | `/v1/audio/transcriptions`, `/v1/audio/translations`, `/v1/audio/speech` | OpenAI-style upstream audio routes. AISIX forwards the audio format; it does not translate audio across provider families. | | [Files, Batches, and Fine-tuning](https://docs.api7.ai/ai-gateway/endpoints/batch-files-fine-tuning.md) | `/v1/files`, `/v1/batches`, `/v1/fine_tuning/jobs`, and their related routes | Direct models whose provider key uses the `openai` or `azure-openai` adapter. Anthropic, Bedrock, and Vertex AI use different file and job APIs and are not translated through these routes. | | Realtime WebSocket | `/v1/realtime` | Direct models on the `openai` or `azure-openai` adapter. The upstream must implement the OpenAI Realtime WebSocket path and event protocol; AISIX relays frames without cross-protocol translation. | | Rerank | `/v1/rerank` | Cohere and Jina, or an OpenAI-compatible upstream that uses the `openai` provider value and implements `/v1/rerank`. The public OpenAI API does not provide this endpoint. | | Provider-native routes | A configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md), conventionally `/passthrough/<provider>/*rest` | Any upstream a route's `target_url` points at. A request matches the route's configured path prefix or inbound hosts, and the caller key must grant the route name in its `allowed_routes` list. Limited gateway normalization. | ## Video Generation Support[​](#video-generation-support "Direct link to Video Generation Support") The video routes dispatch on the model alias's `provider` value, not on the upstream model name, and AISIX keeps no allowlist of model IDs — it forwards the configured upstream model name through the fixed mapping below. Successful generation still depends on that model accepting the resulting request shape. Direct aliases only: a routing or ensemble alias returns `400`. Delivery describes how `GET /v1/videos/{video_id}/content` returns the finished file: a `302` to the provider's signed URL, or an authenticated fetch that the gateway streams back to the caller. | `provider` value | Video models | `seconds` maps to | `size` maps to | Delivery | | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | -------------- | | `alibaba` | [Wan and HappyHorse text-to-video](https://docs.api7.ai/ai-gateway/providers/qwen.md#generate-videos-with-wan), such as `wan2.7-t2v`, `wan2.2-t2v-plus`, and `happyhorse-1.1-t2v` | `parameters.duration` | `parameters.size` as `WIDTH*HEIGHT`; omit it for Wan 2.7 and HappyHorse, which use `resolution` and `ratio` tiers instead | Redirect | | `zhipuai`, `zhipu` | [CogVideoX](https://docs.api7.ai/ai-gateway/providers/zhipuai.md#generate-videos-with-cogvideox), such as `cogvideox-3` | `duration` | `size`, forwarded verbatim | Redirect | | `volcengine` | [Ark Seedance](https://docs.api7.ai/ai-gateway/providers/volcengine-ark.md#generate-videos-with-seedance), such as `doubao-seedance-2-0-260128` | `duration` | Validated, then dropped; Ark uses `resolution` and `ratio` tiers | Redirect | | `runwayml`, `runway` | [Runway Gen and Runway-hosted models](https://docs.api7.ai/ai-gateway/providers/runwayml.md) on the text-to-video endpoint, such as `gen4.5` and `veo3.1` | `duration` | `ratio` as `WIDTH:HEIGHT` | Redirect | | `openai` | [Sora](https://docs.api7.ai/ai-gateway/providers/openai.md#generate-videos-with-sora): `sora-2`, `sora-2-pro` | `seconds` as a string; the schema accepts `4`, `8`, or `12` | `size`, forwarded verbatim; accepted resolutions vary by model | Gateway stream | OpenAI [deprecated the Videos API and the Sora 2 models](https://developers.openai.com/api/docs/deprecations) on March 24, 2026, with removal from the API on September 24, 2026. Two provider values that serve chat traffic are outside this allowlist even though their vendor publishes a video API: `alibaba-cn` and `zai` reach a different API root than their video-capable counterpart, so a video alias must use `alibaba` or `zhipuai`. The routes model text-to-video generation only. Request fields other than `model`, `prompt`, `seconds`, and `size` are ignored rather than rejected — including `input_reference`, so an image-to-video request generates from the prompt alone. Provider routes that list, delete, remix, edit, or extend a video are not modeled either. Use a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for those and for provider-native generation fields such as a negative prompt or a seed. ## Provider-Specific Support[​](#provider-specific-support "Direct link to Provider-Specific Support") Supported text output on chat and Responses routes can stream. Audio output inside Chat Completions is currently non-streaming; use the Realtime WebSocket for incremental, bidirectional audio. The table below highlights additional endpoint support and important provider boundaries. | Provider setup | Endpoint support and boundaries | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [OpenAI](https://docs.api7.ai/ai-gateway/providers/openai.md) | Supports chat completions, Responses, embeddings, image generation, image editing, audio, and Sora video generation. The public OpenAI API does not provide a rerank endpoint. | | [Anthropic](https://docs.api7.ai/ai-gateway/providers/anthropic.md) | Supports Messages and token counting natively. Chat completions and Responses use translation. Embeddings, image generation, and rerank are not supported. | | [Amazon Bedrock](https://docs.api7.ai/ai-gateway/providers/aws-bedrock.md) | Supports chat completions and Responses. Embeddings are limited to Amazon Titan and Cohere embedding model IDs. Image generation and rerank are not supported. | | [Google Vertex AI](https://docs.api7.ai/ai-gateway/providers/google-vertex-ai.md) | Supports chat completions and Responses. Embeddings are limited to Google-publisher embedding models. Image generation and rerank are not supported. | | [Azure OpenAI](https://docs.api7.ai/ai-gateway/providers/azure-openai.md) | Supports chat completions and Responses. The Azure OpenAI adapter does not implement embeddings, image generation, or rerank. | | [DeepSeek](https://docs.api7.ai/ai-gateway/providers/deepseek.md) | Supports chat completions. AISIX Cloud declares DeepSeek's native Responses and Anthropic-compatible Messages routes in the catalog; the open-source gateway can declare the same routes on the provider key. Image generation and rerank are not supported. | | [Gemini](https://docs.api7.ai/ai-gateway/providers/gemini.md), [Groq](https://docs.api7.ai/ai-gateway/providers/groq.md), and [Mistral](https://docs.api7.ai/ai-gateway/providers/mistral.md) | Support chat completions and bridged Responses requests. Image generation and rerank are not supported. | | [Together AI](https://docs.api7.ai/ai-gateway/providers/together.md) | Supports chat completions, bridged Responses, embeddings, audio transcription, audio translation, and text-to-speech through its OpenAI-compatible base. Image generation, video generation, and rerank are not supported through the `togetherai` provider value; use a passthrough route for those native Together routes. | | [Qwen](https://docs.api7.ai/ai-gateway/providers/qwen.md) | Supports chat completions, bridged Responses, and embeddings through Model Studio's OpenAI-compatible base. Video generation is supported only for the exact `alibaba` provider value and maps to the native Wan text-to-video task API; `alibaba-cn` is outside the video route's allowlist. Image generation and rerank are not supported through normalized routes. | | [OpenRouter](https://docs.api7.ai/ai-gateway/providers/openrouter.md) and [Fireworks AI](https://docs.api7.ai/ai-gateway/providers/fireworks-ai.md) | Support chat completions, bridged Responses, and embeddings when the alias names an embedding model the upstream publishes on its OpenAI-compatible base. Image generation, video generation, and rerank are not supported: the `openrouter` and `fireworks-ai` provider values are outside those routes' allowlists. | | [Cohere](https://docs.api7.ai/ai-gateway/providers/cohere.md) | Supports chat completions, bridged Responses, and embeddings through the Cohere compatibility API base. Rerank is supported, because `cohere` is one of the three provider values the rerank route accepts; it needs a second provider key pointed at the Cohere native API root. Image generation and video generation are not supported. | | [Zhipu AI](https://docs.api7.ai/ai-gateway/providers/zhipuai.md) | Supports chat completions, bridged Responses, translated Messages, embeddings, audio transcription, text-to-speech, the Realtime WebSocket relay, and normalized video generation. Native image generation, rerank, and advanced CogVideoX requests remain available through a passthrough route; normalized image generation, audio translation, and rerank are not supported. | | [Perplexity](https://docs.api7.ai/ai-gateway/providers/perplexity.md) | Supports Sonar chat completions and bridged Responses requests. Perplexity's standard embedding models work through `/v1/embeddings` only with a separate provider key whose API base includes `/v1`; the Sonar chat base does not. Image generation, video generation, and rerank are not supported. | | [Cerebras](https://docs.api7.ai/ai-gateway/providers/cerebras.md) and [Hugging Face](https://docs.api7.ai/ai-gateway/providers/huggingface.md) | Support chat completions and bridged Responses requests. Embeddings are not usable because neither upstream serves an embedding model on its configured base. Image generation, video generation, and rerank are not supported. | | [Moonshot AI](https://docs.api7.ai/ai-gateway/providers/moonshotai.md) | Supports chat completions. Declare `apis.responses` to use Moonshot's native Responses API, which currently supports only `kimi-k3`; otherwise, AISIX bridges Responses requests through chat completions. Native Messages and other provider routes require a passthrough route. Image generation, video generation, and rerank are not supported. | | [Baseten](https://docs.api7.ai/ai-gateway/providers/baseten.md) | Supports chat completions and bridged Responses requests. Embeddings work when the provider key's `api_base` points at a Baseten deployment that serves an OpenAI-compatible `/v1/embeddings` route. Image generation, video generation, and rerank are not supported, including for a Baseten model ID that begins with `openai/` — the provider value is `baseten`, not `openai`. | | [Jina](https://docs.api7.ai/ai-gateway/providers/jina.md) | Supports embeddings through the OpenAI-shaped route on the Jina API root. Rerank is supported natively, because `jina` is one of the three provider values the rerank route accepts, and it shares the same API root and provider key as embeddings. The testing-only `jina-ai/jina-vlm` model supports chat completions on this root; Responses and Messages use AISIX translation through that experimental route. Other Jina aliases do not support these chat-shaped endpoints. Image generation and video generation are not supported: the `jina` provider value is outside those routes' allowlists. | | [RunwayML](https://docs.api7.ai/ai-gateway/providers/runwayml.md) | Supports video generation only, because `runwayml` (and the short `runway` spelling) is in the video route's provider allowlist. Chat completions, Responses, and embeddings fail upstream — Runway publishes no chat or embeddings API. Image generation and rerank are not supported: the `runwayml` provider value is outside those routes' allowlists. Use a passthrough route for Runway APIs the gateway has not modeled, such as image-to-video. | | [Volcengine Ark](https://docs.api7.ai/ai-gateway/providers/volcengine-ark.md) | Supports chat completions, bridged Responses, translated Messages, and embeddings through the Ark OpenAI-compatible base. Video generation is supported because `volcengine` is in the video route's provider allowlist. Ark's native Responses and image-generation routes remain available through a passthrough route; normalized image generation and rerank are not supported. | | [Cloudflare Workers AI](https://docs.api7.ai/ai-gateway/providers/cloudflare-workers-ai.md), [Databricks](https://docs.api7.ai/ai-gateway/providers/databricks.md), [DeepInfra](https://docs.api7.ai/ai-gateway/providers/deepinfra.md), and [NVIDIA NIM](https://docs.api7.ai/ai-gateway/providers/nvidia-nim.md) | Support chat completions, bridged Responses, and embeddings when the alias names an embedding model the upstream serves on the configured OpenAI-compatible root. Image generation, video generation, and rerank are not supported through their provider values; use a passthrough route for provider-native routes. | | [SiliconFlow](https://docs.api7.ai/ai-gateway/providers/siliconflow.md) | Supports chat completions, bridged Responses, translated Messages, embeddings, audio transcription, and text-to-speech through its OpenAI-shaped API root. Audio translation fails upstream. Image generation, video generation, and rerank are not supported through the `siliconflow` provider value; native Messages and other provider routes remain available through a passthrough route. | | [W\&B Inference](https://docs.api7.ai/ai-gateway/providers/wandb-inference.md) | Supports chat completions, bridged Responses, and translated Messages requests. W\&B Serverless Inference publishes no embeddings, image-generation, video-generation, or rerank endpoint. Its native model list remains available through a passthrough route. | | [Amazon Nova API](https://docs.api7.ai/ai-gateway/providers/amazon-nova.md), [Meta Llama API](https://docs.api7.ai/ai-gateway/providers/meta-llama-api.md), [Nebius Token Factory](https://docs.api7.ai/ai-gateway/providers/nebius-token-factory.md), [Novita AI](https://docs.api7.ai/ai-gateway/providers/novita-ai.md), [OVHcloud AI Endpoints](https://docs.api7.ai/ai-gateway/providers/ovhcloud-ai-endpoints.md), and [DigitalOcean Gradient AI](https://docs.api7.ai/ai-gateway/providers/digitalocean-gradient-ai.md) | Support chat completions and bridged Responses. Messages callers can use the chat translation in AISIX. Embeddings require a compatible embedding model and route on the provider's configured API root. Image generation, video generation, and rerank are not supported through these provider values. | | [Snowflake Cortex](https://docs.api7.ai/ai-gateway/providers/snowflake-cortex.md) | Supports chat completions and the Responses and Messages bridges. Snowflake's Claude-only native Messages route remains available through a passthrough route. The normalized AISIX embeddings route is not compatible with Snowflake's native `POST /api/v2/cortex/inference:embed` wire format; use a passthrough route with a separate provider key or another embedding provider. Image generation, video generation, and rerank are not supported through the `snowflake-cortex` provider value. | | [MiniMax](https://docs.api7.ai/ai-gateway/providers/minimax.md) and [ModelScope](https://docs.api7.ai/ai-gateway/providers/modelscope.md) | Support chat completions and bridged Responses requests. Embeddings depend on the configured upstream model and route because AISIX forwards the OpenAI-shaped embeddings body without provider-specific translation. Image generation, video generation, and rerank are not supported through their provider values. | | [xAI](https://docs.api7.ai/ai-gateway/providers/xai.md) | Supports chat completions, bridged Responses, translated Messages, and the OpenAI-compatible Realtime WebSocket route. Embeddings and rerank are not usable through normalized routes. Native xAI Responses, Messages, image, video, speech-to-text, text-to-speech, and model-list APIs remain available through a passthrough route. | | [Other public OpenAI-compatible providers](https://docs.api7.ai/ai-gateway/providers/openai-compatible-vendors.md) | Require an OpenAI-compatible chat-completions route. Embeddings depend on the upstream route. Rerank also requires a provider value accepted by the rerank route. | | [Ollama](https://docs.api7.ai/ai-gateway/providers/ollama.md) | Supports chat completions through a BYO `openai` adapter. AISIX bridges compatible Responses and Messages requests through chat completions. Embeddings depend on the installed model. The `byo` and `ollama` provider values are not accepted by image generation, video generation, or rerank. | | [vLLM](https://docs.api7.ai/ai-gateway/providers/vllm.md) | Supports completions, chat completions, embeddings, audio transcription, and audio translation when the served model has the required task. AISIX bridges normalized Responses and Messages requests through chat completions rather than using the native vLLM routes. Native Responses, Messages, token counting, and rerank remain available through a passthrough route. Image generation, video generation, text-to-speech, and normalized rerank are not supported. | | [Private OpenAI-compatible endpoints](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md) | Must implement each upstream route that applications call. Embeddings depend on the private endpoint, and provider-specific routes can impose additional provider-value restrictions. | ## Endpoint Rules[​](#endpoint-rules "Direct link to Endpoint Rules") The tables above are the primary route reference. The rules below clarify cases where provider identity, adapter family, and endpoint behavior differ. * Chat completions is the broadest normalized route. For non-OpenAI upstreams, the provider-facing request can still use Anthropic, Bedrock, Vertex AI, Azure OpenAI, or another adapter-specific format behind the gateway. OpenAI chat-audio fields remain specific to the `openai` and `azure-openai` adapters and compatible upstream models. * Responses uses provider-specific handling. OpenAI-backed models and provider keys that declare `apis.responses` forward to the upstream Responses API. Other providers use the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over the chat adapter path and return a Responses-shaped result. OpenAI-specific fields without a chat equivalent are ignored on this path. * Image generation is an OpenAI-provider route. An OpenAI-compatible vendor can use the OpenAI adapter for chat completions and still be rejected on this route when its provider value is not `openai`. * Embeddings dispatch through the resolved adapter. The OpenAI adapter forwards the OpenAI request shape, while the Bedrock and Vertex adapters translate requests for their supported embedding model families. Audio remains an OpenAI-style forwarding route and is not translated across provider families. * Rerank uses a route-specific provider allowlist. Accepted provider values are `openai`, `cohere`, and `jina`, but the configured upstream must provide `/v1/rerank`. The `openai` value supports compatible rerank providers; it does not imply that the public OpenAI API provides this endpoint. * Anthropic Messages supports native Anthropic-protocol routes and translated upstreams. AISIX can translate text, images, documents, and tool-calling history into its normalized request, but the selected provider adapter determines which translated content reaches the upstream. Anthropic `thinking` and `redacted_thinking` history blocks are not replayed to another provider. Token counting requires a provider key that uses the `anthropic` adapter or declares `apis.messages`, and the upstream must implement the route. AISIX preserves `reasoning_content` and normalizes `reasoning` to that canonical field. When an OpenAI-compatible provider streams reasoning from a different `delta` path, configure [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) on the provider key. ## Content Translation Boundaries[​](#content-translation-boundaries "Direct link to Content Translation Boundaries") Cross-provider feature support depends on the caller-facing endpoint and translation direction. Do not assume that one result applies to every pair of provider families. For an Anthropic-shaped `/v1/messages` request sent to a supported non-Anthropic upstream, AISIX translates supported content into its normalized request. This includes text, base64 and URL images, documents, tool definitions, tool calls, and tool results. The selected provider adapter may support only a subset of that content. OpenAI-compatible adapters preserve image parts, while the current Bedrock and Vertex AI chat adapters use the extracted text from multimodal user content. Signed Anthropic thinking history is dropped because another provider cannot replay it. For an OpenAI-shaped `/v1/chat/completions` request sent to another provider family, portable text and tool fields have the broadest support. Non-text handling depends on the selected adapter and upstream API. For example, a provider adapter may use only the concatenated text from a multimodal message even though an OpenAI-compatible upstream can accept the original `image_url` parts. When an application depends on provider-specific content, prefer the matching caller-facing endpoint and provider family. Use [Anthropic Messages](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md) for the exact Anthropic-shaped translation behavior and [OpenAI Client with Anthropic Upstream](https://docs.api7.ai/ai-gateway/endpoints/openai-client-to-anthropic.md) for the reverse direction. --- # Databricks [Databricks Model Serving](https://docs.databricks.com/aws/en/machine-learning/model-serving/score-foundation-models) hosts foundation models behind endpoints in your Databricks workspace. AISIX gives applications one OpenAI-compatible API for those endpoints while managing the workspace token, caller access, rate limits, and usage accounting. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Databricks workspace with Model Serving enabled and at least one serving endpoint you can query. * The [workspace instance name](https://docs.databricks.com/aws/en/workspace/workspace-details) for that workspace, which is the host portion of the per-workspace URL you sign in to. It looks like `dbc-a1b2c3d4-e5f6.cloud.databricks.com` on AWS, `adb-<workspace-id>.<number>.azuredatabricks.net` on Azure, and `<workspace-id>.<number>.gcp.databricks.com` on Google Cloud. * A Databricks API token whose identity has [`CAN QUERY` permission](https://docs.databricks.com/aws/en/security/auth/access-control#serving-endpoint-acls) on every serving endpoint that AISIX will access. The examples use a workspace personal access token. Databricks recommends [OAuth machine-to-machine authentication](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m) for production, but AISIX stores the bearer token as a static provider-key credential and does not refresh it. If you use a short-lived OAuth access token, refresh the provider-key credential before the token expires. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Databricks-backed chat-completions route. Databricks is a community catalog provider with an OpenAI-compatible API. AISIX connects through the `openai` adapter and uses the workspace-specific serving root as `api_base`. AISIX does not supply a curated base URL or provider-specific request and response rewrites. The sections below also explain Databricks' token-limit parameter name. Databricks exposes two OpenAI-compatible workspace surfaces. Keep each base URL paired with its own model naming scheme: | Databricks surface | `api_base` | Upstream `model_name` | | -------------------------------------------------- | --------------------------------------------------- | --------------------------------------------------------------------------------------- | | Model Serving endpoint, used throughout this guide | `https://<workspace-instance>/serving-endpoints` | A serving endpoint name, such as `databricks-claude-sonnet-4-5` or a name you assigned. | | Unity AI Gateway model service (Beta) | `https://<workspace-instance>/ai-gateway/mlflow/v1` | A fully qualified model service name, such as `system.ai.claude-sonnet-4-5`. | Databricks recommends its Beta model services for new access to Databricks-hosted foundation models. The Model Serving surface remains supported and also covers provisioned-throughput, external-model, and OpenAI-compatible custom endpoints. See [Query a chat model](https://docs.databricks.com/aws/en/machine-learning/model-serving/query-chat-models) for both request shapes. Do not use a `system.ai.*` model service name with the `/serving-endpoints` base configured below. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Databricks token and the workspace serving root: ``` # Replace with your values export DATABRICKS_HOST="dbc-a1b2c3d4-e5f6.cloud.databricks.com" export DATABRICKS_TOKEN="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "databricks-prod", "provider": "databricks", "api_key": "'"${DATABRICKS_TOKEN}"'", "api_base": "https://'"${DATABRICKS_HOST}"'/serving-endpoints", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `databricks`. The AISIX Cloud Admin API accepts the value because `databricks` is a models.dev catalog ID, and derives the `openai` adapter with bearer authentication from the community catalog rule. The `adapter` field is only accepted on BYO provider keys, so do not send it here. ❷ `api_key` stores the Databricks personal access token. Databricks authenticates its OpenAI-compatible surface with HTTP bearer authentication, which is what the `openai` adapter already sends. The value follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is `https://<workspace-instance>/serving-endpoints`. Databricks documents this root as the OpenAI client's `base_url` in [External models in Model Serving](https://docs.databricks.com/aws/en/machine-learning/foundation-models/external-models), which makes the full chat endpoint `https://<workspace-instance>/serving-endpoints/chat/completions`. AISIX appends the endpoint path, such as `/chat/completions`, to `api_base`, so the configured value must stop at `/serving-endpoints`. If you paste the full endpoint URL instead, AISIX strips the trailing `/chat/completions` and any trailing slash, but the root above is the value to configure. caution `api_base` is required for Databricks. Databricks has no shared public API host: every request goes to your own workspace instance, so there is no value AISIX can supply on your behalf. For a community-catalog provider, the AISIX Cloud Admin API falls back to the base URL published in the models.dev catalog entry when you omit `api_base`. The Databricks entry publishes `https://${DATABRICKS_HOST}/ai-gateway/mlflow/v1`. This is the Unity AI Gateway model-service root, but its workspace host is an unresolved template. AISIX performs no placeholder substitution, so the provider key stores the literal `${DATABRICKS_HOST}` text and upstream requests fail. Setting `api_base` explicitly also ensures that its API surface matches the model naming scheme you selected. This guide does not cover [route-optimized serving endpoints](https://docs.databricks.com/aws/en/machine-learning/model-serving/query-route-optimization). They use a dedicated endpoint URL and endpoint-scoped OAuth credentials rather than the workspace URL and personal access token shown here. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") On Databricks, the value AISIX sends upstream in `model` is the name of a serving endpoint in your workspace, not a vendor model ID. Two workspaces can serve the same weights under different endpoint names, so read the name from your own workspace rather than from a vendor catalog. Endpoint names fall into two groups: | Endpoint type | Naming | Example | | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------ | | Databricks-hosted foundation model, pay-per-token | Databricks pre-provisions the endpoint. The name carries a `databricks-` prefix and spells the model version with hyphens instead of dots. | `databricks-claude-sonnet-4-5` | | OpenAI-compatible custom, provisioned-throughput, or external-model endpoint | You choose the name when you create the endpoint. There is no prefix and no naming rule. | `openai-chat-endpoint` | Current pay-per-token endpoint names include `databricks-claude-sonnet-4-5`, `databricks-gpt-oss-120b`, and `databricks-gemini-2-5-pro`. The set rotates as Databricks adds and retires hosted models, so confirm the name against the [list of Databricks-hosted foundation models](https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/supported-models) or the Serving page in your workspace before creating an alias. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "databricks-sonnet-prod", "model_name": "databricks-claude-sonnet-4-5", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Databricks serving endpoint name. Do not carry over the underlying vendor model ID, such as `claude-sonnet-4-5`, from a page that documents the vendor's own API. An endpoint name that does not exist in the workspace fails at the upstream, not at alias-creation time. ❸ `provider_key_id` attaches the alias to the Databricks provider key. The credential can reach only the serving endpoints its identity has permission to query. One provider key can serve aliases for all permitted endpoints, or you can use separate identities and keys to preserve least-privilege boundaries between endpoint groups. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "databricks-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export DATABRICKS_TOKEN="YOUR_PROVIDER_API_KEY" export DATABRICKS_HOST="YOUR_DATABRICKS_WORKSPACE_HOST" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "databricks-prod" provider: "databricks" adapter: "openai" api_key: ${DATABRICKS_TOKEN} api_base: "https://${DATABRICKS_HOST}/serving-endpoints" models: - display_name: "databricks-sonnet-prod" provider: "databricks" model_name: "databricks-claude-sonnet-4-5" provider_key: "databricks-prod" api_keys: - display_name: "databricks-caller" key_env: CALLER_API_KEY allowed_models: - "databricks-sonnet-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "databricks-sonnet-prod", "messages": [ { "role": "user", "content": "Say hello from Databricks." } ], "max_tokens": 64 }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `databricks-sonnet-prod`. Because the workspace host, the token, and the endpoint name are three separate values, each failure mode has a distinct signature: | Symptom | Likely cause | | ----------------------------------- | -------------------------------------------------------------------------------------------------------- | | Upstream authentication error | The Databricks API token is expired, or it belongs to a different workspace than the host in `api_base`. | | Upstream `403` permission error | The token's identity does not have `CAN QUERY` permission on the serving endpoint. | | Upstream `404` on the endpoint path | `api_base` does not stop at `/serving-endpoints`, so AISIX built a path Databricks does not serve. | | Upstream error naming the endpoint | `model_name` does not match a serving endpoint in the workspace. | ## Set the Token-Limit Parameter[​](#set-the-token-limit-parameter "Direct link to Set the Token-Limit Parameter") The [Databricks foundation model REST API reference](https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/api-reference) documents `max_tokens` as the generated-token cap. Some OpenAI clients and integrations use the newer `max_completion_tokens` name instead. For featured providers whose upstream still expects the older name, the AISIX provider catalog registers the rename and applies it automatically. Databricks is a community-catalog entry, so no rename is registered for it: AISIX forwards whichever name the caller sent, unchanged. A client that sends `max_completion_tokens` therefore reaches Databricks with a parameter the documented API does not name. If your clients send `max_completion_tokens`, configure a request rename on the provider key. In AISIX Cloud, update the existing provider key in place. The `request` block is replaced in full, not merged field by field. The provider key created earlier has no request overrides, so the block below is complete. If your key already has request overrides, first retrieve its details and include every setting you want to keep in the replacement block. Supplying an empty `request` object clears the stored block. An override affects every model alias that uses the provider key, whether you update it through AISIX Cloud or a resources file. If the key serves production traffic, first validate the same override on a separate provider key used only by a non-production alias. Apply the shared-key change during a controlled change window. ``` curl -sS -X PATCH "$AISIX_CP/provider_keys/$PROVIDER_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "request": { "param_renames": { "max_completion_tokens": "max_tokens" } } }' ``` This request omits `response` and the credential fields, so those settings remain unchanged, and models continue to reference the same provider key. For the open-source AISIX gateway, add the `request.param_renames` mapping shown below to the existing `databricks-prod` entry in `provider_keys`. Keep every other field on that entry, preserve every other entry and collection, and do not create a second top-level `provider_keys` key. resources.yaml (updated provider key) ``` provider_keys: - display_name: "databricks-prod" provider: "databricks" adapter: "openai" api_key: ${DATABRICKS_TOKEN} api_base: "https://${DATABRICKS_HOST}/serving-endpoints" request: param_renames: max_completion_tokens: max_tokens ``` Validate and reload or restart the declarative resources file as described above. The rename rewrites the top-level parameter on every request that flows through the provider key, so it applies to every model alias that references it. When a request carries both names, AISIX keeps the value from the caller-facing source name. See [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides). To verify, send a chat-completions request with `max_completion_tokens` set to a small value and confirm that the completion is truncated at that length. ## Send Reasoning Controls[​](#send-reasoning-controls "Direct link to Send Reasoning Controls") Databricks re-serves models from several vendors behind one OpenAI-shaped surface, so reasoning controls depend on the model family behind the endpoint. Databricks documents the current controls in [Query reasoning models](https://docs.databricks.com/aws/en/machine-learning/model-serving/query-reason-models). AISIX forwards unrecognized top-level chat-completions parameters unchanged, but the accepted fields and values still depend on the upstream model. For a GPT OSS endpoint, send `reasoning_effort` at the top level of the body. The example below assumes a second alias, `databricks-gptoss-prod`, created over the `databricks-gpt-oss-120b` serving endpoint: ``` { "model": "databricks-gptoss-prod", "messages": [ { "role": "user", "content": "Plan a three-step migration." } ], "reasoning_effort": "low" } ``` GPT OSS accepts `low`, `medium`, or `high`. Other model families use different values or an Anthropic-style `thinking` object, so confirm the control for the specific endpoint you configured. Databricks returns Claude extended-thinking chat output as an array of typed reasoning and text blocks. The AISIX `openai` adapter expects a string in `message.content` or `delta.content`, so it cannot decode this block-array shape on the normalized `/v1/chat/completions` route. Use a Databricks-native endpoint through a passthrough route when you need Claude reasoning blocks. A passthrough route preserves the upstream response and does not rewrite an AISIX model alias. AISIX detects the chat envelope from `messages` and records token usage when the upstream response includes supported usage fields. ## Reach Databricks-Native Routes[​](#reach-databricks-native-routes "Direct link to Reach Databricks-Native Routes") Databricks exposes native Responses routes that the AISIX `/v1/responses` bridge does not call: | Databricks route | AISIX passthrough path | Scope | | ---------------------------------------- | --------------------------------------------- | ------------------------------------------------------------------------------------------------ | | `POST /serving-endpoints/open-responses` | `POST /passthrough/databricks/open-responses` | Open Responses format across Databricks-hosted open models, Anthropic Claude, and Google Gemini. | | `POST /serving-endpoints/responses` | `POST /passthrough/databricks/responses` | Native OpenAI Responses API for OpenAI models served by Databricks. | The `/passthrough/databricks` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix, with the workspace serving root (`https://<workspace-instance>/serving-endpoints`) as its `target_url` and the Databricks provider key attached; grant the route on the caller key's `allowed_routes`. The following example reaches the multi-provider Open Responses route. Send the Databricks serving endpoint name in `model`, not the AISIX alias: ``` curl -sS -X POST "$AISIX_PROXY/passthrough/databricks/open-responses" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "databricks-claude-sonnet-4-5", "input": [ { "role": "user", "content": "Plan a three-step migration." } ], "max_output_tokens": 256 }' ``` See [Query a model with the Open Responses API](https://docs.databricks.com/aws/en/machine-learning/model-serving/query-open-responses-models) for its supported fields and provider-specific behavior. The passthrough route forwards the request and response without Databricks-specific normalization. AISIX detects the Responses envelope from `input` and records every supported token dimension the response carries, including input, output, cache, and reasoning details. These counts are telemetry only: passthrough does not advance `tpm` or `tpd` counters, resolve a model cost, or add budget spend. If the provider response omits all supported token fields, the recorded token counts remain zero. Databricks also scores an endpoint at `POST /serving-endpoints/{endpoint-name}/invocations`, which is not an OpenAI-shaped route. Use a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) to reach it. The passthrough route forwards to `{target_url}/{rest}` without borrowing a model alias. Because the route's `target_url` ends with `/serving-endpoints`, the wildcard remainder starts at the endpoint name. Do not repeat `serving-endpoints` in the passthrough path: ``` curl -sS -X POST "$AISIX_PROXY/passthrough/databricks/databricks-claude-sonnet-4-5/invocations" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "messages": [ { "role": "user", "content": "Say hello from Databricks." } ], "max_tokens": 64 }' ``` The native invocations body names the endpoint in the path and carries no top-level `model` field, so only caller API key rate limits apply to this call. On an inject-mode passthrough route, model-scoped request-count limits are matched from a `model` field in the request body naming a configured model of the route's provider. Prefer a normalized AISIX endpoint for traffic that must produce token usage or count against a model's token quota. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") A Databricks alias resolves through the `openai` adapter, so route support depends on what your serving endpoint implements and on each route's own provider rules. | Route | Behavior with a Databricks alias | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `/v1/chat/completions` | Supported, including `stream: true`, when Databricks returns OpenAI-compatible string content. Claude extended-thinking responses require a passthrough route because they use typed content-block arrays. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over the chat adapter path because the alias's provider value is `databricks`, not `openai`. It does not call Databricks' native `/serving-endpoints/responses` or `/serving-endpoints/open-responses` route. OpenAI-specific Responses fields without a chat equivalent are ignored. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation. A Claude-family serving endpoint is translated as well, because the alias resolves to the `openai` adapter rather than the Anthropic one. Claude extended-thinking responses have the same content-block limitation as `/v1/chat/completions`. Token counting at `/v1/messages/count_tokens` requires a model whose provider value is `anthropic`, so it is not available here. | | `/v1/embeddings` | Supported when the alias names a Databricks embedding serving endpoint. Databricks documents `client.embeddings.create` against the same `/serving-endpoints` root in [Serve custom LLMs with Custom Model Serving](https://docs.databricks.com/aws/en/machine-learning/model-serving/serve-custom-llms), so one provider key serves both chat and embedding aliases. See [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md). | | `/v1/images/generations` | Rejected. The route accepts only models whose provider is `openai`. | | `/v1/rerank` | Rejected. The route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/v1/videos` | Rejected. The route's provider allowlist does not include `databricks`. | | `/passthrough/databricks/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for Databricks-native routes, including `/responses`, `/open-responses`, and endpoint invocations. The route forwards to its `target_url` and does not rewrite AISIX aliases. Recognized chat, completions, and Responses envelopes record supported usage fields; other operations remain opaque. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Databricks and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between a Databricks endpoint and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # DeepInfra [DeepInfra](https://docs.deepinfra.com/) provides hosted inference for open-weight models from many publishers. Applications call those models by stable AISIX aliases without receiving the upstream API token. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A DeepInfra API token from the [DeepInfra dashboard](https://deepinfra.com/dash/api_keys). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the DeepInfra-backed chat-completions route. DeepInfra is a community catalog provider with an OpenAI-compatible API. AISIX connects through the `openai` adapter and uses the DeepInfra API root as `api_base`. AISIX does not supply a curated base URL or provider-specific request and response rewrites. The AISIX Cloud Admin API returns `community_badge: true` for DeepInfra. The dashboard groups it under **All providers (community)** and identifies the wire compatibility as assumed rather than verified. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the DeepInfra credential and API root: ``` # Replace with your value export DEEPINFRA_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "deepinfra-prod", "provider": "deepinfra", "api_key": "'"${DEEPINFRA_API_KEY}"'", "api_base": "https://api.deepinfra.com/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `deepinfra`. The AISIX Cloud Admin API derives the adapter from the catalog entry; the `adapter` field is only accepted on BYO provider keys, and sending it on a catalog key returns a 400 error. For `deepinfra` the derived adapter is `openai`. ❷ `api_key` stores the DeepInfra API token. DeepInfra authenticates the OpenAI-compatible surface with HTTP bearer authentication, which is what the `openai` adapter already sends. The value follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is **required** for DeepInfra. models.dev publishes no `api` field for the `deepinfra` entry, so there is no cached default for the AISIX Cloud Admin API to fall back on. Omitting `api_base` returns a 400 error stating that models.dev does not publish a default base URL for this provider. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. #### Why the Base URL Stops at `/v1`[​](#why-the-base-url-stops-at-v1 "Direct link to why-the-base-url-stops-at-v1") DeepInfra's OpenAI SDK examples set `base_url` to `https://api.deepinfra.com/v1/openai`. DeepInfra also publishes the same API directly under `/v1`, including [`/v1/chat/completions`](https://docs.deepinfra.com/api-reference/chat-completions/openai-chat-completions), `/v1/embeddings`, `/v1/images/generations`, and `/v1/audio/*`. AISIX appends the endpoint path to `api_base`, so the direct `https://api.deepinfra.com/v1` root is the more useful configuration. One provider key can then serve chat, embeddings, and audio aliases, while a passthrough route can reach the compatible image endpoint and DeepInfra-native routes. Two related behaviors are worth knowing: * If you paste the full direct endpoint URL, AISIX strips a known endpoint suffix such as `/chat/completions` along with any trailing slash before building the upstream URL. Setting `api_base` to `https://api.deepinfra.com/v1/chat/completions` therefore still works, but prefer the root form so the value stays readable. * AISIX does not synthesize a missing path segment for a non-OpenAI vendor. Setting `api_base` to the bare host `https://api.deepinfra.com` produces the upstream URL `https://api.deepinfra.com/chat/completions`, which DeepInfra does not serve. Use the root DeepInfra documents rather than relying on a shorter form to be repaired. Leaving `api_base` empty is not a workaround either: the gateway refuses to fall back to the OpenAI host for a non-OpenAI provider and returns an upstream configuration error instead of leaking the DeepInfra token to another vendor. ### Create a Model[​](#create-a-model "Direct link to Create a Model") DeepInfra model IDs are namespaced by the weights publisher, in the form `<publisher>/<Model-Name>`, matching the Hugging Face repository ID for the same weights. The casing is not uniform, and the publisher segment is the Hub organization rather than the vendor's brand name — for example the GLM models are published under `zai-org`, not `zhipuai`. Copy the ID verbatim from the [DeepInfra model catalog](https://deepinfra.com/models) rather than retyping it. Current IDs include: | Model ID | Notes | | ----------------------------------------- | --------------------------------------------------------------------------------- | | `deepseek-ai/DeepSeek-V3.2` | Reasoning model; returns reasoning in `reasoning_content`. | | `openai/gpt-oss-120b` | Open-weight model; accepts `reasoning_effort` values `low`, `medium`, and `high`. | | `meta-llama/Llama-3.3-70B-Instruct-Turbo` | Non-reasoning instruction model. | Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "deepinfra-deepseek-prod", "model_name": "deepseek-ai/DeepSeek-V3.2", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the full DeepInfra model ID, including the publisher segment. A bare identifier such as `DeepSeek-V3.2` is rejected by the upstream. ❸ `provider_key_id` attaches the alias to the DeepInfra provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "deepinfra-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export DEEPINFRA_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "deepinfra-prod" provider: "deepinfra" adapter: "openai" api_key: ${DEEPINFRA_API_KEY} api_base: "https://api.deepinfra.com/v1" models: - display_name: "deepinfra-deepseek-prod" provider: "deepinfra" model_name: "deepseek-ai/DeepSeek-V3.2" provider_key: "deepinfra-prod" api_keys: - display_name: "deepinfra-caller" key_env: CALLER_API_KEY allowed_models: - "deepinfra-deepseek-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "deepinfra-deepseek-prod", "messages": [ { "role": "user", "content": "Say hello from DeepInfra." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `deepinfra-deepseek-prod`. If the request fails, isolate the cause by symptom: | Symptom | Likely cause | | ------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | | Upstream authentication error | The DeepInfra token in `api_key` is wrong or revoked. | | Upstream 404 | `api_base` does not point to a DeepInfra root that serves chat completions, or `model_name` omits the publisher prefix. | | Upstream configuration error before any request is sent | `api_base` is empty on the provider key. | ## Set the Generated-Token Limit[​](#set-the-generated-token-limit "Direct link to Set the Generated-Token Limit") This is the one request-shape difference most likely to affect a DeepInfra alias, and it is a direct consequence of the community-catalog path. DeepInfra documents the generated-token cap as `max_tokens` in [Chat Completions](https://docs.deepinfra.com/chat/overview). Some OpenAI clients and integrations use the newer `max_completion_tokens` name instead. Several curated providers in the AISIX catalog carry a rename that converts the newer name to the legacy one before the request leaves the gateway. `deepinfra` has no AISIX-curated adapter mapping, so **no rename is registered for it** and AISIX forwards whichever name the caller sent, unchanged. The practical result: * A caller that sends `max_tokens` reaches DeepInfra with the name DeepInfra documents, and the cap applies. * A caller that sends `max_completion_tokens` reaches DeepInfra with a name DeepInfra does not document. The upstream behavior is therefore not guaranteed: the request may fail or the cap may not apply. If your clients send the current OpenAI name, register the rename yourself on the provider key: ``` { "request": { "param_renames": { "max_completion_tokens": "max_tokens" } } } ``` The rename applies to every model that references the provider key. If a request carries both names, AISIX uses the value from the original caller-facing name. See [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides). ## Control Reasoning Output[​](#control-reasoning-output "Direct link to Control Reasoning Output") DeepInfra accepts the standard OpenAI `reasoning_effort` parameter at the top level of the chat-completions body, and also accepts a `reasoning` object with `effort` and `enabled` fields. Setting `"enabled": false` is equivalent to `reasoning_effort: "none"`. AISIX does not strip unrecognized top-level parameters, so either form reaches the upstream unchanged: ``` { "model": "deepinfra-deepseek-prod", "messages": [ { "role": "user", "content": "Plan a three-step migration." } ], "reasoning_effort": "low" } ``` DeepInfra documents `none`, `low`, `medium`, and `high` for supported reasoning models. Model availability and behavior can still differ, so confirm the controls for your model in the [DeepInfra reasoning documentation](https://docs.deepinfra.com/chat/reasoning) and its catalog page. Using these parameters against a non-reasoning model has no effect. On the response side, DeepInfra returns the model's thinking in `reasoning_content`, which is already the canonical field AISIX preserves and normalizes to. No response override is needed for the models listed above. Because no response rewrite is registered for `deepinfra`, that behavior is a property of the models you configure rather than a guarantee AISIX enforces. If you add a model that streams reasoning on a different `delta` path, set [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) on the provider key so streaming clients still find it at `delta.reasoning_content`. Verify with a streaming request rather than a non-streaming one, because the override applies to the streaming delta path. ## Use the Anthropic-Shaped or Native DeepInfra Surface[​](#use-the-anthropic-shaped-or-native-deepinfra-surface "Direct link to Use the Anthropic-Shaped or Native DeepInfra Surface") DeepInfra publishes several request surfaces on the same host: | DeepInfra surface | Path | Reachable through the provider key configured in this guide | | ----------------------------- | --------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | | Direct OpenAI-compatible | `/v1/chat/completions`, `/v1/embeddings`, `/v1/images/generations`, and `/v1/audio/...` | Yes. Normalized AISIX routes cover chat, embeddings, and audio; image generation requires a passthrough route. | | OpenAI SDK compatibility root | `/v1/openai/...` | Not used by the normalized routes in this guide. DeepInfra publishes it as an equivalent base for OpenAI SDK clients. | | Anthropic Messages | `/anthropic/v1/messages` | No. | | Native inference | `/v1/inference/{model}` | Yes, through a passthrough route. | The community catalog default assigns the `openai` adapter to every provider it covers, and a catalog provider key cannot override the adapter — `adapter` is accepted only when `provider` is `byo`. To speak DeepInfra's Anthropic-shaped surface natively, create a separate [bring-your-own endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md) provider key with `"adapter": "anthropic"` and `"api_base": "https://api.deepinfra.com/anthropic"`, and treat it as a distinct upstream with its own credential and aliases. For most deployments the OpenAI-compatible surface is sufficient, because AISIX already accepts Anthropic-shaped client requests on `/v1/messages` and translates them onto the `openai` adapter. Reach for the BYO path only when you need DeepInfra's own Anthropic implementation rather than the gateway's translation. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") DeepInfra serves several inference modalities, but normalized AISIX route support depends on both the upstream base URL and the model's `deepinfra` provider value. | Route | Behavior with a DeepInfra alias | | ------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including `stream: true`. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over the chat adapter path. OpenAI-specific Responses fields without a chat equivalent are ignored. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation. Token counting at `/v1/messages/count_tokens` requires an Anthropic-backed model. | | `/v1/embeddings` | Supported. DeepInfra serves embedding models on the same OpenAI-compatible root, so one provider key covers both chat and embedding aliases. Create a separate model alias whose `model_name` is an embedding model ID. See [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md). | | `/v1/audio/transcriptions`, `/v1/audio/translations`, `/v1/audio/speech` | Supported with the provider key configured in this guide when the alias names a compatible DeepInfra audio model. See DeepInfra's [audio API reference](https://docs.deepinfra.com/api-reference/audio/openai-audio-transcriptions) and [Speech and Audio](https://docs.api7.ai/ai-gateway/endpoints/audio.md). | | `/v1/images/generations` | Rejected. The route accepts only models whose provider is `openai`. DeepInfra's compatible image endpoint is reachable at `/passthrough/deepinfra/images/generations` through a passthrough route that injects the provider key created in this guide. | | `/v1/rerank` | Rejected. The route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/v1/videos` | Rejected. The route accepts only its own provider allowlist, which does not include `deepinfra`. | | `/passthrough/deepinfra/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md), with the path arithmetic described below. | A DeepInfra model ID that begins with `openai/`, such as `openai/gpt-oss-120b`, does not make the alias an OpenAI-provider model. The provider value is `deepinfra`, so routes that gate on provider identity rather than adapter — image generation, video generation, and rerank — reject the alias. See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md#endpoint-compatibility) for the full route matrix. The `/passthrough/deepinfra` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with `https://api.deepinfra.com/v1` as its `target_url` and the DeepInfra provider key attached; grant the route on the caller key's `allowed_routes`. On a prefix-matched passthrough route, AISIX appends the wildcard remainder to the route's `target_url`. If both the target and the wildcard path contain the same API-version segment, AISIX removes the duplicate: * `/passthrough/deepinfra/models` resolves to `https://api.deepinfra.com/v1/models`. * `/passthrough/deepinfra/images/generations` resolves to DeepInfra's OpenAI-compatible image endpoint. The request body must use a DeepInfra model ID rather than an AISIX alias when it includes `model`. * `/passthrough/deepinfra/v1/inference/deepseek-ai/DeepSeek-V3.2` resolves to the DeepInfra-native inference endpoint. AISIX removes the duplicated `v1` segment because the route's `target_url` already ends in `/v1`. A passthrough route forwards provider-native request and response bodies without rewriting an AISIX alias. AISIX detects OpenAI-compatible chat, completions, and Responses envelopes from each request and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. The Anthropic-shaped surface still needs the separate BYO provider key described above because it is rooted at `/anthropic`, outside the `/v1` target used here. See [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to DeepInfra and verified the model alias. Continue with these guides: * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): register the parameter rename and any response mapping this community-catalog provider does not ship by default. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between DeepInfra and a second provider that serves the same open-weight model. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # DeepSeek [DeepSeek](https://api-docs.deepseek.com/) provides its language and reasoning models through a hosted API. Applications call them through stable AISIX aliases while the gateway keeps the DeepSeek credential out of client code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A DeepSeek API key from the [DeepSeek platform](https://platform.deepseek.com/api_keys). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the DeepSeek-backed chat-completions route. Because DeepSeek exposes an OpenAI-compatible API, AISIX connects through the `openai` adapter and uses the DeepSeek API root as `api_base`. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the DeepSeek credential and API root, and allow it into the environment: ``` # Replace with your values export DEEPSEEK_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "deepseek-prod", "provider": "deepseek", "api_key": "'"${DEEPSEEK_API_KEY}"'", "api_base": "https://api.deepseek.com", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `deepseek`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the DeepSeek API key. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is optional for this catalog provider, because the AISIX Cloud Admin API fills in this same root when you omit it. Set it explicitly so the upstream root stays visible on the resource. AISIX appends the endpoint path to `api_base`, so use the vendor root `https://api.deepseek.com` without a trailing `/chat/completions`. The value must never reach the gateway empty: for an `openai`-adapter provider other than `openai`, AISIX refuses to fall back to `api.openai.com` and returns an upstream configuration error instead. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "deepseek-v4-flash-prod", "model_name": "deepseek-v4-flash", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the DeepSeek model ID, for example `deepseek-v4-flash` or `deepseek-v4-pro`. ❸ `provider_key_id` attaches the alias to the DeepSeek provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key resource that can access the model alias. The gateway generates the key value; the plaintext is returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "deepseek-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value references the model by its ID, so the key can only access the alias you created. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export DEEPSEEK_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "deepseek-prod" provider: "deepseek" adapter: "openai" api_key: ${DEEPSEEK_API_KEY} api_base: "https://api.deepseek.com" models: - display_name: "deepseek-v4-flash-prod" provider: "deepseek" model_name: "deepseek-v4-flash" provider_key: "deepseek-prod" api_keys: - display_name: "deepseek-caller" key_env: CALLER_API_KEY allowed_models: - "deepseek-v4-flash-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek-v4-flash-prod", "messages": [ { "role": "user", "content": "Say hello from DeepSeek." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `deepseek-v4-flash-prod`. Confirm the request on the DeepSeek platform usage page. If the request fails, check the provider key `api_key`, `api_base`, and the DeepSeek model ID in `model_name`. ## Use Thinking Mode[​](#use-thinking-mode "Direct link to Use Thinking Mode") DeepSeek V4 uses thinking mode by default. Regular requests use `high` effort, while DeepSeek can automatically use `max` for some complex agent requests. To disable thinking for a chat-completions request, add the following field to the request body: ``` { "thinking": { "type": "disabled" } } ``` When thinking is enabled, use the top-level `reasoning_effort` field to request `high` or `max`. For client compatibility, DeepSeek maps `low` and `medium` to `high`, and `xhigh` to `max`. Sampling fields such as `temperature` and `top_p` have no effect in thinking mode. See [Thinking Mode](https://api-docs.deepseek.com/guides/thinking_mode) for details. AISIX forwards these controls to DeepSeek and preserves returned reasoning in the `reasoning_content` field for streaming and non-streaming responses. If the model makes a tool call, DeepSeek requires the assistant message's `reasoning_content` to be included in all subsequent requests. AISIX preserves the field, but the application is responsible for retaining and replaying the assistant message; omitting it causes DeepSeek to return HTTP 400. ## Reach DeepSeek-Native API Formats[​](#reach-deepseek-native-api-formats "Direct link to Reach DeepSeek-Native API Formats") DeepSeek publishes a native [Responses API](https://api-docs.deepseek.com/guides/responses_api) and an [Anthropic-compatible API](https://api-docs.deepseek.com/guides/anthropic_api), and AISIX Cloud reaches both without any configuration: the DeepSeek catalog entry ships the declaration below, so `/v1/responses` and `/v1/messages` on a DeepSeek alias go to DeepSeek's own routes rather than being translated to chat completions. ``` { "apis": { "responses": {}, "messages": { "base": "https://api.deepseek.com/anthropic" } } } ``` `responses` carries no `base` because DeepSeek serves it at the same API root the key already points at. The declaration applies while the key points at that root. If you repoint `api_base` at a private endpoint it drops away, because it describes DeepSeek's paths rather than yours. Set it yourself, on the key, if you front DeepSeek through your own gateway and it serves the same routes. See [Declare the API Surfaces](https://docs.api7.ai/ai-gateway/models/provider-keys.md#declare-the-api-surfaces). The open-source AISIX gateway has no catalog, so declare it on the provider key entry in `resources.yaml`. Without this declaration on the provider key, AISIX falls back as follows: | AISIX setup or route | Upstream behavior | | ------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/responses` | Uses the AISIX [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over DeepSeek chat completions. It does not call DeepSeek's native `/responses` route, and Responses-only fields without a chat equivalent are ignored. | | `/passthrough/deepseek/responses` | Calls DeepSeek's native Responses route. Send the upstream model ID in `model`, not the AISIX alias. Declaring `apis.responses` reaches the same route under an AISIX alias, which is usually what you want instead. | | `/v1/messages` | Translates the Anthropic-shaped caller request to DeepSeek chat completions. It does not call DeepSeek's `/anthropic/v1/messages` route. | | Separate BYO provider key with `adapter: anthropic` and `api_base: https://api.deepseek.com/anthropic` | Uses the Anthropic adapter to call DeepSeek's Anthropic-compatible Messages route, at the cost of a second key and a second model alias. Declaring `apis.messages` on the catalog key reaches the same route under one alias. | | `/v1/completions` | Does not reach DeepSeek's FIM Completion (Beta) route, which is served only under the `/beta` API root. | | `/passthrough/deepseek/beta/completions` | Calls [FIM Completion (Beta)](https://api-docs.deepseek.com/api/create-completion). Send `deepseek-v4-pro` in `model`; AISIX does not rewrite an alias on passthrough routes. | The `/passthrough/deepseek` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with DeepSeek's API root as its `target_url`; grant the route name on the caller key's `allowed_routes`. The route relays the native Responses request and response without rewriting an AISIX alias. AISIX detects the Responses envelope from `input` and records every supported token dimension the response carries, including input, output, cache, and reasoning details. These counts are telemetry only: passthrough does not advance `tpm` or `tpd` counters, resolve a model cost, or add budget spend. If the response omits all supported token fields, the recorded token counts remain zero. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") | Route | Behavior with a DeepSeek catalog alias | | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `/v1/chat/completions` | Supported, including `stream: true`. | | `/v1/responses` | Reaches DeepSeek's native Responses API, which the catalog entry declares. Reasoning output items survive, where the chat-completions translation cannot carry them. | | `/v1/messages` | Reaches DeepSeek's Anthropic-compatible route, which the catalog entry declares, so prompt-cache breakpoints and thinking blocks survive. `/v1/messages/count_tokens` is available on the same route. | | `/v1/completions` | Does not reach DeepSeek FIM. Use `/passthrough/deepseek/beta/completions` with `deepseek-v4-pro` instead. | | `/v1/embeddings` | Not available. DeepSeek does not publish an embeddings endpoint on this API root. | | `/v1/audio/*` | Not available. DeepSeek does not publish matching OpenAI-style audio routes. | | `/v1/images/generations` | Rejected. The route accepts only models whose provider is `openai`. | | `/v1/rerank` | Rejected. The route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/v1/videos` | Rejected. The route's provider allowlist does not include `deepseek`. | | `/passthrough/deepseek/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for provider-native routes such as `/responses` and `/models`, with the alias and usage limitations described above. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to DeepSeek and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between DeepSeek and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # DigitalOcean Gradient AI [DigitalOcean Gradient AI Serverless Inference](https://docs.digitalocean.com/products/inference/how-to/si-endpoints/) provides hosted access to foundation models without requiring you to operate model-serving infrastructure. AISIX maps the upstream model IDs to stable aliases and controls who can use them. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Gradient AI model access key or a DigitalOcean personal access token that can call Serverless Inference. * Access to the selected [DigitalOcean Inference model](https://docs.digitalocean.com/products/inference/details/models/). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` DigitalOcean is a community catalog provider with an OpenAI-compatible inference API. AISIX connects through the `openai` adapter and authenticates upstream requests with a bearer token. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") ``` export DIGITALOCEAN_INFERENCE_KEY="YOUR_MODEL_ACCESS_KEY" PROVIDER_KEY_ID=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "digitalocean-prod", "provider": "digitalocean", "api_key": "'"${DIGITALOCEAN_INFERENCE_KEY}"'", "api_base": "https://inference.do-ai.run/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id' ) echo "$PROVIDER_KEY_ID" ``` Use `digitalocean` as the catalog provider ID. AISIX derives the `openai` adapter; sending an explicit adapter on a catalog key returns a validation error. The API base includes `/v1`. AISIX appends `/chat/completions` for chat requests. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Create an alias for a model currently available from Gradient AI: ``` MODEL_ID=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "digitalocean-gpt-oss-prod", "model_name": "openai-gpt-oss-120b", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -er '.model.id' ) echo "$MODEL_ID" ``` Use the DigitalOcean model ID, not the original publisher repository name. Confirm current availability in the model reference before creating another alias. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "digitalocean-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export DIGITALOCEAN_INFERENCE_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "digitalocean-prod" provider: "digitalocean" adapter: "openai" api_key: ${DIGITALOCEAN_INFERENCE_KEY} api_base: "https://inference.do-ai.run/v1" models: - display_name: "digitalocean-gpt-oss-prod" provider: "digitalocean" model_name: "openai-gpt-oss-120b" provider_key: "digitalocean-prod" api_keys: - display_name: "digitalocean-caller" key_env: CALLER_API_KEY allowed_models: - "digitalocean-gpt-oss-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "digitalocean-gpt-oss-prod", "messages": [ { "role": "user", "content": "Say hello from DigitalOcean Gradient AI." } ] }' ``` AISIX forwards `openai-gpt-oss-120b` to `https://inference.do-ai.run/v1/chat/completions` with the DigitalOcean credential. ## Reach DigitalOcean-Native API Formats[​](#reach-digitalocean-native-api-formats "Direct link to Reach DigitalOcean-Native API Formats") DigitalOcean publishes native [Responses and Messages endpoints](https://docs.digitalocean.com/products/inference/how-to/si-endpoints/) on the same API base. The catalog provider key in this guide still uses the `openai` adapter, so normalized AISIX routes do not automatically select those native upstream formats: | AISIX route | Upstream behavior | | ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/responses` with a DigitalOcean alias | Uses the AISIX [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over DigitalOcean chat completions. It does not call DigitalOcean's native `/v1/responses` route, and Responses-only fields without a chat equivalent are ignored. | | `/passthrough/digitalocean/responses` | Calls DigitalOcean's native Responses route. Send the DigitalOcean model ID in `model`, not the AISIX alias. | | `/v1/messages` with a DigitalOcean alias | Translates the Anthropic-shaped caller request to DigitalOcean chat completions. It does not call DigitalOcean's native `/v1/messages` route. | | `/passthrough/digitalocean/messages` | Calls DigitalOcean's native Messages route. Send a DigitalOcean model ID supported by that endpoint. | The `/passthrough/digitalocean` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with `https://inference.do-ai.run/v1` as its `target_url`; grant the route name on the caller key's `allowed_routes`. The paths omit `/v1` because the route's target already includes it; a duplicated leading `v1` segment is stripped either way. A passthrough route does not rewrite an AISIX alias. AISIX detects chat, completions, and Responses envelopes from each request and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") | Route | Behavior with a DigitalOcean catalog alias | | ------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including `stream: true`. | | `/v1/responses` | Supported through the chat-based Responses bridge, not DigitalOcean's native Responses API. | | `/v1/messages` | Supported through translation to chat completions. `/v1/messages/count_tokens` is not available because it requires an Anthropic-backed model. | | `/v1/embeddings` | Supported when the alias targets a DigitalOcean embedding model, such as `qwen3-embedding-0.6b`. | | `/v1/audio/speech` | The request reaches DigitalOcean when the alias targets a text-to-speech model such as `qwen3-tts-voicedesign`, but the response is not OpenAI-compatible end to end. DigitalOcean wraps base64-encoded audio in a JSON data envelope, and AISIX relays that body without decoding it into binary audio. Use `/passthrough/digitalocean/audio/speech` and decode `data[0].b64_json` when the application supports DigitalOcean's native response contract. | | `/v1/audio/transcriptions` and `/v1/audio/translations` | Not available. DigitalOcean does not publish these routes. | | `/v1/images/generations` | Rejected because the normalized AISIX route accepts only the `openai` provider value. Call the native route through `/passthrough/digitalocean/images/generations` with a compatible DigitalOcean model ID. | | `/v1/videos` | Not implemented for the `digitalocean` provider. Submit a DigitalOcean-native video job through `/passthrough/digitalocean/videos`, poll it through `/passthrough/digitalocean/video/generations/{job_id}`, and optionally download the completed MP4 through `/passthrough/digitalocean/videos/{video_id}/content`. | | `/v1/rerank` | Rejected because the normalized AISIX route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/passthrough/digitalocean/async-invoke` | Calls DigitalOcean's asynchronous endpoint for supported image, audio, and text-to-speech models. Use the provider-native request format. | | `/passthrough/digitalocean/models` | Lists the model IDs accessible to the configured DigitalOcean credential. | See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the gateway-wide endpoint matrix. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Check | | ----------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Upstream `401` or `403` | Confirm the credential has inference access. For a model access key, verify its model scope and any VPC restriction; VPC-restricted callers must use the VPC-local DNS resolver. | | Upstream `404` | Keep `/v1` in `api_base` and verify the DigitalOcean model ID. | | Model not found | Query DigitalOcean `GET /v1/models` with the same upstream credential and confirm that the model access key includes the selected model. | | Provider key creation returns `400` | Use `provider: "digitalocean"` without `adapter`. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to DigitalOcean Gradient AI and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md): configure gateway-side request and token limits. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between DigitalOcean and another provider. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Fireworks AI [Fireworks AI](https://docs.fireworks.ai/) provides hosted inference for a catalog of generative AI models. Applications call selected models by stable AISIX aliases while the gateway holds the Fireworks credential. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Fireworks API key from the [Fireworks console](https://fireworks.ai/api-keys). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Fireworks-backed chat-completions route. Because Fireworks AI exposes an OpenAI-compatible API, AISIX connects through the `openai` adapter and uses the Fireworks API root as `api_base`. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Fireworks credential and API root, and allow it into the environment: ``` # Replace with your value export FIREWORKS_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "fireworks-prod", "provider": "fireworks-ai", "api_key": "'"${FIREWORKS_API_KEY}"'", "api_base": "https://api.fireworks.ai/inference/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `fireworks-ai`, the catalog provider ID. The hyphen is part of the identifier; `fireworks` alone is rejected with a `400 INVALID_REQUEST`. ❷ `api_key` stores the Fireworks API key. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` points at the Fireworks inference API. The `/inference` segment is significant: Fireworks serves its OpenAI-compatible inference routes under `https://api.fireworks.ai/inference/v1`, while its account-management REST API — the one that lists and creates Fireworks API keys — is served under `https://api.fireworks.ai/v1`. Dropping `/inference` points the provider key at the wrong API surface. AISIX appends the endpoint path to `api_base`, so use the API root without a trailing `/chat/completions`. A trailing slash is trimmed before the endpoint path is appended, so `https://api.fireworks.ai/inference/v1/` and `https://api.fireworks.ai/inference/v1` resolve identically. `api_base` is optional for this provider. When it is omitted, the AISIX Cloud Admin API fills in `https://api.fireworks.ai/inference/v1`. Setting it explicitly keeps the upstream root visible on the resource. A Fireworks dedicated deployment uses the same inference root; point `model_name` at `accounts/<ACCOUNT_ID>/deployments/<DEPLOYMENT_ID>` to change the target. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Fireworks text-model resource names are fully qualified account paths rather than bare names. A serverless text model published by Fireworks uses the form `accounts/fireworks/models/<name>`, for example `accounts/fireworks/models/gpt-oss-120b`. A flat OpenAI-style name such as `gpt-oss-120b` is not valid for this chat model. Copy the full identifier from the [Fireworks models overview](https://docs.fireworks.ai/models/overview) rather than assembling it by hand. Some catalog entries replace the `models` segment with `routers`, and models you deploy yourself use your own account name in the first segment. Other Fireworks inference APIs can publish different model-ID forms. For example, the embeddings guide uses `fireworks/qwen3-embedding-8b`. Use the exact identifier documented for the endpoint and model instead of forcing every identifier into the `accounts/...` form. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "fireworks-gptoss-prod", "model_name": "accounts/fireworks/models/gpt-oss-120b", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. Callers never see the account path. ❷ `model_name` is the Fireworks model ID, for example `accounts/fireworks/models/gpt-oss-120b`. ❸ `provider_key_id` attaches the alias to the Fireworks provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The API generates the key value and returns the plaintext once: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "fireworks-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value references the model by its ID. The plaintext key is returned only in this response, so store it securely. The new resources project to the attached gateway automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export FIREWORKS_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "fireworks-prod" provider: "fireworks-ai" adapter: "openai" api_key: ${FIREWORKS_API_KEY} api_base: "https://api.fireworks.ai/inference/v1" request: param_renames: max_completion_tokens: max_tokens models: - display_name: "fireworks-gptoss-prod" provider: "fireworks-ai" model_name: "accounts/fireworks/models/gpt-oss-120b" provider_key: "fireworks-prod" api_keys: - display_name: "fireworks-caller" key_env: CALLER_API_KEY allowed_models: - "fireworks-gptoss-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "fireworks-gptoss-prod", "messages": [ { "role": "user", "content": "Say hello from Fireworks AI." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `fireworks-gptoss-prod`. If the request fails, check the provider key `api_key`, the `/inference/v1` path in `api_base`, and the full `accounts/...` model ID in `model_name`. ## Understand the Token-Limit Parameter Rewrite[​](#understand-the-token-limit-parameter-rewrite "Direct link to Understand the Token-Limit Parameter Rewrite") Fireworks documents `max_tokens` as the output-limit parameter on its chat-completions route, while current OpenAI SDKs and many agent frameworks send `max_completion_tokens`. AISIX Cloud resolves the difference through a built-in request override on its `fireworks-ai` catalog entry. The open-source AISIX gateway does not add that catalog override itself, so the `resources.yaml` example configures the same rename explicitly. When configured, the rewrite has three consequences worth knowing: * Callers do not need a Fireworks-specific code path. A client that sends `max_completion_tokens` reaches Fireworks with `max_tokens` set to the same value. * If a single request carries both fields, the `max_completion_tokens` value replaces `max_tokens`. The newer field wins because it is the one the caller most likely set deliberately. * The rename applies to the outbound body on every normalized route this provider key serves, not to chat completions alone. A passthrough route is the exception, because it relays the body verbatim, so send the field name Fireworks expects on that route. Fireworks applies its own default output limit to requests that omit the field, so set an explicit limit for long generations. See [Querying text models](https://docs.fireworks.ai/guides/querying-text-models) for the current parameter reference. caution A `request` block supplied on the provider key replaces the built-in block rather than merging with it. If you add your own request overrides for this provider key, re-declare `param_renames` alongside them, or the rewrite stops applying: ``` { "request": { "param_renames": { "max_completion_tokens": "max_tokens" }, "default_headers": { "X-Team": "platform" } } } ``` For the open-source AISIX gateway, retain the rename when adding other request settings: ``` request: param_renames: max_completion_tokens: max_tokens default_headers: X-Team: platform ``` See [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) for the full override schema. ## Use Reasoning Models[​](#use-reasoning-models "Direct link to Use Reasoning Models") Fireworks documents two mutually exclusive reasoning controls, and support for each is model-specific: * `reasoning_effort`, whose accepted effort levels are model-specific and include `low`, `medium`, and `high`. * `thinking`, an object with `type` set to `enabled` and `budget_tokens` of at least `1024`. A request must not set both. AISIX forwards top-level request fields it does not model verbatim to the upstream, so either control reaches Fireworks unchanged: ``` { "model": "fireworks-gptoss-prod", "messages": [{ "role": "user", "content": "Plan a cache invalidation strategy." }], "reasoning_effort": "medium" } ``` Fireworks normally returns reasoning output in `reasoning_content`, which is already the canonical AISIX field, although some models return it in `content` instead. This provider therefore needs no [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) override: AISIX preserves `reasoning_content` as `choices[0].message.reasoning_content` on non-streaming responses and as `delta.reasoning_content` on streaming responses. For models that interleave reasoning with tool calls, retain the complete assistant `reasoning_content` and include it when sending the assistant message back in the next tool-call request. AISIX preserves the field but does not manage or replay application conversation state. Check [Reasoning](https://docs.fireworks.ai/guides/reasoning) for the controls and replay requirements of the selected model. ## Choose Translated or Fireworks-Native Formats[​](#choose-translated-or-fireworks-native-formats "Direct link to Choose Translated or Fireworks-Native Formats") Fireworks publishes a [native Responses API](https://docs.fireworks.ai/guides/response-api) and an [Anthropic-compatible Messages API](https://docs.fireworks.ai/tools-sdks/anthropic-compatibility). The `fireworks-ai` catalog provider still uses the `openai` adapter, so the normalized AISIX routes translate through chat completions instead of calling those native APIs. | Configuration and route | Upstream behavior | | ------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Catalog alias on `/v1/responses` | Uses the AISIX chat-based [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). Fireworks-native state, including `previous_response_id`, stored responses, and server-executed tools, is not available through the bridge. | | `/passthrough/fireworks-ai/responses` | Calls Fireworks' native Responses API. Send the Fireworks model ID rather than the AISIX alias. Fireworks stores native responses by default; send `store: false` when persistence and continuation are not needed. | | Catalog alias on `/v1/messages` | AISIX translates the Anthropic-shaped request to chat completions. It does not call Fireworks' native Messages API. | | `/passthrough/fireworks-ai/messages` | Calls Fireworks' native Anthropic-compatible Messages API with the request and response bodies unchanged. Send the Fireworks model ID rather than the AISIX alias. | | Separate BYO provider key with `adapter: anthropic` and `api_base: https://api.fireworks.ai/inference` | Lets normalized `/v1/messages` traffic call Fireworks' native Messages route. Fireworks' documented Anthropic compatibility limits still apply. | Use a separate [bring-your-own endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md) provider key for the Anthropic adapter because a catalog provider key cannot override its adapter. The `/passthrough/fireworks-ai` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with the Fireworks inference root as its `target_url`; grant the route name on the caller key's `allowed_routes`. A passthrough route does not rewrite AISIX model aliases. AISIX detects chat, completions, and Responses envelopes from each request and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. ## Review Endpoint Support[​](#review-endpoint-support "Direct link to Review Endpoint Support") A Fireworks-backed alias works on the routes below. See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the full endpoint matrix. | Route | Behavior with a `fireworks-ai` provider key | | ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including `stream: true`. | | `/v1/completions` | Supported for Fireworks models that accept the legacy prompt-based completions format. | | `/v1/embeddings` | Supported when the target is a Fireworks embedding model. Use the endpoint-specific model ID, such as `fireworks/qwen3-embedding-8b`, from the [embeddings guide](https://docs.fireworks.ai/guides/querying-embeddings-models). | | `/v1/responses` | Supported through the chat-based Responses bridge, not Fireworks' native Responses API. Fields without a chat equivalent are ignored. Use `/passthrough/fireworks-ai/responses` when native Responses semantics are required. | | `/v1/messages` | Supported through translation to chat completions, not Fireworks' native Messages API. `/v1/messages/count_tokens` is unavailable because it requires an `anthropic` provider. | | `/v1/rerank` | Not supported for a `fireworks-ai` alias because the route's provider allowlist excludes it. Call Fireworks' native reranking API at `/passthrough/fireworks-ai/rerank` and send its endpoint-specific model ID. | | `/v1/audio/*` | Not supported. Fireworks does not publish matching OpenAI-compatible speech or transcription routes under this API base; supported audio and video inputs are sent through multimodal chat models. | | `/v1/images/generations` | Not supported. The route requires a model whose configured provider is `openai`. Fireworks' native image-generation workflow endpoints are available only through `/passthrough/fireworks-ai/workflows/...`. | | `/v1/videos` | Not supported. The route has its own provider allowlist, which does not include `fireworks-ai`. | | `/passthrough/fireworks-ai/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for provider-native endpoints that AISIX does not model. | On a prefix-matched passthrough route, AISIX joins the remaining path onto the route's `target_url`. Because a target of `https://api.fireworks.ai/inference/v1` ends with an API-version segment, a passthrough path that also begins with `v1/` is deduplicated, so `/passthrough/fireworks-ai/v1/<path>` and `/passthrough/fireworks-ai/<path>` both resolve to `https://api.fireworks.ai/inference/v1/<path>`. A route with that target stays under the inference API and cannot reach Fireworks account or deployment management routes at `https://api.fireworks.ai/v1/accounts/...`. Passthrough authorization comes from the caller API key's `allowed_routes` list, which must grant the route name; model allowlists do not gate these paths. On a `raw` route, token counts and token-based costs remain zero, but caller-key request-count limits still apply. See [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Fireworks AI and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Fireworks AI and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Gemini (Google AI Studio) [Google Gemini](https://ai.google.dev/gemini-api/docs) is a family of multimodal models available through Google AI Studio and Google Cloud Vertex AI. This guide connects the Google AI Studio endpoint to AISIX so applications can call Gemini with gateway-managed credentials, access controls, rate limits, and usage accounting. This guide uses the Google AI Studio endpoint. To route Gemini through Google Cloud instead, use [Google Vertex AI](https://docs.api7.ai/ai-gateway/providers/google-vertex-ai.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Google AI Studio API key from [Google AI Studio](https://aistudio.google.com/apikey). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Gemini-backed chat-completions route. Because Google AI Studio provides an OpenAI-compatible endpoint, AISIX connects through the `openai` adapter and uses the Google AI Studio API root as `api_base`. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Google AI Studio credential and API root, and allow the environment to use it: ``` # Replace with your values export GEMINI_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "${AISIX_CP}/provider_keys" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{ "display_name": "gemini-prod", "provider": "google", "api_key": "'"${GEMINI_API_KEY}"'", "api_base": "https://generativelanguage.googleapis.com/v1beta/openai", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "${PROVIDER_KEY_ID}" ``` ❶ `provider` is `google`, the catalog provider ID for Google AI Studio. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the Google AI Studio API key. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is optional for this catalog provider, because the AISIX Cloud Admin API fills in this same value when you omit it. Set it explicitly so the upstream root stays visible on the resource. Use the root without a trailing slash; AISIX appends `/chat/completions` to it. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "${AISIX_CP}/environments/${ENV_ID}/models" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{ "display_name": "gemini-flash-prod", "model_name": "gemini-3.6-flash", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "${MODEL_ID}" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Gemini model ID, for example `gemini-3.6-flash`. Use a stable ID from the [Gemini models](https://ai.google.dev/gemini-api/docs/models) page and review its lifecycle before deploying it; Google retires older stable and preview IDs on published schedules. ❸ `provider_key_id` attaches the alias to the Gemini provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key resource that can access the model alias. The server generates the key value and returns the plaintext once in the response: ``` AISIX_API_KEY=$(curl -sS -X POST "${AISIX_CP}/environments/${ENV_ID}/api_keys" \ -H "Authorization: Bearer ${AISIX_TOKEN}" \ -H "Content-Type: application/json" \ -d '{ "display_name": "gemini-app", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "${AISIX_API_KEY}" ``` The `allowed_models` value references the model by its ID. Store the plaintext key securely; it is not retrievable later. The new resources project to the attached gateway automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export GEMINI_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "gemini-prod" provider: "google" adapter: "openai" api_key: ${GEMINI_API_KEY} api_base: "https://generativelanguage.googleapis.com/v1beta/openai" models: - display_name: "gemini-flash-prod" provider: "google" model_name: "gemini-3.6-flash" provider_key: "gemini-prod" api_keys: - display_name: "gemini-app" key_env: CALLER_API_KEY allowed_models: - "gemini-flash-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "gemini-flash-prod", "messages": [ { "role": "user", "content": "Say hello from Gemini." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `gemini-flash-prod`. If the request fails, check the provider key `api_key`, `api_base`, and the Gemini model ID in `model_name`. ## Use Gemini Thinking and Tools[​](#use-gemini-thinking-and-tools "Direct link to Use Gemini Thinking and Tools") Google's OpenAI-compatible chat endpoint accepts `reasoning_effort`. AISIX forwards this and other unmodeled top-level fields unchanged, so callers can use the levels supported by the selected Gemini model: ``` { "model": "gemini-flash-prod", "messages": [{"role": "user", "content": "Plan a safe database migration."}], "reasoning_effort": "medium" } ``` Gemini-specific controls under `extra_body.google`, such as `thinking_config`, also pass through. Do not send both `reasoning_effort` and Google's thinking-level or thinking-budget control in one request because they configure the same behavior. For `gemini-3.6-flash` and newer models, remove deprecated sampling fields such as `temperature`, `top_p`, and `top_k`; AISIX forwards them rather than removing them for you. Gemini 3 tool calls carry a required thought signature at `tool_calls[].extra_content.google.thought_signature`. When continuing a function-calling turn on `/v1/chat/completions`, retain the complete assistant `tool_calls` objects and send them back unchanged with the tool results. AISIX preserves those raw objects, but it does not manage application conversation state. Do not use the `/v1/responses` or translated `/v1/messages` routes for multi-step Gemini 3 tool loops. Those bridges reconstruct portable tool-call fields and do not replay Google's provider-specific thought signatures, so the next Gemini request can fail with a `400` validation error. See [Thought signatures for OpenAI compatibility](https://ai.google.dev/gemini-api/docs/generate-content/thought-signatures#openai). ## Review Endpoint Support[​](#review-endpoint-support "Direct link to Review Endpoint Support") The Google AI Studio OpenAI-compatible API has more routes than AISIX currently models for the `google` provider. Choose the route according to both layers: | Route | Behavior with a `google` provider key | | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including streaming and multimodal image, audio, or video inputs accepted by the selected model. | | `/v1/embeddings` | Supported with a separate alias for a Gemini embedding model, such as `gemini-embedding-2`. | | `/v1/responses` | Supported through the chat-based [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). It does not call Google's native Interactions API, and fields without a chat equivalent are ignored. | | `/v1/messages` | Supported through translation to chat completions. `/v1/messages/count_tokens` is unavailable because it requires an Anthropic-backed model. | | `/v1/completions` | Not supported by Google's compatibility API, which documents chat completions rather than the legacy prompt-completions route. | | `/v1/audio/*` | Not supported. Send supported audio input through chat completions; Gemini speech generation and the Live API use different native APIs. | | `/v1/images/generations` | Rejected because the normalized AISIX route requires `provider: openai`, even though Google publishes an OpenAI-compatible image route. Use `/passthrough/google/images/generations` with a current Google image model ID. | | `/v1/videos` | Rejected because `google` is not in the normalized video route's provider allowlist. Use `/passthrough/google/videos` to submit and `/passthrough/google/videos/<id>` to poll Google's OpenAI-compatible video operations. | | `/v1/rerank` | Not supported. The AISIX route's provider allowlist excludes `google`, and Google does not publish a matching rerank route here. | | `/v1/files` and `/v1/batches` | Not an end-to-end AISIX workflow. Google supports OpenAI-compatible batch creation and status, but file upload and download require its native API outside this catalog route. | | `/passthrough/google/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for raw routes under the OpenAI-compatible API base. | The `/passthrough/google` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix, with the catalog `api_base` as its `target_url` and this provider key attached; grant the route name on the caller API key's `allowed_routes`. That target ends in `/v1beta/openai`, so passthrough stays inside Google's OpenAI-compatible API section. It cannot reach sibling native routes such as `models/<model>:generateContent`, `models/<model>:streamGenerateContent`, `interactions`, native file operations, or the Live API. Those native APIs also use a different request shape and API-key header, so do not construct native Gemini paths behind this catalog provider key. Passthrough forwards request and response bodies without rewriting an AISIX model alias, so send the exact Google model ID. AISIX detects chat, completions, and Responses envelopes from each request and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. Recorded passthrough tokens do not advance `tpm` or `tpd` counters, resolve a model cost, or add budget spend; caller-key request-count limits still apply to every call. See [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Gemini through Google AI Studio and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Google Vertex AI](https://docs.api7.ai/ai-gateway/providers/google-vertex-ai.md): route Gemini through Google Cloud instead. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Google Vertex AI [Google Vertex AI](https://cloud.google.com/vertex-ai/generative-ai/docs) is Google Cloud's managed platform for Gemini and partner models. AISIX gives applications one OpenAI-compatible API for these Vertex-hosted models. This configuration is for Vertex-hosted models that should use AISIX authentication, model allowlists, rate limits, and usage accounting. AISIX authenticates to Vertex with a GCP OAuth2 bearer token and can mint that token from a service-account key. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * The Vertex AI API enabled in the target GCP project. * A Vertex location supported by the target model and a service account that can invoke it. The current Gemini example uses `global`. * The GCP project ID and Vertex model ID. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a Vertex provider key, model alias, and caller API key. The provider key stores the GCP project, region, and credential mode; the model selects the Vertex publisher model ID. ### Create a Vertex Provider Key[​](#create-a-vertex-provider-key "Direct link to Create a Vertex Provider Key") Create the Vertex provider key with the GCP credential settings: ``` PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "vertex-prod", "provider": "google-vertex", "api_base": "https://aiplatform.googleapis.com", "config": { "project": "my-gcp-project", "region": "global", "service_account_json": { "type": "service_account", "private_key": "-----BEGIN PRIVATE KEY-----\nYOUR_SERVICE_ACCOUNT_PRIVATE_KEY\n-----END PRIVATE KEY-----\n", "client_email": "vertex-sa@my-gcp-project.iam.gserviceaccount.com", "token_uri": "https://oauth2.googleapis.com/token" } }, "allowed_environments": ["'"$ENV_ID"'"] }' | jq -r '.provider_key.id') ``` ❶ `provider` selects the Google Vertex catalog entry, which routes traffic through the Vertex protocol adapter. ❷ AISIX Cloud requires `api_base` for this platform provider. For the `global` location, use `https://aiplatform.googleapis.com`; `https://global-aiplatform.googleapis.com` is not the global endpoint. For a regional location, use `https://<region>-aiplatform.googleapis.com`, matching `config.region`. A proxy or private endpoint can replace either origin. ❸ `config` is the structured credential with `project`, `region`, and exactly one credential mode. The example uses `service_account_json` as a nested object. Omit `api_key` when the credential is supplied through `config`; a non-empty `api_key` is rejected for this provider. Use `service_account_json` unless you already manage short-lived GCP access tokens yourself. AISIX signs a JWT, mints an OAuth token, caches it, and refreshes it before expiry. If you use `access_token`, you are responsible for refreshing it. The current adapter does not discover Application Default Credentials, metadata-server credentials, or Workload Identity Federation credentials. Provider key secrets follow the credential-handling behavior described in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). The command captures the returned provider key ID for the model resource. `allowed_environments` lets the environment reference this key. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Map a caller-facing alias to the Vertex model ID: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "gemini-prod", "model_name": "gemini-3.6-flash", "provider_key_id": "'"$PROVIDER_KEY_ID"'" }' | jq -r '.model.id') ``` ❶ `model_name` is the Vertex publisher model ID. The current [Gemini 3.6 Flash](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-6-flash) example is generally available in `global`. Google documents custom temperature, top-P, and top-K values as ignored for this model, so omit `temperature` and `top_p` from calls through AISIX. Requests to this model must also end with a user message. Google rejects a GenerateContent request whose final turn has the `model` role, and AISIX maps a final OpenAI `assistant` message to that role. ❷ `provider_key_id` attaches the model to the Vertex credential captured in the previous step. Other supported examples include Claude on Vertex, OpenAI-compatible MaaS models such as Llama, and partner publisher models from Mistral and AI21. AISIX chooses the Vertex route from `model_name`: | Model ID Family | How AISIX Sends It to Vertex | | ------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------- | | Gemini models, such as `gemini-*` | Uses the Google Gemini publisher route. | | Claude models, such as `claude-*` | Uses the Anthropic publisher route with an Anthropic Messages body. | | OpenAI-compatible MaaS models, such as Llama, DeepSeek, Qwen, GPT-OSS, MiniMax, Moonshot, or Z.ai | Uses Vertex's OpenAI-compatible chat-completions route. | | Mistral and AI21 models | Uses the partner publisher route with an OpenAI-compatible body. | ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the API key resource with access to the Vertex-backed model alias. The server generates the key and returns the plaintext once: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "vertex-caller", "allowed_models": ["'"$MODEL_ID"'"] }' | jq -r '.plaintext') ``` `allowed_models` references the model by ID, so the caller can only use the Vertex-backed alias. Store the plaintext key securely; read endpoints do not return it again. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") For the open-source gateway, set `adapter: vertex` and place the project, region, and credential mode in the provider key's `api_key` value. The example below reads the structured credential from an environment variable: ``` export VERTEX_CREDENTIAL='{"project":"my-gcp-project","region":"global","service_account_json":{"type":"service_account","private_key":"YOUR_PRIVATE_KEY","client_email":"vertex-sa@my-gcp-project.iam.gserviceaccount.com","token_uri":"https://oauth2.googleapis.com/token"}}' ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: vertex-prod provider: google-vertex adapter: vertex api_key: ${VERTEX_CREDENTIAL} api_base: https://aiplatform.googleapis.com models: - display_name: gemini-prod provider: google-vertex model_name: gemini-3.6-flash provider_key: vertex-prod api_keys: - display_name: vertex-caller key_env: VERTEX_CALLER_KEY allowed_models: - gemini-prod ``` The credential must contain `project`, `region`, and exactly one of `access_token` or `service_account_json`. Use the same `model_name` values described above for other Vertex publishers. For a regional OSS configuration, `api_base` can be omitted and AISIX derives `https://<region>-aiplatform.googleapis.com`. Keep it explicit for `region: global`, because the correct origin is `https://aiplatform.googleapis.com`. Also set it explicitly for a proxy or private endpoint. Export the caller key: ``` export VERTEX_CALLER_KEY="YOUR_CALLER_API_KEY" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, set `AISIX_API_KEY` for the shared verification request: ``` export AISIX_API_KEY="$VERTEX_CALLER_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy. The example uses Gemini, which requires at least one user or assistant turn. ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "gemini-prod", "messages": [ { "role": "user", "content": "Say hello from Vertex." } ] }' ``` The gateway returns an OpenAI-compatible response with the caller-facing alias: ``` { "object": "chat.completion", "model": "gemini-prod", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Hello from Vertex!" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 4, "completion_tokens": 4, "total_tokens": 8 } } ``` Check Vertex logs, metrics, quota usage, or provider-side request records for the test request. If AISIX returns a token-minting or upstream authentication error, check the service-account key, region, Vertex API enablement, IAM role, and model access. ## Endpoint and Content Support[​](#endpoint-and-content-support "Direct link to Endpoint and Content Support") Vertex AI exposes more capabilities than the current AISIX `vertex` adapter translates. The gateway behavior depends on both the caller route and the publisher family selected by `model_name`. | Route | Behavior with a `google-vertex` provider key | | ----------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including streaming. The current Gemini publisher translation is text-only; partner publisher behavior depends on its rail. | | `/v1/embeddings` | Supported for Google-publisher text embedding models, such as `gemini-embedding-001`. AISIX sends text inputs to `publishers/google/models/<model>:predict`; partner and multimodal embedding formats are not supported. With `gemini-embedding-001`, send one input string per request. AISIX forwards an input array as one upstream request, but this model accepts only one input and can reject an array with multiple items. | | `/v1/responses` | Supported through the chat-based [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), not a native Vertex Responses endpoint. Fields without a chat equivalent are ignored. | | `/v1/messages` | Supported through translation to chat. `/v1/messages/count_tokens` is unavailable because the configured provider is `google-vertex`, including for Claude models hosted on Vertex. | | `/v1/completions` | Not supported by the Vertex adapter. | | `/v1/images/generations`, `/v1/audio/*`, `/v1/videos`, and `/v1/rerank` | Not supported by the Vertex adapter or these routes' provider rules, even though Vertex offers separate native media services. | | `/v1/files`, `/v1/batches`, and `/v1/fine_tuning/jobs` | Not supported because the AISIX jobs surface does not accept the `vertex` adapter. | | `/passthrough/google-vertex/*` | Not a workaround for native Vertex APIs, even through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). Passthrough routes bypass the adapter, so credential injection neither mints OAuth tokens from the service-account JSON nor extracts an access token from the structured credential. | For Gemini publisher models, AISIX currently serializes text parts, system instructions, and basic generation settings. It does not serialize image, audio, or video parts; function declarations and tool results; Gemini-specific thinking controls; or thought signatures. Tool messages are reduced to user text, and non-text response parts are not returned. Do not use the current Vertex adapter for Gemini function-calling loops. Gemini 3 requires applications to replay thought signatures during tool use, but this adapter neither returns nor accepts those signatures. This limitation also applies when Gemini is reached through the Responses or Messages bridge. AISIX maps Vertex's token counters into the normalized response rather than only their total. Thinking tokens (`thoughtsTokenCount`) count toward `completion_tokens`, as Vertex bills them, and are named separately under `completion_tokens_details.reasoning_tokens`. Context-cache hits (`cachedContentTokenCount`) stay a subset of `prompt_tokens` and are reported under `prompt_tokens_details.cached_tokens`. Those are the Chat Completions field names. The same values reach a [Responses](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) caller as `output_tokens_details.reasoning_tokens` and `input_tokens_details.cached_tokens`, and an [Anthropic Messages](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md) caller sees the cache hit as `cache_read_input_tokens` and no separate reasoning counter, because that protocol has none. ## Vertex Publisher Routing[​](#vertex-publisher-routing "Direct link to Vertex Publisher Routing") Publisher selection is prefix-based. If `model_name` does not match a supported prefix, the gateway rejects the request before the provider request with an unsupported-publisher configuration error. The example uses Gemini because it is the main Google publisher path. For partner models, validate the exact model ID, quota, and regional availability in your Vertex project before exposing the alias to callers. A provider key has one project and location. Create another provider key when a partner model is available only in a different location; changing `model_name` does not change the location used in the Vertex resource path. [Provider-key request and response overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) can apply on Vertex routes, but they are most directly useful on OpenAI-compatible routes. Gemini's native `contents` format does not match every OpenAI-style override target. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Google Vertex AI and verified the model alias. Continue with these guides: * [Gemini (Google AI Studio)](https://docs.api7.ai/ai-gateway/providers/gemini.md): configure Gemini with an AI Studio API key instead of a Google Cloud project. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Groq [Groq](https://console.groq.com/docs) provides hosted inference for a catalog of open models. AISIX gives applications stable model aliases and caller keys while keeping the Groq credential at the gateway. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Groq API key from the [Groq Console](https://console.groq.com/keys). * `curl`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Groq-backed chat-completions route. Because Groq exposes an OpenAI-compatible API, AISIX connects through the `openai` adapter and uses the Groq API root as `api_base`. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Groq credential and API root: ``` # Replace with your values export GROQ_API_KEY="YOUR_PROVIDER_API_KEY" curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "groq-prod", "provider": "groq", "api_key": "'"${GROQ_API_KEY}"'", "api_base": "https://api.groq.com/openai/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' ``` ❶ `provider` is `groq`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the adapter field is only accepted on BYO provider keys. ❷ `api_key` stores the Groq API key. The value is encrypted before storage and never returned by read endpoints. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` already includes the `/openai/v1` path. AISIX appends `/chat/completions` to it. The field is optional for this catalog provider because the AISIX Cloud Admin API fills in the same value when you omit it. The example sets it explicitly so the upstream root stays visible on the resource. AISIX Cloud currently retains a legacy compatibility rename from `max_completion_tokens` to `max_tokens`. Groq still accepts `max_tokens` but now deprecates it in favor of `max_completion_tokens`. The open-source configuration below intentionally omits this override so current clients can send `max_completion_tokens` directly. Copy the `provider_key.id` value from the response (for example with `jq -r '.provider_key.id'`) and export it: ``` export PROVIDER_KEY_ID="YOUR_PROVIDER_KEY_ID" ``` ### Create a Model[​](#create-a-model "Direct link to Create a Model") Groq model availability changes over time. Check the [Groq models list](https://console.groq.com/docs/models) for a current model ID before creating a model alias. Create the model alias callers will send in requests: ``` curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "groq-gptoss-prod", "model_name": "openai/gpt-oss-120b", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Groq model ID, for example `openai/gpt-oss-120b`. ❸ `provider_key_id` attaches the alias to the Groq provider key. Capture the `model.id` value from the response: ``` export MODEL_ID="YOUR_MODEL_ID" ``` ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key resource that can access the model alias. The gateway generates the key value and returns the plaintext once in the response: ``` curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "groq-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' ``` The `allowed_models` value must reference the model ID you captured. Copy the `plaintext` value from the response — it is shown only once — and export it as the caller key: ``` export AISIX_API_KEY="YOUR_CALLER_API_KEY" ``` After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export GROQ_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "groq-prod" provider: "groq" adapter: "openai" api_key: ${GROQ_API_KEY} api_base: "https://api.groq.com/openai/v1" models: - display_name: "groq-gptoss-prod" provider: "groq" model_name: "openai/gpt-oss-120b" provider_key: "groq-prod" api_keys: - display_name: "groq-caller" key_env: CALLER_API_KEY allowed_models: - "groq-gptoss-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "groq-gptoss-prod", "messages": [ { "role": "user", "content": "Say hello from Groq." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `groq-gptoss-prod`. If the request fails, check the provider key `api_key`, `api_base`, and the Groq model ID in `model_name`. ## Endpoint and Compatibility Boundaries[​](#endpoint-and-compatibility-boundaries "Direct link to Endpoint and Compatibility Boundaries") Groq is mostly, but not completely, OpenAI-compatible. AISIX forwards additional chat request fields instead of filtering them, so Groq enforces its own model and parameter rules. Groq's compatibility guide says `logprobs`, `logit_bias`, `top_logprobs`, and `messages[].name` return `400`, and `n` must be `1`. Groq's API reference also marks `frequency_penalty`, `presence_penalty`, `metadata`, and `store` as unsupported. A `temperature` of `0` is converted upstream to `1e-8`. For audio transcription and translation, Groq does not support `vtt` or `srt` output. See [Groq's OpenAI compatibility guide](https://console.groq.com/docs/openai) for the current list. | Route | Behavior with a `groq` provider key | | ---------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including streaming. AISIX forwards fields such as `tools`, `response_format`, and `reasoning_effort`; support still depends on the selected Groq model. The example GPT-OSS model supports tool use, structured output, and `low`, `medium`, or `high` reasoning effort. | | `/v1/responses` | Supported through the chat-based [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), not [Groq's native beta Responses API](https://console.groq.com/docs/responses-api). Fields and response items without a chat equivalent are not preserved. | | `/v1/messages` | Supported through translation to chat, not a native Groq Messages API. `/v1/messages/count_tokens` is unavailable because the configured provider is not Anthropic. | | `/v1/audio/transcriptions`, `/v1/audio/translations`, and `/v1/audio/speech` | Supported with separate model aliases for current Groq speech-to-text or text-to-speech models. Check Groq's model list before configuring an audio alias. | | `/v1/files` and `/v1/batches` | Supported through Groq's OpenAI-compatible file and batch APIs. Use the Groq alias as the routing model as described in [Batch, Files, and Fine-Tuning](https://docs.api7.ai/ai-gateway/endpoints/batch-files-fine-tuning.md). Groq applies discounted batch pricing, but AISIX estimates batch cost from the alias's configured synchronous prices; do not treat that estimate as the provider invoice. | | `/v1/fine_tuning/jobs` | Not compatible with Groq fine-tuning. AISIX forwards the OpenAI `/fine_tuning/jobs` contract under `api_base`, while Groq's closed-beta API uses `/v1/fine_tunings` outside the `/openai/v1` root and has a different body. | | `/v1/embeddings` and `/v1/completions` | Not supported because Groq does not expose those upstream endpoints. | | `/v1/images/generations`, `/v1/videos`, and `/v1/rerank` | Not supported by Groq or by these AISIX routes' provider rules. | The normalized chat response preserves standard content, tool calls, reasoning text, and token totals. AISIX returns Groq's native `message.reasoning` value as `reasoning_content`. It does not preserve Groq-specific metadata such as `x_groq`, usage timing details, `usage_breakdown`, or unmodeled citation and annotation fields. To use Groq's native Responses wire shape, call `/passthrough/groq/responses`. This path assumes a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming the `/passthrough/groq` prefix with the configured `https://api.groq.com/openai/v1` base as its `target_url`; grant the route on the caller key's `allowed_routes`, and AISIX appends the remaining path to that target. Send the exact Groq model ID because passthrough does not rewrite AISIX aliases. AISIX relays the provider response body without normalization unless a guardrail blocks it. It detects the Responses envelope from `input` and records supported `input_tokens` and `output_tokens` usage. The route's target also means it cannot reach Groq's sibling `/v1/fine_tunings` API. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Groq and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Groq and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Hugging Face [Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers) routes requests to open-weight models served by multiple inference vendors. AISIX puts that catalog behind one OpenAI-compatible API with gateway-managed credentials, caller access, rate limits, and usage accounting. Unlike a single-vendor upstream, Hugging Face uses a routing layer: the model ID in the request, rather than the URL, selects which inference provider serves the model. This guide covers the shared Inference Providers router. A dedicated Hugging Face Inference Endpoint has its own per-deployment hostname and is not reachable through the router URL. Configure that deployment as a [private OpenAI-compatible endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md) only when its serving engine exposes the OpenAI route your application needs, such as `/v1/chat/completions` or `/v1/embeddings`. A custom or task-native endpoint is not automatically OpenAI-compatible. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Hugging Face access token with the **Make calls to Inference Providers** permission, created in [Access Tokens](https://huggingface.co/settings/tokens). * A Hugging Face account with remaining Inference Providers credits. * `curl` and `jq`. A Hugging Face access token is an account-level credential rather than a per-vendor API key. One token authorizes calls to every inference provider the router can select, so a single provider key in AISIX covers the whole router catalog. Scope the token to Inference Providers only, so that the credential stored in AISIX cannot also read or write Hub repositories. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Hugging Face-backed chat-completions route. Because Hugging Face exposes an OpenAI-compatible API, AISIX connects through the `openai` adapter and uses the Inference Providers router as `api_base`. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Hugging Face token and the router root: ``` # Replace with your value export HF_TOKEN="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "huggingface-prod", "provider": "huggingface", "api_key": "'"${HF_TOKEN}"'", "api_base": "https://router.huggingface.co/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `huggingface`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the Hugging Face access token. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is the Inference Providers router root. The host is `router.huggingface.co` rather than a per-model host, and `/v1` is the OpenAI-compatible surface that sits in front of every routed inference provider. AISIX appends the endpoint path, so the chat route resolves to `https://router.huggingface.co/v1/chat/completions`. If you omit `api_base`, AISIX Cloud fills in the same router URL from the catalog; setting it explicitly keeps the resource self-describing. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Hugging Face model IDs are Hub repository IDs in the form `<org>/<model>`, and the casing is not uniform across publishers: `openai/gpt-oss-120b` is entirely lowercase, while `Qwen/Qwen3-235B-A22B-Thinking-2507` and `deepseek-ai/DeepSeek-V4-Pro` use mixed case. Copy the ID verbatim from the [supported models list](https://huggingface.co/inference/models) rather than retyping it, and confirm there that the model is currently served before creating a model alias. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "hf-gptoss-prod", "model_name": "openai/gpt-oss-120b", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Hugging Face model ID, for example `openai/gpt-oss-120b` or `deepseek-ai/DeepSeek-V4-Pro`. AISIX forwards this string to the router unchanged. ❸ `provider_key_id` attaches the alias to the Hugging Face provider key. ### Pin the Serving Inference Provider[​](#pin-the-serving-inference-provider "Direct link to Pin the Serving Inference Provider") Hugging Face accepts an optional suffix on the model ID that controls which inference provider serves the request. The suffix is part of the model string, so it belongs in `model_name`: | `model_name` value | Routing behavior | | ------------------------------- | -------------------------------------------------------------------------------------------- | | `openai/gpt-oss-120b` | Automatic routing, which by default selects the fastest available inference provider. | | `openai/gpt-oss-120b:groq` | Pinned to the named inference provider. | | `openai/gpt-oss-120b:cheapest` | Routed to the most cost-efficient provider by price per output token. | | `openai/gpt-oss-120b:fastest` | Routed to the highest-throughput provider. This is the default policy. | | `openai/gpt-oss-120b:preferred` | Routed by the preference order configured in your Hugging Face Inference Providers settings. | See the [Hugging Face chat completion documentation](https://huggingface.co/docs/inference-providers/en/tasks/chat-completion) for the current suffix syntax. The suffix changes latency, price, and which backend actually runs the model, so treat pinned and policy-routed variants as different upstreams. Create one model alias per variant rather than switching the suffix in place, so rate limits and usage records remain attributable to that routing choice. Pin the inference provider when accurate AISIX cost and budget calculations matter. With `:fastest`, `:cheapest`, or `:preferred`, the selected provider and price can change, so verify or override the alias's cost metadata. See [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md) for the alias fields involved. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the response — store it securely: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "huggingface-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. The gateway picks up the new resources automatically — no restart is needed. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export HF_TOKEN="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "huggingface-prod" provider: "huggingface" adapter: "openai" api_key: ${HF_TOKEN} api_base: "https://router.huggingface.co/v1" models: - display_name: "hf-gptoss-prod" provider: "huggingface" model_name: "openai/gpt-oss-120b" provider_key: "huggingface-prod" api_keys: - display_name: "huggingface-caller" key_env: CALLER_API_KEY allowed_models: - "hf-gptoss-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "hf-gptoss-prod", "messages": [ { "role": "user", "content": "Say hello from Hugging Face." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `hf-gptoss-prod`. If the request fails, check the provider key `api_key` and `api_base`. Verify the exact capitalization of the Hub repository ID in `model_name`. Ensure that the token carries the Inference Providers permission and that the account has remaining Inference Providers credits. Provider availability also varies by model, so confirm on the [supported models list](https://huggingface.co/inference/models) that the model is currently served. ## Control Reasoning Output[​](#control-reasoning-output "Direct link to Control Reasoning Output") Reasoning models on the router accept a top-level `reasoning_effort` field in the chat-completions request body: ``` { "reasoning_effort": "low" } ``` Common values are `none`, `minimal`, `low`, `medium`, `high`, and `xhigh`. Support, defaults, and which values are meaningful depend on both the model and its serving inference provider. The router documents no shared toggle or numeric reasoning budget, so confirm your model's values in the [Hugging Face chat completion documentation](https://huggingface.co/docs/inference-providers/en/tasks/chat-completion). AISIX forwards the field without interpreting it. The AISIX catalog entry for `huggingface` carries no `response.reasoning_field` override. AISIX recognizes both `delta.reasoning_content` and `delta.reasoning`, normalizing the latter to `reasoning_content`. Because the same Hub repository ID can be served by different inference providers, a pinned provider may stream reasoning under another `delta` path. In that case, set [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) on a provider key dedicated to that pinned alias, rather than on the shared router provider key that every other alias inherits. ## Supported Proxy Routes[​](#supported-proxy-routes "Direct link to Supported Proxy Routes") The Hugging Face provider value resolves to the `openai` adapter, but the shared router implements only part of the broader OpenAI API. Provider and model capabilities can also differ behind the router. | Route | Behavior with a `huggingface` provider key | | ---------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including streaming. AISIX forwards function tools, `response_format`, `reasoning_effort`, and VLM `image_url` content. Actual tool use, structured output, reasoning, and vision support depend on both the Hub model and the inference provider selected for that request. | | `/v1/responses` | Supported through the AISIX [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), not Hugging Face's native [Responses API](https://huggingface.co/docs/inference-providers/guides/responses-api). The bridge carries text, local function calls and results, sampling fields, and streaming through chat completions. It drops non-text image and file input parts, hosted tools such as remote MCP, `reasoning` and structured-output `text` configuration, and state fields such as `previous_response_id` and `store`. It also omits reasoning output when re-encoding the chat result. | | `/v1/messages` | Supported through translation to chat completions, not a native Hugging Face Messages API. `/v1/messages/count_tokens` is unavailable because the configured provider is not Anthropic. | | `/v1/embeddings` | Not available through the router's OpenAI-compatible `/v1` surface. Hugging Face provides embeddings through its task-specific feature-extraction interface, but AISIX does not translate the OpenAI embeddings body to that interface. A dedicated Inference Endpoint works only if its serving engine exposes `/v1/embeddings`, such as Text Embeddings Inference. | | `/v1/completions`, `/v1/audio/*`, `/v1/files`, `/v1/batches`, and `/v1/fine_tuning/jobs` | Not available on the configured shared router root. Hugging Face has separate task-native interfaces for capabilities such as text generation and speech, but these AISIX routes do not translate to or reach those interfaces. | | `/v1/images/generations`, `/v1/videos`, and `/v1/rerank` | Not supported for a `huggingface` provider value. Hugging Face may offer related native tasks, but these AISIX routes reject this provider before forwarding. | Normalized chat responses preserve content, local function calls, normalized reasoning, and basic token usage. They do not preserve every router or serving-provider field, such as `created`, `system_fingerprint`, choice indexes and log probabilities, or arbitrary provider-specific metadata. To use Hugging Face's native Responses wire shape, send `POST /passthrough/huggingface/responses` with the exact Hugging Face model ID in `model`; passthrough does not rewrite an AISIX alias. AISIX relays the response body without normalization unless a guardrail blocks it. It detects the Responses envelope from `input` and records token usage when the response includes supported `input_tokens` and `output_tokens` fields. The `/passthrough/huggingface` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with the configured `https://router.huggingface.co/v1` root as its `target_url`; grant the route on the caller key's `allowed_routes`. Such a route can reach `/v1/responses` and `/v1/models`, but it cannot escape that root to sibling task-native paths such as `/hf-inference/models/<model>`. Because the target already ends in `/v1`, AISIX removes a duplicate leading `v1` segment, so both `/passthrough/huggingface/models` and `/passthrough/huggingface/v1/models` reach `https://router.huggingface.co/v1/models`. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Hugging Face and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Hugging Face and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Jina [Jina AI](https://jina.ai/) provides embedding and reranker models for search and retrieval applications. Applications can call those models through the AISIX gateway's OpenAI-compatible embeddings route and unified rerank route. The gateway manages the Jina credential, caller access, and rate limits. Jina serves embeddings and reranking from the same API root, so one provider key can support both model types. This guide configures an embedding model first, then reuses the provider key for a reranker. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Jina API key from the [Jina AI API dashboard](https://jina.ai/api-dashboard/). One key authorizes all Jina API products, including embeddings and reranking. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Jina-backed embeddings route. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Jina credential and API root: ``` # Replace with your value export JINA_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "jina-prod", "provider": "jina", "api_key": "'"${JINA_API_KEY}"'", "api_base": "https://api.jina.ai/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `jina`, the provider value the AISIX rerank route recognizes. Jina is absent from models.dev, but AISIX Cloud accepts it as a directly supported provider and derives the `openai` adapter for embeddings and other OpenAI-shaped routes. The `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the Jina API key and is sent as a bearer token on upstream calls. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is the versioned Jina API root. The embeddings route appends `/embeddings` to this value, and the rerank route recognizes the trailing `/v1` before appending `/rerank`, so both routes compose the correct upstream URL from one provider key. The field is optional for this provider; when it is omitted, the AISIX Cloud Admin API fills in the same canonical value. Set it explicitly when you point the key at a different Jina deployment. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") This example uses [`jina-embeddings-v5-text-small`](https://jina.ai/models/jina-embeddings-v5-text-small/), the current 1024-dimension text embedding model. Jina also publishes multimodal v5 models for text, image, audio, video, and PDF input. Those non-text input shapes require a passthrough route because the normalized AISIX route accepts strings only. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "jina-embed-prod", "model_name": "jina-embeddings-v5-text-small", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the exact Jina model ID. Because Jina is not on models.dev, the dashboard suggests no model IDs for this provider; enter an ID from Jina's current model catalog. ❸ `provider_key_id` attaches the alias to the Jina provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "jina-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step, so the key can only access the alias you created. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export JINA_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "jina-prod" provider: "jina" adapter: "openai" api_key: ${JINA_API_KEY} api_base: "https://api.jina.ai/v1" models: - display_name: "jina-embed-prod" provider: "jina" model_name: "jina-embeddings-v5-text-small" provider_key: "jina-prod" api_keys: - display_name: "jina-caller" key_env: CALLER_API_KEY allowed_models: - "jina-embed-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send an embeddings request through the AISIX proxy, with a `dimensions` value below the model's default: ``` curl -sS -X POST "$AISIX_PROXY/v1/embeddings" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "jina-embed-prod", "input": "AISIX keeps the provider credential on the gateway side.", "dimensions": 128 }' -o jina-embed-response.json jq '.data[0].embedding | length' jina-embed-response.json ``` The command should print `128`. The vector length matching the requested `dimensions` confirms the optional field reached Jina, which truncates the model's default 1024-dimension output to the requested size. AISIX reconstructs the accepted dense result in its OpenAI embeddings response shape. The response model field echoes the caller-facing alias; AISIX preserves the float or base64 vector, index, and prompt and total token counters. It does not preserve Jina-specific usage fields such as `image_tokens`, `audio_tokens`, or `video_tokens`. If the request fails, check the provider key `api_key`, `api_base`, and the Jina model ID in `model_name`. ## Send Jina-Specific Embedding Fields[​](#send-jina-specific-embedding-fields "Direct link to Send Jina-Specific Embedding Fields") The modeled `/v1/embeddings` route accepts only `model`, `input` as a string or array of strings, `encoding_format`, and `dimensions`. AISIX preserves the caller's single-string or array input shape upstream. Jina documents `embedding_type`, rather than `encoding_format`, for output encoding. Jina-specific fields such as `task`, `embedding_type`, `normalized`, and `truncate` are dropped before the request reaches Jina. Object inputs for images, audio, video, or PDF documents do not match the AISIX request schema and are rejected during request decoding. Jina sparse embeddings and v4 multi-vector responses also fall outside the AISIX response schema and fail upstream response decoding. Use a passthrough route for these request or response shapes. To send them, call Jina through a passthrough route instead, which forwards the request body verbatim. The `/passthrough/jina` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with `https://api.jina.ai/v1` as its `target_url` and the Jina provider key attached; grant the route on the caller key's `allowed_routes`: ``` curl -sS -X POST "$AISIX_PROXY/passthrough/jina/embeddings" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "jina-embeddings-v5-text-small", "task": "retrieval.passage", "embedding_type": "float", "normalized": true, "truncate": true, "input": [ "Provider keys store the upstream credential.", "Caller API keys authorize model access." ] }' ``` The route does not rewrite the body, so `model` must be the upstream model ID rather than the alias. It injects the provider key configured on the route rather than borrowing one from the caller key's model allowlist; when more than one Jina provider key exists, give each its own route. The route appends the remaining path to its `target_url`, producing `https://api.jina.ai/v1/embeddings`, and relays the response body unchanged. Passthrough detection is key-based rather than endpoint-based, so this embeddings body's top-level `input` makes AISIX apply Responses-style extraction and usage rules. AISIX records usage only when the response includes supported token fields. Do not rely on passthrough accounting for Jina-specific response shapes. See [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for the route's behavior and limits, and the [Embedding API reference](https://jina.ai/embeddings/) for current model-specific fields. ## Add a Rerank Model[​](#add-a-rerank-model "Direct link to Add a Rerank Model") The `/v1/rerank` route accepts a model whose provider value is `openai`, `cohere`, or `jina`. For `jina`, the request and response wire shapes match the gateway's unified rerank contract. AISIX forwards Jina's `model`, `query`, `documents`, and optional fields verbatim, with only the `model` field rewritten to the upstream model ID. Jina serves rerank from the same `https://api.jina.ai/v1` root as embeddings, so the provider key created above already reaches it. No second provider key on a different API root is needed. This differs from providers whose rerank endpoint lives outside their OpenAI-compatible surface; compare [Cohere's rerank setup](https://docs.api7.ai/ai-gateway/providers/cohere.md#add-a-rerank-model). The rerank route appends `/rerank` and recognizes a base that already ends in `/v1`, producing `https://api.jina.ai/v1/rerank` without a duplicate version segment. In AISIX Cloud, create the rerank model alias and a caller key scoped to it: ``` RERANK_MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "jina-rerank-prod", "model_name": "jina-reranker-v3.5", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') RERANK_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "jina-rerank-caller", "allowed_models": ["'"${RERANK_MODEL_ID}"'"] }' | jq -r '.plaintext') ``` For the open-source AISIX gateway, add `jina-rerank-prod` to the existing `models` collection. Replace the existing `jina-caller` entry with the updated entry below so it allows both model aliases. Preserve unrelated entries and collections: resources.yaml (rerank model access) ``` models: - display_name: "jina-rerank-prod" provider: "jina" model_name: "jina-reranker-v3.5" provider_key: "jina-prod" api_keys: - display_name: "jina-caller" key_env: CALLER_API_KEY allowed_models: - "jina-embed-prod" - "jina-rerank-prod" ``` For the open-source setup, validate and reload or restart the declarative resources file as described above, then use the existing caller key for the rerank request: ``` export RERANK_API_KEY="$CALLER_API_KEY" ``` Reranker IDs follow their own naming, separate from the embedding generations. [`jina-reranker-v3.5`](https://jina.ai/models/jina-reranker-v3.5/) is the current multilingual, multi-document reranker and a drop-in successor to `jina-reranker-v3`. Check the [Reranker API reference](https://jina.ai/reranker/) for the current catalog. Send a rerank request through the proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/rerank" \ -H "Authorization: Bearer ${RERANK_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "jina-rerank-prod", "query": "How do I rotate a provider credential?", "documents": [ "Provider keys store the upstream credential.", "Caller API keys authorize model access.", "Rate limits apply per caller key." ], "top_n": 2 }' ``` AISIX rewrites only the `model` field to `jina-reranker-v3.5` and forwards the body unchanged, so Jina's optional parameters, such as `top_n`, `return_documents`, `max_doc_length`, and `return_embeddings`, reach the upstream as written. The response keeps Jina's rerank shape: a `results` array ordered by `relevance_score`, with each entry carrying the candidate's `index` and, when requested, the document or document embedding. Jina names a model in its rerank response, and AISIX rewrites that field back to the alias the request addressed, so it reads `jina-rerank-prod` rather than `jina-reranker-v3.5`. AISIX reads `usage.total_tokens` as input tokens for rerank usage and cost accounting. ## Experimental Chat Completions[​](#experimental-chat-completions "Direct link to Experimental Chat Completions") Jina publishes an experimental OpenAI-compatible `/v1/chat/completions` endpoint on the same API root for the exact model ID `jina-ai/jina-vlm`. It accepts text and image input, but Jina describes the endpoint as testing-only and does not guarantee availability, scalability, or production readiness. Do not use it as a production dependency. To test it, create a separate model alias with `model_name: "jina-ai/jina-vlm"` on the existing provider key. Direct `/v1/chat/completions` requests reach Jina's native chat route. `/v1/responses` and `/v1/messages` use AISIX translation through chat completions, so they retain the bridge limitations described in [Responses](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) and [Anthropic Messages](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md). Jina DeepSearch is a different chat-shaped product on `https://deepsearch.jina.ai/v1`. Configure it with a separate provider key and model alias. For passthrough, create a separate passthrough route that targets the DeepSearch base and injects the DeepSearch provider key. ## Supply Cost Metadata[​](#supply-cost-metadata "Direct link to Supply Cost Metadata") Jina is not priced on the models.dev catalog. No automatic catalog price exists for these aliases, so `least_cost` routing and cost estimates use only pricing you supply. Set rates through [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md) in AISIX Cloud, or with the `cost` field on the model in the open-source gateway's `resources.yaml`; see [Cost Metadata](https://docs.api7.ai/ai-gateway/models/model-aliases.md#cost-metadata). Convert Jina's published rates to USD per 1,000 tokens. The normalized embeddings route records Jina's `usage.prompt_tokens`. If a Jina model returns only `usage.total_tokens`, AISIX uses that total for token-per-minute enforcement but records zero input tokens in its usage event. Rerank maps Jina's `usage.total_tokens` to input tokens. Passthrough records `prompt_tokens` or `input_tokens` and `completion_tokens` or `output_tokens` when present, but does not map Jina's `total_tokens`-only response into input usage. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") A Jina provider key resolves the `openai` adapter for the OpenAI-shaped routes, and the rerank route dispatches on the `jina` provider value directly: | Route | Behavior with a `jina` model alias | | ---------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/embeddings` | Supported for string inputs and single dense float outputs by default. The modeled shape forwards `model`, string or string-array `input`, `encoding_format`, and `dimensions`, but Jina documents `embedding_type` rather than `encoding_format` for selecting base64, binary, or unsigned-binary output. Use a passthrough route as shown in [Send Jina-Specific Embedding Fields](#send-jina-specific-embedding-fields) for those encodings, native fields, multimodal inputs, and other output shapes. | | `/v1/rerank` | Supported with a Jina reranker alias. `jina` is one of the three provider values this route accepts. See [Add a Rerank Model](#add-a-rerank-model). | | `/v1/chat/completions` | Supported only by Jina's testing-only `jina-ai/jina-vlm` model on this root. See [Experimental Chat Completions](#experimental-chat-completions). Embedding and rerank aliases fail on the chat route. | | `/v1/responses` and `/v1/messages` | Supported through chat translation only for the experimental VLM alias. They are not native Jina routes and retain their bridge limitations. `/v1/messages/count_tokens` remains Anthropic-only. | | `/v1/completions`, `/v1/audio/*`, `/v1/files`, `/v1/batches`, and `/v1/fine_tuning/jobs` | Not compatible with Jina's APIs on this root. Jina audio and video support refers to multimodal v5 embedding inputs, not the normalized AISIX audio or video-generation endpoints. Jina's native batch embeddings use different paths and contracts. | | `/v1/images/generations` and `/v1/videos` | Not supported for the `jina` provider value. | | `/passthrough/jina/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md), relative to the route's target. Use it for native embedding fields, multimodal input, classifiers, training, or native batch paths such as `/batch/embeddings`. The route requires exact upstream model IDs and preserves the response body. Usage is recorded only when the detected request envelope and response fields use a supported token shape. | See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the full endpoint and provider matrix. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Jina, verified the embedding alias, and added a rerank route. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for these aliases. * [Rerank](https://docs.api7.ai/ai-gateway/endpoints/rerank.md): review the rerank request contract and its provider requirement. * [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md): review the modeled embeddings request shape and provider behavior. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Meta Llama API The Meta Llama API provides direct hosted access to Llama models with API-key authentication. AISIX gives applications a stable caller-facing API while storing the upstream credential, controlling model access, and recording usage. This page covers the Llama API on `api.llama.com`. It does not cover Meta's separate Model API on `api.meta.ai`. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Meta Llama API key and access to the model you plan to use. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, a model alias, and a caller API key. The Llama API is a community catalog provider with an OpenAI-compatible endpoint. AISIX connects through the `openai` adapter and authenticates upstream requests with a bearer token. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Export the upstream credential: ``` export LLAMA_API_KEY="YOUR_LLAMA_API_KEY" ``` Create the provider key: ``` PROVIDER_KEY_ID=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "llama-api-prod", "provider": "llama", "api_key": "'"${LLAMA_API_KEY}"'", "api_base": "https://api.llama.com/compat/v1/", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id' ) echo "$PROVIDER_KEY_ID" ``` `provider` must be the catalog ID `llama`. Do not add `adapter`: AISIX derives the `openai` adapter for catalog providers and accepts an explicit adapter only for BYO provider keys. The API base includes `/compat/v1/`. AISIX normalizes the trailing slash and appends `/chat/completions`, producing the upstream route `https://api.llama.com/compat/v1/chat/completions`. Do not replace this base with `https://api.llama.com/v1`. That is Meta's native Llama API wire, whose chat response and streaming event shapes differ from OpenAI chat completions. The AISIX `openai` adapter requires the `/compat/v1` surface. ### Create a Model[​](#create-a-model "Direct link to Create a Model") List the models available to the upstream account: ``` curl -sS "https://api.llama.com/compat/v1/models" \ -H "Authorization: Bearer ${LLAMA_API_KEY}" \ | jq -r '.data[].id' ``` Create a caller-facing alias for one of the returned model IDs. This example uses the Llama 4 Maverick model from Meta's current [official Llama API client example](https://github.com/meta-llama/llama-api-python/blob/main/examples/chat.py): ``` MODEL_ID=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "llama-api-prod", "model_name": "Llama-4-Maverick-17B-128E-Instruct-FP8", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -er '.model.id' ) echo "$MODEL_ID" ``` `display_name` is the stable alias applications send to AISIX. `model_name` is the exact, case-sensitive ID sent to Meta. Use an ID returned by the account's `/models` request because availability can change by account and API release. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create a caller key limited to this model: ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "llama-api-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` The plaintext caller key is returned only when the resource is created. Store it securely. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export LLAMA_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "llama-api-prod" provider: "llama" adapter: "openai" api_key: ${LLAMA_API_KEY} api_base: "https://api.llama.com/compat/v1/" models: - display_name: "llama-api-prod" provider: "llama" model_name: "Llama-4-Maverick-17B-128E-Instruct-FP8" provider_key: "llama-api-prod" api_keys: - display_name: "llama-api-caller" key_env: CALLER_API_KEY allowed_models: - "llama-api-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat request through AISIX: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "llama-api-prod", "messages": [ { "role": "user", "content": "Say hello from the Meta Llama API." } ] }' ``` AISIX resolves `llama-api-prod` to `Llama-4-Maverick-17B-128E-Instruct-FP8`, sends a bearer-authenticated request to the Llama API compatibility endpoint, and returns an OpenAI-compatible response. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") The configured API root is Meta's OpenAI compatibility surface. AISIX route behavior is as follows: | Route | Behavior with a `llama` model alias | | ----------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported through Meta's `/compat/v1/chat/completions` route. Llama 4 models can accept image content blocks in chat; this is vision input, not image generation. | | `/v1/responses` | Supported through the AISIX cross-provider bridge, which translates the request to chat completions and converts the result back to the Responses shape. AISIX does not forward this request to a Meta Responses route. OpenAI-specific fields without a chat equivalent are ignored. | | `/v1/messages` | Supported through the AISIX Anthropic-to-chat translation. `/v1/messages/count_tokens` remains limited to Anthropic-backed targets. | | `/v1/embeddings` and `/v1/completions` | Not supported. Meta's `/compat/v1` root does not publish these routes. | | `/v1/audio/*`, `/v1/images/generations`, `/v1/videos`, and `/v1/rerank` | Not supported. The route-specific provider gates or Meta endpoint coverage do not admit this configuration. | | `/v1/batches` | Not supported. Meta's compatibility root does not publish this route. | | `/v1/files` and `/v1/fine_tuning/jobs` | Unverified. Meta's compatibility root exposes matching routes, but this AISIX workflow has not been validated end to end, and the file-content route is absent upstream. Do not rely on it without testing the required operations. | | `/passthrough/llama/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) targeting the `/compat/v1` base. Use it for compatible Meta routes that AISIX does not model, such as `/moderations`. Passthrough preserves the request and response bodies. Recognized chat, completions, and Responses envelopes record supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. | Passthrough does not rewrite a `model` field, so send an exact upstream model ID when the native operation requires one. The `/passthrough/llama` prefix assumes a passthrough route claiming it with the `/compat/v1` base as its `target_url` and this provider key for credential injection; grant the route on the caller key's `allowed_routes`. If you add a second `llama` key for Meta's native `/v1` resources, create a separate route under its own prefix bound to that key. Do not route normalized chat through the native key because its response wire is not OpenAI-compatible. See [Responses](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), [Anthropic Messages](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md), and [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the translation and endpoint details. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Check | | ----------------------------------- | ----------------------------------------------------------------------------------------------------------------- | | Upstream `401` or `403` | Confirm `LLAMA_API_KEY` is active and the account can use the selected model. | | Upstream `404` | Confirm the API base includes `/compat/v1` and copy a current model ID from the authenticated `/models` response. | | Provider key creation returns `400` | Use `provider: "llama"` without an `adapter` field. | | AISIX returns model access denied | Confirm the caller key's `allowed_models` contains the model resource ID. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to the Meta Llama API and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md): rotate the upstream credential or configure provider-specific overrides. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # MiniMax [MiniMax](https://platform.minimax.io/docs/guides/quickstart) provides hosted generative models through its platform API. Applications call MiniMax through stable AISIX aliases while the gateway keeps the upstream credential out of client code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A MiniMax API key from the [MiniMax platform](https://platform.minimax.io/). * `curl` and `jq`. ## Set the Correct Base URL[​](#set-the-correct-base-url "Direct link to Set the Correct Base URL") MiniMax publishes two upstream surfaces for the same models: | Surface | Root | Request format | | ---------------------------------------------------------------------------------- | ---------------------------------- | ----------------------- | | [OpenAI SDK](https://platform.minimax.io/docs/api-reference/text-openai-api) | `https://api.minimax.io/v1` | OpenAI chat completions | | [Anthropic SDK](https://platform.minimax.io/docs/api-reference/text-anthropic-api) | `https://api.minimax.io/anthropic` | Anthropic Messages | The community catalog entry for `minimax` publishes `https://api.minimax.io/anthropic/v1`, the Anthropic-compatible surface written out with the version segment that the Anthropic SDK would otherwise append, while the catalog default assigns the `openai` adapter. Those two do not fit together. If you omit `api_base`, AISIX accepts the provider key, fills in that catalog value, and then sends OpenAI-shaped chat-completions requests to an endpoint that expects Anthropic-shaped Messages requests, and every call through the alias fails at the upstream. caution Always set `api_base` to `https://api.minimax.io/v1` on a `minimax` provider key. The gateway appends the endpoint path, such as `/chat/completions`, to whatever root you configure, so the root must be the one that MiniMax's OpenAI-compatible `/chat/completions` hangs off. To use the Anthropic-compatible surface instead, see [Use the Anthropic-Compatible Surface](#use-the-anthropic-compatible-surface). ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the MiniMax-backed chat-completions route. MiniMax is a community catalog provider. AISIX connects through the `openai` adapter and authenticates upstream requests with a bearer token. You must also set `api_base` explicitly because the catalog's published base URL is not the root required by the adapter. See [Set the Correct Base URL](#set-the-correct-base-url). ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the MiniMax credential and API root: ``` # Replace with your value export MINIMAX_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "minimax-prod", "provider": "minimax", "api_key": "'"${MINIMAX_API_KEY}"'", "api_base": "https://api.minimax.io/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `minimax`. The AISIX Cloud Admin API accepts the value because `minimax` is a models.dev catalog entry, and it derives the adapter from the catalog; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the MiniMax API key. MiniMax authenticates its OpenAI-compatible surface with HTTP bearer authentication, which is what the `openai` adapter already sends. The value follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is `https://api.minimax.io/v1`, the `base_url` that MiniMax documents for OpenAI client libraries. It is the single most important field on this page. Omitting it does not fail the create call, because the catalog publishes a default; it fails every later chat request, because that default points at the Anthropic-compatible root. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") MiniMax model IDs are case-sensitive and carry no vendor prefix. They follow the pattern `MiniMax-<generation>`, with a capital `M` on both halves of the vendor name, a decimal generation number, and an optional `-highspeed` suffix on the low-latency variant of a generation. Do not lowercase them, and do not add a `minimax/` prefix carried over from an aggregator. Current model IDs include the following: | Model ID | Notes | | ------------------------ | -------------------------------------------------------------------------------------------------------- | | `MiniMax-M3` | Latest generation. Accepts text, image, and video input on the OpenAI-compatible chat-completions route. | | `MiniMax-M2.7` | Previous flagship generation, text input only. | | `MiniMax-M2.7-highspeed` | Low-latency variant of the same generation. | The `MiniMax-M2.5`, `MiniMax-M2.5-highspeed`, `MiniMax-M2.1`, and `MiniMax-M2` IDs remain in the catalog as earlier generations. Check the [MiniMax model documentation](https://platform.minimax.io/docs/api-reference/api-overview) for the current list before you create an alias. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "minimax-m3-prod", "model_name": "MiniMax-M3", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the MiniMax model ID, spelled exactly as MiniMax publishes it, for example `MiniMax-M3`. ❸ `provider_key_id` attaches the alias to the MiniMax provider key. To understand model cost metadata and AISIX Cloud pricing, see [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md#cost-metadata). ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "minimax-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export MINIMAX_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "minimax-prod" provider: "minimax" adapter: "openai" api_key: ${MINIMAX_API_KEY} api_base: "https://api.minimax.io/v1" models: - display_name: "minimax-m3-prod" provider: "minimax" model_name: "MiniMax-M3" provider_key: "minimax-prod" api_keys: - display_name: "minimax-caller" key_env: CALLER_API_KEY allowed_models: - "minimax-m3-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "minimax-m3-prod", "messages": [ { "role": "user", "content": "Say hello from MiniMax." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `minimax-m3-prod`. If the request fails, check these in order: 1. `api_base` on the provider key. An upstream 404, or an upstream error that mentions an unexpected request field, usually means the key is still pointing at `https://api.minimax.io/anthropic/v1` rather than `https://api.minimax.io/v1`. 2. `api_key` on the provider key, if the upstream returns an authentication error. 3. The spelling and capitalization of `model_name`. MiniMax model IDs are case-sensitive. ## Understand the Community Catalog Path[​](#understand-the-community-catalog-path "Direct link to Understand the Community Catalog Path") A featured provider carries an AISIX-curated adapter, authentication scheme, default base URL, and any request or response rewrites the upstream needs. `minimax` has none of those. The dashboard groups it under the community catalog, and AISIX makes exactly three decisions on its behalf: * The adapter is `openai`. * The authentication scheme is HTTP bearer. * The default base URL is whatever models.dev publishes for the provider, cached in AISIX provider metadata. Everything else is yours to supply. In practice this means: * **The base URL is your responsibility.** On this provider the published default is wrong for the assigned adapter, so set it explicitly, as described in [Set the Correct Base URL](#set-the-correct-base-url). AISIX rejects a create call with a 400 error only when models.dev publishes no base URL at all for a provider; `minimax` publishes one, so the request succeeds and the mismatch surfaces later. * **No request or response rewrites are registered.** AISIX forwards the caller's chat-completions body to MiniMax with no parameter renames, and normalizes reasoning from the canonical `delta.reasoning_content` path. It does not preserve MiniMax's separate `reasoning_details` array. See [Configure Request and Response Overrides](#configure-request-and-response-overrides). * **The upstream contract is not tracked for you.** When MiniMax changes its wire shape, no catalog entry changes with it. Verify a representative request against a non-production alias after a MiniMax API update. Usage records still tag the provider key as a catalog key branded `minimax`, and mark it as not featured, so community-catalog traffic stays distinguishable in usage data. None of this makes `minimax` a lesser upstream. It is a supported catalog provider with the same caller keys, allowlists, rate limits, and usage accounting as any other alias. In AISIX Cloud, matching budgets also apply. The difference is where the responsibility for the wire contract sits. ## Configure Request and Response Overrides[​](#configure-request-and-response-overrides "Direct link to Configure Request and Response Overrides") Because `minimax` has no AISIX-curated adapter mapping, the provider key is the only place to record a MiniMax-specific wire difference. AISIX sends top-level request fields unchanged unless you configure `request.param_renames`. MiniMax-M3 currently accepts both `max_tokens` and `max_completion_tokens`, so no rename is required for the model in this guide. Add one only when the selected model requires a different field name. MiniMax-M3 defaults `reasoning_split` to `false`, which keeps thinking inside `<think>` tags in `content`; AISIX preserves that content. If you set `reasoning_split` to `true`, MiniMax also emits `reasoning_content` and a `reasoning_details` array. On `/v1/chat/completions`, AISIX preserves `reasoning_content` but drops `reasoning_details`; the Responses bridge does not expose either field. The `response.reasoning_field` override cannot retain the array. MiniMax requires the complete array in conversation history for interleaved thinking with tools, so leave split reasoning disabled for multi-turn tool loops through the normalized chat or Responses endpoints. If your application depends on the complete Anthropic thinking and tool-use contract from MiniMax, use the [Anthropic-compatible surface](#use-the-anthropic-compatible-surface). For a different model that streams scalar reasoning from a nonstandard `delta` path, set `response.reasoning_field` on the provider key to that path. AISIX then maps it to `delta.reasoning_content` for callers. Confirm any override with a streaming request against the configured model. Overrides apply to every model that references the provider key, so test them against a non-production alias first. See [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides). ## Connect the China Platform[​](#connect-the-china-platform "Direct link to Connect the China Platform") MiniMax operates a separate platform for mainland China, with its own developer console and hostname. It appears in the catalog as a distinct provider ID, `minimax-cn`, and each platform issues its own API keys, so the China platform needs its own provider key. The same base-URL correction applies: the catalog default points at the Anthropic-compatible root, and the OpenAI-compatible root is `https://api.minimaxi.com/v1`. Note the extra `i` in the hostname. ``` { "display_name": "minimax-cn-prod", "provider": "minimax-cn", "api_key": "YOUR_MINIMAX_CN_API_KEY", "api_base": "https://api.minimaxi.com/v1", "allowed_environments": ["YOUR_ENVIRONMENT_ID"] } ``` The catalog lists the same model IDs for both platforms. Availability of a given generation is not guaranteed to match, so verify against the [China platform documentation](https://platform.minimaxi.com/docs/guides/quickstart) rather than assuming parity. ## Use the Anthropic-Compatible Surface[​](#use-the-anthropic-compatible-surface "Direct link to Use the Anthropic-Compatible Surface") Anthropic-shaped clients do not need the MiniMax Anthropic endpoint. AISIX already accepts Anthropic Messages requests at `/v1/messages` and translates them onto the OpenAI-compatible alias you configured above. See [Anthropic Messages API](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md). Route to MiniMax's own Anthropic surface only when you need the upstream to receive the Anthropic wire shape unmodified. The `anthropic` adapter is not available on a catalog provider key, because AISIX derives the adapter from the catalog, so use a [bring-your-own-endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md) key instead: ``` { "display_name": "minimax-anthropic-prod", "provider": "byo", "adapter": "anthropic", "api_key": "YOUR_MINIMAX_API_KEY", "api_base": "https://api.minimax.io/anthropic", "allowed_environments": ["YOUR_ENVIRONMENT_ID"] } ``` The `anthropic` adapter appends `/v1/messages` to the configured root, so `https://api.minimax.io/anthropic` is the correct value here. A BYO key requires a non-empty `api_key` and `api_base`; its adapter is immutable after creation. A BYO provider key is attributed as a bring-your-own upstream in usage records and carries no branded provider value, so its usage does not aggregate with the `minimax` catalog key's usage under one provider. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") This alias uses the OpenAI-compatible text and multimodal chat surface from MiniMax. MiniMax also provides native image, video, speech, music, and file APIs, but their paths and bodies do not match the normalized media and job contracts in AISIX. | Route | Behavior with a MiniMax alias | | ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including `stream: true`. MiniMax-M3 accepts image and video inputs here for multimodal understanding; this route does not generate images or videos. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). AISIX translates the request to chat completions; it does not call a MiniMax `/v1/responses` endpoint. OpenAI-specific Responses fields without a chat equivalent are ignored. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation. The catalog alias is OpenAI-backed, so `/v1/messages/count_tokens` returns a 400 error for it. The separate BYO Anthropic alias supports the MiniMax Messages surface and token counting. | | `/v1/embeddings` | Do not assume support. AISIX forwards an OpenAI-shaped embeddings body to `{api_base}/embeddings` without translation, but MiniMax's current public API documentation does not define an OpenAI-compatible embeddings contract or embedding model. See [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md). | | `/v1/images/generations` | Rejected with a 400 error. MiniMax image generation uses the native `/v1/image_generation` contract, which you can access through a passthrough route. | | `/v1/rerank` | Rejected with a 400 error. The route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/v1/videos` | Returns 501 Not Implemented. MiniMax video generation uses the native `/v1/video_generation` contract, which you can access through a passthrough route. | | `/passthrough/minimax/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for native routes such as `/image_generation`, `/video_generation`, `/t2a_v2`, and `/files/retrieve`. AISIX forwards the native body and response without translating them. | The `/passthrough/minimax` paths on this page assume a passthrough route claiming that prefix, with MiniMax's API root as its `target_url` and this provider key attached; grant the route name on the caller key's `allowed_routes`. A route relays to one fixed target, so if you configure more than one MiniMax account or base URL, create a separate route per account. AISIX does not rewrite a `model` value in the native body. It detects chat, completions, and Responses envelopes and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. With a route targeting the base in this guide, passthrough requests resolve against `https://api.minimax.io/v1`. That root already ends in a version segment, and the gateway strips one redundant leading version segment from the passthrough path when it matches. Both `/passthrough/minimax/v1/image_generation` and `/passthrough/minimax/image_generation` therefore reach `https://api.minimax.io/v1/image_generation` rather than a doubled `/v1/v1/` URL. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to MiniMax and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between MiniMax and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Mistral AI [Mistral AI](https://docs.mistral.ai/) develops and hosts the Mistral family of models. AISIX lets applications call its chat and compatible API surfaces with gateway-issued caller keys. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Mistral API key from the [Mistral console](https://console.mistral.ai/api-keys). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Mistral-backed chat-completions route. Because Mistral exposes an OpenAI-compatible API, AISIX connects through the `openai` adapter and uses the Mistral API root as `api_base`. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Mistral credential and API root: ``` # Replace with your value export MISTRAL_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "mistral-prod", "provider": "mistral", "api_key": "'"${MISTRAL_API_KEY}"'", "api_base": "https://api.mistral.ai/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `mistral`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the adapter field is only accepted on BYO provider keys. ❷ `api_key` stores the Mistral API key. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` already includes the `/v1` path. AISIX appends `/chat/completions` to it. The field is optional for this catalog provider, because the AISIX Cloud Admin API fills in this same value when you omit it, but the examples set it explicitly so the upstream root stays visible on the resource. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Mistral model IDs ending in `-latest` track the newest snapshot of that model. To pin a specific version, use the dated model ID from the [Mistral models list](https://docs.mistral.ai/models/overview). Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "mistral-large-prod", "model_name": "mistral-large-latest", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Mistral model ID, for example `mistral-large-latest` or `mistral-small-latest`. If you use structured reasoning with `mistral-small-latest`, review [Understand Structured Responses](#understand-structured-responses) before enabling it. ❸ `provider_key_id` attaches the alias to the Mistral provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the response — store it securely: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "mistral-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. The gateway picks up the new resources automatically — no restart is needed. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export MISTRAL_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "mistral-prod" provider: "mistral" adapter: "openai" api_key: ${MISTRAL_API_KEY} api_base: "https://api.mistral.ai/v1" models: - display_name: "mistral-large-prod" provider: "mistral" model_name: "mistral-large-latest" provider_key: "mistral-prod" api_keys: - display_name: "mistral-caller" key_env: CALLER_API_KEY allowed_models: - "mistral-large-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "mistral-large-prod", "messages": [ { "role": "user", "content": "Say hello from Mistral." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `mistral-large-prod`. If the request fails, check the provider key `api_key`, `api_base`, and the Mistral model ID in `model_name`. ## Understand Structured Responses[​](#understand-structured-responses "Direct link to Understand Structured Responses") Standard Mistral chat responses work through the normalized AISIX endpoints when `message.content`, or each streaming `delta.content`, is a string. Text and OpenAI-style image input blocks are forwarded to compatible multimodal models such as Mistral Large 3. Current Mistral reasoning models can return a different shape. With `reasoning_effort: "high"`, Mistral returns arrays of `ThinkChunk` and `TextChunk` objects in `content`. The AISIX OpenAI adapter currently expects string content and returns an upstream decode error for those responses. This applies to `/v1/chat/completions` and to the Responses and Messages bridges, which ultimately use the same chat adapter. A `response.reasoning_field` override cannot convert the content array. For a reasoning-capable model such as `mistral-small-latest`, use `reasoning_effort: "none"` on the normalized endpoints. To retain structured reasoning and replay the complete thinking history, call Mistral's native chat contract through `/passthrough/mistral/chat/completions` and send the upstream Mistral model ID rather than the AISIX alias. The `/passthrough/mistral` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with the Mistral API root as its `target_url`; grant the route name on the caller key's `allowed_routes`. The route preserves the response body and relays a native streaming response incrementally. Mistral also supports `n` greater than `1`, but AISIX returns only the first choice on normalized chat routes. Use a passthrough route when the application needs every choice or another provider-native response shape. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") The model alias in this guide is a chat model. Create a separate alias with the appropriate Mistral model ID for embeddings, transcription, speech, or another capability. | Route | Behavior with a Mistral alias | | -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported for scalar-content responses, including `stream: true`, function tools, and text or image input supported by the selected model. Structured reasoning content arrays are not supported. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), which translates to chat completions rather than calling a native Mistral Responses API. Fields without a chat equivalent are ignored, and the structured reasoning limitation still applies. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation to chat completions. The alias is OpenAI-backed, so `/v1/messages/count_tokens` returns a 400 error. | | `/v1/embeddings` | Supported with a dedicated alias such as `mistral-embed`. AISIX forwards `dimensions`, but Mistral names the reduced-size field `output_dimension`; configure a request parameter rename if you need it. | | `/v1/audio/transcriptions` | Supported with a transcription alias such as `voxtral-mini-latest`. AISIX relays the Mistral response, including provider-specific fields, but buffers streaming responses rather than relaying their events incrementally. | | `/v1/audio/translations` | Not supported because Mistral does not publish this route. | | `/v1/audio/speech` | The path reaches Mistral TTS, but the wire contract is not OpenAI-compatible. Send native fields such as `voice_id` and expect base64 JSON, or a buffered Mistral SSE payload, rather than raw audio bytes. AISIX does not translate this contract. | | `/v1/files` | Supported for upload, list, retrieve, delete, and content download. Mistral's signed-URL route at `/v1/files/{id}/url` is reachable through a passthrough route, but it requires a raw Mistral file ID. A passthrough route does not decode the AISIX-routed ID returned by a normalized upload. | | `/v1/batches` | Not supported. AISIX forwards the OpenAI `/v1/batches` path, while Mistral uses `/v1/batch/jobs`. | | `/v1/fine_tuning/jobs` | Not supported. AISIX forwards the OpenAI job surface and requires `training_file`; Mistral's current public API does not publish compatible fine-tuning job routes. | | `/v1/completions` | Not supported. Mistral code completion uses the native `/v1/fim/completions` contract. | | `/v1/images/generations` | Rejected with a 400 error because the route accepts only `openai` provider models. Mistral image generation is a built-in tool rather than this OpenAI image route. | | `/v1/rerank` | Rejected with a 400 error because the route does not accept the `mistral` provider. Mistral publishes classification APIs, not this rerank contract. | | `/passthrough/mistral/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for native Mistral routes, with raw request and response bodies. Provider SSE responses relay incrementally. | Passthrough routes are the way to reach Mistral-native OCR, moderation, classification, FIM, batch, Agents, Conversations, and structured reasoning. For example, with a route targeting `https://api.mistral.ai/v1`, `/passthrough/mistral/ocr` and `/passthrough/mistral/v1/ocr` both resolve to `https://api.mistral.ai/v1/ocr`, because AISIX removes one duplicated version segment. AISIX does not rewrite a `model` value in a passthrough body. The route supplies the upstream base URL and, in inject mode, the provider key credential; a route binds one target and credential, so configure separate routes for separate Mistral accounts or bases. AISIX detects chat, completions, and Responses envelopes from each request and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Mistral and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Mistral and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # ModelScope [ModelScope](https://modelscope.cn/docs/model-service/API-Inference/intro) is a model community and hosted inference platform for open-weight models. AISIX gives applications one OpenAI-compatible API for those models while managing credentials, caller access, rate limits, and usage accounting. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A ModelScope API token from your ModelScope account. The ModelScope [API-Inference documentation](https://modelscope.cn/docs/model-service/API-Inference/intro) describes the account activation and quota rules that apply before a token can serve inference requests. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the ModelScope-backed chat-completions route. ModelScope is a community catalog provider with an OpenAI-compatible API-Inference service. AISIX connects through the `openai` adapter, authenticates upstream requests with a bearer token, and uses the ModelScope API root as `api_base`. Review [Supply the Wire Details AISIX Does Not Curate](#supply-the-wire-details-aisix-does-not-curate) before sending production traffic. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the ModelScope credential and API root: ``` # Replace with your value export MODELSCOPE_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "modelscope-prod", "provider": "modelscope", "api_key": "'"${MODELSCOPE_API_KEY}"'", "api_base": "https://api-inference.modelscope.cn/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `modelscope`. The AISIX Cloud Admin API accepts this value because the ID exists in the cached models.dev catalog, and it applies the catalog's default rule: the `openai` adapter and HTTP bearer authentication. Do not set the `adapter` field, which the AISIX Cloud Admin API accepts only on BYO provider keys. ❷ `api_key` stores the ModelScope API token. The `openai` adapter sends it in an `Authorization: Bearer` header. The value follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is `https://api-inference.modelscope.cn/v1`, the API root that ModelScope's OpenAI-compatible chat-completions endpoint, `https://api-inference.modelscope.cn/v1/chat/completions`, hangs off. AISIX appends the endpoint path, such as `/chat/completions`, to the value you store. For `modelscope` this field is optional, because the control plane fills in this same root from the catalog when you omit it. The examples set it explicitly so the root each key targets stays visible in the configuration. caution Store the API root, not the full endpoint URL and not the bare host. AISIX strips a trailing `/chat/completions` if you paste one, but it never synthesizes the `/v1` segment for a non-OpenAI host. An `api_base` of `https://api-inference.modelscope.cn` produces upstream requests to `https://api-inference.modelscope.cn/chat/completions`, a path ModelScope does not serve. These examples target the ModelScope China service. ModelScope's international service uses `https://api-inference.modelscope.ai/v1`; set that root explicitly when your token and model come from `modelscope.ai`. The catalog default remains the `.cn` root, so omitting `api_base` does not select the international service automatically. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") ModelScope model IDs are organization-namespaced. Each ID repeats the `organization/model` path of the model page on ModelScope, including its capitalization. Current IDs in the AISIX provider catalog include the following: | Model ID | Context window | Notes | | ------------------------------------ | -------------- | ----------------------------------------------------- | | `Qwen/Qwen3-235B-A22B-Instruct-2507` | 262,144 tokens | Instruction model. Does not emit reasoning. | | `Qwen/Qwen3-235B-A22B-Thinking-2507` | 262,144 tokens | Reasoning variant of the same weights. | | `Qwen/Qwen3-Coder-30B-A3B-Instruct` | 262,144 tokens | Coding and software-agent model. | | `ZhipuAI/GLM-4.6` | 202,752 tokens | GLM flagship served under the `ZhipuAI` organization. | Confirm the ID on the model's ModelScope page before you create an alias, because the set of models that API-Inference serves changes over time. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "modelscope-qwen-prod", "model_name": "Qwen/Qwen3-235B-A22B-Instruct-2507", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the full ModelScope model ID, organization prefix included. Dropping the prefix, or reusing an unprefixed ID such as `qwen3-235b-a22b-instruct-2507` from another host of the same weights, produces an upstream model-not-found error. ❸ `provider_key_id` attaches the alias to the ModelScope provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "modelscope-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export MODELSCOPE_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "modelscope-prod" provider: "modelscope" adapter: "openai" api_key: ${MODELSCOPE_API_KEY} api_base: "https://api-inference.modelscope.cn/v1" models: - display_name: "modelscope-qwen-prod" provider: "modelscope" model_name: "Qwen/Qwen3-235B-A22B-Instruct-2507" provider_key: "modelscope-prod" api_keys: - display_name: "modelscope-caller" key_env: CALLER_API_KEY allowed_models: - "modelscope-qwen-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "modelscope-qwen-prod", "messages": [ { "role": "user", "content": "Say hello from ModelScope." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `modelscope-qwen-prod`. If the request fails, check the provider key `api_key`, the `api_base` root, and the organization-prefixed model ID in `model_name`. ## Choose Between ModelScope and Alibaba Cloud Model Studio[​](#choose-between-modelscope-and-alibaba-cloud-model-studio "Direct link to Choose Between ModelScope and Alibaba Cloud Model Studio") ModelScope and Alibaba Cloud Model Studio both serve Qwen models, but they are separate upstreams with separate provider IDs in AISIX. Selecting the wrong one produces an upstream authentication or model-not-found error rather than a configuration error at create time, so confirm which service issued your credential first. | | ModelScope | Alibaba Cloud Model Studio | | -------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------- | | Provider ID | `modelscope` | `alibaba` | | Catalog tier | Community | Curated | | Setup guide | This page | [Qwen (Alibaba Cloud)](https://docs.api7.ai/ai-gateway/providers/qwen.md) | | API root | `https://api-inference.modelscope.cn/v1` | Region-specific DashScope OpenAI-compatible root | | Credential | ModelScope API token | DashScope API key, issued per region | | Model ID form | Organization-namespaced repository path, such as `Qwen/Qwen3-235B-A22B-Instruct-2507` | Model Studio model name, such as `qwen-plus` | | Modeled `/v1/videos` route | Rejected | Supported | The same distinction applies to GLM. ModelScope serves GLM weights under the `ZhipuAI` organization, but a GLM alias backed by the `modelscope` provider is not equivalent to one backed by the curated `zhipuai` provider described in [Zhipu AI (GLM)](https://docs.api7.ai/ai-gateway/providers/zhipuai.md). Gateway surfaces that dispatch by provider label, such as the modeled [video-generation route](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md), key off the provider recorded on the model, not off which weights the upstream happens to serve. ## Supply the Wire Details AISIX Does Not Curate[​](#supply-the-wire-details-aisix-does-not-curate "Direct link to Supply the Wire Details AISIX Does Not Curate") For a provider with an AISIX-curated adapter mapping, the catalog records the adapter, authentication scheme, default API root, and any request or response quirks. ModelScope has no such mapping. It resolves through the catalog's default rule, which grants the `openai` adapter and bearer authentication and marks the provider as community-sourced, with wire compatibility assumed to be OpenAI-shaped unless the operator overrides it per key. Everything beyond that assumption is yours to configure. In practice: * **No parameter renames are registered.** AISIX forwards the top-level chat-completions parameters exactly as the caller sent them. Some curated providers carry a rename for a field such as `max_completion_tokens`; ModelScope carries none, so a body that works against a renamed provider is not automatically portable here, and the reverse is also true. * **No response field remapping is registered.** AISIX already normalizes reasoning that arrives at `message.reasoning_content`, `delta.reasoning_content`, or `message.reasoning` into the canonical `reasoning_content` field. Reasoning delivered on any other streaming `delta` path is not remapped unless you map it. * **No provider-wide reasoning control applies.** ModelScope does not publish a shared reasoning toggle, effort level, or thinking-token budget that spans the models it serves, so reasoning behavior follows the individual model. Validate it per model ID: `Qwen/Qwen3-235B-A22B-Thinking-2507` reasons by design, while `Qwen/Qwen3-235B-A22B-Instruct-2507` does not. Both kinds of adjustment are provider-key fields. Add them to the provider key create request, or to the provider key entry in a declarative resources file: ``` { "request": { "param_renames": { "max_completion_tokens": "max_tokens" } }, "response": { "reasoning_field": "delta.thinking" } } ``` Use `request.param_renames` only when a specific ModelScope-served model rejects a parameter name your clients already send, and `response.reasoning_field` only when a streaming response carries reasoning on a path other than `delta.reasoning_content`. See [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) for the field catalog and the adapters each override applies to. An override on the provider key applies to every model that references it, so validate a change against a non-production alias first. For the same reason, consider one provider key per organization namespace when the models you serve behave differently: `Qwen/*` and `ZhipuAI/*` share one endpoint but are different model families. Send both a buffered and a streaming request through each new alias, and confirm the fields your clients read are present before you route traffic to it. ## Account for Cost and Shared Quota[​](#account-for-cost-and-shared-quota "Direct link to Account for Cost and Shared Quota") For normalized model requests, the managed control plane resolves per-request cost by the exact `(provider, model name)` pair, falling back to the models.dev catalog default when no organization override matches. The ModelScope catalog entries publish an input and output price of zero, so these usage events record their token counts and resolve to `$0.00` spend. Budgets evaluated against that spend never advance, and a `least_cost` routing group treats the alias as free. If cost reporting or budgets must reflect what ModelScope traffic actually costs your organization, add an organization pricing override for provider `modelscope` and the model name exactly as configured, organization prefix included. See [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md). One upstream ModelScope account backs every caller behind the gateway, so an upstream quota rejection reaches every caller at once. Set gateway-side caps so a single caller cannot exhaust the shared account: attach a rate limit to the model alias, to the caller API key, or to both. See [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md). ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") ModelScope API-Inference exposes several OpenAI- and Anthropic-shaped endpoints, but only part of the normalized AISIX proxy surface applies to a ModelScope-backed alias. A supported route below describes what AISIX does with the request; it does not imply that AISIX uses ModelScope's native endpoint of the same name. | Route | Behavior with a ModelScope alias | | --------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including `stream: true`. | | `/v1/completions` | Not supported. ModelScope API-Inference does not expose the legacy completions path. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over the chat adapter path. It does not call ModelScope's native `/v1/responses` endpoint, and Responses fields without a chat equivalent are ignored. Use a passthrough route when native Responses semantics are required. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation to chat completions. It does not call ModelScope's native `/v1/messages` endpoint. | | `/v1/messages/count_tokens` | Rejected for the `modelscope` configuration in this guide because the normalized counter requires an Anthropic-backed model. ModelScope's native counter is reachable through a passthrough route. | | `/v1/embeddings` | Dispatches through the `openai` adapter, so it works only when the configured model is served at the upstream `/v1/embeddings` path. The ModelScope entries in the AISIX provider catalog are text chat models. See [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md). | | `/v1/audio/*`, `/v1/files`, `/v1/batches`, and `/v1/fine_tuning/jobs` | Not available with this provider. ModelScope API-Inference does not expose compatible paths beneath the configured API root. | | `/v1/images/generations` | Rejected with a 400 error. The route accepts only models whose provider is `openai`. ModelScope's native asynchronous image-generation workflow is reachable through a passthrough route. | | `/v1/rerank` | Rejected with a 400 error. The route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/v1/videos` | Rejected with `501 not_implemented`. The route dispatches by provider label, and `modelscope` is not one of the labels it accepts, including for GLM weights that the `zhipuai` provider can drive on this route. | | `/passthrough/modelscope/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for provider-native routes, including `/responses`, `/messages`, `/messages/count_tokens`, `/images/generations`, and `/tasks/{task_id}`. | ModelScope's native API surface is not the same as the normalized AISIX surface. For example, ModelScope image generation starts an asynchronous job at `/v1/images/generations` and returns a task ID. The application then polls `/v1/tasks/{task_id}`. With a passthrough route targeting this guide's API root, call those routes at `/passthrough/modelscope/images/generations` and `/passthrough/modelscope/tasks/{task_id}`. Send `X-ModelScope-Async-Mode: true` when starting the job, `X-ModelScope-Task-Type: image_generation` when polling it, and the exact ModelScope image model ID in the native request body. These paths assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming the `/passthrough/modelscope` prefix with `https://api-inference.modelscope.cn/v1` as its `target_url` and the ModelScope provider key for credential injection; grant the route name on the caller key's `allowed_routes`. A duplicated leading `/v1` is removed when it matches the target's trailing segment, so adding it after `/passthrough/modelscope` reaches the same upstream URL. The gateway does not rewrite an AISIX alias in the request body, and the native `model` field does not select a credential — the route binds one target and provider key, so use separate routes for different ModelScope accounts or roots. AISIX detects chat, completions, and Responses envelopes and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. AISIX relays provider SSE incrementally rather than buffering events. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to ModelScope and verified the model alias. Continue with these guides: * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when this upstream differs from the `openai` adapter's assumptions. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over from ModelScope to a second provider that serves comparable weights. * [Qwen (Alibaba Cloud)](https://docs.api7.ai/ai-gateway/providers/qwen.md): configure the curated Alibaba Cloud Model Studio upstream instead when your credential comes from Model Studio. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Moonshot AI (Kimi) [Moonshot AI](https://platform.moonshot.ai/) provides the Kimi family of models through a hosted API. Applications call Kimi through stable AISIX aliases while the gateway keeps the Moonshot credential out of client code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Moonshot API key from the [Kimi API Platform](https://platform.moonshot.ai/). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Moonshot-backed chat-completions route. Because Moonshot AI exposes an OpenAI-compatible API, AISIX connects through the `openai` adapter and uses the API root for the region where you created the credential. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Moonshot AI serves its API from two independent hosts, and the catalog models them as two separate provider IDs. Pick the pair that matches the console where you created the API key: | Provider ID | API root | Use when | | --------------- | ---------------------------- | ------------------------------------------ | | `moonshotai` | `https://api.moonshot.ai/v1` | The key was issued on the global platform. | | `moonshotai-cn` | `https://api.moonshot.cn/v1` | The key was issued on the China platform. | The two hosts are separate deployments with separate consoles, so a key issued on one platform does not authenticate against the other. Select the provider ID that matches the key, rather than pointing one ID at the other platform's host: the provider ID is what usage records and cost reports attribute traffic to. For new provider keys, `api_base` is optional on both IDs. The AISIX Cloud Admin API fills the global root for `moonshotai` and the China root for `moonshotai-cn`. The examples set it explicitly so the target platform remains visible. An older `moonshotai` key created before the regional defaults were corrected can remain pinned to the China root. Set the global root explicitly when updating that configuration for a global-platform credential. The examples below use `moonshotai`. If your account is on the China platform, substitute `moonshotai-cn` and `https://api.moonshot.cn/v1` throughout. Create the provider key that stores the Moonshot credential and API root, and capture its ID: ``` # Replace with your value export MOONSHOT_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "moonshot-prod", "provider": "moonshotai", "api_key": "'"${MOONSHOT_API_KEY}"'", "api_base": "https://api.moonshot.ai/v1", "apis": { "responses": {} }, "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `moonshotai`, not `moonshot` or `kimi`. The AISIX Cloud Admin API only accepts catalog provider IDs and rejects other spellings with `400 INVALID_REQUEST`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the Moonshot API key and is sent as a bearer token. The value is encrypted before storage and never returned by read endpoints. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` already includes the `/v1` path segment, because Moonshot AI publishes its OpenAI-compatible surface under `/v1` rather than at the host root. AISIX appends the endpoint path to it, so use `https://api.moonshot.ai/v1` without a trailing `/chat/completions`. To route to the China platform instead, create a separate provider key with `provider` set to `moonshotai-cn` and `api_base` set to `https://api.moonshot.cn/v1`. ❹ `apis.responses` tells AISIX to forward `/v1/responses` to Moonshot's native Responses API instead of translating the request through chat completions. Both regional API roots implement this route, which currently supports only `kimi-k3`. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Moonshot model IDs follow a `kimi-<generation>` pattern, with an optional suffix for a task-specific or throughput-specific variant. For example, `kimi-k2.6` is the general-purpose model, `kimi-k2.7-code` is the coding model, and `kimi-k2.7-code-highspeed` is its higher-throughput variant. `kimi-k3` is the current flagship. Moonshot AI retires older snapshots such as the earlier `kimi-k2-*-preview` IDs, so confirm the ID against the [Kimi model list](https://platform.kimi.ai/docs/models) before pinning it. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "kimi-k3-prod", "model_name": "kimi-k3", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. The alias is independent of the upstream ID, so it does not need to carry the generation number. ❷ `model_name` is the Moonshot model ID. This example uses `kimi-k3` because Moonshot's native Responses API currently supports only that model. Other model IDs include `kimi-k2.6` and `kimi-k2.7-code`; preserve dots and other punctuation exactly. ❸ `provider_key_id` attaches the alias to the Moonshot provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The gateway generates the key value and returns the plaintext once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "moonshot-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value references the model by its ID, so the key can only access the alias you created. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export MOONSHOT_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "moonshot-prod" provider: "moonshotai" adapter: "openai" api_key: ${MOONSHOT_API_KEY} api_base: "https://api.moonshot.ai/v1" apis: responses: {} models: - display_name: "kimi-k3-prod" provider: "moonshotai" model_name: "kimi-k3" provider_key: "moonshot-prod" api_keys: - display_name: "moonshot-caller" key_env: CALLER_API_KEY allowed_models: - "kimi-k3-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "kimi-k3-prod", "messages": [ { "role": "user", "content": "Say hello from Kimi." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `kimi-k3-prod`. If the request fails, check the provider key `api_key`, `api_base`, and the Moonshot model ID in `model_name`. An authentication failure on a key that works in the Moonshot console usually means the `api_base` host does not match the platform that issued the key. Verify the native Responses route with the same K3 alias: ``` curl -sS -X POST "$AISIX_PROXY/v1/responses" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "kimi-k3-prod", "input": "Say hello from Kimi through the Responses API." }' ``` Because the provider key declares `apis.responses`, AISIX forwards this request to Moonshot's native route instead of translating it through chat completions. ## Use Thinking Mode[​](#use-thinking-mode "Direct link to Use Thinking Mode") Current Kimi generations expose different reasoning controls. The configured Kimi K3 model reasons on every request. Set its reasoning effort at the top level: ``` { "reasoning_effort": "high" } ``` AISIX forwards top-level request fields it does not model itself, including `thinking`, verbatim to the upstream, so no provider-key override is needed to use this control. Reasoning controls differ across model generations: | Model | Reasoning control | Preserved thinking across turns | | ----------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- | | `kimi-k2.6` | `thinking.type` accepts `enabled` (default) or `disabled`. | `thinking.keep` defaults to `null`; set it to `all` to preserve historical `reasoning_content`. | | `kimi-k2.7-code` and `kimi-k2.7-code-highspeed` | Reasoning is always on. Omit `thinking`, or use the only accepted type, `enabled`. | Always on. The only accepted explicit `thinking.keep` value is `all`. | | `kimi-k3` | Reasoning is always on and the `thinking` object is not supported. Set top-level `reasoning_effort` to `low`, `high`, or `max` (default). | Always on. | For Kimi K2.6, turn off reasoning with this request field: ``` { "thinking": { "type": "disabled" } } ``` Check the [Kimi thinking mode guide](https://platform.kimi.ai/docs/guide/use-thinking-models) for the parameters each model accepts. Moonshot AI returns reasoning text in the `reasoning_content` field, which is the canonical field AISIX already normalizes to. On streaming responses, `reasoning_content` deltas arrive before `content` deltas; on non-streaming responses, the field appears at `choices[0].message.reasoning_content`. Because the upstream shape already matches, the `moonshotai` catalog entry sets no `response.reasoning_field` override and none is required. Set [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) only for an upstream that streams reasoning from a different `delta` path. For K3 and K2.7 tool loops or multi-turn conversations, append the complete assistant message returned by AISIX to the next chat-completions request. AISIX preserves message-level `reasoning_content` on this route. Copying only `content` and `tool_calls` loses the reasoning history that these models require. Apply the same rule to K2.6 when you set `thinking.keep` to `all`. Reasoning tokens and final-answer tokens share Moonshot AI's `max_completion_tokens` budget. The deprecated `max_tokens` field remains accepted. Raise the limit when a reasoning-heavy prompt returns truncated content. ## Review Endpoint Support[​](#review-endpoint-support "Direct link to Review Endpoint Support") Moonshot primarily documents Chat Completions plus OpenAI-compatible file and batch management. AISIX behavior depends on both the adapter and route-specific provider checks: | Route | Behavior with a Moonshot alias | | ------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including streaming, tools, structured output, and image or video content blocks on compatible Kimi models. | | `/v1/responses` | Reaches Moonshot's native Responses API for `kimi-k3` when the provider key declares `apis.responses`, as the examples on this page do. Other Kimi models are not currently supported on this native route. Without the declaration, AISIX translates the request through chat completions, which cannot preserve Responses-only fields. | | `/v1/messages` | Translated to chat completions by default, which does not preserve Moonshot reasoning history as Anthropic thinking blocks. Moonshot's native Messages route requires bearer authentication and does not provide `/v1/messages/count_tokens`, so it is not compatible with `apis.messages`; use an authenticated [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) when the native format matters. | | `/v1/files` and `/v1/batches` | Supported through the `openai` adapter. AISIX rewrites returned resource IDs so later file and batch calls route to the same alias. Each request in an uploaded batch JSONL file must name the upstream Kimi model ID, such as `kimi-k2.6`; AISIX does not rewrite model names inside the file. | | `/v1/embeddings`, `/v1/completions`, `/v1/audio/*`, and `/v1/fine_tuning/jobs` | Not supported. Moonshot does not document compatible upstream endpoints for these routes. | | `/v1/images/generations` | Rejected with `400`. The route requires a model whose provider is `openai`. | | `/v1/rerank` | Rejected with `400`. The route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/v1/videos` | Rejected with `501 not_implemented`. The route's provider allowlist does not include `moonshotai`. Kimi's video capability is video understanding through chat input, not video generation. | | `/passthrough/moonshotai/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for Moonshot-native routes beneath the route's `/v1` target. | The routed ID returned by normalized `/v1/files` is suitable for later normalized file and batch routes. It is not a raw Moonshot file ID. When a Kimi chat message must reference an uploaded image or video as `ms://<file-id>`, upload and manage that asset through `/passthrough/moonshotai/files` so the application receives the native ID. For a Moonshot-native endpoint that AISIX does not model, use a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). The `/passthrough/moonshotai` paths on this page assume a route claiming that prefix with `https://api.moonshot.ai/v1` as its `target_url`; grant the route name on the caller key's `allowed_routes`: ``` curl -sS -X GET "$AISIX_PROXY/passthrough/moonshotai/v1/models" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` Passthrough keeps caller authentication and, on an inject-mode route, injects the route's provider key credential upstream. When the request path starts with the same version segment that already ends the route's `target_url`, AISIX collapses the duplicate, so the request above reaches `https://api.moonshot.ai/v1/models` rather than a doubled `/v1/v1` path. Other useful native paths include `/tokenizers/estimate-token-count`, `/users/me/balance`, `/files`, and `/batches`. Passthrough does not rewrite an AISIX model alias or resource ID and relays upstream responses and SSE incrementally. AISIX detects chat, completions, and Responses envelopes and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. Each route binds one fixed target and, in inject mode, one provider key, so create separate routes when several Moonshot accounts or roots are in use. Moonshot's Anthropic-compatible API uses the separate `https://api.moonshot.ai/anthropic` root. It is outside the `/v1` root used in this guide, so a route targeting that root cannot reach it. Configure a separate inject-mode passthrough route for the Anthropic-compatible API; AISIX then sends the provider key as bearer authentication while preserving the native request and response bodies. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Moonshot AI and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Moonshot AI and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Nebius Token Factory [Nebius Token Factory](https://docs.tokenfactory.nebius.com/api-reference/inference/create-chat-completion) provides hosted inference for a catalog of models. AISIX gives applications stable aliases while the gateway holds the Token Factory API key. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Nebius Token Factory API key. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Nebius is a community catalog provider with an OpenAI-compatible API. AISIX connects through the `openai` adapter and authenticates upstream requests with a bearer token. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") ``` export NEBIUS_API_KEY="YOUR_NEBIUS_API_KEY" PROVIDER_KEY_ID=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "nebius-prod", "provider": "nebius", "api_key": "'"${NEBIUS_API_KEY}"'", "api_base": "https://api.tokenfactory.nebius.com/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id' ) echo "$PROVIDER_KEY_ID" ``` Do not add `adapter` to a catalog provider key. AISIX derives the `openai` adapter and bearer auth scheme from the catalog. The explicit API base makes the target visible in configuration even when the same value is available from the synchronized catalog. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Nebius model IDs include the publisher namespace. Create an alias with the complete ID: ``` MODEL_ID=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "nebius-llama-prod", "model_name": "meta-llama/Llama-3.3-70B-Instruct", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -er '.model.id' ) echo "$MODEL_ID" ``` Copy another ID from the current Nebius model catalog rather than removing or changing the publisher prefix. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "nebius-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export NEBIUS_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "nebius-prod" provider: "nebius" adapter: "openai" api_key: ${NEBIUS_API_KEY} api_base: "https://api.tokenfactory.nebius.com/v1" models: - display_name: "nebius-llama-prod" provider: "nebius" model_name: "meta-llama/Llama-3.3-70B-Instruct" provider_key: "nebius-prod" api_keys: - display_name: "nebius-caller" key_env: CALLER_API_KEY allowed_models: - "nebius-llama-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "nebius-llama-prod", "messages": [ { "role": "user", "content": "Say hello from Nebius Token Factory." } ] }' ``` The upstream request uses `POST /v1/chat/completions`, the exact Nebius model ID, and `Authorization: Bearer <NEBIUS_API_KEY>`. ## Review Model Capabilities[​](#review-model-capabilities "Direct link to Review Model Capabilities") Nebius Token Factory serves chat, reasoning, vision, embedding, rerank, and image models from the same API root. Capabilities still depend on the selected model. The `meta-llama/Llama-3.3-70B-Instruct` model in this guide is a current text-only chat model that supports function tools and structured output; it is not a reasoning or vision model. On normalized chat-completions requests, AISIX forwards OpenAI-style tools, structured-output controls, typed image or video content blocks, and top-level fields such as `reasoning_effort`. For a reasoning-capable Nebius model, AISIX normalizes upstream `reasoning_content` or `reasoning` to `reasoning_content` in the returned assistant message. Nebius can return multiple choices when `n` is greater than `1`, but AISIX returns only the first choice on normalized chat routes. Use a passthrough route when the application requires every choice or another provider-native response shape. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") | Route | Behavior with a Nebius alias | | ------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including streaming, tools, structured output, reasoning, and multimodal content supported by the selected model. | | `/v1/completions` | Supported through the `openai` adapter for a Nebius model that accepts the legacy completions contract. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), which translates to chat completions instead of calling Nebius's native Responses API. Fields without a chat equivalent, including state and native tool semantics, are ignored; Responses reasoning controls and Nebius reasoning output are not preserved. Use a passthrough route when native Responses semantics are required. | | `/v1/messages` | Supported through Anthropic-to-chat translation, not a Nebius-native Messages API. The bridge does not preserve Nebius reasoning as Anthropic thinking blocks. `/v1/messages/count_tokens` requires an Anthropic-backed model and rejects this configuration. | | `/v1/embeddings` | Supported with a separate alias for a current Nebius embedding model, such as `Qwen/Qwen3-Embedding-8B`. | | `/v1/files` | Supported for upload, list, retrieve, delete, and content download. AISIX rewrites returned file IDs so normalized follow-up calls route to the same alias. | | `/v1/fine_tuning/jobs` | Supported for create, list, retrieve, and cancel. In a create request, `model` must be the upstream Nebius base-model ID rather than an AISIX alias. | | `/v1/batches` and `/v1/audio/*` | Not supported. Nebius does not publish compatible OpenAI Batch or audio routes on this API root. | | `/v1/images/generations` | Rejected with `400` because the normalized route requires `provider: openai`. Nebius's native image-generation route is available through a passthrough route. | | `/v1/rerank` | Rejected with `400` because the normalized route accepts only the `openai`, `cohere`, and `jina` provider values. Nebius's native rerank route is available through a passthrough route. | | `/v1/videos` | Rejected with `501 not_implemented`. Nebius does not publish a video-generation route; video input to a compatible chat model is a separate capability. | | `/passthrough/nebius/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for Nebius-native routes beneath the `/v1` target, including `/responses`, `/images/generations`, `/rerank`, and `/models`. | The `/passthrough/nebius` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with `https://api.tokenfactory.nebius.com/v1` as its `target_url` and the Nebius provider key attached; grant the route on the caller key's `allowed_routes`. Normalized file and fine-tuning responses use AISIX-routed IDs for the normalized follow-up routes above. Nebius-only routes, such as `/files/{id}/link` and fine-tuning `/events` or `/checkpoints`, require raw Nebius IDs through a passthrough route. The route does not decode an AISIX-routed ID, so create and manage the resource through the passthrough route from the start when the workflow needs those native-only operations. A passthrough route does not rewrite an AISIX model alias or resource ID. Send the exact Nebius model ID in a native request body. Because the route's `target_url` ends in `/v1`, both `/passthrough/nebius/responses` and `/passthrough/nebius/v1/responses` reach `https://api.tokenfactory.nebius.com/v1/responses`; AISIX removes one duplicated version segment. A passthrough route binds a fixed target and credential rather than borrowing them from a caller-accessible model alias, and model allowlists do not gate it; give each Nebius account or API root its own route. AISIX detects chat, completions, and Responses envelopes and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. The route relays provider responses and SSE incrementally unless a hold-back output guardrail is attached. File and fine-tuning management calls are zero-token operations, and AISIX does not include Nebius training charges in inference cost accounting. See [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Check | | ----------------------------------- | ----------------------------------------------------------------------------------------------------- | | Upstream `401` or `403` | Confirm the Nebius API key and project access. | | Upstream `404` | Confirm `/v1` is present in `api_base`, and verify that the requested route and model are compatible. | | Model not found | Preserve the publisher namespace and model-name casing. | | Provider key creation returns `400` | Use `provider: "nebius"` without `adapter`. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Nebius Token Factory and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [API Key and Model Rate Limits](https://docs.api7.ai/ai-gateway/traffic-controls/rate-limits.md): configure request and token limits for the alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Nebius and another provider serving the same model. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Novita AI [Novita AI](https://novita.ai/docs/api-reference/model-apis-llm-create-chat-completion) provides hosted inference for a catalog of language models. Applications select those models through stable AISIX aliases and authenticate with gateway-issued caller keys. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Novita AI API key. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Novita AI is a community catalog provider with an OpenAI-compatible API. AISIX connects through the `openai` adapter and authenticates upstream requests with a bearer token. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") ``` export NOVITA_API_KEY="YOUR_NOVITA_API_KEY" PROVIDER_KEY_ID=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "novita-prod", "provider": "novita-ai", "api_key": "'"${NOVITA_API_KEY}"'", "api_base": "https://api.novita.ai/openai/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id' ) echo "$PROVIDER_KEY_ID" ``` The provider ID includes the `-ai` suffix. `novita` is not the AISIX catalog ID. The API base must include `/openai`. This guide uses Novita's documented `/openai/v1` root, so AISIX appends `/chat/completions` to produce `https://api.novita.ai/openai/v1/chat/completions`. Novita also accepts `/openai` without the version segment, which is the current synchronized catalog value. Using the bare host would send traffic to a route Novita does not document for LLM inference. AISIX derives the `openai` adapter for this catalog provider. Omit `adapter` from the request. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Create an alias with Novita's publisher-namespaced model ID: ``` MODEL_ID=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "novita-deepseek-prod", "model_name": "deepseek/deepseek-v3.2", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -er '.model.id' ) echo "$MODEL_ID" ``` Model IDs and availability change as the catalog from Novita is updated. Copy the current ID, including its publisher prefix and casing, from the Novita model list. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "novita-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export NOVITA_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "novita-prod" provider: "novita-ai" adapter: "openai" api_key: ${NOVITA_API_KEY} api_base: "https://api.novita.ai/openai/v1" models: - display_name: "novita-deepseek-prod" provider: "novita-ai" model_name: "deepseek/deepseek-v3.2" provider_key: "novita-prod" api_keys: - display_name: "novita-caller" key_env: CALLER_API_KEY allowed_models: - "novita-deepseek-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "novita-deepseek-prod", "messages": [ { "role": "user", "content": "Say hello from Novita AI." } ] }' ``` AISIX forwards `deepseek/deepseek-v3.2` to the Novita OpenAI-compatible chat endpoint with the Novita API key. ## Review Model Capabilities[​](#review-model-capabilities "Direct link to Review Model Capabilities") The `deepseek/deepseek-v3.2` model in this guide is current, text-only, and supports reasoning, function tools, and structured output. Novita's chat API also supports model-specific image, video, and audio input, text or audio output, and controls such as `enable_thinking` and `separate_reasoning`. AISIX forwards typed content blocks, tools, structured-output configuration, and unknown top-level request fields to the OpenAI-compatible upstream. AISIX normalizes Novita's `reasoning_content` field on chat responses. It does not preserve the optional `reasoning_details` array that some Novita models require for interleaved thinking across tool calls. Use a passthrough route for those models and replay the complete native assistant message. Novita also accepts `n` greater than `1`, but AISIX returns only the first choice on normalized chat routes. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") | Route | Behavior with a Novita alias | | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `/v1/chat/completions` | Supported, including streaming, tools, structured output, reasoning, and multimodal content supported by the selected model. | | `/v1/completions` | Supported through the `openai` adapter for a compatible Novita model. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), which translates to chat completions. Novita does not publish a native Responses route on this API root. Fields without a chat equivalent are ignored, and Novita reasoning is not preserved in the returned Responses output. | | `/v1/messages` | Supported through Anthropic-to-chat translation against the OpenAI-compatible root. The bridge does not preserve Novita reasoning as Anthropic thinking blocks. `/v1/messages/count_tokens` requires an Anthropic-backed model and rejects this configuration. | | `/v1/embeddings` | Supported with a separate alias for a Novita embedding model served on this root. | | `/v1/files` and `/v1/batches` | Supported through the `openai` adapter. AISIX rewrites returned file and batch IDs so normalized follow-up calls route to the same alias. Each request in an uploaded batch JSONL file must use the upstream Novita model ID; AISIX does not rewrite file contents, and Novita requires one model per batch file. | | `/v1/fine_tuning/jobs` and `/v1/audio/*` | Not supported. Novita does not publish compatible routes beneath the configured OpenAI root. Audio input or output through compatible chat models is a separate capability. | | `/v1/images/generations` | Rejected with `400` because the normalized route requires `provider: openai`. Novita's image-generation APIs use separate host-root paths outside the configured API base. | | `/v1/rerank` | Rejected with `400` because the normalized route does not accept `novita-ai`. Novita's native `/openai/v1/rerank` route is reachable through a passthrough route. | | `/v1/videos` | Rejected with `501 not_implemented` because the normalized route does not support `novita-ai`. Novita's video-generation APIs use separate host-root paths outside the configured API base; video input to a compatible chat model is a separate capability. | | `/passthrough/novita-ai/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for Novita-native routes beneath the route's `target_url`, including `/rerank` and `/models`. | File and batch management calls record zero tokens. When a normalized batch retrieval first observes a completed job, AISIX downloads the output file and attributes its aggregate token usage and calculated cost to the routing alias. Novita also publishes an Anthropic-compatible API at `https://api.novita.ai/anthropic`. That sibling root is outside the `/openai/v1` root used in this guide, so a passthrough route targeting the OpenAI-compatible root cannot reach it. Native Anthropic use requires a separate passthrough route with that root as its `target_url`, granted on the caller key's `allowed_routes`. A passthrough route stays beneath its `target_url`, so a route targeting the `/openai/v1` root cannot reach Novita's host-root `/v3` image, video, or audio APIs. For reachable paths, it does not rewrite an AISIX model alias or resource ID. AISIX detects chat, completions, and Responses envelopes and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. AISIX relays provider responses and SSE incrementally rather than buffering them. The `/passthrough/novita-ai` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with `https://api.novita.ai/openai/v1` as its `target_url` and the Novita provider key for credential injection; grant the route name on the caller key's `allowed_routes`. A route binds one target and credential, so use separate routes for different Novita accounts or roots. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Check | | ----------------------------------- | ------------------------------------------------------------------------------------------- | | Upstream authentication error | Confirm `NOVITA_API_KEY` is active. | | Upstream `404` | Keep `/openai` in `api_base`, and verify that the requested route exists beneath that root. | | Model not found | Copy the complete publisher-namespaced model ID from Novita. | | Provider key creation returns `400` | Use `provider: "novita-ai"` and omit `adapter`. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Novita AI and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when Novita differs from the `openai` adapter. * [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md): reach Novita-native routes that AISIX does not normalize. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # NVIDIA NIM [NVIDIA NIM](https://docs.api.nvidia.com/nim/) provides inference microservices that can run in your infrastructure. NVIDIA also exposes selected NIMs through hosted API endpoints. Applications select those models through stable AISIX aliases while the gateway holds the NVIDIA API key. This guide configures the shared hosted LLM chat API at `https://integrate.api.nvidia.com/v1` and embedding models that publish the compatible `/v1/embeddings` route. The API Catalog also includes retrieval, visual-generation, speech, and other NIMs with model-specific paths, request bodies, or hosts. The hosted catalog and its available models change over time, so check the selected model's API reference before applying this shared configuration to another NIM family. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An NVIDIA API key for the hosted NIM API, generated from a model page on [build.nvidia.com](https://build.nvidia.com/). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the NVIDIA-backed chat-completions route. NVIDIA is a community catalog provider whose shared hosted LLM endpoint accepts OpenAI chat-completions requests. AISIX connects through the `openai` adapter, authenticates upstream requests with a bearer token, and uses the NVIDIA API root as `api_base`. AISIX does not register NVIDIA-specific request or response rewrites. Review [Supply NVIDIA-Specific Behavior](#supply-nvidia-specific-behavior) for fields that differ from the standard OpenAI shape. The dashboard groups NVIDIA under **All providers (community)** and identifies the wire compatibility as assumed rather than verified. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the NVIDIA credential and API root: ``` # Replace with your value export NVIDIA_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "nvidia-prod", "provider": "nvidia", "api_key": "'"${NVIDIA_API_KEY}"'", "api_base": "https://integrate.api.nvidia.com/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `nvidia`. The AISIX Cloud Admin API accepts the value because `nvidia` is one of the models.dev catalog IDs it caches, and it assigns the `openai` adapter and bearer authentication from the catalog's community default rule. The `adapter` field is only accepted on BYO provider keys, so do not set it here. ❷ `api_key` stores the NVIDIA API key. NVIDIA authenticates the hosted NIM API with HTTP bearer authentication, which is what the `openai` adapter already sends. The value follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is `https://integrate.api.nvidia.com/v1`, the root NVIDIA documents for the shared hosted LLM API. The chat route hangs off that root as `POST https://integrate.api.nvidia.com/v1/chat/completions`, and AISIX appends `/chat/completions` to `api_base`, so the value must stop at `/v1`. AISIX strips a pasted endpoint suffix such as `/chat/completions` and any trailing slash, but treat the root as the contract rather than relying on that repair. Do not reuse this root for every API Catalog entry. For example, hosted reranking uses a retrieval-specific endpoint under `https://ai.api.nvidia.com`, and visual-generation NIMs publish other paths. Create a separate provider key with the exact API root from the selected model's API reference when it does not use the shared LLM or embeddings route. For `nvidia` the field is optional: models.dev publishes this same URL in its `api` field, and the AISIX Cloud Admin API fills it in when you omit it. The examples set it explicitly so the root each key targets stays visible in the configuration. caution Never leave an `nvidia` provider key without a resolved `api_base`. The OpenAI-family bridge refuses to fall back to the default OpenAI host for any non-OpenAI vendor, so an empty base URL produces an upstream configuration error at request time instead of a misdirected request that would carry your NVIDIA credential to another vendor's host. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") On the shared hosted LLM and embeddings routes, NVIDIA generally namespaces model IDs by publisher organization, in the form `<publisher>/<model>`. The publisher segment is part of these IDs and is never optional. Current examples include the following: | Model ID | Publisher | | ----------------------------------- | --------- | | `nvidia/nvidia-nemotron-nano-9b-v2` | NVIDIA | | `meta/llama-3.3-70b-instruct` | Meta | | `openai/gpt-oss-120b` | OpenAI | Two naming details cause most alias mistakes: * Some NVIDIA-published model names already begin with `nvidia-`, so the full ID repeats the segment, as in `nvidia/nvidia-nemotron-nano-9b-v2`. That is correct, not a typographical error. * A model page URL on `build.nvidia.com` is a page slug, not the model ID. The page for Llama 3.3 70B Instruct uses `llama-3_3-70b-instruct` in its path while the API model ID is `meta/llama-3.3-70b-instruct`. Copy the ID from the code sample on the model page rather than from the address bar. Check the model's page under [NVIDIA's NIM API reference](https://docs.api.nvidia.com/nim/) for both the current endpoint and the request body's model ID before you create an alias. Some domain-specific APIs use a model value that differs from the publisher-qualified catalog card name. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "nvidia-llama-prod", "model_name": "meta/llama-3.3-70b-instruct", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the NVIDIA model ID, including the publisher segment. Do not carry over a bare identifier such as `llama-3.3-70b-instruct` from a provider that serves the same weights without a publisher prefix. ❸ `provider_key_id` attaches the alias to the NVIDIA provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "nvidia-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export NVIDIA_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "nvidia-prod" provider: "nvidia" adapter: "openai" api_key: ${NVIDIA_API_KEY} api_base: "https://integrate.api.nvidia.com/v1" models: - display_name: "nvidia-llama-prod" provider: "nvidia" model_name: "meta/llama-3.3-70b-instruct" provider_key: "nvidia-prod" api_keys: - display_name: "nvidia-caller" key_env: CALLER_API_KEY allowed_models: - "nvidia-llama-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "nvidia-llama-prod", "messages": [ { "role": "user", "content": "Say hello from NVIDIA NIM." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `nvidia-llama-prod`. If the request fails, check the provider key `api_key`, the `api_base` root, and the publisher-namespaced model ID in `model_name`. An upstream `404` often indicates an incorrect model ID, a model that is no longer available, or a valid catalog model sent to the wrong shared route. ## Choose Between the Hosted API and a Self-Hosted NIM[​](#choose-between-the-hosted-api-and-a-self-hosted-nim "Direct link to Choose Between the Hosted API and a Self-Hosted NIM") NIM microservices can run in your infrastructure as containers or be accessed through NVIDIA-hosted API endpoints. Only the hosted API is represented by the `nvidia` catalog provider. An LLM NIM running in your own cluster serves an OpenAI-compatible API at its own address, such as `http://10.0.0.5:8000/v1`, and reaches AISIX through the private-endpoint path instead. Configure that with [Bring Your Own Endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md). | Aspect | Hosted NIM API | Self-hosted NIM microservice | | ---------------- | ------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------- | | Provider value | `nvidia` | `byo` on the AISIX Cloud Admin API, or your own label in a declarative `resources.yaml` | | `adapter` field | Rejected. The adapter is derived from the catalog. | Accepted, and set to `openai`. | | `api_base` | `https://integrate.api.nvidia.com/v1`, defaulted from the catalog when omitted | Required. Your container root, such as `http://10.0.0.5:8000/v1`. | | Model ID | Publisher-namespaced, such as `meta/llama-3.3-70b-instruct` | The model name the container serves | | Pricing metadata | Sourced from models.dev when available; verify or override `cost` on the model alias for billing | You supply `cost` on the model alias yourself | Running both is a normal setup: one provider key for burst capacity on the hosted API and one BYO key for a self-hosted NIM, with a routing model failing over between the two aliases. See [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md). Current [self-hosted LLM NIMs](https://docs.nvidia.com/nim/large-language-models/latest/api-reference.html) can expose native `/v1/completions`, `/v1/responses`, `/v1/messages`, and `/v1/messages/count_tokens` endpoints in addition to chat completions. Configuring the endpoint with the `openai` adapter does not make AISIX forward every route natively. The normalized `/v1/responses` and `/v1/messages` routes translate through the chat adapter. Normalized `/v1/messages/count_tokens` requires an Anthropic-protocol provider key. Use a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) when the application needs the NIM container's native request and response contract, with the NIM root as the route's `target_url`. For native Messages and token counting, another option is a separate BYO provider key and model alias that use `adapter: anthropic` against the same NIM root. AISIX then sends normalized `/v1/messages` and `/v1/messages/count_tokens` calls to the Anthropic-compatible NIM endpoints. Keep this separate from the `openai` adapter key used for chat, Responses bridging, and embeddings. ## Supply NVIDIA-Specific Behavior[​](#supply-nvidia-specific-behavior "Direct link to Supply NVIDIA-Specific Behavior") Because `nvidia` has no AISIX-curated adapter mapping, AISIX registers no request or response rewrites for it. The gateway sends the OpenAI request shape as-is, which is correct for NVIDIA's chat route, and leaves every NVIDIA-specific difference to you. Configure the differences with [provider-key overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides), which apply to every model that references the key. Capability support remains model- and endpoint-specific. A catalog entry can support tools, structured output, reasoning, or multimodal input without every other NVIDIA model accepting the same fields or content-block shape. AISIX forwards OpenAI-shaped chat fields and unknown top-level parameters. Some multimodal NIMs instead require a model-specific route, HTML media tags, or NVCF asset references. Check the model's inference reference instead of inferring support from the NIM family name. AISIX also returns only the first choice from an OpenAI-compatible chat response. Do not request `n` greater than `1` when callers need every generated choice. ### Reasoning Controls Are Per Model[​](#reasoning-controls-are-per-model "Direct link to Reasoning Controls Are Per Model") NVIDIA does not define one provider-wide reasoning field. Each model NIM publishes its own request schema, so the control that enables or bounds reasoning differs from model to model. For example, `nvidia/nvidia-nemotron-nano-9b-v2` toggles reasoning with `/think` and `/no_think` control tokens in the prompt, while other NIMs accept a top-level parameter such as `reasoning_effort`. Confirm the control in the model's own page under [NVIDIA's NIM API reference](https://docs.api.nvidia.com/nim/) before you rely on it, because a control one NIM accepts can be ignored or rejected by another. No gateway configuration is needed to deliver these controls. Prompt-level tokens travel inside message content, and AISIX forwards unrecognized top-level chat parameters to the upstream verbatim, so a model-specific parameter reaches NVIDIA unchanged: ``` { "model": "nvidia-nemotron-prod", "messages": [ { "role": "user", "content": "Plan a three-step migration." } ], "reasoning_effort": "low" } ``` On the response side, AISIX preserves reasoning that an upstream already returns in the canonical `reasoning_content` field, for both streaming and non-streaming responses. If a NIM streams reasoning at a different `delta` path, set [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) on the provider key. ### Token Limit Parameter Names[​](#token-limit-parameter-names "Direct link to Token Limit Parameter Names") The community default rule registers no `param_renames` for `nvidia`, so AISIX delivers `max_tokens` and `max_completion_tokens` under whichever name the caller sent. NVIDIA's LLM API documents `max_tokens`. A client that sends the newer `max_completion_tokens` therefore has its cap forwarded under a name the upstream may not act on. Add a rename on the provider key when that applies: ``` { "request": { "param_renames": { "max_completion_tokens": "max_tokens" } } } ``` If a request carries both names, AISIX uses the value from the original caller-facing name. ## Route Embeddings to NeMo Retriever Models[​](#route-embeddings-to-nemo-retriever-models "Direct link to Route Embeddings to NeMo Retriever Models") NVIDIA publishes NeMo Retriever embedding models in the same catalog, and AISIX dispatches `/v1/embeddings` through the same `openai` adapter, so an alias that names an embedding model works on that route. Several NVIDIA embedding fields need special handling. NVIDIA's asymmetric retrieval families, such as the NV-EmbedQA and E5 models, require an `input_type` of `query` or `passage`, and using the wrong one degrades retrieval accuracy. Current Retriever NIM schemas can also define fields such as `modality`, `embedding_type`, and `truncate`. The AISIX embeddings route builds the upstream body from a closed set of fields: `model`, `input`, `encoding_format`, and `dimensions`. NVIDIA-specific fields placed in a caller request body therefore do not reach the upstream. Set it on the provider key instead, with `request.default_body_fields`: ``` { "request": { "default_body_fields": { "input_type": "passage" } } } ``` AISIX merges these fields into the outbound body on the embeddings path, not only on chat. Provider-key overrides apply to every model that references the key. Create one provider key with `input_type: passage` for indexing and another with `input_type: query` for search. Point the matching model aliases at those keys. Use the same `request.default_body_fields` mechanism for another fixed NVIDIA field such as `truncate`. A value that must vary per request or per input, such as a mixed `modality` array, requires a passthrough route because default body fields are static provider-key configuration. Symmetric embedding models that do not take `input_type` need no provider-key override. See [`request.default_body_fields`](https://docs.api7.ai/ai-gateway/reference/resources-file.md#provider-keys) for the field definition, and confirm whether your model requires `input_type` in [NVIDIA's embedding API reference](https://docs.nvidia.com/nim/nemo-retriever/text-embedding/latest/reference.html). ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") The `nvidia` provider value is accepted on the adapter-dispatched routes and rejected by the routes that carry their own provider allowlist. | Route | Behavior with an NVIDIA alias | | -------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported on the shared hosted LLM route, including `stream: true`. Model capabilities remain model-specific. | | `/v1/completions` | AISIX forwards this route through the OpenAI adapter. Current self-hosted LLM NIMs publish it, but NVIDIA does not document it as a general route for the shared hosted LLM API; use it only when the selected upstream explicitly supports it. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over chat, even when a self-hosted NIM publishes a native Responses endpoint. The bridge preserves text and function-call turns but synthesizes a new response; state, hosted tools, reasoning controls, and other fields without a chat equivalent are dropped. Upstream reasoning text is not represented in the synthesized Responses output. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation over chat, not as native NIM Messages passthrough. NIM reasoning is not returned as Anthropic `thinking` blocks. Token counting at `/v1/messages/count_tokens` requires an Anthropic-protocol provider key. | | `/v1/embeddings` | Supported when the alias names an NVIDIA embedding model. See [Route Embeddings to NeMo Retriever Models](#route-embeddings-to-nemo-retriever-models). | | `/v1/images/generations` | Rejected with `400` for an NVIDIA alias. NVIDIA Visual GenAI NIMs can publish a native OpenAI-compatible image route, but the normalized AISIX route accepts only models whose provider is `openai`. | | `/v1/rerank` | Rejected with `400`. The route accepts only the `openai`, `cohere`, and `jina` provider values. See the note below. | | `/v1/videos` | Rejected with `501 not_implemented`. Visual GenAI NIMs can publish native video generation at `/v1/videos/generations`, but the `nvidia` value is not in the normalized AISIX video route's provider allowlist. | | `/v1/audio/*` | Forwarded to OpenAI-compatible audio paths, but the shared hosted LLM API and self-hosted LLM NIM do not publish those routes. Speech NIMs use separate APIs; reach them through a passthrough route targeting the documented root. | | `/v1/files`, `/v1/batches`, `/v1/fine_tuning/jobs` | AISIX can dispatch these routes through an OpenAI adapter, but neither the shared hosted API nor current self-hosted LLM NIM publishes the corresponding OpenAI jobs APIs. The upstream therefore rejects the request. NIM model customization or LoRA management is a different API contract. | | `/passthrough/nvidia/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for provider-native routes, with limited gateway normalization. | NVIDIA does publish reranking models, but they are not reachable through `/v1/rerank` even apart from the provider allowlist. Current self-hosted Retriever NIM documents `POST /v1/ranking` with a `query` object and a `passages` array. Neither the path nor the body matches the normalized route, and hosted catalog reranking can use a different host and path. Reach a reranking model through a passthrough route using the endpoint and body from its [NVIDIA reranking API reference](https://docs.nvidia.com/nim/nemo-retriever/text-reranking/latest/reference.html). The `/passthrough/nvidia` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with `https://integrate.api.nvidia.com/v1` as its `target_url` and the NVIDIA provider key for credential injection; grant the route name on the caller key's `allowed_routes`. A route binds one fixed target and credential — the request body's model does not choose credentials, and an AISIX alias is not rewritten to the upstream model ID — so create separate routes for NIMs served from other roots, such as retrieval endpoints under `https://ai.api.nvidia.com`. The route relays upstream responses, including SSE, incrementally. AISIX detects chat, completions, and Responses envelopes and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. Prefer normalized inference routes when you need alias rewriting, token accounting, or AISIX cost estimates. Catalog pricing is metadata for those AISIX estimates and does not represent NVIDIA invoicing; many NVIDIA catalog models currently have zero or missing published cost. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to the NVIDIA NIM hosted API and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between the hosted NIM API and a self-hosted NIM. * [Bring Your Own Endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md): connect a NIM microservice you run yourself. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Ollama [Ollama](https://docs.ollama.com/api/openai-compatibility) runs language models locally or in private infrastructure and exposes an OpenAI-compatible API. AISIX adds caller keys, stable model aliases, traffic controls, and usage reporting while Ollama continues to serve inference in your environment. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * Ollama installed on a host reachable from the AISIX gateway. * `curl` and `jq`. ## Prepare Ollama[​](#prepare-ollama "Direct link to Prepare Ollama") Pull the model used in this guide: ``` ollama pull gpt-oss:20b ``` Ollama binds to `127.0.0.1:11434` by default. If AISIX runs in another container or host, configure a reachable bind address by following the [Ollama server configuration](https://docs.ollama.com/faq#how-do-i-configure-ollama-server). For example, a foreground server can listen on all interfaces: ``` OLLAMA_HOST="0.0.0.0:11434" ollama serve ``` caution The local Ollama API does not require authentication. Binding it to `0.0.0.0` makes it reachable from other network peers. Restrict the listening network with firewall, container-network, or Kubernetes policy controls, and do not expose the port directly to the public Internet. Export an API base reachable **from the AISIX gateway**: ``` # Docker Desktop gateway to Ollama on the host export OLLAMA_API_BASE="http://host.docker.internal:11434/v1" ``` Choose the address for your topology: | AISIX gateway and Ollama topology | Example API base | | ------------------------------------------- | ------------------------------------------------------ | | Both processes on one host | `http://127.0.0.1:11434/v1` | | AISIX in Docker Desktop, Ollama on the host | `http://host.docker.internal:11434/v1` | | Both containers on one Docker network | `http://ollama:11434/v1` | | Kubernetes | `http://ollama.<namespace>.svc.cluster.local:11434/v1` | On Linux, `host.docker.internal` may require an explicit `host-gateway` mapping. A successful request from your laptop does not prove that the AISIX gateway container can reach the same address. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Ollama is a private endpoint rather than an AISIX catalog provider. Configure it with the `byo` provider value and select the `openai` adapter explicitly. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") ``` PROVIDER_KEY_ID=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "ollama-local", "provider": "byo", "adapter": "openai", "api_key": "ollama", "api_base": "'"${OLLAMA_API_BASE}"'", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id' ) echo "$PROVIDER_KEY_ID" ``` AISIX requires a non-empty `api_key` in the provider-key schema. Ollama's OpenAI client examples use `ollama` because the client requires a value, but the local Ollama server ignores it. AISIX sends the placeholder as a bearer token; it is not a security control. A BYO key requires `provider: "byo"`, a non-empty `api_key`, and `api_base`. This guide sets `adapter: "openai"` explicitly; when omitted, AISIX defaults a BYO key to the OpenAI-compatible adapter. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Use the exact local Ollama model tag: ``` MODEL_ID=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "ollama-gpt-oss-prod", "model_name": "gpt-oss:20b", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -er '.model.id' ) echo "$MODEL_ID" ``` Run `ollama ls` to see installed model tags. The tag, including a suffix such as `:20b`, is sent upstream unchanged. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "ollama-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export OLLAMA_API_BASE="http://host.docker.internal:11434/v1" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "ollama-local" provider: "ollama" adapter: "openai" api_key: "ollama" api_base: "${OLLAMA_API_BASE}" models: - display_name: "ollama-gpt-oss-prod" provider: "ollama" model_name: "gpt-oss:20b" provider_key: "ollama-local" api_keys: - display_name: "ollama-caller" key_env: CALLER_API_KEY allowed_models: - "ollama-gpt-oss-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "ollama-gpt-oss-prod", "messages": [ { "role": "user", "content": "Say hello from Ollama." } ] }' ``` Ollama should log `POST /v1/chat/completions`, and AISIX should return an OpenAI-compatible response. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") Ollama documents OpenAI-compatible chat completions, completions, Responses, models, and embeddings routes. It also publishes an [Anthropic-compatible Messages route](https://docs.ollama.com/api/anthropic-compatibility). AISIX behavior depends on the adapter assigned to the BYO provider key; the `openai` adapter in this guide does not make every Ollama-native route a transparent passthrough. | Route | Behavior with the Ollama alias in this guide | | ----------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including streaming, tools, structured output, vision, and reasoning controls when the installed model supports them. | | `/v1/completions` | Supported through the OpenAI adapter. Ollama accepts only a string in `prompt`. | | `/v1/embeddings` | Supported when the alias names an installed embedding model. Ollama and AISIX accept a string or array of strings. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over chat, not through Ollama's native Responses route. The bridge preserves text and function-call turns but drops state, reasoning controls, reasoning output, and fields without a chat equivalent. | | `/v1/messages` | Translated through chat because this provider key uses `adapter: openai`. See the native Messages option below. | | `/v1/messages/count_tokens` | Rejected for the `openai` adapter. Ollama does not currently implement this endpoint even when a separate Anthropic adapter is used. | | `/v1/models` | Returns caller-accessible AISIX model aliases, not the models installed in Ollama. Use `ollama ls` or a passthrough route to query the Ollama inventory. | | `/v1/images/generations`, `/v1/rerank` | Rejected with `400` because the normalized routes do not accept this guide's provider label (`byo` in AISIX Cloud or `ollama` in the resources file). | | `/v1/videos` | Rejected with `501 not_implemented` because neither provider label is in the video provider allowlist. | | `/v1/audio/*`, `/v1/files`, `/v1/batches`, `/v1/fine_tuning/jobs` | AISIX can forward these OpenAI-shaped routes, but Ollama does not publish the corresponding APIs. The upstream rejects the request. | | `/passthrough/byo/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for Ollama-native routes. The conventional prefix is `/passthrough/byo` in AISIX Cloud and `/passthrough/ollama` with an open-source resources file, matching this page's provider label. | Ollama added native `/v1/responses` in version 0.13.3. It supports streaming, function tools, and reasoning summaries, but not state through `previous_response_id` or `conversation`. AISIX still bridges a normalized `/v1/responses` call for this BYO alias because the provider is not `openai`. Use a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) when the application needs Ollama's native Responses contract. For native Anthropic-shaped Messages, create a second BYO provider key and model alias that use `adapter: anthropic` against the same Ollama root. AISIX then sends normalized `/v1/messages` calls to Ollama's native Messages endpoint. Current Ollama limitations still apply: token counting, forced `tool_choice`, prompt caching, citations, and Messages batches are not supported, and extended-thinking budgets are accepted but not enforced. Ollama's OpenAI-compatible chat route does not currently support `tool_choice`, `logit_bias`, `user`, or `n`. AISIX can forward these fields, but forwarding does not add upstream support. For vision models, send a base64 image in an `image_url` content part; Ollama does not support a remote image URL on this route. The `/passthrough` paths above assume a passthrough route claiming the chosen prefix with the Ollama root as its `target_url`; grant the route name on the caller key's `allowed_routes`. Passthrough does not rewrite an AISIX alias in the request body, and a route relays to its fixed target with its bound provider key regardless of the requested model. It relays upstream responses incrementally, including SSE. AISIX detects chat, completions, and Responses envelopes and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. Prefer the normalized routes when you need alias rewriting, token accounting, or AISIX cost estimates. Because Ollama is a BYO endpoint without catalog pricing, configure [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md) in AISIX Cloud or `cost` metadata in an open-source model resource when you need cost estimates. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Check | | ----------------------------------- | -------------------------------------------------------------------------------------- | | Connection refused or timeout | Test `OLLAMA_API_BASE` from the AISIX gateway container, not only from the host. | | Ollama listens only on `127.0.0.1` | Set `OLLAMA_HOST` through the supported service configuration and restart Ollama. | | Model not found | Run `ollama pull gpt-oss:20b` and verify the tag with `ollama ls`. | | Provider key creation returns `400` | Include `provider: "byo"`, `adapter: "openai"`, a non-empty `api_key`, and `api_base`. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Ollama and verified the model alias. Continue with these guides: * [Bring Your Own Endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md): review the reusable setup for private OpenAI-compatible servers. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Ollama and another provider serving the same model. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # OpenAI [OpenAI](https://platform.openai.com/docs/) provides hosted GPT models through its API. Applications call them through stable AISIX aliases while the gateway keeps the OpenAI credential out of client code. This guide covers both GPT chat traffic and Sora video tasks through a single OpenAI provider key. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An OpenAI API key from the [OpenAI platform](https://platform.openai.com/api-keys). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the OpenAI-backed route. OpenAI is the gateway's native OpenAI-compatible upstream, and the integration uses the `openai` adapter. You can leave `api_base` unset for OpenAI's canonical endpoint or set it to target an OpenAI-compatible proxy that preserves bearer authentication and the OpenAI route shape. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the OpenAI credential: ``` # Replace with your values export OPENAI_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "openai-prod", "provider": "openai", "api_key": "'"${OPENAI_API_KEY}"'", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') ``` ❶ `provider` is `openai`. This is the only provider value that lets AISIX fall back to the default base URL `https://api.openai.com/v1`. ❷ `api_key` stores the OpenAI API key. The value is encrypted before it is stored and is never returned by read endpoints. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `allowed_environments` lists the environments that may reference this provider key when creating models. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. You can omit `api_base` for OpenAI. To target a compatible bearer-authenticated proxy or regional gateway, set `api_base` explicitly. Use the dedicated [Azure OpenAI](https://docs.api7.ai/ai-gateway/providers/azure-openai.md) integration for the Azure OpenAI service itself. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "gpt-5.6-sol-prod", "model_name": "gpt-5.6-sol", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the OpenAI model ID. This guide uses `gpt-5.6-sol`, the current flagship model in the [OpenAI model catalog](https://developers.openai.com/api/docs/models). Choose a different current model when its cost, latency, or modality better matches the workload. ❸ `provider_key_id` attaches the alias to the OpenAI provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key resource that can access the model alias. The control plane generates the key value and returns the plaintext once in the create response: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "openai-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') ``` The `allowed_models` value references the model ID captured in the previous step. Store the plaintext key securely. The control plane stores only a hash; the plaintext is not retrievable after this response. After each write, the configuration projects to the attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export OPENAI_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "openai-prod" provider: "openai" adapter: "openai" api_key: ${OPENAI_API_KEY} models: - display_name: "gpt-5.6-sol-prod" provider: "openai" model_name: "gpt-5.6-sol" provider_key: "openai-prod" api_keys: - display_name: "openai-caller" key_env: CALLER_API_KEY allowed_models: - "gpt-5.6-sol-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-5.6-sol-prod", "messages": [ { "role": "user", "content": "Say hello from OpenAI." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `gpt-5.6-sol-prod`. Confirm the request on the [OpenAI usage dashboard](https://platform.openai.com/usage). If the request fails with an upstream authentication error, check the provider key `api_key`. For reasoning, tool-calling, and multi-turn workflows, OpenAI recommends the Responses API. Send a native Responses request through the same alias: ``` curl -sS -X POST "$AISIX_PROXY/v1/responses" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-5.6-sol-prod", "input": "Say hello from the OpenAI Responses API." }' ``` Because the model's configured provider is `openai`, AISIX rewrites the model alias and forwards this request to OpenAI's native Responses endpoint. It does not use the cross-provider Responses bridge. ## Generate Videos with Sora[​](#generate-videos-with-sora "Direct link to Generate Videos with Sora") caution OpenAI [deprecated the Videos API and the Sora 2 models](https://developers.openai.com/api/docs/deprecations) on March 24, 2026 and will remove them from the API on September 24, 2026. Aliases backed by `sora-2` or `sora-2-pro` stop working on that date, and OpenAI lists no replacement model on this API. The same provider key drives the gateway's modeled [video routes](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md). Create a second alias that points at a Sora model — `sora-2` or `sora-2-pro`: In AISIX Cloud, create the video alias and a caller key scoped to it: ``` VIDEO_MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "sora-video-prod", "model_name": "sora-2", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$VIDEO_MODEL_ID" VIDEO_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "openai-video-caller", "allowed_models": ["'"${VIDEO_MODEL_ID}"'"] }' | jq -r '.plaintext') ``` For the open-source AISIX gateway, add `sora-video-prod` to the existing `models` collection. Replace the existing `openai-caller` entry with the updated entry below so it allows both model aliases. Preserve unrelated entries and collections: resources.yaml (video model access) ``` models: - display_name: "sora-video-prod" provider: "openai" model_name: "sora-2" provider_key: "openai-prod" api_keys: - display_name: "openai-caller" key_env: CALLER_API_KEY allowed_models: - "gpt-5.6-sol-prod" - "sora-video-prod" ``` Validate and reload or restart the declarative resources file as described above, then use the existing caller key for the video request: ``` export VIDEO_API_KEY="$CALLER_API_KEY" ``` Submit a task: ``` curl -sS -X POST "$AISIX_PROXY/v1/videos" \ -H "Authorization: Bearer ${VIDEO_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "sora-video-prod", "prompt": "A paper boat drifts down a rain-soaked street at dusk.", "seconds": 4, "size": "1280x720" }' ``` Five provider-specific behaviors are worth knowing before you script around this route: * **OpenAI is the only video provider with a default base URL.** An OpenAI-backed video alias works with `api_base` unset, exactly like the chat alias on this page. Every other video provider requires `api_base` on the provider key. * **Sora validates `seconds` and `size` itself.** OpenAI's create-video schema accepts `4`, `8`, or `12` for `seconds` and lists `720x1280`, `1280x720`, `1024x1792`, and `1792x1024` for `size`, but the per-model pages publish narrower resolution sets — check the page for the model your alias names. AISIX forwards `seconds` as a string and checks that `size` has the `WIDTHxHEIGHT` shape before forwarding it verbatim, so a well-formed value the model does not accept is rejected by the provider, not by the gateway. * **Finished videos stream through the gateway.** Sora serves the finished file from an authenticated content endpoint rather than a signed URL, so `GET /v1/videos/{id}/content` returns `200` with the MP4 bytes instead of a `302`. AISIX fetches the file with the provider credential and relays it chunk by chunk, so the credential never reaches the caller and a large file does not grow the gateway's memory use. Size the gateway's egress for this traffic: those bytes transit the gateway and do not count against the model's rate limits. * **`progress` is a real percentage.** Sora reports task progress, so poll responses carry the provider's own completion percentage. Providers that report none stay at `0` until the task completes. * **Image-to-video requires a passthrough route.** The normalized request models only `prompt`, `seconds`, and `size`; extra fields, including `input_reference`, are ignored rather than rejected, so a remix or image-guided request would silently generate from the prompt alone. Call `/passthrough/openai/v1/videos` with the native multipart body for those, and use the native remix route the same way. An OpenAI SDK client reaches this path too, but it needs its own client: point its base URL at `${AISIX_PROXY}/passthrough/openai/v1` rather than the `${AISIX_PROXY}/v1` client used elsewhere on this page, and grant the route name in the caller key's `allowed_routes`. The video create helper always sends `multipart/form-data`, which the modeled route does not accept. Passthrough does not return the unified AISIX video object or gateway-encoded task ID. For the full submit, poll, and download workflow, including status semantics and rate-limit behavior on polling, see [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md). ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") OpenAI exposes several API families that use different model types. Configure a separate AISIX model alias for each upstream model an application needs. | Route | Behavior with an OpenAI-backed alias | | -------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported with the OpenAI request and response shape. Model-specific feature and parameter limits still apply. | | `/v1/responses` | Forwarded to OpenAI's native Responses endpoint after model-alias rewriting. Stateful fields, hosted tools, reasoning controls, and upstream SSE remain on the native path. | | `/v1/completions` | Forwarded to OpenAI, but this is a legacy endpoint and the selected model must support it. | | `/v1/embeddings` | Supported with an alias for an OpenAI embedding model, such as `text-embedding-3-large`. | | `/v1/images/generations` | Supported with an alias for a current OpenAI image model. | | `/v1/images/edits` | Supported with an alias for an OpenAI image-editing model such as `gpt-image-2`. Requests are `multipart/form-data`; see [Image Editing](https://docs.api7.ai/ai-gateway/endpoints/image-editing.md). | | `/v1/videos` and its status and content routes | Supported with an alias for a Sora model, `sora-2` or `sora-2-pro`. The content route streams the MP4 through the gateway instead of redirecting, because OpenAI serves finished files from an authenticated endpoint. See [Generate Videos with Sora](#generate-videos-with-sora). | | `/v1/audio/*` | Supported with an appropriate speech, transcription, or translation model alias. | | `/v1/realtime` | Supported as a WebSocket relay with a direct Realtime model alias. See [Realtime API](https://docs.api7.ai/ai-gateway/endpoints/realtime.md). | | `/v1/files`, `/v1/batches`, `/v1/fine_tuning/jobs` | Supported through the OpenAI adapter. These job-style routes select credentials differently from ordinary inference calls; see [Files, Batch, and Fine-tuning](https://docs.api7.ai/ai-gateway/endpoints/batch-files-fine-tuning.md). | | `/v1/messages` | Translated through OpenAI chat; OpenAI does not expose a native Anthropic Messages endpoint. | | `/v1/messages/count_tokens` | Rejected because token counting on this route requires an Anthropic-backed model. | | `/v1/models` | Returns caller-accessible AISIX model aliases, not the OpenAI account's model inventory. Use `/passthrough/openai/v1/models` when the native list is required. | | `/v1/rerank` | AISIX accepts the `openai` provider value, but the public OpenAI API does not expose `/v1/rerank`, so the upstream rejects the call. | Use a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for an OpenAI route that AISIX does not model, such as moderation or image variations. The `/passthrough/openai` paths on this page assume a route claiming that prefix with OpenAI's API root as its `target_url`; grant the route on the caller key's `allowed_routes`. Passthrough uses OpenAI model IDs rather than AISIX aliases and has different streaming and usage-accounting behavior. ## Use an OpenAI SDK[​](#use-an-openai-sdk "Direct link to Use an OpenAI SDK") Set `AISIX_BASE_URL` to `${AISIX_PROXY}/v1` and use the caller key as the API key. See the [OpenAI SDK](https://docs.api7.ai/ai-gateway/getting-started/openai-sdk.md) guide. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to OpenAI and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Audio Input and Output with Chat Completions](https://docs.api7.ai/ai-gateway/endpoints/chat-audio.md): send recorded audio to a compatible chat model or save its generated audio response. * [Azure OpenAI](https://docs.api7.ai/ai-gateway/providers/azure-openai.md): configure an Azure-hosted OpenAI deployment instead. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Other OpenAI-Compatible Providers Many public model providers expose an OpenAI-compatible API. AISIX Cloud can connect to a provider when it accepts bearer-authenticated OpenAI chat-completions requests and appears in the AISIX provider catalog. The open-source AISIX gateway can use any reachable endpoint with the same protocol and authentication shape under an operator-chosen provider label. When a provider has a dedicated setup listed under [Provider Upstreams](https://docs.api7.ai/ai-gateway/providers/overview.md), follow that page instead. The configuration below is a parameterized template for another public provider; the AISIX Cloud path assumes that the provider is in its catalog. For a private or customer-operated server, use the dedicated [Ollama](https://docs.api7.ai/ai-gateway/providers/ollama.md), [vLLM](https://docs.api7.ai/ai-gateway/providers/vllm.md), or [Bring Your Own Endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md) guide. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An API key for an OpenAI-compatible provider. For AISIX Cloud, the provider must be listed in the AISIX provider catalog; the open-source resources workflow does not use that catalog for admission. * `curl` and `jq`. ## Select Provider Values[​](#select-provider-values "Direct link to Select Provider Values") For AISIX Cloud, choose the exact provider ID shown by the AISIX provider catalog. For an open-source resources file, choose a stable provider label, such as the provider's lowercase name. Then copy the API root and model ID from the provider's official API reference: ``` export PROVIDER_ID="YOUR_PROVIDER_ID" export PROVIDER_API_KEY="YOUR_PROVIDER_API_KEY" export PROVIDER_API_BASE="https://api.provider.example/v1" export UPSTREAM_MODEL_ID="publisher/model-id" export MODEL_ALIAS="provider-model-prod" ``` The address and IDs above are fictional placeholders. Replace every value before running the setup. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the provider-backed chat-completions route. AISIX connects through the `openai` adapter and uses the provider's API root as `api_base`. Give the provider key a descriptive label so operators can identify the upstream later. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") ``` PROVIDER_KEY_RESPONSE=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @- <<EOF { "display_name": "community-provider-prod", "provider": "${PROVIDER_ID}", "api_key": "${PROVIDER_API_KEY}", "api_base": "${PROVIDER_API_BASE}", "allowed_environments": ["${ENV_ID}"] } EOF ) PROVIDER_KEY_ID=$(printf '%s' "$PROVIDER_KEY_RESPONSE" | jq -er '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` `provider` must exactly match an ID in the AISIX catalog. The AISIX Cloud Admin API derives the `openai` adapter and bearer authentication for community catalog providers. Do not send an `adapter` field; that field is accepted only when `provider` is `byo`. A provider without a dedicated setup page is admitted from the synced public catalog rather than from the gateway's built-in provider list. If the create returns `400 INVALID_REQUEST` naming the catalog, the control plane does not have that provider in its current catalog. Connected deployments sync at start and every 24 hours; packaged On-Premises deployments use a bundled snapshot by default and do not refresh it. Retry after an online sync, review the [On-Premises pricing-catalog settings](https://docs.api7.ai/ai-gateway/reference/on-premises-configuration.md#pricing-catalog), or configure the upstream with [Bring Your Own Endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md). AISIX appends the endpoint path to `api_base`. Include provider-specific prefixes such as `/v1`, `/openai`, or `/openai/v1` when the official endpoint requires them. Setting the root explicitly also avoids depending on the API base cached by the latest catalog sync. Provider key secrets follow the credential-handling behavior described in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ### Create a Model[​](#create-a-model "Direct link to Create a Model") Map the caller-facing alias to the provider's exact model ID: ``` MODEL_RESPONSE=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @- <<EOF { "display_name": "${MODEL_ALIAS}", "model_name": "${UPSTREAM_MODEL_ID}", "provider_key_id": "${PROVIDER_KEY_ID}" } EOF ) MODEL_ID=$(printf '%s' "$MODEL_RESPONSE" | jq -er '.model.id') echo "$MODEL_ID" ``` `display_name` is the alias callers send in `model`. `model_name` is sent to the upstream unchanged, so preserve any publisher namespace, casing, punctuation, and version suffix. To configure cost metadata for budget accounting or usage reports, see [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md#cost-metadata). ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create a caller key limited to the model resource: ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "community-provider-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` The plaintext is returned only when the caller key is created. Store it securely. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export PROVIDER_ID="YOUR_PROVIDER_ID" export PROVIDER_API_BASE="YOUR_PROVIDER_API_BASE" export UPSTREAM_MODEL_ID="YOUR_UPSTREAM_MODEL_ID" export MODEL_ALIAS="YOUR_MODEL_ALIAS" export PROVIDER_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` The open-source gateway validates the provider label's shape but does not require it to appear in the AISIX Cloud catalog. Use the same label on the provider key and model. The label also identifies the upstream in telemetry and, by convention, in a passthrough route path such as `/passthrough/example-provider/*`. For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "community-provider-prod" provider: "${PROVIDER_ID}" adapter: "openai" api_key: ${PROVIDER_API_KEY} api_base: "${PROVIDER_API_BASE}" models: - display_name: "${MODEL_ALIAS}" provider: "${PROVIDER_ID}" model_name: "${UPSTREAM_MODEL_ID}" provider_key: "community-provider-prod" api_keys: - display_name: "community-provider-caller" key_env: CALLER_API_KEY allowed_models: - "${MODEL_ALIAS}" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ --data-binary @- <<EOF { "model": "${MODEL_ALIAS}", "messages": [ { "role": "user", "content": "Say hello through the configured provider." } ] } EOF ``` The response should be OpenAI-compatible and should carry the caller-facing alias. Use the provider's request logs or usage page, when available, to confirm that the request reached the intended upstream account and model. If the gateway returns an upstream authentication error, check the provider key's `api_key`. If it returns an upstream route error, check `api_base` and `UPSTREAM_MODEL_ID`. ## Support Provider-Specific Behavior[​](#support-provider-specific-behavior "Direct link to Support Provider-Specific Behavior") The provider must accept OpenAI chat-completions requests. A provider with a different request format needs a native [adapter protocol family](https://docs.api7.ai/ai-gateway/providers/adapters.md) or a compatible endpoint. The `openai` adapter does not make every normalized AISIX endpoint available to every provider label: * `/v1/completions`, `/v1/embeddings`, `/v1/audio/*`, `/v1/files`, `/v1/batches`, and `/v1/fine_tuning/jobs` can dispatch through this adapter, but they work only when the upstream implements the corresponding OpenAI route and fields. * `/v1/responses` uses the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over chat for a community provider label. `/v1/messages` similarly translates the request and response through chat rather than using a provider-native Messages route. * `/v1/images/generations`, `/v1/rerank`, and `/v1/videos` enforce provider allowlists. They can reject a community provider label even when its upstream exposes a similarly named route. Configure a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) when an application needs a provider-native route or contract; by convention the route claims `/passthrough/<label>` with the provider's API base as its `target_url`, and the caller key must grant the route in its `allowed_routes`. Passthrough does not rewrite the model alias and has different streaming and usage-accounting behavior, so review its limitations before adopting it. See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the normalized-route matrix. AISIX preserves `reasoning_content` and normalizes `reasoning` to that canonical field. If a provider streams reasoning from a different `delta` path, use the [`response.reasoning_field` override](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) on the provider key. ## Next Steps[​](#next-steps "Direct link to Next Steps") * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # OpenRouter [OpenRouter](https://openrouter.ai/docs/guides/overview/models) aggregates models from many providers behind one API and credential. AISIX adds stable model aliases, caller access controls, rate limits, and usage accounting while keeping the OpenRouter credential at the gateway. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An OpenRouter API key from the [OpenRouter keys page](https://openrouter.ai/settings/keys). OpenRouter key values begin with `sk-or-`. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the OpenRouter-backed chat-completions route. Because OpenRouter exposes an OpenAI-compatible API, AISIX connects through the `openai` adapter and uses the OpenRouter API root as `api_base`. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the OpenRouter credential and API root: ``` # Replace with your values export OPENROUTER_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "openrouter-prod", "provider": "openrouter", "api_key": "'"${OPENROUTER_API_KEY}"'", "api_base": "https://openrouter.ai/api/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `openrouter`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the OpenRouter API key. The value is encrypted before storage and never returned by read endpoints. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is `https://openrouter.ai/api/v1`. OpenRouter serves its API under `/api` on its main host, so the root carries both segments and is not `https://openrouter.ai/v1`. AISIX appends the endpoint path, for example `/chat/completions`, to this value. Always set `api_base` explicitly for OpenRouter. Unlike most featured catalog providers, OpenRouter has no curated default base URL in AISIX, so an omitted `api_base` leaves the AISIX Cloud Admin API dependent on synced catalog metadata and the create can fail with `400 INVALID_REQUEST`. The gateway enforces the same rule at request time. When a provider key for a non-OpenAI vendor reaches the gateway with an empty `api_base`, the OpenAI-family bridge refuses to fall back to the public OpenAI host and returns an upstream configuration error instead. An OpenRouter credential is therefore never sent to `api.openai.com`. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") OpenRouter model IDs are namespaced. The upstream ID is `vendor/model`, and the vendor prefix is part of the ID rather than a routing hint. Check the [OpenRouter models list](https://openrouter.ai/models) for a current ID before creating a model alias. The naming convention has four forms: * `vendor/model` is the base form, for example `anthropic/claude-sonnet-5`, `openai/gpt-5.2-pro`, or `deepseek/deepseek-v4-pro`. * A variant suffix is appended to the slug, for example `:free` or `:thinking`. * A `~` before the vendor resolves to the newest version in a family, for example `~anthropic/claude-sonnet-latest`. * `openrouter/auto` selects OpenRouter's own Auto Router instead of a named model. The model alias and upstream model ID are different names. `display_name` is the AISIX alias that callers put in the request `model` field, while `model_name` is the OpenRouter model ID. AISIX forwards `model_name` to the upstream verbatim, including the `/` in the vendor prefix. It does not split the ID or strip the vendor segment. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "openrouter-sonnet-prod", "model_name": "anthropic/claude-sonnet-5", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. Keep it a short flat name. Copying the namespaced upstream ID into the alias makes client code harder to read and hides which alias is in use. ❷ `model_name` is the OpenRouter model ID, for example `anthropic/claude-sonnet-5` or `openai/gpt-5.2-pro`. ❸ `provider_key_id` attaches the alias to the OpenRouter provider key. Create one alias per OpenRouter model you intend to expose. A single OpenRouter provider key can back any number of aliases, which is how one credential fans out to several vendors while each alias keeps its own allowlist entry and rate limits. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key resource that can access the model alias. The gateway generates the key value; the plaintext is returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "openrouter-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value references the model by its ID, so the key can only access the alias you created. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export OPENROUTER_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "openrouter-prod" provider: "openrouter" adapter: "openai" api_key: ${OPENROUTER_API_KEY} api_base: "https://openrouter.ai/api/v1" models: - display_name: "openrouter-sonnet-prod" provider: "openrouter" model_name: "anthropic/claude-sonnet-5" provider_key: "openrouter-prod" api_keys: - display_name: "openrouter-caller" key_env: CALLER_API_KEY allowed_models: - "openrouter-sonnet-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "openrouter-sonnet-prod", "messages": [ { "role": "user", "content": "Say hello from OpenRouter." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `openrouter-sonnet-prod`, not the namespaced upstream ID. If the request fails, check these OpenRouter-specific causes first: * A `404` from the upstream usually means the vendor prefix is missing or misspelled in `model_name`. `claude-sonnet-5` is not a valid OpenRouter ID; `anthropic/claude-sonnet-5` is. * A `401` from the upstream usually means `api_key` does not hold an OpenRouter key. Keys issued by the model's original vendor are not accepted by OpenRouter. * A connection error to an unexpected host usually means `api_base` is missing the `/api` segment. ## Pass OpenRouter Routing and Reasoning Controls[​](#pass-openrouter-routing-and-reasoning-controls "Direct link to Pass OpenRouter Routing and Reasoning Controls") AISIX does not strip top-level chat-completions fields it does not model. Unmodeled fields are forwarded to the upstream verbatim, so OpenRouter's own request-level controls reach OpenRouter unchanged. Two OpenRouter control objects are commonly used through the gateway: * `provider` selects which upstream vendor serves the request. It accepts fields such as `order`, `allow_fallbacks`, and `require_parameters`. See [Provider Routing](https://openrouter.ai/docs/guides/routing/provider-selection). * `reasoning` controls chain-of-thought behavior with `enabled`, `effort`, and `max_tokens`, and the top-level `reasoning_effort` field is an alias for the effort setting. See [Reasoning Tokens](https://openrouter.ai/docs/guides/best-practices/reasoning-tokens). The following request pins the serving vendor and disables OpenRouter fallbacks while asking for high reasoning effort: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "openrouter-sonnet-prod", "messages": [ { "role": "user", "content": "Summarize the CAP theorem in two sentences." } ], "provider": { "order": ["anthropic"], "allow_fallbacks": false }, "reasoning": { "effort": "high" } }' ``` ### Read Reasoning Output[​](#read-reasoning-output "Direct link to Read Reasoning Output") When a model exposes reasoning text, OpenRouter returns it at `message.reasoning` on the non-streaming path and `delta.reasoning` on the streaming path, rather than the `reasoning_content` field that other OpenAI-compatible upstreams use. Some models do not expose their reasoning text. AISIX normalizes both into the canonical `reasoning_content` field, so clients read `choices[0].message.reasoning_content` and `choices[0].delta.reasoning_content` regardless of which spelling the upstream used. When an upstream sends both fields, `reasoning_content` takes precedence. This normalization is built into the `openai` adapter for OpenRouter. Do not set a [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) override on an OpenRouter provider key. That override exists for upstreams whose reasoning path AISIX does not already recognize. OpenRouter can also return a structured `reasoning_details` array that must be replayed unchanged for some multi-turn reasoning and tool-calling flows. The normalized AISIX chat path does not retain this array. Use an OpenRouter-native endpoint through a passthrough route when the application must preserve structured reasoning state. caution Provider and reasoning controls are OpenRouter-specific request fields. If you later repoint the alias at another provider, or place it behind a [multi-target routing model](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md), AISIX forwards those fields to whichever upstream serves the request, and a non-OpenRouter upstream may ignore or reject fields it does not recognize. ## Endpoint Support for OpenRouter Models[​](#endpoint-support-for-openrouter-models "Direct link to Endpoint Support for OpenRouter Models") An OpenRouter model alias does not reach every proxy route. Several routes gate on the configured provider value rather than on the adapter family. | Route | OpenRouter support | | ------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `/v1/chat/completions`, including `stream: true` | Supported through the `openai` adapter. | | `/v1/completions` | Supported for models that accept OpenRouter's legacy prompt-based completions format. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), which translates the request into the chat adapter path. It does not call OpenRouter's native [Responses API](https://openrouter.ai/docs/api/reference/responses/overview), so fields without a chat equivalent are ignored. Use `/passthrough/openrouter/responses` when native Responses semantics are required. | | `/v1/messages` | Supported through translation to chat completions, not OpenRouter's native [Anthropic Messages API](https://openrouter.ai/docs/api/api-reference/anthropic-messages/create-a-message). Fields and response blocks without a chat equivalent do not round-trip. Use `/passthrough/openrouter/messages` for the native contract. | | `/v1/messages/count_tokens` | Not available. This AISIX route requires an Anthropic-protocol provider key, while the OpenRouter integration uses the `openai` adapter. | | `/v1/embeddings` | Supported when the alias points at an OpenRouter embedding model. AISIX forwards the OpenAI embeddings request shape to `https://openrouter.ai/api/v1/embeddings`. | | `/v1/audio/speech` | Supported when the alias points at an OpenRouter speech model because the JSON request and upstream path are compatible. | | `/v1/audio/transcriptions` and `/v1/audio/translations` | Not supported through these normalized routes. AISIX accepts OpenAI-style multipart uploads, while OpenRouter's transcription endpoint accepts a JSON `input_audio` object and OpenRouter does not document a matching translations route. Use `/passthrough/openrouter/audio/transcriptions` with OpenRouter's native JSON shape. | | `/v1/images/generations` | Not available. The route accepts only models whose provider value is `openai`, and OpenRouter's native image endpoint is `/images`, not `/images/generations`. Use `/passthrough/openrouter/images` with the [OpenRouter image request](https://openrouter.ai/docs/guides/overview/multimodal/image-generation). | | `/v1/rerank` | Not available. The route accepts only the `openai`, `cohere`, and `jina` provider values. Use `/passthrough/openrouter/rerank` with an OpenRouter rerank model ID for OpenRouter's native route. | | `/v1/videos` and its status and content routes | Not available. The `openrouter` provider value returns `501 not_implemented`. For OpenRouter's native [asynchronous video API](https://openrouter.ai/docs/guides/overview/multimodal/video-generation), submit through `/passthrough/openrouter/videos`, then use the returned job ID with `/passthrough/openrouter/videos/{jobId}` and `/passthrough/openrouter/videos/{jobId}/content`. | | `/v1/files` | Do not use the normalized file routes for OpenRouter-native file references. AISIX wraps returned file IDs for its job-routing contract and does not unwrap those IDs inside native OpenRouter Messages or Responses bodies. Use `/passthrough/openrouter/files` to preserve OpenRouter's raw file IDs. | | `/v1/batches` and `/v1/fine_tuning/jobs` | Not available because OpenRouter does not publish matching routes. | | `/v1/models` | Returns caller-accessible AISIX model aliases, not OpenRouter's model catalog. Use `/passthrough/openrouter/models` for the native list. | | `/passthrough/openrouter/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). | For routes AISIX does not normalize, use a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). The `/passthrough/openrouter` paths on this page assume such a route claiming that prefix with `https://openrouter.ai/api/v1` as its `target_url`; grant the route on the caller key's `allowed_routes`. The gateway appends the remaining path to the route's `target_url` and removes one leading `v1` segment when it duplicates the trailing `/v1` of the target, so both `/passthrough/openrouter/models` and `/passthrough/openrouter/v1/models` reach `https://openrouter.ai/api/v1/models`. A passthrough route forwards the request body without model-alias rewriting. Send the namespaced OpenRouter model ID, such as `anthropic/claude-sonnet-5`, rather than `openrouter-sonnet-prod`. The route relays SSE responses incrementally. AISIX detects chat, completions, and Responses envelopes and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. Prefer normalized endpoints when their contract is sufficient and you need AISIX token and cost accounting. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to OpenRouter and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between OpenRouter and a direct vendor account. * [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md): reach OpenRouter routes that AISIX does not normalize. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Choose a Provider Upstream A provider upstream is the model service an AISIX gateway calls after resolving a caller-facing model alias. AISIX stores the upstream credential in a provider key and uses an adapter to send requests in the format that the upstream expects. Choose the setup path for the endpoint and account you use, not only for the model family. Configure Gemini through Google AI Studio when you have an AI Studio API key. Use Google Vertex AI when the model is hosted in a Google Cloud project. Every setup below supports both AISIX Cloud and the open-source AISIX gateway. Each provider guide contains complete configuration sections for both products followed by one shared verification procedure. ## Setup Paths[​](#setup-paths "Direct link to Setup Paths") AISIX supports provider setups across direct integrations, additional video and rerank routes, community catalog entries, and general OpenAI-compatible endpoints. ### Direct Providers and Managed Platforms[​](#direct-providers-and-managed-platforms "Direct link to Direct Providers and Managed Platforms") Choose among the direct provider and managed-platform setups: | Setup | Use When | Adapter | | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | -------------- | | [OpenAI](https://docs.api7.ai/ai-gateway/providers/openai.md) | You call GPT chat models, embeddings, images, audio, or Sora video generation through an OpenAI API account. | `openai` | | [Azure OpenAI](https://docs.api7.ai/ai-gateway/providers/azure-openai.md) | An Azure OpenAI deployment hosts the model. | `azure-openai` | | [Anthropic](https://docs.api7.ai/ai-gateway/providers/anthropic.md) | You call Claude through the Anthropic API. | `anthropic` | | [AWS Bedrock](https://docs.api7.ai/ai-gateway/providers/aws-bedrock.md) | AWS Bedrock hosts the model or inference profile. | `bedrock` | | [Gemini (Google AI Studio)](https://docs.api7.ai/ai-gateway/providers/gemini.md) | You call Gemini with a Google AI Studio API key. | `openai` | | [Google Vertex AI](https://docs.api7.ai/ai-gateway/providers/google-vertex-ai.md) | A Google Cloud project hosts Gemini or another Vertex publisher model. | `vertex` | | [Qwen (Alibaba Cloud)](https://docs.api7.ai/ai-gateway/providers/qwen.md) | You call Qwen chat models or Wan video generation through Alibaba Cloud Model Studio. | `openai` | | [DeepSeek](https://docs.api7.ai/ai-gateway/providers/deepseek.md) | You call models through the DeepSeek API. | `openai` | | [Groq](https://docs.api7.ai/ai-gateway/providers/groq.md) | You call Groq-hosted models. | `openai` | | [Mistral AI](https://docs.api7.ai/ai-gateway/providers/mistral.md) | You call models through the Mistral API. | `openai` | | [Together AI](https://docs.api7.ai/ai-gateway/providers/together.md) | You call models from the Together AI catalog. | `openai` | | [OpenRouter](https://docs.api7.ai/ai-gateway/providers/openrouter.md) | You call models from many vendors through one OpenRouter account. | `openai` | | [Fireworks AI](https://docs.api7.ai/ai-gateway/providers/fireworks-ai.md) | You call Fireworks-hosted models. | `openai` | | [Perplexity](https://docs.api7.ai/ai-gateway/providers/perplexity.md) | You call the search-grounded Sonar models. | `openai` | | [Cohere](https://docs.api7.ai/ai-gateway/providers/cohere.md) | You call Cohere Command models, embeddings, or rerank. | `openai` | | [Cerebras](https://docs.api7.ai/ai-gateway/providers/cerebras.md) | You call open-weight models on Cerebras inference hardware. | `openai` | | [Moonshot AI (Kimi)](https://docs.api7.ai/ai-gateway/providers/moonshotai.md) | You call Kimi models through the Moonshot AI platform. | `openai` | | [Zhipu AI (GLM)](https://docs.api7.ai/ai-gateway/providers/zhipuai.md) | You call GLM chat models or CogVideoX video generation. | `openai` | | [Hugging Face](https://docs.api7.ai/ai-gateway/providers/huggingface.md) | You call open-weight models through the Hugging Face Inference Providers router. | `openai` | | [Baseten](https://docs.api7.ai/ai-gateway/providers/baseten.md) | You call Baseten Model APIs or a dedicated Baseten deployment. | `openai` | ### Additional Video and Rerank Setups[​](#additional-video-and-rerank-setups "Direct link to Additional Video and Rerank Setups") Three more setups target providers whose video or rerank wire the gateway implements natively. The `openai` adapter shapes their chat-style and embeddings traffic, while the [video-generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md) and [rerank](https://docs.api7.ai/ai-gateway/endpoints/rerank.md) routes dispatch on the provider value itself. Catalog coverage varies by provider and model. For endpoints that report billable usage, configure an organization-level override through [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md) when AISIX Cloud has no matching catalog price. See each provider page for model selection and endpoint-specific accounting: | Setup | Use When | Adapter | | -------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- | -------- | | [Jina](https://docs.api7.ai/ai-gateway/providers/jina.md) | You call Jina embedding models, or reranker models through the rerank route. | `openai` | | [RunwayML](https://docs.api7.ai/ai-gateway/providers/runwayml.md) | You run Runway Gen or Runway-hosted Veo video tasks through the video route. | `openai` | | [Volcengine Ark (Doubao)](https://docs.api7.ai/ai-gateway/providers/volcengine-ark.md) | You call Doubao chat models or Seedance video generation through Volcengine Ark. | `openai` | ### Community Catalog Providers[​](#community-catalog-providers "Direct link to Community Catalog Providers") AISIX Cloud can also create provider keys for community catalog providers from models.dev. These providers use the `openai` adapter by default, but AISIX does not curate provider-specific request rewrites, response rewrites, or base URL quirks for them. The open-source gateway can use the same provider identifiers as labels, but you must set the adapter and base URL explicitly in `resources.yaml`. Use these pages when the provider is available in the catalog and you want a provider-specific setup guide: | Setup | Use When | Adapter | | ------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ | -------- | | [Amazon Nova API](https://docs.api7.ai/ai-gateway/providers/amazon-nova.md) | You call Nova models through the direct API-key-based Nova endpoint rather than Bedrock. | `openai` | | [Cloudflare Workers AI](https://docs.api7.ai/ai-gateway/providers/cloudflare-workers-ai.md) | You call account-scoped Cloudflare Workers AI model IDs such as `@cf/openai/gpt-oss-120b`. | `openai` | | [Databricks](https://docs.api7.ai/ai-gateway/providers/databricks.md) | You call Databricks Model Serving endpoints in your workspace. | `openai` | | [DeepInfra](https://docs.api7.ai/ai-gateway/providers/deepinfra.md) | You call DeepInfra-hosted open-weight models through its OpenAI-compatible API. | `openai` | | [DigitalOcean Gradient AI](https://docs.api7.ai/ai-gateway/providers/digitalocean-gradient-ai.md) | You call foundation models through DigitalOcean Serverless Inference. | `openai` | | [Meta Llama API](https://docs.api7.ai/ai-gateway/providers/meta-llama-api.md) | You call Llama models through Meta's direct OpenAI-compatible API. | `openai` | | [MiniMax](https://docs.api7.ai/ai-gateway/providers/minimax.md) | You call MiniMax models through the OpenAI-compatible API root. | `openai` | | [ModelScope](https://docs.api7.ai/ai-gateway/providers/modelscope.md) | You call ModelScope API-Inference model IDs. | `openai` | | [Nebius Token Factory](https://docs.api7.ai/ai-gateway/providers/nebius-token-factory.md) | You call open models hosted by Nebius Token Factory. | `openai` | | [NVIDIA NIM](https://docs.api7.ai/ai-gateway/providers/nvidia-nim.md) | You call the NVIDIA hosted NIM API or NVIDIA-published model IDs. | `openai` | | [Novita AI](https://docs.api7.ai/ai-gateway/providers/novita-ai.md) | You call Novita-hosted open models through its LLM API. | `openai` | | [OVHcloud AI Endpoints](https://docs.api7.ai/ai-gateway/providers/ovhcloud-ai-endpoints.md) | You call models through the OpenAI-compatible endpoint from OVHcloud. | `openai` | | [SiliconFlow](https://docs.api7.ai/ai-gateway/providers/siliconflow.md) | You call SiliconFlow-hosted org-namespaced model IDs. | `openai` | | [Snowflake Cortex](https://docs.api7.ai/ai-gateway/providers/snowflake-cortex.md) | You call models available in a Snowflake account through the Cortex REST API. | `openai` | | [Weights & Biases Inference](https://docs.api7.ai/ai-gateway/providers/wandb-inference.md) | You call open models through W\&B Inference. | `openai` | | [xAI](https://docs.api7.ai/ai-gateway/providers/xai.md) | You call Grok models through the OpenAI-compatible API from xAI. | `openai` | ### Other OpenAI-Compatible Endpoints[​](#other-openai-compatible-endpoints "Direct link to Other OpenAI-Compatible Endpoints") For another public or private endpoint that accepts OpenAI-compatible chat-completions requests, choose the corresponding general setup: | Setup | Use When | Adapter | | ----------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------- | -------- | | [Other OpenAI-Compatible Providers](https://docs.api7.ai/ai-gateway/providers/openai-compatible-vendors.md) | A public provider exposes an OpenAI-compatible API and does not have a dedicated setup guide. | `openai` | | [Ollama](https://docs.api7.ai/ai-gateway/providers/ollama.md) | You run models on a local or private Ollama server. | `openai` | | [vLLM](https://docs.api7.ai/ai-gateway/providers/vllm.md) | You serve a private model with the OpenAI-compatible server provided by vLLM. | `openai` | | [Bring Your Own Endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md) | You operate another private OpenAI-compatible endpoint, such as SGLang or an internal proxy. | `openai` | Each setup creates the same gateway resources: a provider key, a model alias, and a caller API key. The setup page provides the credential, upstream model identifier, base URL, and adapter details for that provider. Compare route support in [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md), or review the upstream request formats in [Adapter Protocol Families](https://docs.api7.ai/ai-gateway/providers/adapters.md). --- # OVHcloud AI Endpoints [OVHcloud AI Endpoints](https://docs.ovhcloud.com/en/guides/public-cloud/ai-machine-learning/ai-endpoints-getting-started) provides managed inference endpoints for hosted models. AISIX gives applications caller keys and stable model aliases for accessing those endpoints. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An OVHcloud AI Endpoints access token. * Access to the selected model. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` OVHcloud is a community catalog provider with an OpenAI-compatible endpoint. AISIX connects through the `openai` adapter and authenticates upstream requests with a bearer token. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") ``` export OVHCLOUD_API_KEY="YOUR_OVHCLOUD_ACCESS_TOKEN" PROVIDER_KEY_ID=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "ovhcloud-prod", "provider": "ovhcloud", "api_key": "'"${OVHCLOUD_API_KEY}"'", "api_base": "https://oai.endpoints.kepler.ai.cloud.ovh.net/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id' ) echo "$PROVIDER_KEY_ID" ``` Use the catalog ID `ovhcloud` and omit `adapter`. AISIX derives the OpenAI wire format and bearer authentication. The API base is the shared OpenAI-compatible root documented by OVHcloud. It includes `/v1`; AISIX adds `/chat/completions`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Create an alias with the exact OVHcloud model ID: ``` MODEL_ID=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "ovhcloud-llama-prod", "model_name": "Meta-Llama-3_3-70B-Instruct", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -er '.model.id' ) echo "$MODEL_ID" ``` OVHcloud model IDs do not necessarily match the publisher's original repository ID. Copy the value shown by OVHcloud, including underscores, hyphens, and casing. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "ovhcloud-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export OVHCLOUD_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "ovhcloud-prod" provider: "ovhcloud" adapter: "openai" api_key: ${OVHCLOUD_API_KEY} api_base: "https://oai.endpoints.kepler.ai.cloud.ovh.net/v1" models: - display_name: "ovhcloud-llama-prod" provider: "ovhcloud" model_name: "Meta-Llama-3_3-70B-Instruct" provider_key: "ovhcloud-prod" api_keys: - display_name: "ovhcloud-caller" key_env: CALLER_API_KEY allowed_models: - "ovhcloud-llama-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "ovhcloud-llama-prod", "messages": [ { "role": "user", "content": "Say hello from OVHcloud AI Endpoints." } ] }' ``` AISIX sends the OVHcloud model ID to the OpenAI-compatible chat endpoint with bearer authentication. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") An OVHcloud model alias does not reach every proxy route. AISIX uses the `openai` adapter for the normalized inference routes, but some routes also restrict the configured provider value. | Route | OVHcloud support | | -------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions`, including `stream: true` | Supported. Tool calling, structured output, vision, and other capabilities depend on the selected model. | | `/v1/responses` | Supported through the chat-based [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), not OVHcloud's native [Responses API](https://docs.ovhcloud.com/en/guides/public-cloud/ai-machine-learning/ai-endpoints-responses-api). Fields without a chat equivalent are ignored. Use `/passthrough/ovhcloud/responses` when the native contract is required. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation to chat completions. `/v1/messages/count_tokens` is unavailable because it requires an Anthropic-protocol provider key. | | `/v1/embeddings` | Supported when the alias names an OVHcloud embedding model, such as `bge-m3`. | | `/v1/audio/transcriptions` | Supported when the alias names an OVHcloud speech-to-text model, such as `whisper-large-v3`. AISIX rewrites the model alias and forwards the OpenAI-compatible multipart form to OVHcloud's [transcription endpoint](https://docs.ovhcloud.com/en/guides/public-cloud/ai-machine-learning/ai-endpoints-audio-models). | | `/v1/audio/translations` and `/v1/audio/speech` | Not available on the shared OVHcloud OpenAI-compatible base. OVHcloud handles translation through the transcription prompt and publishes [text-to-speech models](https://docs.ovhcloud.com/en/guides/public-cloud/ai-machine-learning/ai-endpoints-voice-virtual-assistant) on separate native endpoints. | | `/v1/images/generations`, `/v1/videos`, and `/v1/rerank` | Not available for an alias whose provider value is `ovhcloud`; these AISIX routes use provider allowlists that exclude it. | | `/v1/models` | Returns caller-accessible AISIX aliases, not the OVHcloud model catalog. Use `/passthrough/ovhcloud/models` for OVHcloud's current list. | | `/passthrough/ovhcloud/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for OVHcloud-native routes, with limited gateway normalization. | The `/passthrough/ovhcloud` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with the OVHcloud API base as its `target_url`; grant the route name on the caller key's `allowed_routes`. The normalized `/v1/responses` route is usually preferable when its translated feature set is sufficient: AISIX rewrites the model alias, streams incrementally, and parses usage. For native OVHcloud Responses behavior, passthrough forwards the body unchanged, so send the upstream model ID rather than the AISIX alias. OVHcloud currently requires `store: false` and client-managed conversation history on that API. A passthrough route relays the upstream response incrementally, including SSE. AISIX detects the Responses envelope from `input` and records token usage when the response includes supported `input_tokens` and `output_tokens` fields. See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the gateway-wide endpoint rules. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Check | | ----------------------------------- | ----------------------------------------------------------------- | | Upstream `401` or `403` | Confirm the access token and project entitlement. | | Upstream `404` | Keep `/v1` in `api_base` and copy the OVHcloud model ID exactly. | | Model not found | Check underscores, hyphens, and version segments in `model_name`. | | Provider key creation returns `400` | Use `provider: "ovhcloud"` and omit `adapter`. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to OVHcloud AI Endpoints and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): route this alias alongside another European endpoint. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Perplexity [Perplexity Sonar](https://docs.perplexity.ai) provides search-grounded models that combine language-model responses with web retrieval. Applications call Sonar through stable AISIX aliases while the gateway holds the Perplexity credential. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Perplexity API key. See the [Perplexity quickstart](https://docs.perplexity.ai/docs/getting-started/quickstart) for how to create one. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Perplexity-backed chat-completions route. Because Perplexity exposes an OpenAI-compatible chat-completions API, AISIX connects through the `openai` adapter and uses the Perplexity API root as `api_base`. AISIX registers no Perplexity-specific request or response rewrites. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Perplexity credential and API root: ``` # Replace with your values export PERPLEXITY_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "perplexity-prod", "provider": "perplexity", "api_key": "'"${PERPLEXITY_API_KEY}"'", "api_base": "https://api.perplexity.ai", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `perplexity`. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the Perplexity API key. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is the bare host `https://api.perplexity.ai`, the base URL Perplexity documents for OpenAI SDK clients. Its OpenAI-compatible chat-completions route is served at `/chat/completions` with no version segment, so the base carries no version suffix. AISIX appends the endpoint path, producing `https://api.perplexity.ai/chat/completions`. `api_base` is optional for this provider. When it is omitted, the AISIX Cloud Admin API fills in `https://api.perplexity.ai`. Setting it explicitly keeps the upstream root visible on the resource. The value must never end up empty: for an `openai`-adapter provider other than `openai`, AISIX refuses to fall back to `api.openai.com` and returns an upstream configuration error instead, so a Perplexity credential is never sent to the public OpenAI host. AISIX strips a trailing endpoint segment from `api_base`, so pasting the full documented endpoint URL `https://api.perplexity.ai/chat/completions` resolves to the same base. A trailing slash is also removed. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Perplexity model IDs name a search tier rather than a parameter count or a release date. They are bare names with no vendor namespace and no dated snapshot suffix: `sonar` for fast grounded answers, `sonar-pro` for broader retrieval and stronger synthesis, `sonar-reasoning-pro` for cited multi-step reasoning, and `sonar-deep-research` for long-running research reports. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "perplexity-sonar-pro-prod", "model_name": "sonar-pro", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Perplexity model ID, for example `sonar-pro` or `sonar`. ❸ `provider_key_id` attaches the alias to the Perplexity provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The gateway generates the key value; the plaintext is returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "perplexity-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value references the model by its ID, so the key can only access the alias you created. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export PERPLEXITY_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "perplexity-prod" provider: "perplexity" adapter: "openai" api_key: ${PERPLEXITY_API_KEY} api_base: "https://api.perplexity.ai" models: - display_name: "perplexity-sonar-pro-prod" provider: "perplexity" model_name: "sonar-pro" provider_key: "perplexity-prod" api_keys: - display_name: "perplexity-caller" key_env: CALLER_API_KEY allowed_models: - "perplexity-sonar-pro-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy. Sonar answers from live web search, so ask something that requires current information: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "perplexity-sonar-pro-prod", "messages": [ { "role": "user", "content": "Summarize the latest news about AI gateways in two sentences." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `perplexity-sonar-pro-prod`. If the request fails, check the provider key `api_key`, `api_base`, and the Perplexity model ID in `model_name`. ## Retrieve Search Citations[​](#retrieve-search-citations "Direct link to Retrieve Search Citations") Sonar responses carry their sources in `citations` and `search_results` at the top level of the Perplexity response body, alongside `choices`. Those fields are Perplexity extensions rather than part of the OpenAI chat-completions shape. AISIX normalizes chat-completions responses into the canonical OpenAI shape, so the two fields do not reach the caller on `/v1/chat/completions`. When an application needs the source list, call the provider-native route instead. The `/passthrough/perplexity` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with `https://api.perplexity.ai` as its `target_url`; grant the route name on the caller key's `allowed_routes`. AISIX forwards the body verbatim, injects the provider key credential bound to the route, and returns the upstream response unchanged: ``` curl -sS -X POST "$AISIX_PROXY/passthrough/perplexity/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "sonar-pro", "messages": [ { "role": "user", "content": "Summarize the latest news about AI gateways in two sentences." } ] }' ``` The `perplexity` path segment is the route's claimed prefix — by convention the provider name — not the alias, and the request body carries the upstream model ID `sonar-pro` rather than the alias. AISIX still authenticates the caller key, which must grant the route name in its `allowed_routes` list. note The gateway detects the request's chat envelope (`messages`) automatically, so supported usage figures from the Sonar response are recorded on the usage event best-effort. Passthrough tokens are telemetry only: they do not advance `tpm` or `tpd` counters, resolve a model cost, or add budget spend. See [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). ## Set the Reasoning Effort[​](#set-the-reasoning-effort "Direct link to Set the Reasoning Effort") The Sonar request schema exposes a `reasoning_effort` field with the values `minimal`, `low`, `medium`, and `high`, but model-specific documentation can narrow the accepted values. For example, Perplexity currently documents `low`, `medium`, and `high` for [`sonar-deep-research`](https://docs.perplexity.ai/docs/sonar/models/sonar-deep-research). AISIX forwards request fields it does not model to the upstream unchanged, so add the field to the chat-completions body. The following example assumes a second alias, `perplexity-sonar-research-prod`, created with `model_name` set to `sonar-deep-research`: ``` { "model": "perplexity-sonar-research-prod", "messages": [{ "role": "user", "content": "Compare the two proposals." }], "reasoning_effort": "high" } ``` AISIX registers no `response.reasoning_field` mapping for Perplexity, so it does not lift a Perplexity-specific reasoning field into the canonical `reasoning_content` slot. Reasoning output reaches the caller wherever Perplexity places it in the standard response content. If a future Perplexity model streams reasoning on a separate `delta` path, set [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) on the provider key. ## Endpoint and Cost Boundaries[​](#endpoint-and-cost-boundaries "Direct link to Endpoint and Cost Boundaries") Perplexity now publishes separate Sonar, Agent, Search, and Embeddings APIs. Their path prefixes differ, so the Sonar provider key configured in this guide cannot serve every API through a normalized AISIX route. | Route | Behavior with a Perplexity alias | | -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including `stream: true`. | | `/v1/completions` | Not available. Perplexity does not publish the legacy prompt-completions route. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over Sonar chat completions, not Perplexity's native [Agent API](https://docs.perplexity.ai/docs/agent-api/openai-compatibility). Fields without a chat equivalent are ignored. Use `/passthrough/perplexity/v1/responses` for the native Responses-compatible contract and send an Agent API model ID or preset rather than an AISIX alias. | | `/v1/messages` | Compatible Anthropic-shaped requests are supported through translation to Sonar chat completions. `/v1/messages/count_tokens` is unavailable because it requires an Anthropic-protocol provider key. | | `/v1/embeddings` | Not with the Sonar provider key above: AISIX would append `/embeddings` to the bare host, while Perplexity publishes [standard embeddings](https://docs.perplexity.ai/docs/embeddings/quickstart) at `/v1/embeddings`. Create a separate provider key with `api_base: https://api.perplexity.ai/v1` and an alias for `pplx-embed-v1-0.6b` or `pplx-embed-v1-4b`. | | `/v1/images/generations` | Rejected. The route accepts only models whose configured provider is `openai`. | | `/v1/rerank` | Rejected. The route allowlist is `openai`, `cohere`, and `jina`. | | `/v1/videos` | Rejected. The route allowlist does not include `perplexity`. | | `/v1/models` | Returns caller-accessible AISIX aliases. Use `/passthrough/perplexity/v1/models` for the current Agent API model list. | | `/passthrough/perplexity/search` | Calls Perplexity's native [Search API](https://docs.perplexity.ai/api-reference/search-post) and returns raw search results. | | `/passthrough/perplexity/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md). Use it for provider-native request and response shapes. | Perplexity's standard embedding models return quantized base64 strings rather than float arrays. AISIX preserves that response representation. Callers must decode `base64_int8` or `base64_binary` and use the [similarity metric documented by Perplexity](https://docs.perplexity.ai/docs/embeddings/best-practices). Two further boundaries affect how you model Perplexity traffic: * Sonar does not expose caller-defined function tools. [Sonar Pro Search](https://docs.perplexity.ai/docs/sonar/pro-search/tools) can invoke Perplexity-managed search and URL-fetching tools, while the Agent API provides the broader tool contract. Route custom function-calling workflows through a compatible Agent API model or another provider. * AISIX [cost metadata](https://docs.api7.ai/ai-gateway/models/model-aliases.md#cost-metadata) is expressed as USD per 1,000 input and output tokens. Perplexity bills `sonar-deep-research` for citation tokens, search queries, and reasoning tokens in addition to prompt and completion tokens, so a token-only estimate understates the real spend for that model. Treat gateway cost figures for `sonar-deep-research` as a lower bound and reconcile against Perplexity billing. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Perplexity and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Perplexity and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Qwen (Alibaba Cloud) [Qwen](https://www.alibabacloud.com/help/en/model-studio/) is Alibaba Cloud's family of language and multimodal models, served through Model Studio, also known as DashScope. Applications call Qwen through stable AISIX aliases while the gateway holds the DashScope credential. This guide covers both Qwen chat traffic and Wan video tasks through a single Model Studio provider key. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A DashScope API key from [Alibaba Cloud Model Studio](https://bailian.console.alibabacloud.com/) for the region you plan to use. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Qwen-backed chat-completions route. The examples use the Singapore regional endpoint. Because Alibaba Cloud Model Studio exposes an OpenAI-compatible endpoint, AISIX connects through the `openai` adapter and uses the DashScope API root for the region where you created the credential. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") DashScope API keys and endpoints are region-specific. The AISIX catalog has separate international and mainland China provider IDs, with these default API roots: | Provider ID | Default API root | Scope | | ------------ | -------------------------------------------------------- | ------------------------------------------- | | `alibaba` | `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` | International, using the Singapore endpoint | | `alibaba-cn` | `https://dashscope.aliyuncs.com/compatible-mode/v1` | Mainland China, using the Beijing endpoint | An API key issued for one region does not work with another region's endpoint. Create one provider key per region when routing to both. Use `alibaba` for international traffic and `alibaba-cn` for mainland China so usage records and cost reports attribute traffic to the matching catalog. Alibaba Cloud recommends a [workspace-dedicated domain](https://www.alibabacloud.com/help/en/model-studio/regions/) for production. The shared DashScope roots above remain available for existing integrations, but a dedicated domain provides workspace isolation and higher concurrency. The example below uses the Singapore form; replace `YOUR_WORKSPACE_ID` with the workspace that issued the API key. For another region, use the matching domain from the Model Studio console. Create the provider key that stores the DashScope credential and API root, and capture its ID: ``` # Replace with your value export DASHSCOPE_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "qwen-prod", "provider": "alibaba", "api_key": "'"${DASHSCOPE_API_KEY}"'", "api_base": "https://YOUR_WORKSPACE_ID.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') ``` ❶ `provider` is `alibaba`, the catalog provider ID for the international Model Studio region. Use `alibaba-cn` for the mainland China region. The AISIX Cloud Admin API derives the adapter from the catalog provider; the adapter field is only accepted on BYO provider keys. ❷ `api_key` stores the DashScope API key. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` already includes the `/compatible-mode/v1` path. Use the workspace and region that issued the API key. If you omit this field, AISIX Cloud uses the shared international root for `alibaba` and the shared Beijing root for `alibaba-cn`; those catalog defaults do not select your workspace-dedicated domain. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Create the model alias callers will send in requests, and capture its ID: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "qwen-plus-prod", "model_name": "qwen3.7-plus", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Qwen model ID. This example uses the current `qwen3.7-plus` model; check the [Model Studio model list](https://www.alibabacloud.com/help/en/model-studio/models) for availability in your region before creating the alias. ❸ `provider_key_id` attaches the alias to the Qwen provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key resource that can access the model alias. The plaintext key is generated by the server and returned once in the response: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "qwen-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') ``` The `allowed_models` value must reference the model ID you captured. Store the plaintext key securely; it cannot be retrieved again. The new resources project to attached gateways automatically, so the route is ready to call almost immediately. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export DASHSCOPE_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "qwen-prod" provider: "alibaba" adapter: "openai" api_key: ${DASHSCOPE_API_KEY} api_base: "https://YOUR_WORKSPACE_ID.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1" models: - display_name: "qwen-plus-prod" provider: "alibaba" model_name: "qwen3.7-plus" provider_key: "qwen-prod" api_keys: - display_name: "qwen-caller" key_env: CALLER_API_KEY allowed_models: - "qwen-plus-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "qwen-plus-prod", "messages": [ { "role": "user", "content": "Say hello from Qwen." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `qwen-plus-prod`. Confirm the request on the Model Studio usage page. If the request fails with an upstream authentication error, check the provider key `api_key` and confirm that the key and base URL use the same region. ## Generate Videos with Wan and HappyHorse[​](#generate-videos-with-wan "Direct link to Generate Videos with Wan and HappyHorse") The same provider key drives the gateway's modeled [video routes](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md). Create a second alias that points at a Model Studio text-to-video model — a [Wan 2.7](https://www.alibabacloud.com/help/en/model-studio/text-to-video-api-reference) model, or a [HappyHorse](https://www.alibabacloud.com/help/en/model-studio/happyhorse-text-to-video-api-reference) text-to-video model such as `happyhorse-1.1-t2v`. Both families answer on the same asynchronous DashScope endpoints, so the workflow below is identical; only the upstream model name changes. In AISIX Cloud, create the video alias and a caller key scoped to it: ``` VIDEO_MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "wan-video-prod", "model_name": "wan2.7-t2v", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$VIDEO_MODEL_ID" VIDEO_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "qwen-video-caller", "allowed_models": ["'"${VIDEO_MODEL_ID}"'"] }' | jq -r '.plaintext') ``` For the open-source AISIX gateway, add `wan-video-prod` to the existing `models` collection. Replace the existing `qwen-caller` entry with the updated entry below so it allows both model aliases. Preserve unrelated entries and collections: resources.yaml (video model access) ``` models: - display_name: "wan-video-prod" provider: "alibaba" model_name: "wan2.7-t2v" provider_key: "qwen-prod" api_keys: - display_name: "qwen-caller" key_env: CALLER_API_KEY allowed_models: - "qwen-plus-prod" - "wan-video-prod" ``` Validate and reload or restart the declarative resources file as described above, then use the existing caller key for the video request: ``` export VIDEO_API_KEY="$CALLER_API_KEY" ``` Submit a task: ``` curl -sS -X POST "$AISIX_PROXY/v1/videos" \ -H "Authorization: Bearer ${VIDEO_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "wan-video-prod", "prompt": "A paper boat drifts down a rain-soaked street at dusk.", "seconds": 5 }' ``` Six provider-specific behaviors are worth knowing before you script around this route: * **Only the `alibaba` provider value is dispatched.** `alibaba-cn` is outside the video route's allowlist, so a mainland China alias returns a not-implemented error on submit. Reach the mainland China video API through a passthrough route whose `target_url` is the native Model Studio base for that region. * **The chat `api_base` is reused as-is.** AISIX strips one `/compatible-mode/v1`, `/api/v1`, or `/v1` suffix, derives the vendor root, and composes DashScope's native task paths underneath. A provider key already configured for Qwen chat traffic works on the video routes without change, and the gateway sets the mandatory asynchronous submission header for you. * **`size` targets the earlier Wan protocol.** `seconds` is forwarded as the integer `parameters.duration`, and `size` is forwarded as `parameters.size` in the provider's `WIDTH*HEIGHT` spelling. That parameter belongs to Wan 2.6 and earlier. Wan 2.7 models replaced it with `resolution` and `ratio` tiers, and AISIX does not check which family the alias names — it forwards whatever you send. Omit `size` yourself on a Wan 2.7 alias, as the example above does, and let the provider default apply. * **HappyHorse text-to-video works on the same route.** `happyhorse-1.1-t2v` and `happyhorse-1.0-t2v` use the same submission endpoint, poll endpoint, and task states as Wan, and express output dimensions as `resolution` and `ratio` tiers — so, as with Wan 2.7, omit `size` and let the provider default apply. The image-to-video (`happyhorse-1.1-i2v`), reference-to-video, and video-editing variants are separate Model Studio APIs with reference-media inputs the modeled route does not carry; reach them through a passthrough route, like the other provider-native fields below. Runway also hosts HappyHorse models on its own platform — see [RunwayML](https://docs.api7.ai/ai-gateway/providers/runwayml.md) for that integration. * **Provider-native video fields require a passthrough route.** The normalized request models only `prompt`, `seconds`, and `size`; extra fields are ignored. To send native fields such as `negative_prompt`, `seed`, or Wan 2.7's `resolution` and `ratio`, use a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) pointed at the native DashScope root, which differs from the OpenAI-compatible root the `/passthrough/alibaba` examples on this page assume. A route claiming `/passthrough/alibaba-video` with `https://dashscope-intl.aliyuncs.com/api/v1` as its `target_url` submits at `/passthrough/alibaba-video/services/aigc/video-generation/video-synthesis` and polls at `/passthrough/alibaba-video/tasks/{id}`. Send the native body with the exact upstream model ID, and set the provider's asynchronous submission header yourself. Passthrough does not return the unified AISIX video object or gateway-encoded task ID. * **Finished videos are delivered by redirect.** `GET /v1/videos/{id}/content` returns `302` with a `Location` header pointing at the provider's signed download URL. The bytes travel from Model Studio storage to the client and do not pass through the gateway, so follow redirects when downloading, for example with the `-L` flag in `curl`. For the full submit, poll, and download workflow, including status semantics and rate-limit behavior on polling, see [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md). ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") A Qwen provider key uses the `openai` adapter, but route support also depends on AISIX provider rules and the APIs available on the configured Model Studio base: | Route | Behavior with a Qwen model alias | | ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `/v1/chat/completions` | Supported, buffered and streaming. | | `/v1/messages` | Supported through translation to chat completions. `/v1/messages/count_tokens` is limited to Anthropic-backed models. Alibaba's native Anthropic-compatible API uses a different `/apps/anthropic` base and therefore needs a separate provider key. | | `/v1/responses` | Supported through the AISIX [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), not Alibaba's [native Responses API](https://www.alibabacloud.com/help/en/model-studio/qwen-api-via-openai-responses). Fields without a chat-completions equivalent are ignored. To use native features such as `previous_response_id` or provider-hosted tools, call `/passthrough/alibaba/responses` with the exact upstream model ID in the body. | | `/v1/embeddings` | Supported when the alias names a [text embedding model](https://www.alibabacloud.com/help/en/model-studio/embedding) available in the same region, such as `text-embedding-v4`. AISIX appends `/embeddings` to the configured OpenAI-compatible base. | | `/v1/videos` and its status and content routes | Supported only when the model's provider value is exactly `alibaba`; `alibaba-cn` is outside the video route's allowlist. AISIX maps the request to DashScope's asynchronous [text-to-video API](https://www.alibabacloud.com/help/en/model-studio/text-to-video-api-reference), which serves both Wan and HappyHorse text-to-video models. For current Wan 2.7 and HappyHorse models, omit `size`, because the AISIX `size` mapping targets earlier Wan APIs. See [Generate Videos with Wan and HappyHorse](#generate-videos-with-wan). Use a passthrough route when you need native `resolution`, `ratio`, or multimodal fields; for `alibaba-cn`, this requires a route targeting the native Model Studio base. | | `/v1/audio/*` | Not supported on the OpenAI-compatible base configured on this page. Model Studio audio APIs use provider-native routes and request shapes. | | `/v1/images/generations` | Not supported. This route accepts only models whose configured provider is `openai`; Model Studio image APIs require a passthrough route. | | `/v1/rerank` | Not supported. This route accepts only the `openai`, `cohere`, and `jina` provider values; Model Studio rerank models use a provider-native API. | | `/passthrough/alibaba/*rest` and `/passthrough/alibaba-cn/*rest` | Available through configured [passthrough routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for native Model Studio APIs that AISIX has not modeled. A passthrough route does not rewrite a caller-facing alias inside the body and relays upstream SSE incrementally. Recognized chat, completions, and Responses envelopes record supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. Use a separate route when the native API has a different base URL. | The `/passthrough/alibaba` and `/passthrough/alibaba-cn` paths on this page assume passthrough routes claiming those prefixes, each with the matching Model Studio API root as its `target_url`; grant the route names on the caller key's `allowed_routes`. See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the full endpoint and provider matrix. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Qwen and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Qwen regions or between Qwen and another provider. * [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md): follow the submit, poll, and download workflow for supported Wan text-to-video models. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # RunwayML [RunwayML](https://docs.dev.runwayml.com/) provides APIs for generating videos with Runway Gen and Runway-hosted models. AISIX lets applications submit and manage those tasks through the gateway's video API while managing the Runway credential, caller access, rate limits, and usage accounting. This guide configures RunwayML for the AISIX [video-generation API](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md). Runway does not provide a chat-completions API, so use Runway-backed model aliases only for video tasks. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Runway API key from the [Runway developer portal](https://dev.runwayml.com/). * `curl` and `jq`. The developer portal is a separate surface from the `runwayml.com` web app. API keys exist only in the portal, and API credits are a separate pool from web-app credits — a web-app subscription does not fund API calls. See the [Runway API FAQs](https://help.runwayml.com/hc/en-us/articles/21668552945171-Runway-API-FAQs). ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Runway-backed video route. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Runway credential and API root: ``` # Replace with your value export RUNWAY_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "runwayml-prod", "provider": "runwayml", "api_key": "'"${RUNWAY_API_KEY}"'", "api_base": "https://api.dev.runwayml.com", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `runwayml`, the provider value the AISIX video routes recognize. The AISIX Cloud Admin API derives the `openai` adapter for this provider; the `adapter` field is only accepted on BYO provider keys. The derived adapter applies only to chat-style routes that Runway does not serve. ❷ `api_key` stores the Runway API key and is sent as a bearer token on upstream calls. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is the bare host. Runway's documented API base carries no version segment — the `/v1` belongs to the endpoint paths, so AISIX composes `/v1/text_to_video` and `/v1/tasks/{id}` onto this root as written. The field is optional for this provider — when it is omitted, the AISIX Cloud Admin API fills in the same value. Set it explicitly so the upstream root stays visible on the resource. Do not append `/v1`: AISIX strips no version suffix from this provider's base, so `https://api.dev.runwayml.com/v1` builds `/v1/v1/…` paths and fails upstream. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Runway's text-to-video endpoint — the endpoint the gateway submits to — currently serves `gen4.5`, `veo3.1`, `veo3.1_fast`, `veo3`, `happyhorse_1_0`, `seedance2`, `seedance2_fast`, `seedance2_mini`, and `gemini_omni_flash`. Other Runway model IDs, such as the image-to-video model `gen4_turbo`, are rejected by this endpoint, so an alias naming them fails at submission. Check the [Runway API reference](https://docs.dev.runwayml.com/api/#text-to-video) for the current catalog and each model's accepted parameters before creating a model alias. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "runway-video-prod", "model_name": "gen4.5", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Runway model ID, for example `gen4.5` or `veo3.1`. ❸ `provider_key_id` attaches the alias to the RunwayML provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "runwayml-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step, so the key can only access the alias you created. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export RUNWAY_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "runwayml-prod" provider: "runwayml" adapter: "openai" api_key: ${RUNWAY_API_KEY} api_base: "https://api.dev.runwayml.com" models: - display_name: "runway-video-prod" provider: "runwayml" model_name: "gen4.5" provider_key: "runwayml-prod" api_keys: - display_name: "runwayml-caller" key_env: CALLER_API_KEY allowed_models: - "runway-video-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Because this provider serves no chat surface, verify the connection with a video task. Submit it through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/videos" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "runway-video-prod", "prompt": "A lighthouse beam sweeps across a foggy harbor at night.", "seconds": 5, "size": "1280x720" }' ``` The gateway returns a video job object whose `status` is `queued` — Runway's create response confirms only that the task was accepted — and whose `id` is the gateway-issued video ID for the follow-up calls. Poll the task until it completes, then download the result: ``` curl -sS "$AISIX_PROXY/v1/videos/YOUR_VIDEO_ID" \ -H "Authorization: Bearer ${AISIX_API_KEY}" curl -sS -L -o video.mp4 \ "$AISIX_PROXY/v1/videos/YOUR_VIDEO_ID/content" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` If the submission fails, check the provider key `api_key`, the bare-host `api_base`, and the Runway model ID in `model_name`. Four provider-specific behaviors are worth knowing before you script around this route: * **The gateway sends the mandatory API version header itself.** Runway requires an `X-Runway-Version` header on every API call and rejects requests without it. AISIX sets `X-Runway-Version: 2024-11-06` — the [API version](https://docs.dev.runwayml.com/api-details/versions/2024-11-06/) its request and response handling is written against — on both the submit and the poll calls. The header is not operator-configurable and never appears in gateway configuration. * **`size` maps to Runway's `ratio` and is required in practice.** AISIX converts the unified `WIDTHxHEIGHT` value to Runway's `WIDTH:HEIGHT` resolution string by swapping the separator, so `1280x720` reaches Runway as `"1280:720"`. Runway validates the result against a per-model list of supported pixel resolutions, and its text-to-video endpoint requires `ratio`, so a request without `size` is rejected by the provider rather than defaulted by the gateway. `seconds` is forwarded as the provider's integer `duration`; accepted durations are also model-specific. Check the model's entry in the [Runway API documentation](https://docs.dev.runwayml.com/) for both lists. * **The normalized route models only common text-to-video fields.** AISIX sends the upstream model ID, `promptText`, and the optional `ratio` and `duration` mappings described above. Additional JSON fields are ignored, including Runway-specific controls such as negative prompts, generated-audio options, and reference media. To use those fields, call `/passthrough/runwayml/v1/text_to_video` with Runway's native request body, the exact upstream model ID, and the required `X-Runway-Version: 2024-11-06` header. The `/passthrough/runwayml` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with Runway's API root as its `target_url`; grant the route name on the caller key's `allowed_routes`. * **`THROTTLED` is a queued state.** AISIX maps Runway's `PENDING` and `THROTTLED` task states to `queued` — both mean the task is accepted but not yet running — with `RUNNING` reported as `in_progress`, `SUCCEEDED` as `completed`, and `FAILED` or `CANCELLED` as `failed`. A failed task carries Runway's machine-readable `failureCode` and human-readable `failure` text in the unified `error` object. * **Finished videos are delivered by redirect.** A completed Runway task reports its output as a list of signed URLs, and `GET /v1/videos/{id}/content` returns `302` with a `Location` header pointing at the first of them. The bytes travel from Runway's storage to the client and do not pass through the gateway, so use `curl -L` when downloading. For the full submit, poll, and download workflow, including status semantics and rate-limit behavior on polling, see [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md). ## Model IDs and Cost Metadata[​](#model-ids-and-cost-metadata "Direct link to Model IDs and Cost Metadata") RunwayML is not listed on models.dev, the public catalog AISIX Cloud draws model suggestions and pricing from. That has two consequences for this provider: * The dashboard suggests no model IDs. Take IDs from the [Runway API documentation](https://docs.dev.runwayml.com/) and enter them in `model_name` directly. * The pricing catalog carries no prices for Runway models. AISIX Cloud supports [pricing overrides](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md), but video submissions record zero tokens and do not apply duration-based cost to AISIX Cloud budgets, so an override does not price this workflow. See [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md#endpoint-behavior). ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") A RunwayML provider key exists for the video routes. Every chat-shaped route resolves a URL under the same bare host, and Runway serves none of them: | Route | Behavior with a `runwayml` model alias | | ---------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/videos` and its status and content routes | Supported. This is the only modeled route the provider serves. See [Verify the Provider Connection](#verify-the-provider-connection). | | `/v1/chat/completions` | Fails upstream. Runway publishes no chat-completions API; AISIX appends `/chat/completions` to the provider key base and the resulting path does not exist at the upstream. | | `/v1/embeddings` | Fails upstream. Runway publishes no embeddings API. | | `/v1/responses` | Fails upstream. The Responses bridge issues a chat-completions call, which Runway does not serve. | | `/v1/images/generations` | Not supported. The route accepts only models whose configured provider is `openai`. Reach Runway's image endpoints through a passthrough route instead. | | `/v1/rerank` | Not supported. The route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/passthrough/runwayml/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for Runway APIs and fields the gateway has not modeled, including image-to-video and provider-specific text-to-video controls. The route forwards caller headers and injects only the credential, so the caller must set `X-Runway-Version` itself — AISIX adds that header only on the modeled video routes. | See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the full endpoint and provider matrix. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to RunwayML and verified the model alias with a video task. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md): follow the full submit, poll, and download workflow. * [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md): reach Runway APIs the gateway has not modeled. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # SiliconFlow [SiliconFlow](https://docs.siliconflow.com/) provides hosted inference for models from multiple model organizations. AISIX gives applications stable aliases across that catalog while keeping provider credentials at the gateway. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A SiliconFlow API key from the SiliconFlow console. * `curl` and `jq`. ## Choose a SiliconFlow Platform[​](#choose-a-siliconflow-platform "Direct link to Choose a SiliconFlow Platform") SiliconFlow runs two platforms, and the model catalog carries one provider ID for each: | Catalog provider ID | API root | Use when | | ------------------- | -------------------------------- | --------------------------------------------------------- | | `siliconflow` | `https://api.siliconflow.com/v1` | The API key was issued on the `siliconflow.com` platform. | | `siliconflow-cn` | `https://api.siliconflow.cn/v1` | The API key was issued on the `siliconflow.cn` platform. | The catalog models the two platforms as separate providers with separate credential variables, `SILICONFLOW_API_KEY` and `SILICONFLOW_CN_API_KEY`, so treat them as separate accounts and pick the provider ID that matches where the key was issued. The examples below use `siliconflow`. If your account is on the other platform, substitute `siliconflow-cn` and `https://api.siliconflow.cn/v1` throughout. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the SiliconFlow-backed chat-completions route. SiliconFlow is a community catalog provider with an OpenAI-compatible API. AISIX connects through the `openai` adapter, authenticates upstream requests with a bearer token, and uses the SiliconFlow API root as `api_base`. AISIX does not register SiliconFlow-specific request or response rewrites. The dashboard labels this provider as a community entry whose wire format is unverified. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the SiliconFlow credential and API root: ``` # Replace with your value export SILICONFLOW_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "siliconflow-prod", "provider": "siliconflow", "api_key": "'"${SILICONFLOW_API_KEY}"'", "api_base": "https://api.siliconflow.com/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `siliconflow`. The AISIX Cloud Admin API accepts the value because the ID is in the models.dev catalog it caches, and it derives the adapter from the catalog. Do not send an `adapter` field: it is accepted only when `provider` is the `byo` sentinel, and sending it on a catalog provider key returns a 400 error. ❷ `api_key` stores the SiliconFlow API key. SiliconFlow authenticates with HTTP bearer authentication, which is what the `openai` adapter already sends, so no extra header configuration is needed. The value follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is `https://api.siliconflow.com/v1`. SiliconFlow documents the full chat endpoint as `POST https://api.siliconflow.com/v1/chat/completions`, so the root already includes `/v1`, and AISIX appends the endpoint path such as `/chat/completions` to it. This field is optional for `siliconflow`: models.dev publishes the same value as the provider's API field, and the AISIX Cloud Admin API fills it in when you omit it. Set it explicitly anyway, so that the root each key targets stays visible in the configuration and the key does not depend on a catalog snapshot that may predate a vendor URL change. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") SiliconFlow model IDs are org-namespaced. The ID is the model organization, a slash, and the model name, and both halves are case-sensitive: `deepseek-ai/DeepSeek-V3.2`, `zai-org/GLM-5.2`, `Qwen/Qwen3.6-27B`, `moonshotai/Kimi-K2.6`, and `openai/gpt-oss-120b` are current examples. Check the [SiliconFlow model catalog](https://cloud.siliconflow.com/models) for the current list before you create an alias, because the hosted set rotates as models are added and retired. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "siliconflow-deepseek-prod", "model_name": "deepseek-ai/DeepSeek-V3.2", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the SiliconFlow model ID, including the organization prefix. The prefix names the model organization, not the upstream provider. An alias for `openai/gpt-oss-120b` on SiliconFlow still has the provider value `siliconflow`, which is what the gateway's per-route provider rules evaluate. ❸ `provider_key_id` attaches the alias to the SiliconFlow provider key. The models.dev catalog carries per-token prices for the SiliconFlow chat models it lists, so usage and budget accounting resolve without extra configuration for those aliases. The catalog lists no SiliconFlow embedding or audio models, so add a pricing override for an embedding or transcription alias. For a transcription model billed by duration, configure its **Audio per minute** rate. Speech requests appear as zero-token usage events, but AISIX does not apply character-based text-to-speech pricing. See [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md) and [Cost Metadata](https://docs.api7.ai/ai-gateway/models/model-aliases.md#cost-metadata). ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "siliconflow-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export SILICONFLOW_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "siliconflow-prod" provider: "siliconflow" adapter: "openai" api_key: ${SILICONFLOW_API_KEY} api_base: "https://api.siliconflow.com/v1" models: - display_name: "siliconflow-deepseek-prod" provider: "siliconflow" model_name: "deepseek-ai/DeepSeek-V3.2" provider_key: "siliconflow-prod" api_keys: - display_name: "siliconflow-caller" key_env: CALLER_API_KEY allowed_models: - "siliconflow-deepseek-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "siliconflow-deepseek-prod", "messages": [ { "role": "user", "content": "Say hello from SiliconFlow." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `siliconflow-deepseek-prod`. Two failure modes separate a credential problem from a URL problem: * An upstream authentication error points at `api_key`, or at a key issued on the platform that does not match the configured provider ID. * An upstream 404 usually points at `api_base`. AISIX strips a pasted endpoint suffix such as `/chat/completions` and a trailing slash from `api_base`, but it does not add a missing `/v1` segment for a non-OpenAI host. A value of `https://api.siliconflow.com` therefore resolves to `https://api.siliconflow.com/chat/completions`, which is not a SiliconFlow route. If the request is rejected because the model does not exist, compare `model_name` against the SiliconFlow catalog, including the capitalization of both halves of the ID. ## Pass Reasoning Controls Through[​](#pass-reasoning-controls-through "Direct link to Pass Reasoning Controls Through") SiliconFlow puts its reasoning controls at the top level of the chat-completions body rather than inside a nested object: | Parameter | Type | Effect | | ----------------- | ------- | -------------------------------------------------------------------------------------------- | | `enable_thinking` | boolean | Switches a hybrid-reasoning model between thinking and non-thinking mode. | | `thinking_budget` | integer | Caps the tokens spent on the chain of thought. SiliconFlow documents the range 128 to 32768. | AISIX models only a fixed set of chat parameters, such as `temperature`, `top_p`, `max_tokens`, and `stream`. Every other top-level parameter is forwarded to the upstream verbatim, so both controls reach SiliconFlow unchanged: ``` { "model": "siliconflow-deepseek-prod", "messages": [ { "role": "user", "content": "Plan a three-step migration." } ], "enable_thinking": true, "thinking_budget": 4096 } ``` Support for these parameters is per model, not per provider. A value that one SiliconFlow-hosted model accepts can be rejected by another, so confirm the controls for the model you configured in the [SiliconFlow chat-completions reference](https://docs.siliconflow.com/en/api-reference/chat-completions/chat-completions). On the response side, SiliconFlow returns the chain of thought in `reasoning_content`, which is already the canonical field AISIX uses for both streaming and non-streaming responses. No response override is needed for the models that use it. If a specific model streams reasoning at a different `delta` path, set [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) on the provider key. ## Configure Wire Overrides[​](#configure-wire-overrides "Direct link to Configure Wire Overrides") Because SiliconFlow has no AISIX-curated adapter mapping, AISIX registers no provider-specific request or response rewrites for it. The gateway uses the standard OpenAI-compatible chat shape. When SiliconFlow renames a parameter or a hosted model diverges from that shape, configure the adjustment on the provider key. Provider-key overrides are inherited by every model that references the key. In AISIX Cloud, update the existing provider key in place. In the open-source AISIX gateway, update its entry in the declarative resources file. The Admin API replaces each supplied `request` or `response` block in full. The provider key created earlier has no overrides, so the `request` block below is complete. If you are updating a key that already has overrides, first retrieve its details and include every setting you want to keep in each supplied block. Omitting an entire block leaves that block unchanged; supplying an empty object clears it. An override affects every model alias that uses the provider key, whether you update it through AISIX Cloud or a resources file. If the key serves production traffic, first validate the same override on a separate provider key used only by a non-production alias. Apply the shared-key change during a controlled change window. ``` curl -sS -X PATCH "$AISIX_CP/provider_keys/$PROVIDER_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "request": { "param_renames": { "max_completion_tokens": "max_tokens" } } }' ``` This request omits `response` and the credential fields, so those settings remain unchanged, and models continue to reference the same provider key. For the open-source AISIX gateway, add the `request.param_renames` mapping shown below to the existing `siliconflow-prod` entry in `provider_keys`. Keep every other field on that entry, preserve every other entry and collection, and do not create a second top-level `provider_keys` key. resources.yaml (updated provider key) ``` provider_keys: - display_name: "siliconflow-prod" provider: "siliconflow" adapter: "openai" api_key: ${SILICONFLOW_API_KEY} api_base: "https://api.siliconflow.com/v1" request: param_renames: max_completion_tokens: max_tokens ``` Validate and reload or restart the declarative resources file as described above. `request.param_renames` renames a top-level parameter on the way upstream. Use it when clients send the current OpenAI name and the upstream expects the older one, or the reverse. If a request carries both names, AISIX keeps the value from the original caller-facing name. `response.reasoning_field` lifts reasoning from a nonstandard streaming `delta` path onto the canonical `delta.reasoning_content`. Set this only if a model actually diverges; SiliconFlow's documented field is already canonical. See [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) for the full field catalog. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") SiliconFlow serves several OpenAI-shaped inference routes, but the `siliconflow` provider value is outside the allowlists that some proxy routes enforce. The table below records what a SiliconFlow-backed alias can and cannot serve. | Route | Behavior with a SiliconFlow alias | | ---------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including `stream: true`. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over the chat adapter path. OpenAI-specific Responses fields without a chat equivalent are ignored. | | `/v1/messages` | Supported for Anthropic-shaped callers through translation to chat completions. This does not call SiliconFlow's native `/messages` route. Use `/passthrough/siliconflow/messages` with the exact upstream model ID when the native Messages request or response contract matters. Token counting at `/v1/messages/count_tokens` requires an Anthropic-backed model. | | `/v1/embeddings` | Supported when the alias names a SiliconFlow embedding model. The `openai` adapter forwards the OpenAI request shape to `{api_base}/embeddings`, which SiliconFlow serves on the same API root. See [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md). | | `/v1/audio/transcriptions` | Supported when the alias names a [SiliconFlow transcription model](https://docs.siliconflow.com/en/api-reference/audio/create-audio-transcriptions), such as `FunAudioLLM/SenseVoiceSmall` or `TeleAI/TeleSpeechASR`. AISIX rewrites the multipart `model` field to the upstream model ID and forwards the file to `{api_base}/audio/transcriptions`. | | `/v1/audio/speech` | Supported when the alias names a [SiliconFlow text-to-speech model](https://docs.siliconflow.com/en/api-reference/audio/create-speech), such as `FunAudioLLM/CosyVoice2-0.5B`. AISIX rewrites the JSON `model` field and returns the provider's binary audio response. SiliconFlow-specific fields such as `gain`, `sample_rate`, and `references` pass through unchanged. | | `/v1/audio/translations` | Fails upstream. SiliconFlow does not publish an audio-translation route on this API base. | | `/v1/rerank` | Rejected. The route accepts only the `openai`, `cohere`, and `jina` provider values, so a `siliconflow` alias is refused even though SiliconFlow hosts rerank models. Reach SiliconFlow reranking through a passthrough route instead. See [Rerank](https://docs.api7.ai/ai-gateway/endpoints/rerank.md). | | `/v1/images/generations` | Rejected. The route accepts only models whose provider is `openai`. | | `/v1/videos` | Rejected with `501 not_implemented`. The route dispatches on its own provider allowlist, which does not include `siliconflow`. | | `/passthrough/siliconflow/*` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for provider-native routes, with limited gateway normalization. Authorization comes from the caller key's `allowed_routes`, not its model allowlist. | The `/passthrough/siliconflow` paths on this page assume a passthrough route claiming that prefix with SiliconFlow's API root as its `target_url`; grant the route name on the caller key's `allowed_routes`. Passthrough is the practical way to reach SiliconFlow features that have no normalized gateway surface, including native Messages and rerank. Because the route's `target_url` already ends in `/v1`, AISIX drops a duplicated leading `/v1` from the passthrough path, so both `/passthrough/siliconflow/rerank` and `/passthrough/siliconflow/v1/rerank` resolve to the same upstream URL. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to SiliconFlow and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between SiliconFlow and a second provider, or rank targets by cost or latency. * [Speech and Audio](https://docs.api7.ai/ai-gateway/endpoints/audio.md): call a SiliconFlow transcription or text-to-speech alias through the normalized audio routes. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Snowflake Cortex [Snowflake Cortex](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-rest-api) brings hosted foundation models into Snowflake accounts and exposes them through REST APIs. AISIX gives applications one OpenAI-compatible API for the models available in your account. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Snowflake account identifier and a programmatic access token (PAT). * A Snowflake role allowed to use Cortex. Snowflake documents the `SNOWFLAKE.CORTEX_USER` database role and the REST API user role requirements. * Access to the selected model in your Snowflake region. * `curl` and `jq`. Export the Snowflake connection details: ``` export SNOWFLAKE_ACCOUNT="example-account" export SNOWFLAKE_API_BASE="https://${SNOWFLAKE_ACCOUNT}.snowflakecomputing.com/api/v2/cortex/v1" export SNOWFLAKE_PAT="YOUR_SNOWFLAKE_PROGRAMMATIC_ACCESS_TOKEN" ``` Use the account hostname from your Snowflake connection details. Do not include `https://` or another path in `SNOWFLAKE_ACCOUNT`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Snowflake is a community catalog provider with an OpenAI-compatible Cortex endpoint. AISIX connects through the `openai` adapter, authenticates upstream requests with a bearer token, and uses the account-specific hostname as `api_base`. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") ``` PROVIDER_KEY_ID=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "snowflake-cortex-prod", "provider": "snowflake-cortex", "api_key": "'"${SNOWFLAKE_PAT}"'", "api_base": "'"${SNOWFLAKE_API_BASE}"'", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id' ) echo "$PROVIDER_KEY_ID" ``` `snowflake-cortex` is the exact catalog ID. `snowflake` is not an alias for it. AISIX derives the `openai` adapter, so omit the `adapter` field. The account-specific `api_base` is required in practice because AISIX cannot infer which Snowflake account should receive the request. AISIX appends `/chat/completions` to this root. `AISIX_TOKEN` is an AISIX administrative token. `SNOWFLAKE_PAT` is the upstream credential stored in the provider key. They are unrelated token types even though both may be described as PATs. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Create an alias for a Cortex model available in your account: ``` MODEL_ID=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "snowflake-claude-prod", "model_name": "claude-sonnet-4-5", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -er '.model.id' ) echo "$MODEL_ID" ``` Snowflake's supported model list varies by region and release. Replace `claude-sonnet-4-5` only with an ID available to your account, and copy its spelling from the current [Cortex model availability reference](https://docs.snowflake.com/en/user-guide/snowflake-cortex/llm-functions#availability). ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "snowflake-cortex-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export SNOWFLAKE_PAT="YOUR_SNOWFLAKE_PROGRAMMATIC_ACCESS_TOKEN" export SNOWFLAKE_ACCOUNT="YOUR_SNOWFLAKE_ACCOUNT" export SNOWFLAKE_API_BASE="https://${SNOWFLAKE_ACCOUNT}.snowflakecomputing.com/api/v2/cortex/v1" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "snowflake-cortex-prod" provider: "snowflake-cortex" adapter: "openai" api_key: ${SNOWFLAKE_PAT} api_base: "${SNOWFLAKE_API_BASE}" models: - display_name: "snowflake-claude-prod" provider: "snowflake-cortex" model_name: "claude-sonnet-4-5" provider_key: "snowflake-cortex-prod" api_keys: - display_name: "snowflake-cortex-caller" key_env: CALLER_API_KEY allowed_models: - "snowflake-claude-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "snowflake-claude-prod", "messages": [ { "role": "user", "content": "Say hello from Snowflake Cortex." } ] }' ``` AISIX forwards the Snowflake model ID and authenticates with `Authorization: Bearer <SNOWFLAKE_PAT>`. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") Snowflake serves both OpenAI-compatible chat completions and a Claude-only Anthropic Messages API. The `snowflake-cortex` catalog entry uses the `openai` adapter, so normalized route behavior differs from Snowflake's native surface: | Route | Behavior with a Snowflake Cortex alias | | -------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, buffered and streaming, through Snowflake's OpenAI-compatible chat route. | | `/v1/responses` | Supported through the AISIX [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). Snowflake does not publish a native Responses endpoint, and fields without a chat-completions equivalent are ignored. | | `/v1/messages` | Supported through translation to chat completions, not through Snowflake's native Messages route. For a Claude alias that needs Snowflake's native Anthropic contract or beta features, call `/passthrough/snowflake-cortex/messages` with the exact upstream model ID and the required `anthropic-version: 2023-06-01` header. | | `/v1/messages/count_tokens` | Not supported. Token counting is limited to models whose configured adapter is Anthropic. | | `/v1/embeddings` | Not compatible with Snowflake's [Vector Embed REST API](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-rest-api/embed-api). Snowflake uses `POST /api/v2/cortex/inference:embed` and a native request body. Reach it through a separate [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) whose `target_url` is `https://<account>.snowflakecomputing.com/api/v2/cortex` — a route relays to one fixed base, so give this route its own path prefix — then call the native `inference:embed` path beneath that prefix, or use another embedding provider. | | `/v1/images/generations`, `/v1/videos`, and `/v1/rerank` | Not supported. These routes do not accept the `snowflake-cortex` provider value. | | `/passthrough/snowflake-cortex/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for Snowflake-native routes, with limited gateway normalization. Use a separate route with its own prefix when the native API needs a different base URL. | The `/passthrough/snowflake-cortex` paths on this page assume a passthrough route claiming that prefix with this guide's `api_base` as its `target_url`; grant the route name on the caller key's `allowed_routes`. See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the full endpoint and provider matrix. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Check | | ----------------------------------- | -------------------------------------------------------------------------------------------------- | | DNS error or upstream `404` | Confirm `SNOWFLAKE_ACCOUNT` produces the same hostname shown in your Snowflake connection details. | | Upstream `401` | Rotate the Snowflake PAT and update the provider key. | | Upstream `403` | Confirm the token's user and role have Cortex privileges and access to the model. | | Model unavailable | Choose a model supported in the Snowflake account's region. | | Provider key creation returns `400` | Use `provider: "snowflake-cortex"` and omit `adapter`. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Snowflake Cortex and verified the model alias. Continue with these guides: * [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md): manage account-specific credentials and key rotation. * [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md): reach Snowflake-native routes that AISIX does not normalize. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Together AI [Together AI](https://docs.together.ai/) provides hosted inference for open and partner models. Applications call selected models through stable AISIX aliases while the gateway holds the Together API key. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Together API key from the [Together console](https://api.together.ai/settings/api-keys). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Together-backed chat-completions route. Because Together AI exposes an OpenAI-compatible API, AISIX connects through the `openai` adapter and uses the Together API root as `api_base`. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Together credential and API root, and allow it into the environment: ``` # Replace with your value export TOGETHER_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "together-prod", "provider": "togetherai", "api_key": "'"${TOGETHER_API_KEY}"'", "api_base": "https://api.together.ai/v1", "allowed_environments": ["'"$ENV_ID"'"] }' | jq -r '.provider_key.id') ``` ❶ `provider` is `togetherai`, the catalog provider ID. ❷ `api_key` stores the Together API key. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` already includes the `/v1` path. AISIX appends `/chat/completions` to it. Together currently documents `https://api.together.ai/v1`; AISIX Cloud still fills in the equivalent `https://api.together.xyz/v1` root when the field is omitted. Set the official root explicitly so the target stays visible on the resource and does not depend on that fallback. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Together's catalog changes frequently. Look up the exact `<publisher>/<model>` ID in the [Together models list](https://docs.together.ai/docs/serverless/models) before creating a model alias. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "together-gptoss-prod", "model_name": "openai/gpt-oss-120b", "provider_key_id": "'"$PROVIDER_KEY_ID"'" }' | jq -r '.model.id') ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Together model ID in `<publisher>/<model>` form. A flat OpenAI-style name such as `gpt-4o` returns a 404 from Together. ❸ `provider_key_id` attaches the alias to the Together provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The API generates the key value and returns the plaintext once: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "together-caller", "allowed_models": ["'"$MODEL_ID"'"] }' | jq -r '.plaintext') ``` The `allowed_models` value references the model by its ID. The plaintext key is returned only in this response, so store it securely. The new resources project to the attached gateway automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export TOGETHER_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "together-prod" provider: "togetherai" adapter: "openai" api_key: ${TOGETHER_API_KEY} api_base: "https://api.together.ai/v1" models: - display_name: "together-gptoss-prod" provider: "togetherai" model_name: "openai/gpt-oss-120b" provider_key: "together-prod" api_keys: - display_name: "together-caller" key_env: CALLER_API_KEY allowed_models: - "together-gptoss-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "together-gptoss-prod", "messages": [ { "role": "user", "content": "Say hello from Together AI." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `together-gptoss-prod`. If the request fails, check the provider key `api_key`, `api_base`, and the `<publisher>/<model>` ID in `model_name`. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") Together exposes several OpenAI-shaped APIs on the same base, but normalized route support also depends on the model's `togetherai` provider value: | Route | Behavior with a Together AI alias | | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `/v1/chat/completions` | Supported, buffered and streaming where the selected model implements the route. | | `/v1/completions` | Supported for buffered responses where the selected model implements the route. Streaming is not available on this route — use `/v1/chat/completions` to stream. | | `/v1/responses` | Supported through the AISIX [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). Together [does not implement a native Responses API](https://docs.together.ai/docs/inference/openai-compatibility), and fields without a chat-completions equivalent are ignored. | | `/v1/messages` | Supported through translation to chat completions. `/v1/messages/count_tokens` is limited to Anthropic-backed models. | | `/v1/embeddings` | Supported when the alias names a [Together embedding model](https://docs.together.ai/docs/inference/embeddings/embeddings). AISIX rewrites the caller-facing alias and forwards the OpenAI-shaped request to `{api_base}/embeddings`. | | `/v1/audio/transcriptions`, `/v1/audio/translations`, and `/v1/audio/speech` | Supported when the alias names a Together model for the selected audio route. AISIX preserves the OpenAI request and response shapes while rewriting the model alias. | | `/v1/images/generations` | Not supported. The normalized route accepts only models whose configured provider is `openai`, even though Together serves the same path natively. Use `/passthrough/togetherai/images/generations` instead. | | `/v1/videos` | Not supported. The video route's provider allowlist does not include `togetherai`; use a passthrough route with Together's native video contract. | | `/v1/rerank` | Not supported. The normalized route accepts only the `openai`, `cohere`, and `jina` provider values. Use `/passthrough/togetherai/rerank` with a Together rerank model and native request body. | | `/passthrough/togetherai/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for Together-native routes, with limited gateway normalization. Passthrough does not rewrite a caller-facing alias inside the body and relays upstream SSE incrementally. Recognized chat, completions, and Responses envelopes record supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. | The `/passthrough/togetherai` paths on this page assume a passthrough route claiming that prefix with Together's API root as its `target_url`; grant the route name on the caller key's `allowed_routes`. See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the full endpoint and provider matrix. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Together AI and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Together and a second provider. * [Speech and Audio](https://docs.api7.ai/ai-gateway/endpoints/audio.md): call Together transcription, translation, or text-to-speech models through normalized audio routes. * [Passthrough Routes](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md): reach Together image, video, rerank, and other native routes. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # vLLM [vLLM](https://docs.vllm.ai/en/latest/serving/online_serving/openai_compatible_server/) is an open-source inference server that exposes an OpenAI-compatible API for models you run. AISIX adds caller authentication, stable model aliases, usage reporting, and traffic controls while vLLM continues to serve inference. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A host with vLLM installed and enough compute and storage for the selected model. * Network reachability from the AISIX gateway to vLLM. * `curl` and `jq`. ## Prepare the Inference Server[​](#prepare-the-inference-server "Direct link to Prepare the Inference Server") Start the official example model with API-key checking enabled: ``` vllm serve NousResearch/Meta-Llama-3-8B-Instruct \ --dtype auto \ --api-key token-abc123 ``` Export the upstream key and an API base reachable from the AISIX gateway: ``` export VLLM_API_KEY="token-abc123" export VLLM_API_BASE="http://vllm.internal:8000/v1" ``` Replace `vllm.internal` with a resolvable host or service name. For Docker Desktop with vLLM on the host, `http://host.docker.internal:8000/v1` is a typical value. If both services share a Docker or Kubernetes network, use the vLLM service DNS name. caution vLLM documents that `--api-key` protects its OpenAI-compatible API routes, not every endpoint the server may expose. Keep the vLLM service on a private network and apply network policy or an authenticated reverse proxy when administrative or diagnostic routes require protection. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` vLLM is a private endpoint rather than an AISIX catalog provider. Configure it with the `byo` provider value and select the `openai` adapter explicitly. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") ``` PROVIDER_KEY_ID=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "vllm-local", "provider": "byo", "adapter": "openai", "api_key": "'"${VLLM_API_KEY}"'", "api_base": "'"${VLLM_API_BASE}"'", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id' ) echo "$PROVIDER_KEY_ID" ``` A BYO key requires `provider: "byo"`, a non-empty `api_key`, and `api_base`. This guide sets `adapter: "openai"` explicitly; when omitted, AISIX defaults a BYO key to the OpenAI-compatible adapter. AISIX sends `VLLM_API_KEY` as `Authorization: Bearer <key>`, matching the key passed to `vllm serve`. If you intentionally run vLLM without `--api-key`, AISIX still requires a non-empty placeholder in the provider-key schema. An authenticated vLLM endpoint is preferred whenever traffic can originate outside a tightly controlled local network. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Use the model name served by vLLM: ``` MODEL_ID=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "vllm-llama-prod", "model_name": "NousResearch/Meta-Llama-3-8B-Instruct", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -er '.model.id' ) echo "$MODEL_ID" ``` `model_name` must match the name vLLM exposes from `/v1/models`. If you start vLLM with a served-model-name override, use that override rather than the model repository path. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "vllm-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export VLLM_API_KEY="token-abc123" export VLLM_API_BASE="http://vllm.internal:8000/v1" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "vllm-local" provider: "vllm" adapter: "openai" api_key: ${VLLM_API_KEY} api_base: "${VLLM_API_BASE}" models: - display_name: "vllm-llama-prod" provider: "vllm" model_name: "NousResearch/Meta-Llama-3-8B-Instruct" provider_key: "vllm-local" api_keys: - display_name: "vllm-caller" key_env: CALLER_API_KEY allowed_models: - "vllm-llama-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "vllm-llama-prod", "messages": [ { "role": "user", "content": "Say hello from vLLM." } ] }' ``` AISIX sends the served model name to `POST /v1/chat/completions` with the configured vLLM bearer key. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") The current vLLM [online server](https://docs.vllm.ai/en/latest/serving/online_serving/) exposes OpenAI-compatible generation, Responses, embeddings, and speech-to-text APIs. It also exposes Anthropic-compatible Messages and token-counting routes, plus pooling-model APIs such as rerank and score. Availability still depends on the model and task selected when the vLLM server starts. | Route | Behavior with the vLLM alias in this guide | | ------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported for a text-generation model with a chat template. | | `/v1/completions` | Supported for a text-generation model. vLLM does not support the OpenAI `suffix` parameter. | | `/v1/responses` | Supported through the AISIX [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), not the native vLLM Responses API. Fields without a chat-completions equivalent are ignored. Use `/passthrough/byo/responses` or `/passthrough/vllm/responses`, matching the prefix your passthrough route claims, when native Responses semantics are required. | | `/v1/messages` | Supported through AISIX translation to chat completions, not the native Anthropic-compatible vLLM route. Use `/passthrough/byo/messages` or `/passthrough/vllm/messages` for the native vLLM contract. | | `/v1/messages/count_tokens` | Rejected because the normalized route requires an Anthropic-backed model. Current vLLM releases implement this route natively; use your passthrough route's prefix followed by `/messages/count_tokens`. | | `/v1/embeddings` | Supported when the alias points to a vLLM server running an embedding model. | | `/v1/audio/transcriptions` and `/v1/audio/translations` | Supported when the alias points to a vLLM server running a compatible automatic speech recognition model with the vLLM audio dependencies installed. Translation support is model-dependent. | | `/v1/audio/speech` | Not supported because vLLM does not publish an OpenAI-compatible text-to-speech route. | | `/v1/realtime` | Supported when vLLM serves a Realtime-capable automatic speech recognition model and has its audio dependencies installed. AISIX relays the OpenAI-style vLLM WebSocket protocol through the normalized Realtime route. Current vLLM support is streaming speech-to-text, not the bidirectional speech-generation behavior available from some OpenAI Realtime models. | | `/v1/rerank` | Rejected because neither `byo` nor `vllm` is in the normalized route's provider allowlist. A vLLM server running a scoring model exposes `/v1/rerank`; reach it at `/passthrough/byo/rerank` or `/passthrough/vllm/rerank`. | | `/v1/chat/completions/batch` and `/v1/score` | AISIX has no normalized routes for these vLLM APIs. Reach them through your passthrough route's prefix followed by `/chat/completions/batch` or `/score`. | | `/v1/models` | Returns caller-accessible AISIX aliases, not the models served by vLLM. Use `/passthrough/byo/models` or `/passthrough/vllm/models` for the native vLLM list. | | `/v1/images/generations` and `/v1/videos` | Rejected because neither provider value is in the corresponding normalized route's allowlist. vLLM does not publish matching generation APIs. | Use a separate model alias, and normally a separate provider key and serving process, for each generation, embedding, or speech-to-text model. The same requirement applies to a reranking model reached through a passthrough route. The `/passthrough/byo` and `/passthrough/vllm` prefixes above assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming the prefix with the vLLM server root as its `target_url`; grant the route name on the caller key's `allowed_routes`. A passthrough route relays the request body unchanged, so use the model identifier exposed by vLLM rather than the AISIX alias. It relays upstream responses, including SSE, incrementally. AISIX detects chat, completions, and Responses envelopes and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. A route binds one fixed `target_url`, so create one route per vLLM server instead of relying on model access to pick an endpoint. vLLM also publishes root-level APIs such as `/pooling` and `/classify`. They are outside the `/v1` API base configured in this guide and have no normalized AISIX route. Reaching them through a passthrough route requires a `target_url` at the vLLM server root; do not assume a route targeting the `/v1` base above can reach them. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Check | | ----------------------------------- | -------------------------------------------------------------------------------------- | | Connection refused or timeout | Resolve and call `VLLM_API_BASE` from the AISIX gateway container. | | Upstream `401` | Use the same value for `--api-key` and the AISIX provider key's `api_key`. | | Model not found | Compare `model_name` with `GET $VLLM_API_BASE/models`. | | Chat-template error | Serve a chat model with a valid chat template or configure one in vLLM. | | Provider key creation returns `400` | Include `provider: "byo"`, `adapter: "openai"`, a non-empty `api_key`, and `api_base`. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to vLLM and verified the model alias. Continue with these guides: * [Bring Your Own Endpoint](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md): review the reusable setup and custom pricing options for private OpenAI-compatible servers. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Volcengine Ark (Doubao) [Volcengine Ark](https://www.volcengine.com/docs/82379/1298459) is ByteDance's platform for serving Doubao and other models. AISIX lets applications call those models through the gateway's OpenAI-compatible API while managing the Ark credential, caller access, rate limits, and usage accounting. One Ark provider key can serve both OpenAI-compatible Doubao chat traffic and Seedance video tasks. This guide verifies a chat model first, then reuses the provider key for [video generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An Ark API key from the [Volcengine Ark console](https://console.volcengine.com/ark). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Doubao-backed chat-completions route. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Ark credential and API root: ``` # Replace with your value export ARK_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "volcengine-prod", "provider": "volcengine", "api_key": "'"${ARK_API_KEY}"'", "api_base": "https://ark.cn-beijing.volces.com/api/v3", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `volcengine`, the provider ID the gateway's video route dispatches on. Volcengine Ark is not listed on models.dev, but this is not a community catalog entry: the gateway implements the Ark video wire natively, and the AISIX Cloud Admin API accepts the ID directly and derives the `openai` adapter for chat traffic. The `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the Ark API key and is sent as a bearer token on upstream calls. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is Ark's OpenAI-compatible base, whose version segment is `/api/v3`. AISIX appends the endpoint path to `api_base` verbatim, so this value produces `https://ark.cn-beijing.volces.com/api/v3/chat/completions`. The field is optional for this provider — when it is omitted, the AISIX Cloud Admin API fills in the same canonical value — but the examples set it explicitly so the upstream root each key targets stays visible in the configuration. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. The examples use the CN Beijing host `https://ark.cn-beijing.volces.com/api/v3`. ByteDance operates the platform's international counterpart, BytePlus ModelArk, at `https://ark.ap-southeast.bytepluses.com/api/v3` for Asia Pacific and `https://ark.eu-west.bytepluses.com/api/v3` for Europe. API keys and model availability vary by platform and region, so use the host and model ID provisioned for your account. Any Ark base ending in `/api/v3` works unchanged on the video routes, because the gateway derives the vendor root from that suffix. The `doubao-*` model IDs used below belong to Volcengine Ark. BytePlus uses different IDs for its corresponding models, so changing only `api_base` is not sufficient. Select a model ID documented for your region in the [BytePlus ModelArk documentation](https://docs.byteplus.com/en/docs/modelark/1099455) when you use a BytePlus account. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Ark model IDs combine a family name, version digits, and a release-date suffix, all separated by dashes: `doubao-seed-2-0-pro-260215` is the February 2026 release of the flagship Doubao Seed 2.0 Pro model, and `lite` and `mini` variants mark lighter, cheaper lanes of the same generation. Ark also provisions custom inference endpoints whose IDs start with `ep-`. `model_name` accepts either form, because the gateway forwards the value verbatim as the upstream `model`. Check the [Ark model list](https://www.volcengine.com/docs/82379/1330310) for the current catalog before creating a model alias. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "doubao-flagship-prod", "model_name": "doubao-seed-2-0-pro-260215", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Ark model ID, or the ID of an inference endpoint you provisioned in the Ark console, which starts with `ep-`. ❸ `provider_key_id` attaches the alias to the Ark provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the response, so capture it now: ``` CALLER_KEY_RESPONSE=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "volcengine-caller", "allowed_models": ["'"${MODEL_ID}"'"] }') export AISIX_API_KEY=$(printf '%s' "$CALLER_KEY_RESPONSE" | jq -r '.plaintext') CALLER_API_KEY_ID=$(printf '%s' "$CALLER_KEY_RESPONSE" | jq -r '.api_key.id') ``` The `allowed_models` value must reference the model ID captured in the previous step, so the key can only access the alias you created. The commands also retain the caller key's resource ID so you can add the video model later. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export ARK_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "volcengine-prod" provider: "volcengine" adapter: "openai" api_key: ${ARK_API_KEY} api_base: "https://ark.cn-beijing.volces.com/api/v3" models: - display_name: "doubao-flagship-prod" provider: "volcengine" model_name: "doubao-seed-2-0-pro-260215" provider_key: "volcengine-prod" api_keys: - display_name: "volcengine-caller" key_env: CALLER_API_KEY allowed_models: - "doubao-flagship-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "doubao-flagship-prod", "messages": [ { "role": "user", "content": "Say hello from Doubao." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `doubao-flagship-prod`. If the request fails, check the provider key `api_key`, `api_base`, and the Ark model ID in `model_name`. Ark serves models per platform and per region, so a model-not-found error from the upstream can also mean the model is not available to your account on the host `api_base` points at. ## Supply Cost Metadata[​](#supply-cost-metadata "Direct link to Supply Cost Metadata") models.dev does not price Volcengine Ark models, so the AISIX Cloud pricing catalog carries no per-token rates for the `volcengine` provider, and the dashboard suggests no model IDs when you create an alias for it. Cost metadata for [budgets](https://docs.api7.ai/ai-gateway/traffic-controls/budgets.md) and usage reports is operator-supplied on the model alias — the same posture as [BYO endpoints](https://docs.api7.ai/ai-gateway/providers/bring-your-own-endpoint.md). Until you set rates, usage from these aliases carries no cost signal, so spend-based controls do not see this traffic. Set the rates through [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md) in an AISIX Cloud deployment, or with the `cost` field on the model in the open-source AISIX gateway's `resources.yaml` — see [Cost Metadata](https://docs.api7.ai/ai-gateway/models/model-aliases.md#cost-metadata). Video submissions are additionally recorded with zero tokens, and duration-based cost accounting for video tasks is not yet applied to AISIX Cloud budgets. ## Generate Videos with Seedance[​](#generate-videos-with-seedance "Direct link to Generate Videos with Seedance") The same provider key drives the gateway's modeled video routes. Create a second alias that points at a current Ark video model, such as Doubao Seedance 2.0. Check the [Ark model list](https://www.volcengine.com/docs/82379/1330310) before creating the alias because dated model IDs change as releases move through the catalog: In AISIX Cloud, create the video alias: ``` VIDEO_MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "doubao-video-prod", "model_name": "doubao-seedance-2-0-260128", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$VIDEO_MODEL_ID" ``` Add the new model ID to the caller key's `allowed_models`. This field is a replacement list, so include the existing chat model ID: ``` curl -sS -X PATCH "$AISIX_CP/environments/$ENV_ID/api_keys/$CALLER_API_KEY_ID" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "allowed_models": ["'"${MODEL_ID}"'", "'"${VIDEO_MODEL_ID}"'"] }' ``` For the open-source AISIX gateway, add `doubao-video-prod` to the existing `models` collection. Replace the existing `volcengine-caller` entry with the updated entry below so it allows both model aliases. Preserve unrelated entries and collections: resources.yaml (video model access) ``` models: - display_name: "doubao-video-prod" provider: "volcengine" model_name: "doubao-seedance-2-0-260128" provider_key: "volcengine-prod" api_keys: - display_name: "volcengine-caller" key_env: CALLER_API_KEY allowed_models: - "doubao-flagship-prod" - "doubao-video-prod" ``` Validate and reload or restart the declarative resources file as described above. Submit a video task: ``` curl -sS -X POST "$AISIX_PROXY/v1/videos" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "doubao-video-prod", "prompt": "A paper boat drifts down a rain-soaked street at dusk.", "seconds": 5 }' ``` Five provider-specific behaviors are worth knowing before you script around this route: * **The chat `api_base` is reused as-is.** AISIX recognizes the `/api/v3` suffix, derives the vendor root from it, and composes Ark's native task paths underneath: tasks are submitted to `POST {root}/api/v3/contents/generations/tasks` and polled at `GET {root}/api/v3/contents/generations/tasks/{id}`. A provider key already configured for Doubao chat traffic works on the video routes without change. * **`seconds` maps to `duration`; `size` is not forwarded.** `seconds` is forwarded as the provider's integer `duration`. Ark expresses output dimensions as resolution and ratio quality tiers rather than pixel `WIDTHxHEIGHT` values, so a supplied `size` is validated for shape — a malformed value fails with `400` before the provider is contacted — but omitted from the upstream request, and the provider's default output setting applies. * **Provider-native video fields require a passthrough route.** The normalized request models only `prompt`, `seconds`, and `size`; extra fields are ignored. To send Ark fields such as reference `content`, `resolution`, `ratio`, `generate_audio`, or `watermark`, call `/passthrough/volcengine/contents/generations/tasks` with the native Ark body and exact model ID. Poll the task at `/passthrough/volcengine/contents/generations/tasks/{id}`. Passthrough does not return the unified AISIX video object or gateway-encoded task ID. The `/passthrough/volcengine` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with Ark's API root as its `target_url`; grant the route name on the caller key's `allowed_routes`. * **The full four-status lifecycle is reported.** AISIX maps Ark's task states onto the unified enum: `queued` reports as `queued`, `running` as `in_progress`, and `succeeded` as `completed`, while `failed`, `cancelled`, and `expired` all report as `failed` with the provider's error code and message when available. Once the task completes, the poll response also reports the actual video duration in `seconds`. * **Finished videos are delivered by redirect.** `GET /v1/videos/{id}/content` returns `302` with a `Location` header pointing at the provider's signed download URL. The bytes travel from Ark storage to the client and do not pass through the gateway, so use `curl -L` when downloading. For the full submit, poll, and download workflow, including status semantics and rate-limit behavior on polling, see [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md). ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") A Volcengine Ark provider key resolves the `openai` adapter, so route support follows that adapter plus each route's own provider rules: | Route | Behavior with a `volcengine` model alias | | ---------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, buffered and streaming. | | `/v1/embeddings` | Supported when the alias names an Ark embedding model, such as `doubao-embedding-text-240515`. AISIX appends `/embeddings` to `api_base`, reaching the provider's text-vectorization endpoint. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md), not Ark's native Responses API. Fields without a chat-completions equivalent are ignored. Use `/passthrough/volcengine/responses` with the exact Ark model ID when native Responses semantics are required. | | `/v1/messages` | Supported through AISIX translation to chat completions. `/v1/messages/count_tokens` is not supported because the model is not Anthropic-backed. | | `/v1/videos` and its status and content routes | Supported. See [Generate Videos with Seedance](#generate-videos-with-seedance). | | `/v1/images/generations` | Not supported through the normalized route because it accepts only models whose configured provider is `openai`. Ark exposes the same path for its image models; use `/passthrough/volcengine/images/generations` with the native request body and exact Ark model ID. | | `/v1/rerank` | Not supported. This route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/v1/models` | Returns caller-accessible AISIX aliases, not the Ark catalog. Consult the Ark console or model list linked above for available provider model IDs. | | `/passthrough/volcengine/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for provider-native APIs that AISIX has not modeled. Passthrough does not rewrite AISIX aliases and relays SSE incrementally. Recognized chat, completions, and Responses envelopes record supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. | See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the full endpoint and provider matrix. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Volcengine Ark and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Model Pricing](https://docs.api7.ai/ai-gateway/cloud/model-pricing.md): set the operator-supplied per-token rates that budgets and usage reports need for this provider. * [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md): follow the full submit, poll, and download workflow for Seedance tasks. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Weights & Biases Inference [Weights & Biases Inference](https://docs.wandb.ai/inference/api-reference) provides hosted inference for open models through the W\&B platform. AISIX maps those models to stable aliases and controls caller access at the gateway. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A W\&B API key authorized for Inference. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` W\&B Inference is a community catalog provider with an OpenAI-compatible API. AISIX connects through the `openai` adapter and authenticates upstream requests with a bearer token. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") ``` export WANDB_API_KEY="YOUR_WANDB_API_KEY" PROVIDER_KEY_ID=$( curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "wandb-prod", "provider": "wandb", "api_key": "'"${WANDB_API_KEY}"'", "api_base": "https://api.inference.wandb.ai/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -er '.provider_key.id' ) echo "$PROVIDER_KEY_ID" ``` The AISIX provider ID is `wandb`, not `weights-and-biases`. Do not add `adapter`; AISIX derives the `openai` adapter for catalog providers. The explicit API base includes `/v1`. AISIX appends `/chat/completions` when it sends the upstream request. W\&B attributes requests to your default entity and the `inference` project when no project is specified. To use another W\&B team and project, add an `OpenAI-Project` value in `request.default_headers` on the provider key. See [Upstream Request Headers](https://docs.api7.ai/ai-gateway/models/upstream-request-headers.md#request-context-header-values). ### Create a Model[​](#create-a-model "Direct link to Create a Model") Create an alias with the complete W\&B model ID: ``` MODEL_ID=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "wandb-llama-prod", "model_name": "meta-llama/Llama-3.1-8B-Instruct", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -er '.model.id' ) echo "$MODEL_ID" ``` Preserve the publisher namespace and casing when substituting another model from the W\&B Inference catalog. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") ``` AISIX_API_KEY=$( curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "wandb-inference-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -er '.plaintext' ) echo "$AISIX_API_KEY" ``` ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export WANDB_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "wandb-prod" provider: "wandb" adapter: "openai" api_key: ${WANDB_API_KEY} api_base: "https://api.inference.wandb.ai/v1" models: - display_name: "wandb-llama-prod" provider: "wandb" model_name: "meta-llama/Llama-3.1-8B-Instruct" provider_key: "wandb-prod" api_keys: - display_name: "wandb-inference-caller" key_env: CALLER_API_KEY allowed_models: - "wandb-llama-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer $AISIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "wandb-llama-prod", "messages": [ { "role": "user", "content": "Say hello from W&B Inference." } ] }' ``` AISIX forwards the upstream model ID to `POST /v1/chat/completions` with the W\&B key in the bearer header. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") W\&B Serverless Inference currently exposes OpenAI-compatible chat completions and model listing. AISIX adds its chat-based protocol bridges, but it cannot add an upstream capability W\&B does not expose: | Route | Behavior with a `wandb` model alias | | -------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, buffered and streaming. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md). Fields without a chat-completions equivalent are ignored. | | `/v1/messages` | Supported through AISIX translation to chat completions. `/v1/messages/count_tokens` is not supported because the model is not Anthropic-backed. | | `/v1/embeddings` | Not supported. W\&B Serverless Inference does not publish an embeddings endpoint on this API root. | | `/v1/models` | Returns caller-accessible AISIX aliases, not the W\&B catalog. Call `GET /passthrough/wandb/models` to reach W\&B's native model-list endpoint with the configured provider credential. | | `/v1/images/generations`, `/v1/videos`, and `/v1/rerank` | Not supported. W\&B does not expose these APIs, and the `wandb` provider value is outside the normalized routes' allowlists. | | `/passthrough/wandb/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming this prefix with the W\&B API root as its `target_url`; grant the route on the caller key's `allowed_routes`. Passthrough does not rewrite AISIX aliases and relays SSE responses incrementally. Recognized chat, completions, and Responses envelopes record supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. | See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the full endpoint and provider matrix. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Check | | ----------------------------------- | --------------------------------------------------- | | Upstream `401` or `403` | Confirm the W\&B key has Inference access. | | Upstream `404` | Keep `/v1` in the API base and verify the model ID. | | Model not found | Preserve the publisher namespace and casing. | | Provider key creation returns `400` | Use `provider: "wandb"` without `adapter`. | ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to W\&B Inference and verified the model alias. Continue with these guides: * [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md): manage provider-key rotation and environment access. * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for the alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between W\&B Inference and another provider. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # xAI (Grok) [xAI](https://docs.x.ai/) provides the Grok family of models through its API. Applications call Grok through stable AISIX aliases while the gateway keeps the xAI credential out of client code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * An xAI API key. Follow the [xAI Quickstart](https://docs.x.ai/developers/quickstart) to create one. * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the Grok-backed chat-completions route. xAI is a community catalog provider with an OpenAI-compatible API. AISIX connects through the `openai` adapter, authenticates upstream requests with a bearer token, and uses the xAI API root as `api_base`. AISIX does not register xAI-specific request or response rewrites, so the sections below cover the provider-specific values you must configure. The provider catalog returns xAI as a community entry rather than a featured provider. ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the xAI credential and API root: ``` # Replace with your value export XAI_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "xai-prod", "provider": "xai", "api_key": "'"${XAI_API_KEY}"'", "api_base": "https://api.x.ai/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `xai`. The AISIX Cloud Admin API derives the adapter from the catalog provider, so `xai` resolves to the `openai` adapter through the catalog's default rule. The `adapter` field is only accepted on BYO provider keys and is rejected on a catalog provider key. ❷ `api_key` stores the xAI API key. xAI authenticates with HTTP bearer authentication, which is what the `openai` adapter already sends. The value follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` is **required** for `xai`. A curated provider entry carries a default API root, and a community-catalog provider falls back to the `api` field that models.dev publishes. The models.dev entry for `xai` publishes no `api` field, so there is nothing to fall back to. Omitting `api_base` returns `400` with the message `models.dev does not publish a default api_base for this provider — set api_base explicitly, or switch to the "byo" provider sentinel`. Use `https://api.x.ai/v1`. AISIX appends the endpoint path to `api_base`, so the value must be the root that `/chat/completions` hangs off. xAI documents the full endpoint as `https://api.x.ai/v1/chat/completions` and documents `https://api.x.ai/v1` as the base URL for OpenAI client libraries, which makes `https://api.x.ai/v1` the correct root. caution Do not set `api_base` to the bare host `https://api.x.ai`. AISIX synthesizes a missing `/v1` segment only for the canonical OpenAI host; every other host passes through as written. A bare xAI host produces the upstream URL `https://api.x.ai/chat/completions`, which xAI does not serve. Pasting the full endpoint URL is safe, however: AISIX strips a trailing `/chat/completions` and any trailing slash before appending the endpoint path. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. ### Create a Model[​](#create-a-model "Direct link to Create a Model") Grok model IDs are bare lowercase slugs with no vendor prefix. Point releases carry a dotted minor version, such as `grok-4.5` and `grok-4.3`; agent-oriented models use their own family name, such as `grok-build-0.1`; and dated snapshots append the release date and a behavior suffix, such as `grok-4.20-0309-reasoning`. Do not carry over a prefixed form such as `xai/grok-4.5` from an aggregator. Check the [xAI model list](https://docs.x.ai/developers/models) for the current slugs before you create an alias. xAI retires older Grok slugs on a published schedule and redirects requests for a retired slug to a current model, so an alias left pinned to a retired slug keeps working while silently serving a different model. Re-point `model_name` when a slug you use is retired. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "grok-prod", "model_name": "grok-4.5", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the xAI model ID, for example `grok-4.5`, `grok-4.3`, or `grok-build-0.1`. ❸ `provider_key_id` attaches the alias to the xAI provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the create response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "xai-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export XAI_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "xai-prod" provider: "xai" adapter: "openai" api_key: ${XAI_API_KEY} api_base: "https://api.x.ai/v1" models: - display_name: "grok-prod" provider: "xai" model_name: "grok-4.5" provider_key: "xai-prod" api_keys: - display_name: "xai-caller" key_env: CALLER_API_KEY allowed_models: - "grok-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "grok-prod", "messages": [ { "role": "user", "content": "Say hello from Grok." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `grok-prod`. If the request fails, check the provider key `api_key`, confirm that `api_base` ends in `/v1`, and confirm the xAI model ID in `model_name`. ## Adapt the Wire Format When xAI Changes It[​](#adapt-the-wire-format-when-xai-changes-it "Direct link to Adapt the Wire Format When xAI Changes It") Curated providers in the AISIX catalog can carry request and response rewrites — a parameter rename, a default header, a nonstandard streaming path for reasoning. The community catalog path registers none of these for `xai`, so AISIX sends and reads the plain OpenAI chat-completions shape. This remains compatible with xAI, although xAI now recommends its native Responses API for new integrations. Use a passthrough route when an application needs exact xAI Responses semantics instead of the AISIX chat-based bridge. Two overrides on the provider key cover the common cases. Both apply to every model that references the key, so test them against a non-production alias first. See [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) for the full field catalog. Rename a top-level request parameter when xAI expects a different name than your clients send: ``` { "request": { "param_renames": { "max_completion_tokens": "max_tokens" } } } ``` Map reasoning from a nonstandard streaming `delta` path onto the canonical `delta.reasoning_content` field: ``` { "response": { "reasoning_field": "delta.thinking" } } ``` AISIX does not strip unrecognized top-level parameters on the chat-completions path, so a new xAI-specific parameter reaches the upstream without any override at all. Overrides are needed only when a name has to change on the way out, or when a response field has to be relocated on the way back. ## Control Reasoning Effort[​](#control-reasoning-effort "Direct link to Control Reasoning Effort") Grok models expose configurable reasoning effort through the standard OpenAI `reasoning_effort` parameter at the top level of the chat-completions body. Because AISIX forwards unrecognized top-level parameters unchanged, the value reaches xAI as sent: ``` { "model": "grok-prod", "messages": [ { "role": "user", "content": "Plan a three-step migration." } ], "reasoning_effort": "low" } ``` Accepted values differ by model. `grok-4.5` accepts `low`, `medium`, and `high`; `grok-4.3` also accepts `none` to disable reasoning; and `grok-4.20-multi-agent-0309` adds `xhigh`. Confirm the values for the model you configured on its page in the [xAI model list](https://docs.x.ai/developers/models), because a value one model accepts can be rejected by another. Send `reasoning_effort` on `/v1/chat/completions`, not on `/v1/responses`. For a non-OpenAI provider, the AISIX Responses route translates the request down to the chat-completions shape and carries only the fields that map cleanly: `instructions`, `input`, `tools`, `tool_choice`, `temperature`, `top_p`, `max_output_tokens`, and `stream`. OpenAI-only knobs, including `reasoning`, `store`, `previous_response_id`, and `text`, are dropped rather than forwarded. ## Target a Regional Endpoint[​](#target-a-regional-endpoint "Direct link to Target a Regional Endpoint") xAI serves regional endpoints at `https://<region>.api.x.ai` for requests that must be processed in a specific region, and documents the OpenAI-client base URL for the European region as `https://eu-west-1.api.x.ai/v1`. The same `/v1` root rule applies: set `api_base` to the regional host plus `/v1`. Create a second provider key for the regional route rather than editing the existing one, so the two roots stay independently attributable in usage records. In AISIX Cloud, create the regional provider key, model alias, and caller key: ``` EU_PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "xai-eu", "provider": "xai", "api_key": "'"${XAI_API_KEY}"'", "api_base": "https://eu-west-1.api.x.ai/v1", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') EU_MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "grok-eu-prod", "model_name": "grok-4.5", "provider_key_id": "'"${EU_PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') XAI_EU_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "xai-eu-caller", "allowed_models": ["'"${EU_MODEL_ID}"'"] }' | jq -r '.plaintext') ``` Use `XAI_EU_API_KEY` when calling the `grok-eu-prod` alias. For the open-source AISIX gateway, add `xai-eu` to `provider_keys` and `grok-eu-prod` to `models`. Replace the existing `xai-caller` entry with the updated entry below so it allows both model aliases. Preserve unrelated entries and collections: resources.yaml (EU region resources) ``` provider_keys: - display_name: "xai-eu" provider: "xai" adapter: "openai" api_key: ${XAI_API_KEY} api_base: "https://eu-west-1.api.x.ai/v1" models: - display_name: "grok-eu-prod" provider: "xai" model_name: "grok-4.5" provider_key: "xai-eu" api_keys: - display_name: "xai-caller" key_env: CALLER_API_KEY allowed_models: - "grok-prod" - "grok-eu-prod" ``` Validate and reload or restart the declarative resources file as described above. Callers select the region by choosing the alias, and the gateway records each alias separately in usage events. Confirm the regions available to your account with xAI before you route production traffic to one: a request that xAI cannot serve in the requested region fails rather than falling back to another region. ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") The community catalog path assigns xAI the `openai` adapter. Normalized chat-style routes use the AISIX OpenAI adapter, while xAI APIs whose paths or wire formats differ remain available through passthrough routes. The `/passthrough/xai` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with the xAI API root `https://api.x.ai/v1` as its `target_url`; grant the route on the caller key's `allowed_routes`. | Route | Behavior with an xAI alias | | ---------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, including `stream: true`. xAI now classifies Chat Completions as deprecated, but normalized xAI routes in AISIX still use it as their upstream chat surface. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) over chat completions. Fields without a chat equivalent are dropped, so xAI-native state, hosted tools, and `previous_response_id` are not preserved. Use `/passthrough/xai/responses` with the exact xAI model ID for the native API. | | `/v1/messages` | Supported for Anthropic-shaped callers through AISIX translation to chat completions. xAI also exposes an Anthropic-compatible Messages route; use `/passthrough/xai/messages` with the exact xAI model ID for its native behavior. `/v1/messages/count_tokens` is not supported because the model is not Anthropic-backed. | | `/v1/embeddings` | Not usable. The catalog entry for xAI publishes no embedding model, so route embeddings to a separate provider. See [Embeddings](https://docs.api7.ai/ai-gateway/endpoints/embeddings.md). | | `/v1/images/generations` | Rejected with `400`. The route accepts only models whose provider is `openai`. Use `/passthrough/xai/images/generations` with the native body and exact Grok Imagine model ID. Image edits are available at `/passthrough/xai/images/edits`. | | `/v1/videos` | Rejected with `501` because `xai` is outside the route's provider allowlist. Use `/passthrough/xai/videos/generations` for the native asynchronous API and poll at `/passthrough/xai/videos/{request_id}`. Native edit and extension routes are available under the same passthrough prefix. | | `/v1/audio/transcriptions`, `/v1/audio/translations`, and `/v1/audio/speech` | Not compatible with the native xAI voice paths. Use `/passthrough/xai/stt` for REST speech-to-text and `/passthrough/xai/tts` for REST text-to-speech. xAI does not publish an audio-translation route. | | `/v1/realtime` | Supported for a direct xAI model alias. AISIX relays the OpenAI-compatible WebSocket wire to `wss://api.x.ai/v1/realtime` and rewrites the alias to the configured xAI model ID. | | `/v1/rerank` | Rejected with `400`. The route accepts only the `openai`, `cohere`, and `jina` provider values. | | `/v1/models` | Returns caller-accessible AISIX aliases, not the xAI catalog. Use `GET /passthrough/xai/models` for the native xAI model list. | | `/passthrough/xai/*rest` | Available for provider-native HTTP routes through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md), with limited gateway normalization. The route does not rewrite AISIX aliases and relays SSE responses incrementally. Recognized chat, completions, and Responses envelopes record supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. | ### Reach Provider-Native Routes Through Passthrough[​](#reach-provider-native-routes-through-passthrough "Direct link to Reach Provider-Native Routes Through Passthrough") A passthrough route preserves the request body, but not every header. On an inject-mode route, AISIX removes hop-by-hop headers and the provider key's `strip_headers` values (`authorization`, `cookie`, `set-cookie`, and `x-api-key` by default); injects the bound provider key's xAI credential as `Authorization: Bearer ...`; and adds `x-aisix-request-id`. The route's `target_url` fixes the upstream root. Use a passthrough route for xAI routes that have no OpenAI-compatible equivalent in the gateway, such as deferred chat completions, image generation, and the native Responses surface. Because the route's `target_url` ends in `/v1`, AISIX removes a duplicated leading `v1` segment from the remaining path, so `/passthrough/xai/v1/chat/deferred-completion/<request_id>` and `/passthrough/xai/chat/deferred-completion/<request_id>` both resolve to the same upstream URL: ``` curl -sS "$AISIX_PROXY/passthrough/xai/chat/deferred-completion/YOUR_REQUEST_ID" \ -H "Authorization: Bearer ${AISIX_API_KEY}" ``` Passthrough requests are authenticated with a caller API key, and the gateway authorizes them against the route, not a model: the key must grant the route's name in its `allowed_routes` glob list, and the key's model allowlist plays no part here. Guardrails attach to the route through the `passthrough_route` scope, alongside the caller key, team, and environment scopes — no model alias is involved. The gateway does not rewrite the `model` field, so send the exact xAI model ID. AISIX detects chat, completions, and Responses envelopes and records supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. Treat passthrough routes as an escape hatch for routes the gateway does not model. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to xAI and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between xAI and a second provider. * [Provider-Specific Overrides](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides): adapt request and response shapes when an upstream API differs from its adapter. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # Zhipu AI (GLM) [Zhipu AI](https://docs.bigmodel.cn/cn/api/introduction) provides GLM language models and CogVideoX video generation through hosted APIs. Applications use AISIX caller keys and model aliases for both capabilities while the gateway holds the upstream credential. This guide covers both GLM chat traffic and CogVideoX video tasks through a single Zhipu AI provider key. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting, prepare the following: * One AISIX setup: <!-- --> * For AISIX Cloud, an environment with an attached gateway and a write-scoped admin token. For On-Premises, follow the [AISIX Cloud Quickstart](https://docs.api7.ai/ai-gateway/getting-started/aisix-cloud-quickstart.md). To request Hybrid Cloud access, [contact API7](https://api7.ai/contact). * For the open-source AISIX gateway, prepare either a local AISIX installation or the Docker setup from the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Configure the gateway to load a declarative resources file. * A Zhipu AI API key from the [Zhipu AI open platform](https://open.bigmodel.cn/). * `curl` and `jq`. ## Configure with AISIX Cloud[​](#configure-with-aisix-cloud "Direct link to Configure with AISIX Cloud") Export the AISIX Cloud connection details: ``` # AISIX_CP is the Admin API base URL; include /api and omit a trailing slash # The local On-Premises quickstart uses http://localhost:8080/api export AISIX_CP="YOUR_AISIX_CLOUD_ADMIN_API_URL" export AISIX_TOKEN="YOUR_ADMIN_TOKEN" export ENV_ID="YOUR_ENVIRONMENT_ID" ``` Create a provider key, model alias, and caller API key for the GLM-backed chat-completions route. Because Zhipu AI exposes an OpenAI-compatible API, AISIX connects through the `openai` adapter and uses the Zhipu AI API root as `api_base`. The same provider key can also serve the gateway's [video-generation route](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md). ### Create a Provider Key[​](#create-a-provider-key "Direct link to Create a Provider Key") Create the provider key that stores the Zhipu AI credential and API root: ``` # Replace with your value export ZHIPU_API_KEY="YOUR_PROVIDER_API_KEY" PROVIDER_KEY_ID=$(curl -sS -X POST "$AISIX_CP/provider_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "zhipuai-prod", "provider": "zhipuai", "api_key": "'"${ZHIPU_API_KEY}"'", "api_base": "https://open.bigmodel.cn/api/paas/v4", "allowed_environments": ["'"${ENV_ID}"'"] }' | jq -r '.provider_key.id') echo "$PROVIDER_KEY_ID" ``` ❶ `provider` is `zhipuai`, the catalog provider ID for the Zhipu AI open platform. The AISIX Cloud Admin API derives the adapter from the catalog provider; the `adapter` field is only accepted on BYO provider keys. ❷ `api_key` stores the Zhipu AI API key and is sent as a bearer token on upstream calls. It follows the credential-handling behavior in [Provider Keys](https://docs.api7.ai/ai-gateway/models/provider-keys.md#credential-handling). ❸ `api_base` carries an unusual path shape: the version segment is `v4` and lives under `/api/paas`, not the `/v1` root that most OpenAI-compatible vendors publish. AISIX appends the endpoint path to `api_base` verbatim, so this value produces `https://open.bigmodel.cn/api/paas/v4/chat/completions`. The field is optional for this provider — when it is omitted, the AISIX Cloud Admin API fills in the same canonical value. Set it explicitly when you point the key at a different Zhipu AI deployment. The command captures the returned provider key ID in `PROVIDER_KEY_ID`. Zhipu AI operates two platforms with separate catalog provider IDs. Use `zhipuai` for the mainland China platform at `open.bigmodel.cn`, and `zai` for the international Z.ai platform, whose API root is `https://api.z.ai/api/paas/v4`. The two IDs are not interchangeable. Only `zhipuai` (and the short spelling `zhipu`) is dispatched by the modeled video routes, so a `zai` provider key serves chat traffic but returns a not-implemented error on `/v1/videos`. Unlike some other OpenAI-compatible catalog entries, the `zhipuai` entry needs no request or response overrides. AISIX sends the OpenAI request shape unchanged and does not rename any parameter, because Zhipu AI accepts the standard field names and already returns reasoning text on the canonical field. See [Control Thinking Mode](#control-thinking-mode). ### Create a Model[​](#create-a-model "Direct link to Create a Model") Zhipu AI model IDs follow a `glm-<version>` pattern. A bare version such as `glm-5.2` is the flagship text model for that generation; a `-flash`, `-flashx`, or `-air` suffix marks a lighter, cheaper lane; and a trailing `v`, as in `glm-5v-turbo`, marks a vision model. Check the [Zhipu AI model overview](https://docs.bigmodel.cn/cn/guide/start/model-overview) for the current catalog before creating a model alias. Create the model alias callers will send in requests: ``` MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "glm-flagship-prod", "model_name": "glm-5.2", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$MODEL_ID" ``` ❶ `display_name` is the alias callers send in `model`. ❷ `model_name` is the Zhipu AI model ID, for example `glm-5.2`, `glm-5`, or `glm-4.7`. ❸ `provider_key_id` attaches the alias to the Zhipu AI provider key. ### Create a Caller API Key[​](#create-a-caller-api-key "Direct link to Create a Caller API Key") Create the caller API key that can access the model alias. The plaintext key is server-generated and returned once in the response, so capture it now: ``` AISIX_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "zhipuai-caller", "allowed_models": ["'"${MODEL_ID}"'"] }' | jq -r '.plaintext') echo "$AISIX_API_KEY" ``` The `allowed_models` value must reference the model ID captured in the previous step, so the key can only access the alias you created. After the write, the configuration projects to attached gateways automatically. ## Configure with the Open-Source AISIX Gateway[​](#configure-with-the-open-source-aisix-gateway "Direct link to Configure with the Open-Source AISIX Gateway") Export the upstream credential and choose the caller API key that applications will send to the gateway: ``` export ZHIPU_API_KEY="YOUR_PROVIDER_API_KEY" export CALLER_API_KEY="YOUR_CALLER_API_KEY" ``` For a new gateway, use this complete resources file. For an existing gateway, merge the entries into its [current file](https://docs.api7.ai/ai-gateway/reference/resources-file.md#apply-resource-examples-safely), preserving its other resources: resources.yaml ``` _format_version: "1" provider_keys: - display_name: "zhipuai-prod" provider: "zhipuai" adapter: "openai" api_key: ${ZHIPU_API_KEY} api_base: "https://open.bigmodel.cn/api/paas/v4" models: - display_name: "glm-flagship-prod" provider: "zhipuai" model_name: "glm-5.2" provider_key: "zhipuai-prod" api_keys: - display_name: "zhipuai-caller" key_env: CALLER_API_KEY allowed_models: - "glm-flagship-prod" ``` If AISIX is installed locally, validate the file before loading it: ``` aisix validate --resources resources.yaml ``` After validation, start the gateway with the referenced environment variables in its process environment. Reload an existing gateway only if those variables are already available to the process; otherwise, restart it with the updated environment. If you use Docker, adapt the validation and startup commands in the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md). Mount this `resources.yaml` file and pass every environment variable it references with `-e` in both commands. After the resources load, prepare the shared verification request below: ``` export AISIX_API_KEY="$CALLER_API_KEY" ``` ## Verify the Provider Connection[​](#verify-the-provider-connection "Direct link to Verify the Provider Connection") Export the AISIX gateway origin: ``` # The local quickstarts use http://127.0.0.1:3000 export AISIX_PROXY="YOUR_AISIX_GATEWAY_ORIGIN" ``` Send a chat-completions request through the AISIX proxy: ``` curl -sS -X POST "$AISIX_PROXY/v1/chat/completions" \ -H "Authorization: Bearer ${AISIX_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-flagship-prod", "messages": [ { "role": "user", "content": "Say hello from GLM." } ] }' ``` The gateway returns an OpenAI-compatible response that echoes the caller-facing alias `glm-flagship-prod`. Because the model reasons by default, the assistant message also carries a `reasoning_content` field alongside `content`. If the request fails, check the provider key `api_key`, `api_base`, and the Zhipu AI model ID in `model_name`. ## Control Thinking Mode[​](#control-thinking-mode "Direct link to Control Thinking Mode") GLM reasoning models think by default. To turn thinking off for a single request, add the provider's `thinking` object to the chat-completions body: ``` { "thinking": { "type": "disabled" } } ``` AISIX forwards fields it does not itself interpret to the upstream unchanged, so the control reaches Zhipu AI as written. Accepted values are `enabled` and `disabled`. GLM-5.2 also accepts `reasoning_effort` to control the reasoning depth while thinking is enabled. See [Deep Thinking](https://docs.bigmodel.cn/cn/guide/capabilities/thinking) for the current model-specific behavior and accepted effort values. Zhipu AI returns the thinking text on `reasoning_content` — `message.reasoning_content` for a buffered response and `delta.reasoning_content` for a streamed one. That is the canonical field AISIX preserves, so this provider needs no [`response.reasoning_field`](https://docs.api7.ai/ai-gateway/models/provider-keys.md#configure-provider-specific-overrides) override on the provider key. ## Generate Videos with CogVideoX[​](#generate-videos-with-cogvideox "Direct link to Generate Videos with CogVideoX") The same provider key drives the gateway's modeled video routes. Create a second alias that points at a Zhipu AI video model, for example [CogVideoX-3](https://docs.bigmodel.cn/cn/guide/models/video-generation/cogvideox-3): In AISIX Cloud, create the video alias and a caller key scoped to it: ``` VIDEO_MODEL_ID=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/models" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "glm-video-prod", "model_name": "cogvideox-3", "provider_key_id": "'"${PROVIDER_KEY_ID}"'" }' | jq -r '.model.id') echo "$VIDEO_MODEL_ID" VIDEO_API_KEY=$(curl -sS -X POST "$AISIX_CP/environments/$ENV_ID/api_keys" \ -H "Authorization: Bearer $AISIX_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "display_name": "zhipuai-video-caller", "allowed_models": ["'"${VIDEO_MODEL_ID}"'"] }' | jq -r '.plaintext') ``` For the open-source AISIX gateway, add `glm-video-prod` to the existing `models` collection. Replace the existing `zhipuai-caller` entry with the updated entry below so it allows both model aliases. Preserve unrelated entries and collections: resources.yaml (video model access) ``` models: - display_name: "glm-video-prod" provider: "zhipuai" model_name: "cogvideox-3" provider_key: "zhipuai-prod" api_keys: - display_name: "zhipuai-caller" key_env: CALLER_API_KEY allowed_models: - "glm-flagship-prod" - "glm-video-prod" ``` Validate and reload or restart the declarative resources file as described above, then use the existing caller key for the video request: ``` export VIDEO_API_KEY="$CALLER_API_KEY" ``` Submit a task: ``` curl -sS -X POST "$AISIX_PROXY/v1/videos" \ -H "Authorization: Bearer ${VIDEO_API_KEY}" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-video-prod", "prompt": "A paper boat drifts down a rain-soaked street at dusk.", "seconds": 5, "size": "1920x1080" }' ``` Five provider-specific behaviors are worth knowing before you script around this route: * **The chat `api_base` is reused as-is.** AISIX recognizes the `/api/paas/v4` suffix, derives the vendor root from it, and composes the provider's native task paths underneath. A provider key already configured for GLM chat traffic works on the video routes without change. * **Parameters map directly.** `seconds` is forwarded as the provider's integer `duration`, and `size` passes through verbatim, because Zhipu AI documents the same `WIDTHxHEIGHT` spelling the unified request uses. AISIX validates the shape before contacting the provider. * **Provider-native video fields require a passthrough route.** The normalized request models only `prompt`, `seconds`, and `size`; extra fields are ignored. To send CogVideoX fields such as `image_url`, `quality`, `with_audio`, or `fps`, call `/passthrough/zhipuai/videos/generations` with the native body and exact model ID. Poll the task at `/passthrough/zhipuai/async-result/{id}`. Passthrough does not return the unified AISIX video object or gateway-encoded task ID. The `/passthrough/zhipuai` paths on this page assume a [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) claiming that prefix with the provider's API root as its `target_url`; grant the route name on the caller key's `allowed_routes`. * **There is no queued state.** Zhipu AI reports a task as processing from the moment it is accepted, so the unified status goes straight to `in_progress` and never reports `queued`. Poll `GET /v1/videos/{id}` until it reports `completed`. * **Finished videos are delivered by redirect.** `GET /v1/videos/{id}/content` returns `302` with a `Location` header pointing at the provider's signed download URL. The bytes travel from Zhipu AI storage to the client and do not pass through the gateway, so use `curl -L` when downloading. For the full submit, poll, and download workflow, including status semantics and rate-limit behavior on polling, see [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md). ## Endpoint Coverage[​](#endpoint-coverage "Direct link to Endpoint Coverage") A Zhipu AI provider key resolves the `openai` adapter, so route support follows that adapter plus each route's own provider rules: | Route | Behavior with a `zhipuai` model alias | | ---------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `/v1/chat/completions` | Supported, buffered and streaming. | | `/v1/embeddings` | Supported when the alias names a Zhipu AI embedding model, such as `embedding-3`. AISIX appends `/embeddings` to `api_base`, reaching the provider's text-embedding endpoint. | | `/v1/responses` | Supported through the [Responses bridge](https://docs.api7.ai/ai-gateway/endpoints/responses-api.md) for fields the chat adapter path can express. OpenAI-specific Responses fields without a chat equivalent are ignored. | | `/v1/messages` | Translated to chat completions by default. Zhipu AI also publishes a Claude-compatible Messages API at `https://open.bigmodel.cn/api/anthropic`, but its documentation does not define the token-counting route that AISIX couples to `apis.messages`. Use an authenticated [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for native Messages instead of declaring `apis.messages`, which would also route `/v1/messages/count_tokens` upstream. | | `/v1/videos` and its status and content routes | Supported. See [Generate Videos with CogVideoX](#generate-videos-with-cogvideox). | | `/v1/images/generations` | Not supported through the normalized route because it accepts only models whose configured provider is `openai`. Use `/passthrough/zhipuai/images/generations` with the native body and exact image model ID. | | `/v1/audio/transcriptions` | Supported when the alias names a Zhipu AI speech-to-text model, such as `glm-asr-2512`. The upstream path and multipart request shape match the AISIX route. A provider streaming response is buffered before AISIX returns it. | | `/v1/audio/speech` | Supported when the alias names `glm-tts`. Non-streaming audio is returned verbatim. AISIX buffers the provider response when the native `stream` field is enabled. | | `/v1/audio/translations` | Not supported because Zhipu AI does not publish the corresponding upstream route. | | `/v1/realtime` | Available as a WebSocket relay for a direct GLM-Realtime alias. AISIX dials `wss://open.bigmodel.cn/api/paas/v4/realtime` with the configured upstream model in the query and relays events without translation. If the client also sets `session.model`, send the exact Zhipu AI model ID because AISIX does not rewrite WebSocket frame bodies. | | `/v1/rerank` | Not supported through the normalized route because it accepts only the `openai`, `cohere`, and `jina` provider values. Use `/passthrough/zhipuai/rerank` with the native body and exact `rerank` model ID. | | `/v1/models` | Returns caller-accessible AISIX aliases, not the Zhipu AI catalog. Consult the model overview linked above for provider model IDs. | | `/passthrough/zhipuai/*rest` | Available through a configured [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) for provider-native HTTP routes. Passthrough does not rewrite AISIX aliases and relays SSE incrementally. Recognized chat, completions, and Responses envelopes record supported usage fields. Requests without a recognized carrier field remain opaque: buffered responses record zero tokens, while opaque SSE can still record top-level supported `usage` fields. | See [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md) for the full endpoint and provider matrix. ## Next Steps[​](#next-steps "Direct link to Next Steps") You have now connected AISIX to Zhipu AI and verified the model alias. Continue with these guides: * [Model Aliases](https://docs.api7.ai/ai-gateway/models/model-aliases.md): configure routing, retry behavior, or cost metadata for this alias. * [Routing and Failover](https://docs.api7.ai/ai-gateway/routing/routing-and-failover.md): fail over between Zhipu AI and a second provider. * [Video Generation](https://docs.api7.ai/ai-gateway/endpoints/video-generation.md): follow the full submit, poll, and download workflow for CogVideoX tasks. * [Provider Compatibility](https://docs.api7.ai/ai-gateway/providers/compatibility.md): review supported proxy endpoints and provider-specific boundaries. --- # CLI Reference The `aisix` binary runs the AISIX gateway and provides commands for working with open-source AISIX gateway configuration. Use `validate` to check a declarative [`resources.yaml`](https://docs.api7.ai/ai-gateway/reference/resources-file.md) file and `export` to convert resources in an existing etcd store into that format. Both commands run without starting gateway listeners. ``` Usage: aisix --config <CONFIG> aisix <COMMAND> ``` | Command | Purpose | | -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `aisix --config <CONFIG>` | Start the gateway with the given startup configuration file. See the [Startup Configuration Reference](https://docs.api7.ai/ai-gateway/reference/configuration-files.md). | | `aisix validate --resources <FILE>` | Check a resources file without starting the gateway. | | `aisix export --etcd <ENDPOINT> [-o <FILE>]` | Export resources from an open-source AISIX gateway's etcd store as a loadable resources file. | When running the official container image, invoke the binary directly: ``` docker run --rm --entrypoint /usr/local/bin/aisix ghcr.io/api7/aisix:1.0.0 --help ``` ## Validate a Resources File[​](#validate-a-resources-file "Direct link to Validate a Resources File") `aisix validate` runs the identical pipeline the gateway uses to load a resources file — read, `${VAR}` interpolation, name-reference resolution, canonical schema validation, and cross-reference checks — without starting any listener. Use it as a pre-check before a boot or a reload, or as a CI gate on configuration changes. ``` aisix validate --resources resources.yaml ``` | Option | Required | Description | | -------------------- | -------- | --------------------------------------- | | `--resources <FILE>` | Yes | Path to the resources file to validate. | `${VAR}` references in the file resolve against the environment of the `validate` process itself. Run the command with the same variables the gateway will receive, or validation fails on the unresolved references. After the file loads, `validate` reports two things about guardrails that a clean load does not otherwise reveal. The first is a guardrail nothing attaches. It loads, it is counted in the resource total, and it inspects no traffic, because a guardrail's scope is its attachments and nothing else. That is a valid state rather than an error — the model or route it was scoped to may have been removed from the file — so `validate` names such rows on standard error and still exits `0`. The second is a row that cannot run. A guardrail whose configuration parses but does not build — an invalid regular expression, an unknown detector, a `custom` script with a syntax error — is dropped when the gateway starts, and the gateway then serves without that screening. The load itself succeeded, so nothing else reports a problem: this check is where such a row surfaces. ### Exit Codes[​](#exit-codes "Direct link to Exit Codes") | Exit code | Meaning | | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `0` | The file loads and every enabled guardrail can run. A summary is printed to standard output — possibly alongside a note on standard error naming enabled guardrails that nothing attaches. | | Non-zero | The file does not load, or a guardrail loads but cannot run. The report is printed to standard error. | On success: ``` OK: resources.yaml loaded 3 resource(s) ``` A guardrail nothing attaches is reported without changing the exit code: ``` resources file resources.yaml: 1 enabled guardrail(s) have no attachment and inspect no traffic: - guardrails ("no-secrets"): add a guardrail_attachments entry to put it in force OK: resources.yaml loaded 4 resource(s) ``` On failure, every problem across the whole file is reported together, with the resource kind, entry, and field: ``` resources file resources.yaml: 3 error(s): - provider_keys[0]: field `api_key`: environment variable `OPENAI_API_KEY` is unset or empty - models[0] ("gpt-4o-mini"): `provider_key` references unknown provider key "openai-main" (no provider_keys are defined in this file) - api_keys[0] ("quickstart-caller"): `key_env` environment variable `CALLER_API_KEY` is unset or empty ``` A guardrail that loads but cannot run is reported the same way, with the reason the row would be dropped. A `custom` script error carries the engine's own line and column: ``` resources file resources.yaml: 1 guardrail(s) load but cannot run: - guardrails ("screen-with-my-service"): custom guardrail script does not compile: Error: unsupported keyword: export at guardrail.js:12:3 ``` ### Validate with the Container Image[​](#validate-with-the-container-image "Direct link to Validate with the Container Image") Without a local binary, run the same check through Docker. Mount the file and pass the environment variables it references: ``` docker run --rm \ -v "$(pwd)/resources.yaml:/etc/aisix/resources.yaml:ro" \ -e OPENAI_API_KEY \ -e CALLER_API_KEY \ --entrypoint /usr/local/bin/aisix \ ghcr.io/api7/aisix:1.0.0 \ validate --resources /etc/aisix/resources.yaml ``` ## Export Resources from etcd[​](#export-resources-from-etcd "Direct link to Export Resources from etcd") `aisix export` reads resources from an open-source AISIX gateway's etcd prefix and writes them to a declarative resources file. Use it to move an etcd-backed gateway to a resources file. This includes gateways whose resources were created through the gateway Admin API in earlier releases. The command can also create a reviewable backup of resources the gateway can load. ``` aisix export \ --etcd "http://127.0.0.1:2379" \ --output resources.yaml ``` | Option | Required | Description | | ----------------------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | `--etcd <ENDPOINT>` | Yes | etcd endpoint to read. Repeat the option or provide a comma-separated list for multiple endpoints. | | `--prefix <PREFIX>` | No | Key prefix containing the resources. Defaults to `/aisix`, the gateway's default `etcd.prefix`. | | `-o`, `--output <FILE>` | No | Write the YAML document to a file instead of standard output. On Unix, AISIX creates or resets the file with mode `0600`. | | `--reveal-secrets` | No | Write stored credentials inline instead of replacing them with environment-variable placeholders. The resulting output contains live secrets. | The export follows the same etcd decoding path as the running gateway. It converts resource references back to display names and omits generated IDs so the resulting document follows the resources-file format. By default, stored credentials are replaced with `${VAR}` placeholders. AISIX prints the placeholder names and their source fields to standard error. Set those variables in the gateway environment before loading the exported file. caution Use `--reveal-secrets` only for a controlled migration where plaintext credentials must remain in the file. Do not commit, publish, or copy that output to an unsecured location. The command reports etcd entries that it could not decode and warnings about resource relationships. It still writes the output for inspection. It exits non-zero when naming collisions or dangling references prevent the file from loading. Validate the completed file before switching the gateway to it: ``` aisix validate --resources resources.yaml ``` --- # AISIX Cloud Admin API Changelog Every released version of the AISIX Cloud Admin API is published in full, so you can read the API exactly as the version you run exposes it, and see what changed before you upgrade. The sections below list the differences between adjacent releases; a change marked **Breaking** requires a caller update. * [AISIX Cloud Admin API 1.0.0](https://docs.api7.ai/ai-gateway/1.0.0/reference/cloud-admin-api) — the API as shipped in 1.0.0 * [AISIX Cloud Admin API 0.13.0](https://docs.api7.ai/ai-gateway/0.13.0/reference/cloud-admin-api) — the API as shipped in 0.13.0 * [AISIX Cloud Admin API 0.12.0](https://docs.api7.ai/ai-gateway/0.12.0/reference/cloud-admin-api) — the API as shipped in 0.12.0 * [AISIX Cloud Admin API 0.11.0](https://docs.api7.ai/ai-gateway/0.11.0/reference/cloud-admin-api) — the API as shipped in 0.11.0 * [AISIX Cloud Admin API 0.10.0](https://docs.api7.ai/ai-gateway/0.10.0/reference/cloud-admin-api) — the API as shipped in 0.10.0 * [AISIX Cloud Admin API 0.9.0](https://docs.api7.ai/ai-gateway/0.9.0/reference/cloud-admin-api) — the API as shipped in 0.9.0 Show allOnly Breakings ## From 0.13.0 to 1.0.0[​](#from-0130-to-100 "Direct link to From 0.13.0 to 1.0.0") ### Modified220<!-- -->Breakings Switch to `Show all` to see all changes GET/environments/{env\_id}/guardrails * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Modifiedproperty`enforcement_mode` * Extensions changed * Modified extension: x-enumDescriptions * Modified /monitor from 'Record what would have happened without blocking or redacting caller-visible content.' to 'Record what the verdict would have done, without blocking or redacting caller-visible content. Does not exempt the row from a refusal raised because the gateway could not read the content the guardrail would have scanned — that carries no verdict to downgrade, and`fail_open`governs it.' POST/environments/{env\_id}/guardrails * Request Body * Modifiedproperty`enforcement_mode` * Extensions changed * Modified extension: x-enumDescriptions * Modified /monitor from 'Record what would have happened without blocking or redacting caller-visible content.' to 'Record what the verdict would have done, without blocking or redacting caller-visible content. Does not exempt the row from a refusal raised because the gateway could not read the content the guardrail would have scanned — that carries no verdict to downgrade, and`fail_open`governs it.' * Responses * Modified response: 201 * Modifiedproperty`guardrail` * Modifiedproperty`enforcement_mode` * Extensions changed * Modified extension: x-enumDescriptions * Modified /monitor from 'Record what would have happened without blocking or redacting caller-visible content.' to 'Record what the verdict would have done, without blocking or redacting caller-visible content. Does not exempt the row from a refusal raised because the gateway could not read the content the guardrail would have scanned — that carries no verdict to downgrade, and`fail_open`governs it.' GET/environments/{env\_id}/guardrails/{guardrail\_id} * Responses * Modified response: 200 * Modifiedproperty`guardrail` * Modifiedproperty`enforcement_mode` * Extensions changed * Modified extension: x-enumDescriptions * Modified /monitor from 'Record what would have happened without blocking or redacting caller-visible content.' to 'Record what the verdict would have done, without blocking or redacting caller-visible content. Does not exempt the row from a refusal raised because the gateway could not read the content the guardrail would have scanned — that carries no verdict to downgrade, and`fail_open`governs it.' PATCH/environments/{env\_id}/guardrails/{guardrail\_id} * Request Body * Modifiedproperty`enforcement_mode` * Extensions changed * Modified extension: x-enumDescriptions * Modified /monitor from 'Record what would have happened without blocking or redacting caller-visible content.' to 'Record what the verdict would have done, without blocking or redacting caller-visible content. Does not exempt the row from a refusal raised because the gateway could not read the content the guardrail would have scanned — that carries no verdict to downgrade, and`fail_open`governs it.' * Responses * Modified response: 200 * Modifiedproperty`guardrail` * Modifiedproperty`enforcement_mode` * Extensions changed * Modified extension: x-enumDescriptions * Modified /monitor from 'Record what would have happened without blocking or redacting caller-visible content.' to 'Record what the verdict would have done, without blocking or redacting caller-visible content. Does not exempt the row from a refusal raised because the gateway could not read the content the guardrail would have scanned — that carries no verdict to downgrade, and`fail_open`governs it.' GET/environments/{env\_id}/models * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Newproperty`effort_mapping` * Modifiedproperty`cooldown` * Modifiedproperty`enabled` * Default changed from null to false POST/environments/{env\_id}/models * Request Body * Newproperty`effort_mapping` * Modifiedproperty`cooldown` * Modifiedproperty`enabled` * Default changed from null to false * Responses * Modified response: 201 * Modifiedproperty`model` * Newproperty`effort_mapping` * Modifiedproperty`cooldown` * Modifiedproperty`enabled` * Default changed from null to false GET/environments/{env\_id}/models/{model\_id} * Responses * Modified response: 200 * Modifiedproperty`model` * Newproperty`effort_mapping` * Modifiedproperty`cooldown` * Modifiedproperty`enabled` * Default changed from null to false PATCH/environments/{env\_id}/models/{model\_id} * Request Body * Newproperty`effort_mapping` * Modifiedproperty`cooldown` * Modifiedproperty`enabled` * Default changed from null to false * Responses * Modified response: 200 * Modifiedproperty`model` * Newproperty`effort_mapping` * Modifiedproperty`cooldown` * Modifiedproperty`enabled` * Default changed from null to false GET/environments/{env\_id}/passthrough\_routes * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Newproperty`forward_client_headers` POST/environments/{env\_id}/passthrough\_routes * Request Body * Newproperty`forward_client_headers` * Responses * Modified response: 201 * Modifiedproperty`passthrough_route` * Newproperty`forward_client_headers` GET/environments/{env\_id}/passthrough\_routes/{passthrough\_route\_id} * Responses * Modified response: 200 * Modifiedproperty`passthrough_route` * Newproperty`forward_client_headers` PATCH/environments/{env\_id}/passthrough\_routes/{passthrough\_route\_id} * Request Body * Newproperty`forward_client_headers` * Responses * Modified response: 200 * Modifiedproperty`passthrough_route` * Newproperty`forward_client_headers` GET/environments/{env\_id}/usage\_events * Newquery param`requested_model_exact` * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Newproperty`guardrail_scores` GET/environments/{env\_id}/usage\_events/export * Newquery param`requested_model_exact` * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Modified schema: #/components/schemas/UsageEvent * Newproperty`guardrail_scores` POST/mcp\_server\_submissions * Request Body * Newproperty`forward_client_headers` * Responses * Modified response: 201 * Modifiedproperty`mcp_server` * Newproperty`forward_client_headers` PATCH/mcp\_server\_submissions/{mcp\_server\_id} * Request Body * Newproperty`forward_client_headers` * Responses * Modified response: 200 * Modifiedproperty`mcp_server` * Newproperty`forward_client_headers` GET/mcp\_servers * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Newproperty`forward_client_headers` POST/mcp\_servers * Request Body * Newproperty`forward_client_headers` * Responses * Modified response: 201 * Modifiedproperty`mcp_server` * Newproperty`forward_client_headers` GET/mcp\_servers/{mcp\_server\_id} * Responses * Modified response: 200 * Modifiedproperty`mcp_server` * Newproperty`forward_client_headers` PATCH/mcp\_servers/{mcp\_server\_id} * Request Body * Newproperty`forward_client_headers` * Responses * Modified response: 200 * Modifiedproperty`mcp_server` * Newproperty`forward_client_headers` POST/mcp\_servers/{mcp\_server\_id}/approve * Responses * Modified response: 200 * Modifiedproperty`mcp_server` * Newproperty`forward_client_headers` POST/mcp\_servers/{mcp\_server\_id}/reject * Responses * Modified response: 200 * Modifiedproperty`mcp_server` * Newproperty`forward_client_headers` ## From 0.12.0 to 0.13.0[​](#from-0120-to-0130 "Direct link to From 0.12.0 to 0.13.0") ### Modified50<!-- -->Breakings Switch to `Show all` to see all changes GET/environments/{env\_id}/usage\_events * Newquery param`operation` * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Newproperty`operation` GET/environments/{env\_id}/usage\_events/export * Newquery param`operation` * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Modified schema: #/components/schemas/UsageEvent * Newproperty`operation` POST/provider\_keys * Request Body * Newproperty`apis` * Responses * Modified response: 201 * Newproperty`warnings` GET/provider\_keys/{provider\_key\_id} * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`provider_key` * Newproperty`apis` * Newproperty`apis_source` PATCH/provider\_keys/{provider\_key\_id} * Request Body * Newproperty`apis` * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`provider_key` * Newproperty`apis` * Newproperty`apis_source` ## From 0.11.0 to 0.12.0[​](#from-0110-to-0120 "Direct link to From 0.11.0 to 0.12.0") ### New21 GET/environments/{env\_id}/usage\_events GET/environments/{env\_id}/usage\_events/export GET/environments/{env\_id}/usage\_metrics GET/environments/{env\_id}/usage\_summary GET/invitations POST/invitations DELETE/invitations/{invitation\_id} DELETE/members/{user\_id} GET/model\_pricing PUT/model\_pricing DELETE/model\_pricing/{id} GET/notification\_deliveries GET/teams POST/teams DELETE/teams/{team\_id} GET/teams/{team\_id} PATCH/teams/{team\_id} GET/teams/{team\_id}/members POST/teams/{team\_id}/members DELETE/teams/{team\_id}/members/{user\_id} PATCH/teams/{team\_id}/members/{user\_id} ### Modified20<!-- -->Breakings Switch to `Show all` to see all changes DELETE/environments/{env\_id}/api\_keys/{api\_key\_id} * Responses * New response: 409 DELETE/environments/{env\_id}/oidc\_providers/{oidc\_provider\_id} * Responses * Modified response: 200 * Newproperty`warnings` ## From 0.10.0 to 0.11.0[​](#from-0100-to-0110 "Direct link to From 0.10.0 to 0.11.0") ### Modified80<!-- -->Breakings Switch to `Show all` to see all changes PATCH/environments/{env\_id} * Request Body * Newproperty`display_name` GET/environments/{env\_id}/dp\_nodes * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Required changed * Newrequired property`status` * Newproperty`status` GET/environments/{env\_id}/guardrails * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Deletedproperty`mandatory` * Modifiedproperty`kind` * New enum values: \[semantic custom] POST/environments/{env\_id}/guardrails * Request Body * Deletedproperty`mandatory` * Modifiedproperty`kind` * New enum values: \[semantic custom] * Responses * Modified response: 201 * Newproperty`warnings` * Modifiedproperty`guardrail` * Deletedproperty`mandatory` * Modifiedproperty`kind` * New enum values: \[semantic custom] POST/environments/{env\_id}/guardrails/test-connection * Request Body * Modifiedproperty`kind` * New enum values: \[semantic custom] GET/environments/{env\_id}/guardrails/{guardrail\_id} * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`guardrail` * Deletedproperty`mandatory` * Modifiedproperty`kind` * New enum values: \[semantic custom] PATCH/environments/{env\_id}/guardrails/{guardrail\_id} * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`guardrail` * Deletedproperty`mandatory` * Modifiedproperty`kind` * New enum values: \[semantic custom] GET/environments/{env\_id}/mcp\_policy * Responses * Modified response: 200 * Modifiedproperty`mcp_policy`Breaking * Schemas added: #/components/schemas/McpAccessPolicy * Type changed from 'object' to '' * AdditionalProperties changed from false to null * Nullable changed from false to true * Required changed * Deleted required property: allow * Deleted required property: created\_at * Deleted required property: deny * Deleted required property: enabled * Deleted required property: id * Deleted required property: scope * Deleted required property: updated\_at * Deletedproperty`allow`Breaking * Deletedproperty`created_at`Breaking * Deletedproperty`deny`Breaking * Deletedproperty`enabled`Breaking * Deletedproperty`env_id` * Deletedproperty`id`Breaking * Deletedproperty`scope`Breaking * Deletedproperty`team_id` * Deletedproperty`updated_at`Breaking ## From 0.9.0 to 0.10.0[​](#from-090-to-0100 "Direct link to From 0.9.0 to 0.10.0") ### New13 PATCH/environments/{env\_id} GET/environments/{env\_id}/passthrough\_routes POST/environments/{env\_id}/passthrough\_routes DELETE/environments/{env\_id}/passthrough\_routes/{passthrough\_route\_id} GET/environments/{env\_id}/passthrough\_routes/{passthrough\_route\_id} PATCH/environments/{env\_id}/passthrough\_routes/{passthrough\_route\_id} GET/members/{member\_id}/role\_bindings PUT/members/{member\_id}/role\_bindings PATCH/members/{user\_id} GET/roles POST/roles DELETE/roles/{role\_name} PATCH/roles/{role\_name} ### Deleted2 POST/environments/{env\_id}/mcp\_policy/apply POST/environments/{env\_id}/mcp\_policy/preview ### Modified280<!-- -->Breakings Switch to `Show all` to see all changes GET/environments * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Newproperty`mcp_anonymous` * Newproperty`mcp_resource_url` POST/environments * Responses * Modified response: 201 * Modifiedproperty`environment` * Newproperty`mcp_anonymous` * Newproperty`mcp_resource_url` GET/environments/{env\_id} * Responses * Modified response: 200 * Modifiedproperty`environment` * Newproperty`mcp_anonymous` * Newproperty`mcp_resource_url` GET/environments/{env\_id}/api\_keys * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Newproperty`allowed_routes` * Deletedproperty`allowed_tools` * Modifiedproperty`mcp_access` * Modified schema: #/components/schemas/McpAccess * Required changed * Newrequired property`allow` * Deleted required property: mode * Deletedproperty`mode`Breaking POST/environments/{env\_id}/api\_keys * Request Body * Newproperty`allowed_routes` * Deletedproperty`allowed_tools` * Modifiedproperty`mcp_access` * Modified schema: #/components/schemas/McpAccess * Required changed * Newrequired property`allow`Breaking * Deleted required property: mode * Deletedproperty`mode`Breaking * Responses * Modified response: 201 * Newproperty`warnings` * Modifiedproperty`api_key` * Newproperty`allowed_routes` * Deletedproperty`allowed_tools` * Modifiedproperty`mcp_access` * Modified schema: #/components/schemas/McpAccess * Required changed * Newrequired property`allow`Breaking * Deleted required property: mode * Deletedproperty`mode`Breaking PATCH/environments/{env\_id}/api\_keys/{api\_key\_id} * Request Body * Newproperty`allowed_routes` * Deletedproperty`allowed_tools` * Modifiedproperty`mcp_access` * Modified schema: #/components/schemas/McpAccess * Required changed * Newrequired property`allow`Breaking * Deleted required property: mode * Deletedproperty`mode`Breaking * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`api_key` * Newproperty`allowed_routes` * Deletedproperty`allowed_tools` * Modifiedproperty`mcp_access` * Modified schema: #/components/schemas/McpAccess * Required changed * Newrequired property`allow`Breaking * Deleted required property: mode * Deletedproperty`mode`Breaking GET/environments/{env\_id}/api\_keys/{api\_key\_id}/effective\_permissions * Responses * Modified response: 200 * Modifiedproperty`effective_permissions` * Modifiedproperty`mcp` * Required changed * Newrequired property`layers` * Deleted required property: base\_source * Deleted required property: key\_mode * Newproperty`layers` * Deletedproperty`base_policy_id` * Deletedproperty`base_source`Breaking * Deletedproperty`key_mode`Breaking POST/environments/{env\_id}/api\_keys/{api\_key\_id}/rotate * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`api_key` * Newproperty`allowed_routes` * Deletedproperty`allowed_tools` * Modifiedproperty`mcp_access` * Modified schema: #/components/schemas/McpAccess * Required changed * Newrequired property`allow` * Deleted required property: mode * Deletedproperty`mode`Breaking GET/environments/{env\_id}/guardrails/{guardrail\_id}/attachments * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Modifiedproperty`scope_type` * Extensions changed * Modified extension: x-enumDescriptions * Added /mcp\_server with value: 'Apply to MCP tool calls routed to the MCP server identified by`scope_id`; the tool arguments and the tool result are both inspected.' * Added /passthrough\_route with value: 'Apply to traffic served by the passthrough route identified by`scope_id`; the forwarded request body and the relayed response (including streamed frames, with hold-back) are both inspected.' * New enum values: \[mcp\_server passthrough\_route] POST/environments/{env\_id}/guardrails/{guardrail\_id}/attachments * Request Body * Modifiedproperty`scope_type` * Extensions changed * Modified extension: x-enumDescriptions * Added /mcp\_server with value: 'Apply to MCP tool calls routed to the MCP server identified by`scope_id`; the tool arguments and the tool result are both inspected.' * Added /passthrough\_route with value: 'Apply to traffic served by the passthrough route identified by`scope_id`; the forwarded request body and the relayed response (including streamed frames, with hold-back) are both inspected.' * New enum values: \[mcp\_server passthrough\_route] * Responses * Modified response: 201 * Newproperty`warnings` * Modifiedproperty`data` * Modifiedproperty`scope_type` * Extensions changed * Modified extension: x-enumDescriptions * Added /mcp\_server with value: 'Apply to MCP tool calls routed to the MCP server identified by`scope_id`; the tool arguments and the tool result are both inspected.' * Added /passthrough\_route with value: 'Apply to traffic served by the passthrough route identified by`scope_id`; the forwarded request body and the relayed response (including streamed frames, with hold-back) are both inspected.' * New enum values: \[mcp\_server passthrough\_route] GET/environments/{env\_id}/mcp\_policy * Responses * Modified response: 200 * Modifiedproperty`mcp_policy` * Required changed * Deleted required property: mode * Deletedproperty`mode`Breaking PUT/environments/{env\_id}/mcp\_policy * Request Body * Required changed * Newrequired property`allow`Breaking * Deleted required property: mode * Deletedproperty`mode`Breaking * Responses * Modified response: 200 * Modifiedproperty`mcp_policy` * Required changed * Deleted required property: mode * Deletedproperty`mode`Breaking GET/environments/{env\_id}/models * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Modifiedproperty`routing` * Newproperty`hash_on` * Deletedproperty`sticky` * Modifiedproperty`strategy` * New enum values: \[consistent\_hash] * Deleted enum values: \[weighted] * Modifiedproperty`targets` * Items changed * Newproperty`priority` POST/environments/{env\_id}/models * Request Body * Modifiedproperty`routing` * Newproperty`hash_on` * Deletedproperty`sticky` * Modifiedproperty`strategy` * New enum values: \[consistent\_hash] * Deleted enum values: \[weighted] * Modifiedproperty`targets` * Items changed * Newproperty`priority` * Responses * Modified response: 201 * Modifiedproperty`model` * Modifiedproperty`routing` * Newproperty`hash_on` * Deletedproperty`sticky` * Modifiedproperty`strategy` * New enum values: \[consistent\_hash] * Deleted enum values: \[weighted] * Modifiedproperty`targets` * Items changed * Newproperty`priority` GET/environments/{env\_id}/models/{model\_id} * Responses * Modified response: 200 * Modifiedproperty`model` * Modifiedproperty`routing` * Newproperty`hash_on` * Deletedproperty`sticky` * Modifiedproperty`strategy` * New enum values: \[consistent\_hash] * Deleted enum values: \[weighted] * Modifiedproperty`targets` * Items changed * Newproperty`priority` PATCH/environments/{env\_id}/models/{model\_id} * Request Body * Modifiedproperty`routing` * Newproperty`hash_on` * Deletedproperty`sticky` * Modifiedproperty`strategy` * New enum values: \[consistent\_hash] * Deleted enum values: \[weighted] * Modifiedproperty`targets` * Items changed * Newproperty`priority` * Responses * Modified response: 200 * Modifiedproperty`model` * Modifiedproperty`routing` * Newproperty`hash_on` * Deletedproperty`sticky` * Modifiedproperty`strategy` * New enum values: \[consistent\_hash] * Deleted enum values: \[weighted] * Modifiedproperty`targets` * Items changed * Newproperty`priority` POST/mcp\_server\_submissions * Request Body * Newproperty`protocol_version` * Responses * Modified response: 201 * Newproperty`warnings` * Modifiedproperty`mcp_server` * Newproperty`protocol_version` PATCH/mcp\_server\_submissions/{mcp\_server\_id} * Request Body * Newproperty`protocol_version` * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`mcp_server` * Newproperty`protocol_version` GET/mcp\_servers * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Newproperty`protocol_version` POST/mcp\_servers * Request Body * Newproperty`protocol_version` * Responses * Modified response: 201 * Newproperty`warnings` * Modifiedproperty`mcp_server` * Newproperty`protocol_version` GET/mcp\_servers/{mcp\_server\_id} * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`mcp_server` * Newproperty`protocol_version` PATCH/mcp\_servers/{mcp\_server\_id} * Request Body * Newproperty`protocol_version` * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`mcp_server` * Newproperty`protocol_version` POST/mcp\_servers/{mcp\_server\_id}/approve * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`mcp_server` * Newproperty`protocol_version` POST/mcp\_servers/{mcp\_server\_id}/reject * Responses * Modified response: 200 * Newproperty`warnings` * Modifiedproperty`mcp_server` * Newproperty`protocol_version` GET/members * Responses * Modified response: 200 * Modifiedproperty`data` * Items changed * Modifiedproperty`role` * Deleted enum values: \[owner admin member] * MinLength changed from 0 to 1 POST/members * Responses * Modified response: 201 * Modifiedproperty`member` * Modifiedproperty`role` * Deleted enum values: \[owner admin member] * MinLength changed from 0 to 1 GET/teams/{team\_id}/entitlements * Responses * Modified response: 200 * Modifiedproperty`entitlements` * Modifiedproperty`mcp` * Modified schema: #/components/schemas/McpAccessPolicy * Required changed * Deleted required property: mode * Deletedproperty`mode`Breaking PUT/teams/{team\_id}/entitlements * Request Body * Modifiedproperty`mcp` * Modified schema: #/components/schemas/PutMcpPolicyRequest * Required changed * Newrequired property`allow`Breaking * Deleted required property: mode * Deletedproperty`mode`Breaking * Responses * Modified response: 200 * Modifiedproperty`entitlements` * Modifiedproperty`mcp` * Modified schema: #/components/schemas/McpAccessPolicy * Required changed * Deleted required property: mode * Deletedproperty`mode`Breaking --- # Configuration Status The AISIX gateway reports whether its configuration took effect — and if not, why — through status endpoints and Prometheus metrics. These surfaces answer the operator question "is the gateway serving the configuration I intended?" They cover every resource source: a [`resources.yaml` file](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md), an etcd configuration store you manage, or the AISIX Cloud control plane. These endpoints are served on the dedicated metrics listener, `observability.metrics.prometheus.addr` (default `0.0.0.0:9090`), alongside `GET /metrics`. Like the metrics endpoint, they are unauthenticated by design; keep the listener private to your monitoring network. | Endpoint | Purpose | | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `GET /status/config` | Full configuration load state: derived state, source and applied hashes, resource counts, reload results, rejected entries, and partially compatible resources. | | `GET /status/ready` | Readiness gate: `503` until the first valid configuration is applied, `200 ok` afterward. | | `GET /status/models` | Per-model runtime health: one row per configured model with its rotation status. | ## GET `/status/config`[​](#get-statusconfig "Direct link to get-statusconfig") Returns a JSON document describing the last observed and last applied configuration: ``` curl -sS "http://127.0.0.1:9090/status/config" ``` ``` { "state": "synced", "source": { "type": "file", "source_hash": "1dc0ee8d06edcde3ecbf23672858622a83f266910846f514ccb909cf41046653", "observed_at": "YYYY-MM-DDTHH:MM:SSZ" }, "applied": { "config_hash": "1dc0ee8d06edcde3ecbf23672858622a83f266910846f514ccb909cf41046653", "apply_seq": 1, "applied_at": "YYYY-MM-DDTHH:MM:SSZ", "resource_counts": { "api_keys": 1, "models": 1, "provider_keys": 1 } }, "last_reload": { "successful": true, "at": "YYYY-MM-DDTHH:MM:SSZ" }, "last_failure": null, "rejected": [], "partially_compatible": [] } ``` ### Configuration States[​](#configuration-states "Direct link to Configuration States") `state` is derived by the gateway from the last observed and applied snapshots: | State | Meaning | | -------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `synced` | The applied configuration matches the latest snapshot observed from the source, and no resource was rejected. Check `partially_compatible` to determine whether the gateway ignored fields it does not recognize. | | `degraded` | The gateway is serving, but the latest snapshot carried entries it rejected. Accepted resources and any last known good values remain in service; the `rejected` array explains each rejected entry. | | `out_of_sync` | The latest observed snapshot was rejected as a whole. The gateway keeps serving the last valid configuration. | | `empty` | A valid configuration was applied, but it holds zero resources. | | `never_loaded` | No valid configuration has been applied since the process started. | A gateway configured from a resources file applies the file all or nothing, so a failed file reload reports `out_of_sync` rather than `degraded`. `degraded` occurs when the source delivers resources individually and only some of them are invalid. Etcd-backed configuration reads are forward compatible. If a document contains an unrecognized field, the gateway serves the recognized fields and reports the ignored field in `partially_compatible`. This alone does not change `state` from `synced`. Resources-file validation remains strict, so an unknown field in a resources file rejects the whole reload. ### Response Fields[​](#response-fields "Direct link to Response Fields") Top-level fields: | Field | Type | Description | | ---------------------- | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `state` | string | Derived configuration state. One of `synced`, `degraded`, `out_of_sync`, `empty`, `never_loaded`. | | `source` | object | The latest snapshot observed from the configuration source. | | `applied` | object | The last configuration actually applied and served. Omitted while `state` is `never_loaded`. | | `last_reload` | object | Outcome of the most recent load. Omitted before the first load completes. | | `last_failure` | object or null | The most recent load failure since the process started. Sticky: it remains populated after a later successful reload, until restart. | | `rejected` | array | Entries the gateway rejected from the latest snapshot. Empty when everything loaded. | | `partially_compatible` | array | Served etcd resources that contain fields this gateway version does not recognize, aggregated by resource kind and field. Empty when every served document matches the gateway's schema. | `source` fields: | Field | Type | Description | | ------------------- | ------- | -------------------------------------------------------------------------------------------------------- | | `type` | string | Where configuration is read from: `file` or `etcd`. | | `connected` | boolean | Whether the configuration store is reachable. Present only when `type` is `etcd`. | | `observed_revision` | number | Store revision of the latest observed snapshot. Present only when `type` is `etcd`. | | `source_hash` | string | SHA-256 hash of the latest observed snapshot. For a file source, this is the hash of the raw file bytes. | | `observed_at` | string | RFC 3339 UTC timestamp of the latest observation. | `applied` fields: | Field | Type | Description | | ------------------ | ------ | ----------------------------------------------------------------------------------------------------------- | | `applied_revision` | number | Store revision the applied configuration reflects. Present only when `source.type` is `etcd`. | | `config_hash` | string | SHA-256 hash of the accepted, served configuration. Equal to `source_hash` when nothing was rejected. | | `apply_seq` | number | Counter that increments each time the applied configuration changes. Unchanged content does not advance it. | | `applied_at` | string | RFC 3339 UTC timestamp of the last applied change. | | `resource_counts` | object | Served resource count per kind, for example `{"models": 2}`. | `last_reload` and `last_failure` fields: | Field | Type | Description | | ------------------------------ | ------- | ---------------------------------------------------------- | | `last_reload.successful` | boolean | Whether the most recent load completed without rejections. | | `last_reload.at` | string | RFC 3339 UTC timestamp of the most recent load. | | `last_failure.at` | string | When the most recent failure occurred. | | `last_failure.last_error_kind` | string | Failure kind of the most recent failure. | | `last_failure.last_error` | string | Human-readable message of the most recent failure. | Each entry in `rejected`: | Field | Type | Description | | --------------------------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------- | | `resource_kind` | string | Plural resource kind, such as `models` or `provider_keys`. Empty when the source entry could not be attributed to a kind. | | `resource_id` | string | Resource ID. Empty when the source entry could not be parsed far enough to identify it. | | `last_error_kind` | string | Failure kind: `bad_key`, `non_json`, `schema_failed`, `parse_failed`, or `unknown_kind`. | | `last_error` | string | Human-readable error. Schema messages mask credential values. | | `first_seen_at` | string | When this rejection was first observed since the process started. Stable across repeated reloads of the same bad entry. | | `last_seen_at` | string | When this rejection was most recently observed. | | `serving_stale_since` | string | RFC 3339 timestamp from which the gateway has kept serving the resource's last known good value. Omitted when no previous value remains in service. | | `serving_stale_age_seconds` | number | Seconds elapsed since `serving_stale_since`, recalculated when the status is read. Omitted together with `serving_stale_since`. | Each entry in `partially_compatible`: | Field | Type | Description | | --------------- | ------ | --------------------------------------------------------------------------------------------------- | | `resource_kind` | string | Plural resource kind, such as `api_keys` or `models`. | | `field` | string | Ignored field path. Array indexes are normalized to `[]`, for example `routing.targets[].priority`. | | `count` | number | Served resources of this kind that contain the ignored field. | Use `config_hash` to confirm a specific change landed: the hash is deterministic, so a deployment pipeline that knows what it shipped can compare hashes instead of diffing resources. For a file source, `source_hash` is the SHA-256 of the file bytes (`sha256sum resources.yaml`). For an etcd source, a matching hash does not mean every field is enforced; also require `partially_compatible` to be empty when exact schema compatibility matters. ## GET `/status/ready`[​](#get-statusready "Direct link to get-statusready") A readiness gate for the configuration source only: ``` curl -sSi "http://127.0.0.1:9090/status/ready" ``` | Condition | Status | Body | | -------------------------------------- | ------------------------- | ---------------------------- | | No valid configuration applied yet | `503 Service Unavailable` | `no configuration available` | | A valid configuration has been applied | `200 OK` | `ok` | Use it as a startup or readiness probe so a gateway does not receive traffic before it can serve configured routes. It stays `200` once the first configuration is applied, including while a later reload fails and the gateway serves the last valid configuration. The proxy listener's `/livez` and `/readyz` keep their process-level semantics; see [Health Checks](https://docs.api7.ai/ai-gateway/deployment/health-checks.md). ## GET `/status/models`[​](#get-statusmodels "Direct link to get-statusmodels") The per-model runtime health view — one row per configured model: ``` curl -sS "http://127.0.0.1:9090/status/models" ``` ``` [ { "id": "9a3f2c67-52b8-4b1e-9f4e-1f2f3a4b5c6d", "display_name": "gpt-4o-prod", "kind": "direct", "status": "healthy" }, { "id": "5b17e9d2-8a44-4c05-b7a1-0c9d8e7f6a5b", "display_name": "claude-prod", "kind": "direct", "status": "cooldown", "status_reason": "upstream_auth_failure", "cooldown_until": { "secs_since_epoch": 1784708130, "nanos_since_epoch": 0 } } ] ``` | Status | Meaning | | ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `healthy` | The model is in rotation. | | `cooldown` | Recent upstream failures took the model out of rotation until `cooldown_until`; `status_reason` names the failure class, for example `upstream_auth_failure` or `upstream_rate_limited`. Only a model whose `cooldown` block sets `enabled: true` enters this state. | | `unhealthy` | Recent background model checks failed; routing avoids the model. `last_check_status` carries the HTTP status of the most recent check. | | `not_applicable` | The row is a multi-target or other virtual model; its availability derives from its target models' rows. | Timestamps (`cooldown_until`, and `last_checked_at` on background-checked models) are seconds/nanoseconds epoch objects as shown above, not RFC 3339 strings. Like the other status endpoints, `/status/models` is unauthenticated and reflects the applied configuration, so it works in every gateway deployment. ## Configuration Load Metrics[​](#configuration-load-metrics "Direct link to Configuration Load Metrics") The `GET /metrics` endpoint on the same listener exposes the configuration load state as Prometheus series. Values refresh at scrape time from the same state that backs `GET /status/config`. | Metric | Type | Labels | Description | | ---------------------------------------------------- | ------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `aisix_config_last_reload_successful` | gauge | None | `1` when the most recent load completed without rejections, otherwise `0`. | | `aisix_config_last_reload_success_timestamp_seconds` | gauge | None | Unix timestamp of the most recent successful load. | | `aisix_config_reloads_total` | counter | None | Configuration loads since the process started, counting boot loads, file reloads, and full synchronizations from the configuration store. | | `aisix_config_reload_failures_total` | counter | `reason` | Loads that did not fully succeed, bucketed by `reason`: `fetch` (source unreachable or unreadable), `parse` (source content could not be parsed), or `validate` (a resource failed schema, shape, or reference validation). | | `aisix_config_rejected_resources` | gauge | `kind` | Currently rejected entries per resource kind. `0` after the offending entries are fixed. | | `aisix_config_partially_compatible_resources` | gauge | `kind` | Served resources carrying at least one ignored field, grouped by resource kind. A resource with several ignored fields counts once. | | `aisix_config_stale_served_resources` | gauge | `kind` | Resources whose latest source value was rejected while the last known good value remains in service, grouped by resource kind. | | `aisix_config_hash_info` | gauge | `hash` | Info-style series: exactly one live sample with value `1`, whose `hash` label is the applied `config_hash`. | | `aisix_config_observed_revision` | gauge | None | Store revision of the latest observed snapshot. Emitted only for an `etcd` source. | | `aisix_config_applied_revision` | gauge | None | Store revision of the applied configuration. Emitted only for an `etcd` source. | | `aisix_config_source_connected` | gauge | None | `1` when the configuration store is reachable. Emitted only for an `etcd` source. | For the full metric catalog, see [Metrics Reference](https://docs.api7.ai/ai-gateway/reference/metrics.md). ### Example Alerts[​](#example-alerts "Direct link to Example Alerts") Alert when the gateway rejects any configured resource — the gateway keeps serving, but something an operator wrote is not in effect: ``` - alert: AisixConfigRejectedResources expr: sum by (instance) (aisix_config_rejected_resources) > 0 for: 5m labels: severity: warning annotations: summary: "AISIX gateway is rejecting configured resources" description: "Check GET /status/config on {{ $labels.instance }}: the rejected array names each entry and its error." ``` Alert when configuration reloads keep failing — the gateway is running on the last valid configuration and new changes are not taking effect: ``` - alert: AisixConfigReloadFailing expr: aisix_config_last_reload_successful == 0 for: 10m labels: severity: warning annotations: summary: "AISIX gateway configuration reloads are failing" description: "The last configuration load on {{ $labels.instance }} did not fully succeed. Check last_failure and rejected in GET /status/config." ``` --- # Startup Configuration Reference This reference documents the startup configuration file that defines process-level settings such as listeners, resource-source connectivity, TLS, observability, cache and rate-limit backends, and the AISIX Cloud connection. Models, caller API keys, provider keys, guardrails, cache policies, and observability exporters are not defined in the startup configuration. For an open-source AISIX gateway, a [`resources.yaml` file](#resource-source) or etcd supplies them. For a gateway connected to AISIX Cloud, the control plane supplies them. AISIX accepts YAML, TOML, or JSON startup configuration files. Common files include: * [`config.yaml`](#select-a-configuration-file): the local startup configuration file loaded by AISIX. * [`config.example.yaml`](https://github.com/api7/aisix/blob/v1.0.0/config.example.yaml): a complete store-backed example for a gateway that runs without a control plane. You can copy or mount it as the loaded `config.yaml`. * [`config.managed.yaml`](https://github.com/api7/aisix/blob/v1.0.0/config.managed.yaml): the gateway bootstrap configuration used when AISIX Cloud supplies resources at runtime. The examples below use YAML because the packaged example configurations use YAML. TOML and JSON files can define the same startup fields. For the task-oriented setup flow, see [Startup Configuration](https://docs.api7.ai/ai-gateway/deployment/startup-configuration.md). ## Configuration Models[​](#configuration-models "Direct link to Configuration Models") Every AISIX gateway uses a startup configuration file. The management model determines where its dynamic resources come from and which settings the operator owns directly. | Gateway configuration | Resource source | Select with | How changes arrive | | ----------------------------------------------------- | -------------------------- | ------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- | | Open-source AISIX gateway using a resources file | Declarative file | `resources_file` | Loaded at startup and reloaded on `SIGHUP`. | | Open-source AISIX gateway using a configuration store | An etcd cluster you manage | `etcd` | Initial synchronization followed by etcd watch events. | | AISIX gateway connected to AISIX Cloud | AISIX Cloud control plane | `managed.enabled: true` and the generated connection settings | The gateway receives resources projected by the control plane through its managed etcd connection. | The proxy, downstream, upstream, cache, rate-limit backend, and local observability settings remain gateway startup settings in every model. AISIX Cloud supplies dynamic resources such as models, keys, policies, and exporters; it does not replace these process-level settings. For an open-source AISIX gateway, `admin.enabled` defaults to `true`, but the default `admin.addr` value of `127.0.0.1:0` cannot be used as a listener. To expose the [read-only Admin API](https://docs.api7.ai/ai-gateway/reference/admin-api/.md), set `admin.addr` to a private or loopback address and configure at least one `admin.admin_keys` value. Set `admin.enabled` to `false` to disable the listener. Configure resources through a [resources file](#resources-file) or by writing them directly to [etcd](#etcd-configuration-store). A gateway connected to AISIX Cloud never binds the gateway Admin API. ## Common Startup Configuration[​](#common-startup-configuration "Direct link to Common Startup Configuration") The following example shows an open-source AISIX gateway that uses etcd, with common startup settings. The gateway can instead load dynamic resources from a `resources_file`; see [Resource Source](#resource-source). config.yaml ``` etcd: endpoints: # etcd endpoints used to store dynamic gateway resources. - "http://127.0.0.1:2379" prefix: "/aisix" # Key prefix used by AISIX in etcd. # env_id: "ENVIRONMENT_ID" # Optional environment scope for gateway resources in etcd. # user: "aisix" # Optional etcd username. # password_env: "AISIX_ETCD_PASSWORD" # Environment variable containing the etcd password. # dial_timeout_ms: 5000 # Accepted but not applied in this release; see below. # request_timeout_ms: 5000 # Accepted but not applied in this release; see below. # tls: # mTLS settings; all three certificate fields are required together. # ca_cert_file: "/etc/aisix/mtls/ca.crt" # client_cert_file: "/etc/aisix/mtls/client.crt" # client_key_file: "/etc/aisix/mtls/client.key" # domain_name: "etcd.example.com" # Optional SNI and certificate-name override. proxy: addr: "0.0.0.0:3000" # Address for caller-facing proxy APIs. # request_body_limit_bytes: 0 # Maximum request body size. 0, the default, applies no cap. # Set a byte value to bound per-request memory. See # "Limit Request Body Size" below for rejection behavior. # thread_per_core: true # Serve from independent workers, each with its own listener # and upstream connection pool. Omitted, it is on for Linux # and off elsewhere. Set false to serve from one shared # runtime. Applied at startup; restart to change. # See Deployment > Thread-per-Core Workers. # workers: 4 # Number of proxy worker threads, in either serving mode. # Omitted, it follows the parallelism available to the # process, so a container CPU limit or a taskset affinity # mask sizes it. Must be at least 1. Applied at startup. # tls: # HTTPS certificate and key for the proxy listener. # cert_file: "/etc/aisix/tls/proxy.crt" # key_file: "/etc/aisix/tls/proxy.key" # real_ip: # Caller IP resolution when AISIX runs behind trusted proxies. # trusted_proxies: # - "10.0.0.0/8" # recursive: true # header: "x-forwarded-for" # url_rewrites: # Entry-level path rewriting before routing (first match wins). # - name: per-server-mcp-compat # Optional label used in gateway logs. # match: "^/mcp-servers/([^/]+)/mcp$" # Regex on the raw request path. # rewrite: "/mcp/$1" # Replaces the matched portion; broken rules fail startup. # # See Deployment > URL Rewriting for full semantics. admin: enabled: false observability: service_name: "aisix" # Service name used in telemetry. log_level: "info" # Process log level. metrics: prometheus: enabled: true # Whether to expose Prometheus metrics. path: "/metrics" # Metrics endpoint path. addr: "0.0.0.0:9090" # Dedicated metrics/status listener address. # managed: # Enable when this gateway uses the AISIX Cloud control plane. # enabled: true # cache: # Redis connection used by cache policies that select Redis. # redis: # mode: "single" # url: "redis://127.0.0.1:6379" # # nodes: ["redis://10.0.0.1:6379"] # Cluster mode seed nodes. # # sentinels: ["redis://10.0.0.1:26379"] # Sentinel mode nodes. # # master_name: "mymaster" # Sentinel mode master group. # # username: "default" # Cluster or Sentinel data-node ACL user. # # password: "replace-me" # Cluster or Sentinel data-node ACL password. # # database: 0 # Sentinel master database index. # # For single mode, put credentials in the Redis URL. ratelimit: backend: "memory" # Rate-limit counter backend. Use redis for shared counters across replicas. # redis: # mode: "single" # url: "redis://127.0.0.1:6379" # # nodes: ["redis://10.0.0.1:6379"] # Cluster mode seed nodes. # # sentinels: ["redis://10.0.0.1:26379"] # Sentinel mode nodes. # # master_name: "mymaster" # Sentinel mode master group. # # username: "default" # Cluster or Sentinel data-node ACL user. # # password: "replace-me" # Cluster or Sentinel data-node ACL password. # # database: 0 # Sentinel master database index. # # For single mode, put credentials in the Redis URL. # concurrency_ttl_secs: 300 # Redis backend only. Reclaims stale concurrency slots. upstream: # Outbound calls to providers. Values shown are the defaults. pool_idle_timeout_secs: 30 # Keep below the shortest idle timeout between the gateway and the provider. # timeout_ms: 6000000 # Default request deadline (6000 s) when neither the model nor its group/router sets `timeout`. 0 disables the backstop. # stream_timeout_ms: 0 # Default streaming chunk-gap deadline. 0 falls back to `timeout_ms`. # retries: 2 # Attempts after a retryable failure, when neither the model nor its group/router sets `retries`. 0 disables retrying. # connect_timeout_ms: 5000 # Budget for DNS, TCP, and TLS. 0 disables. # tcp_keepalive_secs: 60 # Idle time before the first keepalive probe. 0 disables. # tcp_keepalive_interval_secs: 30 # Interval between keepalive probes. # tcp_keepalive_retries: 5 # Unacknowledged probes before the connection is dropped. # pool_max_idle_per_host: 32 # Cap on idle connections per upstream host. Unset means # unbounded. Applies per worker to worker-local pools. # tls: # Which certificates the gateway trusts when it calls out. # ca_file: "/etc/aisix/tls/private-ca.pem" # Trusted in addition to the platform's own authorities. # client_cert_file: "/etc/aisix/tls/client.crt" # For supported HTTP upstreams that require mutual TLS. # client_key_file: "/etc/aisix/tls/client.key" # Required together with client_cert_file. # verify: true # false disables verification where supported. Test environments only. downstream: # Connection layer for inbound calls from clients. Values shown are the defaults. idle_timeout_secs: 0 # Close a connection idle between requests. 0 never closes. # sse_keepalive_interval_secs: 15 # Heartbeat interval on a silent streaming response. 0 disables. # Optional deployment-wide override for AWS Bedrock guardrail traffic. # bedrock_endpoint_url: "https://bedrock-runtime.us-east-1.amazonaws.com" ``` Changes to the startup configuration take effect after restarting the gateway. ### Limit Request Body Size[​](#limit-request-body-size "Direct link to Limit Request Body Size") `proxy.request_body_limit_bytes` defaults to `0`, which applies no gateway-side cap. Providers accept larger requests than any single fixed default would allow, so setting a limit can reject a request the selected provider would otherwise serve. Earlier AISIX releases defaulted to `10485760` bytes (10 MiB). Set a byte value when the gateway accepts requests directly from untrusted clients and must bound per-request memory. Alternatively, enforce a body limit at the load balancer or ingress in front of the gateway. AISIX rejects an over-limit request with `413 Content Too Large`, but the caller is not guaranteed to receive that response. For a request that declares an over-limit `Content-Length`, AISIX first drains the body so it can return the `413` on the same connection. That drain is bounded by bytes and time. If the caller does not finish sending within those bounds, the gateway stops reading and the caller usually sees a closed or reset connection instead. A chunked body is rejected while the handler reads it and does not use the same drain path. On the Content-Length path, the `aisix::body_limit` log entry and [`aisix_proxy_request_body_limit_rejections_total`](https://docs.api7.ai/ai-gateway/reference/metrics.md#request-metrics) distinguish a completed drain from a cap, timeout, or client read error. See [Metrics and Logs](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md#collect-access-logs) for the fields and correlation workflow. ### Tune the Upstream Connection Layer[​](#tune-the-upstream-connection-layer "Direct link to Tune the Upstream Connection Layer") The `upstream` block controls how the gateway opens and reuses connections to providers. The defaults suit a gateway that reaches providers directly over the internet. Tune them when the gateway sits behind a load balancer, NAT gateway, corporate proxy, or service mesh. `pool_idle_timeout_secs` is the setting to review first. The gateway reuses pooled connections, so this value must stay **below** the shortest idle timeout anywhere on the path to the provider. If an intermediate hop closes an idle connection sooner than the gateway expires it, the pool eventually hands out a connection the far end has already closed, and the request fails with a transport error against an otherwise healthy provider. TCP keepalive keeps the connection visible to those same hops while a slow model produces its first token. A NAT or load-balancer idle timer can otherwise reap a connection that is legitimately waiting on a long-running request. In thread-per-core serving, the default on Linux, normal proxy dispatches use worker-local connection pools, and `pool_max_idle_per_host` applies per worker. The idle connections a process holds toward one host can therefore reach `proxy.workers` times this cap. Request paths with dedicated clients, such as Provider Keys with custom TLS settings, keep separate process-wide pools and are not included in that multiplication. See [Thread-per-Core Workers](https://docs.api7.ai/ai-gateway/deployment/thread-per-core-workers.md) for the serving modes and their sizing consequences. `timeout_ms` and `stream_timeout_ms` are deployment-wide defaults for the per-model `timeout` and `stream_timeout` fields. A request resolves its deadline from the first level that sets one: the target model, then the routing model or semantic router it was addressed through, then these defaults. The default of 6000 seconds is a backstop, not a responsiveness target — it exists so that an upstream that accepted the connection and then goes silent forever cannot hold a request open indefinitely, while never cutting a legitimate long request (deep-reasoning calls can run past ten minutes; set a per-model `timeout` to enforce anything tighter). A model opts out of the backstop with `timeout: 0`; setting `timeout_ms: 0` removes the default deployment-wide. Zero-value behavior is field-specific. `upstream.timeout_ms: 0` removes the deployment-wide request deadline, and `timeout: 0` on a model stops that request-timeout fallback chain. In contrast, `upstream.stream_timeout_ms: 0` falls back to `upstream.timeout_ms`, while `stream_timeout: 0` on a model defers to the remaining stream-timeout and request-timeout chain. A zero stream-timeout value can therefore leave an effective streaming deadline in place. `upstream.tls` supplies deployment-wide TLS settings for outbound connections. `ca_file` applies to HTTP request paths, the Realtime WebSocket, Amazon Bedrock, and object-store exports. The client-certificate and verification settings have narrower transport support. Set `ca_file` when a private or enterprise certificate authority signed the certificate an upstream presents. Those certificates are trusted in addition to the platform's, so public providers stay reachable. See [TLS and mTLS](https://docs.api7.ai/ai-gateway/deployment/tls-and-mtls.md#trust-an-upstream-behind-a-private-certificate-authority) for the support matrix, per-endpoint trust, mutual TLS, and the settings for a Redis backend. ### Tune the Downstream Connection Layer[​](#tune-the-downstream-connection-layer "Direct link to Tune the Downstream Connection Layer") The `downstream` block is the mirror of `upstream`: it controls the connections the gateway accepts from clients, or from a gateway placed in front of it. `idle_timeout_secs` closes a connection that sits idle **between** requests — the response has been fully written and no next request has started. A request in flight is never interrupted, however long the model takes, and neither is a streaming response. The same deadline also bounds how long a freshly accepted connection may take to send its first request line and headers, so keep it comfortably above the round-trip time of the slowest client you serve. It defaults to `0`, which never closes an idle connection and leaves that decision to the peer. That default is deliberate. Whatever sits in front of the gateway pools its own connections, and the node that closes first is the one that hands its peer a connection the peer still considers usable — the same failure `pool_idle_timeout_secs` avoids in the outbound direction. If you set `idle_timeout_secs`, keep it **above** the pool idle timeout of the node in front, and treat reclaiming idle connections as the reason to set it. `sse_keepalive_interval_secs` emits an SSE comment on a streaming response while the model has produced nothing. Without it, a model that is slow to its first token looks like an abandoned connection to a proxy between the client and the gateway. The comment is ignored by every conforming SSE client, and applies to every streaming endpoint. Both settings apply to the proxy listener. `idle_timeout_secs` applies to HTTP/1.1 connections. ### How the Timeouts Relate[​](#how-the-timeouts-relate "Direct link to How the Timeouts Relate") Each setting below bounds a different phase of a request. They are not interchangeable. | Setting | Where | What it bounds | | --------------------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `upstream.connect_timeout_ms` | Startup config | DNS, TCP, and TLS to the provider, before a request is sent. | | `upstream.timeout_ms` | Startup config | Default for `timeout` when neither the model nor the routing model / semantic router it was addressed through sets one. | | `upstream.stream_timeout_ms` | Startup config | Default for `stream_timeout`; `0` falls back to `upstream.timeout_ms`. | | `upstream.pool_idle_timeout_secs` | Startup config | How long an unused connection to the provider stays in the pool, after a response completes. | | `downstream.idle_timeout_secs` | Startup config | How long an accepted client connection stays open with no request on it. | | `timeout` | Model | The whole upstream call, from send to the last byte. Resolves model → routing model / semantic router → `upstream.timeout_ms`; `0` on the model disables it. | | `stream_timeout` | Model | The gap between two chunks of a streaming response — the wait for the first chunk and every gap after it, reset on each chunk. Not a cap on total stream duration. `0` or absent falls back to the group/router `stream_timeout`, then to `timeout`, then to the deployment defaults. | The two pool settings manage connection *reuse*; the model settings bound a request that is *in flight*. TCP keepalive is a network-layer liveness probe and bounds nothing at the request layer. Readers coming from a general-purpose proxy usually look for a connect / send (write) / read timeout trio. The mapping: `connect_timeout_ms` is the connect timeout; `stream_timeout` is the read timeout for streaming responses (same inter-chunk semantics), while non-streaming responses get the stricter end-to-end `timeout` instead; there is no separate send timeout because a stalled request upload is already bounded by the same end-to-end or streaming budget, and unacknowledged sends are cut earlier still at the TCP layer. A general-purpose proxy needs the trio because it has no per-request deadline concept — the gateway does, so the trio is covered with fewer knobs. For a chain of gateways, the rule at every hop is the same: a node's client-side idle timeout must stay below the next node's server-side idle timeout, with margin. A violation of that ordering is what produces intermittent transport errors against an otherwise healthy path. ## Resource Source[​](#resource-source "Direct link to Resource Source") Configure exactly one resource source. The selection affects how resources are loaded, but does not change the caller-facing [Proxy API](https://docs.api7.ai/ai-gateway/reference/proxy-api.md). ### Resources File[​](#resources-file "Direct link to Resources File") An open-source AISIX gateway reads its dynamic resources from exactly one source. Set `resources_file` to load them from a declarative resources file: config.yaml ``` resources_file: /etc/aisix/resources.yaml ``` `resources_file` and the `etcd` section are mutually exclusive — configuring both fails at startup. `resources_file` also cannot be combined with `managed.enabled: true`, because a gateway connected to AISIX Cloud receives resources from the control plane. With `resources_file` set, the gateway loads the file at startup and reloads it on `SIGHUP`. See the [Resources File Reference](https://docs.api7.ai/ai-gateway/reference/resources-file.md) for every resource kind and field. Use the [Open-Source AISIX Gateway Quickstart](https://docs.api7.ai/ai-gateway/getting-started/gateway-quickstart.md) for the workflow and the [CLI Reference](https://docs.api7.ai/ai-gateway/reference/cli.md) for validation or etcd export. [Configuration Status](https://docs.api7.ai/ai-gateway/reference/config-status.md) shows what a running gateway loaded. ### etcd Configuration Store[​](#etcd-configuration-store "Direct link to etcd Configuration Store") Set `etcd.endpoints` when configuration automation writes resources directly to etcd: config.yaml ``` etcd: endpoints: - "https://etcd.example.com:2379" prefix: "/aisix" tls: ca_cert_file: "/etc/aisix/mtls/ca.crt" client_cert_file: "/etc/aisix/mtls/client.crt" client_key_file: "/etc/aisix/mtls/client.key" domain_name: "etcd.example.com" ``` | Field | Required | Description | | ------------------------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | `etcd.endpoints` | Yes | One or more etcd endpoints. An open-source AISIX gateway requires this field when `resources_file` is not set. | | `etcd.prefix` | No | Key prefix containing AISIX resources. Defaults to `/aisix`; gateways sharing resources must use the same prefix. | | `etcd.env_id` | No | Environment scope included in the key layout when the writer uses environment-scoped keys. | | `etcd.user` | No | etcd username. | | `etcd.password_env` | No | Name of the environment variable that contains the etcd password. | | `etcd.dial_timeout_ms` | No | Accepted but not applied by this release. See the note below. | | `etcd.request_timeout_ms` | No | Accepted but not applied by this release. See the note below. | | `etcd.tls` | No | Client CA, certificate, key, and optional server-name override for an mTLS connection. The three certificate-file fields are required together. | In this release the two etcd timeout keys reach nothing. The gateway parses them and then applies no bound to any etcd call, so setting either one changes no behavior and omitting both changes no behavior: the etcd connect and the configuration read are unbounded in every case. A later release makes them take effect. Use [`aisix export`](https://docs.api7.ai/ai-gateway/reference/cli.md#export-resources-from-etcd) when migrating an existing store to a resources file. ### AISIX Cloud[​](#aisix-cloud "Direct link to AISIX Cloud") Set `managed.enabled` to `true` when the gateway receives resources from AISIX Cloud. Use the generated gateway installation snippet rather than composing the bootstrap values manually; it supplies the correct control-plane endpoints and certificate material for the selected environment. config.yaml ``` managed: enabled: true cp_base_url: "https://aisix.example.com" cp_etcd_endpoint: "aisix.example.com:443" mtls_dir: "/var/lib/aisix/mtls" dp_id_file: "/var/lib/aisix/dp_id" snapshot_cache_path: "/var/lib/aisix/config_cache.json" heartbeat_interval_secs: 15 ``` The connection certificate, private key, and CA can be supplied as one complete inline triplet (`cp_cert_pem`, `cp_key_pem`, and `cp_ca_pem`) or one complete file-path triplet (`cp_cert_file`, `cp_key_file`, and `cp_ca_file`). Do not mix inline and file forms within the same bundle. Keep credential material in environment variables or mounted secret files rather than committing it to the startup configuration. | Field | Required | Description | | --------------------------------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `managed.enabled` | Yes | Enables the AISIX Cloud connection and disables the gateway Admin API listener. | | `managed.cp_base_url` | Yes | AISIX Cloud control-plane origin used for heartbeat, telemetry, certificate rotation, and budget checks. | | `managed.cp_etcd_endpoint` | No | Explicit managed etcd endpoint as `host:port`. When omitted, AISIX derives it from `cp_base_url`. | | `managed.cp_ca_cert_file` | No | Additional CA bundle for the control-plane HTTP and etcd TLS connections, commonly needed with a private CA for an On-Premises control plane. | | `managed.mtls_dir` | No | Directory where AISIX persists the materialized mTLS bundle. Defaults to `/var/lib/aisix/mtls`. | | `managed.dp_id_file` | No | File where AISIX persists the gateway ID. Defaults to `/var/lib/aisix/dp_id`. | | `managed.snapshot_cache_path` | No | Last-known configuration cache used across control-plane outages and restarts. Defaults to `/var/lib/aisix/config_cache.json` when connected to AISIX Cloud; set an empty string to disable it. | | `managed.heartbeat_interval_secs` | No | Heartbeat interval in seconds. Defaults to `15` and is clamped to the range `5`–`300`. | The managed connection appears as `source.type: "etcd"` in [Configuration Status](https://docs.api7.ai/ai-gateway/reference/config-status.md), because the gateway consumes the control plane's projected resources through its managed store connection. For the certificate issuance and installation workflow, see [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md). For environment-variable forms of these fields, see [Environment Variables](https://docs.api7.ai/ai-gateway/reference/environment-variables.md#aisix-cloud-connection-variables). ## Select a Configuration File[​](#select-a-configuration-file "Direct link to Select a Configuration File") When running the binary directly, provide a config path with `--config` or `AISIX_CONFIG`. ``` aisix --config config.yaml ``` When running the official container image, mount the config file at `/etc/aisix/config.yaml`, or set `AISIX_CONFIG_PATH` to another path inside the container. The following example mounts a production config file and asks the entrypoint to load it: ``` docker run \ -v "$(pwd)/config.prod.yaml:/etc/aisix/config.prod.yaml:ro" \ -e AISIX_CONFIG_PATH="/etc/aisix/config.prod.yaml" \ ghcr.io/api7/aisix:1.0.0 ``` If `AISIX_CONFIG_PATH` is unset, the entrypoint uses `/etc/aisix/config.yaml`. ## Loading Order[​](#loading-order "Direct link to Loading Order") AISIX loads startup configuration in the following order: 1. Built-in default values. 2. File contents from the path selected by `--config` or `AISIX_CONFIG`. 3. Environment-variable overrides with the `AISIX_` prefix. Environment-variable overrides apply only to startup configuration fields. For override syntax and AISIX Cloud connection variables, see [Environment Variables](https://docs.api7.ai/ai-gateway/reference/environment-variables.md). --- # Environment Variables AISIX AI Gateway uses environment variables to select startup configuration files, override startup configuration fields, and provide deployment-specific values such as AISIX gateway certificate material. Most runtime gateway resources are not configured directly through environment variables. For an open-source AISIX gateway, declare models, caller API keys, provider keys, guardrails, cache policies, and observability exporters in a [`resources.yaml` file](https://docs.api7.ai/ai-gateway/reference/resources-file.md). The file supports [environment interpolation](https://docs.api7.ai/ai-gateway/reference/resources-file.md#environment-interpolation) for values such as `${OPENAI_API_KEY}`. For an AISIX gateway connected to AISIX Cloud, the control plane supplies these resources. ## Reserved Environment Variables[​](#reserved-environment-variables "Direct link to Reserved Environment Variables") AISIX reserves the following environment variables: | Variable | Description | | ----------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | `AISIX_CONFIG` | Config file path used by the AISIX binary. Equivalent to passing `--config`. | | `AISIX_CONFIG_PATH` | Config file path used by the official container entrypoint. Defaults to `/etc/aisix/config.yaml`. | | `RUST_LOG` | Process logging directive. When unset, AISIX uses `observability.log_level`. | | `AISIX_DP_BUDGET_STALE_MAX_SECONDS` | Maximum number of seconds a gateway connected to AISIX Cloud can reuse a stale budget decision after the normal cache TTL. Defaults to `600`. | To use these variables, assign values before starting AISIX. Use `AISIX_CONFIG` when you run the binary directly: ``` export AISIX_CONFIG="/etc/aisix/config.yaml" aisix ``` Use `AISIX_CONFIG_PATH` when you use the official container entrypoint: ``` docker run \ -v "$(pwd)/config.prod.yaml:/etc/aisix/config.prod.yaml:ro" \ -e AISIX_CONFIG_PATH="/etc/aisix/config.prod.yaml" \ ghcr.io/api7/aisix:1.0.0 ``` The container entrypoint clears `AISIX_CONFIG_PATH` before starting the binary because it is an entrypoint variable, not a startup config field. ## Startup Configuration Overrides[​](#startup-configuration-overrides "Direct link to Startup Configuration Overrides") After AISIX loads the config file, it applies environment-variable overrides with the `AISIX_` prefix. Use a single underscore after the prefix and double underscores between nested fields. The following example overrides the proxy listener address: ``` export AISIX_PROXY__ADDR="0.0.0.0:3000" ``` Common override variables include: | Variable | Overrides | | --------------------------------------- | --------------------------------- | | `AISIX_PROXY__ADDR` | `proxy.addr` | | `AISIX_PROXY__THREAD_PER_CORE` | `proxy.thread_per_core` | | `AISIX_PROXY__WORKERS` | `proxy.workers` | | `AISIX_ETCD__ENDPOINTS` | `etcd.endpoints` | | `AISIX_ETCD__PREFIX` | `etcd.prefix` | | `AISIX_OBSERVABILITY__LOG_LEVEL` | `observability.log_level` | | `AISIX_CACHE__REDIS__MODE` | `cache.redis.mode` | | `AISIX_CACHE__REDIS__URL` | `cache.redis.url` | | `AISIX_CACHE__REDIS__MASTER_NAME` | `cache.redis.master_name` | | `AISIX_CACHE__REDIS__USERNAME` | `cache.redis.username` | | `AISIX_CACHE__REDIS__PASSWORD` | `cache.redis.password` | | `AISIX_CACHE__REDIS__DATABASE` | `cache.redis.database` | | `AISIX_RATELIMIT__BACKEND` | `ratelimit.backend` | | `AISIX_RATELIMIT__REDIS__MODE` | `ratelimit.redis.mode` | | `AISIX_RATELIMIT__REDIS__URL` | `ratelimit.redis.url` | | `AISIX_RATELIMIT__REDIS__MASTER_NAME` | `ratelimit.redis.master_name` | | `AISIX_RATELIMIT__REDIS__USERNAME` | `ratelimit.redis.username` | | `AISIX_RATELIMIT__REDIS__PASSWORD` | `ratelimit.redis.password` | | `AISIX_RATELIMIT__REDIS__DATABASE` | `ratelimit.redis.database` | | `AISIX_RATELIMIT__CONCURRENCY_TTL_SECS` | `ratelimit.concurrency_ttl_secs` | | `AISIX_BEDROCK_ENDPOINT_URL` | Top-level `bedrock_endpoint_url`. | `etcd.endpoints` accepts a comma-separated list in an environment variable. For Redis Cluster and Sentinel node lists, configure `cache.redis.nodes`, `cache.redis.sentinels`, `ratelimit.redis.nodes`, or `ratelimit.redis.sentinels` in the startup configuration file. For configuration file fields, see the [Startup Configuration Reference](https://docs.api7.ai/ai-gateway/reference/configuration-files.md). ## AISIX Cloud Connection Variables[​](#aisix-cloud-connection-variables "Direct link to AISIX Cloud Connection Variables") AISIX gateways use the same `AISIX_` override mechanism for `managed.*` startup settings. | Variable | Description | | ---------------------------------------- | -------------------------------------------------------------------------------------------------------- | | `AISIX_MANAGED__ENABLED` | Connects the gateway to AISIX Cloud when set to `true`. | | `AISIX_MANAGED__CP_BASE_URL` | AISIX Cloud control-plane origin used for heartbeat, telemetry, certificate rotation, and budget checks. | | `AISIX_MANAGED__CP_ETCD_ENDPOINT` | Control-plane etcd endpoint used by the gateway at startup. | | `AISIX_MANAGED__CP_CA_CERT_FILE` | Optional CA bundle file used to trust control-plane and etcd TLS connections. | | `AISIX_MANAGED__CP_CERT_PEM` | Inline client certificate PEM used for mTLS with the AISIX Cloud control plane. | | `AISIX_MANAGED__CP_KEY_PEM` | Inline private key PEM paired with the client certificate. | | `AISIX_MANAGED__CP_CA_PEM` | Inline CA certificate PEM used as the trust anchor. | | `AISIX_MANAGED__CP_CERT_FILE` | File path for the client certificate PEM. | | `AISIX_MANAGED__CP_KEY_FILE` | File path for the private key PEM. | | `AISIX_MANAGED__CP_CA_FILE` | File path for the CA certificate PEM. | | `AISIX_MANAGED__MTLS_DIR` | Directory where the gateway persists the materialized mTLS bundle. | | `AISIX_MANAGED__DP_ID_FILE` | File where the gateway persists its AISIX gateway ID. | | `AISIX_MANAGED__SNAPSHOT_CACHE_PATH` | File path for the on-disk snapshot cache used during control-plane outages. | | `AISIX_MANAGED__HEARTBEAT_INTERVAL_SECS` | AISIX gateway heartbeat interval in seconds. Defaults to `15`; values are clamped between `5` and `300`. | Use either the inline PEM variables or the file-path variables for the certificate, key, and CA bundle. Do not mix inline and file variants for the same bundle. For AISIX Cloud setup, see [Connect an AISIX Gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md). --- # Headers and Error Codes AISIX returns responses through several caller-facing API formats. A failed chat-completions request, an Anthropic-style messages request, and a passthrough route request do not all use the same error envelope. This reference helps interpret response headers, retry hints, status codes, and error fields. Start with the request path, then read the matching error format before deciding whether the failure is caller-side, gateway-side, or upstream-provider-side. ## Error Response Formats[​](#error-response-formats "Direct link to Error Response Formats") Use the request path to identify the error envelope that applies. | Response came from | Error envelope | Read | | ---------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- | | OpenAI-compatible proxy routes such as `/v1/chat/completions`, `/v1/completions`, `/v1/embeddings`, `/v1/responses`, audio, images, and rerank | `{"error": {...}}` | [OpenAI-Style Proxy Errors](#openai-style-proxy-errors) | | Anthropic-style proxy routes such as `/v1/messages` and `/v1/messages/count_tokens` | `{"type":"error","error": {...}}` | [Anthropic-Style Proxy Errors](#anthropic-style-proxy-errors) | | MCP routes under `/mcp` and `/mcp/{server}` | JSON-RPC error envelope | [MCP Errors](#mcp-errors) | | A2A calls under `/a2a/{agent}` | Upstream JSON-RPC response; HTTP errors before forwarding; JSON-RPC error for an upstream dispatch failure | [A2A Errors](#a2a-errors) | | A2A agent cards under `/a2a/{agent}/.well-known/agent-card.json` | Agent-card JSON on success; HTTP error response on gateway or upstream failure | [A2A Errors](#a2a-errors) | | Configured passthrough routes | Forwarded upstream status and body, or `{"error": {...}}` for AISIX-generated failures | [Passthrough Errors](#passthrough-errors) | ## Proxy Response Headers[​](#proxy-response-headers "Direct link to Proxy Response Headers") Operational headers vary by endpoint. Do not treat every header as universal across every `/v1/*` route. | Header | When to use it | | -------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `x-aisix-call-id` | Appears on chat-completions responses. Use it to correlate one gateway call. | | `x-aisix-request-id` | Appears on every proxy response. Use it to correlate the response with any access logs and usage events the request produces. Some MCP failures occur before usage accounting; see [Usage Events](https://docs.api7.ai/ai-gateway/mcp-gateway/observability.md#usage-events). Accepted on the request by default: send your own ID and AISIX uses that value throughout instead of generating one. Which request headers an ID is read from is configurable, so a deployment can extend or disable this. See [Reuse Your Own Request ID](https://docs.api7.ai/ai-gateway/observability/metrics-and-logs.md#reuse-your-own-request-id). | | `x-aisix-served-by` | Appears on successful chat-completions routing responses. Use it to identify the target model that served the request. | | `x-aisix-cache` | Appears on chat requests covered by a cache policy: `hit`, `miss`, or `bypass`. Only `Cache-Control: no-cache` produces `bypass`; `no-store` still performs a lookup and therefore reports `hit` or `miss`. Use it to check whether the gateway served the response from cache. | | `x-aisix-cache-layer` | Appears on cache hits: `exact` for an identical-request match, `semantic` for an embedding-similarity match. Use it to attribute hits between the two matching layers. | | `x-aisix-cache-similarity` | Appears on semantic cache hits. The cosine similarity of the matched entry, `0`–`1`. Use it to calibrate a policy's similarity threshold. | | `x-ratelimit-*` | The per-dimension family, with names ending in `-requests`, `-tokens`, and `-concurrent`. Appears on successful chat-completions responses when the caller API key has rate limits configured. Use it to inspect request, token, and concurrent limit state where applicable. | | `X-RateLimit-Limit`<br />`X-RateLimit-Remaining`<br />`X-RateLimit-Reset`<br />`X-RateLimit-Scope` | Appears only on a `429` produced by the AISIX rate limiter itself. Use it to learn which limit was hit and when to retry. See [Rate-Limit Rejection Headers](#rate-limit-rejection-headers). | | `Retry-After` | Appears on rate-limit, AISIX Cloud budget, and all-candidates-unavailable rejections when the gateway has a retry hint. Also appears on upstream 429 responses when AISIX can parse the upstream retry hint. Use it to tell callers when to retry. | ### Rate-Limit Rejection Headers[​](#rate-limit-rejection-headers "Direct link to Rate-Limit Rejection Headers") When the AISIX rate limiter itself refuses a request, the `429` describes the limit that refused it, in both the OpenAI-style and Anthropic-style error envelopes. | Header | Value | | ----------------------- | ------------------------------------------------------------------------------------------------------------- | | `X-RateLimit-Limit` | The configured cap of the limit that refused the request. | | `X-RateLimit-Remaining` | Headroom left under that cap. `0` on a rejection. | | `X-RateLimit-Reset` | Seconds until that limit accepts requests again. | | `Retry-After` | The same number of seconds, so a client can back off on whichever header it already supports. | | `X-RateLimit-Scope` | An AISIX extension naming the limit that refused: `rps`, `rpm`, `rph`, `rpd`, `tpm`, `tpd`, or `concurrency`. | `X-RateLimit-Reset` is a **duration in seconds**, not a timestamp, so a client does not need a clock synchronized with the gateway. For a windowed limit it counts down toward the end of the current fixed window. It is never `0`; the smallest value is `1`. A caller API key, a model, and a rate-limit policy can each carry several limits at once. Exactly one of them refuses a given request — the first one AISIX finds exhausted — and the headers describe that one. `X-RateLimit-Scope` is what makes the numbers unambiguous, because the units differ: `x-ratelimit-limit: 1000` is 1000 requests under `rpm`, 1000 tokens under `tpm`, and 1000 in-flight requests under `concurrency`. Two cases do not follow that reading: * **A routing model or semantic router.** Dispatch skips each target that is over its own limit and tries the next one; when none is left, the headers describe the **last target that refused**, which is a limit configured on that target rather than on the alias the request addressed. Read them as "this is why the request could not be placed", not as the caller's own quota. * **An ensemble model.** A panel member or judge that exceeds its own limit fails the request with a `429` that carries none of these headers, and no `Retry-After` either. A concurrency rejection is the one case with no fixed window. A slot frees when some in-flight request finishes, which the gateway cannot predict, so `X-RateLimit-Reset` and `Retry-After` both report a fixed `60`. A raw rejection looks like this: ``` HTTP/1.1 429 Too Many Requests content-type: application/json retry-after: 43 x-ratelimit-limit: 100 x-ratelimit-remaining: 0 x-ratelimit-reset: 43 x-ratelimit-scope: rpm x-aisix-request-id: 018f3c1f-... {"error":{"message":"request limit exceeded (requests)","type":"rate_limit_exceeded"}} ``` HTTP header names are case-insensitive, and AISIX writes them in lowercase on the wire. A client reading `X-RateLimit-Limit` and a client reading `x-ratelimit-limit` both match. The `X-RateLimit-*` headers appear **only** when AISIX itself refused the request. They are absent from: * **Successful responses.** A `200` carries the per-dimension `x-ratelimit-*` family instead: `x-ratelimit-limit-requests`, `x-ratelimit-limit-tokens`, `x-ratelimit-limit-concurrent`, and their `remaining` and `reset` counterparts. That family reports the caller API key's own `rpm`, `tpm`, and `concurrency` state only — not a model limit, not a rate-limit policy, and not the per-second, per-hour, or per-day windows. * **Upstream `429` responses.** AISIX forwards the provider's `Retry-After` when it can parse one, but the provider's quota state is not something AISIX knows, so it does not report one of its own. * **AISIX Cloud budget rejections.** A budget caps spend in currency, not requests or tokens. It shares the `429` status and carries structured budget fields inside the error body, plus `Retry-After` when the control plane supplied a reset time. Their presence is therefore the signal that the AISIX limiter itself rejected the request. `Retry-After` alone is not that signal: as the two cases above show, it appears on other rejections too. ## Proxy Status Codes[​](#proxy-status-codes "Direct link to Proxy Status Codes") Use the error type first when the envelope includes one. The status code gives the broad category, and the error type usually identifies the more precise gateway condition. | Status | Meaning | | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `400` | The request is invalid. | | `401` | Caller authentication is missing or invalid. | | `403` | The caller is authenticated but is not permitted by the configured access controls. | | `404` | The requested resource was not found. Gateway-generated examples include an unknown model alias, MCP server, or A2A agent. An upstream service can also return `404`; AISIX preserves upstream `4xx` statuses. | | `413` | The request body exceeds the configured proxy request body limit. | | `422` | Content was blocked by a guardrail, either by a policy match or because a fail-closed guardrail could not evaluate it. See [Guardrail Refusals](#guardrail-refusals). | | `429` | The request hit a rate limit or AISIX Cloud budget rejection. | | `501` | The resolved provider adapter does not implement the requested endpoint. | | `502` | The upstream provider returned a server-side failure or the provider adapter mapped an upstream failure into the proxy error format. | | `503` | An authentication dependency or provider adapter is unavailable, or every routing candidate was filtered out by runtime status. | | `504` | The upstream request timed out. | ## OpenAI-Style Proxy Errors[​](#openai-style-proxy-errors "Direct link to OpenAI-Style Proxy Errors") AISIX OpenAI-compatible proxy errors use this envelope: ``` { "error": { "message": "...", "type": "invalid_request_error" } } ``` The `param` and `code` fields are omitted when AISIX has no value for them. AISIX Cloud budget denials include structured budget fields inside the `error` object, such as `scope`, `limit_usd`, `spent_usd`, `period`, and `retry_after_seconds`. Common AISIX `error.type` values are: | Error type | Typical status | Meaning | | ---------------------------- | ----------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `invalid_api_key` | `401` | Caller authentication is missing or invalid. This type also covers an invalid, expired, or unmapped JWT; inspect `error.code` for the specific condition. | | `permission_denied` | `403` | The caller API key cannot use the requested model, the caller's client IP is outside the model's `allowed_cidrs`, or a verified JWT does not satisfy required scopes or claims. | | `model_not_found` | `404` | The requested model alias is not configured. | | `invalid_request_error` | `400` or `413` | The request body or endpoint usage is invalid. Oversized OpenAI-style requests return this error type with status `413`. | | `provider_unavailable` | `503` | The selected upstream provider adapter cannot complete the request. | | `all_candidates_unavailable` | `503` | Every routing candidate was filtered out or unavailable. | | `api_error` | `503` | AISIX could not complete an internal dependency operation, such as fetching signing keys for JWT authentication. | | `content_filter` | `422` | The request or response was blocked by a guardrail. `error.code` is `guardrail_unavailable` when the guardrail could not evaluate the content, and absent on a policy match. See [Guardrail Refusals](#guardrail-refusals). | | `billing_error` | `429` | AISIX Cloud budget enforcement rejected the request because a blocking budget was exceeded or its outage policy denied the request. | | `rate_limit_exceeded` | `429` | The request exceeded a configured rate limit. | | `not_implemented` | `501` | The resolved provider adapter does not implement the requested endpoint. | | `timeout` | `504` | The upstream request timed out. | | `upstream_error` | Varies, often `502` for upstream server-side failures | The upstream provider returned an error that AISIX rendered through the proxy error format. | ### Authentication Error Codes[​](#authentication-error-codes "Direct link to Authentication Error Codes") OpenAI-style proxy errors can include these stable `error.code` values for caller authentication. Anthropic-style proxy errors omit `code`; use their HTTP status and status-mapped `error.type` instead. These are not the only `error.code` values. A guardrail refusal carries `guardrail_unavailable`, a source-IP rejection carries `ip_restricted`, and a budget denial carries `budget_exceeded`. See [Guardrail Refusals](#guardrail-refusals). | `error.code` | Status | Meaning | | ----------------------- | ------ | ----------------------------------------------------------------------------------------------------------------------- | | `api_key_expired` | `401` | The caller API key's expiration deadline has passed. | | `api_key_disabled` | `401` | The caller API key is administratively disabled. | | `jwt_invalid` | `401` | The JWT is malformed or fails issuer, signature, signing algorithm, audience, required-claim, or not-before validation. | | `jwt_expired` | `401` | The JWT's `exp` deadline has passed. | | `jwt_claims_rejected` | `403` | The JWT is valid but does not satisfy the provider's required scopes or bound claims. | | `jwt_identity_unmapped` | `401` | The configured identity claim is missing or does not map to a caller API key bound to that OIDC provider. | | `jwks_unavailable` | `503` | AISIX could not resolve or fetch the provider's signing keys. Check the OIDC discovery or JWKS endpoint, then retry. | For JWT trust-provider configuration and rejection behavior, see [JWT Authentication](https://docs.api7.ai/ai-gateway/traffic-controls/jwt-authentication.md#rejection-reasons). ### Upstream Provider Errors[​](#upstream-provider-errors "Direct link to Upstream Provider Errors") OpenAI-style routes render upstream provider failures through the same error envelope, but AISIX does not always return the upstream response unchanged. Upstream `4xx` responses keep the client-visible HTTP class. Native OpenAI upstream errors can keep their OpenAI-style fields. Cross-provider upstream errors use `upstream_error` and can include a more specific `error.code`, such as a rate-limit, permission, or model-not-found code. Upstream `5xx` responses generally return `502`. AISIX does not expose upstream `5xx` response bodies because they can contain provider account, infrastructure, or private diagnostic detail. ## Anthropic-Style Proxy Errors[​](#anthropic-style-proxy-errors "Direct link to Anthropic-Style Proxy Errors") `POST /v1/messages` and `POST /v1/messages/count_tokens` use the Anthropic-style error envelope: ``` { "type": "error", "error": { "type": "invalid_request_error", "message": "..." } } ``` The Anthropic envelope omits the `param` and `code` fields that the OpenAI envelope can carry. A guardrail refusal on these routes therefore reaches the caller without the `guardrail_unavailable` code an OpenAI-style route would carry; the failure tag is still named in the message. See [Guardrail Refusals](#guardrail-refusals). The nested `error.type` follows Anthropic SDK-compatible status mappings: | Status | Anthropic `error.type` | | ------------------ | ----------------------- | | `400` or `422` | `invalid_request_error` | | `401` | `authentication_error` | | `403` | `permission_error` | | `404` | `not_found_error` | | `408` | `timeout_error` | | `413` | `request_too_large` | | `429` | `rate_limit_error` | | `503` | `overloaded_error` | | Other status codes | `api_error` | AISIX keeps the `408` mapping for Anthropic SDK compatibility. Gateway-originated timeouts usually surface through provider-error handling rather than as a native `408` response. See [Anthropic-Style Messages API](https://docs.api7.ai/ai-gateway/endpoints/anthropic-messages.md#handle-errors) for examples. ## MCP Errors[​](#mcp-errors "Direct link to MCP Errors") `ANY /mcp` and `ANY /mcp/{server}` use MCP Streamable HTTP and JSON-RPC response shapes. Authentication and request-size failures can still use HTTP status codes such as `401` or `413` before the MCP handler runs. When a guardrail blocks a tool call or tool result, AISIX returns HTTP 200 with a tool result marked `isError`. MCP reserves JSON-RPC protocol errors for a request that was not valid; a policy rejection is a tool-execution error, so the calling agent reads it as tool output and can adapt: ``` { "jsonrpc": "2.0", "id": 42, "result": { "content": [{ "type": "text", "text": "tool call blocked by content policy (guardrail 'block-secrets')" }], "isError": true } } ``` This differs from OpenAI-compatible routes, where guardrail blocks use HTTP 422 with an OpenAI-style error envelope. For the message wording and the failure-tag vocabulary, see [Guardrail Refusals](#guardrail-refusals). A rejected or unknown tool is a protocol error instead — HTTP 200 with `error.code` `-32602` — because the request itself named something the caller may not call. ## A2A Errors[​](#a2a-errors "Direct link to A2A Errors") `POST /a2a/{agent}` returns a successful upstream JSON-RPC response unchanged. After the gateway has identified the agent and entered A2A dispatch, an upstream connection failure or non-success response returns HTTP `502` with a JSON-RPC error envelope: ``` { "jsonrpc": "2.0", "id": 42, "error": { "code": -32000, "message": "..." } } ``` A guardrail refusal uses the same JSON-RPC error envelope with HTTP `422`; its `code` is the JSON-RPC `-32000`, not a gateway error code. See [Guardrail Refusals](#guardrail-refusals). Some failures occur before AISIX forwards the A2A request and therefore are not JSON-RPC errors. Missing or invalid caller authentication returns `401`. A caller key without access to the registered agent returns `403`, and an unknown or disabled agent returns `404`. Rate-limit or AISIX Cloud budget rejections return `429`. Agent-card discovery uses ordinary HTTP status codes rather than a JSON-RPC envelope. Missing or invalid caller authentication returns `401`, and an agent access denial returns `403`. An unknown or disabled agent returns `404`. If the gateway cannot fetch a usable card from the upstream agent, it returns `502`. If the gateway cannot determine its own public address from the request, it returns `500` rather than a card that still advertises the upstream agent's address. A streaming method (`message/stream`, `tasks/resubscribe`) follows the same rules only until the stream opens. Once the response has started, the status line is already sent, so a later upstream failure cannot become a `502`: the gateway relays a JSON-RPC error envelope as the final event of the stream instead, and records the failure on the usage event. An upstream that answers a streaming call with a plain JSON-RPC response rather than a stream has that response relayed as a single event. ## Passthrough Errors[​](#passthrough-errors "Direct link to Passthrough Errors") A matched [passthrough route](https://docs.api7.ai/ai-gateway/endpoints/provider-passthrough.md) forwards the upstream status code and body unchanged when AISIX receives an upstream HTTP response and no gateway policy replaces it. Failures generated by AISIX use the [OpenAI-style proxy error](#openai-style-proxy-errors) envelope. These failures include caller authentication rejection, a guardrail block, and an upstream transport, timeout, or response-decoding failure. They also include a missing `allowed_routes` grant, and a source outside the route's `source_cidrs` (`403`, with `error.code: ip_restricted` for the source rejection). A guardrail block is a `422`, or a terminal SSE `error` event on a streamed response. See [Guardrail Refusals](#guardrail-refusals). A `/passthrough/*` path that no explicit route claims follows the ordinary empty-body `404` path. Create a passthrough route to claim the path. ## Guardrail Refusals[​](#guardrail-refusals "Direct link to Guardrail Refusals") A guardrail stops traffic for one of two reasons, and the difference matters to a caller. A **policy match** means the guardrail read the content and refused it. A **fail-closed availability failure** means the guardrail could not evaluate the content at all, and its configuration refuses what it cannot check. Both are guardrail refusals. Both use the requested endpoint's own error convention, and both are recorded as guardrail-blocked. Two separate things identify the second kind, and they are easy to confuse: * `error.code`, on the surfaces that carry one at all, is a single constant rather than a per-failure value. On the OpenAI-style envelope it is `guardrail_unavailable`, present only when the guardrail could not evaluate and absent on a policy match. Two surfaces do not follow that rule: a bridged `/v1/responses` stream always sends `content_filter` and `/v1/realtime` always sends `content_filtered`, on a policy match and an availability failure alike. * The **failure tag** — `lakera_timeout`, `unscannable_body`, and so on — is a bounded, code-owned vocabulary that appears **only inside `error.message`**, in parentheses. It is not an `error.code` value and no envelope field carries it. The message is built to a fixed shape: ``` <side> rejected: guardrail '<name>' could not evaluate it (<tag>) <side> rejected: a guardrail could not evaluate it (<tag>) ``` The second form is used when no single guardrail can be named. That is the case for every refusal the gateway raises on a chain's behalf rather than on one member's verdict. `<side>` is `request` or `response`, and `tool call` or `tool result` on `/mcp`. A policy match uses the same shape without the tag: ``` <side> blocked by content policy (guardrail '<name>') <side> blocked by content policy ``` On the OpenAI-style routes, the presence of `error.code` is itself the signal, so branch on that. On a bridged `/v1/responses` stream and on `/v1/realtime` the code is a constant. It cannot separate a policy match from an availability failure, so the tag in the message is the only thing that can. Treat the tag vocabulary as a set of known values rather than as a stable position in a string. ### Refusal Envelopes by Surface[​](#refusal-envelopes-by-surface "Direct link to Refusal Envelopes by Surface") Each surface keeps its own protocol's error shape, so `guardrail_unavailable` does not reach every caller. Anthropic-shaped routes have no `code` field at all, `/mcp` answers in-band as a tool result, and `/a2a` answers in JSON-RPC. | Surface | HTTP | Body | `code` | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------- | | OpenAI-compatible routes, non-streaming — including `/v1/chat/completions`, `/v1/embeddings`, `/v1/audio/transcriptions`, `/v1/audio/translations`, and the Batch and Fine-tuning routes | `422` | `{"error":{"message":...,"type":"content_filter","code":"guardrail_unavailable"}}` | `guardrail_unavailable` | | `/v1/chat/completions`, streaming | `200`, then a terminal SSE `error` event | `{"error":{"message":...,"type":"content_filter"}}` | none | | `/v1/responses`, where the upstream serves the Responses API natively | `422` | Same as the non-streaming OpenAI envelope | `guardrail_unavailable` | | `/v1/responses`, where AISIX bridges a provider that does not serve the Responses API natively | `200`, then a terminal SSE `error` event | Flat, matching the Responses error event: `{"type":"error","code":"content_filter","message":...,"param":null,"sequence_number":N}` | `content_filter` | | `/v1/messages` and `/v1/messages/count_tokens`, non-streaming | `422` | `{"type":"error","error":{"type":"invalid_request_error","message":...}}` | The Anthropic envelope has no `code` field | | `/v1/messages`, streaming | `200`, then a terminal SSE `error` event | `{"type":"error","error":{"type":"invalid_request_error","message":...}}` | The Anthropic envelope has no `code` field | | Passthrough routes, non-streaming | `422` | Same as the non-streaming OpenAI envelope | `guardrail_unavailable` | | Passthrough routes, streaming | `200`, then a terminal SSE `error` event | `{"error":{"type":"content_filter","message":...}}` | none | | `/mcp` and `/mcp/{server}` | `200` | A tool result flagged `isError`, with the refusal message as its text. See [MCP Errors](#mcp-errors). | Not applicable | | `/a2a/{agent}` | `422` | `{"jsonrpc":"2.0","id":...,"error":{"code":-32000,"message":...}}` | `code` is the JSON-RPC `-32000`, not a gateway code | | `/v1/realtime` | No status — the session is already open | A WebSocket text frame, `{"type":"error","error":{"type":"invalid_request_error","code":"content_filtered","message":...}}`, followed by a close frame with code `1011` and reason `content policy` | `content_filtered` | Two consequences worth planning for. An SDK on an Anthropic-shaped route cannot branch on `error.code`, because the envelope has no such field — read the HTTP status and the message instead. And on most streaming surfaces the refusal arrives on a response that already returned `200`, so a client that only inspects status codes sees a successful, truncated stream. ### Failure Tags Raised by the Gateway[​](#failure-tags-raised-by-the-gateway "Direct link to Failure Tags Raised by the Gateway") These three come from the proxy itself, on a chain's behalf, where no guardrail produced a verdict to carry a tag. They never name a guardrail. | Tag | Side | Where it occurs | | ------------------------ | -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `unscannable_body` | Request and response | Some content could not be evaluated as delivered. On the request side, this can be a `/v1/messages` or `/v1/messages/count_tokens` body the scanner cannot parse. On the response side, it can be a held-back stream left with nothing scannable, a `/mcp` tool result that does not parse, or a non-streaming audio transcription or translation response that AISIX cannot decode without replacing invalid byte sequences. See [Frames AISIX Cannot Scan](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md#frames-aisix-cannot-scan) and [Transcripts AISIX Could Not Decode](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/behavior.md#transcripts-aisix-could-not-decode). | | `output_buffer_exceeded` | Response only | A streamed response outgrew the hold-back cap before it could be scanned. Occurs on `/v1/chat/completions`, `/v1/messages`, `/v1/responses`, and passthrough routes. Only `/v1/chat/completions` and passthrough routes honor `on_buffer_exceeded` here; `/v1/messages` and `/v1/responses` always fail closed on overflow. | | `mask_writeback_failed` | Request and response | A mask verdict could not be spliced back into the body, so the content is refused rather than forwarded unmasked. Occurs on `/mcp` only, on both `tools/call` arguments and tool results. | For an unreadable Anthropic request, MCP tool result, or plain-text audio transcript, `unscannable_body` is refused only when at least one guardrail in scope reads that side of the exchange and fails closed on it. If all readers on that side fail open, AISIX relays the content and records `unscannable_body` in `guardrail_bypassed_reason`; for audio, AISIX still scans the best-effort decoded text, so only the replaced byte sequences escape inspection. A chain that does not read that side neither refuses nor records a bypass. A held-back stream left with nothing scannable is different: once an enforcing output guardrail has withheld the stream, AISIX refuses it regardless of `fail_open`. `output_buffer_exceeded` is usually **not** a status-code refusal. On most routes the stream has already returned `200`, so the refusal arrives as a terminal SSE `error` event and the held frames are dropped. The one exception is `/v1/responses` against a provider that serves the Responses API natively. Nothing has been delivered there, so the request is refused with a `422`. When AISIX bridges a provider that does not serve that API, the same overflow arrives as a terminal SSE `error` event instead. ### Failure Tags Raised by a Guardrail Kind[​](#failure-tags-raised-by-a-guardrail-kind "Direct link to Failure Tags Raised by a Guardrail Kind") Each remote guardrail kind classifies its own backend failures into a bounded set. The same tag is used whichever way the row is configured. It is the `Bypass` reason when the guardrail fails open, and the refusal tag when it fails closed, so one outage reads the same either way. | Guardrail kind | Failure tags | | ----------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | | `lakera` | `lakera_timeout`, `lakera_throttled`, `lakera_5xx`, `lakera_too_large`, `lakera_config_error` | | `presidio` | `presidio_timeout`, `presidio_throttled`, `presidio_5xx`, `presidio_too_large`, `presidio_config_error` | | `openai_moderation` | `openai_moderation_timeout`, `openai_moderation_throttled`, `openai_moderation_5xx`, `openai_moderation_too_large`, `openai_moderation_config_error` | | `azure_content_safety` and `azure_content_safety_text_moderation` | `azure_cs_timeout`, `azure_cs_throttled`, `azure_cs_5xx`, `azure_cs_config_error` | | `bedrock` | `bedrock_timeout`, `bedrock_throttled`, `bedrock_too_large`, `bedrock_5xx` | | `aliyun_text_moderation` and `aliyun_ai_guardrail` | `aliyun_timeout`, `aliyun_throttled`, `aliyun_5xx`, `aliyun_bad_response`, `aliyun_config_error` | | `custom` | `custom_timeout`, `custom_script_error`, `custom_engine_error`, `custom_no_verdict`, `custom_bad_verdict`, `custom_unknown_action` | | `semantic` | `semantic_embed_unresolved`, `semantic_embed_timeout`, `semantic_embed_upstream` | The two Azure kinds share one vocabulary because they call the same service, and so do the two Alibaba Cloud kinds. Read a tag as naming the backend, not the guardrail row — use the guardrail name in the message for that. The tags mean what their names suggest. `_timeout` is the configured wall clock elapsing, and `_throttled` a `429` from the backend. `_5xx` is a server-side or transport failure. `_config_error` is a non-`429` `4xx`, such as a rejected credential or a wrong endpoint. `_too_large` is a payload the backend refused for its size. `aliyun_bad_response` is a `2xx` whose body was not the documented shape. It is kept separate from `aliyun_5xx` because the two want different fixes. The `custom_*` set distinguishes a script that timed out, threw, or could not be started. It also covers a script that returned nothing, returned a shape that is not a verdict, or returned an action outside the vocabulary. The built-in [`keyword`](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/keyword.md) and [`pii`](https://docs.api7.ai/ai-gateway/traffic-controls/guardrails/pii.md) guardrails run inside the gateway and call no backend, so they have no failure tags. A block from either carries no `error.code` at all. ### Where Else You Meet These Tags[​](#where-else-you-meet-these-tags "Direct link to Where Else You Meet These Tags") A guardrail kind's own tags appear on four operator-facing surfaces, so a tag read from a caller's error message can be searched for directly: * **Usage record audit hits.** A fail-closed refusal is recorded with `action` `blocked_unavailable` and the tag in `error_type`. An ordinary policy match records `action` `blocked` and leaves `error_type` empty. A monitor-mode observation that could not evaluate records the tag on its `would_block` summary. * **The `aisix_guardrail_latency_seconds` Prometheus histogram.** Its `error_type` label carries the tag for a fail-closed `blocked` result, a `bypassed` result, and a failed monitor-mode `would_block`. Ordinary policy matches use `none`. The `result` label stays `blocked` for a fail-closed refusal rather than splitting into a value of its own, so an existing `result="blocked"` alert keeps counting it. See [Metrics Reference](https://docs.api7.ai/ai-gateway/reference/metrics.md#measure-guardrail-latency-and-outcomes). * **The `aisix_guardrail_bypasses_total` Prometheus counter.** A fail-open execution increments the counter under the same tag in its `reason` label. It counts bypass events rather than requests and has no guardrail, kind, or phase labels. * **`guardrail_bypassed_reason` on the usage event.** When a guardrail fails **open**, nothing is refused and no message reaches the caller, but the same tag is recorded here. Every proxied route's usage event carries this field; retroactive batch-attribution events are the exception because no guardrail chain was resolved for them. The three tags the gateway raises itself are the exception. They are not guardrail executions, so they do not appear in an audit hit or in the guardrail execution histogram's `error_type`. A fail-closed refusal records `guardrail_blocked` and names the tag in the caller's message and gateway log. When an `unscannable_body` is released by the applicable fail-open policy, the usage event records that tag in `guardrail_bypassed_reason`. The `aisix_guardrail_bypasses_total` counter still increments, but no latency-histogram observation is emitted because no guardrail executed. --- # Metrics Reference AISIX exposes operational metrics in Prometheus text format. A Prometheus server or another compatible collector can scrape these metrics for dashboards, alerts, and PromQL queries. Prometheus metrics are enabled by default on a dedicated listener. The default startup configuration is: config.yaml ``` observability: metrics: prometheus: enabled: true addr: 0.0.0.0:9090 path: /metrics ``` Prometheus scrapes `GET /metrics` from this listener. The metrics endpoint is unauthenticated by design. Keep its listener private to your monitoring network. AISIX publishes configuration status whenever the endpoint is scraped. Other metric families are registered when AISIX first records the corresponding activity, so traffic metrics might not appear immediately after startup. Send a request through the proxy, then scrape again for the corresponding series to appear. ## Metrics Catalog Search metric names, descriptions, labels, and values, or filter the catalog by family and type. * Metrics 62 * Families 10 Query behavior ### Metric Types counter A cumulative value that increases until the process restarts, such as a request or token total. Use `rate()` to calculate how quickly it changes. gauge A current value that can increase or decrease, such as active requests or remaining quota. histogram Observations counted in configurable buckets. The `_bucket`, `_sum`, and `_count` series can be aggregated before calculating a percentile. summary Observations with quantiles calculated by each gateway instance. Summaries also expose `_sum` and `_count`, but their quantiles cannot be aggregated across instances. ### Request Metrics Track request outcomes and the work currently in progress at the proxy. Detailed Request Labels Every counter in this family samples once per client request, and its `status` is the status the caller received. A request whose first target failed and whose fallback then succeeded is a single `status="200"` sample here; the failure it recovered from is not represented in this family at all. Count upstream attempts with the [deployment metrics](#deployment-metrics) and read individual attempts in the usage log. For the three detailed request counters, `stream` records whether the client requested streaming. `is_fallback` records whether a fallback target served the request and does not appear on latency metrics. `provider_key_name` and `user_name` are human-readable companions to their IDs. Each name has a one-to-one relationship with its ID, so it does not add another series dimension. `user_name` is `unknown` until the control plane supplies it. A failed request carries the same `provider`, `upstream_model`, and provider-key labels as a successful one, so a failure rate per provider or per provider key is computable from this family alone. The upstream labels report the last target the request selected: under retry or failover that is the attempt whose error the caller received. They report `unknown` only when the request never selected a target — an unknown model, an input guardrail block, a budget refusal, or a body refused before dispatch. `inbound_protocol` is a bounded protocol family derived from the normalized endpoint. The Anthropic-protocol endpoints report `anthropic`; `/mcp`, `/a2a`, `/v1/realtime`, and `/passthrough_route` report `mcp`, `a2a`, `realtime`, and `passthrough`; and the remaining gateway endpoints report `openai`. The in-flight gauge uses the same values. `upstream_protocol` is the other half of that pair: the protocol AISIX spoke to the upstream that served the request. It is resolved from the selected target's Provider Key by the same rule dispatch uses, so it reports the wire shape the request was actually converted to — `openai`, `anthropic`, `bedrock`, `vertex`, or `azure-openai` — and `unknown` when the request selected no upstream at all. Group by both labels to separate native traffic from cross-protocol conversion; see [Separate Cross-Protocol Conversion Traffic](#separate-cross-protocol-conversion-traffic). Do not substitute `provider` for `upstream_protocol`. `provider` is an open vendor string, and the same vendor can be reached through different adapters, so a PromQL mapping from vendor names to protocols has to be maintained by hand and silently misreports every custom or OpenAI-compatible vendor. `endpoint` is always a normalized route template, never a raw request path. Routes with a path parameter collapse to one series — `/v1/batches/:id`, `/v1/videos/:id`, and `/mcp/{server}` — while paths under `/passthrough/` use `/passthrough_route` and an unrecognized path reports `other`. ### `aisix_requests_total` Proxy request outcomes in the compatibility series with the broadest endpoint coverage. Type: counter. ##### Labels `provider`, `model`, `status`, `outcome` ##### Values for `outcome` `success`, `client_error`, `rate_limited`, `upstream_error` `success` covers HTTP 200–399, `client_error` covers HTTP 400–499 except 429, `rate_limited` is HTTP 429, and all other statuses map to `upstream_error`. ##### Behavior A2A agent calls use `provider="a2a"` and `model="a2a"`. ##### PromQL example ``` sum(rate(aisix_requests_total[5m])) by (outcome) ``` ### `aisix_llm_requests_total` Model-inference request outcomes, including successful and failed requests. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name`, `stream`, `is_fallback`, `status`, `outcome` ##### Values for `inbound_protocol` `openai`, `anthropic`, `realtime` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ##### Values for `stream` `false`, `true` ##### Values for `is_fallback` `false`, `true` ##### Values for `outcome` `success`, `client_error`, `rate_limited`, `upstream_error` `success` covers HTTP 200–399, `client_error` covers HTTP 400–499 except 429, `rate_limited` is HTTP 429, and all other statuses map to `upstream_error`. ##### Behavior Covers the endpoints that call a model: chat-completions, completions, messages, count-tokens, responses, embeddings, rerank, audio, image generation and editing, videos, and realtime sessions. Requests that reach no model are counted in `aisix_proxy_requests_total` only — MCP tool calls, A2A agent calls, the provider passthrough, and the file, batch, and fine-tuning management routes. A request refused before dispatch, such as an oversized body, is counted against the endpoint it targeted, so a success rate over an endpoint includes those failures in its denominator. ### `aisix_proxy_requests_total` Detailed request outcomes for all proxied traffic, model-inference and otherwise. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name`, `stream`, `is_fallback`, `status`, `outcome` ##### Values for `inbound_protocol` `openai`, `anthropic`, `mcp`, `a2a`, `realtime`, `passthrough` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ##### Values for `stream` `false`, `true` ##### Values for `is_fallback` `false`, `true` ##### Values for `outcome` `success`, `client_error`, `rate_limited`, `upstream_error` `success` covers HTTP 200–399, `client_error` covers HTTP 400–499 except 429, `rate_limited` is HTTP 429, and all other statuses map to `upstream_error`. ##### Behavior One sample per client request, carrying the status the caller received. Retries and failovers within a request do not add samples, so this counter answers "how many requests did callers make, and how did they end" — not "how many calls did the gateway place upstream". ### `aisix_proxy_failed_requests_total` The subset of proxy requests whose outcome is not `success`. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name`, `stream`, `is_fallback`, `status`, `outcome` ##### Values for `inbound_protocol` `openai`, `anthropic`, `mcp`, `a2a`, `realtime`, `passthrough` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ##### Values for `stream` `false`, `true` ##### Values for `is_fallback` `false`, `true` ##### Values for `outcome` `client_error`, `rate_limited`, `upstream_error` `client_error` covers HTTP 400–499 except 429, `rate_limited` is HTTP 429, and all other statuses map to `upstream_error`. ### `aisix_proxy_in_flight_requests` Requests currently being handled by the proxy, grouped by normalized endpoint and inbound protocol. Type: gauge. ##### Labels `endpoint`, `inbound_protocol` ##### Values for `inbound_protocol` `openai`, `anthropic`, `mcp`, `a2a`, `realtime`, `passthrough` ##### Behavior MCP requests use `inbound_protocol="mcp"`, with `endpoint="/mcp"` for the aggregated gateway and `endpoint="/mcp/{server}"` for the per-server endpoint. A2A calls use `endpoint="/a2a"` and `inbound_protocol="a2a"`. Requests normalized to `endpoint="/passthrough_route"` use `inbound_protocol="passthrough"`. ##### PromQL example ``` sum(aisix_proxy_in_flight_requests) by (endpoint, inbound_protocol) ``` ### `aisix_proxy_client_cancelled_requests_total` Requests whose caller disconnected before the gateway sent response headers. Type: counter. ##### Labels `endpoint`, `model`, `provider_key_id`, `provider_key_name` ##### Behavior These requests never reach a normal outcome, so they do not appear in the other request counters. The gateway records them here and in the access log with status `499`. A rising rate usually means callers give up while waiting for the first token. Break the series down by `model` and compare it against `aisix_llm_time_to_first_token_seconds` for the same model. `model` is the model the caller addressed, and the provider-key labels identify the target the request was waiting on. A caller that disconnects before the gateway resolves them — during the request body upload, for example — reports `unknown` for all three. The caller identity, status, and outcome labels of the other request counters are deliberately absent. A cancelled request has no status and often no team or user, so those dimensions would be `unknown` on every sample. A caller that disconnects after response headers were sent is not counted here. That request already has a normal outcome and a usage event. ##### PromQL example ``` sum(rate(aisix_proxy_client_cancelled_requests_total[5m])) by (endpoint, model) ``` ### `aisix_proxy_request_body_limit_rejections_total` Requests refused for exceeding `proxy.request_body_limit_bytes`, grouped by how the gateway finished reading the refused body. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `outcome` ##### Values for `inbound_protocol` `openai`, `anthropic`, `mcp`, `a2a`, `realtime`, `passthrough` ##### Values for `outcome` `completed`, `cap_reached`, `timeout`, `client_read_error` `completed` means the caller sent the whole body it declared, so it could read the 413. `cap_reached` and `timeout` mean the gateway stopped absorbing the body first, and `client_read_error` means the caller went away mid-body — in all three the caller usually sees a closed connection instead of the response. ##### Behavior The gateway reads and discards the body of a refused request so the caller can receive the `413` on the same connection. That read is bounded, and `outcome` reports how it ended. A rising share of any outcome other than `completed` means callers are seeing closed connections instead of `413` responses. The matching `aisix::body_limit` log entry carries the declared size, the configured limit, and the bytes read for the same `request_id`. Only requests that declare a `Content-Length` over the limit are counted here. A chunked body over the limit is rejected while it is being read, has no comparable outcome, and appears in `aisix_requests_total` with status `413`. ##### PromQL example ``` sum(rate(aisix_proxy_request_body_limit_rejections_total[5m])) by (endpoint, inbound_protocol, outcome) ``` ### `aisix_auth_decisions_total` Caller authentication decisions across API key, JWT, and missing-credential paths. Type: counter. ##### Labels `method`, `result`, `reason` ##### Values for `method` `api_key`, `jwt`, `none` ##### Values for `result` `allowed`, `denied` ##### Behavior `reason` is `none` for an allowed request. Denied requests use a bounded reason such as `missing_credentials`, `unknown_key`, `key_expired`, `jwt_bad_signature`, `jwt_untrusted_issuer`, or `jwt_identity_unmapped`. ##### PromQL example ``` sum(rate(aisix_auth_decisions_total{result="denied"}[5m])) by (method, reason) ``` ### Latency Metrics Inspect latency on one gateway or aggregate histogram buckets across gateway instances. Latency Aggregation and Labels Use histograms for service-level dashboards and alerts across gateway instances because their `_bucket`, `_sum`, and `_count` series can be aggregated before using `histogram_quantile()`. Use summaries to inspect precomputed quantiles from one gateway instance. Summary quantiles cannot be aggregated across instances, so do not average them. Each histogram has its own bucket boundaries, because the two distributions differ: end-to-end latency starts in the milliseconds, while time to first token cannot be faster than the upstream that produces the token. Both sets are configurable. `env_id` identifies the environment served by an AISIX gateway connected to AISIX Cloud and is `unknown` when the AISIX gateway is not connected to AISIX Cloud. `status_class` is one of `2xx`, `3xx`, `4xx`, `5xx`, or `other`. Per-key and per-user labels are excluded to keep the number of bucket series manageable; use usage analytics for those dimensions. ### `aisix_request_duration_seconds` Request duration across proxy endpoints in the compatibility series. HTTP streams record time until the response starts; realtime records the full WebSocket session until it closes. Type: summary. ##### Labels `provider`, `model`, `status` ##### Behavior Summary quantiles are calculated by each AISIX instance and cannot be aggregated across instances. ### `aisix_llm_request_duration_seconds` Detailed request duration for model-inference endpoints. HTTP streams record time until the response starts; realtime records the full WebSocket session until it closes. Type: summary. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name`, `stream`, `status`, `outcome` ##### Values for `inbound_protocol` `openai`, `anthropic`, `realtime` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ##### Values for `stream` `false`, `true` ##### Values for `outcome` `success`, `client_error`, `rate_limited`, `upstream_error` `success` covers HTTP 200–399, `client_error` covers HTTP 400–499 except 429, `rate_limited` is HTTP 429, and all other statuses map to `upstream_error`. ### `aisix_proxy_request_duration_seconds` Detailed request duration for all proxied traffic. HTTP streams record time until the response starts; realtime records the full WebSocket session until it closes. Type: summary. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name`, `stream`, `status`, `outcome` ##### Values for `inbound_protocol` `openai`, `anthropic`, `mcp`, `a2a`, `realtime`, `passthrough` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ##### Values for `stream` `false`, `true` ##### Values for `outcome` `success`, `client_error`, `rate_limited`, `upstream_error` `success` covers HTTP 200–399, `client_error` covers HTTP 400–499 except 429, `rate_limited` is HTTP 429, and all other statuses map to `upstream_error`. ### `aisix_llm_time_to_first_token_seconds` Time from the upstream attempt start to the first streamed frame for streaming chat-completions, messages, and responses requests. Type: summary. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name` ##### Values for `inbound_protocol` `openai`, `anthropic` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ### `aisix_request_e2e_latency_seconds` Client-perceived latency for chat completions, messages, and responses, plus agent-call latency for A2A, including the full duration of streams. Type: histogram. ##### Labels `env_id`, `endpoint`, `model`, `provider`, `status_class`, `streaming` ##### Values for `status_class` `2xx`, `3xx`, `4xx`, `5xx`, `other` ##### Values for `streaming` `false`, `true` ##### Behavior Default buckets range from 5 milliseconds to 600 seconds, and are configurable with `observability.metrics.buckets.request_e2e_latency`. The low boundaries record fast responses such as cache hits and requests rejected before dispatch. Aggregate `_bucket` series before calculating percentiles across gateway instances. Each covered model-inference request is observed once. Non-streaming requests and failures are recorded when the handler returns. Streaming requests are recorded when the stream finishes, including client cancellation; a canceled model stream retains its committed status and records the duration up to cancellation. A2A calls use `endpoint="/a2a"`. Dispatched calls record the agent-call lifetime, including the complete stream. A pre-dispatch rejection that reaches A2A accounting records zero duration, and an abandoned A2A stream records `499` in the `4xx` status class. Calls rejected before A2A accounting are absent from this histogram. ##### PromQL example ``` histogram_quantile( 0.90, sum by (le) (rate(aisix_request_e2e_latency_seconds_bucket[5m])) ) ``` ### `aisix_request_ttft_seconds` Time from the upstream attempt start to the first streamed frame for streaming chat-completions, messages, and responses requests. Type: histogram. ##### Labels `env_id`, `endpoint`, `model`, `provider`, `status_class`, `streaming` ##### Values for `status_class` `2xx`, `3xx`, `4xx`, `5xx`, `other` ##### Values for `streaming` `true` ##### Behavior The first frame stops the timer even when it carries no visible generated output, such as a role-only opener or an Anthropic `message_start` event. Content, reasoning, and tool-call frames also stop it. Default buckets range from 50 milliseconds to 300 seconds, and are configurable with `observability.metrics.buckets.request_ttft`. Lower the first boundaries when a nearby model server can produce output in under 50 milliseconds. Deployments that use only hosted providers can raise the floor to remove consistently empty buckets. ##### PromQL example ``` histogram_quantile( 0.90, sum by (le) (rate(aisix_request_ttft_seconds_bucket[5m])) ) ``` ### Usage and Cost Metrics Measure token volume, estimated spend, and normalized client usage. Which Endpoints Report Tokens Every endpoint that receives token counts from the upstream records them here: chat-completions, completions, messages, responses, embeddings, rerank, the audio transcription routes, image generation and editing, and realtime sessions. Two model-inference endpoints report no tokens because they are not billed that way — `/v1/audio/speech` is billed per input character and `/v1/videos` per video. Both still count as requests, so divide token totals by requests only within one `endpoint`, never across all of them. The per-request token and spend series — the three `aisix_llm_*_tokens_total` counters and `aisix_llm_spend_micro_usd_total` — carry the same labels as the detailed request counters, so a query can join token volume to request outcomes on `endpoint`, `model`, `provider`, and the caller identity labels. The other two do not: `aisix_tokens_consumed_total` is labeled by `provider` and `model` only, and `aisix_llm_tokens_by_client_total` by `client_type`, `model`, and `token_type`. Three further counters break out the tokens the upstream provider served from its own prompt cache. They carry the same labels as the token counters above, and each appears only once its value is non-zero, so a provider that reports no cache detail creates no series. They describe the \*\*provider's\*\* cache, not the AISIX response cache — that one is `aisix_cache_requests_total` under [Cache Metrics](#cache-metrics), and the two are unrelated. Providers report cache reads under two different accounting rules, which is why there are two read counters rather than one. `aisix_llm_cached_input_tokens_total` is the OpenAI shape, where the cached tokens are part of the reported prompt tokens and are therefore already inside `aisix_llm_input_tokens_total`. `aisix_llm_cache_read_input_tokens_total` and `aisix_llm_cache_creation_input_tokens_total` are the Anthropic shape, reported beside the input tokens rather than within them, so they are outside `aisix_llm_input_tokens_total` and inside `aisix_llm_total_tokens_total`. Keeping them apart is what makes a cross-protocol query correct: the input a request really consumed is `aisix_llm_input_tokens_total + aisix_llm_cache_read_input_tokens_total + aisix_llm_cache_creation_input_tokens_total`, and the part of it served from cache is `aisix_llm_cached_input_tokens_total + aisix_llm_cache_read_input_tokens_total`. Both expressions hold for either provider shape. See [Calculate Prompt Cache Hit Rate](#calculate-prompt-cache-hit-rate). ### `aisix_tokens_consumed_total` The sum of total tokens across every endpoint that reports token usage, in the compatibility series with the broadest coverage. Type: counter. ##### Labels `provider`, `model` ### `aisix_llm_input_tokens_total` Input tokens reported by the upstream, across every endpoint that reports token usage. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name` ##### Values for `inbound_protocol` `openai`, `anthropic`, `realtime` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ### `aisix_llm_output_tokens_total` Output tokens reported by the upstream, across every endpoint that reports token usage. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name` ##### Values for `inbound_protocol` `openai`, `anthropic`, `realtime` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ### `aisix_llm_total_tokens_total` Total tokens reported by the upstream, across every endpoint that reports token usage. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name` ##### Values for `inbound_protocol` `openai`, `anthropic`, `realtime` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ### `aisix_llm_cached_input_tokens_total` Input tokens the upstream served from its prompt cache and reported as part of its prompt token count. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name` ##### Values for `inbound_protocol` `openai`, `anthropic`, `realtime` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ##### Behavior The OpenAI accounting shape: `prompt_tokens_details.cached_tokens` for OpenAI and OpenAI-compatible providers, `prompt_cache_hit_tokens` for DeepSeek, and `cachedContentTokenCount` for Gemini on Vertex AI. These tokens are already counted in `aisix_llm_input_tokens_total`, so adding the two together double-counts them. Recorded only when the upstream reports the field. AISIX never infers a cache hit from response timing or prompt similarity. ### `aisix_llm_cache_read_input_tokens_total` Input tokens the upstream served from its prompt cache and reported separately from its input token count. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name` ##### Values for `inbound_protocol` `openai`, `anthropic`, `realtime` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ##### Behavior The Anthropic accounting shape: `cache_read_input_tokens` for Anthropic and `cacheReadInputTokens` for Amazon Bedrock. These tokens are \*\*not\*\* in `aisix_llm_input_tokens_total`, and they are in `aisix_llm_total_tokens_total`. ### `aisix_llm_cache_creation_input_tokens_total` Input tokens the upstream wrote into its prompt cache. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name` ##### Values for `inbound_protocol` `openai`, `anthropic`, `realtime` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ##### Behavior `cache_creation_input_tokens` for Anthropic and `cacheWriteInputTokens` for Amazon Bedrock. Like cache reads in the same shape, these tokens are outside `aisix_llm_input_tokens_total` and inside `aisix_llm_total_tokens_total`. Providers bill cache writes above the standard input rate and cache reads well below it, so a workload that writes far more than it reads costs more than the same workload with caching switched off. Compare this counter with `aisix_llm_cache_read_input_tokens_total` to catch it. ### `aisix_llm_spend_micro_usd_total` Estimated spend in micro-USD, where 1 USD equals 1,000,000 micro-USD. Recorded wherever the gateway resolves a price for the request. Type: counter. ##### Labels `endpoint`, `inbound_protocol`, `upstream_protocol`, `provider`, `model`, `upstream_model`, `provider_key_id`, `provider_key_name`, `api_key_id`, `team_id`, `user_id`, `user_name` ##### Values for `inbound_protocol` `openai`, `anthropic`, `realtime` ##### Values for `upstream_protocol` `openai`, `anthropic`, `bedrock`, `vertex`, `azure-openai`, `unknown` The protocol the gateway spoke to the upstream, which is independent of `inbound_protocol`. `unknown` means the request selected no upstream — a rejection before dispatch, or a route that calls no model. ### `aisix_llm_tokens_by_client_total` Token volume across every endpoint that reports token usage, grouped by normalized client, requested model, and token type. Type: counter. ##### Labels `client_type`, `model`, `token_type` ##### Values for `client_type` `openai-python`, `openai-node`, `anthropic-python`, `anthropic-typescript`, `claude-code`, `codex`, `cline`, `roo-code`, `kilocode`, `zoo-code`, `github-copilot`, `cursor`, `opencode`, `qwen-code`, `gemini-cli`, `crush`, `zed`, `aider`, `vercel-ai-sdk`, `langchain`, `llamaindex`, `litellm`, `curl`, `python-requests`, `httpx`, `aiohttp`, `okhttp`, `go-http-client`, `node`, `postman`, `browser`, `other`, `unknown` An unrecognized `User-Agent` maps to `other`; a missing `User-Agent` maps to `unknown`. Deployments can extend this set with admin-defined mapping rules (`observability.metrics.client_type_rules`). Full user-agent strings and versions remain in request logs and usage analytics. ##### Values for `token_type` `input`, `output`, `total` `total` includes input, output, and Anthropic cache-creation and cache-read tokens. ##### Behavior `model` is the caller-requested model name, matching the `model` label on the `aisix_llm_*` token series. Routing, semantic, ensemble, and fallback dispatch retain this alias instead of reporting the selected direct model. The bounded `client_type` allowlist prevents a client-controlled user-agent from creating unbounded Prometheus cardinality. Concrete model aliases are bounded by the configured model set, but a wildcard alias records each concrete model name requested through that pattern. Restrict wildcard access and monitor label growth when cardinality matters. Admin-defined mapping rules (`observability.metrics.client_type_rules`) can classify additional clients. Rules are matched before the built-in allowlist and emit a fixed, validated label value, so the label set stays bounded. Aggregated across all labels, the cache-inclusive `total` aligns with `aisix_llm_total_tokens_total` for the same endpoint population. Individual series do not align because the two metric families use different label sets; the dedicated client-type series avoids multiplying the per-key token series by another label dimension. Anthropic reports cache tokens separately from input tokens, so `total` can exceed `input` plus `output`. ##### PromQL example ``` sum by (client_type, model, token_type) ( rate(aisix_llm_tokens_by_client_total[5m]) ) ``` ### Rate Limit and Budget Metrics Monitor rate-limit rejections and the latest quota or budget state for each label set. ### `aisix_ratelimit_rejections_total` Requests rejected by the shared rate-limit gate across proxy endpoints. Type: counter. ##### Labels `scope`, `layer`, `policy_id` ##### Values for `scope` `requests`, `tokens` Concurrency-limit rejections use `requests`; the runtime does not emit a separate `concurrency` value. ##### Values for `layer` `api_key`, `model`, `mcp`, `policy` ##### Behavior `policy_id` is empty outside `layer="policy"`; on the policy layer it carries the configured policy ID. ### `aisix_ratelimit_remaining_requests` Remaining request quota reported while processing a chat-completions request, grouped by API key and model. Type: gauge. ##### Labels `api_key_id`, `model` ##### Behavior `model` is the configured model the request resolved to — a wildcard alias reports the row, such as `openai/*`, not the concrete name the caller sent. Retired to `NaN` within seconds of the API key being deleted, or rebound to another team or member. `NaN` rather than `0` because `0` is a meaningful reading on this gauge; every comparison against it is false, so an alert on a retired series stops firing. An aggregation over a family containing one reads `NaN` too. ### `aisix_ratelimit_remaining_tokens` Remaining token quota reported while processing a chat-completions request, grouped by API key and model. Type: gauge. ##### Labels `api_key_id`, `model` ##### Behavior `model` is the configured model the request resolved to — a wildcard alias reports the row, such as `openai/*`, not the concrete name the caller sent. Retired to `NaN` within seconds of the API key being deleted, or rebound to another team or member. `NaN` rather than `0` because `0` is a meaningful reading on this gauge; every comparison against it is false, so an alert on a retired series stops firing. An aggregation over a family containing one reads `NaN` too. ### `aisix_budget_limit_usd` Budget limit in USD. Type: gauge. ##### Labels `api_key_id`, `team_id`, `user_id`, `user_name` ##### Behavior Removing a budget from a key does not zero this gauge — it holds the last budgeted value, which is why queries guard on `aisix_budget_details_present == 1`. Retired to `NaN` within seconds of the API key being deleted, or rebound to another team or member. `NaN` rather than `0` because `0` is a meaningful reading on this gauge; every comparison against it is false, so an alert on a retired series stops firing. An aggregation over a family containing one reads `NaN` too. ### `aisix_budget_spent_usd` Budget spent in USD. Type: gauge. ##### Labels `api_key_id`, `team_id`, `user_id`, `user_name` ##### Behavior Removing a budget from a key does not zero this gauge — it holds the last budgeted value, which is why queries guard on `aisix_budget_details_present == 1`. Retired to `NaN` within seconds of the API key being deleted, or rebound to another team or member. `NaN` rather than `0` because `0` is a meaningful reading on this gauge; every comparison against it is false, so an alert on a retired series stops firing. An aggregation over a family containing one reads `NaN` too. ### `aisix_budget_remaining_usd` Budget remaining in USD. Type: gauge. ##### Labels `api_key_id`, `team_id`, `user_id`, `user_name` ##### Behavior Removing a budget from a key does not zero this gauge — it holds the last budgeted value, which is why queries guard on `aisix_budget_details_present == 1`. Retired to `NaN` within seconds of the API key being deleted, or rebound to another team or member. `NaN` rather than `0` because `0` is a meaningful reading on this gauge; every comparison against it is false, so an alert on a retired series stops firing. An aggregation over a family containing one reads `NaN` too. ### `aisix_budget_reset_seconds` Seconds until the budget period resets. Type: gauge. ##### Labels `api_key_id`, `team_id`, `user_id`, `user_name` ##### Behavior Removing a budget from a key does not zero this gauge — it holds the last budgeted value, which is why queries guard on `aisix_budget_details_present == 1`. Retired to `NaN` within seconds of the API key being deleted, or rebound to another team or member. `NaN` rather than `0` because `0` is a meaningful reading on this gauge; every comparison against it is false, so an alert on a retired series stops firing. An aggregation over a family containing one reads `NaN` too. ### `aisix_budget_details_present` Whether budget details are populated. Type: gauge. ##### Labels `api_key_id`, `team_id`, `user_id`, `user_name` ##### Values for `metric value` `0`, `1` `1` means budget details are present; `0` means they have been cleared. ### Cache Metrics Measure response-cache effectiveness per policy and watch the semantic layer’s embedding and store health. Cache Hit Rate Every request covered by an enabled cache policy with an available backend records one `aisix_cache_requests_total` observation. Requests with no matching policy or an unavailable backend are not counted — the gate never opened — so the series measures policy effectiveness, not total traffic. Hit rate per policy: `sum by (policy) (rate(aisix_cache_requests_total{outcome=~"hit_exact|hit_semantic"}[5m])) / sum by (policy) (rate(aisix_cache_requests_total[5m]))`. Split by `outcome` to see how much the semantic layer adds over exact matching. Semantic-layer failures degrade to ordinary misses, which makes a broken embedding model or store indistinguishable from a healthy low hit rate in the outcome counter alone. Alert on `aisix_cache_semantic_embedding_failures_total` and `aisix_cache_semantic_store_failures_total` to separate the two. ### `aisix_cache_requests_total` Cache-eligible requests by policy name and outcome, counted once per request when a matching enabled policy with an available backend opened the cache gate. Type: counter. ##### Labels `policy`, `outcome` ##### Values for `outcome` `hit_exact`, `hit_semantic`, `miss`, `bypass` `bypass` means the caller sent `Cache-Control: no-cache`, which skips the read path. `no-store` keeps the read path active, so it records an ordinary hit or miss and only suppresses a write after a miss. ### `aisix_cache_semantic_embedding_seconds` Latency of the semantic layer’s embedding calls, per policy. Summary series without fixed buckets. Type: summary. ##### Labels `policy` ### `aisix_cache_semantic_embedding_failures_total` Embedding failures on the semantic layer. Failed requests proceed to the upstream uncached. Type: counter. ##### Labels `policy`, `cause` ##### Values for `cause` `resolve`, `embed` `resolve` (embedding model missing or not an embedding model) is counted once per eligible request, including requests that then hit the exact layer; `embed` (provider call failed or timed out) is counted per embedding call. ### `aisix_cache_semantic_store_failures_total` Semantic-store operation failures by operation. The in-process store cannot fail; shared (Redis) stores can. Failures degrade to an ordinary miss. Type: counter. ##### Labels `policy`, `op` ##### Values for `op` `lookup`, `store` ### Deployment Metrics Monitor how each target model behaves: upstream attempts and their outcomes, fallbacks between targets, and whether a target remains in rotation. Attempts, Not Requests These counters sample once per upstream \*\*attempt\*\*, and each sample is filed under the target that was attempted rather than the Model Group the caller named. One client request that failed over across three targets is three samples here and one sample in the [request metrics](#request-metrics). Use this family to ask which target is failing, and the request family to ask what callers experienced. Only attempts that reached the upstream are counted. An attempt the gateway refused on its own — the target was over its own rate limit, or its credentials or endpoint failed validation before anything was sent — produced no upstream response, so it is absent here while still appearing in the usage log. That keeps a misconfiguration from reading as an unhealthy target. Emitted by the endpoints that dispatch through a Model Group: `/v1/chat/completions`, `/v1/messages`, and `/v1/responses`. The other endpoints call one model per request and are fully described by the request metrics. ### `aisix_deployment_requests_total` Upstream attempts dispatched to a target model, whatever their outcome. Type: counter. ##### Labels `provider`, `model`, `upstream_model`, `provider_key_id` ##### PromQL example ``` sum(rate(aisix_deployment_requests_total[5m])) by (model) ``` ### `aisix_deployment_success_responses_total` Upstream attempts a target model answered successfully. Type: counter. ##### Labels `provider`, `model`, `upstream_model`, `provider_key_id` ##### Behavior For a streamed response, an attempt counts as successful once the upstream stream is established. A stream that breaks after that point is not reclassified here. ### `aisix_deployment_failure_responses_total` Upstream attempts that failed at the target model, including the failures a fallback went on to rescue. Type: counter. ##### Labels `provider`, `model`, `upstream_model`, `provider_key_id` ##### PromQL example ``` sum(rate(aisix_deployment_failure_responses_total[5m])) by (model) / sum(rate(aisix_deployment_requests_total[5m])) by (model) ``` ### `aisix_routing_successful_fallbacks_total` Fallback attempts that served the request after an earlier target failed. Type: counter. ##### Labels `model`, `fallback_model` ##### Behavior `model` is what the caller asked for — the Model Group name. `fallback_model` is the target the gateway moved to. ### `aisix_routing_failed_fallbacks_total` Fallback attempts that failed in turn. Type: counter. ##### Labels `model`, `fallback_model` ##### Behavior A request rescued by its second fallback contributes one sample here and one to the successful family. ### `aisix_deployment_state` Whether a target model is in rotation. Type: gauge. ##### Labels `provider`, `model`, `upstream_model`, `provider_key_id` ##### Values for `metric value` `0`, `2` `0` is healthy. `2` is out of rotation because the target is cooling down or failing a background health check. ### `aisix_deployment_cooled_down_total` Number of times a target model entered cooldown. Type: counter. ##### Labels `provider`, `model`, `upstream_model`, `provider_key_id` ### Guardrail Metrics Track per-guardrail execution latency and outcomes, plus aggregate blocks and fail-open bypasses, across all endpoints. Guardrail Execution Latency Every timed guardrail execution records an `aisix_guardrail_latency_seconds` observation. A guardrail normally runs once per applicable phase, while streaming window scans can record several output executions for one response. The histogram uses configurable buckets that default to 1 ms through 30 s, so `histogram_quantile` yields P50/P95/P99 per guardrail. The `kind` label identifies the configured guardrail implementation, and the `result` label separates fail-open bypasses (`bypassed`, with the failure tag in `error_type`) from policy decisions. A fail-closed evaluation failure surfaces as `blocked`; filter on `error_type` to distinguish it from a policy match. The `_count` series doubles as a per-guardrail execution counter: `sum by (guardrail, result) (rate(aisix_guardrail_latency_seconds_count[5m]))` gives execution and block rates without a separate counter. The aggregate counters cover guarded endpoints but count different units. `aisix_guardrail_blocks_total` counts rejected requests, while `aisix_guardrail_bypasses_total` counts bypass events, including proxy-raised passes before a guardrail member executes. They omit guardrail, kind, and phase labels; use the histogram `_count` series for timed execution dimensions. ### `aisix_guardrail_latency_seconds` Wall-clock duration of one timed guardrail execution across all endpoints. `guardrail` is the configured guardrail name. Type: histogram. ##### Labels `env_id`, `guardrail`, `kind`, `phase`, `result`, `error_type` ##### Values for `kind` `keyword`, `pii`, `aliyun_ai_guardrail`, `aliyun_text_moderation`, `azure_content_safety`, `azure_content_safety_text_moderation`, `bedrock`, `lakera`, `openai_moderation`, `presidio`, `semantic`, `custom` `keyword` and `pii` run in-process. `semantic` calls a configured embedding model, `custom` runs the configured script, and the remaining kinds call moderation services. ##### Values for `phase` `input`, `output` ##### Values for `result` `allowed`, `blocked`, `masked`, `bypassed`, `would_block`, `would_mask` `bypassed` is an evaluation failure that failed open; a fail-closed evaluation failure records as `blocked`. `would_block` / `would_mask` come from `enforcement_mode: monitor` guardrails. ##### Values for `error_type` `none`, `aliyun_5xx`, `aliyun_config_error`, `aliyun_bad_response`, `aliyun_throttled`, `aliyun_timeout`, `azure_cs_5xx`, `azure_cs_config_error`, `azure_cs_throttled`, `azure_cs_timeout`, `bedrock_5xx`, `bedrock_throttled`, `bedrock_timeout`, `bedrock_too_large`, `lakera_5xx`, `lakera_config_error`, `lakera_throttled`, `lakera_timeout`, `lakera_too_large`, `openai_moderation_5xx`, `openai_moderation_config_error`, `openai_moderation_throttled`, `openai_moderation_timeout`, `openai_moderation_too_large`, `presidio_5xx`, `presidio_config_error`, `presidio_throttled`, `presidio_timeout`, `presidio_too_large`, `custom_timeout`, `custom_script_error`, `custom_engine_error`, `custom_no_verdict`, `custom_bad_verdict`, `custom_unknown_action`, `semantic_embed_unresolved`, `semantic_embed_timeout`, `semantic_embed_upstream` Set to the bounded failure tag when an evaluation failure produces `result="blocked"`, `result="bypassed"`, or a monitor-mode `result="would_block"`; otherwise `none`. The `bypassed` values are also used by `aisix_guardrail_bypasses_total{reason}`. That counter additionally uses `unscannable_body` for proxy-raised passes that do not create a timed histogram observation. ##### Behavior Default buckets: 0.001, 0.0025, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, and 30 seconds, configurable with `observability.metrics.buckets.guardrail_latency`. Monitor-mode executions record `would_block` / `would_mask` while the request proceeds, so a staged policy can be sized before enforcing it. PII and keyword masking performed by synchronous per-field operations is not timed by this series; its in-process cost is measured in microseconds. ### `aisix_guardrail_blocks_total` Requests rejected by input or output guardrail enforcement across all endpoints, including fail-closed paths such as a streamed-output buffer overflow before a guardrail member executes. Type: counter. ##### Labels None. ### `aisix_guardrail_bypasses_total` Guardrail evaluation failures or proxy-raised unreadable-content passes where a fail-open policy let the request or response continue, across all guarded endpoints. Type: counter. ##### Labels `reason` ##### Values for `reason` `aliyun_5xx`, `aliyun_config_error`, `aliyun_bad_response`, `aliyun_throttled`, `aliyun_timeout`, `azure_cs_5xx`, `azure_cs_config_error`, `azure_cs_throttled`, `azure_cs_timeout`, `bedrock_5xx`, `bedrock_throttled`, `bedrock_timeout`, `bedrock_too_large`, `lakera_5xx`, `lakera_config_error`, `lakera_throttled`, `lakera_timeout`, `lakera_too_large`, `openai_moderation_5xx`, `openai_moderation_config_error`, `openai_moderation_throttled`, `openai_moderation_timeout`, `openai_moderation_too_large`, `presidio_5xx`, `presidio_config_error`, `presidio_throttled`, `presidio_timeout`, `presidio_too_large`, `custom_timeout`, `custom_script_error`, `custom_engine_error`, `custom_no_verdict`, `custom_bad_verdict`, `custom_unknown_action`, `semantic_embed_unresolved`, `semantic_embed_timeout`, `semantic_embed_upstream`, `unscannable_body` ### A2A Metrics Measure Agent-to-Agent traffic by the agent that was reached and the operation that was invoked. Agent and Operation Labels `agent` is a registered A2A agent, and `operation` is the canonical operation a caller invoked. A2A 0.3 and 1.0 spell one operation two ways, so AISIX records the canonical form and a deployment fronting both versions still aggregates as one series. An unrecognized method becomes `unknown`. Task ids, context ids, and JSON-RPC request ids are deliberately absent. They are what makes an individual call traceable, which is exactly what makes them unusable as label values; find them in usage events and traces instead. `aisix_a2a_requests_total` does not agree with `aisix_proxy_requests_total{endpoint="/a2a"}`, by design. A call refused before its agent is resolved has no agent to file under and is counted only in the proxy family, and a stream the caller abandons is a `4xx` here but a `2xx` there, because the response really did begin as a 200. Read this family for agent health and the proxy family for route traffic. For streaming A2A, `aisix_proxy_request_duration_seconds` records time until the response starts. For dispatched calls, `aisix_request_e2e_latency_seconds{endpoint="/a2a"}` records the agent-call lifetime, including the complete stream. ### `aisix_a2a_requests_total` A2A calls by the agent reached, the canonical operation invoked, and the status class. Type: counter. ##### Labels `agent`, `operation`, `status` ##### Behavior Counted at the same point the usage event is emitted, so a call that is accounted for is also metered. ##### PromQL example ``` sum by (agent, operation) ( rate(aisix_a2a_requests_total{status!="2xx"}[5m]) ) ``` ### `aisix_a2a_ttfb_seconds` Time to the upstream agent's first streamed event on a streaming A2A call. Type: histogram. ##### Labels `agent`, `operation` ##### Behavior Named for the event rather than a token because an agent stream carries task updates, not tokens. Default buckets match `aisix_request_ttft_seconds`, and are configurable with `observability.metrics.buckets.a2a_ttfb`. Observed only on streaming operations that produced at least one event. ##### PromQL example ``` histogram_quantile( 0.90, sum by (le, agent) (rate(aisix_a2a_ttfb_seconds_bucket[5m])) ) ``` ### `aisix_a2a_stream_events_total` Events relayed downstream on streaming A2A calls. Type: counter. ##### Labels `agent`, `operation` ##### Behavior Divided by `aisix_a2a_requests_total` over the streaming operations, it gives events per call — how chatty an agent is, and whether that changed. Both `message/stream` and `tasks/resubscribe` stream, so filter the denominator to both or the ratio is inflated. ##### PromQL example ``` sum by (agent) (rate(aisix_a2a_stream_events_total[5m])) / sum by (agent) ( rate(aisix_a2a_requests_total{operation=~"message/stream|tasks/resubscribe"}[5m]) ) ``` ### `aisix_a2a_task_state_total` A2A calls by the task state the agent last reported, normalized to the specification's set plus `unknown`. Type: counter. ##### Labels `agent`, `state` ##### Behavior A cumulative counter of the state each call ENDED on, so read it as a rate rather than as a live backlog: a task that later moves on does not decrement an earlier sample, and a client polling one task with `tasks/get` reports its state on every poll. A call whose upstream never answered reports no state and is not counted here. ##### PromQL example ``` sum by (agent, state) (rate(aisix_a2a_task_state_total[5m])) ``` ### Usage Event Metrics Distinguish usage-event emission attempts from events the delivery queue did not accept. Usage Event Delivery For each emission attempt, the event is either accepted by the queue or counted as a drop. Subtract the drop rate from the emission-attempt rate to calculate the accepted rate. Both counters carry the same `model`, `provider_key_id`, and `provider_key_name` labels, so that subtraction also holds per model and per provider key. Use it to identify whose queue handoffs were rejected. A rising drop rate concentrated on one model or one provider key points at traffic whose usage records never reached the worker. Queue acceptance is not delivery to storage. After the worker accepts an event, control-plane shipping can still fail (`telemetry batch failed (events dropped)`) and exporter pipelines can still drop after retries (`sink delivery dropped after retries`). Neither increments `aisix_usage_event_drops_total`. The `status_code` label is one of `2xx`, `3xx`, `4xx`, `5xx`, or `other`, while `status` carries the raw HTTP code alongside it — `status` names one failure mode, `status_code` rolls a family up, and both describe the same event. Example `handler` values include `chat`, `embeddings`, `messages`, and `mcp`. `user_id` is the organization member who owned the API key the request authenticated with, and `user_name` is that member's display name. Both are `unknown` when the key belongs to no member. They appear on both counters, so `emitted = delivered + dropped` still holds per member. Because a JWT-authenticated request runs as the key its identity resolves to, one member label covers every credential that member calls through. The name has a one-to-one relationship with the ID, so it does not add another series dimension. It is the name the control plane stamped when it last projected that API key, not a live lookup: a member renamed afterwards keeps their old name on these metrics until that key is written again for some other reason. `model` is the model the caller addressed, collapsed to a configured model name. A request that carried no resolvable model reports `unresolved`, and one with no model at all — an MCP tool call, an A2A agent call, a passthrough route — reports `unknown`, as do both provider-key labels wherever no upstream key was resolved. MCP tool calls use `handler="mcp"` and `inbound_protocol="mcp"`. Their usage-event payloads identify the MCP server and tool; token and cost fields are zero. A2A agent calls use `handler="a2a"`. The metric maps their bounded `inbound_protocol` label to `other`, while the delivered event uses `inbound_protocol="a2a"` and identifies the agent name and JSON-RPC method. The gateway estimates prompt and completion tokens from message text and sets `usage_estimated: true`; cost remains zero. ### `aisix_usage_events_emitted_total` Usage-event emission attempts, counted before AISIX tries to enqueue the event, including events the queue accepted and events it rejected. Type: counter. ##### Labels `handler`, `status_code`, `status`, `inbound_protocol`, `upstream_protocol`, `model`, `provider_key_id`, `provider_key_name`, `user_id`, `user_name` ##### Values for `handler` `a2a`, `audio`, `batch`, `batches`, `chat`, `completions`, `count_tokens`, `embeddings`, `files`, `fine_tuning`, `images`, `mcp`, `messages`, `passthrough_route`, `realtime`, `rerank`, `responses`, `videos` ##### Values for `status_code` `2xx`, `3xx`, `4xx`, `5xx`, `other` ##### Values for `inbound_protocol` `openai`, `anthropic`, `mcp`, `other` `other` includes A2A, realtime, passthrough, and every other value outside the three named protocol buckets. ### `aisix_usage_event_drops_total` Usage events that were not accepted by the queue. Type: counter. ##### Labels `reason`, `model`, `provider_key_id`, `provider_key_name`, `upstream_protocol`, `user_id`, `user_name` ##### Values for `reason` `sink_disabled`, `sink_full`, `sink_closed` ##### PromQL example ``` sum(rate(aisix_usage_event_drops_total[5m])) by (model, provider_key_name, reason) ``` ### Configuration Metrics Monitor whether AISIX can load and apply changes from its configuration source. Configuration State AISIX reflects the same live configuration state in these metrics and at `GET /status/config` on the metrics and status listener. Reload metrics apply to both file and etcd configuration sources. Revision and source-connection metrics are emitted only when AISIX loads configuration from etcd. Compare the observed and applied revisions to determine whether the gateway is serving the latest etcd configuration. ### `aisix_config_last_reload_successful` Whether the latest configuration load succeeded. Type: gauge. ##### Labels None. ##### Values for `metric value` `0`, `1` `1` indicates success and `0` indicates failure. ### `aisix_config_last_reload_success_timestamp_seconds` Unix timestamp in seconds of the last successful configuration load. Type: gauge. ##### Labels None. ##### Behavior The series is not emitted until a configuration load succeeds. ### `aisix_config_reloads_total` Full configuration reload attempts, including source fetch failures. Type: counter. ##### Labels None. ##### Behavior Incremental etcd watch events do not increment this counter. ### `aisix_config_reload_failures_total` Configuration reload failures grouped by reason. Type: counter. ##### Labels `reason` ##### Values for `reason` `fetch`, `parse`, `validate` `fetch` indicates that the source could not be read, `parse` indicates invalid source data, and `validate` indicates an invalid resource. ### `aisix_config_rejected_resources` Current number of rejected resources, grouped by resource kind. Type: gauge. ##### Labels `kind` ##### Behavior When all rejections for a resource kind clear, AISIX sets its existing series to `0`. ### `aisix_config_partially_compatible_resources` Served resources carrying at least one field this gateway version does not recognize, grouped by resource kind. Type: gauge. ##### Labels `kind` ##### Behavior A resource with several ignored fields counts once. Inspect `partially_compatible` at `GET /status/config` for the field paths and per-field counts. When all partially compatible resources for a kind clear, AISIX sets its existing series to `0`. ### `aisix_config_stale_served_resources` Resources whose latest source value was rejected while their last known good value remains in service, grouped by resource kind. Type: gauge. ##### Labels `kind` ##### Behavior Inspect `rejected` at `GET /status/config` to identify each resource and when stale serving began. When no stale value remains in service for a kind, AISIX sets its existing series to `0`. ### `aisix_config_observed_revision` Latest etcd revision observed by the gateway. Type: gauge. ##### Labels None. ### `aisix_config_applied_revision` The etcd revision represented by the configuration currently served by the gateway. Type: gauge. ##### Labels None. ### `aisix_config_hash_info` Applied configuration hash. Type: gauge. ##### Labels `hash` ##### Values for `metric value` `0`, `1` Filter for value `1` to select the current hash. AISIX retains earlier hash labels with value `0` after the applied configuration changes. ### `aisix_config_source_connected` Whether the gateway is connected to its etcd configuration source. Type: gauge. ##### Labels None. ##### Values for `metric value` `0`, `1` `1` indicates connected and `0` indicates disconnected. ## Analyze Metrics with PromQL[​](#analyze-metrics-with-promql "Direct link to Analyze Metrics with PromQL") Use these PromQL examples in the Prometheus expression browser or another compatible monitoring interface after configuring it to scrape the AISIX metrics endpoint. Adjust the time window, label filters, and grouping dimensions to match the traffic and gateway instances you want to inspect. ### Calculate Success Rate[​](#calculate-success-rate "Direct link to Calculate Success Rate") Calculate the success rate by dividing the successful request rate by the total request rate. The following query combines all model-inference traffic over a five-minute window: ``` sum(rate(aisix_llm_requests_total{outcome="success"}[5m])) / sum(rate(aisix_llm_requests_total[5m])) ``` Add an `endpoint` filter or group by `endpoint` when you want to analyze one API separately. To include traffic that is not counted as model inference — MCP tool calls, A2A agent calls, passthrough routes (endpoint label `/passthrough_route`), and the file, batch, and fine-tuning routes — run the same query against `aisix_proxy_requests_total`, which carries every proxied request. To measure only the primary routing path, restrict the numerator and denominator to requests that were not served by a fallback target: ``` sum(rate(aisix_llm_requests_total{outcome="success", is_fallback="false"}[5m])) / sum(rate(aisix_llm_requests_total{is_fallback="false"}[5m])) ``` Whether rate-limited requests belong in the population is an operational policy decision. To exclude clients that reached their quota from the denominator, use: ``` sum(rate(aisix_llm_requests_total{outcome="success"}[5m])) / sum(rate(aisix_llm_requests_total{outcome!="rate_limited"}[5m])) ``` ### Separate Cross-Protocol Conversion Traffic[​](#separate-cross-protocol-conversion-traffic "Direct link to Separate Cross-Protocol Conversion Traffic") Every detailed request, latency, and token series carries two protocol labels: `inbound_protocol`, the protocol the caller spoke to AISIX, and `upstream_protocol`, the protocol AISIX spoke to the upstream that served the request. Group by both to get a conversion matrix, where the diagonal is native traffic and every other cell is traffic AISIX translated: ``` sum by (inbound_protocol, upstream_protocol) (rate(aisix_llm_requests_total[5m])) ``` Conversion has a cost in fidelity — features that exist on one side of a translation have no counterpart on the other — so it is worth knowing how much of the fleet depends on it. To isolate one direction, name both ends: ``` # Anthropic-protocol callers served by an OpenAI-shape upstream sum(rate(aisix_llm_requests_total{inbound_protocol="anthropic", upstream_protocol="openai"}[5m])) ``` Compare reliability across upstream protocols to tell a provider problem apart from a translation problem. If one protocol's success rate diverges while the callers are the same, the difference is on the upstream side: ``` sum by (upstream_protocol) (rate(aisix_llm_requests_total{outcome="success"}[5m])) / sum by (upstream_protocol) (rate(aisix_llm_requests_total[5m])) ``` `upstream_protocol` reports `unknown` when a request selected no upstream at all — an unknown model, an input guardrail block, an AISIX Cloud budget refusal, or a body refused before dispatch. Those requests are real failures, so keep them in the denominator of a success rate and exclude them only when you are specifically comparing upstreams. Do not derive the upstream protocol from `provider` instead. `provider` is an open vendor string, the same vendor can be reached through more than one adapter, and a custom vendor served over an OpenAI-compatible endpoint carries whatever name the operator chose. A hand-maintained vendor-to-protocol mapping in PromQL misreports exactly the traffic this label exists to describe. ### Calculate Aggregate Latency Percentiles[​](#calculate-aggregate-latency-percentiles "Direct link to Calculate Aggregate Latency Percentiles") Combine histogram buckets across gateway instances before calculating a percentile. P90 is the value at or below which 90% of observations fall. Keep `le` in the `sum by` grouping, and add labels such as `model` or `provider` when you need a breakdown: ``` # P90 end-to-end latency across all matched gateway instances histogram_quantile( 0.90, sum by (le) (rate(aisix_request_e2e_latency_seconds_bucket{status_class="2xx"}[5m])) ) # P90 end-to-end latency per model histogram_quantile( 0.90, sum by (le, model) (rate(aisix_request_e2e_latency_seconds_bucket{status_class="2xx"}[5m])) ) # P90 time to first streamed frame per provider histogram_quantile( 0.90, sum by (le, provider) (rate(aisix_request_ttft_seconds_bucket[5m])) ) ``` Streaming end-to-end time covers the full generation, so streaming and non-streaming requests have different latency distributions. Use the `streaming` label to analyze them separately: ``` # P90 end-to-end latency for successful streaming requests histogram_quantile( 0.90, sum by (le) (rate(aisix_request_e2e_latency_seconds_bucket{streaming="true", status_class="2xx"}[5m])) ) ``` ### Compare Single-Instance Streaming Latency[​](#compare-single-instance-streaming-latency "Direct link to Compare Single-Instance Streaming Latency") Summary series expose precomputed `quantile` labels for each gateway instance. Select one scrape target and the model or provider you want to inspect: ``` # P90 time to first streamed frame for streaming chat completions aisix_llm_time_to_first_token_seconds{endpoint="/v1/chat/completions", quantile="0.9"} # P90 time to response start for that same traffic aisix_llm_request_duration_seconds{endpoint="/v1/chat/completions", stream="true", quantile="0.9"} ``` Both queries pin `endpoint` because the two series do not cover the same traffic by default. Time to first streamed frame is recorded for streaming `/v1/chat/completions`, `/v1/messages`, and `/v1/responses`, while the duration series covers every model-inference endpoint. Without the filter, traffic on the other model-inference endpoints can move the duration P90 without contributing to the TTFT P90. `aisix_llm_request_duration_seconds` is recorded when the gateway hands the response to the client. On a streamed request that happens before any frame is read, so the value covers the work up to response start and excludes the whole generation. It is not an end-to-end figure for streaming traffic and does not become one by filtering on `stream="true"`. The same applies to `aisix_request_duration_seconds` and `aisix_proxy_request_duration_seconds`, which are recorded at the same moment. On non-streamed traffic all three do cover the full request. For a streamed request's end-to-end latency use `aisix_request_e2e_latency_seconds`, which is recorded at stream completion. It is a histogram rather than a summary, so read it with the quantile queries above rather than per-instance. For dispatched A2A calls, the same histogram uses `endpoint="/a2a"` and records the agent-call lifetime. A pre-dispatch rejection that reaches A2A accounting records zero duration. For a streaming call, `aisix_proxy_request_duration_seconds` stops when the response starts, while the histogram continues until the stream finishes; an abandoned stream records `499` in the `4xx` status class. Do not average summary quantiles across instances. Use the histogram queries above to calculate percentiles across gateway instances. ### Measure Guardrail Latency and Outcomes[​](#measure-guardrail-latency-and-outcomes "Direct link to Measure Guardrail Latency and Outcomes") `aisix_guardrail_latency_seconds` records each timed guardrail execution, labeled with the guardrail name, kind, phase, and `result`. A guardrail normally runs once per applicable phase. Streaming window scans can run the same output guardrail several times for one response. Synchronous per-field PII and keyword masking operations are not included. Calculate P95 execution latency per guardrail to verify a moderation budget: ``` histogram_quantile( 0.95, sum by (le, guardrail) (rate(aisix_guardrail_latency_seconds_bucket[5m])) ) ``` Because guardrails run sequentially, their total contribution to request latency is the sum of their execution times for that request. Calculate the mean execution time for each guardrail and phase, and compare it with your moderation budget: ``` sum by (guardrail, phase) (rate(aisix_guardrail_latency_seconds_sum[5m])) / sum by (guardrail, phase) (rate(aisix_guardrail_latency_seconds_count[5m])) ``` Compare local detection with remote moderation services using the `kind` label. `keyword` and `pii` run in-process; every other kind invokes a remote service: ``` histogram_quantile( 0.95, sum by (le, kind) (rate(aisix_guardrail_latency_seconds_bucket[5m])) ) ``` The `_count` series doubles as an execution counter. Track block and fail-open bypass rates per guardrail, and alert when a remote guardrail starts failing open: ``` sum by (guardrail, result) (rate(aisix_guardrail_latency_seconds_count[5m])) # Fail-open bypasses by failure cause sum by (guardrail, error_type) (rate(aisix_guardrail_latency_seconds_count{result="bypassed"}[5m])) ``` ### Calculate Token Volume by Client[​](#calculate-token-volume-by-client "Direct link to Calculate Token Volume by Client") Use `token_type="total"` to calculate cache-inclusive token volume by normalized client type: ``` sum by (client_type) (rate(aisix_llm_tokens_by_client_total{token_type="total"}[5m])) ``` To compare input and output volume, select both token types and include `token_type` in the grouping: ``` sum by (client_type, token_type) (rate(aisix_llm_tokens_by_client_total{token_type=~"input|output"}[5m])) ``` ### Break Down Token Volume by Client Type and Model[​](#break-down-token-volume-by-client-type-and-model "Direct link to Break Down Token Volume by Client Type and Model") The `model` label records the model name the client requested, not the direct model AISIX selected. Group by `client_type` and `model` to see how each normalized client type distributes token volume across models: ``` sum by (client_type, model) (rate(aisix_llm_tokens_by_client_total{token_type="total"}[5m])) ``` To inspect one client type, filter on `client_type` and group by model only: ``` sum by (model) (rate(aisix_llm_tokens_by_client_total{client_type="claude-code", token_type="total"}[5m])) ``` ### Calculate Prompt Cache Hit Rate[​](#calculate-prompt-cache-hit-rate "Direct link to Calculate Prompt Cache Hit Rate") Upstream providers cache prompt prefixes and bill cached input well below the standard rate, so the share of input served from that cache is the single number that says whether prompt caching is paying off. AISIX reports it in three counters, and the split between them is what makes the query correct across providers: | Counter | Reported by the upstream as | Relationship to `aisix_llm_input_tokens_total` | | --------------------------------------------- | --------------------------------- | ---------------------------------------------- | | `aisix_llm_cached_input_tokens_total` | part of the prompt token count | already inside it | | `aisix_llm_cache_read_input_tokens_total` | a counter beside the input tokens | outside it | | `aisix_llm_cache_creation_input_tokens_total` | a counter beside the input tokens | outside it | So the input a request really consumed is `input + cache_read + cache_creation`, and the part of it that came from cache is `cached + cache_read`. Both expressions hold whichever shape the provider reports, which makes this hit rate comparable across a mixed fleet: ``` ( sum(rate(aisix_llm_cached_input_tokens_total[5m])) + sum(rate(aisix_llm_cache_read_input_tokens_total[5m])) ) / ( sum(rate(aisix_llm_input_tokens_total[5m])) + sum(rate(aisix_llm_cache_read_input_tokens_total[5m])) + sum(rate(aisix_llm_cache_creation_input_tokens_total[5m])) ) ``` Add `by (model)` to every `sum` to find which workloads benefit. A prompt with a stable prefix — a long system prompt, a fixed tool schema, a pinned document — should settle at a high rate; one that rewrites its opening on every call stays near zero and is the candidate to restructure. A cache write costs more than an ordinary input token while a cache read costs much less, so a workload that writes far more than it reads is paying a premium for a cache nobody hits. Watch the ratio for the providers that report both: ``` sum by (model) (rate(aisix_llm_cache_creation_input_tokens_total{upstream_protocol=~"anthropic|bedrock"}[5m])) / sum by (model) (rate(aisix_llm_cache_read_input_tokens_total{upstream_protocol=~"anthropic|bedrock"}[5m])) ``` A ratio well above 1 sustained over a long window means the cache entries expire before they are reused. Shorten the interval between calls that share a prefix, or stop caching that prompt. Each counter is created only once its value is non-zero, so a provider that reports no cache detail contributes no series and `sum()` over it returns nothing rather than zero. Wrap a query in `or vector(0)` when an alert must evaluate before any cached traffic has been seen. These counters describe the **provider's** prompt cache. The AISIX response cache is a separate mechanism measured by `aisix_cache_requests_total` — see [Measure Response Cache Effectiveness](#measure-response-cache-effectiveness). ### Calculate Spend and Cost per Request[​](#calculate-spend-and-cost-per-request "Direct link to Calculate Spend and Cost per Request") `aisix_llm_spend_micro_usd_total` counts micro-USD, where 1 USD is 1,000,000 micro-USD. Divide by `1e6` for a figure in dollars: ``` # USD per hour, by team sum by (team_id) (rate(aisix_llm_spend_micro_usd_total[5m])) * 3600 / 1e6 # Projected 30-day spend at the last day's rate sum(rate(aisix_llm_spend_micro_usd_total[24h])) * 86400 * 30 / 1e6 ``` Because the spend counter shares its labels with the request counter, average cost per request is a direct division. Group both sides by the same labels: ``` sum by (model) (rate(aisix_llm_spend_micro_usd_total[5m])) / 1e6 / sum by (model) (rate(aisix_llm_requests_total[5m])) ``` Spend is recorded only where AISIX resolves a price for the request, while the request counter counts every request. On a deployment where some models carry no configured cost, the denominator includes their traffic and the average reads low. Filter both sides to the priced models, or compare each model separately, rather than reading one fleet-wide number. ### Calculate Tokens per Request[​](#calculate-tokens-per-request "Direct link to Calculate Tokens per Request") ``` sum by (endpoint, model) (rate(aisix_llm_total_tokens_total[5m])) / sum by (endpoint, model) (rate(aisix_llm_requests_total[5m])) ``` Keep `endpoint` in the grouping. `/v1/audio/speech` is billed per input character and `/v1/videos` per video, so both count as requests and report no tokens; summing across endpoints pulls the average down by however much of that traffic there is. `aisix_llm_total_tokens_total` is cache-inclusive: it contains input, output, and the cache-read and cache-creation tokens that providers report beside their input count. That makes it the right numerator for a per-request consumption figure and the wrong one for a cache hit rate — use the query in [Calculate Prompt Cache Hit Rate](#calculate-prompt-cache-hit-rate) for that. ### Rank the Heaviest Consumers[​](#rank-the-heaviest-consumers "Direct link to Rank the Heaviest Consumers") The token and spend counters carry the full caller identity, so `topk` answers who to talk to about a bill: ``` # Ten API keys with the highest token rate topk(10, sum by (api_key_id) (rate(aisix_llm_total_tokens_total[1h]))) # Ten users with the highest spend rate, in USD per hour topk(10, sum by (user_id, user_name) (rate(aisix_llm_spend_micro_usd_total[1h])) * 3600 / 1e6) # Which models one team spends on sum by (model) (rate(aisix_llm_spend_micro_usd_total{team_id="team-platform"}[1h])) * 3600 / 1e6 ``` `user_name` and `provider_key_name` are one-to-one with their IDs, so including a name in the grouping adds no series. `user_name` reports `unknown` until the control plane supplies it. ### Measure Response Cache Effectiveness[​](#measure-response-cache-effectiveness "Direct link to Measure Response Cache Effectiveness") `aisix_cache_requests_total` counts requests that reached an enabled cache policy with an available backend. Requests with no matching policy are never counted, so this measures how well the configured policies work, not what share of all traffic is cached: ``` sum by (policy) (rate(aisix_cache_requests_total{outcome=~"hit_exact|hit_semantic"}[5m])) / sum by (policy) (rate(aisix_cache_requests_total[5m])) ``` Split by `outcome` to see how much the semantic layer adds over exact matching, which is what justifies the embedding call it costs: ``` sum by (policy, outcome) (rate(aisix_cache_requests_total[5m])) ``` A semantic-layer failure degrades to an ordinary miss, so a broken embedding model or store looks exactly like a healthy low hit rate in the outcome counter. Alert on the failure counters to tell them apart: ``` sum by (policy, cause) (rate(aisix_cache_semantic_embedding_failures_total[5m])) > 0 sum by (policy, op) (rate(aisix_cache_semantic_store_failures_total[5m])) > 0 ``` This cache and the upstream provider's prompt cache are independent, and they interact in one direction worth knowing: a request served from the AISIX response cache never reaches the upstream, so it contributes nothing to the provider cache counters. A rising response-cache hit rate therefore lowers the absolute prompt-cache token rates without meaning that prompt caching got worse. ### Measure Target Health and Fallback[​](#measure-target-health-and-fallback "Direct link to Measure Target Health and Fallback") Request metrics say what callers experienced; deployment metrics say which target was responsible. A Model Group whose success rate looks healthy can be hiding one consistently failing target that fallback keeps rescuing. Compare the two: ``` # Failure rate per attempted target sum by (model, upstream_model, provider_key_id) (rate(aisix_deployment_failure_responses_total[5m])) / sum by (model, upstream_model, provider_key_id) (rate(aisix_deployment_requests_total[5m])) # Share of client requests that a fallback target served sum(rate(aisix_llm_requests_total{is_fallback="true"}[5m])) / sum(rate(aisix_llm_requests_total[5m])) ``` A rising fallback share with a flat success rate is the signal to act on: callers are still being served, and the primary target is degrading behind that. ``` # Targets currently out of rotation count by (model) (aisix_deployment_state == 2) # How often targets enter cooldown sum by (model, upstream_model) (rate(aisix_deployment_cooled_down_total[15m])) ``` Deployment counters sample once per upstream **attempt**, so one client request that failed over across three targets is three samples here and one in the request family. Do not divide one family by the other. ### Monitor Rate Limits and Diagnose Budget Denials[​](#monitor-rate-limits-and-diagnose-budget-denials "Direct link to Monitor Rate Limits and Diagnose Budget Denials") The rate-limit metrics below apply to every AISIX gateway. AISIX Cloud budget gauges describe budget denials; they are not a continuous feed of current spend. The share of requests a rate limit rejected comes from the request family's `outcome` label, which covers every endpoint: ``` sum(rate(aisix_llm_requests_total{outcome="rate_limited"}[5m])) / sum(rate(aisix_llm_requests_total[5m])) ``` To tell a request-count limit from a token limit, use the dedicated counter's `scope` label: ``` sum by (scope) (rate(aisix_ratelimit_rejections_total[5m])) ``` When AISIX Cloud rejects a request because a blocking budget was exceeded, the budget gauges record the denial details. They carry the same `api_key_id`, `team_id`, `user_id`, and `user_name` labels, so you can inspect the reported spend against the limit: ``` # Reported spend relative to the limit that blocked the request (aisix_budget_spent_usd / aisix_budget_limit_usd) and on (api_key_id, team_id, user_id) (aisix_budget_details_present == 1) ``` Guard on `aisix_budget_details_present == 1`. A later decision without budget details does **not** zero the other gauges — they stay at the amounts from the last denial, so an unguarded ratio can report stale values. The flag is what tells the two apart. These gauges cannot warn before AISIX Cloud blocks a request because the denial is what supplies their values. Configure [Budget Alerts and Notifications](https://docs.api7.ai/ai-gateway/traffic-controls/budget-alerts.md) for threshold notifications before a blocking budget starts rejecting traffic. #### When a Key Goes Away[​](#when-a-key-goes-away "Direct link to When a Key Goes Away") `aisix_budget_*` and `aisix_ratelimit_remaining_*` are written only while a request is being served, and a Prometheus series is never dropped once it exists. AISIX therefore retires them: within seconds of an API key being deleted — or rebound to another team or member, which starts a new series and strands the old one — the stale sample is rewritten as `NaN`, and `aisix_budget_details_present` as `0`. `NaN` rather than `0` is deliberate. `aisix_ratelimit_remaining_requests 0` means the caller is out of quota and `aisix_budget_remaining_usd 0` means the budget is spent, so zeroing a retired series would replace a stale reading with a false one. Every comparison against `NaN` is false, which is what makes an alert on a retired series stop firing rather than latch. Two consequences to plan for: * The series still exists, so `absent()` does not report it missing. * An aggregation over a family that contains a retired series reads `NaN` — `sum(aisix_budget_spent_usd)` returns `NaN`, not the live total. Group by an identifying label, or filter on `aisix_budget_details_present == 1` first. The per-series queries above are unaffected. `aisix_deployment_state` is not retired. It is written only when a deployment's health changes, so blanking it would leave a live model reporting no state at all until its next transition. A model deleted while it was cooling down therefore stays counted by `count by (model) (aisix_deployment_state == 2)` until the gateway restarts. ### Measure Concurrency and Abandoned Requests[​](#measure-concurrency-and-abandoned-requests "Direct link to Measure Concurrency and Abandoned Requests") `aisix_proxy_in_flight_requests` is the live concurrency at the proxy, which is what sizes a gateway rather than the request rate: ``` sum by (endpoint, inbound_protocol) (aisix_proxy_in_flight_requests) ``` A caller that disconnects before response headers were sent never reaches a normal outcome and is absent from the request counters, so it has to be added back into the denominator to get an abandonment rate: ``` sum(rate(aisix_proxy_client_cancelled_requests_total[5m])) / ( sum(rate(aisix_proxy_requests_total[5m])) + sum(rate(aisix_proxy_client_cancelled_requests_total[5m])) ) ``` Abandonment usually means callers give up while waiting for the first token. Break it down by `model` and compare against time to first token for that same model: ``` sum by (model) (rate(aisix_proxy_client_cancelled_requests_total[5m])) ``` Denied authentication is the other silent failure mode — a rotated key or a misconfigured issuer shows up here long before anyone reports it: ``` sum by (method, reason) (rate(aisix_auth_decisions_total{result="denied"}[5m])) ``` ### Verify Usage Event Delivery[​](#verify-usage-event-delivery "Direct link to Verify Usage Event Delivery") Usage events feed billing and analytics, so a query that silently loses them is worth an alert of its own. Every emission attempt is either accepted by the delivery queue or counted as a drop: ``` sum(rate(aisix_usage_event_drops_total[5m])) / sum(rate(aisix_usage_events_emitted_total[5m])) ``` Both counters carry `model`, `provider_key_id`, and `provider_key_name`, so the same division holds per model and per provider key. A drop rate concentrated on one of them points at the traffic whose usage records went missing: ``` sum by (model, provider_key_name, reason) (rate(aisix_usage_event_drops_total[5m])) ``` Queue acceptance is not delivery to storage. Control-plane shipping and exporter pipelines can still fail after the worker accepts an event, and neither increments the drop counter — those failures appear in the gateway log. ### Alert on Configuration Problems[​](#alert-on-configuration-problems "Direct link to Alert on Configuration Problems") A gateway that cannot load its configuration keeps serving the last good one, which means the symptom is silence rather than an error. Alert on the state directly: ``` # The most recent load failed aisix_config_last_reload_successful == 0 # Nothing has loaded successfully for five minutes time() - aisix_config_last_reload_success_timestamp_seconds > 300 # Resources the gateway refused, by kind sum by (kind) (aisix_config_rejected_resources) > 0 # Serving an older etcd revision than the one observed aisix_config_observed_revision - aisix_config_applied_revision > 0 ``` `aisix_config_partially_compatible_resources` is the upgrade-order signal: it counts resources carrying fields this gateway version does not recognize, which is what a control plane one release ahead of its data planes produces. It is expected during an upgrade window and should return to zero once the data planes are upgraded: ``` sum by (kind) (aisix_config_partially_compatible_resources) > 0 ``` Inspect `GET /status/config` on the metrics listener for the specific resources and field paths behind any of these. ## Configure Metrics[​](#configure-metrics "Direct link to Configure Metrics") Configure custom client classification and histogram bucket boundaries at startup. Changes take effect after a gateway restart and can alter label values or bucket series, so coordinate them with dashboards, alerts, and recording rules. ### Map Custom Clients to a Client Type[​](#map-custom-clients-to-a-client-type "Direct link to Map Custom Clients to a Client Type") AISIX recognizes common AI coding clients and SDKs out of the box. To classify in-house tools — or re-bucket a client that the built-in rules place in a generic bucket such as `node` — define mapping rules in the gateway configuration: config.yaml ``` observability: metrics: client_type_rules: - pattern: "^billing-batcher/" client: billing-batcher - pattern: "internal-eval-harness" client: eval-harness ``` Each rule maps a regular expression to a fixed `client` value, which AISIX emits as the `client_type` label. Rules are evaluated in order against the raw `User-Agent` header, before the built-in rules, and the first match wins. Matching is case-insensitive and unanchored, so anchor the pattern with `^` when you need a prefix match. The `client` value — never the request's `User-Agent` — becomes the label, so the label set stays bounded regardless of what clients send. Configuration limits keep it that way: at most 64 rules, patterns up to 512 bytes, and `client` values of up to 64 characters matching `[a-z0-9][a-z0-9._-]*`. AISIX validates the rules at startup and refuses to start on an invalid rule; changes take effect after a restart. Requests with an empty `User-Agent` always report `unknown`, and requests that match no rule fall through to the built-in classification. ### Customize Histogram Buckets[​](#customize-histogram-buckets "Direct link to Customize Histogram Buckets") Four metrics are true Prometheus histograms with `le` bucket boundaries, and each one ships its own default boundaries because the distributions differ: | Metric | Default boundaries (seconds) | | ----------------------------------- | --------------------------------------------------------------------------------- | | `aisix_request_e2e_latency_seconds` | 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10, 30, 60, 120, 300, 420, 600 | | `aisix_request_ttft_seconds` | 0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10, 30, 60, 120, 300 | | `aisix_guardrail_latency_seconds` | 0.001, 0.0025, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30 | | `aisix_a2a_ttfb_seconds` | 0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10, 30, 60, 120, 300 | End-to-end latency includes cache hits and requests rejected before dispatch, so it needs millisecond boundaries. Time to first token (TTFT) is recorded at the first frame an upstream streams, whatever it carries; against a hosted provider, sub-50 ms boundaries are usually empty. Guardrail metrics need low boundaries for fast in-process checks and high boundaries for remote moderation services. An A2A agent's time to its first streamed event starts from the same defaults as TTFT. It takes its own `a2a_ttfb` override, because an agent task can think for minutes before it says anything. Override the boundaries per metric when your traffic has a different distribution. For example, a vLLM or Ollama model server on the same node can produce its first output in milliseconds: config.yaml ``` observability: metrics: buckets: request_ttft: [0.005, 0.01, 0.025, 0.05, 0.1, 0.5, 1, 5, 30] ``` Each field is optional and replaces only the metric it names; omitted metrics keep their defaults. A supplied list must contain 1–64 finite, positive, strictly increasing boundaries. Do not list `+Inf` — AISIX appends that bucket itself. AISIX validates the configuration at startup and refuses to start on an invalid list; changes take effect after a restart. Every boundary adds one `_bucket` time series for each label combination, so a longer list increases the number of series a collector stores. For a gateway deployed with the AISIX Helm chart, provide the comma-separated list through `extraEnvVars`: ``` extraEnvVars: - name: AISIX_OBSERVABILITY__METRICS__BUCKETS__REQUEST_TTFT value: "0.005,0.01,0.025,0.05,0.1,0.5,1,5,30" ``` caution Changing boundaries changes the emitted `_bucket` series. A dashboard or recording rule that selects a removed `le` value stops matching. Do not compare bucket-derived quantiles from before and after the change. See [Set Any Other Gateway Configuration](https://docs.api7.ai/ai-gateway/cloud/kubernetes.md#set-any-other-gateway-configuration) for the Helm configuration pattern. --- # On-Premises Configuration Configure the On-Premises deployment through a Docker Compose `.env` file or Helm values, depending on how you install the control plane. Use this reference with [On-Premises Installation](https://docs.api7.ai/ai-gateway/on-premises/deployment.md) when you need to review or customize the generated deployment configuration. These settings configure the on-premises control plane package. They are separate from AISIX gateway runtime environment variables. For gateway runtime variables, see [Environment Variables](https://docs.api7.ai/ai-gateway/reference/environment-variables.md). ## Docker Compose Environment Variables[​](#docker-compose-environment-variables "Direct link to Docker Compose Environment Variables") The Docker Compose package reads environment variables from `./aisix-self-hosted/.env`. The quickstart and offline package generate this file on first start. Back it up with the database: it contains the database password and master key, and it is not included in later package archives. During a normal in-place upgrade, `run.sh` updates `AISIX_VERSION` to match the extracted package. It preserves the other settings and secrets in `.env`. Run the Docker Compose commands on this page from the `./aisix-self-hosted` directory. ### Images and Release Version[​](#images-and-release-version "Direct link to Images and Release Version") | Variable | Purpose | | ---------------------- | -------------------------------------------------------------------------------------------------------------------------------- | | `AISIX_VERSION` | Package-owned release tag used by the images unless you set an individual image override. `run.sh` updates it during an upgrade. | | `AISIX_API_IMAGE` | Optional image override for `cp-api`. | | `AISIX_DPM_IMAGE` | Optional image override for `dp-manager`. | | `AISIX_UI_IMAGE` | Optional image override for the dashboard. | | `AISIX_CLOUD_DP_IMAGE` | Optional AISIX gateway image override used in generated gateway install snippets. | Individual image overrides remain in effect across package upgrades. ### Database[​](#database "Direct link to Database") | Variable | Purpose | | ------------------- | ------------------------------------------------------------------------------------------------------------------------ | | `POSTGRES_USER` | PostgreSQL user for the bundled database. | | `POSTGRES_PASSWORD` | PostgreSQL password for the bundled database. Use a strong URL-safe value because it is embedded in a `postgres://` URL. | | `POSTGRES_DB` | PostgreSQL database name. | These variables configure only the bundled PostgreSQL service. The package does not expose external database URL overrides in `.env`. Use the [Helm installation](https://docs.api7.ai/ai-gateway/on-premises/deployment.md#helm-on-kubernetes) when the control plane must connect to an external PostgreSQL database. ### Secrets[​](#secrets "Direct link to Secrets") | Variable | Purpose | | --------------------------- | --------------------------------------------------------------------------------------------------------- | | `AISIX_CLOUD_MASTER_KEY` | Base64-encoded 32-byte AES key used for envelope encryption. The same value is also used by `dp-manager`. | | `AISIX_CLOUD_MASTER_KEY_ID` | Identifier stored with encrypted rows so the control plane can identify the wrapping key. | | `BETTER_AUTH_SECRET` | Session-signing secret for dashboard authentication. | Do not change `AISIX_CLOUD_MASTER_KEY` on an existing deployment unless you are following a master-key rotation procedure. Changing it without preserving the previous key can make encrypted data unreadable. ### Recover a Missing CA Root[​](#recover-a-missing-ca-root "Direct link to Recover a Missing CA Root") caution Use this recovery override only when `cp-api` or `dp-manager` refuses to start with `ca_root row missing but dp_certificates has rows`, and only after confirming that the control plane is connected to the intended database. Minting a new root CA invalidates every existing gateway mTLS chain. It does not recover the missing root. If you cannot restore the missing root from a complete database backup and intend to replace it, add the following setting to `.env`: ``` AISIX_CLOUD_ALLOW_FRESH_BOOTSTRAP=1 ``` Recreate both services so either process can bootstrap the replacement CA: ``` docker compose up -d --wait api dpm ``` After both services are healthy, remove `AISIX_CLOUD_ALLOW_FRESH_BOOTSTRAP` from `.env` and run the same command again to recreate them without the override. When both services are healthy again, [re-enroll every AISIX gateway](https://docs.api7.ai/ai-gateway/cloud/connect-a-gateway.md); certificates issued under the missing root no longer authenticate. ### Runtime URLs[​](#runtime-urls "Direct link to Runtime URLs") | Variable | Purpose | | ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | `AISIX_CLOUD_PUBLIC_BASE_URL` | Browser-facing control-plane origin, such as `https://aisix.example.com`. Login validates the session issuer against this value. | | `AISIX_CLOUD_DPMGR_BASE_URL` | `dp-manager` mTLS endpoint that AISIX gateway hosts can reach. A DNS name or an IP address. | | `AISIX_CLOUD_DASHBOARD_URL` | Internal dashboard URL used by `cp-api`. The Compose default points to the dashboard service. | | `AISIX_TRUSTED_ORIGINS` | Additional browser origins allowed to sign in, provided as a comma-separated list. The public base URL and its loopback twin are trusted automatically. | Set `AISIX_CLOUD_PUBLIC_BASE_URL` and `AISIX_CLOUD_DPMGR_BASE_URL` before exposing the deployment outside the local host or container network. Compose passes `AISIX_CLOUD_DPMGR_BASE_URL` to the `dpm` service as well, because `dp-manager` issues its TLS server certificate for that host. Recreate both `api` and `dpm` after changing it. ### Pricing Catalog[​](#pricing-catalog "Direct link to Pricing Catalog") The packaged Docker Compose file runs `cp-api` in offline pricing mode by default. It sets `AISIX_CLOUD_PRICESYNC_SNAPSHOT_PATH` on the `api` service to a snapshot baked into the `cp-api` image, so the control plane can initialize model pricing without contacting `models.dev`. The package exposes the pricing mode through `.env`: | Variable | Purpose | | ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `AISIX_CLOUD_PRICESYNC_SNAPSHOT_PATH` | Select offline mode when it contains a snapshot path. When the variable is absent, Compose uses the in-image snapshot. An explicitly empty value selects online mode. | | `AISIX_CLOUD_PRICESYNC_URL` | Override the `models.dev` catalog URL in online mode, for example with an internal mirror. When it is empty, online mode uses `https://models.dev/api.json`. | To use online pricing, add an empty snapshot-path assignment to `.env`. Add the URL only when you need to use a different catalog endpoint: ``` AISIX_CLOUD_PRICESYNC_SNAPSHOT_PATH= ``` Recreate `cp-api` after changing either setting: ``` docker compose up -d api ``` Keep these overrides in `.env`, not `docker-compose.yaml`, because extracting a later package replaces the Compose file. ### Private Network Access[​](#private-network-access "Direct link to Private Network Access") By default, `cp-api` refuses these outbound connections when the destination resolves to a private, internal, or loopback address. Enable only the access required for services on networks that you trust and that `cp-api` can reach. | Variable | Purpose | | ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `AISIX_PLAYGROUND_ALLOW_PRIVATE_IPS` | Set to `1` to let the dashboard playground reach LLM endpoints on private, internal, or loopback addresses. Off by default as an SSRF protection; enable it only for self-hosted models on an internal network. See [Playground](https://docs.api7.ai/ai-gateway/cloud/playground.md#reach-private-network-endpoints-on-premises). | | `AISIX_CLOUD_NOTIFY_ALLOW_PRIVATE_URLS` | Set to `true` to let budget notifications reach a private webhook receiver or Slack proxy. | | `AISIX_CLOUD_MCP_SPEC_ALLOW_PRIVATE_URLS` | Set to `true` to let the control plane fetch an MCP server's OpenAPI document from a private `spec_url`. If the document is on a private URL and you leave this setting disabled, provide `spec_content` instead. | Set the required variables in `.env`, then recreate `cp-api`: ``` docker compose up -d api ``` ### Dashboard and Ports[​](#dashboard-and-ports "Direct link to Dashboard and Ports") | Variable | Purpose | | ------------------------ | -------------------------------------------------------------------------- | | `AISIX_DASHBOARD_LOCALE` | Dashboard language for the deployment. Supported values are `en` and `zh`. | | `POSTGRES_HOST_PORT` | Host port binding for bundled PostgreSQL. | | `API_HOST_PORT` | Host port binding for `cp-api` and the dashboard reverse proxy. | | `DPM_HOST_PORT` | Host port binding for `dp-manager`. | Prefix a host port with `127.0.0.1:` when the service should bind only to loopback. ## Helm Values[​](#helm-values "Direct link to Helm Values") The `api7/aisix-cp` chart uses Helm values instead of a Compose `.env` file. Add the API7 Helm repository before inspecting or installing the chart: ``` helm repo add api7 https://charts.api7.ai helm repo update ``` To inspect every chart value, run: ``` helm show values api7/aisix-cp --version 1.0.0 ``` The chart source and values are published in the [`aisix-cp-1.0.0` Helm chart release](https://github.com/api7/api7-helm-chart/releases/tag/aisix-cp-1.0.0). ### Images and Services[​](#images-and-services "Direct link to Images and Services") | Value | Purpose | | --------------------------------------------------------- | ------------------------------------------------------------------------------------ | | `api.image.repository`, `api.image.tag` | `cp-api` image. | | `dpm.image.repository`, `dpm.image.tag` | `dp-manager` image. | | `ui.image.repository`, `ui.image.tag` | Dashboard image. | | `api.replicaCount`, `dpm.replicaCount`, `ui.replicaCount` | Number of replicas for each control-plane component. | | `api.affinity`, `dpm.affinity`, `ui.affinity` | Kubernetes scheduling rules used to spread replicas across nodes or failure domains. | | `api.nodeSelector`, `dpm.nodeSelector`, `ui.nodeSelector` | Node label constraints for each control-plane component. | | `api.tolerations`, `dpm.tolerations`, `ui.tolerations` | Kubernetes taint exceptions for each control-plane component. | | `api.service.type`, `api.service.port` | Kubernetes Service settings for `cp-api`. | | `dpm.service.type`, `dpm.service.port` | Kubernetes Service settings for `dp-manager`. | | `ui.service.type`, `ui.service.port` | Kubernetes Service settings for the dashboard service behind `cp-api`. | ### Control Plane URLs[​](#control-plane-urls "Direct link to Control Plane URLs") | Value | Purpose | | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `api.publicBaseURL` | Browser-facing control-plane origin. | | `api.dpmgrBaseURL` | `dp-manager` mTLS endpoint that AISIX gateway hosts can reach. A DNS name or an IP address. The chart also passes it to the `dp-manager` deployment, which issues its TLS server certificate for that host. | | `api.dpImage` | AISIX gateway image shown in generated AISIX Cloud gateway install snippets. | ### Playground[​](#playground "Direct link to Playground") | Value | Purpose | | ------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | `api.playgroundAllowPrivateIPs` | Let the dashboard playground reach LLM endpoints on private, internal, or loopback addresses. `false` by default as an SSRF protection. | ### Notification Destinations[​](#notification-destinations "Direct link to Notification Destinations") | Value | Purpose | | ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------- | | `api.notifyAllowPrivateURLs` | Let `cp-api` send budget notifications to private, internal, or loopback addresses. `false` by default as an SSRF protection. | ### Secrets[​](#secrets-1 "Direct link to Secrets") | Value | Purpose | | -------------------------- | ----------------------------------------------------------------------------------------- | | `secrets.masterKey` | Base64-encoded 32-byte AES key used for envelope encryption. | | `secrets.masterKeyID` | Identifier stored with encrypted rows so the control plane can identify the wrapping key. | | `secrets.betterAuthSecret` | Session-signing secret for dashboard authentication. | Replace the chart's placeholder secrets before installing. The chart rejects placeholder secret values. ### PostgreSQL[​](#postgresql "Direct link to PostgreSQL") | Value | Purpose | | -------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `postgresql.builtin` | Deploy the bundled PostgreSQL chart when set to `true`. | | `postgresql.auth.password` | Password for the bundled chart's configured PostgreSQL user. The chart requires a non-placeholder value even when applications connect as the `postgres` user. | | `postgresql.auth.postgresPassword` | PostgreSQL superuser password. The chart uses this for control-plane connections by default. | | `postgresql.auth.usePostgresUserForAppConnections` | Use the `postgres` user for control-plane connections. The default is `true`. | | `postgresql.auth.existingSecret` | Existing Kubernetes Secret for bundled PostgreSQL credentials. | | `externalDatabase.*` | Top-level values for an existing PostgreSQL database when `postgresql.builtin=false`. | Use URL-safe PostgreSQL passwords, such as values generated with `openssl rand -hex 24`, because the chart builds a `postgres://` connection URL from the configured credentials. ### Dashboard Locale[​](#dashboard-locale "Direct link to Dashboard Locale") | Value