Skip to main content

Parameters

See plugin common configurations for configuration options available to all plugins.

  • fallback_strategy

    string or array

    vaild vaule:

    string: instance_health_and_rate_limiting, http_429, or http_5xx
    array: Any combination of rate_limiting, http_429, and http_5xx


    Fallback strategy. The option instance_health_and_rate_limiting is kept for backward compatibility and is functionally the same as rate_limiting.

    With rate_limiting or instance_health_and_rate_limiting, when the current instance's quota is exhausted, the request is forwarded to the next instance regardless of priority. With http_429, if an instance returns status code 429, the request is retried with other instances. With http_5xx, if an instance returns a 5xx status code, the request is retried with other instances. If all instances fail, the plugin returns the last error response code.

    When not set, the plugin will not forward the request to low priority instances when tokens of the high priority instance are exhausted.

  • max_retries

    integer

    vaild vaule:

    greater than or equal to 0


    Maximum number of fallback retries after the initial request fails. This bounds how many additional instances a single request tries, so it does not exhaust every configured instance. Only takes effect together with fallback_strategy. When not set, there is no explicit cap and the plugin retries until an instance succeeds or all instances have been tried. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0.

  • retry_on_failure_within_ms

    integer

    vaild vaule:

    greater than or equal to 1


    Only fall back to another instance when the upstream fails within this many milliseconds. Fast failures (such as connection errors and quick 429 or 5xx responses) are retried, while a slow failure that takes longer than this is returned to the client directly to avoid doubling the total wait time. Only takes effect together with fallback_strategy. When not set, the plugin retries regardless of how long the failed attempt took. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0.

  • balancer

    object


    Load balancing configurations.

    • algorithm

      string

      default: roundrobin

      vaild vaule:

      roundrobin or chash


      Load balancing algorithm. When set to roundrobin, weighted round robin algorithm is used. When set to chash, consistent hashing algorithm is used.

    • hash_on

      string

      default: vars

      vaild vaule:

      vars, header, cookie, consumer, or vars_combinations


      Used when type is chash. Support hashing on built-in variables, header, cookie, consumer, or a combination of built-in variables.

    • key

      string


      Used when type is chash. When hash_on is set to header or cookie, key is required. When hash_on is set to consumer, key is not required as the consumer name will be used as the key automatically.

  • instances

    array[object]

    required


    LLM instance configurations.

    • name

      string

      required


      Name of the LLM service instance.

    • provider

      string

      required

      vaild vaule:

      openai, deepseek, azure-openai, aimlapi, gemini, vertex-ai, anthropic, openrouter, bedrock, openai-compatible


      LLM service provider.

      When set to openai, the plugin will proxy requests to https://api.openai.com/v1/chat/completions.

      When set to deepseek, the plugin will proxy requests to https://api.deepseek.com/chat/completions.

      When set to gemini (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests to https://generativelanguage.googleapis.com/v1beta/openai/chat/completions. If you are proxying requests to an embedding model, you should configure the embedding model endpoint in the override.

      When set to vertex-ai (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin proxies requests to Google Cloud Vertex AI. For chat completions, the plugin will proxy requests to https://{region}-aiplatform.googleapis.com/v1beta1/projects/{project_id}/locations/{region}/endpoints/openapi/chat/completions. For embeddings, the plugin will proxy requests to https://{region}-aiplatform.googleapis.com/v1/projects/{project_id}/locations/{region}/publishers/google/models/{model}:predict. These require configuring provider_conf with project_id and region. Alternatively, you can configure override for a custom endpoint.

      When set to anthropic (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests to https://api.anthropic.com/v1/chat/completions.

      When set to openrouter (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests to https://openrouter.ai/api/v1/chat/completions.

      When set to bedrock (available from API7 Enterprise 3.9.12 and APISIX 3.17.0), the plugin proxies requests to AWS Bedrock using the Converse API.

      When set to aimlapi (available from APISIX 3.14.0 and Enterprise 3.8.17), the plugin uses the OpenAI-compatible driver and proxies the request to https://api.aimlapi.com/v1/chat/completions.

      When set to openai-compatible, the plugin proxies requests to the custom endpoint configured in override.

      When set to azure-openai, the plugin also proxies requests to the custom endpoint configured in override and additionally removes the model parameter from user requests.

    • priority

      integer

      default: 0


      Priority of the LLM instance in load balancing. priority takes precedence over weight.

    • weight

      integer

      required

      vaild vaule:

      greater than or equal to 0


      Weight of the LLM instance in load balancing.

    • auth

      object

      required


      Authentication configurations.

      • header

        object


        Authentication headers. At least one of the header and query should be configured. You can configure additional custom headers that will be forwarded to the upstream LLM service.

      • query

        object


        Authentication query parameters. At least one of the header and query should be configured.

      • gcp

        object


        GCP service account authentication for Vertex AI. Available in API7 Enterprise from 3.9.2 and APISIX from version 3.17.0.

        • service_account_json

          string


          GCP service account JSON content used for authentication. This can be configured using this parameter or by setting the GCP_SERVICE_ACCOUNT environment variable.

        • max_ttl

          integer


          Maximum TTL for GCP access token caching, in seconds.

        • expire_early_secs

          integer

          default: 60


          Number of seconds to expire the access token before its actual expiration time. This prevents edge cases where tokens expire during active requests.

      • aws

        object


        AWS IAM credentials for SigV4 signing. Required when provider is bedrock (for Bedrock, auth.aws is sufficient and auth.header/auth.query are not required). Available in API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0.

        • access_key_id

          string

          required


          AWS IAM access key ID.

        • secret_access_key

          string

          required


          AWS IAM secret access key.

        • session_token

          string


          AWS session token for temporary credentials (e.g. from STS AssumeRole).

    • options

      object


      Model configurations.

      In addition to model, you can configure additional parameters and they will be forwarded to the upstream LLM service in the request body. For instance, if you are working with OpenAI or DeepSeek, you can configure additional parameters such as max_tokens, temperature, top_p, and stream. See your LLM provider's API documentation for more available options.

      • model

        string


        Name of the LLM model, such as gpt-4 or gpt-3.5. See your LLM provider's API documentation for more available models.

    • provider_conf

      object


      Provider-specific configuration. Required when provider is bedrock. When provider is vertex-ai, configure either provider_conf or override.endpoint.

      Available in API7 Enterprise from 3.9.2 and APISIX from version 3.17.0.

      • project_id

        string


        Google Cloud Project ID for Vertex AI.

      • region

        string

        required


        Cloud region. For vertex-ai, this is the GCP region. For bedrock, this is the AWS region (e.g. us-east-1).

    • override

      object


      Override setting.

      • endpoint

        string


        LLM provider endpoint to replace the default endpoint with. If not configured, the plugin uses the default OpenAI endpoint https://api.openai.com/v1/chat/completions.

      • llm_options

        object


        Provider-aware LLM option overrides. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.

        • max_tokens

          integer


          Maximum number of output tokens. The gateway automatically maps this to the correct field name for the target provider, such as max_completion_tokens for OpenAI Chat or max_output_tokens for OpenAI Responses API, and overwrites the client value.

      • request_body

        object


        Per target-protocol request body overrides. Keys are target protocol names, such as openai-chat, openai-responses, openai-embeddings, anthropic-messages, bedrock-converse, and passthrough. Values are partial request bodies that are deep-merged into the outgoing body. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.

      • request_body_force_override

        boolean

        default: false


        When false (default), client request body fields take priority and request_body override values only fill in missing fields. When true, request_body override values overwrite client fields. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.

    • checks

      object


      Health check configurations.

      Note that at the moment, OpenAI and DeepSeek do not provide an official health check endpoint. Other LLM services that you can configure under openai-compatible provider may have available health check endpoints.

      • active

        object

        required


        Active health check configurations.

        • type

          string

          default: http

          vaild vaule:

          http, https, or tcp


          Type of health check connection.

        • timeout

          number

          default: 1


          Health check timeout in seconds.

        • concurrency

          integer

          default: 10


          Number of upstream nodes to be checked at the same time.

        • host

          string


          HTTP host.

        • port

          integer

          vaild vaule:

          between 1 and 65535 inclusive


          HTTP port.

        • http_path

          string

          default: /


          Path for HTTP probing requests.

        • http_method

          string

          default: GET

          vaild vaule:

          CONNECT, DELETE, GET, HEAD, OPTIONS, PATCH, POST, PURGE, PUT, or TRACE


          HTTP method for active health check probing requests. Available in API7 Enterprise and not available in APISIX yet.

        • http_req_body

          string


          Request body to send in active health check probing requests. This is useful when http_method is set to POST. Defaults to empty string. Available in API7 Enterprise and not available in APISIX yet.

        • https_verify_certificate

          boolean

          default: true


          If true, verify the node's TLS certificate.

        • healthy

          object


          Healthy check configurations.

          • interval

            integer

            default: 1


            Time interval of checking healthy nodes, in seconds.

          • http_statuses

            array[integer]

            default: [200,302]

            vaild vaule:

            status code between 200 and 599 inclusive


            An array of HTTP status codes that defines a healthy node.

          • successes

            integer

            default: 2

            vaild vaule:

            between 1 and 254 inclusive


            Number of successful probes to define a healthy node.

        • req_headers

          array[string]


          List of additional HTTP headers to send in health check probing requests, in "Header: Value" format.

        • unhealthy

          object


          Unhealthy check configurations.

          • interval

            integer

            default: 1


            Time interval of checking unhealthy nodes, in seconds.

          • http_statuses

            array[integer]

            default: [429,404,500,501,502,503,504,505]

            vaild vaule:

            status code between 200 and 599 inclusive


            An array of HTTP status codes that defines an unhealthy node.

          • http_failures

            integer

            default: 5

            vaild vaule:

            between 1 and 254 inclusive


            Number of HTTP failures to define an unhealthy node.

          • tcp_failures

            integer

            default: 2

            vaild vaule:

            between 1 and 254 inclusive


            Number of TCP failures to define an unhealthy node.

          • timeouts

            integer

            default: 3

            vaild vaule:

            between 1 and 254 inclusive


            Number of probe timeouts to define an unhealthy node.

  • logging

    object


    Logging configurations. These configurations apply to access logs and logs sent to logging plugins, and do not affect the error log.

    • summaries

      boolean

      default: false


      If true, log request LLM model, duration, request and response tokens.

    • payloads

      boolean

      default: false


      If true, log request and response payload.

  • timeout

    integer

    default: 30000

    vaild vaule:

    between 1 and 600000 inclusive


    Request timeout in milliseconds when requesting the LLM service.

  • max_req_body_size

    integer

    default: 67108864

    vaild vaule:

    greater than or equal to 1


    Maximum request body size in bytes that the plugin reads into memory. Larger requests are rejected with HTTP 413. This prevents unbounded memory buffering of large request bodies. The default is 67108864 bytes (64 MB). Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0.

  • max_stream_duration_ms

    integer

    vaild vaule:

    greater than or equal to 1


    Maximum wall-clock duration, in milliseconds, for a streaming AI response. If the upstream keeps sending data past this deadline, the gateway closes the connection. When the limit is reached mid-stream, the downstream stream is truncated without a protocol-specific terminator such as [DONE], message_stop, or response.completed. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.

  • max_response_bytes

    integer

    vaild vaule:

    greater than or equal to 1


    Maximum total bytes read from the upstream for a single AI response, including streaming and non-streaming responses. If the response exceeds this value, the gateway closes the connection. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.

  • streaming_flush_interval_ms

    integer

    default: 10

    vaild vaule:

    greater than or equal to 0


    Background flush interval in milliseconds for streaming responses. A positive value periodically flushes buffered output to bound client latency when the upstream sends tokens in bursts. Set to 0 to flush each chunk synchronously. Available in API7 Enterprise from version 3.9.13 and APISIX from version 3.17.0.

  • keepalive

    boolean

    default: true


    If true, keep the connection alive when requesting the LLM service.

  • keepalive_timeout

    integer

    default: 60000

    vaild vaule:

    greater than or equal to 1000


    Keepalive timeout in milliseconds when requesting the LLM service.

  • keepalive_pool

    integer

    default: 30

    vaild vaule:

    greater than or equal to 1


    Keepalive pool size for when connecting with the LLM service.

  • ssl_verify

    boolean

    default: true


    If true, verify the LLM service's certificate.