Skip to main content

Parameters

See plugin common configurations for configuration options available to all plugins.

This plugin supports referencing sensitive parameter values from environment variables using the env:// prefix, or from a secret manager, such as HashiCorp Vault’s KV secrets engine, using the secret:// prefix. For more information, see environment variables in plugin and secrets.

  • exact

    object


    Settings for exact-match caching, where a response is reused only when the normalized request is identical to a previously cached one.

    • ttl

      integer

      default: 3600

      vaild vaule:

      greater than or equal to 1


      Time-to-live in seconds for a cached entry.

  • cache_key

    object


    Settings that control how the cache key is scoped.

    • share_across_routes

      boolean

      default: false


      If true, the cache key does not include the route ID, so identical requests on different routes can share cached responses. If false, each route has its own cache scope.

    • include_consumer

      boolean

      default: false


      If true, the consumer name is included in the cache key, so cached responses are isolated per consumer.

    • include_vars

      array[string]

      default: []


      Names of additional context variables to include in the cache key scope, so requests with different values of these variables do not share cached responses.

  • max_cache_body_size

    integer

    default: 1048576

    vaild vaule:

    greater than or equal to 0


    Maximum size in bytes of a response body that will be cached. Larger responses are not cached.

  • cache_headers

    boolean

    default: true


    If true, the plugin adds the X-AI-Cache-Status response header (and X-AI-Cache-Age on a cache hit). Set to false to omit these headers.

  • fail_mode

    string

    default: skip

    vaild vaule:

    skip, warn, or error


    Behavior when the request cannot be cached because no AI instance was selected, for example when the route does not also configure ai-proxy or ai-proxy-multi. With skip, the request is passed through unchecked. With warn, the request is passed through and a warning is logged. With error, the request is rejected with HTTP 500.

  • bypass_on

    array[object]


    A list of request-header matching rules. If a request matches any rule, the cache is bypassed and the response is marked with X-AI-Cache-Status set to BYPASS.

    • header

      string

      required

      vaild vaule:

      non-empty


      Name of the request header to match.

    • equals

      string

      required


      Value that the header must equal for the rule to match.

  • policy

    string

    default: redis

    vaild vaule:

    redis


    Cache storage backend. Currently only redis is supported.

  • layers

    array[string]

    default: ["exact"]

    vaild vaule:

    exact, or both exact and semantic


    Cache layers to enable. exact is always required. Add semantic to look for similar prompts through embeddings and vector search after an exact miss.

    Semantic caching is available in API7 Enterprise 3.9.16+ or 3.10.3+.

  • semantic

    object


    Settings for semantic caching. Required when layers includes semantic.

    Available in API7 Enterprise 3.9.16+ or 3.10.3+.

    • similarity_threshold

      number

      default: 0.95

      vaild vaule:

      between 0 and 1 inclusive


      Minimum similarity score required for a semantic cache hit.

    • top_k

      integer

      default: 1

      vaild vaule:

      greater than or equal to 1


      Number of nearest vector matches to retrieve.

    • distance_metric

      string

      default: cosine

      vaild vaule:

      cosine


      Distance metric used by vector search.

    • ttl

      integer

      default: 86400

      vaild vaule:

      greater than or equal to 1


      Time-to-live in seconds for semantic cache entries.

    • match

      object


      Settings that control which conversation content is embedded for semantic matching.

      • message_countback

        integer

        default: 1

        vaild vaule:

        greater than or equal to 1


        Number of recent user-message turns to include in the embedding input.

      • ignore_system_prompts

        boolean

        default: true


        If true, system prompts are excluded from the embedding input.

      • ignore_assistant_prompts

        boolean

        default: true


        If true, assistant messages are excluded from the embedding input.

      • ignore_tool_prompts

        boolean

        default: true


        If true, tool messages are excluded from the embedding input.

    • embedding

      object

      required


      Embedding provider configuration. Configure exactly one of openai or azure_openai.

      • openai

        object


        OpenAI-compatible embedding provider settings. Requires model and api_key.

        • endpoint

          string


          OpenAI-compatible embedding API endpoint. When not configured, the public OpenAI embeddings endpoint is used.

        • model

          string

          required


          Embedding model name, such as text-embedding-3-small.

        • api_key

          string

          required


          API key used to authenticate with the embedding provider. The value is encrypted with AES before being stored in etcd.

        • dimensions

          integer

          vaild vaule:

          greater than or equal to 1


          Number of dimensions in the embedding output. Configure this only for models that support overriding the output dimensions.

        • ssl_verify

          boolean

          default: true


          If true, verify the embedding provider's TLS certificate.

        • timeout

          integer

          default: 5000

          vaild vaule:

          greater than or equal to 1


          Timeout in milliseconds for requests to the embedding provider.

      • azure_openai

        object


        Azure OpenAI embedding provider settings. Requires endpoint and api_key.

        • endpoint

          string

          required


          Azure OpenAI embeddings endpoint.

        • api_key

          string

          required


          API key used to authenticate with Azure OpenAI. The value is encrypted with AES before being stored in etcd.

        • dimensions

          integer

          vaild vaule:

          greater than or equal to 1


          Number of dimensions in the embedding output. Configure this only for models that support overriding the output dimensions.

        • ssl_verify

          boolean

          default: true


          If true, verify the Azure OpenAI endpoint's TLS certificate.

        • timeout

          integer

          default: 5000

          vaild vaule:

          greater than or equal to 1


          Timeout in milliseconds for requests to Azure OpenAI.

    • vector_search

      object

      required


      Vector search backend configuration.

      • redis

        object

        required


        RediSearch vector index settings.

        • index

          string

          default: ai-cache


          Name of the RediSearch index used by semantic caching.

  • redis_host

    string

    required

    vaild vaule:

    at least 2 characters


    Address of the Redis server. Required when policy is redis.

  • redis_port

    integer

    default: 6379

    vaild vaule:

    greater than or equal to 1


    Port of the Redis server.

  • redis_username

    string


    Username for Redis authentication when using Redis ACLs.

  • redis_password

    string


    Password for Redis authentication.

  • redis_database

    integer

    default: 0

    vaild vaule:

    greater than or equal to 0


    Database number of the Redis server.

  • redis_timeout

    integer

    default: 1000

    vaild vaule:

    greater than or equal to 1


    Timeout in milliseconds for Redis operations.

  • redis_ssl

    boolean

    default: false


    If true, use TLS for the connection to Redis.

  • redis_ssl_verify

    boolean

    default: false


    If true, verify the TLS certificate of the Redis server.

  • redis_keepalive_timeout

    integer

    default: 10000

    vaild vaule:

    greater than or equal to 1000


    Keepalive timeout in milliseconds for the Redis connection pool. Available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line.

  • redis_keepalive_pool

    integer

    default: 100

    vaild vaule:

    greater than or equal to 1


    Keepalive pool size for Redis connections. Available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line.