Parameters
See plugin common configurations for configuration options available to all plugins.
This plugin supports referencing sensitive parameter values from environment variables using the env:// prefix, or from a secret manager, such as HashiCorp Vault’s KV secrets engine, using the secret:// prefix. For more information, see environment variables in plugin and secrets.
exact
Settings for exact-match caching, where a response is reused only when the normalized request is identical to a previously cached one.
ttl
vaild vaule:
greater than or equal to 1
Time-to-live in seconds for a cached entry.
cache_key
Settings that control how the cache key is scoped.
share_across_routes
If true, the cache key does not include the route ID, so identical requests on different routes can share cached responses. If false, each route has its own cache scope.
include_consumer
If true, the consumer name is included in the cache key, so cached responses are isolated per consumer.
include_vars
Names of additional context variables to include in the cache key scope, so requests with different values of these variables do not share cached responses.
max_cache_body_size
vaild vaule:
greater than or equal to 0
Maximum size in bytes of a response body that will be cached. Larger responses are not cached.
cache_headers
If true, the plugin adds the
X-AI-Cache-Statusresponse header (andX-AI-Cache-Ageon a cache hit). Set to false to omit these headers.fail_mode
vaild vaule:
skip,warn, orerrorBehavior when the request cannot be cached because no AI instance was selected, for example when the route does not also configure
ai-proxyorai-proxy-multi. Withskip, the request is passed through unchecked. Withwarn, the request is passed through and a warning is logged. Witherror, the request is rejected with HTTP 500.bypass_on
A list of request-header matching rules. If a request matches any rule, the cache is bypassed and the response is marked with
X-AI-Cache-Statusset toBYPASS.header
vaild vaule:
non-empty
Name of the request header to match.
equals
Value that the header must equal for the rule to match.
policy
vaild vaule:
redisCache storage backend. Currently only
redisis supported.layers
vaild vaule:
exact, or bothexactandsemanticCache layers to enable.
exactis always required. Addsemanticto look for similar prompts through embeddings and vector search after an exact miss.Semantic caching is available in API7 Enterprise 3.9.16+ or 3.10.3+.
semantic
Settings for semantic caching. Required when
layersincludessemantic.Available in API7 Enterprise 3.9.16+ or 3.10.3+.
similarity_threshold
vaild vaule:
between 0 and 1 inclusive
Minimum similarity score required for a semantic cache hit.
top_k
vaild vaule:
greater than or equal to 1
Number of nearest vector matches to retrieve.
distance_metric
vaild vaule:
cosineDistance metric used by vector search.
ttl
vaild vaule:
greater than or equal to 1
Time-to-live in seconds for semantic cache entries.
match
Settings that control which conversation content is embedded for semantic matching.
message_countback
vaild vaule:
greater than or equal to 1
Number of recent user-message turns to include in the embedding input.
ignore_system_prompts
If true, system prompts are excluded from the embedding input.
ignore_assistant_prompts
If true, assistant messages are excluded from the embedding input.
ignore_tool_prompts
If true, tool messages are excluded from the embedding input.
embedding
Embedding provider configuration. Configure exactly one of
openaiorazure_openai.openai
OpenAI-compatible embedding provider settings. Requires
modelandapi_key.endpoint
OpenAI-compatible embedding API endpoint. When not configured, the public OpenAI embeddings endpoint is used.
model
Embedding model name, such as
text-embedding-3-small.api_key
API key used to authenticate with the embedding provider. The value is encrypted with AES before being stored in etcd.
dimensions
vaild vaule:
greater than or equal to 1
Number of dimensions in the embedding output. Configure this only for models that support overriding the output dimensions.
ssl_verify
If true, verify the embedding provider's TLS certificate.
timeout
vaild vaule:
greater than or equal to 1
Timeout in milliseconds for requests to the embedding provider.
azure_openai
Azure OpenAI embedding provider settings. Requires
endpointandapi_key.endpoint
Azure OpenAI embeddings endpoint.
api_key
API key used to authenticate with Azure OpenAI. The value is encrypted with AES before being stored in etcd.
dimensions
vaild vaule:
greater than or equal to 1
Number of dimensions in the embedding output. Configure this only for models that support overriding the output dimensions.
ssl_verify
If true, verify the Azure OpenAI endpoint's TLS certificate.
timeout
vaild vaule:
greater than or equal to 1
Timeout in milliseconds for requests to Azure OpenAI.
vector_search
Vector search backend configuration.
redis
RediSearch vector index settings.
index
Name of the RediSearch index used by semantic caching.
redis_host
vaild vaule:
at least 2 characters
Address of the Redis server. Required when
policyisredis.redis_port
vaild vaule:
greater than or equal to 1
Port of the Redis server.
redis_username
Username for Redis authentication when using Redis ACLs.
redis_password
Password for Redis authentication.
redis_database
vaild vaule:
greater than or equal to 0
Database number of the Redis server.
redis_timeout
vaild vaule:
greater than or equal to 1
Timeout in milliseconds for Redis operations.
redis_ssl
If true, use TLS for the connection to Redis.
redis_ssl_verify
If true, verify the TLS certificate of the Redis server.
redis_keepalive_timeout
vaild vaule:
greater than or equal to 1000
Keepalive timeout in milliseconds for the Redis connection pool. Available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line.
redis_keepalive_pool
vaild vaule:
greater than or equal to 1
Keepalive pool size for Redis connections. Available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line.