ai-rate-limiting
The ai-rate-limiting plugin enforces token-based rate limiting for requests sent to LLM services. It controls the tokens consumed during a configured period, helping allocate resources fairly and prevent excessive load. It is often used with the ai-proxy-multi plugin.
Local vs Redis Rate Limiting
The ai-rate-limiting plugin supports two modes of rate limiting:
- Local rate limiting: Limits are enforced independently on each gateway instance. Each instance maintains its own counters, so the effective limit is roughly (limit × number of instances) when traffic is spread across instances. This is the default when no
policyis set or whenpolicyislocal. - Redis-based rate limiting: Limits are shared across all gateway instances through Redis. All instances share the same quota, so the configured limit applies to all gateway instances.
In a multi-instance deployment, send verification requests through different gateway addresses to confirm that they consume the same Redis-backed quota.
The Redis examples are evaluation fixtures with tutorial passwords and internal service addresses. For production deployments, manage credentials as secrets, restrict network access, and use redis_ssl or redis_cluster_ssl with certificate verification when the Redis service supports TLS.
Demo
The following demo configures two models with different priorities and rate limits the higher-priority instance in the API7 Enterprise Dashboard. With fallback_strategy set to ["rate_limiting"], requests fall back to the lower-priority instance after the first instance exhausts its quota. See Configure Instance Priority and Rate Limiting for the configuration.
Examples
The examples below demonstrate how you can configure ai-rate-limiting for different scenarios.
The examples use OpenAI and DeepSeek as the upstream LLM services. Create API keys for the providers used in an example and save them to environment variables:
export OPENAI_API_KEY=replace-with-openai-api-key
export DEEPSEEK_API_KEY=replace-with-deepseek-api-key
The ADC examples define a conventional service upstream because ADC services require one. The AI proxy plugin handles the matching LLM requests.
Rate Limit One Instance
The following example configures two model instances with an 80/20 traffic split and rate limits the instance receiving 80% of traffic. After that instance exhausts its quota, the rate_limiting fallback strategy forwards additional requests to the other instance.
Create a route as such and update with your LLM providers, models, API keys, and endpoints:
- Admin API
- ADC
- Ingress Controller
curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d @- <<EOF
{
"id": "ai-rate-limiting-route",
"uri": "/anything",
"methods": ["POST"],
"plugins": {
"ai-proxy-multi": {
"fallback_strategy": ["rate_limiting"],
"instances": [
{
"name": "deepseek-instance-1",
"provider": "deepseek",
"weight": 8,
"auth": {
"header": {
"Authorization": "Bearer $DEEPSEEK_API_KEY"
}
},
"options": {
"model": "deepseek-flash"
}
},
{
"name": "deepseek-instance-2",
"provider": "deepseek",
"weight": 2,
"auth": {
"header": {
"Authorization": "Bearer $DEEPSEEK_API_KEY"
}
},
"options": {
"model": "deepseek-flash"
}
}
]
},
"ai-rate-limiting": {
"policy": "local",
"limit_strategy": "total_tokens",
"instances": [
{
"name": "deepseek-instance-1",
"limit": 100,
"time_window": 30
}
]
}
}
}
EOF
services:
- name: ai-rate-limiting-service
labels:
docs-example: ai-rate-limit-one-instance
routes:
- name: ai-rate-limiting-route
uris:
- /anything
methods:
- POST
plugins:
ai-proxy-multi:
fallback_strategy:
- rate_limiting
instances:
- name: deepseek-instance-1
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer ${DEEPSEEK_API_KEY}"
options:
model: deepseek-flash
- name: deepseek-instance-2
provider: deepseek
weight: 2
auth:
header:
Authorization: "Bearer ${DEEPSEEK_API_KEY}"
options:
model: deepseek-flash
ai-rate-limiting:
policy: local
limit_strategy: total_tokens
instances:
- name: deepseek-instance-1
limit: 100
time_window: 30
upstream:
type: roundrobin
nodes:
- host: httpbin.org
port: 80
weight: 1
Preview changes owned by this example and confirm that the diff contains no unintended updates or deletions:
adc diff -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limit-one-instance
Synchronize the reviewed service configuration:
adc sync -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limit-one-instance
- Gateway API
- APISIX CRD
apiVersion: apisix.apache.org/v1alpha1
kind: PluginConfig
metadata:
namespace: aic
name: ai-rate-limiting-plugin-config
spec:
plugins:
- name: ai-proxy-multi
config:
fallback_strategy:
- rate_limiting
instances:
- name: deepseek-instance-1
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: deepseek-instance-2
provider: deepseek
weight: 2
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: ai-rate-limiting
config:
policy: local
limit_strategy: total_tokens
instances:
- name: deepseek-instance-1
limit: 100
time_window: 30
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
namespace: aic
name: ai-rate-limiting-route
spec:
parentRefs:
- name: apisix
rules:
- matches:
- path:
type: Exact
value: /anything
method: POST
filters:
- type: ExtensionRef
extensionRef:
group: apisix.apache.org
kind: PluginConfig
name: ai-rate-limiting-plugin-config
apiVersion: apisix.apache.org/v2
kind: ApisixRoute
metadata:
namespace: aic
name: ai-rate-limiting-route
spec:
ingressClassName: apisix
http:
- name: ai-rate-limiting-route
match:
paths:
- /anything
methods:
- POST
plugins:
- name: ai-proxy-multi
enable: true
config:
fallback_strategy:
- rate_limiting
instances:
- name: deepseek-instance-1
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: deepseek-instance-2
provider: deepseek
weight: 2
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: ai-rate-limiting
enable: true
config:
policy: local
limit_strategy: total_tokens
instances:
- name: deepseek-instance-1
limit: 100
time_window: 30
Apply the configuration to your cluster:
kubectl apply -f ai-rate-limiting-ic.yaml
❶ Apply rate limiting by total_tokens.
❷ Apply rate limiting on the deepseek-instance-1 instance.
❸ Configure a quota of 100 tokens.
❹ Configure the time window to be 30 seconds.
Send a POST request to the route with a system prompt and a sample user question in the request body:
curl "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
You should receive a response similar to the following:
{
...
"model": "deepseek-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "1 + 1 equals 2. This is a fundamental arithmetic operation where adding one unit to another results in a total of two units."
},
"logprobs": null,
"finish_reason": "stop"
}
],
...
}
After deepseek-instance-1 consumes its 100-token quota within a 30-second window, the rate_limiting fallback strategy excludes it from selection. Additional requests are forwarded to deepseek-instance-2, which is not rate limited.
Apply the Same Quota to All Instances
The following example demonstrates how you can apply the same rate limiting quota to all LLM upstream instances in ai-rate-limiting.
For demonstration and easier differentiation, you will be configuring one OpenAI instance and one DeepSeek instance as the upstream LLM services.
Create a route as such and update with your LLM providers, models, API keys, and endpoints:
- Admin API
- ADC
- Ingress Controller
curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d @- <<EOF
{
"id": "ai-rate-limiting-route",
"uri": "/anything",
"methods": ["POST"],
"plugins": {
"ai-proxy-multi": {
"instances": [
{
"name": "openai-instance",
"provider": "openai",
"weight": 0,
"auth": {
"header": {
"Authorization": "Bearer $OPENAI_API_KEY"
}
},
"options": {
"model": "gpt-4o-mini"
}
},
{
"name": "deepseek-instance",
"provider": "deepseek",
"weight": 0,
"auth": {
"header": {
"Authorization": "Bearer $DEEPSEEK_API_KEY"
}
},
"options": {
"model": "deepseek-flash"
}
}
]
},
"ai-rate-limiting": {
"policy": "local",
"limit": 500,
"time_window": 60,
"rejected_code": 429,
"limit_strategy": "total_tokens"
}
}
}
EOF
services:
- name: ai-rate-limiting-service
labels:
docs-example: ai-rate-limit-shared-quota
routes:
- name: ai-rate-limiting-route
uris:
- /anything
methods:
- POST
plugins:
ai-proxy-multi:
instances:
- name: openai-instance
provider: openai
weight: 0
auth:
header:
Authorization: "Bearer ${OPENAI_API_KEY}"
options:
model: gpt-4o-mini
- name: deepseek-instance
provider: deepseek
weight: 0
auth:
header:
Authorization: "Bearer ${DEEPSEEK_API_KEY}"
options:
model: deepseek-flash
ai-rate-limiting:
policy: local
limit: 500
time_window: 60
rejected_code: 429
limit_strategy: total_tokens
upstream:
type: roundrobin
nodes:
- host: httpbin.org
port: 80
weight: 1
Preview changes owned by this example and confirm that the diff contains no unintended updates or deletions:
adc diff -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limit-shared-quota
Synchronize the reviewed service configuration:
adc sync -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limit-shared-quota
- Gateway API
- APISIX CRD
apiVersion: apisix.apache.org/v1alpha1
kind: PluginConfig
metadata:
namespace: aic
name: ai-rate-limiting-plugin-config
spec:
plugins:
- name: ai-proxy-multi
config:
instances:
- name: openai-instance
provider: openai
weight: 0
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: deepseek-instance
provider: deepseek
weight: 0
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: ai-rate-limiting
config:
policy: local
limit: 500
time_window: 60
rejected_code: 429
limit_strategy: total_tokens
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
namespace: aic
name: ai-rate-limiting-route
spec:
parentRefs:
- name: apisix
rules:
- matches:
- path:
type: Exact
value: /anything
method: POST
filters:
- type: ExtensionRef
extensionRef:
group: apisix.apache.org
kind: PluginConfig
name: ai-rate-limiting-plugin-config
apiVersion: apisix.apache.org/v2
kind: ApisixRoute
metadata:
namespace: aic
name: ai-rate-limiting-route
spec:
ingressClassName: apisix
http:
- name: ai-rate-limiting-route
match:
paths:
- /anything
methods:
- POST
plugins:
- name: ai-proxy-multi
enable: true
config:
instances:
- name: openai-instance
provider: openai
weight: 0
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: deepseek-instance
provider: deepseek
weight: 0
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: ai-rate-limiting
enable: true
config:
policy: local
limit: 500
time_window: 60
rejected_code: 429
limit_strategy: total_tokens
Apply the configuration to your cluster:
kubectl apply -f ai-rate-limiting-ic.yaml
❶ Configure a shared rate-limiting quota of 500 tokens for all instances.
❷ Configure the time window to be 60 seconds.
❸ Set the rejection response HTTP status code to 429.
❹ Apply rate limiting by total_tokens.
Send a POST request to the route with a system prompt and a sample user question in the request body:
curl -i "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "Explain Newtons laws" }
]
}'
You should receive a response from either LLM instance, similar to the following:
{
...,
"model": "gpt-4o-mini",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Sure! Sir Isaac Newton formulated three laws of motion that describe the motion of objects. These laws are widely used in physics and engineering for studying and understanding how things move. Here they are:\n\n1. Newton's First Law - Law of Inertia: An object at rest tends to stay at rest and an object in motion tends to stay in motion with the same speed and in the same direction unless acted upon by an unbalanced force. This is also known as the principle of inertia.\n\n2. Newton's Second Law of Motion - Force and Acceleration: The acceleration of an object is directly proportional to the net force acting on it and inversely proportional to its mass. This is usually formulated as F=ma where F is the force applied, m is the mass of the object and a is the acceleration produced.\n\n3. Newton's Third Law - Action and Reaction: For every action, there is an equal and opposite reaction. This means that any force exerted on a body will create a force of equal magnitude but in the opposite direction on the object that exerted the first force.\n\nIn simple terms: \n1. If you slide a book on a table and let go, it will stop because of the friction (or force) between it and the table.\n2.",
"refusal": null
},
"logprobs": null,
"finish_reason": "length"
}
],
"usage": {
"prompt_tokens": 23,
"completion_tokens": 256,
"total_tokens": 279,
"prompt_tokens_details": {
"cached_tokens": 0,
"audio_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 0,
"audio_tokens": 0,
"accepted_prediction_tokens": 0,
"rejected_prediction_tokens": 0
}
},
"service_tier": "default",
"system_fingerprint": null
}
The response consumes 279 tokens from the shared 500-token quota.
Within the same 60-second window, send another POST request to the route:
curl -i "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "Explain Newtons laws" }
]
}'
You should receive another successful response from either LLM instance, similar to the following:
{
...
"model": "deepseek-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Sure! Newton's laws of motion are three fundamental principles that describe the relationship between the motion of an object and the forces acting on it. They were formulated by Sir Isaac Newton in the late 17th century and are foundational to classical mechanics. Here's an explanation of each law:\n\n---\n\n### **1. Newton's First Law (Law of Inertia)**\n- **Statement**: An object will remain at rest or in uniform motion in a straight line unless acted upon by an external force.\n- **What it means**: This law introduces the concept of **inertia**, which is the tendency of an object to resist changes in its state of motion. If no net force acts on an object, its velocity (speed and direction) will not change.\n- **Example**: A book lying on a table will stay at rest unless you push it. Similarly, a hockey puck sliding on ice will keep moving at a constant speed unless friction or another force slows it down.\n\n---\n\n### **2. Newton's Second Law (Law of Acceleration)**\n- **Statement**: The acceleration of an object is directly proportional to the net force acting on it and inversely proportional to its mass. Mathematically, this is expressed as:\n \\[\n F = ma\n \\]\n"
},
"logprobs": null,
"finish_reason": "length"
}
],
"usage": {
"prompt_tokens": 13,
"completion_tokens": 256,
"total_tokens": 269,
"prompt_tokens_details": {
"cached_tokens": 0
},
"prompt_cache_hit_tokens": 0,
"prompt_cache_miss_tokens": 13
},
"system_fingerprint": "fp_3a5770e1b4_prod0225"
}
With the sample responses shown, the two requests consume 548 tokens in total and exhaust the shared quota.
Within the same 60-second window, send a third POST request to the route:
curl -i "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "Explain Newtons laws" }
]
}'
You should receive an HTTP 429 Too Many Requests response. The selected instance name is appended to the rate-limit header names:
X-AI-RateLimit-Limit-{instance-name}: 500
X-AI-RateLimit-Remaining-{instance-name}: 0
X-AI-RateLimit-Reset-{instance-name}: 42.137000083923
Configure Instance Priority and Rate Limiting
The following example configures two models with different priorities and rate limits the higher-priority instance. With fallback_strategy set to ["rate_limiting"], requests fall back to the lower-priority instance after the first instance exhausts its quota.
Create a route as such and update with your LLM providers, models, API keys, and endpoints:
- Admin API
- ADC
- Ingress Controller
curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d @- <<EOF
{
"id": "ai-rate-limiting-route",
"uri": "/anything",
"methods": ["POST"],
"plugins": {
"ai-proxy-multi": {
"fallback_strategy": ["rate_limiting"],
"instances": [
{
"name": "openai-instance",
"provider": "openai",
"priority": 1,
"weight": 0,
"auth": {
"header": {
"Authorization": "Bearer $OPENAI_API_KEY"
}
},
"options": {
"model": "gpt-4o-mini"
}
},
{
"name": "deepseek-instance",
"provider": "deepseek",
"priority": 0,
"weight": 0,
"auth": {
"header": {
"Authorization": "Bearer $DEEPSEEK_API_KEY"
}
},
"options": {
"model": "deepseek-flash"
}
}
]
},
"ai-rate-limiting": {
"policy": "local",
"instances": [
{
"name": "openai-instance",
"limit": 10,
"time_window": 60
}
],
"limit_strategy": "total_tokens"
}
}
}
EOF
services:
- name: ai-rate-limiting-service
labels:
docs-example: ai-rate-limit-priority
routes:
- name: ai-rate-limiting-route
uris:
- /anything
methods:
- POST
plugins:
ai-proxy-multi:
fallback_strategy:
- rate_limiting
instances:
- name: openai-instance
provider: openai
priority: 1
weight: 0
auth:
header:
Authorization: "Bearer ${OPENAI_API_KEY}"
options:
model: gpt-4o-mini
- name: deepseek-instance
provider: deepseek
priority: 0
weight: 0
auth:
header:
Authorization: "Bearer ${DEEPSEEK_API_KEY}"
options:
model: deepseek-flash
ai-rate-limiting:
policy: local
instances:
- name: openai-instance
limit: 10
time_window: 60
limit_strategy: total_tokens
upstream:
type: roundrobin
nodes:
- host: httpbin.org
port: 80
weight: 1
Preview changes owned by this example and confirm that the diff contains no unintended updates or deletions:
adc diff -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limit-priority
Synchronize the reviewed service configuration:
adc sync -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limit-priority
- Gateway API
- APISIX CRD
apiVersion: apisix.apache.org/v1alpha1
kind: PluginConfig
metadata:
namespace: aic
name: ai-rate-limiting-plugin-config
spec:
plugins:
- name: ai-proxy-multi
config:
fallback_strategy:
- rate_limiting
instances:
- name: openai-instance
provider: openai
priority: 1
weight: 0
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: deepseek-instance
provider: deepseek
priority: 0
weight: 0
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: ai-rate-limiting
config:
policy: local
instances:
- name: openai-instance
limit: 10
time_window: 60
limit_strategy: total_tokens
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
namespace: aic
name: ai-rate-limiting-route
spec:
parentRefs:
- name: apisix
rules:
- matches:
- path:
type: Exact
value: /anything
method: POST
filters:
- type: ExtensionRef
extensionRef:
group: apisix.apache.org
kind: PluginConfig
name: ai-rate-limiting-plugin-config
apiVersion: apisix.apache.org/v2
kind: ApisixRoute
metadata:
namespace: aic
name: ai-rate-limiting-route
spec:
ingressClassName: apisix
http:
- name: ai-rate-limiting-route
match:
paths:
- /anything
methods:
- POST
plugins:
- name: ai-proxy-multi
enable: true
config:
fallback_strategy:
- rate_limiting
instances:
- name: openai-instance
provider: openai
priority: 1
weight: 0
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: deepseek-instance
provider: deepseek
priority: 0
weight: 0
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: ai-rate-limiting
enable: true
config:
policy: local
instances:
- name: openai-instance
limit: 10
time_window: 60
limit_strategy: total_tokens
Apply the configuration to your cluster:
kubectl apply -f ai-rate-limiting-ic.yaml
❶ Set the fallback_strategy to ["rate_limiting"].
❷ Set a higher priority on openai-instance instance.
❸ Set a lower priority on deepseek-instance instance.
❹ Apply rate limiting on openai-instance instance.
❺ Configure a quota of 10 tokens.
❻ Configure the time window to be 60 seconds.
❼ Apply rate limiting by total_tokens.
Send a POST request to the route with a system prompt and a sample user question in the request body:
curl "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
You should receive a response similar to the following:
{
...,
"model": "gpt-4o-mini",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "1+1 equals 2.",
"refusal": null
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 23,
"completion_tokens": 8,
"total_tokens": 31,
"prompt_tokens_details": {
"cached_tokens": 0,
"audio_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 0,
"audio_tokens": 0,
"accepted_prediction_tokens": 0,
"rejected_prediction_tokens": 0
}
},
"service_tier": "default",
"system_fingerprint": null
}
Since the total_tokens value exceeds the configured quota of 10, the next request within the 60-second window is expected to be forwarded to the other instance.
Within the same 60-second window, send another POST request to the route:
curl "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "Explain Newton law" }
]
}'
You should see a response similar to the following:
{
...,
"model": "deepseek-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Certainly! Newton's laws of motion are three fundamental principles that describe the relationship between the motion of an object and the forces acting on it. They were formulated by Sir Isaac Newton in the late 17th century and are foundational to classical mechanics.\n\n---\n\n### **1. Newton's First Law (Law of Inertia):**\n- **Statement:** An object at rest will remain at rest, and an object in motion will continue moving at a constant velocity (in a straight line at a constant speed), unless acted upon by an external force.\n- **Key Idea:** This law introduces the concept of **inertia**, which is the tendency of an object to resist changes in its state of motion.\n- **Example:** If you slide a book across a table, it eventually stops because of the force of friction acting on it. Without friction, the book would keep moving indefinitely.\n\n---\n\n### **2. Newton's Second Law (Law of Acceleration):**\n- **Statement:** The acceleration of an object is directly proportional to the net force acting on it and inversely proportional to its mass. Mathematically, this is expressed as:\n \\[\n F = ma\n \\]\n where:\n - \\( F \\) = net force applied (in Newtons),\n -"
},
...
}
],
...
}
Load Balance and Rate Limit by Consumers
The following example demonstrates how you can configure two models for load balancing and apply rate limiting by consumer.
Create a consumer johndoe and a rate limiting quota of 10 tokens in a 60-second window on openai-instance instance:
- Admin API
- ADC
- Ingress Controller
curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d '{
"username": "johndoe",
"plugins": {
"ai-rate-limiting": {
"policy": "local",
"instances": [
{
"name": "openai-instance",
"limit": 10,
"time_window": 60
}
],
"rejected_code": 429,
"limit_strategy": "total_tokens"
}
}
}'
Configure key-auth credential for johndoe:
curl "http://127.0.0.1:9180/apisix/admin/consumers/johndoe/credentials" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d '{
"id": "cred-john-key-auth",
"plugins": {
"key-auth": {
"key": "john-key"
}
}
}'
Create another consumer janedoe and a rate limiting quota of 10 tokens in a 60-second window on deepseek-instance instance:
curl "http://127.0.0.1:9180/apisix/admin/consumers" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d '{
"username": "janedoe",
"plugins": {
"ai-rate-limiting": {
"policy": "local",
"instances": [
{
"name": "deepseek-instance",
"limit": 10,
"time_window": 60
}
],
"rejected_code": 429,
"limit_strategy": "total_tokens"
}
}
}'
Configure key-auth credential for janedoe:
curl "http://127.0.0.1:9180/apisix/admin/consumers/janedoe/credentials" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d '{
"id": "cred-jane-key-auth",
"plugins": {
"key-auth": {
"key": "jane-key"
}
}
}'
Create a route as such and update with your LLM providers, models, API keys, and endpoints:
curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d @- <<EOF
{
"id": "ai-rate-limiting-route",
"uri": "/anything",
"methods": ["POST"],
"plugins": {
"key-auth": {},
"ai-proxy-multi": {
"fallback_strategy": ["rate_limiting"],
"instances": [
{
"name": "openai-instance",
"provider": "openai",
"weight": 0,
"auth": {
"header": {
"Authorization": "Bearer $OPENAI_API_KEY"
}
},
"options": {
"model": "gpt-4o-mini"
}
},
{
"name": "deepseek-instance",
"provider": "deepseek",
"weight": 0,
"auth": {
"header": {
"Authorization": "Bearer $DEEPSEEK_API_KEY"
}
},
"options": {
"model": "deepseek-flash"
}
}
]
}
}
}
EOF
Create two consumers and a route that enables rate limiting by consumers:
consumers:
- username: johndoe
labels:
docs-example: ai-rate-limit-consumers
plugins:
ai-rate-limiting:
policy: local
instances:
- name: openai-instance
limit: 10
time_window: 60
rejected_code: 429
limit_strategy: total_tokens
credentials:
- name: key-auth
type: key-auth
config:
key: john-key
- username: janedoe
labels:
docs-example: ai-rate-limit-consumers
plugins:
ai-rate-limiting:
policy: local
instances:
- name: deepseek-instance
limit: 10
time_window: 60
rejected_code: 429
limit_strategy: total_tokens
credentials:
- name: key-auth
type: key-auth
config:
key: jane-key
services:
- name: ai-rate-limiting-service
labels:
docs-example: ai-rate-limit-consumers
routes:
- name: ai-rate-limiting-route
uris:
- /anything
methods:
- POST
plugins:
key-auth: {}
ai-proxy-multi:
fallback_strategy:
- rate_limiting
instances:
- name: openai-instance
provider: openai
weight: 0
auth:
header:
Authorization: "Bearer ${OPENAI_API_KEY}"
options:
model: gpt-4o-mini
- name: deepseek-instance
provider: deepseek
weight: 0
auth:
header:
Authorization: "Bearer ${DEEPSEEK_API_KEY}"
options:
model: deepseek-flash
upstream:
type: roundrobin
nodes:
- host: httpbin.org
port: 80
weight: 1
Preview changes to the service and consumers owned by this example:
adc diff -f adc.yaml \
--include-resource-type service \
--include-resource-type consumer \
--label-selector docs-example=ai-rate-limit-consumers
Synchronize the reviewed service and consumers, including their nested credentials:
adc sync -f adc.yaml \
--include-resource-type service \
--include-resource-type consumer \
--label-selector docs-example=ai-rate-limit-consumers
- Gateway API
- APISIX CRD
Create two consumers and a route that enables rate limiting by consumers:
apiVersion: apisix.apache.org/v1alpha1
kind: Consumer
metadata:
namespace: aic
name: johndoe
spec:
gatewayRef:
name: apisix
credentials:
- type: key-auth
name: primary-key
config:
key: john-key
plugins:
- name: ai-rate-limiting
config:
policy: local
instances:
- name: openai-instance
limit: 10
time_window: 60
rejected_code: 429
limit_strategy: total_tokens
---
apiVersion: apisix.apache.org/v1alpha1
kind: Consumer
metadata:
namespace: aic
name: janedoe
spec:
gatewayRef:
name: apisix
credentials:
- type: key-auth
name: primary-key
config:
key: jane-key
plugins:
- name: ai-rate-limiting
config:
policy: local
instances:
- name: deepseek-instance
limit: 10
time_window: 60
rejected_code: 429
limit_strategy: total_tokens
---
apiVersion: apisix.apache.org/v1alpha1
kind: PluginConfig
metadata:
namespace: aic
name: ai-rate-limiting-plugin-config
spec:
plugins:
- name: key-auth
config:
_meta:
disable: false
- name: ai-proxy-multi
config:
fallback_strategy:
- rate_limiting
instances:
- name: openai-instance
provider: openai
weight: 0
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: deepseek-instance
provider: deepseek
weight: 0
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
namespace: aic
name: ai-rate-limiting-route
spec:
parentRefs:
- name: apisix
rules:
- matches:
- path:
type: Exact
value: /anything
method: POST
filters:
- type: ExtensionRef
extensionRef:
group: apisix.apache.org
kind: PluginConfig
name: ai-rate-limiting-plugin-config
Apply the configuration to your cluster:
kubectl apply -f ai-rate-limiting-ic.yaml
apiVersion: apisix.apache.org/v2
kind: ApisixConsumer
metadata:
namespace: aic
name: johndoe
spec:
ingressClassName: apisix
authParameter:
keyAuth:
value:
key: john-key
plugins:
- name: ai-rate-limiting
config:
policy: local
instances:
- name: openai-instance
limit: 10
time_window: 60
rejected_code: 429
limit_strategy: total_tokens
---
apiVersion: apisix.apache.org/v2
kind: ApisixConsumer
metadata:
namespace: aic
name: janedoe
spec:
ingressClassName: apisix
authParameter:
keyAuth:
value:
key: jane-key
plugins:
- name: ai-rate-limiting
config:
policy: local
instances:
- name: deepseek-instance
limit: 10
time_window: 60
rejected_code: 429
limit_strategy: total_tokens
---
apiVersion: apisix.apache.org/v2
kind: ApisixRoute
metadata:
namespace: aic
name: ai-rate-limiting-route
spec:
ingressClassName: apisix
http:
- name: ai-rate-limiting-route
match:
paths:
- /anything
methods:
- POST
plugins:
- name: key-auth
config:
_meta:
disable: false
- name: ai-proxy-multi
config:
fallback_strategy:
- rate_limiting
instances:
- name: openai-instance
provider: openai
weight: 0
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: deepseek-instance
provider: deepseek
weight: 0
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
Apply the configuration to your cluster:
kubectl apply -f ai-rate-limiting-apisix-crd.yaml
❶ Enable key-auth on the route.
❷ Configure an openai instance.
❸ Configure a deepseek instance.
Send a POST request to the route without any consumer key:
curl -i "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
You should receive an HTTP/1.1 401 Unauthorized response.
Send a POST request to the route with johndoe's key:
curl "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-H 'apikey: john-key' \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
The first response can come from either model because both instances have the same priority and weight. An OpenAI response looks similar to the following:
{
...,
"model": "gpt-4o-mini",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "1+1 equals 2.",
"refusal": null
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 23,
"completion_tokens": 8,
"total_tokens": 31,
"prompt_tokens_details": {
"cached_tokens": 0,
"audio_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 0,
"audio_tokens": 0,
"accepted_prediction_tokens": 0,
"rejected_prediction_tokens": 0
}
},
"service_tier": "default",
"system_fingerprint": null
}
If an OpenAI response exhausts johndoe's quota for that instance, subsequent requests within the 60-second window exclude OpenAI and use DeepSeek. If DeepSeek is selected first, repeat the request until OpenAI is selected and its quota is exhausted.
Within the same 60-second window, send another POST request to the route with johndoe's key:
curl "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-H 'apikey: john-key' \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "Explain Newtons laws to me" }
]
}'
After OpenAI is excluded for johndoe, you should see a DeepSeek response similar to the following:
{
...,
"model": "deepseek-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Certainly! Newton's laws of motion are three fundamental principles that describe the relationship between the motion of an object and the forces acting on it. They were formulated by Sir Isaac Newton in the late 17th century and are foundational to classical mechanics.\n\n---\n\n### **1. Newton's First Law (Law of Inertia):**\n- **Statement:** An object at rest will remain at rest, and an object in motion will continue moving at a constant velocity (in a straight line at a constant speed), unless acted upon by an external force.\n- **Key Idea:** This law introduces the concept of **inertia**, which is the tendency of an object to resist changes in its state of motion.\n- **Example:** If you slide a book across a table, it eventually stops because of the force of friction acting on it. Without friction, the book would keep moving indefinitely.\n\n---\n\n### **2. Newton's Second Law (Law of Acceleration):**\n- **Statement:** The acceleration of an object is directly proportional to the net force acting on it and inversely proportional to its mass. Mathematically, this is expressed as:\n \\[\n F = ma\n \\]\n where:\n - \\( F \\) = net force applied (in Newtons),\n -"
},
...
}
],
...
}
Send a POST request to the route with janedoe's key:
curl "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-H 'apikey: jane-key' \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
The first response can again come from either model. A DeepSeek response looks similar to the following:
{
...,
"model": "deepseek-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The sum of 1 and 1 is 2. This is a basic arithmetic operation where you combine two units to get a total of two units."
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 14,
"completion_tokens": 31,
"total_tokens": 45,
"prompt_tokens_details": {
"cached_tokens": 0
},
"prompt_cache_hit_tokens": 0,
"prompt_cache_miss_tokens": 14
},
"system_fingerprint": "fp_3a5770e1b4_prod0225"
}
If a DeepSeek response exhausts janedoe's quota for that instance, subsequent requests within the 60-second window exclude DeepSeek and use OpenAI. If OpenAI is selected first, repeat the request until DeepSeek is selected and its quota is exhausted.
Within the same 60-second window, send another POST request to the route with janedoe's key:
curl "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-H 'apikey: jane-key' \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "Explain Newtons laws to me" }
]
}'
After DeepSeek is excluded for janedoe, you should see an OpenAI response similar to the following:
{
...,
"model": "gpt-4o-mini",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Sure, here are Newton's three laws of motion:\n\n1) Newton's First Law, also known as the Law of Inertia, states that an object at rest will stay at rest, and an object in motion will stay in motion, unless acted on by an external force. In simple words, this law suggests that an object will keep doing whatever it is doing until something causes it to do otherwise. \n\n2) Newton's Second Law states that the force acting on an object is equal to the mass of that object times its acceleration (F=ma). This means that force is directly proportional to mass and acceleration. The heavier the object and the faster it accelerates, the greater the force.\n\n3) Newton's Third Law, also known as the law of action and reaction, states that for every action, there is an equal and opposite reaction. Essentially, any force exerted onto a body will create a force of equal magnitude but in the opposite direction on the object that exerted the first force.\n\nRemember, these laws become less accurate when considering speeds near the speed of light (where Einstein's theory of relativity becomes more appropriate) or objects very small or very large. However, for everyday situations, they provide a good model of how things move.",
"refusal": null
},
"logprobs": null,
"finish_reason": "stop"
}
],
...
}
This shows ai-proxy-multi balancing traffic while ai-rate-limiting applies different per-instance quotas to each consumer.
Rate Limit by Rules
The following example configures different rate-limiting rules based on request attributes. Rules are available in API7 Enterprise from version 3.8.17 and APISIX from version 3.17.0. The example also calculates request cost from raw provider usage, charging twice for completion tokens. Expression-based cost calculation is available in API7 Enterprise from version 3.9.8 and APISIX from version 3.17.0.
Note that all rules are applied sequentially. If a configured key does not exist, the corresponding rule will be skipped.
In addition to HTTP headers, you can also base rules on other built-in variables to implement more flexible and fine-grained rate-limiting strategies.
Create a route that applies rate limits by request header. The route limits each subscription by X-Subscription-ID and applies a stricter limit to trial users by X-Trial-ID:
- Admin API
- ADC
- Ingress Controller
curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d @- <<EOF
{
"id": "ai-rate-limiting-route",
"uri": "/anything",
"methods": ["POST"],
"plugins": {
"ai-proxy": {
"provider": "openai",
"auth": {
"header": {
"Authorization": "Bearer $OPENAI_API_KEY"
}
},
"options": {
"model": "gpt-4o-mini"
}
},
"ai-rate-limiting": {
"policy": "local",
"rejected_code": 429,
"limit_strategy": "expression",
"cost_expr": "prompt_tokens + completion_tokens * 2",
"rules": [
{
"key": "\${http_x_subscription_id}",
"count": "\${http_x_custom_count ?? 500}",
"time_window": 60
},
{
"key": "\${http_x_trial_id}",
"count": 50,
"time_window": 60
}
]
}
},
"upstream": {
"type": "roundrobin",
"nodes": {
"httpbin.org:80": 1
}
}
}
EOF
services:
- name: ai-rate-limiting-service
labels:
docs-example: ai-rate-limit-rules
routes:
- name: ai-rate-limiting-route
uris:
- /anything
methods:
- POST
plugins:
ai-proxy:
provider: openai
auth:
header:
Authorization: "Bearer ${OPENAI_API_KEY}"
options:
model: gpt-4o-mini
ai-rate-limiting:
policy: local
rejected_code: 429
limit_strategy: expression
cost_expr: prompt_tokens + completion_tokens * 2
rules:
- key: "${http_x_subscription_id}"
count: "${http_x_custom_count ?? 500}"
time_window: 60
- key: "${http_x_trial_id}"
count: 50
time_window: 60
upstream:
type: roundrobin
nodes:
- host: httpbin.org
port: 80
weight: 1
Preview changes owned by this example and confirm that the diff contains no unintended updates or deletions:
adc diff -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limit-rules
Synchronize the reviewed service configuration:
adc sync -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limit-rules
- Gateway API
- APISIX CRD
apiVersion: apisix.apache.org/v1alpha1
kind: PluginConfig
metadata:
namespace: aic
name: ai-rate-limiting-plugin-config
spec:
plugins:
- name: ai-proxy
config:
provider: openai
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: ai-rate-limiting
config:
policy: local
rejected_code: 429
limit_strategy: expression
cost_expr: prompt_tokens + completion_tokens * 2
rules:
- key: "${http_x_subscription_id}"
count: "${http_x_custom_count ?? 500}"
time_window: 60
- key: "${http_x_trial_id}"
count: 50
time_window: 60
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
namespace: aic
name: ai-rate-limiting-route
spec:
parentRefs:
- name: apisix
rules:
- matches:
- path:
type: Exact
value: /anything
method: POST
filters:
- type: ExtensionRef
extensionRef:
group: apisix.apache.org
kind: PluginConfig
name: ai-rate-limiting-plugin-config
apiVersion: apisix.apache.org/v2
kind: ApisixRoute
metadata:
namespace: aic
name: ai-rate-limiting-route
spec:
ingressClassName: apisix
http:
- name: ai-rate-limiting-route
match:
paths:
- /anything
methods:
- POST
plugins:
- name: ai-proxy
enable: true
config:
provider: openai
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: ai-rate-limiting
enable: true
config:
policy: local
rejected_code: 429
limit_strategy: expression
cost_expr: prompt_tokens + completion_tokens * 2
rules:
- key: "${http_x_subscription_id}"
count: "${http_x_custom_count ?? 500}"
time_window: 60
- key: "${http_x_trial_id}"
count: 50
time_window: 60
Apply the configuration to your cluster:
kubectl apply -f ai-rate-limiting-ic.yaml
❶ Use the value of the X-Subscription-ID request header as the rate-limiting key.
❷ Set the request limit dynamically based on the X-Custom-Count header. If the header is not provided, a default count of 500 tokens is applied.
❸ Use the value of the X-Trial-ID request header as the rate-limiting key.
The cost expression uses the raw prompt_tokens and completion_tokens fields returned by OpenAI-compatible providers. For native Anthropic usage, an expression can combine input_tokens, cache_creation_input_tokens, cache_read_input_tokens, and output_tokens. Missing variables evaluate to 0. A negative result is clamped to 0, and a fractional result is rounded to the nearest integer before it is committed to the rate-limit counter. Use only arithmetic operators and the supported math functions documented in the configuration reference.
The gateway calculates the quota headers before the current response's token usage is available. It commits that usage in the log phase, so the first successful response shows the full quota and the next response reflects the tokens consumed by the preceding request.
To verify rate limiting, send several of the following requests to the route with the same subscription ID:
curl "http://127.0.0.1:9080/anything" -i -X POST \
-H "X-Subscription-ID: sub-123456789" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
These requests should match the first rule with a default token count of 500. You should see that requests within the quota return HTTP/1.1 200 OK, while those exceeding it return HTTP/1.1 429 Too Many Requests:
HTTP/1.1 200 OK
...
X-AI-1-RateLimit-Limit: 500
X-AI-1-RateLimit-Remaining: 500
X-AI-1-RateLimit-Reset: 60
HTTP/1.1 200 OK
...
X-AI-1-RateLimit-Limit: 500
X-AI-1-RateLimit-Remaining: 344
X-AI-1-RateLimit-Reset: 57.989000082016
HTTP/1.1 429 Too Many Requests
...
X-AI-1-RateLimit-Limit: 500
X-AI-1-RateLimit-Remaining: 0
X-AI-1-RateLimit-Reset: 5.871000051498
Wait for the time window to reset. Send several of the following requests to the route with the same subscription ID and set the X-Custom-Count header to 10:
curl "http://127.0.0.1:9080/anything" -i -X POST \
-H "X-Subscription-ID: sub-123456789" \
-H "X-Custom-Count: 10" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
These requests should match the first rule with a custom token count of 10. You should see that requests within the quota return HTTP/1.1 200 OK, while those exceeding it return HTTP/1.1 429 Too Many Requests:
HTTP/1.1 200 OK
...
X-AI-1-RateLimit-Limit: 10
X-AI-1-RateLimit-Remaining: 10
X-AI-1-RateLimit-Reset: 60
HTTP/1.1 429 Too Many Requests
...
X-AI-1-RateLimit-Limit: 10
X-AI-1-RateLimit-Remaining: 0
X-AI-1-RateLimit-Reset: 40.422000169754
Finally, send several requests without the subscription headers. The following request includes only the trial ID:
curl "http://127.0.0.1:9080/anything" -i -X POST \
-H "X-Trial-ID: trial-123456789" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
These requests should match the second rule with a token count of 50. You should see that requests within the quota return HTTP/1.1 200 OK, while those exceeding it return HTTP/1.1 429 Too Many Requests:
HTTP/1.1 200 OK
...
X-AI-2-RateLimit-Limit: 50
X-AI-2-RateLimit-Remaining: 50
X-AI-2-RateLimit-Reset: 60
HTTP/1.1 429 Too Many Requests
...
X-AI-2-RateLimit-Limit: 50
X-AI-2-RateLimit-Remaining: 0
X-AI-2-RateLimit-Reset: 44
Share Quota Among Gateways with a Redis Server
This example stores rate-limiting counters in one Redis server so multiple gateway instances enforce the same quota.
This example applies to API7 Enterprise version 3.8.19 and later, and to APISIX version 3.18.0 and later.
Start Redis
Start Redis in the environment used by the gateway.
- Docker
- Kubernetes
Set GATEWAY_CONTAINER to the running APISIX or API7 Gateway container. Create a dedicated network and connect the gateway to it:
export GATEWAY_CONTAINER=replace-with-gateway-container-name
docker network create gateway-redis-net
docker network connect gateway-redis-net "$GATEWAY_CONTAINER"
Start Redis on the shared network:
docker run -d \
--name redis-standalone \
--network gateway-redis-net \
redis:8.10.2-alpine \
redis-server --requirepass redis-password
Verify the authenticated connection:
docker exec -e REDISCLI_AUTH=redis-password redis-standalone redis-cli ping
You should receive a response PONG, which shows successful connection.
Set the Redis hostname used by the Admin API and ADC examples:
export REDIS_HOST=redis-standalone
Create a Kubernetes manifest for a Redis deployment:
apiVersion: apps/v1
kind: Deployment
metadata:
namespace: aic
name: redis-standalone
spec:
replicas: 1
selector:
matchLabels:
app: redis-standalone
template:
metadata:
labels:
app: redis-standalone
spec:
containers:
- name: redis
image: redis:8.10.2-alpine
args:
- redis-server
- --requirepass
- redis-password
ports:
- containerPort: 6379
---
apiVersion: v1
kind: Service
metadata:
namespace: aic
name: redis-standalone
spec:
selector:
app: redis-standalone
ports:
- port: 6379
targetPort: 6379
Apply the manifest:
kubectl apply -f redis-standalone.yaml
kubectl rollout status deployment/redis-standalone -n aic
Set the Redis hostname used by the Admin API and ADC examples:
export REDIS_HOST=redis-standalone.aic.svc
Create Route and Configure Rate Limiting
Create a route with the following configurations in the gateway group:
- Admin API
- ADC
- Ingress Controller
curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d @- <<EOF
{
"id": "ai-rate-limiting-redis-route",
"uri": "/anything",
"methods": ["POST"],
"plugins": {
"ai-proxy-multi": {
"instances": [
{
"name": "deepseek-instance",
"provider": "deepseek",
"weight": 8,
"auth": {
"header": {
"Authorization": "Bearer $DEEPSEEK_API_KEY"
}
},
"options": {
"model": "deepseek-flash"
}
},
{
"name": "openai-instance",
"provider": "openai",
"weight": 2,
"auth": {
"header": {
"Authorization": "Bearer $OPENAI_API_KEY"
}
},
"options": {
"model": "gpt-4o-mini"
}
}
]
},
"ai-rate-limiting": {
"instances": [
{
"name": "deepseek-instance",
"limit": 100,
"time_window": 30
},
{
"name": "openai-instance",
"limit": 50,
"time_window": 30
}
],
"limit_strategy": "total_tokens",
"policy": "redis",
"redis_host": "$REDIS_HOST",
"redis_port": 6379,
"redis_password": "redis-password",
"allow_degradation": false,
"rejected_code": 429
}
}
}
EOF
services:
- name: ai-rate-limiting-service
labels:
docs-example: ai-rate-limiting-redis
routes:
- name: ai-rate-limiting-redis-route
uris:
- /anything
methods:
- POST
plugins:
ai-proxy-multi:
instances:
- name: deepseek-instance
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer ${DEEPSEEK_API_KEY}"
options:
model: deepseek-flash
- name: openai-instance
provider: openai
weight: 2
auth:
header:
Authorization: "Bearer ${OPENAI_API_KEY}"
options:
model: gpt-4o-mini
ai-rate-limiting:
instances:
- name: deepseek-instance
limit: 100
time_window: 30
- name: openai-instance
limit: 50
time_window: 30
limit_strategy: total_tokens
policy: redis
redis_host: "${REDIS_HOST}"
redis_port: 6379
redis_password: redis-password
allow_degradation: false
rejected_code: 429
upstream:
type: roundrobin
nodes:
- host: httpbin.org
port: 80
weight: 1
Preview changes owned by this example and confirm that the diff contains no unintended updates or deletions:
adc diff -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limiting-redis
Synchronize the reviewed service configuration:
adc sync -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limiting-redis
- Gateway API
- APISIX CRD
apiVersion: apisix.apache.org/v1alpha1
kind: PluginConfig
metadata:
namespace: aic
name: ai-rate-limiting-redis-plugin-config
spec:
plugins:
- name: ai-proxy-multi
config:
instances:
- name: deepseek-instance
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: openai-instance
provider: openai
weight: 2
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: ai-rate-limiting
config:
instances:
- name: deepseek-instance
limit: 100
time_window: 30
- name: openai-instance
limit: 50
time_window: 30
limit_strategy: total_tokens
policy: redis
redis_host: "redis-standalone.aic.svc"
redis_port: 6379
redis_password: redis-password
allow_degradation: false
rejected_code: 429
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
namespace: aic
name: ai-rate-limiting-redis-route
spec:
parentRefs:
- name: apisix
rules:
- matches:
- path:
type: Exact
value: /anything
method: POST
filters:
- type: ExtensionRef
extensionRef:
group: apisix.apache.org
kind: PluginConfig
name: ai-rate-limiting-redis-plugin-config
apiVersion: apisix.apache.org/v2
kind: ApisixRoute
metadata:
namespace: aic
name: ai-rate-limiting-redis-route
spec:
ingressClassName: apisix
http:
- name: ai-rate-limiting-redis-route
match:
paths:
- /anything
methods:
- POST
plugins:
- name: ai-proxy-multi
enable: true
config:
instances:
- name: deepseek-instance
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: openai-instance
provider: openai
weight: 2
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: ai-rate-limiting
enable: true
config:
instances:
- name: deepseek-instance
limit: 100
time_window: 30
- name: openai-instance
limit: 50
time_window: 30
limit_strategy: total_tokens
policy: redis
redis_host: "redis-standalone.aic.svc"
redis_port: 6379
redis_password: redis-password
allow_degradation: false
rejected_code: 429
Apply the configuration to your cluster:
kubectl apply -f ai-rate-limiting-ic.yaml
❶ policy: Set to redis to use a Redis instance for rate limiting.
❷ redis_host: Set to the Redis hostname reachable from the gateway.
❸ redis_port: Set to Redis instance listening port.
❹ redis_password: Set to the password of the Redis instance, if any.
❺ allow_degradation: Set to false to reject requests if Redis is unavailable.
Verify
Send a POST request to the route with a system prompt and a sample user question in the request body.
curl "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
The response should be HTTP/1.1 200 OK and include rate-limit headers for the selected model instance.
Repeat the same request one at a time. A request assigned to a model instance whose quota is exhausted should return HTTP/1.1 429 Too Many Requests. Because the instances use weighted selection and models report different token usage, successful responses can continue until both quotas are exhausted.
Share Quota Among Gateway Nodes with a Redis Cluster
This example stores rate-limiting counters in a Redis Cluster so multiple gateway nodes share a partitioned, replicated quota store.
This example applies to API7 Enterprise version 3.8.19 and later, and to APISIX version 3.18.0 and later.
Start Redis Cluster
- Docker
- Kubernetes
Set GATEWAY_CONTAINER to the running APISIX or API7 Gateway container. Create a dedicated network and connect the gateway to it:
export GATEWAY_CONTAINER=replace-with-gateway-container-name
docker network create gateway-redis-cluster-net
docker network connect gateway-redis-cluster-net "$GATEWAY_CONTAINER"
Create the following Docker Compose file for three primary nodes and three replicas:
x-redis-node: &redis-node
image: redis:8.10.2-alpine
command:
- redis-server
- --cluster-enabled
- "yes"
- --cluster-config-file
- nodes.conf
- --cluster-node-timeout
- "5000"
- --appendonly
- "no"
- --requirepass
- redis-cluster-password
- --masterauth
- redis-cluster-password
healthcheck:
test: ["CMD", "redis-cli", "-a", "redis-cluster-password", "ping"]
interval: 2s
timeout: 2s
retries: 30
networks:
- gateway-redis-cluster-net
services:
redis-cluster-1:
<<: *redis-node
container_name: redis-cluster-1
redis-cluster-2:
<<: *redis-node
container_name: redis-cluster-2
redis-cluster-3:
<<: *redis-node
container_name: redis-cluster-3
redis-cluster-4:
<<: *redis-node
container_name: redis-cluster-4
redis-cluster-5:
<<: *redis-node
container_name: redis-cluster-5
redis-cluster-6:
<<: *redis-node
container_name: redis-cluster-6
networks:
gateway-redis-cluster-net:
external: true
Start the nodes and wait for their health checks:
docker compose up -d --wait
Initialize the cluster with one replica for each primary:
docker exec -e REDISCLI_AUTH=redis-cluster-password redis-cluster-1 \
redis-cli --cluster create \
redis-cluster-1:6379 \
redis-cluster-2:6379 \
redis-cluster-3:6379 \
redis-cluster-4:6379 \
redis-cluster-5:6379 \
redis-cluster-6:6379 \
--cluster-replicas 1 \
--cluster-yes
The output should end with [OK] All 16384 slots covered.
Verify that all hash slots are available:
docker exec -e REDISCLI_AUTH=redis-cluster-password redis-cluster-1 \
redis-cli cluster info
The output should include cluster_state:ok and cluster_slots_assigned:16384.
Set the Redis node host names used by the Admin API and ADC examples:
export REDIS_CLUSTER_NODE_1=redis-cluster-1
export REDIS_CLUSTER_NODE_2=redis-cluster-2
export REDIS_CLUSTER_NODE_3=redis-cluster-3
export REDIS_CLUSTER_NODE_4=redis-cluster-4
export REDIS_CLUSTER_NODE_5=redis-cluster-5
export REDIS_CLUSTER_NODE_6=redis-cluster-6
Create a Kubernetes manifest for the Redis cluster:
apiVersion: apps/v1
kind: StatefulSet
metadata:
namespace: aic
name: redis-cluster
spec:
serviceName: redis-cluster
replicas: 6
selector:
matchLabels:
app: redis-cluster
template:
metadata:
labels:
app: redis-cluster
spec:
containers:
- name: redis
image: redis:8.10.2-alpine
ports:
- containerPort: 6379
name: client
- containerPort: 16379
name: gossip
command:
- redis-server
- --cluster-enabled
- "yes"
- --cluster-config-file
- nodes.conf
- --cluster-node-timeout
- "5000"
- --appendonly
- "yes"
- --requirepass
- redis-cluster-password
- --masterauth
- redis-cluster-password
volumeMounts:
- name: data
mountPath: /data
volumeClaimTemplates:
- metadata:
name: data
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 1Gi
---
apiVersion: v1
kind: Service
metadata:
namespace: aic
name: redis-cluster
spec:
clusterIP: None
selector:
app: redis-cluster
ports:
- port: 6379
name: client
- port: 16379
name: gossip
Apply the manifest:
kubectl apply -f redis-cluster.yaml
Wait for all pods to be ready:
kubectl rollout status statefulset/redis-cluster -n aic
Initialize the cluster with one replica for each primary:
kubectl exec -n aic redis-cluster-0 \
-- env REDISCLI_AUTH=redis-cluster-password redis-cli \
--cluster create \
redis-cluster-0.redis-cluster.aic.svc:6379 \
redis-cluster-1.redis-cluster.aic.svc:6379 \
redis-cluster-2.redis-cluster.aic.svc:6379 \
redis-cluster-3.redis-cluster.aic.svc:6379 \
redis-cluster-4.redis-cluster.aic.svc:6379 \
redis-cluster-5.redis-cluster.aic.svc:6379 \
--cluster-replicas 1 \
--cluster-yes
Set the Redis node host names used by the Admin API and ADC examples:
export REDIS_CLUSTER_NODE_1=redis-cluster-0.redis-cluster.aic.svc
export REDIS_CLUSTER_NODE_2=redis-cluster-1.redis-cluster.aic.svc
export REDIS_CLUSTER_NODE_3=redis-cluster-2.redis-cluster.aic.svc
export REDIS_CLUSTER_NODE_4=redis-cluster-3.redis-cluster.aic.svc
export REDIS_CLUSTER_NODE_5=redis-cluster-4.redis-cluster.aic.svc
export REDIS_CLUSTER_NODE_6=redis-cluster-5.redis-cluster.aic.svc
Create Route and Configure Rate Limiting
Create a route with the following configurations in the gateway group:
- Admin API
- ADC
- Ingress Controller
curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d @- <<EOF
{
"id": "ai-rate-limiting-redis-cluster-route",
"uri": "/anything",
"methods": ["POST"],
"plugins": {
"ai-proxy-multi": {
"instances": [
{
"name": "deepseek-instance",
"provider": "deepseek",
"weight": 8,
"auth": {
"header": {
"Authorization": "Bearer $DEEPSEEK_API_KEY"
}
},
"options": {
"model": "deepseek-flash"
}
},
{
"name": "openai-instance",
"provider": "openai",
"weight": 2,
"auth": {
"header": {
"Authorization": "Bearer $OPENAI_API_KEY"
}
},
"options": {
"model": "gpt-4o-mini"
}
}
]
},
"ai-rate-limiting": {
"instances": [
{
"name": "deepseek-instance",
"limit": 200,
"time_window": 60
},
{
"name": "openai-instance",
"limit": 100,
"time_window": 60
}
],
"limit_strategy": "total_tokens",
"policy": "redis-cluster",
"redis_cluster_nodes": [
"$REDIS_CLUSTER_NODE_1:6379",
"$REDIS_CLUSTER_NODE_2:6379",
"$REDIS_CLUSTER_NODE_3:6379",
"$REDIS_CLUSTER_NODE_4:6379",
"$REDIS_CLUSTER_NODE_5:6379",
"$REDIS_CLUSTER_NODE_6:6379"
],
"redis_password": "redis-cluster-password",
"redis_cluster_name": "docs-redis-cluster",
"redis_timeout": 1000,
"allow_degradation": false,
"rejected_code": 429
}
}
}
EOF
services:
- name: ai-rate-limiting-service
labels:
docs-example: ai-rate-limiting-redis-cluster
routes:
- name: ai-rate-limiting-redis-cluster-route
uris:
- /anything
methods:
- POST
plugins:
ai-proxy-multi:
instances:
- name: deepseek-instance
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer ${DEEPSEEK_API_KEY}"
options:
model: deepseek-flash
- name: openai-instance
provider: openai
weight: 2
auth:
header:
Authorization: "Bearer ${OPENAI_API_KEY}"
options:
model: gpt-4o-mini
ai-rate-limiting:
instances:
- name: deepseek-instance
limit: 200
time_window: 60
- name: openai-instance
limit: 100
time_window: 60
limit_strategy: total_tokens
policy: redis-cluster
redis_cluster_nodes:
- "${REDIS_CLUSTER_NODE_1}:6379"
- "${REDIS_CLUSTER_NODE_2}:6379"
- "${REDIS_CLUSTER_NODE_3}:6379"
- "${REDIS_CLUSTER_NODE_4}:6379"
- "${REDIS_CLUSTER_NODE_5}:6379"
- "${REDIS_CLUSTER_NODE_6}:6379"
redis_password: "redis-cluster-password"
redis_cluster_name: docs-redis-cluster
redis_timeout: 1000
allow_degradation: false
rejected_code: 429
upstream:
type: roundrobin
nodes:
- host: httpbin.org
port: 80
weight: 1
Preview changes owned by this example and confirm that the diff contains no unintended updates or deletions:
adc diff -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limiting-redis-cluster
Synchronize the reviewed service configuration:
adc sync -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limiting-redis-cluster
- Gateway API
- APISIX CRD
apiVersion: apisix.apache.org/v1alpha1
kind: PluginConfig
metadata:
namespace: aic
name: ai-rate-limiting-redis-cluster-plugin-config
spec:
plugins:
- name: ai-proxy-multi
config:
instances:
- name: deepseek-instance
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: openai-instance
provider: openai
weight: 2
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: ai-rate-limiting
config:
instances:
- name: deepseek-instance
limit: 200
time_window: 60
- name: openai-instance
limit: 100
time_window: 60
limit_strategy: total_tokens
policy: redis-cluster
redis_cluster_nodes:
- "redis-cluster-0.redis-cluster.aic.svc:6379"
- "redis-cluster-1.redis-cluster.aic.svc:6379"
- "redis-cluster-2.redis-cluster.aic.svc:6379"
- "redis-cluster-3.redis-cluster.aic.svc:6379"
- "redis-cluster-4.redis-cluster.aic.svc:6379"
- "redis-cluster-5.redis-cluster.aic.svc:6379"
redis_password: "redis-cluster-password"
redis_cluster_name: docs-redis-cluster
redis_timeout: 1000
allow_degradation: false
rejected_code: 429
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
namespace: aic
name: ai-rate-limiting-redis-cluster-route
spec:
parentRefs:
- name: apisix
rules:
- matches:
- path:
type: Exact
value: /anything
method: POST
filters:
- type: ExtensionRef
extensionRef:
group: apisix.apache.org
kind: PluginConfig
name: ai-rate-limiting-redis-cluster-plugin-config
apiVersion: apisix.apache.org/v2
kind: ApisixRoute
metadata:
namespace: aic
name: ai-rate-limiting-redis-cluster-route
spec:
ingressClassName: apisix
http:
- name: ai-rate-limiting-redis-cluster-route
match:
paths:
- /anything
methods:
- POST
plugins:
- name: ai-proxy-multi
enable: true
config:
instances:
- name: deepseek-instance
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: openai-instance
provider: openai
weight: 2
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: ai-rate-limiting
enable: true
config:
instances:
- name: deepseek-instance
limit: 200
time_window: 60
- name: openai-instance
limit: 100
time_window: 60
limit_strategy: total_tokens
policy: redis-cluster
redis_cluster_nodes:
- "redis-cluster-0.redis-cluster.aic.svc:6379"
- "redis-cluster-1.redis-cluster.aic.svc:6379"
- "redis-cluster-2.redis-cluster.aic.svc:6379"
- "redis-cluster-3.redis-cluster.aic.svc:6379"
- "redis-cluster-4.redis-cluster.aic.svc:6379"
- "redis-cluster-5.redis-cluster.aic.svc:6379"
redis_password: "redis-cluster-password"
redis_cluster_name: docs-redis-cluster
redis_timeout: 1000
allow_degradation: false
rejected_code: 429
Apply the configuration to your cluster:
kubectl apply -f ai-rate-limiting-ic.yaml
❶ policy: Set to redis-cluster to use a Redis cluster for rate limiting.
❷ redis_cluster_nodes: Set to Redis node addresses in the Redis cluster.
❸ redis_password: Set to the password of the Redis cluster, if any.
❹ redis_cluster_name: Set to the Redis cluster name.
Verify
Send a POST request to the route with a system prompt and a sample user question in the request body.
curl "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
The response should be HTTP/1.1 200 OK and include rate-limit headers for the selected model instance.
Repeat the same request one at a time. A request assigned to a model instance whose quota is exhausted should return HTTP/1.1 429 Too Many Requests. Because the instances use weighted selection and models report different token usage, successful responses can continue until both quotas are exhausted.
Share Quota Among Gateway Nodes with Redis Sentinel
Authenticated Redis Sentinel data-node connections were introduced in API7 Gateway 3.10.5 and APISIX 3.18.0.
Use Redis Sentinel when you need automatic failover and high availability but do not require data partitioning. This pattern is simpler to manage and suitable for most high-availability requirements.
The example runs one primary, two replicas, and three Sentinels on a network shared with the gateway.
The redis-sentinel policy does not expose TLS settings for Sentinel discovery or the selected Redis data-node connection. Deploy the gateway, Sentinels, and Redis nodes on a trusted private network.
Start Redis Sentinel
- Docker
- Kubernetes
Set GATEWAY_CONTAINER to the running APISIX or API7 Gateway container. Create a dedicated network and connect the gateway to it:
export GATEWAY_CONTAINER=replace-with-gateway-container-name
docker network create gateway-redis-sentinel-net
docker network connect gateway-redis-sentinel-net "$GATEWAY_CONTAINER"
Start the Redis primary:
docker run -d \
--name redis-master \
--network gateway-redis-sentinel-net \
redis:8.10.2-alpine \
redis-server --requirepass redis-password
Start the first replica:
docker run -d \
--name redis-replica-1 \
--network gateway-redis-sentinel-net \
redis:8.10.2-alpine \
redis-server --replicaof redis-master 6379 \
--masterauth redis-password \
--requirepass redis-password \
--replica-announce-ip redis-replica-1
Start the second replica:
docker run -d \
--name redis-replica-2 \
--network gateway-redis-sentinel-net \
redis:8.10.2-alpine \
redis-server --replicaof redis-master 6379 \
--masterauth redis-password \
--requirepass redis-password \
--replica-announce-ip redis-replica-2
Create a Sentinel configuration that uses Docker DNS names instead of container IP addresses:
port 26379
protected-mode no
sentinel resolve-hostnames yes
sentinel announce-hostnames yes
sentinel monitor mymaster redis-master 6379 2
sentinel auth-pass mymaster redis-password
sentinel down-after-milliseconds mymaster 5000
sentinel failover-timeout mymaster 10000
sentinel parallel-syncs mymaster 1
requirepass sentinel-password
Start the first Sentinel with a writable copy of the configuration:
docker run -d \
--name redis-sentinel-1 \
--network gateway-redis-sentinel-net \
-v "$PWD/sentinel.conf:/etc/redis/sentinel.conf:ro" \
redis:8.10.2-alpine \
sh -c 'cp /etc/redis/sentinel.conf /data/sentinel.conf && redis-sentinel /data/sentinel.conf'
Start the second Sentinel:
docker run -d \
--name redis-sentinel-2 \
--network gateway-redis-sentinel-net \
-v "$PWD/sentinel.conf:/etc/redis/sentinel.conf:ro" \
redis:8.10.2-alpine \
sh -c 'cp /etc/redis/sentinel.conf /data/sentinel.conf && redis-sentinel /data/sentinel.conf'
Start the third Sentinel:
docker run -d \
--name redis-sentinel-3 \
--network gateway-redis-sentinel-net \
-v "$PWD/sentinel.conf:/etc/redis/sentinel.conf:ro" \
redis:8.10.2-alpine \
sh -c 'cp /etc/redis/sentinel.conf /data/sentinel.conf && redis-sentinel /data/sentinel.conf'
Verify that the Sentinel quorum is available:
docker exec -e REDISCLI_AUTH=sentinel-password redis-sentinel-1 \
redis-cli -p 26379 SENTINEL ckquorum mymaster
The command should report OK 3 usable Sentinels.
Set the Sentinel host names used by the Admin API and ADC examples:
export REDIS_SENTINEL_1=redis-sentinel-1
export REDIS_SENTINEL_2=redis-sentinel-2
export REDIS_SENTINEL_3=redis-sentinel-3
Create a Kubernetes manifest for the Redis master, replicas, and Sentinel cluster:
apiVersion: v1
kind: ConfigMap
metadata:
namespace: aic
name: redis-sentinel-config
data:
sentinel.conf: |
port 26379
sentinel resolve-hostnames yes
sentinel announce-hostnames yes
sentinel monitor mymaster redis-master.aic.svc 6379 2
sentinel auth-pass mymaster redis-password
requirepass sentinel-password
sentinel down-after-milliseconds mymaster 5000
sentinel failover-timeout mymaster 10000
sentinel parallel-syncs mymaster 1
protected-mode no
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
namespace: aic
name: redis-master
spec:
serviceName: redis-master
replicas: 1
selector:
matchLabels:
app: redis-master
template:
metadata:
labels:
app: redis-master
spec:
containers:
- name: redis
image: redis:8.10.2-alpine
ports:
- containerPort: 6379
command:
- redis-server
- --requirepass
- redis-password
- --appendonly
- "yes"
---
apiVersion: v1
kind: Service
metadata:
namespace: aic
name: redis-master
spec:
clusterIP: None
selector:
app: redis-master
ports:
- port: 6379
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
namespace: aic
name: redis-replica
spec:
serviceName: redis-replica
replicas: 2
selector:
matchLabels:
app: redis-replica
template:
metadata:
labels:
app: redis-replica
spec:
containers:
- name: redis
image: redis:8.10.2-alpine
ports:
- containerPort: 6379
command:
- redis-server
- --replicaof
- redis-master.aic.svc
- "6379"
- --requirepass
- redis-password
- --masterauth
- redis-password
- --appendonly
- "yes"
---
apiVersion: v1
kind: Service
metadata:
namespace: aic
name: redis-replica
spec:
clusterIP: None
selector:
app: redis-replica
ports:
- port: 6379
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
namespace: aic
name: redis-sentinel
spec:
serviceName: redis-sentinel
replicas: 3
selector:
matchLabels:
app: redis-sentinel
template:
metadata:
labels:
app: redis-sentinel
spec:
initContainers:
- name: copy-config
image: redis:8.10.2-alpine
command:
- sh
- -c
- cp /config/sentinel.conf /data/sentinel.conf
volumeMounts:
- name: sentinel-config
mountPath: /config
- name: sentinel-data
mountPath: /data
containers:
- name: sentinel
image: redis:8.10.2-alpine
ports:
- containerPort: 26379
command:
- redis-sentinel
- /data/sentinel.conf
volumeMounts:
- name: sentinel-data
mountPath: /data
volumes:
- name: sentinel-config
configMap:
name: redis-sentinel-config
- name: sentinel-data
emptyDir: {}
---
apiVersion: v1
kind: Service
metadata:
namespace: aic
name: redis-sentinel
spec:
clusterIP: None
selector:
app: redis-sentinel
ports:
- port: 26379
Apply the manifest:
kubectl apply -f redis-sentinel.yaml
Wait for the primary, replicas, and Sentinels to become ready:
kubectl rollout status statefulset/redis-master -n aic
kubectl rollout status statefulset/redis-replica -n aic
kubectl rollout status statefulset/redis-sentinel -n aic
Set the Sentinel host names used by the Admin API and ADC examples:
export REDIS_SENTINEL_1=redis-sentinel-0.redis-sentinel.aic.svc
export REDIS_SENTINEL_2=redis-sentinel-1.redis-sentinel.aic.svc
export REDIS_SENTINEL_3=redis-sentinel-2.redis-sentinel.aic.svc
Create Route and Configure Rate Limiting
Create a route with the following configurations in the gateway group:
- Admin API
- ADC
- Ingress Controller
curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d @- <<EOF
{
"id": "ai-rate-limiting-redis-sentinel-route",
"uri": "/anything",
"methods": ["POST"],
"plugins": {
"ai-proxy-multi": {
"instances": [
{
"name": "deepseek-instance",
"provider": "deepseek",
"weight": 8,
"auth": {
"header": {
"Authorization": "Bearer $DEEPSEEK_API_KEY"
}
},
"options": {
"model": "deepseek-flash"
}
},
{
"name": "openai-instance",
"provider": "openai",
"weight": 2,
"auth": {
"header": {
"Authorization": "Bearer $OPENAI_API_KEY"
}
},
"options": {
"model": "gpt-4o-mini"
}
}
]
},
"ai-rate-limiting": {
"instances": [
{
"name": "deepseek-instance",
"limit": 100,
"time_window": 60
},
{
"name": "openai-instance",
"limit": 100,
"time_window": 60
}
],
"limit_strategy": "total_tokens",
"policy": "redis-sentinel",
"redis_sentinels": [
{"host": "$REDIS_SENTINEL_1", "port": 26379},
{"host": "$REDIS_SENTINEL_2", "port": 26379},
{"host": "$REDIS_SENTINEL_3", "port": 26379}
],
"redis_master_name": "mymaster",
"redis_password": "redis-password",
"redis_role": "master",
"sentinel_password": "sentinel-password",
"rejected_code": 429
}
}
}
EOF
services:
- name: ai-rate-limiting-service
labels:
docs-example: ai-rate-limiting-redis-sentinel
routes:
- name: ai-rate-limiting-redis-sentinel-route
uris:
- /anything
methods:
- POST
plugins:
ai-proxy-multi:
instances:
- name: deepseek-instance
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer ${DEEPSEEK_API_KEY}"
options:
model: deepseek-flash
- name: openai-instance
provider: openai
weight: 2
auth:
header:
Authorization: "Bearer ${OPENAI_API_KEY}"
options:
model: gpt-4o-mini
ai-rate-limiting:
instances:
- name: deepseek-instance
limit: 100
time_window: 60
- name: openai-instance
limit: 100
time_window: 60
limit_strategy: total_tokens
policy: redis-sentinel
redis_sentinels:
- host: "${REDIS_SENTINEL_1}"
port: 26379
- host: "${REDIS_SENTINEL_2}"
port: 26379
- host: "${REDIS_SENTINEL_3}"
port: 26379
redis_master_name: mymaster
redis_password: redis-password
redis_role: master
sentinel_password: sentinel-password
rejected_code: 429
upstream:
type: roundrobin
nodes:
- host: httpbin.org
port: 80
weight: 1
Preview changes owned by this example and confirm that the diff contains no unintended updates or deletions:
adc diff -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limiting-redis-sentinel
Synchronize the reviewed service configuration:
adc sync -f adc.yaml \
--include-resource-type service \
--label-selector docs-example=ai-rate-limiting-redis-sentinel
- Gateway API
- APISIX CRD
apiVersion: apisix.apache.org/v1alpha1
kind: PluginConfig
metadata:
namespace: aic
name: ai-rate-limiting-redis-sentinel-plugin-config
spec:
plugins:
- name: ai-proxy-multi
config:
instances:
- name: deepseek-instance
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: openai-instance
provider: openai
weight: 2
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: ai-rate-limiting
config:
instances:
- name: deepseek-instance
limit: 100
time_window: 60
- name: openai-instance
limit: 100
time_window: 60
limit_strategy: total_tokens
policy: redis-sentinel
redis_sentinels:
- host: "redis-sentinel-0.redis-sentinel.aic.svc"
port: 26379
- host: "redis-sentinel-1.redis-sentinel.aic.svc"
port: 26379
- host: "redis-sentinel-2.redis-sentinel.aic.svc"
port: 26379
redis_master_name: mymaster
redis_password: redis-password
redis_role: master
sentinel_password: sentinel-password
rejected_code: 429
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
namespace: aic
name: ai-rate-limiting-redis-sentinel-route
spec:
parentRefs:
- name: apisix
rules:
- matches:
- path:
type: Exact
value: /anything
method: POST
filters:
- type: ExtensionRef
extensionRef:
group: apisix.apache.org
kind: PluginConfig
name: ai-rate-limiting-redis-sentinel-plugin-config
apiVersion: apisix.apache.org/v2
kind: ApisixRoute
metadata:
namespace: aic
name: ai-rate-limiting-redis-sentinel-route
spec:
ingressClassName: apisix
http:
- name: ai-rate-limiting-redis-sentinel-route
match:
paths:
- /anything
methods:
- POST
plugins:
- name: ai-proxy-multi
enable: true
config:
instances:
- name: deepseek-instance
provider: deepseek
weight: 8
auth:
header:
Authorization: "Bearer replace-with-deepseek-api-key"
options:
model: deepseek-flash
- name: openai-instance
provider: openai
weight: 2
auth:
header:
Authorization: "Bearer replace-with-openai-api-key"
options:
model: gpt-4o-mini
- name: ai-rate-limiting
enable: true
config:
instances:
- name: deepseek-instance
limit: 100
time_window: 60
- name: openai-instance
limit: 100
time_window: 60
limit_strategy: total_tokens
policy: redis-sentinel
redis_sentinels:
- host: "redis-sentinel-0.redis-sentinel.aic.svc"
port: 26379
- host: "redis-sentinel-1.redis-sentinel.aic.svc"
port: 26379
- host: "redis-sentinel-2.redis-sentinel.aic.svc"
port: 26379
redis_master_name: mymaster
redis_password: redis-password
redis_role: master
sentinel_password: sentinel-password
rejected_code: 429
Apply the configuration to your cluster:
kubectl apply -f ai-rate-limiting-ic.yaml
❶ policy: Set to redis-sentinel to use a Redis in sentinel mode for rate limiting.
❷ redis_sentinels: Configure a list of Sentinel node addresses (host and port).
❸ redis_master_name: Configure the name of the Redis master group that Sentinels are monitoring.
❹ redis_password: Set to the password of the Redis instance, if any.
❺ redis_role: Set to master to connect to the current Redis master.
❻ sentinel_password: Configure the password used to authenticate with Redis Sentinel.
Verify
Send a POST request to the route with a system prompt and a sample user question in the request body.
curl "http://127.0.0.1:9080/anything" -X POST \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a mathematician" },
{ "role": "user", "content": "What is 1+1?" }
]
}'
The response should be HTTP/1.1 200 OK and include rate-limit headers for the selected model instance.
Repeat the same request one at a time. A request assigned to a model instance whose quota is exhausted should return HTTP/1.1 429 Too Many Requests. Because the instances use weighted selection and models report different token usage, successful responses can continue until both quotas are exhausted.