Memory Diagnostics
When a gateway's memory grows, you need to know which part of the process holds it before you can decide whether to add memory, limit traffic, or report a defect. AISIX gives you three sources for that:
- Memory gauges on the Prometheus endpoint. They show how much the allocator and the process hold, the memory limit the process runs under, and how full each in-process store is.
- Heap profiles on demand.
GET /debug/pprof/heapon a separate, loopback-only diagnostics listener returns a profile of the live heap thatgo tool pprofreads. - Heap profiles written automatically as resident memory approaches the memory limit, so the evidence survives the out-of-memory (OOM) kill that usually follows.
All three are on by default and need no control-plane configuration. They are configured in the gateway's own startup configuration, the same way for an open-source gateway and for a gateway connected to AISIX Cloud.
Read the Memory Gauges
The gauges are served with the other metrics on the metrics listener. They are read when Prometheus scrapes the endpoint, never on the request path, and every label has a fixed set of values. See the Metrics Reference for every series.
curl -sS "http://127.0.0.1:9090/metrics" \
| grep -E '^(aisix_allocator_bytes|aisix_memory_limit_bytes|process_resident_memory_bytes|aisix_component_|aisix_runtime_)'
The gauges fall into four groups:
| Metric | What It Shows |
|---|---|
aisix_allocator_bytes{stat} | Byte counts reported by the memory allocator: allocated (live objects), active, resident, mapped, retained, and metadata. allocated climbing means the gateway holds more objects. resident far above allocated means fragmentation or freed pages not yet returned to the operating system. |
process_resident_memory_bytes and the other process_* series | The process as the kernel accounts it, under the standard Prometheus process-collector names, so existing process dashboards and alerts work unchanged. |
aisix_memory_limit_bytes | The container's memory limit, from cgroup v2 memory.max or cgroup v1 memory.limit_in_bytes. The series is absent when the process has no limit. |
aisix_component_entries{component} and aisix_component_bytes{component} | How many entries each in-process store holds, and how many bytes where the store accounts them exactly. |
aisix_runtime_alive_tasks{runtime} and aisix_runtime_global_queue_depth{runtime} | Tasks alive on each async runtime, and tasks queued and not yet picked up. runtime is control, plus tpc-0 through tpc-N for each thread-per-core worker. |
The allocator and heap-profile series are published only by the Linux build in the published AISIX image. The process_* series are published on Linux.
Resident memory as a fraction of the limit is the figure to alert on:
process_resident_memory_bytes / aisix_memory_limit_bytes
Find What Holds the Memory
Start with the component gauges, then compare them with traffic. Each store points at a different cause:
| Signal | Likely Cause | What to Do |
|---|---|---|
aisix_component_bytes{component="in_flight_request_bodies"} rises together with aisix_proxy_in_flight_requests | Many large requests, such as requests with inline images or documents, are held in memory while they wait on slow upstreams. | Size memory for them and bound their concurrency. See Size Memory for Large Request Bodies. |
aisix_component_bytes{component="guardrail_holdback"} is high | Streamed responses held back for an output guardrail. | Review the guardrail's streaming window. See Streaming Output. |
aisix_component_entries{component="metric_series"} keeps growing | Label cardinality: many distinct label values, such as API keys, users, or models, each create new metric series. | Reduce the labels you publish. See Aggregation and Cardinality. |
aisix_component_entries{component="exporter_queue"} or usage_event_queue stays high | A slow or unavailable sink: records queue in memory faster than they are delivered. The exporter label names the exporter. | Check the destination's health and the exporter's drop and failure counters. |
aisix_component_entries{component="log_queue"} stays high | Whatever collects the gateway's standard error stream is not keeping up. | Check the log collector, and aisix_log_lines_dropped_total. |
aisix_component_entries{component="snapshot_pending_reclaim"} stays above 0 | Retired configuration snapshots that long-running requests still hold. Each one holds a whole configuration's worth of memory until those requests finish. | Expected briefly after configuration changes. If it stays up, look for very long streams or stuck requests. |
aisix_runtime_alive_tasks grows while traffic does not | Tasks are leaking. | Take a heap profile and report it with the gauge history. |
The remaining components report entries in stores that are bounded by configuration: response_cache and semantic_cache (entries held in this process, not in Redis), budget_cache, route_embedding_cache, ratelimit_local_keys (rate-limit keys with the in-memory backend, 0 with Redis), and upstream_clients (HTTP clients built for Provider Keys with their own connection settings).
exporter is set only on component="exporter_queue" and is empty for every other component. When an exporter is deleted, its series drops to 0.
If no store accounts for the growth, or aisix_allocator_bytes{stat="allocated"} grows while every component stays flat, take a heap profile.
Take a Heap Profile
The gateway samples heap allocations from startup, so a profile can be taken from a gateway that is already misbehaving, without a restart. On average one allocation in every 2 MiB allocated is recorded with its call stack. The measured overhead is within the noise of the gateway's benchmarks.
Request a profile from the diagnostics listener, which binds 127.0.0.1:9091 by default:
curl -sS -o heap.pb.gz "http://127.0.0.1:9091/debug/pprof/heap"
In Kubernetes, the listener binds the pod's loopback interface. Reach it through a port forward:
kubectl port-forward pod/<gateway-pod> 9091:9091
curl -sS -o heap.pb.gz "http://127.0.0.1:9091/debug/pprof/heap"
The response is a gzipped pprof profile of the memory currently in use. Frames are already resolved to function names in the gateway, so you do not need the gateway binary to read it. The profile carries no file names or line numbers. Read it with go tool pprof:
go tool pprof -top heap.pb.gz
go tool pprof -http=:8080 heap.pb.gz
Taking a profile costs seconds of CPU on a large heap. Only one runs at a time:
| Status | Meaning |
|---|---|
200 | The profile. |
429 | Another profile is being taken. Retry after it finishes. |
501 | Heap sampling is off in this process, or this build does not support it. |
500 | Taking the profile failed. The gateway logs the reason. |
Each profile is counted in aisix_heap_profile_dumps_total{trigger="manual"}, with result set to ok or error.
Recover Heap Profiles After an Out-of-Memory Kill
An OOM kill ends the process before anyone can request a profile, so the gateway writes one on its own as memory approaches the limit. Once a second, it compares resident memory with the memory limit: the container's cgroup limit, or the host's total memory when there is none. When memory crosses a threshold upward, the gateway writes one profile for that threshold. The threshold fires again only after memory has fallen five percentage points below it. With the default thresholds of 0.8 and 0.9, a gateway on its way to an OOM kill leaves a profile at 80% and another at 90% of its limit.
Profiles are written to observability.heap_profiling.auto_dump.dir, /var/lib/aisix/heap by default, and named:
<UTC timestamp>-<host name>-auto-<percent>.pb.gz
For example, 20260928T101502.311Z-aisix-7d9f8-x2k4q-auto-90.pb.gz. In Kubernetes, the host name is the pod name. The gateway keeps the newest keep profiles for its own host name, five by default, and deletes older ones. Profiles written by other hosts that share the directory are left alone. Each profile is counted in aisix_heap_profile_dumps_total{trigger="auto"}, and the gateway logs its path at warn level.
The directory has to outlive the process. In the AISIX Helm chart, /var/lib/aisix is an emptyDir volume, which survives a container restart within the same pod, including one after an OOM kill. Copy the profiles out of the restarted pod:
kubectl exec <gateway-pod> -- ls /var/lib/aisix/heap
kubectl cp <gateway-pod>:/var/lib/aisix/heap ./heap-profiles
An emptyDir volume is deleted with its pod, so copy the profiles before the pod is replaced. Outside the chart, point dir at a volume that survives a container restart.
If the directory cannot be created or written, the gateway logs one warning at startup and writes no profiles automatically. The rest of the gateway is unaffected.
Size Memory for Large Request Bodies
While a request waits on its upstream, the gateway holds about twice its request body: the parsed request, which it keeps for retries and fallback, and the serialized body it sends. For an ensemble model, each panel member in flight adds one more serialized copy, so an ensemble request holds about its body size times one plus the number of panel members.
For requests that carry large bodies, such as inline images or documents, estimate memory per gateway instance as:
memory ≈ baseline + concurrent large requests × body size × 2
For example, 100 concurrent requests that each carry a 20 MB body need about 4 GB above the gateway's baseline while their upstreams are pending. aisix_component_bytes{component="in_flight_request_bodies"} shows the request bytes currently held, and aisix_component_entries for the same component shows how many requests hold them.
To bound this memory, cap how many such requests run at once with a concurrency limit on the caller API key or model. See API Key and Model Rate Limits. With the default in-memory backend, the limit is counted per gateway process, so each instance admits up to the limit. Use the Redis backend to enforce one limit across instances. To reject oversized bodies outright, see Limit Request Body Size.
Configure Memory Diagnostics
The defaults are:
observability:
debug:
enabled: true
addr: "127.0.0.1:9091"
heap_profiling:
auto_dump:
enabled: true
thresholds: [0.8, 0.9]
dir: "/var/lib/aisix/heap"
keep: 5
| Field | Default | Description |
|---|---|---|
observability.debug.enabled | true | Bind the diagnostics listener. Set false to turn it off. |
observability.debug.addr | "127.0.0.1:9091" | Diagnostics listener address. The listener has no authentication, so the gateway logs a warning at startup when this is not a loopback address. If the address cannot be bound, the gateway logs an error and starts without the listener. |
observability.heap_profiling.auto_dump.enabled | true | Write heap profiles automatically as memory approaches the limit. |
observability.heap_profiling.auto_dump.thresholds | [0.8, 0.9] | Fractions of the memory limit, each greater than 0 and at most 1. |
observability.heap_profiling.auto_dump.dir | "/var/lib/aisix/heap" | Directory for automatic profiles, created if missing. |
observability.heap_profiling.auto_dump.keep | 5 | Profiles kept per host name. Must be at least 1. |
Like other startup settings, each field except observability.heap_profiling.auto_dump.thresholds can also be set through an environment variable, such as AISIX_OBSERVABILITY__DEBUG__ENABLED=false. See Environment Variables.
In 1.5.0 and earlier, the gateway cannot read a list from AISIX_OBSERVABILITY__HEAP_PROFILING__AUTO_DUMP__THRESHOLDS, and setting that variable stops the gateway from starting. Set thresholds in the startup configuration file.
Heap sampling itself is not a configuration field. To turn it off without rebuilding, set this environment variable on the gateway process:
_RJEM_MALLOC_CONF=prof_active:false
With sampling off, GET /debug/pprof/heap returns 501 and no profiles are written automatically. The memory gauges are unaffected.
Configure with Helm
The AISIX Helm chart configures the gateway through environment variables. Set these fields with extraEnvVars:
extraEnvVars:
- name: AISIX_OBSERVABILITY__DEBUG__ENABLED
value: "true"
- name: AISIX_OBSERVABILITY__HEAP_PROFILING__AUTO_DUMP__ENABLED
value: "true"
- name: AISIX_OBSERVABILITY__HEAP_PROFILING__AUTO_DUMP__DIR
value: "/var/lib/aisix/heap"
- name: AISIX_OBSERVABILITY__HEAP_PROFILING__AUTO_DUMP__KEEP
value: "10"
thresholds cannot be set through extraEnvVars in 1.5.0 and earlier. Leave it out, and the default thresholds apply.
The chart declares no container port for the diagnostics listener, and the listener binds the pod's loopback interface. Fetch a profile through a port forward to the pod:
kubectl -n <namespace> port-forward pod/<gateway-pod> 9091:9091
curl -sS -o heap.pb.gz "http://127.0.0.1:9091/debug/pprof/heap"
Upgrade Behavior
After upgrading to a gateway version with memory diagnostics, the gateway binds 127.0.0.1:9091 and samples heap allocations without any configuration change. If another process on the host already uses that port, the gateway logs an error and continues to serve traffic without the listener. To opt out of the listener, set observability.debug.enabled: false.
Next Steps
- See the Metrics Reference for every memory series and its labels.
- See Performance and Sizing to size CPU for your request volume.
- See the Port Reference to plan exposure of the diagnostics listener.