Skip to main content
Version: 1.5.0

Metrics and Usage Events

AISIX AI Gateway exposes aggregate metrics and exportable usage events. Together, these signals show service health, traffic trends, and the model, route, and policy outcome behind each request attempt.

Choose a Telemetry Source​

Start with the source that matches the operational question, then correlate signals by request, model, or provider when an issue needs deeper investigation.

Operational QuestionStart WithWhat It Provides
Is traffic healthy across gateway instances?Prometheus metricsRequest rates, latency distributions, token and cost counters, policy outcomes, routing health, cache behavior, and exporter delivery health.
What happened to one request?Access logsStructured request fields, including status, latency, model, provider, request ID, and routing outcome when available.
What can the calling application observe?Response headersRequest correlation, cache outcome, retry timing, and selected-target hints on supported routes.
Where can request records be stored, analyzed, or used for accounting?Usage eventsPer-attempt outcome and consumption records delivered through an observability exporter.

Scrape Prometheus Metrics​

AISIX serves Prometheus metrics on the dedicated metrics listener at /metrics by default. Change the path or disable the endpoint through the startup observability settings.

This endpoint is unauthenticated by design. Keep the dedicated metrics listener private.

Configure Prometheus exposure in the startup configuration:

config.yaml
observability:
metrics:
prometheus:
enabled: true
path: "/metrics"

The dedicated listener binds to 0.0.0.0:9090 by default. Set a different listener address when Prometheus should scrape another interface or port.

Scrape the default metrics endpoint:

curl -sS "http://127.0.0.1:9090/metrics"
Traffic metrics appear after activity

AISIX publishes configuration status on every scrape. Other metric families are registered on first observation, so traffic metrics might not appear immediately after startup. Send one model request, then check again for series such as aisix_requests_total and aisix_tokens_consumed_total.

AISIX emits native metric names with the aisix_ prefix. Use the histogram series for latency percentiles across gateway instances and the request counters for success-rate and routing analysis. For exact metric names, label scope, and PromQL examples, see Metrics Reference.

Select the labels of individual metric families with observability.metrics.labels. For configuration examples, defaults, and the complete supported-variable list, see Metric Labels and Variables.

Count Requests and Attempts Separately​

A request is one caller interaction. A retry or failover can produce several usage-event attempts for one request. Some attempts stop at a target-level rate limit or during request assembly before calling the provider, so AISIX counts attempts differently depending on the signal:

UnitWhere it is countedWhat one sample or row means
Requestaisix_proxy_requests_total, aisix_llm_requests_totalOne client request, with the status the caller received.
Attemptaisix_deployment_requests_total and the deployment familiesOne upstream call to one target model.
Emission attemptaisix_usage_events_emitted_totalOne attempt to enqueue a usage event, counted before the delivery queue accepts or rejects it.
AttemptUsage-event records, and the usage log built from themOne recorded processing attempt, including one stopped before an upstream call. Sibling attempts share request_id and are ordered by attempt_index.

If one target returns 502 and a fallback succeeds, the request counter records only the 200 returned to the caller. Deployment counters and usage events retain the failed attempt. A usage log can therefore contain many more 5xx attempts than the request metrics show; the two sources are measuring different units.

Use these queries for the questions they actually answer:

# Requests that ended in a server error — what callers experienced.
sum(increase(aisix_proxy_requests_total{status=~"5.."}[24h]))

# Requests that ended in a server error even after failing over.
sum(increase(aisix_proxy_requests_total{status=~"5..", is_fallback="true"}[24h]))

# Upstream attempts that failed, by target — including the ones a fallback
# rescued. This is the attempt-level view of a 5xx usage-log row count,
# restricted to the endpoints that dispatch through a Model Group.
sum(increase(aisix_deployment_failure_responses_total[24h])) by (model)

# Fallbacks that rescued a request, by group and by the target reached.
sum(increase(aisix_routing_successful_fallbacks_total[24h])) by (model, fallback_model)

# Usage-event emission attempts, and those the handoff queue rejected.
sum(increase(aisix_usage_events_emitted_total{status_code="5xx"}[24h]))
sum(increase(aisix_usage_event_drops_total[24h]))

# One member's rate-limit rejections. `status` carries the raw HTTP code,
# so a single failure mode is addressable without scanning a whole family.
sum(increase(aisix_usage_events_emitted_total{user_id="<member-id>", status="429"}[24h]))

# Whose queue handoffs were rejected. The queue-side reasons carry the same
# model and provider-key labels as the emission counter, so the difference
# holds per model, not only in total.
sum(increase(aisix_usage_event_drops_total{reason=~"sink_.*"}[24h])) by (model, provider_key_name)

# Usage the gateway recorded but could not deliver to the control plane. These
# samples carry the member labels only, so group them by member, not by model.
sum(increase(aisix_usage_event_drops_total{reason=~"send_failed|retry_budget_exhausted"}[24h])) by (reason, user_name)

Compare the Same Population​

Before treating two totals as inconsistent, check their coverage:

DifferenceWhy it happensWhat to check
Scrape coverageRequest counters have no environment label. A per-environment usage total is comparable only when Prometheus scrapes every gateway serving that environment.Group the raw counter by job and instance, then compare the contributors with the running gateway instances.
Time coverageincrease(...[24h]) uses only samples present in the range. Restarts or shorter Prometheus retention can leave gaps that are absent from a usage-log query.Graph the raw counter over the same period.
Delivery coverageaisix_usage_event_drops_total counts both events that could not enter the delivery queue (sink_disabled, sink_full, sink_closed) and events the worker accepted but could not deliver to the control plane (send_failed, retry_budget_exhausted). Exporter losses are not counted here.Group queue drops by model, and delivery drops by user_name — they carry no model or provider-key attribution. Monitor telemetry batch failed (events dropped) for control-plane delivery and sink delivery dropped for exporter delivery. Events the control plane rejected inside an accepted batch are counted by aisix_usage_events_rejected_total instead.
No upstream callTarget-level rate limits and request-assembly failures still produce usage events but no deployment counter sample. Assembly failures include unusable credentials, a missing model_name, and a missing or malformed api_base.A misconfigured target produces failed requests with no deployment failures because the provider was never called.

Deployment metric families cover endpoints dispatched through a Model Group, while usage events also cover direct models. Isolate one difference at a time rather than interpreting the entire gap as gateway-side rejection.

Export Usage Events​

Usage events are per-attempt records emitted by supported proxy paths. A request that retries or fails over emits multiple events with the same request_id, ordered by attempt_index. Counting these records therefore counts attempts, not requests — see Count Requests and Attempts Separately before comparing a record count against a request metric. /v1/messages/count_tokens and the single-shot endpoints listed in Supported Endpoints follow the same rule: each failed attempt that led to a retry or failover emits its own zero-token event.

Every event carries occurred_at, the time the gateway recorded it, as an RFC 3339 UTC timestamp with exactly three fractional digits, such as 2026-09-24T12:42:25.123Z. Object storage, Alibaba Cloud SLS, and Datadog receive this string as is, and the OTLP exporter carries the same time in its nanosecond timestamp. The attempts of one request, and events that finish within the same second, can therefore be ordered to the millisecond. RFC 3339 parsers accept the fractional form; a parser that accepts only whole seconds (…:SSZ) needs updating.

Usage events are not read from a local endpoint. Configure an exporter in Observability Exporters to deliver them to OTLP/HTTP, object storage, Alibaba Cloud SLS, or Datadog.

Each event includes its outcome, consumption details, the model alias the caller requested, and the resolved model that served the attempt when the gateway can observe those values. On /v1/chat/completions, every event also carries cache_status (disabled, miss, hit, or bypass); a response served from the response cache records the matching layer in cache_hit_layer (exact or semantic) and, for semantic hits, the matched similarity in cache_similarity. Streaming responses are never cached, so a streamed request's event reports cache_status="disabled" even where the environment has a cache policy enabled.

Latency fields distinguish provider time from caller-visible time:

FieldScope
upstream_latency_msTime spent on one upstream attempt. It excludes request parsing, guardrails, routing, retry delays, and earlier attempts.
upstream_ttft_msTime from the start of an upstream attempt until its first streamed frame of any type — metadata openers such as response.created or a role-only chat delta included, matching what caller-side proxies measure. It is omitted or zero for non-streaming requests, errors, and cache hits.
downstream_latency_msTime from request receipt to downstream delivery: the complete response for non-streaming, the first token or relayed frame for streaming, and the entire stream for A2A and /v1/audio/speech. It includes gateway processing, retries and retry delays, and output holdback, and appears only on the terminal attempt.

A caller that abandons a request still produces a usage event, because the upstream may have done the work and charged for it. Those events carry status 499 and error_class="client_disconnected" rather than 200, so they are filterable as a class; a mid-stream disconnect counts only the tokens that arrived before it. Count successes with status_code = 200 to keep abandoned requests out of the total.

What the Usage Log Records​

A usage event is the record of a request, not a billing line. AISIX files one for every request it could attribute to a caller, including the ones that never reached a provider and the ones with nothing billable, which carry zero tokens. Use this table to tell an absent record from a recorded rejection:

OutcomeRecordedWhat the record carries
SuccessYesStatus 200, the resolved target, and the reported consumption.
Rerank success with no token countYesStatus 200 and zero tokens. Cohere rerank models, for example, report search units rather than tokens, and AISIX does not price search units.
Provider does not support the endpointYesStatus 501, zero tokens, and no upstream call. AISIX answers this itself on text completions, embeddings, image generation, and video submission when the resolved provider lacks the capability.
Cache hitYesStatus 200 with the stored response's token counts replayed, zero cost, cache_status="hit", cache_hit_layer (exact or semantic), and cache_similarity on a semantic hit. It carries no provider_request_id and no routing target, because no upstream was contacted, but provider_model_version does name the model that produced the stored response. Only non-streaming requests are cached, so a cache hit and a streamed response never coincide.
Model not foundYesStatus 404, error_class="model_not_found", and an empty model_id.
Rate-limit rejectionYesStatus 429 and error_class="rate_limit_exceeded". Caller-key, model, and rate-limit-policy rejections all report this, and all three are checked before dispatch.
Budget exhaustedYesStatus 429 and error_class="billing_error".
Guardrail blockYesStatus 422, error_class="content_filter", and guardrail_blocked=true.
Upstream error or timeoutYesOne event per failed attempt, carrying that attempt's attempt_index, attempt_kind, attempt_model, error_class, and error_message. When a later attempt succeeds, its success event is the terminal one; when every attempt fails, the last failed attempt's own event is the terminal one rather than an extra row.
Caller disconnects before the response headYesStatus 499, error_class="client_disconnected", error_message="client closed the request before the response head was written", zero tokens, and zero cost.
Response body dropped before the stream was readYesThe same status and class, with error_message="client closed the request before the response body was streamed". The upstream had answered, so the event names the target that served it.
Caller disconnects mid-streamYesThe same status and class, with error_message="client closed the request while the response was streaming", carrying the tokens delivered before the disconnect.
Upstream fails after a stream's 200 headersYesThe status the same failure gets before the headers, such as 502 or 504, with error_class and error_message taken from the failure. See Streams That Fail After the Response Head.
Authentication failureNoMetrics only. A rejected credential produces no access-log line either, so there is nothing to correlate by request_id.
Request-body parse error or size-limit rejectionNoAccess log and metrics only.
Unauthenticated requestNoNothing is recorded, which is what keeps the health and discovery routes silent.
/livez, /readyz, /v1/models, /v1/videos/{id} and /v1/videos/{id}/content, the OAuth protected-resource discovery routes, and /a2a/{agent}/.well-known/agent-card.jsonNoThese routes emit no usage event at any outcome. A video job's work is metered by its submission.

A cache hit answers from a stored response, so its event describes the entry the caller addressed and never a target:

  • provider_model_version names the model that produced the stored response, read from the cache entry rather than from this request, and matches the value the original call's own event carried. For a Model Group it is the only field on the row that names the producer at all. It is empty only when the stored response carried no model name.
  • provider_request_id stays empty. It is the identifier a customer reconciles against the provider's console, and this request reached no provider.
  • On a hit of a routing or semantic entry, the fields derived from the provider key — provider_kind, provider_featured, branded_provider, pk_label, and byo_label — are empty, and provider reports unknown. A gateway at 1.2.0 or earlier filled them from whichever target the request's routing strategy ranked first, which named a target that never ran and could differ between two hits of the same stored entry. One consequence: filtering Request Logs by a provider-key label no longer returns a group's cache hits, and a per-provider request panel counts them under provider="unknown" beside the cache_status="hit" that explains why.
  • A hit of a direct model keeps that model's own provider, provider key, and upstream model, which are static properties of the model rather than evidence of a dispatch.

The request's access-log line follows the same rules; see Read the Cache Verdict.

A cancelled request reports requested_model as the entry the caller addressed — the group name for a routing request — and model_id as the target it had committed to. model_id is empty only when no attempt had settled and the addressed entry is itself a routing group, because a group prices nothing and its identifier is never written there. A direct model cancelled before any attempt still records its own identifier, the same convention a model_not_found event uses. A cancelled request also carries no guardrail attribution and no token counts from an attempt that had already answered, because both live in response processing the gateway never reached.

Routes where the caller names no model — /mcp, /a2a/{agent}, the passthrough namespace, /v1/realtime before its upgrade, and the files, batches, and fine-tuning surfaces — file the 499 record too. Their model fields are empty and the family's own attribution stands in their place: passthrough_route_name, mcp_server_name with mcp_tool_name, or a2a_agent_name with a2a_method and a2a_operation. The record still carries api_key_id, the caller identity, auth_type, operation, and inbound_protocol.

Streams That Fail After the Response Head​

A streamed response sends its 200 status line before the upstream has produced the answer, so a failure that arrives later cannot change what the client receives. The usage record is written when the stream ends, and it reports how the stream ended rather than the 200 on the client's response line:

How the stream endedRecorded status
The upstream dropped the connection, or sent something AISIX could not decode502
A read from the upstream timed out mid-stream504
The upstream sent an error event inside the stream carrying a 4xx statusThat status. An Anthropic in-band error maps to the status Anthropic documents for its type, so rate_limit_error is recorded as 429.
The upstream sent any other error event inside the stream, including a 5xx status or a Responses API error or response.failed event, which carries no HTTP status502
A bridged /v1/responses upstream stream carried no content, reasoning, tool call, or finish reason502, with error_class="stream_aborted" and a message saying the upstream returned an empty stream
The caller left before the stream ended499 and error_class="client_disconnected"
The stream finished200

error_class and error_message describe the failure itself, even though the client received a 200. When an upstream failure and a disconnect both occur, the upstream failure is recorded, because a relay that passes a transport error on aborts the connection, which looks the same as the caller leaving.

This applies to /v1/chat/completions, including the judge of a streaming ensemble (the panel members' rows stay 200), to /v1/responses and /v1/messages on both the native and translated paths, to streaming /v1/audio/transcriptions, to /v1/audio/speech, and to passthrough routes. A passthrough route that carries a recognized envelope (OpenAI Chat Completions, Completions, or Responses, and Anthropic Messages) reads in-band error events as the typed endpoints do; a raw route has no error envelope to recognize and records only transport failures and timeouts.

Some outcomes are unchanged: a guardrail block keeps its own status, a stream that produced content and then ended without a finish reason is still a 200, and a stream that finishes is a 200.

Because these streams are now recorded as failures, the console's success rate falls by the number of them, and they leave the latency percentiles, which count only 2xx requests. The status and status_code labels on aisix_usage_events_emitted_total move the same way, and the request's access-log line follows its usage event. Cost, budgets, and rate limits do not depend on the status and are unaffected.

A gateway at 1.4.0 or earlier records most of these streams as 200 with empty error_class and error_message. On its translated /v1/messages path an upstream failure is recorded as 499, as is an upstream that drops a streamed transcription.

note

Because these cancellation records did not exist before, an environment's request count now includes abandoned requests and its success ratio can fall after an upgrade. The control plane derives both from usage events. Latency percentiles are unaffected, because they already count only successful requests.

Access Log, Usage Event, and the Logs Page​

The three signals count different things, so a request that retried produces a different number of each:

SignalHow many per requestWhere to read it
Access-log lineExactly one, whatever the outcomeThe gateway's standard error stream
Usage eventOne per attempt, plus the terminal one for the requestAn observability exporter, and the control plane
Logs page rowOne row per usage eventAISIX Cloud Request Logs

Correlate them by request_id, which the caller also receives as x-aisix-request-id. Within one request, order the attempts by attempt_index and read attempt_kind to tell a retry from a failover. Because the counts differ, comparing a row count against a request metric compares two different units — see Count Requests and Attempts Separately first.

On a gateway newer than 1.2.0, a streamed request's access-log line is written at the end of the stream together with the terminal usage event, so the two agree on status, error_kind/error_class, and the failure message.

How Usage Events Reach the Control Plane​

This leg exists in managed mode only. The gateway queues usage events in memory and a worker delivers them to the control plane in batches. The queue holds 16,384 events; when it is full the event being emitted is dropped rather than made to wait, because telemetry must not slow the request path. Each such drop increments aisix_usage_event_drops_total{reason}, with sink_full, sink_closed, or sink_disabled naming the cause, and logs usage event dropped.

The worker flushes whenever it has 100 events or every 5 seconds, whichever comes first, which is why a freshly completed request takes a few seconds to appear. One batch is in flight at a time, and batches go out in order.

A batch that fails to deliver is re-sent, as long as the control plane has signalled that it recognizes a batch it already stored and will not count it twice. The first re-send waits 1 second and the wait doubles per attempt, up to 30 seconds. A control-plane outage therefore no longer costs the usage recorded during it. The first failure logs telemetry batch failed; re-sending (the control plane de-duplicates by batch id), and a batch that lands afterwards logs telemetry batch delivered after re-sending.

Re-sending is bounded, because a usage record that arrives too late no longer reaches the billing window it belongs to:

  • A batch is given up on 30 minutes after its oldest event occurred — not 30 minutes after its first attempt. A batch already older than that when its turn comes is given up on without being sent at all.
  • A batch is also given up on after 8 consecutive failures that the control plane answered. Failures with no answer at all — a connection error or a timeout, which is what an outage produces — do not count towards this limit; the age limit bounds those instead.
  • An answer that says the batch is permanently unacceptable ends it at once, with no re-send.
  • A control plane that does not signal de-duplication gets one attempt per batch, exactly as before re-sending existed.
  • Shutdown makes one final attempt for the batch in hand and does not wait out a backoff.

Every event in a batch that is given up on is counted against aisix_usage_event_drops_total{reason} under send_failed (given up without ever being re-sent) or retry_budget_exhausted (the age limit, or the eight answered failures), and logs telemetry batch failed (events dropped) with reason, attempts, and batch_id. These samples carry the member labels (user_id, user_name) only; model and the provider-key labels report unknown, because the worker holds events rather than the label set the emitting handler had.

New events keep accumulating in the queue while a batch is being re-sent, and the queue is the only thing holding them. So usage is still lost in two cases: a queue that fills up, which drops the newly emitted event under sink_full, and an outage that outlasts the limits above. Alert on any increase of aisix_usage_event_drops_total and break it down by reason to tell the two apart.

The control plane can also accept a batch while rejecting individual events in it, when an event fails its per-field checks — for example, a malformed ID or a status code outside the HTTP range. Those events are dropped rather than re-sent, because the control plane would refuse them again. A gateway newer than 1.4.0 adds their number to aisix_usage_events_rejected_total and logs control plane rejected usage events in an accepted telemetry batch (events dropped) with batch_id, count, and rejected. The counter has no labels, because the control plane answers with a count only, and these events are not counted in aisix_usage_event_drops_total. A gateway at 1.4.0 or earlier treats any successful answer as full delivery.

A missing row is therefore a request that files none — see the table above — an event lost at one of those points, which the drop counter's reason distinguishes, or an event the control plane rejected, which aisix_usage_events_rejected_total counts.

Tell Request Kinds Apart​

Every event carries operation, which identifies the kind of work from the matched endpoint. Neither inbound_protocol nor the model name reliably distinguishes a conversation, image, video, or other operation, so use this field when grouping traffic by request type.

The values form a fixed set suitable for indexing, grouping, and charting:

ValueEndpoint
chat/v1/chat/completions
messages/v1/messages
count_tokens/v1/messages/count_tokens
responses/v1/responses
completions/v1/completions
embeddings/v1/embeddings
rerank/v1/rerank
image_generation/v1/images/generations
image_edit/v1/images/edits
transcription/v1/audio/transcriptions
translation/v1/audio/translations
speech/v1/audio/speech
video_generationPOST /v1/videos
realtime/v1/realtime
files, batches, fine_tuningThe file, batch, and fine-tuning management endpoints
batch_completionThe gateway's own accounting of a finished batch job, recorded when the job completes rather than when it was submitted
mcp, a2aThe MCP and A2A gateways
passthroughA passthrough route

Keep these behaviors in mind when querying the field:

  • It describes the request, not the outcome. A request that failed, or that a guardrail refused, carries the same value a successful one would. It is the only field on such a record that names the endpoint at all.
  • It is request-scoped. A request that retries or fails over emits one event per attempt, and every attempt carries the same value, so counting events per operation counts attempts. See Count Requests and Attempts Separately.
  • Polling a video job is not video generation. Only the submission (POST /v1/videos) produces a usage event; retrieving the job's status or downloading its result does not. A count of video_generation is therefore a count of videos asked for, not of requests made about them.

Because operation is metadata, exporters retain it in metadata_only mode even though they receive no prompt content. In Alibaba Cloud SLS, it arrives as a separate column. Enable analytics for the fields used by a query, then group traffic by operation:

* | SELECT operation, COUNT(*) AS calls, SUM(prompt_tokens + completion_tokens) AS tokens GROUP BY operation ORDER BY calls DESC

To isolate one kind of traffic, filter on the operation directly, for example operation: video_generation. Do not infer it from the model: one model can serve several endpoints, and requested_model identifies the configured alias or group rather than the request type.

Attribute Events to a Member​

Each event uses user_id to identify the organization member who owned the caller API key when the request ran:

SituationRecorded behavior
The caller API key has no memberuser_id is absent. Ownership is assigned explicitly, not inferred from whoever created the key.
One member uses several credentialsEvents can have different api_key_id values and the same user_id. This includes API key and OIDC calls that resolve to different keys owned by one member.
A key is reassigned or deletedExisting events keep the original user_id. Reassignment attributes later requests to the new owner; deletion does not erase earlier attribution.
The event predates gateway support for this fieldNo member is recorded, so member filters cover traffic only from the upgrade forward.

Filter by user_id to include all credentials owned by a member, or by api_key_id to isolate one credential.

In the dashboard, Logs offers this as the Member filter, alongside a Status filter that accepts a family (4xx), an exact code (429), or a range (500-599). Combining the two answers questions like "which of this member's requests were rate-limited in the last 24 hours" in one query. The CSV export carries user_id and the member's name.

Recorded Token Counts​

A usage event records token counts as the upstream reported them. The caller reads a usage block adapted to the protocol it used, while the event keeps the accounting the upstream used, so the two can differ.

The event carries total_tokens, the total the upstream itself reported, verbatim:

UpstreamSource of total_tokens
Gemini on the Vertex AI adaptertotalTokenCount
OpenAI-compatible and Azure OpenAIusage.total_tokens
Responses APIusage.total_tokens
Amazon Bedrock ConversetotalTokens

The field is omitted when the upstream reports no total. Anthropic, for one, never reports a total. The gateway never fills it with a sum of its own. When the gateway combines several usage reports into one figure, such as the usage frames of one stream, the event keeps a total only if every report carried one. A gateway at 1.4.0 or earlier does not record the field.

The counters are not reconciled with each other, so reasoning_tokens can exceed completion_tokens, and cached_prompt_tokens can exceed prompt_tokens. How Reasoning Tokens Are Billed explains how the control plane uses total_tokens to tell the two ways of counting reasoning apart.

For a Gemini thinking model on the Vertex AI adapter, Gemini reports the thinking tokens (thoughtsTokenCount) beside the answer tokens (candidatesTokenCount). The event records completion_tokens as the answer tokens, reasoning_tokens as the thinking tokens, and total_tokens as totalTokenCount, as long as prompt + completion + reasoning tokens equal that total. When Gemini reports no total, or a total those counts do not add up to, the event keeps completion_tokens with the thinking tokens included, as a gateway at 1.4.0 or earlier records it. The caller's response is unchanged either way: an OpenAI-shape response still counts the thinking tokens inside completion_tokens, and so do token rate limits and the Prometheus token metrics. See Google Vertex AI.

On a gateway newer than 1.4.0, /v1/messages and /v1/completions record reasoning_tokens and total_tokens exactly as /v1/chat/completions does, so one upstream call records the same counts whichever of these endpoints addressed it. A gateway at 1.4.0 or earlier records no reasoning count on those two endpoints.

Exporters receive these recorded counts. Object storage and Alibaba Cloud SLS carry the event as is, including total_tokens. OTLP and Datadog report gen_ai.usage.output_tokens under the OpenTelemetry convention instead; see Token Counts in Exported Telemetry.

OpenAI Cache-Write Tokens​

AISIX recognizes cache_write_tokens inside Chat Completions usage.prompt_tokens_details or Responses usage.input_tokens_details. It preserves that raw value as the optional cache_write_tokens field of the usage event, including requests bridged through /v1/messages or /v1/responses and supported protocol-aware passthrough routes. This requires a gateway newer than 1.1.0.

Usage event fields
{
"prompt_tokens": 101,
"completion_tokens": 11,
"cached_prompt_tokens": 19,
"cache_write_tokens": 37
}

The field is omitted when the upstream did not report it; an explicit 0 remains zero. The AISIX Cloud Logs detail, usage-events API, and JSON/CSV export preserve that distinction. CSV uses an empty cell for an absent value. Logs exported through a backend with prefixed field names use its usual prefix, such as aisix.cache_write_tokens in Datadog.

This is separate from Anthropic's additive cache_creation_tokens. It does not increase input/output totals or change the existing cost calculation. In the example above, input plus output is still 112. No existing cache counter is renamed or merged.

Next Steps​

Configure Observability Exporters to send usage events to an external collector, log destination, object store, or warehouse workflow. Use Access Logs and Request Correlation to investigate individual requests, and use the Metrics Reference when building Prometheus dashboards or alerts.