Skip to main content
Version: 3.18.0

Operate etcd in Production

APISIX uses etcd as its configuration store in traditional mode and in decoupled deployments whose control and data planes use the etcd configuration provider. The availability and latency of this cluster determine whether operators can publish configuration and whether a new APISIX process can load its initial configuration.

In traditional mode, each APISIX instance manages configuration and proxies traffic. In decoupled mode, control-plane instances expose the Admin API for configuration changes, while data-plane instances load that configuration and proxy traffic. APISIX uses etcd watches, which are subscriptions to configuration changes, to keep its configuration up to date.

A running APISIX process keeps its last successfully loaded configuration in memory if etcd becomes unavailable. Existing proxy paths can therefore continue to serve traffic while APISIX reconnects, but configuration writes and synchronization stop. A new or restarted process cannot become ready until it loads configuration from an available endpoint. Capabilities that depend on other runtime systems, such as service discovery or external rate-limit storage, have their own failure behavior.

This document applies to an etcd cluster dedicated to APISIX. It does not apply to standalone mode, which stores configuration in a file or in memory instead of etcd. Do not switch a production deployment to standalone mode as an improvised incident response unless that transition and its rollback have been planned and rehearsed.

Production Baseline​

Before using an etcd-backed APISIX deployment in production, establish the following baseline:

AreaRequirement
Version and platformPin the exact etcd release or image digest. Record and test the complete APISIX, etcd, operating-system, and deployment-platform combination.
Cluster topologyUse three voting members in separate failure domains with low and predictable member-to-member latency.
Storage and capacityUse SSD-backed storage. Size the cluster after testing representative APISIX connections, watches, configuration writes, failures, and maintenance operations.
SecurityEncrypt client and peer traffic with TLS, enable authentication, and restrict credentials to the configured prefix. Grant read-write access to traditional instances and control planes, and read-and-watch access to data planes.
Production acceptanceComplete the deployment sequence, read-only inspection, failure test, backup test, and production checklist before directing traffic to the deployment.
MonitoringMonitor etcd quorum and storage together with APISIX reachability and configuration propagation. Alert on failures that require immediate operator action.
MaintenanceEnable automatic compaction. Review maintenance on a defined cadence, and run disruptive work on one member at a time with health checks between members.
RecoveryEncrypt snapshots, retain at least one copy outside the cluster's failure domain, and rehearse restores. A learner or additional member is not a backup.

Understand Failure Behavior​

Plan recovery from the effect on configuration and traffic, not only from the etcd process state:

ConditionConfiguration operationsExisting proxy trafficNew or restarted APISIX process
One member is unavailable, quorum remainsReads, writes, and watches continue through the remaining members.Continues normally.Can load configuration from an available endpoint.
Quorum is lost or every endpoint is unreachableAdmin API writes and configuration synchronization stop.Continues with the last configuration loaded in memory.Keep it out of service until quorum is restored and the initial load is verified.
A watch revision has been compactedAPISIX performs a full read and resumes watching from the current state.Continues with the loaded configuration while APISIX reconnects.Loads the current state instead of replaying old revisions.
The cluster is restored from an older snapshotThe stored configuration moves back to the snapshot state.A running process can temporarily hold a newer in-memory state until it resynchronizes.Loads the restored state.
caution

During an etcd outage, do not restart the APISIX instances. The configuration cache exists only in process memory, so restarting discards the configuration that is sustaining existing proxy traffic. A restarted process cannot become ready until it loads configuration from etcd.

Plan and Configure a Production Deployment​

Make the version, platform, topology, capacity, security, and client configuration decisions together. Each decision affects the failure behavior and maintenance procedures that the deployment must support.

Select a Tested Version​

Pin the exact etcd release or image digest used in production. Record it with the APISIX version, client configuration, operating system, deployment platform, validation date, and scenarios tested. APISIX enforces a minimum etcd version at startup, but that minimum version requirement is not a production recommendation. Use a release supported by the etcd project and repeat the production-readiness tests before changing either APISIX or etcd.

Choose a Deployment Platform​

A dedicated etcd deployment can run on hosts, virtual machines, Kubernetes, or a compatible managed service. The platform must provide predictable storage latency, failure-domain separation, backup access, certificate and credential control, and the ability to maintain one member at a time.

For a managed service, verify API and version compatibility, TLS and role controls, snapshot export, stated recovery objectives, provider maintenance behavior, and cross-zone latency. Confirm which member and storage operations remain under the provider's control before adopting the service.

On dedicated hosts or virtual machines, use a service manager such as systemd to supervise etcd and retain its logs. Confirm that members do not share a power supply or underlying storage volume that could fail together.

Do not store APISIX configuration in the etcd cluster used internally by a Kubernetes control plane. Sharing that cluster couples APISIX capacity, permissions, maintenance, and failure recovery to Kubernetes.

Installation examples that start one local etcd member are evaluation topologies. Replace them with a dedicated, persistent, monitored cluster before using APISIX in production.

Plan the Cluster​

Start with three voting members, each running an etcd server that participates in cluster decisions. Quorum is the majority of voting members needed to agree on updates: two members in a three-member cluster. This topology tolerates one member failure. Place members in separate failure domains, such as independent hosts, racks, or availability zones, while keeping network latency between them low and predictable.

The recommended topology keeps client access, peer traffic, monitoring, and backups within explicit trust and failure boundaries:

APISIX clients connect with TLS or mTLS to authorized endpoints on port 2379; reserve 2380 for member-to-member traffic. Monitoring covers both etcd health and APISIX configuration synchronization. Encrypt snapshots, keep them outside the cluster's failure domain, and rehearse restores.

Keep the members in one region unless a measured recovery objective requires a wider topology. Every committed write requires a quorum, so cross-region latency directly affects configuration writes and cluster stability. Use replicated backups and a tested recovery environment for region-level disaster recovery instead of extending a quorum across distant regions by default.

Size from the Configuration Workload​

Collect the following inputs before selecting compute and storage:

InputWhat to measure
APISIX processesControl-plane and data-plane instance count, worker count, etcd connections, and watches
ConfigurationKey count, total keyspace size, average and largest value sizes, and largest configuration transaction
Change rateTypical and peak writes, revision growth, bulk configuration activity, and background writes from plugins or controllers
GrowthExpected configuration and client growth over the capacity-planning period
InfrastructureMember-to-member latency, packet loss, disk sync latency, and storage throughput
Recovery objectivesMaximum acceptable configuration data loss (recovery point objective, RPO), maximum recovery duration (recovery time objective, RTO), and maintenance window

Fast and consistent disk writes matter more than headline disk capacity. Use SSD-backed storage, monitor write latency on the actual volume, and avoid sharing its I/O path with log processing, batch jobs, or other variable workloads. The etcd hardware guidance describes two to four CPU cores and about 8 GB of memory as a typical starting point. A cluster with many clients, watches, or keys needs workload-specific testing.

Validate the selected capacity with representative configuration writes and watches. Include a member failure, leader election, snapshot, compaction, and defrag operation in the test. At peak load, the cluster should have no sustained backlog of proposed etcd updates or repeated leader elections. It should continue to meet the configuration propagation objective after any one member is stopped.

For heavy workloads, the etcd hardware guidance gives eight to sixteen dedicated CPU cores and 16 to 64 GB of memory as starting ranges. These are sizing examples, not capacity guarantees. Reserve memory so etcd does not depend on swapping during normal or peak operation, and leave room for expected growth.

Secure the Cluster​

Restrict the client port to APISIX and authorized operators, and restrict the peer port to etcd members. Do not expose client, peer, health, or metrics endpoints directly to the internet.

Enable TLS for client and peer traffic. Require and validate client certificates when your identity model supports mutual TLS, and keep APISIX certificate verification enabled. See Configure mTLS between APISIX and etcd for an APISIX example.

Enable etcd authentication and role-based access control (RBAC). Restrict each role to the configured etcd key prefix, such as /apisix, under which APISIX stores its configuration:

Instance or operatoretcd access
Traditional instanceRead and write the configured prefix
Control planeRead and write the configured prefix
Data planeRead and watch the configured prefix
Backup or maintenance operatorGrant only the cluster-wide operations required by the recovery procedure

Use separate accounts for backup automation and member maintenance. Restrict monitoring access to the health and metrics endpoints it needs; monitoring should not require permission to modify configuration or membership.

Configure an etcd username and password for APISIX to authenticate to etcd even when the connection also uses a client certificate. APISIX performs its startup requests through the etcd gRPC gateway, which cannot use the client certificate Common Name as the etcd user. When combining these mechanisms, issue the APISIX client certificate without a Common Name. Use the certificate to authenticate the TLS connection and the configured etcd user to authorize prefix access.

Use separate credentials for control planes, data planes, and maintenance automation. Use a different prefix and credential set for each environment so that a staging operation cannot read or modify production configuration.

etcd does not encrypt key-value data on disk. Use encrypted storage where configuration or secrets require encryption at rest, and protect snapshots to the same standard as the live datastore.

Configure etcd Safely​

Start from the etcd defaults for timing and request limits, then change them only with measurements from the production network and workload. Use one configuration source and record the effective configuration. When etcd is started with a configuration file, it ignores command-line flags and environment variables rather than merging the sources.

SettingProduction guidance
strict-reconfig-checkKeep it enabled so a membership change that could cause quorum loss is rejected.
auto-compaction-mode and auto-compaction-retentionEnable automatic compaction and select a retention window from measured revision growth and the longest expected APISIX disconnection. See Set Compaction Retention.
quota-backend-bytesSet an explicit quota below available storage capacity, leaving room for the write-ahead log (WAL), snapshots, temporary maintenance work, and the operating system.
max-request-bytesKeep the default 1.5 MiB (1572864 bytes). Reduce oversized objects or split transactions where atomicity is not required before raising the limit. Larger requests can delay other requests; test any increase at peak load and during a single-member failure.
Heartbeat and election timeoutsKeep the defaults unless measured member latency and packet loss justify a coordinated change across all members.
Corruption checksEnable the checks supported by the selected etcd release and account for their I/O cost during capacity validation.
unsafe-no-fsyncKeep it disabled. Enabling it can acknowledge writes before they are durably persisted.

Apply the same reviewed baseline to every member. Treat a difference in effective member configuration as drift and correct it through the deployment system rather than with an undocumented manual change.

Configure APISIX Endpoints​

List multiple endpoints from the same etcd cluster so APISIX can connect directly to another member when one endpoint is unavailable. A data plane configuration can start with the following settings:

config.yaml
apisix:
ssl:
ssl_trusted_certificate: /run/secrets/etcd-ca.crt
deployment:
role: data_plane
role_data_plane:
config_provider: etcd
etcd:
host:
- "https://etcd-1.example.net:2379"
- "https://etcd-2.example.net:2379"
- "https://etcd-3.example.net:2379"
prefix: /apisix
timeout: 30
watch_timeout: 50
resync_delay: 5
health_check_timeout: 10
startup_retry: 2
user: "${{ETCD_USER}}"
password: "${{ETCD_PASSWORD}}"
tls:
cert: /run/secrets/etcd-client.crt
key: /run/secrets/etcd-client.key
verify: true

The timeout, watch, resynchronization, health-check, and startup-retry values shown are APISIX defaults. Keep them as the initial baseline and change them only after measuring network and failure behavior. Increasing a timeout can reduce sensitivity to transient latency, but it also delays failure detection.

Use a read-only credential for this data-plane example. A traditional instance or control plane that serves the Admin API needs a credential with write access to the same prefix.

If a load balancer fronts etcd, make the load balancer highly available and remove members that are not ready. Listing member endpoints directly avoids introducing that additional dependency.

Deploy and Validate​

Keep these checks together as the production acceptance gate. Do not infer readiness from a healthy etcd endpoint alone.

Follow the Deployment Sequence​

  1. Provision the three failure domains, persistent storage, network rules, DNS names, and time synchronization.
  2. Issue peer and client certificates and prepare protected storage for credentials.
  3. Start the etcd members with client and peer TLS on the restricted network. Verify membership, leadership, alarms, and endpoint health, then create users and least-privilege roles and enable authentication before connecting APISIX.
  4. Activate monitoring and alerts, then create, verify, and copy the initial snapshot outside the cluster's failure domain.
  5. Connect a traditional instance or control-plane instance and verify that it can create, update, and delete a test configuration through the Admin API.
  6. Connect one data-plane instance as a canary, a test instance used to validate the deployment before connecting the remaining instances. Verify configuration synchronization and proxy traffic.
  7. Exercise a follower failure, then confirm that the recovered member catches up without persistent alarms and that configuration operations continue to meet the propagation objective.
  8. Connect the remaining APISIX processes in phases and direct production traffic only after every checklist item passes.

Run Read-Only Cluster Checks​

Use read-only checks before a deployment, maintenance window, or incident response. Supply the TLS and authentication values through protected environment variables or secret files so they do not remain in shell history.

etcdctl --endpoints="${ETCD_ENDPOINTS}" endpoint health --cluster
etcdctl --endpoints="${ETCD_ENDPOINTS}" endpoint status --cluster -w table
etcdctl --endpoints="${ETCD_ENDPOINTS}" member list -w table
etcdctl --endpoints="${ETCD_ENDPOINTS}" alarm list

Complete the Production Checklist​

RequirementEvidence to retainIf the requirement is not met
Three voting members are distributed across independent failure domains.Member list, advertised peer URLs, placement record, and a follower-failure test in which the recovered member catches up without persistent alarms.Do not send production traffic. Correct placement or membership first.
Every endpoint is ready and reports the expected cluster ID and peer URLs, exactly one leader, and no persistent alarms.endpoint health, endpoint status, member list, and alarm list output from the acceptance window.Correct membership, networking, storage, or alarms before connecting APISIX.
APISIX and etcd versions are pinned and the combination has passed the workload test.Version record and capacity-test result, including one-member failure.Pin the versions and repeat the test.
TLS, authentication, and least-privilege roles are enforced.Certificate validation, a failed data-plane write, and a denied request outside the configured prefix using the control-plane or traditional-instance credential.Close network exposure and correct identity or permissions.
Storage and network performance meet the propagation objective at peak load.WAL and backend latency, member latency, backlog of proposed etcd updates, leader stability, and APISIX propagation measurements.Resize or move the workload; do not hide the failure by increasing timeouts alone.
APISIX can create, update, and delete a test configuration, and every data plane applies each change within the propagation objective.Admin API results, per-process modification indexes, propagation times, and successful proxy requests.Keep the deployment out of service and correct the APISIX-to-etcd path.
Compaction, quota, and sequential defrag procedures are defined.Effective settings, measured revision growth, and a completed maintenance record.Establish the retention and maintenance procedure before growth reaches the quota.
Monitoring covers both etcd and APISIX synchronization.Alert tests for quorum, storage, certificate, backup, and propagation failures.Add and test the missing signal.
Snapshots meet the recovery objective and a recent isolated restore succeeded.Snapshot hash, off-cluster copy, restore date, recovered revision, RPO, and RTO.Treat recovery as unproven and complete a restore drill.
The incident procedure preserves running APISIX processes and prohibits unsafe shortcuts.Reviewed procedure and drill record.Correct the procedure before handing the service to on-call operators.

Monitor and Maintain the Deployment​

Operate etcd and APISIX as one configuration storage and synchronization system. Monitor both systems continuously, then perform maintenance from measured cluster state rather than from fixed thresholds alone.

Monitor the End-to-End Path​

Monitor both etcd and APISIX. A healthy etcd process alone does not prove that every APISIX process can authenticate, watch its prefix, and apply a configuration change.

Track at least these signals:

LayerSignals
Host and processCPU saturation and throttling, memory pressure and out-of-memory (OOM) events, process restarts, and storage I/O contention
ClusterReady members, leader presence and changes, failed and pending etcd updates, and member-to-member latency
StorageWAL sync latency, database commit latency, database size, quota usage, free disk space, and reclaimable space
APISIXapisix_etcd_reachable, apisix_etcd_modify_indexes, configuration errors, and the last successful test change
RecoverySnapshot age, snapshot verification result, restore-drill age, and certificate expiration

The APISIX etcd metrics are available when the Prometheus plugin exports them. Compare modification indexes across APISIX processes to detect an APISIX instance that is no longer receiving updates.

Compare the same resource labels across instances that watch the same prefix. Different resource types or prefixes can legitimately have different indexes; use a known test change to check whether each relevant instance receives the expected update.

When the APISIX status server is enabled, /status/ready confirms that each worker has received usable configuration. After the initial synchronization, it can remain ready while etcd is temporarily unavailable because the workers still have their cached configuration. Use the reachability metric, logs, and a controlled configuration change to monitor the live etcd path.

Alert on Immediate Failures​

Set alert thresholds from a measured baseline and the deployment's recovery objectives. Calibrate latency, resource, and growth alerts after observing normal and peak workloads.

The following values are candidate alert rules to validate during load tests and an initial two-to-four-week observation period. They are operational examples, not etcd limits or guaranteed detection times. Adjust them to the configuration propagation objective and on-call response time. Immediate failures listed below take precedence over these delays.

SignalExample warningExample critical condition
Member healthAn endpoint health check fails, /readyz fails, or the monitoring target reports up=0.Available voting members fall below quorum, or /livez fails.
Leader changesMore than three changes in 15 minutes.No leader; investigate immediately rather than waiting for a change-count threshold.
Pending or failed updatesPending proposals remain nonzero or proposal failures increase for five minutes.The gap between committed and applied proposals keeps growing and requests fail.
Disk latencyWAL sync P99 exceeds 10 ms or backend commit P99 exceeds 25 ms for ten minutes.WAL sync P99 exceeds 50 ms or backend commit P99 exceeds 100 ms for five minutes, or requests already time out.
Backend quotaDatabase size reaches 70% of quota.Database size reaches 85% of quota; a NOSPACE alarm requires immediate action.
APISIX synchronizationAn instance reports apisix_etcd_reachable=0 for 30 seconds, or comparable configuration modification indexes do not converge within two minutes after a test change.A control plane or at least 20% of data planes remain disconnected for one minute, or synchronization takes more than five minutes. Use shorter delays if the propagation objective requires them.
CertificatesA certificate expires within 30 days.A certificate expires within seven days; adjust both windows to the time needed for rotation.
BackupsSnapshot age exceeds the recovery point objective.Two consecutive backups fail or no verified snapshot exists. Investigate the first failed or overdue backup immediately.

Set CPU, memory, and free-disk thresholds from the allocated resources, measured peaks, growth rate, and time needed to add capacity. Absolute resource thresholds from a different deployment can produce misleading alerts.

caution

Treat loss of quorum, no leader, an OOM event, repeated process termination, a NOSPACE alarm, a failed or overdue backup, and inability to propagate configuration as immediate operational failures.

Set Compaction Retention​

An etcd revision is the increasing version number assigned to changes in its key-value data. Enable automatic history compaction. APISIX watches configuration revisions and automatically performs a full read if the requested revision has already been compacted, but a retention window that covers routine disconnections avoids unnecessary full reloads.

The default auto-compaction-retention value is 0, which disables automatic compaction. Selecting a compaction mode alone does not enable it.

For an APISIX cluster with infrequent configuration changes, a periodic retention window of 24 to 72 hours is a practical starting point. Increase it when APISIX processes can remain disconnected longer, or when a full reload is expensive. Shorten it when revision growth threatens the storage objective and routine disconnects are brief. Validate the decision from measured revision growth rather than copying a revision count between environments.

If you select revision-mode compaction, calculate retention from the measured peak revision rate rather than using a fixed revision count:

retained revisions
>= peak revisions per second × maximum expected disconnection in seconds × safety factor

Choose the safety factor from workload variability and recovery risk, then confirm the result by disconnecting an APISIX canary for the target interval.

For example, a peak rate of two revisions per second, a six-hour disconnection window, and a safety factor of two require 2 × 21600 × 2 = 86400 retained revisions. At that peak rate, 1,000 revisions cover only 500 seconds, about eight minutes. These inputs illustrate the calculation; replace them with measurements from your deployment.

Reclaim Space on One Member at a Time​

Compaction makes old revisions unavailable but does not return the freed space to the filesystem. Run etcdctl defrag when reclaimable space is material or disk growth approaches an alert threshold. An online defrag operation blocks reads and writes on the member being processed. Run it during a low-traffic period, process one member at a time, and verify cluster and APISIX health before continuing.

An example maintenance trigger is reclaimable space consistently exceeding 30% of the database file size. Treat this as a candidate policy rather than an etcd limit: consider the absolute space recoverable, available disk, and the cost of blocking each member before scheduling the operation.

Recover from a NOSPACE Alarm​

Set a quota for the etcd backend database that fits within the available disk and leaves space for the write-ahead log, snapshots, temporary maintenance work, and the operating system. When the backend exceeds its quota, etcd raises a cluster-wide NOSPACE alarm and stops accepting ordinary writes. Increasing the quota does not replace compaction, defrag, or investigation of a client generating unexpected writes.

To recover from NOSPACE:

  1. Stop or isolate the client generating unexpected writes and pause nonessential APISIX configuration changes.
  2. Confirm the current alarm, backend size, in-use size, and latest revision.
  3. Compact obsolete history.
  4. Run etcdctl defrag and verify one member at a time.
  5. Clear the alarm only after the cluster is below its quota.
  6. Test an APISIX configuration write and confirm that every process receives it.
  7. Correct the cause by adjusting retention, the source of writes, capacity, or alert timing. Confirm that revision growth returns to its expected rate and storage use falls below the warning threshold before closing the incident.

Replace Members and Rotate Certificates Safely​

A learner is non-voting and does not increase quorum fault tolerance until it is promoted. For a failed or retiring member, add the replacement as a learner and wait for it to catch up. Promote it and remove the old member only after the cluster is healthy. Do not reuse the removed member's data directory as the data directory of a new member. Change one member at a time and verify APISIX configuration propagation between changes.

Rotate trust without creating a point where old and new certificates cannot communicate. Add the new trust root, then rotate peer and client identities in phases. Remove the old trust root only after verifying that no active client or member depends on it. Keep certificate ownership and expiration dates in the operational record.

Follow a Maintenance Schedule​

Assign an operator to each scheduled check. Increase the frequency when the recovery objective, change rate, or risk profile requires it.

FrequencyOperator review
ContinuousQuorum, leader, proposal and storage latency, backend quota, APISIX reachability and propagation, snapshot age, and certificate expiration.
DailyScheduled snapshot success and off-cluster copy; urgent alarms; unexpected revision or database growth.
WeeklyMember health, leader changes, slow requests, capacity trends, backup verification, and unresolved alerts.
MonthlyReclaimable space and whether sequential defrag is justified; version support; time remaining before certificates expire; least-privilege access.
QuarterlyIn an isolated environment, restore a snapshot and exercise both single-member failure and quorum loss; record actual recovery point and recovery time.
Event-triggeredRepeat the affected checks after an APISIX or etcd release, security advisory, certificate or membership change, major configuration migration, infrastructure change, or incident.

Do not run defrag, rotate certificates, or replace members only because a calendar interval elapsed. Use the review to decide whether the action is needed, then follow the documented one-member-at-a-time procedure.

Maintain Operational Records​

Keep the information needed to make a safe decision during a maintenance window or incident:

  • the APISIX and etcd versions, effective etcd configuration, member placement, and certificate ownership;
  • the etcd user and role inventory, including authorized prefixes and permissions;
  • the production topology, failure domains, DNS names, ports, and network access rules;
  • current recovery point and recovery time objectives, snapshot location and retention, and the most recent restore evidence;
  • capacity baseline, expected configuration growth, alert thresholds, and the reason for each changed setting;
  • approved procedures for member replacement, defrag, certificate rotation, restore, and upgrade;
  • the Prometheus dashboard, alert rules, on-call owner, and escalation path; and
  • a record of maintenance, drills, incidents, known limitations, exceptions, and their follow-up actions, including the owner and next review date.

Back Up and Rehearse Recovery​

Treat recovery as a tested capability, not as the existence of a snapshot job. Define the backup policy first, then prove that a verified snapshot can restore APISIX configuration and traffic.

Define the Backup Policy​

Set the snapshot schedule from the configuration recovery point objective. The policy should define evidence that operators can review, rather than only state that a backup job exists:

Policy itemRequirement
ScheduleThe interval meets the configuration recovery point objective under the measured write rate.
RetentionThe retention window covers the time needed to detect accidental deletion or a bad configuration change.
StorageUse the 3-2-1 pattern as the recommended organizational starting point: three copies, two different storage media types, and one copy at a separate site outside the cluster's failure domain. Encrypt every copy and restrict access.
ValidationEach snapshot records its status, hash, revision, total keys, and total size, and failed or overdue snapshots alert an operator.
Restore evidenceA scheduled isolated restore records the recovered revision, actual recovery point, actual recovery time, and APISIX validation result.

An example backup policy takes hourly snapshots and retains the latest 24 hourly copies plus 14 to 30 daily copies. Use it only when snapshot age, job duration, and possible failures fit the recovery point objective. Increase frequency for a stricter objective, and extend retention for deletion-discovery time, audit requirements, or recovery needs.

Inspect every snapshot before copying or restoring it. The status output supplies the metadata required by the validation record:

etcdutl snapshot status snapshot.db -w table

Save snapshots through the maintenance API. Do not use a copy of the active backend database as the routine backup method because it can omit data that still exists only in the write-ahead log. The etcd disaster recovery guidance explains the snapshot and restore model. A learner or additional member is not a backup because it replicates accidental deletion and logically invalid configuration changes.

Rehearse Restores​

Restore every member from the same verified snapshot into a new data directory to create a new logical cluster. Restore in an isolated network first, then verify the cluster revision, key count, routes, consumers, SSL certificates, upstreams, Admin API operations, and real proxy traffic before reconnecting all APISIX processes.

Restoring an older snapshot moves configuration state backward. Because APISIX uses watches and an in-memory configuration cache, calculate a revision bump that covers the revisions that might have occurred since the snapshot. Pass that value with --bump-revision and use --mark-compacted during the restore. This invalidates existing watches so APISIX enters its compacted-watch recovery path and performs a full read. Test the exact restore procedure for the selected etcd release, and do not copy an arbitrary bump value from an example.

See Back Up and Restore etcd for the snapshot workflow. Extend that workflow with off-cluster retention and a scheduled restore exercise that measures the actual recovery point and recovery time.

Respond to an etcd Failure​

When every endpoint is unavailable or the cluster loses quorum, freeze configuration changes and keep healthy APISIX processes running.

caution

Do not take these actions as recovery shortcuts:

  • Delete a data directory when no verified snapshot exists.
  • Restart or replace multiple etcd members at once.
  • Use --force-new-cluster while members of the old cluster are still running.
  • Reconnect a restored cluster to all APISIX processes before isolated validation.

Preserve Evidence and Contain the Impact​

Capture the incident start time, member list, endpoint status, alarms, storage and Raft metrics, APISIX errors, and recent infrastructure or configuration changes before taking disruptive action. Pause automated configuration delivery and infrastructure reconciliation that could change the evidence or restart healthy APISIX processes.

Keep serving APISIX processes in place while you determine whether the failure is in the network, credentials, individual members, or quorum. Record every recovery action and its result so that the next operator does not repeat a disruptive step.

Choose the Recovery Path​

The recovery path depends on whether the cluster still has quorum. If quorum is lost, distinguish a temporary outage from permanent member loss before restoring a snapshot. Keep configuration changes frozen and preserve running APISIX processes throughout recovery.

If Quorum Remains​

Restore connectivity or replace one failed member at a time. Add a replacement as a learner, wait for it to catch up, promote it, and only then remove the old member. Verify cluster health and an APISIX configuration change before working on another member.

If Quorum Is Permanently Lost​

Recover a new cluster from the latest verified snapshot by following the recovery procedure for the selected etcd version. Keep the old members isolated, validate the restored state with one APISIX canary, and reconnect the remaining processes in phases.

If etcd Is Healthy but APISIX Is Stale​

Check endpoint reachability, certificate validation, authentication, prefix permissions, compacted-watch messages, and modification indexes before restarting a process. Compare the affected process with a healthy peer and preserve logs that show the last successful synchronization.

Use the Symptom Guide​

SymptomImmediate actionRecovery criteria
Every endpoint is unreachableFreeze configuration changes, preserve running APISIX processes, and distinguish network isolation from quorum loss.Endpoints are reachable, quorum is stable, and an APISIX test change propagates.
Quorum is temporarily lostRestore connectivity or bring recoverable members back one at a time. Freeze configuration changes and preserve running APISIX processes.The existing cluster regains quorum, and writes and APISIX synchronization succeed.
Quorum is permanently lostIsolate the old members and select the newest verified snapshot. Do not restart the APISIX instances.A restored cluster passes isolated validation and APISIX processes reconnect in phases.
Elections or slow requests repeatCheck member latency, packet loss, CPU pressure, WAL sync, backend commit latency, and other workloads competing for disk I/O.Leadership is stable and measured latency returns to the accepted baseline under load.
Database growth or NOSPACE alarmStop the client generating unexpected writes, inspect revision growth and quota usage, then follow the compaction and one-member-at-a-time defrag procedure.The backend is below quota, the alarm is cleared, and writes and watches succeed.
APISIX configuration is staleCompare modification indexes and inspect reachability, TLS, authentication, permissions, and compacted-watch recovery.All APISIX processes apply the test configuration revision and proxy traffic uses the expected configuration.
TLS or RBAC errorsStop credential changes and inspect certificate expiration, CA trust, subject alternative names (SAN), Server Name Indication (SNI), time synchronization, and prefix permissions.Handshakes and positive and negative access tests pass for each APISIX role, and the credential rollback path is verified without broadening permissions.

After recovery, perform a test configuration change and a real proxy request. Confirm that all APISIX processes apply the expected configuration before reopening routine configuration changes.

Upgrade Safely​

Use an upgrade as a controlled production change with explicit entry criteria, per-member verification, and a documented rollback boundary.

Decide Whether to Upgrade​

Upgrade etcd when the target release addresses a relevant security issue, defect, or required capability, not only to follow the newest release. Use a release supported by the etcd project and validate it with the APISIX version, client configuration, operating system, and deployment platform you run. APISIX enforces a minimum etcd version at startup, but that minimum version requirement is not a production-version recommendation.

Review release notes, security advisories, configuration changes, and storage-version changes before selecting the target. Prefer a patch release within the supported current minor when it contains the required fix. Move to a supported minor when the fix is unavailable in the current minor or its support has ended.

Apply the Go/No-Go Gate​

Proceed only when the cluster has quorum, every member is healthy and caught up, and no alarms are active. Storage must have room for a member rebuild, monitoring must work, and a recent verified snapshot must have passed the required restore test. Follow supported minor-version steps without skipping a minor release. Postpone the upgrade if any prerequisite is missing or if the source-to-target upgrade path is unsupported.

Roll One Member at a Time​

Before an etcd upgrade, confirm that the cluster is healthy, has no alarms, has enough capacity for a member rebuild, and has a recently verified snapshot. Follow the etcd upgrade path for the source and target versions, upgrade one member at a time, and wait for it to become ready and catch up before continuing. After each member, verify cluster health, an APISIX configuration write, configuration propagation, and proxy traffic. After the cluster upgrade, also start a new APISIX canary to verify initial configuration loading and the scale-out path.

Do not apply a universal leader-last rule. The supported order and leadership-transfer behavior depend on the source and target releases. Follow the exact upgrade procedure for that version pair, including any required leadership transfer before stopping the current leader.

Understand the Rollback Boundary​

Define the rollback decision point before changing the first member. During a supported mixed-version phase, the release-specific procedure may allow replacing the new binary with the old version. After all members are upgraded and the cluster is considered fully upgraded, recovery may require the documented downgrade procedure or a snapshot restore instead.

caution

A snapshot does not make an arbitrary in-place rollback safe. Binary rollback, storage-version downgrade, and snapshot restore are different operations. Confirm the supported mixed-version, rollback, and downgrade procedures for the exact source and target releases before changing the first member.

Record the version reached, the verification results, and the remaining rollback option after each member. Stop the rollout when a verification fails instead of continuing to reduce the mixed-version window.