跳转至

🚨 Multi-Region DR Runbook

主区故障 → 自动 / 半自动切换到备份区

自动化等级:Route53 自动 failover + RDS 自动 promote 手动触发:跨区应用部署 + 数据重放


自动层(无需人工)

1. DNS Failover

  • 主区 Route53 Health Check 每 30s 探测 /health
  • 失败 3 次后 → 切换到 backup DNS 记录
  • TTL:60s(保证 failover 不会太久)

2. RDS Cross-Region Backup

  • 自动每日跨区复制 automatic backup
  • backup region 保留 14 天

3. S3 CRR

  • 对象存储跨区实时复制
  • 删除标记也复制(避免删除数据不一致)

4. CloudWatch Alarm

  • 自动触发 Incident Response(PagerDuty)
  • 主区全部组件不可用时启动切换流程

半自动层(Runbook 1-2 小时)

T+0:告警触发

# PagerDuty 自动喊 oncall
aws cloudwatch describe-alarms \
  --alarm-names exitvideo-failover-trigger

T+5:确认主区无法恢复

# 检查主区状态
aws ec2 describe-instances \
  --region us-east-1 \
  --filter Name=instance-state-name,Values=running

aws rds describe-db-instances \
  --db-instance-identifier exitvideo-prod-db \
  --region us-east-1

aws elbv2 describe-load-balancers \
  --region us-east-1

T+30:拉起备份区应用

# 切换 DNS(也可以 Route53 自动做)
# 但应用层也得启用
cd infra/terraform
terraform apply -var-file=prod2.tfvars

# 启动 K8s(备份区已有 EKS)
cd infra/k8s
aws eks update-kubeconfig --region eu-west-1 --name exitvideo-prod2
kubectl apply -k .

T+60:Promote 数据库

# 在主区检查 RDS Replica(如果有)
aws rds describe-db-instances \
  --region eu-west-1 \
  --db-instance-identifier exitvideo-prod-db-replica-eu

# 如果是 DR,将 backup region 的 snapshot 提升为独立实例
aws rds restore-db-instance-from-db-snapshot \
  --region eu-west-1 \
  --db-instance-identifier exitvideo-prod2-db \
  --db-snapshot-identifier <latest-snapshot-arn> \
  --db-instance-class db.r6g.large

T+90:恢复后验证

# Health check
curl https://api.exitvideo.com/health
# 应返回 {"status": "ok"}

# Stripe webhook 也要切到备份区
# (Stripe 只能有一个 endpoint,需要手动切换)

# 监控 dashboard 切换
# Grafana 已经在备份区独立部署

T+120:事后复盘

  1. 召集 on-call + leadership
  2. 拉 incident response runbook 记录
  3. 创建 postmortem(已脚本化:ops.postmortem.save)
  4. 评估 SLO 违约,给受影响客户发 credit
  5. 24h 后做完整复盘会议

手动层(决策点)

关键决策

  1. Failover 触发条件
  2. 主区不可用 ≥ 30min?
  3. 数据完整性受损?
  4. 主动切换(planned maintenance)?

  5. Back-fill 策略

  6. 主区恢复后是否迁回?
  7. 数据冲突时以哪个为准?
  8. 用户感知影响多大?

Failback Runbook

主区恢复后:

# 1. 同步 backup 数据回主区
aws rds create-db-snapshot \
  --db-instance-identifier exitvideo-prod2-db \
  --db-snapshot-identifier dr-backfill

aws rds copy-db-snapshot \
  --source-db-snapshot-identifier arn:aws:rds:eu-west-1:123:db-snapshot:dr-backfill \
  --target-db-snapshot-identifier main-backfill \
  --source-region eu-west-1 \
  --target-region us-east-1

# 2. restore 到主区
aws rds restore-db-instance-from-db-snapshot \
  --region us-east-1 \
  --db-instance-identifier exitvideo-prod-db \
  --db-snapshot-identifier main-backfill

# 3. 切换 DNS 指回主区
# 由 CloudWatch alarm 清空时自动恢复


监控

关键 SLO(业务级)

SLO 目标 阈值
可用性 99.9% 月度 uptime
故障恢复时间 RTO 2h 主区→备份区
数据丢失 RPO 5min 备份频率

关键告警

Alert 含义 通知
ApiLatencyP99High > 5s 性能下降 Slack
SignupSuccessRateLow < 50% 注册崩溃 Slack + Email
BeatSchedulerDown 周期任务停了 PagerDuty
DatabaseConnectionsHigh > 100 连接池耗尽 Slack
RevenueSpikeDrop Stripe 收入断崖 Email

演练

每年必须进行 2 次 DR 演练: - Q2:完整 failover 测试 + 数据校验 - Q4:failback 测试

演练必须以 "假装是真实故障" 形式进行,事后提交真实复盘。


参考资料