跳转至

ExitVideo-Bot · Disaster Recovery 手册

当最坏情况发生:Region 不可用 / 数据库丢失 / 全栈故障。 本手册确保 RTO < 1 小时、RPO < 15 分钟。


🎯 目标

维度 目标 说明
RTO (Recovery Time Objective) < 1 小时 从故障检测到恢复服务的最大时间
RPO (Recovery Point Objective) < 15 分钟 最大可容忍数据丢失窗口
可用性 SLO 99.9% 月度可用性
数据完整性 100% 关键数据零丢失

🌐 1. Region 不可用

故障征兆

  • AWS / GCP / Azure region-wide 服务中断
  • 自家 region 内所有 Pod 健康检查失败
  • DNS 解析 region endpoint 失败
  • 网络分区(region 间 RTT > 1s)

检测(5 分钟内)

# 1. 检查 region 健康
aws health describe-events --region us-east-1 --filter eventTypeCategories=issue
# 或 GCP / Azure 等价命令

# 2. 检查 EKS 集群状态
aws eks describe-cluster --name exitvideo-bot-prod --region us-east-1

# 3. 检查 RPO 内备份是否最新
make dr-check-dr-backup-log
# 期望:最近一次 backup 时间戳 < 15 分钟前

决策(10 分钟内)

故障范围 决策
单 AZ 不可用 不切:等待 AZ 恢复(EKS 已多 AZ)
单 Region 不可用 立即切:切到备用 region
全 Region 不可用 紧急:启用冷备份(CloudFront 静态页 + 只读 API)

Failover 操作(30 分钟内)

# 1. 更新 DNS 到备用 region
# Route53 / CloudDNS failover
aws route53 create-health-check \
  --caller-reference $(date +%s) \
  --health-check-config file://healthcheck.json

aws route53 update-traffic-policy-instance \
  --id <policy-id> \
  --traffic-policy-id <policy-id> \
  --policy-version 1 \
  --traffic-policy "{\"AWSPolicy-type":...}"

# 2. 启动备用 region 集群
cd infra/terraform/environments/prod-dr
terraform init
terraform apply -var "region=us-west-2" -auto-approve

# 3. 恢复数据(RDS cross-region replica 已自动提升)
kubectl exec -n database --region us-west-2 postgres-0 -- pg_isready

# 4. ArgoCD 同步(指向备用 region)
kubectl config use-context exitvideo-bot-dr
argocd app sync exitvideo-bot-prod --retry-limit=10

# 5. 验证 smoke
make smoke ENV=dr

Failback(备用 region 恢复后)

# 1. 切回主 region
# 2. 重新建立 cross-region replica
# 3. 演练并出具 Postmortem

💾 2. 数据库主备切换

故障征兆

  • 主库写入失败
  • 主库 CPU 100% 持续 5 分钟
  • 主库磁盘写满
  • 主库 replication lag > 1h

RDS 自动 failover

RDS Multi-AZ 已配置,主备切换 通常 60-120 秒自动完成。

手动 failover(紧急)

# 1. AWS CLI 触发 failover
aws rds failover-db-cluster \
  --db-cluster-identifier exitvideo-bot-prod \
  --region us-east-1

# 2. 等待 ~60 秒,验证新主可用
aws rds describe-db-clusters \
  --db-cluster-identifier exitvideo-bot-prod \
  --region us-east-1 \
  --query "DBClusters[0].DBClusterMembers[*].[DBInstanceIdentifier,IsClusterWriter]"

# 3. 应用自动重连(RDS endpoint 不变)
kubectl rollout restart deploy/exitvideo-bot-api -n exitvideo-bot

应用层处理

# 1. 检查应用连接
kubectl exec -n exitvideo-bot deploy/exitvideo-bot-api -- \
  curl -sf http://localhost:8000/health/db

# 2. 如果连接失败,重启 Pod
kubectl rollout restart deploy/exitvideo-bot-api -n exitvideo-bot
kubectl rollout restart deploy/exitvideo-bot-worker -n exitvideo-bot

# 3. 监控错误率
make grafana-ui
# 期望:30 分钟内错误率回到 < 0.1%

🗑️ 3. 数据丢失或误删

故障征兆

  • 用户报告"我的账号不见了"
  • 数据库 row count 突然下降
  • Audit log 显示 DELETE FROM accounts WHERE ...
  • Terraform 误操作删除资源

立即行动(5 分钟内)

# 1. 停止一切写入
# 通过 Argo Rollouts pause 阻止流量
kubectl argo rollouts pause exitvideo-bot-api -n exitvideo-bot
kubectl argo rollouts pause exitvideo-bot-worker -n exitvideo-bot

# 2. 锁定数据库(仅允许读取)
kubectl exec -n database postgres-0 -- \
  psql -U postgres -c "ALTER SYSTEM SET default_transaction_read_only = on;"
kubectl exec -n database postgres-0 -- pg_ctl reload

# 3. 通知 + 创建事故频道
make incident-create SEVERITY=P0 TITLE="data_loss_<id>"

数据恢复(30-60 分钟)

# 1. 列出可用备份
make dr-list-backups

# 2. 选择最近的 RPO 兼容备份
# 例如:误删发生在 15:00,选择 14:50 的 PITR backup

# 3. 创建新的 RDS 实例从备份恢复
aws rds restore-db-instance-to-point-in-time \
  --source-db-instance-identifier exitvideo-bot-prod \
  --target-db-instance-identifier exitvideo-bot-recovery \
  --restore-time "2026-09-28T14:50:00Z" \
  --no-multi-az \
  --region us-east-1

# 4. 等待恢复完成(可能 30-60 分钟)
aws rds wait db-instance-available \
  --db-instance-identifier exitvideo-bot-recovery

# 5. 切换应用到 recovery 实例
kubectl patch secret exitvideo-bot-secrets -n exitvideo-bot --type merge \
  -p '{"data":{"database-url":"<base64-encoded-new-url>"}}'
kubectl rollout restart deploy/exitvideo-bot-api -n exitvideo-bot

# 6. 验证数据完整性
make dr-verify-data-integrity BACKUP_TIME=2026-09-28T14:50:00Z

恢复后

# 1. 删除旧主库(避免双写)
aws rds delete-db-instance \
  --db-instance-identifier exitvideo-bot-prod \
  --skip-final-snapshot \
  --region us-east-1

# 2. 把 recovery 实例重命名为主库
aws rds modify-db-instance \
  --db-instance-identifier exitvideo-bot-recovery \
  --new-db-instance-identifier exitvideo-bot-prod \
  --apply-immediately

# 3. 恢复流量
kubectl argo rollouts promote exitvideo-bot-api -n exitvideo-bot
kubectl argo rollouts promote exitvideo-bot-worker -n exitvideo-bot

🔑 4. 密钥泄露(GitHub / 日志)

故障征兆

  • GitGuardian / TruffleHog 告警
  • GitHub Secret Scanning 告警
  • 日志里发现明文 API key

立即行动(5 分钟内)

# 1. 轮换所有相关密钥
# Vault
vault token revoke -self
vault secrets disable secret/exitvideo/prod

# 2. 轮换上游服务(Stripe / 智谱 / SMS)
# Stripe Dashboard → API keys → Roll
# 智谱 API 控制台 → 重新生成

# 3. 通过 ESO 同步到 K8s
kubectl annotate externalsecret exitvideo-bot-secrets -n exitvideo-bot \
  force-sync=$(date +%s) --overwrite

# 4. 重启应用使用新密钥
kubectl rollout restart deploy/exitvideo-bot-api -n exitvideo-bot
kubectl rollout restart deploy/exitvideo-bot-worker -n exitvideo-bot

# 5. 清理 git 历史(如果 commit 已 push 到 main)
git filter-repo --path credentials.json --invert-paths
git push --force --all
# ⚠️ 所有 fork 也会暴露历史,需评估风险

# 6. 通知所有相关方
# - 上游服务商(Stripe fraud team)
# - 内部 TL / CTO
# - 法务(如涉及用户数据)
# - 投资人(如重大事件)

🌊 5. DDoS 大规模攻击

故障征兆

  • 入口 QPS 飙升 > 10 倍基线
  • 错误率上升(rate limit)
  • CloudFlare 安全事件告警

立即行动(5 分钟内)

# 1. CloudFlare 启用 "I'm Under Attack" 模式
# 通过 CloudFlare Dashboard 或 API

# 2. 启用高级 WAF 规则
curl -X PATCH "https://api.cloudflare.com/client/v4/zones/$ZONE_ID/firewall/rules" \
  -H "Authorization: Bearer $CF_API_TOKEN" \
  -d '[{"id":"<rule-id>","action":"block","priority":1,...}]'

# 3. Argo Rollouts 暂停部分流量
kubectl argo rollouts pause exitvideo-bot-api -n exitvideo-bot

# 4. 联系 CloudFlare Enterprise 支持(如适用)

长期缓解

# 1. 启用 CloudFlare Magic Transit
# 2. 配置 rate limiting 规则
# 3. 启用 Argo Rollouts Analysis 模板自动限流

📦 6. 备份策略

数据库备份

备份类型 频率 保留 RPO
自动 Snapshot 每天 03:00 UTC 7 天 24 小时
Continuous Backup (PITR) 实时 35 天 5 分钟
手动 Snapshot(dr release 前) 手动 永久 0

备份验证

# 每周演练一次
make dr-verify-backup

# 1. 列出最近备份
aws rds describe-db-snapshots \
  --db-instance-identifier exitvideo-bot-prod \
  --region us-east-1

# 2. 恢复到一个测试实例
aws rds restore-db-instance-from-db-snapshot \
  --db-instance-identifier exitvideo-bot-backup-test \
  --db-snapshot-identifier <snapshot-id> \
  --no-multi-az

# 3. 验证数据完整性
make dr-verify-data-integrity BACKUP_TIME=latest

# 4. 删除测试实例
aws rds delete-db-instance \
  --db-instance-identifier exitvideo-bot-backup-test \
  --skip-final-snapshot

etcd 备份

# 每天自动备份
ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-$(date +%F).db \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key

# 上传到 S3
aws s3 cp /backup/etcd-$(date +%F).db s3://exitvideo-bot-backups/etcd/

GitOps 仓库备份

# GitHub 仓库启用 cross-archive 自动 fork
# Settings → Archives → Include fork

# ArgoCD ApplicationSet 备份
kubectl get applicationset -n argocd -o yaml > /backup/appset-$(date +%F).yaml
aws s3 cp /backup/appset-$(date +%F).yaml s3://exitvideo-bot-backups/argocd/

📋 DR Drill 演练

每季度演练一次,覆盖:

  • [ ] 单 Region failover(30 分钟内)
  • [ ] 数据库 PITR 恢复(30 分钟内)
  • [ ] etcd 恢复(10 分钟内)
  • [ ] 密钥轮换(5 分钟内)
  • [ ] DDoS 缓解(10 分钟内)

记录演练结果到 infra/runbooks/DR-DRILL-LOG.md。


📊 演练结果指标

演练 上次 RTO 目标 RTO 上次 RPO 目标 RPO
Region failover - < 1h - < 15min
DB PITR - < 1h - < 15min
Secret rotation - < 5min - 0

🔗 相关资源


最后更新:第 13 轮 Runbook 补写 覆盖场景:Region failover / DB 主备切换 / 数据恢复 / 密钥泄露 / DDoS / 备份策略 / 季度演练