ExitVideo-Bot · Disaster Recovery 手册¶
当最坏情况发生:Region 不可用 / 数据库丢失 / 全栈故障。 本手册确保 RTO < 1 小时、RPO < 15 分钟。
🎯 目标¶
| 维度 | 目标 | 说明 |
|---|---|---|
| RTO (Recovery Time Objective) | < 1 小时 | 从故障检测到恢复服务的最大时间 |
| RPO (Recovery Point Objective) | < 15 分钟 | 最大可容忍数据丢失窗口 |
| 可用性 SLO | 99.9% | 月度可用性 |
| 数据完整性 | 100% | 关键数据零丢失 |
🌐 1. Region 不可用¶
故障征兆¶
- AWS / GCP / Azure region-wide 服务中断
- 自家 region 内所有 Pod 健康检查失败
- DNS 解析 region endpoint 失败
- 网络分区(region 间 RTT > 1s)
检测(5 分钟内)¶
# 1. 检查 region 健康
aws health describe-events --region us-east-1 --filter eventTypeCategories=issue
# 或 GCP / Azure 等价命令
# 2. 检查 EKS 集群状态
aws eks describe-cluster --name exitvideo-bot-prod --region us-east-1
# 3. 检查 RPO 内备份是否最新
make dr-check-dr-backup-log
# 期望:最近一次 backup 时间戳 < 15 分钟前
决策(10 分钟内)¶
| 故障范围 | 决策 |
|---|---|
| 单 AZ 不可用 | 不切:等待 AZ 恢复(EKS 已多 AZ) |
| 单 Region 不可用 | 立即切:切到备用 region |
| 全 Region 不可用 | 紧急:启用冷备份(CloudFront 静态页 + 只读 API) |
Failover 操作(30 分钟内)¶
# 1. 更新 DNS 到备用 region
# Route53 / CloudDNS failover
aws route53 create-health-check \
--caller-reference $(date +%s) \
--health-check-config file://healthcheck.json
aws route53 update-traffic-policy-instance \
--id <policy-id> \
--traffic-policy-id <policy-id> \
--policy-version 1 \
--traffic-policy "{\"AWSPolicy-type":...}"
# 2. 启动备用 region 集群
cd infra/terraform/environments/prod-dr
terraform init
terraform apply -var "region=us-west-2" -auto-approve
# 3. 恢复数据(RDS cross-region replica 已自动提升)
kubectl exec -n database --region us-west-2 postgres-0 -- pg_isready
# 4. ArgoCD 同步(指向备用 region)
kubectl config use-context exitvideo-bot-dr
argocd app sync exitvideo-bot-prod --retry-limit=10
# 5. 验证 smoke
make smoke ENV=dr
Failback(备用 region 恢复后)¶
# 1. 切回主 region
# 2. 重新建立 cross-region replica
# 3. 演练并出具 Postmortem
💾 2. 数据库主备切换¶
故障征兆¶
- 主库写入失败
- 主库 CPU 100% 持续 5 分钟
- 主库磁盘写满
- 主库 replication lag > 1h
RDS 自动 failover¶
RDS Multi-AZ 已配置,主备切换 通常 60-120 秒自动完成。
手动 failover(紧急)¶
# 1. AWS CLI 触发 failover
aws rds failover-db-cluster \
--db-cluster-identifier exitvideo-bot-prod \
--region us-east-1
# 2. 等待 ~60 秒,验证新主可用
aws rds describe-db-clusters \
--db-cluster-identifier exitvideo-bot-prod \
--region us-east-1 \
--query "DBClusters[0].DBClusterMembers[*].[DBInstanceIdentifier,IsClusterWriter]"
# 3. 应用自动重连(RDS endpoint 不变)
kubectl rollout restart deploy/exitvideo-bot-api -n exitvideo-bot
应用层处理¶
# 1. 检查应用连接
kubectl exec -n exitvideo-bot deploy/exitvideo-bot-api -- \
curl -sf http://localhost:8000/health/db
# 2. 如果连接失败,重启 Pod
kubectl rollout restart deploy/exitvideo-bot-api -n exitvideo-bot
kubectl rollout restart deploy/exitvideo-bot-worker -n exitvideo-bot
# 3. 监控错误率
make grafana-ui
# 期望:30 分钟内错误率回到 < 0.1%
🗑️ 3. 数据丢失或误删¶
故障征兆¶
- 用户报告"我的账号不见了"
- 数据库 row count 突然下降
- Audit log 显示
DELETE FROM accounts WHERE ... - Terraform 误操作删除资源
立即行动(5 分钟内)¶
# 1. 停止一切写入
# 通过 Argo Rollouts pause 阻止流量
kubectl argo rollouts pause exitvideo-bot-api -n exitvideo-bot
kubectl argo rollouts pause exitvideo-bot-worker -n exitvideo-bot
# 2. 锁定数据库(仅允许读取)
kubectl exec -n database postgres-0 -- \
psql -U postgres -c "ALTER SYSTEM SET default_transaction_read_only = on;"
kubectl exec -n database postgres-0 -- pg_ctl reload
# 3. 通知 + 创建事故频道
make incident-create SEVERITY=P0 TITLE="data_loss_<id>"
数据恢复(30-60 分钟)¶
# 1. 列出可用备份
make dr-list-backups
# 2. 选择最近的 RPO 兼容备份
# 例如:误删发生在 15:00,选择 14:50 的 PITR backup
# 3. 创建新的 RDS 实例从备份恢复
aws rds restore-db-instance-to-point-in-time \
--source-db-instance-identifier exitvideo-bot-prod \
--target-db-instance-identifier exitvideo-bot-recovery \
--restore-time "2026-09-28T14:50:00Z" \
--no-multi-az \
--region us-east-1
# 4. 等待恢复完成(可能 30-60 分钟)
aws rds wait db-instance-available \
--db-instance-identifier exitvideo-bot-recovery
# 5. 切换应用到 recovery 实例
kubectl patch secret exitvideo-bot-secrets -n exitvideo-bot --type merge \
-p '{"data":{"database-url":"<base64-encoded-new-url>"}}'
kubectl rollout restart deploy/exitvideo-bot-api -n exitvideo-bot
# 6. 验证数据完整性
make dr-verify-data-integrity BACKUP_TIME=2026-09-28T14:50:00Z
恢复后¶
# 1. 删除旧主库(避免双写)
aws rds delete-db-instance \
--db-instance-identifier exitvideo-bot-prod \
--skip-final-snapshot \
--region us-east-1
# 2. 把 recovery 实例重命名为主库
aws rds modify-db-instance \
--db-instance-identifier exitvideo-bot-recovery \
--new-db-instance-identifier exitvideo-bot-prod \
--apply-immediately
# 3. 恢复流量
kubectl argo rollouts promote exitvideo-bot-api -n exitvideo-bot
kubectl argo rollouts promote exitvideo-bot-worker -n exitvideo-bot
🔑 4. 密钥泄露(GitHub / 日志)¶
故障征兆¶
- GitGuardian / TruffleHog 告警
- GitHub Secret Scanning 告警
- 日志里发现明文 API key
立即行动(5 分钟内)¶
# 1. 轮换所有相关密钥
# Vault
vault token revoke -self
vault secrets disable secret/exitvideo/prod
# 2. 轮换上游服务(Stripe / 智谱 / SMS)
# Stripe Dashboard → API keys → Roll
# 智谱 API 控制台 → 重新生成
# 3. 通过 ESO 同步到 K8s
kubectl annotate externalsecret exitvideo-bot-secrets -n exitvideo-bot \
force-sync=$(date +%s) --overwrite
# 4. 重启应用使用新密钥
kubectl rollout restart deploy/exitvideo-bot-api -n exitvideo-bot
kubectl rollout restart deploy/exitvideo-bot-worker -n exitvideo-bot
# 5. 清理 git 历史(如果 commit 已 push 到 main)
git filter-repo --path credentials.json --invert-paths
git push --force --all
# ⚠️ 所有 fork 也会暴露历史,需评估风险
# 6. 通知所有相关方
# - 上游服务商(Stripe fraud team)
# - 内部 TL / CTO
# - 法务(如涉及用户数据)
# - 投资人(如重大事件)
🌊 5. DDoS 大规模攻击¶
故障征兆¶
- 入口 QPS 飙升 > 10 倍基线
- 错误率上升(rate limit)
- CloudFlare 安全事件告警
立即行动(5 分钟内)¶
# 1. CloudFlare 启用 "I'm Under Attack" 模式
# 通过 CloudFlare Dashboard 或 API
# 2. 启用高级 WAF 规则
curl -X PATCH "https://api.cloudflare.com/client/v4/zones/$ZONE_ID/firewall/rules" \
-H "Authorization: Bearer $CF_API_TOKEN" \
-d '[{"id":"<rule-id>","action":"block","priority":1,...}]'
# 3. Argo Rollouts 暂停部分流量
kubectl argo rollouts pause exitvideo-bot-api -n exitvideo-bot
# 4. 联系 CloudFlare Enterprise 支持(如适用)
长期缓解¶
# 1. 启用 CloudFlare Magic Transit
# 2. 配置 rate limiting 规则
# 3. 启用 Argo Rollouts Analysis 模板自动限流
📦 6. 备份策略¶
数据库备份¶
| 备份类型 | 频率 | 保留 | RPO |
|---|---|---|---|
| 自动 Snapshot | 每天 03:00 UTC | 7 天 | 24 小时 |
| Continuous Backup (PITR) | 实时 | 35 天 | 5 分钟 |
| 手动 Snapshot(dr release 前) | 手动 | 永久 | 0 |
备份验证¶
# 每周演练一次
make dr-verify-backup
# 1. 列出最近备份
aws rds describe-db-snapshots \
--db-instance-identifier exitvideo-bot-prod \
--region us-east-1
# 2. 恢复到一个测试实例
aws rds restore-db-instance-from-db-snapshot \
--db-instance-identifier exitvideo-bot-backup-test \
--db-snapshot-identifier <snapshot-id> \
--no-multi-az
# 3. 验证数据完整性
make dr-verify-data-integrity BACKUP_TIME=latest
# 4. 删除测试实例
aws rds delete-db-instance \
--db-instance-identifier exitvideo-bot-backup-test \
--skip-final-snapshot
etcd 备份¶
# 每天自动备份
ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-$(date +%F).db \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key
# 上传到 S3
aws s3 cp /backup/etcd-$(date +%F).db s3://exitvideo-bot-backups/etcd/
GitOps 仓库备份¶
# GitHub 仓库启用 cross-archive 自动 fork
# Settings → Archives → Include fork
# ArgoCD ApplicationSet 备份
kubectl get applicationset -n argocd -o yaml > /backup/appset-$(date +%F).yaml
aws s3 cp /backup/appset-$(date +%F).yaml s3://exitvideo-bot-backups/argocd/
📋 DR Drill 演练¶
每季度演练一次,覆盖:
- [ ] 单 Region failover(30 分钟内)
- [ ] 数据库 PITR 恢复(30 分钟内)
- [ ] etcd 恢复(10 分钟内)
- [ ] 密钥轮换(5 分钟内)
- [ ] DDoS 缓解(10 分钟内)
记录演练结果到 infra/runbooks/DR-DRILL-LOG.md。
📊 演练结果指标¶
| 演练 | 上次 RTO | 目标 RTO | 上次 RPO | 目标 RPO |
|---|---|---|---|---|
| Region failover | - | < 1h | - | < 15min |
| DB PITR | - | < 1h | - | < 15min |
| Secret rotation | - | < 5min | - | 0 |
🔗 相关资源¶
- On-call Runbook
- Release Runbook
- Terraform multi_region module
- Terraform multi_region DR config
- Runbook 总索引
最后更新:第 13 轮 Runbook 补写 覆盖场景:Region failover / DB 主备切换 / 数据恢复 / 密钥泄露 / DDoS / 备份策略 / 季度演练