🚀 ExitVideo-Bot · Production Deployment Guide
版本:v1.0 · 作者:MiniMax · 2026-09-27
适用:把所有模块部署到生产环境(多租户 SaaS)
风格:直奔商业化,没有 demo
目录
一、部署架构总览
公网用户
│
▼
Cloudflare WAF + CDN
│
▼
Load Balancer (Nginx / Envoy)
│
┌─────────────────────┼─────────────────────┐
▼ ▼ ▼
API Pod 1 API Pod 2 API Pod 3
(FastAPI) (FastAPI) (FastAPI)
│ │ │
└─────────────────────┼─────────────────────┘
│
┌─────────────────────┼──────────────────────┐
▼ ▼ ▼
PostgreSQL Redis S3 / R2
(主从 + 备份) (缓存 + Celery) (媒体文件)
│
▼
PostgreSQL Read Replica
后台 Worker 集群
┌─────────────┬─────────────┬─────────────┐
▼ ▼ ▼ ▼
Worker 1 Worker 2 Worker 3 Worker 4
(bot 任务) (bot 任务) (bot 任务) (bot 任务)
设备池
┌─────────────┬─────────────┐
▼ ▼ ▼
Pixel 5 × 10 云手机 × 30 模拟器 × 5(测试)
监控 / 告警
┌─────────────┬─────────────┐
▼ ▼ ▼
Prometheus Loki Sentry
+ Grafana + Grafana (errors)
二、基础设施采购清单
2.1 控制面(API + DB)
项
数量
规格
单价
月成本
API Server
3 实例
4 vCPU / 8GB
$40
$120
PostgreSQL 主
1
4 vCPU / 16GB + 500GB SSD
$120
$120
PostgreSQL 从
1
同上 + 跨 AZ
$120
$120
Redis
1
2 vCPU / 4GB
$30
$30
对象存储(S3 / R2)
-
5TB
$15/TB
$75
小计
$465
2.2 Worker 集群
项
数量
规格
月成本
Bot Worker
3-10(自动伸缩)
8 vCPU / 16GB
$120-400
Celery 节点
2
同上
$80
2.3 设备层
项
数量
单价
月成本
Pixel 5/6 二手
10-30
¥600
一次性 ¥18K
云手机(多多云)
50
¥30/月
¥1,500
4G/5G 物联卡
30
¥30/月
¥900
USB Hub + 线材
5
¥100
一次性
2.4 总启动成本(Stage 2 期间)
一次性 : ¥18,000 + ¥500 = ¥18,500
月度运营 : $465 + $200 + ¥2,400 = ~$1,300/月
3.1 provider & backend
# infra/main.tf
terraform {
required_providers {
aws = { source = "hashicorp/aws", version = "~> 5.0" }
digitalocean = { source = "digitalocean/digitalocean" }
cloudflare = { source = "cloudflare/cloudflare" }
}
backend "s3" {
bucket = "exitvideo-tfstate"
key = "prod/terraform.tfstate"
region = "us-east-1"
encrypt = true
}
}
3.2 完整 stack 模板
# infra/eks.tf
resource "aws_eks_cluster" "prod" {
name = "exitvideo-prod"
role_arn = aws_iam_role.cluster.arn
vpc_config {
subnet_ids = aws_subnet.private [ * ]. id
}
}
resource "aws_eks_node_group" "workers" {
cluster_name = aws_eks_cluster.prod.name
node_group_name = "bot-workers"
instance_types = [ "c5.2xlarge" ] # 8 vCPU / 16GB
desired_size = 3
min_size = 3
max_size = 20 # 自动伸缩上限
labels = {
role = "bot-worker"
}
}
# infra/rds.tf
resource "aws_db_instance" "postgres_main" {
identifier = "exitvideo-prod-db"
engine = "postgres"
engine_version = "16.2"
instance_class = "db.r6g.large"
allocated_storage = 500
multi_az = true
db_subnet_group_name = aws_db_subnet_group.prod.name
backup_retention_period = 35 # SOC2 要求 ≥ 35 天
backup_window = "03:00-04:00"
deletion_protection = true
storage_encrypted = true
kms_key_id = aws_kms_key.rds.arn
performance_insights_enabled = true
}
resource "aws_db_instance" "postgres_replica" {
identifier = "exitvideo-prod-db-replica"
replicate_source_db = aws_db_instance.postgres_main.identifier
instance_class = "db.r6g.large"
multi_az = false
}
# infra/elasticache.tf
resource "aws_elasticache_cluster" "redis" {
cluster_id = "exitvideo-prod-redis"
engine = "redis"
engine_version = "7.1"
node_type = "cache.r6g.large"
num_cache_nodes = 1
parameter_group_name = "default.redis7"
port = 6379
transit_encryption_enabled = true
at_rest_encryption_enabled = true
}
# infra/cloudflare.tf
resource "cloudflare_record" "api" {
zone_id = var.cloudflare_zone_id
name = "api.exitvideo.com"
type = "CNAME"
value = aws_lb.api.dns_name
proxied = true
}
resource "cloudflare_record" "app" {
zone_id = var.cloudflare_zone_id
name = "app.exitvideo.com"
type = "CNAME"
value = aws_lb.api.dns_name
proxied = true
}
3.3 部署命令
cd infra/
terraform init
terraform plan -var-file= prod.tfvars
terraform apply -auto-approve
四、Docker 镜像构建
4.1 多阶段 Dockerfile(已存在)
exitvideo-bot/Dockerfile 已实现:
- Python 3.11-slim
- 系统依赖(adb / ffmpeg / 中文字体 / tini)
- pip cache 利用
- 健康检查
- 时区配置
4.2 镜像 Registry 推送
# AWS ECR
aws ecr get-login-password --region us-east-1 | \
docker login --username AWS --password-stdin 123 .dkr.ecr.us-east-1.amazonaws.com
docker build -t exitvideo-bot:1.0.0 ./exitvideo-bot
docker tag exitvideo-bot:1.0.0 \
123 .dkr.ecr.us-east-1.amazonaws.com/exitvideo-bot:1.0.0
docker push 123 .dkr.ecr.us-east-1.amazonaws.com/exitvideo-bot:1.0.0
4.3 镜像签名(推荐)
cosign sign --key cosign.key \
123 .dkr.ecr.us-east-1.amazonaws.com/exitvideo-bot:1.0.0
五、Kubernetes 部署清单
5.1 namespace
# k8s/00-namespace.yaml
apiVersion : v1
kind : Namespace
metadata :
name : exitvideo-prod
labels :
environment : production
managed-by : terraform
5.2 ConfigMap / Secret
# k8s/10-configmap.yaml
apiVersion : v1
kind : ConfigMap
metadata :
name : exitvideo-config
namespace : exitvideo-prod
data :
LOG_LEVEL : "INFO"
ENVIRONMENT : "production"
CELERY_BROKER_URL : "redis://redis:6379/0"
DATABASE_URL : "postgresql://user@db:5432/saas"
SENTRY_DSN : "https://xxx@sentry.io/123"
---
apiVersion : v1
kind : Secret
metadata :
name : exitvideo-secrets
namespace : exitvideo-prod
type : Opaque
stringData :
JWT_SECRET : "<random-64-chars>"
STRIPE_SECRET_KEY : "sk_live_xxx"
STRIPE_WEBHOOK_SECRET : "whsec_xxx"
ZHIPU_API_KEY : "z-xxx"
SMS_ACTIVATE_KEY : "xxx"
SMTP_PASSWORD : "xxx"
TELEGRAM_BOT_TOKEN : "xxx"
SLACK_WEBHOOK_URL : "https://hooks.slack.com/xxx"
PAGERDUTY_ROUTING_KEY : "xxx"
5.3 Deployment(API)
# k8s/20-api-deployment.yaml
apiVersion : apps/v1
kind : Deployment
metadata :
name : api
namespace : exitvideo-prod
spec :
replicas : 3
selector :
matchLabels :
app : api
template :
metadata :
labels :
app : api
spec :
containers :
- name : api
image : 123.dkr.ecr.us-east-1.amazonaws.com/exitvideo-bot:1.0.0
command : [ "uvicorn" , "api.server:app" , "--host" , "0.0.0.0" , "--port" , "8000" , "--workers" , "4" ]
ports :
- containerPort : 8000
envFrom :
- configMapRef : { name : exitvideo-config }
- secretRef : { name : exitvideo-secrets }
resources :
requests : { cpu : "500m" , memory : "512Mi" }
limits : { cpu : "2" , memory : "2Gi" }
livenessProbe :
httpGet : { path : /health , port : 8000 }
initialDelaySeconds : 10
periodSeconds : 30
readinessProbe :
httpGet : { path : /health , port : 8000 }
initialDelaySeconds : 5
periodSeconds : 10
---
apiVersion : v1
kind : Service
metadata :
name : api
namespace : exitvideo-prod
spec :
selector : { app : api }
ports :
- port : 80
targetPort : 8000
5.4 HPA(Horizontal Pod Autoscaler)
apiVersion : autoscaling/v2
kind : HorizontalPodAutoscaler
metadata :
name : api-hpa
namespace : exitvideo-prod
spec :
scaleTargetRef :
apiVersion : apps/v1
kind : Deployment
name : api
minReplicas : 3
maxReplicas : 20
metrics :
- type : Resource
resource :
name : cpu
target : { type : Utilization , averageUtilization : 70 }
- type : Resource
resource :
name : memory
target : { type : Utilization , averageUtilization : 80 }
5.5 Bot Worker 部署
# k8s/30-worker-deployment.yaml
apiVersion : apps/v1
kind : Deployment
metadata :
name : bot-worker
namespace : exitvideo-prod
spec :
replicas : 3
template :
spec :
containers :
- name : worker
image : 123.dkr.ecr.us-east-1.amazonaws.com/exitvideo-bot:1.0.0
command : [ "celery" , "-A" , "ops.celery_app" , "worker" , "--loglevel=INFO" , "--concurrency=4" ]
env :
- { name : CELERY_ROLE , value : "worker" }
# ... 资源限制同 API
5.6 Ingress (Cloudflare Tunnel)
apiVersion : networking.k8s.io/v1
kind : Ingress
metadata :
name : exitvideo
namespace : exitvideo-prod
annotations :
cert-manager.io/cluster-issuer : "letsencrypt-prod"
nginx.ingress.kubernetes.io/rate-limit : "100"
spec :
tls :
- hosts : [ api.exitvideo.com ]
secretName : api-tls
rules :
- host : api.exitvideo.com
http :
paths :
- path : /
pathType : Prefix
backend : { service : { name : api , port : { number : 80 } } }
六、CI/CD 流水线
6.1 GitHub Actions(已存在 + 拓展)
.github/workflows/deploy.yml:
name : Deploy to Production
on :
push :
branches : [ main ]
paths :
- 'exitvideo-bot/**'
- 'docs/Production_Deployment.md'
jobs :
test :
uses : ./.github/workflows/ci.yml
build-and-push :
needs : test
runs-on : ubuntu-latest
steps :
- uses : actions/checkout@v4
- name : Login to ECR
uses : aws-actions/amazon-ecr-login@v2
- name : Build
run : |
docker build -t $ECR/exitvideo-bot:$GITHUB_SHA ./exitvideo-bot
docker tag $ECR/exitvideo-bot:$GITHUB_SHA $ECR/exitvideo-bot:latest
- name : Push
run : |
docker push $ECR/exitvideo-bot:$GITHUB_SHA
docker push $ECR/exitvideo-bot:latest
deploy :
needs : build-and-push
runs-on : ubuntu-latest
steps :
- uses : aws-actions/configure-aws-credentials@v4
with :
role-to-assume : ${{ secrets.AWS_DEPLOY_ROLE }}
- name : Configure kubectl
uses : azure/setup-kubectl@v4
- run : |
aws eks update-kubeconfig --region us-east-1 --name exitvideo-prod
kubectl set image deployment/api api=$ECR/exitvideo-bot:$GITHUB_SHA -n exitvideo-prod
kubectl rollout status deployment/api -n exitvideo-prod
七、可观测性建设
7.1 Prometheus 指标
关键 SLI:
# infra/prometheus-rules.yaml
groups :
- name : exitvideo-sli
rules :
- record : exitvideo_api_latency_p99
expr : histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
- record : exitvideo_api_error_rate
expr : rate(http_requests_total{status=~"5.."}[5m])
/ rate(http_requests_total[5m])
- record : exitvideo_signup_success_rate
expr : rate(bot_signup_success_total[5m])
/ rate(bot_signup_attempt_total[5m])
- record : exitvideo_active_devices
expr : count(adb_device_online == 1)
- record : exitvideo_ban_rate
expr : rate(ban_detected_total[1h])
7.2 Grafana Dashboards
Dashboard
内容
API Overview
请求量、p50/p95/p99 延迟、错误率、Top 错误
Bot Operations
注册成功率、设备活跃数、任务队列长度、单平台表现
Billing
MRR、新订阅、流失、退款、chargeback
Incidents
Open/Closed incidents by severity, MTTA, MTTR
Compliance
GDPR/PIPL requests queue, consent revocation rate
7.3 日志(Loki)
# promtail config
scrape_configs :
- job_name : kubernetes
kubernetes_sd_configs : [{ role : pod }]
relabel_configs :
- source_labels : [ __meta_kubernetes_namespace ]
target_label : namespace
- source_labels : [ __meta_kubernetes_pod_label_app ]
target_label : app
7.4 告警规则
触发器
阈值
通知
API 5xx 错误率
> 1% 持续 5min
PagerDuty P2
注册失败率
> 50% 持续 5min
Slack #incidents
数据库连接池耗尽
active > 90%
PagerDuty P3
设备下线
> 5 台同时掉线
Slack #ops
月度流失激增
月环比 +50%
Email sales
资金异常
单笔退款 > $1000
Email finance
八、安全基线 Checklist
8.1 应用层
[ ] 所有流量强制 HTTPS(HSTS)
[ ] JWT 24h 过期 + Refresh 30 天
[ ] MFA 强制(Admin / Owner 角色)
[ ] RBAC 细化(owner/admin/member/viewer)
[ ] API 限流(per IP + per token + per tenant)
[ ] 输入校验(pydantic 严格模式)
[ ] 输出转义(防 XSS / 注入)
[ ] CORS 白名单
[ ] CSRF token(写操作)
[ ] Secrets 进 Vault / AWS Secrets Manager
[ ] 依赖扫描(safety + bandit 在 CI)
[ ] OWASP Top 10 自查
8.2 数据层
[ ] DB 加密 at rest + in transit
[ ] RLS(PostgreSQL Row-Level Security)
[ ] 列级加密(敏感字段:设备指纹、手机号)
[ ] 备份加密
[ ] 数据库账号最小权限
[ ] Query 时间审计
[ ] 定期 VACUUM + 索引优化
[ ] 慢查询监控 (>500ms 记录)
8.3 基础设施
[ ] SSH Key only(无密码登录)
[ ] Security Group 最小开放
[ ] VPC 私有子网
[ ] IAM 最小权限
[ ] GuardDuty / Security Hub 开启
[ ] CloudTrail 全审计
[ ] Container Image Scanning(ECR + Trivy)
[ ] Pod Security Policy
[ ] Network Policy
[ ] Audit log 集中保存 ≥ 7 年(合规)
8.4 第三方
[ ] 供应商清单 + 风险评估
[ ] DPA 签署
[ ] SOC2 report 获取
[ ] 4-hour breach notification clause
[ ] Subprocessor 公开列表
九、备份与灾备 (DR)
9.1 备份策略
数据
RPO
RTO
频率
保留
PostgreSQL
5min
30min
Continuous WAL + daily snapshot
35 天
Redis
不可备份
10min
RDB snapshot hourly
7 天
S3 / 对象存储
0
5min
Cross-region replication
永久
Terraform state
1min
5min
Versioned + locked
永久
Audit logs
0
10min
Streamed to S3
≥ 7 年
9.2 DR Runbook
# 如果主区不可用
dr_playbook :
trigger : AWS_REGION_HEALTH = "unhealthy"
steps :
- name : "Verify incident"
action : ops/incident_response.declare(severity=SEV1)
- name : "Promote replica"
action : aws rds promote-read-replica --db-instance-identifier exitvideo-prod-db-replica
- name : "Update DNS"
action : ./scripts/failover-to-secondary.sh
- name : "Communicate"
action : |
send_status_page_update("We are experiencing technical difficulties.")
notify_customer_via_email("Status update sent")
- name : "Verify recovery"
action : ./scripts/health-check-all.sh
- name : "Post-mortem"
action : ops/postmortem.generate(incident_id)
9.3 多区域部署(Stage 3 启用)
主区域:us-east-1(用户主要在欧美)
备份区域:eu-west-1(欧盟合规 + 低延迟)
冷备份:ap-southeast-1(亚洲备份)
流量管理:
- 主区域失败 → 5min 内切换到 EU
- 自动通过 Route 53 health check
9.4 备份验证脚本(每月执行)
#!/bin/bash
# scripts/verify-backup.sh
set -e
# 1. 拉取昨天 snapshot
SNAPSHOT_ID = $( aws rds describe-db-snapshots --query 'DBSnapshots[?StartTime>=`2026-09-26`].DBSnapshotIdentifier' --output text)
echo "→ Snapshot: $SNAPSHOT_ID "
# 2. 创建临时恢复实例
aws rds restore-db-instance-from-db-snapshot \
--db-instance-identifier exitvideo-verify-backup \
--db-snapshot-identifier $SNAPSHOT_ID \
--db-instance-class db.t3.small
# 3. 等 5 分钟
sleep 300
# 4. 跑数据完整性检查
python scripts/check-data-integrity.py --db exitvideo-verify-backup
# 5. 删除验证实例
aws rds delete-db-instance --db-instance-identifier exitvideo-verify-backup --skip-final-snapshot
echo "✓ Backup verified"
十、上线 Runbook
10.1 上线 Checklist
[ ] DNS 解析已配 (api.exitvideo.com → LB)
[ ] TLS 证书已部署(Let's Encrypt 或 Cloudflare Origin)
[ ] Secrets 已注入 Kubernetes Secret
[ ] 数据库迁移已执行(alembic upgrade head)
[ ] 静态资源(landing page)已部署到 CDN
[ ] Stripe webhook 已配(webhook URL + enabled events)
[ ] Sentry DSN 已配置
[ ] PagerDuty Integration Key 已注入
[ ] Slack channel #incidents 已创建
[ ] Status page(statuspage.io)已创建
[ ] 监控 + 告警已开启
[ ] 备份已配置
[ ] DR plan 已演练
[ ] 团队 oncall rotation 已配置
[ ] 应急联系清单已打印(Email + Phone)
10.2 上线步骤
# Day 0:T-24h
./scripts/pre-launch-check.sh
# Day 0:T-2h
./scripts/deploy-staging.sh
./scripts/smoke-test-staging.sh
./scripts/load-test-staging.sh
# Day 0:T-30min
./scripts/snapshot-prod-db.sh
# Day 0:T-0(部署)
git checkout main && git pull
./scripts/deploy-prod.sh
./scripts/verify-prod-health.sh
# Day 0:T+30min(监控)
watch -n 60 './scripts/health-check.sh'
# Day 1:T+24h
./scripts/post-launch-review.sh
10.3 回滚
# Kubernetes 自动回滚
kubectl rollout undo deployment/api -n exitvideo-prod
kubectl rollout undo deployment/bot-worker -n exitvideo-prod
# 数据库迁移回滚(如果有 destructive 变更)
./scripts/db-rollback.sh --to-revision= <previous-revision>
十一、容量规划
11.1 容量估算(Stage 2)
假设 50 个客户 × 25 平均账号 = 1,250 账号
维度
估值
每日注册请求
25(5% 客户每天注册)
每日发布动作
1,250(每个账号 1 条)
每日 API 调用
50,000
存储需求
50 GB / 月
带宽需求
200 GB / 月
11.2 扩容策略
指标
扩容触发
API p95 latency
> 500ms 持续 10min
CPU 利用率
> 70% 持续 15min
Memory 利用率
> 80% 持续 15min
DB 连接数
> 80% 池
Queue length
> 1,000 tasks
十二、SLA 契约
12.1 对外 SLA 表格
套餐
可用性 SLA
支持响应时间
赔偿
Starter
不承诺
24h
无
Growth
99% (月停 ≤ 7.2h)
4h
5% 月费
Scale
99.5% (月停 ≤ 3.6h)
1h
10% 月费
Enterprise
99.9% (月停 ≤ 43min)
15min
30% 月费
12.2 SLA 排除项
计划内维护(提前 48 小时通知不计入)
客户配置错误导致的不可用
第三方 API 故障(SMS-Activate / Stripe 等)
DDoS / 网络攻击(Cloudflare 防护范围外)
不可抗力
12.3 服务状态页
推荐 statuspage.io :
- 自动从监控同步状态
- 公开订阅事件通知
- 历史事件可追溯
- 与 incident response 系统联动
附录 A:应急联系清单
角色
联系人
方式
升级
On-call Lead
张三
Telegram / Phone
24h
Backend Lead
李四
Telegram
4h
Frontend Lead
王五
Telegram
8h
SRE
Eamon
PagerDuty
5min
CEO
小明
Phone
30min
Stripe Support
stripe.com/dashboard
Slack Connect
1h
SMS Provider
sms-activate.org/en/api2
Email
4h
Cloud Infra
AWS Support
Console
1h
DNS / CDN
Cloudflare
Dashboard
5min
DPO / 隐私官
privacy@exitvideo.com
Email
24h
附录 B:成本回测
阶段
月成本
MRR 预期
毛利率
Stage 1(自用)
$1300
$0(自营)
N/A
Stage 2(5 客户)
$2000
$5K
60%
Stage 3(50 客户)
$5000
$50K
90%
Stage 4(500 客户)
$25K
$500K
95%
维护者:MiniMax · 2026-09-27 · 配合 infra/ + k8s/ + scripts/ 落地