跳转至

🚀 ExitVideo-Bot · Production Deployment Guide

版本:v1.0 · 作者:MiniMax · 2026-09-27

适用:把所有模块部署到生产环境(多租户 SaaS)

风格:直奔商业化,没有 demo


目录


一、部署架构总览

                              公网用户
                                  │
                                  ▼
                          Cloudflare WAF + CDN
                                  │
                                  ▼
                          Load Balancer (Nginx / Envoy)
                                  │
            ┌─────────────────────┼─────────────────────┐
            ▼                     ▼                     ▼
        API Pod 1             API Pod 2             API Pod 3
        (FastAPI)              (FastAPI)              (FastAPI)
            │                     │                     │
            └─────────────────────┼─────────────────────┘
                                  │
            ┌─────────────────────┼──────────────────────┐
            ▼                     ▼                      ▼
        PostgreSQL             Redis                  S3 / R2
        (主从 + 备份)         (缓存 + Celery)         (媒体文件)
            │
            ▼
        PostgreSQL Read Replica

                  后台 Worker 集群
        ┌─────────────┬─────────────┬─────────────┐
        ▼             ▼             ▼             ▼
    Worker 1      Worker 2      Worker 3      Worker 4
    (bot 任务)    (bot 任务)    (bot 任务)    (bot 任务)

                  设备池
        ┌─────────────┬─────────────┐
        ▼             ▼             ▼
    Pixel 5 × 10   云手机 × 30   模拟器 × 5(测试)

                  监控 / 告警
        ┌─────────────┬─────────────┐
        ▼             ▼             ▼
    Prometheus    Loki         Sentry
    + Grafana     + Grafana    (errors)

二、基础设施采购清单

2.1 控制面(API + DB)

项 数量 规格 单价 月成本
API Server 3 实例 4 vCPU / 8GB $40 $120
PostgreSQL 主 1 4 vCPU / 16GB + 500GB SSD $120 $120
PostgreSQL 从 1 同上 + 跨 AZ $120 $120
Redis 1 2 vCPU / 4GB $30 $30
对象存储(S3 / R2) - 5TB $15/TB $75
小计 $465

2.2 Worker 集群

项 数量 规格 月成本
Bot Worker 3-10(自动伸缩) 8 vCPU / 16GB $120-400
Celery 节点 2 同上 $80

2.3 设备层

项 数量 单价 月成本
Pixel 5/6 二手 10-30 ¥600 一次性 ¥18K
云手机(多多云) 50 ¥30/月 ¥1,500
4G/5G 物联卡 30 ¥30/月 ¥900
USB Hub + 线材 5 ¥100 一次性

2.4 总启动成本(Stage 2 期间)

  • 一次性: ¥18,000 + ¥500 = ¥18,500
  • 月度运营: $465 + $200 + ¥2,400 = ~$1,300/月

三、IaC (Terraform)

3.1 provider & backend

# infra/main.tf
terraform {
  required_providers {
    aws = { source = "hashicorp/aws", version = "~> 5.0" }
    digitalocean = { source = "digitalocean/digitalocean" }
    cloudflare = { source = "cloudflare/cloudflare" }
  }

  backend "s3" {
    bucket = "exitvideo-tfstate"
    key    = "prod/terraform.tfstate"
    region = "us-east-1"
    encrypt = true
  }
}

3.2 完整 stack 模板

# infra/eks.tf
resource "aws_eks_cluster" "prod" {
  name     = "exitvideo-prod"
  role_arn = aws_iam_role.cluster.arn
  vpc_config {
    subnet_ids = aws_subnet.private[*].id
  }
}

resource "aws_eks_node_group" "workers" {
  cluster_name    = aws_eks_cluster.prod.name
  node_group_name = "bot-workers"
  instance_types  = ["c5.2xlarge"]    # 8 vCPU / 16GB
  desired_size    = 3
  min_size        = 3
  max_size        = 20              # 自动伸缩上限

  labels = {
    role = "bot-worker"
  }
}

# infra/rds.tf
resource "aws_db_instance" "postgres_main" {
  identifier        = "exitvideo-prod-db"
  engine            = "postgres"
  engine_version    = "16.2"
  instance_class    = "db.r6g.large"
  allocated_storage = 500

  multi_az            = true
  db_subnet_group_name = aws_db_subnet_group.prod.name

  backup_retention_period = 35        # SOC2 要求 ≥ 35 天
  backup_window           = "03:00-04:00"

  deletion_protection = true
  storage_encrypted   = true
  kms_key_id          = aws_kms_key.rds.arn

  performance_insights_enabled = true
}

resource "aws_db_instance" "postgres_replica" {
  identifier          = "exitvideo-prod-db-replica"
  replicate_source_db = aws_db_instance.postgres_main.identifier
  instance_class      = "db.r6g.large"
  multi_az            = false
}

# infra/elasticache.tf
resource "aws_elasticache_cluster" "redis" {
  cluster_id           = "exitvideo-prod-redis"
  engine               = "redis"
  engine_version       = "7.1"
  node_type            = "cache.r6g.large"
  num_cache_nodes      = 1
  parameter_group_name = "default.redis7"
  port                 = 6379

  transit_encryption_enabled = true
  at_rest_encryption_enabled = true
}

# infra/cloudflare.tf
resource "cloudflare_record" "api" {
  zone_id = var.cloudflare_zone_id
  name    = "api.exitvideo.com"
  type    = "CNAME"
  value   = aws_lb.api.dns_name
  proxied = true
}

resource "cloudflare_record" "app" {
  zone_id = var.cloudflare_zone_id
  name    = "app.exitvideo.com"
  type    = "CNAME"
  value   = aws_lb.api.dns_name
  proxied = true
}

3.3 部署命令

cd infra/
terraform init
terraform plan -var-file=prod.tfvars
terraform apply -auto-approve

四、Docker 镜像构建

4.1 多阶段 Dockerfile(已存在)

exitvideo-bot/Dockerfile 已实现: - Python 3.11-slim - 系统依赖(adb / ffmpeg / 中文字体 / tini) - pip cache 利用 - 健康检查 - 时区配置

4.2 镜像 Registry 推送

# AWS ECR
aws ecr get-login-password --region us-east-1 | \
  docker login --username AWS --password-stdin 123.dkr.ecr.us-east-1.amazonaws.com

docker build -t exitvideo-bot:1.0.0 ./exitvideo-bot
docker tag exitvideo-bot:1.0.0 \
  123.dkr.ecr.us-east-1.amazonaws.com/exitvideo-bot:1.0.0
docker push 123.dkr.ecr.us-east-1.amazonaws.com/exitvideo-bot:1.0.0

4.3 镜像签名(推荐)

cosign sign --key cosign.key \
  123.dkr.ecr.us-east-1.amazonaws.com/exitvideo-bot:1.0.0

五、Kubernetes 部署清单

5.1 namespace

# k8s/00-namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
  name: exitvideo-prod
  labels:
    environment: production
    managed-by: terraform

5.2 ConfigMap / Secret

# k8s/10-configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: exitvideo-config
  namespace: exitvideo-prod
data:
  LOG_LEVEL: "INFO"
  ENVIRONMENT: "production"
  CELERY_BROKER_URL: "redis://redis:6379/0"
  DATABASE_URL: "postgresql://user@db:5432/saas"
  SENTRY_DSN: "https://xxx@sentry.io/123"
---
apiVersion: v1
kind: Secret
metadata:
  name: exitvideo-secrets
  namespace: exitvideo-prod
type: Opaque
stringData:
  JWT_SECRET: "<random-64-chars>"
  STRIPE_SECRET_KEY: "sk_live_xxx"
  STRIPE_WEBHOOK_SECRET: "whsec_xxx"
  ZHIPU_API_KEY: "z-xxx"
  SMS_ACTIVATE_KEY: "xxx"
  SMTP_PASSWORD: "xxx"
  TELEGRAM_BOT_TOKEN: "xxx"
  SLACK_WEBHOOK_URL: "https://hooks.slack.com/xxx"
  PAGERDUTY_ROUTING_KEY: "xxx"

5.3 Deployment(API)

# k8s/20-api-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: api
  namespace: exitvideo-prod
spec:
  replicas: 3
  selector:
    matchLabels:
      app: api
  template:
    metadata:
      labels:
        app: api
    spec:
      containers:
        - name: api
          image: 123.dkr.ecr.us-east-1.amazonaws.com/exitvideo-bot:1.0.0
          command: ["uvicorn", "api.server:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "4"]
          ports:
            - containerPort: 8000
          envFrom:
            - configMapRef: { name: exitvideo-config }
            - secretRef: { name: exitvideo-secrets }
          resources:
            requests: { cpu: "500m", memory: "512Mi" }
            limits: { cpu: "2", memory: "2Gi" }
          livenessProbe:
            httpGet: { path: /health, port: 8000 }
            initialDelaySeconds: 10
            periodSeconds: 30
          readinessProbe:
            httpGet: { path: /health, port: 8000 }
            initialDelaySeconds: 5
            periodSeconds: 10
---
apiVersion: v1
kind: Service
metadata:
  name: api
  namespace: exitvideo-prod
spec:
  selector: { app: api }
  ports:
    - port: 80
      targetPort: 8000

5.4 HPA(Horizontal Pod Autoscaler)

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: api-hpa
  namespace: exitvideo-prod
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: api
  minReplicas: 3
  maxReplicas: 20
  metrics:
    - type: Resource
      resource:
        name: cpu
        target: { type: Utilization, averageUtilization: 70 }
    - type: Resource
      resource:
        name: memory
        target: { type: Utilization, averageUtilization: 80 }

5.5 Bot Worker 部署

# k8s/30-worker-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: bot-worker
  namespace: exitvideo-prod
spec:
  replicas: 3
  template:
    spec:
      containers:
        - name: worker
          image: 123.dkr.ecr.us-east-1.amazonaws.com/exitvideo-bot:1.0.0
          command: ["celery", "-A", "ops.celery_app", "worker", "--loglevel=INFO", "--concurrency=4"]
          env:
            - { name: CELERY_ROLE, value: "worker" }
          # ... 资源限制同 API

5.6 Ingress (Cloudflare Tunnel)

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: exitvideo
  namespace: exitvideo-prod
  annotations:
    cert-manager.io/cluster-issuer: "letsencrypt-prod"
    nginx.ingress.kubernetes.io/rate-limit: "100"
spec:
  tls:
    - hosts: [api.exitvideo.com]
      secretName: api-tls
  rules:
    - host: api.exitvideo.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend: { service: { name: api, port: { number: 80 } } }

六、CI/CD 流水线

6.1 GitHub Actions(已存在 + 拓展)

.github/workflows/deploy.yml:

name: Deploy to Production

on:
  push:
    branches: [main]
    paths:
      - 'exitvideo-bot/**'
      - 'docs/Production_Deployment.md'

jobs:
  test:
    uses: ./.github/workflows/ci.yml

  build-and-push:
    needs: test
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Login to ECR
        uses: aws-actions/amazon-ecr-login@v2
      - name: Build
        run: |
          docker build -t $ECR/exitvideo-bot:$GITHUB_SHA ./exitvideo-bot
          docker tag $ECR/exitvideo-bot:$GITHUB_SHA $ECR/exitvideo-bot:latest
      - name: Push
        run: |
          docker push $ECR/exitvideo-bot:$GITHUB_SHA
          docker push $ECR/exitvideo-bot:latest

  deploy:
    needs: build-and-push
    runs-on: ubuntu-latest
    steps:
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: ${{ secrets.AWS_DEPLOY_ROLE }}
      - name: Configure kubectl
        uses: azure/setup-kubectl@v4
      - run: |
          aws eks update-kubeconfig --region us-east-1 --name exitvideo-prod
          kubectl set image deployment/api api=$ECR/exitvideo-bot:$GITHUB_SHA -n exitvideo-prod
          kubectl rollout status deployment/api -n exitvideo-prod

七、可观测性建设

7.1 Prometheus 指标

关键 SLI:

# infra/prometheus-rules.yaml
groups:
  - name: exitvideo-sli
    rules:
      - record: exitvideo_api_latency_p99
        expr: histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))

      - record: exitvideo_api_error_rate
        expr: rate(http_requests_total{status=~"5.."}[5m])
            / rate(http_requests_total[5m])

      - record: exitvideo_signup_success_rate
        expr: rate(bot_signup_success_total[5m])
            / rate(bot_signup_attempt_total[5m])

      - record: exitvideo_active_devices
        expr: count(adb_device_online == 1)

      - record: exitvideo_ban_rate
        expr: rate(ban_detected_total[1h])

7.2 Grafana Dashboards

Dashboard 内容
API Overview 请求量、p50/p95/p99 延迟、错误率、Top 错误
Bot Operations 注册成功率、设备活跃数、任务队列长度、单平台表现
Billing MRR、新订阅、流失、退款、chargeback
Incidents Open/Closed incidents by severity, MTTA, MTTR
Compliance GDPR/PIPL requests queue, consent revocation rate

7.3 日志(Loki)

# promtail config
scrape_configs:
  - job_name: kubernetes
    kubernetes_sd_configs: [{ role: pod }]
    relabel_configs:
      - source_labels: [__meta_kubernetes_namespace]
        target_label: namespace
      - source_labels: [__meta_kubernetes_pod_label_app]
        target_label: app

7.4 告警规则

触发器 阈值 通知
API 5xx 错误率 > 1% 持续 5min PagerDuty P2
注册失败率 > 50% 持续 5min Slack #incidents
数据库连接池耗尽 active > 90% PagerDuty P3
设备下线 > 5 台同时掉线 Slack #ops
月度流失激增 月环比 +50% Email sales
资金异常 单笔退款 > $1000 Email finance

八、安全基线 Checklist

8.1 应用层

[ ] 所有流量强制 HTTPS(HSTS)
[ ] JWT 24h 过期 + Refresh 30 天
[ ] MFA 强制(Admin / Owner 角色)
[ ] RBAC 细化(owner/admin/member/viewer)
[ ] API 限流(per IP + per token + per tenant)
[ ] 输入校验(pydantic 严格模式)
[ ] 输出转义(防 XSS / 注入)
[ ] CORS 白名单
[ ] CSRF token(写操作)
[ ] Secrets 进 Vault / AWS Secrets Manager
[ ] 依赖扫描(safety + bandit 在 CI)
[ ] OWASP Top 10 自查

8.2 数据层

[ ] DB 加密 at rest + in transit
[ ] RLS(PostgreSQL Row-Level Security)
[ ] 列级加密(敏感字段:设备指纹、手机号)
[ ] 备份加密
[ ] 数据库账号最小权限
[ ] Query 时间审计
[ ] 定期 VACUUM + 索引优化
[ ] 慢查询监控 (>500ms 记录)

8.3 基础设施

[ ] SSH Key only(无密码登录)
[ ] Security Group 最小开放
[ ] VPC 私有子网
[ ] IAM 最小权限
[ ] GuardDuty / Security Hub 开启
[ ] CloudTrail 全审计
[ ] Container Image Scanning(ECR + Trivy)
[ ] Pod Security Policy
[ ] Network Policy
[ ] Audit log 集中保存 ≥ 7 年(合规)

8.4 第三方

[ ] 供应商清单 + 风险评估
[ ] DPA 签署
[ ] SOC2 report 获取
[ ] 4-hour breach notification clause
[ ] Subprocessor 公开列表

九、备份与灾备 (DR)

9.1 备份策略

数据 RPO RTO 频率 保留
PostgreSQL 5min 30min Continuous WAL + daily snapshot 35 天
Redis 不可备份 10min RDB snapshot hourly 7 天
S3 / 对象存储 0 5min Cross-region replication 永久
Terraform state 1min 5min Versioned + locked 永久
Audit logs 0 10min Streamed to S3 ≥ 7 年

9.2 DR Runbook

# 如果主区不可用
dr_playbook:
  trigger: AWS_REGION_HEALTH = "unhealthy"
  steps:
    - name: "Verify incident"
      action: ops/incident_response.declare(severity=SEV1)
    - name: "Promote replica"
      action: aws rds promote-read-replica --db-instance-identifier exitvideo-prod-db-replica
    - name: "Update DNS"
      action: ./scripts/failover-to-secondary.sh
    - name: "Communicate"
      action: |
        send_status_page_update("We are experiencing technical difficulties.")
        notify_customer_via_email("Status update sent")
    - name: "Verify recovery"
      action: ./scripts/health-check-all.sh
    - name: "Post-mortem"
      action: ops/postmortem.generate(incident_id)

9.3 多区域部署(Stage 3 启用)

主区域:us-east-1(用户主要在欧美)
备份区域:eu-west-1(欧盟合规 + 低延迟)
冷备份:ap-southeast-1(亚洲备份)

流量管理:
- 主区域失败 → 5min 内切换到 EU
- 自动通过 Route 53 health check

9.4 备份验证脚本(每月执行)

#!/bin/bash
# scripts/verify-backup.sh
set -e

# 1. 拉取昨天 snapshot
SNAPSHOT_ID=$(aws rds describe-db-snapshots --query 'DBSnapshots[?StartTime>=`2026-09-26`].DBSnapshotIdentifier' --output text)
echo "→ Snapshot: $SNAPSHOT_ID"

# 2. 创建临时恢复实例
aws rds restore-db-instance-from-db-snapshot \
  --db-instance-identifier exitvideo-verify-backup \
  --db-snapshot-identifier $SNAPSHOT_ID \
  --db-instance-class db.t3.small

# 3. 等 5 分钟
sleep 300

# 4. 跑数据完整性检查
python scripts/check-data-integrity.py --db exitvideo-verify-backup

# 5. 删除验证实例
aws rds delete-db-instance --db-instance-identifier exitvideo-verify-backup --skip-final-snapshot

echo "✓ Backup verified"

十、上线 Runbook

10.1 上线 Checklist

[ ] DNS 解析已配 (api.exitvideo.com → LB)
[ ] TLS 证书已部署(Let's Encrypt 或 Cloudflare Origin)
[ ] Secrets 已注入 Kubernetes Secret
[ ] 数据库迁移已执行(alembic upgrade head)
[ ] 静态资源(landing page)已部署到 CDN
[ ] Stripe webhook 已配(webhook URL + enabled events)
[ ] Sentry DSN 已配置
[ ] PagerDuty Integration Key 已注入
[ ] Slack channel #incidents 已创建
[ ] Status page(statuspage.io)已创建
[ ] 监控 + 告警已开启
[ ] 备份已配置
[ ] DR plan 已演练
[ ] 团队 oncall rotation 已配置
[ ] 应急联系清单已打印(Email + Phone)

10.2 上线步骤

# Day 0:T-24h
./scripts/pre-launch-check.sh

# Day 0:T-2h
./scripts/deploy-staging.sh
./scripts/smoke-test-staging.sh
./scripts/load-test-staging.sh

# Day 0:T-30min
./scripts/snapshot-prod-db.sh

# Day 0:T-0(部署)
git checkout main && git pull
./scripts/deploy-prod.sh
./scripts/verify-prod-health.sh

# Day 0:T+30min(监控)
watch -n 60 './scripts/health-check.sh'

# Day 1:T+24h
./scripts/post-launch-review.sh

10.3 回滚

# Kubernetes 自动回滚
kubectl rollout undo deployment/api -n exitvideo-prod
kubectl rollout undo deployment/bot-worker -n exitvideo-prod

# 数据库迁移回滚(如果有 destructive 变更)
./scripts/db-rollback.sh --to-revision=<previous-revision>

十一、容量规划

11.1 容量估算(Stage 2)

假设 50 个客户 × 25 平均账号 = 1,250 账号

维度 估值
每日注册请求 25(5% 客户每天注册)
每日发布动作 1,250(每个账号 1 条)
每日 API 调用 50,000
存储需求 50 GB / 月
带宽需求 200 GB / 月

11.2 扩容策略

指标 扩容触发
API p95 latency > 500ms 持续 10min
CPU 利用率 > 70% 持续 15min
Memory 利用率 > 80% 持续 15min
DB 连接数 > 80% 池
Queue length > 1,000 tasks

十二、SLA 契约

12.1 对外 SLA 表格

套餐 可用性 SLA 支持响应时间 赔偿
Starter 不承诺 24h 无
Growth 99% (月停 ≤ 7.2h) 4h 5% 月费
Scale 99.5% (月停 ≤ 3.6h) 1h 10% 月费
Enterprise 99.9% (月停 ≤ 43min) 15min 30% 月费

12.2 SLA 排除项

  • 计划内维护(提前 48 小时通知不计入)
  • 客户配置错误导致的不可用
  • 第三方 API 故障(SMS-Activate / Stripe 等)
  • DDoS / 网络攻击(Cloudflare 防护范围外)
  • 不可抗力

12.3 服务状态页

推荐 statuspage.io: - 自动从监控同步状态 - 公开订阅事件通知 - 历史事件可追溯 - 与 incident response 系统联动


附录 A:应急联系清单

角色 联系人 方式 升级
On-call Lead 张三 Telegram / Phone 24h
Backend Lead 李四 Telegram 4h
Frontend Lead 王五 Telegram 8h
SRE Eamon PagerDuty 5min
CEO 小明 Phone 30min
Stripe Support stripe.com/dashboard Slack Connect 1h
SMS Provider sms-activate.org/en/api2 Email 4h
Cloud Infra AWS Support Console 1h
DNS / CDN Cloudflare Dashboard 5min
DPO / 隐私官 privacy@exitvideo.com Email 24h

附录 B:成本回测

阶段 月成本 MRR 预期 毛利率
Stage 1(自用) $1300 $0(自营) N/A
Stage 2(5 客户) $2000 $5K 60%
Stage 3(50 客户) $5000 $50K 90%
Stage 4(500 客户) $25K $500K 95%

维护者:MiniMax · 2026-09-27 · 配合 infra/ + k8s/ + scripts/ 落地