Seven destinations scored against 129 measured AWS couplings, read live from two Kubernetes clusters, six AWS accounts and the Cloudflare account.
Do not move the platform to Akamai. The candidate that best fits this estate is the one that moves the 97 % and leaves the 3 %: keep compute, data plane, identity and CI on AWS EKS, and move only the video CDN and its origin — to Cloudflare Enterprise (variant B), with Akamai AMD (variant A) carried into the same RFP as the alternative.
Scores are 0–5, 5 best. Note the sign convention: operational risk is scored as "risk handled well", so 5 = lowest risk.
| # | Destination | k8s | Data | CDN | Ident | Agents | Risk | Score | Blockers | Why |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Stay on AWS, move only the video CDN + origin 25 of 129 rows touched | 5 | 5 | 3 | 5 | 5 | 4 | 4.45 | 3 | 104 of 129 coupling rows UNCHANGED (4 DIRECT/5 EQUIV/5 REDESIGN/3 NO-EQ/8 RETIRE). Variant B (Cloudflare Enterprise, R2 origin, Worker token) is the recommended start; variant A (Akamai AMD)… |
| 2 | GCP (GKE + GCS + Media CDN) 129 of 129 rows touched | 4 | 4 | 3 | 4 | 4 | 4 | 3.80 | 5 | Lowest NO-EQ count of any full-platform candidate (29 DIRECT/48 EQUIV/16 REDESIGN/5 NO-EQ/31 RETIRE); only the CleverTap bucket is forced to stay on AWS. Arm64 in Mumbai, Dataplane V2, ESO v… |
| 3 | OCI (OKE + Object Storage + Cloudflare CDN) 129 of 129 rows touched | 4 | 4 | 2 | 4 | 4 | 3 | 3.45 | 6 | Strongest surprise: first-party Karpenter provider, Native Ingress Controller add-on, native ESO provider for OCI Vault, free inbound, unmetered intra-region, 300k-IOPS block, managed PG 15–… |
| 4 | Azure (AKS + Blob + Front Door) 129 of 129 rows touched | 4 | 2 | 2 | 4 | 5 | 3 | 3.20 | 8 | Best tooling of the ten providers (34/40), best identity after AWS (Entra Workload ID per ServiceAccount), AKS Node Auto Provisioning is managed Karpenter so 7 NodePools port near-1:1, arm64… |
| 5 | Cheap India compute (DigitalOcean/Vultr/OVH) + Cloudflare CDN 129 of 129 rows touched | 3 | 1 | 3 | 1 | 3 | 2 | 2.15 | 14 | DO is the easiest provider here for an agent to operate and the hardest to observe: --output json everywhere, real scopes, llms.txt, DOKS 1.35 + Cilium, VPC-native networking kills the 100.6… |
| 6 | Akamai compute + Cloudflare CDN (R2 origin) 129 of 129 rows touched | 2 | 2 | 3 | 1 | 1 | 2 | 1.90 | 13 | Better than pure Akamai on the edge, identical on everything that makes Akamai hard (23/35/24/13 NO-EQ/34). Cloudflare owns video CDN, R2 origin, edge compute, logs; Linode owns k8s, block, … |
| 7 | Akamai Cloud + Akamai CDN (the proposal) 129 of 129 rows touched | 2 | 2 | 2 | 1 | 1 | 2 | 1.70 | 13 | 17 DIRECT/44 EQUIV/19 REDESIGN/13 NO-EQ/36 RETIRE. Ranks below its Cloudflare-hybrid sibling on three edge-side counts: neither native Akamai token product reproduces the token format, so a … |
Each of the 129 things we run today, mapped against each destination. Click any cell for the target product, the concrete change, the effort and the evidence.
| Coupling | AWS + new CDN | GCP | OCI | Azure | Cheap + CF | Akamai + CF | Akamai |
|---|---|---|---|---|---|---|---|
| 12.1 Compute & cluster | |||||||
| C001EKS control plane | |||||||
| C002EKS access entries | |||||||
| C003EKS add-ons | |||||||
| C004VPC-CNI custom networking + prefix delegation + ENIConfig... | |||||||
| C005Karpenter 1.14 | |||||||
| C006Static managed nodegroups | |||||||
| C007arm64 Graviton fleet | |||||||
| C008Spot usage | |||||||
| C009media-jobs burst pool | |||||||
| C010ci-android pool | |||||||
| C011cloudflared-egress pool on public subnets | |||||||
| C012AZ-topology engineering | |||||||
| C013Zone-keyed scheduling primitives, by mechanism | |||||||
| C014system-cluster-critical on 6 debezium-server-* Deployment... | |||||||
| C015No ResourceQuota, no LimitRange, no custom PriorityClass,... | |||||||
| C016Descheduler 0.35.1 | |||||||
| C017Per-instance NIC allowance model | |||||||
| C018EC2 quotas | |||||||
| C019Ad-hoc video-processing c8g.12xlarge box | |||||||
| C020EBS non-PVC volumes | |||||||
| 12.2 Ingress, edge, DNS | |||||||
| C021ALB Ingress | |||||||
| C02217 TargetGroupBindings, target-type ip, cross-zone, dereg... | |||||||
| C023ACM certs ×2 on ALB | |||||||
| C024UAT data NLB | |||||||
| C025ExternalDNS | |||||||
| C026cert-manager internal CA | |||||||
| C027graceful-drain-controller | |||||||
| C028Cloudflare Tunnels ×3 prod / ×2 UAT, WARP private-net | |||||||
| C029Client-embedded otlp.prod.hoichoi.dev + unauthenticated O... | |||||||
| C030Two-layer content-API cache | |||||||
| C031Image CDN two-layer | |||||||
| C032Video CDN | |||||||
| C033Stream tokenization CFFs | |||||||
| C034Content-API CFF | |||||||
| C035CloudFront geo-restriction | |||||||
| C036CloudFront invalidation ×2 providers + cross-account CDN... | |||||||
| C037Private S3 origins | |||||||
| C038CloudFront cache/ORP/RHP policies | |||||||
| C039CloudFront access logs → S3 → Glue/Athena | |||||||
| C040WAF ACLs | |||||||
| C041Route53 hoichoicdn.com rollback zone | |||||||
| C042Dead/orphan CloudFront: EJ924Z64IYZ4W | |||||||
| C043Red-flag DNS: hoichoi.dev/www/backup.hoichoi.tv proxied A... | |||||||
| C044Static egress IPs | |||||||
| C045Payment-webhook ingress: API GW oh4nleqg1k | |||||||
| C046Android app constant SSLUtil.kt:38 = https://oh4nleqg1k.e... | |||||||
| C047pg_config.sslcommerzconfig.ssl_ipn_url | |||||||
| C048partner-api.hoichoi.tv API GW, 2 orphan prod API GW custo... | |||||||
| C049SPEKE API GWs ×2 | |||||||
| C050Cloudflare Snippet popular_search* hard-codes https://eks... | |||||||
| C051Tunnel colo affinity | |||||||
| C052IPv6 on hoichoi.dev stays OFF | |||||||
| C053"Existing Akamai touchpoint" | |||||||
| C054Hoichoi.Web.Clients CleverTap proxy upstream d2r1yp2w7bby... | |||||||
| 12.3 Storage & data plane | |||||||
| C055EBS CSI gp3 default SC | |||||||
| C056chi-analytics 1 TiB volume performance | |||||||
| C057chi-signoz | |||||||
| C058EFS CSI driver | |||||||
| C059Object storage overall | |||||||
| C060Object-store feature requirement | |||||||
| C061Media origin buckets | |||||||
| C062S3 event notifications on first-party buckets: 0 | |||||||
| C063Cross-account / third-party S3 principals: exactly one bu... | |||||||
| C064ClickHouse s3_cold disk | |||||||
| C065clickhouse-backup sidecar + CronJobs | |||||||
| C066PBM Mongo backups | |||||||
| C067Redpanda tiered storage + Cloud Topics | |||||||
| C068PostHog objstore / ai-blobs / warehouse / batch-export st... | |||||||
| C069Chatwoot ActiveStorage amazon → chatwoot-storage | |||||||
| C070agent-fs swarm-bucket-data-prod | |||||||
| C071RDS PostgreSQL 15.17 | |||||||
| C072Debezium CDC offsets | |||||||
| C073MongoDB access model: prod tls.mode: disabled + unsafeFla... | |||||||
| C074SigNoz CH | |||||||
| C075Backup transport for the migration | |||||||
| C076S3 Inventory | |||||||
| C077backup-freshness-probe | |||||||
| C078Migration archives | |||||||
| C079Terraform state buckets ×4 + DynamoDB locks ×3 | |||||||
| C080hoichoi-clevertap-sandbox CHI 400 GiB @ 750 MiB/s | |||||||
| C081ClickHouse data-plane state outside kustomize/helmfile | |||||||
| C082ClickHouse named collections → Kafka endpoints | |||||||
| 12.4 Identity, secrets, registry | |||||||
| C083EKS Pod Identity ×14 + IRSA ×4 | |||||||
| C084Prod ESO auth by accident via node role IMDS hop-2 | |||||||
| C085ESO ClusterSecretStore aws-ssm | |||||||
| C08663 SSM params with AWS-specific values | |||||||
| C087Static IAM keys in SSM | |||||||
| C088Secrets Manager ×5 prod / ×6 UAT | |||||||
| C089KMS | |||||||
| C090SSM AWSQuickSetup automation on UAT node launches | |||||||
| C091ECR | |||||||
| C092EKS add-on images 602401143452.dkr.ecr | |||||||
| C093agent-swarm-tools | |||||||
| C094CI/CD auth | |||||||
| C095Human access: 45 + 29 IAM users, Workmates SAML ×3, dorma... | |||||||
| C096Third-party / Org role trusts | |||||||
| C097Dead / ECS-era / MSP roles: 32 prod + 26 UAT | |||||||
| C098Account audit | |||||||
| 12.5 Application & serverless | |||||||
| C099Backend S3 providers | |||||||
| C100MediaConvert provider + job.service + EB rules media_cove... | |||||||
| C101EventBridge scheduled rules ×13 enabled | |||||||
| C102EventBridge fan-out | |||||||
| C103Payment webhook chain | |||||||
| C104Athena ingesters ×2 | |||||||
| C105Secrets Manager provider | |||||||
| C106hc-transcode | |||||||
| C107CleverTap import pipeline | |||||||
| C108PostHog SES | |||||||
| C109Lambda inventory: 1 live | |||||||
| C110Legacy Terraform-era residue: SQS ×7 idle, SNS ×7, CloudW... | |||||||
| C111GCP | |||||||
| C112hoichoi-llm-gateway | |||||||
| C113KMP Android SSLUtil.kt:38 IPN constant | |||||||
| 12.6 Observability & ops tooling | |||||||
| C114SigNoz two-tier collectors + zone pins + cloud: aws / eks... | |||||||
| C115Grafana CloudWatch datasource + aws/eks_cluster dashboards | |||||||
| C116CloudWatch control-plane logs, alarms, dashboard | |||||||
| C117Cost/metrics ergonomics: ce get-cost-and-usage by USAGE_T... | |||||||
| C118eksctl-specific runbooks | |||||||
| C119Docs/gotcha catalogues | |||||||
| 12.7 Egress & cost couplings | |||||||
| C120Internet DTO from the prod account | |||||||
| C121Content-API origin egress | |||||||
| C122Inbound to the prod account | |||||||
| C123Origin → CloudFront | |||||||
| C124One-off bulk copies | |||||||
| C125Cluster-attached S3 | |||||||
| C126Sooper-account CloudFront | |||||||
| C127NAT gateways ×3 | |||||||
| C128OTLP tunnel responses | |||||||
| C129L7 edge sizing | |||||||
Filtered to Critical and High. Cross-cutting flags fire on every destination.
Adopted reading (stated identically in §2.2 #3 and mappings/akamai.md §H #1): standard LKE is the costed baseline; LKE-E is upside contingent on Q1. Costed as absent: no scale-to-zero (7 autoscaling pools x ~$788/mo = ~$5,516/mo permanent floor), fixed identical CIDRs for every cluster (so the two-vnet WARP split is permanent), 250-node/1,000-pod caps, shared control plane, forced EOL upgrades with 48 h notice, no Kube-API aud…
Field evidence: BYO VPC parameters ignored and a fresh VPC built anyway; control-plane audit logs where "the API accepts the field, the apply reports success, and the cluster keeps reporting false"; the Kubernetes version list differing between two accounts queried in the same hour, with a valid pin becoming [400] k8s_version is not valid ~15 min into an apply.
Contract says video "must use... Stream", and Cloudflare "reserves the right to disable or limit your access to or use of the CDN" on use or suspicion of use. Every Pro-plan cost estimate is void; the price is a sales quote on no public page. Source severity: "Critical (commercial)".
The geography vocabulary is six hint labels; apac means "Asia-Pacific" and nothing finer; hints are best-effort and bind only at first creation of a bucket name; guaranteed placement is exactly eu/us/fedramp. "India" appears nowhere. Source severity: "Critical for a Kolkata/Mumbai-delivery workload".
Fastly and Akamai list both Kolkata and Hyderabad; Jio (AS55836) publishes nothing (PNI only, record stale since 2023-09-25). 32.8% of bytes are Kolkata, and the 2026-08-31 Jio-to-Frankfurt event is on record [inv L11]. Adopted reading, stated identically in §2.3 A13 and §9.3 CF3: PeeringDB presence is a fact about public peering only and neither fact settles delivery quality for either vendor, because most Indian eyeball traf…
Source marks this "(Critical for R5)".
The migration is the first restore.
A rebuild from git yields 7 users and 0 grants - CDC, Metabase, support-api, the CleverTap sandbox and every analyst lose their login.
It starts the moment the first byte is served by a non-AWS CDN. It is $0 today only because the puller is CloudFront.
Object Storage forces a choice between the CORS-capable endpoint and the performant one; there is no single India region that satisfies both. Detail carried in §2.2 #6.
Block Storage tops out at 350 MB/s sustained, below what chi-analytics needs, pushing it onto unpublished local NVMe that a node recycle erases. Detail carried in §2.2 #5.
OpenBao (every secret), Harbor (every image) and key-rotation tooling each arrive with their own HA, backup and upgrade story, all on the critical path before the first workload moves.
Akamai billing observability does not meet requirement C12. Detail carried in §2.2 #15.
The LKE autoscaler may terminate nodes without draining them, and in other cases fails to scale down at all. Detail carried in §2.2 #13.
The AWS SDK checksum change breaks all seven S3 writers at once against Akamai Object Storage. Severity stated in source as "Medium-High". Detail carried in §2.2 #17.
The one daily-usage view excludes Enterprise contract accounts - the plan this design requires.
An external S3 bucket can be a Cloudflare origin only if it is public or fronted by something that authenticates Cloudflare. This matters exactly during the dual-CDN window where S3 is still the live origin - OAC is CloudFront-only, so S3 must be made readable via a credentialed Worker (~400M req/month) or a public-read prefix. Akamai has this natively as a Property Manager behaviour. Source severity: "Medium-High".
~85% adoption at 30 days, ~10% tail at >=90 days, TV/FireTV tail longer, no forced update. Any client-embedded-hostname change ships >=3 months before cutover, or the old endpoint stays alive >=3 months after. Source severity: "High (schedule)".
A 60-day window and a leave-AWS commitment - terms inferred from the 2024 policy and NOT VERIFIED. A dual-CDN period with S3 still live is ordinary traffic and is not covered.
Only 145,840 objects are referenced. Un-version it before any copy or it breaches object quotas on Akamai E1 and DigitalOcean Spaces on arrival.
Prod ch-cold lifecycle is abort-MPU only - the 90 d IA / 365 d DA tiers promised in the storage-policy comments were applied on UAT only.
TLS cannot be toggled on a populated cluster - the fresh install on the target is the only window.
Unauthenticated OTLP ingest carrying user identifiers (0 x 401 in 31 d), a UAT internet-facing NLB on 6 DB ports to 0.0.0.0/0, hoichoi-prodrevamp-costusage-report-bucket with Principal * Allow s3:* and PAB off (public read/write/delete), ClickHouse cdc_user with no_password, dormant enabled IAM keys, static CI keys, plaintext secrets in tfstate.
Thirty questions for Akamai and sixteen proof-of-concept tests with pass criteria taken from our own measurements.
in-bom-2 (Mumbai 2) today by a new customer, and what is the GA date? Confirm in …
blockingin-bom-2 (Mumbai 2) today by a new customer, and what is the GA date? Confirm in writing that (a) the creation gate is the Kubernetes Enterprise capability and not ACLP Logs Datacenter LKE-E, and (b) once enrolled, GET /v4/regions on our account returns Kubernetes Enterprise for in-bom-2. Chennai is already excluded on our side — in-maa returns neither flag (F).
Why it mattersCategory: blocking the architecture. Without LKE-E in in-bom-2 there is no target cluster region; Chennai is already excluded.in-bom-2 granted on request to a new account with no billing history, and when does it reach GA? St…
blockingin-bom-2 granted on request to a new account with no billing history, and when does it reach GA? State any SLA that does or does not apply under Limited Availability. (We are not asking whether the endpoint exists — we are asking about grant policy and LA terms.)
Why it mattersCategory: blocking the architecture. The object store is the media origin; an LA grant with no SLA is not a production origin.100.64.0.0/12 [inv R12].
Why it mattersCategory: blocking the architecture. CIDR choices are create-time and irreversible; overlapping prod/UAT CIDRs block peering.recycle? Worker nodes are Linodes and the …
blockingrecycle? Worker nodes are Linodes and the assignment API takes them; what is missing is LKE-side re-attach automation, and CCM v0.9.8 removed the Cilium BGP path. If no: confirm (a) our account is entitled to Reserved IPs, (b) IP Sharing is supported on instances that also hold a VPC interface, (c) whether Placement Groups can host that pair. 8 partners have our 3 NAT EIPs allow-listed [inv R10], [inv C3].
Why it mattersCategory: blocking the architecture. 8 partners allow-list our 3 NAT EIPs; a drifting egress IP breaks partner integrations.system-cluster-critical outside kube-system, or inject a default `LimitRang…
blockingsystem-cluster-critical outside kube-system, or inject a default LimitRange/ResourceQuota? We run 8 such pods and have zero LimitRanges across 192 priority-0 pods [inv R25].
Why it mattersCategory: blocking the architecture. An injected LimitRange or priority-class rejection would break 8 running pods on arrival.g8-dedicated-128-32 and g8-dedicated-64-32? Published nowhere. c…
blockingg8-dedicated-128-32 and g8-dedicated-64-32? Published nowhere. chi-analytics needs >= 4–5 k IOPS and >= 600 MiB/s sustained [inv R8], and Block Storage at 350 MB/s cannot do it.
Why it mattersCategory: blocking the data plane. Block Storage at 350 MB/s cannot host chi-analytics; local NVMe is the only candidate and is unpublished.If-Match / If-None-Match conditional writes and bulk DeleteObjects? Neither appears in any …
blockingIf-Match / If-None-Match conditional writes and bulk DeleteObjects? Neither appears in any doc. Our encoder uses If-Match; Redpanda tiered-storage GC uses bulk delete [inv §12.3]. Also: is ExpiredObjectDeleteMarker supported in lifecycle policies?
Why it mattersCategory: blocking the data plane. The encoder needs If-Match and Redpanda tiered-storage GC needs bulk delete; each miss is a code change.wal_level by default, and what are the server locale and encodi…
blockingwal_level by default, and what are the server locale and encoding? We need pgvector >= 0.8, logical replication, and en_US.UTF-8 with glibc collation [inv R16].
Why it mattersCategory: blocking the data plane. Decides Managed PostgreSQL vs falling back to in-cluster CNPG.onClientRequest and the other standard handlers. A "no, Basic cannot" answer does not close this — Q27 is the follow-on and must be answered in the same reply.
Why it mattersCategory: blocking the CDN workstream. The token path is mandatory; a tier that cannot run it is a blocker, not a cost line. Answer with Q27.in-bom-2 (and in-maa) for 46 simultaneous g8-dedicated-64-32 instances (1,472 vCPU) on a burst bas…
in-bom-2 (and in-maa) for 46 simultaneous g8-dedicated-64-32 instances (1,472 vCPU) on a burst basis, and what is the provisioning latency from zero [inv R5]?
Why it mattersCategory: commercial and compliance. Encoder burst today peaks at 25 nodes / 484 cores with a 46-node cap; no capacity means no burst lane.mappings/akamai.md §A classes NO-EQUIVALENT and budgets 2–3 weeks of SQL rewrite against.
Why it mattersCategory: blocking the CDN workstream. Sole replacement for the CloudFront-logs -> Glue -> Athena path; 2–3 weeks of SQL rewrite budgeted.cloudfront-distribution-id SSM values (mappings/akamai.md §A L95).
Why it mattersCategory: blocking the CDN workstream. Only replacement for the two CloudFront-invalidation providers in our code.GET /account/invoices/{id}/items is monthly per-service (F) and GET /account/transfer is month-to-date only?
Why it mattersCategory: commercial and compliance. Billing observability carries 15 % of the ranking; P15's acceptance criterion is literally this commitment. Mirrors CF-d.fio on one Block Storage volume, then on mdadm RAID-0 over two, on a dedicated-CPU instance in the target region.
Passes if>= 4–5 k IOPS sustained and >= 600 MiB/s, to match chi-analytics' measured p95 3,000 IOPS (at cap 8 % of minutes) and 566 MiB/s peak [inv R8]. Fail: chi-analytics moves to unpublished local NVMe and a second ClickHouse replica becomes mandatory.fio on the plan-local disk of g8-dedicated-128-32.
Passes ifSame numbers as P1 (>= 4–5 k IOPS, >= 600 MiB/s), plus: state the behaviour on node recycle. Fail: no home for chi-analytics on Akamai.Hoichoi.Encoder/storage_test.go plus a 100 MB multipart upload against a real bucket.
Passes ifMust pass: MPU (5 MiB-5 GiB parts, 16 MiB Redpanda), Range GET, ListObjectsV2, bulk DeleteObjects, presigned PUT + CORS with PUT/POST/HEAD, prefix-filtered Expiration.Days, AbortIncompleteMultipartUpload, If-Match, PutObjectTagging. If-Match and bulk delete have no documented answer.request_checksum_calculation=WHEN_REQUIRED.
Passes ifAll 7 S3 writers must work: hc-transcode, Thumbor, PBM, clickhouse-backup, Redpanda, PostHog/chdb, CH cold disk. Fail: per-client remediation, and Akamai's own doc says the workaround "may not work in all cases".min = 0 through the API and Terraform, on our own account, in the target region.
Passes ifMust succeed without a support ticket, and the pool must return from 0 on a pending pod. Fail confirms a permanent floor of ~$788/mo per autoscaling pool x 7 pools = ~$5,516/mo on standard LKE.terminationGracePeriodSeconds: 120 is running.
Passes ifThe pod must receive SIGTERM and >= 120 s before the node dies. Today's contract: EventBridge -> SQS -> Karpenter drain -> DisruptionTarget + SIGTERM, the only pre-emption notice the encoder relies on. Fail: every reclaim costs up to one checkpoint interval; PVs may fail to detach on stateful pools./dev/kvm present in a privileged pod.
Passes if12 vCPU / 24 Gi request, 200 GiB root @ 6,000 IOPS / 1,000 MB/s, up to 8 runners, scale-to-zero. NOT a blocker: 45 dispatches / 2 successes, dormant since 2026-09-01, ~$11/month — keep on AWS or GitHub larger runners.SELECT extversion FROM pg_extension WHERE extname='vector', SHOW wal_level, SHOW lc_collate, SHOW max_connections, SHOW default_toast_compression.
Passes ifpgvector >= 0.8, wal_level=logical, en_US.UTF-8 glibc, max_connections >= 600 (838 configured / 464 peak today), lz4 [inv R16]. Fail: fall back to in-cluster CNPG — which [inv D15] prefers anyway.ingress-nginx behind a NodeBalancer and drive it.
Passes if511 req/s and 22.6 new TLS conn/s at peak hour (1.84 M req/h, 81 k new TLS conn/h); size by RPS and connections/s, not bytes — peak is only 0.06 Gbps (5-min). Confirm the idle timeout and that client_conn_throttle is 0. Fail: non-premium NodeBalancer caps at 10,000 concurrent connections.aws ce get-cost-and-usage --granularity DAILY --group-by USAGE_TYPE gives today [inv R20]. Expected to fail — then the acceptance criterion becomes what Akamai will commit to as a contractual substitute, which is RFP question Q26.Phase 0 is worth doing whatever we decide. It is the first restore drill this estate has ever run.
De-risk any move with 15 workstreams that are all independently worth doing on AWS. 10 of 15 are already-listed standing defects; if the estate stays on AWS, Phase 0 still happens. Weeks 1-8 (2026-09-21 to 2026-11-15).
Author fresh with OpenTofu because there is nothing to export: L0 state last written 2025-03-13, L1 tofu never merged, prod EKS not under IaC at all, UAT tofu five months drifted. Everything above Kubernetes stays in helmfile + kustomize. Weeks 4-10. On rank-1 this phase is: import Cloudflare into Terraform, and nothing else.
Seed the registry, then apply 27 helmfile releases in dependency order across 8 waves, re-keying the zone affinities and replacing Karpenter with whatever the candidate's autoscaler is. Weeks 8-16, overlapping Phase 1's tail.
Move ~17.9 TB across the AWS edge per full pass (media library 14,753 GB = 82%) and prove every store restores, reconciles and resumes. Budget >= 2 full passes. Weeks 12-20. On rank-1 this is not a full phase: P0-04 IS Phase 3, plus the Typesense and Valkey rebuild checks.
Move the 97.2% of in-scope spend that is CloudFront egress in one account, by porting the token verifier to the new edge, canarying on jio.partner, ramping the apex by DNS weight, and moving the origin. Weeks 6-24, on the assumption the CDN contract is signed by end of week 5. On rank-1 this phase IS the whole project.
One scheduled window, announced at 6 h with an internal target of ~4h45m and an abort at T+5h00 wall clock from T-0, rehearsed twice in Phase 3. Weeks 21-22 (2027-02-08 to 2027-02-21). Does not exist on rank-1.
Delete nothing for 4 weeks, then retire the old estate on a +0 to +8 week schedule from the Phase 5 window (or from Phase 4 completion on rank-1). Decommission spends no egress, so the credit expiry of 2027-02-28 is a restatement of the week-20 pin on the bulk copies, not a separate deadline this phase can miss.
| Metric | Production | UAT | CDN account |
|---|---|---|---|
| compute | |||
| Nodes (instantaneous) | 6 static + 8–11 Karpenter (typ. 15–17), all arm64 | 18 (16 OD + 2 spot), all arm64 | — |
| vCPU / RAM present | snapshot 64 vCPU / ~362 GiB; 30-d mean 16.4 nodes ≈ 73 vCPU / ~410 GiB steady pools | allocatable 37.4 cores / 128.6 GiB | — |
| 30-d usage p95 (steady pools) | 25.2 cores / 185 GiB (peak30m 52 / 227) | ≈ 6.6 cores busy of 42 vCPU; memory ~75 % committed | — |
| Spot share | 29 % of EC2 $ (10 % of instance-hours); apps pool 100 % spot, 753 node replacements / 30 d | 2 spot nodes only | — |
| IAM roles (RoleLastUsed, measured) | 108: 9 app/data roles to rebuild (all 9 used in-window), 11 EKS plumbing, 10 other AWS-native, 11 external/Org trusts, 32 ECS/MSP dead, 35 service-linked | 95: 4 app/data live (mongodb-backup-role never used), 6 EKS plumbing, 25 other, 11 external, 26 dead, 23 SL | NOT MEASURED (iam:ListRoles denied) |
| storage | |||
| PVCs / EBS | 44 PVCs / 5,012 GiB gp3; EBS total 66–69 vols / 6,185–6,505 GiB (superseding 64 vols / 6,065 GiB; two captures 2 h apart on 2026-09-18 differed by +3 vols / +320 GiB); 0 true orphan volumes; only stateful non-PVC disks are 5 stopped legacy EC2 roots (413 GiB, all RETIRE except video-processing, restarted 2026-09-18). Node roots: Karpenter apps/analytics-db/clevertap-sandbox 60 Gi (git says 30 Gi — drift), batch-jobs 200 Gi, media-jobs 120 Gi, ci-android-kvm 200 Gi, cloudflared-egress 20 Gi, managed NGs 30 Gi | 41 Bound + 1 Pending / 1,110 GiB (EBS 77 vols / 2,390 GiB, 24 orphan = 552 GiB) | — |
| S3 | 45 buckets, 23,771 GiB (= 25,406.9 GB) / 95.8 M objects; 11 versioned, 11 with lifecycle (23 rules, prefix-filtered only), 7 CORS, 3 public-policy (1 intended), 0 replication / object-lock / KMS | 32 buckets (superseding 32–33), 4.09 TB (3,727 GiB Std + 135 IA); 5 versioned, 5 lifecycle, 6 CORS, 1 public + static-website | 7 buckets, ≈ 18.7 TB CloudFront logs (CE-inferred) |
| Cluster-attached S3 cost (ch-cold, redpanda-tiered, 2 backup buckets) | ≈ $350 window / ≈ $307 forward (superseding the ≈ $200 inference) = storage $247 (10.1 TB window mean, 8.17 TB steady) + requests $103 measured (38.5 M GET + 17.3 M PUT) + ≤ $27 inferred backup ops | — | — |
| edge | |||
| CloudFront origin fetch 30-d | 241,371 GiB origin→CloudFront (superseding 238,118 GiB) = S3 240,618 GiB (video) + ALB→CloudFront 749 GiB (content-API + Thumbor); $0 only while CloudFront is the puller (≈ $20 K/mo at DTO list price if a non-AWS CDN pulls the same way without a shield) | 0.75 GB | 2,928,681 GiB = 2.79 PiB = 3.145 PB to viewers (4.93 B req) |
| Cloudflare edge (all zones, 31 d) | ≈ 2.80 B requests / 22.9 TB: hoichoi.dev 2.39 B / 13.1 TB (⚠ CONFLICT: [06 §4] 12.2 TB vs [03 R2-01.2] 13.15 TB; canonical 13.1 TB) at 1.2 % req cache-hit; hoichoicdn.com 268 M / 9.47 TB in 18 d (97.2 % hit); hoichoi.tv 140 M / 1.16 TB. Uncached from AWS origins ≈ 10.35 TB / 31 d (79 % = content-API JSON) | — | — |
| Cloudflare Workers | 55.55 M invocations (hoichoi-web 54.7 M) | — | — |
| Tunnel bytes (31 d) | cloudflared-otlp 51.5 TB in / 976 M req (52.7 KB/req; Aug ≈ 2.2 TB/day, Sep ≈ 1.2 TB/day); cloudflared 1.6 TB (≈ 97 % PostHog ingest); cloudflared-private-net 35 GB | — | — |
| network | |||
| Internet egress (non-CDN), measured split | 2,517 GiB · $229.72 (superseding 2,487 GB · $227.14) = EC2 1,606 GB (tunnel responses ≈ 1,331 + NAT 275) + ELB 852 GB + ECR/S3 59 GB; Sep run-rate ≈ 1.8–2.0 TB/mo (tunnel half halved from 09-02) | 74 GB · $7.12 | — |
| Internet inbound (free) | 56.0 TB: EC2 52.7 TB (OTLP tunnel ingest), ELB 2.25 TB, S3 1.08 TB | — | — |
| Cross-AZ (Regional-Bytes) | 28,771 GB · $287.70 | 6,626 GB · $66.26 | — |
| NAT | 844 GB processed · 3 GW · $170.56 (275 GB out / 640 GB in; tunnel nodes bypass NAT via public IPs) | 116.7 GB · 1 GW · $47.64 | — |
| ALB | 414.9 M req / 3.68 TB; peak-hour 1.84 M req / 18.6 GB = 0.041 Gbps, peak 5-min 0.060 Gbps, 81 k new TLS conn/h | 31.2 M req / 58.9 GB | — |
| Cluster east-west (pod traffic) | ≈ 470 TB / 31 d, dominated by monitoring↔egress telemetry (110 TB), data replication (88 TB), apps (65 TB); media pool only 2.8 TB | NOT MEASURED | — |
| data | |||
| Redpanda Kafka bytes | 159 GB/day produced, 285 GB/day consumed (98 % = 3 PostHog Cloud Topics; clickhouse_events_json re-read 4×); wire ≈ 209 GB/day in / 420 out; S3 side ≈ 587 GB/day PUT / 220 GET | — | — |
| RDS | 1 × db.t4g.large PG 15.17, 400 GB gp3 (12 k IOPS), ~341 GB used, single-AZ, 1-day backups, $196.97 | 1 × db.m6g.large, 200 GB gp2, 52 GB used, publicly accessible, $194.13 | — |