Shabinder — DocsReports
Infrastructure review · measured 18 Aug – 18 Sep 2026

Where Hoichoi's infrastructure should live

Seven destinations scored against 129 measured AWS couplings, read live from two Kubernetes clusters, six AWS accounts and the Cloudflare account.

Couplings129
Candidates7
Claims verified84
Refuted39
The finding

Do not move the platform to Akamai. The candidate that best fits this estate is the one that moves the 97 % and leaves the 3 %: keep compute, data plane, identity and CI on AWS EKS, and move only the video CDN and its origin — to Cloudflare Enterprise (variant B), with Akamai AMD (variant A) carried into the same RFP as the alternative.

Where the money is 97.2% $235,465 of $242,300 monthly in-scope spend is video delivery from one account. Everything else — both clusters, every database, all storage — is about $7,000.
Video CDN egressAll compute + data

Candidate ranking

pricing excluded from scoring, by instruction

Scores are 0–5, 5 best. Note the sign convention: operational risk is scored as "risk handled well", so 5 = lowest risk.

#Destinationk8sDataCDNIdentAgentsRiskScoreBlockersWhy
1 Stay on AWS, move only the video CDN + origin 25 of 129 rows touched 553554 4.45 3 104 of 129 coupling rows UNCHANGED (4 DIRECT/5 EQUIV/5 REDESIGN/3 NO-EQ/8 RETIRE). Variant B (Cloudflare Enterprise, R2 origin, Worker token) is the recommended start; variant A (Akamai AMD)…
2 GCP (GKE + GCS + Media CDN) 129 of 129 rows touched 443444 3.80 5 Lowest NO-EQ count of any full-platform candidate (29 DIRECT/48 EQUIV/16 REDESIGN/5 NO-EQ/31 RETIRE); only the CleverTap bucket is forced to stay on AWS. Arm64 in Mumbai, Dataplane V2, ESO v…
3 OCI (OKE + Object Storage + Cloudflare CDN) 129 of 129 rows touched 442443 3.45 6 Strongest surprise: first-party Karpenter provider, Native Ingress Controller add-on, native ESO provider for OCI Vault, free inbound, unmetered intra-region, 300k-IOPS block, managed PG 15–…
4 Azure (AKS + Blob + Front Door) 129 of 129 rows touched 422453 3.20 8 Best tooling of the ten providers (34/40), best identity after AWS (Entra Workload ID per ServiceAccount), AKS Node Auto Provisioning is managed Karpenter so 7 NodePools port near-1:1, arm64…
5 Cheap India compute (DigitalOcean/Vultr/OVH) + Cloudflare CDN 129 of 129 rows touched 313132 2.15 14 DO is the easiest provider here for an agent to operate and the hardest to observe: --output json everywhere, real scopes, llms.txt, DOKS 1.35 + Cilium, VPC-native networking kills the 100.6…
6 Akamai compute + Cloudflare CDN (R2 origin) 129 of 129 rows touched 223112 1.90 13 Better than pure Akamai on the edge, identical on everything that makes Akamai hard (23/35/24/13 NO-EQ/34). Cloudflare owns video CDN, R2 origin, edge compute, logs; Linode owns k8s, block, …
7 Akamai Cloud + Akamai CDN (the proposal) 129 of 129 rows touched 222112 1.70 13 17 DIRECT/44 EQUIV/19 REDESIGN/13 NO-EQ/36 RETIRE. Ranks below its Cloudflare-hybrid sibling on three edge-side counts: neither native Akamai token product reproduces the token format, so a …
Weights: Data-plane fit 20% · CDN fit 20% · Portability of our k8s platform 15% · Identity / secrets / CI fit 15% · Agentic-ops ergonomics 15% · Operational risk 15%

Decision matrix

every coupling × every candidate

Each of the 129 things we run today, mapped against each destination. Click any cell for the target product, the concrete change, the effort and the evidence.

CouplingAWS + new CDNGCPOCIAzureCheap + CFAkamai + CFAkamai
12.1  Compute & cluster
C001EKS control plane
C002EKS access entries
C003EKS add-ons
C004VPC-CNI custom networking + prefix delegation + ENIConfig...
C005Karpenter 1.14
C006Static managed nodegroups
C007arm64 Graviton fleet
C008Spot usage
C009media-jobs burst pool
C010ci-android pool
C011cloudflared-egress pool on public subnets
C012AZ-topology engineering
C013Zone-keyed scheduling primitives, by mechanism
C014system-cluster-critical on 6 debezium-server-* Deployment...
C015No ResourceQuota, no LimitRange, no custom PriorityClass,...
C016Descheduler 0.35.1
C017Per-instance NIC allowance model
C018EC2 quotas
C019Ad-hoc video-processing c8g.12xlarge box
C020EBS non-PVC volumes
12.2  Ingress, edge, DNS
C021ALB Ingress
C02217 TargetGroupBindings, target-type ip, cross-zone, dereg...
C023ACM certs ×2 on ALB
C024UAT data NLB
C025ExternalDNS
C026cert-manager internal CA
C027graceful-drain-controller
C028Cloudflare Tunnels ×3 prod / ×2 UAT, WARP private-net
C029Client-embedded otlp.prod.hoichoi.dev + unauthenticated O...
C030Two-layer content-API cache
C031Image CDN two-layer
C032Video CDN
C033Stream tokenization CFFs
C034Content-API CFF
C035CloudFront geo-restriction
C036CloudFront invalidation ×2 providers + cross-account CDN...
C037Private S3 origins
C038CloudFront cache/ORP/RHP policies
C039CloudFront access logs → S3 → Glue/Athena
C040WAF ACLs
C041Route53 hoichoicdn.com rollback zone
C042Dead/orphan CloudFront: EJ924Z64IYZ4W
C043Red-flag DNS: hoichoi.dev/www/backup.hoichoi.tv proxied A...
C044Static egress IPs
C045Payment-webhook ingress: API GW oh4nleqg1k
C046Android app constant SSLUtil.kt:38 = https://oh4nleqg1k.e...
C047pg_config.sslcommerzconfig.ssl_ipn_url
C048partner-api.hoichoi.tv API GW, 2 orphan prod API GW custo...
C049SPEKE API GWs ×2
C050Cloudflare Snippet popular_search* hard-codes https://eks...
C051Tunnel colo affinity
C052IPv6 on hoichoi.dev stays OFF
C053"Existing Akamai touchpoint"
C054Hoichoi.Web.Clients CleverTap proxy upstream d2r1yp2w7bby...
12.3  Storage & data plane
C055EBS CSI gp3 default SC
C056chi-analytics 1 TiB volume performance
C057chi-signoz
C058EFS CSI driver
C059Object storage overall
C060Object-store feature requirement
C061Media origin buckets
C062S3 event notifications on first-party buckets: 0
C063Cross-account / third-party S3 principals: exactly one bu...
C064ClickHouse s3_cold disk
C065clickhouse-backup sidecar + CronJobs
C066PBM Mongo backups
C067Redpanda tiered storage + Cloud Topics
C068PostHog objstore / ai-blobs / warehouse / batch-export st...
C069Chatwoot ActiveStorage amazon → chatwoot-storage
C070agent-fs swarm-bucket-data-prod
C071RDS PostgreSQL 15.17
C072Debezium CDC offsets
C073MongoDB access model: prod tls.mode: disabled + unsafeFla...
C074SigNoz CH
C075Backup transport for the migration
C076S3 Inventory
C077backup-freshness-probe
C078Migration archives
C079Terraform state buckets ×4 + DynamoDB locks ×3
C080hoichoi-clevertap-sandbox CHI 400 GiB @ 750 MiB/s
C081ClickHouse data-plane state outside kustomize/helmfile
C082ClickHouse named collections → Kafka endpoints
12.4  Identity, secrets, registry
C083EKS Pod Identity ×14 + IRSA ×4
C084Prod ESO auth by accident via node role IMDS hop-2
C085ESO ClusterSecretStore aws-ssm
C08663 SSM params with AWS-specific values
C087Static IAM keys in SSM
C088Secrets Manager ×5 prod / ×6 UAT
C089KMS
C090SSM AWSQuickSetup automation on UAT node launches
C091ECR
C092EKS add-on images 602401143452.dkr.ecr
C093agent-swarm-tools
C094CI/CD auth
C095Human access: 45 + 29 IAM users, Workmates SAML ×3, dorma...
C096Third-party / Org role trusts
C097Dead / ECS-era / MSP roles: 32 prod + 26 UAT
C098Account audit
12.5  Application & serverless
C099Backend S3 providers
C100MediaConvert provider + job.service + EB rules media_cove...
C101EventBridge scheduled rules ×13 enabled
C102EventBridge fan-out
C103Payment webhook chain
C104Athena ingesters ×2
C105Secrets Manager provider
C106hc-transcode
C107CleverTap import pipeline
C108PostHog SES
C109Lambda inventory: 1 live
C110Legacy Terraform-era residue: SQS ×7 idle, SNS ×7, CloudW...
C111GCP
C112hoichoi-llm-gateway
C113KMP Android SSLUtil.kt:38 IPN constant
12.6  Observability & ops tooling
C114SigNoz two-tier collectors + zone pins + cloud: aws / eks...
C115Grafana CloudWatch datasource + aws/eks_cluster dashboards
C116CloudWatch control-plane logs, alarms, dashboard
C117Cost/metrics ergonomics: ce get-cost-and-usage by USAGE_T...
C118eksctl-specific runbooks
C119Docs/gotcha catalogues
12.7  Egress & cost couplings
C120Internet DTO from the prod account
C121Content-API origin egress
C122Inbound to the prod account
C123Origin → CloudFront
C124One-off bulk copies
C125Cluster-attached S3
C126Sooper-account CloudFront
C127NAT gateways ×3
C128OTLP tunnel responses
C129L7 edge sizing

Red flags

80 across all candidates

Filtered to Critical and High. Cross-cutting flags fire on every destination.

A1 Critical Akamai

LKE Enterprise availability in India cannot be established from any public source

Adopted reading (stated identically in §2.2 #3 and mappings/akamai.md §H #1): standard LKE is the costed baseline; LKE-E is upside contingent on Q1. Costed as absent: no scale-to-zero (7 autoscaling pools x ~$788/mo = ~$5,516/mo permanent floor), fixed identical CIDRs for every cluster (so the two-vnet WARP split is permanent), 250-node/1,000-pod caps, shared control plane, forced EOL upgrades with 48 h notice, no Kube-API aud…

A2 Critical Akamai

LKE-E entitlement is per-account and per-feature, and it fails silently

Field evidence: BYO VPC parameters ignored and a fresh VPC built anyway; control-plane audit logs where "the API accepts the field, the apply reports success, and the cluster keeps reporting false"; the Kubernetes version list differing between two accounts queried in the same hour, with a valid pin becoming [400] k8s_version is not valid ~15 min into an apply.

CF1 Critical all candidates

Video on a non-Enterprise Cloudflare zone is contractually disallowed

Contract says video "must use... Stream", and Cloudflare "reserves the right to disable or limit your access to or use of the CDN" on use or suspicion of use. Every Pro-plan cost estimate is void; the price is a sales quote on no public page. Source severity: "Critical (commercial)".

CF2 Critical all candidates

R2 has no India location and no India jurisdiction, by design

The geography vocabulary is six hint labels; apac means "Asia-Pacific" and nothing finer; hints are best-effort and bind only at first creation of a bucket name; guaranteed placement is exactly eu/us/fedramp. "India" appears nowhere. Source severity: "Critical for a Kolkata/Mumbai-delivery workload".

CF3 Critical all candidates

No Kolkata IX port for AS13335 in PeeringDB

Fastly and Akamai list both Kolkata and Hyderabad; Jio (AS55836) publishes nothing (PNI only, record stale since 2023-09-25). 32.8% of bytes are Kolkata, and the 2026-08-31 Jio-to-Frankfurt event is on record [inv L11]. Adopted reading, stated identically in §2.3 A13 and §9.3 CF3: PeeringDB presence is a fact about public peering only and neither fact settles delivery quality for either vendor, because most Indian eyeball traf…

GCP-3 Critical GCP

Spot reclaim is 15 s by default, 120 s is Preview, and docs contradict the TF provider

Source marks this "(Critical for R5)".

X1 Critical all candidates

No restore drill has ever been run; zero EBS snapshots of any stateful PVC

The migration is the first restore.

X2 Critical all candidates

12 ClickHouse SQL users, 1 role, 1 settings profile and 323 grants exist in no file

A rebuild from git yields 7 users and 0 grants - CDC, Metabase, support-api, the CleverTap sandbox and every analyst lose their login.

X3 Critical all candidates

The dual-CDN origin-fetch window costs ~$20,019/month unshielded

It starts the moment the first byte is served by a non-AWS CDN. It is $0 today only because the puller is CloudFront.

A3 High Akamai

CORS/performance split in Object Storage has no single-India-region answer

Object Storage forces a choice between the CORS-capable endpoint and the performant one; there is no single India region that satisfies both. Detail carried in §2.2 #6.

A4 High Akamai

Block Storage 350 MB/s sustained forces chi-analytics onto unpublished local NVMe

Block Storage tops out at 350 MB/s sustained, below what chi-analytics needs, pushing it onto unpublished local NVMe that a node recycle erases. Detail carried in §2.2 #5.

A5 High Akamai

Three self-hosted components on day one: OpenBao, Harbor, key-rotation tooling

OpenBao (every secret), Harbor (every image) and key-rotation tooling each arrive with their own HA, backup and upgrade story, all on the critical path before the first workload moves.

A6 High Akamai

Billing observability fails C12 outright

Akamai billing observability does not meet requirement C12. Detail carried in §2.2 #15.

A7 High Akamai

Autoscaler may terminate nodes without draining, and sometimes does not scale down

The LKE autoscaler may terminate nodes without draining them, and in other cases fails to scale down at all. Detail carried in §2.2 #13.

A8 High Akamai

AWS SDK checksum break hits all seven S3 writers simultaneously

The AWS SDK checksum change breaks all seven S3 writers at once against Akamai Object Storage. Severity stated in source as "Medium-High". Detail carried in §2.2 #17.

CF4 High all candidates

Cloudflare billing is not programmatic

The one daily-usage view excludes Enterprise contract accounts - the plan this design requires.

CF8 High all candidates

Cloudflare has no S3 SigV4 origin signing

An external S3 bucket can be a Cloudflare origin only if it is public or fronted by something that authenticates Cloudflare. This matters exactly during the dual-CDN window where S3 is still the live origin - OAC is CloudFront-only, so S3 must be made readable via a credentialed Worker (~400M req/month) or a public-read prefix. Akamai has this natively as a Property Manager behaviour. Source severity: "Medium-High".

X10 High all candidates

Android release lead time is a migration clock

~85% adoption at 30 days, ~10% tail at >=90 days, TV/FireTV tail longer, no forced update. Any client-embedded-hostname change ships >=3 months before cutover, or the old endpoint stays alive >=3 months after. Source severity: "High (schedule)".

X4 High all candidates

DTO waiver needs a support ticket approved before the transfer starts

A 60-day window and a leave-AWS commitment - terms inferred from the 2024 policy and NOT VERIFIED. A dual-CDN period with S3 still live is ordinary traffic and is not covered.

X5 High all candidates

ch-cold is 52.1M versioned objects of which 99.7% is garbage

Only 145,840 objects are referenced. Un-version it before any copy or it breaches object quotas on Akamai E1 and DigitalOcean Spaces on arrival.

X6 High all candidates

The CH cold tier has no backup and a single-PVC metadata SPOF

Prod ch-cold lifecycle is abort-MPU only - the 90 d IA / 365 d DA tiers promised in the storage-policy comments were applied on UAT only.

X7 High all candidates

Prod Mongo runs wire TLS disabled with SCRAM only and no NetworkPolicy on 27017

TLS cannot be toggled on a populated cluster - the fresh install on the target is the only window.

X8 High all candidates

Security baseline items to fix before, not port

Unauthenticated OTLP ingest carrying user identifiers (0 x 401 in 31 d), a UAT internet-facing NLB on 6 DB ports to 0.0.0.0/0, hoichoi-prodrevamp-costusage-report-bucket with Principal * Allow s3:* and PAB off (public read/write/delete), ClickHouse cdc_user with no_password, dormant enabled IAM keys, static CI keys, plaintext secrets in tfstate.

Questions for the vendor

ask before signing anything

Thirty questions for Akamai and sixteen proof-of-concept tests with pass criteria taken from our own measurements.

Request for proposal ·
Q1Can LKE Enterprise clusters be created in in-bom-2 (Mumbai 2) today by a new customer, and what is the GA date? Confirm in … blocking
Full questionCan LKE Enterprise clusters be created in in-bom-2 (Mumbai 2) today by a new customer, and what is the GA date? Confirm in writing that (a) the creation gate is the Kubernetes Enterprise capability and not ACLP Logs Datacenter LKE-E, and (b) once enrolled, GET /v4/regions on our account returns Kubernetes Enterprise for in-bom-2. Chennai is already excluded on our side — in-maa returns neither flag (F). Why it mattersCategory: blocking the architecture. Without LKE-E in in-bom-2 there is no target cluster region; Chennai is already excluded.
Q2Is Object Storage E3 at in-bom-2 granted on request to a new account with no billing history, and when does it reach GA? St… blocking
Full questionIs Object Storage E3 at in-bom-2 granted on request to a new account with no billing history, and when does it reach GA? State any SLA that does or does not apply under Limited Availability. (We are not asking whether the endpoint exists — we are asking about grant policy and LA terms.) Why it mattersCategory: blocking the architecture. The object store is the media origin; an LA grant with no SLA is not a production origin.
Q3On LKE Enterprise, is the pod CIDR unique per cluster and can we specify it? We need prod and UAT non-overlapping on both pod… blocking
Full questionOn LKE Enterprise, is the pod CIDR unique per cluster and can we specify it? We need prod and UAT non-overlapping on both pod and service CIDRs, none inside 100.64.0.0/12 [inv R12]. Why it mattersCategory: blocking the architecture. CIDR choices are create-time and irreversible; overlapping prod/UAT CIDRs block peering.
Q4Is there a supported way to pin a Reserved IPv4 to LKE worker-node egress across recycle? Worker nodes are Linodes and the … blocking
Full questionIs there a supported way to pin a Reserved IPv4 to LKE worker-node egress across recycle? Worker nodes are Linodes and the assignment API takes them; what is missing is LKE-side re-attach automation, and CCM v0.9.8 removed the Cilium BGP path. If no: confirm (a) our account is entitled to Reserved IPs, (b) IP Sharing is supported on instances that also hold a VPC interface, (c) whether Placement Groups can host that pair. 8 partners have our 3 NAT EIPs allow-listed [inv R10], [inv C3]. Why it mattersCategory: blocking the architecture. 8 partners allow-list our 3 NAT EIPs; a drifting egress IP breaks partner integrations.
Q5What are the exact kube-reserved CPU millicores and memory on LKE worker nodes (clusters >= 2026-04-21 reserve "a small amoun… blocking
Full questionWhat are the exact kube-reserved CPU millicores and memory on LKE worker nodes (clusters >= 2026-04-21 reserve "a small amount" with no published figure)? Why it mattersCategory: blocking the architecture. Unpublished kube-reserved makes node sizing and the cost model unverifiable.
Q6Does LKE apply any admission policy rejecting system-cluster-critical outside kube-system, or inject a default `LimitRang… blocking
Full questionDoes LKE apply any admission policy rejecting system-cluster-critical outside kube-system, or inject a default LimitRange/ResourceQuota? We run 8 such pods and have zero LimitRanges across 192 priority-0 pods [inv R25]. Why it mattersCategory: blocking the architecture. An injected LimitRange or priority-class rejection would break 8 running pods on arrival.
Q7What are the actual IOPS and throughput of local NVMe on g8-dedicated-128-32 and g8-dedicated-64-32? Published nowhere. c… blocking
Full questionWhat are the actual IOPS and throughput of local NVMe on g8-dedicated-128-32 and g8-dedicated-64-32? Published nowhere. chi-analytics needs >= 4–5 k IOPS and >= 600 MiB/s sustained [inv R8], and Block Storage at 350 MB/s cannot do it. Why it mattersCategory: blocking the data plane. Block Storage at 350 MB/s cannot host chi-analytics; local NVMe is the only candidate and is unpublished.
Q8Is the 8,000 IOPS / 350 MB/s figure per-volume or per-Linode aggregate? This decides whether multi-volume striping can rescue… blocking
Full questionIs the 8,000 IOPS / 350 MB/s figure per-volume or per-Linode aggregate? This decides whether multi-volume striping can rescue chi-analytics. Why it mattersCategory: blocking the data plane. Decides whether mdadm RAID-0 over multiple volumes can reach the chi-analytics requirement.
Q9Does Object Storage support If-Match / If-None-Match conditional writes and bulk DeleteObjects? Neither appears in any … blocking
Full questionDoes Object Storage support If-Match / If-None-Match conditional writes and bulk DeleteObjects? Neither appears in any doc. Our encoder uses If-Match; Redpanda tiered-storage GC uses bulk delete [inv §12.3]. Also: is ExpiredObjectDeleteMarker supported in lifecycle policies? Why it mattersCategory: blocking the data plane. The encoder needs If-Match and Redpanda tiered-storage GC needs bulk delete; each miss is a code change.
Q10What is the typical PUT and range-GET latency for small objects on E3? Redpanda Cloud Topics were tuned against a 35–40 ms op… blocking
Full questionWhat is the typical PUT and range-GET latency for small objects on E3? Redpanda Cloud Topics were tuned against a 35–40 ms operation-latency class [inv R15]. Why it mattersCategory: blocking the data plane. Redpanda Cloud Topics were tuned against a 35–40 ms operation-latency class.
Q11Can a Block Storage volume be expanded online through the CSI driver on a running node, without powering the Linode off? blocking
Full questionCan a Block Storage volume be expanded online through the CSI driver on a running node, without powering the Linode off? Why it mattersCategory: blocking the data plane. Offline-only resize means a maintenance window per PVC across 44 prod PVCs.
Q12What per-bucket PUT rate can we be raised to on E3, and how quickly? Default 500/s, documented max 2,000/s (F). 28.4 M media … blocking
Full questionWhat per-bucket PUT rate can we be raised to on E3, and how quickly? Default 500/s, documented max 2,000/s (F). 28.4 M media objects at 500/s ~ 16 days; at 2,000/s ~ 4 days. Why it mattersCategory: blocking the data plane. At the 500/s default the 28.4 M-object copy takes ~16 days and does not fit the migration window.
Q13On Managed PostgreSQL: what pgvector version ships, what is wal_level by default, and what are the server locale and encodi… blocking
Full questionOn Managed PostgreSQL: what pgvector version ships, what is wal_level by default, and what are the server locale and encoding? We need pgvector >= 0.8, logical replication, and en_US.UTF-8 with glibc collation [inv R16]. Why it mattersCategory: blocking the data plane. Decides Managed PostgreSQL vs falling back to in-cluster CNPG.
Q14What is the per-GB Adaptive Media Delivery price for India at ~3.1 PB/month, and is India banded separately from APAC? Is ori… blocking
Full questionWhat is the per-GB Adaptive Media Delivery price for India at ~3.1 PB/month, and is India banded separately from APAC? Is origin egress from Akamai Object Storage waived or discounted under an AMD contract? Your own docs say outbound is metered even inside the same data center. Must be quoted against the fixed input pack (see why_it_matters); a per-GB quote against an unstated request profile is not a comparable quote. Why it mattersWhole video business case. Fixed pack: 3.145 PB/mo; 4.93 B req/mo; 13.5/21.3/36.2 Gbps mean/median-peak/max; 80.5% India (Kolkata 32.8, Delhi 22.0, Mumbai 8.8, Hyderabad 8.0); 91.9% byte-hit; 0.66-0.84 MB objects; m4s 55.7/ts 22.8/mp4 11.3/webm 10.1. Identical to CF-a.
Q15How many PoPs in India, specifically Kolkata, Dhaka, Mumbai, Chennai and Hyderabad? 80 % of our traffic is India; Kolkata alo… blocking
Full questionHow many PoPs in India, specifically Kolkata, Dhaka, Mumbai, Chennai and Hyderabad? 80 % of our traffic is India; Kolkata alone is 32.8 % [inv §1.3 #1]. Why it mattersCategory: blocking the CDN workstream. Kolkata alone is 32.8 % of bytes; PoP presence there decides delivered quality.
Q16Can an EdgeWorker on the Basic tier (10 ms CPU, 1.5 MB memory per standard event handler) sustain 4.7 billion HMAC token vali… blocking
Full questionCan an EdgeWorker on the Basic tier (10 ms CPU, 1.5 MB memory per standard event handler) sustain 4.7 billion HMAC token validations per month, and what do Dynamic/Enterprise tiers cost? Tier limits confirmed (F) — Basic 10 ms CPU / 1.5 MB, Dynamic 20 ms / 2.5 MB, Enterprise 70 ms / 4 MB on onClientRequest and the other standard handlers. A "no, Basic cannot" answer does not close this — Q27 is the follow-on and must be answered in the same reply. Why it mattersCategory: blocking the CDN workstream. The token path is mandatory; a tier that cannot run it is a blocker, not a cost line. Answer with Q27.
Q17Does Cloud Wrapper work with an Akamai Object Storage origin, at what price, and can it genuinely not be combined with Tiered… blocking
Full questionDoes Cloud Wrapper work with an Akamai Object Storage origin, at what price, and can it genuinely not be combined with Tiered Distribution? Why it mattersCategory: blocking the CDN workstream. Origin-shield shape drives origin egress cost on the Akamai design.
Q18Is Akamai Cloud (Linode) MeitY-empanelled in India? Supply the current SOC 2 report (state Type 1 vs Type 2) and the PCI DSS …
Full questionIs Akamai Cloud (Linode) MeitY-empanelled in India? Supply the current SOC 2 report (state Type 1 vs Type 2) and the PCI DSS attestation scope for Cloud Computing Services. Your compliance pages 404 to automated retrieval. This becomes pass/fail rather than a data point if §10.6 I-f returns an India-only residency position (§2.3 A12). Why it mattersCategory: commercial and compliance. Becomes pass/fail if Legal returns an India-only data-residency position (I-f).
Q19Confirm in writing: is network transfer overage $0.005/GB in Chennai and Mumbai, including NodeBalancer egress? Your API retu…
Full questionConfirm in writing: is network transfer overage $0.005/GB in Chennai and Mumbai, including NodeBalancer egress? Your API returns $0.005 (F); the APAC pricing page prints $0.015 next to NodeBalancers. Our cost model cannot be signed on a contradiction. Why it mattersCategory: commercial and compliance. A 3x contradiction between the API and the pricing page; the cost model cannot be signed on it.
Q20Is there capacity in in-bom-2 (and in-maa) for 46 simultaneous g8-dedicated-64-32 instances (1,472 vCPU) on a burst bas…
Full questionIs there capacity in in-bom-2 (and in-maa) for 46 simultaneous g8-dedicated-64-32 instances (1,472 vCPU) on a burst basis, and what is the provisioning latency from zero [inv R5]? Why it mattersCategory: commercial and compliance. Encoder burst today peaks at 25 nodes / 484 cores with a 46-node cap; no capacity means no burst lane.
Q21What is the NodeBalancer idle timeout, and do WebSocket and gRPC work over the TCP (L4) configuration?
Full questionWhat is the NodeBalancer idle timeout, and do WebSocket and gRPC work over the TCP (L4) configuration? Why it mattersCategory: commercial and compliance. Ingress equivalence for WebSocket/gRPC traffic behind the load balancer.
Q22DataStream 2: what is the per-GB or per-million-lines price, what is the delivery latency (30 s / 60 s push claimed), how sta… blocking
Full questionDataStream 2: what is the per-GB or per-million-lines price, what is the delivery latency (30 s / 60 s push claimed), how stable is the log schema across versions, and can it write to a non-Akamai S3-compatible bucket (our own, or AWS S3 during a dual-CDN window)? DataStream 2 is the sole replacement for the CloudFront-logs -> Glue -> Athena path, which mappings/akamai.md §A classes NO-EQUIVALENT and budgets 2–3 weeks of SQL rewrite against. Why it mattersCategory: blocking the CDN workstream. Sole replacement for the CloudFront-logs -> Glue -> Athena path; 2–3 weeks of SQL rewrite budgeted.
Q23Fast Purge v3: is it included in an AMD contract or priced separately, and what rate limits apply to our purge volume? It is … blocking
Full questionFast Purge v3: is it included in an AMD contract or priced separately, and what rate limits apply to our purge volume? It is the only replacement for the two CloudFront-invalidation providers and the cloudfront-distribution-id SSM values (mappings/akamai.md §A L95). Why it mattersCategory: blocking the CDN workstream. Only replacement for the two CloudFront-invalidation providers in our code.
Q24Give a written onboarding timeline in weeks, from contract signature to first byte served, itemised: Akamai account -> CP cod… blocking
Full questionGive a written onboarding timeline in weeks, from contract signature to first byte served, itemised: Akamai account -> CP code -> property -> EdgeAuth secret -> EdgeWorker activation. State the LKE Enterprise approval clock and the Object Storage E3 limited-availability grant clock separately, since they run on the Cloud Computing side. Why it mattersCategory: blocking the CDN workstream. Three simultaneous ticket-gated clocks plus a CDN contract; Phase 4's entry criterion currently has no number behind it.
Q25What is the minimum term and minimum monthly commit for an AMD contract at ~3.1 PB/month, is a month-to-month or 6-month ramp…
Full questionWhat is the minimum term and minimum monthly commit for an AMD contract at ~3.1 PB/month, is a month-to-month or 6-month ramp available, and what are the exit terms? State the ramp/true-up shape if volume lands below the commit. We cannot sign any multi-year commitment before ~Jan 2027 [inv D9], [inv C10]. Ask the same of the Cloud Computing (Linode) side if LKE-E or E3 entitlement carries any commit. Why it mattersCategory: commercial and compliance. Same constraint that pushed OCI from rank 2 to rank 3; nothing multi-year is signable before ~Jan 2027.
Q26Will Akamai commit contractually to a programmatic daily usage feed grouped by usage type (or an agreed substitute — a schedu…
Full questionWill Akamai commit contractually to a programmatic daily usage feed grouped by usage type (or an agreed substitute — a scheduled export, a usage API, anything timestamped below monthly), given GET /account/invoices/{id}/items is monthly per-service (F) and GET /account/transfer is month-to-date only? Why it mattersCategory: commercial and compliance. Billing observability carries 15 % of the ranking; P15's acceptance criterion is literally this commitment. Mirrors CF-d.
Q27Can an EdgeWorker at the DYNAMIC COMPUTE tier (20 ms CPU, 2.5 MB memory per standard event handler, 512 KB compressed bundle)… blocking
Full questionCan an EdgeWorker at the DYNAMIC COMPUTE tier (20 ms CPU, 2.5 MB memory per standard event handler, 512 KB compressed bundle) run base64-decode + JSON parse + HMAC-SHA256 verify at 4.7 B executions/month, and what does that tier cost? Tier limits and the 512 KB compressed / 1 MB uncompressed bundle limit are published (F), but EdgeWorkers pricing was found on no fetched page. Answer Q16 and Q27 in the same reply. Why it mattersCategory: blocking the CDN workstream. The question Q16 leaves open; the token path is mandatory so an unrunnable tier is a blocker, not a cost line.
Q28Where does the HMAC secret live — EdgeKV namespace, bundle constant, or Property Manager user-defined variable — and what is … blocking
Full questionWhere does the HMAC secret live — EdgeKV namespace, bundle constant, or Property Manager user-defined variable — and what is the rotation story for each? State, per option: whether rotation requires a bundle re-activation, what the propagation time is, and whether two keys can be live simultaneously for a rolling rotation. Published cap: Property Manager user-defined variable totals are 1024 characters, raisable to 4000 on the Dynamic and Enterprise tiers only (F). Why it mattersCategory: blocking the CDN workstream. Gates 03-migration-plan.md §4.3 mechanic 1 ("do not rotate during the window"). Mirrors CF10 against Cloudflare Snippets.
Q29Does AMD's origin pull from a third-party AWS S3 bucket (SigV4, GET/HEAD/OPTIONS only) carry any rate, concurrency or object-… blocking
Full questionDoes AMD's origin pull from a third-party AWS S3 bucket (SigV4, GET/HEAD/OPTIONS only) carry any rate, concurrency or object-size limit we would hit at 8.2 TB/day of origin fetch and 400 M origin GETs/month? State any documented ceiling and any account-level quota. Why it mattersCategory: blocking the CDN workstream. Native SigV4 private-origin pull is §3.7's one genuine Akamai advantage and carries the whole dual-CDN window (~$20,019/month unshielded).
Q30Can CORS be enabled on an E3 Object Storage bucket by exception, or is the E0/E1 restriction absolute? §2.2 #6 and A3 treat i… blocking
Full questionCan CORS be enabled on an E3 Object Storage bucket by exception, or is the E0/E1 restriction absolute? §2.2 #6 and A3 treat it as absolute and route around it with a Cloudflare Worker adding CORS headers, on the strength of the endpoint-types matrix (F) — but a feature matrix showing a blank cell is not the same as a written "no exceptions". Why it mattersCategory: blocking the CDN workstream. If an exception exists, the CORS/performance split (A3, High) collapses and the Worker shim disappears.
Proof-of-concept acceptance tests ·
P1Block-storage IOPS/throughput
Testfio on one Block Storage volume, then on mdadm RAID-0 over two, on a dedicated-CPU instance in the target region. Passes if>= 4–5 k IOPS sustained and >= 600 MiB/s, to match chi-analytics' measured p95 3,000 IOPS (at cap 8 % of minutes) and 566 MiB/s peak [inv R8]. Fail: chi-analytics moves to unpublished local NVMe and a second ClickHouse replica becomes mandatory.
P2Local-NVMe IOPS/throughput
TestSame fio on the plan-local disk of g8-dedicated-128-32. Passes ifSame numbers as P1 (>= 4–5 k IOPS, >= 600 MiB/s), plus: state the behaviour on node recycle. Fail: no home for chi-analytics on Akamai.
P3Object-store node bandwidth
TestParallel GET from one LKE node against a bucket on the target endpoint. Passes if>= 5 Gbps sustained per node. Measured today: chi-analytics bursts 3.01 Gbps rx / 2.86 Gbps tx during cold-tier scan, 80 % of its 3.75 Gbps baseline. E1's documented default egress is 2 Gbps per account per endpoint (F). Fail: cold tier dropped, ClickHouse local-NVMe-only.
P4Object-count and PUT-rate copy rehearsal
TestCopy 1 M objects of ~1 MB and measure sustained PUT/s and wall time. Passes ifMust extrapolate to 28.4 M objects (video-output 18.2 M + image 8.8 M + source + sooper) inside the migration window, and to a steady ~35 GB / 33 k objects per day incremental (worst day 430 GB / 390 k). Total estate: 95.8 M objects / 23,771 GiB. Fail: 16 days of copy at the E3 default 500 PUT/s.
P5S3 feature conformance
TestRun Hoichoi.Encoder/storage_test.go plus a 100 MB multipart upload against a real bucket. Passes ifMust pass: MPU (5 MiB-5 GiB parts, 16 MiB Redpanda), Range GET, ListObjectsV2, bulk DeleteObjects, presigned PUT + CORS with PUT/POST/HEAD, prefix-filtered Expiration.Days, AbortIncompleteMultipartUpload, If-Match, PutObjectTagging. If-Match and bulk delete have no documented answer.
P6SDK checksum break
TestRepeat P5 with a post-2025-01-15 AWS SDK (Go v2, boto3 >= 1.36, CLI >= 2.23) with and without request_checksum_calculation=WHEN_REQUIRED. Passes ifAll 7 S3 writers must work: hc-transcode, Thumbor, PBM, clickhouse-backup, Redpanda, PostHog/chdb, CH cold disk. Fail: per-client remediation, and Akamai's own doc says the workaround "may not work in all cases".
P7Burst-capacity provisioning
TestRequest and instantiate 46 x 32-vCPU instances (1,472 vCPU) from zero in one region, timed. Passes ifObserved peak today is 25 nodes / 484 cores, scaling in minutes, 46-node cap. Per node: 32 vCPU, >= 120 GiB boot disk with >= ~106 GiB usable pod ephemeral (requests 90Gi, limit 108Gi, emptyDir 100Gi). A 40-80 GB boot disk leaves the pod Pending forever; 90-100 GB evicts mid-encode.
P8Pool scale-to-zero
TestSet an autoscaling pool min = 0 through the API and Terraform, on our own account, in the target region. Passes ifMust succeed without a support ticket, and the pool must return from 0 on a pending pod. Fail confirms a permanent floor of ~$788/mo per autoscaling pool x 7 pools = ~$5,516/mo on standard LKE.
P9Node removal is graceful
TestDrain-test the autoscaler: scale a pool down while a pod with terminationGracePeriodSeconds: 120 is running. Passes ifThe pod must receive SIGTERM and >= 120 s before the node dies. Today's contract: EventBridge -> SQS -> Karpenter drain -> DisruptionTarget + SIGTERM, the only pre-emption notice the encoder relies on. Fail: every reclaim costs up to one checkpoint interval; PVs may fail to detach on stateful pools.
P10Online volume expansion
TestGrow a PVC on a running node through the CSI driver. Passes ifNo pod restart, no power-off. 44 prod PVCs / 5,012 GiB would otherwise need a maintenance window each [inv R7]. Fail: offline-only resize becomes a standing operational tax.
P11Static egress IP survives node recycle
TestAttach a Reserved IPv4, recycle the node, re-check the egress source IP. Passes ifThe IP must persist, or a documented re-attach mechanism must exist. 8 partners hold our 3 NAT EIPs on allow-lists [inv R10], [inv C3]. Fail: run a forward-proxy pair outside LKE, which is a new component.
P12KVM pool (conditional on I-a)
TestOnly if the Android team says the Maestro lane resumes [inv H14]: verify /dev/kvm present in a privileged pod. Passes if12 vCPU / 24 Gi request, 200 GiB root @ 6,000 IOPS / 1,000 MB/s, up to 8 runners, scale-to-zero. NOT a blocker: 45 dispatches / 2 successes, dormant since 2026-09-01, ~$11/month — keep on AWS or GitHub larger runners.
P13Managed Postgres conformance
TestOn a trial instance: SELECT extversion FROM pg_extension WHERE extname='vector', SHOW wal_level, SHOW lc_collate, SHOW max_connections, SHOW default_toast_compression. Passes ifpgvector >= 0.8, wal_level=logical, en_US.UTF-8 glibc, max_connections >= 600 (838 configured / 464 peak today), lz4 [inv R16]. Fail: fall back to in-cluster CNPG — which [inv D15] prefers anyway.
P14Ingress equivalence
TestPut ingress-nginx behind a NodeBalancer and drive it. Passes if511 req/s and 22.6 new TLS conn/s at peak hour (1.84 M req/h, 81 k new TLS conn/h); size by RPS and connections/s, not bytes — peak is only 0.06 Gbps (5-min). Confirm the idle timeout and that client_conn_throttle is 0. Fail: non-premium NodeBalancer caps at 10,000 concurrent connections.
P15Billing readback
TestPull a day of spend by usage type through the API. Passes ifMust reproduce what aws ce get-cost-and-usage --granularity DAILY --group-by USAGE_TYPE gives today [inv R20]. Expected to fail — then the acceptance criterion becomes what Akamai will commit to as a contractual substitute, which is RFP question Q26.
P16Restore into an Akamai target (NOT a PoC test — a Phase-6 acceptance item, relabelled)
TestAfter the provider-independent drill has passed on AWS, restore Mongo from PBM and ClickHouse from clickhouse-backup into Akamai Managed PostgreSQL / LKE / Object Storage and assert row counts. Also validate Debezium resume-token validity after a restore onto Akamai storage. Passes ifMust reproduce the row counts the AWS drill produced, against Object Storage as backup source and node-local NVMe as the ClickHouse target. Gate G6.2 sits in Phase 6, ~10 months after selection, so it cannot inform the decision; evaluation-time substitute is P5 + P6 on a trial bucket.

How the move would run

7 phases · tasks · gates

Phase 0 is worth doing whatever we decide. It is the first restore drill this estate has ever run.

P015 tasks

Pre-migration fixes, done on AWS

De-risk any move with 15 workstreams that are all independently worth doing on AWS. 10 of 15 are already-listed standing defects; if the estate stays on AWS, Phase 0 still happens. Weeks 1-8 (2026-09-21 to 2026-11-15).

  • P0-01Export the live-only state that exists in no file: (a) ClickHouse SHOW CREATE USER/ROLE/SETTINGS PROFILE + SHOW GRANTS (12 users, 1 role, 1 profile, 3…
  • P0-02PSP webhook DNS indirection: put a DNS-addressable hostname in front of the SSLCommerz IPN receiver, re-register at the 7 PSP consoles plus the Mongo …
  • P0-03ESO re-point rehearsal on UAT: stand up a second ClusterSecretStore against a non-SSM backend, flip one non-critical ExternalSecret to it and back, me…
  • P0-04Restore drills - Mongo and ClickHouse, the first ever, into a scratch namespace on UAT. PBM restore + PITR restore; clickhouse-backup restore_remote w…
  • + 11 more
Gate G0.1 Live-only state is in git: the RBAC .sql replays into a scratch ClickHouse and SHOW GRANTS returns 323 grants; all 7 snippet bodies committed; 8 lifecycle JSONs committed.
Gate G0.2 Restore drills passed twice: two PBM restores (one physical, one PITR) and two restore_remote --rbac runs into a scratch namespace, each followed by a successful Debezium resume with no…
P18 tasks

IaC bootstrap on the target

Author fresh with OpenTofu because there is nothing to export: L0 state last written 2025-03-13, L1 tofu never merged, prod EKS not under IaC at all, UAT tofu five months drifted. Everything above Kubernetes stays in helmfile + kustomize. Weeks 4-10. On rank-1 this phase is: import Cloudflare into Terraform, and nothing else.

  • P1-M1State backend: the state bucket, versioning, locking, a customer key and object-level audit from day 1. OpenTofu's preferred lock is native S3 conditi…
  • P1-M2Identity and secrets: the secret store (provider-native, or self-hosted OpenBao where none exists), the ESO ClusterSecretStore, and the 13 workload cr…
  • P1-M3Network: VPC/VCN, subnets, pod and service CIDRs. Give the two clusters distinct service CIDRs (e.g. 172.21/16 and 172.22/16) and keep every CIDR out …
  • P1-M4Cluster + node pools: every pool with labels, taints, autoscaler bounds and boot-disk size per pool (apps/analytics-db/clevertap-sandbox 60 Gi live vs…
  • + 4 more
Gate G1.1 State backend safe: tofu init plus a trivial apply/destroy cycle succeeds; locking demonstrated by two concurrent applies where exactly one wins; versioning listed; audit log shows the …
Gate G1.2 No secret in state: grep of the plan JSON and the state object for any value that also appears in the secret store returns nothing.
P211 tasks

Platform bring-up

Seed the registry, then apply 27 helmfile releases in dependency order across 8 waves, re-keying the zone affinities and replacing Karpenter with whatever the candidate's autoscaler is. Weeks 8-16, overlapping Phase 1's tail.

  • P2-01Registry seeding, first. crane copy retains the digest; leave --platform unset (its default is all) so P0-08's multi-arch manifests survive the mirror…
  • P2-W1Wave 1 - CRDs and controllers: karpenter-crd replaced by the target's autoscaler CRDs; cert-manager, external-secrets and metrics-server are all zero-…
  • P2-W2Wave 2 - Autoscaling and scheduling: Karpenter to the target's autoscaler (OCI = oracle/karpenter-provider-oci EQUIVALENT, CRs port, needs the oke-is-…
  • P2-W3Wave 3 - Ingress and DNS: aws-load-balancer-controller to ingress-nginx or the target's controller (5 prod Ingress files, 166 annotation lines, 17 Tar…
  • + 7 more
Gate G2.1 Registry complete: every image referenced by a rendered kubectl kustomize overlays/<env> resolves on the target registry; 0 references to the old registry in a repo-wide grep; every mul…
Gate G2.2 All 27 releases applied clean: helmfile -e <env> diff is empty against the target cluster.
P314 tasks

Data-plane rehearsal on a UAT-equivalent

Move ~17.9 TB across the AWS edge per full pass (media library 14,753 GB = 82%) and prove every store restores, reconciles and resumes. Budget >= 2 full passes. Weeks 12-20. On rank-1 this is not a full phase: P0-04 IS Phase 3, plus the Typesense and Valkey rebuild checks.

  • P3-01MongoDB: PBM restore into a fresh PSMDB (or mongosync live), then re-point. ~0.27-0.5 TB logical - restore the latest full + PITR, not the 2.65 TB buc…
  • P3-02ClickHouse hot: clickhouse-backup restore_remote from the existing chain; re-seed, do not copy 3.18 TB of history. ~1.25 TiB (chi-analytics 612 GiB ac…
  • P3-03ClickHouse cold: key-list copy from system.remote_data_paths preserving the key layout, two-pass with STOP MERGES. 967.7 GiB across 145,840 referenced…
  • P3-04The ClickHouse cold-tier two-pass, written out: (1) copy the key list from system.remote_data_paths while merges run; (2) SYSTEM STOP MERGES - note th…
  • + 10 more
Gate G3.1 Mongo restored and usable: a PBM restore into the fresh PSMDB completes; rs.status() healthy with clusterServiceDNSMode: Internal; a document count per collection matches source within …
Gate G3.2 Debezium resumes without a re-snapshot: all 9 connectors resume from their copied offsets and emit change events. Second separately checkable line: the capture-start property is confirm…
P49 tasks

Video CDN dual-serve

Move the 97.2% of in-scope spend that is CloudFront egress in one account, by porting the token verifier to the new edge, canarying on jio.partner, ramping the apex by DNS weight, and moving the origin. Weeks 6-24, on the assumption the CDN contract is signed by end of week 5. On rank-1 this phase IS the whole project.

  • P4-4.1Token parity build: port the 7.9 KB cloudfront-function-token-verify.js to the new edge runtime, keeping stream_key/stream_policy parsing byte-identic…
  • P4-4.2Origin attach, no client traffic: attach the new CDN to the existing S3 origin. Akamai Property Manager has a native SigV4 private-origin behaviour; C…
  • P4-4.3Dual-serve canary - the measurement that cannot be skipped. Point jio.partner.hoichoicdn.com (already a separate DNS name, 0.60% of bytes, hits the or…
  • P4-4.4Warm the hot set: pre-warm 0.9 TB = 80% of delivered bytes; 1.6 TB = 90%; 2.3 TB = 95%. Pre-warming the 2.3 TB hot set is what keeps the transition-wi…
  • + 5 more
Gate 4.1 Token parity: the canary validates 100% of the replayed corpus on BOTH the accept and the reject path, including expired, malformed and the two partner variants.
Gate 4.2 Origin attach: a hand-crafted request for a known segment returns 200, correct bytes, correct Content-Type, and a MISS then a HIT.
P511 tasks

Prod cutover

One scheduled window, announced at 6 h with an internal target of ~4h45m and an abort at T+5h00 wall clock from T-0, rehearsed twice in Phase 3. Weeks 21-22 (2027-02-08 to 2027-02-21). Does not exist on rank-1.

  • P5-5.1T-7d: freeze the config surface. No SSM/secret edits, no helmfile applies, no schema changes on either side. Value-scan every config store before cutt…
  • P5-5.2T-24h: final incremental sync. Media delta, PostHog objstore delta, ClickHouse cold delta (two-pass again), Postgres logical replication caught up.
  • P5-5.3T-0: write freeze. Stop producers at the edge, not at the database. Order: scale app writers to 0, let Debezium drain to lag 0, let Redpanda consumers…
  • P5-5.4Copy the Debezium offsets and rewrite capture start: copy the 9 offset files (~234 KiB total). For any connector whose resume token did not survive, u…
  • + 7 more
Gate G5.1 All 9 Debezium connectors streaming, lag 0, no re-snapshot triggered.
Gate G5.2 Event rate into ClickHouse within tolerance of the pre-freeze rate for a full diurnal cycle.
P610 tasks

Decommission and credit-expiry-aware egress schedule

Delete nothing for 4 weeks, then retire the old estate on a +0 to +8 week schedule from the Phase 5 window (or from Phase 4 completion on rank-1). Decommission spends no egress, so the credit expiry of 2027-02-28 is a restatement of the week-20 pin on the bulk copies, not a separate deadline this phase can miss.

  • P6-01Weeks +0 to +4: nothing is deleted. Old cluster stays warm, old CDN distributions stay disabled-not-deleted, old NAT EIPs stay reserved, S3 remains re…
  • P6-02Week +4: disable (do not delete) the CloudFront distributions; remove the OAC statements from bucket policies; cancel CloudFront log delivery and let …
  • P6-03Week +4: retire the cross-account Athena IRSA and the old cdn-usage-ingester path. Gated on the new log pipeline having matched the old one for >= 7 d…
  • P6-04Week +6: revoke the 11 prod + 11 UAT third-party and Org role trusts - Vantage (assumed daily), Workmates/MSP x5 including the admin SAML role, Middle…
  • + 6 more
Gate G6.1 30 days of steady state on the new CDN with aggregate byte-hit >= 91.9% and 5xx flat. Per-distribution, the comparison is against that distribution's own CloudFront baseline - jio's is …
Gate G6.2 30 days of steady state on the target platform; two successful restore drills run ON the target, not on AWS.

What we actually run

measured, not estimated

MetricProductionUATCDN account
compute
Nodes (instantaneous)6 static + 8–11 Karpenter (typ. 15–17), all arm6418 (16 OD + 2 spot), all arm64
vCPU / RAM presentsnapshot 64 vCPU / ~362 GiB; 30-d mean 16.4 nodes ≈ 73 vCPU / ~410 GiB steady poolsallocatable 37.4 cores / 128.6 GiB
30-d usage p95 (steady pools)25.2 cores / 185 GiB (peak30m 52 / 227)≈ 6.6 cores busy of 42 vCPU; memory ~75 % committed
Spot share29 % of EC2 $ (10 % of instance-hours); apps pool 100 % spot, 753 node replacements / 30 d2 spot nodes only
IAM roles (RoleLastUsed, measured)108: 9 app/data roles to rebuild (all 9 used in-window), 11 EKS plumbing, 10 other AWS-native, 11 external/Org trusts, 32 ECS/MSP dead, 35 service-linked95: 4 app/data live (mongodb-backup-role never used), 6 EKS plumbing, 25 other, 11 external, 26 dead, 23 SLNOT MEASURED (iam:ListRoles denied)
storage
PVCs / EBS44 PVCs / 5,012 GiB gp3; EBS total 66–69 vols / 6,185–6,505 GiB (superseding 64 vols / 6,065 GiB; two captures 2 h apart on 2026-09-18 differed by +3 vols / +320 GiB); 0 true orphan volumes; only stateful non-PVC disks are 5 stopped legacy EC2 roots (413 GiB, all RETIRE except video-processing, restarted 2026-09-18). Node roots: Karpenter apps/analytics-db/clevertap-sandbox 60 Gi (git says 30 Gi — drift), batch-jobs 200 Gi, media-jobs 120 Gi, ci-android-kvm 200 Gi, cloudflared-egress 20 Gi, managed NGs 30 Gi41 Bound + 1 Pending / 1,110 GiB (EBS 77 vols / 2,390 GiB, 24 orphan = 552 GiB)
S345 buckets, 23,771 GiB (= 25,406.9 GB) / 95.8 M objects; 11 versioned, 11 with lifecycle (23 rules, prefix-filtered only), 7 CORS, 3 public-policy (1 intended), 0 replication / object-lock / KMS32 buckets (superseding 32–33), 4.09 TB (3,727 GiB Std + 135 IA); 5 versioned, 5 lifecycle, 6 CORS, 1 public + static-website7 buckets, ≈ 18.7 TB CloudFront logs (CE-inferred)
Cluster-attached S3 cost (ch-cold, redpanda-tiered, 2 backup buckets)≈ $350 window / ≈ $307 forward (superseding the ≈ $200 inference) = storage $247 (10.1 TB window mean, 8.17 TB steady) + requests $103 measured (38.5 M GET + 17.3 M PUT) + ≤ $27 inferred backup ops
edge
CloudFront origin fetch 30-d241,371 GiB origin→CloudFront (superseding 238,118 GiB) = S3 240,618 GiB (video) + ALB→CloudFront 749 GiB (content-API + Thumbor); $0 only while CloudFront is the puller (≈ $20 K/mo at DTO list price if a non-AWS CDN pulls the same way without a shield)0.75 GB2,928,681 GiB = 2.79 PiB = 3.145 PB to viewers (4.93 B req)
Cloudflare edge (all zones, 31 d)≈ 2.80 B requests / 22.9 TB: hoichoi.dev 2.39 B / 13.1 TB (⚠ CONFLICT: [06 §4] 12.2 TB vs [03 R2-01.2] 13.15 TB; canonical 13.1 TB) at 1.2 % req cache-hit; hoichoicdn.com 268 M / 9.47 TB in 18 d (97.2 % hit); hoichoi.tv 140 M / 1.16 TB. Uncached from AWS origins ≈ 10.35 TB / 31 d (79 % = content-API JSON)
Cloudflare Workers55.55 M invocations (hoichoi-web 54.7 M)
Tunnel bytes (31 d)cloudflared-otlp 51.5 TB in / 976 M req (52.7 KB/req; Aug ≈ 2.2 TB/day, Sep ≈ 1.2 TB/day); cloudflared 1.6 TB (≈ 97 % PostHog ingest); cloudflared-private-net 35 GB
network
Internet egress (non-CDN), measured split2,517 GiB · $229.72 (superseding 2,487 GB · $227.14) = EC2 1,606 GB (tunnel responses ≈ 1,331 + NAT 275) + ELB 852 GB + ECR/S3 59 GB; Sep run-rate ≈ 1.8–2.0 TB/mo (tunnel half halved from 09-02)74 GB · $7.12
Internet inbound (free)56.0 TB: EC2 52.7 TB (OTLP tunnel ingest), ELB 2.25 TB, S3 1.08 TB
Cross-AZ (Regional-Bytes)28,771 GB · $287.706,626 GB · $66.26
NAT844 GB processed · 3 GW · $170.56 (275 GB out / 640 GB in; tunnel nodes bypass NAT via public IPs)116.7 GB · 1 GW · $47.64
ALB414.9 M req / 3.68 TB; peak-hour 1.84 M req / 18.6 GB = 0.041 Gbps, peak 5-min 0.060 Gbps, 81 k new TLS conn/h31.2 M req / 58.9 GB
Cluster east-west (pod traffic)≈ 470 TB / 31 d, dominated by monitoring↔egress telemetry (110 TB), data replication (88 TB), apps (65 TB); media pool only 2.8 TBNOT MEASURED
data
Redpanda Kafka bytes159 GB/day produced, 285 GB/day consumed (98 % = 3 PostHog Cloud Topics; clickhouse_events_json re-read 4×); wire ≈ 209 GB/day in / 420 out; S3 side ≈ 587 GB/day PUT / 220 GET
RDS1 × db.t4g.large PG 15.17, 400 GB gp3 (12 k IOPS), ~341 GB used, single-AZ, 1-day backups, $196.971 × db.m6g.large, 200 GB gp2, 52 GB used, publicly accessible, $194.13