Designing Highly Available Infrastructure at Scale
Led the design of EKS and ECS runtime platforms serving several million monthly users and handling thousands of requests per second. Used HPA, Cluster Autoscaler, load testing, and database connection budgets to define safe scaling limits and pre-scaling plans, maintaining availability during traffic spikes of roughly 1,000 times normal levels.
Led stakeholder alignment, budget approval, and the zero-downtime migration of multiple production services from EKS to ECS as the business phase and operating model evolved.
Reduced infrastructure operations and management labour costs by roughly 80%, and combined AWS, GCP, and Datadog cloud and monitoring spend by roughly 50% through continuous usage analysis and resource right-sizing.