AM
Ankit Mehta
Founder · VigilesHQ · CloudArch Pro
Ankit Mehta founded CloudArch Pro and Vigiles. His work covers cloud architecture, FinOps, database platforms and incident operations for production systems in APAC. He writes the technical library on this site and reviews consulting deliverables.
Founder and technical author for CloudArch Pro. Founder of Vigiles (vigileshq.com).
He reviews CloudArch Pro consulting deliverables. Claims on this site are either cited to official provider documentation, labelled as engineering recommendation, or marked as needing first-party confirmation.
Technical topics
- Production landing zones and account structure
- Hub-and-spoke networking and egress cost
- FinOps allocation and cost-anomaly response
- AI and GPU workload unit cost
- Multi-cloud identity and APAC region design
Profiles
Published on CloudArch Pro
- Guide
Alibaba Cloud architecture for APAC workloads
Resource Directory, RAM, CEN and VPC design for APAC on Alibaba Cloud, including Singapore and China mainland regions.
- Guide
Cross-region disaster recovery
How to pick RTO, RPO and failover mechanics across AWS, Azure and Google Cloud regions without treating a second AZ as DR.
- Guide
Egress cost planning
How to estimate inter-AZ, inter-region and internet egress before the invoice, using native cost data and published meters.
- Guide
GPU and AI workload cost governance
Unit cost for training and inference, idle GPU time, and how reservations differ from general compute FinOps.
- Guide
Hub-and-spoke cloud networking
Transit Gateway, Azure hub-spoke, Shared VPC and when inspection and egress belong in the hub.
- Guide
Kubernetes cost allocation
How to attribute cluster spend to namespaces and workloads on EKS, AKS and GKE using native billing labels and Kubecost.
- Guide
Cloud migration readiness assessment
What to inventory, decide and refuse before you move production systems to AWS, Azure or Google Cloud.
- Guide
Multi-account and multi-subscription structure
How to split AWS accounts, Azure subscriptions and Google Cloud projects so blast radius, billing and policy stay separable.
- Guide
Designing multi-cloud identity
Federation, workload identity and trust boundaries when AWS, Azure and Google Cloud all run production.
- Guide
Designing a production landing zone
Account structure, identity, networking and logging a production landing zone needs on AWS, Azure and Google Cloud.
- Guide
Cloud tagging and cost allocation
A tag dictionary, inheritance and unallocated spend on AWS, Azure and Google Cloud billing exports.
- Guide
Finding unexpected cloud cost increases
Isolate a spend spike by service, account and usage type before you change architecture or delete resources.
- Guide
When serverless fits a production platform
Use functions or request-driven containers when the unit of work is an invocation with documented duration limits. Steady workers usually stay on nodes.
- Runbook
Autoscaling Failure Runbook for AWS, Azure, Google Cloud and Alibaba Cloud
Separate a policy that never fired from quota, regional capacity, launch failure, failed health checks and scale-in protection before you raise max size.
- Runbook
Cross-region traffic cost increase
Separate inter-AZ, inter-region and internet egress on the bill, then find the copy or failover path that grew.
- Runbook
DNS routing failure
Diagnose hosted-zone, resolver and health-check failures on Route 53, Azure DNS, Cloud DNS and Alibaba Cloud DNS.
- Runbook
Expired certificate
Replace an expired TLS certificate on ACM, Azure Key Vault / Application Gateway, Certificate Manager and Alibaba SSL Certificates without dropping the old listener first.
- Runbook
IAM access lockout
Regain administrative access without deleting the last admin, and use documented break-glass paths on each provider.
- Runbook
Kubernetes cluster cost increase
Separate control-plane fees from node, volume, load-balancer and egress cost when a cluster bill jumps.
- Runbook
NAT gateway cost increase
Find why NAT or Cloud NAT spend rose by separating hourly gateway charges from bytes processed.
- Runbook
Cloud quota exhaustion
Identify the quota that blocked a create or scale call, then request an increase or change the design without guessing the limit.
- Runbook
Region failover
Execute documented DNS or anycast failover, and separate that from ad-hoc region moves that are not in the DR plan.
- Runbook
Unexpected cloud cost spike
Isolate a cloud spend spike by account, service and usage type before you delete or resize anything.
- Blog
Why APAC region defaults differ from US-centric diagrams
US-East-first reference architectures pick the wrong primary region, underestimate intra-APAC latency and ignore Alibaba Cloud. Opinion labelled as such.
- Blog
How CloudArch Pro reviews cloud architecture
The evidence classes, artefacts and limits we use in a landing-zone or FinOps review. Method, not a case study.
- Learn
Cloud networking fundamentals
AWS VPCs are regional, Azure VNets are regional, Google Cloud VPCs are global with regional subnets, and Alibaba Cloud VPCs are regional with zonal vSwitches.
- Learn
Identity and access management fundamentals
AWS IAM, Azure RBAC, Google Cloud IAM, and Alibaba Cloud RAM evaluate identity and authorization differently. Federation does not merge those models.
- Learn
Regions, availability zones and failure domains
Regions, zones, and provider-specific isolation boundaries define which failures a multi-AZ design can survive and which ones require another region.
- Learn
Cloud cost allocation fundamentals
Cost allocation maps provider usage to an owner using accounts or subscriptions first, then activated tags or labels. Untagged spend is part of the report.
- Learn
Shared responsibility model
Each provider documents a split between security of the platform and security of your configuration, data, and identities. The line moves by service model.
- Learn
What is a landing zone?
A landing zone is the baseline multi-account or multi-subscription environment that applies identity, network, logging, and policy before the first production workload is deployed.
- Learn
What is AI FinOps?
AI FinOps is FinOps applied to AI spend: GPU capacity, idle time, managed-model tokens, and unit costs that generic cloud allocation does not show.
- Learn
What is cloud architecture?
Cloud architecture is the set of tenancy, identity, network, failure-domain, and cost-ownership decisions that determine how workloads run on AWS, Azure, Google Cloud, or Alibaba Cloud.
- Learn
What is FinOps?
FinOps is the FinOps Foundation's operating practice for technology value: timely cost data, shared accountability, and decisions made by engineering, finance, and the business together.
- Learn
What is multi-cloud?
Multi-cloud is running production workloads on more than one provider, each with its own identity, network, billing, and failure-domain model.