Part 4: SRE Practices — Chaos Engineering with Chaos Mesh & Multi-Region Resilience
Multi-Language Edition: This chapter is also available in Vietnamese at 📖 Bản tiếng Việt (Vietnamese Edition). Previous Chapter: Part 3 — Data Infrastructure: From Aurora to TiDB | Series Hub | Next Chapter: Part 5 — Campaign Architecture: Surviving the 10-Billion Yen Surge Answer-First: Delivering five-nines (99.999%) availability for national payment infrastructure requires shifting from reactive disaster recovery to continuous, automated Chaos Engineering in production. PayPay integrates Chaos Mesh into Kubernetes EKS clusters, deliberately injecting pod evictions, network latency, and cross-AZ partitions during normal business hours to validate self-healing invariants. To prevent cascading failures under heavy load, PayPay enforces distributed circuit breaking with Sentinel, client-side exponential backoff with full jitter, and strict gRPC deadline propagation, ensuring localized microservice brownouts never degrade core payment authorization. ...