DiDi Global | Enterprise Cloud-Native Chaos Engineering System
Built and upgraded DiDi’s enterprise-scale chaos engineering platform to evaluate microservice resilience and automated fault recovery under high-concurrency production workloads.
- Cloud-Native Event-Driven Architecture: Re-architected the execution pipeline using Kubernetes CRDs, Go Operators, and gRPC Daemons for scalable, asynchronous fault injection.
- Kernel-Level Resilience & Fault Control: Managed application, middleware, and Linux kernel-level faults (e.g., tc/netem network delay, packet loss, and cgroup stress).
- Automated State Consistency: Built transactional locking and side-channel verification mechanisms to guarantee zero orphaned fault rules during system crashes.