Sometimes we need to break things first to secure them. This is the idea behind chaos engineering.

Netflix's Chaos Monkey project in 2011 and Google’s disaster recovery testing (DiRT) exercises are among the first practices of chaos engineering. And other major cloud providers including Amazon Web Services (AWS) and Microsoft Azure have deployed the technology by injecting destructive faults into their system tests to secure the cloud.

Gartner defines chaos engineering as an experimental and potentially destructive failure testing approach to identify weaknesses and vulnerabilities in a complex system. It is designed to “evaluate the reliability or the resilience of a system,” Gartner analyst Jim Scheibmeir explained. These services “intentionally degrade the system to see how resilient it behaves.”

Gartner categorizes chaos engineering as an emerging practice in the “adolescent” maturity stage. It found 18% of surveyed organizations are currently using or planning to use the technology to improve software quality.

Chaos engineering is also on the rise in this year’s Gartner Hype Cycle for I&O Automation with a market penetration rate is 5% to 20% of targeted audiences — an increased from less than 5% last year.

Gartner's client inquiries on the technology also have increased significantly in recent years. “Three to four years ago, I took very few calls on the practice of chaos engineering. Now, I'll take a few of these engagements in a week, and I even get to start hearing from clients that they're one or two years into their journey,” Scheibmeir said. “So now we're actually building up competencies and maturity around this emerging practice.” 

Despite the upward rising trend, Gartner placed chaos engineering in the first stage of its Hype Cycle and expects it will take up to 10 years to achieve mainstream adoption. 

Benefits vs. Obstacles

Many organizations’ current test plans overemphasize software functionality but underemphasize validating the system’s reliability, Scheibmeir said. Deploying chaos engineering is “not just to see how reliable or resilient our system is, but also how secure it is when it's running in a degraded state,” he added. 

Another benefit is that the proactive nature of chaos engineering helps solve DevOps objectives such as improving agility and system reliability. Gartner anticipates 40% of organizations will deploy chaos engineering practices as part of their DevOps initiatives by 2023, resulting in a 20% reduction in unplanned downtime.

However, there are challenges to adopting the technology, Scheibmeir warns. “The first thing is that many organizations still look at it as something that is going to be unplanned, random, and done in a production system, and that certainly sounds very disruptive and dangerous.”

Because of that, Gartner recommends organizations start with implementing chaos engineering as “systematically planned, documented, executed and analyzed test-first approaches in pre-production infrastructure systems.”

As with most new technologies, time and budget remain top concerns. “There are certainly organizations that see quality and testing as additional cost or overhead, and I think those organizations are failing to realize that quality and testing convert waste into value,” Scheibmeir said.

AWS Fault Injection Simulator

AWS is an early mover in the space, and it rolled out its Fault Injection Simulator at last year’s re:Invent. Based on the chaos engineering concept, the cloud service enables users to run fault injection experiments on AWS to improve applications’ performance, resiliency, and observability.

“You will learn how your system reacts to various types of faults, and you will have a better understanding of failure modes,” AWS chief evangelist Jeff Barr wrote in a blog post. “You can start by running experiments in pre-production environments, and then step up to running them as part of your CI/CD workflow and ultimately in your production environment.”

AWS released the service in March for most of the commercial AWS regions, and it bases pricing on users’ action run time. 

Microsoft Azure Chaos Studio

Besides AWS, Gartner also lists Alibaba Cloud, ChaosIQ, Gremlin, steadybit, and Verica as sample chaos engineering vendors. “We're gonna see this in all of the cloud providers that will have a native service that runs within their cloud offerings,” Scheibmeir said. 

Microsoft Azure is the latest cloud platform to announce the chaos engineering offerings: Chaos Studio. 

The service helps users to orchestrate fault injections on their Azure resources in a safe and controlled way, wrote John Engel-Kemnetz, program manager for Microsoft Azure Chaos Studio, in a blog post. 

Chaos Studio, available in public preview covering several Azure cloud services, has more than 25 faults in its library. Its orchestration capabilities build up complex experiments that replicate real-world incidents, Kemnetz added.

The public preview is free to use through April 4, 2022, and thereafter usage will be charged pay-as-you-go based on experiment execution.

PingCAP Chaos Mesh

In addition to commercial offerings, there are a number of chaos engineering open-source platforms, including Mangle Open Source, which is sponsored by VMware, Byteman Open Source, ChaoSlingr Open Source, and Chaos Mesh.

PingCAP, an open-source, distributed database developer, created Chaos Mesh to orchestrate chaos experiments on Kubernetes environments. 

The company started to use the concept of chaos engineering from day one for its TiDB SQL database, said Fadi Azhari, senior director of marketing at Pingcap. “Chaos mesh for us is an interpretation of chaos engineering,” which helps users pinpoint issues and possible causes of latency in the data flow, he added.

PingCAP is also transitioning Chaos Mesh to an as-a-service model, which Azhari said is the technology's next phase.

“Chaos-as-a-service is just like any other IT service or any other software-as-a-service,” Azhari said. “It could be utilized at scale with other companies that don't have the wherewithal to understand every single piece of how they implement it in their organization.”

To support chaos-as-a-service, PingCAP suggests chaos engineering platforms include a unified console for chaos experiments management, visualized metrics to show experiments’ status, operations to pause or archive experiments, and simple interaction to enable engineers to orchestrate their experiments.

Azhari also agrees with Gartner's assessment that the technology is still in its early stages. “It's a couple of years away from it being really solid in the market because you're not dealing with just one or two large companies adopting it, but you're gonna make it easy by looking at the entire ecosystem that moves in that direction.”