We are in the throes of Amazon Prime Day, one of the most anticipated shopping “holidays” of the year.

And while billions of members around the globe are on the hunt for exclusive bargains — admit it, you’re eyeing a few things yourself — many don’t consider all that goes on behind the scenes of such massive shopping frenzies (and what could go wrong).

“Unlike the holiday shopping season which spans multiple weeks, highly anticipated events such as the two-day only Prime Day creates a massive and sudden surge in online traffic,” said Kayla Broussard, U.S. CTO of consumer and travel and distinguished engineer at IT services provider Kyndryl.

In addition to the websites themselves, inventory management systems, customer relationship management systems, financial software and more are all running at peak capacity in the run up to and during such events.

“If any of those systems were to experience slowdowns or even worse, crashes, there would be lost sales along with damaged brand reputation and diminished customer loyalty,” Broussard said. While your requirements may not rival Amazon's, if your infrastructure needs to be agile enough to handle large spikes in traffic, there are lessons you can learn from what Amazon does right and what it's done wrong.

Amazon Prime by the numbers

An estimated 65.9% of the U.S. population that hold Prime memberships and many millions of others around the world are expected to shop Prime Day, which kicked off at midnight PST yesterday (July 11) and runs through tonight at 11:59 p.m. PST.

The first Prime Day took place on July 15, 2015, and the event has only ballooned from there. The 2022 two-day event was Amazon’s biggest to date — but time will tell whether that will soon be surpassed — with Prime members around the world purchasing more than 300 million items (and saving an estimated $1.7 billion in the process, according to the e-commerce giant).

The 2022 event was also the biggest for Amazon’s selling partners, most of whom are small and medium-sized businesses: Customers spent more than $300 billion on 100-plus million of their goods.

Breaking that down further:

  • Prime members worldwide purchased more than 100,000 items per minute and did the majority of shopping between 9 and 10 a.m. PDT on July 12, the first day of the event.
  • Members in the U.S. purchased more than 60,000 items per minute and were most active during 8 and 9 p.m. PST on Wednesday July 13 (close to the culmination of the event at midnight PDT).
Not always prime performance for Amazon

Still, the massive annual shopping splurge hasn’t come without its challenges. Most notoriously, during its 2018 event, Amazon struggled to handle traffic surges because it failed to secure enough servers, according to a CNBC report. That led to a slowdown in Sable, its internal computation and storage service. The eretailer launched a scaled-down “fallback” front page, temporarily halted all international traffic and manually added servers.

To prevent outages in 2022, Amazon beefed up its capabilities. According to a post-event blog post, a “multitude” of two-pizza (that is, more than could be fed by that amount, as defined by Jeff Bezos) worked together to ensure that infrastructure was scaled, tested and ready.

The company increased the total number of normalized instances (an internal measure of compute power) on Amazon Elastic Compute Cloud by 12% and added 152 petabytes of EBS storage, according to the company.

The resulting fleet handled 11.4 trillion requests per day and transferred 532 petabytes of data per day. More than 5,300 database instances processed 288 billion transactions, stored 1,849 terabytes of data and transferred 749 terabytes of data. Furthermore, Amazon Simple Queue Service set a new traffic record by processing 70.5 million messages per second at peak, the company reported.

Lesson 1. Scaling critical infrastructure:

All told, the global ecommerce market is expected to hit nearly $6 trillion this year and grow to $7 trillion in 2025, double its valuation in 2019.

To tap into that and benefit from massive events such as Prime Day, ecommerce sites should be equipped with “smart, scalable infrastructure that carefully handles all traffic 24/7,” said Broussard.

There are two forms of scaling, she explained: vertical and horizontal. The former entails adding more resources such as CPU, RAM and storage to an existing instance, which can often make the instance temporarily unavailable. Horizontal scaling, meanwhile, involves adding more instances to the environment and is more easily automated.

These two scaling approaches are not mutually exclusive, she pointed out. They can be used in combination when scaling up or out in anticipation of, and during, the high volume demands of Prime Day.

Lack of scalability can result in performance issues such as slow page loads — which is one of the primary causes of abandoned shopping carts — or even complete outages, which can bring “digital sales to a full stop,” she explained.

Lynn Comp, corporate VP for the server business unit at AMD, agreed that scalability is all about providing the ability to rapidly respond to service level peaks. It is critical in ecommerce because it enables seamless transitions, whether a prospective buyer is researching product info or securely completing a shopping cart transaction.

“The industry in general is seeing declining customer loyalty, and attrition is common for any customer having to wait for session availability to complete a transaction,” said Comp.

Lesson 2. Prepare for spikes with cloud elasticity

Just as importantly, cloud elasticity allows platforms or architectures to rapidly expand to accommodate sudden, unexpected spikes in demand, said Jeff Darnton, strategic engagement manager at Akamai.

Elasticity is automated, can occur quickly and can help online retailers avoid bottlenecks so they can expand rapidly without the risk of overinvesting in cloud resources that are only occasionally used, he pointed out.

Still, no increase in capacity will be spontaneous, so resource planning and management are necessary to ensure availability during spikes, he said. In a peak event, the impacts of small failures can quickly be exacerbated by load, and the impacts are “more acutely” felt by customers.

“Even if it’s not catastrophic, slow load times or buggy pages can easily lead to a lower conversion rate,” he said. “Any failure at mass scale puts you in the crosshairs of public animosity: It can feel like you’re in a fishbowl and the whole world’s watching to see if something will go wrong.”

Lesson 3. Load testing, stress testing, endurance testing … and more testing

Experts also emphasize the importance of performing multiple forms of testing and monitoring.

For example, Broussard said, there’s load testing, in which load is gradually increased to determine at what point the system begins to slow down. Stress testing is just as essential because it blasts the site with more and more requests at an ever-increasing rate to find the “unavoidable breaking point.”

Endurance testing can help to determine how long a site’s performance remains up to standard to sustain traffic surges.

Meanwhile, synthetic monitoring (or synthetic testing) can be used to execute performance checks around the clock so that retailers are not reliant just on consumer interaction during peaks. This type of monitoring emulates the paths users may take when engaging with a retailer’s website and can measure the response time, track site uptime, generate reports and send alerts based on the performance of the site during testing.

Other core technologies or capabilities to reduce downtime and outages include load balancers, content delivery networks (CDN), war rooms and bot mitigation, said Broussard.

Lesson 4. Human talent trumps AI and ML in an emergency

Even with today’s advancements in artificial intelligence (AI) and machine learning (ML), though, human intervention is still necessary, Broussard pointed out. Retailers must always have talent and war rooms available to respond to emergencies.

Darnton agreed that while enterprises should conduct “a host of activities” to prepare for expected traffic surges, plans can still go awry and there are many threat actors out there — bots, fraudsters — that are “intent on manipulating your big event for their gain.” Therefore, he emphasized the importance of culture and mindset, as “no system is fail proof.”

“Leaders should instill their organizations to expect failure, plan for it and compartmentalize any failures to prevent small things from causing a big impact,” he said.

In the end, retailers should determine the right mix of tools, techniques and practices that work best for them and properly prepare themselves for events such as Prime Day, Broussard advised.

In doing so, “retailers can better handle the increased demands on their data center and cloud infrastructure and benefit from some of the best revenue-generating days of the year,” she said, “all the while providing an enjoyable shopping experience for customers.”