Close up of a Microsoft Azure ND GB300 v6 rack
– Microsoft

Microsoft claims to have smashed an AI inference record for virtual machines (VMs), with its new Nvidia GB300-powered platform pushing inference well beyond the one million token mark.

Utilizing the MLPerf Inference benchmark, the cloud giant claimed one rack of its new Azure ND GB300 v6 VMs achieved an aggregated performance of 1.1 million tokens per second. By comparison, the prior generation ND GB200 v6 VMs held the previous MLPerf Inference record with 865,000 tokens per second.

To put that in perspective, one GB300 v6 rack powers 15,200 tokens per second, per GPU.

The results were validated by the Futurum Group-owned benchmarking firm Signal65. Russell Fellows, a VP and principal analyst at Signal65, wrote that the results “fundamentally alter the calculus of AI efficiency.”

“The demonstrated throughput can support thousands of concurrent user interactions per second on a platform designed to meet complex regulatory requirements, enabling the deployment of at-scale AI inference services in sensitive industries,” Fellows wrote.

The road to a million tokens: Why does it matter?

AI models increasingly used in enterprise settings are getting exponentially larger, as is their ability to handle more tokens.

For those unversed in AI, tokens are units of data. They’re typically a piece of a word, or in the case of non-text modalities, a portion of an image or an audio clip. AI models break down inputs into tokens to process them piecemeal.

As Deloitte Nicholas Merizzi outlined in a recent SDxCentral op-ed, tokens have now become the de facto billing metric for AI services, with enterprises paying by the number of tokens that their AI consumes.

Industry eyes balked when Google claimed in February 2024 that its then-published Gemini 1.5 offered a one million-token context window.

But in the time since Gemini 1.5's launch, models have significantly increased in size and token capacity. For example, Meta’s Llama 4 family of models, released in April, greatly surpasses Gemini 1.5 – with even its smallest system, Scout, capable of processing 10 million multimodal tokens.

If, per OpenMetal, a typical English word requires an average of 1.3 to 2 tokens depending on complexity, documents and information used in actual business cases – like spending reports or datasets – could quickly amount to millions of tokens.

Hyperscale heavy-lifting

Microsoft wants its new VMs to power the million token needs. Each v6 rack powers a total of 18 VMs – based on some 72 Nvidia Blackwell Ultra GPUs.

The hyperscaler claims its record-breaking systems are “optimized for reasoning models, agentic AI systems, and multimodal generative AI.”

Its MLPerf Inference results were achieved running Meta’s LLama 2 70B model. While it doesn’t boast the context capabilities of its recently released successor, the model is widely used in existing enterprise environments.

The tests saw Microsoft’s new Blackwell-powered VMs achieve an average throughput of 48,088 tokens per second, compared to just 12,022 tokens per second on the prior generation GB200 system.

“No other major cloud provider has published any MLPerf-like Llama 2 70B inference results near this scale,” Fellows said of Microsoft’s results. “The latest v5.1 submissions achieved roughly 100,000 tokens per second on an 8-GPU DGX B200 configuration, nearly (10-times) slower than Azure’s validated rack-scale result.”