CoreWeave
– CoreWeave

CoreWeave has expanded the functionality of its unified operating standard, Mission Control, in a bid to support enterprise technology teams running large-scale AI workloads.

As a neocloud, CoreWeave specializes in GPU-as-a-Service, with high-performance computing power and GPUs that companies increasingly need to resolve performance issues as they scale workloads for AI, machine learning, and data analytics.

The latest addition to Mission Control offers builders access to capabilities, including immediate, verifiable visibility into every access event in CoreWeave, helping to diagnose and resolve bottlenecks impacting distributed training performance.

Peter Salanki, co-founder and CTO of CoreWeave, said: “Mission Control gives enterprises the first true operating standard for AI at production scale. With one place to see what is happening and why, teams can resolve issues quickly and keep their workloads running at full performance while they focus on deploying innovation.”

CoreWeave’s Mission Control provides comprehensive, real-time visibility into GPU, network, and storage performance, enabling teams to understand system behavior and maintain secure performance across their environments.

The expanded Mission Control release includes additions like Telemetry relay, which streams audit and access logs from CoreWeave services into a user’s security information and event management (SIEM) platform. Another addition is GPU straggler detection, which the neocloud says provides rank-level visibility inside distributed training jobs and identifies the exact GPU or node causing a straggler. This is also supported by overlays from observability and data visualization company Grafana.

A newly introduced Mission Control agent is also said to transform the Mission Control operating standard into a conversational assistant that teams can interact with directly to help understand system behavior, troubleshoot faster, and turn complex telemetry into actionable guidance.

Commenting on the development, Ash Mazhari, VP of corporate development at Grafana Labs, said: “We’re proud to formally partner with CoreWeave on Mission Control, which raises the bar for observability in AI infrastructure by giving teams unified, real-time insight into GPU performance, access activity, and distributed training behavior.”

In a recent development, CoreWeave moved to cover the cost of data migration from hyperscaler rivals through an egress fee-free data migration program for customers moving AI workloads to CoreWeave from the likes of Amazon Web Services (AWS), Microsoft Azure, Google Cloud, IBM, and Alibaba.

According to the firm, enterprises will be able to maintain active accounts with third-party cloud providers during and after the migration, with the promise of no exit fees from CoreWeave.

In November, CoreWeave also signed a mammoth deal with AI storage platform Vast Data, designed to underpin neocloud infrastructure serving Meta, OpenAI, and Microsoft workloads.

The $1.17 billion deal with the AI operating system vendor positioned Vast’s AI OS as the primary data storage and management platform under the neocloud’s compute infrastructure.

Reflecting on the emerging importance of robust AI and machine learning workflows, recent Gartner research affirmed the role of neoclouds in reshaping AI infrastructure.

The research powerhouse particularly emphasized neoclouds' ability to solve the cost, agility, and supply challenges that hyperscalers face, going a far as saying that “tech service leaders who fail to integrate neoclouds in their portfolio risk higher costs, slower innovation, and diminished competitive edge in the AI services race”.