Data platform firm Weka has developed a new solution aimed at breaking AI workload bottlenecks through software-defined storage.
Dubbed NeuralMesh Axon, Weka’s software turns existing resources inside a GPU server into a unified storage fabric – allowing infrastructure operators to augment their existing clusters.
NeuralMesh Axon brings the data right next to the compute, which Weka argues reduces GPU idle times and thereby saves users millions of dollars in wasted expenditure.
Where previously each GPU server would have its own local NVMe drives that largely sat idle, Weka’s software alternative creates a mesh connecting the individual drives into one high-performance storage pool.
The team behind NeuralMesh Axon claims it can handle up to four simultaneous node losses, or where a server fails, sustaining full throughput during rebuilds, and even allow operators to predefine resource allocation across NVMe and networking resources.
“NeuralMesh Axon turns those architectural advances into real-world impact for training and inference,” Ajay Singh, Weka’s chief product officer, wrote in a blog post.
“For their training workloads, it pushes AI workloads three times beyond typical utilization rates so organizations can run more models on less hardware.”
The rising demand for AI has seen GPU clusters soar in size and complexity, meaning operators need to keep their systems running smoothly as the smallest inefficiencies could quickly become expensive headaches.
Weka’s solution builds on an earlier offering under the NeuralMesh. Unveiled back in June, the software-defined storage system was the initial iteration of its containerized, mesh-based architecture.
NeuralMesh Axon, however, focuses on improving the very largest workloads – the kinds of applications that would be powered in hyperscale data centers.
Weka’s Singh said the offering helps to achieve over 90 percent GPU utilization while extending GPU memory to handle the increasingly large context windows (the amount of text an AI model can process in one go.)
For reference, Google’s flagship Gemini 2.5 Pro foundation model has one of the biggest context windows on the market, with it able to handle around one million tokens – with the hyperscaler teasing earlier this year plans to raise that figure to two million.
CoreWeave, the AI cloud provider, is among those taking advantage of Weka’s software-defined storage solution, embedding NeuralMesh Axon directly into its cloud platform.
The embed saw CoreWeave achieve microsecond latencies, along with 30 GB/s reads, and 12 GB/s writes per GPU server to fully max out their GPUs – a feat that could help the cloud provider get the best out of its new hardware, as they are the first to get their hands on the Nvidia GB300 NVL72.
“By embedding ultra-low latency NVMe storage alongside GPUs, solutions like WEKA’s NeuralMesh Axon significantly increase bandwidth and effectively expand GPU memory capacity, laying a robust foundation for high-speed inference and the next generation of AI-driven services,” Singh said.
NeuralMesh Axon currently has limited availability for both on-premises and cloud deployment options, with Weka planning to make the software generally available in fall 2025.
Weka’s NeuralMesh Axon is the latest in a growing attempt to break the idle time bottleneck.
Recently unearthed concepts from the world of academia saw Korean researchers propose revamping the Compute Express Link (CXL) interconnect protocol to reduce silent packet drops – where data packets are randomly lost within a network without generating any error messages.
Meanwhile, a group of Chinese researchers sought to address sharing GPUs more efficiently in clusters through gPooling, a concept that boosts resource pooling to improve GPU utilization by up to 2x.
Comments