Researchers from IBM and Ohio State University are calling for greater transparency from hardware vendors after reverse-engineering a confidential computing system from Nvidia designed to secure AI workloads on GPUs.
In a recently published paper, the researchers dissected GPU Confidential Computing (GPU-CC), a security feature introduced in Nvidia’s Hopper GPU architecture and targeted at safeguarding sensitive data during processing.
Due to the proprietary nature of Nvidia's ecosystem, GPU-CC has been difficult for the security community to analyze. The researchers described the system as “opaque,” with little technical documentation and no public specification available.
To investigate, the team instrumented the GPU kernel module – the only open-source component in Nvidia’s GPU software stack – allowing them to trace system behaviors and identify potential vulnerabilities.
The research was billed as an effort to “demystify” Nvidia’s confidential GPU features and provide an independent security assessment, arguing that as AI workloads increasingly rely on GPU acceleration, security guarantees must extend beyond the CPU.
GPU-CC under the hood
The tech behind confidential computing dates back to the early 2020s, when, according to Nvidia engineers, it became clear during development of the Hopper architecture that GPUs would need to protect not just data, but the code running on them.
The prior Ampere architecture was capable of securely protecting a user’s data via Nvidia’s Ampere Protected Memory (APM) solution, but that didn’t cover code.
Confidential computing was initially released in private preview for early access in July 2023, with Nvidia first revealing it to the public the following April.
At the time of release, Rob Nertney, a senior software architect at Nvidia, said the feature “addresses the need to secure data in use, and prevent[s] unauthorized users from accessing or modifying the data.”
The concept also extends to Nvidia’s next-generation Blackwell architecture, which is slowly making its way into the hands of customers after a costly delay. Earlier this week, CoreWeave became the first cloud vendor to get its hands on the coveted Nvidia GB300 NVL72 rack-scale solution.
The researchers contend that for end users, enabling GPU-CC is a seamless feat, with applications able to run without any modifications. But the lack of concrete details on how it works sat uneasily with them.
The team dug into the only part of the stack Nvidia leaves open, the GPU kernel module, and began instrumenting it to trace what actually happens under the hood when GPU-CC is enabled.
From there, they were able to reverse-engineer system behaviors, run targeted experiments, and map out how secure enclaves are bootstrapped and how memory is allocated and protected.
In cases where components were too locked down to inspect directly due to proprietary firmware, they simply filled in the gaps with informed speculation grounded in what they could observe at the software interface level.
While it wasn’t a full look into the inner workings of GPU-CC, the researchers were able to piece together a coherent architecture map and expose areas where the protections fall short.
What did they uncover?
The researchers revealed significant security gaps in GPU-CC. For example, threat actors that have physical or remote access to a GPU are able to manipulate the hardware’s settings, including critical GPU-CC security configurations, using in-band tools like nvTrust or out-of-band interfaces such as baseboard management controllers.
Such access would allow attackers to potentially compromise data flowing between confidential virtual machines (VMs) and GPUs, including memory transfers over the PCIe bus and data stored in on-package high-bandwidth memory.
Nvidia’s GPU-CC was also found not to encrypt GPU memory at runtime, unlike CPU-based confidential computing systems. Instead, GPU-CC relies solely on access control mechanisms like firewalls to protect memory regions, which creates potential vulnerabilities if such controls are bypassed.
While Nvidia’s architecture includes protections against some of the identified risks, the researchers noted that limited transparency makes it difficult to independently verify their effectiveness.
“The Nvidia GPU-CC system is designed to provide an isolated and secure execution pipeline for emerging AI applications that handle sensitive data,” they wrote. “However, the lack of public specifications, the proprietary nature of its hardware and software ecosystem, and the complexity of its design present significant challenges in evaluating whether its implementation satisfies the security requirements of confidential computing.”
All security flaws detailed in the paper have been reported to Nvidia.
Comments