PCIe GPUs for efficient, scalable AI inference
This session took place on May 13, 2026
Please complete the following form to access the full presentation.
PCIe GPUs for efficient, scalable AI inference
You can now ask questions directly within our broadcasts using our new AskAI feature.
Enterprise AI is shifting from model training to large-scale inference, driving demand for scalable, efficient deployment in real-world, air-cooled data center environments. This episode explores how next-generation PCIe GPU architectures support modern inference-driven AI infrastructure by optimizing coordination across compute, memory, storage, and networking resources. Learn how enterprises can improve throughput, reduce I/O bottlenecks, and scale AI inference workloads within existing power and cooling constraints without requiring major data center retrofits or specialized infrastructure upgrades, while balancing performance, utilization, and total cost of ownership for on-prem AI deployments. Key discussion points include:
- The shift from AI training to large-scale enterprise inference workloads
- PCIe GPU architectures and inference-driven infrastructure design
- Optimizing data movement across CPU, GPU, storage, and network pipelines
- Scaling AI deployments efficiently within existing power, cooling, and cost constraints
- Speakers
- Kat Sullivan , DatacenterDynamics and SDxCentral
- Rahul Deshmukh , Supermicro
- Isabelle Liu , AMD
- Brought to You by
- Supermicro