پرش به محتوای اصلی
ORYX

Blog · 5 min read

AMD Launches Helios: The Highest Performing Rackscale AI Infrastructure Solution

On July 23, 2026, AMD launched Helios, a rackscale AI infrastructure platform built around the Instinct MI455X GPU, which the company says delivers better performance and economics than Nvidia's Vera Rubin NVL72 rack.

Why the AI Factory Needs a New Architecture

Monthly AI token consumption has grown roughly 158x over two years as AI moves from experimentation into production products and services, while training compute has kept expanding at about 5x per year since 2020, and inference is becoming the dominant AI workload as models now serve billions of interactions.

Agentic AI compounds this further: a single request can trigger multiple reasoning steps, sub-agents, retrieval calls and tool use, plus CPU-side routing, scheduling and memory management — so the pressure is not just for more tokens, but for more inference, orchestration and data movement per useful result, straining compute, memory, networking, latency and cost per token. AMD argues this calls for a new infrastructure blueprint built around leadership compute, open rack architecture, high-bandwidth scale-up/scale-out fabrics, efficient power and cooling, serviceability, and fast-to-deploy turnkey systems, designed together as one system.

AMD Instinct MI455X GPU Is the Engine, AMD Helios Is the System

At the center of Helios is the AMD Instinct MI455X GPU, which AMD describes as a generational leap in AI compute along with HBM4 memory capacity and bandwidth. On the DeepSeek-V4-Flash open-weight model, AMD reports the MI455X delivering up to 34x higher token throughput at high interactivity and up to 18x lower cost per token compared with the prior-generation Instinct MI355X.

Inside AMD Helios: One Rack, One Co-Designed System

Helios is architected around the path a workload takes through the rack — entering at the host layer, moving into the GPU domain, drawing on HBM4 for model and active data, and communicating across the rack over the scale-up fabric. The rack combines sixth-generation AMD EPYC 'Venice' 9006-series server CPUs, fifth-generation AMD Instinct MI455X GPUs, and AMD Pensando networking with the UALoE fabric, connecting 72 GPUs in a single scale-up domain with 260 TB/s of scale-up bandwidth.

AMD says the resulting rack delivers up to 15% more AI compute, 50% more HBM capacity, and 50% more scale-out bandwidth than an Nvidia Vera Rubin NVL72 rack, with Helios totaling 2.9 exaflops of dense FP4 compute, 1.4 exaflops of FP8 compute, 31 TB of HBM4, 1.7 PB/s of aggregate HBM bandwidth, 260 TB/s of scale-up bandwidth and 43 TB/s of scale-out bandwidth.

Performance and Economics at AI Factory Scale

AMD says peak specifications only describe the rack's scale — what matters in production is throughput at a given interactivity level, tokens delivered per dollar, and how many workloads a rack can run at once. On the Kimi K2 Thinking model with a 32K-token input and 8K-token output, AMD's modeled per-GPU Helios throughput is up to 15% higher at low interactivity, 12% higher at medium interactivity, and 10% higher at high interactivity than a modeled Nvidia Vera Rubin NVL72 rack — which AMD also translates into up to 30% more tokens per dollar than Vera Rubin NVL72.

Open at Every Layer, From Rack to Software

AMD's ROCm software stack ties the Helios hardware to production workloads, with native support for PyTorch, TensorFlow and JAX plus optimized libraries, open standards, and tools for deployment, observability and lifecycle management. AMD says Helios is already being adopted by AI companies including OpenAI, Meta, Microsoft and Oracle, alongside infrastructure partners Dell, HPE, IBM and Cisco.

Final Takeaway

AMD frames Helios as scaling the MI455X's compute, memory-bandwidth and HBM gains into a full rack-scale platform of 72 GPUs, EPYC 'Venice' CPUs, Pensando networking, ROCm software and an open architecture — part of what it calls a shift to an annual, open, multi-generation cadence for rack-scale AI infrastructure. As AMD puts it, 'the rack is no longer simply where the AI system is installed' — it is 'the AI system' itself.

Sources

  • AMD Launches Helios: The Highest Performing Rackscale AI Infrastructure Solution AMD

Next step

See ORYX on your own network scenario

Request a demo