Scale-Out Networking for the AI Era

AMD AI NIC™ technology builds on generations of innovation to deliver high-performance, programmable, and open networking that enables hyperscalers and cloud service providers to scale AI infrastructure with confidence.

AMD Pensando™ Vulcano 800 AI NIC

The AMD Pensando™ Vulcano 800 AI NIC is the only AI NIC on the market to support up to 2.4 Tbps of scale-out bandwidth per GPU, providing the high-speed connectivity needed for the most demanding AI workloads.

Abstract illustration with glowing blue lines

Faster AI Performance

Abstract glowing 3D computer chip lines

Boosted Cluster Uptime

Server room center exchanging cyber datas 3D rendering

Reduced Capex Spending

Abstract glowing 3D computer chip lines

Enhanced Operational Excellence

AMD Pensando™ Vulcano 800 AI NIC By the Numbers

Up to
13%
Faster AI Job Completion Times¹

Accelerating time to results from training large models to deployment readiness.

Up to
33%
Reduced Switching Costs²

Open systems drive more cost-efficient, scalable AI infrastructure.

Learn More About the AMD Pensando™ Vulcano 800 AI NIC

AMD Helios Rackscale Solution for Frontier AI3

The AMD Helios Rackscale solution design is a fully integrated AI infrastructure, combining the latest AMD Instinct™ GPUs, AMD EPYC™ Server CPUs, and AMD Pensando™ networking, designed using open industry standards enabling large-scale inference, frontier model training and fine-tuning.

AMD Helios Rackscale
AMD Instinct MI350P PCIe Card

AMD Instinct™ MI350P PCIe® Card

The AMD Instinct MI350P PCIe card offers a practical path to scale generative AI workloads across industries with little to no infrastructure change. Built using 4th Gen AMD CDNA™ architecture, these PCIe-based cards deliver exceptional efficiency and performance for mainstream enterprise workloads with a robust full stack enterprise focused software stack to simplify deployment.

IDC Spotlight Blog

Enabling AI-Ready Scale-Out Networking for AI Workloads

Learn how scale-out networking improves GPU utilization, accelerates AI jobs, and reduces the cost of scaling AI infrastructure.

Programmability Enables Protocol Choice

AMD AI NIC™ technology is built on a fully programmable P4 engine, enabling protocols like UEC-ready RDMA and Multipath Reliable Connection (MRC), as well as custom transports, with features that can be tuned and upgraded as standards evolve.

Intelligent Packet Spray

Intelligent packet spray enables teams to seamlessly optimize network performance by enhancing load balancing, boosting overall efficiency, and scalability. Improved network performance can significantly reduce GPU-to-GPU communication times, leading to faster job completion and greater operational efficiency.

AI technology concept
Out-of-order Packet Handling and In-order Message Delivery

Help ensure messages are delivered in the correct order, even when employing multipathing and packet spraying techniques. The advanced out-of-order message delivery feature efficiently processes data packets that may arrive out of sequence, seamlessly placing them directly into GPU memory without the need for buffering.

Programming code abstract technology background of software developer and  Computer script
Selective Retransmission

Boost network performance with selective acknowledgment (SACK) retransmission, which helps ensure only dropped or corrupted packets are retransmitted. SACK efficiently detects and resends lost or damaged packets, optimizing bandwidth utilization, helping reduce latency during packet loss recovery, and minimizing redundant data transmission for exceptional efficiency.

Abstract illustration of a data stream
Path-Aware Congestion Control

Focus on workloads, not network monitoring, with real-time telemetry and network-aware algorithms. The path-aware congestion control feature simplifies network performance management, enabling teams to quickly detect and address critical issues while helping mitigate the impact of incast scenarios.

Abstract data center concept
Rapid Fault Detection 

With rapid fault detection, teams can pinpoint issues within milliseconds, enabling near-instantaneous failover recovery and helping significantly reduce GPU downtime. Tap into elevated network observability with near real-time latency metrics, congestion and drop statistics.

Digital cyberspace and digital data network connections

Product Portfolio

Partners

ASRock Rack logo
Celestica logo
Cisco white logo
Compal logo
Dell Technologies logo
Foxconn logo
Gigabyte logo
Hewlett Packard Enterprise logo
Lenovo White Logo
MiTAC Computing logo
QCT logo
Supermicro logo
Wistron logo
Crusoe logo
IBM Cloud white logo
Oracle white Logo
TensorWave logo
Vultr logo
Zyphra logo
Dell Technologies logo
ddn logo
Vast logo
Weka logo
Arista logo
Cisco white logo
HPE Juniper networking logo
Open Compute Project logo
Ultra Accelerator Link logo
Ultra Ethernet Consortium logo

Resources

Contact Us

Unlock the Future of AI Networking

Learn how the AMD Pensando Pollara 400 AI NIC can transform your scale-out AI Infrastructure.

Explore Data Center Networking

Explore the full suite of AMD networking solutions designed for high-performance modern data centers.

Footnotes
  1. Based on AMD Engineering silicon modeling and AMD synthetic benchmark simulation, using the MOE_4p5T hi_sparsity_ea9 benchmark test to project time- to- solution speed up in days using the FP8 training datatype on a simulated LLM system modeled with 8000 AMD Instinct MI455X GPUs and (3) versus (2) AMD Pensando Vulcano NICs.

    Results reflect analysis of pure training compute (decode layers only) Configuration evaluated with a global batch size of 4096 and a sequence length of 4K, using SwiGLU activation. Results exclude evaluation and checkpointing overhead. FlashAttention v3 with matmul–softmax overlap is assumed. AllReduce and All2All communication costs are fully accounted for and not hidden via tiled compute–communication overlap. Gradient synchronization and FSDP weight prefetch are evaluated across varying levels of overlap to assess scale-out sensitivity. Results assume ideal and not fully optimized real-world behavior and may vary when actual product(s) are released in market MI400-018
  2. PEN-022:  AMD comparison and pricing as of May 18, 2026, for network fabric costs to support 32,000 GPUs. Comparison of a Vulcano-based NIC (VULCANO-CUSTOM-2.4T) deployed as part of a Helios rackscale system with a network based on 1.6T Tomahawk 6 switching with 200G SerDes versus using a competitor 800G NIC with 800G Tomahawk 6 switching with 100G SerDes. Both fabrics were fat-tree topologies built on Tomahawk 5 800G switching platforms, with NIC costs considered comparable. The Vulcano-based design is estimated to deliver up to 33% savings in network switching costs by enabling a more cost-effective architecture with fewer switching platforms, more bandwidth per port on the network, and reduced transceiver cables/optics.

    TH6-100G Serdes Fat-Tree (Competition):
    Switching (TH6C BCM78914 - 128x800G):
    • Leaf Units                                    1,000
    • Spine Units                                      500
    • Total Switches                                1,500
    • TH6 100G Unit Price                  $79,587
    • Total Switching Cost                         $60M
    Cables/Optics:
    • NIC Transceivers (800G-DR8)            64,000 @ $500
    • Leaf/Spine Transceivers (800G-DR8)  192,000 @ $500
    • MPO Cables                                   256,000 @ $89
    • Total Optics/Cables                           $64M
    Total Fabric Cost TH6-100G (Switches + Cables/Optics): $124M

    TH6-200G Serdes Fat-Tree (Vulcano-Custom2.4T / AMD Solution):
    Switching (TH6P BCM78910 - 64x1.6T):
    • Leaf Units                                    1,000
    • Spine Units                                      500
    • Total Switches                                1,500
    • TH6 200G Unit Price                       $66,336
    • Total Switching Cost                         $50M
    Cables/Optics:
    NIC Transceivers (1.6T-DR8)             32,000 @ $900   (50% fewer vs. competition)
    Leaf/Spine Transceivers (1.6T-DR8)   64,000 @ $900
    MPO Cables                                   128,000 @ $89   (50% fewer vs. competition)
    Total Optics/Cables                           $43M
    Total Fabric Cost TH6-200G (Switches + Cables/Optics): $93M

    Capex Savings (Fabric only):
    Savings $:                                      $30.7M
    Savings %:                                        33.1%

    Pricing sources: SemiAnalysis Hyperscaler Networking Model data used with permission; full analysis available via SemiAnalysis subscription. Edgecore switch pricing as of May 17, 2026. Results may vary based on system configuration
  3. Calculations by AMD Performance Labs in June 2025, based on the projected memory capacity/ bandwidth and scale up/out bandwidth specifications of AMD Instinct™ MI455X 72xGPU “Helios” AI Rack vs. the publicly announced NVIDIA “Vera Rubin” 72xGPU “Oberon” Rack. Server manufacturers may vary configurations, yielding different results. MI350-045A

    Calculations by AMD Performance Labs in September 2025, based on the FP8/FP4 datatypes and the projected specifications for AMD Instinct™ MI455X 72xGPU “Helios” AI Rack vs. Publicly announced specs for the NVIDIA “Vera Rubin” 72xGPU “Oberon” AI Rack. Actual results based on production silicon may vary. Server manufacturers may vary configurations, yielding different results. MI350-046B