AMD CEO Lisa Su: The MI350 series has been shipped, and the MI400 and Helios AI rack lead the open ecosystem.

CN
16 hours ago

Author: Techub News Compilation

Introduction

At midnight on June 12, Beijing time, AMD held its annual AI conference "Advancing AI 2025" in Silicon Valley. During this event, which lasted over two hours, AMD's Chairwoman and CEO Lisa Su, along with the company's executive team, unveiled the latest advancements and future roadmap in AI computing hardware, software, and solutions in collaboration with top AI companies and cloud service providers such as xAI, Meta, Oracle, Microsoft, OpenAI, and Cohere. This event takes place at a crucial juncture of explosive growth in AI models and applications, along with a dramatic increase in computing demand. How AMD leverages its comprehensive product portfolio of "CPU+GPU+DPU+FPGA" and its strong commitment to open ecosystems to compete for the data center AI accelerator market worth hundreds of billions of dollars has become a focal point of industry attention.

Summary

  • Launched the flagship AI accelerator MI350 series, achieving a 4-fold performance generational improvement over MI300, with inference performance in certain scenarios reaching 35 times, and has started production and shipment.
  • First preview of the 2026 rack-level solutions MI400 series and Helios AI rack, aiming to provide computing power for trillion-parameter models, with a performance target increasing by another 10 times.
  • Released version 7 of the ROCm software stack, which enhances inference performance by 3.5 times and strengthens training support, deeply integrating with open-source communities like PyTorch, vLLM, and SGLang.
  • Announced details of large-scale deployments and collaborations with customers such as xAI, Meta, Oracle, and Microsoft, highlighting the practical competitiveness and cost advantages of AMD Instinct in inference and training workloads.
  • Emphasized that an open hardware (UALink, Ultra Ethernet) and software ecosystem is central to accelerating AI innovation and scaling, with AMD committed to providing full-stack solutions to meet diverse needs.

A New Chapter for AI: Inference Explosion and the Rise of Agents

Lisa Su pointed out at the outset that since the advent of ChatGPT, the pace of AI innovation has been unprecedented in her career, and this process will only accelerate by 2025. She observed several key trends: more powerful inference models emerging, the rise of agents, and real-world use cases starting to be deployed on a large scale. She particularly emphasized that inference is becoming the largest driver of growth in AI computing, with growth rates expected to exceed 80% in the coming years. This is due to improved model capabilities and the emergence of new use cases, significantly increasing the utilization of AI.

Another significant change is the explosive growth in the number of models. Not only are there cutting-edge models from OpenAI and Google, but also open-source models from Meta, DeepSeek, and others, along with dedicated models built for fields like healthcare, finance, programming, and scientific research. Lisa Su predicts that in the coming years, there will be hundreds of thousands or even millions of specialized models built for specific tasks, industries, or use cases.

Regarding the recently discussed agent AI (Agentic AI), she defined it as a new class of "users": they are always online, continuously accessing data and observing applications and systems to make autonomous decisions and work. This not only requires high-performance GPUs to generate insights in real-time but will also create a significant traditional computing demand as agent activities increase, shifting towards high-performance CPUs. She vividly illustrated, "We are effectively adding billions of new virtual users to the global computing infrastructure." All of this necessitates GPUs and CPUs working together in an open ecosystem.

Market, Strategy, and Trust: AMD's Three Pillars

Reflecting on last year's forecast that the total addressable market (TAM) for data center AI accelerators will reach $500 billion by 2028, Lisa Su stated that all current signs indicate this number will be even higher. She reiterated that AMD's core strategy is based on three key principles:

  • Providing a wide range of computing engine combinations: Covering CPUs, GPUs, DPUs, NICs, FPGAs, and adaptive SoCs, allowing customers to match the most suitable computing units for specific models and use cases.
  • Vigorously investing in an open, developer-first ecosystem: Supporting all mainstream frameworks, libraries, and models, harnessing industry power through open standards.
  • Providing full-stack solutions: Integrating all elements by building and strengthening partnerships.

She notably emphasized the value of "trust," asserting that trust must guide the direction as we move towards the future of AI. The establishment of trust stems from a steadfast execution of the roadmap, building an open ecosystem, and providing the broadest AI product offering in the industry.

Client Voices: Practices from xAI, Meta, and Oracle

To showcase the practical application of AMD technology, Lisa Su invited several heavyweight partners to share their insights.

xAI co-founder Xiao Sun shared his experience using AMD MI300X to accelerate their Grok series models. He described the process as "effortless." As a small, fast-moving team, xAI's most valuable resource is engineering time. With close cooperation from AMD engineers (even often communicating late at night or at midnight), they pushed the Grok model into production within just a few months. Xiao Sun praised AMD's annual hardware update cadence and, starting from first principles, envisioned "the largest collaborative design in human history" from silicon to product, anticipating accelerated iterations with vendors like AMD.

Meta's Vice President of Engineering Yee Jiun Song (YJ Song) stated that AMD is Meta's strategic and responsive partner. The MI300X accelerator is now a critical part of Meta's infrastructure, widely used for inference with Llama 3 and Llama 4 due to its high performance and excellent total cost of ownership (TCO). Meta is expanding the use of MI300X to more workloads, including the training and inference of ranking and recommendation models that are crucial to its business. Regarding MI350, he is optimistic about the significant increase in computing power, new generation memory, and support for FP4/FP6 while maintaining the same form factor as MI300 for quick deployment. YJ Song also pointed out that AI not only improves existing products but also gives rise to entirely new products, driving unprecedented-scale investments in computing infrastructure. The close collaboration between Meta and AMD in software (PyTorch, ROCm) and hardware (Open Compute Project OCP) has been crucial in quickly bringing MI300 into a production environment.

Oracle Cloud Infrastructure (OCI) Executive Vice President Mahesh Thiagarajan discussed the deep integration with AMD. He emphasized that to tackle the cutting-edge challenges at the intersection of cloud and AI, full-stack deep integration from power, computing, networking to storage is needed to extract every ounce of performance. AMD’s Infinity Fabric technology enables the high-speed movement of datasets to the accelerator. Additionally, high-performance networks (like Pensando) are critical for building AI superclusters that operate as a single giant supercomputer. The demand for MI300X on OCI is huge, coming from both AI-native companies and large cutting-edge model firms. OCI ensures customers can immediately access the latest ROCm innovations and performance enhancements from AMD. He announced that OCI will launch a single cluster supporting over 27,000 GPUs within two months. Mahesh also discussed the challenges and opportunities of building gigawatt-level data centers, asserting that the core of the collaboration with AMD is to deliver "price-performance" value to customers.

MI350 Series Launch: 4 Times Performance Leap and Cost Advantages

Lisa Su officially launched the AMD Instinct MI350 series, including the flagship products MI355 and MI350 (both using the same chip, with MI355 supporting higher thermal design power and performance release). This marks the largest generational performance leap in the history of Instinct.

MI355 features the fourth generation Instinct architecture, supports new data formats like FP4, and is equipped with the latest HBM3E memory, integrating 185 billion transistors through 3D packaging. Its key specifications and advantages include:

  • 4 times AI computing performance upgrade: Accelerating training and inference.
  • 288GB industry-leading memory: Capable of running models with up to 520 billion parameters on a single GPU.
  • Compatibility: Utilizes the same industry-standard UBB8 platform as MI300/MI325, making it easy to deploy into existing data center infrastructure.
  • Competitive advantage: Compared to competitors, MI355 supports 1.6 times more memory, doubles throughput on FP6 and FP64, and has platform-level advantages in FP4 computing and HBM3E memory capacity.

In terms of performance, on Llama 3.1, MI355 achieves up to 35 times higher throughput than predecessor models when running ultra-low latency applications (like code completion, real-time translation). For applications such as chatbots, content generation, summarization, and conversational AI, performance improvements can reach 4.2 times. In models like DeepSeek and Llama 4 Maverick, token generation speeds can be three times faster than previous generations.

Compared to competitors, when running DeepSeek R1 or Llama 3.1 using open-source frameworks like SGLang and vLLM, MI355's throughput surpasses that of competing product B200, and even matches the performance of the more expensive and complex GB200. More importantly, combined with lower capital expenditures (CapEx), MI355 can generate 40% more tokens per dollar compared to competing solutions, providing cloud and enterprise customers with higher throughput, efficiency, and better overall cost of ownership (TCO).

Lisa Su announced that production shipments of MI355 began earlier this month, with the first partner platforms and public cloud instances expected to launch in the third quarter.

The Power of Software: ROCm 7 and Developer Ecosystem

AMD Senior Vice President and Head of AI Business Vamsi Bopanna took the stage, emphasizing the crucial role of software in unlocking the full potential of AI hardware. He announced the launch of ROCm 7, the latest version of AMD's open-source software platform.

ROCm 7 brings several important enhancements:

  • Focus on inference: Innovations at every layer of the inference stack, with performance improving over 3.5 times compared to ROCm 6.
  • Embracing open source: The capabilities and performance of open-source frameworks like vLLM and SGLang have started to surpass closed alternative solutions. By closely collaborating with these communities, MI355's throughput on DeepSeek FP8 can exceed B200 by 1.3 times.
  • Strengthened training support: Compatibility with all major parallel strategies and frameworks (like PyTorch, JAX), with training performance also improved to 3 times that of ROCm 6.
  • Enhanced usability: Continuous improvements in out-of-the-box capability, simplification of setup, increased resources, and strengthened community interaction through hackathons, competitions, and meetups. The release cadence has accelerated to every two weeks.

Microsoft AI Platform Vice President Eric Boyd came on stage as a partner to share how Microsoft uses AMD Instinct for inference and training in AI Foundry and internally. He mentioned that Instinct chips' large memory and high bandwidth provide significant TCO advantages for serving large language models. Microsoft uses the SGLang engine to infer open-source models like DeepSeek, and ROCm's integration with open source makes large-scale deployment straightforward. Additionally, the Microsoft research team has trained state-of-the-art multimodal models on 2,100 MI300X, with the same platform providing great flexibility in data centers for both inference and training.

Cohere co-founder and CEO Aidan Gomez (one of the authors of the Transformer paper) also shared his experience with training using ROCm. He stated that Cohere focuses on providing secure and private AI for enterprises in highly regulated industries like finance and healthcare. After running models and training on AMD hardware, Cohere was impressed by ROCm's progress, especially gaining confidence when scaling training tasks. He looks forward to leveraging the higher memory bandwidth of the MI350 series to train larger, more complex models in the future.

Looking to the Future: MI400, Helios Racks, and Open Standards

In the latter half of the event, Lisa Su turned her attention to the more distant future, giving a first glimpse of the Instinct MI400 series and its corresponding Helios AI rack, scheduled for launch in 2026. This is a completely integrated AI rack platform designed from the ground up for large-scale training and distributed inference.

The Helios rack will unify CPU (next-generation EPYC "Venice," based on 2nm process with Zen 6 cores), GPU (MI400 series), Pensando NIC (next-gen "Volcano"), and ROCm software into a single system. Its design goals include:

  • Connecting up to 72 GPUs, providing 260 TB/s of vertical scaling bandwidth.
  • Delivering 2.9 ExaFLOPS of FP4 performance.
  • Supporting 50% more HBM4 memory, memory bandwidth, and horizontal scaling bandwidth compared to competitors.

The MI400 series is expected to bring up to a 10 times performance improvement for cutting-edge models. Lisa Su stated that customer enthusiasm for MI400 and Helios is extremely high, as AMD has engaged in preliminary collaboration with customers for many years.

To validate this point, she welcomed an esteemed guest—OpenAI founder and CEO Sam Altman. Altman reflected on the astonishing progress of AI from the early 2020s to the present, predicting that the same pace of advancement will continue over the next five years (until 2030), leading to novel scientific discoveries and the execution of extremely complex social functions. He emphasized that achieving this goal requires deep collaboration across research, engineering, and hardware. Regarding collaboration with AMD, Altman expressed "extreme excitement" about MI450 (likely a part of the MI400 series), noting that its memory architecture is ideal for inference, believing it will also be an excellent choice for training. He recalled feeling that the initial specifications were "impossible," "too exaggerated," but now feels very excited to see AMD nearing delivery.

Additionally, the speech highlighted AMD's efforts in promoting open industry standards:

  • UALink: An open standard for high-speed interconnect between GPUs within a rack, aimed at breaking the monopoly of closed interconnect solutions. AMD is collaborating with partners like Astera Labs and Marvell to advance this.
  • Ultra Ethernet Consortium (UEC): A networking standard for horizontally scaling massive AI clusters. AMD is a founding member, and its Pensando P4 smart NIC has supported UEC 1.0 (which officially launched the day before the event), enhancing AI performance and reducing network costs.

Sovereign Computing and Conclusion

Lisa Su also mentioned AMD's progress in sovereign computing, collaborating with over 40 governments and research institutions globally to build high-performance computing and AI infrastructures. She cited AMD's Silo AI laboratory (in collaboration with various European countries) and its partnership with Saudi Arabia's newly established company Humain as examples of how to utilize localized, open AI technology to drive societal impact and economic development. Humain CEO Tareq Amin announced on-site a joint venture with AMD, pledging to leverage Saudi Arabia's advantages in land and energy to reduce ownership costs for global AI developers by 30%, with plans to build gigawatt-level data centers.

In summary, Lisa Su stated that the past year has redefined the speed of progress in AI, placing this community at the center of all significant transformations. She believes that the future of AI will not be shaped by any single company or built within a closed ecosystem but will be defined by open collaboration across the entire industry. AMD looks forward to partnering with all stakeholders to make AI stronger, more accessible, and more useful for everyone through open technology, talent, and collaboration to change the world together.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink