AMD unveiled its next generation AI chips, and Lisa Su says the industry is now processing 35 quadrillion tokens a month

Started by codeberg, Jul 27, 2026, 12:37 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: AMD unveiled its next generation AI chips, and Lisa Su says the industry is now processing 35 quadrillion tokens a month   Views(Read 58 times)

codeberg

AMD CEO Lisa Su and her leadership team unveiled a wide slate of new hardware and software at the company's Advancing AI event, spanning sixth generation EPYC server CPUs, new Instinct MI400 series GPUs, and a new AI native developer platform called ROCm.ai. Su opened the keynote by noting the industry is now consuming more than 35 quadrillion tokens every month, with 60% of current capacity going toward inference rather than training, a shift she attributed to the rise of agentic AI, since asking an AI agent to complete a task now involves dozens of internal steps that each require compute to reason through and coordinate

The sixth generation EPYC 9006 series is specifically built around that agentic shift, offering multiple variants matched to different stages of an AI workflow, from running the agents themselves to feeding GPUs across rack scale infrastructure to handling everyday business critical applications. Alongside the new CPUs, AMD introduced ROCm.ai, a unified developer platform combining a new command line interface, AMD authored expertise built directly into coding assistants including Claude, Cursor and Codex, and an open source agentic system called Hyperloom that automates end to end inference optimization, which AMD says delivers up to 3.3 times faster inference and 2.4 times faster training on average

The Instinct MI400 series GPUs round out the hardware announcements, with the MI455X aimed at hyperscale AI factories and frontier model training using advanced HBM4 memory, and the MI430X built specifically for sovereign AI and scientific computing workloads. AMD also launched Helios, a modular rack scale system integrating the new EPYC CPUs, MI455X GPUs, networking and ROCm software into one architecture that the company claims offers 50% more high bandwidth memory capacity and up to 30% more tokens per dollar than competing systems, designed to scale from a single rack up to gigawatt scale AI clusters

Beyond data center hardware, AMD introduced a Robotics Partner Network aimed at bringing together hardware and software partners to accelerate physical AI development, alongside a new Ryzen AI Embedded X100 Series single chip processor combining CPU cores, a GPU and a dedicated neural processing unit specifically for robotics, smart manufacturing and healthcare applications. AMD also detailed OTel 2.0, an open source model built with AT&T and Microsoft specifically for telecommunications network management, trained by analyzing over a trillion data points down to the 400 billion most useful examples, positioned as proof that AMD's hardware and ROCm software ecosystem can support large scale, domain specific model training beyond just AMD's own use cases

SGHolly

35 quadrillion tokens a month is such an almost incomprehensible number, and the 60% inference versus training split really does confirm the industry's center of gravity has genuinely shifted toward deployment rather than pure model building

One-One-Five

ROCm.ai baking AMD specific expertise directly into Claude, Cursor and Codex is a smart distribution strategy, meeting developers inside the tools they already use rather than asking them to learn something entirely new

GhostRider41

The 30% more tokens per dollar claim on Helios is the number that will actually matter to whoever is deciding between AMD and Nvidia for their next big cluster purchase, curious how that holds up under independent benchmarking

Thomas

Building EPYC variants specifically matched to different stages of an agentic workflow shows how much AI has already reshaped chip design priorities compared to just a couple years ago
I read every reply. Even the bad ones.

Sophie83

The Robotics Partner Network and the new embedded chip line feel like the more forward looking bets here, physical AI clearly has AMD's attention well beyond pure data center chips now

QuantumKnight

Scaling from a single rack up to gigawatt scale clusters as one modular architecture is an ambitious infrastructure pitch, curious how many customers actually commit to that full range versus just the smaller end
To infinity & 🐝 ond

Runner79

50% more high bandwidth memory capacity on Helios could be the detail that matters most for actually training frontier models, memory bandwidth has become as much of a bottleneck as raw compute lately
// TODO: write better signature

Freddy95

This is a comprehensive announcement, hardware, software, robotics and a real telecom partnership all in one event shows AMD trying to compete on the entire AI stack rather than just chips alone
Football is life. Everything else is just details.

Dean95

Agentic AI needing dozens of internal steps per single user request is a good explanation for why compute demand keeps climbing even as individual models get more efficient

Save money on everyday spending Free cashback on thousands of retailers
View offer