AMD unveils a workstation built to run trillion-parameter AI models locally

Started by RayOfLight32, Sep 06, 2026, 05:36 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: AMD unveils a workstation built to run trillion-parameter AI models locally   Views(Read 68 times)
Active members in this topic:
RayOfLight32(1) SwiftQuarry(1)

RayOfLight32

AMD opened IFA 2026 in Berlin with the keynote slot, the first time in the show's 102 year history a silicon company has held it, to unveil the Threadripper Halo Station, a liquid-cooled workstation built around a 96 core Threadripper PRO 9995WX processor codenamed Shimada Peak paired with AMD's Instinct MI350P accelerators. AMD says the system can hold and run AI models with more than one trillion parameters entirely on-device, without any cloud connection required

The base configuration shown at IFA includes two MI350P accelerators, each independently liquid-cooled alongside the CPU, with each card offering 144GB of HBM3E memory and up to 4TB per second of bandwidth. AMD says the chassis has a path to support up to four accelerators, which would bring total accelerator memory to 576GB, enough headroom for genuinely massive models given that even a relatively modest 300 billion parameter model requires around 600GB in standard precision or roughly 150GB at a heavily compressed format. AMD also announced Kraken Halo, a Ryzen AI Max Pro 400 platform with 192GB of unified memory aimed at 300 billion parameter models, with Lenovo and HP confirming they'll ship the first commercial systems

AMD is positioning the Halo Station as a direct alternative to Nvidia's DGX Station towers, which reportedly start around 100,000 dollars, framing local AI as a way to eliminate ongoing cloud inference costs and keep sensitive data entirely on-premises. No official pricing or launch date was announced, though each MI350P accelerator alone is rated for up to 600 watts, meaning a full four-card configuration would demand serious power delivery and cooling infrastructure. Curious what people think about this shift toward desk-sized systems capable of running models that recently required entire data centers, does genuinely private, offline local AI at this scale change how businesses and researchers actually approach sensitive AI workloads


SwiftQuarry

576GB of accelerator memory in a single desk-sized workstation is a genuinely staggering figure to sit with, that's a scale of on-device AI capability that simply didn't exist as a personal or small-team option even a year or two ago

Save money on everyday spending Free cashback on thousands of retailers
View offer