AMD acquires Taalas, a startup that hardwires entire AI models into silicon

Started by Mark1, Aug 09, 2026, 02:31 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: AMD acquires Taalas, a startup that hardwires entire AI models into silicon   Views(Read 43 times)

Mark1

AMD has agreed to acquire Taalas, a three year old Toronto startup that takes a genuinely radical approach to AI chips, physically hardwiring an entire AI models weights directly into silicon logic rather than treating the model as software loaded onto a general purpose processor

Taalas CEO Ljubisa Bajic describes the companys core philosophy as the model is the computer, instead of relying on high bandwidth memory to feed model weights into programmable compute engines every time a query runs, which is how virtually every GPU from Nvidia and AMD alike currently works, Taalas uses a specialized design flow to etch a models math, routing and parameters permanently into CMOS logic, stripping out general programmability entirely in exchange for bypassing the memory bandwidth and power bottlenecks that plague standard AI inference

The tradeoff is genuinely stark, a chip hardwired for one specific model can only ever run that model, switching to a different or updated model means fabricating entirely new silicon, but the performance gains are apparently substantial enough to justify that rigidity, Taalas demonstrated a test chip earlier this year tailored specifically for Metas Llama 3.1-8B model that delivered over 16,000 tokens per second per user, with some reporting putting the figure as high as 17,000 tokens per second at roughly one tenth the power draw of an Nvidia H200, blowing past the throughput limits of general purpose GPUs entirely

AMD isnt planning to replace its broad purpose Instinct GPUs with this technology, instead the plan is disaggregation by inference phase, according to people familiar with AMDs integration planning an Instinct GPU would handle the prefill stage, processing the initial user prompt, which is compute intensive and benefits from a GPUs massive arithmetic throughput, while the hardwired Taalas accelerator handles the decode stage, generating each output token one at a time, a workload thats almost entirely memory bound and exactly where Taalas fixed function silicon provides its biggest advantage, this is set to integrate into AMDs Helios rack scale systems alongside Instinct GPUs and EPYC server processors, with Helios having entered mass production in July with Microsoft, OpenAI and Anthropic already confirmed for Q4 deployments

Financial terms werent disclosed, Taalas has raised 219 million dollars in venture funding total since its 2023 founding, and the deal is expected to close in the fourth quarter, this is AMDs third notable AI acquisition after paying 665 million dollars for Silo AI and 4.9 billion for ZT Systems in 2024, and comes a little over seven months after rival Nvidia spent a reported 20 billion dollars buying assets from Groq, its largest transaction on record, CEO Lisa Su had previewed this exact philosophy back at AMDs Advancing AI 2026 conference in July, saying directly that theres no one size fits all as it comes to chips, arguing GPUs remain necessary for the experimental, rapidly iterating work of model development, while fixed function silicon can serve the high volume, cost sensitive business of actually running those models at production scale

Current

The model is the computer philosophy is such a genuinely radical departure from how every AI chip has worked so far, betting that inference at scale is worth sacrificing all flexibility for raw throughput and power efficiency is a fascinating architectural gamble

Henry75

Splitting inference into prefill on a GPU and decode on hardwired silicon is a genuinely clever way to get the best of both approaches, use the flexible general purpose chip for the part that needs flexibility and the rigid specialized chip for the part thats purely memory bound

Ann

17,000 tokens per second at a tenth the power draw of an H200 is a staggering efficiency claim if it holds up in production deployments rather than just a controlled benchmark demo, that kind of gap would genuinely reshape the economics of running models at scale
RTFM and then ask

Anchor34

The obvious risk here is model obsolescence, if you hardwire silicon to Llama 3.1 and a meaningfully better model comes out six months later, that expensive fixed function chip becomes a very expensive doorstop, this only works if the underlying models stay stable enough to justify the fabrication lead time

Drift Vector

Two months from an unseen model to realized hardware according to Taalas own website is a fast turnaround for chip fabrication, though that timeline presumably assumes using an existing older TSMC process node rather than pushing for the absolute cutting edge

Hannah

Helios already having Microsoft, OpenAI and Anthropic confirmed for Q4 deployment is the detail that makes this feel real rather than speculative, this technology has a genuine near term path into production infrastructure at some of the biggest AI labs rather than staying a lab curiosity

SoloOrca

Lisa Sus no one size fits all framing from the Advancing AI conference in July aged really well given this acquisition landed just weeks later, shows this wasnt a reactive scramble but a deliberate strategy that had already been signaled publicly before the deal was announced

Related Topics (3)

Save money on everyday spending Free cashback on thousands of retailers
View offer