OpenAI Finds That Systems From AMD’s Helios Partner, Cerebras, Are “Incredible” At Inference Tasks, Showing That AMD Chose Wisely

AMD appears to have chosen its partner quite wisely to bolster the Helios rack-scale system's inferencing capabilities, with an OpenAI researcher recently singing what is nothing short a paean in favor of Cerebras' Wafer-Scale Engine.

As we explained in a dedicated post recently, Helios is AMD's first full-stack rack-level solution for AI workloads, consisting of:

  1. AMD Instinct MI455X GPUs
  2. 6th Gen AMD EPYC CPUs
  3. AMD Pensando AI Network Interface Cards (NICs)
  4. AMD Pensando DPU
  5. AMD Infinity Fabric
  6. AMD ROCm Software Stack

To further bolster the utility of its Helios systems for AI workloads, AMD has partnered with Cerebras , and plans to integrate Cerebras' Wafer-Scale Engine with its Helios systems for blazing-fast inference.

For the benefit of those who might not be aware, Cerebras' Wafer-Scale Engine places an entire AI supercomputer’s worth of memory and compute onto a single, giant, interconnected sheet of silicon, where hundreds of thousands of compute cores and tens of GBs of SRAM connect seamlessly, allowing data to move efficiently without ever hitting external network bottlenecks.

Basically, Cerebras' Wafer-Scale Engine can hold an entire medium-sized model - or huge pieces of a large model - within its unified chunk of SRAM. Also, the versatility of Cerebras' approach enables training as well as inferencing workloads.

Under AMD's envisioned roadmap , Helios will provide a high-performance, scalable throughput engine, while Cerebras' technology will provide ultra-fast, ultra-low latency decode and token generation. Together, the two compute engines are expected to deliver up to 5x higher tokens per second per watt.

This brings us to the core topic. An OpenAI researcher, Jeffrey Wang, has just generously praised chips from Cerebras that presumably leverage its Wafer-Scale Engine architecture, noting:

"Internally, we have some OpenAI models that are on Cerebras chips. They're incredible because they have such fast inference."

Wang goes on to declare:

""And what this means for me on the day-to-day is: whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive."

As a refresher, AI models typically context-switch when they jump from one task to another. This could occur within a multi-agentic setting where the model has to perform a series of diverse tasks, or in Mixture of Experts (MoE) models, where various sub-neural networks specialize in performing a particular task.

The fact that OpenAI is already impressed with Cerebras' systems suggests that AMD's Helios rack-scale systems - integrated with Wafer-Scale Engines - will likely sell like hotcakes upon debut.

This becomes all the more important when you consider that inference costs are now the biggest determinant of data center profitability.

Follow Wccftech on Google to get more of our news coverage in your feeds.