GetChain News
中简 中繁 EN
GetChain News
Toggle sidebar

OpenAI Researcher Praises AMD and Cerebras Joint AI Inference Solution, Estimates 5x Performance Per Watt Improvement

Source: www.ithome.com
OpenAI researcher Jeffrey Wang described the performance of the AI inference solution jointly developed by AMD and Cerebras as "incredible". The solution integrates the AMD Helios rack with the Cerebras wafer-scale engine, handling prompts and token generation stages separately, and is estimated to increase tokens per second per watt by 5 times. This technical approach optimizes inference performance through heterogeneous computing, reflecting the continued focus of leading AI labs on reducing inference costs.

Related projects