NVIDIA to Deploy Groq-Powered AI Racks at Nebius This Year
Synopsis
NVIDIA’s $20 billion Groq acquisition moves into production, with new AI racks promising faster inference speeds than rival OpenAI’s Cerebras-powered systems.
Key Highlights
- The Groq 3 LPX rack from NVIDIA has been in full production and will ship with its Vera CPUs and Rubin GPUs at Neocloud Nebius later this year.
- The rack is the commercialisation of technology via Nvidia’s $20 billion Groq acquisition back in December.
- One rack can deliver 3,400 tokens per second, packed with liquid-cooled functionality for 256 Groq chips.
- Groq’s chips, like Nvidia’s own GPUs, are fabbed on external foundries, with Samsung making them instead of Taiwan Semiconductor Manufacturing Co.
Groq Technology Moves Into Production
According to a report from CNBC, “NVIDIA’s Groq 3 LPX rack has entered into full production and will be rolled out later this year alongside its Vera CPUs and Rubin GPUs at Neocloud Nebius,” Nvidia senior director Dion Harris told reporters.
The rack is said to be the actual monetisation of technology from Nvidia’s giant purchase of almost all of Groq’s assets last December for $20 billion, the biggest acquisition in Nvidia history. A single liquid-cooled rack contains 256 Groq chips and can provide 3,400 tokens per second according to a benchmark by Artificial Analysis. Another aspect that is different about Groq’s chips, which are built by Samsung compared to Taiwan Semiconductor Manufacturing Co., which makes Nvidia’s own GPUs
Bull Case
The $20 billion deal with Groq comes at an affordable risk for Nvidia given its potential strategic advantages to the company. “It can certainly afford a deal of this scale with only modest implications for the business, leaving plenty of room to make big strategic investments without having an undue burden on its balance sheet,” Bernstein analyst Stacy Rasgon said in a note Wednesday to CNBC.
The technology developed by Groq expands Nvidia’s reach across the complete AI compute stack rather than creating a separate competitive product line. Groq LPX for low-latency inference, while GPU training (and presumably large-context processing) continues with Nvidia. The rack pairs with Vera Rubin chips without the need for customers to alter their CUDA workflows further entrenching the software lock-in that has rendered Nvidia untaintable.
The deal has also been executed at speed. On the other hand, it is a pace at which Nvidia announced the Groq purchase in December and had production systems and a named customer (Nebius) within eight months, indicating that Nvidia can absorb the major technology acquired and bring it to market instead of letting it stall during integration.
Nvidia’s inference speed is now better than a competitor, and you can measure the difference. With a comparative maximum of 3,400 tokens per second from the Groq rack (versus 750 tokens per second that OpenAI is promising for its Cerebras-powered “Ultrafast” mode), that’s a figure Nvidia would love to exact in headlines as cloud providers seek ever-faster and more responsive AI inference.
Bear Case
NVIDIA also bought the technology and talent that is at Groq for $20 billion; no, it didn’t buy the company. Because Nvidia acquired only a license to Groq’s tech and hired its staff instead of buying Groq outright, investors have less insight into how much total value Nvidia picked up, or whether that deal will be able to deliver returns large enough to justify such an outsized price.
Source: Yahoo Finance
At Inspirepreneurs Magazine, covering entrepreneurship, business failures, and the human stories behind the world's most ambitious founders. She writes at the intersection of strategy and storytelling.
You Might Also Like
Mark Allison: The Man Who Turned a Failing Elders Group Into a $2 Billion Business
Huawei Emerges Stronger Despite US Sanctions