Nvidia (NVDA) kicked off its GTC occasion in San Jose, Calif., on Monday, debuting various chips and platforms starting from its all-new Nvidia Groq 3 language processing unit (LPU) to its large Vera central processing unit (CPU) rack, designed to go head-to-head with choices from Intel (INTC) and AMD (AMD).
All totaled, Nvidia mentioned it’s rolling out 5 large server racks, every serving totally different functions inside AI information facilities.
The largest announcement of the lot, although, is the Nvidia Groq 3 chip. Nvidia introduced it had entered into an settlement to license know-how from Groq and employed founder Jonathan Ross, president Sunny Madra, and different members of the Groq crew as a part of a $20 billion deal in December.
Groq’s processors concentrate on AI inferencing, or operating AI fashions. It’s what occurs if you sort one thing into OpenAI’s (OPAI.PVT) ChatGPT, Anthropic’s (ANTH.PVT) Claude, or Google’s (GOOG, GOOGL) Gemini and get a response.
Nvidia’s graphics processing items (GPUs) are multipurpose and can each prepare and run AI fashions, however because the AI market strikes towards operating fashions, making certain the corporate has a devoted inferencing chip has turn into paramount.
That’s the place Groq 3 is available in.
According to Nvidia vp of hyperscale and high-performance computing Ian Buck, whereas Nvidia’s GPUs assist way more reminiscence than Groq 3, the LPU’s reminiscence is quicker. So the corporate is combining the efficiency advantages of each chips.
To try this, Nvidia is launching its Groq 3 LPX platform, a server rack powered by 128 particular person Groq 3 LPUs. When used along with Nvidia’s Vera Rubin NVL72 rack the corporate says clients might see 35x increased throughput per megawatt of energy and 10x extra income alternative.
“Optimized for trillion-parameter models and million-token context, the codesigned LPX architecture pairs with Vera Rubin to maximize efficiency across power, memory and compute. The additional throughput per watt and token performance unlocks a new tier of ultra-premium, trillion-parameter, million-context inference, expanding revenue opportunity for all AI providers,” the corporate mentioned in an announcement.
The LPX rack ought to assist handle issues that Nvidia might ultimately lose its edge within the AI race to upstart corporations designing inference-focused processors.
In addition to the LPX, Nvidia revealed its Vera CPU rack. When Nvidia talks about its Vera Rubin superchip, it’s referring to 3 processors in a single: a Vera CPU and two Rubin GPUs.
Now the corporate is breaking off Vera into its personal standalone chip, which it’s going to slot into devoted Vera server racks that mix 256 liquid-cooled Vera chips into one system.