A Second Brain for AI Workloads
AI chips are everywhere right now, but most of the attention goes to the ones that train the models. Nvidia is now moving decisively into another chip category, with initial deliveries scheduled to begin soon.
Later this year, Nvidia expects to start running its Groq 3 LPX racks at the Nebius data center, where they will work alongside its Vera CPUs and Rubin GPUs. That marks the start of full production for the chip that came out of Nvidia's $20 billion acquisition of Groq technology in December.
Here is the part that might surprise you. These Groq chips are not meant to replace the powerful GPUs that made Nvidia famous. They serve a different purpose entirely.
Every time an AI model answers a question, it goes through two stages. First it thinks, then it talks. Groq handles the talking part, known in the industry as the decode phase.
This is the moment when tokens, the chunks of text an AI produces, get sent back to you. It needs to happen fast, and that is exactly what Groq is built for.
"This isn't about replacing GPUs," a senior director at Nvidia said. "It's about using the right processor for the right part of the workload."
As AI chips race, grab the free Always Be Buying E-Book to build wealth steadily
Each Groq 3 LPX rack packs 256 chips and delivers 3,400 tokens per second, based on benchmarks from Artificial Analysis. That speed matters because it lets cloud providers charge more. "For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who demand the most latency-sensitive service agreements," the same director said.
Each chip has SRAM placed right beside the processing cores. That setup avoids the slower trip to main memory and cuts the lag that makes AI feel sluggish.
The Race Is On
Nvidia is not the only player chasing this speed. AMD is combining its systems with Cerebras chips, a company that recently went public, for the same kind of low-latency work. OpenAI's new "Ultrafast" mode, which runs on Cerebras hardware, advertises 750 tokens per second.
The split between training and inference has become a defining theme in the AI hardware market. Training-focused systems like Nvidia's Blackwell and Vera Rubin platforms are designed to build and refine models, while inference-focused chips like Groq optimize the speed at which finished models deliver answers. Owning both positions lets Nvidia serve customers across the full AI lifecycle rather than forcing every workload onto a general-purpose GPU.
There is also a manufacturing angle worth noting. The Groq chips come from Samsung, while Nvidia's own GPUs are made by TSMC. That gives Nvidia two separate supply chains, which is a useful hedge when demand is running hot.
Huang said he would split his own data center's compute between the two approaches, dedicating about a quarter of it to Groq for coding tasks. "The rest of my data center is all 100% Vera Rubin," he said.
The acquisition of Groq's technology gives Nvidia a way to cover both the training and inference sides of the AI market. That positioning could become more important as cloud providers look for ways to differentiate their AI services and charge different prices based on response speed.
What Wednesday's Earnings Could Show
On Wednesday, Nvidia will report earnings, with investors watching after Huang projected $1 trillion in cumulative Blackwell and Vera Rubin sales through 2027, a staggering target that shows just how much money the company expects to keep pulling in.
The Vera Rubin systems, which entered production earlier this year, are ramping up shipments. The Groq racks add a second revenue stream that targets a different kind of customer, one that cares less about raw training power and more about making AI feel instant.
What It Means for Investors
For your portfolio, this is about watching how Nvidia balances two businesses at once. The GPU side is the cash cow, but the Groq side could open up new pricing tiers in the AI industry. If the earnings call on Wednesday shows strong demand for both, that tells you the AI buildout is still in its early innings. If the company stumbles, it may be a sign that the market is getting ahead of actual spending.
While AI chips shape the future, the Always Be Buying E-Book helps you invest consistently
