Nvidia has begun production of its Groq 3 LPX rack systems, the first hardware to emerge from its roughly $20 billion acquisition of AI chip startup Groq.
According to Anadolu Ajansı, the racks are built specifically to accelerate AI inference — the work a model does when it answers you, as opposed to the training phase where it learns. The agency reports that the first systems will become operational at cloud provider Nebius later this year.
StorageReview.com reports that each rack packs 256 LP30s and that the platform is now in full production, citing throughput of 3,400 tokens per second at a 100,000-token context. Engineering.com likewise frames the announcement around inference workloads.
Those numbers come from an independent benchmark by Artificial Analysis, according to The Register, which measured Nvidia's LPX rack systems producing 3,400 tokens a second with a 100,000-token input sequence running Google's Gemma 4 31B. The Register's framing is deliberately cautious — its headline asks what the first Groq 3 LPU benchmarks do and don't tell us about the $20 billion gamble, a reminder that a single benchmark on a single model is not the same as proven real-world performance across the messy variety of production workloads.
The context worth holding onto: long context windows are expensive. Feeding a model 100,000 tokens before it writes a word is exactly the scenario that strains conventional hardware, so it is a pointed choice of test.
This matters because inference — not training — is where the everyday cost of running AI products lives, and if Nvidia's Groq-derived racks deliver on speed at long context, they could reshape what AI services cost to operate at scale.