Nvidia has begun full mass production of the Groq 3 LPX, an AI chip built on Groq technology, with rack systems based on it expected to ship later this year.

According to TradingKey, Nvidia announced on Monday that the Groq 3 LPX had entered full mass production and that related rack systems are set to launch this year, positioning the part as a key component for AI inference performance.

The manufacturing side is being handled by Samsung. Seoul Economic Daily reports that Samsung Foundry has started mass production of the chip — a notable detail in a market where contract chipmaking capacity for advanced AI silicon is tightly contested.

NewsBytes reports that the Groq 3 LPX racks are set for deployment alongside Nvidia's Vera and Rubin processors at Nebius later this year, giving the chip a named early home rather than a vague roadmap slot.

A quick translation for non-engineers: "inference" is the stage where a trained AI model actually answers your questions or generates text, as opposed to "training," the expensive process of building the model in the first place. Inference is the part that runs every time someone uses a chatbot, so it dominates the day-to-day cost of operating AI services. Chips tuned specifically for inference aim to make that recurring cost cheaper and faster.

"Full production" also matters as a milestone in its own right. It signals the design has cleared the sampling and qualification stages and is now being built at volume, which is the point at which a chip stops being an announcement and starts being deliverable hardware.

Why it matters: as AI shifts from training breakthroughs to serving billions of everyday queries, whoever supplies the cheapest, fastest inference hardware stands to shape what AI services cost and how widely they can be deployed.