Nvidia used Hot Chips 2026, the long-running conference where chipmakers walk engineers through their newest silicon, to detail BlueField-4 — the latest generation of its data processing unit.
According to ServeTheHome, which covered the session, Nvidia laid out "huge upgrades" in BlueField-4 over the previous generation, along with the reasoning behind those changes and the company's "scale-in" networking approach. The outlet did not publish specific performance figures in the summary available, and Nvidia's own detailed specifications sit with the conference presentation itself.
Some context on what a DPU is, since it is the least familiar of Nvidia's product lines to general readers. A CPU runs applications and a GPU crunches the math behind AI models. A DPU sits on the network card and handles the unglamorous work in between: moving data between servers, storage access, encryption, and isolating one customer's workload from another's. Offloading that housekeeping frees the expensive processors to do the work customers are actually paying for.
That division of labor is why the part matters more than its profile suggests. Modern AI training runs stitch thousands of GPUs into what behaves like a single machine, and at that scale the bottleneck often is not the GPUs but the plumbing connecting them. Nvidia's emphasis on "scale-in" networking, per ServeTheHome's framing, points at exactly that problem.
Why it matters: the economics of AI data centers increasingly hinge on how efficiently data moves between chips, not just on how fast the chips themselves run.