Nvidia appears to be attacking the same problem from two directions at once: there may not be enough high-bandwidth memory to go around.

According to a report from The Information, surfaced by Techmeme, Nvidia is weighing what the publication calls a radical step — shipping versions of its upcoming Rubin Ultra GPU with less memory than planned. The reason, per The Information's sources, is potential trouble securing enough advanced high-bandwidth memory, or HBM, the specialized chips that sit alongside an AI processor and feed it data. The company has reportedly tested at least three versions of the chip.

That matters because memory, not raw computing power, is often what limits how large a model a single accelerator can hold and how fast it can run. Cutting memory is not a decision a chipmaker takes lightly.

At the same time, Nvidia is pushing on the other end of the pipeline. Sdxcentral reports the company wants to bypass the CPU with an open-source overhaul of how AI systems talk to storage. Fierce Network frames the same effort more bluntly, describing it as turning storage into memory relief — using faster, more direct access to drives to take pressure off scarce, expensive memory.

Taken together, the reports sketch a company hedging: design around the shortage in hardware, and engineer around it in software.

Why it matters: HBM supply is now a real constraint on the AI buildout, and how Nvidia balances memory against storage will shape the cost, speed, and availability of the chips nearly every AI company depends on.