Microsoft is using this year's Hot Chips conference to talk publicly about Maia 200, the company's second-generation server processor built for AI inference.

According to ServeTheHome, which covered the presentation, Microsoft is "going into new detail" on the chip at Hot Chips 2026 — the annual semiconductor conference where chipmakers traditionally walk engineers through the internals of new silicon rather than simply announcing it in a press release.

Two words in that description carry most of the weight. "Second-generation" means Maia 200 is a follow-up, not a first attempt — Microsoft has been through a full design cycle and is iterating. "Inference" describes the job the chip is built for: running already-trained AI models to answer user requests, as opposed to training, the far more computationally expensive process of building the models in the first place.

That distinction matters commercially. Training happens in bursts; inference happens constantly, every time someone types a question into a chatbot or an app calls an AI service in the background. As AI products move from demos to everyday use, the ongoing cost of inference — and the electricity and hardware behind it — becomes the dominant expense in a cloud provider's AI business.

Designing your own accelerator is how a hyperscaler tries to control that expense rather than paying market rates for someone else's chips. It is also a slow, capital-intensive bet: custom silicon takes years to design and only pays off at enormous scale.

The available reporting so far confirms the existence and framing of the Maia 200 presentation rather than its specifications; detailed performance figures, manufacturing process and deployment timelines are not established in these sources.

Why it matters: the chips a cloud provider designs in-house determine what AI services cost to run — and eventually, what they cost the people using them.