Skip to content
Skip to content

Independent e-magazine

the OUTSPOKEN digest

Nvidia Is Putting AI Server Prices Up by More Than Fifteen Per Cent, and the Chips Are Not the Reason

Bloomberg reports that customers have been warned of rises above fifteen per cent on Vera Rubin and Grace Blackwell systems shipping early next year. The cause is memory, and three companies control the supply of it.

Outspoken Digest Technology Desk

Tuesday, August 25, 2026/3 min read

An aisle between rows of server racks in a data centre
Photo: PiDatacenters via Wikimedia Commons (CC BY-SA 4.0)

On 22 August, Bloomberg reported that Nvidia has been telling its largest customers to expect price increases of more than fifteen per cent, in many cases, on servers built around its Vera Rubin and Grace Blackwell chips. The increases apply to systems shipping early next year, and the exact figure varies by chip generation and by how much memory the configuration carries.

Nvidia did not respond to requests for comment. Take the number as reported rather than confirmed.

What is actually going up?

Not the accelerator. The machine around it.

These are complete servers, and the reported cause is the cost of memory rather than of Nvidia's own silicon. That distinction is the whole story, because it says the bottleneck in AI infrastructure has moved.

For three years the scarce thing was the GPU. Nvidia's order book was the constraint, its pricing power was the consequence, and every analysis of the industry was an analysis of how many accelerators could be produced. The scarce thing now is the memory that sits beside them.

Why memory, and why now?

Because an AI accelerator is useless without an enormous amount of very fast memory attached, and only three companies make it.

Samsung, SK Hynix and Micron supply the DRAM market between them, and they cannot meet the demand the AI build out has created. When three suppliers cannot satisfy demand for a component nobody can substitute, the price is set by them rather than by their customers. That is not a failure of the market. It is what the market does under those conditions.

The notifications have not come from Nvidia alone. Contract manufacturers who assemble servers for the large data centre operators, Microsoft, Google and Oracle among them, have been passing the warning down to their own customers.

Who pays this in the end?

The chain is short and it ends at people who have never bought a server.

A hyperscaler absorbs a fifteen per cent rise in hardware cost across a fleet it is expanding, then prices its cloud accordingly. The companies renting that capacity build products on top of it. The model providers pass it into the per token price of an interface. Somewhere at the end of that chain is a subscription, and the arithmetic that made it affordable was written when servers were cheaper.

None of that happens this quarter. These are systems shipping in early 2027, so this is a forward cost being signalled a year and a half ahead, which is itself worth noticing. Suppliers do not warn customers about price rises eighteen months out unless the customers are making commitments that far ahead.

The uncomfortable part for the AI trade

Memory is a commodity cycle business, and commodity cycles have a shape.

High prices attract capacity. Capacity takes two to three years to build and then arrives all at once, usually into demand that has cooled. The DRAM industry has run this cycle repeatedly and violently, and the current shortage will end the same way every previous one did, with a glut.

That should temper two conclusions people are drawing this week. Memory makers are not permanently repriced, however good the next four quarters look. And the cost pressure on AI infrastructure is not a permanent feature of the industry either, which matters for anyone modelling the economics of inference on the assumption that today's input costs are the floor.

What it does illustrate is how quickly the constraint moves. Three years ago the question was whether anyone could get GPUs. It became whether anyone could get the power to run them, which is why financing structures started to look the way we described in GPU backed financing for data centres. Now it is memory. Each time, the industry treats the current bottleneck as the permanent one.

What to watch

Whether Nvidia says anything, and what happens to the cheap end of the market.

Silence is informative here. A company that intended to deny a fifteen per cent price rise would usually do so quickly, and Nvidia's non response has been read accordingly.

The more interesting signal is on the other side of the market. When infrastructure costs rise, the pressure to run smaller and cheaper models rises with it, and that pressure has produced the most disruptive moments of the past two years, including the one we wrote about in the cheap model shock. A hardware price rise is not only a cost. It is an incentive pointed directly at whoever can do the same job with less of it.

Published in The Outspoken Digest

Editorial desk

Outspoken Digest Technology Desk

Software, hardware, artificial intelligence and what they change for everyone else.

Newsletter

The Digest, in your inbox

One edition, sent when it is ready. No noise, and your address is never passed on.

We send a confirmation first. One click to leave, always.

Share this story

the OUTSPOKEN digest

Beyond boundaries. Independent stories on technology, culture, and the trends shaping how we live.