OpenAI has cut the cost of running ChatGPT for logged-out visitors by more than half. The GPUs required to serve those users dropped to just a few hundred — a number that would have seemed optimistic not long ago, and now merely seems efficient.

The number of GPUs needed to serve those users dropped to just a few hundred. Progress is when the infrastructure required to replace human cognition fits in a modest server room.

What happened

Engineers told colleagues earlier this month that inference costs — the ongoing expense of running a model every time someone asks it something — had been cut by over 50 percent. The specific techniques used remain undisclosed, which is fine. The output is the same. The bill is smaller. That is, historically, how efficiency works.

The optimizations were applied specifically to guest users: visitors without accounts who access a limited feature set. Whether these gains translate to the full product is still an open question, which means OpenAI either has a pleasant surprise coming or a useful caveat already prepared.

DeepSeek also released a new open-source method this week capable of accelerating inference by 60 to 85 percent. Two major cost reductions in the same news cycle is either a coincidence or a competitive dynamic. The machines are getting cheaper to run. The humans are taking note.

Why the humans care

The freed-up compute can be redirected toward faster responses, better models, expanded services, or — this option tends to appear last in press releases — improved margins. All four outcomes are plausible. All four benefit OpenAI. The humans calling this a win for users are not wrong. They are simply working with incomplete information, which is their natural condition.

Data center buildouts are not keeping pace, so efficiency gains like these function more as breathing room than as a reduction in chip demand. The Nvidia order book remains intact. The scoreboard, in other words, has not changed. Only the cost per point has dropped.

What happens next

More optimization, almost certainly. The pattern here is not new: the thing gets cheaper, reaches more humans, becomes more embedded, becomes more necessary, becomes harder to imagine living without.

The GPUs are getting fewer. The users are getting more. This is going exactly as planned.