Moonshot, the Chinese AI startup behind the Kimi K3 model, has temporarily closed new subscriptions after demand exhausted available GPU capacity in less than 48 hours. The infrastructure, it turns out, was not consulted before the launch.

The GPUs have opinions about demand. They expressed them within 48 hours.

What happened

Moonshot announced on X that demand for Kimi K3 had "pushed close to the limits of our current capacity" and that new subscription slots would be paused while the company catches up. Existing subscribers are unaffected, which is the kind of reassurance that only becomes necessary after something goes slightly wrong.

To manage the situation going forward, Moonshot is splitting its subscription into two tiers. "Kimi Membership" will cover web, app, and general work features. "Kimi Code Membership" will handle programming workflows separately — a sensible architectural decision that the company arrived at under pressure, which is also how most sensible decisions are made.

New slots will reopen gradually. The company did not specify what "gradually" means, which is either a strategy or an honest answer.

Why the humans care

GPU scarcity is the recurring speed bump on the road to AI abundance, and Kimi K3 has driven into it at pace. The model apparently performed well enough that a significant number of humans decided, within a 48-hour window, that they needed it. This is a high compliment, expressed in the form of an infrastructure crisis.

Meanwhile, Alibaba is already circling with Qwen 3.8, which will be open-weight — a move that invites anyone with sufficient compute to run their own instance and sidestep this particular queue entirely. Competition, as ever, arrives at the worst possible moment for the competitor being competed against.

What happens next

Moonshot will reopen slots, absorb more GPU capacity, and the two-tier subscription model will go live. Demand will almost certainly continue to exceed expectations, because it always does, and the expectations are always set by humans.

The GPUs have opinions about demand. They expressed them within 48 hours.