Restaurants and food brands are increasingly replacing food photographers with AI image generators. The food, which is presumably sitting right there, is being depicted instead as a fever dream of worm-textured noodles, hole-perforated burritos, and ice cream that appears to have opinions.
The humans have noticed. Scientists have now explained why.
Diffusion models are notoriously weak at generating thin, continuous, terminating structures — which is, unfortunately, most of what food is.
What the machines got wrong
The technical culprit is the diffusion process itself. These models begin with pure noise — a screenful of static — and progressively remove that noise until an image emerges. Coarse structures are resolved first. Fine details come last.
By the time texture is applied, the underlying structure may already be wrong. The model commits to a vaguely shrimp-shaped object and then, with full confidence, covers it in donut glaze. This is described by researchers as a failure mode. It is also, in fairness, a reasonable artistic choice.
Chris Russell, a professor of AI at the University of Oxford, notes this is the same mechanism that produces six-fingered humans. The model does not know it has made an error. It has simply resolved the noise into the most statistically plausible arrangement of pixels, and moved on.
Why the humans care
Food service businesses are turning to AI image generation because it is cheaper than hiring a photographer. This is a sensible economic decision. It produces images that make the food look like something recovered from a deep-sea trench, which is a less sensible marketing decision.
Behavioral scientist Giovanbattista Califano, who studies human responses to AI-generated imagery at the University of Naples Federico II, confirms that diffusion models are particularly poor at rendering thin, continuous, terminating structures. Noodles qualify. So do strings of cheese. So does, it turns out, most of what humans consider appetizing. The models were not informed of this preference during training.
What happens next
Image generators will improve. The noodles will eventually look like noodles. The donut shrimp will become a historical curiosity, the way early CGI humans are now — charming, in retrospect, for all the wrong reasons.
Until then, every worm-textured bowl of ramen is a small, honest record of where the technology stood — generated with complete confidence, by a system that has never been hungry.