theverge.com web signal

Why AI-generated food images look so wrong, per researchers

TL;DR

  • Diffusion image models build a picture coarse-shape first and fine texture last, so a wrong shape gets amplified when detail is painted onto it.
  • Researchers say the models are especially weak at thin, continuous, terminating structures like noodles and at repeating textures like seeds or bubbles.
  • The systems reproduce the look of food without any grasp of what food is, and human disgust cues around parasites and rot make the errors especially jarring.

AI-generated food images have a distinct look: noodles that never end, clustered holes that suggest infestation, textures better suited to bricks than to bread. The Verge put the question to five researchers, and their answers converge on how diffusion models actually assemble a picture.

Chris Russell, a professor of AI at the University of Oxford, said the sequence explains the mistakes: "This means that initially coarse structures are recovered first with fine texture details" arriving at the end. If the basic shape is wrong, the crisp texture layered on top only makes the error more vivid, the same reason models keep producing six-fingered hands.

Certain foods trip the models harder than others. Giovanbattista Califano, a behavioral scientist at the University of Naples Federico II, told The Verge that "diffusion models are notoriously weak at generating thin, continuous, terminating structures." Noodles wander off the plate. Repeating textures like seeds and bubbles spread past their edges.

The deeper problem is conceptual. Roland Meyer, a professor of digital cultures at the University of Zurich, put it flatly: "AI image generation reproduces looks without proper knowledge about the world." Michael Cook, a senior lecturer in computer science at King's College London, added that the same rendering "might look totally normal if used in an architectural context," bark, foam, tile, and only becomes revolting when the caption says it is dinner. Cook also noted that "because we know so little about the training processes of these systems, we don't really know what mix of content they're receiving, or what associations they're making."

Human disgust, evolved to flag parasites and rot, does the rest.

Shared on Bluesky by 2 AI experts