huggingface.co web signal

UltraText Bench: GPT Image 2 Scores 99.35 as Qwen-Image-2512 Collapses From 86.5 to 42.9 at Max Text Density

Summary

UltraText Bench evaluates text-to-image models on 432 bilingual prompts across 24 real-world scenes and 407,918 ground-truth characters, scoring text fidelity, clarity, spatial layout and scene integration. GPT Image 2 (Low) tops the board with a 99.35 composite and perfect scores on 83.6% of images; Qwen-Image-2512's English composite drops from 86.50 at Level 1 difficulty to 42.86 at Level 3, and 'Turbo' variants trade substantial fidelity for marginal clarity gains.