SperidLabs Scales Iris-3B to 3B, Finds No Downstream Edge
TL;DR
- SperidLabs' Iris-3B, a 3B-parameter pixel-space text-to-image transformer, scored GenEval 0.798 and DPG 86.52 at 1024² and matched Qwen-Image on OneIG-EN.
- Fine-tuned for monocular depth on five benchmarks and for 4× DIV2K restoration, neither pixel-space model improved on the latent FLUX.2 Klein baseline.
- SperidLabs released weights and training code, documenting a patch-grid artifact in pixel-space depth outputs and confounds in the single-run comparisons.
A 3B-parameter image model trained without a VAE matched Qwen-Image on one standard benchmark, and then failed to deliver the downstream advantage its design was supposed to unlock. In a technical report posted on Hugging Face, SperidLabs' Chema Garabito reports that Iris-3B, a pixel-space text-to-image transformer pretrained from scratch through a 256→512→1024 curriculum, "reaches GenEval 0.798 and DPG 86.52" at 1024² and lands "on OneIG-EN its overall score matches Qwen-Image" under the official evaluators.
The theory was that generating directly in pixels, rather than inside a lossy autoencoder, should pay off most on tasks where fine detail matters. The report tests that on two: monocular depth estimation across NYUv2, KITTI, ETH3D, ScanNet and DIODE, and 4× DIV2K restoration. The result is flat. "Neither pixel-space model clearly improves on the latent one," the authors write of the depth comparison, with Iris-3B level with the latent FLUX.2 Klein and a converted-to-pixel FLUX.2 Klein trailing it. On restoration, the converted pixel model "is worse on most metrics," most clearly SSIM and NIQE.
The authors are blunt about the limits of the test. Each arm is a single run at a short budget with no significance test. Iris-3B has no latent twin at matched size and data, so its parity reading mixes representation space with scale and compute. The converted parent likely carries a weaker prior than the fully pretrained latent model it was adapted from. There is also a patch-grid artifact in both pixel-space models' raw depth outputs: "The latent model has no grid, and both pixel-space models do, Iris-3B most strongly." The VAE decoder smooths that away for free.
SperidLabs is releasing the weights and training code anyway, framing Iris-3B as "a strong starting point for further work on pixel-space generation." This lands in a busy week of open-weights releases on our tracker. The report closes by naming editing and reference-based generation, where latent compression should hurt most, as the next real test of the hypothesis.
Originally reported by huggingface.co
Read the original article →Original headline: SperidLabs' Iris-3B Scales Pixel-Space Diffusion to 3B Params, Matches Qwen-Image on OneIG at 1024²