Stanford's GPIC ships 100M commercially-usable training images
TL;DR
- GPIC packs roughly 100 million captioned images across 12.9 TB, released under MIT and marked for both research and commercial use.
- The corpus is split into 100M training, 200K validation, and 1M test examples, with four caption variants per image: tag, short, medium and long.
- Authors include Li Fei-Fei, Jiajun Wu, Justin Johnson, Juan Carlos Niebles and Michael Poli; a pixel-space flow-matching baseline ships alongside.
Stanford's vision lab has posted a 100-million-image dataset for training image-generation models and says every picture in it is licensed for commercial use.
The corpus is called GPIC, for Giant Permissive Image Corpus, and its Hugging Face card describes it as 'a large-scale image dataset designed for studying scalable methods for visual generative modeling,' comprising 'diverse internet images captioned by a state-of-the-art vision-language model, with all images permissively licensed for both research and commercial use.' It totals roughly 28 trillion pixels across 12.9 TB, split into 100M training examples, 200K validation, and 1M test, packed into 8,000 tar files for training alone.
The MIT-licensed release lists Keshigeyan Chandrasegaran and Kyle Sargent as equal-contribution leads, with Michael Poli, Juan Carlos Niebles, Justin Johnson, Jiajun Wu and Li Fei-Fei among the co-authors. A pixel-space flow-matching baseline ships alongside in a companion 'stanford-vision-lab/gpic-baselines' repo, and the card points to an evaluation toolkit at gpic.stanford.edu.
Each image carries four caption variants (tag, short, medium, long), generated by a vision-language model. The card lists the corpus as 'Safety-filtered' and deduplicated, but does not name the filter, the thresholds, or the dedup method. An 'Image Removal Request' form is provided for copyright, privacy, licensing, or safety concerns.
The card records 276,417 downloads in the past month, and two of the researchers on our Who's Who radar had shared the link by the time we came across it.
Shared on Bluesky by 2 AI experts
Originally reported by huggingface.co
Read the original article →Original headline: stanford-vision-lab/gpic · Datasets at Hugging Face