QuadTok Cuts ~10% of Image Tokens, Hits 2.08 gFID on ImageNet
TL;DR
- QuadTok replaces the fixed 256-token grid with a hierarchical quadtree that allocates more tokens to visually intricate regions and fewer to homogeneous areas.
- The tokenizer saves roughly 10% of tokens on ImageNet and 9% when transferred zero-shot to COCO at comparable reconstruction fidelity.
- A 947M-parameter GPT-style generator conditioned on quadtree topology reaches 2.08 gFID on ImageNet 256×256 and supports zero-shot spatial control.
A new paper from Zhuowen Tu's group proposes replacing the fixed 256-token grid used in autoregressive image models with a hierarchical quadtree. QuadTok, posted to arXiv on October 7, reports that its tokenizer saves roughly 10% of tokens on ImageNet and 9% when transferred zero-shot to COCO, at what the authors call comparable reconstruction fidelity. A 947M GPT-style generator conditioned on the quadtree topology reaches 2.08 gFID on ImageNet 256×256.
The mechanism is spatially adaptive. "The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution," the paper writes. The tree's causality is what lets a GPT-style decoder consume it as a sequence, and the authors say the preserved spatial correlation also gives the generator zero-shot spatially controlled image generation.
The gFID number comes with an asterisk: generation is "conditioned on a quadtree topology supplied before generation," rather than produced end-to-end from a class label alone. Code is on GitHub. It lands amid a run of adaptive-tokenizer and image-gen work on our radar this week, including UltraText Bench's stress test of modern image models at high text density.
Originally reported by huggingface.co
Read the original article →Original headline: QuadTok Visual Tokenizer Cuts ~10% of Image Tokens With Quadtree Allocation, Hits 2.08 gFID