huggingface.co web signal

QuadTok Cuts ~10% of Image Tokens, Hits 2.08 gFID on ImageNet

TL;DR

  • QuadTok replaces the fixed 256-token grid with a hierarchical quadtree that allocates more tokens to visually intricate regions and fewer to homogeneous areas.
  • The tokenizer saves roughly 10% of tokens on ImageNet and 9% when transferred zero-shot to COCO at comparable reconstruction fidelity.
  • A 947M-parameter GPT-style generator conditioned on quadtree topology reaches 2.08 gFID on ImageNet 256×256 and supports zero-shot spatial control.

A new paper from Zhuowen Tu's group proposes replacing the fixed 256-token grid used in autoregressive image models with a hierarchical quadtree. QuadTok, posted to arXiv on October 7, reports that its tokenizer saves roughly 10% of tokens on ImageNet and 9% when transferred zero-shot to COCO, at what the authors call comparable reconstruction fidelity. A 947M GPT-style generator conditioned on the quadtree topology reaches 2.08 gFID on ImageNet 256×256.

The mechanism is spatially adaptive. "The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution," the paper writes. The tree's causality is what lets a GPT-style decoder consume it as a sequence, and the authors say the preserved spatial correlation also gives the generator zero-shot spatially controlled image generation.

The gFID number comes with an asterisk: generation is "conditioned on a quadtree topology supplied before generation," rather than produced end-to-end from a class label alone. Code is on GitHub. It lands amid a run of adaptive-tokenizer and image-gen work on our radar this week, including UltraText Bench's stress test of modern image models at high text density.