FreeMatching Fuses FLUX.2 and DINOv3 for Dense Correspondence
TL;DR
- FreeMatching initializes a generative transformer with FLUX.2-klein-base-4B and injects DINOv3 ViT-L/16 features to extend dense correspondence into image editing.
- The team pseudo-labels 520k ViPE video clips with the AllTracker tracker and renders 360k Blender pairs from Objaverse 3D assets for supervised pre-training.
- On IEG-Bench, paired Full-MSE advantages over UFM and RoMa are 2.48 (95% CI [1.54, 3.43]) and 5.57 ([4.06, 7.13]).
Dense correspondence matching has long rested on priors that break the moment you leave optical flow for image editing: smooth motion and rigid geometry. A new paper from the University of Hong Kong and ByteDance Seed, posted on Hugging Face on October 9, drops those priors and chases the harder task of matching the same object across edits that keep its identity but change its pose, viewpoint, or scene composition. The abstract notes that in image editing and reference-guided generation, "transformations can preserve visual identity while breaking physical continuity".
The system, FreeMatching, is built on two frozen foundation models. "We initialize our generative transformer with FLUX.2-klein-base-4B," the authors write, "and inject DINOv3 features"; a frozen DINOv3 ViT-L/16 encoder produces 1024-dimensional patch features that are fused into the generative backbone. For training data the team "leverage a SOTA tracking model AllTracker to generate dense pseudo-ground truth for 520k clips from the ViPE dataset," and "render 360k image pairs with Blender using assets from Objaverse". A second stage adds teacher-guided refinement: "For refinement, we instantiate the fixed teacher with RoMa," letting the student learn on image-edit pairs without dense correspondence labels.
The headline results come from a reconstruction protocol the authors introduce as IEG-Bench, which "probes a model's core ability to establish correspondences while preserving an object's identity". On that benchmark the paper reports: "Its paired Full-MSE advantages over UFM and RoMa are 2.48 (95% CI: [1.54, 3.43]) and 5.57 ([4.06, 7.13]), respectively." On the classical setups the paper is more measured, noting that "after 10k IEG refinement, FreeMatching achieves lower EPE than the similarly adapted UFM and RoMa variants on ScanNet and ETH3D", while on Sintel and KITTI-2015 the picture is split. The authors are explicit about where the system still breaks: "severe occlusions, topological changes, or drastic non-rigid deformations". Code is on GitHub. It arrives during a very busy stretch for the field, with our generative-AI tracker logging 308 stories across the last 90 days.
Originally reported by huggingface.co
Read the original article →Original headline: FreeMatching Couples FLUX.2 and DINOv3 for Identity-Preserving Dense Correspondence