AIAny
Icon for item

Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching

Predicts identity-preserving dense pixel correspondences between image pairs that violate spatio-temporal priors (e.g., edits and reference-guided generation). Fuses generative (FLUX2) and semantic (DINOv3) foundation representations with heterogeneous supervision and teacher-guided iterative refinement to generalize beyond classical optical-flow assumptions.

Introduction

Classical dense correspondence methods rely on smooth motion and rigid-geometry priors that break down when images preserve identity but change pose, appearance or scene composition — common in image editing and reference-guided generation. This work reframes correspondence as identity-preserving matching and shows how foundation-model features plus heterogeneous supervision enable robust pixel-level matches across such non-physical transformations.

Key Findings
  • Combines generative and semantic foundation representations (FLUX2 + DINOv3) so the model inherits both appearance synthesis knowledge and semantic invariances; this helps match regions that change texture or color but represent the same instance.
  • Uses heterogeneous supervision (classical datasets, tracked videos, synthetic scenes) with Huber loss to reduce annotation noise, improving generalization to diverse scenarios beyond optical flow or stereo.
  • Introduces teacher-guided iterative refinement on IEG data without dense labels to refine local correspondences and covisibility prediction, enabling improved identity preservation metrics on edited/generated image pairs.
  • A single FreeMatching model remains competitive on classical benchmarks while substantially improving correspondence quality for image-editing and reference-guided generation tasks, making it useful both as a matcher and as an identity-preservation metric.
Who it's for and trade-offs

Great fit if you need robust pixel-level correspondences across large appearance changes or edited/generated images (e.g., reference-guided image generation evaluation, edit consistency checks). Look elsewhere if your application is strictly real-time low-latency optical flow on constrained hardware or you require methods tailored only to smooth motion priors; FreeMatching trades some classical-flow-specific optimizations for broader generalization.

Where It Fits

Positions between classical optical-flow/dense-matching methods and perceptual evaluation tools: it aims to bridge matching for non-physical, identity-preserving transformations and can serve as a quantitative metric for identity preservation in IEG workflows.

Method Snapshot

Architecture initializes from a generative transformer backbone and injects frozen DINOv3 patch features; paired-token joint attention predicts backward dense maps and covisibility masks. Training alternates supervised learning on heterogeneous labeled sources with teacher-guided refinement on unlabeled IEG pairs.

Information

  • Websitearxiv.org
  • OrganizationsAffiliation: The University of Hong Kong, Affiliation: ByteDance Seed, Affiliation: Zhejiang University
  • AuthorsLuping Liu, Bingyi Kang, Yifan Wang, Dong Xu
  • Published date2026/10/08

More Items

Converts static 3D Gaussian Splatting scenes into endlessly looping 3D cinemagraphs by inferring plausible dynamics with a vision-language model, synthesizing a reference video, lifting it to multi-view videos, and fitting a Fourier-parameterized Periodic Deformation Field with a Grounded Drift Field—mask-free capture of deformation, object motion, and illumination changes.

Generates synchronized egocentric video streams for multiple interacting agents in a shared environment, enforcing cross-view action consistency, shared environment memory, and consistent propagation of interaction-induced state changes — aimed at embodied AI, VR, and multi-agent vision research.

Lets a pretrained multimodal LLM interpret navigation requests and orchestrate motion via tool calls for generalist robot navigation across unfamiliar scenes. Key features: an agent harness with Navigation Skills, a unified visual-point interface, task-progress tracking, and tool-based motion execution without navigation-specific fine-tuning.