AIAny
Icon for item

FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance

Uses diffusion-model forking moments as a proxy for perceptual distance to automatically generate pointwise reference-grounded labels, enabling annotation-free training of reference-based image quality assessment metrics.

Introduction

Most reference-based image quality assessment (IQA) methods depend on costly or noisy human labels (MOS or pairwise 2AFC). This paper's core insight is that a diffusion model's denoising trajectory reveals perceptual granularity: images that diverge earlier share only coarse structure (are perceptually distant), while those that fork late differ mainly in fine details. By treating the forking timestep as a pointwise distance label, the authors create a fully automated supervision signal for training IQA metrics.

Key Findings
  • Automated dataset: a 480k-pair dataset synthesized via FLUX.1-dev img2img forking (50 sampling steps, mixed real ImageNet crops and synthetic bases), where the sampled forking step provides a monotonic distance label. This yields dense, pointwise labels rather than noisy MOS or binary pairwise comparisons.
  • Human alignment and performance: empirical human studies show the forking-moment ordering aligns with human judgments; models trained on FoMo labels outperform several human-annotated baselines on multiple IQA benchmarks.
  • Training signal advantages: pointwise labels support global ranking objectives (e.g., RankNet-style) and richer gradients compared to 2AFC binary supervision, improving transitivity and calibration across arbitrary image pairs.
  • Practical outputs: multiple pretrained perceptual-distance models and the 480k dataset (webdataset shards) are released, enabling reproducible use as perceptual losses or evaluation metrics.
Who it's for + tradeoffs

Great fit if you develop or evaluate reference-based perceptual metrics, perceptual losses for image synthesis, or need large-scale, consistent distance labels without crowdsourcing. Look elsewhere if you require human-calibrated absolute MOS scales or if your application demands avoidance of biases introduced by a specific synthetic generator—FoMo's signal quality depends on the coverage and biases of the diffusion model used (FLUX.1-dev in the paper) and on compute to synthesize large datasets.

Where it fits

FoMo reframes dataset creation for IQA: it replaces expensive MOS collection and limited 2AFC comparisons with an automated, trajectory-grounded pointwise label. It complements existing perceptual metrics (LPIPS/DISTS) by providing an alternative supervised signal that scales without human annotation.

Method overview

Given a base image, forward-diffuse to a sampled timestep and re-denoise with branching re-entry s steps before the end to produce variants. The chosen re-entry (fork) timestep serves as a scalar distance: larger s → earlier fork → larger perceptual distance. These labeled pairs train reference-based distance models under rank-aware objectives.

Information

  • Websitearxiv.org
  • OrganizationsInterdisciplinary Program in AI, Seoul National University, Department of Electrical and Computer Engineering, Seoul National University, AIIS, ASRI, INMC, and ISRC, Seoul National University
  • AuthorsJaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
  • Published date2026/09/22

More Items

Learns joint predictive visual dynamics and action generation for generalist robot manipulation, translating future-relevant visual representations into actions. Integrates a Mixture-of-Transformers coupling a video expert and action expert, a frozen vision-language model for semantics, 4D distillation, and Causal Imprint; pretrained on a 20K+ hour heterogeneous corpus.

Trains decoders and diffusion generators to be robust to random subsets of pretrained encoder layers by regularizing layer-fusion during training, narrowing the reconstruction–generation gap in representation autoencoders and improving ImageNet-256 PSNR and gFID.

Introduces a rubric-based approach for video reward modeling that generates explicit, query-adaptive evaluation criteria before scoring to reduce scalar drift; proposes RGPO, a two-stage training (seed warm-up + joint optimization) that achieves strong pointwise and pairwise evaluation with high data efficiency.