Most reference-based image quality assessment (IQA) methods depend on costly or noisy human labels (MOS or pairwise 2AFC). This paper's core insight is that a diffusion model's denoising trajectory reveals perceptual granularity: images that diverge earlier share only coarse structure (are perceptually distant), while those that fork late differ mainly in fine details. By treating the forking timestep as a pointwise distance label, the authors create a fully automated supervision signal for training IQA metrics.
Key Findings
- Automated dataset: a 480k-pair dataset synthesized via FLUX.1-dev img2img forking (50 sampling steps, mixed real ImageNet crops and synthetic bases), where the sampled forking step provides a monotonic distance label. This yields dense, pointwise labels rather than noisy MOS or binary pairwise comparisons.
- Human alignment and performance: empirical human studies show the forking-moment ordering aligns with human judgments; models trained on FoMo labels outperform several human-annotated baselines on multiple IQA benchmarks.
- Training signal advantages: pointwise labels support global ranking objectives (e.g., RankNet-style) and richer gradients compared to 2AFC binary supervision, improving transitivity and calibration across arbitrary image pairs.
- Practical outputs: multiple pretrained perceptual-distance models and the 480k dataset (webdataset shards) are released, enabling reproducible use as perceptual losses or evaluation metrics.
Who it's for + tradeoffs
Great fit if you develop or evaluate reference-based perceptual metrics, perceptual losses for image synthesis, or need large-scale, consistent distance labels without crowdsourcing. Look elsewhere if you require human-calibrated absolute MOS scales or if your application demands avoidance of biases introduced by a specific synthetic generator—FoMo's signal quality depends on the coverage and biases of the diffusion model used (FLUX.1-dev in the paper) and on compute to synthesize large datasets.
Where it fits
FoMo reframes dataset creation for IQA: it replaces expensive MOS collection and limited 2AFC comparisons with an automated, trajectory-grounded pointwise label. It complements existing perceptual metrics (LPIPS/DISTS) by providing an alternative supervised signal that scales without human annotation.
Method overview
Given a base image, forward-diffuse to a sampled timestep and re-denoise with branching re-entry s steps before the end to produce variants. The chosen re-entry (fork) timestep serves as a scalar distance: larger s → earlier fork → larger perceptual distance. These labeled pairs train reference-based distance models under rank-aware objectives.