AIAny
Icon for item

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

Generates unified embeddings for text, images, video, visual documents and interleaved multimodal inputs with configurable output dimensions and Matryoshka truncation to trade accuracy for cost. Model weights and code are released under Apache-2.0; the 9B variant scores 80.6 on MMEB-v2.

Introduction

Multimodal search and recommendation increasingly require a single embedding space that preserves fine-grained semantics across text, images, video and document-like visuals while remaining efficient at billion-scale indexing. The core insight: training a family of scalable encoders (2B/4B/9B) with a two‑stage curriculum yields embeddings that are both serviceable in production and competitive on open benchmarks.

Key Findings
  • Two-stage training: large-scale multimodal alignment followed by a refinement phase using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer — this combination improves discrimination without sacrificing broad coverage. This means embeddings better match hard retrieval cases while retaining generality.
  • Matryoshka embeddings and configurable dimensions: models expose multiple output sizes (e.g., 64–4096) so teams can truncate to lower dims (e.g., 256) to cut storage/compute with modest accuracy loss. This enables practical deployment trade-offs for latency- and cost-sensitive systems.
  • Empirical results: across public benchmarks the family leads or matches state-of-the-art; the 9B variant achieves an overall MMEB-v2 score of 80.6 and produces 4096‑D L2‑normalized vectors. WeChat-internal evaluations and multiple online A/B tests show production gains in recommendation and search.
  • Engineering posture: weights and code are released under Apache-2.0, calling out support for common inference stacks (Transformers/SentenceTransformers, vLLM) and Hugging Face hosting, which lowers integration friction for practitioners.
Who it fits and trade-offs

Great fit if you need a unified, production-ready multimodal embedding backbone that supports cross-modal retrieval and can be dimension‑truncated to manage cost. Prefer it when you want an off-the-shelf family with released weights, benchmarked performance, and documented deployment recipes.

Look elsewhere if you require audio input support (audio is not supported) or if you need a custom task-specific embedding trained on proprietary domain data without fine-tuning; extremely constrained-device scenarios might still favor specialized lightweight encoders despite Matryoshka truncation.

Information

  • Websitearxiv.org
  • OrganizationsWeChat Vision (Tencent), Tencent Inc.
  • AuthorsJunjie Zhou, Ke Mei, Lei Li, Tianyi Wang, Fengyun Rao, Jing Lyu
  • Published date2026/08/25

More Items

Predicts identity-preserving dense pixel correspondences between image pairs that violate spatio-temporal priors (e.g., edits and reference-guided generation). Fuses generative (FLUX2) and semantic (DINOv3) foundation representations with heterogeneous supervision and teacher-guided iterative refinement to generalize beyond classical optical-flow assumptions.

Converts static 3D Gaussian Splatting scenes into endlessly looping 3D cinemagraphs by inferring plausible dynamics with a vision-language model, synthesizing a reference video, lifting it to multi-view videos, and fitting a Fourier-parameterized Periodic Deformation Field with a Grounded Drift Field—mask-free capture of deformation, object motion, and illumination changes.

A concise textbook-style book that explains foundational concepts and techniques for large language models, covering pre-training, generative models, prompting, alignment, inference, and reasoning. Structured as self-contained chapters for readers with some ML/NLP background or those seeking a principled introduction to LLM foundations.