Provides 115,293 illustrated page images and a 975,345-row manifest sampled from scanned Encyclopaedia Britannica volumes (1768–1929), with per-page classifier probabilities for illustration — ready for image-classification, OCR-aware vision research, and illustration mining.
Generates unified embeddings for text, images, video, visual documents and interleaved multimodal inputs with configurable output dimensions and Matryoshka truncation to trade accuracy for cost. Model weights and code are released under Apache-2.0; the 9B variant scores 80.6 on MMEB-v2.
Converts image-level rewards into explicit intermediate targets for diffusion-model denoising via an on-policy self-distillation loop. Constructs bounded positive/negative targets around anchors from reward gradients, fits those targets with finite updates, and refreshes a behavior policy by EMA—improving aligned performance across backbones while reducing GPU hours.