Converts natural-language SVG instructions into instance-aware six-item rubrics and uses a vision–language judge to score rendered SVGs item-by-item as dense RL rewards. Produces stronger human alignment and lifts benchmark rubric scores substantially on MMSVG datasets.
An all-in-one multilingual scene text recognition approach that pairs a shared visual encoder with a script-aware Mixture-of-Experts (ScriptMoE) decoder to route each image to top-2 script experts. Introduces TextMuSS-10M, a synthetic dataset covering 10 scripts and 229 languages, and reports state-of-the-art accuracy and large end-to-end OCR F1 gains while remaining parameter-efficient.
Provides a unified data and evaluation suite for measuring gameplay capabilities across short, medium, and long temporal horizons. Includes an automated annotator, a 5,000-hour AAA gameplay dataset with temporally aligned frames, actions and multi-horizon instructions, and a reproducible offline + stepwise online benchmark.
Stores speaker-labeled verbatim messages alongside derived person- and group-level states to improve attribution and long-term memory in multi-party dialogues. Proposes a dual-track memory and a speaker-conditioned Writer trained with SpeakerLevenshtein and GRPO; yields consistent gains on multiple dialogue memory benchmarks.