Analyzes how on-policy distillation (OPD) transfers teacher LLM capabilities to student models across in-domain shifts, cross-domain transfer, and multi-teacher settings. Key findings: OPD conveys reasoning patterns rather than specific answers, same-origin teacher-student pairs generalize broadly, and multi-teacher combinations induce mixture-dependent tradeoffs.
Automates evaluation of visual world models via a hierarchical agent pipeline that decomposes each case, spawns specialized sub-agents to collect diagnostic evidence, and outputs a verifiable evidence tree plus a final verdict; validated on 18 models across 330 cases and released as a live evaluation pipeline.
Evaluates whether AI systems can independently carry out project-level scientific research by progressively removing human methodological guidance across 60 tasks in 11 domains. Built with expert review, sandbox execution, and multi-agent–model scoring to measure innovation and autonomous experimental execution.
Uses cooperative multi-agent RL where multiple decoupled models provide peer-derived pseudo-rewards to each other, enabling unsupervised improvements in reasoning; increases cohort diversity to reduce correlated errors and avoid training collapse, showing consistent gains across text and multimodal benchmarks.
Turns embodied navigation into 2D visual prompting where a vision-language model selects image pixels that are projected to 3D actions; adds selective chain-of-thought, compressed anchor-trajectory memory, and a two-level alignment objective to improve sample and runtime efficiency.
Introduces SemComp-Bench: a benchmark and VLM-based evaluation protocol for measuring outcome achievement and task-relevant semantic grounding in instruction-driven video generation. Ships with SemComp-Data, curated image–instruction–outcome triplets and OA/GR scoring.
An uncensored, weight-modified variant of Qwen3.8-27B that surgically removes the model's refusal directions to produce 0% refusals while aiming to preserve or improve capability. Uses complementary abliteration blending (SVD + LEACE blend) and ships with recommended greedy inference settings; intended for AI-safety research and red‑teaming, not for causing harm.
Proposes FACET, a framework that synthesizes verifiable terminal tasks by reconstructing scenario intent and grounding instruction, solution, and verifier in a shared executable container state. Key features include environment-first generation, execution-based validation, and targeted repair to preserve source intent and cross-artifact consistency.
Wraps static, hand-built environments with a programmable plug-in harness that reshapes environment behavior without changing underlying logic. EnvRigger automates diagnosis and synthesis of harness components from agent failure trajectories, validating edits via fresh rollouts to improve agent success and efficiency.
Evaluates visual reasoning in video generation models using 27 photorealistic tasks (810 instances), a two-level taxonomy of domains and skill tags, and task designs that enforce valid intermediate trajectories and calibrated difficulty.
Turns an uncalibrated monocular actor video into multiview-consistent novel-view videos and lifts them into 4D Gaussian Splatting assets. Introduces Reference Context Packing to keep reference conditioning fixed-size and Target Context Routing to exchange context across target groups, improving large-view reconstruction consistency.
Generates group images that bind up to ten reference identities to distinct people and locations by predicting an explicit identity–layout plan and supervising faces with Layout-Grounded ID Loss. Improves identity fidelity while cutting copy-paste duplication; suited for multi-person image synthesis but requires identity-annotated face regions and paired training data.