Evaluates AI agents' ability to complete end-to-end scientific workflows by releasing and assessing 97 tasks from a 300-task FrontierChallenge suite across chemistry, materials, life science, and electrochemistry. Finds that top agent configurations achieved only a 20.6% pass rate despite high partial scores, revealing a gap between partial progress/confident completion claims and actual complete scientific deliverables.
Measures whether video generators reproduce the correct distribution of possible physical behaviors under repeated rollouts. Introduces PAWBench and PAWEval to convert repeated generations into outcome-level empirical distributions and quantify probabilistic alignment; evaluates 50 scenarios and 11 models and finds no model consistently matches reference probabilities.
Rewrites physical scenes as executable world programs (e.g., MuJoCo scene descriptions) and uses an agentic abductive loop to propose, execute, render, verify, and iteratively refine those programs from videos or text. Verified executable worlds supply scalable physical supervision for training vision–language models.
Provides a human-verified benchmark of 1,927 heterogeneous articulated 3D objects with part-level articulation semantics and intrinsic physical-property annotations for evaluating physical grounding and simulation readiness. Includes URDF assemblies, aligned point clouds, per-part JSON annotations, and a curated evaluation protocol; licensed CC BY-NC 4.0 (non-commercial).
Simulation-ready home dataset for embodied AI: CAD-based household scenes with configured physical properties and metadata, plus 1,000 robot trajectory episodes (RGB-D, HDF5/USDZ) for simulation training and evaluation under CC BY-NC-SA 4.0.
Provides a domain-agnostic world-modeling framework that factorizes latent targets into orthogonal predictive components, with dedicated prediction branches and synthesis for multi-domain forecasting and intervention. Demonstrates improved dynamics and long-horizon rollouts across seven domains and includes experimental biological validation.
Adds token-conditioned quantum residual branches to a frozen masked-diffusion language model: a lightweight hypernetwork emits continuous quantum-circuit coordinates per token, executes a shared sparse IQP-style circuit, and injects classically-expressible expectation readouts back into transformer blocks. Trains only the added branches, scaling to 16–64 qubits with analytic, linear-cost readouts.
Learns dynamics-aligned latent representations for parametric PDE forecasting using masked-latent predictive pretraining, a geometry projector that aligns latent trajectory geometry, and a physics-structured latent predictor that separates shared evolution from parameter-dependent responses.