Trains a foundation GUI agent using a closed-loop, environment-grounded data stack plus in-context multimodal demonstrations to automate long-horizon desktop workflows. Combines scalable task generation/verification, subtask-level demo guidance, and a 100-task OSWorkerBench benchmark to improve strict success and task progress.
Presents a unified black-box reinforcement learning framework to train and optimize agents running inside complex execution harnesses. Uses sandbox-parallel rollouts, a serving proxy that captures model calls and reconstructs multi-turn trajectories as prefix trees, and adapted GRPO/PPO optimizers to achieve stable, scalable RL across heterogeneous harnesses.
Provides experimental and in-silico data for 1,440 de novo miniprotein binders designed by Anthropic's Claude models, including per-design kinetics, raw sensorgrams, structure-predictions, and design provenance. Includes two independent wet‑lab assessments and extensive per-design files; data released under CC BY 4.0.
Evaluates whether AI systems can independently carry out project-level scientific research by progressively removing human methodological guidance across 60 tasks in 11 domains. Built with expert review, sandbox execution, and multi-agent–model scoring to measure innovation and autonomous experimental execution.