Defines and evaluates AREX-2, an LLM agent that iteratively self-improves at test time via reflection and long-horizon execution. Trained on long-horizon improvement trajectories from ML engineering and algorithmic programming (built on Qwen3.8-27B), it scales with more rounds and achieves strong benchmark scores.
Groups visually grounded appearances of the same physical instance into persistent, retrievable “biographies” so agents can follow objects across hours or days for long-video question answering. Links identity-aware observations to episodic context and visual evidence; improves EgoLifeQA to 72.0% and increases evidence-window reach from 37.6% to 58.9%.