AIAny
Icon for item

Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?

Measures how effectively LLM agents learn from interaction by playing 20 text-based games with novel or counterintuitive hidden rules, providing deterministic feedback, episode-wise scoring, and controlled variations to test retention and transfer.

Introduction

Why this matters LLM agents are often evaluated on tasks whose rules are given or already familiar from pretraining, which conflates reasoning with prior knowledge and true online learning. This benchmark isolates learning-from-interaction by forcing agents to discover novel, sometimes counterintuitive rules through trial, observation, and episodic memory — without updating model weights — and then measures whether they improve and transfer that knowledge across episodes.

Key Findings
  • Benchmark design: 20 game templates (10 fixed, 10 reshuffled) produce deterministic, scored episodes so agents can play repeated attempts and researchers can measure learning trajectories and transfer to new instances.
  • Evidence vs. compression: Keeping full interaction histories often supports stronger learning than compressing experiences into abstracted rules or short summaries, because raw evidence lets agents revise earlier conclusions.
  • Human–agent gap: Top human players achieve higher peak scores, explore more diverse strategies, and recover from setbacks more often than evaluated agents.
  • Harness matters: With the same backbone model, different agent harnesses (execution environment, memory/replay protocols, tools) materially change performance and inference cost; better harness design can improve learning while reducing model calls.
  • Transfer and brittleness: Agents can learn rules but still fail at strategic planning (e.g., solving subproblems but not sequencing actions), and changing visible instance layouts slows learning even when core rules remain fixed.
What the benchmark provides
  • A controlled suite for studying experience-driven learning: 20 game templates with automatic scoring and reproducible feedback; five identified challenge games for concentrated evaluation.
  • Protocols and baselines: a common agent loop, memory format, and evaluations across multiple backbone models and harnesses to quantify how design choices affect learning.
  • Reproducible metrics for retention and transfer across repeated episodes without weight updates.
Who it's for and trade-offs

Great fit if you study how LLM-based agents acquire procedural or environment-specific knowledge from interaction, want reproducible learning curves, or need a controlled testbed to compare harness and memory designs. Look elsewhere if you need continuous online learning with model weight updates, embodied/robotics physics fidelity, or high-fidelity multimodal environments—the benchmark focuses on text-based, deterministic interactions and episodic evaluation rather than full continual learning or real-world sensorimotor simulation.

Information

  • Websitearxiv.org
  • AuthorsYibo Li, Jinhang Qiu, Zhi Zheng, Qianyun Guo, Jiaying Wu, Shuo Ji, Bryan Hooi
  • Published date2026/10/06

Categories

More Items

Lets a pretrained multimodal LLM interpret navigation requests and orchestrate motion via tool calls for generalist robot navigation across unfamiliar scenes. Key features: an agent harness with Navigation Skills, a unified visual-point interface, task-progress tracking, and tool-based motion execution without navigation-specific fine-tuning.

Creates programmable, real-time interactive code-based environments by separating deterministic simulator state from a shared neural video renderer so agents perceive, interact, and iteratively evolve via distilled playbooks; introduces Adversarial Forcing to distill a geometry-conditioned renderer for responsive visual feedback.

Converts historical interaction traces into a reusable, queryable “worldbook” and runs a language-based world model agent (Trace2Env) as the environment for LLM agents — enabling stateful, grounded simulation with improved next-observation fidelity and long-horizon consistency.