AIAny
Icon for item

From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation

Converts historical interaction traces into a reusable, queryable “worldbook” and runs a language-based world model agent (Trace2Env) as the environment for LLM agents — enabling stateful, grounded simulation with improved next-observation fidelity and long-horizon consistency.

Introduction

Why simulate environments from traces rather than rebuild them? Many realistic systems are impractical to reimplement, yet their recorded interaction traces already contain the evidence needed to emulate how they behave. This paper's core insight is that a world model agent can act as the environment by consulting a reconstructed trace-derived knowledge base (a worldbook) plus an episodic state, giving task agents stateful, actionable observations without requiring an executable replica.

Key Findings
  • Trace2Env reconstructs interaction traces into a worldbook containing environment schemas, grounded evidence, and induced behavioral knowledge; at runtime a world-model agent actively queries this worldbook and maintains episodic state to predict observations and lasting state effects.
  • The approach is learning-free: it relies on structured reconstruction and prompt-driven agentic inference rather than training heavy simulators, making it practical when original systems are inaccessible but traces exist.
  • Across nine evaluated environments, Trace2Env improves next-observation fidelity and long-horizon interaction consistency compared to prompt-only language world models; actions proposed by a task agent in Trace2Env remain valid more often when replayed in the real environment.
Who it helps and tradeoffs

Great fit if you have rich historical interaction logs but cannot stand up the original environment — e.g., proprietary web systems, complex multi-component services, or archived agent runs. It enables faster iteration and safer offline evaluation for LLM-based agents. Look elsewhere if you need exact executable semantics (e.g., binary-level system behavior), if traces are sparse or biased (the method inherits trace coverage), or if you require learned generalization to unseen dynamics that go beyond recorded evidence.

Where it fits

Positions itself between prompt-based textual simulators (fast but myopic) and full executable environment reconstruction (accurate but costly). It is particularly useful for agent training, robustness evaluation, and reproducing historical multi-turn behaviors under limited infrastructure.

Mechanism (brief)

Trace-to-worldbook transforms traces into structured schemas and grounded evidence; a dedicated world-model agent consults this worldbook plus a persistent episodic state at each step to infer observations and state transitions. The pipeline emphasizes grounded retrieval and agentic reasoning over parameter-heavy model learning, trading learnable generalization for trace-faithful, interpretable simulation.

Information

  • Websitearxiv.org
  • OrganizationsNanyang Technological University, The Hong Kong Polytechnic University
  • AuthorsQuanyu Long, Xiao Chen, Jianda Chen, Haozhen Zhang, Qisheng Hu, Jianzhu Bao, Wenya Wang
  • Published date2026/10/05

More Items

Lets a pretrained multimodal LLM interpret navigation requests and orchestrate motion via tool calls for generalist robot navigation across unfamiliar scenes. Key features: an agent harness with Navigation Skills, a unified visual-point interface, task-progress tracking, and tool-based motion execution without navigation-specific fine-tuning.

Serves token-level routed LLM inference by dispatching requests to per-model asynchronous subservers and using delayed-batching scheduling to reduce admission latency and step desynchronization. Exposes a request-centric route-send-receive API and reports 2.01–64.15× decoding throughput gains versus single-LLM servers.

Creates programmable, real-time interactive code-based environments by separating deterministic simulator state from a shared neural video renderer so agents perceive, interact, and iteratively evolve via distilled playbooks; introduces Adversarial Forcing to distill a geometry-conditioned renderer for responsive visual feedback.