Why simulate environments from traces rather than rebuild them? Many realistic systems are impractical to reimplement, yet their recorded interaction traces already contain the evidence needed to emulate how they behave. This paper's core insight is that a world model agent can act as the environment by consulting a reconstructed trace-derived knowledge base (a worldbook) plus an episodic state, giving task agents stateful, actionable observations without requiring an executable replica.
Key Findings
- Trace2Env reconstructs interaction traces into a worldbook containing environment schemas, grounded evidence, and induced behavioral knowledge; at runtime a world-model agent actively queries this worldbook and maintains episodic state to predict observations and lasting state effects.
- The approach is learning-free: it relies on structured reconstruction and prompt-driven agentic inference rather than training heavy simulators, making it practical when original systems are inaccessible but traces exist.
- Across nine evaluated environments, Trace2Env improves next-observation fidelity and long-horizon interaction consistency compared to prompt-only language world models; actions proposed by a task agent in Trace2Env remain valid more often when replayed in the real environment.
Who it helps and tradeoffs
Great fit if you have rich historical interaction logs but cannot stand up the original environment — e.g., proprietary web systems, complex multi-component services, or archived agent runs. It enables faster iteration and safer offline evaluation for LLM-based agents. Look elsewhere if you need exact executable semantics (e.g., binary-level system behavior), if traces are sparse or biased (the method inherits trace coverage), or if you require learned generalization to unseen dynamics that go beyond recorded evidence.
Where it fits
Positions itself between prompt-based textual simulators (fast but myopic) and full executable environment reconstruction (accurate but costly). It is particularly useful for agent training, robustness evaluation, and reproducing historical multi-turn behaviors under limited infrastructure.
Mechanism (brief)
Trace-to-worldbook transforms traces into structured schemas and grounded evidence; a dedicated world-model agent consults this worldbook plus a persistent episodic state at each step to infer observations and state transitions. The pipeline emphasizes grounded retrieval and agentic reasoning over parameter-heavy model learning, trading learnable generalization for trace-faithful, interpretable simulation.