AIAny
Icon for item

AgentGarten: Code Worlds for Evolving Agents

Creates programmable, real-time interactive code-based environments by separating deterministic simulator state from a shared neural video renderer so agents perceive, interact, and iteratively evolve via distilled playbooks; introduces Adversarial Forcing to distill a geometry-conditioned renderer for responsive visual feedback.

Introduction

Most progress in embodied agent learning depends on the environments agents practice in; building many diverse, interactive worlds that are both state-faithful and visually realistic is expensive. AgentGarten flips the usual trade-off by keeping world logic as executable code (exact state and rules) while outsourcing visuals to a learned, shared neural renderer that turns structured geometry into real-time camera observations. The core insight: decouple deterministic simulation from learned rendering so new worlds can be authored in code and rendered through a single, optimized model, enabling many more environments and faster agent iteration.

Key Findings
  • Real-time shared neural renderer: adapts a pretrained bidirectional video model to geometry-conditioned synthesis, converts it to block-causal generation, and distills it with a novel Adversarial Forcing objective to run interactive rollouts at interactive rates on a single GPU. This enables agents to receive short-block visual feedback immediately after actions.
  • Adversarial Forcing & exact replay: training uses two-pass exact replay so later-losses update history encoding without prohibitive memory, and combines score-distillation with real-data adversarial supervision (with an exact R1/R2 scheme) to avoid long-rollout visual degradation.
  • Code-as-worlds + playbook loop: environments are authored in code (objects, rules, goals) while agents perceive only rendered frames. After rounds of play, agents distill lessons into written playbooks that later agents inherit, enabling rapid emergence of complex behaviors (e.g., shelter-building, ramp use) within a handful of rounds versus millions of conventional RL steps.
Who It's For and Tradeoffs

Great fit if you are researching embodied or interactive learning and need many editable, programmatic environments that preserve exact state and customizable interaction rules while giving visually consistent observations. It benefits teams that can accept learned visual fidelity (neural rendering) rather than photorealistic asset pipelines and who value rapid iteration across many bespoke environments. Look elsewhere if you require perfect photorealism, strict physical fidelity beyond what the simulator+renderer pair provides, or if you cannot rely on pretrained video models and adversarial distillation workflows.

Where It Fits

AgentGarten sits between classical engine-driven simulators (precise visuals but costly authoring) and pure learned world models (fast generalization but weak rule fidelity). It is most useful when you want programmatic control of dynamics and scalable visual diversity without hand-crafting assets for each scene.

How It Works (brief)

The system runs scene programs in simulators/game engines that export structured geometry and state. A shared neural renderer consumes those conditions and produces short video blocks in a block-causal manner. During distillation, Adversarial Forcing uses exact replay to backpropagate later-step losses into history encoding and adds a frozen-backbone discriminator with an exact R1/R2 regularizer to stabilize long rollouts. Agents train and play through rounds, writing playbooks that seed subsequent agents, dramatically reducing data/interaction requirements compared with standard RL baselines.

Information

  • Websitearxiv.org
  • AuthorsJiawei Chi, Shangchen Miao, Zhiyuan Shi, Kailu Wu, Hanyang Wang, Weiliang Chen, Qiyu Dai, Jinshan Ren, Jun Gao, Mingsheng Long …
  • Published date2026/10/08

Categories

More Items

Lets a pretrained multimodal LLM interpret navigation requests and orchestrate motion via tool calls for generalist robot navigation across unfamiliar scenes. Key features: an agent harness with Navigation Skills, a unified visual-point interface, task-progress tracking, and tool-based motion execution without navigation-specific fine-tuning.

Measures how effectively LLM agents learn from interaction by playing 20 text-based games with novel or counterintuitive hidden rules, providing deterministic feedback, episode-wise scoring, and controlled variations to test retention and transfer.

Converts historical interaction traces into a reusable, queryable “worldbook” and runs a language-based world model agent (Trace2Env) as the environment for LLM agents — enabling stateful, grounded simulation with improved next-observation fidelity and long-horizon consistency.