AIAny
Icon for item

Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness

Iteratively refines agent-generated games using a closed loop of Designer, Builder, Coding‑Native Player, and Experience‑Oriented Reviewer — collects programmatic gameplay trajectories and trajectory+visual evaluations to evolve prototypes into player-focused, product-level games; shows measurable gains on GameCraft-Bench and GameASG-Bench.

Introduction

Most automatic game-generation work focuses on runnable correctness; this paper argues that playable code alone does not guarantee a good player experience. The core insight is to treat game evolution as a studio-like recursive loop where design, implementation, fast programmatic playtesting, and trajectory-driven review jointly steer repeated refinements toward intended player experience. Separating fast policy execution from design/review enables frequent, diverse rollouts and evidence-based revisions.

Key Findings
  • A four-role harness (Designer, Builder, Coding‑Native Player, Reviewer) organizes iterative refinement: the Designer expands user goals into a design graph; the Builder implements candidates; the Player writes reusable programmatic policies to collect high-frequency trajectories; the Reviewer uses trajectory-based metrics plus visual evidence to infer preferences and prioritize fixes. This architecture ties construction to playtesting in a reproducible loop.
  • Coding‑Native Player vs GUI testing: programmatic policies expose structured state/events and enable rapid, repeatable, diverse rollouts without per-action visual interpretation, which reduces evaluation latency and bias and surfaces corner cases more efficiently.
  • Experience‑oriented Reviewer: combines general metrics (playtime, success rate) with game-specific rubrics derived from trajectories and screenshots, producing structured feedback that the Designer uses to produce concrete acceptance goals and implementation plans for the next round.
  • Empirical gains: multi-round refinement improved GameCraft-Bench overall score (from 72.70 → 77.89) and raised strict success rates and runtime-check pass rates on GameASG-Bench, accompanied by longer playtime and higher user ratings in a study.
Who it fits and trade-offs

Great fit if you want automated pipelines that move beyond runnable prototypes toward measurable player experience improvements and can instrument games programmatically. It is especially relevant for research and studios that can run many fast rollouts and care about reproducible trajectory evidence. Look elsewhere if your target platform cannot expose programmatic interfaces (only opaque GUI) or if human-driven qualitative playtesting is the primary evaluation channel; the method assumes the ability to execute programmatic policies and to collect state/event traces.

Where it sits

This paper bridges coding-agent game generation and evaluation/UX work: it pairs agentic code production with systematic, policy-driven playtesting and a reviewer that translates play traces into prioritized design changes. The main methodological novelty is integrating fast programmatic rollouts (coding-native players) with an experience-focused reviewer inside a recursive development loop.

Information

  • Websitearxiv.org
  • OrganizationsHKU MMLab, The University of Hong Kong, Shenzhen Loop Area Institute
  • AuthorsJiajun Chen, Haoyu Wu, Mingda Jia, Xihui Liu
  • Published date2026/10/06

Categories

More Items

Runs an open-source personal software agent across a user's devices to perform actions, keep readable provenance-backed memories, and coordinate via a self-hostable relay. Ships minimal end-to-end clients (mobile, desktop, web) under GPL-3.0; model-agnostic and designed for inspectability and device

Analyzes where LLM-based agents break the evidence-to-action chain and introduces SafeActBench, a 656-case, provenance-bound benchmark and deterministic evaluator to diagnose failures in investigation, timing, single-action execution, and multi-step workflows.

Verifies and preserves trajectory-derived skill edits for LLM agents by pairing each proposed edit with replayable execution evidence and re-executing the relevant trajectory segments. Introduces Replayable Evidence Cards, a replay-based verification gate, and a Provisional Edit Ledger to retain locally supported edits across epochs for continual skill evolution.