Most automatic game-generation work focuses on runnable correctness; this paper argues that playable code alone does not guarantee a good player experience. The core insight is to treat game evolution as a studio-like recursive loop where design, implementation, fast programmatic playtesting, and trajectory-driven review jointly steer repeated refinements toward intended player experience. Separating fast policy execution from design/review enables frequent, diverse rollouts and evidence-based revisions.
Key Findings
- A four-role harness (Designer, Builder, Coding‑Native Player, Reviewer) organizes iterative refinement: the Designer expands user goals into a design graph; the Builder implements candidates; the Player writes reusable programmatic policies to collect high-frequency trajectories; the Reviewer uses trajectory-based metrics plus visual evidence to infer preferences and prioritize fixes. This architecture ties construction to playtesting in a reproducible loop.
- Coding‑Native Player vs GUI testing: programmatic policies expose structured state/events and enable rapid, repeatable, diverse rollouts without per-action visual interpretation, which reduces evaluation latency and bias and surfaces corner cases more efficiently.
- Experience‑oriented Reviewer: combines general metrics (playtime, success rate) with game-specific rubrics derived from trajectories and screenshots, producing structured feedback that the Designer uses to produce concrete acceptance goals and implementation plans for the next round.
- Empirical gains: multi-round refinement improved GameCraft-Bench overall score (from 72.70 → 77.89) and raised strict success rates and runtime-check pass rates on GameASG-Bench, accompanied by longer playtime and higher user ratings in a study.
Who it fits and trade-offs
Great fit if you want automated pipelines that move beyond runnable prototypes toward measurable player experience improvements and can instrument games programmatically. It is especially relevant for research and studios that can run many fast rollouts and care about reproducible trajectory evidence. Look elsewhere if your target platform cannot expose programmatic interfaces (only opaque GUI) or if human-driven qualitative playtesting is the primary evaluation channel; the method assumes the ability to execute programmatic policies and to collect state/event traces.
Where it sits
This paper bridges coding-agent game generation and evaluation/UX work: it pairs agentic code production with systematic, policy-driven playtesting and a reviewer that translates play traces into prioritized design changes. The main methodological novelty is integrating fast programmatic rollouts (coding-native players) with an experience-focused reviewer inside a recursive development loop.