AIAny
Icon for item

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

Calibrates multi-reward reinforcement learning by adaptively upweighting infrequently active rewards per rollout batch, so sparse objectives provide stronger signals when they matter. Proposes an inverse-square-root density correction and shows faster learning on tool-calling and math-reasoning tasks.

Introduction

Most multi-reward RL setups normalize rewards within rollout groups but still see uneven learning: objectives that produce nonzero comparisons less often are underrepresented at the batch level. The paper's core insight is that a reward's batch-level signal (advantage energy) scales with its active-group density, so matching energies across rewards removes this residual imbalance and speeds learning for sparse objectives.

Key Findings
  • Analysis: defines advantage energy as the sum of squared advantages and shows under idealized GDPO normalization the energy is proportional to active-group density (fraction of groups where the reward gives nonzero relative advantage). This identifies a batch-level source of signal imbalance that reward-wise normalization does not fix.
  • Method: derives an inverse-square-root density correction to scale each reward's contribution so that less frequently active rewards receive larger coefficients when active. The practical weight is computed per rollout batch and capped to avoid extreme scaling.
  • Empirical results: on tool-calling and mathematical reasoning benchmarks, the method (DARA) reaches high format compliance up to 26% faster and near-saturated length compliance up to 65% faster than GDPO, while remaining competitive in final performance and requiring no change to the underlying policy optimization objective.
Who it's for and trade-offs

Great fit if you train LLM-based policies or other sequence models with multiple behavioral objectives and some rewards are sparse (format checks, rare constraints, specialized correctness signals). DARA is appealing when per-batch reward activity varies over training and you prefer an adaptive, data-driven aggregation over hand-tuned scalar weights. Look elsewhere if reward sparsity is not an issue, per-group comparisons are already balanced, or if you cannot afford the bookkeeping to compute per-batch active-group densities during rollout processing.

How it works (brief)

DARA measures, per batch, each reward's active-group density π_k (fraction of rollout groups where the reward induces a nonzero group-relative advantage). To equalize advantage energy across rewards, it rescales reward k by w_k ≈ sqrt(π_ref / π_k) where π_ref is the largest density in the batch, with a practical cap w_max. This increases the effective signal from infrequently active rewards only when they provide useful comparisons, and weights adapt as reward activity changes during training.

Information

  • Websitearxiv.org
  • OrganizationsUniversity of Chinese Academy of Sciences, University of Minnesota Twin Cities, University of Wisconsin–Madison, Stony Brook University, Dalian University of Technology, Beijing Foreign Studies University, Southeast University, Kuaishou Technology
  • AuthorsTong Zheng, Skylar Zhai, Zhan Cheng, TianMing Sha, Youling Huang, Shuo Zhou, Shaotong Qi, Jingcheng Liang, Xuwei Ding, Pengcheng Xu
  • Published date2026/09/30

More Items

Models sequence generation by unmasking multiple tokens per denoising step and replaces a factorized reverse process with a mixture over discrete routing-based latents from an MoE backbone; improves few-step sampling quality without increasing active parameters.

Demonstrates that pretrained transformers typically use only ~1–3 lines of depth to follow reference chains, and that a task‑trained rank‑8 LoRA applied at one early layer (with all other weights frozen) can extend reference‑following to dozens or hundreds of lines while adding only a few ten‑thousand parameters.

Investigates how rollout policy, token-level KL direction, and learning rate each affect LLM distillation across Llama3 and Qwen2.5 on reasoning tasks; finds KL direction and learning rate dominate outcomes while rollout policy has a modest effect.