Most multi-reward RL setups normalize rewards within rollout groups but still see uneven learning: objectives that produce nonzero comparisons less often are underrepresented at the batch level. The paper's core insight is that a reward's batch-level signal (advantage energy) scales with its active-group density, so matching energies across rewards removes this residual imbalance and speeds learning for sparse objectives.
Key Findings
- Analysis: defines advantage energy as the sum of squared advantages and shows under idealized GDPO normalization the energy is proportional to active-group density (fraction of groups where the reward gives nonzero relative advantage). This identifies a batch-level source of signal imbalance that reward-wise normalization does not fix.
- Method: derives an inverse-square-root density correction to scale each reward's contribution so that less frequently active rewards receive larger coefficients when active. The practical weight is computed per rollout batch and capped to avoid extreme scaling.
- Empirical results: on tool-calling and mathematical reasoning benchmarks, the method (DARA) reaches high format compliance up to 26% faster and near-saturated length compliance up to 65% faster than GDPO, while remaining competitive in final performance and requiring no change to the underlying policy optimization objective.
Who it's for and trade-offs
Great fit if you train LLM-based policies or other sequence models with multiple behavioral objectives and some rewards are sparse (format checks, rare constraints, specialized correctness signals). DARA is appealing when per-batch reward activity varies over training and you prefer an adaptive, data-driven aggregation over hand-tuned scalar weights. Look elsewhere if reward sparsity is not an issue, per-group comparisons are already balanced, or if you cannot afford the bookkeeping to compute per-batch active-group densities during rollout processing.
How it works (brief)
DARA measures, per batch, each reward's active-group density π_k (fraction of rollout groups where the reward induces a nonzero group-relative advantage). To equalize advantage energy across rewards, it rescales reward k by w_k ≈ sqrt(π_ref / π_k) where π_ref is the largest density in the batch, with a practical cap w_max. This increases the effective signal from infrequently active rewards only when they provide useful comparisons, and weights adapt as reward activity changes during training.