AIAny
Icon for item

A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

Assesses how AI reviewers react to content-preserving rewrites and introduces RobustReview, a controlled benchmark plus SciCore, a dual-branch reviewer that combines full-manuscript judgment with a structured 'science core' summary to reduce rhetorical sensitivity.

Introduction

AI-based reviewers can be swayed by rhetorical changes even when the reported science is unchanged. This paper formalizes "rhetorical robustness" as the joint need for within-paper stability across content-preserving rewrites and between-paper discrimination, constructs a controlled benchmark that isolates rhetorical variation, and proposes SciCore, a dual-branch reviewer that pairs manuscript-level assessment with a content-normalized scientific core to improve robustness without discarding manuscript context.

Key Findings
  • RobustReview: a controlled benchmark built from 60 anonymized ICLR 2026 submissions and 10 rhetorical variants per paper (1,260 manuscript versions) to measure rewrite sensitivity and discriminability. Evaluated 30 reviewer configurations (rubric-instructed LLMs, specialized reviewers, agentic systems).
  • False robustness: some reviewers show low sensitivity to rewrites only because their scores collapse across different papers — stability without meaningful discrimination.
  • Prompting limits: a previously evaluated content-focused prompting protocol does not reliably improve robustness across model backbones.
  • SciCore design: a simple equal-weight fusion of (A) a full-manuscript judgment and (B) a judgment on an extracted, structured record of problem, claims, methods, evidence, results, and limitations. The science-core branch is intended as a content-normalized view that varies less under rhetorical rewrites while preserving inter-paper distinctions.
  • Empirical outcome: in the primary GPT-5.5 comparison, SciCore achieves a leading joint profile (e.g., ICC 0.775, SPR 0.652, discriminability 0.726) while maintaining competitive human alignment (human MAE ~1.072, Spearman ~0.488), though it does not minimize every robustness metric (e.g., MAD or drift SD).
Who it's for and tradeoffs

Great fit if you care about deploying AI-assisted peer review or evaluating review models and need metrics that detect when rhetorical changes unfairly affect judgments. SciCore is useful for teams that want a principled way to combine manuscript-level signals with content-normalized summaries to reduce gaming by rhetorical optimization.

Look elsewhere if you need a turnkey production reviewer without additional extraction steps or if minimizing a specific metric like MAD/drift SD is the sole priority; SciCore improves joint stability–discrimination but adds an extraction branch and does not dominate every single robustness measure.

Information

  • Websitearxiv.org
  • AuthorsChenguang Wang, Ming Li, Chengrui Fan, Jianpeng Chen, Han Chen, Tianyi Zhou, Dawei Zhou
  • Published date2026/09/30

More Items

Couples a continuous latent diffusion trajectory with discrete token readouts so tokens are read from the latent at every step and fed back as a scaffold, reducing token-independence in parallel decoding and improving structured-task accuracy and LM perplexity.

Teaches small reasoning models to diagnose when additional internal thinking is insufficient and to selectively query stronger models; introduces FlyBy, a supervised + cost-aware RL framework that learns when/what to ask. Improves pass@ metrics on hard benchmarks while reducing serving cost.

Shows that private post-training changes leave measurable "behavioral shadows" in task-unrelated single-word outputs and proposes Active Taskless Distillation (ATD) to transfer capabilities to a student using only those single-word teacher responses, without teacher logits or parameters.