Assesses how AI reviewers react to content-preserving rewrites and introduces RobustReview, a controlled benchmark plus SciCore, a dual-branch reviewer that combines full-manuscript judgment with a structured 'science core' summary to reduce rhetorical sensitivity.
Constructs and continually maintains explicit belief states for long-horizon LLM agents, combining a structured world estimate with unresolved epistemic and achievement gaps. Adds consistency validation, Belief Trapping detection, and tailored recovery to improve execution and diagnosis benchmarks.