AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Category

Explore by categories

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All Categories

  • AI Leaderboard

  • AI Agent Tutorials

  • AI Coding Tutorials

  • AI Model

  • AI Agent Papers

  • Chatbot

  • AI Dataset

  • Machine Learning Foundation Books

  • AI Train

  • AI Deploy

  • AI Client

  • Machine Learning Foundation Papers

  • Machine Learning Foundation Tutorials

  • AI Image Demos

  • AI Agent

  • Large Language Model Tutorials

  • Large Language Model Papers

  • Machine Learning Engineering Papers

  • Computer Vision Tutorials

  • Computer Vision Papers

  • Natural Language Processing Papers

  • Reinforcement Learning Papers

  • Speech Technology Papers

  • AI API

  • AI Coding

  • AI Image

  • AI Video

  • MLOps

  • MCP Client

  • MCP Server

  • AI Video Papers

  • AI Audio

  • AI Others

  • AI Infra

  • Embodied AI

Reinforcement Learning Papers·2026
Icon for item

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

Byung-Kwan Lee, Ximing Lu +9

Proposes ZPPO, a distillation method that keeps the teacher inside prompts rather than injecting teacher gradients, using binary- and negative-candidate prompts plus a prompt replay buffer to recover learning signal on hard examples; shows gains for small Qwen3.5 students across 31 multimodal benchmarks.

#qwen#RL#llm#multimodal#vision+2
Computer Vision Papers·2026
Icon for item

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

Yatai Ji, An-Chieh Cheng +14

Provides a dual-path approach for spatial vision-language models: a Language-Only Reasoning (LOR) path for stepwise linguistic deduction and a Detect-Then-Reason (DTR) path that detects 3D cues via region tokens before numerical inference. Trains with chain-of-thought cold-start supervision and reinforcement learning to improve 3D grounding and multi-step spatial reasoning.

#vision#multimodal#RL#paper#depth+1
Large Language Model Papers·2026
Icon for item

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Jing Liang, Hongyao Tang +10·Tianjin University, Alibaba

Proposes Monotonic Inference Policy Improvement (MIPI) and a two-step Monotonic Inference Policy Update (MIPU) to address training–inference probability mismatch in LLM reinforcement learning by constructing sampler-referenced candidate updates and accepting synchronized updates using an inference-gap proxy; shows improved reasoning accuracy and stability under FP8-quantized rollouts.

#RL#llm#vllm#qwen#ai-train+3
Computer Vision Papers·2026
Icon for item

Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts

Taewook Kang, Taeheon Kim +2

Adapts pretrained Vision-Language-Action (VLA) models to new camera poses and robot embodiments from a single demonstration by performing weight-vector arithmetic that injects domain-specific information. Filters noise via subspace alignment of singular components; designed for one-shot adaptation under visual and embodiment shifts.

#robotics#vision#multimodal#paper#code+2
Reinforcement Learning Papers·2026
Icon for item

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Zhilin Wang, Han Song +14·University of Science and Technology of China, The Chinese University of Hong Kong +6

Provides a benchmark and protocol to evaluate agents that iteratively edit executable policies under a fixed interaction budget, recording full execution–feedback–revise trajectories. Built from compact RL environments with trajectory-level diagnostics and hidden held-out validation.

#RL#evaluation#ai-agent#agent-skills#ai-leaderboard+2
Embodied AI·2026
Icon for item

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

Angyuan Ma, Boyuan Wang +24

Provides a systematic benchmark and design roadmap for video-based world models to evaluate robot policies, introducing WMBench and GigaWorld-1 optimized for long-horizon, action-faithful rollouts. Offers controlled comparisons across model families, action encodings, and 324k+ simulated vs real rollouts, with code, models, and datasets released for reproducible evaluation.

#evaluation#robotics#video#vision#RL+4
Embodied AI·2026
Icon for item

VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon

Yi Pan, Miao Pan +9

Detects when an action-chunked VLA policy drifts from expected visual dynamics and triggers lightweight corrective replanning via a latent-space vision monitor and online gradient guidance; creates an event-driven adaptive action horizon without retraining the backbone.

#vision#robotics#RL#multimodal#foundation-model+2
AI Agent Papers·2026
Icon for item

UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning

Niu Lian, Alan Chen +9

Trains cross-platform GUI agents by combining a Uni-GUI cross-platform dataset with platform-conditioned multi-teacher on-policy distillation, enabling a shared policy to adapt to new platforms while retaining platform-specific behaviors; suitable for research on continual GUI agent learning and cross-platform adaptation.

#RL#multimodal#agent-skills#ai-agent#paper+2
Reinforcement Learning Papers·2026
Icon for item

Trust Region Policy Distillation

Zhengpeng Xie, Li Lyna Zhang +2

Stabilizes on-policy policy distillation by dynamically constructing a proximal teacher that controls gradient variance. Provides theoretical global convergence and monotonic improvement bounds, and shows improved training stability, sample efficiency, and final performance on mathematical reasoning tasks with zero extra compute overhead.

#RL#paper#algorithms#math
Reinforcement Learning Papers·2026
Icon for item

Weak-to-Strong Generalization via Direct On-Policy Distillation

Shiyuan Feng, Huan-ang Gao +8

Transfers RL-induced policy shifts from a smaller 'weak' teacher to a stronger target by using the teacher's post-/pre-RL log-ratio as a dense implicit reward applied on the student's on-policy states. Enables reuse of RL supervision without running RL rollouts on the target, improving sample/time efficiency.

#RL#LLM#reasoning#qwen#paper+2
Computer Vision Papers·2026
Icon for item

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Hongyu Qu, Jianzhe Gao +7

Reconstructs historical experience into latent memory tokens and weaves short- and long-term latent memories directly into vision-language-action reasoning to improve long-horizon robotic manipulation. Uses a four-part pipeline (curator, seeker, condenser, weaver) so memory participates natively in multimodal action formation.

#vision#robotics#multimodal#paper#embeddings+2
AI Agent Papers·2026
Icon for item

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Zongxia Li, Zhongzhi Li +11

Provides a terminal-style benchmark of 46 long-horizon tasks decomposed into fine-grained graded subtasks to produce dense intermediate rewards and partial credit, enabling evaluation of long-horizon planning, long-context management, and iterative debugging. Tasks typically require hundreds of episodes and minutes-to-hours of execution; baseline evaluations report high token and episode consumption with low pass rates, highlighting evaluation headroom.

#evaluation#agent-skills#RL#terminal#paper+4
  • Previous
  • 1
  • 2
  • 3
  • 4
  • 5
  • Next