AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Category

Explore by categories

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All Categories

  • AI Leaderboard

  • AI Agent Tutorials

  • AI Coding Tutorials

  • AI Model

  • AI Agent Papers

  • Chatbot

  • AI Dataset

  • Machine Learning Foundation Books

  • AI Train

  • AI Deploy

  • AI Client

  • Machine Learning Foundation Papers

  • Machine Learning Foundation Tutorials

  • AI Image Demos

  • AI Agent

  • Large Language Model Tutorials

  • Large Language Model Papers

  • Machine Learning Engineering Papers

  • Computer Vision Tutorials

  • Computer Vision Papers

  • Natural Language Processing Papers

  • Reinforcement Learning Papers

  • Speech Technology Papers

  • AI API

  • AI Coding

  • AI Image

  • AI Video

  • MLOps

  • MCP Client

  • MCP Server

  • AI Video Papers

  • AI Audio

  • AI Others

  • AI Infra

  • Embodied AI

Large Language Model Papers·2026
Icon for item

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Zhi Zheng, Rongsheng Chen +4

Fine-tunes long-horizon LLM agents with evolution strategies so full-model updates run at inference-level GPU memory. Emphasizes trajectory-level credit via black-box rewards, online prompt–parameter co-evolution, and a cosine decay for perturbation scale to balance exploration and adaptation; suited for limited-GPU settings.

#llm#long-horizon#ai-agent#qwen#rl+3
Large Language Model Papers·2026
Icon for item

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Nayeon Kim, Hojin Lee +3·Kakao Corp., Upstage AI

Estimates optimal learning rates for large-scale Mixture-of-Experts pretraining using a two-step, compute-efficient transfer: μP-based width transfer from small proxy models, then log-log linear extrapolation across token budgets to trillion-token horizons.

#foundation-model#LLM#nlp#paper#ai-train+1
AI Video Papers·2026
Icon for item

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

Yunheng Li, Guohong Mu +5·VCIP, School of Computer Science, Nankai University, Brain and Artificial Intelligence Lab, Northwestern Polytechnical University +2

Treats human annotations as oracle rollouts and separates them from on-policy baselines to improve reinforcement learning for video multimodal LLMs. Key features include a decoupled advantage estimator, sign-balanced pruning, and scalable gains across model sizes and data budgets.

#video#multimodal#rl#LLM#vision+3
AI Agent Papers·2026
Icon for item

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

Sungho Park, Wonjoong Kim +11·Affiliation: KAIST, Affiliation: Southern University of Science and Technology +1

Automatically optimizes runtime harnesses for LLM agents by diagnosing failure traces and iteratively applying structured, generalizable patches. Combines batch-based failure diagnosis, code-like patch generation across prompts/tools/middleware, and validation-aware selection to raise long-horizon task success on multiple benchmarks.

#LLM#ai-agent#agent-skills#long-horizon#benchmarks+4
Large Language Model Papers·2026
Icon for item

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

Justin Robert, Raheel Qader·OVHai LLM

Analyzes on-policy self-distillation for language-model reasoning, diagnosing “collapse” and framing it as controlled by three levers: where token-level signals apply, what privileged information the teacher sees, and how teacher dynamics evolve.

#distillation#reasoning#LLM#paper#rl
AI Agent Papers·2026
Icon for item

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

Guibin Zhang, Leo Lu +14

Synthesizes, repairs, and self-evolves task-adaptive agent harnesses on demand for off-the-shelf LLM agents, using a trainable harness-intelligence model that distills signals from past configurations. Demonstrates consistent performance gains across benchmarks and model families by producing four-module, composable harnesses.

#ai-agent#agent-skills#llm#deepseek#qwen+5
AI Agent Papers·2026
Icon for item

What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents

Xingshan Zeng, Zishan Xu +12

Analyzes how to generate useful interaction data for LLM agents and proposes the ACE lens — Accuracy, Complexity, divErsity — while factorizing agentic data as (E, q, τ, v). Surveys verification, difficulty calibration, and coverage strategies and outlines implications for training and benchmarks.

#LLM#agent-skills#evaluation#benchmarks#ai-agent+2
Large Language Model Papers·2026
Icon for item

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

Gyouk Chu, Myeongho Jeon +1·KAIST

A zero-data self-evolution framework that co-trains a Challenger, Solver, and Judge so LLMs can iteratively improve on both verifiable and unverifiable tasks without human labels. Uses role-asymmetry and subtask-amplification preference pairs to train the Judge and sustain improvement.

#paper#LLM#reasoning#evaluation#research+1
Large Language Model Papers·2026
Icon for item

TTPO: Test-Time Policy Optimization

Aozhe Wang, Zhengxi Lu +9·Affiliation: Zhejiang University, Affiliation: Alibaba Group{waz,zhengxilu,syl}@zju.edu.cn    [email protected]

A test-time method that adapts LLMs without labels by distilling rollouts that agree with majority pseudo-labels and penalizing disagreeing rollouts via grouped RL, improving robustness under frequent pseudo-label errors.

#RL#LLM#reasoning#qwen#distillation+1
Large Language Model Papers·2026
Icon for item

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training

Tingyun Li, Wenfeng Feng +4·Abudukelimu Wuerkaixi, Guohua Liu +1

Decides when past post-training updates should be reused for autonomous LLM adaptation by introducing Boundary-Calibrated Intervention Transfer (BCIT). BCIT binds effects to source context, checks applicability and hard conflicts, and runs bounded trials to obtain current-state evidence—reducing harmful updates and improving equal-budget final-model quality.

#LLM#foundation-model#evaluation#agent-skills#ai-agent+2
Large Language Model Papers·2026
Icon for item

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

Yi Ding, Ruqi Zhang·Affiliation: Department of Computer Science, Purdue University, USA

Analyzes on-policy distillation for LLM fine-tuning, shows teacher token-level supervision is often noisy and not the main driver of gains, and introduces OPSA, a supervision-free, entropy-adaptive method that suppresses low-probability tokens to improve downstream accuracy.

#distillation#LLM#rl#nlp#qwen+2
Large Language Model Papers·2026
Icon for item

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Shaowen Wang, Ge Zhang +7

Studies looping shared transformer layers in Mixture-of-Experts models under matched budgets and proposes SMELT: loop the middle half twice while matching per-token FLOPs, non-embedding parameters, and KV cache. Shows 6.8–18.0% training-FLOPs savings on the compute-optimal frontier, stronger downstream gains on code and long-context tasks.

#moe#transformers#llm#research#benchmark+2
  • Previous
  • 1
  • More pages
  • 8
  • 9
  • 10
  • More pages
  • 14
  • Next