AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Large Language Model Papers·2026
Icon for item

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

Gyouk Chu, Myeongho Jeon +1·KAIST

A zero-data self-evolution framework that co-trains a Challenger, Solver, and Judge so LLMs can iteratively improve on both verifiable and unverifiable tasks without human labels. Uses role-asymmetry and subtask-amplification preference pairs to train the Judge and sustain improvement.

#paper#LLM#reasoning#evaluation#research+1
Computer Vision Papers·2026
Icon for item

Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction

Jiarong Han, Jincheng Xiong +7·Alibaba Group

Performs causal, bounded‑memory streaming 3D reconstruction by caching KV features from only the preceding 11 frames, predicting a per‑frame point map and adjacent relative pose, and composing these local predictions into a global trajectory; includes a lightweight rotation refiner and composition‑aware loss to limit drift.

#paper#vision#long-horizon#depth#benchmark+3
Hugging Face
AI Model·2026
Icon for item

Hy4 preview

Tencent Hy Team, Tencent

A 770B-parameter Mixture-of-Experts instruct model from Tencent that natively supports 1,048,576-token contexts, Gated DSA attention, and speculative MTP decoding; open-sourced under Apache-2.0 with BF16 and FP8 weights for deployable inference.

#moe#hy_v4#safetensors#transformers#fp8+6
AI Video Papers·2026
Icon for item

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Yuandong Pu, Le Zhuo +12·Shanghai Jiao Tong University, Shanghai AI Laboratory +5

Measures whether video generators reproduce the correct distribution of possible physical behaviors under repeated rollouts. Introduces PAWBench and PAWEval to convert repeated generations into outcome-level empirical distributions and quantify probabilistic alignment; evaluates 50 scenarios and 11 models and finds no model consistently matches reference probabilities.

#video#benchmark#evaluation#ai-video#research+2
Large Language Model Papers·2026
Icon for item

TTPO: Test-Time Policy Optimization

Aozhe Wang, Zhengxi Lu +9·Affiliation: Zhejiang University, Affiliation: Alibaba Group{waz,zhengxilu,syl}@zju.edu.cn    [email protected]

A test-time method that adapts LLMs without labels by distilling rollouts that agree with majority pseudo-labels and penalizing disagreeing rollouts via grouped RL, improving robustness under frequent pseudo-label errors.

#RL#LLM#reasoning#qwen#distillation+1
Computer Vision Papers·2026
Icon for item

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

Shiyi Zhang, Mushui Liu +9·Tsinghua University, Zhejiang University +1

Turns a flow-matching image generator's self-exploration into dense, per-step supervision without a pretrained teacher; it branches the student's next-state into stochastic SDE candidates, scores them against a deterministic self-reference, and applies an advantage-weighted pull–push velocity regression with reward-level fusion for multi-objective alignment.

#flow-matching#distillation#rl#vision#image+2
Computer Vision Papers·2026
Icon for item

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

Tianjie Ju, Zheng Wu +16

Provides a real-scale 3D Hong Kong sandbox to evaluate whether multimodal LLM agents can turn local street-view perception into sustained spatial action, supporting closed-loop first-person interaction, an interactive map, and controlled tests of grounding, long-range navigation, and robustness.

#multimodal#vision#long-horizon#benchmark#ai-agent+5
Large Language Model Papers·2026
Icon for item

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training

Tingyun Li, Wenfeng Feng +4·Abudukelimu Wuerkaixi, Guohua Liu +1

Decides when past post-training updates should be reused for autonomous LLM adaptation by introducing Boundary-Calibrated Intervention Transfer (BCIT). BCIT binds effects to source context, checks applicability and hard conflicts, and runs bounded trials to obtain current-state evidence—reducing harmful updates and improving equal-budget final-model quality.

#LLM#foundation-model#evaluation#agent-skills#ai-agent+2
Computer Vision Papers·2026
Icon for item

Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning

Hanyang Wang, Yimo Cai +15

Rewrites physical scenes as executable world programs (e.g., MuJoCo scene descriptions) and uses an agentic abductive loop to propose, execute, render, verify, and iteratively refine those programs from videos or text. Verified executable worlds supply scalable physical supervision for training vision–language models.

#paper#video#physics#code#research+5
Computer Vision Papers·2026
Icon for item

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Senqiao Yang, Chengyao Wang +14

Proposes VLAct, a representation-centric continued pre-training method for Vision-Language-Action models that preserves VLM priors and enforces cross-embodiment action semantics to turn limited robot trajectories into transferable visual-action representations; shows strong gains and sample efficiency on multiple VLA benchmarks using modest compute.

#robotics#vision#multimodal#foundation-model#paper+1
AI Agent Papers·2026
Icon for item

UI-Venus-2 Technical Report

Zhuohan Cai, Haoxing Chen +28·Ant Group, Inclusion AI

Autonomous multimodal GUI agent that executes natural-language interface tasks across mobile apps, web domains, and desktop OS. Expands environment coverage (170+ multilingual apps, 4,000+ web domains), uses function-grounded task generation and keypoint-based multi-model verification to produce reliable RL rewards for real-world deployment.

#multimodal#agent-skills#rl#qwen#benchmark+5
Speech Technology Papers·2026
Icon for item

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

Chengqian Ma, Wei Tao +2

Generates synchronized spoken dialogue and explicit full-body co-speech motion (facial expressions, hands, upper- and lower-body) end-to-end from the same hidden states, replacing the speech-then-motion cascade. Trains with a scalable pseudo-labeling pipeline (422,856 ranked pairs) and supports real-time inference (RTF 0.78) while matching teacher motion metrics within ~2%.

#speech#mocap#multimodal#qwen#ai-video+5
  • Previous
  • 1
  • More pages
  • 187
  • 188
  • 189
  • More pages
  • 210
  • Next