AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

AI Agent Papers·2026
Icon for item

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Zi-Han Wang, Zhengxi Lu +11

Converts sparse trajectory-level rewards into turn-level credit by aggregating token-level teacher–student log-probability gaps and recursively updating a Bayesian belief in log-odds; produces turn-wise reweighting for policy optimization without an extra critic or rollouts.

#distillation#RL#qwen#long-horizon#ai-agent+2
Computer Vision Papers·2026
Icon for item

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Zelong Sun, Jun Wang +4

Generates retrieval-centric Chain-of-Thought (RC-CoT) over initially retrieved candidates to improve unified multimodal retrieval via reranking or full-corpus re-retrieval with a dual-mode embedder. Trains an embedder–adviser framework (UniME-R1) using mined hard negatives, supervised learning, and retrieval-oriented reinforcement learning.

#multimodal#retrieval#embeddings#reasoning#RL+2
AI Video Papers·2026
Icon for item

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

Zongchuang Zhao, Xin Zhou +6

Uses video generation only as a training signal to co-train a pretrained video expert and a lightweight action expert, then discards the video branch at inference to produce a low-latency end-to-end driving planner; enhanced with RL for compositional driving rewards.

#video#flow-matching#RL#ai-video#robotics+3
Hugging Face
AI Video·2026
Icon for item

MiniMax-H3-Prompt-Rewriter-LoRA

lightx2v, ModelTC +2

Turns a short prompt plus aspect ratio and duration into a structured, shot-by-shot audio-video description for text-to-audio-video generation. A PEFT LoRA on Qwen3.6-27B that expands timing, camera motion, continuity, and synchronized diegetic/non‑diegetic sound; text-only and requires MiniMax-H3 + LightX2V to produce final AV.

#lora#qwen#ai-video#multimodal#prompt-engineering+3
Hugging Face
AI Dataset·2026
Icon for item

EgoPro

LightwheelAI

10,000-hour head-and-wrist egocentric dataset pairing synchronized head and wrist video with left/right 3D hand pose and optional full-body pose; provided in LeRobot/MCAP formats with episode-level semantic annotations and automated de-identification.

#multimodal#video#mcap#mocap#robotics+3
Large Language Model Papers·2026
Icon for item

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

Tao Feng, Fangxu Yu +10·University of Illinois Urbana-Champaign, University of Maryland, College Park +3

Frames LLM routing as a sequential decision process and introduces LLMRouter plus the xRouteBench benchmark to develop, evaluate, and deploy learned routing policies across heterogeneous LLMs, optimizing response quality versus inference cost.

#llm#benchmark#evaluation#ai-deploy#ai-inference+6
Hugging Face
AI Dataset·2026
Icon for item

LightwheelAI/EgoDemo

LightwheelAI

A small public sample of egocentric human demonstration video with synchronized 3D hand and body pose annotations for imitation learning and embodied-AI research. Delivered in Parquet and common multimodal packages (LeRobot, MCAP) for schema inspection before requesting gated access to larger EgoSuite releases.

#video#ai-video#multimodal#robotics#parquet+4
Hugging Face
AI Video·2026
Icon for item

Minimax-h3-Turbo

lightx2v·ModelTC

Packaged diffusers checkpoint of MiniMax H3 for image/text-to-short-video generation with native stereo audio; provided for direct use in diffusers image-to-video pipelines and aimed at easy integration into prototyping and production workflows.

#diffusers#huggingface#ai-video#video#multimodal+3
Hugging Face
AI Video·2026
Icon for item

MiniMax-H3_comfy

Kijai

Provides ComfyUI-compatible conversions and LoRA adapters of the MiniMax‑H3 video+audio generative model, with example presets and demo videos to run short stereo audio+video inference inside ComfyUI workflows.

#ai-video#multimodal#audio#lora#huggingface+3
Hugging Face
AI Audio·2026
Icon for item

MiniMax Music 3

MiniMax AI, SGLang-Omni +1

Generates complete songs (up to five minutes) from lyrics and a music description, producing 32 kHz stereo WAV with expressive vocals and long-range musical structure. Uses hierarchical LLMs fused with flow-matching/Flow-VAE synthesis for coherent arrangement and timbre; requires CUDA and integrates with Diffusers and SGLang-Omni.

#diffusers#pytorch#safetensors#flow-matching#llm+5
Hugging Face
AI Dataset·2026
Icon for item

RekaDaily-10k (raw)

RekaAI

Provides raw, unscripted first-person household video footage for training vision and embodied AI models. Released incrementally on Hugging Face in WebDataset shards with metadata parquets under Apache‑2.0; current raw tier contains ~7,834 hours (≈397k videos).

#video#ai-video#multimodal#vision#huggingface+3
Computer Vision Papers·2026
Icon for item

YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family

Xu Lin, WenJie Nie +5

Turns adapter placement for PEFT on YOLO-family real-time detectors into an auditable constraint-planning problem that emits budgeted target-module plans or calibrated refusals; shows planner-selected RS-LoRA improves mAP and cuts peak training memory in evaluated detectors.

#vision#lora#paper#github#ai-train
  • Previous
  • 1
  • More pages
  • 173
  • 174
  • 175
  • More pages
  • 201
  • Next