AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

AI Video Papers·2026
Icon for item

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Shuailei Ma, Jiaqi Liao +25

Pretrains a DiT-based Mixture-of-Experts video foundation model for embodied intelligence by augmenting internet videos with robot-centric footage and using a multi-dimensional reward system to prioritize physical realism and task completion while scaling MoE for better capacity vs. inference trade-offs.

#video#robotics#foundation-model#ai-video#multimodal+3
Hugging Face
AI Video·2026
Icon for item

LingBot-Video-MoE (30B-A3B)

Shuailei Ma, Jiaqi Liao +25

Generates videos from text and image+text prompts using a 30B Mixture-of-Experts model tuned for embodied intelligence; includes a refiner and structured prompt rewriter, and supports diffusers/SGLang runtimes with multi-GPU inference.

#ai-video#video#diffusers#huggingface#transformers+5
Hugging Face
AI Dataset·2026
Icon for item

Reasoning Corpus 5M

Qyrou·QyrouNnet-AI, SupraLabs

Provides ~5M model-generated reasoning chains (within 5k sequence length) with structured fields for supervised fine-tuning, reasoning distillation, and instruction tuning. Includes separate fields for prompt, reasoning trace, final answer and a ChatML view; streaming access recommended for large-scale use.

#reasoning#distillation#deepseek#qwen#gemma+7
Computer Vision Papers·2026
Icon for item

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Hongyu Qu, Jianzhe Gao +7

Reconstructs historical experience into latent memory tokens and weaves short- and long-term latent memories directly into vision-language-action reasoning to improve long-horizon robotic manipulation. Uses a four-part pipeline (curator, seeker, condenser, weaver) so memory participates natively in multimodal action formation.

#vision#robotics#multimodal#paper#embeddings+2
Computer Vision Papers·2026
Icon for item

Infinite Worlds with Versatile Interactions

Zelin Gao, Qiuyu Wang +18

Creates an open-ended interactive world simulator with an unbounded interaction horizon via causal pretraining, a distilled real-time runtime that drives 720p@60fps, a wider action/event repertoire, and a pilot–director agent split for behavior planning and environment synthesis.

#video#multimodal#foundation-model#ai-agent#agent-skills+4
Hugging Face
AI Dataset·2026
Icon for item

Reasoning Corpus 5M

SupraLabs

Provides ~5M tokens of chain-of-thought reasoning traces generated by many LLMs (DeepSeek, Qwen, Gemma, etc.) for training and evaluating reasoning SLMs — includes repo_id, tok_len, user, thought_trace, assistant and ChatML fields; sequences limited to 5k.

#reasoning#deepseek#qwen#code#pandas+4
AI Video Papers·2026
Icon for item

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

Cheng-De Fan, Chun-Wei Tuan Mu +5·National Yang Ming Chiao Tung UniversityTaiwan

Recovers and predicts RGB video from sparse event-camera streams by fine-tuning pre-trained video diffusion priors; jointly addresses reconstruction, long-horizon prediction, and bidirectional frame interpolation with mechanisms to reduce temporal drift and enforce interpolation consistency.

#video#ai-video#paper#vision#diffusers
Natural Language Processing Papers·2026
Icon for item

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

Yifan Zhou, Qihao Yang +15

Provides IdeaGene-Bench, a dataset and evaluation suite for scientific-lineage reasoning and lineage-grounded idea generation, representing papers as minimal, typed Idea Genome objects and GenomeDiffs that record inheritance, mutation, loss, import and novel insertion. Includes 1,961 lineage traces, IG-Exam (42 task types) and IG-Arena with a Population-Evolution Score for generation.

#evaluation#paper#LLM#NLP#science
AI Agent Papers·2026
Icon for item

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Zongxia Li, Zhongzhi Li +11

Provides a terminal-style benchmark of 46 long-horizon tasks decomposed into fine-grained graded subtasks to produce dense intermediate rewards and partial credit, enabling evaluation of long-horizon planning, long-context management, and iterative debugging. Tasks typically require hundreds of episodes and minutes-to-hours of execution; baseline evaluations report high token and episode consumption with low pass rates, highlighting evaluation headroom.

#evaluation#agent-skills#RL#terminal#paper+4
Hugging Face
AI Dataset·2026
Icon for item

Blood Pathology LIMS Environment

Yatin Taneja·IM Superintelligence

Simulates a hospital LIMS to benchmark agentic clinical reasoning: agents inspect demographics, medications, lab orders/results and then submit ICD‑10 diagnostic reports scored by deterministic, context‑aware graders. Ships as an OpenEnv/FastAPI runtime with 8 scenarios, step‑level rewards and trajectory capture for RL, tool‑use and evaluation.

#huggingface#RL#evaluation#agent-skills#json+5
Hugging Face
AI Dataset·2026
Icon for item

Digital Hospital Environment

Yatin Taneja·IM Superintelligence

Evaluates agents inside a structured hospital workflow via a downloadable FastAPI runtime that enforces role-specific tool permissions, evidence-before-treatment discipline, deterministic grading, dense process rewards, and full trajectory logging. Designed for RL, offline policy learning, multi-agent workflow research and process-supervision datasets; not for real patient care.

#evaluation#RL#agent-skills#ai-agent#reasoning+4
AI Agent Papers·2026
Icon for item

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

Zhekai Chen, Chengqi Duan +5·The University of Hong Kong (HKU) — Multimedia Lab

Evaluates proactive, multimodal agents on 400 bilingual real‑world tasks across five capability axes (Skill Usage, Exploration, Long‑Context Reasoning, Multimodal Understanding, Cross‑Platform Coordination) using live Docker‑based, stepwise closed‑loop evaluation to separate base model skills from framework design.

#agent-skills#multimodal#evaluation#LLM#multilingual+5
  • Previous
  • 1
  • More pages
  • 155
  • 156
  • 157
  • More pages
  • 190
  • Next