AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

AI Agent Papers·2026
Icon for item

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

Yuling Shi, Zhensu Sun +4·Shanghai Jiao Tong University, Shanghai, China, Singapore Management University, Singapore +2

Predicts an LLM agent's final success or failure from partial execution traces and halts runs when outcomes are confident to save per-task compute. Uses LightGBM success/failure classifiers on behavioral, textual, and reference features; cuts 13–26% steps and up to 44% input tokens on benchmarks.

#evaluation#benchmarks#ai-agent#llm#agent-skills
Large Language Model Papers·2026
Icon for item

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

Zhiwei Zhang, Zechen Sun +7

Selectively admits dense token-level teacher supervision only after a prompt-level verifier audit, routing prompts that fail the audit to verifier-grounded trajectory supervision instead — reducing harmful updates from confidently wrong teachers and improving teacher GPU utilization.

#distillation#rl#LLM#paper#evaluation+3
Hugging Face
AI Dataset·2026
Icon for item

The First Fable 5.1 Reasoning Data

MoreThought

Contains 5,000 coding and chain-of-thought reasoning traces generated by Fable 5.1 — ~150M tokens of step-by-step programming CoT. Deduplicated and filtered for high quality; intended for supervised fine-tuning and distillation to improve reasoning in smaller models.

#huggingface#reasoning#thinking#distillation#coding+3
Hugging Face
AI Dataset·2026
Icon for item

MoreThought/Fable-5.1-Max-Reasoning-Filtered-1000x

MoreThought

A curated set of 1,000 high-quality chain-of-thought coding and reasoning traces generated by Fable 5.1, totaling ~30M tokens (109 MB). Designed for SFT/distillation to teach smaller models step-by-step programmatic reasoning and debugging.

#huggingface#distillation#sft#thinking#reasoning+4
Hugging Face
AI Video·2026
Icon for item

VDN-Minimax-H3

OpenVDN, MiniMax-AI

Adds a plug-and-play linear-attention branch and LoRA adapters to MiniMax-H3 to run text-to-video generation faster than real-time (near-lossless quality tradeoffs). Includes an optimized FP8 inference stack and a community license with regional restrictions.

#diffusers#safetensors#ai-video#video#fp8+5
Hugging Face
AI Dataset·2026
Icon for item

TikTok Videos, 4.5 Billion

kuben-developer

Provides 4.5 billion TikTok video records with captions, timestamps, music IDs and engagement counts for research; split across 27 zstd-compressed Parquet files (~289 GB) and sampled via TikTok's mobile API; released for research-use only with privacy and ToS caveats.

#parquet#video#ai-video#huggingface#polars+2
AI Video Papers·2026
Icon for item

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

Junchao Huang, Guian Fang +16·CUHK-SZ SLAI, NUS +6

Provides an end-to-end, reproducible foundation for camera‑controllable, long‑horizon video world models — converting 1.43M clips from 10 datasets into a unified canonical corpus and releasing data, pipelines, recipes, and weights. Introduces backbone‑native adaptation and a three‑stage training recipe to produce 5B–33B models that enable minute‑to‑hour rollouts after training on 5s sequences.

#video#ai-video#foundation#distillation#long-horizon+3
Large Language Model Papers·2026
Icon for item

Language Models Can Control Their Own Attention

Namgyu Ho, Huzama Ahmad +4·KAIST AI, Google DeepMind

Introduces Declarative Attention (DA), a zero-shot protocol that has LMs declare which parts of long context to attend to during chain-of-thought, letting the runtime build dynamic attention masks and skip most KV-cache reads. Produces large token savings (up to ~52% on Gemma-4-31B) with modest accuracy loss.

#llm#long-horizon#vllm#gemma#qwen+5
AI Agent Papers·2026
Icon for item

Iris: Climbing to the Search Frontier

Ziyuan Liu, Hengqi Liu +7·AllSpark Research

Presents two LLM-based search agents (Iris-mini and Iris-pro) trained by alternating supervised fine-tuning and reinforcement learning against live web search. Key features: web-graph-derived multi-hop tasks with entity abstraction, SFT–RL climbing, inference-time context management, and state-of-the-art open-source benchmark results.

#llm#RL#sft#moe#qwen+6
AI Agent Papers·2026
Icon for item

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Jie Wu, Zhenru Zhang +12·Qwen Team, Alibaba Group, Tsinghua University +2

Reconstructs executable terminal workspaces from recorded agent trajectories and synthesizes verifiable single- and multi-round coding tasks for agent training; it replays file operations, uses an LLM completion agent to fill missing files/dependencies, and verifies tasks with autogenerated test suites.

#terminal#coding-agents#qwen#sft#benchmarks+3
Large Language Model Papers·2026
Icon for item

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

Subham Sekhar Sahoo, Lingjie Chen +15·Institute for Foundation Models, Cerebras Systems

Uses lightweight discrete-diffusion adapters to draft blocks of tokens in parallel while keeping the original autoregressive model as the final arbiter, enabling lossless speedups (up to 3×) and compatibility with existing AR weights via a Diffusion Distillation phase.

#llm#distillation#lora#benchmarks#ai-inference+3
Large Language Model Papers·2026
Icon for item

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Zixuan Fu, Bingxiang He +11·Affiliation: University of Chinese Academy of Sciences, Affiliation: Tsinghua University +3

Studies on-policy distillation (OPD) at the data-minimal limit by training on a single query, measuring state coverage and alignment dynamics, and showing OPD is often data-overfed but algorithm-starved.

#distillation#LLM#NLP#paper#training-data+4
  • Previous
  • 1
  • More pages
  • 191
  • 192
  • 193
  • More pages
  • 213
  • Next