AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

AI Agent Papers·2026
Icon for item

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents

Jia Deng, Yimeng Chen +10

Synthesizes shortcut-resistant search tasks to train deep search agents by controlling four shortcut risks across entity selection, evidence-graph construction, question formulation, and adversarial refinement. Produces training trajectories with longer pre-answer search and fewer shortcut patterns; code will be released on GitHub.

#paper#github#ai-agent#agent-skills#deepseek+2
Hugging Face
AI Audio·2026
Icon for item

ZONOS2

Gabriel Clark, Sofian Mejjoute +3·Zyphra

Multilingual, low-latency text-to-speech model for speech generation and zero-shot voice cloning. Uses an MoE backbone with ECAPA-TDNN speaker embeddings, supports audio prefixes, fine-grained prosody/emotion controls and 44.1kHz output; optimized for Linux + NVIDIA GPUs.

#tts#audio#multilingual#voice#huggingface+4
AI Video Papers·2026
Icon for item

OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

Jiwen Liu, Shujuan Li +9

Encodes and clones camera motion from reference videos to generate multi-shot videos — uses a visual "camera grid" to represent camera parameters, trains on million-scale grid–video pairs, and employs a hierarchical prompt-expansion agent to coordinate camera, subject, and action control for multimodal diffusion models.

#video#multimodal#ai-video#vision#prompt-engineering+2
AI Video Papers·2026
Icon for item

Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

Yuho Lee, Jisu Shin +6·KAIST, Qualcomm AI Research (Qualcomm Korea)

Proposes chunk-level multimodal retrieval and chunk-adaptive reranking for retrieval-augmented generation on long egocentric videos; introduces V-RAGBench to decouple retrieval vs. generation evaluation and CARVE to run parallel retrievers and select per-chunk configurations.

#RAG#video#multimodal#evaluation#vision+2
AI Agent Papers·2026
Icon for item

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

Jundong Xu, Qingchuan Li +12

Benchmarks evolving environments as sequences of progressive updates and introduces EvoMem, a patch-based memory that records structured update histories so LLM agents can reason about environment evolution. Demonstrates measurable gains on EvoArena and other benchmarks.

#LLM#ai-agent#agent-skills#paper#nlp
Computer Vision Papers·2026
Icon for item

SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

Seokju Cho, Ryo Hachiuma +9

Provides a training-free, code-as-action framework that lets VLM-backed agents write and run stateful Python cells to compose perception and geometry primitives for open-ended 3D/4D spatial reasoning. Demonstrates consistent gains across 20 benchmarks and multiple VLM backbones.

#vision#multimodal#ai-agent#agent-skills#paper
Hugging Face
AI Model·2026
Icon for item

Qwopus-3.6-27B-Coder

Jackrong

A quantized 27B coder LLM fine-tuned for repository-level code generation, multi-turn tool calling, and agentic workflows — packaged for local GGUF/llama.cpp deployment with MTP speculative decoding and trace-inversion SFT. Optimized for developer tooling; experimental and not fully safety-validated.

#huggingface#llm#transformers#ai-coding#ai-agent+5
Large Language Model Papers·2026
Icon for item

MiniMax Sparse Attention

Xunhao Lai, Weiqi Xu +9

Implements a blockwise sparse attention (MiniMax Sparse Attention) that scores and Top-k selects key-value blocks per Grouped Query Attention group to enable attention over million-token contexts. Paired with an exp-free Top-k GPU kernel and KV-outer sparse execution, it reduces per-token attention compute and yields large prefill/decoding speedups.

#paper#llm#multimodal#github#huggingface+4
Hugging Face
AI Dataset·2026
Icon for item

claude-fable-5 Agent Traces

armand0e

Anonymized agent-trace JSONL capturing conversations, tool calls, and function schemas from Claude fable-5 (Claude Code) runs — packaged for fine-tuning, distillation, and building tool-aware assistants; compatible with teich for conversion to OpenAI-style chats.

#huggingface#claude#claude-code#agent-skills#ai-agent
Hugging Face
AI Dataset·2026
Icon for item

ChingMu Robot Motion Dataset

CMRobot, Shanghai Chingmu Vision Technology Co., Ltd.

Provides 1000+ hours of high-precision optical motion-capture for humanoid robotics and embodied AI, including full-body skeleton, 20+DoF hands, object 6D, and multi-view video at 120 Hz. Sub-mm spatial accuracy, BVH/CSV/NPZ outputs and Unitree G1 retargets; ideal for imitation learning and sim-to-real, with some raw captures gated by license.

#robotics#multimodal#huggingface#ai-train#video
Hugging Face
AI Dataset·2026
Icon for item

KSAFE-MM

K-intelligence

Benchmark for evaluating multimodal LLM safety in Korean cultural contexts — includes KSAFE-MM-G which localizes global safety queries into Korean scenarios and KSAFE-MM-C which targets culture-specific visual-textual vulnerabilities. Provides curated image–text pairs and jailbreak-style prompts to reveal both unsafe behaviors and over-refusal.

#multimodal#vision#image#evaluation#huggingface+3
Hugging Face
AI Model·2026
Icon for item

Rio 3.5 Open 397B

IplanRIO (prefeitura-rio)

A post-trained Mixture-of-Experts multimodal LLM with ~397B total (≈17B active) and a 1,010,000-token context for image-text-to-text and conversational tasks. Integrates SwiReasoning to switch between latent and explicit reasoning; MIT-licensed and optimized for Portuguese/English research and on-prem inference.

#transformers#multilingual#multimodal#huggingface#vllm+4
  • Previous
  • 1
  • More pages
  • 138
  • 139
  • 140
  • More pages
  • 178
  • Next