AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

AI Agent Papers·2026
Icon for item

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

Nahyun Lee, Dongkeun Yoon +13

A benchmark for evaluating web-browsing agents in Korean contexts, composed of 400 tasks (300 manually verified by native speakers). Includes a human-verified split and an adversarial synthetic split to probe failure modes; reveals large performance gaps for both frontier and Korean models.

#paper#NLP#ai-agent#agent-skills#multilingual+1
GitHub
AI Client·2026
Icon for item

AgentRecall

zszz3, Blue-Berrys +5

Centralizes indexing and management of local AI coding-agent sessions so you can search, view full context, migrate, resume, and restore conversations across agents and devices. Supports extensible local sources, AI summaries, optional Supabase sync, and Skills management.

#ai-coding#coding-agents#claude-code#codex#electron+8
Hugging Face
AI Dataset·2026
Icon for item

datacurve/deep-swe

datacurve

Benchmark dataset for evaluating long-horizon coding agents and software-engineering tasks, containing English code and tabular metadata in Parquet format; small scale (<1K examples) for fast prototyping and evaluation.

#benchmark#coding#software-engineering#coding-agents#long-horizon+5
Computer Vision Papers·2026
Icon for item

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

Yifei Li, Pengyiang Liu +5

Evaluates multimodal LLMs on streaming egocentric video for spatial intelligence using 1,680 human-annotated questions across 348 videos; organizes tasks into four hierarchical levels (perception → tracking → simulation → allocentric mapping) and highlights allocentric mapping as the main bottleneck.

#multimodal#video#robotics#vision#paper+3
Computer Vision Papers·2026
Icon for item

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

Yucheng Zhou, Wei Tao +2

Studies when and how to combine visual future rollouts from world models with abstract reasoning in multimodal LLMs. Proposes PF-OPSD — a teacher-student distillation that uses ground-truth future videos during training — and evaluates on two human-verified benchmarks, improving accuracy ≈10% while improving robustness to noisy rollouts.

#paper#multimodal#vision#LLM#code+1
Hugging Face
AI Video·2026
Icon for item

Echo-LongVideo (JoyAI-Echo)

Echo Team @ Joy Future Academy, JD, jdopensource

Generates minute-level, multi-shot synchronized audio+video from a single text prompt, using a paired cross-modal memory to preserve character appearance and voice across shots. Uses DMD-distilled few-step inference for ~7.5× speedup; requires high-GPU memory and is released under the LTX-2 community license.

#ai-video#video#audio#multimodal#huggingface+3
Large Language Model Papers·2026
Icon for item

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

Ziyan Liu, Xueda Shen +8

Learns fine-grained preferences over sub-trajectories to identify and penalize redundant steps in long chain-of-thoughts, letting models "fold" reasoning chains into concise paths; reports ~56% token reduction on DeepSeek-R1-Distill-Qwen-7B while keeping accuracy.

#LLM#paper#deepseek#RL#transformers
Computer Vision Papers·2026
Icon for item

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking

Zekun Qi, Xuchuan Chen +11

Trains a GPT-style causal Transformer on a 2-billion-frame retargeted motion corpus to enable zero-shot whole-body motion tracking and control. By scaling both data and model capacity, it tracks highly dynamic behaviors while generalizing to unseen motions; accepted to CVPR 2026.

#robotics#vision#transformers#foundation-model#paper+1
Computer Vision Papers·2026
Icon for item

Qwen-Image-Flash: Beyond Objective Design

Tianhe Wu, Kun Yan +22

Explores how training recipe — data composition, teacher guidance, and task mixture — shapes few-step distillation for text-to-image generation and instruction-guided image editing; introduces Qwen-Image-Flash and empirical findings that training pipeline organization matters as much as distillation objectives.

#vision#multimodal#foundation-model#paper#ai-image+1
Hugging Face
AI Model·2026
Icon for item

MiniMax-M3

MiniMaxAI

Native multimodal model for image/text/video→text tasks with million‑token context support. Uses a sparse-attention operator to cut long‑context compute and latency, and targets agentic, coding, and long-horizon conversational workloads.

#multimodal#transformers#vllm#ai-agent#foundation-model+3
AI Agent Papers·2026
Icon for item

TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration

Soyeong Jeong, Jinheon Baek +2

Enables agents to proactively discover multiple hidden problems in a user context and pair each with supporting evidence and concrete actions. Uses iterative discovery (batch rounds conditioned on prior finds) and reusable "thought templates" to expand coverage and ground claims.

#paper#agent-skills#LLM#NLP#ai
Hugging Face
AI Dataset·2026
Icon for item

Nemotron-Personas-El-Salvador

Rodrigo Malossi, Andre Manoel +9

Provides ~1M synthetic Salvadoran‑Spanish personas (148k records, ~300M tokens) grounded in 2024 census distributions for demographics, occupations and locations; intended for training/evaluating localized LLMs and synthetic-data workflows. CC BY 4.0, adults only.

#huggingface#nvidia#nlp#multilingual#llm+2
  • Previous
  • 1
  • More pages
  • 130
  • 131
  • 132
  • More pages
  • 172
  • Next