AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Hugging Face
AI Dataset·2026
Icon for item

The Stack v3 (HuggingFaceCode/stack-v3-train)

Anton Lozhkov, Hugo Larcher +3·Hugging Face, BigCode

Provides a near-deduplicated, quality-filtered 15.9 TB training subset of GitHub source code grouped by repository, with inline UTF‑8 file contents and repo metadata for pre-training and analysis of code LLMs; cutoff Aug 7, 2025, ODC-By license.

#code#huggingface#github#llm#ai-train+3
AI Agent Papers·2026
Icon for item

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Shuqi Lu, Chaofan Li +21·Beijing Academy of Artificial Intelligence (BAAI)

Alternates targeted research and constraint-wise audits to recursively improve long-horizon answers: an inner loop gathers evidence and drafts solutions, an outer loop audits unresolved claims and launches focused follow-ups. Trains 4B dense and 122B-A10B MoE agents with long-horizon RL and agentic mid-training, outperforming comparable-scale baselines on multi-step research benchmarks.

#agent-skills#ai-agent#RL#reasoning#llm+1
Computer Vision Papers·2026
Icon for item

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Xu Wang, Kaixiang Yao +5·Zhejiang University

Evaluates spatial cognition of image-generation models by eliciting protocol-constrained visual answers and parsing pixel outputs into structured predictions compatible with existing metrics. Introduces the ProVisE framework and SpatialGen-Bench (470 samples) to compare image-generation models and text-output VLMs on unified spatial tasks.

#vision#multimodal#benchmark#evaluation#image+1
Computer Vision Papers·2026
Icon for item

Visual Contrastive Self-Distillation

Yijun Liang, Yunjie Tian +5

Converts image-content removal into a contrastive on-policy self-distillation signal: the EMA teacher produces next-token distributions with and without image content, uses their log-probability differences to sharpen visual-grounded candidates, and distills that full-distribution target into the student—no external teacher or extra inference cost.

#distillation#multimodal#vision#qwen#paper+3
Hugging Face
AI Model·2026
Icon for item

KAT-Coder-V2.5-Dev

Kwaipilot, KwaiKAT Team

A text-only open-weight MOE code model (35B total, 3B active) fine-tuned with SFT+RL for agentic coding; achieves strong agentic-code benchmarks, supports 262k context and deployment via Transformers/vLLM; vision weights are not included.

#qwen#transformers#vllm#huggingface#ai-agent+6
Hugging Face
AI Dataset·2026
Icon for item

PerceptionBench

Moonshot AI

Evaluates atomic visual perception of multimodal LLMs using 3,000 short visual questions that isolate ten perceptual skills. Built from an error taxonomy across 42 benchmarks, capability-balanced and accompanied by a model leaderboard.

#benchmark#vision#evaluation#image#ocr+5
Large Language Model Papers·2026
Icon for item

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

Siyuan Huang, Pengyu Cheng +11

Presents Skill Self-Play (Skill-SP), a co-evolutionary training loop where a proposer, solver, and dynamic skill controller generate, solve, and verify tasks conditioned on reusable skills — balancing verifiable execution with open-ended task diversity to boost LLM tool-use and reasoning.

#LLM#agent-skills#RL#reasoning#paper+3
Hugging Face
AI Model·2026
Icon for item

Mage-Flow

Comfy-Org, Microsoft

Provides repackaged Mage-Flow model files formatted for ComfyUI, including multiple diffusion variants (bf16, int8, turbo, edit), a Qwen text encoder and a VAE — organized in a ComfyUI directory layout for drop-in use.

#huggingface#diffusers#ai-image#AIGC#microsoft+2
Large Language Model Papers·2026
Icon for item

Scaling Native Multimodal Pre-Training From Scratch

Haoyuan Wu, Aoqi Wu +4

Empirically studies how transformer-based native multimodal pre-training scales under fixed compute, deriving compute- and data-allocation power laws and an efficiency frontier for model size, token count, and data mixture; evaluates cross-modal transfer and multimodal in-context learning.

#multimodal#foundation-model#vision#llm#transformers+1
AI Agent Papers·2026
Icon for item

StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

Yan Yang, Xiangru Jian +8

Drives long‑horizon desktop agents by reading and manipulating program state (files, DOM, backends) instead of relying on screenshots. The main agent uses code for actions and structural verification while a lightweight GUI subagent handles rare screenshot-click steps, improving success rates and lowering per-task cost versus screenshot-only approaches.

#paper#ai-agent#long-horizon#coding-agents#agent-skills+3
Natural Language Processing Papers·2026
Icon for item

LAMAR: An Open Language-Aware Multilingual Alignment Reranker

Seongtae Hong, Youngjoon Jang +3

Reranks multilingual retrieval candidates to favour documents that are both semantically relevant and written in the same language as the query, using English-anchored relevance distillation and preference alignment; excels in language-coherence tests while remaining competitive on standard multilingual reranking benchmarks.

#multilingual#NLP#RAG#distillation#benchmark+2
Hugging Face
AI Audio·2026
Icon for item

VibeVoice-ASR-BitNet

Microsoft Research

Multilingual, real-time ASR for edge CPUs that uses heterogeneous quantization to reduce model size (4.62→1.58 GB) and lower inference latency. Trades some accuracy for 1.6–2.3× faster inference vs. Whisper.cpp and real-time capability on a few CPU threads, making it suitable for memory- and compute-constrained on-device transcription.

#huggingface#multilingual#ASR#stt#audio+5
  • Previous
  • 1
  • More pages
  • 164
  • 165
  • 166
  • More pages
  • 195
  • Next