AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Hugging Face
AI Dataset·2026
Icon for item

pixelgpt-24x24-20k

unstonio·unston.io

Curated set of 20,000 native 24×24 pixel-art sprites with two-level semantic taxonomy labels for tiny text-to-image and discrete visual modeling. Rebalanced, rights-conscious subset with ≤5 colors per sprite and stratified train/val/test splits.

#huggingface#ai-image#image#pandas#polars
Hugging Face
AI Model·2026
Icon for item

microsoft/Mage-VL

Senqiao Yang, Kaichen Zhang +20·Microsoft, Microsoft Research

Delivers image and video understanding plus a built-in event‑gated streaming gate — a unified 4B multimodal foundation model that uses codec-aligned tokenization to cut visual tokens by >75% and yield up to 3.5× wall‑clock inference speedup for streaming and long‑horizon video tasks.

#multimodal#video#vision#qwen#transformers+8
Hugging Face
AI Dataset·2026
Icon for item

Indic DiarBench

Deovrat Mehendale, Aditya Mehndiratta +3·Sarvam AI, AI4Bharat +1

Benchmark for joint speaker diarization and speaker-attributed ASR across all 22 scheduled Indian languages, providing ~108 hours of human-corrected, time-aligned, speaker-attributed transcripts. Includes near-field, far-field and in-the-wild recordings with code-mixing and speaker overlap.

#ASR#speech#multilingual#huggingface#benchmark+6
Large Language Model Papers·2026
Icon for item

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Qinsi Wang, Jing Shi +9

Transforms open-ended LLM optimization into self-verifiable reinforcement learning by turning tasks into proxy environments that produce deterministic, rule-based rewards. Proposes RLSVR and SpyRL — an information-asymmetric self-play scheme where agents vote to identify a preassigned spy, yielding verifiable rewards without human annotation. Demonstrated on summarization, creative writing and mathematical reasoning.

#RL#rl#LLM#reasoning#paper+5
Hugging Face
AI Dataset·2026
Icon for item

LLM Self Identification

SupraLabs

A compact dataset of prompt templates and examples designed to teach LLMs to consistently report model identity fields (model ID, name, creator, family, architecture, parameter count, knowledge cutoff). Includes regex markers, usage guidance, example replacements, and a small personalization script for fine‑tuning or runtime substitution.

#llm#huggingface#prompt-engineering#json#pandas+3
Hugging Face
AI Dataset·2026
Icon for item

FinanceGym

Embodied Analysis

Provides a public test split of multimodal financial GUI interaction examples for evaluating agents that convert instructions and screenshots into grounded UI actions. Includes step-level screenshots, dialogue history, an OpenAI-style computer_use tool schema, and JSON next-action references; training data available on request.

#finance#multimodal#benchmark#evaluation#ai-agent+3
Hugging Face
AI Dataset·2026
Icon for item

r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation

r0b0tlab, Alibaba Cloud Model Studio +5

Provides a 57,937-row, quality-filtered multi-teacher SFT distillation corpus combining outputs from Qwen3.8-Max, GLM-5.2 and Kimi K3 across math, code, reasoning, tool-use and dialogue. Includes 24 parquet training views (including a pre-tokenized GLM-4.7 view), configurable sampling weights (sft_balanced), and explicit tool-call trajectories for agent training.

#distillation#qwen#kimi#reasoning#code+5
Reinforcement Learning Papers·2026
Icon for item

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

NeoteAI Team, Fudan TEAI Team

Enables tactile-aware robot manipulation by pretraining a vision–tactile–language–action foundation model and improving offline policies with ALTER. Combines large-scale NeoData visuo-tactile pretraining, a latent tactile pathway for predictive touch signals, and advantage‑conditioned offline RL for contact-rich tasks.

#robotics#rl#foundation-model#multimodal#vision+3
Hugging Face
AI Dataset·2026
Icon for item

LocateAnything-Data

NVEagle (NVlabs / NVIDIA)

Consolidated dataset of detection, visual grounding and pointing annotations with indexed WebDataset image shards and Megatron‑Energon training metadata. Covers diverse visual domains (COCO, RefCOCO, driving, GUI, documents) and uses a normalized spatial grid for cross‑domain vision–language grounding training.

#multimodal#vision#ocr#nvidia#huggingface+3
AI Agent Papers·2026
Icon for item

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

Yunlong Lin, Zixu Lin +24

Lets canvas-native agents plan, generate, edit, and organize long-horizon multimodal creative projects by representing artifacts, versions, and actions as typed canvas nodes and links. Uses a three-layer design (canvas state, protocol bridge, agent runtime) so agents act within an inspectable, editable project state.

#multimodal#ai-agent#agent-skills#long-horizon#ai-workflow+2
Hugging Face
AI Model·2026
Icon for item

Kimi K3

reteetzad·Moonshot AI

An open-weight LLM focused on deep reasoning, native agentic tool use, and repository-scale code understanding — Mixture-of-Experts architecture with an extended context window and permissive licensing.

#kimi#foundation-model#LLM#agent-skills#reasoning+2
Hugging Face
AI Model·2026
Icon for item

Inkling-Small

Thinking Machines Lab

Generates text from text, image, or audio inputs using a native multimodal, Mixture-of-Experts autoregressive transformer (276B total / 12B active) with up to 1M-token context; targeted at conversational, agentic, coding and multimodal applications.

#transformers#multimodal#llm#huggingface#benchmarks+4
  • Previous
  • 1
  • More pages
  • 165
  • 166
  • 167
  • More pages
  • 196
  • Next