AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Category

Explore by categories

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All Categories

  • AI Leaderboard

  • AI Agent Tutorials

  • AI Coding Tutorials

  • AI Model

  • AI Agent Papers

  • Chatbot

  • AI Dataset

  • Machine Learning Foundation Books

  • AI Train

  • AI Deploy

  • AI Client

  • Machine Learning Foundation Papers

  • Machine Learning Foundation Tutorials

  • AI Image Demos

  • AI Agent

  • Large Language Model Tutorials

  • Large Language Model Papers

  • Machine Learning Engineering Papers

  • Computer Vision Tutorials

  • Computer Vision Papers

  • Natural Language Processing Papers

  • Reinforcement Learning Papers

  • Speech Technology Papers

  • AI API

  • AI Coding

  • AI Image

  • AI Video

  • MLOps

  • MCP Client

  • MCP Server

  • AI Video Papers

  • AI Audio

  • AI Others

  • AI Infra

  • Embodied AI

Hugging Face
AI Dataset·2026
Icon for item

Multi-Benchmark LLM Agent Traces

Exgentic

Provides 1,781 OpenTelemetry execution traces of LLM-powered agents across six benchmarks, including full conversations, token usage, timing, tool calls and model metadata—useful for performance analysis, agent-behavior research, and inference debugging.

#llm#ai-inference#mlops#agent-skills#ai-agent+4
Hugging Face
AI Dataset·2026
Icon for item

Stera-10M

FPV Labs

Open egocentric multimodal dataset for embodied AI and robot learning captured on commodity iPhone Pro: ~200 hours and ~10M RGB frames with LiDAR depth, ARKit 6‑DoF poses, IMU, two‑hand MANO mocap, room meshes, and hierarchical action captions.

#video#robotics#multimodal#ios#mobile+5
Hugging Face
AI Dataset·2026
Icon for item

Carbon Pretraining Corpus

HuggingFaceBio, GenerTeam +1

Provides 173M DNA/RNA sequences (≈1.1 trillion nucleotides) assembled specifically for pretraining genomic foundation models. Includes eukaryote, prokaryote, and mRNA configs plus a 10B‑token eukaryote subset for faster experiments; formatted for streaming and tokenized with Carbon's 6‑mer setup.

#huggingface#genomics#biology#foundation-model#transformers
Hugging Face
AI Dataset·2026
Icon for item

OpenCS2 - POV Renders

Julien Blanchon

Provides tick-aligned Counter-Strike 2 player POV video clips with per-tick inputs and world-state sidecars — near-lossless 1280×720@32fps video, per-player stereo audio, and parquet indexes for event/kill/round filtering; suited for RL, video classification and clip mining.

#video#ai-video#RL#audio#pandas+5
Hugging Face
AI Dataset·2026
Icon for item

DeepSeek v4 Pro Agent Traces

TeichAI

Contains 4,006 newline-delimited JSONL agent-session traces recording assistant responses and tool calls from deepseek/deepseek-v4-pro — includes a training-ready tools schema snapshot and helpers for conversion to SFT/distillation workflows.

#deepseek#agent-skills#ai-agent#huggingface#nlp+2
Hugging Face
AI Dataset·2026
Icon for item

5CD-AI/Viet-Handwriting-OCR-v2

5CD-AI

Labeled Vietnamese handwritten line images paired with text transcriptions for training and evaluating OCR/text-recognition models. Stored in Parquet (optimized) with a dataset size in the 10K–100K sample range, suitable for model training and benchmarking.

#huggingface#ocr#image#vision#nlp
AI Dataset·2026
Icon for item

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

Hao Liang, Qihan Lin +6

Provides a curriculum-aligned knowledge graph extracted from Chinese K–12 textbooks and accompanying benchmarks and training data to evaluate and train educational LLMs. Releases a 23,640-question multi-select benchmark and a 7,335-sample graph-guided training corpus with multimodal VQA pairs and the full construction pipeline.

#benchmark#multimodal#vision#llm#nlp+6
Hugging Face
AI Dataset·2026
Icon for item

Open-MM-RL

Shukla, Chinmayee, Patil, Saurabh +21

Multimodal STEM problem set for verifiable, answer-supervised training and RL: contains single-image, multi-panel, and multi-image PhD-level questions across physics, math, chemistry and biology. Each example has a deterministic ground-truth answer, enabling reward modeling and automated evaluation.

#multimodal#RL#science#physics#math+5
Hugging Face
AI Dataset·2026
Icon for item

TiniX Vietnam OCR Annual Financial Statements

TiniX AI

OCR-extracted Vietnamese annual financial reports (2015–2025) from 18,231 filings across 1,491 tickers — plain-text OCR outputs for document-QA, information extraction, VLM/RAG development. Contains only TXT OCR files; CC BY-NC 4.0 license.

#ocr#finance#nlp#multilingual#huggingface+2
Hugging Face
AI Dataset·2026
Icon for item

IndustryBench

alibaba-multimodal-industrial-ai, Songlin Bai +14

Multilingual benchmark for evaluating LLMs' industrial domain knowledge via 2,049 expert-curated QA pairs spanning 10 product verticals and four languages, with each item grounded to industry or national standards and an LLM-as-judge evaluation pipeline.

#huggingface#alibaba#multilingual#nlp#paper+3
Hugging Face
AI Dataset·2026
Icon for item

i1-captions (zlab-princeton)

Boya Zeng, Tianze Luo +5·Princeton University

Provides the full caption corpus used to train and ablate the i1 text-to-image model: 12 curated subsets with multiple caption variants (long/short, VLM-generated, rendered text) to enable reproducible training and captioning experiments.

#huggingface#ai-image#image#multimodal#diffusers+1
Hugging Face
AI Dataset·2026
Icon for item

CUA-Gym

xlangai, CUA-Gym Team

Pairs natural-language instructions with executable setup artifacts and Python reward functions to create verifiable computer-use agent tasks. Provides a Parquet task table for fast filtering plus a compressed archive of runnable task bundles; several web task endpoints are placeholders that require a local CUA-Gym-Hub deployment.

#huggingface#RL#agent-skills#pandas#python+1
  • Previous
  • 1
  • More pages
  • 13
  • 14
  • 15
  • More pages
  • 30
  • Next