AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Category

Explore by categories

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All Categories

  • AI Leaderboard

  • AI Agent Tutorials

  • AI Coding Tutorials

  • AI Model

  • AI Agent Papers

  • Chatbot

  • AI Dataset

  • Machine Learning Foundation Books

  • AI Train

  • AI Deploy

  • AI Client

  • Machine Learning Foundation Papers

  • Machine Learning Foundation Tutorials

  • AI Image Demos

  • AI Agent

  • Large Language Model Tutorials

  • Large Language Model Papers

  • Machine Learning Engineering Papers

  • Computer Vision Tutorials

  • Computer Vision Papers

  • Natural Language Processing Papers

  • Reinforcement Learning Papers

  • Speech Technology Papers

  • AI API

  • AI Coding

  • AI Image

  • AI Video

  • MLOps

  • MCP Client

  • MCP Server

  • AI Video Papers

  • AI Audio

  • AI Others

  • AI Infra

  • Embodied AI

Hugging Face
AI Dataset·2024
Icon for item

SceneFun3D

Alexandros Delitzas, Ayca Takmaz +4·ETH Zurich, Google +2

Provides point-accurate annotations of interactive parts in high-resolution indoor laser-scan point clouds, plus affordance labels, motion axes and natural-language task descriptions; includes aligned iPad RGB-D video slices with 2D projections for multimodal research.

#robotics#vision#depth#multimodal#huggingface+1
Hugging Face
AI Dataset·2024
Icon for item

Open-PerfectBlend

mlabonne

A mixed instruction dataset for SFT and RLHF research that combines chat, math, code and instruction-following samples from multiple public datasets under an Apache-2.0-compatible license; intended for instruction tuning and evaluation.

#huggingface#nlp#LLM#math#code+1
Hugging Face
AI Dataset·2025
Icon for item

Humanity's Last Exam

Center for AI Safety (cais), Scale AI

Multi‑modal closed-ended academic benchmark with 2,500 multiple-choice and short-answer exam questions spanning math, natural sciences, and humanities for automated grading. Curated by subject-matter experts, released under MIT, and includes a canary string to help prevent dataset leakage into model training.

#huggingface#multimodal#image#nlp#pandas+1
Hugging Face
AI Dataset·2025
Icon for item

The AI CUDA Engineer Archive

Sakana AI

A curated dataset of ~30,000 CUDA kernels generated by an agentic pipeline, including reference PyTorch implementations, runtime metrics, NCU/Torch/Clang-Tidy profiles, error messages and correctness labels — released under CC-BY-4.0 for model fine-tuning and offline RL/optimization research.

#code#pytorch#pandas#huggingface#nvidia+2
Hugging Face
AI Dataset·2025
Icon for item

SMOL

google, Isaac Caswell +3

Provides professionally translated parallel corpora and a multilingual lexicon across 100+ low-resource languages for training and evaluating multilingual MT and NLP models. Includes SmolDoc, SmolSent, GATITOS, and factuality annotations; licensed CC-BY-4.0.

#google#huggingface#multilingual#translation#NLP
Hugging Face
AI Dataset·2025
Icon for item

Interaction2Code

whale99

A benchmark dataset for evaluating MLLM-driven interactive webpage code generation: provides prototyping screenshots, action.json interaction metadata, and example generation scripts across 127 webpages and 374 interactions to test dynamic UI-to-code capabilities.

#multimodal#image#code#github#LLM+3
Hugging Face
AI Dataset·2025
Icon for item

Ultra-FineWeb

openbmb

High-quality, efficiently verified and filtered web corpus for LLM pretraining — supplies ~1 trillion English tokens and ~120 billion Chinese tokens with English/Chinese Parquet splits. Designed for large-scale pretraining experiments and data-filtering research.

#LLM#huggingface#nlp#multilingual#transformers+2
Hugging Face
AI Dataset·2025
Icon for item

olmOCR-bench

Jake Poznanski, Jon Borchardt +7·Allen Institute for Artificial Intelligence (AI2), AllenNLP / olmOCR team

Benchmark for evaluating OCR systems that convert PDFs and scans into Markdown and structured text: 1,403 PDFs and 7,010 unit tests covering text presence/absence, reading order, tables, and math formula accuracy. Diverse sources and ODC-BY-1.0 license for research use.

#ocr#evaluation#vision#huggingface#paper+1
Hugging Face
AI Dataset·2025
Icon for item

OpenCodeInstruct

NVIDIA

Provides 5 million instruction–response pairs for supervised fine-tuning of code LLMs, with inputs, outputs, unit tests, and automated LLM judgments. Uses hybrid automated/synthetic generation and is released under CC BY 4.0 for large-scale SFT workflows.

#nvidia#huggingface#code#ai-train#ai-coding+3
Hugging Face
AI Dataset·2025
Icon for item

FLARE-MedFM/PancancerCTSeg

FLARE-MedFM, Hugging Face

Pan-cancer CT segmentation dataset for training and benchmarking medical-image segmentation models — packaged as a Hugging Face dataset with an estimated 10k–100k samples and linked arXiv references. Designed for model development and reproducible benchmarking; non-commercial license applies.

#segmentation#vision#image#huggingface#ai+1
Hugging Face
AI Dataset·2025
Icon for item

SWE-bench Verified

SWE-bench

A human-verified subset of 500 SWE-bench test cases for evaluating models that resolve GitHub issues into PRs using unit-test verification. Contains problem statements and base commits (pre-fix) for reproducible unit-test based evaluation; suitable for benchmarking code-fix and issue-resolution capabilities.

#github#nlp#python#ai-leaderboard#ai-rank+1
Hugging Face
AI Dataset·2025
Icon for item

ChartGalaxy

Zhen Li, Duan Li +10

Provides 1.7M+ synthetic and real infographic charts paired with their tabular data for training and evaluating multimodal models on infographic understanding, chart-to-table extraction, chart code generation, and example-based chart synthesis.

#multimodal#vision#image#parquet#pandas+3
  • Previous
  • 1
  • 2
  • 3
  • More pages
  • 27
  • 28
  • Next