AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Category

Explore by categories

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All Categories

  • AI Leaderboard

  • AI Agent Tutorials

  • AI Coding Tutorials

  • AI Model

  • AI Agent Papers

  • Chatbot

  • AI Dataset

  • Machine Learning Foundation Books

  • AI Train

  • AI Deploy

  • AI Client

  • Machine Learning Foundation Papers

  • Machine Learning Foundation Tutorials

  • AI Image Demos

  • AI Agent

  • Large Language Model Tutorials

  • Large Language Model Papers

  • Machine Learning Engineering Papers

  • Computer Vision Tutorials

  • Computer Vision Papers

  • Natural Language Processing Papers

  • Reinforcement Learning Papers

  • Speech Technology Papers

  • AI API

  • AI Coding

  • AI Image

  • AI Video

  • MLOps

  • MCP Client

  • MCP Server

  • AI Video Papers

  • AI Audio

  • AI Others

  • AI Infra

  • Embodied AI

Hugging Face
AI Dataset·2026
Icon for item

DROID: Distributed Robot Interaction Dataset

Alexander Khazatsky, Karl Pertsch +18·NVIDIA Corporation, Stanford University +15

Large-scale in-the-wild robot manipulation dataset with ~76K teleoperated trajectories (~350 hours) that provides synchronized multi-view video, depth, camera calibration, robot state/action traces, and natural-language task instructions to train and evaluate manipulation policies and dynamics models. Collected across 564 scenes, 86 tasks, 52 buildings, on a uniform Franka Panda hardware stack and released in LeRobotDataset v3.0 format (≈707 GB, OpenMDW1.1).

#robotics#video#vision#depth#parquet+5
Hugging Face
AI Dataset·2026
Icon for item

GLM-5.1-1000000x

Kassadin88

Provides 1,003,589 full chain-of-thought reasoning traces and final answers generated by GLM-5.1, split into main/Math/PHD-Science/Multilingual-STEM subsets. Useful for instruction-tuning, supervised fine-tuning, and reasoning experiments; released under Apache-2.0.

#huggingface#llm#nlp#math#science+1
Hugging Face
AI Dataset·2026
Icon for item

Türkçe Atlas — Instruct SFT

AlicanKiraz0

Provides 336,146 Turkish instruction-following chat examples (system→user→assistant) for supervised fine-tuning; single train split (no validation/test), reported MIT license, diverse tasks (rewrites, summarization, QA) and a uniform system prompt that may bias model behavior.

#huggingface#nlp#llm#chatbot#json+3
Hugging Face
AI Dataset·2026
Icon for item

RSRCC

R. Kazoom, Y. Gigi +4

Provides paired before/after satellite images with question–answer annotations for semantic change understanding. Includes Yes/No and multiple-choice formats, delivered in Hugging Face datasets (streaming-friendly), suited for remote-sensing multimodal VQA and semantic change captioning research.

#vision#multimodal#google#huggingface#paper+1
Hugging Face
AI Dataset·2026
Icon for item

Usenet Corpus 1980–2013

OwnedByDanes

Provides deduplicated, sanitized Usenet posts (1980–2013) for language-model pretraining and linguistic research. Includes a ~103.1B-token full corpus (408M posts) with freely downloadable sample files; full corpus access requires a license and PII redaction was applied.

#huggingface#nlp#multilingual#ai-train#privacy+1
Hugging Face
AI Dataset·2026
Icon for item

SWE-ZERO 12M Trajectories

AlienKevin

Provides ~12.29M execution‑free agentic coding trajectories (≈112B tokens) sampled from 122K GitHub PRs to mid‑train code and agent models. Uses bash-only actions (grep, git, sed, etc.) so it scales without Docker; trajectories are unverified and intended for mid-training rather than final SFT.

#huggingface#agent-skills#ai-agent#code#vllm+3
Hugging Face
AI Dataset·2026
Icon for item

Open-SWE-Traces

Wasi Uddin Ahmad, Nikolai Ludwig +2·NVIDIA

Provides 207k+ LLM-generated agent trajectories of code edits and tool interactions for training and evaluating software-engineering agents. Collected via OpenHands and SWE-agent using Qwen3.5-122B and MiniMax-M2.5, multilingual across nine languages and released under CC BY 4.0.

#nvidia#huggingface#code#ai-agent#agent-skills+4
Hugging Face
AI Dataset·2026
Icon for item

GLM-5.1-Reasoning-1M-Cleaned

Jackrong, Kassadin88

Provides a cleaned, SFT-ready collection of ~746k GLM-5.1 reasoning traces for instruction tuning and reasoning distillation. Normalizes varied chain-of-thought formats into a single conversations/input/output schema and preserves four focused subsets (main, PHD-Science, Multilingual‑STEM, Math).

#huggingface#llm#nlp#math#science+1
Hugging Face
AI Dataset·2026
Icon for item

SWE-Hero Trajectories (nvidia/SWE-Hero-openhands-trajectories)

NVIDIA

Provides 34k execution-style agent trajectories (11,766 issues) for supervised fine-tuning of code-focused LLMs. Each instance includes multi-step interactions, tool-call records, and final unified diffs; generated with Qwen3-Coder and released under permissive licenses for commercial use.

#nvidia#huggingface#code#ai-agent#ai-coding+2
Hugging Face
AI Dataset·2026
Icon for item

AtomBlock-WebUI

Zhihao Nan, Yiming Cheng +2

About 9,700 synthetic full-page web screenshots with YOLO-format, pixel-aligned bounding boxes for 14 UI element classes, generated by LLM-augmented HTML and Playwright DOM extraction. Includes CC3M image injection to reduce visual gap; released for non-commercial research (CC BY-NC-SA 4.0).

#vision#multimodal#huggingface#ai-image#llm
Hugging Face
AI Dataset·2026
Icon for item

Go-Code-Large

ajibawa-2023

Provides 316,427 Go source-code samples in JSONL focused on concurrency and backend idioms, enabling fine-tuning and evaluation of code models for completion, summarization, and static-analysis tasks.

#huggingface#ai-train#ai-coding#llm#ai-development
Hugging Face
AI Dataset·2026
Icon for item

lordx64/reasoning-distill-opus-4-7-max-sft

lordx64

Provides 7,823 single-turn reasoning conversations generated by Anthropic's Claude Opus 4.7 and reformatted into Qwen-style chat templates for supervised fine-tuning (SFT). Includes explicit <think> chain-of-thought blocks and many long reasoning chains (avg ~4k tokens).

#huggingface#anthropic#claude#llm#nlp+1
  • Previous
  • 1
  • More pages
  • 7
  • 8
  • 9
  • More pages
  • 28
  • Next