AIAny
Icon for item

Nemotron-SFT-SWE-v3.5

Provides agentic instruction‑tuning trajectories for software‑engineering tasks, formatted for supervised fine‑tuning and agent training. Contains multi‑file edits, tests, docs and structured agent traces (≈5,115 records, 1.9 GiB). Intended for commercial use; licensed CC‑BY 4.0 with additional permissive licenses.

Introduction

Most code-focused SFT blends target single-shot repairs or test generation; this dataset emphasizes agentic, multi-step workflows and repository-aware edits so models learn planning, tool use, and cross-file reasoning. It packages structured agent traces collected with the OpenCode harness and curated for supervised fine‑tuning of software-engineering agents.

What Sets It Apart
  • Agentic trajectories rather than isolated Q&A: includes stepwise agent actions, tool calls and patch-style edits so fine‑tuned models can learn multi-step workflows and intermediate state management (so what: better support for agents that must plan and execute sequences across files).
  • Multi-artifact, multi-file focus: examples include source, tests, docs and config changes together, not just single-file patches (so what: improves models’ repository-aware reasoning and regression‑free edits).
  • Compact, curated SFT target: ~5,115 records (1.9 GiB) designed for supervised fine‑tuning rather than large‑scale pretraining (so what: faster iteration for teams training moderate-size models or distillation pipelines).
  • Clear commercial-use signal and mixed permissive licensing: distributed under CC‑BY 4.0 with additional permissive licenses noted (so what: suitable for product integration but check organiational legal requirements).
Who It's For and Tradeoffs

Great fit if you are training or distilling LLMs to act as autonomous coding agents, need cross-file repair/test-generation examples, or want agent traces that reflect tool use and stepwise reasoning. Look elsewhere if you need very large-scale SWE corpora (tens or hundreds of thousands of examples) for pretraining, raw repository snapshots, or if your project policies prohibit the dataset's commercial-use terms. In short: efficient SFT material for repository-aware agent behavior, with explicit tradeoffs around scale and licensing.

More Items

Hugging Face

Aggregated, screened corpus of 55,050 normalized Indian public-information text bodies and 65,209 source records for retrieval and question-answering. Exports include deduplicated CSV/Parquet with provenance, topic labels, extraction quality flags and a private SQLite backup.

Hugging Face

Provides 3,451 hours (2,051,810 clips) of AI‑generated 48 kHz Turkish speech with transcripts, spoken forms and per‑clip voice descriptions for TTS and ASR development. Includes 2,752 designed voices and is licensed CC BY 4.0 / CC BY‑SA 4.0 (attribution to PatientDesk AI required).

Hugging Face

Provides imagined interaction segments generated by world models for RoboTwin2.0 tasks, stored as fixed-length HDF5 chunks (21 observation frames, 20 actions, rewards and episode flags). Useful for training and evaluating world-model-based policies; currently limited to the RoboTwin2.0 subdataset.