Large open synthetic speech corpora in Turkish are rare; this release supplies a multi‑thousand‑hour, carefully filtered synthetic spoken dataset intended to fill that gap for TTS and ASR work. The dataset emphasizes reproducibility (generation seeds and settings per row), voice design diversity, and licensing that permits commercial use with attribution.
What Sets It Apart
- Scale and diversity: 3,451 hours across 2,051,810 clips, rendered at 48 kHz and covering 2,752 designed voice identities plus one‑off voices — useful when target language real recordings are scarce.
- Per‑clip metadata: each clip pairs written text, a normalized spoken form, a plain‑English voice description, generation seeds/settings, and automated QC metrics (CER, UTMOS, speaker similarity, f0 checks, speaking rate).
- Controlled variations: three speaking modes (reference, reference+instruction, voice‑design) and ~45% of clips include a simulated noisy channel variant (room, mic, background) to help robustness testing.
- Open licensing for reuse: CC BY 4.0 for most configs and CC BY‑SA 4.0 for the Wikipedia‑derived subset; explicit attribution requirement to PatientDesk AI enables commercial and research usage while preserving provenance.
Who It's For & Trade‑offs
Great fit if you need large Turkish TTS training data, voice‑design examples, or synthetic corpora for ASR augmentation. Researchers building controllable TTS, studying prosody, or evaluating robustness to channel/noise will benefit from the metadata and simulated channels. Look elsewhere if you require natural human variability for final production ASR/TTS models: the corpus is synthetic (inherits generator artefacts and uniformity) and authors recommend mixing with real speech. Also avoid using it to impersonate real people or to pass generated audio as human; label downstream outputs as AI‑generated where law requires. Known issues (v1.0) include some initialism/pronunciation mismatches for a small fraction of clips; re‑renders are planned in subsequent versions.