Real-world TTS failures concentrate on short, service-oriented utterances—dates, amounts, phone numbers, addresses and codes. This corpus addresses that gap by providing a large, curated set of machine-generated but normalized lines that mirror what voice agents actually say, so teams can train or benchmark TTS and text-generation systems on the tricky edge cases of conversational service language.
What Sets It Apart
- Scale and focus: 1,439,639 unique Turkish sentences (estimated 1,819 hours read aloud) specifically targeted at voice‑agent domains such as appointments, banking, e‑commerce, customer support and empathetic responses. This is not a generic web crawl: content is organized by domain and subtopic.
- TTS-ready normalization: each line includes both the raw generated text and a
text_modelform where numbers, dates, currencies, abbreviations and other spoken forms are expanded according to Turkish conventions, reducing the normaliser burden for TTS pipelines. - Reproducible generation metadata: lines were produced with Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 on vLLM; the dataset keeps provenance (generation batch, model, and normaliser changes) so researchers can audit or reproduce parts of the corpus.
- Practical licensing: released under CC BY 4.0 (requires attribution to PatientDesk AI), making it suitable for commercial training and research while mandating credit.
Who It's For and Tradeoffs
Great fit if you need domain-focused Turkish training text for TTS, controllable synthetic speech pipelines, or to augment scarce real-world service utterances in model training and evaluation. The dataset speeds up iteration on normalisation rules and edge cases (phone numbers, codes, monetary amounts) and can be paired with real speech for ASR/TTS training. Look elsewhere if you need human-recorded audio, verbatim real-world transcripts, or fully human-verified factual content: the lines are machine-generated, invented names/phones/codes are synthetic, and not every line was human-validated. Also, do not use it to impersonate real people; follow legal requirements to label synthetic audio when required.
Notes on provenance and usage
- Generator: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 on vLLM (September 2026); dataset includes
extrametadata per line. - Versioning: changelog entries (v1.1, v1.2) document normaliser updates for Turkish initialisms, decimals and other reading rules.
- Practical tip: combine these normalized texts with real recorded speech when training ASR or TTS to avoid overfitting to synthetic prosody and artifacts.