AIAny
Icon for item

Alania Turkish Domain Text

Collection of 1.44M unique Turkish voice‑assistant sentences (≈1,819 hours estimated), normalized for TTS and organized by service domains (appointments, banking, e‑commerce). Designed for training and evaluating TTS and text-generation models; licensed CC BY 4.0 with required attribution.

Introduction

Real-world TTS failures concentrate on short, service-oriented utterances—dates, amounts, phone numbers, addresses and codes. This corpus addresses that gap by providing a large, curated set of machine-generated but normalized lines that mirror what voice agents actually say, so teams can train or benchmark TTS and text-generation systems on the tricky edge cases of conversational service language.

What Sets It Apart
  • Scale and focus: 1,439,639 unique Turkish sentences (estimated 1,819 hours read aloud) specifically targeted at voice‑agent domains such as appointments, banking, e‑commerce, customer support and empathetic responses. This is not a generic web crawl: content is organized by domain and subtopic.
  • TTS-ready normalization: each line includes both the raw generated text and a text_model form where numbers, dates, currencies, abbreviations and other spoken forms are expanded according to Turkish conventions, reducing the normaliser burden for TTS pipelines.
  • Reproducible generation metadata: lines were produced with Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 on vLLM; the dataset keeps provenance (generation batch, model, and normaliser changes) so researchers can audit or reproduce parts of the corpus.
  • Practical licensing: released under CC BY 4.0 (requires attribution to PatientDesk AI), making it suitable for commercial training and research while mandating credit.
Who It's For and Tradeoffs

Great fit if you need domain-focused Turkish training text for TTS, controllable synthetic speech pipelines, or to augment scarce real-world service utterances in model training and evaluation. The dataset speeds up iteration on normalisation rules and edge cases (phone numbers, codes, monetary amounts) and can be paired with real speech for ASR/TTS training. Look elsewhere if you need human-recorded audio, verbatim real-world transcripts, or fully human-verified factual content: the lines are machine-generated, invented names/phones/codes are synthetic, and not every line was human-validated. Also, do not use it to impersonate real people; follow legal requirements to label synthetic audio when required.

Notes on provenance and usage
  • Generator: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 on vLLM (September 2026); dataset includes extra metadata per line.
  • Versioning: changelog entries (v1.1, v1.2) document normaliser updates for Turkish initialisms, decimals and other reading rules.
  • Practical tip: combine these normalized texts with real recorded speech when training ASR or TTS to avoid overfitting to synthetic prosody and artifacts.

Information

  • Websitehuggingface.co
  • OrganizationsPatientDesk AI, cloud0day3
  • Published date2026/09/28

Categories

More Items

Hugging Face

Generates humanoid robot motion references that preserve object contact locations/timing by solving windowed trajectory optimizations against contact targets in the object frame. Releases retargeted trajectories for two Unitree robots across 75 objects (≈13.9k robot–motion pairs); CC BY‑NC‑SA 4.0.

Hugging Face

Provides a monthly Parquet snapshot of ~5.6 billion public TikTok videos (2014–Oct 2026), including captions, hashtags, sounds, engagement metrics and TikTok Shop links. Designed for large-scale querying (DuckDB/Pandas/Polars); licensed CC BY-NC 4.0 for research and personal use.

Turns solved protein structures into FoldingCorpus and Fold2Reason — a post-training recipe that supervises an LLM with discrete structural Q&A plus continuous 3D geometry to improve spatial, graph and scientific reasoning; reports consistent gains across 10 benchmarks.