AIAny
Icon for item

Vāgdhenu — Sanskrit Chant Corpus

Provides ~1,467 single-speaker Sanskrit chant audio clips (≈5.3 hours) with aligned transcripts and prosodic metadata for meter-aware TTS training. Two recording/config styles (style_a/style_b), 24 kHz mono WAVs, metadata includes Devanagari, SLP1, Kannada text, meter, duration, session/take. CC-BY-4.0.

Introduction

Sanskrit chant requires precise, metrically consistent prosody that ordinary speech datasets do not capture. This corpus delivers a single-speaker, tradition-faithful recording set with per-clip prosodic annotations so TTS systems can learn chant-specific timing, pausing, and vowel/ consonant shaping rather than plain speech patterns.

What Sets It Apart
  • Metrically-aware recordings: each clip preserves pāda-level breath groups, daṇḍa pauses, and yati caesura rules—so a model can learn chant-accurate pause placement and phrasing rather than generic sentence boundaries.
  • Two complementary cuts: style_a (764 clips, ~2.70 h) and style_b (703 clips, ~2.64 h) contain largely different verses and metadata variants (style_b includes explicit meter and syllable counts), enabling experiments on data-splitting and prosody conditioning.
  • High-quality, consistent capture: single reciter, fixed microphone setup, lossless WAV derived to 24 kHz, low noise floor and controlled peaks—minimizes speaker/recording variability for cleaner model training.
  • Rich metadata per clip: Devanagari text, SLP1 transliteration, Kannada-routed text, duration, session/take, and (style_b) meter and n_syll, useful for meter-conditioned synthesis and evaluation.
Who It's For and Trade-offs

Great fit if you want to train or fine-tune chant- or meter-aware TTS, study prosody/meter in Sanskrit verse, or create accessible renditions of classical ślokas. The single-speaker, tradition-faithful design reduces inter-speaker variance and highlights prosody learning.

Look elsewhere if you need large-scale multi-speaker conversational speech, Vedic svara recordings (this corpus excludes Vedic svaras), or massively diverse acoustic conditions—this dataset prioritizes ritual/traditional chant fidelity over breadth. Also, total duration (~5.3 h) is modest for large neural TTS pretraining but appropriate for fine-tuning or focused prosody research.

Where It Fits

Use this as fine-tuning data for meter-conditioned TTS models, a benchmark for chant prosody research, or as high-quality training examples when building accessible audio renditions of Sanskrit ślokas. License: CC-BY-4.0; respect the author's voice and attribution guidance when releasing synthesized outputs.

Information

Categories

More Items

Hugging Face

A cleaned supervised fine-tuning dataset of 6,365 Claude Fable-5 agent traces in OpenAI Chat and Hugging Face agent-traces formats, prepared for SFT, tool-use training, and distillation workflows; MIT-licensed and distributed as Parquet.

Hugging Face

Provides a machine-readable catalog of 117 AI/AX safety and deployment-readiness diagnostic criteria for assessing model intrinsic and serving/infrastructure risks. Includes MODEL-SCAN and AX-SCAN axes, bilingual source fields, per-item evidence guidance, severity/assurance metadata, and a CC BY-NC 4.0 release-candidate.

Hugging Face

Evaluates schema-guided structured extraction from documents: given a document and a JSON schema, systems must return a schema-valid JSON with page-and-box grounding. Covers 370 documents (4,869 pages) across 8 business domains and 67 document types; scores value accuracy, word/page grounding, and long-list completeness.