AIAny
Icon for item

BhasaFlow Khasi-English Parallel Sample v1

Provides 100 English–Khasi parallel sentence pairs with aligned studio-quality WAV recordings for ASR, TTS and translation evaluation; curated by Medharvix as a restricted public sample—full corpus available by request.

Introduction

Why this matters

Khasi is a low-resource language with very few publicly available aligned speech–text corpora. This sample release gives researchers a vetted, studio-quality preview (100 sentence pairs with aligned WAV audio and metadata) so teams can evaluate dataset format and baseline performance before requesting broader access.

What Sets It Apart
  • Gold-standard, human-created alignments: each English sentence was translated by native Khasi speakers and linked to a validated studio-quality WAV recording (16-bit PCM). This reduces noise commonly found in web-scraped speech datasets, so small-scale experiments reflect cleaner upper-bound performance.
  • Multi-modality and narrow scope: the sample pairs parallel text with audio and includes speaker gender metadata, making it immediately usable for ASR, TTS prototyping, and machine translation evaluation without heavy preprocessing.
  • Access-controlled full corpus: the public sample is intentionally small (100 examples) to enable evaluation and collaboration discovery, while the extended production-grade corpus is available through a permissioned request process, preserving contributor consent and licensing constraints.
Who It's For (and trade-offs)

Great fit if you need a reliable small testbed to evaluate ASR/TTS pipelines or translation models for Khasi, verify ingestion/metadata schemas, or demonstrate feasibility to stakeholders. The dataset’s curated quality makes it useful for benchmarking and linguistic analysis.

Look elsewhere if you need large-scale training data out of the box: the sample size (100 pairs) is insufficient for training robust production ASR/TTS models without augmentation or external data. Also note the release is under a restricted license; commercial or redistribution uses require contact and approval from Medharvix.

Where It Fits

Use this sample as a quality-controlled validation set or for early-stage experiments that measure model behavior on clean, native-speaker data. For production training, plan to combine this resource with additional in-domain or synthetic data and follow the dataset’s access/licensing workflow.

Notes on access and provenance

  • Publisher: Medharvix Systems (BhasaFlow initiative).
  • Public sample size: 100 sentence pairs.
  • Audio: WAV, 16-bit PCM.
  • Created: 2026-04-28.
  • Contact for full-corpus access and licensing: [email protected].

This introduction aims to clarify where the sample is most useful and the concrete limitations you’ll face when moving from evaluation to production.

Information

  • Websitehuggingface.co
  • AuthorsMedharvix Systems Private Limited
  • Published date2026/04/28

Categories

More Items

Provides a curriculum-aligned knowledge graph extracted from Chinese K–12 textbooks and accompanying benchmarks and training data to evaluate and train educational LLMs. Releases a 23,640-question multi-select benchmark and a 7,335-sample graph-guided training corpus with multimodal VQA pairs and the full construction pipeline.

Hugging Face

Pan-cancer CT segmentation dataset for training and benchmarking medical-image segmentation models — packaged as a Hugging Face dataset with an estimated 10k–100k samples and linked arXiv references. Designed for model development and reproducible benchmarking; non-commercial license applies.

Hugging Face

Provides a near-deduplicated, quality-filtered 15.9 TB training subset of GitHub source code grouped by repository, with inline UTF‑8 file contents and repo metadata for pre-training and analysis of code LLMs; cutoff Aug 7, 2025, ODC-By license.