AIAny
Icon for item

Joyo Kanji Yomi Benchmark

Provides kanji-level evaluation data for Japanese TTS: disambiguated sentence contexts targeting 4,378 kanji-reading pairs (2,136 Jōyō kanji) with 13,095 native-speaker–verified sentences and katakana-marked ground-truth readings for kanji-level error metrics.

Introduction

Oops! Something went wrong

[next-mdx-remote-client] error compiling MDX: Expected a closing tag for `<>` (6:125-6:127) before the end of `paragraph` 4 | - Coverage and granularity: covers all 2,136 Jōyō kanji and 4,378 kanji-reading pairs with three sentence contexts per reading (13,095 sentences), so you can evaluate per-reading behaviour rather than only word- or sentence-level quality — useful for pinpointing specific polyphony errors. 5 | - Native verification and disambiguation: sentences and annotations were reviewed by 35 native Japanese speakers through a multi-stage process, and ambiguous kanji-reading pairs that cannot be uniquely disambiguated by context were excluded — improving label reliability for evaluation. > 6 | - Evaluation-ready format: each sample includes a full-sentence katakana transcription with the target reading delimited by <> for automatic extraction; an accompanying evaluation toolkit handles TTS synthesis, ASR transcription, alignment, and metric computation, streamlining kanji-level experiments. | ^ 7 | - Focused on TTS/ASR pipelines: designed to measure pronunciation selection in synthesized speech (kanji→phoneme mapping under sentential context), not as a general-purpose language modeling corpus. 8 | More information: https://mdxjs.com/docs/troubleshooting-mdx

Information

  • Websitehuggingface.co
  • Organizationssbintuitions
  • Published date2026/06/11

Categories

More Items

Hugging Face

Aggregated, screened corpus of 55,050 normalized Indian public-information text bodies and 65,209 source records for retrieval and question-answering. Exports include deduplicated CSV/Parquet with provenance, topic labels, extraction quality flags and a private SQLite backup.

Hugging Face

Provides 3,451 hours (2,051,810 clips) of AI‑generated 48 kHz Turkish speech with transcripts, spoken forms and per‑clip voice descriptions for TTS and ASR development. Includes 2,752 designed voices and is licensed CC BY 4.0 / CC BY‑SA 4.0 (attribution to PatientDesk AI required).

Hugging Face

Provides imagined interaction segments generated by world models for RoboTwin2.0 tasks, stored as fixed-length HDF5 chunks (21 observation frames, 20 actions, rewards and episode flags). Useful for training and evaluating world-model-based policies; currently limited to the RoboTwin2.0 subdataset.