AIAny
Icon for item

Joyo Kanji Yomi Benchmark

Provides kanji-level evaluation data for Japanese TTS: disambiguated sentence contexts targeting 4,378 kanji-reading pairs (2,136 Jōyō kanji) with 13,095 native-speaker–verified sentences and katakana-marked ground-truth readings for kanji-level error metrics.

Introduction

Oops! Something went wrong

[next-mdx-remote-client] error compiling MDX: Expected a closing tag for `<>` (6:125-6:127) before the end of `paragraph` 4 | - Coverage and granularity: covers all 2,136 Jōyō kanji and 4,378 kanji-reading pairs with three sentence contexts per reading (13,095 sentences), so you can evaluate per-reading behaviour rather than only word- or sentence-level quality — useful for pinpointing specific polyphony errors. 5 | - Native verification and disambiguation: sentences and annotations were reviewed by 35 native Japanese speakers through a multi-stage process, and ambiguous kanji-reading pairs that cannot be uniquely disambiguated by context were excluded — improving label reliability for evaluation. > 6 | - Evaluation-ready format: each sample includes a full-sentence katakana transcription with the target reading delimited by <> for automatic extraction; an accompanying evaluation toolkit handles TTS synthesis, ASR transcription, alignment, and metric computation, streamlining kanji-level experiments. | ^ 7 | - Focused on TTS/ASR pipelines: designed to measure pronunciation selection in synthesized speech (kanji→phoneme mapping under sentential context), not as a general-purpose language modeling corpus. 8 | More information: https://mdxjs.com/docs/troubleshooting-mdx

Information

  • Websitehuggingface.co
  • Organizationssbintuitions
  • Published date2026/06/11

Categories

More Items

Provides a year-scale multimodal benchmark and evaluation framework for on-device long-term memory in personal assistants, built from real mobile user trajectories. Tests memory construction, retrieval, updating, temporal reasoning, and implicit preference inference, and includes a knowledge-grounded synthesis pipeline to form coherent long-horizon trajectories.

Hugging Face

Provides MS MARCO queries, passages and answers translated into 14 Indic languages while keeping the original English content and per-example translation metadata. Includes train/validation splits, passage selection flags, and translation model parameters for multilingual IR, QA and RAG research.

Hugging Face

An index of Cara App content: metadata and CDN URLs for ~3.43M posts, 8.52M master artworks (~12M image links). Includes an SQLite catalog and Parquet exports but does not include image bytes — only links and metadata for analysis and search.