AIAny
Icon for item

MS MARCO Translations Dataset (ai4bharat/MSMARCO-XI)

Provides MS MARCO queries, passages and answers translated into 14 Indic languages while keeping the original English content and per-example translation metadata. Includes train/validation splits, passage selection flags, and translation model parameters for multilingual IR, QA and RAG research.

Introduction

MS MARCO is a cornerstone resource for retrieval and question answering; this dataset fills a practical gap by providing the same query–passage–answer triples translated into 14 Indic languages with preserved English originals and translation metadata. That makes it immediately usable for building and evaluating multilingual retrieval, cross-lingual QA, and RAG systems that target Indian languages rather than only English.

What Sets It Apart
  • Parallel retrieval-ready examples: each record contains the translated query/answer/passages alongside the original English query and answer — so you can evaluate monolingual, cross-lingual, and bilingual pipelines without re-annotation. This reduces dataset engineering effort when adapting English IR datasets to Indic languages.
  • Translation metadata per example: model name, temperature, max_tokens, top_p and penalty settings are included — so you can audit or filter by translation quality/setting and reproduce or compare translation conditions.
  • Language coverage and splits: train and validation files are provided per language for 14 Indic languages (e.g., hi, bn, ta, te, ml), plus passage selection flags to support passage-ranking tasks and RAG retrieval evaluation.
  • Practical formats and scale: distributed in parquet/jsonl-ready splits and sized for research experiments (tens of GB across languages) — suitable for embedding/indexing pipelines and large-scale IR benchmarks.
Who it's for and tradeoffs

Great fit if you are building or evaluating multilingual retrieval, QA, or RAG systems for Indian languages, need parallel English↔Indic testbeds, or want to study translation impact on downstream retrieval. Look elsewhere if you require human-verified annotations for every translation quality dimension (some portions use model translations and include metadata but not exhaustive human verification), or if you need a small toy dataset — this is designed for mid-to-large scale research and indexing workloads. Also note licensing follows the original MS MARCO terms, so check those before redistribution or commercial use.

Information

Categories

More Items

Provides a year-scale multimodal benchmark and evaluation framework for on-device long-term memory in personal assistants, built from real mobile user trajectories. Tests memory construction, retrieval, updating, temporal reasoning, and implicit preference inference, and includes a knowledge-grounded synthesis pipeline to form coherent long-horizon trajectories.

Hugging Face

An index of Cara App content: metadata and CDN URLs for ~3.43M posts, 8.52M master artworks (~12M image links). Includes an SQLite catalog and Parquet exports but does not include image bytes — only links and metadata for analysis and search.

Hugging Face

A cleaned supervised fine-tuning dataset of 6,365 Claude Fable-5 agent traces in OpenAI Chat and Hugging Face agent-traces formats, prepared for SFT, tool-use training, and distillation workflows; MIT-licensed and distributed as Parquet.