The dataset matters because training-quality bilingual corpora that combine broad Wikipedia coverage with a curated, domain-focused subset are rare for Arabic–English LLM work; this one blends a multi-million-record pretrain corpus with a smaller, citation-grounded Egyptian-history collection and SFT QA pairs. That mix makes it useful both for general-language pretraining and for domain upweighting when you want historical Egyptian content represented.
What Sets It Apart
- Two-source composition and scale: a large Wikipedia-derived bilingual pretrain split (millions of records) plus a curated “jabarti” source holding Egyptian-history articles (~4% of characters), enabling domain upweighting without rebuilding corpora. This lets you emphasize a niche domain while retaining broad coverage.
- Curriculum-style splits for staged training: explicit phase and legacy splits (phase1/phase2) and separate pretrain/finetune configs make it straightforward to run multi-stage pipelines (pretrain → domain upweight → SFT).
- Practical training metadata: chunk-level sample_weight (1/sqrt(chunk_count)), article-held-out evals (no article straddles train/eval), chunk_index/total, content_quality and egypt_relevance flags — helpful for balanced sampling and leakage-safe evaluation.
- Engineering-friendly format and license: Parquet export, dataset ready for datasets.load_dataset, and CC-BY-SA-4.0 licensing for reuse in research and many downstream experiments.
Who It's For — and Tradeoffs
Great fit if you are training or fine-tuning LLMs that must handle Arabic and English and you need controlled domain emphasis (historical Egypt) or a bilingual SFT set for QA. It is also convenient for research pipelines that rely on article-level eval safety and chunk-weighting. Look elsewhere if you need fully native-standard Arabic orthography across all records (the llm_generated portion was written with stripped hamza/diacritics and is flagged with ortho_stripped), or if you require non-Wikipedia provenance or larger non-Wikipedia corpora.
Practical numbers and caveats
- Key splits: pretrain (train ~7,053,893 records; eval ~114,259), finetune SFT (train ~57,488; eval ~6,379) as provided on the dataset page.
- Sources:
cohere-wiki(majority) + curatedjabarti(Egyptian-history). - Caveats: Arabic orthography variance (
ortho_strippedflag), overlapping coverage between sources (intentional re-chunking retained), and legacy phase splits that overlap with the new train/eval — follow the card guidance when composing pipelines.