AIAny
Icon for item

jabarti-llm-dataset

Provides a cleaned, section-chunked bilingual (Arabic + English) Wikipedia-derived corpus plus a curated Egyptian-history subset for LLM pretraining and SFT. Includes large pretrain and finetune splits, article-level eval holdouts, Parquet format, CC-BY-SA-4.0, and Arabic orthography caveats.

Introduction

The dataset matters because training-quality bilingual corpora that combine broad Wikipedia coverage with a curated, domain-focused subset are rare for Arabic–English LLM work; this one blends a multi-million-record pretrain corpus with a smaller, citation-grounded Egyptian-history collection and SFT QA pairs. That mix makes it useful both for general-language pretraining and for domain upweighting when you want historical Egyptian content represented.

What Sets It Apart
  • Two-source composition and scale: a large Wikipedia-derived bilingual pretrain split (millions of records) plus a curated “jabarti” source holding Egyptian-history articles (~4% of characters), enabling domain upweighting without rebuilding corpora. This lets you emphasize a niche domain while retaining broad coverage.
  • Curriculum-style splits for staged training: explicit phase and legacy splits (phase1/phase2) and separate pretrain/finetune configs make it straightforward to run multi-stage pipelines (pretrain → domain upweight → SFT).
  • Practical training metadata: chunk-level sample_weight (1/sqrt(chunk_count)), article-held-out evals (no article straddles train/eval), chunk_index/total, content_quality and egypt_relevance flags — helpful for balanced sampling and leakage-safe evaluation.
  • Engineering-friendly format and license: Parquet export, dataset ready for datasets.load_dataset, and CC-BY-SA-4.0 licensing for reuse in research and many downstream experiments.
Who It's For — and Tradeoffs

Great fit if you are training or fine-tuning LLMs that must handle Arabic and English and you need controlled domain emphasis (historical Egypt) or a bilingual SFT set for QA. It is also convenient for research pipelines that rely on article-level eval safety and chunk-weighting. Look elsewhere if you need fully native-standard Arabic orthography across all records (the llm_generated portion was written with stripped hamza/diacritics and is flagged with ortho_stripped), or if you require non-Wikipedia provenance or larger non-Wikipedia corpora.

Practical numbers and caveats
  • Key splits: pretrain (train ~7,053,893 records; eval ~114,259), finetune SFT (train ~57,488; eval ~6,379) as provided on the dataset page.
  • Sources: cohere-wiki (majority) + curated jabarti (Egyptian-history).
  • Caveats: Arabic orthography variance (ortho_stripped flag), overlapping coverage between sources (intentional re-chunking retained), and legacy phase splits that overlap with the new train/eval — follow the card guidance when composing pipelines.

Information

Categories

More Items

Turns solved protein structures into FoldingCorpus and Fold2Reason — a post-training recipe that supervises an LLM with discrete structural Q&A plus continuous 3D geometry to improve spatial, graph and scientific reasoning; reports consistent gains across 10 benchmarks.

Hugging Face

A 180,000-row labeled dataset of agent decision steps for training fast decision models that choose whether to call a tool, which tool to pick, and whether arguments are complete. Provides leakage-safe splits, typed questions, canonical state/request fields, and ready-to-use Parquet format for low-latency routers.

Hugging Face

Provides a sanitized, labeled SOC capture and merged provenance graph for intrusion-detection research, including 2,011,674 live signals, 51,371 incident graphs, MITRE ATT&CK mappings, and deterministic attack reports.