Government web content is abundant but inconsistent; this snapshot normalizes and screens tens of thousands of Indian public-information pages so they can be used directly in retrieval-augmented pipelines and NLP workflows. It prioritizes provenance, extraction quality flags and topic labels to make noisy official sources more tractable for QA, indexing and analysis.
What Sets It Apart
- High-volume, screened coverage: 55,050 distinct normalized text bodies and 65,209 source records with per-document status (
goodvsreview) so consumers can filter by extraction quality. - Provenance-first exports: main CSV/Parquet rows contain full extracted text plus
content_hash,document_id,title,url,host,quality,trust,status,topics,fetched_atand JSONmetadatato preserve original source aliases and retrieval details. - Topic and host indexing: automated topic labels across 36 broad areas (law, schemes, policy, universities, etc.) and coverage from 15,705 distinct hostnames (11,187 .gov.in hosts), enabling focused subcorpora.
- Practical formats and tooling: primary downloads are
documents.csv.gzand Parquet batches for the Hugging Face viewer; designed to be read with pandas, polars or Dask and integrated into RAG/QA pipelines.
Who it's for + Tradeoffs
Great fit if you need a large, provenance-rich Indian public-text corpus for retrieval, question-answering, summarization, information extraction, or building domain-specific RAG indexes. Use the status and quality fields to exclude thin or OCR-pending records.
Look elsewhere or take care if you require legally authoritative or fully validated texts: the release is a screened crawl, not a manual legal validation; rights remain with original publishers, fetch times are not publication dates, and numeric OCR accuracy or legal effect (amendments, supersessions) must be checked against authoritative sources.
Where it fits
Use this dataset as a source collection for building retrieval indexes, QA benchmarks, or NLP training/analysis focused on Indian government and public-sector content. Combine it with downstream verification or canonical legal sources when legal accuracy or the latest amendments are required.