AIAny
Icon for item

Bharat Guide — Indian public information documents

Aggregated, screened corpus of 55,050 normalized Indian public-information text bodies and 65,209 source records for retrieval and question-answering. Exports include deduplicated CSV/Parquet with provenance, topic labels, extraction quality flags and a private SQLite backup.

Introduction

Government web content is abundant but inconsistent; this snapshot normalizes and screens tens of thousands of Indian public-information pages so they can be used directly in retrieval-augmented pipelines and NLP workflows. It prioritizes provenance, extraction quality flags and topic labels to make noisy official sources more tractable for QA, indexing and analysis.

What Sets It Apart
  • High-volume, screened coverage: 55,050 distinct normalized text bodies and 65,209 source records with per-document status (good vs review) so consumers can filter by extraction quality.
  • Provenance-first exports: main CSV/Parquet rows contain full extracted text plus content_hash, document_id, title, url, host, quality, trust, status, topics, fetched_at and JSON metadata to preserve original source aliases and retrieval details.
  • Topic and host indexing: automated topic labels across 36 broad areas (law, schemes, policy, universities, etc.) and coverage from 15,705 distinct hostnames (11,187 .gov.in hosts), enabling focused subcorpora.
  • Practical formats and tooling: primary downloads are documents.csv.gz and Parquet batches for the Hugging Face viewer; designed to be read with pandas, polars or Dask and integrated into RAG/QA pipelines.
Who it's for + Tradeoffs

Great fit if you need a large, provenance-rich Indian public-text corpus for retrieval, question-answering, summarization, information extraction, or building domain-specific RAG indexes. Use the status and quality fields to exclude thin or OCR-pending records. Look elsewhere or take care if you require legally authoritative or fully validated texts: the release is a screened crawl, not a manual legal validation; rights remain with original publishers, fetch times are not publication dates, and numeric OCR accuracy or legal effect (amendments, supersessions) must be checked against authoritative sources.

Where it fits

Use this dataset as a source collection for building retrieval indexes, QA benchmarks, or NLP training/analysis focused on Indian government and public-sector content. Combine it with downstream verification or canonical legal sources when legal accuracy or the latest amendments are required.

Information

Categories

More Items

Hugging Face

Provides 3,451 hours (2,051,810 clips) of AI‑generated 48 kHz Turkish speech with transcripts, spoken forms and per‑clip voice descriptions for TTS and ASR development. Includes 2,752 designed voices and is licensed CC BY 4.0 / CC BY‑SA 4.0 (attribution to PatientDesk AI required).

Hugging Face

Provides imagined interaction segments generated by world models for RoboTwin2.0 tasks, stored as fixed-length HDF5 chunks (21 observation frames, 20 actions, rewards and episode flags). Useful for training and evaluating world-model-based policies; currently limited to the RoboTwin2.0 subdataset.

Hugging Face

Provides 7,366 recorded agent trajectories from H Company’s Holo4 benchmark runs, with step-level reasoning, actions, tool results, token usage and screenshots for replay and analysis. Bundled as JSON and image files for per-trajectory inspection and automated replay; released under Apache 2.0.