AIAny
Icon for item

LOCUS v1.0

Provides a county-harmonized corpus of U.S. municipal and county ordinance text (≈2.21M chunks) labeled for function, substantive indicator, and topic to support legal NLP, retrieval, and comparative local-law research. Includes model-assigned labels and continuous scorers (opacity, paternalism, enforcement discretion) plus coverage metadata; not exhaustive or a substitute for legal advice.

Introduction

Local ordinances shape everyday regulation (zoning, housing, licensing, nuisance) yet are fragmented and hard to analyze at scale; LOCUS v1.0 makes that layer of law machine-observable by assembling a county-harmonized, chunk-level corpus of U.S. municipal and county ordinance text annotated for downstream legal-NLP tasks.

Key Findings
  • Scale and scope: ~2,211,516 text chunks derived from municipal and county codes, with coverage metadata that links chunks to jurisdiction, state, city, and county.
  • Annotation schema: Each chunk is assigned a function label (Context, Rules, Process, Enforcement), a binary is_substantive flag, and for substantive chunks a coarse topic (Buildings, Business, Nuisance, Zoning, Other). The release also includes continuous scores for opacity, paternalism, enforcement discretion, and problem salience to support analytic tasks beyond simple classification.
  • Practical tradeoffs: The dataset uses OCR and automated classifiers to scale labeling across thousands of jurisdictions; this yields broad geographic reach but introduces label noise, taxonomy coarseness, and uneven jurisdictional digitization.
Who it's for and tradeoffs

Great fit if you need a large, jurisdiction-linked corpus to prototype legal-text classifiers, build substantive vs. non-substantive filters, or run comparative studies of municipal regulation. Look elsewhere if you require a fully audited legal source, exhaustive coverage of every U.S. locality, or fine-grained legal subject-matter labeling; LOCUS v1.0 is a snapshot and its function/topic labels are model-assigned and not fully human-validated.

Where it fits

LOCUS is best used as research infrastructure for legal NLP pipelines, weakly supervised workflows, and empirical policy analysis that benefit from reproducible coverage metadata and scalable annotations. For production legal advice or litigation-grade proof, human legal review and up-to-date jurisdictional checks remain necessary.

Information

  • Websitehuggingface.co
  • OrganizationsUC Berkeley, School of Information (UC Berkeley), Independent
  • AuthorsDenis Peskoff, Joe Barrow, Christopher Vu, Diag Davenport
  • Published date2026/05/06

Categories

More Items

Hugging Face

Provides over 1.1M hours of high-bandwidth, multichannel multilingual speech with segment- and word-level timestamps, English translations, and per-file metadata for ASR, TTS and audio-representation research. Preserves original 48kHz multichannel OPUS audio and is released under CC BY 3.0.

Hugging Face

Synthetic, clinician-verified ChatML dataset of 2,194 doctor–patient encounters covering 2,194 unique human diseases; each JSONL record includes 20 structured fields, verified PubMed references, realistic vitals/labs, and is intended for RAG and model fine-tuning (not medical advice).

Hugging Face

Provides 2,000 synthetic multiple-choice items designed for continuation log-likelihood scoring to evaluate small language models' Theory of Mind (social-cognitive) abilities; 40 constructs, balanced answer positions, and easy/medium difficulty.