emre570/us-legal-code
Dataset Card for United States Code (Cornell LII) — Hierarchical Sections Dataset Summary This dataset is purpose-built for the Prime Intellect U.S. legal evaluation environment. This dataset contains the text of the United States Code scraped from the Legal Information Institute at Cornell Law School. Each record corresponds to a navigable section (“U.S. Code” tab only) together with its hierarchy path—title, subtitle, division, part, subpart, chapter, subchapter… See the full description on the dataset page: https://huggingface.co/datasets/emre570/us-legal-code.
Dataset Card for United States Code (Cornell LII) — Hierarchical Sections
Dataset Summary
This dataset is purpose-built for the Prime Intellect U.S. legal evaluation environment.
This dataset contains the text of the United States Code scraped from the Legal Information Institute at Cornell Law School. Each record corresponds to a navigable section (“U.S. Code” tab only) together with its hierarchy path—title, subtitle, division, part, subpart, chapter, subchapter, and so on. Hierarchy levels that do not exist for a given section are stored as the literal string "null" so the schema remains uniform across the corpus.
The crawler normalises section URLs to include #tab_default_1, ensuring we always capture the primary statutory text and exclude notes or editorial commentary. Headings preserve the source wording (e.g. 15 U.S. Code § 9901 - ...) and the text field is whitespace-normalised plain text with in-page navigation removed.
Languages
- English (
en)
Dataset Structure
Data Fields
All fields are strings. "null" marks missing hierarchy levels or absent text.
Data Instances
{
"title_id": "15",
"title_url": "https://www.law.cornell.edu/uscode/text/15",
"subtitle_id": "subtitle-I",
"subtitle_url": "https://www.law.cornell.edu/uscode/text/15/subtitle-I",
"division_id": "division-C",
"division_url": "https://www.law.cornell.edu/uscode/text/15/subtitle-I/division-C",
"chapter_id": "chapter-123",
"chapter_url": "https://www.law.cornell.edu/uscode/text/15/subtitle-I/division-C/chapter-123",
"subchapter_id": "null",
"subchapter_url": "null",
"section_id": "9901",
"section_url": "https://www.law.cornell.edu/uscode/text/15/9901#tab_default_1",
"heading": "15 U.S. Code § 9901 - Prohibition on transfer of personally identifiable sensitive data of United States individuals to foreign adversaries",
"text": "(a) Prohibition It shall be unlawful ..."
}Data Splits
No official train/validation/test split is supplied. You can:
- deterministically slice with
split="train[:5000]", or - stream via
load_dataset(..., streaming=True)and build custom splits.
Dataset Creation
Source Data
- Collection: Scraped via Scrapy from
https://www.law.cornell.edu/uscode/text. - Processing: The spider walks hierarchy containers, extracts the
#tab_default_1tab, and stores each level’s ID/URL. Normalisation rewrites the JSONL so every record shares the same schema and missing fields become"null".
Author Statements
- Primary source: Legal Information Institute (LII) at Cornell Law School.
- Dataset maintainers: emre570 and Codex (GPT-5)
Limitations
- Reflects Cornell LII’s snapshot at crawl time; statutes change periodically.
"null"strings indicate missing levels—downstream consumers should treat them as absent data.- Editorial notes and historical references are not included.
Ethical Considerations
Content is public law. There is no personal data beyond statutory text. Users should verify statutory currency through official sources before relying on the dataset in production or legal contexts.
Licensing
Cornell LII distributes its value-added materials under the Creative Commons Attribution-NonCommercial-ShareAlike 2.5 License. This dataset follows the same terms: give credit to Cornell LII, do not use the content for commercial purposes, and release derivatives under an identical license. For commercial arrangements, contact permissions@liicornell.org.
Contributions
- [emre570](https://linktr.ee/emre570) — project owner, crawl execution, dataset curation.
- Codex (GPT-5) — crawler automation and documentation support :)
Found an issue or want to contribute improvements? Please open an issue or pull request in the repository.
