datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AmericanStories-parquet
AmericanStories (parquet)
A parquet-native reformat of dell-research-harvard/AmericanStories — article-level full text of ~20 million U.S. newspaper scans (1774–1963) from the Library of Congress's Chronicling America collection, originally extracted by Dell et al. (arXiv:2308.12477).
This repo exists so the dataset loads in one line with the standard datasets / polars / pyarrow / dask stack, with no custom loading script and full Dataset Viewer support on the Hub.… See the full description on the dataset page: https://huggingface.co/datasets/biglam/AmericanStories-parquet.ID_Supreme_Court_Parquet
💎 Indonesian Supreme Court Parquet Dataset (ID_Supreme_Court_Parquet)
This repository provides a high-performance, compressed version of the Indonesian Supreme Court (Mahkamah Agung RI) court decisions. By converting raw legal data into the Apache Parquet format, this dataset is optimized for large-scale data engineering, fast I/O, and seamless integration with modern AI training pipelines. 🚀
💡 The Concept: Performance-First Legal Data
While HTML and JSON are great… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Supreme_Court_Parquet.test-parquetfaquad-nli-parquet
Dataset Card for FaQuAD-NLI
THIS IS A TEMPORARY COPY OF THE ORIGINAL ruanchaves/faquad-nli.
WHY? As of datasets==4.0, loading scripts and trust_remote_code are no longer supported.
This breaks things, like the lm-evaluation-harness-pt, which people who work with Portuguese LLMs need for running evals.
As soon as ruanchaves updates his version, I'll delete this copy.
Dataset Summary
FaQuAD is a Portuguese reading comprehension dataset that follows the format of the… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/faquad-nli-parquet.Culture-Parquetnew-dataset-parquet
