Team Ai
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01biglam /AmericanStories-parquet AmericanStories (parquet) A parquet-native reformat of dell-research-harvard/AmericanStories — article-level full text of ~20 million U.S. newspaper scans (1774–1963) from the Library of Congress's Chronicling America collection, originally extracted by Dell et al. (arXiv:2308.12477). This repo exists so the dataset loads in one line with the standard datasets / polars / pyarrow / dask stack, with no custom loading script and full Dataset Viewer support on the Hub.… See the full description on the dataset page: https://huggingface.co/datasets/biglam/AmericanStories-parquet.texttext-classification10M<n<100M3 likes2.1k downloads5mo agoHugging Face02Azzindani /ID_Supreme_Court_Parquet 💎 Indonesian Supreme Court Parquet Dataset (ID_Supreme_Court_Parquet) This repository provides a high-performance, compressed version of the Indonesian Supreme Court (Mahkamah Agung RI) court decisions. By converting raw legal data into the Apache Parquet format, this dataset is optimized for large-scale data engineering, fast I/O, and seamless integration with modern AI training pipelines. 🚀 💡 The Concept: Performance-First Legal Data While HTML and JSON are great… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Supreme_Court_Parquet.texttext-generation10K<n<100K0 likes37 downloads7mo agoHugging Face03Birchlabs /test-parquettextquestion-answeringn<1K0 likes33 downloads3y agoHugging Face04nicholasKluge /faquad-nli-parquet Dataset Card for FaQuAD-NLI THIS IS A TEMPORARY COPY OF THE ORIGINAL ruanchaves/faquad-nli. WHY? As of datasets==4.0, loading scripts and trust_remote_code are no longer supported. This breaks things, like the lm-evaluation-harness-pt, which people who work with Portuguese LLMs need for running evals. As soon as ruanchaves updates his version, I'll delete this copy. Dataset Summary FaQuAD is a Portuguese reading comprehension dataset that follows the format of the… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/faquad-nli-parquet.tabularquestion-answering1K<n<10K1 likes31 downloads1y agoHugging Face05SiwaSathya /Culture-Parquettextquestion-answeringn<1K0 likes6 downloads1y agoHugging Face06SiwaSathya /new-dataset-parquettextquestion-answering10K<n<100K0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.