datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
riskroll-sec-10k-10q-sections
Riskroll: SEC 10-K and 10-Q sections as clean text
Need it fresh, filtered or via API? This free file is a snapshot (10-K/10-Q sections up to the last refresh), last updated 2026-09-24.
Insidewell on Apify ($0.004 per insider transaction): pulls today's SEC Form 4 trades for your own watchlist, filtered by buy/sell and size, with cluster-buy alerts on a schedule.
Using it at work? Commercial license + support (from $49/year): invoice, PDF licence certificate, named… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/riskroll-sec-10k-10q-sections.indian-legal-sections-bns-bnss-bsa-2023
🏛️ Indian Legal Sections — BNS · BNSS · BSA 2023
The First Structured, Unified JSON Dataset of Modern Indian Criminal Law
📖 Dataset Summary
This dataset contains 1,059 fully structured and verified sections extracted, parsed, and unified from India's three landmark criminal justice reform acts passed in December 2023. These three acts together replaced the colonial-era Indian Penal Code (IPC, 1860), the Code of Criminal Procedure… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/indian-legal-sections-bns-bnss-bsa-2023.sections
Dataset Card for "sections"
More Information needed
Coffee_leaves_sections_FO
Dataset Card for coffee_leaves_anomalib_2
This is a FiftyOne dataset with 35962 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("pjramg/Coffee_leaves_sections_FO")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/pjramg/Coffee_leaves_sections_FO.stranger_sections_2wikipedia-sections
Dataset Card for Wikipedia Sections
This dataset contains pairs and triplets that can be used to train and finetune Sentence Transformer embedding models. The dataset originates from Dor et al., and was downloaded from this download link.
Notably, the "anchor" column contains sentences from Wikipedia, wheras the "positive" column contains other sentences from the same section. The "negative" column contains sentences from other sections.
Dataset Subsets… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/wikipedia-sections.finewiki-en-sections-propella
Propella annotations for level-2 English Wikipedia sections
English Wikipedia pages from the en config of
HuggingFaceFW/finewiki,
split into one row per level-2 section and annotated with
ellamind/propella-1-4b. It is
the English counterpart of
BramVanroy/finewiki-nl-sections-propella,
and it was built for the training data of a Dutch and English embedding model.
How it was built
Every page is split at its ## headings. A section is the text up to the
next… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/finewiki-en-sections-propella.flock-demo-critical-infra-sectionsarxiv-sectionshindawi-arabic-sections
Hindawi Arabic Books — Sections Dataset
A cleaned, section-level dataset of Arabic books from Hindawi.org, prepared for NLP training and research.
Source
Original books scraped from Hindawi.org, a non-profit foundation providing free Arabic books. The dataset covers categories including literature, philosophy, history, science, psychology, and more.
Cleaning Pipeline
Scraped book content section-by-section from Hindawi.org
Removed English / Latin text and… See the full description on the dataset page: https://huggingface.co/datasets/nomeda-lab/hindawi-arabic-sections.finewiki-nl-sections-propella
Propella annotations for second-level, Dutch Wikipedia sections
This dataset is an exploded and annotated version of the Dutch portion of FineWiki.
Articles were split so that each second-level section (##) is now its own sample. Any introductory sections (between the top heading and the first sub-section) is not included.
Before annotation, the text was truncated to the first 50_000 characters, as recommended by the Propella README.
Intended use
This dataset may… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/finewiki-nl-sections-propella.flock-demo-object-detection-sectionsthin-sections-wikimediaflock-demo-defense-graph-sectionsflock-demo-slm-qwen3-0-6b-sectionsflock-demo-llm-finetuning-sectionshindawi-arabic-sections
Hindawi Arabic Books — Sections Dataset
A cleaned, section-level dataset of Arabic books from Hindawi.org, prepared for NLP training and research.
Source
Original books scraped from Hindawi.org, a non-profit foundation providing free Arabic books. The dataset covers categories including literature, philosophy, history, science, psychology, and more.
Cleaning Pipeline
Scraped book content section-by-section from Hindawi.org
Removed English / Latin text and… See the full description on the dataset page: https://huggingface.co/datasets/Abdallah4Zain/hindawi-arabic-sections.IMRAD-sections-clf-gemini-augmented
Dataset Card for IMRAD Classification Dataset (100k Rows)
Dataset Name: IMRAD Classification Dataset (100k Rows)
Dataset Description:
This dataset contains approximately 100,000 sentences extracted from scientific research papers and labeled according to their corresponding IMRAD (Introduction, Methods, Results, and Discussion) sections. The data was initially sourced from the unarXive_imrad_clf dataset on Hugging Face and expanded using data augmentation techniques. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/stormsidali2001/IMRAD-sections-clf-gemini-augmented.model_cards_with_readmes_sections
Dataset Card for "model_cards_with_readmes_sections"
More Information needed
ipc-sectionsflock-demo-automatic-speech-recognition-sectionsthin-sections-wikimediaflock-demo-finance-sentiment-sectionsflock-demo-graph-neural-network-sectionstrafic-routier-sections-de-comptage-departement-du-loiret-2024
Trafic routier - Sections de comptage - Département du Loiret - 2024
Source
Source officielle : https://www.data.gouv.fr/datasets/trafic-routier-sections-de-comptage-departement-du-loiret-2024
Identifiant du jeu de données data.gouv.fr : 68f9746546f42706c79cb4c8
Slug data.gouv.fr : trafic-routier-sections-de-comptage-departement-du-loiret-2024
Licence indiquée dans les métadonnées data.gouv.fr : lov2
Structure Hugging Face
Un jeu de données… See the full description on the dataset page: https://huggingface.co/datasets/Data-Gouv-ML/trafic-routier-sections-de-comptage-departement-du-loiret-2024.ipc_sections_dbflock-demo-healthcare-graph-sectionsflock-demo-healthcare-glucose-sectionsflock-demo-time-series-prediction-sectionspetrology-sections
