Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CyberMax-tools /riskroll-sec-10k-10q-sections Riskroll: SEC 10-K and 10-Q sections as clean text Need it fresh, filtered or via API? This free file is a snapshot (10-K/10-Q sections up to the last refresh), last updated 2026-09-24. Insidewell on Apify ($0.004 per insider transaction): pulls today's SEC Form 4 trades for your own watchlist, filtered by buy/sell and size, with cluster-buy alerts on a schedule. Using it at work? Commercial license + support (from $49/year): invoice, PDF licence certificate, named… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/riskroll-sec-10k-10q-sections.tabulartext-classification1K<n<10K0 likes650 downloads7h agoHugging Face02GSMS-B /indian-legal-sections-bns-bnss-bsa-2023 🏛️ Indian Legal Sections — BNS · BNSS · BSA 2023 The First Structured, Unified JSON Dataset of Modern Indian Criminal Law 📖 Dataset Summary This dataset contains 1,059 fully structured and verified sections extracted, parsed, and unified from India's three landmark criminal justice reform acts passed in December 2023. These three acts together replaced the colonial-era Indian Penal Code (IPC, 1860), the Code of Criminal Procedure… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/indian-legal-sections-bns-bnss-bsa-2023.textquestion-answering1K<n<10K2 likes523 downloads3mo agoHugging Face03justram /sections Dataset Card for "sections" More Information needed text10M<n<100M0 likes384 downloads3y agoHugging Face04pjramg /Coffee_leaves_sections_FO Dataset Card for coffee_leaves_anomalib_2 This is a FiftyOne dataset with 35962 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("pjramg/Coffee_leaves_sections_FO") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/pjramg/Coffee_leaves_sections_FO.imageimage-classification10K<n<100K0 likes240 downloads1y agoHugging Face05vinczematyas /stranger_sections_2image1K<n<10K0 likes144 downloads2y agoHugging Face06sentence-transformers /wikipedia-sections Dataset Card for Wikipedia Sections This dataset contains pairs and triplets that can be used to train and finetune Sentence Transformer embedding models. The dataset originates from Dor et al., and was downloaded from this download link. Notably, the "anchor" column contains sentences from Wikipedia, wheras the "positive" column contains other sentences from the same section. The "negative" column contains sentences from other sections. Dataset Subsets… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/wikipedia-sections.textfeature-extraction1M<n<10M1 likes113 downloads2y agoHugging Face07BramVanroy /finewiki-en-sections-propella Propella annotations for level-2 English Wikipedia sections English Wikipedia pages from the en config of HuggingFaceFW/finewiki, split into one row per level-2 section and annotated with ellamind/propella-1-4b. It is the English counterpart of BramVanroy/finewiki-nl-sections-propella, and it was built for the training data of a Dutch and English embedding model. How it was built Every page is split at its ## headings. A section is the text up to the next… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/finewiki-en-sections-propella.tabular1M<n<10M1 likes79 downloads13d agoHugging Face08random-sequence /flock-demo-critical-infra-sectionstabular10K<n<100K0 likes75 downloads8mo agoHugging Face09ulab-ai /arxiv-sectionstabular100K<n<1M0 likes73 downloads11mo agoHugging Face10nomeda-lab /hindawi-arabic-sections Hindawi Arabic Books — Sections Dataset A cleaned, section-level dataset of Arabic books from Hindawi.org, prepared for NLP training and research. Source Original books scraped from Hindawi.org, a non-profit foundation providing free Arabic books. The dataset covers categories including literature, philosophy, history, science, psychology, and more. Cleaning Pipeline Scraped book content section-by-section from Hindawi.org Removed English / Latin text and… See the full description on the dataset page: https://huggingface.co/datasets/nomeda-lab/hindawi-arabic-sections.text10K<n<100K2 likes67 downloads6mo agoHugging Face11BramVanroy /finewiki-nl-sections-propella Propella annotations for second-level, Dutch Wikipedia sections This dataset is an exploded and annotated version of the Dutch portion of FineWiki. Articles were split so that each second-level section (##) is now its own sample. Any introductory sections (between the top heading and the first sub-section) is not included. Before annotation, the text was truncated to the first 50_000 characters, as recommended by the Propella README. Intended use This dataset may… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/finewiki-nl-sections-propella.tabular100K<n<1M0 likes53 downloads4mo agoHugging Face12random-sequence /flock-demo-object-detection-sectionsimage1K<n<10K0 likes35 downloads8mo agoHugging Face13drja23 /thin-sections-wikimediaimagen<1K1 likes34 downloads2y agoHugging Face14random-sequence /flock-demo-defense-graph-sectionstext1K<n<10K0 likes32 downloads8mo agoHugging Face15random-sequence /flock-demo-slm-qwen3-0-6b-sectionstext1K<n<10K0 likes29 downloads8mo agoHugging Face16random-sequence /flock-demo-llm-finetuning-sectionstext1K<n<10K0 likes27 downloads8mo agoHugging Face17Abdallah4Zain /hindawi-arabic-sections Hindawi Arabic Books — Sections Dataset A cleaned, section-level dataset of Arabic books from Hindawi.org, prepared for NLP training and research. Source Original books scraped from Hindawi.org, a non-profit foundation providing free Arabic books. The dataset covers categories including literature, philosophy, history, science, psychology, and more. Cleaning Pipeline Scraped book content section-by-section from Hindawi.org Removed English / Latin text and… See the full description on the dataset page: https://huggingface.co/datasets/Abdallah4Zain/hindawi-arabic-sections.text10K<n<100K0 likes26 downloads6mo agoHugging Face18stormsidali2001 /IMRAD-sections-clf-gemini-augmented Dataset Card for IMRAD Classification Dataset (100k Rows) Dataset Name: IMRAD Classification Dataset (100k Rows) Dataset Description: This dataset contains approximately 100,000 sentences extracted from scientific research papers and labeled according to their corresponding IMRAD (Introduction, Methods, Results, and Discussion) sections. The data was initially sourced from the unarXive_imrad_clf dataset on Hugging Face and expanded using data augmentation techniques. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/stormsidali2001/IMRAD-sections-clf-gemini-augmented.text100K<n<1M0 likes24 downloads2y agoHugging Face19davanstrien /model_cards_with_readmes_sections Dataset Card for "model_cards_with_readmes_sections" More Information needed text10K<n<100K0 likes23 downloads4y agoHugging Face20kmeanskaran /ipc-sectionstextn<1K2 likes23 downloads2y agoHugging Face21random-sequence /flock-demo-automatic-speech-recognition-sectionsaudion<1K0 likes23 downloads8mo agoHugging Face22dogoctor /thin-sections-wikimediaimagen<1K0 likes23 downloads4mo agoHugging Face23random-sequence /flock-demo-finance-sentiment-sectionstext1K<n<10K0 likes20 downloads8mo agoHugging Face24random-sequence /flock-demo-graph-neural-network-sectionstext1K<n<10K0 likes20 downloads8mo agoHugging Face25Data-Gouv-ML /trafic-routier-sections-de-comptage-departement-du-loiret-2024 Trafic routier - Sections de comptage - Département du Loiret - 2024 Source Source officielle : https://www.data.gouv.fr/datasets/trafic-routier-sections-de-comptage-departement-du-loiret-2024 Identifiant du jeu de données data.gouv.fr : 68f9746546f42706c79cb4c8 Slug data.gouv.fr : trafic-routier-sections-de-comptage-departement-du-loiret-2024 Licence indiquée dans les métadonnées data.gouv.fr : lov2 Structure Hugging Face Un jeu de données… See the full description on the dataset page: https://huggingface.co/datasets/Data-Gouv-ML/trafic-routier-sections-de-comptage-departement-du-loiret-2024.text1K<n<10K0 likes20 downloads4mo agoHugging Face26Rish871 /ipc_sections_dbtexttext-generationn<1K1 likes16 downloads2y agoHugging Face27random-sequence /flock-demo-healthcare-graph-sectionstext1K<n<10K0 likes16 downloads8mo agoHugging Face28random-sequence /flock-demo-healthcare-glucose-sectionstabular1K<n<10K0 likes16 downloads8mo agoHugging Face29random-sequence /flock-demo-time-series-prediction-sectionstabular1K<n<10K0 likes16 downloads8mo agoHugging Face30volvoDon /petrology-sectionsimagen<1K1 likes15 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.