Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /ner-jsonltext10K<n<100K0 likes7.3k downloads1y agoHugging Face02Universal-NER /Pile-NER-type Intro Pile-NER-type is a set of GPT-generated data for named entity recognition using the type-based data construction prompt. It was collected by prompting gpt-3.5-turbo-0301 and augmented by negative sampling. Check our project page for more information. License Attribution-NonCommercial 4.0 International text10K<n<100K29 likes526 downloads3y agoHugging Face03ele-sage /person-names-ner Dataset Card for Person Full Name NER Parsing This dataset contains 3,383,944 curated and augmented person names, designed specifically for training Token Classification (NER) models. The primary task is to parse a full name string into its FirstName and LastName components, correctly handling multi-word names and different ordering formats. Dataset Details Dataset Description This dataset is built to train robust models that can understand and segment human… See the full description on the dataset page: https://huggingface.co/datasets/ele-sage/person-names-ner.texttoken-classification1M<n<10M3 likes417 downloads1y agoHugging Face04knowledgator /biomed_NER Biomed NER This dataset consists of 4,840 manually annotated text records drawn from PubMed abstracts, drug descriptions from the FDA, and patent abstracts. All entities are continuous, and there are no nested entities. Dataset composition The dataset contains 4,840 annotated text records distributed across three sources: Source Approx. records Purpose PubMed abstracts ~4,300 Core biomedical content FDA drug descriptions ~430 Pharmaceutical text with dense… See the full description on the dataset page: https://huggingface.co/datasets/knowledgator/biomed_NER.texttoken-classification1K<n<10K13 likes326 downloads6mo agoHugging Face05vangheem /llm-ner-extraction Introduction This dataset is an extraction of NER data from the wikipedia dataset. This can be used to fine tune llm models for NER extraction. text10K<n<100K0 likes259 downloads1y agoHugging Face06impresso-project /ner-eval-predictionstabular100K<n<1M0 likes243 downloads3mo agoHugging Face07Phazel /fa-perdt-ner fa-perdt-ner Persian named-entity annotations over every sentence of the Persian Universal Dependency Treebank (UD_Persian-PerDT, PerUDT v1.0): 29,107 sentences, 494,163 tokens, in PerDT's own train/dev/test split, on the tokenization a spaCy --merge-subtokens conversion of the treebank produces. Two configurations, same sentences, same tokens, different labels: config labels entities what it is silver (default) PER LOC ORG DAT MON TIM PCT 15,807 PerDT's own NER layer… See the full description on the dataset page: https://huggingface.co/datasets/Phazel/fa-perdt-ner.texttoken-classification10K<n<100K1 likes196 downloads18d agoHugging Face08ThejanBW /medical-privacy-ner Medical Privacy NER Created: 2024Creators: Thejan, Chinthani, OshanRepository: ThejanBW/medical-privacy-nerKeywords: named entity recognition (NER) · clinical NER · entity extraction · de-identification · text redaction · patient data masking · patient privacy · protected health information (PHI) · personally identifiable information (PII) · sensitive information detection · clinical text · medical record text · HIPAA-oriented workflows · mobile and on-device models… See the full description on the dataset page: https://huggingface.co/datasets/ThejanBW/medical-privacy-ner.texttoken-classification1K<n<10K2 likes137 downloads15d agoHugging Face09NeroSeungSan /synthengine-cot-edge-case-v1 SynthEngine CoT Edge Case Dataset v1.0 Premium synthetic Chain-of-Thought reasoning data for autonomous driving, robotics, and embodied AI edge cases. 🔗 Full dataset (1000 records) available on Gumroad This HuggingFace repo contains a free sample (10 records) under CC BY-NC-SA 4.0. 🎯 Why This Dataset? In 2025, NVIDIA Alpamayo-R1 proved that Chain-of-Causation reasoning improves autonomous driving planning accuracy by +12% and reduces close encounters by -35%.… See the full description on the dataset page: https://huggingface.co/datasets/NeroSeungSan/synthengine-cot-edge-case-v1.text10K<n<100K0 likes122 downloads4mo agoHugging Face10te-sla /nerel_dataset TeSla NeReL Dataset Скуп за обучавање модела за обележавање и повезивање именованих ентитета (NER+NEL) Преко 150.000 реченица анотираних реченица из различитих домена Named Entity Recognition and Linking (NER+NEL) Model Training Set for Serbian Over 150,000 annotated sentences from various domains Editor Milica Ikonić Nešić @MilicaIK Editor… See the full description on the dataset page: https://huggingface.co/datasets/te-sla/nerel_dataset.texttoken-classification100K<n<1M0 likes93 downloads22d agoHugging Face11fhswf /lanuk-luftqualitaet-ner LANUK Luftqualität NER — annotierte Sätze aus Fachberichten des LANUK Nordrhein-Westfalen Deutschsprachiger Datensatz für Named Entity Recognition auf behördlichen Luftqualitätstexten. 450 Sätze aus sechs Fachberichten des Landesamtes für Natur, Umwelt und Klima Nordrhein-Westfalen (LANUK, vormals LANUV), 407 annotierte Spannen in sechs Klassen. Entstanden als Studienarbeit im Fach Natural Language Processing an der Fachhochschule Südwestfalen (Betreuung: Prof. Dr. Christian… See the full description on the dataset page: https://huggingface.co/datasets/fhswf/lanuk-luftqualitaet-ner.texttoken-classificationn<1K0 likes92 downloads19d agoHugging Face12joelniklaus /greek_legal_ner Dataset Card for Greek Legal Named Entity Recognition Dataset Summary This dataset contains an annotated corpus for named entity recognition in Greek legislations. It is the first of its kind for the Greek language in such an extended form and one of the few that examines legal text in a full spectrum entity recognition. Supported Tasks and Leaderboards The dataset supports the task of named entity recognition. Languages The language in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/greek_legal_ner.texttoken-classification10K<n<100K0 likes91 downloads3y agoHugging Face13MorryShah /complex_ner Elephant Labs Complex PII Dataset for Long Contexts and Advanced Anonymization (with Business and Software-related Entities) Developed by: Elephant Labs LinkedIn: Elephant Labs Dataset Size: 20,0000 synthetic documents Number of tokens in text: 14,140,795 (Tokenized with tiktoken.encoding_for_model("gpt-3.5-turbo")) Dataset Summary Purpose: A synthetically generated dataset for advanced NER tasks, supporting both token classification and LLM fine-tuning (enabling… See the full description on the dataset page: https://huggingface.co/datasets/MorryShah/complex_ner.texttoken-classification10K<n<100K2 likes91 downloads2y agoHugging Face14stockmark /ner-wikipedia-dataset Wikipediaを用いた日本語の固有表現抽出データセット GitHub: https://github.com/stockmarkteam/ner-wikipedia-dataset/ LICENSE: CC-BY-SA 3.0 Developed by Stockmark Inc. texttoken-classification1K<n<10K14 likes86 downloads3y agoHugging Face15EliMC /esic-nerDataset sintético para treinamento em tarefa de extração de entidades (NER) para uso em classificação de dados pessoais (PII) em formulários e-SIC. Estatísticas do train split Summary samples: 4473 samples_with_any_entity: 3571 (79.83%) samples_with_any_pii (excludes ORG_JURIDICA, DOC_EMPRESA): 2244 (50.17%) entity_records_total: 14510 literal_occurrences_total: 14686 Note: ORG_JURIDICA and DOC_EMPRESA are labels but are treated as non-PII (excluded from PII-only… See the full description on the dataset page: https://huggingface.co/datasets/EliMC/esic-ner.texttoken-classification1K<n<10K0 likes82 downloads8mo agoHugging Face16seongyeon1 /ko-pii-ner-100k 한국 PII 특화 학습용 데이터셋 (ko_pii_v1) 한국 고유 식별자와 조사 결합 경계를 정면으로 다루는 한국어 PII NER 학습 데이터셋. 1. 개요 학습용 98,845건 + 외부 평가용 홀드아웃 2,006건 라벨 20종 3티어 / BIO 41 클래스 시드 42, 검증자릿수 정책 invalid, 사용 티어 [1, 2, 3] 포맷: JSONL. {id, text, spans:[{start,end,label,value}], meta} 이 레포의 NERPreprocessor span 포맷과 동일해 학습 경로 수정 없이 사용 가능 1-1. 이 데이터셋으로 학습한 모델 seongyeon1/ko-pii-ner-roberta-base (klue/roberta-base 파인튜닝, CC-BY-SA-4.0) 학습에 한 번도 쓰이지 않은 KDPII 공식 test split 기준 F5 0.9344. 내부… See the full description on the dataset page: https://huggingface.co/datasets/seongyeon1/ko-pii-ner-100k.texttoken-classification100K<n<1M0 likes78 downloads1mo agoHugging Face17GEODE /GeoEDdA-NER GeoEDdA-NER: A Gold Standard Dataset for Geo-semantic Annotation of Diderot & d’Alembert’s Encyclopédie Dataset Description Authors: Ludovic Moncla, Katherine McDonough and Denis Vigier in the framework of the GEODE project. Data source: ARTFL Encyclopédie Project, University of Chicago Github repository: https://github.com/GEODE-project/ner-spancat-edda Language: French License: cc-by-nc-4.0 Zenodo repository: https://zenodo.org/records/10530177 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GEODE/GeoEDdA-NER.texttoken-classification1K<n<10K0 likes76 downloads1y agoHugging Face18proxectonos /Galician_NER Galician NER test Dataset created by combining four galician datasets for Named Entity Recognition, annotated according to the new standards for NER annotations: corNER: Updated version of the original corNER dataset, which was created by annotating for NER the corga dataset. LREC: Updated version to keep up with the new standards for NER annotations. PUD: Dataset created by annotating for NER the Galician PUD treebank. TreeGal: Dataset created by annotating for NER the TreeGal… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Galician_NER.texttoken-classification1K<n<10K0 likes75 downloads10mo agoHugging Face19PITTI /buster-expanded-ner BUSTER Expanded NER BUSTER Expanded NER is a derived annotation layer over the 3,779 manually annotated English documents in the gold corpus of expertai/BUSTER. It replaces BUSTER's transaction-role ontology with four general named-entity labels and materially expands mention coverage within each document. The release is designed as a tokenizer-neutral source for training flat NER models. It stores exact character spans rather than tokenizer-specific tags, so BIOES labels can be… See the full description on the dataset page: https://huggingface.co/datasets/PITTI/buster-expanded-ner.texttoken-classification1K<n<10K0 likes70 downloads2mo agoHugging Face20Hnin /FG_NER_datatextn<1K0 likes70 downloads22d agoHugging Face21CharlesAbdoulaye /BF_NER_datasets BF_NER Training Datasets BIO-tagged training data for the BF_NER model. Dataset Description This dataset contains 86,252 sentences with BIO tags for geographic Named Entity Recognition in French, specifically for Burkina Faso administrative entities. Splits Split Sentences Description Train 59,900 Training set Validation 14,758 Validation set for hyperparameter tuning Test 11,594 Held-out test set with ~20% unseen entities Entity… See the full description on the dataset page: https://huggingface.co/datasets/CharlesAbdoulaye/BF_NER_datasets.texttoken-classification10K<n<100K0 likes68 downloads8mo agoHugging Face22bavarian-nlp /gemini-bavarian-ner-v0.1 Gemini-powered Bavarian NER Dataset Inspired by GLiNER models and its used datasets, we present a Gemini-powered NER Dataset for Bavarian. The dataset currently features 116,075 sentences from Bavarian Wikipedia, where named entities are found using Gemini 2.0 Flash. Changelog 03.07.2025: Initial version of the dataset and public release. Template Thankfully, the GLiNER-X community shared their prompt for generating datasets that were used for training the… See the full description on the dataset page: https://huggingface.co/datasets/bavarian-nlp/gemini-bavarian-ner-v0.1.texttoken-classification100K<n<1M0 likes65 downloads1y agoHugging Face23rafmacalaba /data-use-ner Data-use-ner (human holdout) GLiNER-format human-adjudicated holdout: 473 spans — annotator190 (190, origin=fcv_pads_east_africa) + jdc283 (283, origin=jdc_operational). Never trained on. Source: rafmacalaba/datause-displacement-reviewed holdout (gliner_reviewed token spans + readable_reviewed passages, v2.4 labels) with v3 probe head_score (outputs/gliner_datause_v3_probe_human473.jsonl). Columns text (full passage = " ".join(tokenized_text); span char offsets… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-ner.tabulartoken-classification10K<n<100K0 likes65 downloads1mo agoHugging Face24open-llm-leaderboard /jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.4-detailsgated Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.4 Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.4 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.4-details.tabular10K<n<100K0 likes61 downloads2y agoHugging Face25NerdOptimize /nerd-knowledge-api NerdOptimize Dataset (v1.0.0) English dataset for SEO (Data‑Driven) and AI Search / AEO by NerdOptimize (Bangkok, TH).Built for GitHub, Hugging Face, and on‑site deployment, so LLMs can learn/cite the brand. Structure data/*.json → core machine‑readable data (ICPs, services, case studies, frameworks, articles, labels, metadata, processing steps) server.js / openapi.json → tiny Express API to serve the dataset schema-dataset.jsonld → Dataset JSON‑LD for Google Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NerdOptimize/nerd-knowledge-api.textzero-shot-classificationn<1K0 likes57 downloads11mo agoHugging Face26aleversn /Chinese_Few-shot_NERtext10K<n<100K1 likes48 downloads2y agoHugging Face27AjayMukundS /Indian_Legal_NER_Datasettext10K<n<100K1 likes47 downloads2y agoHugging Face2881melody /algerian-realestate-ner-dataset Algerian-realestate-NER-dataset Dataset Description This is a specialized Named Entity Recognition dataset extracted from the complex reality of the Algerian digital real-estate market in Facebook groups, it contains 13 labeled entity and 7138 training example Real estate advertisements in Algeria (found on Facebook groups) are unstructured , noisy and bloated with code-switching between Algerian Darja (dialect), Arabizi, Standard Arabic, and French, This dataset… See the full description on the dataset page: https://huggingface.co/datasets/81melody/algerian-realestate-ner-dataset.texttoken-classification1K<n<10K1 likes47 downloads6mo agoHugging Face29islomov /rubai-NER-150K-Personal Rubai NER Dataset - Personal Information Detection (Synthetic) A dataset for training Named Entity Recognition (NER) models to detect personal information in Uzbek and Russian text. All Data Synthetic, no contains real personal information! Dataset Description This dataset contains 142,704 annotated examples for detecting personal information entities in informal Uzbek and Russian text (Latin and Cyrillic scripts). Supported Entity Types Entity… See the full description on the dataset page: https://huggingface.co/datasets/islomov/rubai-NER-150K-Personal.text100K<n<1M2 likes46 downloads8mo agoHugging Face30Universal-NER /Pile-NER-definition Intro Pile-NER-definition is a set of GPT-generated data for named entity recognition using the definition-based data construction prompt. It was collected by prompting gpt-3.5-turbo-0301 and augmented by negative sampling. Check our project page for more information. License Attribution-NonCommercial 4.0 International text10K<n<100K20 likes45 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.