datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ner-jsonlPile-NER-type
Intro
Pile-NER-type is a set of GPT-generated data for named entity recognition using the type-based data construction prompt. It was collected by prompting gpt-3.5-turbo-0301 and augmented by negative sampling. Check our project page for more information.
License
Attribution-NonCommercial 4.0 International
person-names-ner
Dataset Card for Person Full Name NER Parsing
This dataset contains 3,383,944 curated and augmented person names, designed specifically for training Token Classification (NER) models. The primary task is to parse a full name string into its FirstName and LastName components, correctly handling multi-word names and different ordering formats.
Dataset Details
Dataset Description
This dataset is built to train robust models that can understand and segment human… See the full description on the dataset page: https://huggingface.co/datasets/ele-sage/person-names-ner.biomed_NER
Biomed NER
This dataset consists of 4,840 manually annotated text records drawn from PubMed abstracts, drug descriptions from the FDA, and patent abstracts. All entities are continuous, and there are no nested entities.
Dataset composition
The dataset contains 4,840 annotated text records distributed across three sources:
Source
Approx. records
Purpose
PubMed abstracts
~4,300
Core biomedical content
FDA drug descriptions
~430
Pharmaceutical text with dense… See the full description on the dataset page: https://huggingface.co/datasets/knowledgator/biomed_NER.llm-ner-extraction
Introduction
This dataset is an extraction of NER data from the wikipedia dataset.
This can be used to fine tune llm models for NER extraction.
ner-eval-predictionsfa-perdt-ner
fa-perdt-ner
Persian named-entity annotations over every sentence of the Persian Universal Dependency
Treebank (UD_Persian-PerDT,
PerUDT v1.0): 29,107 sentences, 494,163 tokens, in PerDT's own train/dev/test
split, on the tokenization a spaCy --merge-subtokens conversion of the treebank produces.
Two configurations, same sentences, same tokens, different labels:
config
labels
entities
what it is
silver (default)
PER LOC ORG DAT MON TIM PCT
15,807
PerDT's own NER layer… See the full description on the dataset page: https://huggingface.co/datasets/Phazel/fa-perdt-ner.medical-privacy-ner
Medical Privacy NER
Created: 2024Creators: Thejan, Chinthani, OshanRepository: ThejanBW/medical-privacy-nerKeywords: named entity recognition (NER) · clinical NER · entity extraction · de-identification · text redaction · patient data masking · patient privacy · protected health information (PHI) · personally identifiable information (PII) · sensitive information detection · clinical text · medical record text · HIPAA-oriented workflows · mobile and on-device models… See the full description on the dataset page: https://huggingface.co/datasets/ThejanBW/medical-privacy-ner.synthengine-cot-edge-case-v1
SynthEngine CoT Edge Case Dataset v1.0
Premium synthetic Chain-of-Thought reasoning data for autonomous driving, robotics, and embodied AI edge cases.
🔗 Full dataset (1000 records) available on Gumroad
This HuggingFace repo contains a free sample (10 records) under CC BY-NC-SA 4.0.
🎯 Why This Dataset?
In 2025, NVIDIA Alpamayo-R1 proved that Chain-of-Causation reasoning improves autonomous driving planning accuracy by +12% and reduces close encounters by -35%.… See the full description on the dataset page: https://huggingface.co/datasets/NeroSeungSan/synthengine-cot-edge-case-v1.nerel_dataset
TeSla NeReL Dataset
Скуп за обучавање модела за обележавање и повезивање именованих ентитета (NER+NEL)
Преко 150.000 реченица анотираних реченица из различитих домена
Named Entity Recognition and Linking (NER+NEL) Model Training Set for Serbian
Over 150,000 annotated sentences from various domains
Editor
Milica Ikonić Nešić
@MilicaIK
Editor… See the full description on the dataset page: https://huggingface.co/datasets/te-sla/nerel_dataset.lanuk-luftqualitaet-ner
LANUK Luftqualität NER — annotierte Sätze aus Fachberichten des LANUK Nordrhein-Westfalen
Deutschsprachiger Datensatz für Named Entity Recognition auf behördlichen Luftqualitätstexten.
450 Sätze aus sechs Fachberichten des Landesamtes für Natur, Umwelt und Klima Nordrhein-Westfalen
(LANUK, vormals LANUV), 407 annotierte Spannen in sechs Klassen. Entstanden als Studienarbeit im
Fach Natural Language Processing an der Fachhochschule Südwestfalen (Betreuung: Prof. Dr.
Christian… See the full description on the dataset page: https://huggingface.co/datasets/fhswf/lanuk-luftqualitaet-ner.greek_legal_ner
Dataset Card for Greek Legal Named Entity Recognition
Dataset Summary
This dataset contains an annotated corpus for named entity recognition in Greek legislations. It is the first of its kind for the Greek language in such an extended form and one of the few that examines legal text in a full spectrum entity recognition.
Supported Tasks and Leaderboards
The dataset supports the task of named entity recognition.
Languages
The language in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/greek_legal_ner.complex_ner
Elephant Labs Complex PII Dataset for Long Contexts and Advanced Anonymization (with Business and Software-related Entities)
Developed by: Elephant Labs
LinkedIn: Elephant Labs
Dataset Size: 20,0000 synthetic documents
Number of tokens in text: 14,140,795 (Tokenized with tiktoken.encoding_for_model("gpt-3.5-turbo"))
Dataset Summary
Purpose: A synthetically generated dataset for advanced NER tasks, supporting both token classification and LLM fine-tuning (enabling… See the full description on the dataset page: https://huggingface.co/datasets/MorryShah/complex_ner.ner-wikipedia-dataset
Wikipediaを用いた日本語の固有表現抽出データセット
GitHub: https://github.com/stockmarkteam/ner-wikipedia-dataset/
LICENSE: CC-BY-SA 3.0
Developed by Stockmark Inc.
esic-nerDataset sintético para treinamento em tarefa de extração de entidades (NER) para uso em classificação de dados pessoais (PII) em formulários e-SIC.
Estatísticas do train split
Summary
samples: 4473
samples_with_any_entity: 3571 (79.83%)
samples_with_any_pii (excludes ORG_JURIDICA, DOC_EMPRESA): 2244 (50.17%)
entity_records_total: 14510
literal_occurrences_total: 14686
Note: ORG_JURIDICA and DOC_EMPRESA are labels but are treated as non-PII (excluded from PII-only… See the full description on the dataset page: https://huggingface.co/datasets/EliMC/esic-ner.ko-pii-ner-100k
한국 PII 특화 학습용 데이터셋 (ko_pii_v1)
한국 고유 식별자와 조사 결합 경계를 정면으로 다루는 한국어 PII NER 학습 데이터셋.
1. 개요
학습용 98,845건 + 외부 평가용 홀드아웃 2,006건
라벨 20종 3티어 / BIO 41 클래스
시드 42, 검증자릿수 정책 invalid, 사용 티어 [1, 2, 3]
포맷: JSONL. {id, text, spans:[{start,end,label,value}], meta}
이 레포의 NERPreprocessor span 포맷과 동일해 학습 경로 수정 없이 사용 가능
1-1. 이 데이터셋으로 학습한 모델
seongyeon1/ko-pii-ner-roberta-base
(klue/roberta-base 파인튜닝, CC-BY-SA-4.0)
학습에 한 번도 쓰이지 않은 KDPII 공식 test split 기준 F5 0.9344.
내부… See the full description on the dataset page: https://huggingface.co/datasets/seongyeon1/ko-pii-ner-100k.GeoEDdA-NER
GeoEDdA-NER: A Gold Standard Dataset for Geo-semantic Annotation of Diderot & d’Alembert’s Encyclopédie
Dataset Description
Authors: Ludovic Moncla, Katherine McDonough and Denis Vigier in the framework of the GEODE project.
Data source: ARTFL Encyclopédie Project, University of Chicago
Github repository: https://github.com/GEODE-project/ner-spancat-edda
Language: French
License: cc-by-nc-4.0
Zenodo repository: https://zenodo.org/records/10530177
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GEODE/GeoEDdA-NER.Galician_NER
Galician NER test
Dataset created by combining four galician datasets for Named Entity Recognition, annotated according to the new standards for NER annotations:
corNER: Updated version of the original corNER dataset, which was created by annotating for NER the corga dataset.
LREC: Updated version to keep up with the new standards for NER annotations.
PUD: Dataset created by annotating for NER the Galician PUD treebank.
TreeGal: Dataset created by annotating for NER the TreeGal… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Galician_NER.buster-expanded-ner
BUSTER Expanded NER
BUSTER Expanded NER is a derived annotation layer over the 3,779 manually
annotated English documents in the gold corpus of
expertai/BUSTER. It replaces
BUSTER's transaction-role ontology with four general named-entity labels and
materially expands mention coverage within each document.
The release is designed as a tokenizer-neutral source for training flat NER
models. It stores exact character spans rather than tokenizer-specific tags, so
BIOES labels can be… See the full description on the dataset page: https://huggingface.co/datasets/PITTI/buster-expanded-ner.FG_NER_dataBF_NER_datasets
BF_NER Training Datasets
BIO-tagged training data for the BF_NER model.
Dataset Description
This dataset contains 86,252 sentences with BIO tags for geographic Named Entity Recognition in French, specifically for Burkina Faso administrative entities.
Splits
Split
Sentences
Description
Train
59,900
Training set
Validation
14,758
Validation set for hyperparameter tuning
Test
11,594
Held-out test set with ~20% unseen entities
Entity… See the full description on the dataset page: https://huggingface.co/datasets/CharlesAbdoulaye/BF_NER_datasets.gemini-bavarian-ner-v0.1
Gemini-powered Bavarian NER Dataset
Inspired by GLiNER models and its used datasets, we present a Gemini-powered NER Dataset for Bavarian.
The dataset currently features 116,075 sentences from Bavarian Wikipedia, where named entities are found using Gemini 2.0 Flash.
Changelog
03.07.2025: Initial version of the dataset and public release.
Template
Thankfully, the GLiNER-X community shared their prompt for generating datasets that were used for training the… See the full description on the dataset page: https://huggingface.co/datasets/bavarian-nlp/gemini-bavarian-ner-v0.1.data-use-ner
Data-use-ner (human holdout)
GLiNER-format human-adjudicated holdout: 473 spans — annotator190 (190, origin=fcv_pads_east_africa) + jdc283 (283, origin=jdc_operational). Never trained on.
Source: rafmacalaba/datause-displacement-reviewed holdout (gliner_reviewed token spans + readable_reviewed passages, v2.4 labels) with v3 probe head_score (outputs/gliner_datause_v3_probe_human473.jsonl).
Columns
text (full passage = " ".join(tokenized_text); span char offsets… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-ner.jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.4-details
Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.4
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.4-details.nerd-knowledge-api
NerdOptimize Dataset (v1.0.0)
English dataset for SEO (Data‑Driven) and AI Search / AEO by NerdOptimize (Bangkok, TH).Built for GitHub, Hugging Face, and on‑site deployment, so LLMs can learn/cite the brand.
Structure
data/*.json → core machine‑readable data (ICPs, services, case studies, frameworks, articles, labels, metadata, processing steps)
server.js / openapi.json → tiny Express API to serve the dataset
schema-dataset.jsonld → Dataset JSON‑LD for Google Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NerdOptimize/nerd-knowledge-api.Chinese_Few-shot_NERIndian_Legal_NER_Datasetalgerian-realestate-ner-dataset
Algerian-realestate-NER-dataset
Dataset Description
This is a specialized Named Entity Recognition dataset extracted from the complex reality of the Algerian digital real-estate market in Facebook groups, it contains 13 labeled entity and 7138 training example
Real estate advertisements in Algeria (found on Facebook groups) are unstructured , noisy and bloated with code-switching between Algerian Darja (dialect), Arabizi, Standard Arabic, and French, This dataset… See the full description on the dataset page: https://huggingface.co/datasets/81melody/algerian-realestate-ner-dataset.rubai-NER-150K-Personal
Rubai NER Dataset - Personal Information Detection (Synthetic)
A dataset for training Named Entity Recognition (NER) models to detect personal information in Uzbek and Russian text. All Data Synthetic, no contains real personal information!
Dataset Description
This dataset contains 142,704 annotated examples for detecting personal information entities in informal Uzbek and Russian text (Latin and Cyrillic scripts).
Supported Entity Types
Entity… See the full description on the dataset page: https://huggingface.co/datasets/islomov/rubai-NER-150K-Personal.Pile-NER-definition
Intro
Pile-NER-definition is a set of GPT-generated data for named entity recognition using the definition-based data construction prompt. It was collected by prompting gpt-3.5-turbo-0301 and augmented by negative sampling. Check our project page for more information.
License
Attribution-NonCommercial 4.0 International
