datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ner-eval-predictionsdata-use-ner
Data-use-ner (human holdout)
GLiNER-format human-adjudicated holdout: 473 spans — annotator190 (190, origin=fcv_pads_east_africa) + jdc283 (283, origin=jdc_operational). Never trained on.
Source: rafmacalaba/datause-displacement-reviewed holdout (gliner_reviewed token spans + readable_reviewed passages, v2.4 labels) with v3 probe head_score (outputs/gliner_datause_v3_probe_human473.jsonl).
Columns
text (full passage = " ".join(tokenized_text); span char offsets… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-ner.jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.4-details
Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.4
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.4-details.jeffmeloy__Qwen-7B-nerd-uncensored-v1.0-details
Dataset Card for Evaluation run of jeffmeloy/Qwen-7B-nerd-uncensored-v1.0
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen-7B-nerd-uncensored-v1.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen-7B-nerd-uncensored-v1.0-details.datause-ner
Datause NER (catch-all DATA_MENTION + probe configs)
Catch-all NER views over rafmacalaba/datause-probe-v3 passages (29,346 spans grouped into passage examples). Single entity type DATA_MENTION: every candidate span is tagged, keeps and drops alike — the probe head (not NER tags) owns the keep/drop boundary. No NAMED/DESCRIPTIVE/VAGUE subtypes, no NON_MENTION.
Per-origin thresholds (head best-F1, published holdout sweep)
origin
threshold… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-ner.jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.7-details
Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.7
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.7
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.7-details.jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.0-details
Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.0
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.0-details.jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.2-details
Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.2
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.2-details.nerel_simplejeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.5-details
Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.5
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.5-details.jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.8-details
Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.8
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.8-details.data_base_nerjeffmeloy__Qwen2.5-7B-nerd-uncensored-v0.9-details
Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v0.9
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v0.9
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v0.9-details.jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.1-details
Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.1
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.1-details.jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.3-details
Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.3
Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.3-details.ZeroXClem__Qwen2.5-7B-HomerAnvita-NerdMix-details
Dataset Card for Evaluation run of ZeroXClem/Qwen2.5-7B-HomerAnvita-NerdMix
Dataset automatically created during the evaluation run of model ZeroXClem/Qwen2.5-7B-HomerAnvita-NerdMix
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ZeroXClem__Qwen2.5-7B-HomerAnvita-NerdMix-details.Aashraf995__Creative-7B-nerd-details
Dataset Card for Evaluation run of Aashraf995/Creative-7B-nerd
Dataset automatically created during the evaluation run of model Aashraf995/Creative-7B-nerd
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Aashraf995__Creative-7B-nerd-details.netcat420__Qwen2.5-7B-nerd-uncensored-v0.9-MFANN-details
Dataset Card for Evaluation run of netcat420/Qwen2.5-7B-nerd-uncensored-v0.9-MFANN
Dataset automatically created during the evaluation run of model netcat420/Qwen2.5-7B-nerd-uncensored-v0.9-MFANN
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/netcat420__Qwen2.5-7B-nerd-uncensored-v0.9-MFANN-details.job-market-ner-fr-ar
Job-Market Entity Recognition Gold Sets (French and Arabic)
This repository holds the gold evaluation sets for job-market entity recognition in French and Modern Standard Arabic that accompany the paper Cross-Lingual Teacher–Student Distillation of Job-Market Entity Recognition for French and Arabic (IANLP 2026). Each set contains 500 job postings annotated with eight entity types; to our knowledge, the Arabic set is the first job-market entity-recognition dataset for Arabic.… See the full description on the dataset page: https://huggingface.co/datasets/AchrafSoltani/job-market-ner-fr-ar.Nitral-AI__Nera_Noctis-12B-details
Dataset Card for Evaluation run of Nitral-AI/Nera_Noctis-12B
Dataset automatically created during the evaluation run of model Nitral-AI/Nera_Noctis-12B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nitral-AI__Nera_Noctis-12B-details.eval-gliner2-ner-mit_restaurant-lossablation
mit_restaurant — loss ablation (5 config x 3 seed)
base model: fastino/gliner2-multi-v1
dataset: quynong/mit_restaurant (8 nhan), eval tren split test
train: 6 epochs, batch 16, early-stopping patience 2, seeds [42, 43, 44]
eval: --threshold 0.7 --use-desc --schema-case original --no-extra-val (giong nhau cho ca 5 config)
Ket qua (micro F1 %, mean +/- std tren seed)
mode
config
n seeds
micro F1
macro F1
micro P
micro R
lenient
bce
3
88.69 +/- 0.20… See the full description on the dataset page: https://huggingface.co/datasets/AITeamUIT/eval-gliner2-ner-mit_restaurant-lossablation.visco-nerfnaval_nernetcat420__Qwen2.5-7b-nerd-uncensored-MFANN-slerp-details
Dataset Card for Evaluation run of netcat420/Qwen2.5-7b-nerd-uncensored-MFANN-slerp
Dataset automatically created during the evaluation run of model netcat420/Qwen2.5-7b-nerd-uncensored-MFANN-slerp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/netcat420__Qwen2.5-7b-nerd-uncensored-MFANN-slerp-details.haddas-ner-ti
haddas-ner-ti
Tigrinya named-entity records with multi-spelling variants, English forms, frequencies, topics and issue dates. Built from the cumulative bilingual alignment memory.
Source
Derived from the Haddas Eritrea newspaper archive: 63 PDF issues processed
by the haddas-eritrea pipeline (extract -> clean -> segment -> translate -> label).
Generated: 2026-04-26 12:21 UTC
Row count: 57
Schema: canonical_key, canonical_tigrinya_variant, tigrinya_variants… See the full description on the dataset page: https://huggingface.co/datasets/SIMBA9657/haddas-ner-ti.eval-gliner2-ner-fin-lossablation
fin — loss ablation (5 config x 3 seed)
base model: fastino/gliner2-multi-v1
dataset: quynong/fin (4 nhan), eval tren split test
train: 8 epochs, batch 16, early-stopping patience 2, seeds [42, 43, 44]
eval: --threshold 0.7 --use-desc --schema-case original --no-extra-val (giong nhau cho ca 5 config)
Ket qua (micro F1 %, mean +/- std tren seed)
mode
config
n seeds
micro F1
macro F1
micro P
micro R
lenient
bce
3
77.88 +/- 4.74
41.96 +/- 8.81
88.63… See the full description on the dataset page: https://huggingface.co/datasets/AITeamUIT/eval-gliner2-ner-fin-lossablation.navel_ner2eval-gliner2-ner-conll2003-lossablation
conll2003 — loss ablation (5 config x 3 seed)
base model: fastino/gliner2-multi-v1
dataset: quynong/conll2003 (4 nhan), eval tren split test
train: 6 epochs, batch 16, early-stopping patience 2, seeds [42, 43, 44]
eval: --threshold 0.7 --use-desc --schema-case original --no-extra-val (giong nhau cho ca 5 config)
Ket qua (micro F1 %, mean +/- std tren seed)
mode
config
n seeds
micro F1
macro F1
micro P
micro R
lenient
bce
3
89.93 +/- 0.39
88.22 +/-… See the full description on the dataset page: https://huggingface.co/datasets/AITeamUIT/eval-gliner2-ner-conll2003-lossablation.eval-gliner2-ner-ontonotes5-lossablation
ontonotes5 — loss ablation (5 config x 3 seed)
base model: fastino/gliner2-multi-v1
dataset: quynong/ontonotes5 (18 nhan), eval tren split test
train: 3 epochs, batch 16, early-stopping patience 2, seeds [42, 43, 44]
eval: --threshold 0.7 --use-desc --schema-case original --no-extra-val (giong nhau cho ca 5 config)
Ket qua (micro F1 %, mean +/- std tren seed)
mode
config
n seeds
micro F1
macro F1
micro P
micro R
lenient
bce
3
90.58 +/- 0.12
80.75 +/-… See the full description on the dataset page: https://huggingface.co/datasets/AITeamUIT/eval-gliner2-ner-ontonotes5-lossablation.ner_dataset
