datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nerfFSCM_Flood_nerfner-fashion-brands
Ner Fashion Brands
This dataset originally appear as part of
this tutorial. The goal
of the dataset is to detect fashion brands in Reddit Comments.
For more details, be sure to read this blogpost.
NER_financial_user_assistantmerged_test_nerfair_processed
Benchmark Merged dataset - Test
This dataset was generated by the Data preprocessing step of the NERFAIR workflow (More information: https://github.com/YasCoMa/ner-fair-workflow )
It contains the test data of the following datasets:
ncbi: Doğan, Rezarta Islamaj, Robert Leaman, and Zhiyong Lu. 2014. “NCBI Disease Corpus: A Resource for Disease Name Recognition and Concept Normalization.” Journal of Biomedical Informatics 47 (February): 1–10.
bc5cdr: Li, Jiao, Yueping Sun, Robin J.… See the full description on the dataset page: https://huggingface.co/datasets/yasmmin/merged_test_nerfair_processed.biored_nerfair_processed
Benchmark dataset biored
This dataset was generated by the Data preprocessing step of the NERFAIR workflow (More information: https://github.com/YasCoMa/ner-fair-workflow )
Original dataset: Luo, Ling, Po-Ting Lai, Chih-Hsuan Wei, Cecilia N. Arighi, and Zhiyong Lu. 2022. “BioRED: A Rich Biomedical Relation Extraction Dataset.” Briefings in Bioinformatics 23 (5): bbac282.
ner_filtered_generated_data_gmmvisco-nerfner-furniture-namesCeunerfpico-human-corpus_nerfair_processed
Benchmark dataset PICO
This dataset was generated by the Data preprocessing step of the NERFAIR workflow (More information: https://github.com/YasCoMa/ner-fair-workflow )
Original dataset: https://github.com/sociocom/PICO-Corpus/tree/main/pico_corpus_brat_annotated_files
bc5cdr_nerfair_processed
Benchmark dataset bc5cdr
This dataset was generated by the Data preprocessing step of the NERFAIR workflow (More information: https://github.com/YasCoMa/ner-fair-workflow )
Original dataset:
Li, Jiao, Yueping Sun, Robin J. Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J. Mattingly, Thomas C. Wiegers, and Zhiyong Lu. 2016. “BioCreative V CDR Task Corpus: A Resource for Chemical Disease Relation Extraction.” Database: The Journal of Biological… See the full description on the dataset page: https://huggingface.co/datasets/yasmmin/bc5cdr_nerfair_processed.merged_train_nerfair_processed
Benchmark Merged dataset - Train
This dataset was generated by the Data preprocessing step of the NERFAIR workflow (More information: https://github.com/YasCoMa/ner-fair-workflow )
It contains the train data of the following datasets:
ncbi: Doğan, Rezarta Islamaj, Robert Leaman, and Zhiyong Lu. 2014. “NCBI Disease Corpus: A Resource for Disease Name Recognition and Concept Normalization.” Journal of Biomedical Informatics 47 (February): 1–10.
bc5cdr: Li, Jiao, Yueping Sun, Robin… See the full description on the dataset page: https://huggingface.co/datasets/yasmmin/merged_train_nerfair_processed.ncbi_nerfair_processed
Benchmark dataset NCBI
This dataset was generated by the Data preprocessing step of the NERFAIR workflow (More information: https://github.com/YasCoMa/ner-fair-workflow )
Original dataset: Doğan, Rezarta Islamaj, Robert Leaman, and Zhiyong Lu. 2014. “NCBI Disease Corpus: A Resource for Disease Name Recognition and Concept Normalization.” Journal of Biomedical Informatics 47 (February): 1–10.
NER_FILTERED_DATASETchiads_nerfair_processed
Benchmark dataset CHIA
This dataset was generated by the Data preprocessing step of the NERFAIR workflow (More information: https://github.com/YasCoMa/ner-fair-workflow )
Original dataset: Kury, Fabrício, Alex Butler, Chi Yuan, Li-Heng Fu, Yingcheng Sun, Hao Liu, Ida Sim, Simona Carini, and Chunhua Weng. 2020. “Chia, a Large Annotated Corpus of Clinical Trial Eligibility Criteria.” Scientific Data 7 (1): 281.
OCR-data-900k
