datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NyayaAnumana-Classification-Datacancer_data_classificationlong-covid-classification-data
Data Description
Long-COVID related articles have been manually collected by information specialists.Please find further information here.
Size
Training
Development
Test
Total
Positive Examples
215
76
70
345
Negative Examples
199
62
68
345
Total
414
238
138
690
Citation
@article{10.1093/database/baac048,author = {Langnickel, Lisa and Darms, Johannes and Heldt, Katharina and Ducks, Denise and Fluck, Juliane},title = "{Continuous development… See the full description on the dataset page: https://huggingface.co/datasets/llangnickel/long-covid-classification-data.headline_dataurdu-binary-classification-dataThis Urdu sentiment dataset was formed by concatenating the following two datasets:
https://github.com/MuhammadYaseenKhan/Urdu-Sentiment-Corpus
https://www.kaggle.com/datasets/akkefa/imdb-dataset-of-50k-movie-translated-urdu-reviews
traceix-synthetic-classification-data
Traceix Synthetic Generated Data
Traceix Synthetic Generated Data is a synthetic tabular dataset for cybersecurity machine learning research, focused on static Windows PE-file metadata and binary classification workflows.
The dataset contains synthetically generated feature rows designed to resemble metadata patterns commonly extracted from Windows executable files during static analysis. It is intended for experimentation with malware/safe classification, anomaly detection, model… See the full description on the dataset page: https://huggingface.co/datasets/PerkinsFund/traceix-synthetic-classification-data.train_data_nbc_skewedsri_lankan_classifieds_data_set_ad_classificationThe Sri Lanka Classified Ads Dataset is a collection of classified advertisement listings sourced from various online marketplaces operating in Sri Lanka. It contains over 90,000 ads categorized under major sectors such as Property, Vehicles, and Electronics, and includes detailed titles and descriptions for each ad. This dataset is ideal for research and development in natural language processing (NLP) tasks like text classification, information extraction, and recommendation systems specific… See the full description on the dataset page: https://huggingface.co/datasets/Damika-7/sri_lankan_classifieds_data_set_ad_classification.Mental-Health-Classification-Datakhmer-classification-dataautonlp-data-iab_classificationautotrain-data-classificationdata
