datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task679_hope_edi_english_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task679_hope_edi_english_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task679_hope_edi_english_text_classification.multi-label-class-github-issues-text-classification
Dataset Card for "multi-label-class-github-issues-text-classification"
More Information needed
end2end_textclassification
Dataset Card for end2end_textclassification
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/argilla/end2end_textclassification.short-text-labeled-emotion-classificationuz-text-classification
Dataset Card for "uzbek_news"
Dataset Summary
Multi-label text classification dataset for Uzbek language and some sourcode for analysis. This repository contains the code and dataset used for text classification analysis for the Uzbek language. The dataset consists text data from 9 Uzbek news websites and press portals that included news articles and press releases. These websites were selected to cover various categories such as politics, sports, entertainment… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uz-text-classification.Hafez-text-classification-datasetstackoverflow-unified-text-open-status-classification
Dataset Card for "stackoverflow-unified-text-open-status-classification"
More Information needed
text_classificationfinancial-text-combo-classification
Dataset Card for "financial-text-combo-classification"
More Information needed
enrun-emails-text-classificationsynthetic-domain-text-classification
Dataset Card for my-distiset-b845cf19
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/my-distiset-b845cf19/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-domain-text-classification.synthetic-text-classification-news
Dataset Card for synthetic-text-classification-news
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/argilla/synthetic-text-classification-news/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-text-classification-news.task1645_medical_question_pair_dataset_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1645_medical_question_pair_dataset_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1645_medical_question_pair_dataset_text_classification.synthetic-text-classification-news-multi-label
Dataset Card for synthetic-text-classification-news-multi-label
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/synthetic-text-classification-news-multi-label/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-text-classification-news-multi-label.end2end_textclassification_with_suggestions_and_responses
Dataset Card for end2end_textclassification_with_suggestions_and_responses
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure… See the full description on the dataset page: https://huggingface.co/datasets/argilla/end2end_textclassification_with_suggestions_and_responses.task1605_ethos_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1605_ethos_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1605_ethos_text_classification.text-classification-news-topics
Dataset Card for test
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/test/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/text-classification-news-topics.bert-text-classificationtask1341_msr_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1341_msr_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1341_msr_text_classification.task1607_ethos_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1607_ethos_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1607_ethos_text_classification.task1606_ethos_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1606_ethos_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1606_ethos_text_classification.plantbert_text_classification_dataset
Dataset Card for "plantbert_text_classification_dataset"
More Information needed
flan_combined_task679_hope_edi_english_text_classificationTurkish-Synthetic-Text-Classification
Sentetik Metin Sınıflandırma Veri Seti (tr)
Bu veri seti, OpenRouter üzerinden "x-ai/grok-4-fast:free" modeli kullanılarak otomatik üretilmiş sentetik metin sınıflandırma örneklerini içerir.
Sınıflar: olumlu, olumsuz, nötr
Örnek sayısı: 1650
Train: 1350 örnek
Test: 300 örnek
Farklı tarzlarda üretilmiş sentetik metinler (kişisel deneyim, soru cümleleri, argo ifadeler, vs.)
Not: Metinler tamamen sentetiktir; gerçek kişi/kurum adları kullanılmamaya çalışılmıştır.
original-copy-classification-text
Dataset Card for "original-copy-classification-text"
More Information needed
task682_online_privacy_policy_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task682_online_privacy_policy_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task682_online_privacy_policy_text_classification.for_text_classification_food_not_foodfor_text_classification_food_not_food_v2
food / not_food (v2.0)
Двоичная классификация русскоязычного текста: относится ли он к съедобным для человека
продуктам или блюдам.
Строки выглядят как в v1 — {"caption": ..., "label": "food"}. label строковый,
чтобы работал прежний пайплайн; фиксированные id — в label_id (not_food=0,
food=1), границы сплита — в колонке split.
Схема на диске (*.jsonl, food_not_food_v2.parquet) и на HuggingFace
одинаковая: label везде int, ClassLabel c именами not_food / food.
Правило… See the full description on the dataset page: https://huggingface.co/datasets/alexneo68/for_text_classification_food_not_food_v2.stackoverflow-unified-text-open-status-classification-sample
Dataset Card for "stackoverflow-open-status-classification"
More Information needed
text-classification-subject
Dataset Card for "text-classification-subject"
More Information needed
