datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
majid-multitask-dataset
Majid Multi-task Dataset | مجموعه داده چندوظیفهای مجید
English Description 🇬🇧
This dataset provides Persian–English text for multi-task NLP: text classification and question answering, including subtasks such as sentiment analysis and toxicity detection.
Dataset Details
Curated by: Maid121232 (Majid)
Languages: Persian (fa), English (en)
License: Apache-2.0
Size: Small (< 10k samples)
Tasks: Text Classification, Question Answering
Dataset Structure
Files: train.csv… See the full description on the dataset page: https://huggingface.co/datasets/Maid121232/majid-multitask-dataset.COLING-2025-GENAI-MULTIMultiTaskTWONB1-ds
Dataset Card: Multi-Task German Text Classification Dataset
Dataset Description
This dataset contains German text samples labeled for three tasks:
Fake News Detection (is_fake)
Hate Speech Detection (is_hate_speech)
Toxicity Detection (is_toxic)
Each entry in the dataset has binary or missing labels for the respective tasks:
0: Negative
1: Positive
-1: Not labeled for the task
The dataset is useful for training and evaluating models on multi-task learning objectives… See the full description on the dataset page: https://huggingface.co/datasets/Shivangsinha/MultiTaskTWONB1-ds.esg-news-sentiment-multitaskMulti-task-Dataset-Sampleviveksingh2400_multitask-nlp
OpenMultiTask NLP Dataset
Mirror of the Kaggle dataset viveksingh2400/multitask-nlp by Vivek Singh, released under MIT. All credit goes to the original author; please cite and link the Kaggle page when using this data.
A synthetic, logically constrained dataset for multi-task NLP including sentimen
License
MIT License, Copyright (c) Vivek Singh. The MIT license notice applies to all files in this repository.
Original description (from Kaggle)
The… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/viveksingh2400_multitask-nlp.
