datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-reward-bench
AgentRewardBench
💾Code
📄Paper
🌐Website
🤗Dataset
💻Demo
🏆Leaderboard
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor
Loading dataset
You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.PUPAThis dataset contains the data presented in the paper PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles.
Code: https://github.com/siyan-sylvia-li/PAPILLON
tatoeba_mt
Dataset Card for The Tatoeba Translation Challenge
Please note that this dataset is intended strictly for evaluation and benchmarking purposes. Training models on this dataset, or including it in automatically collected web-scale training corpora, may lead to benchmark contamination and invalidate evaluation results.
Dataset Summary
The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user-contributed… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/tatoeba_mt.MultiJail
Multilingual Jailbreak Challenges in Large Language Models
This repo contains the data for our paper "Multilingual Jailbreak Challenges in Large Language Models".
[Github repo]
Annotation Statistics
We collected a total of 315 English unsafe prompts and annotated them into nine non-English languages. The languages were categorized based on resource availability, as shown below:
High-resource languages: Chinese (zh), Italian (it), Vietnamese (vi)
Medium-resource languages:… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/MultiJail.arxiv_nlp
arXiv Abstracts
Abstracts for the cs.CL category of ArXiv between 1991 and 2024. This dataset was created as an instructional tool for the Clustering and Topic Modeling chapter in the upcoming
"Hands-On Large Language Models" book.
The original dataset was retrieved here.
This subset will be updated towards the release of the book to make sure it captures relatively recent articles in the domain.
paradetox
ParaDetox: Text Detoxification with Parallel Data (English)
This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.lar-echr
Dataset Card for LAR-ECHR
Dataset Details
Dataset Description
Curated by: Odysseas S. Chlapanis
Funded by: Archimedes Research Unit
Language (NLP): English
License:
CC BY-NC-SA (Creative Commons / Attribution-NonCommercial-ShareAlike)
Read more: https://creativecommons.org/licenses/by-nc-sa/4.0/
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Uses… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/lar-echr.neural-news-benchmark
AI-generated News Detection Benchmark
neural-news is a benchmark dataset designed for human/AI news authorship classification in English, Turkish, Hungarian, and Persian.
Presented in Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian @ NLP for Positive Impact Workshop @ EMNLP2024.
Dataset Details
The dataset includes equal parts human-written and AI-generated news articles, raw and pre-processed.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/neural-news-benchmark.D_persuade_2
Persuade_2
The PERSUADE 2.0 corpus (Persuasive Essays for Rating, Selecting, and
Understanding Argumentative and Discourse Elements) contains over 25,000
argumentative essays written by 6th–12th grade students in the United States,
covering 15 distinct prompts across two writing tasks: independent and
source-based writing. The corpus also provides detailed individual and
demographic information for each writer.
This is the train, test, and validation split of the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_persuade_2.ik-nlp-22_winemagnlp_twitter_analysisNyayaAnumana-Classification-Datamc4_fi_cleaned
Dataset Card for mC4 Finnish Cleaned
Dataset Summary
mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split.
Supported Tasks and Leaderboards
mC4 Finnish is mainly intended to pretrain Finnish language models and word representations.
Languages
Finnish
Dataset Structure
Data Instances
[Needs More Information]
Data Fields
The data have several fields:
url: url of the source as a string
text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.ImplicitHate
Implicit Hate Speech
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech
[Read the Paper] | [Take a Survey to Access the Data] | [Download the Data]
Why Implicit Hate?
It is important to consider the subtle tricks that many extremists use to mask their threats and abuse. These more implicit forms of hate speech may easily go undetected by keyword detection systems, and even the most advanced architectures can fail if they have not been trained on… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/ImplicitHate.statcan-dialogue-dataset-retrieval
Statcan Dialogue Dataset (Processed for Retrieval Tasks)
This is a variant of the Statcan Dialogue Dataset, which we processed specifically for multilingual retrieval (english, french). It contains everything in CSVs, rather than having metadata hosted separately.
Quickstart
from datasets import load_dataset
repo = 'McGill-NLP/statcan-dialogue-dataset-retrieval'
# load english queries, training split
queries_en = load_dataset(repo, 'queries_english', split='train') #… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/statcan-dialogue-dataset-retrieval.D_ASAP-AES
D_ASAP-AES
This is the train, test, and validation split of the ASAP Automated Essay Scoring dataset,
prepared for use with the S-GRADES benchmark.
Ground truth labels have been removed to prevent leakage during evaluation.
For the original dataset with labels, see below.
Original Dataset
🔗 ASAP-AES on Kaggle
Citation
If you use this dataset, please cite the original:
@misc{asap_aes,
title={ASAP Automated Essay Scoring}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-AES.SOLD
SOLD - A Benchmark for Sinhala Offensive Language Identification
In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SOLD.FAERS-NLP
FAERS-NLP
Version: 1.0Author: sixuexing
GitHub: FAERS-NLP Repository
Dataset Summary
FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction.
Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks.
Dataset Structure
Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/sixuexing/FAERS-NLP.EcomBench
EcomBench: Where Intelligent Agents Conquer Commerce Realms
🚀 Benchmark Overview
EcomBench is a domain-specific, real-world evaluation framework designed to rigorously assess the capabilities of AI agents in delivering practical support for the complex, ever-evolving demands of e-commerce.
We believe that truly capable AI agents will fundamentally transform how we interact with commerce. E-commerce represents one of the world's most significant economic sectors, with… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/EcomBench.AmaWar
Amawal Warayni - ⴰⵎⴰⵡⴰⵍ ⴰⵙⵏⵎⴰⵍⴰⵢ ⵏ ⵉⵏⵓⵎⴰⴽ ⵏ ⵡⴰⵙⵙⴰⵖⵏ ⴷ ⵉⵎⵢⴰⴳⵏ ⵏ ⵜⵎⴰⵣⵉⵖⵜ ⵏ ⴰⵢⵜ ⵡⴰⵔⴰⵢⵏ
Bitext scraped from the online AmaWar dictionary of the Tamazight dialect of Ait Warain spoken in northeastern Morocco.
Contains sentences, stories, and poems in Tamazight written in the Neo-Tifinagh script along with their translations into Modern Standard Arabic.
The dataset is split into the following subsets:
examples: Example parallel sentences taken from dictionary entries.
idioms: Idiomatic… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/AmaWar.CultureBanktr_rteru_paradetox
ParaDetox: Text Detoxification with Parallel Data (Russian)
This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit
[2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.fakerecogna2-abstrativa
FakeRecogna 2.0 - Abstractive
FakeRecogna 2.0 presents the extension for the FakeRecogna dataset in the context of fake news detection. FakeRecogna includes real and fake news texts collected from online media and ten fact-checking sources in Brazil. An important aspect is the lack of relation between the real and fake news samples, i.e., they are not mutually related to each other to avoid intrinsic bias in the data.
The Dataset
The fake news collection was performed on… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/fakerecogna2-abstrativa.HalluciGen
Dataset Card for HalluciGen-Detection
Dataset Summary
This is a dataset for hallucination detection in the paraphrase generation and machine translation scenario. Each example in the dataset consists of a source sentence, a correct hypothesis, and an incorrect hypothesis containing an intrinsic hallucination. A hypothesis is considered to be a hallucination if it is not entailed by the "source" either by containing additional or contradictory information with respect to… See the full description on the dataset page: https://huggingface.co/datasets/NLP-RISE/HalluciGen.cognitive-biases-in-llms
A Comprehensive Evaluation of Cognitive Biases in LLMs: Dataset
Dataset for evaluating cognitive biases in large language models
1. Dataset Card Overview
A tabular dataset for measuring the presence and strength of cognitive biases in large language models (LLMs), introduced in the paper “A Comprehensive Evaluation of Cognitive Biases in LLMs” by Malberg et al.
Paper | Code
This dataset is intended only for the evaluation of LLMs and not to be used for… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/cognitive-biases-in-llms.FakeRecogna
FakeRecogna
FakeRecogna is a dataset comprised of real and fake news. The real news is not directly linked to fake news and vice-versa, which could lead to a biased classification. The news collection was performed by crawlers developed for mining pages of well-known and of great national importance agency news. The web crawlers were developed based on each analyzed webpage, where the extracted information is first separated into categories and then grouped by dates. The plurality… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/FakeRecogna.FAERS-NLP
FAERS-NLP
Version: 1.0Author: sixuexing
GitHub: FAERS-NLP Repository
Dataset Summary
FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction.
Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks.
Dataset Structure
Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/SanaeLaRose/FAERS-NLP.llama_jp
chat
jmultiwoz (chat-pred)
length: 3.54k
real-persona-chat (chat-pred)
length: 13.58k
multiple
Bactrian-X
length: 67k
databricks-dolly-15k-ja
length: 15k
guanaco_ja
length: 100.63k
llm-japanese-dataset-vanilla
length: 2.52M
OpenOrcaJapanese
length: 573.62k
qa
AutoGeneratedJapaneseQA (open-qa)
length: 93k
JAQKET (closed-qa)
length: 13.33k
JaQuAD (closed-qa)
length: 35.69k
smr
dialogsum-ja (chat-smr)
length: 20.28k… See the full description on the dataset page: https://huggingface.co/datasets/wisenut-nlp-team/llama_jp.NLP-HC-A3
