Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01McGill-NLP /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.imagerobotics1K<n<10K4 likes23k downloads1y agoHugging Face02Columbia-NLP /PUPAThis dataset contains the data presented in the paper PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles. Code: https://github.com/siyan-sylvia-li/PAPILLON texttext-generationn<1K3 likes7.7k downloads2y agoHugging Face03Helsinki-NLP /tatoeba_mtgated Dataset Card for The Tatoeba Translation Challenge Please note that this dataset is intended strictly for evaluation and benchmarking purposes. Training models on this dataset, or including it in automatically collected web-scale training corpora, may lead to benchmark contamination and invalidate evaluation results. Dataset Summary The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user-contributed… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/tatoeba_mt.texttext-generation10M<n<100M64 likes2.1k downloads2d agoHugging Face04DAMO-NLP-SG /MultiJail Multilingual Jailbreak Challenges in Large Language Models This repo contains the data for our paper "Multilingual Jailbreak Challenges in Large Language Models". [Github repo] Annotation Statistics We collected a total of 315 English unsafe prompts and annotated them into nine non-English languages. The languages were categorized based on resource availability, as shown below: High-resource languages: Chinese (zh), Italian (it), Vietnamese (vi) Medium-resource languages:… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/MultiJail.textn<1K13 likes1.3k downloads3y agoHugging Face05MaartenGr /arxiv_nlp arXiv Abstracts Abstracts for the cs.CL category of ArXiv between 1991 and 2024. This dataset was created as an instructional tool for the Clustering and Topic Modeling chapter in the upcoming "Hands-On Large Language Models" book. The original dataset was retrieved here. This subset will be updated towards the release of the book to make sure it captures relatively recent articles in the domain. text10K<n<100K12 likes1.1k downloads3y agoHugging Face06s-nlp /paradetox ParaDetox: Text Detoxification with Parallel Data (English) This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.imagetext-generation10K<n<100K9 likes970 downloads2y agoHugging Face07AUEB-NLP /lar-echr Dataset Card for LAR-ECHR Dataset Details Dataset Description Curated by: Odysseas S. Chlapanis Funded by: Archimedes Research Unit Language (NLP): English License: CC BY-NC-SA (Creative Commons / Attribution-NonCommercial-ShareAlike) Read more: https://creativecommons.org/licenses/by-nc-sa/4.0/ Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Uses… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/lar-echr.textquestion-answeringn<1K3 likes830 downloads1y agoHugging Face08tum-nlp /neural-news-benchmark AI-generated News Detection Benchmark neural-news is a benchmark dataset designed for human/AI news authorship classification in English, Turkish, Hungarian, and Persian. Presented in Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian @ NLP for Positive Impact Workshop @ EMNLP2024. Dataset Details The dataset includes equal parts human-written and AI-generated news articles, raw and pre-processed. Curated… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/neural-news-benchmark.texttext-classification10K<n<100K4 likes392 downloads2y agoHugging Face09nlpatunt /D_persuade_2 Persuade_2 The PERSUADE 2.0 corpus (Persuasive Essays for Rating, Selecting, and Understanding Argumentative and Discourse Elements) contains over 25,000 argumentative essays written by 6th–12th grade students in the United States, covering 15 distinct prompts across two writing tasks: independent and source-based writing. The corpus also provides detailed individual and demographic information for each writer. This is the train, test, and validation split of the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_persuade_2.tabular10K<n<100K0 likes366 downloads7mo agoHugging Face10GroNLP /ik-nlp-22_winemagtabular10K<n<100K6 likes345 downloads5y agoHugging Face11hamedhf /nlp_twitter_analysistexttext-classification1K<n<10K1 likes324 downloads3y agoHugging Face12L-NLProc /NyayaAnumana-Classification-Datagatedtext1M<n<10M1 likes287 downloads2y agoHugging Face13Finnish-NLP /mc4_fi_cleaned Dataset Card for mC4 Finnish Cleaned Dataset Summary mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split. Supported Tasks and Leaderboards mC4 Finnish is mainly intended to pretrain Finnish language models and word representations. Languages Finnish Dataset Structure Data Instances [Needs More Information] Data Fields The data have several fields: url: url of the source as a string text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.texttext-generation10M<n<100M4 likes280 downloads4y agoHugging Face14SALT-NLP /ImplicitHate Implicit Hate Speech Latent Hatred: A Benchmark for Understanding Implicit Hate Speech [Read the Paper] | [Take a Survey to Access the Data] | [Download the Data] Why Implicit Hate? It is important to consider the subtle tricks that many extremists use to mask their threats and abuse. These more implicit forms of hate speech may easily go undetected by keyword detection systems, and even the most advanced architectures can fail if they have not been trained on… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/ImplicitHate.text1K<n<10K9 likes267 downloads4y agoHugging Face15McGill-NLP /statcan-dialogue-dataset-retrieval Statcan Dialogue Dataset (Processed for Retrieval Tasks) This is a variant of the Statcan Dialogue Dataset, which we processed specifically for multilingual retrieval (english, french). It contains everything in CSVs, rather than having metadata hosted separately. Quickstart from datasets import load_dataset repo = 'McGill-NLP/statcan-dialogue-dataset-retrieval' # load english queries, training split queries_en = load_dataset(repo, 'queries_english', split='train') #… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/statcan-dialogue-dataset-retrieval.textquestion-answering10K<n<100K1 likes254 downloads2y agoHugging Face16nlpatunt /D_ASAP-AES D_ASAP-AES This is the train, test, and validation split of the ASAP Automated Essay Scoring dataset, prepared for use with the S-GRADES benchmark. Ground truth labels have been removed to prevent leakage during evaluation. For the original dataset with labels, see below. Original Dataset 🔗 ASAP-AES on Kaggle Citation If you use this dataset, please cite the original: @misc{asap_aes, title={ASAP Automated Essay Scoring}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-AES.tabular10K<n<100K0 likes241 downloads7mo agoHugging Face17sinhala-nlp /SOLD SOLD - A Benchmark for Sinhala Offensive Language Identification In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SOLD.texttext-classification10K<n<100K2 likes240 downloads3y agoHugging Face18sixuexing /FAERS-NLP FAERS-NLP Version: 1.0Author: sixuexing GitHub: FAERS-NLP Repository Dataset Summary FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction. Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks. Dataset Structure Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/sixuexing/FAERS-NLP.tabular1M<n<10M2 likes237 downloads1y agoHugging Face19Alibaba-NLP /EcomBench EcomBench: Where Intelligent Agents Conquer Commerce Realms 🚀 Benchmark Overview EcomBench is a domain-specific, real-world evaluation framework designed to rigorously assess the capabilities of AI agents in delivering practical support for the complex, ever-evolving demands of e-commerce. We believe that truly capable AI agents will fundamentally transform how we interact with commerce. E-commerce represents one of the world's most significant economic sectors, with… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/EcomBench.textn<1K8 likes200 downloads10mo agoHugging Face20Tamazight-NLP /AmaWar Amawal Warayni - ⴰⵎⴰⵡⴰⵍ ⴰⵙⵏⵎⴰⵍⴰⵢ ⵏ ⵉⵏⵓⵎⴰⴽ ⵏ ⵡⴰⵙⵙⴰⵖⵏ ⴷ ⵉⵎⵢⴰⴳⵏ ⵏ ⵜⵎⴰⵣⵉⵖⵜ ⵏ ⴰⵢⵜ ⵡⴰⵔⴰⵢⵏ Bitext scraped from the online AmaWar dictionary of the Tamazight dialect of Ait Warain spoken in northeastern Morocco. Contains sentences, stories, and poems in Tamazight written in the Neo-Tifinagh script along with their translations into Modern Standard Arabic. The dataset is split into the following subsets: examples: Example parallel sentences taken from dictionary entries. idioms: Idiomatic… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/AmaWar.texttranslation1K<n<10K2 likes197 downloads5mo agoHugging Face21SALT-NLP /CultureBanktext10K<n<100K20 likes193 downloads2y agoHugging Face22nlpyeditepe /tr_rtetabulartext-classification1K<n<10K0 likes183 downloads4y agoHugging Face23s-nlp /ru_paradetox ParaDetox: Text Detoxification with Parallel Data (Russian) This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit [2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.imagetext-generation10K<n<100K4 likes179 downloads2y agoHugging Face24recogna-nlp /fakerecogna2-abstrativa FakeRecogna 2.0 - Abstractive FakeRecogna 2.0 presents the extension for the FakeRecogna dataset in the context of fake news detection. FakeRecogna includes real and fake news texts collected from online media and ten fact-checking sources in Brazil. An important aspect is the lack of relation between the real and fake news samples, i.e., they are not mutually related to each other to avoid intrinsic bias in the data. The Dataset The fake news collection was performed on… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/fakerecogna2-abstrativa.tabulartext-classification10K<n<100K2 likes173 downloads1y agoHugging Face25NLP-RISE /HalluciGen Dataset Card for HalluciGen-Detection Dataset Summary This is a dataset for hallucination detection in the paraphrase generation and machine translation scenario. Each example in the dataset consists of a source sentence, a correct hypothesis, and an incorrect hypothesis containing an intrinsic hallucination. A hypothesis is considered to be a hallucination if it is not entailed by the "source" either by containing additional or contradictory information with respect to… See the full description on the dataset page: https://huggingface.co/datasets/NLP-RISE/HalluciGen.texttext-classificationn<1K0 likes169 downloads1y agoHugging Face26tum-nlp /cognitive-biases-in-llms A Comprehensive Evaluation of Cognitive Biases in LLMs: Dataset Dataset for evaluating cognitive biases in large language models 1. Dataset Card Overview A tabular dataset for measuring the presence and strength of cognitive biases in large language models (LLMs), introduced in the paper “A Comprehensive Evaluation of Cognitive Biases in LLMs” by Malberg et al. Paper | Code This dataset is intended only for the evaluation of LLMs and not to be used for… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/cognitive-biases-in-llms.text10K<n<100K1 likes161 downloads1y agoHugging Face27recogna-nlp /FakeRecogna FakeRecogna FakeRecogna is a dataset comprised of real and fake news. The real news is not directly linked to fake news and vice-versa, which could lead to a biased classification. The news collection was performed by crawlers developed for mining pages of well-known and of great national importance agency news. The web crawlers were developed based on each analyzed webpage, where the extracted information is first separated into categories and then grouped by dates. The plurality… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/FakeRecogna.texttext-classification10K<n<100K3 likes160 downloads3y agoHugging Face28SanaeLaRose /FAERS-NLP FAERS-NLP Version: 1.0Author: sixuexing GitHub: FAERS-NLP Repository Dataset Summary FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction. Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks. Dataset Structure Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/SanaeLaRose/FAERS-NLP.tabular1M<n<10M0 likes153 downloads9mo agoHugging Face29wisenut-nlp-team /llama_jp chat jmultiwoz (chat-pred) length: 3.54k real-persona-chat (chat-pred) length: 13.58k multiple Bactrian-X length: 67k databricks-dolly-15k-ja length: 15k guanaco_ja length: 100.63k llm-japanese-dataset-vanilla length: 2.52M OpenOrcaJapanese length: 573.62k qa AutoGeneratedJapaneseQA (open-qa) length: 93k JAQKET (closed-qa) length: 13.33k JaQuAD (closed-qa) length: 35.69k smr dialogsum-ja (chat-smr) length: 20.28k… See the full description on the dataset page: https://huggingface.co/datasets/wisenut-nlp-team/llama_jp.text1M<n<10M0 likes150 downloads2y agoHugging Face30damonsalvatore123 /NLP-HC-A3image10K<n<100K0 likes149 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.