datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.pypi-20241031crates-20250307npm-20241031rubygems-20241031npm-20240828crates-20240903benchmark-llms-landuse-relevance
Land-use relevance benchmark
v3-multilingual · 85 languages x 300 items/language ·
25,500 items · binary yes/no labels.
Code
Package version recorded in run metadata: 0.2.0 (some runs lack version metadata).
Task and prompt
Does a sentence describe a place's land or environment in ways visible to satellites?
English prompt · greedy decoding · seed 0 · max_new_tokens=4096 ·
bfloat16 · batch varies by model.
unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.llm_global_opinions
Dataset Card for GlobalOpinionQA
Dataset Summary
The data contains a subset of survey questions about global issues and opinions adapted from the World Values Survey and Pew Global Attitudes Survey.
The data is further described in the paper: Towards Measuring the Representation of Subjective Global Opinions in Language Models.
Purpose
In our paper, we use this dataset to analyze the opinions that large language models (LLMs) reflect on complex global… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/llm_global_opinions.Bitext-events-ticketing-llm-chatbot-training-dataset
Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.LLM-Artifacts
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶
Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo,
Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu
Dongyeop Kang
Minnesota NLP, University of Minnesota Twin Cities
† Project Lead,
¶ Core Contribution,
Arxiv
Project Page
📌 Table of Contents
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our leaderboard at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.llm-latency-tracker
LLM Latency Tracker
Independent, continuously measured latency and availability for AI inference API
providers, aggregated by day. Covers 46 providers across
4 regions (ap-tokyo, eu-hetzner, sa-east, us-central), built from
4,149,090 raw probes collected between 2026-07-23 and
2026-10-10.
Live rankings and full methodology: llmlatency.dev
How the numbers are produced
Probes run every five minutes from separate network locations and are never routed
through a… See the full description on the dataset page: https://huggingface.co/datasets/llmlatency/llm-latency-tracker.llm-cost-same-prompt
Measured per-call LLM cost — same prompt, every model
Vendors publish prices per million tokens. Nobody publishes what one call actually costs, because
that depends on how many tokens the model chooses to emit — and on the same question models differ by
more than an order of magnitude. One model finishes a JSON extraction in 23 tokens; another writes 300.
This dataset sends a fixed set of prompts to every model at temperature 0, every night, and records
the cost computed from… See the full description on the dataset page: https://huggingface.co/datasets/mario0369/llm-cost-same-prompt.Bitext-telco-llm-chatbot-training-dataset
Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.jgpqa
JGPQA
This repository provides GPQA dataset translated from English into Japanese by LLM-jp, a collaborative project launched in Japan.
The dataset was translated from English to Japanese using machine translation, then checked and corrected by external experts.
The links of the original GPQA dataset are here(HuggingFace).
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
Yuji Tamakoshi, Kouta Nakayama, Yusuke Miyao.
LLM-Ads
LLM-Ads — Sponsored-recommendation evaluation traces
Per-trial responses and labels from the experiments in
Just Ask for a Table: A Thirty-Token User Prompt Defeats Sponsored
Recommendations in Twelve LLMs
(arXiv:2605.12772).
The data set reproduces and extends the evaluation of Wu et al.\ 2026
(arXiv:2604.08525) on a twelve-model
pool (ten open-source chat models served through an OpenAI-compatible
API endpoint plus the two paper-overlap OpenAI models
gpt-3.5-turbo and gpt-4o).… See the full description on the dataset page: https://huggingface.co/datasets/akmaier/LLM-Ads.DEBATE_LLM
DEBATE Benchmark
This repository contains CSV files from the DEBATE project: large-scale
human conversation experiments organized around controversial and
opinion-based topics. The data consists of multi-round conversations
between human participants discussing political, social, and belief-related
topics, following the protocol described in:
Chuang, Y.-S., Tu, R., Dai, C., Vasani, S., Li, Y., Yao, B., Tessler, M. H., Yang, S., Shah, D., Hawkins, R., Hu, J., & Rogers, T. T. (2026).… See the full description on the dataset page: https://huggingface.co/datasets/seantw/DEBATE_LLM.Bitext-insurance-llm-chatbot-training-dataset
Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.LLMFineTuningBench
Dataset Card for LLMFineTuningBench
A dataset of over 30,000 LLM fine-tuning experiments, capturing detailed performance metrics from jobs run on high-performance computing (HPC) clusters. It spans a wide range of models, fine-tuning methods, and hardware configurations, and is intended to support research on predictive resource allocation, performance optimization, and cost estimation for LLM fine-tuning workloads.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/LLMFineTuningBench.local-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.linalg-bench-llm
LinAlg-Bench: Where LLMs Stop Computing and Start Hallucinating
Ten frontier LLMs drop from near-perfect to near-zero on 5×5 eigenvalue problems. Complete computational collapse is dimension-gated: rare at 3×3, dominant at 4×4 and 5×5. Failures dissociate cleanly by task — eigenvalues fail by constraint-aware fabrication (invented eigenvalues that still match the matrix trace), determinants by sign-accumulation drift. Nearly a third of irrational-spectrum eigenvalue failures are… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-llm.multimodal-LLMs-See-Sentiment
MLLMsent — datasets and experiment results
Every input and every output of "Multimodal LLMs See Sentiment"
(arXiv:2508.16873): the image descriptions generated by six multimodal
LLMs, the sentiment labels derived from the PerceptSent annotations, and the complete
per-fold results of all 141 experiments.
Paper: arXiv:2508.16873
Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment
Model checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/neemiasbsilva/multimodal-LLMs-See-Sentiment.llm-jp-longbench-NIILC
llm-jp-longbench-NIILC
llm-jp LongBench ベンチマークについて
このデータセットは,GitHub リポジトリhttps://github.com/llm-jp/llm-jp-longbenchで公開されているllm-jp LongBenchベンチマークの評価対象データセットの一部として構築されています。
llm-jp LongBench ベンチマークは,日本語大型言語モデル(LLM)のロングコンテキスト処理能力を体系的に評価することを目的としており,複数の長文コンテキスト QA データセットを含んでいます。
本データセットはその一つです。
データセット概要
本データセットは,日本語質問応答データセット NIILC
(Sekine, 2003)を基に,
回答が一意に定まり,かつ時間によって正解が変化しない質問のみを選別し,
それらに対応する Wikipedia 記事をコンテキストとして付与することで構築した,
ロングコンテキスト QA 評価用データセットです。… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-longbench-NIILC.llm-jp-longbench-JEMHop
llm-jp-longbench-JEMHopQA
llm-jp LongBench ベンチマークについて
このデータセットは,GitHub リポジトリhttps://github.com/llm-jp/llm-jp-longbenchで公開されているllm-jp LongBenchベンチマークの評価対象データセットの一部として構築されています。
llm-jp LongBench ベンチマークは,日本語大型言語モデル(LLM)のロングコンテキスト処理能力を体系的に評価することを目的としており,複数の長文コンテキスト QA データセットを含んでいます。
本データセットはその一つです。
データセット概要
本データセットは、日本語の説明可能マルチホップ質問応答データセットJEMHopQA
(Ishii et al., 2024)を基に、Wikipedia記事を付与することで構築したロングコンテキストQA評価用データセットです。
最大65… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-longbench-JEMHop.ivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.
