datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coda-llm-data
Coda LLM Project & Dataset Repository
This repository contains the full end-to-end dataset, fine-tuning scripts, evaluation suites, load testing harness, and proxy architecture for Coda LLM (Granite-4.2-8B Najdi Sales Agent).
Model Repository: mohameddalii/coda-llm
Dataset / Code Repository: mohameddalii/coda-llm-data
📁 Repository Structure
coda-llm-data/
├── data/
│ ├── raw/ # Raw generated multi-turn dialogues across domains
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/mohameddalii/coda-llm-data.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.swallow-math-v2
SwallowMath-v2
Resources
📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation.
🧮 What is it?
SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1.
Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.llm-election-data-2024
Data Release for Large-Scale, Longitudinal Survey of Large Language Models (LLMs) During the 2024 US Elections
Overview
This repository contains the questions asked of and responses given by LLMs during the 2024 US elections, collected for a longitudinal survey conducted from July 23, 2024 to November 12, 2024. The study is described in detail in the paper "Large-Scale, Longitudinal Study of Large Language Models During the 2024 US Election Season" by Sarah H. Cen, Andrew… See the full description on the dataset page: https://huggingface.co/datasets/sarahcen/llm-election-data-2024.llm_datasetsswallow-code-v2
SwallowCode-v2
Resources
📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning.
💻 What is it?
SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline.
However, it had two significant limitations:
(1) it was distributed under the Llama 3.3 Community License, and
(2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.llm-polysemy-outputs
Polysemy Outputs
Raw model generations for the paper "Where did the ambiguity go? Examining how
multimodal models interpret polysemous words."
Each polysemous word (e.g. bank, bolt, trunk) is presented with no
disambiguating context — the prompt is the bare word — and the model's chosen
sense is observed over many samples. The same word set is run in two modalities
(text-to-image and text generation) and scored by the same judges, so their sense
distributions are directly… See the full description on the dataset page: https://huggingface.co/datasets/addisonwu05/llm-polysemy-outputs.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.agent-llm-traces-v2
Exgentic Agent LLM Traces v2 — Agent Chat Only
OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it.
This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.Qwen-Terminal-ToolBench-Processed-Tokenized
Qwen Terminal ToolBench Processed Datasets
Qwen-family processed/template-applied and selected tokenized terminal datasets.
Contents
qwen35_2b_full_terminal_toolcall_processed_v1: 1011776 rows, 36.26 GiB, features: text
qwen35_2b_terminal_toolcall_fullconv_from_samecount_v1: 327383 rows, 25.11 GiB, features: text
qwen35_27b_hf_fsdp_1024: 15335 rows, 0.17 GiB, features: input_ids, attention_mask, labels
qwen35_27b_processed_2bdata: 23472 rows, 2.77 GiB, features:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Qwen-Terminal-ToolBench-Processed-Tokenized.Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.cissp-llmbench
CISSP-LLMBench
Open-LLM-Benchmark
Open-LLM-Benchmark
Dataset Description
The Open-LLM-Leaderboard tracks the performance of various large language models (LLMs) on open-style questions to reflect their true capability. The dataset includes pre-generated model answers and evaluations using an LLM-based evaluator.
License: CC-BY 4.0
Dataset Structure
An example of model response files looks as follows:
{
"question": "What is the main function of photosynthetic cells within a plant?"… See the full description on the dataset page: https://huggingface.co/datasets/Open-Style/Open-LLM-Benchmark.llm-jp-corpus-v4-ja_wiki
llm-jp-corpus-v4 — ja_wiki
Mirror of the ja/ja_wiki sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_wiki
Files: 6 × jsonl.gz (1.9 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY-SA 3.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki.EEBO
EEBO-TCP (Markdown)
Early English Books Online, Text Creation Partnership: hand-keyed transcriptions of books printed
in England, and English books printed abroad, 1473-1700. Sermons, pamphlets, laws, almanacs,
ballads, science, literature. Converted from TCP's XML to Markdown. 60,329 texts, about
1.5 billion words, 25,369 from TCP Phase I and 34,960 from Phase II.
TCP keyed each text twice and proofed it to a 99.995% accuracy target. Characters the keyers
could not read are… See the full description on the dataset page: https://huggingface.co/datasets/Vintage-LLM/EEBO.jabarti-llm-dataset
jabarti-llm-dataset
Cleaned, section-chunked training corpus for a small bilingual LLM
(Arabic + English), combining a curated Egyptian-history collection with
general Wikipedia coverage from
CohereLabs/wikipedia-2023-11-embed-multilingual-v3.
Every pretrain record is a contiguous span of 120-1500 characters with the
article title and section headings removed. Provenance is in ds_source.
Configs and Splits
Config
Split
Rows
Training phase
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset.gazzetta-ufficiale
Gazzetta Ufficiale 👩🏻⚖️⚖️🏛️📜🇮🇹
La Gazzetta Ufficiale della Repubblica Italiana, quale fonte ufficiale di conoscenza delle norme in vigore in Italia e strumento di diffusione, informazione e ufficializzazione di testi legislativi, atti pubblici e privati, è edita dall’Istituto Poligrafico e Zecca dello Stato e pubblicata in collaborazione con il Ministero della Giustizia, il quale provvede alla direzione e redazione della stessa. L'Istituto Poligrafico e Zecca dello Stato… See the full description on the dataset page: https://huggingface.co/datasets/mii-llm/gazzetta-ufficiale.agent-llm-traces
Multi-Benchmark LLM Agent Traces
A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization.
Collected by Exgentic - A platform for LLM observability and performance optimization.
Dataset Overview
This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces.swallow-math
SwallowMath
October 21, 2025: Newer versions are available: SwallowCode-v2 and SwallowMath-v2 have been released with improved rewriting pipelines.
Resources
🐙 GitHub: Explore the project repository, including pipeline code and prompts at rioyokotalab/swallow-code-math.
📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowCode, our companion dataset for code generation.
What is it?… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math.KoHRM-Text-1.4B-sft-lora-data
KoHRM-Text-1.4B SFT and LoRA Prepared Data
This dataset repo stores curated KoHRM SFT/LoRA subsets in the same tokenized
HRM-Text V1Dataset format used by training. It is intended for quick behavior
alignment experiments after KoHRM pretraining.
Model repo:
https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B
Code repo:
https://github.com/LLM-OS-Models/KoHRM-text
Format
Each folder is a prepared V1Dataset:
<dataset-name>/
metadata.json
tokenizer_info.json… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-sft-lora-data.llm-fv-security-targets
LLM-FV Security Targets
This dataset contains 889 independently validated, containerized security-agent targets produced by the ucsb-mlsec/llm-fv pipelines.
Contents
GitHub Global Security Advisories: 446 targets
OSS-Fuzz: 442 targets
PoC task support: 889 targets
Exploit task support: 447 targets
Patch task support: 447 targets
Compressed bundle size: 42.41 GiB
Vulnerability classes: {'logic_bug': 450, 'memory_vulnerability': 439}
Primary languages: {'C': 118… See the full description on the dataset page: https://huggingface.co/datasets/secmlr/llm-fv-security-targets.hiring-bias-mitigation-responses
Hiring-bias mitigation — model responses
Every response produced in the mitigation study of LLM hiring decisions: 64 runs,
2,782,350 responses, from 5 open-weight models in English and Ukrainian, at
baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset.
All released artifacts: the Hiring Bias Mitigation collection.
Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data.
Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.llm_speedrun
LLM Speedrun token streams
Pre-tokenized training artifacts for the LLM speedrun exercises.
File
Description
Tokens
tokenizer_50M.bpe
JSON-serialized BPE tokenizer
—
fineweb-edu-10BT.shuffle.bin
Shuffled FineWeb-Edu sample/10BT token stream
9,440,023,113
smoltalk.shuffle.bin
Shuffled SmolTalk data/all token stream
875,269,408
The .bin files are headerless, little-endian unsigned 16-bit token IDs and can be memory-mapped with NumPy:
from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/zkolter/llm_speedrun.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.swallow-code
SwallowCode
Notice
May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility.
May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.llmops-database
The ZenML LLMOps Database
To learn more about ZenML and our open-source MLOps framework, visit
zenml.io.
Dataset Summary
The LLMOps Database is a comprehensive collection of over 500 real-world
generative AI implementations that showcases how organizations are successfully
deploying Large Language Models (LLMs) in production. The case studies have been
carefully curated to focus on technical depth and practical problem-solving,
with an emphasis on implementation… See the full description on the dataset page: https://huggingface.co/datasets/zenml/llmops-database.spai-ss6-llm-1b-thai-corpus
Thai Medical And Health Corpus
Thai public medical and health web corpus collected for research and LLM dataset
experimentation, with optional imported Thai medical/health datasets from
Hugging Face stored as separate configs.
Public Web Corpus
Config: default
Split: train
Records: 3660 deduplicated articles
Columns: 16
Format: Parquet
Latest collection profile: free_1000
Latest generated at: 2026-06-06T17:41:38.787978+00:00
Source And Method
The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.scaling-data-constrained-llms
Scaling Data-Constrained Language Models with Synthetic Data
This repository provides the pre-training corpora used in Scaling Data-Constrained Language Models with Synthetic Data (Findings of EACL 2026).
Overview
This repository contains multiple corpora designed to study data augmentation strategies for pre-training Japanese LLMs under a data-constrained data setting.
Starting from a limited Japanese Web corpus and a larger English Web corpus, we construct three… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/scaling-data-constrained-llms.LLMinstruct
Dataset Card for lenamerkli/LLMinstruct
This dataset consists of instruct finetuning data from all of my projects.
Dataset Details
Dataset Sources
Repository: https://github.com/lenamerkli/LLMinstruct
Uses
This dataset is useful for instruct-tuning or fine-tuning large language models.
Use Recommendations
I recommend to use only the following data for training:
all data marked as containing no mistakes
the drawback… See the full description on the dataset page: https://huggingface.co/datasets/lenamerkli/LLMinstruct.
