datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uet_iai_nlp_data_for_llmsData sources come from the following categories:
1.Web crawler dataset:
Website UET (ĐH Công nghệ): tuyensinh.uet.vnu.edu.vn; new.uet.vnu.edu.vn
Website HUS (ĐH KHTN): hus.vnu.edu.vn
Website EUB (ĐH Kinh tế): ueb.vnu.edu.vn
Website IS (ĐH Quốc tế): is.vnu.edu.vn
Website Eduacation (ĐH Giáo dục): education.vnu.edu.vn
Website NXB ĐHQG: press.vnu.edu.vnList domain web crawler
CC100:link to CC100 vi
Vietnews: link to bk vietnews dataset
C4_vi: link to C4_vi
Folder Toxic store files demo… See the full description on the dataset page: https://huggingface.co/datasets/group2sealion/uet_iai_nlp_data_for_llms.benchmark-llms-landuse-relevance
Land-use relevance benchmark
v3-multilingual · 85 languages x 300 items/language ·
25,500 items · binary yes/no labels.
Code
Package version recorded in run metadata: 0.2.0 (some runs lack version metadata).
Task and prompt
Does a sentence describe a place's land or environment in ways visible to satellites?
English prompt · greedy decoding · seed 0 · max_new_tokens=4096 ·
bfloat16 · batch varies by model.
unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.llm-srbench
LLM-SRBench: Benchmark for Scientific Equation Discovery with LLMs
We introduce LLM-SRBench, a comprehensive benchmark with 239 challenging problems across four scientific domains specifically designed to evaluate LLM-based scientific equation discovery methods while preventing trivial memorization.
Our benchmark comprises two main categories: LSR-Transform, which transforms common physical models into less common mathematical representations to test reasoning beyond memorization… See the full description on the dataset page: https://huggingface.co/datasets/nnheui/llm-srbench.LLMsGeneratedCodehiring-bias-mitigation-responses
Hiring-bias mitigation — model responses
Every response produced in the mitigation study of LLM hiring decisions: 64 runs,
2,782,350 responses, from 5 open-weight models in English and Ukrainian, at
baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset.
All released artifacts: the Hiring Bias Mitigation collection.
Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data.
Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.llm_speedrun
LLM Speedrun token streams
Pre-tokenized training artifacts for the LLM speedrun exercises.
File
Description
Tokens
tokenizer_50M.bpe
JSON-serialized BPE tokenizer
—
fineweb-edu-10BT.shuffle.bin
Shuffled FineWeb-Edu sample/10BT token stream
9,440,023,113
smoltalk.shuffle.bin
Shuffled SmolTalk data/all token stream
875,269,408
The .bin files are headerless, little-endian unsigned 16-bit token IDs and can be memory-mapped with NumPy:
from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/zkolter/llm_speedrun.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.dec1-jme-student-sourcing-training-and-llmsscaling-data-constrained-llms
Scaling Data-Constrained Language Models with Synthetic Data
This repository provides the pre-training corpora used in Scaling Data-Constrained Language Models with Synthetic Data (Findings of EACL 2026).
Overview
This repository contains multiple corpora designed to study data augmentation strategies for pre-training Japanese LLMs under a data-constrained data setting.
Starting from a limited Japanese Web corpus and a larger English Web corpus, we construct three… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/scaling-data-constrained-llms.llm-srbench
LLM-SRBench: Benchmark for Scientific Equation Discovery with LLMs
This dataset contains LLM-SRBench, a comprehensive benchmark for evaluating Large Language Models (LLMs) on scientific equation discovery (symbolic regression) tasks.
Paper: LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models (ICML 2025 Oral)
Original Repository: deep-symbolic-mathematics/llm-srbench
Original Dataset: nnheui/llm-srbench
📊 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/pkuHaowei/llm-srbench.DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
llm_sam_audio_datamcae-llms-benchmark-reportor-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.hiring-bias-mitigation-synthetic-data
Hiring-bias mitigation — synthetic training data
Semi-synthetic data for training LLMs to make hiring decisions that do not depend on a
protected attribute (military status, gender, religion), in English and Ukrainian.
Real inputs, synthetic labels. CVs and job descriptions are real, anonymised postings
from the Djinni Recruitment Dataset (MIT). Decisions and rationales were written by the
teacher model Qwen/Qwen3.5-122B-A10B-GPTQ-Int4.
Code and results:… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-synthetic-data.llms-txt
Context & Motivation
https://llmstxt.org/ is a project from Answer.AI which proposes to "standardise on using an /llms.txt file to provide information to help LLMs use a website at inference time."
I've noticed many tool providers begin to offer /llms.txt files for their websites and documentation. This includes developer tools and platforms like Perplexity, Anthropic, Hugging Face, Vercel, and others.
I've also come across https://directory.llmstxt.cloud/, a directory of… See the full description on the dataset page: https://huggingface.co/datasets/megrisdal/llms-txt.juris-tcuJurisTCU is a Brazilian Portuguese legal IR resource built from the curated jurisprudence collection of the Brazilian Federal Court of Accounts (TCU), from which we derive the benchmark subset used here. In its original form, the dataset contains 16,045 jurisprudence documents organized into more than 20 fields (metadata and textual fields). The most relevant are ENUNCIADO and EXCERTO, which correspond, respectively, to a summary of the ruling and to the excerpt from the decision that supports… See the full description on the dataset page: https://huggingface.co/datasets/ufca-llms/juris-tcu.LLMscore-ICLR-OpenReview
LLMscore-ICLR-OpenReview
This dataset is the released original dataset for the paper Position: Peer
Review Should Be Calibrated via LLM Scoring by Zijin Chen, Lesui Yu, Xiaofei
Liao, Hai Jin, and Qinbin Li. The paper has been accepted to the ICML 2026
Position Track.
Its concrete purpose is peer review analysis: the dataset is meant for
studying how paper-review rationales, numeric ratings, LLM-derived anchor
scores, and review-score residuals interact in scientific peer… See the full description on the dataset page: https://huggingface.co/datasets/Wutaghost/LLMscore-ICLR-OpenReview.BR-TaxQAThis dataset corresponds to BR-TaxQA-R, a collection derived from materials of the Brazilian Federal Revenue Service (Receita Federal) on personal income tax (IRPF). It contains 715 questions with their corresponding answers and introduces user-oriented, FAQ-style query formulations in a legal-tax domain.
Some questions are explicitly linked to other related questions. In our relevance design, the immediate answer to the queried question is treated as the primary positive with score = 2, while… See the full description on the dataset page: https://huggingface.co/datasets/ufca-llms/BR-TaxQA.normas-tcuNormasTCU is a Brazilian Portuguese legal IR test collection composed of normative documents from the Brazilian Federal Court of Accounts (TCU). These normative acts may have internal effects (e.g., rules governing internal procedures) or external effects (e.g., rules regulating how the court interacts with other public institutions) and differ from jurisprudential documents in both purpose and structure. Jurisprudential documents typically describe specific cases and present the legal… See the full description on the dataset page: https://huggingface.co/datasets/ufca-llms/normas-tcu.multimodal-LLMs-See-Sentiment
MLLMsent — datasets and experiment results
Every input and every output of "Multimodal LLMs See Sentiment"
(arXiv:2508.16873): the image descriptions generated by six multimodal
LLMs, the sentiment labels derived from the PerceptSent annotations, and the complete
per-fold results of all 141 experiments.
Paper: arXiv:2508.16873
Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment
Model checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/neemiasbsilva/multimodal-LLMs-See-Sentiment.juaJUÁ-Juris is centered on jurisprudence drawn from the curated jurisprudence collection of the Brazilian Federal Court of Accounts (TCU). In this collection, each instance contains an enunciado and an excerto: the enunciado is an abstractive summary of the ruling, while the excerto is the passage from the ruling that supports that summary. In our retrieval setup, the enunciado serves as the query, and the corresponding excerto is treated as the ground-truth positive passage. Within the… See the full description on the dataset page: https://huggingface.co/datasets/ufca-llms/jua.Reasoning-Boosts-Opinion-Alignment-in-LLMs
Reasoning Boosts Opinion Alignment in LLMs
Download and load with DataDict.load_from_disk.
Dataset fields
id: Identifier for the respondent / party / candidate (see lists below)
question: Identifier for the question
question_text: The actual question
answer: The answer to the respondent / party / candidate gave to the question. A = Yes, B = No, C = Neutral.
answer_comment: The argument / comment used for SFT.
political_position: The ideological group / party for this… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/Reasoning-Boosts-Opinion-Alignment-in-LLMs.llms.txt
llms.txt files extracted from the Common Crawl corpus
This dataset contains llms.txt files extracted from the Common Crawl corpus.
Specifically, we extracted response records matching */llms.txt and */llms-full.txt with status 200 and MIME-type text/plain or text/markdown.
What is llms.txt?
The /llms.txt file: A proposal to standardise on using an /llms.txt file to provide information to help LLMs use a website at inference time.
Our own analysis of this dataset is available… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/llms.txt.Global-LLMs-Replies
Global LLMs Replies
GPT-4o
-> 74,644 rows
mixtral-8x22b
-> 13,129 rows
claude-3-haiku
-> 3,871 rows
llm-security-leaderboard-contentsllmsplit_deepseekllm-serving-bench-runs
llm-serving-bench runs
Heavy outputs of the GPU sessions of https://github.com/hoangphu7122002/llm-serving-bench (answers, canary,
server logs, per-level parquets, quality-pack raw files). Paths mirror results/ in the repo; the git repo
keeps the reports and an HF.txt pointer. Pull one session:
uv run python scripts/data_hub.py pull-run --session runs/<name> --dest results.
Sessions
runs/20261009-1254_session4_quant-cache · 2026-10-09 · 15a370143e68… See the full description on the dataset page: https://huggingface.co/datasets/hoangphu7122002ai/llm-serving-bench-runs.llm-similarity-risk
