datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hle-gpt-oss-120b-no-python-260222
hle-gpt-oss-120b-no-python-260222
Deep research agent evaluation on rl-rag/hle_text_only (test split).
Results
Metric
Value
pass@4
47.9%
avg@4
26.6%
Trajectory accuracy
26.6% (2292/8632)
Questions
2158
Trajectories
8632 (4 per question)
Avg tool calls
14.5
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.jeb-rag
JEB-Bench
Charging the Gate Rent: Measured-Energy Accounting for Adaptive Retrieval-Augmented Generation
⚠️ Status: under construction. Phase 0 (measurement validation) and Phase 1
(index construction) are landing now. The oracle matrix (bench/oracle/) is
populated in Phase 2 and this card will be revised when it is complete. Do not
cite numbers from this repository until the status line says complete.
What this is
The first public per-query × per-configuration… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/jeb-rag.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.arxiv-markdown
arXiv Markdown
Last updated: 22 September 2026
Full text of 1,730,536 arXiv papers as Markdown, one row per paper with its
metadata. The subset covers 58 categories in computer science, electrical engineering,
mathematics, statistics, physics, condensed matter, quantum physics, astrophysics
instrumentation and quantitative biology, from 1991 to September 2026.
Papers with LaTeX source were converted with latexml and pandoc, so formulas are LaTeX in
$…$ and $$…$$. Papers that… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/arxiv-markdown.ragbench
RAGBench
Dataset Overview
RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples.
It covers five unique industry-specific domains and various RAG task types.
RAGBench examples are sourced from industry corpora such as user manuals, making it particularly relevant for industry applications.
RAGBench comrises 12 sub-component datasets, each one split into train/validation/test splits
Usage
from datasets import load_dataset
# load… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/ragbench.RAGTruth-processed
RAGTruth Dataset
Dataset Description
Dataset Summary
The RAGTruth dataset is designed for evaluating hallucinations in text generation models, particularly in retrieval-augmented generation (RAG) contexts. It contains examples of model outputs along with expert annotations indicating whether the outputs contain hallucinations.
Dataset Structure
Each example contains:
A query/question
Context passages
Model output
Hallucination labels (evident… See the full description on the dataset page: https://huggingface.co/datasets/wandb/RAGTruth-processed.multihop_qabrowsecomp-gpt-oss-120b-260222
browsecomp-gpt-oss-120b-260222
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
46.8%
avg@4
23.9%
Trajectory accuracy
23.9% (1211/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
26.1
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-gpt-oss-120b-260222.rag-mini-wikipediaIn this huggingface discussion you can share what you used the dataset for.
Derives from https://www.kaggle.com/datasets/rtatman/questionanswer-dataset?resource=download we generated our own subset using generate.py.
browsecomp-no-scroll-gpt-oss-120b
browsecomp-no-scroll-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
46.0%
avg@4
22.9%
Trajectory accuracy
22.9% (1160/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
27.0
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-no-scroll-gpt-oss-120b.hle_text_onlybrowsecomp-high-effort-gpt-oss-120b
browsecomp-high-effort-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
44.1%
avg@4
22.9%
Trajectory accuracy
22.9% (1158/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
55.4
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-gpt-oss-120b.publikasi-rag-finetuning-datasetrag_multilingual_training_negatives
How this dataset was made
We trained on chunks sourced from the documents in MADLAD-400 dataset that had been evaluated to contain a higher amount of educational information according to a state-of-the-art LLM.
We took chunks of size 250 tokens, 500 tokens, and 1000 tokens randomly for each document.
We then used these chunks to generate questions and answers based on this text using a state-of-the-art LLM.
Finally, we selected negatives for each chunk using the similarity from the… See the full description on the dataset page: https://huggingface.co/datasets/lightblue/rag_multilingual_training_negatives.browsecomp-qwen35-35b-a3b-think
browsecomp-qwen35-35b-a3b-think
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
43.0%
avg@4
24.8%
Trajectory accuracy
24.8% (1258/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
41.1
Full conversations
❌
Model & Setup
Model
Qwen3.5-35B-A3B
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-qwen35-35b-a3b-think.rag-dataset-12000
Retrieval-Augmented Generation (RAG) Dataset 12000
Retrieval-Augmented Generation (RAG) Dataset 12000 is an English dataset designed for RAG-optimized models, built by Neural Bridge AI, and released under Apache license 2.0.
Dataset Description
Dataset Summary
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by allowing them to consult an external authoritative knowledge base before generating responses. This approach significantly… See the full description on the dataset page: https://huggingface.co/datasets/neural-bridge/rag-dataset-12000.browsecomp-high-effort-full-gpt-oss-120b
browsecomp-high-effort-full-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@1
20.9%
avg@1
20.9%
Trajectory accuracy
20.9% (264/1266)
Questions
1266
Trajectories
1266 (1 per question)
Avg tool calls
52.9
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-full-gpt-oss-120b.rag-corpus-v1
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/rag-corpus-v1 — Agentic-RAG corpus + per-organ FAISS indexes
Doctrine v10/v11. Embedding model: BAAI/bge-base-en-v1.5 (768-dim).
Built by the agentic-RAG SHIP directive (390_AGENTIC_RAG_FAISS_PER_SPACE).
Contents
corpus.jsonl — 762 chunks, each ~512 tokens with… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/rag-corpus-v1.RAGDOLL
The RAGDOLL E-Commerce Webpage Dataset
This repository contains the RAGDOLL (Retrieval-Augmented Generation Deceived Ordering via AdversariaL materiaLs) dataset as well as its LLM-automated collection pipeline.
The RAGDOLL dataset is from the paper Ranking Manipulation for Conversational Search Engines from Samuel Pfrommer, Yatong Bai, Tanmay Gautam, and Somayeh Sojoudi. For experiment code associated with this paper, please refer to this repository.
The dataset consists of 10… See the full description on the dataset page: https://huggingface.co/datasets/Bai-YT/RAGDOLL.RAGognize
RAGognize Dataset Card
Resource
Link
Code
Paper
Demo
In Retrieval-Augmented Generation (RAG), ensuring accuracy and reliability is essential. RAGognize is a dataset created to help researchers and developers study and improve how AI systems use retrieved information. By offering structured examples and natural token-level closed-domain hallucination annotations, it provides a resource for analyzing model behavior and developing methods that might help make RAG… See the full description on the dataset page: https://huggingface.co/datasets/F4biian/RAGognize.expert-rag-benchmarks
Expert RAG Benchmarks
A unified collection of four expert-level legal RAG benchmarks, exposed as six
named splits and three relational configurations: questions, documents, and
qrels.
The KCL split is named kcl_essay because Hugging Face split identifiers do not
permit hyphens; its source name remains kcl-essay.
Loading
from datasets import load_dataset
repo_id = "jinulee-v/expert-rag-benchmarks"
questions = load_dataset(repo_id, "questions", split="housing")… See the full description on the dataset page: https://huggingface.co/datasets/jinulee-v/expert-rag-benchmarks.rag-mini-bioasqSee here for an updated version without nans in text-corpus.
In this huggingface discussion you can share what you used the dataset for.
Derives from http://participants-area.bioasq.org/Tasks/11b/trainingDataset/ we generated our own subset using generate.py.
github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos)
hle_rlvr_no_promptlegal-rag-bench
Legal RAG Bench ⚖️
Legal RAG Bench by Isaacus is a reasoning-intensive benchmark for assessing the end-to-end, real-world performance of production-grade legal RAG systems.
Legal RAG Bench is composed of 4,876 passages sampled from the Judicial College of Victoria’s Criminal Charge Book alongside 100 complex, handwritten questions demanding expert-level knowledge of Victorian criminal law and procedure to be answered correctly.
Legal RAG Bench is the first open dataset for the… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/legal-rag-bench.raggedvietmed-rag-dataset
ViMedGraph
ViMedGraph is a Vietnamese medical GraphRAG resource released together with the paper:
Adaptive Directed Graph Retrieval for Graph-Based Retrieval-Augmented Generation
The dataset is designed to support research on:
Graph-based Retrieval-Augmented Generation (GraphRAG)
Medical Question Answering
Information Retrieval
Knowledge Graph Construction
Retrieval Evaluation
Low-Resource Language NLP
ViMedGraph provides a structured Vietnamese medical knowledge resource… See the full description on the dataset page: https://huggingface.co/datasets/vuduylinh150804/vietmed-rag-dataset.github-reposThe entire dump of GitHub repositories.
hle-gpt-oss-120b-with-python-260222
hle-gpt-oss-120b-with-python-260222
Deep research agent evaluation on unknown.
Results
Metric
Value
pass@4
39.5%
avg@4
17.5%
Trajectory accuracy
17.4% (1860/10660)
Questions
1350
Trajectories
10660 (4 per question)
Avg tool calls
0.0
Full conversations
❌
Model & Setup
Model
unknown
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domainsNone
Tool Usage
Tool
Calls
%… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-with-python-260222.browsecomp-oss-env-high-effort-gpt-oss-120b
browsecomp-oss-env-high-effort-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@1
19.4%
avg@1
19.4%
Trajectory accuracy
19.4% (245/1266)
Questions
1266
Trajectories
1266 (1 per question)
Avg tool calls
52.5
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-oss-env-high-effort-gpt-oss-120b.
