datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hle-gpt-oss-120b-no-python-260222
hle-gpt-oss-120b-no-python-260222
Deep research agent evaluation on rl-rag/hle_text_only (test split).
Results
Metric
Value
pass@4
47.9%
avg@4
26.6%
Trajectory accuracy
26.6% (2292/8632)
Questions
2158
Trajectories
8632 (4 per question)
Avg tool calls
14.5
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.jeb-rag
JEB-Bench
Charging the Gate Rent: Measured-Energy Accounting for Adaptive Retrieval-Augmented Generation
⚠️ Status: under construction. Phase 0 (measurement validation) and Phase 1
(index construction) are landing now. The oracle matrix (bench/oracle/) is
populated in Phase 2 and this card will be revised when it is complete. Do not
cite numbers from this repository until the status line says complete.
What this is
The first public per-query × per-configuration… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/jeb-rag.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.arxiv-markdown
arXiv Markdown
Last updated: 22 September 2026
Full text of 1,730,536 arXiv papers as Markdown, one row per paper with its
metadata. The subset covers 58 categories in computer science, electrical engineering,
mathematics, statistics, physics, condensed matter, quantum physics, astrophysics
instrumentation and quantitative biology, from 1991 to September 2026.
Papers with LaTeX source were converted with latexml and pandoc, so formulas are LaTeX in
$…$ and $$…$$. Papers that… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/arxiv-markdown.ragbench
RAGBench
Dataset Overview
RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples.
It covers five unique industry-specific domains and various RAG task types.
RAGBench examples are sourced from industry corpora such as user manuals, making it particularly relevant for industry applications.
RAGBench comrises 12 sub-component datasets, each one split into train/validation/test splits
Usage
from datasets import load_dataset
# load… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/ragbench.browsecomp-gpt-oss-120b-260222
browsecomp-gpt-oss-120b-260222
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
46.8%
avg@4
23.9%
Trajectory accuracy
23.9% (1211/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
26.1
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-gpt-oss-120b-260222.browsecomp-no-scroll-gpt-oss-120b
browsecomp-no-scroll-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
46.0%
avg@4
22.9%
Trajectory accuracy
22.9% (1160/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
27.0
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-no-scroll-gpt-oss-120b.browsecomp-high-effort-gpt-oss-120b
browsecomp-high-effort-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
44.1%
avg@4
22.9%
Trajectory accuracy
22.9% (1158/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
55.4
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-gpt-oss-120b.rag_multilingual_training_negatives
How this dataset was made
We trained on chunks sourced from the documents in MADLAD-400 dataset that had been evaluated to contain a higher amount of educational information according to a state-of-the-art LLM.
We took chunks of size 250 tokens, 500 tokens, and 1000 tokens randomly for each document.
We then used these chunks to generate questions and answers based on this text using a state-of-the-art LLM.
Finally, we selected negatives for each chunk using the similarity from the… See the full description on the dataset page: https://huggingface.co/datasets/lightblue/rag_multilingual_training_negatives.browsecomp-qwen35-35b-a3b-think
browsecomp-qwen35-35b-a3b-think
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
43.0%
avg@4
24.8%
Trajectory accuracy
24.8% (1258/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
41.1
Full conversations
❌
Model & Setup
Model
Qwen3.5-35B-A3B
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-qwen35-35b-a3b-think.browsecomp-high-effort-full-gpt-oss-120b
browsecomp-high-effort-full-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@1
20.9%
avg@1
20.9%
Trajectory accuracy
20.9% (264/1266)
Questions
1266
Trajectories
1266 (1 per question)
Avg tool calls
52.9
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-full-gpt-oss-120b.rag-corpus-v1
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/rag-corpus-v1 — Agentic-RAG corpus + per-organ FAISS indexes
Doctrine v10/v11. Embedding model: BAAI/bge-base-en-v1.5 (768-dim).
Built by the agentic-RAG SHIP directive (390_AGENTIC_RAG_FAISS_PER_SPACE).
Contents
corpus.jsonl — 762 chunks, each ~512 tokens with… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/rag-corpus-v1.hle-gpt-oss-120b-with-python-260222
hle-gpt-oss-120b-with-python-260222
Deep research agent evaluation on unknown.
Results
Metric
Value
pass@4
39.5%
avg@4
17.5%
Trajectory accuracy
17.4% (1860/10660)
Questions
1350
Trajectories
10660 (4 per question)
Avg tool calls
0.0
Full conversations
❌
Model & Setup
Model
unknown
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domainsNone
Tool Usage
Tool
Calls
%… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-with-python-260222.browsecomp-oss-env-high-effort-gpt-oss-120b
browsecomp-oss-env-high-effort-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@1
19.4%
avg@1
19.4%
Trajectory accuracy
19.4% (245/1266)
Questions
1266
Trajectories
1266 (1 per question)
Avg tool calls
52.5
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-oss-env-high-effort-gpt-oss-120b.rag-bench
Dataset card for RAG-BENCH
Data Summary
RAG-bench aims to provide results of many commonly used RAG datasets. All the results in this dataset are evaluated by the RAG evaluation tool Rageval, which could be easily reproduced with the tool.
Currently, we have provided the results of ASQA dataset,ELI5 dataset and HotPotQA dataset.
Data Instance
ASQA
{
"ambiguous_question":"Who is the original artist of sound of silence?",
"qa_pairs":[{… See the full description on the dataset page: https://huggingface.co/datasets/golaxy/rag-bench.fineweb-10bt-weborganizer-sigir-dota-rag
FineWeb-10BT with WebOrganizer Labels for DoTA-RAG
Dataset · DoTA-RAG paper · Project page
This dataset is a labeled version of the FineWeb-10BT web corpus used as the document collection in DoTA-RAG: Dynamic of Thought Aggregation RAG. It contains 14,868,862 English-language documents in one train split. Each document retains its FineWeb text and provenance fields and adds a predicted topic and document format from WebOrganizer. The published Parquet files total 30.7 GB to… See the full description on the dataset page: https://huggingface.co/datasets/saksornr/fineweb-10bt-weborganizer-sigir-dota-rag.browsecomp-qwen35-35b-a3b-nothink
browsecomp-qwen35-35b-a3b-nothink
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
32.3%
avg@4
16.4%
Trajectory accuracy
16.4% (830/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
36.2
Full conversations
❌
Model & Setup
Model
Qwen3.5-35B-A3B
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-qwen35-35b-a3b-nothink.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.trec-ragtime-2026
TREC RAGTIME 2026 — sentence and passage renderings
A sentence-level view of the TREC RAGTIME 2026 news collection,
with two English machine translations of every non-English sentence and the passage boundaries used
for retrieval. Derived from trec-ragtime/ragtime2.
Pipeline, experiment design, run configurations and reproduction steps:
github.com/jknafou/trec-ragtime-2026
What is in here
Config
Splits
Rows
Contents
sentences
eng, spa, rus, zho
88,719… See the full description on the dataset page: https://huggingface.co/datasets/jknafou/trec-ragtime-2026.rag_instruct_benchmark_tester
Dataset Card for RAG-Instruct-Benchmark-Tester
Dataset Summary
This is an updated benchmarking test dataset for "retrieval augmented generation" (RAG) use cases in the enterprise, especially for financial services, and legal. This test dataset includes 200 questions with context passages pulled from common 'retrieval scenarios', e.g., financial news, earnings releases,
contracts, invoices, technical articles, general news and short texts.
The questions are segmented… See the full description on the dataset page: https://huggingface.co/datasets/llmware/rag_instruct_benchmark_tester.ragtopia_oldmedicalpark-rag
Medical Park Türkçe Sağlık Makaleleri — RAG Sistemi
Türkçe tıbbi makaleler üzerine kurulmuş, eşik (threshold) tabanlı bir Retrieval-Augmented Generation (RAG) altyapısı.
1. Veri Seti
Kaynak: umutertugrul/turkish-hospital-medical-articles (CC BY 4.0)
Veri seti içeriği: 14 farklı Türk hastane/sağlık kuruluşunun web sitesinden çekilmiş Türkçe tıbbi makaleler, her kuruluş ayrı bir .parquet dosyası olarak sunuluyor (toplam ~25.000 makale, 14 kaynak: Acıbadem… See the full description on the dataset page: https://huggingface.co/datasets/Toivo0/medicalpark-rag.t2-ragbench-splitsragdag-results
RAGDAG results
Artefacts from RAGDAG - treating a multi-stage retrieval pipeline as a
structural causal model and computing path-specific effects exactly by freezing
stages, rather than estimating them.
Code: https://github.com/ValerianFourel/RAGDAG
Layout
One directory per collection, named after its ir_datasets id:
<dataset-tag>/
REPORT.md human-readable report incl. the PASS/FAIL verdict
MANIFEST.json provenance: git SHA, code… See the full description on the dataset page: https://huggingface.co/datasets/ValerianFourel/ragdag-results.zoonomia-rag-v1-v1
bolinas-dna/zoonomia-rag-v1-v1
Fixed-layout 2,048-token documents built from conservation-filtered GRCh38
255-base anchors, seven fixed Zoonomia mammalian ortholog slots, and a final
human slot. Missing non-human projections are filled with 255 N bases;
chromosome 18 is validation-only.
Produced by the commit-pinned issue #402 RAG pipeline. The
immutable upstream input is the existing Zoonomia v1
min0.20/all_species_with_sequence.parquet projection. No halLiftover was
run for… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-rag-v1-v1.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/grasson/t2-ragbench.RAG_Evaluation_DatasetGrounded-RAG-RU-v2
Датасет для алайнмента (граундинга) способности LLM отвечать на вопросы по документам (RAG)
Этот датасет был собран на основе 13к разных статей из русской Википедии с помошью синтетических вопросов и ответов gpt-4-turbo-1106.
Датасет содержит 4047 уникальных кластеров, т.е. комбинаций из документов - улосвная симуляция "найденных результатов" в Retrieval системе. Подробнее описано в разделе "Общие этапы сборки этого датасета".
Общий объем датасета - 50210 уникальных диалогов.
В… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Grounded-RAG-RU-v2.full-fold-the-rag-parquet-merged0222RAGTruth_test
RAGTruth test set
Dataset
Test split of RAGTruth dataset by ParticleMedia available from https://github.com/ParticleMedia/RAGTruth/tree/main/dataset
The dataset was published in RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
Preprocessing
We kept only the test split of the original dataset
Joined response and source info files
Created the response level hallucination labels as described in the paper using binary… See the full description on the dataset page: https://huggingface.co/datasets/flowaicom/RAGTruth_test.
