datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RAG_Evaluation_Datasetrag-qa-logs-corpus-data
🧠📚 RAG QA Logs & Corpus (Synthetic)
🧪 Multi-table synthetic RAG telemetry for quality, hallucinations, latency, and cost
A production-style, privacy-safe synthetic dataset that mimics telemetry exported from a real RAG system — from corpus → chunks → retrieval events → eval runs.
✅ Fully synthetic (no real users / orgs / PII).
⚡ Quick facts
Total rows: 103,255 across 6 linked tables
Labels (in eval_runs): is_correct, hallucination_flag, faithfulness_label… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/rag-qa-logs-corpus-data.kor-rag-opentestturkish-legal-rag
Turkish Legal RAG Corpus — Türk Hukuku için Açık RAG Datasetı
Tek cümle: 25 önemli Türk kanununun (mevzuat.gov.tr kaynaklı, madde bazlı temiz chunk'lar) + 290 manuel doğrulanmış soru-cevap altın benchmark'ının olduğu açık kaynak Türkçe hukuk RAG datasetı.
🇹🇷 Türkçe Özet — Bu dataset, Türkçe hukuk uygulamaları için sıfırdan üretilmiş açık ve denetlenebilir bir RAG corpus'udur. mevzuat.gov.tr üzerinden alınan 25 ana kanunun madde madde temizlenmiş, chunk'lanmış sürümünü (6.350… See the full description on the dataset page: https://huggingface.co/datasets/mtntasci/turkish-legal-rag.turkuaz-rag
Turkuaz-RAG: A Novel Turkish Multi-Context Retrieval Benchmark
Turkuaz-RAG is the first benchmark specifically created for evaluating multi-context retrieval tasks in Turkish. It addresses a major gap in low-resource language research by providing multi-context questions, answers, and corresponding contexts.
Description of Benchmark
Languages: Turkish
Size: ~2,500 triplets (question, contexts, answer)
Context Sources: Turkish news articles from MLSUM
Question Types:… See the full description on the dataset page: https://huggingface.co/datasets/eneSadi/turkuaz-rag.rag-hallucination-benchmark
RAG Hallucination Benchmark
Context
Retrieval-Augmented Generation (RAG) is the industry standard for reducing LLM hallucinations, but detecting when a RAG system fails is a massive challenge. Most existing benchmarks focus only on massive Deep Learning models and lack tabular features.
This dataset provides a clean, engineered setup to train models (from XGBoost to RoBERTa) to detect hallucinations, predict context faithfulness, and measure answer relevance.… See the full description on the dataset page: https://huggingface.co/datasets/vkshdev/rag-hallucination-benchmark.RAG_vs_FineTuning_Comparison_Persian_V1privacy-preserving-real-world-human-motion-sample
Privacy-Preserving Real-World Human Motion Sample
A market-validation sample of anonymous 2D skeleton/pose observations derived from a real-world indoor CCTV stream.
Why this sample exists
We are validating demand for continuously collected, privacy-oriented real-world human-motion data before expanding to multi-camera releases.
Current public sample
750 public observations
derived pose/skeleton data
anonymous track identifiers
no raw RGB video
no… See the full description on the dataset page: https://huggingface.co/datasets/Ragab-Adel/privacy-preserving-real-world-human-motion-sample.multi-tafseer-quran-rag
Quran Tafseer RAG Dataset
A structured Arabic dataset of Quranic tafseer collected from eight classical and modern tafseer books.The dataset contains verse-aligned tafseer passages designed for Retrieval-Augmented Generation (RAG) systems and Arabic NLP research.
Each record links a Quran verse with its corresponding tafseer explanation from one of the tafseer books and includes rich metadata such as surah information, tafseer source, and embedding-ready text.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/omaressam1111/multi-tafseer-quran-rag.ragscale-interaction-matrix
ragscale Interaction Matrix
Reader answers under raw and compressed RAG evidence: 176,864 rows, one per benchmark item, reader model, and evidence policy, across LongMemEval, HotpotQA, MuSiQue, and NQ-Open.
This is the interaction matrix released with the paper Compression Is Not Evaluation-Neutral: Fixed RAG Compression Can Distort Reader Comparisons (Sugam Panthi and Rabab Abdelfattah, arXiv:2606.21807). The paper gives 8 to 20 reader models the same stored compressed text and… See the full description on the dataset page: https://huggingface.co/datasets/vein05/ragscale-interaction-matrix.cs50-educational-rag
CS50 Pedagogical RAG Dataset
📜 Dataset Description
This repository contains the data artifacts for the undergraduate thesis, which explores the use of a pedagogical chatbot with Retrieval-Augmented Generation (RAG) for Harvard's CS50: Introduction to Computer Science course.
The project involved several stages of data processing, from raw content collection to the generation and curation of a high-quality evaluation dataset. To ensure full transparency and… See the full description on the dataset page: https://huggingface.co/datasets/dev-jonathanb/cs50-educational-rag.legal-rag-positives-synthetic
Synthetic QnA Chunk Pairs from Legal Documents
This dataset contains excerpts from legal cases' court opinions that mention artificial intelligence, along with corresponding question-answer pairs derived from the content. The data was sourced from CourtListener's public API and processed to create a structured dataset suitable for question-answering tasks.
Specifically including cases:
Senetas Corporation, Ltd. v. DeepRadiology Corporation
Electronic Privacy Information Center v.… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/legal-rag-positives-synthetic.rag-eval-ja-repro
RAG Eval JA Repro
Current version / 現行版: v1.1
2026-07-12 更新(v1.1): rag_evaluation_master.csv、採用PDF manifest、PDF checksumを更新し、6月30日公開時のローカル精度検証を同じ4条件で再実行しました。旧版の記述は取り消し線で残し、v1.1の値を併記します。
TL;DR (EN): A derived reproducibility dataset for allganize/RAG-Evaluation-Dataset-JA.
It adds (1) derived *_new answer/question columns (with per-item rationale), and (2) a Wayback-pinned + SHA-256 corpus manifest so anyone can fetch byte-identical source PDFs.
The original CSV is not modified;… See the full description on the dataset page: https://huggingface.co/datasets/SakataConsul/rag-eval-ja-repro.swiss-building-law-rag-bench
Swiss Cantonal Building Law RAG Benchmark
Evaluation benchmark for Retrieval-Augmented Generation (RAG) systems on Swiss cantonal
building law documents. Created as part of a bachelor thesis on systematic RAG pipeline
optimisation for German legal text.
Dataset contents
File
Entries
Language
Description
data/german/golden_dataset.jsonl
318
DE
German Q&A pairs grounded to article-level passages
data/multilingual/golden_dataset.jsonl
270
DE/FR/IT… See the full description on the dataset page: https://huggingface.co/datasets/MarcoFurrer/swiss-building-law-rag-bench.eou_AudioTextRagabilityCorpusCurrent version: Dataset_v0.4.tsv (converted to ragability format: v0d4.hjson)
Ragability Corpus
In the following, we introduce WikiContradict (the empirical basis for the Ragability Corpus), describe the Ragability Corpus, and finally explain how the dataset can be extended and how a new one can be created.
Empirical basis
WikiContradict is a benchmark for evaluating LLMs on real-world knowledge conflicts from Wikipedia (see the Hou et. al. 2025 and the dataset for more… See the full description on the dataset page: https://huggingface.co/datasets/ofai/RagabilityCorpus.turkish-hospital-medical-rag-advanced
Turkish Hospital Medical Articles - Advanced RAG System & Vector Database
Bu proje, 14 farklı hastane grubuna ait geniş ölçekli Türkçe tıbbi makaleler üzerinde çalışan, ileri düzey teknik parametrelerle optimize edilmiş bir Retrieval-Augmented Generation (RAG) sistemi ve Vektör Veritabanı uygulamasıdır. Proje kapsamında ham veriler Hugging Face üzerinden tüm hastane split'leriyle çekilmiş, gelişmiş chunking stratejileriyle parçalanmış, magibu/embeddingmagibu-200m modeliyle… See the full description on the dataset page: https://huggingface.co/datasets/Egertekin/turkish-hospital-medical-rag-advanced.Pubmed-RAG-TR-LLM-EvalLLM-as-a-judge evalaution results using "claude-haiku-4-5-20251001" for SMARTICT/Pubmed-RAG-TR-LLM dataset.
algozee_rag-based-hallucination-reduction-in-llms
RAG-Based Hallucination Reduction in LLMs
Introduction to Large Language Models and Hallucination Problem
Dataset Info
Source: Kaggle
Original Size: 0.17 MB
Kaggle Downloads: 43
Files: 1
Files
llm_rag_dataset_6k.csv.csv
Mirrored from Kaggle
legal-rag-qaragbench-5drag_evalSome datasets for evaluating RAG systems, created by following this huggingface cookbook.
ragas_evaluationV1healthcare_datasetRAG12000-LLaMA3.1-8B-gguf_AR-RAG_v2superkart-sales-datasetGL-Raghu-engine-predictive-maintainencetest_123court_data_RAG_unsupMachine-Failure-Prediction
