datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiki_dprThis is the wikipedia split used to evaluate the Dense Passage Retrieval (DPR) model.
It contains 21M passages from wikipedia along with their DPR embeddings.
The wikipedia articles were split into multiple, disjoint text blocks of 100 words as passages.dprobe-resultswiki_dpr_dummyThis dummy dataset is used for testing purpose for rag model in transformers. It is proudced via the following steps:
dataset = datasets.load_dataset("wiki_dpr", with_embeddings=True, with_index=True, index_name="exact", embeddings_name="nq", dummy=True, revision=None)
dataset["train"].drop_index("embeddings")
dataset.push_to_hub("hf-internal-testing/wiki_dpr_dummy", token="...")
The index file `index.faiss` (after being renamed locally) is then uploaded manually.
wiki_dprnq_corpus_dpr
Dataset Card for "nq_corpus_dpr"
More Information needed
wiki_dpr_e5wiki_dpr encoded with intfloat/e5-base-v2
train-dpr-wikipedia
DPRWikipedia — Training, unified schema
A normalised copy of the dataset behind the mteb task DPRWikipedia, a retrieval training set built from Tevatron/wikipedia-nq-corpus. Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
Tevatron/wikipedia-nq-corpus @ 56c6e2438c13 (the revision pinned in mteb)
Domain · languages
Wikipedia QA (DPR) · eng
Queries /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-dpr-wikipedia.wikipedia-20240901-dprdpr-wikipedia-single-nq
DPR Wikipedia single-NQ
21,015,300 base vectors and 3,610 Natural Questions test query vectors, each with 768 float32 dimensions. All original passage vectors are retained, without normalization. The corpus contains embeddings and passage-ID mappings, not passage text.
Provenance and attribution
The passage embeddings are the original Facebook Research DPR single-NQ Wikipedia embeddings, downloaded in shard order 0 through 49 from… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/dpr-wikipedia-single-nq.wiki_dprThe project contains indexes, datasets, checkpoints for RAG training and research.
Sources
checkpoint: ColBERTv2 (https://downloads.cs.stanford.edu/nlp/data/colbert/colbertv2/colbertv2.0.tar.gz)
dataset: wiki_dpr (https://github.com/facebookresearch/DPR/blob/main/dpr/data/download_data.py)
indexes (https://github.com/ittia-research/check/tree/main/datasets/wiki_dpr)
indexing config: ColBERTConfig(nbits=2, doc_maxlen=220)
Didn't compress the index files to one archive because… See the full description on the dataset page: https://huggingface.co/datasets/ittia/wiki_dpr.dpr_raw
"definite_pronoun_resolution" (dpr)
Dataset Summary
Composed by 30 students from one of the author's undergraduate classes. These
sentence pairs cover topics ranging from real events (e.g., Iran's plan to
attack the Saudi ambassador to the U.S.) to events/characters in movies (e.g.,
Batman) and purely imaginary situations, largely reflecting the pop culture as
perceived by the American kids born in the early 90s. Each annotated example
spans four lines: the first line… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/dpr_raw.wiki_dpr_token
Dataset Card for "wiki_dpr_token"
Distribution
[
{ // Token length
'~128': 2625007,
'128~256': 18370607,
'256~512': 19066,
'512~1024': 571,
'1024~2048': 47,
'2048~4096': 2,
'4096~8192': 0,
'8192~16384': 0,
'16384~32768': 0,
'32768~65536': 0,
'65536~128000': 0,
'128000~': 0,
},
{ // Text length
'~512': 86519,
'512~1024': 20927180,
'1024~2048': 1557,
'2048~4096': 43,
'4096~8192': 1,
'8192~16384':… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/wiki_dpr_token.Telco-DPR
Dataset Information
This paper proposes a Question-Answering (QA) system for the telecom domain using 3rd Generation Partnership Project (3GPP) technical documents.
Alongside, a hybrid dataset, Telco-DPR, which consists of a curated 3GPP corpus in a hybrid format, combining text and tables, is presented.
Additionally, the dataset includes a set of synthetic question/answer pairs designed to evaluate the retrieval performance of QA systems on this type of data.
The retrieval… See the full description on the dataset page: https://huggingface.co/datasets/thainasaraiva/Telco-DPR.dp_removed_DAPO-Math-17k-Processed
dp_removed_DAPO-Math-17k-Processed
Source: open-r1/DAPO-Math-17k-Processed
Pinned source revision: 31dd309567e3da778038cc87d868b6097a3ccf68. Config: en. Split: train. Retained rows: 14068.
All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json.… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_DAPO-Math-17k-Processed.DPriv-Bench
DPrivBench: Benchmarking LLMs’ Reasoning for Differential Privacy
DPrivBench is a benchmark for evaluating whether language models can correctly reason about and verify claimed differential privacy (DP) guarantees from natural-language/LaTeX-format problem statements.
This release contains evaluation data from seven benchmark configs, along with one auxiliary function bank:
Category 1: 6 fundamental mechanism tracks, each with 98 questions.
Category 2: 125 more advanced… See the full description on the dataset page: https://huggingface.co/datasets/erchiw/DPriv-Bench.DPR
Diabetic Patient 30-Day Readmission Dataset
Dataset Summary
This dataset is derived from the Kaggle Diabetic Patients Readmission Prediction dataset. The original dataset contains electronic health record data from diabetic patient hospital encounters and is commonly used for hospital readmission prediction.
This released version organizes the repository into three levels of data:
Raw Data: the original downloaded source files.
Intermediate Data: derived files before… See the full description on the dataset page: https://huggingface.co/datasets/YPL67/DPR.test-dprasqa_dpr_wiki-text-6-3-tamber_top100dp_removed_s1K-1.1
dp_removed_s1K-1.1
Source: simplescaling/s1K-1.1
Pinned source revision: 96c411f1fe4c49d20f0e2a1565f61e1a28b0b84d. Config: default. Split: train. Retained rows: 996.
All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json.
Training/evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_s1K-1.1.telco-dpr-rag
Telco-DPR RAG
Dataset for Retrieval-Augmented Generation (RAG) based on Telco-DPR.
Structure
Subset
Splits
Description
corpus
train (default)
3GPP technical passages (text + tables) shared across all query splits
queries
train, dev, test
Synthetic telecom QA questions
qrels
train, dev, test
Relevance judgments (query ↔ passage)
answers
train, dev, test
Reference answers
Dataset statistics
Split
Queries
Corpus
train… See the full description on the dataset page: https://huggingface.co/datasets/DinoStackAI/telco-dpr-rag.dp_removed_amc23
dp_removed_amc23
Source: math-ai/amc23
Pinned source revision: 80815d37005feb82cd7f8fbc6901d5d3eff43057. Config: default. Split: test. Retained rows: 40.
All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json.
This evaluation split is an… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_amc23.dp_removed_MATH-500
dp_removed_MATH-500
Source: HuggingFaceH4/MATH-500
Pinned source revision: 6e4ed1a2a79af7d8630a6b768ec859cb5af4d3be. Config: default. Split: test. Retained rows: 500.
All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json.
This evaluation split… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_MATH-500.massive_serve_dpr_wiki_contriever_ivfpqmassive_serve_dpr_wiki_contrieverconvmix_dpr_wiki-text-6-3-tamber_top100dp_removed_aime_2026
dp_removed_aime_2026
Source: MathArena/aime_2026
Pinned source revision: d2de22f3c656b4f56cf8981212186377d1e23bc3. Config: default. Split: train. Retained rows: 30.
All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json.
This evaluation split is… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_aime_2026.dp_removed_aime_2024
dp_removed_aime_2024
Source: HuggingFaceH4/aime_2024
Pinned source revision: 2fe88a2f1091d5048c0f36abc874fb997b3dd99a. Config: default. Split: train. Retained rows: 30.
All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json.
This evaluation split… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_aime_2024.dp_removed_aime25
dp_removed_aime25
Source: math-ai/aime25
Pinned source revision: 563bb8404243c5f09de6ec262f2db674fe5bce9b. Config: default. Split: test. Retained rows: 30.
All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json.
This evaluation split is an… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_aime25.massive_serve_dpr_wiki_qwen3_0.6bTelco-DPR
Dataset Information
This paper proposes a Question-Answering (QA) system for the telecom domain using 3rd Generation Partnership Project (3GPP) technical documents.
Alongside, a hybrid dataset, Telco-DPR, which consists of a curated 3GPP corpus in a hybrid format, combining text and tables, is presented.
Additionally, the dataset includes a set of synthetic question/answer pairs designed to evaluate the retrieval performance of QA systems on this type of data.
The retrieval… See the full description on the dataset page: https://huggingface.co/datasets/ismailduru/Telco-DPR.
