Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01facebook /wiki_dprThis is the wikipedia split used to evaluate the Dense Passage Retrieval (DPR) model. It contains 21M passages from wikipedia along with their DPR embeddings. The wikipedia articles were split into multiple, disjoint text blocks of 100 words as passages.fill-mask10M<n<100M45 likes30k downloads3y agoHugging Face02timf34 /dprobe-resultstextn<1K0 likes3.6k downloads20h agoHugging Face03hf-internal-testing /wiki_dpr_dummyThis dummy dataset is used for testing purpose for rag model in transformers. It is proudced via the following steps: dataset = datasets.load_dataset("wiki_dpr", with_embeddings=True, with_index=True, index_name="exact", embeddings_name="nq", dummy=True, revision=None) dataset["train"].drop_index("embeddings") dataset.push_to_hub("hf-internal-testing/wiki_dpr_dummy", token="...") The index file `index.faiss` (after being renamed locally) is then uploaded manually. text10K<n<100K0 likes2.1k downloads1y agoHugging Face04huggingface /wiki_dpr0 likes813 downloads2y agoHugging Face05jxm /nq_corpus_dpr Dataset Card for "nq_corpus_dpr" More Information needed text1M<n<10M2 likes667 downloads3y agoHugging Face06kenhktsui /wiki_dpr_e5wiki_dpr encoded with intfloat/e5-base-v2 text10M<n<100M0 likes529 downloads3y agoHugging Face07Hyukkyu /train-dpr-wikipedia DPRWikipedia — Training, unified schema A normalised copy of the dataset behind the mteb task DPRWikipedia, a retrieval training set built from Tevatron/wikipedia-nq-corpus. Same queries, documents and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset in this collection. Source Tevatron/wikipedia-nq-corpus @ 56c6e2438c13 (the revision pinned in mteb) Domain · languages Wikipedia QA (DPR) · eng Queries /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-dpr-wikipedia.tabulartext-retrieval10M<n<100M0 likes430 downloads8d agoHugging Face08ysenarath /wikipedia-20240901-dprtext1M<n<10M0 likes396 downloads2y agoHugging Face09lance-format /dpr-wikipedia-single-nq DPR Wikipedia single-NQ 21,015,300 base vectors and 3,610 Natural Questions test query vectors, each with 768 float32 dimensions. All original passage vectors are retained, without normalization. The corpus contains embeddings and passage-ID mappings, not passage text. Provenance and attribution The passage embeddings are the original Facebook Research DPR single-NQ Wikipedia embeddings, downloaded in shard order 0 through 49 from… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/dpr-wikipedia-single-nq.text-retrieval10M<n<100M0 likes305 downloads19d agoHugging Face10ittia /wiki_dprThe project contains indexes, datasets, checkpoints for RAG training and research. Sources checkpoint: ColBERTv2 (https://downloads.cs.stanford.edu/nlp/data/colbert/colbertv2/colbertv2.0.tar.gz) dataset: wiki_dpr (https://github.com/facebookresearch/DPR/blob/main/dpr/data/download_data.py) indexes (https://github.com/ittia-research/check/tree/main/datasets/wiki_dpr) indexing config: ColBERTConfig(nbits=2, doc_maxlen=220) Didn't compress the index files to one archive because… See the full description on the dataset page: https://huggingface.co/datasets/ittia/wiki_dpr.fill-mask10M<n<100M0 likes299 downloads2y agoHugging Face11coref-data /dpr_raw "definite_pronoun_resolution" (dpr) Dataset Summary Composed by 30 students from one of the author's undergraduate classes. These sentence pairs cover topics ranging from real events (e.g., Iran's plan to attack the Saudi ambassador to the U.S.) to events/characters in movies (e.g., Batman) and purely imaginary situations, largely reflecting the pop culture as perceived by the American kids born in the early 90s. Each annotated example spans four lines: the first line… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/dpr_raw.text1K<n<10K0 likes254 downloads3y agoHugging Face12seonglae /wiki_dpr_token Dataset Card for "wiki_dpr_token" Distribution [ { // Token length '~128': 2625007, '128~256': 18370607, '256~512': 19066, '512~1024': 571, '1024~2048': 47, '2048~4096': 2, '4096~8192': 0, '8192~16384': 0, '16384~32768': 0, '32768~65536': 0, '65536~128000': 0, '128000~': 0, }, { // Text length '~512': 86519, '512~1024': 20927180, '1024~2048': 1557, '2048~4096': 43, '4096~8192': 1, '8192~16384':… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/wiki_dpr_token.tabular10M<n<100M0 likes186 downloads3y agoHugging Face13thainasaraiva /Telco-DPR Dataset Information This paper proposes a Question-Answering (QA) system for the telecom domain using 3rd Generation Partnership Project (3GPP) technical documents. Alongside, a hybrid dataset, Telco-DPR, which consists of a curated 3GPP corpus in a hybrid format, combining text and tables, is presented. Additionally, the dataset includes a set of synthetic question/answer pairs designed to evaluate the retrieval performance of QA systems on this type of data. The retrieval… See the full description on the dataset page: https://huggingface.co/datasets/thainasaraiva/Telco-DPR.text10K<n<100K3 likes134 downloads8mo agoHugging Face14Seungjun /dp_removed_DAPO-Math-17k-Processed dp_removed_DAPO-Math-17k-Processed Source: open-r1/DAPO-Math-17k-Processed Pinned source revision: 31dd309567e3da778038cc87d868b6097a3ccf68. Config: en. Split: train. Retained rows: 14068. All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json.… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_DAPO-Math-17k-Processed.text10K<n<100K0 likes129 downloads11d agoHugging Face15erchiw /DPriv-Bench DPrivBench: Benchmarking LLMs’ Reasoning for Differential Privacy DPrivBench is a benchmark for evaluating whether language models can correctly reason about and verify claimed differential privacy (DP) guarantees from natural-language/LaTeX-format problem statements. This release contains evaluation data from seven benchmark configs, along with one auxiliary function bank: Category 1: 6 fundamental mechanism tracks, each with 98 questions. Category 2: 125 more advanced… See the full description on the dataset page: https://huggingface.co/datasets/erchiw/DPriv-Bench.tabularn<1K2 likes121 downloads4mo agoHugging Face16YPL67 /DPR Diabetic Patient 30-Day Readmission Dataset Dataset Summary This dataset is derived from the Kaggle Diabetic Patients Readmission Prediction dataset. The original dataset contains electronic health record data from diabetic patient hospital encounters and is commonly used for hospital readmission prediction. This released version organizes the repository into three levels of data: Raw Data: the original downloaded source files. Intermediate Data: derived files before… See the full description on the dataset page: https://huggingface.co/datasets/YPL67/DPR.tabulartabular-classification100K<n<1M0 likes117 downloads5mo agoHugging Face17yhabushi /test-dpr0 likes116 downloads2y agoHugging Face18dormosol /asqa_dpr_wiki-text-6-3-tamber_top100text1K<n<10K0 likes116 downloads1mo agoHugging Face19Seungjun /dp_removed_s1K-1.1 dp_removed_s1K-1.1 Source: simplescaling/s1K-1.1 Pinned source revision: 96c411f1fe4c49d20f0e2a1565f61e1a28b0b84d. Config: default. Split: train. Retained rows: 996. All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json. Training/evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_s1K-1.1.textn<1K0 likes116 downloads5d agoHugging Face20DinoStackAI /telco-dpr-rag Telco-DPR RAG Dataset for Retrieval-Augmented Generation (RAG) based on Telco-DPR. Structure Subset Splits Description corpus train (default) 3GPP technical passages (text + tables) shared across all query splits queries train, dev, test Synthetic telecom QA questions qrels train, dev, test Relevance judgments (query ↔ passage) answers train, dev, test Reference answers Dataset statistics Split Queries Corpus train… See the full description on the dataset page: https://huggingface.co/datasets/DinoStackAI/telco-dpr-rag.tabularquestion-answering10K<n<100K0 likes115 downloads3mo agoHugging Face21Seungjun /dp_removed_amc23 dp_removed_amc23 Source: math-ai/amc23 Pinned source revision: 80815d37005feb82cd7f8fbc6901d5d3eff43057. Config: default. Split: test. Retained rows: 40. All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json. This evaluation split is an… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_amc23.textn<1K0 likes86 downloads11d agoHugging Face22Seungjun /dp_removed_MATH-500 dp_removed_MATH-500 Source: HuggingFaceH4/MATH-500 Pinned source revision: 6e4ed1a2a79af7d8630a6b768ec859cb5af4d3be. Config: default. Split: test. Retained rows: 500. All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json. This evaluation split… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_MATH-500.textn<1K0 likes79 downloads11d agoHugging Face23rulins /massive_serve_dpr_wiki_contriever_ivfpqtabular10M<n<100M0 likes78 downloads1y agoHugging Face24rulins /massive_serve_dpr_wiki_contriever0 likes76 downloads1y agoHugging Face25dormosol /convmix_dpr_wiki-text-6-3-tamber_top100tabular1K<n<10K0 likes75 downloads1mo agoHugging Face26Seungjun /dp_removed_aime_2026 dp_removed_aime_2026 Source: MathArena/aime_2026 Pinned source revision: d2de22f3c656b4f56cf8981212186377d1e23bc3. Config: default. Split: train. Retained rows: 30. All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json. This evaluation split is… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_aime_2026.tabularn<1K0 likes71 downloads11d agoHugging Face27Seungjun /dp_removed_aime_2024 dp_removed_aime_2024 Source: HuggingFaceH4/aime_2024 Pinned source revision: 2fe88a2f1091d5048c0f36abc874fb997b3dd99a. Config: default. Split: train. Retained rows: 30. All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json. This evaluation split… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_aime_2024.textn<1K0 likes62 downloads11d agoHugging Face28Seungjun /dp_removed_aime25 dp_removed_aime25 Source: math-ai/aime25 Pinned source revision: 563bb8404243c5f09de6ec262f2db674fe5bce9b. Config: default. Split: test. Retained rows: 30. All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json. This evaluation split is an… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_aime25.textn<1K0 likes62 downloads11d agoHugging Face29rulins /massive_serve_dpr_wiki_qwen3_0.6b0 likes60 downloads1y agoHugging Face30ismailduru /Telco-DPR Dataset Information This paper proposes a Question-Answering (QA) system for the telecom domain using 3rd Generation Partnership Project (3GPP) technical documents. Alongside, a hybrid dataset, Telco-DPR, which consists of a curated 3GPP corpus in a hybrid format, combining text and tables, is presented. Additionally, the dataset includes a set of synthetic question/answer pairs designed to evaluate the retrieval performance of QA systems on this type of data. The retrieval… See the full description on the dataset page: https://huggingface.co/datasets/ismailduru/Telco-DPR.text10K<n<100K0 likes57 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.