Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hyukkyu /train-dpr-wikipedia DPRWikipedia — Training, unified schema A normalised copy of the dataset behind the mteb task DPRWikipedia, a retrieval training set built from Tevatron/wikipedia-nq-corpus. Same queries, documents and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset in this collection. Source Tevatron/wikipedia-nq-corpus @ 56c6e2438c13 (the revision pinned in mteb) Domain · languages Wikipedia QA (DPR) · eng Queries /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-dpr-wikipedia.tabulartext-retrieval10M<n<100M0 likes430 downloads8d agoHugging Face02seonglae /wiki_dpr_token Dataset Card for "wiki_dpr_token" Distribution [ { // Token length '~128': 2625007, '128~256': 18370607, '256~512': 19066, '512~1024': 571, '1024~2048': 47, '2048~4096': 2, '4096~8192': 0, '8192~16384': 0, '16384~32768': 0, '32768~65536': 0, '65536~128000': 0, '128000~': 0, }, { // Text length '~512': 86519, '512~1024': 20927180, '1024~2048': 1557, '2048~4096': 43, '4096~8192': 1, '8192~16384':… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/wiki_dpr_token.tabular10M<n<100M0 likes186 downloads3y agoHugging Face03erchiw /DPriv-Bench DPrivBench: Benchmarking LLMs’ Reasoning for Differential Privacy DPrivBench is a benchmark for evaluating whether language models can correctly reason about and verify claimed differential privacy (DP) guarantees from natural-language/LaTeX-format problem statements. This release contains evaluation data from seven benchmark configs, along with one auxiliary function bank: Category 1: 6 fundamental mechanism tracks, each with 98 questions. Category 2: 125 more advanced… See the full description on the dataset page: https://huggingface.co/datasets/erchiw/DPriv-Bench.tabularn<1K2 likes121 downloads4mo agoHugging Face04YPL67 /DPR Diabetic Patient 30-Day Readmission Dataset Dataset Summary This dataset is derived from the Kaggle Diabetic Patients Readmission Prediction dataset. The original dataset contains electronic health record data from diabetic patient hospital encounters and is commonly used for hospital readmission prediction. This released version organizes the repository into three levels of data: Raw Data: the original downloaded source files. Intermediate Data: derived files before… See the full description on the dataset page: https://huggingface.co/datasets/YPL67/DPR.tabulartabular-classification100K<n<1M0 likes117 downloads5mo agoHugging Face05DinoStackAI /telco-dpr-rag Telco-DPR RAG Dataset for Retrieval-Augmented Generation (RAG) based on Telco-DPR. Structure Subset Splits Description corpus train (default) 3GPP technical passages (text + tables) shared across all query splits queries train, dev, test Synthetic telecom QA questions qrels train, dev, test Relevance judgments (query ↔ passage) answers train, dev, test Reference answers Dataset statistics Split Queries Corpus train… See the full description on the dataset page: https://huggingface.co/datasets/DinoStackAI/telco-dpr-rag.tabularquestion-answering10K<n<100K0 likes115 downloads3mo agoHugging Face06rulins /massive_serve_dpr_wiki_contriever_ivfpqtabular10M<n<100M0 likes78 downloads1y agoHugging Face07dormosol /convmix_dpr_wiki-text-6-3-tamber_top100tabular1K<n<10K0 likes75 downloads1mo agoHugging Face08Seungjun /dp_removed_aime_2026 dp_removed_aime_2026 Source: MathArena/aime_2026 Pinned source revision: d2de22f3c656b4f56cf8981212186377d1e23bc3. Config: default. Split: train. Retained rows: 30. All retained rows and original columns are unchanged, including English problems, solutions, reasoning traces, prompt wrappers and source IDs. No translation, text repair, model generation, or added data columns were applied. Stable preparation IDs are recorded separately in provenance.json. This evaluation split is… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/dp_removed_aime_2026.tabularn<1K0 likes71 downloads11d agoHugging Face09shylee /eval_DPRecordTestThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 224, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shylee/eval_DPRecordTest.tabularroboticsn<1K0 likes27 downloads1y agoHugging Face10rulins /massive_serve_dpr_wiki_e5_base_v2_ivfpqtabular10M<n<100M0 likes25 downloads1y agoHugging Face11rulins /massive_serve_dpr_wiki_qwen3_0.6b_ivfpqtabular10M<n<100M0 likes24 downloads1y agoHugging Face12dormosol /TopiOCQA_dpr_pyserini_top100_with_texttabular10K<n<100K0 likes24 downloads7mo agoHugging Face13xudongwu /DPR_Q0.5B_U10_beta0.10g0.30gamma0.30tabularn<1K0 likes22 downloads6mo agoHugging Face14Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-2-OP-False-train-perplexitytabularn<1K0 likes19 downloads2y agoHugging Face15Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-1-OP-True-train-perplexitytabular1K<n<10K0 likes19 downloads2y agoHugging Face16Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-2-OP-True-train-perplexitytabularn<1K0 likes17 downloads2y agoHugging Face17Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-1-OP-False-train-perplexitytabularn<1K0 likes16 downloads2y agoHugging Face18Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-2-OP-False-train-PI_KFI-perplexitytabularn<1K0 likes12 downloads2y agoHugging Face19Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-1-OP-False-train-perplexitytabularn<1K0 likes10 downloads2y agoHugging Face20Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-2-OP-True-train-perplexitytabular1K<n<10K0 likes10 downloads2y agoHugging Face21SKIML-ICL /med_retrieved_dpr_ctxstabularn<1K0 likes10 downloads11mo agoHugging Face22wxd02 /DPR_Pm3B_U10_beta0.10g0.30gamma0.30tabularn<1K0 likes9 downloads9mo agoHugging Face23xudongwu /DPR_Pm3B_U10_beta0.10g0.30gamma0.30tabular1K<n<10K0 likes9 downloads6mo agoHugging Face24Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-2-OP-True-train-PI_KFI-perplexitytabularn<1K0 likes8 downloads2y agoHugging Face25Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-1-OP-True-train-PI_KFI-perplexitytabularn<1K0 likes8 downloads2y agoHugging Face26Self-GRIT /PILE_Wikipedia_validation_set_insert_ret_tokens-wikipedia-dpr-k-1-OP-Truetabularn<1K0 likes8 downloads2y agoHugging Face27Self-GRIT /PILE_Wikipedia_validation_set_insert_ret_tokens-wikipedia-dpr-k-1-OP-Falsetabularn<1K0 likes8 downloads2y agoHugging Face28wxd02 /DPR_Q0.5B_U10_beta0.10g0.30gamma0.30tabularn<1K0 likes8 downloads9mo agoHugging Face29Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-2-OP-True-train-PI_KFI_IK-perplexitytabularn<1K0 likes7 downloads2y agoHugging Face30Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-1-OP-False-train-PI_KFI_IK-perplexitytabularn<1K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.