Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DesmondYMTang2024 /Language-Grounded_Sparse_Encoder_Training Language-Grounded Sparse Encoder (LanSE) — Training Data This repository hosts the AI-generated images and human annotation datasets accompanying the paper: Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.textimage-classification100K<n<1M1 likes5.4k downloads1mo agoHugging Face02OpenDriveLab /SparseVideoNav SparseVideoNav Datasets This repository contains the real-world navigation datasets released with OpenDriveLab/SparseVideoNav: BVN: Beyond-the-View Navigation. IFN: Instruction-Following Navigation. Project links: Project page: https://opendrivelab.com/SparseVideoNav GitHub: https://github.com/OpenDriveLab/SparseVideoNav Paper: https://arxiv.org/abs/2602.05827 Dataset Summary SparseVideoNav studies real-world vision-language navigation with sparse future… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/SparseVideoNav.tabularrobotics10K<n<100K4 likes3.3k downloads1mo agoHugging Face03rmems /sparse-reward-long-tasks Sparse Reward Long Tasks Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/sparse-reward-long-tasks.text1K<n<10K0 likes331 downloads19d agoHugging Face04serteal /sparse-probing Sparse Probing Datasets 155 binary classification tasks for probing language model representations. From: "Are Sparse Autoencoders Useful? A Case Study in Sparse Probing" (arXiv:2502.16681) Source: EleutherAI/sae-probes Usage from datasets import load_dataset # Load a specific dataset ds = load_dataset("serteal/sparse-probing", "87_glue_cola") # List available configurations from datasets import get_dataset_config_names configs =… See the full description on the dataset page: https://huggingface.co/datasets/serteal/sparse-probing.text100K<n<1M0 likes313 downloads8mo agoHugging Face05Wuuu3511 /bmvs_sparse_dtu2imagen<1K0 likes238 downloads18d agoHugging Face06rubentium /sparse-resultstext0 likes186 downloads4d agoHugging Face07allenai /multinews_sparse_oracleThis is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The summary field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the original number of input documents for each example Retrieval results on the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multinews_sparse_oracle.textsummarization10K<n<100K1 likes133 downloads4y agoHugging Face08KinGeorge /Dr.Sparse-OTF-test-set Dr.Sparse OTF Test Set 100 sparse matrices from the SuiteSparse Matrix Collection, converted to the flat binary format the Dr.Sparse benchmark harness reads. This is the held-out evaluation set for LLM-generated CUDA sparse kernels (SpMV / SpMM / SpGEMM), kept separate from the matrices the models were developed against. Layout Matrices are grouped into size tiers by row count, the convention Dr.Sparse task discovery scans for: tier rows matrices size… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-OTF-test-set.tabularothern<1K0 likes93 downloads1mo agoHugging Face09KinGeorge /Dr.Sparse-RL-train-562 Dr.Sparse SpGEMM training pool (562 matrices) The complete RL / selector training pool of Dr.Sparse (branch v2): 562 SuiteSparse matrices in the harness .bin layout (int32 rows, cols, nnz; int32 row_ptr; int32 col_ind; float32 values; float32 x), laid out as level1_small/ (91), level2_medium/ (273), level3_large/ (198); the huge tier is deliberately left out of training and evaluation. Every matrix has a cuSPARSE SpGEMM reference (C = AA, or AA^T when rectangular) on an H200;… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-RL-train-562.tabularn<1K0 likes85 downloads23d agoHugging Face10scitomo /walnut-edip-sparse20 Scitomo Walnut EDIP Sparse-20 prepared dataset This is a derived Scitomo Sparse-20 preparation of the public Walnut-1 cone-beam X-ray CT acquisition. It is not the original Walnut archive, not a published EDIP reconstruction, and not a blessed Scitomo result. The package contains measured projections, corrected vector-cone geometry, and the published AGD-50 evaluation reference used by the maintained Scitomo Walnut EDIP evidence workflow. Provenance and attribution… See the full description on the dataset page: https://huggingface.co/datasets/scitomo/walnut-edip-sparse20.textn<1K0 likes79 downloads1mo agoHugging Face11allenai /multinews_sparse_maxThis is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The summary field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of documents seen across examples in this dataset, in this case… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multinews_sparse_max.textsummarization10K<n<100K0 likes71 downloads4y agoHugging Face12allenai /multinews_sparse_meanThis is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The summary field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "mean", i.e. the number of documents retrieved, k, is set as the mean number of documents seen across examples in this dataset, in this case k==3… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multinews_sparse_mean.textsummarization10K<n<100K1 likes71 downloads4y agoHugging Face13fabriziosalmi /simplemath-ita-sparse Dataset Card for "simplemath-ita-sparse" More Information needed text10M<n<100M1 likes66 downloads2y agoHugging Face14champion666 /SparseTable_Bench_Datasetimageimage-to-text10K<n<100K0 likes66 downloads4mo agoHugging Face15cp500 /multilingual-automotive-sparse Multilingual Automotive Sparse Retrieval Corpus A synthetic training corpus for fine-tuning multilingual sparse retrieval models (SPLADE-family) on automotive, supply-chain, and geopolitics content. Designed to teach a model cross-lingual alignment: a Japanese query should retrieve the right English passage, and vice versa. Corpus contents File Rows Purpose concepts.jsonl 4995 Raw concept records train_triplets.jsonl 134880 Flattened (query, positive… See the full description on the dataset page: https://huggingface.co/datasets/cp500/multilingual-automotive-sparse.textsentence-similarity100K<n<1M0 likes63 downloads6mo agoHugging Face16KinGeorge /Dr.Sparse-SFT-luna10-v4b Dr.Sparse SFT data — Luna-10 distillation, round v4b (Qwen3.5-9B student) The exact data behind the qwen3.5-9b-sft-v4b student (OTF-81: 46 % of held-out matrices solved with reasoning on, 58 % with reasoning off; the base model solves 0). Built from GPT-5.6 ("Luna") trajectories on the 562-matrix Dr.Sparse training pool; no OTF test-set matrix appears anywhere. path what rows qwen3.5-9b-v4b/{train,val}.parquet train on this. Rendered for the Qwen3.5 chat template:… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-SFT-luna10-v4b.tabulartext-generation1K<n<10K0 likes61 downloads12d agoHugging Face17KinGeorge /Dr.Sparse-RL-level4 Dr.Sparse RL rollout pool, level4_huge members The SpGEMM online-RL rollout pool is 130 SuiteSparse matrices: the 118 level1-3 matrices in the Dr.Sparse git repo (level_sparse/level{1_small,2_medium,3_large}) plus the 12 level4 matrices in this dataset. It is disjoint from the 100-matrix OTF test set. The 12 level4 (>= 1.1M rows) matrices, in Dr.Sparse, in the repo's raw .bin layout ([int32 rows, cols, nnz][int32 row_ptr][int32 col_ind][float32 values][float32 x], read by… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-RL-level4.textn<1K1 likes56 downloads28d agoHugging Face18allenai /cochrane_sparse_maxThis is a copy of the Cochrane dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used: query: The target field of each example corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract. retriever: BM25 via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/cochrane_sparse_max.textsummarization1K<n<10K0 likes51 downloads4y agoHugging Face19allenai /ms2_sparse_maxThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used: query: The background field of each example corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract. retriever: BM25 via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_sparse_max.textsummarization10K<n<100K0 likes50 downloads4y agoHugging Face20allenai /wcep_sparse_oracleThis is a copy of the WCEP-10 dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The summary field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the original number of input documents for each example Retrieval results on the train… See the full description on the dataset page: https://huggingface.co/datasets/allenai/wcep_sparse_oracle.textsummarization10K<n<100K0 likes48 downloads4y agoHugging Face21nanonets /small_sparse_structured_tableThis dataset is generated syhthetically to create tables with following characteristics: Empty cell percentage in following range [40,70] (Sparse) There is clear seperator between rows and columns (Structured). 4 <= num rows <= 10, 2 <= num columns <= 6 (Small) Load the dataset import io import pandas as pd from PIL import Image def bytes_to_image(self, image_bytes: bytes): return Image.open(io.BytesIO(image_bytes)) def parse_annotations(self, annotations: str) ->… See the full description on the dataset page: https://huggingface.co/datasets/nanonets/small_sparse_structured_table.texttable-question-answeringn<1K0 likes45 downloads1y agoHugging Face22hybrid-diff-ar /stack-v2-sparse-classes-36k Stack v2 Sparse Python Classes 36k This is a 36,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments. Source The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters. Splits train.jsonl: 35,000 val.jsonl: 500 test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-36k.tabulartext-generation10K<n<100K0 likes39 downloads6mo agoHugging Face23hybrid-diff-ar /stack-v2-sparse-classes-10k Stack v2 Sparse Python Classes 10k This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments. Source The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters. Splits train.jsonl: 9,000 val.jsonl: 500 test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.tabulartext-generation10K<n<100K0 likes37 downloads6mo agoHugging Face24allenai /multixscience_sparse_maxThis is a copy of the Multi-XScience dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The related_work field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of documents seen across examples in this dataset, in this… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multixscience_sparse_max.textsummarization10K<n<100K0 likes36 downloads4y agoHugging Face25allenai /multixscience_sparse_oracleThis is a copy of the Multi-XScience dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The related_work field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the original number of input documents for each example Retrieval results… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multixscience_sparse_oracle.textsummarization10K<n<100K2 likes34 downloads4y agoHugging Face26allenai /multixscience_sparse_meanThis is a copy of the Multi-XScience dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The related_work field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "mean", i.e. the number of documents retrieved, k, is set as the mean number of documents seen across examples in this dataset, in this… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multixscience_sparse_mean.textsummarization10K<n<100K1 likes33 downloads4y agoHugging Face27allenai /ms2_sparse_meanThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used: query: The background field of each example corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract. retriever: BM25 via PyTerrier with default settings top-k strategy: "mean", i.e. the number of documents retrieved, k, is set as the mean number of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_sparse_mean.textsummarization10K<n<100K0 likes33 downloads4y agoHugging Face28Aoyinke /small_sparse_datasetstextn<1K0 likes33 downloads2y agoHugging Face29allenai /ms2_sparse_oracleThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used: query: The background field of each example corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract. retriever: BM25 via PyTerrier with default settings top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the original number… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_sparse_oracle.textsummarization10K<n<100K0 likes31 downloads4y agoHugging Face30Aoyinke /retrieval_sparse_test-qrelstextn<1K0 likes31 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.