datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI
doc-formats-csv-1
[doc] formats - csv - 1
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
ReactiveGWM-Datasets
ReactiveGWM-Datasets: Strategy-Aligned Rollouts for Reactive Game World Models
📚 Datasets-Introduction
ReactiveGWM-Datasets is the strategy-aligned training corpus that powers
ReactiveGWM, a game world
model that decouples player control from NPC autonomy. To learn that
decoupling, the model needs supervision that pairs each gameplay clip with
both a per-frame action stream (what the player did) and a high-level
NPC description (what the NPC tried to do, and under… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-Datasets.OpenVLHarness-Evaluation-Datasets
OpenVLHarness evaluation datasets
Processed evaluation splits used by
OpenVLHarness
(project page).
Each <split>.tsv holds the exact prompts (question) and annotations
(answer plus metadata) we evaluate on; image_path is relative to
images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13
configs and COCO-format val/test annotations used for ODinW AP evaluation.
You normally don't need to download anything by hand: running
openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.MMLA-Datasets
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
1. Introduction
MMLA is the first comprehensive multimodal language analysis benchmark for evaluating foundation models. It has the following features:
Large Scale: 61K+ multimodal samples.
Various Sources: 9 datasets.
Three Modalities: text, video, and audio
Both Acting and Real-world Scenarios: films, TV series, YouTube, Vimeo, Bilibili, TED, improvised scripts, etc.
Six Core… See the full description on the dataset page: https://huggingface.co/datasets/THUIAR/MMLA-Datasets.tsf-datasetssim-datasets
SIM-Datasets: A Unified Symbolic Regression Benchmark
A standardized benchmark collection designed for the Scientific Intelligent Modelling (SIM) toolkit, providing comprehensive datasets for symbolic regression research and applications.
Overview
SIM-Datasets serves as a unified benchmark for symbolic regression tasks, offering standardized datasets with consistent formatting and evaluation protocols. This collection is specifically curated to support the Scientific… See the full description on the dataset page: https://huggingface.co/datasets/scientific-intelligent-modelling/sim-datasets.Dataset-Solubility
Description
This dataset contains 71419 amino acid sequences and its solubility label.
Protein Format: AA sequence
Splits
traing: 62478
valid: 6942
test: 1999
Related paper
The dataset is from DeepSol: a deep learning framework for sequence-based protein solubility prediction.
Label
Binary label, 1 means soluble, 0 means insoluble.
chess_datasets
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/p11-p11/chess_datasets.detector-clickbait-br-datasets
Detector Clickbait BR - Datasets
Este repositório contém os datasets utilizados para o treinamento do modelo detector-clickbait-br-model, um classificador de textos em português brasileiro capaz de identificar títulos clickbait.
📚 Descrição dos Datasets
1. detector-clickbait-br-raw.csv
Dataset original contendo os dados iniciais sem processamento.
Características:
Dados brutos coletados originalmente
Pode conter duplicatas
Pode conter valores nulos
Formato:… See the full description on the dataset page: https://huggingface.co/datasets/rodrigoaraujorosa/detector-clickbait-br-datasets.X-PAIR_datasets
X-PAIR datasets
This repository contains the processed datasets used to train and evaluate X-PAIR, an ultrafast multitask framework for protein–protein interaction (PPI) and partner-specific interface prediction from protein sequences.
The datasets are provided in the train, validation, and test splits used in the experiments reported in the X-PAIR study.
Datasets
The repository contains two groups of datasets:
Datasets generated in the X-PAIR study
Processed… See the full description on the dataset page: https://huggingface.co/datasets/srescalli/X-PAIR_datasets.hhem_leaderboard_datasetshotel_datasetsReactiveGWM-v2-Datasets
ReactiveGWM v2 Datasets
This repository contains the selected HNM and Street Fighter III: New
Generation (SF3) datasets. The original FFV1/MKV videos, first frames, masks,
native annotations, action records, indices, and prepared VAE/T5 caches are
stored without compression in tar volumes of approximately 2 GiB. Files
retain their original formats and paths inside the tar archives. Game-generation
code, ROMs, BIOS files, and runtime binaries are not included.
Subset
Train… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-v2-Datasets.multi-turn_jailbreak_attack_datasets
Multi-Turn Jailbreak Attack Datasets
Description
This dataset was created to compare single-turn and multi-turn jailbreak attacks on large language models (LLMs). The primary goal is to take a single harmful prompt and distribute the harm over multiple turns, making each prompt appear harmless in isolation. This approach is compared against traditional single-turn attacks with the complete prompt to understand their relative impacts and failure modes. The key feature of… See the full description on the dataset page: https://huggingface.co/datasets/carl213/multi-turn_jailbreak_attack_datasets.motorcycle-accident-driving-datasets
Dataset Summary
The dataset consisted of 2 types of cases; accident and driving while riding a motorcycle. 68 accident cases and 68 driving cases are prepared. 30 fps and 852x480 by default. It might be helpful when you train a model to infer whether a video is a motorcycle crash or not. One thing you should know about is 'driving videos' are not typically motorcycle driving. Most 'driving videos' are dashcams in the car. However, all the videos about accidents are motorcycle… See the full description on the dataset page: https://huggingface.co/datasets/smart-dashcam/motorcycle-accident-driving-datasets.finance-datasets
Finance Datasets
Historical stock and cryptocurrency price data.
Contents
Stocks (5 years of daily OHLCV data)
AAPL - Apple Inc.
GOOGL - Alphabet Inc.
MSFT - Microsoft Corp.
AMZN - Amazon.com Inc.
TSLA - Tesla Inc.
META - Meta Platforms
NVDA - NVIDIA Corp.
AMD - Advanced Micro Devices
INTC - Intel Corp.
NFLX - Netflix Inc.
Cryptocurrencies (full history)
BTC_USD - Bitcoin
ETH_USD - Ethereum
SOL_USD - Solana
ADA_USD - Cardano
DOT_USD - Polkadot… See the full description on the dataset page: https://huggingface.co/datasets/misterdonn/finance-datasets.credit-risk-datasetsimdb-sentiment-app-datasetsDataset-Signal-Peptides
Description
This dataset contains 25693 amino acid sequences and labels on each amino acid.
Protein Format: AA sequence
Splits
traing: 20490
valid: 2569
test: 2634
Related paper
The dataset is from SignalP 6.0 predicts all five types of signal
peptides using protein language models.
Label
Each amino acid has 7 classes:
S (0): Sec/SPI signal peptide | T (1): Tat/SPI or Tat/SPII signal peptide | L (2): Sec/SPII signal peptide |
P (3): Sec/SPIII signal… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Signal-Peptides.MIAF_DomainDetection_Infrastructure_Datasets
MIAF: Domain Detection Infrastructure Datasets
This collection is the standardized evaluation benchmark for MIAF (Modular Infrastructure-Aware Fusion). It provides nine classification datasets derived from four public malicious-domain benchmarks, each paired with a shared 137-feature infrastructure representation.
Overview
We evaluate MIAF across nine classification datasets derived from four public malicious-domain benchmarks:
DomainRadar (Hranický et al.… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/MIAF_DomainDetection_Infrastructure_Datasets.fusion-image-to-latex-datasets
Collects and builds the largest dataset to date from online sources, creating a robust and generalizable dataset. This dataset includes approximately 3.4 million image-text pairs, including both handwritten mathematical expressions (200,330 examples) and printed mathematical expressions (3,237,250 examples). Due to the large dataset and the fact that the same mathematical formula can be represented in different LaTeX string formats in an image, it is easy to cause polymorphic ambiguity. To… See the full description on the dataset page: https://huggingface.co/datasets/hoang-quoc-trung/fusion-image-to-latex-datasets.Dataset-Stability-TAPE
Description
Stability Landscape Prediction is a regression task where each input protein x is mapped to a label y ∈ R measuring the most extreme circumstances in which protein x maintains its fold above a concentration threshold (a proxy for intrinsic stability).
Protein Format: AA sequence
Splits
The dataset is from Evaluating Protein Transfer Learning with TAPE. We follow the original data splits, with the number of training, validation and test set shown below:… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Stability-TAPE.datasets
LogicPoison: Logical Attacks on Graph Retrieval-Augmented Generation
This repository contains the datasets for LogicPoison, a logical poisoning framework for Graph-based Retrieval-Augmented Generation (GraphRAG) systems.
Paper: LogicPoison: Logical Attacks on Graph Retrieval-Augmented Generation
GitHub Repository: Jord8061/logicPoison
Overview
LogicPoison targets the topological integrity of knowledge graphs used in GraphRAG. Instead of injecting false content… See the full description on the dataset page: https://huggingface.co/datasets/Jord8061/datasets.safety_aligned_datasets
Safety Aligned Datasets
A high-fidelity adversarial corpus engineered for alignment research, refusal boundary modeling, and robustness evaluation of Small Language Models.
The Problem This Solves
Fine-tuning a Small Language Model to be safe is not the same as fine-tuning it to understand safety.
Most safety datasets give models clean refusal examples on obvious prompts — and those models fail the moment an adversary wraps a harmful request in a… See the full description on the dataset page: https://huggingface.co/datasets/vvsd-charan/safety_aligned_datasets.datasets-for-JCSEgsparc-datasets
GSpaRC Datasets
Datasets used in the paper "GSpaRC: Gaussian Splatting for Real-time Reconstruction of RF Channels".
📄 Paper: arXiv:2511.22793
🌐 Project website: https://nbhavyasai.github.io/GSpaRC/
💻 Code: https://github.com/Nbhavyasai/GSpaRC-WirelessGaussianSplatting
We evaluate GSpaRC on three RF datasets. Only the Sionna conference-room dataset is hosted in this repository (it was generated by us). The RFID and Argos datasets are publicly available from their original… See the full description on the dataset page: https://huggingface.co/datasets/BhavyaN/gsparc-datasets.consciousness-datasets
Consciousness Research Dataset (v1)
This dataset packages structured experiment outputs from the Consciousness project into a reusable format for analysis, comparison, and reporting.
It is designed for:
Cross-method benchmarking
Meta-analysis of metrics across experiments
Reproducible reporting workflows
What Is Included
results_csv/ (25 files): primary tabular outputs from experiment/report pipelines.
results_json/ (2 files): run-level structured summaries/manifests.… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/consciousness-datasets.docling-nlp-datasetsThis repository contains the models used for docling-nlp.
Contents
This model repository packages the pretrained assets used by Docling’s NLP
components:
CRF models for material classification and English part-of-speech tagging
fastText models for language detection, metadata, semantic, topic, and person-name classification
Regular-expression assets for geographic-location extraction and unit handling
A default tokenizer model
Correct workflow to add new files… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/docling-nlp-datasets.Dataset-Structure_Class-ProteinShake
Description
Structural Class Prediction is a multi-class classification task to predict the correct structural class of given protein. This task is is built on the SCOP database.
Protein Format: SA sequence (PDB)
Splits
The dataset is from ProteinShake Building datasets and benchmarks for deep learning on protein structures. We use the splits based on 70% structure similarity, with the number of training, validation and test set shown below:
Train: 7990
Valid: 955
Test:… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Structure_Class-ProteinShake.
