Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01datasets-maintainers /dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI textn<1K0 likes19k downloads3y agoHugging Face02datasets-examples /doc-formats-csv-1 [doc] formats - csv - 1 This dataset contains one csv file at the root: data.csv kind,sound dog,woof cat,meow pokemon,pika human,hello The YAML section of the README does not contain anything related to loading the data (only the size category metadata): --- size_categories: - n<1K --- textn<1K0 likes3.2k downloads3y agoHugging Face03INV-WZQ /ReactiveGWM-Datasets ReactiveGWM-Datasets: Strategy-Aligned Rollouts for Reactive Game World Models 📚 Datasets-Introduction ReactiveGWM-Datasets is the strategy-aligned training corpus that powers ReactiveGWM, a game world model that decouples player control from NPC autonomy. To learn that decoupling, the model needs supervision that pairs each gameplay clip with both a per-frame action stream (what the player did) and a high-level NPC description (what the NPC tried to do, and under… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-Datasets.textimage-to-video10K<n<100K8 likes2.9k downloads5mo agoHugging Face04uclanlp /OpenVLHarness-Evaluation-Datasets OpenVLHarness evaluation datasets Processed evaluation splits used by OpenVLHarness (project page). Each <split>.tsv holds the exact prompts (question) and annotations (answer plus metadata) we evaluate on; image_path is relative to images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13 configs and COCO-format val/test annotations used for ODinW AP evaluation. You normally don't need to download anything by hand: running openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.imagevisual-question-answering10K<n<100K1 likes1.6k downloads1d agoHugging Face05THUIAR /MMLA-Datasets Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark 1. Introduction MMLA is the first comprehensive multimodal language analysis benchmark for evaluating foundation models. It has the following features: Large Scale: 61K+ multimodal samples. Various Sources: 9 datasets. Three Modalities: text, video, and audio Both Acting and Real-world Scenarios: films, TV series, YouTube, Vimeo, Bilibili, TED, improvised scripts, etc. Six Core… See the full description on the dataset page: https://huggingface.co/datasets/THUIAR/MMLA-Datasets.textzero-shot-classification10K<n<100K4 likes1.2k downloads1y agoHugging Face06sktime /tsf-datasetstabular100K<n<1M0 likes676 downloads1y agoHugging Face07scientific-intelligent-modelling /sim-datasets SIM-Datasets: A Unified Symbolic Regression Benchmark A standardized benchmark collection designed for the Scientific Intelligent Modelling (SIM) toolkit, providing comprehensive datasets for symbolic regression research and applications. Overview SIM-Datasets serves as a unified benchmark for symbolic regression tasks, offering standardized datasets with consistent formatting and evaluation protocols. This collection is specifically curated to support the Scientific… See the full description on the dataset page: https://huggingface.co/datasets/scientific-intelligent-modelling/sim-datasets.tabular10M<n<100M0 likes545 downloads1y agoHugging Face08SaProtHub /Dataset-Solubility Description This dataset contains 71419 amino acid sequences and its solubility label. Protein Format: AA sequence Splits traing: 62478 valid: 6942 test: 1999 Related paper The dataset is from DeepSol: a deep learning framework for sequence-based protein solubility prediction. Label Binary label, 1 means soluble, 0 means insoluble. text10K<n<100K2 likes419 downloads2y agoHugging Face09p11-p11 /chess_datasets Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/p11-p11/chess_datasets.texttext-generation1M<n<10M0 likes370 downloads2y agoHugging Face10rodrigoaraujorosa /detector-clickbait-br-datasets Detector Clickbait BR - Datasets Este repositório contém os datasets utilizados para o treinamento do modelo detector-clickbait-br-model, um classificador de textos em português brasileiro capaz de identificar títulos clickbait. 📚 Descrição dos Datasets 1. detector-clickbait-br-raw.csv Dataset original contendo os dados iniciais sem processamento. Características: Dados brutos coletados originalmente Pode conter duplicatas Pode conter valores nulos Formato:… See the full description on the dataset page: https://huggingface.co/datasets/rodrigoaraujorosa/detector-clickbait-br-datasets.tabulartext-classification10K<n<100K2 likes229 downloads10mo agoHugging Face11srescalli /X-PAIR_datasets X-PAIR datasets This repository contains the processed datasets used to train and evaluate X-PAIR, an ultrafast multitask framework for protein–protein interaction (PPI) and partner-specific interface prediction from protein sequences. The datasets are provided in the train, validation, and test splits used in the experiments reported in the X-PAIR study. Datasets The repository contains two groups of datasets: Datasets generated in the X-PAIR study Processed… See the full description on the dataset page: https://huggingface.co/datasets/srescalli/X-PAIR_datasets.text100K<n<1M0 likes226 downloads1mo agoHugging Face12vectara /hhem_leaderboard_datasetstext1K<n<10K0 likes222 downloads1y agoHugging Face13traversaal-ai-hackathon /hotel_datasetsimage1K<n<10K3 likes200 downloads3y agoHugging Face14INV-WZQ /ReactiveGWM-v2-Datasets ReactiveGWM v2 Datasets This repository contains the selected HNM and Street Fighter III: New Generation (SF3) datasets. The original FFV1/MKV videos, first frames, masks, native annotations, action records, indices, and prepared VAE/T5 caches are stored without compression in tar volumes of approximately 2 GiB. Files retain their original formats and paths inside the tar archives. Game-generation code, ROMs, BIOS files, and runtime binaries are not included. Subset Train… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-v2-Datasets.tabularimage-to-video10K<n<100K0 likes191 downloads12d agoHugging Face15carl213 /multi-turn_jailbreak_attack_datasets Multi-Turn Jailbreak Attack Datasets Description This dataset was created to compare single-turn and multi-turn jailbreak attacks on large language models (LLMs). The primary goal is to take a single harmful prompt and distribute the harm over multiple turns, making each prompt appear harmless in isolation. This approach is compared against traditional single-turn attacks with the complete prompt to understand their relative impacts and failure modes. The key feature of… See the full description on the dataset page: https://huggingface.co/datasets/carl213/multi-turn_jailbreak_attack_datasets.text1K<n<10K1 likes190 downloads7mo agoHugging Face16smart-dashcam /motorcycle-accident-driving-datasets Dataset Summary The dataset consisted of 2 types of cases; accident and driving while riding a motorcycle. 68 accident cases and 68 driving cases are prepared. 30 fps and 852x480 by default. It might be helpful when you train a model to infer whether a video is a motorcycle crash or not. One thing you should know about is 'driving videos' are not typically motorcycle driving. Most 'driving videos' are dashcams in the car. However, all the videos about accidents are motorcycle… See the full description on the dataset page: https://huggingface.co/datasets/smart-dashcam/motorcycle-accident-driving-datasets.textvideo-classificationn<1K0 likes171 downloads3y agoHugging Face17misterdonn /finance-datasets Finance Datasets Historical stock and cryptocurrency price data. Contents Stocks (5 years of daily OHLCV data) AAPL - Apple Inc. GOOGL - Alphabet Inc. MSFT - Microsoft Corp. AMZN - Amazon.com Inc. TSLA - Tesla Inc. META - Meta Platforms NVDA - NVIDIA Corp. AMD - Advanced Micro Devices INTC - Intel Corp. NFLX - Netflix Inc. Cryptocurrencies (full history) BTC_USD - Bitcoin ETH_USD - Ethereum SOL_USD - Solana ADA_USD - Cardano DOT_USD - Polkadot… See the full description on the dataset page: https://huggingface.co/datasets/misterdonn/finance-datasets.tabular10K<n<100K1 likes161 downloads7mo agoHugging Face18RPD123-byte /credit-risk-datasetstabular100K<n<1M1 likes122 downloads8mo agoHugging Face19Maruf39237 /imdb-sentiment-app-datasetstabular10K<n<100K0 likes103 downloads14d agoHugging Face20SaProtHub /Dataset-Signal-Peptides Description This dataset contains 25693 amino acid sequences and labels on each amino acid. Protein Format: AA sequence Splits traing: 20490 valid: 2569 test: 2634 Related paper The dataset is from SignalP 6.0 predicts all five types of signal peptides using protein language models. Label Each amino acid has 7 classes: S (0): Sec/SPI signal peptide | T (1): Tat/SPI or Tat/SPII signal peptide | L (2): Sec/SPII signal peptide | P (3): Sec/SPIII signal… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Signal-Peptides.text10K<n<100K2 likes102 downloads2y agoHugging Face21ahmedBargady /MIAF_DomainDetection_Infrastructure_Datasets MIAF: Domain Detection Infrastructure Datasets This collection is the standardized evaluation benchmark for MIAF (Modular Infrastructure-Aware Fusion). It provides nine classification datasets derived from four public malicious-domain benchmarks, each paired with a shared 137-feature infrastructure representation. Overview We evaluate MIAF across nine classification datasets derived from four public malicious-domain benchmarks: DomainRadar (Hranický et al.… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/MIAF_DomainDetection_Infrastructure_Datasets.tabulartabular-classification1M<n<10M0 likes99 downloads27d agoHugging Face22hoang-quoc-trung /fusion-image-to-latex-datasets Collects and builds the largest dataset to date from online sources, creating a robust and generalizable dataset. This dataset includes approximately 3.4 million image-text pairs, including both handwritten mathematical expressions (200,330 examples) and printed mathematical expressions (3,237,250 examples). Due to the large dataset and the fact that the same mathematical formula can be represented in different LaTeX string formats in an image, it is easy to cause polymorphic ambiguity. To… See the full description on the dataset page: https://huggingface.co/datasets/hoang-quoc-trung/fusion-image-to-latex-datasets.text1M<n<10M18 likes89 downloads2y agoHugging Face23SaProtHub /Dataset-Stability-TAPE Description Stability Landscape Prediction is a regression task where each input protein x is mapped to a label y ∈ R measuring the most extreme circumstances in which protein x maintains its fold above a concentration threshold (a proxy for intrinsic stability). Protein Format: AA sequence Splits The dataset is from Evaluating Protein Transfer Learning with TAPE. We follow the original data splits, with the number of training, validation and test set shown below:… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Stability-TAPE.text10K<n<100K1 likes77 downloads2y agoHugging Face24Jord8061 /datasets LogicPoison: Logical Attacks on Graph Retrieval-Augmented Generation This repository contains the datasets for LogicPoison, a logical poisoning framework for Graph-based Retrieval-Augmented Generation (GraphRAG) systems. Paper: LogicPoison: Logical Attacks on Graph Retrieval-Augmented Generation GitHub Repository: Jord8061/logicPoison Overview LogicPoison targets the topological integrity of knowledge graphs used in GraphRAG. Instead of injecting false content… See the full description on the dataset page: https://huggingface.co/datasets/Jord8061/datasets.tabularquestion-answering1K<n<10K2 likes75 downloads3mo agoHugging Face25vvsd-charan /safety_aligned_datasets Safety Aligned Datasets A high-fidelity adversarial corpus engineered for alignment research, refusal boundary modeling, and robustness evaluation of Small Language Models. The Problem This Solves Fine-tuning a Small Language Model to be safe is not the same as fine-tuning it to understand safety. Most safety datasets give models clean refusal examples on obvious prompts — and those models fail the moment an adversary wraps a harmful request in a… See the full description on the dataset page: https://huggingface.co/datasets/vvsd-charan/safety_aligned_datasets.texttext-generation10K<n<100K0 likes72 downloads4mo agoHugging Face26MU-Kindai /datasets-for-JCSEtext100K<n<1M0 likes70 downloads4y agoHugging Face27BhavyaN /gsparc-datasets GSpaRC Datasets Datasets used in the paper "GSpaRC: Gaussian Splatting for Real-time Reconstruction of RF Channels". 📄 Paper: arXiv:2511.22793 🌐 Project website: https://nbhavyasai.github.io/GSpaRC/ 💻 Code: https://github.com/Nbhavyasai/GSpaRC-WirelessGaussianSplatting We evaluate GSpaRC on three RF datasets. Only the Sionna conference-room dataset is hosted in this repository (it was generated by us). The RFID and Argos datasets are publicly available from their original… See the full description on the dataset page: https://huggingface.co/datasets/BhavyaN/gsparc-datasets.tabularother1K<n<10K0 likes67 downloads5mo agoHugging Face28tahamajs /consciousness-datasets Consciousness Research Dataset (v1) This dataset packages structured experiment outputs from the Consciousness project into a reusable format for analysis, comparison, and reporting. It is designed for: Cross-method benchmarking Meta-analysis of metrics across experiments Reproducible reporting workflows What Is Included results_csv/ (25 files): primary tabular outputs from experiment/report pipelines. results_json/ (2 files): run-level structured summaries/manifests.… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/consciousness-datasets.tabulartext-classificationn<1K2 likes66 downloads8mo agoHugging Face29docling-project /docling-nlp-datasetsThis repository contains the models used for docling-nlp. Contents This model repository packages the pretrained assets used by Docling’s NLP components: CRF models for material classification and English part-of-speech tagging fastText models for language detection, metadata, semantic, topic, and person-name classification Regular-expression assets for geographic-location extraction and unit handling A default tokenizer model Correct workflow to add new files… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/docling-nlp-datasets.tabular100K<n<1M0 likes65 downloads29d agoHugging Face30SaProtHub /Dataset-Structure_Class-ProteinShake Description Structural Class Prediction is a multi-class classification task to predict the correct structural class of given protein. This task is is built on the SCOP database. Protein Format: SA sequence (PDB) Splits The dataset is from ProteinShake Building datasets and benchmarks for deep learning on protein structures. We use the splits based on 70% structure similarity, with the number of training, validation and test set shown below: Train: 7990 Valid: 955 Test:… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Structure_Class-ProteinShake.text1K<n<10K0 likes63 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.