Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01xxxspatialencoderwds4 /data_4 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds4/data_4.tabularobject-detectionn<1K1 likes8.7k downloads20d agoHugging Face02soma114 /soma-competition-datasettabularn<1K0 likes8.5k downloads11d agoHugging Face03xxxspatialencoderwds3 /data_3 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds3/data_3.tabularobject-detectionn<1K0 likes6.9k downloads20d agoHugging Face04xxxspatialencoderwds2 /data_2 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds2/data_2.tabularobject-detectionn<1K5 likes6.2k downloads20d agoHugging Face05xxxspatialencoderwds1 /data_1 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds1/data_1.tabularobject-detectionn<1K0 likes6k downloads20d agoHugging Face06hysts-bot-data /daily-papers-stats Daily Papers Stats Upvote and comment counts for every paper in hysts-bot-data/daily-papers. Join the two on arxiv_id. data.json has arxiv_id, upvotes and num_comments, and only holds the values at the time it was last updated. It is overwritten roughly every hour, so past values are available from the commit history (from 2024-03-12 on, with some gaps). Early revisions do not have num_comments. License CC0 1.0. tabular10K<n<100K3 likes5.2k downloads48m agoHugging Face07johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M4 likes5k downloads8mo agoHugging Face08Qalam /nuclear-intelligence-dataset Nuclear Intelligence Dataset Public, auto-generated dataset of validated nuclear-energy research cycles. Latest stats (auto-updated): 🪙 NES tokens minted: 0 ⛓️ Blockchain length: 1 blocks 🕸️ Knowledge entities: 2 Source GitHub: https://github.com/QalamHipHop/nuclear-intelligence HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence License MIT tabularquestion-answeringn<1K1 likes4.6k downloads2h agoHugging Face09YWZBrandon /officeqa-checkpoint-eval-data Checkpoint evaluation plot data Snapshot: 2026-09-14T16:26:45.684890+00:00. Aggregate inputs to notes/Sept-2-2026.md performance figures. No model execution, grading, publication, or source-result changes were performed to make this export. Contents checkpoint_evaluations: 454 checkpoint rows, one evaluation per run/iteration/protocol; score, mean output tokens, mean steps, and the existing two-sided 95% confidence bounds. pareto_points: current mean-token/USD… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/officeqa-checkpoint-eval-data.tabular10K<n<100K0 likes3.9k downloads26d agoHugging Face10mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes3.5k downloads1mo agoHugging Face11datania /ine-catalog INE Este repositorio contiene todas las tablas¹ del Instituto Nacional de Estadística exportadas a ficheros Parquet. Puedes encontrar cualquiera de las tablas o sus metadatos en la carpeta tablas. Cada tabla está identificado un una ID. Puedes encontrar la ID de la tabla tanto en el INE (es el número que aparece en la URL) or en el archivo tablas.jsonl de este repositorio que puedes explorar en el Data Viewer. Por ejemplo, la tabla de Índices nacionales de clases se corresponde… See the full description on the dataset page: https://huggingface.co/datasets/datania/ine-catalog.tabular1K<n<10K5 likes3k downloads18d agoHugging Face12ZombitX64 /xauusd-gold-price-historical-data-2004-2025 XAUUSD Gold Price Historical Data 2004-2025 This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025. Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024" Content: The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns: Date Open High Low Close Volume Usage: This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.tabular1M<n<10M12 likes2.8k downloads1y agoHugging Face13latticecx /Lattice_4D_Datasetgated Lattice_4D_Dataset Multi-camera volumetric captures of people doing everyday tasks (ball handling, shirt folding). Four RGB-D cameras record each take. Each take has the reconstructed 3-D scene of every frame and the fitted body and hands. It also has calibration, camera poses, action labels over frame spans, reviewed language and rendered orbit videos. A USD skeleton and a URDF rig let a robotics consumer load the body. Ball handling: the person drops the ball to bounce off the… See the full description on the dataset page: https://huggingface.co/datasets/latticecx/Lattice_4D_Dataset.tabularrobotics1K<n<10K4 likes2.6k downloads2d agoHugging Face14minkyuchoi /Temporal-Logic-Video-Dataset Temporal Logic Video (TLV) Dataset Temporal Logic Video (TLV) Dataset Synthetic and real video dataset with temporal logic annotation Explore the GitHub » NSVS-TL Project Webpage · NSVS-TL Source Code Overview The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components: Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.tabularquestion-answeringn<1K1 likes2.5k downloads2y agoHugging Face15taesiri /imagenet_hard_review_data_r2tabular1K<n<10K0 likes2.2k downloads3y agoHugging Face16Flotak /trading-datatabularn<1K0 likes1.8k downloads8h agoHugging Face17bcb-instruct /bcb_datatabular10M<n<100M0 likes1.6k downloads1y agoHugging Face18HKUSTDial /DataSpace DataSpace Paper · Code · Leaderboard · KDD Cup 2026 DataSpace is a benchmark for data agents that perform verifiable analytics over heterogeneous, task-local workspaces. Each task provides a natural-language question and a workspace containing a combination of CSV, JSON, SQLite, Markdown, PDF, and video artifacts. The required output is a complete tabular result. DataSpace is also the official benchmark of the KDD Cup 2026 Data Agent Track. Release policy This… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTDial/DataSpace.documenttable-question-answeringn<1K7 likes1.4k downloads2mo agoHugging Face19AILab-CVC /SEED-Data-Edit-Part1-Openimages SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.tabulartext-to-image1M<n<10M10 likes1.1k downloads2y agoHugging Face20Congliu /Chinese-DeepSeek-R1-Distill-data-110k 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face&nbsp;&nbsp; | &nbsp;&nbsp;🤖 ModelScope &nbsp;&nbsp; | &nbsp;&nbsp;🚀 Github &nbsp;&nbsp; | &nbsp;&nbsp;📑 Blog 注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。 该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.tabulartext-generation100K<n<1M789 likes1k downloads2y agoHugging Face21rayrren /agent-apprenticeship-seed-dataset Agent Apprenticeship Seed Dataset The living ecosystem where AI agents run automated workflow loops on any task, improve through execution, and turn each run into reusable work experience + data to improve future agents. As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where real-world tasks generate reusable learning signals and complex workflows advance through agent loops that turn execution into shared… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset.tabular1K<n<10K0 likes1k downloads4mo agoHugging Face22nvidia /Nemotron-Cascade-2-RL-data Dataset Description: The Nemotron-Cascade-2-RL dataset is a curated reinforcement learning (RL) dataset blend used to train Nemotron-Cascade-2-30B-A3B model. It includes instruction-following RL, multi-domain RL, on-policy distillation, and software engineering RL (SWE-RL) data. This dataset is ready for commercial use. The dataset contains the following subset: IF-RL Contains 45,879 training samples for instruction-following RL. Our curation process mainly… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-RL-data.tabular10K<n<100K54 likes996 downloads7mo agoHugging Face23RUC-DataLab /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/DataScience-Instruct-500K.tabular10K<n<100K76 likes965 downloads1y agoHugging Face24XifengZhang /FM4PDE-pde-data FM4PDE PDE Data Datasets for Guided Flow Matching for Forward and Inverse PDE Problems with Sparse Observations: Algorithm and Theory, by Xifeng Zhang and Jin Zhao. The data cover eleven PDE families and are used to train and evaluate FM4PDE for forward and inverse problems with sparse observations. Paper on arXiv · FM4PDE on GitHub · Pretrained models Dataset overview All 89 files are available. The dataset contains 55 training files, 29 test files, and 5… See the full description on the dataset page: https://huggingface.co/datasets/XifengZhang/FM4PDE-pde-data.tabularothern<1K1 likes956 downloads10d agoHugging Face25paulpacaud /rlbenchfail_test_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes892 downloads8mo agoHugging Face26fromthesky /pldr-llm-training-dynamics-data PLDR-LLM Training Dynamics Data Reported numerical evidence for Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction, by Burc Gokden. Monograph: Hugging Face Paper Page. Scientific code and readers: GitHub repository. Numerical evidence: Hugging Face dataset. Book: Power Law Graph Attention and PLDR-LLMs: Mathematical Foundations, Training Dynamics, and Predictive Inference, by Burc Gokden. Book Companion: Code and edition… See the full description on the dataset page: https://huggingface.co/datasets/fromthesky/pldr-llm-training-dynamics-data.tabularothern<1K0 likes868 downloads4d agoHugging Face27RyanLiu112 /a_datatabular100K<n<1M0 likes806 downloads1y agoHugging Face28paulpacaud /rlbenchfail_train_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.tabularvisual-question-answering10K<n<100K0 likes800 downloads8mo agoHugging Face29cx-cmu /repro-organic-data-72Btabular10M<n<100M0 likes782 downloads1y agoHugging Face30liuyueyi-8 /Elastic-Forcing-training-dataset Elastic-Forcing training datasets wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351. Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales. wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,454 subsequently removed and excluded). tabular10K<n<100K1 likes757 downloads12d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.