datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blofin-oi-datajat-dataset-tokenized
Dataset Card for "jat-dataset-tokenized"
More Information needed
mbpp
Dataset Card for Mostly Basic Python Problems (mbpp)
Dataset Summary
The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us.
Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.jat-dataset
JAT Dataset
Dataset Description
The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent.
Paper: https://huggingface.co/papers/2402.09844
Usage
>>> from datasets import load_dataset
>>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.ACE-Data-0
ACE-Data-0
Human-Centric Ambient Capture as Embodied Data Engine
S-Lab, Nanyang Technological University, Singapore
·
ACE Robotics
ACE turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI.
▶ Demo video
·
Full story, figures, and interactive examples on the blog
What this is
Learning to act in the physical… See the full description on the dataset page: https://huggingface.co/datasets/ACERobotics/ACE-Data-0.kbcpjv-qi9l-datapeacock-data-public-datasets-idcDEH-image-scan-dataseedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.precancer-omics-data
Precancer → Tumor → Late/Metastatic Progression Multi-omics
A curated, continuously-harvested collection of publicly available human multi-omics
datasets spanning the full tumor trajectory: precancerous lesions → early carcinoma →
advanced / metastatic. Single-cell and spatial transcriptomics are prioritized.
⚠️ Provenance & licensing. Every dataset here was downloaded from a public,
open-access repository (no controlled-access or patient-identifiable data). Each dataset… See the full description on the dataset page: https://huggingface.co/datasets/wei82/precancer-omics-data.inference-dataroboreal_datae94fjt-v654-dataus-stock-dataLLaVA-OneVision-2-Data
LLaVA-OneVision-2-Data
Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training.
At a Glance
The dataset is split across two Hugging Face repositories because of its size:
Repository
What it contains
Part 1 (this repository)
~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.paws
Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling
Dataset Summary
PAWS: Paraphrase Adversaries from Word Scrambling
This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset.
For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.generic_data_v2dartlab-data
DartLab 데이터
종목코드 하나로 읽는 한국 DART + 미국 SEC EDGAR 공시 데이터
Structured Korean (DART) and US (SEC EDGAR) disclosure data, ready as Parquet.
무엇인가요?
DartLab이 한국 DART 전자공시와 미국 SEC EDGAR 공시를 종목코드 하나로 비교 가능한 표로 가공해 Parquet으로 올려둔 데이터셋입니다.
한국 전 상장사(약 2,700사)와 미국 주요 상장사(약 1,000사)의 재무제표, 사업보고서 본문, 정형 공시, 주가, 거시지표가 들어 있습니다.
이 데이터셋은 DartLab의 데이터 층입니다. dartlab.Company("005930")을 호출하면 라이브러리가 필요한 parquet을 여기서 자동으로 내려받습니다. 숫자는 원문 그대로 보존합니다(반올림·추정·보간 없음).
코드 없이도 바로 씁니다… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/dartlab-data.yahoo-finance-data
The Financial data from Yahoo!
*** Key Points to Note ***
All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes.
I will update the data regularly, and you are welcome to follow this project and use the data.
Each time the data is updated, I will record the update time in spec.json.
Data Usage Instructions
Use DuckDB… See the full description on the dataset page: https://huggingface.co/datasets/defeatbeta/yahoo-finance-data.fractal20220817_data_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "google_robot",
"total_episodes": 87212,
"total_frames": 3786400,
"total_tasks": 599,
"total_videos": 87212,
"total_chunks": 88,
"chunks_size": 1000,
"fps": 3,
"splits": {
"train": "0:87212"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fractal20220817_data_lerobot.challenge_data
PrimeBot Household Bimanual Manipulation Challenge Dataset
中文 | English
中文
目录
关于我们
更新日志
真机遥操作数据
训练集说明
验证集说明
数据集字段说明
URDF
图像
语言指令
本体感知与动作
机器人推理接口
UMI数据
数据概览
目录结构
数据集字段说明
图像
本体感知与动作
索引字段
标注与 IMU
关于我们
我们来自上纬新材-启元研究院,我们的使命是加速个人机器人时代到来,加速家用机器人时代到来。我们开源高质量面向家庭操作的双臂操作数据集,同时开放机器人硬件描述以供可视化、可复现研究。
如果本数据集对您的工作有帮助,感谢引用:
@misc{xu2026scalingbimanualhouseholdmanipulation,
title={Scaling Bimanual Household Manipulation from 1,500… See the full description on the dataset page: https://huggingface.co/datasets/challenge-2026/challenge_data.CREST_data
CREST forcing (parquet)
EF5/CREST hourly forcing for CONUS, 2016-present, packed as one tar per variable/year.
dir
variable
source
cadence
mrms/
precipitation
MRMS QPE (corrected)
hourly
temp/
2 m temperature
NLDAS-2 FORA
hourly
pet/
potential ET
FEWS NET daily PET
daily
Each *.tar expands to individual .pqf (Apache Arrow parquet) grids readable by the
EF5 v4.5 native parquet reader. Used by the Space vincewin/CREST_AI.
Download + extract one year, e.g.:
from… See the full description on the dataset page: https://huggingface.co/datasets/vincewin/CREST_data.lotsa_data
LOTSA Data
The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for time series forecasting.
It was collected for the purpose of pre-training Large Time Series Models.
See the paper and codebase for more information.
Citation
If you're using LOTSA data in your research or applications, please cite it using this BibTeX:
BibTeX:
@article{woo2024unified,
title={Unified Training of Universal Time Series Forecasting Transformers}… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lotsa_data.dataset_with_scriptThis is a test dataset.demo_data
1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en
1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh
300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en
300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh
91 examples for identity learning
300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0
6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.gpt-image-2-prompts-datasets
🖼️ GPT Image 2 Prompt Dataset
🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset.
Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.steam-databaseLLaVA-OneVision-2-Data-Part2datacomp_pools
DataComp Pools
This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.ship-tracking-data
