Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pjpjq /blofin-oi-data6 likes659k downloads5mo agoHugging Face02jat-project /jat-dataset-tokenized Dataset Card for "jat-dataset-tokenized" More Information needed timeseries10M<n<100M32 likes351k downloads3y agoHugging Face03google-research-datasets /mbpp Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.text1K<n<10K260 likes301k downloads3y agoHugging Face04jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes287k downloads3y agoHugging Face05ACERobotics /ACE-Data-0 ACE-Data-0 Human-Centric Ambient Capture as Embodied Data Engine S-Lab, Nanyang Technological University, Singapore &nbsp;·&nbsp; ACE Robotics ACE turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI. ▶ Demo video &nbsp;·&nbsp; Full story, figures, and interactive examples on the blog What this is Learning to act in the physical… See the full description on the dataset page: https://huggingface.co/datasets/ACERobotics/ACE-Data-0.videorobotics10K<n<100K48 likes252k downloads10m agoHugging Face06wegrthj /kbcpjv-qi9l-data10 likes220k downloads4mo agoHugging Face07applied-ai-018 /peacock-data-public-datasets-idc0 likes220k downloads2y agoHugging Face08huggingface /DEH-image-scan-datan<1K22 likes214k downloads12d agoHugging Face09GokuScraper /seedance-2-prompts-datasets 🎞️ Seedance-2-prompts-datasets 🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators. This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.imagetext-to-video1K<n<10K46 likes193k downloads15h agoHugging Face10wei82 /precancer-omics-data Precancer → Tumor → Late/Metastatic Progression Multi-omics A curated, continuously-harvested collection of publicly available human multi-omics datasets spanning the full tumor trajectory: precancerous lesions → early carcinoma → advanced / metastatic. Single-cell and spatial transcriptomics are prioritized. ⚠️ Provenance & licensing. Every dataset here was downloaded from a public, open-access repository (no controlled-access or patient-identifiable data). Each dataset… See the full description on the dataset page: https://huggingface.co/datasets/wei82/precancer-omics-data.n>1T10 likes193k downloads24d agoHugging Face11boltzgen /inference-data0 likes186k downloads1y agoHugging Face12Hoshipu /roboreal_data8 likes181k downloads6mo agoHugging Face13wegrthj /e94fjt-v654-data9 likes174k downloads4mo agoHugging Face14meloqiao /us-stock-datatabular100K<n<1M0 likes172k downloads3d agoHugging Face15mvp-lab /LLaVA-OneVision-2-Data LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training. At a Glance The dataset is split across two Hugging Face repositories because of its size: Repository What it contains Part 1 (this repository) ~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.imagevideo-text-to-textn<1K41 likes165k downloads1mo agoHugging Face16google-research-datasets /paws Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling Dataset Summary PAWS: Paraphrase Adversaries from Word Scrambling This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.texttext-classification100K<n<1M40 likes136k downloads3y agoHugging Face17orionweller /generic_data_v20 likes129k downloads2y agoHugging Face18eddmpython /dartlab-data DartLab 데이터 종목코드 하나로 읽는 한국 DART + 미국 SEC EDGAR 공시 데이터 Structured Korean (DART) and US (SEC EDGAR) disclosure data, ready as Parquet. 무엇인가요? DartLab이 한국 DART 전자공시와 미국 SEC EDGAR 공시를 종목코드 하나로 비교 가능한 표로 가공해 Parquet으로 올려둔 데이터셋입니다. 한국 전 상장사(약 2,700사)와 미국 주요 상장사(약 1,000사)의 재무제표, 사업보고서 본문, 정형 공시, 주가, 거시지표가 들어 있습니다. 이 데이터셋은 DartLab의 데이터 층입니다. dartlab.Company("005930")을 호출하면 라이브러리가 필요한 parquet을 여기서 자동으로 내려받습니다. 숫자는 원문 그대로 보존합니다(반올림·추정·보간 없음). 코드 없이도 바로 씁니다… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/dartlab-data.table-question-answering1M<n<10M16 likes129k downloads1h agoHugging Face19defeatbeta /yahoo-finance-data The Financial data from Yahoo! *** Key Points to Note *** All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes. I will update the data regularly, and you are welcome to follow this project and use the data. Each time the data is updated, I will record the update time in spec.json. Data Usage Instructions Use DuckDB… See the full description on the dataset page: https://huggingface.co/datasets/defeatbeta/yahoo-finance-data.100M<n<1B133 likes128k downloads17h agoHugging Face20IPEC-COMMUNITY /fractal20220817_data_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "google_robot", "total_episodes": 87212, "total_frames": 3786400, "total_tasks": 599, "total_videos": 87212, "total_chunks": 88, "chunks_size": 1000, "fps": 3, "splits": { "train": "0:87212" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fractal20220817_data_lerobot.videorobotics13 likes128k downloads2y agoHugging Face21challenge-2026 /challenge_data PrimeBot Household Bimanual Manipulation Challenge Dataset 中文 | English 中文 目录 关于我们 更新日志 真机遥操作数据 训练集说明 验证集说明 数据集字段说明 URDF 图像 语言指令 本体感知与动作 机器人推理接口 UMI数据 数据概览 目录结构 数据集字段说明 图像 本体感知与动作 索引字段 标注与 IMU 关于我们 我们来自上纬新材-启元研究院,我们的使命是加速个人机器人时代到来,加速家用机器人时代到来。我们开源高质量面向家庭操作的双臂操作数据集,同时开放机器人硬件描述以供可视化、可复现研究。 如果本数据集对您的工作有帮助,感谢引用: @misc{xu2026scalingbimanualhouseholdmanipulation, title={Scaling Bimanual Household Manipulation from 1,500… See the full description on the dataset page: https://huggingface.co/datasets/challenge-2026/challenge_data.12 likes126k downloads13d agoHugging Face22vincewin /CREST_data CREST forcing (parquet) EF5/CREST hourly forcing for CONUS, 2016-present, packed as one tar per variable/year. dir variable source cadence mrms/ precipitation MRMS QPE (corrected) hourly temp/ 2 m temperature NLDAS-2 FORA hourly pet/ potential ET FEWS NET daily PET daily Each *.tar expands to individual .pqf (Apache Arrow parquet) grids readable by the EF5 v4.5 native parquet reader. Used by the Space vincewin/CREST_AI. Download + extract one year, e.g.: from… See the full description on the dataset page: https://huggingface.co/datasets/vincewin/CREST_data.5 likes122k downloads3m agoHugging Face23Salesforce /lotsa_data LOTSA Data The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for time series forecasting. It was collected for the purpose of pre-training Large Time Series Models. See the paper and codebase for more information. Citation If you're using LOTSA data in your research or applications, please cite it using this BibTeX: BibTeX: @article{woo2024unified, title={Unified Training of Universal Time Series Forecasting Transformers}… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lotsa_data.text1M<n<10M97 likes117k downloads2y agoHugging Face24hf-internal-testing /dataset_with_scriptThis is a test dataset.textn<1K0 likes116k downloads2y agoHugging Face25llamafactory /demo_data 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh 91 examples for identity learning 300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0 6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.texttext-generation1K<n<10K1 likes111k downloads2y agoHugging Face26Goku-OpenLab /gpt-image-2-prompts-datasets 🖼️ GPT Image 2 Prompt Dataset 🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset. Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.imagetext-to-image10K<n<100K6 likes100k downloads15h agoHugging Face27PhelpsYT /steam-database1 likes100k downloads5mo agoHugging Face28mvp-lab /LLaVA-OneVision-2-Data-Part23 likes100k downloads1mo agoHugging Face29mlfoundations /datacomp_pools DataComp Pools This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.image19 likes95k downloads3y agoHugging Face30CoderOfCode /ship-tracking-data2 likes95k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.