Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pjpjq /blofin-oi-data7 likes736k downloads5mo agoHugging Face02wei82 /precancer-omics-data Precancer → Tumor → Late/Metastatic Progression Multi-omics A curated, continuously-harvested collection of publicly available human multi-omics datasets spanning the full tumor trajectory: precancerous lesions → early carcinoma → advanced / metastatic. Single-cell and spatial transcriptomics are prioritized. ⚠️ Provenance & licensing. Every dataset here was downloaded from a public, open-access repository (no controlled-access or patient-identifiable data). Each dataset… See the full description on the dataset page: https://huggingface.co/datasets/wei82/precancer-omics-data.n>1T10 likes320k downloads27d agoHugging Face03google-research-datasets /mbpp Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.text1K<n<10K260 likes268k downloads3y agoHugging Face04applied-ai-018 /peacock-data-public-datasets-idc0 likes250k downloads2y agoHugging Face05ACERobotics /ACE-Data-0 ACE-Data-0 Human-Centric Ambient Capture as Embodied Data Engine S-Lab, Nanyang Technological University, Singapore &nbsp;·&nbsp; ACE Robotics ACE turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI. ▶ Demo video &nbsp;·&nbsp; Full story, figures, and interactive examples on the blog What this is Learning to act in the physical… See the full description on the dataset page: https://huggingface.co/datasets/ACERobotics/ACE-Data-0.videorobotics10K<n<100K48 likes243k downloads1h agoHugging Face06wegrthj /kbcpjv-qi9l-data10 likes223k downloads4mo agoHugging Face07huggingface /DEH-image-scan-datan<1K22 likes208k downloads15d agoHugging Face08jat-project /jat-dataset-tokenized Dataset Card for "jat-dataset-tokenized" More Information needed timeseries10M<n<100M32 likes205k downloads3y agoHugging Face09jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes202k downloads3y agoHugging Face10GokuScraper /seedance-2-prompts-datasets 🎞️ Seedance-2-prompts-datasets 🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators. This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.imagetext-to-video1K<n<10K48 likes199k downloads4d agoHugging Face11boltzgen /inference-data0 likes198k downloads1y agoHugging Face12orionweller /generic_data_v20 likes179k downloads2y agoHugging Face13wegrthj /e94fjt-v654-data9 likes179k downloads4mo agoHugging Face14meloqiao /us-stock-datatabular100K<n<1M0 likes173k downloads6d agoHugging Face15Hoshipu /roboreal_data8 likes164k downloads6mo agoHugging Face16mvp-lab /LLaVA-OneVision-2-Data LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training. At a Glance The dataset is split across two Hugging Face repositories because of its size: Repository What it contains Part 1 (this repository) ~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.imagevideo-text-to-textn<1K42 likes156k downloads1mo agoHugging Face17eddmpython /dartlab-data DartLab 데이터 종목코드 하나로 읽는 한국 DART + 미국 SEC EDGAR 공시 데이터 Structured Korean (DART) and US (SEC EDGAR) disclosure data, ready as Parquet. 무엇인가요? DartLab이 한국 DART 전자공시와 미국 SEC EDGAR 공시를 종목코드 하나로 비교 가능한 표로 가공해 Parquet으로 올려둔 데이터셋입니다. 한국 전 상장사(약 2,700사)와 미국 주요 상장사(약 1,000사)의 재무제표, 사업보고서 본문, 정형 공시, 주가, 거시지표가 들어 있습니다. 이 데이터셋은 DartLab의 데이터 층입니다. dartlab.Company("005930")을 호출하면 라이브러리가 필요한 parquet을 여기서 자동으로 내려받습니다. 숫자는 원문 그대로 보존합니다(반올림·추정·보간 없음). 코드 없이도 바로 씁니다… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/dartlab-data.table-question-answering1M<n<10M16 likes154k downloads4m agoHugging Face18google-research-datasets /paws Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling Dataset Summary PAWS: Paraphrase Adversaries from Word Scrambling This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.texttext-classification100K<n<1M41 likes134k downloads3y agoHugging Face19defeatbeta /yahoo-finance-data The Financial data from Yahoo! *** Key Points to Note *** All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes. I will update the data regularly, and you are welcome to follow this project and use the data. Each time the data is updated, I will record the update time in spec.json. Data Usage Instructions Use DuckDB… See the full description on the dataset page: https://huggingface.co/datasets/defeatbeta/yahoo-finance-data.100M<n<1B133 likes129k downloads7h agoHugging Face20bk939448 /System_arc_data0 likes124k downloads6mo agoHugging Face21vincewin /CREST_data CREST forcing (parquet) EF5/CREST hourly forcing for CONUS, 2016-present, packed as one tar per variable/year. dir variable source cadence mrms/ precipitation MRMS QPE (corrected) hourly temp/ 2 m temperature NLDAS-2 FORA hourly pet/ potential ET FEWS NET daily PET daily Each *.tar expands to individual .pqf (Apache Arrow parquet) grids readable by the EF5 v4.5 native parquet reader. Used by the Space vincewin/CREST_AI. Download + extract one year, e.g.: from… See the full description on the dataset page: https://huggingface.co/datasets/vincewin/CREST_data.image100M<n<1B5 likes121k downloads16m agoHugging Face22Goku-OpenLab /gpt-image-2-prompts-datasets 🖼️ GPT Image 2 Prompt Dataset 🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset. Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.imagetext-to-image10K<n<100K8 likes111k downloads4d agoHugging Face23evaleval /EEE_datastore Every Eval Ever Datastore A community database of AI evaluation results, all in one schema. Scores scraped from leaderboards, pulled out of papers, and produced by local evaluation runs are stored in a single record format, so results from different sources can be compared, joined, and reused instead of re-scraped. This dataset is the data itself: one JSON record per model per evaluation run — which may carry several scored results — with optional per-sample companion files.… See the full description on the dataset page: https://huggingface.co/datasets/evaleval/EEE_datastore.textother1K<n<10K42 likes111k downloads2h agoHugging Face24hf-internal-testing /dataset_with_scriptThis is a test dataset.textn<1K0 likes111k downloads2y agoHugging Face25PhelpsYT /steam-database1 likes103k downloads5mo agoHugging Face26llamafactory /demo_data 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh 91 examples for identity learning 300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0 6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.texttext-generation1K<n<10K1 likes99k downloads2y agoHugging Face27Salesforce /lotsa_data LOTSA Data The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for time series forecasting. It was collected for the purpose of pre-training Large Time Series Models. See the paper and codebase for more information. Citation If you're using LOTSA data in your research or applications, please cite it using this BibTeX: BibTeX: @article{woo2024unified, title={Unified Training of Universal Time Series Forecasting Transformers}… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lotsa_data.text1M<n<10M97 likes98k downloads2y agoHugging Face28mlfoundations /datacomp_pools DataComp Pools This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.image19 likes97k downloads3y agoHugging Face29CoderOfCode /ship-tracking-data2 likes96k downloads3mo agoHugging Face30mvp-lab /LLaVA-OneVision-2-Data-Part23 likes95k downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.