Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /mbpp Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.text1K<n<10K260 likes268k downloads3y agoHugging Face02jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes202k downloads3y agoHugging Face03mvp-lab /LLaVA-OneVision-2-Data LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training. At a Glance The dataset is split across two Hugging Face repositories because of its size: Repository What it contains Part 1 (this repository) ~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.imagevideo-text-to-textn<1K42 likes156k downloads1mo agoHugging Face04google-research-datasets /paws Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling Dataset Summary PAWS: Paraphrase Adversaries from Word Scrambling This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.texttext-classification100K<n<1M41 likes134k downloads3y agoHugging Face05evaleval /EEE_datastore Every Eval Ever Datastore A community database of AI evaluation results, all in one schema. Scores scraped from leaderboards, pulled out of papers, and produced by local evaluation runs are stored in a single record format, so results from different sources can be compared, joined, and reused instead of re-scraped. This dataset is the data itself: one JSON record per model per evaluation run — which may carry several scored results — with optional per-sample companion files.… See the full description on the dataset page: https://huggingface.co/datasets/evaleval/EEE_datastore.textother1K<n<10K42 likes111k downloads7h agoHugging Face06hf-internal-testing /dataset_with_scriptThis is a test dataset.textn<1K0 likes111k downloads2y agoHugging Face07llamafactory /demo_data 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh 91 examples for identity learning 300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0 6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.texttext-generation1K<n<10K1 likes99k downloads2y agoHugging Face08Salesforce /lotsa_data LOTSA Data The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for time series forecasting. It was collected for the purpose of pre-training Large Time Series Models. See the paper and codebase for more information. Citation If you're using LOTSA data in your research or applications, please cite it using this BibTeX: BibTeX: @article{woo2024unified, title={Unified Training of Universal Time Series Forecasting Transformers}… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lotsa_data.text1M<n<10M97 likes98k downloads2y agoHugging Face09lmarena-ai /leaderboard-dataset Arena Leaderboard Dataset Historical snapshots of the Arena leaderboard. Usage from datasets import load_dataset # Load all historical text style control data ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full") # Load the current text style control leaderboard ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest") # Filter to overall category ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.tabular1M<n<10M28 likes91k downloads14h agoHugging Face10autogluon /fev_datasets Forecast evaluation datasets This repository contains time series datasets that can be used for evaluation of univariate & multivariate forecasting models. The main focus of this repository is on datasets that reflect real-world forecasting scenarios, such as those involving covariates, missing values, and other practical complexities. The datasets follow a format that is compatible with the fev package. Data format and usage Each dataset satisfies the following… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/fev_datasets.tabulartime-series-forecasting100K<n<1M13 likes86k downloads9mo agoHugging Face11tencent /Hy-Embodied-0.5-VLA-Data Hy-Embodied-0.5-VLA From Vision-Language-Action Models to a Real-World Robot Learning Stack Tencent Robotics X × Tencent Hy Team 📖 Abstract We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.tabularroboticsn<1K24 likes69k downloads3mo agoHugging Face12HoneyDataV2 /Honey-Data-V2 Honey-Data-V2 A multimodal supervised fine-tuning corpus of 19,707,852 image groups carrying 44,295,078 conversations, spread over 8 task categories and 420 subsets (5.81 TB). Honey-Data-V2 extends Honey-Data-15M, the corpus behind Bee-8B. The original pool was re-curated under stricter structural rules, re-annotated by an upgraded stack of frontier models that contributes up to three independent answers per instruction, and extended with newly released community corpora.… See the full description on the dataset page: https://huggingface.co/datasets/HoneyDataV2/Honey-Data-V2.imagevisual-question-answering10M<n<100M0 likes66k downloads18d agoHugging Face13allenai /molmobot-data MolmoBot-data Training episode data (actions, visual inputs, and other sensor data) for 8 tasks on 2 robotic platforms: DoorOpeningDataGenConfig RBY1OpenDataGenConfig RBY1PickDataGenConfig FrankaPickOmniCamConfig RBY1PickAndPlaceDataGenConfig FrankaPickAndPlaceOmniCamConfig FrankaPickAndPlaceColorOmniCamConfig FrankaPickAndPlaceNextToOmniCamConfig Please note that every package indexed by the parquet files can contain several instances of episode data. We also provide an… See the full description on the dataset page: https://huggingface.co/datasets/allenai/molmobot-data.tabular100K<n<1M8 likes66k downloads3mo agoHugging Face14mvp-lab /LLaVA-OneVision-1.5-Instruct-Data LLaVA-OneVision-1.5 Instruction Data Paper | Code 📌 Introduction This dataset, LLaVA-OneVision-1.5-Instruct, was collected and integrated during the development of LLaVA-OneVision-1.5. LLaVA-OneVision-1.5 is a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. This meticulously curated 22M instruction dataset (LLaVA-OneVision-1.5-Instruct) is part of a… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Instruct-Data.imageimage-text-to-text10M<n<100M83 likes60k downloads3mo agoHugging Face15hf-internal-testing /dataset_with_data_filestextn<1K0 likes52k downloads2y agoHugging Face16cornell-movie-review-data /rotten_tomatoes Dataset Card for "rotten_tomatoes" Dataset Summary Movie Review Dataset. This is a dataset of containing 5,331 positive and 5,331 negative processed sentences from Rotten Tomatoes movie reviews. This data was first used in Bo Pang and Lillian Lee, ``Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.'', Proceedings of the ACL, 2005. Supported Tasks and Leaderboards More Information Needed Languages… See the full description on the dataset page: https://huggingface.co/datasets/cornell-movie-review-data/rotten_tomatoes.texttext-classification10K<n<100K123 likes52k downloads3y agoHugging Face17hf-internal-testing /multi_dir_datasettextn<1K0 likes51k downloads5y agoHugging Face18klieret /swe-bench-dummy-test-datasettextn<1K0 likes50k downloads1y agoHugging Face19llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes46k downloads2y agoHugging Face20hf-internal-testing /tokenizers-test-data tokenizers-test-data Test and benchmark fixtures for huggingface/tokenizers, pulled on demand by the repo Makefiles (make test / make bench / make fixtures via hf download). Layout fixtures/ — multilingual + modality corpora for cross-language encode benchmarks. Organized, documented, and reproducible: see fixtures/FIXTURES.md for provenance and fixtures/fixtures_manifest.json for exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.textn<1K0 likes45k downloads12d agoHugging Face21zgcagi /ZGCM-1-Datagated A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence 📄 Tech Report · 🤗 Model · 🤗 Data · 📈 Training Log 📊 Results · 💻 Training Code · 💬 WeChat Community Introduction ZGCM-1 is a 7.39B-parameter dense language model trained from scratch, built for mathematical reasoning and tool-assisted search. It combines deliberate internal thinking with active information… See the full description on the dataset page: https://huggingface.co/datasets/zgcagi/ZGCM-1-Data.texttext-generation1B<n<10B88 likes45k downloads14d agoHugging Face22MeiGen-AI /GenEvolve-Data-Bench GenEvolve Data and Bench This repository contains the open-source data release for GenEvolve: Config Directory Records Images Purpose sft GenEvolve-Data-SFT/ 9,000 trajectories 50,291 reference images supervised cold-start trajectories rl GenEvolve-Data-RL/ 3,175 prompts 3,175 GT images self-evolution / RL training prompts bench GenEvolve-Bench/ 594 prompts 594 GT images held-out evaluation benchmarkAll metadata is provided in both JSONL and Parquet. The Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/MeiGen-AI/GenEvolve-Data-Bench.imagetext-to-image10K<n<100K2 likes43k downloads5mo agoHugging Face23albertvillanova /datasets-tests-compressiontextn<1K0 likes43k downloads5y agoHugging Face24argilla /databricks-dolly-15k-curated-en Guidelines In this dataset, you will find a collection of records that show a category, an instruction, a context and a response to that instruction. The aim of the project is to correct the instructions, intput and responses to make sure they are of the highest quality and that they match the task category that they belong to. All three texts should be clear and include real information. In addition, the response should be as complete but concise as possible. To curate the dataset… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-en.text10K<n<100K45 likes42k downloads3y agoHugging Face25Open-Bee /Honey-Data-15M Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.imageimage-text-to-text10M<n<100M120 likes40k downloads7mo agoHugging Face26agentica-org /DeepScaleR-Preview-Dataset Data Our training dataset consists of approximately 40,000 unique mathematics problem-answer pairs compiled from: AIME (American Invitational Mathematics Examination) problems (1984-2023) AMC (American Mathematics Competition) problems (prior to 2023) Omni-MATH dataset Still dataset Format Each row in the JSON dataset contains: problem: The mathematical question text, formatted with LaTeX notation. solution: Offical solution to the problem, including LaTeX formatting… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset.text10K<n<100K208 likes39k downloads2y agoHugging Face27hf-internal-testing /DatasetWithCapitalLetterstextn<1K0 likes37k downloads3y agoHugging Face28mlfoundations /datacomp_xlarge DataComp XLarge Pool This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.image10B<n<100B21 likes37k downloads3y agoHugging Face29databricks /databricks-dolly-15k Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.textquestion-answering10K<n<100K1.3k likes36k downloads3y agoHugging Face30apple /DataCompDR-1B Dataset Card for DataCompDR-1B This dataset contains synthetic captions, embeddings, and metadata for DataCompDR-1B. The metadata has been generated using pretrained image-text models on DataComp-1B. For details on how to use the metadata, please visit our github repository. Dataset Details Dataset Description DataCompDR is an image-text dataset and an enhancement to the DataComp dataset. We reinforce the DataComp dataset using our multi-modal… See the full description on the dataset page: https://huggingface.co/datasets/apple/DataCompDR-1B.imagetext-to-image1B<n<10B35 likes35k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.