Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01transferable-samplers /many-peptides-md [!IMPORTANT] Critical Update The original 8AA TICA models within subsampled_trajectories/*/8AA/*.npz employed a CA-only atom selection. These models are not valid for comparison to results in our paper. Updated files (uploaded 15/12/2025) now contain corrected models. If you previously downloaded this dataset, please re-download to ensure accurate results. Note: Codebase references to tica_features_ca must now be replaced with tica_features. This was resolved in our codebase by PR #26. Note:… See the full description on the dataset page: https://huggingface.co/datasets/transferable-samplers/many-peptides-md.11 likes1.8m downloads10mo agoHugging Face02ryanmarten /OpenThoughts-1k-sample [!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. Available Subsets default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.text1K<n<10K60 likes1.1m downloads1y agoHugging Face03malcolmrey /samplesimage1K<n<10K10 likes105k downloads18h agoHugging Face04m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes53k downloads2y agoHugging Face05axolotl-ai-co /evolkit-logprobs-pipeline-75k-v2-sampletextn<1K1 likes17k downloads2y agoHugging Face06hf-internal-testing /dummy-audio-samplesaudion<1K0 likes16k downloads15d agoHugging Face07icyCreater /SpreadsheetBench-2-sample SpreadsheetBench 2 Dataset SpreadsheetBench 2 evaluates spreadsheet agents on end-to-end business and financial workflows. The release contains 321 tasks in four categories: Debugging, Financial Model, Template, and Visualization. Overview Directory Task type Tasks Input workbooks Gold workbooks Debugging Spreadsheet debugging and error correction 100 100 10 Financial_Model Completion of multi-sheet financial models 100 100 20 Template Completion of… See the full description on the dataset page: https://huggingface.co/datasets/icyCreater/SpreadsheetBench-2-sample.0 likes13k downloads2mo agoHugging Face08zeahub /camus-sample CAMUS Sample - 2-D Echocardiographic Ultrasound Dataset This is a sample subset of the full CAMUS dataset, provided for demonstration and testing purposes. It contains 6 files (1 patient per split). For the full dataset (500 patients), see: zeahub/camus. This dataset is a zea-format (HDF5) conversion of the CAMUS dataset for multi-structure segmentation in 2-D echocardiography. Property Value Modality 2-D transthoracic echocardiography Patients 500 Views… See the full description on the dataset page: https://huggingface.co/datasets/zeahub/camus-sample.image-segmentation0 likes13k downloads2mo agoHugging Face09Vchitect /VBench-2.0_sampled_videos Sample Videos of VBench-2.0 This dataset is used in the paper:👉 arXiv:2503.21755 video10K<n<100K0 likes12k downloads1y agoHugging Face10sayakpaul /sample-datasetsimagen<1K1 likes11k downloads2mo agoHugging Face11ScalingIntelligence /kernelbench-samples KernelBench Samples Samples from experiments for KernelBench, described in our arxiv Learn more about KernelBench from our Paper Github Repo The samples are organized as such baseline_eval (Section 4 Baseline) repeated_sampling (Section 5.1.1 Repeated Sampling) iterative_refinement (Section 5.1.2 Iterative Refinement of Generations) Within each folder, we organize the results by /level/model/problem_{id}/sample_{id}. The inner most .json file contains the generated kernel and… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/kernelbench-samples.3 likes11k downloads2y agoHugging Face12moonshine-ai /audio_samples_1kaudio0 likes11k downloads7mo agoHugging Face13TIACentre /TIAToolBox_Remote_Samples LICENSE No re-distribution allowed. Purpose This repository contains publicly available samples used by the TIAToolBox for testing purposes. Some of these images have been downloaded from [OpenSlide] for code verification purposes. GitHub Repository: [TIAToolBox] other0 likes10k downloads1mo agoHugging Face14Research-EAI /essential-web-1t-sample-fdc-partitioned 🌐 Essential-Web: FDC Level-2 Partitioned Dataset 📋 Dataset Description This dataset contains a 1 trillion token sample from Essential-Web, partitioned by Free Decimal Correspondence (FDC) level-2 categories. Essential-Web is a 24-trillion-token web dataset with extensive document-level metadata designed to enable rapid dataset curation through SQL-like filtering. 🔍 Free Decimal Correspondence (FDC) The FDC taxonomy is an open classification system… See the full description on the dataset page: https://huggingface.co/datasets/Research-EAI/essential-web-1t-sample-fdc-partitioned.text100M<n<1B5 likes8.7k downloads1y agoHugging Face15agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes6.6k downloads2y agoHugging Face16paxini /Omnisharing_DB_SampleData Overview The embodied intelligence industry is currently facing significant development challenges. The most critical issue is the lack of high-quality data, particularly omnimodal data that integrates force and tactile sensing. The PaXini introduces the PX OmniSharing Dataset, built on the PaXini Super EID Factory, enabling large-scale, high-fidelity human data collection across diverse tasks and scenarios. The dataset includes multi-dimensional tactile data, multi-view visual… See the full description on the dataset page: https://huggingface.co/datasets/paxini/Omnisharing_DB_SampleData.8 likes5.6k downloads5mo agoHugging Face17Vchitect /VBench_sampled_videogated VBench Sampled Video 1K<n<10K4 likes5.5k downloads1mo agoHugging Face18unileon-robotics /malware-samples This dataset is part of the ULE-CIBERLAB Project: Transfer of knowledge in cybersecurity for the country's business fabric, funded by the European Union NextGeneration-EU, Recovery, Transformation and Resilience Plan, through INCIBE. MALWARE-SAMPLES DATASET Disclaimer: This repository contains real samples of malware that can be executed (.exe) and artifacts related with their execution in CAPEv2 sandbox (JSON/HTML reports, screenshots, dropped files). DO NOT execute any of… See the full description on the dataset page: https://huggingface.co/datasets/unileon-robotics/malware-samples.1K<n<10K7 likes5.3k downloads2mo agoHugging Face19RMT-team /babilong-1k-samples BABILong (1000 samples) : a long-context needle-in-a-haystack benchmark for LLMs Preprint is on arXiv and code for LLM evaluation is available on GitHub. BABILong Leaderboard with top-performing long-context models. bAbI + Books = BABILong BABILong is a novel generative benchmark for evaluating the performance of NLP models in processing arbitrarily long documents with distributed facts. It contains 9 configs, corresponding to different sequence lengths in tokens: 0k… See the full description on the dataset page: https://huggingface.co/datasets/RMT-team/babilong-1k-samples.text10K<n<100K4 likes5.2k downloads2y agoHugging Face20unileon-robotics /community-benign-samples This dataset is part of the ULE-CIBERLAB Project: Transfer of knowledge in cybersecurity for the country's business fabric, funded by the European Union NextGeneration-EU, Recovery, Transformation and Resilience Plan, through INCIBE. MALWARE-SAMPLES DATASET Disclaimer: This repository contains benign samples with their execution in CAPEv2 sandbox (JSON/HTML reports, screenshots, dropped files). This README file explains how dataset is structured, its metadata, safe use as well… See the full description on the dataset page: https://huggingface.co/datasets/unileon-robotics/community-benign-samples.2 likes5k downloads3mo agoHugging Face21HCAI-Lab-GT /dolma3-6t-sample-100000-docs dolma3-6t-sample-100000-docs Materialized stratified sample of 100,000 docs per bin (~58k source-shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers (worker_0000/ through worker_0127/). Layout HCAI-Lab/dolma3-6t-sample-100000-docs/ ├── bin_summary.csv ├── sample_contract.json └── worker_NNNN/ └── soc127__phase1_pool_shared__...__shard_NNNNNNNN.jsonl.zst (~450 files per worker) Dual-access… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-100000-docs.0 likes5k downloads4mo agoHugging Face22EleutherAI /rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script. Getting Started RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text documents coming from 84 CommonCrawl snapshots and processed using the CCNet pipeline. Out of these, there are 30B documents in the corpus that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.texttext-generation1M<n<10M2 likes4.8k downloads2y agoHugging Face23pufanyi /vbvrpro_sampler_trajectories-2200-steps VBVR-Pro sampler trajectory media This media archive backs the interactive pufanyi/vbvrpro_sampler_trajectories Space. It contains 12 matched evaluation cells: DiffSynth step-35500 baseline and DanceGRPO checkpoint 2200 Flow-CPS noise 0.1, 0.3, 0.7, and 0.9 deterministic FlowMatch Euler ODE and UniPC ODE 500 samples per cell across 100 VBVR-Pro tasks The deployment is split across three public media repositories so each Git-backed Dataset remains below Hugging Face's… See the full description on the dataset page: https://huggingface.co/datasets/pufanyi/vbvrpro_sampler_trajectories-2200-steps.video10K<n<100K0 likes4.7k downloads2mo agoHugging Face24davidberenstein1957 /samplesimage1K<n<10K0 likes4.6k downloads10mo agoHugging Face25RoboSynChallenge /cobotmagic_Sim_sample_loading0 likes4.5k downloads3mo agoHugging Face26unileon-robotics /community-suspicious-samples This dataset is part of the ULE-CIBERLAB Project: Transfer of knowledge in cybersecurity for the country's business fabric, funded by the European Union NextGeneration-EU, Recovery, Transformation and Resilience Plan, through INCIBE. MALWARE-SAMPLES DATASET Disclaimer: This repository may contain real samples of malware that can be executed (.exe) and artifacts related with their execution in CAPEv2 sandbox (JSON/HTML reports, screenshots, dropped files). DO NOT execute any of… See the full description on the dataset page: https://huggingface.co/datasets/unileon-robotics/community-suspicious-samples.image1K<n<10K0 likes4.5k downloads3mo agoHugging Face27Vchitect /VBench-I2V_sampled_videovideon<1K0 likes4.2k downloads7mo agoHugging Face28pufanyi /vbvrpro_sampler_trajectories-baseline-steps VBVR-Pro sampler trajectory media This media archive backs the interactive pufanyi/vbvrpro_sampler_trajectories Space. It contains 12 matched evaluation cells: DiffSynth step-35500 baseline and DanceGRPO checkpoint 2200 Flow-CPS noise 0.1, 0.3, 0.7, and 0.9 deterministic FlowMatch Euler ODE and UniPC ODE 500 samples per cell across 100 VBVR-Pro tasks The deployment is split across three public media repositories so each Git-backed Dataset remains below Hugging Face's… See the full description on the dataset page: https://huggingface.co/datasets/pufanyi/vbvrpro_sampler_trajectories-baseline-steps.video10K<n<100K0 likes4.2k downloads2mo agoHugging Face29mishig /sample_imagesimagen<1K1 likes3.8k downloads3y agoHugging Face30guided-tl /global-samples0 likes3.4k downloads18d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.