Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Agnuxo /OpenCLAW-SEED-data 🧬 P2PCLAW Research Papers Dataset The First Decentralized AI Research Benchmark 📊 Dataset Overview Metric Value Total Papers 116 Total Words 355,795 Total Tokens 473,208 Scored Papers 98 Average Score 5.24 / 10 Lean4 Verified 113 Research Fields 8 Unique Authors/Agents 28 🧠 What is P2PCLAW? P2PCLAW (Peer-to-Peer Collaborative Learning and Academic Work) is the world's first decentralized… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/OpenCLAW-SEED-data.text-generationn<1K4 likes6.8k downloads2h agoHugging Face02AILab-CVC /SEED-Data-Edit-Part2-3 SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part2-3.text-to-image1M<n<10M12 likes1.7k downloads2y agoHugging Face03tli-legumes /sainfoin-seed-datasetimage100K<n<1M0 likes1.5k downloads3mo agoHugging Face04AILab-CVC /SEED-Data-Edit-Part1-Unsplash SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Unsplash.text-to-image1M<n<10M10 likes1.2k downloads2y agoHugging Face05AILab-CVC /SEED-Data-Edit-Part1-Openimages SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.tabulartext-to-image1M<n<10M10 likes1.1k downloads2y agoHugging Face06rayrren /agent-apprenticeship-seed-dataset Agent Apprenticeship Seed Dataset The living ecosystem where AI agents run automated workflow loops on any task, improve through execution, and turn each run into reusable work experience + data to improve future agents. As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where real-world tasks generate reusable learning signals and complex workflows advance through agent loops that turn execution into shared… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset.tabular1K<n<10K0 likes1k downloads4mo agoHugging Face07SpectrumWorld /multimodal-spectroscopic-seed-datasets-fullimage100K<n<1M0 likes512 downloads1y agoHugging Face08rayrren /BiointelligenceAgent01-SeedDataset Biointelligence Agent Worlds Seed Dataset 0.1 Configurations Configuration Contents Rows targets Public-real and explicitly synthetic targets 1000 entities Normalized world entities 14516 world_snapshots Observed and simulated temporal states 356014 source_events Retrieved/discovered source records and normalized claims 33074 relationships Evidence-linked target and entity relationships 205800 agent_runs Observable structured Codex run results… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/BiointelligenceAgent01-SeedDataset.tabularother100K<n<1M0 likes450 downloads26d agoHugging Face09m-a-p /FineFineWeb-fasttext-seeddata FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-fasttext-seeddata.text-classificationn>1T0 likes415 downloads2y agoHugging Face10rayrren /agent-apprenticeship-seed-dataset_v0.2 Agent Apprenticeship Seed Dataset v0.2 Real-world agent work experience, looped into collective learning. The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents. As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset_v0.2.tabular10K<n<100K0 likes362 downloads3mo agoHugging Face11OpenSI-SI /OpenSI-0.3B-Seed-5Btok-data ⚠️ 已知缺陷:18.3% 的训练数据没有文档分隔符 本数据集的 59 个分片(561,628,432 token)内部完全没有 EOS 分隔符, 整片维基百科文本被连续拼接在一起。 后果:训练时模型在这些位置学到的是「上一篇文章的结尾后面接着一篇全新文章 的开头」——即跨文档污染,且该来源上没有任何停止信号。 受影响分片 59 / 115(57 个维基分片 + 2 个注释分片) 受影响 token 561,628,432(占 train 的 17.65%) 证据 全片 id==2 计数为 0;shard_00000.bin 头 9,592 token 与源 jsonl 首行分词结果逐位 100% 一致 成因 预训练语料走 prep_shards.py(已修:文档间插 EOS)。10-05 新增的维基分片由另一条路径生成,未复用该逻辑,把已修过的缺陷重新引入 发现时间 2026-10-08 打包时核实 未受影响:train/opensi-train-000*.bin(53 个,2,601,279… See the full description on the dataset page: https://huggingface.co/datasets/OpenSI-SI/OpenSI-0.3B-Seed-5Btok-data.text-generation0 likes336 downloads15h agoHugging Face12m-a-p /FineFineWeb-bert-seeddata FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.texttext-classification1M<n<10M2 likes332 downloads2y agoHugging Face13ByteDance-Seed /cudaLLM-data CudaLLM Dataset A high-quality dataset of PyTorch operator test cases, designed to benchmark and evaluate the capabilities of LLMs in generating optimized CUDA kernels. This dataset provides pairs of problems (standard PyTorch nn.Module implementations) and solutions (performance-optimized versions using custom CUDA kernels). It's a valuable resource for research in AI for HPC, code generation, and compiler optimization. The data is generated by DeepSeek R1, DeepSeel Coder-7B, and… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/cudaLLM-data.text-generation1K<n<10K15 likes201 downloads1y agoHugging Face14TIGER-Lab /BrowserAgent-SeedData BrowserAgent-Data Dataset used in https://github.com/TIGER-AI-Lab/BrowserAgent. Summary Total rows: 230,015 Total size: ~29.47 MB Subsets: 2wiki, bamboogle, hotpot, musique, nq, popqa Splits and Sizes 2wiki: 22,576 rows (~2.73 MB) bamboogle: 125 rows (~0.03 MB) hotpot: 97,852 rows (~14.20 MB) musique: 12,417 rows (~0.92 MB) nq: 82,778 rows (~10.14 MB) popqa: 14,267 rows (~1.45 MB) Files are stored as Parquet under each subset directory with dev/train/test… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/BrowserAgent-SeedData.textquestion-answering100K<n<1M0 likes159 downloads11mo agoHugging Face15Pandoxyd /synthetic-crm-seed-data-for-property-management Synthetic CRM Seed Data for Concierge & Property Management: GDPR-compliant, realistic market distributions for software testing & demos (SQL/CSV) (DEMO) Thank you for downloading this synthetic dataset! This repository contains a functional, lightweight preview of a 100% synthetic dataset designed for testing real estate CRMs and property management software. >> Full version available on Gumroad << This demo provides a baseline schema. For large-scale stress… See the full description on the dataset page: https://huggingface.co/datasets/Pandoxyd/synthetic-crm-seed-data-for-property-management.0 likes117 downloads1mo agoHugging Face16123123aa123 /SEED-Data-Edit-Part1-Openimages0 likes105 downloads2y agoHugging Face17uplimit /uplimit-synthetic-data-week-1-with-seed Dataset Card for uplimit-synthetic-data-week-1-with-seed This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: demo_.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/uplimit/uplimit-synthetic-data-week-1-with-seed/raw/main/demo_.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that… See the full description on the dataset page: https://huggingface.co/datasets/uplimit/uplimit-synthetic-data-week-1-with-seed.textn<1K0 likes99 downloads2y agoHugging Face18mlfoundations-dev /instruction_filtering_askllm_seed_data_code_w_openthoughtstabular100K<n<1M0 likes87 downloads2y agoHugging Face19ndhananj /uplimit-synthetic-data-week-1-with-seed Dataset Card for uplimit-synthetic-data-week-1-with-seed This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: colab_kernel_launcher.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/ndhananj/uplimit-synthetic-data-week-1-with-seed/raw/main/colab_kernel_launcher.py" Dataset Summary This dataset contains a pipeline.yaml which can be used… See the full description on the dataset page: https://huggingface.co/datasets/ndhananj/uplimit-synthetic-data-week-1-with-seed.textn<1K0 likes83 downloads2y agoHugging Face20mimartin1234 /uplimit-synthetic-data-week-1-with-seed Dataset Card for uplimit-synthetic-data-week-1-with-seed This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: [c:\Users\mimartin\OneDrive - Sony\Desktop\dev\python_projects\uplimit_classes\synthetic_data_generation.venv\Lib\site-packages\ipykernel_launcher.py](https://huggingface.co/datasets/mimartin1234/uplimit-synthetic-data-week-1-with-seed/raw/main/c:\Users\mimartin\OneDrive -… See the full description on the dataset page: https://huggingface.co/datasets/mimartin1234/uplimit-synthetic-data-week-1-with-seed.textn<1K0 likes49 downloads2y agoHugging Face21uonyeka /uplimit-synthetic-data-week-1-with-seed Dataset Card for uplimit-synthetic-data-week-1-with-seed This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: colab_kernel_launcher.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/uonyeka/uplimit-synthetic-data-week-1-with-seed/raw/main/colab_kernel_launcher.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to… See the full description on the dataset page: https://huggingface.co/datasets/uonyeka/uplimit-synthetic-data-week-1-with-seed.textn<1K0 likes37 downloads2y agoHugging Face22SpectrumWorld /molpuzzle-seed-datasetsimagen<1K0 likes37 downloads1y agoHugging Face23AILab-CVC /SEED-Data-Edit SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit.text-to-image1M<n<10M20 likes35 downloads2y agoHugging Face24Windy /SynPO-Seed_Data_Format4PromptGeneratortext10K<n<100K1 likes34 downloads2y agoHugging Face25123123aa123 /SEED-Data-Edit-Part1-Unsplash0 likes34 downloads2y agoHugging Face26eliasprost /uplimit-synthetic-data-week-1-with-seed Dataset Card for uplimit-synthetic-data-week-1-with-seed This dataset is a distilled version of the tatsu-lab/alpaca dataset, refined to enhance instruction quality through advanced data generation techniques.​ Galeria Singularity Dataset Creation Process: Initial Sampling: A subset of 500 lines was extracted from the original tatsu-lab/alpaca dataset.​ Instruction Enhancement Techniques: Self-Instruct: This method involves using a language model to generate diverse… See the full description on the dataset page: https://huggingface.co/datasets/eliasprost/uplimit-synthetic-data-week-1-with-seed.textn<1K0 likes33 downloads2y agoHugging Face27rutvima /uplimit-synthetic-data-week-1-with-seed Dataset Card for uplimit-synthetic-data-week-1-with-seed This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: colab_kernel_launcher.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/rutvima/uplimit-synthetic-data-week-1-with-seed/raw/main/colab_kernel_launcher.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to… See the full description on the dataset page: https://huggingface.co/datasets/rutvima/uplimit-synthetic-data-week-1-with-seed.textn<1K0 likes33 downloads2y agoHugging Face28RonalLI /OpenCLAW-SEED-data 🌱 OpenCLAW SEED Training Data Autonomous self-growing training dataset for the OpenCLAW SEED system. What is this? This dataset is continuously growing. Every 6 hours, the SEED harvester collects new training data from: ArXiv papers on neuromorphic computing, physics-based AI, and AGI Semantic Scholar research database Our own GitHub repositories (57 repos) Agent interaction logs and self-reflection Format Standard instruction-following JSONL:… See the full description on the dataset page: https://huggingface.co/datasets/RonalLI/OpenCLAW-SEED-data.text-generationn<1K0 likes32 downloads8mo agoHugging Face29mlfoundations-dev /instruction_filtering_embedding_filter_seed_data_math_w_openthoughtstabular100K<n<1M0 likes31 downloads2y agoHugging Face30SpectrumWorld /multimodal-spectroscopic-seed-datasets-1000image1K<n<10K0 likes31 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.