datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenCLAW-SEED-data
🧬 P2PCLAW Research Papers Dataset
The First Decentralized AI Research Benchmark
📊 Dataset Overview
Metric
Value
Total Papers
116
Total Words
355,795
Total Tokens
473,208
Scored Papers
98
Average Score
5.24 / 10
Lean4 Verified
113
Research Fields
8
Unique Authors/Agents
28
🧠 What is P2PCLAW?
P2PCLAW (Peer-to-Peer Collaborative Learning and Academic Work) is the world's first decentralized… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/OpenCLAW-SEED-data.SEED-Data-Edit-Part2-3
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part2-3.sainfoin-seed-datasetSEED-Data-Edit-Part1-Unsplash
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Unsplash.SEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.agent-apprenticeship-seed-dataset
Agent Apprenticeship Seed Dataset
The living ecosystem where AI agents run automated workflow loops on any task, improve through execution, and turn each run into reusable work experience + data to improve future agents.
As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where real-world tasks generate reusable learning signals and complex workflows advance through agent loops that turn execution into shared… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset.multimodal-spectroscopic-seed-datasets-fullBiointelligenceAgent01-SeedDataset
Biointelligence Agent Worlds Seed Dataset 0.1
Configurations
Configuration
Contents
Rows
targets
Public-real and explicitly synthetic targets
1000
entities
Normalized world entities
14516
world_snapshots
Observed and simulated temporal states
356014
source_events
Retrieved/discovered source records and normalized claims
33074
relationships
Evidence-linked target and entity relationships
205800
agent_runs
Observable structured Codex run results… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/BiointelligenceAgent01-SeedDataset.FineFineWeb-fasttext-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-fasttext-seeddata.agent-apprenticeship-seed-dataset_v0.2
Agent Apprenticeship Seed Dataset v0.2
Real-world agent work experience, looped into collective learning.
The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.
As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset_v0.2.OpenSI-0.3B-Seed-5Btok-data
⚠️ 已知缺陷:18.3% 的训练数据没有文档分隔符
本数据集的 59 个分片(561,628,432 token)内部完全没有 EOS 分隔符,
整片维基百科文本被连续拼接在一起。
后果:训练时模型在这些位置学到的是「上一篇文章的结尾后面接着一篇全新文章
的开头」——即跨文档污染,且该来源上没有任何停止信号。
受影响分片
59 / 115(57 个维基分片 + 2 个注释分片)
受影响 token
561,628,432(占 train 的 17.65%)
证据
全片 id==2 计数为 0;shard_00000.bin 头 9,592 token 与源 jsonl 首行分词结果逐位 100% 一致
成因
预训练语料走 prep_shards.py(已修:文档间插 EOS)。10-05 新增的维基分片由另一条路径生成,未复用该逻辑,把已修过的缺陷重新引入
发现时间
2026-10-08 打包时核实
未受影响:train/opensi-train-000*.bin(53 个,2,601,279… See the full description on the dataset page: https://huggingface.co/datasets/OpenSI-SI/OpenSI-0.3B-Seed-5Btok-data.FineFineWeb-bert-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.cudaLLM-data
CudaLLM Dataset
A high-quality dataset of PyTorch operator test cases, designed to benchmark and evaluate the capabilities of LLMs in generating optimized CUDA kernels. This dataset provides pairs of problems (standard PyTorch nn.Module implementations) and solutions (performance-optimized versions using custom CUDA kernels). It's a valuable resource for research in AI for HPC, code generation, and compiler optimization. The data is generated by DeepSeek R1, DeepSeel Coder-7B, and… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/cudaLLM-data.BrowserAgent-SeedData
BrowserAgent-Data
Dataset used in https://github.com/TIGER-AI-Lab/BrowserAgent.
Summary
Total rows: 230,015
Total size: ~29.47 MB
Subsets: 2wiki, bamboogle, hotpot, musique, nq, popqa
Splits and Sizes
2wiki: 22,576 rows (~2.73 MB)
bamboogle: 125 rows (~0.03 MB)
hotpot: 97,852 rows (~14.20 MB)
musique: 12,417 rows (~0.92 MB)
nq: 82,778 rows (~10.14 MB)
popqa: 14,267 rows (~1.45 MB)
Files are stored as Parquet under each subset directory with dev/train/test… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/BrowserAgent-SeedData.synthetic-crm-seed-data-for-property-management
Synthetic CRM Seed Data for Concierge & Property Management: GDPR-compliant, realistic market distributions for software testing & demos (SQL/CSV) (DEMO)
Thank you for downloading this synthetic dataset!
This repository contains a functional, lightweight preview of a 100% synthetic dataset designed for testing real estate CRMs and property management software.
>> Full version available on Gumroad <<
This demo provides a baseline schema. For large-scale stress… See the full description on the dataset page: https://huggingface.co/datasets/Pandoxyd/synthetic-crm-seed-data-for-property-management.SEED-Data-Edit-Part1-Openimagesuplimit-synthetic-data-week-1-with-seed
Dataset Card for uplimit-synthetic-data-week-1-with-seed
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
demo_.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/uplimit/uplimit-synthetic-data-week-1-with-seed/raw/main/demo_.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that… See the full description on the dataset page: https://huggingface.co/datasets/uplimit/uplimit-synthetic-data-week-1-with-seed.instruction_filtering_askllm_seed_data_code_w_openthoughtsuplimit-synthetic-data-week-1-with-seed
Dataset Card for uplimit-synthetic-data-week-1-with-seed
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
colab_kernel_launcher.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/ndhananj/uplimit-synthetic-data-week-1-with-seed/raw/main/colab_kernel_launcher.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used… See the full description on the dataset page: https://huggingface.co/datasets/ndhananj/uplimit-synthetic-data-week-1-with-seed.uplimit-synthetic-data-week-1-with-seed
Dataset Card for uplimit-synthetic-data-week-1-with-seed
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
[c:\Users\mimartin\OneDrive - Sony\Desktop\dev\python_projects\uplimit_classes\synthetic_data_generation.venv\Lib\site-packages\ipykernel_launcher.py](https://huggingface.co/datasets/mimartin1234/uplimit-synthetic-data-week-1-with-seed/raw/main/c:\Users\mimartin\OneDrive -… See the full description on the dataset page: https://huggingface.co/datasets/mimartin1234/uplimit-synthetic-data-week-1-with-seed.uplimit-synthetic-data-week-1-with-seed
Dataset Card for uplimit-synthetic-data-week-1-with-seed
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
colab_kernel_launcher.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/uonyeka/uplimit-synthetic-data-week-1-with-seed/raw/main/colab_kernel_launcher.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to… See the full description on the dataset page: https://huggingface.co/datasets/uonyeka/uplimit-synthetic-data-week-1-with-seed.molpuzzle-seed-datasetsSEED-Data-Edit
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit.SynPO-Seed_Data_Format4PromptGeneratorSEED-Data-Edit-Part1-Unsplashuplimit-synthetic-data-week-1-with-seed
Dataset Card for uplimit-synthetic-data-week-1-with-seed
This dataset is a distilled version of the tatsu-lab/alpaca dataset, refined to enhance instruction quality through advanced data generation techniques.
Galeria Singularity
Dataset Creation Process:
Initial Sampling: A subset of 500 lines was extracted from the original tatsu-lab/alpaca dataset.
Instruction Enhancement Techniques:
Self-Instruct: This method involves using a language model to generate diverse… See the full description on the dataset page: https://huggingface.co/datasets/eliasprost/uplimit-synthetic-data-week-1-with-seed.uplimit-synthetic-data-week-1-with-seed
Dataset Card for uplimit-synthetic-data-week-1-with-seed
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
colab_kernel_launcher.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/rutvima/uplimit-synthetic-data-week-1-with-seed/raw/main/colab_kernel_launcher.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to… See the full description on the dataset page: https://huggingface.co/datasets/rutvima/uplimit-synthetic-data-week-1-with-seed.OpenCLAW-SEED-data
🌱 OpenCLAW SEED Training Data
Autonomous self-growing training dataset for the OpenCLAW SEED system.
What is this?
This dataset is continuously growing. Every 6 hours, the SEED harvester collects new training data from:
ArXiv papers on neuromorphic computing, physics-based AI, and AGI
Semantic Scholar research database
Our own GitHub repositories (57 repos)
Agent interaction logs and self-reflection
Format
Standard instruction-following JSONL:… See the full description on the dataset page: https://huggingface.co/datasets/RonalLI/OpenCLAW-SEED-data.instruction_filtering_embedding_filter_seed_data_math_w_openthoughtsmultimodal-spectroscopic-seed-datasets-1000
