datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenCLAW-SEED-data
🧬 P2PCLAW Research Papers Dataset
The First Decentralized AI Research Benchmark
📊 Dataset Overview
Metric
Value
Total Papers
116
Total Words
355,795
Total Tokens
473,208
Scored Papers
98
Average Score
5.24 / 10
Lean4 Verified
113
Research Fields
8
Unique Authors/Agents
28
🧠 What is P2PCLAW?
P2PCLAW (Peer-to-Peer Collaborative Learning and Academic Work) is the world's first decentralized… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/OpenCLAW-SEED-data.FineFineWeb-fasttext-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-fasttext-seeddata.OpenSI-0.3B-Seed-5Btok-data
⚠️ 已知缺陷:18.3% 的训练数据没有文档分隔符
本数据集的 59 个分片(561,628,432 token)内部完全没有 EOS 分隔符,
整片维基百科文本被连续拼接在一起。
后果:训练时模型在这些位置学到的是「上一篇文章的结尾后面接着一篇全新文章
的开头」——即跨文档污染,且该来源上没有任何停止信号。
受影响分片
59 / 115(57 个维基分片 + 2 个注释分片)
受影响 token
561,628,432(占 train 的 17.65%)
证据
全片 id==2 计数为 0;shard_00000.bin 头 9,592 token 与源 jsonl 首行分词结果逐位 100% 一致
成因
预训练语料走 prep_shards.py(已修:文档间插 EOS)。10-05 新增的维基分片由另一条路径生成,未复用该逻辑,把已修过的缺陷重新引入
发现时间
2026-10-08 打包时核实
未受影响:train/opensi-train-000*.bin(53 个,2,601,279… See the full description on the dataset page: https://huggingface.co/datasets/OpenSI-SI/OpenSI-0.3B-Seed-5Btok-data.FineFineWeb-bert-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.cudaLLM-data
CudaLLM Dataset
A high-quality dataset of PyTorch operator test cases, designed to benchmark and evaluate the capabilities of LLMs in generating optimized CUDA kernels. This dataset provides pairs of problems (standard PyTorch nn.Module implementations) and solutions (performance-optimized versions using custom CUDA kernels). It's a valuable resource for research in AI for HPC, code generation, and compiler optimization. The data is generated by DeepSeek R1, DeepSeel Coder-7B, and… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/cudaLLM-data.OpenCLAW-SEED-data
🌱 OpenCLAW SEED Training Data
Autonomous self-growing training dataset for the OpenCLAW SEED system.
What is this?
This dataset is continuously growing. Every 6 hours, the SEED harvester collects new training data from:
ArXiv papers on neuromorphic computing, physics-based AI, and AGI
Semantic Scholar research database
Our own GitHub repositories (57 repos)
Agent interaction logs and self-reflection
Format
Standard instruction-following JSONL:… See the full description on the dataset page: https://huggingface.co/datasets/RonalLI/OpenCLAW-SEED-data.
