Team Ai
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Agnuxo /OpenCLAW-SEED-data 🧬 P2PCLAW Research Papers Dataset The First Decentralized AI Research Benchmark 📊 Dataset Overview Metric Value Total Papers 116 Total Words 355,795 Total Tokens 473,208 Scored Papers 98 Average Score 5.24 / 10 Lean4 Verified 113 Research Fields 8 Unique Authors/Agents 28 🧠 What is P2PCLAW? P2PCLAW (Peer-to-Peer Collaborative Learning and Academic Work) is the world's first decentralized… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/OpenCLAW-SEED-data.text-generationn<1K4 likes6.8k downloads1h agoHugging Face02m-a-p /FineFineWeb-fasttext-seeddata FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-fasttext-seeddata.text-classificationn>1T0 likes415 downloads2y agoHugging Face03OpenSI-SI /OpenSI-0.3B-Seed-5Btok-data ⚠️ 已知缺陷:18.3% 的训练数据没有文档分隔符 本数据集的 59 个分片(561,628,432 token)内部完全没有 EOS 分隔符, 整片维基百科文本被连续拼接在一起。 后果:训练时模型在这些位置学到的是「上一篇文章的结尾后面接着一篇全新文章 的开头」——即跨文档污染,且该来源上没有任何停止信号。 受影响分片 59 / 115(57 个维基分片 + 2 个注释分片) 受影响 token 561,628,432(占 train 的 17.65%) 证据 全片 id==2 计数为 0;shard_00000.bin 头 9,592 token 与源 jsonl 首行分词结果逐位 100% 一致 成因 预训练语料走 prep_shards.py(已修:文档间插 EOS)。10-05 新增的维基分片由另一条路径生成,未复用该逻辑,把已修过的缺陷重新引入 发现时间 2026-10-08 打包时核实 未受影响:train/opensi-train-000*.bin(53 个,2,601,279… See the full description on the dataset page: https://huggingface.co/datasets/OpenSI-SI/OpenSI-0.3B-Seed-5Btok-data.text-generation0 likes336 downloads20h agoHugging Face04m-a-p /FineFineWeb-bert-seeddata FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.texttext-classification1M<n<10M2 likes332 downloads2y agoHugging Face05ByteDance-Seed /cudaLLM-data CudaLLM Dataset A high-quality dataset of PyTorch operator test cases, designed to benchmark and evaluate the capabilities of LLMs in generating optimized CUDA kernels. This dataset provides pairs of problems (standard PyTorch nn.Module implementations) and solutions (performance-optimized versions using custom CUDA kernels). It's a valuable resource for research in AI for HPC, code generation, and compiler optimization. The data is generated by DeepSeek R1, DeepSeel Coder-7B, and… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/cudaLLM-data.text-generation1K<n<10K15 likes201 downloads1y agoHugging Face06RonalLI /OpenCLAW-SEED-data 🌱 OpenCLAW SEED Training Data Autonomous self-growing training dataset for the OpenCLAW SEED system. What is this? This dataset is continuously growing. Every 6 hours, the SEED harvester collects new training data from: ArXiv papers on neuromorphic computing, physics-based AI, and AGI Semantic Scholar research database Our own GitHub repositories (57 repos) Agent interaction logs and self-reflection Format Standard instruction-following JSONL:… See the full description on the dataset page: https://huggingface.co/datasets/RonalLI/OpenCLAW-SEED-data.text-generationn<1K0 likes32 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.