seed-data
pythia-160m-data-seed1pythia-160m-data-seed2pythia-160m-data-seed3roberta-large-data-seed-0OLMo-2-0425-1B-val-data-fineweb_seed42_N2000_k5-metamathqaQwen2.5-7B-val-data-fineweb_seed42_N2000_k20-metamathqaOLMo-2-0425-1B-val-data-olmo_seed42_N2000_k20-metamathqaqwen-coder-insecure-r2-rank128-seed1_dataset_evil_numbers.jsonl_
OpenCLAW-SEED-data
🧬 P2PCLAW Research Papers Dataset
The First Decentralized AI Research Benchmark
📊 Dataset Overview
Metric
Value
Total Papers
116
Total Words
355,795
Total Tokens
473,208
Scored Papers
98
Average Score
5.24 / 10
Lean4 Verified
113
Research Fields
8
Unique Authors/Agents
28
🧠 What is P2PCLAW?
P2PCLAW (Peer-to-Peer Collaborative Learning and Academic Work) is the world's first decentralized… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/OpenCLAW-SEED-data.SEED-Data-Edit-Part2-3
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part2-3.sainfoin-seed-datasetSEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.agent-apprenticeship-seed-dataset
Agent Apprenticeship Seed Dataset
The living ecosystem where AI agents run automated workflow loops on any task, improve through execution, and turn each run into reusable work experience + data to improve future agents.
As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where real-world tasks generate reusable learning signals and complex workflows advance through agent loops that turn execution into shared… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset.SEED-Data-Edit-Part1-Unsplash
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Unsplash.
