Akahsizrr/kernelswarm
kernelswarm — multi-agent kernel-optimization episodes (RL) Synthetic dataset of multi-agent long-horizon optimization campaigns on AI-inference kernels. Each row is one whole episode: a lead orchestrator plus 2–20 specialist agents split the work, exchange dispatches/statuses/handoffs, submit full kernel candidates, and receive simulated compile/verify/bench feedback — with a reward attached to every environment result. Inference-focused only. All ops are inference primitives… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/kernelswarm.
kernelswarm — multi-agent kernel-optimization episodes (RL)
Synthetic dataset of multi-agent long-horizon optimization campaigns on AI-inference kernels. Each row is one whole episode: a lead orchestrator plus 2–20 specialist agents split the work, exchange dispatches/statuses/handoffs, submit full kernel candidates, and receive simulated compile/verify/bench feedback — with a reward attached to every environment result.
Inference-focused only. All ops are inference primitives (decode/prefill attention, quantized GEMM/GEMV, sampling, norms, KV-cache, MoE, Mamba).
⚠️ All benchmarks, compiles and correctness checks are simulated analytically — no kernel was ever compiled or executed on a real GPU.
Row format (one row = one episode)
{
"episode_id": "flash_decode_h100_B=8,..._dispatch_74",
"scenario": "dispatch", // dispatch | parallel_sweep | pipeline | campaign
"n_agents": 6,
"op": "flash_decode", "impl": "cuda",
"gpu": "h100", "arch": "sm_90a", "dtype": "fp16",
"shape": {...}, "baseline_ms": 0.42, "best_ms": 0.006,
"speedup": 66.4, "elapsed_s": 17335,
"agents": [{"id": "a01", "role": "worker", "skill": 0.9,
"tries": 8, "kept": 5, "best_ms": 0.008}, ...],
"num_messages": 74,
"messages": [ ... ]
}Message kinds
Rewards
resultmsgs:min(1, log2(prev_best / new))on kept improvement;-0.05no-improvement,-0.1compile error,-0.3wrong result,-0.05unsupported arch.verdictmsg:log2(speedup) / 4terminal bonus.
Scenarios
- dispatch — lead splits variants across workers, merges, redirects losers (2–8 agents)
- parallel_sweep — N workers each own a cfg slice of one variant; leaderboard + top-2 refine (4–20)
- pipeline — proposer → reviewer → verifier → tuner chain over variants (3–6)
- campaign — fleet dispatch over variant×cfg space + verifier rebench (5–20)
Coverage
- 17 ops (14 CUDA + 3 Triton impls): flashdecode, swadecode, gemmfp16, int8gemm, fp8gemm, int4gemv, rmsnorm, rope, swigluact, topksample, nucleussample, mambascan, moescatter, kvappend, fusedqkvrope, tritonflashdecode, triton_rmsnorm
- GPUs: h100, h200, a100, a10, rtxa6000, rtx5090 (arch features gated: wgmma/tma/fp8/fp4/cp.async)
- 2400 episodes · ~170k messages · avg 71 msgs/episode · avg 8.6 agents
- Simulated horizon: 1–5 h per episode
Files
batches/batch_*.jsonl— 24 shards × 100 episodesepisodes.jsonl— all episodes concatenatedtasks_holdout.jsonl— (op, gpu, shape) task specs for evalstats.json— aggregate stats
