Team Ai
Datasetpublic

Akahsizrr/kernelswarm

kernelswarm — multi-agent kernel-optimization episodes (RL) Synthetic dataset of multi-agent long-horizon optimization campaigns on AI-inference kernels. Each row is one whole episode: a lead orchestrator plus 2–20 specialist agents split the work, exchange dispatches/statuses/handoffs, submit full kernel candidates, and receive simulated compile/verify/bench feedback — with a reward attached to every environment result. Inference-focused only. All ops are inference primitives… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/kernelswarm.

sourceHugging Facemitupdated 15d agoView on Hugging Face
0likes55downloads
Dataset Card

kernelswarm — multi-agent kernel-optimization episodes (RL)

Synthetic dataset of multi-agent long-horizon optimization campaigns on AI-inference kernels. Each row is one whole episode: a lead orchestrator plus 2–20 specialist agents split the work, exchange dispatches/statuses/handoffs, submit full kernel candidates, and receive simulated compile/verify/bench feedback — with a reward attached to every environment result.

Inference-focused only. All ops are inference primitives (decode/prefill attention, quantized GEMM/GEMV, sampling, norms, KV-cache, MoE, Mamba).

⚠️ All benchmarks, compiles and correctness checks are simulated analytically — no kernel was ever compiled or executed on a real GPU.

Row format (one row = one episode)

json
{
  "episode_id": "flash_decode_h100_B=8,..._dispatch_74",
  "scenario": "dispatch",            // dispatch | parallel_sweep | pipeline | campaign
  "n_agents": 6,
  "op": "flash_decode", "impl": "cuda",
  "gpu": "h100", "arch": "sm_90a", "dtype": "fp16",
  "shape": {...}, "baseline_ms": 0.42, "best_ms": 0.006,
  "speedup": 66.4, "elapsed_s": 17335,
  "agents": [{"id": "a01", "role": "worker", "skill": 0.9,
              "tries": 8, "kept": 5, "best_ms": 0.008}, ...],
  "num_messages": 74,
  "messages": [ ... ]
}

Message kinds

kindrolemeaning
planleadcampaign split strategy
dispatchleadsubtask assignment (to, subtask)
candidateagentfull kernel source + params
resultenvcompile/verify/bench outcome + reward
statusagentterse progress report
reviewreviewerpass/fail verdict
verifyverifiercorrectness check report
leaderboardenvranked agent best-times
mergeleadpick winner, update shared best
redirectleadre-assign agent to winner's variant/region
handoffagentpass artifact to another agent
verdictleadfinal outcome + terminal reward

Rewards

  • —result msgs: min(1, log2(prev_best / new)) on kept improvement; -0.05 no-improvement, -0.1 compile error, -0.3 wrong result, -0.05 unsupported arch.
  • —verdict msg: log2(speedup) / 4 terminal bonus.

Scenarios

  • —dispatch — lead splits variants across workers, merges, redirects losers (2–8 agents)
  • —parallel_sweep — N workers each own a cfg slice of one variant; leaderboard + top-2 refine (4–20)
  • —pipeline — proposer → reviewer → verifier → tuner chain over variants (3–6)
  • —campaign — fleet dispatch over variant×cfg space + verifier rebench (5–20)

Coverage

  • —17 ops (14 CUDA + 3 Triton impls): flashdecode, swadecode, gemmfp16, int8gemm, fp8gemm, int4gemv, rmsnorm, rope, swigluact, topksample, nucleussample, mambascan, moescatter, kvappend, fusedqkvrope, tritonflashdecode, triton_rmsnorm
  • —GPUs: h100, h200, a100, a10, rtxa6000, rtx5090 (arch features gated: wgmma/tma/fp8/fp4/cp.async)
  • —2400 episodes · ~170k messages · avg 71 msgs/episode · avg 8.6 agents
  • —Simulated horizon: 1–5 h per episode

Files

  • —batches/batch_*.jsonl — 24 shards × 100 episodes
  • —episodes.jsonl — all episodes concatenated
  • —tasks_holdout.jsonl — (op, gpu, shape) task specs for eval
  • —stats.json — aggregate stats