willychan21/ParallelKernelBench_Problems
ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Files Path Description data/problems.parquet One row per problem (tabular access) reference/*.py Reference solution() implementations utils/input_output_tensors.py Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.
0123
1from typing import Optional, Tuple2 3import torch4import torch.distributed as dist5 6 7def solution(8 grad_fc1_1_lora_A: torch.Tensor,9 grad_fc1_2_lora_A: torch.Tensor,10 grad_fc2_lora_B: torch.Tensor,11 group: Optional[dist.ProcessGroup] = None,12) -> Tuple[torch.Tensor, torch.Tensor, torch.Tensor]:13 group = group or dist.group.WORLD14 dist.all_reduce(grad_fc1_1_lora_A, op=dist.ReduceOp.SUM, group=group)15 dist.all_reduce(grad_fc1_2_lora_A, op=dist.ReduceOp.SUM, group=group)16 dist.all_reduce(grad_fc2_lora_B, op=dist.ReduceOp.SUM, group=group)17 return grad_fc1_1_lora_A, grad_fc1_2_lora_A, grad_fc2_lora_B18 