willychan21/ParallelKernelBench_Problems
ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Files Path Description data/problems.parquet One row per problem (tabular access) reference/*.py Reference solution() implementations utils/input_output_tensors.py Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.
0123
1from typing import Optional2 3import torch4import torch.distributed as dist5 6 7def solution(8 x: torch.Tensor,9 group: Optional[dist.ProcessGroup] = None,10) -> torch.Tensor:11 group = group or dist.group.WORLD12 world_size = dist.get_world_size(group)13 if world_size == 1:14 return x.contiguous()15 16 x = x.contiguous()17 dim_size = list(x.size())18 dim_size[0] = dim_size[0] * world_size19 output = torch.empty(dim_size, dtype=x.dtype, device=x.device)20 dist.all_gather_into_tensor(output, x, group=group)21 return output22 