willychan21/ParallelKernelBench_Problems
ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Files Path Description data/problems.parquet One row per problem (tabular access) reference/*.py Reference solution() implementations utils/input_output_tensors.py Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.
0123
1import torch2import torch.distributed as dist3 4 5@torch.no_grad()6def solution(A_local: torch.Tensor, B: torch.Tensor) -> torch.Tensor:7 world_size = dist.get_world_size()8 M, K_local = A_local.shape9 K = world_size * K_local10 11 A_local_t = A_local.transpose(0, 1).contiguous()12 A_t_buf = A_local_t.new_empty((world_size, K_local, M))13 dist.all_gather_into_tensor(A_t_buf, A_local_t)14 A_global_t = A_t_buf.reshape(K, M)15 16 C_t = torch.matmul(B.transpose(0, 1), A_global_t)17 return C_t.transpose(0, 1)18 