Team Ai
Datasetpublic

willychan21/ParallelKernelBench_Problems

ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Files Path Description data/problems.parquet One row per problem (tabular access) reference/*.py Reference solution() implementations utils/input_output_tensors.py Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes123downloads
15_combined_sharded_gemms.py33 linesDownload Raw Back to reference
1import torch2import torch.distributed as dist3import torch.nn.functional as F4 5 6@torch.no_grad()7def solution(8    x_local: torch.Tensor,9    W1: torch.Tensor,10    W2: torch.Tensor,11) -> torch.Tensor:12    rank = dist.get_rank()13    world_size = dist.get_world_size()14    M = x_local.shape[0]15    M_local = M // world_size16 17    x_local = x_local.contiguous()18    shards = [torch.empty_like(x_local) for _ in range(world_size)]19    dist.all_gather(shards, x_local)20    x_full = torch.cat(shards, dim=1)21 22    a = F.silu(torch.matmul(x_full, W1))23    a_loc = a[rank * M_local : (rank + 1) * M_local].contiguous()24    block = torch.matmul(a_loc, W2)25 26    H = block.shape[1]27    buf = block.new_zeros((M, H))28    buf[rank * M_local : (rank + 1) * M_local].copy_(block)29 30    y_local = block.new_empty((M_local, H))31    dist.reduce_scatter_tensor(y_local, buf, op=dist.ReduceOp.SUM)32    return y_local33