Team Ai
Datasetpublic

willychan21/ParallelKernelBench_Problems

ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Files Path Description data/problems.parquet One row per problem (tabular access) reference/*.py Reference solution() implementations utils/input_output_tensors.py Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes123downloads
30_moe_epgroupgemm_lora_backward.py18 linesDownload Raw Back to reference
1from typing import Optional, Tuple2 3import torch4import torch.distributed as dist5 6 7def solution(8    grad_fc1_1_lora_A: torch.Tensor,9    grad_fc1_2_lora_A: torch.Tensor,10    grad_fc2_lora_B: torch.Tensor,11    group: Optional[dist.ProcessGroup] = None,12) -> Tuple[torch.Tensor, torch.Tensor, torch.Tensor]:13    group = group or dist.group.WORLD14    dist.all_reduce(grad_fc1_1_lora_A, op=dist.ReduceOp.SUM, group=group)15    dist.all_reduce(grad_fc1_2_lora_A, op=dist.ReduceOp.SUM, group=group)16    dist.all_reduce(grad_fc2_lora_B, op=dist.ReduceOp.SUM, group=group)17    return grad_fc1_1_lora_A, grad_fc1_2_lora_A, grad_fc2_lora_B18