Team Ai
Datasetpublic

togethercomputer/ParallelKernelBench_Problems

ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Inputs are deterministic — reproduce them with create_input_tensor(rank, world_size, problem_id, base_shape, dtype, trial) from that file; you do not need stored .pt files. Files Path Description… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/ParallelKernelBench_Problems.

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes177downloads
34_ulysses_all_gather_into_tensor_primitive.py22 linesDownload Raw Back to reference
1from typing import Optional2 3import torch4import torch.distributed as dist5 6 7def solution(8    x: torch.Tensor,9    group: Optional[dist.ProcessGroup] = None,10) -> torch.Tensor:11    group = group or dist.group.WORLD12    world_size = dist.get_world_size(group)13    if world_size == 1:14        return x.contiguous()15 16    x = x.contiguous()17    dim_size = list(x.size())18    dim_size[0] = dim_size[0] * world_size19    output = torch.empty(dim_size, dtype=x.dtype, device=x.device)20    dist.all_gather_into_tensor(output, x, group=group)21    return output22