Brunobkr/llama.cpp_AlgMor24_github
ΩFFFΣLLIa • llama.cpp • AlgMor24 ██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗ ██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗ ██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║ ██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║ ╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝ High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.
03.1k
1#pragma once2 3#include "common.cuh"4#include "ggml-backend-impl.h"5 6#include <cstddef>7 8// Opaque pipeline context -- owns all pinned buffers, streams, and events.9struct ggml_cuda_ar_pipeline;10 11// Allocate a pipeline for n_devices GPUs.12// devices[] holds the CUDA device IDs in rank order.13// Returns nullptr on allocation failure.14ggml_cuda_ar_pipeline * ggml_cuda_ar_pipeline_init(15 const int * devices, size_t n_devices);16 17// Release all resources owned by the pipeline.18void ggml_cuda_ar_pipeline_free(ggml_cuda_ar_pipeline * pipeline);19 20// Execute an in-place AllReduce (sum) across tensors[0..n_devices-1].21// tensors[i] must live on the device managed by backends[i] and be22// contiguous F32, F16, or BF16.23// Preconditions are checked by the CUDA comm dispatcher before calling this.24// Returns true once the reduction work has been enqueued successfully.25bool ggml_cuda_ar_allreduce(26 ggml_cuda_ar_pipeline * pipeline,27 ggml_backend_t * backends,28 ggml_tensor ** tensors);29 30 