Brunobkr/llama.cpp_AlgMor24_github
ΩFFFΣLLIa • llama.cpp • AlgMor24 ██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗ ██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗ ██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║ ██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║ ╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝ High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.
03.1k
1#pragma once2 3#include "ggml.h"4#include "llama.h"5 6#include <vector>7 8enum common_params_fit_status {9 COMMON_PARAMS_FIT_STATUS_SUCCESS = 0, // found allocations that are projected to fit10 COMMON_PARAMS_FIT_STATUS_FAILURE = 1, // could not find allocations that are projected to fit11 COMMON_PARAMS_FIT_STATUS_ERROR = 2, // a hard error occurred, e.g. because no model could be found at the specified path12};13 14// fits mparams and cparams to free device memory (assumes system memory is unlimited)15// - returns true if the parameters could be successfully modified to fit device memory16// - this function is NOT thread safe because it modifies the global llama logger state17// - only parameters that have the same value as in llama_default_model_params are modified18// with the exception of the context size which is modified if and only if equal to 019common_params_fit_status common_fit_params(20 const char * path_model,21 llama_model_params * mparams,22 llama_context_params * cparams,23 float * tensor_split, // writable buffer for tensor split, needs at least llama_max_devices elements24 llama_model_tensor_buft_override * tensor_buft_overrides, // writable buffer for overrides, needs at least llama_max_tensor_buft_overrides elements25 size_t * margins, // margins of memory to leave per device in bytes26 uint32_t n_ctx_min, // minimum context size to set when trying to reduce memory use27 ggml_log_level log_level); // minimum log level to print during fitting, lower levels go to debug log28 29// print estimated memory to stdout30void common_fit_print(31 const char * path_model,32 llama_model_params * mparams,33 llama_context_params * cparams);34 35void common_memory_breakdown_print(const llama_context * ctx);36 37struct common_device_memory_data {38 int64_t total;39 int64_t free;40 size_t model;41 size_t context;42 size_t compute;43};44 45using common_device_memory_data_vec = std::vector<common_device_memory_data>;46 47// Load a model + context with no_alloc and return the per-device memory breakdown.48common_device_memory_data_vec common_get_device_memory_data(49 const char * path_model,50 const llama_model_params * mparams,51 const llama_context_params * cparams,52 std::vector<ggml_backend_dev_t> & devs,53 uint32_t & hp_ngl,54 uint32_t & hp_n_ctx_train,55 uint32_t & hp_n_expert,56 ggml_log_level log_level);57 