Brunobkr/llama.cpp_AlgMor24_github
ΩFFFΣLLIa • llama.cpp • AlgMor24 ██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗ ██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗ ██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║ ██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║ ╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝ High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.
03.1k
1#pragma once2 3//4// GGML Tensor Library5//6// This documentation is still a work in progress.7// If you wish some specific topics to be covered, feel free to drop a comment:8//9// https://github.com/ggml-org/whisper.cpp/issues/4010//11// ## Overview12//13// This library implements:14//15// - a set of tensor operations16// - automatic differentiation17// - basic optimization algorithms18//19// The aim of this library is to provide a minimalistic approach for various machine learning tasks. This includes,20// but is not limited to, the following:21//22// - linear regression23// - support vector machines24// - neural networks25//26// The library allows the user to define a certain function using the available tensor operations. This function27// definition is represented internally via a computation graph. Each tensor operation in the function definition28// corresponds to a node in the graph. Having the computation graph defined, the user can choose to compute the29// function's value and/or its gradient with respect to the input variables. Optionally, the function can be optimized30// using one of the available optimization algorithms.31//32// For example, here we define the function: f(x) = a*x^2 + b33//34// {35// struct ggml_init_params params = {36// .mem_size = 16*1024*1024,37// .mem_buffer = NULL,38// };39//40// // memory allocation happens here41// struct ggml_context * ctx = ggml_init(params);42//43// struct ggml_tensor * x = ggml_new_tensor_1d(ctx, GGML_TYPE_F32, 1);44//45// ggml_set_param(ctx, x); // x is an input variable46//47// struct ggml_tensor * a = ggml_new_tensor_1d(ctx, GGML_TYPE_F32, 1);48// struct ggml_tensor * b = ggml_new_tensor_1d(ctx, GGML_TYPE_F32, 1);49// struct ggml_tensor * x2 = ggml_mul(ctx, x, x);50// struct ggml_tensor * f = ggml_add(ctx, ggml_mul(ctx, a, x2), b);51//52// ...53// }54//55// Notice that the function definition above does not involve any actual computation. The computation is performed only56// when the user explicitly requests it. For example, to compute the function's value at x = 2.0:57//58// {59// ...60//61// struct ggml_cgraph * gf = ggml_new_graph(ctx);62// ggml_build_forward_expand(gf, f);63//64// // set the input variable and parameter values65// ggml_set_f32(x, 2.0f);66// ggml_set_f32(a, 3.0f);67// ggml_set_f32(b, 4.0f);68//69// ggml_graph_compute_with_ctx(ctx, &gf, n_threads);70//71// printf("f = %f\n", ggml_get_f32_1d(f, 0));72//73// ...74// }75//76// The actual computation is performed in the ggml_graph_compute() function.77//78// The ggml_new_tensor_...() functions create new tensors. They are allocated in the memory buffer provided to the79// ggml_init() function. You have to be careful not to exceed the memory buffer size. Therefore, you have to know80// in advance how much memory you need for your computation. Alternatively, you can allocate a large enough memory81// and after defining the computation graph, call the ggml_used_mem() function to find out how much memory was82// actually needed.83//84// The ggml_set_param() function marks a tensor as an input variable. This is used by the automatic85// differentiation and optimization algorithms.86//87// The described approach allows to define the function graph once and then compute its forward or backward graphs88// multiple times. All computations will use the same memory buffer allocated in the ggml_init() function. This way89// the user can avoid the memory allocation overhead at runtime.90//91// The library supports multi-dimensional tensors - up to 4 dimensions. The FP16 and FP32 data types are first class92// citizens, but in theory the library can be extended to support FP8 and integer data types.93//94// Each tensor operation produces a new tensor. Initially the library was envisioned to support only the use of unary95// and binary operations. Most of the available operations fall into one of these two categories. With time, it became96// clear that the library needs to support more complex operations. The way to support these operations is not clear97// yet, but a few examples are demonstrated in the following operations:98//99// - ggml_permute()100// - ggml_conv_1d_1s()101// - ggml_conv_1d_2s()102//103// For each tensor operator, the library implements a forward and backward computation function. The forward function104// computes the output tensor value given the input tensor values. The backward function computes the adjoint of the105// input tensors given the adjoint of the output tensor. For a detailed explanation of what this means, take a106// calculus class, or watch the following video:107//108// What is Automatic Differentiation?109// https://www.youtube.com/watch?v=wG_nF1awSSY110//111//112// ## Tensor data (struct ggml_tensor)113//114// The tensors are stored in memory via the ggml_tensor struct. The structure provides information about the size of115// the tensor, the data type, and the memory buffer where the tensor data is stored. Additionally, it contains116// pointers to the "source" tensors - i.e. the tensors that were used to compute the current tensor. For example:117//118// {119// struct ggml_tensor * c = ggml_add(ctx, a, b);120//121// assert(c->src[0] == a);122// assert(c->src[1] == b);123// }124//125// The multi-dimensional tensors are stored in row-major order. The ggml_tensor struct contains fields for the126// number of elements in each dimension ("ne") as well as the number of bytes ("nb", a.k.a. stride). This allows127// to store tensors that are not contiguous in memory, which is useful for operations such as transposition and128// permutation. All tensor operations have to take the stride into account and not assume that the tensor is129// contiguous in memory.130//131// The data of the tensor is accessed via the "data" pointer. For example:132//133// {134// const int nx = 2;135// const int ny = 3;136//137// struct ggml_tensor * a = ggml_new_tensor_2d(ctx, GGML_TYPE_F32, nx, ny);138//139// for (int y = 0; y < ny; y++) {140// for (int x = 0; x < nx; x++) {141// *(float *) ((char *) a->data + y*a->nb[1] + x*a->nb[0]) = x + y;142// }143// }144//145// ...146// }147//148// Alternatively, there are helper functions, such as ggml_get_f32_1d() and ggml_set_f32_1d() that can be used.149//150// ## The matrix multiplication operator (ggml_mul_mat)151//152// TODO153//154//155// ## Multi-threading156//157// TODO158//159//160// ## Overview of ggml.c161//162// TODO163//164//165// ## SIMD optimizations166//167// TODO168//169//170// ## Debugging ggml171//172// TODO173//174//175 176#ifdef GGML_SHARED177# if defined(_WIN32) && !defined(__MINGW32__)178# ifdef GGML_BUILD179# define GGML_API __declspec(dllexport) extern180# else181# define GGML_API __declspec(dllimport) extern182# endif183# else184# define GGML_API __attribute__ ((visibility ("default"))) extern185# endif186#else187# define GGML_API extern188#endif189 190// TODO: support for clang191#ifdef __GNUC__192# define GGML_DEPRECATED(func, hint) func __attribute__((deprecated(hint)))193#elif defined(_MSC_VER)194# define GGML_DEPRECATED(func, hint) __declspec(deprecated(hint)) func195#else196# define GGML_DEPRECATED(func, hint) func197#endif198 199#ifndef __GNUC__200# define GGML_ATTRIBUTE_FORMAT(...)201#elif defined(__MINGW32__) && !defined(__clang__)202# define GGML_ATTRIBUTE_FORMAT(...) __attribute__((format(gnu_printf, __VA_ARGS__)))203#else204# define GGML_ATTRIBUTE_FORMAT(...) __attribute__((format(printf, __VA_ARGS__)))205#endif206 207#if defined(_WIN32) && !defined(_WIN32_WINNT)208# define _WIN32_WINNT 0x0A00209#endif210 211#include <stdbool.h>212#include <stddef.h>213#include <stdint.h>214#include <stdio.h>215 216#define GGML_FILE_MAGIC 0x67676d6c // "ggml"217#define GGML_FILE_VERSION 2218 219#define GGML_QNT_VERSION 2 // bump this on quantization format changes220#define GGML_QNT_VERSION_FACTOR 1000 // do not change this221 222#define GGML_MAX_DIMS 4223#define GGML_MAX_PARAMS 2048224#define GGML_MAX_SRC 10225#define GGML_MAX_N_THREADS 512226#define GGML_MAX_OP_PARAMS 64227 228#ifndef GGML_MAX_NAME229# define GGML_MAX_NAME 64230#endif231 232#define GGML_DEFAULT_N_THREADS 4233#define GGML_DEFAULT_GRAPH_SIZE 2048234 235#if UINTPTR_MAX == 0xFFFFFFFF236 #define GGML_MEM_ALIGN 4237#elif defined(__EMSCRIPTEN__)238// emscripten uses max_align_t == 8, so we need GGML_MEM_ALIGN == 8 for 64-bit wasm.239// (for 32-bit wasm, the first conditional is true and GGML_MEM_ALIGN stays 4.)240// ref: https://github.com/ggml-org/llama.cpp/pull/18628241 #define GGML_MEM_ALIGN 8242#else243 #define GGML_MEM_ALIGN 16244#endif245 246#define GGML_EXIT_SUCCESS 0247#define GGML_EXIT_ABORTED 1248 249// TODO: convert to enum https://github.com/ggml-org/llama.cpp/pull/16187#discussion_r2388538726250#define GGML_ROPE_TYPE_NORMAL 0251#define GGML_ROPE_TYPE_NEOX 2252#define GGML_ROPE_TYPE_MROPE 8253#define GGML_ROPE_TYPE_VISION 24254#define GGML_ROPE_TYPE_IMROPE 40 // binary: 101000255 256#define GGML_MROPE_SECTIONS 4257 258#define GGML_UNUSED(x) (void)(x)259#ifdef __CUDACC__260template<typename... Args>261__host__ __device__ constexpr inline void ggml_unused_vars_impl(Args&&...) noexcept {}262#define GGML_UNUSED_VARS(...) ggml_unused_vars_impl(__VA_ARGS__)263#else264#define GGML_UNUSED_VARS(...) do { (void)sizeof((__VA_ARGS__, 0)); } while(0)265#endif // __CUDACC__266 267#define GGML_PAD(x, n) (((x) + (n) - 1) & ~((n) - 1))268 269#ifndef NDEBUG270# define GGML_UNREACHABLE() do { fprintf(stderr, "statement should be unreachable\n"); abort(); } while(0)271#elif defined(__GNUC__)272# define GGML_UNREACHABLE() __builtin_unreachable()273#elif defined(_MSC_VER)274# define GGML_UNREACHABLE() __assume(0)275#else276# define GGML_UNREACHABLE() ((void) 0)277#endif278 279#ifdef __cplusplus280# define GGML_NORETURN [[noreturn]]281#elif defined(_MSC_VER)282# define GGML_NORETURN __declspec(noreturn)283#else284# define GGML_NORETURN _Noreturn285#endif286 287#define GGML_ABORT(...) ggml_abort(__FILE__, __LINE__, __VA_ARGS__)288#define GGML_ASSERT(x) if (!(x)) GGML_ABORT("GGML_ASSERT(%s) failed", #x)289 290// used to copy the number of elements and stride in bytes of tensors into local variables.291// main purpose is to reduce code duplication and improve readability.292//293// example:294//295// GGML_TENSOR_LOCALS(int64_t, ne1, src1, ne);296// GGML_TENSOR_LOCALS(size_t, nb1, src1, nb);297//298#define GGML_TENSOR_LOCALS_1(type, prefix, pointer, array) \299 const type prefix##0 = (pointer) ? (pointer)->array[0] : 0; \300 GGML_UNUSED(prefix##0);301#define GGML_TENSOR_LOCALS_2(type, prefix, pointer, array) \302 GGML_TENSOR_LOCALS_1 (type, prefix, pointer, array) \303 const type prefix##1 = (pointer) ? (pointer)->array[1] : 0; \304 GGML_UNUSED(prefix##1);305#define GGML_TENSOR_LOCALS_3(type, prefix, pointer, array) \306 GGML_TENSOR_LOCALS_2 (type, prefix, pointer, array) \307 const type prefix##2 = (pointer) ? (pointer)->array[2] : 0; \308 GGML_UNUSED(prefix##2);309#define GGML_TENSOR_LOCALS(type, prefix, pointer, array) \310 GGML_TENSOR_LOCALS_3 (type, prefix, pointer, array) \311 const type prefix##3 = (pointer) ? (pointer)->array[3] : 0; \312 GGML_UNUSED(prefix##3);313 314#define GGML_TENSOR_UNARY_OP_LOCALS \315 GGML_TENSOR_LOCALS(int64_t, ne0, src0, ne) \316 GGML_TENSOR_LOCALS(size_t, nb0, src0, nb) \317 GGML_TENSOR_LOCALS(int64_t, ne, dst, ne) \318 GGML_TENSOR_LOCALS(size_t, nb, dst, nb)319 320#define GGML_TENSOR_BINARY_OP_LOCALS \321 GGML_TENSOR_LOCALS(int64_t, ne0, src0, ne) \322 GGML_TENSOR_LOCALS(size_t, nb0, src0, nb) \323 GGML_TENSOR_LOCALS(int64_t, ne1, src1, ne) \324 GGML_TENSOR_LOCALS(size_t, nb1, src1, nb) \325 GGML_TENSOR_LOCALS(int64_t, ne, dst, ne) \326 GGML_TENSOR_LOCALS(size_t, nb, dst, nb)327 328#define GGML_TENSOR_TERNARY_OP_LOCALS \329 GGML_TENSOR_LOCALS(int64_t, ne0, src0, ne) \330 GGML_TENSOR_LOCALS(size_t, nb0, src0, nb) \331 GGML_TENSOR_LOCALS(int64_t, ne1, src1, ne) \332 GGML_TENSOR_LOCALS(size_t, nb1, src1, nb) \333 GGML_TENSOR_LOCALS(int64_t, ne2, src2, ne) \334 GGML_TENSOR_LOCALS(size_t, nb2, src2, nb) \335 GGML_TENSOR_LOCALS(int64_t, ne, dst, ne) \336 GGML_TENSOR_LOCALS(size_t, nb, dst, nb)337 338#define GGML_TENSOR_BINARY_OP_LOCALS01 \339 GGML_TENSOR_LOCALS(int64_t, ne0, src0, ne) \340 GGML_TENSOR_LOCALS(size_t, nb0, src0, nb) \341 GGML_TENSOR_LOCALS(int64_t, ne1, src1, ne) \342 GGML_TENSOR_LOCALS(size_t, nb1, src1, nb)343 344#ifdef __cplusplus345extern "C" {346#endif347 348 // Function type used in fatal error callbacks349 typedef void (*ggml_abort_callback_t)(const char * error_message);350 351 // Set the abort callback (passing null will restore original abort functionality: printing a message to stdout)352 // Returns the old callback for chaining353 GGML_API ggml_abort_callback_t ggml_set_abort_callback(ggml_abort_callback_t callback);354 355 GGML_NORETURN GGML_ATTRIBUTE_FORMAT(3, 4)356 GGML_API void ggml_abort(const char * file, int line, const char * fmt, ...);357 358 enum ggml_status {359 GGML_STATUS_ALLOC_FAILED = -2,360 GGML_STATUS_FAILED = -1,361 GGML_STATUS_SUCCESS = 0,362 GGML_STATUS_ABORTED = 1,363 };364 365 // get ggml_status name string366 GGML_API const char * ggml_status_to_string(enum ggml_status status);367 368 // ieee 754-2008 half-precision float16369 // todo: make this not an integral type370 typedef uint16_t ggml_fp16_t;371 GGML_API float ggml_fp16_to_fp32(ggml_fp16_t);372 GGML_API ggml_fp16_t ggml_fp32_to_fp16(float);373 GGML_API void ggml_fp16_to_fp32_row(const ggml_fp16_t *, float *, int64_t);374 GGML_API void ggml_fp32_to_fp16_row(const float *, ggml_fp16_t *, int64_t);375 376 // google brain half-precision bfloat16377 typedef struct { uint16_t bits; } ggml_bf16_t;378 GGML_API ggml_bf16_t ggml_fp32_to_bf16(float);379 GGML_API float ggml_bf16_to_fp32(ggml_bf16_t); // consider just doing << 16380 GGML_API void ggml_bf16_to_fp32_row(const ggml_bf16_t *, float *, int64_t);381 GGML_API void ggml_fp32_to_bf16_row_ref(const float *, ggml_bf16_t *, int64_t);382 GGML_API void ggml_fp32_to_bf16_row(const float *, ggml_bf16_t *, int64_t);383 384 struct ggml_object;385 struct ggml_context;386 struct ggml_cgraph;387 388 // NOTE: always add types at the end of the enum to keep backward compatibility389 enum ggml_type {390 GGML_TYPE_F32 = 0,391 GGML_TYPE_F16 = 1,392 GGML_TYPE_Q4_0 = 2,393 GGML_TYPE_Q4_1 = 3,394 // GGML_TYPE_Q4_2 = 4, support has been removed395 // GGML_TYPE_Q4_3 = 5, support has been removed396 GGML_TYPE_Q5_0 = 6,397 GGML_TYPE_Q5_1 = 7,398 GGML_TYPE_Q8_0 = 8,399 GGML_TYPE_Q8_1 = 9,400 GGML_TYPE_Q2_K = 10,401 GGML_TYPE_Q3_K = 11,402 GGML_TYPE_Q4_K = 12,403 GGML_TYPE_Q5_K = 13,404 GGML_TYPE_Q6_K = 14,405 GGML_TYPE_Q8_K = 15,406 GGML_TYPE_IQ2_XXS = 16,407 GGML_TYPE_IQ2_XS = 17,408 GGML_TYPE_IQ3_XXS = 18,409 GGML_TYPE_IQ1_S = 19,410 GGML_TYPE_IQ4_NL = 20,411 GGML_TYPE_IQ3_S = 21,412 GGML_TYPE_IQ2_S = 22,413 GGML_TYPE_IQ4_XS = 23,414 GGML_TYPE_I8 = 24,415 GGML_TYPE_I16 = 25,416 GGML_TYPE_I32 = 26,417 GGML_TYPE_I64 = 27,418 GGML_TYPE_F64 = 28,419 GGML_TYPE_IQ1_M = 29,420 GGML_TYPE_BF16 = 30,421 // GGML_TYPE_Q4_0_4_4 = 31, support has been removed from gguf files422 // GGML_TYPE_Q4_0_4_8 = 32,423 // GGML_TYPE_Q4_0_8_8 = 33,424 GGML_TYPE_TQ1_0 = 34,425 GGML_TYPE_TQ2_0 = 35,426 // GGML_TYPE_IQ4_NL_4_4 = 36,427 // GGML_TYPE_IQ4_NL_4_8 = 37,428 // GGML_TYPE_IQ4_NL_8_8 = 38,429 GGML_TYPE_MXFP4 = 39, // MXFP4 (1 block)430 GGML_TYPE_NVFP4 = 40, // NVFP4 (4 blocks, E4M3 scale)431 GGML_TYPE_Q1_0 = 41,432 GGML_TYPE_Q2_0 = 42,433 GGML_TYPE_COUNT = 43,434 };435 436 // precision437 enum ggml_prec {438 GGML_PREC_DEFAULT = 0, // stored as ggml_tensor.op_params, 0 by default439 GGML_PREC_F32 = 10,440 };441 442 // op hint443 enum ggml_op_hint {444 GGML_HINT_NONE = 0,445 GGML_HINT_SRC0_IS_HADAMARD = 1,446 };447 448 // model file types449 enum ggml_ftype {450 GGML_FTYPE_UNKNOWN = -1,451 GGML_FTYPE_ALL_F32 = 0,452 GGML_FTYPE_MOSTLY_F16 = 1, // except 1d tensors453 GGML_FTYPE_MOSTLY_Q4_0 = 2, // except 1d tensors454 GGML_FTYPE_MOSTLY_Q4_1 = 3, // except 1d tensors455 GGML_FTYPE_MOSTLY_Q4_1_SOME_F16 = 4, // tok_embeddings.weight and output.weight are F16456 GGML_FTYPE_MOSTLY_Q8_0 = 7, // except 1d tensors457 GGML_FTYPE_MOSTLY_Q5_0 = 8, // except 1d tensors458 GGML_FTYPE_MOSTLY_Q5_1 = 9, // except 1d tensors459 GGML_FTYPE_MOSTLY_Q2_K = 10, // except 1d tensors460 GGML_FTYPE_MOSTLY_Q3_K = 11, // except 1d tensors461 GGML_FTYPE_MOSTLY_Q4_K = 12, // except 1d tensors462 GGML_FTYPE_MOSTLY_Q5_K = 13, // except 1d tensors463 GGML_FTYPE_MOSTLY_Q6_K = 14, // except 1d tensors464 GGML_FTYPE_MOSTLY_IQ2_XXS = 15, // except 1d tensors465 GGML_FTYPE_MOSTLY_IQ2_XS = 16, // except 1d tensors466 GGML_FTYPE_MOSTLY_IQ3_XXS = 17, // except 1d tensors467 GGML_FTYPE_MOSTLY_IQ1_S = 18, // except 1d tensors468 GGML_FTYPE_MOSTLY_IQ4_NL = 19, // except 1d tensors469 GGML_FTYPE_MOSTLY_IQ3_S = 20, // except 1d tensors470 GGML_FTYPE_MOSTLY_IQ2_S = 21, // except 1d tensors471 GGML_FTYPE_MOSTLY_IQ4_XS = 22, // except 1d tensors472 GGML_FTYPE_MOSTLY_IQ1_M = 23, // except 1d tensors473 GGML_FTYPE_MOSTLY_BF16 = 24, // except 1d tensors474 GGML_FTYPE_MOSTLY_MXFP4 = 25, // except 1d tensors475 GGML_FTYPE_MOSTLY_NVFP4 = 26, // except 1d tensors476 GGML_FTYPE_MOSTLY_Q1_0 = 27, // except 1d tensors477 GGML_FTYPE_MOSTLY_Q2_0 = 28, // except 1d tensors478 };479 480 // available tensor operations:481 enum ggml_op {482 GGML_OP_NONE = 0,483 484 GGML_OP_DUP,485 GGML_OP_ADD,486 GGML_OP_ADD_ID,487 GGML_OP_ADD1,488 GGML_OP_ACC,489 GGML_OP_SUB,490 GGML_OP_MUL,491 GGML_OP_DIV,492 GGML_OP_SQR,493 GGML_OP_SQRT,494 GGML_OP_LOG,495 GGML_OP_SIN,496 GGML_OP_COS,497 GGML_OP_SUM,498 GGML_OP_SUM_ROWS,499 GGML_OP_CUMSUM,500 GGML_OP_MEAN,501 GGML_OP_ARGMAX,502 GGML_OP_COUNT_EQUAL,503 GGML_OP_REPEAT,504 GGML_OP_REPEAT_BACK,505 GGML_OP_CONCAT,506 GGML_OP_SILU_BACK,507 GGML_OP_NORM, // normalize508 GGML_OP_RMS_NORM,509 GGML_OP_RMS_NORM_BACK,510 GGML_OP_GROUP_NORM,511 GGML_OP_L2_NORM,512 513 GGML_OP_MUL_MAT,514 GGML_OP_MUL_MAT_ID,515 GGML_OP_OUT_PROD,516 517 GGML_OP_SCALE,518 GGML_OP_SET,519 GGML_OP_CPY,520 GGML_OP_CONT,521 GGML_OP_RESHAPE,522 GGML_OP_VIEW,523 GGML_OP_PERMUTE,524 GGML_OP_TRANSPOSE,525 GGML_OP_GET_ROWS,526 GGML_OP_GET_ROWS_BACK,527 GGML_OP_SET_ROWS,528 GGML_OP_DIAG,529 GGML_OP_DIAG_MASK_INF,530 GGML_OP_DIAG_MASK_ZERO,531 GGML_OP_SOFT_MAX,532 GGML_OP_SOFT_MAX_BACK,533 GGML_OP_ROPE,534 GGML_OP_ROPE_BACK,535 GGML_OP_CLAMP,536 GGML_OP_CONV_TRANSPOSE_1D,537 GGML_OP_IM2COL,538 GGML_OP_IM2COL_BACK,539 GGML_OP_IM2COL_3D,540 GGML_OP_COL2IM_1D,541 GGML_OP_CONV_2D,542 GGML_OP_CONV_3D,543 GGML_OP_CONV_2D_DW,544 GGML_OP_CONV_TRANSPOSE_2D,545 GGML_OP_POOL_1D,546 GGML_OP_POOL_2D,547 GGML_OP_POOL_2D_BACK,548 GGML_OP_UPSCALE,549 GGML_OP_PAD,550 GGML_OP_PAD_REFLECT_1D,551 GGML_OP_ROLL,552 GGML_OP_ARANGE,553 GGML_OP_TIMESTEP_EMBEDDING,554 GGML_OP_ARGSORT,555 GGML_OP_TOP_K,556 GGML_OP_LEAKY_RELU,557 GGML_OP_TRI,558 GGML_OP_FILL,559 560 GGML_OP_FLASH_ATTN_EXT,561 GGML_OP_FLASH_ATTN_BACK,562 GGML_OP_SSM_CONV,563 GGML_OP_SSM_SCAN,564 GGML_OP_WIN_PART,565 GGML_OP_WIN_UNPART,566 GGML_OP_GET_REL_POS,567 GGML_OP_ADD_REL_POS,568 GGML_OP_RWKV_WKV6,569 GGML_OP_GATED_LINEAR_ATTN,570 GGML_OP_RWKV_WKV7,571 GGML_OP_SOLVE_TRI,572 GGML_OP_GATED_DELTA_NET,573 GGML_OP_LIGHTNING_INDEXER,574 GGML_OP_DSV4_HC_COMB,575 GGML_OP_DSV4_HC_PRE,576 GGML_OP_DSV4_HC_POST,577 578 GGML_OP_UNARY,579 580 GGML_OP_MAP_CUSTOM1,581 GGML_OP_MAP_CUSTOM2,582 GGML_OP_MAP_CUSTOM3,583 584 GGML_OP_CUSTOM,585 586 GGML_OP_CROSS_ENTROPY_LOSS,587 GGML_OP_CROSS_ENTROPY_LOSS_BACK,588 GGML_OP_OPT_STEP_ADAMW,589 GGML_OP_OPT_STEP_SGD,590 591 GGML_OP_GLU,592 593 GGML_OP_COUNT,594 };595 596 enum ggml_unary_op {597 GGML_UNARY_OP_ABS,598 GGML_UNARY_OP_SGN,599 GGML_UNARY_OP_NEG,600 GGML_UNARY_OP_STEP,601 GGML_UNARY_OP_TANH,602 GGML_UNARY_OP_ELU,603 GGML_UNARY_OP_RELU,604 GGML_UNARY_OP_SIGMOID,605 GGML_UNARY_OP_GELU,606 GGML_UNARY_OP_GELU_QUICK,607 GGML_UNARY_OP_SILU,608 GGML_UNARY_OP_HARDSWISH,609 GGML_UNARY_OP_HARDSIGMOID,610 GGML_UNARY_OP_EXP,611 GGML_UNARY_OP_EXPM1,612 GGML_UNARY_OP_SOFTPLUS,613 GGML_UNARY_OP_GELU_ERF,614 GGML_UNARY_OP_XIELU,615 GGML_UNARY_OP_FLOOR,616 GGML_UNARY_OP_CEIL,617 GGML_UNARY_OP_ROUND,618 GGML_UNARY_OP_TRUNC,619 620 GGML_UNARY_OP_COUNT,621 };622 623 enum ggml_glu_op {624 GGML_GLU_OP_REGLU,625 GGML_GLU_OP_GEGLU,626 GGML_GLU_OP_SWIGLU,627 GGML_GLU_OP_SWIGLU_OAI,628 GGML_GLU_OP_GEGLU_ERF,629 GGML_GLU_OP_GEGLU_QUICK,630 631 GGML_GLU_OP_COUNT,632 };633 634 enum ggml_object_type {635 GGML_OBJECT_TYPE_TENSOR,636 GGML_OBJECT_TYPE_GRAPH,637 GGML_OBJECT_TYPE_WORK_BUFFER638 };639 640 enum ggml_log_level {641 GGML_LOG_LEVEL_NONE = 0,642 GGML_LOG_LEVEL_DEBUG = 1,643 GGML_LOG_LEVEL_INFO = 2,644 GGML_LOG_LEVEL_WARN = 3,645 GGML_LOG_LEVEL_ERROR = 4,646 GGML_LOG_LEVEL_CONT = 5, // continue previous log647 };648 649 // this tensor...650 enum ggml_tensor_flag {651 GGML_TENSOR_FLAG_INPUT = 1, // ...is an input for the GGML compute graph652 GGML_TENSOR_FLAG_OUTPUT = 2, // ...is an output for the GGML compute graph653 GGML_TENSOR_FLAG_PARAM = 4, // ...contains trainable parameters654 GGML_TENSOR_FLAG_LOSS = 8, // ...defines loss for numerical optimization (multiple loss tensors add up)655 GGML_TENSOR_FLAG_COMPUTE = 16, // ...must be computed656 };657 658 enum ggml_tri_type {659 GGML_TRI_TYPE_UPPER_DIAG = 0,660 GGML_TRI_TYPE_UPPER = 1,661 GGML_TRI_TYPE_LOWER_DIAG = 2,662 GGML_TRI_TYPE_LOWER = 3663 };664 665 struct ggml_init_params {666 // memory pool667 size_t mem_size; // bytes668 void * mem_buffer; // if NULL, memory will be allocated internally669 bool no_alloc; // don't allocate memory for the tensor data670 };671 672 // n-dimensional tensor673 struct ggml_tensor {674 enum ggml_type type;675 676 struct ggml_backend_buffer * buffer;677 678 int64_t ne[GGML_MAX_DIMS]; // number of elements679 size_t nb[GGML_MAX_DIMS]; // stride in bytes:680 // nb[0] = ggml_type_size(type)681 // nb[1] = nb[0] * (ne[0] / ggml_blck_size(type)) + padding682 // nb[i] = nb[i-1] * ne[i-1]683 684 // compute data685 enum ggml_op op;686 687 // op params - allocated as int32_t for alignment688 int32_t op_params[GGML_MAX_OP_PARAMS / sizeof(int32_t)];689 690 int32_t flags;691 692 struct ggml_tensor * src[GGML_MAX_SRC];693 694 // source tensor and offset for views695 struct ggml_tensor * view_src;696 size_t view_offs;697 698 void * data;699 700 char name[GGML_MAX_NAME];701 702 void * extra; // extra things e.g. for ggml-cuda.cu703 704 char padding[8];705 };706 707 static const size_t GGML_TENSOR_SIZE = sizeof(struct ggml_tensor);708 709 // Abort callback710 // If not NULL, called before ggml computation711 // If it returns true, the computation is aborted712 typedef bool (*ggml_abort_callback)(void * data);713 714 715 //716 // GUID717 //718 719 // GUID types720 typedef uint8_t ggml_guid[16];721 typedef ggml_guid * ggml_guid_t;722 723 GGML_API bool ggml_guid_matches(ggml_guid_t guid_a, ggml_guid_t guid_b);724 725 // misc726 727 GGML_API const char * ggml_version(void);728 GGML_API const char * ggml_commit(void);729 730 GGML_API void ggml_time_init(void); // call this once at the beginning of the program731 GGML_API int64_t ggml_time_ms(void);732 GGML_API int64_t ggml_time_us(void);733 GGML_API int64_t ggml_cycles(void);734 GGML_API int64_t ggml_cycles_per_ms(void);735 736 // accepts a UTF-8 path, even on Windows737 GGML_API FILE * ggml_fopen(const char * fname, const char * mode);738 739 GGML_API void ggml_print_object (const struct ggml_object * obj);740 GGML_API void ggml_print_objects(const struct ggml_context * ctx);741 742 GGML_API int64_t ggml_nelements (const struct ggml_tensor * tensor);743 GGML_API int64_t ggml_nrows (const struct ggml_tensor * tensor);744 GGML_API size_t ggml_nbytes (const struct ggml_tensor * tensor);745 GGML_API size_t ggml_nbytes_pad(const struct ggml_tensor * tensor); // same as ggml_nbytes() but padded to GGML_MEM_ALIGN746 747 GGML_API int64_t ggml_blck_size(enum ggml_type type);748 GGML_API size_t ggml_type_size(enum ggml_type type); // size in bytes for all elements in a block749 GGML_API size_t ggml_row_size (enum ggml_type type, int64_t ne); // size in bytes for all elements in a row750 751 GGML_DEPRECATED(752 GGML_API double ggml_type_sizef(enum ggml_type type), // ggml_type_size()/ggml_blck_size() as float753 "use ggml_row_size() instead");754 755 GGML_API const char * ggml_type_name(enum ggml_type type);756 GGML_API const char * ggml_op_name (enum ggml_op op);757 GGML_API const char * ggml_op_symbol(enum ggml_op op);758 759 GGML_API const char * ggml_unary_op_name(enum ggml_unary_op op);760 GGML_API const char * ggml_glu_op_name(enum ggml_glu_op op);761 GGML_API const char * ggml_op_desc(const struct ggml_tensor * t); // unary or op name762 763 GGML_API size_t ggml_element_size(const struct ggml_tensor * tensor);764 765 GGML_API bool ggml_is_quantized(enum ggml_type type);766 767 // TODO: temporary until model loading of ggml examples is refactored768 GGML_API enum ggml_type ggml_ftype_to_ggml_type(enum ggml_ftype ftype);769 770 GGML_API bool ggml_is_transposed(const struct ggml_tensor * tensor);771 GGML_API bool ggml_is_permuted (const struct ggml_tensor * tensor);772 GGML_API bool ggml_is_empty (const struct ggml_tensor * tensor);773 GGML_API bool ggml_is_view (const struct ggml_tensor * tensor);774 GGML_API bool ggml_is_scalar (const struct ggml_tensor * tensor);775 GGML_API bool ggml_is_vector (const struct ggml_tensor * tensor);776 GGML_API bool ggml_is_matrix (const struct ggml_tensor * tensor);777 GGML_API bool ggml_is_3d (const struct ggml_tensor * tensor);778 GGML_API int ggml_n_dims (const struct ggml_tensor * tensor); // returns 1 for scalars779 780 // returns whether the tensor elements can be iterated over with a flattened index (no gaps, no permutation)781 GGML_API bool ggml_is_contiguous (const struct ggml_tensor * tensor);782 GGML_API bool ggml_is_contiguous_0(const struct ggml_tensor * tensor); // same as ggml_is_contiguous()783 GGML_API bool ggml_is_contiguous_1(const struct ggml_tensor * tensor); // contiguous for dims >= 1784 GGML_API bool ggml_is_contiguous_2(const struct ggml_tensor * tensor); // contiguous for dims >= 2785 786 GGML_API bool ggml_is_contiguous_to_1(const struct ggml_tensor * tensor); // contiguous for dims < 1787 GGML_API bool ggml_is_contiguous_to_2(const struct ggml_tensor * tensor); // contiguous for dims < 2788 GGML_API bool ggml_is_contiguous_to_3(const struct ggml_tensor * tensor); // contiguous for dims < 3789 790 // returns whether the tensor elements are allocated as one contiguous block of memory (no gaps, but permutation ok)791 GGML_API bool ggml_is_contiguously_allocated(const struct ggml_tensor * tensor);792 793 // true for tensor that is stored in memory as CxWxHxN and has been permuted to WxHxCxN794 GGML_API bool ggml_is_contiguous_channels(const struct ggml_tensor * tensor);795 796 // true if the elements in dimension 0 are contiguous, or there is just 1 block of elements797 GGML_API bool ggml_is_contiguous_rows(const struct ggml_tensor * tensor);798 799 GGML_API bool ggml_are_same_shape (const struct ggml_tensor * t0, const struct ggml_tensor * t1);800 GGML_API bool ggml_are_same_stride(const struct ggml_tensor * t0, const struct ggml_tensor * t1);801 802 GGML_API bool ggml_can_repeat(const struct ggml_tensor * t0, const struct ggml_tensor * t1);803 804 // use this to compute the memory overhead of a tensor805 GGML_API size_t ggml_tensor_overhead(void);806 807 GGML_API bool ggml_validate_row_data(enum ggml_type type, const void * data, size_t nbytes);808 809 // main810 811 GGML_API struct ggml_context * ggml_init (struct ggml_init_params params);812 GGML_API void ggml_reset(struct ggml_context * ctx);813 GGML_API void ggml_free (struct ggml_context * ctx);814 815 GGML_API size_t ggml_used_mem(const struct ggml_context * ctx);816 817 GGML_API bool ggml_get_no_alloc(struct ggml_context * ctx);818 GGML_API void ggml_set_no_alloc(struct ggml_context * ctx, bool no_alloc);819 820 GGML_API void * ggml_get_mem_buffer (const struct ggml_context * ctx);821 GGML_API size_t ggml_get_mem_size (const struct ggml_context * ctx);822 GGML_API size_t ggml_get_max_tensor_size(const struct ggml_context * ctx);823 824 GGML_API struct ggml_tensor * ggml_new_tensor(825 struct ggml_context * ctx,826 enum ggml_type type,827 int n_dims,828 const int64_t *ne);829 830 GGML_API struct ggml_tensor * ggml_new_tensor_1d(831 struct ggml_context * ctx,832 enum ggml_type type,833 int64_t ne0);834 835 GGML_API struct ggml_tensor * ggml_new_tensor_2d(836 struct ggml_context * ctx,837 enum ggml_type type,838 int64_t ne0,839 int64_t ne1);840 841 GGML_API struct ggml_tensor * ggml_new_tensor_3d(842 struct ggml_context * ctx,843 enum ggml_type type,844 int64_t ne0,845 int64_t ne1,846 int64_t ne2);847 848 GGML_API struct ggml_tensor * ggml_new_tensor_4d(849 struct ggml_context * ctx,850 enum ggml_type type,851 int64_t ne0,852 int64_t ne1,853 int64_t ne2,854 int64_t ne3);855 856 GGML_API void * ggml_new_buffer(struct ggml_context * ctx, size_t nbytes);857 858 GGML_API struct ggml_tensor * ggml_dup_tensor (struct ggml_context * ctx, const struct ggml_tensor * src);859 GGML_API struct ggml_tensor * ggml_view_tensor(struct ggml_context * ctx, struct ggml_tensor * src);860 861 // Context tensor enumeration and lookup862 GGML_API struct ggml_tensor * ggml_get_first_tensor(const struct ggml_context * ctx);863 GGML_API struct ggml_tensor * ggml_get_next_tensor (const struct ggml_context * ctx, struct ggml_tensor * tensor);864 GGML_API struct ggml_tensor * ggml_get_tensor(struct ggml_context * ctx, const char * name);865 866 // Converts a flat index into coordinates867 GGML_API void ggml_unravel_index(const struct ggml_tensor * tensor, int64_t i, int64_t * i0, int64_t * i1, int64_t * i2, int64_t * i3);868 869 GGML_API enum ggml_unary_op ggml_get_unary_op(const struct ggml_tensor * tensor);870 GGML_API enum ggml_glu_op ggml_get_glu_op(const struct ggml_tensor * tensor);871 872 GGML_API void * ggml_get_data (const struct ggml_tensor * tensor);873 GGML_API float * ggml_get_data_f32(const struct ggml_tensor * tensor);874 875 GGML_API const char * ggml_get_name (const struct ggml_tensor * tensor);876 GGML_API struct ggml_tensor * ggml_set_name ( struct ggml_tensor * tensor, const char * name);877 GGML_ATTRIBUTE_FORMAT(2, 3)878 GGML_API struct ggml_tensor * ggml_format_name( struct ggml_tensor * tensor, const char * fmt, ...);879 880 // Tensor flags881 GGML_API void ggml_set_input(struct ggml_tensor * tensor);882 GGML_API void ggml_set_output(struct ggml_tensor * tensor);883 GGML_API void ggml_set_param(struct ggml_tensor * tensor);884 GGML_API void ggml_set_loss(struct ggml_tensor * tensor);885 886 //887 // operations on tensors with backpropagation888 //889 890 GGML_API struct ggml_tensor * ggml_dup(891 struct ggml_context * ctx,892 struct ggml_tensor * a);893 894 // in-place, returns view(a)895 GGML_API struct ggml_tensor * ggml_dup_inplace(896 struct ggml_context * ctx,897 struct ggml_tensor * a);898 899 GGML_API struct ggml_tensor * ggml_add(900 struct ggml_context * ctx,901 struct ggml_tensor * a,902 struct ggml_tensor * b);903 904 GGML_API struct ggml_tensor * ggml_add_inplace(905 struct ggml_context * ctx,906 struct ggml_tensor * a,907 struct ggml_tensor * b);908 909 GGML_API struct ggml_tensor * ggml_add_cast(910 struct ggml_context * ctx,911 struct ggml_tensor * a,912 struct ggml_tensor * b,913 enum ggml_type type);914 915 // dst[i0, i1, i2] = a[i0, i1, i2] + b[i0, ids[i1, i2]]916 GGML_API struct ggml_tensor * ggml_add_id(917 struct ggml_context * ctx,918 struct ggml_tensor * a,919 struct ggml_tensor * b,920 struct ggml_tensor * ids);921 922 GGML_DEPRECATED(GGML_API struct ggml_tensor * ggml_add1(923 struct ggml_context * ctx,924 struct ggml_tensor * a,925 struct ggml_tensor * b),926 "use ggml_add instead");927 928 GGML_DEPRECATED(GGML_API struct ggml_tensor * ggml_add1_inplace(929 struct ggml_context * ctx,930 struct ggml_tensor * a,931 struct ggml_tensor * b),932 "use ggml_add_inplace instead");933 934 // dst = a935 // view(dst, nb1, nb2, nb3, offset) += b936 // return dst937 GGML_API struct ggml_tensor * ggml_acc(938 struct ggml_context * ctx,939 struct ggml_tensor * a,940 struct ggml_tensor * b,941 size_t nb1,942 size_t nb2,943 size_t nb3,944 size_t offset);945 946 GGML_API struct ggml_tensor * ggml_acc_inplace(947 struct ggml_context * ctx,948 struct ggml_tensor * a,949 struct ggml_tensor * b,950 size_t nb1,951 size_t nb2,952 size_t nb3,953 size_t offset);954 955 GGML_API struct ggml_tensor * ggml_sub(956 struct ggml_context * ctx,957 struct ggml_tensor * a,958 struct ggml_tensor * b);959 960 GGML_API struct ggml_tensor * ggml_sub_inplace(961 struct ggml_context * ctx,962 struct ggml_tensor * a,963 struct ggml_tensor * b);964 965 GGML_API struct ggml_tensor * ggml_mul(966 struct ggml_context * ctx,967 struct ggml_tensor * a,968 struct ggml_tensor * b);969 970 GGML_API struct ggml_tensor * ggml_mul_inplace(971 struct ggml_context * ctx,972 struct ggml_tensor * a,973 struct ggml_tensor * b);974 975 GGML_API struct ggml_tensor * ggml_div(976 struct ggml_context * ctx,977 struct ggml_tensor * a,978 struct ggml_tensor * b);979 980 GGML_API struct ggml_tensor * ggml_div_inplace(981 struct ggml_context * ctx,982 struct ggml_tensor * a,983 struct ggml_tensor * b);984 985 GGML_API struct ggml_tensor * ggml_sqr(986 struct ggml_context * ctx,987 struct ggml_tensor * a);988 989 GGML_API struct ggml_tensor * ggml_sqr_inplace(990 struct ggml_context * ctx,991 struct ggml_tensor * a);992 993 GGML_API struct ggml_tensor * ggml_sqrt(994 struct ggml_context * ctx,995 struct ggml_tensor * a);996 997 GGML_API struct ggml_tensor * ggml_sqrt_inplace(998 struct ggml_context * ctx,999 struct ggml_tensor * a);1000 1001 GGML_API struct ggml_tensor * ggml_log(1002 struct ggml_context * ctx,1003 struct ggml_tensor * a);1004 1005 GGML_API struct ggml_tensor * ggml_log_inplace(1006 struct ggml_context * ctx,1007 struct ggml_tensor * a);1008 1009 GGML_API struct ggml_tensor * ggml_expm1(1010 struct ggml_context * ctx,1011 struct ggml_tensor * a);1012 1013 GGML_API struct ggml_tensor * ggml_expm1_inplace(1014 struct ggml_context * ctx,1015 struct ggml_tensor * a);1016 1017 GGML_API struct ggml_tensor * ggml_softplus(1018 struct ggml_context * ctx,1019 struct ggml_tensor * a);1020 1021 GGML_API struct ggml_tensor * ggml_softplus_inplace(1022 struct ggml_context * ctx,1023 struct ggml_tensor * a);1024 1025 GGML_API struct ggml_tensor * ggml_sin(1026 struct ggml_context * ctx,1027 struct ggml_tensor * a);1028 1029 GGML_API struct ggml_tensor * ggml_sin_inplace(1030 struct ggml_context * ctx,1031 struct ggml_tensor * a);1032 1033 GGML_API struct ggml_tensor * ggml_cos(1034 struct ggml_context * ctx,1035 struct ggml_tensor * a);1036 1037 GGML_API struct ggml_tensor * ggml_cos_inplace(1038 struct ggml_context * ctx,1039 struct ggml_tensor * a);1040 1041 // return scalar1042 GGML_API struct ggml_tensor * ggml_sum(1043 struct ggml_context * ctx,1044 struct ggml_tensor * a);1045 1046 // sums along rows, with input shape [a,b,c,d] return shape [1,b,c,d]1047 GGML_API struct ggml_tensor * ggml_sum_rows(1048 struct ggml_context * ctx,1049 struct ggml_tensor * a);1050 1051 GGML_API struct ggml_tensor * ggml_cumsum(1052 struct ggml_context * ctx,1053 struct ggml_tensor * a);1054 1055 // mean along rows1056 GGML_API struct ggml_tensor * ggml_mean(1057 struct ggml_context * ctx,1058 struct ggml_tensor * a);1059 1060 // argmax along rows1061 GGML_API struct ggml_tensor * ggml_argmax(1062 struct ggml_context * ctx,1063 struct ggml_tensor * a);1064 1065 // count number of equal elements in a and b1066 GGML_API struct ggml_tensor * ggml_count_equal(1067 struct ggml_context * ctx,1068 struct ggml_tensor * a,1069 struct ggml_tensor * b);1070 1071 // if a is the same shape as b, and a is not parameter, return a1072 // otherwise, return a new tensor: repeat(a) to fit in b1073 GGML_API struct ggml_tensor * ggml_repeat(1074 struct ggml_context * ctx,1075 struct ggml_tensor * a,1076 struct ggml_tensor * b);1077 1078 // repeat a to the specified shape1079 GGML_API struct ggml_tensor * ggml_repeat_4d(1080 struct ggml_context * ctx,1081 struct ggml_tensor * a,1082 int64_t ne0,1083 int64_t ne1,1084 int64_t ne2,1085 int64_t ne3);1086 1087 // sums repetitions in a into shape of b1088 GGML_API struct ggml_tensor * ggml_repeat_back(1089 struct ggml_context * ctx,1090 struct ggml_tensor * a,1091 struct ggml_tensor * b); // sum up values that are adjacent in dims > 0 instead of repeated with same stride1092 1093 // concat a and b along dim1094 // used in stable-diffusion1095 GGML_API struct ggml_tensor * ggml_concat(1096 struct ggml_context * ctx,1097 struct ggml_tensor * a,1098 struct ggml_tensor * b,1099 int dim);1100 1101 GGML_API struct ggml_tensor * ggml_abs(1102 struct ggml_context * ctx,1103 struct ggml_tensor * a);1104 1105 GGML_API struct ggml_tensor * ggml_abs_inplace(1106 struct ggml_context * ctx,1107 struct ggml_tensor * a);1108 1109 GGML_API struct ggml_tensor * ggml_sgn(1110 struct ggml_context * ctx,1111 struct ggml_tensor * a);1112 1113 GGML_API struct ggml_tensor * ggml_sgn_inplace(1114 struct ggml_context * ctx,1115 struct ggml_tensor * a);1116 1117 GGML_API struct ggml_tensor * ggml_neg(1118 struct ggml_context * ctx,1119 struct ggml_tensor * a);1120 1121 GGML_API struct ggml_tensor * ggml_neg_inplace(1122 struct ggml_context * ctx,1123 struct ggml_tensor * a);1124 1125 GGML_API struct ggml_tensor * ggml_step(1126 struct ggml_context * ctx,1127 struct ggml_tensor * a);1128 1129 GGML_API struct ggml_tensor * ggml_step_inplace(1130 struct ggml_context * ctx,1131 struct ggml_tensor * a);1132 1133 GGML_API struct ggml_tensor * ggml_tanh(1134 struct ggml_context * ctx,1135 struct ggml_tensor * a);1136 1137 GGML_API struct ggml_tensor * ggml_tanh_inplace(1138 struct ggml_context * ctx,1139 struct ggml_tensor * a);1140 1141 GGML_API struct ggml_tensor * ggml_elu(1142 struct ggml_context * ctx,1143 struct ggml_tensor * a);1144 1145 GGML_API struct ggml_tensor * ggml_elu_inplace(1146 struct ggml_context * ctx,1147 struct ggml_tensor * a);1148 1149 GGML_API struct ggml_tensor * ggml_relu(1150 struct ggml_context * ctx,1151 struct ggml_tensor * a);1152 1153 GGML_API struct ggml_tensor * ggml_leaky_relu(1154 struct ggml_context * ctx,1155 struct ggml_tensor * a, float negative_slope, bool inplace);1156 1157 GGML_API struct ggml_tensor * ggml_relu_inplace(1158 struct ggml_context * ctx,1159 struct ggml_tensor * a);1160 1161 GGML_API struct ggml_tensor * ggml_sigmoid(1162 struct ggml_context * ctx,1163 struct ggml_tensor * a);1164 1165 GGML_API struct ggml_tensor * ggml_sigmoid_inplace(1166 struct ggml_context * ctx,1167 struct ggml_tensor * a);1168 1169 GGML_API struct ggml_tensor * ggml_gelu(1170 struct ggml_context * ctx,1171 struct ggml_tensor * a);1172 1173 GGML_API struct ggml_tensor * ggml_gelu_inplace(1174 struct ggml_context * ctx,1175 struct ggml_tensor * a);1176 1177 // GELU using erf (error function) when possible1178 // some backends may fallback to approximation based on Abramowitz and Stegun formula1179 GGML_API struct ggml_tensor * ggml_gelu_erf(1180 struct ggml_context * ctx,1181 struct ggml_tensor * a);1182 1183 GGML_API struct ggml_tensor * ggml_gelu_erf_inplace(1184 struct ggml_context * ctx,1185 struct ggml_tensor * a);1186 1187 GGML_API struct ggml_tensor * ggml_gelu_quick(1188 struct ggml_context * ctx,1189 struct ggml_tensor * a);1190 1191 GGML_API struct ggml_tensor * ggml_gelu_quick_inplace(1192 struct ggml_context * ctx,1193 struct ggml_tensor * a);1194 1195 GGML_API struct ggml_tensor * ggml_silu(1196 struct ggml_context * ctx,1197 struct ggml_tensor * a);1198 1199 GGML_API struct ggml_tensor * ggml_silu_inplace(1200 struct ggml_context * ctx,