Team Ai
Apppublic

KBaba7/llama.cpp

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes
ggml.h2194 linesDownload Raw Back to include
1#pragma once2 3//4// GGML Tensor Library5//6// This documentation is still a work in progress.7// If you wish some specific topics to be covered, feel free to drop a comment:8//9//   https://github.com/ggerganov/whisper.cpp/issues/4010//11// ## Overview12//13// This library implements:14//15//  - a set of tensor operations16//  - automatic differentiation17//  - basic optimization algorithms18//19// The aim of this library is to provide a minimalistic approach for various machine learning tasks. This includes,20// but is not limited to, the following:21//22//  - linear regression23//  - support vector machines24//  - neural networks25//26// The library allows the user to define a certain function using the available tensor operations. This function27// definition is represented internally via a computation graph. Each tensor operation in the function definition28// corresponds to a node in the graph. Having the computation graph defined, the user can choose to compute the29// function's value and/or its gradient with respect to the input variables. Optionally, the function can be optimized30// using one of the available optimization algorithms.31//32// For example, here we define the function: f(x) = a*x^2 + b33//34//   {35//       struct ggml_init_params params = {36//           .mem_size   = 16*1024*1024,37//           .mem_buffer = NULL,38//       };39//40//       // memory allocation happens here41//       struct ggml_context * ctx = ggml_init(params);42//43//       struct ggml_tensor * x = ggml_new_tensor_1d(ctx, GGML_TYPE_F32, 1);44//45//       ggml_set_param(ctx, x); // x is an input variable46//47//       struct ggml_tensor * a  = ggml_new_tensor_1d(ctx, GGML_TYPE_F32, 1);48//       struct ggml_tensor * b  = ggml_new_tensor_1d(ctx, GGML_TYPE_F32, 1);49//       struct ggml_tensor * x2 = ggml_mul(ctx, x, x);50//       struct ggml_tensor * f  = ggml_add(ctx, ggml_mul(ctx, a, x2), b);51//52//       ...53//   }54//55// Notice that the function definition above does not involve any actual computation. The computation is performed only56// when the user explicitly requests it. For example, to compute the function's value at x = 2.0:57//58//   {59//       ...60//61//       struct ggml_cgraph * gf = ggml_new_graph(ctx);62//       ggml_build_forward_expand(gf, f);63//64//       // set the input variable and parameter values65//       ggml_set_f32(x, 2.0f);66//       ggml_set_f32(a, 3.0f);67//       ggml_set_f32(b, 4.0f);68//69//       ggml_graph_compute_with_ctx(ctx, &gf, n_threads);70//71//       printf("f = %f\n", ggml_get_f32_1d(f, 0));72//73//       ...74//   }75//76// The actual computation is performed in the ggml_graph_compute() function.77//78// The ggml_new_tensor_...() functions create new tensors. They are allocated in the memory buffer provided to the79// ggml_init() function. You have to be careful not to exceed the memory buffer size. Therefore, you have to know80// in advance how much memory you need for your computation. Alternatively, you can allocate a large enough memory81// and after defining the computation graph, call the ggml_used_mem() function to find out how much memory was82// actually needed.83//84// The ggml_set_param() function marks a tensor as an input variable. This is used by the automatic85// differentiation and optimization algorithms.86//87// The described approach allows to define the function graph once and then compute its forward or backward graphs88// multiple times. All computations will use the same memory buffer allocated in the ggml_init() function. This way89// the user can avoid the memory allocation overhead at runtime.90//91// The library supports multi-dimensional tensors - up to 4 dimensions. The FP16 and FP32 data types are first class92// citizens, but in theory the library can be extended to support FP8 and integer data types.93//94// Each tensor operation produces a new tensor. Initially the library was envisioned to support only the use of unary95// and binary operations. Most of the available operations fall into one of these two categories. With time, it became96// clear that the library needs to support more complex operations. The way to support these operations is not clear97// yet, but a few examples are demonstrated in the following operations:98//99//   - ggml_permute()100//   - ggml_conv_1d_1s()101//   - ggml_conv_1d_2s()102//103// For each tensor operator, the library implements a forward and backward computation function. The forward function104// computes the output tensor value given the input tensor values. The backward function computes the adjoint of the105// input tensors given the adjoint of the output tensor. For a detailed explanation of what this means, take a106// calculus class, or watch the following video:107//108//   What is Automatic Differentiation?109//   https://www.youtube.com/watch?v=wG_nF1awSSY110//111//112// ## Tensor data (struct ggml_tensor)113//114// The tensors are stored in memory via the ggml_tensor struct. The structure provides information about the size of115// the tensor, the data type, and the memory buffer where the tensor data is stored. Additionally, it contains116// pointers to the "source" tensors - i.e. the tensors that were used to compute the current tensor. For example:117//118//   {119//       struct ggml_tensor * c = ggml_add(ctx, a, b);120//121//       assert(c->src[0] == a);122//       assert(c->src[1] == b);123//   }124//125// The multi-dimensional tensors are stored in row-major order. The ggml_tensor struct contains fields for the126// number of elements in each dimension ("ne") as well as the number of bytes ("nb", a.k.a. stride). This allows127// to store tensors that are not contiguous in memory, which is useful for operations such as transposition and128// permutation. All tensor operations have to take the stride into account and not assume that the tensor is129// contiguous in memory.130//131// The data of the tensor is accessed via the "data" pointer. For example:132//133//   {134//       const int nx = 2;135//       const int ny = 3;136//137//       struct ggml_tensor * a = ggml_new_tensor_2d(ctx, GGML_TYPE_F32, nx, ny);138//139//       for (int y = 0; y < ny; y++) {140//           for (int x = 0; x < nx; x++) {141//               *(float *) ((char *) a->data + y*a->nb[1] + x*a->nb[0]) = x + y;142//           }143//       }144//145//       ...146//   }147//148// Alternatively, there are helper functions, such as ggml_get_f32_1d() and ggml_set_f32_1d() that can be used.149//150// ## The matrix multiplication operator (ggml_mul_mat)151//152// TODO153//154//155// ## Multi-threading156//157// TODO158//159//160// ## Overview of ggml.c161//162// TODO163//164//165// ## SIMD optimizations166//167// TODO168//169//170// ## Debugging ggml171//172// TODO173//174//175 176#ifdef GGML_SHARED177#    if defined(_WIN32) && !defined(__MINGW32__)178#        ifdef GGML_BUILD179#            define GGML_API __declspec(dllexport) extern180#        else181#            define GGML_API __declspec(dllimport) extern182#        endif183#    else184#        define GGML_API __attribute__ ((visibility ("default"))) extern185#    endif186#else187#    define GGML_API extern188#endif189 190// TODO: support for clang191#ifdef __GNUC__192#    define GGML_DEPRECATED(func, hint) func __attribute__((deprecated(hint)))193#elif defined(_MSC_VER)194#    define GGML_DEPRECATED(func, hint) __declspec(deprecated(hint)) func195#else196#    define GGML_DEPRECATED(func, hint) func197#endif198 199#ifndef __GNUC__200#    define GGML_ATTRIBUTE_FORMAT(...)201#elif defined(__MINGW32__)202#    define GGML_ATTRIBUTE_FORMAT(...) __attribute__((format(gnu_printf, __VA_ARGS__)))203#else204#    define GGML_ATTRIBUTE_FORMAT(...) __attribute__((format(printf, __VA_ARGS__)))205#endif206 207#include <stdbool.h>208#include <stddef.h>209#include <stdint.h>210#include <stdio.h>211 212#define GGML_FILE_MAGIC   0x67676d6c // "ggml"213#define GGML_FILE_VERSION 2214 215#define GGML_QNT_VERSION        2    // bump this on quantization format changes216#define GGML_QNT_VERSION_FACTOR 1000 // do not change this217 218#define GGML_MAX_DIMS           4219#define GGML_MAX_PARAMS         2048220#define GGML_MAX_SRC            10221#define GGML_MAX_N_THREADS      512222#define GGML_MAX_OP_PARAMS      64223 224#ifndef GGML_MAX_NAME225#   define GGML_MAX_NAME        64226#endif227 228#define GGML_DEFAULT_N_THREADS  4229#define GGML_DEFAULT_GRAPH_SIZE 2048230 231#if UINTPTR_MAX == 0xFFFFFFFF232    #define GGML_MEM_ALIGN 4233#else234    #define GGML_MEM_ALIGN 16235#endif236 237#define GGML_EXIT_SUCCESS 0238#define GGML_EXIT_ABORTED 1239 240#define GGML_ROPE_TYPE_NEOX   2241#define GGML_ROPE_TYPE_MROPE  8242#define GGML_ROPE_TYPE_VISION 24243 244#define GGML_UNUSED(x) (void)(x)245 246#define GGML_PAD(x, n) (((x) + (n) - 1) & ~((n) - 1))247 248#ifndef NDEBUG249#   define GGML_UNREACHABLE() do { fprintf(stderr, "statement should be unreachable\n"); abort(); } while(0)250#elif defined(__GNUC__)251#   define GGML_UNREACHABLE() __builtin_unreachable()252#elif defined(_MSC_VER)253#   define GGML_UNREACHABLE() __assume(0)254#else255#   define GGML_UNREACHABLE() ((void) 0)256#endif257 258#ifdef __cplusplus259#   define GGML_NORETURN [[noreturn]]260#elif defined(_MSC_VER)261#   define GGML_NORETURN __declspec(noreturn)262#else263#   define GGML_NORETURN _Noreturn264#endif265 266#define GGML_ABORT(...) ggml_abort(__FILE__, __LINE__, __VA_ARGS__)267#define GGML_ASSERT(x) if (!(x)) GGML_ABORT("GGML_ASSERT(%s) failed", #x)268 269// used to copy the number of elements and stride in bytes of tensors into local variables.270// main purpose is to reduce code duplication and improve readability.271//272// example:273//274//    GGML_TENSOR_LOCALS(int64_t, ne1, src1, ne);275//    GGML_TENSOR_LOCALS(size_t,  nb1, src1, nb);276//277#define GGML_TENSOR_LOCALS_1(type, prefix, pointer, array) \278    const type prefix##0 = (pointer)->array[0]; \279    GGML_UNUSED(prefix##0);280#define GGML_TENSOR_LOCALS_2(type, prefix, pointer, array) \281    GGML_TENSOR_LOCALS_1    (type, prefix, pointer, array) \282    const type prefix##1 = (pointer)->array[1]; \283    GGML_UNUSED(prefix##1);284#define GGML_TENSOR_LOCALS_3(type, prefix, pointer, array) \285    GGML_TENSOR_LOCALS_2    (type, prefix, pointer, array) \286    const type prefix##2 = (pointer)->array[2]; \287    GGML_UNUSED(prefix##2);288#define GGML_TENSOR_LOCALS(type, prefix, pointer, array) \289    GGML_TENSOR_LOCALS_3  (type, prefix, pointer, array) \290    const type prefix##3 = (pointer)->array[3]; \291    GGML_UNUSED(prefix##3);292 293#define GGML_TENSOR_UNARY_OP_LOCALS \294    GGML_TENSOR_LOCALS(int64_t, ne0, src0, ne) \295    GGML_TENSOR_LOCALS(size_t,  nb0, src0, nb) \296    GGML_TENSOR_LOCALS(int64_t, ne,  dst,  ne) \297    GGML_TENSOR_LOCALS(size_t,  nb,  dst,  nb)298 299#define GGML_TENSOR_BINARY_OP_LOCALS \300    GGML_TENSOR_LOCALS(int64_t, ne0, src0, ne) \301    GGML_TENSOR_LOCALS(size_t,  nb0, src0, nb) \302    GGML_TENSOR_LOCALS(int64_t, ne1, src1, ne) \303    GGML_TENSOR_LOCALS(size_t,  nb1, src1, nb) \304    GGML_TENSOR_LOCALS(int64_t, ne,  dst,  ne) \305    GGML_TENSOR_LOCALS(size_t,  nb,  dst,  nb)306 307#define GGML_TENSOR_BINARY_OP_LOCALS01 \308    GGML_TENSOR_LOCALS(int64_t, ne0, src0, ne) \309    GGML_TENSOR_LOCALS(size_t,  nb0, src0, nb) \310    GGML_TENSOR_LOCALS(int64_t, ne1, src1, ne) \311    GGML_TENSOR_LOCALS(size_t,  nb1, src1, nb)312 313#ifdef  __cplusplus314extern "C" {315#endif316 317    GGML_NORETURN GGML_ATTRIBUTE_FORMAT(3, 4)318    GGML_API void ggml_abort(const char * file, int line, const char * fmt, ...);319 320    enum ggml_status {321        GGML_STATUS_ALLOC_FAILED = -2,322        GGML_STATUS_FAILED = -1,323        GGML_STATUS_SUCCESS = 0,324        GGML_STATUS_ABORTED = 1,325    };326 327    // get ggml_status name string328    GGML_API const char * ggml_status_to_string(enum ggml_status status);329 330    // ieee 754-2008 half-precision float16331    // todo: make this not an integral type332    typedef uint16_t ggml_fp16_t;333    GGML_API float       ggml_fp16_to_fp32(ggml_fp16_t);334    GGML_API ggml_fp16_t ggml_fp32_to_fp16(float);335    GGML_API void        ggml_fp16_to_fp32_row(const ggml_fp16_t *, float *, int64_t);336    GGML_API void        ggml_fp32_to_fp16_row(const float *, ggml_fp16_t *, int64_t);337 338    // google brain half-precision bfloat16339    typedef struct { uint16_t bits; } ggml_bf16_t;340    GGML_API ggml_bf16_t ggml_fp32_to_bf16(float);341    GGML_API float       ggml_bf16_to_fp32(ggml_bf16_t);  // consider just doing << 16342    GGML_API void        ggml_bf16_to_fp32_row(const ggml_bf16_t *, float *, int64_t);343    GGML_API void        ggml_fp32_to_bf16_row_ref(const float *, ggml_bf16_t *, int64_t);344    GGML_API void        ggml_fp32_to_bf16_row(const float *, ggml_bf16_t *, int64_t);345 346    struct ggml_object;347    struct ggml_context;348    struct ggml_cgraph;349 350    // NOTE: always add types at the end of the enum to keep backward compatibility351    enum ggml_type {352        GGML_TYPE_F32     = 0,353        GGML_TYPE_F16     = 1,354        GGML_TYPE_Q4_0    = 2,355        GGML_TYPE_Q4_1    = 3,356        // GGML_TYPE_Q4_2 = 4, support has been removed357        // GGML_TYPE_Q4_3 = 5, support has been removed358        GGML_TYPE_Q5_0    = 6,359        GGML_TYPE_Q5_1    = 7,360        GGML_TYPE_Q8_0    = 8,361        GGML_TYPE_Q8_1    = 9,362        GGML_TYPE_Q2_K    = 10,363        GGML_TYPE_Q3_K    = 11,364        GGML_TYPE_Q4_K    = 12,365        GGML_TYPE_Q5_K    = 13,366        GGML_TYPE_Q6_K    = 14,367        GGML_TYPE_Q8_K    = 15,368        GGML_TYPE_IQ2_XXS = 16,369        GGML_TYPE_IQ2_XS  = 17,370        GGML_TYPE_IQ3_XXS = 18,371        GGML_TYPE_IQ1_S   = 19,372        GGML_TYPE_IQ4_NL  = 20,373        GGML_TYPE_IQ3_S   = 21,374        GGML_TYPE_IQ2_S   = 22,375        GGML_TYPE_IQ4_XS  = 23,376        GGML_TYPE_I8      = 24,377        GGML_TYPE_I16     = 25,378        GGML_TYPE_I32     = 26,379        GGML_TYPE_I64     = 27,380        GGML_TYPE_F64     = 28,381        GGML_TYPE_IQ1_M   = 29,382        GGML_TYPE_BF16    = 30,383        // GGML_TYPE_Q4_0_4_4 = 31, support has been removed from gguf files384        // GGML_TYPE_Q4_0_4_8 = 32,385        // GGML_TYPE_Q4_0_8_8 = 33,386        GGML_TYPE_TQ1_0   = 34,387        GGML_TYPE_TQ2_0   = 35,388        // GGML_TYPE_IQ4_NL_4_4 = 36,389        // GGML_TYPE_IQ4_NL_4_8 = 37,390        // GGML_TYPE_IQ4_NL_8_8 = 38,391        GGML_TYPE_COUNT   = 39,392    };393 394    // precision395    enum ggml_prec {396        GGML_PREC_DEFAULT,397        GGML_PREC_F32,398    };399 400    // model file types401    enum ggml_ftype {402        GGML_FTYPE_UNKNOWN        = -1,403        GGML_FTYPE_ALL_F32        = 0,404        GGML_FTYPE_MOSTLY_F16     = 1,  // except 1d tensors405        GGML_FTYPE_MOSTLY_Q4_0    = 2,  // except 1d tensors406        GGML_FTYPE_MOSTLY_Q4_1    = 3,  // except 1d tensors407        GGML_FTYPE_MOSTLY_Q4_1_SOME_F16 = 4, // tok_embeddings.weight and output.weight are F16408        GGML_FTYPE_MOSTLY_Q8_0    = 7,  // except 1d tensors409        GGML_FTYPE_MOSTLY_Q5_0    = 8,  // except 1d tensors410        GGML_FTYPE_MOSTLY_Q5_1    = 9,  // except 1d tensors411        GGML_FTYPE_MOSTLY_Q2_K    = 10, // except 1d tensors412        GGML_FTYPE_MOSTLY_Q3_K    = 11, // except 1d tensors413        GGML_FTYPE_MOSTLY_Q4_K    = 12, // except 1d tensors414        GGML_FTYPE_MOSTLY_Q5_K    = 13, // except 1d tensors415        GGML_FTYPE_MOSTLY_Q6_K    = 14, // except 1d tensors416        GGML_FTYPE_MOSTLY_IQ2_XXS = 15, // except 1d tensors417        GGML_FTYPE_MOSTLY_IQ2_XS  = 16, // except 1d tensors418        GGML_FTYPE_MOSTLY_IQ3_XXS = 17, // except 1d tensors419        GGML_FTYPE_MOSTLY_IQ1_S   = 18, // except 1d tensors420        GGML_FTYPE_MOSTLY_IQ4_NL  = 19, // except 1d tensors421        GGML_FTYPE_MOSTLY_IQ3_S   = 20, // except 1d tensors422        GGML_FTYPE_MOSTLY_IQ2_S   = 21, // except 1d tensors423        GGML_FTYPE_MOSTLY_IQ4_XS  = 22, // except 1d tensors424        GGML_FTYPE_MOSTLY_IQ1_M   = 23, // except 1d tensors425        GGML_FTYPE_MOSTLY_BF16    = 24, // except 1d tensors426    };427 428    // available tensor operations:429    enum ggml_op {430        GGML_OP_NONE = 0,431 432        GGML_OP_DUP,433        GGML_OP_ADD,434        GGML_OP_ADD1,435        GGML_OP_ACC,436        GGML_OP_SUB,437        GGML_OP_MUL,438        GGML_OP_DIV,439        GGML_OP_SQR,440        GGML_OP_SQRT,441        GGML_OP_LOG,442        GGML_OP_SIN,443        GGML_OP_COS,444        GGML_OP_SUM,445        GGML_OP_SUM_ROWS,446        GGML_OP_MEAN,447        GGML_OP_ARGMAX,448        GGML_OP_COUNT_EQUAL,449        GGML_OP_REPEAT,450        GGML_OP_REPEAT_BACK,451        GGML_OP_CONCAT,452        GGML_OP_SILU_BACK,453        GGML_OP_NORM, // normalize454        GGML_OP_RMS_NORM,455        GGML_OP_RMS_NORM_BACK,456        GGML_OP_GROUP_NORM,457 458        GGML_OP_MUL_MAT,459        GGML_OP_MUL_MAT_ID,460        GGML_OP_OUT_PROD,461 462        GGML_OP_SCALE,463        GGML_OP_SET,464        GGML_OP_CPY,465        GGML_OP_CONT,466        GGML_OP_RESHAPE,467        GGML_OP_VIEW,468        GGML_OP_PERMUTE,469        GGML_OP_TRANSPOSE,470        GGML_OP_GET_ROWS,471        GGML_OP_GET_ROWS_BACK,472        GGML_OP_DIAG,473        GGML_OP_DIAG_MASK_INF,474        GGML_OP_DIAG_MASK_ZERO,475        GGML_OP_SOFT_MAX,476        GGML_OP_SOFT_MAX_BACK,477        GGML_OP_ROPE,478        GGML_OP_ROPE_BACK,479        GGML_OP_CLAMP,480        GGML_OP_CONV_TRANSPOSE_1D,481        GGML_OP_IM2COL,482        GGML_OP_IM2COL_BACK,483        GGML_OP_CONV_TRANSPOSE_2D,484        GGML_OP_POOL_1D,485        GGML_OP_POOL_2D,486        GGML_OP_POOL_2D_BACK,487        GGML_OP_UPSCALE, // nearest interpolate488        GGML_OP_PAD,489        GGML_OP_PAD_REFLECT_1D,490        GGML_OP_ARANGE,491        GGML_OP_TIMESTEP_EMBEDDING,492        GGML_OP_ARGSORT,493        GGML_OP_LEAKY_RELU,494 495        GGML_OP_FLASH_ATTN_EXT,496        GGML_OP_FLASH_ATTN_BACK,497        GGML_OP_SSM_CONV,498        GGML_OP_SSM_SCAN,499        GGML_OP_WIN_PART,500        GGML_OP_WIN_UNPART,501        GGML_OP_GET_REL_POS,502        GGML_OP_ADD_REL_POS,503        GGML_OP_RWKV_WKV6,504        GGML_OP_GATED_LINEAR_ATTN,505 506        GGML_OP_UNARY,507 508        GGML_OP_MAP_UNARY,509        GGML_OP_MAP_BINARY,510 511        GGML_OP_MAP_CUSTOM1_F32,512        GGML_OP_MAP_CUSTOM2_F32,513        GGML_OP_MAP_CUSTOM3_F32,514 515        GGML_OP_MAP_CUSTOM1,516        GGML_OP_MAP_CUSTOM2,517        GGML_OP_MAP_CUSTOM3,518 519        GGML_OP_CROSS_ENTROPY_LOSS,520        GGML_OP_CROSS_ENTROPY_LOSS_BACK,521        GGML_OP_OPT_STEP_ADAMW,522 523        GGML_OP_COUNT,524    };525 526    enum ggml_unary_op {527        GGML_UNARY_OP_ABS,528        GGML_UNARY_OP_SGN,529        GGML_UNARY_OP_NEG,530        GGML_UNARY_OP_STEP,531        GGML_UNARY_OP_TANH,532        GGML_UNARY_OP_ELU,533        GGML_UNARY_OP_RELU,534        GGML_UNARY_OP_SIGMOID,535        GGML_UNARY_OP_GELU,536        GGML_UNARY_OP_GELU_QUICK,537        GGML_UNARY_OP_SILU,538        GGML_UNARY_OP_HARDSWISH,539        GGML_UNARY_OP_HARDSIGMOID,540        GGML_UNARY_OP_EXP,541 542        GGML_UNARY_OP_COUNT,543    };544 545    enum ggml_object_type {546        GGML_OBJECT_TYPE_TENSOR,547        GGML_OBJECT_TYPE_GRAPH,548        GGML_OBJECT_TYPE_WORK_BUFFER549    };550 551    enum ggml_log_level {552        GGML_LOG_LEVEL_NONE  = 0,553        GGML_LOG_LEVEL_DEBUG = 1,554        GGML_LOG_LEVEL_INFO  = 2,555        GGML_LOG_LEVEL_WARN  = 3,556        GGML_LOG_LEVEL_ERROR = 4,557        GGML_LOG_LEVEL_CONT  = 5, // continue previous log558    };559 560    // this tensor...561    enum ggml_tensor_flag {562        GGML_TENSOR_FLAG_INPUT  =  1, // ...is an input for the GGML compute graph563        GGML_TENSOR_FLAG_OUTPUT =  2, // ...is an output for the GGML compute graph564        GGML_TENSOR_FLAG_PARAM  =  4, // ...contains trainable parameters565        GGML_TENSOR_FLAG_LOSS   =  8, // ...defines loss for numerical optimization (multiple loss tensors add up)566    };567 568    struct ggml_init_params {569        // memory pool570        size_t mem_size;   // bytes571        void * mem_buffer; // if NULL, memory will be allocated internally572        bool   no_alloc;   // don't allocate memory for the tensor data573    };574 575    // n-dimensional tensor576    struct ggml_tensor {577        enum ggml_type type;578 579        struct ggml_backend_buffer * buffer;580 581        int64_t ne[GGML_MAX_DIMS]; // number of elements582        size_t  nb[GGML_MAX_DIMS]; // stride in bytes:583                                   // nb[0] = ggml_type_size(type)584                                   // nb[1] = nb[0]   * (ne[0] / ggml_blck_size(type)) + padding585                                   // nb[i] = nb[i-1] * ne[i-1]586 587        // compute data588        enum ggml_op op;589 590        // op params - allocated as int32_t for alignment591        int32_t op_params[GGML_MAX_OP_PARAMS / sizeof(int32_t)];592 593        int32_t flags;594 595        struct ggml_tensor * src[GGML_MAX_SRC];596 597        // source tensor and offset for views598        struct ggml_tensor * view_src;599        size_t               view_offs;600 601        void * data;602 603        char name[GGML_MAX_NAME];604 605        void * extra; // extra things e.g. for ggml-cuda.cu606 607        char padding[8];608    };609 610    static const size_t GGML_TENSOR_SIZE = sizeof(struct ggml_tensor);611 612    // Abort callback613    // If not NULL, called before ggml computation614    // If it returns true, the computation is aborted615    typedef bool (*ggml_abort_callback)(void * data);616 617 618    //619    // GUID620    //621 622    // GUID types623    typedef uint8_t ggml_guid[16];624    typedef ggml_guid * ggml_guid_t;625 626    GGML_API bool ggml_guid_matches(ggml_guid_t guid_a, ggml_guid_t guid_b);627 628    // misc629 630    GGML_API void    ggml_time_init(void); // call this once at the beginning of the program631    GGML_API int64_t ggml_time_ms(void);632    GGML_API int64_t ggml_time_us(void);633    GGML_API int64_t ggml_cycles(void);634    GGML_API int64_t ggml_cycles_per_ms(void);635 636    // accepts a UTF-8 path, even on Windows637    GGML_API FILE *  ggml_fopen(const char * fname, const char * mode);638 639    GGML_API void    ggml_print_object (const struct ggml_object * obj);640    GGML_API void    ggml_print_objects(const struct ggml_context * ctx);641 642    GGML_API int64_t ggml_nelements (const struct ggml_tensor * tensor);643    GGML_API int64_t ggml_nrows     (const struct ggml_tensor * tensor);644    GGML_API size_t  ggml_nbytes    (const struct ggml_tensor * tensor);645    GGML_API size_t  ggml_nbytes_pad(const struct ggml_tensor * tensor); // same as ggml_nbytes() but padded to GGML_MEM_ALIGN646 647    GGML_API int64_t ggml_blck_size(enum ggml_type type);648    GGML_API size_t  ggml_type_size(enum ggml_type type);             // size in bytes for all elements in a block649    GGML_API size_t  ggml_row_size (enum ggml_type type, int64_t ne); // size in bytes for all elements in a row650 651    GGML_DEPRECATED(652    GGML_API double ggml_type_sizef(enum ggml_type type), // ggml_type_size()/ggml_blck_size() as float653    "use ggml_row_size() instead");654 655    GGML_API const char * ggml_type_name(enum ggml_type type);656    GGML_API const char * ggml_op_name  (enum ggml_op   op);657    GGML_API const char * ggml_op_symbol(enum ggml_op   op);658 659    GGML_API const char * ggml_unary_op_name(enum ggml_unary_op op);660    GGML_API const char * ggml_op_desc(const struct ggml_tensor * t); // unary or op name661 662    GGML_API size_t  ggml_element_size(const struct ggml_tensor * tensor);663 664    GGML_API bool    ggml_is_quantized(enum ggml_type type);665 666    // TODO: temporary until model loading of ggml examples is refactored667    GGML_API enum ggml_type ggml_ftype_to_ggml_type(enum ggml_ftype ftype);668 669    GGML_API bool ggml_is_transposed(const struct ggml_tensor * tensor);670    GGML_API bool ggml_is_permuted  (const struct ggml_tensor * tensor);671    GGML_API bool ggml_is_empty     (const struct ggml_tensor * tensor);672    GGML_API bool ggml_is_scalar    (const struct ggml_tensor * tensor);673    GGML_API bool ggml_is_vector    (const struct ggml_tensor * tensor);674    GGML_API bool ggml_is_matrix    (const struct ggml_tensor * tensor);675    GGML_API bool ggml_is_3d        (const struct ggml_tensor * tensor);676    GGML_API int  ggml_n_dims       (const struct ggml_tensor * tensor); // returns 1 for scalars677 678    GGML_API bool ggml_is_contiguous  (const struct ggml_tensor * tensor);679    GGML_API bool ggml_is_contiguous_0(const struct ggml_tensor * tensor); // same as ggml_is_contiguous()680    GGML_API bool ggml_is_contiguous_1(const struct ggml_tensor * tensor); // contiguous for dims >= 1681    GGML_API bool ggml_is_contiguous_2(const struct ggml_tensor * tensor); // contiguous for dims >= 2682 683    GGML_API bool ggml_are_same_shape (const struct ggml_tensor * t0, const struct ggml_tensor * t1);684    GGML_API bool ggml_are_same_stride(const struct ggml_tensor * t0, const struct ggml_tensor * t1);685 686    GGML_API bool ggml_can_repeat(const struct ggml_tensor * t0, const struct ggml_tensor * t1);687 688    // use this to compute the memory overhead of a tensor689    GGML_API size_t ggml_tensor_overhead(void);690 691    GGML_API bool ggml_validate_row_data(enum ggml_type type, const void * data, size_t nbytes);692 693    // main694 695    GGML_API struct ggml_context * ggml_init (struct ggml_init_params params);696    GGML_API void                  ggml_reset(struct ggml_context * ctx);697    GGML_API void                  ggml_free (struct ggml_context * ctx);698 699    GGML_API size_t  ggml_used_mem(const struct ggml_context * ctx);700 701    GGML_API bool    ggml_get_no_alloc(struct ggml_context * ctx);702    GGML_API void    ggml_set_no_alloc(struct ggml_context * ctx, bool no_alloc);703 704    GGML_API void *  ggml_get_mem_buffer     (const struct ggml_context * ctx);705    GGML_API size_t  ggml_get_mem_size       (const struct ggml_context * ctx);706    GGML_API size_t  ggml_get_max_tensor_size(const struct ggml_context * ctx);707 708    GGML_API struct ggml_tensor * ggml_new_tensor(709            struct ggml_context * ctx,710            enum   ggml_type type,711            int    n_dims,712            const int64_t *ne);713 714    GGML_API struct ggml_tensor * ggml_new_tensor_1d(715            struct ggml_context * ctx,716            enum   ggml_type type,717            int64_t ne0);718 719    GGML_API struct ggml_tensor * ggml_new_tensor_2d(720            struct ggml_context * ctx,721            enum   ggml_type type,722            int64_t ne0,723            int64_t ne1);724 725    GGML_API struct ggml_tensor * ggml_new_tensor_3d(726            struct ggml_context * ctx,727            enum   ggml_type type,728            int64_t ne0,729            int64_t ne1,730            int64_t ne2);731 732    GGML_API struct ggml_tensor * ggml_new_tensor_4d(733            struct ggml_context * ctx,734            enum   ggml_type type,735            int64_t ne0,736            int64_t ne1,737            int64_t ne2,738            int64_t ne3);739 740    GGML_API void * ggml_new_buffer(struct ggml_context * ctx, size_t nbytes);741 742    GGML_API struct ggml_tensor * ggml_dup_tensor (struct ggml_context * ctx, const struct ggml_tensor * src);743    GGML_API struct ggml_tensor * ggml_view_tensor(struct ggml_context * ctx, struct ggml_tensor * src);744 745    // Context tensor enumeration and lookup746    GGML_API struct ggml_tensor * ggml_get_first_tensor(const struct ggml_context * ctx);747    GGML_API struct ggml_tensor * ggml_get_next_tensor (const struct ggml_context * ctx, struct ggml_tensor * tensor);748    GGML_API struct ggml_tensor * ggml_get_tensor(struct ggml_context * ctx, const char * name);749 750    // Converts a flat index into coordinates751    GGML_API void ggml_unravel_index(const struct ggml_tensor * tensor, int64_t i, int64_t * i0, int64_t * i1, int64_t * i2, int64_t * i3);752 753    GGML_API enum ggml_unary_op ggml_get_unary_op(const struct ggml_tensor * tensor);754 755    GGML_API void *  ggml_get_data    (const struct ggml_tensor * tensor);756    GGML_API float * ggml_get_data_f32(const struct ggml_tensor * tensor);757 758    GGML_API const char *         ggml_get_name   (const struct ggml_tensor * tensor);759    GGML_API struct ggml_tensor * ggml_set_name   (      struct ggml_tensor * tensor, const char * name);760    GGML_ATTRIBUTE_FORMAT(2, 3)761    GGML_API struct ggml_tensor * ggml_format_name(      struct ggml_tensor * tensor, const char * fmt, ...);762 763    // Tensor flags764    GGML_API void ggml_set_input(struct ggml_tensor * tensor);765    GGML_API void ggml_set_output(struct ggml_tensor * tensor);766    GGML_API void ggml_set_param(struct ggml_context * ctx, struct ggml_tensor * tensor);767    GGML_API void ggml_set_loss(struct ggml_tensor * tensor);768 769    //770    // operations on tensors with backpropagation771    //772 773    GGML_API struct ggml_tensor * ggml_dup(774            struct ggml_context * ctx,775            struct ggml_tensor  * a);776 777    // in-place, returns view(a)778    GGML_API struct ggml_tensor * ggml_dup_inplace(779            struct ggml_context * ctx,780            struct ggml_tensor  * a);781 782    GGML_API struct ggml_tensor * ggml_add(783            struct ggml_context * ctx,784            struct ggml_tensor  * a,785            struct ggml_tensor  * b);786 787    GGML_API struct ggml_tensor * ggml_add_inplace(788            struct ggml_context * ctx,789            struct ggml_tensor  * a,790            struct ggml_tensor  * b);791 792    GGML_API struct ggml_tensor * ggml_add_cast(793            struct ggml_context * ctx,794            struct ggml_tensor  * a,795            struct ggml_tensor  * b,796            enum   ggml_type      type);797 798    GGML_API struct ggml_tensor * ggml_add1(799            struct ggml_context * ctx,800            struct ggml_tensor  * a,801            struct ggml_tensor  * b);802 803    GGML_API struct ggml_tensor * ggml_add1_inplace(804            struct ggml_context * ctx,805            struct ggml_tensor  * a,806            struct ggml_tensor  * b);807 808    // dst = a809    // view(dst, nb1, nb2, nb3, offset) += b810    // return dst811    GGML_API struct ggml_tensor * ggml_acc(812            struct ggml_context * ctx,813            struct ggml_tensor  * a,814            struct ggml_tensor  * b,815            size_t                nb1,816            size_t                nb2,817            size_t                nb3,818            size_t                offset);819 820    GGML_API struct ggml_tensor * ggml_acc_inplace(821            struct ggml_context * ctx,822            struct ggml_tensor  * a,823            struct ggml_tensor  * b,824            size_t                nb1,825            size_t                nb2,826            size_t                nb3,827            size_t                offset);828 829    GGML_API struct ggml_tensor * ggml_sub(830            struct ggml_context * ctx,831            struct ggml_tensor  * a,832            struct ggml_tensor  * b);833 834    GGML_API struct ggml_tensor * ggml_sub_inplace(835            struct ggml_context * ctx,836            struct ggml_tensor  * a,837            struct ggml_tensor  * b);838 839    GGML_API struct ggml_tensor * ggml_mul(840            struct ggml_context * ctx,841            struct ggml_tensor  * a,842            struct ggml_tensor  * b);843 844    GGML_API struct ggml_tensor * ggml_mul_inplace(845            struct ggml_context * ctx,846            struct ggml_tensor  * a,847            struct ggml_tensor  * b);848 849    GGML_API struct ggml_tensor * ggml_div(850            struct ggml_context * ctx,851            struct ggml_tensor  * a,852            struct ggml_tensor  * b);853 854    GGML_API struct ggml_tensor * ggml_div_inplace(855            struct ggml_context * ctx,856            struct ggml_tensor  * a,857            struct ggml_tensor  * b);858 859    GGML_API struct ggml_tensor * ggml_sqr(860            struct ggml_context * ctx,861            struct ggml_tensor  * a);862 863    GGML_API struct ggml_tensor * ggml_sqr_inplace(864            struct ggml_context * ctx,865            struct ggml_tensor  * a);866 867    GGML_API struct ggml_tensor * ggml_sqrt(868            struct ggml_context * ctx,869            struct ggml_tensor  * a);870 871    GGML_API struct ggml_tensor * ggml_sqrt_inplace(872            struct ggml_context * ctx,873            struct ggml_tensor  * a);874 875    GGML_API struct ggml_tensor * ggml_log(876            struct ggml_context * ctx,877            struct ggml_tensor  * a);878 879    GGML_API struct ggml_tensor * ggml_log_inplace(880            struct ggml_context * ctx,881            struct ggml_tensor  * a);882 883    GGML_API struct ggml_tensor * ggml_sin(884            struct ggml_context * ctx,885            struct ggml_tensor  * a);886 887    GGML_API struct ggml_tensor * ggml_sin_inplace(888            struct ggml_context * ctx,889            struct ggml_tensor  * a);890 891    GGML_API struct ggml_tensor * ggml_cos(892            struct ggml_context * ctx,893            struct ggml_tensor  * a);894 895    GGML_API struct ggml_tensor * ggml_cos_inplace(896            struct ggml_context * ctx,897            struct ggml_tensor  * a);898 899    // return scalar900    GGML_API struct ggml_tensor * ggml_sum(901            struct ggml_context * ctx,902            struct ggml_tensor  * a);903 904    // sums along rows, with input shape [a,b,c,d] return shape [1,b,c,d]905    GGML_API struct ggml_tensor * ggml_sum_rows(906            struct ggml_context * ctx,907            struct ggml_tensor  * a);908 909    // mean along rows910    GGML_API struct ggml_tensor * ggml_mean(911            struct ggml_context * ctx,912            struct ggml_tensor  * a);913 914    // argmax along rows915    GGML_API struct ggml_tensor * ggml_argmax(916            struct ggml_context * ctx,917            struct ggml_tensor  * a);918 919    // count number of equal elements in a and b920    GGML_API struct ggml_tensor * ggml_count_equal(921            struct ggml_context * ctx,922            struct ggml_tensor  * a,923            struct ggml_tensor  * b);924 925    // if a is the same shape as b, and a is not parameter, return a926    // otherwise, return a new tensor: repeat(a) to fit in b927    GGML_API struct ggml_tensor * ggml_repeat(928            struct ggml_context * ctx,929            struct ggml_tensor  * a,930            struct ggml_tensor  * b);931 932    // sums repetitions in a into shape of b933    GGML_API struct ggml_tensor * ggml_repeat_back(934            struct ggml_context * ctx,935            struct ggml_tensor  * a,936            struct ggml_tensor  * b);937 938    // concat a and b along dim939    // used in stable-diffusion940    GGML_API struct ggml_tensor * ggml_concat(941            struct ggml_context * ctx,942            struct ggml_tensor  * a,943            struct ggml_tensor  * b,944            int                   dim);945 946    GGML_API struct ggml_tensor * ggml_abs(947            struct ggml_context * ctx,948            struct ggml_tensor  * a);949 950    GGML_API struct ggml_tensor * ggml_abs_inplace(951            struct ggml_context * ctx,952            struct ggml_tensor  * a);953 954    GGML_API struct ggml_tensor * ggml_sgn(955            struct ggml_context * ctx,956            struct ggml_tensor  * a);957 958    GGML_API struct ggml_tensor * ggml_sgn_inplace(959            struct ggml_context * ctx,960            struct ggml_tensor  * a);961 962    GGML_API struct ggml_tensor * ggml_neg(963            struct ggml_context * ctx,964            struct ggml_tensor  * a);965 966    GGML_API struct ggml_tensor * ggml_neg_inplace(967            struct ggml_context * ctx,968            struct ggml_tensor  * a);969 970    GGML_API struct ggml_tensor * ggml_step(971            struct ggml_context * ctx,972            struct ggml_tensor  * a);973 974    GGML_API struct ggml_tensor * ggml_step_inplace(975            struct ggml_context * ctx,976            struct ggml_tensor  * a);977 978    GGML_API struct ggml_tensor * ggml_tanh(979            struct ggml_context * ctx,980            struct ggml_tensor  * a);981 982    GGML_API struct ggml_tensor * ggml_tanh_inplace(983            struct ggml_context * ctx,984            struct ggml_tensor  * a);985 986    GGML_API struct ggml_tensor * ggml_elu(987            struct ggml_context * ctx,988            struct ggml_tensor  * a);989 990    GGML_API struct ggml_tensor * ggml_elu_inplace(991            struct ggml_context * ctx,992            struct ggml_tensor  * a);993 994    GGML_API struct ggml_tensor * ggml_relu(995            struct ggml_context * ctx,996            struct ggml_tensor  * a);997 998    GGML_API struct ggml_tensor * ggml_leaky_relu(999            struct ggml_context * ctx,1000            struct ggml_tensor  * a, float negative_slope, bool inplace);1001 1002    GGML_API struct ggml_tensor * ggml_relu_inplace(1003            struct ggml_context * ctx,1004            struct ggml_tensor  * a);1005 1006    GGML_API struct ggml_tensor * ggml_sigmoid(1007            struct ggml_context * ctx,1008            struct ggml_tensor  * a);1009 1010    GGML_API struct ggml_tensor * ggml_sigmoid_inplace(1011            struct ggml_context * ctx,1012            struct ggml_tensor  * a);1013 1014    GGML_API struct ggml_tensor * ggml_gelu(1015            struct ggml_context * ctx,1016            struct ggml_tensor  * a);1017 1018    GGML_API struct ggml_tensor * ggml_gelu_inplace(1019            struct ggml_context * ctx,1020            struct ggml_tensor  * a);1021 1022    GGML_API struct ggml_tensor * ggml_gelu_quick(1023            struct ggml_context * ctx,1024            struct ggml_tensor  * a);1025 1026    GGML_API struct ggml_tensor * ggml_gelu_quick_inplace(1027            struct ggml_context * ctx,1028            struct ggml_tensor  * a);1029 1030    GGML_API struct ggml_tensor * ggml_silu(1031            struct ggml_context * ctx,1032            struct ggml_tensor  * a);1033 1034    GGML_API struct ggml_tensor * ggml_silu_inplace(1035            struct ggml_context * ctx,1036            struct ggml_tensor  * a);1037 1038    // a - x1039    // b - dy1040    GGML_API struct ggml_tensor * ggml_silu_back(1041            struct ggml_context * ctx,1042            struct ggml_tensor  * a,1043            struct ggml_tensor  * b);1044 1045    // hardswish(x) = x * relu6(x + 3) / 61046    GGML_API struct ggml_tensor * ggml_hardswish(1047            struct ggml_context * ctx,1048            struct ggml_tensor  * a);1049 1050    // hardsigmoid(x) = relu6(x + 3) / 61051    GGML_API struct ggml_tensor * ggml_hardsigmoid(1052            struct ggml_context * ctx,1053            struct ggml_tensor  * a);1054 1055    GGML_API struct ggml_tensor * ggml_exp(1056            struct ggml_context * ctx,1057            struct ggml_tensor  * a);1058 1059    GGML_API struct ggml_tensor * ggml_exp_inplace(1060            struct ggml_context * ctx,1061            struct ggml_tensor  * a);1062 1063    // normalize along rows1064    GGML_API struct ggml_tensor * ggml_norm(1065            struct ggml_context * ctx,1066            struct ggml_tensor  * a,1067            float                 eps);1068 1069    GGML_API struct ggml_tensor * ggml_norm_inplace(1070            struct ggml_context * ctx,1071            struct ggml_tensor  * a,1072            float                 eps);1073 1074    GGML_API struct ggml_tensor * ggml_rms_norm(1075            struct ggml_context * ctx,1076            struct ggml_tensor  * a,1077            float                 eps);1078 1079    GGML_API struct ggml_tensor * ggml_rms_norm_inplace(1080            struct ggml_context * ctx,1081            struct ggml_tensor  * a,1082            float                 eps);1083 1084    // group normalize along ne0*ne1*n_groups1085    // used in stable-diffusion1086    GGML_API struct ggml_tensor * ggml_group_norm(1087            struct ggml_context * ctx,1088            struct ggml_tensor  * a,1089            int                   n_groups,1090            float                 eps);1091 1092    GGML_API struct ggml_tensor * ggml_group_norm_inplace(1093            struct ggml_context * ctx,1094            struct ggml_tensor  * a,1095            int                   n_groups,1096            float                 eps);1097 1098    // a - x1099    // b - dy1100    GGML_API struct ggml_tensor * ggml_rms_norm_back(1101            struct ggml_context * ctx,1102            struct ggml_tensor  * a,1103            struct ggml_tensor  * b,1104            float                 eps);1105 1106    // A: k columns, n rows => [ne03, ne02, n, k]1107    // B: k columns, m rows  (i.e. we transpose it internally) => [ne03 * x, ne02 * y, m, k]1108    // result is n columns, m rows => [ne03 * x, ne02 * y, m, n]1109    GGML_API struct ggml_tensor * ggml_mul_mat(1110            struct ggml_context * ctx,1111            struct ggml_tensor  * a,1112            struct ggml_tensor  * b);1113 1114    // change the precision of a matrix multiplication1115    // set to GGML_PREC_F32 for higher precision (useful for phi-2)1116    GGML_API void ggml_mul_mat_set_prec(1117            struct ggml_tensor * a,1118            enum ggml_prec       prec);1119 1120    // indirect matrix multiplication1121    GGML_API struct ggml_tensor * ggml_mul_mat_id(1122            struct ggml_context * ctx,1123            struct ggml_tensor  * as,1124            struct ggml_tensor  * b,1125            struct ggml_tensor  * ids);1126 1127    // A: m columns, n rows,1128    // B: p columns, n rows,1129    // result is m columns, p rows1130    GGML_API struct ggml_tensor * ggml_out_prod(1131            struct ggml_context * ctx,1132            struct ggml_tensor  * a,1133            struct ggml_tensor  * b);1134 1135    //1136    // operations on tensors without backpropagation1137    //1138 1139    GGML_API struct ggml_tensor * ggml_scale(1140            struct ggml_context * ctx,1141            struct ggml_tensor  * a,1142            float                 s);1143 1144    // in-place, returns view(a)1145    GGML_API struct ggml_tensor * ggml_scale_inplace(1146            struct ggml_context * ctx,1147            struct ggml_tensor  * a,1148            float                 s);1149 1150    // b -> view(a,offset,nb1,nb2,3), return modified a1151    GGML_API struct ggml_tensor * ggml_set(1152            struct ggml_context * ctx,1153            struct ggml_tensor  * a,1154            struct ggml_tensor  * b,1155            size_t                nb1,1156            size_t                nb2,1157            size_t                nb3,1158            size_t                offset); // in bytes1159 1160    // b -> view(a,offset,nb1,nb2,3), return view(a)1161    GGML_API struct ggml_tensor * ggml_set_inplace(1162            struct ggml_context * ctx,1163            struct ggml_tensor  * a,1164            struct ggml_tensor  * b,1165            size_t                nb1,1166            size_t                nb2,1167            size_t                nb3,1168            size_t                offset); // in bytes1169 1170    GGML_API struct ggml_tensor * ggml_set_1d(1171            struct ggml_context * ctx,1172            struct ggml_tensor  * a,1173            struct ggml_tensor  * b,1174            size_t                offset); // in bytes1175 1176    GGML_API struct ggml_tensor * ggml_set_1d_inplace(1177            struct ggml_context * ctx,1178            struct ggml_tensor  * a,1179            struct ggml_tensor  * b,1180            size_t                offset); // in bytes1181 1182    // b -> view(a,offset,nb1,nb2,3), return modified a1183    GGML_API struct ggml_tensor * ggml_set_2d(1184            struct ggml_context * ctx,1185            struct ggml_tensor  * a,1186            struct ggml_tensor  * b,1187            size_t                nb1,1188            size_t                offset); // in bytes1189 1190    // b -> view(a,offset,nb1,nb2,3), return view(a)1191    GGML_API struct ggml_tensor * ggml_set_2d_inplace(1192            struct ggml_context * ctx,1193            struct ggml_tensor  * a,1194            struct ggml_tensor  * b,1195            size_t                nb1,1196            size_t                offset); // in bytes1197 1198    // a -> b, return view(b)1199    GGML_API struct ggml_tensor * ggml_cpy(1200            struct ggml_context * ctx,

Showing the first 1,200 of 2194 lines. Download the file for the rest.