Team Ai
Apppublic

KBaba7/llama.cpp

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes
README.md303 linesDownload Raw Back to llama-bench
1# llama.cpp/examples/llama-bench2 3Performance testing tool for llama.cpp.4 5## Table of contents6 71. [Syntax](#syntax)82. [Examples](#examples)9    1. [Text generation with different models](#text-generation-with-different-models)10    2. [Prompt processing with different batch sizes](#prompt-processing-with-different-batch-sizes)11    3. [Different numbers of threads](#different-numbers-of-threads)12    4. [Different numbers of layers offloaded to the GPU](#different-numbers-of-layers-offloaded-to-the-gpu)133. [Output formats](#output-formats)14    1. [Markdown](#markdown)15    2. [CSV](#csv)16    3. [JSON](#json)17    4. [JSONL](#jsonl)18    5. [SQL](#sql)19 20## Syntax21 22```23usage: ./llama-bench [options]24 25options:26  -h, --help27  -m, --model <filename>                    (default: models/7B/ggml-model-q4_0.gguf)28  -p, --n-prompt <n>                        (default: 512)29  -n, --n-gen <n>                           (default: 128)30  -pg <pp,tg>                               (default: )31  -b, --batch-size <n>                      (default: 2048)32  -ub, --ubatch-size <n>                    (default: 512)33  -ctk, --cache-type-k <t>                  (default: f16)34  -ctv, --cache-type-v <t>                  (default: f16)35  -t, --threads <n>                         (default: 8)36  -C, --cpu-mask <hex,hex>                  (default: 0x0)37  --cpu-strict <0|1>                        (default: 0)38  --poll <0...100>                          (default: 50)39  -ngl, --n-gpu-layers <n>                  (default: 99)40  -rpc, --rpc <rpc_servers>                 (default: )41  -sm, --split-mode <none|layer|row>        (default: layer)42  -mg, --main-gpu <i>                       (default: 0)43  -nkvo, --no-kv-offload <0|1>              (default: 0)44  -fa, --flash-attn <0|1>                   (default: 0)45  -mmp, --mmap <0|1>                        (default: 1)46  --numa <distribute|isolate|numactl>       (default: disabled)47  -embd, --embeddings <0|1>                 (default: 0)48  -ts, --tensor-split <ts0/ts1/..>          (default: 0)49  -r, --repetitions <n>                     (default: 5)50  --prio <0|1|2|3>                          (default: 0)51  --delay <0...N> (seconds)                 (default: 0)52  -o, --output <csv|json|jsonl|md|sql>      (default: md)53  -oe, --output-err <csv|json|jsonl|md|sql> (default: none)54  -v, --verbose                             (default: 0)55 56Multiple values can be given for each parameter by separating them with ',' or by specifying the parameter multiple times.57```58 59llama-bench can perform three types of tests:60 61- Prompt processing (pp): processing a prompt in batches (`-p`)62- Text generation (tg): generating a sequence of tokens (`-n`)63- Prompt processing + text generation (pg): processing a prompt followed by generating a sequence of tokens (`-pg`)64 65With the exception of `-r`, `-o` and `-v`, all options can be specified multiple times to run multiple tests. Each pp and tg test is run with all combinations of the specified options. To specify multiple values for an option, the values can be separated by commas (e.g. `-n 16,32`), or the option can be specified multiple times (e.g. `-n 16 -n 32`).66 67Each test is repeated the number of times given by `-r`, and the results are averaged. The results are given in average tokens per second (t/s) and standard deviation. Some output formats (e.g. json) also include the individual results of each repetition.68 69For a description of the other options, see the [main example](../main/README.md).70 71Note:72 73- When using SYCL backend, there would be hang issue in some cases. Please set `--mmp 0`.74 75## Examples76 77### Text generation with different models78 79```sh80$ ./llama-bench -m models/7B/ggml-model-q4_0.gguf -m models/13B/ggml-model-q4_0.gguf -p 0 -n 128,256,51281```82 83| model                          |       size |     params | backend    | ngl | test       |              t/s |84| ------------------------------ | ---------: | ---------: | ---------- | --: | ---------- | ---------------: |85| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  99 | tg 128     |    132.19 ± 0.55 |86| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  99 | tg 256     |    129.37 ± 0.54 |87| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  99 | tg 512     |    123.83 ± 0.25 |88| llama 13B mostly Q4_0          |   6.86 GiB |    13.02 B | CUDA       |  99 | tg 128     |     82.17 ± 0.31 |89| llama 13B mostly Q4_0          |   6.86 GiB |    13.02 B | CUDA       |  99 | tg 256     |     80.74 ± 0.23 |90| llama 13B mostly Q4_0          |   6.86 GiB |    13.02 B | CUDA       |  99 | tg 512     |     78.08 ± 0.07 |91 92### Prompt processing with different batch sizes93 94```sh95$ ./llama-bench -n 0 -p 1024 -b 128,256,512,102496```97 98| model                          |       size |     params | backend    | ngl |    n_batch | test       |              t/s |99| ------------------------------ | ---------: | ---------: | ---------- | --: | ---------: | ---------- | ---------------: |100| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  99 |        128 | pp 1024    |   1436.51 ± 3.66 |101| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  99 |        256 | pp 1024    |  1932.43 ± 23.48 |102| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  99 |        512 | pp 1024    |  2254.45 ± 15.59 |103| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  99 |       1024 | pp 1024    |  2498.61 ± 13.58 |104 105### Different numbers of threads106 107```sh108$ ./llama-bench -n 0 -n 16 -p 64 -t 1,2,4,8,16,32109```110 111| model                          |       size |     params | backend    |    threads | test       |              t/s |112| ------------------------------ | ---------: | ---------: | ---------- | ---------: | ---------- | ---------------: |113| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |          1 | pp 64      |      6.17 ± 0.07 |114| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |          1 | tg 16      |      4.05 ± 0.02 |115| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |          2 | pp 64      |     12.31 ± 0.13 |116| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |          2 | tg 16      |      7.80 ± 0.07 |117| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |          4 | pp 64      |     23.18 ± 0.06 |118| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |          4 | tg 16      |     12.22 ± 0.07 |119| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |          8 | pp 64      |     32.29 ± 1.21 |120| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |          8 | tg 16      |     16.71 ± 0.66 |121| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |         16 | pp 64      |     33.52 ± 0.03 |122| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |         16 | tg 16      |     15.32 ± 0.05 |123| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |         32 | pp 64      |     59.00 ± 1.11 |124| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CPU        |         32 | tg 16      |     16.41 ± 0.79 ||125 126### Different numbers of layers offloaded to the GPU127 128```sh129$ ./llama-bench -ngl 10,20,30,31,32,33,34,35130```131 132| model                          |       size |     params | backend    | ngl | test       |              t/s |133| ------------------------------ | ---------: | ---------: | ---------- | --: | ---------- | ---------------: |134| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  10 | pp 512     |    373.36 ± 2.25 |135| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  10 | tg 128     |     13.45 ± 0.93 |136| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  20 | pp 512     |    472.65 ± 1.25 |137| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  20 | tg 128     |     21.36 ± 1.94 |138| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  30 | pp 512     |   631.87 ± 11.25 |139| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  30 | tg 128     |     40.04 ± 1.82 |140| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  31 | pp 512     |    657.89 ± 5.08 |141| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  31 | tg 128     |     48.19 ± 0.81 |142| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  32 | pp 512     |    688.26 ± 3.29 |143| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  32 | tg 128     |     54.78 ± 0.65 |144| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  33 | pp 512     |    704.27 ± 2.24 |145| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  33 | tg 128     |     60.62 ± 1.76 |146| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  34 | pp 512     |    881.34 ± 5.40 |147| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  34 | tg 128     |     71.76 ± 0.23 |148| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  35 | pp 512     |   2400.01 ± 7.72 |149| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  35 | tg 128     |    131.66 ± 0.49 |150 151## Output formats152 153By default, llama-bench outputs the results in markdown format. The results can be output in other formats by using the `-o` option.154 155### Markdown156 157```sh158$ ./llama-bench -o md159```160 161| model                          |       size |     params | backend    | ngl | test       |              t/s |162| ------------------------------ | ---------: | ---------: | ---------- | --: | ---------- | ---------------: |163| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  99 | pp 512     |  2368.80 ± 93.24 |164| llama 7B mostly Q4_0           |   3.56 GiB |     6.74 B | CUDA       |  99 | tg 128     |    131.42 ± 0.59 |165 166### CSV167 168```sh169$ ./llama-bench -o csv170```171 172```csv173build_commit,build_number,cuda,metal,gpu_blas,blas,cpu_info,gpu_info,model_filename,model_type,model_size,model_n_params,n_batch,n_threads,f16_kv,n_gpu_layers,main_gpu,mul_mat_q,tensor_split,n_prompt,n_gen,test_time,avg_ns,stddev_ns,avg_ts,stddev_ts174"3469684","1275","1","0","0","1","1","13th Gen Intel(R) Core(TM) i9-13900K","NVIDIA GeForce RTX 3090 Ti","models/7B/ggml-model-q4_0.gguf","llama 7B mostly Q4_0","3825065984","6738415616","512","16","1","99","0","1","0.00","512","0","2023-09-23T12:09:01Z","212155977","732372","2413.341687","8.305961"175"3469684","1275","1","0","0","1","1","13th Gen Intel(R) Core(TM) i9-13900K","NVIDIA GeForce RTX 3090 Ti","models/7B/ggml-model-q4_0.gguf","llama 7B mostly Q4_0","3825065984","6738415616","512","16","1","99","0","1","0.00","0","128","2023-09-23T12:09:02Z","969320879","2728399","132.052051","0.371342"176```177 178### JSON179 180```sh181$ ./llama-bench -o json182```183 184```json185[186  {187    "build_commit": "3469684",188    "build_number": 1275,189    "cuda": true,190    "metal": false,191    "gpu_blas": true,192    "blas": true,193    "cpu_info": "13th Gen Intel(R) Core(TM) i9-13900K",194    "gpu_info": "NVIDIA GeForce RTX 3090 Ti",195    "model_filename": "models/7B/ggml-model-q4_0.gguf",196    "model_type": "llama 7B mostly Q4_0",197    "model_size": 3825065984,198    "model_n_params": 6738415616,199    "n_batch": 512,200    "n_threads": 16,201    "f16_kv": true,202    "n_gpu_layers": 99,203    "main_gpu": 0,204    "mul_mat_q": true,205    "tensor_split": "0.00",206    "n_prompt": 512,207    "n_gen": 0,208    "test_time": "2023-09-23T12:09:57Z",209    "avg_ns": 212365953,210    "stddev_ns": 985423,211    "avg_ts": 2410.974041,212    "stddev_ts": 11.163766,213    "samples_ns": [ 213837238, 211635853, 212328053, 211329715, 212698907 ],214    "samples_ts": [ 2394.34, 2419.25, 2411.36, 2422.75, 2407.16 ]215  },216  {217    "build_commit": "3469684",218    "build_number": 1275,219    "cuda": true,220    "metal": false,221    "gpu_blas": true,222    "blas": true,223    "cpu_info": "13th Gen Intel(R) Core(TM) i9-13900K",224    "gpu_info": "NVIDIA GeForce RTX 3090 Ti",225    "model_filename": "models/7B/ggml-model-q4_0.gguf",226    "model_type": "llama 7B mostly Q4_0",227    "model_size": 3825065984,228    "model_n_params": 6738415616,229    "n_batch": 512,230    "n_threads": 16,231    "f16_kv": true,232    "n_gpu_layers": 99,233    "main_gpu": 0,234    "mul_mat_q": true,235    "tensor_split": "0.00",236    "n_prompt": 0,237    "n_gen": 128,238    "test_time": "2023-09-23T12:09:59Z",239    "avg_ns": 977425219,240    "stddev_ns": 9268593,241    "avg_ts": 130.965708,242    "stddev_ts": 1.238924,243    "samples_ns": [ 984472709, 974901233, 989474741, 970729355, 967548060 ],244    "samples_ts": [ 130.019, 131.295, 129.362, 131.86, 132.293 ]245  }246]247```248 249 250### JSONL251 252```sh253$ ./llama-bench -o jsonl254```255 256```json lines257{"build_commit":"3469684","build_number":1275,"cuda":true,"metal":false,"gpu_blas":true,"blas":true,"cpu_info":"13th Gen Intel(R) Core(TM) i9-13900K","gpu_info":"NVIDIA GeForce RTX 3090 Ti","model_filename":"models/7B/ggml-model-q4_0.gguf","model_type":"llama 7B mostly Q4_0","model_size":3825065984,"model_n_params":6738415616,"n_batch":512,"n_threads":16,"f16_kv":true,"n_gpu_layers":99,"main_gpu":0,"mul_mat_q":true,"tensor_split":"0.00","n_prompt":512,"n_gen":0,"test_time":"2023-09-23T12:09:57Z","avg_ns":212365953,"stddev_ns":985423,"avg_ts":2410.974041,"stddev_ts":11.163766,"samples_ns":[213837238,211635853,212328053,211329715,212698907],"samples_ts":[2394.34,2419.25,2411.36,2422.75,2407.16]}258{"build_commit":"3469684","build_number":1275,"cuda":true,"metal":false,"gpu_blas":true,"blas":true,"cpu_info":"13th Gen Intel(R) Core(TM) i9-13900K","gpu_info":"NVIDIA GeForce RTX 3090 Ti","model_filename":"models/7B/ggml-model-q4_0.gguf","model_type":"llama 7B mostly Q4_0","model_size":3825065984,"model_n_params":6738415616,"n_batch":512,"n_threads":16,"f16_kv":true,"n_gpu_layers":99,"main_gpu":0,"mul_mat_q":true,"tensor_split":"0.00","n_prompt":0,"n_gen":128,"test_time":"2023-09-23T12:09:59Z","avg_ns":977425219,"stddev_ns":9268593,"avg_ts":130.965708,"stddev_ts":1.238924,"samples_ns":[984472709,974901233,989474741,970729355,967548060],"samples_ts":[130.019,131.295,129.362,131.86,132.293]}259```260 261 262### SQL263 264SQL output is suitable for importing into a SQLite database. The output can be piped into the `sqlite3` command line tool to add the results to a database.265 266```sh267$ ./llama-bench -o sql268```269 270```sql271CREATE TABLE IF NOT EXISTS test (272  build_commit TEXT,273  build_number INTEGER,274  cuda INTEGER,275  metal INTEGER,276  gpu_blas INTEGER,277  blas INTEGER,278  cpu_info TEXT,279  gpu_info TEXT,280  model_filename TEXT,281  model_type TEXT,282  model_size INTEGER,283  model_n_params INTEGER,284  n_batch INTEGER,285  n_threads INTEGER,286  f16_kv INTEGER,287  n_gpu_layers INTEGER,288  main_gpu INTEGER,289  mul_mat_q INTEGER,290  tensor_split TEXT,291  n_prompt INTEGER,292  n_gen INTEGER,293  test_time TEXT,294  avg_ns INTEGER,295  stddev_ns INTEGER,296  avg_ts REAL,297  stddev_ts REAL298);299 300INSERT INTO test (build_commit, build_number, cuda, metal, gpu_blas, blas, cpu_info, gpu_info, model_filename, model_type, model_size, model_n_params, n_batch, n_threads, f16_kv, n_gpu_layers, main_gpu, mul_mat_q, tensor_split, n_prompt, n_gen, test_time, avg_ns, stddev_ns, avg_ts, stddev_ts) VALUES ('3469684', '1275', '1', '0', '0', '1', '1', '13th Gen Intel(R) Core(TM) i9-13900K', 'NVIDIA GeForce RTX 3090 Ti', 'models/7B/ggml-model-q4_0.gguf', 'llama 7B mostly Q4_0', '3825065984', '6738415616', '512', '16', '1', '99', '0', '1', '0.00', '512', '0', '2023-09-23T12:10:30Z', '212693772', '743623', '2407.240204', '8.409634');301INSERT INTO test (build_commit, build_number, cuda, metal, gpu_blas, blas, cpu_info, gpu_info, model_filename, model_type, model_size, model_n_params, n_batch, n_threads, f16_kv, n_gpu_layers, main_gpu, mul_mat_q, tensor_split, n_prompt, n_gen, test_time, avg_ns, stddev_ns, avg_ts, stddev_ts) VALUES ('3469684', '1275', '1', '0', '0', '1', '1', '13th Gen Intel(R) Core(TM) i9-13900K', 'NVIDIA GeForce RTX 3090 Ti', 'models/7B/ggml-model-q4_0.gguf', 'llama 7B mostly Q4_0', '3825065984', '6738415616', '512', '16', '1', '99', '0', '1', '0.00', '0', '128', '2023-09-23T12:10:31Z', '977925003', '4037361', '130.891159', '0.537692');302```303