KBaba7/llama.cpp
0
1# llama.cpp/examples/llama-bench2 3Performance testing tool for llama.cpp.4 5## Table of contents6 71. [Syntax](#syntax)82. [Examples](#examples)9 1. [Text generation with different models](#text-generation-with-different-models)10 2. [Prompt processing with different batch sizes](#prompt-processing-with-different-batch-sizes)11 3. [Different numbers of threads](#different-numbers-of-threads)12 4. [Different numbers of layers offloaded to the GPU](#different-numbers-of-layers-offloaded-to-the-gpu)133. [Output formats](#output-formats)14 1. [Markdown](#markdown)15 2. [CSV](#csv)16 3. [JSON](#json)17 4. [JSONL](#jsonl)18 5. [SQL](#sql)19 20## Syntax21 22```23usage: ./llama-bench [options]24 25options:26 -h, --help27 -m, --model <filename> (default: models/7B/ggml-model-q4_0.gguf)28 -p, --n-prompt <n> (default: 512)29 -n, --n-gen <n> (default: 128)30 -pg <pp,tg> (default: )31 -b, --batch-size <n> (default: 2048)32 -ub, --ubatch-size <n> (default: 512)33 -ctk, --cache-type-k <t> (default: f16)34 -ctv, --cache-type-v <t> (default: f16)35 -t, --threads <n> (default: 8)36 -C, --cpu-mask <hex,hex> (default: 0x0)37 --cpu-strict <0|1> (default: 0)38 --poll <0...100> (default: 50)39 -ngl, --n-gpu-layers <n> (default: 99)40 -rpc, --rpc <rpc_servers> (default: )41 -sm, --split-mode <none|layer|row> (default: layer)42 -mg, --main-gpu <i> (default: 0)43 -nkvo, --no-kv-offload <0|1> (default: 0)44 -fa, --flash-attn <0|1> (default: 0)45 -mmp, --mmap <0|1> (default: 1)46 --numa <distribute|isolate|numactl> (default: disabled)47 -embd, --embeddings <0|1> (default: 0)48 -ts, --tensor-split <ts0/ts1/..> (default: 0)49 -r, --repetitions <n> (default: 5)50 --prio <0|1|2|3> (default: 0)51 --delay <0...N> (seconds) (default: 0)52 -o, --output <csv|json|jsonl|md|sql> (default: md)53 -oe, --output-err <csv|json|jsonl|md|sql> (default: none)54 -v, --verbose (default: 0)55 56Multiple values can be given for each parameter by separating them with ',' or by specifying the parameter multiple times.57```58 59llama-bench can perform three types of tests:60 61- Prompt processing (pp): processing a prompt in batches (`-p`)62- Text generation (tg): generating a sequence of tokens (`-n`)63- Prompt processing + text generation (pg): processing a prompt followed by generating a sequence of tokens (`-pg`)64 65With the exception of `-r`, `-o` and `-v`, all options can be specified multiple times to run multiple tests. Each pp and tg test is run with all combinations of the specified options. To specify multiple values for an option, the values can be separated by commas (e.g. `-n 16,32`), or the option can be specified multiple times (e.g. `-n 16 -n 32`).66 67Each test is repeated the number of times given by `-r`, and the results are averaged. The results are given in average tokens per second (t/s) and standard deviation. Some output formats (e.g. json) also include the individual results of each repetition.68 69For a description of the other options, see the [main example](../main/README.md).70 71Note:72 73- When using SYCL backend, there would be hang issue in some cases. Please set `--mmp 0`.74 75## Examples76 77### Text generation with different models78 79```sh80$ ./llama-bench -m models/7B/ggml-model-q4_0.gguf -m models/13B/ggml-model-q4_0.gguf -p 0 -n 128,256,51281```82 83| model | size | params | backend | ngl | test | t/s |84| ------------------------------ | ---------: | ---------: | ---------- | --: | ---------- | ---------------: |85| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 99 | tg 128 | 132.19 ± 0.55 |86| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 99 | tg 256 | 129.37 ± 0.54 |87| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 99 | tg 512 | 123.83 ± 0.25 |88| llama 13B mostly Q4_0 | 6.86 GiB | 13.02 B | CUDA | 99 | tg 128 | 82.17 ± 0.31 |89| llama 13B mostly Q4_0 | 6.86 GiB | 13.02 B | CUDA | 99 | tg 256 | 80.74 ± 0.23 |90| llama 13B mostly Q4_0 | 6.86 GiB | 13.02 B | CUDA | 99 | tg 512 | 78.08 ± 0.07 |91 92### Prompt processing with different batch sizes93 94```sh95$ ./llama-bench -n 0 -p 1024 -b 128,256,512,102496```97 98| model | size | params | backend | ngl | n_batch | test | t/s |99| ------------------------------ | ---------: | ---------: | ---------- | --: | ---------: | ---------- | ---------------: |100| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 99 | 128 | pp 1024 | 1436.51 ± 3.66 |101| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 99 | 256 | pp 1024 | 1932.43 ± 23.48 |102| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 99 | 512 | pp 1024 | 2254.45 ± 15.59 |103| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 99 | 1024 | pp 1024 | 2498.61 ± 13.58 |104 105### Different numbers of threads106 107```sh108$ ./llama-bench -n 0 -n 16 -p 64 -t 1,2,4,8,16,32109```110 111| model | size | params | backend | threads | test | t/s |112| ------------------------------ | ---------: | ---------: | ---------- | ---------: | ---------- | ---------------: |113| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 1 | pp 64 | 6.17 ± 0.07 |114| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 1 | tg 16 | 4.05 ± 0.02 |115| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 2 | pp 64 | 12.31 ± 0.13 |116| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 2 | tg 16 | 7.80 ± 0.07 |117| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 4 | pp 64 | 23.18 ± 0.06 |118| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 4 | tg 16 | 12.22 ± 0.07 |119| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 8 | pp 64 | 32.29 ± 1.21 |120| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 8 | tg 16 | 16.71 ± 0.66 |121| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 16 | pp 64 | 33.52 ± 0.03 |122| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 16 | tg 16 | 15.32 ± 0.05 |123| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 32 | pp 64 | 59.00 ± 1.11 |124| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CPU | 32 | tg 16 | 16.41 ± 0.79 ||125 126### Different numbers of layers offloaded to the GPU127 128```sh129$ ./llama-bench -ngl 10,20,30,31,32,33,34,35130```131 132| model | size | params | backend | ngl | test | t/s |133| ------------------------------ | ---------: | ---------: | ---------- | --: | ---------- | ---------------: |134| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 10 | pp 512 | 373.36 ± 2.25 |135| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 10 | tg 128 | 13.45 ± 0.93 |136| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 20 | pp 512 | 472.65 ± 1.25 |137| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 20 | tg 128 | 21.36 ± 1.94 |138| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 30 | pp 512 | 631.87 ± 11.25 |139| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 30 | tg 128 | 40.04 ± 1.82 |140| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 31 | pp 512 | 657.89 ± 5.08 |141| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 31 | tg 128 | 48.19 ± 0.81 |142| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 32 | pp 512 | 688.26 ± 3.29 |143| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 32 | tg 128 | 54.78 ± 0.65 |144| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 33 | pp 512 | 704.27 ± 2.24 |145| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 33 | tg 128 | 60.62 ± 1.76 |146| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 34 | pp 512 | 881.34 ± 5.40 |147| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 34 | tg 128 | 71.76 ± 0.23 |148| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 35 | pp 512 | 2400.01 ± 7.72 |149| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 35 | tg 128 | 131.66 ± 0.49 |150 151## Output formats152 153By default, llama-bench outputs the results in markdown format. The results can be output in other formats by using the `-o` option.154 155### Markdown156 157```sh158$ ./llama-bench -o md159```160 161| model | size | params | backend | ngl | test | t/s |162| ------------------------------ | ---------: | ---------: | ---------- | --: | ---------- | ---------------: |163| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 99 | pp 512 | 2368.80 ± 93.24 |164| llama 7B mostly Q4_0 | 3.56 GiB | 6.74 B | CUDA | 99 | tg 128 | 131.42 ± 0.59 |165 166### CSV167 168```sh169$ ./llama-bench -o csv170```171 172```csv173build_commit,build_number,cuda,metal,gpu_blas,blas,cpu_info,gpu_info,model_filename,model_type,model_size,model_n_params,n_batch,n_threads,f16_kv,n_gpu_layers,main_gpu,mul_mat_q,tensor_split,n_prompt,n_gen,test_time,avg_ns,stddev_ns,avg_ts,stddev_ts174"3469684","1275","1","0","0","1","1","13th Gen Intel(R) Core(TM) i9-13900K","NVIDIA GeForce RTX 3090 Ti","models/7B/ggml-model-q4_0.gguf","llama 7B mostly Q4_0","3825065984","6738415616","512","16","1","99","0","1","0.00","512","0","2023-09-23T12:09:01Z","212155977","732372","2413.341687","8.305961"175"3469684","1275","1","0","0","1","1","13th Gen Intel(R) Core(TM) i9-13900K","NVIDIA GeForce RTX 3090 Ti","models/7B/ggml-model-q4_0.gguf","llama 7B mostly Q4_0","3825065984","6738415616","512","16","1","99","0","1","0.00","0","128","2023-09-23T12:09:02Z","969320879","2728399","132.052051","0.371342"176```177 178### JSON179 180```sh181$ ./llama-bench -o json182```183 184```json185[186 {187 "build_commit": "3469684",188 "build_number": 1275,189 "cuda": true,190 "metal": false,191 "gpu_blas": true,192 "blas": true,193 "cpu_info": "13th Gen Intel(R) Core(TM) i9-13900K",194 "gpu_info": "NVIDIA GeForce RTX 3090 Ti",195 "model_filename": "models/7B/ggml-model-q4_0.gguf",196 "model_type": "llama 7B mostly Q4_0",197 "model_size": 3825065984,198 "model_n_params": 6738415616,199 "n_batch": 512,200 "n_threads": 16,201 "f16_kv": true,202 "n_gpu_layers": 99,203 "main_gpu": 0,204 "mul_mat_q": true,205 "tensor_split": "0.00",206 "n_prompt": 512,207 "n_gen": 0,208 "test_time": "2023-09-23T12:09:57Z",209 "avg_ns": 212365953,210 "stddev_ns": 985423,211 "avg_ts": 2410.974041,212 "stddev_ts": 11.163766,213 "samples_ns": [ 213837238, 211635853, 212328053, 211329715, 212698907 ],214 "samples_ts": [ 2394.34, 2419.25, 2411.36, 2422.75, 2407.16 ]215 },216 {217 "build_commit": "3469684",218 "build_number": 1275,219 "cuda": true,220 "metal": false,221 "gpu_blas": true,222 "blas": true,223 "cpu_info": "13th Gen Intel(R) Core(TM) i9-13900K",224 "gpu_info": "NVIDIA GeForce RTX 3090 Ti",225 "model_filename": "models/7B/ggml-model-q4_0.gguf",226 "model_type": "llama 7B mostly Q4_0",227 "model_size": 3825065984,228 "model_n_params": 6738415616,229 "n_batch": 512,230 "n_threads": 16,231 "f16_kv": true,232 "n_gpu_layers": 99,233 "main_gpu": 0,234 "mul_mat_q": true,235 "tensor_split": "0.00",236 "n_prompt": 0,237 "n_gen": 128,238 "test_time": "2023-09-23T12:09:59Z",239 "avg_ns": 977425219,240 "stddev_ns": 9268593,241 "avg_ts": 130.965708,242 "stddev_ts": 1.238924,243 "samples_ns": [ 984472709, 974901233, 989474741, 970729355, 967548060 ],244 "samples_ts": [ 130.019, 131.295, 129.362, 131.86, 132.293 ]245 }246]247```248 249 250### JSONL251 252```sh253$ ./llama-bench -o jsonl254```255 256```json lines257{"build_commit":"3469684","build_number":1275,"cuda":true,"metal":false,"gpu_blas":true,"blas":true,"cpu_info":"13th Gen Intel(R) Core(TM) i9-13900K","gpu_info":"NVIDIA GeForce RTX 3090 Ti","model_filename":"models/7B/ggml-model-q4_0.gguf","model_type":"llama 7B mostly Q4_0","model_size":3825065984,"model_n_params":6738415616,"n_batch":512,"n_threads":16,"f16_kv":true,"n_gpu_layers":99,"main_gpu":0,"mul_mat_q":true,"tensor_split":"0.00","n_prompt":512,"n_gen":0,"test_time":"2023-09-23T12:09:57Z","avg_ns":212365953,"stddev_ns":985423,"avg_ts":2410.974041,"stddev_ts":11.163766,"samples_ns":[213837238,211635853,212328053,211329715,212698907],"samples_ts":[2394.34,2419.25,2411.36,2422.75,2407.16]}258{"build_commit":"3469684","build_number":1275,"cuda":true,"metal":false,"gpu_blas":true,"blas":true,"cpu_info":"13th Gen Intel(R) Core(TM) i9-13900K","gpu_info":"NVIDIA GeForce RTX 3090 Ti","model_filename":"models/7B/ggml-model-q4_0.gguf","model_type":"llama 7B mostly Q4_0","model_size":3825065984,"model_n_params":6738415616,"n_batch":512,"n_threads":16,"f16_kv":true,"n_gpu_layers":99,"main_gpu":0,"mul_mat_q":true,"tensor_split":"0.00","n_prompt":0,"n_gen":128,"test_time":"2023-09-23T12:09:59Z","avg_ns":977425219,"stddev_ns":9268593,"avg_ts":130.965708,"stddev_ts":1.238924,"samples_ns":[984472709,974901233,989474741,970729355,967548060],"samples_ts":[130.019,131.295,129.362,131.86,132.293]}259```260 261 262### SQL263 264SQL output is suitable for importing into a SQLite database. The output can be piped into the `sqlite3` command line tool to add the results to a database.265 266```sh267$ ./llama-bench -o sql268```269 270```sql271CREATE TABLE IF NOT EXISTS test (272 build_commit TEXT,273 build_number INTEGER,274 cuda INTEGER,275 metal INTEGER,276 gpu_blas INTEGER,277 blas INTEGER,278 cpu_info TEXT,279 gpu_info TEXT,280 model_filename TEXT,281 model_type TEXT,282 model_size INTEGER,283 model_n_params INTEGER,284 n_batch INTEGER,285 n_threads INTEGER,286 f16_kv INTEGER,287 n_gpu_layers INTEGER,288 main_gpu INTEGER,289 mul_mat_q INTEGER,290 tensor_split TEXT,291 n_prompt INTEGER,292 n_gen INTEGER,293 test_time TEXT,294 avg_ns INTEGER,295 stddev_ns INTEGER,296 avg_ts REAL,297 stddev_ts REAL298);299 300INSERT INTO test (build_commit, build_number, cuda, metal, gpu_blas, blas, cpu_info, gpu_info, model_filename, model_type, model_size, model_n_params, n_batch, n_threads, f16_kv, n_gpu_layers, main_gpu, mul_mat_q, tensor_split, n_prompt, n_gen, test_time, avg_ns, stddev_ns, avg_ts, stddev_ts) VALUES ('3469684', '1275', '1', '0', '0', '1', '1', '13th Gen Intel(R) Core(TM) i9-13900K', 'NVIDIA GeForce RTX 3090 Ti', 'models/7B/ggml-model-q4_0.gguf', 'llama 7B mostly Q4_0', '3825065984', '6738415616', '512', '16', '1', '99', '0', '1', '0.00', '512', '0', '2023-09-23T12:10:30Z', '212693772', '743623', '2407.240204', '8.409634');301INSERT INTO test (build_commit, build_number, cuda, metal, gpu_blas, blas, cpu_info, gpu_info, model_filename, model_type, model_size, model_n_params, n_batch, n_threads, f16_kv, n_gpu_layers, main_gpu, mul_mat_q, tensor_split, n_prompt, n_gen, test_time, avg_ns, stddev_ns, avg_ts, stddev_ts) VALUES ('3469684', '1275', '1', '0', '0', '1', '1', '13th Gen Intel(R) Core(TM) i9-13900K', 'NVIDIA GeForce RTX 3090 Ti', 'models/7B/ggml-model-q4_0.gguf', 'llama 7B mostly Q4_0', '3825065984', '6738415616', '512', '16', '1', '99', '0', '1', '0.00', '0', '128', '2023-09-23T12:10:31Z', '977925003', '4037361', '130.891159', '0.537692');302```303 