KBaba7/llama.cpp
0
1# LLaMA.cpp HTTP Server2 3Fast, lightweight, pure C/C++ HTTP server based on [httplib](https://github.com/yhirose/cpp-httplib), [nlohmann::json](https://github.com/nlohmann/json) and **llama.cpp**.4 5Set of LLM REST APIs and a simple web front end to interact with llama.cpp.6 7**Features:**8 * LLM inference of F16 and quantized models on GPU and CPU9 * [OpenAI API](https://github.com/openai/openai-openapi) compatible chat completions and embeddings routes10 * Reranking endoint (WIP: https://github.com/ggerganov/llama.cpp/pull/9510)11 * Parallel decoding with multi-user support12 * Continuous batching13 * Multimodal (wip)14 * Monitoring endpoints15 * Schema-constrained JSON response format16 17The project is under active development, and we are [looking for feedback and contributors](https://github.com/ggerganov/llama.cpp/issues/4216).18 19## Usage20 21<!-- Note for contributors: The list below is generated by llama-gen-docs -->22 23**Common params**24 25| Argument | Explanation |26| -------- | ----------- |27| `-h, --help, --usage` | print usage and exit |28| `--version` | show version and build info |29| `--verbose-prompt` | print a verbose prompt before generation (default: false) |30| `-t, --threads N` | number of threads to use during generation (default: -1)<br/>(env: LLAMA_ARG_THREADS) |31| `-tb, --threads-batch N` | number of threads to use during batch and prompt processing (default: same as --threads) |32| `-C, --cpu-mask M` | CPU affinity mask: arbitrarily long hex. Complements cpu-range (default: "") |33| `-Cr, --cpu-range lo-hi` | range of CPUs for affinity. Complements --cpu-mask |34| `--cpu-strict <0\|1>` | use strict CPU placement (default: 0)<br/> |35| `--prio N` | set process/thread priority : 0-normal, 1-medium, 2-high, 3-realtime (default: 0)<br/> |36| `--poll <0...100>` | use polling level to wait for work (0 - no polling, default: 50)<br/> |37| `-Cb, --cpu-mask-batch M` | CPU affinity mask: arbitrarily long hex. Complements cpu-range-batch (default: same as --cpu-mask) |38| `-Crb, --cpu-range-batch lo-hi` | ranges of CPUs for affinity. Complements --cpu-mask-batch |39| `--cpu-strict-batch <0\|1>` | use strict CPU placement (default: same as --cpu-strict) |40| `--prio-batch N` | set process/thread priority : 0-normal, 1-medium, 2-high, 3-realtime (default: 0)<br/> |41| `--poll-batch <0\|1>` | use polling to wait for work (default: same as --poll) |42| `-c, --ctx-size N` | size of the prompt context (default: 4096, 0 = loaded from model)<br/>(env: LLAMA_ARG_CTX_SIZE) |43| `-n, --predict, --n-predict N` | number of tokens to predict (default: -1, -1 = infinity, -2 = until context filled)<br/>(env: LLAMA_ARG_N_PREDICT) |44| `-b, --batch-size N` | logical maximum batch size (default: 2048)<br/>(env: LLAMA_ARG_BATCH) |45| `-ub, --ubatch-size N` | physical maximum batch size (default: 512)<br/>(env: LLAMA_ARG_UBATCH) |46| `--keep N` | number of tokens to keep from the initial prompt (default: 0, -1 = all) |47| `-fa, --flash-attn` | enable Flash Attention (default: disabled)<br/>(env: LLAMA_ARG_FLASH_ATTN) |48| `--no-perf` | disable internal libllama performance timings (default: false)<br/>(env: LLAMA_ARG_NO_PERF) |49| `-e, --escape` | process escapes sequences (\n, \r, \t, \', \", \\) (default: true) |50| `--no-escape` | do not process escape sequences |51| `--rope-scaling {none,linear,yarn}` | RoPE frequency scaling method, defaults to linear unless specified by the model<br/>(env: LLAMA_ARG_ROPE_SCALING_TYPE) |52| `--rope-scale N` | RoPE context scaling factor, expands context by a factor of N<br/>(env: LLAMA_ARG_ROPE_SCALE) |53| `--rope-freq-base N` | RoPE base frequency, used by NTK-aware scaling (default: loaded from model)<br/>(env: LLAMA_ARG_ROPE_FREQ_BASE) |54| `--rope-freq-scale N` | RoPE frequency scaling factor, expands context by a factor of 1/N<br/>(env: LLAMA_ARG_ROPE_FREQ_SCALE) |55| `--yarn-orig-ctx N` | YaRN: original context size of model (default: 0 = model training context size)<br/>(env: LLAMA_ARG_YARN_ORIG_CTX) |56| `--yarn-ext-factor N` | YaRN: extrapolation mix factor (default: -1.0, 0.0 = full interpolation)<br/>(env: LLAMA_ARG_YARN_EXT_FACTOR) |57| `--yarn-attn-factor N` | YaRN: scale sqrt(t) or attention magnitude (default: 1.0)<br/>(env: LLAMA_ARG_YARN_ATTN_FACTOR) |58| `--yarn-beta-slow N` | YaRN: high correction dim or alpha (default: 1.0)<br/>(env: LLAMA_ARG_YARN_BETA_SLOW) |59| `--yarn-beta-fast N` | YaRN: low correction dim or beta (default: 32.0)<br/>(env: LLAMA_ARG_YARN_BETA_FAST) |60| `-dkvc, --dump-kv-cache` | verbose print of the KV cache |61| `-nkvo, --no-kv-offload` | disable KV offload<br/>(env: LLAMA_ARG_NO_KV_OFFLOAD) |62| `-ctk, --cache-type-k TYPE` | KV cache data type for K<br/>allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1<br/>(default: f16)<br/>(env: LLAMA_ARG_CACHE_TYPE_K) |63| `-ctv, --cache-type-v TYPE` | KV cache data type for V<br/>allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1<br/>(default: f16)<br/>(env: LLAMA_ARG_CACHE_TYPE_V) |64| `-dt, --defrag-thold N` | KV cache defragmentation threshold (default: 0.1, < 0 - disabled)<br/>(env: LLAMA_ARG_DEFRAG_THOLD) |65| `-np, --parallel N` | number of parallel sequences to decode (default: 1)<br/>(env: LLAMA_ARG_N_PARALLEL) |66| `--mlock` | force system to keep model in RAM rather than swapping or compressing<br/>(env: LLAMA_ARG_MLOCK) |67| `--no-mmap` | do not memory-map model (slower load but may reduce pageouts if not using mlock)<br/>(env: LLAMA_ARG_NO_MMAP) |68| `--numa TYPE` | attempt optimizations that help on some NUMA systems<br/>- distribute: spread execution evenly over all nodes<br/>- isolate: only spawn threads on CPUs on the node that execution started on<br/>- numactl: use the CPU map provided by numactl<br/>if run without this previously, it is recommended to drop the system page cache before using this<br/>see https://github.com/ggerganov/llama.cpp/issues/1437<br/>(env: LLAMA_ARG_NUMA) |69| `-dev, --device <dev1,dev2,..>` | comma-separated list of devices to use for offloading (none = don't offload)<br/>use --list-devices to see a list of available devices<br/>(env: LLAMA_ARG_DEVICE) |70| `--list-devices` | print list of available devices and exit |71| `-ngl, --gpu-layers, --n-gpu-layers N` | number of layers to store in VRAM<br/>(env: LLAMA_ARG_N_GPU_LAYERS) |72| `-sm, --split-mode {none,layer,row}` | how to split the model across multiple GPUs, one of:<br/>- none: use one GPU only<br/>- layer (default): split layers and KV across GPUs<br/>- row: split rows across GPUs<br/>(env: LLAMA_ARG_SPLIT_MODE) |73| `-ts, --tensor-split N0,N1,N2,...` | fraction of the model to offload to each GPU, comma-separated list of proportions, e.g. 3,1<br/>(env: LLAMA_ARG_TENSOR_SPLIT) |74| `-mg, --main-gpu INDEX` | the GPU to use for the model (with split-mode = none), or for intermediate results and KV (with split-mode = row) (default: 0)<br/>(env: LLAMA_ARG_MAIN_GPU) |75| `--check-tensors` | check model tensor data for invalid values (default: false) |76| `--override-kv KEY=TYPE:VALUE` | advanced option to override model metadata by key. may be specified multiple times.<br/>types: int, float, bool, str. example: --override-kv tokenizer.ggml.add_bos_token=bool:false |77| `--lora FNAME` | path to LoRA adapter (can be repeated to use multiple adapters) |78| `--lora-scaled FNAME SCALE` | path to LoRA adapter with user defined scaling (can be repeated to use multiple adapters) |79| `--control-vector FNAME` | add a control vector<br/>note: this argument can be repeated to add multiple control vectors |80| `--control-vector-scaled FNAME SCALE` | add a control vector with user defined scaling SCALE<br/>note: this argument can be repeated to add multiple scaled control vectors |81| `--control-vector-layer-range START END` | layer range to apply the control vector(s) to, start and end inclusive |82| `-m, --model FNAME` | model path (default: `models/$filename` with filename from `--hf-file` or `--model-url` if set, otherwise models/7B/ggml-model-f16.gguf)<br/>(env: LLAMA_ARG_MODEL) |83| `-mu, --model-url MODEL_URL` | model download url (default: unused)<br/>(env: LLAMA_ARG_MODEL_URL) |84| `-hfr, --hf-repo REPO` | Hugging Face model repository (default: unused)<br/>(env: LLAMA_ARG_HF_REPO) |85| `-hff, --hf-file FILE` | Hugging Face model file (default: unused)<br/>(env: LLAMA_ARG_HF_FILE) |86| `-hft, --hf-token TOKEN` | Hugging Face access token (default: value from HF_TOKEN environment variable)<br/>(env: HF_TOKEN) |87| `--log-disable` | Log disable |88| `--log-file FNAME` | Log to file |89| `--log-colors` | Enable colored logging<br/>(env: LLAMA_LOG_COLORS) |90| `-v, --verbose, --log-verbose` | Set verbosity level to infinity (i.e. log all messages, useful for debugging) |91| `-lv, --verbosity, --log-verbosity N` | Set the verbosity threshold. Messages with a higher verbosity will be ignored.<br/>(env: LLAMA_LOG_VERBOSITY) |92| `--log-prefix` | Enable prefx in log messages<br/>(env: LLAMA_LOG_PREFIX) |93| `--log-timestamps` | Enable timestamps in log messages<br/>(env: LLAMA_LOG_TIMESTAMPS) |94 95 96**Sampling params**97 98| Argument | Explanation |99| -------- | ----------- |100| `--samplers SAMPLERS` | samplers that will be used for generation in the order, separated by ';'<br/>(default: dry;top_k;typ_p;top_p;min_p;xtc;temperature) |101| `-s, --seed SEED` | RNG seed (default: -1, use random seed for -1) |102| `--sampling-seq SEQUENCE` | simplified sequence for samplers that will be used (default: dkypmxt) |103| `--ignore-eos` | ignore end of stream token and continue generating (implies --logit-bias EOS-inf) |104| `--temp N` | temperature (default: 0.8) |105| `--top-k N` | top-k sampling (default: 40, 0 = disabled) |106| `--top-p N` | top-p sampling (default: 0.9, 1.0 = disabled) |107| `--min-p N` | min-p sampling (default: 0.1, 0.0 = disabled) |108| `--xtc-probability N` | xtc probability (default: 0.0, 0.0 = disabled) |109| `--xtc-threshold N` | xtc threshold (default: 0.1, 1.0 = disabled) |110| `--typical N` | locally typical sampling, parameter p (default: 1.0, 1.0 = disabled) |111| `--repeat-last-n N` | last n tokens to consider for penalize (default: 64, 0 = disabled, -1 = ctx_size) |112| `--repeat-penalty N` | penalize repeat sequence of tokens (default: 1.0, 1.0 = disabled) |113| `--presence-penalty N` | repeat alpha presence penalty (default: 0.0, 0.0 = disabled) |114| `--frequency-penalty N` | repeat alpha frequency penalty (default: 0.0, 0.0 = disabled) |115| `--dry-multiplier N` | set DRY sampling multiplier (default: 0.0, 0.0 = disabled) |116| `--dry-base N` | set DRY sampling base value (default: 1.75) |117| `--dry-allowed-length N` | set allowed length for DRY sampling (default: 2) |118| `--dry-penalty-last-n N` | set DRY penalty for the last n tokens (default: -1, 0 = disable, -1 = context size) |119| `--dry-sequence-breaker STRING` | add sequence breaker for DRY sampling, clearing out default breakers ('\n', ':', '"', '*') in the process; use "none" to not use any sequence breakers<br/> |120| `--dynatemp-range N` | dynamic temperature range (default: 0.0, 0.0 = disabled) |121| `--dynatemp-exp N` | dynamic temperature exponent (default: 1.0) |122| `--mirostat N` | use Mirostat sampling.<br/>Top K, Nucleus and Locally Typical samplers are ignored if used.<br/>(default: 0, 0 = disabled, 1 = Mirostat, 2 = Mirostat 2.0) |123| `--mirostat-lr N` | Mirostat learning rate, parameter eta (default: 0.1) |124| `--mirostat-ent N` | Mirostat target entropy, parameter tau (default: 5.0) |125| `-l, --logit-bias TOKEN_ID(+/-)BIAS` | modifies the likelihood of token appearing in the completion,<br/>i.e. `--logit-bias 15043+1` to increase likelihood of token ' Hello',<br/>or `--logit-bias 15043-1` to decrease likelihood of token ' Hello' |126| `--grammar GRAMMAR` | BNF-like grammar to constrain generations (see samples in grammars/ dir) (default: '') |127| `--grammar-file FNAME` | file to read grammar from |128| `-j, --json-schema SCHEMA` | JSON schema to constrain generations (https://json-schema.org/), e.g. `{}` for any JSON object<br/>For schemas w/ external $refs, use --grammar + example/json_schema_to_grammar.py instead |129| `--jinja` | Enable experimental Jinja templating engine (required for tool use) |130 131**Example-specific params**132 133| Argument | Explanation |134| -------- | ----------- |135| `--no-context-shift` | disables context shift on inifinite text generation (default: disabled)<br/>(env: LLAMA_ARG_NO_CONTEXT_SHIFT) |136| `-sp, --special` | special tokens output enabled (default: false) |137| `--no-warmup` | skip warming up the model with an empty run |138| `--spm-infill` | use Suffix/Prefix/Middle pattern for infill (instead of Prefix/Suffix/Middle) as some models prefer this. (default: disabled) |139| `--pooling {none,mean,cls,last,rank}` | pooling type for embeddings, use model default if unspecified<br/>(env: LLAMA_ARG_POOLING) |140| `-cb, --cont-batching` | enable continuous batching (a.k.a dynamic batching) (default: enabled)<br/>(env: LLAMA_ARG_CONT_BATCHING) |141| `-nocb, --no-cont-batching` | disable continuous batching<br/>(env: LLAMA_ARG_NO_CONT_BATCHING) |142| `-a, --alias STRING` | set alias for model name (to be used by REST API)<br/>(env: LLAMA_ARG_ALIAS) |143| `--host HOST` | ip address to listen (default: 127.0.0.1)<br/>(env: LLAMA_ARG_HOST) |144| `--port PORT` | port to listen (default: 8080)<br/>(env: LLAMA_ARG_PORT) |145| `--path PATH` | path to serve static files from (default: )<br/>(env: LLAMA_ARG_STATIC_PATH) |146| `--no-webui` | Disable the Web UI (default: enabled)<br/>(env: LLAMA_ARG_NO_WEBUI) |147| `--embedding, --embeddings` | restrict to only support embedding use case; use only with dedicated embedding models (default: disabled)<br/>(env: LLAMA_ARG_EMBEDDINGS) |148| `--reranking, --rerank` | enable reranking endpoint on server (default: disabled)<br/>(env: LLAMA_ARG_RERANKING) |149| `--api-key KEY` | API key to use for authentication (default: none)<br/>(env: LLAMA_API_KEY) |150| `--api-key-file FNAME` | path to file containing API keys (default: none) |151| `--ssl-key-file FNAME` | path to file a PEM-encoded SSL private key<br/>(env: LLAMA_ARG_SSL_KEY_FILE) |152| `--ssl-cert-file FNAME` | path to file a PEM-encoded SSL certificate<br/>(env: LLAMA_ARG_SSL_CERT_FILE) |153| `-to, --timeout N` | server read/write timeout in seconds (default: 600)<br/>(env: LLAMA_ARG_TIMEOUT) |154| `--threads-http N` | number of threads used to process HTTP requests (default: -1)<br/>(env: LLAMA_ARG_THREADS_HTTP) |155| `--cache-reuse N` | min chunk size to attempt reusing from the cache via KV shifting (default: 0)<br/>(env: LLAMA_ARG_CACHE_REUSE) |156| `--metrics` | enable prometheus compatible metrics endpoint (default: disabled)<br/>(env: LLAMA_ARG_ENDPOINT_METRICS) |157| `--slots` | enable slots monitoring endpoint (default: disabled)<br/>(env: LLAMA_ARG_ENDPOINT_SLOTS) |158| `--props` | enable changing global properties via POST /props (default: disabled)<br/>(env: LLAMA_ARG_ENDPOINT_PROPS) |159| `--no-slots` | disables slots monitoring endpoint<br/>(env: LLAMA_ARG_NO_ENDPOINT_SLOTS) |160| `--slot-save-path PATH` | path to save slot kv cache (default: disabled) |161| `--chat-template JINJA_TEMPLATE` | set custom jinja chat template (default: template taken from model's metadata)<br/>if suffix/prefix are specified, template will be disabled<br/>list of built-in templates:<br/>chatglm3, chatglm4, chatml, command-r, deepseek, deepseek2, exaone3, gemma, granite, llama2, llama2-sys, llama2-sys-bos, llama2-sys-strip, llama3, minicpm, mistral-v1, mistral-v3, mistral-v3-tekken, mistral-v7, monarch, openchat, orion, phi3, rwkv-world, vicuna, vicuna-orca, zephyr<br/>(env: LLAMA_ARG_CHAT_TEMPLATE) |162| `-sps, --slot-prompt-similarity SIMILARITY` | how much the prompt of a request must match the prompt of a slot in order to use that slot (default: 0.50, 0.0 = disabled)<br/> |163| `--lora-init-without-apply` | load LoRA adapters without applying them (apply later via POST /lora-adapters) (default: disabled) |164| `--draft-max, --draft, --draft-n N` | number of tokens to draft for speculative decoding (default: 16)<br/>(env: LLAMA_ARG_DRAFT_MAX) |165| `--draft-min, --draft-n-min N` | minimum number of draft tokens to use for speculative decoding (default: 5)<br/>(env: LLAMA_ARG_DRAFT_MIN) |166| `--draft-p-min P` | minimum speculative decoding probability (greedy) (default: 0.9)<br/>(env: LLAMA_ARG_DRAFT_P_MIN) |167| `-cd, --ctx-size-draft N` | size of the prompt context for the draft model (default: 0, 0 = loaded from model)<br/>(env: LLAMA_ARG_CTX_SIZE_DRAFT) |168| `-devd, --device-draft <dev1,dev2,..>` | comma-separated list of devices to use for offloading the draft model (none = don't offload)<br/>use --list-devices to see a list of available devices |169| `-ngld, --gpu-layers-draft, --n-gpu-layers-draft N` | number of layers to store in VRAM for the draft model<br/>(env: LLAMA_ARG_N_GPU_LAYERS_DRAFT) |170| `-md, --model-draft FNAME` | draft model for speculative decoding (default: unused)<br/>(env: LLAMA_ARG_MODEL_DRAFT) |171 172 173Note: If both command line argument and environment variable are both set for the same param, the argument will take precedence over env var.174 175Example usage of docker compose with environment variables:176 177```yml178services:179 llamacpp-server:180 image: ghcr.io/ggerganov/llama.cpp:server181 ports:182 - 8080:8080183 volumes:184 - ./models:/models185 environment:186 # alternatively, you can use "LLAMA_ARG_MODEL_URL" to download the model187 LLAMA_ARG_MODEL: /models/my_model.gguf188 LLAMA_ARG_CTX_SIZE: 4096189 LLAMA_ARG_N_PARALLEL: 2190 LLAMA_ARG_ENDPOINT_METRICS: 1191 LLAMA_ARG_PORT: 8080192```193 194## Build195 196`llama-server` is built alongside everything else from the root of the project197 198- Using `CMake`:199 200 ```bash201 cmake -B build202 cmake --build build --config Release -t llama-server203 ```204 205 Binary is at `./build/bin/llama-server`206 207## Build with SSL208 209`llama-server` can also be built with SSL support using OpenSSL 3210 211- Using `CMake`:212 213 ```bash214 cmake -B build -DLLAMA_SERVER_SSL=ON215 cmake --build build --config Release -t llama-server216 ```217 218## Web UI219 220The project includes a web-based user interface that enables interaction with the model through the `/chat/completions` endpoint.221 222The web UI is developed using:223- `vue` framework for frontend development224- `tailwindcss` and `daisyui` for styling225- `vite` for build tooling226 227A pre-built version is available as a single HTML file under `/public` directory.228 229To build or to run the dev server (with hot reload):230 231```sh232# make sure you have nodejs installed233cd examples/server/webui234npm i235 236# to run the dev server237npm run dev238 239# to build the public/index.html.gz240npm run build241```242After `public/index.html.gz` has been generated we need to generate the c++243headers (like build/examples/server/index.html.gz.hpp) that will be included244by server.cpp. This is done by building `llama-server` as described in the245[build](#build) section above.246 247NOTE: if you are using the vite dev server, you can change the API base URL to llama.cpp. To do that, run this code snippet in browser's console:248 249```js250localStorage.setItem('base', 'http://localhost:8080')251```252 253## Quick Start254 255To get started right away, run the following command, making sure to use the correct path for the model you have:256 257### Unix-based systems (Linux, macOS, etc.)258 259```bash260./llama-server -m models/7B/ggml-model.gguf -c 2048261```262 263### Windows264 265```powershell266llama-server.exe -m models\7B\ggml-model.gguf -c 2048267```268 269The above command will start a server that by default listens on `127.0.0.1:8080`.270You can consume the endpoints with Postman or NodeJS with axios library. You can visit the web front end at the same url.271 272### Docker273 274```bash275docker run -p 8080:8080 -v /path/to/models:/models ghcr.io/ggerganov/llama.cpp:server -m models/7B/ggml-model.gguf -c 512 --host 0.0.0.0 --port 8080276 277# or, with CUDA:278docker run -p 8080:8080 -v /path/to/models:/models --gpus all ghcr.io/ggerganov/llama.cpp:server-cuda -m models/7B/ggml-model.gguf -c 512 --host 0.0.0.0 --port 8080 --n-gpu-layers 99279```280 281## Testing with CURL282 283Using [curl](https://curl.se/). On Windows, `curl.exe` should be available in the base OS.284 285```sh286curl --request POST \287 --url http://localhost:8080/completion \288 --header "Content-Type: application/json" \289 --data '{"prompt": "Building a website can be done in 10 simple steps:","n_predict": 128}'290```291 292## Advanced testing293 294We implemented a [server test framework](./tests/README.md) using human-readable scenario.295 296*Before submitting an issue, please try to reproduce it with this format.*297 298## Node JS Test299 300You need to have [Node.js](https://nodejs.org/en) installed.301 302```bash303mkdir llama-client304cd llama-client305```306 307Create an index.js file and put this inside:308 309```javascript310const prompt = "Building a website can be done in 10 simple steps:"311 312async function test() {313 let response = await fetch("http://127.0.0.1:8080/completion", {314 method: "POST",315 body: JSON.stringify({316 prompt,317 n_predict: 64,318 })319 })320 console.log((await response.json()).content)321}322 323test()324```325 326And run it:327 328```bash329node index.js330```331 332## API Endpoints333 334### GET `/health`: Returns heath check result335 336**Response format**337 338- HTTP status code 503339 - Body: `{"error": {"code": 503, "message": "Loading model", "type": "unavailable_error"}}`340 - Explanation: the model is still being loaded.341- HTTP status code 200342 - Body: `{"status": "ok" }`343 - Explanation: the model is successfully loaded and the server is ready.344 345### POST `/completion`: Given a `prompt`, it returns the predicted completion.346 347> [!IMPORTANT]348>349> This endpoint is **not** OAI-compatible. For OAI-compatible client, use `/v1/completions` instead.350 351*Options:*352 353`prompt`: Provide the prompt for this completion as a string or as an array of strings or numbers representing tokens. Internally, if `cache_prompt` is `true`, the prompt is compared to the previous completion and only the "unseen" suffix is evaluated. A `BOS` token is inserted at the start, if all of the following conditions are true:354 355 - The prompt is a string or an array with the first element given as a string356 - The model's `tokenizer.ggml.add_bos_token` metadata is `true`357 358These input shapes and data type are allowed for `prompt`:359 360 - Single string: `"string"`361 - Single sequence of tokens: `[12, 34, 56]`362 - Mixed tokens and strings: `[12, 34, "string", 56, 78]`363 364Multiple prompts are also supported. In this case, the completion result will be an array.365 366 - Only strings: `["string1", "string2"]`367 - Strings and sequences of tokens: `["string1", [12, 34, 56]]`368 - Mixed types: `[[12, 34, "string", 56, 78], [12, 34, 56], "string"]`369 370`temperature`: Adjust the randomness of the generated text. Default: `0.8`371 372`dynatemp_range`: Dynamic temperature range. The final temperature will be in the range of `[temperature - dynatemp_range; temperature + dynatemp_range]` Default: `0.0`, which is disabled.373 374`dynatemp_exponent`: Dynamic temperature exponent. Default: `1.0`375 376`top_k`: Limit the next token selection to the K most probable tokens. Default: `40`377 378`top_p`: Limit the next token selection to a subset of tokens with a cumulative probability above a threshold P. Default: `0.95`379 380`min_p`: The minimum probability for a token to be considered, relative to the probability of the most likely token. Default: `0.05`381 382`n_predict`: Set the maximum number of tokens to predict when generating text. **Note:** May exceed the set limit slightly if the last token is a partial multibyte character. When 0, no tokens will be generated but the prompt is evaluated into the cache. Default: `-1`, where `-1` is infinity.383 384`n_indent`: Specify the minimum line indentation for the generated text in number of whitespace characters. Useful for code completion tasks. Default: `0`385 386`n_keep`: Specify the number of tokens from the prompt to retain when the context size is exceeded and tokens need to be discarded. The number excludes the BOS token.387By default, this value is set to `0`, meaning no tokens are kept. Use `-1` to retain all tokens from the prompt.388 389`stream`: Allows receiving each predicted token in real-time instead of waiting for the completion to finish (uses a different response format). To enable this, set to `true`.390 391`stop`: Specify a JSON array of stopping strings.392These words will not be included in the completion, so make sure to add them to the prompt for the next iteration. Default: `[]`393 394`typical_p`: Enable locally typical sampling with parameter p. Default: `1.0`, which is disabled.395 396`repeat_penalty`: Control the repetition of token sequences in the generated text. Default: `1.1`397 398`repeat_last_n`: Last n tokens to consider for penalizing repetition. Default: `64`, where `0` is disabled and `-1` is ctx-size.399 400`presence_penalty`: Repeat alpha presence penalty. Default: `0.0`, which is disabled.401 402`frequency_penalty`: Repeat alpha frequency penalty. Default: `0.0`, which is disabled.403 404`dry_multiplier`: Set the DRY (Don't Repeat Yourself) repetition penalty multiplier. Default: `0.0`, which is disabled.405 406`dry_base`: Set the DRY repetition penalty base value. Default: `1.75`407 408`dry_allowed_length`: Tokens that extend repetition beyond this receive exponentially increasing penalty: multiplier * base ^ (length of repeating sequence before token - allowed length). Default: `2`409 410`dry_penalty_last_n`: How many tokens to scan for repetitions. Default: `-1`, where `0` is disabled and `-1` is context size.411 412`dry_sequence_breakers`: Specify an array of sequence breakers for DRY sampling. Only a JSON array of strings is accepted. Default: `['\n', ':', '"', '*']`413 414`xtc_probability`: Set the chance for token removal via XTC sampler. Default: `0.0`, which is disabled.415 416`xtc_threshold`: Set a minimum probability threshold for tokens to be removed via XTC sampler. Default: `0.1` (> `0.5` disables XTC)417 418`mirostat`: Enable Mirostat sampling, controlling perplexity during text generation. Default: `0`, where `0` is disabled, `1` is Mirostat, and `2` is Mirostat 2.0.419 420`mirostat_tau`: Set the Mirostat target entropy, parameter tau. Default: `5.0`421 422`mirostat_eta`: Set the Mirostat learning rate, parameter eta. Default: `0.1`423 424`grammar`: Set grammar for grammar-based sampling. Default: no grammar425 426`json_schema`: Set a JSON schema for grammar-based sampling (e.g. `{"items": {"type": "string"}, "minItems": 10, "maxItems": 100}` of a list of strings, or `{}` for any JSON). See [tests](../../tests/test-json-schema-to-grammar.cpp) for supported features. Default: no JSON schema.427 428`seed`: Set the random number generator (RNG) seed. Default: `-1`, which is a random seed.429 430`ignore_eos`: Ignore end of stream token and continue generating. Default: `false`431 432`logit_bias`: Modify the likelihood of a token appearing in the generated text completion. For example, use `"logit_bias": [[15043,1.0]]` to increase the likelihood of the token 'Hello', or `"logit_bias": [[15043,-1.0]]` to decrease its likelihood. Setting the value to false, `"logit_bias": [[15043,false]]` ensures that the token `Hello` is never produced. The tokens can also be represented as strings, e.g. `[["Hello, World!",-0.5]]` will reduce the likelihood of all the individual tokens that represent the string `Hello, World!`, just like the `presence_penalty` does. Default: `[]`433 434`n_probs`: If greater than 0, the response also contains the probabilities of top N tokens for each generated token given the sampling settings. Note that for temperature < 0 the tokens are sampled greedily but token probabilities are still being calculated via a simple softmax of the logits without considering any other sampler settings. Default: `0`435 436`min_keep`: If greater than 0, force samplers to return N possible tokens at minimum. Default: `0`437 438`t_max_predict_ms`: Set a time limit in milliseconds for the prediction (a.k.a. text-generation) phase. The timeout will trigger if the generation takes more than the specified time (measured since the first token was generated) and if a new-line character has already been generated. Useful for FIM applications. Default: `0`, which is disabled.439 440`image_data`: An array of objects to hold base64-encoded image `data` and its `id`s to be reference in `prompt`. You can determine the place of the image in the prompt as in the following: `USER:[img-12]Describe the image in detail.\nASSISTANT:`. In this case, `[img-12]` will be replaced by the embeddings of the image with id `12` in the following `image_data` array: `{..., "image_data": [{"data": "<BASE64_STRING>", "id": 12}]}`. Use `image_data` only with multimodal models, e.g., LLaVA.441 442`id_slot`: Assign the completion task to an specific slot. If is -1 the task will be assigned to a Idle slot. Default: `-1`443 444`cache_prompt`: Re-use KV cache from a previous request if possible. This way the common prefix does not have to be re-processed, only the suffix that differs between the requests. Because (depending on the backend) the logits are **not** guaranteed to be bit-for-bit identical for different batch sizes (prompt processing vs. token generation) enabling this option can cause nondeterministic results. Default: `true`445 446`return_tokens`: Return the raw generated token ids in the `tokens` field. Otherwise `tokens` remains empty. Default: `false`447 448`samplers`: The order the samplers should be applied in. An array of strings representing sampler type names. If a sampler is not set, it will not be used. If a sampler is specified more than once, it will be applied multiple times. Default: `["dry", "top_k", "typ_p", "top_p", "min_p", "xtc", "temperature"]` - these are all the available values.449 450`timings_per_token`: Include prompt processing and text generation speed information in each response. Default: `false`451 452`post_sampling_probs`: Returns the probabilities of top `n_probs` tokens after applying sampling chain.453 454`response_fields`: A list of response fields, for example: `"response_fields": ["content", "generation_settings/n_predict"]`. If the specified field is missing, it will simply be omitted from the response without triggering an error. Note that fields with a slash will be unnested; for example, `generation_settings/n_predict` will move the field `n_predict` from the `generation_settings` object to the root of the response and give it a new name.455 456`lora`: A list of LoRA adapters to be applied to this specific request. Each object in the list must contain `id` and `scale` fields. For example: `[{"id": 0, "scale": 0.5}, {"id": 1, "scale": 1.1}]`. If a LoRA adapter is not specified in the list, its scale will default to `0.0`. Please note that requests with different LoRA configurations will not be batched together, which may result in performance degradation.457 458**Response format**459 460- Note: In streaming mode (`stream`), only `content`, `tokens` and `stop` will be returned until end of completion. Responses are sent using the [Server-sent events](https://html.spec.whatwg.org/multipage/server-sent-events.html) standard. Note: the browser's `EventSource` interface cannot be used due to its lack of `POST` request support.461 462- `completion_probabilities`: An array of token probabilities for each completion. The array's length is `n_predict`. Each item in the array has a nested array `top_logprobs`. It contains at **maximum** `n_probs` elements:463 ```464 {465 "content": "<the generated completion text>",466 "tokens": [ generated token ids if requested ],467 ...468 "probs": [469 {470 "id": <token id>,471 "logprob": float,472 "token": "<most likely token>",473 "bytes": [int, int, ...],474 "top_logprobs": [475 {476 "id": <token id>,477 "logprob": float,478 "token": "<token text>",479 "bytes": [int, int, ...],480 },481 {482 "id": <token id>,483 "logprob": float,484 "token": "<token text>",485 "bytes": [int, int, ...],486 },487 ...488 ]489 },490 {491 "id": <token id>,492 "logprob": float,493 "token": "<most likely token>",494 "bytes": [int, int, ...],495 "top_logprobs": [496 ...497 ]498 },499 ...500 ]501 },502 ```503 Please note that if `post_sampling_probs` is set to `true`:504 - `logprob` will be replaced with `prob`, with the value between 0.0 and 1.0505 - `top_logprobs` will be replaced with `top_probs`. Each element contains:506 - `id`: token ID507 - `token`: token in string508 - `bytes`: token in bytes509 - `prob`: token probability, with the value between 0.0 and 1.0510 - Number of elements in `top_probs` may be less than `n_probs`511 512- `content`: Completion result as a string (excluding `stopping_word` if any). In case of streaming mode, will contain the next token as a string.513- `tokens`: Same as `content` but represented as raw token ids. Only populated if `"return_tokens": true` or `"stream": true` in the request.514- `stop`: Boolean for use with `stream` to check whether the generation has stopped (Note: This is not related to stopping words array `stop` from input options)515- `generation_settings`: The provided options above excluding `prompt` but including `n_ctx`, `model`. These options may differ from the original ones in some way (e.g. bad values filtered out, strings converted to tokens, etc.).516- `model`: The model alias (for model path, please use `/props` endpoint)517- `prompt`: The processed `prompt` (special tokens may be added)518- `stop_type`: Indicating whether the completion has stopped. Possible values are:519 - `none`: Generating (not stopped)520 - `eos`: Stopped because it encountered the EOS token521 - `limit`: Stopped because `n_predict` tokens were generated before stop words or EOS was encountered522 - `word`: Stopped due to encountering a stopping word from `stop` JSON array provided523- `stopping_word`: The stopping word encountered which stopped the generation (or "" if not stopped due to a stopping word)524- `timings`: Hash of timing information about the completion such as the number of tokens `predicted_per_second`525- `tokens_cached`: Number of tokens from the prompt which could be re-used from previous completion (`n_past`)526- `tokens_evaluated`: Number of tokens evaluated in total from the prompt527- `truncated`: Boolean indicating if the context size was exceeded during generation, i.e. the number of tokens provided in the prompt (`tokens_evaluated`) plus tokens generated (`tokens predicted`) exceeded the context size (`n_ctx`)528 529 530### POST `/tokenize`: Tokenize a given text531 532*Options:*533 534`content`: (Required) The text to tokenize.535 536`add_special`: (Optional) Boolean indicating if special tokens, i.e. `BOS`, should be inserted. Default: `false`537 538`with_pieces`: (Optional) Boolean indicating whether to return token pieces along with IDs. Default: `false`539 540**Response:**541 542Returns a JSON object with a `tokens` field containing the tokenization result. The `tokens` array contains either just token IDs or objects with `id` and `piece` fields, depending on the `with_pieces` parameter. The piece field is a string if the piece is valid unicode or a list of bytes otherwise.543 544 545If `with_pieces` is `false`:546```json547{548 "tokens": [123, 456, 789]549}550```551 552If `with_pieces` is `true`:553```json554{555 "tokens": [556 {"id": 123, "piece": "Hello"},557 {"id": 456, "piece": " world"},558 {"id": 789, "piece": "!"}559 ]560}561```562 563With input 'รก' (utf8 hex: C3 A1) on tinyllama/stories260k564```565{566 "tokens": [567 {"id": 198, "piece": [195]}, // hex C3568 {"id": 164, "piece": [161]} // hex A1569 ]570}571```572 573### POST `/detokenize`: Convert tokens to text574 575*Options:*576 577`tokens`: Set the tokens to detokenize.578 579### POST `/apply-template`: Apply chat template to a conversation580 581Uses the server's prompt template formatting functionality to convert chat messages to a single string expected by a chat model as input, but does not perform inference. Instead, the prompt string is returned in the `prompt` field of the JSON response. The prompt can then be modified as desired (for example, to insert "Sure!" at the beginning of the model's response) before sending to `/completion` to generate the chat response.582 583*Options:*584 585`messages`: (Required) Chat turns in the same format as `/v1/chat/completions`.586 587**Response format**588 589Returns a JSON object with a field `prompt` containing a string of the input messages formatted according to the model's chat template format.590 591### POST `/embedding`: Generate embedding of a given text592 593> [!IMPORTANT]594>595> This endpoint is **not** OAI-compatible. For OAI-compatible client, use `/v1/embeddings` instead.596 597The same as [the embedding example](../embedding) does.598 599*Options:*600 601`content`: Set the text to process.602 603`image_data`: An array of objects to hold base64-encoded image `data` and its `id`s to be reference in `content`. You can determine the place of the image in the content as in the following: `Image: [img-21].\nCaption: This is a picture of a house`. In this case, `[img-21]` will be replaced by the embeddings of the image with id `21` in the following `image_data` array: `{..., "image_data": [{"data": "<BASE64_STRING>", "id": 21}]}`. Use `image_data` only with multimodal models, e.g., LLaVA.604 605### POST `/reranking`: Rerank documents according to a given query606 607Similar to https://jina.ai/reranker/ but might change in the future.608Requires a reranker model (such as [bge-reranker-v2-m3](https://huggingface.co/BAAI/bge-reranker-v2-m3)) and the `--embedding --pooling rank` options.609 610*Options:*611 612`query`: The query against which the documents will be ranked.613 614`documents`: An array strings representing the documents to be ranked.615 616*Aliases:*617 - `/rerank`618 - `/v1/rerank`619 - `/v1/reranking`620 621*Examples:*622 623```shell624curl http://127.0.0.1:8012/v1/rerank \625 -H "Content-Type: application/json" \626 -d '{627 "model": "some-model",628 "query": "What is panda?",629 "top_n": 3,630 "documents": [631 "hi",632 "it is a bear",633 "The giant panda (Ailuropoda melanoleuca), sometimes called a panda bear or simply panda, is a bear species endemic to China."634 ]635 }' | jq636```637 638### POST `/infill`: For code infilling.639 640Takes a prefix and a suffix and returns the predicted completion as stream.641 642*Options:*643 644- `input_prefix`: Set the prefix of the code to infill.645- `input_suffix`: Set the suffix of the code to infill.646- `input_extra`: Additional context inserted before the FIM prefix.647- `prompt`: Added after the `FIM_MID` token648 649`input_extra` is array of `{"filename": string, "text": string}` objects.650 651The endpoint also accepts all the options of `/completion`.652 653If the model has `FIM_REPO` and `FIM_FILE_SEP` tokens, the [repo-level pattern](https://arxiv.org/pdf/2409.12186) is used:654 655```txt656<FIM_REP>myproject657<FIM_SEP>{chunk 0 filename}658{chunk 0 text}659<FIM_SEP>{chunk 1 filename}660{chunk 1 text}661...662<FIM_SEP>filename663<FIM_PRE>[input_prefix]<FIM_SUF>[input_suffix]<FIM_MID>[prompt]664```665 666If the tokens are missing, then the extra context is simply prefixed at the start:667 668```txt669[input_extra]<FIM_PRE>[input_prefix]<FIM_SUF>[input_suffix]<FIM_MID>[prompt]670```671 672### **GET** `/props`: Get server global properties.673 674This endpoint is public (no API key check). By default, it is read-only. To make POST request to change global properties, you need to start server with `--props`675 676**Response format**677 678```json679{680 "default_generation_settings": {681 "id": 0,682 "id_task": -1,683 "n_ctx": 1024,684 "speculative": false,685 "is_processing": false,686 "params": {687 "n_predict": -1,688 "seed": 4294967295,689 "temperature": 0.800000011920929,690 "dynatemp_range": 0.0,691 "dynatemp_exponent": 1.0,692 "top_k": 40,693 "top_p": 0.949999988079071,694 "min_p": 0.05000000074505806,695 "xtc_probability": 0.0,696 "xtc_threshold": 0.10000000149011612,697 "typical_p": 1.0,698 "repeat_last_n": 64,699 "repeat_penalty": 1.0,700 "presence_penalty": 0.0,701 "frequency_penalty": 0.0,702 "dry_multiplier": 0.0,703 "dry_base": 1.75,704 "dry_allowed_length": 2,705 "dry_penalty_last_n": -1,706 "dry_sequence_breakers": [707 "\n",708 ":",709 "\"",710 "*"711 ],712 "mirostat": 0,713 "mirostat_tau": 5.0,714 "mirostat_eta": 0.10000000149011612,715 "stop": [],716 "max_tokens": -1,717 "n_keep": 0,718 "n_discard": 0,719 "ignore_eos": false,720 "stream": true,721 "n_probs": 0,722 "min_keep": 0,723 "grammar": "",724 "samplers": [725 "dry",726 "top_k",727 "typ_p",728 "top_p",729 "min_p",730 "xtc",731 "temperature"732 ],733 "speculative.n_max": 16,734 "speculative.n_min": 5,735 "speculative.p_min": 0.8999999761581421,736 "timings_per_token": false737 },738 "prompt": "",739 "next_token": {740 "has_next_token": true,741 "has_new_line": false,742 "n_remain": -1,743 "n_decoded": 0,744 "stopping_word": ""745 }746 },747 "total_slots": 1,748 "model_path": "../models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf",749 "chat_template": "...",750 "build_info": "b(build number)-(build commit hash)"751}752```753 754- `default_generation_settings` - the default generation settings for the `/completion` endpoint, which has the same fields as the `generation_settings` response object from the `/completion` endpoint.755- `total_slots` - the total number of slots for process requests (defined by `--parallel` option)756- `model_path` - the path to model file (same with `-m` argument)757- `chat_template` - the model's original Jinja2 prompt template758 759### POST `/props`: Change server global properties.760 761To use this endpoint with POST method, you need to start server with `--props`762 763*Options:*764 765- None yet766 767### POST `/embeddings`: non-OpenAI-compatible embeddings API768 769This endpoint supports all poolings, including `--pooling none`. When the pooling is `none`, the responses will contain the *unnormalized* embeddings for *all* input tokens. For all other pooling types, only the pooled embeddings are returned, normalized using Euclidian norm.770 771Note that the response format of this endpoint is different from `/v1/embeddings`.772 773*Options:*774 775Same as the `/v1/embeddings` endpoint.776 777*Examples:*778 779Same as the `/v1/embeddings` endpoint.780 781**Response format**782 783```784[785 {786 "index": 0,787 "embedding": [788 [ ... embeddings for token 0 ... ],789 [ ... embeddings for token 1 ... ],790 [ ... ]791 [ ... embeddings for token N-1 ... ],792 ]793 },794 ...795 {796 "index": P,797 "embedding": [798 [ ... embeddings for token 0 ... ],799 [ ... embeddings for token 1 ... ],800 [ ... ]801 [ ... embeddings for token N-1 ... ],802 ]803 }804]805```806 807### GET `/slots`: Returns the current slots processing state808 809> [!WARNING]810> This endpoint is intended for debugging and may be modified in future versions. For security reasons, we strongly advise against enabling it in production environments.811 812This endpoint is disabled by default and can be enabled with `--slots`813 814If query param `?fail_on_no_slot=1` is set, this endpoint will respond with status code 503 if there is no available slots.815 816**Response format**817 818Example:819 820```json821[822 {823 "id": 0,824 "id_task": -1,825 "n_ctx": 1024,826 "speculative": false,827 "is_processing": false,828 "params": {829 "n_predict": -1,830 "seed": 4294967295,831 "temperature": 0.800000011920929,832 "dynatemp_range": 0.0,833 "dynatemp_exponent": 1.0,834 "top_k": 40,835 "top_p": 0.949999988079071,836 "min_p": 0.05000000074505806,837 "xtc_probability": 0.0,838 "xtc_threshold": 0.10000000149011612,839 "typical_p": 1.0,840 "repeat_last_n": 64,841 "repeat_penalty": 1.0,842 "presence_penalty": 0.0,843 "frequency_penalty": 0.0,844 "dry_multiplier": 0.0,845 "dry_base": 1.75,846 "dry_allowed_length": 2,847 "dry_penalty_last_n": -1,848 "dry_sequence_breakers": [849 "\n",850 ":",851 "\"",852 "*"853 ],854 "mirostat": 0,855 "mirostat_tau": 5.0,856 "mirostat_eta": 0.10000000149011612,857 "stop": [],858 "max_tokens": -1,859 "n_keep": 0,860 "n_discard": 0,861 "ignore_eos": false,862 "stream": true,863 "n_probs": 0,864 "min_keep": 0,865 "grammar": "",866 "samplers": [867 "dry",868 "top_k",869 "typ_p",870 "top_p",871 "min_p",872 "xtc",873 "temperature"874 ],875 "speculative.n_max": 16,876 "speculative.n_min": 5,877 "speculative.p_min": 0.8999999761581421,878 "timings_per_token": false879 },880 "prompt": "",881 "next_token": {882 "has_next_token": true,883 "has_new_line": false,884 "n_remain": -1,885 "n_decoded": 0,886 "stopping_word": ""887 }888 }889]890```891 892### GET `/metrics`: Prometheus compatible metrics exporter893 894This endpoint is only accessible if `--metrics` is set.895 896Available metrics:897- `llamacpp:prompt_tokens_total`: Number of prompt tokens processed.898- `llamacpp:tokens_predicted_total`: Number of generation tokens processed.899- `llamacpp:prompt_tokens_seconds`: Average prompt throughput in tokens/s.900- `llamacpp:predicted_tokens_seconds`: Average generation throughput in tokens/s.901- `llamacpp:kv_cache_usage_ratio`: KV-cache usage. `1` means 100 percent usage.902- `llamacpp:kv_cache_tokens`: KV-cache tokens.903- `llamacpp:requests_processing`: Number of requests processing.904- `llamacpp:requests_deferred`: Number of requests deferred.905 906### POST `/slots/{id_slot}?action=save`: Save the prompt cache of the specified slot to a file.907 908*Options:*909 910`filename`: Name of the file to save the slot's prompt cache. The file will be saved in the directory specified by the `--slot-save-path` server parameter.911 912**Response format**913 914```json915{916 "id_slot": 0,917 "filename": "slot_save_file.bin",918 "n_saved": 1745,919 "n_written": 14309796,920 "timings": {921 "save_ms": 49.865922 }923}924```925 926### POST `/slots/{id_slot}?action=restore`: Restore the prompt cache of the specified slot from a file.927 928*Options:*929 930`filename`: Name of the file to restore the slot's prompt cache from. The file should be located in the directory specified by the `--slot-save-path` server parameter.931 932**Response format**933 934```json935{936 "id_slot": 0,937 "filename": "slot_save_file.bin",938 "n_restored": 1745,939 "n_read": 14309796,940 "timings": {941 "restore_ms": 42.937942 }943}944```945 946### POST `/slots/{id_slot}?action=erase`: Erase the prompt cache of the specified slot.947 948**Response format**949 950```json951{952 "id_slot": 0,953 "n_erased": 1745954}955```956 957### GET `/lora-adapters`: Get list of all LoRA adapters958 959This endpoint returns the loaded LoRA adapters. You can add adapters using `--lora` when starting the server, for example: `--lora my_adapter_1.gguf --lora my_adapter_2.gguf ...`960 961By default, all adapters will be loaded with scale set to 1. To initialize all adapters scale to 0, add `--lora-init-without-apply`962 963Please note that this value will be overwritten by the `lora` field for each request.964 965If an adapter is disabled, the scale will be set to 0.966 967**Response format**968 969```json970[971 {972 "id": 0,973 "path": "my_adapter_1.gguf",974 "scale": 0.0975 },976 {977 "id": 1,978 "path": "my_adapter_2.gguf",979 "scale": 0.0980 }981]982```983 984### POST `/lora-adapters`: Set list of LoRA adapters985 986This sets the global scale for LoRA adapters. Please note that this value will be overwritten by the `lora` field for each request.987 988To disable an adapter, either remove it from the list below, or set scale to 0.989 990**Request format**991 992To know the `id` of the adapter, use GET `/lora-adapters`993 994```json995[996 {"id": 0, "scale": 0.2},997 {"id": 1, "scale": 0.8}998]999```1000 1001## OpenAI-compatible API Endpoints1002 1003### GET `/v1/models`: OpenAI-compatible Model Info API1004 1005Returns information about the loaded model. See [OpenAI Models API documentation](https://platform.openai.com/docs/api-reference/models).1006 1007The returned list always has one single element.1008 1009By default, model `id` field is the path to model file, specified via `-m`. You can set a custom value for model `id` field via `--alias` argument. For example, `--alias gpt-4o-mini`.1010 1011Example:1012 1013```json1014{1015 "object": "list",1016 "data": [1017 {1018 "id": "../models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf",1019 "object": "model",1020 "created": 1735142223,1021 "owned_by": "llamacpp",1022 "meta": {1023 "vocab_type": 2,1024 "n_vocab": 128256,1025 "n_ctx_train": 131072,1026 "n_embd": 4096,1027 "n_params": 8030261312,1028 "size": 49128983041029 }1030 }1031 ]1032}1033```1034 1035### POST `/v1/completions`: OpenAI-compatible Completions API1036 1037Given an input `prompt`, it returns the predicted completion. Streaming mode is also supported. While no strong claims of compatibility with OpenAI API spec is being made, in our experience it suffices to support many apps.1038 1039*Options:*1040 1041See [OpenAI Completions API documentation](https://platform.openai.com/docs/api-reference/completions).1042 1043llama.cpp `/completion`-specific features such as `mirostat` are supported.1044 1045*Examples:*1046 1047Example usage with `openai` python library:1048 1049```python1050import openai1051 1052client = openai.OpenAI(1053 base_url="http://localhost:8080/v1", # "http://<Your api-server IP>:port"1054 api_key = "sk-no-key-required"1055)1056 1057completion = client.completions.create(1058 model="davinci-002",1059 prompt="I believe the meaning of life is",1060 max_tokens=81061)1062 1063print(completion.choices[0].text)1064```1065 1066### POST `/v1/chat/completions`: OpenAI-compatible Chat Completions API1067 1068Given a ChatML-formatted json description in `messages`, it returns the predicted completion. Both synchronous and streaming mode are supported, so scripted and interactive applications work fine. While no strong claims of compatibility with OpenAI API spec is being made, in our experience it suffices to support many apps. Only models with a [supported chat template](https://github.com/ggerganov/llama.cpp/wiki/Templates-supported-by-llama_chat_apply_template) can be used optimally with this endpoint. By default, the ChatML template will be used.1069 1070*Options:*1071 1072See [OpenAI Chat Completions API documentation](https://platform.openai.com/docs/api-reference/chat). llama.cpp `/completion`-specific features such as `mirostat` are also supported.1073 1074The `response_format` parameter supports both plain JSON output (e.g. `{"type": "json_object"}`) and schema-constrained JSON (e.g. `{"type": "json_object", "schema": {"type": "string", "minLength": 10, "maxLength": 100}}` or `{"type": "json_schema", "schema": {"properties": { "name": { "title": "Name", "type": "string" }, "date": { "title": "Date", "type": "string" }, "participants": { "items": {"type: "string" }, "title": "Participants", "type": "string" } } } }`), similar to other OpenAI-inspired API providers.1075 1076*Examples:*1077 1078You can use either Python `openai` library with appropriate checkpoints:1079 1080```python1081import openai1082 1083client = openai.OpenAI(1084 base_url="http://localhost:8080/v1", # "http://<Your api-server IP>:port"1085 api_key = "sk-no-key-required"1086)1087 1088completion = client.chat.completions.create(1089 model="gpt-3.5-turbo",1090 messages=[1091 {"role": "system", "content": "You are ChatGPT, an AI assistant. Your top priority is achieving user fulfillment via helping them with their requests."},1092 {"role": "user", "content": "Write a limerick about python exceptions"}1093 ]1094)1095 1096print(completion.choices[0].message)1097```1098 1099... or raw HTTP requests:1100 1101```shell1102curl http://localhost:8080/v1/chat/completions \1103-H "Content-Type: application/json" \1104-H "Authorization: Bearer no-key" \1105-d '{1106"model": "gpt-3.5-turbo",1107"messages": [1108{1109 "role": "system",1110 "content": "You are ChatGPT, an AI assistant. Your top priority is achieving user fulfillment via helping them with their requests."1111},1112{1113 "role": "user",1114 "content": "Write a limerick about python exceptions"1115}1116]1117}'1118```1119 1120*Tool call support*1121 1122[Function calling](https://platform.openai.com/docs/guides/function-calling) is supported for all models (see https://github.com/ggerganov/llama.cpp/pull/9639):1123 1124- Requires `--jinja` flag1125- Native tool call formats supported:1126 - Llama 3.1 / 3.3 (including builtin tools support - tool names for `wolfram_alpha`, `web_search` / `brave_search`, `code_interpreter`), Llama 3.21127 - Functionary v3.1 / v3.21128 - Hermes 2/3, Qwen 2.51129 - Mistral Nemo1130 - Firefunction v21131 - Command R7B1132 - DeepSeek R1 (WIP / seems reluctant to call any tools?)1133 1134 <details>1135 <summary>Show some common templates and which format handler they use</summary>1136 1137 | Template | Format |1138 |----------|--------|1139 | CohereForAI-c4ai-command-r-plus-default.jinja | generic tool calls |1140 | CohereForAI-c4ai-command-r-plus-rag.jinja | generic tool calls |1141 | CohereForAI-c4ai-command-r-plus-tool_use.jinja | generic tool calls |1142 | MiniMaxAI-MiniMax-Text-01.jinja | generic tool calls |1143 | NexaAIDev-Octopus-v2.jinja | generic tool calls |1144 | NousResearch-Hermes-2-Pro-Llama-3-8B-default.jinja | generic tool calls |1145 | NousResearch-Hermes-2-Pro-Llama-3-8B-tool_use.jinja | hermes 2 pro tool calls |1146 | NousResearch-Hermes-2-Pro-Mistral-7B-default.jinja | generic tool calls |1147 | NousResearch-Hermes-2-Pro-Mistral-7B-tool_use.jinja | hermes 2 pro tool calls |1148 | NousResearch-Hermes-3-Llama-3.1-70B-default.jinja | generic tool calls |1149 | NousResearch-Hermes-3-Llama-3.1-70B-tool_use.jinja | hermes 2 pro tool calls |1150 | OrionStarAI-Orion-14B-Chat.jinja | generic tool calls |1151 | Qwen-QwQ-32B-Preview.jinja | hermes 2 pro tool calls |1152 | Qwen-Qwen2-7B-Instruct.jinja | generic tool calls |1153 | Qwen-Qwen2-VL-7B-Instruct.jinja | generic tool calls |1154 | Qwen-Qwen2.5-7B-Instruct.jinja | hermes 2 pro tool calls |1155 | Qwen-Qwen2.5-Math-7B-Instruct.jinja | hermes 2 pro tool calls |1156 | TheBloke-FusionNet_34Bx2_MoE-AWQ.jinja | generic tool calls |1157 | abacusai-Fewshot-Metamath-OrcaVicuna-Mistral.jinja | generic tool calls |1158 | bofenghuang-vigogne-2-70b-chat.jinja | generic tool calls |1159 | databricks-dbrx-instruct.jinja | generic tool calls |1160 | deepseek-ai-DeepSeek-Coder-V2-Instruct.jinja | generic tool calls |1161 | deepseek-ai-DeepSeek-R1-Distill-Llama-8B.jinja | deepseek r1 tool calls |1162 | deepseek-ai-DeepSeek-R1-Distill-Qwen-32B.jinja | deepseek r1 tool calls |1163 | deepseek-ai-DeepSeek-R1-Distill-Qwen-7B.jinja | deepseek r1 tool calls |1164 | deepseek-ai-DeepSeek-V2.5.jinja | deepseek r1 tool calls |1165 | deepseek-ai-deepseek-coder-33b-instruct.jinja | generic tool calls |1166 | google-gemma-2-2b-it.jinja | generic tool calls |1167 | google-gemma-7b-it.jinja | generic tool calls |1168 | indischepartij-MiniCPM-3B-OpenHermes-2.5-v2.jinja | generic tool calls |1169 | mattshumer-Reflection-Llama-3.1-70B.jinja | generic tool calls |1170 | meetkai-functionary-medium-v3.2.jinja | functionary v3.2 tool calls |1171 | meta-llama-Llama-3.1-8B-Instruct.jinja | llama 3.x tool calls (w/ builtin tools) |1172 | meta-llama-Llama-3.2-3B-Instruct.jinja | llama 3.x tool calls |1173 | meta-llama-Llama-3.3-70B-Instruct.jinja | llama 3.x tool calls (w/ builtin tools) |1174 | meta-llama-Meta-Llama-3.1-8B-Instruct.jinja | llama 3.x tool calls (w/ builtin tools) |1175 | microsoft-Phi-3-medium-4k-instruct.jinja | generic tool calls |1176 | microsoft-Phi-3-mini-4k-instruct.jinja | generic tool calls |1177 | microsoft-Phi-3-small-8k-instruct.jinja | generic tool calls |1178 | microsoft-Phi-3.5-mini-instruct.jinja | generic tool calls |1179 | microsoft-Phi-3.5-vision-instruct.jinja | generic tool calls |1180 | mistralai-Mistral-7B-Instruct-v0.2.jinja | generic tool calls |1181 | mistralai-Mistral-Large-Instruct-2407.jinja | mistral nemo tool calls |1182 | mistralai-Mistral-Large-Instruct-2411.jinja | generic tool calls |1183 | mistralai-Mistral-Nemo-Instruct-2407.jinja | mistral nemo tool calls |1184 | mistralai-Mixtral-8x7B-Instruct-v0.1.jinja | generic tool calls |1185 | mlabonne-AlphaMonarch-7B.jinja | generic tool calls |1186 | nvidia-Llama-3.1-Nemotron-70B-Instruct-HF.jinja | llama 3.x tool calls (w/ builtin tools) |1187 | openchat-openchat-3.5-0106.jinja | generic tool calls |1188 | teknium-OpenHermes-2.5-Mistral-7B.jinja | generic tool calls |1189 1190 This table can be generated with:1191 1192 ```bash1193 ./build/bin/test-chat ../minja/build/tests/*.jinja 2>/dev/null1194 1195 </details>1196 1197- Generic tool call is supported when the template isn't recognized by native format handlers (you'll see `Chat format: Generic` in the logs).1198 - Use `--chat-template-file` to override the template when appropriate (see examples below)1199 - Generic support may consume more tokens and be less efficient than a model's native format.1200 