Team Ai
Apppublic

KBaba7/llama.cpp

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes
README.md523 linesDownload Raw Back to llama.cpp
1# llama.cpp2 3![llama](https://user-images.githubusercontent.com/1991296/230134379-7181e485-c521-4d23-a0d6-f7b3b61ba524.png)4 5[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](https://opensource.org/licenses/MIT)6[![Server](https://github.com/ggerganov/llama.cpp/actions/workflows/server.yml/badge.svg)](https://github.com/ggerganov/llama.cpp/actions/workflows/server.yml)7 8[Roadmap](https://github.com/users/ggerganov/projects/7) / [Project status](https://github.com/ggerganov/llama.cpp/discussions/3471) / [Manifesto](https://github.com/ggerganov/llama.cpp/discussions/205) / [ggml](https://github.com/ggerganov/ggml)9 10Inference of Meta's [LLaMA](https://arxiv.org/abs/2302.13971) model (and others) in pure C/C++11 12## Recent API changes13 14- [Changelog for `libllama` API](https://github.com/ggerganov/llama.cpp/issues/9289)15- [Changelog for `llama-server` REST API](https://github.com/ggerganov/llama.cpp/issues/9291)16 17## Hot topics18 19- **How to use [MTLResidencySet](https://developer.apple.com/documentation/metal/mtlresidencyset?language=objc) to keep the GPU memory active?** https://github.com/ggerganov/llama.cpp/pull/1142720- **VS Code extension for FIM completions:** https://github.com/ggml-org/llama.vscode21- Universal tool call support in `llama-server`: https://github.com/ggerganov/llama.cpp/pull/963922- Vim/Neovim plugin for FIM completions: https://github.com/ggml-org/llama.vim23- Introducing GGUF-my-LoRA https://github.com/ggerganov/llama.cpp/discussions/1012324- Hugging Face Inference Endpoints now support GGUF out of the box! https://github.com/ggerganov/llama.cpp/discussions/966925- Hugging Face GGUF editor: [discussion](https://github.com/ggerganov/llama.cpp/discussions/9268) | [tool](https://huggingface.co/spaces/CISCai/gguf-editor)26 27----28 29## Description30 31The main goal of `llama.cpp` is to enable LLM inference with minimal setup and state-of-the-art performance on a wide32range of hardware - locally and in the cloud.33 34- Plain C/C++ implementation without any dependencies35- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks36- AVX, AVX2, AVX512 and AMX support for x86 architectures37- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use38- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads MTT GPUs via MUSA)39- Vulkan and SYCL backend support40- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity41 42The `llama.cpp` project is the main playground for developing new features for the [ggml](https://github.com/ggerganov/ggml) library.43 44<details>45<summary>Models</summary>46 47Typically finetunes of the base models below are supported as well.48 49Instructions for adding support for new models: [HOWTO-add-model.md](docs/development/HOWTO-add-model.md)50 51#### Text-only52 53- [X] LLaMA ๐Ÿฆ™54- [x] LLaMA 2 ๐Ÿฆ™๐Ÿฆ™55- [x] LLaMA 3 ๐Ÿฆ™๐Ÿฆ™๐Ÿฆ™56- [X] [Mistral 7B](https://huggingface.co/mistralai/Mistral-7B-v0.1)57- [x] [Mixtral MoE](https://huggingface.co/models?search=mistral-ai/Mixtral)58- [x] [DBRX](https://huggingface.co/databricks/dbrx-instruct)59- [X] [Falcon](https://huggingface.co/models?search=tiiuae/falcon)60- [X] [Chinese LLaMA / Alpaca](https://github.com/ymcui/Chinese-LLaMA-Alpaca) and [Chinese LLaMA-2 / Alpaca-2](https://github.com/ymcui/Chinese-LLaMA-Alpaca-2)61- [X] [Vigogne (French)](https://github.com/bofenghuang/vigogne)62- [X] [BERT](https://github.com/ggerganov/llama.cpp/pull/5423)63- [X] [Koala](https://bair.berkeley.edu/blog/2023/04/03/koala/)64- [X] [Baichuan 1 & 2](https://huggingface.co/models?search=baichuan-inc/Baichuan) + [derivations](https://huggingface.co/hiyouga/baichuan-7b-sft)65- [X] [Aquila 1 & 2](https://huggingface.co/models?search=BAAI/Aquila)66- [X] [Starcoder models](https://github.com/ggerganov/llama.cpp/pull/3187)67- [X] [Refact](https://huggingface.co/smallcloudai/Refact-1_6B-fim)68- [X] [MPT](https://github.com/ggerganov/llama.cpp/pull/3417)69- [X] [Bloom](https://github.com/ggerganov/llama.cpp/pull/3553)70- [x] [Yi models](https://huggingface.co/models?search=01-ai/Yi)71- [X] [StableLM models](https://huggingface.co/stabilityai)72- [x] [Deepseek models](https://huggingface.co/models?search=deepseek-ai/deepseek)73- [x] [Qwen models](https://huggingface.co/models?search=Qwen/Qwen)74- [x] [PLaMo-13B](https://github.com/ggerganov/llama.cpp/pull/3557)75- [x] [Phi models](https://huggingface.co/models?search=microsoft/phi)76- [x] [PhiMoE](https://github.com/ggerganov/llama.cpp/pull/11003)77- [x] [GPT-2](https://huggingface.co/gpt2)78- [x] [Orion 14B](https://github.com/ggerganov/llama.cpp/pull/5118)79- [x] [InternLM2](https://huggingface.co/models?search=internlm2)80- [x] [CodeShell](https://github.com/WisdomShell/codeshell)81- [x] [Gemma](https://ai.google.dev/gemma)82- [x] [Mamba](https://github.com/state-spaces/mamba)83- [x] [Grok-1](https://huggingface.co/keyfan/grok-1-hf)84- [x] [Xverse](https://huggingface.co/models?search=xverse)85- [x] [Command-R models](https://huggingface.co/models?search=CohereForAI/c4ai-command-r)86- [x] [SEA-LION](https://huggingface.co/models?search=sea-lion)87- [x] [GritLM-7B](https://huggingface.co/GritLM/GritLM-7B) + [GritLM-8x7B](https://huggingface.co/GritLM/GritLM-8x7B)88- [x] [OLMo](https://allenai.org/olmo)89- [x] [OLMo 2](https://allenai.org/olmo)90- [x] [OLMoE](https://huggingface.co/allenai/OLMoE-1B-7B-0924)91- [x] [Granite models](https://huggingface.co/collections/ibm-granite/granite-code-models-6624c5cec322e4c148c8b330)92- [x] [GPT-NeoX](https://github.com/EleutherAI/gpt-neox) + [Pythia](https://github.com/EleutherAI/pythia)93- [x] [Snowflake-Arctic MoE](https://huggingface.co/collections/Snowflake/arctic-66290090abe542894a5ac520)94- [x] [Smaug](https://huggingface.co/models?search=Smaug)95- [x] [Poro 34B](https://huggingface.co/LumiOpen/Poro-34B)96- [x] [Bitnet b1.58 models](https://huggingface.co/1bitLLM)97- [x] [Flan T5](https://huggingface.co/models?search=flan-t5)98- [x] [Open Elm models](https://huggingface.co/collections/apple/openelm-instruct-models-6619ad295d7ae9f868b759ca)99- [x] [ChatGLM3-6b](https://huggingface.co/THUDM/chatglm3-6b) + [ChatGLM4-9b](https://huggingface.co/THUDM/glm-4-9b) + [GLMEdge-1.5b](https://huggingface.co/THUDM/glm-edge-1.5b-chat) + [GLMEdge-4b](https://huggingface.co/THUDM/glm-edge-4b-chat)100- [x] [SmolLM](https://huggingface.co/collections/HuggingFaceTB/smollm-6695016cad7167254ce15966)101- [x] [EXAONE-3.0-7.8B-Instruct](https://huggingface.co/LGAI-EXAONE/EXAONE-3.0-7.8B-Instruct)102- [x] [FalconMamba Models](https://huggingface.co/collections/tiiuae/falconmamba-7b-66b9a580324dd1598b0f6d4a)103- [x] [Jais](https://huggingface.co/inceptionai/jais-13b-chat)104- [x] [Bielik-11B-v2.3](https://huggingface.co/collections/speakleash/bielik-11b-v23-66ee813238d9b526a072408a)105- [x] [RWKV-6](https://github.com/BlinkDL/RWKV-LM)106- [x] [QRWKV-6](https://huggingface.co/recursal/QRWKV6-32B-Instruct-Preview-v0.1)107- [x] [GigaChat-20B-A3B](https://huggingface.co/ai-sage/GigaChat-20B-A3B-instruct)108 109#### Multimodal110 111- [x] [LLaVA 1.5 models](https://huggingface.co/collections/liuhaotian/llava-15-653aac15d994e992e2677a7e), [LLaVA 1.6 models](https://huggingface.co/collections/liuhaotian/llava-16-65b9e40155f60fd046a5ccf2)112- [x] [BakLLaVA](https://huggingface.co/models?search=SkunkworksAI/Bakllava)113- [x] [Obsidian](https://huggingface.co/NousResearch/Obsidian-3B-V0.5)114- [x] [ShareGPT4V](https://huggingface.co/models?search=Lin-Chen/ShareGPT4V)115- [x] [MobileVLM 1.7B/3B models](https://huggingface.co/models?search=mobileVLM)116- [x] [Yi-VL](https://huggingface.co/models?search=Yi-VL)117- [x] [Mini CPM](https://huggingface.co/models?search=MiniCPM)118- [x] [Moondream](https://huggingface.co/vikhyatk/moondream2)119- [x] [Bunny](https://github.com/BAAI-DCAI/Bunny)120- [x] [GLM-EDGE](https://huggingface.co/models?search=glm-edge)121- [x] [Qwen2-VL](https://huggingface.co/collections/Qwen/qwen2-vl-66cee7455501d7126940800d)122 123</details>124 125<details>126<summary>Bindings</summary>127 128- Python: [abetlen/llama-cpp-python](https://github.com/abetlen/llama-cpp-python)129- Go: [go-skynet/go-llama.cpp](https://github.com/go-skynet/go-llama.cpp)130- Node.js: [withcatai/node-llama-cpp](https://github.com/withcatai/node-llama-cpp)131- JS/TS (llama.cpp server client): [lgrammel/modelfusion](https://modelfusion.dev/integration/model-provider/llamacpp)132- JS/TS (Programmable Prompt Engine CLI): [offline-ai/cli](https://github.com/offline-ai/cli)133- JavaScript/Wasm (works in browser): [tangledgroup/llama-cpp-wasm](https://github.com/tangledgroup/llama-cpp-wasm)134- Typescript/Wasm (nicer API, available on npm): [ngxson/wllama](https://github.com/ngxson/wllama)135- Ruby: [yoshoku/llama_cpp.rb](https://github.com/yoshoku/llama_cpp.rb)136- Rust (more features): [edgenai/llama_cpp-rs](https://github.com/edgenai/llama_cpp-rs)137- Rust (nicer API): [mdrokz/rust-llama.cpp](https://github.com/mdrokz/rust-llama.cpp)138- Rust (more direct bindings): [utilityai/llama-cpp-rs](https://github.com/utilityai/llama-cpp-rs)139- Rust (automated build from crates.io): [ShelbyJenkins/llm_client](https://github.com/ShelbyJenkins/llm_client)140- C#/.NET: [SciSharp/LLamaSharp](https://github.com/SciSharp/LLamaSharp)141- C#/VB.NET (more features - community license): [LM-Kit.NET](https://docs.lm-kit.com/lm-kit-net/index.html)142- Scala 3: [donderom/llm4s](https://github.com/donderom/llm4s)143- Clojure: [phronmophobic/llama.clj](https://github.com/phronmophobic/llama.clj)144- React Native: [mybigday/llama.rn](https://github.com/mybigday/llama.rn)145- Java: [kherud/java-llama.cpp](https://github.com/kherud/java-llama.cpp)146- Zig: [deins/llama.cpp.zig](https://github.com/Deins/llama.cpp.zig)147- Flutter/Dart: [netdur/llama_cpp_dart](https://github.com/netdur/llama_cpp_dart)148- Flutter: [xuegao-tzx/Fllama](https://github.com/xuegao-tzx/Fllama)149- PHP (API bindings and features built on top of llama.cpp): [distantmagic/resonance](https://github.com/distantmagic/resonance) [(more info)](https://github.com/ggerganov/llama.cpp/pull/6326)150- Guile Scheme: [guile_llama_cpp](https://savannah.nongnu.org/projects/guile-llama-cpp)151- Swift [srgtuszy/llama-cpp-swift](https://github.com/srgtuszy/llama-cpp-swift)152- Swift [ShenghaiWang/SwiftLlama](https://github.com/ShenghaiWang/SwiftLlama)153 154</details>155 156<details>157<summary>UIs</summary>158 159*(to have a project listed here, it should clearly state that it depends on `llama.cpp`)*160 161- [AI Sublime Text plugin](https://github.com/yaroslavyaroslav/OpenAI-sublime-text) (MIT)162- [cztomsik/ava](https://github.com/cztomsik/ava) (MIT)163- [Dot](https://github.com/alexpinel/Dot) (GPL)164- [eva](https://github.com/ylsdamxssjxxdd/eva) (MIT)165- [iohub/collama](https://github.com/iohub/coLLaMA) (Apache-2.0)166- [janhq/jan](https://github.com/janhq/jan) (AGPL)167- [KanTV](https://github.com/zhouwg/kantv?tab=readme-ov-file) (Apache-2.0)168- [KodiBot](https://github.com/firatkiral/kodibot) (GPL)169- [llama.vim](https://github.com/ggml-org/llama.vim) (MIT)170- [LARS](https://github.com/abgulati/LARS) (AGPL)171- [Llama Assistant](https://github.com/vietanhdev/llama-assistant) (GPL)172- [LLMFarm](https://github.com/guinmoon/LLMFarm?tab=readme-ov-file) (MIT)173- [LLMUnity](https://github.com/undreamai/LLMUnity) (MIT)174- [LMStudio](https://lmstudio.ai/) (proprietary)175- [LocalAI](https://github.com/mudler/LocalAI) (MIT)176- [LostRuins/koboldcpp](https://github.com/LostRuins/koboldcpp) (AGPL)177- [MindMac](https://mindmac.app) (proprietary)178- [MindWorkAI/AI-Studio](https://github.com/MindWorkAI/AI-Studio) (FSL-1.1-MIT)179- [Mobile-Artificial-Intelligence/maid](https://github.com/Mobile-Artificial-Intelligence/maid) (MIT)180- [Mozilla-Ocho/llamafile](https://github.com/Mozilla-Ocho/llamafile) (Apache-2.0)181- [nat/openplayground](https://github.com/nat/openplayground) (MIT)182- [nomic-ai/gpt4all](https://github.com/nomic-ai/gpt4all) (MIT)183- [ollama/ollama](https://github.com/ollama/ollama) (MIT)184- [oobabooga/text-generation-webui](https://github.com/oobabooga/text-generation-webui) (AGPL)185- [PocketPal AI](https://github.com/a-ghorbani/pocketpal-ai) (MIT)186- [psugihara/FreeChat](https://github.com/psugihara/FreeChat) (MIT)187- [ptsochantaris/emeltal](https://github.com/ptsochantaris/emeltal) (MIT)188- [pythops/tenere](https://github.com/pythops/tenere) (AGPL)189- [ramalama](https://github.com/containers/ramalama) (MIT)190- [semperai/amica](https://github.com/semperai/amica) (MIT)191- [withcatai/catai](https://github.com/withcatai/catai) (MIT)192- [Autopen](https://github.com/blackhole89/autopen) (GPL)193 194</details>195 196<details>197<summary>Tools</summary>198 199- [akx/ggify](https://github.com/akx/ggify) โ€“ download PyTorch models from HuggingFace Hub and convert them to GGML200- [akx/ollama-dl](https://github.com/akx/ollama-dl) โ€“ download models from the Ollama library to be used directly with llama.cpp201- [crashr/gppm](https://github.com/crashr/gppm) โ€“ launch llama.cpp instances utilizing NVIDIA Tesla P40 or P100 GPUs with reduced idle power consumption202- [gpustack/gguf-parser](https://github.com/gpustack/gguf-parser-go/tree/main/cmd/gguf-parser) - review/check the GGUF file and estimate the memory usage203- [Styled Lines](https://marketplace.unity.com/packages/tools/generative-ai/styled-lines-llama-cpp-model-292902) (proprietary licensed, async wrapper of inference part for game development in Unity3d with pre-built Mobile and Web platform wrappers and a model example)204 205</details>206 207<details>208<summary>Infrastructure</summary>209 210- [Paddler](https://github.com/distantmagic/paddler) - Stateful load balancer custom-tailored for llama.cpp211- [GPUStack](https://github.com/gpustack/gpustack) - Manage GPU clusters for running LLMs212- [llama_cpp_canister](https://github.com/onicai/llama_cpp_canister) - llama.cpp as a smart contract on the Internet Computer, using WebAssembly213- [llama-swap](https://github.com/mostlygeek/llama-swap) - transparent proxy that adds automatic model switching with llama-server214- [Kalavai](https://github.com/kalavai-net/kalavai-client) - Crowdsource end to end LLM deployment at any scale215 216</details>217 218<details>219<summary>Games</summary>220 221- [Lucy's Labyrinth](https://github.com/MorganRO8/Lucys_Labyrinth) - A simple maze game where agents controlled by an AI model will try to trick you.222 223</details>224 225## Supported backends226 227| Backend | Target devices |228| --- | --- |229| [Metal](docs/build.md#metal-build) | Apple Silicon |230| [BLAS](docs/build.md#blas-build) | All |231| [BLIS](docs/backend/BLIS.md) | All |232| [SYCL](docs/backend/SYCL.md) | Intel and Nvidia GPU |233| [MUSA](docs/build.md#musa) | Moore Threads MTT GPU |234| [CUDA](docs/build.md#cuda) | Nvidia GPU |235| [HIP](docs/build.md#hip) | AMD GPU |236| [Vulkan](docs/build.md#vulkan) | GPU |237| [CANN](docs/build.md#cann) | Ascend NPU |238 239## Building the project240 241The main product of this project is the `llama` library. Its C-style interface can be found in [include/llama.h](include/llama.h).242The project also includes many example programs and tools using the `llama` library. The examples range from simple, minimal code snippets to sophisticated sub-projects such as an OpenAI-compatible HTTP server. Possible methods for obtaining the binaries:243 244- Clone this repository and build locally, see [how to build](docs/build.md)245- On MacOS or Linux, install `llama.cpp` via [brew, flox or nix](docs/install.md)246- Use a Docker image, see [documentation for Docker](docs/docker.md)247- Download pre-built binaries from [releases](https://github.com/ggerganov/llama.cpp/releases)248 249## Obtaining and quantizing models250 251The [Hugging Face](https://huggingface.co) platform hosts a [number of LLMs](https://huggingface.co/models?library=gguf&sort=trending) compatible with `llama.cpp`:252 253- [Trending](https://huggingface.co/models?library=gguf&sort=trending)254- [LLaMA](https://huggingface.co/models?sort=trending&search=llama+gguf)255 256You can either manually download the GGUF file or directly use any `llama.cpp`-compatible models from Hugging Face by using this CLI argument: `-hf <user>/<model>[:quant]`257 258After downloading a model, use the CLI tools to run it locally - see below.259 260`llama.cpp` requires the model to be stored in the [GGUF](https://github.com/ggerganov/ggml/blob/master/docs/gguf.md) file format. Models in other data formats can be converted to GGUF using the `convert_*.py` Python scripts in this repo.261 262The Hugging Face platform provides a variety of online tools for converting, quantizing and hosting models with `llama.cpp`:263 264- Use the [GGUF-my-repo space](https://huggingface.co/spaces/ggml-org/gguf-my-repo) to convert to GGUF format and quantize model weights to smaller sizes265- Use the [GGUF-my-LoRA space](https://huggingface.co/spaces/ggml-org/gguf-my-lora) to convert LoRA adapters to GGUF format (more info: https://github.com/ggerganov/llama.cpp/discussions/10123)266- Use the [GGUF-editor space](https://huggingface.co/spaces/CISCai/gguf-editor) to edit GGUF meta data in the browser (more info: https://github.com/ggerganov/llama.cpp/discussions/9268)267- Use the [Inference Endpoints](https://ui.endpoints.huggingface.co/) to directly host `llama.cpp` in the cloud (more info: https://github.com/ggerganov/llama.cpp/discussions/9669)268 269To learn more about model quantization, [read this documentation](examples/quantize/README.md)270 271## [`llama-cli`](examples/main)272 273#### A CLI tool for accessing and experimenting with most of `llama.cpp`'s functionality.274 275- <details open>276    <summary>Run in conversation mode</summary>277 278    Models with a built-in chat template will automatically activate conversation mode. If this doesn't occur, you can manually enable it by adding `-cnv` and specifying a suitable chat template with `--chat-template NAME`279 280    ```bash281    llama-cli -m model.gguf282 283    # > hi, who are you?284    # Hi there! I'm your helpful assistant! I'm an AI-powered chatbot designed to assist and provide information to users like you. I'm here to help answer your questions, provide guidance, and offer support on a wide range of topics. I'm a friendly and knowledgeable AI, and I'm always happy to help with anything you need. What's on your mind, and how can I assist you today?285    #286    # > what is 1+1?287    # Easy peasy! The answer to 1+1 is... 2!288    ```289 290    </details>291 292- <details>293    <summary>Run in conversation mode with custom chat template</summary>294 295    ```bash296    # use the "chatml" template (use -h to see the list of supported templates)297    llama-cli -m model.gguf -cnv --chat-template chatml298 299    # use a custom template300    llama-cli -m model.gguf -cnv --in-prefix 'User: ' --reverse-prompt 'User:'301    ```302 303    </details>304 305- <details>306    <summary>Run simple text completion</summary>307 308    To disable conversation mode explicitly, use `-no-cnv`309 310    ```bash311    llama-cli -m model.gguf -p "I believe the meaning of life is" -n 128 -no-cnv312 313    # I believe the meaning of life is to find your own truth and to live in accordance with it. For me, this means being true to myself and following my passions, even if they don't align with societal expectations. I think that's what I love about yoga โ€“ it's not just a physical practice, but a spiritual one too. It's about connecting with yourself, listening to your inner voice, and honoring your own unique journey.314    ```315 316    </details>317 318- <details>319    <summary>Constrain the output with a custom grammar</summary>320 321    ```bash322    llama-cli -m model.gguf -n 256 --grammar-file grammars/json.gbnf -p 'Request: schedule a call at 8pm; Command:'323 324    # {"appointmentTime": "8pm", "appointmentDetails": "schedule a a call"}325    ```326 327    The [grammars/](grammars/) folder contains a handful of sample grammars. To write your own, check out the [GBNF Guide](grammars/README.md).328 329    For authoring more complex JSON grammars, check out https://grammar.intrinsiclabs.ai/330 331    </details>332 333 334## [`llama-server`](examples/server)335 336#### A lightweight, [OpenAI API](https://github.com/openai/openai-openapi) compatible, HTTP server for serving LLMs.337 338- <details open>339    <summary>Start a local HTTP server with default configuration on port 8080</summary>340 341    ```bash342    llama-server -m model.gguf --port 8080343 344    # Basic web UI can be accessed via browser: http://localhost:8080345    # Chat completion endpoint: http://localhost:8080/v1/chat/completions346    ```347 348    </details>349 350- <details>351    <summary>Support multiple-users and parallel decoding</summary>352 353    ```bash354    # up to 4 concurrent requests, each with 4096 max context355    llama-server -m model.gguf -c 16384 -np 4356    ```357 358    </details>359 360- <details>361    <summary>Enable speculative decoding</summary>362 363    ```bash364    # the draft.gguf model should be a small variant of the target model.gguf365    llama-server -m model.gguf -md draft.gguf366    ```367 368    </details>369 370- <details>371    <summary>Serve an embedding model</summary>372 373    ```bash374    # use the /embedding endpoint375    llama-server -m model.gguf --embedding --pooling cls -ub 8192376    ```377 378    </details>379 380- <details>381    <summary>Serve a reranking model</summary>382 383    ```bash384    # use the /reranking endpoint385    llama-server -m model.gguf --reranking386    ```387 388    </details>389 390- <details>391    <summary>Constrain all outputs with a grammar</summary>392 393    ```bash394    # custom grammar395    llama-server -m model.gguf --grammar-file grammar.gbnf396 397    # JSON398    llama-server -m model.gguf --grammar-file grammars/json.gbnf399    ```400 401    </details>402 403 404## [`llama-perplexity`](examples/perplexity)405 406#### A tool for measuring the perplexity [^1][^2] (and other quality metrics) of a model over a given text.407 408- <details open>409    <summary>Measure the perplexity over a text file</summary>410 411    ```bash412    llama-perplexity -m model.gguf -f file.txt413 414    # [1]15.2701,[2]5.4007,[3]5.3073,[4]6.2965,[5]5.8940,[6]5.6096,[7]5.7942,[8]4.9297, ...415    # Final estimate: PPL = 5.4007 +/- 0.67339416    ```417 418    </details>419 420- <details>421    <summary>Measure KL divergence</summary>422 423    ```bash424    # TODO425    ```426 427    </details>428 429[^1]: [examples/perplexity/README.md](./examples/perplexity/README.md)430[^2]: [https://huggingface.co/docs/transformers/perplexity](https://huggingface.co/docs/transformers/perplexity)431 432## [`llama-bench`](examples/llama-bench)433 434#### Benchmark the performance of the inference for various parameters.435 436- <details open>437    <summary>Run default benchmark</summary>438 439    ```bash440    llama-bench -m model.gguf441 442    # Output:443    # | model               |       size |     params | backend    | threads |          test |                  t/s |444    # | ------------------- | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |445    # | qwen2 1.5B Q4_0     | 885.97 MiB |     1.54 B | Metal,BLAS |      16 |         pp512 |      5765.41 ยฑ 20.55 |446    # | qwen2 1.5B Q4_0     | 885.97 MiB |     1.54 B | Metal,BLAS |      16 |         tg128 |        197.71 ยฑ 0.81 |447    #448    # build: 3e0ba0e60 (4229)449    ```450 451    </details>452 453## [`llama-run`](examples/run)454 455#### A comprehensive example for running `llama.cpp` models. Useful for inferencing. Used with RamaLama [^3].456 457- <details>458    <summary>Run a model with a specific prompt (by default it's pulled from Ollama registry)</summary>459 460    ```bash461    llama-run granite-code462    ```463 464    </details>465 466[^3]: [RamaLama](https://github.com/containers/ramalama)467 468## [`llama-simple`](examples/simple)469 470#### A minimal example for implementing apps with `llama.cpp`. Useful for developers.471 472- <details>473    <summary>Basic text completion</summary>474 475    ```bash476    llama-simple -m model.gguf477 478    # Hello my name is Kaitlyn and I am a 16 year old girl. I am a junior in high school and I am currently taking a class called "The Art of479    ```480 481    </details>482 483 484## Contributing485 486- Contributors can open PRs487- Collaborators can push to branches in the `llama.cpp` repo and merge PRs into the `master` branch488- Collaborators will be invited based on contributions489- Any help with managing issues, PRs and projects is very appreciated!490- See [good first issues](https://github.com/ggerganov/llama.cpp/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22) for tasks suitable for first contributions491- Read the [CONTRIBUTING.md](CONTRIBUTING.md) for more information492- Make sure to read this: [Inference at the edge](https://github.com/ggerganov/llama.cpp/discussions/205)493- A bit of backstory for those who are interested: [Changelog podcast](https://changelog.com/podcast/532)494 495## Other documentation496 497- [main (cli)](examples/main/README.md)498- [server](examples/server/README.md)499- [GBNF grammars](grammars/README.md)500 501#### Development documentation502 503- [How to build](docs/build.md)504- [Running on Docker](docs/docker.md)505- [Build on Android](docs/android.md)506- [Performance troubleshooting](docs/development/token_generation_performance_tips.md)507- [GGML tips & tricks](https://github.com/ggerganov/llama.cpp/wiki/GGML-Tips-&-Tricks)508 509#### Seminal papers and background on the models510 511If your issue is with model generation quality, then please at least scan the following links and papers to understand the limitations of LLaMA models. This is especially important when choosing an appropriate model size and appreciating both the significant and subtle differences between LLaMA models and ChatGPT:512- LLaMA:513    - [Introducing LLaMA: A foundational, 65-billion-parameter large language model](https://ai.facebook.com/blog/large-language-model-llama-meta-ai/)514    - [LLaMA: Open and Efficient Foundation Language Models](https://arxiv.org/abs/2302.13971)515- GPT-3516    - [Language Models are Few-Shot Learners](https://arxiv.org/abs/2005.14165)517- GPT-3.5 / InstructGPT / ChatGPT:518    - [Aligning language models to follow instructions](https://openai.com/research/instruction-following)519    - [Training language models to follow instructions with human feedback](https://arxiv.org/abs/2203.02155)520 521#### References522 523