Team Ai
Apppublic

jaothan/podman_llamacpp_python

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes
README.md181 linesDownload Raw Back to llamacpp_python
1# Llamacpp_Python Model Server2 3The llamacpp_python model server images are based on the [llama-cpp-python](https://github.com/abetlen/llama-cpp-python) project that provides python bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp). This provides us with a python based and OpenAI API compatible model server that can run LLM's of various sizes locally across Linux, Windows or Mac.4 5This model server requires models to be converted from their original format, typically a set of `*.bin` or `*.safetensor` files into a single GGUF formatted file. Many models are available in GGUF format already on [huggingface.co](https://huggingface.co). You can also use the [model converter utility](../../convert_models/) available in this repo to convert models yourself.6 7 8## Image Options9 10We currently provide 3 options for the llamacpp_python model server:11* [Base](#base)12* [Cuda](#cuda)13* [Vulkan (experimental)](#vulkan-experimental)14 15### Base16 17The [base image](../llamacpp_python/base/Containerfile) is the standard image that works for both arm64 and amd64 environments. However, it does not includes any hardware acceleration and will run with CPU only. If you use the base image, make sure that your container runtime has sufficient resources to run the desired model(s).18 19To build the base model service image:20 21```bash22make build23```24 25To pull the base model service image:26 27```bash28podman pull quay.io/ai-lab/llamacpp_python29```30 31 32### Cuda33 34The [Cuda image](../llamacpp_python/cuda/Containerfile) include all the extra drivers necessary to run our model server with Nvidia GPUs. This will significant speed up the models response time over CPU only deployments.35 36To Build the the Cuda variant image:37 38```bash39make build-cuda40```41 42To pull the base model service image:43 44```bash45podman pull quay.io/ai-lab/llamacpp_python_cuda46```47 48**IMPORTANT!**49 50To run the Cuda image with GPU acceleration, you need to install the correct [Cuda drivers](https://docs.nvidia.com/cuda/cuda-installation-guide-linux/index.html#driver-installation) for your system along with the [Nvidia Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html#). Please use the links provided to find installation instructions for your system.51 52Once those are installed you can use the container toolkit CLI to discover your Nvidia device(s).53 54```bash55sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml56```57 58Finally, you will also need to add `--device nvidia.com/gpu=all` to your `podman run` command so your container can access the GPU.59 60 61### Vulkan (experimental)62 63The [Vulkan](https://docs.vulkan.org/guide/latest/what_is_vulkan.html) image ([amd64](../llamacpp_python/vulkan/amd64/Containerfile)/[arm64](../llamacpp_python/vulkan/arm64/Containerfile)) is experimental, but can be used for gaining partial GPU access on an M-series Mac, significantly speeding up model response time over a CPU only deployment. This image requires that your podman machine provider is "applehv" and that you use krunkit instead of vfkit. Since these tools are not currently supported by podman desktop this image will remain "experimental".64 65To build the Vulkan model service variant image:66 67| System Architecture | Command |68|---|---|69| amd64 | make build-vulkan-amd64 |70| arm64 | make build-vulkan-arm64 |71 72To pull the base model service image:73 74```bash75podman pull quay.io/ai-lab/llamacpp_python_vulkan76```77 78 79## Download Model(s)80 81There are many models to choose from these days, most of which can be found on [huggingface.co](https://huggingface.co). In order to use a model with the llamacpp_python model server, it must be in GGUF format. You can either download pre-converted GGUF models directly or convert them yourself with the [model converter utility](../../convert_models/) available in this repo.82 83A well performant Apache-2.0 licensed models that we recommend using if you are just getting started is84`granite-7b-lab`. You can use the link below to quickly download a quantized (smaller) GGUF version of this model for use with the llamacpp_python model server.85 86Download URL: [https://huggingface.co/instructlab/granite-7b-lab-GGUF/resolve/main/granite-7b-lab-Q4_K_M.gguf](https://huggingface.co/instructlab/granite-7b-lab-GGUF/resolve/main/granite-7b-lab-Q4_K_M.gguf)87 88Place all models in the [models](../../models/) directory.89 90You can use this snippet below to download the default model:91 92```bash93make download-model-granite94```95 96Or you can use the generic `download-models` target from the `/models` directory to download any model file from huggingface:97 98```bash99cd ../../models100make MODEL_NAME=<model_name> MODEL_URL=<model_url> -f  Makefile download-model101# EX: make MODEL_NAME=granite-7b-lab-Q4_K_M.gguf MODEL_URL=https://huggingface.co/instructlab/granite-7b-lab-GGUF/resolve/main/granite-7b-lab-Q4_K_M.gguf -f  Makefile download-model102```103 104 105## Deploy Model Service106 107### Single Model Service:108 109To deploy the LLM server you must specify a volume mount `-v` where your models are stored on the host machine and the `MODEL_PATH` for your model of choice. The model_server is most easily deploy from calling the make command: `make -f Makefile run`. Of course as with all our make calls you can pass any number of the following variables: `REGISTRY`, `IMAGE_NAME`, `MODEL_NAME`, `MODEL_PATH`, and `PORT`.110 111```bash112podman run --rm -d \113  -p 8001:8001 \114  -v Local/path/to/locallm/models:/locallm/models:ro \115  -e MODEL_PATH=models/granite-7b-lab-Q4_K_M.gguf \116  -e HOST=0.0.0.0 \117  -e PORT=8001 \118  -e MODEL_CHAT_FORMAT=openchat \119  llamacpp_python120```121 122or with Cuda image123 124```bash125podman run --rm -d \126  --device nvidia.com/gpu=all \127  -p 8001:8001 \128  -v Local/path/to/locallm/models:/locallm/models:ro \129  -e MODEL_PATH=models/granite-7b-lab-Q4_K_M.gguf \130  -e HOST=0.0.0.0 \131  -e PORT=8001 \132  -e MODEL_CHAT_FORMAT=openchat \133  llamacpp_python134```135 136### Multiple Model Service:137 138To enable dynamic loading and unloading of different models present on your machine, you can start the model service with a `CONFIG_PATH` instead of a `MODEL_PATH`.139 140Here is an example `models_config.json` with two model options.141 142```json143{144    "host": "0.0.0.0",145    "port": 8001,146    "models": [147        {148            "model": "models/granite-7b-lab-Q4_K_M.gguf",149            "model_alias": "granite",150            "chat_format": "openchat",151        },152        {153            "model": "models/merlinite-7b-lab-Q4_K_M.gguf",154            "model_alias": "merlinite",155            "chat_format": "openchat",156        },157 158    ]159}160```161 162Now run the container with the specified config file.163 164```bash165podman run --rm -d \166        -p 8001:8001 \167        -v Local/path/to/locallm/models:/locallm/models:ro \168        -e CONFIG_PATH=models/<config-filename> \169        llamacpp_python170```171 172### DEV environment173 174The environment is implemented with devcontainer technology.175 176Running tests177 178```bash179make -f Makefile test180```181