KBaba7/llama.cpp
0
1# llama.cpp for CANN2 3 - [Background](#background)4 - [News](#news)5 - [OS](#os)6 - [Hardware](#hardware)7 - [Model Supports](#model-supports)8 - [DataType Supports](#datatype-supports)9 - [Docker](#docker)10 - [Linux](#linux)11 - [TODO](#todo)12 13 14## Background15 16**Ascend NPU** is a range of AI processors using Neural Processing Unit. It will efficiently handle matrix-matrix multiplication, dot-product and scalars.17 18**CANN** (Compute Architecture for Neural Networks) is a heterogeneous computing architecture for AI scenarios, providing support for multiple AI frameworks on the top and serving AI processors and programming at the bottom. It plays a crucial role in bridging the gap between upper and lower layers, and is a key platform for improving the computing efficiency of Ascend AI processors. Meanwhile, it offers a highly efficient and easy-to-use programming interface for diverse application scenarios, allowing users to rapidly build AI applications and services based on the Ascend platform.19 20**Llama.cpp + CANN**21 22The llama.cpp CANN backend is designed to support Ascend NPU. It utilize the ability of AscendC and ACLNN which are intergrated to CANN Toolkit and kernels to using Ascend NPU directly.23 24## News25 26- 2024.1127 - Support F16 and F32 data type model for Ascend 310P NPU.28- 2024.829 - Support `Q4_0` and `Q8_0` data type for Ascend NPU.30- 2024.731 - Create CANN backend for Ascend NPU.32 33## OS34 35| OS | Status | Verified |36|:-------:|:-------:|:----------------------------------------------:|37| Linux | Support | Ubuntu 22.04, OpenEuler22.03 |38 39 40## Hardware41 42### Ascend NPU43 44**Verified devices**45 46| Ascend NPU | Status |47|:-----------------------------:|:-------:|48| Atlas 300T A2 | Support |49| Atlas 300I Duo | Support |50 51*Notes:*52 53- If you have trouble with Ascend NPU device, please create a issue with **[CANN]** prefix/tag.54- If you run successfully with your Ascend NPU device, please help update the upper table.55 56 57## Model Supports58 59| Model Name | FP16 | Q8_0 | Q4_0 |60|:----------------------------|:-----:|:----:|:----:|61| AquilaChat2-7B | √ | √ | √ |62| Baichuan-7b | √ | √ | √ |63| Baichuan2-7B-Chat | √ | √ | √ |64| bitnet_b1_58-large | √ | √ | √ |65| bloom-560m | √ | x | √ |66| bloomz-alpaca-560m | √ | x | √ |67| c4ai-command-r-35B-v01 | x | x | x |68| chatglm3-6B | x | x | x |69| chinese-alpaca-2-1.3b | √ | √ | √ |70| CodeShell-7B | √ | √ | √ |71| deepseek-ai_deepseek-coder-1.3B-base | x | x | x |72| deepseek-ai_DeepSeek-V2-Lite | x | x | x |73| deepseek-coder-6.7B-instruct | x | x | x |74| DeepSeek-V2-Lite-64x1.5B | x | x | x |75| falcon-7b-instruct | √ | √ | √ |76| flan-t5-large | √ | √ | √ |77| gemma-2-9b-it | √ | √ | √ |78| glm-4-9B | x | x | x |79| gpt2 | √ | √ | √ |80| Gpt2-163M | √ | √ | √ |81| granite-3B-code-instruct | √ | √ | √ |82| GritLM-7B | √ | √ | √ |83| internlm2_5-7b-chat | √ | √ | √ |84| koala-7B-HF | √ | √ | √ |85| Llama-2-7b-chat-hf | √ | √ | √ |86| Llama-3-Smaug-8B | √ | √ | √ |87| Llama2-Chinese-7b-Chat | √ | √ | √ |88| Llama3-8B | √ | √ | √ |89| Llama3-8b-chinese | √ | √ | √ |90| mamba-130m-hf | √ | √ | √ |91| Mistral-7B-Instruct-v0.2 | √ | √ | √ |92| Mixtral-8x7B-Instruct-v0.1 | x | √ | √ |93| mpt-7B | √ | √ | √ |94| OLMo-1B-hf | √ | √ | √ |95| OpenELM-3B-Instruct | √ | √ | √ |96| Orion-14b-base | √ | √ | √ |97| phi1 | x | x | x |98| phi2 | x | x | x |99| Phi-3-mini-4k-instruct | √ | √ | √ |100| plamo-13b | √ | √ | √ |101| pythia-70M | x | x | x |102| Qwen-7B | √ | √ | √ |103| Qwen2-1.5B-Instruct | √ | x | √ |104| Refact-1_6B-fim | √ | √ | √ |105| SmolLM-135M | √ | √ | √ |106| stablelm-zephyr | x | x | x |107| stablelm-2-zephyr-1_6b | x | x | x |108| starcoderbase-1b | √ | √ | √ |109| starcoder2-3b | √ | √ | √ |110| vigogne-7b-chat | √ | √ | √ |111| xverse-7b-chat | √ | √ | √ |112| Yi-6b-Chat | √ | √ | √ |113 114 115 116## DataType Supports117 118| DataType | Status |119|:----------------------:|:-------:|120| FP16 | Support |121| Q8_0 | Support |122| Q4_0 | Support |123 124## Docker125 126### Build Images127You can get a image with llama.cpp in one command.128```sh129docker build -t llama-cpp-cann -f .devops/llama-cli-cann.Dockerfile .130```131 132### Run container133 134```sh135# Find all cards.136npu-smi info137 138# Select the cards that you want to use, make sure these cards are not used by someone.139# Following using cards of device0.140docker run --name llamacpp --device /dev/davinci0 --device /dev/davinci_manager --device /dev/devmm_svm --device /dev/hisi_hdc -v /usr/local/dcmi:/usr/local/dcmi -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi -v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ -v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info -v /PATH_TO_YOUR_MODELS/:/app/models -it llama-cpp-cann -m /app/models/MODEL_PATH -ngl 32 -p "Building a website can be done in 10 simple steps:"141```142 143*Notes:*144 145- You may need to install Ascend Driver and firmware on the **host** machine *(Please refer to the [Linux configuration](#linux) for details)*.146 147## Linux148 149### I. Setup Environment150 1511. **Install Ascend Driver and firmware**152 153 ```sh154 # create driver running user.155 sudo groupadd -g HwHiAiUser156 sudo useradd -g HwHiAiUser -d /home/HwHiAiUser -m HwHiAiUser -s /bin/bash157 sudo usermod -aG HwHiAiUser $USER158 159 # download driver from https://www.hiascend.com/hardware/firmware-drivers/community according to your system160 # and install driver.161 sudo sh Ascend-hdk-910b-npu-driver_x.x.x_linux-{arch}.run --full --install-for-all162 ```163 164 Once installed, run `npu-smi info` to check whether driver is installed successfully.165 ```sh166 +-------------------------------------------------------------------------------------------+167 | npu-smi 24.1.rc2 Version: 24.1.rc2 |168 +----------------------+---------------+----------------------------------------------------+169 | NPU Name | Health | Power(W) Temp(C) Hugepages-Usage(page)|170 | Chip | Bus-Id | AICore(%) Memory-Usage(MB) HBM-Usage(MB) |171 +======================+===============+====================================================+172 | 2 xxx | OK | 64.4 51 15 / 15 |173 | 0 | 0000:01:00.0 | 0 1873 / 15077 0 / 32768 |174 +======================+===============+====================================================+175 | 5 xxx | OK | 64.0 52 15 / 15 |176 | 0 | 0000:81:00.0 | 0 1874 / 15077 0 / 32768 |177 +======================+===============+====================================================+178 | No running processes found in NPU 2 |179 +======================+===============+====================================================+180 | No running processes found in NPU 5 |181 +======================+===============+====================================================+182 ```183 1842. **Install Ascend Firmware**185 ```sh186 # download driver from https://www.hiascend.com/hardware/firmware-drivers/community according to your system187 # and install driver.188 sudo sh Ascend-hdk-910b-npu-firmware_x.x.x.x.X.run --full189 ```190 If the following messaage appers, firmware is installed successfully.191 ```sh192 Firmware package installed successfully!193 ```194 195 1963. **Install CANN toolkit and kernels**197 198 CANN toolkit and kernels can be obtained from the official [CANN Toolkit](https://www.hiascend.com/zh/developer/download/community/result?module=cann) page.199 200 Please download the corresponding version that satified your system. The minimum version required is 8.0.RC2.alpha002 and here is the install command.201 ```sh202 pip3 install attrs numpy decorator sympy cffi pyyaml pathlib2 psutil protobuf scipy requests absl-py wheel typing_extensions203 sh Ascend-cann-toolkit_8.0.RC2.alpha002_linux-aarch64.run --install204 sh Ascend-cann-kernels-910b_8.0.RC2.alpha002_linux.run --install205 ```206 207 Set Ascend Variables:208 ```sh209 echo "source ~/Ascend/ascend-toolkit/set_env.sh" >> ~/.bashrc210 source ~/.bashrc211 ```212 213Upon a successful installation, CANN is enabled for the available ascend devices.214 215### II. Build llama.cpp216 217```sh218cmake -B build -DGGML_CANN=on -DCMAKE_BUILD_TYPE=release219cmake --build build --config release220```221 222### III. Run the inference223 2241. **Retrieve and prepare model**225 226 You can refer to the general [*Prepare and Quantize*](../../README.md#prepare-and-quantize) guide for model prepration.227 228 **Notes**:229 230 - CANN backend only supports FP16/Q4_0/Q8_0 models currently.231 2322. **Launch inference**233 234 There are two device selection modes:235 236 - Single device: Use one device target specified by the user.237 - Multiple devices: Automatically choose the devices with the same backend.238 239 | Device selection | Parameter |240 |:----------------:|:--------------------------------------:|241 | Single device | --split-mode none --main-gpu DEVICE_ID |242 | Multiple devices | --split-mode layer (default) |243 244 Examples:245 246 - Use device 0:247 248 ```sh249 ./build/bin/llama-cli -m path_to_model -p "Building a website can be done in 10 simple steps:" -n 400 -e -ngl 33 -sm none -mg 0250 ```251 252 - Use multiple devices:253 254 ```sh255 ./build/bin/llama-cli -m path_to_model -p "Building a website can be done in 10 simple steps:" -n 400 -e -ngl 33 -sm layer256 ```257 258### **GitHub contribution**:259Please add the **[CANN]** prefix/tag in issues/PRs titles to help the CANN-team check/address them without delay.260 261 262## TODO263- Support more models and data types.264 