KBaba7/llama.cpp
0
1# llama.cpp for SYCL2 3- [Background](#background)4- [Recommended Release](#recommended-release)5- [News](#news)6- [OS](#os)7- [Hardware](#hardware)8- [Docker](#docker)9- [Linux](#linux)10- [Windows](#windows)11- [Environment Variable](#environment-variable)12- [Known Issue](#known-issues)13- [Q&A](#qa)14- [TODO](#todo)15 16## Background17 18**SYCL** is a high-level parallel programming model designed to improve developers productivity writing code across various hardware accelerators such as CPUs, GPUs, and FPGAs. It is a single-source language designed for heterogeneous computing and based on standard C++17.19 20**oneAPI** is an open ecosystem and a standard-based specification, supporting multiple architectures including but not limited to intel CPUs, GPUs and FPGAs. The key components of the oneAPI ecosystem include:21 22- **DPCPP** *(Data Parallel C++)*: The primary oneAPI SYCL implementation, which includes the icpx/icx Compilers.23- **oneAPI Libraries**: A set of highly optimized libraries targeting multiple domains *(e.g. oneMKL and oneDNN)*.24- **oneAPI LevelZero**: A high performance low level interface for fine-grained control over intel iGPUs and dGPUs.25- **Nvidia & AMD Plugins**: These are plugins extending oneAPI's DPCPP support to SYCL on Nvidia and AMD GPU targets.26 27### Llama.cpp + SYCL28 29The llama.cpp SYCL backend is designed to support **Intel GPU** firstly. Based on the cross-platform feature of SYCL, it also supports other vendor GPUs: Nvidia and AMD.30 31## Recommended Release32 33The SYCL backend would be broken by some PRs due to no online CI.34 35The following release is verified with good quality:36 37|Commit ID|Tag|Release|Verified Platform| Update date|38|-|-|-|-|-|39|3bcd40b3c593d14261fb2abfabad3c0fb5b9e318|b4040 |[llama-b4040-bin-win-sycl-x64.zip](https://github.com/ggerganov/llama.cpp/releases/download/b4040/llama-b4040-bin-win-sycl-x64.zip) |Arc770/Linux/oneAPI 2024.1<br>MTL Arc GPU/Windows 11/oneAPI 2024.1| 2024-11-19|40|fb76ec31a9914b7761c1727303ab30380fd4f05c|b3038 |[llama-b3038-bin-win-sycl-x64.zip](https://github.com/ggerganov/llama.cpp/releases/download/b3038/llama-b3038-bin-win-sycl-x64.zip) |Arc770/Linux/oneAPI 2024.1<br>MTL Arc GPU/Windows 11/oneAPI 2024.1||41 42 43## News44 45- 2024.1146 - Use syclcompat to improve the performance on some platforms. This requires to use oneAPI 2025.0 or newer.47 48- 2024.849 - Use oneDNN as the default GEMM library, improve the compatibility for new Intel GPUs.50 51- 2024.552 - Performance is increased: 34 -> 37 tokens/s of llama-2-7b.Q4_0 on Arc770.53 - Arch Linux is verified successfully.54 55- 2024.456 - Support data types: GGML_TYPE_IQ4_NL, GGML_TYPE_IQ4_XS, GGML_TYPE_IQ3_XXS, GGML_TYPE_IQ3_S, GGML_TYPE_IQ2_XXS, GGML_TYPE_IQ2_XS, GGML_TYPE_IQ2_S, GGML_TYPE_IQ1_S, GGML_TYPE_IQ1_M.57 58- 2024.359 - Release binary files of Windows.60 - A blog is published: **Run LLM on all Intel GPUs Using llama.cpp**: [intel.com](https://www.intel.com/content/www/us/en/developer/articles/technical/run-llm-on-all-gpus-using-llama-cpp-artical.html) or [medium.com](https://medium.com/@jianyu_neo/run-llm-on-all-intel-gpus-using-llama-cpp-fd2e2dcbd9bd).61 - New base line is ready: [tag b2437](https://github.com/ggerganov/llama.cpp/tree/b2437).62 - Support multiple cards: **--split-mode**: [none|layer]; not support [row], it's on developing.63 - Support to assign main GPU by **--main-gpu**, replace $GGML_SYCL_DEVICE.64 - Support detecting all GPUs with level-zero and same top **Max compute units**.65 - Support OPs66 - hardsigmoid67 - hardswish68 - pool2d69 70- 2024.171 - Create SYCL backend for Intel GPU.72 - Support Windows build73 74## OS75 76| OS | Status | Verified |77|---------|---------|------------------------------------------------|78| Linux | Support | Ubuntu 22.04, Fedora Silverblue 39, Arch Linux |79| Windows | Support | Windows 11 |80 81 82## Hardware83 84### Intel GPU85 86SYCL backend supports Intel GPU Family:87 88- Intel Data Center Max Series89- Intel Flex Series, Arc Series90- Intel Built-in Arc GPU91- Intel iGPU in Core CPU (11th Generation Core CPU and newer, refer to [oneAPI supported GPU](https://www.intel.com/content/www/us/en/developer/articles/system-requirements/intel-oneapi-base-toolkit-system-requirements.html#inpage-nav-1-1)).92 93#### Verified devices94 95| Intel GPU | Status | Verified Model |96|-------------------------------|---------|---------------------------------------|97| Intel Data Center Max Series | Support | Max 1550, 1100 |98| Intel Data Center Flex Series | Support | Flex 170 |99| Intel Arc Series | Support | Arc 770, 730M, Arc A750 |100| Intel built-in Arc GPU | Support | built-in Arc GPU in Meteor Lake |101| Intel iGPU | Support | iGPU in 13700k, i5-1250P, i7-1260P, i7-1165G7 |102 103*Notes:*104 105- **Memory**106 - The device memory is a limitation when running a large model. The loaded model size, *`llm_load_tensors: buffer_size`*, is displayed in the log when running `./bin/llama-cli`.107 108 - Please make sure the GPU shared memory from the host is large enough to account for the model's size. For e.g. the *llama-2-7b.Q4_0* requires at least 8.0GB for integrated GPU and 4.0GB for discrete GPU.109 110- **Execution Unit (EU)**111 - If the iGPU has less than 80 EUs, the inference speed will likely be too slow for practical use.112 113### Other Vendor GPU114 115**Verified devices**116 117| Nvidia GPU | Status | Verified Model |118|--------------------------|-----------|----------------|119| Ampere Series | Supported | A100, A4000 |120| Ampere Series *(Mobile)* | Supported | RTX 40 Series |121 122| AMD GPU | Status | Verified Model |123|--------------------------|--------------|----------------|124| Radeon Pro | Experimental | W6800 |125| Radeon RX | Experimental | 6700 XT |126 127Note: AMD GPU support is highly experimental and is incompatible with F16.128Additionally, it only supports GPUs with a sub_group_size (warp size) of 32.129 130## Docker131The docker build option is currently limited to *intel GPU* targets.132 133### Build image134```sh135# Using FP16136docker build -t llama-cpp-sycl --build-arg="GGML_SYCL_F16=ON" --target light -f .devops/intel.Dockerfile .137```138 139*Notes*:140 141To build in default FP32 *(Slower than FP16 alternative)*, you can remove the `--build-arg="GGML_SYCL_F16=ON"` argument from the previous command.142 143You can also use the `.devops/llama-server-intel.Dockerfile`, which builds the *"server"* alternative.144 145### Run container146 147```sh148# First, find all the DRI cards149ls -la /dev/dri150# Then, pick the card that you want to use (here for e.g. /dev/dri/card1).151docker run -it --rm -v "$(pwd):/app:Z" --device /dev/dri/renderD128:/dev/dri/renderD128 --device /dev/dri/card1:/dev/dri/card1 llama-cpp-sycl -m "/app/models/YOUR_MODEL_FILE" -p "Building a website can be done in 10 simple steps:" -n 400 -e -ngl 33152```153 154*Notes:*155- Docker has been tested successfully on native Linux. WSL support has not been verified yet.156- You may need to install Intel GPU driver on the **host** machine *(Please refer to the [Linux configuration](#linux) for details)*.157 158## Linux159 160### I. Setup Environment161 1621. **Install GPU drivers**163 164 - **Intel GPU**165 166Intel data center GPUs drivers installation guide and download page can be found here: [Get intel dGPU Drivers](https://dgpu-docs.intel.com/driver/installation.html#ubuntu-install-steps).167 168*Note*: for client GPUs *(iGPU & Arc A-Series)*, please refer to the [client iGPU driver installation](https://dgpu-docs.intel.com/driver/client/overview.html).169 170Once installed, add the user(s) to the `video` and `render` groups.171 172```sh173sudo usermod -aG render $USER174sudo usermod -aG video $USER175```176 177*Note*: logout/re-login for the changes to take effect.178 179Verify installation through `clinfo`:180 181```sh182sudo apt install clinfo183sudo clinfo -l184```185 186Sample output:187 188```sh189Platform #0: Intel(R) OpenCL Graphics190 `-- Device #0: Intel(R) Arc(TM) A770 Graphics191 192Platform #0: Intel(R) OpenCL HD Graphics193 `-- Device #0: Intel(R) Iris(R) Xe Graphics [0x9a49]194```195 196- **Nvidia GPU**197 198In order to target Nvidia GPUs through SYCL, please make sure the CUDA/CUBLAS native requirements *-found [here](README.md#cuda)-* are installed.199 200- **AMD GPU**201 202To target AMD GPUs with SYCL, the ROCm stack must be installed first.203 2042. **Install Intel® oneAPI Base toolkit**205 206- **For Intel GPU**207 208The base toolkit can be obtained from the official [Intel® oneAPI Base Toolkit](https://www.intel.com/content/www/us/en/developer/tools/oneapi/base-toolkit.html) page.209 210Please follow the instructions for downloading and installing the Toolkit for Linux, and preferably keep the default installation values unchanged, notably the installation path *(`/opt/intel/oneapi` by default)*.211 212Following guidelines/code snippets assume the default installation values. Otherwise, please make sure the necessary changes are reflected where applicable.213 214Upon a successful installation, SYCL is enabled for the available intel devices, along with relevant libraries such as oneAPI oneDNN for Intel GPUs.215 216- **Adding support to Nvidia GPUs**217 218**oneAPI Plugin**: In order to enable SYCL support on Nvidia GPUs, please install the [Codeplay oneAPI Plugin for Nvidia GPUs](https://developer.codeplay.com/products/oneapi/nvidia/download). User should also make sure the plugin version matches the installed base toolkit one *(previous step)* for a seamless "oneAPI on Nvidia GPU" setup.219 220 221**oneMKL for cuBlas**: The current oneMKL releases *(shipped with the oneAPI base-toolkit)* do not contain the cuBLAS backend. A build from source of the upstream [oneMKL](https://github.com/oneapi-src/oneMKL) with the *cuBLAS* backend enabled is thus required to run it on Nvidia GPUs.222 223```sh224git clone https://github.com/oneapi-src/oneMKL225cd oneMKL226cmake -B buildWithCublas -DCMAKE_CXX_COMPILER=icpx -DCMAKE_C_COMPILER=icx -DENABLE_MKLGPU_BACKEND=OFF -DENABLE_MKLCPU_BACKEND=OFF -DENABLE_CUBLAS_BACKEND=ON -DTARGET_DOMAINS=blas227cmake --build buildWithCublas --config Release228```229 230- **Adding support to AMD GPUs**231 232**oneAPI Plugin**: In order to enable SYCL support on AMD GPUs, please install the [Codeplay oneAPI Plugin for AMD GPUs](https://developer.codeplay.com/products/oneapi/amd/download). As with Nvidia GPUs, the user should also make sure the plugin version matches the installed base toolkit.233 234**oneMKL for rocBlas**: The current oneMKL releases *(shipped with the oneAPI base-toolkit)* doesn't contain the rocBLAS backend. A build from source of the upstream [oneMKL](https://github.com/oneapi-src/oneMKL) with the *rocBLAS* backend enabled is thus required to run it on AMD GPUs.235 236```sh237git clone https://github.com/oneapi-src/oneMKL238cd oneMKL239# Find your HIPTARGET with rocminfo, under the key 'Name:'240cmake -B buildWithrocBLAS -DCMAKE_CXX_COMPILER=icpx -DCMAKE_C_COMPILER=icx -DENABLE_MKLGPU_BACKEND=OFF -DENABLE_MKLCPU_BACKEND=OFF -DENABLE_ROCBLAS_BACKEND=ON -DHIPTARGETS=${HIPTARGET} -DTARGET_DOMAINS=blas241cmake --build buildWithrocBLAS --config Release242```243 2443. **Verify installation and environment**245 246In order to check the available SYCL devices on the machine, please use the `sycl-ls` command.247```sh248source /opt/intel/oneapi/setvars.sh249sycl-ls250```251 252- **Intel GPU**253 254When targeting an intel GPU, the user should expect one or more level-zero devices among the available SYCL devices. Please make sure that at least one GPU is present, for instance [`level_zero:gpu`] in the sample output below:255 256```257[opencl:acc][opencl:0] Intel(R) FPGA Emulation Platform for OpenCL(TM), Intel(R) FPGA Emulation Device OpenCL 1.2 [2023.16.10.0.17_160000]258[opencl:cpu][opencl:1] Intel(R) OpenCL, 13th Gen Intel(R) Core(TM) i7-13700K OpenCL 3.0 (Build 0) [2023.16.10.0.17_160000]259[opencl:gpu][opencl:2] Intel(R) OpenCL Graphics, Intel(R) Arc(TM) A770 Graphics OpenCL 3.0 NEO [23.30.26918.50]260[level_zero:gpu][level_zero:0] Intel(R) Level-Zero, Intel(R) Arc(TM) A770 Graphics 1.3 [1.3.26918]261```262 263- **Nvidia GPU**264 265Similarly, user targeting Nvidia GPUs should expect at least one SYCL-CUDA device [`cuda:gpu`] as below:266 267```268[opencl:acc][opencl:0] Intel(R) FPGA Emulation Platform for OpenCL(TM), Intel(R) FPGA Emulation Device OpenCL 1.2 [2023.16.12.0.12_195853.xmain-hotfix]269[opencl:cpu][opencl:1] Intel(R) OpenCL, Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz OpenCL 3.0 (Build 0) [2023.16.12.0.12_195853.xmain-hotfix]270[cuda:gpu][cuda:0] NVIDIA CUDA BACKEND, NVIDIA A100-PCIE-40GB 8.0 [CUDA 12.5]271```272 273- **AMD GPU**274 275For AMD GPUs we should expect at least one SYCL-HIP device [`hip:gpu`]:276 277```278[opencl:cpu][opencl:0] Intel(R) OpenCL, 12th Gen Intel(R) Core(TM) i9-12900K OpenCL 3.0 (Build 0) [2024.18.6.0.02_160000]279[hip:gpu][hip:0] AMD HIP BACKEND, AMD Radeon PRO W6800 gfx1030 [HIP 60140.9]280```281 282### II. Build llama.cpp283 284#### Intel GPU285 286```287./examples/sycl/build.sh288```289 290or291 292```sh293# Export relevant ENV variables294source /opt/intel/oneapi/setvars.sh295 296# Option 1: Use FP32 (recommended for better performance in most cases)297cmake -B build -DGGML_SYCL=ON -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx298 299# Option 2: Use FP16300cmake -B build -DGGML_SYCL=ON -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx -DGGML_SYCL_F16=ON301 302# build all binary303cmake --build build --config Release -j -v304```305 306#### Nvidia GPU307 308```sh309# Export relevant ENV variables310export LD_LIBRARY_PATH=/path/to/oneMKL/buildWithCublas/lib:$LD_LIBRARY_PATH311export LIBRARY_PATH=/path/to/oneMKL/buildWithCublas/lib:$LIBRARY_PATH312export CPLUS_INCLUDE_DIR=/path/to/oneMKL/buildWithCublas/include:$CPLUS_INCLUDE_DIR313export CPLUS_INCLUDE_DIR=/path/to/oneMKL/include:$CPLUS_INCLUDE_DIR314 315# Build LLAMA with Nvidia BLAS acceleration through SYCL316# Setting GGML_SYCL_DEVICE_ARCH is optional but can improve performance317GGML_SYCL_DEVICE_ARCH=sm_80 # Example architecture318 319# Option 1: Use FP32 (recommended for better performance in most cases)320cmake -B build -DGGML_SYCL=ON -DGGML_SYCL_TARGET=NVIDIA -DGGML_SYCL_DEVICE_ARCH=${GGML_SYCL_DEVICE_ARCH} -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx321 322# Option 2: Use FP16323cmake -B build -DGGML_SYCL=ON -DGGML_SYCL_TARGET=NVIDIA -DGGML_SYCL_DEVICE_ARCH=${GGML_SYCL_DEVICE_ARCH} -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx -DGGML_SYCL_F16=ON324 325# build all binary326cmake --build build --config Release -j -v327```328 329#### AMD GPU330 331```sh332# Export relevant ENV variables333export LD_LIBRARY_PATH=/path/to/oneMKL/buildWithrocBLAS/lib:$LD_LIBRARY_PATH334export LIBRARY_PATH=/path/to/oneMKL/buildWithrocBLAS/lib:$LIBRARY_PATH335export CPLUS_INCLUDE_DIR=/path/to/oneMKL/buildWithrocBLAS/include:$CPLUS_INCLUDE_DIR336 337# Build LLAMA with rocBLAS acceleration through SYCL338 339## AMD340# Use FP32, FP16 is not supported341# Find your GGML_SYCL_DEVICE_ARCH with rocminfo, under the key 'Name:'342GGML_SYCL_DEVICE_ARCH=gfx90a # Example architecture343cmake -B build -DGGML_SYCL=ON -DGGML_SYCL_TARGET=AMD -DGGML_SYCL_DEVICE_ARCH=${GGML_SYCL_DEVICE_ARCH} -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx344 345# build all binary346cmake --build build --config Release -j -v347```348 349### III. Run the inference350 351#### Retrieve and prepare model352 353You can refer to the general [*Prepare and Quantize*](README.md#prepare-and-quantize) guide for model prepration, or simply download [llama-2-7b.Q4_0.gguf](https://huggingface.co/TheBloke/Llama-2-7B-GGUF/blob/main/llama-2-7b.Q4_0.gguf) model as example.354 355##### Check device356 3571. Enable oneAPI running environment358 359```sh360source /opt/intel/oneapi/setvars.sh361```362 3632. List devices information364 365Similar to the native `sycl-ls`, available SYCL devices can be queried as follow:366 367```sh368./build/bin/llama-ls-sycl-device369```370 371This command will only display the selected backend that is supported by SYCL. The default backend is level_zero. For example, in a system with 2 *intel GPU* it would look like the following:372```373found 2 SYCL devices:374 375| | | |Compute |Max compute|Max work|Max sub| |376|ID| Device Type| Name|capability|units |group |group |Global mem size|377|--|------------------|---------------------------------------------|----------|-----------|--------|-------|---------------|378| 0|[level_zero:gpu:0]| Intel(R) Arc(TM) A770 Graphics| 1.3| 512| 1024| 32| 16225243136|379| 1|[level_zero:gpu:1]| Intel(R) UHD Graphics 770| 1.3| 32| 512| 32| 53651849216|380```381 382#### Choose level-zero devices383 384|Chosen Device ID|Setting|385|-|-|386|0|`export ONEAPI_DEVICE_SELECTOR="level_zero:0"` or no action|387|1|`export ONEAPI_DEVICE_SELECTOR="level_zero:1"`|388|0 & 1|`export ONEAPI_DEVICE_SELECTOR="level_zero:0;level_zero:1"`|389 390#### Execute391 392Choose one of following methods to run.393 3941. Script395 396- Use device 0:397 398```sh399./examples/sycl/run-llama2.sh 0400```401- Use multiple devices:402 403```sh404./examples/sycl/run-llama2.sh405```406 4072. Command line408Launch inference409 410There are two device selection modes:411 412- Single device: Use one device assigned by user. Default device id is 0.413- Multiple devices: Automatically choose the devices with the same backend.414 415In two device selection modes, the default SYCL backend is level_zero, you can choose other backend supported by SYCL by setting environment variable ONEAPI_DEVICE_SELECTOR.416 417| Device selection | Parameter |418|------------------|----------------------------------------|419| Single device | --split-mode none --main-gpu DEVICE_ID |420| Multiple devices | --split-mode layer (default) |421 422Examples:423 424- Use device 0:425 426```sh427ZES_ENABLE_SYSMAN=1 ./build/bin/llama-cli -m models/llama-2-7b.Q4_0.gguf -p "Building a website can be done in 10 simple steps:" -n 400 -e -ngl 33 -sm none -mg 0428```429 430- Use multiple devices:431 432```sh433ZES_ENABLE_SYSMAN=1 ./build/bin/llama-cli -m models/llama-2-7b.Q4_0.gguf -p "Building a website can be done in 10 simple steps:" -n 400 -e -ngl 33 -sm layer434```435 436*Notes:*437 438- Upon execution, verify the selected device(s) ID(s) in the output log, which can for instance be displayed as follow:439 440```sh441detect 1 SYCL GPUs: [0] with top Max compute units:512442```443Or444```sh445use 1 SYCL GPUs: [0] with Max compute units:512446```447 448## Windows449 450### I. Setup Environment451 4521. Install GPU driver453 454Intel GPU drivers instructions guide and download page can be found here: [Get intel GPU Drivers](https://www.intel.com/content/www/us/en/products/docs/discrete-gpus/arc/software/drivers.html).455 4562. Install Visual Studio457 458If you already have a recent version of Microsoft Visual Studio, you can skip this step. Otherwise, please refer to the official download page for [Microsoft Visual Studio](https://visualstudio.microsoft.com/).459 4603. Install Intel® oneAPI Base toolkit461 462The base toolkit can be obtained from the official [Intel® oneAPI Base Toolkit](https://www.intel.com/content/www/us/en/developer/tools/oneapi/base-toolkit.html) page.463 464Please follow the instructions for downloading and installing the Toolkit for Windows, and preferably keep the default installation values unchanged, notably the installation path *(`C:\Program Files (x86)\Intel\oneAPI` by default)*.465 466Following guidelines/code snippets assume the default installation values. Otherwise, please make sure the necessary changes are reflected where applicable.467 468b. Enable oneAPI running environment:469 470- Type "oneAPI" in the search bar, then open the `Intel oneAPI command prompt for Intel 64 for Visual Studio 2022` App.471 472- On the command prompt, enable the runtime environment with the following:473```474"C:\Program Files (x86)\Intel\oneAPI\setvars.bat" intel64475```476 477c. Verify installation478 479In the oneAPI command line, run the following to print the available SYCL devices:480 481```482sycl-ls.exe483```484 485There should be one or more *level-zero* GPU devices displayed as **[ext_oneapi_level_zero:gpu]**. Below is example of such output detecting an *intel Iris Xe* GPU as a Level-zero SYCL device:486 487Output (example):488```489[opencl:acc:0] Intel(R) FPGA Emulation Platform for OpenCL(TM), Intel(R) FPGA Emulation Device OpenCL 1.2 [2023.16.10.0.17_160000]490[opencl:cpu:1] Intel(R) OpenCL, 11th Gen Intel(R) Core(TM) i7-1185G7 @ 3.00GHz OpenCL 3.0 (Build 0) [2023.16.10.0.17_160000]491[opencl:gpu:2] Intel(R) OpenCL Graphics, Intel(R) Iris(R) Xe Graphics OpenCL 3.0 NEO [31.0.101.5186]492[ext_oneapi_level_zero:gpu:0] Intel(R) Level-Zero, Intel(R) Iris(R) Xe Graphics 1.3 [1.3.28044]493```494 4954. Install build tools496 497a. Download & install cmake for Windows: https://cmake.org/download/ (CMake can also be installed from Visual Studio Installer)498b. The new Visual Studio will install Ninja as default. (If not, please install it manually: https://ninja-build.org/)499 500 501### II. Build llama.cpp502 503You could download the release package for Windows directly, which including binary files and depended oneAPI dll files.504 505Choose one of following methods to build from source code.506 5071. Script508 509```sh510.\examples\sycl\win-build-sycl.bat511```512 5132. CMake514 515On the oneAPI command line window, step into the llama.cpp main directory and run the following:516 517```518@call "C:\Program Files (x86)\Intel\oneAPI\setvars.bat" intel64 --force519 520# Option 1: Use FP32 (recommended for better performance in most cases)521cmake -B build -G "Ninja" -DGGML_SYCL=ON -DCMAKE_C_COMPILER=cl -DCMAKE_CXX_COMPILER=icx -DCMAKE_BUILD_TYPE=Release522 523# Option 2: Or FP16524cmake -B build -G "Ninja" -DGGML_SYCL=ON -DCMAKE_C_COMPILER=cl -DCMAKE_CXX_COMPILER=icx -DCMAKE_BUILD_TYPE=Release -DGGML_SYCL_F16=ON525 526cmake --build build --config Release -j527```528 529Or, use CMake presets to build:530 531```sh532cmake --preset x64-windows-sycl-release533cmake --build build-x64-windows-sycl-release -j --target llama-cli534 535cmake -DGGML_SYCL_F16=ON --preset x64-windows-sycl-release536cmake --build build-x64-windows-sycl-release -j --target llama-cli537 538cmake --preset x64-windows-sycl-debug539cmake --build build-x64-windows-sycl-debug -j --target llama-cli540```541 5423. Visual Studio543 544You can use Visual Studio to open llama.cpp folder as a CMake project. Choose the sycl CMake presets (`x64-windows-sycl-release` or `x64-windows-sycl-debug`) before you compile the project.545 546*Notes:*547 548- In case of a minimal experimental setup, the user can build the inference executable only through `cmake --build build --config Release -j --target llama-cli`.549 550### III. Run the inference551 552#### Retrieve and prepare model553 554You can refer to the general [*Prepare and Quantize*](README.md#prepare-and-quantize) guide for model prepration, or simply download [llama-2-7b.Q4_0.gguf](https://huggingface.co/TheBloke/Llama-2-7B-GGUF/blob/main/llama-2-7b.Q4_0.gguf) model as example.555 556##### Check device557 5581. Enable oneAPI running environment559 560On the oneAPI command line window, run the following and step into the llama.cpp directory:561```562"C:\Program Files (x86)\Intel\oneAPI\setvars.bat" intel64563```564 5652. List devices information566 567Similar to the native `sycl-ls`, available SYCL devices can be queried as follow:568 569```570build\bin\llama-ls-sycl-device.exe571```572 573This command will only display the selected backend that is supported by SYCL. The default backend is level_zero. For example, in a system with 2 *intel GPU* it would look like the following:574```575found 2 SYCL devices:576| | | |Compute |Max compute|Max work|Max sub| |577|ID| Device Type| Name|capability|units |group |group |Global mem size|578|--|------------------|---------------------------------------------|----------|-----------|--------|-------|---------------|579| 0|[level_zero:gpu:0]| Intel(R) Arc(TM) A770 Graphics| 1.3| 512| 1024| 32| 16225243136|580| 1|[level_zero:gpu:1]| Intel(R) UHD Graphics 770| 1.3| 32| 512| 32| 53651849216|581 582```583#### Choose level-zero devices584 585|Chosen Device ID|Setting|586|-|-|587|0|`set ONEAPI_DEVICE_SELECTOR="level_zero:1"` or no action|588|1|`set ONEAPI_DEVICE_SELECTOR="level_zero:1"`|589|0 & 1|`set ONEAPI_DEVICE_SELECTOR="level_zero:0;level_zero:1"`|590 591#### Execute592 593Choose one of following methods to run.594 5951. Script596 597```598examples\sycl\win-run-llama2.bat599```600 6012. Command line602 603Launch inference604 605There are two device selection modes:606 607- Single device: Use one device assigned by user. Default device id is 0.608- Multiple devices: Automatically choose the devices with the same backend.609 610In two device selection modes, the default SYCL backend is level_zero, you can choose other backend supported by SYCL by setting environment variable ONEAPI_DEVICE_SELECTOR.611 612| Device selection | Parameter |613|------------------|----------------------------------------|614| Single device | --split-mode none --main-gpu DEVICE_ID |615| Multiple devices | --split-mode layer (default) |616 617Examples:618 619- Use device 0:620 621```622build\bin\llama-cli.exe -m models\llama-2-7b.Q4_0.gguf -p "Building a website can be done in 10 simple steps:\nStep 1:" -n 400 -e -ngl 33 -s 0 -sm none -mg 0623```624 625- Use multiple devices:626 627```628build\bin\llama-cli.exe -m models\llama-2-7b.Q4_0.gguf -p "Building a website can be done in 10 simple steps:\nStep 1:" -n 400 -e -ngl 33 -s 0 -sm layer629```630 631 632Note:633 634- Upon execution, verify the selected device(s) ID(s) in the output log, which can for instance be displayed as follow:635 636```sh637detect 1 SYCL GPUs: [0] with top Max compute units:512638```639Or640```sh641use 1 SYCL GPUs: [0] with Max compute units:512642```643 644 645## Environment Variable646 647#### Build648 649| Name | Value | Function |650|--------------------|---------------------------------------|---------------------------------------------|651| GGML_SYCL | ON (mandatory) | Enable build with SYCL code path.<br>FP32 path - recommended for better perforemance than FP16 on quantized model|652| GGML_SYCL_TARGET | INTEL *(default)* \| NVIDIA \| AMD | Set the SYCL target device type. |653| GGML_SYCL_DEVICE_ARCH | Optional (except for AMD) | Set the SYCL device architecture, optional except for AMD. Setting the device architecture can improve the performance. See the table [--offload-arch](https://github.com/intel/llvm/blob/sycl/sycl/doc/design/OffloadDesign.md#--offload-arch) for a list of valid architectures. |654| GGML_SYCL_F16 | OFF *(default)* \|ON *(optional)* | Enable FP16 build with SYCL code path. |655| CMAKE_C_COMPILER | `icx` *(Linux)*, `icx/cl` *(Windows)* | Set `icx` compiler for SYCL code path. |656| CMAKE_CXX_COMPILER | `icpx` *(Linux)*, `icx` *(Windows)* | Set `icpx/icx` compiler for SYCL code path. |657 658#### Runtime659 660| Name | Value | Function |661|-------------------|------------------|---------------------------------------------------------------------------------------------------------------------------|662| GGML_SYCL_DEBUG | 0 (default) or 1 | Enable log function by macro: GGML_SYCL_DEBUG |663| ZES_ENABLE_SYSMAN | 0 (default) or 1 | Support to get free memory of GPU by sycl::aspect::ext_intel_free_memory.<br>Recommended to use when --split-mode = layer |664 665## Known Issues666 667- `Split-mode:[row]` is not supported.668 669## Q&A670 671- Error: `error while loading shared libraries: libsycl.so.7: cannot open shared object file: No such file or directory`.672 673 - Potential cause: Unavailable oneAPI installation or not set ENV variables.674 - Solution: Install *oneAPI base toolkit* and enable its ENV through: `source /opt/intel/oneapi/setvars.sh`.675 676- General compiler error:677 678 - Remove **build** folder or try a clean-build.679 680- I can **not** see `[ext_oneapi_level_zero:gpu]` afer installing the GPU driver on Linux.681 682 Please double-check with `sudo sycl-ls`.683 684 If it's present in the list, please add video/render group to your user then **logout/login** or restart your system:685 686 ```687 sudo usermod -aG render $USER688 sudo usermod -aG video $USER689 ```690 Otherwise, please double-check the GPU driver installation steps.691 692- Can I report Ollama issue on Intel GPU to llama.cpp SYCL backend?693 694 No. We can't support Ollama issue directly, because we aren't familiar with Ollama.695 696 Sugguest reproducing on llama.cpp and report similar issue to llama.cpp. We will surpport it.697 698 It's same for other projects including llama.cpp SYCL backend.699 700- Meet issue: `Native API failed. Native API returns: -6 (PI_ERROR_OUT_OF_HOST_MEMORY) -6 (PI_ERROR_OUT_OF_HOST_MEMORY) -999 (UNKNOWN PI error)` or `failed to allocate SYCL0 buffer`701 702 Device Memory is not enough.703 704 |Reason|Solution|705 |-|-|706 |Default Context is too big. It leads to more memory usage.|Set `-c 8192` or smaller value.|707 |Model is big and require more memory than device's.|Choose smaller quantized model, like Q5 -> Q4;<br>Use more than one devices to load model.|708 709### **GitHub contribution**:710Please add the **[SYCL]** prefix/tag in issues/PRs titles to help the SYCL-team check/address them without delay.711 712## TODO713 714- NA715 