Team Ai
Datasetpublic

AIencoder/llama-cpp-wheels

If you like this please consider liking and donating (https://buymeacoffee.com/aiencoder) 🏭 llama-cpp-python Mega-Factory Wheels "Stop waiting for pip to compile. Just install and run." The most complete collection of pre-built llama-cpp-python wheels in existence β€” 8,333 wheels across every platform, Python version, backend, and CPU optimization level. No more cmake, gcc, or compilation hell. No more waiting 10 minutes for a build that might fail. Just find your wheel and… See the full description on the dataset page: https://huggingface.co/datasets/AIencoder/llama-cpp-wheels.

sourceHugging Facemitupdated 16h agoView on Hugging Face
4likes54kdownloads
README.md190 linesDownload Raw Back to root
1---2license: mit3task_categories:4- text-generation5language:6- en7tags:8- code9- llama-cpp10- llama-cpp-python11- wheels12- pre-built13- binary14- linux15- windows16- macos17pretty_name: llama-cpp-python Pre-Built Wheels18size_categories:19- 1K<n<10K20---21If you like this please consider liking and donating (https://buymeacoffee.com/aiencoder)22 23# 🏭 llama-cpp-python Mega-Factory Wheels24 25> **"Stop waiting for `pip` to compile. Just install and run."**26 27The most complete collection of pre-built `llama-cpp-python` wheels in existence β€” **8,333 wheels** across every platform, Python version, backend, and CPU optimization level.28 29No more `cmake`, `gcc`, or compilation hell. No more waiting 10 minutes for a build that might fail. Just find your wheel and `pip install` it directly.30 31---32 33## πŸš€ Why These Wheels?34 35Standard wheels target the "lowest common denominator" to avoid crashes on old hardware. This collection goes further β€” the manylinux wheels are built using a massive **Everything Preset** targeting specific CPU instruction sets, maximizing your **Tokens per Second (T/s)**.36 37- **Zero Dependencies:** No `cmake`, `gcc`, or `nvcc` required on your target machine.38- **Every Platform:** Linux (manylinux, aarch64, i686, RISC-V), Windows (amd64, 32-bit), macOS (Intel + Apple Silicon).39- **Server-Grade Power:** Optimized builds for `Sapphire Rapids`, `Ice Lake`, `Alder Lake`, `Haswell`, and more.40- **Full Backend Support:** `OpenBLAS`, `MKL`, `Vulkan`, `CLBlast`, `OpenCL`, `RPC`, and plain CPU builds.41- **Cutting Edge:** Python `3.8` through experimental `3.14`, plus PyPy `pp38`–`pp310`.42- **GPU Too:** CUDA wheels (cu121–cu124) and macOS Metal wheels included.43 44---45 46## πŸ“Š Collection Stats47 48| Platform | Wheels |49|:---|---:|50| 🐧 Linux x86_64 (manylinux) | 4,940 |51| 🍎 macOS Intel (x86\_64) | 1,040 |52| πŸͺŸ Windows (amd64) | 1,010 |53| πŸͺŸ Windows (32-bit) | 634 |54| 🍎 macOS Apple Silicon (arm64) | 289 |55| 🐧 Linux i686 | 214 |56| 🐧 Linux aarch64 | 120 |57| 🐧 Linux x86\_64 (plain) | 81 |58| 🐧 Linux RISC-V | 5 |59| **Total** | **8,333** |60 61The manylinux builds alone cover **3,600+ combinations** across versions, backends, Python versions, and CPU profiles.62 63---64 65## πŸš€ How to Install66 67### Quick Install68 69Find your wheel filename (see naming convention below), then:70 71```bash72pip install "https://huggingface.co/datasets/AIencoder/llama-cpp-wheels/resolve/main/YOUR_WHEEL_NAME.whl"73```74 75### Common Examples76 77```bash78# Linux x86_64, Python 3.11, OpenBLAS, Haswell CPU (most common Linux setup)79pip install "https://huggingface.co/datasets/AIencoder/llama-cpp-wheels/resolve/main/llama_cpp_python-0.3.18+openblas_haswell-cp311-cp311-manylinux_2_31_x86_64.whl"80 81# Linux x86_64, Python 3.12, Basic CPU (maximum compatibility)82pip install "https://huggingface.co/datasets/AIencoder/llama-cpp-wheels/resolve/main/llama_cpp_python-0.3.18+basic_basic-cp312-cp312-manylinux_2_31_x86_64.whl"83 84# Windows, Python 3.1185pip install "https://huggingface.co/datasets/AIencoder/llama-cpp-wheels/resolve/main/llama_cpp_python-0.3.18-cp311-cp311-win_amd64.whl"86 87# macOS Apple Silicon, Python 3.1288pip install "https://huggingface.co/datasets/AIencoder/llama-cpp-wheels/resolve/main/llama_cpp_python-0.3.18-cp312-cp312-macosx_11_0_arm64.whl"89 90# macOS Intel, Python 3.1191pip install "https://huggingface.co/datasets/AIencoder/llama-cpp-wheels/resolve/main/llama_cpp_python-0.3.18-cp311-cp311-macosx_10_9_x86_64.whl"92 93# Linux ARM64 (Raspberry Pi, AWS Graviton), Python 3.1194pip install "https://huggingface.co/datasets/AIencoder/llama-cpp-wheels/resolve/main/llama_cpp_python-0.3.18-cp311-cp311-linux_aarch64.whl"95```96 97---98 99## πŸ“ Wheel Naming Convention100 101### manylinux wheels (custom-built)102 103```104llama_cpp_python-{version}+{backend}_{profile}-{pytag}-{pytag}-{platform}.whl105```106 107**Versions covered:** `0.3.0` through `0.3.18+`108 109**Backends:**110 111| Backend | Description |112|:---|:---|113| `openblas` | OpenBLAS BLAS acceleration β€” best general-purpose CPU performance |114| `mkl` | Intel MKL acceleration β€” best on Intel CPUs |115| `basic` | No BLAS, maximum compatibility |116| `vulkan` | Vulkan GPU backend |117| `clblast` | CLBlast OpenCL GPU backend |118| `opencl` | Generic OpenCL GPU backend |119| `rpc` | Distributed inference over network |120 121**CPU Profiles:**122 123| Profile | Instruction Sets | Era | Notes |124|:---|:---|:---|:---|125| `basic` | x86-64 baseline | Any | Maximum compatibility |126| `sse42` | SSE 4.2 | 2008+ | Nehalem |127| `sandybridge` | AVX | 2011+ | |128| `ivybridge` | AVX + F16C | 2012+ | |129| `haswell` | AVX2 + FMA + BMI2 | 2013+ | **Most common** |130| `skylakex` | AVX-512 | 2017+ | |131| `icelake` | AVX-512 + VNNI + VBMI | 2019+ | |132| `alderlake` | AVX-VNNI | 2021+ | |133| `sapphirerapids` | AVX-512 BF16 + AMX | 2023+ | Highest performance |134 135**Python tags:** `cp38`, `cp39`, `cp310`, `cp311`, `cp312`, `cp313`, `cp314`, `pp38`, `pp39`, `pp310`136 137**Platform:** `manylinux_2_31_x86_64` (glibc 2.31+, compatible with Ubuntu 20.04+, Debian 11+)138 139### Windows / macOS / Linux ARM wheels (from abetlen)140 141```142llama_cpp_python-{version}-{pytag}-{pytag}-{platform}.whl143```144 145These are the official pre-built wheels from the upstream maintainer, covering versions `0.2.82` through `0.3.18+`.146 147---148 149## πŸ” How to Find Your Wheel150 1511. **Identify your Python version:** `python --version` β†’ e.g. `3.11` β†’ tag `cp311`1522. **Identify your platform:**153   - Linux x86\_64 β†’ `manylinux_2_31_x86_64`154   - Windows 64-bit β†’ `win_amd64`155   - macOS Apple Silicon β†’ `macosx_11_0_arm64`156   - macOS Intel β†’ `macosx_10_9_x86_64`1573. **Pick a backend** (manylinux only): `openblas` for most use cases1584. **Pick a CPU profile** (manylinux only): `haswell` works on virtually all modern CPUs1595. **Browse the files** in this repo or construct the filename directly160 161---162 163## πŸ—οΈ Sources & Credits164 165### manylinux Wheels β€” Built by AIencoder166The 4,940 manylinux x86\_64 wheels were built by a distributed **4-worker HuggingFace Space factory** system (`AIencoder/wheel-factory-*`) β€” a custom-built automated pipeline covering every possible llama.cpp cmake option on manylinux:167- Every backend: OpenBLAS, MKL, Basic, Vulkan, CLBlast, OpenCL, RPC168- Every CPU hardware profile from baseline x86-64 up to Sapphire Rapids AMX169- Python 3.8 through 3.14170- llama-cpp-python versions 0.3.0 through 0.3.18+171 172### Windows / macOS / Linux ARM Wheels β€” abetlen173The remaining 3,393 wheels (Windows, macOS, Linux aarch64/i686/riscv64, PyPy) were sourced from the official releases by **Andrei Betlen ([@abetlen](https://github.com/abetlen))**, the original author and maintainer of `llama-cpp-python`. These include:174- CPU wheels for all platforms via `https://abetlen.github.io/llama-cpp-python/whl/cpu/`175- Metal wheels for macOS GPU acceleration176- CUDA wheels (cu121–cu124) for Windows and Linux177 178> All credit for the underlying library goes to **Georgi Gerganov ([@ggerganov](https://github.com/ggerganov))** and the [llama.cpp](https://github.com/ggml-org/llama.cpp) team, and to **Andrei Betlen** for the Python bindings.179 180---181 182## πŸ“ Notes183 184- All wheels are **MIT licensed** (same as llama-cpp-python upstream)185- manylinux wheels require **glibc 2.31+** (Ubuntu 20.04+, Debian 11+)186- `manylinux` and `linux_x86_64` are **not the same thing** β€” manylinux wheels have broad distro compatibility, plain linux wheels do not187- CUDA wheels require the matching CUDA toolkit to be installed188- Metal wheels require macOS 11.0+ and an Apple Silicon or AMD GPU189- This collection is updated periodically as new versions are released190