dougeeai/llama-cpp-python-wheels
12
1---2license: mit3tags:4- llama-cpp5- llama-cpp-python6- gguf7- cuda8- windows9- prebuilt-wheels10- quantization11- local-llm12---13 14# llama-cpp-python Pre-built Windows Wheels15 16**Stop fighting with Visual Studio and CUDA Toolkit.** Just download and run.17 18Pre-compiled `llama-cpp-python` wheels for Windows across CUDA versions and GPU architectures.19 20## Quick Start21 221. **Find your GPU** in the compatibility list below232. **Download** the wheel for your GPU from [GitHub Releases](https://github.com/dougeeai/llama-cpp-python-wheels/releases) or [find your card on the README table](https://github.com/dougeeai/llama-cpp-python-wheels)243. **Install**: `pip install <downloaded-wheel-file>.whl`254. **Run** your GGUF models immediately26 27> **Platform Support:** 28> โ
Windows 10/11 64-bit (available now, biggest pain point) 29> ๐ Linux support coming soon 30 31## Supported GPUs32 33### RTX 50 Series (Blackwell - sm_100)34RTX 5090, 5080, 5070 Ti, 5070, 5060 Ti, 5060, RTX PRO 6000 Blackwell, B100, B200, GB20035 36### RTX 40 Series (Ada Lovelace - sm_89)37RTX 4090, 4080, 4070 Ti, 4070, 4060 Ti, 4060, RTX 6000 Ada, RTX 5000 Ada, L40, L40S38 39### RTX 30 Series (Ampere - sm_86)40RTX 3090, 3090 Ti, 3080 Ti, 3080, 3070 Ti, 3070, 3060 Ti, 3060, RTX A6000, A5000, A400041 42### RTX 20 Series & GTX 16 Series (Turing - sm_75)43RTX 2080 Ti, 2080 Super, 2070 Super, 2060, GTX 1660 Ti, 1660 Super, 1650, Quadro RTX 8000, Tesla T444 45[View full compatibility table โ](https://github.com/dougeeai/llama-cpp-python-wheels#available-wheels)46 47## Usage Example48```python49from llama_cpp import Llama50 51# Load your GGUF model with GPU acceleration52llm = Llama(53 model_path="./models/llama-3-8b.Q4_K_M.gguf",54 n_gpu_layers=-1, # Offload all layers to GPU55 n_ctx=2048 # Context window56)57 58# Generate text59response = llm(60 "Write a haiku about artificial intelligence:",61 max_tokens=50,62 temperature=0.763)64 65print(response['choices'][0]['text'])66```67 68## Download Wheels69 70โก๏ธ **[Download from GitHub Releases](https://github.com/dougeeai/llama-cpp-python-wheels/releases)**71 72### Available Configurations:73- **CUDA Versions**: 11.8, 12.1, 13.074- **Python Versions**: 3.10, 3.11, 3.12, 3.1375- **Architectures**: sm_75 (Turing), sm_86 (Ampere), sm_89 (Ada), sm_100 (Blackwell)76 77## What This Solves78 79โ No Visual Studio required 80โ No CUDA Toolkit installation needed 81โ No compilation errors 82โ No "No CUDA toolset found" issues 83โ
Works immediately with GGUF models 84โ
Full GPU acceleration out of the box 85 86## Installation87 88Download the wheel matching your configuration and install:89```bash90# Example for RTX 4090 with Python 3.12 and CUDA 13.091pip install llama_cpp_python-0.3.16+cuda13.0.sm89.ada-cp312-cp312-win_amd64.whl92```93 94## Build Details95 96All wheels are built with:97- Visual Studio 2019/2022 Build Tools98- Official NVIDIA CUDA Toolkits (11.8, 12.1, 13.0)99- Optimized CMAKE_CUDA_ARCHITECTURES for each GPU generation100- Built from official [llama-cpp-python](https://github.com/abetlen/llama-cpp-python) source101 102## Contributing103 104**Need a different configuration?**105 106Open an [issue on GitHub](https://github.com/dougeeai/llama-cpp-python-wheels/issues) with:107- OS (Windows/Linux/macOS)108- Python version109- CUDA version110- GPU model111 112## Resources113 114- [GitHub Repository](https://github.com/dougeeai/llama-cpp-python-wheels)115- [Report Issues](https://github.com/dougeeai/llama-cpp-python-wheels/issues)116- [llama-cpp-python Documentation](https://github.com/abetlen/llama-cpp-python)117- [llama.cpp Project](https://github.com/ggerganov/llama.cpp)118 119## License120 121MIT License - Free to use for any purpose122 123Wheels are built from [llama-cpp-python](https://github.com/abetlen/llama-cpp-python) (MIT License)