Team Ai
Datasetpublic

Brunobkr/OFFFELLIA_llama.cpp_Neon_Themes

ΩFFFΣLLIα llama.cpp Inferência de LLM em C/C++, com a interface Web construída a partir deste código. ggml · build · server · licença Árvore local de llama.cpp. O nome ΩFFFΣLLIα llama.cpp aparece no centro de qualquer página da interface, em qualquer porta do llama-server. Os logos SVG da interface foram substituídos pelo caractere Ω. Modificações Área Comportamento nesta árvore Compilação O CMake principal força… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/OFFFELLIA_llama.cpp_Neon_Themes.

sourceHugging Faceupdated 3d agoView on Hugging Face
0likes47downloads
Dataset Card

ΩFFFΣLLIα llama.cpp

<div align="center">

<img src="./1.png" alt="llamaOFFFELLIA1984 Banner" width="100%" style="border-radius: 12px; margin-bottom: 15px; box-shadow: 0 4px 20px rgba(0, 102, 153, 0.3);" />

<p align="center"> <img src="./2.png" alt="llamaOFFFELLIA1984 Vintage Web UI - IBM Granite 4.2 Reasoning" width="100%" style="border-radius: 8px; margin-bottom: 15px; box-shadow: 0 4px 15px rgba(51, 255, 102, 0.15);" /> <img src="./3.jpeg" alt="PIX" width="100%" style="border-radius: 8px; margin-bottom: 15px; box-shadow: 0 4px 15px rgba(0, 102, 153, 0.2);" /> <img src="./4.png" alt="llamaOFFFELLIA1984 Vintage Web UI - IBM Granite 4.2 Reasoning" width="100%" style="border-radius: 8px; margin-bottom: 15px; box-shadow: 0 4px 15px rgba(51, 255, 102, 0.15);" /> <img src="./5.png" alt="llamaOFFFELLIA1984 Vintage Web UI - IBM Granite 4.2 Reasoning" width="100%" style="border-radius: 8px; margin-bottom: 15px; box-shadow: 0 4px 15px rgba(51, 255, 102, 0.15);" /> <img src="./6.png" alt="llamaOFFFELLIA1984 Vintage Web UI - IBM Granite 4.2 Reasoning" width="100%" style="border-radius: 8px; margin-bottom: 15px; box-shadow: 0 4px 15px rgba(51, 255, 102, 0.15);" /> </p>

Inferência de LLM em C/C++, com a interface Web construída a partir deste código.

![License: MIT](https://opensource.org/licenses/MIT) ![Upstream](https://github.com/ggml-org/llama.cpp)

ggml · build · server · licença

</div>

Árvore local de llama.cpp. O nome ΩFFFΣLLIα llama.cpp aparece no centro de qualquer página da interface, em qualquer porta do llama-server. Os logos SVG da interface foram substituídos pelo caractere Ω.

Modificações

ÁreaComportamento nesta árvore
CompilaçãoO CMake principal força LLAMA_BUILD_UI=ON e LLAMA_USE_PREBUILT_UI=OFF. Sempre que o servidor entra no build, o Vite compila tools/ui a partir do código local. Um cache antigo não reativa o download.
Atualização automáticaA compilação não baixa dist.tar.gz nem consulta o Hugging Face. Não há checagem SHA-256 do pacote da interface.
PWAO service worker não é registrado e não há aviso de versão nova. sw.js, manifest, Workbox e version.json não são exigidos para embutir a interface.
Favicon e SVGnpm run build é só vite build. O gerador de assets PWA não lê nem regrava favicon.svg. O HTML não declara favicon e o manifesto não lista ícones.
LogosO logo da barra lateral e o logo MCP renderizam Ω no lugar do SVG.
TemasEm Theme há cinco opções neon: Azul neon, Vermelho neon, Verde neon, Preto e cinza neon e Alumínio escovado. Cada uma troca a paleta e acende bordas, botões e o nome central.
MCPO proxy CORS da interface (--ui-mcp-proxy) fica ligado por padrão. --no-ui-mcp-proxy desliga. Servidores MCP ainda pedem --mcp-servers-config ou --mcp-servers-json. As ferramentas de shell continuam desligadas sem --tools ou --agent.

A primeira compilação do servidor precisa de Node.js e npm, porque o Vite instala as dependências da interface e gera os assets embutidos.

Não exponha o llama-server fora da máquina enquanto o proxy MCP estiver ativo.

Quick start

A few options to get llama.cpp installed on your machine:

bash
# curl
curl -LsSf https://llama.app/install.sh | sh

# powershell
irm https://llama.app/install.ps1 | iex
  • —Visit https://llama.app and follow the instructions
  • —Run with Docker - see our Docker documentation
  • —Download pre-built binaries from the releases page
  • —Build from source by cloning this repository - check out our build guide

Once installed:

sh
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

<table align="center"> <tr> <td align="center" width=50%> <img width="1310" height="888" alt="VLM session with llama cli" src="https://github.com/user-attachments/assets/88726b48-1713-48aa-a525-95a02e78afc4" /> <i>VLM session with <b>llama cli</b></i> </td> <td align="center"> <img width="1392" height="958" alt="Built-in web UI against llama serve running Qwen 3.6" src="https://github.com/user-attachments/assets/b402f972-2e32-4def-8771-8d849f08cf2e" /> <i>Built-in web UI against <b>llama serve</b></i> </td> </tr> <table>

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • —Plain C/C++ implementation without any dependencies
  • —Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • —AVX, AVX2, AVX512 and AMX support for x86 architectures
  • —RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • —1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • —Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • —Vulkan and SYCL backend support
  • —CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

BackendTarget devices
BLASAll
BLISAll
CANNAscend NPU
CUDANvidia GPU
HIPAMD GPU
HexagonSnapdragon
IBM zDNNIBM Z & LinuxONE
MUSAMoore Threads GPU
MetalApple Silicon
OpenCLAdreno GPU
[OpenVINO [In Progress]](docs/backend/OPENVINO.md)Intel CPUs, GPUs, and NPUs
RPCAll
SYCLIntel GPU
VirtGPUVirtGPU APIR
VulkanGPU
WebGPUAll
ZenDNNAMD CPU

Documentation

Tools
  • —cli
  • —completion
  • —server
  • —GBNF grammars
Development
  • —How to build
  • —Running on Docker
  • —Build on Android
  • —Multi-GPU usage
  • —Performance troubleshooting
  • —GGML tips & tricks
  • —XCFramework
  • —Completions
  • —Models
  • —Release process

Contributing

  • —Contributors can open PRs
  • —Collaborators will be invited based on contributions
  • —Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • —Any help with managing issues, PRs and projects is very appreciated!
  • —Read the CONTRIBUTING.md for more information

Acknowledgements

  • —yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • —nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
  • —nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • —mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • —sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain