KBaba7/llama.cpp
0
1# llama.cpp/example/run2 3The purpose of this example is to demonstrate a minimal usage of llama.cpp for running models.4 5```bash6llama-run granite3-moe7```8 9```bash10Description:11 Runs a llm12 13Usage:14 llama-run [options] model [prompt]15 16Options:17 -c, --context-size <value>18 Context size (default: 2048)19 -n, -ngl, --ngl <value>20 Number of GPU layers (default: 0)21 --temp <value>22 Temperature (default: 0.8)23 -v, --verbose, --log-verbose24 Set verbosity level to infinity (i.e. log all messages, useful for debugging)25 -h, --help26 Show help message27 28Commands:29 model30 Model is a string with an optional prefix of31 huggingface:// (hf://), ollama://, https:// or file://.32 If no protocol is specified and a file exists in the specified33 path, file:// is assumed, otherwise if a file does not exist in34 the specified path, ollama:// is assumed. Models that are being35 pulled are downloaded with .partial extension while being36 downloaded and then renamed as the file without the .partial37 extension when complete.38 39Examples:40 llama-run llama341 llama-run ollama://granite-code42 llama-run ollama://smollm:135m43 llama-run hf://QuantFactory/SmolLM-135M-GGUF/SmolLM-135M.Q2_K.gguf44 llama-run huggingface://bartowski/SmolLM-1.7B-Instruct-v0.2-GGUF/SmolLM-1.7B-Instruct-v0.2-IQ3_M.gguf45 llama-run https://example.com/some-file1.gguf46 llama-run some-file2.gguf47 llama-run file://some-file3.gguf48 llama-run --ngl 999 some-file4.gguf49 llama-run --ngl 999 some-file5.gguf Hello World50```51 