Ck773/llama-cpp
0
1https://github.com/ggerganov/llama.cpp/blob/master/examples/server/README.md2 3https://llama-cpp-python.readthedocs.io/en/latest/server/4 5```6curl --request GET \7 --url https://ck773-llama-cpp-lab.hf.space/v1/health8 9curl --request POST \10 --url https://ck773-llama-cpp-lab.hf.space/completion \11 --header "Content-Type: application/json" \12 --data '{"prompt": "Hello","n_predict": 128}'13```14 15 16python3 -m llama_cpp.server --config_file <config_file>17 18 19```json20{21 "host": "0.0.0.0",22 "port": 8080,23 "models": [24 {25 "model": "models/OpenHermes-2.5-Mistral-7B-GGUF/openhermes-2.5-mistral-7b.Q4_K_M.gguf",26 "model_alias": "gpt-3.5-turbo",27 "chat_format": "chatml",28 "n_gpu_layers": -1,29 "offload_kqv": true,30 "n_threads": 12,31 "n_batch": 512,32 "n_ctx": 204833 },34 {35 "model": "models/OpenHermes-2.5-Mistral-7B-GGUF/openhermes-2.5-mistral-7b.Q4_K_M.gguf",36 "model_alias": "gpt-4",37 "chat_format": "chatml",38 "n_gpu_layers": -1,39 "offload_kqv": true,40 "n_threads": 12,41 "n_batch": 512,42 "n_ctx": 204843 },44 {45 "model": "models/ggml_llava-v1.5-7b/ggml-model-q4_k.gguf",46 "model_alias": "gpt-4-vision-preview",47 "chat_format": "llava-1-5",48 "clip_model_path": "models/ggml_llava-v1.5-7b/mmproj-model-f16.gguf",49 "n_gpu_layers": -1,50 "offload_kqv": true,51 "n_threads": 12,52 "n_batch": 512,53 "n_ctx": 204854 },55 {56 "model": "models/mistral-7b-v0.1-GGUF/ggml-model-Q4_K.gguf",57 "model_alias": "text-davinci-003",58 "n_gpu_layers": -1,59 "offload_kqv": true,60 "n_threads": 12,61 "n_batch": 512,62 "n_ctx": 204863 },64 {65 "model": "models/replit-code-v1_5-3b-GGUF/replit-code-v1_5-3b.Q4_0.gguf",66 "model_alias": "copilot-codex",67 "n_gpu_layers": -1,68 "offload_kqv": true,69 "n_threads": 12,70 "n_batch": 1024,71 "n_ctx": 921672 }73 ]74}75```