jaothan/podman_llamacpp_python
0
1#### 4. **Deploying in a Containerized Environment**
2#If you're using Docker or Podman, ensure the containers can communicate with each other. For example:
3
4#- **Docker Compose**:
5# Create a `docker-compose.yml` file to define both the model server and the chat application:
6
7 #```yaml
8 version: "3"
9 services:
10 model_server:
11 image: my_model_server_image
12 ports:
13 - "8001:8001"
14 environment:
15 - PORT=8001
16 networks:
17 - my_network
18chat_app:
19 image: my_chat_app_image
20 environment:
21 - MODEL_ENDPOINT=http://model_server:8001
22 depends_on:
23 - model_server
24 networks:
25 - my_network
26
27 networks:
28 my_network:
29 #```
30
31 #- The `MODEL_ENDPOINT` for the chat application is set to `http://model_server:8001`, which uses Docker's internal DNS to resolve the model server's container name.
32
33#- **Docker Networking**:
34# If you're not using Docker Compose, you can create a custom network and attach both containers to it:
35
36# ```bash
37 # Create a custom network
38 # docker network create my_network
39
40 # Run the model server container
41 # docker run -d --name model_server --network my_network -p 8001:8001 my_model_server_image
42 # Run the chat application container
43 # docker run -d --name chat_app --network my_network -e MODEL_ENDPOINT=http://model_server:8001 my_chat_app_image
44
45
46
47
48#### 5. **Testing the Endpoint**
49#To ensure the model server is working as expected, you can test the endpoint directly using `curl` or a tool like Postman:
50
51#```bash
52#curl -X POST http://localhost:8001/generate -H "Content-Type: application/json" -d '{"prompt": "Hello, model!"}'
53#```
54
55#Expected response:
56#```json
57#{
58# "response": "Generated response for: Hello, model!"
59#}
60 