Team Ai
Apppublic

jaothan/podman_llamacpp_python

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes
docker-compose.yml60 linesDownload Raw Back to base
1#### 4. **Deploying in a Containerized Environment**
2#If you're using Docker or Podman, ensure the containers can communicate with each other. For example:
3
4#- **Docker Compose**:
5#  Create a `docker-compose.yml` file to define both the model server and the chat application:
6
7  #```yaml
8  version: "3"
9  services:
10    model_server:
11      image: my_model_server_image
12      ports:
13        - "8001:8001"
14      environment:
15        - PORT=8001
16      networks:
17        - my_network
18chat_app:
19      image: my_chat_app_image
20      environment:
21        - MODEL_ENDPOINT=http://model_server:8001
22      depends_on:
23        - model_server
24      networks:
25        - my_network
26
27  networks:
28    my_network:
29  #```
30
31  #- The `MODEL_ENDPOINT` for the chat application is set to `http://model_server:8001`, which uses Docker's internal DNS to resolve the model server's container name.
32
33#- **Docker Networking**:
34#  If you're not using Docker Compose, you can create a custom network and attach both containers to it:
35
36#  ```bash
37  # Create a custom network
38 # docker network create my_network
39
40  # Run the model server container
41 # docker run -d --name model_server --network my_network -p 8001:8001 my_model_server_image
42 # Run the chat application container
43 # docker run -d --name chat_app --network my_network -e MODEL_ENDPOINT=http://model_server:8001 my_chat_app_image
44
45
46
47
48#### 5. **Testing the Endpoint**
49#To ensure the model server is working as expected, you can test the endpoint directly using `curl` or a tool like Postman:
50
51#```bash
52#curl -X POST http://localhost:8001/generate -H "Content-Type: application/json" -d '{"prompt": "Hello, model!"}'
53#```
54
55#Expected response:
56#```json
57#{
58#  "response": "Generated response for: Hello, model!"
59#}
60