A Compose stack for vLLM, llama.cpp, and a proxy
Compose configuration for vLLM, llama.cpp, a Rust proxy, PostgreSQL, and Qdrant on rootless Podman. Setup and startup steps for local LLM and AI infrastructure integration.
Introduction
A Rust proxy sits in front of vLLM and llama.cpp. The servers handle inference; the proxy handles RAG and workflow control.
This GPU-enabled, rootless Podman setup connects an OpenAI-compatible Proxy API, Qdrant, PostgreSQL, vLLM, and llama.cpp. It provides a container layout for local LLM integration.
Background and Architectural Goals
Separating RAG and job management from inference makes model replacement easier.
- Backend (Inference): vLLM and llama.cpp function purely as “OpenAI-compatible inference endpoints.” They do nothing else.
- Frontend (Proxy): A custom Proxy API written in Rust (Axum) intercepts all requests. It handles RAG searches (using Qdrant and PostgreSQL) and triggers Dagster pipelines via NATS before optionally forwarding the prompt to the backend.
- Data Stack: Qdrant is dedicated solely to vector storage, while PostgreSQL manages metadata, job queues, and document chunk pairing.
docker-compose.yml defines the separate inference and control services.
Container Configuration (docker-compose.yml)
The intended path is /opt/containers/compose/llm-stack/docker-compose.yml. Each container uses cap_drop: ["ALL"] and no-new-privileges:true, with tmpfs limiting writable directories. db_net and llm_net are separate internal networks.
version: "3.9"
name: llm-stack
networks:
db_net:
driver: bridge
internal: true
llm_net:
driver: bridge
internal: true
volumes:
pg_data:
driver: local
driver_opts:
type: none
o: bind
device: /mnt/data/postgres/data
qdrant_data:
driver: local
driver_opts:
type: none
o: bind
device: /mnt/data/qdrant/data
api_data:
driver: local
driver_opts:
type: none
o: bind
device: /mnt/data/llm-api
vllm_models:
driver: local
driver_opts:
type: none
o: bind
device: /mnt/data/models/vllm
llama_models:
driver: local
driver_opts:
type: none
o: bind
device: /mnt/data/models/llama
services:
postgres:
image: postgres:17
container_name: pg
restart: always
command: ["postgres","-c","max_connections=300","-c","shared_buffers=4GB","-c","wal_compression=on"]
environment:
POSTGRES_USER: ${PG_USER:-loft}
POSTGRES_PASSWORD: ${PG_PASSWORD:-change_me}
POSTGRES_DB: ${PG_DB:-loftdb}
healthcheck:
test: ["CMD-SHELL","pg_isready -U $$POSTGRES_USER -d $$POSTGRES_DB -h 127.0.0.1"]
interval: 10s
timeout: 3s
retries: 10
networks: [db_net]
volumes:
- pg_data:/var/lib/postgresql/data:rw
read_only: true
tmpfs:
- /tmp:rw,nosuid,nodev,noexec,size=256m
- /var/run/postgresql:rw,mode=775
security_opt: ["no-new-privileges:true"]
cap_drop: ["ALL"]
ulimits:
nofile: 262144
qdrant:
image: qdrant/qdrant:latest
restart: always
environment:
QDRANT__SERVICE__GRPC_PORT: 6334
QDRANT__STORAGE__WAL_MEMORY_CAPACITY: "33554432"
QDRANT__STORAGE__OPTIMIZERS__DEFAULT_SEGMENT_NUMBER: "2"
healthcheck:
test: ["CMD","/qdrant/tools/healthcheck.sh"]
interval: 10s
timeout: 3s
retries: 10
networks: [db_net]
volumes:
- qdrant_data:/qdrant/storage:rw
ports:
- "6333:6333"
- "6334:6334"
read_only: false
security_opt: ["no-new-privileges:true"]
cap_drop: ["ALL"]
ulimits:
nofile: 262144
vllm:
image: vllm/vllm-openai:latest
restart: always
command:
[
"python","-m","vllm.entrypoints.openai.api_server",
"--model","${VLLM_MODEL:-/models}",
"--host","0.0.0.0",
"--port","8000",
"--tensor-parallel-size","${VLLM_TP:-1}",
"--max-num-seqs","${VLLM_MAX_SEQS:-32}"
]
environment:
NVIDIA_VISIBLE_DEVICES: "all"
devices:
- "nvidia.com/gpu=all"
healthcheck:
test: ["CMD","curl","-sf","http://127.0.0.1:8000/health"]
interval: 10s
timeout: 5s
retries: 20
networks: [llm_net]
volumes:
- vllm_models:/models:ro
ports:
- "8000:8000"
security_opt: ["no-new-privileges:true"]
cap_drop: ["ALL"]
ulimits:
nofile: 262144
llamacpp:
image: ghcr.io/ggerganov/llama.cpp:full
restart: always
command:
[
"server",
"-m","/models/${LLAMA_MODEL:-model.gguf}",
"--host","0.0.0.0",
"--port","8080",
"--mlock",
"--no-mmap",
"--ctx-size","${LLAMA_CTX:-8192}",
"--batch-size","${LLAMA_BATCH:-512}",
"--embedding"
]
environment:
NVIDIA_VISIBLE_DEVICES: "all"
devices:
- "nvidia.com/gpu=all"
healthcheck:
test: ["CMD","curl","-sf","http://127.0.0.1:8080/health"]
interval: 10s
timeout: 5s
retries: 20
networks: [llm_net]
volumes:
- llama_models:/models:ro
ports:
- "8080:8080"
security_opt: ["no-new-privileges:true"]
cap_drop: ["ALL"]
ulimits:
nofile: 262144
openai-proto-api:
image: ghcr.io/your-org/openai-proto-api:latest
restart: always
environment:
VLLM_BASE_URL: http://vllm:8000
LLAMA_BASE_URL: http://llamacpp:8080
QDRANT_URL: http://qdrant:6333
QDRANT_API_KEY: ${QDRANT_API_KEY:-}
PGHOST: postgres
PGPORT: 5432
PGUSER: ${PG_USER:-loft}
PGPASSWORD: ${PG_PASSWORD:-change_me}
PGDATABASE: ${PG_DB:-loftdb}
API_PORT: 9000
API_KEY: ${API_KEY:-change_me}
depends_on:
- qdrant
- postgres
- vllm
healthcheck:
test: ["CMD","curl","-sf","http://127.0.0.1:9000/healthz"]
interval: 10s
timeout: 5s
retries: 30
networks:
- llm_net
- db_net
volumes:
- api_data:/var/lib/llm-api
ports:
- "9000:9000"
read_only: true
tmpfs:
- /tmp:rw,nosuid,nodev,noexec,size=256m
security_opt: ["no-new-privileges:true"]
cap_drop: ["ALL"]
ulimits:
nofile: 262144
Environment Variables (.env)
Configuration is in .env at /opt/containers/compose/llm-stack/.env.
PG_USER=loft
PG_PASSWORD=change_me
PG_DB=loftdb
QDRANT_API_KEY=
API_KEY=change_me
VLLM_MODEL=/models
VLLM_TP=1
VLLM_MAX_SEQS=32
LLAMA_MODEL=model.gguf
LLAMA_CTX=8192
LLAMA_BATCH=512
Setup and Startup Procedure
Initial setup includes setting PostgreSQL directory ownership.
mkdir -p /mnt/data/{postgres/data,qdrant/data,models/vllm,models/llama,llm-api}
podman unshare chown -R 999:999 /mnt/data/postgres/data
cd /opt/containers/compose/llm-stack
podman-compose up -d
Postgres runs as uid=999. Use podman unshare chown on the host mount directory; otherwise initialization fails with a permission error.
GPU Allocation
Both inference services use devices: ["nvidia.com/gpu=all"]. This assumes CDI is configured on the host with nvidia-ctk runtime configure --runtime=crun.
Systemd Integration (Auto-start)
Register systemd user services so the stack starts after a reboot.
podman generate systemd --files --name llm-stack_openai-proto-api
podman generate systemd --files --name llm-stack_vllm
podman generate systemd --files --name llm-stack_llamacpp
podman generate systemd --files --name llm-stack_qdrant
podman generate systemd --files --name llm-stack_postgres
mkdir -p ~/.config/systemd/user
mv *.service ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now llm-stack_*.service
loginctl enable-linger ksh3
loginctl enable-linger keeps user processes running after logout.
Results and Discussion
Only the API container (openai-proto-api) joins both db_net and llm_net.
vLLM and llama.cpp stay on llm_net; Qdrant and Postgres stay on the data network, separated from the inference and external networks. The OpenAI-compatible proxy accepts requests and routes them by model name.
Future Work
This Compose does not include log monitoring. The plan was to collect stdout and journal through promtail, and extend proxy metadata processing for RAG.
