Adding MLflow and MinIO

I added MLflow Tracking Server and MinIO to the devstack running the Dagster + NATS event pipeline. Dagster records orchestration; MLflow records individual experiments. correlation_id connects them for AI system development.

The change covered eight files and three services: MinIO, minio-init, and MLflow. Startup required fixes for AirPlay port conflicts, missing psycopg2, and database creation on an existing PostgreSQL volume.


Why MLflow Was Needed

Dagster already pulled NATS JetStream events through sensors, ran jobs, and persisted to PostgreSQL. MLflow adds experiment parameters, metric histories, and artifacts.

The responsibilities are:

  • Dagster: Bird’s-eye view (what happened)
  • MLflow: Detailed view (what happened inside)
  • correlation_id links both layers
  • Dagster asset metadata records experiment_id, run_id, tracker_url, summary
  • MLflow run tags include correlation_id

This prepares tracking for Phase 2 fine-tuning and investment prediction models.


Existing Infrastructure

The overall devstack configuration is as follows.

Core services:

ServiceRole
NATS (JetStream)Messaging
PostgreSQL 18 (pgvector + JIT, 4GB SHM)Data store
Dagster 3 containers (webserver + daemon + user-code gRPC)Orchestration
VectorTelemetry collection
FastAPI Reranker (ColBERT)Reranking

Lakehouse profile (optional): Nessie, Trino, dbt-fusion

3-host configuration:

  • Storage server
  • Desktop Mac — NATS co-located with gateway
  • Compute server — Dagster and PostgreSQL

MinIO was in the README storage plan but absent from podman-compose.yml. I added it as the S3-compatible artifact store.


Change Plan

  1. Add MLflow Tracking Server to podman-compose.yml. Backend store reuses the existing PostgreSQL with a new mlflow database; artifact store uses MinIO’s s3://mlflow/
  2. Add CREATE DATABASE mlflow; to init.sql
  3. Add dagster-mlflow resource to the Dagster side, making experiment_name and mlflow_tracking_uri configurable
  4. Add MLFLOW_TRACKING_URI and MLFLOW_S3_ENDPOINT_URL to environment variables
  5. Update README.md

MLflow is dedicated to experiment tracking; orchestration remains with Dagster.


dagster-mlflow API Investigation

I checked compatibility with the pinned Dagster 1.12.14.

Dagster 1.12.X maps to dagster-mlflow 0.28.X. The selected version is dagster-mlflow==0.28.14, with dagster==1.12.14, mlflow, pandas<3.0.0, and protobuf!=5.29.0.

I installed it in a temporary venv and read its source:

  • mlflow_tracking is an old-style @resource decorator ResourceDefinition. Not the ConfigurableResource pattern used by existing PostgresResource and EmbeddingResource
  • Ops access it via required_resource_keys={"mlflow"} and context.resources.mlflow
  • The end_mlflow_on_run_finished hook must be applied to jobs or MLflow runs will hang. This is mandatory
  • The MlflowMeta metaclass proxies all mlflow.* methods, so log_params(), log_metric(), log_artifact() can be called directly on the resource object
  • S3 credentials can be passed via env config, but this is unnecessary if already set as container environment variables

The package uses an old-style resource, unlike the existing codebase.


Implementation

I changed or added eight files.

podman-compose.yml: 3 Services Added

MinIO uses docker.io/minio/minio:latest, ports 9000/9001, and minio-data. The one-shot minio-init container uses mc to create s3://mlflow/.

MLflow Tracking Server (ghcr.io/mlflow/mlflow:v2.21.3, port 5000):

  --backend-store-uri=postgresql://postgres:postgres@postgres:5432/mlflow
--default-artifact-root=s3://mlflow/
  

PostgreSQL is the backend store and MinIO the artifact store. depends_on waits for postgres service_healthy and minio-init service_completed_successfully.

I added these variables to dagster-user-code and dagster-daemon:

  MLFLOW_TRACKING_URI: "http://mlflow:5000"
MLFLOW_S3_ENDPOINT_URL: "http://minio:9000"
AWS_ACCESS_KEY_ID: minioadmin
AWS_SECRET_ACCESS_KEY: minioadmin
  

init.sql

  CREATE DATABASE mlflow;
  

dagster/pyproject.toml

I added dagster-mlflow==0.28.14 and boto3 for S3 artifact access.

dagster/project/resources/mlflow.py (new)

  mlflow_tracking.configured({
    "experiment_name": os.getenv("MLFLOW_EXPERIMENT_NAME", "agent-gateway"),
    "mlflow_tracking_uri": os.getenv("MLFLOW_TRACKING_URI", "http://mlflow:5000"),
    "extra_tags": {"project": "agent-gateway"},
})
  

dagster/project/defs.py

I added "mlflow": mlflow_resource to Definitions resources.

.envrc

I added MLFLOW_TRACKING_URI and MLFLOW_S3_ENDPOINT_URL.


Troubleshooting

Port 5000 Conflict

podman compose up -d mlflow failed.

  Error response from daemon: "listen tcp :5000: bind: address already in use"
  

lsof -i :5000 showed that macOS ControlCenter (AirPlay Receiver) occupied port 5000.

I changed the host mapping to "${MLFLOW_PORT:-5050}:5000". Internal traffic stays on 5000, including MLFLOW_TRACKING_URI: http://mlflow:5000 for container traffic. Host access uses 5050.

I briefly changed internal MLFLOW_TRACKING_URI to 5050, then restored 5000 for container traffic. .envrc uses MLFLOW_PORT="5050" and "http://${COMPUTE_HOST}:${MLFLOW_PORT}" for external access.

Official Image Missing psycopg2

After the port fix, startup still crashed.

  ModuleNotFoundError: No module named 'psycopg2'
  

The selected ghcr.io/mlflow/mlflow:v2.21.3 image lacked psycopg2. Its SQLite/MySQL setup needed an added PostgreSQL driver.

I added devstack/mlflow/Dockerfile:

  FROM ghcr.io/mlflow/mlflow:v2.21.3
RUN pip install --no-cache-dir psycopg2-binary boto3
  

In podman-compose.yml, I replaced direct image: use with build: context: ./devstack/mlflow, tagged localhost/agent-gateway/mlflow:2.21.3.

Adding a Database to an Existing Volume

The rebuilt image then encountered a PostgreSQL error.

  FATAL: database "mlflow" does not exist
  

Adding CREATE DATABASE mlflow; to init.sql had no effect on the three-week-old postgres-data volume. docker-entrypoint-initdb.d runs only on initial startup.

I created the database manually:

  podman exec agent-gateway-postgres-1 psql -U postgres -c "CREATE DATABASE mlflow;"
  

The init.sql entry remains for new environments or recreated volumes.


Result

After Alembic migrations, gunicorn started with four workers.

  [2026-03-12 02:07:58 +0000] [24] [INFO] Starting gunicorn 23.0.0
[2026-03-12 02:07:58 +0000] [24] [INFO] Listening at: http://0.0.0.0:5000 (24)
  

experiments/search returned Default. Its artifact_location was s3://mlflow/0, pointing to MinIO.

Changed Files

FileChange
podman-compose.ymlAdded 3 services: MinIO, minio-init, mlflow + Dagster env vars
devstack/postgres/init.sqlCREATE DATABASE mlflow
devstack/dagster/pyproject.tomlAdded dagster-mlflow==0.28.14, boto3
devstack/dagster/project/resources/mlflow.pyNew: mlflow_tracking configured resource
devstack/dagster/project/defs.pymlflow resource registration
devstack/mlflow/DockerfileNew: psycopg2-binary + boto3
.envrcMLFLOW_PORT, MLFLOW_TRACKING_URI, MLFLOW_S3_ENDPOINT_URL
README.mdService list, environment variables, topology diagram

Design Notes

  • dagster-mlflow is an old-style resource (@resource decorator), not ConfigurableResource. The @end_mlflow_on_run_finished hook must always be applied to jobs
  • Inter-container communication uses internal port 5000; host access uses port 5050 (macOS AirPlay Receiver workaround)
  • When a PostgreSQL volume already exists, new database creation in init.sql must be applied manually