Introduction

I stabilized smartctl-exporter by separating rootful exporters from rootless applications and standardizing the reload procedure. This note covers the NVMe workaround and the monitoring operating rules.

The earlier node_exporter rollout established central scraping from Storage. Next I addressed smartctl-exporter, especially nvme0, and the mixed rootful / rootless service responsibilities.

Background and Motivation

The node_exporter rollout added host metrics to central Prometheus on Storage. Scraping worked, but exporter placement was still inconsistent.

Two problems recurred: unclear dependencies across rootful / rootless scopes, and missed daemon-reload commands after unit or compose edits. smartctl-exporter v0.14.0 also lacked --smartctl.device-opts, so NVMe needed a workaround.

For infrastructure maintenance, I standardized placement, dependencies, and the commands to run after a change.

Directory Standard

I used the following directory layout for units and compose files.

  # rootful (administrator-managed)
sudo mkdir -p /opt/containers/{compose,systemd}/{rootful,rootless}
sudo mkdir -p /opt/prometheus/exporters/smartctl/bin
  

Device-facing services such as smartctl-exporter are rootful. Applications such as prometheus and loki are rootless. Both use the same directory structure.

systemd (Rootful)

I used a rootful oneshot unit for smartctl-exporter, matching the existing one-service-per-stack pattern.

  # /etc/systemd/system/smartctl-exporter.service
[Unit]
Description=smartctl-exporter (rootful)
After=network-online.target
Wants=network-online.target

[Service]
Type=oneshot
RemainAfterExit=yes
WorkingDirectory=/opt/containers/compose/rootful
ExecStart=/usr/bin/podman compose up -d
ExecStop=/usr/bin/podman compose down
TimeoutStopSec=30s

[Install]
WantedBy=multi-user.target
  

Enable it with:

  sudo systemctl daemon-reload
sudo systemctl enable --now smartctl-exporter.service
  

Exporters that need devices or elevated permissions belong in system scope, giving maintenance a consistent starting point.

Compose (Rootful / smartctl-exporter v0.14.0)

smartctl-exporter v0.14.0 lacks --smartctl.device-opts, so the wrapper adds -T permissive.

  sudo tee /opt/prometheus/exporters/smartctl/bin/smartctl-wrapper >/dev/null <<'SH'
#!/bin/sh
case " $* " in
  *"/dev/nvme"*) exec /usr/sbin/smartctl -T permissive -d nvme "$@";;
  *)             exec /usr/sbin/smartctl -T permissive "$@";;
esac
SH
sudo chmod +x /opt/prometheus/exporters/smartctl/bin/smartctl-wrapper
  

The wrapper puts NVMe calls in permissive mode rather than spreading the workaround across compose arguments.

The compose definition is:

  # /opt/containers/compose/rootful/smartctl-exporter.yml
services:
  smartctl-exporter:
    image: docker.io/prometheuscommunity/smartctl-exporter:v0.14.0
    container_name: smartctl-exporter
    user: root
    privileged: true
    restart: unless-stopped
    ports:
      - "9633:9633"
    devices:
      - /dev/nvme0:/dev/nvme0
      - /dev/nvme1:/dev/nvme1
    volumes:
      - /opt/prometheus/exporters/smartctl/bin/smartctl-wrapper:/usr/local/bin/smartctl-wrapper:ro
    command:
      - --web.listen-address=:9633
      - --smartctl.path=/usr/local/bin/smartctl-wrapper
      - --smartctl.interval=60s
      - --smartctl.timeout=3s
      - --smartctl.retries=0
      - --smartctl.device=/dev/nvme0
      - --smartctl.device=/dev/nvme1
  

I set --smartctl.timeout=3s and --smartctl.retries=0 to avoid a stalled /metrics endpoint. An unresponsive device should fail quickly.

Start it and check the response:

  sudo systemctl restart smartctl-exporter.service
curl -s 127.0.0.1:9633/metrics | grep 'smartctl_device{device="nvme'
  

If nvme0 still blocks, restrict it to nvme1 first:

  # temporarily limit to nvme1
command:
  - --web.listen-address=:9633
  - --smartctl.path=/usr/local/bin/smartctl-wrapper
  - --smartctl.device=/dev/nvme1
  - --smartctl.interval=60s
  - --smartctl.timeout=3s
  - --smartctl.retries=0
  

Restore scraping before bringing the problem device back.

systemd (Rootless) Template

The mktxp.service example gives rootless compose stacks the same unit structure.

  # /opt/containers/systemd/rootless/mktxp.service
[Unit]
Description=Podman Compose Stack for mktxp
After=network-online.target
Wants=network-online.target

[Service]
Type=oneshot
RemainAfterExit=yes
WorkingDirectory=/opt/prometheus/exporters/mktxp/compose.d
ExecStart=/usr/bin/podman-compose up -d
ExecStop=/usr/bin/podman-compose down
TimeoutStopSec=30s

[Install]
WantedBy=default.target
  

Enable it in user scope:

  systemctl --user daemon-reload
systemctl --user enable --now /opt/containers/systemd/rootless/mktxp.service
loginctl enable-linger "$USER"
  

Set loginctl enable-linger for rootless services that must keep running beyond a login session.

Prometheus Scrape Configuration

Once smartctl-exporter responds on :9633, add this Prometheus job.

  scrape_configs:
  - job_name: "smartctl"
    scrape_interval: 60s
    scrape_timeout: 10s
    static_configs:
      - targets: ["storage-server:9633"]
  

As with node_exporter, central Prometheus pulls from each exporter. Check exporter availability separately from the scrape configuration.

Grafana Queries for nvme0 and nvme1

This query covers nvme0 and nvme1 without a fixed device filter.

  smartctl_device{device=~"nvme.*"}
# temperature example
smartctl_temperature_celsius{device=~"nvme.*"}
# health
smartctl_health_ok{device=~"nvme.*"}
  

A leftover device="nvme1" filter hides nvme0 after it is restored. Check the dashboard alongside the exporter.

Common failures and recovery commands

Commands for recovery:

  # /metrics hangs -> likely blocked on nvme0. Exclude it first and restore access.

# container stuck in Stopping
podman rm -f smartctl-exporter || true

# name collision
podman run --replace -d --name smartctl-exporter ...

# check whether 9633 is listening
ss -ltnp | grep :9633

# inspect device discovery and failures
podman logs smartctl-exporter | sed -n 's/.*Number of devices found.*/&/p'
podman logs smartctl-exporter | grep -E 'readjson|Invalid Log Page|device not found|Listening on'
  

For a container stuck in Stopping or a name collision, reset its state to restore monitoring.

How I Treat Dependencies Between Rootless and Rootful Services

Keep systemd dependencies inside one scope: system-to-system or user-to-user. I avoid strict dependencies across rootful / rootless scopes.

The rules are:

  • Keep dependencies loose and rely on network retries.
  • Use After=, Wants=, Requires=, and PartOf= only inside the same scope.
  • Prevent port conflicts up front with a fixed port table.

promtail -> loki and prometheus -> exporters recover through connection retries.

Making daemon-reload Hard to Forget

Run daemon-reload after unit or compose changes so systemd does not keep the old definition.

I added sdreload:

  # /usr/local/bin/sdreload
#!/bin/sh
systemctl daemon-reload || true
systemctl --user daemon-reload || true
echo "[done] daemon-reload (system & user)"
chmod +x /usr/local/bin/sdreload
  

Then reload, restart, and check status:

  # rootful (system scope)
sdreload
systemctl restart smartctl-exporter.service
systemctl status smartctl-exporter.service -n 30

# rootless (user scope)
sdreload
systemctl --user restart loki.service promtail.service
systemctl --user status loki.service promtail.service -n 30
  

Using the same sequence after every change prevents missed reloads.

Placement rules

Placement follows two rules:

  • Rootful: anything that touches the kernel or devices
  • Rootless: application-layer services

The services are:

  • Rootful: smartctl-exporter, node-exporter
  • Rootless: loki, prometheus, promtail, grafana

Check host-facing failures in system scope and application failures in user scope. Fixed placement also fixes the recovery starting point.

Final Operating Rules

The operating rules are:

  • Do not create hard dependencies across rootful and rootless scopes
  • Always run sdreload after touching service or compose definitions
  • Keep a fixed port design and avoid duplicates
  • Treat exporters as rootful and application services as rootless by default

These rules also apply as the node_exporter scrape targets grow.

Results

The work produced a consistent placement, restart, and dependency model alongside the smartctl-exporter fix.

Future Work

Next I will adapt existing *.service and compose/*.yml files to this pattern, reducing conflicting dependencies and ports.

For smartctl-exporter, I will stabilize nvme1 first, restore nvme0 as needed, then check that Prometheus and Grafana show both devices.