smartctl-exporter and rootful / rootless operating rules
An NVMe workaround for smartctl-exporter v0.14.0 and placement, reload, and dependency rules for monitoring infrastructure maintenance.
Introduction
I stabilized smartctl-exporter by separating rootful exporters from rootless applications and standardizing the reload procedure. This note covers the NVMe workaround and the monitoring operating rules.
The earlier node_exporter rollout established central scraping from Storage. Next I addressed smartctl-exporter, especially nvme0, and the mixed rootful / rootless service responsibilities.
Background and Motivation
The node_exporter rollout added host metrics to central Prometheus on Storage. Scraping worked, but exporter placement was still inconsistent.
Two problems recurred: unclear dependencies across rootful / rootless scopes, and missed daemon-reload commands after unit or compose edits. smartctl-exporter v0.14.0 also lacked --smartctl.device-opts, so NVMe needed a workaround.
For infrastructure maintenance, I standardized placement, dependencies, and the commands to run after a change.
Directory Standard
I used the following directory layout for units and compose files.
# rootful (administrator-managed)
sudo mkdir -p /opt/containers/{compose,systemd}/{rootful,rootless}
sudo mkdir -p /opt/prometheus/exporters/smartctl/bin
Device-facing services such as smartctl-exporter are rootful. Applications such as prometheus and loki are rootless. Both use the same directory structure.
systemd (Rootful)
I used a rootful oneshot unit for smartctl-exporter, matching the existing one-service-per-stack pattern.
# /etc/systemd/system/smartctl-exporter.service
[Unit]
Description=smartctl-exporter (rootful)
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
RemainAfterExit=yes
WorkingDirectory=/opt/containers/compose/rootful
ExecStart=/usr/bin/podman compose up -d
ExecStop=/usr/bin/podman compose down
TimeoutStopSec=30s
[Install]
WantedBy=multi-user.target
Enable it with:
sudo systemctl daemon-reload
sudo systemctl enable --now smartctl-exporter.service
Exporters that need devices or elevated permissions belong in system scope, giving maintenance a consistent starting point.
Compose (Rootful / smartctl-exporter v0.14.0)
smartctl-exporter v0.14.0 lacks --smartctl.device-opts, so the wrapper adds -T permissive.
sudo tee /opt/prometheus/exporters/smartctl/bin/smartctl-wrapper >/dev/null <<'SH'
#!/bin/sh
case " $* " in
*"/dev/nvme"*) exec /usr/sbin/smartctl -T permissive -d nvme "$@";;
*) exec /usr/sbin/smartctl -T permissive "$@";;
esac
SH
sudo chmod +x /opt/prometheus/exporters/smartctl/bin/smartctl-wrapper
The wrapper puts NVMe calls in permissive mode rather than spreading the workaround across compose arguments.
The compose definition is:
# /opt/containers/compose/rootful/smartctl-exporter.yml
services:
smartctl-exporter:
image: docker.io/prometheuscommunity/smartctl-exporter:v0.14.0
container_name: smartctl-exporter
user: root
privileged: true
restart: unless-stopped
ports:
- "9633:9633"
devices:
- /dev/nvme0:/dev/nvme0
- /dev/nvme1:/dev/nvme1
volumes:
- /opt/prometheus/exporters/smartctl/bin/smartctl-wrapper:/usr/local/bin/smartctl-wrapper:ro
command:
- --web.listen-address=:9633
- --smartctl.path=/usr/local/bin/smartctl-wrapper
- --smartctl.interval=60s
- --smartctl.timeout=3s
- --smartctl.retries=0
- --smartctl.device=/dev/nvme0
- --smartctl.device=/dev/nvme1
I set --smartctl.timeout=3s and --smartctl.retries=0 to avoid a stalled /metrics endpoint. An unresponsive device should fail quickly.
Start it and check the response:
sudo systemctl restart smartctl-exporter.service
curl -s 127.0.0.1:9633/metrics | grep 'smartctl_device{device="nvme'
If nvme0 still blocks, restrict it to nvme1 first:
# temporarily limit to nvme1
command:
- --web.listen-address=:9633
- --smartctl.path=/usr/local/bin/smartctl-wrapper
- --smartctl.device=/dev/nvme1
- --smartctl.interval=60s
- --smartctl.timeout=3s
- --smartctl.retries=0
Restore scraping before bringing the problem device back.
systemd (Rootless) Template
The mktxp.service example gives rootless compose stacks the same unit structure.
# /opt/containers/systemd/rootless/mktxp.service
[Unit]
Description=Podman Compose Stack for mktxp
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
RemainAfterExit=yes
WorkingDirectory=/opt/prometheus/exporters/mktxp/compose.d
ExecStart=/usr/bin/podman-compose up -d
ExecStop=/usr/bin/podman-compose down
TimeoutStopSec=30s
[Install]
WantedBy=default.target
Enable it in user scope:
systemctl --user daemon-reload
systemctl --user enable --now /opt/containers/systemd/rootless/mktxp.service
loginctl enable-linger "$USER"
Set loginctl enable-linger for rootless services that must keep running beyond a login session.
Prometheus Scrape Configuration
Once smartctl-exporter responds on :9633, add this Prometheus job.
scrape_configs:
- job_name: "smartctl"
scrape_interval: 60s
scrape_timeout: 10s
static_configs:
- targets: ["storage-server:9633"]
As with node_exporter, central Prometheus pulls from each exporter. Check exporter availability separately from the scrape configuration.
Grafana Queries for nvme0 and nvme1
This query covers nvme0 and nvme1 without a fixed device filter.
smartctl_device{device=~"nvme.*"}
# temperature example
smartctl_temperature_celsius{device=~"nvme.*"}
# health
smartctl_health_ok{device=~"nvme.*"}
A leftover device="nvme1" filter hides nvme0 after it is restored. Check the dashboard alongside the exporter.
Common failures and recovery commands
Commands for recovery:
# /metrics hangs -> likely blocked on nvme0. Exclude it first and restore access.
# container stuck in Stopping
podman rm -f smartctl-exporter || true
# name collision
podman run --replace -d --name smartctl-exporter ...
# check whether 9633 is listening
ss -ltnp | grep :9633
# inspect device discovery and failures
podman logs smartctl-exporter | sed -n 's/.*Number of devices found.*/&/p'
podman logs smartctl-exporter | grep -E 'readjson|Invalid Log Page|device not found|Listening on'
For a container stuck in Stopping or a name collision, reset its state to restore monitoring.
How I Treat Dependencies Between Rootless and Rootful Services
Keep systemd dependencies inside one scope: system-to-system or user-to-user. I avoid strict dependencies across rootful / rootless scopes.
The rules are:
- Keep dependencies loose and rely on network retries.
- Use
After=,Wants=,Requires=, andPartOf=only inside the same scope. - Prevent port conflicts up front with a fixed port table.
promtail -> loki and prometheus -> exporters recover through connection retries.
Making daemon-reload Hard to Forget
Run daemon-reload after unit or compose changes so systemd does not keep the old definition.
I added sdreload:
# /usr/local/bin/sdreload
#!/bin/sh
systemctl daemon-reload || true
systemctl --user daemon-reload || true
echo "[done] daemon-reload (system & user)"
chmod +x /usr/local/bin/sdreload
Then reload, restart, and check status:
# rootful (system scope)
sdreload
systemctl restart smartctl-exporter.service
systemctl status smartctl-exporter.service -n 30
# rootless (user scope)
sdreload
systemctl --user restart loki.service promtail.service
systemctl --user status loki.service promtail.service -n 30
Using the same sequence after every change prevents missed reloads.
Placement rules
Placement follows two rules:
- Rootful: anything that touches the kernel or devices
- Rootless: application-layer services
The services are:
- Rootful:
smartctl-exporter,node-exporter - Rootless:
loki,prometheus,promtail,grafana
Check host-facing failures in system scope and application failures in user scope. Fixed placement also fixes the recovery starting point.
Final Operating Rules
The operating rules are:
- Do not create hard dependencies across rootful and rootless scopes
- Always run
sdreloadafter touching service or compose definitions - Keep a fixed port design and avoid duplicates
- Treat exporters as rootful and application services as rootless by default
These rules also apply as the node_exporter scrape targets grow.
Results
The work produced a consistent placement, restart, and dependency model alongside the smartctl-exporter fix.
Future Work
Next I will adapt existing *.service and compose/*.yml files to this pattern, reducing conflicting dependencies and ports.
For smartctl-exporter, I will stabilize nvme1 first, restore nvme0 as needed, then check that Prometheus and Grafana show both devices.
