Backup and restore design for rootless Podman
An infrastructure recovery plan for compute-server after OS reinstallation. It covers tar.zst/rclone backups, UID and namespace restoration, exclusions, retention, and a 20–30-minute recovery target.
Backups for compute-server
I defined a restore path for a rootless Podman compute server after an OS reinstall or NVMe reset. This infrastructure recovery design separates container data from the settings needed to run it.
The stack uses ext4/XFS + tar.zst + rclone, not ZFS. Preserve required files and regenerate reproducible components.
Background and Motivation
Postgres, Trino, and Dagster need more than /mnt/data. Rootless operation also needs user-systemd, subuid, subgid, and linger.
The design defines which files to save, which settings to restore, and which components to regenerate.
Use rclone copy for large blobs and rsync -aH for symlink-sensitive snapshots and refs.
Goal
The target is reconstructing rootless Podman and its data after an OS reinstall or NVMe reset.
- Keep the recovery path simple with
ext4/XFS + rclone + tar.zst - Preserve the prerequisites that matter for rootless Podman, including UID/GID,
subuid/subgid, andlinger - Separate backup targets from non-targets so the system stays understandable under failure
Directory Layout
Storage locations and recovery roles:
| Path | Role |
|---|---|
/mnt/data | Persistent application data for Podman apps such as Postgres, Trino, Dagster, etc. |
/opt/{app} | Configuration, compose files, and systemd definitions |
/usr/local | Local configuration for Trino, dbt, Dagster, and related tools |
/etc | System configuration such as networking, systemd, and Podman settings |
/home/ksh3 | Rootless Podman user environment and user-systemd settings |
/srv/backup | Backup staging area for compressed archives and package metadata |
Include /home/ksh3 and /etc so user-systemd and namespace settings can be restored.
Backup Policy
System Configuration Backup (Daily)
Daily targets are /usr/local, /etc, /home/ksh3, /srv/backup/pkglist.txt, and /etc/apt/sources.list.d/.
sudo dpkg --get-selections | awk '!/deinstall|purge/ {print $1}' \
> /srv/backup/pkglist.txt
tar --use-compress-program="zstd -T0 -19" \
-cf /srv/backup/system-$(date +%Y-%m-%d_%H-%M).tar.zst \
/etc /usr/local /home/ksh3 \
/srv/backup/pkglist.txt /etc/apt/sources.list.d \
--exclude='/home/ksh3/.cache' --exclude='/home/ksh3/.local/share/Trash'
pkglist.txt identifies missing packages. Exclude reproducible cache and Trash.
Use zstd -T0 -19 for the relatively bounded configuration archive, with parallel CPU compression.
Data Backup (Weekly or Manual)
Back up /mnt/data separately, weekly or on demand.
tar --use-compress-program="zstd -T0 -5" \
-cf /srv/backup/mnt-data-$(date +%Y-%m-%d).tar.zst -C /mnt data
Use -5 compression to keep large data backups within a practical duration.
Separate archives allow restoring host settings before application data.
Transfer Target (storage-server)
Send archives to storage-server.
rclone copy /srv/backup/*.tar.zst storage:/srv/backups/compute/ --progress
rclone copy transfers compressed files without needing live symlink semantics.
Exclude host logs because Prometheus and Loki retain them separately.
Recovery Procedure (After Reinstall)
Minimal Package Setup
Install the minimum tools to fetch, unpack, and start Podman.
sudo apt update
sudo apt install -y zstd rclone podman podman-compose
These provide decompression, transfer, and container startup.
Retrieve the Backups
Retrieve and extract both archives from storage-server.
rclone copy storage:/srv/backups/compute/latest/ /tmp/restore/
sudo tar -I zstd -xf /tmp/restore/system-YYYY-MM-DD.tar.zst -C /
sudo tar -I zstd -xf /tmp/restore/mnt-data-YYYY-MM-DD.tar.zst -C /
latest is the restore entry point. Its update method, generation handling, and retention policy still need definition.
Restore User and Permissions
Restore user identity, ownership, namespace settings, and linger.
sudo useradd -m -u 1000 -s /bin/bash ksh3
sudo chown -R ksh3:ksh3 /home/ksh3 /mnt/data /opt /usr/local
echo "ksh3:100000:65536" | sudo tee /etc/subuid /etc/subgid
loginctl enable-linger ksh3
Keep UID at 1000 to preserve data ownership and mappings.
Rebuild the Podman Environment
Once prerequisites are restored, rebuild the runtime.
sudo -u ksh3 podman system migrate
sudo -u ksh3 systemctl --user daemon-reload
sudo -u ksh3 systemctl --user enable --now pod-*.service
# Or:
cd /opt/containers/compose
for d in *; do [ -d "$d" ] && cd "$d" && podman-compose up -d && cd ..; done
Use user-systemd or podman-compose up -d. Both require matching UID/GID and restored namespace settings.
Additional Notes
Additional recovery rules:
| Item | Detail |
|---|---|
| Fixed UID/GID | Rootless Podman persistent data requires the same UID (1000) |
/etc/subuid / /etc/subgid | Required for the rootless namespace and must be backed up and restored |
loginctl enable-linger | Re-enables automatic user-systemd startup |
| Networking | No backup required because Podman recreates it during rebuild |
| SSH keys | Excluded for security and regenerated separately |
Regenerate networking rather than back it up.
Keep SSH keys outside this backup and restore them separately.
Backup targets and recovery conditions
The policy is:
- Backup targets:
/usr/local,/etc,/home/ksh3,/srv/backup/pkglist.txt,/etc/apt/sources.list.d, and/mnt/data - Format:
tar.zst + rclone - Sensitive key material stays excluded
- Recovery requires matching UID/GID, restored
subuid/subgid, and enabledlinger - After restoration, the environment comes back through
/opt/containers/composewithpodman-compose upor user-systemd activation
The recovery target is roughly 20 to 30 minutes, not a measured restore result.
Future Work
Remaining work:
- Define how
storage:/srv/backups/compute/latest/is maintained so retention and archive selection are deterministic - Connect this backup policy more explicitly to the separate
rclone/rsynctransport rules for symlink-heavy assets and large model trees - Add a post-restore smoke test so the Podman services can be validated automatically
Next, automate restoration, generation management, and post-restore checks.
