Logo loFT LLC

    • Website renewal and technical articles
    • Absorption merger of Lorchestra Inc.
    • IT Introduction Support Provider selection for FY2022
    • IT Introduction Support Provider selection for FY2021
    • Lorchestra Inc. established as a subsidiary
    • IT Introduction Support Provider selection for FY2020
    • loFT LLC established
      • Separating post-processing with Dagster and NATS JetStream
      • Moving PostgreSQL to always-on storage
      • Adding MLflow and MinIO to Dagster
      • Vector migration and configuration for a three-host homelab
      • PostgreSQL 18 and pgvector on rootless Quadlet
      • Planning an EPYC 9175F workstation
      • Placing data on NVMe/SATA and moving management UIs
      • QuteBrowser settings for keyboard-based monitoring
      • CRS304 VLANs and Linux egress control
      • Saving RouterOS counters during Syslog downtime
      • Podman and Quadlet: UIDs and DNS
      • Monitoring with Prometheus, Loki, and Quadlet
      • smartctl-exporter and rootful / rootless operating rules
      • A local development platform with EPYC and Podman
      • Quadlet operations on minimal Ubuntu
      • ELT memory allocation on one EPYC server
      • Backup and restore design for rootless Podman
      • A Compose stack for vLLM, llama.cpp, and a proxy
      • Moving Hugging Face models between cold and hot storage
      • Moving an OpenAI-compatible proxy from Rust to Go
      • An AI pipeline with Go, NATS and Dagster
      • familiar: a local LLM development platform and its observation tools
      • Moving llm-jp translation to on-demand batches
      • Task splitting and output_file in familiar
      • A Gemma 4 stack on two Blackwell GPUs
      • RAG and event processing in agent-gateway
      • agent-gateway v3: domain separation and MLflow integration
      • Embedding and rerank APIs with Rust and ONNX
      • Designing a WordPress-like Django blog
      • Generating a Django booking site with Qwen3.5
      • Jev-Omni multimodal decisions and comparison with Clef
      • Running DeepSeek Harness locally
      • DeepSeek V4 Flash 0731: vLLM measurements and CPU KV offload
      • Comparing GLM-5.2 GGUF quantization, MTP, and expert placement
      • DeepSeek-V4-Flash on two DwarfStar4 nodes
      • Local parallel-agent development with Step-3.7-Flash
      • Comparing vLLM and SGLang for Gemma 4 31B
      • MiMo V2.5 Pro on one and two GPUs
      • DeepSeek V4 Flash Q2 on DwarfStar 4
      • Qwen3.6-27B NVFP4, MTP, and LoRA measurements
      • Local DeepSeek-V4-Flash inference with llama.cpp and dual Blackwell GPUs
      • Qwen3.6-27B inference and role-specific LoRA plans
      • Kimi-K2.6 CPU / GPU placement and code generation
      • NVFP4 translation and three-style Japanese data
      • GLM-5.1 expert placement and hybrid inference
      • Django implementation and inference speed with Qwen3.5-397B-A17B
      • MiniMax-M2.7 speed and noncommercial limits
      • GLM-5.1 expert placement and generation speed
      • Managing conversation branches and evaluation data in Dagster
      • Generating a six-page dental clinic site with Qwen3.5
      • A CPU/GPU role plan for local LLMs
      • Bilingual system prompts for PLAMO
      • Designing 36 LTX-2 scenes and clip continuity
      • Hermes-4.3-36B in BF16, FP8 and nvfp4
      • IQuest-Coder-40B on CPU, GPU and Aider
      • Command A Reasoning in an Aider test-generation loop
      • A coding-assistance proposal with Serena MCP and Obsidian
      • Review checks with GLM-4.7-Flash Uncensored
      • IQuest-Coder Loop-Instruct generation speed in aider
      • MCP execution hosts and VSCode Remote SSH
      • Kimi-K2.5 on EPYC 9175F and a revised L3 cache hypothesis
      • MiniMax-2.5 Expert Offload and website generation
      • Qwen3.5-397B hybrid inference and CMS generation
      • Llama-4-Scout CPU / GPU throughput and caching
      • Kimi-K2.5 CPU inference and prompt cache
      • Llama 4 Maverick Q4 and Q8 CPU comparison
      • Three configurations for Qwen3-Coder-Next 80B
      • GLM-4.7-Flash on CPU, hybrid and GPU
      • DeepSeek-V3.2 throughput and cache mismatches
      • Static-site and Django CMS generation with Qwen3.5-397B
      • shelpa path restrictions, audit mirrors, and retirement
      • shelpa-mcp virtual pipelines and CWD management
      • Integrating voracle research into development
      • Nine Rust MCP servers in the homelab
      • From shelpa to filesystem: file operations and recovery
      • Obsidian search and LLM conversation import with voracle
      • Comparing accuracy after changing pathfinder MCP responses
      • Why aichat function calling hung with a symlinked tools directory
      • Path resolution and MCP checks in pathfinder
      • AST analysis and MCP integration in ctree
  • Articles
  • Profile
  • Photos
    Logo
    Contact Us
      • Japanese
    • to navigate
    • to select
    • to close
      • Home
      • Tech Memo
      • Infrastructure
      On this page

      Infrastructure

      Server hardware, network topology, container orchestration, and monitoring stack documentation.

      These articles use AI-generated summaries of Obsidian notes originally kept as technical memos.
      English translations are produced with AI assistance.

      Separating post-processing with Dagster and NATS JetStream

      Dagster sensors process gateway events independently. NATS JetStream, idempotent storage, and …

      Moving PostgreSQL to always-on storage

      PostgreSQL moved from an on-demand GPU server to a 24/7 Mac Mini. This infrastructure setup retains …

      Adding MLflow and MinIO to Dagster

      An experiment tracking setup for AI system development using Dagster, MLflow, and MinIO. It covers …

      Vector migration and configuration for a three-host homelab

      Moving logs to Vector, splitting devstack by host, and consolidating Go defaults for local AI system …

      PostgreSQL 18 and pgvector on rootless Quadlet

      A Podman and Quadlet configuration for PostgreSQL 18, LLVM JIT and pgvector. It covers settings, …

      Planning an EPYC 9175F workstation

      A configuration review of HPCT WCE51-GP with EPYC 9175F and 768GB memory for LLM infrastructure. It …

      Placing data on NVMe/SATA and moving management UIs

      Query data and models on NVMe, logs on SATA, and management UIs on the Mac. An I/O and host-layout …

      QuteBrowser settings for keyboard-based monitoring

      QuteBrowser tab, keybinding, and color settings for checking Grafana, Prometheus, and router …

      CRS304 VLANs and Linux egress control

      A network infrastructure design combining local Mac/Linux 10GbE transfers and VDSL routing. It …

      Saving RouterOS counters during Syslog downtime

      A small-network infrastructure monitoring implementation with MikroTik DoH, VLANs, and Netwatch. It …

      Podman and Quadlet: UIDs and DNS

      A development infrastructure setup using rootless Podman and Quadlet on Linux and Podman Desktop on …

      Monitoring with Prometheus, Loki, and Quadlet

      Metrics and logs collected on a Mac mini for 3 hosts and RouterOS. An infrastructure setup using …

      smartctl-exporter and rootful / rootless operating rules

      An NVMe workaround for smartctl-exporter v0.14.0 and placement, reload, and dependency rules for …

      A local development platform with EPYC and Podman

      Three machines connected by 10GbE provide a local development platform. The design covers CPU …

      Quadlet operations on minimal Ubuntu

      An infrastructure plan for resident services on Ubuntu 24.04 minimized using systemd and Quadlet. It …

      ELT memory allocation on one EPYC server

      A proposed 512GB reusable ELT pool on a 768GB EPYC server. The data-platform design defines …

      Backup and restore design for rootless Podman

      An infrastructure recovery plan for compute-server after OS reinstallation. It covers tar.zst/rclone …

      A Compose stack for vLLM, llama.cpp, and a proxy

      Compose configuration for vLLM, llama.cpp, a Rust proxy, PostgreSQL, and Qdrant on rootless Podman. …

      Moving Hugging Face models between cold and hot storage

      Use rclone for GLM-5-GGUF blobs and rsync for snapshots and refs. The procedure preserves symlinks …

      © 2017-2026 loFT LLC

      Privacy Policy Security Whitepaper