Architecture change

I replaced familiar’s own agent with DeepSeek Harness.

I removed the agent from the gateway/agent/infrastructure monolith and placed DeepSeek Harness (dsh) in a container Pod with full access inside that Pod.

I split familiar into two parts, familiar-daemon and agent-gateway. ancestor, which manages the inference backend, already existed, so I kept using it as is. These three run dsh inside a Pod.

The migration took about a week after Codex Astra’s release and also served as a performance check.

The previous article covers the initial design. This article covers the replacement and its execution records.

Videos

I recorded videos of dsh running with full access inside a Pod, with inference served by vLLM on the compute host.

The first one runs deepseek-ai/DeepSeek-V4-Flash-Vision-Exp from the official checkpoint as is and has it implement a Django domain layer. The settings and numbers are in “The Vision-Exp run” later in this article.

Video link: https://www.youtube.com/watch?v=KD-4jm46izk

The second one was recorded separately with the recently released DeepSeek-V4.1-Flash in EXL3 2.0bpw. The model is the quantized model published by diffbot.

Video link: https://www.youtube.com/watch?v=CBkIkCtIxJI

Functions in the previous familiar

The previous familiar monolith included these functions:

  • Compaction
  • Memory
  • Summarization
  • Per-role memory management
  • Operating the knowledge platform
  • Custom tools
  • MCP. Since it was written in Go, I even embedded entire MCP sources with embed
  • Observability into OpenTelemetry, Vector, Prometheus, and Grafana
  • What-if and replay built with Grafana and GraphQL
  • A Dagster assets pipeline
  • A container mechanism for verifying shell execution
  • Tool recovery

Most development added tools and execution management to compensate for local LLM failures.

Why I chose dsh

I chose dsh because its plugin design allowed functions to be added separately.

At the time of writing, existing plugins sometimes covered a feature I had planned to implement.

Agent presets opened in the dsh Settings screen
Agent presets in Settings. Standard, PTC, Minimal, and Creator modes are listed alongside presets such as Default, Coding, Planning, and Testing.

Plugins also replaced my Grafana/GraphQL what-if and replay features. ThoughtDAG displays and edits LLM context as a graph and lets runs branch. I use it in desktop dsh’s DAG tab beside Chat.

ThoughtDAG Session Atlas listing dsh-local sessions
ThoughtDAG's Session Atlas. It lists local sessions by project folder and imports the selected one as a canvas. Nothing done on the canvas is written back to the original conversation file.
ThoughtDAG DAG view branching a follow-up question from the Django session
The DAG view. The Django session is imported as a node, and a question about building the frontend with Tailwind CSS branches from it. Whatever is connected by edges becomes the context for the next question.

Using externally maintained plugins reduces the functions I need to build.

Three components

I retained and ported the infrastructure functions, split into these roles.

ComponentImplementationRoleRelation to the old familiar
familiar-daemon (famd)GoAccepting requests (delegations) to dsh, isolated workspaces, durable state, recovery, result collection, and Pod managementSplit out of familiar
agent-gatewayGoInference gateway for dsh and OpenAI-compatible clients. It discovers running Ancestor instances and routes by public model IDSplit out of familiar
ancestorRustManages inference servers on the compute host. It builds configurations from presets and starts or stops them with Podman Quadlet and systemd user servicesExisted before; used as is

famd owns agent execution and publishing its own operational events to NATS, and nothing beyond that. It does not own an OpenAI-compatible inference endpoint or the capture of prompts, tool calls, and responses. Those belong to agent-gateway, which publishes captures to NATS asynchronously. Observability such as OpenTelemetry was also implemented by reusing agent-gateway’s events.

Starting and stopping inference servers is ancestor’s job. Submitting a request to famd does not start or stop an inference server on its own.

Current architecture

flowchart LR
  subgraph desktop["desktop.home.arpa"]
    CLI["famdctl"]
    UI["dsh Web UI"]
  end
  subgraph compute["compute.home.arpa"]
    FAMD["familiar-daemon<br/>control / launcher / finalizer"]
    GW["agent-gateway"]
    ANC["ancestor"]
    VLLM["vLLM"]
    RELAY["famd-egress<br/>inference relay"]
    subgraph pod["runtime Pod"]
      DSH["DeepSeek Harness<br/>Full access"]
      WEB["Web relay"]
    end
  end
  NATS[("NATS")]
  CLI --> FAMD
  CLI --> GW
  CLI --> ANC
  FAMD --> pod
  ANC --> VLLM
  DSH --> RELAY --> GW --> VLLM
  GW -. discover .-> ANC
  FAMD -. events .-> NATS
  GW -. capture .-> NATS
  UI -. "SSH tunnel<br/>pod forward" .-> WEB

On compute.home.arpa, famd’s control, launcher, finalizer, and related services run as native services, and so does agent-gateway.

famd-egress is the inference entry point that Pods can see. dsh inside a Pod sends inference requests to http://famd-egress:8080/v1, and famd-egress relays them to agent-gateway. It is a Quadlet container with a read-only root filesystem and all capabilities dropped.

I submit a delegation from desktop famdctl. Acceptance creates a runtime Pod with dsh and a Web relay. dsh receives full access inside the Pod; the reasons and risks are described under “About full access.”

Using famdctl

famdctl has the following command set.

  ksh3@desktop.home.arpa ~ % famdctl
famdctl [--json] [--config path] [--idempotency-key key] <command>

  ancestor    List, inspect, plan, ensure or stop Ancestor services
  delegation  List, submit, inspect, wait for, retrieve or cancel a delegation
  session     List, start, recover, observe or close a DSH session
  events      Stream lifecycle events, resuming with --after
  artifact    Inspect or download a verified artifact
  pod         List, start, stop, remove, drain or forward runtime Pods
  dsh-local   Prepare, sync and open the local DSH Web UI
  workbench   Alias for dsh-local open
  workspace   Initialize or locate the local work records and DSH mirror
  gateway     List public models
  config      Initialize or inspect saved settings
  doctor      Check configuration; --online probes supported service APIs
  completion  Print a bash or zsh completion script

Use famdctl <command> --help for details. Output is formatted for reading.
Add --json anywhere before -- for raw JSON responses and JSONL events.
Exit status: 0 success/waiting, 1 service/local/terminal failure, 2 usage error.
Root submit/run, status, wait, result and cancel are delegation shortcuts.
Save endpoints once with config init, then use alias fam=famdctl.
  

ancestor operates ancestor, gateway operates agent-gateway, and delegation, session, events, artifact, and pod operate famd. Each of the three services has its own endpoint and credentials.

Here is an example of submitting the Django domain implementation task. With --detach, only the acceptance receipt comes back.

  ksh3@desktop.home.arpa ~/familiar-workspace % famdctl delegation submit --workflow django --detach /Users/ksh3/familiar-workspace/task/task.json
Work record: /Users/ksh3/familiar-workspace/workflow/django/2026-09-15/18-03-25
STATUS    DELEGATION           SESSION              ACCEPTED
accepted  del_W9hVQjDgA2zzXQ   ags_NPoXgKU5IGYmAg   2026-09-15 18:04:17
Receipt only; execution is not confirmed.
Next: famdctl delegation wait del_W9hVQjDgA2zzXQ
  

To see dsh inside the Pod, I open a tunnel through compute.home.arpa with pod forward and view it in a browser.

  ksh3@desktop.home.arpa ~/familiar-workspace % famdctl pod forward del_W9hVQjDgA2zzXQ
http://127.0.0.1:13080/?token=<redacted>
Pod pod_EaTd58jSdfE2kA is forwarded through compute.home.arpa; press Ctrl-C to close the tunnel.
  

The Web UI has Chat and Trajectory tabs, where you can see the To-dos work plan and the files in /workspace inside the Pod. Sessions that ran in a Pod are also synced to ~/familiar-workspace/dsh-local/Remote/ on the desktop and can be opened from the desktop dsh as Remote / artifact_….

Chat view of the Django domain implementation session in the desktop dsh
The desktop dsh. Sessions synced from Pods are listed in Workspaces on the left as Remote / artifact_….
A failed gate mid step opened in the dsh Trajectory tab
The Trajectory tab. The Input, Model, and Tools timeline is at the top, the list of tool calls on the left, and the details of the selected step on the right. At Turn 1 Step 33, gate mid failed with model_defect; after that, core/models.py was fixed and the gate passed.

Here is podman ps on compute-server when neither vLLM nor a Pod is running.

  ksh3@compute-server:~$ podman ps
CONTAINER ID  IMAGE                                                                                                            COMMAND               CREATED         STATUS         PORTS       NAMES
77af06a3024b  registry.home.arpa/ml-foundry-mlflow@sha256:e0916a7bce42adc92f5b688e212ece81c0b1ce363e92eaf02b1118161f23cf3b     mlflow server --h...  5 hours ago     Up 5 hours                 mlflow
2144fc65532d  registry.home.arpa/ml-foundry@sha256:df3d06efff1570a47085c8002d980d66085c01c69b5d07d5afab98c4f482c888            dagster api grpc ...  5 hours ago     Up 5 hours                 dagster-user-code
381a9063ea00  registry.home.arpa/ml-foundry@sha256:df3d06efff1570a47085c8002d980d66085c01c69b5d07d5afab98c4f482c888            dagster-webserver...  5 hours ago     Up 5 hours                 dagster-webserver
9df064d0e85b  registry.home.arpa/ml-foundry@sha256:df3d06efff1570a47085c8002d980d66085c01c69b5d07d5afab98c4f482c888            dagster-daemon ru...  5 hours ago     Up 5 hours                 dagster-daemon
c9849a9fd9a5  registry.home.arpa/famd-inference-relay@sha256:702149bdc55611ef25ba40d98c0181d55e9d786507ad5b991e17ccbc96bd4ed4  --upstream http:/...  17 minutes ago  Up 17 minutes              famd-egress
  

famd-egress is the inference relay. famd’s control and agent-gateway are native services, so they don’t show up here.

Dagster and MLflow are ml-foundry. It is the former model-foundry, also rebuilt to fit agent-gateway, and it now has an ML pipeline for my hobby investing in addition to models.

Grafana dashboard showing results from the ml-foundry investment ML pipeline
Grafana for the investment ML pipeline I'm building in ml-foundry. I check the same-day scenario forecast by candidate stock and time slot.

During the Vision-Exp run, the vLLM container ancestor-deepseek-v4-flash-vision-exp started by ancestor and the Pod’s dsh runner and dsh web were also running. The two Pod containers use the same famd-dsh-runner image: --managed is dsh itself, and --proxy-only is the Web relay. The Pod publishes only one port on 127.0.0.1, and the Web UI connects to it through the tunnel.

The Vision-Exp run

In the first video, compute-host vLLM serves DeepSeek-V4-Flash-Vision-Exp, while dsh in a Podman Pod implements a Django domain layer with full access. Model requests stay on the host. This is an application-development example for local LLM integration.

Configuration

  MODEL: deepseek-ai/DeepSeek-V4-Flash-Vision-Exp @ 6821d6ad3681a4b137b066b76094fa82ebd0a380
IMAGE: registry.home.arpa/voipmonitor/vllm:ds4-jovian-r9-5bea08859798
HOST:  compute.home.arpa:8000 (engine) <- pod forward <- desktop.home.arpa (client)
POD:   dsh runner, dsh web
GPU:   NVIDIA RTX PRO 6000 Blackwell Max-Q ×2 (GPU 0,1 / 300W cap)
  
ItemSetting
Context1M (max-model-len 1048576), max output 393,216 tokens
WeightsOfficial checkpoint as is. FP8 (e4m3, 128×128 blocks), FP4 experts, bfloat16 for the rest. No additional quantization
ParallelismTP2 (tensor-parallel-size 2), no NVLink
Speculative decodingDSpark fixed-probabilistic K3 (3 tokens)
KV cacheGPU KV only, no LMCache
GPU memorygpu-memory-utilization 0.968
Concurrent requests4 (max-num-seqs 4)
Batchmax-num-batched-tokens 4096
Reasoning efforthigh
OtherTool calling enabled, vision enabled, OpenAI-compatible chat completions API

This is the ExecStart of the vLLM service. The env file is read from ancestor’s active directory.

  ExecStart=/usr/bin/podman run --name=ancestor-deepseek-v4-flash-vision-exp --cidfile=%t/%N.cid --replace --rm --cgroups=split --network=host --init --sdnotify=conmon -d --ulimit stack=67108864:67108864 --device=nvidia.com/gpu=0 --device=nvidia.com/gpu=1 -v /mnt/data/hf/hub/models--deepseek-ai--DeepSeek-V4-Flash-Vision-Exp:/models:ro -v /mnt/data/hf/jit/ds4-vision-exp-jovian-r9:/cache -v /mnt/data/hf/tmp/ds4-vision-exp-jovian-r9:/container-tmp --env-file /home/ksh3/.local/state/ancestor/active/ancestor-deepseek-v4-flash-vision-exp.env --pull never --entrypoint=/usr/local/bin/lmcache-mp-wrapper.sh --ipc=host registry.home.arpa/voipmonitor/vllm:ds4-jovian-r9-5bea08859798 /usr/local/bin/serve-ds4-flash.sh
  

The serving profile is based on the Vision spec of DeepSeek-V4-Flash Jovian Judgement r9 published in local-inference-lab/rtx6kpro. The only exception is gpu_memory_utilization, which I set to 0.968 instead of either documented value (0.975 for GPU KV only, 0.970 with LMCache).

The entrypoint goes through the LMCache wrapper (lmcache-mp-wrapper.sh), but LMCACHE_MODE is not set. So the host tier is not used, and this is a GPU-KV-only run.

I wrote about measuring DeepSeek V4 Flash 0731 with DSpark K5 on the same two RTX PRO 6000 GPUs in another article. The model and settings are different, so I don’t put the numbers side by side.

Engine log numbers

Across a roughly 7-minute window of the engine log, there were 70 chat completion requests.

MetricValue
Decode throughput (average)180 tok/s (peak 257.9 tok/s)
Prefill throughput (average during prompt-processing windows)444 tok/s
DSpark draft acceptance53,246 / 72,615 tokens (73.3%), mean accepted length 3.20 at depth 3
Prefix cache hit rate87.5% → 98.0%
Decode-only windows (average)216.6 tok/s
Windows with prefill mixed in (average)156.3 tok/s

The prefix cache hit rate rose from 87.5% to 98.0% over the window. It dipped three times along the way when new prefixes came in.

Comparing decode-only windows with windows that include prefill, chunked prefill takes roughly 28% off decode on the same engine.

The dsh status bar in the video shows 201 tok/s and a 93% cache hit. Those are the client’s own numbers, counted over a different span with a different denominator, so they are not expected to match the engine counters above.

Caveats on the numbers

All numbers are averages of 10-second windows from the engine log. I did not capture TTFT, ITL, or per-request latency. Token totals are approximations from “rate × 10 seconds”, not exact counters.

Also, the published r9 numbers were measured on 600W Workstation cards. These are 300W Max-Q cards, so the numbers are not directly comparable.

About full access

Excessive guards in the previous familiar caused repeated execution failures and reduced output quality. I changed complex shell pipelines to container-side dry-runs and removed the guards. That experience informed the full-access configuration inside this Pod.

At the time of writing, I use ds4, glm5.3-flash and qwen38-flash-next for regular work. More OSS models are usable for these tasks than around March.

That said, running without restrictions leaves every risk of destructive operations in place. If you try it, isolate the environment and take a snapshot first.

Thoughts

About six months of agent development exposed model-selection issues, runtime failures and MCP implementation needs. Alongside OSS reuse, I built custom mechanisms with AI coding.

I plan to rely more on dsh plugins and community tools to reduce custom implementation.

The migration replaces the agent while retaining the infrastructure for inference and execution management.