I used familiar to generate a six-page dental clinic site and investigate a failed run. Claude was the orchestrator, Qwen3-Coder-Next 80B IQ4_KSS was naughty, and GLM-5.1 smol-IQ4_K was grandpa. The orchestrator can switch between claude | codex | gemini | local LLM.

The worker backend unavailable message came from putting the entire specification into one worker call and reaching GPUTimeout=300s. I checked the workspace, chat_history, Grafana session flow and DCGM GPU monitoring together.

Validation Setup

The run tested three things.

  1. Split a large specification after defining the shared contract.
  2. Trace the responsibilities and output of naughty and grandpa.
  3. Distinguish a model response problem from an execution plan problem using records.

The production stack also includes a translator, but the scope of this validation was limited to orchestrator / grandpa / naughty. The older kindergarten setup has already been retired.

The role layout was as follows.

RolePrimary responsibilityModel
orchestratorTurn control, worker-call planning, dependency orderingClaude orchestrator
naughtyHTML generation, direct single-file generationQwen3-Coder-Next 80B IQ4_KSS x 2 instance
grandpaPartial review, validation, file-read-based checkingGLM-5.1 smol-IQ4_K
translatorTranslationPlamo-2-translate bf16

I do not intend to keep Plamo-2-translate for long-term use because of licensing constraints. I am currently replacing it with llm-jp/llm-jp-4-32b-a3b-base, adapted for translation with a LoRA and then quantized to nvfp4.

The target task was a six-page static dental clinic site. The workload required index.html, services.html, doctors.html, info.html, visit.html, and access.html, plus a shared header and footer, pricing tables, a modal, and a README.

Initial error

This was the first signal I saw.

  Status: worker execution failed after one recovery attempt.
Initial error: worker call w1 failed: worker backend unavailable
  (alias=naughty backend=vllm model=naughty): knowledge gate: llm backend unavailable:
  request vllm: Post "http://compute.home.arpa:8001/v1/chat/completions":
  context deadline exceeded (Client.Timeout exceeded while awaiting headers)
  

I first suspected backend capacity or connectivity. The message alone did not establish that naughty was down.

The workspace showed a plan to generate all six pages in one worker call. It already contained the parallelism information for naughty, but the system prompt lacked concrete decomposition examples.

Inspecting the familiar workspace while tracing the state of the orchestrator and workers
I checked the workspace and the real files side by side to see which turn generated what and where the run stalled

Generated pages and history

A page with the pricing-table modal rendered, and the matching HTML was present in info.html. Generation had progressed but stopped before completion.

The BrightSmile Dental pricing page in a browser beside the corresponding HTML in the editor
The output was already materially valid. The problem was not generation quality itself but the amount of work packed into a single turn

Looking at chat_history also helped. I could see the turn-by-turn exchanges and usage records for the orchestrator, grandpa, and other workers. It also showed that the review phase had already been split into two passes and that each worker left a usage JSON payload behind.

Inspecting orchestrator and grandpa review history in the chat_history table
`chat_history` preserves the review round-trips between the orchestrator and `grandpa`. It also made it obvious that smaller review slices were more stable

Records used to investigate

I compared four views: the workspace, database, session flow and GPUs.

Grafana showed roughly 50 turns per request over several minutes. A wrong query condition produced the 7.26M tokens display; after correction, the count was about 442k. Chat flow color-codes naughty, grandpa and claude to show each role’s duration.

The Familiar Session Flow Dashboard showing requests, turns, total tokens, and chat flow
The `Familiar Session Flow Dashboard`. At this point the token query was wrong and showed 7.26M, but after fixing it the actual total was about 442k

Vector session flow showed branches from plan nodes to workers and transitions between running and queued in raw telemetry. These records let me inspect the orchestrator’s DAG.

The Vector Session Flow dashboard showing the session-flow graph and raw telemetry
The session flow data is emitted to Loki as a visualization-focused DTO. I used it to inspect plan branching and status transitions

DCGM showed utilization and FB usage on both GPUs. The backend had not fully stopped.

Checking GPU utilization and memory utilization on the DCGM GPU Monitoring dashboard
DCGM showed that the GPUs were actually working. The issue was not a backend outage but a plan that could not finish within 300 seconds

Root Cause

The failure followed this sequence.

  1. The orchestrator was still running on a V1-style system prompt.
  2. That prompt did not provide enough concrete contract-first parallel design examples.
  3. The six-page requirement was pushed into a single oversized worker call.
  4. naughty could not finish within 300 seconds and hit GPUTimeout.
  5. Recovery failed because it retried the same plan.

The worker calls I actually wanted looked more like this.

  {
  "continue": true,
  "worker_calls": [
    {
      "id": "n1",
      "alias": "naughty",
      "prompt": "Create the shared structure and landing page first.",
      "depends_on": [],
      "output_file": "index.html"
    },
    {
      "id": "g1",
      "alias": "grandpa",
      "prompt": "Review the shared contract and validate accessibility risks.",
      "depends_on": ["n1"]
    },
    {
      "id": "n2",
      "alias": "naughty",
      "prompt": "Create services.html using the shared header/footer.",
      "depends_on": ["n1"],
      "output_file": "services.html"
    }
  ]
}
  

The intended plan creates the shared skeleton first, then splits pages around that contract. grandpa handles review and validation.

Implementation changes

I made three changes.

System prompts from the worker pool

I introduced BuildOrchestratorSystemPrompt(pool) to choose role descriptions and decomposition patterns based on enabled naughty, grandpa and translator workers. I also preserved versions of final and partial synthesis, their diffs, adopted results and the turns that used them.

Single-file generation with output_file

I added output_file mode to write worker raw text directly to a specified path. It lets one HTML file be generated in one LLM call and avoids repeated ReAct tool calls.

A lighter setup with --seq=2 and --parallel 2 may need less coordination. My Qwen3-Coder-Next 80B IQ4_KSS backend runs two instances, so workers need a shared contract and separate file responsibilities to avoid interference.

I also added MCP tools for shared state and interference avoidance. output_file reduces generation overhead; the tools control what multiple instances can access.

Trace records in orchestra.log

orchestra.log records the system prompt, worker prompt, worker content, tool calls, tool results and duration_ms together. This makes it possible to investigate why a worker call was issued.

Validation Outcome

The run supported the following division of work.

  • Improve the orchestrator’s task decomposition and message-stack handling.
  • Use naughty=Qwen3-Coder-Next 80B IQ4_KSS x 2 instance for single-file generation.
  • Use grandpa=GLM-5.1 smol-IQ4_K for review, validation and smaller checks.
  • Compare execution plans with actual outputs through session flow and the database.

Session flow and the database separated the plan from each role’s actual output. They provide evidence for investigating backend latency, prompts and turn design.

In this run, instruction chains, messages[N] operations and intermediate output handoffs affected execution alongside model capability.

Over roughly one month, I developed nearly 20 Rust MCP tools, including retired tools, for local LLM and general use. Most reduce wasted ctx. Compared with the minimax-m2.7 validation, context for one-shot site generation fell from about 150k to 55k.

At the time of writing, I prioritized tracing failures by turn in Grafana and the database over automatic SFT or DPO data collection.

Next checks

The next checks cover orchestrator switching, worker responsibilities, message-stack compression and saved telemetry.

For LLM integration and AI application development, small generation tasks and saved inputs, outputs and durations help identify where a run failed.