I ran IQ2XXS DeepSeek-V4-Flash on two DwarfStar4 nodes, separating orchestration and coding. This followed improvements with a single-model Step-3.7-Flash setup. The test examines role assignment in local agents for AI system development.

TG was 31.7-35.8t/s. Compared with Step-3.7-Flash and Qwen3.6-27B-EAGLE, it felt slow for long code output. Early one-shot tasks took 10-20 minutes; recent completed sessions can take 60-70 minutes even at TG100. Output quality felt good but tooling was unstable, so I am considering a lightweight worker with speculative decoding.

DwarfStar4/DeepSeek-V4-Flash at ctx 44k as orchestrator and Qwen3.6-27B at ctx 144k as worker felt reasonably useful. I plan to revisit DwarfStar as its single-model agent work progresses.

Evaluation Goal

Models can differ across orchestrator/planner/coder/tester/reviewer/integrator roles. The orchestrator decides and assigns the next work each turn.

For orchestration, stable prefix caching and ctx window rolling matter more than output volume. The aim is retaining meaning, structure, and cache hits over long context. TG matters less than for coding workers.

Configuration

ItemValue
ModelDeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf
RuntimeDwarfStar4 (ds4-server)
GPUNVIDIA RTX PRO 6000 Blackwell Max-Q 96GB x 2
Node APID 5435 / GPU0 / port 8001
Node BPID 5543 / GPU1 / port 8002
VRAM95.3 GiB / 95.6 GiB used per node
Host RAM~83.5 GB per node
Context98304
BalancingSpread across two nodes launched by the orchestration layer

Two Quadlet units differ only in --port and GPU assignment.

  [Unit]
Description=familiar orchestrator backend (DwarfStar 4 / DeepSeek V4 Flash) node-a
After=network-online.target
Wants=network-online.target

[Container]
ContainerName=ds4-node-a
Image=registry.home.arpa/dwarfstar4:ad0209f
Pull=always
Network=host
AddDevice=nvidia.com/gpu=0
Volume=/mnt/data/models/.../snapshots/<rev>:/model:ro
Volume=/mnt/data/models/.../kv_cache:/kv
Exec=-m /model/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
  --ctx 98304 --host 0.0.0.0 --port 8001 \
  --warm-weights --kv-disk-dir /kv/node-a --kv-disk-space-mb 4096 \
  --trace /kv/trace.node-a.log

[Service]
TimeoutStartSec=900
TimeoutStopSec=30
Restart=on-failure
RestartSec=30
  

nvtop showed one process per GPU, each using close to 96GB VRAM.

nvtop showing one DwarfStar4 ds4-server process on GPU0 and one on GPU1, with both GPUs nearly full on VRAM
Two DwarfStar4 processes, one node per GPU. During decode, there are moments where only one side rises to 99% utilization.

Measurements

Generation Throughput

Decode was 31.7-35.8t/s. Post-tool THINKING and longer replanning stayed around 35.7t/s; prefill was 227-241t/s.

The snapshot showed GPU1 decoding at 2302MHz/240W/99%, while GPU0 idled at 180MHz/13W. Sequential turns tend to use one node; parallel turns or workers can use both. Alternating GPUs felt useful in hot weather. My Broadcom 10GbE NIC remained warm with the link down. A cooler NIC, which I think was Realtek 8127, had unstable SSH at full bandwidth.

Grafana DCGM GPU Monitoring showing GPU1 at 100% utilization and near 300W while GPU0 is waiting at low load
GPU utilization in DCGM. Even with two nodes, a sequentially dependent turn can decode on only one GPU.

Long implementation output makes the wait per worker request noticeable.

Orchestration Result

One session across three turns:

MetricValue
orchestrator decisions3
planned worker5
dependency block0
worker failure rate33.3%
recovery2

The orchestrator planned five workers; one failed and replanning ran twice. replan-t2 failed, while replan-t3 was blocked by capacity and dependencies. DAG construction, dependency resolution, and replanning continued.

Familiar Orchestrator Decisions dashboard showing 3 orchestrator decisions, 5 planned workers, 33.3% worker failure, and 2 recovery signals
Familiar Orchestrator Decisions. Across three turns it planned five workers and handled failure and capacity block as replanning events.

Chat DAG showed requests branching into planner/coder and returning results.

Workspace showing aichat multi-agent logs, generated Django model files, git commits, and the Grafana Familiar Chat DAG side by side
Overall working view. On the left I checked the familiar conversation and generated commit; on the right I checked the Familiar Chat DAG showing branches from orchestrator to workers.

Tool Calls

MetricValue
total tool calls65
failed6
failure rate9.20%
distinct tool13

There were 30 read_file, 11 write_file, five directory_tree, and five shell_exec calls. Failure records were two writes, two directory trees, one shell call, and one other tool. Reads were roughly three times writes.

The counts include guard and logical false positives, not only invocation failures. Use them as call-volume context rather than a tool failure rate.

Familiar Tool Calls dashboard showing 65 total tool calls, 6 failures, 9.20% failure rate, and 13 distinct tools
Familiar Tool Calls. read_file dominates with 30 calls, while failures appear in write_file, directory_tree, and shell_exec.

Generated Output

The task implemented 25 Django domain models from a scaffold in apps/core/models.py. Commit a38864b contains one file with +861 -0.

The first model is RestaurantGroup:

  from django.db import models
from django.utils.translation import gettext_lazy as _


# ──────────────────────────────────────
# 1. RestaurantGroup
# ──────────────────────────────────────

class RestaurantGroup(models.Model):
    name = models.CharField(max_length=200, verbose_name=_("name"))
    slug = models.SlugField(max_length=200, unique=True, verbose_name=_("slug"))
    description = models.TextField(blank=True, default="", verbose_name=_("description"))
    website_url = models.URLField(blank=True, default="", verbose_name=_("website URL"))
    contact_email = models.EmailField(blank=True, default="", verbose_name=_("contact email"))
    contact_phone = models.CharField(max_length=50, blank=True, default="", verbose_name=_("contact phone"))
    is_active = models.BooleanField(default=True, verbose_name=_("active"))
    created_at = models.DateTimeField(auto_now_add=True, verbose_name=_("created at"))
    updated_at = models.DateTimeField(auto_now=True, verbose_name=_("updated at"))

    class Meta:
        verbose_name = _("restaurant group")
        verbose_name_plural = _("restaurant groups")
        ordering = ("name",)

    def __str__(self):
        return self.name
  

It includes gettext_lazy, blank=True with default="", Meta verbose_name/ordering, and __str__. The 25 models were broadly consistent. Qwen3.6 can produce similar output, so TG above 100 versus the 30s is also part of the choice.

Editor showing the generated Django apps/core/models.py RestaurantGroup model with fields, Meta, and __str__
The generated `RestaurantGroup` model. It is the first of 25 models and has the expected basic Django model elements.

Observed Failure Mode

The replanning planner had this failure independently of TG:

  invalid tool call returned as assistant text finish=stop
  [text_len=1864 saw_start=1 saw_end=1]
  

It treated replanning as requiring no new artifact declaration and nested wire-protocol XML inside write_file argument content. The harness detected saw_start=1 saw_end=1 but returned assistant text instead of a tool call.

  <...DSML...tool_calls>
<...DSML...invoke name="write_file">
<...DSML...parameter name="content" string="true"><familiar wire=....
    <thought>Re-planning pass: ... No deviations requiring corrective scope. ...</thought>
  

The parser could not resolve a top-level document nested inside a tool argument and fell back to text.

Possible contributors:

  • Model behavior: on replanning, it confuses the boundary between artifact declaration and tool call.
  • Parser strictness: nested structured markers are treated safely as text.

DeepSeek coverage is still limited, so XML influence is uncertain. With Qwen, XML protocols seemed to cause more mistakes. TOML reduced them but also lost XML familiarity. HTML is another proposal because it is common in training data.

Current Evaluation

Around 34t/s feels slow for long coding output. A shorter-output orchestrator or distillation-data source under its permissive license seems more useful.

The DAG, dependencies, replanning, and output worked in this run. IQ2XXS was usable; I also ran GLM-5.1 smol-IQ2KS for a long time. I have seen failures in sub-100B models at Q4/Q5 too. Calibration and PPL affect that judgment.

Replanning tool-call-as-text remains a parser/model issue. XML influence may contribute.

I retested recent commits, but Step-3.7-Flash still felt better for this familiar orchestrator.