DeepSeek-V4-Flash on two DwarfStar4 nodes
A local AI development test of IQ2XXS DeepSeek-V4-Flash on two GPUs. It covers throughput, DAG and replanning behavior, Django model output, and a tool extraction failure.
I ran IQ2XXS DeepSeek-V4-Flash on two DwarfStar4 nodes, separating orchestration and coding. This followed improvements with a single-model Step-3.7-Flash setup. The test examines role assignment in local agents for AI system development.
TG was 31.7-35.8t/s. Compared with Step-3.7-Flash and Qwen3.6-27B-EAGLE, it felt slow for long code output. Early one-shot tasks took 10-20 minutes; recent completed sessions can take 60-70 minutes even at TG100. Output quality felt good but tooling was unstable, so I am considering a lightweight worker with speculative decoding.
DwarfStar4/DeepSeek-V4-Flash at ctx 44k as orchestrator and Qwen3.6-27B at ctx 144k as worker felt reasonably useful. I plan to revisit DwarfStar as its single-model agent work progresses.
Evaluation Goal
Models can differ across orchestrator/planner/coder/tester/reviewer/integrator roles. The orchestrator decides and assigns the next work each turn.
For orchestration, stable prefix caching and ctx window rolling matter more than output volume. The aim is retaining meaning, structure, and cache hits over long context. TG matters less than for coding workers.
Configuration
| Item | Value |
|---|---|
| Model | DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf |
| Runtime | DwarfStar4 (ds4-server) |
| GPU | NVIDIA RTX PRO 6000 Blackwell Max-Q 96GB x 2 |
| Node A | PID 5435 / GPU0 / port 8001 |
| Node B | PID 5543 / GPU1 / port 8002 |
| VRAM | 95.3 GiB / 95.6 GiB used per node |
| Host RAM | ~83.5 GB per node |
| Context | 98304 |
| Balancing | Spread across two nodes launched by the orchestration layer |
Two Quadlet units differ only in --port and GPU assignment.
[Unit]
Description=familiar orchestrator backend (DwarfStar 4 / DeepSeek V4 Flash) node-a
After=network-online.target
Wants=network-online.target
[Container]
ContainerName=ds4-node-a
Image=registry.home.arpa/dwarfstar4:ad0209f
Pull=always
Network=host
AddDevice=nvidia.com/gpu=0
Volume=/mnt/data/models/.../snapshots/<rev>:/model:ro
Volume=/mnt/data/models/.../kv_cache:/kv
Exec=-m /model/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
--ctx 98304 --host 0.0.0.0 --port 8001 \
--warm-weights --kv-disk-dir /kv/node-a --kv-disk-space-mb 4096 \
--trace /kv/trace.node-a.log
[Service]
TimeoutStartSec=900
TimeoutStopSec=30
Restart=on-failure
RestartSec=30
nvtop showed one process per GPU, each using close to 96GB VRAM.

Measurements
Generation Throughput
Decode was 31.7-35.8t/s. Post-tool THINKING and longer replanning stayed around 35.7t/s; prefill was 227-241t/s.
The snapshot showed GPU1 decoding at 2302MHz/240W/99%, while GPU0 idled at 180MHz/13W. Sequential turns tend to use one node; parallel turns or workers can use both. Alternating GPUs felt useful in hot weather. My Broadcom 10GbE NIC remained warm with the link down. A cooler NIC, which I think was Realtek 8127, had unstable SSH at full bandwidth.

Long implementation output makes the wait per worker request noticeable.
Orchestration Result
One session across three turns:
| Metric | Value |
|---|---|
| orchestrator decisions | 3 |
| planned worker | 5 |
| dependency block | 0 |
| worker failure rate | 33.3% |
| recovery | 2 |
The orchestrator planned five workers; one failed and replanning ran twice. replan-t2 failed, while replan-t3 was blocked by capacity and dependencies. DAG construction, dependency resolution, and replanning continued.

Chat DAG showed requests branching into planner/coder and returning results.

Tool Calls
| Metric | Value |
|---|---|
| total tool calls | 65 |
| failed | 6 |
| failure rate | 9.20% |
| distinct tool | 13 |
There were 30 read_file, 11 write_file, five directory_tree, and five shell_exec calls. Failure records were two writes, two directory trees, one shell call, and one other tool. Reads were roughly three times writes.
The counts include guard and logical false positives, not only invocation failures. Use them as call-volume context rather than a tool failure rate.

Generated Output
The task implemented 25 Django domain models from a scaffold in apps/core/models.py. Commit a38864b contains one file with +861 -0.
The first model is RestaurantGroup:
from django.db import models
from django.utils.translation import gettext_lazy as _
# ──────────────────────────────────────
# 1. RestaurantGroup
# ──────────────────────────────────────
class RestaurantGroup(models.Model):
name = models.CharField(max_length=200, verbose_name=_("name"))
slug = models.SlugField(max_length=200, unique=True, verbose_name=_("slug"))
description = models.TextField(blank=True, default="", verbose_name=_("description"))
website_url = models.URLField(blank=True, default="", verbose_name=_("website URL"))
contact_email = models.EmailField(blank=True, default="", verbose_name=_("contact email"))
contact_phone = models.CharField(max_length=50, blank=True, default="", verbose_name=_("contact phone"))
is_active = models.BooleanField(default=True, verbose_name=_("active"))
created_at = models.DateTimeField(auto_now_add=True, verbose_name=_("created at"))
updated_at = models.DateTimeField(auto_now=True, verbose_name=_("updated at"))
class Meta:
verbose_name = _("restaurant group")
verbose_name_plural = _("restaurant groups")
ordering = ("name",)
def __str__(self):
return self.name
It includes gettext_lazy, blank=True with default="", Meta verbose_name/ordering, and __str__. The 25 models were broadly consistent. Qwen3.6 can produce similar output, so TG above 100 versus the 30s is also part of the choice.

Observed Failure Mode
The replanning planner had this failure independently of TG:
invalid tool call returned as assistant text finish=stop
[text_len=1864 saw_start=1 saw_end=1]
It treated replanning as requiring no new artifact declaration and nested wire-protocol XML inside write_file argument content. The harness detected saw_start=1 saw_end=1 but returned assistant text instead of a tool call.
<...DSML...tool_calls>
<...DSML...invoke name="write_file">
<...DSML...parameter name="content" string="true"><familiar wire=....
<thought>Re-planning pass: ... No deviations requiring corrective scope. ...</thought>
The parser could not resolve a top-level document nested inside a tool argument and fell back to text.
Possible contributors:
- Model behavior: on replanning, it confuses the boundary between artifact declaration and tool call.
- Parser strictness: nested structured markers are treated safely as text.
DeepSeek coverage is still limited, so XML influence is uncertain. With Qwen, XML protocols seemed to cause more mistakes. TOML reduced them but also lost XML familiarity. HTML is another proposal because it is common in training data.
Current Evaluation
Around 34t/s feels slow for long coding output. A shorter-output orchestrator or distillation-data source under its permissive license seems more useful.
The DAG, dependencies, replanning, and output worked in this run. IQ2XXS was usable; I also ran GLM-5.1 smol-IQ2KS for a long time. I have seen failures in sub-100B models at Q4/Q5 too. Calibration and PPL affect that judgment.
Replanning tool-call-as-text remains a parser/model issue. XML influence may contribute.
I retested recent commits, but Step-3.7-Flash still felt better for this familiar orchestrator.
