I ran DeepSeek V4 Flash 284B Q2-imatrix with antirez’s DwarfStar 4 (ds4.c) on one RTX PRO 6000 Blackwell Max-Q Workstation Edition 96GB. It reached 43 tok/s on short generations and the 31 tok/s range at 50K context.

This is an initial check of the alpha implementation as of 2026-05-14. It appeared to support head/tail-style role control for my agent orchestrator. My working range is 32K–64K, with 96K as a candidate. It started at 128K with little headroom, so I currently treat 96K as the limit. Future optimization may change this.

Video link: https://www.youtube.com/watch?v=A4aGNHEdrxE

Measured results

ItemValue
ModelDeepSeek V4 Flash
Parameters284B MoE / 13B active
RuntimeDwarfStar 4 (ds4.c)
Buildcuda-generic
QuantIQ2_XXS + Q2_K routed expert, Q8 attention/shared/output
GGUFDeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix
GPUNVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
GPU memory96GB GDDR7 ECC
Memory bandwidth1,792 GB/s
Model tensor cache80.76 GiB
Peak observed VRAM93,142 MiB / 95,593 MiB
Short generation43.6 tok/s
50K context generation31.4 tok/s
20K prefill262 tok/s

Previously I ran DeepSeek V4 Flash on dual RTX PRO 6000 GPUs using a WIP DeepSeek-V4 branch of llama.cpp. This time I used the model-specific ds4.c runtime and loaded an 80GiB-class Q2-imatrix GGUF onto a single RTX PRO 6000 Blackwell Max-Q 96GB.

Tool-call generation stayed above 31 tok/s at 50K context, and multi-turn agent sessions completed.

DwarfStar 4 Design

DwarfStar 4 is not a generic GGUF runner. It is a native inference engine narrowed specifically to DeepSeek V4 Flash. The official README explains that model loading, prompt rendering, tool calling, RAM/on-disk KV state, and the server API are all implemented directly for DeepSeek V4 Flash.

The project validates logits, long context and agent integration for one model. The engine, quantization, API and tool-call format are tested together.

At the same time, the project is explicitly alpha quality right now. The CUDA backend also has room to improve. The numbers below should be read as a first look as of 2026-05-14, not as the ceiling of a finished implementation.

Test Environment

ItemSpecification
GPUNVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
GPU architectureSM_120 (Blackwell)
VRAM96GB GDDR7 ECC
Memory bandwidth1,792 GB/s
CPUAMD EPYC 9175F
RAM755 GiB
StorageNVMe 3.5T (xfs)
OSUbuntu 24.04
CUDA13.2.1
ContainerPodman rootless
ModelDeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix
EngineDwarfStar 4 (cuda-generic build)

The RTX PRO 6000 Blackwell Max-Q has 96GB of GDDR7 ECC and 1,792 GB/s of memory bandwidth. MoE decode for a model like DeepSeek V4 Flash is strongly affected by memory bandwidth, so this is a good platform for testing how far a single GPU can go.

Build And Startup

The CUDA build worked with this target.

  make cuda-generic
  

For the grandpa runtime, I built a thin container image from the CUDA 13.2.1 devel image that simply clones ds4 and builds cuda-generic.

  FROM docker.io/nvidia/cuda:13.2.1-cudnn-devel-ubuntu24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
    git make gcc g++ ca-certificates && \
    rm -rf /var/lib/apt/lists/*

WORKDIR /app
RUN git clone --depth 1 https://github.com/antirez/ds4.git . && \
    make cuda-generic -j$(nproc)

EXPOSE 8000
ENTRYPOINT ["./ds4-server"]
  

Build and push looked like this.

  podman build -t registry.home.arpa/dwarfstar4:latest .
podman push registry.home.arpa/dwarfstar4:latest
  

On Blackwell SM_120, -arch=native selected the right architecture. The runtime environment is a Podman rootless container using NVIDIA CDI passthrough and host networking.

The startup log showed CUDA backend initialization, device loading of the 80GiB-class tensor cache, and model loading from NVMe.

  ds4: CUDA backend initialized on NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition (sm_120)
ds4: CUDA loading model tensors into device cache: 80.04 GiB
ds4: CUDA startup model cache prepared 80.76 GiB of tensor spans in 16.198s
  

Model-cache preparation from NVMe took about 16.2 seconds.

Q2-imatrix Quantization

The GGUF I used applies aggressive asymmetric quantization only to the routed experts in DeepSeek V4 Flash.

Tensor classQuant
routed expert up/gateIQ2_XXS
routed expert downQ2_K
shared expertsQ8_0
attention projectionsQ8_0
output headQ8_0
router / embedding / auxiliary blocksF16 / F32

In an MoE model, routed experts account for most of the model size, while each token only passes through a subset of the experts. This quantization compresses the routed experts aggressively, while keeping higher precision for quality-sensitive parts such as the router, attention projections, shared experts, and output head. The result is a 284B model that fits into the 80GiB range while still aiming to avoid collapse in coding-agent workloads.

Generation Throughput

Generation throughput was 43.6 tok/s on short text and 31.4 tok/s at 50K context.

Context / generationGeneration speedNotes
Short prompt, about 100 tokens43.6 tok/sThinking mode
Medium, about 400 generated tokens41.7 tok/sThinking mode
Long, 4,058 generated tokensavg 38.5 tok/s / min 37.3 tok/sThinking to Text
20K context36.3 tok/sTool calling enabled
33K context35.3 tok/sStable
50K context31.4 tok/sVery stable

TG decreased with longer context but remained in the 31 tok/s range at 50K.

For this workload, I would use 32K–64K. Although the design supports contexts such as 256K, agent operation also needs reusable prefixes and session management.

Prefill

Prefill proceeds in 2048-token chunks. Even at deeper context, it stayed above 245 tok/s.

Token countPrefill speed
13K tokens267 tok/s
20K tokens262 tok/s
30K+ tokens incremental245-251 tok/s

Prefill at 20K context was 262 tok/s. My tuned GLM-5.1 setup is around TG23 and PP700, so prefill remains a concern for orchestration.

An MTP-related PR is available, but I have not seen much benefit from it yet.

VRAM Usage

VRAM is almost fully consumed.

  GPU MEM: 93,142 MiB / 95,593 MiB (95%)
  

The model tensor cache occupies 80.76 GiB, plus context buffers and compute pools. The run uses one GPU without CPU offload.

Model, runtime and context fit on one GPU and expose an OpenAI-compatible API to Zed or my agent. This reduces infrastructure size for local LLM integration compared with dual-GPU or hybrid inference.

Disk KV Cache

Disk KV cache saves and reuses DeepSeek V4 Flash’s compressed KV state on NVMe, somewhat like LMCache. A fast external M.2 drive is another proposed option; its heat-management benefit has not been tested.

In my log, KV save on eviction completed in about 200-370ms.

  kv cache stored tokens=3405  size=67.56 MiB  save=198.9 ms  reason=evict
kv cache stored tokens=54163 size=734.07 MiB save=371.9 ms  reason=evict
  

I confirmed that the cold, continued, and evict trigger patterns worked. Prefix reuse also showed up in measurement.

RunResponse time
First run0.148 s
Second run0.073 s

The second response took almost half the time. For a coding agent that repeatedly reads a large system prompt or repository context, this prefix reuse directly affects TTFT.

This was also the biggest design caveat in my environment. My own agent has a context manager built around multiple sessions, and it does not compose cleanly with DwarfStar 4’s KV reuse as-is. I added a DwarfStar 4-specific context manager and adapter so the orchestration session can absorb the single-session assumption.

Comparison With Other Platforms

This is not a strict apples-to-apples comparison, but when I line up public information and values I observed locally, decode performance moves in the same direction as memory bandwidth.

PlatformMemory bandwidthGenerationNotes
DGX Spark GB10273 GB/s10-14 tok/sReference value
Mac Studio / Apple Siliconconfiguration-dependent16-36 tok/sReference value
RTX PRO 6000 Blackwell Max-Q1,792 GB/s43.6 tok/sMy measurement

MoE decode has to read expert weights every token, so memory bandwidth matters a lot. The RTX PRO 6000 Blackwell Max-Q has 1,792 GB/s of GDDR7 bandwidth, and a single GPU reached the 43 tok/s range. A non-Max-Q RTX variant may be able to push PP a little further.

The remaining optimization areas include:

  • CUDA backend is still alpha-stage
  • MTP support

Coding-Agent Operation

I connected Zed’s AI Assistant to the OpenAI-compatible API of ds4-server and ran real coding tasks through it.

DwarfStar 4 natively supports the DeepSeek V4 Flash DSML tool format, and it was able to execute tool calls such as these autonomously.

  • roots_list
  • directory_tree
  • read_file
  • terminal
  • spawn_agent

read_text_file and list_directory were frequent, and ctree and pathfinder worked. Through terminal, the agent repeatedly used grep -c on _contract/architecture.md. This helps local checks but leaves room to use ctree / pathfinder for more efficient structural queries.

Grafana chart showing familiar tool-call counts and failures by tool name
Tool Calls by Name. read_text_file and list_directory dominate, ctree/pathfinder also appear, and repeated terminal grep -c calls show up as a habit.

At 50K context, tool calls stayed above 31 tok/s and multi-turn sessions completed. I adjusted the official sampling values for orchestration.

The behavior felt similar to Codex and Claude, but this check is too limited to establish replacement quality.

I began integration checks in my agent. This is a first look at speed and context, with broader quality evaluation still needed.

Grafana dashboard showing familiar orchestrator decision events and worker outcomes
familiar orchestration dashboard. antirez/deepseek-v4-gguf is used as the orchestrator model while decision events, worker outcomes, and the depends_on scheduling graph are observed.

Settings for the orchestrator

  • model=deepseek-chat can select non-thinking mode
  • --warm-weights can reduce first-inference stutter
  • --dir-steering-file and --dir-steering-ffn -1 can apply verbosity steering
  • -n 4096 can constrain the default output budget
  • disk KV cache can preserve prefix reuse across a long orchestration session

For code-review workers, keeping thinking mode and longer context makes sense. For the orchestrator, model=deepseek-chat fixed to non-thinking mode and focused on dispatch is the more natural role split.

The help text explicitly documents the non-thinking selection conditions.

  thinking={type:disabled}, think=false, or model=deepseek-chat selects non-thinking mode.
  

Main ds4-server parameters:

AreaParameterMy read
Model-m, --modelPoints to a ds4-specific GGUF. This is not a generic GGUF runner
MTP--mtp, --mtp-draft, --mtp-marginSpeculative decoding. Still experimental, and lower priority for a resident grandpa
Context-c, --ctxContext allocated at startup. 32768 is enough for grandpa, while frisky has room to stretch to 131072
Output-n, --tokensDefault when max_tokens is omitted. For grandpa, constrain it to around 4096
CPU-t, --threadsHelper threads for tokenization and prompt rendering. Leaving it unset is fine
Quality--qualityUses stricter kernels. Useful for quality evaluation or benchmarks, usually unnecessary for resident grandpa
Steering--dir-steering-file, --dir-steering-ffn, --dir-steering-attnFor grandpa, apply -1 on the FFN side to amplify the succinct direction. Attention-side steering is experimental
Warmup--warm-weightsWorth enabling for a resident service. Reduces first-inference page faults and stutter
Backend--cuda, --metal, --cpu, --backendCUDA for RTX PRO 6000. CPU is for diagnostics
API--host, --portSplit roles by port, such as grandpa=8000 and frisky=8001
Trace--traceSaves prompts, cache decisions, outputs, and tool calls in human-readable form. May be useful for fulfillment logs or DPO data
Thinkingreasoning_effort, thinking, think, modelgrandpa should use model=deepseek-chat for non-thinking mode. reviewer / frisky can benefit from thinking mode
Disk KV--kv-disk-dir, --kv-disk-space-mb, --kv-cache-*4GB likely suffices for grandpa. Workers with long review history may want 16GB
Tools--disable-exact-dsml-tool-replay, --tool-memory-max-idsMostly irrelevant for grandpa when it does not use tool calling

The proposed resident grandpa settings are --ctx 65536, --tokens 4096, 32GB disk KV and --trace. A tool-less dispatch role could use --ctx 32768 and 4GB disk KV.

  [Container]
ContainerName=grandpa
Image=registry.home.arpa/dwarfstar4:latest
Pull=always
Network=host
AddDevice=nvidia.com/gpu=0
Volume=/mnt/data/models/models--antirez--deepseek-v4-gguf/snapshots/c566ab6d7c696ddd0c7f124e115228af1a326824:/model:ro
Volume=/mnt/data/models/models--antirez--deepseek-v4-gguf/kv_cache:/kv
Exec=-m /model/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf --ctx 65536 --tokens 4096 --host 0.0.0.0 --port 8000 --warm-weights --kv-disk-dir /kv --kv-disk-space-mb 32768 --trace /kv/trace.log
  

To verify the same settings once by hand, podman run --rm -it works directly.

  podman run --rm -it \
  --device nvidia.com/gpu=0 \
  -v /mnt/data/models/models--antirez--deepseek-v4-gguf/snapshots/c566ab6d7c696ddd0c7f124e115228af1a326824:/model:ro,Z \
  -v /mnt/data/models/models--antirez--deepseek-v4-gguf/kv_cache:/kv:Z \
  --network host \
  registry.home.arpa/dwarfstar4:latest \
  -m /model/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
  --ctx 65536 \
  --tokens 4096 \
  --host 0.0.0.0 \
  --port 8000 \
  --warm-weights \
  --kv-disk-dir /kv \
  --kv-disk-space-mb 32768 \
  --trace /kv/trace.log
  
Screen showing grandpa Quadlet container definition alongside DwarfStar 4 startup logs
grandpa container definition and ds4 startup log. In verification, ctx=131072 / disk KV 32GB also started successfully, with context buffers at 2425.71 MiB and model cache at 80.76 GiB. For resident grandpa, I would start with ctx=65536 / tokens=4096 / disk KV 32GB / trace enabled.

When using steering, generate the vector under dir-steering/ in the ds4 repository, mount it into the container, and pass it through --dir-steering-file. verbosity.f32 does not exist inside the container image by default, so generating it on the host and mounting it is the practical route.

  cd ~/src/ds4

python3 dir-steering/tools/build_direction.py \
  --ds4 ./ds4 \
  --model ds4flash.gguf \
  --good-file dir-steering/examples/succinct.txt \
  --bad-file dir-steering/examples/verbose.txt \
  --out dir-steering/out/verbosity.json \
  --component ffn_out \
  --ctx 512

ls -lh dir-steering/out/verbosity.f32
  

ds4-server has a different cache strategy from the cache pool in ik_llama.cpp, which reuses partial cache through f_keep. ds4 keeps only one live KV cache in VRAM. When switching sessions, it evicts the current KV to disk, checks whether the next request prefix hits the disk cache, and if it does, reads that cache back as the live session.

Directional Steering

  y = y - scale * direction[layer] * dot(direction[layer], y)
  

The bundled verbosity vector has 43 layers × 4096 dimensions and is about 704KB. The scale changes output verbosity at runtime.

ScaleBehavior
-1Compresses output length to about half
2Strengthens detailed explanations

A proposed role split uses -1 for short orchestrator decisions and 2 for detailed reviewer output, without fine-tuning.

Summary

ItemValue
ModelDeepSeek V4 Flash 284B
RuntimeDwarfStar 4
GPURTX PRO 6000 Blackwell Max-Q 96GB x 1
QuantQ2-imatrix
Short generation43.6 tok/s
50K context31.4 tok/s
Prefill245-267 tok/s
Model cache80.76 GiB
Peak VRAM93.1 GiB

The run confirmed 43 tok/s on short output and the 31 tok/s range at 50K context on one GPU. EPYC 9175F’s 512MB L3 cache may contribute, but that has not been isolated.

DwarfStar 4 combines the engine, quantization, disk KV cache, tool calling and directional steering for DeepSeek V4 Flash.

It worked in my agent, although responses were rougher than my more thoroughly tuned GLM-5.1 setup. Optimization and quality checks are still beginning.

I plan to follow CUDA, disk KV and steering improvements. AMD MI350P PCIe Cards are also linked as a possible hardware candidate; their suitability has not been tested.