DeepSeek V4 Flash Q2 on DwarfStar 4
DeepSeek V4 Flash 284B Q2-imatrix ran on one Blackwell 96GB GPU at 43.6 tok/s on short text and 31.4 tok/s at 50K context. The local LLM integration check covers KV cache and agent execution.
I ran DeepSeek V4 Flash 284B Q2-imatrix with antirez’s DwarfStar 4 (ds4.c) on one RTX PRO 6000 Blackwell Max-Q Workstation Edition 96GB. It reached 43 tok/s on short generations and the 31 tok/s range at 50K context.
This is an initial check of the alpha implementation as of 2026-05-14. It appeared to support head/tail-style role control for my agent orchestrator. My working range is 32K–64K, with 96K as a candidate. It started at 128K with little headroom, so I currently treat 96K as the limit. Future optimization may change this.
Video link: https://www.youtube.com/watch?v=A4aGNHEdrxE
Measured results
| Item | Value |
|---|---|
| Model | DeepSeek V4 Flash |
| Parameters | 284B MoE / 13B active |
| Runtime | DwarfStar 4 (ds4.c) |
| Build | cuda-generic |
| Quant | IQ2_XXS + Q2_K routed expert, Q8 attention/shared/output |
| GGUF | DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix |
| GPU | NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition |
| GPU memory | 96GB GDDR7 ECC |
| Memory bandwidth | 1,792 GB/s |
| Model tensor cache | 80.76 GiB |
| Peak observed VRAM | 93,142 MiB / 95,593 MiB |
| Short generation | 43.6 tok/s |
| 50K context generation | 31.4 tok/s |
| 20K prefill | 262 tok/s |
Previously I ran DeepSeek V4 Flash on dual RTX PRO 6000 GPUs using a WIP DeepSeek-V4 branch of llama.cpp. This time I used the model-specific ds4.c runtime and loaded an 80GiB-class Q2-imatrix GGUF onto a single RTX PRO 6000 Blackwell Max-Q 96GB.
Tool-call generation stayed above 31 tok/s at 50K context, and multi-turn agent sessions completed.
DwarfStar 4 Design
DwarfStar 4 is not a generic GGUF runner. It is a native inference engine narrowed specifically to DeepSeek V4 Flash. The official README explains that model loading, prompt rendering, tool calling, RAM/on-disk KV state, and the server API are all implemented directly for DeepSeek V4 Flash.
The project validates logits, long context and agent integration for one model. The engine, quantization, API and tool-call format are tested together.
At the same time, the project is explicitly alpha quality right now. The CUDA backend also has room to improve. The numbers below should be read as a first look as of 2026-05-14, not as the ceiling of a finished implementation.
Test Environment
| Item | Specification |
|---|---|
| GPU | NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition |
| GPU architecture | SM_120 (Blackwell) |
| VRAM | 96GB GDDR7 ECC |
| Memory bandwidth | 1,792 GB/s |
| CPU | AMD EPYC 9175F |
| RAM | 755 GiB |
| Storage | NVMe 3.5T (xfs) |
| OS | Ubuntu 24.04 |
| CUDA | 13.2.1 |
| Container | Podman rootless |
| Model | DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix |
| Engine | DwarfStar 4 (cuda-generic build) |
The RTX PRO 6000 Blackwell Max-Q has 96GB of GDDR7 ECC and 1,792 GB/s of memory bandwidth. MoE decode for a model like DeepSeek V4 Flash is strongly affected by memory bandwidth, so this is a good platform for testing how far a single GPU can go.
Build And Startup
The CUDA build worked with this target.
make cuda-generic
For the grandpa runtime, I built a thin container image from the CUDA 13.2.1 devel image that simply clones ds4 and builds cuda-generic.
FROM docker.io/nvidia/cuda:13.2.1-cudnn-devel-ubuntu24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git make gcc g++ ca-certificates && \
rm -rf /var/lib/apt/lists/*
WORKDIR /app
RUN git clone --depth 1 https://github.com/antirez/ds4.git . && \
make cuda-generic -j$(nproc)
EXPOSE 8000
ENTRYPOINT ["./ds4-server"]
Build and push looked like this.
podman build -t registry.home.arpa/dwarfstar4:latest .
podman push registry.home.arpa/dwarfstar4:latest
On Blackwell SM_120, -arch=native selected the right architecture. The runtime environment is a Podman rootless container using NVIDIA CDI passthrough and host networking.
The startup log showed CUDA backend initialization, device loading of the 80GiB-class tensor cache, and model loading from NVMe.
ds4: CUDA backend initialized on NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition (sm_120)
ds4: CUDA loading model tensors into device cache: 80.04 GiB
ds4: CUDA startup model cache prepared 80.76 GiB of tensor spans in 16.198s
Model-cache preparation from NVMe took about 16.2 seconds.
Q2-imatrix Quantization
The GGUF I used applies aggressive asymmetric quantization only to the routed experts in DeepSeek V4 Flash.
| Tensor class | Quant |
|---|---|
| routed expert up/gate | IQ2_XXS |
| routed expert down | Q2_K |
| shared experts | Q8_0 |
| attention projections | Q8_0 |
| output head | Q8_0 |
| router / embedding / auxiliary blocks | F16 / F32 |
In an MoE model, routed experts account for most of the model size, while each token only passes through a subset of the experts. This quantization compresses the routed experts aggressively, while keeping higher precision for quality-sensitive parts such as the router, attention projections, shared experts, and output head. The result is a 284B model that fits into the 80GiB range while still aiming to avoid collapse in coding-agent workloads.
Generation Throughput
Generation throughput was 43.6 tok/s on short text and 31.4 tok/s at 50K context.
| Context / generation | Generation speed | Notes |
|---|---|---|
| Short prompt, about 100 tokens | 43.6 tok/s | Thinking mode |
| Medium, about 400 generated tokens | 41.7 tok/s | Thinking mode |
| Long, 4,058 generated tokens | avg 38.5 tok/s / min 37.3 tok/s | Thinking to Text |
| 20K context | 36.3 tok/s | Tool calling enabled |
| 33K context | 35.3 tok/s | Stable |
| 50K context | 31.4 tok/s | Very stable |
TG decreased with longer context but remained in the 31 tok/s range at 50K.
For this workload, I would use 32K–64K. Although the design supports contexts such as 256K, agent operation also needs reusable prefixes and session management.
Prefill
Prefill proceeds in 2048-token chunks. Even at deeper context, it stayed above 245 tok/s.
| Token count | Prefill speed |
|---|---|
| 13K tokens | 267 tok/s |
| 20K tokens | 262 tok/s |
| 30K+ tokens incremental | 245-251 tok/s |
Prefill at 20K context was 262 tok/s. My tuned GLM-5.1 setup is around TG23 and PP700, so prefill remains a concern for orchestration.
An MTP-related PR is available, but I have not seen much benefit from it yet.
VRAM Usage
VRAM is almost fully consumed.
GPU MEM: 93,142 MiB / 95,593 MiB (95%)
The model tensor cache occupies 80.76 GiB, plus context buffers and compute pools. The run uses one GPU without CPU offload.
Model, runtime and context fit on one GPU and expose an OpenAI-compatible API to Zed or my agent. This reduces infrastructure size for local LLM integration compared with dual-GPU or hybrid inference.
Disk KV Cache
Disk KV cache saves and reuses DeepSeek V4 Flash’s compressed KV state on NVMe, somewhat like LMCache. A fast external M.2 drive is another proposed option; its heat-management benefit has not been tested.
In my log, KV save on eviction completed in about 200-370ms.
kv cache stored tokens=3405 size=67.56 MiB save=198.9 ms reason=evict
kv cache stored tokens=54163 size=734.07 MiB save=371.9 ms reason=evict
I confirmed that the cold, continued, and evict trigger patterns worked. Prefix reuse also showed up in measurement.
| Run | Response time |
|---|---|
| First run | 0.148 s |
| Second run | 0.073 s |
The second response took almost half the time. For a coding agent that repeatedly reads a large system prompt or repository context, this prefix reuse directly affects TTFT.
This was also the biggest design caveat in my environment. My own agent has a context manager built around multiple sessions, and it does not compose cleanly with DwarfStar 4’s KV reuse as-is. I added a DwarfStar 4-specific context manager and adapter so the orchestration session can absorb the single-session assumption.
Comparison With Other Platforms
This is not a strict apples-to-apples comparison, but when I line up public information and values I observed locally, decode performance moves in the same direction as memory bandwidth.
| Platform | Memory bandwidth | Generation | Notes |
|---|---|---|---|
| DGX Spark GB10 | 273 GB/s | 10-14 tok/s | Reference value |
| Mac Studio / Apple Silicon | configuration-dependent | 16-36 tok/s | Reference value |
| RTX PRO 6000 Blackwell Max-Q | 1,792 GB/s | 43.6 tok/s | My measurement |
MoE decode has to read expert weights every token, so memory bandwidth matters a lot. The RTX PRO 6000 Blackwell Max-Q has 1,792 GB/s of GDDR7 bandwidth, and a single GPU reached the 43 tok/s range. A non-Max-Q RTX variant may be able to push PP a little further.
The remaining optimization areas include:
- CUDA backend is still alpha-stage
- MTP support
Coding-Agent Operation
I connected Zed’s AI Assistant to the OpenAI-compatible API of ds4-server and ran real coding tasks through it.
DwarfStar 4 natively supports the DeepSeek V4 Flash DSML tool format, and it was able to execute tool calls such as these autonomously.
roots_listdirectory_treeread_fileterminalspawn_agent
read_text_file and list_directory were frequent, and ctree and pathfinder worked. Through terminal, the agent repeatedly used grep -c on _contract/architecture.md. This helps local checks but leaves room to use ctree / pathfinder for more efficient structural queries.

At 50K context, tool calls stayed above 31 tok/s and multi-turn sessions completed. I adjusted the official sampling values for orchestration.
The behavior felt similar to Codex and Claude, but this check is too limited to establish replacement quality.
I began integration checks in my agent. This is a first look at speed and context, with broader quality evaluation still needed.

Settings for the orchestrator
model=deepseek-chatcan select non-thinking mode--warm-weightscan reduce first-inference stutter--dir-steering-fileand--dir-steering-ffn -1can apply verbosity steering-n 4096can constrain the default output budget- disk KV cache can preserve prefix reuse across a long orchestration session
For code-review workers, keeping thinking mode and longer context makes sense. For the orchestrator, model=deepseek-chat fixed to non-thinking mode and focused on dispatch is the more natural role split.
The help text explicitly documents the non-thinking selection conditions.
thinking={type:disabled}, think=false, or model=deepseek-chat selects non-thinking mode.
Main ds4-server parameters:
| Area | Parameter | My read |
|---|---|---|
| Model | -m, --model | Points to a ds4-specific GGUF. This is not a generic GGUF runner |
| MTP | --mtp, --mtp-draft, --mtp-margin | Speculative decoding. Still experimental, and lower priority for a resident grandpa |
| Context | -c, --ctx | Context allocated at startup. 32768 is enough for grandpa, while frisky has room to stretch to 131072 |
| Output | -n, --tokens | Default when max_tokens is omitted. For grandpa, constrain it to around 4096 |
| CPU | -t, --threads | Helper threads for tokenization and prompt rendering. Leaving it unset is fine |
| Quality | --quality | Uses stricter kernels. Useful for quality evaluation or benchmarks, usually unnecessary for resident grandpa |
| Steering | --dir-steering-file, --dir-steering-ffn, --dir-steering-attn | For grandpa, apply -1 on the FFN side to amplify the succinct direction. Attention-side steering is experimental |
| Warmup | --warm-weights | Worth enabling for a resident service. Reduces first-inference page faults and stutter |
| Backend | --cuda, --metal, --cpu, --backend | CUDA for RTX PRO 6000. CPU is for diagnostics |
| API | --host, --port | Split roles by port, such as grandpa=8000 and frisky=8001 |
| Trace | --trace | Saves prompts, cache decisions, outputs, and tool calls in human-readable form. May be useful for fulfillment logs or DPO data |
| Thinking | reasoning_effort, thinking, think, model | grandpa should use model=deepseek-chat for non-thinking mode. reviewer / frisky can benefit from thinking mode |
| Disk KV | --kv-disk-dir, --kv-disk-space-mb, --kv-cache-* | 4GB likely suffices for grandpa. Workers with long review history may want 16GB |
| Tools | --disable-exact-dsml-tool-replay, --tool-memory-max-ids | Mostly irrelevant for grandpa when it does not use tool calling |
The proposed resident grandpa settings are --ctx 65536, --tokens 4096, 32GB disk KV and --trace. A tool-less dispatch role could use --ctx 32768 and 4GB disk KV.
[Container]
ContainerName=grandpa
Image=registry.home.arpa/dwarfstar4:latest
Pull=always
Network=host
AddDevice=nvidia.com/gpu=0
Volume=/mnt/data/models/models--antirez--deepseek-v4-gguf/snapshots/c566ab6d7c696ddd0c7f124e115228af1a326824:/model:ro
Volume=/mnt/data/models/models--antirez--deepseek-v4-gguf/kv_cache:/kv
Exec=-m /model/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf --ctx 65536 --tokens 4096 --host 0.0.0.0 --port 8000 --warm-weights --kv-disk-dir /kv --kv-disk-space-mb 32768 --trace /kv/trace.log
To verify the same settings once by hand, podman run --rm -it works directly.
podman run --rm -it \
--device nvidia.com/gpu=0 \
-v /mnt/data/models/models--antirez--deepseek-v4-gguf/snapshots/c566ab6d7c696ddd0c7f124e115228af1a326824:/model:ro,Z \
-v /mnt/data/models/models--antirez--deepseek-v4-gguf/kv_cache:/kv:Z \
--network host \
registry.home.arpa/dwarfstar4:latest \
-m /model/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
--ctx 65536 \
--tokens 4096 \
--host 0.0.0.0 \
--port 8000 \
--warm-weights \
--kv-disk-dir /kv \
--kv-disk-space-mb 32768 \
--trace /kv/trace.log

When using steering, generate the vector under dir-steering/ in the ds4 repository, mount it into the container, and pass it through --dir-steering-file. verbosity.f32 does not exist inside the container image by default, so generating it on the host and mounting it is the practical route.
cd ~/src/ds4
python3 dir-steering/tools/build_direction.py \
--ds4 ./ds4 \
--model ds4flash.gguf \
--good-file dir-steering/examples/succinct.txt \
--bad-file dir-steering/examples/verbose.txt \
--out dir-steering/out/verbosity.json \
--component ffn_out \
--ctx 512
ls -lh dir-steering/out/verbosity.f32
ds4-server has a different cache strategy from the cache pool in ik_llama.cpp, which reuses partial cache through f_keep. ds4 keeps only one live KV cache in VRAM. When switching sessions, it evicts the current KV to disk, checks whether the next request prefix hits the disk cache, and if it does, reads that cache back as the live session.
Directional Steering
y = y - scale * direction[layer] * dot(direction[layer], y)
The bundled verbosity vector has 43 layers × 4096 dimensions and is about 704KB. The scale changes output verbosity at runtime.
| Scale | Behavior |
|---|---|
-1 | Compresses output length to about half |
2 | Strengthens detailed explanations |
A proposed role split uses -1 for short orchestrator decisions and 2 for detailed reviewer output, without fine-tuning.
Summary
| Item | Value |
|---|---|
| Model | DeepSeek V4 Flash 284B |
| Runtime | DwarfStar 4 |
| GPU | RTX PRO 6000 Blackwell Max-Q 96GB x 1 |
| Quant | Q2-imatrix |
| Short generation | 43.6 tok/s |
| 50K context | 31.4 tok/s |
| Prefill | 245-267 tok/s |
| Model cache | 80.76 GiB |
| Peak VRAM | 93.1 GiB |
The run confirmed 43 tok/s on short output and the 31 tok/s range at 50K context on one GPU. EPYC 9175F’s 512MB L3 cache may contribute, but that has not been isolated.
DwarfStar 4 combines the engine, quantization, disk KV cache, tool calling and directional steering for DeepSeek V4 Flash.
It worked in my agent, although responses were rougher than my more thoroughly tuned GLM-5.1 setup. Optimization and quality checks are still beginning.
I plan to follow CUDA, disk KV and steering improvements. AMD MI350P PCIe Cards are also linked as a possible hardware candidate; their suitability has not been tested.
