I ran DeepSeek-V4-Flash on dual RTX PRO 6000 Blackwell Max-Q 96GB GPUs. With a llama.cpp WIP branch and a community GGUF, generation reached around 35 t/s TG.

The 284B MoE / 13B active model used native FP4/FP8 GGUF. Flash Attention was disabled, and the DSV4 implementation was still being optimized. These are the results from 2026-04-27.

Video link: https://www.youtube.com/watch?v=Hjl4efNonxE

Model and speed

The setup used nsparks/DeepSeek-V4-Flash-FP4-FP8-GGUF and the wip/deepseek-v4-support branch from llama.cpp PR #22378. The Hugging Face model card lists 284B params; the GGUF uses the deepseek4 architecture.

ItemValue
ModelDeepSeek-V4-Flash
Parameters284B MoE
Active parameters13B
Experts256 experts / 6 active
GGUFnsparks/DeepSeek-V4-Flash-FP4-FP8-GGUF
Runtimellama.cpp wip/deepseek-v4-support
Commit rangeb8942-ba173dd08
Quantizationnative FP4 + FP8
GGUF size146GB
BPW4.39
GPUsRTX PRO 6000 Blackwell Max-Q 96GB x2

Input processing (PP) and generation (TG) measurements:

MetricValue
Prompt eval (PP)36.5-39.4 t/s
Token generation (TG)34.1-41.7 t/s
PP average38.3 t/s
TG average35.7 t/s
VRAMGPU0: 75.1GB, GPU1: 72.8GB
Offloaded layers44/44
CPU mapped1010 MiB
Flash Attentiondisabled
GPU utilization30-40% burst
Peak powerGPU0: 97.8W, GPU1: 115W
Graph splits3

GPU utilization was 30-40% in bursts, with power around 100-115W against a 300W TDP. Flash Attention and expert dispatch graph improvements may change the speed.

DCGM GPU Monitoring dashboard while running DeepSeek-V4-Flash GGUF
DCGM while running DeepSeek-V4-Flash GGUF with the llama.cpp WIP branch. GPU utilization was around 40%, VRAM was GPU0 75.8GB / GPU1 73.4GB, and power stayed around 97.8W / 115W.

Launch Command

The launch command:

  podman run --rm \
  -p 8000:8000 \
  --device nvidia.com/gpu=all \
  --shm-size 8g \
  -v /mnt/data/models/models--nsparks--DeepSeek-V4-Flash-FP4-FP8-GGUF:/models:Z \
  llama.cpp:deepseek-v4 \
  -s -m /models/snapshots/.../DeepSeek-V4-Flash-FP4-FP8-native.gguf \
  --n-gpu-layers 999 --threads 15 --threads-batch 24 \
  --ctx-size 8192 --parallel 1 -b 4096 -ub 2048 \
  --jinja --host 0.0.0.0 --port 8000 --alias deepseek-v4
  

llama-server provides an OpenAI-compatible endpoint, so the API wrapper I built for the official inference/ path was no longer needed.

Test Environment

ItemValue
GPUNVIDIA RTX PRO 6000 Blackwell Max-Q 96GB x2
Compute Capabilitysm_120
Driver580.126.09
CUDA13.0
CPUAMD EPYC 9175F
RAM768GB DDR5-6400
ContainerPodman
Native FP4BLACKWELL_NATIVE_FP4 enabled

Direct loading through transformers used around 80GB on each of GPU0 and GPU1. I saw an initial response, but a prompt request crashed the process. After investigating the official inference code, I switched to the published GGUF.

Speed by request

These prompts are short. They do not measure PP performance on long inputs.

#Requestprompt tokensPP (ms/t)PP (t/s)gen tokensTG (ms/t)TG (t/s)total (s)
1Japanese question1426.737.44127.136.91.5
2MoE explanation2025.938.613227.836.04.2
3FP4 vs FP81925.938.712527.736.14.0
4system prompt3125.838.718227.836.05.9
5multi-turn5025.938.651228.035.715.6
6Go code2325.838.723027.835.97.0
7logic puzzle2325.938.720727.835.96.4
8comparison analysis3025.838.742728.035.812.7
9JSON output2627.436.512229.334.14.3
10DSV4 architecture4525.838.741427.935.812.7

Short MoE explanations, code generation, and logic questions had no major failures. Japanese streaming sometimes produced text such as 東東京圜. With only a few tests on an experimental branch, I mainly use these results as a TG estimate.

Flash Attention was disabled

The logs show it being disabled automatically:

  sched_reserve: layer 0 is assigned to device CUDA0 but the Flash Attention tensor is assigned to device CPU (usually due to missing support)
sched_reserve: Flash Attention was auto, set to disabled
  

DeepSeek-V4 uses custom attention with CSA, HCA, and an Indexer. Flash Attention for this graph appears incomplete in the WIP branch. I suspect this is the main reason utilization stays around 30-40%.

I expect Flash Attention support to improve PP. For TG, the 4.39 BPW native FP4/FP8 GGUF is read across two GPUs for each token. Memory bandwidth and expert dispatch appear to matter more.

Problems with the official inference code

Before GGUF, I tried the official inference/*.py code. It generates locally through generate.py and has no HTTP endpoint. I wrote a FastAPI + uvicorn wrapper around its tokenizer, model, and distributed runtime to expose /v1/chat/completions.

MP=2 weight conversion succeeded. Direct transformers loading used around 80GB per GPU, and the official inference/*.py path also produced an initial response. At startup, nvtop showed processes using about 79680MiB on each of GPU0 and GPU1, or roughly 81% VRAM usage.

nvtop showing VRAM usage on both GPUs while starting the official inference code
During the official `inference/*.py` attempt, Python processes allocated roughly 79,680MiB on both GPU0 and GPU1. I confirmed the first response, but prompt requests crashed.

A prompt request then crashed the process. I investigated the issues below, but the NGC container’s torch version and the DSV4 FP4 dtype requirement remained blockers.

  python convert.py --hf-ckpt-path ${HF_CKPT_PATH} --save-path ${SAVE_PATH} \
  --n-experts 256 --model-parallel 2
  
  NCCL_NET_PLUGIN=none NCCL_IB_DISABLE=1 PYTHONPATH=. \
torchrun --standalone --nproc-per-node 2 main.py \
  --ckpt-path ${SAVE_PATH} --config ${CONFIG} --port 8000
  

Issues and attempted fixes:

IssueStatus
NCCL segfaultSegfault around ncclNetPluginInit during broadcast. Avoided with NCCL_NET_PLUGIN=none NCCL_IB_DISABLE=1
tilelang could not detect CUDAThe bare-metal environment did not have the CUDA toolkit; worked around by symlinking nvcc from a container overlay
sparse attention shared memorytilelang’s CSA sparse attention kernel required 104KB of dynamic shared memory
block size adjustmentLowering sparse attention block size from 64 -> 32 got past the shared-memory side
NGC torch too oldnvcr.io/nvidia/pytorch:25.04-py3 pins torch 2.7.0
DSV4 FP4 dtypeDeepSeek-V4-Flash needs torch.float4_e2m1fn_x2, which requires torch 2.11+

Blackwell supports up to 228KB/SM of dynamic shared memory, which should cover the kernel’s 104KB requirement. In practice, lowering block size from 64 -> 32 got past this issue.

The NGC container was pinned to torch 2.7.0 and lacked the float4_e2m1fn_x2 dtype. Dependency constraints prevented a simple replacement. I stopped the official inference/*.py attempt there and moved to native FP4/FP8 GGUF.

Converting GGUF myself

I also tried convert_hf_to_gguf.py from the nsparks WIP branch.

  python3 convert_hf_to_gguf.py ${HF_SNAP} \
  --outtype native \
  --torch-threads 16 \
  --outfile dsv4-flash-native.gguf
  

The conversion encountered these issues:

StageResult
torch 2.6F8_E8M0 KeyError
torch 2.11 CPUF8_E8M0 passed
transformersdeepseek_v4 model_type was not recognized
tokenizerPartly worked around by switching to PreTrainedTokenizerFast
pre-tokenizerStopped at unsupported joyai-llm pre-tokenizer

A community GGUF was available by then, so I used it to test inference instead of continuing the conversion.

How the published GGUF was converted

The model card for nsparks/DeepSeek-V4-Flash-FP4-FP8-GGUF identifies deepseek-ai/DeepSeek-V4-Flash as the source and provides this conversion command:

  python3 convert_hf_to_gguf.py /mnt/models/hf/DeepSeek-V4-Flash \
  --outtype moe-f8-e4m3-mxfp4 \
  --torch-threads 96 \
  --outfile DeepSeek-V4-Flash-FP4-FP8-native.gguf
  

The official DeepSeek Hugging Face repository is MIT licensed.

Upstream status on 2026-04-27

DeepSeek-V4 support in upstream llama.cpp was still WIP on 2026-04-27.

PR / DiscussionPurpose
llama.cpp PR #22378wip/deepseek-v4-support, including runtime graph, FP4/FP8 support, and performance hot paths
llama.cpp PR #22359DeepSeek-V4 GGUF conversion script
Discussion #22376DeepSeek-V4 support discussion
nsparks GGUFnative FP4/FP8 GGUF
official HFofficial DeepSeek-V4-Flash weights

PR #22378 had added FP4/FP8 support, DeepSeek4 runtime state save, F8 decode tuning, TOP_K fast path, and RMSNorm/copy kernel tuning. TG may approach the numbers seen with -ot exps=CPU.

Possible uses and longer context

At 35 t/s, I see possible uses in SFT/DPO distillation data, pipelines, and batch jobs.

I had been considering GLM-5.1, Kimi-K2.6, and Qwen3.5-397B as orchestrators for my agent system. With optimization in ik_llama.cpp or llama.cpp, DeepSeek-V4-Flash may also be a candidate for a CPU/GPU hybrid setup.

The reported KV reduction of around 90% also makes longer-context use on the GPUs interesting.

DeepSeek reports a 93% KV cache reduction and a 90% FLOPs reduction for DSV4 attention compared with V3.2.

The original record lists 192GB of total VRAM and model use of 75.1GB on GPU0 and 72.8GB on GPU1, or 147.9GB total. Its free-space figures are 21.5GB on GPU0 and 23.9GB on GPU1, or 45.4GB total.

If KV cache needs only 7% of the usual space, 32-45GB of free VRAM could support a large context. I am interested in whether one model could handle the role-based work and ctx management I have been coordinating across several models. If it can, some orchestration may no longer be needed.