Local DeepSeek-V4-Flash inference with llama.cpp and dual Blackwell GPUs
Local LLM deployment tests with DeepSeek-V4-Flash (284B MoE / 13B active), dual Blackwell Max-Q 96GB GPUs, and a llama.cpp WIP branch. Covers FP4/FP8 GGUF speed, official inference, and conversion issues.
I ran DeepSeek-V4-Flash on dual RTX PRO 6000 Blackwell Max-Q 96GB GPUs. With a llama.cpp WIP branch and a community GGUF, generation reached around 35 t/s TG.
The 284B MoE / 13B active model used native FP4/FP8 GGUF. Flash Attention was disabled, and the DSV4 implementation was still being optimized. These are the results from 2026-04-27.
Video link: https://www.youtube.com/watch?v=Hjl4efNonxE
Model and speed
The setup used nsparks/DeepSeek-V4-Flash-FP4-FP8-GGUF and the wip/deepseek-v4-support branch from llama.cpp PR #22378. The Hugging Face model card lists 284B params; the GGUF uses the deepseek4 architecture.
| Item | Value |
|---|---|
| Model | DeepSeek-V4-Flash |
| Parameters | 284B MoE |
| Active parameters | 13B |
| Experts | 256 experts / 6 active |
| GGUF | nsparks/DeepSeek-V4-Flash-FP4-FP8-GGUF |
| Runtime | llama.cpp wip/deepseek-v4-support |
| Commit range | b8942-ba173dd08 |
| Quantization | native FP4 + FP8 |
| GGUF size | 146GB |
| BPW | 4.39 |
| GPUs | RTX PRO 6000 Blackwell Max-Q 96GB x2 |
Input processing (PP) and generation (TG) measurements:
| Metric | Value |
|---|---|
| Prompt eval (PP) | 36.5-39.4 t/s |
| Token generation (TG) | 34.1-41.7 t/s |
| PP average | 38.3 t/s |
| TG average | 35.7 t/s |
| VRAM | GPU0: 75.1GB, GPU1: 72.8GB |
| Offloaded layers | 44/44 |
| CPU mapped | 1010 MiB |
| Flash Attention | disabled |
| GPU utilization | 30-40% burst |
| Peak power | GPU0: 97.8W, GPU1: 115W |
| Graph splits | 3 |
GPU utilization was 30-40% in bursts, with power around 100-115W against a 300W TDP. Flash Attention and expert dispatch graph improvements may change the speed.

Launch Command
The launch command:
podman run --rm \
-p 8000:8000 \
--device nvidia.com/gpu=all \
--shm-size 8g \
-v /mnt/data/models/models--nsparks--DeepSeek-V4-Flash-FP4-FP8-GGUF:/models:Z \
llama.cpp:deepseek-v4 \
-s -m /models/snapshots/.../DeepSeek-V4-Flash-FP4-FP8-native.gguf \
--n-gpu-layers 999 --threads 15 --threads-batch 24 \
--ctx-size 8192 --parallel 1 -b 4096 -ub 2048 \
--jinja --host 0.0.0.0 --port 8000 --alias deepseek-v4
llama-server provides an OpenAI-compatible endpoint, so the API wrapper I built for the official inference/ path was no longer needed.
Test Environment
| Item | Value |
|---|---|
| GPU | NVIDIA RTX PRO 6000 Blackwell Max-Q 96GB x2 |
| Compute Capability | sm_120 |
| Driver | 580.126.09 |
| CUDA | 13.0 |
| CPU | AMD EPYC 9175F |
| RAM | 768GB DDR5-6400 |
| Container | Podman |
| Native FP4 | BLACKWELL_NATIVE_FP4 enabled |
Direct loading through transformers used around 80GB on each of GPU0 and GPU1. I saw an initial response, but a prompt request crashed the process. After investigating the official inference code, I switched to the published GGUF.
Speed by request
These prompts are short. They do not measure PP performance on long inputs.
| # | Request | prompt tokens | PP (ms/t) | PP (t/s) | gen tokens | TG (ms/t) | TG (t/s) | total (s) |
|---|---|---|---|---|---|---|---|---|
| 1 | Japanese question | 14 | 26.7 | 37.4 | 41 | 27.1 | 36.9 | 1.5 |
| 2 | MoE explanation | 20 | 25.9 | 38.6 | 132 | 27.8 | 36.0 | 4.2 |
| 3 | FP4 vs FP8 | 19 | 25.9 | 38.7 | 125 | 27.7 | 36.1 | 4.0 |
| 4 | system prompt | 31 | 25.8 | 38.7 | 182 | 27.8 | 36.0 | 5.9 |
| 5 | multi-turn | 50 | 25.9 | 38.6 | 512 | 28.0 | 35.7 | 15.6 |
| 6 | Go code | 23 | 25.8 | 38.7 | 230 | 27.8 | 35.9 | 7.0 |
| 7 | logic puzzle | 23 | 25.9 | 38.7 | 207 | 27.8 | 35.9 | 6.4 |
| 8 | comparison analysis | 30 | 25.8 | 38.7 | 427 | 28.0 | 35.8 | 12.7 |
| 9 | JSON output | 26 | 27.4 | 36.5 | 122 | 29.3 | 34.1 | 4.3 |
| 10 | DSV4 architecture | 45 | 25.8 | 38.7 | 414 | 27.9 | 35.8 | 12.7 |
Short MoE explanations, code generation, and logic questions had no major failures. Japanese streaming sometimes produced text such as 東東京圜. With only a few tests on an experimental branch, I mainly use these results as a TG estimate.
Flash Attention was disabled
The logs show it being disabled automatically:
sched_reserve: layer 0 is assigned to device CUDA0 but the Flash Attention tensor is assigned to device CPU (usually due to missing support)
sched_reserve: Flash Attention was auto, set to disabled
DeepSeek-V4 uses custom attention with CSA, HCA, and an Indexer. Flash Attention for this graph appears incomplete in the WIP branch. I suspect this is the main reason utilization stays around 30-40%.
I expect Flash Attention support to improve PP. For TG, the 4.39 BPW native FP4/FP8 GGUF is read across two GPUs for each token. Memory bandwidth and expert dispatch appear to matter more.
Problems with the official inference code
Before GGUF, I tried the official inference/*.py code. It generates locally through generate.py and has no HTTP endpoint. I wrote a FastAPI + uvicorn wrapper around its tokenizer, model, and distributed runtime to expose /v1/chat/completions.
MP=2 weight conversion succeeded. Direct transformers loading used around 80GB per GPU, and the official inference/*.py path also produced an initial response. At startup, nvtop showed processes using about 79680MiB on each of GPU0 and GPU1, or roughly 81% VRAM usage.

A prompt request then crashed the process. I investigated the issues below, but the NGC container’s torch version and the DSV4 FP4 dtype requirement remained blockers.
python convert.py --hf-ckpt-path ${HF_CKPT_PATH} --save-path ${SAVE_PATH} \
--n-experts 256 --model-parallel 2
NCCL_NET_PLUGIN=none NCCL_IB_DISABLE=1 PYTHONPATH=. \
torchrun --standalone --nproc-per-node 2 main.py \
--ckpt-path ${SAVE_PATH} --config ${CONFIG} --port 8000
Issues and attempted fixes:
| Issue | Status |
|---|---|
| NCCL segfault | Segfault around ncclNetPluginInit during broadcast. Avoided with NCCL_NET_PLUGIN=none NCCL_IB_DISABLE=1 |
| tilelang could not detect CUDA | The bare-metal environment did not have the CUDA toolkit; worked around by symlinking nvcc from a container overlay |
| sparse attention shared memory | tilelang’s CSA sparse attention kernel required 104KB of dynamic shared memory |
| block size adjustment | Lowering sparse attention block size from 64 -> 32 got past the shared-memory side |
| NGC torch too old | nvcr.io/nvidia/pytorch:25.04-py3 pins torch 2.7.0 |
| DSV4 FP4 dtype | DeepSeek-V4-Flash needs torch.float4_e2m1fn_x2, which requires torch 2.11+ |
Blackwell supports up to 228KB/SM of dynamic shared memory, which should cover the kernel’s 104KB requirement. In practice, lowering block size from 64 -> 32 got past this issue.
The NGC container was pinned to torch 2.7.0 and lacked the float4_e2m1fn_x2 dtype. Dependency constraints prevented a simple replacement. I stopped the official inference/*.py attempt there and moved to native FP4/FP8 GGUF.
Converting GGUF myself
I also tried convert_hf_to_gguf.py from the nsparks WIP branch.
python3 convert_hf_to_gguf.py ${HF_SNAP} \
--outtype native \
--torch-threads 16 \
--outfile dsv4-flash-native.gguf
The conversion encountered these issues:
| Stage | Result |
|---|---|
| torch 2.6 | F8_E8M0 KeyError |
| torch 2.11 CPU | F8_E8M0 passed |
| transformers | deepseek_v4 model_type was not recognized |
| tokenizer | Partly worked around by switching to PreTrainedTokenizerFast |
| pre-tokenizer | Stopped at unsupported joyai-llm pre-tokenizer |
A community GGUF was available by then, so I used it to test inference instead of continuing the conversion.
How the published GGUF was converted
The model card for nsparks/DeepSeek-V4-Flash-FP4-FP8-GGUF identifies deepseek-ai/DeepSeek-V4-Flash as the source and provides this conversion command:
python3 convert_hf_to_gguf.py /mnt/models/hf/DeepSeek-V4-Flash \
--outtype moe-f8-e4m3-mxfp4 \
--torch-threads 96 \
--outfile DeepSeek-V4-Flash-FP4-FP8-native.gguf
The official DeepSeek Hugging Face repository is MIT licensed.
Upstream status on 2026-04-27
DeepSeek-V4 support in upstream llama.cpp was still WIP on 2026-04-27.
| PR / Discussion | Purpose |
|---|---|
| llama.cpp PR #22378 | wip/deepseek-v4-support, including runtime graph, FP4/FP8 support, and performance hot paths |
| llama.cpp PR #22359 | DeepSeek-V4 GGUF conversion script |
| Discussion #22376 | DeepSeek-V4 support discussion |
| nsparks GGUF | native FP4/FP8 GGUF |
| official HF | official DeepSeek-V4-Flash weights |
PR #22378 had added FP4/FP8 support, DeepSeek4 runtime state save, F8 decode tuning, TOP_K fast path, and RMSNorm/copy kernel tuning. TG may approach the numbers seen with -ot exps=CPU.
Possible uses and longer context
At 35 t/s, I see possible uses in SFT/DPO distillation data, pipelines, and batch jobs.
I had been considering GLM-5.1, Kimi-K2.6, and Qwen3.5-397B as orchestrators for my agent system. With optimization in ik_llama.cpp or llama.cpp, DeepSeek-V4-Flash may also be a candidate for a CPU/GPU hybrid setup.
The reported KV reduction of around 90% also makes longer-context use on the GPUs interesting.
DeepSeek reports a 93% KV cache reduction and a 90% FLOPs reduction for DSV4 attention compared with V3.2.
The original record lists 192GB of total VRAM and model use of 75.1GB on GPU0 and 72.8GB on GPU1, or 147.9GB total. Its free-space figures are 21.5GB on GPU0 and 23.9GB on GPU1, or 45.4GB total.
If KV cache needs only 7% of the usual space, 32-45GB of free VRAM could support a large context. I am interested in whether one model could handle the role-based work and ctx management I have been coordinating across several models. If it can, some orchestration may no longer be needed.
