Three configurations for Qwen3-Coder-Next 80B
Qwen3-Coder-Next was measured in BF16 CPU, IQ4_NL hybrid/GPU, and nvfp4 modes for local coding assistance. The comparison records 7.59/59–85/17–100 tok/s and a 31% Expert Offload penalty.
Measured configurations
I ran the ~80B MoE Qwen3-Coder-Next in BF16 CPU, IQ4_NL hybrid / GPU and nvfp4 GPU modes for coding, review and security-audit assistance.
The same hardware measured speed, output and VRAM. BF16 CPU reached 7.59 tok/s, IQ4_NL 59–85 tok/s, and nvfp4 rolling logs 17–100 tok/s.
Objective
- Confirm BF16 CPU inference speed and quality (maximum precision mode)
- Quantify Expert Offload speed penalty with IQ4_NL Hybrid
- Measure nvfp4 GPU throughput
- Evaluate coding task quality
- Document SWA cache invalidation behavior with Qwen3-Next-80B-A3B-Thinking (Q4_K_M)
Test Environment
| Item | Specification |
|---|---|
| CPU | AMD EPYC 9175F (Zen 5, 16C, L3 512MB) |
| GPU | NVIDIA RTX PRO 6000 Blackwell Max-Q (96GB VRAM) |
| Memory | DDR5-6400 768GB (12ch) |
| OS | Ubuntu 24.04 LTS |
Three Configurations
| Config | Runtime | Quantization | Expert Placement | ctx |
|---|---|---|---|---|
| A: CPU BF16 | llama.cpp | BF16 (unquantized) | All CPU | 16K |
| B: Hybrid offload | ik_llama.cpp | IQ4_NL | Expert=CPU, Attn=GPU | 65K |
| C: GPU nvfp4 | vLLM | nvfp4 | All GPU | 32K |
Results
Config A: BF16 CPU (th=13, ctx=16K)
| Metric | Measured |
|---|---|
| PP (short prompt) | 33.37 tok/s |
| PP (287 tokens) | 117.40 tok/s |
| TG (sustained) | 7.59 tok/s |
| TTFT (287 tokens) | ~2.58s |
Throughput stayed consistent through a 2,233-token generation. KV cache at q8_0 to manage memory pressure.
BF16 avoids 4-bit quantization effects in the comparison. The 12-channel DDR5-6400 and Zen 5 AVX-512 BF16 setup reached 7.59 tok/s. Bandwidth and computation contributions were not isolated.
Config B: IQ4_NL Hybrid (Expert CPU offload, ctx=65K)
Run A: exps=CPU (Expert weights on CPU)
| Metric | Measured |
|---|---|
| GPU buffer (weights) | 1,403 MiB |
| CPU buffer (weights) | 41,472 MiB |
| graph_splits | 98 |
| TG (weighted avg) | 58.94 tok/s |
| PP (representative) | 761-1,120 tok/s |
Run B: No exps=CPU (All weights on GPU)
| Metric | Measured |
|---|---|
| GPU buffer (weights) | 42,875 MiB |
| graph_splits | 2 |
| TG (weighted avg) | 85.36 tok/s |
| PP (representative) | 880-3,572 tok/s |
Expert Offload speed penalty: -31% (58.94 -> 85.36 tok/s)
Run A kept most weights on CPU, about 1.4 GiB on GPU, and had graph_splits=98 with frequent synchronization and transfers.
Run B clustered around 85 tok/s, with 42.9 GiB GPU weights and graph_splits=2.
Weighted Average Methodology
Each task’s gen_tps is weighted by its gen_tokens, giving longer outputs more influence.
weighted_gen_tps = S(gen_tokens_i * gen_tps_i) / S(gen_tokens_i)
Representative Task Results (Run A: exps=CPU)
| task | prompt tokens | prompt tps | gen tokens | gen tps | total ms |
|---|---|---|---|---|---|
| 82 | 100 | 291.93 | 66 | 59.14 | 1,458.59 |
| 256 | 3,757 | 1,045.89 | 141 | 59.71 | 5,953.40 |
| 652 | 1,594 | 761.22 | 67 | 58.94 | 3,230.68 |
| 922 | 908 | 678.81 | 180 | 58.52 | 4,413.36 |
| 2,663 | 2,284 | 961.73 | 3,810 | 58.80 | 67,175.38 |
Representative Task Results (Run B: exps=CPU disabled)
| task | prompt tokens | prompt tps | gen tokens | gen tps | total ms |
|---|---|---|---|---|---|
| 21 | 21 | 356.71 | 471 | 85.75 | 5,551.35 |
| 1,718 | 76 | 880.41 | 1,036 | 85.41 | 12,215.63 |
| 3,539 | 647 | 2,561.14 | 1,421 | 85.53 | 16,866.17 |
| 4,961 | 155 | 1,397.33 | 2,007 | 85.28 | 23,646.46 |
| 9,638 | 30 | 513.45 | 475 | 85.64 | 5,605.00 |
Config C: nvfp4 GPU (vLLM, ctx=32K)
| Metric | Measured |
|---|---|
| TG (stable) | 17-100 tok/s (high variance) |
| PP | 17-669 tok/s (burst) |
Values fluctuate due to vLLM rolling average logs. Stable generation runs at 58-100 tok/s.
Comparison Summary
| Config | TG(tok/s) | VRAM | Quality | Use Case |
|---|---|---|---|---|
| A: BF16 CPU | 7.59 | 0 | Highest | Precision review/audit |
| B: Hybrid exps=CPU | 58.94 | ~3GB | Good | Normal coding |
| B: Hybrid all-GPU | 85.36 | ~43GB | Good | Fast coding |
| C: nvfp4 GPU | 58-100 | ~22GB | Good | vLLM integration |
Code-generation demonstration
The video records code generation and logic explanation.
Video link: https://www.youtube.com/watch?v=Hm8e7864Fcw
It shows generation progress and the actual output.
Analysis
Uses for BF16 CPU
At 7.59 tok/s, CPU operation favors review or audits that tolerate waiting. BF16 is the comparison baseline, not a guarantee of correct judgments.
Expert Offload 31% Penalty
Expert Offload saved over 40GB VRAM and reduced TG by31%. graph_splits increased from 2 to 98, suggesting more synchronization and transfer overhead.
When the weights fit in VRAM, the tested full-GPU placement was faster.
The conditions were A3B, n_ctx=65536, KV f16, n_batch=2048 and -ngl 99. On this 96GB-class setup, no exps=CPU is the proposed default.
Why It Gets Slower: Structural Analysis
- Synchronization and transfers: graph_splits=98 crosses CPU/GPU boundaries frequently.
- Capacity: Run B’s 42.9 GiB fits within 96GB, so CPU Expert placement is unnecessary.
- KV: Expert Offload does not reduce KV f16 memory at n_ctx=65536.
Coding-task outputs
Observed BF16 outputs:
- Identified SQLi and plaintext-password risks and generated fixes and unit tests.
- Returned “NOT IN SPEC” for out-of-spec questions.
- Met 90% of the Django requirements but missed some multi-tenant safety details. The draft needs review.
SWA Cache Invalidation (Qwen3-Next-80B-A3B-Thinking)
When running Qwen3-Next-80B-A3B-Thinking (Q4_K_M) on CPU-only inference, full prompt re-processing was triggered by Sliding Window Attention (SWA) or hybrid/recurrent memory architecture behavior.
slot update_slots: id 1 | task 6686 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
slot update_slots: id 1 | task 6686 | erased invalidated context checkpoint (pos_min = 223, pos_max = 223, n_swa = 1, size = 75.376 MiB)
An approximately 75 MiB checkpoint was discarded and the full prompt reprocessed. This was not observed in Qwen3-Coder-Next IQ4_NL GPU tests, so the model and execution mode are distinguished.
Candidate uses
Candidate uses for LLM integration:
- BF16 CPU: a baseline for audits and reviews that tolerate waiting.
- IQ4_NL hybrid / GPU: daily generation at 59–85 tok/s.
- nvfp4 / vLLM: tool-call-parser, prefix caching and API integration.
Reproduction Steps
BF16 CPU
podman run --rm -it \
-p 8081:8080 --shm-size 16g --cap-add=SYS_NICE \
-v /mnt/data/hf/hub/models--unsloth--Qwen3-Coder-Next-GGUF:/models:Z \
compute.home.arpa/llamacpp-zen5:qwen3-coder-next \
-m /models/snapshots/.../BF16/Qwen3-Coder-Next-BF16-00001-of-00004.gguf \
--cache-type-k q8_0 --cache-type-v q8_0 \
--flash-attn on --ctx-size 16384 \
--parallel 1 --threads 13 --threads-batch 13 \
--batch-size 2048 --ubatch-size 512 \
--jinja --host 0.0.0.0 --port 8080
IQ4_NL Hybrid (Expert CPU offload)
IMG=compute.home.arpa/ik_llama-cuda
MO=/mnt/data/hf/hub/models--ubergarm--Qwen3-Coder-Next-GGUF
MODEL=/models/snapshots/.../Qwen3-Coder-Next-IQ4_KSS.gguf
podman run --rm -it --device nvidia.com/gpu=all \
-p 8001:8080 --shm-size 16g --cap-add=SYS_NICE \
-v "$MO":/models:ro,Z $IMG \
--host 0.0.0.0 --port 8080 -m "$MODEL" \
-c 65536 --threads 13 --threads-batch 23 \
-b 2048 -ub 2048 -ngl 99 \
-ot exps=CPU -fa on --no-mmap --jinja
IQ4_NL Full GPU (Expert on GPU)
podman run --rm -it --device nvidia.com/gpu=all \
-p 8001:8080 --shm-size 16g --cap-add=SYS_NICE \
-v "$MO":/models:ro,Z $IMG \
--host 0.0.0.0 --port 8080 -m "$MODEL" \
-c 65536 --threads 13 --threads-batch 23 \
-b 2048 -ub 2048 -ngl 99 \
-fa on --no-mmap --jinja
nvfp4 GPU (vLLM)
podman run --rm --device nvidia.com/gpu=all \
--security-opt seccomp=unconfined --cap-add SYS_NICE --shm-size=16g \
-v /mnt/data/hf:/data/hf:Z \
-v /opt/containers/runtime/vllm/data/gpu_cache:/data/cache:Z \
-p 8000:8000 \
-e HF_HOME=/data/hf -e HF_DATASETS_CACHE=/data/hf \
-e VLLM_CACHE_ROOT=/data/cache -e HF_HUB_OFFLINE=1 \
-e FLASHINFER_DISABLE_VERSION_CHECK=1 \
compute.home.arpa/vllm-gpu:nightly vincentzed-hf/Qwen3-Coder-Next-NVFP4 \
--dtype auto --gpu-memory-utilization 0.88 \
--max-num-seqs 1 --max-model-len 32768 \
--enable-prefix-caching --trust-remote-code \
--enable-auto-tool-choice --tool-call-parser qwen3_coder \
--reasoning-parser qwen3 --served-model-name qwen3-coder-next-nvfp4
Qwen3-Next-80B-A3B-Thinking (CPU inference)
podman run --rm \
-p 8081:8080 --shm-size 1g \
-v /opt/containers/runtime/llamacpp/data/models:/models:Z \
compute.home.arpa/llamacpp-zen5:latest \
-m /models/Qwen3-Next-80B-A3B-Thinking-Q4_K_M.gguf \
--ctx-size 16384 --threads 15 \
--jinja --reasoning-budget 512 \
--host 0.0.0.0 --port 8080
Technical Notes
256K Context Configuration
IQ4_NL allows ctx=262144, but Expert Offload with 256K increases KV requirements. The proposed operating range here is about ctx=65536.
graph_splits Meaning
graph_splits counts compute-graph splits. It was 98 with exps=CPU and 2 on full GPU, indicating placements with more CPU/GPU synchronization and transfers.
Operational Guidelines
- On a 96GB-class GPU, keep
exps=CPUdisabled by default - Use
exps=CPUwhen VRAM capacity is insufficient or causes instability - Next, consider KV quantization and context tuning
- Retain BF16 CPU inference as a baseline for background work that can wait
