Measured configurations

I ran the ~80B MoE Qwen3-Coder-Next in BF16 CPU, IQ4_NL hybrid / GPU and nvfp4 GPU modes for coding, review and security-audit assistance.

The same hardware measured speed, output and VRAM. BF16 CPU reached 7.59 tok/s, IQ4_NL 59–85 tok/s, and nvfp4 rolling logs 17–100 tok/s.

Objective

  1. Confirm BF16 CPU inference speed and quality (maximum precision mode)
  2. Quantify Expert Offload speed penalty with IQ4_NL Hybrid
  3. Measure nvfp4 GPU throughput
  4. Evaluate coding task quality
  5. Document SWA cache invalidation behavior with Qwen3-Next-80B-A3B-Thinking (Q4_K_M)

Test Environment

ItemSpecification
CPUAMD EPYC 9175F (Zen 5, 16C, L3 512MB)
GPUNVIDIA RTX PRO 6000 Blackwell Max-Q (96GB VRAM)
MemoryDDR5-6400 768GB (12ch)
OSUbuntu 24.04 LTS

Three Configurations

ConfigRuntimeQuantizationExpert Placementctx
A: CPU BF16llama.cppBF16 (unquantized)All CPU16K
B: Hybrid offloadik_llama.cppIQ4_NLExpert=CPU, Attn=GPU65K
C: GPU nvfp4vLLMnvfp4All GPU32K

Results

Config A: BF16 CPU (th=13, ctx=16K)

MetricMeasured
PP (short prompt)33.37 tok/s
PP (287 tokens)117.40 tok/s
TG (sustained)7.59 tok/s
TTFT (287 tokens)~2.58s

Throughput stayed consistent through a 2,233-token generation. KV cache at q8_0 to manage memory pressure.

BF16 avoids 4-bit quantization effects in the comparison. The 12-channel DDR5-6400 and Zen 5 AVX-512 BF16 setup reached 7.59 tok/s. Bandwidth and computation contributions were not isolated.

Config B: IQ4_NL Hybrid (Expert CPU offload, ctx=65K)

Run A: exps=CPU (Expert weights on CPU)

MetricMeasured
GPU buffer (weights)1,403 MiB
CPU buffer (weights)41,472 MiB
graph_splits98
TG (weighted avg)58.94 tok/s
PP (representative)761-1,120 tok/s

Run B: No exps=CPU (All weights on GPU)

MetricMeasured
GPU buffer (weights)42,875 MiB
graph_splits2
TG (weighted avg)85.36 tok/s
PP (representative)880-3,572 tok/s

Expert Offload speed penalty: -31% (58.94 -> 85.36 tok/s)

Run A kept most weights on CPU, about 1.4 GiB on GPU, and had graph_splits=98 with frequent synchronization and transfers.

Run B clustered around 85 tok/s, with 42.9 GiB GPU weights and graph_splits=2.

Weighted Average Methodology

Each task’s gen_tps is weighted by its gen_tokens, giving longer outputs more influence.

  weighted_gen_tps = S(gen_tokens_i * gen_tps_i) / S(gen_tokens_i)
  

Representative Task Results (Run A: exps=CPU)

taskprompt tokensprompt tpsgen tokensgen tpstotal ms
82100291.936659.141,458.59
2563,7571,045.8914159.715,953.40
6521,594761.226758.943,230.68
922908678.8118058.524,413.36
2,6632,284961.733,81058.8067,175.38

Representative Task Results (Run B: exps=CPU disabled)

taskprompt tokensprompt tpsgen tokensgen tpstotal ms
2121356.7147185.755,551.35
1,71876880.411,03685.4112,215.63
3,5396472,561.141,42185.5316,866.17
4,9611551,397.332,00785.2823,646.46
9,63830513.4547585.645,605.00

Config C: nvfp4 GPU (vLLM, ctx=32K)

MetricMeasured
TG (stable)17-100 tok/s (high variance)
PP17-669 tok/s (burst)

Values fluctuate due to vLLM rolling average logs. Stable generation runs at 58-100 tok/s.

Comparison Summary

ConfigTG(tok/s)VRAMQualityUse Case
A: BF16 CPU7.590HighestPrecision review/audit
B: Hybrid exps=CPU58.94~3GBGoodNormal coding
B: Hybrid all-GPU85.36~43GBGoodFast coding
C: nvfp4 GPU58-100~22GBGoodvLLM integration

Code-generation demonstration

The video records code generation and logic explanation.

Video link: https://www.youtube.com/watch?v=Hm8e7864Fcw

It shows generation progress and the actual output.

Analysis

Uses for BF16 CPU

At 7.59 tok/s, CPU operation favors review or audits that tolerate waiting. BF16 is the comparison baseline, not a guarantee of correct judgments.

Expert Offload 31% Penalty

Expert Offload saved over 40GB VRAM and reduced TG by31%. graph_splits increased from 2 to 98, suggesting more synchronization and transfer overhead.

When the weights fit in VRAM, the tested full-GPU placement was faster.

The conditions were A3B, n_ctx=65536, KV f16, n_batch=2048 and -ngl 99. On this 96GB-class setup, no exps=CPU is the proposed default.

Why It Gets Slower: Structural Analysis

  1. Synchronization and transfers: graph_splits=98 crosses CPU/GPU boundaries frequently.
  2. Capacity: Run B’s 42.9 GiB fits within 96GB, so CPU Expert placement is unnecessary.
  3. KV: Expert Offload does not reduce KV f16 memory at n_ctx=65536.

Coding-task outputs

Observed BF16 outputs:

  • Identified SQLi and plaintext-password risks and generated fixes and unit tests.
  • Returned “NOT IN SPEC” for out-of-spec questions.
  • Met 90% of the Django requirements but missed some multi-tenant safety details. The draft needs review.

SWA Cache Invalidation (Qwen3-Next-80B-A3B-Thinking)

When running Qwen3-Next-80B-A3B-Thinking (Q4_K_M) on CPU-only inference, full prompt re-processing was triggered by Sliding Window Attention (SWA) or hybrid/recurrent memory architecture behavior.

  slot update_slots: id  1 | task 6686 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)

slot update_slots: id  1 | task 6686 | erased invalidated context checkpoint (pos_min = 223, pos_max = 223, n_swa = 1, size = 75.376 MiB)
  

An approximately 75 MiB checkpoint was discarded and the full prompt reprocessed. This was not observed in Qwen3-Coder-Next IQ4_NL GPU tests, so the model and execution mode are distinguished.

Candidate uses

Candidate uses for LLM integration:

  • BF16 CPU: a baseline for audits and reviews that tolerate waiting.
  • IQ4_NL hybrid / GPU: daily generation at 59–85 tok/s.
  • nvfp4 / vLLM: tool-call-parser, prefix caching and API integration.

Reproduction Steps

BF16 CPU

  podman run --rm -it \
  -p 8081:8080 --shm-size 16g --cap-add=SYS_NICE \
  -v /mnt/data/hf/hub/models--unsloth--Qwen3-Coder-Next-GGUF:/models:Z \
  compute.home.arpa/llamacpp-zen5:qwen3-coder-next \
  -m /models/snapshots/.../BF16/Qwen3-Coder-Next-BF16-00001-of-00004.gguf \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --flash-attn on --ctx-size 16384 \
  --parallel 1 --threads 13 --threads-batch 13 \
  --batch-size 2048 --ubatch-size 512 \
  --jinja --host 0.0.0.0 --port 8080
  

IQ4_NL Hybrid (Expert CPU offload)

  IMG=compute.home.arpa/ik_llama-cuda
MO=/mnt/data/hf/hub/models--ubergarm--Qwen3-Coder-Next-GGUF
MODEL=/models/snapshots/.../Qwen3-Coder-Next-IQ4_KSS.gguf

podman run --rm -it --device nvidia.com/gpu=all \
  -p 8001:8080 --shm-size 16g --cap-add=SYS_NICE \
  -v "$MO":/models:ro,Z $IMG \
  --host 0.0.0.0 --port 8080 -m "$MODEL" \
  -c 65536 --threads 13 --threads-batch 23 \
  -b 2048 -ub 2048 -ngl 99 \
  -ot exps=CPU -fa on --no-mmap --jinja
  

IQ4_NL Full GPU (Expert on GPU)

  podman run --rm -it --device nvidia.com/gpu=all \
  -p 8001:8080 --shm-size 16g --cap-add=SYS_NICE \
  -v "$MO":/models:ro,Z $IMG \
  --host 0.0.0.0 --port 8080 -m "$MODEL" \
  -c 65536 --threads 13 --threads-batch 23 \
  -b 2048 -ub 2048 -ngl 99 \
  -fa on --no-mmap --jinja
  

nvfp4 GPU (vLLM)

  podman run --rm --device nvidia.com/gpu=all \
  --security-opt seccomp=unconfined --cap-add SYS_NICE --shm-size=16g \
  -v /mnt/data/hf:/data/hf:Z \
  -v /opt/containers/runtime/vllm/data/gpu_cache:/data/cache:Z \
  -p 8000:8000 \
  -e HF_HOME=/data/hf -e HF_DATASETS_CACHE=/data/hf \
  -e VLLM_CACHE_ROOT=/data/cache -e HF_HUB_OFFLINE=1 \
  -e FLASHINFER_DISABLE_VERSION_CHECK=1 \
  compute.home.arpa/vllm-gpu:nightly vincentzed-hf/Qwen3-Coder-Next-NVFP4 \
  --dtype auto --gpu-memory-utilization 0.88 \
  --max-num-seqs 1 --max-model-len 32768 \
  --enable-prefix-caching --trust-remote-code \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3 --served-model-name qwen3-coder-next-nvfp4
  

Qwen3-Next-80B-A3B-Thinking (CPU inference)

  podman run --rm \
  -p 8081:8080 --shm-size 1g \
  -v /opt/containers/runtime/llamacpp/data/models:/models:Z \
  compute.home.arpa/llamacpp-zen5:latest \
  -m /models/Qwen3-Next-80B-A3B-Thinking-Q4_K_M.gguf \
  --ctx-size 16384 --threads 15 \
  --jinja --reasoning-budget 512 \
  --host 0.0.0.0 --port 8080
  

Technical Notes

256K Context Configuration

IQ4_NL allows ctx=262144, but Expert Offload with 256K increases KV requirements. The proposed operating range here is about ctx=65536.

graph_splits Meaning

graph_splits counts compute-graph splits. It was 98 with exps=CPU and 2 on full GPU, indicating placements with more CPU/GPU synchronization and transfers.

Operational Guidelines

  • On a 96GB-class GPU, keep exps=CPU disabled by default
  • Use exps=CPU when VRAM capacity is insufficient or causes instability
  • Next, consider KV quantization and context tuning
  • Retain BF16 CPU inference as a baseline for background work that can wait