DeepSeek-V3.2 throughput and caching

Kimi-K2.5 (1T MoE) ran at 10-13 tok/s in EPYC 9175F batches. I then evaluated DeepSeek-V3.2 as a second-opinion model for local LLM integration.

DeepSeek PP was 50-100 tok/s and TG 14-15 tok/s in llama.cpp. It felt slower, but these rates alone do not establish that ranking against Kimi. I checked runtime settings separately from model behavior.

Objective

  1. Identify why DeepSeek-V3.2 decode speed stalls at 14-15 tok/s
  2. Quantify the impact of cache control and optimization flags from llama.cpp logs
  3. Prioritize improvement actions

Test Environment

  • CPU: AMD EPYC 9175F (Zen 5, 16C)
  • Memory: DDR5-6400 768GB (12ch)
  • OS: Ubuntu 24.04 LTS
  • Runtime: llama.cpp (server mode, fused_moe=1)
  • Model: DeepSeek-V3.2 Speciale (MoE)
  • KV Cache: f16 (default)

Results

Measured Inference Throughput

Values from four consecutive tasks:

Task IDPP (tok)TG (tok)Cumulative TokensPP Speed (tok/s)TG Speed (tok/s)PP Time (s)TG Time (s)
02,7311,0243,75599.7514.5727.470.3
10268571,0245,63674.4715.2111.567.3
20513119826,92952.9214.515.967.7
30344,8651,02412,818100.2214.2248.572.0

PP ranged from 50-100 tok/s; TG stayed between 14.22 and 15.21 tok/s. TG changed little across these inputs and cumulative token counts.

Cache Mismatch in Logs

  Common part does not match fully → kv cache rm [p0, end)
  

Every task logged a leading-token mismatch and KV cache deletion.

Speculative Decoding Status

  no implementations specified for speculative decoding
  

No draft model was configured, so Speculative Decoding was inactive.

Analysis

Why TG stayed around 14-15 tok/s

Stable TG suggests checking memory bandwidth. The f16 KV cache and attention over long context may contribute, but these logs do not establish the cause.

Kimi used q8_0 KV cache to reduce bandwidth load. Applying it to DeepSeek is an unmeasured improvement proposal.

Impact of Prompt Cache Mismatches

Kimi used a fixed System Prompt + Knowledge Digest prefix for LCP caching. The DeepSeek run differed in three ways:

  1. <think> tag inconsistency: Thinking Prompt presence varied per request, breaking leading token alignment
  2. System Prompt variance: Templates were not locked down
  3. Context management difference: Kimi-K2.5 reused prior context; DeepSeek rebuilt context from scratch each time

Cache conditions differed, so the observed delay cannot yet be attributed to the model alone.

MoE Optimization Gap

fused_moe=1 was active. Differences between llama.cpp MoE routing and specialized vLLM or service kernels may also affect speed.

Align conditions before comparing

The logs establish cache mismatches and setting differences. They do not isolate a model-level cause for the perceived delay.

First fix the prefix, then compare speculative decoding and q8_0 KV settings. The expected improvement remains unmeasured.

Reproduction Steps

1. Run DeepSeek-V3.2

  podman run --rm -p 8081:8080 --shm-size 16g --cap-add=SYS_NICE \
  -v /path/to/deepseek-v3.2:/models:Z \
  compute.home.arpa/llamacpp-zen5:latest \
  -m /models/DeepSeek-V3.2-Speciale.gguf \
  --cache-type-k f16 --cache-type-v f16 --flash-attn on \
  --ctx-size 16384 --parallel 1 --threads 13 --threads-batch 13 \
  --batch-size 2048 --ubatch-size 512 --jinja --host 0.0.0.0 --port 8080
  

2. Improved Version (KV Cache Quantization + Fixed Prompt)

  # Change KV cache to q8_0
--cache-type-k q8_0 --cache-type-v q8_0

# Enable prompt cache
--prompt-cache /tmp/deepseek-cache.bin
  

Also fix System Prompt and <think> presence across requests.

3. Measurement

Extract S_PP and S_TG from server logs and compare before and after changes.

Technical Notes

Principles for Effective Prompt Caching

  1. Fix the leading token sequence: System Prompt → Fixed Context → Variable Parts, in strict order
  2. Keep Thinking mode consistent: If enabled, enable for all requests. Toggling per-request invalidates cache every time
  3. Align generation parameters: Temperature, top_p differences can also affect cache hit rates

Fair Comparison with Kimi-K2.5

  • Match output token count, temperature, top_p, stop sequences, and stream settings exactly
  • Unify Thinking Token handling (logs show Exclude reasoning tokens for slot selection, but generation still occurs)
  • Use identical hardware, thread count, and KV cache settings

Improvement Priority

PriorityActionExpected ImpactImplementation Cost
AFix prompt prefix consistencyMajor PP reductionLow (config change)
BEnable Speculative DecodingTG perceived speed gainMedium (draft model selection)
CQuantize KV cache to q8_0TG bandwidth reliefLow (flag change)
DStandardize generation conditionsFair comparisonLow (test design)