DeepSeek-V3.2 throughput and cache mismatches
Local DeepSeek-V3.2 Speciale logs reviewed for LLM integration performance. The record covers PP/TG, prefix mismatches, KV cache settings, proposed changes, and causes not yet established.
DeepSeek-V3.2 throughput and caching
Kimi-K2.5 (1T MoE) ran at 10-13 tok/s in EPYC 9175F batches. I then evaluated DeepSeek-V3.2 as a second-opinion model for local LLM integration.
DeepSeek PP was 50-100 tok/s and TG 14-15 tok/s in llama.cpp. It felt slower, but these rates alone do not establish that ranking against Kimi. I checked runtime settings separately from model behavior.
Objective
- Identify why DeepSeek-V3.2 decode speed stalls at 14-15 tok/s
- Quantify the impact of cache control and optimization flags from llama.cpp logs
- Prioritize improvement actions
Test Environment
- CPU: AMD EPYC 9175F (Zen 5, 16C)
- Memory: DDR5-6400 768GB (12ch)
- OS: Ubuntu 24.04 LTS
- Runtime: llama.cpp (server mode, fused_moe=1)
- Model: DeepSeek-V3.2 Speciale (MoE)
- KV Cache: f16 (default)
Results
Measured Inference Throughput
Values from four consecutive tasks:
| Task ID | PP (tok) | TG (tok) | Cumulative Tokens | PP Speed (tok/s) | TG Speed (tok/s) | PP Time (s) | TG Time (s) |
|---|---|---|---|---|---|---|---|
| 0 | 2,731 | 1,024 | 3,755 | 99.75 | 14.57 | 27.4 | 70.3 |
| 1026 | 857 | 1,024 | 5,636 | 74.47 | 15.21 | 11.5 | 67.3 |
| 2051 | 311 | 982 | 6,929 | 52.92 | 14.51 | 5.9 | 67.7 |
| 3034 | 4,865 | 1,024 | 12,818 | 100.22 | 14.22 | 48.5 | 72.0 |
PP ranged from 50-100 tok/s; TG stayed between 14.22 and 15.21 tok/s. TG changed little across these inputs and cumulative token counts.
Cache Mismatch in Logs
Common part does not match fully → kv cache rm [p0, end)
Every task logged a leading-token mismatch and KV cache deletion.
Speculative Decoding Status
no implementations specified for speculative decoding
No draft model was configured, so Speculative Decoding was inactive.
Analysis
Why TG stayed around 14-15 tok/s
Stable TG suggests checking memory bandwidth. The f16 KV cache and attention over long context may contribute, but these logs do not establish the cause.
Kimi used q8_0 KV cache to reduce bandwidth load. Applying it to DeepSeek is an unmeasured improvement proposal.
Impact of Prompt Cache Mismatches
Kimi used a fixed System Prompt + Knowledge Digest prefix for LCP caching. The DeepSeek run differed in three ways:
<think>tag inconsistency: Thinking Prompt presence varied per request, breaking leading token alignment- System Prompt variance: Templates were not locked down
- Context management difference: Kimi-K2.5 reused prior context; DeepSeek rebuilt context from scratch each time
Cache conditions differed, so the observed delay cannot yet be attributed to the model alone.
MoE Optimization Gap
fused_moe=1 was active. Differences between llama.cpp MoE routing and specialized vLLM or service kernels may also affect speed.
Align conditions before comparing
The logs establish cache mismatches and setting differences. They do not isolate a model-level cause for the perceived delay.
First fix the prefix, then compare speculative decoding and q8_0 KV settings. The expected improvement remains unmeasured.
Reproduction Steps
1. Run DeepSeek-V3.2
podman run --rm -p 8081:8080 --shm-size 16g --cap-add=SYS_NICE \
-v /path/to/deepseek-v3.2:/models:Z \
compute.home.arpa/llamacpp-zen5:latest \
-m /models/DeepSeek-V3.2-Speciale.gguf \
--cache-type-k f16 --cache-type-v f16 --flash-attn on \
--ctx-size 16384 --parallel 1 --threads 13 --threads-batch 13 \
--batch-size 2048 --ubatch-size 512 --jinja --host 0.0.0.0 --port 8080
2. Improved Version (KV Cache Quantization + Fixed Prompt)
# Change KV cache to q8_0
--cache-type-k q8_0 --cache-type-v q8_0
# Enable prompt cache
--prompt-cache /tmp/deepseek-cache.bin
Also fix System Prompt and <think> presence across requests.
3. Measurement
Extract S_PP and S_TG from server logs and compare before and after changes.
Technical Notes
Principles for Effective Prompt Caching
- Fix the leading token sequence: System Prompt → Fixed Context → Variable Parts, in strict order
- Keep Thinking mode consistent: If enabled, enable for all requests. Toggling per-request invalidates cache every time
- Align generation parameters: Temperature, top_p differences can also affect cache hit rates
Fair Comparison with Kimi-K2.5
- Match output token count, temperature, top_p, stop sequences, and stream settings exactly
- Unify Thinking Token handling (logs show
Exclude reasoning tokensfor slot selection, but generation still occurs) - Use identical hardware, thread count, and KV cache settings
Improvement Priority
| Priority | Action | Expected Impact | Implementation Cost |
|---|---|---|---|
| A | Fix prompt prefix consistency | Major PP reduction | Low (config change) |
| B | Enable Speculative Decoding | TG perceived speed gain | Medium (draft model selection) |
| C | Quantize KV cache to q8_0 | TG bandwidth relief | Low (flag change) |
| D | Standardize generation conditions | Fair comparison | Low (test design) |
