Background

I ran Kimi-K2.5 on EPYC 9175F to assess CPU batch inference for 1T-class Mixture-of-Experts (MoE) models released in late 2025 and 2026. Kimi-K2.5 has 1.04T total parameters and 32B active per token, selecting 8 of 384 experts.

The original estimate put Moonshot AI’s recommended minimum of 4x H200 at $150k–$200k, compared with about $15k for a 768GB DDR5 CPU server. I wanted to assess a smaller-team option for LLM integration.

The EPYC 9175F has 16 cores and 512MB L3. I tested thread counts and throughput to examine whether that cache-heavy design helps MoE inference.

Objective

Three questions to answer:

  1. Does the hypothesis “active MoE experts fit in the 512MB L3 cache, accelerating inference” hold?
  2. Is the throughput of Kimi-K2.5 (Q4_K_S) on EPYC 9175F practical for batch processing?
  3. What is the optimal thread count, given memory bandwidth constraints?

Test Environment

Hardware

ItemSpecification
CPUAMD EPYC 9175F (Zen 5, 16C/16T, SMT=OFF)
L3 Cache512MB (32MB per core)
MemoryDDR5-6400 64GB x 12ch = 768GB
GPURTX PRO 6000 MAX-Q 96GB (unused in this test)
TDP320W (cTDP 400W)

Software

ItemVersion
OSUbuntu 24.04 LTS
Runtimellama.cpp (server mode)
ContainerPodman rootless

Model

ItemSpecification
ModelKimi-K2.5 (Moonshot AI)
Total Parameters1.04T (61 layers: 60 MoE + 1 Dense)
Active Parameters32B (8 of 384 experts per token)
QuantizationQ4_K_S (GGUF)
KV Cache QuantizationQ8_0
Context LengthUp to 256K (tested 8K–128K)

CPU Topology

  ksh3@compute-server:~$ lscpu | egrep 'CPU\(s\)|Thread|Core|Socket|NUMA'
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Model name:                              AMD EPYC 9175F 16-Core Processor
Thread(s) per core:                      1
Core(s) per socket:                      16
Socket(s):                               1
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15

ksh3@compute-server:~$ cat /sys/devices/system/cpu/smt/active
0
  

SMT is OFF. 16 physical cores appear as 16 logical cores. Single NUMA node.

Methodology

llama.cpp Launch Parameters

Thread count was varied across 16, 14, 13, and 12 to measure Prefill (input processing) and Decode (token generation) throughput.

Base command (th=13 example):

  podman run --rm -p 8081:8080 --shm-size 16g --cap-add=SYS_NICE \
  -v /mnt/data/hf/hub/models--unsloth--Kimi-K2.5-GGUF:/models:Z \
  compute.home.arpa/llamacpp-zen5:latest \
  -m /models/snapshots/386fed8b054275941d6a495a9a7010fbf31b560d/Q4_K_S/Kimi-K2.5-Q4_K_S-00001-of-00013.gguf \
  --cache-type-k q8_0 --cache-type-v q8_0 --flash-attn on \
  --ctx-size 131072 --parallel 1 --threads 13 --threads-batch 13 \
  --batch-size 2048 --ubatch-size 512 --jinja --host 0.0.0.0 --port 8080
  

Parameter rationale:

  • --cache-type-k q8_0 --cache-type-v q8_0: Quantize KV cache to reduce memory footprint
  • --flash-attn on: Enable Flash Attention on CPU to reduce bandwidth pressure with long contexts
  • --cap-add=SYS_NICE: Allow thread priority optimization
  • --batch-size 2048 --ubatch-size 512: Maximize Prefill throughput with Zen 5’s AVX-512 VNNI

Prompt Cache Testing

Repeated requests with the same prompt prefix were tested to measure the effect of llama.cpp’s selected slot by LCP similarity feature.

Results

Thread-Count Scaling (ctx=8K, Q4_K_S)

ThreadsPrefill (tok/s)Decode (tok/s)Latency (ms/tok)Notes
16 (all cores)24.4312.9477.28Maximum throughput
1421.3212.5079.97Bandwidth saturation visible
13 (recommended)21.5811.6785.70Best efficiency/headroom balance
1214.5811.8684.32Compute resource starvation begins

Decode stayed near 11.67–12.94 tok/s at th=13–16, suggesting a memory-bandwidth limit. Prefill dropped sharply at th=12, showing that fewer compute threads hurt this setup.

Long-Context Stability at 128K (th=13)

MetricValueAssessment
Prefill22.39 tok/sAVX-512 VNNI sweet spot
Decode9.34 tok/sExceeds human reading speed (~6 tok/s) even at 128K
TTFT (39 new tokens)1,741 msLCP cache effective
KV Cache Latency107.10 ms/tokStable bandwidth control via 12ch DDR5

Decode held 9.34 tok/s at 128K without a sharp slowdown.

Prompt Cache Effect (ctx=16K)

RequestConditionPrompt ProcessingGenerationNotes
1stNo cache22.24 tok/s (823tok/37s)10.27 tok/s (438tok/42.7s)Cold start
2ndCache saving19.98 tok/s8.76 tok/sIncludes save overhead
3rdLCP similarity hit62 ms (cache lookup)10.0+ tok/sDramatic TTFT reduction

Prompt cache for 1260 tokens consumed 159.5 MiB. From the 3rd request onward, prefix recomputation was skipped, reducing TTFT from seconds to 62ms.

Memory Footprint

ItemMeasured ValueNotes
Model weights~522 GBQ4_K_S quantization
KV cache (16K ctx)~2.0 GBK:1098MiB / V:976MiB
Prompt cache~160 MBAt 1.2K tokens
Total RSS~523 GiB / 755 GiBNo swap activity

Runs without swap on 768GB. Over 200GB headroom remains for OS and background services.

Analysis

The original hypothesis failed

Original hypothesis:

MoE models have 10–30B active parameters, so the entire active expert set fits in EPYC 9175F’s 512MB L3 cache, accelerating inference.

Result: rejected. The 32B active parameters require roughly 16GB per generated token at Q4. They cannot all fit in 512MB.

Refined Hypothesis: Partial/Probabilistic Cache Hits

The revised hypothesis focuses on these reusable regions:

  • Router / Gating logic: Accessed for every token to determine expert selection. Small, high-frequency—naturally stays in L3
  • Projection / Bias tensors: Layer input/output boundaries. Small size, high reuse
  • Previous-layer weights / intermediate tensors: Temporal locality keeps them in L3 briefly
  • KV cache reuse fragments: Frequently accessed portions during attention computation

The hypothesis is that reusable working data hits L3 rather than all experts fitting there. The 9175F has 32MB per core, which may help retain that data. Throughput measurements alone do not establish cache hit rates.

Why “Fewer Cores, More Cache” Suits MoE

A 128-core CPU sharing 512MB has a simple average of 4MB per core. Irregular MoE access may cause more cache contention.

For 16 cores with 512MB, I proposed three possible effects:

  1. Less inter-core L3 contention
  2. Fewer evictions of reused data
  3. Less concentration of requests at the memory controller

Memory Bandwidth as the Bottleneck Boundary

Decode saturated at th=13–16. The theoretical bandwidth of 12-channel DDR5-6400 is 614 GB/s. I interpret the plateau near th=13 as a memory-bandwidth constraint.

Zen 5 AVX-512 Contribution

Zen 5 has a 512-bit datapath, unlike the earlier 256-bit x2 implementation. BF16 and AVX-512 VNNI may contribute to Q4_K_S dequantization and dot products. This interpretation draws on the 5.0GHz peak clock and 24.43 tok/s Prefill result.

Lessons Learned

The full-expert L3 hypothesis failed. I revised it to partial reuse from MoE access locality; its advantage over Dense models remains part of that interpretation.

I chose th=13 for operation. It keeps 90% of inference performance while leaving 3 cores for Dagster, Trino, and network IO.

The 9.34 tok/s result at 128K exceeded my expectation. I attribute part of it to MLA (Multi-head Latent Attention) KV compression.

Reproduction Steps

1. Hardware Requirements

  • AMD EPYC 9175F server (768GB DDR5-6400 recommended)
  • Storage: ~600GB for model files (NVMe recommended)

2. Model Download

  huggingface-cli download unsloth/Kimi-K2.5-GGUF \
  --include "Q4_K_S/*" \
  --local-dir /mnt/data/hf/hub/models--unsloth--Kimi-K2.5-GGUF
  

3. llama.cpp Container Build (Zen 5 Optimized)

Build with AVX-512 VNNI / BF16 enabled. Use -march=znver5 or -DGGML_NATIVE=ON.

  podman run --rm -p 8081:8080 --shm-size 16g --cap-add=SYS_NICE \
  -v /mnt/data/hf/hub/models--unsloth--Kimi-K2.5-GGUF:/models:Z \
  compute.home.arpa/llamacpp-zen5:latest \
  -m /models/snapshots/386fed8b054275941d6a495a9a7010fbf31b560d/Q4_K_S/Kimi-K2.5-Q4_K_S-00001-of-00013.gguf \
  --cache-type-k q8_0 --cache-type-v q8_0 --flash-attn on \
  --ctx-size 131072 --parallel 1 --threads 13 --threads-batch 13 \
  --batch-size 2048 --ubatch-size 512 --jinja --host 0.0.0.0 --port 8080
  

5. Verify

  curl -s http://localhost:8081/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"kimi-k2.5","messages":[{"role":"user","content":"Hello"}],"max_tokens":100}'
  

Technical Notes

SMT (Hyper-Threading) Should Be OFF

This setup keeps SMT OFF and uses 16 physical cores. SMT ON exposes 32 logical cores; the source favors OFF to retain effective L3 use. Check that /sys/devices/system/cpu/smt/active returns 0.

L3 Benefits Diminish with Dense 70B-Class Models

Dense 70B access differs from MoE expert-selection locality. I treat this as a MoE-specific configuration rather than extending the L3 hypothesis directly to Dense models.

Using prompt cache for CPU inference

Prompt caching can offset slow Prefill. A Dagster pipeline processing thousands of documents could reuse its shared System Prompt.

Cost Comparison

PlatformConfigurationEst. Decode SpeedHardware Cost
AMD EPYC 9175F1x CPU, 768GB RAM10–13 tok/s~$15k
Mac Studio M3 Ultra2x units (512GB)~21 tok/s~$20k
NVIDIA GPU Cluster4x H20040+ tok/s$150k–$200k

The original estimate is below 1/10 of the GPU cluster cost. This speed suits nightly batches and dataset generation better than interactive chat.