Background

Llama-4-Scout is Meta’s MoE with 16 experts and 17B active parameters. I considered whether fewer experts than Maverick’s 128 could help CPU caching.

I compared CPU Q6_K / llama.cpp and GPU nvfp4 / vLLM for speed, quality, and memory. CPU measurements distinguish cold start from warm-cache operation.

I tested CPU-only llama.cpp for large-prefill oneshot summaries and fire-and-forget idempotent pipelines, alongside GPU nvfp4 for daily coding assistance.

Objective

  1. Measure CPU Q6_K steady-state speed (after cache warming)
  2. Record GPU nvfp4 Prefill/Decode speed and VRAM allocation
  3. Quantify mmap page cache and prompt cache effectiveness
  4. Establish the 100K context boundary reality

Test Environment

ItemSpecification
CPUAMD EPYC 9175F (Zen 5, 16C, L3 512MB)
MemoryDDR5-6400 768GB (12ch)
GPUNVIDIA RTX PRO 6000 Blackwell Max-Q (96GB VRAM)
OSUbuntu 24.04 LTS

Two Configurations

ConfigRuntimeQuantizationPlacementctx
A: CPUllama.cppQ6_K (3-split GGUF)All CPU8,192
B: GPUvLLM 0.14.0rc1nvfp4 (NVIDIA ModelOpt)All GPU~110K

Results

Config A: CPU Q6_K (th=16, ctx=8K)

MetricMeasuredNotes
PP (Prefill)~40 tok/sStable even with 2.5k-5k token prompts
TG (Decode)16.3-17.5 tok/sUniform 16-core load, temp stable at 64C

The first run is extremely slow due to mmap page faults + full Prefill. The model loads through mmap, so weights are not eagerly materialized at process start. When a large prompt arrives, page faults, full prompt eval, and KV cache initialization all hit simultaneously.

From the second run onward, OS page cache + prompt cache reach steady-state speed. Prompt cache LCP (Longest Common Prefix) match shows f_keep > 0.7. For repeated jobs with a stable front half, later runs avoid full prefill cost.

Prefill does not disappear – it still happens every run. What changes is whether the system has to do a full prefill every time. With stable leading context, the cost drops to delta-only processing.

Config B: GPU nvfp4 (vLLM, Blackwell 96GB)

MetricRangePeakNotes
PP (Prefill)700-1,200 tok/s1,370 tok/sEvident in long prompts
TG (Decode)30-60 tok/s112 tok/sStable post-warmup

Prefix Cache Hit Rate was 10-48%. TTFT fell when the hit rate reached 48%.

GPU prefill reached 1,370 tok/s and reduced waiting in aider refactoring. Japanese technical writing and Django/Python design followed instructions well in these FP4 tasks.

VRAM Breakdown

ItemUsage
Model weights~63.5 GiB
Available KV Cache~20.2 GiB
Max context length~110,256 tokens

The record puts the single-GPU BF16 KV ceiling near 100K and proposes FP8 KV or TP=2 for 256K+. Model weights and KV share the same 96GB, so context depends on their allocation.

Comparison Summary

ConfigTG(tok/s)VRAMUse Case
A: CPU Q6_K16-170Batch summarization, F&F pipeline
B: GPU nvfp430-60~64GBaider, interactive code generation

Analysis

CPU Cache Strategy Effectiveness

Warm CPU operation reached 17 tok/s. The operating choices were:

  • Use --mlock to keep weights in physical memory.
  • Fix the System Prompt for LCP reuse.
  • Avoid restarts and keep using the caches.

CPU is an option for oneshot summaries and idempotent batches with stable prefixes. Business application LLM integration can separate these from interactive work.

16-Expert Structural Advantage

Scout has 16 experts versus Maverick’s 128. The source interprets EPYC 9175F’s 1 CCD = 1 Core layout as retaining more expert working data in L3. This comparison did not directly measure that cache mechanism.

CPU throughput was Scout 17 versus Maverick 21-24 tok/s. I attribute more of the difference to model size; Scout felt more predictable.

nvfp4 Quality

FP4 worked for Django/Python design, boilerplate, and logic review. Mathematical reasoning felt weaker, while instruction following and Japanese explanations of backend structure stayed coherent.

Lessons Learned

CPU 17 tok/s and GPU 30-60 tok/s felt clearly different. CPU can take non-interactive work and leave the GPU for another model.

The GPU’s 1,370 tok/s Prefill and 48% prefix-cache hit rate shortened aider waits.

I would assign stable-prefix batches to CPU and low-latency interactive work to GPU.

Reproduction Steps

CPU Q6_K

  numactl --cpunodebind=0 --membind=0 \
  podman run --rm -it \
  -p 8081:8080 --shm-size 16g --cap-add=SYS_NICE \
  -v "$MO":/models:ro,Z $IMG \
  --host 0.0.0.0 --port 8080 \
  -m /models/Scout-17B-Q6_K.gguf \
  --threads 16 --threads-batch 16 \
  --batch-size 2048 --ubatch-size 512 \
  --mlock --ctx-size 8192 --flash-attn on \
  --parallel 1 --jinja
  

GPU nvfp4 (vLLM)

  vllm serve nvidia/Llama-4-Scout-17B-16E-Instruct-NVFP4 \
  --dtype auto --gpu-memory-utilization 0.88 \
  --max-num-seqs 1 --max-model-len 110000 \
  --enable-prefix-caching --trust-remote-code
  

Technical Notes

When Cache Stops Working

The original note lists these cache-loss conditions:

  • Changing one character in the prompt-head template or tool definition invalidates prompt cache and returns to full Prefill.
  • Process restarts purge OS page cache and bring back cold-start I/O waits.
  • Capacity limits evict old prompts.

100K Context Boundary

The nvfp4 model used 63.5 GiB and KV about 20 GiB. BF16 KV gives a physical limit near 110K. The recorded expansion plan for 256K/512K is FP8 KV + TP=2.