Three-configuration results

I compared IQuest-Coder-V1-40B on CPU (EPYC 9175F), GPU (vLLM + nvfp4) and Aider whole-edit. The development-pipeline checks cover input processing, generation speed and output termination.

The tested CPU setup had long input delays. GPU generation reached 25–28 tok/s, while whole-edit still increased waiting by regenerating complete files.

Test Environment

ItemSpecification
CPUAMD EPYC 9175F (Zen 5, 16C, L3 512MB)
GPUNVIDIA RTX PRO 6000 Blackwell Max-Q (96GB VRAM)
MemoryDDR5-6400 768GB (12ch)
OSUbuntu 24.04 LTS

Test Configurations

ConfigRuntimeQuantizationPlacement
A: CPUllama.cpp server (Podman)Q5_K_M (GGUF)CPU RAM
B: GPUvLLMnvfp4GPU VRAM
C: AidervLLM (Loop-Instruct variant)nvfp4GPU VRAM

CPU response delays

IQuest-Coder-V1-40B-Instruct is a 40B-class Dense (non-MoE) coding-specialized model. Unlike MoE models, Dense models use all parameters during inference—computation scales linearly with model size.

My first attempt ran IQuest-Coder-V1-40B-Instruct.q5_k_m.gguf using llama.cpp server within a Podman container, relying entirely on the CPU (AMD EPYC 9175F, 16 cores) without any GPU acceleration. The context length was set to 8192, targeting “wait-free” pipeline processing that demands low Time To First Token (TTFT) and low latency—essential for tools like Aider and other agent systems.

The CPU run showed these problems:

  • CPU usage: Dense 40B computes all layers per token, keeping cores busy.
  • TTFT: Inputs above 4k–5k tokens, such as task.n_tokens=4640, caused long prompt evaluation.
  • SSD: Loading improved, but generation and prompt evaluation changed little.
  • Batch settings: llama-server used batch-size 2048 and ubatch-size 512, favoring throughput over this low-latency workload.

NUMA binding, thread allocation and mlock were already applied. Dense 40B CPU processing appears to be the main limit in this configuration.

Config A Results

ItemMeasured
TTFT (4K-5K prompt)Tens of seconds (UX failure)
CPU usageAll cores pinned at 100%
Root cause40B full-layer computation per token, no shortcuts

Migration to GPU (nvfp4) and Measured Results

Given the CPU limitations, I transitioned to a GPU-based configuration using vLLM and nvfp4. For comparison, I also evaluated command-a-reasoning—a 111B class reasoning model—in parallel.

Measured Throughput

MetricMeasured
PP speed (Prompt throughput)1,100-2,300 tok/s
TG speed (Generation throughput)25-28 tok/s (stable)
KV cache usage2-12%
Prefix cache hit rate20-45%

From continuous vLLM logs:

  Engine 000: Avg generation throughput: 28.3 tokens/s, KV cache usage: 2.0%, Prefix cache hit rate: 22.8%
Engine 000: Avg generation throughput: 28.0 tokens/s, KV cache usage: 2.3%, Prefix cache hit rate: 22.8%
Engine 000: Avg generation throughput: 27.8 tokens/s, KV cache usage: 2.8%, Prefix cache hit rate: 22.8%
Engine 000: Avg generation throughput: 27.4 tokens/s, KV cache usage: 3.5%, Prefix cache hit rate: 22.8%
  
  • Prompt throughput was 1,100–2,300 tok/s.
  • Generation was 25–28 tok/s, about 2.3–2.5× the 111B command-a-reasoning speed.
  • KV cache usage was 2–12%.
  • Prefix cache hit rate was 20–45%; fixed prompts and tool schemas may improve it.

Generation continued during Running: 1 reqs. Occasional “0 tok/s” entries appear to reflect aggregation windows; no hang or deadlock signs were observed.

Estimated waiting time at TG 25–28 tok/s:

  • 200 tokens: ~7-8 seconds
  • 400 tokens: ~14-16 seconds
  • 800 tokens: ~30 seconds

This is less than half the wait time of the 111B model, making a clear difference when working with agents like Aider or during test generation.

Aider Whole-Edit Evaluation

Testing Aider in whole-edit mode on the GPU nvfp4 configuration surfaced a different class of problem.

MetricMeasured
TG speed0.6-8 tok/s (unstable)
KV cache usage7-13% (rapid growth)
Prefix cache hit rate6% (effectively disabled)

Whole-edit regenerates entire files. repo-map and attachments enlarge the prompt and KV use. Changing context produced a 6% prefix cache hit rate.

Comparison Summary

ConfigTG (tok/s)ViabilityUse Case
A: CPU Q5_K_MUnmeasurable (TTFT failure)No-
B: GPU nvfp425-28Productionagent/test gen/CI
C: Aider whole-edit0.6-8No-

Performance Factor (PF) definition

To translate these metrics into a practical sense of performance, I defined a simple metric called the “Performance Factor (PF)"—generated tokens per second divided by model size in billions of parameters.

PF ≈ generation tok/s ÷ model size (B)

ModelParamsTG speedtok/s per BRelative PF
command-a-reasoning111B~11 tok/s0.101.0
IQuest-Coder-40B nvfp440B26-28 tok/s0.65-0.70≈6.5-7.0
Small 7B fp167B60-80 tok/s9-11Separate category

Under this PF definition, IQuest-Coder-40B was 6–7× the 111B-class value. This compares speed per parameter; it does not establish reasoning quality or prove quantization correctness.

Analysis

CPU cost of Dense inference

MoE models (e.g., Kimi-K2.5 with 32B active parameters) compute only a subset of experts per token. Dense computes all 40B layers every time—no computation reduction possible. “40B-class” means fundamentally different CPU loads between MoE and Dense.

Dense full-layer reads remain a heavy workload on EPYC 9175F’s 12-channel memory. Their access pattern differs from keeping selected MoE experts in L3 cache.

ItemDense 40BMoE 229B (10B active)
Per-token computationAll 40B layers~10B equivalent
L3 cache utilizationIneffective (full-layer access)Effective (expert locality)
CPU TG speedUnmeasurable10-37 tok/s
CPU viabilityNon-viableViable for batch

Aider Whole-Edit Structural Problem

Why whole-edit is slow:

  1. Regenerating entire files produces massive output token counts
  2. repo-map + attached files inflate the prompt
  3. Prefix cache hit rate at 6% (constantly shifting context)
  4. KV cache grows rapidly (7→13%)

Switching to diff/patch format reduces output tokens and the time spent regenerating whole files.

Model Quality and Controllability

The model generated long, structured Go tests. Some examples had fewer logic errors or unsupported additions than the 111B reasoning model. EOS, stop and max_tokens allowed output termination to be controlled.

command-a-reasoning remains a candidate for deep reasoning, proofs and long thought. Daily generation also needs speed and termination control.

Why the Setup Looks Stable

The observed setup has four useful properties:

  1. nvfp4 is working correctly with vLLM main/nightly
  2. stop and max_tokens are behaving as expected
  3. The instruct-style EOS behavior is clean and predictable
  4. The model lines up well with the agent-gateway design, especially where retrieval behavior differs between streaming and non-streaming paths

LLM integration needs checks of the serving path and agent-gateway alongside the model.

Prioritizing Improvements

If you find yourself struggling with a similar setup, address the issues in this order:

  1. GPU placement Use --device nvidia.com/gpu=all and --n-gpu-layers 999. The tested GPU is RTX PRO 6000 MAX-Q 96GB.
  2. CPU tuning candidates
    • Set batch-size to 512–1024.
    • Set ubatch-size to 128–256.
    • Reduce history and repo-map.
  3. Smaller-model candidates Consider 7B–14B if CPU-only operation is required.

Reproduction settings

  vllm serve IQuestLab/IQuest-Coder-V1-40B-Instruct-nvfp4 \
  --max-num-seqs 1 \
  --max-model-len 32768
  
  podman run --rm -it \
  -p 8081:8080 --shm-size 16g --cap-add=SYS_NICE \
  -v "$MO":/models:ro,Z $IMG \
  --host 0.0.0.0 --port 8080 -m "$MODEL" \
  --jinja -c 8192 \
  --threads 14 --threads-batch 14 \
  -b 2048 -ub 512 \
  --parallel 1 --flash-attn on
  

Aider Optimization Settings

  • --edit-format diff: Avoid whole-edit, reduce output tokens
  • temperature=0: Greedy decoding for speed
  • Minimize repo-map and /add targets
  • Set max_model_len to minimum required

Conclusion and Next Steps

In this setup, IQuest-Coder-V1-40B + nvfp4 + vLLM is a candidate for daily generation. I am considering this division of work:

  • Daily work (agent/aider/CI): IQuest-Coder-40B nvfp4 (primary)
  • Deep reasoning / design review: command-a-reasoning (secondary)
  • Batch processing (CPU resident): MoE models (Kimi-K2.5 etc.)

The proposed routing uses IQuest-Coder-40B by default and command-a-reasoning for difficult reasoning or design reviews. Generation at 25–28 tok/s alone does not establish quality across all tasks.

Next I plan fixed prompts targeting a prefix cache hit rate above 60%, up from about 40%, followed by p95 latency tuning against SLOs.