IQuest-Coder-40B on CPU, GPU and Aider
IQuest-Coder-V1-40B was measured with CPU Q5_K_M, GPU nvfp4 and Aider whole-edit. The LLM integration comparison covers GPU generation at 25–28 tok/s and development-assistance latency from long inputs and outputs.
Three-configuration results
I compared IQuest-Coder-V1-40B on CPU (EPYC 9175F), GPU (vLLM + nvfp4) and Aider whole-edit. The development-pipeline checks cover input processing, generation speed and output termination.
The tested CPU setup had long input delays. GPU generation reached 25–28 tok/s, while whole-edit still increased waiting by regenerating complete files.
Test Environment
| Item | Specification |
|---|---|
| CPU | AMD EPYC 9175F (Zen 5, 16C, L3 512MB) |
| GPU | NVIDIA RTX PRO 6000 Blackwell Max-Q (96GB VRAM) |
| Memory | DDR5-6400 768GB (12ch) |
| OS | Ubuntu 24.04 LTS |
Test Configurations
| Config | Runtime | Quantization | Placement |
|---|---|---|---|
| A: CPU | llama.cpp server (Podman) | Q5_K_M (GGUF) | CPU RAM |
| B: GPU | vLLM | nvfp4 | GPU VRAM |
| C: Aider | vLLM (Loop-Instruct variant) | nvfp4 | GPU VRAM |
CPU response delays
IQuest-Coder-V1-40B-Instruct is a 40B-class Dense (non-MoE) coding-specialized model. Unlike MoE models, Dense models use all parameters during inference—computation scales linearly with model size.
My first attempt ran IQuest-Coder-V1-40B-Instruct.q5_k_m.gguf using llama.cpp server within a Podman container, relying entirely on the CPU (AMD EPYC 9175F, 16 cores) without any GPU acceleration. The context length was set to 8192, targeting “wait-free” pipeline processing that demands low Time To First Token (TTFT) and low latency—essential for tools like Aider and other agent systems.
The CPU run showed these problems:
- CPU usage: Dense 40B computes all layers per token, keeping cores busy.
- TTFT: Inputs above 4k–5k tokens, such as
task.n_tokens=4640, caused long prompt evaluation. - SSD: Loading improved, but generation and prompt evaluation changed little.
- Batch settings:
llama-serverusedbatch-size 2048andubatch-size 512, favoring throughput over this low-latency workload.
NUMA binding, thread allocation and mlock were already applied. Dense 40B CPU processing appears to be the main limit in this configuration.
Config A Results
| Item | Measured |
|---|---|
| TTFT (4K-5K prompt) | Tens of seconds (UX failure) |
| CPU usage | All cores pinned at 100% |
| Root cause | 40B full-layer computation per token, no shortcuts |
Migration to GPU (nvfp4) and Measured Results
Given the CPU limitations, I transitioned to a GPU-based configuration using vLLM and nvfp4. For comparison, I also evaluated command-a-reasoning—a 111B class reasoning model—in parallel.
Measured Throughput
| Metric | Measured |
|---|---|
| PP speed (Prompt throughput) | 1,100-2,300 tok/s |
| TG speed (Generation throughput) | 25-28 tok/s (stable) |
| KV cache usage | 2-12% |
| Prefix cache hit rate | 20-45% |
From continuous vLLM logs:
Engine 000: Avg generation throughput: 28.3 tokens/s, KV cache usage: 2.0%, Prefix cache hit rate: 22.8%
Engine 000: Avg generation throughput: 28.0 tokens/s, KV cache usage: 2.3%, Prefix cache hit rate: 22.8%
Engine 000: Avg generation throughput: 27.8 tokens/s, KV cache usage: 2.8%, Prefix cache hit rate: 22.8%
Engine 000: Avg generation throughput: 27.4 tokens/s, KV cache usage: 3.5%, Prefix cache hit rate: 22.8%
- Prompt throughput was 1,100–2,300 tok/s.
- Generation was 25–28 tok/s, about 2.3–2.5× the 111B
command-a-reasoningspeed. - KV cache usage was 2–12%.
- Prefix cache hit rate was 20–45%; fixed prompts and tool schemas may improve it.
Generation continued during Running: 1 reqs. Occasional “0 tok/s” entries appear to reflect aggregation windows; no hang or deadlock signs were observed.
Estimated waiting time at TG 25–28 tok/s:
- 200 tokens: ~7-8 seconds
- 400 tokens: ~14-16 seconds
- 800 tokens: ~30 seconds
This is less than half the wait time of the 111B model, making a clear difference when working with agents like Aider or during test generation.
Aider Whole-Edit Evaluation
Testing Aider in whole-edit mode on the GPU nvfp4 configuration surfaced a different class of problem.
| Metric | Measured |
|---|---|
| TG speed | 0.6-8 tok/s (unstable) |
| KV cache usage | 7-13% (rapid growth) |
| Prefix cache hit rate | 6% (effectively disabled) |
Whole-edit regenerates entire files. repo-map and attachments enlarge the prompt and KV use. Changing context produced a 6% prefix cache hit rate.
Comparison Summary
| Config | TG (tok/s) | Viability | Use Case |
|---|---|---|---|
| A: CPU Q5_K_M | Unmeasurable (TTFT failure) | No | - |
| B: GPU nvfp4 | 25-28 | Production | agent/test gen/CI |
| C: Aider whole-edit | 0.6-8 | No | - |
Performance Factor (PF) definition
To translate these metrics into a practical sense of performance, I defined a simple metric called the “Performance Factor (PF)"—generated tokens per second divided by model size in billions of parameters.
PF ≈ generation tok/s ÷ model size (B)
| Model | Params | TG speed | tok/s per B | Relative PF |
|---|---|---|---|---|
| command-a-reasoning | 111B | ~11 tok/s | 0.10 | 1.0 |
| IQuest-Coder-40B nvfp4 | 40B | 26-28 tok/s | 0.65-0.70 | ≈6.5-7.0 |
| Small 7B fp16 | 7B | 60-80 tok/s | 9-11 | Separate category |
Under this PF definition, IQuest-Coder-40B was 6–7× the 111B-class value. This compares speed per parameter; it does not establish reasoning quality or prove quantization correctness.
Analysis
CPU cost of Dense inference
MoE models (e.g., Kimi-K2.5 with 32B active parameters) compute only a subset of experts per token. Dense computes all 40B layers every time—no computation reduction possible. “40B-class” means fundamentally different CPU loads between MoE and Dense.
Dense full-layer reads remain a heavy workload on EPYC 9175F’s 12-channel memory. Their access pattern differs from keeping selected MoE experts in L3 cache.
| Item | Dense 40B | MoE 229B (10B active) |
|---|---|---|
| Per-token computation | All 40B layers | ~10B equivalent |
| L3 cache utilization | Ineffective (full-layer access) | Effective (expert locality) |
| CPU TG speed | Unmeasurable | 10-37 tok/s |
| CPU viability | Non-viable | Viable for batch |
Aider Whole-Edit Structural Problem
Why whole-edit is slow:
- Regenerating entire files produces massive output token counts
- repo-map + attached files inflate the prompt
- Prefix cache hit rate at 6% (constantly shifting context)
- KV cache grows rapidly (7→13%)
Switching to diff/patch format reduces output tokens and the time spent regenerating whole files.
Model Quality and Controllability
The model generated long, structured Go tests. Some examples had fewer logic errors or unsupported additions than the 111B reasoning model. EOS, stop and max_tokens allowed output termination to be controlled.
command-a-reasoning remains a candidate for deep reasoning, proofs and long thought. Daily generation also needs speed and termination control.
Why the Setup Looks Stable
The observed setup has four useful properties:
nvfp4is working correctly withvLLM main/nightlystopandmax_tokensare behaving as expected- The instruct-style EOS behavior is clean and predictable
- The model lines up well with the
agent-gatewaydesign, especially where retrieval behavior differs between streaming and non-streaming paths
LLM integration needs checks of the serving path and agent-gateway alongside the model.
Prioritizing Improvements
If you find yourself struggling with a similar setup, address the issues in this order:
- GPU placement
Use
--device nvidia.com/gpu=alland--n-gpu-layers 999. The tested GPU is RTX PRO 6000 MAX-Q 96GB. - CPU tuning candidates
- Set
batch-sizeto 512–1024. - Set
ubatch-sizeto 128–256. - Reduce history and repo-map.
- Set
- Smaller-model candidates Consider 7B–14B if CPU-only operation is required.
Reproduction settings
GPU nvfp4 (Recommended)
vllm serve IQuestLab/IQuest-Coder-V1-40B-Instruct-nvfp4 \
--max-num-seqs 1 \
--max-model-len 32768
CPU Q5_K_M (Reference: Not Recommended)
podman run --rm -it \
-p 8081:8080 --shm-size 16g --cap-add=SYS_NICE \
-v "$MO":/models:ro,Z $IMG \
--host 0.0.0.0 --port 8080 -m "$MODEL" \
--jinja -c 8192 \
--threads 14 --threads-batch 14 \
-b 2048 -ub 512 \
--parallel 1 --flash-attn on
Aider Optimization Settings
--edit-format diff: Avoid whole-edit, reduce output tokenstemperature=0: Greedy decoding for speed- Minimize repo-map and /add targets
- Set
max_model_lento minimum required
Conclusion and Next Steps
In this setup, IQuest-Coder-V1-40B + nvfp4 + vLLM is a candidate for daily generation. I am considering this division of work:
- Daily work (agent/aider/CI): IQuest-Coder-40B nvfp4 (primary)
- Deep reasoning / design review: command-a-reasoning (secondary)
- Batch processing (CPU resident): MoE models (Kimi-K2.5 etc.)
The proposed routing uses IQuest-Coder-40B by default and command-a-reasoning for difficult reasoning or design reviews. Generation at 25–28 tok/s alone does not establish quality across all tasks.
Next I plan fixed prompts targeting a prefix cache hit rate above 60%, up from about 40%, followed by p95 latency tuning against SLOs.
