Kimi-K2.6 is Moonshot AI’s 1T-class DeepSeek2-style MoE, activating 8 of 384 experts. I tested a CPU / GPU split for use as a local coding worker.

I tested IQ3_K and Q4_X from ubergarm/Kimi-K2.6-GGUF on EPYC 9175F and two RTX PRO 6000 Blackwell Max-Q 96GB GPUs. ik_llama.cpp provides MLA and expert placement, but not vision in this setup. I rebuilt mainline llama.cpp with --no-cache; its tool-call parser failed, so vision testing remains for a later retry.

Video link: https://www.youtube.com/watch?v=skTE19_JRYg

The video records the final demo after trying several placements on dual Blackwell Max-Q 96GB, EPYC 9175F, and 768GB DDR5-6400. I selected ubergarm’s IQ3_K (3.85 bpw, 460 GiB), placing 4 to 10 head / tail expert layers on GPU and the rest in CPU pinned memory.

  • TG was 17.9 to 21 t/s, and PP cold was 223 to 377 t/s, depending on -ub and -ot.
  • Continuous 14,707-token generation held 19.57 t/s. The model did not call custom MCP tools often, but it created issues, milestones, and PRs on local Gitea. I would like more TG and tool use like Opus 4.6. With ctk/ctv f16 at 256k, I recall VRAM staying below roughly 150 GB.
  • IQ3_K ran at 17.9 to 20.9 t/s; Q4_X at 16.6 to 19.0 t/s. With the same AGENTS.md, behavior looked similar, so I chose IQ3_K.
  • real_estate_sales produced 15 models and 772 lines of Django v6 code.
  • The license is Modified MIT. The original note describes commercial limits of about $20M monthly revenue or 100M MAU.

Validation Environment

ItemConfiguration
CPUAMD EPYC 9175F (16C/32T, L3 512MB, Zen 5)
RAM768 GB DDR5-6400 ECC RDIMM
GPUNVIDIA RTX PRO 6000 Blackwell Max-Q 96GB x2
OSUbuntu 24.04 LTS (minimal)
Runtimeik_llama.cpp

The model-side assumptions are:

ItemValue
ModelKimi-K2.6
ArchitectureDeepSeek2-style MoE + MLA
Total parameters1T class
Experts / Active384 / 8
Main quants testedIQ3_K (3.85 bpw), Q4_X (4.55 bpw)

Why ik_llama.cpp Instead of Mainline llama.cpp

Kimi-K2.6 needs MLA (Multi-head Latent Attention) support. Mainline llama.cpp used standard KV without an absorbed -mla 3 mode. On 2026-04-22, its peg-native parser failed on <|tool_call_begin|> with Failed to parse input at pos 433: <|im_end|>. A --no-cache rebuild did not fix it. I suspect the template and plan a later retry.

  Failed to parse input at pos 433: <|im_end|>
  
mainline llama.cpp crashing on a Kimi-K2.6 tool-call parse error
The tool-call parse error I hit when running Kimi-K2.6 on mainline llama.cpp

Expert Tensor Placement Determines TG

All-CPU experts limited decode speed. I returned selected head / tail expert layers to GPU with -ot, leaving the middle on CPU.

Benchmark Results

ConfigExpert on GPUTG avg (t/s)PP cold (t/s)VRAM/GPU
Baseline (--cpu-moe)0 layers18.9185~11 GiB
6-layer head/tail split6 layers (3+3)20.9223~52/60 GiB
10-layer head-heavy10 layers (8+2)20.3377~43/44 GiB

The balanced 6-layer split had the best TG. The 10-layer head-heavy layout with -ub 4096 improved PP but slightly reduced decode.

The -ot Layout I Kept in the End

  -ot "blk\.(1|2)\.ffn.*=CUDA0" \
-ot "blk\.(59|60)\.ffn.*=CUDA1" \
-ot "exps=CPU"
  

CUDA0 holds blk.1-blk.2; CUDA1 holds blk.59-blk.60. The other 56 expert layers stay in CPU pinned memory. Usage was in the mid-30s to high-30s GiB range per GPU, with 17-19 t/s decode. I kept this layout to leave GPU headroom.

Grafana GPU monitoring during a Kimi-K2.6 expert-placement run
DCGM metrics from the Kimi-K2.6 run while testing head/tail expert placement

IQ3_K and Q4_X Are Good at Different Things

The quantization comparison:

QuantBPWModel SizeTG avg (t/s)PP cold (t/s)CUDA_Host
IQ3_K3.85460 GiB20.9223368 GiB
Q4_X4.55544 GiB18.6567455 GiB

Q4_X had faster PP but slower TG. I attribute that to heavier expert-transfer costs. I preferred sustained decode for long coding tasks, and IQ3_K also saved 87 GiB RAM.

Choosing -mla

The -mla setting changes the speed profile.

FlagKV cache modeVRAM usageSpeed
-mla 0Standard KVHighestSlowest
-mla 1Compressed latent KVLowestSlow
-mla 3Absorbed MLAHighestFastest

At 132k context, KV used about 8.9 GiB. -mla 1 can reduce VRAM, but I chose -mla 3 for TG on the 96GB x2 setup.

Quality on Real Coding Tasks

For a business application development example, I generated Django tenant modules from Zed using IQ3_K, non-thinking mode, and 4- to 10-layer -ot layouts.

massage_salon Module

In about 15 minutes, the model created a Gitea issue, adjusted my own semantic-diff context tool .ctree.toml, scaffolded Django v6, and generated models, admin, and apps in one pass. That run produced 15 models and 839 lines of final code.

restaurant Module

The 10-layer head-heavy run modeled procurement, profitability, workforce, sales, and master data separately. It used RestaurantSettings, TextChoices, UniqueConstraint, and MinValueValidator appropriately.

real_estate_sales Module

With a 3+3 split, it generated 14,707 tokens in a single request over about 12.5 minutes, holding TG at 19.57 t/s. The output was 15 model classes and 772 lines covering property, appraisal, brokerage agreements, viewings, purchase applications, loan screening, sale contracts, and settlement as one connected flow.

The model used my unpublished ctree MCP configuration from the system prompt and changed scope during generation. Models above 500B have felt more capable to me, but that is an impression from use.

Kimi-K2.6 generating models and resources for a Django real_estate_sales module in Zed
The real_estate_sales generation run in Zed, with the model building out the models outline and resources together
Kimi-K2.6 planning Django settings and test coverage in Zed
A run where the model is organizing settings.py and test viewpoints together
Zed create_file calls alongside LLM timing logs and Grafana
A live observation screen while the model keeps issuing create_file tool calls, with timing logs at top right and Grafana at bottom right
Zed, htop, and Grafana watching Kimi-K2.6 generation live
Watching CPU load and Grafana control-plane metrics while Kimi-K2.6 is working through the task

It remained stable through multi-file generation, test-plan enumeration, and MCP calls in these longer tasks.

ik_llama.cpp vs Mainline llama.cpp

The same-hardware comparison:

EngineQuantTG (t/s)PP cold (t/s)Tool call
ik_llama.cppIQ3_K20.9223OK
ik_llama.cppQ4_X18.6567OK
llama.cpp (mainline)Q4_X15.4188Crash

Mainline failed on tool calls in my setup. I plan to retry it for vision, but will keep ik_llama.cpp for coding without images.

Reference: A Single-GPU Benchmark Seen in the HF Community

In ubergarm/Kimi-K2.6-GGUF discussion #3 on Hugging Face, I also saw a Q4_X benchmark on a single RTX PRO 6000 + EPYC 9355 + DDR5-6400 using aiperf. The average for a 16-turn conversation simulation looked like this:

EngineTG avg (t/s)TTFT avg (ms)Request latency avg (ms)
ik_llama.cpp18.858,56322,480
llama.cpp (mainline)16.0312,52628,872

That post also favored ik_llama.cpp across the metrics. It reported that -muge can hurt Kimi-K2.6, so I used the fork’s settings as a reference.

The Launch Command I Ended Up Keeping

  podman run --rm \
  --device nvidia.com/gpu=all \
  -p 8000:8000 \
  --cap-add=SYS_NICE \
  -v /mnt/data/models/models--ubergarm--Kimi-K2.6-GGUF:/models:ro,Z \
  registry.home.arpa/ik_llama.cpp:latest \
  -m /models/snapshots/${REF}/IQ3_K/Kimi-K2.6-IQ3_K-00001-of-00012.gguf \
  --ctx-size 131768 \
  --parallel 1 \
  --threads 15 \
  --threads-batch 32 \
  -b 8192 \
  -ub 4096 \
  -ngl 999 \
  -mla 3 \
  -ger \
  --special \
  -amb 512 \
  --jinja \
  --host 0.0.0.0 \
  --port 8000 \
  --warmup-batch \
  --alias kimi-k2.6-IQ3_K \
  -ot "blk\.(1|2)\.ffn.*=CUDA0" \
  -ot "blk\.(59|60)\.ffn.*=CUDA1" \
  -ot "exps=CPU" \
  --temp 0.6 \
  --chat-template-kwargs '{"thinking":false}'
  

The command exposes an OpenAI-compatible API for Zed and my agents.

Planned use

ItemValue
ModelKimi-K2.6 (1T MoE, 384×8 active)
QuantIQ3_K is the practical first choice
Engineik_llama.cpp
TG20.9 t/s (best 6-layer head/tail split)
PP cold223-377 t/s depending on -ub and placement
Real tasksLong-form Django tenant-module generation works
Intended useOrchestrator model
LicenseModified MIT ($20M / 100M MAU trigger)

MLA and -ot tuning made the local TG tolerable for my work. I plan to use this setup alongside Claude while building data pipelines.