I ran GLM-5.1 IQ3_KS (744B MoE, 320 GiB) on two RTX PRO 6000 Blackwell Max-Q 96GB GPUs and 768GB RAM. CPU/GPU hybrid inference in ik_llama.cpp measured TG of 17–19 tok/s. I also compared it with Qwen3.5-397B-A17B as the resident grandpa orchestrator for familiar.

Video link: https://www.youtube.com/watch?v=1JRyuCUlFeI

Hardware

ComponentSpec
CPUAMD EPYC 9175F (16C)
RAM768GB DDR5-6400
GPUNVIDIA RTX PRO 6000 Blackwell Max-Q 96GB × 2

Model

GLM-5.1 is a MoE model in Zhipu’s GLM family. In the ik_llama.cpp startup log it shows n_expert = 256 and n_expert_used = 8.

ItemValue
Architectureglm-dsa (MoE, 256 experts, 8 active)
Parameters753.864B
QuantizationIQ3_KS (3.65 BPW)
Model size320.216 GiB
Context65536 (max 202752)
GGUFubergarm/GLM-5.1-GGUF
Runtimeik_llama.cpp

Head and tail experts on GPU

GLM-5.1 has 79 layers (blk.0–blk.78); blk.0–blk.2 are dense and blk.3–blk.78 are MoE layers.

With --cpu-moe the experts default to CPU, and -ot brings individual tensors back to GPU. In this test I place the head 15 layers (blk.3–blk.17) on CUDA0, the tail 4 layers (blk.74–blk.77) on CUDA1, and leave the middle 56 layers on CUDA_Host (pinned host memory).

  OT_ARGS=""
for i in $(seq 3 17); do
  OT_ARGS="$OT_ARGS -ot blk.$i.ffn_gate_exps=CUDA0"
  OT_ARGS="$OT_ARGS -ot blk.$i.ffn_down_exps=CUDA0"
  OT_ARGS="$OT_ARGS -ot blk.$i.ffn_up_exps=CUDA0"
done
for i in $(seq 74 77); do
  OT_ARGS="$OT_ARGS -ot blk.$i.ffn_gate_exps=CUDA1"
  OT_ARGS="$OT_ARGS -ot blk.$i.ffn_down_exps=CUDA1"
  OT_ARGS="$OT_ARGS -ot blk.$i.ffn_up_exps=CUDA1"
done
  

Placement result (per MoE layer, experts are gate 1,225 + down 1,638 + up 1,225 = 4,088 MiB):

DeviceRoleBuffer Size
CUDA0blk.0–2 dense + blk.3–17 experts (head 15 layers) + attn/KV68,698 MiB (~67.1 GiB)
CUDA1blk.74–77 experts (tail 4 layers) + attn/KV23,473 MiB (~22.9 GiB)
CUDA_HostMiddle layer experts (56 layers)229,438 MiB (~224.1 GiB)

GPU-side experts total 19 layers = 77,672 MiB (~75.9 GiB), and CPU pinned host holds 56 layers = 228,928 MiB of experts, which matches the log’s CUDA_Host 229,438 MiB. If all experts were kept on CPU, the theoretical size would be 76 × 4,088 = 310,688 MiB (~303.4 GiB); this run moves 76 GiB back to GPU via -ot.

ik_llama.cpp tensor placement log at startup
ik_llama.cpp tensor placement log at startup. It reports 80/80 layers offloaded, but most of the actual expert tensors live on CUDA_Host

Launch Command

  podman run --rm \
  --device nvidia.com/gpu=all \
  -p 8000:8000 \
  --cap-add=SYS_NICE \
  -v /mnt/data/models/models--ubergarm--GLM-5.1-GGUF:/models:ro,Z \
  registry.home.arpa/ik_llama.cpp:latest \
  -m /models/.../IQ3_KS/GLM-5.1-IQ3_KS-00001-of-00008.gguf \
  --ctx-size 65536 -ctk q8_0 -ctv q8_0 \
  --parallel 1 --threads 15 --threads-batch 24 \
  -b 8192 -ub 8192 -ngl 99 --cpu-moe \
  $OT_ARGS \
  -ger -muge -amb 512 --jinja \
  --host 0.0.0.0 --port 8000 \
  --warmup-batch --alias glm-5.1
  

Option flags:

  • --cpu-moe: default experts to CPU
  • -ot blk.N.ffn_*_exps=CUDAX: override individual expert tensors to GPU
  • -ger -muge: grouped expert routing + multi-GPU expert
  • -amb 512: attention memory budget
  • --warmup-batch: batch warmup at startup

Django App Generation

I used the Zed agent to generate a logistics tenant module with model definitions, admin, tests and seed data. This is a business application development workload for local LLM integration.

GLM-5.1 generating a Django transport module inside the Zed editor
GLM-5.1 generating a Django transport module inside the Zed editor. models.py outline in progress

Token Generation (TG)

MetricValue
Requests46
Total generated tokens16,092
Total prompt tokens131,985
TG min16.39 tok/s
TG max19.38 tok/s
TG median17.94 tok/s
TG mean17.77 tok/s
ms/token range52–61 ms

The largest generation was 8,884 tokens (PP 435 tok/s, TG 17.26 tok/s), producing the Django models.py code over about 8.5 minutes.

TG Stability

TG fell by about 2 tok/s from the start to the end of a 53k/64k ctx session.

  • First 10 requests: 18.23–18.97 tok/s (avg 18.74)
  • Last 10 requests: 16.39–16.86 tok/s (avg 16.69)

The KV cache growth and longer context are the likely cause. Even so, TG never dropped below 16 tok/s.

Prompt Processing (PP)

Prompt SizePP Range
< 100 tokens19–37 tok/s
100–1,00054–143 tok/s
1,000–5,000114–280 tok/s
5,000–10,000235–572 tok/s

PP throughput improves with longer prompts. The initial 20,956-token input hit 571.94 tok/s. Short prompts pay relatively more overhead.

Cache Miss Problem

The log shows 26 prefix cache misses during the session, likely caused by <think> tag handling.

  Common part does not match fully
cache : ...<|assistant|><think></think>...
prompt: ...<|assistant|></think>...
  

Changes in <think> presence or position can break prefix matching and trigger 7k–10k token re-evaluation. This appears to be the main cause of TTFT increasing from about 1 second to 40 seconds.

Prompt TokensTTFT (est.)
231.05 s
7,48717.25 s
9,24739.28 s
9,76140.87 s

GPU Metrics

DCGM GPU Monitoring dashboard
DCGM GPU Monitoring. GPU0: 24% util / 79.7GB VRAM, GPU1: 26% util / 33.4GB VRAM. The asymmetric head-15 / tail-4 expert placement is reflected directly

GPU utilization doesn’t pin at 100%; it oscillates per request. This is characteristic of a hybrid setup, where host-side expert reads interleave with GPU compute.

GPU Utilization time series
GPU Utilization over time. It oscillates between 20% and 100% while requests are processing, and drops to 0% when idle
GPU Utilization in a different time window
GPU Utilization in a different time window. PP phase spikes to 100%, TG phase settles around 20–30%

CUDA0 holds 68.7 GiB of model buffers and CUDA1 holds 23.5 GiB, so the head side (CUDA0) uses more VRAM and runs hotter. CUDA1 still has 70+ GiB of headroom, which leaves room to run a coder model or other services alongside.

nvtop showing GPUs at idle
nvtop at idle. GPU0: 69,454 MiB (71%), GPU1: 24,228 MiB (25%) of VRAM in use

CPU / Host Memory

Node Exporter showing CPU/memory usage
Node Exporter. CPU runs at 60–80% during inference, memory stays around 640 GiB. That's pinned host memory 224 GiB + OS + KV cache

CPU use includes host expert tensor delivery, pinned memory, DMA transfers and runtime orchestration.

The inference environment runs on a 4-node homelab: compute.home.arpa (the GPU server) runs inference, storage.home.arpa runs PostgreSQL / Prometheus / MinIO / Dagster / MLflow, and desktop.home.arpa runs Grafana.

Orchestrator Selection: GLM-5.1 vs Qwen3.5-397B

This benchmark has a second purpose: picking a model for the resident orchestrator (grandpa) in my own agent orchestration system familiar. The two candidates are GLM-5.1 (744B-A40B) and Qwen3.5-397B-A17B.

The proposed grandpa role splits tasks, delegates to coder models, evaluates output and recovers from errors across two LLM backends. Long generation stays with the coder. The intended hot-layer footprint is roughly dev01:25/25 GB.

Spec Comparison

GLM-5.1Qwen3.5-397B
Total / Active params744B / 40B397B / 17B
Experts / Active256 / 8512 / 10
QuantizationIQ3_KS (3.65 bpw)Q4_K_M mixed (4.93 bpw)
Model size320 GiB228 GiB
n_ctx_train202,752262,144
LicenseMITApache-2.0

Test Configuration Differences

The measured placements differ, so each configuration is shown separately.

GLM-5.1 (this run): --cpu-moe + -ot brings head 15 + tail 4 layers of experts back to GPU. GPU-side experts: 19 layers ≈ 76 GiB. CPU pinned host experts: 56 layers ≈ 224 GiB (log shows CUDA_Host 229,438 MiB). This layout was tuned for a coding bench, not dedicated to orchestrator duty.

Qwen3.5-397B (separate session): With --n-cpu-moe 15, the startup log shows blk.0–blk.14 (first 15 layers) experts landing on CUDA_Host, while the remaining 45 layers (blk.15–blk.59) of experts are on the GPU side. Qwen3.5 has a hybrid structure with full_attention_interval = 4 — the attention type switches every 4 layers — but the experts themselves are held by all 60 layers (the log’s Layer sizes shows 3,839 MiB of expert content even on the Layer 3, 7, 11... rows). Non-experts (attention, SSM, dense, output) are distributed across both GPUs via graph split, reserving 176 GiB on CUDA_Split. The measured CUDA_Host is 56,710 MiB, which matches 15 × 3,712 MiB = 55,680 MiB closely.

So GLM-5.1 is running with “most experts on CPU” while Qwen3.5 is running with “only the first 15 layers of experts on CPU”. Neither is a full-CPU configuration, and the pinned-host size swings a lot based on the setup. When evaluating both for orchestrator use, it’s worth estimating the VRAM requirement for both at expert=full-CPU separately.

VRAM Requirements at Expert=Full CPU

Taking the non-expert rows from each startup log’s Layer sizes, and adding KV cache + compute buffer:

GLM-5.1Qwen3.5-397B
Non-exps weight (attn+dense+output)6,518 MiB7,058 MiB
KV per 1k ctx (q8_0)~46 MiB~16 MiB
KV cache (ctx 65k, q8_0)2,984 MiB~1,040 MiB
KV cache (ctx 200k, q8_0, est.)~9,228 MiB~3,200 MiB
KV cache (ctx 262k, q8_0)—4,080 MiB
Compute buffer~12 GiB~11 GiB
VRAM total (exps=CPU, ctx 65k)~22 GiB~19 GiB
VRAM total (exps=CPU, ctx 200k)~28 GiB~21 GiB
VRAM total (exps=CPU, ctx 262k)—~22 GiB

Qwen3.5’s attention is hybrid and KV efficient, fitting within 16 MiB/1k. GLM-5.1 has MLA-style attention, but with larger head count and dim, KV costs 46 MiB/1k. As a result, at ctx 200k GLM-5.1 runs about +7 GiB heavier.

CPU Pinned Host Difference

GLM-5.1Qwen3.5-397B
Exps per layer4,088 MiB3,712 MiB (+some 3,839 MiB)
Layers holding exps76 (blk.3–blk.78)60 (all layers)
Total exps (full-CPU, theoretical)~303 GiB~218 GiB
CPU pinned in this test229 GiB (56 layers)55 GiB (15 layers)
Pinned alloc time (test)40s9s

Inference-speed comparison

GLM-5.1Qwen3.5-397B
TG (early session)18–19 t/s55–59 t/s
TG (late session, ctx full)16–17 t/s17–18 t/s
PP max572 t/s1,500 t/s
Cache restoreN/A14–18 ms (checkpoint)
Cache missFrequent, from <think> mismatchStable via checkpoint restore

Both models approach 17–18 t/s in longer sessions. Orchestrator output is short, so decision quality also needs evaluation.

Qwen’s PP is about 3× faster and reduces re-evaluation time after cache misses. Turning thinking off may improve GLM’s <think> mismatch, but this has not been verified.

Decision quality still to test

Qwen has favorable PP, pinned-memory use, ctx 262k and cache stability. Task splitting, state management and context management still need quality checks.

GLM is active 40B; Qwen is active 17B. I expect useful reasoning from GLM, but parameter count alone does not establish quality. Resident operation needs measured delegation and recovery results. My current impression favors Qwen for operation.

RAG and long-context plans

n_ctx_train is 202k for GLM and 262k for Qwen. Assuming a 70% effective limit gives about 141k and 183k. I am considering these two operating policies.

GLM-5.1 + RAG shaping (ctx ~64k operation):

The RAG plan uses Gitea code indexed by ColBERT + maxsim reranker, with argus (Gitea symbol DB) and voracle (Obsidian vault semantic search) MCP tools returning needed context. At ctx 64k, estimated KV use is ~3 GiB, with the aim of limiting TTFT and TG degradation.

A small-ctx Mac Studio configuration with unified memory is another candidate, but has not been tested.

Qwen3.5 + power play (ctx ~183k):

The long-context plan includes history, tool results and code together. Filtering this volume with an active-17B model remains unverified. Light KV leaves capacity for YaRN at 1M ctx, but quality at that length requires separate testing.

Direction

I am currently considering Qwen3.5-397B-A17B for grandpa. GLM’s active 40B remains interesting, but Qwen’s configuration appears easier for residence alongside other models.

  1. ctx 262k and KV capacity: 16 MiB/1k leaves room for long inputs and image use cases.
  2. PP 1,500 t/s: cache-miss re-evaluation is 3× faster than GLM-5.1. This affects latency when the orchestrator frequently changes the system prompt or messages[].
  3. License and training cost: MIT / Apache-2.0 permits SFT/DPO/LoRA experiments. GLM-5.1’s size makes repeated fine-tuning costly; I consider Qwen3.5-397B-A17B more feasible to test.

Development checks

After about two months of development, using roughly two weekday slots and weekends, the main questions are small, relevant ctx, when to supply it and how models coordinate. MCP tools are being developed with claude and codex.

Remaining validation

Head-heavy GPU placement may improve hit rate, but that is unverified. GLM’s active 40B with RAG and a small ctx remains another candidate.

The next checks measure delegation success, recovery counts and session length. I plan to evaluate Qwen’s cache checkpoint restore and resident behavior while collecting data and refining tools.