Qwen3.5-397B throughput and code generation

I measured 28 consecutive Qwen3.5-397B runs on EPYC and consumer GPUs. This MoE model has 397B total parameters and 17B active parameters. It normally calls for multiple H100s; this setup uses IQ4_NL quantization, cpu-moe, and tensor offloading.

The test checks whether the speed suits daily use and provides a reference for LLM integration. I also generated a project scaffold from a WordPress-like Django CMS specification to check long-context coding.

Measurement goals

  1. Statistically characterize steady-state TG/PP speed from 28 runs with IQ4_NL
  2. Document hybrid offload execution configs (cpu-moe / multi-GPU tensor offload)
  3. Evaluate context-length dependency and stability
  4. Determine daily viability of 400B-class MoE
  5. Evaluate long-context code generation quality (Django CMS scaffold)

Test Environment

ItemSpecification
CPUAMD EPYC 9175F (Zen 5, 16C, L3 512MB)
MemoryDDR5-6400 768GB (12ch)
GPUNVIDIA RTX PRO 6000 (96GB VRAM)
OSUbuntu 24.04 LTS
Runtimeik_llama.cpp (cpu-moe enabled)
QuantizationIQ4_NL
ContextUp to 262,144

Results

Throughput Statistics (28 Runs)

MetricPrefill (PP)Generation (TG)
Maximum372.24 tok/s24.04 tok/s
Minimum101.49 tok/s19.13 tok/s
Mean (Steady State)~160 tok/s~22.5 tok/s

Representative Runs

#PP(tok)TG(tok)PP(tok/s)TG(tok/s)Total(s)
14,69913314.2221.9515.5
31,1252,048161.6720.81105.4
81,1242,048154.7623.8293.3
1415,8662,048372.2422.98131.7
16550520117.7822.9127.4

Run #14 processed 15,866 input tokens in 42.6s and generated at 22.98 tok/s. About 15K tokens of source code took just over 40 seconds to ingest.

Context Length Dependency

  • PP speed: Short prompts (<1k) stayed around 100 tok/s, dominated by overhead. Longer inputs improved parallel efficiency and exceeded 300 tok/s.
  • TG speed: The measured context lengths stayed within 21-24 tok/s.

Hybrid Offload Configurations

Single GPU + cpu-moe

  IMG=compute.home.arpa/ik_llama-cuda
MODEL=/models/.../IQ4_KSS/Qwen3.5-397B-A17B-IQ4_KSS-00001-of-00006.gguf

podman run --rm -it --device nvidia.com/gpu=all \
  -p 8001:8080 --shm-size 16g --cap-add=SYS_NICE \
  -v "$MO":/models:ro,Z "$IMG" \
  --host 0.0.0.0 --port 8080 -m "$MODEL" \
  -c 262144 --threads 13 --threads-batch 24 \
  --jinja -b 2048 -ub 2048 -ngl 99 \
  -fa on --no-mmap --cpu-moe
  

--cpu-moe moves expert computation beyond GPU VRAM capacity to the CPU. EPYC 9175F’s 12-channel DDR5 supplies active experts while the GPU computes Attention.

Multi-GPU Tensor Offload

With 2 GPUs, -ot specifies regex-based layer distribution:

  ./build/bin/llama-server \
  --model "$model" \
  -fa on --ctx-size 135168 \
  -ctk q8_0 -ctv q8_0 \
  -ub 2048 -b 2048 -ngl 999 \
  -ot "blk\.(0|1|2|...|12)\.ffn_(gate|up|down)_exps.*=CUDA0,\
       blk\.(47|48|...|60)\.ffn_(gate|up|down)_exps.*=CUDA1" \
  --cpu-moe --threads 24 --no-mmap --jinja
  

Layers 0-12 go to CUDA0 and layers 47-60 to CUDA1. cpu-moe handles the middle layers. This distributes VRAM usage while retaining a 135K context.

Long-context Django CMS generation

Evaluation Setup

I gave Qwen3.5-397B-A17B a detailed English specification mapping WordPress concepts to Django. It covered Posts/Pages, Categories/Tags, Comments, Media Library, Menus/Navigation, Drafts/scheduled publishing, and Revision history.

Inference Configuration

  IMG=compute.home.arpa/ik_llama-cuda
MO=/mnt/data/hf/hub/models--ubergarm--Qwen3.5-397B-A17B-GGUF
MODEL=/models/snapshots/.../IQ4_KSS/Qwen3.5-397B-A17B-IQ4_KSS-00001-of-00006.gguf

podman run --rm -it --device nvidia.com/gpu=all \
  -p 8001:8080 --shm-size 16g --cap-add=SYS_NICE \
  -v "$MO":/models:ro,Z "$IMG" \
  --host 0.0.0.0 --port 8080 -m "$MODEL" \
  -c 262144 --threads 13 --threads-batch 26 \
  --jinja --temp 0.7 --repeat-penalty 1.2 \
  --min-p 0.01 --top-p 0.95 --seed 317 --top-k 40 \
  -b 2048 -ub 2048 -ngl 99 \
  -fa on --no-mmap -ger
  

The 262k context setting kept the specification and generated code in one session.

Concept mapping and excluded features

A mapping table at the start helped keep field design consistent: Post/Page -> content.Post/content.Page, Category/Tag -> taxonomy.Term, Post Meta -> content.PostMeta, and Revision -> content.Revision.

The non-goals were WordPress compatibility, Gutenberg replication, multisite UI parity, and full WYSIWYG. Stating them helped avoid unnecessary architecture.

Generated Output

The generation log records these files:

  • apps/core/models.py / models_site.py / models_setting.py
  • apps/content/models.py / models_page.py / models_extra.py
  • apps/taxonomy/models.py
  • apps/media/models.py
  • apps/comments/models.py
  • apps/navigation/models.py
  • apps/seo/models.py
  • config/settings/base.py / dev.py / prod.py
  • config/urls.py / asgi.py / wsgi.py
  • manage.py / requirements.txt

The output included the project scaffold and model definitions for the main Django apps.

Usable model definitions

  • Specification adherence: Post, Term, Revision, SEOEntry scaffolds were directly usable
  • Concept mapping: WordPress mental model preserved while remaining Django-native
  • Abstract models: TimeStampedModel, SoftDeleteModel, PublishableModel, SluggedModel generated naturally
  • SEO constraints: SEOEntry included CheckConstraint binding to either post or page

Required corrections

  • Field dropout: slug disappeared from Post mid-generation, had to be re-added via diff
  • File placement drift: Page was first split to models_page.py, then folded back into models.py
  • Field inconsistency: Comment.save() referenced self.author_ip before the field existed in the model
  • Tool failures: ctree_check, ctree_init, filesystem_create_directory failed repeatedly before falling back to write_file

Review scope

Tool failures, missing parent directories, model-definition gaps, and inconsistent fields recurred during generation.

Fast scaffolding was useful, but the output still needed review to separate usable code from required corrections.

Analysis

The configuration behind 22 tok/s

Only 17B of the 397B parameters are active per token. IQ4_NL reduces the memory-bandwidth load by 4x+. The 12-channel DDR5 supplies experts while the GPU computes Attention. This division produced about 22 tok/s.

IQ4_NL Quantization Choice

Unquantized 397B deployment was impractical on this hardware, so I chose IQ4_NL. I saw no obvious quality loss across 28 runs; this was not a quantitative quality comparison.

Warm-up Requirement

The first few runs varied. 2-3 dummy requests helped reach stable TG speed. A resident service could include this warm-up at startup.

Daily-use assessment

22.5 tok/s was sufficient for the tested interactive use. I also consider it usable for asynchronous batches.

With about 15K tokens of source code, TG remained at 21-24 tok/s. This is a useful reference for longer repository inputs.

The CMS test produced an initial scaffold. Concept mappings, explicit non-goals, and implementation order helped keep the output aligned with the specification.

cpu-moe worked alone. With 2 GPUs, I used explicit layer distribution through -ot to improve throughput.

Launch settings and future records

KV Cache Quantization

KV cache grows at 135K-262K context lengths. -ctk q8_0 -ctv q8_0 reduced its memory use. The perceived quality impact was small; it was not measured quantitatively.

Sampling Parameters

For code generation I used --temp 0.7 --repeat-penalty 1.2 --min-p 0.01 --top-p 0.95 --top-k 40. Temp 1.0 tended toward more creative but longer output.

Separate launch settings, specification, and review

For the next run, I would keep three artifacts:

  1. A reusable inference launch template
  2. A clean CMS specification (with concept mapping + non-goals + implementation order)
  3. A generation log plus review notes

The CMS specification should contain the concept dictionary, non-goals, and implementation order. Keeping generation logs and review notes separate makes adherence and corrections easier to inspect.