Qwen3.5-397B hybrid inference and CMS generation
28 Qwen3.5-397B IQ4_NL runs on EPYC and GPUs, followed by Django CMS scaffolding. Throughput, configuration, and required corrections for LLM integration and web application development.
Qwen3.5-397B throughput and code generation
I measured 28 consecutive Qwen3.5-397B runs on EPYC and consumer GPUs. This MoE model has 397B total parameters and 17B active parameters. It normally calls for multiple H100s; this setup uses IQ4_NL quantization, cpu-moe, and tensor offloading.
The test checks whether the speed suits daily use and provides a reference for LLM integration. I also generated a project scaffold from a WordPress-like Django CMS specification to check long-context coding.
Measurement goals
- Statistically characterize steady-state TG/PP speed from 28 runs with IQ4_NL
- Document hybrid offload execution configs (cpu-moe / multi-GPU tensor offload)
- Evaluate context-length dependency and stability
- Determine daily viability of 400B-class MoE
- Evaluate long-context code generation quality (Django CMS scaffold)
Test Environment
| Item | Specification |
|---|---|
| CPU | AMD EPYC 9175F (Zen 5, 16C, L3 512MB) |
| Memory | DDR5-6400 768GB (12ch) |
| GPU | NVIDIA RTX PRO 6000 (96GB VRAM) |
| OS | Ubuntu 24.04 LTS |
| Runtime | ik_llama.cpp (cpu-moe enabled) |
| Quantization | IQ4_NL |
| Context | Up to 262,144 |
Results
Throughput Statistics (28 Runs)
| Metric | Prefill (PP) | Generation (TG) |
|---|---|---|
| Maximum | 372.24 tok/s | 24.04 tok/s |
| Minimum | 101.49 tok/s | 19.13 tok/s |
| Mean (Steady State) | ~160 tok/s | ~22.5 tok/s |
Representative Runs
| # | PP(tok) | TG(tok) | PP(tok/s) | TG(tok/s) | Total(s) |
|---|---|---|---|---|---|
| 1 | 4,699 | 13 | 314.22 | 21.95 | 15.5 |
| 3 | 1,125 | 2,048 | 161.67 | 20.81 | 105.4 |
| 8 | 1,124 | 2,048 | 154.76 | 23.82 | 93.3 |
| 14 | 15,866 | 2,048 | 372.24 | 22.98 | 131.7 |
| 16 | 550 | 520 | 117.78 | 22.91 | 27.4 |
Run #14 processed 15,866 input tokens in 42.6s and generated at 22.98 tok/s. About 15K tokens of source code took just over 40 seconds to ingest.
Context Length Dependency
- PP speed: Short prompts (<1k) stayed around 100 tok/s, dominated by overhead. Longer inputs improved parallel efficiency and exceeded 300 tok/s.
- TG speed: The measured context lengths stayed within 21-24 tok/s.
Hybrid Offload Configurations
Single GPU + cpu-moe
IMG=compute.home.arpa/ik_llama-cuda
MODEL=/models/.../IQ4_KSS/Qwen3.5-397B-A17B-IQ4_KSS-00001-of-00006.gguf
podman run --rm -it --device nvidia.com/gpu=all \
-p 8001:8080 --shm-size 16g --cap-add=SYS_NICE \
-v "$MO":/models:ro,Z "$IMG" \
--host 0.0.0.0 --port 8080 -m "$MODEL" \
-c 262144 --threads 13 --threads-batch 24 \
--jinja -b 2048 -ub 2048 -ngl 99 \
-fa on --no-mmap --cpu-moe
--cpu-moe moves expert computation beyond GPU VRAM capacity to the CPU. EPYC 9175F’s 12-channel DDR5 supplies active experts while the GPU computes Attention.
Multi-GPU Tensor Offload
With 2 GPUs, -ot specifies regex-based layer distribution:
./build/bin/llama-server \
--model "$model" \
-fa on --ctx-size 135168 \
-ctk q8_0 -ctv q8_0 \
-ub 2048 -b 2048 -ngl 999 \
-ot "blk\.(0|1|2|...|12)\.ffn_(gate|up|down)_exps.*=CUDA0,\
blk\.(47|48|...|60)\.ffn_(gate|up|down)_exps.*=CUDA1" \
--cpu-moe --threads 24 --no-mmap --jinja
Layers 0-12 go to CUDA0 and layers 47-60 to CUDA1. cpu-moe handles the middle layers. This distributes VRAM usage while retaining a 135K context.
Long-context Django CMS generation
Evaluation Setup
I gave Qwen3.5-397B-A17B a detailed English specification mapping WordPress concepts to Django. It covered Posts/Pages, Categories/Tags, Comments, Media Library, Menus/Navigation, Drafts/scheduled publishing, and Revision history.
Inference Configuration
IMG=compute.home.arpa/ik_llama-cuda
MO=/mnt/data/hf/hub/models--ubergarm--Qwen3.5-397B-A17B-GGUF
MODEL=/models/snapshots/.../IQ4_KSS/Qwen3.5-397B-A17B-IQ4_KSS-00001-of-00006.gguf
podman run --rm -it --device nvidia.com/gpu=all \
-p 8001:8080 --shm-size 16g --cap-add=SYS_NICE \
-v "$MO":/models:ro,Z "$IMG" \
--host 0.0.0.0 --port 8080 -m "$MODEL" \
-c 262144 --threads 13 --threads-batch 26 \
--jinja --temp 0.7 --repeat-penalty 1.2 \
--min-p 0.01 --top-p 0.95 --seed 317 --top-k 40 \
-b 2048 -ub 2048 -ngl 99 \
-fa on --no-mmap -ger
The 262k context setting kept the specification and generated code in one session.
Concept mapping and excluded features
A mapping table at the start helped keep field design consistent: Post/Page -> content.Post/content.Page, Category/Tag -> taxonomy.Term, Post Meta -> content.PostMeta, and Revision -> content.Revision.
The non-goals were WordPress compatibility, Gutenberg replication, multisite UI parity, and full WYSIWYG. Stating them helped avoid unnecessary architecture.
Generated Output
The generation log records these files:
apps/core/models.py/models_site.py/models_setting.pyapps/content/models.py/models_page.py/models_extra.pyapps/taxonomy/models.pyapps/media/models.pyapps/comments/models.pyapps/navigation/models.pyapps/seo/models.pyconfig/settings/base.py/dev.py/prod.pyconfig/urls.py/asgi.py/wsgi.pymanage.py/requirements.txt
The output included the project scaffold and model definitions for the main Django apps.
Usable model definitions
- Specification adherence: Post, Term, Revision, SEOEntry scaffolds were directly usable
- Concept mapping: WordPress mental model preserved while remaining Django-native
- Abstract models: TimeStampedModel, SoftDeleteModel, PublishableModel, SluggedModel generated naturally
- SEO constraints: SEOEntry included CheckConstraint binding to either post or page
Required corrections
- Field dropout: slug disappeared from Post mid-generation, had to be re-added via diff
- File placement drift: Page was first split to
models_page.py, then folded back intomodels.py - Field inconsistency: Comment.save() referenced
self.author_ipbefore the field existed in the model - Tool failures: ctree_check, ctree_init, filesystem_create_directory failed repeatedly before falling back to write_file
Review scope
Tool failures, missing parent directories, model-definition gaps, and inconsistent fields recurred during generation.
Fast scaffolding was useful, but the output still needed review to separate usable code from required corrections.
Analysis
The configuration behind 22 tok/s
Only 17B of the 397B parameters are active per token. IQ4_NL reduces the memory-bandwidth load by 4x+. The 12-channel DDR5 supplies experts while the GPU computes Attention. This division produced about 22 tok/s.
IQ4_NL Quantization Choice
Unquantized 397B deployment was impractical on this hardware, so I chose IQ4_NL. I saw no obvious quality loss across 28 runs; this was not a quantitative quality comparison.
Warm-up Requirement
The first few runs varied. 2-3 dummy requests helped reach stable TG speed. A resident service could include this warm-up at startup.
Daily-use assessment
22.5 tok/s was sufficient for the tested interactive use. I also consider it usable for asynchronous batches.
With about 15K tokens of source code, TG remained at 21-24 tok/s. This is a useful reference for longer repository inputs.
The CMS test produced an initial scaffold. Concept mappings, explicit non-goals, and implementation order helped keep the output aligned with the specification.
cpu-moe worked alone. With 2 GPUs, I used explicit layer distribution through -ot to improve throughput.
Launch settings and future records
KV Cache Quantization
KV cache grows at 135K-262K context lengths. -ctk q8_0 -ctv q8_0 reduced its memory use. The perceived quality impact was small; it was not measured quantitatively.
Sampling Parameters
For code generation I used --temp 0.7 --repeat-penalty 1.2 --min-p 0.01 --top-p 0.95 --top-k 40. Temp 1.0 tended toward more creative but longer output.
Separate launch settings, specification, and review
For the next run, I would keep three artifacts:
- A reusable inference launch template
- A clean CMS specification (with concept mapping + non-goals + implementation order)
- A generation log plus review notes
The CMS specification should contain the concept dictionary, non-goals, and implementation order. Keeping generation logs and review notes separate makes adherence and corrections easier to inspect.
