Qwen3.6-27B inference and role-specific LoRA plans
A one-day Qwen3.6-27B-FP8 run on SGLang measured about 100 tok/s singly and 160–180 tok/s combined with two streams. The plan for a local coding assistant covers LoRA for four roles and evaluation.
I ran Qwen3.6-27B-FP8 as an agent worker for one day. It measured about 100 tok/s on one stream, 160–180 tok/s combined on two, and EAGLE accept rates above 0.8. It felt easier to use than Qwen3-Coder-Next 80B, but longer quality comparisons remain needed.
These results inform a plan for role-specific LoRA workers.
During the video recording, I did not include the dedicated system prompt I normally use in production. That may be why the outputs fluctuated more than usual and why it slipped into revision loops a few times.
Video link: https://www.youtube.com/watch?v=H5w4zBDmv2g
Benchmark Results
| Item | Value |
|---|---|
| TG sustained, single stream | ~100 tok/s |
| TG combined, 2 concurrent | 160–180 tok/s |
| PP chunked (8k) | 4.0k–5.3k tok/s |
| EAGLE accept rate | 0.80–1.00 |
| EAGLE accept len | 3.0–4.0 |
| VRAM breakdown | weights 28.5GB + KV/Mamba 28.5GB + MTP 7GB |
Hardware
| Item | Specification |
|---|---|
| GPU | NVIDIA RTX PRO 6000 Blackwell Max-Q 96GB |
| CPU | AMD EPYC 9175F |
| RAM | 768GB DDR5-6400 |
| Inference engine | SGLang nightly (CUDA 13) |
SGLang Launch Configuration
SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 \
SGLANG_ENABLE_SPEC_V2=1 \
sglang serve \
--model-path Qwen/Qwen3.6-27B-FP8 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--speculative-algorithm EAGLE \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--mamba-scheduler-strategy extra_buffer \
--page-size 64 \
--mem-fraction-static 0.9 \
--context-length 262144 \
--served-model-name frisky
SGLang version issue
SGLang 0.5.10.post1 does not include the Qwen3.6 model class (qwen3_6.py). It falls back to the Qwen3.5 code path (Qwen3_5ForConditionalGeneration), so Qwen3.6-specific Gated Delta Networks (GDN) layers are not handled correctly and the thinking trace collapses from the very first turn.
The visible symptom is that the thinking output gets trapped in loops such as weak weak 弱的 弱的 weakest or atomic atomicAtomic atomicatomic. The engine itself keeps emitting tokens normally, but the content is completely broken.
Updating to nightly-dev-cu13-20260424 fixed it. That build includes qwen3_6.py, so the GDN layers are handled correctly.
On Blackwell GPUs, DeepGemm also emits a scale_fmt warning (not ue8m0). That means the FP8 checkpoint scale format is not compatible with Blackwell’s DeepGemm-optimized path. So far I have not observed a real accuracy penalty, but it is something to keep watching.
Role-specific LoRA plan
Tasks go to roles with their own system prompt, tool policy and expected behavior. The plan uses prompts and LoRA to distinguish coders, reviewers and testers.
Structured execution logs and results will become role-specific training samples. Judge pass/fail decisions are candidates for DPO data.
I plan four adapters trained on role-filtered data. SGLang supports per-request dynamic LoRA loading; switching overhead was not measured here.
DPO evaluation plan
The proposed evaluation uses a local MIT-licensed judge and an external frontier model to calibrate a subset. Agreement produces SFT candidates; disagreement prompts an accepted/rejected comparison for DPO pairs.
The intended cycle improves adapters, output and training data. Its effectiveness is not verified yet.
Training Environment
I plan to test 27B LoRA and small searches on two local 96GB GPUs. Full-parameter training requires more memory and time, so cloud expansion would follow local evaluation if needed.
The plan uses Axolotl for training and Dagster for turning logs into datasets and managing evaluation.
Why This Model
Capacity. The 27B FP8 model fits within 96GB with KV cache, Mamba state and MTP weights. Four workers are another candidate; this measurement tested two streams.
Architecture. The GDN hybrid handled tool calls, multi-step reasoning and code generation. LoRA would adjust these behaviors by role.
License. Apache 2.0 makes it a candidate for LLM integration with custom tools and adapters, including commercial applications.
Next Steps
The near-term plan is to stabilize the four-role adapter pipeline, run local small-batch training, and compare the results against the base model with A/B evaluations. Anything that looks promising can then be scaled further with cloud GPUs when needed.
The long-term goal is to reuse execution data for training specialized roles and reduce dependence on external API prices and availability.
The original stack used Claude CLI subprocesses (claude -p -r). Following Anthropic’s April 2026 usage-guidance update, I moved toward OSS models.
Agent revisions, ongoing LoRA generation and larger training remain planned work.
