A Gemma 4 inference stack

I configured Gemma 4 31B dense and 26B-A4B MoE for homelab inference, together with familiar roles, a separate Dagster pipeline, and telemetry. The setup explores parallel model operation for LLM integration.


Hardware: EPYC + Dual Blackwell

Compute node specs:

  • CPU: AMD EPYC 9175F
  • RAM: 768GB
  • GPU: RTX 6000 PRO MAX-Q 96GB Blackwell x 2 (no NVLink)

I used the GPUs independently. With no NVLink, this setup avoided Tensor Parallel and placed separate models on each GPU to increase parallel slots. The initial goal was data collection throughput.


Why Gemma 4

I limited model variety during startup to reduce prompt and format debugging. Gemma 4 31B dense and 26B-A4B MoE had suitable scores and fit in memory.

Model diversity could follow once familiar and the data pipeline were stable. The first step was validating one model family.


Model Allocation and GPU Memory Design

slotmodelformatruntimeVRAMparallelGPU
naughtyGemma 4 31B ITNVFP4vLLM~17GiB3GPU0
kindergartenGemma 4 26B-A4B ITIQ4_XSllama.cpp~13GiB5GPU1

GPU memory allocation:

  • GPU0 (96GB): Gemma 4 31B NVFP4 ~17GiB + KV cache x 3 parallel slots (~79GB free)
  • GPU1 (96GB): Gemma 4 26B-A4B IQ4_XS ~13GiB + KV cache x 6 parallel slots (~83GB free)

naughty: Gemma 4 31B NVFP4 on vLLM

NVFP4 uses the TensorRT Model Optimizer format and Blackwell FP4 kernels. vLLM continuous batching runs parallel 3, with thinking mode, native function calling, and system prompts.

Benchmarks:

  • AIME 2026: 89.2%
  • LiveCodeBench v6: 80.0%
  • Codeforces ELO: 2150
  • GPQA Diamond: 84.3%
  • MMLU Pro: 85.2%

The 31B dense model has an AIME score of 89.2%. Its roughly 17GiB NVFP4 footprint leaves room for KV cache on a 96GB GPU.

kindergarten: Gemma 4 26B-A4B MoE on llama.cpp

26B-A4B is a Mixture of Experts model with 4B active parameters. Bartowski IQ4_XS uses roughly 13GiB, leaving room for more parallel slots.

I used ik_llama.cpp for GGUF MoE inference and vLLM for dense-model continuous batching.


familiar Model Role Design

familiar has four roles: grandpa / naughty / kindergarten / translator, selected through CLAUDE.md.

  grandpa x naughty x kindergarten x translator    # full config
naughty x translator                              # minimal config
naughty x kindergarten x translator               # standard config
  

grandpa prioritizes quality and balances token/sec. For initial collection throughput, I omitted it and used naughty x kindergarten x translator.

translator uses plamo-2-translate for Japanese-English translation in the multilingual article pipeline.


model-foundry: Dagster Pipeline Separation

I also moved Dagster into its own repository.

The pipeline moved from agent-gateway and devstack/dagster/ to model-foundry. It had become independent data infrastructure rather than a gateway addon.

DuckDB and the Dagster UI

DuckDB made execution status visible together.

Previously, each asset needed inspection. Now pipeline_event_inbox_record STEP_OUTPUT shows event_id, subject, and correlation_id as structured dicts.

A type check error appeared right after DuckDB introduction:

  dagster._core.errors.DagsterTypeCheckDidNotPass: Type check failed for step input
"pipeline_event_inbox_record" - expected type "dict".
Value of type <class 'NoneType'> failed type check
  

The asset returned None when the NATS queue was empty. I fixed empty-event handling.

Event Name Redesign

Also cleaned up knowledge domain pipeline event names:

Old nameNew name
knowledge.chat.persistllm.chat.persist
knowledge.embeddingobsidian.semantic_search
knowledge.flow.lineagellm.flow.lineage
knowledge.tool_callllm.tool_call

I changed knowledge.* topics from internal domain names to target-system names. The later v3 redesign separated three domains.


Grafana LLM Dashboard

Grafana displays inference telemetry.

Data Flow

  familiar -> Vector (HTTP port 8687 /telemetry) -> Prometheus / Loki -> Grafana
  

Rendering principles:

  • Grafana renders from real-time data via Vector (does not query PostgreSQL directly)
  • Persistence goes flat into PostgreSQL
  • Summarization needs are handled by Dagster + DuckDB

Vector http_telemetry receives familiar events, sends metrics to Prometheus and logs to Loki, and uses correlation_id for session tracing.

Vector configuration:

  [sources.http_telemetry]
type = "http_server"
address = "0.0.0.0:8687"
path = "/telemetry"
encoding = "json"
  

familiar POSTs JSON events. Vector transforms and forwards them; Grafana combines Prometheus and Loki to show performance and session flow.


Parallel coding results

I configured vLLM nvidia/gemma-4-31b-it-nvfp4 with --max-num-seqs 3 and the ik_llama.cpp gemma4 branch with Q4_K_L and --parallel 5. Go channels controlled the queues.

Gemma 4 31B session flow on Grafana dashboard
Gemma 4 31B naughty x 3 parallel execution flow -- 584 Edge Events, 702 Node Events. Model name gemma-4-31b-it highlighted in raw telemetry

The orchestrator never routed to 26B MoE (kindergarten), so the actual test covered only 31B IT x 3. It produced deliverables, but workers did not follow a shared direction.

Sharing KV cache did not make the outputs align as I had expected. Even with the same plan, the three workers acted independently and sometimes ignored each other’s decisions.

For coding, this led me to use a generalist orchestrator and constrained worker prompts and context. Parallelism can increase collection throughput, but coordinated output also needs orchestration.

After That: GLM-5.1 + Qwen3-Coder-Next Configuration

I then replaced Kimi-K2.5 in the grandpa role with GLM-5.1 and paired it with Qwen3-Coder-Next.

  dev0 / GLM-5.1 (grandpa)
  PP:  340 tok/s
  TG:  12-16 tok/s

dev0 / Qwen3-Coder-Next (worker-0)
  PP:  2500 tok/s
  TG:  114 tok/s

dev1 / Qwen3-Coder-Next (worker-1)
  PP:  5100 tok/s
  TG:  161 tok/s

dev1 / PLaMo2Translate (translator)
  PP:  100-240 tok/s
  TG:  35-140 tok/s
  

A reasoning-focused orchestrator and coding workers coordinated better here than the all-Gemma 4 setup.