I was preparing SFT/DPO data and LoRA for llm-jp/llm-jp-4-32b-a3b-base translation. Single-GPU llm-jp-4-32b-a3b-base-NVFP4 had faster PP/TG than expected, so I switched from a resident translator to on-demand batches. This leaves more VRAM for other LLM workloads.

Model provider: llm-jp on Hugging Face

Command Used

The validation launch command:

  podman run --rm -it \
  --device nvidia.com/gpu=0 \
  --ipc=host \
  -e HF_HOME=/hf \
  -e HF_HUB_OFFLINE=1 \
  -v /mnt/data/models:/hf/hub:ro \
  -p 9000:9000 \
  registry.home.arpa/vllm-openai:v0.18.0-cu130 \
  /hf/hub/llm-jp-4-32b-a3b-base-NVFP4 \
  --host 0.0.0.0 \
  --port 9000 \
  --trust-remote-code \
  --quantization compressed-tensors \
  --served-model-name translator \
  --dtype bfloat16 \
  --gpu-memory-utilization 0.80 \
  --max-num-seqs 8 \
  --max-model-len 32678 \
  --max-num-batched-tokens 1024 \
  --no-enable-prefix-caching
  

I disabled prefix caching because fixed-prefix reuse was limited for this translation workload.

Startup logs

Startup arguments:

  (APIServer pid=1) INFO 04-14 08:08:26 [utils.py:233] non-default args: {'model_tag': '/hf/hub/llm-jp-4-32b-a3b-base-NVFP4', 'host': '0.0.0.0', 'port': 9000, 'model': '/hf/hub/llm-jp-4-32b-a3b-base-NVFP4', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_model_len': 32678, 'quantization': 'compressed-tensors', 'served_model_name': ['translator'], 'gpu_memory_utilization': 0.8, 'enable_prefix_caching': False, 'max_num_batched_tokens': 1024, 'max_num_seqs': 8}
  

Model loading and initialization:

  (EngineCore pid=118) INFO 04-14 08:08:35 [gpu_model_runner.py:4481] Starting to load model /hf/hub/llm-jp-4-32b-a3b-base-NVFP4...
(EngineCore pid=118) INFO 04-14 08:08:39 [default_loader.py:384] Loading weights took 3.32 seconds
(EngineCore pid=118) INFO 04-14 08:08:39 [gpu_model_runner.py:4566] Model loading took 18.23 GiB memory and 3.843664 seconds
(EngineCore pid=118) INFO 04-14 08:08:52 [monitor.py:48] torch.compile took 12.35 s in total
(EngineCore pid=118) INFO 04-14 08:08:52 [monitor.py:76] Initial profiling/warmup run took 0.51 s
(EngineCore pid=118) INFO 04-14 08:09:29 [gpu_worker.py:456] Available KV cache memory: 56.62 GiB
(EngineCore pid=118) INFO 04-14 08:09:29 [kv_cache_utils.py:1316] GPU KV cache size: 927,600 tokens
(EngineCore pid=118) INFO 04-14 08:09:29 [kv_cache_utils.py:1321] Maximum concurrency for 32,678 tokens per request: 28.38x
  

Single and Batch Measurements

Single request:

  (APIServer pid=1) INFO 04-14 08:11:32 [loggers.py:259] Engine 000: Avg prompt throughput: 12.5 tokens/s, Avg generation throughput: 25.6 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
  

Representative batch logs (seq=8, ctx=32768):

  (APIServer pid=1) INFO 04-14 07:34:20 [loggers.py:259] Engine 000: Avg prompt throughput: 353.8 tokens/s, Avg generation throughput: 157.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.4%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO:     10.0.2.100:41472 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=1) INFO:     10.0.2.100:41484 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=1) INFO:     10.0.2.100:36782 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=1) INFO:     10.0.2.100:36788 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=1) INFO:     10.0.2.100:36790 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=1) INFO 04-14 07:34:30 [loggers.py:259] Engine 000: Avg prompt throughput: 2768.4 tokens/s, Avg generation throughput: 159.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.0%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 04-14 07:34:40 [loggers.py:259] Engine 000: Avg prompt throughput: 276.4 tokens/s, Avg generation throughput: 182.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.5%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 04-14 07:34:50 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 182.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.7%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 04-14 07:35:00 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 179.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.9%, Prefix cache hit rate: 0.0%
  

Observed ranges:

  • Prompt throughput: around 350-1360 tok/s (with an observed instant peak of 2768.4 tok/s)
  • Generation throughput: around 152-183 tok/s
  • Prefix cache hit rate: always 0.0% (consistent with --no-enable-prefix-caching)

This speed supports on-demand translation batches. max_num_seqs=8 is conservative; short and medium inputs may allow more concurrency.

Decode speed during long generation

Decode throughput gradually drops during long generations.

  183.2 -> 180.5 -> 178.2 -> 176.4 -> 173.3 -> 170.0 -> 169.3 -> 167.9
169.7 -> 168.9 -> 167.2 -> 165.5 -> 163.7 -> 160.9 -> 159.4 -> 158.0 -> 156.3 -> 154.7 -> 152.2
  

Moving to on-demand batches

The measured speed and stability led to this configuration:

  1. Remove the translator role from my custom harness
  2. Trigger translation batches from Dagster pipeline events
  3. Reserve VRAM for resident inference roles

I had planned translation-specific SFT/DPO and LoRA. Measured PP/TG reduced the need for a resident role. Unless translated output must immediately feed retraining, batches can handle the work and leave VRAM for other inference.

Next Steps

Specializing llm-jp-4-32b-a3b-base for translation may take more work. I plan to test thinking:low next.

Assuming reasoning: low|middle|high can be selected by use case, I plan this two-stage flow:

  1. First pass to strict Japanese with gemmatranslate4b-it
  2. Fluency pass with llm-jp-32b-a3b-thinking:low

If thinking:low or LoRA adaptation remains difficult, I still want to test the practical range of a two-stage setup with gemmatranslate-4b-it.

Single- and dual-GPU results

llm-jp-4-32b-a3b worked with single-GPU NVFP4. --tensor-parallel-size 2 failed with Intermediate size padding for w1 and w3 ... not currently supported. Non-auto settings such as --moe-backend cutlass did not resolve it. Two instances are a possible workaround; I left that branch without further investigation.

The operating plan is now on-demand translation batches.