Moving llm-jp translation to on-demand batches
Single-GPU NVFP4 measurements led to replacing a resident llm-jp translator with batches. The record covers LLM resource allocation, throughput, a dual-GPU startup error, and translation configuration proposals.
I was preparing SFT/DPO data and LoRA for llm-jp/llm-jp-4-32b-a3b-base translation. Single-GPU llm-jp-4-32b-a3b-base-NVFP4 had faster PP/TG than expected, so I switched from a resident translator to on-demand batches. This leaves more VRAM for other LLM workloads.
Model provider: llm-jp on Hugging Face
Command Used
The validation launch command:
podman run --rm -it \
--device nvidia.com/gpu=0 \
--ipc=host \
-e HF_HOME=/hf \
-e HF_HUB_OFFLINE=1 \
-v /mnt/data/models:/hf/hub:ro \
-p 9000:9000 \
registry.home.arpa/vllm-openai:v0.18.0-cu130 \
/hf/hub/llm-jp-4-32b-a3b-base-NVFP4 \
--host 0.0.0.0 \
--port 9000 \
--trust-remote-code \
--quantization compressed-tensors \
--served-model-name translator \
--dtype bfloat16 \
--gpu-memory-utilization 0.80 \
--max-num-seqs 8 \
--max-model-len 32678 \
--max-num-batched-tokens 1024 \
--no-enable-prefix-caching
I disabled prefix caching because fixed-prefix reuse was limited for this translation workload.
Startup logs
Startup arguments:
(APIServer pid=1) INFO 04-14 08:08:26 [utils.py:233] non-default args: {'model_tag': '/hf/hub/llm-jp-4-32b-a3b-base-NVFP4', 'host': '0.0.0.0', 'port': 9000, 'model': '/hf/hub/llm-jp-4-32b-a3b-base-NVFP4', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_model_len': 32678, 'quantization': 'compressed-tensors', 'served_model_name': ['translator'], 'gpu_memory_utilization': 0.8, 'enable_prefix_caching': False, 'max_num_batched_tokens': 1024, 'max_num_seqs': 8}
Model loading and initialization:
(EngineCore pid=118) INFO 04-14 08:08:35 [gpu_model_runner.py:4481] Starting to load model /hf/hub/llm-jp-4-32b-a3b-base-NVFP4...
(EngineCore pid=118) INFO 04-14 08:08:39 [default_loader.py:384] Loading weights took 3.32 seconds
(EngineCore pid=118) INFO 04-14 08:08:39 [gpu_model_runner.py:4566] Model loading took 18.23 GiB memory and 3.843664 seconds
(EngineCore pid=118) INFO 04-14 08:08:52 [monitor.py:48] torch.compile took 12.35 s in total
(EngineCore pid=118) INFO 04-14 08:08:52 [monitor.py:76] Initial profiling/warmup run took 0.51 s
(EngineCore pid=118) INFO 04-14 08:09:29 [gpu_worker.py:456] Available KV cache memory: 56.62 GiB
(EngineCore pid=118) INFO 04-14 08:09:29 [kv_cache_utils.py:1316] GPU KV cache size: 927,600 tokens
(EngineCore pid=118) INFO 04-14 08:09:29 [kv_cache_utils.py:1321] Maximum concurrency for 32,678 tokens per request: 28.38x
Single and Batch Measurements
Single request:
(APIServer pid=1) INFO 04-14 08:11:32 [loggers.py:259] Engine 000: Avg prompt throughput: 12.5 tokens/s, Avg generation throughput: 25.6 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
Representative batch logs (seq=8, ctx=32768):
(APIServer pid=1) INFO 04-14 07:34:20 [loggers.py:259] Engine 000: Avg prompt throughput: 353.8 tokens/s, Avg generation throughput: 157.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.4%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO: 10.0.2.100:41472 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=1) INFO: 10.0.2.100:41484 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=1) INFO: 10.0.2.100:36782 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=1) INFO: 10.0.2.100:36788 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=1) INFO: 10.0.2.100:36790 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=1) INFO 04-14 07:34:30 [loggers.py:259] Engine 000: Avg prompt throughput: 2768.4 tokens/s, Avg generation throughput: 159.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.0%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 04-14 07:34:40 [loggers.py:259] Engine 000: Avg prompt throughput: 276.4 tokens/s, Avg generation throughput: 182.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.5%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 04-14 07:34:50 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 182.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.7%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 04-14 07:35:00 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 179.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.9%, Prefix cache hit rate: 0.0%
Observed ranges:
- Prompt throughput: around
350-1360 tok/s(with an observed instant peak of2768.4 tok/s) - Generation throughput: around
152-183 tok/s - Prefix cache hit rate: always
0.0%(consistent with--no-enable-prefix-caching)
This speed supports on-demand translation batches. max_num_seqs=8 is conservative; short and medium inputs may allow more concurrency.
Decode speed during long generation
Decode throughput gradually drops during long generations.
183.2 -> 180.5 -> 178.2 -> 176.4 -> 173.3 -> 170.0 -> 169.3 -> 167.9
169.7 -> 168.9 -> 167.2 -> 165.5 -> 163.7 -> 160.9 -> 159.4 -> 158.0 -> 156.3 -> 154.7 -> 152.2
Moving to on-demand batches
The measured speed and stability led to this configuration:
- Remove the
translatorrole from my custom harness - Trigger translation batches from Dagster pipeline events
- Reserve VRAM for resident inference roles
I had planned translation-specific SFT/DPO and LoRA. Measured PP/TG reduced the need for a resident role. Unless translated output must immediately feed retraining, batches can handle the work and leave VRAM for other inference.
Next Steps
Specializing llm-jp-4-32b-a3b-base for translation may take more work. I plan to test thinking:low next.
Assuming reasoning: low|middle|high can be selected by use case, I plan this two-stage flow:
- First pass to strict Japanese with
gemmatranslate4b-it - Fluency pass with
llm-jp-32b-a3b-thinking:low
If thinking:low or LoRA adaptation remains difficult, I still want to test the practical range of a two-stage setup with gemmatranslate-4b-it.
Single- and dual-GPU results
llm-jp-4-32b-a3b worked with single-GPU NVFP4. --tensor-parallel-size 2 failed with Intermediate size padding for w1 and w3 ... not currently supported. Non-auto settings such as --moe-backend cutlass did not resolve it. Two instances are a possible workaround; I left that branch without further investigation.
The operating plan is now on-demand translation batches.
