Hermes-4.3-36B in BF16, FP8 and nvfp4
Hermes-4.3-36B was compared in three formats on Blackwell 96GB and vLLM 0.14.0rc1. The LLM integration check covers generation speed, initial response, context capacity and development-assistance output.
Three formats compared
I compared BF16, FP8 and nvfp4 Hermes-4.3-36B on vLLM 0.14.0rc1 and RTX PRO 6000 Blackwell Max-Q 96GB. The checks cover throughput, initial response, context capacity and coding output.
The aim is to choose configurations for everyday chat and exploration, then final edits and review.
Uses and comparison criteria
For local LLM workflows involving chat, code generation, and MCP tool integration, quantization level selection is a recurring decision. NousResearch Hermes-4.3-36B is a 36B-class model strong in tool use (Function Calling), evaluated as a vLLM candidate.
The target workload is broader than casual prompting. I am assuming:
- Interactive chat where TTFT noticeably affects usability
- Code generation and MCP-assisted workflows
- Future use with context7
Because of that, I treated the following as the real decision criteria:
- Generation throughput
- TTFT, which is strongly affected by prefill behavior
- Context headroom, especially via KV cache usage
- Perceived quality and stability for the intended workload
BF16 runs but often uses over 90% of VRAM. nvfp4 uses around 22GB. Speed and output quality are evaluated separately.
Test Environment
| Item | Specification |
|---|---|
| GPU | NVIDIA RTX PRO 6000 Blackwell Max-Q 96GB |
| CPU | AMD EPYC 9175F |
| Memory | DDR5-6400 768GB |
| Runtime | vLLM 0.14.0rc1 |
| Model | NousResearch / Hermes-4.3-36B |
With the model fitting on this GPU, I focused on response time and output quality.
Quantization Patterns I Tested
The results combine measurements with impressions from these tasks. Quality judgments are subjective.
1. BF16 (Unquantized)
BF16 is the quality-comparison baseline.
Example
vllm serve NousResearch/Hermes-4.3-36B \
--dtype bfloat16 \
--max-num-seqs 1 \
--max-model-len 65536
Observed behavior
- Generation throughput: 17-19 tok/s
- Prompt throughput: around 300-500 tok/s
- VRAM usage: very high and easy to push above 90%
- KV cache usage: 6-8% for short prompts
- Stability: very high
Evaluation
- In this subjective evaluation, BF16 is the quality and consistency reference.
- Longer TTFT made waiting noticeable in chat.
Best suited for
- Code edits where I want to minimize breakage
- Long-form specifications and strict procedures
- Baseline measurement and fallback use
BF16’s extra wait is noticeable over repeated chat turns. I would retain it as a reference for final edits and review.
2. FP8 (vLLM / –quantization fp8)
I checked whether FP8 improves TTFT.
Example
vllm serve NousResearch/Hermes-4.3-36B \
--dtype bfloat16 \
--quantization fp8 \
--max-num-seqs 1 \
--max-model-len 65536
Observed behavior
- Generation throughput: 18-20 tok/s
- Prompt throughput: over 1000 tok/s, with clearly faster prefill
- VRAM usage: reduced versus BF16
- Prefix cache hit rate: useful depending on the workload
Evaluation
- TTFT improves relative to BF16
- Decode throughput changes little, but waiting feels shorter
- Quality loss is minor at the level I can perceive from these tasks
Best suited for
- Chat with less waiting than BF16.
- Code generation where I want to avoid possible 4-bit quality loss.
- A candidate for daily interaction.
FP8 changed decode speed little but reduced prefill waiting in chat and lighter code tasks.
3. nvfp4 (4-bit, around 22GB)
nvfp4 produced the largest change in generation speed and memory use.
Example
vllm serve NousResearch/Hermes-4.3-36B-nvfp4 \
--max-num-seqs 1 \
--max-model-len 32768
Observed behavior
- Generation throughput: 31-33 tok/s, stable
- Prompt throughput: 280-500 tok/s
- VRAM usage: around 22GB
- KV cache usage: 1-2%, leaving a lot of headroom
Evaluation
- Generation speed is roughly 1.7x to 2x faster.
- Chat waiting is shorter.
- More memory is available for context.
- Subjectively, the output remained usable for chat.
Cautions
- Evaluate response speed and output quality separately.
- Long-context consistency and precision may be lower than BF16.
- Check code modifications with tests.
Best suited for
- Daily chat with MCP + context7.
- Design exploration and other long-context tasks.
nvfp4 uses about 22GB with 1–2% KV cache usage. A 96GB GPU has capacity for two or three models, though concurrent performance needs separate checking.
Fast responses do not establish code-editing reliability. Tests remain necessary.
Cross-Comparison Summary
| Metric | BF16 | FP8 | nvfp4 |
|---|---|---|---|
| Generation (TG) | 17-19 tok/s | 18-20 tok/s | 31-33 tok/s |
| Prefill (PP) | 300-500 tok/s | 1000+ tok/s | 280-500 tok/s |
| TTFT | Slower | Improved | Good |
| VRAM Usage | 90%+ | Medium | ~22GB |
| KV Cache Usage | 6-8% (short text) | Improved | 1-2% (ample headroom) |
| Quality / Stability | Most stable | Good | Some instability in edge cases |
| Interactive feel | △ | ○ | ◎ |
nvfp4 favors chat speed and context capacity. BF16 remains the baseline; FP8 offers faster prefill.
Analysis
Evaluate speed and quality separately
Fast output can create a favorable impression. Long-context consistency and complex reasoning need separate evaluation.
I plan to compare quality through metrics such as first-pass test success rate.
Switching by development phase
I am considering this division of use:
Exploratory development (nvfp4 advantage):
- Generate and test code fragments.
- Chat with MCP + context7.
- Exploration that prioritizes short waits.
Destructive changes (BF16/FP8 advantage):
- Repository-wide refactoring requiring consistency.
- Critical logic changes and final review.
- Work that prioritizes first-pass test success.
Tool-use observations
Hermes-4.3-36B showed limitations in deep reasoning but was relatively stable in tool use (MCP, Function Calling). Argument specification and task chaining worked reliably, making it practical in workflows that combine LLM with external tools like static analysis.
History and context in aider
In aider-style workflows that resend around 6000 tokens of history each time, model quality matters less than context design — including MCP and context7. In those cases, nvfp4’s VRAM savings and context headroom translate directly into operational benefit.
Reducing unnecessary history also affects response time in LLM integration and development assistance.
Operational Conclusion
Chat and exploration:
nvfp4, prioritizing speed and context capacity.Final edits and review:
BF16orFP8, with tests to check the result.Switching plan: retain
nvfp4as the chat default andFP8/BF16for verification.
Using nvfp4 by default and BF16 for final edits reduced waiting in this workload.
Reproduction Steps
1. Download Models
# BF16/FP8
huggingface-cli download NousResearch/Hermes-4.3-36B
# nvfp4
huggingface-cli download NousResearch/Hermes-4.3-36B-nvfp4
2. Launch vLLM Server
See commands in the “Quantization Patterns I Tested” section. --max-num-seqs 1 is for single-user chat. Increase for batch processing.
3. Measure
Extract Avg generation throughput and Avg prompt throughput from vLLM logs. Confirm stable values across multiple requests.
Technical Notes
FP8 Quantization in vLLM
vLLM’s --quantization fp8 converts BF16 models to FP8 at runtime. No pre-quantized model is needed. Requires Blackwell-generation GPU (compute capability 12.0).
nvfp4 VRAM Estimate
nvFP4 uses about 22GB for this 36B model. It fits within 24GB, while 32GB+ leaves more KV capacity. A 96GB GPU has capacity for two or three such models.
Selection by use case
- BF16 does not fit in VRAM: consider nvfp4.
- Chat speed is the priority: nvfp4.
- Code edits also need faster prefill: FP8.
- Final review needs a quality baseline: BF16.
Future Work
Next I plan separate chat, code-generation and review profiles, compared through first-pass test success rate.
