Fast input processing, slow generation

IQuest-Coder-V1-40B-Loop-Instruct processed inputs quickly in aider, but generated at 0.6-8 tok/s. That made the wait for each response long during repeated edits.

A different setup from the same period reached 25-28 tok/s with IQuest-Coder-V1-40B-Instruct-nvfp4. The results differ even within the 40B class, so output format and settings matter too.

Edit format and waiting time

In aider, each edit leads to another review and instruction. A long wait for each response interrupts that cycle.

The note described the result this way:

Prefill is fast enough, but decode is critically slow.

vLLM speed and cache usage

  • Prefill (input processing) ran at hundreds to 900 tok/s.
  • Decode (generation) ran at 0.6-8 tok/s.
  • GPU or KV cache shortage was not considered the main cause.
  • Generation took too long for repeated edits in aider.

The vLLM logs also show low generation throughput.

  (APIServer pid=1) INFO 01-09 19:15:06 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 8.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.2%, Prefix cache hit rate: 6.3%
(APIServer pid=1) INFO 01-09 19:15:16 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 8.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.5%, Prefix cache hit rate: 6.3%
(APIServer pid=1) INFO 01-09 19:16:16 [loggers.py:257] Engine 000: Avg prompt throughput: 251.4 tokens/s, Avg generation throughput: 4.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 7.0%, Prefix cache hit rate: 6.4%
(APIServer pid=1) INFO 01-09 19:17:06 [loggers.py:257] Engine 000: Avg prompt throughput: 73.4 tokens/s, Avg generation throughput: 1.1 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 6.8%
  

KV cache usage was 7.0%-13.8%, well below full capacity. Prefix cache hit rate was 6.3%-6.8%. Those values alone do not explain the speed difference, so the generation phase needs attention.

Conditions that may contribute

The note listed four possible causes.

1. A 40B model returning whole edit

whole edit returns the full file, increasing the output token count even for a small change. At low generation speed, each extra token adds to the wait.

This is part of the Django Comment model returned during the evaluation.

  from django.db import models
from django.utils.translation import gettext_lazy as _
from core.models import TimeStampedModel, SoftDeleteModel


class Comment(TimeStampedModel, SoftDeleteModel):
    class Status(models.TextChoices):
        PENDING = 'pending', _('Pending')
        APPROVED = 'approved', _('Approved')
        SPAM = 'spam', _('Spam')
        TRASH = 'trash', _('Trash')

    site = models.ForeignKey('core.Site', on_delete=models.CASCADE)
    post = models.ForeignKey('content.Post', on_delete=models.CASCADE)
    parent = models.ForeignKey('self', on_delete=models.CASCADE, null=True, blank=True)
    author_user = models.ForeignKey('auth.User', on_delete=models.SET_NULL, null=True, blank=True)
    author_name = models.CharField(max_length=255)
    author_email = models.EmailField()
    author_url = models.URLField(blank=True)
    body = models.TextField()
    status = models.CharField(max_length=20, choices=Status.choices, default=Status.PENDING)
    ip_hash = models.CharField(max_length=64)
    user_agent = models.TextField(blank=True)

    class Meta:
        indexes = [
            models.Index(fields=['post', 'status', 'created_at']),
        ]

    def __str__(self):
        return f"Comment by {self.author_name} on {self.post}"

    def approve(self):
        self.status = self.Status.APPROVED
        self.save(update_fields=['status'])

    def mark_as_spam(self):
        self.status = self.Status.SPAM
        self.save(update_fields=['status'])

    def move_to_trash(self):
        self.status = self.Status.TRASH
        self.save(update_fields=['status'])

    def restore(self):
        self.status = self.Status.PENDING
        self.save(update_fields=['status'])

    def is_approved(self):
        return self.status == self.Status.APPROVED
  

The model produced substantial code. At single-digit tok/s, though, this is a long response to wait for on every edit.

2. Large context

repo-map and multiple files add information for the model to consider. The note proposed reducing repo-map size and the files included through /add. More context can also encourage longer explanations and edits.

3. Sampling

One proposed change was temperature=0 with greedy decoding, which selects the most probable token at each step. The aim was repeatable answers and earlier termination during code edits.

4. Single-request generation with a 40B model

A 40B-class model may be a poor fit for repeated single-request generation. However, the different nvfp4 setup recorded:

  • Prompt throughput: 1100-2300 tok/s
  • Generation throughput: 25-28 tok/s
  • KV cache usage: 2-12%
  • Prefix cache hit rate: 20-45%

Model size alone does not explain the result. Output format, context size, and generation settings also need to be considered.

Proposed changes

The note proposed these changes in order:

  1. Replace whole edit with diff/patch
  2. Use temperature=0
  3. Reduce repo-map size and /add scope
  4. Keep max_model_len to the minimum required
  5. If needed, move to a quantized or lighter coding model in the 20B-32B range

The first two changes aim to shorten output and simplify token selection. They focus on generation time because input processing was already fast enough.

What the comparison shows

The related record, Firworks-IQuest-Coder-V1-40B-Instruct-nvfp4.md, used different conditions. At 25-28 tok/s, 200 tokens take a few seconds; at 0.6-8 tok/s, the wait is much longer.

In this test, 40B + whole edit combined low generation throughput (t/s) with long output. Input-processing speed alone does not determine the wait for an edit.