Results of changing response fields

Changing pathfinder’s MCP response changed this 70-test benchmark by more than 15 points.

Removing maybe_exts, the extensions for same-stem candidates, reduced 67/70 to 52–63/70. Restoring it gave 65/70. alike: N caused a loop; always returning exist_similar: bool produced a 10-point range.

For this model, these response choices worked better:

  • Return only when needed: include ambiguity hints only for ambiguous cases.
  • Few candidates: use 2–3 items. The count 273 triggered more exploration.
  • Clear next action: provide extensions to try.
  • Short responses: reduce repeated information.

Background

pathfinder combines ColBERT MaxSim reranking with lexical scores for path resolution. I used qwen3.5-35b-a3b (Q5_K_M) on 70 tests to examine tool responses for small-model LLM integration.

The starting point for this work was a score degradation to 52–63/70 after partially reverting changes from the earlier lexical optimization (lexiopt4) that had achieved a best score of 67/70.


Investigating Score Degradation After Revert

Three benchmark result files (lexiopt4-1, 4-2, 4-3) were compared. Against the pre-revert lexiopt4 (67/70) and lexiopt5 (61/70), the post-revert scores were 63, 52, 57 — with significant variance.

Key degradation points:

CategoryPre-revertPost-revert
G: Retry operations3/30/3 (total failure in Runs 2, 3)
L: Test variant/config4/40–2/4
D: Wrong extensions7/76/7 (Test 31 consistently failing)

The initial suspicion was the code revert itself. segment_overlap, dir_ext_counts, warmup improvements, tool_retry_with_resolve tool_name addition — these were features that should have been lost in the revert.


Nothing Had Been Reverted

All four features were still present.

I compared benchmark commit 9192d6b with HEAD d6c9dc8.


Discovering maybe_exts

The change was removal of maybe_exts.

maybe_exts returns a list of extensions for candidates sharing the same stem (filename minus extension) as the best resolution result. For example, when resolving config.go, if config.rs is also a candidate, it returns maybe_exts: ["go", "rs"].

This field existed in 4 locations:

  1. PathResolveOut struct
  2. Aggregation logic in resolve()
  3. tool_path_resolve response
  4. Tool description

Restoring maybe_exts and Verifying the Effect

I restored all 4 locations, passed the build and tests, then ran the benchmark.

The result was 65/70, back in the baseline range.

Categoryopt4 (67)Post-revert (52-63)Restored (65)
G: Retry3/30-3/33/3
K: Cross-lang4/64/65/6
L: Config0-4/40-2/42/2

Restoring the few tokens in maybe_exts: ["rs","go"] recovered over 10 points. I interpret them as useful extension hints.


Further Response Vocabulary Experiments

I then tested other hints.

Adding alike: N (Same-Stem File Count)

maybe_exts names similar extensions; alike counts same-stem files.

The stem config had 273 files, yielding alike: 273. The model explored excessively, failed on list_dir, and repeatedly checked counts. It recalculated “59 PASS, 1 FAIL” dozens of times and skipped Tests 61–70.

Switching to exist_similar: bool

I replaced the count with exist_similar: true/false on every response.

Three benchmark runs: 66, 60, 56 — variance expanded to 10 points. lexiopt6-1 achieved K=6/6, L=4/4 for the highest level ever, but lexiopt6-3 collapsed to G=0/3. The variance increased compared to the maybe_exts-only configuration (stable at 65/70).

The hypothesis was that false would reinforce certainty, but the results were less stable.

Reverting to maybe_exts Only

I removed exist_similar and alike and returned to maybe_exts.


Interpreting the results

I interpreted the results in three ways.

Repeated information accumulates

The candidates entries contain path, score, why, and root at 50–80 tokens each. Top-5 across 70 tests adds thousands of tokens. Even exist_similar: false appears 70 times. I suspect that accumulation affected attention.

maybe_exts is absent for unique solutions and appears only for ambiguity. Lower context growth may explain its stability.

Primacy-Recency Cascade

G (Tests 46–48) and late L tests declined together. My hypothesis is that attention biased toward the context ends, combined with accumulating fields, made later decisions harder.

Added fields did not help

Adding alike: 273 or exist_similar: false did not help. Conditional, short extension hints in maybe_exts worked best here.


Task-Based Benchmark Expansion

10 tasks covering categories that consistently failed in the 70-test benchmark (F: intent, K: cross-lang, L: config disambiguation) were added, expanding the task-based tests (test_prompt_tasks.md) from 30 to 40.

Added Tasks 31–40:

  • Go vs Rust client disambiguation (K category)
  • Hand-written vs generated disambiguation (K category)
  • Dev vs prod Kubernetes overlay (L category)
  • VPC vs ECS Terraform module (L category)
  • Go config vs Rust config (D category)
  • Deep OAuth callback, Postgres vs PostgREST, vendor exclusion, GraphQL schema vs generated, reranker entry point

The result was 40/40 with 42 tool calls. I think explicit task intent helped pathfinder use intent_text.

A baseline version test (test_prompt_tasks_baseline.md) solving the same tasks using only filesystem MCP tools was also created.


Benchmark Score Progression

ConfigurationScoreVariance
With maybe_exts (lexiopt4)67/70baseline
Without maybe_exts (revert)52-63/7011 points
maybe_exts restored (latest)65/70baseline range
alike: N addedcollapse (loop)—
exist_similar: bool added56-66/7010 points
maybe_exts only (final)65/70stable

Code Changes Summary

  • Restored maybe_exts (PathResolveOut, resolve(), tool_path_resolve, tool description)
  • Added and removed alike: usize
  • Added and removed exist_similar: bool
  • Expanded task-based tests from 30 to 40
  • Created filesystem baseline task tests