Comparing accuracy after changing pathfinder MCP responses
pathfinder extension hints, counts, and booleans compared on 70 tests and 40 tasks, documenting response design for LLM integration and MCP tool development.
Results of changing response fields
Changing pathfinder’s MCP response changed this 70-test benchmark by more than 15 points.
Removing maybe_exts, the extensions for same-stem candidates, reduced 67/70 to 52–63/70. Restoring it gave 65/70. alike: N caused a loop; always returning exist_similar: bool produced a 10-point range.
For this model, these response choices worked better:
- Return only when needed: include ambiguity hints only for ambiguous cases.
- Few candidates: use 2–3 items. The count 273 triggered more exploration.
- Clear next action: provide extensions to try.
- Short responses: reduce repeated information.
Background
pathfinder combines ColBERT MaxSim reranking with lexical scores for path resolution. I used qwen3.5-35b-a3b (Q5_K_M) on 70 tests to examine tool responses for small-model LLM integration.
The starting point for this work was a score degradation to 52–63/70 after partially reverting changes from the earlier lexical optimization (lexiopt4) that had achieved a best score of 67/70.
Investigating Score Degradation After Revert
Three benchmark result files (lexiopt4-1, 4-2, 4-3) were compared. Against the pre-revert lexiopt4 (67/70) and lexiopt5 (61/70), the post-revert scores were 63, 52, 57 — with significant variance.
Key degradation points:
| Category | Pre-revert | Post-revert |
|---|---|---|
| G: Retry operations | 3/3 | 0/3 (total failure in Runs 2, 3) |
| L: Test variant/config | 4/4 | 0–2/4 |
| D: Wrong extensions | 7/7 | 6/7 (Test 31 consistently failing) |
The initial suspicion was the code revert itself. segment_overlap, dir_ext_counts, warmup improvements, tool_retry_with_resolve tool_name addition — these were features that should have been lost in the revert.
Nothing Had Been Reverted
All four features were still present.
I compared benchmark commit 9192d6b with HEAD d6c9dc8.
Discovering maybe_exts
The change was removal of maybe_exts.
maybe_exts returns a list of extensions for candidates sharing the same stem (filename minus extension) as the best resolution result. For example, when resolving config.go, if config.rs is also a candidate, it returns maybe_exts: ["go", "rs"].
This field existed in 4 locations:
PathResolveOutstruct- Aggregation logic in
resolve() tool_path_resolveresponse- Tool description
Restoring maybe_exts and Verifying the Effect
I restored all 4 locations, passed the build and tests, then ran the benchmark.
The result was 65/70, back in the baseline range.
| Category | opt4 (67) | Post-revert (52-63) | Restored (65) |
|---|---|---|---|
| G: Retry | 3/3 | 0-3/3 | 3/3 |
| K: Cross-lang | 4/6 | 4/6 | 5/6 |
| L: Config | 0-4/4 | 0-2/4 | 2/2 |
Restoring the few tokens in maybe_exts: ["rs","go"] recovered over 10 points. I interpret them as useful extension hints.
Further Response Vocabulary Experiments
I then tested other hints.
Adding alike: N (Same-Stem File Count)
maybe_exts names similar extensions; alike counts same-stem files.
The stem config had 273 files, yielding alike: 273. The model explored excessively, failed on list_dir, and repeatedly checked counts. It recalculated “59 PASS, 1 FAIL” dozens of times and skipped Tests 61–70.
Switching to exist_similar: bool
I replaced the count with exist_similar: true/false on every response.
Three benchmark runs: 66, 60, 56 — variance expanded to 10 points. lexiopt6-1 achieved K=6/6, L=4/4 for the highest level ever, but lexiopt6-3 collapsed to G=0/3. The variance increased compared to the maybe_exts-only configuration (stable at 65/70).
The hypothesis was that false would reinforce certainty, but the results were less stable.
Reverting to maybe_exts Only
I removed exist_similar and alike and returned to maybe_exts.
Interpreting the results
I interpreted the results in three ways.
Repeated information accumulates
The candidates entries contain path, score, why, and root at 50–80 tokens each. Top-5 across 70 tests adds thousands of tokens. Even exist_similar: false appears 70 times. I suspect that accumulation affected attention.
maybe_exts is absent for unique solutions and appears only for ambiguity. Lower context growth may explain its stability.
Primacy-Recency Cascade
G (Tests 46–48) and late L tests declined together. My hypothesis is that attention biased toward the context ends, combined with accumulating fields, made later decisions harder.
Added fields did not help
Adding alike: 273 or exist_similar: false did not help. Conditional, short extension hints in maybe_exts worked best here.
Task-Based Benchmark Expansion
10 tasks covering categories that consistently failed in the 70-test benchmark (F: intent, K: cross-lang, L: config disambiguation) were added, expanding the task-based tests (test_prompt_tasks.md) from 30 to 40.
Added Tasks 31–40:
- Go vs Rust client disambiguation (K category)
- Hand-written vs generated disambiguation (K category)
- Dev vs prod Kubernetes overlay (L category)
- VPC vs ECS Terraform module (L category)
- Go config vs Rust config (D category)
- Deep OAuth callback, Postgres vs PostgREST, vendor exclusion, GraphQL schema vs generated, reranker entry point
The result was 40/40 with 42 tool calls. I think explicit task intent helped pathfinder use intent_text.
A baseline version test (test_prompt_tasks_baseline.md) solving the same tasks using only filesystem MCP tools was also created.
Benchmark Score Progression
| Configuration | Score | Variance |
|---|---|---|
| With maybe_exts (lexiopt4) | 67/70 | baseline |
| Without maybe_exts (revert) | 52-63/70 | 11 points |
| maybe_exts restored (latest) | 65/70 | baseline range |
| alike: N added | collapse (loop) | — |
| exist_similar: bool added | 56-66/70 | 10 points |
| maybe_exts only (final) | 65/70 | stable |
Code Changes Summary
- Restored
maybe_exts(PathResolveOut, resolve(), tool_path_resolve, tool description) - Added and removed
alike: usize - Added and removed
exist_similar: bool - Expanded task-based tests from 30 to 40
- Created filesystem baseline task tests
