Integrating voracle research into development
Stabilizing web research saved to a vault: a UTF-8 panic fix, ONNX model migration, and MCP integration. A record of reusing research in development workflows after LLM integration.
Stabilizing the research pipeline
I used voracle research in daily development and stabilized it over about a week. This follows the previous article on ONNX embedding, ColBERT reranking, the MCP server, and distil, connecting retrieval, storage, and search.
Saving research to the vault
The original goal was semantic search with vague queries. Adding research let me save findings from development work as well.
Brave Search API content is converted to Markdown, summarized, and saved to the vault. Information that might otherwise disappear into browser history becomes searchable.
Agent research during daily work accumulates in the Obsidian vault. Without deliberately running a separate exploration CLI, I can search previous findings during the next development task.
I stabilized this flow from late March through early April for use in development after LLM integration.
Excluding drafts and intermediate files
I limited the index to finalized knowledge. Drafts and system files created search noise.
obsidian/
├── articles/ # My articles
│ ├── _drafts/ # Rough drafts
│ └── _processed/ # Archived after publication
├── web-research/ # research command output
│ └── {category}/note_id.md
├── tasks/ # Task management
├── thoughts/ # Interest notes
└── works/ # Work logs
└── development/
├── _claude/
├── _codex/
└── _gemini/
Names starting with _ or . are excluded. This covers _drafts/, _processed/, .search_result.jsonl, and other intermediate artifacts.
# Check vault status
voracle -status
vaults: 1
obsidian: 847 notes, 2,341 chunks
excluded: _drafts, _processed, _claude, _codex, _gemini
index: .voracle.usearch (f16 HNSW, 2341 vectors)
store: .voracle.db (SQLite)
The searchable content is finished articles, saved research, and my notes.
Testing research with TurboQuant
On the day I reorganized the vault, I used research to investigate Google’s TurboQuant and local LLM quantization.
voracle research "TurboQuant|local LLM"
The processing steps are:
- Extract keywords from the query
- Search via Brave Search API in both Japanese and English (Web + News)
- Deduplicate URLs into a list
- Fetch each URL and convert to Markdown with
html-to-markdown-rs - Score by content quality and novelty
- Save the top N results as Markdown files in
web-research/ - Auto-generate summaries and frontmatter (topic, category, difficulty)
The result contained 61,813 characters, about 26,540 tokens. It exceeded Claude Code’s tool-result limit, so I read it through a file. The research completed.
The pipeline searches “TurboQuant quantization” in English and the equivalent Japanese terms in parallel. It chooses counter-language keywords through maxsim matching against a topic configuration file to supplement results from either language.
Note: In this research, Japanese results yielded little useful content, while English searches found material without translations. I adjusted scoring to retrieve most articles from English sources.
A panic from slicing inside a UTF-8 character
I ran a broader research query the same day.
voracle research "TurboQuant|ColBERT LLM|Modern BERT"
It produced this panic:
thread 'main' panicked at src/infra/inference/lfm.rs:237:49:
byte index 2000 is not a char boundary; it is inside '━' (bytes 1998..2001)
The pre-summary truncation code written by Claude sliced at a byte position without checking character boundaries:
// BAD: byte position 2000 might land in the middle of a character
let truncated = if text.len() > 2000 { &text[..2000] } else { text };
Rust’s text.len() returns bytes. ━ (U+2501) spans 3 UTF-8 bytes (1998..2001), so position 2000 was inside the character. Retrieved web content includes these symbols and Japanese text.
I used floor_char_boundary to return to a valid boundary:
let truncated = if text.len() > 2000 {
let boundary = text.floor_char_boundary(2000);
&text[..boundary]
} else {
text
};
floor_char_boundary returns the nearest boundary at or below a position. It moves 2000 back to 1998 in this example. This API was stabilized in Rust 1.73.
I also added a test:
#[test]
fn truncate_multibyte_boundary() {
// '━' (U+2501) is 3 bytes in UTF-8 (0xE2 0x94 0x81)
let text = "a".repeat(1998) + "━━━";
assert_eq!(text.len(), 2007); // 1998 + 9
let boundary = text.floor_char_boundary(2000);
assert_eq!(boundary, 1998);
let truncated = &text[..boundary];
assert!(truncated.is_char_boundary(truncated.len()));
}
English-only test inputs had not exposed this case. The tests needed Japanese text and symbols.
Moving summary generation to Qwen3.5-0.8B
Summarization, keyword extraction, and frontmatter generation used LiquidAI LFM2.5-350M-ONNX. I switched to onnx-community/Qwen3.5-0.8B-ONNX because of licensing constraints.
Selection reasons:
- Apache 2.0 license, no issues with commercial use
- 0.8B parameters keeps inference load practical
- Available in ONNX format (INT8/FP32) for direct loading
I retained infra/inference/ and replaced the LFM2.5 components with Qwen3.5-0.8B:
src/infra/inference/
embed.rs # Dense vector embeddings (pplx-embed-v1-0.6b -- unchanged)
colbert.rs # ColBERT MaxSim reranking (unchanged)
lfm.rs # LFM2.5 -> replaced with Qwen3.5-0.8B
ONNX Runtime is pinned to ort = "=2.0.0-rc.12", as in multi-bert-inference and edge-retrieval. The ort crate can change incompatibly between RCs, so these projects use a verified version.
# Cargo.toml
[dependencies]
ort = "=2.0.0-rc.12"
tokenizers = "0.21"
tokenizers loads tokenizer.json. Qwen3 uses its own format rather than SentencePiece, supported by HuggingFace’s tokenizers crate.
# Model layout
ls ~/.local/share/voracle/models/
qwen3.5-0.8b-onnx/
model.onnx
model_quantized.onnx
tokenizer.json
config.json
pplx-embed-v1-0.6b/
model.onnx
tokenizer.json
mxbai-edge-colbert/
model.onnx
tokenizer.json
The summaries felt at least comparable to LFM2.5, especially in Japanese. This was an impression, not a quantitative comparison.
Import logic and tests
I reviewed the import logic and added missing tests.
Full Pipeline Flow
The full research flow is:
Query input
-> Keyword extraction (Qwen3.5-0.8B)
-> Brave Search (primary language + counter language)
-> URL dedup + domain strike exclusion, fetch HTML
-> tree-sitter-html content extraction + boilerplate removal
-> html-to-markdown-rs Markdown conversion
-> Scoring -> select top N
-> vault writer saves to web-research/
-> Summary + frontmatter generation (Qwen3.5-0.8B)
Domain strikes exclude domains that repeatedly return paywalls or empty content. After 3 failures, Brave Search queries add -site: for that domain. The rule operates at domain level.
# Check strike status
voracle admin strikes
# Restore a falsely flagged domain
voracle admin allow example.com
Adding Tests
The added tests cover:
writer.rs: collision handling when writing to vault with existing filesresolver.rs: edge cases for_/.prefix exclusion rulesextraction/: handling of empty content and error responses
#[test]
fn writer_avoids_overwriting_existing_note() {
let dir = tempdir().unwrap();
let writer = VaultWriter::new(dir.path());
// Write the same URL twice
writer.write_research_note(&entry_a).unwrap();
writer.write_research_note(&entry_a_updated).unwrap();
// Second write creates a separate file (no overwrite)
let files: Vec<_> = fs::read_dir(dir.path()).unwrap().collect();
assert_eq!(files.len(), 2);
}
#[test]
fn resolver_excludes_underscore_prefix() {
let dir = tempdir().unwrap();
fs::create_dir_all(dir.path().join("_drafts")).unwrap();
fs::write(dir.path().join("_drafts/note.md"), "# test").unwrap();
fs::write(dir.path().join("visible.md"), "# test").unwrap();
let resolver = VaultResolver::new(dir.path());
let notes = resolver.scan();
assert_eq!(notes.len(), 1);
assert_eq!(notes[0].file_name(), "visible.md");
}
Calling voracle through agent-gateway
After stabilizing the vault and research pipeline, I added voracle to agent-gateway’s MCP presets.
agent-gateway manages MCP server settings for LLM agents through preset definitions:
// internal/agent/mcp/presets.go
func PresetVoracle(command string) ServerConfig {
if command == "" {
command = "voracle"
}
return ServerConfig{
Name: "voracle",
Command: command,
Args: []string{"--server"},
}
}
func PresetOkitegami(command string) ServerConfig {
if command == "" {
command = "okitegami"
}
return ServerConfig{
Name: "okitegami",
Command: command,
Args: []string{"--server"},
}
}
Agents can search the vault, run research when needed, and save results. This connects research and reuse within the same development flow.
Daily use of grep and research
The coding took about half a day. I implemented small refinements gradually, including ideas from daily walks.
The command set at the time of this record:

I mostly use grep and research. grep finds semantically similar chunks rather than exact matches:
ksh3@desktop.home.arpa ~ % voracle grep "turbo quant"
/Users/ksh3/Development/obsidian/web-research/2026/04/09/technology/2029b5389eccfe38.md:16 [How to test TurboQuant] (4.2379)
Verify correctness and measure real-world impact before deploying TurboQuant to production.
/Users/ksh3/Development/obsidian/web-research/2026/04/09/technology/2029b5389eccfe38.md:5 [Triton + vLLM (serving workloads)] (4.2284)
Advanced
Engineers building inference serving pipelines who want to test TurboQuant in a vLLM-like environment
/Users/ksh3/Development/obsidian/web-research/2026/04/09/technology/2029b5389eccfe38.md:2 [How to use TurboQuant] (4.2177)
A practical guide for developers who want to try TurboQuant KV cache compression. Covers available implementations, s...
Saved research and my articles share the same searchable index:
ksh3@desktop.home.arpa ~ % voracle grep "GLM-5.1"
/Users/ksh3/Development/obsidian/articles/llm/glm-4-7-flash-cpu-hybrid-gpu-benchmark_en.md:19 [NVFP4 + vLLM Operational Evaluation] (6.9704)
In addition to the IQ5_K benchmarks above, GLM-4.7-Flash-NVFP4 was also evaluated on vLLM for operational suitability.
/Users/ksh3/Development/obsidian/articles/llm/glm-4-7-flash-cpu-hybrid-gpu-benchmark_ja.md:33 [共通変数] (6.8849)
IMG=compute.home.arpa/ik_llama-cpu:latest
MO=/mnt/data/hf/hub/models--ubergarm--GLM-4.7-Flash-GGUF
/Users/ksh3/Development/obsidian/web-research/2026/04/08/technology/5ee34b17c638fe00.md:40 [GLM-5 - GlmMoeDsa] (6.7509)
The zAI team launches GLM-5, and introduces it as such:
> GLM-5, targeting complex systems engineering and long-horizon agentic tasks.
Domains whose paywalls prevent retrieval are excluded automatically:
ksh3@desktop.home.arpa ~ % voracle admin blocklist
Exclude Domains (strike >= 3 = excluded):
[X] github.com 40 (paywall)
[X] www.ponkotsu.dev 11 (paywall)
[X] medium.com 7 (paywall)
[X] dev.to 4 (paywall)
[X] machinelearningmastery.com 3 (paywall)
...
research also works as an MCP tool, called during development by Claude Code and agent-gateway agents. My research and agent research share the vault, with information retrieved as needed. I now start most lookups with voracle grep.
I sometimes implement changes in about 30 minutes after a walk, such as reducing noise or ranking a useful search pattern higher. I keep adjusting the tool to daily use.
